> A model with a 10M context window that retains it's intelligence up to 2M tokens would be a big breakthrough.
so the true next frontier might not be just raw intelligence but rather larger context window?
> A model with a 10M context window that retains it's intelligence up to 2M tokens would be a big breakthrough.
so the true next frontier might not be just raw intelligence but rather larger context window?
Or better context curation - less lossy compression saving back to context. Maybe even jettisoning part context into an external semantic store instead of conpression. Or placing less data into context to start with.
Or a combination of all those things.
[dead]