128k context is not a limit of the model, that's a limit of implementation:
"Context Length: 262,144 natively and extensible up to 1,000,000 tokens."
128k context is not a limit of the model, that's a limit of implementation:
"Context Length: 262,144 natively and extensible up to 1,000,000 tokens."
We're talking about the Cerebras implementation, which is limited to 128K.
It's in the link.
TPM means Tokens per Minute.
GP is referring to GGP’s last paragraph. 150k t/m, yes, and 128k context.