Ridiculously fast on OpenRouter, just subjectively it's a really strange experience because I've never seen a model respond or execute that quickly.
Ridiculously fast on OpenRouter, just subjectively it's a really strange experience because I've never seen a model respond or execute that quickly.
Try https://chatjimmy.ai/ from Taalas. There is an emergent space for super-fast-models esp finetuned or guardrailed to solve very specific latency sensitive tasks.
Generation is insanely fast, the other side is presentation, which can be slower and more controlled. Blasting the end-user with text walls is a UX problem now.
Try Cerebras. When I think about how speed of generation is another variable to tweak for "intelligence", it seems like this speed is best used for searching for solutions in a problem space and then validating and discarding and keeping what is best. Being intelligent at the Fable level, but what if the Fable level machine could think at 100x? What does that mean: perhaps it means more parallel "experiments" for solutions in the token/generation/hyper-dimensions of the latent space.