I had the same reaction but then I showed it to my partner. She completely didn't get it, in her words "how can it be thinking of a good answer when it's that quick?"
I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.
to be fair, the model used for Chat Jimmy is not very smart, but the world where it is smart is very interesting.
It’s going to be really crazy when the bottle neck for agents is the speed of the tool calls rather than the speed of inference. Imagine an agent interacting with the terminal near instantly…
I asked it some old hardware command line questions I'd recently asked Gemini, it hallucinated parts of the answer.
The characters in the 3-act Shakespearean play had very little depth, many of the names were similar, and they were not very smart, but the simple plot was cohesive.
It’s not reasoning, the hardware demo uses a 3.-something generation Llama 8B.
But it’s proven they can automate this (they didn’t etch eight billion weights by hand after all, obviously), so now the interesting question is whether they can scale it to more recent aka bigger models.
After all, there’s already very useful models even for productivity at 27 or 35B.
It's not really "instant", i.e. the text is still generated token-by-token, it's just super fast. Reasoning would work with this model without any changes to the chip but it's disabled for speed.
I had the same reaction but then I showed it to my partner. She completely didn't get it, in her words "how can it be thinking of a good answer when it's that quick?"
I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.
It's not thinking. Not in the way she probably meant. It can "think" that fast the same way a calculator can "think" that fast (kind of).
Because it's not human and not "thinking", it's a mathematical algorithm
to be fair, the model used for Chat Jimmy is not very smart, but the world where it is smart is very interesting.
It’s going to be really crazy when the bottle neck for agents is the speed of the tool calls rather than the speed of inference. Imagine an agent interacting with the terminal near instantly…
I had the chance to try out MiMo v2.5 Pro Ultraspeed (600-1000tok/s) for a couple weeks and it is amazing.
Developing software becomes 95% about intent and requirements. Can’t wait for the next iteration of that.
I asked it some old hardware command line questions I'd recently asked Gemini, it hallucinated parts of the answer.
The characters in the 3-act Shakespearean play had very little depth, many of the names were similar, and they were not very smart, but the simple plot was cohesive.
Wait, is it even thinking? Or is it an instant model?
It’s not reasoning, the hardware demo uses a 3.-something generation Llama 8B.
But it’s proven they can automate this (they didn’t etch eight billion weights by hand after all, obviously), so now the interesting question is whether they can scale it to more recent aka bigger models.
After all, there’s already very useful models even for productivity at 27 or 35B.
My concern is that reasoning could involve some sequential steps that instant models don't.
Not sure if modern models "think" only by outputting <thinking> blocks, or there is a more complex mechanism at play.
It's not really "instant", i.e. the text is still generated token-by-token, it's just super fast. Reasoning would work with this model without any changes to the chip but it's disabled for speed.
I feel like Ray Kroc in the McDonald's movie trying to figure out how his hamburger could possibly be done when he just ordered it
For those old enough to remember, this is like dial up internet to broadband. So fast it creates new markets