Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?
Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?
They make a giant inference chip. Their inference service is basically just advertising for their core value prop: hardware.
The CEO was on Gradient Dissent a couple years ago: https://www.youtube.com/watch?v=qNXebAQ6igs
The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).
They can host larger models by pipelining it on multiple wafers. Each wafer stores one layer and N layers can serve an N * 44 gb model with N concurrency. The limitation would of course be inter-wafer I/O, which my comment was getting at. That's probably how they can serve bigger models like GPT 5.6 Sol [1].
[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...
I never said offloading was impossible. It will result in a large slowdown.
It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.
[dead]
Mostly economics I'm sure