Try Cerebras. When I think about how speed of generation is another variable to tweak for "intelligence", it seems like this speed is best used for searching for solutions in a problem space and then validating and discarding and keeping what is best. Being intelligent at the Fable level, but what if the Fable level machine could think at 100x? What does that mean: perhaps it means more parallel "experiments" for solutions in the token/generation/hyper-dimensions of the latent space.