so many inference project, omlx already supports all of this and has a 1000 people trying to optimize it constantly

Both projects are different in scope. Think of slotstream as optimizing for memory and for this specific model for now, my intention is not to build an inference engine the same as oMLX

Interesting! I'll check it out