does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?
does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?
Yes that is exactly what this does.
Could you explain what happens when you try to shoehorn a 2.4T parameter model into a 24gb m4 mac?
Wondering the same thing but for 48gb M5 Max.