0.01 tk/s on an M1 Max is not "nearly". This is completely unusable, and in no way cost effective.
0.01 tokens per second means 1 million tokens ($3 worth of API usage [1]) takes 3.2 YEARS.
0.01 tk/s on an M1 Max is not "nearly". This is completely unusable, and in no way cost effective.
0.01 tokens per second means 1 million tokens ($3 worth of API usage [1]) takes 3.2 YEARS.