GPT Astra did some benchmarking on the DGX Spark. Speed: 34.38 tokens/sec for generation.
Seems like we don't have a drafter model yet so it could not test with speculative decoding on. ngram speculative decoding did not help too much either - not enough accepted tokens.
Smaller size I suppose does not mean better performance in this case - we maybe limited by Spark's low memory bandwidth.
That's a rough place to land on a spark. It seems unlikely to be memory bandwidth at this model size, but maybe just lack of tuned kernels? The chip is missing some CUDA features but with tuning you should be able to hit way more than that even without a drafter.
I was wondering if those ternary bits get expanded into full floats internally in the kernels - you’re probably right about lack of tuned kernels. I’m not sure you could just tune your way out of that easily though. Any suggestions on trying particular solutions?
What are you getting for prefill?
450 with PTQ_01 and 900 with the other PQ2_0.