I absolutely love this comment. I wished there was a website where people would post their working command lines as well as what hardware they are using to run that stuff on + tokens / sec prefill + gen.

As a MBP and DGX spark owner, I would love such a site... (Feel like it would be a low effort feature of hugging face).

Searching through Reddit and forums for best commands is annoying.

Problem with that is I think that it quickly devolves into cargo culting, nonsense and noise.

Arguably, what I am doing is also very very close to that, with the only difference being that I am somewhat less of an idiot than the average internet dweller you'd get on such a site. Or rather a different flavor of idiot.

Ideally, the people building the tools build them in a way that just does the right thing - which I am confident that llama.cpp does or will do in the future.

So you encode that knowledge not in language and online comments but in code and with a filter for actual expertise.

And, frankly, there's really not all that much to it. It's like maybe 3 parameters to play around with.

The valid solution space is pretty small, but people will want to make it "theirs" regardless, so you get non-solutions just so that everyone could also be a part of it. The usual social dynamics foo.

I've found that once you factor in multiple GPUs things can get complex quite quickly because the default packing routine in the LLM runners tends to be very coarse resulting in substantial amounts of VRAM wasted. More so if you start running drafters and multiple models at the same time.

What stackoverflow should have become.

pretty sure this exists already...

[dead]