I've posted a few times about my project that's a collection of 30k-250k webapps that are served from a WebDAV server. The apps know how to write updated copies of themselves back to the server.
My family uses it. I have gallery apps (yearbooks for each year are a lot of fun!) of us on trips and just living, an outlining app that's a mesh of Workflowy and Org Mode (it's called Fluxtral), a markdown-backed app (it uses marked.min.js, and is called Dextral) that offers documents, logs, calendars, and kanban boards, all parsed from markdown. I have a List app for gear, trips, shopping, etc. that we all can contribute to. There are utilities (world clock, calendar) and games (an oracle for RPGs, a KenKen implementation), and apps (a diagram editor that exports to SVG, a web-launcher that uses pneumonics, a Scheme-based hacking environment, and a spreadsheet that does most of what you'd expect aside from Solver and Pivot tables).
I started these projects before AI, and made slow progress over the years, but the modern versions of all this stuff have been built with Deepseek V4 Flash. I've also used Gemini in the very early days, and Kimi K2.6 later on, but these days, since I can now host Deepseek v4 Flash 0731 in a 2-bit quant on my Strix Halo box (128GB, but only about 250GB/s of memory bandwidth, so 15t/s), I used Deepseek with omp for almost everything. It's a very capable model, and I'm amazed I can run it locally and get good results. It's really revolutionary for my (small) use cases.
I too have Strix Halo. Two of them. The model have been awesome for me too.
With unsloth's Q3_S quant + kyuz0 'llama-vulkan-radv-performance' toolbox, I am getting 280+tps (batch and ubatch at 2048) for PP and 18+tps for TG. I really only need 256k context so it all fits.
If I go down to the Q3_XXS quant + dpsark + 'llama-vulkan-radv-performance', I can get about the same PP and 25+tps for TG with draft set to 2 or 3. Fits about the same as above.
Edit: I did notice the 25+tps quickly degrades down to 20+ after the first few hundred tokens.
Makes sense, I'll look out for it! although of course most Show HNs these days get lost in a deluge of submissions.. you might actually be better off omitting the Show HN when you submit it..
I've posted a few times about my project that's a collection of 30k-250k webapps that are served from a WebDAV server. The apps know how to write updated copies of themselves back to the server.
My family uses it. I have gallery apps (yearbooks for each year are a lot of fun!) of us on trips and just living, an outlining app that's a mesh of Workflowy and Org Mode (it's called Fluxtral), a markdown-backed app (it uses marked.min.js, and is called Dextral) that offers documents, logs, calendars, and kanban boards, all parsed from markdown. I have a List app for gear, trips, shopping, etc. that we all can contribute to. There are utilities (world clock, calendar) and games (an oracle for RPGs, a KenKen implementation), and apps (a diagram editor that exports to SVG, a web-launcher that uses pneumonics, a Scheme-based hacking environment, and a spreadsheet that does most of what you'd expect aside from Solver and Pivot tables).
I started these projects before AI, and made slow progress over the years, but the modern versions of all this stuff have been built with Deepseek V4 Flash. I've also used Gemini in the very early days, and Kimi K2.6 later on, but these days, since I can now host Deepseek v4 Flash 0731 in a 2-bit quant on my Strix Halo box (128GB, but only about 250GB/s of memory bandwidth, so 15t/s), I used Deepseek with omp for almost everything. It's a very capable model, and I'm amazed I can run it locally and get good results. It's really revolutionary for my (small) use cases.
I too have Strix Halo. Two of them. The model have been awesome for me too.
With unsloth's Q3_S quant + kyuz0 'llama-vulkan-radv-performance' toolbox, I am getting 280+tps (batch and ubatch at 2048) for PP and 18+tps for TG. I really only need 256k context so it all fits.
If I go down to the Q3_XXS quant + dpsark + 'llama-vulkan-radv-performance', I can get about the same PP and 25+tps for TG with draft set to 2 or 3. Fits about the same as above.
Edit: I did notice the 25+tps quickly degrades down to 20+ after the first few hundred tokens.
Sounds fascinating! A blog write-up about your platform would be a fun read, if you're up to it
For sure! I'll be doing a Show HN at some point, just want to feel a bit more confident about certain aspects first.
Makes sense, I'll look out for it! although of course most Show HNs these days get lost in a deluge of submissions.. you might actually be better off omitting the Show HN when you submit it..
A collection of 30k-250k apps? Like individual unique apps?
Sorry, a collection of apps whose size is between 30kb and 250kb.
Haha, I misinterpreted as well. That would be a lot of apps!
[flagged]
> which in your case is?
oh, they're mad.