Qwen3.6-35B is my daily driver for AI, and what convinced me to cancel my Claude subscription back in April. The Qwen3.6 line is easily the best local model I've tried, and I've tried a lot. I've got it diligently grinding away on my laptop right now, reviewing and fixing some bugs in my F# code.

Qwen-3.6-35B-A3B was our "gateway drug" into switching our organisation to agent/harness-first coding.

Particularly, I had one team member who was extremely sceptical of AIs/LLMs/harnesses and refused to use them. One day he said "Well, I have an RTX 5090 doing nothing... should I try to get something up on it?" and a few minutes later he had 3.6-35B loaded up, running OpenCode.

It continues to be a workhorse to this day, running on both my local Mac for various types of jobs, an AMD R9700 at the office, and said teammember still uses it on his 5090, although in practical terms we do a lot more with DS-V4-Flash-0731 these days.

I’ve run 3.6-27B and 3.6-35B on 32GB locally for a lot of bulk non-code tasks. Let it run overnight and wake up to millions of output tokens worth of results without data having left my house, all for the price of electricity.

I haven’t found it very useful for code. It can do some code, but I’ve tried a dozen different quants and context lengths and the output is always bad enough that it has to be discarded for anything other than really easy tasks. It has been useful for exploring codebases for search and summary, though.

DS Flash is where local models begin to feel useful for coding, but the quants we run locally are sharply reduced in intelligence from the benchmarks for the full models.

For applications where data cannot leave the local network it’s good to have them. For actual coding work I can’t actually justify the power of electricity and cooling, let alone the expensive hardware, compared to hosted APIs.

But I admit I do enjoy playing with them anyway. I think it’s one of those hobbies where it’s most fun if you never do the math on how much you’re paying for the privilege. If someone has a requirement that data stay local then it’s different, of course.

> I’ve run 3.6-27B and 3.6-35B on 32GB locally for a lot of bulk non-code tasks.

Do you mind sharing your use cases?

Not OP, but I use it for a ton of smaller things. I have it hooked into Hermes and have been using it to help bulk rename my media folders so they all follow a common format, add titles that sort of thing which wouldn't be easy to 'script'. Another thing I use it for is comparing data sets, looking at my exported Spotify artists and compare to what I have locally, and letting me know where there are missing artists, or albums, and recommendations based on similar artists that I may not have locally.

Sure a lot of this could be done without AI, but it's certainly quicker and easier, and since my AI box is on solar, it's just the power of the sun to keep it going.

Not OP, but driving knowledge bases is the poster child use case for me https://github.com/aka-rider/llm-wiki

I started with Karpathy's LLM wiki, and did everything he said not to do - downgraded the model to mere tool usage and summarization, and it works great.

I am a data hoarder, and finally I can just dump all the content I remotely like, and get something interesting to browse for the price of electricity.

Agentic long-running tasks, as others have mentioned:

- Groom and triage tickets for agentic SWE workflows

- bug hunt — the probability of Qwen fixing a complex bug is 50/50 but often it is capable of identifying the root cause or at least laying the ground work for a more capable model to pick it up.

Love the workflow. Is there an existing way to do this inside Obsidian -- where I have all my knowledge?

SmartConnections but heads up, Connections Pro asks $300/year, or, more for the plugin than for an LLM subscription, more than Microsoft Office for that matter.

https://smartconnections.app

Keep in mind Obsidian is open standard JS plugins... You know what can write open standard JS plugins?

Classic which comes first, LLM or the plugin, though. :-)

Not OP, but I use it for personal tasks that are just not worth the claude tokens -- rooting through historical medical records to unify prescription history, super-OCR'ing thousands of PDF pages (i.e. beyond PDF dumping -- vision means it can look at tables, understand tricky things like a continuation of a block quote or aside on the next page, etc.), and when power is cheap I'll just let it noodle on little projects on my data. During my agent's "free time" I give it with a tremendously open ended prompt last week, it did a linguistic analysis of how my texting changes in the lead up to, initiation, maintenance, and ending of romantic relationships.

What kind of hardware do you run that massive beast on (DS-V4-Flash-0731)?

That is exactly what got me past just enough of my cynicism to get started. I am still cynical but now I have meaningful knowledge.

> Qwen-3.6-35B-A3B

The A3B models are super fast but I found the A3B Q4 model ran in circles a lot and ended up taking longer to complete tasks that 27B Q6 because it kept having to redo/rethink/fix something.

I was writing extensive prompts to rein it in and it would still ignore basic directives like "never force push on the repo, ask me instead". I ended up switching back to 27B after about a week of frustration and lost productivity.

We’re using 6bit quants since we have 32GB cards.

Gemma QAT is an honourable mention.

What!? You are skeptical of AI but will go through the manual process of hosting a model that’s less than frontier intelligence (talking about Qwen 3.6)? Anti-AI folks are always odd to me

A local model needs 0 investment and 0 commitment, takes literal minutes to get started (especially if you have someone who is into that stuff showing you the ropes) and if you end up disliking the experience of using AI you can just `rm -fr` it and forget the whole thing existed.

Needs 0 investment and 0 committment?

- You at least need a capable machine, so that's not 0 monetary investment. - You need to spend at least an hour decicding between ollama, llamacp, mlx, etc. - You need to find the correct quantized version of the model that works for you based on the architecture. - You need to figure out the correct context window size to get reasonable performance. - You need to setup a harness that works against your model - You might need to setup additional websearch tools, image tools, etc since harnesses like pi don't come with the model. Ofc you can't use codex and claude code, because those aren't opensource and you are anti-AI.

Or, you could sign up for Opencode for $10 and just be productive.

I'm particularly calling out the hypocrisy of the original comment. Being Anti-AI, and then spending hours on setting up a less than frontier AI model.

> Or, you could sign up for Opencode for $10 and just be productive.

You forgot the step before where you spend months waiting for security to vet it, legal to sign off and procurement to approve it.

Or you could use hardware your team has lying around. Everyone isn’t working on cloud-hosted CRUD APIs.

You might be anti-ai in the sense you aren’t comfortable with all your data being shipped back and forth to a third party.

Im anti AI in the sense of VC backed global warming far right accelerationists.

Local AI is almost perfect. But its like all democracy: its history is marred with lots of crap.

All of your objections have already been addressed by the previous comments.

The original comment states that the person in question already had a suitable graphics card to hand, so it did not require a monetary investment.

GP clearly states that "someone who is into that that stuff" was guiding the process, so it did not require a significant time investment.

> I'm particularly calling out the hypocrisy of the original comment. Being Anti-AI, and then spending hours on setting up a less than frontier AI model.

I see no hypocrisy in the original comment.

You've also assumed the skeptic in question doubts the capabilities of AI. That may be the case (like you, I have no idea), but they may also have privacy concerns, in which case a local model is the appropriate choice.

There are plenty of reasons to be skeptical of AI.

install LM Studio, download the automatically selected quant based on your hardware, start a conversation with the automatic context size. 10 minutes at best and zero effort

Maybe not '0 investment and 0 commitment', but incredibly little depending on what you have laying around. It takes less than 5 minutes to download say LM Studio and an Open Model and as long as you have the hardware to support it, you start moving along. If you are on AMD in some ways it's even 'easier', you can download Lemonade and it will tell you exactly what will fit and best options based on what you are trying to do.

For me at least the local AI stuff, powered with solar has been pretty great. Would that scale to a large business? Goodness no, but for my tinkering and learning, it works great.

> You at least need a capable machine, so that's not 0 monetary investment

It is 0 monetary investment if I already have said machine lying around doing nothing.

Which is exactly the story OP talked about.

But most people don't have an RTX 5090 lying around, so the story doesn't apply to them, right?

You don't need a 5090 to run local AI. A whole lot of people out there are doing it with Macs. Unified ram is the biggest thing.

Correct. If the premise doesn't hold, it's ex falso quodlibet for anyone.

Back in “the day” nerds just bought the hardware to fuck with. Some of us still do. Claiming that compute is the barrier to entry just means you’re not a nerd. That’s ok.

[deleted]

I'm running Qwen 27B no problem with an AMD 9070XT + 24gb DDR5 ram. Does basic web search for me (tool call with tavily, costs nothing I get 1000 searches a month) and is great for creative writing (primarily breaking writer's block). Until the recent surge in ram costs, that wouldn't be hard to do. I built the computer for ~$1600 a year ago.

I am getting ~13-15 tps with my 9070XT for the 27B (~35tps for the 35B-A3B), but I think for me the main bottleneck is the 64gb of DDR4 3600 memory. What kinda speeds are you getting with what speed of DDR5?

> You need to setup a harness that works against your model - You might need to setup additional websearch tools, image tools, etc since harnesses like pi don't come with the model

Pi has a nice guide on it (https://pi.dev/docs/latest/llama-cpp) and it is really not that hard.

How is that hypocrisy? Self hosting is somehow anti AI? Its not anti AI. Its literally using AI!

…and honestly, at a higher technical level than slapping your wallet against a token provider and running prompts in a hosted sandbox you can't even see the prompts in.

[deleted]

Look man I am incredibly skeptical of how LLM’s have been rolled out and all the promises people make (it’s so much snake oil and pipedreams), but I also found it very trivial to hop on LM studio and start tinkering with models. If you’ve already got a decent midtier computer on hand, which I imagine a lot of us already do, then it’s really not hard to get started and get immediate results.

Local models on regular hardware aren't really capable of anything. Whatever you're testing is nowhere near a measly $20/mo subscription, so it's of limited use.

i really like the idea of running local models but i'm always in the position of wanting the best model(s) available and i don't have any severe privacy concerns. as such i have yet to justify ever using local models.

This is the diametric opposite of the rent-vs-buy scenario that this entails.

Local: You need to invest $thousands into GPU and/or very-high-end CPU+Memory hardware.

Vendor: You can use any existing device, even a phone or tablet. A very low-end laptop is fine.

> takes literal minutes to get started

Local: Typical scenario is hours just to download the software, the model weights, and then faffing around with CUDA and matching your GPU drivers.

Vendor: Free-tier available instantly on a web URL. Even local agents have free tiers from multiple vendors. Install is a single command and/or download and "next,next,next,finish" wizard that takes ~1 minute.

> you can just `rm -fr` it and forget the whole thing existed.

I'm still cleaning up multi-GB model weights floating around in hidden subdirectories under my user profile from months ago when I was experimenting with local models!

Meanwhile I simply... stopped using Gemini. That was the entire process: I no longer actively use it. They stopped billing me for my token usage, because it is now zero. That's... it.

You have it totally backwards.

> I'm still cleaning up multi-GB model weights floating around in hidden subdirectories under my user profile from months ago when I was experimenting with local models!

Are you trying to say that local models are hard to use because... you're having issues handling files properly? I am not sure I get the argument.

I get the rest of the comment: local models require an investment upfront, and it is less convenient. It doesn't say that it is not cheaper, though.

> you're having issues handling files properly?

I guess they were using ollama, which does not tell you where it puts the models it downloads.

Filelight / ncdu are my friends for finding random 30GB directories containing cached models.

Personally, I prefer QDirStat. I just tried to use FileLight to compare, but the package seems to be broken on Lubuntu.

You two have just listed several tools needs to clean up after another tool: not a strong argument for “simpler”.

It took me about three hours total to set up a local model. I already have a GPU and I have fiber for the download. llama.cpp is not difficult to compile and has many backends. It can run parts of the model on different backends, like in the common case that the GPU doesn't have enough VRAM for everything. There are many step-by-step guides available.

Takes even less depending on your system. LM Studio or Lemonade and you are set up in minutes and now they can even tell you what models will fit with the memory you have.

And it would be in seconds if models weren’t that large and slow-ish to download! LM studio is such a noob friendly experience, pretty neat first experience!

Three hours is a lot longer than one minute.

> "I'm still cleaning up multi-GB model weights floating around in hidden subdirectories under my user profile from months ago when I was experimenting with local models!"

I used to deal with these kinds of frustrations too.

    fd --unrestricted --size +1G

    fd --help

    -u, --unrestricted...
        Perform an unrestricted search, including ignored and hidden files. This is an alias for
       '--no-ignore --hidden'.

    -S, --size size
        Limit  results  based  on  the  size  of  files  using  the  format
        <+-><NUM><UNIT>

Huh what? Qwen3.5-35B-A3B runs just fine with maximum context, on an RTX SUPER 12 GB, with offloading of some expert layers to DDR4-3200.

Same story on an RTX 4060 Ti 16 GB. MTP is a serious boost to tg.

Downloading the model is a simple hf command that HuggingFace's web UI even gives you.

llama.cpp is trivial to use, and so is llama-swap, if you want to use other models too.

If you don't know what arguments to run it with, you download ggrun and use that.

Local LLMs are incredibly capable and don't need expensive hardware. A $500 GPU will do. Or even cheaper.

This is all trivial.

Full model or a 4-bit quant? I have a 5090 and I'm not sure whether I should use a quant that fits within the VRAM or a much bigger version where I'd have to offload a lot to 64GB RAM and a beefy CPU (but still a CPU)

I personally run the Q6 quant on my RX 9070 XT (16GB VRAM). On r/LocalLlama there was a post recently as well, which talked about the degradation of different quants (for the 27B version)[0]

[0]: https://www.reddit.com/r/LocalLLaMA/comments/1vef79c/quantiz...

RTX Super 12GB costs $700. An openrouter account costs nothing.

A lot of people already have 12GB+ GPUs lying around for playing games, doing video editing, etc. I would not get a GPU or mac just to run LLMs personally, but if one wants to get such a device for other tasks too, it may make sense to eg choose a slightly higher (v)RAM variant if they want to run some bigger models. Then what you pay for the local llms is just the difference.

> Local: Typical scenario is hours just to download the software, the model weights, and then faffing around with CUDA and matching your GPU drivers.

Download LM studio, search models, click download, wait minutes, prompt and have fun

“If you have the prerequisite hardware, then… know which model you want out of thousands of a variants… and your drivers are up to date, then it is fast!”

This largely describes me. I'm skeptical of AI in that it's capabilities, while very impressive, are vastly oversold and overblown. Being skeptical of AI is not being "Anti-AI". That's largely the AI data centers are using up all the water and electricity types.

Maybe you're anti-AI because you're really anti-outsourcing your thinking to some remote corporation you don't control?

That's one of my main issues with AI anyways, the thought of having all my data go through some sketchy foreign (to me) entity with questionable motives and under a questionable regime.

Local AI solves for all of those.

I'm not against AI. I'm calling out the hypocrisy in the comment. I'm anti-AI, but will spend hours trying to setup a local model, instead of just getting access to frontier intelligence in 15 mins, and actually getting useful work done.

If you're learning about model inference, then it's a different and you are definitely not anti-AI in that case.

I think learning how to set up a local AI now requires less time than learning about potential pitfalls of token-plans and processing payments in corporate environments.

It's not that billing is complicated, but learning to set up a local AI is a lot more useful and more rewarding.

35B MoE is certainly a good and fast local model. I find 27B dense to be quite a bit smarter, so I daily drive that. I wish there was a ~100B MoE with maybe 10B active. It would be super smart and fast!

There was a 3.5 122B 10A release -

https://huggingface.co/Qwen/Qwen3.5-122B-A10B

I tried it for a bit, and It was not really worth its size. It got swept up in all the other AI news recently, but laguna s 2.1 I think is the best ~100B moe model right now

I didn't mention it above, but Laguna S is my other favorite model. I use Qwen a lot more, it's smaller and faster, but I like to switch to Laguna when I feel like I need a "heavy hitter" for certain huge or complex tasks.

What on earth hardwares do you guys have to be able to run 100gb models locally?! That's crazy! I'm here struggling to even get 27b models to run in somewhat usable way

Strix Halo, 128GB RAM. I got a refurbished Corsair AI Workstation for a smoking price ($2100) about two months ago. Lucky timing that it was in stock.

Strix Halo as well. Bought it for $1,800 new on sale and shoved an extra 4tb drive into it. Been amazing for local AI. Maybe not the absolute fastest thing (usually around 30t/s depending on the task) but has been awesome for a local AI box that I can solar power.

Nice. Mind sharing the solar side of your setup?

Couple of rack mount batteries and roughly 5kw of solar panels. Feeds into a subpanel so I can flip it when I want a couple rooms of solar on the house, or hook a generator up if needed. Can't power the entire house, but works well for thinks like computers, lighting, etc. And if I want to expand, just throw on more panels, or realistically, just throw on more batteries to store the juice.

Haha I'm on an Mac Studio with an M1 Ultra, 64gb ram. I bought it when it first came out, it just happens to be good for local LLMs. I have to use a smaller quant of Laguna S though (I think 4-bit? Not at my machine to check), as 8-bit and full size definitely don't fit in the 64gb I have.

brb, going to see if 2nd hand mac studios are available!

Yeah, a good rule of thumb is that the weights take up ~100% of the size of the model, so 100B bytes (8-bit quant) would be, well, 100GB and a 4-bit quant would be half that.

I don't know if that's, well, a rule of thumb, it might be, well, straight multiplication.

Oh, that is a useful rule to know! Thanks!

That's not exactly the math. Theres also vram needed for context. I operate several 72-128 GB machines and the larger the context the slower they go.

And the context takes space +kv cache. KV cache drives usefulness as your context grows, it needs to pull the kv cache.

Simplified, the context has to be run on every turn, so the KV cache supplies the processed tokens, so it just needs the new inpute.

Yeah, here I am sitting deeply deeply deeply regretting not buying couple CMP 170HX at $200 or $350, knowing I could just flip them ethically at purchase price if nothing came of it... I could have just casually built a 128GB dual A100 local AI monster

I'm working with a lab that has a few Ampere GPUs on infiniband and they are just not compatible with the latest quants and vLLM updates. FP8 is about as low as you can go.

But they're reportedly a soft nerfed GA100 64GB/40GB at $1200, that's not more expensive and certainly can't be slower than a Mac Studio.

quantized + offload

I have an RX 6700 XT with 12gb vram and 64gb system ram. running dense models like 27b is difficult, but i can run IQ4/IQ5 qwen 122b-a10b or 35b-a3b at ~20tok/s

IQ4 qwen 122b-a10b would mean 61GB total size and 5GB active, so about 5GB of the model loaded into GPURAM plus any generated context, and 61GB of weights loaded into system RAM? I don't know if that math is correct, but does that run well? Wouldn't that only leave 3GB of system RAM?

Usually theyre quantized. Also, there was a window where AMD 395+ W/128GB was just a high end $2500 hardware with unified gpu memory.

dgx spark, nvfp4 so I have spare room for KV cache (context)

MoE models can use system memory along with a GPU.

and get high token bandwidth?

Not badly so because MoE models(identifiable by "CoolName-xxxB-AxxB" naming scheme) have bunch of branches in the middle that only one out of all gets non-zero values. Each of branches aka "Experts" as well as top/bottom parts are significantly smaller than the whole, and so CPU emulation of CUDA operations mixed with GPU taking as much as possible become not so out of question, unlike for dense models("CoolName-xxxB" without "-AxxB")

Similar to a spark, which isn't blazing fast but usable.

Using qwen 3.6 27b for local coding as well and downloaded Laguna s 2.1 but haven't had time to give it a full spin yet.

Curious for any more experiences

I agree. 27b dense really did seem like the sweet spot.

A 35 A3B as smart as previous gen 27B would be a sweet point

> 35B MoE is certainly a good and fast local model. I find 27B dense to be quite a bit smarter

Isn't that just the definition of MoE vs dense ?

Full name is 35B-A3B. 3B Is the token generator thats selected out of the 35b available in the model, which is some layered jazz.

So it can be dumber but its quite capablr.

I use 27B in plan mode and 35B MoE in act mode. I noticed that is the best balance for me for consistent tool calls and intelligent planning. Takes some time to switch, but it's worth it for me.

Wonder if it's possible to share a common cache via cachyllama between 27B and 35B-A3B

I've heard 27B is smarter! I tried it some time ago but couldn't get it working with my oMLX. I need to try it again.

Honestly the 27b dense one punches way above its weight in a lot of domains, especially coding in my testing, so I think you will probably be disappointed.

In my case I would say they are comparable but moe models are looping and getting lost a lot more than dense models.

On the other hand having 90t/s with any local model is nice and Pi with loop police extension can prevent looping a lot.

Looping seems related to quantization and not the model itself. If youre digging deep into quants to get working context then yeah.

It was for me too but the new deepseek pricing is too good to ignore for now.

I honestly think that with my electricity prices running qwen 36B myself is more expensive than hitting the cache rate at deepseek.

Can you elaborate on DeepSeek (deepseek-v4-flash, i assume?). What does your typical usage pattern look like and what is your weekly/monthly spend?

I gave it a try for a few days (pi + openrouter + deepseek-v4-flash via deepinfra) and ended up paying ~$18 for rather light usage. Yes it's still cheap, yes it's fast, but i feel i would still get a better deal with a Claude subscription plan.

DeepSeek without OpenRouter is wayyyy cheaper

Is it? OpenRouter shows DeepInfra being cheaper than DeepSeek directly

https://openrouter.ai/deepseek/deepseek-v4-flash-20260731#pr...

I agree with parent. OpenRouter might be cheaper list-price, but i have been using 10$ on DS platform since April/May, still have 2$ left. Using OpenRouter i depleted the same dollar-amount in a 1-2 weeks with same usage pattern. No idea why.

[deleted]

What does your usage look like?

DeepInfra's Cache Read is 6.5 times more expensive than DeepSeek.

And in many scenarios, cache hits are 99% of tokens, so the price difference in caching really adds up quickly.

Interesting. My anecdotal experience is that it is.

Going directly though DeepSeek's API and hitting all day long with light/medium tasks I'm at $5/month. That's with pretty vanilla ohmypi.

[dead]

I would recommend looking into Ornith1.0 - it's using Qwen3.6 35B-A3B and excels in coding, at least for my coding needs, Python, web-dev, SQL scripting and some C#. Using Pi harness.

I have been using Qwen3.6-35B-A3B as my daily driver as well and its been phenomenal when it comes to coding

How do you use a 72GB model as your daily driver locally?

Apple Silicon. But: there's no need to use the FP16 version. At 8-bit precision the quality loss is almost imperceptible. That cuts the footprint to 36GB. Which is great for a 64GB Mac, because you have room for plenty of context. 6-bit also works nicely at 26GB + context.

You want to use the newer quantization formats like Unsloth's UD quants or oQe, where the weights are selectively quantized using a calibration dataset so that important weights are left at/closer to full precision.

use a quantized version. since it's MoE, what matters is that the 3b parameters that are used for every token fit in gpu vram, the rest can stay in system ram. really great if you don't have unified memory.

At 4 bit it easily fits in 32GB. That's what most people use.

What kind of machine do you have running that? My attempts at local have always resulted in a very hot lap

I host the models on my Mac Studio, an M1 Ultra with 64gb ram (I bought it when it came out, just happens to be good at LLMs). So when I work on my laptop, I have my oh-my-pi setup configured to use the models on my Mac over my local "bonjour" network or whatever Apple calls it. That way I have a nice cool lap, while using models that my M4 MacBook Air with its 16gb ram couldn't possibly run.

Cool yeah. I got a M3 with 32GB ram and it’s a little iffy. I’ve considered getting a MacMini to act as an in-house

Strix Halo for me. If I am running something on my laptop, it's a much smaller usually around 12b model, but those are a bit less functional. I mean I think there is a a ROG FLow Z that has the Strix Halo setup, but that thing was super expensive.

With laptop being ...?

It's just an M4 MacBook Air with 16gb ram – probably incapable of running models itself. I actually run the models on my Mac Studio which is an M1 Ultra with 64gb, and oh-my-pi on my laptop is configured to use the models over the local network.

Is your Qwen3.6 locally run on your laptop? What kind of tokens/s are you getting from your laptop GPU?

Qwen is running on my Mac Studio, an M1 Ultra 64gb. My harness (oh-my-pi) on my laptop is configured to use the models hosted on my local network, since it's just a MacBook Air 16gb and probably incapable of running anything useful itself.

I get about 45-55 tokens per second using Qwen with this setup. I could probably squeeze out more if I messed around with the settings, but I'm mostly using oMLX's defaults for the model.

I have Qwen3.6 35B-A3B on my laptop and it does 60 tokens/s

Could you share the specs of your laptop?

M5 Pro/64 GB RAM

compared to claude - how 'fast' is it in terms of throughput on your laptop?

On my SpacemiT K3 SBC with 32GB RAM (where models run on the eight A100 RISC-V cores with 1024 bit vectors) doing the same task I got 5, 5.8, 6.5 tok/s using gemma-4-26B-A4B-it-QAT-Q4_0.gguf, Qwen3.6-35B-A3B-Q4_K_M.gguf, Qwen3.5-35B-A3B-Q4_K_M.gguf. The corresponding dense models are more in the 2.5-3 tok/s range.

Kind of slow, but using only 14W of electricity so the Wh per task is twice as good as using my i9-13900 laptop with 4060 GPU.

I use it with a strix halo server. 35B runs stupidly fast. 27B is about 700 TPS prefill and 30 TPS token generation. Which interestedly is about what Kimi K3 gives me depending on provider.

what hardware do you use or recommend for this? never heard of it until today.

Strix Halo is the unified memory platform from AMD. Similar to the DGX Spark from NVIDIA or the M series Macs.

I personally have the Framework Desktop, but there's also systems from other brands like Bosgame

You can also get it in a laptop form factor that feels like a MBP with a nicer keyboard if you get an HP Zbook G1A!

Huge fan of that thing, it's th e Linux MBP I've always wanted.

While the laptop option is nice, for an inference server you're probably going to want the desktop form factor as it has significantly more thermal overhead and thus better performance. In the desktop models most of the internal volume is a gigantic heatsink

I have a framework desktop, but depending on your need, DGX spark might be better. The prefill and NVFP4 is a significant advantage. But framework desktop is a better general computer. I expect to be able to use it for years to come. Where as DGX Spark you’re at the mercy of NVIDIA BSP.

RTX 4060 and above. Ideally RTX 50 Series, because you can run NVFP4-quantized GGUFs that give you better prefill AND better quality.

On an 8GB GPU and 32GB laptop: ~5 words/s while running in Qubes via ollama with completely default settings (I don't have an install at the moment that'll tell me tokens/s). Not exactly a highly tuned setup, but it's a ballpark at least :)

Tolerable and usable for some things, though thinking makes it take about a minute to reply in many cases. But getting this kind of thing to run on 8GB of VRAM is the main benefit of the mix-of-experts setup: it can do partial GPU loading and get a ton better throughput than a similarly-sized dense model (like 5-10x, sometimes more).

It's pretty fast, faster than I could type anyway, but not as fast as Claude of course. My oMLX dashboard says I get about 45 tokens per second from the Qwen model I'm running (I host it on my M1 Mac Studio, not on my laptop).

[deleted]

[dead]

I'm a big ole noob when it comes to local AI. What are you using for a harness? Or platform to interact with it?

I recommend trying pi.dev as your agent harness for local models. In my experience it has been the sweet spot of functionality (which you can and should extend with plugins) vs performance (OpenCode just swamps local models on my hardware).

you can try ollama, omlx or llama.cpp for instance to download a model and get an inference server running locally. They expose „open ai compatible“ endpoints, so you can configure almost any harness to use them.

What do you use to pair it with web search?

Depends on how you are doing it. LM Studio and tool calling models can use the web, or you could go for something like Perplexica, or if you want to go real crazy, something like Hermes or OpenClaw.

I use searxng.

On what hardware do you run the model locally, if so?

Not the GP, but I run this model as daily driver too. It runs great on a Macbook Pro 64GB (M3 Max). Token generation speed can be about 100 tokens/sec with multi-token prediction, although it depends on the context. Worst case speed is around 50 tokens/sec.

The weaker point is prompt prefill, which starts at 1,400 tokens/sec but decreases significantly at high contexts. That said, for agentic scenarios, if you're using a harness that doesn't needlessly bust the cache, it doesn't feel slow.

I really hope they release a Qwen 3.8 35B, although the lack of a mention seems ominous.

What are the specs of your laptop and what tokens per second do you get?

It's just a Macbook Air with the base M4 and 16gb ram, but I'm hosting the models on a Mac Studio with M1 Ultra and 64gb ram that I had purchased when it came out. I get about 45-55 tokens per second with this setup. I think I could get more if I spent some time fiddling with the parameters, but I don't really know what I'm doing there so I've just left most of it on oMLX's defaults.

[deleted]

I'm on the verge over here, the new Anthropic models have been a disappointment. I've tried the A3B variant, but had mixed results. What do you use as the coding agent, and have you heavily customized your workflows?

What laptop?

It's just a MacBook Air with an M4, cheap and nothing special. I host Qwen on my Mac Studio, an M1 with 64gb ram. The model uses around 20-25gb ram depending on what it's doing.

I wonder if I could get this running on my 48gb M4 Pro. Haven't been able to load anything beyond 27B

[deleted]