Tried yesterday on my own laptop (a UltraCore 7 255H without dedicated GPU,with 32 GB RAM), it wasn't even starting thinking, even on a small context window (65k)
Tried yesterday on my own laptop (a UltraCore 7 255H without dedicated GPU,with 32 GB RAM), it wasn't even starting thinking, even on a small context window (65k)
I think not much can run without a dedicated GPU
Till now I was using successfully Qwen 3.5 and Gemma 4 at a reasonable speed
I think (maybe I missed something) that identical size and quant versions of Qwen 3.5 and 3.8 should run at the same speed. It’s the exact same architecture.
Tried the 4 Bit versions. It loads, bit the <thinking> output isn't even coming out.
There is no way you would run a dense 27b model on that spec. I ran 3.6 27b on a 64gb ram, 24 gb vram, and it felt like the lower limit for this model with a decent context window.
If you want a better experience, maybe wait for either a moe model (like 3.6 35b A3) or a model with less parameters (like 9b). Qwen has been releasing those in the past, so maybe we’ll have them for 3.8 too.
From my experience if it doesn't fit on vram it is rarely worth to bother except for a few narrow tasks.
For example make an essay about something where you don't actively engage with the LLM after the initial prompt. So mostly one-shot prompts.
were you running the MoE models? those perform better speed wise
What's the story with Mac laptops? Worth a try?
The author tested in on an M5 laptop too:
> It feels pretty slow on both the M5 Mac and the DGX Spark.
Dense ones like this are more bandwidth-hungry, so you want to try MoE ones like Qwen3.6-35B-A3B (35 Billion params but only 3 Billion Active) or Gemma 4. Unfortunately it seems like we might not be getting a 3.8 MoE.
Yes. mtplx runs it at 25 tok/sec on a M4 Max with 48GB RAM.
> it wasn't even starting thinking
Probably stuck in prompt processing which is compute bound especially for iGPUs.
You've mentioned 3.5 - but it's actually the same model the only differences are training and implicit MTP support (affects prompt processing - can be disabled)
Have you tried with different amounts for the "reasoning_effort (xhigh|medium|low)" parameter?
Or the "<|think_xhigh|> | <|think_low|> | <|think_off|>" tags: apart from this template detail, it is not immediately clear if reasoning_effort is deterministic (API) or is prompt engineering.
No, good point. I will have to tried it
What was your prompt length? It's possible it was just processing it and it's likely not fast on your setup.
Really small (<100 tokens), I wanted to test its capabilities
i have same 255h and i was able to run it with low token speed 6-8tg/s with approx similar context window 60k
Interesting. What are you using? I was using ollama
Don't bother with plans