What stack do you use for that home inference? I’m rolling my own with gemma4, but even with Claude helping with the implementation it’s a lot of moving pieces to figure out.
What stack do you use for that home inference? I’m rolling my own with gemma4, but even with Claude helping with the implementation it’s a lot of moving pieces to figure out.
Ollama running Meta's Glimmer model at the moment. I find it really capable for what I'm doing and it doesn't suffer censorship problems.
Are you using the API or is it self hosted as well?
(sorry if this question doesn't make sense)
It is self hosted.
Ollama serves models on an OpenAI compatible API directly on the inference machine, which most open source harness software will connect to. The inference machine has a name like inferencemachine.local on my home network so the Ollama URL looks like http://inference machine:11434/v1 to connect to any model I've downloaded.
> AI agent on my phone I can long-press power button and give spoken commands ("add milk to the shopping list"). Uses private home inference machine, see below.
For this part, are you using a tool-calling harness like Pi or Codex? Or did you sort of hack it together with prompts that tell it what json to emit?
I just got the rest of the “application layer” for my setup working yesterday, with a small set of evals. Auto-filing voice notes is what’s next for me.