We’ve literally saved tens of millions of dollars already (no exaggeration! already 8 digits) by switching to Luna for many workloads at my company.

The amount of workloads we can shift with an advisor model pattern continues to grow.

It’s seriously amazing.

Luna came out about a month ago, you're saying that the cost saving from switching to Luna has saved your company $20 000 000+ in 1 months spending on API usage?

OpenAI and Anthropic are believed to be making billions per month in revenue selling tokens to enterprises. It’s not that surprising for enterprise SaaS they would have individual customers making up ~1% of their total sales. Although, perhaps a bit more surprising those few customers are posting about it here.

Another interpretation would this is a counterfactual savings, like they previously paid $1M for y tokens, and now that tokens are cheaper they increased usage and paid $1M for 20*y tokens.

A second angle on the counterfactual savings would be Luna telling them not to pursue a potential session with the expected/extrapolated (from which sessions they did ignore Luna on, e.g. just to keep efficiency statistics current) sunk costs at time of getting shut down used to derive the quoted number.

A third angle is that it is not true and that's just some attempt at influencing an HN thread, which is happening quite regularly pro and against AI. We know that the pro-AI has a way bigger budget to burn with this kind of operations tho.

It makes sense at a company that sells AI as a service. Perhaps to make presentations...

The ones with terrible slides on their MTA ads?

Canva generates billions in annual revenue, it's not unbelievable

By "advisor model pattern" you mean https://claude.com/blog/the-advisor-strategy ?

How do you implement that outside of claude code?

the easier way to do this, with any harness, is to use an expensive main agent that is told to delegate all code reading, writing, exploration, research etc to weaker subagents to conserve tokens. It's an inversion of the pattern but the resulting split is the same

Yeah, that one is easy, but I was thinking of the other way around, where the weak model calls stronger one.

I append this to many of my opus claude code prompts

`You may use a Fable subagent to answer questions, solve problems, and provide an adversarial review of your ideas and code`

You can use a similar pattern in most any harness, and you can tell them to use other harnesses. In claude you can write `Use codex cli to have Sol56 Xhigh provide an adversarial review to your plan before presenting it to me` or `Use opencode cli with GLM 5.3 to verify all code review findings before presenting` or whatever you're doing, as long as those other tools are setup and ready to be called.

IMO: This isn't useful as a token saving pattern in my experience with agentic engineering, but it is useful as a quality-enhancer.

what has a single company accomplished with tens of millions of token spend?

A pelican on a bike accurate to a subatomic level

But the knees still bend the wrong way

higher valuation

What are you using it for?

[dead]