The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.
I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.
This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
Yes. You have to find the provider with pricing that suits your usage.
I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.
If your tasks are write-heavy, find a provider with cheaper output.
If you build a customer-facing app, pay a bit extra for 400+ tok/s e.g. on Lithos.
What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?
Sheesh... what's in a name? :p
Bio shows: https://twin.so/
Codes day and night a snake eating its own tail…
Way too harsh
Lots of guys buying articles on TechCrunch saying they’ll build this, he’s bootstrapped
You have all the struggles for the price of Anthropic / cursor subscription. I use the first one I code large chunks some PR are 50k LOC and I have at least 2-3 like this a week . It’s a greenfield project .
I still have quotas left I use it for home things build 3d model of my renovation projects, alerts for shopping list etc . And yeah I use cutting edge of cutting edge of models that saves me time and money , only discount monitor saved me ~$2k on my renovation project
I am curious how you managed to spend that much on Deepseek via OpenRouter. I loaded $100 back in July while using v4-flash or whatever the cheap good model was at the time, and have upgraded as the new ones came out from Deepseek. I still have $16 and some of that spend also goes towards the AI usage from my customers (the context they need to load in is quite large too).
And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.
It's wild how different usage patterns are between users.
I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.
People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.
Those usage patterns don't correlate to output.
Now that Microsoft allows you to do /cost for individual tasks (or whatever they call them today). So I tasked Sol, Astra and Fable in cowork with exactly the same vibe coding task on the exact same zip file containing a code project I needed an update for. Astra used 20x and Fable used 15x of what Sol did.
Fable changed a lot of things I had explicitly told it not to change. Arguably a lot of them would've been correct if you didn't work in a place where abstractions are directly against the core principles, but what it produced was basically unusable. I'm not sure if Sol or Astra did best, they produced rather similar code outputs. Astra's was better, but Sol didn't do so bad. It forgot to clean up a few places after it's refactor and it made two bugs I had to correct but other than that it was fine. Astra on the flip-side might have produced code that didn't need changes but it also rewrote every piece of documentation so that it became horrible.
As far as the "experiment" goes, it just shows you that the credit consumption is basically pure magic. You'd think that the Microsoft AI admin tools and the Agent365 FOMO DLC license they sell might give you some sort of reporting, but it doesn't. What you can see is how many tokens a user consumes and the total number of tasks they've initiated as well as whatever running agents they have. You can't see what models they use or which tasks are expensive, which makes it very hard to help them. Early on we had an employee who hit their limit in an hour, and it turned out they had basically uploaded a lot of information and run it in a single long task that kept going over it again and again. We told them it might be a good idea to only give it what it needed and to create more tasks, and even though it's been three months, they have yet to consume as many credits as they did that first hour.
But that's how you support and track it. You see a user spend a lot, then you go to their computer and now that you can actually do the /cost thing, you go through their tasks and try and figure out where they're spending money...
It's obviously improving. A month ago /cost wasn't there and they just released a new dashboard for cowork, but it's still black magic that is impossible to govern.
Likely the user doesn't know what they're doing or has extermely bad workflows. They're prob not managing their cache, and dont use compaction.. Letting context get to 500k and invalidating their cache every 10 tool calls because they have no providor fallback settings.
I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool/mcp usage and finding it very difficult to spend more than $75 in a month.
Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.
opencode DeepSeek v4.1 Flash isn’t us/eu hosted as of recently, so not sure how this impacts privacy / model training
But aren't you developing bad habits and learning patterns that won't work long term? Or do you think things will get cheap enough that you will be able to keep going with your current patterns post-subsidies?
Compared to what a lot of companies spend on software for chip and electronics design (we're talking about $10k-200k/seat per year), AI coding assistants have a long way to go in cost before companies won't be willing to pay for them. Companies pay a fortune for software when it enables their engineers to be productive.
For my company, I'd honestly pay $4-8k/month for Claude if I had to (it would be painful, and I'd try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it's at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it's well worth it.
If they stop subsidising Claude Code for the pro/max users, there will be a lot of people priced out of it, especially the casual developer. But I don't see it going away for commercial use, even with a large price increase.
Nah dude I think it’s worse than that
Old coding is done, as a workflow in teams. It’s the top down executive pressure of being non competitive as a company, and the bottom up pressure of human laziness
Show me people handwriting code à la NASA
And I mean we as coders have been trying to do this workflow for a while, I personally would refuse to code without IntelliJ magic complete
No one involved in this is thinking about the long term
> But aren't you developing bad habits and learning patterns that won't work long term?
2 reasons - there's an advantage now, use it. 2nd the frontier providers, this is the "early cheap days" like when uber was initially cheap to compete vs standard cabs. they want you to become hooked and boy are we hooked.
I expect by that point we'll have local models that can do a decent job, I would guess give it a decade and we'll be running custom accelerators that are smarter than current frontier models.
In the same way that only supercomputers used to have multiple processors and caches but it's now standard.
I use the frontier openai/anthropic models at work but exclusively open weight models (on cloud/hosted inference) for personal stuff and I think about it like this; 1) I don't see any reason GLM and DeepSeek won't eventually be as good as Claude, it's just a matter of time and 2) the open models are well and truly capable enough for most of what I'd want to do. I don't need nor want an LLM chewing away on a horrible enterprise spaghetti codebase, my employers can pay for that privilege.
Long term, we will see what happens and adapt. At worst we all go back coding by hand. Meanwhile what can I do, tell my customers that I'm raising my fee because I have to pay for token? The Claude Pro $20 plan is good enough for me and even in auto mode I never had to wait for the 5 hours reset.
How's the caching? I have 99.5% cache hit rate with deepseek when using their own API, it's dirt cheap.
OpenRouter is complete garbage.
Buy directly from DeepSeek's API.
You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).
[delayed]
I'm all in for saving money and _can_ move to using DS directly from them, but maybe I am missing something here:
OpenRouter Pricing:
$0.02/M input tokens $0.60/M output tokens
DeepSeek Pricing (cache miss, off-peak):
$0.15/M Input $0.60/m output
When 98.5% of my requests are cache hits (according to Pi for the last week), the cache miss price isn’t that important to me, and $0.003-0.006 per 1M input tokens is shockingly cheap.
It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?
Interesting, if the cache hit is that good, I think HN convinced me to toss $20 at DS official, and see how long that lasts.
It will of course depend on what you’re doing with it, but right now my session at work has a 99.8% cache hit rate, and I’ve been running this session for hours with 23M tokens read and 713K tokens written (Opus 5.5 in this case though)
I've heard that certain inference providers may have different quality of caching implementations, so even if the listed numbers are as you say, the practical cache hit % you get might be significantly different/incur significantly different costs.
There’s a big difference in speed & quality between using DeepSeek API directly with DSH vs. DeepSeek in Opencode Go with Opencode CLI. Can’t tell if it’s the provider or the harness - but worth to give it a try.
One of the reasons I use OpenRouter is because they offer zero data retention. As far as I can tell, DeepSeek's own API doesn't support ZDR.
DeepInfra does and it's the same price. That's what I use.
Youre doing something so special you need that?
I've got a toilet cam to install in your bathroom
DeepSeek trains on your inputs. That's why people go on OpenRouter and choose ZDR providers.
Let me get this straight you guys really like deep seek because it’s open but you don’t wanna help them improve.
No, we just want a choice on how to license our work.
Or.. BYOK Deepseek because OpenRouter's UX is much nicer?
Zero Data Retention and not having company source code leak to "CHINA!" (said in Trumps annoying voice) would be two reasons not to
Which provider was that?
Also DeepSeek usage is subsidized as well, it’s a power hungry model.
Interesting. Are all of the providers on OpenRouter simply losing money? How does that even work out?
No, it’s outrageously profitable above x% utilization without stealing any prompts. Provider economics still pretty good. Acquiring hardware is the current limiter.
You're the RLHF.
Not if you're not giving feedback.
With ZDR-only enabled? That seems illegal.
Don't ask and dance as long as the music keeps playing.
Are you sure about that? My impression was most providers on openrouter were purely selling tokens for profit...
Have y'all tried an Ollama Cloud subscription? Their off-hours pricing for V4.1 Flash is extremely competitive.
You should basically never pay API prices, they are always several times higher than subscriptions.
There are several open weight subscription providers. OpenCode Go used to be good but now it's complete shit. Charm Hyper is really great and the best value. Other subscriptions have a more limited model selection or provide less value but are still decent.