I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition.
Baking models onto silicon would've been the next logical move to get a moat.
Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.
The Taalas chips are not physically small. And part of their secret (if you look at the design) is just locating a bunch of memory soldered on the edges ( I belive higher amounts of SRAM ? )
Baking the base models on to ROM makes a lot of economic sense. SRAM for the KV cache & fine-tunes, not so much. Sure you’d get incredible speeds but it’s not scalable from a die-size or cost perspective.
Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.
I think the lifecycle for these chips could stretch far longer. If you're offering these models on a two year lifecycle, then you'd be able to stand up your top tier (wouldn't need to be frontier) at high speed. Run (for example) Kimi K3 on it and give it a brand name:
AcmeAI Carbon
Market it as your premier (only) model at high throughput. Two years later you stand up MSICs for the new state of the art with entirely new hardware, your lineup becomes:
AcmeAI Nitrogen (top tier) AcmeAI Carbon (mid tier)
If you just kept pushing the same model down your pricing tier over time you could still extract a lot of value from an old model, even years after it's been set in stone. Working on brand new code/frameworks? Pay to use the newest model. Working on legacy code? Use the lower tier models that will already know your legacy frameworks, pay far less and still get massive throughput. I've worked on a lot of government projects that this would be absolutely brilliant for.
The other side of this is that agent harnesses are NOT set in stone, so even a legacy model with a knowledge cut-off that's years out of date can likely still be helped quite a bit by harness and fetch behaviours that are still developing rapidly. Especially at this kind of throughput.
I'm somewhat doubtful that we will be seeing something as large as Kimi K3 in silicon any time soon.
This tech can definitely scale up from the current 8B prototype, but - at least as far as my limited understanding of the tech involved goes - you cannot just ASIC a trillion weights model due to physical size constraints.
___
Specification HC1
Model Llama 3.1 8B (hardwired)
Process TSMC 6nm
Die size 815mm²
___
So the current prototype already pushes the limits of what we can fit on a single die, and that is already likely going to limit your yield.
This is an architectural limitation that may be overcome by how you bake the MoE (mixture-of-experts) onto silicon.
If you could manage a per-die expert somehow and keep the expert routing gate relatively fast (through an interposer interconnect or doing wafer-scale Cerebras type shit) you don't need to keep the whole thing on the same die. Small dies with one expert per die on an interposer, and a very tiny router might be sufficient.
Kimi K3 is huge, though. Deepseek V4 Flash is a much more moderate model (284B total), and it works extremely well. Models of that size, and smaller, are just going to keep getting better and better. Presumably there's a threshold below which models are not generally useful or competitive, but if models-on-silicon can scale up to just 256B, that would be really remarkable.
My 2 cents to for your point:
- Deepseek V4 Flash is impressively capable. Sonnet still beats it out by a thin margin, but the real kicker is that a typical session with Sonnet at current API costs is ~$2. The same session with Deepseek is 2 cents (ha). Its even allowed me to consider offering free-with-limits API usage on my own app. - Taalas (or competitors) have a lot going for them. If anything I feel like they need to join hands with these smaller model makers and converge in 2028
AMD can let the SRAM be on a different chip. Maybe even something similar to their 3D cache. that could increase density to 20B[1]. They could also move from 6nm to 2nm. that would probably increase density by another 3x to 60B.
Add a bunch of chips together, and you get to a server that can run a 800B model, very fast and probably significantly cheaper than others.
[1]https://www.eetimes.com/taalas-specializes-to-extremes-for-e...
Baking the base models on to ROM makes a lot of economic sense.
Less so for consumers though, because it'd mean the phone is out of date in 3 months when a better model comes along.
It's a perfect reason to get consumers to buy a new phone every year again! They got bored of the camera.
Right now the models are doubling in performance (by the METR time horizon metric at least) every 4 months, so 3 doublings in a year; conversely, I hear (not my field) it takes around a year to make a prototype IC and another year to turn that into mass production, i.e. if the next (late-2026 model) iPhone has a chip like this, it will likely be with, at best, a late-2024 set of weights. I think you can get open-weights models today that have performance equivalent to the SOTA-late-2024 while fitting in the RAM of a (high end) 2025-26 phone.
At some point the music will stop on training bigger models, and when that happens it will make sense to have ROM weights (or 100% analog circuits given how noise-resistant LLMs are), but we'll know when that is because the investment bubble funding the training of new models will have burst.
wait, analog circuits? can you elaborate this?
Transistors can be used to amplify signals, they are not limited to acting as binary switches. If you use analog rather than digital, using transistors in this way means you can replace however many transistors it would have taken for multiplying two n-bit numbers with just one; I understand capacitors can be used for accumulation, but don't know how many additional components that needs as I'm an electronics noob.
The reason we don't do this in general (any more) is that for long chains between input and output it has been much too difficult to avoid accumulation of errors. LLMs happen to be extremely resilient to errors like this, which is also why we can use e.g. 4-bit weights.
Circuits that aren't restricted to two particular voltage levels.
At some point it’s got to be good enough for the normal “phone stuff” that appeal to most users. So they wouldn’t suffer from FOMO because they didn’t wait for the next model. Every phone gimmick went through the same evolution curve until it passed the “good enough” point and eventually plateaued.
Yes, but irrelevant. While these models are improving at the present rate, the manufacturer can save money at no loss of feature-bullet-point-on-website by letting you download a model after you bought the thing and running it on normal hardware.
The rate of change to the models has to be slower than the hardware roll-out to be worth a hardware solution. If "good enough" happens before then, that just means the user gets a software solution.
You might be right but hard to tell without analyzing costs and benefits. Is a cutting edge model for phone stuff worth the slower performance and battery drain for example?
The rate of change by itself doesn’t tell you the whole story because of costs and diminishing returns. So what if your model is twice as good if it’s 10x the cost and it saves you 1ms? Everything else about phones reached “good enough for a phone” levels in years, and then got minimal generational improvements.
This stopped working some years ago.
Now you sell the same phone with higher price tag.
I don’t think average user _needs_ to solve frontier challenges. ”Call to Jane”, ”turn on the lights” and ”what’s the weather this afternoon” is more like it I would guess.
Ofc if the model has some critical bugs that’s another matter.
Your examples worked on phones for over a decade.
Maybe baking in a model that is "certified" to have some unconditioned truths + rest is pulled from external models/store could make sense. But AFAIK that doesn't exist and I'm not sure it can possibly be made. Perhaps society as a whole at least can work on an open corpus of training data, but I'm not holding my breath on this.
It barely works even today, like Siri is laughably bad.
I mean the examples he gave definitely work. Mostly well I'd say as they are pretty primitive.
What Siri is missing is more logical solutions and answers for recipes, etc (still suck even with chatgpt integration).
No, they don't work. Just asked Siri the other day "what's the weather tomorrow in $LOCATION" (where $LOCATION is a broader zone and not strictly a city) and the answer was the weather in a street called "$LOCATION Avenue" in a city 150km away.
Works, sometimes. But they can fail spectacularly and unexpectedly even on very basic questions/instructions, like so simple that a hand-coded word-matching style logic could get them right 20 years ago.
I just asked Siri
"Hey Siri, what's the weather in <nearby town with a generic name> tomorrow" and it gave me a town with the same name ~800 miles from me.
>Your examples worked on phones for over a decade.
Nope. And not only not a decade ago, right now.
If you have an Android or iPhone, you can give it clear and easy to understand instructions that Gemma 4 could complete[1] if it had tool calls on it, and that 100.00% of Claude, ChatGPT, Grok, Kimi, you name it, could understand and all complete if they had the access.
The phones will fail to complete it. I just tried Siri. I said "hey Siri", waited for Siri to come up, and then I asked one of the exact sentences you replied to: "what's the weather this afternoon?" It thought for around 20 seconds, and said "Something went wrong. Please try again."[2]
I have Wifi, I have mobile Internet, I have free storage space, I have up to date software. What went wrong is that phones have never properly connected agents, not ten years ago, not last year, not this year, and probably not next year.
But don't settle for what Google could do in 1999 by hotlinking the keyword "weather" in any query to the weather being shown in the results.
Tell your phone (any phone): "Please call back the last number that called me that is not an unlisted number, regardless of who it came from."
0 out of any phone will complete that today, tomorrow, a year from now, five years from now, ever, because phone makers are not going to let them do that.
Meanwhile, 100% of all frontier agents could complete it if they had tool calls on the phone. Which they don't, and won't ever, thanks to the duopoly.
Okay, that's a bit dismissive, I would love to be wrong!
[1] after any voice recognition to text - which does work really well on both Android and iPhone! [2] screenshot: https://ibb.co/21rtDnfV
I’m on the IOS 27 beta and Siri did those two tasks (weather/phone) flawlessly. It’s a lot better than it used to be.
Thanks for trying that! Very interesting.
Can you say this to it: "Hey Siri [wait for it to come up] - please send me an email with the temperature right now so I have it for my records." and see if it can complete the task without any backtalk or misunderstanding, and if you get exactly what you asked for. (It's a really clear request.) Should be 1 statement, no clarification, conversation, random search results, ("Here's what I found!"), etc.
A normal frontier model can do that - or Siri can do it if it is properly connected to Claude, ChatGPT, Gemini, Grok, or any other frontier AI - but previously it was never properly connected.
If it can do this task, I might have to look into this again. It counts as a success if it sends yourself any email with the current temperature and you actually get it (it can include whatever other text in the email), and a failure if it talks back, says "here's what I found", says it can't, asks you any question, sends you an email that doesn't actually contain the current temperature, just reads you the temperature and then asks if you want it to send an email, etc. Should be 1 shot.
let me know if it works!
It brings up a preview of the email and you have to tap or tell it to send it from there but otherwise it worked for this as well.
Subject: Current Temperature Body: The current temperature is 27°C in <my city>.
thanks! useful.
In my experience, they had a lot of stuff working well in the first few years they rolled out the home voice assistants - Alexa, google home, etc. But for whatever reason, they've spent the last eight(?) years silently breaking things that used to work. Stuff like audiobook playing, music alarms, or even messaging people.
Once they started seeing useful (if niche) functionality as a cost center, there wasn't really a world in which these could usefully exist. Their big bet now seems to be that LLMs will lead them to profitability - but whether that's from increased data harvesting, cheaper integrations, or because it'll be useful enough to charge subscription fees, I couldn't tell you.
And the customers can wait for the new phone released next year. These are edge models - the average customer doesn’t need the latest frontier model. Just needs to be good enough for the features you promised.
That's a software engineering problem. They just need to figure out how to fine-tune for alignment and tool usage.
That's the only thing the normie consumer cares for really.
There's in-betweens; '1T SRAM' or eDRAM.
Of course, 1T SRAM isn't really SRAM, but my understanding is it doesn't require external refresh like eDRAM, is a bit easier to fab on-die than eDRAM, and is half the mm2 per Megabit compared to real SRAM (15% more die size than eDRAM)...
We also have ReRAM (Analog Computing), which also holds a promising future given its efficiency and low power. Though ReRAM of larger size is still a research area.
It is my understanding that just baking the model itself into silicon only gives moderate gains because memory bandwidth remains a bottleneck.
Did you try using the the talaas chat? Something stupid like 18k tokens/second.
Think it's called Askjimmy or similar.
What is the model they're using there though? Interrogated, it claims it's a BERT variant and has capabilities around GPT-3 and below GPT-4.
(Not that I believe it, it writes too well for GPT-3.)
Hosted frontier models from two years ago would be much faster today, too.
They run Llama 3.1 8B.
Oh boy, thanks for sharing this, truly mind blowing. It was chatjimmy.ai
The big benefit is ROM cells require fewer components than DRAM. So the chips would be tiny, dense, cheap and consume far less power.
It's not even just that. If you just built the rom chips separately and swapped them for the RAM of a normal accelerator, it would not help at all.
The trick is that every compute element in their system has it's own small pool of ROM, instead of putting all the ram behind a common pipe. ROM is just used because it's the densest kind of memory that can be fabricated on the same process as their logic.
I thought DRAM was pretty dense already. Is mask ROM that much denser?
Yes, each rom bit can be a transistor or even a diode with a decoder circuit. Simplest Dram cell is capacitor+transistor - and you need a clock, refresh circuit etc.
Someday, I imagine model weights could even be encoded as analog resistors (memristors or similar) for even greater density
Hm. I wonder how many relays I'd need to make a physical MNIST classifier. That'd be dope
Their PoC chips are big, but then it's ridiculously fast (have you seen chatjimmy.ai?). Also they must be holding a bunch of patents.
Its a cool demo, but its gpt-3.5 level stupid, or worse.
edit: Ok, I will self-apologize. Its apparently a 3B model. Mighty impressive for what it does.
It's a quantised 8B model (Llama 3.1 8B to be exact).
[1] https://taalas.com/the-path-to-ubiquitous-ai/
How time flies. Just 3.5 years ago, gpt-3.5 was touted as almost AGI, we're all going to be replaced by machines and worst case they will kill us all. And here we are, not much later, and it serves as the benchmark for "stupid"..
Sure but the first step to having something that is physically small, small enough to cram into an iPhone, is to have something that, at first, isn't.
I don't think this works out from a cost/silicon perspective. Small models already run pretty well in software (since the weights fit in cache) and big models require silicon area proportional to the size of weights. On a mobile device putting a chip like this is competing directly in BOM and power against a whole lot more l3 cache, and the l3 cache makes everything faster
My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.
Did you use chatjimmy? It's somewhat terrifying to use when you think of the potential results with a better model.
Ok, real life example: I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). What if it gave back the same excellent results, but instantaneously? Why, then, I certainly would become the bottleneck. So, quite possibly, my last work task would be to plug this agent directly into the ticket system where the domain experts input their feature requests. Maybe we still need 1 developer out of 100, to coordinate releases and all that (ok, say 1 out of 10).
But that's not taking things far enough: why do we need these domain experts at all? Our pitch is clear, and all software-enabled, though it took years to develop. We can just have the clients express their concerns to the AI, directly or indirectly. Have multiple lighting-fast agents with different roles (refactoring agent, new features agent, debugger agent, domain expert agent, etc.). So we fire everyone, maybe keep 1 product owner / devops to keep the trolls out. The cost is still probably 100 times less than it used to be (beyond the initial cost of acquisition of the magic machine or whatever).
But one of these clients, surely, will realize that these 10 years of manual and slowly-automated development can now be emulated in very, very little time. Why not just, say, take screenshots of the entire app and feed them into the magic machine? Why, this way, they could have the service for a tenth of the yearly cost, forever!
And then the economy implodes.
I'm not saying it's THE most likely version of things, I'm saying that at a certain level, quantity (or rather, speed) is a quality all its own. And this new quality might change the world. Let's hope it's for the better!
Errors compound, and making 1000 wrong decisions per hour, will not result in something useful. Maybe you‘ve tried setting up guardrails for good design or architecture at some point? I think it’s simply not possible to do that.
It would certainly be an accelerator for people who know exactly what they want. And it would remove multi tasking, which I‘d appreciate.
I don't have a great answer but you pose a great question.
Obviously a CTO is not going to walk away from the technology just because it's not good enough. That much more incentive for someone to create a powerful enough harness that can direct that power safely and productively. Like a nuclear core, we'll need to come up with the graphite rods and water tank. And if tokens are essentially free, why not, for every million tokens, spend 10x tokens on code review, testing, etc?
I do spend 5x more tokens on planning and reviewing, than for implementation. But architecture is still nothing I can delegate.
If your task has incremental rewards/feedback, you can push the "intelligence rate" simply by sampling the reward function faster. That's not fake, even if it not a substitute either.
This is the "dumber but honest person that works harder" phenomenon, vs "lazy genius".
That's a good way to put it, but still my experience is that worse code bases are non-linearly harder to maintain and improve in the future, software tends to break down without a good enough base.
Sure, in the future full rewrites and stuff like that will be just another "throw money at it" problem, but fundamentally software can get arbitrary complex and we barely know how to write large, maintainable code bases.
Nonetheless, I think testing (and maybe proofs) will have its long-awaited time to shine, as being the "reward function".
I totally agree with you on the first bit, but I also think that I am way better at deciding on how to refactor code bases than the LLM is.
Right now, I put models in low thinking mode during my refactors and hate waiting. I would much rather have a faster model that that maybe was slightly stupider, and I would wait far less long between prompts where it needs my valuable input.
Models that are dumb, but humble and fast, can be fine.
AI helps you but also your competition, and gets factored in by investors while customers can use it to find better deals. The whole market is different even if a company did nothing.
Whatever you can cheaply do with AI is not a moat, if there is profit in there there will be quick imitation and competition will eat away those profits.
Models can be replaced easily, harnesses & AI tools too. And if cloud inference gets too expensive there are local models keeping the cloud prices hard capped.
Probably AI won't make anyone very rich.
This is the same pitch that people make about AI today. Speed isn’t the differentiator, quality is
Speed will be one of the killer features once you get closer to instant speeds of 300ms. Just remember what changes were made possible simply by upgrading from ADSL to broadband.
If inference speed goes up, I can launch the same query 5 times, evaluate the best result and proceed from there. Of course, evaluation is also instant, so in seconds I can get a near perfect solution. Or maybe 10 and I can pick what I like the best.
They are both the differentiator.
AI previously provided speed but not quality. As soon as quality reached an acceptable threshold, the speed became the reigning factor.
In my opinion the quality is still much lower, but speed means the cost is significantly lower also.
AI is already fast enough that human is a bottleneck. Hell, typing speed became a bottleneck like it was never before.
I mean, if an agent can do half-decent work in less time than it takes the user to prompt them (and "user" in this context is a fast touch-typist like most programmers are), it's obvious it's not the agent that's the bottleneck anymore.
Because finishing someone else's (or something else's) "half decent work" to the point of "actually decent" becomes the bottleneck.
This has always been the case for human project management, and LLMs just aren't at that level yet.
It's more like everyone is speed running to how fast they can convince others that "half decent" is good enough. And for sure, newer models of LLM seem to be getting better at that.
> It's more like everyone is speed running to how fast they can convince others that "half decent" is good enough.
But that's what Agile is all about, isn't it? We've been speedrunning delivering increasingly smelly shit at increased velocity ever since SaaS became a thing, because ubiquitous Internet access is what allowed our industry to adopt the "lob feces over the fence for users to deal with" release model.
AI does speed that up, true (though since the market - and management - didn't catch up with it yet, we have a brief moment where we can use AI to increase quality while keeping usual delivery rate.)
So for every work produced by AI have ten separate agents review it thoroughly.
I'm not sure inference speed is always the slowest thing for me right now. The agent is running tests, loading webpages, etc, which all take time. I don't know if a fast agent would speed things up in all cases.
That said, it obviously depends on the project.
> "The agent is running tests, loading webpages, etc, which all take time"
A frustrating vision of the future would be when we've been asking for faster loading lighter web pages for years and then companies start caring about it and improving it not for us humans but for LLMs.
It's already kind of that way with MCP servers popping up everywhere. The JIRA MCP server is like a couple orders of magnitude faster to work with than the website itself.
That’s their API with extra steps, or am I missing something? That was always faster.
They finally cared about clear requirements and documentation when that meant getting rid of devs.
That happened at corpo work for each of: * Build times * CI latency * Developer tooling * Documentation * Modularity
I think this reads like Ray Kurzwheil (sorry not able to spell that off top of my head, that bloke who wrote that book about the future) .. But yeah very dystopian and totally realistic. Not if but when..
I LOVE Kurzwheil! Thank you for the compliment, I'm very far from having his writing skills. But yes sci-fi is looking more and more like, well, sci.
> I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless).
you need to launch 10-15 more terminals, who is waiting these days? :)
You sound like my boss! I'm not really into the whole "burnout" thing though.
how can you get burned out just watching the work being done for you?? :)
I tried it. I asked where Bruce Lee was born. It stated he was born in Hong Kong. I challenged it and it went further naming a hospital there. I stated he was born in San Francisco and it apologized and then said his father was a missionary traveling in America, which was also wrong. Bruce’s father was a famous Cantonese Opera singer and actor.
This model had zero information right, while being fast in responding.
Unacceptable.
It gave the correct answers to both questions for me:
> Bruce Lee was born in San Francisco, California, USA on November 27, 1940.
> Bruce Lee's father was a Chinese opera singer
That being said, this is not a good test. It is a language model (a very small one), not an encyclopedia.
ChatJimmy interface is just a tech demo. Without tool calling functionality we can't expect it to be factually correct.
if it's baked into silicon how can you two get different answers?
It still works the same way other LLMs do, by outputting the probability distribution over the possible completions (The weather is ... (sunny (50%), cloudy (50%))). Then the next token is sampled from this probability distribution (in our example the next word could be "sunny" or "cloudy" equally likely), which can result in different outputs every run.
Could the model or algorithm be changed to make it deterministic somehow? It could help a lot if there were reproduceable outputs from deterministic baked-in silicon.
You can make any LLM deterministic by dropping the temperature hyperparameter to zero.
This will generally make them suck, though, a little bit of randomness is necessary for proper function.
You can also use a fixed seed for your prng. A hash of the input text (up to the current turn) should do.
But since it's so fast you can just ask it 100 times where Bruce Lee was born, and statistically you'll get the correct answer. We could call it "mixture of idiots". /s
That's not what speed is useful for.
I just pasted your comment and its whole inheritance chain to it, started my comment, and asked to generate a total of 9 completions, 3 from each of {current & next word, current paragraph, current paragraph + rewrite the entire paragraph}.
Half of the answers were perfectly good (ironically, not the "next word" ones!), but the important bit, they came back near-instantly ("Generated in 0.024s - 14,163 tok/s", the page says). Slightly more powerful model while keeping this under a second, and this could easily become a qualitatively different form of autocomplete/text suggestion. Running in the background every couple keystrokes, or every time user stops typing for more than 500ms.
>That's not what speed is useful for.
>I just pasted your comment and its whole inheritance chain to it,
Good idea. Only problem is it doesn't work. I just did the same thing with exactly this prompt:
>did the user IOT_Apprentice participate in the thread below and if, number and quote all of their comments. Only just number and quote the comments or write "Did not participate", do not add any commentary. Quote any comments by this user verbatim, exactly as input. Thread:
followed by pasting the thread[1]
And received the answer "IOT_Apprentice did not participate in the thread."[2] in 0.001s, even though they have literally the last comment in my quote and it's clearly legible.
It's particularly insidious because the understanding and thinking that is required to follow my requested answer format exactly is substantial - so based on the fact that it gets the format right and clearly understood the assignment, I would be inclined to believe that it would also be correct!
So to use your example, it's not just autocomplete, it's autocomplete that confidently returns "No matching results" in 0.001 seconds, even though there is a search term matching what you put in, right in the prompt itself that was sent to it. That is much worse than useless.
[1] prompt: https://ibb.co/CKVmRvtd
[2] result: https://ibb.co/BKdRKmyD
Using LLMs for information retrieval is the most stupid thing one can do. Especially when old methods work much better.
It's pretty clear if you see what's happening on current phones.
Autocorrect that works. Reply suggestions that almost work, just need to be tad more accurate (probably more of a data access issue than model) and a tad faster to look completely seamless. Screenshots with automated text detection and OCR and automatic interpretation (different suggested actions for when something on the picture looks like a web link, phone number, postal address, e-mail, or QR code, or an event poster). That's just a fraction of things I saw showing up on my Samsung phone over the last 6 months.
For over a year now, you could get a much better autocorrect and spell/grammar check, and a translator all in one, if you just pasted your text to a frontier model and asked it to check for errors or translate into target language. Now imagine being able to go through a round of such checks in a 1/100 of a second. You could have this running every keystroke, and suddenly the inline autocorrect/checks would not suck anymore.
Auto-linkifying that can correct for typos and doesn't need careful regex tuning because it understands from context what is meant to be a link or not. That's just one of many obvious things possible once you get local models running fast enough. Tip of an iceberg, and the first step to imagining all the other potential uses is to let go of the two mistaken beliefs people hold on to:
1. That LLMs are about written language. They're not; ever since "multimodal models" became a thing, tokenization extended to visual and audio space, and now textual and visual and aural inputs are all just regular, first-class tokens.
2. That chatting with the models is the only optimal way for end users to interact with AI. That's just artificially limiting yourself to the space of chat-based UI.
[flagged]
[dead]
The best way I can explain it is that it's the same feeling when I upgraded from 56k dialup to cable broadband.
In the case on on-device/self-hosted LLMs. You ask your agent to implement xyz feature 10 times and use a model to compare the outputs and combine the best results.
Raw intelligence becomes slightly less important when you can iterate and improve automatically. You can still claim it was "one shot" even when 30 different implementations were made then combined.
Problem is, there exists no judge model that will really pick the same winner that you would.
That likely isn't as relevant for on-device iPhone usage as it is for Real Work™. I won't notice the difference between 50tps and 1000tps when asking Siri a question.
I don't know. As others have said, the Taalas chip wasn't small, or particularly low power, so it's hard to "imagine" what that tech in an cell phone chip might look like.
But if the basic premise of "good enough LLM at insane throughput" holds, I think it could qualitatively change local uses of LLMs. At a certain speed point, you're able to move from request -> response to a cascade of tool calling and "subagents", which could allow a small model to be much more useful, if provided with a lot of local data and tool calls.
That said, this is assuming you could stuff a "good enough" model into a phone with Taalas-like technology. The Taalas tech demo was an 8B parameter model and required hundreds of watts (IIRC) to run. The efficiency was good given the speed (as I understand), but it's not clear at all that the approach scales small enough to be a sensible coprocessor on an iPhone or whatever.
> At a certain speed point, you're able to move from request -> response to a cascade of tool calling and "subagents"
That also needs server class hardware though. A phone won’t happily service the insane amount of IO, compute, and network that this cascade would require.
Box that plugs into my desktop would be fine. Or perhaps in SSF form factor.
But wouldn't higher tps allow for more reasoning or other hidden processes, potententially making a smarter model?
This is my thought as well. Models have to be intentional about which tokens they burn because there's a real lag time. If you can just fork out 10 different reasoning sessions at once with no regard for token waste/lag, you can compensate a smaller model with just doing more at once with it. No idea if this is reasonably true though.
I think that only works if you have checkpoints where all that reasoning can be checked against reality. Otherwise you get an army of armchair experts. LLMs are hilariously bad at home improvement advice btw, where reasoning alone won’t get you far.
That order of magnitude could be the difference between "the users wants me to open the notes app, let's open it" and "I've scanned all your notes before you could blink and found what you're looking for".
If Siri is using a 3T model in high reasoning mode to answer your question you will.
Massive economic simulations with thousands if not millions of agents to front run the global economy and stock market.
Fully interactive realtime NPCs in videogames at scale.
Recommender systems that simulate individual consumers.
Crazy shit
About your first example, isn’t the butterfly effect preventing this from being useful? One agent in your simulation decides to sell, and starts an avalanche, that won’t happen in reality?
Run the sim many times... faster than it can run on actual humans and compute a probability density for specific events.
Better yet use it to dimulate counterfactual phenomena like market manipulations ypu intend to enact...
When you run tens of thousands of simulations for complex economic models, you actually do want to see the extreme outliers too. I can't recall who said it, but in finance the interconnected incentives make so-called Black Swan events much more likely and frequent than models or theories can comfortably account for.
In a way... when it's finance, they should be maybe called Gray'ish Swans?
What, even if it means you can run models without relying on the currently backlogged DRAM production?
The size of model we're talking about running doesn't need much if any dram.
The chatjimmy demo is using a model that needs 6-18GB of VRAM. That's not exactly trivial.
I could see it being feasible to get a Qwen-3.6-27b type of model done on something like this. Qwen-3.6-27b at 18tok/s would be a game changer.
right, but that's a reticle size chip. to put something in a phone it has to be ~10-30x smaller
Works great from a press release perspective though.
The weights might fit in cache, if you're using a small model. If you wanted to have a 20B+ parameter model, that's just going in RAM. You could put more RAM in the device and pay the perf cost or have a dedicated chip. Most devices already have a dedicated chip, this just changes which silicon you're spending the money on.
That math doesn't really work.
8B model (FP4) = 4 GB DRAM = 32 Gb DRAM = 80 mm2
8B model (Taalas) = 4 GB ROM = ~800 mm2
Slightly besides your point, but it's interesting how many here naturally ponder about how the current winner could or "should" keep winning, instead of how another company could become a competitor by doing the more clever thing the incumbent isn't thinking about.
It is not a “should”. At least not in the “we wish it were so” sense.
It is more that there are multiple reasons why this idea (burning an LLM into silicone and deploying it into a device in people’s pockets) requires huge piles of cash and the kind of engineering chops only a few company posesses.
Of course i would like it if a small upstart would do this, but it doesn’t seem likely as a posibility. They won’t have the funds to fab the IC. They won’t have the funds to train and validate the model before burning it into silicone. They can’t absorb the risk of the first tape out going wrong. They can’t absorb the risk of the model being faulty in some subtle way. They don’t have a device to integrate the IC into. They won’t have the funds to develop one. If they somehow would make a device they don’t have the marketing and sales channels built out to get the device into people’s hands in sufficient numbers to justify the development cost.
Basically this idea feels ruinously expensive. Apple has deep pockets, they already have working well-regarded phones, and an ethos of privacy preserving innovation. This is why this idea feels well suited for them and not many others.
Do i want the winners to keep winning? No. But not many others can pay for a moonshot crossed with a manhattan project. They just can’t.
Good point. IMHO we're thinking of too much in single entities doing everything.
Right now it's kinda the world we live in, with Apple or Google doing the total vertical integration from chip design to retail shops, but it doesn't have to be that way.
It could be done the traditional way with for instance a joint venture receiving funds and expertise from several players in each of the field and collaborating with external companies to get to the final package.
Yes, it's an interesting register (sorry for the claudism; blame lesswrong-weighted training) for the use of the word 'should'. I agree with your assessment and it is rarely articulated. Sometimes I think that the HN set is abused by big tech both from above on the employer side and the consumer usage side (all the T&C's, VC incentives and M&A taking away once-good-things). So they adopt the only sliver of agency-salving language available, like 'big company that I have no scope over should X'.
It would be cool if the future was a standard fairphone like module system where you could replace the model chip when you felt like it without having to shell out 1-2k $$$s for a new phone
Its wild but if chip is something like $30 and provides frontier intelligence then just throwing them away every 3 months isn't that big of a deal when a lot of us pay $50 to $150 to $1.5k per month on AI tools.
I don't think it needs to be on phone per-se. It can keep chugging in cloud - plenty of people use cheaper older models.
And I suspect the growth will slow eventually making taalas interations slower.
That's actually a really good point... There's currently zero incentive to buying more hardware, and that's one very good reason do have a new one.
But this is already happening with iPhones. Apple is touting on-device AI and only the latest phones offer the full capabilities. Newer phones will be able to run better models, so the incentive is there as soon as someone makes the killer app that only makes sense when the model is running locally on your phone.
> as soon as someone makes the killer app that only makes sense when the model is running locally on your phone.
I expect this to be around the time when we're finally ready to travel to Mars.
From what I remember, these chips are not mobile size yet
A small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local.
> A small model would be [mobile size]
A ~30mm side for the HC1 tech for an 8b model (still unclear the planned HC2)?
Is that analogue or are they baking floating points into the silicon?
It's entirely possible they're using something like block floating point, where most of the hardware is simply fixed point. AMD's NPU does this, for example.
Nope, a small model would be larger than the whole iPhone SoC.
> if you could burn a gemma4 class model into an iphone
... you would still have a mediocre phone with half-assed barely working features driven by locked down proprietary software
[flagged]
Dwarkesh recently pointed out [0] that these guys are almost forced to spend most of their compute on training instead of inference. This is because they need to maintain the appearance (which may also be the truth) that future models will make current models obsolete and be much more valuable.
Completely fixed-function HW can't be used for training, it's inherently a statement that "this model is Good Enough and we are now gonna start just extracting its value instead of extending it". So yeah it's an inference moat but it's not a growth moat.
Makes perfect sense for a company trying to get into the compute business, not companies who wanna be in the creating-ASI business.
Still, I guess/hope they have teams doing it in-house anyway. Just not something they'd wanna make a huge amount of noise about, it doesn't look good for To The Moon valuations.
[0] https://www.dwarkesh.com/p/why-compute-might-get-10x-more-ex...
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
If you’re only running models for frontier capabilities, yeah. For tasks where current models are smart enough, running them 100x faster is the most impactful improvement you can make. Consider all the things you could use a model for, but don’t, because the latency is just a bit too high.
I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
But Claude Opus 4.6 is not really practical. Taalas' process seems targeted for edge models. Their proof of concept model, for example, is a heavily quantized version of Llama 3.1 8B and even then they acknowledge their custom 3-bit/6-bit representation causes model quality degradation.
Taalas is going to have a tough time putting a trillion-parameter model on one conventional die. Their HC1 die is already near the maximum size that conventional lithography can expose. They claim they could partition the model across many chips, but I'm not sure if they have tested this process or what it means for compute. The basic storage arithmetic is unforgiving: for a one trillion parameters model at four bits it will take 50–100 chips. To service a sizable customer base will take thousands of 100-chip fabs.
That all said, I'm bullish on this technology, and look forward to seeing it evolve.
A really fast qwen-3.6-27B type of model could be useful. With a specialized harness and this speed I 'd expect it to find many applications. Implementing a coding plan is the minimum I can think of.
I'm sure life would find a way. I'd love to see what kind of power-harnesses people have to come up with to steer 16k tps QPU's (Qwen Processing Units) productively.
Yeah. But this kinda feels like a bandaid.
Eventually someone will have to solve compute in memory at scale.
With thousands of token per second output it would be an enormous waste of resources. Such chips are clearly made to process thousands of conversations simultaneously. Not necessarily in parallel. All LLM workflows are turn based right now, there are often seconds between turns until tool calls finish or users type the next message.
If the LLM response only takes a few milliseconds, the chip can process hundreds of other requests until the first conversation becomes active again.
> it would be an enormous waste of resources
Sounds a lot like "640Kb ought to be enough for anybody"
With those speeds I can benchmark a batch of different approaches, compare the results and serve the results all within a second. It's quite amazing.
It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.
I'm expecting the Taalas MSIC version to cost a fraction of that. Then probably have some kind of cheap subscription to Anthropic for updates (yes, Taalas chips can receive a certain kind of updates: they have a small SRAM).
> It costs something like $300,000 for the hardware to run a model of that size
You did not compute that as the cost for a speculative card from Taalas, right?
It's the cost of the current nvidia hardware used to run these models. Of course all bets are off if you are accounting for some future chip that doesn't exist yet which could cost less.
Not so long ago, I was good enough for many coding tasks. But I found that things can change in a hurry.
Yes, a cheap and fast Opus4.6 can drive a lot of value in current context. But if we continue to craft bigger-and-bigger balls of mud, Opus 4.6 may end up hitting its conceptual ceiling and unable to contribute.
Winding the clock back on your statement gives:
> I'd gladly pay for a Claude Sonnet 3.5 in silicon and use it for 1-2 years.
Man, I dunno.
Assuming moore's law like progress, which I'm 100% sure isn't going to happen - I think we're at the top of the S curve already. But assuming dramatically increased intelligence every year this is still the exact same position as anyone who bought a computer in the last 5 decades. Yet, people did very much buy computers.
But that's what people said 6 months ago about whichever model was current 6 months ago, but you hate that model now.
Depends on how much it costs the consumer. If I could buy a "cartridge" of Kimi K3 for 300 bucks I 100% would buy that shit asap. Even if it's "no good" after lets say 4 months still would be worth it IMO.
That's definitely super-enthousiast territory. Paying 80 bucks a month for AI is more than 99.99% of people would be willing to do
Do think about b2b. Companies are already paying much more for AI. a new K3 (or similar model) every 6 months for a monthly rate of ~100$ per month is something MANY businesses would pay for. Then they could even sell them at half the price to consumers.
This will be considered very cheap within the year IMO. The value you get from AI is exponentially increasing and like all tech just takes some time to ramp up. Cell phones, internet and many other amenities when they came out many people were not willing to pay for but that all changed and considering how important AI tech is this will also be the case especially considering if its 100% private such as for that cartridge.
> The value you get from AI is exponentially increasing.
Perhaps in some cases, but the value I personally and professionally got out of LLMs reached a limit a while ago and has since kind of fluctuated between that limit and a bit less.
If the best model was instant, like the demo here, it could certainly provide more value, I guess, but I think the limit I'd quickly hit is the same one as now, which is how much of it do I want to produce, for what reasons?
That's because the super-enthusiast will upgrade in 4 months when a better model is released. The casual user would keep it for years. A year of claude at the lowest plan is almost $300
"seems like baking models into silicon is speed-running obsolescence"
Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.
Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)
Yes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now.
One of the underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware: it is clear that there is a real application for far greater density and lower power demand, and people are now pursuing this much harder than they had been. That will yield what it has always yielded; orders of magnitude jumps in capacity and performance.
Yes but have we considered employing, like, a really big block of ice? Like old-timey surgeries? What if we put a big block of ice on the 2.5 cubic meter CPU what happens then?
Phones were getting too thin anyways.
Or autonomous weapon systems, missiles, and drones.
Why would they need multi TB frontier models?
For pondering trolley problems, maybe?
I could see this making sense when model development start to settle down ... it's going to settle down, right? ...
Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.
Surely that added flexibility negatively impacts the density/parameter count of the model you could etch?
I think they already do that, except it's not 1980 so you don't fix the upper mask, you fix the lowest metal layer (the upper layer is very coarse and is only useful for power). But even a single mask is still quite expensive.
But I suppose the interconnect masks don't have the resolution requirements of the masks for transistors. Therefore it could be a lot cheaper.
(Yes, you could fix a number of masks, e.g. entire logic gates, of course).
Or do a hybrid
Compute the cost of producing n of them devices, imagine a fair price based on that, and see if that local, blazing fast card* can be an asset that could be replaced periodically.
*(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.)
It depends on how quickly you can bake new architectures.
Text diffusion might be a disruptor here, but let me just say the most cutting edhe form of image diffusion (JiT and DiT) right now is just a big fat stack of alternating attention and MLP matmulls. Not theoretically hard to bake
Which is exactly what companies and shareholders want to increase sales.
Look at it the other way: compared to the cost of training a model, the cost of making a custom ASIC is trivial.
obsolescence is the whole point. apple gets to sell a new phone very 6-12 months because of it.
i have written about this:
"For device makers
Packaging models with laptops and smartphones will let application access near free, low latency inference and potentially offer users a better experience with the option of preserving data on-device. This is viable under the condition that tasks that do require larger expert models that run in the cloud can be routed to external models. A side-effect of local models and what will let Apple cut upgrade cycles from ~4 years (?) down to 12-18 months is specialized hardware to run them. For almost a decade, smartphones have been trying to compete on better cameras. This coming decade will see them selling better GPUs, NPUs, ASICs and whatever other things they'll be calling the inference chips, to drive re-purchase. Every six months will see a better model on new hardware, which will enable better performance in certain applications."
https://try.works/role-model-the-case-for-a-model-routing-pr...
No, the point is inference speed and power.
you don't understand what I wrote.
I do. The point is inference speed and power, making previously impossible local inference possible. A side effect of that hardware optimization is fixed capabilities.
You've confused engineering compromise for malice, and reversed the purpose. For the model capabilities and inference power draw, what alternative do you see to a (at least mostly) fixed hardware model?
What I'm saying is that Apple will use these type of models etched into chips, and they will do it because it drives obsolescence, so they can shorten the upgrade cycle. They will do it because they figure out it's good for them.
You've confused engineering compromise for malice, and reversed the purpose. For the model capabilities and inference power draw, what alternative do you see to a (at least mostly) fixed hardware model?
I wish we lived in a reality where Framework was anywhere near rich enough to acquire them. I'd love to have models on a chip that I could swap in at a whim. That would do the opposite by eliminating moats.
I guess there's a tiny chance AMD makes something like that happen. It seems like a great way to get people and orgs to pay a few hundred bucks every 6 months or so.
ASICs is what took over Bitcoin mining, cheaper in all ways, and lasts longer than Nvidia GPUs for inference.
> cheaper in all ways,
Bitcoin mining doesn't have large memory requirements, but does have huge compute requirements. ASICs work great there because it's very straightforward to add some circuits for computing hashes. If you _also_ have to add many GB of memory, then suddenly ASICs will cost as much or more than comparable off-the-shelf hardware and they won't be faster unless you've also invested in huge memory bandwidth.
My understanding is an ASIC can last 10+ years, where are Nvidia enterprise GPUs are rated for 5...
Most enterprise GPUs are scrap after 5 years because they're so inefficient compared to newer models. It's entirely possible to make them last longer by undervolting them, people just don't because it doesn't make sense.
Bitcoin OTOH has used the same PoW algorithm for a decade. Barring some really exciting discoveries about the nature of computation, new ASICs are not that much more efficient than old ones.
BTC mining is also not exactly competitive anymore; the nature of the PoW algorithm means that it's dominated by a few large players who've set up shop next to a dam and who pay very little for electricity.
New entrants are highly discouraged because the mining rewards are constantly halving, it's hard to find cheap power, and the price of BTC is now so volatile that a yearslong investment is very likely to lose money.
You need to find customers for several-generations-ago models before this makes any sense. AMD is a lot more incentivized to look than mr vanilla llm is
They are: https://openai.com/index/cerebras-partnership/ My guess is they only consider Luna "good enough" to justify the immense up-front investment to put it onto silicon, but Luna at 10x the current speed would be killer. If they're really pursuing live voice conversations with a hardware assistant, latency is more important than accuracy (for complex questions the assistant could always say something like "wait a minute, I need to think about this" and hand over to another model).
Cerebras doesn't etch the model onto silicon though. They're basically just wafer scale GPUs. They're more flexible than etched silicon though because they can just run the next version of the model almost straight away.
Interesting. I did not know that. And still they got 15x speedup (https://www.cerebras.ai/blog/openai-gpt-oss-120b-runs-fastes...). I wonder how much additional speedup would be possible by really etching the model.
How ? Do LLMs actually "know' when they don't "know" ?
They do this all the time, I'm using ChatGPT in Instant mode and it auto updates to thinking if my question is complex. Most of the time this works.
To answer your question: A large language model itself does not know this (afaik). But chatbots are not "just LLMs" but a whole bunch of systems (and models) around them.
Ok, but the article is about etching the model, not a "whole bunch of systems". So far, I still don't if its actually doable or if it is just unsubstantiated speculation.
How do humans?
Always the same trick of not answering the question and deflecting to „what about humans“. Can you folks not evaluate LLMs as the system they are, without vague gestures at how a different system behaves?
Evaluating LLMs is incredibly difficult. They are categorically different from any other system we have intuition about.
That said, I read the question I am replying to as a rhetorical one. If it was meant as a genuine question, curious about the question of meta knowledge, then I misread. Certainly the question is extremely interesting, for both LLMs and humans! But it's also obviously a very difficult one, as we don't even have a clear theory on how "knowing" works in the base case.
Their thing is improving the models; it would be extremely counter-company-culture to bet on models plateau-ing. Maybe wise in terms of hedging, but still difficult to pull of as a company decision.
They are busy capturing market first, and I think that makes more sense. 'premature optimization etc.'
Isn’t that kind of useless for the stock? It sounds complicated, unlike having number of CPUs go up.
It’s like talking about anything else than Megapixels when everyone was convinced that megapixels must go up in certain periods of the smartphone boom.
Chinese already start making DUV which can do the lower end 7nm. They are winning. Once that 7nm and up market cornered by Chinese, AMD Intel and TSMC and Samsung will have to burn thru bleeding edge depreciation faster perhaps from 7yr down to just 18mths. The CPU they generated will be incredibly expensive. Meanwhile Chinese just keep minting the AI cheaply and more efficiently and inching upwards towards 1.4nm.
They can do 7nm. But they use multiple patterning(printing the same pattern multiple times to get to 7nm), which is expensive. So it's not comparable on cost to western single-patterning 7nm, and of course not to the leading edge on cost/size/power.
I’m surprised Nvidia hasn’t partnered to make a Claude chip yet. It’s a win/win you can license them out, sell them when they become obsolete, etc.
Apparently Anthropic is moving that way: https://arstechnica.com/ai/2026/08/anthropic-confirms-plans-...
Not necessarily: it is relevant to Taalas only if it is a compute-in-memory architecture.
The Jalapeño mentioned («Anthropic is not alone in walking this path») in the article is still a classical Von Neumann architecture.
And Taalas' idea makes sense in a perspective of scale - producing a large number of cards; "for internal use" (a lower order of items) means a high production cost.
This seems like a very bad and dangerous direction for our society.
We don’t really known (from the outside) how exactly do they move. Besides it may have not been truly viable 1-2 years ago…
I guess I'm not understanding why this makes sense for AMD to buy Taalas unless they plan to get into hosting. It doesn't seem like a great fit.
Just to see how fast it is try chatjimmy.ai
It is really fast and ... really hallucinates. I asked "Does the Wang corporation still exist? If not, what happened to it?" and it replied (in part):
"Yes, the Wang Corporation, the company that originally developed and marketed the Wang 2200 computer, still exists as a rebranded company under the name PPL (Precision Pencil and Label), but it has undergone significant changes and challenges over the years.
Here's a brief overview of what happened:
In fact, Wang labs was founded in 1951. PPL seems to be a made up entity. But it did generate those "facts" in 0.033 seconds. If people value speed over accuracy then I can write an LLM that is 100x faster than chatjimmy.ai and make big bucks by responding one of N canned responses to any question.to be fair, it's running an 8b model from like 2 years ago. Taalas just does the chip design, not the model architecture.
their tech is a mere demo to open up a new path, the day we can have some asics running a Qwen3.6 27b, this would open up new doors
Pretty incredible to see. It reminds me of when I first used the Groq chatbot, except in this case it's a full response instantly.
They were busy buying open source frameworks and teams behind those.
Didn’t Anthropic acquire Cerebras? Seems like a move into the same direction.
I also think that etching models into ASICs may be a bit too inflexible for what OpenAI and Anthropic want.
No, that's backwards. OpenAI are the ones investing in Cerebras. Part of the deal is that they can't sell to Anthropic.
Because Openai and anthropic are not hardware companies. They outsource that to Broadcom and AWS' Annapurna labs.
OpenAI and Anthropic are both designing ASICs.
So they have decided that putting a small LLM on a phone would backfire because people would have a negative perception of their cloud models. Pretty sure AMD will use these taalas chips in data centers, not phones
>Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Google is no longer a serious player in frontier AI. I doubt they will ever hit a SOTA model again.
A model can't be updated, and a chip that is only relevant for 6 months at max?
Depends what you mean by relevant. If you use AI primarily as a search/knowledge engine, it makes no sense. If it's your capable assistant that has a lot of general knowledge, can do tool calls, and has a big context window, very doable.
Indeed, for some kinds of applications involving secure/legal data etc. I can see the consistency of silicon winning out, because it combines performance with immutability and guardrails in hardware. Some chips have write-once PROMs to store password hashes and similar, you could do the same thing with prompt hashing to absolutely force or forbid certain behaviors. A model that can't be updated is also a model that can't be hacked.
People already buy new phones every year, this just creates even more reason to do so
Outside of this website I've never met a person who buys a new phone every year. It's closer to every 3-4 years for most people.
Closer to every 5-6 years these days and with ram prices going up it will be even longer. Especially with the low/mid range phones, which are most phones outside some developed countries, people will keep their phones as long as they can.
Would depend on the income levels, but yeah, buying a new phone these days is entirely a non essential luxury. An iphone easily lasts 7 years so the moment money is tight, it's a very easy choice to not buy a new one.
I mean are there even any reasons to buy a new phone?
If I compare the Pixel 6 Pro I'm using at the moment to current models, they are functionally identical. The only reason to upgrade might be getting a fresh battery and access to firmware updates.
Otherwise I'd be happy to continue using it for the next 10 years.
I only buy phones if the current one starts showing signs of deterioration, mostly battery.
I swear a midrange Chinese phone from 2017 would be enough for me in 2026 to read HN/Whatsapp and some Youtube.
Your location/income bias is showing. Most people do not buy new phones every year.
I live in the bay area and buy a phone maybe every 3 years? Why do people waste so much money :D
I never said most people. But it's not uncommon. I personally find it ridiculous, and hold onto my phone until it's unusable, but plenty of people in middle class Australia seem convinced that they need the new one whenever it comes out.
Base model sure, but the stack will be hybrid. It’s still early days here. Too bad FPGAs have such large feature size.
One of these chips smart enough to take orders at a drive-thru would be relevant for a decade, minimum.
It's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further.
Not so if it's embedded in something smart enough for its intended purpose.
Think vision, spatial reasoning, speech synthesis, even some speech analysis. Think self-driving cars (and drones) that need 10x less power for the brain, and can think at 10x situation per second.
This is only true for people who are solely focused on performance. There is absolutely a market for acceptable performance combined with predictability.
True, but predictability cuts both ways.
We're all used to having to constantly update our browsers and phones to keep up with the security arms race. If a frozen model can't be updated, it will predictably remain vulnerable to any "exploits" or idiosyncratic quirks that people discover over time.
Let's say, as somebody suggested in another comment, that you buy 100,000 of these chips and deploy them to run fast-food drive-thrus. And then somebody discovers the model has a fondness for goblins[1], and if you role-play convincingly enough, you can get it to accept payment in shiny buttons and rodent skulls instead of cash.
What do you do then? I guess your options are to try and fix the behavior with a better prompt, or put some kind of filter in front of the model to catch attempted exploits. If the filter is cheap and dumb it probably won't work well enough, and if you use another model as a filter, you've negated the cost and speed benefits of putting the first model in hardware.
Of course the real answer is to just never expose the model to situations where an adversarial input could possibly lead to an undesired output. But that drastically limits what you can do with it.
[1]: https://openai.com/index/where-the-goblins-came-from/
I see your argument but your example seems highly contrived. I can't think why you'd want to use something like this for something as dynamic as takeout ordering, where you might have to deal with bad customers, supply chain breakages, public health recalls, or any of many other probabilistic events.
I think it's far more likely to see them used in safety critical applications where you need a capable model that can run on low power and doesn't have multiple layers of operating abstractions between the model and the hardware.
What safety critical applications would be a good fit for LLMs?
I'm not thinking of language models specifically, but large neural networks in silico. I feel like a 27B parameter model would likely be capable of flying and landing an airliner, for example.
> Of course the real answer is to just never expose the model to situations where an adversarial input could possibly lead to an undesired output. But that drastically limits what you can do with it.
Does it though? Isn't that what CPUs are, very fast-not-so-clever computing brain surrounded by layers that protect it?
If a model is good enough today, it's still gonna be good enough in a year. Except you'll be able to serve it 1/100 of the price. Or 100x the speed.
I think we will eventually reach a point where this is the case, but at the moment it seems like you can throw virtually any non-trivial use-case at a model today and end up being more satisfied with the results that a model tomorrow gives.
I may just be closed minded as to what use-cases we have that current models are truly "good enough" (i.e. won't be dissatisfied when comparing results of today's model to tomorrow's model)
Might be worse because it takes time between designing the silicon and having the first usable chips. So they're outdated the moment they hit the market or even before that.
OTOH, people get a new iPhone every year and they are ok with it.
How is that in any way related to a consumer device? This method doesn't reduce physical memory requirements, so still results in huge die area. This isn't a for-end-user thing, probably for decades.
Ok, how long until nvidia gives us a new GPU?
I don't follow. How is that related? GPUs don't have fixed memory. You don't throw them away when you want to load a new model.
NVIDIA will probably give us a new GPU when someone competent in the free market decides they want wheelbarrows full of money. Unfortunately, AMD is entirely, incomprehensibly, incompetent, to the point where I can only assume they're colluding with Nvidia, behind the scenes.
It googles models suck