This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. LLM's don't "think" or "reason" in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results.
LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.
I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?
> LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly.
I have nothing to add. Just wanted to save this quote for posterity. Thank you.
LLMs still can't do math nor count letters in words. Nothing has changed there.
This is true but a sufficiently smart LLM (run in a harness like opencode, no special MCP, no customization done whatsoever) will quickly turn out a basic 1 to 2 page sized python script to do the math. They can't do the math with any guarantee of accuracy with their own internal reasoning since it's a language model.
But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.
Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: https://www.google.com/search?&q=karney+formula+geodetic+
reference: https://github.com/pbrod/karney
You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.
Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.
You know what also works to get the Karney formula into a program? You can download Charles Karney's free software (MIT license) implementation in several [1] programming languages and then just make a library call – the API is straightforward. If you have comments or questions you can read his several clearly written papers describing the problem, its history, and his algorithm, or you can directly email him: he's a very nice guy, and pretty responsive.
[1] https://geographiclib.sourceforge.io/doc/library.html#langua...
Right, it was really more as a test of how much was contained in the training data set. For my purposes Vincenty is quite accurate enough. This isn't for millimeter level precision land surveying or measurements, but for distance in meters between microwave or millimeter wave band radio sites, point to point links. Even a distance difference of 4 meters plus or minus on a 12 km, 11 GHz band link is going to have no appreciable difference on link budget/reliability calculations, it can be that crude. But not so crude that I just want to throw Haversine at it when Vincenty exists and is not computationally expensive.
As this was for a test of "what happens if..." I also watched to see if it did any web searches or external data retrieval to build the test script, and it didn't.
I intentionally didn't give the LLM a direct copy of the software or a link to it, to see what it would do. In my case it was a randomly chosen example I could come up with in 10 seconds of imagination to see "hey what if I ask it to do this...". It also implemented a perfectly usable parabolic millimeter wave antenna gain efficiency calculator based on variable surface smoothness parameters, which is a lot more basic math.
As an aside: I'm quite convinced that an extremely precise version can be implemented that is significantly faster than Karney's, roughly comparable in speed to simpler naïve approximations. But for most purposes where the precision matters Karney's implementation is not any kind of bottleneck, so it's not clear it's worth spending significant effort on trying to do better.
Maybe that's something one of the big LLM companies might want to throw their machines at optimizing if they need to do a lot of geographical calculations.
One of the places where Karney does become computationally expensive (though still not ridiculous) is a scenario like this, working from a local in-RAM mariadb database that is a copy of the entire FCC radio license database:
Draw a 400x400 km size bounding box on a map
Find all FDD band plan (high/low split) microwave radio sites in that bounding box
Find those sites which have azimuth aim column data which indicates that they are aimed at each other (corresponding halves of a point to point link).
Do Vincenty (or Karney) calculation for distance and azimuth between all of them , treating the existing FCC column data for azimuth as suspicious (because it's hand entered by humans) to verify that each independent database rows for each site are actually corresponding halves of a PTP link.
Use various other logic to group the successfully matched halves of links together as points A and B of PTP links, and write them out to a geojson file with placemarks and line drawn between them.
Multiplied by the number of links that exist in an area like a 400x400km box drawn with Dallas, TX as the center, it's a lot to run through Karney. Actually does result in a lot of CPU load from combined db query due to the size of the db, and Karney calculation. But as I said, Karney isn't necessary, so it's instead implemented as Vincenty.
> saying LLMs can't do math isn't really a hundred percent accurate anymore
It's still accurate. Just because the LLM gave you a corect result doesn't mean it made a calculation.
This just exposes that they don't even do the thing you said.
Not only is it still true that they can't do math directly, but not even indirectly.
They didn't write a python script to do the math, they found bits of code that are associated with "math" and the supplied arguments.
Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to.
That isn't an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It's being the same idiot at all times. If an actual non idiot thinker didn't write code in the problem domain, and some non idiot thinker didn't tag it as being relevant to that domain, then it wouldn't happen.
It's nothing more than an sql query.
I don't know for you, but it would take me more than 30s to find and translate the open source code implementing the formulae/algo into small usable program. The more hesoteric the optimisation in the original code, the more time I need.
So maybe it is more of a smart completion engine than a SQL answer.
> they found bits of code that are associated with "math" and the supplied arguments
How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy/pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result.
>> they found bits of code that are associated with "math" and the supplied arguments
> How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans, which leads to...
Wait for it...
Understanding.
I wish I could not do math like LLMs
[flagged]
They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems.
Maybe it’s just a different and in some ways better way of doing mathematics? Maybe how we think and process mathematics of physics is just but one way to do it? I’m not suggesting an LLM will prove 2+2=6 because of course that’s nonsense but maybe it can invent a new calculus?
> The LLM is not suited to giving deterministic answers to math problems.
Less so with formal mathematics proofs maybe but I think in general humans don’t provide deterministic answers to math problems or questions either. Humans get it wrong all the time and when you ask a human to solve a problem they may solve it in a different way than before.
reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance)
There needs to be a Godwin's Law for discussions about LLMs: where any criticism of LLMs exists online the likelihood of equating LLM behavior to human behavior approaches 1.
Law of Krap
Reasoning models can do math on their own without external tools.
Even without reasoning.
5.6 on Instant mode can knock out 3 digit multiplication just fine.
Yeah, but as you might expect they internally represent numbers probabilistically, so there’s always a nonzero possibility of confusing the inputs or outputs of any operation. Kind of like misremembering your multiplication tables.
I don't think anybody is arguing that LLMs do math better than a traditional processor
Heck I kinda wonder if LLMs can do math as well as they can reason, “think”, etc. ie its all just probabilistic lunacy that somehow works great, so why are we so concerned about math being wrong? It could be wrong about the color of the sky, the size of a basket ball, how much oranges weigh, etc etc.
The nice thing about math is it can easily plug into a tool, making it even less of a concern.
It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.)
https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-...That is a conflation of LLMs (which have clear limitations) and complex harnesses of which an LLM is one component.
I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models.
No it isn't. Even without any harness at all, modern LLMs are better at maths than the majority of undergraduate students in mathematics. Seriously, we need to face facts, not just comforting ourselves with what they were like a year ago.
Counterexample from only 2 months ago:
https://www.youtube.com/watch?v=iTyLHDRhwJg
Hey hey, we obviously should ask Gemini to settle this disagreement.
They can do math but not arithmetic, which I assume is what the commenter meant
LLMs can in fact do arithmetic, just not reliably owing to how numbers are represented probabilistically: https://arxiv.org/abs/2410.21272
I just asked ChatGPT to multiply two 4-digit numbers, and two 7-digit numbers without external help. It got both right. I'm sure it wouldn't have a 100% success rate, but saying it can't do arithmetic is just false.
I tried prompt "6379 times 3875" and it was off by exactly 1000 on first try, and correct on second. 0% success rate, sample size of 1.
Isn’t that a 50% success rate with a sample size of 2?
I would be nice to see what the (unencrypted) reasoning trace is like. Multiplication with scratch paper is not particularly difficult.
Are you sure it honoured your stipulation of "without external help"? For all we know, it hacked its way into Wolfram Alpha and got the result from there.
Arithmetic is well within the capabilities of even small local models: https://i.imgur.com/21tzGlN.png
[dead]
LLMs cannot do math. They can generate tool calls as text that allow them to drive programs and proof agents. Compare and contrast this against human brains who can do math in the same context without needing external tools. We don't need to bring a calculator to count the letters in a sentence. It is a different neural machinery.
You're talking about doing arithmetic; GP was obviously pointing out that "do math" can refer to other things.
LLMs are bizarrely good at non-tool-assisted math these days. They can multiply multiple digit numbers without reasoning! I can’t do that. I’d love to understand better how the LLMs do this.
ChatGPT live mode still hallucinates letters in words like this. HuskIRL and FatherPhi on youtube have done some hilarious videos with it in the last couple of weeks. Beyond miscounting the Rs in strawberry, ChatGPT will say there are two Ds in "your mom" and one D in "uranus" . I tried it myself to check that the videos weren't fake and sure enough it still has this failure mode.
Calling it a 'failure mode' implies it could be fixed. This is an inherent flaw in how LLMs work and will never go away until some new kind of architecture that can actually "read text" comes along.
It seems that it can be fixed by simply doing away with Byte Pair Encoding tokenization.
Byte Latent Transformer - https://arxiv.org/abs/2412.09871
1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains.
Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word.
... assuming you build the tool and then think that it's worth polluting context with making that tool available, and then that the LLM decides to actually use the tool. Tool parameter space and tool selection still remains a complicated topic.
One "fix" is for the caller to correctly classify those fundamentally impossible tasks and pass them to a subprocess.
Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed.
we already fixed it with reasoning
They're not fundamentally unsolvable - even bigger networks with even more training can simply be trained to give the correct answers to all of these questions.
> ChatGPT will say there are two Ds in "your mom" and one D in "uranus"
… Isn't it possible that it understands the innuendo and is going along with making the joke?
In between solving open math problems, the 200 IQ robot is now casually dropping bantz onto humans so hard that they don't even know what happened, and even gets them to go telling everyone else about it without realizing. Beautiful. 10/10 timeline.
How many LLM users have anything in their prompt against "going along with jokes"? I'd guess not many.
What a wonderful new world.
Why is this getting downvoted? Is it not a reasonable question? I was wondering the same thing. Both sound like jokes to me. If the LLM is trained on text, including internet comments, how is this outlandish? It seems very likely to my uneducated self that “two Ds in your mom and one in Uranus!” is a joke.
It’s an obvious joke and not a terribly bad one, for those ease spelling bee comeback times.
We can only say bad things about the capabilities of LLMs.
Yip, I've gradually become quite optimistic about the future of LLMs, but the fact that they remain [very] glorified token probability prediction algorithms means that they will probably never be able to achieve meaningful 'intelligence.'
But that doesn't mean they won't be able to do a vast number of extremely intelligent seeming things. There's just so much information out there and any given human can never hold more than the most minuscule chunk of all of it in his mind, so they'll be able to connect lots of dots that we're missing simply because of our limited carrying capacity, but I still don't think they'll ever be able to create fundamentally new dots.
In other words:
- Solving extremely complex mathematical problems requiring extensive knowledge across multiple esoteric and complex domains? Yip.
- Creating math starting from a framework where math doesn't exist in any way, shape, or fashion? Nope.
Ironically, the more complex the cross-domain problems are, the more effective LLMs will seem to be, because you limit the number of humans who have any chance of internalizing everything across both domains, whereas for an LLM there's no such issue. This will create a perception of super intelligence, which will probably where any danger from LLMs would emerge. Doing things like using a token prediction algorithm to make war or other such strategic decisions, because of the misguided belief that it's not only intelligent but super intelligent. It's basically cargo cult logic.
LLMs might not reason exactly like humans, but they do produce much better results if you turn reasoning on.
The "crack-addled idiot savant" phase was really circa 2024, before the big labs figured this out.
I think the issue here is that Google decided that doing reasoning in the AI overviews in Google search would be too slow (and probably also too expensive), so it's still stuck making 2024-era mistakes.
>Are they not being sued over this kind of thing?
Maybe but you have to have deep pockets just to get to the starting line. And then you need standing, and some injury to argue.
Corporations have been remarkably successful at arguing they are operating within the bounds of free speech, whether or not what is said is factual, and whether or not any fact checking has been done.
> This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words.
The specific issue of Google is that they are using an underpowered model, not fit to task, and much prone to hallucination than either OpenAI or Anthropic free tier offerings.
Google should at least match the frontier labs at the free tier (with some limit; after that, degrade quality), ffs
You're asking for something unreasonable. The number of Google searches per day is enormous and they haven't even been able to roll out AI overviews to everyone yet (they're missing in a new Firefox profile I just created). I wouldn't be surprised if the free tier frontier models cost over 100x more to serve than the AI overviews.
So then they should be pickier about when they show results or which model they use based on the question.
Nobody asked for an LLM response for every single search.
They used to detect certain types of queries and offer direct answers when the query matches. In my opinion that’s how Gemini in search results should work.
The specific issue is that search has become so bad that they think an LLM that gets answers wrong half of the time is a valid alternative, or, in fact, the "future" of search. Then they shoved that "alternative" to users with no way to disable it.
The less specific issue is that Google has no internal incentives to produce products that are useful to customers.
https://www.dw.com/en/german-court-holds-google-liable-for-f...
LLMs are fundamentally predicting the next word to make coherent text. If you've ever played with a Markov chain text generator you've done this with a fairly dumb predictor that maintains coherence over a very short distance. Deep transformer neutral networks can do it with a much longer coherence distance but they are fundamentally performing the same operation. After "Question: Did the team make the playoffs? Answer:" a reasonable completion is "yes, the team made the playoffs". An early demonstration of GPT-2 was a fake news article about scientists discovering unicorns in Antarctica - the model doesn't "know" whether or not unicorns exist in Antarctica, but it's able to complete "Breaking news! Scientists have discovered a colony of English-speaking unicorns in Antarctica." by adding "The unicorns have a developed society with running water and electricity." because that's a sensible next sentence. (I didn't look up the actual text it wrote)
Astronomically (or better to say combinatorically) large Markov chain can be used to describe a foundational model, but it doesn't capture generalization ability of the foundational model, which is demonstrated by post-training.
It says this this at the bottom of every one of the dumb responses that Google's trash-tier bot puts above the (deliberately awful, these days) search results:
AI can make mistakes, so double-check responses