> Delivering the benefits of scientific progress and economic growth that very intelligent machines enable.
I think we're very close to the point where AI-driven breakthroughs outside of pure math and software start to really affect the world.
We evaluated GPT-6 Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks.
It's the new frontier model by a landslide. It's even more dominant than the Fable 5 release, because not only does it wipe the floor with the second best model (Fable 5.1), it was also ~80% cheaper and 30% faster in agentic coding[1].
Astra is a groundbreaking model. The biggest breakthrough since Opus 4.5, maybe even since GPT 4. It broke AAII, which is hitting the limits of what most popular benchmarks can measure -- it's definitely fair to call it AGI.
Data at https://gertlabs.com/rankings
(1) Note that we used the "OpenAI Flex" endpoint on openrouter, which is half the price and didn't cause any delays in our testing (this is different from the batch endpoint)
>I think we're very close to the point where AI-driven breakthroughs outside of pure math and software start to really affect the world.
What are you basing this on? What specific breakthroughs have convinced you of this trajectory?
The results we've been seeing internally on our physics and circuit design environments are expert-level and beyond-expert-level results from models that Astra completely outclasses across the board on our evaluation suite (Fable 5+/Opus 5/Grok 4.6 were all worthy of being called AGI in my opinion). That's hard tech that will translate to real product innovation.
But you don't need any kind of insider information to see how fast the world is changing. ChatGPT launched less than 4 years ago and the advances in robotics, unsolved maths, and software are all riding the steepest exponential improvement curve any of us have seen. Interesting times we live in.
I mean honestly, that's the problem. I'm actually not seeing the world changing. What specific advances in robotics, unsolved maths, and software have LLMs provided? What is the finished result that affects everyday life? In all categories, it's been hype with little actual real results. The robots are still doing the things they did before 2022. The maths are a handful of fairly insignificant proofs that have no significant applications. Software seems to be buggier than ever, but that aside, we certainly aren't seeing a lot of new innovative applications. We're using the same applications as ever. The same operating systems. They've all changed very little.
I'm not trying to be a pain here, but I keep seeing people saying "look at the massive change all around us" and back here in reality, there is none. Give me concrete, real world examples. Name software. Name products. Name the breakthroughs specifically. This should be easy.
There are several hundred thousand mathematicians producing hundreds of thousands of new results in math each year. Why can't most people name any human contributions to mathematics from the past decade? What is the finished result that affects everyday life? Do you consider all those human mathematicians to be useless?
Most of the biggest breakthroughs in mathematics, breakthroughs that win Fields Medals like sphere packing in dimensions 8 and 24, have no applications in everyday life. Probably the only new mathematics results that people notice affecting their daily lives are the ones that enabled AI.
Nevermind mathematicians. What about the millions of programmers? Are they all hype too because people pre-2023 were griping on hn that software is buggier than ever and people are still using the same operating systems as always? Why couldn't the 30 million human programmers make something better in the past decade?
You set your bar so high that all the world's human experts in math and programming combined would fail to meet it.
The top LLMs in 2024 were Sonnet 3.5 and GPT 4o. You couldn't have expected those much weaker models to be making breakthroughs in math. The models that are making breakthroughs haven't been around very long.
They asked for a specific example, you've still not provided one, please provide the example.
Why?
1. I didn't say LLMs have made any breakthroughs in math, not because they haven't, but because it's irrelevant to my point. The parent comment is using the same argument academic research opponents have long used against research. The vast majority of research fails to meet their bar. How would your daily life be different if we had no humanities papers published since 2023? Or math?
2. You can google this in 10 seconds and see a dozen results in math. This is not a good-faith demand.
Or you can google and find out that those breakthroughs were not as revolutionary as they are presented.
For the last 60 years, whenever AI achieves something revolutionary, some people immediately say "well that wasn't particulary revolutionary".
Every single time.
It's a tired argument, and we should strive for the intellectual humility to do better in this forum.
Maybe AI isn't that incredible, the more you use it, the more you realize it's a tool, like a VCR, maybe that's why?
What was sold as AI was basically a "computer person". Maybe that isn't the reality so when people are like, "here's the self coding machine" everyone is a bit disappointed because it's not C3PO?
We're not going to suddenly have new robots or operating systems. Those things take time. Just because there's massive change afoot, doesn't mean it's widely adopted or applied at lower-level products. ChatGPT is a product, and software, and a breakthrough. You can have a freeform conversation with your computer about anything, in human language, and ask it to do or make stuff and it will at least try, sometimes with surprising results. That wasn't possible until recently. The robots and products are coming, rest assured.
The number of bug fixes to important programs has skyrocketed. Look at what Google are saying about how many Chrome security bugs they've been fixing lately. Other big software firms have been doing the same thing - AI has been finding and fixing a ton of bugs. I know of one big program where thousands and thousands of security bugs are being found and fixed.
It may not feel like this to you because a lot of the dollars right now are going into security bugs which you can't perceive. But it's definitely happening.
At the company I own I've got AI employees autonomously triaging backlogs and fixing long tail bugs. The software is definitely getting better, although by definition long tail bugs aren't ones you are likely to encounter. The subjective "feel" of how robust the software is won't change quickly.
New tech is always applied in apparently boring ways because we are imagination constrained and people harvest the low hanging fruits first. Remember claims there was worldwide demand for only about four computers? When Gates said he wanted a computer on every desk and in every home people laughed at him. What would people do with all those computers, they asked. But he was right about where the world was heading.
Amazon’s chatbot processed a price adjustment the other day without forcing me to call or chat with a human agent.
Constructed human life is just more complicated than the AI capitalists would want you to believe. For example, even if an AI model can design a circuit-board, does that mean it's inherently useful? You need to source the wafers, cut them, package them, advertise them, etc. Given LLMs by their nature are confined to language and language-adjacent tasks, that is a very small percentage of the overall reasoning needed to make changes in the real world. In reality, LLMs are the intended way to extract maximal surplus-value from white-collar workers. We may see an increase in innovation as a result of that, but not because AI necessarily did it, in the same way that the power loom didn't create computers because its textiles clothed the computer scientists.
I talked to an LLM at burger king the other day until it couldn't figure out that I wanted to change the drink on my previous order.
Incredible... software engineers will be joining the breadline soon as managers, executives and PMs take over deliverables.
The world will look very different on Jan 1st 2027.
Software engineers have been trying to put themselves out of a job ever since the profession first came into being. Whenever an engineer gets a task their very first thought is "how can I automate this?" Going by mainstream consensus we should all have been unemployed by now. Yet every new leap into automation opens up a whole new tree of possibilities with an order of magnutude more jobs. So no, the profession will be fine. The only requirement is that you keep up with the new advancements. The people losing jobs will be the ones who still go "I don't trust this AI thing to write code for me".
> Software engineers have been trying to put themselves out of a job ever since the profession first came into being
That's by design. Software is all about optimizing effort and people who want to do this generally correlate with world view that better, faster, smarter humans are better for the world. If coding is gone, but humanity is 20% _better_, then ideal software engineer would be happy with this sacrifice. Surely people who cracked coding before LLMs can crack other professions and if anything a lot of this knowledge is transferable.
This time it is different. Because in the past, setting up that automation needed a, drumroll, qualified engineer. Now you can get a 14 year old halfway around the world who knows how to prompt alright enough to ship. There is no more moat.
"Low code" has been a dream of the industry for longer than I've been alive. There are reasons SQL and COBOL look superficially like English even when it's inefficient to do so. There are reasons Excel is the most popular programming language. Programmers have always been trying to enable non-programmers to write software.
Experience, culture, and domain knowledge are still somewhat of a moat. That foreign youth is unlikely to be able to write a good prompt for building, let's say, the software in an FDA-regulated medical device or custom Fortune 500 ERP application. The LLMs are great at building what you ask for but it's still garbage in / garbage out.
Increasingly less so though as these american companies themselves offshore not low skill work, but high skill work now brought on from general upskilling of the general population in recent decades along with massive investment in world class R&D campus facilities no different than what you see in that sort of facility stateside. Scary times ahead for the high skill american...
Would the Americans investing in SPY be saved?
Maybe I just lack imagination, but I don't really know how jobs are supposed to solidify around the role of giving prompts to agents and then looking at the results. I mean, engineers will be in the breadline because their role was simply to prompt the agents.. only to be superseded by managers or executives who no longer manage engineers but themselves prompt the agents? And, for this previously considered obsolete function which they do presumably by copy/pasting requirements from their email inbox, they will be paid by someone who doesn't know that they could just be talking to their own agents?
Sorry if I misunderstand the point, just trying to understand.
Regardless of the imagination quandary, this second, RIGHT NOW is the worst these systems will ever be. They are only going to get better.
I don't know if he's right, but Peter Zeihan thinks the breakdown in globalization will negatively affect the ability to continue to improve the chips that AI depends on[0]. Too many steps in the supply chain, too widespread, too vulnerable to deglobalization.
0: https://zeihan.com/the-ai-race-to-regression/
An invasion of Taiwan would definitely slow progress but it wouldn’t stop it.
We already have sufficient hardware that algorithmic (software) improvements alone should get us to GPT 7 / Greek Reference 6 even if not a single new chip is delivered to an AI data center ever again, starting today.
Maybe, maybe not. It's not unreasonable that these systems cap out at some point, or perhaps fizzle away entirely.
The businesses that create these systems are not profitable and run at a massive historical and go-forward loss.
New data centers required to operate these systems are facing increasing pushback at local levels. New construction is not guaranteed. Energy and power grid constraints exist as well.
Government regulation is way behind. What happens when (if) mass layoffs due to AI occur? How does the population react? Theoretically AI can be regulated out of significant progress, or outright existence for many purposes. At the end of the day, US and other prominent governments make the calls, not corporations.
For better or for worse, this technology isn't going away any more than search engines, smartphones, or social media have gone away.
Those technologies reached profitability relatively early, if not immediately.
https://en.wikipedia.org/wiki/Productivity_paradox
Computers never got profitable. They just made the alternative infeasible.
Also, if you believe Amazon accounting, e-commerce only very recently got somewhat profitable.
These inventions all stopped disrupting the world and just became a part of it. The question is whether LLMs are going to just take their quiet place, or profoundly change (or eliminate) humanity in a self-feeding frenzy towards singularity.
> It's not unreasonable that these systems cap out at some point, or perhaps fizzle away entirely.
Yeah, like computers and mobile phones did. Things that have utility, even if not immediate or initially obvious, don't fizzle out.
That's survivor bias to the extreme.
For sure, and for that reason I mean to say that I wouldn't feel great as a manager/executive/etc either.
Those nuclear powered flying cars envisioned in the 50s were also inevitable progress of the automobile.
There are two things SOTA LLMs fundamentally cannot do. They cannot take financial or legal responsibility for mistakes, and they cannot learn new things without forgetting things (except to a limited degree by adding it to their context). This is clear to anyone who has used even the smartest models for tasks requiring domain knowledge outside of math and coding, for which it's not possible to generate an infinite amount of synthetic training data: they still make stupid mistakes, and have limited ability to learn from those mistakes.
Humans also have a limit on the amount of domain knowledge they can acquire, albeit a much larger one. Executives hence cannot just replace all knowledge workers with LLMs, because executives have neither the domain knowledge to prompt and check the LLMs' work nor the bandwidth to keep on top of such a large volume of ongoing work.
In the US, Business' are treated like people with free speech rights. If it would be cheaper for them in the long run to use ai and robots instead of humans, they will figure out a way to make it so.
For the moment that may be true. They are getting better and better at acquiring, retaining, and processing domain knowledge. I wonder what this will look like in a few more years.
The responsibility side is a different matter of course.
>There are two things SOTA LLMs fundamentally cannot do.
I would say there’s a third thing. They seem to be very bad at being creative. Maybe they will eventually fix that, but if you ask it to come up with a list of business names or business ideas, for example, what you’ll get is the most generic, boring answer you could think of. They seem to be terrible at extrapolating outside of their training data. To me, this is the most significant difference.
> they cannot learn new things without forgetting things
Where did you get that idea from? Basically last few years was them constantly learning new things while improving their capability on the things they already knew.
I hope this is satire.
I hope so too!
Hope is all we have left to cling to, now.
It's not like managers and executives and PM's are the only people who can prompt an AI. And experienced software developer will be much more effective at using an AI to generate code compared to someone who isn't. So why would we expect the former in the breadline and the latter not?
If anything, I'd be more concerned about the leadership team being out in the cold. Why do I need a PM, or a manager, or a CEO if I can ship products myself?