I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines.
I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries
I feel like there are so many cloistered people at these companies that they are left scratching their heads about what normies even want. Like, they literally can't fathom basic stuff that isn't just highly consumer-oriented. I dunno, like applying for government services, paying your gas/elec bill without being confused af, keeping the dr up to date with your dad's illness, or how to get your newborn to sleep at 2am.
I will happily shout AGI from the rooftops the day I can turn on voice mode in ChatGPT and have the model calm down my toddler for a tantrum, or keep him from opening all the bananas instead of eating one.
People in this thread arguing about AGI relating to Einstein problems and physics. Yeah no. When AI can handle toddlers then we're getting somewhere.
All those stupid Bell Labs researchers not inventing Uber or Tinder. How come they didn't just build the obviously popular and profitable businesses that became possible once they invented the internet?
Many VCs also dislike these examples, I believe. I'm doubtful this is what they're being pitched.
As for public releases: I wonder if it's because these examples are easy to relate to. Many websites are just a long tail of industry or use-case specific stuff. What's valuable to me probably means nothing to you. This is unlikely to resonate with people-wit-large (and LLMs are marketed broadly) or requires the reader to think (and marketing that requires thinking is bad these days).
Second, it's arguably a good litmus test. If it still can't do the worn out examples of plane tickets and shopping, which would be a good assumption since we've been demo'd these use-cases for 2 years at this point, then ...
This is the funniest part of it all for me.
Ok we have AGI, so where are the _things_?!
There's a lot of truth to this: Execs want their own use case covered. In addition, everyone wants to be better than the competition at basic use cases, because that's what people compare first - even though we all know they are not a good reflection of reality. I'm currently working on exactly a project like this.
Maybe this is why they NEED AGI (does it come with a soul?)
Don’t forget making podcasts and telling you what to bake with your kids
tarpit ideas
That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?
As a realistic construct:
When I tell a bot to find the best value per volume for a reasonable quantity of unscented Dawn dish soap [so I can buy that], then: It often makes a complete mess of this seemingly-simple operation.
(And yeah, that is an actual thing that I've tried to accomplish with voice commands while standing in my kitchen and doing some dishes. It seems very simple, and it did not go well.
Maybe when we get the basics figured out we can start worrying about how inept it is at doing vacation planning.
It seems that this kind of thing isn't sorted at all, and that this is a very real problem for those who are in the bot business: These missed opportunities leave money on the table.)
I think we can safely assume that AGI is still a good ways out then.
These AI LLM products really do well on the illusion of seeming intelligent to regular folks, but I'm not fully convinced they can replace human brains... yet. ;)
I mean: That sample was only demonstrative of shopping for dish soap. :)
I often get some seemingly-good results in the non-shopping research department, where I'm exploring science and physics; things that don't generally change rapidly.
But unlike dish soap, I can't buy the science results that I want[].
[]: er. well, ackshually... let's just not talk about that concept right now.
I wanted some banana plugs for my speakers, as well as I needed some longer speaker wire. I kinda wanted it for the weekend so I could watch movies with some friends. I didn't want to order online as that wouldn't arrive in time. Three stores around me had some, but they were quite far apart and I didn't want to drive to all of them, so I needed to figure out which store had ones that were actually decent and not too crappy.
LLMs completely and utterly failed at this.
The Mensa test results are impressive but are in no way shape or form relevant to real world scenarios.
I wanna see benchmarks on successfully completing benefits applications, on finding the best option to purchase X from, on correctly identifying a piece of furniture and its condition and setting a correct price on it and selling it on craigslist/FB market place and so forth.
Funny, I had a similar task for an LLM recently but it succeeded, so I'll describe it:
I asked the chatbot (librechat, tavily-mcp connected, kimi-k3) if anybody tests dish soaps rigorously. It said yes. I asked it to find the consumer reports from Germany. It did so, full access is paywalled but it did see the top positions. The I asked it what term to put in evay.de if I want to order. It gave me a string that found the auctions for me. I also asked it what the actual difference is and it described what info the EU law forces onto the label and told me what to look for in the ingredient lists.
The process is still involved. I wish I could just do it all by saying "find me another dishsoap, this one sucks". But then would everybody saying this phrase except the same approach to the issue that I took?
Because some people do.
Corporate travel is an example. In many organisations, you tell someone in the travel department "I need to be in Tokyo for this conference from Tuesday to Sunday, and charge it to this cost code", and they figure out flights, accommodation, etc for you, with minimal input from you.
If an agent planned a flight for me with an overnight layover, I'm unplugging it. I don't care whose dime it is lol.
“Find me a cheap ticket from Seattle to LA”
… 23 hours in Denver later…
Claude code planned my recent trip to China. I'm a very experienced traveller but don't enjoy planning. It was a great trip.
Yeah but the contention is here automating the bookings. Did Claude book the flights, hotels, tickets and attractions for you or would you let it without oversight?
I think it depends on what you do for work. I'm not going to ask an agent to book my flight for my vacation to French Polynesia. I want to pick my seat and potentially find a deal making an upgrade worth it, choose an airline, etc.
But my routine business trips in the CONUS with strictly defined booking options... let me just email an agent "Get there by meeting on day A, leave after meeting day B" and have it sort it all out without the drudgery of the corporate travel portal. YES PLEASE!
By that point, you don't need an AI that boils the oceans and hopefully doesn't confabulate or misinterpret your words, all you need is a better corporate travel portal, which is a legitimate workplace productivity discussion to have, and a solved problem with traditional methods. Ours is effectively close to the workflow you describe: picking dates and time brackets, destination, fine tuning flights and hotel options, sending for validation, less than 10 clicks through information-dense and effective/predictable screens which I wouldn't want to trade for a chatbot and it's usually over the top words salad.
Most AI use cases can be reduced to a deterministic listing/form/report format (on the web) or a cli tools.
We have had idea of expert systems for decades at this point. And spend however much on software development. Some how we have failed to reach this in way too many places. Even simplest things like cancelling some service might be broken...
And now we want to add some random factor into middle of it all...
A good agent would call you to ask clarifying questions and then actually get you that perfect seat, etc. ideally
And I could bet money on that in short time after someone builds that sort of system the next step is to make it worse. Push worse and more expensive options to user. Or at least those from highest bidder... Anyone involved just can't keep themselves honest so it is doomed to be exploitative.
For me, it is not a matter of trust but that I actually like shopping, planning a trip, deciding what restaurant to go to. Deciding what to buy when shopping is a matter of personal taste and not intelligence.
A human assistant is largely a status symbol. Most people are not really that busy. The real problem with an agentic assistant is if everyone can have one then it no longer acts as a status symbol.
Plan a holiday, definitely. There are many esoteric things one has to research to properly plan a holiday that I’d rather just not. Some things require reservations months in advance and I’d rather an AI just figure all that out for me ahead of time.
Plan, sure, many people ask this of AI already, but not actually ask it to buy autonomously.
The rich fucks who run the show do.
(This holds generally) maybe because the people most likely to be convinced to buy this particular thing are the people who buy a lot of things. The people who only buy things they need, they’re a harder demographic to convince via ads, and they tend to pay less since they rarely need extras. Every ad caters to whales.
Maybe we are not the real audience. I wonder if they are trying to convince the advertising industry/investors that in the future it won't be Google search that stands between the consumer and the product, but rather their LLM.
Maybe because the people making them are workaholic types who really don't care? I've certainly been in situations where I didn't really care what shows up for a meal. Someone was tasked with getting food and we leave it up to them. Sure, I can imagine such a meal being bad and it has been a few times in my life but 24 of 25 times, maybe more, it's fine. Further, the AI knows your preferences.
You can't even get many people to buy things online at all and if you can it's less profitable than retail, because you need to spend a lot of money to convince people, advertise to be seen, and account for returns. I think this is also due to the factors you mention.
One quick example: In fashion, Inditex and Shein have about the same revenue (€39.9bn and $41.8bn in 2025), but Inditex is more than three times as profitable. I don't see how there is a demand for agentic commerce that would remove even more control from the customer when shopping. Part of why we shop is for the experience. For B2B producurement platforms like Alibaba I can see the appeal though.
I recently needed to buy some hardware for a piece of furniture.
Ran Codex, it found it for 18% less than what I found in the top Google results. It did it by finding smaller shops, applying a discount code, subscribing to a newsletter for a better code after approval, and took into account the shipping (by placing it in the cart and going to checkout) all to get me the best price.
I’m guessing without it I would have spent much more time on it and paid the original price I saw.
If you use AI agents well, they can easily save you more money than they cost, and saving money is something most people are pretty excited about.
(Disclosure: OpenAI employee)
>to get me the best price
How do you know it's the best price ?
Found the evals fan!
I mean, that's cool and all, but the numbers are really going to shift when it accidentally goes off and orders that same hardware from every vendor in your local region and the top 5 online results for comparison.
It's the same problem as all other LLM solutions (that I hope OpenAI is working on!) it's non-deterministic, and there's no way for the user (or model provider) to know what the distribution of possible outcomes is. This just gets compounded when multi-call harnesses come onto play.
its going to get worse though as cloudflare keeps blocking more and more
I‘m using mostly the browser automations for things like that now. Same for research - ChatGPT is banned from reading many pages, but Codex can read anything I can.
But isn’t it funny that Cloudflare is blocking AI on their pages, but on the other hand is researching and marketing things like „you can put a browser in a CF worker“
Their browser workers won’t get blocked. Same with all big vendors, keep out small competitors enjoy access yourself and sell it to a select few partners.
The sites that do that won't be getting money from my and others' agents. Guessing that's going to become more and more of a problem for those sites.
I expect a small site that undercuts the top Google results by 18% with a sign-up discount probably isn't profitable on those orders, so blocking agents would save them money - it's not like someone using agents in that way is going to have any loyalty to shopping from that site in the future.
If no one can find your site because you block agents, that won’t do much good either.
And yours and other's agents will probably remain an insignificant and invisible customer base anyway.
My crystal ball is as good as anyone's, but if "agentic shopping" ever becomes mainstream, you can be sure that the vast majority will ask their phone (i.e. Google, i.e. Google Shopping) what the best price is anyways.
Sure, Google could be their agent, or ChatGPT, or Claude, or whatever local model the person is running. The big guys might all have the results cached so they don't have to rerun the crawl. Whatever their choice of agent, it seems pretty clear that almost everyone's going to use them, they're way too useful not to.
Thanks. This is genuinely a cool usage example.
How did you run this? Web interface, desktop app, CLI?
How did you complete the final transaction?
I really wonder what the demo people are thinking as there are many great uses of LLMs in day-to-day lives but all we get is these 3 rehashed use cases. No wonder normal people think LLMs still can't do anything.
The overwhelming majority of things I buy are things I've bought before. Alexa having access to my Amazon order history means I can just say "order a new water filter for my fridge" and the correct item shows up the next day. Far from life changing, but it's a feature I use somewhat frequently these days. Similarly, I would trust an AI to put in my usual Chipotle order or pizza from my local pizza joint.
I wouldn't want it to pick food for me from a place I've never been, though to be honest with enough order history it could probably do a decent job at it.
There was amazon dash button for this
This isn't something you need an AI to do for you though...
"Somebody is wrong on the internet" will get you more, and more reliable, results on what people actually want from AI. You are now participating in a very high-value survey.
Man I can think of so many reasons why companies want “agentic commerce” to catch on - and none of them are ethical.
I work for a larger german retail chain and agentic shopping is already on the "near future vision". No one thinks this will be used but somehow shareholders love it.
Lol same story here, I work in payments and 0 people within the company (including the team working on it!) are convinced at all about the viability
True, but they're still friction to be reduced here.
What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow.
I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.
because "people will let our AI spend their money for them" is the workflow that makes their valuations reasonable.
Agreed. There's not many things I don't want AI to help with, but buying stuff autonomously is high up on the list of things I don't want. Brockman's latest interview was something like: "AGI would be able to say oh this band is playing, I bought the tickets for you and arranged your flights - I hope you don't mind" (paraphrasing here). I definitely don't want AGI running my life like that so I can be a mindless consumer. I'm sure the advertising/marketing companies would love it though, so they can make closed-room deals with AI providers to shill you garbage you don't need. Just another reason why open-weight models need to keep up.
If it was Google's marketing department it would be: booking a table at a restaurant.
close enough.
The second to last line is "book it" for some tennis thing, and the scene before that has the guy eating the food the ai ordered.
I’m guessing the marketing must mean the AI-shops-for-you use case is a pretty big market, much bigger than AI-makes-life-easier.
I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate.
"Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip!
I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming up with such demos.
(Note: I live in SF)
I mean, not auto-purchasing with the card, no. But my wife has definitely used chatGPT to plan activities on the trips we've already booked.