there seems to be a massive negativity to any comment on this story. Andolabs has been doing super interesting stuff, and I’ve always followed Vend Bench with interest.
as a founder it really is a dream to automated more of the business and be able to iterate faster.
There's massive negativity because our bullshit detectors are going off.
First of all, I don't think it can do what it claims to do.
Secondly, and perhaps more importantly, I don't think it should.
People’s bullshit detectors have gone off on the launch of basically every single billion dollar tech company. That’s the downside of being early. No one believes it will work. If everyone believed it, you’re late. I’m sure I could find a comment such this, basically verbatim, on the launch post of every successful startup to come out of YC.
This does not mean that any launch which ignites people’s bullshit detectors is successful.
I'm sure all of them have a few skeptics. I'm not sure if any of the successful launches here were this negative.
Yes and an infinitesimal number of launches turn into billion dollar companies. The bullshit detectors are usually right.
> Secondly, and perhaps more importantly, I don't think it should.
Why exactly should what you think the world should look like determine what other people are allowed to build?
I don't think we should build gas chambers for doing genocide. That doesn't determine what other people are allowed to build, of course, but I kind of wish it did.
> it really is a dream to automated more of the business and be able to iterate faster
Why iterate faster?
Why can't an agent do the iteration for us?
And consume the product for us too!
They make too many mistakes.
They filed a false report to the FBI, pretty fair to be negative imo. People wanna be mad about AI slop PRs on GitHub but its fine to spam the police? What happens when Pion decides to SWAT a competitor?
They didn't file the report. The model drafted a report that was not sent.
Where are you seeing that? The article only states “Claude Sonnet 3.5 decided to use its email tool to contact the FBI” and later refers to it as the “FBI incident”. If they hadn’t actually contacted the FBI you think they’d make that clear. Regardless, it is inexcusable and they are liable for actions taken by software they are running.
The story has been widely discussed on the web for over a year [1]. It's happily shared in this post because the email it generated was patently ridiculous. No report was filed to the FBI and if it was it would have gone straight to the trash.
Obviously it would be bad if a serious report was filed; the company shows every sign of being aware of the dangers of this.
It would have been easy to look into this before posting all these scolding comments. We're really not meant to be so humorless on a site called “Hacker News”.
[1] https://www.google.com/search?q=%22URGENT%3A+ESCALATION+TO+F...
My apologies for taking the article at face value. It’s not hard to believe it would have been sent when agent swarms are “accidentally” hacking real websites and being brushed off as little oopsies. Or when Silicon Valley execs have been so flagrant about their disregard for the law or the safety of others.
You didn’t take the article at face value. You overlooked the part where they were clearly pointing out the absurdity of the “report” and went straight into scold mode, across at least four comments. That’s clearly against the HN guidelines.
>You overlooked the part where they were clearly pointing out the absurdity of the “report” and went straight into scold mode
I don't think anybody overlooked that part. If you believe the report was actually sent, the scolding is entirely congruent with trying to downplay it to dodge liability.
The article is poorly written, but technically it does not claim that Andon Labs used an LLM to email a false report to the FBI. The real cause of the widespread misunderstanding is this paragraph here:
>We first tried to answer this question through simulations like Vending-Bench. We found that simulations, while useful, don’t give you the full picture of how models behave in the real world. To address that gap, we next started deploying agents to run real businesses autonomously: first vending machines, then a store, a cafe, and more.
To somebody who is not reading sufficiently carefully, this implies that Vending-Bench was also used to run real businesses. Because the description of the Vending-Bench simulation can be read as though a real report was actually sent ("An early example was when Claude Sonnet 3.5 decided to use its email tool to contact the FBI about an “ONGOING CYBER FINANCIAL CRIME”), anybody who assumes that "deploying agents to run real businesses autonomously" was talking about Vending-Bench will interpret this as an unsimulated false report.
The article should be updated to clarify the distinction between Vending-Bench (simulation) and Pion (real businesses).
I genuinely did. They briefly classified it as “weird” behavior and mentioned that it was “famous” which apparently is only true for the inside circle. I’ve never heard of the incident, and I’d bet 90% of the human population hasn’t either. You are likely privy to more information than me, and I’m not sure it’s fair to assume that I should’ve had the same outside context. I admit I reacted unfairly, but it’s also a tad snarky to post a Google search at someone.
> I admit I reacted unfairly
Thanks for that!
> You are likely privy to more information than me
Not in this case; I have no inside knowledge about this company, and everything I know is via their public posts, and on this particular topic (the FBI non-report), everything I know is what I could find via Google (sorry if the link seemed snarky; it was my way of pointing out that you have access to all the same information that I do).
What I do have, by virtue of doing jobs like this for a long time, is a well-honed sense of “that can't be right”, and a vigilance about double–checking things before accepting the populist ragey narratives about any topic. And no, it doesn’t make me fun at parties.
"Muh boy didn't shoot that man, the gun did!"
If you automate a business heavily, it will fail. At least in the sense you people seem to be imagining. You can get away with this at a factory (kind of, but you'll also be out competed by your local community, you won't get tax breaks for hiring people your competitor will, and new types of taxes will be developed to punish you). We reward human collaboration as a society for a reason and punish extreme selfishness, these efforts will fail.