I am in the process of attempting to have AI run my business. I'm actually making very good progress, but it's happening in pieces - I document some task and have it take over, or I give it something to handle while staying in the loop and providing feedback until it's in good shape. At this point it's handling large swathes of my operations, marketing and finance.

The experience of getting it there makes me pretty skeptical of the idea of a general business agent like this. First, because I still find myself having to review some categories of work for errors. These are decreasing over time, but they're still there. I fully believe that as models get better, errors will decline, but I am somewhat surprised to see some of the errors current models make given their intelligence. A common set is having Claude take some product photos and turn them into lifestyle ones using ChatGPT via Claude in Chrome. It's a pretty well-honed workflow at this point, but it'll still return images where the product is obviously not correct and seemingly not notice them.

Anyway, even if the agents are "perfect" in terms of their ability to execute tasks, there's still just an enormous amount of nuance and context in each business that takes a ton of time to convey. I've been at this for a couple of years now, and I'm still clarifying things. Now maybe an agent starting a business from scratch would have a better time since it's not inheriting all of this, but to have one run an existing business requires a very extended handoff, even if the agent is objectively amazing at all aspects of running a business.

> I am somewhat surprised to see some of the errors current models make given their intelligence

They have no intelligence. These are very very very refined prediction engines.

> A common set is having Claude take some product photos and turn them into lifestyle ones using ChatGPT via Claude in Chrome. It's a pretty well-honed workflow at this point, but it'll still return images where the product is obviously not correct and seemingly not notice them.

Obvious to you or I, or someone with actual intelligence. But things like this slip by a frontier model in the same way AI from a few years ago would generate an image with seven fingers. They do not count. They do not understand. They do not consider.

No matter how good these models appear to be at intelligent tasks, it's foolish to give them "a company" to run, because they cannot understand when they've made a mistake the way even the least competent human can.

> They have no intelligence. These are very very very refined prediction engines.

Silly and pointless criticism. Why do the semantics of the word intelligence matter?

Also, you need to understand that language evolves over time, and people often use a word to describe a thing that is newly discovered or invented that's similar to the word being used, because it's a helpful way of describing it that the listener will understand better than a longwinded technical description. The purpose of language is to communicate thoughts, not engage in pedantry.

> But things like this slip by a frontier model in the same way AI from a few years ago would generate an image with seven fingers.

Yes, you make an excellent point here that over time, the capabilities of these models to recognize certain categories of problems have increased dramatically. I expect this will continue!

> it's foolish to give them "a company" to run, because they cannot understand when they've made a mistake the way even the least competent human can.

The words of someone who has not spent time with the least competent human, or anything close to it.

> Silly and pointless criticism. Why do the semantics of the word intelligence matter?

It matters because, we're still sorting out what intelligence means for an AI agent. As you pointed out, language evolves over time. The question remains whether or not attributing intelligence to the current iteration of models is correct. This is not settled and I don't see why it's wrong to bring it up.

It's not wrong to bring it up, but the other comment did not "bring it up". It purported to correct someone by categorically stating "it is not intelligence, you are wrong and I am right", while not once saying what, then, intelligence is.

One thing is clear: LLMs at least already are capable of 1) not making the same mistake that person made, and 2) clearly seeing why the other person's post was "wrong" in both reasoning and tone.

[deleted]

I mean, here's a pretty good starting point then: https://aclanthology.org/2020.acl-main.463.pdf

That's all the more reason to stop talking about it and actually discuss concrete capabilities or lack thereof.

We're all well aware that we have different thresholds for what we consider intelligence, and that these thresholds are constantly changing.

Even if we agree that "LLMs are AI!!!" or "LLMs are not AI!!!" in this thread, all we've established is that this particular set of commentors share a similar enough definition at this point in time.

Language is indeed evolving.

Being good in chess was (and still is) associated with being intelligent. But if a computer does it, it is just a calculator (and it is).

Then go, the great game for intelligence, too complex for calculators to have a chance against intelligent humans. Until it was solved.

And then text, the original Turing test solved, AI capable enough to fool humans. And now already replacing humans in jobs strongly associated with intelligence - programming.

I find it hard to debate, that we don't have created artificial intelligence, by the way we used to use the word "intelligence" before.

So if AI is really understanding something?

Likely not in the way we use that term. But it definitely shows intelligent behavior and actions.

Go isn't solved.

Go is solved.

Hey, look I can make unsupported claims with just as much evidence as you!

Maybe you'd like to add some details as to why AlphaGo and its successors haven't "solved" Go (insofar as a game like Go can ever be "solved")?

When people talk about solving a game like go or chess, they mean proving mathematically if there exists a way to guarantee a win, or if the second player can always guarantee a draw, among related questions. Current game engines have no such knowledge, they just pick statistically what the think the best move is and hope for the best.

No, that is not what "people" in general mean, this is only what some people mean.

Other people know there is no solving go with calculating, but using statistics to achieve the goal of becoming better at humans. And they are. (With a recent unexpected exception unlikely to be repeated more often)

I don't think people "in general" refer to games as "solved" or "not solved" at all. Among people who do, chess and go are definitely not referred to as (fully) solved.

"to have a chance against intelligent humans. Until it was solved."

But this was the original statement - so solved references "chance against intelligent humans". And this is clearly solved. No comment on that there cannot be better go engines, but no matter how painful it is, they beat the best humans. (And it was painful, they did cry)

So let's say we make up a new word, machilligence to serve as a parallel term to intelligence but strictly for machines. How is the world different in that case vs. if we call it intelligence?

Or I guess to put it another way, if we want to coin a new term for the intelligence-esque thing that AI has, but the key differentiator is that it's an AI thing and not a human thing, then what linguistic value does the new word have? If I said "Claude is intelligent" then the fact that we're talking about AI "intelligence" is already captured in the sentence anyway; no new word needed

I think "intelligence" carries baggage. When most people hear it, they're thinking of the constellation of things: judgment, consistency, moral reasoning, the ability to decide something and stick with it. An intelligent person has an internal model of the world, values they apply consistently, and the capacity to learn from mistakes in a meaningful way. LLMs don't do any of that. They perform statistical pattern matching on text at an extraordinary scale. They're shockingly good at mimicking the surface features of intelligent behavior. I think the overuse of the word intelligence is something to criticise as grandma is not across the tech details.

It matters the same way that calling a dog or a cat intelligent matters. It humanizes the thing, even if that isn't your intention. Right now that's not a big deal since most people agree that machines should not have rights, but I wouldn't take that for granted.

It's not the word I'm taking issue with, it's the concept. From your post you are surprised that the model makes mistakes and does not recognize things, and these are only surprising if you imagine the models to be thinking about things.

We can argue about words like thinking and intelligence all day long, and so can an LLM. But at the end of the day the LLM is only mimicking the processes you or I use.

Last week I tried out gpt sol 5.6 and asked it to count the letter r in a massive string of letters, without using an external app. It succeeded until I made the garbage sentence suitably large and then it consistently, confidently failed. Each time the "thinking" showed that it was teething to find a "gotcha" each time. "Ah, the first time I forgot to count the letters in the instruction itself" etc. at no point did it just understand that it had miscounted. It seems to be incapable of considering that it just made a regular mistake. Even a six year old child would just try again the same way and end up with the right answer

I know people who have run this experiment, with companies in the 7figs ARR, indie devs. Both of them have reverted to hiring humans.

[dead]

"Judgement" is the #1 thing missing from AI if you ask me.

You give AI an input, and then an output is created...a very coherent and deep output in fect...but, there is very little understanding of alternatives or downstream implications in my experience.

I think its a very good text/image generator on any subject. But I am not finding much in the way of understanding tradeoffs or downstream implications in a non-pro's and con's list. It's like a human that has a particular type of brain damage...they can make an infinite pro/con list, but not reasonable decision is reliably made.

> They have no intelligence. These are very very very refined prediction engines.

"Very refined prediction engine" is not a bad working definition for intelligence. It is not the only component you need, but it might be the most important over-all capability.

If its making lots of errors, its not doing a very good job of predicting the outcome of its actions.

It's a terrible definition of intelligence. Intelligence is far more than just predicting things based on a massive set of similar things. We do do pattern recognition, but I don't need to have seen 10k dogs to be able to recognize a dog. And it's not a difference of degree, either. "Meaning" is something I am able to derive from my experiences, and it is not something an LLM does or will ever be able to do (nor is training data even analogous to having experiences in any way shape or form.)

I didn't say anything about training methods. I said an intelligence can predict things accurately, most particularly if it can predict the outcomes of possible actions it can take then it can do planning. Its especially important that it should be able to predict outcomes in cases that are not exactly the same as things it has seen before, but how it gets to that capability level is not important for understanding what the capability unlocks.

This is one of the most distinctive qualities of human intelligence compared to many animals - is our ability to adapt to new situations and make accurate predictions of the outcomes of our actions. Other animals can also do this but over very narrow time horizons and situations.

AI agents are better at it than any animal, and better at it than humans in some domains.

Is "meaning" inherent to intelligence? Is emotion and the ability to have a subjective experience inherent to meaning? Genuine questions that I'm not sure we have good answers for yet.

Your other point that humans need very few examples for learning vs. what LLMs need is an interesting one. A few researchers talked about this very thing on the most recent episode of Dwarkesh's podcast.

My own view is that it seems unfair to compare an LLM's training with just what a human gets over the course of a lifetime, because our brains have been trained by a billion years of evolution for pattern matching. I'm not at all confident that AI models won't catch up.

I'm curious what your working definition of "actual intelligence" is.

> I am in the process of attempting to have AI run my business. I'm actually making very good progress, but it's happening in pieces

Very cool! What's been the hardest part? Have you successfully automated non-trivial communications? (for example with prospects, customers, vendors, partners, etc.)

> maybe an agent starting a business from scratch would have a better time

This describes the project I've been working on since last year, with agents in company roles, deciding on strategy, collaborating, and making progress (admittedly nonlinear). Some parts of the platform are stronger than others. So far the company agents have developed a company strategy and plans for executing it, launched a website and blog, and built and operate two products, one free and one paid.

But I would be wary of using a third party like Pion as a platform for running a business. It's one thing to know that your prompts and responses will be used to train models in the future, by companies that have a huge sea of relatively unstructured data. But it's another to hand over to another company every aspect of your strategy and operations, available to that company in real time and structured in such a way it's quick and easy to understand what you're up to and where you are going. Especially if it's being offered for free and at scale. With a free offering, value creation will likely come from either customer data or from escalating prices once customers are locked in. And even if neither of those happen, you now have a huge single point of failure for your entire organization. Seems like there are lots of strategic risks in there for the users.

> At this point it's handling large swathes of my operations, marketing and finance.

Can you expand on this in terms of what it is handling specifically and how?

He writes about it on his substack. https://theautomatedoperator.substack.com/

Cool that he's writing about this publicly, but for someone that's so business-minded I find it odd that he never renders any of these AI experiments in terms of ROI. Hard to find any signal in those posts on whether it's profitable to automate these processes with AI vs. alternatives.

That's a fair criticism for some of it, but in a lot of cases I try to automate things that don't have alternatives because they're unique to me (e.g. all the brands I own in my fund have one bank account but need to be tracked separately for investor payback purposes, so I have a Claude skill that takes all of my transactions and assigns them to the various Google Sheets ledgers I have).

In other cases I'm doing stuff that's complex enough that there's no real "alternative" at this point that doesn't involve hiring one or more people and working with them for an extended period. My next set of posts is going to be about my latest acquisition and how AI redesigned a Shopify store then designed and launched Meta ads, all of which only took a couple of days but is going to easily push revenue up by something like 40% this year. The money made is great, as is the money saved on people I would've hired to do this sort of thing, but the real cherry on top is that even hiring someone would've required me to spend waaaaay more time than I did here. So big ROI on the money but also the time, which is arguably more important, since time savings permit me to acquire more brands.

> unique to me

Ah, so it’s ”No Silver Bullet” (1986) [1] all over again.

Thinking out loud here...

AI mainly helps reduce accidental complexity. It can help one understand essential complexity, but essential complexity must still be paid.

You describe doing the essential, irreducible part. And that’s specific to your needs, depends on the problem you’ve chosen, understanding reality of your domain, evaluating tradeoffs, and being accountable for the outcome.

Right?

[1] https://cekrem.github.io/posts/there-is-still-no-silver-bull...

For anything visual, they are still effectively blind, right? Still just working off the image embedding, sometimes using scripts to actually inspect individual pixels.

Blind seems too far here - Claude's increasingly able to diagnose visual issues on its own. I used to have to check every single image it generated, but now it's at the point where it'll catch most bad ones and regenerate them before it gets to the review step. Still misses some, though.

no. I just used Gemini to update a website I'm managing for a charity, and one of those steps involved Gemini, on its own, offering to scan the website for one of our beneficiaries, specifically the header images, to see if there was more information to include in our writeup about the beneficiary. And it did it quite well.

Right, I forgot that they also read and parse text in images well.

How are you automating all the parts? OpenClaw/Hermes?

Mostly Claude Code and a lot of internal tools that it's built. I've honestly still not had a chance to play with those kinds of harnesses yet, but my impression is they're just a better way for you to instruct an assistant to take care of things in order to get them done. My goal isn't to be managing an assistant, it's to get things automated without my input (and where my input is needed, have the assistant escalate to me).

What is your target audience and what do you sell them?

From their blog [1]

"I acquire e-commerce brands that sell on Amazon"

Honestly it sounds like the OP is part of the machine that makes Amazon such a trashy marketplace these days.

[1]: https://theautomatedoperator.substack.com/p/15-ways-im-using...

Well that's not very nice!

And for the record, I only buy brands that sell high-quality products. Marketing and operations I can fix, but if the stuff they sell is no good, there's no point.

How do you inspect the quality of the products you're selling? From your blog it appears you just read through their sales information, and don't ever see or handle the product. And yet we all know Amazon sellers juice their sales metrics and ratings with scammy practices, and now I see selling their brand is part of the incentive.

And more, you are AI-generating listings and images for your acquired brands. You're filling Amazon with slop, and automating the process.

You should seriously reflect on your role in polluting these marketplaces.

How do you identify them and choose what to invest in? I guess you do not really have access to their technical specifics or sales volumes.

Of course I do. No one buying a business would do it without that information. I get their P&L before I even get on a call with them, and during the due diligence period I get access to Shopify, Amazon, etc. to get first-party data to validate their claims.

[deleted]

One step up from “vending machine” as far as business complexity goes.

I suspect you mean this as an insult, but it's not too far off. That was one of the main points of the business - each one is simple to run, so I can manage a lot of them. And that was before AI got really good.

Anyway, I'm sure the app you makes that lets you get a phone number that can send image messages is very complex, and I think that's very nice.

Open to sharing what workflow you're using on the lifestyle image creation? I've tried the Higgsfield MCP, and direct Figma integrations - but keep running into hallucinations on product dimensions, etc.