The AI is not remotely ready for this - this is an absolute delusion they are selling.

"By the way don't do this again" <- as if the AI has the ability to ingest and systemically diffuse this.

I think Zuckerberg himself is deeply into the Koolaid, and is likely himself unaware of the limits of this tech.

He's probably surrounded by enablers.

Had that problem here. Our product lead thinks it’s infallible and is surrounded by yes men.

Day one our agentic platform goes live it causes a reportable compliance issue. Massive clean up. Reputational impact. Turned off.

No one held accountable still. Assuming that adding more guardrails will fix everything.

You should write about that, because it's the primary systems failure of AI right now.

'Managers Delusion' - which includes techies as well, to be fair, at least they have the excuse they are one step removed.

I'd rather everyone burned for it than heeded a warning.

They are morally bankrupt and will try and do it again and again if it's stopped too early.

Lyrics come to mind: "It's an addiction bound to stick around, a junkie won't bounce until they hit the ground"

'Say Hey There' by Atmosphere

[deleted]

They're building for future capabilities. It's a logical approach to leveraging the data they have.

If it uses "memory" files like Claude Code there's a good chance it works.

If it was Claude I'd expect it to now put some form of this into every output even when not really related to the task: "No offers were accepted without consulting you and I haven't shared your address or availability."

Is it guaranteed to work? No.

And obviously it's a terrible idea to set up a chatbot to communicate, negotiate deals, and handle logistics on your behalf.

"If it uses "memory" files like Claude Code there's a good chance it works."

You have explained literally why it would not work.

The AI can absolutely not depend on 'arbitrary statements in some file' as operational policy.

For a very, very narrow scope of work, when it's well defined, when the information is rigorously applied, sure ...

But they don't have that.

The are throwing agents out there like they can handle this degree of complexity and nuance, when they cannot.

100% failure rate over any period of time.

I didn't say it would work with high reliability, or that it would fix this clearly unsuitable use case.

However, it is untrue that it doesn't have the capability to memorize an instruction and diffuse it to new sessions.

Simply saying "do not ever do this again" can result in the behavior not reoccurring with any likelihood.

You'd have to benchmark whether with the instruction in place it would violate it, and in how many cases, so that you can understand the risk better.

I don't think it's reasonable to really talk about the fact that in theory, AI can take a user message, save it in some random place, and ingest it later.

I think we get that.

It's completley unreliable, which is the issue.

It's not in theory, this capability exists in practice, and you have no clue whether this specific instruction will work in 90, 99, or 99.9% of cases.

I think it's unreasonable to pretend that this couldn't possibly work, that you understand to what degree it does work, and that the only thing we should be discussing here was that it isn't deterministic, which every reader already knows.

No, this is upside down.

Your argument exemplified by 'the ai can write to a file and use that information in context later' - implies a 'capability' to do theoretically do something, but not with any consistency at all.

Not in any scenario - even with human oversight - would be 'trust' any system to reliably work on the basis of arbitrary memory.

'It can possibly work' does not map to any reasonable conclusion that it will work with any degree of fidelity.

There are almost no examples where we rely on AI in this way. Chatbots for customer engagements etc. are merely thin language wrappers around deterministic systems.

It's 'possible' that Meta is doing something rigorous and deterministic but it's not remotely reasonable to make that assertion because we literally have the evidence right in front of us. That's literally what the article is about.

The Wright Brothers plane 'technically took flight' but it's not reasonable to conclude that it can fly passengers safely in any way, especially when the story is about a flight crashing.

It would be novel actually, if there were a story about Meta's unique use of AI in these kinds of information flows, that transcended what we all recognize as AI's inability to reliably process information, given what we know about AI and how it works aka 'it can write to memory'.

A substantive discussion would focus on whether Meta's Muse agent does indeed have a memory feature, whether the user's request to 'never do this again' triggers it, and how good their Muse model and harness is at following such instructions.

The discussion was good to read. You seem to not want to understand that appending text to a file called MEMORY.md or whatever name doesn't reliably work and would cause another "that's on me"

No, the millionth rehash of "LLMs are not deterministic" was not interesting to read, and if you want to hear rants about how Zuckerberg markets his products, there are surely better platforms than HN.

Here I expect factual discussion rather than misleading statements that suggest there was no way for the Muse assistant to follow an instruction never to do something again.

Vague arguments that such instructions are not deterministic are uninteresting, because it is obvious.

[deleted]

The AI is not ready for this _yet_, but it will be, and FB wanting getting ahead of the game here is potentially good business. It’s all in the public perception of utility vs fuck-up, and it’s far too early to say Zuck got that wrong, and indicators are he got that right.

> The AI is not ready for this _yet_

It should be ready in about two more weeks! How many models have a Ph.D level intelligence now? I feel like I’ve been hearing that for about a year at this point.

I’m guessing you’re not a programmer. If you were, you’d have seen models go from “kinda helpful for programming” to “usable as a daily driver” about 9 months ago, for example.

“Models aren’t improving incredibly fast” seems a very odd point to be making.

Yes, it would be absolutely unthinkable that a certain automation architecture might plateau somewhere. They will probably be performing open heart surgery sometime next year.

I think if you are equating “performing heart surgery” with “a modest drop in already successful inventory management”, we may not have a shared basis in reality from which to converse.

[flagged]