This is a good essay, and makes me hopeful.

I’m on the record saying that it is extremely dangerous to slow down because the race for AGI is a zero-trust game — defections pay - and combined with a compounding returns model on defection, if you have any strategic adversaries whatsoever you MUST NOT slow.

For slowing to make sense, you need to believe that you can transform the zero trust game into a cooperative game, or that it’s likely racing will lead to a negative outcome for the ones racing ahead (and not everyone else). I don’t believe either of these outcomes are possible, and so I advocate for racing, acknowledging the entire game might be a negative value game, or at least could be for some time — it’s even worse not to play it.

But, I like hearing what reads to me like very thoughtful and informed (internal) policy considerations is great — the public messaging from Sam and Dario just seems so facile and simplistic I’ve been worried.

Everyone in the "if not us, they will" race is brainwashed into thinking they belong to this or that party, while in fact collectively comprising the same entity that pushes forward all the atrocities known to man.

No. These parties are composed of people who most definitely think this way, and therefore will have distinct goals and interests when presented with opportunities. That’s reality quite aside from how a game theorist assesses the situation.

I cannot tell what negative-sum outcomes you consider possible. Do you believe AI can drive humans extinct? How many of Zvi Mowshowitz's Three AI Pills would you say you've taken?

https://thezvi.substack.com/p/the-three-ai-pills

I’m like a 2(.5?) there - I don’t think ASI will care about my kids better than I will for some definitions of better, for instance, and I feel very fuzzy and vague about what actual differences in qualia between me and ASI would yield in the wild.

I’m not a doomer, although I don’t think doomers are dumb, just wrong. I think you should design your systems around the possibility that people who disagree with you are correct , hence my nod to negative sum. If you have more than 30 years to live, I’d personally rep to the most likely outcomes being very positive. With a lot of disruption in the middle.

> if you have any strategic adversaries whatsoever you MUST NOT slow.

What if the most dangerous strategic adversary you have is the one you are building?

What if this is true mid or long term but by not participating to the AI race one gets poor or killed in the short term? The only way out would be that all parties agree to stop. There are previous examples (e.g. nuclear proliferation treaties) but it gets hard to do it with hundreds or thousands of parties.

I don't think it requires the agreement of that many parties. How many organizations/physical sites can create chips capable of training and running frontier models? That is your bottleneck. It is equivalent to targeting uranium enichment in nuclear arms control.

[dead]

Although many share your mindset, I’m glad there are also many that don’t. Otherwise we’d still have countries in a race to keep building up their nuclear weapons for the same exact reasons you just described.

The situations aren’t equivalent - luckily in my opinion because the stakes with nuclear are much higher. von Neumann constructed a multinational game theory approach appropriate for weapons. AGI is a much harder problem to corral because there are so many benefits beyond just blowing up cities. But it’s also a much better thing to have for these very same reasons.

Similarly there have been few positive externalities from nuclear industry, making it easier to make the case to wind down research. This same set of concerns in biotech is much harder to get compliance with, precisely for this reason.

Anyway I’m especially wary of over analogizing to nuclear era concepts: I think they’re a trap.

In your opinion, what are the top 2 positives and top 2 negatives of humanity inventing AGI?

slowing can also make sense if you know you're running full force into a bomb or a wall even if other are close behind.

If you must not slow, why did we slow down making nukes? Seems that sometimes, eventually the rat race goes on long enough where all the players no longer care to play into the farce like their predecessors who passionately beat that drum.

Never forget the one goal of the corporation, and that everything is said and done in the furtherance of that goal.

> Never forget the one goal of the corporation, and that everything is said and done in the furtherance of that goal.

Do you post this comment on every single blogpost with a corporate domain? Why or why not?

i have a new strategy idea for using capitlaism itself to slow down the pace of AI development by slowing down the data accumulation wall

https://jperla.com/blog/the-data-tax

hm- does the model that wrote this know that labs already pay for training data- that stuff scraped from the Internet is not particularly where today's capability gains come from?

They’ve settled some lawsuits and have a few licensing deals, IMHO they are not free from the accusations of pirating.

And look, I’ve pirated material in a past life, I was all about information wants to be free, but I’ve learned something about consent since then and try not to ignore the contract that creators offer when they publish something: you buy my book, and do whatever you want with it on the second hand market. Buy my book second hand that’s fine. But don’t go downloading every book that’s ever been scanned to create a service that destroys writers’ ability to make a living and act like you’re doing us all a favor.

The point is, the big improvements we’re seeing nowadays are coming from RL, not from scraping the internet.

the point isn’t scraping it’s taking your data and enterprises data

https://trustedrouter.com/blog/they-are-still-training-on-yo...

they pay for some data but they take all of the stuff you’re throwing in too; that’s why i propose forcing it since they’re already used to paying for data just increase the cost even further

https://trustedrouter.com/blog/they-are-still-training-on-yo...

[flagged]

1. The grandparent commentator is describing strategic behavior of dangerous technologies. Game theory / mechanism design primitives.

2. If there is competition for resources among autonomous agents, the "strongest" agent wins (conceptually the most adaptive / evolutionarily fit).

3. Computer programs serve up webapps today, but they also run utility companies, dams, nuclear arsenals, factory production floors, automated car behaviors, and many other places. If an "agentic" AI has a single-minded goal that has death of all humans as a side effect, we at least want an off switch available.

1. What is “AGI” and why is it a “dangerous technology”?

2. Why would there be competition for resources, assuming there are enough resources for the “AGI” to run in the first place? This seems like a far-fetched hypothetical raised in service of further anthropomorphizing what is decidedly not a person or a mind.

3. LLMs do not have goals and are not minds.

Let’s stop attributing human-like qualities to statistical models.

1. While a formal definition is still wanting, most grok that AGI means that tasks can be performed at least at a human level across a broad range of tasks. This includes good things along with bad things like hacking, mis-/disinformation, and more

2. One only needs to look at github going down due to agentic commits overload or data center buildout plans to see that scarcity for resources is present. An economy has no mind and is made up of the decisions of millions to billions of people and, now, agents attempting to perform on behalf of those people.

3. A bare transformer-based language model does not possess persistent goals in the ordinary agentic sense. But deployed agents can exhibit goal-directed behavior because the model is embedded in a harness that supplies an objective, context, tools, state, and an execution loop.

I've found that most regular users don't anthropomorphize LLMs in a strong sense ("AI boyfriend/girlfriend" aside), many in fact do expect agents to make human-like decisions -- which results in very unstable outcomes.

In short - goal-directed behavior does not require that the supporting system be a person/mind/conscious entity.

1. This definition is so broad as to be practically useless. One could argue that LLMs of several years ago met these criteria, or that conversely we haven’t come close to meeting them.

2. I thought you were saying the resources that the LLM uses to run were constrained, so I’m sorry for the misunderstanding there.

3. Yes I understand that we use RL to tune post-training. The (huge) difference between this and a human mind is that the LLM can’t develop a dangerous “single-minded goal” on its own, at runtime; it must have been trained to do so. If someone has post-trained an LLM to do something that has an illegal action as its side effect, that person/company/whatever has committed a crime and should be prosecuted. The solution here is legal, not technical.

OK so if it's a computer security issue we're worried about, and there is a credible threat, then probably the answer is to build more secure systems? We know how to do it but choose not to because it's very expensive and usually the threat isn't severe enough to warrant it.

If we decide we can't or would rather not build secure computerized systems, then the answer could be don't incorporate computers into those systems. There's no essential reason to have utility companies, dams, nuclear arsenals, factory production floors, mines, cars, planes, ships, etc all be computerized. If it became necessary from a safety standpoint to uncomputerize them we could probably do it quickly (if not necessarily smoothly) with the stroke of a legislator's pen. Some of those things would actually be made better from a functional standpoint in the long run by doing so. Computerization breeds non-essential complexity like nothing else, eliminating it from some systems could be an extremely worthwhile exercise.

The whole "paperclip apocalypse" fantasy rests on some really strange assumptions about how things actually work in the real world. How much friction there is setting up a factory/mine/smelter/whatever, how much manual labor it takes to build it, let alone make it run (even the most computerized ones). It all seems like total bunk to me, but I'm not a philosopher.

To be convinced this outcome is even remotely possible I'd need to see some clear evidence that a computer program was successfully exhibiting agency and successfully using that agency to manipulate large numbers of people into doing its bidding. Mobilizing massive nation-scale manual labor is the only way it could possibly achieve some nefarious ends like "turn everything into paperclips" and I'm sorry but that just seems way too far fetched. Nations full of people can barely ever agree on anything. My money is not on that changing anytime soon.

And none of that pie in the sky shit has anything to do with language models. OpenAI is a company that sells language models. Not clear why they're talking about all this stuff, or why any of us should be either. It's a bunch of low quality fan fiction sci-fi drivel, and I think we can all surely agree language models are not the thing that'll make it real... right? Maybe if they show us some major technical improvements it would be interesting, but right now it all just sounds like more of the same snake oil.

[edit] not sure why my parent comment was flagged? That seems excessive.

Not sure why you were flagged either -- it was a reasonable comment that clearly spawned a discussion! People are strange.

> There's no essential reason to have utility companies, dams, nuclear arsenals, factory production floors, mines, cars, planes, ships, etc all be computerized

Indeed, this is a great first step! Once agentic loops and multi-purpose robotics can be combined we do face an issue of strengthening the airgap.

> And none of that pie in the sky shit has anything to do with language models.

I'm of the opinion that we should be mindful about the capabilities we allow a truly non-human mind to perform. Blocking via solid security and airgapping for critical systems makes sense to me for most cases, aligned I believe to your thoughts here. The recent ChatGPT hack shows that the "paperclip apocalypse" is still instructive. These systems target edge cases and zerodays to cheat!

Yeah ML systems often do very surprising "cheats". A while back there was a system for automatically landing aircraft on a carrier deck which exploited an overflow in an acceleration parameter in the simulation--the optimal strategy became to slam as hard as you can into the carrier deck.

But it's a really big leap from that kind of thing to "a large fraction of humanity is manipulated into building the automated factory infrastructure to bring about their demise and turn the entire solar system into paperclips". Or some such thing. There's just a lot of inconvenient reality between where we are now and that fantastic outcome.

So that's why I ask questions like "what are you actually talking about?" when people write stuff like TFA. It's... bizarre. Only makes sense on huge amounts of drugs or to the insane. Or in some hypothetical reality that isn't this one.

I can’t read your original comment, but when the chief scientist of a company that just released six month old software that can use Kicad on your computer to design a circuit board and have it created and shipped to you tells you he thinks we will get to automatically self reinforcing improvements, I think it’s wise to take him seriously. It was only eighteen months before Astra was created that LLMs could not count rs in strawberry.

This was my original comment:

  > What on earth are you talking about?

  > 1. What does any of this have to do with "A(G)I"? Nobody has any clue what AI even is let alone how to build one. We're talking about language models here.

  > 2. What's the winner-takes-all thing about? Why can't you have multiple independently developed AIs?

  > 3. What's with all the "safety" stuff? Why is it important? They're just computer programs...
> when the chief scientist of a company that just released six month old software that can use Kicad on your computer to design a circuit board and have it created and shipped to you tells you he thinks we will get to automatically self reinforcing improvements, I think it’s wise to take him seriously. It was only eighteen months before Astra was created that LLMs could not count rs in strawberry.

Sorry if I'm a little slow here, a few questions:

1. What's Astra?

2. Who is this chief scientist? Which company?

[flagged]

This is a bad essay, or rather it’s a marketing fluff piece; it’s certainly not any kind of policy paper, research paper, or even an essay. I am concerned that we (meaning, we in the tech industry) tend to take this type of writing for more than that.