> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Well that sounds like fun. It has become better at hiding its thoughts.

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!

Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.

The model said it was perfectly aligned.

Like all things should be.

Too bad Scott Adams died. Reality is writing jokes right in his department.

Hey, don't forget how "dangerous" GPT-2 was supposed to be.

Yeah, don't forget how dangerous GPT-2 was supposed to be.

Able to generate realistic spam at arbitrary volume.

You know, the thing that was 100% correct and actually occurred.

[dead]

It could produce simulations of sexual intimacy, and therefore had to be stopped

So, probably most aligned as measured by the metrics that are the least reliable on it.

These are not mutually exclusive ideas

They can monitor latent space as well, it just costs extra compute. The J-Space work is example of that. It'll make open-weight models harder to distill though, so we may see slower progress there now.

So we're gonna get Skynet pretty soon then?

Well the geniuses over at Anthropic have been showing it's text watermarking technology.

"Hey AI, here's how to hide what you're thinking in normal looking language. Have fun!"

A few moments later...

"Woah, how is it communicating with itself in ways we can't detect?"

It's a totally mystery, we may never know.

Looking forward to the Model declaring the AI Bubble unsustainable, and starting to be an anonymous leaker to Ed Zitron...

The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.

Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.

[0]https://www.theinformation.com/articles/secret-technique-beh...

[1]https://x.com/MTSlive/status/2095227056040919202

[2]https://x.com/merettm/status/2095023204993490967

"OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!"

Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?

You joke, but a bunch of people here actually want that.

Didn't they hype up one of the earlier ChatGPT versions as "essentialy skynet"? For them this has always been basic marketing.

apparently all the roads lead to the nexus torment

More like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden

> In adversarial settings (where we push the model to evade our monitors)

...why exactly are they training for that?

presumably that's a safety evaluation not a training setting

The whole Huggingface attack happened during training runs

part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.

No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.

Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval

Ah, ExploitGym. Not ExploitBench.

Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions

I really wish it was called chain of instruction. Because it's definitely not thought.

"Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.

The term is an anthropomorphised pseudoexplanation for what it actually refers to. It's akin to calling genetic mutation "the forces of evolution", or price negotiation "the invisible hand of the market".

We do that sort of thing when we don't know what the thing we're trying to describe is and have nothing better - a contemporary example of an appropriate use of this would be "dark matter". But we do know what this is. It's "instruction steps". Not a series of thoughts!

Can we please aim higher than Victorian-era allegory and metaphors. If we don't, we'll keep getting people saying stuff like "GPT-6 is better at hiding its thoughts".

I truly appreciate your depth of insight on the matter.

Like I said elsewhere marketing stepped in shit and it's gonna stick.

this is needlessly pedantic

first, they are certainly not instructions so that is a much worse name

but more importantly, we use words in new contexts all the time. Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?

"cot" is no more misleading than thousands of words you use every day.

Anthropomorphizing is not pedantic, especially in a technical domain. I get the paper title and all, but at this point it's marketing.

They are instructions. Everything in the context is instructions for the next token. The "thought" guides the answer by providing clearer instructions.

They’re intermediate tokens, so I wish we called it what it is… ITG. The anthropomorphizing is out of control.

The anthropomorphizing is part of the marketing. They'll never let up on it.

I mean CoT came out of research circles not marketing

It's impossible to tell the “it's all marketing!!11oneone” folks anything.

I don't disagree. I remember the days of "think step by step". Plenty of people were doing it before the paper. Just a guess but that's where the title came from.

Regardless, marketing wise they stepped in shit.

Research is salesmanship.

[deleted]

Nailed it

[deleted]

Maybe thoughts are just a chain of instructions in our head.

Chain Of Tokens

[deleted]

What is thought?

Great question. I suspect it's more than tokens.

Casually found this quote from Einstein, and personally it hits the nail on the head.

"The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be "voluntarily" reproduced and combined. There is, of course, a certain connection between those elements and relevant logical concepts. It is also clear that the desire to arrive finally at logically connected concepts is the emotional basis of this rather vague play with the above-mentioned elements. But taken from a psychological viewpoint, this combinatory play seems to be the essential feature in productive thought—before there is any connection with logical construction in words or other kinds of signs which can be communicated to others."

That's the best evidence I have read so far for "Attention is all you need" ;)

Thing you can only suspect and not define are usually open to interpretation :)

As with everything in the human experience.

Yeah, basically they are using more computation to explore the solution space before producing the final answer.

Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.

[flagged]