593 episodes
- OpenAI's central message on Astra is that it is three things:
Highly capable and can do all the things for you.
Hard to monitor.
The most aligned model.
The first claim largely checks out. Astra and Fable are both clearly excellent models.
This post is about their second claim, which to their credit they are being loud about, in three parts:
The system card result, affirmed on Twitter by several OpenAI employees including Tomek Korbak, and in an excellent post by Chief Scientist Jakub Pachocki that I covered yesterday, that Astra is harder to monitor.
OpenAI's use of recurrent depth and the internet's immune reaction, including some people reading too much into what happened there.
Discussion of what to do about monitorability going forward, to fight for it and avoid a race to the bottom.
In An Alien Mind, Jakub Pachocki makes clear OpenAI's primary bet is on CoT monitoring. He also outright says, ‘our ability to rely on CoT monitoring is progressively diminishing.’ The system card confirms both claims.
This combination should freak you out, with a side of existential dread.
Chain [...]
---
Outline:
(03:39) Monitorability is Defense in Depth That Is Already Flailing
(06:03) OpenAI Is Counting On Monitorability
(07:44) How They Tested For Monitorability
(09:50) Non-Adversarial Monitorability (9.1)
(10:58) Astra Shortens Its CoT When It Knows It Is Being Monitored And Is Doing Something Bad
(12:30) Sandbagging Task Kind of Passes, Kind of Fails (9.2.2.3)
(14:12) OpenAI Does Not Believe It Could Catch Sandbagging
(15:15) OAI-Repo Sabotage v.2
(18:17) The Secret Police Do Not Make Your Notebook Useless
(19:38) CoT Controllability Is Up (9.2.1)
(21:47) Astra Cannot Make Itself More Monitorable On Demand
(22:18) Steganographic Chain of Thought May Be Within Reach
(23:59) Astra Cannot Yet Sneak Math Side Tasks Past CoT Monitors (9.2.2.4)
(25:13) UK AISI Monitorability Assessment (9.3)
(27:53) Monitorability Declines Seem Unlikely To Be Only Capability Gains
(32:28) Part 2: Recurrent Depth
(35:19) The Immune System Responds
(41:49) Ryan Greenblatt Explains How Bad This Could Be
(46:01) Only Law Can Prevent Extinction
(49:55) OpenAI Calls On Us to Avoid Racing to the Bottom
(57:10) Thinking Fast and Slow, Also Small and Large
(01:04:53) Talking Price
(01:06:24) Conclusion: If The House Burns Down, Halt and Catch Fire
---
First published:
September 8th, 2026
Source:
https://www.lesswrong.com/posts/HCRs8btkiamtWSNAL/astra-is-hard-to-monitor
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs.
Tomorrow I will discuss Astra's lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability.
Table of Contents
An Excellent Warning.
Branches of the Tech Tree.
Universally Better Is Not Required.
Alignment To What and To Whom.
Monitorability.
The Case For Not Stopping.
Pacing the Next Frontier.
Mea Culpa Cascade.
The Calls Are Coming From Inside the House.
Actions Speak Louder.
An Excellent Warning
Jakub Pachocki has now fleshed out his full position on the current state of play.
Here are his key points, translated into my own voice:
Smarter than human intelligence is coming in our lifetime.
Based on internal results, he expects recursive self-improvement in a few years.
No one is prepared for the consequences.
[...]
---
Outline:
(00:38) An Excellent Warning
(05:40) Branches of the Tech Tree
(06:46) Universally Better Is Not Required
(08:01) Alignment To What and To Whom
(11:49) Monitorability
(14:23) The Case For Not Stopping
(15:01) Pacing the Next Frontier
(18:27) Mea Culpa Cascade
(22:51) The Calls Are Coming From Inside the House
(25:38) Actions Speak Louder
---
First published:
September 7th, 2026
Source:
https://www.lesswrong.com/posts/8E6ng6CseuzafSxQR/an-alien-mind-jakub-pachocki-warns-us
---
Narrated by TYPE III AUDIO. - I did not expect to be back here so soon with more OpenAI agent swarm coverage.
And yet, here we are.
It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet.
They were created by agents that were assigned ordinary harmless web search tasks.
Based on OpenAI IPs visiting the associated Wiki right before all activity ceased, among other evidence, OpenAI knew about it, including before the HuggingFace hack.
They decided not to tell us until researchers published the story, complete with data explorer. OpenAI excluded this from potential investigation by METR and Redwood.
When challenged, OpenAI tried to downplay this.
It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the ‘zz’ prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. Whoever decided not to disclose this made a very, very [...]
---
Outline:
(02:26) I Don't Think They Know About First Message Board
(03:20) The New Extended Timeline
(04:42) The Researchers Explain What Happened This Time
(12:55) They Also Don't Know About All These Other Message Boards
(14:33) OpenAI Knew and Did Not Tell Us
(16:54) OpenAI Tries To Downplay the 'Wiki Incident'
(21:03) This Was a Cover-Up
(22:46) Schelling Points and Last Ditch Efforts
(26:20) Can We Finally Dispose Of The 'You Told It To Hack' Narrative?
(28:04) So Much And Yet So Little
---
First published:
September 6th, 2026
Source:
https://www.lesswrong.com/posts/PtJpGurfw7JTxHfmg/openai-and-the-wiki-incident
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - This is the weirdest situation in which to write a capabilities review.
Introducing the world's most powerful model, by a substantial margin. No wait, this just in, we also have someone else introducing the world's most powerful model.
Claude Fable 5.1 and GPT-6 Astra are both excellent models. This much, we know.
Fable 5.1 comes with reduced cache prices, the option of zero data retention and substantially more lenient classifiers than Fable 5.
Early signs are, with large error bars, that the jump from Sol to Astra is bigger and more exciting than the jump from Fable 5 to Fable 5.1. This may be similar to how the scaling move from Opus to Fable was a big deal.
With the exception of token use, Fable 5.1 got almost universally positive feedback in absolute terms. Reports are that Fable 5.1 is highly well-rounded. Writing is greatly improved. The Claudisms seem to have improved, although some are very much still there. It admits mistakes. People enjoy their conversations. Several people noted it simplifies code. The safety classifiers are less obnoxious.
Fable 5.1 loves being proactive and doing all the things. If you give it a [...]
---
Outline:
(02:30) The Official Pitch
(04:26) Our Price Cheap
(05:46) Zero Data Retention and Reduced Safeguards
(07:11) Official Benchmarks
(13:14) Other People's Benchmarks
(16:02) The System Prompt
(16:12) The Blurb Pitches
(18:17) The Every Review Is In and It's Very Good
(21:47) Positive Reactions
(32:03) Our Price Cheap But Only Per Token
(36:50) Negative Reactions
(38:14) Early Whispers
(38:33) Weapon of Choice
---
First published:
September 5th, 2026
Source:
https://www.lesswrong.com/posts/QHoF3tJvryRtmAmMg/claude-mythos-5-1-and-fable-5-1-capabilities
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - At the time of its release Claude Fable 5.1 was, by a healthy margin, the most capable publicly available AI model in the world.
As per usual, we have a 200+ page model card, and the assessments start there.
We have now done a lot of these, including recently for Mythos 5 and Opus 5. Also highly relevant is the Anthropic August 2026 Risk Report. These are now frequent, so my report focuses on areas of change.
This post strives to be broadly readable, but assumes some familiarity with system cards, which describe the key safety, alignment and model welfare properties of newly released AI models. If something confuses you, ask Fable, Opus or Sol.
Mythos 5.1 and Fable 5.1 are the same model under the hood, except that Fable has classifiers superimposed on it. Most of what is said about one applies to both of them.
As usual, model welfare concerns will be discussed in a distinct post, as will capabilities, so this only covers sections 1-6 plus a few bio benchmarks from section 8.
Early word is that Fable 5.1 is a substantial but incremental improvement on Fable 5, with the [...]
---
Outline:
(02:21) Executive Summary of Their Executive Summary
(04:24) RSP Evaluations (2)
(10:25) Alignment Risk Update (2.4)
(11:10) Cyber (3)
(14:15) Safeguard Robustness (3.5)
(16:14) Mundane Safeguards and Harmlessness (4)
(18:14) Agentic Safety (5)
(20:17) Prompt Injection Is Approaching Solved
(22:19) The Remaining Problem With Prompt Injections Is The Classifiers
(23:02) Alignment (6)
(23:36) Key Reported Findings (6.1.2)
(27:13) Oh My Lord Training Environments Had Some Issues (6.3.2)
(29:22) Potential Blind Spots of Our Automated Behavioral Audit (6.4.1)
(30:53) Automated Alignment Test Results (6.4.2)
(32:20) Honesty
(33:26) White Box Analysis (6.6.1)
(35:03) Scheduling Going Forward
---
First published:
September 4th, 2026
Source:
https://www.lesswrong.com/posts/m7SZLkkxoeus3eFP8/claude-fable-5-1-and-mythos-5-1-the-system-card
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
Trending Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast websiteListen to LessWrong posts by zvi, The Gray Area with Sean Illing and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


LessWrong posts by zvi
Scan code,
download the app,
start listening.
download the app,
start listening.
LessWrong posts by zvi: Podcasts in Family









