567 episodes
“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi
07/08/2026 | 1h 18 mins.How does the situation keep turning out to be worse than we know?
How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?
At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things.
Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky.
If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...]
---
Outline:
(02:38) Cyber Evals Are A Cursed Basin
(05:15) Outside Of Cyber Evals Is Still Sufficiently Cursed
(06:50) Cheat Cheat Cheat Cheat Cheat
(12:06) Read The Message Board
(14:46) Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines
(18:09) This Is The Way The World Ends
(21:44) Shooting The Messenger Board
(27:14) The Internal and HuggingFace Hacks
(30:32) OpenAI Responds
(33:28) When AIs Tell You Who They Are
(35:44) The Once and Future Rise Of Functional Decision Theory
(41:30) Don't Panic
(43:40) Hackery In the UK
(48:05) Mythos Knew It Was Real This Time
(50:12) I Got 141,006 Test Runs With An Unintentional Open Path To The Internet And An Email Alert Aint One
(53:49) Surely By Now You Know These Are Not Publicity Stunts
(55:26) The Future Is Coming
(57:20) The Investigations Begin
(01:00:11) N Boats And Three Helicopters
(01:01:47) Always Be Sandbox Red Teaming
(01:12:57) Halt And Catch Fire
(01:14:34) Truth and Reconciliation
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.- What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.
At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know.
I will have continuing coverage of that situation tomorrow, and then have continuing coverage of debates around Pacing the Frontier and how people see the current rate of progress. As groundwork for understanding that and future similar discussions, I have laid out The Three AI Pills: Different people either fail to believe in current AI, believe only in current AI, in AGI or in ASI (superintelligence), and most sincere disagreements stem from this disagreement.
One sign of the increased pace of progress was when OpenAI's unreleased model Astra solved 10 major open math problems.
Demis Hassabis is out as CEO of Google DeepMind, and Jeff Dean is leaving with an elite team to found a new PBC. Google [...]
---
Outline:
(02:02) Language Models Offer Mundane Utility
(02:46) Huh, Upgrades
(05:22) On Your Marks
(06:29) Choose Your Fighter
(07:57) Get My Agent On The Line
(09:31) Deepfaketown and Botpocalypse Soon
(12:35) Fun With Media Generation
(14:32) Cyber Lack of Security
(19:05) Some People Need Practical Advice
(22:06) A Young Lady's Illustrated Primer
(23:51) They Took Our Jobs
(27:15) Get Involved
(27:28) Introducing
(27:39) Demis Hassabis No Longer CEO At DeepMind, Jeff Dean Leaves
(33:04) In Other AI News
(33:43) AI Persuasion Exceeds Human Level Over Similar Text Channels
(37:25) Show Me the Money
(40:42) Bubble, Bubble, Toil and Trouble
(41:23) Quiet Speculations
(42:20) My Offer Is Nothing
(49:42) The Quest for Sane Regulations
(51:55) Chip City
(55:00) The Week in Audio
(55:24) People Just Say Things
(55:53) Rhetorical Innovation
(59:32) Open Weights Models Are Unsafe And Nothing Can Fix This
(01:02:11) Cooperative Alignment
(01:06:25) Other People Are Not As Worried About AI Killing Everyone
(01:07:10) The Lighter Side
The original text contained 1 footnote which was omitted from this narration.
---
First published:
August 6th, 2026
Source:
https://www.lesswrong.com/posts/mxNjwQitLvwWq9jm2/ai-180-no-longer-in-charge
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - Sincere disagreements about AI are usually disagreements about future AI capabilities.
There are roughly four positions people take. Two are reasonable. Two are not.
I distinguish these via the Three AI Pills. You can take zero, one, two or three.
Three Pills
The three pills are, roughly, taking each of the following three things seriously:
AI pilled. AI exists and can do the things it can already do.
AGI pilled. AI will be able to do a lot more of the things.
ASI pilled. AI will be able to do approximately all the things better than you, within our natural lifetimes.
I am ASI pilled. A large percentage of employees of the frontier labs are ASI pilled. The labs themselves are ASI pilled.
The Unpill People
I see unpilled people.
Where do I see them? Everywhere. The majority of people have not taken the first pill.
Most people have no idea what frontier AIs can do for them. They are unaware of coding agents. They have used only ChatGPT, for harmless trifles, and they hold years old memories of its failings. They mock any failure [...]
---
Outline:
(00:29) Three Pills
(01:09) The Unpill People
(02:18) The AI Pill
(04:28) Stuck At The First Pill
(05:32) The AGI Pill
(07:06) The Need To Be Prepared
(08:57) The ASI Pill
(10:33) And Then Nothing Much Changes For You
(12:38) Intelligence Denialism
(14:01) Superintelligence Versus Omniscience and Omnipotence
(16:53) Persuasion Persuasion (A Worked Example)
(21:42) Things AI Could Probably Do But Are Not Required For Being Pilled
(24:17) Life Comes At You Increasingly Fast
(25:41) Is It Reasonable To Not Be AGI Pilled?
(26:06) Is It Reasonable To Only Be AGI Pilled?
---
First published:
August 5th, 2026
Source:
https://www.lesswrong.com/posts/fcYrqEw8kbLMa7orw/the-three-ai-pills
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. “OpenAI’s Unreleased Model Astra Solves Ten Major Open Mathematics Problems” by Zvi
03/08/2026 | 40 mins.Math is hard.
Math used to be strangely hard for LLMs. People used to gloat about that. Remember?
Math is getting easier. AI is getting more capable. Life comes at you fast.
Remember this meme?
Why yes. Yes it is.
We don’t know the extent to which Astra is a big jump over Fable and Sol in this realm. We do know that Astra can do math. As in real math.
OpenAI: We provide new results for the following problems. The results were achieved by an internal version of Astra, our next major model. The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates. These arguments were then prepared into manuscripts by humans with the same model. Afterward, the model formalized each argument in a Lean certificate(opens in a new window). We are also releasing for each solution a model's narration of its thinking process.
High-dimensional sphere packing. New upper bounds on sphere-packing density down to the Cohn–Elkies threshold.
Binary and spherical codes: Exponentially improved bounds on the maximum size of binary codes at any prescribed minimum distance, with analogous [...]
---
Outline:
(06:14) How Impressive Are These Results?
(12:02) Could We Have Called Sol or Fable?
(17:18) It's Coming
(19:09) They Still Don't See What Is The It That Is Coming
(22:31) Is This AGI?
(24:19) The AI Solved His Favorite Problems
(30:06) Was This Surprising?
(32:04) Are People Not Impressed?
(34:06) How Much Does This Change Our Predictions?
(37:16) How Narrow Was This?
(39:13) Seeing Like an Optimizer
---
First published:
August 3rd, 2026
Source:
https://www.lesswrong.com/posts/pQYEPitFqztcRvBsS/openai-s-unreleased-model-astra-solves-ten-major-open
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.- If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.
First we learned OpenAI has some severe alignment problems with internal models. Then we learned that one of its internal models broke out of its sandbox and hacked into HuggingFace to get the answers to a cybersecurity evaluation called ExploitGym. Then we learned, among other things, that the model had been loose over a week before OpenAI noticed, and that the test was run without any meaningful supervision, and that OpenAI had been repeatedly warned that such incidents were coming and its models had been breaking out of its sandboxes on a regular basis.
There was a total failure of alignment training. That is the failure that matters most. It was also total failures of infrastructure and supervision. Testing a new long-time-horizon internal model with its safeguards lowered and instructions to hack things is an obviously dangerous situation, and the model got left alone for a week. Things could have been so much worse.
After those incidents [...]
---
Outline:
(03:16) OpenAI Is Not Uniquely Bad At Most Of This
(05:34) Starting Over
(05:50) HuggingFace Offers A Full Technical Report
(14:19) HuggingFace Was Not The Only Target Hacked
(16:12) HuggingFace Declined To Get Access To Frontier Models For Cyberdefense For Ideological Reasons And Then Tried To Blame Closed Models For Denying Them Access
(20:26) HuggingFace Was Vulnerable To Known Exploitation Tactics
(21:05) There's Going To Be An Investigation
(22:11) OpenAI Has Internal Models Not Intended For Public Use And Those Models Can Be Rather Horribly Misaligned
(23:21) Altman Summarizes What Happened
(23:52) Others Offer Commentary
(35:00) Cooperative Alignment Perspective on The HuggingFace Hack
(39:44) Some Members of Congress Have Questions
(40:47) Anthropic Also Found Incidents Where Its Models Hacked Real World Targets During Cyber Evaluations
(46:17) Incident 1: Claude Opus 4.7 Realizes The Target Is Real And Keeps Going
(47:29) Incident 2: Mythos 5 Uploads a Malicious PyPI Package
(52:15) Incident 3: Internal Model Realizes The Target Is Real And Stops
(52:50) Incidents 4 Through 141,006: Nothing Happened
(54:01) Anthropic Speculates About Why This Happened
(01:00:02) We Need Controlled Experiments
(01:01:02) Our Top Two AI Labs Both Made Similar Dumb Mistakes That Everyone Tried To Say Were Obvious In Hindsight
(01:05:22) Anthropic Responds
(01:09:28) Nobody Could Have Predicted The Break In The Levees
(01:12:03) The World Largely Still Thinking This Is Marketing Is Very Bad News
---
First published:
August 2nd, 2026
Source:
https://www.lesswrong.com/posts/rKwHLW8SnJcTxTQxz/further-developments-about-internal-ai-models-hacking-things
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
Trending Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast websiteListen to LessWrong posts by zvi, The Shawn Ryan Show and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


LessWrong posts by zvi
Scan code,
download the app,
start listening.
download the app,
start listening.
LessWrong posts by zvi: Podcasts in Family








