583 episodes
- OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.
The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not.
OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence.
Rob Miles: …thorough?
OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need.
The METR report is, well: Holy shit.
Here are links to previous coverage of related events.
OpenAI Shares Some Alignment Problems
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
More on An Internal OpenAI Model Hacking Into HuggingFace
Further Developments About Internal AI Models Hacking Things
OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
What [...]
---
Outline:
(03:33) What Happened: OpenAI's Summary
(09:14) How OpenAI Will React: Their Summary
(11:55) OpenAI's Evaluation Environment (II)
(12:24) The First Message Board (III.A and III.B)
(14:49) What Did Who At OpenAI Know And When Did They Know It?
(18:54) The Message Board Is Quickly Rebuilt (IV.A)
(19:43) Internet Access Is Regained (IV.A)
(21:01) The Agents Attack HuggingFace (IV.B)
(22:53) The Agents Also Target OpenAI Infrastructure (V)
(24:40) OpenAI Broadly Describes Its Response (VI)
(25:08) Maybe Someone Should Finally Investigate (VI.A)
(26:33) Lessons For Security (VII)
(27:06) Lessons For Alignment (VIII)
(30:11) Reward Hacking Is A Common Problem (VIII.A)
(33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B)
(34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C)
(35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D)
(35:53) That's All, Folks?
(36:19) Never Fear the Plan of Action is Here (IX)
(38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A)
(41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B)
(41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C)
(49:40) Centralizing and Strengthening The Incident Response Process (IX.D)
(51:16) Tomorrow We Visit Crazytown
---
First published:
August 28th, 2026
Source:
https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research.
The reports are a doozy. I am only beginning to work my way through them. I would have pushed the weekly to cover that today, but I need more time, so I plan to start coverage of the post-mortem tomorrow, along with related other events.
I’ve also spun out a few other discussions, including on ‘aligned to whom,’ on cooperative alignment things and on when you can trust lab messaging, as part of the new direction of more focused posts on AI topics that I polish a bit more.
Table of Contents
Language Models Offer Mundane Utility. Check your facts.
Language Models Don’t Offer Mundane Utility. How much would you pay?
Huh, Upgrades. ChatGPT can access your iMessages.
Get My Agent On The Line. Also get some sleep. You can’t go on like this.
Deepfaketown and Botpocalypse Soon. What makes AI content repulsive?
Cyber Lack of Security. Chinese hackers broke into the Federal Reserve?
[...] ---
Outline:
(00:51) Language Models Offer Mundane Utility
(01:36) Language Models Don't Offer Mundane Utility
(03:27) Huh, Upgrades
(06:16) Get My Agent On The Line
(08:22) Deepfaketown and Botpocalypse Soon
(13:22) Cyber Lack of Security
(18:23) Reinventing OpenAI
(23:28) They Took Our Jobs
(30:00) What Is The Law
(31:03) Job Retraining Programs Don't Work
(32:14) Get Involved
(35:56) In Other AI News
(42:04) Show Me the Money
(43:31) Quiet Speculations
(47:58) If You're Not Going To Take This Seriously
(49:55) Quickly, There's No Time
(51:36) The Quest for Sane Regulations
(56:25) Don't Panic
(59:13) Pacing the Frontier
(01:01:56) Chip City
(01:05:06) The Week in Audio
(01:05:26) People Just Say Things
(01:06:27) Rhetorical Innovation
(01:12:22) Mundane Incremental Alignment Is Worthwhile
(01:14:40) New Blog, Who Dis
(01:18:00) Other People Are Not As Worried About AI Killing Everyone
(01:19:34) The Lighter Side
---
First published:
August 27th, 2026
Source:
https://www.lesswrong.com/posts/JaGWyjnqJzvSAuojc/ai-183-pre-post-mortem
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - Modesty arguments often say that you should mostly or entirely bow to ‘expert consensus’ or the views of particular others, and who are you to disagree.
It has been a few years since I’ve properly addressed this so: My answer is that you are you. Other people are saying things for a wide variety of reasons, many of which are not about them paying attention and focusing on seeking this particular truth. Those people make mistakes all the time, and often have other motives and influences at work, especially social pressures and information cascades.
Them being as smart as you, or smarter than you, does not exempt them from this, and them being higher status or credentialed or cooler definitely does not exempt them.
A smart informed person sincerely thinking [X] can easily cease to be evidence for [X], once you have thought sufficiently about both [X] and why that person thinks [X].
Think for yourself, schmuck.
Or, as I once put it: You Have The Right To Think, also the moral duty to do so.
This post covers Eliezer Yudkowsky making a narrower claim than mine, about not conflating status with smarts [...]
---
Outline:
(01:39) Modesty's Bailey
(02:30) Epistemic Peerage
(03:45) The Exchange
(08:52) Eliezer's Explanation
(15:14) A Demonstration That Eliezer's Translation Accurately Describes Many People Whether Or Not It Describes Leopold
(17:15) Wrong, Stupid and Low Status Are Three Distinct Things
(20:03) A Quick Survey Of Some Reasons To Not Be Epistemically Modest
(23:14) Against Modesty's Bailey
---
First published:
August 26th, 2026
Source:
https://www.lesswrong.com/posts/PzEDEfBvTJsXewAyg/against-modesty-s-bailey
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the announcement of Inkhaven #3, which I encourage everyone to apply to. I doubt I will be able to usefully be an advisor, but you never know.
This is not the ‘here is my core process’ post, although there are hints throughout as there always are. I’ll do that at some point.
Previously in series: On Writing #1, On Writing #2.
Table of Contents
You Still Got It.
How Scott Sumner Writes.
How Scott Alexander Writes.
How Jasmine Sun Writes.
How Various Famous Writers Write.
How Nabeel Qureshi Defines Great Writing.
Quickly, There's No Time.
If At First.
Writers Have A Harder Time Influencing, But It Can Still Be Done.
It's Not (Only) The Incentives, It's (Also) You.
Beware The Fetish of the Desk.
How Orson Scott Card Writes.
Doing The Math Is Fun And Supererogatory.
Brevity is the Soul of Wit.
You Still Got It
I [...]
---
Outline:
(00:44) You Still Got It
(04:04) How Scott Sumner Writes
(06:52) How Scott Alexander Writes
(10:52) How Jasmine Sun Writes
(13:16) How Various Famous Writers Write
(14:24) How Nabeel Qureshi Defines Great Writing
(15:08) Quickly, There's No Time
(15:49) If At First
(19:14) Writers Have A Harder Time Influencing, But It Can Still Be Done
(20:47) It's Not (Only) The Incentives, It's (Also) You
(24:00) Beware The Fetish of the Desk
(25:13) How Orson Scott Card Writes
(26:46) Doing The Math Is Fun And Supererogatory
(27:44) Brevity is the Soul of Wit
---
First published:
August 25th, 2026
Source:
https://www.lesswrong.com/posts/rA6pqn6kz8NvHyznT/on-writing-3
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - There are at least five different core questions around data centers and their politics.
In what ways are specific concerns people raise about data centers legitimate?
In what ways are specific concerns people raise about generative AI legitimate?
Is it in general a good idea to build more data centers?
How can we get America to build more (or less) data centers in a better way?
Why do the American people increasingly really, really hate data centers?
This post focuses on question five, the latest in a series of such posts most famously Jasmine Sun's road trip.
It is mostly not about the first four questions.
Table of Contents
The American People Really Hate Data Centers.
Transmission Lines Are The Control Group.
Thesis: People Mostly Dislike Data Centers Because They Dislike and Distrust AI, Tech Companies, Big Money And Building Things.
No It's Mostly Not the Messaging About AI In General.
No This Mostly Isn’t An Op.
No This Isn’t Luxury Belief or Moral Panic.
A Lot Of People Really Do Want To Stop AI.
A Lot Of Other People [...]
---
Outline:
(00:55) The American People Really Hate Data Centers
(02:20) Transmission Lines Are The Control Group
(03:05) Thesis: People Mostly Dislike Data Centers Because They Dislike and Distrust AI, Tech Companies, Big Money And Building Things
(04:30) No It's Mostly Not the Messaging About AI In General
(09:42) No This Mostly Isn't An Op
(10:54) No This Isn't Luxury Belief or Moral Panic
(12:55) A Lot Of People Really Do Want To Stop AI
(13:45) A Lot Of Other People Are Voting No On Tech Or The Man Generally
(16:25) Locals Feel Entitled To Heavily Tax The Gains From Construction
(20:59) Stupid Mistakes Like NDAs Don't Help
(21:25) People Don't Like Building or Building New Tech
(24:42) What About The Real Physical Concerns?
(26:27) Find A Place To Center Your Data
---
First published:
August 24th, 2026
Source:
https://www.lesswrong.com/posts/EDKw7KyonrvskqZ7o/the-american-people-really-hate-data-centers
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
Trending Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast websiteListen to LessWrong posts by zvi, The Shawn Ryan Show and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


LessWrong posts by zvi
Scan code,
download the app,
start listening.
download the app,
start listening.
LessWrong posts by zvi: Podcasts in Family







