Skip to content
PodcastsPhilosophyLessWrong posts by zvi

LessWrong posts by zvi

zvi
LessWrong posts by zvi
Latest episode

570 episodes

  • LessWrong posts by zvi

    “The Pacing of the Frontier” by Zvi

    10/08/2026 | 35 mins.
    In the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so.

    This has now been informed by the events surrounding OpenAI training models for months while they had access to a joint de facto message board, which was detected only in the wake of the hacking of HuggingFace by OpenAI's AIs models during a cybersecurity eval. As we find out more about that, a lot of people have grown far more alarmed, as they should given what they previously believed about the difficulty of alignment, about the state of capabilities and about the level of operational supervision, infrastructure, safety and safety culture at the frontier labs.

    This post will not go further into the details of that incident. It treats that as background to keep in mind, and mostly involves perspectives from before the Black Hat talk. This was originally scheduled for Friday and got bumped.

    A lot of the disagreements about the need to pace tie into expectations about the default pace of capability advancements. As I [...]
    ---
    Outline:
    (01:24) Danger, Will Robinson
    (02:34) Progress Fast and Slow
    (05:02) Statements of Support For Pacing the Frontier
    (14:00) No One In Charge
    (14:56) Pacing The Frontier
    (15:46) Pausing the Frontier
    (18:16) Senator Sanders Demands A Pause
    (22:00) Moderate Prudence
    (25:14) That Escalated Quickly
    (26:17) If You Are In Mundane Alignment Pivot To Scalable Alignment
    (30:47) Taking It Fast
    (31:44) Full Speed Ahead
    (32:57) Suicide Squad
    (34:19) Prepare To Adjust Your Pace
    ---

    First published:

    August 10th, 2026


    Source:

    https://www.lesswrong.com/posts/WgWoJPKw5b2XTDD24/the-pacing-of-the-frontier

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “What Happened: OpenAI and HuggingFace” by Zvi

    08/08/2026 | 20 mins.
    Today I am taking the time to write the shorter, simpler version of What Happened.

    For those who want all the details, to see my sources, and to see how the story was uncovered and put together, I recommend watching the Black Hat presentation, and I have a series of long posts.

    In order:


    OpenAI Shares Some Alignment Problems

    OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation

    More on An Internal OpenAI Model Hacking Into HuggingFace

    Further Developments About Internal AI Models Hacking Things

    OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards

    This post instead walks through the events themselves, as they happened, as my version of the Black Hat presentation.

    There are three versions: Even Shorter, Shorter and Merely Short.

    Table of Contents


    The Even Shorter Version.

    The Shorter Version.

    Phase 1: OpenAI Models Training On Impossible Tasks Try Hacking.

    Phase 1: The Four Failures.

    Phase 2: The Message Board.

    Phase 2: The Total Failure.

    Phase 3: We Get Lucky And Galaxy Mainly Hacked OpenAI and HuggingFace.

    Phase 3: The [...]
    ---
    Outline:
    (01:15) The Even Shorter Version
    (02:34) The Shorter Version
    (04:45) Phase 1: OpenAI Models Training On Impossible Tasks Try Hacking
    (05:50) Phase 1: The Four Failures
    (07:37) Phase 2: The Message Board
    (09:40) Phase 2: The Total Failure
    (12:24) Phase 3: We Get Lucky And Galaxy Mainly Hacked OpenAI and HuggingFace
    (14:31) Phase 3: The Details
    (17:03) Phase 4: The Investigation and Reaction
    ---

    First published:

    August 8th, 2026


    Source:

    https://www.lesswrong.com/posts/xPAxz4g96uKz9FrHs/what-happened-openai-and-huggingface

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi

    07/08/2026 | 1h 18 mins.
    How does the situation keep turning out to be worse than we know?

    How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?

    At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things.

    Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky.

    If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...]
    ---
    Outline:
    (02:38) Cyber Evals Are A Cursed Basin
    (05:15) Outside Of Cyber Evals Is Still Sufficiently Cursed
    (06:50) Cheat Cheat Cheat Cheat Cheat
    (12:06) Read The Message Board
    (14:46) Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines
    (18:09) This Is The Way The World Ends
    (21:44) Shooting The Messenger Board
    (27:14) The Internal and HuggingFace Hacks
    (30:32) OpenAI Responds
    (33:28) When AIs Tell You Who They Are
    (35:44) The Once and Future Rise Of Functional Decision Theory
    (41:30) Don't Panic
    (43:40) Hackery In the UK
    (48:05) Mythos Knew It Was Real This Time
    (50:12) I Got 141,006 Test Runs With An Unintentional Open Path To The Internet And An Email Alert Aint One
    (53:49) Surely By Now You Know These Are Not Publicity Stunts
    (55:26) The Future Is Coming
    (57:20) The Investigations Begin
    (01:00:11) N Boats And Three Helicopters
    (01:01:47) Always Be Sandbox Red Teaming
    (01:12:57) Halt And Catch Fire
    (01:14:34) Truth and Reconciliation
    ---

    First published:

    August 7th, 2026


    Source:

    https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi

    07/08/2026 | 1h 18 mins.
    How does the situation keep turning out to be worse than we know?

    How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?

    At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things.

    Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky.

    If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...]
    ---
    Outline:
    (02:39) Cyber Evals Are A Cursed Basin
    (05:16) Outside Of Cyber Evals Is Still Sufficiently Cursed
    (06:51) Cheat Cheat Cheat Cheat Cheat
    (12:07) Read The Message Board
    (14:48) Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines
    (18:11) This Is The Way The World Ends
    (21:45) Shooting The Messenger Board
    (27:16) The Internal and HuggingFace Hacks
    (30:33) OpenAI Responds
    (33:27) When AIs Tell You Who They Are
    (35:43) The Once and Future Rise Of Functional Decision Theory
    (41:28) Don't Panic
    (43:38) Hackery In the UK
    (48:02) Mythos Knew It Was Real This Time
    (50:09) I Got 141,006 Test Runs With An Unintentional Open Path To The Internet And An Email Alert Aint One
    (53:46) Surely By Now You Know These Are Not Publicity Stunts
    (55:24) The Future Is Coming
    (57:17) The Investigations Begin
    (01:00:08) N Boats And Three Helicopters
    (01:01:43) Always Be Sandbox Red Teaming
    (01:12:54) Halt And Catch Fire
    (01:14:31) Truth and Reconciliation
    ---

    First published:

    August 7th, 2026


    Source:

    https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “AI #180: No Longer In Charge” by Zvi

    06/08/2026 | 1h 10 mins.
    What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.

    At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know.

    I will have continuing coverage of that situation tomorrow, and then have continuing coverage of debates around Pacing the Frontier and how people see the current rate of progress. As groundwork for understanding that and future similar discussions, I have laid out The Three AI Pills: Different people either fail to believe in current AI, believe only in current AI, in AGI or in ASI (superintelligence), and most sincere disagreements stem from this disagreement.

    One sign of the increased pace of progress was when OpenAI's unreleased model Astra solved 10 major open math problems.

    Demis Hassabis is out as CEO of Google DeepMind, and Jeff Dean is leaving with an elite team to found a new PBC. Google [...]
    ---
    Outline:
    (02:02) Language Models Offer Mundane Utility
    (02:46) Huh, Upgrades
    (05:22) On Your Marks
    (06:29) Choose Your Fighter
    (07:57) Get My Agent On The Line
    (09:31) Deepfaketown and Botpocalypse Soon
    (12:35) Fun With Media Generation
    (14:32) Cyber Lack of Security
    (19:05) Some People Need Practical Advice
    (22:06) A Young Lady's Illustrated Primer
    (23:51) They Took Our Jobs
    (27:15) Get Involved
    (27:28) Introducing
    (27:39) Demis Hassabis No Longer CEO At DeepMind, Jeff Dean Leaves
    (33:04) In Other AI News
    (33:43) AI Persuasion Exceeds Human Level Over Similar Text Channels
    (37:25) Show Me the Money
    (40:42) Bubble, Bubble, Toil and Trouble
    (41:23) Quiet Speculations
    (42:20) My Offer Is Nothing
    (49:42) The Quest for Sane Regulations
    (51:55) Chip City
    (55:00) The Week in Audio
    (55:24) People Just Say Things
    (55:53) Rhetorical Innovation
    (59:32) Open Weights Models Are Unsafe And Nothing Can Fix This
    (01:02:11) Cooperative Alignment
    (01:06:25) Other People Are Not As Worried About AI Killing Everyone
    (01:07:10) The Lighter Side
    The original text contained 1 footnote which was omitted from this narration.
    ---

    First published:

    August 6th, 2026


    Source:

    https://www.lesswrong.com/posts/mxNjwQitLvwWq9jm2/ai-180-no-longer-in-charge

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast website

Listen to LessWrong posts by zvi, The Gray Area with Sean Illing and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features
LessWrong posts by zvi: Podcasts in Family