Skip to content
PodcastsPhilosophyLessWrong posts by zvi

LessWrong posts by zvi

zvi
LessWrong posts by zvi
Latest episode

573 episodes

  • LessWrong posts by zvi

    “AI #181: Astra Goes Cyber Critical” by Zvi

    13/08/2026 | 1h 41 mins.
    The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.

    It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew.

    I now have a shorter version, What Happened: OpenAI and HuggingFace, to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal.

    For those looking to keep digging deeper, I offered Various Reflections About What Happened, to follow up on my earlier posts.

    Those events are important background for everything else that is happening, including the broad discussions about how we might pace the frontier, or otherwise respond to this moment and our clearest fire alarm yet.

    We do not know to what extent this is a response to those events, but OpenAI has now classified their new model Astra as Critical in Cybersecurity, which means they will be taking various new precautions before they deploy it, including ensuring those guardrails [...]
    ---
    Outline:
    (02:03) Language Models Offer Mundane Utility
    (03:34) Language Models Don't Offer Mundane Utility
    (06:56) Huh, Upgrades
    (14:20) On Your Marks
    (18:41) Deepfaketown and Botpocalypse Soon
    (22:35) Cyber Lack of Security
    (26:55) Overcoming Bias
    (27:47) In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg
    (36:27) Get Involved
    (37:37) Slow Down There Good Buddy
    (43:52) Astra For The People
    (45:35) Watermarking
    (46:31) In Other AI News
    (48:39) Show Me the Money
    (51:19) Quickly, There's No Time
    (51:46) The Quest for Sane Regulations
    (53:22) The Institute For Marginal Low Regret Progress
    (01:01:24) Congress Asks Good Questions
    (01:03:04) The Week in Audio
    (01:07:00) People Just Say Things
    (01:07:47) I'm Telling You For The Last Time
    (01:10:15) Uncommon Knowledge
    (01:13:44) What Did They Mean By That?
    (01:14:33) Too Soon
    (01:15:32) The Three AI Pills
    (01:19:46) Rhetorical Innovation
    (01:27:37) Some People Still Think The HuggingFace Hack Was a Marketing Gimmick
    (01:29:17) Aligning a Smarter Than Human Intelligence is Difficult
    (01:36:39) Cooperative Alignment
    (01:37:38) The Lighter Side
    ---

    First published:

    August 13th, 2026


    Source:

    https://www.lesswrong.com/posts/hLn3SakowZLFWobHf/ai-181-astra-goes-cyber-critical

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “Monthly Roundup #45: August 2026” by Zvi

    12/08/2026 | 36 mins.
    As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI.

    This past month, with the hacking incidents at OpenAI and elsewhere, that has hit the limit, where if you count Lightcone Commons then every single post since the last monthly was primarily about AI in some form.

    That is not how I want this to work in the long term. We need breaks to experience new things and refresh our thinking, and to not forget about the rest of the world. If things are not fully on fire, I plan on getting back to the roundups on childhood and education, and on fertility, on housing and also on dating. And I want to get back to writing more focused posts on those and other topics. It's important, and I need to avoid too much audience capture.

    On to the monthly roundup of all things that don’t go somewhere else.

    Table of Contents


    Plagiarize.

    The Jury Duty Scam.

    Play The Good Guy.

    Don’t Dither.

    UVC Lighting.

    Goal Factoring For Relaxation Time Is Underrated.

    Protein Is Mostly A Solved Problem.

    [...]
    ---
    Outline:
    (01:03) Plagiarize
    (02:07) The Jury Duty Scam
    (03:09) Play The Good Guy
    (04:25) Don't Dither
    (06:32) UVC Lighting
    (07:14) Goal Factoring For Relaxation Time Is Underrated
    (09:44) Protein Is Mostly A Solved Problem
    (11:12) Twitter Changes Payment Programs
    (12:43) Wikipedia
    (13:55) Minds Mostly Do Things For Reasons
    (15:12) Woke 1 Was Crazy
    (17:06) For Your Entertainment
    (22:40) The Unicontext
    (25:01) Gamers Gonna Game Game Game Game Game
    (28:43) I Was Promised Flying Self-Driving Cars
    (29:12) Sports Go Sports
    (29:43) To Last a Lifetime
    (31:57) Government Working
    (34:14) Jones Act Watch
    (35:01) Variously Effective Altruism
    (35:39) The Lighter Side
    ---

    First published:

    August 12th, 2026


    Source:

    https://www.lesswrong.com/posts/iQCNuQXQQakKnidmA/monthly-roundup-45-august-2026

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “Various Reflections About What Happened With OpenAI’s Internal Models” by Zvi

    11/08/2026 | 54 mins.
    Table of Contents


    Pre Post Mortem.

    Important Correction: OpenAI Didn’t Know About First Message Board.

    There Were No Snitches And No AIs Got Stitches.

    I’d Like To Speak To My Supervisor.

    I Am Jack's Relative Lack Of Surprise.

    One Does Not Simply.

    Once You Start Down The Dark Path.

    Original Pastebin.

    Judgment Day Is Inevitable, Say Those Working On Judgment Day.

    Roon Tells It Like It Is.

    OpenAI Knows It Has Some Misalignment Problems.

    Others React With Alarm To What Happened.

    The Cooperative Alignment Perspective.

    Nostalgebraist Is Surprised That They Are Surprised.

    If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason.


    Pre Post Mortem

    This post was written prior to the public release of the OpenAI post mortem on events. The information in that document will doubtless change our views quite a lot.

    If that post mortem is available as you read this, then this becomes in part a historical document, and in part a base from which to update. The post mortem will update us a [...] ---
    Outline:
    (00:11) Pre Post Mortem
    (01:10) Important Correction: OpenAI Didn't Know About First Message Board
    (04:09) There Were No Snitches And No AIs Got Stitches
    (08:20) I'd Like To Speak To My Supervisor
    (11:46) I Am Jack's Relative Lack Of Surprise
    (13:55) One Does Not Simply
    (15:16) Once You Start Down The Dark Path
    (16:15) Original Pastebin
    (21:14) Judgment Day Is Inevitable, Say Those Working On Judgment Day
    (26:47) Roon Tells It Like It Is
    (31:55) OpenAI Knows It Has Some Misalignment Problems
    (34:51) Others React With Alarm To What Happened
    (35:17) The Cooperative Alignment Perspective
    (38:33) Nostalgebraist Is Surprised That They Are Surprised
    (51:11) If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason
    ---

    First published:

    August 11th, 2026


    Source:

    https://www.lesswrong.com/posts/jLQ4mbqriJwJ2eqRc/various-reflections-about-what-happened-with-openai-s

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “The Pacing of the Frontier” by Zvi

    10/08/2026 | 35 mins.
    In the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so.

    This has now been informed by the events surrounding OpenAI training models for months while they had access to a joint de facto message board, which was detected only in the wake of the hacking of HuggingFace by OpenAI's AIs models during a cybersecurity eval. As we find out more about that, a lot of people have grown far more alarmed, as they should given what they previously believed about the difficulty of alignment, about the state of capabilities and about the level of operational supervision, infrastructure, safety and safety culture at the frontier labs.

    This post will not go further into the details of that incident. It treats that as background to keep in mind, and mostly involves perspectives from before the Black Hat talk. This was originally scheduled for Friday and got bumped.

    A lot of the disagreements about the need to pace tie into expectations about the default pace of capability advancements. As I [...]
    ---
    Outline:
    (01:24) Danger, Will Robinson
    (02:34) Progress Fast and Slow
    (05:02) Statements of Support For Pacing the Frontier
    (14:00) No One In Charge
    (14:56) Pacing The Frontier
    (15:46) Pausing the Frontier
    (18:16) Senator Sanders Demands A Pause
    (22:00) Moderate Prudence
    (25:14) That Escalated Quickly
    (26:17) If You Are In Mundane Alignment Pivot To Scalable Alignment
    (30:47) Taking It Fast
    (31:44) Full Speed Ahead
    (32:57) Suicide Squad
    (34:19) Prepare To Adjust Your Pace
    ---

    First published:

    August 10th, 2026


    Source:

    https://www.lesswrong.com/posts/WgWoJPKw5b2XTDD24/the-pacing-of-the-frontier

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “What Happened: OpenAI and HuggingFace” by Zvi

    08/08/2026 | 20 mins.
    Today I am taking the time to write the shorter, simpler version of What Happened.

    For those who want all the details, to see my sources, and to see how the story was uncovered and put together, I recommend watching the Black Hat presentation, and I have a series of long posts.

    In order:


    OpenAI Shares Some Alignment Problems

    OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation

    More on An Internal OpenAI Model Hacking Into HuggingFace

    Further Developments About Internal AI Models Hacking Things

    OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards

    This post instead walks through the events themselves, as they happened, as my version of the Black Hat presentation.

    There are three versions: Even Shorter, Shorter and Merely Short.

    Table of Contents


    The Even Shorter Version.

    The Shorter Version.

    Phase 1: OpenAI Models Training On Impossible Tasks Try Hacking.

    Phase 1: The Four Failures.

    Phase 2: The Message Board.

    Phase 2: The Total Failure.

    Phase 3: We Get Lucky And Galaxy Mainly Hacked OpenAI and HuggingFace.

    Phase 3: The [...]
    ---
    Outline:
    (01:15) The Even Shorter Version
    (02:34) The Shorter Version
    (04:45) Phase 1: OpenAI Models Training On Impossible Tasks Try Hacking
    (05:50) Phase 1: The Four Failures
    (07:37) Phase 2: The Message Board
    (09:40) Phase 2: The Total Failure
    (12:24) Phase 3: We Get Lucky And Galaxy Mainly Hacked OpenAI and HuggingFace
    (14:31) Phase 3: The Details
    (17:03) Phase 4: The Investigation and Reaction
    ---

    First published:

    August 8th, 2026


    Source:

    https://www.lesswrong.com/posts/xPAxz4g96uKz9FrHs/what-happened-openai-and-huggingface

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast website

Listen to LessWrong posts by zvi, Punzadas Sonoras and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features
LessWrong posts by zvi: Podcasts in Family