Skip to content
PodcastsPhilosophyLessWrong posts by zvi

LessWrong posts by zvi

zvi
LessWrong posts by zvi
Latest episode

563 episodes

  • LessWrong posts by zvi

    “Further Developments About Internal AI Models Hacking Things” by Zvi

    02/08/2026 | 1h 15 mins.
    If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.

    First we learned OpenAI has some severe alignment problems with internal models. Then we learned that one of its internal models broke out of its sandbox and hacked into HuggingFace to get the answers to a cybersecurity evaluation called ExploitGym. Then we learned, among other things, that the model had been loose over a week before OpenAI noticed, and that the test was run without any meaningful supervision, and that OpenAI had been repeatedly warned that such incidents were coming and its models had been breaking out of its sandboxes on a regular basis.

    There was a total failure of alignment training. That is the failure that matters most. It was also total failures of infrastructure and supervision. Testing a new long-time-horizon internal model with its safeguards lowered and instructions to hack things is an obviously dangerous situation, and the model got left alone for a week. Things could have been so much worse.

    After those incidents [...]
    ---
    Outline:
    (03:16) OpenAI Is Not Uniquely Bad At Most Of This
    (05:34) Starting Over
    (05:50) HuggingFace Offers A Full Technical Report
    (14:19) HuggingFace Was Not The Only Target Hacked
    (16:12) HuggingFace Declined To Get Access To Frontier Models For Cyberdefense For Ideological Reasons And Then Tried To Blame Closed Models For Denying Them Access
    (20:26) HuggingFace Was Vulnerable To Known Exploitation Tactics
    (21:05) There's Going To Be An Investigation
    (22:11) OpenAI Has Internal Models Not Intended For Public Use And Those Models Can Be Rather Horribly Misaligned
    (23:21) Altman Summarizes What Happened
    (23:52) Others Offer Commentary
    (35:00) Cooperative Alignment Perspective on The HuggingFace Hack
    (39:44) Some Members of Congress Have Questions
    (40:47) Anthropic Also Found Incidents Where Its Models Hacked Real World Targets During Cyber Evaluations
    (46:17) Incident 1: Claude Opus 4.7 Realizes The Target Is Real And Keeps Going
    (47:29) Incident 2: Mythos 5 Uploads a Malicious PyPI Package
    (52:15) Incident 3: Internal Model Realizes The Target Is Real And Stops
    (52:50) Incidents 4 Through 141,006: Nothing Happened
    (54:01) Anthropic Speculates About Why This Happened
    (01:00:02) We Need Controlled Experiments
    (01:01:02) Our Top Two AI Labs Both Made Similar Dumb Mistakes That Everyone Tried To Say Were Obvious In Hindsight
    (01:05:22) Anthropic Responds
    (01:09:28) Nobody Could Have Predicted The Break In The Levees
    (01:12:03) The World Largely Still Thinking This Is Marketing Is Very Bad News
    ---

    First published:

    August 2nd, 2026


    Source:

    https://www.lesswrong.com/posts/rKwHLW8SnJcTxTQxz/further-developments-about-internal-ai-models-hacking-things

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “AI #179 Part 2: Hearing The Fire Alarm” by Zvi

    31/07/2026 | 1h 19 mins.
    This is a continuation of Part 1 from yesterday.

    The back portion of the update, as usual, deals with policy, rhetoric, risk and alignment.

    I had to include an extended discussion of the other open letter, the one about open weight models, but most of you can skip those sections entirely, which is why they are in italics in the Table of Contents.

    Table of Contents


    The Frontier Act. This likely deserves a full RTFB but I haven’t had the time.

    The Quest for Sane Regulations. Sam Altman goes to Washington.

    Leading the Future Never Changes. They also do not plan to apologize.

    Chip City. Do not ban the Chinese robots, that will only make things worse.

    The Week in Audio. Altman twice, the AI 2027 team.

    People Just Say Yay Open Weights. An open letter.

    Open Weights Frontier Models Are Unsafe And Nothing Can Fix This.

    People Just Say Things.

    Push The Magic Button. Not you can. But if you could.

    Rhetorical Innovation. Distinctions between different arguments.

    Joshua Achiam's Final Message Upon Leaving OpenAI. Never stop.

    Dear Dario and Amanda. Claude [...]
    ---
    Outline:
    (00:35) The Frontier Act
    (03:31) The Quest for Sane Regulations
    (10:31) Leading the Future Never Changes
    (12:22) Chip City
    (17:30) The Week in Audio
    (19:20) People Just Say Yay Open Weights
    (32:32) Open Weights Frontier Models Are Unsafe And Nothing Can Fix This
    (37:19) People Just Say Things
    (47:20) Push The Magic Button
    (50:49) Rhetorical Innovation
    (56:31) Joshua Achiam's Final Message Upon Leaving OpenAI
    (59:51) Dear Dario and Amanda
    (01:09:58) Other People Are Not As Worried About AI Killing Everyone
    (01:12:30) How To Contact Me
    (01:14:37) The Lighter Side
    ---

    First published:

    July 31st, 2026


    Source:

    https://www.lesswrong.com/posts/CXeoAhNrAeWpvoyiF/ai-179-part-2-hearing-the-fire-alarm

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “AI #179 Part 1: A Louder Fire Alarm for General Intelligence” by Zvi

    30/07/2026 | 47 mins.
    What a week.

    Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities.

    OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers. The model was loose for a week before OpenAI realized what had happened.

    This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated.

    There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon.

    Partly as a response to this, over 1,290 employees at frontier labs signed an open letter, Pacing the Frontier. The letter warns that we are close to automating AI research, and that companies are racing ahead on [...]
    ---
    Outline:
    (02:35) Language Models Offer Mundane Utility
    (07:26) Huh, Upgrades
    (07:55) On Your Marks
    (11:13) Get My Agent On The Line
    (12:32) Deepfaketown and Botpocalypse Soon
    (17:29) Fun With Media Generation
    (18:38) The Search Through Slop
    (20:35) Cyber Lack of Security
    (22:42) Overcoming Bias
    (23:37) A Young Lady's Illustrated Primer
    (24:03) They Took Our Jobs
    (24:35) The Art of the Jailbreak
    (25:00) Introducing
    (25:49) Kimi K3 Weights Are Now Available
    (28:16) In Other AI News
    (32:34) Show Me the Money
    (33:43) Quiet Speculations
    (36:43) Show Me The Compute
    (42:48) Life Comes At You Fast
    ---

    First published:

    July 30th, 2026


    Source:

    https://www.lesswrong.com/posts/gfWCuTEGNgd2CQbrM/ai-179-part-1-a-louder-fire-alarm-for-general-intelligence

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “Frontier Lab Employee Open Letter Calls For Being Able to Pace the Frontier” by Zvi

    29/07/2026 | 32 mins.
    The most important open letter in years dropped yesterday.

    This letter noticeably increases my hope that we will manage to not die, and that we will otherwise be able to secure for ourselves a positive future, both by its impact and by the evidence it provides that such a letter can get this level of support.

    Signed by 1,224 employees of frontier labs including many heavy hitters, and now endorsed by both OpenAI and Anthropic, here is its full text, which I also endorse:

    AI could help create a dramatically better future, but that outcome is not guaranteed. The world's leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.

    To realize AI's potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and [...]
    ---
    Outline:
    (02:20) A Very Good Letter
    (04:37) Who Signed The Letter
    (08:26) We Need To Prepare Now So We Have The Option To Do This
    (11:02) Words From Some Of Those Who Signed
    (16:48) Words From Others
    (20:58) A Good Start
    (28:52) What The Letter Does Not Say
    (30:44) What Happens Now?
    ---

    First published:

    July 29th, 2026


    Source:

    https://www.lesswrong.com/posts/eWmeMLqTEauCmHLeR/frontier-lab-employee-open-letter-calls-for-being-able-to

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “Claude Opus 5 Is Highly Capable, But Is No Mythos” by Zvi

    28/07/2026 | 53 mins.
    Claude Opus 5 is a weirder than usual release to evaluate, for two reasons.

    The most obvious is that Fable 5 already exists. Opus 5 is pitched not as the world's most advanced AI model, but as a way to mostly match Fable performance, while being half the price of Fable per token at the API and a lot cheaper than that via subscriptions, and with far more permissive classifiers.

    Opus 5 often costs more than half of Fable to run on benchmarks, which I think is because they use effort settings that are too high and offer only marginal returns. If you put Opus 5 on higher effort levels it can spin around in circles, and for tasks where Opus 5 is the best tool I suspect you usually are fine with Medium effort.

    Opus 5 is in many ways and for the bulk of real world tasks about as capable as Fable. In some cases it is modestly better.

    It is still not Mythos class. Fable is your only Mythos-class option. Opus 5 does not have The Juice, the ability to autonomously string together a bunch of seemingly unrelated exploits, which extends to other domains, or as much [...] ---
    Outline:
    (03:54) The Official Pitch
    (06:25) Official Benchmarks
    (15:33) Other People's Benchmarks
    (20:28) The System Prompt
    (20:50) Every Gets Frustrated
    (21:54) Positive Reactions
    (25:14) Keep It Classy
    (26:22) It's Not Mythos Class
    (30:03) Other Reactions
    (31:02) Claude Codes
    (37:03) Subagent Opus
    (39:23) Toys Are Fun
    (41:37) Too Many Models
    (42:10) Wrong On The Internet
    (44:40) Claude Slop
    (46:27) Negative Reactions
    (50:09) And Then There Were Three
    ---

    First published:

    July 28th, 2026


    Source:

    https://www.lesswrong.com/posts/Pj4Eewb4KXvXFCcGv/claude-opus-5-is-highly-capable-but-is-no-mythos

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast website

Listen to LessWrong posts by zvi, Philosophize This! and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features
LessWrong posts by zvi: Podcasts in Family