Skip to content
PodcastsPhilosophyLessWrong posts by zvi

LessWrong posts by zvi

zvi
LessWrong posts by zvi
Latest episode

613 episodes

  • LessWrong posts by zvi

    “The Quest for Embedded Evaluators” by Zvi

    27/09/2026 | 17 mins.
    Dario Amodei's essay We Must Pace the Frontier committed Anthropic to embedded evaluators, who would be placed inside Anthropic and given employee-level access, so they could provide outside perspective and also reports on what was happening.

    There is only one problem. Who will be the evaluators?

    OpenAI followed suit on committing to the evaluators, and also issued a milquetoast but welcome call for international coordination. I will cover that here as well.

    What I won’t cover today, but hope to cover tomorrow, is the latest torrent of new AI hacking incidents that came to light over the weekend, which highlights that we badly need at least embedded evaluators, and plausibly far harsher measures.

    For now, you need to know that there were a lot more incidents that OpenAI did not disclosed, and also a new incident at OpenAI that just happened that forced them to again pause their most advanced model. I’ll get right on sorting all that out.

    Table of Contents


    Look, All I’m Asking For Is That You Find A Highly-Qualified, Experienced, Trustworthy, Non-Conflicted Source of Embedded Evaluators That Will Work Entirely For Free, Without Government Assistance or Money from [...]
    ---
    Outline:
    (01:15) Look, All I'm Asking For Is That You Find A Highly-Qualified, Experienced, Trustworthy, Non-Conflicted Source of Embedded Evaluators That Will Work Entirely For Free, Without Government Assistance or Money from EA Sources Not Chosen By the Lab
    (04:44) Anthropic Partners with Accenture for Embedded Evaluation, also Plans to Include METR
    (11:04) Reading the METR
    (14:55) OpenAI Suggests Doing The Least We Can Do
    ---

    First published:

    September 27th, 2026


    Source:

    https://www.lesswrong.com/posts/uLmf3GmBywsmG8LLZ/the-quest-for-embedded-evaluators

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “Claude Opus 5.5 Should Raise Your Ambitions” by Zvi

    26/09/2026 | 36 mins.
    When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level. They should have raised yours, too.

    Claude Opus 5.5 should raise your ambition levels again. It just works, and it persists, like Astra does. It does the things. And it is highly pleasant to talk to, and its writing is pleasant to read, while you are at it. The game has been changed, again.

    Feedback is almost universally positive. Claude was never gone, but also is so back.

    The benchmarks are excellent, but ignore the benchmarks. Be ambitious. Go out and do things. Get curious. Have more interesting conversations.

    If one of those things is Pacing the Frontier or otherwise ensuring that AI does not kill everyone, leaving us to enjoy our bounty? That's even better.

    By Claude Opus 5.5, for this post

    The Official Pitch

    The pitch is Fable-5.1-level performance at lower Opus-level price. Good pitch.

    We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than [...]
    ---
    Outline:
    (01:22) The Official Pitch
    (04:38) Our Price Cheap
    (06:18) Official Benchmarks
    (07:26) Other People's Benchmarks
    (10:54) Claude Classifies
    (12:27) The System Prompt
    (12:34) Reaction Rules
    (13:07) Vision In 3D
    (15:29) Claude Creates
    (18:11) Claude Composes
    (18:53) Positive Reactions
    (24:50) Good Talk
    (27:03) On Writing
    (30:30) Big Model Smell
    (32:30) Check Your Work
    (32:56) Negative Reactions
    (34:02) Not So Fast
    (34:55) Some People Need Practical Advice
    ---

    First published:

    September 26th, 2026


    Source:

    https://www.lesswrong.com/posts/rtPiip9igy3QvxYdM/claude-opus-5-5-should-raise-your-ambitions

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “On Ezra Klein’s Podcast With Jensen Huang” by Zvi

    25/09/2026 | 52 mins.
    Jensen Huang accidentally called for shutting down OpenAI and intentionally called for spending vastly more on safety.

    This is why we say that some podcasts are self-recommending. Here we go.

    As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.

    If I am quoting directly I use quote marks, otherwise assume paraphrases.

    Section titles are from the transcript whenever possible, to aid in navigation, but here we don’t have those so I chose the section titles.

    Jensen Huang very much does not believe in ASI (superintelligence). He doesn’t think AI can ever be a different kind of thing from software. He thinks demand can rise by a billion times and we can ‘accelerate the living daylights out of’ AI, but it will never be more than a ‘new abstraction level’ and thus won’t fundamentally change anything. This is not a coherent position under reflection, but that is the position he holds.

    The ‘intro’ sections are fine, but the real meat starts with the HuggingFace Incident.

    What we see is Jensen Huang on tilt and caught in loops [...]
    ---
    Outline:
    (02:59) Jensen Gives His AI Speech
    (04:39) They Took Our Jobs
    (11:52) Open Weights Models Are Good For Nvidia
    (13:39) The HuggingFace Incident
    (15:16) Jensen Huang Says Keep Your AIs From Harming the World
    (19:38) Jensen Huang Accidentally Calls For Shutting Down OpenAI
    (22:51) Jensen's Arguments Prove Too Much
    (31:47) Astra Is Hard To Monitor
    (33:04) Jensen Huang Seems Legitimately Confused In Confusing Ways
    (37:00) Solve Your Other Problems First and Get Back to Me
    (40:07) Explicit Denial of Existential Risk
    (44:34) A Short Summary
    (45:50) Chip City
    (48:46) Jensen Huang
    ---

    First published:

    September 25th, 2026


    Source:

    https://www.lesswrong.com/posts/j3xefrWrNqsmMfJEi/on-ezra-klein-s-podcast-with-jensen-huang

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “AI #187: Coming Into Play” by Zvi

    24/09/2026 | 1h 52 mins.
    Opus 5.5 was released on Tuesday. I covered the system card yesterday, and will cover its capabilities soon. By all reports it is an excellent model. There are lots of fun videos going around that Opus has generated, which I will include as part of that.

    OpenAI released a new cheaper and improved Sol and Luna. No one is talking about them due to Opus 5.5, but these should be an important upgrade under the hood.

    Bernie Sanders and Greg Casar have formally introduced the Ban Artificial Superintelligence Act. That means we get to read (RTFB) it. As always, I reserve judgment on particular bills until I can read them in detail. MIRI did so, and endorses the bill as directly confronting the extinction threat. I hope to do an RTFB soon.

    I have spun two things off the weekly:


    Coverage of the quest for the right embedded evaluators and related questions and attacks, which will become its own post.

    Some issues related to cooperative alignment, which may get folded into the model welfare post.

    I also might, in addition to a potential RTFB on the Sanders bill, do full podcast [...]
    ---
    Outline:
    (01:52) On The Terms Superintelligence and 'Super Intelligence'
    (03:53) Language Models Offer Mundane Utility
    (04:26) Language Models Don't Offer Mundane Utility
    (05:53) Language Models Can Only Work With What You Give Them
    (08:53) Huh, Upgrades
    (10:42) On Your Marks
    (13:16) Get My Agent On The Line
    (15:53) Deepfaketown and Botpocalypse Soon
    (17:43) Fun With Media Generation
    (18:27) Copyright Confrontation
    (19:07) Cyber Lack of Security
    (20:04) Hugging the Face
    (24:15) Hacking Into OpenAI
    (28:11) They Took Our Jobs
    (28:33) Get Involved
    (28:41) Anthropic Has a Wet Lab and a Potential Gene Editing Technique
    (34:13) Introducing
    (35:11) In Other AI News
    (37:08) Show Me the Money
    (37:35) Bubble, Bubble, Toil and Trouble
    (39:02) Anthropic Approaches Recursive Self-Improvement
    (45:40) Others Approach Recursive Self-Improvement
    (48:56) Burden of Proof
    (49:28) Quickly, There's No Time
    (52:46) Left Wing Americans Really Hate AI For Different Reasons
    (54:41) Chip City
    (54:49) Pick Up the Phone
    (55:48) The Week in Audio
    (01:01:13) People Just Say Things
    (01:06:07) Venkatesh Rao Stops Writing
    (01:07:35) A Call for Control of Frontier AI Models
    (01:11:50) Calls For Pacing The Frontier
    (01:13:56) A Matter of Antitrust
    (01:14:23) A Matter of Liability
    (01:18:26) Quest for Sane Regulations
    (01:19:20) Rhetorical Innovation
    (01:24:08) Tap the Sign
    (01:24:49) A Matter of Some Debate
    (01:28:43) Astra Is Hard to Monitor
    (01:29:14) Anticipating What a Smarter Intelligence Can Do Is Impossible
    (01:34:56) Would You Look At All These Goalposts
    (01:39:30) I, Robot
    (01:43:13) People Are Worried About AI Killing Everyone
    (01:44:40) Other People Are Not As Worried About AI Killing Everyone
    (01:45:50) Joe Rogan
    (01:47:19) The Lighter Network Graph
    (01:51:27) The Lighter Side
    ---

    First published:

    September 24th, 2026


    Source:

    https://www.lesswrong.com/posts/o7pYWzWWwGDePoC5E/ai-187-coming-into-play

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “Claude Opus 5.5: The System Card” by Zvi

    23/09/2026 | 52 mins.
    Introducing the world's most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5.

    Anthropic is claiming Opus 5.5 is outright as good or better than Fable 5.1, while being actively cheaper than Opus 5.

    That means it's time for a good old system card reading.

    Due to the situation becoming increasingly hard to monitor, I never got a chance to publish my model welfare review for Claude Fable 5.1.

    My plan is to combine that with my welfare review for Claude Opus 5.5, once we have had time to get experience with Opus 5.5.

    The capabilities review will arrive in the next few days as per usual. The quick feedback from the internet is that Opus 5.5 is very good. I need more time before I am willing to offer comment.

    Areas that duplicate previous cards or otherwise contain no useful info are skipped.

    Opus 5.5 Self-Portrait (fully self-created using code)

    Table of Contents


    Classifiers (1.5).

    RSP Evaluations (2).

    Biological Evaluations (2.2).

    AI R&D (2.3).

    Alignment Risk (2.4).

    Cyber (3).

    Cyber Capability [...]
    ---
    Outline:
    (01:24) Classifiers (1.5)
    (02:23) RSP Evaluations (2)
    (03:17) Biological Evaluations (2.2)
    (07:54) AI R&D (2.3)
    (12:34) Alignment Risk (2.4)
    (13:22) Cyber (3)
    (15:17) Cyber Capability Evals (3.3)
    (17:05) Safeguards (3.4)
    (17:39) Safeguards Robustness Training (3.5)
    (20:12) Safeguards and Harmlessness (4)
    (21:52) Agentic Safety (5)
    (22:56) Malicious Agentic Influence Campaigns (5.1.3)
    (23:51) Prompt Injection Risk (5.2)
    (25:16) Alignment (6)
    (28:24) Negotiating With Your Local Claude Auditor (6.1.3)
    (29:33) Internal Misalignment Cases (6.3.1)
    (30:46) Automated Behavioral Audit (6.4)
    (33:24) Wherever Did These Evals Come From (6.4.8 and 6.4.9)
    (35:51) Potential Blind Spots (6.4.11)
    (38:24) Targeted alignment and honesty evaluations (6.5)
    (41:48) White Box Analysis (6.6)
    (43:37) Verbalized Grader Awareness (6.6.2)
    (44:56) Sandbagging (6.6.3)
    (45:54) Capabilities to Evade Safeguards (6.6.4)
    (49:27) Intentionally Taking Actions Very Rarely (6.6.4.3)
    (50:28) Chain of Thought Controllability (6.6.4.4)
    (51:37) It's A Good Model, Sir
    ---

    First published:

    September 23rd, 2026


    Source:

    https://www.lesswrong.com/posts/vMNTWTDWLorDqd3LS/claude-opus-5-5-the-system-card

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast website

Listen to LessWrong posts by zvi, everybody has a secret and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features
LessWrong posts by zvi: Podcasts in Family