Skip to content
PodcastsPhilosophyLessWrong posts by zvi

LessWrong posts by zvi

zvi
LessWrong posts by zvi
Latest episode

551 episodes

  • LessWrong posts by zvi

    “On Kimi K3: Its Capabilities And Related Discontents” by Zvi

    20/07/2026 | 1h 17 mins.
    Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability, the strongest open model.

    Do not get carried away. Do not judge Kimi K3 only its relative strengths. In aggregate it is several months behind the closed model frontier, at least four and my median guess is six, with the post-training closer and the pre-training farther out. This is less months than before, but the months are denser now.

    It is somewhat distilled. It likely outperforms on benchmarks relative to practical performance. All its benchmarks are scored at maximum effort, typically a lot more tokens than are used in similar tests by Fable or Sol. Performance looks jagged. Kimi will be excellent at some things, less so at other things.

    We will know more over the coming weeks. For now access is spotty and not that many people have actually had the chance to try Kimi K3, so I have larger error bars than usual around its capabilities. Alas, time waits for no one, so we press on.

    It is the largest open model so far at 2.8T, on [...]
    ---
    Outline:
    (03:07) DeepSeek Moments: Here We Go Again
    (05:47) We Had a Moment (Reprise from June 2025)
    (10:03) The Story Since Then
    (16:19) The Kimi K3 Announcement, Pitch and Basic Facts
    (19:34) On Modern Benchmaxxing
    (21:16) Other People's Benchmarks
    (26:15) Benchmarks Are Not The Real World
    (27:17) Technical Safeguards? What Are Those?
    (30:53) Things Kimi Can Do
    (32:06) Things Kimi Cannot Do
    (33:40) Things It Is Not Easy To Get Kimi To Do
    (37:02) Open Weight Models Are Unsafe And Nothing Can Fix This
    (40:34) Dean Ball Attempts To Be Constructive
    (58:24) Trump Administration Considering Executive Order Banning Chinese Open Models Within the United States
    (01:01:53) OpenAI Employees Are Relatively Bullish On This One
    (01:03:30) Kimi K3 Is Relatively Strongest At Typical Agentic Coding, Front End Work and 3D
    (01:06:06) Reactions
    (01:10:14) Who Are You?
    (01:12:09) How Did They Do It?
    (01:15:00) Conclusion
    ---

    First published:

    July 20th, 2026


    Source:

    https://www.lesswrong.com/posts/t7oZyAFej8FZrfbtY/on-kimi-k3-its-capabilities-and-related-discontents

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “Demis Hassabis on the New Coming Age” by Zvi

    19/07/2026 | 21 mins.
    Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that essay and various responses to it in Part 1.

    Part 2 of this post then covers Alex Turner's resignation, and his story about how he tried and failed to prevent Google from signing up to allow the Department of War to use its models for essentially whatever the government wants, including autonomous weapons.

    Demis Hassabis sold DeepMind to Google on condition that something like this would not happen. Yet here it is, happening. A cautionary tale.

    I will cover Kimi K3 tomorrow. I am hoping to know more by then. Please do share any reactions or info about it in the comments here.

    The Core Statement and Request

    He saying we are standing in the foothills of the singularity.

    His ask is a Frontier AI Standards Body within the US Government, similar to FINRA, that would govern ‘frontier labs,’ defined as any company that produces a frontier model based on various technical benchmarks. Evaluations would be updated regularly, and vulnerabilities would be addressed, both before [...]
    ---
    Outline:
    (01:04) The Core Statement and Request
    (02:38) Things Left Unsaid
    (04:05) The Proposal
    (06:04) A Good Start But Insufficient
    (08:52) Skeptics Of Future AI Capabilities
    (10:43) DeepMind On Bioresilience
    (12:39) Part 2: DeepMind Folds To The Department of War
    (13:28) DeepMind Leadership Failed Us
    (19:45) This Was a Failure We Must Learn From
    ---

    First published:

    July 19th, 2026


    Source:

    https://www.lesswrong.com/posts/3RfJLcmkztSTq9afc/demis-hassabis-on-the-new-coming-age

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “AI #177 Part 1: Tip of the Iceberg” by Zvi

    18/07/2026 | 47 mins.
    This week saw the releases of, among other things:


    GPT-5-6 Sol. It is a very good model, sir.

    Plan A, the follow up to AI 2027. It is a good plan worthy of discussion, sir.

    Kimi K3. This is only rolling out now, and will be covered next week.

    Muse Spark 1.1, the new Meta model. It is not frontier, but it is progress for them.

    Inkling, the first model from Thinking Machines.

    A call for regulatory action by Demis Hassabis, which I’ll cover soon.

    A new brief open letter call to action on AI regulation.

    That's on top of everything else, and an Opus 5 announcement is likely coming soon.

    The weekly once again got out of hand, so we’re splitting it once again into two, and once again saying we’ll be raising the bar for inclusion. And this time I mean it, as in enough to actually matter.

    Table of Contents


    Language Models Offer Mundane Utility. Whatever ye seek, ye shall find.

    Language Models Don’t Offer Mundane Utility. Gemini app needs some work.

    Language Models Upload Your Git Repository. Big problems [...]
    ---
    Outline:
    (01:17) Language Models Offer Mundane Utility
    (04:46) Language Models Don't Offer Mundane Utility
    (05:28) Language Models Upload Your Git Repository
    (08:35) Huh, Upgrades
    (09:30) Muse Spark 1.1
    (11:47) First Hit Free
    (15:36) On Your Marks
    (18:06) Choose Your Fighter
    (19:57) Get My Agent On The Line
    (23:23) Deepfaketown and Botpocalypse Soon
    (24:44) Fun With Media Generation
    (25:45) Copyright Confrontation
    (27:37) OpenAI Strikes Again
    (32:26) A Young Lady's Illustrated Primer
    (32:45) Recommendations for Policymakers
    (34:13) They Took Our Jobs
    (38:23) The Art of the Jailbreak
    (39:27) Get Involved
    (40:32) Introducing
    (41:12) In Other AI News
    (43:46) New Short Obviously True Statement About AI Just Dropped
    (46:03) Show Me the Money
    (46:21) The Lighter Side
    ---

    First published:

    July 16th, 2026


    Source:

    https://www.lesswrong.com/posts/who9xZ7DxuprsJoTr/ai-177-part-1-tip-of-the-iceberg

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “AI #177 Part 2: Wish You Were Here” by Zvi

    17/07/2026 | 1h 15 mins.
    As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions.

    Xi gave an important speech yesterday, so this post opens with that.

    There is talk that Kimi K3 is sufficiently strong that it upends many of these questions. It is clearly a candidate for another DeepSeek Moment, complete with stock drops for Google and SpaceX and (once again in a clear wrong-way move, the same as last time) Nvidia.

    Kimi K3 is clearly a very good model, exceeding expectations. Some are saying it is close to the frontier. The Artificial Analysis intelligence index has it at 57, a point ahead of Claude Opus 4.8, two behind Sol and three behind Fable. My presumption is that this number overstates its capabilities, but as always unless and until we have extensively tried the model ourselves, which I do not plan to do, we need to withhold judgment for at least a few days. I will be covering Kimi K3 in its own post at some point early next week.

    I have pushed further discussions involving Plan A and related issues into next week, as well as discussions around Demis Hassabis and Google [...]
    ---
    Outline:
    (01:29) Xi Gives A Good Speech on AI
    (15:33) Quiet Speculations
    (19:08) Tyler Cowen On Rebuilding The Future
    (22:39) The Quest for Sane Regulations
    (24:15) Wish You Were Here
    (26:55) The Week in Audio
    (27:29) New York Issues Moratorium On Data Centers
    (30:26) People Just Say Things
    (31:24) Rhetorical Innovation
    (39:29) Imagine Asking Questions
    (41:46) Anthropic Surveys Things It Calls Misalignment
    (54:30) Aligning a Smarter Than Human Intelligence is Difficult
    (58:30) The Most Forbidden Technique
    (01:01:04) Cooperative Alignment
    (01:12:38) The Lighter Side
    ---

    First published:

    July 17th, 2026


    Source:

    https://www.lesswrong.com/posts/Zjj3PTEng8GDqfK6j/ai-177-part-2-wish-you-were-here

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • LessWrong posts by zvi

    “AI #177 Part 1: Tip of the Iceberg” by Zvi

    16/07/2026 | 47 mins.
    This week saw the releases of, among other things:


    GPT-5-6 Sol. It is a very good model, sir.

    Plan A, the follow up to AI 2027. It is a good plan worthy of discussion, sir.

    Kimi K3. This is only rolling out now, and will be covered next week.

    Muse Spark 1.1, the new Meta model. It is not frontier, but it is progress for them.

    Inkling, the first model from Thinking Machines.

    A call for regulatory action by Demis Hassabis, which I’ll cover soon.

    A new brief open letter call to action on AI regulation.

    That's on top of everything else, and an Opus 5 announcement is likely coming soon.

    The weekly once again got out of hand, so we’re splitting it once again into two, and once again saying we’ll be raising the bar for inclusion. And this time I mean it, as in enough to actually matter.

    Table of Contents


    Language Models Offer Mundane Utility. Whatever ye seek, ye shall find.

    Language Models Don’t Offer Mundane Utility. Gemini app needs some work.

    Language Models Upload Your Git Repository. Big problems [...]
    ---
    Outline:
    (01:17) Language Models Offer Mundane Utility
    (04:45) Language Models Don't Offer Mundane Utility
    (05:27) Language Models Upload Your Git Repository
    (08:34) Huh, Upgrades
    (09:29) Muse Spark 1.1
    (11:48) First Hit Free
    (15:36) On Your Marks
    (18:06) Choose Your Fighter
    (19:57) Get My Agent On The Line
    (23:23) Deepfaketown and Botpocalypse Soon
    (24:45) Fun With Media Generation
    (25:46) Copyright Confrontation
    (27:38) OpenAI Strikes Again
    (32:26) A Young Lady's Illustrated Primer
    (32:46) Recommendations for Policymakers
    (34:13) They Took Our Jobs
    (38:23) The Art of the Jailbreak
    (39:27) Get Involved
    (40:32) Introducing
    (41:12) In Other AI News
    (43:47) New Short Obviously True Statement About AI Just Dropped
    (46:04) Show Me the Money
    (46:21) The Lighter Side
    ---

    First published:

    July 16th, 2026


    Source:

    https://www.lesswrong.com/posts/who9xZ7DxuprsJoTr/ai-177-part-1-tip-of-the-iceberg

    ---

    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast website

Listen to LessWrong posts by zvi, History of Philosophy Without Any Gaps and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features
LessWrong posts by zvi: Podcasts in Family