578 episodes
- Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.
Here is how his solution works, or see Tenobrus's version.
AI outputs are not deterministic. The AI's job is to pick the probability of each potential next token. The token is then chosen at random.
By default you use a source of pseudo-randomness for each choice, since actual true randomness is annoying.
To apply the watermark, you use an otherwise identical private source of pseudo-randomness derived from a secret key.
Then, given enough text, a score is derived for howe well the choices fit with that particular pseudo-randomness source, versus a different source.
You provide an API that lets anyone check for the watermark.
If you want to dig deeper, here is a full paper. The method has very nice properties:
This has no practical impact on outputs. Humans cannot tell the difference, at all.
The marginal cost of doing this is very close to zero.
The watermark can be removed by rewriting in your own words, and appears in proportion to how many of the AI's detail choices you [...]
---
Outline:
(03:51) This Is Fine
(04:37) Anthropic Derangement Syndrome
(07:34) People Don't Understand LLM Outputs Are Already Random
(08:47) People Don't Trust The Method To Be Costless
(12:20) People Are Suspicious Of Any Alteration On Principle
(14:16) Maybe It's Partly The Word Watermark
(15:14) A Lot Of People Don't Want To Get Caught
(16:04) There Are Some Times You Prefer Not To Be Recognized
(16:18) There Are Some Good Reasons To Be Concerned
(16:37) Cheat Cheat Cheat Cheat Cheat
(18:38) The Writing In The Middle and Error Rates
(21:00) Millions For Defense But Not One Cent For Tribute
---
First published:
August 21st, 2026
Source:
https://www.lesswrong.com/posts/3mKuPHmaK7NW3QypR/ai-text-watermarking-is-free-and-good
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward.
OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger problems are in alignment, infrastructure and supervision, and in its training pipeline. OpenAI has now taken initial steps to address What Happened leading up to HuggingFace attack, including pauses to development while new safeguards are put in place and problems are diagnosed. These are promising early signs, but it is early. We will see if they follow through, and we still await the post-mortem of the HuggingFace attack.
Anthropic revenue continues to climb as they prepare for their IPO, although growth has slowed somewhat recently. However, they too have plenty of problems under the hood. They shared many of them in the August 2026 Anthropic Risk Report.
This week also offered time to cover Dwarkesh Patel's Podcast With Ryan Greenblatt, centrally on the potential for AI recursive self-improvement. I am working on a follow-up post to some other issues raised during that podcast.
Table of Contents
Language Models Offer Mundane Utility. The token [...]
---
Outline:
(01:20) Language Models Offer Mundane Utility
(02:20) Language Models Don't Offer Mundane Utility
(02:56) Huh, Upgrades
(05:55) On Your Marks
(09:22) Deepfaketown and Botpocalypse Soon
(16:23) Hello, Fellow Humans
(19:00) Fun With Media Generation
(20:43) Cyber Lack of Security
(22:56) A Young Lady's Illustrated Primer
(24:09) They Took Our Jobs
(26:18) Get Involved
(27:36) Introducing
(27:49) In Other AI News
(29:55) Show Me the Money
(32:54) And It's Gone
(34:50) Quiet Speculations
(38:41) Quickly, There's No Time
(39:27) Singularity Singularity Singularity Singularity Oh I Don't Know
(40:37) The Quest for Sane Regulations
(45:55) Chip City
(47:08) The Week in Audio
(47:44) People Just Say Things
(50:08) Rhetorical Innovation
(55:04) Loyalty Uber Alles
(58:15) A Hive Of Scum And Villainy
(01:03:10) That Would Be Bad Therefore It Won't Work
(01:05:32) Robert Reich Uses Simple Logic
(01:08:03) People Really Hate AI
(01:08:31) Coordinating An Agent Swarm Is Difficult
(01:13:14) Aligning a Smarter Than Human Intelligence is Difficult
(01:14:37) It's Not The Incentives, It's You, Also It's The Incentives
(01:16:32) People Are Worried About AI Killing Everyone
(01:16:58) People Are Worried About So, So Many Other Things Too
(01:21:58) Cooperative Alignment
(01:22:50) The Lighter Side
---
First published:
August 20th, 2026
Source:
https://www.lesswrong.com/posts/JSZkzsi8cD4pW6ffA/ai-182-pause-for-reflection
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.
I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere:
OpenAI Shares Some Alignment Problems
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
More on An Internal OpenAI Model Hacking Into HuggingFace
Further Developments About Internal AI Models Hacking Things
OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
What Happened: OpenAI and HuggingFace.
Various Reflections About What Happened With OpenAI's Internal Models.
If you do not know the basics, read What Happened. It is necessary context for basically everything that is happening in the AI world.
It is important to get this right and understand how big a deal it was, whereas many such as the Financial Times get this centrally wrong.
We are still awaiting the full post-mortem on What Happened. I plan to cover that in depth once we have it.
OpenAI is now taking active, expensive steps to try and fix the problem going forward.
As usual, I am simultaneously happy to see [...]
---
Outline:
(02:07) OpenAI Has Some Alignment Problems
(04:22) Slow Down There Good Buddy
(10:12) What Exactly Is Paused?
(12:12) Three Pillars
(14:45) I've Got My Eye On You
(18:07) The Most Forbidden Technique
(20:03) Monitoring Is Only Defense-In-Depth
(23:32) Security
(24:15) Alignment
(30:37) A Crisis of Culture
(32:24) Closer Collaboration
(33:28) Reports of Death of Preparedness Team Greatly Exaggerated
(35:40) The OpenAI Foundation Just Funds Things
(37:51) Quickly, There's No Time
---
First published:
August 19th, 2026
Source:
https://www.lesswrong.com/posts/X3p8cFAzCgRErEcJr/openai-takes-initial-steps-to-address-its-alignment-problems
---
Narrated by TYPE III AUDIO. - I am grateful that Anthropic is producing periodic Risk Reports.
At first I was skeptical. It turns out I was wrong. Anthropic is revealing a lot of new information, some of it rather alarming, that it did not have to disclose, and is providing detailed insight into how they think about things. This is very cool.
Thus I found this report to be a moderately positive update overall, if we presume they are not silently omitting the worst of it. There are a bunch of not great things we find out about, but I would have expected some set of mistakes at least as bad, and I wouldn’t have expected them to choose to tell us about all of it.
It does mean one more set of 186 page documents I have to read every so often, almost all of which is meaningfully new material this time around.
The other revelation is the existence of the world's likely best model, ‘Model 2.’
This was a rough one to fully get through, so apologies in advance for any errors of interpretation.
Table of Contents
Agent Model 1 and Agent Model 2.
[...]
---
Outline:
(01:11) Agent Model 1 and Agent Model 2
(02:49) Executive Summary (1)
(04:13) The Rules Are Serious But Not Literal
(06:24) Misalignment Is a State of Mind (2.5)
(11:40) Autonomy Threat Model 1: Misalignment in High-Stakes Settings (2)
(14:28) Some Strange Uses Of The Word Safe I Wasn't Previously Aware Of
(15:53) Now Versus Future (2.17)
(16:29) The Core Claims And Argument (2.6)
(26:44) The Rest of the Important Arguments In Section 2
(28:59) Risk Assessment (2.19)
(29:47) Pre-Internal-Deployment Review (2.18)
(30:49) A Guide To Internal Use Monitoring (2.23.1)
(37:11) Blocking Interventions (2.23.2)
(38:28) The Power Seeking Environment Evaluation (2.24)
(39:25) Opus 4.8-Reward-Hacker (2.25)
(41:31) Autonomy threat model 2: Risks from automated R&D (3)
(42:29) Yes That Does Seem Kind Of Risky
(44:53) Could We Replace Our Researchers?
(46:27) How Much Could We Be Accelerating Our AI Researchers?
(49:43) What Could Possibly Go Wrong If We Replaced Our Researchers?
(50:23) Risk Mitigations For AI R&D Automation
(53:42) Overall Risk From Automation of AI R&D
(53:56) Biological and Technically Also Chemical Weapons Production
(54:45) The Threat Models for Biological and Chemical Weapons
(59:52) Model Capabilities (4.4)
(01:01:13) Classifiers (4.5)
(01:03:04) Acceleration Dynamics (5.1)
(01:04:06) Distillation (5.1.1)
(01:05:17) Safety Process Failures (5.2)
(01:05:31) Refusing To Find Innovative Misalignment Techniques (5.2.2)
(01:06:47) Exposing the Chain of Thought Reasoning To Grading Pressure Quite a Lot (5.2.3)
(01:08:15) Directly Training On Misaligned Behavior During a Production Training Run (5.2.4)
(01:10:42) An instance of unmonitored unrestricted agents with access tosensitive resources (5.2.5)
(01:11:43) Repeated training on alignment-faking transcript datasets (5.2.6)
(01:14:17) Benefits From Anthropic's Operating as a Frontier AI company (5.3)
(01:16:20) Model Weight Security (6.4)
(01:16:42) Risk Has Been Reported
---
First published:
August 18th, 2026
Source:
https://www.lesswrong.com/posts/dA8gohzABk6vT7yzP/anthropic-risk-report-august-2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go.
The vibes have shifted, contrast this to the lit recursion when he talked to Huang
As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.
If I am quoting directly I use quote marks, otherwise assume paraphrases.
Section titles are from the transcript whenever possible, to aid in navigation.
Introduction
The discussion is interesting throughout, although often frustrating, especially in the (mostly isolated) discussion about ‘aligned to whom?’ As usual, one could expand many responses into full posts, and maybe one should.
This podcast exists in light of recent misalignment and hacking events at OpenAI, Anthropic and UK AISI. You’ll want basic knowledge of that as background.
Ryan and Dwarkesh both have views of the situation different from my own, but are attempting to see where their positions lead, and try to balance educating people who start at zero with having a high level discussion.
[...] ---
Outline:
(00:56) Introduction
(03:08) Is AI R&D Verifiable Enough To Unlock Recursive Self-Improvement?
(10:03) Is AI progress bottlenecked by human expert data?
(19:07) Flat token prices suggest scaling has been slow
(21:54) Skills AI can't train on: does it even need them?
(22:36) Aligned to whom?
(31:46) Recent incidents of AIs colluding and deceiving humans
(34:39) What could possibly go wrong? A concrete scenario
(41:57) From reward hacking to takeover
(46:28) Time To Update
---
First published:
August 15th, 2026
Source:
https://www.lesswrong.com/posts/BZW8CeAHHJ52EvwYt/on-dwarkesh-patel-s-podcast-with-ryan-greenblatt
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
Trending Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast websiteListen to LessWrong posts by zvi, Alan Watts Being in the Way and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


LessWrong posts by zvi
Scan code,
download the app,
start listening.
download the app,
start listening.
LessWrong posts by zvi: Podcasts in Family






