621 episodes
- It is kind of a huge deal. OpenAI dumped a broad range of huge new mathematical results produced by an internal frontier model. They just put it all on GitHub.
This included 90 of the top 500 open problems in all of math, as per Proof of Atlas. In total there were 722 manuscripts (now 719 after three withdraws from the cluster that did not have Lean proofs) organized into 372 families.
This was the result of a single model, presumably the same one that produced the Navier-Stokes proof (as per their link back to that post), mostly on a single prompt (quasi-RH was one of the few exceptions), working an average of three hours’ worth of compute per solution found, after being asked to try its luck at about 4,000 problems. The prompt included lines like ‘Even if the problem is “open,” the intention is that you should resolve it and present a full solution.’ OpenAI was trying a lot less than maximally hard.
Levant and others called October 6, 2026, ‘obviously the most significant moment in mathematical history.’
Table of Contents
What Did We Prove?
The Mathocalypse.
The World [...]
---
Outline:
(01:16) What Did We Prove?
(04:09) The Mathocalypse
(07:10) The World Does Not Understand
(10:17) For Now You Can Still Do Math
(11:10) You Will Need To Find A New Problem
(14:40) Verification or Evaluation Is Not Always Easier Than Generation
(15:56) Cracking the Code
(20:08) Never Change
(20:28) The Mathematicians Are Not Okay
(29:27) The Situation Turns Ugly
(32:21) The Advisory Group Responds
(34:15) Another Mathematical Group Responds
(37:03) A Different Approach
(40:07) Three Withdraws
(41:00) Leaning Into Lean
---
First published:
October 9th, 2026
Source:
https://www.lesswrong.com/posts/eyqxBuxXCaamNSxBB/new-math-from-openai
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - The big drop of this week was not a new AI model. It was instead the biggest day (so far!) in the history of mathematics, as OpenAI dropped solutions to 90 of the top 500 open math problems, along with many others, reached on an average budget of three hours of Pro-level compute per question. This was kind of a big deal and I plan to cover it tomorrow.
We did see Claude Haiku 5.5, which looks promising given it is only 0.10 dollars/0.50 dollars.
My week was largely spent at The Curve. The entire conference was Chatham House, so I can’t give as many details as I would like, but I have a write-up here.
I finally had a chance to post my coverage of model welfare for both Mythos/Fable 5.1 and Opus 5.5. I continue to think this is an important issue for anyone wanting to understand today's AIs, even if you are highly confident that model welfare does not matter directly, as it has many practical implications on top of that.
Jay Clayton is the new AI Czar, and is at the head of a new taskforce. Given who else was in [...]
---
Outline:
(02:36) Language Models Offer Mundane Utility
(02:46) Consumers Use AI
(09:48) Language Models Don't Offer Mundane Utility
(12:02) Huh, Upgrades
(14:42) On Your Marks
(16:53) Gamers Gonna Game Game Game Game Game
(18:49) Choose Your Fighter
(19:48) Get My Agent On The Line
(25:15) Deepfaketown and Botpocalypse Soon
(29:17) Fun With Media Generation
(31:10) Cyber Lack of Security
(34:12) Hugging the Face
(35:18) Misaligned!
(37:16) A Young Lady's Illustrated Primer
(38:28) They Took Our Jobs
(40:45) Corporations Are Not Superintelligences
(42:56) Get Involved
(43:11) Introducing
(45:26) In Other AI News
(45:59) Show Me the Money
(48:18) Quiet Speculations
(50:26) Quickly, There's No Time
(52:26) He's Putting Together a Team
(55:14) The Quest for Sane Regulations
(57:37) Well, At Least They're Forecasters
(58:41) Chip City
(59:16) The Week in Audio
(01:03:48) Stop, Stop, He's Already Dead
(01:08:32) People Just Say Things
(01:10:46) Take a Moment
(01:12:24) [Artificial Intelligence]
(01:13:07) The American People Really Hate AI
(01:16:29) Rhetorical Innovation
(01:19:55) Aligning a Smarter Than Human Intelligence is Difficult
(01:20:11) Open Weight Models Are Unsafe And Nothing Can Fix This
(01:23:22) Cooperative Alignment
(01:33:45) Building the Field
(01:36:05) People Are Worried About AI Killing Everyone
(01:40:03) Other People Are Not As Worried About AI Killing Everyone
(01:40:38) Please Speak Directly Into This Microphone
(01:42:54) The Lighter Side
---
First published:
October 8th, 2026
Source:
https://www.lesswrong.com/posts/DZPodskqJJNEkMLvH/ai-189-new-math
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - The plan is no plan.
That is not the worst possible plan. But it is close.
Since before the transformer, we have warned that the most suicidal thing you could do would be to ask your AI to do your alignment homework, and automate the process. Alignment is complex and interacts deeply with every aspect of the world, and is one of the hardest possible things for an AI to get right even if it means maximally well and is itself functionally aligned. Mistakes get amplified up the chain, you get exactly what you optimized for, and you are rather doomed. Never go full RSI.
Yet, now more than ever, this seems to be the plan:
Solve prosaic issues and have operational excellence, to align current AI.
Have current AI do automated alignment work and figure it all out, aka ????.
Profit.
The good news is, most people realize this is at best a no-good terrible plan and would prefer a better one.
The bad news is that some people still think this is a plan that, while not great, still has a solid chance of working, with only [...]
---
Outline:
(02:59) Welcome to the Chatham House
(03:32) Overall Impressions
(05:29) AI Is Kind Of A Big Deal, Sir
(07:49) Quickly, There's No Time
(10:37) The Situation is Grim
(11:34) Track Trouble
(15:00) AI Is Not a Normal Technology
(16:00) The Plan is No Plan
(19:54) The Plan is to Pace
(21:45) Other Tracks
(25:30) The Plan is the President
(27:20) The Plan is Politics
(29:39) The Plan is to Post
(30:37) Are Alignment Evals Doomed?
(32:24) Eternal September
(34:06) Man's Search for Meaning
(35:48) The Food
(37:32) Chill Pill
---
First published:
October 7th, 2026
Source:
https://www.lesswrong.com/posts/qFa5qpw9bJiQMpJks/the-curve-bends-you
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - First you get to college via the wrong metrics, as previously discussed in #12. Then things get easier.
I’m still digging out from under everything that happened the four days I was gone. Normal AI posting should resume tomorrow, likely with a report from the conference.
Is Our Children Learning
The declines in NAEP scores and other tests seem to be concentrated on the weakest students. The bottom of the distribution is falling through the floor and the rest is mostly holding up. This does not match anecdotal reports coming out of schools, which suggest widespread declines, although those are presumably untrustworthy and it's traditional to think children are in general getting dumber. That doesn’t mean they aren’t but we need systematic evidence.
It's Bad, But It's Not That Bad
This really would be far worse than I think if it was true.
Tyler Cowen: It is worse than you think:
Of 360,000 children aged 15 in Zambia only five (not 5%, 5 total) could read at “globally proficient levels.”
The source is PISA.
I mean, in addition to this meaning essentially college-level reading skills, it's very obviously not true. [...]
---
Outline:
(00:30) Is Our Children Learning
(00:59) It's Bad, But It's Not That Bad
(02:10) Standardized Tests Help Disadvantaged Students
(04:26) Do Not Saturate Your Benchmarks
(05:23) Holistic Admissions Turn Childhood Into Kayfabe
(08:30) Holistic Admissions Should Mostly Be Positive Selection
(09:40) Beware Stolen Valor
(10:39) The Name Game
(11:00) Fair Weather College
(11:31) Disability Accommodations Are Now Mostly A Scam
(13:19) Harvard Has Some Grade Inflation
(21:06) Fighting Grade Inflation with GAAP Accounting
(24:38) Not Fighting Grade Inflation
(25:11) Is Our Children Learning?
(26:22) Biting All The Bullets
(28:00) Good Luck, Have Fun
(28:53) Cheaters Gonna Cheat Cheat Cheat Cheat Cheat
(35:06) Most College Students Fake Wokeness
---
First published:
October 6th, 2026
Source:
https://www.lesswrong.com/posts/QXmyxaQnB57reu7ie/childhood-and-education-21-grades-and-standards
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. - Is Google back?
They claim that they are back. Gemini 4 Argon is rolling out, with competitive frontier-level benchmarks, at 2 dollars/10 dollars.
What we don’t have is access to the model, because Google Fails Marketing Forever. So it is far too early to say what we have here. When I know more, so will you.
OpenAI was forced to pull what would have been GPT-6.1 Astra due to alignment failures. They did offer us GPT-6.1 Sol, which is pitched as approaching Astra quality at the much lower price of 2 dollars/10 dollars, the same as Gemini 4 Argon.
The rest of OpenAI's big Dev Day announcements were Ultrafast mode and Dots, your always-on AI agent based on Astra, which comes with your Pro subscription. I’m trying it out and will report back over time if I find it useful.
The new hotness remains Claude Opus 5.5. This model rocks. It has made me considerably more productive and made my day more pleasant. It should raise your ambitions. There are some particular reasons to call upon Fable 5.1 or Astra, and sometimes a cheaper model will do, but pending Argon I consider Opus 5.5 [...]
---
Outline:
(03:04) Language Models Offer Mundane Utility
(06:12) Huh, Upgrades
(07:38) Better Call Sol
(11:39) Gotta Go Ultrafast
(13:13) On Your Marks
(16:42) Choose Your Fighter
(19:16) Get My Agent On The Line
(22:34) The Warner Sister
(26:33) Deepfaketown and Botpocalypse Soon
(30:58) Fun With Media Generation
(31:53) Cyber Lack of Security
(33:24) A Young Lady's Illustrated Primer
(33:59) They Took Our Jobs
(41:04) Levels of Friction
(43:50) Get Involved
(45:54) Introducing
(47:54) In Other AI News
(48:57) Show Me the Money
(50:43) Quickly, There's No Time
(52:38) Pick Up the Phone
(53:36) Quest for Sane Regulations
(53:58) Chip City
(54:07) The Open Model Frontier Is Largely Massive Fraudulent Distillation Attacks
(56:23) The Week in Audio
(57:38) People Just Say Things
(59:17) Rhetorical Innovation
(01:03:55) Greetings From the Department of War
(01:06:55) The Department of Autonomous Warfare
(01:08:16) Aligning a Smarter Than Human Intelligence is Difficult
(01:11:03) Cooperative Alignment
(01:16:42) I'm Upping My p(doom), the Future Goes Foom
(01:24:16) No, You Make a Good Point, You're Not That Persuasive
(01:25:27) Muddling Through
(01:27:34) The Lighter Side
---
First published:
October 1st, 2026
Source:
https://www.lesswrong.com/posts/S2EAn9v4BwRdptsom/ai-188-gemini-dot-argon
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
More Philosophy podcasts
Trending Philosophy podcasts
About LessWrong posts by zvi
Audio narrations of LessWrong posts by zvi
Podcast websiteListen to LessWrong posts by zvi, Within Reason and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


LessWrong posts by zvi
Scan code,
download the app,
start listening.
download the app,
start listening.
LessWrong posts by zvi: Podcasts in Family









