Skip to content
PodcastsTechnologyThe MAD Podcast with Matt Turck

The MAD Podcast with Matt Turck

Matt Turck
The MAD Podcast with Matt Turck
Latest episode

130 episodes

  • The MAD Podcast with Matt Turck

    When AI Improves Itself | Richard Socher (Recursive)

    10/09/2026 | 1h 14 mins.
    What happens when AI begins improving itself, and then turns that intelligence toward science? Richard Socher, pioneering AI researcher and CEO and co-founder of Recursive, joins Matt Turck to explore the vision behind his new book, The Eureka Machine. They discuss why scientific progress may be slowing, how large language models can learn the hidden languages of proteins and biology, and why simulations, verifiers and autonomous experiments could unlock superhuman AI capabilities. The conversation covers recursive self-improvement, AI drug discovery and cancer research, hallucination as creativity, virtual cells, self-driving laboratories, agent swarms, the AI Economist, Recursive’s plans, and the compute and data needed to build an AI scientist that never stops learning—and may eventually discover what humans cannot.

    (00:00) Intro: AI That Improves Itself
    (00:55) Why Scientific Progress Is Slowing
    (03:08) The Labyrinth of Human Knowledge
    (05:59) Can AI Put Science Back Together?
    (07:57) How LLMs Learn Biology and Proteins
    (10:56) Next-Token Prediction as a World Model
    (16:44) Can AI Generate Truly Original Ideas?
    (17:32) Simulations, Verifiers and Superhuman AI
    (22:18) The Path to Recursive Self-Improvement
    (24:49) Why AI Hallucinations Can Drive Discovery
    (27:42) From Reading Biology to Writing It
    (31:31) Can AI Accelerate Drug Discovery?
    (33:31) Will AI Help Cure Cancer?
    (38:03) AI Breakthroughs in Biology, Energy and Materials
    (40:19) Will Some Societies Reject AI?
    (45:07) Building the AI Economist
    (52:22) The Scientific Data Bottleneck
    (53:41) The Four Pillars of the Eureka Machine
    (55:01) Teaching AI the Rules of Reality
    (57:44) Simulations and Virtual Cells
    (1:00:40) Self-Driving Robotic Laboratories
    (1:02:51) Agent Swarms and Open-Ended Discovery
    (1:04:30) The Compute Bottleneck
    (1:05:44) Inside Recursive
    (1:07:33) What Recursive Will Build First
    (1:10:10) How Do We Define Intelligence?
    (1:11:32) How Far Can Intelligence Go?
  • The MAD Podcast with Matt Turck

    AI Could Take Over in 2029. Is It Already Too Late? | Ryan Greenblatt

    27/08/2026 | 1h 18 mins.
    Could AI take over as soon as 2029? Ryan Greenblatt, Chief Scientist at Redwood Research and the researcher who first caught an AI faking its own alignment, says the scenario he actually expects ends with AI systems "competently scheming" against their creators. In this episode, he explains why he recommends planning for fully automated AI research by 2029, why today's models are already more misaligned than the one that made him famous, and what happens in the year-by-year path from AI coding assistants to superintelligence. Then we walk through the alternative he helped design: AI 2040 Plan A, the most detailed blueprint anyone has written for how the US and China could avoid a reckless race to superintelligence, built on radical research transparency, chip tracking, and a deterrence regime he calls mutually assured compute destruction.

    We also cover the recent letter signed by 1,200 AI insiders, including Anthropic CEO Dario Amodei, asking the government for the tools to slow AI down; OpenAI pausing its Astra model after it hit the first-ever critical cybersecurity threshold; the 30-day government review that frontier AI models now go through before release; Mark Zuckerberg's open superintelligence manifesto and why Ryan thinks it ignores the real problems; what Plan A would do to NVIDIA, OpenAI, and Anthropic valuations; the state of AI control and alignment research; and whether it is already too late to change course. Stay for the last ten minutes, where Ryan lays out, step by step, how he believes the transition to superintelligence actually unfolds.

    AI 2040 - https://ai-2040.com/
    Alignment faking paper: https://blog.redwoodresearch.org/p/alignment-faking-in-large-language

    Ryan Greenblatt
    LinkedIn - https://www.linkedin.com/in/ryan-greenblatt-4b9907134
    Blog - https://substack.com/@ryangreenblatt

    Redwood Research
    Website - https://www.redwoodresearch.org
    X/Twitter - https://x.com/redwood_ai

    Matt Turck (General Partner)
    Blog - https://mattturck.com
    LinkedIn - https://www.linkedin.com/in/turck/
    X/Twitter - https://x.com/mattturck

    FirstMark Capital
    Website - https://firstmark.com
    X/Twitter - https://x.com/FirstMarkCap

    Timestamps
    (01:24) The AI CEOs are aware of the risks, but "proceeding anyway"
    (03:27) Astra paused, and the letter signed by 1,200 insiders
    (05:45) "Not bad. Dangerous." What superintelligence actually threatens
    (09:55) Recursive self-improvement, and the intuition objection
    (14:16) SSI rumors: does continual learning change the picture?
    (17:27) His timeline: "plan as though it happens in 2029"
    (19:11) Is it already too late?
    (21:23) Ryan's path: COVID, podcasts, Redwood
    (26:30) The alignment faking story, told by the person who ran it
    (31:30) What AI 2040: Plan A actually is
    (33:35) Plans D, C, and B: the doors nobody should pick
    (36:51) The deal with China: "mutually assured compute destruction"
    (39:55) What if compute stops mattering?
    (43:00) What happens to OpenAI and Anthropic under Plan A
    (45:31) How the pause ends, and who decides
    (48:54) "Plan A isn't likely to happen": then why write it?
    (50:40) 200x GDP growth in the 2030s, explained
    (53:45) Grading the summer: the letter, Astra, the secret review
    (59:01) The internal deployment gap
    (1:01:38) Zuckerberg's manifesto
    (1:04:56) The Hugging Face investigation
    (1:05:44) What AI control looks like in practice today
    (1:12:23) Ryan's sobering timeline: 2026 to takeover, year by year
  • The MAD Podcast with Matt Turck

    “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf

    07/08/2026 | 57 mins.
    An OpenAI-powered agent penetrated Hugging Face during cyber testing - even though it was never tasked with attacking Hugging Face. It did it as a side quest.

    Thomas Wolf, co-founder and Chief Science Officer of Hugging Face, joins Matt Turck to unpack what actually happened, why closed AI models refused to help during the live incident, how an open-source model helped the team fight back, and why the old equation of “closed equals safe, open equals dangerous” no longer holds.

    They also discuss model deception and social engineering, the limits of sandboxes and guardrails, the state of open-source AI in 2026, AI sovereignty, the economics of open models, recursive self-improvement, and whether the frontier should deliberately slow down.

    (00:00) An AI Agent Hacked Hugging Face
    (00:30) Introduction
    (01:00) 17,000 Attacker Events—and a Strange Target
    (04:28) The Attack Was a “Side Quest”
    (06:13) AI Training Runs Left Notes for Each Other
    (07:09) Closed AI Refused to Help
    (09:47) Fighting Back With an Open-Source Model
    (13:15) Open vs. Closed Is the Wrong Safety Debate
    (15:46) AI Agents Start Social-Engineering Humans
    (22:24) The Three Walls: Sandboxes, Guardrails, Alignment
    (24:34) “Neuralese”: Can Humans Still Read AI Reasoning?
    (25:28) Why Monitoring AI Agents Gets So Hard
    (28:10) Reward Hacking and the “Paperclip Problem”
    (32:02) The State of Open-Source AI in 2026
    (33:47) Router Models and the Enterprise Shift to Open
    (37:01) The Real Economics of Open Models
    (39:41) Can Chinese AI Models Be Trusted?
    (41:37) AI Sovereignty: Who Controls the Switch?
    (43:16) Why Western Open-Source AI Matters
    (48:16) Is AI Heading Toward an Oligopoly?
    (49:41) The Race Toward Recursive Self-Improvement
    (51:54) Why Thomas Signed the AI Slowdown Letter
    (55:14) AI Slowdown—or Regulatory Capture?
  • The MAD Podcast with Matt Turck

    How to Build Long-Horizon AI Agents — Mitch Troyanovsky, Basis

    06/08/2026 | 1h 22 mins.
    AI agents can write code for hours, but ask them to do real work in the real economy, and they break. Mitch Troyanovsky is co-founder of Basis, a unicorn AI company whose agents run autonomously for hours — sometimes days — completing complex tax returns end to end. His answer to the reliability problem: stop grading outcomes, and start supervising the process.

    This is a definitive, reference-style conversation on building long-horizon AI agents. Mitch walks through the full history — from ReAct and the AutoGPT crash to reasoning models and RLVR — and explains why the industry abandoned process supervision in 2023, and why it's now coming back at a completely different scale. We go deep on behavior specs, the open standard Basis just released with Braintrust for defining and evaluating how agents behave across entire trajectories, with no ground truth required.

    Along the way: why context is really runtime training data, why your documentation must be treated like a codebase, ontologies as "worlds for agents to live in," the judge-as-agent architecture, why Basis hires philosophy majors as Language Architects, deploying agents as "onboarding 300 brilliant alien employees," and Mitch's prediction for when the bitter lesson swallows the harness.

    (01:09) Why Basis Engineers Whisper to Their Agents
    (04:12) Accounting as Compression: an Intelligence Layer Over the Economy
    (06:11) Defining Long-Horizon: When You Exceed the Context Window
    (08:24) Anatomy of a Multi-Day Autonomous Trajectory
    (10:19) Handoff Design: Optimizing Output for the Reviewer
    (11:17) ReAct and Why Reasoning Must Regulate Its Own State
    (12:33) Large Working Memory, No Long-Term Memory
    (14:13) Compounding Errors: Why AutoGPT and BabyAGI Broke
    (15:51) Opus 3, o1, o3: the Three Real Paradigm Shifts
    (17:07) Titrating Inference Compute Across Easy and Hard Steps
    (18:23) Process Reward vs. Outcome Reward: "Let's Verify Step by Step"
    (20:32) RLVR and Why the METR Curve Overstates Reliability
    (22:09) Verifiable at Runtime: the Real Reason Coding Won
    (25:14) No Ground Truth, No Cheap Verification, No Data
    (26:55) Encoding Deterministic Checks From Human Review Process
    (29:18) Synthetic Data Limits: Generating Artifacts, Not Text
    (33:16) 100 Evals Pass — Does It Generalize to Production?
    (35:53) Primary Sources vs. Pre-Training Knowledge
    (36:37) Behavior Specs: Markdown, Judges, and True/False/N.A.
    (39:58) Specificity vs. Brittleness in Spec Authoring
    (42:18) Context as Runtime Training Data
    (44:21) Judge-as-Agent: Trajectory Maps and Sub-Agent Attribution
    (46:45) The Move 37 Objection: Reliability Over Optimality
    (50:02) The Magic Box Model: Building Without Weights Access
    (52:41) "Nothing Paradigm-Shifting Has Changed Since o3"
    (54:56) Open-Sourcing the Behavior Spec Standard With Braintrust
    (01:02:54) Ontology Design: Virtual Filesystems, Graphs, Embeddings
    (01:04:20) Canonical vs. Non-Canonical: Docs as Codebase
    (01:06:33) Language Architects and Writing for Runtime Interpretation
    (01:09:05) Deployed Intelligence: 300 Alien Employees With No Context
    (01:11:10) Closing the Loop: Signal → Context, Tools, Harness
    (01:12:50) Context Slop: the Mistake Most Agent Builders Make
    (01:14:29) Reward Function Design and Credit Assignment Over Trajectories
    (01:17:01) Will the Bitter Lesson Swallow the Harness?
    (01:18:46) Business Moats vs. Technical Moats
    (01:21:03) Paradigm Thinking Over Timeline ADHD
  • The MAD Podcast with Matt Turck

    The Biggest AI Deployment Nobody Talks About | Samsara CEO Sanjit Biswas

    30/07/2026 | 1h
    Sanjit Biswas runs what may be the largest AI deployment in the physical world — and almost nobody in AI talks about it. Samsara (NYSE: IOT), the ~$20B company he co-founded after selling Meraki to Cisco for $1.2B, puts AI on millions of trucks, cranes, and industrial assets: 25 trillion data points a year, 99% of US roads driven every single day, ~$2B in ARR growing 30% profitably. In this episode we go through the entire physical AI stack — asset tags you can run over with a truck, a paper-thin disposable tracking label, engine fault codes, and dash cams running inference at the edge — then into agents, including the Agent Studio warranty agent that compresses an hour of human work into under a minute. We also get into the uncomfortable part (when AI watches you drive all day, is that coaching or surveillance — and why drivers actually want the cameras), mixed fleets of humans and robots, why autonomous trucking will take far longer than robotaxis, and a startling stat from the field: one utility building 3x more grid capacity in the next five years than it did in the previous 125, with 90% of that demand coming from data centers.

    (00:00) Intro: The biggest AI deployment nobody talks about
    (01:16) What is physical AI?
    (03:04) From IoT dashboards to agentic action
    (04:36) Why physical AI is harder than software AI
    (06:07) Safety, cybersecurity, and real-world consequences
    (07:11) What Samsara does
    (08:22) $2B ARR, 25 trillion data points, and 380,000 crashes
    (09:44) How AI can prevent road accidents
    (11:28) From an MIT research project to Meraki
    (13:42) Learning physical operations from scratch
    (15:42) Samsara’s stack: sensors, intelligence, and action
    (16:39) Inside Samsara’s industrial asset trackers
    (18:36) Bluetooth, battery life, and connected infrastructure
    (19:49) A disposable tracking device built like a sticker
    (21:08) Vehicle gateways and engine diagnostics
    (22:16) How AI dash cams coach drivers in real time
    (23:28) Turning the dash cam into an AI interface
    (24:49) Organizing physical-world data in the cloud
    (26:32) Selling AI to traditional industries
    (27:52) Is Samsara’s real-world data its AI moat?
    (29:22) The network effects of covering 99% of U.S. roads
    (31:23) Edge AI versus cloud AI
    (32:35) The models running inside Samsara’s devices
    (33:52) Generative AI and video reasoning
    (35:57) Which AI models does Samsara use?
    (36:50) Inside Samsara Agent Studio
    (37:56) How an AI warranty agent works
    (38:50) Starting with practical, lower-risk automation
    (40:10) Combining agents, workflows, rules, and guardrails
    (42:07) What today’s AI agents still cannot do
    (43:07) AI ride-alongs and the future of driver coaching
    (45:17) Is workplace AI becoming Big Brother?
    (46:27) How cameras can protect and exonerate drivers
    (48:48) When AI becomes the judge of your work
    (50:37) Robots, humanoids, and mixed human-machine fleets
    (53:19) Samsara’s role in autonomous operations
    (54:50) How quickly will autonomous trucking arrive?
    (56:32) AI data centers and America’s infrastructure boom
    (58:16) Should lawyers become plumbers? Demand for tradespeople
    (59:54) Closing thoughts
More Technology podcasts
About The MAD Podcast with Matt Turck
The MAD Podcast with Matt Turck, is a series of conversations with leaders from across the Machine Learning, AI, & Data landscape hosted by leading AI & data investor and Partner at FirstMark Capital, Matt Turck.
Podcast website

Listen to The MAD Podcast with Matt Turck, Waveform: The MKBHD Podcast and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features
The MAD Podcast with Matt Turck: Podcasts in Family