1041 episodes
- 当AI学会了“活在当下”,不再被历史包袱拖累时,我们人类自己又该如何避免被它悄悄“废掉”核心能力呢?本期节目,我们不仅要探讨如何用一把特制的尺子去衡量AI是否真的懂我们的“不开心”,还将揭秘如何培养出一个靠谱的AI“批评家”,让它实现高效的自我进化。最后,我们会一起探寻训练AI时那个神秘的“档位”,看看这些最新论文将如何刷新我们对人机协作与AI成长的认知。
00:00:33 让AI学会“活在当下”
00:05:00 AI越来越聪明,但它真的懂你的“不开心”吗?
00:09:24 那个替你干活的AI,正在悄悄“废掉”你
00:13:48 AI的成长烦恼,一个“批评家”的自我修养
00:20:39 训练AI的秘密“档位”
本期介绍的几篇论文:
[AI] SKILL.state: Scalable Long-Horizon Agent Skills
[Google LLC & Purdue University]
https://arxiv.org/abs/2608.26263
---
[CL] HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
[Beth Israel Deaconess Medical Center]
https://arxiv.org/abs/2608.25071
---
[AI] AI Agents Push Humans Out of the Loop
[Hugging Face]
https://arxiv.org/abs/2608.23642
---
[LG] Best Practice Critic Optimization
[National University of Singapore & Tencent Hunyuan]
https://arxiv.org/abs/2608.23566
---
[LG] Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
[Peking University & Ant Group]
https://arxiv.org/abs/2608.24814
在小宇宙查看该单集文稿 - 我们总惊叹AI越来越聪明,但你有没有想过,一个能理解世间万物的“意义图书馆”和一个离完美交付总差一步的“95分陷阱”同时存在于AI身上?本期几篇最新论文将带我们一探究竟,看看如何为AI装上外部“工作室”和独立的“世界大脑”,甚至揭示出AI群体“不靠说话”的协作奥秘。准备好了吗?让我们一起看看AI如何从一个聪明的“答题者”,进化成一个可靠的“行动派”。
00:00:32 AI的“意义图书馆”是怎么建成的?
00:06:25 AI的大考,为什么「差不多」等于「差很多」
00:10:45 人工智能的“外挂”,到底有多厉害?
00:15:23 给AI游戏世界装上一个“大脑”
00:20:36 人多,到底是力量大,还是乱糟糟?
本期介绍的几篇论文:
[CV] WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
[WeChat Vision, Tencent Inc.]
https://arxiv.org/abs/2608.24053
---
[AI] FrontierChallenge: Evaluating Scientific Workflow Completion
[Apodex Team]
https://arxiv.org/abs/2608.24979
---
[AI] Prime Agent: A Self-Improving RLM Harness
[Princeton University & Prime Intellect]
https://arxiv.org/abs/2608.23552
---
[CV] Code World Model: Coding Agent as World Brain
[Westlake University & Nanyang Technological University]
https://arxiv.org/abs/2608.25927
---
[AI] SwarmWorld: Stigmergic technological evolution in societies of language-model agents
[MIT]
https://arxiv.org/abs/2608.26081
在小宇宙查看该单集文稿 - 本期我们要聊的几篇最新论文,简直就像是AI上演了一出精彩的“进化三重奏”。你有没有想过,AI不仅能亲自下场做实验,还能通过扔掉海量信息反而学得更快?我们还会看到,AI如何像一个不眠不休的科研团队那样在失败中进化,像军队一样高效分工,以及这一切的背后,如何靠一本“备忘录”将所有经验沉淀为真正的智慧。准备好了吗?让我们一起探索AI正在解锁的全新可能性!
00:00:34 AI 不再只是“思想家”,它开始“动手”了
00:04:59 AI 进化新思路,扔掉 95% 的信息,反而学得更好?
00:10:55 AI的“试错”进化论
00:17:42 将军与士兵,人工智能的完美分工
00:22:55 给AI装个“备忘录”,为什么笨办法反而是真聪明?
本期介绍的几篇论文:
[AI] Accelerating Scientific Research with Gemini in the Real-World
[Google DeepMind & Duke University & Columbia University]
https://arxiv.org/abs/2608.26701
---
[CV] LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
[German Cancer Research Center & Mila]
https://arxiv.org/abs/2608.27395
---
[AI] AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design
[Hunan University & Nanjing University & The Chinese University of Hong Kong]
https://arxiv.org/abs/2608.26747
---
[AI] Decoupling Planning and Control for Instructable Agents
[UC Berkeley & Google DeepMind]
https://arxiv.org/abs/2608.26788
---
[AI] WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
[Google Research]
https://arxiv.org/abs/2608.27454
在小宇宙查看该单集文稿 - 你有没有想过,让AI变得更聪明,关键可能不是让它知道得更多,而是教会它如何更高效地“思考”?本期我们要聊的几篇最新论文,就深入到了AI的思维深处:从让AI懂得“选择性遗忘”以实现长时间推理,到揭开决定AI学习成败的三个神秘“开关”。我们甚至会看到,机器人是如何通过“自言自语”来规划复杂任务的。准备好一起探索AI大脑的内部运作机制了吗?我们马上开始!
00:00:33 如何让AI长时间思考,还不“累”?
00:05:05 给你一个确定性的菜谱,靠谱吗?
00:10:44 你关心的问题,AI能比专家更快找到答案吗?
00:16:24 拆开AI的“黑箱”,决定它聪明的三个开关
00:22:50 机器人会思考,需要分几步?
本期介绍的几篇论文:
[CL] Prefix Sliding for efficient test-time scaling
[Stanford University & University of California at Santa Barbara & University of Washington]
https://arxiv.org/abs/2608.26070
---
[LG] Beyond Optimal Rates in Stochastic Optimization: Trajectory-Adaptive Stopping Rules
[UC Berkeley & PSL Research University]
https://arxiv.org/abs/2608.25551
---
[AI] Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
[Google Research]
https://arxiv.org/abs/2608.26088
---
[LG] Demystifying Reinforcement Learning Post-Training of Language Models
[University of Washington]
https://arxiv.org/abs/2608.24949
---
[RO] R^3: Training Robots to Reason in Natural Language via Reinforcement Learning
[Carnegie Mellon University (CMU)]
https://arxiv.org/abs/2608.26053
在小宇宙查看该单集文稿 - 本期我们要聊点脑洞大开的:如果让一群AI自己组建科研社区,甚至给它们“放假”,会涌现出怎样的科学发现?我们会看到,AI真正的成长秘诀,不在于修正答案,而在于递归式地优化自己的“思考方法”,甚至学会像孙悟空一样用“分身术”同时探索多种可能。接着,当AI团队犯错时,我们将化身侦探,精准定位“责任人”,并揭秘一个让AI提速的妙招——不是靠堆算力,而是靠精明的“预算”分配。准备好了吗?让我们一起从几篇最新论文中,探寻这些关于AI工作流、团队协作与自我进化的深刻洞见。
00:00:43 AI也需要“放假”?科学发现的新模式
00:06:03 成长的秘密,不是优化答案,而是优化方法
00:10:58 让AI学会“分身术”,我们能快多少?
00:16:06 AI犯错,我们应该怪谁?
00:20:55 AI 为什么那么慢?这篇论文给了个巧妙的答案
本期介绍的几篇论文:
[AI] Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
[DualverseAI & University of California San Diego]
https://arxiv.org/abs/2608.23691
---
[AI] Metan^n: Recursive Self-Improvement through Emergent Depth
[University of Minnesota & Seoul National University]
https://arxiv.org/abs/2608.24735
---
[AI] Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning
[Tsinghua University & NVIDIA]
https://arxiv.org/abs/2608.24658
---
[CL] Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
[Bar-Ilan University & UNC Chapel Hill]
https://arxiv.org/abs/2608.24306
---
[CL] AgentSpec: Speculative Decoding for Batch Inference of LLM Agents
[The Ohio State University & Microsoft Research & University of Michigan]
https://arxiv.org/abs/2608.24004
在小宇宙查看该单集文稿
More Technology podcasts
Trending Technology podcasts
About AI可可AI生活
来自 @爱可可-爱生活 的第一手AI快报,用最简单易懂的语言,带你直击最前沿的人工智能科研动态。无论你是科技小白,还是行业达人,这里都有你想知道的AI故事和未来趋势。跟着我们,轻松解锁人工智能的无限可能!
#人工智能 #科技前沿
Podcast websiteListen to AI可可AI生活, The Vergecast and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


AI可可AI生活
Scan code,
download the app,
start listening.
download the app,
start listening.





























