[人人能懂AI前沿] 从状元策略、长链陷阱到悬崖学习

[人人能懂AI前沿] 从状元策略、长链陷阱到悬崖学习

28分钟 ·
播放数73
·
评论数0

本期,我们来聊聊AI如何从一个“普通学生”被系统地培养成编程竞赛的世界冠军,甚至超越了人类状元。但与此同时,为什么我们身边的AI助理,处理复杂任务时却常常“走着走着就散架”了?我们又该如何教会AI管理自己的“注意力”,像人一样划重点?以及,如何通过精准定位它“第一次犯错的瞬间”,让它的学习效率实现飞跃?四篇最新论文,带我们深入AI的“学霸心法”,揭示智能背后的策略、局限与成长之道。

00:00:37 AI学会考试了,而且比状元考得还好

00:06:06 你的AI助理,为啥走着走着就“散架”了?

00:11:30 AI的注意力,该由谁做主?

00:16:47 如何让机器学会聪明,抓住第一次犯错的瞬间

00:22:08 知识的“断舍离”,我们究竟该记住什么?

本期介绍的几篇论文:

[LG] Post-Training Language Models for Gold-Medal Performance in Coding Competitions

[NVIDIA]

arxiv.org

---

[AI] How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

[Microsoft AI]

arxiv.org

---

[CL] Language Models Can Control Their Own Attention

[KAIST AI & Google DeepMind]

arxiv.org

---

[LG] Cliff: Learning Process Rewards from the First Mistake

[Amazon Web Services]

arxiv.org

---

[LG] What Is Worth Representing? Representational Empowerment for Continual Model Construction

[UC Berkeley & University of Tübingen]

arxiv.org