[人人能懂AI前沿] 从精准反馈、高效协作到群体智慧

[人人能懂AI前沿] 从精准反馈、高效协作到群体智慧

28分钟 ·
播放数114
·
评论数0

你有没有觉得,最聪明的AI有时也会犯一些“低级错误”?本期节目,我们就从几篇最新论文出发,去看看AI那些意想不到的“脆弱时刻”。我们将一起探索,为什么AI合作有时会“1+1<2”,甚至被少数派“带偏”;又为什么一个不起眼的错别字,就能让它瞬间“走神儿”。更重要的是,我们将看到科学家们如何像一位“自动马鞍匠”一样,为AI打造不断进化的外部装备,又如何通过“字斟句酌”的反馈,教会AI抵御外界的恶意指令。

00:00:36 如何给AI装上一个“自动升级”的马鞍?

00:05:13 为什么笼统的批评没用?从教AI“防骗”的底层逻辑说起

00:10:36 1+1 < 2?合作的隐形成本

00:15:48 一个好汉三个帮,AI为何越帮越忙?

00:20:58 为什么一个错别字,就能让AI“走神儿”?

本期介绍的几篇论文:

[AI] AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

[POSTECH & KAIST & Southern University of Science and Technology]

arxiv.org

---

[AI] SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

[UC Berkeley]

arxiv.org

---

[CL] The Collaboration Tax: How Much LLM Multi-Agent Systems Pay to Coordinate

[University of Notre Dame & Meta Superintelligence Labs]

arxiv.org

---

[CL] Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations

[ETH Zurich & Tel Aviv University]

arxiv.org

---

[CL] Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

[Missouri University of Science and Technology & University of North Texas]

arxiv.org