2026.10.01 | RIDE外推教师残差;UniEvo-VL在线自蒸馏多模态

2026.10.01 | RIDE外推教师残差;UniEvo-VL在线自蒸馏多模态

12分钟 ·
播放数11
·
评论数1

【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 www.xiaoyuzhoufm.com

【目录】
本期的 15 篇论文如下:

00:27 🧭 The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation(教师是方向,而非终点:在在线策略蒸馏中外推强化学习诱导的表征残差)
01:11 🪞 UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement(UniEvo-VL:面向多模态模型自我改进的在线策略自蒸馏训练方案)
01:54 🕵 False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents(虚假前沿:诊断与缓解自演化搜索智能体中的共同作弊)
02:42 🤖 AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks(AREX-2:通过长时程反思任务推进自我改进智能体)
03:29 🖥 Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents(Mid-Harness:在模型与执行框架之间扩展终端智能体的动作)
04:09 🧬 EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery(EvoDuet:面向科学发现的网络搜索与任务求解双层协同演化)
05:01 🕵 WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents(WorldAuditBench:使用多模态智能体进行交互式3D世界审计)
05:47 🛠 Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI(测试时 AI4AI 中面向智能体执行框架设计的元技能学习)
06:26 🧠 EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making(EVOKE:激发智能体中的世界知识以实现可迁移决策)
07:12 🎮 RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement(RSIGame:具备递归自我改进能力的自主智能体游戏开发)
07:53 📉 More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models(更多选择,更少决策:类JEV直接决策模型中的序数尺度偏差)
08:45 🧠 Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering(Imagine3D-LLM:教会多模态大语言模型在回答前想象3D场景)
09:31 🛠 Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training(智能体错误数据集:面向失败分析与错误感知后训练,规模化构建5万条错误—诊断配对)
10:21 🏮 LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models(LANTERN:照亮语言模型中的隐藏数学知识)
11:05 🖼 It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them(并非图像所示:无关上下文会扰乱 VLM 评判模型却不为其提供信息)

【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递

展开Show Notes
Neverlandrvr
Neverlandrvr
5小时前
10:16 突然变声了