

2026.08.20 | 闭环进化提升具身智能;验证门控保障工业代码【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:34] 🤖 Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence(Zetta ζ:面向自进化物理智能的高效闭环具身智能体框架) [01:29] ✅ SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation(SemaPLC:一种基于项目、以验证为门控的PLC代码生成智能体框架) [02:31] 🎯 SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation(SemComp-Bench:视频生成中语义任务完成的基准测试) [03:28] 🔬 OmniScientist: An Omni-Modal Omni-Discipline AI Scientist(全能科学家:一个全模态、全学科的人工智能科学家) [04:26] 🧠 Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL(Co-RL:多智能体强化学习中多样化群体催生无监督推理) [05:13] 🎮 SPADE: Self-Play in Adaptive Synthetic Executable Environments(SPADE:自适应合成可执行环境中的自博弈) [06:08] 🧪 Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis(训练面向单步逆合成的化学合理性感知大语言模型) [07:04] 🎯 Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning(潜在世界模型中的决策度量对齐:诊断与面向MPC规划的动作条件目标) [07:58] 🧬 Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification(训练留下痕迹:面向语言模型谱系验证的中心化残差签名) [08:50] 🔁 Looped Language Models Improve Compositional Tool Calling(循环语言模型提升组合式工具调用能力) [09:36] 🖐 SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation(SoftVTBench:面向可变形物体操作的变形感知视触觉数据集与基准) [10:33] ⚽ FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents(FM-Bench:面向竞争智能体的长时程管理基准) [11:25] ✍ Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion(借助属性引导的体裁扩展,将创意写作扩展到故事中心数据之外) [12:15] 🔥 The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning(越热门越难遗忘:大语言模型遗忘的自适应流行度方法) [13:01] 🔍 Temporal Multi-Signal Fusion for Token-Level Hallucination Detection(面向Token级幻觉检测的时序多信号融合) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.19 | 进化策略微调省显存;技能应用有双刃剑【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:34] 🧬 Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements(Agentic ESOpt:以极低GPU需求微调长程LLM智能体) [01:46] 🧩 Demystifying Agent Skills: Why They Work-Until They Don't(揭秘智能体技能:它们为何有效——直到失效) [02:35] 🔬 ASI-Bench: At the Dawn of Artificial Superintelligence(ASI-Bench:人工超级智能的黎明) [03:26] 💻 FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution(FreeToken:高效的边缘原生MoE服务与带宽自适应执行) [04:20] 🧭 Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation(具身导航器:指向、思考、记忆与对齐实现高效导航) [05:22] 🎬 AVA-Encoder: Towards Agent-Native Video Representation Learning(AVA-编码器:迈向智能体原生的视频表示学习) [06:12] 🖼 EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing(EDITBRIDGE:迈向忠实且高效的超高分辨率图像编辑) [07:13] ⚡ Agent Lightning v1.0: Towards Harnessed Agentic RL(Agent Lightning v1.0:迈向框架化的智能体强化学习) [08:12] 🎥 V-RAE: Rethinking Video Latent Spaces for Generation(V-RAE:重新思考用于生成的视频潜空间) [09:17] 🎬 CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing(CoinVE-200K:面向组合式指令引导视频编辑的大规模高质量数据集) [10:17] 🛡 DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization(DiSCO:通过分布引导的对比提示优化防御文本到图像生成) [11:22] 🧠 Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents(驾驭记忆:对记忆智能体中记忆底层介质的整体评估) [12:32] 🎨 From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation(从语料到协同演进的能力:以能力为中心的通用图像生成数据设计) [13:31] 📊 StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows(StartupBench:对通用智能体在市场验证的端到端工作流上的基准测评) [14:34] ⚡ Energy-Guided Flow Matching(能量引导的流匹配) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.18 | 智能体化评测让世界模型可诊断;多模态三维生成仍有瓶颈【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🕵 HarnessEval-W: Agentifying the Evaluation of Visual Worlds(HarnessEval-W:使视觉世界的评估智能体化) [01:32] 🌍 VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?(VibeWorlding:多模态智能体能否端到端构建3D开放世界?) [02:10] ⚡ MOSS-VL Technical Report(MOSS-VL 技术报告) [03:05] 🤖 ClawGym II: Exploring Black-Box RL on Agent Harness(ClawGym II:在智能体框架上探索黑盒强化学习) [03:56] 🎯 Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization(学习尚未掌握的,而非已经精通的:面向多奖励策略优化的饱和感知优势重加权) [04:51] 🤖 UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations(UI-Mate:利用上下文演示推进开放权重的基础图形用户界面智能体) [05:46] 🔬 Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search(大型发现模型:基于经验建模的开放式搜索) [06:37] 🎨 An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models(训练像素空间文本到图像扩散模型的实证研究) [07:31] 🤖 Agentic Transaction: Towards ACID-Compliant Agent Systems(智能体事务:迈向ACID合规的智能体系统) [08:29] 🔬 How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks(智能体如何在自动研究中失败:基于100个真实前沿研究任务的端到端诊断评估) [09:24] ⚡ GenRouter: Unified Workflow Routing for Agentic Image Generation(GenRouter:用于智能体图像生成的统一工作流路由) [10:29] 🧩 MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling(MegaParts:通过词元高效的自回归建模将部件感知的3D物体生成扩展到300个部件) [11:22] 🧠 Understanding Cognition-Induced Risks in Agentic AI Systems(理解智能体AI系统中认知引发的风险) [12:21] 🔗 Advancing Open and Reproducible Relational Learning: RelArena-$α$, TabPFN-Rel and RPI(推进开放可复现的关系学习:RelArena-α、TabPFN-Rel与RPI) [13:24] 🛡 Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs(Ventor-QTest:威胁模型驱动的供应商托管LLM API验证) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.17 | 视频检测难防伪;自监督蒸馏促提升【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🛡 Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination(我们能防御针对现实世界危机事件的AI生成视频攻击吗?对检测器、生成器与社会传播的系统评估) [01:27] 👁 Self-Supervised Visual On-Policy Distillation(自监督视觉同策略蒸馏) [02:21] 🤖 Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development(超越最终得分:对长周期AI研究与开发智能体的系统评估) [03:15] 🧠 Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning(Intern-S2-Mobius:知识与推理解耦的基础模型) [04:09] 🎮 Marionette: Predicting World States, Rendering Geometry, Painting Appearance(Marionette:预测世界状态,渲染几何,绘制外观) [05:00] 🧠 SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning(SimpleOPD:面向长上下文推理的简单分词器无关在线策略蒸馏) [06:03] 🧠 MobileMem: Learning from a Year of Mobile Experiences(移动记忆:从一年的移动体验中学习) [06:54] 🤖 DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data(DFM Mimir v1:仅使用合规后训练数据、以1B参数实现前沿性能的开放HRM) [07:50] 🤸 HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark(HumanTracker:迈向全面且与人类感知一致的运动追踪基准) [08:50] 🎨 CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing(CPI-Bench:一个面向真实世界图像编辑的全面、实用且智能的基准) [09:48] 🧠 Latent On-Policy Self-Distillation(潜在同策略自蒸馏) [10:53] 🔍 Claim-Level Reliability Assessment for Efficient Test-Time Reasoning(面向高效测试时推理的声明级可靠性评估) [11:43] 🤖 PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment(PRM-as-a-Judge 1.5:机器人过程评估工具包) [12:42] 📉 Forecast Collapse in Time-Series Foundation Models(时间序列基础模型中的预测崩溃) [13:41] 🤔 Second Thought: Reasoning in Parallel as LLM Agents Act and Observe(第二思考:LLM智能体在行动与观察时并行推理) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
【周末特辑】8月第3周最火AI论文 | 潜空间推理降本;持续学习促进化【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 5 篇论文如下: [00:46] TOP1(🔥639) | 🧠 BDH-CQ: In-Context Learning with Recurrent Latent Reasoning(BDH-CQ:基于循环潜在推理的上下文学习) [03:27] TOP2(🔥332) | 🔄 Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA(Macaron-V1:迈向具备自我改进和LoRA混合的开放持续学习) [06:39] TOP3(🔥278) | 📄 Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill(从火花到论文:作为可组合技能的端到端研究论文生成) [10:15] TOP4(🔥255) | 🧬 OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution(OpenART:通过开放式环境演化扩展智能体红队测试) [13:01] TOP5(🔥208) | 🧠 On-Policy Self-Distillation without Any Supervision(无需任何监督的在线策略自蒸馏) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.14 | 动作条件视频世界模型引入几何感知;长时记忆外部化实现无尽世界【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🤖 DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation(DreamX-Phi 1.0:面向机器人操作的动作条件视频世界模型) [01:24] 🌍 Alaya-EVOKE: From Linear-Scaling Supervision to Endless World(Alaya-EVOKE:从线性扩展监督到无尽世界) [02:30] 🔀 LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers(LLMRouter:开发、评估和部署LLM路由器的统一基础设施) [03:25] 🧬 DarwinX: Evolving Agent Harnesses Through Natural Selection(DarwinX:通过自然选择进化智能体框架) [04:26] 🔬 Intern-S2-Preview: Scientific Agentic Foundation Model(Intern-S2-Preview:科学智能体基础模型) [05:20] 🎮 PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives(PlayWorld:使用智能体玩家在长程目标上对世界模型进行基准测试) [06:20] 🤖 AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design(AutoDesign:面向长时程智能体设计的元框架优化) [07:23] 🧠 Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence(空间记忆智能体:基于经验的程序记忆实现空间智能) [08:12] ⚡ Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus(混合线性注意力大语言模型中的大规模激活:注意力前尖峰与尖峰间平台) [09:03] 🎭 UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos(UniSwap:面向说话视频的流式音频-视觉身份交换) [10:12] ⚡ LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time(LiveAnimate:实时稳定长格式流式人体动画生成) [11:08] 🤖 How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review(修辞何以能对AI审稿人进行奖励黑客?解析基于AI的同行评审中的修辞敏感性) [12:00] ✂ An AI4AI Framework for Visual Token Pruning(面向视觉Token剪枝的AI4AI框架) [12:56] 🤖 H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models(H2R-Bench:在世界模型中评估人类到机器人的操作视频生成) [14:08] 🔄 Full-bandwidth transformer(全带宽Transformer) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.13 | 演化环境揭示智能体风险;组合技能实现论文生成【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:26] 🧬 OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution(OpenART:通过开放式环境演化扩展智能体红队测试) [01:35] 📄 Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill(从火花到论文:作为可组合技能的端到端研究论文生成) [02:43] 🧩 AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses(测试时AI4AI:通过推理支架实现强到弱能力迁移) [03:37] 🔬 Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence(Mechanist:将人工智能作为揭示智能机制的科学仪器) [04:33] 🎭 Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives(大语言模型智能体能否坚守剧本?交互式叙事中长程一致性的基准测试) [05:33] 🌍 StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization(StateFlow:为预可视化构建、演化与访问3D世界状态) [06:28] 📐 Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models(自几何:面向几何一致的3D视觉基础模型的无真值即插即用测试时自适应) [07:25] 🪞 From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection(从合成到去除:物理驱动的反射模拟与基于扩散模型的视频去反射) [08:22] 🛡 ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents(ToolHazard:扩展对抗性环境,用于基于大语言模型的智能体的安全评估与对齐) [09:15] 🔍 The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images(视觉工具使用的幻象:图像思维的因果审计) [10:16] 🤖 Self-Evolving Embodied Agents via Skill-Harness Evolution(基于技能与执行框架演化的自进化具身智能体) [11:19] 🧩 SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries(SkillZip:面向可扩展智能体技能库的合约保持图压缩) [12:14] 🛡 Agent Safety Should Be a Runtime Contract(智能体安全应当是一种运行时契约) [13:01] 🧬 Persistent Recursive Worlds Enable Autonomous Software Evolution(持久递归世界赋能自主软件演化) [13:46] 💡 MBA: Multimodal Benchmark and Agents for Real-World Business Ideation(MBA:面向真实世界商业构思的多模态基准与智能体) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.12 | 共体智能体以人为中心助人成长;智能体与环境共演化迈向自我导向【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:35] 🤖 ComBodied Agents: a New Paradigm of Human-Centric Agentic AI(共体智能体:以人为中心的智能体人工智能新范式) [01:22] 🧬 Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design(智能体系统中的共同演化:迈向超越人类设计的自我导向演化) [02:22] 🌍 Beyond Pixels: From Video Priors to 4D Worlds(超越像素:从视频先验到4D世界) [03:12] 🧩 Articulated Object Reconstruction from Rest-State Observation(基于静止状态观测的铰接物体重建) [04:09] ⚔ AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss(AdvFD:通过对抗性弗雷歇距离损失提升视觉生成) [05:05] 🧬 Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution(孟德尔·哥德尔机:通过比较进化实现递归自我改进的编码智能体) [06:00] 🎭 Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence(Ex-Omni-2D:具备原生视觉临场感的表现性全模态对话模型) [06:48] 🌍 VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?(VibeLifeBench:你的生活智能体能否在动态世界中主动且持久地行动?) [07:46] 🚫 Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness(解码级禁忌:大语言模型鲁棒性的诊断性压力测试) [08:33] 📦 SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure(SkillZip:通过发现可复用结构实现自我进化智能体的免评估技能压缩) [09:33] 📱 SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information(SPIEval:评估大型语言模型作为移动助手处理分散个人信息的能力) [10:36] ✂ Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents(并不值得再投入一个词元:高效深度研究智能体的边际价值估计) [11:34] 🌐 Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation(开放大语言模型用于多语言机器翻译的无参考后训练) [12:28] 🔍 InSight-doc: Agentic Visual Perception for Long-Document Understanding(InSight-doc:面向长文档理解的智能体视觉感知) [13:27] 🔀 UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models(UniMoMo:基于专家合并的大型推荐模型MoE加速方法) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.11 | 自进化混合专家赋能持续学习;代码重构基准揭示智能体局限【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🔄 Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA(Macaron-V1:迈向具备自我改进和LoRA混合的开放持续学习) [01:24] 🔧 SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring(SWE-Bench ProMax:面向大规模多语言代码重构的智能体基准评测) [02:21] 🐍 Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution(Ouroboros:通过核心评审进化实现自我发展的前沿编程智能体) [03:31] 🧠 BDH-CQ: In-Context Learning with Recurrent Latent Reasoning(BDH-CQ:基于循环潜在推理的上下文学习) [04:13] 🧠 Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory(智能体记忆蒸馏:利用分层教师记忆赋能小型大语言模型智能体) [05:11] 🧠 Motif 3: Technical Report(Motif 3:技术报告) [05:59] 🔬 Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains(Sci-VBench:评估科学领域中知识与推理密集型视频生成) [06:49] 🖼 What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems(下一步编辑什么:对话系统中的视觉对齐图像编辑后续建议) [07:53] 🎯 SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation(SPOT:面向同策略蒸馏的稀疏探测与结果校准) [08:58] ⚡ OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching(OasisKV:通过前瞻稀疏预取将解码期KV缓存扩展到HBM之外) [09:49] 🧠 RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States(RoMeRL:通过降阶效用状态平衡自进化智能体记忆中的反馈覆盖与记忆-奖励陷阱) [10:43] 🔍 Evidence-RL: Towards Evidence-intensive Visual Reasoning(证据强化学习:迈向证据密集型视觉推理) [11:42] 🧠 Scaling Inherently Interpretable Language Models(扩展内在可解释的语言模型) [12:40] 🧬 Evo-Bench: Can Language Models Improve Agent Harness?(Evo-Bench:语言模型能否改进智能体运行框架?) [13:40] 🔓 Stealing Reasoning Traces from Proprietary LLM APIs(从专有大语言模型API中窃取推理轨迹) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.10 | 多模态智能体环境设计重质轻量;强化学习利于多任务共存【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🌍 Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning(超越单纯的环境规模扩展:为多模态智能体学习设计有效的环境分布) [01:36] ⚔ SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs(SFT冲突、RL共存:大语言模型多任务学习的理论与实证分析) [02:42] 🚗 SimWAM: A Simple World Action Model for End-to-End Autonomous Driving(SimWAM:面向端到端自动驾驶的简单世界动作模型) [03:40] 🎯 YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family(YOLO-PEFT:YOLO系列上的参数高效微调) [04:43] 🎥 StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding(StreamArena:迈向连续、交互与长时程的智能体流式视频理解) [05:36] 🎧 Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning(以演化评分标准作为奖励的音频推理强化学习) [06:19] 🙈 When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles(当激活预言机学会不读取:微调预言机中的概念特异性盲区) [07:09] ⚡ Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss(面向大语言模型的高效知识蒸馏:离线Top-K Logits与融合分块KL损失) [08:16] 🚁 Uncertainty-Aware World Model for Aerial Image-Goal Navigation(面向航拍图像目标导航的不确定性感知世界模型) [09:19] 🎯 Douyin Multimodal Embedding Model Technical Report(抖音多模态嵌入模型技术报告) [10:19] 📈 Skaling: Chinchilla's Exponents Meet Kaplan's Coupling(Skaling:Chinchilla指数与Kaplan耦合的融合) [11:11] 🧠 Addressable Memory for Video World Models(视频世界模型的可寻址记忆) [12:05] 🧩 Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression(相关但不完整:硬提示压缩中作为范式级失败模式的指称悬空) [13:01] 🔄 Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors(往返一致性:双向扩散模型可预测自身的展开误差) [14:01] 🧠 The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows(优化器即智能体:跨提示、程序与机器学习工作流的推理驱动搜索) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
【周末特辑】8月第2周最火AI论文 | 递归合成扩数据;状态管理提长程【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 5 篇论文如下: [00:41] TOP1(🔥221) | 🔁 Recursive Synthesis for Long-Horizon Terminal Tasks(面向长时程终端任务的递归式合成) [03:39] TOP2(🔥162) | 🧭 LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks(LongHorizon-Harness:推动面向真实世界任务的长时程智能体) [06:29] TOP3(🔥154) | 🎙 SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks(SwanTale:面向指令与零样本任务的统一多说话人语音与音频生成) [09:28] TOP4(🔥140) | 🚗 Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs(自动驾驶视觉语言模型中用于可验证推理的未来轨迹延迟暴露) [12:30] TOP5(🔥108) | ⚓ DAPD: Dual-Anchored Policy Distillation(双锚定策略蒸馏) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.07 | 递归自蒸馏重塑智能体信用;开源裁判低成本评估操作【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🎯 AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning(AgentOPSD:面向智能体强化学习的递归自蒸馏) [01:35] 🤖 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models(OSReward:为跨平台计算机使用奖励模型制定标准化评估) [02:24] 🌍 WorldClaw: Agentic 3D Open-World Generation at Scale(WorldClaw:大规模智能体式3D开放世界生成) [03:12] 🗺 GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?(GST-Bench:视觉语言模型能否从视频中形成全局空间意识?) [04:17] 💭 EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning(EnvACE:通过世界预演将环境动态内化于智能体强化学习) [05:10] 🔍 Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval(从失败中学习:基于硬负样本的检索中心思维链用于统一多模态检索) [06:08] ⏳ ChronoVision: Temporal Reasoning via Latent State Reconstruction(ChronoVision:通过潜在状态重建实现时序推理) [07:08] 🌐 From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models(从经济主体到主体经济:经济世界模型的系统蓝图) [08:10] 🧮 On-Policy Delta Distillation for Multilingual Math Reasoning(面向多语言数学推理的同策略差值蒸馏) [08:59] ⚙ HarnessOpt-Bench: Evaluating LLMs at Harness Optimization(HarnessOpt-Bench:评估大语言模型在智能体运行框架优化上的表现) [09:43] 🇬 Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains(教Nemotron希腊语:面向专业领域的现代希腊语语料挖掘、检索适配与有据生成) [10:43] 🤖 DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation(DyPES-VLA:学习共享动力学先验与具身特定控制以实现跨具身操作) [11:38] 🤖 World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation(世界到手腕:面向精细机器人操作的任务条件化未来手腕建模) [12:44] 📊 DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces(DataSpace:面向异构工作空间的可验证分析数据代理基准测试) [13:44] 🪄 EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal(EffectLearner:面向真实世界视频目标移除的世界感知对象-效应推理) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.06 | 回溯答案训练搜索智能体;统一工具实现图像生成【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🔍 ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment(ABSeeker:通过答案回溯的信用分配训练长程搜索智能体) [01:25] 🎨 ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation(ToolArtist:面向智能体图像生成的工具使用统一多模态模型) [02:16] 🪞 The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads(个性化幻象:大语言模型如何捏造用户画像,以及自我监控为何具有误导性) [03:17] ⚛ Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes(迈向多模态预训练的物理机制:知识流动、模态协同、早期统一与配方) [04:20] 🤖 OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents(OneDayAgent:面向自主智能体的长周期任务执行框架) [05:14] 🧬 GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks(GDPevo:在真实业务任务中评估智能体的自我进化) [06:05] 🎯 When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation(当教师误导:虚假信号感知的同策略蒸馏) [06:56] 🧩 Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning(迈向技能原生的大语言模型:用于长程推理评测与训练的技能熵) [08:02] 🧩 NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap(NOLLI:用于诊断英韩性能差距的难度校准谜题基准) [08:57] 🤖 Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data(Ego2Robot:从第一人称人类数据中可扩展合成机器人数据) [09:54] 🎬 AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities(AVE-Compass:面向音视频编辑能力的全面评估) [10:55] 🧠 When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents(当记忆说谎:VLM智能体中空间记忆陈旧性的实证研究) [11:53] 🎯 Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance(在失败处蒸馏:利用自适应教师指导恢复负强化学习组的学习信号) [12:49] 🧠 FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory(FocusMem:分解潜在GUI记忆中的内容、读取与信任) [13:52] 👋 HelloWorld: Enabling Socially Interactive Characters in Video World Models(HelloWorld:在视频世界模型中实现社交互动角色) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.05 | 智能体长期运营后净资产不足人类三成;实时视频编辑达高清流畅【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:27] 🛒 MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations(MerchantBench:电商运营中LLM智能体长期一致性的基准评测) [01:19] 🎬 JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion(JoyAI-Video-Edit:利用自回归扩散的实时开放式视频编辑) [02:19] 🧊 Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing(Hunyuan3D-Buffalo 1.0:一种可扩展的三维生成、理解与编辑统一多模态模型) [03:15] 🌌 AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling(AURORA-LM:自编码统一表示用于连续潜变量扩散语言建模) [04:18] 🎥 Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent(视频深度研究:迈向下一代多模态深度研究智能体) [05:15] 🔄 Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation(知识-几何解耦:面向流式推荐的可刷新预训练迁移) [06:20] 🤖 PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning(PCSD:智能体强化学习中自蒸馏的持久一致性) [07:08] 🌍 Quo Vadis, World Modeling?(世界建模,何去何从?) [07:59] 🔁 PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents(PAST-Bench:对个人智能体中递归自我改进的基础进行基准测试) [08:56] 🌉 Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging(Any-OPD:通过表示空间桥接实现异构流匹配模型的在策略蒸馏) [09:57] 🗜 OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models(OmniPack:用于高效全模态大语言模型的统一令牌压缩) [11:02] 📈 LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models(LLaDA MoE v2:扩展混合专家扩散语言模型) [12:00] 🎯 CAPEval: A Decoupled Caption Evaluation across Understanding and Generation(CAPEval:面向理解与生成的解耦式图像描述评估) [12:55] 🧩 UniWorld-Design: From Pixel Generation to Layer-Native Design(UniWorld-Design:从像素生成到图层原生设计) [13:56] 🧠 SkillJack: Persistent Skill Backdoors in Self-Evolving Agents(SkillJack:自进化智能体中的持久性技能后门) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.04 | 长时程智能体成功率提升至八成;多说话人语音音频统一生成【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🧭 LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks(LongHorizon-Harness:推动面向真实世界任务的长时程智能体) [01:30] 🎙 SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks(SwanTale:面向指令与零样本任务的统一多说话人语音与音频生成) [02:31] 🎯 VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation(VAD:在多模态在策略蒸馏中为目标重建归因视觉证据) [03:37] 🤖 Progressive Agent Skill Generation via Reinforcement Learning(基于强化学习的渐进式智能体技能生成) [04:39] ⚓ DAPD: Dual-Anchored Policy Distillation(双锚定策略蒸馏) [05:37] 🧲 UEmbed: Unified Sparse and Dense Multimodal Embeddings(UEmbed:统一稀疏与稠密多模态嵌入) [06:44] 🌍 WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity(WorldExam:从表象外观到内在反应性的世界模型基准评测) [07:54] 🔗 CADENA: Stepwise CAD Reverse Engineering(CADENA:逐步式CAD逆向工程) [08:55] 🛠 SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation(SKT:通过经验证的合成数据生成实现规模化技能使用训练) [09:58] 🤖 SWE-Touch: Benchmarking Coding Agents When Users Touch the Code(SWE-Touch:在用户改动代码时对编码代理的基准测试) [11:08] 🚗 Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs(自动驾驶视觉语言模型中用于可验证推理的未来轨迹延迟暴露) [12:09] 🧠 WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning(WCM:面向视觉-语言-动作强化学习的世界评论家模型) [13:05] 🔄 Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations(超越形态的运动:从抽象运动表征引导跨类别运动迁移) [14:04] 🧠 GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning(GradCuit:信用分配的梯度流实现稳健且可解释的测试时潜在推理) [14:59] 🛋 Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis(Roomer:面向三维室内布局合成的反思式对象级模型编辑与修复) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递