

【周末特辑】10月第1周最火AI论文 | 自演化搜索防共作弊;Raven编排可组合智能体【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 5 篇论文如下: [00:50] TOP1(🔥609) | 🕵 False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents(虚假前沿:诊断与缓解自演化搜索智能体中的共同作弊) [02:49] TOP2(🔥563) | 🤖 Raven: The Harness of Harnesses for Composable Agentic Intelligence(Raven:面向可组合智能体智能的“框架之框架”) [05:15] TOP3(🔥532) | 🧭 The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation(教师是方向,而非终点:在在线策略蒸馏中外推强化学习诱导的表征残差) [07:52] TOP4(🔥467) | 🔁 LoopVL: Recurrent Visual Intelligence(LoopVL:循环视觉智能) [09:59] TOP5(🔥411) | 🎨 MaLiang-Harness: A Programmable Path to Image and Video Generation(MaLiang-Harness:通往图像与视频生成的可编程路径) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.10.02 | 流式视频主动记忆;音视频扩散奖励路由【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:29] 🎥 OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction(OneStreamer:统一流式视频交互中的感知、记忆与主动响应) [01:11] 🔀 Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL(自适应奖励路由:通过前向过程强化学习实现联合音视频扩散的动态多奖励优化) [01:55] 🧠 Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States(超越记忆:利用显式信念状态驾驭长时程智能体) [02:38] 🤖 Agent Priors-guided Policy Learning(智能体先验引导的策略学习) [03:25] 🌀 Hierarchical Continuous Diffusion Language Models(分层连续扩散语言模型) [04:03] 👁 World Observer: Joint Actor-Observer Generation for Persistent World Modeling(World Observer:面向持久世界建模的联合行动者-观察者生成) [04:54] 📉 Sharpening Tax in Post-Training(后训练中的锐化税) [05:39] 🎯 ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization(ActiveSaddler:面向智能体执行框架优化的自动化课程学习) [06:22] 🤖 A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review(可信AI审稿人缺失的一环:从修辞鲁棒性基准测试到SciCore评审) [07:08] 🎮 ROWBench: Do Video Models Render What the Program Specifies?(ROWBench:视频模型能否渲染程序所指定的内容?) [07:55] 🤖 AutoGUIWorld: Image Generators as Visual World Models for GUI Agent(AutoGUIWorld:将图像生成器用作 GUI 智能体的视觉世界模型) [08:39] 🔍 Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation(基于跨执行框架适配的检索增强技能优化) [09:23] ⚖ Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL(让稀疏奖励算数:多奖励强化学习中的密度感知奖励聚合) [10:05] 🤖 Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding(去中心化 Master-Mind:多智能体路径规划中通过迭代意图去噪的联合动作精炼) [11:02] 🧩 E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models(E-MoE:面向非因子化扩散语言模型的增强混合专家) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.10.01 | RIDE外推教师残差;UniEvo-VL在线自蒸馏多模态【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:27] 🧭 The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation(教师是方向,而非终点:在在线策略蒸馏中外推强化学习诱导的表征残差) [01:11] 🪞 UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement(UniEvo-VL:面向多模态模型自我改进的在线策略自蒸馏训练方案) [01:54] 🕵 False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents(虚假前沿:诊断与缓解自演化搜索智能体中的共同作弊) [02:42] 🤖 AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks(AREX-2:通过长时程反思任务推进自我改进智能体) [03:29] 🖥 Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents(Mid-Harness:在模型与执行框架之间扩展终端智能体的动作) [04:09] 🧬 EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery(EvoDuet:面向科学发现的网络搜索与任务求解双层协同演化) [05:01] 🕵 WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents(WorldAuditBench:使用多模态智能体进行交互式3D世界审计) [05:47] 🛠 Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI(测试时 AI4AI 中面向智能体执行框架设计的元技能学习) [06:26] 🧠 EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making(EVOKE:激发智能体中的世界知识以实现可迁移决策) [07:12] 🎮 RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement(RSIGame:具备递归自我改进能力的自主智能体游戏开发) [07:53] 📉 More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models(更多选择,更少决策:类JEV直接决策模型中的序数尺度偏差) [08:45] 🧠 Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering(Imagine3D-LLM:教会多模态大语言模型在回答前想象3D场景) [09:31] 🛠 Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training(智能体错误数据集:面向失败分析与错误感知后训练,规模化构建5万条错误—诊断配对) [10:21] 🏮 LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models(LANTERN:照亮语言模型中的隐藏数学知识) [11:05] 🖼 It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them(并非图像所示:无关上下文会扰乱 VLM 评判模型却不为其提供信息) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
【月末特辑】9月最火AI论文 | LimiX-2结构化智能;Vidu S2实时可编辑空间视频【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 10 篇论文如下: [00:38] TOP1(🔥810) | 🧠 LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence(LimiX-2:面向通用结构化数据智能的上下文机制网络) [02:36] TOP2(🔥704) | 🎬 Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation(Vidu S2:实时交互、可编辑与空间视频生成) [04:41] TOP3(🔥494) | 🤖 Raven: The Harness of Harnesses for Composable Agentic Intelligence(Raven:面向可组合智能体智能的“框架之框架”) [06:43] TOP4(🔥493) | 🎓 StudentSim: Training LLM-based Student Simulators(StudentSim:训练基于大语言模型的学生模拟器) [08:51] TOP5(🔥484) | 🤖 Scaling Automatic Research Agents via World Models(通过世界模型扩展自动研究智能体) [11:10] TOP6(🔥431) | 🤖 Atria Dawn: The Dawn of Agentic Superintelligence(Atria Dawn:智能体超级智能的黎明) [13:35] TOP7(🔥396) | 🤖 Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills(从仓库到技能:将GitHub代码库蒸馏为AI4AI技能) [15:38] TOP8(🔥382) | 🎨 MaLiang-Harness: A Programmable Path to Image and Video Generation(MaLiang-Harness:通往图像与视频生成的可编程路径) [17:42] TOP9(🔥376) | 🧠 Continual Learning Mechanisms Compose for Long-Horizon Memorization(持续学习机制组合用于长时程记忆) [19:55] TOP10(🔥358) | 🤖 In-Context Learning for Robots: Methods and Applications(面向机器人的上下文学习:方法与应用) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
【周末特辑】9月第4周最火AI论文 | 世界模型训练客体永久性;OmniEdu 教学基础模型【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 5 篇论文如下: [00:48] TOP1(🔥236) | 🧠 Training Object Permanence in World Models(在世界模型中训练客体永久性) [03:10] TOP2(🔥235) | 🎓 OmniEdu: Open Foundation Models for Learning and Teaching(OmniEdu:面向学习与教学的开放基础模型) [05:22] TOP3(🔥218) | 🧬 RRSI: Regularized Recursive Self-Improvement of Agent Harnesses(RRSI:智能体运行框架的正则化递归自我改进) [07:44] TOP4(🔥217) | 🎙 Realtime-Venus: A full-duplex interaction system with asynchronous delegation(Realtime-Venus:一种支持异步委派的全双工交互系统) [09:47] TOP5(🔥160) | 🧭 The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks(有品味的智能体:长时程任务中的品味测量与提升) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.09.28 | 层融合正则缩小重建生成差;血缘数据流加速保序容错【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:28] 🧩 FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders(FuseReg:正则化层融合缓解表征自编码器中的重建-生成差距) [01:12] 🔗 RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation(RayOrch:面向基础模型数据准备的受血缘控制多粒度数据流编程与执行) [01:58] ⚡ Block Sparse Attention with Log-Linear Complexity(对数线性复杂度的块稀疏注意力) [02:40] 🤖 InternW0-$Δ$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data(InternW0-Δ:一个连接预测动态与动作、基于20K+小时开放数据的世界动作模型) [03:24] 🖐 Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors(Tactile-JEPA:面向分布式触觉传感器的拓扑感知自监督表示学习) [04:11] 🛰 Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning(利用预训练扩散模型与多模态条件增强摄影测量数字表面模型) [04:58] 👁 FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance(FoMo:生成轨迹中的分叉时刻作为感知距离) [05:42] 📊 Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem(真实环境中的 Jev:Jev 模型功能、应用与生态的数据驱动分析) [06:30] 🎯 TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations(TrackEverything:通过去重3D场景表示实现长时程密集跟踪) [07:21] 🧩 SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL(SLCA-GRPO:解决工具调用强化学习中的跨段信用误归因) [08:08] 🎯 CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation(CARD:面向个性化文本生成的聚类级适配与奖励引导解码) [08:54] ⚖ Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs(隐式个性化与显式风格会冲突吗?PsPLUG:用于平衡定制化 LLM 中个性化与风格的轻量级插件) [09:41] 🤝 AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs(AgentWorld:多智能体大语言模型长时程协作基准测试) [10:24] 🎮 Game Arena: Strategic LLM Evaluation in Competitive Environments(游戏竞技场:竞争环境中的大语言模型策略评估) [11:08] 🏦 IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking(IndicBankBench:评估印度零售银行中语言模型助手的安全性与可靠性) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.09.25 | 世界模型物理推理可评测;大模型线性叠加双想法【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:27] 🧠 Training Object Permanence in World Models(在世界模型中训练客体永久性) [01:12] 🧠 Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs(你的Transformer能同时容纳两个想法:LLM中线性叠加的证据) [01:53] 🎬 WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation(WanPE:面向现代文本到视频生成的电影级提示增强) [02:40] 🧭 OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents(OmniEcho:面向全模态具身智能体的音视频空间理解) [03:18] 🤖 Agent-Editing World Model: Rethinking World Modeling for LLM Agents(智能体编辑世界模型:为LLM智能体重新思考世界建模) [04:05] 🧩 Parts-of-Speech as Emergent Categories in SAE Latent Space(词性作为SAE潜空间中的涌现类别) [04:46] 🤖 Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents(Qwen-Planner-Agent:面向真实世界移动规划智能体的闭环 AI-for-AI 框架) [05:20] 🔍 IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis(IterSynth:通过角色解耦的迭代合成重新思考深度搜索智能体) [06:04] 🤖 Coding Agents for Generalized Task and Motion Planning Problems(面向泛化任务与运动规划问题的编码智能体) [06:49] 🧠 Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone(神经谱容量:仅凭网络规格测量与设计架构) [07:28] 🛡 AgentKernel: The Trust-Native Agentic Operating System(AgentKernel:信任原生的智能体操作系统) [08:15] 🧪 Rufus-Air: An Open LLM Post-Training Recipe(Rufus-Air:开放的大语言模型后训练配方) [08:59] 🤖 PUBG Ally: A Conversational Embodied Agent as an AI Teammate(PUBG Ally:作为AI队友的对话式具身智能体) [09:44] 🤖 World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal(世界动作智能体:利用视觉语言模型通过世界动作预演实现机器人操作) [10:27] 🛸 ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds(ExplorationBench:测量 AI 系统在可验证异星世界中的探索能力) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.09.24 | 说话者双轨记忆;交互学习空间推理【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:28] 🧠 SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue(SpeakerMem-R1:以说话者为中心的多方对话双轨记忆) [01:13] 🤖 Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World(Spatial-Interactor:通过与可观测物理世界交互学习空间推理) [01:59] 🌍 HappyWorld-Bench(快乐世界基准(HappyWorld-Bench)) [02:40] 🧠 The Past Frames the Future: Memory for Autoregressive Video Generation(过往帧塑造未来:自回归视频生成中的记忆机制) [03:29] 🧠 Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents(即时记忆:学习为LLM智能体整理任务自适应记忆) [04:11] 🎬 RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling(RewardVerse:面向视频奖励建模的评分量规引导策略优化) [05:00] 🎯 PACT: From Credit Assignment to Critic Alignment(PACT:从信用分配到评论家对齐) [05:43] 🧪 Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?(薛定谔的代码仓库:LLM 是学会了 SWE-bench,还是记住了它?) [06:24] 📐 GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression(GeoPair:用于免训练 Transformer 压缩的几何保持跨层因子分解) [07:00] 📦 PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing(PackLab:在机器人装箱中开发、训练与评估多模态大语言模型的综合框架) [07:53] 🧠 MemBodied: Recurrent Associative Memory for Vision-Language-Action Models(MemBodied:面向视觉-语言-动作模型的循环联想记忆) [08:39] 🧪 WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents(WhatWorkedBench:基准测试AI智能体的实验理解能力) [09:21] 🧠 Hunyuan-A13B Technical Report(混元-A13B 技术报告) [10:03] 🎮 Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms(可验证隐藏动力学游戏:从已求解机制生成智能体强化学习环境) [10:45] 🎥 All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation(所有模态都是平等的,但视频更平等:弥合联合视频生成中的交叉注意力差距) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.09.22 | VLM智能迁移至机器人控制;隐式3D记忆构建视频世界模型【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:28] 🤖 Transferring the Intelligence of VLMs to Robotic Control(将VLM的智能迁移至机器人控制) [01:11] 🎥 WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory(WorldCrafter:具有隐式3D感知记忆的一致视频世界模型) [01:53] 🎮 GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay(GameHorizon 套件:游戏玩法中的多时间跨度数据与评估) [02:35] 🧬 RRSI: Regularized Recursive Self-Improvement of Agent Harnesses(RRSI:智能体运行框架的正则化递归自我改进) [03:20] 📄 Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion(文档检索感知分块(D-RAC):通过 PDF 规范化与多模态 Markdown 转换实现企业文档的通用检索感知摄取) [04:01] 🎓 OmniEdu: Open Foundation Models for Learning and Teaching(OmniEdu:面向学习与教学的开放基础模型) [04:47] 🎬 VideoGen-Agent: Reinforcing Video Generation Agents(VideoGen-Agent:强化视频生成智能体) [05:37] 🐼 onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction(onPanda:通过Token级纠正为LLM与智能体高效标注同策略对齐数据) [06:19] 🤖 Grounded Action Model: 3D Grounding as a Foundation for Robotics(接地动作模型:以3D接地作为机器人学基础) [07:05] 🤖 One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents(从一到多,从多到一:面向软件工程智能体的类别感知迭代专家训练) [07:57] 🎭 Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations(Deep Persona:面向角色扮演智能体与模拟的心理学基础架构与评估框架) [08:36] 🧠 Harness-Zero: Harness Distillation via Agent-as-Harness(Harness-Zero:通过智能体作为外壳进行外壳蒸馏) [09:18] 🤖 CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies(CARE:面向视觉-语言-动作策略的经验引导式原子纠正执行) [10:13] 🧠 Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents(Jev-Mem:面向高效 AI 智能体的 System-One 控制型智能体记忆) [10:52] 🎥 Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms(为什么视频扩散模型会违背物理规律?揭示注意力机制中的缺陷) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.09.21 | 代码合成有据技能;源码扩展编程RL环境【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🧩 Grounded Skill Synthesis from Code at Scale for Agentic Intelligence(面向智能体智能的从大规模代码中合成有依据技能) [01:21] 🤖 CodeMidas: Scaling Agentic Coding RL Environments from Code Itself(CodeMidas:从代码本身扩展智能体编程强化学习环境) [02:03] 🧬 EvoOntology: A Self-Evolving Ontology Layer for Data Agents(EvoOntology:面向数据智能体的自进化本体层) [02:48] 🤖 RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents(RecreationWorld:面向混合计算机使用智能体的可扩展且可验证环境) [03:28] 🧩 IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts(IntBMoE:将块级条件融入专家组合以实现全参与混合专家) [04:14] 🎥 OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue(OmniVChat:面向原生音视频对话的合成、基准测试与训练) [04:59] 🎨 Paint-Anything: Unified Any-Color Control for Image Generation and Editing(Paint-Anything:面向图像生成与编辑的统一任意颜色控制) [05:45] 🎬 OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation(OmniVBench:面向全能参考到视频生成的基准与大规模数据集) [06:25] 🧬 GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills(GraphSkillEvo:图结构智能体技能的进化优化) [07:06] 🤖 MintAct: A Unified Visual Agent for Digital Environments(MintAct:面向数字环境的统一视觉智能体) [07:49] 🎨 Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design(Designer-RSI:从用户流量中演化程序性记忆以支持智能体图形设计) [08:34] 🤖 When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation(当AI评审训练AI评审者:科学判断坍缩与缓解) [09:15] ⚖ Calibrating Teacher--Student Discrepancy for On-Policy Distillation(面向在线策略蒸馏的教师—学生差异校准) [09:57] 🛡 FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection(FRAUDSkill:面向音频反欺诈检测的结构化冻结权重技能优化) [10:41] 📞 TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection(TeleAntiFraud 2.0:一个可刷新、以画像为依据且基于音频的电信诈骗检测基准) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
【周末特辑】9月第3周最火AI论文 | 实时可编辑空间视频;智能体超级智能【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 5 篇论文如下: [00:50] TOP1(🔥683) | 🎬 Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation(Vidu S2:实时交互、可编辑与空间视频生成) [03:11] TOP2(🔥405) | 🤖 Atria Dawn: The Dawn of Agentic Superintelligence(Atria Dawn:智能体超级智能的黎明) [05:04] TOP3(🔥328) | 🧠 Continual Learning Mechanisms Compose for Long-Horizon Memorization(持续学习机制组合用于长时程记忆) [07:23] TOP4(🔥246) | 🤖 ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search(ZGCM-1:一个完全开放且极其高效、面向数学与智能体搜索的基础模型) [09:31] TOP5(🔥226) | 🧠 Dream-RSI: Recursive Self-Improvement through Evolving Worlds(Dream-RSI:通过演化世界实现递归自我改进) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.09.18 | V4.1-Flash压缩KV提效;SoL-Pi优化智能体降本【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🗜 DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression(DeepSeek-V4.1-Flash:将 KV 缓存压缩推向极限) [01:15] ⚡ SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness(SoL-Pi:递归扩展自动化研究循环以实现高效智能体执行框架) [01:55] 🛑 When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation(当 EOS token 不一致:理解在线策略蒸馏中的长度膨胀) [02:33] 🧪 An Empirical Study of Harness Design for Coding Agents(面向编码智能体的 Harness 设计实证研究) [03:20] 🌍 JEPA-Anything: Learning Predictive Models across Different Worlds(JEPA-Anything:跨不同世界学习预测模型) [04:09] 🕵 RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation(RiskChainBench:面向混淆平台消息还原与证据支撑网络调查的基准) [04:56] 🎓 RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning(RetireOPD:面向智能体强化学习的自退场在线策略蒸馏) [05:40] 📄 WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing(WeVisDoc:从覆盖到能力,实现鲁棒的端到端文档解析) [06:29] 🔄 Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents(反思、修订、复用:面向GUI智能体的免训练技能演化) [07:10] 🤖 VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control(VABench:通过视觉演示、主动感知和度量控制测量具身空间智能) [07:56] 🎥 Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation(Video DeltaNet:面向直播视频生成的视频原生混合注意力) [08:42] 🦾 FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations(FAMOS:基于稀疏观测的前馈式三维铰接建模) [09:24] 🖼 UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation(UFO:面向多模态图像生成全条件对齐的评估链) [10:07] 🌍 Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model(MiniMax-H3 能否推理物理世界?一项全模态生成模型评估) [10:54] 🧠 When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models(When2Think:面向高效混合推理模型的难度感知长度控制学习) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.09.17 | 科学代码库转智能体环境;上下文机制网络赋能结构化数据智能【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:31] 🧪 ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments(ScienceIDE:将全球科学代码库转化为智能体可学习环境) [01:23] 🧠 LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence(LimiX-2:面向通用结构化数据智能的上下文机制网络) [02:10] 📉 Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening(重新思考PPO中的评论家学习:理解与缓解价值平坦化) [02:56] 🧠 Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents(置信度源于经验:从推理到智能体的经验性置信度估计) [03:36] 🤖 ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks(ProgramDistill:从交互式 Web 应用到可验证的参考引导软件工程任务) [04:21] 🤖 ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models(ActionPiece:重新思考自回归视觉-语言-动作模型的动作标记化) [05:01] 🧠 Agora: Git as Shared Memory for Collective AutoResearch(Agora:将 Git 作为集体自动研究的共享记忆) [05:43] ⚡ VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention(VC-Attention:面向低比特注意力的值平滑与 Softmax 转换) [06:22] 📈 EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents(EvolveTrade:面向自演化LLM交易智能体的经验驱动策略精炼) [07:02] 🎮 Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control(Zing-0.5:迈向具备实时联合动作与文本控制的可玩世界) [07:51] 👀 Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX(注视作为共同基础证据:MapTask与MUNDEX的跨语料库分析) [08:35] 🎯 A Zeroth-Order Paradigm for LLM Preference Alignment(面向LLM偏好对齐的零阶范式) [09:17] 🧬 HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses(HypoEvolve:遗传算法使多智能体大语言模型能够发现科学假设) [10:02] ✋ EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset(EventEgoHands++:基于事件的第一人称3D手部网格重建与真实数据集) [10:55] 🔭 SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization(SpectralShift:通过谱重参数化有效扩展 Gated DeltaNet 的上下文窗口) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.09.16 | 持续学习组合提升保留率;游戏AI六类角色待整合【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:29] 🧠 Continual Learning Mechanisms Compose for Long-Horizon Memorization(持续学习机制组合用于长时程记忆) [01:17] 🎮 AI for Games in the Foundation Model Era(基础模型时代的游戏人工智能) [02:05] 🎙 StepAudio 3 Realtime Technical Report(StepAudio 3 Realtime 技术报告) [02:45] 🎵 StepAudio 3 Music Technical Report(StepAudio 3 Music 技术报告) [03:22] 🤖 ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents(ScienceBuddy:面向交互式科学智能体的递归嵌套式自我改进) [04:09] 🧭 HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness(HarnessVLN:通过智能体框架统一免训练具身导航) [04:52] 🤖 The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement(人类建造的最后一个AI:迈向真正的递归自我改进) [05:40] 🏗 Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?(墙里的另一张蓝图:如何像孩子一样向前沿AI提问?) [06:26] 🧠 Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States(Mind2Dialogue:通过模拟用户心理状态训练人类感知语言模型) [07:10] 🧩 ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement(ModularRSI:模块化且可泛化的递归式执行框架自我改进) [07:54] 🤖 Modality-Autoregressive World-Action Models(模态自回归世界动作模型) [08:38] 📐 Disentangling Representation Evolution in Transformers through Directional Decomposition(通过方向分解解耦 Transformer 中的表征演化) [09:25] 🎥 PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control(PhysStream:具备结构化场景记忆与细粒度运动控制的流式物理基础视频生成) [10:06] 🧪 ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals(ImpossibleRubrics:对作为奖励信号生成的评分标准进行压力测试) [10:55] 🧠 Convergent Emergence of In-Context Learning Across Modalities(跨模态上下文学习的趋同涌现) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.09.15 | 智能体超级智能黎明;实时可编辑空间视频生成【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🤖 Atria Dawn: The Dawn of Agentic Superintelligence(Atria Dawn:智能体超级智能的黎明) [01:13] 🎬 Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation(Vidu S2:实时交互、可编辑与空间视频生成) [01:58] 🤖 ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search(ZGCM-1:一个完全开放且极其高效、面向数学与智能体搜索的基础模型) [02:47] 🧠 Dream-RSI: Recursive Self-Improvement through Evolving Worlds(Dream-RSI:通过演化世界实现递归自我改进) [03:24] 🤖 PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models(PhysBrain 1.5:从视觉语言模型到物理基础模型) [04:06] 🗜 Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction(分组值注意力:通过按需键重建实现高效KV缓存) [04:51] 🎬 LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows(LynnReal-Omni:面向智能体视觉工作流的原生多模态视频生成) [05:45] 🤖 RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments(RSIAgent:新环境中面向递归自我改进的自主探索) [06:30] 🔬 Discovery Foundation Models: Toward Open-Ended Discovery Intelligence(发现基础模型:迈向开放式发现智能) [07:13] 🎬 BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender(BVB:在 Blender 中通过程序化重建对智能体视频理解进行基准测试) [07:56] ⚖ How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus(无损投机解码究竟有多无损?数值精度在 Orthrus 中的作用) [08:39] 🌐 AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video(AlayaVista:从全景状态到透视视频的流式世界建模) [09:26] 🧩 Kaininja: Extending Native 3D Generators to the Part Level(Kaininja:将原生3D生成器扩展至部件级) [10:13] 🧠 Omni-Streaming Thinking(全模态流式思维) [11:00] 🖥 LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents(LLaDA-UI:将块级扩散引入视觉语言 GUI 智能体) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递