

2026.08.14 | 动作条件视频世界模型引入几何感知;长时记忆外部化实现无尽世界【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🤖 DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation(DreamX-Phi 1.0:面向机器人操作的动作条件视频世界模型) [01:24] 🌍 Alaya-EVOKE: From Linear-Scaling Supervision to Endless World(Alaya-EVOKE:从线性扩展监督到无尽世界) [02:30] 🔀 LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers(LLMRouter:开发、评估和部署LLM路由器的统一基础设施) [03:25] 🧬 DarwinX: Evolving Agent Harnesses Through Natural Selection(DarwinX:通过自然选择进化智能体框架) [04:26] 🔬 Intern-S2-Preview: Scientific Agentic Foundation Model(Intern-S2-Preview:科学智能体基础模型) [05:20] 🎮 PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives(PlayWorld:使用智能体玩家在长程目标上对世界模型进行基准测试) [06:20] 🤖 AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design(AutoDesign:面向长时程智能体设计的元框架优化) [07:23] 🧠 Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence(空间记忆智能体:基于经验的程序记忆实现空间智能) [08:12] ⚡ Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus(混合线性注意力大语言模型中的大规模激活:注意力前尖峰与尖峰间平台) [09:03] 🎭 UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos(UniSwap:面向说话视频的流式音频-视觉身份交换) [10:12] ⚡ LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time(LiveAnimate:实时稳定长格式流式人体动画生成) [11:08] 🤖 How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review(修辞何以能对AI审稿人进行奖励黑客?解析基于AI的同行评审中的修辞敏感性) [12:00] ✂ An AI4AI Framework for Visual Token Pruning(面向视觉Token剪枝的AI4AI框架) [12:56] 🤖 H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models(H2R-Bench:在世界模型中评估人类到机器人的操作视频生成) [14:08] 🔄 Full-bandwidth transformer(全带宽Transformer) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.13 | 演化环境揭示智能体风险;组合技能实现论文生成【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:26] 🧬 OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution(OpenART:通过开放式环境演化扩展智能体红队测试) [01:35] 📄 Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill(从火花到论文:作为可组合技能的端到端研究论文生成) [02:43] 🧩 AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses(测试时AI4AI:通过推理支架实现强到弱能力迁移) [03:37] 🔬 Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence(Mechanist:将人工智能作为揭示智能机制的科学仪器) [04:33] 🎭 Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives(大语言模型智能体能否坚守剧本?交互式叙事中长程一致性的基准测试) [05:33] 🌍 StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization(StateFlow:为预可视化构建、演化与访问3D世界状态) [06:28] 📐 Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models(自几何:面向几何一致的3D视觉基础模型的无真值即插即用测试时自适应) [07:25] 🪞 From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection(从合成到去除:物理驱动的反射模拟与基于扩散模型的视频去反射) [08:22] 🛡 ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents(ToolHazard:扩展对抗性环境,用于基于大语言模型的智能体的安全评估与对齐) [09:15] 🔍 The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images(视觉工具使用的幻象:图像思维的因果审计) [10:16] 🤖 Self-Evolving Embodied Agents via Skill-Harness Evolution(基于技能与执行框架演化的自进化具身智能体) [11:19] 🧩 SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries(SkillZip:面向可扩展智能体技能库的合约保持图压缩) [12:14] 🛡 Agent Safety Should Be a Runtime Contract(智能体安全应当是一种运行时契约) [13:01] 🧬 Persistent Recursive Worlds Enable Autonomous Software Evolution(持久递归世界赋能自主软件演化) [13:46] 💡 MBA: Multimodal Benchmark and Agents for Real-World Business Ideation(MBA:面向真实世界商业构思的多模态基准与智能体) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.12 | 共体智能体以人为中心助人成长;智能体与环境共演化迈向自我导向【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:35] 🤖 ComBodied Agents: a New Paradigm of Human-Centric Agentic AI(共体智能体:以人为中心的智能体人工智能新范式) [01:22] 🧬 Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design(智能体系统中的共同演化:迈向超越人类设计的自我导向演化) [02:22] 🌍 Beyond Pixels: From Video Priors to 4D Worlds(超越像素:从视频先验到4D世界) [03:12] 🧩 Articulated Object Reconstruction from Rest-State Observation(基于静止状态观测的铰接物体重建) [04:09] ⚔ AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss(AdvFD:通过对抗性弗雷歇距离损失提升视觉生成) [05:05] 🧬 Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution(孟德尔·哥德尔机:通过比较进化实现递归自我改进的编码智能体) [06:00] 🎭 Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence(Ex-Omni-2D:具备原生视觉临场感的表现性全模态对话模型) [06:48] 🌍 VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?(VibeLifeBench:你的生活智能体能否在动态世界中主动且持久地行动?) [07:46] 🚫 Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness(解码级禁忌:大语言模型鲁棒性的诊断性压力测试) [08:33] 📦 SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure(SkillZip:通过发现可复用结构实现自我进化智能体的免评估技能压缩) [09:33] 📱 SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information(SPIEval:评估大型语言模型作为移动助手处理分散个人信息的能力) [10:36] ✂ Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents(并不值得再投入一个词元:高效深度研究智能体的边际价值估计) [11:34] 🌐 Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation(开放大语言模型用于多语言机器翻译的无参考后训练) [12:28] 🔍 InSight-doc: Agentic Visual Perception for Long-Document Understanding(InSight-doc:面向长文档理解的智能体视觉感知) [13:27] 🔀 UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models(UniMoMo:基于专家合并的大型推荐模型MoE加速方法) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.11 | 自进化混合专家赋能持续学习;代码重构基准揭示智能体局限【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🔄 Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA(Macaron-V1:迈向具备自我改进和LoRA混合的开放持续学习) [01:24] 🔧 SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring(SWE-Bench ProMax:面向大规模多语言代码重构的智能体基准评测) [02:21] 🐍 Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution(Ouroboros:通过核心评审进化实现自我发展的前沿编程智能体) [03:31] 🧠 BDH-CQ: In-Context Learning with Recurrent Latent Reasoning(BDH-CQ:基于循环潜在推理的上下文学习) [04:13] 🧠 Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory(智能体记忆蒸馏:利用分层教师记忆赋能小型大语言模型智能体) [05:11] 🧠 Motif 3: Technical Report(Motif 3:技术报告) [05:59] 🔬 Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains(Sci-VBench:评估科学领域中知识与推理密集型视频生成) [06:49] 🖼 What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems(下一步编辑什么:对话系统中的视觉对齐图像编辑后续建议) [07:53] 🎯 SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation(SPOT:面向同策略蒸馏的稀疏探测与结果校准) [08:58] ⚡ OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching(OasisKV:通过前瞻稀疏预取将解码期KV缓存扩展到HBM之外) [09:49] 🧠 RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States(RoMeRL:通过降阶效用状态平衡自进化智能体记忆中的反馈覆盖与记忆-奖励陷阱) [10:43] 🔍 Evidence-RL: Towards Evidence-intensive Visual Reasoning(证据强化学习:迈向证据密集型视觉推理) [11:42] 🧠 Scaling Inherently Interpretable Language Models(扩展内在可解释的语言模型) [12:40] 🧬 Evo-Bench: Can Language Models Improve Agent Harness?(Evo-Bench:语言模型能否改进智能体运行框架?) [13:40] 🔓 Stealing Reasoning Traces from Proprietary LLM APIs(从专有大语言模型API中窃取推理轨迹) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.10 | 多模态智能体环境设计重质轻量;强化学习利于多任务共存【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🌍 Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning(超越单纯的环境规模扩展:为多模态智能体学习设计有效的环境分布) [01:36] ⚔ SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs(SFT冲突、RL共存:大语言模型多任务学习的理论与实证分析) [02:42] 🚗 SimWAM: A Simple World Action Model for End-to-End Autonomous Driving(SimWAM:面向端到端自动驾驶的简单世界动作模型) [03:40] 🎯 YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family(YOLO-PEFT:YOLO系列上的参数高效微调) [04:43] 🎥 StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding(StreamArena:迈向连续、交互与长时程的智能体流式视频理解) [05:36] 🎧 Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning(以演化评分标准作为奖励的音频推理强化学习) [06:19] 🙈 When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles(当激活预言机学会不读取:微调预言机中的概念特异性盲区) [07:09] ⚡ Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss(面向大语言模型的高效知识蒸馏:离线Top-K Logits与融合分块KL损失) [08:16] 🚁 Uncertainty-Aware World Model for Aerial Image-Goal Navigation(面向航拍图像目标导航的不确定性感知世界模型) [09:19] 🎯 Douyin Multimodal Embedding Model Technical Report(抖音多模态嵌入模型技术报告) [10:19] 📈 Skaling: Chinchilla's Exponents Meet Kaplan's Coupling(Skaling:Chinchilla指数与Kaplan耦合的融合) [11:11] 🧠 Addressable Memory for Video World Models(视频世界模型的可寻址记忆) [12:05] 🧩 Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression(相关但不完整:硬提示压缩中作为范式级失败模式的指称悬空) [13:01] 🔄 Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors(往返一致性:双向扩散模型可预测自身的展开误差) [14:01] 🧠 The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows(优化器即智能体:跨提示、程序与机器学习工作流的推理驱动搜索) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
【周末特辑】8月第2周最火AI论文 | 递归合成扩数据;状态管理提长程【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 5 篇论文如下: [00:41] TOP1(🔥221) | 🔁 Recursive Synthesis for Long-Horizon Terminal Tasks(面向长时程终端任务的递归式合成) [03:39] TOP2(🔥162) | 🧭 LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks(LongHorizon-Harness:推动面向真实世界任务的长时程智能体) [06:29] TOP3(🔥154) | 🎙 SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks(SwanTale:面向指令与零样本任务的统一多说话人语音与音频生成) [09:28] TOP4(🔥140) | 🚗 Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs(自动驾驶视觉语言模型中用于可验证推理的未来轨迹延迟暴露) [12:30] TOP5(🔥108) | ⚓ DAPD: Dual-Anchored Policy Distillation(双锚定策略蒸馏) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.07 | 递归自蒸馏重塑智能体信用;开源裁判低成本评估操作【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🎯 AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning(AgentOPSD:面向智能体强化学习的递归自蒸馏) [01:35] 🤖 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models(OSReward:为跨平台计算机使用奖励模型制定标准化评估) [02:24] 🌍 WorldClaw: Agentic 3D Open-World Generation at Scale(WorldClaw:大规模智能体式3D开放世界生成) [03:12] 🗺 GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?(GST-Bench:视觉语言模型能否从视频中形成全局空间意识?) [04:17] 💭 EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning(EnvACE:通过世界预演将环境动态内化于智能体强化学习) [05:10] 🔍 Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval(从失败中学习:基于硬负样本的检索中心思维链用于统一多模态检索) [06:08] ⏳ ChronoVision: Temporal Reasoning via Latent State Reconstruction(ChronoVision:通过潜在状态重建实现时序推理) [07:08] 🌐 From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models(从经济主体到主体经济:经济世界模型的系统蓝图) [08:10] 🧮 On-Policy Delta Distillation for Multilingual Math Reasoning(面向多语言数学推理的同策略差值蒸馏) [08:59] ⚙ HarnessOpt-Bench: Evaluating LLMs at Harness Optimization(HarnessOpt-Bench:评估大语言模型在智能体运行框架优化上的表现) [09:43] 🇬 Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains(教Nemotron希腊语:面向专业领域的现代希腊语语料挖掘、检索适配与有据生成) [10:43] 🤖 DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation(DyPES-VLA:学习共享动力学先验与具身特定控制以实现跨具身操作) [11:38] 🤖 World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation(世界到手腕:面向精细机器人操作的任务条件化未来手腕建模) [12:44] 📊 DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces(DataSpace:面向异构工作空间的可验证分析数据代理基准测试) [13:44] 🪄 EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal(EffectLearner:面向真实世界视频目标移除的世界感知对象-效应推理) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.06 | 回溯答案训练搜索智能体;统一工具实现图像生成【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🔍 ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment(ABSeeker:通过答案回溯的信用分配训练长程搜索智能体) [01:25] 🎨 ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation(ToolArtist:面向智能体图像生成的工具使用统一多模态模型) [02:16] 🪞 The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads(个性化幻象:大语言模型如何捏造用户画像,以及自我监控为何具有误导性) [03:17] ⚛ Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes(迈向多模态预训练的物理机制:知识流动、模态协同、早期统一与配方) [04:20] 🤖 OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents(OneDayAgent:面向自主智能体的长周期任务执行框架) [05:14] 🧬 GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks(GDPevo:在真实业务任务中评估智能体的自我进化) [06:05] 🎯 When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation(当教师误导:虚假信号感知的同策略蒸馏) [06:56] 🧩 Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning(迈向技能原生的大语言模型:用于长程推理评测与训练的技能熵) [08:02] 🧩 NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap(NOLLI:用于诊断英韩性能差距的难度校准谜题基准) [08:57] 🤖 Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data(Ego2Robot:从第一人称人类数据中可扩展合成机器人数据) [09:54] 🎬 AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities(AVE-Compass:面向音视频编辑能力的全面评估) [10:55] 🧠 When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents(当记忆说谎:VLM智能体中空间记忆陈旧性的实证研究) [11:53] 🎯 Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance(在失败处蒸馏:利用自适应教师指导恢复负强化学习组的学习信号) [12:49] 🧠 FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory(FocusMem:分解潜在GUI记忆中的内容、读取与信任) [13:52] 👋 HelloWorld: Enabling Socially Interactive Characters in Video World Models(HelloWorld:在视频世界模型中实现社交互动角色) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.05 | 智能体长期运营后净资产不足人类三成;实时视频编辑达高清流畅【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:27] 🛒 MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations(MerchantBench:电商运营中LLM智能体长期一致性的基准评测) [01:19] 🎬 JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion(JoyAI-Video-Edit:利用自回归扩散的实时开放式视频编辑) [02:19] 🧊 Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing(Hunyuan3D-Buffalo 1.0:一种可扩展的三维生成、理解与编辑统一多模态模型) [03:15] 🌌 AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling(AURORA-LM:自编码统一表示用于连续潜变量扩散语言建模) [04:18] 🎥 Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent(视频深度研究:迈向下一代多模态深度研究智能体) [05:15] 🔄 Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation(知识-几何解耦:面向流式推荐的可刷新预训练迁移) [06:20] 🤖 PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning(PCSD:智能体强化学习中自蒸馏的持久一致性) [07:08] 🌍 Quo Vadis, World Modeling?(世界建模,何去何从?) [07:59] 🔁 PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents(PAST-Bench:对个人智能体中递归自我改进的基础进行基准测试) [08:56] 🌉 Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging(Any-OPD:通过表示空间桥接实现异构流匹配模型的在策略蒸馏) [09:57] 🗜 OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models(OmniPack:用于高效全模态大语言模型的统一令牌压缩) [11:02] 📈 LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models(LLaDA MoE v2:扩展混合专家扩散语言模型) [12:00] 🎯 CAPEval: A Decoupled Caption Evaluation across Understanding and Generation(CAPEval:面向理解与生成的解耦式图像描述评估) [12:55] 🧩 UniWorld-Design: From Pixel Generation to Layer-Native Design(UniWorld-Design:从像素生成到图层原生设计) [13:56] 🧠 SkillJack: Persistent Skill Backdoors in Self-Evolving Agents(SkillJack:自进化智能体中的持久性技能后门) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.08.04 | 长时程智能体成功率提升至八成;多说话人语音音频统一生成【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🧭 LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks(LongHorizon-Harness:推动面向真实世界任务的长时程智能体) [01:30] 🎙 SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks(SwanTale:面向指令与零样本任务的统一多说话人语音与音频生成) [02:31] 🎯 VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation(VAD:在多模态在策略蒸馏中为目标重建归因视觉证据) [03:37] 🤖 Progressive Agent Skill Generation via Reinforcement Learning(基于强化学习的渐进式智能体技能生成) [04:39] ⚓ DAPD: Dual-Anchored Policy Distillation(双锚定策略蒸馏) [05:37] 🧲 UEmbed: Unified Sparse and Dense Multimodal Embeddings(UEmbed:统一稀疏与稠密多模态嵌入) [06:44] 🌍 WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity(WorldExam:从表象外观到内在反应性的世界模型基准评测) [07:54] 🔗 CADENA: Stepwise CAD Reverse Engineering(CADENA:逐步式CAD逆向工程) [08:55] 🛠 SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation(SKT:通过经验证的合成数据生成实现规模化技能使用训练) [09:58] 🤖 SWE-Touch: Benchmarking Coding Agents When Users Touch the Code(SWE-Touch:在用户改动代码时对编码代理的基准测试) [11:08] 🚗 Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs(自动驾驶视觉语言模型中用于可验证推理的未来轨迹延迟暴露) [12:09] 🧠 WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning(WCM:面向视觉-语言-动作强化学习的世界评论家模型) [13:05] 🔄 Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations(超越形态的运动:从抽象运动表征引导跨类别运动迁移) [14:04] 🧠 GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning(GradCuit:信用分配的梯度流实现稳健且可解释的测试时潜在推理) [14:59] 🛋 Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis(Roomer:面向三维室内布局合成的反思式对象级模型编辑与修复) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
【月末特辑】7月最火AI论文 | 虎鲸模型预测世界状态;Kimi K3开源逼近顶尖【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 10 篇论文如下: [00:42] TOP1(🔥473) | 🌍 Orca: The World is in Your Mind(虎鲸:世界在你心中) [03:50] TOP2(🔥438) | 🧠 Kimi K3: Open Frontier Intelligence(Kimi K3:开放前沿智能) [06:31] TOP3(🔥309) | 🧩 Program-as-Weights: A Programming Paradigm for Fuzzy Functions(程序即权重:面向模糊函数的编程范式) [09:09] TOP4(🔥307) | 🌍 ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU(ABot-World-0:在单个桌面GPU上实现无限交互式世界展开) [12:34] TOP5(🔥293) | 🧪 AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis(AskChem:以论断为中心的化学文献综合基础设施) [15:21] TOP6(🔥289) | 🤖 Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents(Qwen-UI-Agent技术报告:迈向下一代以真实世界为中心的基础GUI智能体) [18:00] TOP7(🔥259) | 🧠 Metis: Memory Foundation Model(Metis:记忆基础模型) [22:07] TOP8(🔥232) | 🧭 Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable(驾驭手册:使不断演化的智能体驾驭系统可读、可导航且可编辑) [25:09] TOP9(🔥205) | 🧠 LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget(长稻草:在固定GPU预算下实现超过200万Token的长上下文强化学习) [28:24] TOP10(🔥198) | 🤖 RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model(RynnBrain 1.1:迈向更强大和更通用的具身基础模型) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
【周末特辑】8月第1周最火AI论文 | Kimi K3开源前沿智能;AskChem论断级化学检索【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 5 篇论文如下: [00:44] TOP1(🔥431) | 🧠 Kimi K3: Open Frontier Intelligence(Kimi K3:开放前沿智能) [03:17] TOP2(🔥292) | 🧪 AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis(AskChem:以论断为中心的化学文献综合基础设施) [05:53] TOP3(🔥281) | 🤖 Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents(Qwen-UI-Agent技术报告:迈向下一代以真实世界为中心的基础GUI智能体) [08:46] TOP4(🔥256) | 🧠 Metis: Memory Foundation Model(Metis:记忆基础模型) [11:40] TOP5(🔥192) | 🤖 Progress Reward Modeling for Robotic Learning: A Comprehensive Survey(机器人学习的进度奖励建模:一项全面综述) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.07.30 | TurboVLA实现消费级显卡实时操控;CoRT精细化信用分配提升指令遵循。【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:33] ⚡ TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM(TurboVLA:在RTX 4090上以32Hz频率运行且显存占用低于1GB的实时视觉-语言-动作模型) [01:36] 🎯 CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization(CoRT:用于令牌级准则引导策略优化的反事实重放) [02:31] 🤖 HumanCLAW: Can Vision-Language Models Act Through a Body?(HumanCLAW:视觉-语言模型能否通过身体行动?) [03:41] 🧬 DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space(DecoEvo:文本空间中求解器与评估生成器技能的解耦协同进化) [04:39] 🧠 CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition(CLBench-V:从基础定位到知识获取的多模态上下文学习评估) [05:32] 🧩 CAST: Game Solvers as Turn-Level Teachers for LLM Agents(CAST:游戏求解器作为LLM智能体的回合级教师) [06:28] 🧠 SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution(技能崛起:面向跨任务技能演化的智能体强化学习) [07:18] 🎮 StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation(StatePlay:状态感知的游戏世界模型用于机制一致的内容生成) [08:05] 📊 OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding(OmegaUse-OfficeVal:基于经济基准评估LLM智能体在长期办公套件任务中的表现) [09:08] 🤖 Can AI agents conduct open-ended AI research? Early evidence from two case studies(AI智能体能否进行开放式的AI研究?来自两个案例研究的早期证据) [10:04] 🛡 GPT-Red: Automated Red Teaming via Self-Play at Scale(GPT-Red:通过大规模自我对弈实现自动化红队测试) [11:04] 🎬 Explicit Layer Modeling for Video Object Insertion and Layer Decomposition(显式层建模用于视频对象插入与层分解) [12:01] 🕵 StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents(StealthBench:衡量自主攻击安全代理的操作隐蔽性) [12:54] 🧠 Memory for Large Language Models(大型语言模型的记忆机制) [13:48] 📜 Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems(为叙述者评级:面向多智能体知识系统中声明级溯源的一种伊斯纳德-里贾尔框架) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.07.29 | 高保真数据训练策略,逼近真机效果;相关性动态引导搜索,精准高效检索【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:31] 🤖 HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone(HiFi-UMI:仅从高保真UMI数据学习可部署的操作策略) [01:31] 🔍 A New Role for Relevance: Guiding Corpus Interaction in Agentic Search(相关性的新角色:在智能体搜索中引导语料库交互) [02:20] 🎨 ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition(ReDesign:通过智能体分解从图像中恢复可编辑的设计结构) [03:07] 🧠 Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory(保持铭记:基准测试代理记忆中的内隐关联盲点) [03:56] 🏃 Pass the Baton: Trajectory-Relayed On-Policy Distillation(传递接力棒:轨迹中继的在线策略蒸馏) [04:47] ⚡ Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model(Mage-VL:一种高效的编解码器原生流式多模态基础模型) [05:49] 🧩 CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents(CodeNib:一种为编码智能体提供仓库上下文服务的多视图数据系统) [06:48] 🌍 Wonder: Video World Model Done Better(Wonder:更优的视频世界模型) [07:51] 👁 PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models(感知基准:评估多模态大语言模型中的原子视觉感知能力) [08:44] 🔍 Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking(新颖主张还是似曾相识?重新思考多模态自动事实核查的“无污染”动态评估) [09:40] 🛡 Shieldstral(盾星) [10:41] 🔀 MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities(MODUS:仅解码器的任意模态到任意模态多样化建模) [11:41] ⚡ Parallel Decoding Distillation for Fast Image and Video Generation(并行解码蒸馏:面向快速图像与视频生成的方法) [12:23] 🎬 Visual prompt engineering for video models(视频模型的视觉提示工程) [13:16] 🎯 OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs(OmniDelta:面向全模态大语言模型令牌压缩的技能驱动预算分配方法) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递
2026.07.28 | Kimi K3开源模型性能领先;JarvisHub画布框架革新创意协作【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:33] 🧠 Kimi K3: Open Frontier Intelligence(Kimi K3:开放前沿智能) [01:22] 🎨 JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents(JarvisHub:一个面向画布原生多模态创意代理的开放框架) [02:20] 🤖 From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search(从专有到开源:通过多智能体协议蒸馏弥合智能搜索中的分布差距) [03:18] 🤖 Progress Reward Modeling for Robotic Learning: A Comprehensive Survey(机器人学习的进度奖励建模:一项全面综述) [04:17] 🤖 StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents(StateAct:面向长周期计算机使用代理,程序状态优先于像素) [05:23] 🧠 Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation(重新思考在线策略扩散蒸馏中的无分类器引导) [06:27] 🗼 Data Pyramid for Embodied Manipulation(具身操作的数据金字塔) [07:33] ⚡ Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification(Sol-Attn:通过即时注意力稀疏化加速视频生成推理) [08:37] 🎬 OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation(OmniVAE:一种具有跨模态对齐的音频-视频VAE,用于联合生成) [09:30] 🧠 The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation(多轮长程规划中的物理学:通过单教师与多教师在线策略智能体蒸馏从预训练到后训练) [10:26] 👗 Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On(氧气试穿:面向时尚的原生基础模型,实现任意物品虚拟试穿) [11:17] 🦎 Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling(Chamaileon:基于情境化建模与混合采样的跨情境结合剂设计) [12:11] 🔮 dRAE: Representation Autoencoder with Hyper-Spherical Codes(dRAE:超球面码的表示自编码器) [13:09] 🧩 DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes(解耦混合:用于可扩展VLM数据配方的解耦比率搜索与凸分配) [14:08] 🏥 ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding(ClinFusion:一种面向整体医学理解的以视觉为中心的多模态大语言模型系统) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递