【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 www.xiaoyuzhoufm.com
【目录】
本期的 15 篇论文如下:
[] 🔍 ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment(ABSeeker:通过答案回溯的信用分配训练长程搜索智能体)
[] 🎨 ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation(ToolArtist:面向智能体图像生成的工具使用统一多模态模型)
[] 🪞 The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads(个性化幻象:大语言模型如何捏造用户画像,以及自我监控为何具有误导性)
[] ⚛ Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes(迈向多模态预训练的物理机制:知识流动、模态协同、早期统一与配方)
[] 🤖 OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents(OneDayAgent:面向自主智能体的长周期任务执行框架)
[] 🧬 GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks(GDPevo:在真实业务任务中评估智能体的自我进化)
[] 🎯 When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation(当教师误导:虚假信号感知的同策略蒸馏)
[] 🧩 Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning(迈向技能原生的大语言模型:用于长程推理评测与训练的技能熵)
[] 🧩 NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap(NOLLI:用于诊断英韩性能差距的难度校准谜题基准)
[] 🤖 Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data(Ego2Robot:从第一人称人类数据中可扩展合成机器人数据)
[] 🎬 AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities(AVE-Compass:面向音视频编辑能力的全面评估)
[] 🧠 When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents(当记忆说谎:VLM智能体中空间记忆陈旧性的实证研究)
[] 🎯 Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance(在失败处蒸馏:利用自适应教师指导恢复负强化学习组的学习信号)
[] 🧠 FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory(FocusMem:分解潜在GUI记忆中的内容、读取与信任)
[] 👋 HelloWorld: Enabling Socially Interactive Characters in Video World Models(HelloWorld:在视频世界模型中实现社交互动角色)

【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
