【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 www.xiaoyuzhoufm.com
【目录】
本期的 15 篇论文如下:
[] 🤖 Show-Harness: Just a VLM Agent Can Play Robots(Show-Harness:仅一个VLM智能体即可玩转机器人)
[] 🎮 Programmable World Model(可编程世界模型)
[] 🤖 AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems(AgentGrad:干预引导的多智能体系统提示优化)
[] ⌚ WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data(WearableQA:面向真实世界可穿戴数据的健康推理基准)
[] ✅ SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents(SWE-Bench Pro Verified:面向软件工程智能体的可靠基准)
[] 🔬 SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?(SAEScientist-Bench:AI智能体能否开展自主SAE可解释性研究?)
[] 🔍 Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents(仅凭分数不能证明发现:用于审计 AI 研究代理的发现认证协议)
[] 🗣 Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation(Puppeteer:基于物体锚定与姿态感知的共语手势生成)
[] ⚗ DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents(DianShi-RxnDB:面向研究人员与AI智能体的、通过全自动流程构建的大规模细粒度有机反应数据平台)
[] 🤖 SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators(SyncWorld:视觉校准使世界模型成为零样本模拟器)
[] 🏗 $Φ$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?(Φ-Bench:大型语言模型能否工程化支撑自身的基础设施?)
[] 🧠 Revisiting Complete Reasoning Traces for Post-Training(重新审视后训练中的完整推理轨迹)
[] 🔄 Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning(训练更智能,而非更费力:主动学习中的切换信号引导训练)
[] 🎥 Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs(为什么视频仍然如此昂贵?视频与音视频大语言模型中的推理效率机制综述)
[] 🤖 Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails(协同演化智能体运行框架与模型:同策略修正帮助较弱模型在模仿失败之处追赶)

【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
