2026.08.31 | LoopArena揭示循环控制难题;DART-SD拓扑感知优化工具调用。

2026.08.31 | LoopArena揭示循环控制难题;DART-SD拓扑感知优化工具调用。

15分钟 ·
播放数51
·
评论数0

【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 www.xiaoyuzhoufm.com

【目录】
本期的 15 篇论文如下:

[00:29] 🔁 LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering(LoopArena:将模型作为循环工程的运行时控制器进行基准测试)
[01:27] 💎 DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents(DART-SD:多轮工具调用代理的菱形拓扑感知检索与自蒸馏调优)
[02:25] 🛠 Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities(智能体工件创建:系统、评估、原则与机遇)
[03:25] 🤖 Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models(超越数据规模:面向视觉-语言-动作模型的以表征为中心的持续预训练)
[04:42] 🌍 Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning(代码即世界:面向物理推理的可执行世界表示的智能体发现)
[05:45] 🔄 J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data(J-Zero:零数据下挑战者—求解者—评判者的统一协同进化)
[06:37] 🎥 Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction(重新审视面向长时程流式三维重建的局部上下文)
[07:35] 🧠 ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL(ContextPilot:通过细粒度强化学习训练智能体进行主动上下文管理)
[08:33] 🧠 LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation(LayerRecall:用于视频生成长时程一致性的状态条件记忆路由器)
[09:38] 💰 Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090(Puro-2B:穷实验室在RTX 5090上以不到5090美元训练出的Qwen2-1.5B)
[10:27] 🐘 Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge(盲人摸象:探究长尾分歧知识下大语言模型的认识论短视)
[11:18] 🛡 StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing(StepGuard:通过可扩展监督与安全效用平衡学习步骤级护栏)
[12:12] 🎨 Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents(画你所见:多模态智能体中灵巧视觉工具使用的基准测试)
[13:08] ⚡ Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding(在视频中定位一切:重新思考高效生成式时空视频定位)
[14:04] 🧠 PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control(PonderPounce:预训练多模态大语言模型作为机器人控制的回合上下文引擎)

【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递