2026.07.29 | 高保真数据训练策略,逼近真机效果;相关性动态引导搜索,精准高效检索

2026.07.29 | 高保真数据训练策略,逼近真机效果;相关性动态引导搜索,精准高效检索

15分钟 ·
播放数70
·
评论数0

【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 www.xiaoyuzhoufm.com

【目录】
本期的 15 篇论文如下:

[00:31] 🤖 HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone(HiFi-UMI:仅从高保真UMI数据学习可部署的操作策略)
[01:31] 🔍 A New Role for Relevance: Guiding Corpus Interaction in Agentic Search(相关性的新角色:在智能体搜索中引导语料库交互)
[02:20] 🎨 ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition(ReDesign:通过智能体分解从图像中恢复可编辑的设计结构)
[03:07] 🧠 Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory(保持铭记:基准测试代理记忆中的内隐关联盲点)
[03:56] 🏃 Pass the Baton: Trajectory-Relayed On-Policy Distillation(传递接力棒:轨迹中继的在线策略蒸馏)
[04:47] ⚡ Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model(Mage-VL:一种高效的编解码器原生流式多模态基础模型)
[05:49] 🧩 CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents(CodeNib:一种为编码智能体提供仓库上下文服务的多视图数据系统)
[06:48] 🌍 Wonder: Video World Model Done Better(Wonder:更优的视频世界模型)
[07:51] 👁 PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models(感知基准:评估多模态大语言模型中的原子视觉感知能力)
[08:44] 🔍 Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking(新颖主张还是似曾相识?重新思考多模态自动事实核查的“无污染”动态评估)
[09:40] 🛡 Shieldstral(盾星)
[10:41] 🔀 MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities(MODUS:仅解码器的任意模态到任意模态多样化建模)
[11:41] ⚡ Parallel Decoding Distillation for Fast Image and Video Generation(并行解码蒸馏:面向快速图像与视频生成的方法)
[12:23] 🎬 Visual prompt engineering for video models(视频模型的视觉提示工程)
[13:16] 🎯 OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs(OmniDelta:面向全模态大语言模型令牌压缩的技能驱动预算分配方法)

【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递