2026.08.13 | 演化环境揭示智能体风险;组合技能实现论文生成

2026.08.13 | 演化环境揭示智能体风险;组合技能实现论文生成

15分钟 ·
播放数34
·
评论数0

【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 www.xiaoyuzhoufm.com

【目录】
本期的 15 篇论文如下:

[00:26] 🧬 OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution(OpenART:通过开放式环境演化扩展智能体红队测试)
[01:35] 📄 Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill(从火花到论文:作为可组合技能的端到端研究论文生成)
[02:43] 🧩 AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses(测试时AI4AI:通过推理支架实现强到弱能力迁移)
[03:37] 🔬 Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence(Mechanist:将人工智能作为揭示智能机制的科学仪器)
[04:33] 🎭 Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives(大语言模型智能体能否坚守剧本?交互式叙事中长程一致性的基准测试)
[05:33] 🌍 StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization(StateFlow:为预可视化构建、演化与访问3D世界状态)
[06:28] 📐 Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models(自几何:面向几何一致的3D视觉基础模型的无真值即插即用测试时自适应)
[07:25] 🪞 From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection(从合成到去除:物理驱动的反射模拟与基于扩散模型的视频去反射)
[08:22] 🛡 ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents(ToolHazard:扩展对抗性环境,用于基于大语言模型的智能体的安全评估与对齐)
[09:15] 🔍 The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images(视觉工具使用的幻象:图像思维的因果审计)
[10:16] 🤖 Self-Evolving Embodied Agents via Skill-Harness Evolution(基于技能与执行框架演化的自进化具身智能体)
[11:19] 🧩 SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries(SkillZip:面向可扩展智能体技能库的合约保持图压缩)
[12:14] 🛡 Agent Safety Should Be a Runtime Contract(智能体安全应当是一种运行时契约)
[13:01] 🧬 Persistent Recursive Worlds Enable Autonomous Software Evolution(持久递归世界赋能自主软件演化)
[13:46] 💡 MBA: Multimodal Benchmark and Agents for Real-World Business Ideation(MBA:面向真实世界商业构思的多模态基准与智能体)

【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递