OneDayAgent 测试长周期自主智能体工作流
这一评测框架在 104 项基准任务中管理任务拆解、记忆与验证。
为什么重要
这项工作瞄准了智能体系统中持续存在的失效模式,包括目标漂移、状态丢失和上下文溢出。其在来自三个模型家族的五个后端 LLM 上报告的表现表明,框架设计或许能在不依赖单一模型的情况下提升长周期可靠性。
核心要点
- 1.OneDayAgent 管理任务拆解、记忆和最终验证。
- 2.它在使用 GLM-5.2 的 104 项 AgentIF-OneDay 任务上取得 0.821 分。
- 3.该框架运行于来自三个模型家族的五个 LLM 后端。
研究人员推出了 OneDayAgent,这是一个面向自主智能体的框架,用于处理工作、学习和生活中开放式的日常请求。该系统会将请求拆解为边界清晰的子任务,在上下文压力下维护执行记忆,并对最终交付物进行验证和修复。在 AgentIF-OneDay 上,它接受了 104 项任务评估,并在使用 GLM-5.2 后端时取得 0.821 的总体得分。
⚡ 今天就能用
在设计必须跨越多个步骤保持目标、记忆和交付质量的长周期智能体之前,先阅读这篇论文。
来源与原始报道
本简报汇总并链接到以下媒体的报道。
- arXiv cs.AIA Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon AgentsAug 6, 12:00 PM↗
- arXiv cs.AIContextual Agentic Memory is a Memo, Not True MemoryAug 6, 12:00 PM↗
- arXiv cs.AIOneDayAgent: Towards a Long-Horizon Harness for Autonomous AgentsAug 6, 12:00 PM↗
- arXiv cs.LGOneDayAgent: Towards a Long-Horizon Harness for Autonomous AgentsAug 6, 12:00 PM↗
- arXiv cs.LGEvolveNet: Collaborative Harness Evolution for Agent Self-ImprovementAug 6, 12:00 PM↗
- arXiv cs.AIBeyond Retrieval: Analytic Memory for Multimodal AgentsAug 5, 12:00 PM↗
- arXiv cs.AIWeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent NetworksAug 5, 12:00 PM↗
- arXiv cs.AIVerifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model AgentsAug 5, 12:00 PM↗
- arXiv cs.CLMetis: Memory Foundation ModelAug 5, 12:00 PM↗
- arXiv cs.LGMetis: Memory Foundation ModelAug 5, 12:00 PM↗
- arXiv cs.AIWhen Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation BoundaryAug 4, 12:00 PM↗
- arXiv cs.AIStop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIAgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI AgentsAug 4, 12:00 PM↗
- arXiv cs.AIMemoryForge: Synthesize Lifelong Memory for Human-Like LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIHarness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure TrajectoriesAug 4, 12:00 PM↗
- arXiv cs.AIMemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIWhen Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation BoundaryAug 4, 12:00 PM↗
- arXiv cs.AIV-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic MemoryAug 4, 12:00 PM↗
- arXiv cs.AITrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue AgentsAug 4, 12:00 PM↗
- arXiv cs.AIPMMC: Prospective Multimodal Memory Compilation for Long-Term LVLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIPersonalizing Large Language Model Agents with Small Policy ModelsAug 4, 12:00 PM↗
- arXiv cs.LGMemoryForge: Synthesize Lifelong Memory for Human-Like LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.LGHarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent HarnessesAug 4, 12:00 PM↗
- arXiv cs.LGStop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.CLHarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent HarnessesAug 4, 12:00 PM↗
- arXiv cs.CLV-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic MemoryAug 4, 12:00 PM↗
- arXiv cs.CLAgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI AgentsAug 4, 12:00 PM↗
- arXiv cs.CLMemoryForge: Synthesize Lifelong Memory for Human-Like LLM AgentsAug 4, 12:00 PM↗
- HF Daily PapersOneDayAgent: Towards a Long-Horizon Harness for Autonomous AgentsAug 4, 4:00 AM↗
觉得这篇简报有用?下一篇直接送到你的邮箱。