OneDayAgent、長期的な自律エージェントのワークフローを検証
このハーネスは、104件のベンチマークタスクで、タスク分解、メモリ、検証を管理する。
なぜ重要か
この研究は、目標の逸脱、状態の喪失、コンテキストのあふれといった、エージェント型システムに残る失敗パターンに焦点を当てている。3つのモデルファミリーに属する5種類のバックエンドLLMで示された性能は、ハーネス設計が単一モデルに依存せず、長期的な信頼性を高めうることを示唆している。
要点
- 1.OneDayAgentは、分解、メモリ、最終検証を管理する。
- 2.GLM-5.2を用いた104件のAgentIF-OneDayタスクで0.821を記録した。
- 3.このハーネスは、3つのモデルファミリーにまたがる5種類のLLMバックエンドで動作した。
研究者らは、仕事、学習、生活にまたがる自由度の高い日常的な依頼に対応する自律エージェント向けハーネス「OneDayAgent」を発表した。このシステムは依頼を範囲の限定されたサブタスクに分解し、コンテキスト圧迫下でも実行メモリを維持し、最終成果物を検証して修復する。AgentIF-OneDayでは104件のタスクで評価され、GLM-5.2をバックエンドに用いた場合、総合スコア0.821を達成した。
⚡ 今日から使える
多くのステップを通じて目標、メモリ、成果物の品質を保つ必要がある長期的エージェントを設計する前に、この論文を読んでおきたい。
出典・一次報道
本記事は以下の媒体の報道を要約し、リンクしています。
- arXiv cs.AIA Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon AgentsAug 6, 12:00 PM↗
- arXiv cs.AIContextual Agentic Memory is a Memo, Not True MemoryAug 6, 12:00 PM↗
- arXiv cs.AIOneDayAgent: Towards a Long-Horizon Harness for Autonomous AgentsAug 6, 12:00 PM↗
- arXiv cs.LGOneDayAgent: Towards a Long-Horizon Harness for Autonomous AgentsAug 6, 12:00 PM↗
- arXiv cs.LGEvolveNet: Collaborative Harness Evolution for Agent Self-ImprovementAug 6, 12:00 PM↗
- arXiv cs.AIBeyond Retrieval: Analytic Memory for Multimodal AgentsAug 5, 12:00 PM↗
- arXiv cs.AIWeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent NetworksAug 5, 12:00 PM↗
- arXiv cs.AIVerifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model AgentsAug 5, 12:00 PM↗
- arXiv cs.CLMetis: Memory Foundation ModelAug 5, 12:00 PM↗
- arXiv cs.LGMetis: Memory Foundation ModelAug 5, 12:00 PM↗
- arXiv cs.AIWhen Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation BoundaryAug 4, 12:00 PM↗
- arXiv cs.AIStop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIAgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI AgentsAug 4, 12:00 PM↗
- arXiv cs.AIMemoryForge: Synthesize Lifelong Memory for Human-Like LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIHarness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure TrajectoriesAug 4, 12:00 PM↗
- arXiv cs.AIMemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIWhen Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation BoundaryAug 4, 12:00 PM↗
- arXiv cs.AIV-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic MemoryAug 4, 12:00 PM↗
- arXiv cs.AITrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue AgentsAug 4, 12:00 PM↗
- arXiv cs.AIPMMC: Prospective Multimodal Memory Compilation for Long-Term LVLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIPersonalizing Large Language Model Agents with Small Policy ModelsAug 4, 12:00 PM↗
- arXiv cs.LGMemoryForge: Synthesize Lifelong Memory for Human-Like LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.LGHarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent HarnessesAug 4, 12:00 PM↗
- arXiv cs.LGStop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.CLHarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent HarnessesAug 4, 12:00 PM↗
- arXiv cs.CLV-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic MemoryAug 4, 12:00 PM↗
- arXiv cs.CLAgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI AgentsAug 4, 12:00 PM↗
- arXiv cs.CLMemoryForge: Synthesize Lifelong Memory for Human-Like LLM AgentsAug 4, 12:00 PM↗
- HF Daily PapersOneDayAgent: Towards a Long-Horizon Harness for Autonomous AgentsAug 4, 4:00 AM↗
この記事が役に立ちましたか?次号をメールでお届けします。
関連: 研究
研究
Google、準同型暗号でプライベートAIを前進
Googleは、準同型暗号を活用してプライベートAIをより実用的にしようとしていると説明している。
Hacker News+8 outlets·8h ago
研究
Google、プライベートAIの実用化を進めていると説明
Google Securityの投稿がHacker Newsで議論を呼び、AI全般への批判とも並んだ。
Hacker News+5 outlets·8h ago
研究
研究者ら、LLMを制御下で研究するためのLittleLearnerを公開
50億パラメータのモデルは、小学5年生以下のカリキュラムに基づく880億トークンのコーパスで訓練された。
r/LocalLLaMA+1 outlet·14h ago