OneDayAgent, 장기 자율 에이전트 워크플로를 시험하다
이 하네스는 104개 벤치마크 과제 전반에서 작업 분해, 메모리, 검증을 관리한다.
왜 중요한가
이번 연구는 목표 이탈, 상태 손실, 컨텍스트 초과 등 에이전트 시스템의 고질적인 실패 양상을 겨냥한다. 세 개 모델 계열의 다섯 가지 백엔드 LLM에서 보고된 성능은 하네스 설계가 특정 단일 모델에 의존하지 않고 장기 작업 신뢰성을 높일 수 있음을 시사한다.
핵심 포인트
- 1.OneDayAgent는 작업 분해, 메모리, 최종 검증을 관리한다.
- 2.GLM-5.2를 사용해 104개 AgentIF-OneDay 과제에서 0.821점을 기록했다.
- 3.이 하네스는 세 개 모델 계열의 다섯 가지 LLM 백엔드에서 실행됐다.
연구진은 업무, 학습, 생활 전반의 개방형 일상 요청을 처리하는 자율 에이전트용 하네스 OneDayAgent를 공개했다. 이 시스템은 요청을 범위가 정해진 하위 작업으로 나누고, 컨텍스트 압박 속에서도 실행 메모리를 유지하며, 최종 산출물을 검증하고 수정한다. AgentIF-OneDay에서는 104개 과제를 대상으로 평가됐고, GLM-5.2 백엔드를 사용해 종합 점수 0.821을 기록했다.
⚡ 오늘 바로 활용
여러 단계를 거치며 목표, 메모리, 산출물 품질을 유지해야 하는 장기 에이전트를 설계하기 전에 이 논문을 읽어볼 필요가 있다.
출처 및 원본 보도
이 브리핑은 아래 매체의 보도를 요약하고 링크합니다.
- arXiv cs.AIA Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon AgentsAug 6, 12:00 PM↗
- arXiv cs.AIContextual Agentic Memory is a Memo, Not True MemoryAug 6, 12:00 PM↗
- arXiv cs.AIOneDayAgent: Towards a Long-Horizon Harness for Autonomous AgentsAug 6, 12:00 PM↗
- arXiv cs.LGOneDayAgent: Towards a Long-Horizon Harness for Autonomous AgentsAug 6, 12:00 PM↗
- arXiv cs.LGEvolveNet: Collaborative Harness Evolution for Agent Self-ImprovementAug 6, 12:00 PM↗
- arXiv cs.AIBeyond Retrieval: Analytic Memory for Multimodal AgentsAug 5, 12:00 PM↗
- arXiv cs.AIWeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent NetworksAug 5, 12:00 PM↗
- arXiv cs.AIVerifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model AgentsAug 5, 12:00 PM↗
- arXiv cs.CLMetis: Memory Foundation ModelAug 5, 12:00 PM↗
- arXiv cs.LGMetis: Memory Foundation ModelAug 5, 12:00 PM↗
- arXiv cs.AIWhen Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation BoundaryAug 4, 12:00 PM↗
- arXiv cs.AIStop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIAgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI AgentsAug 4, 12:00 PM↗
- arXiv cs.AIMemoryForge: Synthesize Lifelong Memory for Human-Like LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIHarness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure TrajectoriesAug 4, 12:00 PM↗
- arXiv cs.AIMemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIWhen Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation BoundaryAug 4, 12:00 PM↗
- arXiv cs.AIV-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic MemoryAug 4, 12:00 PM↗
- arXiv cs.AITrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue AgentsAug 4, 12:00 PM↗
- arXiv cs.AIPMMC: Prospective Multimodal Memory Compilation for Long-Term LVLM AgentsAug 4, 12:00 PM↗
- arXiv cs.AIPersonalizing Large Language Model Agents with Small Policy ModelsAug 4, 12:00 PM↗
- arXiv cs.LGMemoryForge: Synthesize Lifelong Memory for Human-Like LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.LGHarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent HarnessesAug 4, 12:00 PM↗
- arXiv cs.LGStop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM AgentsAug 4, 12:00 PM↗
- arXiv cs.CLHarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent HarnessesAug 4, 12:00 PM↗
- arXiv cs.CLV-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic MemoryAug 4, 12:00 PM↗
- arXiv cs.CLAgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI AgentsAug 4, 12:00 PM↗
- arXiv cs.CLMemoryForge: Synthesize Lifelong Memory for Human-Like LLM AgentsAug 4, 12:00 PM↗
- HF Daily PapersOneDayAgent: Towards a Long-Horizon Harness for Autonomous AgentsAug 4, 4:00 AM↗
이 브리핑이 유용했나요? 다음 소식을 메일로 받아보세요.
관련: 연구
연구
Google, 동형암호로 프라이빗 AI 진전
Google은 동형암호를 활용해 프라이빗 AI를 더 실용적으로 만들고 있다고 밝혔다.
Hacker News+8 outlets·6h ago
연구
Google, private AI를 실용화하고 있다고 밝혀
Google Security 게시물이 Hacker News에서 논의를 불러왔고, 더 넓은 AI 비판과도 맞물렸다.
Hacker News+5 outlets·6h ago
연구
연구진, 통제된 LLM 연구용 LittleLearner 공개
50억 개 파라미터 모델로, 초등학교 5학년 이하 수준의 커리큘럼 말뭉치 880억 토큰으로 학습됐다.
r/LocalLLaMA+1 outlet·11h ago