AAI News Hub
ResearchTue, August 4, 2026·Aug 42 sources corroborating

New papers target long-horizon agent memory

Researchers propose harnesses and memory designs for agents that must work across extended tasks.

Why it matters

The papers reflect a shift from isolated model benchmarks toward agent infrastructure that preserves goals, manages accumulated experience and improves behavior over time. They also highlight unresolved risks, including limits of retrieval-based memory and vulnerability to persistent memory poisoning.

The key points

  • 1.OneDayAgent reports 0.821 on 104 long-horizon tasks.
  • 2.Several papers frame memory as core agent infrastructure, not just storage.
  • 3.Authors warn retrieval-only memory may limit generalization and security.

A set of recent arXiv papers examines how LLM-based agents can handle long-horizon tasks that exceed fixed context windows. The work includes OneDayAgent, a harness that decomposes open-ended requests, maintains execution memory and verifies deliverables, reporting a 0.821 score on 104 AgentIF-OneDay tasks with a GLM-5.2 backend. Other papers argue current vector-store and scratchpad systems are closer to lookup than true memory, and propose approaches such as analytic multimodal memory and collaborative harness evolution.

Try this today

Before building long-horizon agents, evaluate memory and harness design separately from the base model, including retrieval limits, verification, and poisoning risks.

Sources & original reporting

This brief summarizes and links to reporting from the publishers below.

Enjoyed this brief? Get the next one in your inbox.

More in Research