PAST-Bench tests whether personal agents improve over time
The benchmark measures retained-experience gains across models, frameworks and agent-loop pathways.
Why it matters
The work gives agent developers a more targeted way to evaluate whether memory and experience retention actually improve behavior, rather than only measuring end-task gains. It also highlights that recursive self-improvement claims need pathway evidence, not just aggregate performance changes.
The key points
- 1.PAST-Bench evaluates retained-experience gains in personal AI agents.
- 2.The benchmark spans 26 scenarios and 204 episodes.
- 3.Observed improvements were real but uneven across models and frameworks.
Researchers introduced PAST-Bench, a benchmark for testing whether personal AI agents use retained experience to improve future behavior across fresh-session task sequences. The benchmark covers 26 scenarios and 204 episodes spanning memory, procedural reuse, information gathering and update, comparing conditions with retained experience turned on and off. Across seven base models and four agent frameworks, the reports say improvement was real but uneven, and similar headline gains did not always reflect the intended save, retrieve and update pathway.
⚡ Try this today
Use PAST-Bench-style controls to test whether an agent’s memory system improves later tasks through the intended save, retrieve and update steps.
Sources & original reporting
This brief summarizes and links to reporting from the publishers below.
Enjoyed this brief? Get the next one in your inbox.
More in Research
Google advances private AI with homomorphic encryption
Google says it is making private AI more practical using homomorphic encryption.
Google says it is making private AI practical
A Google Security post drew Hacker News discussion alongside broader AI critiques.
Researchers release LittleLearner training sandbox
The 5B-parameter model was trained on an 88B-token Grade 5-limited corpus.