AAI News Hub
ResearchTue, August 4, 2026·Aug 42 sources corroborating

GROVE adds layered memory for streaming video assistants

The training-free framework shares one video memory for reactive QA and proactive assistance.

Why it matters

The work points to memory architectures as a key lever for wearable and video-language assistants, where models must use past visual evidence before future questions are known. It also shows growing interest in improving Video-LLM behavior without changing the underlying model.

The key points

  • 1.GROVE builds temporal memory from continuous video streams.
  • 2.It supports both reactive QA and proactive assistance.
  • 3.ObjectStream uses latent objects as persistent video-memory anchors.

Researchers introduced GROVE, a training-free framework for streaming video experience that builds memory causally from a continuous video stream. The system keeps fine-grained perceptual evidence and consolidates it into time-stamped moments, coherent episodes and recurring cross-day patterns, with retrieval skills matched to each level. The authors report that GROVE achieved the best results among compared methods on multiple benchmarks, including MM-lifelong and EgoServe. A related arXiv paper, ObjectStream, also targets streaming video understanding with training-free memory, using latent objects as persistent anchors.

Try this today

Read the GROVE and ObjectStream papers before designing long-context memory for streaming video assistants.

Sources & original reporting

This brief summarizes and links to reporting from the publishers below.

Enjoyed this brief? Get the next one in your inbox.

More in Research