AAI News Hub
ResearchThu, August 6, 2026·Aug 62 sources corroborating

New studies highlight VLM gaps in spatial reasoning

Recent papers test ways to measure and improve visual detail and global spatial awareness in VLMs.

Why it matters

The work points to spatial perception as a major bottleneck for VLMs used in embodied agents and video understanding. It also suggests that adding visual-token reasoning, tool use or better labeling of visual details may be necessary for reliable multimodal systems.

The key points

  • 1.VLMs still struggle with dense visual perception and spatial awareness.
  • 2.GST-Bench reports a large zero-shot gap between models and humans.
  • 3.COVT and SpatialCLI propose training methods to improve perceptual reasoning.

Several recent papers examine a common weakness in vision-language models: they can reason in language but often miss fine-grained visual and spatial details. COVT proposes continuous visual tokens distilled from vision experts to capture appearance, geometry, layout and edges within roughly 20 tokens. GST-Bench evaluates global spatial awareness in video and reports the best zero-shot model at 42.68 versus a human score of 79.08, while SpatialCLI trains VLMs to use and internalize spatial tools.

Try this today

Read the GST-Bench and SpatialCLI papers before relying on a VLM for long-horizon spatial or embodied-agent tasks.

Sources & original reporting

This brief summarizes and links to reporting from the publishers below.

Enjoyed this brief? Get the next one in your inbox.

More in Research