#vlm
8 briefs tagged #vlm.
Researchers target hidden-state reasoning in video VLMs
New arXiv papers propose latent training methods to cut video reasoning overhead and improve visual grounding.
New papers probe spatial intelligence in AI vision models
SpaRRTa, SMA and PinpointQA target spatial reasoning gaps in visual and multimodal systems.
New papers target long-document VQA bottlenecks
InSight-doc, DocAtlas and DocMemo propose adaptive ways to find evidence across visually rich documents.
Researchers target visual evidence gaps in VLMs
New arXiv papers propose evidence selection, retrieval and token-pruning methods for multimodal QA.
EffectLearner targets object effects in video removal
The framework pairs VLM-based effect reasoning with a DiT video eraser for real-world video editing.
PosterMELD automates editable scientific posters
The multi-agent pipeline turns papers into print-ready PPTX and PNG posters with design controls.
Researchers target visual-token bottlenecks in VLMs
New arXiv papers propose training-free and adaptive methods to cut image and video inference costs.
Researchers target trajectory bias in driving VLMs
The paper proposes AD-MCQ and DEFT-RLVR to make autonomous-driving reasoning more verifiable.