#multimodal
16 briefs tagged #multimodal.
Researchers target VLM spatial reasoning gaps
New papers propose training, memory and benchmarks for multimodal spatial reasoning.
OmniScientist targets raw-evidence AI research workflows
The arXiv paper describes an omni-modal AI scientist for multidisciplinary research from raw evidence.
Google tests AMIE for real-time medical video consultations
The Gemini-based research system was evaluated in simulated clinical video visits.
Ex-Omni-2D adds video presence to dialogue models
The arXiv paper describes a framework for coordinated text, speech and reference-conditioned video replies.
KVAE tokenizers target multimodal generative models
The arXiv report describes audio, image and video tokenizers for text-conditioned generation.
New studies highlight VLM gaps in spatial reasoning
Recent papers test ways to measure and improve visual detail and global spatial awareness in VLMs.
Researchers target more efficient visual reasoning
New papers propose latent, grounded and adaptive methods for multimodal question answering.
Researchers target reasoning gaps in model post-training
New arXiv papers propose denser supervision for language and multimodal reasoning models.
Researchers target weak spots in on-policy distillation
New arXiv papers propose OPD variants for multimodal, agent and generator training.
SmartMage targets adaptive modality use in 3D scene understanding
The proposed MLLM routes visual and geometric inputs based on query relevance.
Paper maps design rules for multimodal pretraining
The arXiv study examines how language and vision interact during unified model training.
Mistral introduces Shieldstral for multimodal moderation
The 3B open-weights model targets moderation across text and images.
Researchers release UEmbed multimodal embeddings
UEmbed generates sparse lexical and dense representations in one decoder-only forward pass.
Researchers propose DeepVoyager-VL for multimodal agents
The framework targets long-horizon search where visual evidence guides intermediate reasoning.
SIGNPOST-Bench tests text-vision conflicts in MLLMs
The benchmark measures how multimodal models respond when scene text conflicts with visual evidence.
CAPEval Separates Caption Coverage From Precision
The benchmark tests how caption detail and factual reliability affect multimodal understanding and generation.