AAI News Hub
商业Tue, August 18, 2026·22h ago2 家媒体交叉佐证

Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

arXiv:2603.

为什么重要

Tracked across 4 sources; see linked reporting for full context.

核心要点

  • 1.arXiv:2603.
  • 2.06697v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks.

arXiv:2603. 06697v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time.

来源与原始报道

本简报汇总并链接到以下媒体的报道。

觉得这篇简报有用?下一篇直接送到你的邮箱。

更多 商业