1 briefs tagged #3d-scene-understanding.
The proposed MLLM routes visual and geometric inputs based on query relevance.