#benchmarking
3 briefs tagged #benchmarking.
Research
TRACE-Bench targets multi-reference image generation
The benchmark decomposes prompts into operators to diagnose model failures by capability.
arXiv cs.AI+1 outlet·1d ago
Research
Cultivar benchmark tests translation localisation robustness
The FLORES subset probes contamination and locale gaps across 32 open-weight models.
arXiv cs.CL+1 outlet·Aug 11
Research
SIGNPOST-Bench tests text-vision conflicts in MLLMs
The benchmark measures how multimodal models respond when scene text conflicts with visual evidence.
HF Daily Papers·Aug 4