1 briefs tagged #ai4ai-bench.
The arXiv benchmark evaluates whether LLM agents can rewrite training algorithms across frozen repositories.