#agent evaluation
2 briefs tagged #agent evaluation.
Research
RUPA targets trajectory-level uncertainty in LLM agents
The arXiv paper proposes graph-based uncertainty propagation across agent execution histories.
arXiv cs.CL+1 outlet·2d ago
Research
New papers scrutinize LLM agent tool-use evaluation
Several arXiv papers propose benchmarks and auditing methods for tool-calling agents.
arXiv cs.CL+1 outlet·Aug 11