#evaluation
5 briefs tagged #evaluation.
Research
Studies highlight costs and failures from agent skills
New arXiv papers test when LLM agent skills help, waste tokens or break tasks.
arXiv cs.AI+1 outlet·3d ago
Research
Researchers introduce Taboo stress test for LLM robustness
The zero-prompt diagnostic masks candidate tokens at runtime to test off-path model behavior.
arXiv cs.CL+1 outlet·4d ago
Research
AI safety testing faces scrutiny as risks broaden
Reports point to gaps in containment and human-subject evidence for AI safety work.
TechCrunch AI+1 outlet·5d ago
Research
AllenAI releases TutorMoments for evaluating AI tutors
The preview tests whether LLM tutors know when to help students and when to hold back.
Hugging Face·Aug 8
Research
Legal IR papers target leakage in case-law retrieval
New work proposes temporally fenced retrieval and a CJEU dataset for more realistic legal search evaluation.
arXiv cs.CL+1 outlet·Aug 4