#benchmarks
40 briefs tagged #benchmarks.
HarnessRisk benchmark tests agent harness safety
The benchmark evaluates failures across tool, state, permission and recovery phases for LLM agents.
GLM-5.3 benchmark page draws developer discussion
Artificial Analysis benchmarks for GLM-5.3 surfaced on Hacker News and r/LocalLLaMA.
Researchers introduce ASI-Bench for AI research tasks
The benchmark tests innovative exploration and autonomous scientific execution across 11 domains.
StartupBench tests agents on real startup workflows
The new benchmark finds the strongest tested model completes only about 30% of tasks.
Qwen3.8 27B scores 52 on Artificial Analysis
Reddit and Hacker News users flagged a sharp benchmark jump for Qwen3.8 27B.
R^3-Bench tests LLM reasoning under shared budgets
The benchmark finds six models often lag simple allocation baselines across reasoning suites.
New papers probe spatial intelligence in AI vision models
SpaRRTa, SMA and PinpointQA target spatial reasoning gaps in visual and multimodal systems.
Researchers target VLM spatial reasoning gaps
New papers propose training, memory and benchmarks for multimodal spatial reasoning.
Researchers improve logical compound-answer reasoning
A new framework decomposes AND, OR and NEITHER/NOR options before composing predictions.
Grok 4.6 scores 61 on Artificial Analysis index
Artificial Analysis benchmarked Grok 4.6, while xAI published its Grok 4.6 announcement.
Researchers introduce MBA-Bench for business ideation agents
The multimodal benchmark includes 30K samples across six business domains.
Researchers introduce Business Arena for LLM agents
The benchmark tests agents running a cross-border shop using Alibaba.com sourcing data.
Study finds LLM search can fingerprint benchmarks
GPU-kernel tests showed 30% of in-distribution wins failed on held-out configurations.
New papers scrutinize LLM agent tool-use evaluation
Several arXiv papers propose benchmarks and auditing methods for tool-calling agents.
DSAgentBench tests agents on real data-science workflows
The benchmark evaluates end-to-end data-science tasks in real computer environments.
New benchmarks expose gaps in mobile AI agents
Researchers test mobile agents on judging, personal data use and real-world interface robustness.
Researchers introduce VibeLifeBench for life agents
The benchmark tests whether LLM agents can act proactively across simulated multi-week tasks.
Local testers benchmark Muse Glimmer 30B
Reddit users report strong local fit and speed, but weaker coding results than Qwen 3.6 27B.
Researchers target more reliable LLM agents
New papers propose training, workflow and reward methods for tool-using AI agents.
Researchers introduce SWE-Bench ProMax
The benchmark tests AI coding agents on 170 multilingual code refactoring tasks.
Sci-VBench tests scientific reasoning in video models
The benchmark evaluates whether video generators can model scientific and causal dynamics.
HarnessOpt-Bench tests LLMs on agent harness optimization
The benchmark measures how LLMs improve prompts, tools, memory and orchestration around agents.
Researchers release Yiddish-focused MameLoshnLM
The open-source 8B model uses a new corpus and benchmark for low-resource Yiddish NLP.
HarnessOpt-Bench evaluates LLM harness optimization
The arXiv benchmark tests how frontier LLMs improve agent prompts, tools, control flow and memory.
Qwen3.8 Max tops Artificial Analysis agentic index
Reddit and Hacker News users flagged the model’s ranking and debated changes to the benchmark.
Researchers adapt Nemotron retrieval stack for Modern Greek
The work adds Greek retrieval training, reranking, grounded generation and a HERA benchmark.
New papers probe VLMs' spatial reasoning limits
Benchmarks and methods point to gaps in dense perception, global spatial awareness and visual detail use.
Researchers scale synthetic tasks for terminal agents
RST and CalibForge propose verified task-generation methods for training terminal agents.
ContinualSkillBench tests whether LLM agents build skills
The benchmark finds gains from sequential tasks, but not always from explicit skill libraries.
GDPevo benchmark targets agent self-evolution
The benchmark tests whether agents can reuse prior experience on enterprise workflows.
Benchmarks test whether LLM agents can improve skills
New papers evaluate skill learning, retrieval and memory in self-evolving LLM agents.
AI spending concerns rise as new capability tests land
Reports highlight heavy AI debt, concentrated cloud demand and fresh math and coding benchmarks.
Papers target better training data for terminal agents
RST and CalibForge synthesize verified terminal tasks for long-horizon agent training.
NOLLI benchmark probes English-Korean model gaps
The new puzzle benchmark tests whether Korean gaps come from language, script or culture-specific tasks.
Researchers propose Skill Entropy for LLM reasoning
Skill^2-Bench tests how models switch across 558 skills in long-horizon tasks.
ScrambleToolBench tests agents without semantic tool cues
The benchmark probes whether agents can infer hidden tool behavior and adapt when mappings change.
New benchmarks test whether LLM agents can evolve skills
Recent arXiv papers probe skill learning, retrieval and memory in self-evolving agents.
SWE-Touch tests coding agents in shared workspaces
The benchmark finds user counter-edits reduce coding-agent resolve rates on SWE-bench Verified.
LongHorizon-Harness targets long-horizon agent state
The research harness separates task state from execution and verifies updates through an audit loop.
WorldExam tests video models for world reactivity
The benchmark evaluates whether controllable video models infer plausible consequences from scene state.