AAI News Hub

#benchmarks

40 briefs tagged #benchmarks.

Research

HarnessRisk benchmark tests agent harness safety

The benchmark evaluates failures across tool, state, permission and recovery phases for LLM agents.

arXiv cs.AI+1 outlet·9h ago
Models

GLM-5.3 benchmark page draws developer discussion

Artificial Analysis benchmarks for GLM-5.3 surfaced on Hacker News and r/LocalLLaMA.

Hacker News+1 outlet·15h ago
Research

Researchers introduce ASI-Bench for AI research tasks

The benchmark tests innovative exploration and autonomous scientific execution across 11 domains.

HF Daily Papers+1 outlet·1d ago
Research

StartupBench tests agents on real startup workflows

The new benchmark finds the strongest tested model completes only about 30% of tasks.

HF Daily Papers+1 outlet·1d ago
Models

Qwen3.8 27B scores 52 on Artificial Analysis

Reddit and Hacker News users flagged a sharp benchmark jump for Qwen3.8 27B.

r/LocalLLaMA+1 outlet·1d ago
Research

R^3-Bench tests LLM reasoning under shared budgets

The benchmark finds six models often lag simple allocation baselines across reasoning suites.

HF Daily Papers+1 outlet·2d ago
Research

New papers probe spatial intelligence in AI vision models

SpaRRTa, SMA and PinpointQA target spatial reasoning gaps in visual and multimodal systems.

arXiv cs.LG+1 outlet·5d ago
Research

Researchers target VLM spatial reasoning gaps

New papers propose training, memory and benchmarks for multimodal spatial reasoning.

arXiv cs.AI+1 outlet·6d ago
Research

Researchers improve logical compound-answer reasoning

A new framework decomposes AND, OR and NEITHER/NOR options before composing predictions.

HF Daily Papers+1 outlet·6d ago
Models

Grok 4.6 scores 61 on Artificial Analysis index

Artificial Analysis benchmarked Grok 4.6, while xAI published its Grok 4.6 announcement.

Hacker News+1 outlet·6d ago
Research

Researchers introduce MBA-Bench for business ideation agents

The multimodal benchmark includes 30K samples across six business domains.

HF Daily Papers+1 outlet·Aug 12
Research

Researchers introduce Business Arena for LLM agents

The benchmark tests agents running a cross-border shop using Alibaba.com sourcing data.

arXiv cs.AI+1 outlet·Aug 11
Research

Study finds LLM search can fingerprint benchmarks

GPU-kernel tests showed 30% of in-distribution wins failed on held-out configurations.

arXiv cs.AI+1 outlet·Aug 11
Research

New papers scrutinize LLM agent tool-use evaluation

Several arXiv papers propose benchmarks and auditing methods for tool-calling agents.

arXiv cs.CL+1 outlet·Aug 11
Research

DSAgentBench tests agents on real data-science workflows

The benchmark evaluates end-to-end data-science tasks in real computer environments.

HF Daily Papers+1 outlet·Aug 11
Research

New benchmarks expose gaps in mobile AI agents

Researchers test mobile agents on judging, personal data use and real-world interface robustness.

HF Daily Papers+1 outlet·Aug 11
Research

Researchers introduce VibeLifeBench for life agents

The benchmark tests whether LLM agents can act proactively across simulated multi-week tasks.

HF Daily Papers+1 outlet·Aug 11
Models

Local testers benchmark Muse Glimmer 30B

Reddit users report strong local fit and speed, but weaker coding results than Qwen 3.6 27B.

r/LocalLLaMA+2 outlets·Aug 10
Research

Researchers target more reliable LLM agents

New papers propose training, workflow and reward methods for tool-using AI agents.

HF Daily Papers+1 outlet·Aug 10
Research

Researchers introduce SWE-Bench ProMax

The benchmark tests AI coding agents on 170 multilingual code refactoring tasks.

HF Daily Papers+1 outlet·Aug 10
Research

Sci-VBench tests scientific reasoning in video models

The benchmark evaluates whether video generators can model scientific and causal dynamics.

HF Daily Papers+1 outlet·Aug 10
Research

HarnessOpt-Bench tests LLMs on agent harness optimization

The benchmark measures how LLMs improve prompts, tools, memory and orchestration around agents.

arXiv cs.AI+1 outlet·Aug 7
Research

Researchers release Yiddish-focused MameLoshnLM

The open-source 8B model uses a new corpus and benchmark for low-resource Yiddish NLP.

arXiv cs.AI+1 outlet·Aug 7
Research

HarnessOpt-Bench evaluates LLM harness optimization

The arXiv benchmark tests how frontier LLMs improve agent prompts, tools, control flow and memory.

arXiv cs.AI+1 outlet·Aug 7
Models

Qwen3.8 Max tops Artificial Analysis agentic index

Reddit and Hacker News users flagged the model’s ranking and debated changes to the benchmark.

r/LocalLLaMA+1 outlet·Aug 7
Research

Researchers adapt Nemotron retrieval stack for Modern Greek

The work adds Greek retrieval training, reranking, grounded generation and a HERA benchmark.

arXiv cs.CL+1 outlet·Aug 6
Research

New papers probe VLMs' spatial reasoning limits

Benchmarks and methods point to gaps in dense perception, global spatial awareness and visual detail use.

arXiv cs.AI+1 outlet·Aug 6
Research

Researchers scale synthetic tasks for terminal agents

RST and CalibForge propose verified task-generation methods for training terminal agents.

HF Daily Papers+1 outlet·Aug 6
Research

ContinualSkillBench tests whether LLM agents build skills

The benchmark finds gains from sequential tasks, but not always from explicit skill libraries.

arXiv cs.AI+1 outlet·Aug 5
Research

GDPevo benchmark targets agent self-evolution

The benchmark tests whether agents can reuse prior experience on enterprise workflows.

arXiv cs.AI+1 outlet·Aug 5
Research

Benchmarks test whether LLM agents can improve skills

New papers evaluate skill learning, retrieval and memory in self-evolving LLM agents.

arXiv cs.CL+1 outlet·Aug 5
Business

AI spending concerns rise as new capability tests land

Reports highlight heavy AI debt, concentrated cloud demand and fresh math and coding benchmarks.

Hacker News+8 outlets·Aug 5
Research

Papers target better training data for terminal agents

RST and CalibForge synthesize verified terminal tasks for long-horizon agent training.

HF Daily Papers+1 outlet·Aug 5
Research

NOLLI benchmark probes English-Korean model gaps

The new puzzle benchmark tests whether Korean gaps come from language, script or culture-specific tasks.

HF Daily Papers+1 outlet·Aug 5
Research

Researchers propose Skill Entropy for LLM reasoning

Skill^2-Bench tests how models switch across 558 skills in long-horizon tasks.

HF Daily Papers+1 outlet·Aug 5
Research

ScrambleToolBench tests agents without semantic tool cues

The benchmark probes whether agents can infer hidden tool behavior and adapt when mappings change.

arXiv cs.CL+1 outlet·Aug 4
Research

New benchmarks test whether LLM agents can evolve skills

Recent arXiv papers probe skill learning, retrieval and memory in self-evolving agents.

arXiv cs.AI+1 outlet·Aug 4
Research

SWE-Touch tests coding agents in shared workspaces

The benchmark finds user counter-edits reduce coding-agent resolve rates on SWE-bench Verified.

arXiv cs.AI+1 outlet·Aug 4
Research

LongHorizon-Harness targets long-horizon agent state

The research harness separates task state from execution and verifies updates through an audit loop.

HF Daily Papers·Aug 3
Research

WorldExam tests video models for world reactivity

The benchmark evaluates whether controllable video models infer plausible consequences from scene state.

HF Daily Papers·Aug 3