AAI News Hub

#llm agents

14 briefs tagged #llm agents.

Research

New papers target long-horizon LLM agent training

Researchers propose TRCA, RLCascadeRouter and Agentic ESOpt for agent credit assignment, routing and fine-tuning.

arXiv cs.AI+1 outlet·1d ago
Research

Researchers propose RUPA for LLM agent uncertainty

The framework models agent execution histories as graphs to estimate trajectory-level confidence.

arXiv cs.CL+1 outlet·1d ago
Research

Researchers target reliability gaps in search agents

New papers propose evidence grounding, memory, training data and workflow methods for LLM agents.

arXiv cs.AI+1 outlet·Aug 12
Research

ToolHazard tests LLM agents against tool-based attacks

The framework synthesizes adversarial environments to evaluate and align tool-using LLM agents.

HF Daily Papers+1 outlet·Aug 12
Research

Researchers introduce Business Arena for LLM agents

The benchmark tests agents running a cross-border shop using Alibaba.com sourcing data.

arXiv cs.AI+1 outlet·Aug 11
Research

New papers scrutinize LLM agent tool-use evaluation

Several arXiv papers propose benchmarks and auditing methods for tool-calling agents.

arXiv cs.CL+1 outlet·Aug 11
Research

RoMeRL targets memory failures in self-evolving agents

The method uses fixed-dimensional utility states to reduce feedback dispersion and misleading memory rewards.

arXiv cs.CL+1 outlet·Aug 11
Research

Researchers introduce VibeLifeBench for life agents

The benchmark tests whether LLM agents can act proactively across simulated multi-week tasks.

HF Daily Papers+1 outlet·Aug 11
Research

Researchers target more reliable LLM agents

New papers propose training, workflow and reward methods for tool-using AI agents.

HF Daily Papers+1 outlet·Aug 10
Research

New papers target cheaper training for LLM agents

State2State, EnvACE and OASE explore alternatives to expert data, executable environments and static baselines.

arXiv cs.CL+1 outlet·Aug 6
Research

ContinualSkillBench tests whether LLM agents build skills

The benchmark finds gains from sequential tasks, but not always from explicit skill libraries.

arXiv cs.AI+1 outlet·Aug 5
Research

Researchers propose SKILL-KD for LLM agent skill distillation

The framework turns teacher-student trajectory gaps into tested textual skill patches.

arXiv cs.AI+1 outlet·Aug 5
Research

Benchmarks test whether LLM agents can improve skills

New papers evaluate skill learning, retrieval and memory in self-evolving LLM agents.

arXiv cs.CL+1 outlet·Aug 5
Research

New benchmarks test whether LLM agents can evolve skills

Recent arXiv papers probe skill learning, retrieval and memory in self-evolving agents.

arXiv cs.AI+1 outlet·Aug 4