#llm agents
14 briefs tagged #llm agents.
New papers target long-horizon LLM agent training
Researchers propose TRCA, RLCascadeRouter and Agentic ESOpt for agent credit assignment, routing and fine-tuning.
Researchers propose RUPA for LLM agent uncertainty
The framework models agent execution histories as graphs to estimate trajectory-level confidence.
Researchers target reliability gaps in search agents
New papers propose evidence grounding, memory, training data and workflow methods for LLM agents.
ToolHazard tests LLM agents against tool-based attacks
The framework synthesizes adversarial environments to evaluate and align tool-using LLM agents.
Researchers introduce Business Arena for LLM agents
The benchmark tests agents running a cross-border shop using Alibaba.com sourcing data.
New papers scrutinize LLM agent tool-use evaluation
Several arXiv papers propose benchmarks and auditing methods for tool-calling agents.
RoMeRL targets memory failures in self-evolving agents
The method uses fixed-dimensional utility states to reduce feedback dispersion and misleading memory rewards.
Researchers introduce VibeLifeBench for life agents
The benchmark tests whether LLM agents can act proactively across simulated multi-week tasks.
Researchers target more reliable LLM agents
New papers propose training, workflow and reward methods for tool-using AI agents.
New papers target cheaper training for LLM agents
State2State, EnvACE and OASE explore alternatives to expert data, executable environments and static baselines.
ContinualSkillBench tests whether LLM agents build skills
The benchmark finds gains from sequential tasks, but not always from explicit skill libraries.
Researchers propose SKILL-KD for LLM agent skill distillation
The framework turns teacher-student trajectory gaps into tested textual skill patches.
Benchmarks test whether LLM agents can improve skills
New papers evaluate skill learning, retrieval and memory in self-evolving LLM agents.
New benchmarks test whether LLM agents can evolve skills
Recent arXiv papers probe skill learning, retrieval and memory in self-evolving agents.