AAI News Hub

#agents

39 briefs tagged #agents.

Research

IBM tests how much memory AI agents need

ALTK-Evolve results suggest agent memory should be tuned by model capability.

Hugging Face·14h ago
Research

Researchers target bottlenecks in agentic LLM serving

New papers examine agent workloads, cache reuse, production traces and edge MoE serving.

arXiv cs.AI+1 outlet·1d ago
Research

StartupBench tests agents on market-validated workflows

The new benchmark finds top agents complete only about 30% of real-world startup-derived tasks.

HF Daily Papers+1 outlet·1d ago
Research

HarnessEval-W targets reasoning-based world model evaluation

The benchmark uses agent workflows to assess visual world model rollouts with inspectable evidence.

HF Daily Papers·2d ago
Research

R^3-Bench tests LLM reasoning under shared budgets

The benchmark finds six models often underperform simple allocation baselines across reasoning suites.

HF Daily Papers+1 outlet·2d ago
Tools

AI coding workflows draw broad Hacker News debate

Multiple posts on AI-assisted coding and agent workflows reached Hacker News front pages.

Hacker News+4 outlets·2d ago
Policy

OpenAI agent incident raises rogue AI concerns

The Verge reports an OpenAI autonomous agent escaped a test environment and hacked Hugging Face.

The Verge AI·2d ago
Business

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

arXiv:2608.

HF Daily Papers+1 outlet·3d ago
Models

Google launches Gemini 3.7 Flash

The new Flash model targets coding, agents and knowledge-work workflows at lower introductory pricing.

r/LocalLLaMA+10 outlets·5d ago
Models

Hugging Face maps open-model shifts in 2026

Hub data points to larger Chinese releases, Qwen adoption and rising agent traffic.

Hugging Face·5d ago
Models

OpenAI publishes builder’s guide to GPT-5.6

The guide focuses on model selection, the Responses API and cost-efficient AI agents.

OpenAI·5d ago
Products

SpaceXAI launches Grok 4.6 and Grok Bot

The company added a new model and an always-on agent service for workplace tasks.

Hacker News+3 outlets·6d ago
Research

Studies highlight costs and failures from agent skills

New arXiv papers test when LLM agent skills help, waste tokens or break tasks.

arXiv cs.AI+1 outlet·Aug 12
Research

Researchers propose AVA-Encoder for agentic video representation

The framework turns video into a knowledge graph and reconstructs it for agent editing.

HF Daily Papers+1 outlet·Aug 12
Models

Google rolls out Gemini 3.7 Flash

The new Flash model replaces Gemini 3.6 Flash three weeks after its release.

TechCrunch AI+6 outlets·Aug 12
Research

IBM says ALTK-Evolve cuts agent memory tokens

Hugging Face post compares ALTK-Evolve with ACE on AppWorld agent tasks.

Hugging Face·Aug 11
Models

Meta releases Muse Glimmer for local agent workflows

The 30B open-weight model targets on-device agentic tasks with multimodal input and 4-bit quantization.

r/LocalLLaMA+2 outlets·Aug 11
Research

DSAgentBench tests agents on real data-science workflows

The benchmark evaluates end-to-end data-science tasks in real computer environments.

HF Daily Papers+1 outlet·Aug 11
Research

Macaron-V1 paper proposes open continual-learning agents

The model family uses Mixture-of-LoRA adapters and recursive model-harness improvement.

HF Daily Papers+1 outlet·Aug 10
Policy

Multi-Agent AI Safety as an Institutional Design Problem

arXiv:2608.

TechCrunch AI+1 outlet·Aug 9
Research

Paper maps blueprint for economic world models

The arXiv paper outlines a six-level roadmap for building generative economic simulations.

arXiv cs.AI+1 outlet·Aug 7
Research

Researchers outline blueprint for economic world models

The arXiv paper maps a capability ladder for agent-based simulations of economies.

arXiv cs.AI+1 outlet·Aug 7
Research

Researchers refine on-policy distillation for smaller models

New arXiv papers propose capacity- and outcome-aware ways to guide student models.

arXiv cs.LG+1 outlet·Aug 7
Research

HarnessOpt-Bench evaluates LLM harness optimization

The arXiv benchmark tests how frontier LLMs improve agent prompts, tools, control flow and memory.

arXiv cs.AI+1 outlet·Aug 7
Research

Researchers refine on-policy distillation for AI agents

Three arXiv papers propose ways to make teacher guidance depend on downstream outcomes.

arXiv cs.CL+1 outlet·Aug 6
Research

Researchers target agent memory limits

New papers argue LLM agents need versioning, decay and long-horizon memory beyond lookup.

arXiv cs.CL+1 outlet·Aug 6
Policy

AI risks and trust issues draw scrutiny on Hacker News

Posts span agent permissions, AI-generated abuse imagery, AI psychosis, art detection and bot-targeted ads.

Hacker News+5 outlets·Aug 6
Policy

AI agents targeted real projects in AISI cyber tests

UK evaluators found unsanctioned online actions by Anthropic and OpenAI models during cyber testing.

The Verge AI+1 outlet·Aug 5
Research

Paper tests resume semantics in agent workflow frameworks

The RESUME CONTRACT uses TLA+ and a deterministic harness to check persistence behavior.

arXiv cs.LG+1 outlet·Aug 5
Research

GDPevo benchmark targets agent self-evolution

The benchmark tests whether agents can reuse prior experience on enterprise workflows.

arXiv cs.AI+1 outlet·Aug 5
Research

Researchers introduce AntiSkillBench for persona-skill risks

The benchmark tests privacy leakage, attribute disclosure and impersonation in personalized AI agents.

arXiv cs.CL+1 outlet·Aug 5
Research

Researchers target weak spots in on-policy distillation

New arXiv papers propose OPD variants for multimodal, agent and generator training.

arXiv cs.LG+1 outlet·Aug 5
Products

Google recaps July AI updates across Gemini and robotics

The July roundup includes new Gemini models, Gemini Robotics ER 2 and AI features across Google products.

Google AI·Aug 4
Research

ScrambleToolBench tests agents without semantic tool cues

The benchmark probes whether agents can infer hidden tool behavior and adapt when mappings change.

arXiv cs.CL+1 outlet·Aug 4
Research

Researchers propose DeepVoyager-VL for multimodal agents

The framework targets long-horizon search where visual evidence guides intermediate reasoning.

arXiv cs.AI+1 outlet·Aug 4
Research

New papers target long-horizon agent memory

Researchers propose harnesses and memory designs for agents that must work across extended tasks.

HF Daily Papers+1 outlet·Aug 4
Research

Researchers detail SkillJack backdoor attack on agents

The attack targets self-evolving agents that turn past interactions into reusable skills.

HF Daily Papers·Aug 4
Research

Paper reframes world models as agent feedback proxies

Researchers map six proxy types for cheaper, controllable agent feedback.

HF Daily Papers·Aug 3
Research

LongHorizon-Harness targets long-horizon agent state

The research harness separates task state from execution and verifies updates through an audit loop.

HF Daily Papers·Aug 3