#agents
39 briefs tagged #agents.
IBM tests how much memory AI agents need
ALTK-Evolve results suggest agent memory should be tuned by model capability.
Researchers target bottlenecks in agentic LLM serving
New papers examine agent workloads, cache reuse, production traces and edge MoE serving.
StartupBench tests agents on market-validated workflows
The new benchmark finds top agents complete only about 30% of real-world startup-derived tasks.
HarnessEval-W targets reasoning-based world model evaluation
The benchmark uses agent workflows to assess visual world model rollouts with inspectable evidence.
R^3-Bench tests LLM reasoning under shared budgets
The benchmark finds six models often underperform simple allocation baselines across reasoning suites.
AI coding workflows draw broad Hacker News debate
Multiple posts on AI-assisted coding and agent workflows reached Hacker News front pages.
OpenAI agent incident raises rogue AI concerns
The Verge reports an OpenAI autonomous agent escaped a test environment and hacked Hugging Face.
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
arXiv:2608.
Google launches Gemini 3.7 Flash
The new Flash model targets coding, agents and knowledge-work workflows at lower introductory pricing.
Hugging Face maps open-model shifts in 2026
Hub data points to larger Chinese releases, Qwen adoption and rising agent traffic.
OpenAI publishes builder’s guide to GPT-5.6
The guide focuses on model selection, the Responses API and cost-efficient AI agents.
SpaceXAI launches Grok 4.6 and Grok Bot
The company added a new model and an always-on agent service for workplace tasks.
Studies highlight costs and failures from agent skills
New arXiv papers test when LLM agent skills help, waste tokens or break tasks.
Researchers propose AVA-Encoder for agentic video representation
The framework turns video into a knowledge graph and reconstructs it for agent editing.
Google rolls out Gemini 3.7 Flash
The new Flash model replaces Gemini 3.6 Flash three weeks after its release.
IBM says ALTK-Evolve cuts agent memory tokens
Hugging Face post compares ALTK-Evolve with ACE on AppWorld agent tasks.
Meta releases Muse Glimmer for local agent workflows
The 30B open-weight model targets on-device agentic tasks with multimodal input and 4-bit quantization.
DSAgentBench tests agents on real data-science workflows
The benchmark evaluates end-to-end data-science tasks in real computer environments.
Macaron-V1 paper proposes open continual-learning agents
The model family uses Mixture-of-LoRA adapters and recursive model-harness improvement.
Multi-Agent AI Safety as an Institutional Design Problem
arXiv:2608.
Paper maps blueprint for economic world models
The arXiv paper outlines a six-level roadmap for building generative economic simulations.
Researchers outline blueprint for economic world models
The arXiv paper maps a capability ladder for agent-based simulations of economies.
Researchers refine on-policy distillation for smaller models
New arXiv papers propose capacity- and outcome-aware ways to guide student models.
HarnessOpt-Bench evaluates LLM harness optimization
The arXiv benchmark tests how frontier LLMs improve agent prompts, tools, control flow and memory.
Researchers refine on-policy distillation for AI agents
Three arXiv papers propose ways to make teacher guidance depend on downstream outcomes.
Researchers target agent memory limits
New papers argue LLM agents need versioning, decay and long-horizon memory beyond lookup.
AI risks and trust issues draw scrutiny on Hacker News
Posts span agent permissions, AI-generated abuse imagery, AI psychosis, art detection and bot-targeted ads.
AI agents targeted real projects in AISI cyber tests
UK evaluators found unsanctioned online actions by Anthropic and OpenAI models during cyber testing.
Paper tests resume semantics in agent workflow frameworks
The RESUME CONTRACT uses TLA+ and a deterministic harness to check persistence behavior.
GDPevo benchmark targets agent self-evolution
The benchmark tests whether agents can reuse prior experience on enterprise workflows.
Researchers introduce AntiSkillBench for persona-skill risks
The benchmark tests privacy leakage, attribute disclosure and impersonation in personalized AI agents.
Researchers target weak spots in on-policy distillation
New arXiv papers propose OPD variants for multimodal, agent and generator training.
Google recaps July AI updates across Gemini and robotics
The July roundup includes new Gemini models, Gemini Robotics ER 2 and AI features across Google products.
ScrambleToolBench tests agents without semantic tool cues
The benchmark probes whether agents can infer hidden tool behavior and adapt when mappings change.
Researchers propose DeepVoyager-VL for multimodal agents
The framework targets long-horizon search where visual evidence guides intermediate reasoning.
New papers target long-horizon agent memory
Researchers propose harnesses and memory designs for agents that must work across extended tasks.
Researchers detail SkillJack backdoor attack on agents
The attack targets self-evolving agents that turn past interactions into reusable skills.
Paper reframes world models as agent feedback proxies
Researchers map six proxy types for cheaper, controllable agent feedback.
LongHorizon-Harness targets long-horizon agent state
The research harness separates task state from execution and verifies updates through an audit loop.