HarnessRisk benchmarks safety failures in agent harnesses
The benchmark tests how agent harness risks emerge across lifecycle phases and configurations.
Why it matters
The results suggest agent safety depends not only on the model but also on harness design, configuration and operational lifecycle controls. The finding that Harness Configuration was the most vulnerable phase highlights a practical risk area for deployed agent systems.
The key points
- 1.HarnessRisk covers six phases of agent harness operation.
- 2.Attack success ranged from 12.6% to 80.9% across tested setups.
- 3.Harness Configuration was the most vulnerable phase.
Researchers introduced HarnessRisk, a lifecycle-oriented benchmark for safety in LLM agent harnesses that manage tools, extensions, persistent state, permissions and external actions. The benchmark includes 128 sandboxed cases pairing benign user objectives with adversarial instructions embedded in untrusted workflow artifacts. Across three harnesses, six language models and 14 model and harness configurations, attack success ranged from 12.6% to 80.9%, while Utility stayed between 75.0% and 97.6%.
⚡ Try this today
Audit agent harness configuration defaults and security-sensitive parameters before expanding tool or permission access.
Sources & original reporting
This brief summarizes and links to reporting from the publishers below.
- arXiv cs.AIHarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness SafetyAug 19, 12:00 PM↗
- arXiv cs.AITask-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure OperationsAug 19, 12:00 PM↗
- HF Daily PapersHarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness SafetyAug 18, 4:00 AM↗
Enjoyed this brief? Get the next one in your inbox.
More in Research
Startup says cancer AI needs better data
A TechCrunch report says the company argues data is the key barrier to cancer breakthroughs.
Agent Lightning v1.0 targets harnessed agentic RL
The framework connects arbitrary agent harnesses to RL training through an LLM endpoint proxy.
Researchers test cross-model transfer for LLM memory
A paper studies moving frozen hashed memory between model backbones using target-side reader training.