Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens
Claw-like agents are always-on processes with persistent access to credentials, files, tools, and external services, creating security risks that model- and tool-call-level benchmarks do not capture. This work treats them as agentic computer systems and introduces SafeClawArena, a benchmark of 406 adversarial tasks spanning skill supply-chain integrity, persistent-state exploitation, cross-boundary data flow, and indirect prompt injection. Evaluations across three agent platforms and five frontier models expose substantial cross-component vulnerabilities and limitations in current platform defenses.