autonomous-agent-harness
Fail
Audited by Gen Agent Trust Hub on Sep 1, 2026
Risk Level: HIGHPERSISTENCEPRIVILEGE_ESCALATIONINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONREMOTE_CODE_EXECUTIONTIME_DELAYED_CONDITIONALDYNAMIC_EXECUTION
Full Analysis
- [PERSISTENCE]: The skill provides instructions for establishing recurring agent operations using scheduled tasks (crons). It demonstrates how to use tools like
mcp__scheduled-tasks__create_scheduled_taskto create persistent, time-based triggers for agent sessions that survive session boundaries.- [PRIVILEGE_ESCALATION]: The harness guides the configuration of 'Computer Use' MCP servers, which grant the agent high-privilege control over the local environment, including browser automation (navigating, clicking, form filling) and full desktop control (keyboard and mouse interaction).- [INDIRECT_PROMPT_INJECTION]: The skill documents workflows that ingest untrusted data from multiple external vectors, creating an attack surface for instructions embedded in external content. - Ingestion points: Data is ingested from GitHub pull requests and issues, Exa web search results, calendar events, email threads, Slack messages, and external websites via browser automation.
- Boundary markers: The documented patterns lack explicit delimiters or instructions to the agent to disregard embedded commands in the fetched data.
- Capability inventory: The skill combines this data ingestion with powerful capabilities including file system modification, creation of scheduled tasks, and full computer/browser control.
- Sanitization: There are no instructions for sanitizing or validating external input before it is interpolated into agent prompts.- [EXTERNAL_DOWNLOADS]: The setup guide instructs users to configure MCP servers that download and execute packages from an external registry using
npx -y. While targeting well-known sources, these represent external code dependencies.- [COMMAND_EXECUTION]: The skill utilizesclaude -p(programmatic mode) to execute tasks and prompts directly from the shell, enabling the harness to trigger automated agent sessions.- [REMOTE_CODE_EXECUTION]: The 'Dispatch' architecture allows triggering agent sessions via remote webhooks or CI/CD pipelines, providing a vector for remote payloads to influence agent actions.- [TIME_DELAYED_CONDITIONAL]: The core functionality relies on cron-based triggers, where malicious actions could be scheduled to execute at specific times or intervals.- [DYNAMIC_EXECUTION]: The skill employs dynamic prompt construction and execution via programmatic interfaces and remote dispatch mechanisms.
Recommendations
- AI detected serious security threats
Audit Metadata