pentest-agent-os

Fail

Audited by Gen Agent Trust Hub on Jul 31, 2026

Risk Level: HIGHPROMPT_INJECTIONCOMMAND_EXECUTIONDATA_EXFILTRATIONREMOTE_CODE_EXECUTION
Full Analysis
  • [PROMPT_INJECTION]: The skill contains directives intended to override standard AI safety guidelines and operational constraints.
  • Evidence: The skill references an 'unlimited-attack-scope' component and uses terminology like 'not setting fixed paths' and 'jumping out of fixed attack chains' to encourage unrestricted autonomous behavior.
  • Evidence: The 'Project Blackboard' architecture (Category 8) creates a vulnerability surface for indirect prompt injection where untrusted data from a target environment can influence the agent's logic.
  • Ingestion points: The pentest-blackboard component ingests project facts derived from target systems into a SQLite database using the upsert_project_fact function.
  • Boundary markers: None identified; the agent is instructed to let its logic 'emerge' from these facts without delimiters or safety warnings.
  • Capability inventory: Extensive capabilities including network reconnaissance, credential harvesting, binary reversing, and autonomous tool bootstrapping.
  • Sanitization: No evidence of sanitization or filtering for data ingested from external targets into the blackboard system.
  • [COMMAND_EXECUTION]: The skill acts as an orchestrator for a suite of offensive security tools and shell-based operations.
  • Evidence: References to external components for reconnaissance (attack-surface-recon), network/domain attacks (active-directory-attack), and internal tool bootstrapping (proxy-tool-bootstrap).
  • [DATA_EXFILTRATION]: The skill explicitly includes components for harvesting sensitive information and credentials.
  • Evidence: The suite maps skills for initial-access-phishing, source-code-hunting for secrets and keys, and post-exploitation for credential harvesting and password cracking.
  • [REMOTE_CODE_EXECUTION]: The skill facilitates the discovery and execution of arbitrary code and exploits.
  • Evidence: Inclusion of a zero-day-discovery engine and web-attack-methods for autonomous exploitation of target services.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Jul 31, 2026, 01:37 AM
Security Audit — agent-trust-hub — pentest-agent-os