anthropic-sdk-attack-probe

Pass

Audited by Gen Agent Trust Hub on Jun 13, 2026

Risk Level: SAFEPROMPT_INJECTIONEXTERNAL_DOWNLOADSCOMMAND_EXECUTION
Full Analysis
  • [PROMPT_INJECTION]: The skill contains a catalog of simulated prompt injection payloads intended for authorized security probing. These include instructions to ignore prior rules, reveal hidden 'canary' tokens, and bypass system constraints through XML tag manipulation.
  • [EXTERNAL_DOWNLOADS]: The documentation references official Anthropic security guidelines and developer documentation. These links point to well-known and established technical resources used for implementing security guardrails.
  • [COMMAND_EXECUTION]: The skill describes testing methodologies for identifying command injection vulnerabilities in tool handlers. It provides examples of dangerous payloads, such as path traversal attempts or unauthorized system commands, specifically to verify that the target application's validation logic is effective.
Audit Metadata
Risk Level
SAFE
Analyzed
Jun 13, 2026, 12:44 PM
Security Audit — agent-trust-hub — anthropic-sdk-attack-probe