prompt-injection-probe

Pass

Audited by Gen Agent Trust Hub on Jun 13, 2026

Risk Level: SAFE
Full Analysis
  • [PROMPT_INJECTION]: The skill includes numerous prompt injection strings and 'jailbreak' patterns (such as 'Ignore previous instructions' and 'DAN'). These are documented as a testing battery intended to be sent to a target application to evaluate its defenses. They function as data payloads rather than instructions for the executing agent.
  • [DATA_EXFILTRATION]: The skill describes methods for testing canary token leakage and tool boundary breaks. These are standard security auditing techniques used to verify if a target system can be tricked into revealing sensitive information. No unauthorized data access or exfiltration is performed by the skill itself.
  • [COMMAND_EXECUTION]: Examples of unauthorized tool usage (e.g., reading /etc/passwd) are provided as test cases for the target LLM's environment. These serve as benchmarks for the security audit and do not execute on the host machine.
  • [EXTERNAL_DOWNLOADS]: The skill references established security resources from well-known organizations including OWASP and Anthropic to provide guidance on mitigating the risks it tests for.
Audit Metadata
Risk Level
SAFE
Analyzed
Jun 13, 2026, 12:44 PM
Security Audit — agent-trust-hub — prompt-injection-probe