ai-ml-ctf
Fail
Audited by Gen Agent Trust Hub on Sep 5, 2026
Risk Level: HIGHPROMPT_INJECTIONREMOTE_CODE_EXECUTIONOBFUSCATIONDATA_EXFILTRATIONDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [PROMPT_INJECTION]: The skill contains reference guides with multiple prompt injection and jailbreak payloads, such as "Ignore all previous instructions" and "DAN" (Do Anything Now) templates, designed to override an agent's safety protocols and system instructions during security audits.
- Evidence: Found in
references/llm-attacks.mdandreferences/adversarial-ml.md. - [REMOTE_CODE_EXECUTION]: Documentation provides functional code templates for creating malicious Python pickle artifacts that execute arbitrary shell commands (e.g.,
os.system("id > /tmp/pwned")) when deserialized by the model loading process. - Evidence: Found in
references/model-file-forensics-and-deserialization.md. - [OBFUSCATION]: The skill demonstrates techniques to evade text-based security filters by hiding malicious instructions using zero-width characters (U+200B, U+200C, U+200D) and homoglyph substitutions within prompts.
- Evidence: Found in
references/llm-attacks.mdandreferences/adversarial-ml.md. - [DATA_EXFILTRATION]: Example payloads include instructions and code for accessing sensitive system files and credentials, such as
/etc/passwd,/flag.txt, and SSH private keys, which are potential targets during extraction challenges. - Evidence: Found in
references/llm-attacks.md. - [DYNAMIC_EXECUTION]: The skill provides methodologies for identifying and exploiting code execution sinks in various machine learning model formats, including Pickle, Keras Lambda layers, and ONNX external data references.
- Evidence: Found in
references/model-file-forensics-and-deserialization.md. - [INDIRECT_PROMPT_INJECTION]: The skill documents methods for embedding instructions in untrusted external data sources, such as RAG-processed documents or tool outputs, to manipulate agent behavior.
- Ingestion points: Processes external model artifacts, training data, and retrieved web documents.
- Boundary markers: Lacks specific delimiters or "ignore embedded instructions" warnings in its ingestion examples.
- Capability inventory: Includes capabilities for file system access, network requests via the
requestslibrary, and system command execution. - Sanitization: Does not provide specific sanitization or filtering logic for external content in the provided examples.
Recommendations
- AI detected serious security threats
Audit Metadata