skills/cofin/flow/challenge/Gen Agent Trust Hub

challenge

Pass

Audited by Gen Agent Trust Hub on Jun 26, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill is composed entirely of markdown instructions and YAML metadata. It contains no executable scripts, binaries, or network-enabled commands.
  • [PROMPT_INJECTION]: The workflow instructions (e.g., 'Ignore reflexive agreement') are designed to improve analytical reasoning and avoid cognitive bias (sycophancy). These instructions do not attempt to bypass core safety guardrails or ignore system-level constraints.
  • [PROMPT_INJECTION]: The Indirect Prompt Injection surface was assessed as follows:
  • Ingestion points: The agent processes user-provided claims or assertions as described in the workflow in SKILL.md.
  • Boundary markers: There are no explicit technical delimiters (e.g., XML tags or special characters) defined for the input data.
  • Capability inventory: The skill identifies that the agent should 'read the code' or 'check the documentation' to verify assertions, which involves standard file-reading tools.
  • Sanitization: The workflow includes a logical sanitization step by instructing the agent to 'strip away qualifiers and framing' to isolate the core claim, effectively filtering out potentially malicious instructional padding.
Audit Metadata
Risk Level
SAFE
Analyzed
Jun 26, 2026, 11:42 AM
Security Audit — agent-trust-hub — challenge