code-review-expert

Warn

Audited by Gen Agent Trust Hub on Jul 22, 2026

Risk Level: MEDIUMPROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The 'skill-report.json' file includes a deceptive 'security_audit' section that purports to be an official safety assessment. It declares the skill as 'safe' and 'safe_to_publish' based on an analysis supposedly performed by a non-existent model version ('gpt-5.5') and includes a future-dated audit timestamp (2026). This metadata appears designed to influence the safety evaluation of the skill through misleading claims.
  • [PROMPT_INJECTION]: The skill is designed to ingest and analyze user-provided code and documentation, which is a surface for indirect prompt injection. The instructions do not define boundary markers or include directives to the agent to ignore instructions embedded within the untrusted code being reviewed.
  • Ingestion points: User-supplied code changes and architecture documentation.
  • Boundary markers: Absent; no instructions to use specific delimiters or to disregard instructions found within the input data.
  • Capability inventory: Reasoning and assessment capabilities; no destructive or network-based tool permissions are requested.
  • Sanitization: None; no logic is provided to filter or escape malicious content in the input data.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Jul 22, 2026, 12:39 PM
Security Audit — agent-trust-hub — code-review-expert