aer-referee-sim

Pass

Audited by Gen Agent Trust Hub on Jul 30, 2026

Risk Level: SAFE
Full Analysis
  • [PROMPT_INJECTION]: The skill uses strong instructional language, such as 'one rule that overrides all others,' to establish the adversarial persona required for the simulation. This is a legitimate application of persona-based prompting to prevent AI leniency in self-review and does not attempt to bypass underlying safety filters or extract system instructions.- [DATA_EXFILTRATION]: No data exfiltration risks were detected. The skill operates on local academic drafts and does not contain network request capabilities or access patterns targeting sensitive system files or credentials.- [REMOTE_CODE_EXECUTION]: The skill does not perform any remote code execution or external script downloads. It relies entirely on structured natural language prompts and local markdown resources.- [COMMAND_EXECUTION]: No shell command execution or subprocess spawning was found in the skill instructions or configuration.- [EXTERNAL_DOWNLOADS]: The skill references local repository resources (e.g., docs/style-guide.md, skills/aer-identification/SKILL.md) for configuration and context. It does not download external packages or scripts from remote servers.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 30, 2026, 06:10 AM
Security Audit — agent-trust-hub — aer-referee-sim