output-scoring-en

Pass

Audited by Gen Agent Trust Hub on Aug 25, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill consists exclusively of instructional text and logical frameworks for evaluating AI outputs. It does not include any executable scripts (Python, JavaScript, shell) or configuration files that could initiate malicious behavior.
  • [SAFE]: Permissions are restricted to the Read tool, and the metadata explicitly declares pii-egress: none and data-residency: local. This configuration prevents the agent from exfiltrating data or performing unauthorized network operations.
  • [SAFE]: External references in the attribution and canonical source fields point to well-known research repositories (Fudan DISC Lab) and the author's own official GitHub organization. No typosquatting, obfuscated URLs, or malicious redirection patterns were found.
  • [SAFE]: The instructions do not contain any prompt injection attempts, persistence mechanisms, or privilege escalation patterns. The logic focuses entirely on the stated goal of legal output scoring.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 25, 2026, 11:09 AM
Security Audit — agent-trust-hub — output-scoring-en