model-evaluation

Pass

Audited by Gen Agent Trust Hub on Jul 14, 2026

Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill workflow involves generating and executing local scripts to calculate performance metrics from user-supplied medical imaging data.\n
  • Evidence: The SKILL.md file explicitly states that the agent generates and executes code to compute metrics like Dice, HD95, and AUROC.\n
  • Evidence: The tool uses Bash to run analysis and validation scripts (e.g., scripts/check_metric_reporting.py).\n- [SAFE]: The skill's components are transparent and restricted to the stated purpose of medical data analysis.\n
  • Verification: Analysis of scripts/check_metric_reporting.py confirms it is a self-contained script using only the Python standard library for regex-based report auditing.\n
  • Verification: No evidence of obfuscation, hardcoded credentials, persistence mechanisms, or unauthorized network operations was found in any of the 19 files.\n
  • Verification: The skill author, Aperivue, uses external references only for established scientific methodology and documentation.
Audit Metadata
Risk Level
SAFE
Analyzed
Jul 14, 2026, 02:11 AM
Security Audit — agent-trust-hub — model-evaluation