compare-results
Pass
Audited by Gen Agent Trust Hub on Jul 19, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides a structured framework for model performance analysis and quantization feasibility assessment without introducing security risks.\n- [DATA_EXFILTRATION]: The mention of 'credentials' in the workflow is restricted to ensuring that the configuration for a baseline evaluation matches the candidate evaluation for scientific consistency. The skill does not access, store, or transmit sensitive information to external parties.\n- [PROMPT_INJECTION]: The skill processes external evaluation artifacts, logs, and MLflow data (Ingestion points: SKILL.md Step 3 and 5). This constitutes a potential surface for indirect prompt injection. However, the workflow mitigates this through mandatory verification steps (Step 4: 'Verify completed evaluation run') and a detailed comparability checklist (Capability inventory: Uses accessing-mlflow and evaluation skills), which ensures the integrity of the data used for the final delta computation.
Audit Metadata