ag2-evaluation
Pass
Audited by Gen Agent Trust Hub on Jul 4, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides documentation and Python code examples for evaluating AG2 agents. All described operations, including running agents against test suites, scoring outputs with prebuilt or custom scorers, and persisting results to local directories for regression analysis, are standard and appropriate for a software evaluation framework. The recommended installation command utilizes the author's own library, and no suspicious network activity, data exfiltration, or obfuscation techniques were identified.
Audit Metadata