fin-kg-define-measurement-test
Pass
Audited by Gen Agent Trust Hub on Jul 17, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill facilitates a governed workflow for knowledge graph benchmarking through deterministic Python scripts that handle initialization, validation, and evidence packaging.
- [SAFE]: All scripts use secure coding practices, such as employing
yaml.safe_loadfor parsing external artifacts and utilizing SHA-256 hashing to ensure data integrity across the specification and run packages. - [SAFE]: The system includes a proactive security feature (
scan_boundary_leakageinmts_common.py) that audits processed data for prohibited keys and phrases, effectively preventing prompt injection or schema confusion that could introduce subjective quality judgments into the measurement layer. - [SAFE]: The skill correctly segregates design-time intent from runtime execution facts and evidence, maintaining clear data lineage and preventing persistence or state manipulation attacks.
- [SAFE]: No unauthorized network connections, access to sensitive local environment paths (e.g., credentials), or obfuscated commands were identified in any of the provided files.
Audit Metadata