llm-evaluation
Pass
Audited by Gen Agent Trust Hub on Sep 15, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill content is entirely safe. It includes documentation and sample implementations for standard evaluation metrics (such as BLEU, ROUGE, BERTScore) and framework integrations (such as LangSmith) without any malicious behaviors, prompt injections, or unsafe operations.
Audit Metadata