skills/leeroo-ai/superml/ml-debug/Gen Agent Trust Hub

ml-debug

Pass

Audited by Gen Agent Trust Hub on Mar 23, 2026

Risk Level: SAFEEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
  • [SAFE]: The skill is designed for the technical diagnosis of machine learning failures (OOM, NaN, etc.). It implements strict procedural controls (referred to as 'Iron Laws') to prevent speculative analysis and ensure responses are based on verifiable documentation. References to 'Leeroopedia KB' and tools like 'diagnose_failure' represent legitimate vendor-specific functionality from Leeroo-AI.\n- [EXTERNAL_DOWNLOADS]: The skill instructs the agent to use a 'WebFetch' capability to retrieve technical facts, release tags, and configurations from well-known and trusted services. Targeted sources include the official GitHub repositories and documentation sites for PyTorch, HuggingFace (Transformers, PEFT), Microsoft (DeepSpeed), vLLM, and Axolotl, as well as the PyPI registry. These downloads are performed to ensure diagnosis is grounded in version-specific technical documentation.\n- [COMMAND_EXECUTION]: The skill provides the user with runnable verification scripts and CLI commands (e.g., Python scripts for concurrent load testing and memory profiling). These commands are intended to be executed by the user to verify the efficacy of a proposed fix and do not execute automatically on the host.\n- [PROMPT_INJECTION]: The skill processes untrusted user-supplied data in the form of error logs and symptoms (Phase 1). This creates a surface for indirect prompt injection.\n
  • Ingestion points: User-provided symptoms and error logs in Phase 1.\n
  • Boundary markers: None explicitly specified to delineate log data from agent instructions.\n
  • Capability inventory: Access to 'WebFetch' for network operations and generating user-executable scripts.\n
  • Sanitization: No specific sanitization or filtering of log content is described.\n Despite this surface, the risk is assessed as low given the skill's specific purpose of log analysis and its requirement for documentation-based grounding.
Audit Metadata
Risk Level
SAFE
Analyzed
Mar 23, 2026, 02:35 PM
Security Audit — agent-trust-hub — ml-debug