skills/leeroo-ai/superml/ml-research/Gen Agent Trust Hub

ml-research

Pass

Audited by Gen Agent Trust Hub on Mar 23, 2026

Risk Level: SAFEPROMPT_INJECTIONEXTERNAL_DOWNLOADSCOMMAND_EXECUTION
Full Analysis
  • [PROMPT_INJECTION]: The skill uses authoritative steering instructions (e.g., 'The Iron Law', 'ZERO-TOLERANCE RULE') to override the AI's default behavior of answering from memory. These instructions act as safety-aligned constraints to minimize hallucinations by forcing grounding in external documentation.
  • [PROMPT_INJECTION]: Output control measures are present (e.g., 'Do NOT write any prose... before completing at least 3 WebFetch calls') to ensure tool use precedes response generation. This enforces a specific operational sequence for factual reliability.
  • [EXTERNAL_DOWNLOADS]: The skill references a registry of documentation URLs from well-known services and trusted organizations, including Hugging Face, PyTorch, LangChain, and vLLM. These downloads are restricted to technical documentation and configuration examples required for the skill's primary research purpose.
  • [COMMAND_EXECUTION]: The skill generates complete Python scripts, Dockerfiles, and CLI commands. While this creates a potential surface for indirect prompt injection from processed web content, the risk is mitigated by the skill's extensive Phase 1 and Phase 2 verification gates, and the generated content is intended for user-side execution.
  • [SAFE]: The skill demonstrates high-quality grounding practices by identifying hallucination-prone areas (e.g., benchmark numbers, model specs, config schemas) and requiring verbatim verification from retrieved sources. Ingestion points (WebFetch) are limited to the provided URL registry, and capability inventory is consistent with a research-oriented tool.
Audit Metadata
Risk Level
SAFE
Analyzed
Mar 23, 2026, 02:35 PM
Security Audit — agent-trust-hub — ml-research