llm-eval-lab

Fail

Audited by Gen Agent Trust Hub on Aug 26, 2026

Risk Level: HIGHCREDENTIALS_UNSAFEDATA_EXFILTRATIONCOMMAND_EXECUTIONEXTERNAL_DOWNLOADSPROMPT_INJECTION
Full Analysis
  • [CREDENTIALS_UNSAFE]: The file grid_eval.py contains a hardcoded authentication token sk-dev-proxy-123 within the _agent_judge function, which is used to authenticate requests to the local evaluation proxy at http://localhost:4001.
  • [DATA_EXFILTRATION]: The shell script run.sh contains logic to source environment variables from a hardcoded absolute path: ${HOME}/workspace/experiments/sparta/.env. This exposes sensitive configurations and credentials from external directories to the skill's execution environment.
  • [COMMAND_EXECUTION]: The eval_app.py module uses subprocess.run to dynamically execute pip install commands at runtime if the typer or rich packages are missing, bypassing standard package management practices.
  • [EXTERNAL_DOWNLOADS]: The pyproject.toml file defines a dependency on the scillm package, which is retrieved directly from a GitHub repository (github.com/grahama1970/scillm).
  • [PROMPT_INJECTION]: The skill facilitates automated LLM evaluation by loading ground-truth JSON files and interpolating their contents into model prompts and automated judge instructions, creating a surface for indirect prompt injection.
  • Ingestion points: External ground-truth JSON files loaded in find_minimum.py and grid_eval.py.
  • Boundary markers: None present; data is interpolated directly into message arrays.
  • Capability inventory: Shell command execution via subprocess.run and network requests via httpx to local service endpoints.
  • Sanitization: Input content is stripped but not sanitized for instruction-carrying patterns.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Aug 26, 2026, 06:00 PM
Security Audit — agent-trust-hub — llm-eval-lab