llm-eval-lab
Fail
Audited by Gen Agent Trust Hub on Aug 26, 2026
Risk Level: HIGHCREDENTIALS_UNSAFEDATA_EXFILTRATIONCOMMAND_EXECUTIONEXTERNAL_DOWNLOADSPROMPT_INJECTION
Full Analysis
- [CREDENTIALS_UNSAFE]: The file
grid_eval.pycontains a hardcoded authentication tokensk-dev-proxy-123within the_agent_judgefunction, which is used to authenticate requests to the local evaluation proxy athttp://localhost:4001. - [DATA_EXFILTRATION]: The shell script
run.shcontains logic to source environment variables from a hardcoded absolute path:${HOME}/workspace/experiments/sparta/.env. This exposes sensitive configurations and credentials from external directories to the skill's execution environment. - [COMMAND_EXECUTION]: The
eval_app.pymodule usessubprocess.runto dynamically executepip installcommands at runtime if thetyperorrichpackages are missing, bypassing standard package management practices. - [EXTERNAL_DOWNLOADS]: The
pyproject.tomlfile defines a dependency on thescillmpackage, which is retrieved directly from a GitHub repository (github.com/grahama1970/scillm). - [PROMPT_INJECTION]: The skill facilitates automated LLM evaluation by loading ground-truth JSON files and interpolating their contents into model prompts and automated judge instructions, creating a surface for indirect prompt injection.
- Ingestion points: External ground-truth JSON files loaded in
find_minimum.pyandgrid_eval.py. - Boundary markers: None present; data is interpolated directly into message arrays.
- Capability inventory: Shell command execution via
subprocess.runand network requests viahttpxto local service endpoints. - Sanitization: Input content is stripped but not sanitized for instruction-carrying patterns.
Recommendations
- AI detected serious security threats
Audit Metadata