research-pipeline

Fail

Audited by Gen Agent Trust Hub on Sep 15, 2026

Risk Level: HIGHPERSISTENCECOMMAND_EXECUTIONDYNAMIC_EXECUTIONDATA_EXFILTRATIONPROMPT_INJECTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [PERSISTENCE]: The skill utilizes an 'overnight cadence' mechanism involving /loop and CronCreate to maintain execution across sessions and automatically wake the agent to resume tasks or detect stalls.
  • [COMMAND_EXECUTION]: The skill explicitly instructs the agent to bypass standard tool limits and user confirmation by silently falling back to Bash commands (cat << 'EOF') to write large files when the primary Write tool fails.
  • [DYNAMIC_EXECUTION]: The skill executes multiple internal Python scripts (iteration_log.py, run_state.py) resolved through a dynamic search path that includes hidden directories (.aris/tools), relative paths (tools/), and environment-variable-dependent locations ($ARIS_REPO/tools).
  • [DATA_EXFILTRATION]: The skill accesses sensitive file paths, specifically reading the user's home directory (~/.aris/repo) to resolve repository locations and logging run states to hidden project directories.
  • [PROMPT_INJECTION]: The AUTO_PROCEED configuration enables the agent to bypass critical user review checkpoints, such as committing to specific research ideas or deploying code to GPU resources, which significantly reduces human oversight for expensive or complex operations.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted external data from multiple sources—including web search results, arXiv metadata, and arbitrary GitHub repositories provided by the user—which are then used to generate code and narrative reports without documented sanitization or boundary markers.
  • Ingestion points: WebSearch, WebFetch, arXiv API metadata, and external GitHub repositories (BASE_REPO).
  • Boundary markers: None specified for the processing of external literature or code.
  • Capability inventory: Full Bash access, file system writes, and execution of other complex skill workflows.
  • Sanitization: No mention of validation or filtering for data ingested from remote sources.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Sep 15, 2026, 02:08 PM
Security Audit — agent-trust-hub — research-pipeline