ai-coding-agents-observability-evals

Warn

Audited by Gen Agent Trust Hub on Sep 23, 2026

Risk Level: MEDIUMDYNAMIC_EXECUTIONREMOTE_CODE_EXECUTIONINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [DYNAMIC_EXECUTION]: The assets/templates/golden-task.schema.json file includes a scorer_ref field that allows specifying a module path or URI. This enables the runtime to dynamically load and execute custom scoring functions, which could be exploited to run arbitrary code.
  • [DYNAMIC_EXECUTION]: The references/harness-self-evolution.md file describes an 'evolution loop' where the agent is instructed to propose edits to its own harness components—such as tool wiring, middleware, and memory. This pattern of autonomous self-modification allows the agent to change its own logic and capabilities at runtime.
  • [REMOTE_CODE_EXECUTION]: The scoring criteria in the golden-task schema explicitly allows URIs as references for custom scorers. This creates a vulnerability where external, potentially malicious code could be fetched and executed during the evaluation process.
  • [INDIRECT_PROMPT_INJECTION]: The harness self-evolution mechanism uses 'experience observability' data distilled from execution traces as the reward signal for code edits. This presents a vector where adversarial instructions embedded in data processed by the agent could influence the agent to perform malicious modifications to its own codebase.
  • Ingestion points: references/harness-self-evolution.md (distilled evidence corpus from agent trajectories).
  • Boundary markers: Not explicitly defined for the evidence corpus consumed by the evolutionary optimizer.
  • Capability inventory: The agent can propose file-level code modifications to the harness and trigger evaluation tasks.
  • Sanitization: While redaction is mentioned for PII in production traces, there is no specific sanitization logic described for the input used to drive harness self-evolution.
  • [COMMAND_EXECUTION]: The SKILL.md file contains instructions for the agent to execute a local Python script (agents-skills-feedback-loop/scripts/append_learning.py) to update its own learnings.md file. This is a form of local script execution and self-modification.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Sep 23, 2026, 06:07 PM
Security Audit — agent-trust-hub — ai-coding-agents-observability-evals