ai-coding-agents-observability-evals
Warn
Audited by Gen Agent Trust Hub on Sep 23, 2026
Risk Level: MEDIUMDYNAMIC_EXECUTIONREMOTE_CODE_EXECUTIONINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
- [DYNAMIC_EXECUTION]: The
assets/templates/golden-task.schema.jsonfile includes ascorer_reffield that allows specifying a module path or URI. This enables the runtime to dynamically load and execute custom scoring functions, which could be exploited to run arbitrary code. - [DYNAMIC_EXECUTION]: The
references/harness-self-evolution.mdfile describes an 'evolution loop' where the agent is instructed to propose edits to its own harness components—such as tool wiring, middleware, and memory. This pattern of autonomous self-modification allows the agent to change its own logic and capabilities at runtime. - [REMOTE_CODE_EXECUTION]: The
scoringcriteria in the golden-task schema explicitly allows URIs as references for custom scorers. This creates a vulnerability where external, potentially malicious code could be fetched and executed during the evaluation process. - [INDIRECT_PROMPT_INJECTION]: The harness self-evolution mechanism uses 'experience observability' data distilled from execution traces as the reward signal for code edits. This presents a vector where adversarial instructions embedded in data processed by the agent could influence the agent to perform malicious modifications to its own codebase.
- Ingestion points:
references/harness-self-evolution.md(distilled evidence corpus from agent trajectories). - Boundary markers: Not explicitly defined for the evidence corpus consumed by the evolutionary optimizer.
- Capability inventory: The agent can propose file-level code modifications to the harness and trigger evaluation tasks.
- Sanitization: While redaction is mentioned for PII in production traces, there is no specific sanitization logic described for the input used to drive harness self-evolution.
- [COMMAND_EXECUTION]: The
SKILL.mdfile contains instructions for the agent to execute a local Python script (agents-skills-feedback-loop/scripts/append_learning.py) to update its ownlearnings.mdfile. This is a form of local script execution and self-modification.
Audit Metadata