model-eval

Fail

Audited by Gen Agent Trust Hub on Apr 5, 2026

Risk Level: HIGHCOMMAND_EXECUTIONDATA_EXFILTRATIONPROMPT_INJECTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill executes high-privilege AWS CLI commands (aws lambda update-function-configuration) to modify the environment variables of live Lambda functions. This allows the agent to change the operational state of infrastructure.\n- [DATA_EXFILTRATION]: The skill retrieves data from extralife_chat_logs.chat_logs via Athena. This involves accessing potentially sensitive user conversation history and tool usage logs.\n- [PROMPT_INJECTION]: The scoring mechanism is vulnerable to indirect prompt injection.\n
  • Ingestion points: Raw user queries are pulled from Athena logs in Phase 0 and model responses are collected during evaluation.\n
  • Boundary markers: The scoring template in references/SCORING-TEMPLATE.md directly interpolates untrusted data into the judge prompt without delimiters or safety instructions.\n
  • Capability inventory: The skill has permissions to modify Lambda configurations and execute SQL queries via Athena.\n
  • Sanitization: No validation or escaping is performed on the ingested messages before they are processed by the judge model.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Apr 5, 2026, 12:44 AM
Security Audit — agent-trust-hub — model-eval