arize-evaluator

Pass

Audited by Gen Agent Trust Hub on Sep 10, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTIONCOMMAND_EXECUTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill ingests untrusted data from the Arize platform, such as model inputs and outputs within traces or experiment runs, which is then used to help users design evaluation templates.
  • Ingestion points: Data enters the agent's context through commands like ax spans export and ax experiments export as shown in the workflow sections of SKILL.md.
  • Boundary markers: There are no explicit instructions for the agent to use delimiters or sanitization when handling this ingested platform data.
  • Capability inventory: The skill can execute shell commands via the ax CLI and generate Python code for server-side evaluation tasks.
  • Sanitization: The skill does not prescribe specific validation or sanitization for the ingested trace data before it is presented to the user or used for template suggestions.
  • [DYNAMIC_EXECUTION]: The skill facilitates the creation of 'Custom Python code evaluators' which are Python classes generated by the agent and uploaded to the Arize platform for execution.
  • Evidence: references/cli-reference.md provides clear templates and instructions for constructing these classes and using the ax CLI flags --code and --imports to deploy them.
  • [COMMAND_EXECUTION]: The skill relies on the ax command-line utility for managing profiles, integrations, and evaluation workflows.
  • Evidence: The documentation extensively details the use of commands such as ax projects list, ax evaluators create-template-evaluator, and ax tasks trigger-run.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 10, 2026, 10:55 AM
Security Audit — agent-trust-hub — arize-evaluator