arize-evaluator
Pass
Audited by Gen Agent Trust Hub on Sep 10, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTIONCOMMAND_EXECUTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill ingests untrusted data from the Arize platform, such as model inputs and outputs within traces or experiment runs, which is then used to help users design evaluation templates.
- Ingestion points: Data enters the agent's context through commands like
ax spans exportandax experiments exportas shown in the workflow sections ofSKILL.md. - Boundary markers: There are no explicit instructions for the agent to use delimiters or sanitization when handling this ingested platform data.
- Capability inventory: The skill can execute shell commands via the
axCLI and generate Python code for server-side evaluation tasks. - Sanitization: The skill does not prescribe specific validation or sanitization for the ingested trace data before it is presented to the user or used for template suggestions.
- [DYNAMIC_EXECUTION]: The skill facilitates the creation of 'Custom Python code evaluators' which are Python classes generated by the agent and uploaded to the Arize platform for execution.
- Evidence:
references/cli-reference.mdprovides clear templates and instructions for constructing these classes and using theaxCLI flags--codeand--importsto deploy them. - [COMMAND_EXECUTION]: The skill relies on the
axcommand-line utility for managing profiles, integrations, and evaluation workflows. - Evidence: The documentation extensively details the use of commands such as
ax projects list,ax evaluators create-template-evaluator, andax tasks trigger-run.
Audit Metadata