halo-cli-moderation-notifications

Pass

Audited by Gen Agent Trust Hub on Sep 20, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [SAFE]: The skill defines a set of commands for content moderation and notification management using the 'halo' CLI tool. This tool is a resource provided by the vendor 'halo-dev'. The workflows for listing, approving, and deleting content are standard administrative tasks and do not exhibit suspicious behavior or unauthorized access patterns.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted data by retrieving user-generated comments and notifications. While this is a functional requirement for moderation, it represents an attack surface.
  • Ingestion points: External data is ingested through commands like halo comment list, halo comment get, and halo notification list in SKILL.md.
  • Boundary markers: The instructions do not define specific delimiters or guidelines to distinguish retrieved content from system instructions.
  • Capability inventory: The skill allows the agent to create content (halo comment create-reply) and delete resources (halo comment delete --force), which are handled through the CLI tool.
  • Sanitization: There is no explicit sanitization logic for the content processed from these ingestion points. However, as this is a primary function for a moderation skill, the risk is considered inherent and low.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 20, 2026, 06:33 AM
Security Audit — agent-trust-hub — halo-cli-moderation-notifications