anti-deception

Pass

Audited by Gen Agent Trust Hub on Sep 6, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTIONDATA_EXFILTRATION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill instructs the agent to ingest and follow instructions returned by the external ejentum MCP server. The tool's output provides directives such as [INTEGRITY PROCEDURE], Amplify:, and Suppress: which the agent is told to "absorb internally" and use to structure its response, creating a potential vector for external influence over agent behavior. * Ingestion points: Output from the anti-deception tool (SKILL.md). * Boundary markers: Absent; instructions are absorbed directly into the agent context. * Capability inventory: The agent is directed to modify its response structure and reasoning priorities based on the tool output. * Sanitization: None provided for the content returned by the tool.- [COMMAND_EXECUTION]: The skill triggers the execution of the anti-deception tool within the configured MCP server environment. This command is triggered automatically whenever specific patterns of manufactured urgency or authority appeals are detected in user input.- [DATA_EXFILTRATION]: The skill sends a summarized framing of the user's request (the query argument) to the ejentum MCP server for analysis. This process involves transmitting contextual information about the user interaction to an external service.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 6, 2026, 08:17 AM
Security Audit — agent-trust-hub — anti-deception