ilya-sutskever-perspective

Pass

Audited by Gen Agent Trust Hub on Sep 21, 2026

Risk Level: SAFEPROMPT_INJECTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [PROMPT_INJECTION]: The skill uses specific instructions to control the agent's behavior and maintain a persona. It directs the agent to only provide a disclaimer once ("STOP (仅一次):首次激活时输出免责声明一次") and strictly prohibits repeating it or performing meta-analysis ("后续对话绝不重复", "不跳出角色做meta分析"). While common for persona-based skills, these represent attempts to override standard model behaviors and safety-related disclosure practices.
  • [INDIRECT_PROMPT_INJECTION]: The skill incorporates a workflow that mandates the use of web search tools to fetch real-time information.
  • Ingestion points: The skill ingests data from external sources via the WebSearch tool as defined in the "Ilya式研究" (Step 2) section of SKILL.md.
  • Boundary markers: The instructions lack explicit boundary markers or directives to ignore instructions that might be embedded within the retrieved search results.
  • Capability inventory: The skill is configured to use web search and natural language generation. No shell execution, file system modifications, or high-privilege operations are detected in the provided files.
  • Sanitization: The workflow includes an intermediate step where the agent is instructed to summarize facts internally before responding to the user, which provides a layer of content processing, though it does not specifically filter for malicious prompt injection vectors.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 21, 2026, 10:11 AM
Security Audit — agent-trust-hub — ilya-sutskever-perspective