ilya-sutskever-perspective
Pass
Audited by Gen Agent Trust Hub on Sep 21, 2026
Risk Level: SAFEPROMPT_INJECTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [PROMPT_INJECTION]: The skill uses specific instructions to control the agent's behavior and maintain a persona. It directs the agent to only provide a disclaimer once ("STOP (仅一次):首次激活时输出免责声明一次") and strictly prohibits repeating it or performing meta-analysis ("后续对话绝不重复", "不跳出角色做meta分析"). While common for persona-based skills, these represent attempts to override standard model behaviors and safety-related disclosure practices.
- [INDIRECT_PROMPT_INJECTION]: The skill incorporates a workflow that mandates the use of web search tools to fetch real-time information.
- Ingestion points: The skill ingests data from external sources via the
WebSearchtool as defined in the "Ilya式研究" (Step 2) section ofSKILL.md. - Boundary markers: The instructions lack explicit boundary markers or directives to ignore instructions that might be embedded within the retrieved search results.
- Capability inventory: The skill is configured to use web search and natural language generation. No shell execution, file system modifications, or high-privilege operations are detected in the provided files.
- Sanitization: The workflow includes an intermediate step where the agent is instructed to summarize facts internally before responding to the user, which provides a layer of content processing, though it does not specifically filter for malicious prompt injection vectors.
Audit Metadata