sentencepiece

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONPRIVILEGE_ESCALATIONEXTERNAL_DOWNLOADS
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill provides instructions for training tokenizers on external datasets, which introduces a surface for indirect prompt injection.
  • Ingestion points: The skill ingests raw text through file paths (e.g., data.txt) or Python iterators passed to the training functions (SKILL.md, references/training.md).
  • Boundary markers: There are no specific instructions or delimiters used to separate the training data from the agent's operating instructions.
  • Capability inventory: The skill utilizes the sentencepiece and transformers libraries to read training corpora and write model and vocabulary files to the local file system.
  • Sanitization: The instructions do not include steps for sanitizing or filtering the content of the training data before processing.
  • [PRIVILEGE_ESCALATION]: The manual C++ installation section includes commands to install the compiled library system-wide using elevated privileges.
  • Evidence: The use of sudo make install in the build workflow in SKILL.md.
  • [EXTERNAL_DOWNLOADS]: The skill fetches the project source code from Google's official GitHub repository during the manual installation process.
  • Evidence: git clone https://github.com/google/sentencepiece.git in SKILL.md.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — sentencepiece