sentencepiece
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONPRIVILEGE_ESCALATIONEXTERNAL_DOWNLOADS
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill provides instructions for training tokenizers on external datasets, which introduces a surface for indirect prompt injection.
- Ingestion points: The skill ingests raw text through file paths (e.g.,
data.txt) or Python iterators passed to the training functions (SKILL.md, references/training.md). - Boundary markers: There are no specific instructions or delimiters used to separate the training data from the agent's operating instructions.
- Capability inventory: The skill utilizes the
sentencepieceandtransformerslibraries to read training corpora and write model and vocabulary files to the local file system. - Sanitization: The instructions do not include steps for sanitizing or filtering the content of the training data before processing.
- [PRIVILEGE_ESCALATION]: The manual C++ installation section includes commands to install the compiled library system-wide using elevated privileges.
- Evidence: The use of
sudo make installin the build workflow inSKILL.md. - [EXTERNAL_DOWNLOADS]: The skill fetches the project source code from Google's official GitHub repository during the manual installation process.
- Evidence:
git clone https://github.com/google/sentencepiece.gitinSKILL.md.
Audit Metadata