huggingface-tokenizers
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill provides instructions for installing official packages (
tokenizers,transformers,datasets) via the standard Python package manager (pip). It also references model and dataset retrieval from HuggingFace Hub, which is a recognized and trusted service in the AI ecosystem. - [DATA_EXPOSURE]: Code examples demonstrate standard local file operations, such as reading text files for training (
tokenizer.train) and saving vocabulary files to the local disk (tokenizer.save). No sensitive file access or credential exposure patterns were found. - [INDIRECT_PROMPT_INJECTION]: The skill involves processing external text data for tokenization, which is an inherent feature of NLP tools and represents a potential surface for indirect prompt injection if the resulting tokens are processed by an LLM without context boundaries.
- Ingestion points: Text inputs for
tokenizer.encode()andtokenizer.train()as seen inSKILL.mdandreferences/training.md. - Boundary markers: Not applicable to the low-level library usage examples provided.
- Capability inventory: The skill facilitates file system writes (
tokenizer.save) and network access to trusted repositories for model loading. - Sanitization: The skill describes the library's normalization pipeline (e.g.,
NFD,Lowercase,StripAccents), which standardizes input but is not a security-focused sanitization mechanism.
Audit Metadata