huggingface-tokenizers

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFEEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTION
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill provides instructions for installing official packages (tokenizers, transformers, datasets) via the standard Python package manager (pip). It also references model and dataset retrieval from HuggingFace Hub, which is a recognized and trusted service in the AI ecosystem.
  • [DATA_EXPOSURE]: Code examples demonstrate standard local file operations, such as reading text files for training (tokenizer.train) and saving vocabulary files to the local disk (tokenizer.save). No sensitive file access or credential exposure patterns were found.
  • [INDIRECT_PROMPT_INJECTION]: The skill involves processing external text data for tokenization, which is an inherent feature of NLP tools and represents a potential surface for indirect prompt injection if the resulting tokens are processed by an LLM without context boundaries.
  • Ingestion points: Text inputs for tokenizer.encode() and tokenizer.train() as seen in SKILL.md and references/training.md.
  • Boundary markers: Not applicable to the low-level library usage examples provided.
  • Capability inventory: The skill facilitates file system writes (tokenizer.save) and network access to trusted repositories for model loading.
  • Sanitization: The skill describes the library's normalization pipeline (e.g., NFD, Lowercase, StripAccents), which standardizes input but is not a security-focused sanitization mechanism.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — huggingface-tokenizers