nanogpt
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill fetches the TinyShakespeare educational dataset from a well-known GitHub repository (
raw.githubusercontent.com/karpathy/char-rnn). This is standard behavior for language modeling tutorials and utilizes a reputable source. - [DYNAMIC_EXECUTION]: The training scripts use
pickle.loadandtorch.loadto manage metadata and model checkpoints. While these serialization methods involve runtime execution, they are the standard mechanism in the PyTorch ecosystem for preserving and restoring training state. - [INDIRECT_PROMPT_INJECTION]: The skill ingests untrusted text data for training purposes, which is a standard surface for data poisoning in machine learning.
- Ingestion points: Data is read from remote URLs and local files in
data/shakespeare_char/prepare.pyanddata/custom/prepare.py. - Boundary markers: No specific boundary markers are employed to isolate text content, as the model is designed to learn directly from the raw data stream.
- Capability inventory: The skill can write files (
tofile,torch.save), perform network operations for logging (wandb), and execute training commands viatorchrun. - Sanitization: Input text is tokenized into numeric identifiers without filtering for instruction-like patterns, which is the intended behavior for training generative language models.
Audit Metadata