long-context
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill provides instructions to download specific research implementations and datasets necessary for context extension.
- Fetches the YaRN implementation from a public GitHub repository (
jquesnelle/yarn). - Downloads standard text datasets from HuggingFace Hub (e.g., PG-19, arXiv, UltraChat).
- Installs common machine learning libraries including
transformers,torch, andeinopsvia standard package managers. - [INDIRECT_PROMPT_INJECTION]: The skill's fine-tuning and evaluation workflows involve processing external datasets which could theoretically contain adversarial content.
- Ingestion points: Data loading functions in
references/fine_tuning.mdthat ingest external document text. - Boundary markers: Includes data validation routines (
validate_training_data) to check for repetition and truncation artifacts. - Capability inventory: The ingested data is used for model training and metric calculation (perplexity, retrieval accuracy).
- Sanitization: Content is passed through standard model tokenizers with specific truncation and formatting logic.
Audit Metadata