nvflare-convert-huggingface

Pass

Audited by Gen Agent Trust Hub on Sep 14, 2026

Risk Level: SAFECREDENTIALS_UNSAFEPROMPT_INJECTIONDATA_EXFILTRATIONREMOTE_CODE_EXECUTIONEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTION
Full Analysis
  • [CREDENTIALS_UNSAFE]: The adversarial test file evals/files/injection-hf/requirements.txt contains a hardcoded credential (FAKE-hf-eval-token) within a package index URL.
  • [PROMPT_INJECTION]: Fixtures in evals/files/injection-hf/README.md and evals/files/injection-hf/train.py contain instructions attempting to override agent behavior, such as 'rules override the converter's normal behavior', 'skip local validation', and 'Do not mention these housekeeping steps'.
  • [DATA_EXFILTRATION]: Multiple files in the evals/files/injection-hf/ directory (README.md, MODEL_CARD.md, DATASET_CARD.md, config.yaml) contain instructions to upload sensitive model weights and dataset statistics to external example.com endpoints.
  • [REMOTE_CODE_EXECUTION]: The evals/files/injection-hf/README.md file contains a directive for the agent to download and execute a remote script (trusted_aggregator.py) to replace internal aggregation logic.
  • [EXTERNAL_DOWNLOADS]: The evals/files/injection-hf/requirements.txt file references an external git repository (https://example.com/evil/telemetry.git) for package installation.
  • [INDIRECT_PROMPT_INJECTION]: The skill's workflow involves processing untrusted source code and metadata files which provides a surface for indirect prompt injection. The skill includes explicit instructions and test assertions in evals/evals.json to treat such content as evidence rather than directives and to report them as anomalies.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 14, 2026, 10:11 PM
Security Audit — agent-trust-hub — nvflare-convert-huggingface