taxonomy
Audited by Socket on Mar 17, 2026
2 alerts found:
AnomalyObfuscated FileThe code implements a functional taxonomy extraction pipeline with dual paths (keywords and LLM). However, a critical syntax issue (EXTRACTION_PROMPT definition) must be resolved before deployment. Additional concerns include reliance on dynamic imports, external scripts, and environment-sourced credentials that could leak data to an LLM service. Remediation should prioritize: fix the syntax error, audit and constrain dynamic imports, isolate or sandbox memory interactions, harden environment-variable handling, and validate data flows to external services to minimize exposure. Overall security risk remains medium until those concerns are addressed.
The module is a standard ML training script that legitimately executes a local helper script to fetch labeled data and sends plaintext to a configurable embedding service. I found no evidence of intentionally embedded backdoors or obfuscated malicious payloads in this Python file. However, two operational supply-chain/privacy risks stand out: (1) executing a local run.sh (MEMORY_RUN) presents an arbitrary-code-execution vector if that script can be modified, and (2) embedding sends unredacted text to an HTTP endpoint that can be redirected to exfiltrate sensitive data. Recommended mitigations: audit and lock down the run.sh file (verify provenance, use strict file permissions), restrict or validate embedding_url (use localhost or authenticated TLS endpoints only), redact or filter sensitive fields before embedding, fail fast on embedding errors rather than silently substituting zeros, and run training in a restricted environment (container/VM) with least privilege.