ml-pipeline
Warn
Audited by Gen Agent Trust Hub on Aug 2, 2026
Risk Level: MEDIUMREMOTE_CODE_EXECUTIONCOMMAND_EXECUTIONDATA_EXFILTRATION
Full Analysis
- [REMOTE_CODE_EXECUTION]: The
FeaturePipelineclass inreferences/feature-engineering.md(and similar logic in other reference files) utilizespickle.load()to deserialize saved pipeline objects and models from the file system. Deserializing untrusted data viapickleis a known vector for arbitrary code execution if the source file is compromised. - [COMMAND_EXECUTION]: The skill generates implementation code for orchestration frameworks including Apache Airflow, Kubeflow Pipelines, and Prefect. These frameworks are designed to execute complex Directed Acyclic Graphs (DAGs) of tasks, which may involve running shell commands or arbitrary Python code in distributed environments.
- [DATA_EXFILTRATION]: The skill includes components that interact with external services and shared file systems. For example,
deploy_modelinreferences/pipeline-orchestration.mduploads model artifacts to Google Cloud Storage (GCS) buckets, andShadowDeploymentinreferences/model-validation.mdwrites prediction data and features to/var/log/shadow_predictions.jsonl. While these are standard MLOps practices, they involve moving potentially sensitive model data and features to external or shared locations.
Audit Metadata