learn-datalake

Warn

Audited by Socket on Aug 26, 2026

9 alerts found:

Anomalyx5Securityx4
AnomalyLOW
learn_datalake.py

The fragment appears to implement a legitimate PDF ingestion and review orchestrator. It contains no clear malware indicators or direct credential theft/exfiltration. It does contain a meaningful command-injection risk because the memory_scope_pdf value is interpolated into bash -lc commands without safe argument passing or quoting. The code is incomplete and depends heavily on imported modules and external scripts, so those components require separate review.

Confidence: 98%Severity: 62%
AnomalyLOW
supervise_learn_datalake_helpers.py

The supplied fragment appears to be a workflow supervisor rather than malware. It contains legitimate-looking local state management, task monitoring, and memory-learning integration. The primary security concern is intentional arbitrary shell execution through bash -lc in _run_alert_hook() and _run_shell_command(); these functions are dangerous if their command parameters are derived from untrusted input. The outbound memory-service POST may disclose workflow metadata, but the default destination is local and no credential harvesting or clear malicious behavior is present. Review the referenced run.sh files and callers before deployment.

Confidence: 93%Severity: 62%
SecurityMEDIUM
extractability_model.py

The code appears to be a legitimate PDF-extractability training and prediction utility, with no clear malicious behavior. It has a significant command-injection risk because a user-controlled PDF path is embedded in a bash command, and a conditional arbitrary-code-execution risk from untrusted pickle deserialization. The subprocess should use an argument list without a shell, and model files should be integrity-protected or replaced with a safe serialization format.

Confidence: 98%Severity: 72%
SecurityMEDIUM
SKILL.md

SUSPICIOUS: the stated purpose broadly matches corpus-learning orchestration, but the skill's true execution boundary is opaque because it delegates to many unprovided local scripts and composed skills. No explicit credential theft or malicious exfiltration is shown, yet the unverifiable internal toolchain, optional fetch workflows, and autonomous long-running behavior make the overall security risk medium-high.

Confidence: 79%Severity: 74%
AnomalyLOW
api.py

The fragment appears to be a legitimate local monitoring and PDF-viewer service, with no clear malicious behavior or supply-chain backdoor. The main security concerns are unauthenticated binding to all network interfaces and insufficient validation of PDF paths read from local quarantine records. These issues warrant review and hardening, but the code provides no evidence of intentional malware.

Confidence: 97%Severity: 55%
AnomalyLOW
worker_pool.py

The fragment appears to be legitimate orchestration code for parallel PDF review. It contains no clear malware indicators or intentional data theft. The principal security issue is unsafe shell command construction: configuration values are interpolated into commands without robust shell escaping, so an attacker who can control cfg.root, cfg.memory_scope, or cfg.taxonomy_collection could potentially execute arbitrary shell commands. Use subprocess argument arrays or shlex.quote for every interpolated value, and validate paths and option values. The displayed fragment also appears syntactically incomplete at the end.

Confidence: 98%Severity: 68%
SecurityMEDIUM
subprocess_exec.py

The code is a readable command-execution and watchdog utility, not apparent malware. Its main security issue is shell command injection and unrestricted local command execution when cmd is influenced by untrusted data. Use argument lists without a shell, strict command allowlisting, or robust validation for externally influenced commands. Logging command strings and output may also leak secrets. No direct exfiltration, persistence, credential theft, or obfuscated payload is present.

Confidence: 98%Severity: 78%
AnomalyLOW
monitor_supervisor.sh

The fragment is a legitimate-looking local watchdog and status-bridge script. It shows no direct malware indicators such as network exfiltration, credential theft, cryptomining, reverse shells, or destructive actions. Security concerns are primarily unsafe path construction from LABEL and REPORT_PATH, disclosure of arbitrary files through the state-controlled run_log field, and potential Python code injection caused by interpolating shell values into python3 -c source. These risks depend on untrusted invocation or tampering with local state/configuration.

Confidence: 96%Severity: 55%
SecurityMEDIUM
viewer/vite.config.ts

The code appears to be a project-specific Vite development configuration, not malware. It does expose local corpus data over the network with wildcard CORS and contains a path containment weakness: startsWith() can accept sibling paths sharing the root prefix, and symlink escapes are possible. Binding to 0.0.0.0 makes the exposure reachable by other network clients. Use a path-boundary check such as path.relative(), reject absolute/traversal paths, resolve realpaths, restrict host access, and avoid wildcard CORS unless required.

Confidence: 97%Severity: 72%
Audit Metadata
Analyzed At
Aug 26, 2026, 06:05 PM
Package URL
pkg:socket/skills-sh/grahama1970%2Fagent-skills%2Flearn-datalake%2F@c4d8a791fd2ca1ab611007e579ac8174dd0438b5
Security Audit — socket — learn-datalake