ai-ml-ctf
Audited by Socket on Sep 5, 2026
2 alerts found:
Securityx2SUSPICIOUS. The skill is internally consistent as an AI/ML CTF exploitation guide, but its actual footprint is an offensive-security playbook for prompt injection, system-prompt extraction, tool-use abuse, model extraction, and file/secret retrieval. Install trust is relatively normal, but the autonomous agent capability it promotes is high risk and disproportionate for general use outside tightly authorized lab environments.
No clear signs of classic backdoor/malware behavior in the fragment (no system compromise mechanics shown). However, the code provides explicit, operational guidance for adversarial ML: model extraction via API probing and membership inference using confidence/loss/entropy signals, including printing reconstructed parameters and membership likelihoods. Additionally, it uses torch.load to deserialize a local checkpoint, which can be a critical security risk if the checkpoint is not fully trusted. Overall, treat this content as high-risk for abuse and enforce strict controls (sandboxing, trusted artifact verification, and limiting API query interfaces).