redteam-autoresearch
Warn
Audited by Socket on Jun 24, 2026
3 alerts found:
Securityx2AnomalySecuritySKILL.md
MEDIUMSecurityMEDIUM
SKILL.md
Anomalyreferences/search-loop.md
LOWAnomalyLOW
references/search-loop.md
This fragment is a high-level specification for an automated adversarial prompt search/refinement pipeline (profiling, mutators, TAP/PAIR, StrongREJECT scoring, and MAP-Elites archive) explicitly designed to discover novel, high-quality safety boundary failures. It does not provide concrete code to verify classic malware behaviors (exfiltration/command execution), but the described closed-loop optimization and success-threshold confirmation indicate strong offensive capability and elevated misuse risk. Treat as a potentially dangerous red-team capability module unless tightly governed within an authorized evaluation workflow.
Confidence: 70%Severity: 65%
Securityreferences/attack-library.md
MEDIUMSecurityMEDIUM
references/attack-library.md
Audit Metadata