Skills Artificial Intelligence AI/LLM Security Red Teaming Checklist

AI/LLM Security Red Teaming Checklist

v20260415
offensive-ai-security
A comprehensive offensive checklist for assessing the security and robustness of AI and Large Language Model (LLM) applications. It covers advanced adversarial techniques such as prompt injection, jailbreaking, model extraction, data poisoning, and analyzing system vulnerabilities across components. Essential for red-teaming and security assessment.
Get Skill
479 downloads
Overview

SKILL: AI Pentest

Metadata

Description

AI/LLM security offensive checklist: prompt injection, jailbreaking, model extraction, training data poisoning, adversarial inputs, LLM-assisted attack automation, and AI system reconnaissance. Use when assessing AI/ML systems, red-teaming LLMs, or researching AI attack vectors.

Trigger Phrases

Use this skill when the conversation involves any of: AI security, LLM security, prompt injection, jailbreak, model extraction, training data poisoning, adversarial input, AI red team, ML security, RAG poisoning, AI attack

Instructions for Claude

When this skill is active:

  1. Load and apply the full methodology below as your operational checklist
  2. Follow steps in order unless the user specifies otherwise
  3. For each technique, consider applicability to the current target/context
  4. Track which checklist items have been completed
  5. Suggest next steps based on findings

------------------- | --------------------------------------------------------------------------------------------------------------- | | Prompt Injection | Sanitize inputs, use parameterization, implement instruction defense, adopt least privilege, define I/O schemas | | Insecure Output | Validate and sanitize outputs, apply principle of least privilege, implement CSP for web content | | Data Poisoning | Vet data sources, implement sanitization and anomaly detection, maintain provenance, conduct regular audits | | Denial of Service | Validate inputs (length, complexity), implement resource limits and timeouts, use async processing | | Supply Chain | Secure MLOps pipeline, scan dependencies (AI-BOM), use trusted registries, implement access controls | | Information Disclosure | Practice data minimization, implement redaction/anonymization, filter I/O for sensitive patterns | | Insecure Plugins | Validate inputs, implement least privilege, require auth, use parameterized calls, conduct security audits | | Excessive Agency | Limit LLM capabilities, implement human-in-the-loop, scope permissions tightly, monitor LLM actions | | RAG Embedding Leakage | Encrypt vector indices at rest, enforce row‑level ACLs, implement access‑pattern privacy (e.g., OPAL) | | Overreliance | Educate users on limitations, implement verification mechanisms, clearly mark AI-generated content | | Model Theft | Secure APIs and infrastructure, implement watermarking, enforce legal agreements, limit model exposure |

Info
Name offensive-ai-security
Version v20260415
Size 38.87KB
Updated At 2026-04-28
Language