detecting-model-extraction-attacks
mukul975/Anthropic-Cybersecurity-Skills
This skill provides comprehensive methods and frameworks to detect sophisticated AI attacks, including model extraction, membership inference, and model inversion, which abuse inference APIs. It focuses on monitoring per-principal query volume, input distribution, and confidence exposure. Use it for securing public or partner inference APIs, or for conducting pre-deployment red-teaming to assess the model's extractability and overall security posture against MITRE ATLAS techniques.