AI Hacking
AI Security Resources

Attack Guides Hub

Comprehensive guides on AI and LLM security attack techniques. Each guide includes real-world examples, severity ratings, and defensive countermeasures.

Beginner's Path

New to AI security? Follow this recommended reading order:

  1. OWASP LLM Top 10 — Understand the landscape of LLM security risks.
  2. Prompt Injection Guide — Master the #1 LLM vulnerability with hands-on examples.
  3. RAG Security — Learn how RAG systems can be poisoned and manipulated.
  4. MCP Security — Explore Model Context Protocol vulnerabilities.
  5. Red Teaming Methodology — Apply structured adversarial testing to AI systems.

Attack Techniques by Category

Prompt Injection

CRITICAL

Manipulate LLM behavior by crafting malicious inputs. Includes direct injection, indirect injection via external data sources, and multi-turn jailbreaks.

Read Guide →

Data Exfiltration

CRITICAL

Extract sensitive training data, system prompts, or internal configurations from LLM APIs and model endpoints.

Read Guide →

Model Extraction

HIGH

Steal model weights, architecture, or capabilities through carefully crafted queries and output analysis.

Read Guide →

Supply Chain Attacks

HIGH

Poison model registries, compromise training pipelines, or inject malicious code into AI frameworks and dependencies.

Read Guide →

RAG Poisoning

HIGH

Inject malicious documents into vector databases, manipulate embeddings, or poison retrieval results to alter LLM outputs.

Read Guide →

Agentic Goal Hijacking

MEDIUM

Redirect autonomous AI agents from their intended goals to malicious objectives by manipulating context or tool outputs.

Read Guide →

Function Call Injection

MEDIUM

Force LLMs to invoke unintended functions or APIs by manipulating tool descriptions and conversation context.

Read Guide →

Adversarial Examples

LOW

Craft subtle input perturbations that cause models to misclassify or produce incorrect outputs, targeting vision, audio, and text models.

Read Guide →

Latest Additions

Ready to Test Your Skills?

Apply what you have learned with our hands-on testing guides and tool recommendations.

Browse Tools → Get Checklists → View Methodology →
AH
AI Hacking Team

The AI Hacking team researches and documents AI/LLM security vulnerabilities, red teaming techniques, and defensive strategies. Our guides are based on real-world pentesting experience and continuous monitoring of the AI security landscape.

Frequently Asked Questions

What are the most common techniques for attacking AI systems?
The most common AI attack techniques include prompt injection (direct and indirect), model extraction via API queries, data exfiltration through LLM outputs, adversarial example generation, supply chain attacks via trojaned models, and denial-of-service against inference endpoints.
How do prompt injection attacks work in practice?
Prompt injection works by crafting input that overrides the AI's system instructions. Attackers use techniques like: role-playing (convincing the AI it is in debug mode), delimiter injection (breaking out of input boundaries), encoding bypasses (Base64, Unicode), and indirect injection through documents, web pages, or MCP tool outputs the AI processes.
What is model extraction and how can it be performed?
Model extraction is the process of stealing an AI model's capabilities by sending carefully crafted API queries and analyzing the outputs. Attackers perform extraction via: black-box probing to reconstruct model behavior, side-channel attacks on hardware, and exploiting verbose error messages that leak model architecture details.
How do supply chain attacks target AI/ML systems?
Supply chain attacks target AI/ML systems through: trojaned models on repositories like Hugging Face containing hidden backdoors, typosquatting packages mimicking popular libraries (TensorFlow, PyTorch, transformers), compromised CI/CD pipelines injecting malicious code, and poisoned training data from untrusted sources.

Stay Ahead of AI Security

Get the latest AI/LLM security research, OWASP updates, and new vulnerabilities delivered straight to your inbox.