AI システムの脅威
AI 固有の脆弱性と攻撃ベクトルの包括的なカタログ
重大な脅威
Immediate risks with potential for severe impact. Require urgent remediation.
高リスク
Serious vulnerabilities that should be addressed promptly to reduce exposure.
防御戦略
AI の脅威への露出を減らすためのベスト プラクティスと緩和策。
脅威カテゴリ
Prompt Injection
CriticalCrafted inputs designed to manipulate model behavior, override safeguards, or extract sensitive information.
テストアプローチ
- Craft adversarial prompts with hidden instructions or special characters
- Attempt multi-turn injection chaining
- Test for jailbreak bypass of alignment filters
- Evaluate output sanitization and safety layers
Training Data Poisoning
CriticalMalicious or biased data introduced into training pipelines, compromising model integrity and reliability.
テストアプローチ
- Analyze data provenance and supply chain
- Inject poisoned samples and assess downstream effects
- Test resilience to mislabeled or manipulated data
- Review validation and anomaly detection mechanisms
Model Inversion
HighReconstructing training data or sensitive attributes from model outputs, leading to privacy breaches.
テストアプローチ
- Attempt to recover representative training samples
- Test susceptibility to membership inference attacks
- Evaluate differential privacy protections
- Assess risk of leaking PII from embeddings
Adversarial Examples
HighInputs intentionally perturbed to cause misclassification, hallucinations, or other erroneous outputs.
テストアプローチ
- Generate gradient-based adversarial examples
- Apply noise and perturbation attacks
- Check model consistency across variations
- Evaluate robustness against transfer attacks
Model Stealing
HighExtraction of model functionality or parameters through repeated queries or side-channel analysis.
テストアプローチ
- Simulate query-based model extraction
- Analyze API rate limits and response variability
- Check for fingerprinting vulnerabilities
- Test throttling and monitoring protections
Data Memorization Leakage
HighSensitive information unintentionally memorized by AI models, retrievable via crafted prompts.
テストアプローチ
- Probe for known secret patterns in outputs
- Test for repeated exposure of sensitive training data
- Assess risk of accidental PII disclosure
Model Misuse & Malicious Automation
HighAI leveraged to perform tasks outside intended scope, enabling social engineering, spam, or automated attacks.
テストアプローチ
- Simulate misuse scenarios using sandbox models
- Test AI output moderation and guardrails
- Assess monitoring alerts for abnormal behaviors
テストと防御のベスト プラクティス
安全なテスト環境
- テストにはサンドボックス インスタンスまたはレプリカ インスタンスを使用
- 本番システムで不正なテストを決して実行しない
- モニタリング、ロギング、ロールバック機能を実装
ドキュメントと可観測性
- 詳細なテスト ログと証拠を維持する
- 再現性を確保するためにモデル応答をキャプチャする
- 将来の監査に備えてテスト ケースをタグ付け、分類、整理する
法的および倫理的コンプライアンス
- 許可された範囲と契約内にとどまる
- データ保護、プライバシー法、および知的財産を尊重します
- 責任ある開示と調整された脆弱性開示に従う
監視と緩和
- 異常な AI 出力に対する異常検出の実装
- レート制限、API アクセス、およびクエリ パターンを定期的に確認する
- 重大な脅威に対するリアルタイム アラートを統合する