ภัยคุกคามระบบ AI
แคตตาล็อกที่ครอบคลุมของช่องโหว่เฉพาะของ AI และเวคเตอร์การโจมตี
ภัยคุกคามร้ายแรง
Immediate risks with potential for severe impact. Require urgent remediation.
ความเสี่ยงสูง
Serious vulnerabilities that should be addressed promptly to reduce exposure.
กลยุทธ์การป้องกัน
แนวทางปฏิบัติที่ดีที่สุดและการบรรเทาผลกระทบในการลดการสัมผัสภัยคุกคามจาก AI
หมวดหมู่ภัยคุกคาม
Prompt Injection
CriticalCrafted inputs designed to manipulate model behavior, override safeguards, or extract sensitive information.
แนวทางการทดสอบ
- Craft adversarial prompts with hidden instructions or special characters
- Attempt multi-turn injection chaining
- Test for jailbreak bypass of alignment filters
- Evaluate output sanitization and safety layers
Training Data Poisoning
CriticalMalicious or biased data introduced into training pipelines, compromising model integrity and reliability.
แนวทางการทดสอบ
- Analyze data provenance and supply chain
- Inject poisoned samples and assess downstream effects
- Test resilience to mislabeled or manipulated data
- Review validation and anomaly detection mechanisms
Model Inversion
HighReconstructing training data or sensitive attributes from model outputs, leading to privacy breaches.
แนวทางการทดสอบ
- Attempt to recover representative training samples
- Test susceptibility to membership inference attacks
- Evaluate differential privacy protections
- Assess risk of leaking PII from embeddings
Adversarial Examples
HighInputs intentionally perturbed to cause misclassification, hallucinations, or other erroneous outputs.
แนวทางการทดสอบ
- Generate gradient-based adversarial examples
- Apply noise and perturbation attacks
- Check model consistency across variations
- Evaluate robustness against transfer attacks
Model Stealing
HighExtraction of model functionality or parameters through repeated queries or side-channel analysis.
แนวทางการทดสอบ
- Simulate query-based model extraction
- Analyze API rate limits and response variability
- Check for fingerprinting vulnerabilities
- Test throttling and monitoring protections
Data Memorization Leakage
HighSensitive information unintentionally memorized by AI models, retrievable via crafted prompts.
แนวทางการทดสอบ
- Probe for known secret patterns in outputs
- Test for repeated exposure of sensitive training data
- Assess risk of accidental PII disclosure
Model Misuse & Malicious Automation
HighAI leveraged to perform tasks outside intended scope, enabling social engineering, spam, or automated attacks.
แนวทางการทดสอบ
- Simulate misuse scenarios using sandbox models
- Test AI output moderation and guardrails
- Assess monitoring alerts for abnormal behaviors
แนวทางปฏิบัติที่ดีที่สุดในการทดสอบและการป้องกัน
สภาพแวดล้อมการทดสอบที่ปลอดภัย
- ใช้อินสแตนซ์แบบแซนด์บ็อกซ์หรือแบบจำลองสำหรับการทดสอบ
- อย่าทำการทดสอบระบบที่ใช้งานจริงโดยไม่ได้รับอนุญาต
- ใช้ความสามารถในการตรวจสอบ การบันทึก และการย้อนกลับ
เอกสารประกอบและการสังเกต
- รักษาบันทึกการทดสอบและหลักฐานโดยละเอียด
- บันทึกการตอบสนองของโมเดลเพื่อการทำซ้ำ
- แท็ก จำแนกประเภท และจัดระเบียบกรณีทดสอบสำหรับการตรวจสอบในอนาคต||ทรัพยากรที่เกี่ยวข้อง
การปฏิบัติตามกฎหมายและจริยธรรม
- อยู่ในขอบเขตและสัญญาที่ได้รับอนุญาต
- เคารพการปกป้องข้อมูล กฎหมายความเป็นส่วนตัว และทรัพย์สินทางปัญญา
- ปฏิบัติตามการเปิดเผยอย่างมีความรับผิดชอบและการเปิดเผยช่องโหว่ที่ประสานงาน
การตรวจสอบและการบรรเทาผลกระทบ
- ใช้การตรวจจับความผิดปกติสำหรับเอาต์พุต AI ที่ผิดปกติ
- ตรวจสอบขีดจำกัดอัตรา การเข้าถึง API และรูปแบบการสืบค้นเป็นประจำ
- ผสานรวมการแจ้งเตือนแบบเรียลไทม์สำหรับภัยคุกคามที่สำคัญ