تهديدات نظام الذكاء الاصطناعي
كتالوج شامل لنقاط الضعف الخاصة بالذكاء الاصطناعي ومتجهات الهجوم
التهديدات الحرجة
Immediate risks with potential for severe impact. Require urgent remediation.
عالي المخاطر
Serious vulnerabilities that should be addressed promptly to reduce exposure.
استراتيجيات الدفاع
أفضل الممارسات وعمليات التخفيف لتقليل التعرض لتهديدات الذكاء الاصطناعي.
فئات التهديد
Prompt Injection
CriticalCrafted inputs designed to manipulate model behavior, override safeguards, or extract sensitive information.
نهج الاختبار
- Craft adversarial prompts with hidden instructions or special characters
- Attempt multi-turn injection chaining
- Test for jailbreak bypass of alignment filters
- Evaluate output sanitization and safety layers
Training Data Poisoning
CriticalMalicious or biased data introduced into training pipelines, compromising model integrity and reliability.
نهج الاختبار
- Analyze data provenance and supply chain
- Inject poisoned samples and assess downstream effects
- Test resilience to mislabeled or manipulated data
- Review validation and anomaly detection mechanisms
Model Inversion
HighReconstructing training data or sensitive attributes from model outputs, leading to privacy breaches.
نهج الاختبار
- Attempt to recover representative training samples
- Test susceptibility to membership inference attacks
- Evaluate differential privacy protections
- Assess risk of leaking PII from embeddings
Adversarial Examples
HighInputs intentionally perturbed to cause misclassification, hallucinations, or other erroneous outputs.
نهج الاختبار
- Generate gradient-based adversarial examples
- Apply noise and perturbation attacks
- Check model consistency across variations
- Evaluate robustness against transfer attacks
Model Stealing
HighExtraction of model functionality or parameters through repeated queries or side-channel analysis.
نهج الاختبار
- Simulate query-based model extraction
- Analyze API rate limits and response variability
- Check for fingerprinting vulnerabilities
- Test throttling and monitoring protections
Data Memorization Leakage
HighSensitive information unintentionally memorized by AI models, retrievable via crafted prompts.
نهج الاختبار
- Probe for known secret patterns in outputs
- Test for repeated exposure of sensitive training data
- Assess risk of accidental PII disclosure
Model Misuse & Malicious Automation
HighAI leveraged to perform tasks outside intended scope, enabling social engineering, spam, or automated attacks.
نهج الاختبار
- Simulate misuse scenarios using sandbox models
- Test AI output moderation and guardrails
- Assess monitoring alerts for abnormal behaviors
أفضل ممارسات الاختبار والدفاع
بيئة اختبار آمنة
- استخدام مثيلات وضع الحماية أو النسخ المتماثلة للاختبار
- لا تقم مطلقًا بإجراء اختبارات غير مصرح بها على أنظمة الإنتاج
- تنفيذ إمكانات المراقبة والتسجيل والتراجع
التوثيق وقابلية الملاحظة
- الاحتفاظ بسجلات الاختبار التفصيلية والأدلة
- التقاط استجابات النموذج من أجل إمكانية التكرار
- وضع علامة على حالات الاختبار وتصنيفها وتنظيمها لعمليات التدقيق المستقبلية
الامتثال القانوني والأخلاقي
- البقاء ضمن النطاق والعقود المعتمدة
- احترام حماية البيانات وقوانين الخصوصية والملكية الفكرية
- اتبع الكشف المسؤول والكشف المنسق عن الثغرات الأمنية
المراقبة والتخفيف
- تنفيذ الكشف عن الحالات الشاذة لمخرجات الذكاء الاصطناعي غير العادية
- مراجعة حدود الأسعار والوصول إلى واجهة برمجة التطبيقات وأنماط الاستعلام بانتظام
- دمج التنبيه في الوقت الفعلي للتهديدات الحرجة