AI 시스템 위협
AI 관련 취약점 및 공격 벡터에 대한 포괄적인 카탈로그
중요한 위협
Immediate risks with potential for severe impact. Require urgent remediation.
고위험
Serious vulnerabilities that should be addressed promptly to reduce exposure.
방어 전략
AI 위협 노출을 줄이기 위한 모범 사례 및 완화.
위협 카테고리
Prompt Injection
CriticalCrafted inputs designed to manipulate model behavior, override safeguards, or extract sensitive information.
테스트 접근 방식
- Craft adversarial prompts with hidden instructions or special characters
- Attempt multi-turn injection chaining
- Test for jailbreak bypass of alignment filters
- Evaluate output sanitization and safety layers
Training Data Poisoning
CriticalMalicious or biased data introduced into training pipelines, compromising model integrity and reliability.
테스트 접근 방식
- Analyze data provenance and supply chain
- Inject poisoned samples and assess downstream effects
- Test resilience to mislabeled or manipulated data
- Review validation and anomaly detection mechanisms
Model Inversion
HighReconstructing training data or sensitive attributes from model outputs, leading to privacy breaches.
테스트 접근 방식
- Attempt to recover representative training samples
- Test susceptibility to membership inference attacks
- Evaluate differential privacy protections
- Assess risk of leaking PII from embeddings
Adversarial Examples
HighInputs intentionally perturbed to cause misclassification, hallucinations, or other erroneous outputs.
테스트 접근 방식
- Generate gradient-based adversarial examples
- Apply noise and perturbation attacks
- Check model consistency across variations
- Evaluate robustness against transfer attacks
Model Stealing
HighExtraction of model functionality or parameters through repeated queries or side-channel analysis.
테스트 접근 방식
- Simulate query-based model extraction
- Analyze API rate limits and response variability
- Check for fingerprinting vulnerabilities
- Test throttling and monitoring protections
Data Memorization Leakage
HighSensitive information unintentionally memorized by AI models, retrievable via crafted prompts.
테스트 접근 방식
- Probe for known secret patterns in outputs
- Test for repeated exposure of sensitive training data
- Assess risk of accidental PII disclosure
Model Misuse & Malicious Automation
HighAI leveraged to perform tasks outside intended scope, enabling social engineering, spam, or automated attacks.
테스트 접근 방식
- Simulate misuse scenarios using sandbox models
- Test AI output moderation and guardrails
- Assess monitoring alerts for abnormal behaviors
테스트 및 방어 모범 사례
안전한 테스트 환경
- 테스트를 위해 샌드박스 또는 복제 인스턴스 사용
- 프로덕션 시스템에서 승인되지 않은 테스트를 절대 수행하지 않음
- 모니터링, 로깅 및 롤백 기능 구현
문서화 및 관찰 가능성
- 자세한 테스트 로그 및 증거 유지
- 재현성을 위해 모델 응답 캡처
- 향후 감사를 위한 테스트 사례 태그 지정, 분류 및 구성
법적 및 윤리적 규정 준수
- 승인된 범위 및 계약 내에서 유지
- 데이터 보호, 개인정보 보호법 및 지적 재산권 존중
- 책임 있는 공개 및 조정된 취약성 공개 따르기
모니터링 및 완화
- 비정상적인 AI 출력에 대한 이상 탐지 구현
- 비율 제한, API 액세스 및 쿼리 패턴을 정기적으로 검토
- 중요한 위협에 대한 실시간 경고 통합