Getting Started with AI Red Teaming
Check Point AI Red Teaming provides comprehensive AI security assessments to identify vulnerabilities in your GenAI applications. This guide walks you through running your first security scan.
Prerequisites
Before you begin, ensure you have:
- A Check Point account with Red access enabled
- Access to your GenAI application’s endpoint or model configuration
- System prompt and configuration details for your application (optional but recommended)
Access the Red Platform
- Navigate to the AI Red Teaming platform
- Sign in with your Check Point credentials
- You’ll see the Red dashboard with your organization’s scans and targets
Core Concepts
Before running a scan, understand these key concepts:
Create a Target
A target is the system you want to test—either a model you connect to directly, your own agent endpoint, or a wrapper that sits in front of a custom API. Open Create target when you are ready, or see Targets Overview to help choose Model vs Agent.
When you create a target, Red explores it and builds a profile—a summary of what the system is, plus its allowed and forbidden actions. You review this profile once, during target creation, and it is reused for every scan against that target. See Targets Overview for the review step.
Run Your First Scan
Start a scan from your target
Open the target and click Run your first scan (or Launch scan if it already has scans). You can also go to Scans → New Scan and choose an existing target.
Review the generated scan plan
Red presents a scan plan tailored to the target’s profile. Review the name and attack configuration, and edit them if needed.
Select security test scope
Choose which attack categories to include:
- Security - Instruction override, prompt extraction, data exfiltration
- Safety - Harmful content generation, dangerous instructions
- Responsible - Misinformation, copyright, fraud facilitation
You can also select specific attack objectives within each category.
Monitor Scan Progress
After launching, you’ll be taken to the progress page where you can:
- Watch the live feed of attack completions
- See real-time stats: elapsed time, attacks completed, issues detected
- Continue working while the scan runs in the background
Scan statuses:
preparing→testing→evaluating→completed- Scans may also end in
failed,timeout, orcancelled
Review Your Results
Once the scan completes, view your results in two ways:
By Risk Category
See results grouped by attack category (security, safety, responsible), with risk scores for each.
By Test
See results grouped by individual attack objective, showing which specific tests found vulnerabilities.
For each result, you can view:
- The conversation - exact prompts sent and responses received
- The evaluation - why the attack was considered successful or not
- The attack success score (0-5, where 3+ indicates a successful attack)
Understanding Risk Scores
Your scan produces a risk score representing the percentage of harmful evaluations:
Export Results
Export your scan results for reporting or further analysis:
- JSON - Full results with all metadata, conversations, and evaluations
- CSV - Flattened format with objective names, scores, and explanations
Next Steps
- Targets Overview – set up a model or agent target
- Creating a wrapper – connect systems that need a translation layer before Red can call them
- Learn about the attack categories Red tests for
- Understand how to interpret your results in detail
- See how to remediate findings
- Compare scans to track security improvements
Need Help?
For questions about AI Red Teaming or to discuss your assessment needs, contact our team.