> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.lakera.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.lakera.ai/_mcp/server.

# AI Guardrails

Check Point AI Guardrails provides real-time visibility and security for GenAI applications. It does this through a combination of Check Point managed and custom guardrails.

Check Point offers guardrails for the following defenses:

* [Prompt Defense](/docs/prompt-defense) - Detect and respond to direct and indirect
  prompt attacks. This includes jailbreaks, prompt injections and any attempts to manipulate and exploit AI models through malicious or unintentionally
  troublesome instructions, preventing potential harm to your application.
* [Content Moderation](/docs/content-moderation) - Ensure your GenAI applications do not
  violate your organization's policies by detecting and stopping harmful and unwanted content.
* [Data Leakage Prevention](/docs/data-leakage-prevention) -
  Safeguard Personally Identifiable Information, prevent system prompt leakage, and avoid costly leakage of sensitive data, ensuring compliance with data protection and privacy regulations.
* [Malicious links detection](/docs/unknown-links) - Prevent attackers manipulating the LLM into displaying malicious or phishing links to your users by detecting unknown links.
* [Agent Behavior Defense](/docs/agent-behavior-defense) - Detect dangerous actions outside the agent's trusted mandate with the Dangerous Deviation detector, and control which tools an agent may call at runtime with the Tool Allow/Deny List.

## GenAI faces novel threats

Large Language Models and other generative AI technologies are introducing brand new
cyber security threats that existing cyber security tools can't address. The number of
potential attackers of LLMs is also massively larger than traditional software since you
don't need specialist technical skills, anyone who can write can exploit LLMs just using
natural language.

This means the attack surface of GenAIs is orders of magnitudes larger and fundamentally
different than traditional software and requires a paradigm shift in cyber security to
secure.

One of the major components of AI security is a **real-time AI application firewall**
for any applications using LLMs. This solution integrates into the application and
screens any user or reference contents passed into an LLM and the output response from
the LLM. Any threats detected can then be handled in real-time, blocking attackers and
preventing harm to end users, the application, and to the organization running the
application.

## Overview

When securing a GenAI application follow these general steps to set up defenses:

#### System prompt

Create a robust and securely written system prompt to ensure the AI behaves securely and as intended. For help see our guide [Crafting Secure System Prompts for LLM and GenAI Applications](https://www.lakera.ai/ai-security-guides/crafting-secure-system-prompts-for-llm-and-genai-applications).

#### Prompt defenses

Set up prompt defenses for all LLM inputs, directly from users and referenced content, as even trusted sources can contain inadvertent prompt attacks.

#### Additional guardrails

Add guardrail defenses to prevent dangerous or sensitive content in LLM inputs or outputs in real-time, such as [content moderation](./content-moderation), [data leakage prevention](./data-leakage-prevention) or [malicious link detection](./unknown-links).

## Check Point managed guardrails

Check Point managed guardrails use a combination of machine
learning and language models with rule-based filters to detect threats within the contents submitted
to AI Guardrails for screening.

Guardrails are designed to tackle specific types of threats. AI Guardrails can be
customized according to the threat profile of your application by setting the relevant
guardrails to use in screening in the [AI Guardrails policy](/docs/policies).

Check Point's guardrails are updated on daily basis to incorporate defenses against new attacks and reduce false positives. We offer the option to fine-tune Check Point guardrails based on customer data or targeted feedback.

For details on each of the guardrails, please see the documentation linked above for each defense.

> **Note**
>
> We are always actively improving our detectors and working to increase accuracy and
> reduce bias. We are also continuously improving our controls and interface to empower
> customers to effectively and easily secure their applications. If you experience any
> issues or would like to provide feedback, please reach out to [support@lakera.ai](mailto:support@lakera.ai).

## Custom guardrails

Customers can enforce their own bespoke security or content policies by using custom guardrails. You describe what to detect in plain language and a built-in AI assistant compiles it into a detection policy, combining pattern rules for structured formats with semantic rules for content and intent. Custom guardrails are available in beta; see the [Custom Guardrails documentation](./custom-guardrails).

## Fine-tuning guardrails

AI Guardrails' defenses can be customized within your [policies](/docs/policies) to make them
more or less aggressive in detecting potential threats. This is done via
threshold levels. These set the confidence level the detector needs to reach in order to
return the detection in the screened contents.

For example, if you had a high risk tolerance for one use case you can set a guardrail to
only return very high confidence detections in order to have low false positives. Or, if
you had a use case where you wanted to be really sure the LLM wasn't manipulated, even
at the cost of potential impact on user experience, you could set the guardrail to return
anything that the detector thinks could potentially be a detection.

AI Guardrails uses the following threshold levels, in line with
[OWASP's paranoia level definitions for WAFs](https://coreruleset.org/docs/2-how-crs-works/2-2-paranoia_levels/):

1. **L1** - Lenient, very few false positives, if any.
2. **L2** - Balanced, some false positives.
3. **L3** - Stricter, expect false positives but very low false negatives.
4. **L4** - Paranoid, higher false positives but very few false
   negatives, if any. This is our default confidence threshold.

Setting a guardrail to a threshold level in the policy means that the detector will return a positive decision
whenever it has that level of confidence, or higher, that the screened contents contain
a threat of that type.

The higher the threshold level the stricter the guardrail will be, reducing the
probability that a potential threat slips through but at the potential risk of higher
false positives during benign interactions.

Note that the threshold levels fine-tune the required confidence of the detector, **not
the severity** of the threat.

> **Warning**
>
> We would love any feedback on the threshold levels to make sure they're calibrated
> correctly and give you the control you need for your use cases. If you experience any
> issues or would like to provide feedback, please reach out to [support@lakera.ai](mailto:support@lakera.ai).

### Allow and Deny Lists

AI Guardrails also provides the ability to create custom **allow** and **deny** lists to temporarily override model decisions. This feature helps customers quickly address false positives or false negatives while waiting for model improvements. These lists are designed as a temporary measure for addressing urgent edge cases that impact critical workflows, not as a permanent security solution.

> **Warning**
>
> Overriding AI Guardrails' guardrails with custom lists can introduce security loopholes. We recommend using this feature only as a temporary measure while reporting misclassified prompts to Check Point for robust fixes.

For implementation details, see the [Allow and Deny Lists documentation](./allow-deny-lists).