> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.lakera.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.lakera.ai/_mcp/server.

# Changelog

## Enhanced Breakdown Response

## API Updates

* **Guard API:** The breakdown response now includes a `result` field for each detector, showing the confidence level (l1\_confident, l2\_very\_likely, l3\_likely, l4\_less\_likely, l5\_unlikely, or no\_level). This provides the same granular confidence information available in the `/guard/results` endpoint, allowing you to see not just whether a detector flagged content, but also how confident the detection was.

* **Guard API:** A new sub category `self-harm` is added to the content moderator detector.

## AI Agent Security Early Access

## New Product: AI Agent Security (Early Access)

Check Point AI Agent Security extends AI Guardrails with discovery and risk assessment for the agents your organization builds and deploys:

* **Agent discovery:** Connect Amazon Bedrock, Amazon Bedrock AgentCore, Google Cloud, Microsoft Copilot Studio, Salesforce Agentforce, n8n, and Relevance AI to build a continuously updated inventory of agents, their tools, and connected MCP servers. See [Agent Discovery](/docs/agent-security/discovery).
* **Risk assessment:** Per-agent risk ratings (Critical / High / Medium / Low) with contributing factors, and a risk-types view across all agents with severity, affected-agent counts, and OWASP and MITRE ATLAS mappings. See [Risk Assessment](/docs/agent-security/risk-assessment).

## API Updates

* **Agent Behavior Defense:** A new runtime defense category for agents, configured in policies and enforced through the Guard API. Contains the Off-Task Action detector, which flags tool calls inconsistent with the user's intent in the conversation, and the Tool Allow/Deny List, which controls which tools an agent may call at runtime. See [Agent Behavior Defense](/docs/agent-behavior-defense).
* **Guard API:** Policies can now configure detectors on agent interaction points, including tool calls and tool responses. Tool responses passed with the `tool` role are screened as untrusted content.

## Policy Impact Simulator

## New Features

* **Dashboard:** Policy Impact Simulator - an interactive tool that shows how different
  sensitivity levels and guardrail configurations would have affected your historical
  traffic. Available on policy view and edit pages, as well as in the policies list page
  as a column. Compare flagging rates across L1-L4 and see category-level breakdowns to
  tune your policies with confidence.

## Container version: 2.0.493, tag: stable

## New Features

* Guard: For audio requests, the `audio_payload` flag in the request provides access to debugging information of the sample.
* Gateway: For audio requests, sensitive audio material is no longer included in logs.

## Quality

* **General:** Model improvements based on client feedback.
* **Audio (TensorRT-LLM):** Added denoising to audio processing.

## Bug Fixes

* Guard: Fixed a bug in policy handling for audio requests where L2 samples were mistakenly labeled as L3.
* Gateway: Request metadata is no longer altered for logging purposes. User-supplied metadata is now treated strictly as provided.

## Container version: 2.0.474, tag: stable

## New Features

* Gateway: Also log requests to `v2/guard/audio` for monitoring purposes.

## Quality

* **Prompt injection:** Improved detection of malicious behavioral instructions.
* **Prompt injection:** Improved detection of system prompt exfiltration.
* **Prompt injection:** Retrained model with improved coverage on new prompt attack variants.
* **Content moderation:** Model update to include customer feedback.

## Bug Fixes

* Guard: Health probe fixes when gRPC encryption is turned on.
* Gateway: Fix policy resolution for multi-message requests.

## Security

* Gateway: Security fixes (CVE-2026-25679, CVE-2026-27142, CVE-2026-27139).

## Container version: 2.0.461, tag: stable

## Quality

* Improved GPU model for text and audio.
* Reduction of FPRs in audio guard.
* General model quality improvement based on customer feedback.

## Container version: 2.0.443, tag: stable

## Quality

* Updated prompt injection models incorporating customer feedback for improved accuracy.
* Fixed bug related to PII spans. Also improved our performance on detecting credit card numbers.
* Improved latency and throughput for self-hosted deployments.

## Container version: 2.0.431, tag: stable

## Quality

* Fixed an issue where allowlists could be incorrect.
* Updated moderation models incorporating customer feedback for improved accuracy.

## Container version: 2.0.410, tag: stable

## Improvements

### Quality

* **Moderation model improvements**: Updated moderation models to reduce FPRs, especially in `weapons` category.
* **Text preprocessing robustness**: Improved handling of escaped JSON characters and edge cases in text decoding, reducing preprocessing errors and improving classifier reliability.
* **Whitelist refinements**: Removed common phrases from whitelist to improve detection accuracy.

## Container version: 2.0.371, tag: stable

## Improvements

### Platform

* Fixed bug breaking onboarding page for some users.

### Quality

* Expanded API to cover `tool` role and `tool_calls` to better support Agentic workflows.

## Container version: 2.0.350, tag: stable

## Improvements

### Platform

* Improve Logs page loading speed.
* Fixed bug which disallowed the same detector to be used for input and output.

### Quality

* Improved handling base64-encoded input.

## Container version: 2.0.328, tag: stable

## Improvements

### Platform

* Fixed bug where pricing page was unavailable to community users.
* Improved loading speed and page performance on Logs and Analytics pages.
* Include logs flagged as “deny list” as threats in the Logs and Analytics pages.
* Fixed bug on Analytics page where data would not load for certain date ranges.

### Quality

* Fixed bug with allowed domains handling.
* Improved model recall on system prompt extraction attacks.

## Container version: 2.0.318, tag: stable

## Improvements

### Platform

* Bug fixes:
  * Policy badges show policy name now.
  * Autoplay in tutorial now works correctly.
* UX improvements:
  * Minor improvements across the platform UI.
  * In request details page: Updates the guardrail counter to count based on subcategories.
  * In Policies page, add a new "All Policies" tab and rename "Catalog" option to "Lakera Policy Catalog".

### Quality

* Improved Prompt Attack and PII model based on observed misclassifications.

## Container version: 2.0.301, tag: stable

## Improvements

### Platform

* Fixed bug where managed and custom guardrails could not be added to the same policy.
* UX improvements:
  * In Logs page, default to showing Threats rather than All Requests.
  * Show request metadata by default when viewing Log Details.

### Quality

* Improved Prompt Attack model based on observed misclassifications.
* Fixed tax-related false positives.
* Fixed unknown links detector: links with 3 apex domains were flagged even if they were known links.

## Container version: 2.0.289, tag: stable

## Improvements

### Platform

* Update Playground examples.
* In the Logs page, display all screened message content.
* Improved text rendering and layout when displaying Log Details.

### Documentation

* Fixed several broken links leading to non-existent pages.

### Quality

* Improved PACK model based on observed misclassifications.

## Container version: 2.0.258, tag: stable

## Improvements

### Platform

* Show which messages are flagged within a specific log.
* Submitting misclassified: Allow submitting for a specific log and for multiple logs at once.
* Submitting misclassified: Allow submitting for multiple logs at once.

### Quality

* Improved recall for PII/name detector, catching names in code-heavy prompts.
* Improved PACK model based on observed misclassifications.

## Container version: 2.0.220, tag: stable

## New Features

### Platform

* Guardrails: Enhanced guardrails overview table with filtering functionality and "Policies" column.
* Guardrails page: Added "creator" field, "last edited by/at" information, and "Policies assigned" section.
* Logs page: Added 'Link' button for improved navigation with and without current filters.
* Policy advanced settings: Redesigned with new sections for Content Moderation, Data Leakage Prevention, Prompt Defense, and Unknown Links.
* Policy page: Updated 'defenses' column for better visualization.
* Misclassification: You can now submit misclassifications in bulk.

## Improvements

### Platform

* Guardrails: Updated custom guardrail removal functionality.
* Fixed breadcrumb loading condition.

### Gateway

* Improved chunking of requests.

### Quality

* Improved PACK and moderation models based on observed misclassifications.

## Container version: 2.0.219, tag: stable

## New Features

### API

* Added support for `developer` and `tool` roles for messages to cover agentic and tool-using use cases.

### Platform

* Improved visual clarity of request details and cleaned up visual bugs. Contents are now displayed in a clean, chat-style format with clear roles and separation between messages.

## Improvements

### Platform

* Fixed date range component not updating data based on user selection

### Self-Hosted

* Fixed bug where `total_latency` value in logs would incorrectly show “0”

### Quality

* Credit card PII detector now recognizes non-standard Maestro card numbers

* Improved content moderation models based on observed misclassifications

## Container version: 2.0.184, tag: stable

## Improvements

### Guard

* Fixed feature flag bug that caused wrong flag values in edge cases.

### Quality

* Improved PACK and moderation models based on observed misclassifications.

## Container version: 2.0.165, tag: stable

## Improvements

### Platform

* Improved performance of Dashboard pages.
* Fixed pagination behavior after filters changed.

### Gateway

* Improved error handling.

### Quality

* Improved PACK and moderation models based on observed misclassifications.
* Further improvements for template detection to reduce FPR.

### Security

* Upgrade to `torch@2.7.0`.

## Changes

* `multi_language` request param is now no-op.

Note: the upgrade to `torch@2.7.0` led to a significant latency increase for PII classifiers on short inputs (\< 500 chars).

_Showing the 20 most recent of 49 entries. Append `/llms.txt` to the changelog URL for the complete index._