# Moderation

Content moderation checks the prompt of every AI call before it reaches the model provider and rejects content that the moderation provider classifies as unsafe. The check runs by default, against the OpenAI moderation API.

You have three choices for your deployment. Keep the OpenAI provider and set the model it calls. Send the content to a moderation endpoint that you run. Or turn the check off and screen the content yourself.

The check runs on the prompt of every AI call: chat messages, actions, reviews, and Document Processing calls. It also runs on images at upload, on prompts saved to the context library, and on text files from the context library that an action, a review, or a Document Processing call attaches. No other attachment reaches the moderation provider, and model output is never checked. [Guardrails](guardrails.md) inspect the files, editor documents, and web pages that a chat message attaches. See [Security and compliance](../../guides/ckeditor-ai/security-and-compliance.md) for the full data flow.

The provider reports six categories: `sexual`, `harassment`, `hate`, `illicit`, `selfHarm`, and `violence`. One flagged category is enough. The service rejects the request with HTTP 422 and the `flaggedCategories`, and the end user sees a message that the content was not allowed. A blocked request leaves no record you can retrieve: the service writes a log line, and there is no endpoint or export for blocked content.

You configure content moderation through the `moderation` option. It has two fields:

* `enabled` (required) – whether content moderation runs. When it is `false`, you do not need the `provider` field.

* `provider` (required when `enabled` is `true`) – the moderation provider. It supports two types:

  * `openai` – calls the [OpenAI moderation API](https://platform.openai.com/docs/guides/moderation) with the API key of the OpenAI provider that you configured in [LLM providers](llm-providers.md). See [Configuring the OpenAI provider](#configuring-the-openai-provider).
  * `custom` – sends the content to an endpoint you run. See [Using a custom moderation endpoint](#using-a-custom-moderation-endpoint).

> **Warning**
>
> The default `openai` provider needs an OpenAI API key. If you do not define the `moderation` option, content moderation is enabled with the `openai` provider, and that provider reads the API key of the OpenAI provider in your `providers` option.
>
> If your deployment configures no OpenAI provider, the check inspects nothing. Every request passes, and the service records one message at `debug` level. A deployment that runs only Anthropic, Google, or models you host yourself stays in this state until you configure an OpenAI provider, point the option at your own endpoint, or set `enabled` to `false`.

> **Warning**
>
> Content moderation fails open: when the check cannot complete, the content goes through. This happens when the moderation endpoint is unreachable, exceeds the request timeout, answers with a status other than 200, or returns a body that the service cannot parse. The service records a warning for each such request, `Error moderating the prompt` for prompt checks. Nothing is inspected until the endpoint answers again.
>
> Alert on that warning. Without an alert, a moderation outage is invisible: end users see normal replies, and the deployment reports no errors. See [Logs](logs.md) for the log format and [Observability](observability.md) for exporting logs to your monitoring backend.
>
> Detection is also probabilistic. The provider uses a model to judge the content, so accuracy depends on that model. Do not make content moderation the only protection for sensitive content.

<a id="configuring-the-openai-provider">

## Configuring the OpenAI provider

A request to the OpenAI moderation API times out after 2 seconds. You cannot change this timeout. The provider accepts the following options:

* `modelId` (optional, default: `omni-moderation-latest`) – the moderation model to call.
* `url` (optional, default: `https://api.openai.com/v1/moderations`) – the moderation API endpoint. Set it to route the calls through a proxy or a gateway.

**JSON**

```json
{
	"moderation": {
		"enabled": true,
		"provider": {
			"type": "openai",
			"modelId": "omni-moderation-latest"
		}
	}
}
```

**Environment variable**

```bash
MODERATION='{"enabled":true,"provider":{"type":"openai","modelId":"omni-moderation-latest"}}'
```

<a id="disabling-moderation">

## Disabling moderation

To turn content moderation off, set `moderation.enabled` to `false`:

**JSON**

```json
{
	"moderation": {
		"enabled": false
	}
}
```

**Environment variable**

```bash
MODERATION='{"enabled":false}'
```

With content moderation off, the service sends every prompt and image straight to the model provider. You screen unsafe content yourself. The usage policies of the model provider still apply, and the provider can reject the content itself.

This option does not affect [guardrails](guardrails.md). You configure guardrails separately, through the `guardrails` option.

<a id="using-a-custom-moderation-endpoint">

## Using a custom moderation endpoint

The `openai` provider sends the content to OpenAI. To use a service that you run instead, set the provider `type` to `custom` and give it the URL of your endpoint. Your endpoint must accept the request format and return the response format below.

To configure a custom moderation endpoint, set the `moderation` option:

**JSON**

```json
{
	"moderation": {
		"enabled": true,
		"provider": {
			"type": "custom",
			"endpoint": "https://your-moderation-endpoint.com",
			"extraHttpHeaders": {
				"Authorization": "Bearer token"
			},
			"requestTimeout": 2000
		}
	}
}
```

**Environment variable**

The value should be provided as a stringified JSON object:

```bash
MODERATION='{
	"enabled": true,
	"provider": {
		"type": "custom",
		"endpoint": "https://your-moderation-endpoint.com",
		"extraHttpHeaders": {
			"Authorization": "Bearer token"
		},
		"requestTimeout": 2000
	}
}'
```

The `custom` provider takes the following fields:

* `type` (required) – set it to `custom`.
* `endpoint` (required) – the URL of your moderation endpoint.
* `extraHttpHeaders` (required) – the headers that the service adds to every moderation request, as an object of string values. Pass `{}` if you need none.
* `requestTimeout` (optional, default: `2000`) – the request timeout in milliseconds.

A custom endpoint keeps the fail-open behavior. If your endpoint is down, the content goes through. If a screening outage must block requests instead, use a [hook](hooks.md): hooks fail closed, so the chat turn fails while your endpoint is down. A hook covers chat messages only, and it runs before content moderation, so your endpoint receives unscreened content. See [Moderate content with your own rules](../../guides/ckeditor-ai/content-moderation.md).

<a id="request-format">

### Request format

Your endpoint must accept `POST` requests with a JSON body. The body has two forms: one for text, and one for an image.

For text content:

```json
{
  "type": "text",
  "text": "content to moderate"
}
```

For image content:

```json
{
  "type": "image",
  "mediaType": "image/png",
  "image": "data:image/png;base64,..."
}
```

Where:

* `type` (required) – the type of the content to moderate. Allowed values: `text`, `image`.
* `text` (required for `text`) – the text content to moderate.
* `mediaType` (required for `image`) – the MIME type of the image (for example, `image/png`).
* `image` (required for `image`) – the image encoded as a base64 data URL.

<a id="response-format">

### Response format

Your endpoint must respond with HTTP 200 and a JSON body in this format:

```json
{
  "flagged": true,
  "categories": {
	"sexual": false,
	"harassment": false,
	"hate": false,
	"illicit": false,
	"selfHarm": false,
	"violence": true
  }
}
```

Where:

* `flagged` (required) – whether the content is unsafe. When `true`, the request is rejected.
* `categories` (optional) – one flag for each category in the example above. Each flag is a boolean and defaults to `false` when you omit it. The service records the flags when it rejects a request.

---

Full index of the Cloud Services documentation: [llms.txt](../../../llms.txt)
