Sign up (with export icon)

Moderation

Show the table of contents

Content moderation checks the prompt of every AI call before it reaches the model provider and rejects content that the moderation provider classifies as unsafe. The check runs by default, against the OpenAI moderation API.

You have three choices for your deployment. Keep the OpenAI provider and set the model it calls. Send the content to a moderation endpoint that you run. Or turn the check off and screen the content yourself.

The check runs on the prompt of every AI call: chat messages, actions, reviews, and Document Processing calls. It also runs on images at upload, on prompts saved to the context library, and on text files from the context library that an action, a review, or a Document Processing call attaches. No other attachment reaches the moderation provider, and model output is never checked. Guardrails inspect the files, editor documents, and web pages that a chat message attaches. See Security and compliance for the full data flow.

The provider reports six categories: sexual, harassment, hate, illicit, selfHarm, and violence. One flagged category is enough. The service rejects the request with HTTP 422 and the flaggedCategories, and the end user sees a message that the content was not allowed. A blocked request leaves no record you can retrieve: the service writes a log line, and there is no endpoint or export for blocked content.

You configure content moderation through the moderation option. It has two fields:

Configuring the OpenAI provider

Copy link

A request to the OpenAI moderation API times out after 2 seconds. You cannot change this timeout. The provider accepts the following options:

  • modelId (optional, default: omni-moderation-latest) – the moderation model to call.
  • url (optional, default: https://api.openai.com/v1/moderations) – the moderation API endpoint. Set it to route the calls through a proxy or a gateway.
{
	"moderation": {
		"enabled": true,
		"provider": {
			"type": "openai",
			"modelId": "omni-moderation-latest"
		}
	}
}
Copy code

Disabling moderation

Copy link

To turn content moderation off, set moderation.enabled to false:

{
	"moderation": {
		"enabled": false
	}
}
Copy code

With content moderation off, the service sends every prompt and image straight to the model provider. You screen unsafe content yourself. The usage policies of the model provider still apply, and the provider can reject the content itself.

This option does not affect guardrails. You configure guardrails separately, through the guardrails option.

Using a custom moderation endpoint

Copy link

The openai provider sends the content to OpenAI. To use a service that you run instead, set the provider type to custom and give it the URL of your endpoint. Your endpoint must accept the request format and return the response format below.

To configure a custom moderation endpoint, set the moderation option:

{
	"moderation": {
		"enabled": true,
		"provider": {
			"type": "custom",
			"endpoint": "https://your-moderation-endpoint.com",
			"extraHttpHeaders": {
				"Authorization": "Bearer token"
			},
			"requestTimeout": 2000
		}
	}
}
Copy code

The custom provider takes the following fields:

  • type (required) – set it to custom.
  • endpoint (required) – the URL of your moderation endpoint.
  • extraHttpHeaders (required) – the headers that the service adds to every moderation request, as an object of string values. Pass {} if you need none.
  • requestTimeout (optional, default: 2000) – the request timeout in milliseconds.

A custom endpoint keeps the fail-open behavior. If your endpoint is down, the content goes through. If a screening outage must block requests instead, use a hook: hooks fail closed, so the chat turn fails while your endpoint is down. A hook covers chat messages only, and it runs before content moderation, so your endpoint receives unscreened content. See Moderate content with your own rules.

Request format

Copy link

Your endpoint must accept POST requests with a JSON body. The body has two forms: one for text, and one for an image.

For text content:

{
  "type": "text",
  "text": "content to moderate"
}
Copy code

For image content:

{
  "type": "image",
  "mediaType": "image/png",
  "image": "data:image/png;base64,..."
}
Copy code

Where:

  • type (required) – the type of the content to moderate. Allowed values: text, image.
  • text (required for text) – the text content to moderate.
  • mediaType (required for image) – the MIME type of the image (for example, image/png).
  • image (required for image) – the image encoded as a base64 data URL.

Response format

Copy link

Your endpoint must respond with HTTP 200 and a JSON body in this format:

{
  "flagged": true,
  "categories": {
	"sexual": false,
	"harassment": false,
	"hate": false,
	"illicit": false,
	"selfHarm": false,
	"violence": true
  }
}
Copy code

Where:

  • flagged (required) – whether the content is unsafe. When true, the request is rejected.
  • categories (optional) – one flag for each category in the example above. Each flag is a boolean and defaults to false when you omit it. The service records the flags when it rejects a request.