# REST API

Every CKEditor AI call is an HTTP request to one of the endpoint families below. This page lists them, says which one you call, and covers the mechanics they all share: versioning, pagination, limits, and error codes. For the base URL, see [SaaS](saas.md) or [On-Premises](on-premises.md).

For the editor integration, start with the [CKEditor 5 AI documentation](../../../../ckeditor5/latest/features/ai/ckeditor-ai-overview.md).

<a id="endpoint-families">

## Endpoint families

| Family                                        | What it does                                                                                        |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| [Conversations](conversations.md)             | A chat that keeps its history, with documents and files attached to it                              |
| [Actions](actions.md)                         | One transformation of a piece of content, with no history                                           |
| [Reviews](reviews.md)                         | A suggestion for each element of a document, anchored by its `data-id`                              |
| [Document Processing](document-processing.md) | One prompt applied to whole documents, returning the edited documents and a summary of what changed |

Conversations, Actions, and Reviews are called by the CKEditor 5 AI features. You do not choose between them. The feature the user runs does. Choose Document Processing when there is no editor. CKEditor 5 can call it too, through a gateway, and [Using CKEditor AI programmatically](../../../../ckeditor5/latest/features/ai/ckeditor-ai-programmatic.md) shows how.

To run the editor itself on a server and write AI results into collaboration documents as suggestions, see [Server-side Editor API](../../developer-resources/server-side-editor-api/editor-scripts.md).

<a id="endpoint-documentation">

## Endpoint documentation

Every endpoint is documented at **[ai.cke-cs.com/v1/docs](https://ai.cke-cs.com/v1/docs)**, with its request schema, response schema, and error responses. The same content is available as an OpenAPI JSON file, for import into a tool such as Postman or to hand to a coding assistant:

```
https://ai.cke-cs.com/v1/doc.json
```

On on-premises deployments, your instance serves the same documentation at `/v1/docs` and the JSON file at `/v1/doc.json`, for example `http://localhost:8000/v1/docs`. The shorter `/docs` redirects there.

<a id="versioning">

## Versioning

The URL path of each endpoint contains its version, for example `/v1/conversations` or `/v1.1/conversations/{conversationId}/messages`. Endpoints are versioned independently. A breaking change ships as a new version of that endpoint. The old version keeps working, and you can migrate when you are ready.

The models endpoint (`GET /v1/models/{compatibilityVersion}`) also takes a compatibility version, which identifies the set of models available to your integration. See [Models](models.md) for how compatibility versions work.

<a id="pagination">

## Pagination

List endpoints use cursor pagination: conversations, files, documents, web resources, contexts, context files, prompts, context libraries, and the admin conversations listing. Each response contains `nextCursor` and `previousCursor`. To fetch the next page, send `nextCursor` in the `cursor` query parameter. Send `previousCursor` for the previous page. The [endpoint documentation](#endpoint-documentation) has the parameters for each endpoint.

<a id="limits">

## Limits

Rate limits apply to requests and tokens per minute. File limits apply to what you attach to a conversation. Each model also has its own context limits.

<a id="rate-limits">

### Rate limits

Rate limits apply to requests that call a model, search the web, or scrape a page. Request limits count requests, and the token limit counts input tokens. Both are enforced per minute, per organization and per user. When you exceed a limit, the API returns a 429 with code `rate-limits-exceeded`. Wait, then send the request again. The response has no `Retry-After` header.

> **Note**
>
> These limits apply to SaaS only. On-premises deployments do not rate limit requests or tokens. The only limits are those of your infrastructure and your LLM provider.

<a id="request-limits">

#### Request limits

| Limit                 | Organization | Per user |
| --------------------- | ------------ | -------- |
| LLM requests          | 1,000 / min  | 75 / min |
| Web search requests   | 300 / min    | 30 / min |
| Web scraping requests | 100 / min    | 5 / min  |

<a id="token-limits">

#### Token limits

| Limit  | Organization     | Per user        |
| ------ | ---------------- | --------------- |
| Tokens | 16,000,000 / min | 1,000,000 / min |

The token limit counts input tokens, weighted per model, so the same request consumes a different share of the limit depending on the model. The weights are internal to the service and not returned by the API. Cached tokens count too, except for Anthropic cache reads.

<a id="file-and-content-limits">

### File and content limits

A file attached to a conversation can be 25 MB at most (PDF, DOCX, PNG, JPEG, Markdown, HTML, plain text). You can attach up to 100 files per conversation, 30 MB in total. The PDFs in a conversation must not have more than 100 pages in total.

Anthropic and Agent models limit images to 5 MB. Other file types keep the 25 MB limit.

Each model has its own context length, file size, and file count limits, returned by `GET /v1/models/{compatibilityVersion}`. See [Model information](models.md#model-information) on the Models page for an example response.

<a id="error-codes">

## Error codes

Every error response carries `statusCode`, `code`, `message`, and `traceId`. Some errors also include a `data` object with details, an `explanation`, and an `action`. Example:

```json
{
  "statusCode": 403,
  "code": "missing-permissions",
  "message": "The user is missing permissions",
  "traceId": "8f3c1c2e-...",
  "data": { "missingPermissions": ["ai:conversations:write"] }
}
```

The **[endpoint documentation](https://ai.cke-cs.com/v1/docs#section/Overview/Error-codes)** lists every error code.

<a id="usage-and-billing">

## Usage and billing

On SaaS, `POST /v1/admin/billing/usage` returns the credits your environment consumed over a period, in total and per user. It requires a token with the `ai:admin` permission. See [Usage and billing](usage-and-billing.md) for examples.

<a id="next-steps">

## Next steps

* **[Authentication](authentication.md)** covers the JWT bearer token every request carries, and the token endpoint that issues it.
* **[Streaming protocol](streaming-protocol.md)** defines the Server-Sent Events that carry generated content.
* **[Permissions](permissions.md)** lists the scopes a token can carry.
* **[Models](models.md)** lists the available models, their capabilities, and compatibility versions.

---

Full index of the Cloud Services documentation: [llms.txt](../../../llms.txt)
