AI inference error codes and retries
Overview
The CloudPe AI Inference API at https://inferapi.cloudpe.com/v1 returns JSON error bodies for most gateway failures using an OpenAI-like error object (message, type, code). Client applications can inspect these fields to handle rate limits, authentication issues, and temporary capacity variations gracefully.
Gateway-generated errors use the shapes documented below. Upstream engine rejections on HTTP 400 are passed through from the serving engine and may use different type or code values. Cloudflare edge errors (for example HTTP 524 after a long idle read) may return HTML instead of JSON.
Before you start
- Client library compatibility: Official OpenAI SDKs automatically parse error payloads into typed exception objects.
- Quickstart guide: See AI inference quickstart to configure valid client credentials.
- Limit policies: Review RPM and concurrency quota guidelines in AI inference rate limits and spend caps.
Steps
Handle errors in application code
- Catch API error exceptions in your client code.
- Check the HTTP status code and error code string.
- For status 429, 502, or 503, read the Retry-After header and execute an exponential backoff retry.
- For status 401, 403, 404, 413, or 422, log the defect and fail without retrying until inputs or credentials are corrected.
API
Error response schema:
{
"error": {
"message": "Key requests-per-minute limit exceeded.",
"type": "rate_limit_error",
"code": "rate_limit_exceeded"
}
}
Gateway error codes reference
The following error codes are emitted by CloudPe AI Gateway nodes:
| HTTP status | Error type | Error code | Cause | Retryable? |
|---|---|---|---|---|
| 400 | (engine-dependent) | (engine-dependent) | Upstream engine rejection (e.g. prompt length exceeds context window) passed through unchanged from the serving engine | No (truncate prompt or reduce max_tokens) |
| 401 | authentication_error | invalid_api_key |
Missing, malformed, unrecognized, or revoked API key | No |
| 402 | insufficient_quota | insufficient_quota |
Prepaid wallet balance exhausted | No (top up wallet) |
| 403 | permission_error | org_suspended |
Organization account is suspended | No (contact support) |
| 403 | permission_error | insufficient_scope |
API key lacks the required inference:invoke scope |
No |
| 403 | permission_error | model_not_visible |
Model is restricted and not visible to organization | No |
| 404 | invalid_request_error | model_not_found |
Requested model slug does not exist or is inactive | No |
| 413 | invalid_request_error | request_too_large |
Request payload exceeds the maximum permitted body size | No (reduce payload) |
| 422 | invalid_request_error | invalid_request |
Malformed JSON body or invalid request parameters | No |
| 429 | rate_limit_error | rate_limit_exceeded |
Requests per minute, TPM, or concurrency limit breached | Yes (honour Retry-After) |
| 429 | rate_limit_error | budget_exceeded |
Monthly spend cap reached for key or organization | Yes (after cap reset/increase) |
| 502 | server_error | upstream_error |
Internal model serving engine failure or timeout | Yes (exponential backoff) |
| 503 | server_error | service_unavailable |
Policy snapshot expired or service undergoing maintenance | Yes (exponential backoff) |
| 503 | server_error | model_unavailable |
Model is not ready or has no healthy serving cluster | Yes (exponential backoff) |
| 503 | server_error | limiter_unavailable |
Rate limiter service temporarily unreachable | Yes (brief backoff) |
| 503 | server_error | not_ready |
Gateway node startup probe in progress | Yes (brief backoff) |
Limits & billing
- Retry guidance:
- Retryable errors: HTTP 429, 502, and 503 represent transient conditions. Implement exponential backoff with randomized jitter. When a
Retry-Afterheader is present, wait at least the duration indicated before dispatching a retry. - Non-retryable errors: HTTP 400, 401, 403, 404, 413, and 422 represent client-side validation, context length overflow, or configuration errors. Retrying identical requests will repeatedly fail.
- Retryable errors: HTTP 429, 502, and 503 represent transient conditions. Implement exponential backoff with randomized jitter. When a
- Long generations: Prefer streaming (
stream: true) for models with long time-to-first-token or extended reasoning (for exampledeepseek-r1-distill-qwen-32b) so connections stay active. Non-streaming calls that exceed the edge read timeout may fail with Cloudflare HTTP 524 and an HTML error page. - Failed request billing: Token quota is not consumed by failed requests (requests returning 4xx or 5xx status codes do not incur token billing charges). However, request-per-minute (RPM) and concurrency counters are consumed for requests that pass the rate limiter before failing upstream (such as an HTTP 502 upstream error).
Troubleshooting
| Message | What it means | What to do |
|---|---|---|
AI Gateway is not enabled |
AI Gateway is disabled in this environment. | Check system status announcements or contact support. |
AI_KEY_MINT_REQUIRES_KYC_OR_FUNDS |
Organization is not KYC-verified and has no qualifying paid wallet top-up when minting keys or playground tokens. | Complete KYC or add a direct paid top-up, then retry. |
key not found |
Key ID queried in management endpoints does not exist. | Refresh keys list or verify the key UUID. |
FAQ
How should my application handle Retry-After headers? The Retry-After HTTP header contains the integer number of seconds to pause before sending another request. Automated clients should read this value and sleep accordingly.
Why does upstream_error return 502 without details? To protect security and privacy, raw engine stack traces or server internals are never exposed to clients. CloudPe returns sanitized standard error messages.
What should I do if budget_exceeded persists? Increase the monthly spend cap in the console under API Keys under AI, or wait for the cap to reset automatically at the beginning of the next calendar month.

