AI inference error codes and retries

Last updated 8 Oct 2026
View as Markdown

Overview

The CloudPe AI Inference API at https://inferapi.cloudpe.com/v1 returns JSON error bodies for most gateway failures using an OpenAI-like error object (message, type, code). Client applications can inspect these fields to handle rate limits, authentication issues, and temporary capacity variations gracefully.

Gateway-generated errors use the shapes documented below. Upstream engine rejections on HTTP 400 are passed through from the serving engine and may use different type or code values. Cloudflare edge errors (for example HTTP 524 after a long idle read) may return HTML instead of JSON.

Before you start

Steps

Handle errors in application code

  1. Catch API error exceptions in your client code.
  2. Check the HTTP status code and error code string.
  3. For status 429, 502, or 503, read the Retry-After header and execute an exponential backoff retry.
  4. For status 401, 403, 404, 413, or 422, log the defect and fail without retrying until inputs or credentials are corrected.

API

Error response schema:

{
  "error": {
    "message": "Key requests-per-minute limit exceeded.",
    "type": "rate_limit_error",
    "code": "rate_limit_exceeded"
  }
}

Gateway error codes reference

The following error codes are emitted by CloudPe AI Gateway nodes:

HTTP status Error type Error code Cause Retryable?
400 (engine-dependent) (engine-dependent) Upstream engine rejection (e.g. prompt length exceeds context window) passed through unchanged from the serving engine No (truncate prompt or reduce max_tokens)
401 authentication_error invalid_api_key Missing, malformed, unrecognized, or revoked API key No
402 insufficient_quota insufficient_quota Prepaid wallet balance exhausted No (top up wallet)
403 permission_error org_suspended Organization account is suspended No (contact support)
403 permission_error insufficient_scope API key lacks the required inference:invoke scope No
403 permission_error model_not_visible Model is restricted and not visible to organization No
404 invalid_request_error model_not_found Requested model slug does not exist or is inactive No
413 invalid_request_error request_too_large Request payload exceeds the maximum permitted body size No (reduce payload)
422 invalid_request_error invalid_request Malformed JSON body or invalid request parameters No
429 rate_limit_error rate_limit_exceeded Requests per minute, TPM, or concurrency limit breached Yes (honour Retry-After)
429 rate_limit_error budget_exceeded Monthly spend cap reached for key or organization Yes (after cap reset/increase)
502 server_error upstream_error Internal model serving engine failure or timeout Yes (exponential backoff)
503 server_error service_unavailable Policy snapshot expired or service undergoing maintenance Yes (exponential backoff)
503 server_error model_unavailable Model is not ready or has no healthy serving cluster Yes (exponential backoff)
503 server_error limiter_unavailable Rate limiter service temporarily unreachable Yes (brief backoff)
503 server_error not_ready Gateway node startup probe in progress Yes (brief backoff)

Limits & billing

  • Retry guidance:
    • Retryable errors: HTTP 429, 502, and 503 represent transient conditions. Implement exponential backoff with randomized jitter. When a Retry-After header is present, wait at least the duration indicated before dispatching a retry.
    • Non-retryable errors: HTTP 400, 401, 403, 404, 413, and 422 represent client-side validation, context length overflow, or configuration errors. Retrying identical requests will repeatedly fail.
  • Long generations: Prefer streaming (stream: true) for models with long time-to-first-token or extended reasoning (for example deepseek-r1-distill-qwen-32b) so connections stay active. Non-streaming calls that exceed the edge read timeout may fail with Cloudflare HTTP 524 and an HTML error page.
  • Failed request billing: Token quota is not consumed by failed requests (requests returning 4xx or 5xx status codes do not incur token billing charges). However, request-per-minute (RPM) and concurrency counters are consumed for requests that pass the rate limiter before failing upstream (such as an HTTP 502 upstream error).

Troubleshooting

Message What it means What to do
AI Gateway is not enabled AI Gateway is disabled in this environment. Check system status announcements or contact support.
AI_KEY_MINT_REQUIRES_KYC_OR_FUNDS Organization is not KYC-verified and has no qualifying paid wallet top-up when minting keys or playground tokens. Complete KYC or add a direct paid top-up, then retry.
key not found Key ID queried in management endpoints does not exist. Refresh keys list or verify the key UUID.

FAQ

How should my application handle Retry-After headers? The Retry-After HTTP header contains the integer number of seconds to pause before sending another request. Automated clients should read this value and sleep accordingly.

Why does upstream_error return 502 without details? To protect security and privacy, raw engine stack traces or server internals are never exposed to clients. CloudPe returns sanitized standard error messages.

What should I do if budget_exceeded persists? Increase the monthly spend cap in the console under API Keys under AI, or wait for the cap to reset automatically at the beginning of the next calendar month.

Related

Did this guide answer your question?If you need customized assistance with your deployment, reach out to our team.
Contact Support