AI inference reliability and architecture

Last updated 8 Oct 2026
View as Markdown

Overview

The CloudPe AI Inference platform provides serverless access to open-weights language models deployed in Indian data centres.

Model execution and gateway processing run in CloudPe data centres in India; the public endpoint is fronted by Cloudflare. Automated load handling helps maintain availability during traffic surges.

Before you start

  • Account eligibility: Your organization must be KYC-verified or funded with a direct paid top-up (promotional or bonus credits do not qualify).
  • Error handling: Ensure your application implements retry logic for transient status codes. See AI inference error codes and retries.
  • Models catalogue: Explore supported models and context windows in Supported AI models and token pricing.

Steps

Verify endpoint connectivity

  1. Test connection to the primary API host https://inferapi.cloudpe.com/v1:
curl https://inferapi.cloudpe.com/v1/models \
  -H "Authorization: Bearer $CLOUDPE_API_KEY"
  1. Confirm model readiness and check response status.

API

Management endpoints to inspect service availability via the console management API. Use a signed-in dashboard session or an unrestricted console API key. Scoped console API keys (including those with ai:keys) cannot call GET /api/v1/ai/enabled today. Inference keys carry only inference:invoke and authenticate to https://inferapi.cloudpe.com/v1 for model inference—they cannot call most app.cloudpe.com management APIs. See Creating and managing AI API keys.

Method and path Permission
GET /api/v1/ai/enabled Authenticated user (unrestricted console key or session)
GET /api/v1/ai/models ai:use

Check service status (unrestricted console API key or dashboard session):

curl https://app.cloudpe.com/api/v1/ai/enabled \
  -H "Authorization: Bearer <UNRESTRICTED_CONSOLE_API_KEY>"

Limits & billing

Data residency and network path

  • Model execution: GPU instances that run inference workloads are located in CloudPe data centres in India.
  • Gateway processing: Authentication, rate limiting, and policy enforcement run on CloudPe gateway nodes in India.
  • Inter-cluster forwarding: When a request is forwarded between CloudPe clusters, traffic stays on CloudPe private network links within India.
  • Public endpoint: https://inferapi.cloudpe.com/v1 is fronted by Cloudflare for TLS termination and DDoS protection. Client traffic therefore passes through Cloudflare's edge (whose location depends on the client) before reaching CloudPe origin infrastructure in India.
  • Billing records: Usage metering records token counts, model identifier, request timing, and status for billing and observability.

Capacity management and quantized fallback disclosure

Model instances are deployed in Zone B with elastic replica policies (min_replicas=0, max_replicas=1). When a model scales from zero, initial requests may encounter cold-start latency while the engine initializes.

Under conditions of high cluster load or capacity saturation, CloudPe implements automated fallback to preserve availability:

  • Quantized serving: A request may be transparently served by an AWQ (Activation-aware Weight Quantization) variant of the requested model.
  • Output quality: Under fallback, output quality may differ slightly from full-precision serving (wording, formatting, or tool-call reliability).
  • Transparent reporting: The API response model field continues to report the exact model slug requested by the client application.
  • Fair billing guarantee: Customers are billed according to the standard published token pricing of the requested model. No price adjustments or surcharges occur during fallback serving.
  • API contract: Response JSON shape and SSE streaming mechanics follow the same OpenAI-compatible schema; validate critical paths in your application if you rely on AWQ fallback under load.

Troubleshooting

Message What it means What to do
AI Gateway is not enabled AI Gateway is disabled in the active environment. Contact support or check service status announcements.
Permission denied: an AI permission is required Account lacks required AI role permissions. Request the ai:use permission from an organization administrator.

FAQ

Does AWQ quantized fallback change model output? The API schema is unchanged, but text and formatting may differ slightly compared with full-precision serving. Test your prompts if output fidelity is critical.

What happens if all serving clusters for a model are down? The gateway returns an HTTP 503 response with error code model_unavailable. Clients should retry with exponential backoff.

Where does my traffic go before it reaches the GPU? Clients connect to inferapi.cloudpe.com through Cloudflare's edge, then to CloudPe gateway and serving infrastructure in Indian data centres. Model execution does not run outside CloudPe's Indian facilities.

Related

Did this guide answer your question?If you need customized assistance with your deployment, reach out to our team.
Contact Support