---
title: "AI inference reliability and architecture"
slug: "ai-inference-reliability"
source: "https://app.cloudpe.com/help/ai-inference-reliability"
updated: "2026-10-08T05:09:58.520Z"
---

# AI inference reliability and architecture

## Overview

The CloudPe AI Inference platform provides serverless access to open-weights language models deployed in Indian data centres.

Model execution and gateway processing run in CloudPe data centres in India; the public endpoint is fronted by Cloudflare. Automated load handling helps maintain availability during traffic surges.

## Before you start

- Account eligibility: Your organization must be KYC-verified or funded with a direct paid top-up (promotional or bonus credits do not qualify).
- Error handling: Ensure your application implements retry logic for transient status codes. See [AI inference error codes and retries](/help/ai-inference-errors).
- Models catalogue: Explore supported models and context windows in [Supported AI models and token pricing](/help/ai-inference-models-pricing).

## Steps

### Verify endpoint connectivity

1. Test connection to the primary API host `https://inferapi.cloudpe.com/v1`:

```bash
curl https://inferapi.cloudpe.com/v1/models \
  -H "Authorization: Bearer $CLOUDPE_API_KEY"
```

2. Confirm model readiness and check response status.

## API

Management endpoints to inspect service availability via the console management API. Use a signed-in dashboard session or an unrestricted console API key. Scoped console API keys (including those with `ai:keys`) cannot call `GET /api/v1/ai/enabled` today. Inference keys carry only `inference:invoke` and authenticate to `https://inferapi.cloudpe.com/v1` for model inference—they cannot call most `app.cloudpe.com` management APIs. See [Creating and managing AI API keys](/help/ai-api-keys).

| Method and path | Permission |
|---|---|
| `GET /api/v1/ai/enabled` | Authenticated user (unrestricted console key or session) |
| `GET /api/v1/ai/models` | `ai:use` |

Check service status (unrestricted console API key or dashboard session):

```bash
curl https://app.cloudpe.com/api/v1/ai/enabled \
  -H "Authorization: Bearer <UNRESTRICTED_CONSOLE_API_KEY>"
```

## Limits & billing

### Data residency and network path

- Model execution: GPU instances that run inference workloads are located in CloudPe data centres in India.
- Gateway processing: Authentication, rate limiting, and policy enforcement run on CloudPe gateway nodes in India.
- Inter-cluster forwarding: When a request is forwarded between CloudPe clusters, traffic stays on CloudPe private network links within India.
- Public endpoint: `https://inferapi.cloudpe.com/v1` is fronted by Cloudflare for TLS termination and DDoS protection. Client traffic therefore passes through Cloudflare's edge (whose location depends on the client) before reaching CloudPe origin infrastructure in India.
- Billing records: Usage metering records token counts, model identifier, request timing, and status for billing and observability.

### Capacity management and quantized fallback disclosure

Model instances are deployed in Zone B with elastic replica policies (`min_replicas=0`, `max_replicas=1`). When a model scales from zero, initial requests may encounter cold-start latency while the engine initializes.

Under conditions of high cluster load or capacity saturation, CloudPe implements automated fallback to preserve availability:

- Quantized serving: A request may be transparently served by an AWQ (Activation-aware Weight Quantization) variant of the requested model.
- Output quality: Under fallback, output quality may differ slightly from full-precision serving (wording, formatting, or tool-call reliability).
- Transparent reporting: The API response `model` field continues to report the exact model slug requested by the client application.
- Fair billing guarantee: Customers are billed according to the standard published token pricing of the requested model. No price adjustments or surcharges occur during fallback serving.
- API contract: Response JSON shape and SSE streaming mechanics follow the same OpenAI-compatible schema; validate critical paths in your application if you rely on AWQ fallback under load.

## Troubleshooting

| Message | What it means | What to do |
|---|---|---|
| `AI Gateway is not enabled` | AI Gateway is disabled in the active environment. | Contact support or check service status announcements. |
| `Permission denied: an AI permission is required` | Account lacks required AI role permissions. | Request the `ai:use` permission from an organization administrator. |

## FAQ

**Does AWQ quantized fallback change model output?**
The API schema is unchanged, but text and formatting may differ slightly compared with full-precision serving. Test your prompts if output fidelity is critical.

**What happens if all serving clusters for a model are down?**
The gateway returns an HTTP 503 response with error code `model_unavailable`. Clients should retry with exponential backoff.

**Where does my traffic go before it reaches the GPU?**
Clients connect to `inferapi.cloudpe.com` through Cloudflare's edge, then to CloudPe gateway and serving infrastructure in Indian data centres. Model execution does not run outside CloudPe's Indian facilities.

## Related

- [Supported AI models and token pricing](/help/ai-inference-models-pricing)
- [AI inference quickstart](/help/ai-inference-quickstart)
- [AI inference error codes and retries](/help/ai-inference-errors)
- [Tracking AI usage and billing](/help/ai-usage-and-billing)