Supported AI models and token pricing
Overview
The CloudPe AI Inference API provides serverless access to high-performance open weights models through an OpenAI-compatible endpoint at https://inferapi.cloudpe.com/v1. Customers pay only for the prompt and completion tokens processed by the inference cluster, with no ongoing VM infrastructure or idle GPU charges.
Models run on dedicated enterprise GPUs in CloudPe data centres in India. The public inference endpoint is fronted by Cloudflare (see AI inference reliability and architecture).
Before you start
- Account eligibility: Your organization must be KYC-verified or funded with a direct paid top-up (promotional or bonus credits do not qualify).
- Permissions: You need the
ai:usepermission to view the model catalogue orai:keysto mint inference keys. - Quickstart guide: See AI inference quickstart to send your first completion request.
- Billing dashboard: Monitor token usage and spend in the console. See Tracking AI usage and billing.
Steps
Browse models in the dashboard
- In the sidebar, open Models under AI.
- Review model cards displaying model display name, slug, capability tags, and context windows.
- Check the pricing section on each card for prompt and completion token rates.
- Select the code generator to copy ready-to-use curl, Python, or Node.js snippets.
API
Retrieve active models and real-time token rates programmatically via the console management API. This endpoint requires a console API key with the ai:use scope (or an active dashboard session). Inference keys carry only inference:invoke and authenticate to https://inferapi.cloudpe.com/v1 for model inference—they cannot call this management route. See Creating and managing AI API keys.
| Method and path | Permission |
|---|---|
GET /api/v1/ai/models |
ai:use |
List active models:
curl https://app.cloudpe.com/api/v1/ai/models \
-H "Authorization: Bearer <CONSOLE_API_KEY>"
Supported models directory
The catalogue offers four primary models optimized for latency, reasoning, code generation, and general-purpose chat:
| Model slug | Model name | Context length | Typical use case | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|---|---|---|
llama-3-1-8b |
Llama 3.1 8B Instruct | 32k (32,768 tokens) | High-throughput chat, summarization, extraction, classification | INR 20 | INR 30 |
gemma-3-27b |
Gemma 3 27B Instruct | 16k (16,384 tokens) | General reasoning, multilingual translation, document analysis | INR 30 | INR 50 |
qwen3-32b |
Qwen3 32B Instruct | 16k (16,384 tokens) | Coding, structured JSON generation, agent workflows | INR 30 | INR 50 |
deepseek-r1-distill-qwen-32b |
DeepSeek R1 Distill Qwen 32B | 16k (16,384 tokens) | Deep reasoning, math, logic, complex code synthesis | INR 30 | INR 50 |
Context length specification: The context lengths listed above reflect each model's deployed context window (32,768 tokens for Llama 3.1 8B, and 16,384 tokens for Gemma 3 27B, Qwen3 32B, and DeepSeek R1 Distill Qwen 32B).
Model selection guidance
llama-3-1-8b: Best choice when latency and cost efficiency are critical. Ideal for high-volume customer support chat, classification, data parsing, and low-latency agent tasks.gemma-3-27b: Excellent balance of reasoning depth and latency. Recommended for general knowledge question answering, multilingual processing, and complex document comprehension.qwen3-32b: Superior performance on programming languages, API tool use, and strict JSON output schemas. Recommended for development tooling, SQL generation, and code review assistants.deepseek-r1-distill-qwen-32b: Built for step-by-step thinking and rigorous mathematical or logical problems. Recommended for technical analysis, theorem verification, and difficult problem-solving.
Limits & billing
- Token metering: Usage is counted by the serving engine on each request and rated in INR per million tokens.
- Cached tokens: Cached prompt tokens are billed at the model's standard input rate.
- Settlement: Prepaid accounts are debited hourly from wallet balances. Postpaid accounts receive itemized lines on monthly invoices.
- Concurrency and rate limits: Organization defaults are 60 requests per minute and 4 concurrent requests. See AI inference rate limits and spend caps.
Troubleshooting
| Message | What it means | What to do |
|---|---|---|
AI Gateway is not enabled |
AI Gateway is disabled in the active environment. | Contact support or check back once the service is live. |
Permission denied: an AI permission is required |
Account lacks required AI role permissions. | Request the ai:use permission from an organization administrator. |
FAQ
Are prompt and completion tokens billed separately? Yes. Input tokens and output tokens carry distinct rates per million tokens as specified in the pricing table.
Does CloudPe support prompt caching? Yes. When supported by model configuration, repetitive prefix tokens utilize KV cache. Cached prompt tokens are billed at the model's standard input rate.
What happens if a model is under heavy load? Under high capacity pressure, a request may be served by an AWQ quantized variant of the same model. Customers are billed the requested model's standard price. See AI inference reliability and architecture.

