---
title: "Supported AI models and token pricing"
slug: "ai-inference-models-pricing"
source: "https://app.cloudpe.com/help/ai-inference-models-pricing"
updated: "2026-10-08T05:09:58.508Z"
---

# Supported AI models and token pricing

## Overview

The CloudPe AI Inference API provides serverless access to high-performance open weights models through an OpenAI-compatible endpoint at `https://inferapi.cloudpe.com/v1`. Customers pay only for the prompt and completion tokens processed by the inference cluster, with no ongoing VM infrastructure or idle GPU charges.

Models run on dedicated enterprise GPUs in CloudPe data centres in India. The public inference endpoint is fronted by Cloudflare (see [AI inference reliability and architecture](/help/ai-inference-reliability)).

## Before you start

- Account eligibility: Your organization must be KYC-verified or funded with a direct paid top-up (promotional or bonus credits do not qualify).
- Permissions: You need the `ai:use` permission to view the model catalogue or `ai:keys` to mint inference keys.
- Quickstart guide: See [AI inference quickstart](/help/ai-inference-quickstart) to send your first completion request.
- Billing dashboard: Monitor token usage and spend in the console. See [Tracking AI usage and billing](/help/ai-usage-and-billing).

## Steps

### Browse models in the dashboard

1. In the sidebar, open **Models** under **AI**.
2. Review model cards displaying model display name, slug, capability tags, and context windows.
3. Check the pricing section on each card for prompt and completion token rates.
4. Select the code generator to copy ready-to-use curl, Python, or Node.js snippets.

## API

Retrieve active models and real-time token rates programmatically via the console management API. This endpoint requires a console API key with the `ai:use` scope (or an active dashboard session). Inference keys carry only `inference:invoke` and authenticate to `https://inferapi.cloudpe.com/v1` for model inference—they cannot call this management route. See [Creating and managing AI API keys](/help/ai-api-keys).

| Method and path | Permission |
|---|---|
| `GET /api/v1/ai/models` | `ai:use` |

List active models:

```bash
curl https://app.cloudpe.com/api/v1/ai/models \
  -H "Authorization: Bearer <CONSOLE_API_KEY>"
```

### Supported models directory

The catalogue offers four primary models optimized for latency, reasoning, code generation, and general-purpose chat:

| Model slug | Model name | Context length | Typical use case | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|---|---|---|
| `llama-3-1-8b` | Llama 3.1 8B Instruct | 32k (32,768 tokens) | High-throughput chat, summarization, extraction, classification | INR 20 | INR 30 |
| `gemma-3-27b` | Gemma 3 27B Instruct | 16k (16,384 tokens) | General reasoning, multilingual translation, document analysis | INR 30 | INR 50 |
| `qwen3-32b` | Qwen3 32B Instruct | 16k (16,384 tokens) | Coding, structured JSON generation, agent workflows | INR 30 | INR 50 |
| `deepseek-r1-distill-qwen-32b` | DeepSeek R1 Distill Qwen 32B | 16k (16,384 tokens) | Deep reasoning, math, logic, complex code synthesis | INR 30 | INR 50 |

Context length specification: The context lengths listed above reflect each model's deployed context window (32,768 tokens for Llama 3.1 8B, and 16,384 tokens for Gemma 3 27B, Qwen3 32B, and DeepSeek R1 Distill Qwen 32B).

### Model selection guidance

- `llama-3-1-8b`: Best choice when latency and cost efficiency are critical. Ideal for high-volume customer support chat, classification, data parsing, and low-latency agent tasks.
- `gemma-3-27b`: Excellent balance of reasoning depth and latency. Recommended for general knowledge question answering, multilingual processing, and complex document comprehension.
- `qwen3-32b`: Superior performance on programming languages, API tool use, and strict JSON output schemas. Recommended for development tooling, SQL generation, and code review assistants.
- `deepseek-r1-distill-qwen-32b`: Built for step-by-step thinking and rigorous mathematical or logical problems. Recommended for technical analysis, theorem verification, and difficult problem-solving.

## Limits & billing

- Token metering: Usage is counted by the serving engine on each request and rated in INR per million tokens.
- Cached tokens: Cached prompt tokens are billed at the model's standard input rate.
- Settlement: Prepaid accounts are debited hourly from wallet balances. Postpaid accounts receive itemized lines on monthly invoices.
- Concurrency and rate limits: Organization defaults are 60 requests per minute and 4 concurrent requests. See [AI inference rate limits and spend caps](/help/ai-inference-limits).

## Troubleshooting

| Message | What it means | What to do |
|---|---|---|
| `AI Gateway is not enabled` | AI Gateway is disabled in the active environment. | Contact support or check back once the service is live. |
| `Permission denied: an AI permission is required` | Account lacks required AI role permissions. | Request the `ai:use` permission from an organization administrator. |

## FAQ

**Are prompt and completion tokens billed separately?**
Yes. Input tokens and output tokens carry distinct rates per million tokens as specified in the pricing table.

**Does CloudPe support prompt caching?**
Yes. When supported by model configuration, repetitive prefix tokens utilize KV cache. Cached prompt tokens are billed at the model's standard input rate.

**What happens if a model is under heavy load?**
Under high capacity pressure, a request may be served by an AWQ quantized variant of the same model. Customers are billed the requested model's standard price. See [AI inference reliability and architecture](/help/ai-inference-reliability).

## Related

- [AI inference quickstart](/help/ai-inference-quickstart)
- [Browsing AI models and using the playground](/help/ai-models-and-playground)
- [Tracking AI usage and billing](/help/ai-usage-and-billing)
- [AI inference reliability and architecture](/help/ai-inference-reliability)