---
title: "AI inference quickstart"
slug: "ai-inference-quickstart"
source: "https://app.cloudpe.com/help/ai-inference-quickstart"
updated: "2026-10-08T05:09:58.554Z"
---

# AI inference quickstart

## Overview

The CloudPe AI Inference API provides an OpenAI-compatible interface hosted at `https://inferapi.cloudpe.com/v1`. You can connect existing applications using official OpenAI SDKs in Python and Node.js or standard HTTP client tools like curl.

Model execution and gateway processing run in CloudPe data centres in India; the public endpoint is fronted by Cloudflare (see [AI inference reliability and architecture](/help/ai-inference-reliability)).

## Before you start

- CloudPe account: You need an active user account in an organization. See [Creating your CloudPe account](/help/account-signup-onboarding).
- Eligibility requirement: Your organization must be KYC-verified or funded with a direct paid top-up (promotional or bonus credits do not qualify) before you can mint inference keys or obtain a playground session token. After you have a valid inference key, calls to `https://inferapi.cloudpe.com/v1` do not re-check this gate on every request.
- Permission: Generating inference keys requires the `ai:keys` permission. Browsing models requires `ai:use`.
- Models: Check available models and context lengths in the model catalogue. See [Supported AI models and token pricing](/help/ai-inference-models-pricing).

## Steps

### Create an inference key

1. In the sidebar, open **API Keys** under **AI**.
2. Select **Create key** to open the creation dialog.
3. Enter a key name (for example, `app-quickstart`).
4. Optionally configure RPM, TPM, concurrency, or monthly spend caps.
5. Select **Create**.
6. Copy the plaintext key starting with `cpk_`. Store it in your local environment as `CLOUDPE_API_KEY`.
7. Select **Done** to close the modal.

### Call the API using curl

Run a chat completion request with streaming enabled:

```bash
curl -X POST https://inferapi.cloudpe.com/v1/chat/completions \
  -H "Authorization: Bearer $CLOUDPE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama-3-1-8b",
    "messages": [
      {"role": "system", "content": "You are a concise assistant."},
      {"role": "user", "content": "Explain serverless computing in two sentences."}
    ],
    "stream": true
  }'
```

### Call the API using OpenAI Python SDK

Install the official OpenAI package:

```bash
pip install openai
```

Initialize the client with the CloudPe base URL and execute a streaming completion:

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://inferapi.cloudpe.com/v1",
    api_key=os.environ.get("CLOUDPE_API_KEY"),
)

stream = client.chat.completions.create(
    model="llama-3-1-8b",
    messages=[
        {"role": "system", "content": "You are a helpful coding assistant."},
        {"role": "user", "content": "Write a quicksort implementation in Python."},
    ],
    stream=True,
)

for chunk in stream:
    # Final SSE chunk has usage only (empty choices) when include_usage is enabled.
    if not chunk.choices:
        continue
    content = chunk.choices[0].delta.content or ""
    print(content, end="", flush=True)
print()
```

### Call the API using OpenAI Node.js SDK

Install the package:

```bash
npm install openai
```

Execute a streaming completion in TypeScript or JavaScript:

```javascript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://inferapi.cloudpe.com/v1",
  apiKey: process.env.CLOUDPE_API_KEY,
});

async function main() {
  const stream = await client.chat.completions.create({
    model: "llama-3-1-8b",
    messages: [
      { role: "system", content: "You are a helpful assistant." },
      { role: "user", content: "Summarize the advantages of prompt caching." },
    ],
    stream: true,
  });

  for await (const chunk of stream) {
    process.stdout.write(chunk.choices[0]?.delta?.content || "");
  }
  process.stdout.write("\n");
}

main();
```

## API

Primary endpoints available on the inference host `https://inferapi.cloudpe.com/v1`:

| Endpoint | Method | Purpose |
|---|---|---|
| `https://inferapi.cloudpe.com/v1/chat/completions` | POST | Create chat completions (streaming and non-streaming) |
| `https://inferapi.cloudpe.com/v1/models` | GET | List models available to your organization |

Management endpoints on `https://app.cloudpe.com`:

| Method and path | Permission |
|---|---|
| `GET /api/v1/ai/keys` | `ai:keys` |
| `POST /api/v1/ai/keys` | `ai:keys` |

## Limits & billing

- Default quotas: Organizations receive 60 requests per minute and 4 concurrent requests by default. See [AI inference rate limits and spend caps](/help/ai-inference-limits).
- Metering: Charges apply per token consumed (prompt, cached, and completion). Token usage events stream from inference engines to billing services.
- Currency: Token consumption is billed in INR to the paisa.
- Billing schedule: Prepaid organizations are debited automatically from wallet balances. Postpaid accounts receive itemized lines on monthly invoices.

## Troubleshooting

| Message | What it means | What to do |
|---|---|---|
| `API key does not have required AI scopes` | An inference API key was used for `GET /api/v1/ai/usage`, which requires the `ai:usage` scope. Inference keys carry only `inference:invoke`. | Query usage with a console API key that includes `ai:usage` or with a signed-in dashboard session. |
| `API key scope does not include a permission required for this route.` | A scoped console API key lacks permission for the route you called (for example `GET /api/v1/ai/keys` or `GET /api/v1/ai/enabled`). | Use an unrestricted console API key or a signed-in dashboard session, or call a route your key's scopes include. |
| `API keys cannot create API keys` | A console API key was used for `POST /api/v1/ai/keys`. | Create inference keys in the dashboard while signed in, or use an interactive session—not a console API key. |
| `API keys cannot manage account credentials` | A console API key was used for `DELETE /api/v1/ai/keys/{key_id}`. | Revoke inference keys in the dashboard while signed in. |
| `AI_KEY_MINT_REQUIRES_KYC_OR_FUNDS` | Organization is not KYC-verified and has no qualifying paid wallet top-up. | Complete KYC or add a direct paid top-up (promotional or bonus credits do not qualify), then retry key creation or playground access. |
| `Permission denied: 'ai:keys' required` | Account lacks permission to create inference keys. | Request the `ai:keys` permission from an organization administrator. |
| `AI Gateway is not enabled` | The AI Gateway service is not active in this environment. | Contact support or check service status announcements. |

## FAQ

**Do I need a separate SDK to use CloudPe AI Inference?**
No. Any standard OpenAI-compatible client library works by setting the base URL to `https://inferapi.cloudpe.com/v1`.

**What is the minimum billing unit?**
Token consumption is rated per 1M tokens in INR down to the individual paisa.

**Where does data processing occur?**
Model execution and gateway processing run in CloudPe data centres in India. The public endpoint `https://inferapi.cloudpe.com/v1` is fronted by Cloudflare for TLS and DDoS protection, so traffic passes through Cloudflare's edge before reaching CloudPe infrastructure.

## Related

- [Creating and managing AI API keys](/help/ai-api-keys)
- [Supported AI models and token pricing](/help/ai-inference-models-pricing)
- [AI inference rate limits and spend caps](/help/ai-inference-limits)
- [AI inference error codes and retries](/help/ai-inference-errors)