AI inference quickstart
Overview
The CloudPe AI Inference API provides an OpenAI-compatible interface hosted at https://inferapi.cloudpe.com/v1. You can connect existing applications using official OpenAI SDKs in Python and Node.js or standard HTTP client tools like curl.
Model execution and gateway processing run in CloudPe data centres in India; the public endpoint is fronted by Cloudflare (see AI inference reliability and architecture).
Before you start
- CloudPe account: You need an active user account in an organization. See Creating your CloudPe account.
- Eligibility requirement: Your organization must be KYC-verified or funded with a direct paid top-up (promotional or bonus credits do not qualify) before you can mint inference keys or obtain a playground session token. After you have a valid inference key, calls to
https://inferapi.cloudpe.com/v1do not re-check this gate on every request. - Permission: Generating inference keys requires the
ai:keyspermission. Browsing models requiresai:use. - Models: Check available models and context lengths in the model catalogue. See Supported AI models and token pricing.
Steps
Create an inference key
- In the sidebar, open API Keys under AI.
- Select Create key to open the creation dialog.
- Enter a key name (for example,
app-quickstart). - Optionally configure RPM, TPM, concurrency, or monthly spend caps.
- Select Create.
- Copy the plaintext key starting with
cpk_. Store it in your local environment asCLOUDPE_API_KEY. - Select Done to close the modal.
Call the API using curl
Run a chat completion request with streaming enabled:
curl -X POST https://inferapi.cloudpe.com/v1/chat/completions \
-H "Authorization: Bearer $CLOUDPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3-1-8b",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain serverless computing in two sentences."}
],
"stream": true
}'
Call the API using OpenAI Python SDK
Install the official OpenAI package:
pip install openai
Initialize the client with the CloudPe base URL and execute a streaming completion:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://inferapi.cloudpe.com/v1",
api_key=os.environ.get("CLOUDPE_API_KEY"),
)
stream = client.chat.completions.create(
model="llama-3-1-8b",
messages=[
{"role": "system", "content": "You are a helpful coding assistant."},
{"role": "user", "content": "Write a quicksort implementation in Python."},
],
stream=True,
)
for chunk in stream:
# Final SSE chunk has usage only (empty choices) when include_usage is enabled.
if not chunk.choices:
continue
content = chunk.choices[0].delta.content or ""
print(content, end="", flush=True)
print()
Call the API using OpenAI Node.js SDK
Install the package:
npm install openai
Execute a streaming completion in TypeScript or JavaScript:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://inferapi.cloudpe.com/v1",
apiKey: process.env.CLOUDPE_API_KEY,
});
async function main() {
const stream = await client.chat.completions.create({
model: "llama-3-1-8b",
messages: [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Summarize the advantages of prompt caching." },
],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || "");
}
process.stdout.write("\n");
}
main();
API
Primary endpoints available on the inference host https://inferapi.cloudpe.com/v1:
| Endpoint | Method | Purpose |
|---|---|---|
https://inferapi.cloudpe.com/v1/chat/completions |
POST | Create chat completions (streaming and non-streaming) |
https://inferapi.cloudpe.com/v1/models |
GET | List models available to your organization |
Management endpoints on https://app.cloudpe.com:
| Method and path | Permission |
|---|---|
GET /api/v1/ai/keys |
ai:keys |
POST /api/v1/ai/keys |
ai:keys |
Limits & billing
- Default quotas: Organizations receive 60 requests per minute and 4 concurrent requests by default. See AI inference rate limits and spend caps.
- Metering: Charges apply per token consumed (prompt, cached, and completion). Token usage events stream from inference engines to billing services.
- Currency: Token consumption is billed in INR to the paisa.
- Billing schedule: Prepaid organizations are debited automatically from wallet balances. Postpaid accounts receive itemized lines on monthly invoices.
Troubleshooting
| Message | What it means | What to do |
|---|---|---|
API key does not have required AI scopes |
An inference API key was used for GET /api/v1/ai/usage, which requires the ai:usage scope. Inference keys carry only inference:invoke. |
Query usage with a console API key that includes ai:usage or with a signed-in dashboard session. |
API key scope does not include a permission required for this route. |
A scoped console API key lacks permission for the route you called (for example GET /api/v1/ai/keys or GET /api/v1/ai/enabled). |
Use an unrestricted console API key or a signed-in dashboard session, or call a route your key's scopes include. |
API keys cannot create API keys |
A console API key was used for POST /api/v1/ai/keys. |
Create inference keys in the dashboard while signed in, or use an interactive session—not a console API key. |
API keys cannot manage account credentials |
A console API key was used for DELETE /api/v1/ai/keys/{key_id}. |
Revoke inference keys in the dashboard while signed in. |
AI_KEY_MINT_REQUIRES_KYC_OR_FUNDS |
Organization is not KYC-verified and has no qualifying paid wallet top-up. | Complete KYC or add a direct paid top-up (promotional or bonus credits do not qualify), then retry key creation or playground access. |
Permission denied: 'ai:keys' required |
Account lacks permission to create inference keys. | Request the ai:keys permission from an organization administrator. |
AI Gateway is not enabled |
The AI Gateway service is not active in this environment. | Contact support or check service status announcements. |
FAQ
Do I need a separate SDK to use CloudPe AI Inference?
No. Any standard OpenAI-compatible client library works by setting the base URL to https://inferapi.cloudpe.com/v1.
What is the minimum billing unit? Token consumption is rated per 1M tokens in INR down to the individual paisa.
Where does data processing occur?
Model execution and gateway processing run in CloudPe data centres in India. The public endpoint https://inferapi.cloudpe.com/v1 is fronted by Cloudflare for TLS and DDoS protection, so traffic passes through Cloudflare's edge before reaching CloudPe infrastructure.

