---
title: "Browsing AI models and using the playground"
slug: "ai-models-and-playground"
source: "https://app.cloudpe.com/help/ai-models-and-playground"
updated: "2026-10-08T05:09:58.536Z"
---

# Browsing AI models and using the playground

## Overview

CloudPe AI Gateway provides a managed, serverless inference service exposing an OpenAI-compatible API endpoint at `https://inferapi.cloudpe.com/v1` backed by high-performance GPU infrastructure. You can run popular open weights chat, completion, and reasoning models without provisioning or managing virtual machines.

This managed service provides shared model endpoints with per-token pricing where you pay only for the prompt, cached, and completion tokens your applications consume. Supported models include Llama 3.1 8B, Gemma 3 27B, Qwen3 32B, and DeepSeek R1 Distill Qwen 32B.

The model catalogue lets you browse supported models, inspect context window limits and capabilities, review token pricing, and copy ready-to-use code snippets. You can also experiment with models in the interactive browser playground before integrating them into applications.

## Before you start

- You need an active CloudPe account and an organization membership. See [Creating your CloudPe account](/help/account-signup-onboarding).
- Viewing the model catalogue and using the interactive playground requires the `ai:use` permission.
- Minting and managing persistent inference API keys requires the `ai:keys` permission. See [Creating and managing AI API keys](/help/ai-api-keys).
- Viewing organization token usage and spend requires the `ai:usage` permission. See [Tracking AI usage and billing](/help/ai-usage-and-billing).
- Account eligibility: Your organization must be KYC-verified or funded with a direct paid top-up (promotional or bonus credits do not qualify) to mint inference keys or obtain a playground session token.

## Steps

### Browse available models

1. In the sidebar, open **Models** under **AI**.
2. Review the list of available models. Each card displays:
   - Model name, slug identifier, and capability tags.
   - Context window length.
   - Per-token pricing for input and output tokens.
3. Use the code snippet selector on any model card to copy ready-to-use request examples for curl, Python, or Node.js.

### Test a model in the playground

1. In the sidebar, open **Playground** under **AI**.
2. Choose a model from the model dropdown selector.
3. Optionally enter custom system instructions in the system prompt field.
4. Adjust runtime parameters such as temperature and max tokens, and choose whether to enable streaming responses.
5. Enter your prompt in the message box and submit the request.
6. View the streamed response and review token consumption reported in the configuration pane.

## API

Retrieve available models or mint temporary playground session tokens programmatically via the console management API. These endpoints require a console API key with the `ai:use` scope (or an active dashboard session). Inference keys carry only `inference:invoke` and authenticate to `https://inferapi.cloudpe.com/v1` for model inference—they cannot call these management routes. See [Creating and managing AI API keys](/help/ai-api-keys).

| Method and path | Permission |
|---|---|
| `GET /api/v1/ai/models` | `ai:use` |
| `POST /api/v1/ai/playground/token` | `ai:use` |

List active models and live token prices:

```bash
curl https://app.cloudpe.com/api/v1/ai/models \
  -H "Authorization: Bearer <CONSOLE_API_KEY>"
```

Mint a temporary browser session token for the playground:

```bash
curl -X POST https://app.cloudpe.com/api/v1/ai/playground/token \
  -H "Authorization: Bearer <CONSOLE_API_KEY>"
```

## Limits & billing

- Token consumption is metered by the inference engine on every request and aggregated into hourly usage intervals.
- Prepaid organizations are debited automatically from wallet balances, while postpaid organizations receive line items on monthly invoices.
- Playground sessions use temporary browser tokens that auto-renew during active sessions and do not require creating a persistent API key.
- Playground conversations are processed in memory for the active session.
- Model execution and gateway processing run in CloudPe data centres in India; the public endpoint is fronted by Cloudflare (see [AI inference reliability and architecture](/help/ai-inference-reliability)). Billing records token counts, model identifier, timing, and request status for usage metering.

## Troubleshooting

| Message | What it means | What to do |
|---|---|---|
| `AI Gateway is not enabled` | The AI Gateway service is not active in this environment. | Contact support or check back later once the service is enabled. |
| `Permission denied: an AI permission is required` | Your account role lacks the required AI permissions. | Request the `ai:use` permission from your organization administrator. |

## FAQ

**What is the difference between AI Gateway and dedicated inference endpoints?**
AI Gateway is a serverless token API where you pay per token consumed. Legacy dedicated inference endpoints ran on dedicated GPU instances billed at continuous hourly compute rates regardless of request volume.

**What usage data does CloudPe record for billing?**
Billing and usage metering record token counts, model identifier, request timing, and status. For architecture and network path, see [AI inference reliability and architecture](/help/ai-inference-reliability).

**Can I use standard OpenAI client libraries with CloudPe AI Gateway?**
Yes. Point the official OpenAI Python or Node.js SDK base URL to `https://inferapi.cloudpe.com/v1` and supply an inference API key.

## Related

- [Supported AI models and token pricing](/help/ai-inference-models-pricing)
- [Creating and managing AI API keys](/help/ai-api-keys)
- [Tracking AI usage and billing](/help/ai-usage-and-billing)
- [AI inference quickstart](/help/ai-inference-quickstart)