AI inference rate limits and spend caps
Overview
The CloudPe AI Inference API enforces request rate limits and concurrency caps to protect shared infrastructure and keep latency predictable under normal load. Rate limits operate automatically across distributed gateway nodes in Indian data centres.
In addition to default organization-level limits, administrators can set per-key rate limits and configure monthly spend caps to prevent unintended budget overruns.
Before you start
- Account eligibility: Your organization must be KYC-verified or funded with a direct paid top-up (promotional or bonus credits do not qualify).
- Permissions: Modifying inference key limits or configuring budget caps requires the
ai:keyspermission. - Monitoring: Review your organization's spend and usage history in the dashboard. See Tracking AI usage and billing.
Steps
Configure rate limits and budget caps for a key
- In the sidebar, open API Keys under AI.
- Locate the key and select Edit limits.
- Configure desired thresholds:
- Requests per minute (RPM)
- Tokens per minute (TPM)
- Maximum concurrency
- Monthly spend cap in INR
- Select Save to apply the changes immediately.
Request a limit increase
Organizations requiring higher throughput can request increased quotas:
- In the sidebar, open Support.
- Select New Ticket and name the AI Inference API in the subject or description.
- Include your organization identifier, target model slugs, required requests per minute, peak concurrency requirements, and estimated monthly token consumption.
- The operations team reviews and provisions custom quota tiers upon verification.
API
Inspect and adjust rate limits programmatically via the console management API. PATCH accepts a scoped console API key with ai:keys or a dashboard session. GET /api/v1/ai/keys requires an unrestricted console API key or a dashboard session today (scoped ai:keys keys return 403). Inference keys carry only inference:invoke and authenticate to https://inferapi.cloudpe.com/v1 for model inference—they cannot call these management routes. See Creating and managing AI API keys.
| Method and path | Permission |
|---|---|
GET /api/v1/ai/keys |
ai:keys (unrestricted console key or session) |
PATCH /api/v1/ai/keys/{key_id} |
ai:keys |
Update key limits via API (key limits must stay within the organization limits, which default to 60 RPM and 4 concurrent requests):
curl -X PATCH https://app.cloudpe.com/api/v1/ai/keys/<key_id> \
-H "Authorization: Bearer <CONSOLE_API_KEY>" \
-H "Content-Type: application/json" \
-d '{
"rpm": 60,
"concurrency": 4,
"monthly_cap": 5000.00
}'
Limits & billing
Default quotas
Every organization receives standard baseline quotas upon activation:
| Limit dimension | Default organization quota | Scope |
|---|---|---|
| Request rate limit | 60 requests per minute | Organization-wide |
| Concurrency limit | 4 concurrent requests | Organization-wide |
| Token rate limit | Uncapped by default | Organization-wide |
| Monthly spend cap | Configurable (defaults apply for postpaid) | Organization-wide |
Budget alerts and spend enforcement
- Spend tracking: Spend is calculated in 10-minute rating intervals from usage records. Budget alert evaluations run every 15 minutes.
- Alert thresholds: Automated email alerts and webhooks trigger when organization or key spend crosses 50 percent, 80 percent, and 100 percent of the configured monthly budget cap.
- Cap enforcement: When an organization or key reaches 100 percent of its monthly budget cap, further inference requests are halted with an HTTP 429 response until the cap is increased or the new billing cycle begins.
- Billing cycle reset: Monthly budget caps reset on the UTC calendar month boundary (00:00 UTC), not on a rolling window.
- Rate limit headers: When a request exceeds RPM or concurrency limits, the response includes a Retry-After header indicating the waiting interval before retrying.
Troubleshooting
| Message | What it means | What to do |
|---|---|---|
key not found |
The specified key identifier does not exist or was deleted. | Refresh the API keys list or verify the key UUID. |
Permission denied: 'ai:keys' required |
Your role lacks permissions to update key limits. | Request the ai:keys permission from an organization administrator. |
FAQ
How quickly do updated key limits take effect? Limit changes and key updates propagate to gateway nodes within about a minute (gateway authorization snapshots refresh on a 30-second interval).
Do unauthenticated requests count against organization rate limits? No. Unauthenticated requests are rejected at edge gateways before reaching organization limit counters.
Can individual keys have higher limits than the organization default? No. Gateway policies enforce key and organization limits independently. An individual key cannot exceed the organization ceiling (default 60 RPM and 4 concurrency) unless the organization quota itself is upgraded.

