AI inference rate limits and spend caps

Last updated 8 Oct 2026
View as Markdown

Overview

The CloudPe AI Inference API enforces request rate limits and concurrency caps to protect shared infrastructure and keep latency predictable under normal load. Rate limits operate automatically across distributed gateway nodes in Indian data centres.

In addition to default organization-level limits, administrators can set per-key rate limits and configure monthly spend caps to prevent unintended budget overruns.

Before you start

  • Account eligibility: Your organization must be KYC-verified or funded with a direct paid top-up (promotional or bonus credits do not qualify).
  • Permissions: Modifying inference key limits or configuring budget caps requires the ai:keys permission.
  • Monitoring: Review your organization's spend and usage history in the dashboard. See Tracking AI usage and billing.

Steps

Configure rate limits and budget caps for a key

  1. In the sidebar, open API Keys under AI.
  2. Locate the key and select Edit limits.
  3. Configure desired thresholds:
    • Requests per minute (RPM)
    • Tokens per minute (TPM)
    • Maximum concurrency
    • Monthly spend cap in INR
  4. Select Save to apply the changes immediately.

Request a limit increase

Organizations requiring higher throughput can request increased quotas:

  1. In the sidebar, open Support.
  2. Select New Ticket and name the AI Inference API in the subject or description.
  3. Include your organization identifier, target model slugs, required requests per minute, peak concurrency requirements, and estimated monthly token consumption.
  4. The operations team reviews and provisions custom quota tiers upon verification.

API

Inspect and adjust rate limits programmatically via the console management API. PATCH accepts a scoped console API key with ai:keys or a dashboard session. GET /api/v1/ai/keys requires an unrestricted console API key or a dashboard session today (scoped ai:keys keys return 403). Inference keys carry only inference:invoke and authenticate to https://inferapi.cloudpe.com/v1 for model inference—they cannot call these management routes. See Creating and managing AI API keys.

Method and path Permission
GET /api/v1/ai/keys ai:keys (unrestricted console key or session)
PATCH /api/v1/ai/keys/{key_id} ai:keys

Update key limits via API (key limits must stay within the organization limits, which default to 60 RPM and 4 concurrent requests):

curl -X PATCH https://app.cloudpe.com/api/v1/ai/keys/<key_id> \
  -H "Authorization: Bearer <CONSOLE_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
        "rpm": 60,
        "concurrency": 4,
        "monthly_cap": 5000.00
      }'

Limits & billing

Default quotas

Every organization receives standard baseline quotas upon activation:

Limit dimension Default organization quota Scope
Request rate limit 60 requests per minute Organization-wide
Concurrency limit 4 concurrent requests Organization-wide
Token rate limit Uncapped by default Organization-wide
Monthly spend cap Configurable (defaults apply for postpaid) Organization-wide

Budget alerts and spend enforcement

  • Spend tracking: Spend is calculated in 10-minute rating intervals from usage records. Budget alert evaluations run every 15 minutes.
  • Alert thresholds: Automated email alerts and webhooks trigger when organization or key spend crosses 50 percent, 80 percent, and 100 percent of the configured monthly budget cap.
  • Cap enforcement: When an organization or key reaches 100 percent of its monthly budget cap, further inference requests are halted with an HTTP 429 response until the cap is increased or the new billing cycle begins.
  • Billing cycle reset: Monthly budget caps reset on the UTC calendar month boundary (00:00 UTC), not on a rolling window.
  • Rate limit headers: When a request exceeds RPM or concurrency limits, the response includes a Retry-After header indicating the waiting interval before retrying.

Troubleshooting

Message What it means What to do
key not found The specified key identifier does not exist or was deleted. Refresh the API keys list or verify the key UUID.
Permission denied: 'ai:keys' required Your role lacks permissions to update key limits. Request the ai:keys permission from an organization administrator.

FAQ

How quickly do updated key limits take effect? Limit changes and key updates propagate to gateway nodes within about a minute (gateway authorization snapshots refresh on a 30-second interval).

Do unauthenticated requests count against organization rate limits? No. Unauthenticated requests are rejected at edge gateways before reaching organization limit counters.

Can individual keys have higher limits than the organization default? No. Gateway policies enforce key and organization limits independently. An individual key cannot exceed the organization ceiling (default 60 RPM and 4 concurrency) unless the organization quota itself is upgraded.

Related

Did this guide answer your question?If you need customized assistance with your deployment, reach out to our team.
Contact Support