Tracking AI usage and billing

Last updated 8 Oct 2026
View as Markdown

Overview

The AI Usage dashboard provides transparent observability into token consumption, request counts, and financial spend across all AI Gateway workloads in your organization. You can inspect hourly token aggregations, breakdown costs across models, keys, users, projects, or calendar days, and export usage records for accounting reconciliation.

Token consumption is rated automatically to the paisa based on active per-model token pricing for prompt, cached, and completion tokens. Charges settle continuously against your prepaid wallet balance or accrue as itemized line items on your monthly postpaid tax invoice. Model execution and gateway processing run in CloudPe data centres in India; the public endpoint is fronted by Cloudflare (see AI inference reliability and architecture).

Before you start

Steps

View and filter AI token usage

  1. In the sidebar, open Usage under AI.
  2. Review the summary card at the top of the page displaying cumulative spend so far in the current billing cycle.
  3. Select an aggregation dimension from the group by selector:
    • Model (group by model slug)
    • Key (group by inference API key)
    • User (group by user account)
    • Project (group by project identifier)
    • Day (group by calendar date)
  4. Adjust the date range using the start and end date pickers to inspect historical windows.
  5. Review the breakdown table displaying:
    • Dimension grouping identifier
    • Prompt token count
    • Completion token count
    • Cached token count
    • Total request count
    • Total rated cost in currency units

Export usage data

  1. In the sidebar, open Usage under AI.
  2. Configure your desired grouping and date filters.
  3. Select the export option in the top action bar to download the filtered dataset as a CSV file.

API

Query token consumption and financial spend programmatically using the console management usage endpoint. This endpoint requires a console API key with the ai:usage scope (or an active dashboard session). Inference keys carry only inference:invoke and authenticate to https://inferapi.cloudpe.com/v1 for model inference—they cannot call this management route. See Creating and managing AI API keys.

Method and path Permission
GET /api/v1/ai/usage ai:usage

Query usage grouped by model for the current calendar month:

curl "https://app.cloudpe.com/api/v1/ai/usage?group_by=model" \
  -H "Authorization: Bearer <CONSOLE_API_KEY>"

Query usage grouped by day across a specific date window:

curl "https://app.cloudpe.com/api/v1/ai/usage?group_by=day&from=2026-09-01&to=2026-09-25" \
  -H "Authorization: Bearer <CONSOLE_API_KEY>"

Limits & billing

  • Requests processed by the AI Gateway emit token usage events containing prompt, completion, and cached token metrics.
  • Usage events are aggregated into intervals per organization, project, key, and model.
  • Intervals are rated every 10 minutes based on model price tiers active at the time of execution. Cached prompt tokens are billed at the model's standard input rate.
  • Prepaid organizations are debited automatically from wallet balances via hourly billing tasks.
  • Postpaid organizations accumulate rated intervals throughout the billing period and settle via consolidated monthly tax invoices.
  • Monthly spend caps configured per key or organization return HTTP 429 when budget limits are breached.

Troubleshooting

Message What it means What to do
Permission denied: an AI permission is required Your user account lacks the necessary AI permissions to access usage analytics. Request the ai:usage permission from your organization administrator.
Permission denied You do not have permission to view organization-wide billing records. Contact your organization owner to upgrade your role permissions.

FAQ

How often is AI token usage updated on the dashboard? Gateway nodes record token usage events, which are rated in 10-minute intervals and aggregated for reporting.

How does token caching affect my bill? When prompt caching is supported by a model, cached prompt tokens are billed at the model's standard input rate.

Where do AI inference charges appear on my invoice? Postpaid invoices include itemized line items under the AI inference category detailing total prompt, completion, and cached token usage per model.

Related

Did this guide answer your question?If you need customized assistance with your deployment, reach out to our team.
Contact Support