Migrating from legacy inference endpoints to serverless API

Last updated 8 Oct 2026
View as Markdown

Overview

The legacy GPU-VM based Inference Endpoints feature (accessed via /dashboard/inference, using cpe-inf- keys) is officially deprecated and retired. Creation of new dedicated GPU endpoints is frozen, and the deployment wizard is hidden in the dashboard navigation.

Sunset date: to be announced. Existing dedicated endpoints continue operating and accruing hourly VM billing until deleted by the user or decommissioned at final sunset. All customers are encouraged to migrate workloads to the serverless CloudPe AI Inference API hosted at https://inferapi.cloudpe.com/v1.

Before you start

  • Deprecation status: Dedicated GPU endpoints created under the legacy architecture cannot be scaled or reconfigured.
  • Account eligibility: Your organization must be KYC-verified or funded via a direct paid wallet top-up (promotional credits and sign-up bonuses do not qualify). Eligibility is checked when minting API keys or playground session tokens.
  • Permissions: Generating serverless inference keys requires the ai:keys permission. Deleting legacy endpoints requires vms:delete.
  • Documentation: Review serverless connection samples in AI inference quickstart.

Steps

Step 1: Create a serverless inference API key

  1. In the sidebar, open API Keys under AI.
  2. Select Create key.
  3. Provide a key name (for example, migrated-service-key).
  4. Optionally configure RPM, TPM, concurrency, or monthly spend caps.
  5. Select Create and copy the secret key token starting with cpk_.
  6. Select Done to close the dialog.

Step 2: Update application configuration

Update your client configuration settings:

  1. Replace the legacy dedicated base URL (<dedicated_endpoint_ip>:8000/v1) with the unified serverless endpoint:
    https://inferapi.cloudpe.com/v1
    
  2. Replace your legacy cpe-inf- key with your new cpk_ API key in the Authorization: Bearer <API_KEY> header.
  3. Update model identifiers to supported serverless model slugs (for example, llama-3-1-8b, gemma-3-27b, or qwen3-32b). See Supported AI models and token pricing.

Step 3: Verify serverless responses

Run a test completion request against the serverless endpoint:

curl -X POST https://inferapi.cloudpe.com/v1/chat/completions \
  -H "Authorization: Bearer $CLOUDPE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama-3-1-8b",
    "messages": [{"role": "user", "content": "Ping"}]
  }'

Step 4: Decommission legacy dedicated endpoints

Once traffic is successfully rerouted to the serverless endpoint, delete your legacy GPU VM endpoints:

  1. Open Legacy Inference under PLATFORM (shown only while you still have endpoints) in the sidebar.
  2. Select the legacy endpoint from the list.
  3. Select delete and confirm. Decommissioning stops continuous hourly compute billing for the associated GPU VM.

API

Serverless endpoint replaces legacy per-VM floating IPs:

Architecture Base URL Auth prefix Billing model
Legacy dedicated <floating_ip>:8000/v1 cpe-inf- Continuous hourly GPU VM
Serverless API https://inferapi.cloudpe.com/v1 Inference key (inference:invoke) Per-token consumed

Limits & billing

  • Pay per token: Serverless inference eliminates the need to pay for idle GPU VM compute time. You are billed strictly for input and output tokens consumed by your requests.
  • No infrastructure maintenance: Serverless removes VM lifecycle management, disk resizing, floating IP allocations, and custom vLLM parameter tuning.
  • Availability: Serverless models run in Zone B with elastic replicas (min_replicas=0, max_replicas=1). Cold starts and single-replica limits apply; see AI inference reliability and architecture.

Troubleshooting

Message What it means What to do
key not found The key identifier queried does not exist or was deleted. Verify the key identifier in the API keys table.
Permission denied: 'ai:keys' required Your account lacks permissions to manage inference keys. Request the ai:keys permission from an organization administrator.

FAQ

Can I deploy custom HuggingFace weights on the serverless API? The serverless API serves curated, high-performance open weights models. If your workload requires bespoke proprietary fine-tuned weights, provision a dedicated GPU instance. See Deploying GPU virtual machines.

What happens to my legacy endpoints if I do not migrate immediately? Sunset date: to be announced. Legacy endpoints remain online during the transition window, continuing hourly VM billing until deleted by the user or retired at sunset.

Do serverless API keys expire? Serverless keys remain active indefinitely until explicitly revoked by an administrator.

Related

Did this guide answer your question?If you need customized assistance with your deployment, reach out to our team.
Contact Support