Billing

Xinference bills postpaid, with no upfront commitments or minimum charges. How usage is measured depends on which kind of model you use — hosted models are billed per token, dedicated deployments per GPU-second.

How Billing Works

There are two billing meters, and which one applies is decided by the kind of model, not by how you call it:

Model kind Meter You pay for
Hosted hosted_tokens Tokens processed — input and output are priced separately, per million tokens. Nothing runs when you are idle, so nothing accrues.
Dedicated gpu_instance_seconds Wall-clock time on the GPU instances behind your deployment, per second — regardless of how many requests you send.

Dedicated deployments

  1. A deployment starts — the billing clock starts when cluster provisioning begins.
  2. GPU instances run — each instance type has a per-second rate.
  3. A deployment terminates — billing stops when the cluster is fully torn down (status: terminated).

Warning

A dedicated deployment is billed for the time it is running, not for the requests it serves. One sitting idle still costs the full instance rate — terminate it when you are done.

Billing Accounts

Every organization has a Billing Account that aggregates usage across all members' deployments. Individual users (without an organization) have a personal billing account.

Billing accounts operate in postpaid mode by default. Usage is tallied and invoiced at the end of each billing period.

Metering

For dedicated deployments, usage is measured as GPU instance-seconds — the product of the number of instances and the duration they run, broken down by instance type. Hosted models are metered in tokens instead.

Example: - 1× g5.xlarge (A10G, 24 GB GPU) running for 1 hour = 3,600 instance-seconds

Each instance type has a separate per-second rate. See Price Plans →.

Viewing Usage

Dashboard

Navigate to Billing to see: - Current period usage summary - Per-deployment usage breakdown - Historical invoices

API

Authentication uses the session cookie set when you sign in (see Authentication).

import requests

r = requests.get(
    "https://api.xinference.co/v1/billing/usage-summary",
    cookies={"session": "<your-session-cookie>"},
)
print(r.json())

Cost Control Tips

  • Terminate idle deployments — the 15-minute idle auto-termination is a safety net, not a substitute for explicit termination.
  • Use quantized models — int4 models run on smaller (cheaper) instances than full-precision equivalents.
  • Monitor usage — set up billing alerts in Billing → Alerts to get notified when usage exceeds a threshold.
  • Choose the right model size — a 7B model is much cheaper per hour than a 70B model and sufficient for many tasks.

Next Steps