Price Plans

A Price Plan defines the rates applied to metered usage — a per-second rate for each GPU instance type behind a dedicated deployment, and per-million-token rates for hosted models. All billing accounts are assigned to the active default price plan.

Understanding Rates

Each rate entry specifies:

Field Description
meter The billing meter — gpu_instance_seconds for dedicated deployments, hosted_tokens for hosted models
unit second
unit_price Cost per second in USD
dimensions Instance type and region this rate applies to

Example Rate

An g5.xlarge at AWS on-demand pricing (~$1.006/hr) would have a per-second rate of approximately:

$1.006 / 3600 ≈ $0.000279 per second

GPU Instance Reference Pricing

The table below covers the instance types Xinference actually launches, as listed in the platform catalog. Where an hourly figure is given it is an approximate AWS on-demand list price used as a baseline for the default price plan; where it is not, check the AWS pricing page rather than assuming.

The full set in the catalog is g4dn.xlarge, g4dn.2xlarge, g4dn.4xlarge, g4dn.8xlarge, g4dn.12xlarge, g4dn.16xlarge, g5.xlarge, g5.2xlarge, g6.2xlarge, g7e.24xlarge and p4de.24xlarge. It moves between releases — the console shows the instance recommended for your model when you launch it.

Instance Type GPU VRAM Approx. $/hr
g4dn.xlarge T4 16 GB ~$0.526
g5.xlarge A10G 24 GB ~$1.006
g5.2xlarge A10G 24 GB ~$1.212
g4dn.12xlarge 4× T4 64 GB see AWS pricing
g6.2xlarge L4 24 GB see AWS pricing
p4de.24xlarge 8× A100 80 GB 640 GB see AWS pricing
g7e.24xlarge RTX PRO 6000 Blackwell — see AWS pricing

Note

These are reference prices. Contact your account manager for current Xinference sell prices.

Default Plan

On first startup, Xinference seeds a default price plan (aws-ondemand-default) derived from AWS on-demand list prices. This is a placeholder — confirm sell prices should be loaded via the admin CLI before going live.

Changing Prices

Prices are immutable once published. To change rates:

  1. Create a new price plan with a new code (e.g. aws-ondemand-2025-q3).
  2. Import it via the CLI or admin API.
  3. All billing accounts are automatically pointed to the new plan.
  4. Already-rated usage retains its historical rate snapshot.

CLI (Self-Hosted)

python -m app.cli.import_price_plan path/to/prices.json

JSON format:

{
  "code": "aws-ondemand-2025-q3",
  "name": "AWS On-Demand Q3 2025",
  "currency": "USD",
  "rates": [
    {
      "meter": "gpu_instance_seconds",
      "unit": "second",
      "unit_price": "0.000280",
      "dimensions": {
        "instance_type": "g5.xlarge",
        "region": "us-east-1"
      }
    }
  ]
}

Pass --no-activate to import without switching accounts to the new plan immediately.

Viewing the Current Plan

Navigate to Billing → Price Plan in the dashboard to see the rates currently applied to your account.