Price Plans
A Price Plan defines the rates applied to metered usage — a per-second rate for each GPU instance type behind a dedicated deployment, and per-million-token rates for hosted models. All billing accounts are assigned to the active default price plan.
Understanding Rates
Each rate entry specifies:
| Field | Description |
|---|---|
meter |
The billing meter — gpu_instance_seconds for dedicated deployments, hosted_tokens for hosted models |
unit |
second |
unit_price |
Cost per second in USD |
dimensions |
Instance type and region this rate applies to |
Example Rate
An g5.xlarge at AWS on-demand pricing (~$1.006/hr) would have a per-second rate of approximately:
$1.006 / 3600 ≈ $0.000279 per second
GPU Instance Reference Pricing
The table below covers the instance types Xinference actually launches, as listed in the platform catalog. Where an hourly figure is given it is an approximate AWS on-demand list price used as a baseline for the default price plan; where it is not, check the AWS pricing page rather than assuming.
The full set in the catalog is g4dn.xlarge, g4dn.2xlarge, g4dn.4xlarge, g4dn.8xlarge, g4dn.12xlarge, g4dn.16xlarge, g5.xlarge, g5.2xlarge, g6.2xlarge, g7e.24xlarge and p4de.24xlarge. It moves between releases — the console shows the instance recommended for your model when you launch it.
| Instance Type | GPU | VRAM | Approx. $/hr |
|---|---|---|---|
g4dn.xlarge |
T4 | 16 GB | ~$0.526 |
g5.xlarge |
A10G | 24 GB | ~$1.006 |
g5.2xlarge |
A10G | 24 GB | ~$1.212 |
g4dn.12xlarge |
4× T4 | 64 GB | see AWS pricing |
g6.2xlarge |
L4 | 24 GB | see AWS pricing |
p4de.24xlarge |
8× A100 80 GB | 640 GB | see AWS pricing |
g7e.24xlarge |
RTX PRO 6000 Blackwell | — | see AWS pricing |
Note
These are reference prices. Contact your account manager for current Xinference sell prices.
Default Plan
On first startup, Xinference seeds a default price plan (aws-ondemand-default) derived from AWS on-demand list prices. This is a placeholder — confirm sell prices should be loaded via the admin CLI before going live.
Changing Prices
Prices are immutable once published. To change rates:
- Create a new price plan with a new code (e.g.
aws-ondemand-2025-q3). - Import it via the CLI or admin API.
- All billing accounts are automatically pointed to the new plan.
- Already-rated usage retains its historical rate snapshot.
CLI (Self-Hosted)
python -m app.cli.import_price_plan path/to/prices.json
JSON format:
{
"code": "aws-ondemand-2025-q3",
"name": "AWS On-Demand Q3 2025",
"currency": "USD",
"rates": [
{
"meter": "gpu_instance_seconds",
"unit": "second",
"unit_price": "0.000280",
"dimensions": {
"instance_type": "g5.xlarge",
"region": "us-east-1"
}
}
]
}
Pass --no-activate to import without switching accounts to the new plan immediately.
Viewing the Current Plan
Navigate to Billing → Price Plan in the dashboard to see the rates currently applied to your account.