Xinference Documentation
Xinference is a cloud platform for deploying and serving open-source AI models at scale. With a single API call you can launch large language models (LLMs) and embedding models on dedicated GPU instances — billed per second, no infrastructure management required.
Start here
Pick the task you are trying to complete.
Launch a model on a GPU instance and query an OpenAI-compatible endpoint in minutes.
Choose a model →Browse the hosted and dedicated model catalogs, and see which GPU instances back them.
Call the inference API →Send chat completion and embedding requests to a running deployment using OpenAI-style clients.
Manage deployments →Follow a deployment through its lifecycle, read its status, and terminate it when you are done.
Understand what you pay →Per-second postpaid billing on the GPU instances behind your deployments — how it is measured.
Set up your team →Invite teammates, assign roles, and share one billing account across an organization.
Look up an endpoint →Reference for the management API: deployments, models, organizations, and billing.
Run it yourself →Self-host Xinference on your own infrastructure with Docker or AWS.
How Xinference works
Xinference provisions GPU-backed inference clusters on demand. When you create a deployment, the platform:
- Selects the optimal GPU instance type for your chosen model
- Provisions the EC2 cluster in the cloud
- Downloads the model weights (or restores them from cache)
- Starts the Xinference inference server
- Returns an OpenAI-compatible endpoint you can query immediately
When you terminate a deployment the cluster is torn down and billing stops.
Core concepts
| Concept | Description |
|---|---|
| Deployable Model | A model variant pre-approved for deployment (e.g. qwen3.5-0.8b on a specific GPU family). |
| Deployment | A running instance of a deployable model assigned to your account, with its own endpoint URL. |
| Cluster | The set of EC2 instances (supervisor + workers) backing a deployment. |
| Organization | A group of users sharing billing and deployment quotas. |
| Billing Account | Tracks usage and balance for an organization. Metered in GPU instance-seconds. |
Supported model types
- LLM (Chat) — instruction-tuned language models for chat and text generation
- Embedding — text embedding models for semantic search and RAG
Key features
- OpenAI-compatible API — drop-in replacement for
openai.ChatCompletionandopenai.Embedding - Per-second billing — pay only for the time your model is running
- Model caching — popular models cached on S3 for fast cold-start times
- Multi-user organizations — invite teammates and share a single billing account
- SSO support — sign in with Google
- Idle auto-termination — clusters shut down automatically after a configurable idle period
Looking for the open-source engine?
These pages cover Xinference Cloud — the hosted platform. For the core inference engine itself (model formats, engines, cluster internals, and the Python API), see the community documentation.