Every GPU. One API. Priced by the market.
Bring idle GPUs and cloud API keys to the marketplace — or tap into hundreds of independent providers through one unified endpoint.
How it works
Two sides. One marketplace.
Providers list OpenAI-compatible capacity. Developers top up Air Credits and call one Air model ID. Air always uses that provider and meters exact PAYG usage.
Get an air_ API key
Sign up, finish onboarding, and copy your key once. Point the OpenAI SDK base URL at Air.
Enable models + buy Credits
Open a marketplace listing, enable the models you need, then top up Air Credits.
Call /api/v1
Send one Air model ID per Chat Completions request. Usage is charged from Credits.
For providers
List your compute. Earn from idle capacity.
Create a listing, add model offers, validate your endpoint, then go live. Developers enable your Air model IDs and pay with Credits.
OpenAI-compatible endpoints
vLLM, llama.cpp server, RunPod wrappers — anything that speaks /chat/completions.
List any capacity
One listing can include multiple models. Each gets its own Air ID and PAYG rates.
You set the rates
Compete on $/M tokens. Developers enable models on your listing page.
Usage & earnings
Track routed traffic and accrued net from the dashboard. Connect payouts come later.
Provider preview
Live listing
PAYG rates · Credits · routed /api/v1 traffic
Requests
2.4M
Fill rate
94%
Avg $/M tok
$0.41
1import OpenAI from "openai";2 3const client = new OpenAI({4 apiKey: process.env.AIR_API_KEY,5 baseURL: "https://www.airinference.com/api/v1",6});7 8const completion = await client.chat.completions.create({9 model: "air_gpt4o_pa_7K4D",10 messages: [{ role: "user", content: "Hello, Air." }],11});For developers
One API. Every model. Marketplace pricing.
Point your SDK at Air with baseURL. Open a listing, enable models, top up Credits, then call each Air model ID.
- OpenAI-compatible chat completions
- Browse listings, enable models you need
- One Air model ID per request
Marketplace pulse
0+
Active Providers
0+
Models Available
0M+
Requests served
0%
Avg. Savings vs Direct
Live listings
Models priced by competing hosts
Rates tick as providers compete. Same model family, different hosts — pick one or let the router choose.
Llama 3.3 70B
Claude 3.5
GPT-4o class
Mistral Large
DeepSeek V3
Qwen 2.5 72B
Gemini Flash
Command R+
Pricing
Competition beats fixed rate cards
Single-platform pricing locks you to one inventory. Marketplace hosts compete — rates move with supply.
From the marketplace
What builders and operators say
Real workflows from both sides — topping up Air Credits, listing OpenAI-compatible capacity, and calling exact Air model IDs with one API key.
Developer
“We pointed our OpenAI SDK at Air, enabled models on a listing, topped up Credits, and shipped the same night.”
Provider
“Two idle H100s were sitting cold after a cancelled contract. Listed an OpenAI-compatible endpoint Friday afternoon; by Tuesday PAYG Credits traffic was covering colo power for the month.”
Developer
“Pay-as-you-go Credits are boring in a good way — we see exact spend per request. Spending limits saved us from a silent overage when an agent loop got chatty.”
Provider
“Transparent $/M rates are where the real volume shows up. A buyer enabled our model, topped up Credits, and traffic hit our vLLM box within an hour.”
Developer
“Enabling a specific Air model ID for latency-sensitive work is the split we wanted. One air_ key, one provider per ID.”
Provider
“Credentials stay encrypted and buyers never see our upstream key. Going live without that would have been a non-starter for our compliance review.”
Early access
Join the compute marketplace
Whether you have spare capacity or need market-priced inference — create an account. Dual-path from day one.