The compute marketplace for AI inference

Every GPU. One API. Priced by the market.

Bring idle GPUs and cloud API keys to the marketplace — or tap into hundreds of independent providers through one unified endpoint.

G24.9on G2
Capterra4.7on Capterra
TrustRadius4.8on TrustRadius
Route any model family
Meta LlamaLlama
Anthropic ClaudeClaude
OpenAI GPTGPT
MistralMistral
DeepSeekDeepSeek
QwenQwen
Google GeminiGemini
llama-3.3-70b$0.4248mslive…·
claude-3.5-sonnet$2.08112msok…·
gpt-4o$1.8289mslive…·
mistral-large$0.6561msok…·
deepseek-v3$0.2754mslive…·
qwen-2.5-72b$0.2971msok…·
gemini-2.0-flash$0.1438mslive…·
llama-3.3-70b$0.4248mslive…·
claude-3.5-sonnet$2.08112msok…·
gpt-4o$1.8289mslive…·
mistral-large$0.6561msok…·
deepseek-v3$0.2754mslive…·
qwen-2.5-72b$0.2971msok…·
gemini-2.0-flash$0.1438mslive…·

How it works

Two sides. One marketplace.

Providers list OpenAI-compatible capacity. Developers top up Air Credits and call one Air model ID. Air always uses that provider and meters exact PAYG usage.

01

Get an air_ API key

Sign up, finish onboarding, and copy your key once. Point the OpenAI SDK base URL at Air.

02

Enable models + buy Credits

Open a marketplace listing, enable the models you need, then top up Air Credits.

03

Call /api/v1

Send one Air model ID per Chat Completions request. Usage is charged from Credits.

For providers

List your compute. Earn from idle capacity.

Create a listing, add model offers, validate your endpoint, then go live. Developers enable your Air model IDs and pay with Credits.

OpenAI-compatible endpoints

vLLM, llama.cpp server, RunPod wrappers — anything that speaks /chat/completions.

List any capacity

One listing can include multiple models. Each gets its own Air ID and PAYG rates.

You set the rates

Compete on $/M tokens. Developers enable models on your listing page.

Usage & earnings

Track routed traffic and accrued net from the dashboard. Connect payouts come later.

Provider preview

Live listing

PAYG rates · Credits · routed /api/v1 traffic

Live

Requests

2.4M

Fill rate

94%

Avg $/M tok

$0.41

inference.ts
OpenAI-compatible
1import OpenAI from "openai";
2 
3const client = new OpenAI({
4 apiKey: process.env.AIR_API_KEY,
5 baseURL: "https://www.airinference.com/api/v1",
6});
7 
8const completion = await client.chat.completions.create({
9 model: "air_gpt4o_pa_7K4D",
10 messages: [{ role: "user", content: "Hello, Air." }],
11});

For developers

One API. Every model. Marketplace pricing.

Point your SDK at Air with baseURL. Open a listing, enable models, top up Credits, then call each Air model ID.

  • OpenAI-compatible chat completions
  • Browse listings, enable models you need
  • One Air model ID per request

Marketplace pulse

0+

Active Providers

0+

Models Available

0M+

Requests served

0%

Avg. Savings vs Direct

Live listings

Models priced by competing hosts

Rates tick as providers compete. Same model family, different hosts — pick one or let the router choose.

H100 Cluster

Llama 3.3 70B

$0.42/M tok
Azure OpenAI

Claude 3.5

$2.10/M tok
Bedrock

GPT-4o class

$1.85/M tok
RunPod

Mistral Large

$0.68/M tok
Lambda

DeepSeek V3

$0.27/M tok
Home GPU

Qwen 2.5 72B

$0.31/M tok
Vertex

Gemini Flash

$0.15/M tok
Custom Rig

Command R+

$0.55/M tok

Pricing

Competition beats fixed rate cards

Single-platform pricing locks you to one inventory. Marketplace hosts compete — rates move with supply.

Fixed platform pricingAir marketplace
Llama 70B$0.90 → $0.42/M tok
Mistral Large$1.20 → $0.68/M tok
GPT-class$2.50 → $1.85/M tok
Claude-class$3.00 → $2.10/M tok

From the marketplace

What builders and operators say

Real workflows from both sides — topping up Air Credits, listing OpenAI-compatible capacity, and calling exact Air model IDs with one API key.

Developer

“We pointed our OpenAI SDK at Air, enabled models on a listing, topped up Credits, and shipped the same night.”

Maya ChenStaff engineer · Northline AIUsing Air for 4 months

Provider

“Two idle H100s were sitting cold after a cancelled contract. Listed an OpenAI-compatible endpoint Friday afternoon; by Tuesday PAYG Credits traffic was covering colo power for the month.”

Derek OkonkwoIndependent GPU operator · Okoye Compute2 live listings

Developer

“Pay-as-you-go Credits are boring in a good way — we see exact spend per request. Spending limits saved us from a silent overage when an agent loop got chatty.”

Elena VasquezPlatform lead · Brightfield LabsAgent workloads

Provider

“Transparent $/M rates are where the real volume shows up. A buyer enabled our model, topped up Credits, and traffic hit our vLLM box within an hour.”

Jonah ParkInfra engineer · Cascade NodesvLLM · us-west

Developer

“Enabling a specific Air model ID for latency-sensitive work is the split we wanted. One air_ key, one provider per ID.”

Aisha RahmanML engineer · Helix HealthExact Air IDs

Provider

“Credentials stay encrypted and buyers never see our upstream key. Going live without that would have been a non-starter for our compliance review.”

Marcus WebbSecurity-minded operator · Webb GPU Co-opCompliance-sensitive

Early access

Join the compute marketplace

Whether you have spare capacity or need market-priced inference — create an account. Dual-path from day one.