API · v1
API documentation
OpenAI-compatible Chat Completions over provider-specific Air model IDs. Enable a PAYG offer, top up Air Credits, then call https://www.airinference.com/api/v1.
Quickstart
- 1Create an
air_key in Dashboard → API keys. - 2Open the Marketplace, choose a provider offer, and enable PAYG.
- 3Copy that offer's Air model ID.
- 4Add Air Credits.
- 5Call
POST /chat/completionswith that model ID.
First request
curl https://www.airinference.com/api/v1/chat/completions \
-H "Authorization: Bearer $AIR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "air_gpt4o_pa_7K4D",
"messages": [{"role": "user", "content": "Hello from Air"}]
}'Air model IDs
Each Air model ID maps to exactly one provider offer (one provider + one upstream model). Air does not automatically switch providers.
Air model ID
↓
exact provider offer
↓
provider model ID
↓
provider endpoint
air_gpt4o_pa_7K4D
→ Provider A
→ upstream model: gpt-4oGet an ID: Marketplace → choose offer → Enable PAYG → Copy Air model ID.
Air model IDs are provider-specific and immutable. To use another provider for the same underlying model, enable that provider's separate Air model ID and send another request.
Authentication
Authorization: Bearer air_...Create keys in Dashboard → API keys. Keep the key server-side — do not expose it in browser or client-side code.
Base URL
https://www.airinference.com/api/v1Use this as baseURL / base_url with the OpenAI SDK. Paths below are relative to this base (API version /v1).
/chat/completions
Auth requiredOpenAI-compatible Chat Completions. Pass one enabled Air model ID. Extra OpenAI fields (for example temperature) are forwarded to the provider, except provider and listing_id, which are ignored.
Request
curl https://www.airinference.com/api/v1/chat/completions \
-H "Authorization: Bearer $AIR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "air_gpt4o_pa_7K4D",
"messages": [{"role": "user", "content": "Hello from Air"}]
}'Parameters
| Field | Required | Description |
|---|---|---|
| model | Yes | Enabled Air model ID for this account. |
| messages | Yes | Non-empty OpenAI-style chat messages array. |
| stream | No | If true, returns SSE when the offer supports streaming. |
| max_tokens | No | Forwarded upstream. Also used when estimating the Credits reservation. |
| tools | No | Forwarded when the offer supports tools; otherwise capability_unsupported. |
| response_format | No | json_object / json_schema require an offer with JSON support. |
Non-stream responses set model to the Air model ID you requested (not the upstream provider model string).
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "air_gpt4o_pa_7K4D",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello from Air" },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 4,
"total_tokens": 16
}
}Streaming
Set stream: true. Air proxies text/event-stream from the provider. Air requests stream_options.include_usage so token usage can be metered after the stream ends; Credits are finalized from that usage (reserve → capture).
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AIR_API_KEY,
baseURL: "https://www.airinference.com/api/v1",
});
const stream = await client.chat.completions.create({
model: "air_gpt4o_pa_7K4D",
messages: [{ role: "user", content: "Say hello" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}If the offer does not support streaming, the API returns 400 capability_unsupported.
/models
PublicPublic catalog of live provider-specific PAYG offers. No API key required. Each id is an Air model ID.
Request
curl https://www.airinference.com/api/v1/models{
"object": "list",
"data": [
{
"id": "air_gpt4o_pa_7K4D",
"object": "model",
"display_name": "gpt-4o",
"billing": "payg",
"provider_model_id": "gpt-4o",
"family": "gpt-4o",
"listing_title": "Provider A",
"listing_slug": "provider-a",
"region": "us-east",
"input_usd_per_1m": "0.18",
"output_usd_per_1m": "0.72",
"health_status": "healthy",
"capabilities": {
"streaming": true,
"tools": true,
"json": true,
"vision": false
}
}
]
}Fetch one offer: GET /models/{id} (same fields; 404 unknown_model if missing).
/me/models
Auth requiredLive offers plus enabled for this developer, and a Credits summary. Disabled rows appear so you can see what is available to enable.
curl https://www.airinference.com/api/v1/me/models \
-H "Authorization: Bearer $AIR_API_KEY"{
"object": "list",
"data": [
{
"id": "air_gpt4o_pa_7K4D",
"object": "model",
"enabled": true,
"display_name": "gpt-4o",
"billing": "payg",
"provider_model_id": "gpt-4o",
"listing_title": "Provider A",
"input_usd_per_1m": "0.18",
"output_usd_per_1m": "0.72",
"capabilities": { "streaming": true, "tools": true, "json": true, "vision": false }
}
],
"credits": {
"available": "25.00",
"reserved": "0",
"total": "25.00",
"currency": "usd"
}
}| Endpoint | Returns |
|---|---|
| /models | Public live offers |
| /me/models | Offers + enabled flag for this key |
Enable a model
Marketplace
→ choose provider / model offer
→ Enable PAYG
→ copy Air model IDEnabling one provider's model does not enable another provider's model, even when both offer the same underlying model.
air_gpt4o_pa_7K4D
does not enable
air_gpt4o_pb_92QFIf you enable both, choose which Air model ID to send per request. Air does not automatically pick between them or fail over.
Billing
PAYG uses Air Credits. Each chat request reserves an estimate, calls the provider, then charges actual usage and releases the unused hold.
Air Credits
→ temporary reservation
→ provider request
→ actual usage
→ final charge
→ unused reservation releasedReservations are temporary holds, not additional charges. Final billing uses measured tokens.
Reserved: $0.001625
Actual: $0.000029
Released: $0.001596
Final charge: $0.000029Pricing
Cost is computed from the offer's rate card (USD per 1M tokens). There is no cent-rounding floor on the charge — amounts use high-precision USD.
cost =
prompt_tokens × input_usd_per_1m / 1,000,000
+ completion_tokens × output_usd_per_1m / 1,000,000Historical usage keeps the rate card that applied at call time. New calls use the current active rate.
Wallet
GET /me/wallet returns available / reserved Credits. Top up in the dashboard. Optional per-account spending limits are enforced when configured. Auto-recharge may run after successful calls when enabled for your account.
curl https://www.airinference.com/api/v1/me/wallet \
-H "Authorization: Bearer $AIR_API_KEY"Errors
Errors use an error object. Important codes for /chat/completions:
| Status | Code | Meaning | What to do |
|---|---|---|---|
| 400 | invalid_json / missing_model / missing_messages | Malformed body | Fix JSON, model, or messages |
| 400 | capability_unsupported | stream / tools / vision / JSON not supported by offer | Disable that feature or pick another offer |
| 400 | reservation_cap_exceeded | Estimated cost above max reservation | Lower max_tokens or split the request |
| 401 | — | Invalid or revoked API key | Create a new air_ key |
| 402 | insufficient_credits | Not enough Air Credits | Top up Credits |
| 402 | spending_limit_exceeded | Account/key spending limit hit | Raise limit or wait for the window |
| 403 | model_not_enabled | Known Air ID not enabled | Enable on Marketplace |
| 404 | unknown_model | Unknown Air model ID | Copy ID from Marketplace /models |
| 409 | idempotency_conflict | Same Idempotency-Key, different body | Use a new key for a new request |
| 429 | — | Over 60 req/min per key | Backoff; check Retry-After |
| 502 | provider_failed | Upstream provider failed | Retry later or use another enabled Air ID |
| 503 | model_unavailable / api_paused / credits_paused | Offer or platform unavailable | Retry later; Air does not fail over |
| 504 | provider_timeout | Upstream timed out | Retry; unused reservation is released |
Examples
{
"error": {
"message": "Model is not enabled for this account. Enable it on the marketplace listing page.",
"code": "model_not_enabled",
"model": "air_gpt4o_pa_7K4D",
"marketplace": "https://www.airinference.com/dashboard/marketplace"
}
}{
"error": {
"message": "Insufficient Air Credits. Top up to continue.",
"type": "insufficient_quota",
"code": "insufficient_credits",
"model": "air_gpt4o_pa_7K4D"
},
"air": {
"available": "0.12",
"reserved": "0",
"topup": "https://www.airinference.com/dashboard/credits"
}
}Idempotency
Optional on POST /chat/completions:
Idempotency-Key: <unique-request-key>- Same key + same request body → safe replay of the stored response.
- Same key + different body →
409 idempotency_conflict.
Rate limits
60 requests / minute / API key across authenticated /api/v1 routes. When exceeded, the API returns 429 with a Retry-After header (seconds).
OpenAI SDK
Air provides an OpenAI-compatible API. Use the standard OpenAI SDK with Air's base URL and an Air model ID. Air supports the fields and behavior documented on this page — not every OpenAI product feature.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AIR_API_KEY,
baseURL: "https://www.airinference.com/api/v1",
});
const response = await client.chat.completions.create({
model: "air_gpt4o_pa_7K4D",
messages: [{ role: "user", content: "Hello from Air" }],
});
console.log(response.choices[0].message.content);Providers
Publish a PAYG offer from the dashboard:
Provider
→ create listing / model offer
→ enter provider model ID + PAYG rates
→ Base URL + upstream API key
→ Validate
→ set LiveValidation checks the endpoint, credentials, model availability, a small inference request, and response shape. After publication, Air monitors offer health. If that exact provider is unavailable, the API returns an error — Air does not silently switch providers.
API version
These docs describe the /api/v1 API. Paths such as /chat/completions are relative to https://www.airinference.com/api/v1.