API · v1

API documentation

OpenAI-compatible Chat Completions over provider-specific Air model IDs. Enable a PAYG offer, top up Air Credits, then call https://www.airinference.com/api/v1.

Quickstart

  1. 1Create an air_ key in Dashboard → API keys.
  2. 2Open the Marketplace, choose a provider offer, and enable PAYG.
  3. 3Copy that offer's Air model ID.
  4. 4Add Air Credits.
  5. 5Call POST /chat/completions with that model ID.

First request

curl https://www.airinference.com/api/v1/chat/completions \
  -H "Authorization: Bearer $AIR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "air_gpt4o_pa_7K4D",
    "messages": [{"role": "user", "content": "Hello from Air"}]
  }'

Air model IDs

Each Air model ID maps to exactly one provider offer (one provider + one upstream model). Air does not automatically switch providers.

Air model ID
      ↓
exact provider offer
      ↓
provider model ID
      ↓
provider endpoint

air_gpt4o_pa_7K4D
→ Provider A
→ upstream model: gpt-4o

Get an ID: Marketplace → choose offer → Enable PAYG → Copy Air model ID.

Air model IDs are provider-specific and immutable. To use another provider for the same underlying model, enable that provider's separate Air model ID and send another request.

Authentication

Authorization: Bearer air_...

Create keys in Dashboard → API keys. Keep the key server-side — do not expose it in browser or client-side code.

Base URL

https://www.airinference.com/api/v1

Use this as baseURL / base_url with the OpenAI SDK. Paths below are relative to this base (API version /v1).

POST

/chat/completions

Auth required

OpenAI-compatible Chat Completions. Pass one enabled Air model ID. Extra OpenAI fields (for example temperature) are forwarded to the provider, except provider and listing_id, which are ignored.

Request

curl https://www.airinference.com/api/v1/chat/completions \
  -H "Authorization: Bearer $AIR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "air_gpt4o_pa_7K4D",
    "messages": [{"role": "user", "content": "Hello from Air"}]
  }'

Parameters

FieldRequiredDescription
modelYesEnabled Air model ID for this account.
messagesYesNon-empty OpenAI-style chat messages array.
streamNoIf true, returns SSE when the offer supports streaming.
max_tokensNoForwarded upstream. Also used when estimating the Credits reservation.
toolsNoForwarded when the offer supports tools; otherwise capability_unsupported.
response_formatNojson_object / json_schema require an offer with JSON support.

Non-stream responses set model to the Air model ID you requested (not the upstream provider model string).

{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "model": "air_gpt4o_pa_7K4D",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Hello from Air" },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 4,
    "total_tokens": 16
  }
}

Streaming

Set stream: true. Air proxies text/event-stream from the provider. Air requests stream_options.include_usage so token usage can be metered after the stream ends; Credits are finalized from that usage (reserve → capture).

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AIR_API_KEY,
  baseURL: "https://www.airinference.com/api/v1",
});

const stream = await client.chat.completions.create({
  model: "air_gpt4o_pa_7K4D",
  messages: [{ role: "user", content: "Say hello" }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

If the offer does not support streaming, the API returns 400 capability_unsupported.

GET

/models

Public

Public catalog of live provider-specific PAYG offers. No API key required. Each id is an Air model ID.

Request

curl https://www.airinference.com/api/v1/models
{
  "object": "list",
  "data": [
    {
      "id": "air_gpt4o_pa_7K4D",
      "object": "model",
      "display_name": "gpt-4o",
      "billing": "payg",
      "provider_model_id": "gpt-4o",
      "family": "gpt-4o",
      "listing_title": "Provider A",
      "listing_slug": "provider-a",
      "region": "us-east",
      "input_usd_per_1m": "0.18",
      "output_usd_per_1m": "0.72",
      "health_status": "healthy",
      "capabilities": {
        "streaming": true,
        "tools": true,
        "json": true,
        "vision": false
      }
    }
  ]
}

Fetch one offer: GET /models/{id} (same fields; 404 unknown_model if missing).

GET

/me/models

Auth required

Live offers plus enabled for this developer, and a Credits summary. Disabled rows appear so you can see what is available to enable.

curl https://www.airinference.com/api/v1/me/models \
  -H "Authorization: Bearer $AIR_API_KEY"
{
  "object": "list",
  "data": [
    {
      "id": "air_gpt4o_pa_7K4D",
      "object": "model",
      "enabled": true,
      "display_name": "gpt-4o",
      "billing": "payg",
      "provider_model_id": "gpt-4o",
      "listing_title": "Provider A",
      "input_usd_per_1m": "0.18",
      "output_usd_per_1m": "0.72",
      "capabilities": { "streaming": true, "tools": true, "json": true, "vision": false }
    }
  ],
  "credits": {
    "available": "25.00",
    "reserved": "0",
    "total": "25.00",
    "currency": "usd"
  }
}
EndpointReturns
/modelsPublic live offers
/me/modelsOffers + enabled flag for this key

Enable a model

Marketplace
→ choose provider / model offer
→ Enable PAYG
→ copy Air model ID

Enabling one provider's model does not enable another provider's model, even when both offer the same underlying model.

air_gpt4o_pa_7K4D
does not enable
air_gpt4o_pb_92QF

If you enable both, choose which Air model ID to send per request. Air does not automatically pick between them or fail over.

Billing

PAYG uses Air Credits. Each chat request reserves an estimate, calls the provider, then charges actual usage and releases the unused hold.

Air Credits
→ temporary reservation
→ provider request
→ actual usage
→ final charge
→ unused reservation released

Reservations are temporary holds, not additional charges. Final billing uses measured tokens.

Reserved: $0.001625
Actual:   $0.000029
Released: $0.001596
Final charge: $0.000029

Pricing

Cost is computed from the offer's rate card (USD per 1M tokens). There is no cent-rounding floor on the charge — amounts use high-precision USD.

cost =
  prompt_tokens × input_usd_per_1m / 1,000,000
+ completion_tokens × output_usd_per_1m / 1,000,000

Historical usage keeps the rate card that applied at call time. New calls use the current active rate.

Wallet

GET /me/wallet returns available / reserved Credits. Top up in the dashboard. Optional per-account spending limits are enforced when configured. Auto-recharge may run after successful calls when enabled for your account.

curl https://www.airinference.com/api/v1/me/wallet \
  -H "Authorization: Bearer $AIR_API_KEY"

Errors

Errors use an error object. Important codes for /chat/completions:

StatusCodeMeaningWhat to do
400invalid_json / missing_model / missing_messagesMalformed bodyFix JSON, model, or messages
400capability_unsupportedstream / tools / vision / JSON not supported by offerDisable that feature or pick another offer
400reservation_cap_exceededEstimated cost above max reservationLower max_tokens or split the request
401—Invalid or revoked API keyCreate a new air_ key
402insufficient_creditsNot enough Air CreditsTop up Credits
402spending_limit_exceededAccount/key spending limit hitRaise limit or wait for the window
403model_not_enabledKnown Air ID not enabledEnable on Marketplace
404unknown_modelUnknown Air model IDCopy ID from Marketplace /models
409idempotency_conflictSame Idempotency-Key, different bodyUse a new key for a new request
429—Over 60 req/min per keyBackoff; check Retry-After
502provider_failedUpstream provider failedRetry later or use another enabled Air ID
503model_unavailable / api_paused / credits_pausedOffer or platform unavailableRetry later; Air does not fail over
504provider_timeoutUpstream timed outRetry; unused reservation is released

Examples

{
  "error": {
    "message": "Model is not enabled for this account. Enable it on the marketplace listing page.",
    "code": "model_not_enabled",
    "model": "air_gpt4o_pa_7K4D",
    "marketplace": "https://www.airinference.com/dashboard/marketplace"
  }
}
{
  "error": {
    "message": "Insufficient Air Credits. Top up to continue.",
    "type": "insufficient_quota",
    "code": "insufficient_credits",
    "model": "air_gpt4o_pa_7K4D"
  },
  "air": {
    "available": "0.12",
    "reserved": "0",
    "topup": "https://www.airinference.com/dashboard/credits"
  }
}

Idempotency

Optional on POST /chat/completions:

Idempotency-Key: <unique-request-key>
  • Same key + same request body → safe replay of the stored response.
  • Same key + different body → 409 idempotency_conflict.

Rate limits

60 requests / minute / API key across authenticated /api/v1 routes. When exceeded, the API returns 429 with a Retry-After header (seconds).

OpenAI SDK

Air provides an OpenAI-compatible API. Use the standard OpenAI SDK with Air's base URL and an Air model ID. Air supports the fields and behavior documented on this page — not every OpenAI product feature.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AIR_API_KEY,
  baseURL: "https://www.airinference.com/api/v1",
});

const response = await client.chat.completions.create({
  model: "air_gpt4o_pa_7K4D",
  messages: [{ role: "user", content: "Hello from Air" }],
});

console.log(response.choices[0].message.content);

Providers

Publish a PAYG offer from the dashboard:

Provider
→ create listing / model offer
→ enter provider model ID + PAYG rates
→ Base URL + upstream API key
→ Validate
→ set Live

Validation checks the endpoint, credentials, model availability, a small inference request, and response shape. After publication, Air monitors offer health. If that exact provider is unavailable, the API returns an error — Air does not silently switch providers.

API version

These docs describe the /api/v1 API. Paths such as /chat/completions are relative to https://www.airinference.com/api/v1.