← Blog

2026-09-19 · 5 min read

How Developers Access AI Inference Models with Air Inference

Explore how developers can enable AI models, purchase Air Credits, and effectively call the marketplace's /api/v1 chat completions endpoint for seamless integration.

How Developers Access AI Inference Models with Air Inference

As artificial intelligence continues to reshape industries and enhance productivity, the demand for accessible AI inference models has never been higher. Air Inference is a two-sided marketplace that connects developers with AI providers, allowing for seamless integration of AI capabilities into applications. In this article, we will explore how developers can enable AI models, purchase Air Credits, and effectively call the marketplace's /api/v1 chat completions endpoint.

Understanding Air Inference

Air Inference serves as a bridge between developers seeking AI functionalities and providers offering OpenAI-compatible endpoints. Developers can easily top up their Air Credits, which serve as the currency for accessing various AI models hosted by providers. These providers can include wrappers for models like vLLM, llama.cpp, RunPod, and more.

The platform operates on a fee model, where providers receive payments off-platform, subject to a documented fee of around 10%. This structure incentivizes providers to list their endpoints while ensuring that developers have access to a wide range of powerful AI models.

Enabling AI Models

Selecting the Right Model

The first step for any developer looking to leverage AI inference is to select the appropriate model that fits their needs. Air Inference offers a variety of models, each with unique capabilities. Here are some factors to consider when choosing a model:

  • Use Case: Understand the specific tasks you want the AI to perform, such as text generation, summarization, or conversational AI.
  • Performance: Different models have varying performance metrics. Look for benchmarks or user reviews to gauge the effectiveness of a model based on your requirements.
  • Compatibility: Ensure that the model you choose is compatible with the OpenAI API standards, as this will facilitate smoother integration into your application.

Listing Your Model

If you are a provider looking to list your AI model on Air Inference, the process is straightforward. You will need to:

  1. Create an Account: Sign up on the Air Inference platform.
  2. Submit Your Endpoint: Provide the necessary details about your AI model's endpoint, including its capabilities and pricing.
  3. Compliance Check: Ensure that your model complies with the platform's guidelines, including the necessary performance and safety requirements.

Once your model is listed, developers can discover and access it easily through the marketplace.

Purchasing Air Credits

To access AI models on Air Inference, developers must first purchase Air Credits. This process is simple and can be broken down into a few straightforward steps:

Step 1: Create an Account

Before you can purchase Air Credits, you need to create an account on Air Inference. This account will allow you to manage your credits and access the APIs.

Step 2: Top Up Your Air Credits

Once you have created your account, you can top up your Air Credits through a secure payment process. The exact payment methods available may vary, but common options include credit cards and other online payment solutions.

Step 3: Manage Your Credits

After purchasing Air Credits, you can manage your balance directly from your account dashboard. You'll be able to see your total credits, transaction history, and any pending purchases. Keeping track of your credits is crucial, as it ensures you can continue to access the AI models you need without interruptions.

Calling the /api/v1 Chat Completions Endpoint

With your Air Credits topped up and your selected AI model in mind, the next step is to call the /api/v1 chat completions endpoint. This is where the integration happens, and you can start utilizing the AI capabilities directly in your application.

Setting Up Your Request

To call the chat completions endpoint, you will typically need to construct an API request that includes the following elements:

  • Endpoint URL: The URL of the API you want to call, which will be in the format of https://api.airinference.com/api/v1/chat/completions.
  • Headers: Include necessary headers, such as Authorization for your API key and Content-Type to specify the type of data you're sending.
  • Payload: This will consist of the input data you want to send to the model, such as prompts or conversation history.

Example API Call

Here is a simplified example of how you might structure a request to the /api/v1 chat completions endpoint using curl:

curl -X POST https://api.airinference.com/api/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
  "model": "your_model_name",
  "messages": [
    {
      "role": "user",
      "content": "Hello, how can I access AI models?"
    }
  ]
}'

In this example, replace YOUR_API_KEY with your actual API key and your_model_name with the name of the model you are using. The messages array contains the conversation context, allowing the model to generate an appropriate response based on the user input.

Handling the Response

Once your request is sent, the API will respond with the model's output. This response typically includes the generated text and any relevant metadata. It's crucial to handle this response effectively in your application to provide users with a seamless experience.

  • Parse the Response: Extract the generated text and any other important information from the response.
  • Display the Output: Present the output to your users in a user-friendly format, whether that's in a chat interface, a webpage, or another medium.

Best Practices for Using Air Inference

To make the most out of Air Inference and its AI models, consider the following best practices:

Optimize Your API Calls

  • Batch Requests: If your application requires multiple requests, consider batching them to reduce latency and enhance performance.
  • Rate Limiting: Be mindful of any rate limits imposed by the API to avoid disruptions in service.

Monitor Usage

Keep track of your Air Credits usage and model performance. This will help you make data-driven decisions about which models to use and when to top up your credits.

Stay Updated

AI technology evolves rapidly, and so does the Air Inference marketplace. Stay informed about new model releases, updates, and best practices by following the official Air Inference blog and community forums.

Conclusion

Air Inference is revolutionizing how developers access AI inference models by providing a user-friendly platform to purchase Air Credits and call powerful APIs. By understanding how to select models, manage credits, and effectively interact with the /api/v1 chat completions endpoint, developers can seamlessly integrate AI capabilities into their applications.

Whether you're a seasoned developer or just starting in the AI space, Air Inference offers the tools you need to enhance your projects and deliver innovative solutions. Embrace the future of AI inference today by exploring what Air Inference has to offer!