← Blog

2026-10-04 · 5 min read

Setting Up OpenAI-Compatible Endpoints with vLLM and llama.cpp

Learn how to configure vLLM and llama.cpp servers as OpenAI-compatible endpoints and successfully list them on the Air Inference marketplace.

Setting Up OpenAI-Compatible Endpoints with vLLM and llama.cpp

In today's rapidly evolving landscape of artificial intelligence, developers are constantly seeking ways to leverage powerful models while maintaining flexibility and cost-effectiveness. Air Inference offers a unique two-sided marketplace that enables developers to utilize OpenAI-compatible APIs while allowing providers to monetize their AI inference endpoints. In this guide, we will walk you through the process of setting up vLLM and llama.cpp servers as OpenAI-compatible endpoints and how to list them on the Air Inference marketplace.

Understanding vLLM and llama.cpp

Before we dive into the setup process, let's briefly discuss what vLLM and llama.cpp are.

vLLM is a versatile library designed for efficient inference of large language models. It allows you to run models with optimized memory usage and faster execution times, making it an excellent choice for developers looking to deploy AI applications.

llama.cpp, on the other hand, is an implementation of the LLaMA (Large Language Model Meta AI) architecture that provides a simple and efficient way to run LLaMA models on various hardware setups. It is particularly useful for developers who want to harness the power of LLaMA while maintaining compatibility with existing tools and frameworks.

By setting up these endpoints as OpenAI-compatible, you can tap into a broader audience on the Air Inference marketplace.

Prerequisites

Before proceeding, ensure that you have the following:

  • A server or local machine capable of running vLLM or llama.cpp.
  • Basic knowledge of Docker and command-line interfaces.
  • An account on Air Inference to list your endpoints.
  • OpenAI-compatible API knowledge (understanding of JSON request and response formats).

Step 1: Setting Up vLLM

1. Install Dependencies

To get started with vLLM, you'll need to install the necessary dependencies. You can do this using Python and pip. If you haven't installed vLLM yet, follow these commands:

pip install vllm

2. Running the vLLM Server

Once you have vLLM installed, you can start the server. Run the following command:

vllm serve --model <model_name> --port <port_number>

Replace <model_name> with the name of the model you wish to use (e.g., gpt-3.5-turbo) and <port_number> with the desired port for your server (e.g., 8080).

3. Exposing the Endpoint

To expose your vLLM server as an OpenAI-compatible endpoint, ensure that your server is accessible over the internet. You may need to configure a reverse proxy using tools like Nginx or Apache to route requests to your vLLM instance.

Here’s an example of a basic Nginx configuration:

server {
    listen 80;
    server_name your-domain.com;

    location /api/v1/ {
        proxy_pass http://localhost:8080/;  # Assuming vLLM runs on port 8080
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

4. Testing Your Endpoint

Once your server is running and exposed, you can test it using curl. Here’s a simple example of how to do this:

curl -X POST http://your-domain.com/api/v1/ \
-H "Content-Type: application/json" \
-d '{"prompt": "Hello, world!", "max_tokens": 50}'

If everything is configured correctly, you should receive a response from your vLLM server.

Step 2: Setting Up llama.cpp

1. Install llama.cpp

For those who prefer using llama.cpp, the installation process is straightforward. Ensure you have a compatible compiler and then clone the repository:

git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
mkdir build
cd build
cmake ..
make

2. Running the llama.cpp Server

You can run the llama.cpp server with the following command:

./llama_cpp_server --model <model_path> --port <port_number>

Replace <model_path> with the path to your LLaMA model files and <port_number> with your chosen port.

3. Exposing the Endpoint

Similar to vLLM, you’ll need to expose the llama.cpp server to the internet. You can use the same Nginx configuration as above, just change the proxy pass line to point to the port on which your llama.cpp server is running.

4. Testing Your Endpoint

You can test the llama.cpp endpoint in the same way as with vLLM:

curl -X POST http://your-domain.com/api/v1/ \
-H "Content-Type: application/json" \
-d '{"prompt": "What is the capital of France?", "max_tokens": 50}'

Step 3: Listing Your Endpoints on Air Inference

Once you have your endpoints running and exposed, it’s time to list them on the Air Inference marketplace. Here’s how to do it:

1. Create an Account

If you haven’t already, sign up for an account on Air Inference. Verify your email and log in.

2. Navigate to the Provider Dashboard

Once logged in, navigate to the Provider Dashboard. Here, you can list your OpenAI-compatible endpoints.

3. Fill Out the Endpoint Details

When listing your endpoint, you will need to provide the following details:

  • Endpoint URL: The URL where your API is accessible (e.g., http://your-domain.com/api/v1/).
  • Description: A brief description of your service and its capabilities.
  • Model Name: Specify the model(s) your endpoint supports (e.g., vLLM, LLaMA).
  • Pricing: Set your pricing model. Remember that Air Inference charges a ~10% fee off-platform.

4. Submit for Review

After filling out all necessary details, submit your endpoint for review. The Air Inference team will evaluate your submission, and once approved, your endpoint will be live on the marketplace.

Best Practices for Managing Your Endpoints

1. Monitor Performance

Regularly monitor the performance of your endpoints. Ensure that they are running efficiently and can handle the expected load. Utilize logging and monitoring tools to track usage metrics and identify potential issues.

2. Keep Models Updated

AI models are continually evolving. Make sure to keep your models updated to leverage the latest advancements and optimizations. This will enhance the performance and reliability of your endpoints.

3. Engage with Users

As a provider on Air Inference, it’s crucial to engage with your users. Collect feedback and be responsive to inquiries. This will help you improve your service and build a loyal customer base.

4. Optimize Costs

Keep an eye on your operational costs. Optimize your server usage and consider scalable solutions to handle varying loads. This will allow you to maintain profitability while providing high-quality service.

Conclusion

Setting up OpenAI-compatible endpoints using vLLM and llama.cpp is an excellent way to contribute to the growing AI ecosystem. By following the steps outlined in this guide, you can successfully expose your models as APIs and list them on the Air Inference marketplace. With proper management and engagement, you can build a sustainable AI service that meets the needs of developers and businesses alike.

As you embark on this journey, remember that the landscape is continually changing, and staying informed about new developments will be key to your success. Happy coding!