← Blog

2026-10-01 · 4 min read

Creating OpenAI-Compatible Endpoints with vLLM and llama.cpp for Marketplace Listing

Learn how to expose vLLM and llama.cpp servers as OpenAI-compatible endpoints and list them on the Air Inference marketplace, maximizing your potential earnings from AI inference.

Creating OpenAI-Compatible Endpoints with vLLM and llama.cpp for Marketplace Listing

In the rapidly evolving landscape of artificial intelligence, developers and providers alike are seeking innovative ways to monetize their AI models. Air Inference serves as a robust marketplace that connects developers in need of AI inference with providers hosting OpenAI-compatible endpoints. This guide will take you through the process of exposing vLLM and llama.cpp servers as OpenAI-compatible endpoints, allowing you to list them on the Air Inference marketplace and maximize your earnings.

Understanding vLLM and llama.cpp

Before diving into the technical details, it’s essential to understand what vLLM and llama.cpp are.

  • vLLM: This is a high-performance implementation designed for large language models. It offers efficient memory management and can handle multiple requests concurrently, making it ideal for serving AI applications.

  • llama.cpp: This is a lightweight wrapper around the LLaMA (Large Language Model Meta AI) model, allowing developers to run inference on systems with limited resources. It’s particularly useful for those who want to deploy LLaMA models without heavy infrastructure.

Both vLLM and llama.cpp are compatible with the OpenAI API specifications, making them excellent candidates for listing on Air Inference.

Setting Up Your Environment

To expose your vLLM or llama.cpp server as an OpenAI-compatible endpoint, follow these steps:

  1. Install Required Software: Ensure you have Python installed, along with the necessary libraries for vLLM or llama.cpp. You can typically install these with pip:

    pip install vllm
    

    or for llama.cpp, you might need to clone the repository and build it according to its instructions.

  2. Set Up Your Model: Load your desired model into vLLM or llama.cpp. For vLLM, you'd typically do something like this:

    from vllm import Model
    
    model = Model("path/to/your/model")
    

    For llama.cpp, refer to its documentation for loading models.

Exposing the Endpoint

Once your model is loaded, you can expose it via an HTTP API. This can typically be done using frameworks like Flask or FastAPI in Python.

Using FastAPI with vLLM

Here’s a basic example of how to set up a FastAPI server with vLLM:

from fastapi import FastAPI
from vllm import Model

app = FastAPI()
model = Model("path/to/your/model")

@app.post("/v1/engines/{engine_id}/completions")
async def create_completion(engine_id: str, prompt: str):
    response = await model.generate(prompt)
    return {"id": engine_id, "choices": [{"text": response}]}

This endpoint listens for POST requests at /v1/engines/{engine_id}/completions and returns generated text based on the prompt received.

Using llama.cpp with Flask

For llama.cpp, you can use a similar approach with Flask:

from flask import Flask, request, jsonify
from llama import Llama

app = Flask(__name__)
model = Llama("path/to/your/model")

@app.route('/v1/engines/<engine_id>/completions', methods=['POST'])
def create_completion(engine_id):
    data = request.json
    prompt = data.get("prompt")
    response = model.generate(prompt)
    return jsonify({"id": engine_id, "choices": [{"text": response}]})

Testing Your Endpoint

Once your server is running, you can test the endpoint using curl:

curl -X POST "http://localhost:8000/v1/engines/my-engine/completions" \
     -H "Content-Type: application/json" \
     -d '{"prompt": "What is AI?"}'

Make sure to replace my-engine with your actual engine ID.

Listing on Air Inference Marketplace

Now that you have your OpenAI-compatible endpoint up and running, the next step is to list it on the Air Inference marketplace. Here’s how you can do that:

  1. Create an Account: If you haven’t already, create an account on Air Inference. This will be your platform for managing your listings.

  2. Top Up Air Credits: Before you can list your endpoint, ensure you have topped up your Air Credits. This is a requirement for providers on the marketplace.

  3. Submit Your Endpoint: Navigate to the provider section of the Air Inference website and submit your endpoint details. This will typically include:

    • The URL of your endpoint
    • Description of the model and its capabilities
    • Pricing information
    • Any specific usage guidelines
  4. Compliance and Approval: Your submission will be reviewed for compliance with Air Inference standards. Be prepared to provide additional information if requested.

Optimizing Your Listing

To maximize your potential earnings from AI inference, consider the following tips when creating your marketplace listing:

  • Clear Descriptions: Provide a clear and concise description of your model, including its strengths and potential use cases. This helps attract the right developers.

  • Competitive Pricing: Analyze similar offerings in the marketplace and set a competitive price. Remember that Air Inference takes a small fee (around 10%) from your earnings.

  • Usage Examples: Include usage examples in your listing. This helps potential users understand how to interact with your endpoint effectively.

  • Performance Metrics: If applicable, provide performance metrics such as response times and accuracy rates. This builds trust with potential users.

Maintaining Your Endpoint

Once your endpoint is live, it’s essential to maintain it effectively:

  • Monitor Performance: Regularly check the performance of your model and endpoint. Use monitoring tools to track response times and error rates.

  • Update Models: As new versions of models become available or as you improve your own models, ensure that you update the endpoint accordingly.

  • Engage with Users: Respond to feedback from developers using your endpoint. This can help you enhance the model and improve user satisfaction.

Conclusion

Creating OpenAI-compatible endpoints with vLLM and llama.cpp is a valuable opportunity for providers to monetize their AI models through the Air Inference marketplace. By following the steps outlined in this guide, you can expose your models as endpoints and effectively list them on the marketplace, maximizing your earning potential.

As the demand for AI inference continues to grow, platforms like Air Inference offer a unique opportunity for developers and providers to connect and thrive in this exciting field. Be sure to keep your endpoints updated, engage with your users, and continuously look for ways to enhance your offerings.