You pay the same per-token price the model provider charges.
Keep your SDK. Change the destination.
The request shape stays familiar. Replace the API key and base URL, then choose a model from the live catalogue.
Before · OpenAI directly
client = OpenAI(
api_key=OPENAI_API_KEY
)After · LLM.API
client = OpenAI(
api_key=LLM_API_KEY,
base_url="https://api.llmapi.ai/v1"
)If one provider route fails, requests can continue on another route for the same model.
Set a budget on each API key so one app or agent cannot overspend.
Your prompts and responses are not stored after the request completes.
Use it in your IDE or agent
Any coding tool with an "OpenAI-compatible" provider option works. Paste these values into its settings.
OpenAI Compatiblehttps://api.llmapi.ai/v1YOUR_LLM_API_KEYclaude-sonnet-5Coding agents burn through tokens fast. A spend cap on the key you give the agent keeps a long session from turning into a surprise bill.
The same values work in frameworks such as LangChain and LlamaIndex through their OpenAI client classes — set the base URL and keep the rest of your code.
Your existing OpenAI client, pointed at LLM.API
Use the official OpenAI package you already know. Your API key belongs on the server, never in browser code.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_LLM_API_KEY",
base_url="https://api.llmapi.ai/v1",
)
response = client.chat.completions.create(
model="claude-sonnet-5",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LLM_API_KEY,
baseURL: "https://api.llmapi.ai/v1",
});
const response = await client.chat.completions.create({
model: "claude-sonnet-5",
messages: [{ role: "user", content: "Hello" }],
});
console.log(response.choices[0].message.content);curl https://api.llmapi.ai/v1/chat/completions \
-H "Authorization: Bearer $LLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5",
"messages": [{"role": "user", "content": "Hello"}]
}'Using the Responses API? That works too
LLM.API also supports OpenAI's newer Responses API, which many GPT-5 and GPT-6 integrations use. Keep the same client and call client.responses.create.
response = client.responses.create(
model="gpt-6-sol",
input="Summarise this ticket in one line.",
)
print(response.output_text)Switch model families with one value
These are current model IDs from the LLM.API catalogue. The rest of the request can stay the same.
Familiar endpoints, model-specific capabilities
Compatibility describes the request format. Features still depend on the model you select, so check each model page before shipping.
Chat completions
SupportedThe standard messages-based request format.
Streaming
SupportedStream incremental output with the SDK's normal streaming option.
Tool calling
Model-dependentChoose a model marked Tools in the catalogue.
Structured output
Model-dependentJSON and schema support vary by model.
Vision input
Model-dependentChoose a model marked Vision for image inputs.
Embeddings
SupportedUse embedding models from the dedicated embeddings catalogue.
A low-risk migration path
Start with one request, confirm the response your application relies on, then widen traffic.
Create a key
Create an LLM.API account and store the key in your server environment.
Set the base URL
Point the OpenAI client at https://api.llmapi.ai/v1.
Choose a model ID
Copy the exact identifier from the live model catalogue.
Test your response path
Check streaming, tools or structured output with the model you selected.
Three checks solve most migration errors
Do not change your whole application first. Confirm the key, endpoint and model identifier in that order.
401 UnauthorizedConfirm that the LLM.API key is present, current and sent as a bearer token.
404 or model not foundCopy the exact model ID from the Models page; display names are not request IDs.
Unexpected response behaviorConfirm that the chosen model supports the feature your code expects, such as tools or vision.
Choose, compare and control spend
Use the live catalogue and independent research pages before deciding which model should receive production traffic.
Browse all models
Compare model capabilities, context and live prices.
Open Models →Buyer's guidesChoose the right model
Use interactive quizzes, research and cost calculators.
Open the guides →Cost controlEstimate your monthly bill
Compare models at your traffic and token volume.
Open AI Cost Control →OpenAI-compatible API FAQ
What compatibility means in practice when you move an existing integration.
What does OpenAI-compatible mean?
It means you can use the familiar OpenAI SDK and request format with LLM.API by setting a different base URL and API key.
Do I need to rewrite my application?
Usually not. Start by changing the API key, base URL and model ID. Test any model-specific features your application uses.
Can I use models from different companies?
Yes. LLM.API exposes supported model families through one API. Change the model ID to switch between them.
Does every model support tools, JSON and vision?
No. Those capabilities depend on the selected model. Check the model catalogue and the individual model page before relying on one.
Can I keep using the official OpenAI SDK?
Yes. The examples above use the official OpenAI Python and JavaScript packages with a custom base URL.
Where should I store my API key?
Store it in a server-side environment variable or secret manager. Do not expose it in public browser code.
Keep the SDK. Expand your model options.
Create an API key, run the example above, and test your first OpenAI-compatible request through LLM.API.