One endpoint. Every model. Automatic fallback.
The Inferix Router is an OpenAI-compatible endpoint that routes each request to the best available model on real GPUs — with a fallback chain so a single unavailable model never fails your call. Something you won’t find on legacy model hubs.
Automatic fallback
Pass a chain of models. If the first is unavailable, the router transparently tries the next — your request still succeeds.
Smart routing aliases
Use auto or cheapest to always hit the smallest capable model, or largest / best for the most capable one. No hard-coding.
One OpenAI-compatible endpoint
A single /v1/chat/completions surface for every routable model. Drop it into any OpenAI SDK by changing the base URL.
Served on real Inferix GPUs
Every completion runs on verified GPU hardware via llama.cpp — no third-party reseller in the path.
Routable models
Live from the router — use any id below, or an alias like auto / cheapest / largest.
Try the router
Send a real request. The response shows exactly which model served it and the full fallback trace.
API reference
OpenAI-compatible — point any OpenAI SDK at the base URL and set model to a routable id or alias.
curl https://inferix.co/api/v0/router/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $INFERIX_TOKEN" \
-d '{
"model": "auto",
"models": ["qwen2.5-1.5b", "llama-3.2-1b"],
"messages": [{"role": "user", "content": "Explain routing in one line."}]
}'{
"object": "chat.completion",
"model": "Qwen/Qwen2.5-0.5B-Instruct-GGUF",
"choices": [{ "index": 0, "message": { "role": "assistant", "content": "..." } }],
"x_inferix_router": {
"served_by": "qwen2.5-0.5b",
"fallback_used": false,
"attempts": [{ "model": "qwen2.5-0.5b", "ok": true }]
}
}Endpoints: GET /api/v0/router/models · POST /api/v0/router/v1/chat/completions. Aliases: auto, cheapest, largest, best.