Infero — from the Latin inferre, to bring in, to conclude — routes your requests across dozens of frontier models through a single OpenAI-compatible surface. One key, one bill, zero rewrites. If a route goes down, the next one picks up mid-sentence.
Frontier proprietary models, open-weight reasoning, fast coders and cheap workhorses — curated, versioned, and reachable through the same /v1/chat/completions you already call.
GPT-5.5, 5.6 Sol/Luna/Terra, GPT-6 Astra — flagship reasoning and vision, served straight from upstream.
Sonnet and Opus class models for long-context agentic work, code review and careful writing.
Gemini Flash tier for high-volume, low-latency pipelines — billed per token like everything else.
Pro and Flash, tools-enabled, million-token context at a fraction of closed-model pricing.
GLM-5 series, Kimi K2.7-Code, Qwen 3.x — the strongest open coders and reasoners, always current.
Specialist and experimental models cycled in as they earn their place on the router.
Swap api.openai.com for api.infero.sbs/v1. OpenAI, OpenRouter and every compatible client just works.
Use one key across the whole catalogue. Combos route to the best live provider automatically.
Dead routes are pulled and retried upstream-side. Your app sees one steady endpoint, not a zoo of providers.
$ curl https://api.infero.sbs/v1/chat/completions \ -H "Authorization: Bearer inf_…" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-5.6-sol", "messages": [ { "role": "user", "content": "Route me somewhere fast." } ], "stream": true }' > data: {"id":"chatcmpl-…","provider":"infero"} > data: "choices":[{"delta":{"content":"Routed."}}]} > data: [DONE]
Transparent per-million-token pricing, prepaid credit, no top-up fee, no per-request surcharge. Route intelligently: cheap models for volume, flagships for the hard turns.
| Model | Context | Input / MTok | Output / MTok |
|---|---|---|---|
| GPT-5.6 Sol gpt-5.6-sol | 1M | $5.00 | $30.00 |
| GPT-5.6 Luna gpt-5.6-luna | 400k | $1.00 | $6.00 |
| Claude Sonnet class claude-sonnet | 200k | $3.00 | $15.00 |
| DeepSeek V4 Pro deepseek-v4-pro | 1M | $1.74 | $3.48 |
| DeepSeek V4 Flash deepseek-v4-flash | 1M | $0.19 | $0.51 |
| Kimi K2.7 Code kimi-k2.7-code | 262k | $0.95 | $4.00 |
| GLM-5 Flash glm-5-flash | 128k | $0.10 | $0.40 |
| Gemini Flash class gemini-flash | 1M | $0.30 | $2.50 |
Point any OpenAI-compatible client at https://api.infero.sbs/v1 with your bearer token and go. Works with OpenAI SDKs, LangChain, Cline, OpenCode, anything that speaks the protocol.
from openai import OpenAI
client = OpenAI(
base_url="https://api.infero.sbs/v1",
api_key="inf_…",
)
r = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Hello, router."}],
)
print(r.choices[0].message.content)