Unified endpoint
One OpenAI-compatible URL for every configured model. Point your existing SDK at it and keep shipping.
One endpoint · configured frontier models
Wholerouter is a single OpenAI-compatible endpoint for the configured Anthropic, DeepSeek, and GLM models your organization can use.
multiple frontier models→1 endpoint
Model roster
Provider-backed models are reachable through the same endpoint, with a clear published roster.
Rates
USD per 1M tokens. Rates below are maintained with the published model configuration.
| Model | Input / 1M | Cached input / 1M | Output / 1M | Discount |
|---|---|---|---|---|
| $0.750 | $0.150 | $2.50 | 25% off | |
| $0.920 | $0.180 | $3.00 | 23% off | |
| $1.00 | $0.800 | $3.30 | 29% off | |
| $1.05 | $0.195 | $3.30 | 25% off | |
| $4.25 | $0.425 | $21.25 | 15% off | |
| $1.70 | $0.170 | $8.50 | 15% off | |
| $8.50 | $0.850 | $42.50 | 15% off | |
| $4.25 | $0.425 | $21.25 | 15% off | |
| $4.25 | $0.425 | $21.25 | 15% off | |
| $4.25 | $0.425 | $21.25 | 15% off | |
| $2.55 | $0.255 | $12.75 | 15% off | |
| $0.850 | $0.085 | $4.25 | 15% off | |
| $0.105 | $0.0021 | $0.210 | 25% off | |
| $0.326 | $0.0027 | $0.652 | 25% off |
How it routes
One OpenAI-compatible URL for every configured model. Point your existing SDK at it and keep shipping.
WholeRouter keeps routing and spend clear: automatic retries and hidden provider fallbacks are disabled.
Every request is logged against your organization with published configured token prices.
Getting started
Start with the partner login flow and access the organization your administrator has provisioned.
Use the dashboard to create the LiteLLM virtual key that authorizes your configured model access.
Keep your OpenAI-compatible request shape and send traffic through the WholeRouter endpoint.
FAQ
The roster is generated from WholeRouter's published configuration and stays in sync with the proxy.
Yes. Use the OpenAI-compatible chat-completions and Responses request shapes with your WholeRouter virtual key.
Cached input pricing applies only when the upstream provider reports cache-read token usage for a request.
No. Automatic retries and provider fallbacks are disabled so provider behavior and spend remain explicit.
Partner access starts with the WholeRouter sign-in flow. Your organization then uses LiteLLM virtual keys for inference.
Ready when you are