One endpoint · configured frontier models

The right model,
at a better rate.

Wholerouter is a single OpenAI-compatible endpoint for the configured Anthropic, DeepSeek, and GLM models your organization can use.

Claude Opus 5Claude Sonnet 5DeepSeek V4 ProDeepSeek V4 FlashGLM 5.2Claude Haiku 4.5Your App

multiple frontier models1 endpoint

Model roster

Every configured model, one key.

Provider-backed models are reachable through the same endpoint, with a clear published roster.

Z.ai logoGLM 5Z.ai logoGLM 5 TurboZ.ai logoGLM 5.1Z.ai logoGLM 5.2Anthropic logoClaude Opus 5Anthropic logoClaude Sonnet 5Anthropic logoClaude Fable 5Anthropic logoClaude Opus 4.8Anthropic logoClaude Opus 4.7Anthropic logoClaude Opus 4.6Anthropic logoClaude Sonnet 4.6Anthropic logoClaude Haiku 4.5DeepSeek logoDeepSeek V4 FlashDeepSeek logoDeepSeek V4 Pro

Rates

Transparent rates, built in.

USD per 1M tokens. Rates below are maintained with the published model configuration.

ModelInput / 1MCached input / 1MOutput / 1MDiscount
Z.ai logoGLM 5Z.ai$0.750$1.00$0.150$0.200$2.50$3.2025% off
Z.ai logoGLM 5 TurboZ.ai$0.920$1.20$0.180$0.240$3.00$4.0023% off
Z.ai logoGLM 5.1Z.ai$1.00$1.40$0.800$0.260$3.30$4.4029% off
Z.ai logoGLM 5.2Z.ai$1.05$1.40$0.195$0.260$3.30$4.4025% off
Anthropic logoClaude Opus 5Anthropic$4.25$5.00$0.425$0.500$21.25$25.0015% off
Anthropic logoClaude Sonnet 5Anthropic$1.70$2.00$0.170$0.200$8.50$10.0015% off
Anthropic logoClaude Fable 5Anthropic$8.50$10.00$0.850$1.00$42.50$50.0015% off
Anthropic logoClaude Opus 4.8Anthropic$4.25$5.00$0.425$0.500$21.25$25.0015% off
Anthropic logoClaude Opus 4.7Anthropic$4.25$5.00$0.425$0.500$21.25$25.0015% off
Anthropic logoClaude Opus 4.6Anthropic$4.25$5.00$0.425$0.500$21.25$25.0015% off
Anthropic logoClaude Sonnet 4.6Anthropic$2.55$3.00$0.255$0.300$12.75$15.0015% off
Anthropic logoClaude Haiku 4.5Anthropic$0.850$1.00$0.085$0.100$4.25$5.0015% off
DeepSeek logoDeepSeek V4 FlashDeepSeek$0.105$0.140$0.0021$0.0028$0.210$0.28025% off
DeepSeek logoDeepSeek V4 ProDeepSeek$0.326$0.435$0.0027$0.0036$0.652$0.87025% off

How it routes

Route once. Reach everything.

Unified endpoint

One OpenAI-compatible URL for every configured model. Point your existing SDK at it and keep shipping.

Explicit provider behavior

WholeRouter keeps routing and spend clear: automatic retries and hidden provider fallbacks are disabled.

One bill, transparent rates

Every request is logged against your organization with published configured token prices.

Getting started

Three steps to your first request.

01

Sign in to WholeRouter

Start with the partner login flow and access the organization your administrator has provisioned.

02

Create a virtual key

Use the dashboard to create the LiteLLM virtual key that authorizes your configured model access.

03

Point your client at one URL

Keep your OpenAI-compatible request shape and send traffic through the WholeRouter endpoint.

FAQ

Clear before you route.

Which models can I use?

The roster is generated from WholeRouter's published configuration and stays in sync with the proxy.

Is the endpoint OpenAI-compatible?

Yes. Use the OpenAI-compatible chat-completions and Responses request shapes with your WholeRouter virtual key.

What is cached input pricing?

Cached input pricing applies only when the upstream provider reports cache-read token usage for a request.

Does Wholerouter automatically fail over requests?

No. Automatic retries and provider fallbacks are disabled so provider behavior and spend remain explicit.

How do I get access?

Partner access starts with the WholeRouter sign-in flow. Your organization then uses LiteLLM virtual keys for inference.

Ready when you are

Stop juggling API keys. Start routing smarter.