Caffeine InferenceEarly access

Cost-effective AI & Smart Routing.

What we're building: Caffeine Inference aims to deliver cost-effective smart routing. You don't need to choose models, our routing does it for you. What exists at the moment: cost-efficient inference.

Open Caffeine Inference

Runs on the Caffeine subscription.

Routing, not a menu.

Once smart routing lands, Caffeine will read the request, weigh what each model would cost and how fast they are answering right now, and send it to one of them. When a model is busy or slow, the request will go somewhere else instead of failing.

your requestroutercost · load · difficultymodelmodelmodelanswer

Point the tools youalready use at Caffeine.

Claude Code and anything else that accepts a custom base URL. Copy the config, paste it, keep working.

Goose

GOOSE_PROVIDER=openai
OPENAI_HOST=...
OPENAI_API_KEY=...

Claude Code

ANTHROPIC_BASE_URL=...
ANTHROPIC_AUTH_TOKEN=...

Anything else

base URL + key
same request shape

Four steps to your first call.

Pick a tool, copy the config, and send one call. Caffeine gives you the settings in the format that tool expects.

1.Overview

Questions we'd rather answer here.

Input is $0.2 per million tokens, cached input $0.05, and output $0.5. Pricing is subject to change as routing improves.

Caffeine Inference

It runs on the Caffeine budget already in your account, and your existing config is three lines away from working.

Get started

Early access. No separate billing, no separate account.