Smart
Inference
Console

THE FIELD GUIDE

GET SET.
MAKE IT SMART.

Choose a model, prepare a request,
and learn to read the receipt.

QUICKSTARTCONFIGURATION

01 / PICK YOUR BRAIN

Choose a model.

Compare capabilities, context and token rates in the catalogue. A developer creates the model; a provider serves it. Check provider conditions alongside prices.

Start with the same task.

Use the same input, cached input and output counts when comparing providers. A lower output rate alone does not mean a lower total request cost.

Explore model rates →

02 / PREPARE YOUR CLIENT

A template for your next request.

Choose a model and client below. Copy the relevant configuration, then replace its placeholders when live access is available.

CONFIGURATION

Replace the base URL and key placeholders when verified live access is available. These templates do not call an API.

03 / READ THE RECEIPT

Three counts. One estimate.

Input

New prompt tokens, excluding cache hits.

Cached input

Reused input billed at the provider’s cache rate, where available.

Output

Tokens generated in the response.

(input × input rate + cache × cache rate + output × output rate) / 1,000,000

Token estimates exclude separately applicable payment and service fees. If cached tokens are used and a cache rate is missing, the total is unavailable.

Explore a request estimate →

04 / CURRENT ACCESS

What you can do today.

  • Available in thisBrowse models, compare historical rates, estimate tasks and copy configuration templates.
  • Not connected yetAuthentication, API key creation, live inference, usage balance and payments.

Client templates are an integration starting point. They are not proof of a working endpoint or tested client compatibility.

05 / INSPECT THE DETAILS

Take a look inside.

The existing Console contains usage and request records. Use it to inspect the relationship between a request, its token usage and its cost.

Open Console ↗