Models
Inference Usage
Inference API Usage is where every Playground, comparison, and API call is totalled — by key, by model, and over time.
Open the usage page
Section titled “Open the usage page”The page needs the Models permission — see Permissions. Three entry points reach it:
| From | Control |
|---|---|
| Playground | Your Usage |
| Billing & Usage | View inference usage |
| API key row | View Usage (opens filtered, ?api_key_id=) |
Under the title sits the staleness note: All dates in UTC · data may lag up to 30 min.

Filter the view
Section titled “Filter the view”| Filter | What it does |
|---|---|
| API Key | One key, or All keys. |
| Time range | Presets, or Custom with Pick dates. A range must stay inside 180 days. |
| Metric | Switches every figure between Cost and tokens or units. |
| Sort by | Cost or Name, with Ascending / Descending. |
| Hide unused (37) | Drops models with no usage; Show 37 unused models brings them back. |
The Text, Image, and Video tabs change the charge unit the page counts in — Tokens for text, Units for image and video.
Read the numbers
Section titled “Read the numbers”Total spend · all types in the header covers every category at once, regardless of the tab you are on. Below it are three tiles:
| Tile | Reads |
|---|---|
| Total cost | Spend in the selected range and category. |
| Total tokens | Tokens or units for the same range. |
| Active models | How many models were used, of 40. |
The Token breakdown card splits Input, Cached, and Output, and gives a Cache hit percentage — 12% of input. Cached tokens are a subset of input tokens, never an addition to them.
Spend over time
Section titled “Spend over time”The chart is titled Spend over time, or Tokens over time when the Metric filter is set to tokens or units. Granularity toggles between daily and weekly; weeks start on Monday. Updating… appears while a new range is being fetched.
Per-model table
Section titled “Per-model table”One row per model, with Spend, Cache hit, and Token breakdown, ordered by the Sort by filter. With nothing to show it reads No models with usage in this range.
Limits
Section titled “Limits”- There is no export or download on this page; read the figures in the console.
- A failed load shows Couldn’t load usage and Please try again. with Retry — the range is not lost, so retrying re-runs the same query.
- To reconcile with an invoice, compare against Usage This Month on Credits & Billing, which splits Compute, API-Token, and Subscription spend. The balance in the header is authoritative; this page is a rollup and can trail it.

