Skip to content

Models

Compare Models

Send one prompt to several models at once, read the answers next to each other, and synthesize them into a single best answer.

  • An account with the Models permission — see Permissions.
  • At least one API key, and credits on the balance: a comparison runs against a real key and bills like a normal API call, once per model.

On the Models Playground toolbar, click Compare models. The page is titled Multi-model comparison; Back returns to the catalogue.

The category tabs Text, Image, and Video decide which models and parameters the page offers, so pick one before you build a lineup.

  1. Under the lineup, click Add a model… and use Search models… to pick each entrant. The picker stops you at the ceiling with You can select a maximum of 6 models.

  2. Choose a key in the API key selector. It reads Select an API key, or No API keys available when you hold none; keys are listed as Key #12.

  3. Type a prompt in the composer and click Send. Shift+Enter starts a new line, and Stop all halts every card at once.

  4. The lineup locks once a conversation starts: The model lineup is locked during a conversation. Start a new comparison to change models. Use Clear to wipe the record and unlock it.

Screenshot of a multi-model comparison with three model cards and the Comparison Overview awards

The settings panel applies to every model: Temperature, Top P, Max Tokens, and Render as (Auto, Text, Markdown). Image and video tabs add Number of Images, Duration (s), Aspect ratio, Resolution, and Negative Prompt, each marked (not supported by all selected models). Reset to global defaults puts them back.

Reference images (or Start frame on video) take Add image, Replace image, and Remove: Up to 10 images. The same references are sent to every selected model.

Each card carries Expand/Collapse, Show rendered output / Show raw response, Copy, Copy entire conversation, Regenerate, Continue with this model, Stop generating, and Remove. Override params for this model takes one model off the global settings — leave a field empty and it falls back to the global value. The header shows Pricing at launch, the rate, latency, and the state Streaming…, Generating…, or Failed.

The Comparison Overview card gives two awards: Fastest (lowest latency) and Lowest cost (cheapest estimated spend, shown as ~cost). Each axis needs at least two valid candidates, otherwise no winner is shown, and models whose pricing is unknown are excluded rather than winning by default.

More actions → Bookmarks saves the current models and settings to reload later. Save current as bookmark… opens Save bookmarkSaves the current category, selected models, and parameters under a name you can recall later. — with the placeholder e.g. My weekly benchmark. Saved lineups list under the same menu with Delete bookmark; an empty list reads No saved bookmarks yet.

Below the cards, the fuse panel says: Pick a model to synthesize the 3 selected answers below into a single best answer. Billed as one normal call to that model.

Tick the answers to include, pick a judge under Fuse with, then click Fuse. Fuse settings carries its own Temperature and Top P plus Fusing instructionsSent as the system prompt of the fuse call. Leave empty to use the default. The result lands in Best fused answer with Copy; Fuse again re-runs it and Stop halts it.

Screenshot of the fuse panel with the selected answers, the Fuse with picker and a best fused answer

Running N models is N charges, and a fuse adds one more. Every charge appears on Inference Usage against the key you selected.