Model cost & benchmark comparison

Token prices, context limits, and benchmark scores for 20 frontier models, sourced from vendor docs and the OpenRouter models API. Click a column header to sort; “—” means no clean score was found.

USD / 1M tokensgood bad
ModelInputCachedOutputContextMax outSWE-ProSWE-VerifiedTerminalOSWorldBrowseCompAA IntelAA CodeAA AgentAgentic benchBenchmark noteStatus

Quick cost calculator

Base token estimate only; excludes long-context surcharges, effort/tool fees, and provider-specific billing.

input tokens

output tokens

Benchmarks used here

  • SWE-bench Pro — harder, real-world multi-file coding tasks; less saturated than the original SWE-bench. scale.com/leaderboard/swe-bench-pro
  • SWE-bench Verified — human-verified subset of real GitHub issues; can the model produce a patch that passes the repo's own tests? swebench.com
  • Terminal-Bench — agentic tasks executed entirely inside a real terminal/CLI environment. tbench.ai
  • OSWorld — computer-use agents completing real desktop/browser GUI tasks end-to-end. os-world.github.io
  • BrowseComp — OpenAI benchmark for agents that have to browse the web to find hard-to-locate facts. openai.com/index/browsecomp
  • AA Intel / Code / Agent — Artificial Analysis' composite 0-100 indices blending multiple benchmarks per capability, served live via the OpenRouter models API. artificialanalysis.ai
  • Design Arena — head-to-head human voting on model-generated UI/creative output, referenced in a few benchmark notes. designarena.ai

Also see OpenRouter's usage rankings — a live leaderboard of which models are actually being run in production across OpenRouter, not a quality benchmark.

Legend / caveats

* = source snippet / vendor chart summaryref. = comparison reference model— = no clean public score found
  • Agent harness, effort level, tool budget, and timeout can change benchmark outcomes materially.
  • OpenAI 1.05M models have surcharge rules above 272K input tokens.
  • Sonnet 5 shows intro pricing through Aug 31, 2026; standard pricing is also noted.
  • AA Intel/Code/Agent columns are Artificial Analysis index scores (0-100) served directly from the OpenRouter models API; not every model has a published index.
  • Default sort uses blended cost = (3×input + output) / 4 per 1M tokens, most expensive first, approximating typical coding/agentic usage where input volume dominates but output costs more per token.

Sources

Recommendation shortcut

  • Hardest coding agents: Fable 5, Opus 4.8, GPT-5.6 Sol.
  • Best cost-performance: Sonnet 5, GPT-5.6 Terra.
  • High-volume cheap tasks: Haiku 4.5, GPT-5.6 Luna.