Compare AI models
Understand the tradeoff between capability and cost.
Capability and cost
Higher score, lower cost
92 options
Method & data
What the axes mean. Output tokens are answer plus reasoning generated per Intelligence Index task; they exclude input and cache traffic. Cost is the publisher’s estimated total task cost. The horizontal axes are logarithmic, so equal spacing represents equal proportional change.
Cohort and frontier. The source has 644 records; 97 meet the non-estimated complete-measure rule, and 92 also report a positive task cost. The curve connects configurations offering the highest score at each resource budget; AI Charts derives it from this cohort. No benchmark families are blended.
Index construction. The publisher’s 10-evaluation index weights agents 30% · coding 20% · scientific 20% · general 30%. These model-level output observations remain separate from coding-agent configurations and total-token measurements.
Source. Publisher methodology · Source terms. Measurements come from the publisher’s public models leaderboard; the snapshot date is a retrieval date, not a model’s evaluation date.
Compare GPT-6 Astra and GPT-5.6 Sol
GPT-6 Astra scores 52.8 and GPT-5.6 Sol scores 47.1 at max effort. Astra generates 7.2% fewer output tokens and costs 63.8% more per task. Astra source · Sol source.
Subscription vs API vs GPUs
One maxed ChatGPT Pro seat implies a monthly token volume. The calculator prices it five ways: the subscription sticker, OpenAI and DeepSeek API rates, GPUs you buy, and GPUs you rent.