English | 中文
AI coding subscriptions sell a monthly fee, not a per-token price. This project works out what each plan actually costs per million tokens, then plots that price against public leaderboard scores to show which plans give the most capability for the money.
Real price = monthly fee ÷ tokens you can actually use in a month.
Open the interactive site → Pick models, filter channels, compare prices and allowances. English / 中文.
How to read the charts
- Each point is one plan × the model it serves. X is the real price in USD per million tokens on a log scale, cheaper to the right. Y is that leaderboard's score.
- Filled squares are subscriptions; hollow diamonds are metered APIs at list price. Both compete on the same frontier.
- The black line is the Pareto frontier: for every point on it, no other point is both cheaper and higher-scoring.
- Subscription prices assume you use the whole allowance. Use half of it and your real price doubles.
Snapshot: 2026-10-01 · 318 plan × model points · all charts, SVG / PNG, both languages
- Monthly allowance. Tokens per month at saturated use. A month is four weeks unless the vendor defines its own monthly pool (Kimi's is 5× the weekly pool). Input, output and cache tokens all count.
- Measured when possible. The best evidence is a direct measurement: tokens used against the change in the dashboard's quota percentage, local usage logs, controlled saturation runs, or an official absolute-token table. Samples that include a token breakdown are converted to dollar worth at public list prices and then to the channel's workload tier; samples without a breakdown and official token tables use raw totals as-is and are flagged "not workload-normalized" on the site.
- Converted when necessary. Dollar or credit pools, and API list prices, are turned into tokens with one standard workload: 97% cache reads, 2.5% fresh input, 0.5% output. This is a comparison convention, not a claim about anyone's real usage. Anthropic models price the fresh-input share at the cache-write rate, and StepFun and Google use a low-cache variant. See CONVENTIONS.md.
- Off-peak pricing (GLM, DeepSeek, MiMo) is shown as separate scenario points, not averaged.
- Confidence. Every row is rated high, medium or low. High means a dashboard back-calculation, a controlled test or an official table. Medium means an official multiplier applied to a high-confidence anchor, or several consistent independent sources. Low means a single report or a cross-plan assumption. Derived values are never presented as measurements.
- One plan, several models. Each model on a plan gets its own point. Those allowances are alternatives and do not add up.
- Scores are copied from each leaderboard and never mixed across boards. Static charts use each model's highest archived configuration.
- Free during a promotion. A model that a plan temporarily doesn't meter is drawn at ≈$0 on a dedicated slot at the right edge. There are 1 such points at the moment: SWE-2 on Devin Pro, until 2026-10-31.
Every adopted value, its evidence and the reason for the choice: DECISIONS.md and the decision_note column of adopted.csv.
Each leaderboard gets its own chart, with its own scores and snapshot date. A model missing from a board is left off that chart but stays in the price and allowance data.
SVG · PNG · 中文 SVG · 中文 PNG · chart shown at the top
Artificial Analysis Intelligence Index v4.3. Scores are not comparable with earlier index versions: a lower number after the version change does not mean a model got worse. Rows marked [AA estimate] are Artificial Analysis's own estimates.
Coding Agent Index v1.5. Each score belongs to a tested harness × model × effort configuration. Higher effort doesn't change the price per token, but it can change how many tokens a task uses.
The WebDev Overall Arena Score. It measures web-app building, not general coding ability.
The 0–100 average task score: requirements 30 + design quality 70. OpenDesign's cost- and speed-weighted recommendation score is not used. All 13 archived models map to adopted points.
The official 66-task leaderboard hosted by Stanford, Harbor and the Laude Institute (snapshot 2026-09-03), with all 22 published configurations. Vendor-reported scores for models the official board doesn't list are added and labelled [self-reported], for example SWE-2 · Devin Pro at 27.3% from Cognition's launch post.
The same 66 tasks run by Artificial Analysis on its own harness (snapshot 2026-09-23). The two Terminal-Bench boards are not interchangeable. On matched configurations the median gap is about 2.6 points, but it can be much larger: Grok 4.7 xhigh scores 37.58 on the official board and 25.76 here.
Pass@1 on 113 tasks, official rows run on mini-swe-agent (snapshot 2026-09-03). Vendor-reported scores are added as supplements and labelled [self-reported].
All 317 priced subscription and API points on one $/MTok scale.
SVG · PNG · Table · 中文 SVG · 中文 PNG · 中文表
The 298 subscription points with a monthly allowance, split into three bands by monthly fee in USD. Each band is ranked on its own. The undivided chart and a hybrid-scale view are in the chart index.
$0–30 · SVG · PNG · Table · 中文 SVG · 中文 PNG · 中文表
Over $30, up to $100 · SVG · PNG · Table · 中文 SVG · 中文 PNG · 中文表
Over $100, up to $300 · SVG · PNG · Table · 中文 SVG · 中文 PNG · 中文表
Chinese charts count tokens in 亿 (100 million): 77.37 亿 = 7.737 billion.
| Points | Count |
|---|---|
| All plan × model points | 318 |
| Subscriptions with a monthly allowance | 298 |
| Free during a promotion (≈$0) | 1 |
| Metered APIs at list price | 19 |
| Leaderboard | Scored points |
|---|---|
| AA Intelligence | 268 |
| AA Coding Agent | 97 |
| Code Arena | 178 |
| Agent Arena | 164 |
| OpenDesign Arena | 88 |
| Terminal-Bench 4.0 | 111 |
| Terminal-Bench 4.0 (AA) | 36 |
| DeepSWE v1.1 | 195 |
The largest plan families are Command Code GOAT (58 points), MiMo Token Plan (32), OpenCode Go (28), Droid Max (27), Ollama (22) and Step Plan (12).
Downloads: adopted values (CSV) · computed points (CSV) / JSON · data notes · dated evidence
Every benchmark configuration, not just the highest per model: the configuration archive (CSV) keeps all 330 records with their original labels, harness, effort, score intervals and task costs. The plan-to-configuration mappings (CSV) hold 1854 explicit references. Unknown harnesses, efforts and intervals stay empty instead of being guessed. The all-configuration interactive chart (Chinese; download and open locally, needs network access for Plotly) lets you switch between configurations and effort levels.
- Real prices are lower bounds. They assume the full allowance is used.
- Benchmark scores are references for a harness × model × effort configuration. They are not tests of each subscription channel, and whether a quota measurement used the same effort and harness is unverified.
- Score intervals are kept (visible on hover in the interactive views) but don't yet affect which points are on the frontier. Confidence labels are qualitative, not error bars.
- Source task costs are kept as separate fields. They are not the cost of the same task on a subscription.
- Tokenizer differences between vendors are not corrected.
- Promotional ≈$0 points have to be re-checked when the promotion ends.
- Rebuild everything: BUILD.md. Rules and conversions: CONVENTIONS.md. Why each value was chosen: DECISIONS.md.
- Have a usage measurement of your own (tokens used vs. quota percentage)? Open a data issue.
- Original software is MIT. Data references include Awesome Coding Plan (CC BY 4.0) and the Caijing article 《Token经济,中国账本》. Attribution, changes and third-party terms: SOURCES.md. Redaction scope of this public edition: PUBLICATION.md.