Model directory · updated July 2026
Context window, per-token pricing, and what each model is actually good for, cross-checked against provider pricing pages and independent benchmarks. Pricing changes often; treat this as a starting point, not a purchase order.
Pricing shown is per 1M tokens on each provider's standard first-party API, cross-checked across provider documentation and independent tracking sites in the first week of July 2026. Cached-input discounts, batch pricing, and subscription-plan bundling (ChatGPT Plus, Claude Pro, Gemini Advanced) are not reflected in the table, those are usually cheaper per use than raw API rates for casual use. Always confirm current pricing on the provider's own page before making a purchasing decision.
A 1M-token window sounds enormous, but recall accuracy for most models drops noticeably past a few hundred thousand tokens when the answer depends on a small detail buried in the middle.
A frequent, costly mistake is routing every request to the most powerful (and expensive) tier. Match model capability to task difficulty, a mid or small tier is correct for most classification, extraction, or routing work.
Every major lab has cut prices at least once in the last year. Build for model-agnostic routing where you can, and revisit vendor choice quarterly rather than locking in once.
The pricing calculator (coming soon) turns this table into a real monthly cost estimate for your workload.
Back to the news feed