Comparisons
Closed-source LLMs compared
Six proprietary LLM APIs, scored uniformly on the same seven criteria - with a source link and date for every entry, so you can verify the numbers yourself. This category is aimed at developers and technical decision-makers building an application on an API - not end users of a finished chat app (see our "AI Chat Assistants" category for that). Switch between the individual and business perspective at the top of the table. At the bottom you'll find scenario recommendations instead of a single "winner": which API fits depends on your specific situation.
A comparison of closed-source LLM APIs weighs offerings like GPT, Claude, Gemini, Grok, Cohere and Kimi K3 from a developer's perspective - from price per token to rate limits to enterprise SLAs - for anyone building their own application on an API.
How we score →At a glance
Calculated automatically from the scores below - the provider highlighted in green has the highest average score across all criteria, not an editorial pick. Toggle providers in the legend on or off.
Last data review: 08/17/2026
Click a row for strengths, weaknesses, and our take.
| Tool | Price | Value for money | Privacy & hosting | Performance & context window | Rate limits & availability | Ecosystem & tooling | Multimodality | Enterprise readiness & compliance | Ideal for | Learn more |
|---|---|---|---|---|---|---|---|---|---|---|
GPT (API) OpenAI | $5/M input · $30/M output tokens (flagship Sol) · cheaper Terra ($2/$12) and Luna ($0.20/$1.20) tiers | 5 of 5 | 4 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | 4 of 5 | 5 of 5 | Developers wanting maximum raw performance at a known price +1 | |
Claude (API) Anthropic | $10/M input · $50/M output tokens (Fable 5) · cheaper Opus 5 ($5/$25) and Sonnet 5/Haiku tiers | 3 of 5 | 2 of 5 | 5 of 5 | 3 of 5 | 5 of 5 | 2 of 5 | 4 of 5 | Coding- and analysis-heavy projects needing maximum text quality +1 | |
Gemini (API) | $2/M input (up to 200k tokens, then $4) · $12/M output tokens (up to 200k tokens, then $18) (3.1 Pro) · cheaper Flash tiers | 4 of 5 | 4 of 5 | 5 of 5 | 3 of 5 | 3 of 5 | 5 of 5 | 5 of 5 | Multimodal projects (image, audio, video in one model) +1 | |
Grok (API) xAI | $2/M input · $6/M output tokens (<200k tokens; unchanged from 4.5, now also applies to 4.6) · not yet available in the EU per xAI · predecessor 4.3 ($1.25/$2.50) usable for EU customers | 5 of 5 | 2 of 5 | 4 of 5 | 3 of 5 | 4 of 5 | 4 of 5 | 2 of 5 | Price-conscious developers outside the EU wanting an OpenAI-compatible API +1 | |
Cohere Cohere | Command A+ now Apache 2.0/open weights, free to download · Model Vault from $4/hour or $2,500/month · older Command R/R+ tiers still paid | 3 of 5 | 5 of 5 | 2 of 5 | 3 of 5 | 4 of 5 | 2 of 5 | 5 of 5 | RAG-/retrieval-heavy projects needing rerank/embed +1 | |
Kimi K3 Moonshot AI | $0.30/M input (cache hit) · $3/M input (cache miss) · $15/M output tokens · currently API-only, no open weights | 4 of 5 | 1 of 5 | 5 of 5 | 2 of 5 | 2 of 5 | 3 of 5 | 1 of 5 | Price-conscious developers with non-critical data wanting maximum raw performance +1 | |
Which tool fits you?
The highest raw-performance benchmark (coding/analysis)
GPT (API)
The highest benchmark score among all six APIs compared (96.2% SWE-bench) at an unchanged flagship price.
Building agents and tools with MCP
Claude (API)
The origin and deepest integration of the MCP standard, plus one of the highest benchmark scores in this comparison.
Multimodal applications from a single provider
Gemini (API)
The broadest multimodality (image, video, audio, four image-generation tiers) and the best certification standing.
Maximum cost efficiency outside the EU
Grok (API)
Still the cheapest flagship pricing tier among the APIs - for EU customers, currently only the predecessor Grok 4.3 is usable.
A cheap, high-performance API for non-critical data
Kimi K3
The third-highest benchmark score in this comparison at a low cost - but without certifications or EU hosting, so suitable only for non-critical data.
Key takeaways
- GPT (API) scores 96.2% on SWE-bench with GPT-5.6 Sol at an unchanged flagship price of $5/$30 per million tokens - though the top of the leaderboard now belongs to Claude Opus 5 at 97.0%.
- Claude (API) fields the strongest model in this comparison with Opus 5 (97.0%) - and at half the price of its own Fable 5 ($5/$25 vs. $10/$50). Anthropic recommends Opus 5 as the default starting point; Fable 5 remains the special case for maximum capability, but without a zero-data-retention option.
- Kimi K3 (Moonshot AI) is new in this comparison: the third-highest benchmark score (93.4%) at a fraction of the cost - but with no certifications, EU hosting, or ZDR option at all, and as of today still without open weights (promised by Moonshot for 2026-07-27).
- Grok (API) made the biggest benchmark jump of any vendor with 4.5, but per xAI itself it isn't yet available in the EU - EU customers should stick with the predecessor Grok 4.3 for now.
- Cohere surprisingly pivoted to open weights with Command A+ (Apache 2.0, free) - but raw performance remains the weakest in this comparison.
- Gemini (API) still offers the broadest multimodality and the best certification coverage among the six APIs.
→ Self-host, use an API, or buy a ready-made solution? Build vs. buy vs. API
→ Find the right approach for your own data: RAG-or-Fine-Tuning Finder
Frequently asked questions
Which LLM API is best for coding applications?
GPT (API) now scores the highest benchmark among the six APIs compared here with GPT-5.6 Sol (96.2% SWE-bench) - just ahead of Claude's new Fable 5 (95.0%) and Kimi K3 (93.4%). Claude, though, remains the origin of the MCP (Model Context Protocol) standard for tool integrations and still has the deepest ecosystem there.
Which LLM API is the cheapest?
Grok (API) still has the cheapest flagship pricing per million tokens among the APIs with a full enterprise offering - though for EU customers, only the predecessor Grok 4.3 is currently available. Kimi K3 undercuts Grok on cache-hit input pricing, but comes in above it on output pricing.
Is there an LLM API with on-premise deployment for regulated industries?
Cohere is still the only option among the six APIs compared with true private/on-premise deployment (Model Vault, BYOC) - since May 2026 also with open weights (Command A+, Apache 2.0) for anyone who wants to self-host.
Is Kimi K3 an open or a closed model?
As of 2026-07-21, Kimi K3 is a pure API model with no publicly available weights - Moonshot AI has announced an open-weight release for 2026-07-27, but hasn't delivered it yet. That's why Kimi K3 currently sits in this category rather than our open-source LLMs category; if the announcement lands, we'll reassess where it belongs.
Why is Cohere listed here even though Command A+ has open weights?
Cohere is evaluated in this category primarily for its API/enterprise business model (Model Vault, managed Rerank/Embed services, a sales process for enterprise terms), not for the weights themselves. Unlike the models in our open-source category, Cohere's managed platform is the focus here - Command A+'s open weights are an additional self-hosting option, not a business-model pivot.
What's the difference between this category and "AI Chat Assistants"?
This category evaluates APIs from a developer's perspective (price per token, context window, SDKs, rate limits) for building custom applications. The "AI Chat Assistants" category instead compares the finished consumer chat apps (ChatGPT, Claude.ai, the Gemini app, etc.) for end users.