llm api pricing

LLM API Pricing 2026

Claude, GPT, and Gemini cost per 1M tokens — plus what 4 common SME workloads actually cost on each model. Vendor-neutral. Sourced 2026-06-02.

Per-1M-token list prices

ModelVendorInput / 1MOutput / 1MCached input
Claude Haiku 3.5Anthropic$0.80$4.00$0.08 input
Claude Sonnet 4.6Anthropic$3.00$15.00$0.30 input
Claude Opus 4Anthropic$15.00$75.00$1.50 input
GPT-4o-miniOpenAI$0.15$0.60$0.075 input
GPT-4oOpenAI$2.50$10.00$1.25 input
GPT-5OpenAI$8.00$32.00$4.00 input
Gemini 2.0 FlashGoogle$0.10$0.40$0.025 input
Gemini 2.0 ProGoogle$1.25$5.00$0.31 input

Monthly cost at 4 common SME workloads

Picking a model in the abstract is hard; picking one against a specific workload is easy. Below: 4 NodeSparks reference architectures and what each costs across 5 models.

Use caseHaikuSonnetGPT-4o-miniGPT-4oGemini Flash
Small outreach agent (500 emails/mo personalized)500k in / 50k out$0.60$2.25$0.11$1.75$0.07
Slack invoice agent (200 invoices/mo with vision)4M in / 200k out$4.00$15.00$0.72$12.00$0.48
Mid-size ops automation (10k workflow runs/mo)50M in / 5M out$60$225$10.50$175$7
Large content engagement agent (1M social actions/mo)500M in / 50M out$600$2,250$105$1,750$70

Per-model notes

  • Claude Haiku 3.5. Fastest Claude. Best for classification, extraction, summarization at scale. 10× prompt-cache discount on cached tokens.
  • Claude Sonnet 4.6. Default workhorse. Best balance of cost, reasoning, and speed for most production use. 200k context window.
  • Claude Opus 4. Highest reasoning capability. Use only when Sonnet visibly struggles. Premium price for premium tasks.
  • GPT-4o-mini. Cheapest competent model. Best for high-volume personalization, simple classification. 128k context.
  • GPT-4o. OpenAI workhorse. Comparable to Claude Sonnet on most tasks; cheaper. 128k context.
  • GPT-5. OpenAI flagship reasoning model. Use sparingly — most tasks do not need GPT-5 quality.
  • Gemini 2.0 Flash. Cheapest in the comparison. 1M context window. Best for ingesting large documents (RAG, code review at scale).
  • Gemini 2.0 Pro. Comparable to Claude Sonnet on cost; weaker on reasoning, stronger on multimodal. 2M context window.

Methodology

List prices captured on 2026-06-02 from each vendor's public pricing page (anthropic.com/pricing, openai.com/api/pricing, ai.google.dev/pricing). Volumes assume single-region US/EU endpoint deployment; cross-region inference is typically ~10% more expensive.

Workload examples assume average prompt size and a reasonable cache hit rate (~60% for the Sonnet/GPT-4o numbers). Real-world cost can deviate ±30% depending on prompt engineering quality.

How to cite

NodeSparks (2026). "LLM API Pricing 2026 — Real Cost at SME Volumes." https://www.nodesparks.com/data/llm-api-pricing. CC BY 4.0.

Related

Download

⬇ Download CSV
Per-1M-token list prices, CC BY 4.0.

Last updated 2026-06-02.

Still paying for tools you could own?
We replace the SaaS stack and the manual ops work eating your team's time. One custom system, owned by you.

Let's start with a real conversation.We’re ready when you are.