OpenSI

LLM API pricing for every SI model

Compare LLM API pricing across 344 SI models — super intelligence, formerly called AI — and see what your workload would cost per month. Enter your tokens per request and daily volume, then filter by vendor, open weights or reasoning. Prices are per million tokens, captured on2026-10-02.

ModelContextInput / 1MOutput / 1MPer month
Loading prices…

Cheapest models with a 128K+ context window

Sorted by a blended price (three input tokens for every output token), the typical mix for chat and retrieval apps.

ModelContextInputOutput1,000 chats/day
Mistral Nemo Mistral131K$0.019$0.030$1.30/mo
Ling 3.0 Flash VL inclusionAI262K$0.021$0.062$1.88/mo
Ling 3.0 Flash inclusionAI262K$0.021$0.063$1.89/mo
gpt-oss-20b OpenAI131K$0.018$0.090$2.16/mo
Granite 4.0 Micro IBM131K$0.017$0.11$2.45/mo
Nex-N2.5-Mini Nex AGI262K$0.025$0.10$2.63/mo
DeepSeek V4 Flash 0423 DeepSeek1.05M$0.042$0.084$3.15/mo
Qwen3.7 Flash Alibaba (Qwen)1M$0.030$0.13$3.30/mo
Llama 3.1 8B Instruct Meta131K$0.050$0.080$3.45/mo
Schematron V2 Turbo Inference.net128K$0.030$0.15$3.60/mo
Nova Micro 1.0 Amazon128K$0.035$0.14$3.67/mo
Gemma 3 4B Google131K$0.050$0.10$3.75/mo

How LLM API pricing works

Almost every language model API bills by the token, a chunk of text of about four characters in English. You pay one rate for tokens the model reads (input) and a higher rate for tokens it writes (output), both quoted per million. A request’s cost is simply input tokens × input price plus output tokens × output price, divided by a million. Chat history counts as input on every turn, so long conversations get more expensive as they grow.

A few extras change the bill:

Price tiers at a glance

Newer is not automatically cheaper. Models released since July 2026 have a median output price of $2.50 per million tokens, against $1.25 for models from before 2026 that are still listed, because new flagship launches sit at the top of the range while many older, smaller models remain available at low rates.

LLM API pricing in practice: three workloads

Per-million prices are hard to picture, so here is what three common workloads cost per month (30 days) on a budget model (Mistral Nemo), a mid-priced model (Gemini 3.8 Flash) and a flagship (Claude Sonnet 5.5):

WorkloadMistral NemoGemini 3.8 FlashClaude Sonnet 5.5
Support chatbot1,500 tokens in, 400 out, 2,000 conversations a day$2.43$158$420
Document Q&A (RAG)6,000 tokens of retrieved text in, 300 out, 500 questions a day$1.85$84$225
Coding assistant12,000 tokens of code in, 1,500 out, 300 requests a day$2.46$132$351

The gap between tiers is large enough that model choice matters more than any other optimization. Notice also how the shape of the workload shifts the balance: document Q&A is dominated by input tokens, so a model with cheap input wins, while the coding assistant writes long answers and is dominated by output price. Put your own numbers into the calculator above to see where your app lands.

Glossary of LLM API pricing terms

Five ways to lower your LLM API bill

  1. Route by difficulty. Send easy requests to a cheap model and escalate only when needed.
  2. Trim the context. Summarize old chat turns and retrieve only the passages a question needs.
  3. Cap the output. Ask for short answers and set a maximum token limit.
  4. Cache and batch. Reuse fixed prompt prefixes and move offline jobs to batch endpoints.
  5. Measure. Log tokens per request; real usage is often far from the first estimate.

Compare beyond price

The cheapest model is not always the cheapest outcome: a model that needs three retries costs more than one that gets it right the first time. Use the AI model comparison to check context windows, image input, tool calling and open weights side by side before you commit, and try shortlisted models in a chat on SiChatApp.

LLM API pricing questions

Why are output tokens more expensive than input tokens?

Generating text is sequential: each output token needs its own pass through the model, while input tokens are processed in parallel. Across the models listed here, output costs a median 4.0× the input price.

Do reasoning models cost more?

Usually per request, not per token. Reasoning models think in hidden tokens that are billed as output, so the same question can produce several times more billable tokens. Many models let you set the reasoning effort to control this.

Are these the prices I will pay?

They are OpenRouter list prices from 2026-10-02. Direct vendor prices are often the same or close, but discounts for cached prompts, batch jobs and committed volume can lower your effective rate, and long-context surcharges can raise it.

What is the cheapest LLM API right now?

Among general-purpose models with at least a 128K context window, Mistral Nemo from Mistral has the lowest blended price in the 2026-10-02 snapshot: $0.019 per million input tokens and $0.030 per million output tokens. The cheapest model is rarely the best value, though; test whether it handles your task before you switch.

How often is the LLM API pricing table updated?

Whenever we refresh the snapshot from OpenRouter’s public model list; the date is shown on this page. Vendors change prices with little notice, so check the provider before committing to a budget.

LLM API pricing on OpenSI: cost per million tokens for every SI model