Cheapest models with a 128K+ context window
Sorted by a blended price (three input tokens for every output token), the typical mix for chat and retrieval apps.
| Model | Context | Input | Output | 1,000 chats/day |
|---|---|---|---|---|
| Mistral Nemo Mistral | 131K | $0.019 | $0.030 | $1.30/mo |
| Ling 3.0 Flash VL inclusionAI | 262K | $0.021 | $0.062 | $1.88/mo |
| Ling 3.0 Flash inclusionAI | 262K | $0.021 | $0.063 | $1.89/mo |
| gpt-oss-20b OpenAI | 131K | $0.018 | $0.090 | $2.16/mo |
| Granite 4.0 Micro IBM | 131K | $0.017 | $0.11 | $2.45/mo |
| Nex-N2.5-Mini Nex AGI | 262K | $0.025 | $0.10 | $2.63/mo |
| DeepSeek V4 Flash 0423 DeepSeek | 1.05M | $0.042 | $0.084 | $3.15/mo |
| Qwen3.7 Flash Alibaba (Qwen) | 1M | $0.030 | $0.13 | $3.30/mo |
| Llama 3.1 8B Instruct Meta | 131K | $0.050 | $0.080 | $3.45/mo |
| Schematron V2 Turbo Inference.net | 128K | $0.030 | $0.15 | $3.60/mo |
| Nova Micro 1.0 Amazon | 128K | $0.035 | $0.14 | $3.67/mo |
| Gemma 3 4B Google | 131K | $0.050 | $0.10 | $3.75/mo |
How LLM API pricing works
Almost every language model API bills by the token, a chunk of text of about four characters in English. You pay one rate for tokens the model reads (input) and a higher rate for tokens it writes (output), both quoted per million. A request’s cost is simply input tokens × input price plus output tokens × output price, divided by a million. Chat history counts as input on every turn, so long conversations get more expensive as they grow.
A few extras change the bill:
- Cached input — repeated prompt prefixes, such as a long system prompt, are often billed at a fraction of the normal input rate.
- Batch processing — jobs that can wait hours are commonly offered at around half price.
- Reasoning tokens — thinking models bill their hidden reasoning as output; 67% of the models here offer a reasoning mode.
- Images and files — vision input is converted to tokens or billed per image, depending on the model.
- Long-context tiers — some models charge more once a request passes a size threshold.
Price tiers at a glance
- Under $0.50 per million (blended): 132 models, including Perceptron Mk1.5, Solar Mini 4, GPT-6 Luna Pro, GPT-6 Luna.
- $0.50 – $3 per million (blended): 123 models, including Pareto 26.10 Preview, Aion 3.5 Mini, Command A+, MiMo-V2.6-Pro.
- $3 and up per million (blended): 89 models, including GPT-6.1 Sol Pro, GPT-6.1 Sol, Claude Sonnet 5.5, Ember-1.
Newer is not automatically cheaper. Models released since July 2026 have a median output price of $2.50 per million tokens, against $1.25 for models from before 2026 that are still listed, because new flagship launches sit at the top of the range while many older, smaller models remain available at low rates.
LLM API pricing in practice: three workloads
Per-million prices are hard to picture, so here is what three common workloads cost per month (30 days) on a budget model (Mistral Nemo), a mid-priced model (Gemini 3.8 Flash) and a flagship (Claude Sonnet 5.5):
| Workload | Mistral Nemo | Gemini 3.8 Flash | Claude Sonnet 5.5 |
|---|---|---|---|
| Support chatbot1,500 tokens in, 400 out, 2,000 conversations a day | $2.43 | $158 | $420 |
| Document Q&A (RAG)6,000 tokens of retrieved text in, 300 out, 500 questions a day | $1.85 | $84 | $225 |
| Coding assistant12,000 tokens of code in, 1,500 out, 300 requests a day | $2.46 | $132 | $351 |
The gap between tiers is large enough that model choice matters more than any other optimization. Notice also how the shape of the workload shifts the balance: document Q&A is dominated by input tokens, so a model with cheap input wins, while the coding assistant writes long answers and is dominated by output price. Put your own numbers into the calculator above to see where your app lands.
Glossary of LLM API pricing terms
- Token: the unit models read and write; about four characters or three quarters of an English word.
- Context window: the most tokens one request can hold, prompt and answer together.
- Max output: the cap on tokens the model may write in a single response.
- Blended price: one number that mixes input and output prices for a typical ratio, handy for rough comparisons.
- Open weights: the model can be downloaded and hosted by anyone, which usually creates price competition between hosts.
Five ways to lower your LLM API bill
- Route by difficulty. Send easy requests to a cheap model and escalate only when needed.
- Trim the context. Summarize old chat turns and retrieve only the passages a question needs.
- Cap the output. Ask for short answers and set a maximum token limit.
- Cache and batch. Reuse fixed prompt prefixes and move offline jobs to batch endpoints.
- Measure. Log tokens per request; real usage is often far from the first estimate.
Compare beyond price
The cheapest model is not always the cheapest outcome: a model that needs three retries costs more than one that gets it right the first time. Use the AI model comparison to check context windows, image input, tool calling and open weights side by side before you commit, and try shortlisted models in a chat on SiChatApp.
LLM API pricing questions
Why are output tokens more expensive than input tokens?
Generating text is sequential: each output token needs its own pass through the model, while input tokens are processed in parallel. Across the models listed here, output costs a median 4.0× the input price.
Do reasoning models cost more?
Usually per request, not per token. Reasoning models think in hidden tokens that are billed as output, so the same question can produce several times more billable tokens. Many models let you set the reasoning effort to control this.
Are these the prices I will pay?
They are OpenRouter list prices from 2026-10-02. Direct vendor prices are often the same or close, but discounts for cached prompts, batch jobs and committed volume can lower your effective rate, and long-context surcharges can raise it.
What is the cheapest LLM API right now?
Among general-purpose models with at least a 128K context window, Mistral Nemo from Mistral has the lowest blended price in the 2026-10-02 snapshot: $0.019 per million input tokens and $0.030 per million output tokens. The cheapest model is rarely the best value, though; test whether it handles your task before you switch.
How often is the LLM API pricing table updated?
Whenever we refresh the snapshot from OpenRouter’s public model list; the date is shown on this page. Vendors change prices with little notice, so check the provider before committing to a budget.
