On this page
The most-linked lists of free LLM APIs live on GitHub and are kept up by hand. They are useful, and they go stale quietly: a limit changes on a docs page and nobody notices for weeks. GitHub Models, a free option many of those lists carried, was retired on 30 July 2026.
So this page does two things differently. Every provider row carries the date I read it on the provider's own page, and every day a script re-reads those pages and flags any row whose wording has changed. The OpenRouter half needs no hand-checking at all: its public API lists every free model, and the tracker logs each one that appears or disappears.
Free tiers matter to me for a plain reason. Gravity lets people start with one free agent, and I wanted to know exactly what free looks like one layer down.
Free LLM APIs on 4 October 2026
- 19 text models are free on OpenRouter on 4 October 2026, in the first snapshot of this tracker; the week-on-week change starts on 11 October 2026. 17 of them are ":free" variants, which OpenRouter caps at 20 requests a minute and 50 a day until an account has bought 10 credits.
- 18 of the 19 free models take 128K tokens of context or more. The largest window is Inkling Small at 1,048,576 tokens, as listed on 4 October 2026.
- 12 providers have a free API allowance I could verify on their own pages: 6 standing free tiers, 2 monthly credits, 2 one-time quotas and 2 trials. Every row in the table below carries the date it was read.
- The highest published daily request cap is Groq's: 1,000 requests a day per model and 200,000 tokens a day on its free plan. Google, Mistral and Cloudflare publish no comparable daily request figure for their free tiers.
- 6 providers I checked have no free API tier. That includes GitHub Models, which its own docs say was fully retired on 30 July 2026.
Updated . OpenRouter models API (466 models listed); provider rows verified on each provider's own page, dates in the table.
Cite this page: "Free LLM API tiers", Gravity, updated , https://gravity.fast/data/free-llm-api-tiers/. The data is licensed CC BY 4.0: reuse it anywhere with a link back. Downloads: data.csv (provider free tiers) and free-models.csv (today's free OpenRouter models).
Provider free tiers
Four kinds of "free" hide behind the same word, and the type column keeps them apart. A standing free tier renews and has no end date. A monthly credit is a small dollar amount that resets. A one-time quota is spent once. A trial expires, and sometimes needs a card first.
| Provider and plan | Type | Free limits | Card needed | Terms and caveats | Verified |
|---|---|---|---|---|---|
| Google AI Studio (Gemini API) Free tier | Standing free tier | Free of charge on the listed models. Google no longer publishes free-tier requests or tokens per minute or day in its docs; your project's limits are shown in AI Studio. Requests per day reset at midnight Pacific. Models: Gemini 3.8, 3.7, 3.6 and 3.5 Flash, 3.5 and 3.1 Flash-Lite, 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite, Gemma 4. Gemini 3.1 Pro is not free. | No billing account (inferred from the tier table) | Free-tier content is used to improve Google's products. Grounding with Google Search is not available on the free tier. Terms page | Wording re-checked 4 October 2026 |
| Groq Free plan | Standing free tier | gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B: 30 requests a minute, 1,000 a day, 8,000 tokens a minute, 200,000 tokens a day each. Models: gpt-oss-120b, gpt-oss-20b, Qwen3.8 27B (plus Whisper speech-to-text) | Not stated | Groq calls the table a high-level summary with possible exceptions; your organisation's exact limits are on the console Limits page. | Wording re-checked 4 October 2026 |
| OpenRouter Free model variants (IDs ending in :free) | Standing free tier | 20 requests a minute. 50 requests a day if you have bought less than 10 credits in total, 1,000 a day once you have bought 10 or more. Models: Every model variant whose ID ends in :free; see the model table below | Not stated; buying 10 credits lifts the daily cap to 1,000 | A negative balance can cause errors even on free models. Free variants can be added or withdrawn by the model host at any time. | Wording re-checked 4 October 2026 |
| Cloudflare Workers AI Workers Free allocation | Standing free tier | 10,000 Neurons a day free (Neurons are Cloudflare's compute unit, converted to tokens per model on the pricing page). Text generation is capped at 300 requests a minute on every plan. Models: Workers AI catalog, except models Cloudflare marks as needing a paid plan (Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 among them) | Not stated | Above the daily allocation you need Workers Paid at $0.011 per 1,000 Neurons. Terms page | Wording re-checked 4 October 2026 |
| Mistral AI Free plan | Monthly free credit | $10 a month in API credits. Rate limits apply but are shown only in the console. Models: Mistral's models through Studio and the API | No card required (per Mistral's docs) | The free plan's model-training setting could not be read from the public pricing table. Terms page | Wording re-checked 4 October 2026 |
| Cohere Trial key | Standing free tier | 1,000 API calls a month in total; chat is limited to 20 requests a minute per model. Models: Command A family, Command R+, Command R, Command R7B, North Mini Code | Not stated | Trial keys are for evaluation; production use needs a paid production key. Terms page | Wording re-checked 4 October 2026 |
| SambaNova Cloud Free tier | Standing free tier | 20 requests a minute, 20 a day and 200,000 tokens a day per model. Models: DeepSeek-V3.1, Llama 3.3 70B, gpt-oss-120b; DeepSeek-V3.2 and Gemma 4 31B in preview | No payment method (the free tier applies until one is linked) | Preview models are for evaluation only, not production, and can be removed at short notice. | Wording re-checked 4 October 2026 |
| Hugging Face Inference Providers Free account monthly credits | Monthly free credit | $0.10 of credit a month, which Hugging Face says is subject to change. PRO accounts get $2.00. Models: Models served through Inference Providers when requests are routed by Hugging Face | Not stated | Credits do not apply when you use your own key for a provider. | Wording re-checked 4 October 2026 |
| Alibaba Cloud Model Studio (international) New-user free quota | One-time free quota | Typically 1,000,000 tokens per model, input and output combined, valid for 90 days from activation. Models: Qwen models in the Singapore region (each model and dated snapshot has its own quota) | Not stated | Usage is billed automatically after the quota runs out unless the Free Quota Only switch is turned on (it is off by default). Real-time inference only. | Wording re-checked 4 October 2026 |
| Scaleway Generative APIs Free tier (serverless) | One-time free quota | Up to 1,000,000 tokens and 60 minutes of audio transcription at no cost, applied to the most expensive tokens first. The FAQ does not say whether this is one-time or monthly. Models: Serverless models billed by tokens | Not stated | Read from Scaleway's own documentation source because the live docs page blocked automated reads on the verification date. Terms page | Wording re-checked 4 October 2026 |
| NVIDIA API Catalog Free trial credits (NVIDIA Developer Program) | Trial | Free credits for prototyping. NVIDIA's docs state no credit amount and no rate limit. Models: NVIDIA-hosted NIM endpoints on build.nvidia.com | Not stated | Positioned for prototyping, not production. | Wording re-checked 4 October 2026 |
| Cerebras Free trial | Trial | $5 of credit that expires 30 days after it is granted. 5 requests a minute, 30,000 uncached tokens a minute, 1,000,000 tokens a day per model. Models: gpt-oss-120b, Qwen 3.8 27B | Card required (verified payment method) | Cerebras's own FAQ answers "Is there a permanently free tier?" with no. | Wording re-checked 4 October 2026 |
Limits as written on each provider's own page on the verified date. "Not stated" means the page does not say; it is not a yes or a no. Context windows are not published on any of these limits pages, so they are left to the model table below. Every day the tracker re-fetches each source page and checks that the verified wording is still there.
Checked: no free API tier
These providers come up in searches for free LLM APIs. On the date shown, their own pages offered no free allowance.
| Provider | What its own pages say | Checked |
|---|---|---|
| GitHub Models | Retired. GitHub's docs say the service was fully retired on 30 July 2026, including the playground and inference API. | |
| Together AI | No free trial; a minimum $5 credit purchase is required. | |
| Fireworks AI | No free credit stated; accounts without a payment method are limited to 10 requests a minute. | |
| DeepSeek | No free tier or free amount stated on the pricing page. | |
| xAI | No free credits mentioned on the rate-limits or pricing pages; the lowest tier starts at $0 spent. | |
| Chutes | The pricing FAQ says there is no free tier at this time. |
Free models on OpenRouter today
OpenRouter lists some models at $0 for input and output, mostly as ":free" variants that hosts sponsor for a while. The list below is read from OpenRouter's public models API each day. Free variants share OpenRouter's free-usage limits, set out in the provider table above.
| Model | Maker | Context | Free how | Listed on OpenRouter | Expiry date |
|---|---|---|---|---|---|
Ling 3.1 Flashinclusionai/ling-3.1-flash | inclusionai | 262K | Free at every route | None listed | |
Apodex 1.1 Miniapodex/apodex-1.1-mini:free | apodex | 262K | Free variant | None listed | |
Space Bunny Alphastealth/space-bunny-alpha | stealth | 1M | Free at every route | 5 October 2026 | |
Ling 3.0 Flash Santeinclusionai/ling-3.0-flash-sante:free | inclusionai | 262K | Free variant | None listed | |
Qwen3.8 27Bqwen/qwen3.8-27b:free | qwen | 262K | Yes, a paid version is also listed | None listed | |
Dots3-Note Previewdots-studio/dots-3-note-preview:free | dots-studio | 512K | Free variant | 31 December 2026 | |
LFM2.5-2.6Bliquid/lfm-2.5-2.6b:free | liquid | 66K | Free variant | None listed | |
Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning:free | nvidia | 1M | Yes, a paid version is also listed | None listed | |
Inkling Smallthinkingmachines/inkling-small:free | thinkingmachines | 1.05M | Yes, a paid version is also listed | None listed | |
Laguna S 2.1poolside/laguna-s-2.1:free | poolside | 262K | Yes, a paid version is also listed | 31 October 2026 | |
Inklingthinkingmachines/inkling:free | thinkingmachines | 1.05M | Yes, a paid version is also listed | None listed | |
Laguna XS 2.1poolside/laguna-xs-2.1:free | poolside | 262K | Yes, a paid version is also listed | 31 October 2026 | |
North Mini Codecohere/north-mini-code:free | cohere | 256K | Free variant | None listed | |
Nemotron 3.5 Content Safetynvidia/nemotron-3.5-content-safety:free | nvidia | 128K | Yes, a paid version is also listed | None listed | |
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b:free | nvidia | 1M | Yes, a paid version is also listed | None listed | |
Nemotron 3 Nano Omninvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | nvidia | 256K | Free variant | None listed | |
Gemma 4 26B A4Bgoogle/gemma-4-26b-a4b-it:free | 262K | Yes, a paid version is also listed | None listed | ||
Gemma 4 31Bgoogle/gemma-4-31b-it:free | 262K | Yes, a paid version is also listed | None listed | ||
Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b:free | nvidia | 262K | Yes, a paid version is also listed | None listed |
Every text-in, text-out model priced at $0 for input and output in the OpenRouter models API on 4 October 2026, newest first. OpenRouter's own routers and image or audio models are excluded. "Expiry date" is the date OpenRouter lists for the model's retirement, where it lists one.
Which one to start with
For prototyping against a strong general model with the least setup, the Gemini API free tier is the obvious first stop, as long as your prompts can be used to improve Google's products. Check that before you send anything sensitive.
For the most published requests per day on open models, Groq's free plan is the one to beat, and its limits are written down per model. SambaNova's free tier covers larger models but allows far fewer requests a day.
For trying many models through one key, OpenRouter's free variants are the flexible choice; buying 10 credits once lifts the daily cap from 50 to 1,000 requests. Expect individual free models to come and go, which the change log below records.
If you are building on Cloudflare already, Workers AI's daily allocation is the natural fit. Mistral's $10 monthly credit and Hugging Face's $0.10 suit small experiments rather than a running product.
Once a project outgrows free limits, the question becomes price. The LLM API pricing history tracks the cheapest paid price for a given level of capability every day, and the floor is a few cents per million tokens.
Change log
Free models added to and removed from OpenRouter, provider pages whose wording changed, and re-verifications, newest first.
- : tracker started. 19 free text models on OpenRouter; 12 provider free tiers and 6 providers without one verified on their own pages; GitHub Models recorded as retired (30 July 2026)
Methodology
What counts as a free tier. An allowance that lets a new account call an LLM text model through an API without paying, stated on the provider's own pricing, limits or docs pages. Chat apps and playgrounds without API access do not count. Each row records the limits as the provider writes them, the card requirement only where a page states it, and the caveats the provider attaches, such as using free-tier content for training or limiting use to evaluation.
How rows are verified. I read every provider row on the provider's own page on the verified date shown. Where a page renders its limits with JavaScript, the figures come from the page source. Scaleway's live docs blocked automated reads on 4 October 2026, so its row comes from Scaleway's own documentation repository on GitHub.
How rows stay current. Every day the tracker fetches each source page and checks that two or three short pieces of the verified wording, such as "1,000 API calls a month", are still present. If one disappears, the row is flagged as re-verifying and the change is logged; I then re-read the page and update the row with a new date. A page can change a limit without touching the checked wording, so the verified date remains the date to cite.
OpenRouter free models. Every model in the OpenRouter models API with a $0 prompt price and a $0 completion price whose input and output are text. OpenRouter's own routers (such as openrouter/free) and music or image models are excluded, which is why the count here can be a few lower than a raw count of $0 entries. Additions and removals compare each day's list with the previous day's.
Known gaps. Google and Mistral show free-tier rate limits only inside their consoles, so those cells say so instead of giving a number. NVIDIA does not publish its trial credit amount. Context windows are not stated on any provider's limits page; the OpenRouter table carries them for the models it lists. For free AI agent products, as opposed to free model APIs, see the best free AI agents and the AI agent platform free tier comparison.
FAQ
1. Which LLM APIs have a free tier in 2026?
On 4 October 2026, 6 providers run a standing free API tier: Google AI Studio (Gemini API), Groq, OpenRouter, Cloudflare Workers AI, Cohere, SambaNova Cloud. Mistral AI and Hugging Face Inference Providers give a small monthly credit instead, Alibaba Cloud Model Studio (international) and Scaleway Generative APIs give a one-time quota, and NVIDIA API Catalog and Cerebras offer trials. OpenRouter also lists 19 models that cost nothing to call.
2. Is the Gemini API free?
Yes, for most models. On 4 October 2026 Google's pricing page marked these as free of charge on the free tier: Gemini 3.8, 3.7, 3.6 and 3.5 Flash, 3.5 and 3.1 Flash-Lite, 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite, Gemma 4. Gemini 3.1 Pro is not free. Google no longer publishes the free tier's request or token limits in its docs; each project sees its own limits in AI Studio. Content sent on the free tier is used to improve Google's products.
3. How many free models does OpenRouter have?
19 text models were priced at $0 on 4 October 2026, 17 of them as ":free" variants. Free variants are limited to 20 requests a minute and 50 a day, or 1,000 a day once an account has bought 10 credits. Free models come and go, which is why this page logs every addition and removal by date.
4. Is the Groq API free?
Groq has a free plan. On 4 October 2026 its rate-limits page listed gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B: 30 requests a minute, 1,000 a day, 8,000 tokens a minute, 200,000 tokens a day each. Groq describes the table as a summary; each organisation's exact limits are on its console Limits page.
5. Which free LLM API has the highest limits?
By published daily requests, Groq: 1,000 requests a day per model on its free plan. OpenRouter allows 1,000 a day on free variants after a 10-credit purchase. Cloudflare Workers AI counts its free allowance in its own compute unit instead: 10,000 Neurons a day free. Google publishes no free-tier numbers to compare.
6. Can I use a free LLM API in production?
Usually not as your only path. Several providers attach conditions to free use: Google uses free-tier content to improve its products, Cohere trial keys are for evaluation, SambaNova preview models are not for production, and free OpenRouter variants can disappear without notice. Free tiers are best for prototypes, tests and low-volume tools, with a paid key ready as a fallback.
7. Is GitHub Models still free?
No. GitHub's docs say the service was fully retired on 30 July 2026, including the playground and inference API. It was a popular free option for developers, so older lists of free LLM APIs still include it.
8. How is this list kept current?
The OpenRouter model list is read from its public API every day and compared with the day before. The provider rows are read by hand on each provider's own page and dated; every day the tracker re-fetches those pages and flags any row whose verified wording has changed. The data is CC BY 4.0.
Sources
- OpenRouter, models API, read daily since 4 October 2026, and API rate limits, read 4 October 2026.
- Google, Gemini Developer API pricing and rate limits, read 4 October 2026.
- Groq, rate limits, read 4 October 2026.
- Cloudflare, Workers AI pricing and limits, read 4 October 2026.
- Mistral AI, pricing, read 4 October 2026.
- Cohere, rate limits, read 4 October 2026.
- SambaNova, rate limits, read 4 October 2026.
- Hugging Face, Inference Providers pricing, read 4 October 2026.
- Alibaba Cloud, Model Studio free quota, read 4 October 2026.
- Scaleway, Generative APIs FAQ (documentation source), read 4 October 2026.
- NVIDIA, API catalog FAQ, read 4 October 2026.
- Cerebras, rate limits, read 4 October 2026.
- GitHub, GitHub Models documentation (retirement notice), read 4 October 2026.
Related reading: the cheapest AI agent platforms, the AI agent pricing index, and how to create your own AI agent for free. Gravity's plans are on the pricing page.