On this page

The most-linked lists of free LLM APIs live on GitHub and are kept up by hand. They are useful, and they go stale quietly: a limit changes on a docs page and nobody notices for weeks. GitHub Models, a free option many of those lists carried, was retired on 30 July 2026.

So this page does two things differently. Every provider row carries the date I read it on the provider's own page, and every day a script re-reads those pages and flags any row whose wording has changed. The OpenRouter half needs no hand-checking at all: its public API lists every free model, and the tracker logs each one that appears or disappears.

Free tiers matter to me for a plain reason. Gravity lets people start with one free agent, and I wanted to know exactly what free looks like one layer down.

Headline numbers

Free LLM APIs on 4 October 2026

  1. 19 text models are free on OpenRouter on 4 October 2026, in the first snapshot of this tracker; the week-on-week change starts on 11 October 2026. 17 of them are ":free" variants, which OpenRouter caps at 20 requests a minute and 50 a day until an account has bought 10 credits.
  2. 18 of the 19 free models take 128K tokens of context or more. The largest window is Inkling Small at 1,048,576 tokens, as listed on 4 October 2026.
  3. 12 providers have a free API allowance I could verify on their own pages: 6 standing free tiers, 2 monthly credits, 2 one-time quotas and 2 trials. Every row in the table below carries the date it was read.
  4. The highest published daily request cap is Groq's: 1,000 requests a day per model and 200,000 tokens a day on its free plan. Google, Mistral and Cloudflare publish no comparable daily request figure for their free tiers.
  5. 6 providers I checked have no free API tier. That includes GitHub Models, which its own docs say was fully retired on 30 July 2026.

Updated . OpenRouter models API (466 models listed); provider rows verified on each provider's own page, dates in the table.

Cite this page: "Free LLM API tiers", Gravity, updated , https://gravity.fast/blog/free-llm-api-tiers/. The data is licensed CC BY 4.0: reuse it anywhere with a link back. Downloads: data.csv (provider free tiers) and free-models.csv (today's free OpenRouter models).

Provider free tiers

Four kinds of "free" hide behind the same word, and the type column keeps them apart. A standing free tier renews and has no end date. A monthly credit is a small dollar amount that resets. A one-time quota is spent once. A trial expires, and sometimes needs a card first.

Provider and planTypeFree limitsCard neededTerms and caveatsVerified
Google AI Studio (Gemini API)
Free tier
Standing free tierFree of charge on the listed models. Google no longer publishes free-tier requests or tokens per minute or day in its docs; your project's limits are shown in AI Studio. Requests per day reset at midnight Pacific.
Models: Gemini 3.8, 3.7, 3.6 and 3.5 Flash, 3.5 and 3.1 Flash-Lite, 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite, Gemma 4. Gemini 3.1 Pro is not free.
No billing account (inferred from the tier table)Free-tier content is used to improve Google's products. Grounding with Google Search is not available on the free tier. Terms page
Wording re-checked 4 October 2026
Groq
Free plan
Standing free tiergpt-oss-120b, gpt-oss-20b and Qwen3.8 27B: 30 requests a minute, 1,000 a day, 8,000 tokens a minute, 200,000 tokens a day each.
Models: gpt-oss-120b, gpt-oss-20b, Qwen3.8 27B (plus Whisper speech-to-text)
Not statedGroq calls the table a high-level summary with possible exceptions; your organisation's exact limits are on the console Limits page.
Wording re-checked 4 October 2026
OpenRouter
Free model variants (IDs ending in :free)
Standing free tier20 requests a minute. 50 requests a day if you have bought less than 10 credits in total, 1,000 a day once you have bought 10 or more.
Models: Every model variant whose ID ends in :free; see the model table below
Not stated; buying 10 credits lifts the daily cap to 1,000A negative balance can cause errors even on free models. Free variants can be added or withdrawn by the model host at any time.
Wording re-checked 4 October 2026
Cloudflare Workers AI
Workers Free allocation
Standing free tier10,000 Neurons a day free (Neurons are Cloudflare's compute unit, converted to tokens per model on the pricing page). Text generation is capped at 300 requests a minute on every plan.
Models: Workers AI catalog, except models Cloudflare marks as needing a paid plan (Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 among them)
Not statedAbove the daily allocation you need Workers Paid at $0.011 per 1,000 Neurons. Terms page
Wording re-checked 4 October 2026
Mistral AI
Free plan
Monthly free credit$10 a month in API credits. Rate limits apply but are shown only in the console.
Models: Mistral's models through Studio and the API
No card required (per Mistral's docs)The free plan's model-training setting could not be read from the public pricing table. Terms page
Wording re-checked 4 October 2026
Cohere
Trial key
Standing free tier1,000 API calls a month in total; chat is limited to 20 requests a minute per model.
Models: Command A family, Command R+, Command R, Command R7B, North Mini Code
Not statedTrial keys are for evaluation; production use needs a paid production key. Terms page
Wording re-checked 4 October 2026
SambaNova Cloud
Free tier
Standing free tier20 requests a minute, 20 a day and 200,000 tokens a day per model.
Models: DeepSeek-V3.1, Llama 3.3 70B, gpt-oss-120b; DeepSeek-V3.2 and Gemma 4 31B in preview
No payment method (the free tier applies until one is linked)Preview models are for evaluation only, not production, and can be removed at short notice.
Wording re-checked 4 October 2026
Hugging Face Inference Providers
Free account monthly credits
Monthly free credit$0.10 of credit a month, which Hugging Face says is subject to change. PRO accounts get $2.00.
Models: Models served through Inference Providers when requests are routed by Hugging Face
Not statedCredits do not apply when you use your own key for a provider.
Wording re-checked 4 October 2026
Alibaba Cloud Model Studio (international)
New-user free quota
One-time free quotaTypically 1,000,000 tokens per model, input and output combined, valid for 90 days from activation.
Models: Qwen models in the Singapore region (each model and dated snapshot has its own quota)
Not statedUsage is billed automatically after the quota runs out unless the Free Quota Only switch is turned on (it is off by default). Real-time inference only.
Wording re-checked 4 October 2026
Scaleway Generative APIs
Free tier (serverless)
One-time free quotaUp to 1,000,000 tokens and 60 minutes of audio transcription at no cost, applied to the most expensive tokens first. The FAQ does not say whether this is one-time or monthly.
Models: Serverless models billed by tokens
Not statedRead from Scaleway's own documentation source because the live docs page blocked automated reads on the verification date. Terms page
Wording re-checked 4 October 2026
NVIDIA API Catalog
Free trial credits (NVIDIA Developer Program)
TrialFree credits for prototyping. NVIDIA's docs state no credit amount and no rate limit.
Models: NVIDIA-hosted NIM endpoints on build.nvidia.com
Not statedPositioned for prototyping, not production.
Wording re-checked 4 October 2026
Cerebras
Free trial
Trial$5 of credit that expires 30 days after it is granted. 5 requests a minute, 30,000 uncached tokens a minute, 1,000,000 tokens a day per model.
Models: gpt-oss-120b, Qwen 3.8 27B
Card required (verified payment method)Cerebras's own FAQ answers "Is there a permanently free tier?" with no.
Wording re-checked 4 October 2026

Limits as written on each provider's own page on the verified date. "Not stated" means the page does not say; it is not a yes or a no. Context windows are not published on any of these limits pages, so they are left to the model table below. Every day the tracker re-fetches each source page and checks that the verified wording is still there.

Checked: no free API tier

These providers come up in searches for free LLM APIs. On the date shown, their own pages offered no free allowance.

ProviderWhat its own pages sayChecked
GitHub ModelsRetired. GitHub's docs say the service was fully retired on 30 July 2026, including the playground and inference API.
Together AINo free trial; a minimum $5 credit purchase is required.
Fireworks AINo free credit stated; accounts without a payment method are limited to 10 requests a minute.
DeepSeekNo free tier or free amount stated on the pricing page.
xAINo free credits mentioned on the rate-limits or pricing pages; the lowest tier starts at $0 spent.
ChutesThe pricing FAQ says there is no free tier at this time.

Free models on OpenRouter today

OpenRouter lists some models at $0 for input and output, mostly as ":free" variants that hosts sponsor for a while. The list below is read from OpenRouter's public models API each day. Free variants share OpenRouter's free-usage limits, set out in the provider table above.

ModelMakerContextFree howListed on OpenRouterExpiry date
Ling 3.1 Flash
inclusionai/ling-3.1-flash
inclusionai262KFree at every routeNone listed
Apodex 1.1 Mini
apodex/apodex-1.1-mini:free
apodex262KFree variantNone listed
Space Bunny Alpha
stealth/space-bunny-alpha
stealth1MFree at every route5 October 2026
Ling 3.0 Flash Sante
inclusionai/ling-3.0-flash-sante:free
inclusionai262KFree variantNone listed
Qwen3.8 27B
qwen/qwen3.8-27b:free
qwen262KYes, a paid version is also listedNone listed
Dots3-Note Preview
dots-studio/dots-3-note-preview:free
dots-studio512KFree variant31 December 2026
LFM2.5-2.6B
liquid/lfm-2.5-2.6b:free
liquid66KFree variantNone listed
Nemotron 3.5 Lightning
nvidia/nemotron-3.5-lightning:free
nvidia1MYes, a paid version is also listedNone listed
Inkling Small
thinkingmachines/inkling-small:free
thinkingmachines1.05MYes, a paid version is also listedNone listed
Laguna S 2.1
poolside/laguna-s-2.1:free
poolside262KYes, a paid version is also listed31 October 2026
Inkling
thinkingmachines/inkling:free
thinkingmachines1.05MYes, a paid version is also listedNone listed
Laguna XS 2.1
poolside/laguna-xs-2.1:free
poolside262KYes, a paid version is also listed31 October 2026
North Mini Code
cohere/north-mini-code:free
cohere256KFree variantNone listed
Nemotron 3.5 Content Safety
nvidia/nemotron-3.5-content-safety:free
nvidia128KYes, a paid version is also listedNone listed
Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b:free
nvidia1MYes, a paid version is also listedNone listed
Nemotron 3 Nano Omni
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free
nvidia256KFree variantNone listed
Gemma 4 26B A4B
google/gemma-4-26b-a4b-it:free
google262KYes, a paid version is also listedNone listed
Gemma 4 31B
google/gemma-4-31b-it:free
google262KYes, a paid version is also listedNone listed
Nemotron 3 Super
nvidia/nemotron-3-super-120b-a12b:free
nvidia262KYes, a paid version is also listedNone listed

Every text-in, text-out model priced at $0 for input and output in the OpenRouter models API on 4 October 2026, newest first. OpenRouter's own routers and image or audio models are excluded. "Expiry date" is the date OpenRouter lists for the model's retirement, where it lists one.

Which one to start with

For prototyping against a strong general model with the least setup, the Gemini API free tier is the obvious first stop, as long as your prompts can be used to improve Google's products. Check that before you send anything sensitive.

For the most published requests per day on open models, Groq's free plan is the one to beat, and its limits are written down per model. SambaNova's free tier covers larger models but allows far fewer requests a day.

For trying many models through one key, OpenRouter's free variants are the flexible choice; buying 10 credits once lifts the daily cap from 50 to 1,000 requests. Expect individual free models to come and go, which the change log below records.

If you are building on Cloudflare already, Workers AI's daily allocation is the natural fit. Mistral's $10 monthly credit and Hugging Face's $0.10 suit small experiments rather than a running product.

Once a project outgrows free limits, the question becomes price. The LLM API pricing history tracks the cheapest paid price for a given level of capability every day, and the floor is a few cents per million tokens.

Change log

Free models added to and removed from OpenRouter, provider pages whose wording changed, and re-verifications, newest first.

  1. : tracker started. 19 free text models on OpenRouter; 12 provider free tiers and 6 providers without one verified on their own pages; GitHub Models recorded as retired (30 July 2026)

Methodology

What counts as a free tier. An allowance that lets a new account call an LLM text model through an API without paying, stated on the provider's own pricing, limits or docs pages. Chat apps and playgrounds without API access do not count. Each row records the limits as the provider writes them, the card requirement only where a page states it, and the caveats the provider attaches, such as using free-tier content for training or limiting use to evaluation.

How rows are verified. I read every provider row on the provider's own page on the verified date shown. Where a page renders its limits with JavaScript, the figures come from the page source. Scaleway's live docs blocked automated reads on 4 October 2026, so its row comes from Scaleway's own documentation repository on GitHub.

How rows stay current. Every day the tracker fetches each source page and checks that two or three short pieces of the verified wording, such as "1,000 API calls a month", are still present. If one disappears, the row is flagged as re-verifying and the change is logged; I then re-read the page and update the row with a new date. A page can change a limit without touching the checked wording, so the verified date remains the date to cite.

OpenRouter free models. Every model in the OpenRouter models API with a $0 prompt price and a $0 completion price whose input and output are text. OpenRouter's own routers (such as openrouter/free) and music or image models are excluded, which is why the count here can be a few lower than a raw count of $0 entries. Additions and removals compare each day's list with the previous day's.

Known gaps. Google and Mistral show free-tier rate limits only inside their consoles, so those cells say so instead of giving a number. NVIDIA does not publish its trial credit amount. Context windows are not stated on any provider's limits page; the OpenRouter table carries them for the models it lists. For free AI agent products, as opposed to free model APIs, see the best free AI agents and the AI agent platform free tier comparison.

FAQ

1. Which LLM APIs have a free tier in 2026?

On 4 October 2026, 6 providers run a standing free API tier: Google AI Studio (Gemini API), Groq, OpenRouter, Cloudflare Workers AI, Cohere, SambaNova Cloud. Mistral AI and Hugging Face Inference Providers give a small monthly credit instead, Alibaba Cloud Model Studio (international) and Scaleway Generative APIs give a one-time quota, and NVIDIA API Catalog and Cerebras offer trials. OpenRouter also lists 19 models that cost nothing to call.

2. Is the Gemini API free?

Yes, for most models. On 4 October 2026 Google's pricing page marked these as free of charge on the free tier: Gemini 3.8, 3.7, 3.6 and 3.5 Flash, 3.5 and 3.1 Flash-Lite, 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite, Gemma 4. Gemini 3.1 Pro is not free. Google no longer publishes the free tier's request or token limits in its docs; each project sees its own limits in AI Studio. Content sent on the free tier is used to improve Google's products.

3. How many free models does OpenRouter have?

19 text models were priced at $0 on 4 October 2026, 17 of them as ":free" variants. Free variants are limited to 20 requests a minute and 50 a day, or 1,000 a day once an account has bought 10 credits. Free models come and go, which is why this page logs every addition and removal by date.

4. Is the Groq API free?

Groq has a free plan. On 4 October 2026 its rate-limits page listed gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B: 30 requests a minute, 1,000 a day, 8,000 tokens a minute, 200,000 tokens a day each. Groq describes the table as a summary; each organisation's exact limits are on its console Limits page.

5. Which free LLM API has the highest limits?

By published daily requests, Groq: 1,000 requests a day per model on its free plan. OpenRouter allows 1,000 a day on free variants after a 10-credit purchase. Cloudflare Workers AI counts its free allowance in its own compute unit instead: 10,000 Neurons a day free. Google publishes no free-tier numbers to compare.

6. Can I use a free LLM API in production?

Usually not as your only path. Several providers attach conditions to free use: Google uses free-tier content to improve its products, Cohere trial keys are for evaluation, SambaNova preview models are not for production, and free OpenRouter variants can disappear without notice. Free tiers are best for prototypes, tests and low-volume tools, with a paid key ready as a fallback.

7. Is GitHub Models still free?

No. GitHub's docs say the service was fully retired on 30 July 2026, including the playground and inference API. It was a popular free option for developers, so older lists of free LLM APIs still include it.

8. How is this list kept current?

The OpenRouter model list is read from its public API every day and compared with the day before. The provider rows are read by hand on each provider's own page and dated; every day the tracker re-fetches those pages and flags any row whose verified wording has changed. The data is CC BY 4.0.

Sources

  1. OpenRouter, models API, read daily since 4 October 2026, and API rate limits, read 4 October 2026.
  2. Google, Gemini Developer API pricing and rate limits, read 4 October 2026.
  3. Groq, rate limits, read 4 October 2026.
  4. Cloudflare, Workers AI pricing and limits, read 4 October 2026.
  5. Mistral AI, pricing, read 4 October 2026.
  6. Cohere, rate limits, read 4 October 2026.
  7. SambaNova, rate limits, read 4 October 2026.
  8. Hugging Face, Inference Providers pricing, read 4 October 2026.
  9. Alibaba Cloud, Model Studio free quota, read 4 October 2026.
  10. Scaleway, Generative APIs FAQ (documentation source), read 4 October 2026.
  11. NVIDIA, API catalog FAQ, read 4 October 2026.
  12. Cerebras, rate limits, read 4 October 2026.
  13. GitHub, GitHub Models documentation (retirement notice), read 4 October 2026.

Related reading: the cheapest AI agent platforms, the AI agent pricing index, and how to create your own AI agent for free. Gravity's plans are on the pricing page.