What a token is
A token is the unit a language model reads and writes. Before a model sees your prompt, a tokenizer cuts the text into pieces from a fixed vocabulary: common words become one token, rarer words split into parts, and punctuation, digits and spaces are folded into tokens of their own. The model never sees letters, only token numbers, and API providers bill for every one of them.
Here is how OpenAI's o200k tokenizer, the one this token counter runs, splits a few strings. A dot (·) marks a space.
- “Hello, world.” is 4 tokens: Hello,·world. The space travels with the word after it.
- “Tokenization is unbelievably useful.” is 6 tokens: Tokenization·is·unbelievably·useful. A long word can be one token if it is common enough.
- The number 1234567 is 3 tokens, 1234567, and the date 2026-10-09 is 6.
- The Hindi greeting नमस्ते is one word but 4 tokens: नमस्ते
That is why the rule of thumb of four characters per token only holds for plain English. Code, numbers, names and other languages use more tokens per word, so it pays to count the real text you plan to send.
How to count tokens for GPT, Claude and Gemini
Each model family has its own tokenizer, so the same text gives a slightly different count depending on where you send it.
GPT and other OpenAI models
OpenAI publishes its tokenizers. GPT-4o, GPT-4.1, the o-series reasoning models and GPT-5 use o200k_base, the tokenizer this page runs in your browser, so for those models it works as an exact OpenAI token counter for the text itself. If you use a newer model, check its documentation for the tokenizer it uses. Chat requests also add a few tokens per message for the role markers, and images and tool definitions are counted separately.
Claude
Anthropic does not publish Claude's tokenizer as a library. Its API has a token counting endpoint that returns the input count for a request, and every response reports the input and output tokens it billed. Treat the number on this page as an estimate for Claude: it can differ by roughly 10 to 20%. To use the table as a Claude API pricing calculator, pick Anthropic in the provider filter.
Gemini, Llama and others
Google's Gemini API has a countTokens method, and open models such as Llama, Qwen and Mistral ship their tokenizer with the model weights. Without those tools, the count here is a fair estimate. Filter the table by Google to use it as a Gemini API pricing calculator.
How API cost is calculated
Providers bill input and output tokens separately, each at a price per million tokens. For one call:
cost per call = input tokens ÷ 1,000,000 × input price + output tokens ÷ 1,000,000 × output price
cost per month = cost per call × calls per month
A worked example: a prompt of 1,500 input tokens that gets a 500-token reply, on a model priced at $2 per million input tokens and $10 per million output tokens. Input costs 1,500 ÷ 1,000,000 × $2 = $0.003. Output costs 500 ÷ 1,000,000 × $10 = $0.005. That is $0.008 a call, and at 3,000 calls a month, $24.
Two things push the real bill up. Output usually costs more than input: in our 9 October 2026 price list the output price was at least double the input price on 54 of 55 models. And reasoning models bill their hidden thinking as output tokens, so set the expected output higher for them. The calculator above runs this formula for every model at once, which makes it an LLM pricing comparison as well as an OpenAI pricing calculator. For a full agent budget, with retries and tool calls, see how much it costs to run an AI agent 24/7.
Ways to cut token spend
- Shorten the system prompt. It is sent with every call. A 2,000-token system prompt at 3,000 calls a month is 6 million input tokens before anyone has typed a word. Cut repeated rules and examples the model does not need, then paste the new version above to check the saving.
- Use prompt caching. When the start of the prompt is the same every time, providers can bill those tokens at a cached rate. In our 9 October 2026 price list, 48 of the 50 models that publish a cache price charged a quarter of the normal input price or less for a cached read. Put the fixed instructions first and the changing content last so the cache can match.
- Send simple steps to smaller models. Sorting, tagging, pulling fields out of an email and routing rarely need the most capable model. Give those steps to a cheaper one and keep the expensive model for the hard reasoning. Filter the table above to compare.
- Limit output length. Set a maximum number of output tokens, ask for a fixed format such as three bullets or a few JSON fields, and tell the model to skip the preamble. Output is the expensive side, so this often saves more than trimming the input.
- Send only what the model needs. Trim old chat turns, and pass the relevant passages of a document instead of the whole file.
For prototypes, several providers have free API tiers; our free LLM API tiers tracker lists the current limits. For a longer playbook, read AI agent cost optimization.
Limits of this tool
- The count is exact only for OpenAI models that use the o200k tokenizer. For Claude, Gemini, Llama and others it is an estimate that can differ by roughly 10 to 20%.
- It counts the text you paste. Chat formatting, tool definitions, images, audio and files add tokens that are not counted here.
- Output tokens are your own estimate. Reasoning models can spend many more output tokens than the visible reply.
- Prices are list prices per million tokens. They leave out cached-input discounts, batch discounts, long-context surcharges, free tiers and taxes, so your invoice can come out lower or higher.
- The table covers the five newest paid text models from each lab our LLM API price tracker follows, not every model on the market.
Questions
What is a token?
A token is the unit a language model reads, writes and bills by. It can be a whole short word, part of a longer word, a number or a punctuation mark. For OpenAI's o200k tokenizer, “Hello, world.” is four tokens: “Hello”, “,”, “ world” and “.”. API providers charge per token, with separate prices for input and output.
How many words is 1,000 tokens?
Roughly 750 English words with OpenAI's tokenizers, or about four characters per token. The ratio moves with the text: code, numbers, names and languages other than English use more tokens per word. Paste your own text into the counter to see its tokens per word.
Does Claude count tokens differently?
Yes. Anthropic uses its own tokenizer for Claude, so the same text can come out as a different number of tokens than OpenAI's count, by roughly 10 to 20%. For an exact figure, use Anthropic's token counting endpoint or the token usage returned with each API response. Gemini, Llama and other model families also have their own tokenizers.
Why do output tokens cost more than input tokens?
The model reads all the input tokens in one parallel pass, but it writes the output one token at a time, with a full pass through the model for each. That makes output slower and more expensive to serve. In our 9 October 2026 price list, output cost at least twice the input price on 54 of 55 models, and four times on the median model.
Is my text sent anywhere?
No. The tokenizer runs inside your browser, so your text never leaves this page. The page downloads the tokenizer and the daily price list, and analytics record only that the tool was used, never what you typed.
How current are the prices?
The table reads our LLM API price tracker, which records list prices every day for the five newest paid text models from each lab it follows, taken from OpenRouter and the labs' published prices. The date above the table shows the last update. Prices are per million tokens, before caching, batch or volume discounts.