Token
A model doesn't read text the way a person does, word by word. Text first gets broken into tokens — small chunks that are often a whole word, but just as often a piece of one, a punctuation mark, or a few characters that show up together a lot. The model reads and generates these tokens, not words directly, and everything measured about a request — how much fits, how much it costs — is counted in tokens, not words or characters.
How text actually becomes tokens
A common word like "the" or "cat" is usually its own single token, because it shows up constantly and earned a dedicated slot. A longer or rarer word often gets split into a few smaller pieces — "tokenization" itself might come apart into something like "token," "iz," and "ation." The general technique behind this, used across most current models, builds its set of tokens by starting from individual characters and repeatedly merging whichever pair shows up together most often in a huge amount of text, until common chunks — whole words, common word-parts like "ing" or "tion" — end up as single tokens and rarer text ends up split into more of them.
A useful approximation, not a rule
For ordinary English text, a rough rule of thumb is about 4 characters per token, or a little under one token per word. It's genuinely just a rough guide: code, URLs, and text in other languages often use noticeably more tokens for the same length, because the common-chunk shortcuts that make English words cheap don't apply the same way. Never assume a word and a token are interchangeable when estimating cost or how much fits in a request.
Why this actually matters
Two things are measured in tokens, not words: how much text fits in a model's context window for one request, and — for most paid APIs — how much a request costs, usually priced separately for the tokens sent in and the tokens generated back. A request that reads as "short" in plain English can still use meaningfully more tokens than expected if it's full of code, unusual names, or non-English text, because those tokenize less efficiently than everyday English prose.
Word count and token count can diverge
"The capital of France is Paris" tokenizes into a small handful of tokens, because every word in it is common. A sentence with a rare technical term, a long URL, or a name in a different alphabet can use several tokens for what looks like one "word" on the page — the visible word count and the actual token count can diverge by a lot, and only the token count is what the model's limits and the bill are based on.
In this guide
FAQ
Is one token always one word?
No. Common short words are often exactly one token, but longer or rarer words frequently split into two or more, and some tokens are just punctuation or a few characters that don't correspond to a word at all. Treat "word count" and "token count" as related but different numbers, not interchangeable ones.
Do all models split text into tokens the same way?
No — different models use different token vocabularies, so the same sentence can come out to a different number of tokens depending on which model's tokenizer is doing the counting. A count from one model's tokenizer isn't a reliable estimate for another model.