Header Ad Banner Area (728x90 / Responsive)

How Does Tokenization Work? How AI Reads Text

Before an AI model can understand your words, it has to chop them into pieces. That chopping — called tokenization — is one of those invisible mechanisms that affects everything: what AI costs you, how well it works in your language, and why it sometimes mangles rare words.

I’ve dug into tokenizers while testing models, and it’s one of those topics that seems boring until you realize it explains half the weird behaviors you’ve seen. Let me make it concrete.

Table of Contents

A laptop showing a token counter increasing as someone types

The Short Answer

Tokenization is the process of splitting text into small chunks called tokens before an AI model processes it. A token is roughly ¾ of an English word — common words are usually one token, rare words get split into pieces. Everything an AI reads and writes is measured, priced, and limited in tokens, not words.

Why AI Can’t Just Read Words

Here’s the thing most people don’t realize: AI models don’t process text. They process numbers. Every piece of text has to be converted into a sequence of numbers (token IDs) before the model can do anything with it.

So why not just use whole words? Two problems:

Vocabulary size. English has hundreds of thousands of words, plus names, typos, and new slang appearing daily. A model with a separate number for every possible word would need an enormous vocabulary — most of which it would rarely use.

Unknown words. The moment someone types a word the model has never seen — a new product name, a typo, a username — a word-based system breaks. It has no number for that word.

Tokens solve both problems. Instead of whole words, the model uses a fixed vocabulary of text chunks — some are whole common words, some are word pieces, some are individual characters. Any text, in any language, can be represented as a sequence of these chunks. Nothing is ever truly “unknown” because worst case, the text gets spelled out character by character.

This conversion happens at the very start of the pipeline I describe in how AI models are trained — tokenization is step zero, before any learning happens.

How Tokenization Actually Works

The most common method is called Byte Pair Encoding (BPE). Here’s the intuition:

  1. Start with individual characters.
  2. Find the most common pair of adjacent pieces in a huge sample of text.
  3. Merge that pair into a new piece. Add it to the vocabulary.
  4. Repeat tens of thousands of times.

After enough merges, the vocabulary contains the most useful chunks: common words (” the”, ” and”), common word parts (“ing”, “pre”, “tion”), and individual characters as fallback.

To tokenize new text, the system applies these merges greedily — always using the longest available chunk. So “unhappiness” might become [“un”, “happiness”] or [“un”, “happi”, “ness”] depending on which chunks made the vocabulary cut.

Modern tokenizers have vocabularies of roughly 30,000 to 200,000 tokens. The exact vocabulary is part of what defines a model — different models tokenize the same text differently.

Letterpress type blocks as an analogy for AI tokens

Real Examples: Watch Text Get Chopped

Let me show you what this looks like in practice (using a typical modern tokenizer):

“Hello, how are you?” → ["Hello", ",", " how", " are", " you", "?"] — 6 tokens. Clean and efficient.

“Tokenization” → ["Token", "ization"] — 2 tokens. The word gets split at a natural seam.

“supercalifragilisticexpialidocious” → ["super", "cal", "ifrag", "ilistic", "exp", "ial", "idoc", "ious"] — 8 tokens. Rare long words shatter into many pieces.

“Hello 👋” → ["Hello", " 👋"] or several tokens for the emoji — emoji are token-expensive.

“antidisestablishmentarianism” → splits into 5-7 pieces depending on the tokenizer.

Notice the pattern: the more common the text, the fewer tokens it needs. This has real consequences, which brings us to…

Why Tokenization Matters to You

Cost. AI APIs charge per token. A prompt in English costs less than the same content in a language that tokenizes poorly. If you’re building on AI APIs, tokenization directly hits your bill. I’ve seen the same task cost 3x more just because of language choice.

Speed. Models generate token by token. More tokens = more generation steps = slower responses. That laggy feeling with long outputs is partly token count.

Context limits. Remember the context window? It’s measured in tokens. Text that tokenizes inefficiently eats your window faster. A “128K context” holds fewer words of Thai than English.

Quality. Here’s the subtle one: the model “thinks” in tokens, not words or letters. This is why AI is weirdly bad at things like counting letters in a word (“how many r’s in strawberry?”) — it never sees the letters, only the token chunks. It’s also why models sometimes struggle with rhyming, wordplay, and precise character-level tasks.

Weird artifacts. Ever seen an AI repeat a strange phrase or get stuck on an unusual word? Sometimes that’s a tokenization edge case — a rare chunk the model barely saw during training behaving unpredictably.

The Language Inequality Problem

This is the part that bothers me most. Tokenization is dramatically less efficient for many non-English languages:

  • English: ~1.3 tokens per word
  • French, Spanish, German: ~1.5-2 tokens per word
  • Arabic, Hindi: ~2-4 tokens per word
  • Thai, Burmese, some African languages: ~5-10+ tokens per word

Same meaning, wildly different token counts. That means speakers of these languages pay more, wait longer, fit less in the context window, and often get worse quality — for the exact same task.

This isn’t anyone’s evil plan; it’s a consequence of tokenizers being trained mostly on English-heavy data. But it’s a real equity issue in AI, and it’s one reason multilingual model work matters. When you see a model advertised as “multilingual,” one of the first things I check is whether its tokenizer was actually designed for those languages or just tolerates them.

How Tokenizers Get Built

Building a tokenizer is part art, part statistics:

  1. Collect training text — ideally representative of what the model will see.
  2. Run BPE merges — typically 30K-200K merges, as described above.
  3. Add special tokens — markers for start/end of text, padding, and control signals.
  4. Test on diverse text — code, multiple languages, emoji, math — checking that nothing explodes into absurd token counts.

Some newer approaches go further: byte-level tokenizers that work directly on raw bytes (no “unknown” anything, ever), or learned tokenizers optimized for specific domains like code or science.

The tokenizer is frozen before model training begins — it’s part of the model’s DNA. You can’t swap tokenizers after training any more than you can change the foundation after building the house.

Two speech bubbles showing how languages tokenize differently

Frequently Asked Questions

How does tokenization work in simple terms?

Tokenization splits text into small chunks called tokens before an AI model processes it. Using a method called Byte Pair Encoding, common words become single tokens while rare words split into pieces. A token is roughly three-quarters of an English word. Everything AI reads and writes is measured and priced in tokens.

How many tokens are in a word?

On average, one English word is about 1.3 tokens — so a token is roughly three-quarters of a word. Common words are usually one token each; rare or long words split into multiple tokens. Other languages vary widely: some need 5-10 tokens per word, making AI more expensive and slower in those languages.

Why does tokenization matter for AI cost?

AI APIs charge per token, so inefficient tokenization directly increases your bill. The same task can cost 2-3x more in a language that tokenizes poorly versus English. Token count also affects speed (models generate token by token) and how much fits in the context window.

Why is AI bad at counting letters in words?

Because the model never sees individual letters — it sees token chunks. The word “strawberry” might be two tokens, so the model has no direct access to its letters. It’s like asking someone to count the bricks in a wall they’ve only seen in a photograph. This token-level view also explains struggles with rhyming and precise wordplay.

Do all AI models use the same tokenizer?

No. Each model family has its own tokenizer with its own vocabulary (typically 30,000 to 200,000 tokens). The same text tokenizes differently across models, which means token counts, costs, and context usage aren’t directly comparable between providers.

What is the language inequality problem in tokenization?

Tokenizers trained mostly on English split other languages into far more tokens — sometimes 5-10x more per word. This means non-English speakers pay more, wait longer, fit less in the context window, and often get lower quality for the same task. It’s a significant equity issue in AI deployment.

Join the Discussion

Your email address will not be published. Required fields are marked *