> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Tokens and Tokenization

> Explains tokens and tokenization for LLMs, their impact on cost, speed, context windows, tokenizer differences, and practical guidance for estimating and optimizing token usage.

Large language models (LLMs) do not read text the way humans do. You see words; an LLM sees tokens.

A token is a chunk of text — sometimes a whole word, sometimes part of a word, and sometimes a single character. Before an LLM processes anything you type, it first breaks your text into these tokens. This step is called tokenization.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/token-chunk-text-slide-kodekloud-presenter.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=872a634b873c7a30e2a174244ecec0af" alt="A presentation slide titled &#x22;A token is a chunk of text&#x22; listing that tokens can be whole words, parts of words, or single characters. A presenter wearing a KodeKloud t-shirt stands to the right explaining the concept." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/token-chunk-text-slide-kodekloud-presenter.jpg" />
</Frame>

Why tokenization matters

* Tokens determine cost: providers bill per token for both input and output.
* Tokens determine speed: models generate text token-by-token.
* Tokens determine what the model can consider at once: the context window limit is measured in tokens.

Concrete example

* Sentence: `The tokenization process is fascinating.`
* One possible tokenization: `The`, `token`, `ization`, `process`, `is`, `fascinating`.

Most common words remain single tokens because they occur frequently in training data. Less common or technical words are split into sub-word pieces. Tokenizers learn these reusable sub-word units so the model can handle words it has never seen before (for example, if it knows `token` and `ization`, it can generalize to `visualization`, `globalization`, etc.). Rare or highly specialized terms usually break into more tokens and therefore cost more to process. Code and punctuation often increase token counts because tokenizers split punctuation and keywords separately; non-English text can also use more tokens per word depending on tokenizer training data.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/tokenization-patterns-slide-token-counts.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=5f7101efa939198ed5b9e8f31d0a8576" alt="A presenter stands on the right next to a slide titled &#x22;TOKENIZATION PATTERNS&#x22; that shows a table of example texts (like &#x22;the&#x22;, &#x22;computer&#x22;, &#x22;tokenization&#x22;) with corresponding token counts and brief explanations. The slide uses colored text and a dark grid background." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/tokenization-patterns-slide-token-counts.jpg" />
</Frame>

Quick rules of thumb

| Concept | Approximate value | Notes |
| - | -: | - |
| Characters per token (English) | ≈ 4 characters | Rough average; depends on language and punctuation. |
| Tokens per word (English) | ≈ 0.75 words | A 1,000-word essay ≈ 1,300–1,500 tokens (approximate). |
| Tokenization behavior | Frequency-based | Common words often map to single tokens; rare words split into more tokens. |

Why tokens directly affect your applications

Cost

* Providers usually charge per token for both input (prompt/context) and output (model-generated tokens).
* Output tokens are often costlier because generation is autoregressive: the model processes input once but generates output token-by-token.

Example pricing (illustrative — check current rates with your provider):

* Input tokens: \~\$2.50 per 1,000,000 tokens
* Output tokens: \~\$10.00 per 1,000,000 tokens

<Callout icon="warning" color="#FF6B6B">
  Token costs can accumulate quickly, especially for agents that make many LLM calls or generate long outputs. Always check the latest pricing and estimate token usage before deploying at scale.
</Callout>

Speed

* LLMs predict one token at a time. Responses with many tokens take longer to produce.
* Shorter outputs and more compact prompts improve response latency.

Context window (input + output)

* Every model has a maximum context window measured in tokens. This is a hard limit that includes both your prompt and the model’s generated output.
* If your prompt plus the desired output exceed the context window, the model cannot process the full request.

To illustrate how context windows have expanded over time:

* GPT-3.5 (2022): \~4,000 tokens
* GPT-4o (2024): variants up to \~128,000 tokens
* Research and vendor announcements describe models supporting contexts on the order of hundreds of thousands to millions of tokens — exact limits vary by model and provider. Always check model specs.

If your combined input and output exceed the model’s context window, you must truncate, summarize, or otherwise reduce the content before sending it to the model. Later sections cover strategies for working within and around context limits.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/context-window-token-sizes-slide.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=1e8a293bf86d72e65f69d61fdb996612" alt="A presenter stands to the right of a slide titled &#x22;CONTEXT WINDOW&#x22; showing a bar chart comparing model years and maximum token sizes (GPT-3.5 4K, GPT-4o 128K, Claude/GPT-5.5 1M, Llama 4 Scout 10M). The speaker is wearing a KodeKloud t-shirt and a caption at the bottom reads &#x22;INPUT + OUTPUT MUST FIT.&#x22;" width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/context-window-token-sizes-slide.jpg" />
</Frame>

Different tokenizers, different counts

* Tokenization is not standardized across all models. Each company designs tokenizers based on its own vocabulary and training data.
* The same sentence may be 15 tokens with one tokenizer and 18 with another. Use model-specific tokenizers/tools to get accurate counts.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/tokenizer-compare-gpt4-llama3-token-counts.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=65ad556f6c693c4d3a811f3b33139504" alt="An infographic comparing how different tokenizers (GPT-4 vs LLaMA 3) split the sentence &#x22;tokenization is interesting&#x22; into tokens with counts. A person wearing a KodeKloud t-shirt stands on the right side of the image gesturing." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/tokenizer-compare-gpt4-llama3-token-counts.jpg" />
</Frame>

Try it yourself

* OpenAI provides a Tokenizer tool where you can paste text and see the exact token splits for their tokenizer.
* Experiment with natural language, code, technical jargon, and non-English text to estimate token costs for your prompts and applications.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/tokenizer-web-tool-screenshot-kodekloud-man.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=40ddf144b73f04b74046845a7547e05e" alt="A screenshot of a &#x22;Tokenizer&#x22; web tool showing text input, token counts, and the header &#x22;TRY IT YOURSELF&#x22; on a dark grid background. A man in a white KodeKloud t-shirt stands on the right side, facing the camera." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Tokens-and-Tokenization/tokenizer-web-tool-screenshot-kodekloud-man.jpg" />
</Frame>

Resources and next steps

* Tokenizer tool: [https://platform.openai.com/tokenizer](https://platform.openai.com/tokenizer)
* Check your target model’s documentation for exact context window and pricing details.
* When building agents, measure and optimize token usage for cost, latency, and context constraints.

<Callout icon="lightbulb" color="#1CB2FE">
  Try several different text samples — natural English, code snippets, technical jargon, and non-English text — to see how tokenization patterns change and to estimate the token cost of your prompts.
</Callout>

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/13d4f7ad-29e5-4bc0-b026-47c4ae43c31c/lesson/a6bc1f0f-b710-40c6-baf6-da5ec50bc9e0" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.