Skip to main content
In this lesson you’ll learn how foundation models convert text into tokens, why tokens matter for limits, cost, and performance, and how to design prompts and workflows that respect token constraints. We’ll cover tokenization basics, model differences, practical examples, and mapping tokens to approximate text size so you can choose the right strategy for your application.

Problem statement

Large prompts can exceed a model’s context window. Every model declares a maximum context size (measured in tokens). If you ignore that limit, requests may fail or produce truncated or unexpected outputs. Token usage also affects cost, latency, and throughput. To manage these factors, developers must understand:
  • how input text is converted into tokens (tokenization),
  • how tokens are counted (input + output),
  • and how the total token count compares to a model’s limits.

What is a token?

A token is the atomic unit a model processes. Unlike human-oriented units (words, sentences), models operate on tokens. A token can be:
  • a whole word (e.g., “report”),
  • a subword or part of a word (e.g., “un”, “believ”, “able”),
  • or punctuation (e.g., ”.”).
Models generate text one token at a time. Both the input you send and the output the model produces are tokenized and counted.
An infographic titled "Solution: Learn How Tokens Work" with four numbered blue gradient cards. Each card explains tokens: they're units of text (words/parts/punctuation), models generate one token at a time, and models think in tokens rather than whole sentences.

Tokenization examples

Tokenization happens automatically when you submit text to a model, but it’s useful to reason about examples:
  • The sentence “Summarize a cybersecurity incident report.” might split into tokens such as “Summarize”, “a”, “cyber”, “security”, “incident”, “report”, ”.” — about seven tokens.
  • Some words split into subwords: “unbelievable” could tokenize to parts like “un”, “believ”, “able” depending on the tokenizer.
  • Special characters, emojis, or non-Latin scripts behave differently across tokenizers.

Why tokens matter

Token awareness affects multiple operational dimensions:
  • Pricing: many providers bill per token (often shown as cost per million tokens). Knowing token counts helps estimate and optimize costs.
  • Context window: models enforce a maximum number of tokens per request (input + expected output).
  • Performance: fewer tokens generally reduce latency and increase throughput in production systems.
Because both input and output tokens count toward the context window and billing, you should constrain both sides when practical (e.g., provide only necessary context; set max output tokens).
A presentation slide titled "Solution: Why Tokens Matter" showing four numbered panels that explain token concepts: pricing is based on tokens, context window measured in tokens, input+output tokens both count, and where tokens connect to money and limits. Each panel includes simple icons and brief captions.
Remember: a single request consumes the tokens in your prompt (input) plus the tokens the model generates (output). The combined total must fit within the model’s context window and determines billing.

A concrete example

Walk through a simple token accounting example.
  • Instruction: “Summarize a cybersecurity incident report.” → ~7 tokens
  • Incident report body: → ~250 tokens
  • Input total: 257 tokens
If the model returns: “The incident started at 10 PM and lasted 15 minutes.” → ~11 tokens Total tokens used = 257 (input) + 11 (output) = 268 tokens You must ensure that 268 tokens fit within the chosen model’s context limit and that the cost matches your budget. Example token calculation in code-like form:
When designing prompts, explicitly constrain desired output length or format (for example, “one sentence” or “two bullet points”) to reduce generated tokens and overall cost.

Tokenizers vary by model

Different models may use distinct tokenization algorithms and vocabularies. A document that tokenizes to 250 tokens on one model could be 282 tokens on another. That affects context usage and billing when you switch models.
A slide titled "Solution: Tokenization Varies by Model" showing three blue panels labeled "Tokenizers," "Counting," and "Test," each with an icon and brief notes about different tokenizers, token counting, and testing token usage when switching models. The design uses a dark background and gradient blue cards.
Always test tokenization with the exact model and tokenizer you plan to use. Token counts (and thus cost and feasibility) can change when you change models or languages.

Mapping tokens to approximate text size

Model docs sometimes report limits in bytes or kilobytes rather than tokens. Use these practical approximations when estimating document sizes: These are rough estimates that vary by language, character set, and tokenizer, but they help when selecting a model for expected document sizes.
A dark-blue presentation slide titled "Workflow: Tokens and Context Window" with a three-column table. The table compares units (1 token, 1,000 tokens, 10,000 tokens) to approximate sizes (~4 characters / ~750 words / ~7,500 words) and text sizes (~0.004 KB / ~4 KB / ~40 KB).

Practical outcomes of understanding tokenization

Designing with tokens in mind leads to concrete benefits:
  1. More efficient prompt design — include only the role, context, instruction, and output format required; every extra sentence costs tokens.
  2. Reduced risk of exceeding model limits — plan mitigation strategies like truncation, chunking, or retrieval-augmented generation.
  3. Better control over AI costs — minimizing unnecessary input and constraining output reduces billing.
  4. Improved performance at scale — lower token counts decrease latency and increase throughput.
A presentation slide titled "Results" showing four numbered dark-blue panels with icons listing benefits: 01 More efficient prompt design, 02 Reduced risk of exceeding model limits, 03 Better control over generative AI costs, and 04 Improved performance of AI applications. Each panel has a small colored top strip and a white icon above the text.

Key takeaway

Foundation models operate on tokens, not whole sentences. Both input and output tokens count toward the model’s context window and billing. Plan prompts, context, and output constraints with token usage in mind to control cost, avoid hitting limits, and improve application performance.
A presentation slide titled "Key Takeaway" that says foundation models process tokens, not full text, and that token limits affect model capability and cost.

Further topics

Next, explore strategies and tooling for working within context windows:
  • chunking large documents and summarizing each chunk,
  • retrieval-augmented generation (RAG) to provide only relevant context,
  • streaming outputs (when supported) to reduce peak memory needs,
  • and programmatic token counting using tokenizer libraries.
A presentation slide titled "What's Next? Understanding Context Window." On the right is a teal circular icon showing a stylized brain with circuit-like connections against a dark curved background.
Links and references

Watch Video