> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Understanding Tokens

> Describes tokenization, token impact on context limits cost and performance, and practical prompt and workflow strategies to manage token usage

In this lesson you'll learn how foundation models convert text into tokens, why tokens matter for limits, cost, and performance, and how to design prompts and workflows that respect token constraints. We'll cover tokenization basics, model differences, practical examples, and mapping tokens to approximate text size so you can choose the right strategy for your application.

## Problem statement

Large prompts can exceed a model's context window. Every model declares a maximum context size (measured in tokens). If you ignore that limit, requests may fail or produce truncated or unexpected outputs. Token usage also affects cost, latency, and throughput. To manage these factors, developers must understand:

* how input text is converted into tokens (tokenization),
* how tokens are counted (input + output),
* and how the total token count compares to a model's limits.

## What is a token?

A token is the atomic unit a model processes. Unlike human-oriented units (words, sentences), models operate on tokens. A token can be:

* a whole word (e.g., "report"),
* a subword or part of a word (e.g., "un", "believ", "able"),
* or punctuation (e.g., ".").

Models generate text one token at a time. Both the input you send and the output the model produces are tokenized and counted.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/learn-how-tokens-work-infographic.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=af08476e1638670c45582d3120ebeb37" alt="An infographic titled &#x22;Solution: Learn How Tokens Work&#x22; with four numbered blue gradient cards. Each card explains tokens: they're units of text (words/parts/punctuation), models generate one token at a time, and models think in tokens rather than whole sentences." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/learn-how-tokens-work-infographic.jpg" />
</Frame>

## Tokenization examples

Tokenization happens automatically when you submit text to a model, but it's useful to reason about examples:

* The sentence "Summarize a cybersecurity incident report." might split into tokens such as "Summarize", "a", "cyber", "security", "incident", "report", "." — about seven tokens.
* Some words split into subwords: "unbelievable" could tokenize to parts like "un", "believ", "able" depending on the tokenizer.
* Special characters, emojis, or non-Latin scripts behave differently across tokenizers.

## Why tokens matter

Token awareness affects multiple operational dimensions:

* Pricing: many providers bill per token (often shown as cost per million tokens). Knowing token counts helps estimate and optimize costs.
* Context window: models enforce a maximum number of tokens per request (input + expected output).
* Performance: fewer tokens generally reduce latency and increase throughput in production systems.

Because both input and output tokens count toward the context window and billing, you should constrain both sides when practical (e.g., provide only necessary context; set max output tokens).

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/solution-why-tokens-matter-slide.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=53df7020a201d452c7d737be0d6a6a39" alt="A presentation slide titled &#x22;Solution: Why Tokens Matter&#x22; showing four numbered panels that explain token concepts: pricing is based on tokens, context window measured in tokens, input+output tokens both count, and where tokens connect to money and limits. Each panel includes simple icons and brief captions." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/solution-why-tokens-matter-slide.jpg" />
</Frame>

<Callout icon="lightbulb" color="#1CB2FE">
  Remember: a single request consumes the tokens in your prompt (input) plus the tokens the model generates (output). The combined total must fit within the model's context window and determines billing.
</Callout>

## A concrete example

Walk through a simple token accounting example.

* Instruction: "Summarize a cybersecurity incident report." → \~7 tokens
* Incident report body: → \~250 tokens
* Input total: 257 tokens

If the model returns: "The incident started at 10 PM and lasted 15 minutes." → \~11 tokens

Total tokens used = 257 (input) + 11 (output) = 268 tokens

You must ensure that 268 tokens fit within the chosen model's context limit and that the cost matches your budget.

Example token calculation in code-like form:

```Python theme={null}
instruction_tokens = 7
document_tokens = 250
expected_output_tokens = 11

total_tokens = instruction_tokens + document_tokens + expected_output_tokens
# total_tokens = 268
```

When designing prompts, explicitly constrain desired output length or format (for example, "one sentence" or "two bullet points") to reduce generated tokens and overall cost.

## Tokenizers vary by model

Different models may use distinct tokenization algorithms and vocabularies. A document that tokenizes to 250 tokens on one model could be 282 tokens on another. That affects context usage and billing when you switch models.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/tokenization-varies-by-model.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=9217b06f5803f9e2865bba67f01982c0" alt="A slide titled &#x22;Solution: Tokenization Varies by Model&#x22; showing three blue panels labeled &#x22;Tokenizers,&#x22; &#x22;Counting,&#x22; and &#x22;Test,&#x22; each with an icon and brief notes about different tokenizers, token counting, and testing token usage when switching models. The design uses a dark background and gradient blue cards." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/tokenization-varies-by-model.jpg" />
</Frame>

<Callout icon="warning" color="#FF6B6B">
  Always test tokenization with the exact model and tokenizer you plan to use. Token counts (and thus cost and feasibility) can change when you change models or languages.
</Callout>

## Mapping tokens to approximate text size

Model docs sometimes report limits in bytes or kilobytes rather than tokens. Use these practical approximations when estimating document sizes:

| Unit | Approximate equivalent |
| - | - |
| 1 token | ≈ 4 characters on average |
| 1,000 tokens | ≈ 750 words ≈ \~4 KB |
| 10,000 tokens | ≈ 7,500 words ≈ \~40 KB |

These are rough estimates that vary by language, character set, and tokenizer, but they help when selecting a model for expected document sizes.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/tokens-context-window-table.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=3dc608644b476eb598b4fa98e90b3edb" alt="A dark-blue presentation slide titled &#x22;Workflow: Tokens and Context Window&#x22; with a three-column table. The table compares units (1 token, 1,000 tokens, 10,000 tokens) to approximate sizes (~4 characters / ~750 words / ~7,500 words) and text sizes (~0.004 KB / ~4 KB / ~40 KB)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/tokens-context-window-table.jpg" />
</Frame>

## Practical outcomes of understanding tokenization

Designing with tokens in mind leads to concrete benefits:

1. More efficient prompt design — include only the role, context, instruction, and output format required; every extra sentence costs tokens.
2. Reduced risk of exceeding model limits — plan mitigation strategies like truncation, chunking, or retrieval-augmented generation.
3. Better control over AI costs — minimizing unnecessary input and constraining output reduces billing.
4. Improved performance at scale — lower token counts decrease latency and increase throughput.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/results-ai-benefits-panels.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=2d67924146abc4a1031c6385155d1fd3" alt="A presentation slide titled &#x22;Results&#x22; showing four numbered dark-blue panels with icons listing benefits: 01 More efficient prompt design, 02 Reduced risk of exceeding model limits, 03 Better control over generative AI costs, and 04 Improved performance of AI applications. Each panel has a small colored top strip and a white icon above the text." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/results-ai-benefits-panels.jpg" />
</Frame>

## Key takeaway

Foundation models operate on tokens, not whole sentences. Both input and output tokens count toward the model's context window and billing. Plan prompts, context, and output constraints with token usage in mind to control cost, avoid hitting limits, and improve application performance.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/foundation-models-process-tokens-limits-cost.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=002d7941c68241a0d02be2356e4ee2eb" alt="A presentation slide titled &#x22;Key Takeaway&#x22; that says foundation models process tokens, not full text, and that token limits affect model capability and cost." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/foundation-models-process-tokens-limits-cost.jpg" />
</Frame>

## Further topics

Next, explore strategies and tooling for working within context windows:

* chunking large documents and summarizing each chunk,
* retrieval-augmented generation (RAG) to provide only relevant context,
* streaming outputs (when supported) to reduce peak memory needs,
* and programmatic token counting using tokenizer libraries.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/whats-next-context-window-brain-icon.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=f2ca1bbf6251896f9df775cf63697c83" alt="A presentation slide titled &#x22;What's Next? Understanding Context Window.&#x22; On the right is a teal circular icon showing a stylized brain with circuit-like connections against a dark curved background." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Understanding-Tokens/whats-next-context-window-brain-icon.jpg" />
</Frame>

Links and references

* Tokenization and tokenizers: [https://huggingface.co/docs/tokenizers/](https://huggingface.co/docs/tokenizers/)
* Practical token counting tool (tiktoken for OpenAI-style tokenization): [https://github.com/openai/tiktoken](https://github.com/openai/tiktoken)
* Overview of context windows and prompt design: [https://platform.openai.com/docs/guides/usage-and-billing](https://platform.openai.com/docs/guides/usage-and-billing) (model-specific docs vary by provider)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/20770e4e-fc7b-49d2-b443-aaefd5e7d3e9/lesson/2568afc7-5842-4a68-ae45-b610c2bfdafb" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.