> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Bedrock Knowledge Bases Introduction Part 3

> Explains chunking strategies, modes, and best practices for Amazon Bedrock Knowledge Bases to improve retrieval in RAG systems and manage embeddings, overlap, and cost.

When you ingest data into a Knowledge Base, the service first connects to your data source and then breaks documents into smaller pieces called chunks before creating embeddings. This prevents, for example, converting a 100‑page manual into one giant embedding vector — instead, the content is split so the most relevant sections can be retrieved for a user query.

Chunking improves retrieval in Retrieval-Augmented Generation (RAG) systems by dividing documents into semantically meaningful or size-bounded segments. Each segment receives an embedding vector and is stored in a vector store. At query time, the user query is embedded and compared to the stored vectors (often using cosine similarity or another distance metric). The highest-scoring chunks are returned to augment the model prompt and produce grounded answers.

Chunking strategy is a foundational design choice in any RAG pipeline. How you split content and whether you include overlap affects retrieval relevance, token usage, and cost.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-3/bedrock-control-chunk-size-workflow.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=575d80c534e2d0cb49263612c2da7334" alt="A presentation slide titled &#x22;Workflow: Control Chunk Size in Bedrock Knowledge Bases&#x22; showing a central blue icon of a head with arrows connected to three colored circles labeled &#x22;Overlap between chunks,&#x22; &#x22;Chunking strategy,&#x22; and &#x22;Maximum chunk size.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-3/bedrock-control-chunk-size-workflow.jpg" />
</Frame>

Because Amazon Bedrock Knowledge Bases is a managed service, you influence chunking at ingestion but do not control the service’s internal chunking algorithm at query time. Key operational points:

* Chunking configuration happens once at Knowledge Base creation via the data source configuration.
* Your application does not supply chunking parameters at query time; you provide guidance up-front when creating the Knowledge Base.
* The Bedrock Knowledge Base service performs the parsing, chunking, embedding, and vector storage during ingestion.

Chunking modes offered by Bedrock Knowledge Bases

| Mode | Description | When to use |
| - | - | - |
| Default chunking | Produces chunks of roughly a few hundred tokens (commonly \~300). Smaller documents are usually left unchanged. | General use when you want a balanced, plug-and-play setting. |
| Fixed-size chunking | You set the exact chunk size (token or character based). | When you need strict control over chunk/token counts for cost or model-context reasons. |
| Hierarchical chunking | Uses source structure (sections, subsections) to create parent/child relationships between chunks. | For well-structured documents (manuals, specs) where section context matters. |
| Semantic chunking | Groups sentences or paragraphs by semantic similarity rather than fixed size. | When preserving topical coherence is important across chunks. |
| No chunking | Leaves documents as-is (useful when sources are already pre-chunked or are short). | When upstream tooling handles segmentation or you want full-document retrieval. |

Why chunk size and overlap matter

* Too large: Retrieved chunks may contain irrelevant content, increasing prompt size, token usage, and cost; answers can become less precise.
* Too small: Important context can be split across chunks, requiring more retrievals and making it harder for the model to form coherent responses.
* Overlap: Purposeful overlap between chunks preserves context across boundaries without making each chunk excessively large. This is a common tactic to improve recall at chunk boundaries.

<Callout icon="lightbulb" color="#1CB2FE">
  Chunking in a managed Knowledge Base is a one-time configuration you should tune during ingestion. Monitor retrieval relevance and token usage, then iterate on chunk size and overlap until you reach the best balance for your dataset.
</Callout>

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-3/chunk-size-too-large-too-small.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=d17d04c14cc2c90a1008254358b16369" alt="An infographic titled &#x22;Workflow: Why Chunk Size Matters&#x22; showing a balance scale with a large orange bag labeled &#x22;Too Large&#x22; and a smaller purple bag labeled &#x22;Too Small.&#x22; Each side lists drawbacks—too large: imprecise retrieval, more irrelevant text and higher token/cost; too small: fragmented context, insufficient information and more chunks to retrieve." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-3/chunk-size-too-large-too-small.jpg" />
</Frame>

Best practices and practical advice

* Start with the default chunking and evaluate retrieval precision and token costs on a representative set of queries.
* For long, structured documents (e.g., manuals, policy documents), prefer hierarchical or semantic chunking to preserve section context.
* Use moderate overlap (e.g., 10–20%) to reduce boundary artifacts without much additional cost.
* Monitor vector store size and average tokens per retrieved prompt to estimate real cost impacts.
* If you need predictable token counts for safety/cost control, consider fixed-size chunking and adjust overlap accordingly.

What you gain from using Bedrock Knowledge Bases

| Benefit | Why it matters |
| - | - |
| Faster RAG implementation | Managed pipeline handles parsing, chunking, embedding, and vector storage so you can build RAG systems faster. |
| Reduced operational complexity | Fewer components to maintain and scale compared to a fully self-managed stack. |
| Scalable, scoped Knowledge Bases | Create multiple Knowledge Bases for teams, products, or domains, each with its own ingestion configuration. |
| Automated ingestion and updates | Keep your vector store current as new content is added or changed. |
| More accurate, context-aware AI responses | Grounding answers in similarity-search results over your proprietary data improves relevance. |

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-3/rag-results-faster-simpler-scalable-contextual.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=8de8307a6eab468d6f36741475483704" alt="A slide titled &#x22;Results&#x22; with five numbered panels listing benefits: faster implementation of RAG solutions, reduced operational complexity, scalable access to enterprise knowledge, automatic data ingestion and updates, and more accurate context-aware AI responses." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-3/rag-results-faster-simpler-scalable-contextual.jpg" />
</Frame>

Summary
Bedrock Knowledge Bases implements the core RAG building blocks for you: parsing, chunking, embedding, and vector management. By offloading these tasks to a managed service, you can implement RAG faster and operate it with less in-house infrastructure expertise. The main lever you control is the ingestion-time chunking configuration — choose the chunking mode, size, and overlap that best align with your content and retrieval goals, and iterate based on measured retrieval relevance and cost.

Next steps and references

* Try a hands-on lab to build a simple RAG-enabled application using Bedrock Knowledge Bases.
* Resources:
  * [Amazon Bedrock documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/) — product guides and API references.
  * [Retrieval-Augmented Generation (RAG) overview](https://www.semanticscholar.org/) — surveys and best practices for RAG architectures.
  * Articles and blogs on chunking, embeddings, and vector search for deeper tuning strategies.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/282d7660-0e67-4e7d-a498-291bc16784e5/lesson/5a9fa7ef-53ae-4b7d-8600-e22e78775cdf" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/282d7660-0e67-4e7d-a498-291bc16784e5/lesson/016dd624-a680-407f-9ff6-f001b2803ce6" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.