
- Chunking configuration happens once at Knowledge Base creation via the data source configuration.
- Your application does not supply chunking parameters at query time; you provide guidance up-front when creating the Knowledge Base.
- The Bedrock Knowledge Base service performs the parsing, chunking, embedding, and vector storage during ingestion.
Why chunk size and overlap matter
- Too large: Retrieved chunks may contain irrelevant content, increasing prompt size, token usage, and cost; answers can become less precise.
- Too small: Important context can be split across chunks, requiring more retrievals and making it harder for the model to form coherent responses.
- Overlap: Purposeful overlap between chunks preserves context across boundaries without making each chunk excessively large. This is a common tactic to improve recall at chunk boundaries.
Chunking in a managed Knowledge Base is a one-time configuration you should tune during ingestion. Monitor retrieval relevance and token usage, then iterate on chunk size and overlap until you reach the best balance for your dataset.

- Start with the default chunking and evaluate retrieval precision and token costs on a representative set of queries.
- For long, structured documents (e.g., manuals, policy documents), prefer hierarchical or semantic chunking to preserve section context.
- Use moderate overlap (e.g., 10–20%) to reduce boundary artifacts without much additional cost.
- Monitor vector store size and average tokens per retrieved prompt to estimate real cost impacts.
- If you need predictable token counts for safety/cost control, consider fixed-size chunking and adjust overlap accordingly.

- Try a hands-on lab to build a simple RAG-enabled application using Bedrock Knowledge Bases.
- Resources:
- Amazon Bedrock documentation — product guides and API references.
- Retrieval-Augmented Generation (RAG) overview — surveys and best practices for RAG architectures.
- Articles and blogs on chunking, embeddings, and vector search for deeper tuning strategies.