Skip to main content
When you ingest data into a Knowledge Base, the service first connects to your data source and then breaks documents into smaller pieces called chunks before creating embeddings. This prevents, for example, converting a 100‑page manual into one giant embedding vector — instead, the content is split so the most relevant sections can be retrieved for a user query. Chunking improves retrieval in Retrieval-Augmented Generation (RAG) systems by dividing documents into semantically meaningful or size-bounded segments. Each segment receives an embedding vector and is stored in a vector store. At query time, the user query is embedded and compared to the stored vectors (often using cosine similarity or another distance metric). The highest-scoring chunks are returned to augment the model prompt and produce grounded answers. Chunking strategy is a foundational design choice in any RAG pipeline. How you split content and whether you include overlap affects retrieval relevance, token usage, and cost.
A presentation slide titled "Workflow: Control Chunk Size in Bedrock Knowledge Bases" showing a central blue icon of a head with arrows connected to three colored circles labeled "Overlap between chunks," "Chunking strategy," and "Maximum chunk size."
Because Amazon Bedrock Knowledge Bases is a managed service, you influence chunking at ingestion but do not control the service’s internal chunking algorithm at query time. Key operational points:
  • Chunking configuration happens once at Knowledge Base creation via the data source configuration.
  • Your application does not supply chunking parameters at query time; you provide guidance up-front when creating the Knowledge Base.
  • The Bedrock Knowledge Base service performs the parsing, chunking, embedding, and vector storage during ingestion.
Chunking modes offered by Bedrock Knowledge Bases Why chunk size and overlap matter
  • Too large: Retrieved chunks may contain irrelevant content, increasing prompt size, token usage, and cost; answers can become less precise.
  • Too small: Important context can be split across chunks, requiring more retrievals and making it harder for the model to form coherent responses.
  • Overlap: Purposeful overlap between chunks preserves context across boundaries without making each chunk excessively large. This is a common tactic to improve recall at chunk boundaries.
Chunking in a managed Knowledge Base is a one-time configuration you should tune during ingestion. Monitor retrieval relevance and token usage, then iterate on chunk size and overlap until you reach the best balance for your dataset.
An infographic titled "Workflow: Why Chunk Size Matters" showing a balance scale with a large orange bag labeled "Too Large" and a smaller purple bag labeled "Too Small." Each side lists drawbacks—too large: imprecise retrieval, more irrelevant text and higher token/cost; too small: fragmented context, insufficient information and more chunks to retrieve.
Best practices and practical advice
  • Start with the default chunking and evaluate retrieval precision and token costs on a representative set of queries.
  • For long, structured documents (e.g., manuals, policy documents), prefer hierarchical or semantic chunking to preserve section context.
  • Use moderate overlap (e.g., 10–20%) to reduce boundary artifacts without much additional cost.
  • Monitor vector store size and average tokens per retrieved prompt to estimate real cost impacts.
  • If you need predictable token counts for safety/cost control, consider fixed-size chunking and adjust overlap accordingly.
What you gain from using Bedrock Knowledge Bases
A slide titled "Results" with five numbered panels listing benefits: faster implementation of RAG solutions, reduced operational complexity, scalable access to enterprise knowledge, automatic data ingestion and updates, and more accurate context-aware AI responses.
Summary Bedrock Knowledge Bases implements the core RAG building blocks for you: parsing, chunking, embedding, and vector management. By offloading these tasks to a managed service, you can implement RAG faster and operate it with less in-house infrastructure expertise. The main lever you control is the ingestion-time chunking configuration — choose the chunking mode, size, and overlap that best align with your content and retrieval goals, and iterate based on measured retrieval relevance and cost. Next steps and references

Watch Video

Practice Lab