> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Embeddings and Vector Stores Part 2

> Explains chunking, embeddings, vector stores and RAG workflows to enable semantic retrieval and integrate retrieved text with foundation models for accurate, scalable search

When you have a long document, embedding the whole file as a single unit usually reduces search accuracy. Instead, split the content into smaller, focused pieces (chunks) and embed each chunk independently—this improves relevance and makes retrieval more efficient.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Embeddings-and-Vector-Stores-Part-2/workflow-embedding-whole-documents-chunking-context.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=1cb442b5f9b547069ca4d80512af6419" alt="An infographic titled &#x22;Workflow: Are Whole Documents Embedded?&#x22; Displayed are five colorful panels noting that documents can be long and multi-topic, searching entire documents reduces accuracy, models have context limits, and chunking content is preferred over embedding whole documents." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Embeddings-and-Vector-Stores-Part-2/workflow-embedding-whole-documents-chunking-context.jpg" />
</Frame>

Example: splitting a product manual

* Consider a 20‑page manual. One straightforward approach is to split it into chunks by page range (for instance, pages 1–6, 7–13, and 14–20) and embed each chunk separately. Each chunk becomes an independent unit stored in your vector store and returned by similarity search when relevant.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Embeddings-and-Vector-Stores-Part-2/workflow-chunking-product-manual-1-20.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=93462f6e373f33e655434da6d1329948" alt="A slide titled &#x22;Workflow: What Is Chunking?&#x22; showing a Product Manual (Pages 1–20) split into three colored chunks: Chunk 1 (Pages 1–6), Chunk 2 (Pages 7–13), and Chunk 3 (Pages 14–20)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Embeddings-and-Vector-Stores-Part-2/workflow-chunking-product-manual-1-20.jpg" />
</Frame>

Chunking strategy matters. Two common approaches:

| Strategy | When to use | Example |
| - | -: | - |
| Fixed-size splits | When format is regular (pages, tokens, or sentence counts) and you want predictable chunk sizes | Split every N tokens or every M pages |
| Semantic/topic splits | When you want to keep related content together (chapters, sections, or topic boundaries) | Split at headings, TOC entries, or using topic-detection algorithms |

Both approaches follow the same principle: embed each chunk separately to increase the likelihood that retrieval returns directly relevant content.

How are embedding vectors created?

* An embedding model converts text into a high-dimensional numerical vector that captures semantic meaning.
* Different embedding models vary in vector dimensionality, training data, and architecture—vectors from different models live in different vector spaces and are not directly comparable.

There are two distinct phases where embeddings are used:

1. Ingestion: convert your document chunks into vectors and store them in a vector database (vector store).
2. Query embedding: convert the user’s query into a vector so you can perform a similarity search against the stored vectors.

<Callout icon="lightbulb" color="#1CB2FE">
  Always use the same embedding model for both ingestion and query embedding. Mixing embedding models will produce vectors in different vector spaces and break similarity comparisons.
</Callout>

Embedding models vs. generation (foundation) models

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Embeddings-and-Vector-Stores-Part-2/workflow-separation-embedding-vs-generation.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=30df49deb8eed6c8d88a5cb7188c7730" alt="Slide titled &#x22;Workflow: Separation&#x22; showing two side-by-side boxes comparing Embedding Models (good at semantic positioning, used for search, not used for generation) and Generation Models (good at language production, only sees text not vectors, better reasoning). The left box is blue and the right box is orange on a dark background." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Embeddings-and-Vector-Stores-Part-2/workflow-separation-embedding-vs-generation.jpg" />
</Frame>

* Embedding models: map text into a vector space for semantic positioning and similarity search. They are not used to generate natural language responses.
* Generation (foundation) models: accept plain text (not vectors) as input and produce text outputs. They reason, synthesize, and produce responses, but they do not operate on vectors directly.

Typical Retrieval-Augmented Generation (RAG) workflow

1. Ingestion: chunk documents, embed each chunk using an embedding model, and store vectors in a vector store.
2. Query embedding: embed the incoming user query using the same embedding model.
3. Retrieval: perform a similarity search in the vector store; retrieve the top-k most relevant chunk(s) as plain text.
4. Augmented prompt assembly: combine the user question and retrieved text chunks into a single prompt (plain text).
5. Generation: send the augmented prompt to a foundation model (e.g., LLaMA, Anthropic Claude, etc.) to produce the final answer.

Important: vectors live in the retrieval layer and are invisible to the LLM. The LLM only receives the retrieved text chunks as context.

Why use a vector database in RAG?

| Benefit | Why it helps | Example |
| - | - | - |
| Better natural language search | Similarity search finds semantically related content even without exact keyword matches | A user asks “how do I reset the device?” and gets the correct procedure even if the manual uses “factory reset” |
| Faster access to proprietary knowledge | Retrieve relevant internal documents without manual filtering | Customer support agents get targeted sections from internal KBs instantly |
| Scalability | Vector stores optimize nearest-neighbor search for large datasets | Millions of vectors can be indexed and searched quickly |
| Incremental ingestion | Add new documents via batch jobs or event-driven pipelines as data arrives | New product manuals are embedded automatically when released |

Key takeaway: embeddings enable meaning-based (semantic) search, making it possible to retrieve the right information for RAG systems even when the user doesn’t use exact keywords.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Embeddings-and-Vector-Stores-Part-2/key-takeaway-embeddings-meaning-search-rag.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=de5fb564bc0a62f586cd7e171f68c11f" alt="A presentation slide titled &#x22;Key Takeaway&#x22; with a short point: &#x22;Embeddings enable meaning-based search, making it possible to retrieve the right information for RAG systems.&#x22; The slide has a dark blue left panel and a large white content area on the right." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Embeddings-and-Vector-Stores-Part-2/key-takeaway-embeddings-meaning-search-rag.jpg" />
</Frame>

Next steps

* The next section walks through an end-to-end RAG implementation: ingestion pipelines, vector store operations (indexing and search), and the exact sequence that runs when a user query hits your application.

Links and references

* [Retrieval-Augmented Generation overview](https://en.wikipedia.org/wiki/Retrieval-Augmented_Generation)
* [Vector database concepts and examples](https://www.arxiv.org/abs/2102.01931)
* [Embedding models and best practices](https://developers.google.com/machine-learning/guides/text-classification/embeddings)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/282d7660-0e67-4e7d-a498-291bc16784e5/lesson/14012742-80ba-4613-9b29-63228c6b0023" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.