Skip to main content
When you have a long document, embedding the whole file as a single unit usually reduces search accuracy. Instead, split the content into smaller, focused pieces (chunks) and embed each chunk independently—this improves relevance and makes retrieval more efficient.
An infographic titled "Workflow: Are Whole Documents Embedded?" Displayed are five colorful panels noting that documents can be long and multi-topic, searching entire documents reduces accuracy, models have context limits, and chunking content is preferred over embedding whole documents.
Example: splitting a product manual
  • Consider a 20‑page manual. One straightforward approach is to split it into chunks by page range (for instance, pages 1–6, 7–13, and 14–20) and embed each chunk separately. Each chunk becomes an independent unit stored in your vector store and returned by similarity search when relevant.
A slide titled "Workflow: What Is Chunking?" showing a Product Manual (Pages 1–20) split into three colored chunks: Chunk 1 (Pages 1–6), Chunk 2 (Pages 7–13), and Chunk 3 (Pages 14–20).
Chunking strategy matters. Two common approaches: Both approaches follow the same principle: embed each chunk separately to increase the likelihood that retrieval returns directly relevant content. How are embedding vectors created?
  • An embedding model converts text into a high-dimensional numerical vector that captures semantic meaning.
  • Different embedding models vary in vector dimensionality, training data, and architecture—vectors from different models live in different vector spaces and are not directly comparable.
There are two distinct phases where embeddings are used:
  1. Ingestion: convert your document chunks into vectors and store them in a vector database (vector store).
  2. Query embedding: convert the user’s query into a vector so you can perform a similarity search against the stored vectors.
Always use the same embedding model for both ingestion and query embedding. Mixing embedding models will produce vectors in different vector spaces and break similarity comparisons.
Embedding models vs. generation (foundation) models
Slide titled "Workflow: Separation" showing two side-by-side boxes comparing Embedding Models (good at semantic positioning, used for search, not used for generation) and Generation Models (good at language production, only sees text not vectors, better reasoning). The left box is blue and the right box is orange on a dark background.
  • Embedding models: map text into a vector space for semantic positioning and similarity search. They are not used to generate natural language responses.
  • Generation (foundation) models: accept plain text (not vectors) as input and produce text outputs. They reason, synthesize, and produce responses, but they do not operate on vectors directly.
Typical Retrieval-Augmented Generation (RAG) workflow
  1. Ingestion: chunk documents, embed each chunk using an embedding model, and store vectors in a vector store.
  2. Query embedding: embed the incoming user query using the same embedding model.
  3. Retrieval: perform a similarity search in the vector store; retrieve the top-k most relevant chunk(s) as plain text.
  4. Augmented prompt assembly: combine the user question and retrieved text chunks into a single prompt (plain text).
  5. Generation: send the augmented prompt to a foundation model (e.g., LLaMA, Anthropic Claude, etc.) to produce the final answer.
Important: vectors live in the retrieval layer and are invisible to the LLM. The LLM only receives the retrieved text chunks as context. Why use a vector database in RAG? Key takeaway: embeddings enable meaning-based (semantic) search, making it possible to retrieve the right information for RAG systems even when the user doesn’t use exact keywords.
A presentation slide titled "Key Takeaway" with a short point: "Embeddings enable meaning-based search, making it possible to retrieve the right information for RAG systems." The slide has a dark blue left panel and a large white content area on the right.
Next steps
  • The next section walks through an end-to-end RAG implementation: ingestion pipelines, vector store operations (indexing and search), and the exact sequence that runs when a user query hits your application.
Links and references

Watch Video