
- Consider a 20‑page manual. One straightforward approach is to split it into chunks by page range (for instance, pages 1–6, 7–13, and 14–20) and embed each chunk separately. Each chunk becomes an independent unit stored in your vector store and returned by similarity search when relevant.

Both approaches follow the same principle: embed each chunk separately to increase the likelihood that retrieval returns directly relevant content.
How are embedding vectors created?
- An embedding model converts text into a high-dimensional numerical vector that captures semantic meaning.
- Different embedding models vary in vector dimensionality, training data, and architecture—vectors from different models live in different vector spaces and are not directly comparable.
- Ingestion: convert your document chunks into vectors and store them in a vector database (vector store).
- Query embedding: convert the user’s query into a vector so you can perform a similarity search against the stored vectors.
Always use the same embedding model for both ingestion and query embedding. Mixing embedding models will produce vectors in different vector spaces and break similarity comparisons.

- Embedding models: map text into a vector space for semantic positioning and similarity search. They are not used to generate natural language responses.
- Generation (foundation) models: accept plain text (not vectors) as input and produce text outputs. They reason, synthesize, and produce responses, but they do not operate on vectors directly.
- Ingestion: chunk documents, embed each chunk using an embedding model, and store vectors in a vector store.
- Query embedding: embed the incoming user query using the same embedding model.
- Retrieval: perform a similarity search in the vector store; retrieve the top-k most relevant chunk(s) as plain text.
- Augmented prompt assembly: combine the user question and retrieved text chunks into a single prompt (plain text).
- Generation: send the augmented prompt to a foundation model (e.g., LLaMA, Anthropic Claude, etc.) to produce the final answer.
Key takeaway: embeddings enable meaning-based (semantic) search, making it possible to retrieve the right information for RAG systems even when the user doesn’t use exact keywords.

- The next section walks through an end-to-end RAG implementation: ingestion pipelines, vector store operations (indexing and search), and the exact sequence that runs when a user query hits your application.
- Retrieval-Augmented Generation overview
- Vector database concepts and examples
- Embedding models and best practices