Skip to main content
In this lesson we explain embeddings and vector stores for retrieval-augmented generation (RAG) systems. When your system must search across thousands of documents, how do you reliably find the right information for a user query? Traditional keyword search matches strings, not meaning. This article covers the real-world problem and the solution—embedding vectors and vector stores—how they work, what to store, common ingestion patterns, and practical tips for chunking and context windows.
A slide titled "Lecture Flow" showing a flowchart from "Real-World Problem" (keyword search is not semantic) to "Solution" (embedding vectors) to "Workflow." Below it are boxes for "Results," "Key Takeaway," and "What's Next," emphasizing embeddings and more relevant, semantic search.

The problem: keyword search is not semantic

If a RAG system relies on keyword matching, it can miss semantically relevant documents. For example, a user asks, “How do I get my money back?” A document titled “Refund Policy” may contain exactly the information the user needs, but a naive string-based search might not match because the query and the document don’t share tokens. Humans map different words and phrases to the same intent; embeddings allow machines to do the same by mapping text to a geometric space where similar meanings are close together.
A presentation slide titled "Problem: Keyword Search Is Not Semantic" with three numbered panels. The panels explain that RAG must search your data using prompts, that "refund policy" can match "getting my money back," and that search should match meaning rather than exact string matches.

How embeddings solve the problem

Embeddings convert text into numeric vectors that capture semantic meaning. The basic workflow:
  • Convert the user’s query into an embedding vector (a high-dimensional numeric array).
  • Convert document text (entire documents or smaller chunks) into embeddings and store them.
  • Perform a similarity (nearest-neighbor) search in vector space to find document chunks whose vectors are near the query vector.
Because semantically similar text produces vectors that are spatially close, you retrieve relevant content even without shared keywords. Common similarity metrics include cosine similarity and dot product; many vector databases use approximate nearest neighbor (ANN) algorithms for speed at scale.
A slide titled "Solution: Embedding Vectors and Stores" with three colored callouts describing enabling semantic search, converting prompts to vectors, and generating similar vectors. To the right is a 3D coordinate diagram showing embedding vectors for "Return" and "Refund" and an example numeric vector.

What is a vector store (vector database) and why it matters

A vector store is the engine that powers semantic search. It stores:
  1. Embedding vectors (numeric representations).
  2. The original text or document chunk.
  3. Metadata (source, document ID, chunk index, timestamps, etc.).
The vector store runs similarity searches (e.g., using cosine similarity or dot product) and returns the most relevant document chunks. Those chunks are then used to augment prompts sent to the language model. Importantly, you retrieve text for the LLM—not raw vectors.
A presentation slide titled "Solution: What Is a Vector Store?" showing five numbered panels that explain features: stores embedding vectors, stores original text, allows similarity search, returns the most relevant chunks, and stores numbers representing meaning. Each panel includes a simple icon and a short caption.

Where to host your vector store (AWS and other options)

Amazon Bedrock provides models and inference but does not itself function as a vector database. You must choose a separate store. Common hosting options include managed services, object stores, relational databases with vector extensions, and graph/analytics platforms. Each option has trade-offs in cost, latency, scalability, and features (indexing, ANN support, replication).
A presentation slide titled "Solution: What Is a Vector Store?" showing four colored icons labeled AWS OpenSearch Serverless, AWS S3 Vectors, AWS PostgreSQL, and AWS Neptune Analytics. The icons represent different AWS services that can act as vector stores.
Recommended reading:

What to store and the ingestion workflow

When populating a vector store you typically persist:
  • The embedding vector.
  • The source text (or document chunk).
  • Metadata describing source, position, file type, timestamps, etc.
Key design questions:
  • Do you embed whole documents or smaller chunks?
  • When are embeddings created and stored?
Two common embedding creation patterns: Typical RAG flow:
  1. At ingestion, split documents into chunks (optional), compute embeddings, and store vectors + text + metadata.
  2. At query time, compute an embedding for the user query.
  3. Query the vector store for the nearest neighbors.
  4. Retrieve the corresponding text chunks and metadata.
  5. Construct an augmented prompt (user query + retrieved text) and send that to the LLM.
Remember: the LLM receives text context, not the numerical vectors.
An infographic slide titled "Workflow: Are Whole Documents Embedded?" showing four colored panels that say documents can be long, long documents can contain multiple topics, searching entire documents reduces accuracy, and models have context limits.

Chunking and model context limits

Embedding very long documents as a single vector can reduce retrieval precision because multiple topics get conflated. Best practices:
  • Split long documents into semantically coherent chunks (paragraphs, sections, or topic-based segments).
  • Store chunk-level metadata: source, chunk index, and optionally a short summary.
  • Keep the model’s context window in mind: the augmented prompt (user query + retrieved text) must fit within the LLM’s token limit.
  • Limit the number of retrieved chunks and tune chunk size to balance relevance, cost, and token consumption.
Chunking tips: choose chunk sizes that capture coherent ideas (e.g., ~200–800 tokens), include moderate overlap between chunks (10–20%) to preserve boundary context, and limit the number of retrieved chunks to stay within the model’s context window.

Quick recap

  • Keyword search matches strings; embeddings enable semantic search by mapping meaning to vectors.
  • Vector stores (vector databases) hold embeddings, text, and metadata and perform nearest-neighbor searches to return the most relevant chunks.
  • Create embeddings at ingestion to populate a store, and create query-time embeddings to search that store.
  • Chunk documents for better retrieval accuracy and to manage the amount of context sent to the model.
Next lesson: we will demonstrate how to generate embeddings programmatically and integrate a vector store into a full RAG pipeline, including example code and deployment considerations.

Watch Video