
The problem: keyword search is not semantic
If a RAG system relies on keyword matching, it can miss semantically relevant documents. For example, a user asks, “How do I get my money back?” A document titled “Refund Policy” may contain exactly the information the user needs, but a naive string-based search might not match because the query and the document don’t share tokens. Humans map different words and phrases to the same intent; embeddings allow machines to do the same by mapping text to a geometric space where similar meanings are close together.
How embeddings solve the problem
Embeddings convert text into numeric vectors that capture semantic meaning. The basic workflow:- Convert the user’s query into an embedding vector (a high-dimensional numeric array).
- Convert document text (entire documents or smaller chunks) into embeddings and store them.
- Perform a similarity (nearest-neighbor) search in vector space to find document chunks whose vectors are near the query vector.

What is a vector store (vector database) and why it matters
A vector store is the engine that powers semantic search. It stores:- Embedding vectors (numeric representations).
- The original text or document chunk.
- Metadata (source, document ID, chunk index, timestamps, etc.).

Where to host your vector store (AWS and other options)
Amazon Bedrock provides models and inference but does not itself function as a vector database. You must choose a separate store. Common hosting options include managed services, object stores, relational databases with vector extensions, and graph/analytics platforms. Each option has trade-offs in cost, latency, scalability, and features (indexing, ANN support, replication).
What to store and the ingestion workflow
When populating a vector store you typically persist:- The embedding vector.
- The source text (or document chunk).
- Metadata describing source, position, file type, timestamps, etc.
- Do you embed whole documents or smaller chunks?
- When are embeddings created and stored?
Typical RAG flow:
- At ingestion, split documents into chunks (optional), compute embeddings, and store vectors + text + metadata.
- At query time, compute an embedding for the user query.
- Query the vector store for the nearest neighbors.
- Retrieve the corresponding text chunks and metadata.
- Construct an augmented prompt (user query + retrieved text) and send that to the LLM.

Chunking and model context limits
Embedding very long documents as a single vector can reduce retrieval precision because multiple topics get conflated. Best practices:- Split long documents into semantically coherent chunks (paragraphs, sections, or topic-based segments).
- Store chunk-level metadata: source, chunk index, and optionally a short summary.
- Keep the model’s context window in mind: the augmented prompt (user query + retrieved text) must fit within the LLM’s token limit.
- Limit the number of retrieved chunks and tune chunk size to balance relevance, cost, and token consumption.
Chunking tips: choose chunk sizes that capture coherent ideas (e.g., ~200–800 tokens), include moderate overlap between chunks (10–20%) to preserve boundary context, and limit the number of retrieved chunks to stay within the model’s context window.
Quick recap
- Keyword search matches strings; embeddings enable semantic search by mapping meaning to vectors.
- Vector stores (vector databases) hold embeddings, text, and metadata and perform nearest-neighbor searches to return the most relevant chunks.
- Create embeddings at ingestion to populate a store, and create query-time embeddings to search that store.
- Chunk documents for better retrieval accuracy and to manage the amount of context sent to the model.
Links and references
- Retrieval-Augmented Generation (RAG) fundamentals
- Amazon Bedrock overview
- Vector database for GenAI (course)
- Amazon S3 basics
- PostgreSQL on AWS RDS