- Ingestion: turning source documents into vectors and storing them.
- Query processing: embedding the user query, retrieving relevant chunks, and augmenting the prompt.
- Generation: sending the augmented prompt to a foundation model to produce a grounded answer.

Core concepts (quick)
- Embeddings: numeric vectors that capture semantic meaning of text.
- Vector store: an index/database that stores vectors + original text + metadata to enable fast semantic search.
- Grounding: supplying retrieved source text to a foundation model so its answer is traceable to document facts.
- Consistency: use the same embedding model (or compatible models) for ingestion and query-time embeddings.
Ingestion: documents → chunks → vectors → vector store
Ingestion prepares your proprietary content for retrieval:- Chunking: split long documents into smaller, meaningful chunks (paragraphs or logical sections).
- Embedding: convert each chunk into a numeric embedding using an embedding model.
- Storage: save the embedding together with the original chunk text and metadata in a vector store (examples: S3 Vectors, Postgres with pgvector, OpenSearch).
Choose an embedding model and remain consistent for ingestion and query so similarity comparisons are meaningful.
Query: user asks a question
When an app receives a user question, the RAG system must embed the question to perform semantic search against the stored vectors. Example user question:- “How long does the iUniverse Pro battery last?”

Embed the user query
Convert the user text into an embedding using the same model used during ingestion. For example, using an Amazon Titan embedding model:Similarity search in the vector store
Run a nearest-neighbor (similarity) search using the query embedding. The vector store returns the most semantically similar chunks. Importantly, what you receive back is the chunk text and metadata — not the vector — because the text will be used for prompt augmentation and grounding.
- The original user question.
- One or more retrieved document chunks (plain text) that are semantically relevant to the question.
Prompt augmentation and generation
Combine the retrieved context with the user question to create an augmented prompt for the foundation model. This structure gives the model explicit facts it can cite. Augmented prompt structure (recommended order):- System/instruction layer (optional): desired response format, safety filters, style constraints.
- Retrieved context: one or more relevant chunks from the vector store, with citations/metadata.
- User question: the original prompt to answer.


Putting the full pipeline together
High-level flow:- App receives a user prompt → passes it to the RAG pipeline.
- (Preceding ingestion step): documents were chunked, embedded, and stored in a vector store.
- Query-time operations: embed the user query → similarity search → retrieve text chunks → augment the prompt → send to the foundation model → generate a grounded answer.

- Use consistent embedding models (or compatible dimensionality/semantics) across ingestion and query.
- Store text + metadata with vectors for traceability.
- Limit retrieved context size to fit model input tokens; prefer high-relevance chunks.
- Add system prompts to constrain format and enforce safety.
- Track provenance (source file, chunk id, timestamp) for every returned chunk.
Key takeaway: RAG has two distinct phases — ingestion (embed and store document chunks) and query (embed query, retrieve similar chunks, augment the prompt, and generate a grounded answer). Ensure you use a consistent embedding approach between ingestion and query for effective retrieval and traceable responses.
Links and References
- Amazon Bedrock documentation: https://docs.aws.amazon.com/bedrock
- Embeddings and semantic search overview: https://en.wikipedia.org/wiki/Semantic_search
- pgvector (vector extension for Postgres): https://github.com/pgvector/pgvector