> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# RAG End to End Workflow

> Explains end to end Retrieval Augmented Generation workflow from document ingestion and vectorization to query retrieval, prompt augmentation, and grounded LLM answer generation.

This lesson walks through a complete Retrieval-Augmented Generation (RAG) workflow—from a user question to a grounded answer. Understanding each step and how they connect helps you design, debug, and optimize reliable RAG systems.

We’ll cover:

* Ingestion: turning source documents into vectors and storing them.
* Query processing: embedding the user query, retrieving relevant chunks, and augmenting the prompt.
* Generation: sending the augmented prompt to a foundation model to produce a grounded answer.

Example scenario: an organization has product documents for a device called the iUniverse Pro. The documents contain facts like battery life (18 hours), warranty (2 years), and Bluetooth support (5.3). When a user asks “How long does the iUniverse Pro battery last?” the RAG pipeline should return a grounded answer such as “18 hours.”

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/rag-workflow-iuniverse-pro-battery.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=46d39081b428d56a5ba3462b812ff590" alt="An infographic titled &#x22;Workflow: Example RAG Process Flow&#x22; showing a three-step process with numbered circular nodes connected to blue rounded content boxes. The boxes include a user question (&#x22;How long does the iUniverse Pro battery last?&#x22;) and document details listing battery life (18 hours), warranty (2 years) and Bluetooth support (5.3)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/rag-workflow-iuniverse-pro-battery.jpg" />
</Frame>

## Core concepts (quick)

* Embeddings: numeric vectors that capture semantic meaning of text.
* Vector store: an index/database that stores vectors + original text + metadata to enable fast semantic search.
* Grounding: supplying retrieved source text to a foundation model so its answer is traceable to document facts.
* Consistency: use the same embedding model (or compatible models) for ingestion and query-time embeddings.

## Ingestion: documents → chunks → vectors → vector store

Ingestion prepares your proprietary content for retrieval:

1. Chunking: split long documents into smaller, meaningful chunks (paragraphs or logical sections).
2. Embedding: convert each chunk into a numeric embedding using an embedding model.
3. Storage: save the embedding together with the original chunk text and metadata in a vector store (examples: S3 Vectors, Postgres with pgvector, OpenSearch).

Store both vector and original text so the vector enables semantic match and the text provides grounded context for the LLM.

Example stored record (JSON):

```json theme={null}
{
  "vector": [0.12, -0.33, 0.91, 0.04, ...],
  "text": "The iUniverse Pro provides up to 18 hours of battery life under typical usage.",
  "metadata": {
    "source": "iUniverse_Pro_product_manual.pdf",
    "chunk_id": 3
  }
}
```

Table — Common vector store options

| Vector store | Use case | Notes |
| - | - | - |
| S3 Vectors | Large-scale storage + custom search | Good for cost-effective object storage with vector index layers |
| Postgres (pgvector) | Relational + semantic search | Integrates with existing SQL workflows |
| OpenSearch | Search + analytics | Useful when mixing text search and vectors |

Choose an embedding model and remain consistent for ingestion and query so similarity comparisons are meaningful.

## Query: user asks a question

When an app receives a user question, the RAG system must embed the question to perform semantic search against the stored vectors.

Example user question:

* "How long does the iUniverse Pro battery last?"

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/workflow-query-step1-user-sends-prompt.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=068043016477b6508f1be5740a85898b" alt="A workflow diagram titled &#x22;Workflow: Query Step 1 – User Sends Prompt&#x22; shows the question &#x22;How long does the iUniverse Pro battery last?&#x22; in a speech bubble alongside a small robot illustration. Below are five numbered steps describing the retrieval/embedding flow: chunking and generating embeddings, storing them, embedding the user query, retrieving similar chunks, and sending the augmented query to an LLM." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/workflow-query-step1-user-sends-prompt.jpg" />
</Frame>

### Embed the user query

Convert the user text into an embedding using the same model used during ingestion. For example, using an Amazon Titan embedding model:

```json theme={null}
{
  "model": "amazon.titan-embed-text-v2",
  "input": "How long does the iUniverse Pro battery last?",
  "output": [0.0123, -0.8834, 0.4421, 0.1945, ...]
}
```

Note: embedding the query is a lightweight operation and does not contact the generative foundation model (LLM) yet — it only prepares for semantic retrieval.

### Similarity search in the vector store

Run a nearest-neighbor (similarity) search using the query embedding. The vector store returns the most semantically similar chunks. Importantly, what you receive back is the chunk text and metadata — not the vector — because the text will be used for prompt augmentation and grounding.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/workflow-query-step3-similarity-search.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=2919f8adf7702894d8b2fcd67608e03e" alt="The image is a presentation slide titled &#x22;Workflow: Query Step 3 — Similarity Search in Vector Store&#x22; showing a diagram of a RAG/vector-search pipeline. It highlights chunking and embedding, finding the most similar chunk, and returning that chunk, with five numbered boxes summarizing the full embedding-and-retrieval steps." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/workflow-query-step3-similarity-search.jpg" />
</Frame>

At this point the RAG pipeline has:

* The original user question.
* One or more retrieved document chunks (plain text) that are semantically relevant to the question.

## Prompt augmentation and generation

Combine the retrieved context with the user question to create an augmented prompt for the foundation model. This structure gives the model explicit facts it can cite.

Augmented prompt structure (recommended order):

* System/instruction layer (optional): desired response format, safety filters, style constraints.
* Retrieved context: one or more relevant chunks from the vector store, with citations/metadata.
* User question: the original prompt to answer.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/workflow-query-step4-prompt-augmentation.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=b9db94977f49f3e4fe9d04b88c842333" alt="An infographic titled &#x22;Workflow: Query Step 4 – Prompt Augmentation&#x22; showing the question &#x22;How long does the iUniverse Pro battery last?&#x22; being combined with product chunks and sent into a RAG system. Below are five numbered steps: generate embeddings, store them, embed the user query, retrieve similar chunks, and send the augmented query to an LLM." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/workflow-query-step4-prompt-augmentation.jpg" />
</Frame>

Send the augmented prompt to your chosen foundation model (for example, an Amazon Bedrock model such as a Meta Llama variant). The model reads the provided context and generates a response that is grounded in the supplied document text — reducing hallucinations and improving traceability.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/workflow-query-step5-foundation-model-answer.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=94f7515a2a8f358b836fa66af2194740" alt="A presentation slide titled &#x22;Workflow: Query Step 5 – Foundation Model Generate Answer&#x22; showing three dark rounded panels numbered 01–03 with icons (book/magnifier, battery, checkmark) and short explanatory text about reading context, the iUniverse Pro's 18‑hour battery, and the answer being grounded. The slide is branded with a © Copyright KodeKloud note." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/workflow-query-step5-foundation-model-answer.jpg" />
</Frame>

Because the model has the retrieved chunks in the prompt, the response can be phrased confidently and traced back to the provided documents — this is what we mean by grounding.

## Putting the full pipeline together

High-level flow:

* App receives a user prompt → passes it to the RAG pipeline.
* (Preceding ingestion step): documents were chunked, embedded, and stored in a vector store.
* Query-time operations: embed the user query → similarity search → retrieve text chunks → augment the prompt → send to the foundation model → generate a grounded answer.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/rag-pipeline-ingestion-embeddings-vectorstore-workflow.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=cc3ab4e78f8251eddd1995a9d997b174" alt="A slide-sized workflow diagram titled &#x22;Workflow: App Actions vs Bedrock Actions&#x22; showing a RAG pipeline with boxes for App, RAG System, embedding model, Vector Store, Ingestion, and a chosen model, connected by arrows for prompts, vectors, and grounded responses. The diagram visualizes how documents are ingested, chunked into embeddings, searched, and used to produce grounded answers." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/RAG-End-to-End-Workflow/rag-pipeline-ingestion-embeddings-vectorstore-workflow.jpg" />
</Frame>

Quick checklist for production RAG systems

* Use consistent embedding models (or compatible dimensionality/semantics) across ingestion and query.
* Store text + metadata with vectors for traceability.
* Limit retrieved context size to fit model input tokens; prefer high-relevance chunks.
* Add system prompts to constrain format and enforce safety.
* Track provenance (source file, chunk id, timestamp) for every returned chunk.

<Callout icon="lightbulb" color="#1CB2FE">
  Key takeaway: RAG has two distinct phases — ingestion (embed and store document chunks) and query (embed query, retrieve similar chunks, augment the prompt, and generate a grounded answer). Ensure you use a consistent embedding approach between ingestion and query for effective retrieval and traceable responses.
</Callout>

## Links and References

* Amazon Bedrock documentation: [https://docs.aws.amazon.com/bedrock](https://docs.aws.amazon.com/bedrock)
* Embeddings and semantic search overview: [https://en.wikipedia.org/wiki/Semantic\_search](https://en.wikipedia.org/wiki/Semantic_search)
* pgvector (vector extension for Postgres): [https://github.com/pgvector/pgvector](https://github.com/pgvector/pgvector)

This RAG explanation applies directly to Amazon Bedrock and Bedrock Knowledge Bases and can be adapted to other embedding models, vector stores, and foundation models.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/282d7660-0e67-4e7d-a498-291bc16784e5/lesson/89f98d8b-83f6-4202-96e3-56de9543cd47" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.