> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Bedrock Knowledge Bases Introduction Part 2

> Tutorial on building and configuring Amazon Bedrock knowledge bases, covering naming, IAM roles, data sources, parsing strategies, embedding models, vector stores, and cost versus quality tradeoffs.

In this lesson we create a knowledge base (KB) from the Amazon Bedrock console and walk through the key configuration choices: KB name, IAM role, data source, parsing strategy, embedding model, and vector store. Each choice directly affects the quality, cost, and performance of your retrieval-augmented generation (RAG) solution.

## Quick overview: configuration sequence

1. KB name — a descriptive identifier for the KB’s purpose and scope.
2. IAM role — the security principal that grants Bedrock access to resources (for example S3).
3. Data source — where documents live (S3, external storage, etc.).
4. Parsing strategy — how content is extracted before chunking and embedding.
5. Embedding model — converts text into vectors.
6. Vector store — where vectors are stored and how similarity search is performed.

## Naming and IAM role

Start with a clear, descriptive name that communicates the KB’s domain and scope. Then select or create an IAM service role that grants Bedrock the minimal permissions required to access the data source.

* If your documents are in an S3 bucket, the role must include S3 read permissions (or cross-account access if the bucket is in another AWS account).
* When creating a service role from the console, the role wizard often pre-populates the base policies necessary for common data sources.

## Data source: where your documents live

Choose the source location for the documents you want to ingest. Typical formats: PDFs, DOCX, CSV, plain text, HTML, or scanned images (OCR required). In this example the documents are stored in a standard S3 bucket.

Example identifiers used in this sample configuration:

```text theme={null}
new-knowledge-base
s3://hodei-bedrock-intro
```

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-2/workflow-kb-embeddings-vector-store.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=1e0725dfb950ec30a1fb493fe2cc5146" alt="A screenshot of a web UI titled &#x22;Workflow: Creating a KB&#x22; showing the &#x22;Configure data storage and processing&#x22; step with an embeddings model selected (Titan Text Emb...) and vector store configuration options. The left sidebar shows the step-by-step navigation for creating the knowledge base." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-2/workflow-kb-embeddings-vector-store.jpg" />
</Frame>

## Parsing strategy: extract before chunking

Parsing determines how text and structural information are extracted from source documents prior to chunking and embedding. Choose a parser that aligns with your document types and budget.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-2/workflow-parsing-strategy-document-parsers.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=cedb69b28a74538c9eb418c58b2f4798" alt="A slide titled &#x22;Workflow: Parsing Strategy&#x22; showing a table that maps document types (plain text, PDFs, forms, HTML, knowledge articles, mixed data) to their typical content and recommended parsers (default parser, data automation, or foundation model parser)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Retrieval-Augmentation-With-Bedrock-Knowledge-Bases/Bedrock-Knowledge-Bases-Introduction-Part-2/workflow-parsing-strategy-document-parsers.jpg" />
</Frame>

Table: parser summary, when to use each, and cost trade-offs

| Parser | Best for | Pros | Cons |
| - | - | - | - |
| Default parser | Plain text, simple PDFs, basic documents | Fast, built-in, no per-page or LLM costs | Limited at extracting complex tables, forms, multi-column layouts |
| Data automation parser | Structured or semi-structured documents (invoices, forms, consistent layouts) | Better extraction for consistent formats | Scans per-page — incurs per-page charges |
| Foundation model parser | Messy or highly heterogeneous documents (HTML-heavy, scanned pages, irregular layouts) | Highest semantic extraction quality using an LLM | Consumes model tokens — higher LLM inference cost |

Summary guidance:

* Use the default parser when it delivers acceptable search quality to minimize costs.
* Use the data automation parser for consistently formatted documents that contain tables/fielded data; accept per-page fees for better extraction.
* Use the foundation model parser when other parsers fail to capture the required semantics; expect higher token-based costs.

<Callout icon="warning" color="#FF6B6B">
  Parser choice affects both data quality and cost. Prefer the default parser for simple datasets to avoid extra charges. Use the data automation parser for consistent, structured documents (expect per-page fees). Reserve the foundation model parser for messy or complex documents when you need deep semantic extraction—be aware of potentially high token-based costs.
</Callout>

## Embedding model

Embedding models convert chunks of text into numeric vectors that capture semantic meaning. At the time of this lesson, embedding options in the Bedrock model catalog are more limited than generative models; Amazon and Cohere are common embedding vendors.

* Choose an embedding model that balances vector quality and inference cost.
* In this example, we select Amazon Titan Text Embeddings for embedding generation.

Why embeddings matter: the retrieval layer performs similarity search (for example cosine similarity) over these vectors to find relevant passages for answer generation. Higher-quality embeddings generally improve retrieval relevance.

## Vector store: where vectors live and how they are searched

Select a vector store that meets your scale, latency, cost, and feature requirements. You can create a new vector store during KB setup or point the KB at an existing one.

Table: common vector store options and typical use cases

| Vector store | Typical use case | Notes |
| - | - | - |
| Amazon OpenSearch Serverless | Low-latency search at scale | Provides integrated vector/semantic search features |
| Amazon S3 Vectors (vector bucket) | Low-cost storage for vectors | Good for archiving or low-query-rate systems |
| Amazon Aurora PostgreSQL | Vector search with relational features | Use when you need SQL-based workflows or transactional guarantees |
| Amazon Neptune Analytics | Graph-based retrieval and analytics | Good for graph relationships and complex traversal queries |

Choosing a vector store affects retrieval latency, scale, and operational complexity. Consider expected query volume, desired latency, and integrations with your analytics or application stack.

## Putting it all together: trade-offs and recommendations

* Cost vs. quality: Default parser + basic embedding model minimizes cost. Foundation model parser + high-quality embeddings maximizes semantic accuracy but increases costs.
* Speed vs. accuracy: Simpler parsing and embeddings produce faster ingestion and lower latency. Use advanced parsers/embeddings when retrieval quality requires it.
* Storage and scaling: Choose a vector store that matches your throughput and operational needs.

Checklist for creating a KB in Bedrock

| Step | Action |
| - | - |
| 1 | Choose a descriptive KB name (e.g., `new-knowledge-base`) |
| 2 | Assign or create an IAM role with the minimal required permissions for your data source |
| 3 | Point the KB to your data source (e.g., `s3://hodei-bedrock-intro`) |
| 4 | Select a parsing strategy based on document complexity and budget |
| 5 | Select an embedding model (e.g., Titan Text Embeddings) |
| 6 | Create or select a vector store (OpenSearch Serverless, S3 Vectors, Aurora, Neptune) |

Choosing the right parser, embedding model, and vector store together determines the effectiveness of your RAG pipeline. Make trade-offs between cost, speed, and extraction quality based on your dataset and the search experience you require.

## Links and references

* [Amazon Bedrock documentation](https://docs.aws.amazon.com/bedrock)
* [Amazon OpenSearch Service](https://aws.amazon.com/opensearch-service/)
* [RAG fundamentals course (example)](https://learn.kodekloud.com/user/courses/fundamentals-of-rag)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/282d7660-0e67-4e7d-a498-291bc16784e5/lesson/840b9efd-db99-45df-813b-c94659db6266" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.