Skip to main content
In this lesson we create a knowledge base (KB) from the Amazon Bedrock console and walk through the key configuration choices: KB name, IAM role, data source, parsing strategy, embedding model, and vector store. Each choice directly affects the quality, cost, and performance of your retrieval-augmented generation (RAG) solution.

Quick overview: configuration sequence

  1. KB name — a descriptive identifier for the KB’s purpose and scope.
  2. IAM role — the security principal that grants Bedrock access to resources (for example S3).
  3. Data source — where documents live (S3, external storage, etc.).
  4. Parsing strategy — how content is extracted before chunking and embedding.
  5. Embedding model — converts text into vectors.
  6. Vector store — where vectors are stored and how similarity search is performed.

Naming and IAM role

Start with a clear, descriptive name that communicates the KB’s domain and scope. Then select or create an IAM service role that grants Bedrock the minimal permissions required to access the data source.
  • If your documents are in an S3 bucket, the role must include S3 read permissions (or cross-account access if the bucket is in another AWS account).
  • When creating a service role from the console, the role wizard often pre-populates the base policies necessary for common data sources.

Data source: where your documents live

Choose the source location for the documents you want to ingest. Typical formats: PDFs, DOCX, CSV, plain text, HTML, or scanned images (OCR required). In this example the documents are stored in a standard S3 bucket. Example identifiers used in this sample configuration:
A screenshot of a web UI titled "Workflow: Creating a KB" showing the "Configure data storage and processing" step with an embeddings model selected (Titan Text Emb...) and vector store configuration options. The left sidebar shows the step-by-step navigation for creating the knowledge base.

Parsing strategy: extract before chunking

Parsing determines how text and structural information are extracted from source documents prior to chunking and embedding. Choose a parser that aligns with your document types and budget.
A slide titled "Workflow: Parsing Strategy" showing a table that maps document types (plain text, PDFs, forms, HTML, knowledge articles, mixed data) to their typical content and recommended parsers (default parser, data automation, or foundation model parser).
Table: parser summary, when to use each, and cost trade-offs Summary guidance:
  • Use the default parser when it delivers acceptable search quality to minimize costs.
  • Use the data automation parser for consistently formatted documents that contain tables/fielded data; accept per-page fees for better extraction.
  • Use the foundation model parser when other parsers fail to capture the required semantics; expect higher token-based costs.
Parser choice affects both data quality and cost. Prefer the default parser for simple datasets to avoid extra charges. Use the data automation parser for consistent, structured documents (expect per-page fees). Reserve the foundation model parser for messy or complex documents when you need deep semantic extraction—be aware of potentially high token-based costs.

Embedding model

Embedding models convert chunks of text into numeric vectors that capture semantic meaning. At the time of this lesson, embedding options in the Bedrock model catalog are more limited than generative models; Amazon and Cohere are common embedding vendors.
  • Choose an embedding model that balances vector quality and inference cost.
  • In this example, we select Amazon Titan Text Embeddings for embedding generation.
Why embeddings matter: the retrieval layer performs similarity search (for example cosine similarity) over these vectors to find relevant passages for answer generation. Higher-quality embeddings generally improve retrieval relevance.

Vector store: where vectors live and how they are searched

Select a vector store that meets your scale, latency, cost, and feature requirements. You can create a new vector store during KB setup or point the KB at an existing one. Table: common vector store options and typical use cases Choosing a vector store affects retrieval latency, scale, and operational complexity. Consider expected query volume, desired latency, and integrations with your analytics or application stack.

Putting it all together: trade-offs and recommendations

  • Cost vs. quality: Default parser + basic embedding model minimizes cost. Foundation model parser + high-quality embeddings maximizes semantic accuracy but increases costs.
  • Speed vs. accuracy: Simpler parsing and embeddings produce faster ingestion and lower latency. Use advanced parsers/embeddings when retrieval quality requires it.
  • Storage and scaling: Choose a vector store that matches your throughput and operational needs.
Checklist for creating a KB in Bedrock Choosing the right parser, embedding model, and vector store together determines the effectiveness of your RAG pipeline. Make trade-offs between cost, speed, and extraction quality based on your dataset and the search experience you require.

Watch Video