Skip to main content
In this lesson we examine data privacy and protection for generative AI (GenAI) applications built on Amazon Bedrock, focusing on practical ways to avoid exposing sensitive data. What you’ll learn
  • The core problem: how sensitive data can be sent to or generated by foundation models in Bedrock.
  • Defensive strategies for Bedrock-enabled applications.
  • How to implement input and output filtering and redaction.
  • Expected results, a key takeaway, and next steps.
Problem statement When a user submits a prompt to Bedrock, that prompt may already contain sensitive data (for example, database records pasted by a user). Sensitive data can enter a GenAI workflow in three common ways:
  1. Direct user input — a user pastes confidential text into the prompt.
  2. Retrieval-Augmented Generation (RAG) — a RAG pipeline performs a semantic query, retrieves document chunks, and inserts those chunks into the model context.
  3. Model output — the model echoes or generates sensitive information that originated in the prompt or the RAG context.
A diagram titled "Problem" that maps interactions between a User, an AI service labeled "Bedrock," and a Knowledge Base with numbered arrows showing prompts, queries, chunk retrieval, and responses. Below the flow are three labeled boxes outlining leakage points: "In the prompt," "In the RAG context," and "In the response."
Why this matters If your application returns the raw model output to users without controls, you risk regulatory, legal, and reputational consequences. Depending on jurisdiction and sector, a leakage could trigger fines or mandatory disclosures under frameworks such as: A single exposure of personally identifiable information (PII), financial data, or health records can cause customer harm, loss of trust, and business impact. Key defensive strategies Apply layered protections at each stage of the pipeline — before calling Bedrock, during retrieval, and after receiving model outputs.
  • Minimize what you send. Only include the information required to perform the task. This reduces both token costs and the surface area for leaks.
  • Mask or redact sensitive data. Detect patterns and fields (credit card numbers, national IDs, account numbers, email addresses, phone numbers, etc.) and redact them prior to sending to the model.
  • Limit RAG retrieval scope and chunk size. Return only targeted snippets instead of whole documents to reduce exposing unrelated sensitive content.
  • Enforce least privilege with IAM and access controls. Restrict which services, users, or roles can call specific models or access data stores.
  • Validate and sanitize model outputs. Verify output formats (e.g., valid JSON) and apply post-generation filters to redact any leaked sensitive values.
Table — Leakage points and recommended mitigations Apply these measures both before calling Bedrock and after the model responds — do not depend solely on the model or a vendor to perform these protections for you. A practical protection flow
  1. User input → Filter and redact sensitive or harmful content before storage or forwarding.
  2. RAG / Data retrieval → Filter and redact retrieved content before including it in the model prompt.
  3. Bedrock runtime → Call the model to generate the response.
  4. Validate and filter model response → Check format (for example, is this valid JSON?), required keys, and redact any leaked sensitive data.
  5. Deliver sanitized response → Return only the cleaned, validated content to the user.
A slide titled "Solution: Protect Data in Bedrock Applications" showing a flowchart from User Input → Filter/Redact → Bedrock Runtime → Validate/Filter Output → User Response, with Secure Data Sources connected into the filtering stage.
Practical examples
  • Redaction via pattern matching (simple example)
  • Output validation (pseudo-code)
Bedrock-specific considerations and assurances Customers often ask whether vendor-hosted models (e.g., ChatGPT, Anthropic’s Claude) retain inputs for model training. For Bedrock, the important points to communicate:
  • Bedrock is a managed service hosting foundation models on AWS-managed infrastructure. Refer to the current AWS Bedrock documentation and service terms for the latest data usage commitments, but AWS does not use customer inference data to train the underlying foundation models.
  • Data sent to Bedrock for inference stays within your AWS account boundary and the managed service workflow.
  • Transport channels such as Direct Connect or VPN provide private connectivity into an AWS Region.
  • Access to models is controlled by AWS IAM. Calls to a model without required permissions will fail.
  • In-transit and at-rest encryption are used: TLS (HTTPS) and AWS encryption mechanisms protect data. Intercepted assets remain encrypted and unusable without keys.
A dark-themed presentation slide titled "Workflow: What Bedrock Already Provides" showing a green Bedrock logo and a list of security features: "Your data isn't used to train models," "Data stays in your AWS account," "IAM controls access to models," and "Encryption at rest and in transit."
Apply layered protections: minimize inputs, redact and mask sensitive data, scope RAG retrievals, validate outputs, and enforce least-privilege access. Your application must implement these safeguards both before and after calls to Bedrock.
Next steps and additional resources Further lessons will provide deeper, hands-on techniques and code patterns for:
  • Implementing robust input/output filtering and redaction.
  • Designing safe RAG pipelines with scoped retrieval and chunking strategies.
  • Integrating IAM-based least-privilege access and secure networking.
References If you want, I can follow up with a code-first lesson that shows an end-to-end example of a safe RAG pipeline (retrieval, redaction, model call, output validation).

Watch Video