> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Data Privacy and Protection Part 1

> Guidance on protecting sensitive data in Amazon Bedrock GenAI applications using redaction, input/output filtering, scoped RAG retrievals, and least privilege controls.

In this lesson we examine data privacy and protection for generative AI (GenAI) applications built on Amazon Bedrock, focusing on practical ways to avoid exposing sensitive data.

What you’ll learn

* The core problem: how sensitive data can be sent to or generated by foundation models in Bedrock.
* Defensive strategies for Bedrock-enabled applications.
* How to implement input and output filtering and redaction.
* Expected results, a key takeaway, and next steps.

Problem statement

When a user submits a prompt to Bedrock, that prompt may already contain sensitive data (for example, database records pasted by a user). Sensitive data can enter a GenAI workflow in three common ways:

1. Direct user input — a user pastes confidential text into the prompt.
2. Retrieval-Augmented Generation (RAG) — a RAG pipeline performs a semantic query, retrieves document chunks, and inserts those chunks into the model context.
3. Model output — the model echoes or generates sensitive information that originated in the prompt or the RAG context.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Governance-Safety-in-Responsible-AI/Data-Privacy-and-Protection-Part-1/user-bedrock-knowledgebase-rag-leakage-diagram.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=63d0de575ecfee2ae3d51519e398f221" alt="A diagram titled &#x22;Problem&#x22; that maps interactions between a User, an AI service labeled &#x22;Bedrock,&#x22; and a Knowledge Base with numbered arrows showing prompts, queries, chunk retrieval, and responses. Below the flow are three labeled boxes outlining leakage points: &#x22;In the prompt,&#x22; &#x22;In the RAG context,&#x22; and &#x22;In the response.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Governance-Safety-in-Responsible-AI/Data-Privacy-and-Protection-Part-1/user-bedrock-knowledgebase-rag-leakage-diagram.jpg" />
</Frame>

Why this matters

If your application returns the raw model output to users without controls, you risk regulatory, legal, and reputational consequences. Depending on jurisdiction and sector, a leakage could trigger fines or mandatory disclosures under frameworks such as:

* GDPR — General Data Protection Regulation: [https://gdpr.eu/](https://gdpr.eu/)
* HIPAA — U.S. Health Insurance Portability and Accountability Act: [https://www.hhs.gov/hipaa/index.html](https://www.hhs.gov/hipaa/index.html)
* PCI DSS — Payment Card Industry Data Security Standard: [https://www.pcisecuritystandards.org/](https://www.pcisecuritystandards.org/)

A single exposure of personally identifiable information (PII), financial data, or health records can cause customer harm, loss of trust, and business impact.

Key defensive strategies

Apply layered protections at each stage of the pipeline — before calling Bedrock, during retrieval, and after receiving model outputs.

* Minimize what you send. Only include the information required to perform the task. This reduces both token costs and the surface area for leaks.
* Mask or redact sensitive data. Detect patterns and fields (credit card numbers, national IDs, account numbers, email addresses, phone numbers, etc.) and redact them prior to sending to the model.
* Limit RAG retrieval scope and chunk size. Return only targeted snippets instead of whole documents to reduce exposing unrelated sensitive content.
* Enforce least privilege with IAM and access controls. Restrict which services, users, or roles can call specific models or access data stores.
* Validate and sanitize model outputs. Verify output formats (e.g., valid JSON) and apply post-generation filters to redact any leaked sensitive values.

Table — Leakage points and recommended mitigations

| Leakage point | Example risk | Recommended mitigations |
| - | - | - |
| In the prompt | User pastes PII into prompt | Input validation, redaction, user warnings |
| In the RAG context | Retrieval returns full documents with sensitive fields | Scoped queries, smaller chunks, snippet filtering |
| In the response | Model echoes secrets or generates new sensitive content | Output validation, redaction, content policies |

Apply these measures both before calling Bedrock and after the model responds — do not depend solely on the model or a vendor to perform these protections for you.

A practical protection flow

1. User input → Filter and redact sensitive or harmful content before storage or forwarding.
2. RAG / Data retrieval → Filter and redact retrieved content before including it in the model prompt.
3. Bedrock runtime → Call the model to generate the response.
4. Validate and filter model response → Check format (for example, is this valid JSON?), required keys, and redact any leaked sensitive data.
5. Deliver sanitized response → Return only the cleaned, validated content to the user.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Governance-Safety-in-Responsible-AI/Data-Privacy-and-Protection-Part-1/bedrock-protect-data-filter-redact.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=b22da79e9ca3511f0923ba929764c1df" alt="A slide titled &#x22;Solution: Protect Data in Bedrock Applications&#x22; showing a flowchart from User Input → Filter/Redact → Bedrock Runtime → Validate/Filter Output → User Response, with Secure Data Sources connected into the filtering stage." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Governance-Safety-in-Responsible-AI/Data-Privacy-and-Protection-Part-1/bedrock-protect-data-filter-redact.jpg" />
</Frame>

Practical examples

* Redaction via pattern matching (simple example)

```javascript theme={null}
// Simple JavaScript example to redact credit-card-like numbers
const redact = (text) =>
  text.replace(/\b(?:\d[ -]*?){13,16}\b/g, '[REDACTED-CARD]');

console.log(redact("Customer card: 4111 1111 1111 1111"));
// Output: Customer card: [REDACTED-CARD]
```

* Output validation (pseudo-code)

```pseudo theme={null}
response = call_model(prompt)
if not is_valid_json(response):
  log_and_fail("Invalid model output format")
else:
  sanitized = redact_sensitive_values(response)
  return sanitized
```

Bedrock-specific considerations and assurances

Customers often ask whether vendor-hosted models (e.g., ChatGPT, Anthropic’s Claude) retain inputs for model training. For Bedrock, the important points to communicate:

* Bedrock is a managed service hosting foundation models on AWS-managed infrastructure. Refer to the current AWS Bedrock documentation and service terms for the latest data usage commitments, but AWS does not use customer inference data to train the underlying foundation models.
* Data sent to Bedrock for inference stays within your AWS account boundary and the managed service workflow.
* Transport channels such as Direct Connect or VPN provide private connectivity into an AWS Region.
* Access to models is controlled by AWS IAM. Calls to a model without required permissions will fail.
* In-transit and at-rest encryption are used: TLS (HTTPS) and AWS encryption mechanisms protect data. Intercepted assets remain encrypted and unusable without keys.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Governance-Safety-in-Responsible-AI/Data-Privacy-and-Protection-Part-1/bedrock-workflow-security-slide.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=b29f2fd7b5fd587f4502e01c4a38abbf" alt="A dark-themed presentation slide titled &#x22;Workflow: What Bedrock Already Provides&#x22; showing a green Bedrock logo and a list of security features: &#x22;Your data isn't used to train models,&#x22; &#x22;Data stays in your AWS account,&#x22; &#x22;IAM controls access to models,&#x22; and &#x22;Encryption at rest and in transit.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Governance-Safety-in-Responsible-AI/Data-Privacy-and-Protection-Part-1/bedrock-workflow-security-slide.jpg" />
</Frame>

<Callout icon="lightbulb" color="#1CB2FE">
  Apply layered protections: minimize inputs, redact and mask sensitive data, scope RAG retrievals, validate outputs, and enforce least-privilege access. Your application must implement these safeguards both before and after calls to Bedrock.
</Callout>

Next steps and additional resources

Further lessons will provide deeper, hands-on techniques and code patterns for:

* Implementing robust input/output filtering and redaction.
* Designing safe RAG pipelines with scoped retrieval and chunking strategies.
* Integrating IAM-based least-privilege access and secure networking.

References

* Bedrock documentation: [https://docs.aws.amazon.com/bedrock](https://docs.aws.amazon.com/bedrock) (check the current AWS docs for up-to-date details)
* GDPR overview: [https://gdpr.eu/](https://gdpr.eu/)
* HIPAA information: [https://www.hhs.gov/hipaa/index.html](https://www.hhs.gov/hipaa/index.html)
* PCI DSS: [https://www.pcisecuritystandards.org/](https://www.pcisecuritystandards.org/)

If you want, I can follow up with a code-first lesson that shows an end-to-end example of a safe RAG pipeline (retrieval, redaction, model call, output validation).

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/1e4cbaa1-e041-4afb-828f-3045e5003b60/lesson/35d63cbd-a54c-4f77-97d7-f999e2ca04e9" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.