> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Gen AI and LLMs Introduction Part 2

> Overview of hosting, tool integration, and use cases for pre-trained generative AI models, including retrieval-augmented patterns and governance for up-to-date private data

Where you run or consume a model is a design choice. Many people are familiar with hosted services such as [ChatGPT](https://chat.openai.com/) that expose web UIs backed by GPT-family models, but you can also run models locally (on a laptop or a private server) or consume managed, hosted offerings. Running very large models locally is possible, but typically slow and GPU-intensive unless you have substantial specialized hardware. Given this course’s focus, examples will use Bedrock to demonstrate hosted patterns and governance.

Key distinctions to keep in mind:

* When a generative AI model answers a question about a historical event correctly, it’s using knowledge learned during pre-training. Pre-training ingests large corpora (public text, licensed data, and other sources), and the model’s responses reflect patterns from that data.
* A local model whose training data stops before a particular event or topic cannot reliably answer questions that postdate its training.
* Hosted services or model configurations can be augmented with external tools—web browsing, search, or plugins—so the system can fetch up-to-date information and combine it with the model’s internal knowledge. A raw local model has no tool access unless you explicitly integrate those capabilities.

When you give a model additional data at inference time (for example, a PDF, CSV, or plain text), the model can combine that supplied context with its pre-trained knowledge to produce responses. For example, feeding a 200‑page PDF and asking, “In one paragraph, what are the 10 most important points?” will produce an answer based on both the document content and the model’s internal knowledge.

If you enable tool access—databases, APIs, or a search index—the model can fetch fresh data at runtime and synthesize it with its reasoning. This pattern is commonly called retrieval-augmented generation (RAG): the model queries a retrieval layer (embedding + vector store), ingests the returned snippets, and generates responses grounded in the retrieved data. Platforms like Bedrock make it easier to grant tool access under governance and guardrails.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Gen-AI-and-LLMs/Gen-AI-and-LLMs-Introduction-Part-2/genai-solution-three-step-summaries.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=d6f5d4c135ea53246ac80f6af62495e7" alt="A dark-themed infographic titled &#x22;Solution: GenAI&#x22; showing a three-step horizontal timeline with numbered blue, orange, and green boxes. Each box summarizes a different type of AI summary: historic world events, provided content, and external data sources." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Gen-AI-and-LLMs/Gen-AI-and-LLMs-Introduction-Part-2/genai-solution-three-step-summaries.jpg" />
</Frame>

Simply running a model locally won’t give you integrated tool access or live internet connectivity by default. You must explicitly wire up retrieval systems, APIs, or browser tools for that capability.

<Callout icon="warning" color="#FF6B6B">
  Pre-training large models requires enormous compute across many GPU-hours and significant VRAM. Training at that scale (often millions of GPU hours) is expensive and beyond most organizations’ budgets—this is why many teams use vendor-hosted models or fine-tune smaller models rather than training from scratch.
</Callout>

What you can build with pre-trained LLMs

* Use cases span customer-facing chatbots, DevOps automation, summarization, code generation, content creation, and cross-silo analysis.
* A single pre-trained model can often perform many different tasks without retraining—prompts and context steer behavior.

Use cases at a glance:

| Resource Type | Typical Use Case | Example |
| - | - | - |
| Chatbots | Customer support or internal assistants | `Customer support chat on a website` |
| DevOps / SRE | Analyze logs/metrics, suggest root causes, orchestrate remediations | `Detect anomaly in logs and suggest remediation` |
| Summarization | Convert long documents into concise bullets | `Summarize a 200-page report into key takeaways` |
| Code generation | Produce code, tests, or IaC templates | `Write Jest tests for this JavaScript function` |
| Content & Presentations | Marketing copy, slide outlines, creative assets | `Create a presentation outline for product launch` |
| Cross-silo analysis | Synthesize PDFs, spreadsheets, and API data | `Extract insights from mixed-source financial reports` |

When you want to serve private or up-to-date data repeatedly and at scale, common integration patterns include:

* Passing documents or context directly in the prompt for ad-hoc queries.
* Building an index: compute embeddings, store them in a vector database, and retrieve relevant context at runtime (RAG).
* Integrating runtime tool access so models can call APIs or databases under governance.

A short table comparing hosting choices:

| Hosting Option | Pros | Cons |
| - | - | - |
| Local (on-prem or laptop) | Full control, reduced external dependency | High latency for large models, limited tool access, heavy compute needs |
| Hosted (vendor-managed) | Scalability, tool integrations, governance features | Costs, potential data residency concerns |
| Hybrid (fine-tuned small models locally + hosted for large models) | Balance of control and capabilities | Operational complexity |

Examples: prompt → output

* Prompt: “Summarize this report” + report text\
  Output: concise summary of main points.

* Prompt: “Write Jest tests for this JavaScript function” + function code\
  Output: JavaScript unit tests in Jest.

* Prompt: “Create a presentation” + topic, audience, length, tone\
  Output: a structured presentation outline or slide contents.

Example prompt and expected output formats (illustrative):

```text theme={null}
Prompt:
"Summarize this report" + [report text]

Expected Output:
- One-paragraph executive summary
- 5 key bullet points with action items
```

```text theme={null}
Prompt:
"Write Jest tests for this function" + [function code]

Expected Output:
- Multiple Jest test cases covering happy path and edge cases
- Mocking examples and assertions
```

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Gen-AI-and-LLMs/Gen-AI-and-LLMs-Introduction-Part-2/model-examples-summarize-js-test-presentation.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=c51f5ce7ab55f7d9e3bbbfcc30641dc0" alt="A dark-themed slide diagram titled &#x22;Solution: What Can We Use It for?&#x22; showing a central &#x22;Model&#x22; (brain icon) with example inputs on the left—&#x22;Summarize with report,&#x22; &#x22;Write a test for this JavaScript function,&#x22; and &#x22;Create presentation&#x22;—and corresponding outputs on the right—&#x22;Text summary,&#x22; &#x22;JavaScript jest code,&#x22; and &#x22;Presentation outline.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Gen-AI-and-LLMs/Gen-AI-and-LLMs-Introduction-Part-2/model-examples-summarize-js-test-presentation.jpg" />
</Frame>

All of the above tasks can typically be achieved with the same underlying pre-trained model: you don’t necessarily need to retrain every time you change tasks. Prompt design is the primary mechanism for steering behavior; later lessons in this course will cover prompt engineering techniques to improve reliability and specificity.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Introduction-to-Gen-AI-and-LLMs/Gen-AI-and-LLMs-Introduction-Part-2/prompt-to-large-pretrained-model-response.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=79d437a59ba0413f1857dc900bf627ca" alt="A flow diagram titled &#x22;Workflow: Model Use Flow&#x22; showing a prompt icon on the left feeding into a brain-shaped &#x22;Large Pre-Trained Neural Network&#x22; in the center. An arrow leads to a speech-bubble icon on the right labeled &#x22;Generate response.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Introduction-to-Gen-AI-and-LLMs/Gen-AI-and-LLMs-Introduction-Part-2/prompt-to-large-pretrained-model-response.jpg" />
</Frame>

<Callout icon="lightbulb" color="#1CB2FE">
  When choosing where to host or run models, consider latency, scale, tool access, governance, and cost. For private, up-to-date data at scale, pair a retrieval layer (embeddings + vector store) with controlled tool integrations to keep responses accurate and auditable.
</Callout>

Summary

* Choose hosting (local vs. hosted) based on performance, scale, and tool needs.
* Provide prompt context or integrate retrieval for private/up-to-date data.
* Use governance and monitoring when enabling models to call external tools.
* These patterns form the foundation for the Bedrock-based examples used later in this course.

Links and references

* [ChatGPT](https://chat.openai.com/)
* [AWS Bedrock (product overview)](https://aws.amazon.com/bedrock/)
* [Retrieval-Augmented Generation (RAG) — overview and patterns](https://en.wikipedia.org/wiki/Retrieval-augmented_generation)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/21e251db-dd75-4627-9910-aa15938adb6b/lesson/41c1bc05-37f5-40b6-92ef-454005098e7b" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.