- When a generative AI model answers a question about a historical event correctly, it’s using knowledge learned during pre-training. Pre-training ingests large corpora (public text, licensed data, and other sources), and the model’s responses reflect patterns from that data.
- A local model whose training data stops before a particular event or topic cannot reliably answer questions that postdate its training.
- Hosted services or model configurations can be augmented with external tools—web browsing, search, or plugins—so the system can fetch up-to-date information and combine it with the model’s internal knowledge. A raw local model has no tool access unless you explicitly integrate those capabilities.

Pre-training large models requires enormous compute across many GPU-hours and significant VRAM. Training at that scale (often millions of GPU hours) is expensive and beyond most organizations’ budgets—this is why many teams use vendor-hosted models or fine-tune smaller models rather than training from scratch.
- Use cases span customer-facing chatbots, DevOps automation, summarization, code generation, content creation, and cross-silo analysis.
- A single pre-trained model can often perform many different tasks without retraining—prompts and context steer behavior.
When you want to serve private or up-to-date data repeatedly and at scale, common integration patterns include:
- Passing documents or context directly in the prompt for ad-hoc queries.
- Building an index: compute embeddings, store them in a vector database, and retrieve relevant context at runtime (RAG).
- Integrating runtime tool access so models can call APIs or databases under governance.
Examples: prompt → output
-
Prompt: “Summarize this report” + report text
Output: concise summary of main points. -
Prompt: “Write Jest tests for this JavaScript function” + function code
Output: JavaScript unit tests in Jest. -
Prompt: “Create a presentation” + topic, audience, length, tone
Output: a structured presentation outline or slide contents.


When choosing where to host or run models, consider latency, scale, tool access, governance, and cost. For private, up-to-date data at scale, pair a retrieval layer (embeddings + vector store) with controlled tool integrations to keep responses accurate and auditable.
- Choose hosting (local vs. hosted) based on performance, scale, and tool needs.
- Provide prompt context or integrate retrieval for private/up-to-date data.
- Use governance and monitoring when enabling models to call external tools.
- These patterns form the foundation for the Bedrock-based examples used later in this course.