Skip to main content
Where you run or consume a model is a design choice. Many people are familiar with hosted services such as ChatGPT that expose web UIs backed by GPT-family models, but you can also run models locally (on a laptop or a private server) or consume managed, hosted offerings. Running very large models locally is possible, but typically slow and GPU-intensive unless you have substantial specialized hardware. Given this course’s focus, examples will use Bedrock to demonstrate hosted patterns and governance. Key distinctions to keep in mind:
  • When a generative AI model answers a question about a historical event correctly, it’s using knowledge learned during pre-training. Pre-training ingests large corpora (public text, licensed data, and other sources), and the model’s responses reflect patterns from that data.
  • A local model whose training data stops before a particular event or topic cannot reliably answer questions that postdate its training.
  • Hosted services or model configurations can be augmented with external tools—web browsing, search, or plugins—so the system can fetch up-to-date information and combine it with the model’s internal knowledge. A raw local model has no tool access unless you explicitly integrate those capabilities.
When you give a model additional data at inference time (for example, a PDF, CSV, or plain text), the model can combine that supplied context with its pre-trained knowledge to produce responses. For example, feeding a 200‑page PDF and asking, “In one paragraph, what are the 10 most important points?” will produce an answer based on both the document content and the model’s internal knowledge. If you enable tool access—databases, APIs, or a search index—the model can fetch fresh data at runtime and synthesize it with its reasoning. This pattern is commonly called retrieval-augmented generation (RAG): the model queries a retrieval layer (embedding + vector store), ingests the returned snippets, and generates responses grounded in the retrieved data. Platforms like Bedrock make it easier to grant tool access under governance and guardrails.
A dark-themed infographic titled "Solution: GenAI" showing a three-step horizontal timeline with numbered blue, orange, and green boxes. Each box summarizes a different type of AI summary: historic world events, provided content, and external data sources.
Simply running a model locally won’t give you integrated tool access or live internet connectivity by default. You must explicitly wire up retrieval systems, APIs, or browser tools for that capability.
Pre-training large models requires enormous compute across many GPU-hours and significant VRAM. Training at that scale (often millions of GPU hours) is expensive and beyond most organizations’ budgets—this is why many teams use vendor-hosted models or fine-tune smaller models rather than training from scratch.
What you can build with pre-trained LLMs
  • Use cases span customer-facing chatbots, DevOps automation, summarization, code generation, content creation, and cross-silo analysis.
  • A single pre-trained model can often perform many different tasks without retraining—prompts and context steer behavior.
Use cases at a glance: When you want to serve private or up-to-date data repeatedly and at scale, common integration patterns include:
  • Passing documents or context directly in the prompt for ad-hoc queries.
  • Building an index: compute embeddings, store them in a vector database, and retrieve relevant context at runtime (RAG).
  • Integrating runtime tool access so models can call APIs or databases under governance.
A short table comparing hosting choices: Examples: prompt → output
  • Prompt: “Summarize this report” + report text
    Output: concise summary of main points.
  • Prompt: “Write Jest tests for this JavaScript function” + function code
    Output: JavaScript unit tests in Jest.
  • Prompt: “Create a presentation” + topic, audience, length, tone
    Output: a structured presentation outline or slide contents.
Example prompt and expected output formats (illustrative):
A dark-themed slide diagram titled "Solution: What Can We Use It for?" showing a central "Model" (brain icon) with example inputs on the left—"Summarize with report," "Write a test for this JavaScript function," and "Create presentation"—and corresponding outputs on the right—"Text summary," "JavaScript jest code," and "Presentation outline."
All of the above tasks can typically be achieved with the same underlying pre-trained model: you don’t necessarily need to retrain every time you change tasks. Prompt design is the primary mechanism for steering behavior; later lessons in this course will cover prompt engineering techniques to improve reliability and specificity.
A flow diagram titled "Workflow: Model Use Flow" showing a prompt icon on the left feeding into a brain-shaped "Large Pre-Trained Neural Network" in the center. An arrow leads to a speech-bubble icon on the right labeled "Generate response."
When choosing where to host or run models, consider latency, scale, tool access, governance, and cost. For private, up-to-date data at scale, pair a retrieval layer (embeddings + vector store) with controlled tool integrations to keep responses accurate and auditable.
Summary
  • Choose hosting (local vs. hosted) based on performance, scale, and tool needs.
  • Provide prompt context or integrate retrieval for private/up-to-date data.
  • Use governance and monitoring when enabling models to call external tools.
  • These patterns form the foundation for the Bedrock-based examples used later in this course.
Links and references

Watch Video