Overview of generative AI and large language models covering modalities, multimodal applications, developer code assistance, creative content generation, and selecting models for enterprise use with Amazon Bedrock context
Often, we adapt the outputs we get from a model primarily by changing the prompt — without changing model code or retraining. Many users consume models that were trained by others, so prompt design and orchestration are the primary levers for getting useful output.A model’s modality determines the kinds of inputs it accepts and the outputs it produces: text, images, audio, or video. Some models generate text only; others generate images, short videos, or audio. Multimodal models can accept and produce multiple types, enabling richer applications that combine text, visuals, and sound.For example:
A text-to-image model can create an original image from a prompt like “play cricket on the moon in zero gravity.”
A video-generation model can synthesize short clips from text prompts such as “show a short clip of a person struggling to board a crowded train.”
An image-input model can analyze a photo and describe what’s happening.
A multimodal model can take an image and a text prompt to produce a story with inline illustrations and narration.
Different models excel at different tasks. Specialist models often outperform general-purpose models on narrow tasks like image-only generation or video synthesis. When selecting a model, verify which modalities it supports and whether it accepts mixed inputs (for example, text + image) or produces mixed outputs (for example, text + audio).Table: Modality overview and common use cases
Multimodal models accept and combine inputs like text, images, and audio, producing richer outputs. Choose a model that supports the modalities required for your application.
One of the most impactful applications of generative AI is developer assistance—code generation that augments engineering productivity rather than replacing people. Generative AI can accelerate many parts of software development:
Editor autocompletion: predict and insert code during authoring.
Function generation: generate a function from a specification (including tests).
Debugging assistance: analyze failing code or errors and suggest fixes.
Operational automation: produce scripts, infrastructure-as-code snippets, or CI/CD pipeline templates.
These capabilities reduce repetitive work, lower trivial errors, and enable private, secure developer assistants that access internal code and context instead of posting to public Q&A sites. As AI-powered developer tools spread, some public forums have seen decreased traffic.There are two common integration patterns for AI-assisted development:
IDE plug-ins that augment existing editors (e.g., VS Code extensions).
Standalone, AI-first development environments where the assistant is central to the workflow.
Table: Developer integration patterns
Integration pattern
Description
Example tools
IDE plug-ins
Add AI features directly into existing editors
GitHub Copilot (VS Code)
Standalone AI-first IDEs
Dedicated environments where AI guides much of the workflow
Claude Code, Cursor, Cline
What organizations can realistically expect from adopting generative AI and LLMs:
Fast document summarization and extraction to speed decision-making.
Conversational question-and-answer experiences backed by enterprise knowledge bases (for example, an HR chatbot that understands multi-year policies).
Code generation, testing, and debugging by enabling LLMs to access codebases or CI/CD contexts securely.
Creative content generation (images, audio, video) for marketing or internal materials without always relying on external agencies.
Intelligent assistants and automated workflows that aggregate data across silos (PDFs, spreadsheets, APIs) and produce recommendations or trigger actions.
LLM-backed assistants can reason over aggregated inputs and act on combined data. For example: “Do we have a customer who bought more than three items and has personalization enabled?” — the assistant can query multiple data sources and return a recommendation or trigger a workflow.
SummaryPre-trained large language models are trained on massive datasets with significant compute and can often be prompted to produce outputs across multiple modalities (text, images, audio, video). These models power a wide range of applications: faster developer workflows, conversational knowledge assistants, and creative content production. Selecting the right model depends on the modalities and task specialization required for your use case.Next stepsNow that you have a foundation in generative AI and LLMs, the next lesson introduces Amazon Bedrock — a platform for building and deploying Gen AI experiences with access to multiple foundation models and modality support.Links and references
For additional reading: search for “multimodal models”, “generative AI code generation”, and “Amazon Bedrock” in official documentation and whitepapers.