Model selection: match size to task
- Small models: fast, low-cost, ideal for classification, sentiment analysis, short summaries, and deterministic transformations.
- Large models: higher latency and cost, best for complex reasoning, multi-step tasks, and situations requiring broad context or advanced synthesis.
- When unsure, benchmark both representative smaller and larger models on your workload and measure accuracy, latency, and cost.
Control output size and format
Explicitly constraining output length and format reduces token usage and prevents verbose, costly responses. When invoking a model, set a hard token limit and provide a strict output template.
maxTokenCount is set to 100. Coupled with an instruction that the output must be exactly three bullet points, this prevents rambling and reduces inference cost.
Prompt design: reduce input token cost
The input context also contributes to cost. Avoid repeatedly sending large or irrelevant documents. In retrieval-augmented generation (RAG), only pass the most relevant chunks into the prompt to keep both input and output token usage low.
When you require a strict structured output (for example JSON), include the exact schema or a template in the prompt. This reduces malformed outputs and prevents extraneous text that consumes tokens. For example, supply a JSON schema and instruct: “Return only valid JSON matching this schema.”
Knowledge base (RAG) billing and chunk-size trade-offs
When you use managed RAG services or Bedrock Knowledge Bases, costs arrive from multiple places. Monitor and optimize across the whole pipeline:
Chunk size directly affects both ingestion and runtime costs:
- Larger chunks → fewer embeddings (lower ingestion cost) but coarser retrieval. Retrieved chunks may include irrelevant information, increasing augmented prompt size and runtime cost.
- Smaller chunks → more embeddings (higher ingestion cost) but more precise retrieval. Precise retrieval injects only tightly relevant content into the prompt and can lower runtime cost.

Evaluate trade-offs carefully: choose chunk sizes and embedding models that balance ingestion vs. runtime costs for your workload. Continuously monitor costs across ingestion, retrieval, and inference and adjust parameters based on observed accuracy and spend.
Quick reference
Summary: favor smaller models for simple tasks, reserve large models for complex reasoning, constrain outputs (token limits and strict formats), design precise prompts, and use RAG intelligently—retrieve only necessary chunks to reduce token usage and cost.