Skip to main content
Time to write tests for your agent. This lesson walks through a three-layer test strategy that balances speed, determinism, and real-world coverage:
  • Unit tests — fast, deterministic, no network or tokens
  • Mock-based integration tests — validate orchestration without external calls
  • Evaluation (end-to-end) tests — exercise the real model and tools before deployment
Each layer targets different classes of bugs. Together they increase confidence before deploying your agent.

Setup

Prerequisites and assumptions:
  • Python, a virtual environment, and the OpenAI SDK are pre-installed.
  • Your working directory is /root/code.
  • agent_components.py already exists and exposes check_calendar and execute_tool.
  • Create a new file named test_agent.py in your working directory.
Create the imports at the top of test_agent.py. We’ll use Python’s unittest and unittest.mock to keep tests deterministic:
Keep tests small and focused. Unit tests should run in milliseconds and never talk to external services or consume API tokens.

1) Unit tests (fast, no network / no tokens)

Purpose: catch low-level tool bugs on every commit. Test only pure functions and interfaces—no network, no tokens. Example unit-test class:
Run unit tests quickly:
Unit tests should finish in milliseconds and provide fast feedback in CI on every commit.

2) Mock-based integration tests (orchestration, no tokens)

Purpose: simulate the agent loop and model/tool orchestration without calling the real API. These tests verify dispatch logic — when the model indicates a tool call, your code dispatches to execute_tool with the expected name and arguments. Example mock-based integration test:
Why mocks?
  • Validate your orchestration code without spending tokens
  • Keep tests deterministic and fast
  • Isolate failures to either orchestration or tools

3) Evaluation tests (end-to-end / integration with API)

Purpose: run an end-to-end scenario against the real OpenAI client to catch prompt regressions, wiring problems, or API compatibility issues before deployment. Run these sparingly (CI gate tests or local pre-deploy checks) because they call the real service and consume tokens. Important: ensure environment variables are set (OPENAI_API_KEY and optionally OPENAI_API_BASE). If not set, skip the test early so CI doesn’t fail unexpectedly.
Eval tests hit the live API and consume tokens. Only run them in CI gates or as an explicit pre-deploy step. Ensure OPENAI_API_KEY is set and correct to avoid misleading failures.
Run the full test suite (verbose):

Test Layers at a Glance


Summary

  • Unit tests: fast feedback for function-level correctness.
  • Mock-based integration tests: validate orchestration and dispatch without consuming tokens.
  • Evaluation tests: ensure end-to-end correctness against the real model and tool definitions before deployment.
These three layers combined provide robust coverage for agent correctness and regression detection.
A retro-style graphic with a green check and the headline "THREE LAYERS, ONE FILE." Below it are three colored boxes labeled "UNIT TESTS" (catch tool bugs on every commit), "MOCK TESTS" (catch orchestration bugs, no tokens), and "EVAL TESTS" (catch prompt regressions pre-deploy).

Watch Video