- Unit tests — fast, deterministic, no network or tokens
- Mock-based integration tests — validate orchestration without external calls
- Evaluation (end-to-end) tests — exercise the real model and tools before deployment
Setup
Prerequisites and assumptions:- Python, a virtual environment, and the OpenAI SDK are pre-installed.
- Your working directory is
/root/code. agent_components.pyalready exists and exposescheck_calendarandexecute_tool.- Create a new file named
test_agent.pyin your working directory.
test_agent.py. We’ll use Python’s unittest and unittest.mock to keep tests deterministic:
Keep tests small and focused. Unit tests should run in milliseconds and never talk to external services or consume API tokens.
1) Unit tests (fast, no network / no tokens)
Purpose: catch low-level tool bugs on every commit. Test only pure functions and interfaces—no network, no tokens. Example unit-test class:2) Mock-based integration tests (orchestration, no tokens)
Purpose: simulate the agent loop and model/tool orchestration without calling the real API. These tests verify dispatch logic — when the model indicates a tool call, your code dispatches toexecute_tool with the expected name and arguments.
Example mock-based integration test:
- Validate your orchestration code without spending tokens
- Keep tests deterministic and fast
- Isolate failures to either orchestration or tools
3) Evaluation tests (end-to-end / integration with API)
Purpose: run an end-to-end scenario against the real OpenAI client to catch prompt regressions, wiring problems, or API compatibility issues before deployment. Run these sparingly (CI gate tests or local pre-deploy checks) because they call the real service and consume tokens. Important: ensure environment variables are set (OPENAI_API_KEY and optionally OPENAI_API_BASE). If not set, skip the test early so CI doesn’t fail unexpectedly.
Eval tests hit the live API and consume tokens. Only run them in CI gates or as an explicit pre-deploy step. Ensure
OPENAI_API_KEY is set and correct to avoid misleading failures.Test Layers at a Glance
Summary
- Unit tests: fast feedback for function-level correctness.
- Mock-based integration tests: validate orchestration and dispatch without consuming tokens.
- Evaluation tests: ensure end-to-end correctness against the real model and tool definitions before deployment.
Links & References
- OpenAI API Documentation
- Python
unittest— https://docs.python.org/3/library/unittest.html unittest.mock— https://docs.python.org/3/library/unittest.mock.html
