> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Lab Walkthrough Test Your Agent

> Guide to a three-layer testing strategy for AI agents covering unit tests, mock-based integration tests, and end-to-end evaluation with setup, examples, and CI considerations.

Time to write tests for your agent. This lesson walks through a three-layer test strategy that balances speed, determinism, and real-world coverage:

* Unit tests — fast, deterministic, no network or tokens
* Mock-based integration tests — validate orchestration without external calls
* Evaluation (end-to-end) tests — exercise the real model and tools before deployment

Each layer targets different classes of bugs. Together they increase confidence before deploying your agent.

***

## Setup

Prerequisites and assumptions:

* Python, a virtual environment, and the OpenAI SDK are pre-installed.
* Your working directory is `/root/code`.
* `agent_components.py` already exists and exposes `check_calendar` and `execute_tool`.
* Create a new file named `test_agent.py` in your working directory.

Create the imports at the top of `test_agent.py`. We'll use Python's `unittest` and `unittest.mock` to keep tests deterministic:

```python theme={null}
import os
import json
import unittest
from unittest.mock import MagicMock, patch

import agent_components
from agent_components import check_calendar, execute_tool
```

<Callout icon="lightbulb" color="#1CB2FE">
  Keep tests small and focused. Unit tests should run in milliseconds and never talk to external services or consume API tokens.
</Callout>

***

## 1) Unit tests (fast, no network / no tokens)

Purpose: catch low-level tool bugs on every commit. Test only pure functions and interfaces—no network, no tokens.

Example unit-test class:

```python theme={null}
class TestUnit(unittest.TestCase):
    def test_imports_and_callables(self):
        # Ensure the functions we've been using actually exist and are callable.
        self.assertTrue(callable(check_calendar))
        self.assertTrue(callable(execute_tool))
```

Run unit tests quickly:

```bash theme={null}
python3 -m unittest
```

Unit tests should finish in milliseconds and provide fast feedback in CI on every commit.

***

## 2) Mock-based integration tests (orchestration, no tokens)

Purpose: simulate the agent loop and model/tool orchestration without calling the real API. These tests verify dispatch logic — when the model indicates a tool call, your code dispatches to `execute_tool` with the expected name and arguments.

Example mock-based integration test:

```python theme={null}
class TestAgentLoop(unittest.TestCase):
    def test_tool_dispatch_on_tool_call(self):
        # Patch the execute_tool function on the agent_components module so we can assert it was called.
        with patch.object(agent_components, "execute_tool") as mock_execute_tool:
            fake_tool_call = MagicMock()
            fake_tool_call.function.name = "check_calendar"
            # The model often returns tool arguments as a JSON string.
            fake_tool_call.function.arguments = json.dumps({})

            # Simulate what your orchestration logic would extract from the model response.
            name = fake_tool_call.function.name
            args = json.loads(fake_tool_call.function.arguments or "{}")

            # Call the dispatch function (using the module reference to ensure we hit the patched function).
            agent_components.execute_tool(name, args)

            # Assert dispatch happened with the expected tool name and arguments.
            mock_execute_tool.assert_called_with("check_calendar", {})
```

Why mocks?

* Validate your orchestration code without spending tokens
* Keep tests deterministic and fast
* Isolate failures to either orchestration or tools

***

## 3) Evaluation tests (end-to-end / integration with API)

Purpose: run an end-to-end scenario against the real OpenAI client to catch prompt regressions, wiring problems, or API compatibility issues before deployment. Run these sparingly (CI gate tests or local pre-deploy checks) because they call the real service and consume tokens.

Important: ensure environment variables are set (`OPENAI_API_KEY` and optionally `OPENAI_API_BASE`). If not set, skip the test early so CI doesn't fail unexpectedly.

```python theme={null}
class TestAgentEval(unittest.TestCase):
    def test_agent_responds_to_calendar_query(self):
        from openai import OpenAI

        # Expect an environment variable with the API key. Fail early with a helpful message if missing.
        api_key = os.environ.get("OPENAI_API_KEY")
        if not api_key:
            self.skipTest("OPENAI_API_KEY not set; skipping eval test")

        base = os.environ.get("OPENAI_API_BASE")

        # Create the client (include base_url only if provided).
        client_kwargs = {"api_key": api_key}
        if base:
            client_kwargs["base_url"] = base

        client = OpenAI(**client_kwargs)

        # Send a calendar-related user message and include the calendar tool definition that your agent uses.
        response = client.chat.completions.create(
            model="openai/gpt-4.1-mini",
            messages=[{"role": "user", "content": "What's on my calendar?"}],
            tools=[agent_components.CHECK_CALENDAR_TOOL],
        )

        choice = response.choices[0]

        # The model may indicate it wants to call a tool (finish_reason can vary across SDKs/versions),
        # or it may reply directly. Accept multiple common finish_reason values used for tool/function calls.
        made_tool_call = getattr(choice, "finish_reason", None) in ("tool_calls", "tool_call", "function_call")

        # Extract content safely from the returned choice. Some SDK versions expose .message, others provide dicts.
        content = ""
        if hasattr(choice, "message"):
            message = choice.message
            raw = getattr(message, "content", "") or ""
            if isinstance(raw, str):
                content = raw
            elif isinstance(raw, dict):
                # Common patterns: plain text, or { "parts": [...] } structures.
                content = raw.get("text") or " ".join(raw.get("parts", [])) or ""
            else:
                content = str(raw)
        else:
            # Fallback for dict-like choices
            content = (getattr(choice, "get", lambda k, d=None: d)("message") or {}).get("content", "") or ""

        content_lower = content.lower()
        has_keyword = any(k in content_lower for k in ["calendar", "meeting", "standup"])

        self.assertTrue(made_tool_call or has_keyword)
```

<Callout icon="warning" color="#FF6B6B">
  Eval tests hit the live API and consume tokens. Only run them in CI gates or as an explicit pre-deploy step. Ensure `OPENAI_API_KEY` is set and correct to avoid misleading failures.
</Callout>

Run the full test suite (verbose):

```bash theme={null}
python3 -m unittest -v
```

***

## Test Layers at a Glance

| Test Layer | Purpose | Characteristics |
| - | - | - |
| Unit tests | Catch tool-level bugs quickly | Fast, deterministic, no network |
| Mock-based integration tests | Validate orchestration and dispatch logic | Uses `unittest.mock`, no tokens |
| Evaluation tests | End-to-end prompt and tool wiring validation | Calls real API; requires `OPENAI_API_KEY` and optionally `OPENAI_API_BASE` |

***

## Summary

* Unit tests: fast feedback for function-level correctness.
* Mock-based integration tests: validate orchestration and dispatch without consuming tokens.
* Evaluation tests: ensure end-to-end correctness against the real model and tool definitions before deployment.

These three layers combined provide robust coverage for agent correctness and regression detection.

## Links & References

* [OpenAI API Documentation](https://platform.openai.com/docs)
* Python `unittest` — [https://docs.python.org/3/library/unittest.html](https://docs.python.org/3/library/unittest.html)
* `unittest.mock` — [https://docs.python.org/3/library/unittest.mock.html](https://docs.python.org/3/library/unittest.mock.html)

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Lab-Walkthrough-Test-Your-Agent/three-layers-one-file-tests.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=95d0169d9941a3308fbc373f5db78a7f" alt="A retro-style graphic with a green check and the headline &#x22;THREE LAYERS, ONE FILE.&#x22; Below it are three colored boxes labeled &#x22;UNIT TESTS&#x22; (catch tool bugs on every commit), &#x22;MOCK TESTS&#x22; (catch orchestration bugs, no tokens), and &#x22;EVAL TESTS&#x22; (catch prompt regressions pre-deploy)." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Lab-Walkthrough-Test-Your-Agent/three-layers-one-file-tests.jpg" />
</Frame>

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/b8b38b25-c4eb-425f-a093-cec426365977/lesson/cfdaa4c6-95ea-4099-a70f-ae35ab25ce2d" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.