Chapter 11

Mocking — checking how an external service was called

Stand a Mock in for a payment or email client and check how it was called. return_value, side_effect, autospec, patch, mocker — and when a fake beats a mock.

50 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

Here is a small checkout. It totals a cart, charges a card through a payment client, and emails the customer.

payments.py:

python
class PaymentDeclined(Exception):
    pass


class PaymentGateway:
    """Talks to the real payment provider over HTTPS."""

    def __init__(self, api_key):
        self.api_key = api_key

    def charge(self, amount_cents, token):
        raise RuntimeError("real network call — never in a test")

    def refund(self, charge_id):
        raise RuntimeError("real network call — never in a test")

checkout.py:

python
from payments import PaymentDeclined


def cart_total(cart):
    return sum(item["price_cents"] * item["qty"] for item in cart)


def place_order(cart, email, token, gateway, mailer):
    total = cart_total(cart)
    if total == 0:
        raise ValueError("cart is empty")

    try:
        receipt = gateway.charge(total, token=token)
    except PaymentDeclined:
        mailer.send(email, "Your payment was declined")
        return {"status": "declined"}

    mailer.send(email, f"Order confirmed: {receipt['id']}")
    return {"status": "paid", "charge_id": receipt["id"], "total_cents": total}

In real life charge takes money from a real card, so a test must never reach it. In the last chapter monkeypatch let you swap a thing out — but swapping it out is only half the job here. The return value of place_order says "paid", while the bugs that matter sit in the calls: was the card charged 2500 or 25? With which token? Once, or twice? Did the email go to the right address? To answer those, the stand-in has to remember how it was called.

You could write a small recording class for that, one per dependency. unittest.mock has already written it for you.

By the end of this chapter you can

  • Use Mock to stand in for a client and check how it was called with assert_called_once_with, call_args, call_count and mock_calls
  • Make a mock return values or raise exceptions with return_value and side_effect
  • Explain why a bare Mock is dangerous, and close the hole with spec= and create_autospec
  • Replace a name with patch, as a decorator and as a context manager, in the right place
  • Use pytest-mock's mocker.patch and mocker.spy
  • Tell a stub, a mock and a fake apart, and know when a fake is the better test

Prerequisites: monkeypatch — replacing things for one test.


Before you write the test

Before any Mock appears, write down what place_order actually promises. It has one input you control (the cart, email and token) and two side effects you cannot see in its return value — a charge and an email. A test for code like this has to check both halves.

The contract. Given a non-empty cart, place_order charges exactly the total, once, with the given token; emails a confirmation; returns "paid". If the card is declined, it charges nothing, emails an apology and returns "declined". An empty cart raises ValueError before anything leaves the building. A network failure is not swallowed.

The setup. pytest in your virtual environment, payments.py and checkout.py importable from where you run pytest, and the pytest-mock plugin for the mocker fixture (pip install pytest-mock). unittest.mock ships with Python. No API key, no network — if a test needs either, it is not a unit test.

The plan.

| Case | Input | Expected | |---|---|---| | Happy path | 2 × 1000 + 1 × 500, card accepted | one charge of 2500 with tok_visa; confirmation email; "paid" | | Declined card | same cart, gateway raises PaymentDeclined | no charge recorded; apology email; "declined" | | Empty cart | [] | ValueError("cart is empty"); gateway and mailer untouched | | Network failure | gateway raises TimeoutError | TimeoutError propagates; no email | | Contract | any order | charge called with the real signature (amount_cents, token) |

What not to test. The payment provider itself — that is its authors' job, and your test cannot reach it anyway. Python's sum. And the internal steps of place_order: whether it calls cart_total once or inlines the arithmetic is its own business. The plan lists outcomes and the one call that crosses the boundary — nothing else. Keep that in mind; it is the difference between a useful mock and a harmful one, and the end of the chapter comes back to it.


Mock — an object that writes down every call

python
from unittest.mock import Mock

gateway = Mock()

result = gateway.charge(2500, token="tok_visa")
gateway.refund("ch_1")

print(result)
print(gateway.charge.called, gateway.charge.call_count)
print(gateway.charge.call_args)
print(gateway.mock_calls)
text
<Mock name='mock.charge()' id='132252166523936'>
True 1
call(2500, token='tok_visa')
[call.charge(2500, token='tok_visa'), call.refund('ch_1')]

gateway has no charge method; nothing was defined. A Mock invents any attribute you touch, and every attribute is itself a Mock that can be called. Each call is recorded: called says whether it happened, call_count how many times, call_args with exactly what. mock_calls on the parent holds every call to every method, in order — useful when the order matters.

Reading those attributes in a bare assert works, but the assertion methods say it in one line and explain themselves when they fail:

python
from unittest.mock import Mock

gateway = Mock()
gateway.charge(2500, token="tok_visa")

gateway.charge.assert_called_once_with(2500, token="tok_visa")
print("first check passed")

gateway.charge.assert_called_once_with(2500, token="tok_amex")
text
first check passed
Traceback (most recent call last):
  ...
AssertionError: expected call not found.
Expected: charge(2500, token='tok_amex')
  Actual: charge(2500, token='tok_visa')

assert_called_once_with checks two things: there was exactly one call, and it had exactly these arguments. Its relatives are assert_called_once() (once, any arguments), assert_called_with(...) (only the last call) and assert_not_called(). (Tracebacks in this chapter are trimmed to ... where only the last line matters.)

Why "once"? Because charging a card twice is a real bug, and assert_called_with alone would not notice a second call.

return_value and side_effect — what the mock gives back

By default a call returns another Mock. Code that does receipt["id"] needs a real answer, so give it one:

python
from unittest.mock import Mock

gateway = Mock()
gateway.charge.return_value = {"id": "ch_1", "status": "succeeded"}

print(gateway.charge(2500, token="tok_visa"))
print(gateway.charge(999, token="tok_other"))
text
{'id': 'ch_1', 'status': 'succeeded'}
{'id': 'ch_1', 'status': 'succeeded'}

The same answer, whatever the arguments. When that is not enough, set side_effect. It has three forms.

An exception — gateway.charge.side_effect = PaymentDeclined("card declined") makes every call raise it. This is how you reach the unhappy paths that are hard to cause for real: a declined card, a timeout, a 500 from the server.

A list — one item per call, in order. An exception in the list is raised; anything else is returned. Good for "fails once, then succeeds":

python
from unittest.mock import Mock

gateway = Mock()
gateway.charge.side_effect = [
    TimeoutError("gateway timed out"),
    {"id": "ch_2", "status": "succeeded"},
]

try:
    gateway.charge(2500, token="tok_visa")
except TimeoutError as exc:
    print("first call:", exc)

print("second call:", gateway.charge(2500, token="tok_visa"))
gateway.charge(2500, token="tok_visa")
text
first call: gateway timed out
second call: {'id': 'ch_2', 'status': 'succeeded'}
Traceback (most recent call last):
  ...
StopIteration

A third call found the list empty. StopIteration from a mock means the code called it more often than you planned for — which is itself worth knowing.

A function — called with the same arguments as the mock, and whatever it returns is the answer. gateway.charge.side_effect = lambda amount_cents, token: {"status": "declined"} if amount_cents > 100_000 else {"id": "ch_1"} gives an answer that depends on the input — and the calls are still recorded and counted.

MagicMock — when the code uses len, with or for

A plain Mock does not support Python's special methods: len(Mock()) raises TypeError: object of type 'Mock' has no len(), and with Mock(): fails too. MagicMock supports them with sensible defaults — len(MagicMock()) is 0, it works in a with block, and m.__len__.return_value = 3 changes the answer.

Reach for MagicMock when the stand-in is used as a context manager, a container or an iterator — a database connection in a with block, say. patch and mocker.patch, below, create MagicMocks by default.

The danger: a Mock says yes to everything

That obliging nature has a price. Misspell a method, and nothing complains:

python
from unittest.mock import Mock

gateway = Mock()
gateway.chrage(2500, token="tok_visa")   # typo in the method name

print("no error — the typo was recorded as a new method")
print(gateway.mock_calls)
text
no error — the typo was recorded as a new method
[call.chrage(2500, token='tok_visa')]

If that typo were in checkout.py, a test with a bare Mock could pass, and the first real customer would get AttributeError. Wrong arguments slip through the same way. Here is the wrong way and the right way side by side — the code under test forgets the token:

python
from unittest.mock import Mock, create_autospec

from payments import PaymentGateway


def charge_order(gateway, total, token):
    return gateway.charge(total)   # bug: the token is never passed


def test_with_a_plain_mock():
    gateway = Mock()
    charge_order(gateway, 2500, "tok_visa")
    gateway.charge.assert_called_once()


def test_with_an_autospec():
    gateway = create_autospec(PaymentGateway, instance=True)
    charge_order(gateway, 2500, "tok_visa")
    gateway.charge.assert_called_once()
text
$ pytest -q --tb=line test_typo.py
.F                                                                       [100%]
=================================== FAILURES ===================================
E   TypeError: missing a required argument: 'token'
/usr/lib/python3.12/inspect.py:3157: TypeError: missing a required argument: 'token'
=========================== short test summary info ============================
FAILED test_typo.py::test_with_an_autospec - TypeError: missing a required ar...
1 failed, 1 passed in 0.07s

The plain Mock let a broken call through — a green test protecting nothing. create_autospec builds a mock from the real PaymentGateway: it has only that class's attributes, and every method checks its arguments against the real signature. instance=True means "an instance of the class", not the class itself.

Three levels of strictness:

  • Mock() — any attribute, any arguments.
  • Mock(spec=PaymentGateway) — only real attribute names (gateway.chrage raises AttributeError: Mock object has no attribute 'chrage'. Did you mean: 'charge'?), but arguments are not checked.
  • create_autospec(PaymentGateway, instance=True), or autospec=True on patch — real names and real signatures.

Why does this matter beyond typos? Because the real client changes. When a library upgrade renames a method or adds a required argument, an autospec mock fails the same day; a bare Mock keeps agreeing with code that no longer works. For anything standing in for a real class, use the strictest level.

Python does guard one kind of typo itself: since 3.12, a misspelled assertion such as assert_called_once_wiht raises AttributeError: 'assert_called_once_wiht' is not a valid assertion. instead of silently passing. That guard only knows the mock's own assertion names. It cannot know that your gateway has charge and not chrage — only a spec can.

patch — replacing a name the code looks up itself

place_order receives its gateway as an argument, so passing a mock in is easy. Often the code fetches its dependency itself:

python
# emailer.py
def send_email(to, subject):
    raise RuntimeError("real SMTP call — never in a test")


# signup.py
from emailer import send_email


def register(email):
    send_email(email, "Welcome aboard")
    return {"email": email, "active": True}

unittest.mock.patch replaces a name with a MagicMock for the length of a test, then puts the original back. It works as a decorator — the mock arrives as an argument — or as a context manager:

python
from unittest.mock import patch

from signup import register


@patch("signup.send_email")
def test_register_as_decorator(send_email):
    user = register("asha@example.com")

    assert user["active"] is True
    send_email.assert_called_once_with("asha@example.com", "Welcome aboard")


def test_register_as_context_manager():
    with patch("signup.send_email") as send_email:
        register("ravi@example.com")

    send_email.assert_called_once_with("ravi@example.com", "Welcome aboard")
text
$ pytest -q test_signup.py
..                                                                       [100%]
2 passed in 0.08s

The target is "signup.send_email", not "emailer.send_email". That is the rule from the last chapter: patch where the name is looked up, not where it is defined. from emailer import send_email copied the name into signup, so register looks there. The wrong way — patching the original module — leaves the copy untouched and the real function runs:

python
from unittest.mock import patch

from signup import register


def test_patched_in_the_wrong_place():
    with patch("emailer.send_email") as send_email:
        register("asha@example.com")

    send_email.assert_called_once()
text
$ pytest -q --tb=line test_signup_wrong.py
F                                                                        [100%]
=================================== FAILURES ===================================
E   RuntimeError: real SMTP call — never in a test
/home/you/shop/emailer.py:2: RuntimeError: real SMTP call — never in a test
=========================== short test summary info ============================
FAILED test_signup_wrong.py::test_patched_in_the_wrong_place - RuntimeError: ...
1 failed in 0.08s

The error comes from emailer.py — the real one. Here the stand-in raises loudly; a real email client would simply have sent the email.

patch takes autospec=True too, for the same reason as before: patch("signup.send_email", autospec=True) makes a mock that only accepts (to, subject).

mocker — the pytest-mock fixture

The pytest-mock plugin wraps all of this in a fixture called mocker. mocker.patch takes the same arguments as patch, needs no decorator or with, and is undone automatically when the test ends — just like monkeypatch. mocker.spy is the odd one out: it does not replace anything. The real method runs; the spy only watches.

python
from checkout import place_order
from fakes import FakeGateway, FakeMailer
from signup import register


def test_register_with_mocker(mocker):
    send_email = mocker.patch("signup.send_email", autospec=True)

    register("asha@example.com")

    send_email.assert_called_once_with("asha@example.com", "Welcome aboard")


def test_spy_watches_a_real_object(mocker):
    gateway = FakeGateway()
    spy = mocker.spy(gateway, "charge")

    place_order([{"price_cents": 700, "qty": 1}], "a@example.com", "tok", gateway, FakeMailer())

    spy.assert_called_once_with(700, token="tok")
    assert spy.spy_return == {"id": "ch_1"}
    assert gateway.charges == [("ch_1", 700, "tok")]
text
$ pytest -q test_signup_mocker.py
..                                                                       [100%]
2 passed in 0.08s

The spy recorded the call like a mock, kept the real result in spy_return, and the real charge still did its work — gateway.charges was filled. (FakeGateway is defined in the next section.) mocker also offers mocker.Mock, mocker.MagicMock and mocker.create_autospec, so a test can get everything from one fixture.

Stub, mock, fake — and when not to mock

Three words get mixed up, and the difference decides what your test can catch:

  • A stub only answers. gateway.charge.return_value = {"id": "ch_1"}, and you never check how it was called. It feeds the code under test.
  • A mock is a stub you interrogate afterwards: assert_called_once_with(...). The test is about the call itself.
  • A fake is a real, working, small implementation — an in-memory gateway that keeps a list instead of talking to a bank.

A test made of mocks checks how the code did its work. Overdo it and the test stops checking what it did. The wrong way first:

python
import checkout
from checkout import place_order


def test_place_order_over_mocked(mocker):
    mocker.patch("checkout.cart_total", return_value=2500)
    gateway = mocker.Mock()
    gateway.charge.return_value = {"id": "ch_1"}
    mailer = mocker.Mock()

    place_order([{"price_cents": 1, "qty": 1}], "a@example.com", "tok", gateway, mailer)

    checkout.cart_total.assert_called_once()
    gateway.charge.assert_called_once_with(2500, token="tok")
    mailer.send.assert_called_once()
text
$ pytest -q test_overmock.py
.                                                                        [100%]
1 passed in 0.07s

It passes. The cart costs one cent, and the test "confirms" a charge of 2500 — because the test itself replaced the arithmetic. Break cart_total and this test still passes. Rename cart_total without changing any behaviour and it fails. That is the over-mocking smell: a test that mirrors the implementation line by line, tests its own mocks, and breaks on every refactor.

The right way follows the plan from the start of the chapter: mock at the edge of your system, and only there. The payment provider and the SMTP server are the edge. cart_total is your own code — let it run. And for collaborators with a small interface, a fake usually beats a mock:

python
from payments import PaymentDeclined


class FakeGateway:
    """An in-memory stand-in for PaymentGateway: same methods, no network."""

    def __init__(self, declines=False):
        self.declines = declines
        self.charges = []

    def charge(self, amount_cents, token):
        if self.declines:
            raise PaymentDeclined("card declined")
        charge_id = f"ch_{len(self.charges) + 1}"
        self.charges.append((charge_id, amount_cents, token))
        return {"id": charge_id}

    def refund(self, charge_id):
        self.charges = [c for c in self.charges if c[0] != charge_id]


class FakeMailer:
    def __init__(self):
        self.sent = []

    def send(self, to, subject):
        self.sent.append((to, subject))

A test with fakes asserts on state — what ended up in the list — rather than on a sequence of calls. It reads like a description of the behaviour, and it survives a refactor as long as the behaviour holds. You write the fake once, in fakes.py or a conftest.py fixture, and every test reuses it. Its one weakness is that a fake can drift from the real client — which is why the complete example keeps one autospec test for the contract.


A complete example

test_checkout.py implements the test plan row by row, with each tool where it belongs — fakes for behaviour, an autospec mock for the contract, a side_effect for the failure:

python
from unittest.mock import create_autospec

import pytest

from checkout import place_order
from fakes import FakeGateway, FakeMailer
from payments import PaymentGateway

CART = [{"price_cents": 1000, "qty": 2}, {"price_cents": 500, "qty": 1}]
EMAIL = "asha@example.com"


# Behaviour, checked with fakes: what happened, not which calls were made.
def test_paid_order_charges_once_and_sends_a_receipt():
    gateway, mailer = FakeGateway(), FakeMailer()

    result = place_order(CART, EMAIL, "tok_visa", gateway, mailer)

    assert result == {"status": "paid", "charge_id": "ch_1", "total_cents": 2500}
    assert gateway.charges == [("ch_1", 2500, "tok_visa")]
    assert mailer.sent == [(EMAIL, "Order confirmed: ch_1")]


def test_declined_card_sends_an_apology_and_charges_nothing():
    gateway, mailer = FakeGateway(declines=True), FakeMailer()

    result = place_order(CART, EMAIL, "tok_visa", gateway, mailer)

    assert result == {"status": "declined"}
    assert gateway.charges == []
    assert mailer.sent == [(EMAIL, "Your payment was declined")]


def test_empty_cart_never_reaches_the_gateway():
    gateway, mailer = FakeGateway(), FakeMailer()

    with pytest.raises(ValueError, match="cart is empty"):
        place_order([], EMAIL, "tok_visa", gateway, mailer)

    assert gateway.charges == []
    assert mailer.sent == []


# The contract with the real client, checked with an autospec mock.
def test_gateway_is_called_with_the_real_signature():
    gateway = create_autospec(PaymentGateway, instance=True)
    gateway.charge.return_value = {"id": "ch_9"}

    place_order(CART, EMAIL, "tok_visa", gateway, FakeMailer())

    gateway.charge.assert_called_once_with(2500, token="tok_visa")
    gateway.refund.assert_not_called()


def test_a_timeout_propagates_and_sends_nothing():
    gateway = create_autospec(PaymentGateway, instance=True)
    gateway.charge.side_effect = TimeoutError("gateway timed out")
    mailer = FakeMailer()

    with pytest.raises(TimeoutError):
        place_order(CART, EMAIL, "tok_visa", gateway, mailer)

    assert mailer.sent == []
text
$ pytest -v test_checkout.py
============================= test session starts ==============================
collected 5 items

test_checkout.py::test_paid_order_charges_once_and_sends_a_receipt PASSED [ 20%]
test_checkout.py::test_declined_card_sends_an_apology_and_charges_nothing PASSED [ 40%]
test_checkout.py::test_empty_cart_never_reaches_the_gateway PASSED       [ 60%]
test_checkout.py::test_gateway_is_called_with_the_real_signature PASSED  [ 80%]
test_checkout.py::test_a_timeout_propagates_and_sends_nothing PASSED     [100%]

============================== 5 passed in 0.08s ===============================

Two things are worth noticing.

First, the three behaviour tests contain no assert_called... at all. They check the outcome — the result, the charges recorded, the emails sent — so place_order could be rewritten from scratch and they would still hold.

Second, the one test that checks a call is the contract with PaymentGateway, and it uses create_autospec. If someone renames charge or adds a required argument, that test fails instead of quietly agreeing. Five rows in the plan, five tests, nothing that tests the inside of place_order.


When it breaks

AssertionError: expected call not found. The mock was called, but not with these arguments. The message prints Expected: and Actual: one above the other — compare them character by character. A positional argument against a keyword argument (charge(2500, "tok") versus charge(2500, token="tok")) counts as a different call.

AssertionError: Expected 'charge' to be called once. Called 2 times. The code called it more than once, and the Calls: line lists every call. Often a retry loop, or one mock reused across two actions in a test.

AttributeError: Mock object has no attribute 'chrage'. Did you mean: 'charge'? A spec doing its job: the code uses a name the real class does not have. Fix the code, not the test.

TypeError: missing a required argument: 'token' An autospec mock caught a call that does not match the real signature. The code would fail the same way against the real client.

AttributeError: <module 'signup' from '...'> does not have the attribute 'send_mail' The patch target names something that does not exist — a typo, or the wrong module. patch refuses to invent new attributes, which is a good thing.

The patch "did nothing" and the real service was called You patched where the name is defined, not where it is looked up. If the module did from emailer import send_email, patch "signup.send_email".

StopIteration from a mock call side_effect was a list, and the code made more calls than the list had items.

A result holds <MagicMock name='mock.charge().__getitem__()' ...> instead of data You never set return_value, and a MagicMock happily answered receipt["id"] with another mock. With a plain Mock the same code fails with TypeError: 'Mock' object is not subscriptable. Set return_value to the shape the real client returns.

fixture 'mocker' not found pytest-mock is not installed in the environment pytest runs from. pip install pytest-mock.