Chapter 28

Testing with pytest

From `assert` to pytest, reading a failure report, `pytest.raises`, `parametrize`, `tmp_path` and fixtures — and which shape of code is easy to test.

37 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

We have been "testing" code in exactly one way so far: run it, read the output, and decide whether it looked right.

That works once. The trouble begins the second time.

Chapter twenty-five's program had TAX_RATE at 0.15. Suppose it becomes 0.18. The program runs, every number prints, nothing raises — and you have no way of knowing whether anything changed, unless you happen to remember the previous output.

And what has to be remembered eventually is not.

python
def line_total(price, quantity):
    return round(price * quantity * 1.15, 2)


assert line_total(15.0, 1) == 17.25
print("ok")
text
ok

assert is a simple sentence: "this is supposed to be true". When it is, nothing happens; when it is not, the program stops.

python
def line_total(price, quantity):
    return round(price * quantity * 1.15, 2)


assert line_total(15.0, 3) == 51.0
print("ok")
text
AssertionError

That is the foundation. This chapter is a tool built on it — pytest — which runs hundreds of such claims and, when one fails, says why.

By the end of this chapter you can

  • Write a test file and run pytest
  • Read a failure report and find the problem from it
  • Test errors with pytest.raises
  • Run one test over many cases with @pytest.mark.parametrize
  • Test file-handling code with tmp_path
  • Say which code is easy to test, and why

Prerequisites: Dataclasses and type hints.


A first test

Two rules, and that is all: the file name starts with test_, and so does the function name.

pricing.py:

python
TAX_RATE = 0.15


def line_total(price: float, quantity: int) -> float:
    if quantity < 1:
        raise ValueError(f"quantity must be at least 1: {quantity}")
    return round(price * quantity * (1 + TAX_RATE), 2)

test_pricing.py:

python
from pricing import line_total


def test_one_item():
    assert line_total(15.0, 1) == 17.25


def test_three_items():
    assert line_total(15.0, 3) == 51.75

Then pytest -q:

text
..                                                                       [100%]
2 passed in 0.01s

Two dots, two tests. pytest finds the files, finds the functions and runs them — nothing has to be registered.

The failure report is the point

Put a deliberately wrong number in one test:

text
F                                                                        [100%]
================================== FAILURES ===================================
______________________________ test_three_items _______________________________

    def test_three_items():
>       assert line_total(15.0, 3) == 51.0
E       assert 51.75 == 51.0
E        +  where 51.75 = line_total(15.0, 3)

test_pricing.py:5: AssertionError
=========================== short test summary info ===========================
FAILED test_pricing.py::test_three_items - assert 51.75 == 51.0
1 failed in 0.01s

A plain assert gave us only AssertionError — one word, carrying nothing. Here we get:

  • Which test (test_three_items) and which line (test_pricing.py:5)
  • The exact claim that broke, marked with >
  • assert 51.75 == 51.0 — what was actually produced, against what was expected
  • where 51.75 = line_total(15.0, 3) — where that number came from

This is called assertion introspection, and it is the main reason for pytest. Many other languages need a separate assertEqual(a, b) for this; here an ordinary == is enough.

Errors are behaviour too

Chapter twenty-four said a function should raise on bad input. That behaviour is a promise as well, and so it can be tested:

python
import pytest

from pricing import line_total


def test_zero_is_rejected():
    with pytest.raises(ValueError):
        line_total(15.0, 0)


def test_message_names_the_value():
    with pytest.raises(ValueError, match="at least 1: -3"):
        line_total(15.0, -3)
text
..                                                                       [100%]
2 passed in 0.01s

The pytest.raises block says "this error is supposed to happen in here". When it does not, the test fails:

text
F                                                                        [100%]
================================== FAILURES ===================================
____________________________ test_one_is_rejected _____________________________

    def test_one_is_rejected():
>       with pytest.raises(ValueError):
E       Failed: DID NOT RAISE <class 'ValueError'>

test_pricing.py:7: Failed
=========================== short test summary info ===========================
FAILED test_pricing.py::test_one_is_rejected - Failed: DID NOT RAISE <class '...
1 failed in 0.01s

The match= part checks that the message contains that text. It is worth writing, because chapter twenty-four's rule — put the offending value in the message — then becomes a tested promise of its own.

One test, many cases

Writing four functions for four cases is dull, and invites copy-paste mistakes.

python
import pytest

from pricing import line_total


@pytest.mark.parametrize(
    "price, quantity, expected",
    [
        (15.0, 1, 17.25),
        (15.0, 3, 51.75),
        (0.0, 5, 0.0),
        (100.0, 2, 230.0),
    ],
)
def test_line_total(price, quantity, expected):
    assert line_total(price, quantity) == expected

With pytest -v:

text
============================= test session starts =============================
collecting ... collected 4 items

test_pricing.py::test_line_total[15.0-1-17.25] PASSED                    [ 25%]
test_pricing.py::test_line_total[15.0-3-51.75] PASSED                    [ 50%]
test_pricing.py::test_line_total[0.0-5-0.0] PASSED                       [ 75%]
test_pricing.py::test_line_total[100.0-2-230.0] PASSED                   [100%]

Four separate tests from one function. And each name carries its values, so a failure identifies itself at once — written as a loop, the first failure would stop the rest and would not say which values were at fault.

When one is wrong, the report shows the values too:

text
________________________ test_line_total[15.0-3-51.0] _________________________

price = 15.0, quantity = 3, expected = 51.0

When a file is needed — tmp_path

Testing chapter twenty-three's code needs a file. Keeping one in the repository is a poor habit: the tests become entangled, and one test changing the file breaks another.

pytest can give each test an empty folder of its own:

python
from reader import read_names


def test_blank_lines_are_skipped(tmp_path):
    path = tmp_path / "names.txt"
    path.write_text("one\n\ntwo\n", encoding="utf-8")

    assert read_names(path) == ["one", "two"]


def test_empty_file_gives_empty_list(tmp_path):
    path = tmp_path / "names.txt"
    path.write_text("", encoding="utf-8")

    assert read_names(path) == []
text
..                                                                       [100%]
2 passed in 0.01s

tmp_path is a pathlib.Path — chapter twenty-three's Path, joined with / exactly as there. The parameter's name is what tells pytest what to supply, and each test gets a new folder, so two tests can use the same file name without meeting.

Repeated setup — fixture

When several tests need the same input, it can be written once and named:

python
import pytest

from pricing import order_total


@pytest.fixture
def order():
    return [("pen", 15.0, 3), ("bag", 850.0, 1)]


def test_total(order):
    assert order_total(order) == 1029.25


def test_one_line_removed(order):
    assert order_total(order[:1]) == 51.75
text
..                                                                       [100%]
2 passed in 0.01s

As with tmp_path, pytest matches by the parameter's name. And the function runs again for every test, so a test that modifies the list still leaves the next one a fresh one — chapter twenty-one's shared-state problem designed out rather than avoided.

The floating-point trap

python
def test_addition():
    assert 0.1 + 0.2 == 0.3
text
F                                                                        [100%]
================================== FAILURES ===================================
________________________________ test_addition ________________________________

    def test_addition():
>       assert 0.1 + 0.2 == 0.3
E       assert (0.1 + 0.2) == 0.3

test_money.py:2: AssertionError

Chapter five's old business, and tests are usually where it bites first.

python
import pytest


def test_addition():
    assert 0.1 + 0.2 == pytest.approx(0.3)
text
.                                                                        [100%]
1 passed in 0.01s

pytest.approx says "close enough will do". Use it in any test involving floats — unless the number has already been pinned down with round(), as line_total does.

What to test

One rule is worth more than all the others: a function that takes values and returns a value is easy to test.

Remember chapter twenty-one's pure functions? This chapter is their reward. Testing line_total needed no file, no input, no setting up — only calling it and looking at the result.

Conversely, a function that prints, asks for input or writes a file needs arrangements made before it can be tested. Which is why chapter twenty-three's report took no file: it took a list and returned a list.

Tests being hard to write is usually not the tests' fault but the code's shape.

What to test, in three kinds: what normally happens, the edges (zero, empty, one), and what is supposed to fail.


A complete example

pricing.py:

python
"""Prices and tax. Pure functions, which is what makes them testable."""

TAX_RATE = 0.15


def line_total(price: float, quantity: int) -> float:
    if quantity < 1:
        raise ValueError(f"quantity must be at least 1: {quantity}")
    return round(price * quantity * (1 + TAX_RATE), 2)


def order_total(lines: list[tuple[str, float, int]]) -> float:
    return round(sum(line_total(price, qty) for _, price, qty in lines), 2)

test_pricing.py:

python
"""Tests for pricing. Each one names the behaviour it protects."""

import pytest

from pricing import line_total, order_total


@pytest.mark.parametrize(
    "price, quantity, expected",
    [
        (15.0, 1, 17.25),
        (15.0, 3, 51.75),
        (0.0, 5, 0.0),
        (0.01, 1, 0.01),
    ],
)
def test_line_total_applies_tax(price, quantity, expected):
    assert line_total(price, quantity) == expected


@pytest.mark.parametrize("quantity", [0, -1, -100])
def test_quantity_below_one_is_rejected(quantity):
    with pytest.raises(ValueError, match="at least 1"):
        line_total(15.0, quantity)


def test_error_message_names_the_value():
    with pytest.raises(ValueError, match="at least 1: -3"):
        line_total(15.0, -3)


@pytest.fixture
def order():
    return [("pen", 15.0, 3), ("bag", 850.0, 1), ("ink", 120.0, 2)]


def test_order_total_sums_the_lines(order):
    assert order_total(order) == 1305.25


def test_empty_order_totals_zero():
    assert order_total([]) == 0


def test_one_bad_line_stops_the_order(order):
    with pytest.raises(ValueError):
        order_total(order + [("clip", 5.0, 0)])
text
...........                                                              [100%]
11 passed in 0.01s

Four things worth looking at.

Eleven tests from seven functions. parametrize made the difference, and each case is counted separately.

The names state behaviour, not code. test_quantity_below_one_is_rejected tells you what the program promises; test_line_total_2 tells you nothing. That name is the first thing a failure report shows, so it is worth making it a sentence.

The 0.01 case is not arbitrary. It is the smallest possible price — an edge, and precisely where round()'s behaviour is most in doubt. Bugs live at the edges, not in the middle.

The last test holds an indirect promise. order_total does no validation of its own — it relies on line_total. If somebody one day puts a try/except in order_total to skip bad lines quietly, this test will fail and ask: "did you really mean that?"


When it breaks

no tests ran The file name does not start with test_, or the function name does not. Both are needed.

ModuleNotFoundError: No module named 'pricing' Run pytest from the folder the files are in.

A test passes and checks nothing The assert is missing. Merely calling a function passes as long as it does not raise.

fixture 'order' not found The @pytest.fixture is missing, or the name does not match the parameter's.

A float comparison fails although the numbers look right Take pytest.approx, or round() in the function itself.

DID NOT RAISE The code inside the pytest.raises block did not fail. Either the code is wrong, or the test's expectation is.

A test passes alone and fails with the others The tests are sharing something — often a module-level list or dictionary. Chapters twenty-two and twenty-six's shared state. Use a fixture.