Chapter 16

Test design and TDD — what makes a good test

Planning a test before writing it, Arrange-Act-Assert, testing behaviour rather than implementation, red-green-refactor, property-based testing with Hypothesis, the causes and cures of flaky tests, and a regression test for every bug.

60 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

You now know every tool this course set out to teach: assert, raises, parametrize, fixtures, mocks, markers, configuration, coverage, plugins. And it is entirely possible to use all of them and still end up with tests that hurt more than they help.

Here is a small shopping cart, with two tests that pass:

python
class Cart:
    def __init__(self):
        self._items = []

    def add(self, name, price, quantity=1):
        self._items.append((name, price, quantity))

    def total(self):
        return sum(price * quantity for _, price, quantity in self._items)
python
from cart import Cart


def test_add():
    cart = Cart()
    cart.add("pen", 15.0, 2)
    assert cart._items == [("pen", 15.0, 2)]


def test_total():
    cart = Cart()
    cart.add("pen", 15.0, 2)
    cart.add("bag", 850.0)
    assert cart.total() == 880.0
text
..                                                                       [100%]
2 passed in 0.01s

Now someone improves the cart. Adding the same item twice should merge into one line, so the list becomes a dictionary keyed by name. The cart still adds things and still totals them correctly — nothing a customer could notice has changed:

python
class Cart:
    def __init__(self):
        self._lines = {}

    def add(self, name, price, quantity=1):
        _, already = self._lines.get(name, (price, 0))
        self._lines[name] = (price, already + quantity)

    def total(self):
        return sum(price * quantity for price, quantity in self._lines.values())
text
F.                                                                       [100%]
=================================== FAILURES ===================================
___________________________________ test_add ___________________________________

    def test_add():
        cart = Cart()
        cart.add("pen", 15.0, 2)
>       assert cart._items == [("pen", 15.0, 2)]
               ^^^^^^^^^^^
E       AttributeError: 'Cart' object has no attribute '_items'

test_cart.py:7: AssertionError
=========================== short test summary info ============================
FAILED test_cart.py::test_add - AttributeError: 'Cart' object has no attribut...
1 failed, 1 passed in 0.01s

A red test, and no bug. The test was not checking what the cart does; it was checking how the cart was built. A test like that fails when you improve the code and stays silent when you break it — the exact opposite of its job.

This last chapter is not about another tool. It is about judgement: what a good test looks like, how to let tests drive the code, how to let the computer invent test cases for you, and how to stop tests from failing at random.

By the end of this chapter you can

  • Plan a test before writing it: the contract, the cases, and what to leave out
  • Structure a test as Arrange, Act, Assert, and name it as a sentence
  • Tell a test of behaviour from a test of implementation, and rewrite the second into the first
  • Work in red-green-refactor cycles, with a real failure at every red step
  • Write property-based tests with Hypothesis and read a shrunk failing example
  • Find and fix the four usual causes of flaky tests
  • Turn every bug report into a regression test

Prerequisites: async tests, hooks and plugins.


Before you write the test

Most bad tests are not badly written. They are badly planned — written before anyone decided what the code promises. So before any test code, four questions.

Through this chapter we will build one small function test-first: slugify, which turns a post title such as "Café au lait" into the URL piece "cafe-au-lait".

1. What is the contract? Say it in one or two sentences, as the caller sees it: given a title (a `str`), return a slug: lowercase ASCII letters and digits, words joined by single hyphens, no hyphen at either end. Accented letters become their plain letter. No side effects — it touches no file, no clock, no network. Notice what the contract does not mention: regular expressions, Unicode normalisation, helper functions. Those are how it is done, and they are free to change.

2. What must be in place? A virtual environment with the tools installed, and the code importable from the tests:

text
$ python -m venv .venv
$ source .venv/bin/activate
$ pip install pytest hypothesis

Then two files side by side — slugs.py and test_slugs.py — and pytest run from that folder, so from slugs import slugify works. This topic needs no fixtures, no temporary files and no environment variables. That is worth noticing too: a pure function is the easiest thing in the world to test, which is a good reason to push logic into pure functions.

3. Which cases? Happy path first, then the edges, then invalid input, then boundaries:

| Case | Input | Expected | |---|---|---| | ordinary words | "Hello World" | "hello-world" | | punctuation | "Hello, World!" | "hello-world" | | accented letters | "Café au lait" | "cafe-au-lait" | | extra spaces at the ends | " Hello World " | "hello-world" | | digits are kept | "Top 10 Tips" | "top-10-tips" | | nothing usable | "!!!" | ? |

That question mark is the most useful cell in the table. Writing the plan has exposed a decision nobody has made: what should happen when a title contains no letters at all? We will leave it open on purpose, and see later in this chapter what an open question costs.

4. What will you not test? Python's re module, unicodedata, and str.lower() already have their own tests; re-testing them only slows you down. Private helpers (anything starting with _) are reached through slugify, never directly. And you will not try to list every Unicode character by hand — that is a job for a property-based test, later in this chapter.

The table becomes the tests. Everything below implements that plan.

Arrange, Act, Assert

Every good test has the same three parts, in the same order:

  • Arrange — build the world the test needs
  • Act — do the one thing being tested
  • Assert — check what came out

Here is the cart tested again — this time through what it promises, not how it stores things:

python
from cart import Cart


def test_an_empty_cart_totals_zero():
    cart = Cart()

    assert cart.total() == 0


def test_total_is_price_times_quantity_summed_over_items():
    # Arrange
    cart = Cart()
    cart.add("pen", 15.0, 2)
    cart.add("bag", 850.0)

    # Act
    total = cart.total()

    # Assert
    assert total == 880.0


def test_adding_the_same_item_twice_adds_up_the_quantity():
    cart = Cart()
    cart.add("pen", 15.0, 2)

    cart.add("pen", 15.0, 1)

    assert cart.total() == 45.0

Against the list-based cart, and again against the dictionary-based one:

text
...                                                                      [100%]
3 passed in 0.01s
...                                                                      [100%]
3 passed in 0.01s

Same tests, two implementations, all green both times. That is the property you want: a refactor that keeps behaviour should keep the tests green. The comments in the middle test are only there to show the shape; most people drop them and separate the three parts with blank lines, as in the third test.

Why does it matter that the Act is a single step? Because when the test fails, you want to know which thing broke. If the Act is five calls, a failure points at five suspects.

Test behaviour, not implementation

The rule of thumb: a test may use only what a caller may use. For the cart that is Cart(), .add() and .total(). Not _items, not _lines. The leading underscore is Python's way of saying "this is mine, and I may change it".

The wrong way couples the test to a decision that was never part of the contract — that items are kept in a list. The right way asks the question a caller would ask: if I add these things, what is the total?

Two signs that a test is coupled to implementation:

  • It reads attributes that start with _, or mocks a function the code calls internally
  • A refactor that no user could notice turns it red

Mocks deserve a note here. Chapter eleven's mocks are the right tool at a boundary — the network, the clock, a payment service. Used inside your own code ("assert that _calculate was called with these arguments"), they become implementation tests by another name.

Names that read as sentences, and one reason to fail

Run the three cart tests with -v:

text
test_cart_behaviour.py::test_an_empty_cart_totals_zero PASSED            [ 33%]
test_cart_behaviour.py::test_total_is_price_times_quantity_summed_over_items PASSED [ 66%]
test_cart_behaviour.py::test_adding_the_same_item_twice_adds_up_the_quantity PASSED [100%]

Read the names alone and you have a specification of the cart. test_add and test_total told you only which method was touched. A good name says the situation and the expected result: an empty cart totals zero. It is long; that is fine. You never type it, and you read it every time it fails.

The second half of the rule is one reason to fail. Here is the opposite — one test that checks everything, run against a cart with a bug (it ignores quantity):

python
from cart import Cart


def test_cart():
    cart = Cart()
    assert cart.total() == 0
    cart.add("pen", 15.0, 2)
    cart.add("bag", 850.0)
    assert cart.total() == 880.0
    cart.add("pen", 15.0, 1)
    assert cart.total() == 895.0
text
$ pytest -q --tb=no
F                                                                        [100%]
=========================== short test summary info ============================
FAILED test_one_big.py::test_cart - assert 865.0 == 880.0
1 failed in 0.01s

(--tb=no hides the tracebacks and leaves only the summary — the part you read first.) test_cart failed — which says nothing — with 865.0 == 880.0, which you now have to puzzle over. And the third assertion never ran, so you do not know whether it would have passed. The same bug against the three focused tests:

text
$ pytest -q --tb=no
.FF                                                                      [100%]
=========================== short test summary info ============================
FAILED test_cart_behaviour.py::test_total_is_price_times_quantity_summed_over_items
FAILED test_cart_behaviour.py::test_adding_the_same_item_twice_adds_up_the_quantity
2 failed, 1 passed in 0.01s

Before opening a single traceback, the summary already reads like a diagnosis: the empty cart is fine, but quantities are going wrong. "One reason to fail" does not mean one assert line — two asserts that check two sides of the same result are fine. It means one behaviour per test.

The test pyramid

Tests come in sizes:

  • Unit tests check one function or class with nothing real around it. Milliseconds each. You have thousands.
  • Integration tests check that pieces fit: your code against a real database, a real file system, a real HTTP app via a test client. Slower; you have dozens to hundreds.
  • End-to-end tests drive the whole system the way a user would — a browser, a deployed API. Slow and fragile; you have a handful, covering the paths that make money.

Drawn by count, that is a pyramid: wide base of unit tests, thin top of end-to-end ones. The reason is cost. When a unit test fails, it points at a few lines. When an end-to-end test fails, it points at the whole system. Push each check as far down the pyramid as it can go, and keep the top for what only the top can see.

Red, green, refactor

Test-driven development (TDD) runs in a tight loop:

  1. Red — write one small test for behaviour that does not exist yet, and watch it fail
  2. Green — write the least code that makes it pass
  3. Refactor — tidy the code, with the tests green the whole time

Why watch it fail? Because a test you have never seen fail might be testing nothing. The red step proves the test can catch the absence of the feature.

We take the plan from the table, one row at a time.

Red. The first row, before slugs.py exists:

python
from slugs import slugify


def test_lowercases_and_joins_words_with_hyphens():
    assert slugify("Hello World") == "hello-world"
text
==================================== ERRORS ====================================
________________________ ERROR collecting test_slugs.py ________________________
ImportError while importing test module '/home/you/blog/test_slugs.py'.
Hint: make sure your test modules/packages have valid Python names.
Traceback:
/usr/lib/python3.12/importlib/__init__.py:90: in import_module
    return _bootstrap._gcd_import(name[level:], package, level)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
test_slugs.py:1: in <module>
    from slugs import slugify
E   ModuleNotFoundError: No module named 'slugs'
=========================== short test summary info ============================
ERROR test_slugs.py
!!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!!
1 error in 0.01s

That is a perfectly good red. The test is asking for something that is not there.

Green. The least code that passes — and it really is the least:

python
def slugify(title):
    return title.lower().replace(" ", "-")
text
.                                                                        [100%]
1 passed in 0.01s

It feels like cheating. It is deliberate: the code does only what a test has demanded so far. Anything more would be code no test is checking.

Red. The second row of the table:

python
def test_drops_punctuation():
    assert slugify("Hello, World!") == "hello-world"
text
.F                                                                       [100%]
=================================== FAILURES ===================================
____________________________ test_drops_punctuation ____________________________

    def test_drops_punctuation():
>       assert slugify("Hello, World!") == "hello-world"
E       AssertionError: assert 'hello,-world!' == 'hello-world'
E
E         - hello-world
E         + hello,-world!
E         ?      +      +

test_slugs.py:9: AssertionError
=========================== short test summary info ============================
FAILED test_slugs.py::test_drops_punctuation - AssertionError: assert 'hello,...
1 failed, 1 passed in 0.01s

The ? line marks exactly the two characters that should not be there.

Green. Now the replace-one-space trick is not enough. Replace every run of "not a letter or digit" with one hyphen, then trim hyphens from the ends:

python
import re


def slugify(title):
    return re.sub(r"[^a-z0-9]+", "-", title.lower()).strip("-")
text
..                                                                       [100%]
2 passed in 0.01s

The rows for extra spaces and digits would already pass with this code. Add them anyway — they are part of the contract — but know that a test that is green the first time it runs has not proved anything about the code you just wrote. Only a red-then-green test has.

Red. Accented letters:

python
def test_turns_accented_letters_into_plain_ones():
    assert slugify("Café au lait") == "cafe-au-lait"
text
..F                                                                      [100%]
=================================== FAILURES ===================================
_________________ test_turns_accented_letters_into_plain_ones __________________

    def test_turns_accented_letters_into_plain_ones():
>       assert slugify("Café au lait") == "cafe-au-lait"
E       AssertionError: assert 'caf-au-lait' == 'cafe-au-lait'
E
E         - cafe-au-lait
E         ?    -
E         + caf-au-lait

test_slugs.py:13: AssertionError
=========================== short test summary info ============================
FAILED test_slugs.py::test_turns_accented_letters_into_plain_ones - Assertion...
1 failed, 2 passed in 0.01s

é is not in a-z, so it was thrown away with the punctuation.

Green. Unicode "NFKD" normalisation splits é into e plus a separate accent mark; encoding to ASCII with "ignore" then drops the mark and keeps the e:

python
import re
import unicodedata


def slugify(title):
    plain = unicodedata.normalize("NFKD", title).encode("ascii", "ignore").decode()
    return re.sub(r"[^a-z0-9]+", "-", plain.lower()).strip("-")
text
...                                                                      [100%]
3 passed in 0.01s

Refactor. It works, but the body is one dense line. Give each step a name, compile the pattern once, and add type hints — changing no behaviour:

python
import re
import unicodedata

NOT_ALLOWED = re.compile(r"[^a-z0-9]+")


def _to_ascii(text: str) -> str:
    """Split accented letters into letter + accent, then drop the accents."""
    decomposed = unicodedata.normalize("NFKD", text)
    return decomposed.encode("ascii", "ignore").decode("ascii")


def slugify(title: str) -> str:
    words = NOT_ALLOWED.sub("-", _to_ascii(title).lower())
    return words.strip("-")
text
test_slugs.py::test_lowercases_and_joins_words_with_hyphens PASSED       [ 33%]
test_slugs.py::test_drops_punctuation PASSED                             [ 66%]
test_slugs.py::test_turns_accented_letters_into_plain_ones PASSED        [100%]

============================== 3 passed in 0.01s ===============================

This is the step that makes TDD pay. You could restructure freely because the tests check behaviour; had they reached into _to_ascii or the regex, the refactor would have broken them.

Property-based testing with Hypothesis

Every test so far has one invented example: "Hello World", "Café au lait". The examples are only as good as your imagination, and your imagination shares its blind spots with the code you just wrote.

A property is a statement that holds for every input. Hypothesis generates the inputs — a hundred per test by default — and tries hard to find one that breaks the statement. @given says where the inputs come from; strategies (always imported as st) describes their shape:

python
from hypothesis import given
from hypothesis import strategies as st


@given(
    prices=st.lists(st.integers(min_value=0, max_value=10_000), max_size=20),
    discount=st.integers(min_value=0, max_value=100),
)
def test_a_discount_never_raises_the_total(prices, discount):
    total = sum(prices)

    discounted = total * (100 - discount) // 100

    assert 0 <= discounted <= total
text
.                                                                        [100%]
1 passed in 0.01s

One dot, but a hundred lists of prices and a hundred discounts went through it. Other strategies you will reach for: st.text(), st.floats(), st.booleans(), st.sampled_from([...]), st.dictionaries(...), st.builds(...).

What properties does slugify have? You cannot say what the slug of a random string is, but you can say what it must look like, and that slugifying a slug changes nothing:

python
import re

from hypothesis import given
from hypothesis import strategies as st

from slugs import slugify


@given(st.text())
def test_a_slug_only_holds_lowercase_letters_digits_and_inner_hyphens(title):
    slug = slugify(title)

    assert re.fullmatch(r"[a-z0-9]+(-[a-z0-9]+)*|", slug)


@given(st.text())
def test_slugifying_a_slug_changes_nothing(title):
    once = slugify(title)

    assert slugify(once) == once
text
..                                                                       [100%]
2 passed in 0.01s

st.text() produces emoji, Chinese, control characters, empty strings — inputs you would never have typed into the plan table.

A round trip finds a real bug

The most productive property is the round trip: if you can encode something and decode it again, decoding the encoding must give back the original. Here is a run-length encoder — "aaab" becomes "3a1b" — with two example tests that pass:

python
import re


def encode(text: str) -> str:
    """'aaab' -> '3a1b': each run of a character becomes count + character."""
    out = []
    for match in re.finditer(r"(.)\1*", text, flags=re.DOTALL):
        run = match.group(0)
        out.append(f"{len(run)}{run[0]}")
    return "".join(out)


def decode(encoded: str) -> str:
    return "".join(char * int(count) for count, char in re.findall(r"(\d+)(\D)", encoded))
python
from hypothesis import given
from hypothesis import strategies as st

from rle import decode, encode


def test_encode_counts_each_run():
    assert encode("aaab") == "3a1b"


def test_decode_expands_each_run():
    assert decode("3a1b") == "aaab"


@given(st.text())
def test_decoding_an_encoding_gives_back_the_original(text):
    assert decode(encode(text)) == text
text
..F                                                                      [100%]
=================================== FAILURES ===================================
______________ test_decoding_an_encoding_gives_back_the_original _______________

    @given(st.text())
>   def test_decoding_an_encoding_gives_back_the_original(text):
                   ^^^

test_rle.py:16:
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _

text = '0'

    @given(st.text())
    def test_decoding_an_encoding_gives_back_the_original(text):
>       assert decode(encode(text)) == text
E       AssertionError: assert '' == '0'
E
E         - 0
E       Failing test case: test_decoding_an_encoding_gives_back_the_original(
E           text='0',
E       )

test_rle.py:17: AssertionError
=========================== short test summary info ============================
FAILED test_rle.py::test_decoding_an_encoding_gives_back_the_original - Asser...
1 failed, 2 passed in 0.01s

The text "0" encodes to "10" — "one zero" — and the decoder reads 10 as a count with no character after it. Any text containing a digit is corrupted. Both example tests used only letters, so they could never see it.

Shrinking

Hypothesis did not first stumble on "0". It found some long, ugly string, then shrank it: tried smaller and simpler inputs, keeping each one that still failed, until nothing simpler failed. You can watch that happen by recording each failing input:

python
from hypothesis import given, seed, settings
from hypothesis import strategies as st

from rle import decode, encode

failing = []


@seed(2026)
@settings(database=None)
@given(st.text())
def round_trip(text):
    if decode(encode(text)) != text:
        if text not in failing:
            failing.append(text)
        raise AssertionError


try:
    round_trip()
except AssertionError:
    pass

for text in failing:
    print(repr(text))
text
'wfê\x9f\x0c\U000b6082Ñ/\x879\x9dÂ\n'
'\x03\U000a1acb\U0003c4fc1'
'´\U000cd3dc\U00081acfÔ÷\U0007122f\x89ê½5'
'l𨺮\nÝÛ\x894\x8b'
'\U00083189𤾨«ñï\xa0^ó8A\x96\x88=\x05'
'0000'
'000'
'00'
'0'

(@seed fixes the random choices so this run is repeatable; database=None stops Hypothesis replaying a remembered failure.) The first failure is thirteen characters of noise with a 9 hidden in it. You would have stared at it for minutes. The shrunk example says the whole story in one character: a digit breaks it.

Hypothesis also saves failing examples in a .hypothesis/ folder, so the next run tries "0" first. Add that folder to .gitignore.

The fix puts a separator between count and character, so a digit can never be mistaken for part of a count. And the example Hypothesis found is pinned with @example, so it runs every time, on every machine:

python
import re


def encode(text: str) -> str:
    """'aaab' -> '3:a1:b': each run becomes count, a colon, then the character."""
    out = []
    for match in re.finditer(r"(.)\1*", text, flags=re.DOTALL):
        run = match.group(0)
        out.append(f"{len(run)}:{run[0]}")
    return "".join(out)


def decode(encoded: str) -> str:
    pairs = re.findall(r"(\d+):(.)", encoded, flags=re.DOTALL)
    return "".join(char * int(count) for count, char in pairs)

The two example tests change to the new format ("3:a1:b"), and the property gains one line:

python
from hypothesis import example, given
from hypothesis import strategies as st

from rle import decode, encode


@given(st.text())
@example("0")  # the input Hypothesis found; now it is checked on every run
def test_decoding_an_encoding_gives_back_the_original(text):
    assert decode(encode(text)) == text
text
...                                                                      [100%]
3 passed in 0.01s

Properties do not replace examples. The two example tests document the format in a way a reader understands at a glance; the property guards the corners nobody thought of. Use both.

Flaky tests

A flaky test passes and fails on the same code. It is worse than no test: people learn to press "re-run" and stop believing red. Nearly every flaky test comes from one of four causes, and each has a fix that removes the cause rather than hiding it.

Shared state and test order

python
_users = set()


def register(name):
    _users.add(name)


def count():
    return len(_users)
python
from registry import count, register


def test_registering_a_user_counts_them():
    register("rahim")

    assert count() == 1


def test_a_new_registry_is_empty():
    assert count() == 0
text
$ pytest -q --tb=no
.F                                                                       [100%]
=========================== short test summary info ============================
FAILED test_registry.py::test_a_new_registry_is_empty - assert 1 == 0
1 failed, 1 passed in 0.01s

$ pytest -q "test_registry.py::test_a_new_registry_is_empty"
.                                                                        [100%]
1 passed in 0.01s

Fails in the suite, passes alone. The module-level set survives from one test to the next, so the second test sees what the first left behind. Change the order, or run tests in parallel with pytest-xdist, and the result changes. The wrong fix is to reorder the tests. The right fix removes the shared state: make the registry an object, and give each test a fresh one through a fixture:

python
class Registry:
    def __init__(self):
        self._users = set()

    def register(self, name):
        self._users.add(name)

    def count(self):
        return len(self._users)
python
import pytest

from registry import Registry


@pytest.fixture
def registry():
    return Registry()  # a fresh one for every test


def test_registering_a_user_counts_them(registry):
    registry.register("rahim")

    assert registry.count() == 1


def test_a_new_registry_is_empty(registry):
    assert registry.count() == 0
text
..                                                                       [100%]
2 passed in 0.01s

When you cannot change the code, an autouse fixture that resets the state before and after each test is the fallback. Every test should pass alone, and in any order.

Time

python
from datetime import datetime


def greeting(now=None):
    now = now or datetime.now()
    return "Good morning" if now.hour < 12 else "Good afternoon"
python
from greeting import greeting


def test_greets_the_morning():
    assert greeting() == "Good morning"  # true only before noon

The same code, run at the same moment, on two machines set to different time zones:

text
$ TZ=Europe/London pytest -q
.                                                                        [100%]
1 passed in 0.01s

$ TZ=Asia/Dhaka pytest -q --tb=no
F                                                                        [100%]
=========================== short test summary info ============================
FAILED test_greeting.py::test_greets_the_morning - AssertionError: assert 'Go...
1 failed in 0.01s

The fix is already in the function's signature: now can be passed in. A test that controls the clock tests both sides of the boundary, and gives the same answer at any hour:

python
from datetime import datetime

from greeting import greeting


def test_before_noon_it_says_good_morning():
    assert greeting(now=datetime(2026, 1, 5, 9, 30)) == "Good morning"


def test_from_noon_on_it_says_good_afternoon():
    assert greeting(now=datetime(2026, 1, 5, 12, 0)) == "Good afternoon"
text
..                                                                       [100%]
2 passed in 0.01s

Passing the time in is simpler than any mock. When you cannot change the signature, monkeypatch from chapter ten replaces the clock instead.

Randomness

python
import random


def pick_winner(names, rng=random):
    return rng.choice(names)
python
from raffle import pick_winner


def test_picks_a_winner():
    assert pick_winner(["rahim", "karim", "salma"]) == "rahim"

Six runs, no change to anything:

text
$ for i in 1 2 3 4 5 6; do pytest -q | tail -1; done
1 failed in 0.01s
1 failed in 0.01s
1 failed in 0.01s
1 passed in 0.01s
1 passed in 0.01s
1 passed in 0.01s

Two fixes, for two kinds of question. Assert what is true for every outcome — the winner is one of the entrants. Or take control of the randomness by passing a seeded generator:

python
import random

from raffle import pick_winner


def test_the_winner_is_one_of_the_entrants():
    names = ["rahim", "karim", "salma"]

    assert pick_winner(names) in names


def test_the_same_seed_picks_the_same_winner():
    names = ["rahim", "karim", "salma"]

    first = pick_winner(names, rng=random.Random(42))
    second = pick_winner(names, rng=random.Random(42))

    assert first == second
text
$ for i in 1 2 3; do pytest -q | tail -1; done
2 passed in 0.01s
2 passed in 0.01s
2 passed in 0.01s

The outside world

The fourth cause is anything you do not control: a real network call, a real server, a sleep(0.1) that "should be enough". The fixes are the ones from earlier chapters — mock at the boundary, use tmp_path instead of a shared folder, wait for a condition instead of a fixed time. One shape covers all four causes: whatever the test depends on, the test should create or control.

A regression test for every bug

Back to the question mark in the plan table. Nobody answered it, and the bug report arrives: a post titled `"!!!"` was saved at `/posts/`, overwriting the index page.

Before touching slugify, write a test that reproduces the report and watch it fail. That proves the test catches this bug:

python
import pytest

from slugs import slugify


def test_a_title_with_no_letters_or_digits_is_rejected():
    # Bug: slugify("!!!") returned "", and the post was saved at /posts/
    with pytest.raises(ValueError, match="no letters or digits"):
        slugify("!!!")
text
$ pytest -q test_slugs_regressions.py
F                                                                        [100%]
=================================== FAILURES ===================================
______________ test_a_title_with_no_letters_or_digits_is_rejected ______________

    def test_a_title_with_no_letters_or_digits_is_rejected():
        # Bug: slugify("!!!") returned "", and the post was saved at /posts/
>       with pytest.raises(ValueError, match="no letters or digits"):
E       Failed: DID NOT RAISE ValueError

test_slugs_regressions.py:8: Failed
=========================== short test summary info ============================
FAILED test_slugs_regressions.py::test_a_title_with_no_letters_or_digits_is_rejected
1 failed in 0.01s

Then fix it — the end of slugify becomes:

python
def slugify(title: str) -> str:
    slug = NOT_ALLOWED.sub("-", _to_ascii(title).lower()).strip("-")
    if not slug:
        raise ValueError(f"title has no letters or digits: {title!r}")
    return slug
text
....                                                                     [100%]
4 passed in 0.01s

The regression test and the three example tests pass. But run the whole suite, and the property tests object:

text
FAILED test_slugs_properties.py::test_a_slug_only_holds_lowercase_letters_digits_and_inner_hyphens
FAILED test_slugs_properties.py::test_slugifying_a_slug_changes_nothing - Val...
2 failed, 4 passed in 0.01s

Hypothesis tried the empty string, and slugify("") now raises. That is not a new bug — it is the contract changing on purpose, and the properties noticing. Good. Update them to state the new contract: titles that contain at least one letter or digit.

python
import re
import string

from hypothesis import given
from hypothesis import strategies as st

from slugs import slugify

# Any text, with at least one plain letter or digit somewhere inside it.
titles = st.builds(
    lambda before, word, after: before + word + after,
    st.text(),
    st.text(alphabet=string.ascii_letters + string.digits, min_size=1),
    st.text(),
)


@given(titles)
def test_a_slug_only_holds_lowercase_letters_digits_and_inner_hyphens(title):
    slug = slugify(title)

    assert re.fullmatch(r"[a-z0-9]+(-[a-z0-9]+)*", slug)

test_slugifying_a_slug_changes_nothing switches to @given(titles) in the same way.

text
......                                                                   [100%]
6 passed in 0.01s

Why a test for every bug? Because a bug that happened once has proved it is easy to make, and the next person to touch that code is likely to make it again. The comment saying what went wrong is part of the test: in a year, it is the only record of why this odd-looking case is here.


A complete example

A password-strength checker, built with everything in this chapter. The plan first:

| Case | Input | Expected | |---|---|---| | empty | "" | weak | | every kind, too short | "Ab1!xyz" (7) | weak | | long enough, one kind | "abcdefgh" | weak | | long enough, two kinds | "abcdefg1" | medium | | 12 long, three kinds | "abcdefghij1!" | strong | | 12 long, two kinds | "abcdefghijk1" | medium |

passwords.py:

python
def _kinds(password: str) -> int:
    """How many of the four kinds of character the password uses."""
    return sum([
        any(c.islower() for c in password),
        any(c.isupper() for c in password),
        any(c.isdigit() for c in password),
        any(not c.isalnum() for c in password),
    ])


def strength(password: str) -> str:
    """Rate a password as 'weak', 'medium' or 'strong'."""
    if len(password) < 8:
        return "weak"
    kinds = _kinds(password)
    if len(password) >= 12 and kinds >= 3:
        return "strong"
    if kinds >= 2:
        return "medium"
    return "weak"

test_passwords.py:

python
import pytest
from hypothesis import given
from hypothesis import strategies as st

from passwords import strength

RANK = {"weak": 0, "medium": 1, "strong": 2}


@pytest.mark.parametrize(
    "password, expected",
    [
        ("", "weak"),                      # nothing at all
        ("Ab1!xyz", "weak"),               # every kind, but only 7 long
        ("abcdefgh", "weak"),              # 8 long, one kind
        ("abcdefg1", "medium"),            # 8 long, two kinds
        ("abcdefghij1!", "strong"),        # 12 long, three kinds
        ("abcdefghijk1", "medium"),        # 12 long, only two kinds
    ],
)
def test_strength_follows_the_length_and_variety_rules(password, expected):
    assert strength(password) == expected


@given(st.text(), st.text())
def test_adding_characters_never_makes_a_password_weaker(password, extra):
    before = strength(password)

    after = strength(password + extra)

    assert RANK[after] >= RANK[before]


def test_a_long_password_of_one_kind_is_still_weak():
    # Bug: "aaaaaaaaaaaaaaaaaaaa" (20 letters) was rated "medium"
    assert strength("a" * 20) == "weak"
text
$ pytest -v
collected 8 items

test_passwords.py::test_strength_follows_the_length_and_variety_rules[-weak] PASSED [ 12%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[Ab1!xyz-weak] PASSED [ 25%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefgh-weak] PASSED [ 37%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefg1-medium] PASSED [ 50%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefghij1!-strong] PASSED [ 62%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefghijk1-medium] PASSED [ 75%]
test_passwords.py::test_adding_characters_never_makes_a_password_weaker PASSED [ 87%]
test_passwords.py::test_a_long_password_of_one_kind_is_still_weak PASSED [100%]

============================== 8 passed in 0.01s ===============================

Three decisions are worth noticing.

The table rows sit on the boundaries — 7 and 8 characters, 11 and 12, two kinds and three. Bugs live at boundaries: a < that should be <= is invisible in the middle of a range.

The property says something no example can: adding characters never makes a password weaker. Nobody would write that test by hand for every password, and Hypothesis checks it against hundreds, including Unicode letters that are neither upper nor lower case.

And no test touches _kinds. It is reached through strength; tomorrow it may be gone.


When it breaks

AttributeError: 'Cart' object has no attribute '_items' after a refactor The test was reading the inside of the object. Rewrite it to check through the public methods — what a caller can see — and it will survive the next refactor too.

A test passes alone but fails in the full run (or the other way round) Shared state: a module-level list, dictionary or cache; a file in a fixed location; an environment variable set and never unset. Make each test create what it needs (fixtures, tmp_path, monkeypatch) instead of reordering the tests.

hypothesis.errors.FailedHealthCheck: It looks like this test is filtering out a lot of inputs. 0 inputs were generated successfully, while 50 inputs were filtered out. A .filter() (or assume()) throws away nearly everything Hypothesis generates — for example st.integers().filter(lambda n: n % 1000 == 7). Build the values you want directly instead: st.integers().map(lambda n: n * 1000 + 7).

hypothesis.errors.FlakyFailure: Hypothesis test_depends_on_earlier_runs(n=-25617) produces unreliable results: Failed on the first call but did not on a subsequent one The test gave a different answer for the same input. Something outside the generated arguments changed between calls — a counter, a module-level list, the clock. Hypothesis replays each failure, so a test with hidden state cannot shrink.

hypothesis.errors.DeadlineExceeded: Test took 300.07ms, which exceeds the deadline of 200.00ms. Each generated example must finish within 200 ms by default, because a hundred slow examples make a slow suite. Make the test faster, or, if the slowness is genuine, raise it with @settings(deadline=...).

A new TDD test passes the first time you run it It is not proving anything about the code you are about to write. Either the behaviour already exists (fine — keep the test as documentation), or the test is not checking what you think. Break the code on purpose for a moment and confirm the test goes red.