Test design and TDD — what makes a good test
Planning a test before writing it, Arrange-Act-Assert, testing behaviour rather than implementation, red-green-refactor, property-based testing with Hypothesis, the causes and cures of flaky tests, and a regression test for every bug.
- 1Encounter
- 2Understand
- 3Worked
- 4Predict
- 5Apply
- 6Stretch
The problem we are solving
You now know every tool this course set out to teach: assert, raises, parametrize, fixtures, mocks, markers, configuration, coverage, plugins. And it is entirely possible to use all of them and still end up with tests that hurt more than they help.
Here is a small shopping cart, with two tests that pass:
class Cart:
def __init__(self):
self._items = []
def add(self, name, price, quantity=1):
self._items.append((name, price, quantity))
def total(self):
return sum(price * quantity for _, price, quantity in self._items)from cart import Cart
def test_add():
cart = Cart()
cart.add("pen", 15.0, 2)
assert cart._items == [("pen", 15.0, 2)]
def test_total():
cart = Cart()
cart.add("pen", 15.0, 2)
cart.add("bag", 850.0)
assert cart.total() == 880.0.. [100%]
2 passed in 0.01sNow someone improves the cart. Adding the same item twice should merge into one line, so the list becomes a dictionary keyed by name. The cart still adds things and still totals them correctly — nothing a customer could notice has changed:
class Cart:
def __init__(self):
self._lines = {}
def add(self, name, price, quantity=1):
_, already = self._lines.get(name, (price, 0))
self._lines[name] = (price, already + quantity)
def total(self):
return sum(price * quantity for price, quantity in self._lines.values())F. [100%]
=================================== FAILURES ===================================
___________________________________ test_add ___________________________________
def test_add():
cart = Cart()
cart.add("pen", 15.0, 2)
> assert cart._items == [("pen", 15.0, 2)]
^^^^^^^^^^^
E AttributeError: 'Cart' object has no attribute '_items'
test_cart.py:7: AssertionError
=========================== short test summary info ============================
FAILED test_cart.py::test_add - AttributeError: 'Cart' object has no attribut...
1 failed, 1 passed in 0.01sA red test, and no bug. The test was not checking what the cart does; it was checking how the cart was built. A test like that fails when you improve the code and stays silent when you break it — the exact opposite of its job.
This last chapter is not about another tool. It is about judgement: what a good test looks like, how to let tests drive the code, how to let the computer invent test cases for you, and how to stop tests from failing at random.
By the end of this chapter you can
- Plan a test before writing it: the contract, the cases, and what to leave out
- Structure a test as Arrange, Act, Assert, and name it as a sentence
- Tell a test of behaviour from a test of implementation, and rewrite the second into the first
- Work in red-green-refactor cycles, with a real failure at every red step
- Write property-based tests with Hypothesis and read a shrunk failing example
- Find and fix the four usual causes of flaky tests
- Turn every bug report into a regression test
Prerequisites: async tests, hooks and plugins.
Before you write the test
Most bad tests are not badly written. They are badly planned — written before anyone decided what the code promises. So before any test code, four questions.
Through this chapter we will build one small function test-first: slugify, which turns a post title such as "Café au lait" into the URL piece "cafe-au-lait".
1. What is the contract? Say it in one or two sentences, as the caller sees it: given a title (a `str`), return a slug: lowercase ASCII letters and digits, words joined by single hyphens, no hyphen at either end. Accented letters become their plain letter. No side effects — it touches no file, no clock, no network. Notice what the contract does not mention: regular expressions, Unicode normalisation, helper functions. Those are how it is done, and they are free to change.
2. What must be in place? A virtual environment with the tools installed, and the code importable from the tests:
$ python -m venv .venv
$ source .venv/bin/activate
$ pip install pytest hypothesisThen two files side by side — slugs.py and test_slugs.py — and pytest run from that folder, so from slugs import slugify works. This topic needs no fixtures, no temporary files and no environment variables. That is worth noticing too: a pure function is the easiest thing in the world to test, which is a good reason to push logic into pure functions.
3. Which cases? Happy path first, then the edges, then invalid input, then boundaries:
| Case | Input | Expected | |---|---|---| | ordinary words | "Hello World" | "hello-world" | | punctuation | "Hello, World!" | "hello-world" | | accented letters | "Café au lait" | "cafe-au-lait" | | extra spaces at the ends | " Hello World " | "hello-world" | | digits are kept | "Top 10 Tips" | "top-10-tips" | | nothing usable | "!!!" | ? |
That question mark is the most useful cell in the table. Writing the plan has exposed a decision nobody has made: what should happen when a title contains no letters at all? We will leave it open on purpose, and see later in this chapter what an open question costs.
4. What will you not test? Python's re module, unicodedata, and str.lower() already have their own tests; re-testing them only slows you down. Private helpers (anything starting with _) are reached through slugify, never directly. And you will not try to list every Unicode character by hand — that is a job for a property-based test, later in this chapter.
The table becomes the tests. Everything below implements that plan.
Arrange, Act, Assert
Every good test has the same three parts, in the same order:
- Arrange — build the world the test needs
- Act — do the one thing being tested
- Assert — check what came out
Here is the cart tested again — this time through what it promises, not how it stores things:
from cart import Cart
def test_an_empty_cart_totals_zero():
cart = Cart()
assert cart.total() == 0
def test_total_is_price_times_quantity_summed_over_items():
# Arrange
cart = Cart()
cart.add("pen", 15.0, 2)
cart.add("bag", 850.0)
# Act
total = cart.total()
# Assert
assert total == 880.0
def test_adding_the_same_item_twice_adds_up_the_quantity():
cart = Cart()
cart.add("pen", 15.0, 2)
cart.add("pen", 15.0, 1)
assert cart.total() == 45.0Against the list-based cart, and again against the dictionary-based one:
... [100%]
3 passed in 0.01s
... [100%]
3 passed in 0.01sSame tests, two implementations, all green both times. That is the property you want: a refactor that keeps behaviour should keep the tests green. The comments in the middle test are only there to show the shape; most people drop them and separate the three parts with blank lines, as in the third test.
Why does it matter that the Act is a single step? Because when the test fails, you want to know which thing broke. If the Act is five calls, a failure points at five suspects.
Test behaviour, not implementation
The rule of thumb: a test may use only what a caller may use. For the cart that is Cart(), .add() and .total(). Not _items, not _lines. The leading underscore is Python's way of saying "this is mine, and I may change it".
The wrong way couples the test to a decision that was never part of the contract — that items are kept in a list. The right way asks the question a caller would ask: if I add these things, what is the total?
Two signs that a test is coupled to implementation:
- It reads attributes that start with
_, or mocks a function the code calls internally - A refactor that no user could notice turns it red
Mocks deserve a note here. Chapter eleven's mocks are the right tool at a boundary — the network, the clock, a payment service. Used inside your own code ("assert that _calculate was called with these arguments"), they become implementation tests by another name.
Names that read as sentences, and one reason to fail
Run the three cart tests with -v:
test_cart_behaviour.py::test_an_empty_cart_totals_zero PASSED [ 33%]
test_cart_behaviour.py::test_total_is_price_times_quantity_summed_over_items PASSED [ 66%]
test_cart_behaviour.py::test_adding_the_same_item_twice_adds_up_the_quantity PASSED [100%]Read the names alone and you have a specification of the cart. test_add and test_total told you only which method was touched. A good name says the situation and the expected result: an empty cart totals zero. It is long; that is fine. You never type it, and you read it every time it fails.
The second half of the rule is one reason to fail. Here is the opposite — one test that checks everything, run against a cart with a bug (it ignores quantity):
from cart import Cart
def test_cart():
cart = Cart()
assert cart.total() == 0
cart.add("pen", 15.0, 2)
cart.add("bag", 850.0)
assert cart.total() == 880.0
cart.add("pen", 15.0, 1)
assert cart.total() == 895.0$ pytest -q --tb=no
F [100%]
=========================== short test summary info ============================
FAILED test_one_big.py::test_cart - assert 865.0 == 880.0
1 failed in 0.01s(--tb=no hides the tracebacks and leaves only the summary — the part you read first.) test_cart failed — which says nothing — with 865.0 == 880.0, which you now have to puzzle over. And the third assertion never ran, so you do not know whether it would have passed. The same bug against the three focused tests:
$ pytest -q --tb=no
.FF [100%]
=========================== short test summary info ============================
FAILED test_cart_behaviour.py::test_total_is_price_times_quantity_summed_over_items
FAILED test_cart_behaviour.py::test_adding_the_same_item_twice_adds_up_the_quantity
2 failed, 1 passed in 0.01sBefore opening a single traceback, the summary already reads like a diagnosis: the empty cart is fine, but quantities are going wrong. "One reason to fail" does not mean one assert line — two asserts that check two sides of the same result are fine. It means one behaviour per test.
The test pyramid
Tests come in sizes:
- Unit tests check one function or class with nothing real around it. Milliseconds each. You have thousands.
- Integration tests check that pieces fit: your code against a real database, a real file system, a real HTTP app via a test client. Slower; you have dozens to hundreds.
- End-to-end tests drive the whole system the way a user would — a browser, a deployed API. Slow and fragile; you have a handful, covering the paths that make money.
Drawn by count, that is a pyramid: wide base of unit tests, thin top of end-to-end ones. The reason is cost. When a unit test fails, it points at a few lines. When an end-to-end test fails, it points at the whole system. Push each check as far down the pyramid as it can go, and keep the top for what only the top can see.
Red, green, refactor
Test-driven development (TDD) runs in a tight loop:
- Red — write one small test for behaviour that does not exist yet, and watch it fail
- Green — write the least code that makes it pass
- Refactor — tidy the code, with the tests green the whole time
Why watch it fail? Because a test you have never seen fail might be testing nothing. The red step proves the test can catch the absence of the feature.
We take the plan from the table, one row at a time.
Red. The first row, before slugs.py exists:
from slugs import slugify
def test_lowercases_and_joins_words_with_hyphens():
assert slugify("Hello World") == "hello-world"==================================== ERRORS ====================================
________________________ ERROR collecting test_slugs.py ________________________
ImportError while importing test module '/home/you/blog/test_slugs.py'.
Hint: make sure your test modules/packages have valid Python names.
Traceback:
/usr/lib/python3.12/importlib/__init__.py:90: in import_module
return _bootstrap._gcd_import(name[level:], package, level)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
test_slugs.py:1: in <module>
from slugs import slugify
E ModuleNotFoundError: No module named 'slugs'
=========================== short test summary info ============================
ERROR test_slugs.py
!!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!!
1 error in 0.01sThat is a perfectly good red. The test is asking for something that is not there.
Green. The least code that passes — and it really is the least:
def slugify(title):
return title.lower().replace(" ", "-"). [100%]
1 passed in 0.01sIt feels like cheating. It is deliberate: the code does only what a test has demanded so far. Anything more would be code no test is checking.
Red. The second row of the table:
def test_drops_punctuation():
assert slugify("Hello, World!") == "hello-world".F [100%]
=================================== FAILURES ===================================
____________________________ test_drops_punctuation ____________________________
def test_drops_punctuation():
> assert slugify("Hello, World!") == "hello-world"
E AssertionError: assert 'hello,-world!' == 'hello-world'
E
E - hello-world
E + hello,-world!
E ? + +
test_slugs.py:9: AssertionError
=========================== short test summary info ============================
FAILED test_slugs.py::test_drops_punctuation - AssertionError: assert 'hello,...
1 failed, 1 passed in 0.01sThe ? line marks exactly the two characters that should not be there.
Green. Now the replace-one-space trick is not enough. Replace every run of "not a letter or digit" with one hyphen, then trim hyphens from the ends:
import re
def slugify(title):
return re.sub(r"[^a-z0-9]+", "-", title.lower()).strip("-").. [100%]
2 passed in 0.01sThe rows for extra spaces and digits would already pass with this code. Add them anyway — they are part of the contract — but know that a test that is green the first time it runs has not proved anything about the code you just wrote. Only a red-then-green test has.
Red. Accented letters:
def test_turns_accented_letters_into_plain_ones():
assert slugify("Café au lait") == "cafe-au-lait"..F [100%]
=================================== FAILURES ===================================
_________________ test_turns_accented_letters_into_plain_ones __________________
def test_turns_accented_letters_into_plain_ones():
> assert slugify("Café au lait") == "cafe-au-lait"
E AssertionError: assert 'caf-au-lait' == 'cafe-au-lait'
E
E - cafe-au-lait
E ? -
E + caf-au-lait
test_slugs.py:13: AssertionError
=========================== short test summary info ============================
FAILED test_slugs.py::test_turns_accented_letters_into_plain_ones - Assertion...
1 failed, 2 passed in 0.01sé is not in a-z, so it was thrown away with the punctuation.
Green. Unicode "NFKD" normalisation splits é into e plus a separate accent mark; encoding to ASCII with "ignore" then drops the mark and keeps the e:
import re
import unicodedata
def slugify(title):
plain = unicodedata.normalize("NFKD", title).encode("ascii", "ignore").decode()
return re.sub(r"[^a-z0-9]+", "-", plain.lower()).strip("-")... [100%]
3 passed in 0.01sRefactor. It works, but the body is one dense line. Give each step a name, compile the pattern once, and add type hints — changing no behaviour:
import re
import unicodedata
NOT_ALLOWED = re.compile(r"[^a-z0-9]+")
def _to_ascii(text: str) -> str:
"""Split accented letters into letter + accent, then drop the accents."""
decomposed = unicodedata.normalize("NFKD", text)
return decomposed.encode("ascii", "ignore").decode("ascii")
def slugify(title: str) -> str:
words = NOT_ALLOWED.sub("-", _to_ascii(title).lower())
return words.strip("-")test_slugs.py::test_lowercases_and_joins_words_with_hyphens PASSED [ 33%]
test_slugs.py::test_drops_punctuation PASSED [ 66%]
test_slugs.py::test_turns_accented_letters_into_plain_ones PASSED [100%]
============================== 3 passed in 0.01s ===============================This is the step that makes TDD pay. You could restructure freely because the tests check behaviour; had they reached into _to_ascii or the regex, the refactor would have broken them.
Property-based testing with Hypothesis
Every test so far has one invented example: "Hello World", "Café au lait". The examples are only as good as your imagination, and your imagination shares its blind spots with the code you just wrote.
A property is a statement that holds for every input. Hypothesis generates the inputs — a hundred per test by default — and tries hard to find one that breaks the statement. @given says where the inputs come from; strategies (always imported as st) describes their shape:
from hypothesis import given
from hypothesis import strategies as st
@given(
prices=st.lists(st.integers(min_value=0, max_value=10_000), max_size=20),
discount=st.integers(min_value=0, max_value=100),
)
def test_a_discount_never_raises_the_total(prices, discount):
total = sum(prices)
discounted = total * (100 - discount) // 100
assert 0 <= discounted <= total. [100%]
1 passed in 0.01sOne dot, but a hundred lists of prices and a hundred discounts went through it. Other strategies you will reach for: st.text(), st.floats(), st.booleans(), st.sampled_from([...]), st.dictionaries(...), st.builds(...).
What properties does slugify have? You cannot say what the slug of a random string is, but you can say what it must look like, and that slugifying a slug changes nothing:
import re
from hypothesis import given
from hypothesis import strategies as st
from slugs import slugify
@given(st.text())
def test_a_slug_only_holds_lowercase_letters_digits_and_inner_hyphens(title):
slug = slugify(title)
assert re.fullmatch(r"[a-z0-9]+(-[a-z0-9]+)*|", slug)
@given(st.text())
def test_slugifying_a_slug_changes_nothing(title):
once = slugify(title)
assert slugify(once) == once.. [100%]
2 passed in 0.01sst.text() produces emoji, Chinese, control characters, empty strings — inputs you would never have typed into the plan table.
A round trip finds a real bug
The most productive property is the round trip: if you can encode something and decode it again, decoding the encoding must give back the original. Here is a run-length encoder — "aaab" becomes "3a1b" — with two example tests that pass:
import re
def encode(text: str) -> str:
"""'aaab' -> '3a1b': each run of a character becomes count + character."""
out = []
for match in re.finditer(r"(.)\1*", text, flags=re.DOTALL):
run = match.group(0)
out.append(f"{len(run)}{run[0]}")
return "".join(out)
def decode(encoded: str) -> str:
return "".join(char * int(count) for count, char in re.findall(r"(\d+)(\D)", encoded))from hypothesis import given
from hypothesis import strategies as st
from rle import decode, encode
def test_encode_counts_each_run():
assert encode("aaab") == "3a1b"
def test_decode_expands_each_run():
assert decode("3a1b") == "aaab"
@given(st.text())
def test_decoding_an_encoding_gives_back_the_original(text):
assert decode(encode(text)) == text..F [100%]
=================================== FAILURES ===================================
______________ test_decoding_an_encoding_gives_back_the_original _______________
@given(st.text())
> def test_decoding_an_encoding_gives_back_the_original(text):
^^^
test_rle.py:16:
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _
text = '0'
@given(st.text())
def test_decoding_an_encoding_gives_back_the_original(text):
> assert decode(encode(text)) == text
E AssertionError: assert '' == '0'
E
E - 0
E Failing test case: test_decoding_an_encoding_gives_back_the_original(
E text='0',
E )
test_rle.py:17: AssertionError
=========================== short test summary info ============================
FAILED test_rle.py::test_decoding_an_encoding_gives_back_the_original - Asser...
1 failed, 2 passed in 0.01sThe text "0" encodes to "10" — "one zero" — and the decoder reads 10 as a count with no character after it. Any text containing a digit is corrupted. Both example tests used only letters, so they could never see it.
Shrinking
Hypothesis did not first stumble on "0". It found some long, ugly string, then shrank it: tried smaller and simpler inputs, keeping each one that still failed, until nothing simpler failed. You can watch that happen by recording each failing input:
from hypothesis import given, seed, settings
from hypothesis import strategies as st
from rle import decode, encode
failing = []
@seed(2026)
@settings(database=None)
@given(st.text())
def round_trip(text):
if decode(encode(text)) != text:
if text not in failing:
failing.append(text)
raise AssertionError
try:
round_trip()
except AssertionError:
pass
for text in failing:
print(repr(text))'wfê\x9f\x0c\U000b6082Ñ/\x879\x9dÂ\n'
'\x03\U000a1acb\U0003c4fc1'
'´\U000cd3dc\U00081acfÔ÷\U0007122f\x89ê½5'
'l𨺮\nÝÛ\x894\x8b'
'\U00083189𤾨«ñï\xa0^ó8A\x96\x88=\x05'
'0000'
'000'
'00'
'0'(@seed fixes the random choices so this run is repeatable; database=None stops Hypothesis replaying a remembered failure.) The first failure is thirteen characters of noise with a 9 hidden in it. You would have stared at it for minutes. The shrunk example says the whole story in one character: a digit breaks it.
Hypothesis also saves failing examples in a .hypothesis/ folder, so the next run tries "0" first. Add that folder to .gitignore.
The fix puts a separator between count and character, so a digit can never be mistaken for part of a count. And the example Hypothesis found is pinned with @example, so it runs every time, on every machine:
import re
def encode(text: str) -> str:
"""'aaab' -> '3:a1:b': each run becomes count, a colon, then the character."""
out = []
for match in re.finditer(r"(.)\1*", text, flags=re.DOTALL):
run = match.group(0)
out.append(f"{len(run)}:{run[0]}")
return "".join(out)
def decode(encoded: str) -> str:
pairs = re.findall(r"(\d+):(.)", encoded, flags=re.DOTALL)
return "".join(char * int(count) for count, char in pairs)The two example tests change to the new format ("3:a1:b"), and the property gains one line:
from hypothesis import example, given
from hypothesis import strategies as st
from rle import decode, encode
@given(st.text())
@example("0") # the input Hypothesis found; now it is checked on every run
def test_decoding_an_encoding_gives_back_the_original(text):
assert decode(encode(text)) == text... [100%]
3 passed in 0.01sProperties do not replace examples. The two example tests document the format in a way a reader understands at a glance; the property guards the corners nobody thought of. Use both.
Flaky tests
A flaky test passes and fails on the same code. It is worse than no test: people learn to press "re-run" and stop believing red. Nearly every flaky test comes from one of four causes, and each has a fix that removes the cause rather than hiding it.
Shared state and test order
_users = set()
def register(name):
_users.add(name)
def count():
return len(_users)from registry import count, register
def test_registering_a_user_counts_them():
register("rahim")
assert count() == 1
def test_a_new_registry_is_empty():
assert count() == 0$ pytest -q --tb=no
.F [100%]
=========================== short test summary info ============================
FAILED test_registry.py::test_a_new_registry_is_empty - assert 1 == 0
1 failed, 1 passed in 0.01s
$ pytest -q "test_registry.py::test_a_new_registry_is_empty"
. [100%]
1 passed in 0.01sFails in the suite, passes alone. The module-level set survives from one test to the next, so the second test sees what the first left behind. Change the order, or run tests in parallel with pytest-xdist, and the result changes. The wrong fix is to reorder the tests. The right fix removes the shared state: make the registry an object, and give each test a fresh one through a fixture:
class Registry:
def __init__(self):
self._users = set()
def register(self, name):
self._users.add(name)
def count(self):
return len(self._users)import pytest
from registry import Registry
@pytest.fixture
def registry():
return Registry() # a fresh one for every test
def test_registering_a_user_counts_them(registry):
registry.register("rahim")
assert registry.count() == 1
def test_a_new_registry_is_empty(registry):
assert registry.count() == 0.. [100%]
2 passed in 0.01sWhen you cannot change the code, an autouse fixture that resets the state before and after each test is the fallback. Every test should pass alone, and in any order.
Time
from datetime import datetime
def greeting(now=None):
now = now or datetime.now()
return "Good morning" if now.hour < 12 else "Good afternoon"from greeting import greeting
def test_greets_the_morning():
assert greeting() == "Good morning" # true only before noonThe same code, run at the same moment, on two machines set to different time zones:
$ TZ=Europe/London pytest -q
. [100%]
1 passed in 0.01s
$ TZ=Asia/Dhaka pytest -q --tb=no
F [100%]
=========================== short test summary info ============================
FAILED test_greeting.py::test_greets_the_morning - AssertionError: assert 'Go...
1 failed in 0.01sThe fix is already in the function's signature: now can be passed in. A test that controls the clock tests both sides of the boundary, and gives the same answer at any hour:
from datetime import datetime
from greeting import greeting
def test_before_noon_it_says_good_morning():
assert greeting(now=datetime(2026, 1, 5, 9, 30)) == "Good morning"
def test_from_noon_on_it_says_good_afternoon():
assert greeting(now=datetime(2026, 1, 5, 12, 0)) == "Good afternoon".. [100%]
2 passed in 0.01sPassing the time in is simpler than any mock. When you cannot change the signature, monkeypatch from chapter ten replaces the clock instead.
Randomness
import random
def pick_winner(names, rng=random):
return rng.choice(names)from raffle import pick_winner
def test_picks_a_winner():
assert pick_winner(["rahim", "karim", "salma"]) == "rahim"Six runs, no change to anything:
$ for i in 1 2 3 4 5 6; do pytest -q | tail -1; done
1 failed in 0.01s
1 failed in 0.01s
1 failed in 0.01s
1 passed in 0.01s
1 passed in 0.01s
1 passed in 0.01sTwo fixes, for two kinds of question. Assert what is true for every outcome — the winner is one of the entrants. Or take control of the randomness by passing a seeded generator:
import random
from raffle import pick_winner
def test_the_winner_is_one_of_the_entrants():
names = ["rahim", "karim", "salma"]
assert pick_winner(names) in names
def test_the_same_seed_picks_the_same_winner():
names = ["rahim", "karim", "salma"]
first = pick_winner(names, rng=random.Random(42))
second = pick_winner(names, rng=random.Random(42))
assert first == second$ for i in 1 2 3; do pytest -q | tail -1; done
2 passed in 0.01s
2 passed in 0.01s
2 passed in 0.01sThe outside world
The fourth cause is anything you do not control: a real network call, a real server, a sleep(0.1) that "should be enough". The fixes are the ones from earlier chapters — mock at the boundary, use tmp_path instead of a shared folder, wait for a condition instead of a fixed time. One shape covers all four causes: whatever the test depends on, the test should create or control.
A regression test for every bug
Back to the question mark in the plan table. Nobody answered it, and the bug report arrives: a post titled `"!!!"` was saved at `/posts/`, overwriting the index page.
Before touching slugify, write a test that reproduces the report and watch it fail. That proves the test catches this bug:
import pytest
from slugs import slugify
def test_a_title_with_no_letters_or_digits_is_rejected():
# Bug: slugify("!!!") returned "", and the post was saved at /posts/
with pytest.raises(ValueError, match="no letters or digits"):
slugify("!!!")$ pytest -q test_slugs_regressions.py
F [100%]
=================================== FAILURES ===================================
______________ test_a_title_with_no_letters_or_digits_is_rejected ______________
def test_a_title_with_no_letters_or_digits_is_rejected():
# Bug: slugify("!!!") returned "", and the post was saved at /posts/
> with pytest.raises(ValueError, match="no letters or digits"):
E Failed: DID NOT RAISE ValueError
test_slugs_regressions.py:8: Failed
=========================== short test summary info ============================
FAILED test_slugs_regressions.py::test_a_title_with_no_letters_or_digits_is_rejected
1 failed in 0.01sThen fix it — the end of slugify becomes:
def slugify(title: str) -> str:
slug = NOT_ALLOWED.sub("-", _to_ascii(title).lower()).strip("-")
if not slug:
raise ValueError(f"title has no letters or digits: {title!r}")
return slug.... [100%]
4 passed in 0.01sThe regression test and the three example tests pass. But run the whole suite, and the property tests object:
FAILED test_slugs_properties.py::test_a_slug_only_holds_lowercase_letters_digits_and_inner_hyphens
FAILED test_slugs_properties.py::test_slugifying_a_slug_changes_nothing - Val...
2 failed, 4 passed in 0.01sHypothesis tried the empty string, and slugify("") now raises. That is not a new bug — it is the contract changing on purpose, and the properties noticing. Good. Update them to state the new contract: titles that contain at least one letter or digit.
import re
import string
from hypothesis import given
from hypothesis import strategies as st
from slugs import slugify
# Any text, with at least one plain letter or digit somewhere inside it.
titles = st.builds(
lambda before, word, after: before + word + after,
st.text(),
st.text(alphabet=string.ascii_letters + string.digits, min_size=1),
st.text(),
)
@given(titles)
def test_a_slug_only_holds_lowercase_letters_digits_and_inner_hyphens(title):
slug = slugify(title)
assert re.fullmatch(r"[a-z0-9]+(-[a-z0-9]+)*", slug)test_slugifying_a_slug_changes_nothing switches to @given(titles) in the same way.
...... [100%]
6 passed in 0.01sWhy a test for every bug? Because a bug that happened once has proved it is easy to make, and the next person to touch that code is likely to make it again. The comment saying what went wrong is part of the test: in a year, it is the only record of why this odd-looking case is here.
A complete example
A password-strength checker, built with everything in this chapter. The plan first:
| Case | Input | Expected | |---|---|---| | empty | "" | weak | | every kind, too short | "Ab1!xyz" (7) | weak | | long enough, one kind | "abcdefgh" | weak | | long enough, two kinds | "abcdefg1" | medium | | 12 long, three kinds | "abcdefghij1!" | strong | | 12 long, two kinds | "abcdefghijk1" | medium |
passwords.py:
def _kinds(password: str) -> int:
"""How many of the four kinds of character the password uses."""
return sum([
any(c.islower() for c in password),
any(c.isupper() for c in password),
any(c.isdigit() for c in password),
any(not c.isalnum() for c in password),
])
def strength(password: str) -> str:
"""Rate a password as 'weak', 'medium' or 'strong'."""
if len(password) < 8:
return "weak"
kinds = _kinds(password)
if len(password) >= 12 and kinds >= 3:
return "strong"
if kinds >= 2:
return "medium"
return "weak"test_passwords.py:
import pytest
from hypothesis import given
from hypothesis import strategies as st
from passwords import strength
RANK = {"weak": 0, "medium": 1, "strong": 2}
@pytest.mark.parametrize(
"password, expected",
[
("", "weak"), # nothing at all
("Ab1!xyz", "weak"), # every kind, but only 7 long
("abcdefgh", "weak"), # 8 long, one kind
("abcdefg1", "medium"), # 8 long, two kinds
("abcdefghij1!", "strong"), # 12 long, three kinds
("abcdefghijk1", "medium"), # 12 long, only two kinds
],
)
def test_strength_follows_the_length_and_variety_rules(password, expected):
assert strength(password) == expected
@given(st.text(), st.text())
def test_adding_characters_never_makes_a_password_weaker(password, extra):
before = strength(password)
after = strength(password + extra)
assert RANK[after] >= RANK[before]
def test_a_long_password_of_one_kind_is_still_weak():
# Bug: "aaaaaaaaaaaaaaaaaaaa" (20 letters) was rated "medium"
assert strength("a" * 20) == "weak"$ pytest -v
collected 8 items
test_passwords.py::test_strength_follows_the_length_and_variety_rules[-weak] PASSED [ 12%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[Ab1!xyz-weak] PASSED [ 25%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefgh-weak] PASSED [ 37%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefg1-medium] PASSED [ 50%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefghij1!-strong] PASSED [ 62%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefghijk1-medium] PASSED [ 75%]
test_passwords.py::test_adding_characters_never_makes_a_password_weaker PASSED [ 87%]
test_passwords.py::test_a_long_password_of_one_kind_is_still_weak PASSED [100%]
============================== 8 passed in 0.01s ===============================Three decisions are worth noticing.
The table rows sit on the boundaries — 7 and 8 characters, 11 and 12, two kinds and three. Bugs live at boundaries: a < that should be <= is invisible in the middle of a range.
The property says something no example can: adding characters never makes a password weaker. Nobody would write that test by hand for every password, and Hypothesis checks it against hundreds, including Unicode letters that are neither upper nor lower case.
And no test touches _kinds. It is reached through strength; tomorrow it may be gone.
When it breaks
AttributeError: 'Cart' object has no attribute '_items' after a refactor The test was reading the inside of the object. Rewrite it to check through the public methods — what a caller can see — and it will survive the next refactor too.
A test passes alone but fails in the full run (or the other way round) Shared state: a module-level list, dictionary or cache; a file in a fixed location; an environment variable set and never unset. Make each test create what it needs (fixtures, tmp_path, monkeypatch) instead of reordering the tests.
hypothesis.errors.FailedHealthCheck: It looks like this test is filtering out a lot of inputs. 0 inputs were generated successfully, while 50 inputs were filtered out. A .filter() (or assume()) throws away nearly everything Hypothesis generates — for example st.integers().filter(lambda n: n % 1000 == 7). Build the values you want directly instead: st.integers().map(lambda n: n * 1000 + 7).
hypothesis.errors.FlakyFailure: Hypothesis test_depends_on_earlier_runs(n=-25617) produces unreliable results: Failed on the first call but did not on a subsequent one The test gave a different answer for the same input. Something outside the generated arguments changed between calls — a counter, a module-level list, the clock. Hypothesis replays each failure, so a test with hidden state cannot shrink.
hypothesis.errors.DeadlineExceeded: Test took 300.07ms, which exceeds the deadline of 200.00ms. Each generated example must finish within 200 ms by default, because a hundred slow examples make a slow suite. Make the test faster, or, if the slowness is genuine, raise it with @settings(deadline=...).
A new TDD test passes the first time you run it It is not proving anything about the code you are about to write. Either the behaviour already exists (fine — keep the test as documentation), or the test is not checking what you think. Break the code on purpose for a moment and confirm the test goes red.
Step 4 of 6 — Predict
Check your understanding
Run as a whole file, the second test fails. What is the last line when only the second test is run on its own?
_users = set()
def register(name):
_users.add(name)
def count():
return len(_users)
def test_registering_a_user_counts_them():
register("rahim")
assert count() == 1
def test_a_new_registry_is_empty():
assert count() == 0
# Run only the second test:
# pytest -q "test_registry.py::test_a_new_registry_is_empty"- A1 failed in 0.01s
- B1 passed in 0.01s
- C1 failed, 1 passed in 0.01s
- D1 error in 0.01s
Cart switches from a list to a dictionary inside, and its behaviour stays the same. Which test turns red?
def test_total_of_two_pens():
cart = Cart()
cart.add("pen", 15.0, 2)
assert cart.total() == 30.0
def test_empty_cart_totals_zero():
assert Cart().total() == 0
def test_one_line_per_item():
cart = Cart()
cart.add("pen", 15.0)
assert len(cart._items) == 1
def test_adding_twice_adds_up():
cart = Cart()
cart.add("pen", 15.0)
cart.add("pen", 15.0)
assert cart.total() == 30.0- A`test_total_of_two_pens`
- B`test_empty_cart_totals_zero`
- C`test_one_line_per_item`
- D`test_adding_twice_adds_up`
This test passes Monday to Friday and fails on Saturday and Sunday. What is the best fix?
from datetime import date
def is_weekend(day=None):
day = day or date.today()
return day.weekday() >= 5
def test_weekdays_are_not_weekend():
assert is_weekend() is False- ARe-run the test automatically whenever it fails
- BSkip the test on Saturdays and Sundays
- CWrite `== False` instead of `is False`
- DPass a fixed date — `is_weekend(date(2026, 10, 7))` — and add a separate test for a Saturday
Answering needs an account
Sign in to check your answers
The questions are above, and working them out in your head is the part that matters. Sign in to see the answers, the explanations and the three-level hints.
Your turn
Build two functions, test-first, in durations.py:
parse_duration(text)turns"1h30m"into5400(seconds). Units areh,m,s, in that order, each optional;"45s","2h"and"1h0m5s"are all valid. Anything else —"","h","1x","30m1h","1.5h"— raisesValueErrorwithnot a durationin the message.format_duration(seconds)goes the other way:5400becomes"1h30m",3605becomes"1h5s",0becomes"0s". A negative number raisesValueError.
Work in this order:
- Write the plan table first: case, input, expected — including the invalid inputs and the boundaries (
59and60seconds, for example). - Go row by row in red-green cycles. Run
pytestafter every change and look at the red before making it green. - Add the round-trip property with Hypothesis: for any whole number of seconds from
0upwards,parse_duration(format_duration(n)) == n. - A user reports that
"90m"was rejected, though people type it all the time. Add a regression test with a comment saying what went wrong.
When you are done, refactor once, and confirm the whole suite stays green.
Solution
durations.py:
import re
UNITS = {"h": 3600, "m": 60, "s": 1}
PATTERN = re.compile(r"(?:(\d+)h)?(?:(\d+)m)?(?:(\d+)s)?")
def parse_duration(text: str) -> int:
"""'1h30m' -> 5400. Units must appear in the order h, m, s."""
match = PATTERN.fullmatch(text)
if not text or match is None:
raise ValueError(f"not a duration: {text!r}")
hours, minutes, seconds = (int(part or 0) for part in match.groups())
return hours * 3600 + minutes * 60 + seconds
def format_duration(seconds: int) -> str:
"""5400 -> '1h30m'. Zero is '0s'."""
if seconds < 0:
raise ValueError(f"duration cannot be negative: {seconds}")
if seconds == 0:
return "0s"
parts = []
for unit, size in UNITS.items():
amount, seconds = divmod(seconds, size)
if amount:
parts.append(f"{amount}{unit}")
return "".join(parts)test_durations.py:
import pytest
from hypothesis import given
from hypothesis import strategies as st
from durations import format_duration, parse_duration
@pytest.mark.parametrize(
"text, seconds",
[
("45s", 45),
("2h", 7200),
("1h30m", 5400),
("1h0m5s", 3605),
("0s", 0),
],
)
def test_parse_reads_hours_minutes_and_seconds(text, seconds):
assert parse_duration(text) == seconds
@pytest.mark.parametrize("text", ["", "h", "1x", "30m1h", "1.5h"])
def test_parse_rejects_text_that_is_not_a_duration(text):
with pytest.raises(ValueError, match="not a duration"):
parse_duration(text)
@pytest.mark.parametrize(
"seconds, text",
[(0, "0s"), (59, "59s"), (60, "1m"), (3605, "1h5s"), (5400, "1h30m")],
)
def test_format_writes_only_the_units_it_needs(seconds, text):
assert format_duration(seconds) == text
def test_format_rejects_a_negative_duration():
with pytest.raises(ValueError, match="negative"):
format_duration(-1)
@given(st.integers(min_value=0, max_value=10**7))
def test_parsing_a_formatted_duration_gives_back_the_seconds(seconds):
assert parse_duration(format_duration(seconds)) == seconds
def test_minutes_over_sixty_are_accepted():
# Bug: "90m" was rejected, though people type it all the time
assert parse_duration("90m") == 5400$ pytest -q
.................. [100%]
18 passed in 0.01sWhy it is built this way:
- The plan became the parametrize tables. Each row of the table is one row of
parametrize, so adding a case later is one line, and the test names in-voutput show which row failed. - Valid and invalid inputs are separate tests. They are separate behaviours with separate reasons to fail. The invalid list holds one example of each kind of mistake: empty, unit without a number, unknown unit, wrong order, decimal.
match="not a duration"checks the message as well as the type, so aValueErrorraised by accident somewhere else —int(""), say — cannot make the test pass.if not textis there because the pattern, with every part optional, matches the empty string. That is exactly the kind of case the plan table forces you to think about before the regex does it for you.59and60are the boundary where seconds roll over into minutes; a test in the middle of the range would never catch an off-by-one there.- The round-trip property is the strongest test in the file. It does not know a single expected value, yet it checks ten million possible durations against each other — any disagreement between the two functions shows up, shrunk to the smallest number that breaks it.
- The regression test passes already with this implementation; it exists to stop a future "tidy-up" that limits minutes to
0-59from bringing the bug back. Its comment says why. - Nothing tests
PATTERNorUNITSdirectly. They are implementation, and you are free to replace the regex with a hand-written parser tomorrow.
Where to go next
This is the end of the course. You can write tests, read their failures, structure them, isolate them, measure them, extend pytest itself, and — after this chapter — decide what is worth testing in the first place. Some directions from here:
- Make it a habit on real code. Take a project you already have and add tests to the part you are most afraid to change. Write the plan table first. The fear goes down as the tests go up.
- Go further with Hypothesis. Its documentation covers stateful testing (
RuleBasedStateMachine), which generates whole sequences of operations against an object, and finds bugs no single input could. - Mutation testing. Tools such as
mutmutmake small changes to your code —<to<=,+to-— and check that some test fails. A mutation that survives is a line your tests do not really check, however high the coverage number is. - Grow the CI set-up from chapter fourteen. Run the suite on every Python version you support, and make a failing test block the merge — a suite nobody is forced to keep green slowly stops being green.
- Read good test suites. The tests of
pytest,requests,attrsandhypothesisitself are open source. Reading how experienced people name, structure and isolate tests teaches more than any rule list.
Whatever you write next, start with the question this chapter opened with: if someone improves this code tomorrow without changing what it does, will my tests stay green — and if someone breaks it, will they go red?
Step 6 of 6
Stretch — the chapter quiz
Ten questions from easy to hard. The last ones are difficult on purpose.
Sign in to take the quiz