Before you write a test — what you actually need
The thinking that comes before any test code: turning a vague requirement into a precise contract, making code testable, choosing cases with boundaries, knowing the expected answer without running the code, and deciding what not to test — all with plain Python and assert.
- 1Encounter
- 2Understand
- 3Worked
- 4Predict
- 5Apply
- 6Stretch
The problem we are solving
Here is a function that averages a list of exam scores, and the way most of us check a function like it — call it a couple of times and look:
def average(scores):
return sum(scores) / len(scores)
print(average([80, 90, 70]))
print(average([100]))80.0
100.0Both answers are right. The function "works", and it goes into the program. A week later a class with no results yet reaches the same line:
def average(scores):
return sum(scores) / len(scores)
print(average([]))Traceback (most recent call last):
File "/home/you/school/report.py", line 5, in <module>
print(average([]))
^^^^^^^^^^^
File "/home/you/school/report.py", line 2, in average
return sum(scores) / len(scores)
~~~~~~~~~~~~^~~~~~~~~~~~~
ZeroDivisionError: division by zero"It works when I tried it" was true. It only ever meant "it works for the two inputs I happened to try". Manual checking has two weaknesses that this exposes.
The first is that it covers only what you thought of in the moment, and an empty list is exactly the kind of thing nobody thinks of in the moment.
The second is that it does not repeat. The two print calls were run once, read once, and deleted. When the function changes next month, nobody runs them again.
Automated tests fix the second weakness, and the rest of this course is about writing them with pytest. But look at the first weakness more closely, because no tool fixes it. Suppose you wanted to write a check for the empty list right now. What would it say? Should average([]) return 0? Return None? Raise an error? Nobody decided. You cannot write a test for a behaviour nobody has decided, and you cannot check an answer you do not already know.
That is what this chapter is about: the thinking that has to happen before the first line of test code. Get it right and the tests almost write themselves. Skip it and you end up with tests that pass and still miss the bug.
By the end of this chapter you can
- Say what a test is in one sentence: a runnable claim about behaviour
- Turn a vague requirement into a precise contract — inputs, output, errors, side effects
- Recognise code that is hard to test, and reshape it into a pure core with a thin shell
- Plan cases with equivalence classes and boundary values, as a table
- Work out expected answers independently of the code — the oracle
- Decide what not to test, and when you have enough tests
- Write the plan as plain
assertchecks, run them, and say why they are not yet enough
Prerequisites: none — this is where the course starts. You need Python 3 and to know how to write a function. You do not need pytest yet; everything in this chapter runs with plain Python.
Before you write the test
Every chapter of this course has a section with this name, and every one of them answers the same six questions for its own topic. This chapter is where the questions come from, so here they are in one place. Keep them; the rest of the course uses them as a checklist.
- What exactly is promised? The contract: for these inputs, this output; for these inputs, this error; and these side effects, or none.
- Can the code be called by a test? Does it take its inputs as arguments and hand back a result — or does it read the keyboard, print, and look at the clock?
- Which cases? One from each kind of input, both sides of every edge, the invalid inputs, the empty and the very large.
- How do I know the right answer? Worked out by hand or from the specification — never by running the code under test.
- What must be in place? The Python version, an environment, where the test files will live — and, in later chapters, files, settings and tools the topic needs.
- What will I not test? Python itself, other people's libraries, code too simple to be wrong, and behaviour nobody promised.
Here they are applied to average from above, which is small enough to do in full.
The contract has a gap, so the first step is to close it. Suppose you ask, and the answer is: "an average of no scores makes no sense; raise ValueError". The function is already testable — it takes a list and returns a number. For this chapter, all that must be in place is Python 3. And the cases:
| Case | Input | Expected | Why | | --- | --- | --- | --- | | Typical | [80, 90, 70] | 80.0 | the ordinary use: (80 + 90 + 70) / 3 | | A single score | [100] | 100.0 | smallest list that has an average | | Empty | [] | ValueError | the case that broke; now decided |
What is not on the list: whether sum adds correctly, or whether / divides — that is Python's job, and it is tested far more thoroughly than anything you or I will write.
The rest of the chapter takes each of the six questions in turn, and ends by running the whole process on one realistic function.
What a test actually is
A test is a runnable claim about behaviour: given this input, expect this result. "Result" comes in exactly three kinds:
- a return value —
average([80, 90, 70])gives80.0 - an error —
average([])raisesValueError - a side effect — after
add_score(scores, 70), the listscoreshas70on the end
Python already has a statement for writing a claim down: assert. It does nothing when the condition is true and raises AssertionError when it is false. Here is one claim of each kind, with average now following its decided contract:
def average(scores):
if not scores:
raise ValueError("average() of an empty list")
return sum(scores) / len(scores)
def add_score(scores, new_score):
scores.append(new_score)
# 1. A return value: given this input, expect this output.
assert average([80, 90, 70]) == 80.0
# 2. An error: given this input, expect this exception.
try:
average([])
except ValueError:
pass
else:
raise AssertionError("average([]) should raise ValueError")
# 3. A side effect: given this call, expect this change to the world.
scores = [80, 90]
add_score(scores, 70)
assert scores == [80, 90, 70]
print("3 claims checked")3 claims checkedNotice what each claim needs: an input you choose, a result you already know, and a way to call the code and look at what it did. The six questions exist to make sure you have all three before you start. Notice too how clumsy the error check is — five lines of try/except/else to say "this should raise". pytest turns that into one line in chapter four.
1. The contract: from a vague requirement to a precise promise
Requirements usually arrive as a sentence: "Members get 10% off big orders." It sounds complete. Give it to two careful developers and watch:
# "Members get 10% off big orders." Two honest readings of the same sentence.
def discount_by_asha(total, is_member):
if is_member and total > 1000:
return total * 0.10
return 0
def discount_by_ravi(total, is_member):
if is_member and total >= 1000:
return round(total * 0.10)
return 0
for total in [500, 1000, 1234.56]:
print(total, discount_by_asha(total, True), discount_by_ravi(total, True))500 0 0
1000 0 100
1234.56 123.456 123Neither of them made a mistake. The sentence did not say whether exactly 1000 counts as "big", or how to round. Each filled the gap differently, and an order of 1000 gets a discount of 0 from one and 100 from the other. No test can say which of them is right, because "right" was never defined.
A contract is the requirement made precise enough to test. To get from the sentence to the contract, ask questions — the same ones every time:
- Inputs. What types? What range? Is there a limit, and is the limit itself included? What about zero, negative, empty,
None? - Output. What exactly is returned — the discount, or the new total? What type? Rounded to what?
- Errors. Which inputs are refused, and how — which exception?
- Side effects. Does it change anything — a list passed in, a file, a database — or print anything? Or nothing at all?
Asked about the discount, those questions produce answers, and the answers go where the code is — in the docstring:
def member_discount(total, is_member):
"""Return the discount on an order, in currency units.
- total: the order total before discount, a number >= 0.
A negative total raises ValueError.
- is_member: True or False.
- Members get 10% of the total when the total is 1000 or more
(1000 itself qualifies). Everyone else gets 0.
- The result is rounded to 2 decimal places.
- Returns the discount amount, not the new total.
- No side effects: reads nothing, prints nothing, changes nothing.
"""
if total < 0:
raise ValueError(f"total must not be negative, got {total}")
if is_member and total >= 1000:
return round(total * 0.10, 2)
return 0
print(member_discount(999.99, True))
print(member_discount(1000, True))
print(member_discount(1234.56, True))
print(member_discount(5000, False))
try:
member_discount(-5, True)
except ValueError as error:
print("ValueError:", error)0
100.0
123.46
0
ValueError: total must not be negative, got -5Every line of the docstring is now something a test can check, and every check has an obvious expected value. That is the test of a good contract: could a stranger write the expected value for any input, without reading the code? If not, there is still a question to ask.
Who answers the questions? Whoever owns the requirement — a product owner, a teacher, a client, or you. When there is nobody to ask, decide, write the decision into the contract, and move on. An explicit decision can be reviewed and changed; a silent one just lives in the code until it surprises someone.
2. Testable code: a pure core and a thin shell
A test calls a function with chosen inputs and compares what comes back. Some functions make that impossible. Here is one, for a café's happy hour — 20% off drinks between 17:00 and 19:00:
from datetime import datetime
def drink_price():
price = float(input("Price: "))
hour = datetime.now().hour
if 17 <= hour < 19:
price = price * 0.8
print(f"You pay {price:.2f}")It works when you run it and type a price. Now try to check it from another file, check_drinks.py:
from drinks import drink_price
result = drink_price()
print("returned:", result)An automated check runs with nobody at the keyboard. < /dev/null gives the program exactly that — an input with nothing in it:
$ python check_drinks.py < /dev/null
Price: Traceback (most recent call last):
File "/home/you/cafe/check_drinks.py", line 3, in <module>
result = drink_price()
^^^^^^^^^^^^^
File "/home/you/cafe/drinks.py", line 5, in drink_price
price = float(input("Price: "))
^^^^^^^^^^^^^^^^
EOFError: EOF when reading a lineFeed it a price, and it runs — but look at what comes back:
$ echo 10 | python check_drinks.py
Price: You pay 10.00
returned: NoneThree separate problems, each one common:
- It reads its input with
input()instead of taking it as an argument, so a test cannot choose the input. - It prints its result instead of returning it, so a test gets
Noneand has nothing to compare. - It reads the clock itself. That run was just before seven in the morning, so the answer was
10.00; the same command at half past five in the evening gives8.00. A check on it would pass or fail depending on when it ran — a flaky test, which is worse than none, because people learn to ignore it.
The fix is not a testing trick. It is a change of shape: separate the decision from the input and output. The decision becomes a pure function — everything it needs comes in as arguments, the answer goes out as the return value, and it touches nothing else. The keyboard, the screen and the clock move into a thin shell around it:
from datetime import datetime
def drink_price(price, hour):
"""Return the price to pay for a drink.
20% off from 17:00 up to, but not including, 19:00.
"""
if 17 <= hour < 19:
return round(price * 0.8, 2)
return price
def main():
# The thin shell: all the input, output and clock reading lives here.
price = float(input("Price: "))
print(f"You pay {drink_price(price, datetime.now().hour):.2f}")
if __name__ == "__main__":
main()Now the checks choose any hour they like, at any time of day:
from drinks import drink_price
assert drink_price(10, 16) == 10
assert drink_price(10, 17) == 8.0
assert drink_price(10, 18) == 8.0
assert drink_price(10, 19) == 10
print("all checks passed")all checks passedThe program behaves exactly as before for the person at the keyboard. What changed is that the part worth testing — the rule — can now be tested in isolation, and the shell is so thin there is almost nothing in it to get wrong.
The warning signs of hard-to-test code, so you can spot them before writing a test:
| Sign | Why it hurts a test | The usual fix | | --- | --- | --- | | input() inside the logic | the test cannot choose the input | take it as an argument | | print() as the only result | nothing comes back to compare | return the value; print in the shell | | datetime.now(), random inside | the answer changes between runs | pass the time or the random value in | | reads or changes a global variable | tests affect one another | pass it in, return the new value | | opens a fixed file or URL inside | the test needs that file or network | pass the data, or the path, in |
You cannot always reshape code — sometimes it belongs to someone else, or the I/O is the behaviour. pytest has tools for those cases: capsys captures printed output (chapter nine), tmp_path gives a throwaway folder (chapter nine), and monkeypatch and mocks replace the clock, the network and the environment (chapters ten and eleven). But reshaping comes first. Code that is easy to test is usually just easier to understand.
3. Choosing the cases
You cannot test every input; drink_price alone accepts every price at every hour. The job is to pick a small set where each case could catch a mistake the others would miss. Two ideas do most of that work.
Equivalence classes. Group the inputs that the contract treats the same way. Any one member of a group stands for the rest: if drink_price(10, 12) is right, drink_price(10, 11) almost certainly is too, because the same line of code handles both. Suppose the contract has been tightened to say that an hour outside 0–23, or a negative price, raises ValueError. Then the hours fall into five classes:
hour: ... -2 -1 | 0 1 ... 15 16 | 17 18 | 19 20 ... 23 | 24 25 ...
invalid | full price | 20% off | full price | invalidBoundary values. Mistakes do not spread evenly across a class. They gather at its edges, because that is where < and <= are chosen, and where "up to 19:00" is turned into code. So for every edge, test the value on each side of it: the last of one class and the first of the next. For a limit L that means looking at L - 1, L and L + 1; for counts and lengths, 0 and 1 as well.
Here is why "one typical value per class" is not enough. Below, the happy-hour test is written with an easy slip, 17 < hour instead of 17 <= hour:
def drink_price(price, hour):
# A slip: 17 < hour instead of 17 <= hour.
if 17 < hour < 19:
return round(price * 0.8, 2)
return price
# Three "typical" cases, one from each valid class:
assert drink_price(10, 12) == 10
assert drink_price(10, 18) == 8.0
assert drink_price(10, 21) == 10
print("typical cases passed")
# The boundary:
assert drink_price(10, 17) == 8.0
print("boundary passed")typical cases passed
Traceback (most recent call last):
File "/home/you/cafe/boundary.py", line 15, in <module>
assert drink_price(10, 17) == 8.0
^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionErrorOne case from each class: all pass, and customers who arrive at five o'clock pay full price. The boundary case catches it at once.
Classes and boundaries cover the shape of the input. Then run through these questions, each of which has found real bugs:
- Invalid input — what the contract says to refuse: a negative price, hour
24. - Empty and zero —
[],"",0, a price of0. Does "nothing" behave? None— only if the contract mentions it. If it does not, it is not a case (see question six).- Large values — a very long list, a very large number. Anything with a limit, test at the limit.
- Rounding — where a calculation produces more decimals than the result keeps.
Put together, that is a test plan: one row per case, written before any test code. The last column matters most — a row you cannot justify is a row you do not need.
| Case | Input (price, hour) | Expected | Why | | --- | --- | --- | --- | | Lowest valid hour | (10, 0) | 10 | edge of the valid range | | Last full-price hour | (10, 16) | 10 | below the 17:00 edge | | First discounted hour | (10, 17) | 8.0 | on the 17:00 edge | | Last discounted hour | (10, 18) | 8.0 | below the 19:00 edge | | First full price again | (10, 19) | 10 | on the 19:00 edge — "not including" | | Highest valid hour | (10, 23) | 10 | edge of the valid range | | Free drink | (0, 17) | 0 | zero price | | Rounding | (2.49, 17) | 1.99 | 2.49 × 0.8 = 1.992, rounded to 2 places | | Hour too low | (10, -1) | ValueError | just outside the range | | Hour too high | (10, 24) | ValueError | just outside the range | | Negative price | (-1, 12) | ValueError | invalid input |
And the plan turned into checks, against the tightened function:
def drink_price(price, hour):
"""Return the price to pay for a drink.
- price: a number >= 0. hour: a whole number 0-23.
- 20% off from 17:00 up to, but not including, 19:00.
- The result is rounded to 2 decimal places.
- An hour outside 0-23 or a negative price raises ValueError.
"""
if not 0 <= hour <= 23:
raise ValueError(f"hour must be 0-23, got {hour}")
if price < 0:
raise ValueError(f"price must not be negative, got {price}")
if 17 <= hour < 19:
return round(price * 0.8, 2)
return price
def raises_value_error(price, hour):
try:
drink_price(price, hour)
except ValueError:
return True
return False
assert drink_price(10, 0) == 10 # lowest valid hour
assert drink_price(10, 16) == 10 # last full-price hour before
assert drink_price(10, 17) == 8.0 # first discounted hour
assert drink_price(10, 18) == 8.0 # last discounted hour
assert drink_price(10, 19) == 10 # first full-price hour after
assert drink_price(10, 23) == 10 # highest valid hour
assert drink_price(0, 17) == 0 # free drink stays free
assert drink_price(2.49, 17) == 1.99 # rounding: 1.992 -> 1.99
assert raises_value_error(10, -1) # just below the valid range
assert raises_value_error(10, 24) # just above the valid range
assert raises_value_error(-1, 12) # negative price
print("11 checks passed")11 checks passedEleven rows, and each one is there for a reason you can say out loud.
4. The oracle: knowing the answer without the code
Every check compares the code's answer with an expected answer. The source of the expected answer is called the oracle, and it has one rule: it must not be the code under test.
That sounds too obvious to need saying, until you see how easily it is broken. Here member_discount has a bug — 1% instead of 10% — and two checks:
def member_discount(total, is_member):
if total < 0:
raise ValueError(f"total must not be negative, got {total}")
if is_member and total >= 1000:
return round(total * 0.01, 2) # bug: 1%, not 10%
return 0
# Wrong: the expected value comes from the code under test.
expected = member_discount(2000, True)
assert member_discount(2000, True) == expected
print("self-check passed, expected was", expected)
# Right: the expected value was worked out by hand. 10% of 2000 is 200.
assert member_discount(2000, True) == 200
print("hand-check passed")self-check passed, expected was 20.0
Traceback (most recent call last):
File "/home/you/shop/oracle.py", line 15, in <module>
assert member_discount(2000, True) == 200
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionErrorThe first check compares the function with itself. It passes for the correct function, for this broken one, and for any other version anyone could write — it cannot fail, so it tells you nothing.
Nobody writes it quite that baldly. The same mistake usually arrives in one of two disguises:
- Copying the output. You run the function, it prints
20.0, and you paste20.0into the check. If the code was already wrong, the check now guards the bug. - Repeating the formula. You write the expected value as
2000 * 0.01— the same sum the code does. A copy of the code's logic shares the code's mistakes. ("When it breaks", below, shows this one running.)
Good oracles, roughly from most to least common:
- Working it out by hand from the contract. Choose inputs that make this easy: 2000 rather than 1873.41, so that "10% of it" is something you can do in your head.
- Examples in the specification. If the requirement says "an order of 1500 gets 150 off", that row is a test.
- Known facts. Water boils at 100 °C and 212 °F; an empty list has length 0.
- A different, trusted method. A simpler, slower calculation, or a standard-library function that does the same job another way:
import statistics
def average(scores):
if not scores:
raise ValueError("average() of an empty list")
return sum(scores) / len(scores)
samples = [[80, 90, 70], [1, 2], [5], [10, 0, 0, 0]]
for scores in samples:
assert average(scores) == statistics.mean(scores), scores
print("agrees with statistics.mean on", len(samples), "samples")agrees with statistics.mean on 4 samplesstatistics.mean was written by other people, in a different way, so agreeing with it means something. (Comparing decimal results with == can bite — 0.1 + 0.2 == 0.3 is False in Python. The values here are safe; chapter five shows how to compare floats properly.)
If you cannot work out the expected answer by any of these means, that is not a testing problem. It means you do not yet know what the code should do — go back to the contract.
5. What must be in place
For this chapter, only Python 3. Check which one you have:
$ python3 --version
Python 3.12.3This course uses Python 3.12, and anything from 3.10 up will run nearly all of it. On Windows the command is usually py --version.
From the next chapter on, three more things have to be in place before a test can run, and chapter one sets each of them up:
- A virtual environment for the project, so that the testing tool is installed in the same Python that runs your code.
- pytest installed into it.
- A place for the tests. Test files sit next to the code, or in a
tests/folder, with names startingtest_. Chapter one starts with the first; chapter three moves to the second and explains why.
Later chapters add their own needs — a configuration file, a plugin, sample data on disk — and each one lists them in its own "Before you write the test".
6. What not to test, and when to stop
A test that cannot catch a mistake in your code costs time to write, time to run and time to maintain, and gives nothing back. Leave these out:
- Python itself. Not whether
sortedsorts orroundrounds. Test that your function sorts the right thing in the right order. - Other people's libraries. Not whether
requestssends a request orpandasreads a CSV. Test what your code does with the result — and if you need to be sure of a library's behaviour, a single small check that documents your assumption is enough. - Code too simple to be wrong. A function that returns a constant, or a class that only stores its arguments. If you cannot imagine the mistake, the test cannot catch it.
- Behaviour nobody promised. If the contract says nothing about passing a string to
drink_price, a test would only freeze whatever it happens to do today. Either add it to the contract — then test it — or leave it out. - How, rather than what. Not which helper the function called internally, or in what order it did its steps. Test what goes in and what comes out; then the inside can be rewritten freely, and the tests prove it still works.
And how many is enough? There is no number, but there is a rule. You have enough when:
- every equivalence class has at least one case,
- every boundary has a case on each side,
- every error the contract promises has a case,
- and every bug you have ever fixed has a case of its own, so it cannot quietly return.
Then stop. For each new case you are tempted to add, ask: what mistake would this catch that the others would not? If you cannot name one, it is a duplicate. drink_price(10, 11) next to drink_price(10, 12) catches nothing new.
A complete example
Now the whole process, start to finish, on one realistic function.
The requirement, as it arrived: "Shipping is 60 for local parcels and 120 for national ones. Heavier parcels cost more — 20 per extra kilo locally, 30 nationally. We don't take anything over 30 kg."
The questions, and the answers they got:
| Question | Answer | | --- | --- | | What does the base price cover? | The first kilogram — 1 kg exactly is still the base price. | | Is 1.2 kg charged as 1 kg or as 2? | Every started kilogram counts: 1.2 kg is charged as 2 kg. | | Is 30 kg itself accepted? | Yes. 30 is allowed; anything above is refused. | | What about 0 kg, or a negative weight? | Refused — that is a data-entry mistake. | | How are zones written? | Exactly "local" or "national". Anything else is refused, including "Local". | | What comes back? | The cost as a whole number. Nothing is printed or saved. |
The contract and the testable function, shipping.py. It takes everything as arguments and returns a number; there is no shell to separate, because there is no input or output in it at all:
import math
BASE = {"local": 60, "national": 120} # covers the first kilogram
PER_EXTRA_KG = {"local": 20, "national": 30}
MAX_WEIGHT_KG = 30
def shipping_cost(weight_kg, zone):
"""Return the shipping cost of one parcel, as a whole number.
- weight_kg: a number. At or below 0, or above 30: ValueError.
- zone: "local" or "national", exactly. Anything else: ValueError.
- The base price covers the first kilogram (1 kg included).
- Every started kilogram after that costs extra: 1.2 kg is charged as 2 kg.
- No side effects: reads nothing, prints nothing, changes nothing.
"""
if zone not in BASE:
raise ValueError(f"unknown zone: {zone!r}")
if not 0 < weight_kg <= MAX_WEIGHT_KG:
raise ValueError(f"weight must be more than 0 and at most 30 kg, got {weight_kg}")
extra_kg = math.ceil(weight_kg) - 1
return BASE[zone] + extra_kg * PER_EXTRA_KG[zone]math.ceil rounds up to the next whole number — math.ceil(1.2) is 2 — which is exactly "every started kilogram counts".
The case table. The weight classes are: invalid (0 and below), the base kilogram (above 0, up to 1), extra kilograms (above 1, up to 30), invalid (above 30). The zones are "local", "national", and anything else. Every expected value is worked out by hand from the contract, not from the code:
| Case | Input | Expected | Why | | --- | --- | --- | --- | | Light parcel | (0.5, "local") | 60 | inside the base kilogram | | Exactly 1 kg | (1, "local") | 60 | top edge of the base: 1 kg is included | | Just over 1 kg | (1.2, "local") | 80 | started kilogram counts: 60 + 1 × 20 | | Whole kilos, local | (3, "local") | 100 | 60 + 2 × 20 | | Whole kilos, national | (3, "national") | 180 | 120 + 2 × 30 — the other zone | | At the limit | (30, "national") | 990 | 120 + 29 × 30; 30 is accepted | | Zero | (0, "local") | ValueError | bottom edge: 0 is refused | | Negative | (-2, "local") | ValueError | invalid input | | Over the limit | (30.5, "local") | ValueError | just above 30 | | Wrong zone | (2, "Local") | ValueError | not exactly "local" |
What is deliberately not in the table: a weight of None or "2". The contract says the weight is a number, so those are outside it — Python raises its own TypeError when they are compared with 0, and we make no promise beyond that. Not math.ceil either: that is Python's, not ours. And not (2, "local") next to (3, "local"): same class, same line of code, nothing new to catch.
The checks, check_shipping.py — the table, row for row:
from shipping import shipping_cost
def raises_value_error(weight_kg, zone):
"""True if shipping_cost raises ValueError for these arguments."""
try:
shipping_cost(weight_kg, zone)
except ValueError:
return True
return False
# Valid weights: one case per class, and both sides of every edge.
assert shipping_cost(0.5, "local") == 60
assert shipping_cost(1, "local") == 60
assert shipping_cost(1.2, "local") == 80
assert shipping_cost(3, "local") == 100
assert shipping_cost(3, "national") == 180
assert shipping_cost(30, "national") == 990
# Invalid input: the contract promises a ValueError for each.
assert raises_value_error(0, "local")
assert raises_value_error(-2, "local")
assert raises_value_error(30.5, "local")
assert raises_value_error(2, "Local")
print("all 10 checks passed")$ python check_shipping.py
all 10 checks passedNow break it. Some months later, someone tidies shipping.py, drops import math and writes int(weight_kg) instead of math.ceil(weight_kg). It reads well, and it runs without an error. Run the checks again:
$ python check_shipping.py
Traceback (most recent call last):
File "/home/you/shop/check_shipping.py", line 14, in <module>
assert shipping_cost(0.5, "local") == 60
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionErrorThe plan did its job: the change was caught the moment the checks ran. int cuts the decimals off instead of rounding up, so 0.5 kg became int(0.5) - 1, which is -1 extra kilograms.
But read the report as if you did not already know that, and notice how little it says:
- What was the actual value? The report shows the line, not the number. You have to run
shipping_cost(0.5, "local")yourself to learn it was40. - What else is broken? The script stopped at the first failure. The 1.2 kg case is wrong too — it now gives
60instead of80— but that check never ran. - Which promises held? Nothing tells you that the error cases still pass.
- And the error checks needed a helper function with
try/exceptjust to say "this should raise".
The thinking was the hard part, and it is done: the contract, the table, the oracle. What is missing is a better runner — one that finds the checks on its own, runs every one even after a failure, shows the actual value next to the expected one, and checks errors in a single line. That is pytest, and installing it is the next chapter.
When it breaks
These are planning mistakes rather than Python errors, so the symptoms are quieter: usually a check that passes when it should not.
The expected value repeats the code's own formula The check is written by doing the same sum the function does. Here the function has the int bug from above, and so does the check:
def shipping_cost(weight_kg, zone):
# Local zone only, to keep the example short. Bug: int() instead of math.ceil().
return 60 + (int(weight_kg) - 1) * 20
# The expected value repeats the code's own sum, so it repeats the code's own bug.
assert shipping_cost(1.2, "local") == 60 + (int(1.2) - 1) * 20
print("passed")
# Worked out by hand from the contract: 1.2 kg is charged as 2 kg, so 60 + 20.
assert shipping_cost(1.2, "local") == 80passed
Traceback (most recent call last):
File "/home/you/shop/rederive.py", line 11, in <module>
assert shipping_cost(1.2, "local") == 80
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionErrorThis is also what "testing the implementation" looks like: the check mirrors how the code works instead of stating what it promises. Write the expected value as a plain number from the contract.
The expected value was copied from the output You ran the function, saw 20.0, and pasted it into the check. It passes — and if 20.0 was wrong, it now protects the bug. Work every expected value out before you run the code; if your number and the code's disagree, find out which one is wrong before writing anything down.
The plan has no boundary cases Every check passes, and the bug sits on the edge: five o'clock, exactly 1 kg, age 12. A table with only "typical" rows is the usual cause. For every <, <=, "up to", "from" and "at least" in the contract, there should be a case on each side.
The function cannot be called by a check EOFError: EOF when reading a line means the function waits for the keyboard; returned: None means it printed its answer instead of returning it; a check that passes in the morning and fails in the evening means it reads the clock. All three are design problems, fixed by the pure-core-and-thin-shell reshaping from section two.
SyntaxWarning: assertion is always true, perhaps remove parentheses? Brackets around an assert and its message turn them into one tuple, and a non-empty tuple is always true:
def drink_price(price, hour):
if 17 <= hour < 19:
return round(price * 0.8, 2)
return price
assert (drink_price(10, 17) == 999, "happy hour price")
print("passed?!")/home/you/cafe/tuple.py:7: SyntaxWarning: assertion is always true, perhaps remove parentheses?
assert (drink_price(10, 17) == 999, "happy hour price")
passed?!999 is plainly wrong, and the check passed. Write assert drink_price(10, 17) == 999, "happy hour price" — no brackets — and it fails as it should. Python warns you here; read the warning rather than scrolling past it.
A check that compares decimal results with == fails for a correct answer 0.1 + 0.2 == 0.3 is False in Python, because 0.1 + 0.2 is 0.30000000000000004. The planning lesson is to say in the contract how results are rounded, as member_discount and drink_price do. The pytest tool for comparing floats is chapter five.
Step 4 of 6 — Predict
Check your understanding
The requirement: "orders of 500 or more ship free." What happens when this script runs?
def free_shipping(total):
return total > 500
assert free_shipping(800) is True
assert free_shipping(100) is False
print("checks passed")
assert free_shipping(500) is True
print("all done")- A`checks passed` and then `all done` are printed
- B`checks passed` is printed, then an `AssertionError` on the last assert
- CAn `AssertionError` on the first assert, and nothing is printed
- DNothing happens — asserts are ignored in a script
Which function can be tested reliably, as it is, with a one-line assert?
- Adef greet(): print('Hi', input())
- Bdef is_weekend(): return datetime.now().weekday() >= 5
- Cdef is_weekend(weekday): return weekday >= 5
- Ddef log_visit(): VISITS.append(1)
This check is meant to guard the rule "1.2 kg is charged as 2 kg". What is wrong with it?
from shipping import shipping_cost
expected = shipping_cost(1.2, "local")
assert shipping_cost(1.2, "local") == expected- AThe expected value comes from the code under test, so the check can never fail
- B`1.2` is not a boundary value, so it is not worth testing
- CIt should use `is` instead of `==`
- DThe function must be called inside `try`/`except`
Answering needs an account
Sign in to check your answers
The questions are above, and working them out in your head is the part that matters. Sign in to see the answers, the explanations and the three-level hints.
Your turn
A cinema sends you this requirement:
"Kids pay less, pensioners get a discount, and babies go free. Adults pay full price, 300."
Do everything in this chapter, in order, before writing a single check:
- Ask the questions. Write down at least five things the sentence leaves open — what "kids" means, where "pensioner" starts, whether the edges are included, what the prices are, what happens with an impossible age. Then decide answers, as the cinema would: under 3 is free; 3 to 12 pays 150; 13 to 64 pays 300; 65 and over pays 200; ages run from 0 to 120, and anything outside that raises
ValueError. - Write the contract as the docstring of a function
ticket_price(age)incinema.py. - Make it testable: arguments in, a return value out, no
input()and noprint(). - Write the case table: case, input, expected, why. Find the equivalence classes first, then put a case on each side of every edge. Work every expected value out from the contract.
- Decide what you will not test, and say why.
- Write
check_cinema.pywith one bareassertper row, and run it.
Then break the function on purpose: change the child rule from "12 and under" to "under 12" — <= to <. Before running, predict which check fails. Run it and compare.
Solution
1. The questions, and the decided answers:
| Question | Answer | | --- | --- | | Up to what age is a baby free? | Ages 0, 1 and 2. From 3 a ticket costs money. | | What is a "kid"? | Ages 3 to 12, both included. Price 150. | | Is a 13-year-old a kid? | No — 13 pays the adult price, 300. | | Where do pensioners start? Is 65 itself included? | 65 and over. Price 200. | | What is an impossible age? | Below 0 or above 120: ValueError. | | What comes back? | The price as a whole number; nothing printed. |
2 and 3. The contract and the function, cinema.py. The edges are named constants, so the code reads like the table:
FREE_UNDER = 3 # ages 0-2 go free
CHILD_UP_TO = 12 # ages 3-12 pay the child price
SENIOR_FROM = 65 # ages 65 and over pay the senior price
MAX_AGE = 120
def ticket_price(age):
"""Return the price of one cinema ticket, as a whole number.
- age: a whole number of years, 0 to 120. Outside that range: ValueError.
- 0-2: 0 (free), 3-12: 150, 13-64: 300, 65 and over: 200.
- No side effects: reads nothing, prints nothing, changes nothing.
"""
if not 0 <= age <= MAX_AGE:
raise ValueError(f"age must be 0-{MAX_AGE}, got {age}")
if age < FREE_UNDER:
return 0
if age <= CHILD_UP_TO:
return 150
if age < SENIOR_FROM:
return 300
return 2004. The case table. The classes are: invalid (below 0), free (0–2), child (3–12), adult (13–64), senior (65–120), invalid (above 120). That gives four typical cases, eight boundary cases and two invalid ones:
| Case | Input | Expected | Why | | --- | --- | --- | --- | | Typical baby | 1 | 0 | free class | | Typical child | 8 | 150 | child class | | Typical adult | 30 | 300 | adult class | | Typical senior | 80 | 200 | senior class | | Newborn | 0 | 0 | lowest valid age | | Last free age | 2 | 0 | below the free/child edge | | First child age | 3 | 150 | on the free/child edge | | Last child age | 12 | 150 | "12 and under" — the edge itself | | First adult age | 13 | 300 | just past the child edge | | Last adult age | 64 | 300 | below the senior edge | | First senior age | 65 | 200 | "65 and over" — the edge itself | | Oldest valid age | 120 | 200 | top of the valid range | | Below the range | -1 | ValueError | just outside, low side | | Above the range | 121 | ValueError | just outside, high side |
5. Not tested: a string or None for the age — the contract says a whole number, and promises nothing else. A decimal age such as 12.5 — the same. The prices of other cinemas, or Python's <= — not ours to test. And a second typical case in the same class, such as 9 next to 8 — it would catch nothing new.
6. The checks, check_cinema.py:
from cinema import ticket_price
def raises_value_error(age):
"""True if ticket_price raises ValueError for this age."""
try:
ticket_price(age)
except ValueError:
return True
return False
# One typical age from each class.
assert ticket_price(1) == 0
assert ticket_price(8) == 150
assert ticket_price(30) == 300
assert ticket_price(80) == 200
# Both sides of every edge between classes.
assert ticket_price(0) == 0
assert ticket_price(2) == 0
assert ticket_price(3) == 150
assert ticket_price(12) == 150
assert ticket_price(13) == 300
assert ticket_price(64) == 300
assert ticket_price(65) == 200
assert ticket_price(120) == 200
# Invalid input, just outside the valid range on each side.
assert raises_value_error(-1)
assert raises_value_error(121)
print("all 14 checks passed")$ python check_cinema.py
all 14 checks passedThe deliberate bug: if age <= CHILD_UP_TO: becomes if age < CHILD_UP_TO:. Only age 12 changes class — it falls through to the adult price — so only the check for 12 should fail:
$ python check_cinema.py
Traceback (most recent call last):
File "/home/you/cinema/check_cinema.py", line 23, in <module>
assert ticket_price(12) == 150
^^^^^^^^^^^^^^^^^^^^^^^
AssertionErrorLine 23 is assert ticket_price(12) == 150, as predicted.
Why it is done this way:
- The questions came first. "Kids" and "pensioners" are not numbers; every answer in the first table turned a word into an edge you can test. Without them,
12and65would be guesses. - The function is pure. Age in, price out, nothing else — so every check is one line, and it gives the same answer at any time of day on any machine.
- Named constants for the edges.
CHILD_UP_TO = 12says in one place what the contract says, and the checks name the same numbers. - The boundary rows did the work. The four typical ages — 1, 8, 30, 80 — all still pass with the bug in place. Only the row for
12, the edge itself, caught it. That is the whole argument for boundary values in one run. - Every expected value came from the contract. None of them was copied from running
ticket_price, so a wrong answer from the code cannot sneak into the checks. - It stopped at the first failure, and told you only the line. You knew which check failed because you predicted it — but it did not tell you the actual price for age 12 was 300. That is the gap the next chapter closes.
Step 6 of 6
Stretch — the chapter quiz
Ten questions from easy to hard. The last ones are difficult on purpose.
Sign in to take the quiz