Chapter 00

Before you write a test — what you actually need

The thinking that comes before any test code: turning a vague requirement into a precise contract, making code testable, choosing cases with boundaries, knowing the expected answer without running the code, and deciding what not to test — all with plain Python and assert.

50 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

Here is a function that averages a list of exam scores, and the way most of us check a function like it — call it a couple of times and look:

python
def average(scores):
    return sum(scores) / len(scores)


print(average([80, 90, 70]))
print(average([100]))
text
80.0
100.0

Both answers are right. The function "works", and it goes into the program. A week later a class with no results yet reaches the same line:

python
def average(scores):
    return sum(scores) / len(scores)


print(average([]))
text
Traceback (most recent call last):
  File "/home/you/school/report.py", line 5, in <module>
    print(average([]))
          ^^^^^^^^^^^
  File "/home/you/school/report.py", line 2, in average
    return sum(scores) / len(scores)
           ~~~~~~~~~~~~^~~~~~~~~~~~~
ZeroDivisionError: division by zero

"It works when I tried it" was true. It only ever meant "it works for the two inputs I happened to try". Manual checking has two weaknesses that this exposes.

The first is that it covers only what you thought of in the moment, and an empty list is exactly the kind of thing nobody thinks of in the moment.

The second is that it does not repeat. The two print calls were run once, read once, and deleted. When the function changes next month, nobody runs them again.

Automated tests fix the second weakness, and the rest of this course is about writing them with pytest. But look at the first weakness more closely, because no tool fixes it. Suppose you wanted to write a check for the empty list right now. What would it say? Should average([]) return 0? Return None? Raise an error? Nobody decided. You cannot write a test for a behaviour nobody has decided, and you cannot check an answer you do not already know.

That is what this chapter is about: the thinking that has to happen before the first line of test code. Get it right and the tests almost write themselves. Skip it and you end up with tests that pass and still miss the bug.

By the end of this chapter you can

  • Say what a test is in one sentence: a runnable claim about behaviour
  • Turn a vague requirement into a precise contract — inputs, output, errors, side effects
  • Recognise code that is hard to test, and reshape it into a pure core with a thin shell
  • Plan cases with equivalence classes and boundary values, as a table
  • Work out expected answers independently of the code — the oracle
  • Decide what not to test, and when you have enough tests
  • Write the plan as plain assert checks, run them, and say why they are not yet enough

Prerequisites: none — this is where the course starts. You need Python 3 and to know how to write a function. You do not need pytest yet; everything in this chapter runs with plain Python.


Before you write the test

Every chapter of this course has a section with this name, and every one of them answers the same six questions for its own topic. This chapter is where the questions come from, so here they are in one place. Keep them; the rest of the course uses them as a checklist.

  1. What exactly is promised? The contract: for these inputs, this output; for these inputs, this error; and these side effects, or none.
  2. Can the code be called by a test? Does it take its inputs as arguments and hand back a result — or does it read the keyboard, print, and look at the clock?
  3. Which cases? One from each kind of input, both sides of every edge, the invalid inputs, the empty and the very large.
  4. How do I know the right answer? Worked out by hand or from the specification — never by running the code under test.
  5. What must be in place? The Python version, an environment, where the test files will live — and, in later chapters, files, settings and tools the topic needs.
  6. What will I not test? Python itself, other people's libraries, code too simple to be wrong, and behaviour nobody promised.

Here they are applied to average from above, which is small enough to do in full.

The contract has a gap, so the first step is to close it. Suppose you ask, and the answer is: "an average of no scores makes no sense; raise ValueError". The function is already testable — it takes a list and returns a number. For this chapter, all that must be in place is Python 3. And the cases:

| Case | Input | Expected | Why | | --- | --- | --- | --- | | Typical | [80, 90, 70] | 80.0 | the ordinary use: (80 + 90 + 70) / 3 | | A single score | [100] | 100.0 | smallest list that has an average | | Empty | [] | ValueError | the case that broke; now decided |

What is not on the list: whether sum adds correctly, or whether / divides — that is Python's job, and it is tested far more thoroughly than anything you or I will write.

The rest of the chapter takes each of the six questions in turn, and ends by running the whole process on one realistic function.

What a test actually is

A test is a runnable claim about behaviour: given this input, expect this result. "Result" comes in exactly three kinds:

  • a return value — average([80, 90, 70]) gives 80.0
  • an error — average([]) raises ValueError
  • a side effect — after add_score(scores, 70), the list scores has 70 on the end

Python already has a statement for writing a claim down: assert. It does nothing when the condition is true and raises AssertionError when it is false. Here is one claim of each kind, with average now following its decided contract:

python
def average(scores):
    if not scores:
        raise ValueError("average() of an empty list")
    return sum(scores) / len(scores)


def add_score(scores, new_score):
    scores.append(new_score)


# 1. A return value: given this input, expect this output.
assert average([80, 90, 70]) == 80.0

# 2. An error: given this input, expect this exception.
try:
    average([])
except ValueError:
    pass
else:
    raise AssertionError("average([]) should raise ValueError")

# 3. A side effect: given this call, expect this change to the world.
scores = [80, 90]
add_score(scores, 70)
assert scores == [80, 90, 70]

print("3 claims checked")
text
3 claims checked

Notice what each claim needs: an input you choose, a result you already know, and a way to call the code and look at what it did. The six questions exist to make sure you have all three before you start. Notice too how clumsy the error check is — five lines of try/except/else to say "this should raise". pytest turns that into one line in chapter four.

1. The contract: from a vague requirement to a precise promise

Requirements usually arrive as a sentence: "Members get 10% off big orders." It sounds complete. Give it to two careful developers and watch:

python
# "Members get 10% off big orders." Two honest readings of the same sentence.

def discount_by_asha(total, is_member):
    if is_member and total > 1000:
        return total * 0.10
    return 0


def discount_by_ravi(total, is_member):
    if is_member and total >= 1000:
        return round(total * 0.10)
    return 0


for total in [500, 1000, 1234.56]:
    print(total, discount_by_asha(total, True), discount_by_ravi(total, True))
text
500 0 0
1000 0 100
1234.56 123.456 123

Neither of them made a mistake. The sentence did not say whether exactly 1000 counts as "big", or how to round. Each filled the gap differently, and an order of 1000 gets a discount of 0 from one and 100 from the other. No test can say which of them is right, because "right" was never defined.

A contract is the requirement made precise enough to test. To get from the sentence to the contract, ask questions — the same ones every time:

  • Inputs. What types? What range? Is there a limit, and is the limit itself included? What about zero, negative, empty, None?
  • Output. What exactly is returned — the discount, or the new total? What type? Rounded to what?
  • Errors. Which inputs are refused, and how — which exception?
  • Side effects. Does it change anything — a list passed in, a file, a database — or print anything? Or nothing at all?

Asked about the discount, those questions produce answers, and the answers go where the code is — in the docstring:

python
def member_discount(total, is_member):
    """Return the discount on an order, in currency units.

    - total: the order total before discount, a number >= 0.
      A negative total raises ValueError.
    - is_member: True or False.
    - Members get 10% of the total when the total is 1000 or more
      (1000 itself qualifies). Everyone else gets 0.
    - The result is rounded to 2 decimal places.
    - Returns the discount amount, not the new total.
    - No side effects: reads nothing, prints nothing, changes nothing.
    """
    if total < 0:
        raise ValueError(f"total must not be negative, got {total}")
    if is_member and total >= 1000:
        return round(total * 0.10, 2)
    return 0


print(member_discount(999.99, True))
print(member_discount(1000, True))
print(member_discount(1234.56, True))
print(member_discount(5000, False))
try:
    member_discount(-5, True)
except ValueError as error:
    print("ValueError:", error)
text
0
100.0
123.46
0
ValueError: total must not be negative, got -5

Every line of the docstring is now something a test can check, and every check has an obvious expected value. That is the test of a good contract: could a stranger write the expected value for any input, without reading the code? If not, there is still a question to ask.

Who answers the questions? Whoever owns the requirement — a product owner, a teacher, a client, or you. When there is nobody to ask, decide, write the decision into the contract, and move on. An explicit decision can be reviewed and changed; a silent one just lives in the code until it surprises someone.

2. Testable code: a pure core and a thin shell

A test calls a function with chosen inputs and compares what comes back. Some functions make that impossible. Here is one, for a café's happy hour — 20% off drinks between 17:00 and 19:00:

python
from datetime import datetime


def drink_price():
    price = float(input("Price: "))
    hour = datetime.now().hour
    if 17 <= hour < 19:
        price = price * 0.8
    print(f"You pay {price:.2f}")

It works when you run it and type a price. Now try to check it from another file, check_drinks.py:

python
from drinks import drink_price

result = drink_price()
print("returned:", result)

An automated check runs with nobody at the keyboard. < /dev/null gives the program exactly that — an input with nothing in it:

text
$ python check_drinks.py < /dev/null
Price: Traceback (most recent call last):
  File "/home/you/cafe/check_drinks.py", line 3, in <module>
    result = drink_price()
             ^^^^^^^^^^^^^
  File "/home/you/cafe/drinks.py", line 5, in drink_price
    price = float(input("Price: "))
                  ^^^^^^^^^^^^^^^^
EOFError: EOF when reading a line

Feed it a price, and it runs — but look at what comes back:

text
$ echo 10 | python check_drinks.py
Price: You pay 10.00
returned: None

Three separate problems, each one common:

  • It reads its input with input() instead of taking it as an argument, so a test cannot choose the input.
  • It prints its result instead of returning it, so a test gets None and has nothing to compare.
  • It reads the clock itself. That run was just before seven in the morning, so the answer was 10.00; the same command at half past five in the evening gives 8.00. A check on it would pass or fail depending on when it ran — a flaky test, which is worse than none, because people learn to ignore it.

The fix is not a testing trick. It is a change of shape: separate the decision from the input and output. The decision becomes a pure function — everything it needs comes in as arguments, the answer goes out as the return value, and it touches nothing else. The keyboard, the screen and the clock move into a thin shell around it:

python
from datetime import datetime


def drink_price(price, hour):
    """Return the price to pay for a drink.

    20% off from 17:00 up to, but not including, 19:00.
    """
    if 17 <= hour < 19:
        return round(price * 0.8, 2)
    return price


def main():
    # The thin shell: all the input, output and clock reading lives here.
    price = float(input("Price: "))
    print(f"You pay {drink_price(price, datetime.now().hour):.2f}")


if __name__ == "__main__":
    main()

Now the checks choose any hour they like, at any time of day:

python
from drinks import drink_price

assert drink_price(10, 16) == 10
assert drink_price(10, 17) == 8.0
assert drink_price(10, 18) == 8.0
assert drink_price(10, 19) == 10
print("all checks passed")
text
all checks passed

The program behaves exactly as before for the person at the keyboard. What changed is that the part worth testing — the rule — can now be tested in isolation, and the shell is so thin there is almost nothing in it to get wrong.

The warning signs of hard-to-test code, so you can spot them before writing a test:

| Sign | Why it hurts a test | The usual fix | | --- | --- | --- | | input() inside the logic | the test cannot choose the input | take it as an argument | | print() as the only result | nothing comes back to compare | return the value; print in the shell | | datetime.now(), random inside | the answer changes between runs | pass the time or the random value in | | reads or changes a global variable | tests affect one another | pass it in, return the new value | | opens a fixed file or URL inside | the test needs that file or network | pass the data, or the path, in |

You cannot always reshape code — sometimes it belongs to someone else, or the I/O is the behaviour. pytest has tools for those cases: capsys captures printed output (chapter nine), tmp_path gives a throwaway folder (chapter nine), and monkeypatch and mocks replace the clock, the network and the environment (chapters ten and eleven). But reshaping comes first. Code that is easy to test is usually just easier to understand.

3. Choosing the cases

You cannot test every input; drink_price alone accepts every price at every hour. The job is to pick a small set where each case could catch a mistake the others would miss. Two ideas do most of that work.

Equivalence classes. Group the inputs that the contract treats the same way. Any one member of a group stands for the rest: if drink_price(10, 12) is right, drink_price(10, 11) almost certainly is too, because the same line of code handles both. Suppose the contract has been tightened to say that an hour outside 0–23, or a negative price, raises ValueError. Then the hours fall into five classes:

text
hour:   ... -2 -1 | 0 1 ... 15 16 | 17 18 | 19 20 ... 23 | 24 25 ...
        invalid   | full price    | 20% off | full price | invalid

Boundary values. Mistakes do not spread evenly across a class. They gather at its edges, because that is where < and <= are chosen, and where "up to 19:00" is turned into code. So for every edge, test the value on each side of it: the last of one class and the first of the next. For a limit L that means looking at L - 1, L and L + 1; for counts and lengths, 0 and 1 as well.

Here is why "one typical value per class" is not enough. Below, the happy-hour test is written with an easy slip, 17 < hour instead of 17 <= hour:

python
def drink_price(price, hour):
    # A slip: 17 < hour instead of 17 <= hour.
    if 17 < hour < 19:
        return round(price * 0.8, 2)
    return price


# Three "typical" cases, one from each valid class:
assert drink_price(10, 12) == 10
assert drink_price(10, 18) == 8.0
assert drink_price(10, 21) == 10
print("typical cases passed")

# The boundary:
assert drink_price(10, 17) == 8.0
print("boundary passed")
text
typical cases passed
Traceback (most recent call last):
  File "/home/you/cafe/boundary.py", line 15, in <module>
    assert drink_price(10, 17) == 8.0
           ^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionError

One case from each class: all pass, and customers who arrive at five o'clock pay full price. The boundary case catches it at once.

Classes and boundaries cover the shape of the input. Then run through these questions, each of which has found real bugs:

  • Invalid input — what the contract says to refuse: a negative price, hour 24.
  • Empty and zero — [], "", 0, a price of 0. Does "nothing" behave?
  • None — only if the contract mentions it. If it does not, it is not a case (see question six).
  • Large values — a very long list, a very large number. Anything with a limit, test at the limit.
  • Rounding — where a calculation produces more decimals than the result keeps.

Put together, that is a test plan: one row per case, written before any test code. The last column matters most — a row you cannot justify is a row you do not need.

| Case | Input (price, hour) | Expected | Why | | --- | --- | --- | --- | | Lowest valid hour | (10, 0) | 10 | edge of the valid range | | Last full-price hour | (10, 16) | 10 | below the 17:00 edge | | First discounted hour | (10, 17) | 8.0 | on the 17:00 edge | | Last discounted hour | (10, 18) | 8.0 | below the 19:00 edge | | First full price again | (10, 19) | 10 | on the 19:00 edge — "not including" | | Highest valid hour | (10, 23) | 10 | edge of the valid range | | Free drink | (0, 17) | 0 | zero price | | Rounding | (2.49, 17) | 1.99 | 2.49 × 0.8 = 1.992, rounded to 2 places | | Hour too low | (10, -1) | ValueError | just outside the range | | Hour too high | (10, 24) | ValueError | just outside the range | | Negative price | (-1, 12) | ValueError | invalid input |

And the plan turned into checks, against the tightened function:

python
def drink_price(price, hour):
    """Return the price to pay for a drink.

    - price: a number >= 0. hour: a whole number 0-23.
    - 20% off from 17:00 up to, but not including, 19:00.
    - The result is rounded to 2 decimal places.
    - An hour outside 0-23 or a negative price raises ValueError.
    """
    if not 0 <= hour <= 23:
        raise ValueError(f"hour must be 0-23, got {hour}")
    if price < 0:
        raise ValueError(f"price must not be negative, got {price}")
    if 17 <= hour < 19:
        return round(price * 0.8, 2)
    return price


def raises_value_error(price, hour):
    try:
        drink_price(price, hour)
    except ValueError:
        return True
    return False


assert drink_price(10, 0) == 10      # lowest valid hour
assert drink_price(10, 16) == 10     # last full-price hour before
assert drink_price(10, 17) == 8.0    # first discounted hour
assert drink_price(10, 18) == 8.0    # last discounted hour
assert drink_price(10, 19) == 10     # first full-price hour after
assert drink_price(10, 23) == 10     # highest valid hour
assert drink_price(0, 17) == 0       # free drink stays free
assert drink_price(2.49, 17) == 1.99 # rounding: 1.992 -> 1.99
assert raises_value_error(10, -1)    # just below the valid range
assert raises_value_error(10, 24)    # just above the valid range
assert raises_value_error(-1, 12)    # negative price
print("11 checks passed")
text
11 checks passed

Eleven rows, and each one is there for a reason you can say out loud.

4. The oracle: knowing the answer without the code

Every check compares the code's answer with an expected answer. The source of the expected answer is called the oracle, and it has one rule: it must not be the code under test.

That sounds too obvious to need saying, until you see how easily it is broken. Here member_discount has a bug — 1% instead of 10% — and two checks:

python
def member_discount(total, is_member):
    if total < 0:
        raise ValueError(f"total must not be negative, got {total}")
    if is_member and total >= 1000:
        return round(total * 0.01, 2)   # bug: 1%, not 10%
    return 0


# Wrong: the expected value comes from the code under test.
expected = member_discount(2000, True)
assert member_discount(2000, True) == expected
print("self-check passed, expected was", expected)

# Right: the expected value was worked out by hand. 10% of 2000 is 200.
assert member_discount(2000, True) == 200
print("hand-check passed")
text
self-check passed, expected was 20.0
Traceback (most recent call last):
  File "/home/you/shop/oracle.py", line 15, in <module>
    assert member_discount(2000, True) == 200
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionError

The first check compares the function with itself. It passes for the correct function, for this broken one, and for any other version anyone could write — it cannot fail, so it tells you nothing.

Nobody writes it quite that baldly. The same mistake usually arrives in one of two disguises:

  • Copying the output. You run the function, it prints 20.0, and you paste 20.0 into the check. If the code was already wrong, the check now guards the bug.
  • Repeating the formula. You write the expected value as 2000 * 0.01 — the same sum the code does. A copy of the code's logic shares the code's mistakes. ("When it breaks", below, shows this one running.)

Good oracles, roughly from most to least common:

  • Working it out by hand from the contract. Choose inputs that make this easy: 2000 rather than 1873.41, so that "10% of it" is something you can do in your head.
  • Examples in the specification. If the requirement says "an order of 1500 gets 150 off", that row is a test.
  • Known facts. Water boils at 100 °C and 212 °F; an empty list has length 0.
  • A different, trusted method. A simpler, slower calculation, or a standard-library function that does the same job another way:
python
import statistics


def average(scores):
    if not scores:
        raise ValueError("average() of an empty list")
    return sum(scores) / len(scores)


samples = [[80, 90, 70], [1, 2], [5], [10, 0, 0, 0]]
for scores in samples:
    assert average(scores) == statistics.mean(scores), scores
print("agrees with statistics.mean on", len(samples), "samples")
text
agrees with statistics.mean on 4 samples

statistics.mean was written by other people, in a different way, so agreeing with it means something. (Comparing decimal results with == can bite — 0.1 + 0.2 == 0.3 is False in Python. The values here are safe; chapter five shows how to compare floats properly.)

If you cannot work out the expected answer by any of these means, that is not a testing problem. It means you do not yet know what the code should do — go back to the contract.

5. What must be in place

For this chapter, only Python 3. Check which one you have:

text
$ python3 --version
Python 3.12.3

This course uses Python 3.12, and anything from 3.10 up will run nearly all of it. On Windows the command is usually py --version.

From the next chapter on, three more things have to be in place before a test can run, and chapter one sets each of them up:

  • A virtual environment for the project, so that the testing tool is installed in the same Python that runs your code.
  • pytest installed into it.
  • A place for the tests. Test files sit next to the code, or in a tests/ folder, with names starting test_. Chapter one starts with the first; chapter three moves to the second and explains why.

Later chapters add their own needs — a configuration file, a plugin, sample data on disk — and each one lists them in its own "Before you write the test".

6. What not to test, and when to stop

A test that cannot catch a mistake in your code costs time to write, time to run and time to maintain, and gives nothing back. Leave these out:

  • Python itself. Not whether sorted sorts or round rounds. Test that your function sorts the right thing in the right order.
  • Other people's libraries. Not whether requests sends a request or pandas reads a CSV. Test what your code does with the result — and if you need to be sure of a library's behaviour, a single small check that documents your assumption is enough.
  • Code too simple to be wrong. A function that returns a constant, or a class that only stores its arguments. If you cannot imagine the mistake, the test cannot catch it.
  • Behaviour nobody promised. If the contract says nothing about passing a string to drink_price, a test would only freeze whatever it happens to do today. Either add it to the contract — then test it — or leave it out.
  • How, rather than what. Not which helper the function called internally, or in what order it did its steps. Test what goes in and what comes out; then the inside can be rewritten freely, and the tests prove it still works.

And how many is enough? There is no number, but there is a rule. You have enough when:

  • every equivalence class has at least one case,
  • every boundary has a case on each side,
  • every error the contract promises has a case,
  • and every bug you have ever fixed has a case of its own, so it cannot quietly return.

Then stop. For each new case you are tempted to add, ask: what mistake would this catch that the others would not? If you cannot name one, it is a duplicate. drink_price(10, 11) next to drink_price(10, 12) catches nothing new.


A complete example

Now the whole process, start to finish, on one realistic function.

The requirement, as it arrived: "Shipping is 60 for local parcels and 120 for national ones. Heavier parcels cost more — 20 per extra kilo locally, 30 nationally. We don't take anything over 30 kg."

The questions, and the answers they got:

| Question | Answer | | --- | --- | | What does the base price cover? | The first kilogram — 1 kg exactly is still the base price. | | Is 1.2 kg charged as 1 kg or as 2? | Every started kilogram counts: 1.2 kg is charged as 2 kg. | | Is 30 kg itself accepted? | Yes. 30 is allowed; anything above is refused. | | What about 0 kg, or a negative weight? | Refused — that is a data-entry mistake. | | How are zones written? | Exactly "local" or "national". Anything else is refused, including "Local". | | What comes back? | The cost as a whole number. Nothing is printed or saved. |

The contract and the testable function, shipping.py. It takes everything as arguments and returns a number; there is no shell to separate, because there is no input or output in it at all:

python
import math

BASE = {"local": 60, "national": 120}       # covers the first kilogram
PER_EXTRA_KG = {"local": 20, "national": 30}
MAX_WEIGHT_KG = 30


def shipping_cost(weight_kg, zone):
    """Return the shipping cost of one parcel, as a whole number.

    - weight_kg: a number. At or below 0, or above 30: ValueError.
    - zone: "local" or "national", exactly. Anything else: ValueError.
    - The base price covers the first kilogram (1 kg included).
    - Every started kilogram after that costs extra: 1.2 kg is charged as 2 kg.
    - No side effects: reads nothing, prints nothing, changes nothing.
    """
    if zone not in BASE:
        raise ValueError(f"unknown zone: {zone!r}")
    if not 0 < weight_kg <= MAX_WEIGHT_KG:
        raise ValueError(f"weight must be more than 0 and at most 30 kg, got {weight_kg}")
    extra_kg = math.ceil(weight_kg) - 1
    return BASE[zone] + extra_kg * PER_EXTRA_KG[zone]

math.ceil rounds up to the next whole number — math.ceil(1.2) is 2 — which is exactly "every started kilogram counts".

The case table. The weight classes are: invalid (0 and below), the base kilogram (above 0, up to 1), extra kilograms (above 1, up to 30), invalid (above 30). The zones are "local", "national", and anything else. Every expected value is worked out by hand from the contract, not from the code:

| Case | Input | Expected | Why | | --- | --- | --- | --- | | Light parcel | (0.5, "local") | 60 | inside the base kilogram | | Exactly 1 kg | (1, "local") | 60 | top edge of the base: 1 kg is included | | Just over 1 kg | (1.2, "local") | 80 | started kilogram counts: 60 + 1 × 20 | | Whole kilos, local | (3, "local") | 100 | 60 + 2 × 20 | | Whole kilos, national | (3, "national") | 180 | 120 + 2 × 30 — the other zone | | At the limit | (30, "national") | 990 | 120 + 29 × 30; 30 is accepted | | Zero | (0, "local") | ValueError | bottom edge: 0 is refused | | Negative | (-2, "local") | ValueError | invalid input | | Over the limit | (30.5, "local") | ValueError | just above 30 | | Wrong zone | (2, "Local") | ValueError | not exactly "local" |

What is deliberately not in the table: a weight of None or "2". The contract says the weight is a number, so those are outside it — Python raises its own TypeError when they are compared with 0, and we make no promise beyond that. Not math.ceil either: that is Python's, not ours. And not (2, "local") next to (3, "local"): same class, same line of code, nothing new to catch.

The checks, check_shipping.py — the table, row for row:

python
from shipping import shipping_cost


def raises_value_error(weight_kg, zone):
    """True if shipping_cost raises ValueError for these arguments."""
    try:
        shipping_cost(weight_kg, zone)
    except ValueError:
        return True
    return False


# Valid weights: one case per class, and both sides of every edge.
assert shipping_cost(0.5, "local") == 60
assert shipping_cost(1, "local") == 60
assert shipping_cost(1.2, "local") == 80
assert shipping_cost(3, "local") == 100
assert shipping_cost(3, "national") == 180
assert shipping_cost(30, "national") == 990

# Invalid input: the contract promises a ValueError for each.
assert raises_value_error(0, "local")
assert raises_value_error(-2, "local")
assert raises_value_error(30.5, "local")
assert raises_value_error(2, "Local")

print("all 10 checks passed")
text
$ python check_shipping.py
all 10 checks passed

Now break it. Some months later, someone tidies shipping.py, drops import math and writes int(weight_kg) instead of math.ceil(weight_kg). It reads well, and it runs without an error. Run the checks again:

text
$ python check_shipping.py
Traceback (most recent call last):
  File "/home/you/shop/check_shipping.py", line 14, in <module>
    assert shipping_cost(0.5, "local") == 60
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionError

The plan did its job: the change was caught the moment the checks ran. int cuts the decimals off instead of rounding up, so 0.5 kg became int(0.5) - 1, which is -1 extra kilograms.

But read the report as if you did not already know that, and notice how little it says:

  • What was the actual value? The report shows the line, not the number. You have to run shipping_cost(0.5, "local") yourself to learn it was 40.
  • What else is broken? The script stopped at the first failure. The 1.2 kg case is wrong too — it now gives 60 instead of 80 — but that check never ran.
  • Which promises held? Nothing tells you that the error cases still pass.
  • And the error checks needed a helper function with try/except just to say "this should raise".

The thinking was the hard part, and it is done: the contract, the table, the oracle. What is missing is a better runner — one that finds the checks on its own, runs every one even after a failure, shows the actual value next to the expected one, and checks errors in a single line. That is pytest, and installing it is the next chapter.


When it breaks

These are planning mistakes rather than Python errors, so the symptoms are quieter: usually a check that passes when it should not.

The expected value repeats the code's own formula The check is written by doing the same sum the function does. Here the function has the int bug from above, and so does the check:

python
def shipping_cost(weight_kg, zone):
    # Local zone only, to keep the example short. Bug: int() instead of math.ceil().
    return 60 + (int(weight_kg) - 1) * 20


# The expected value repeats the code's own sum, so it repeats the code's own bug.
assert shipping_cost(1.2, "local") == 60 + (int(1.2) - 1) * 20
print("passed")

# Worked out by hand from the contract: 1.2 kg is charged as 2 kg, so 60 + 20.
assert shipping_cost(1.2, "local") == 80
text
passed
Traceback (most recent call last):
  File "/home/you/shop/rederive.py", line 11, in <module>
    assert shipping_cost(1.2, "local") == 80
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AssertionError

This is also what "testing the implementation" looks like: the check mirrors how the code works instead of stating what it promises. Write the expected value as a plain number from the contract.

The expected value was copied from the output You ran the function, saw 20.0, and pasted it into the check. It passes — and if 20.0 was wrong, it now protects the bug. Work every expected value out before you run the code; if your number and the code's disagree, find out which one is wrong before writing anything down.

The plan has no boundary cases Every check passes, and the bug sits on the edge: five o'clock, exactly 1 kg, age 12. A table with only "typical" rows is the usual cause. For every <, <=, "up to", "from" and "at least" in the contract, there should be a case on each side.

The function cannot be called by a check EOFError: EOF when reading a line means the function waits for the keyboard; returned: None means it printed its answer instead of returning it; a check that passes in the morning and fails in the evening means it reads the clock. All three are design problems, fixed by the pure-core-and-thin-shell reshaping from section two.

SyntaxWarning: assertion is always true, perhaps remove parentheses? Brackets around an assert and its message turn them into one tuple, and a non-empty tuple is always true:

python
def drink_price(price, hour):
    if 17 <= hour < 19:
        return round(price * 0.8, 2)
    return price


assert (drink_price(10, 17) == 999, "happy hour price")
print("passed?!")
text
/home/you/cafe/tuple.py:7: SyntaxWarning: assertion is always true, perhaps remove parentheses?
  assert (drink_price(10, 17) == 999, "happy hour price")
passed?!

999 is plainly wrong, and the check passed. Write assert drink_price(10, 17) == 999, "happy hour price" — no brackets — and it fails as it should. Python warns you here; read the warning rather than scrolling past it.

A check that compares decimal results with == fails for a correct answer 0.1 + 0.2 == 0.3 is False in Python, because 0.1 + 0.2 is 0.30000000000000004. The planning lesson is to say in the contract how results are rounded, as member_discount and drink_price do. The pytest tool for comparing floats is chapter five.