Chapter 14

Coverage and CI — which lines did your tests never run?

Measuring coverage with pytest-cov, reading the Missing column, branch coverage, enforcing a minimum, running tests in parallel with pytest-xdist, and running the whole suite on every push with GitHub Actions. And why 100% coverage does not mean correct code.

50 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

Here is a small function that prices a parcel, and two tests for it.

shop/shipping.py:

python
def shipping_cost(weight_kg: float, country: str) -> float:
    if weight_kg <= 0:
        raise ValueError(f"weight must be positive: {weight_kg}")
    if country == "BD":
        base = 60.0
    elif country == "IN":
        base = 90.0
    else:
        base = 500.0
    if weight_kg > 5:
        base += (weight_kg - 5) * 20
    return base

tests/test_shipping.py:

python
from shop.shipping import shipping_cost


def test_light_parcel_to_bangladesh():
    assert shipping_cost(1, "BD") == 60.0


def test_heavy_parcel_to_bangladesh():
    assert shipping_cost(7, "BD") == 100.0
text
$ pytest -q
..                                                                       [100%]
2 passed in 0.07s

Green. And the green says nothing about the parts of the function the tests never touched. Does a parcel to India cost 90? Does a weight of zero really raise? Nobody asked. A typo in the "IN" branch would sit there, passing, until a customer found it.

With twelve lines you can check by eye. With twelve thousand you cannot. You need the computer to answer one question: which lines did my tests never run? That measurement is called coverage, and this chapter is about getting it, reading it, and not being fooled by it — and then about running the whole suite fast and automatically on every push.

By the end of this chapter you can

  • Measure coverage with pytest-cov and read the Missing column
  • Open the HTML report to see uncovered lines in context
  • Explain why line coverage can say 100% when a branch was never taken, and turn on branch coverage
  • Fail the run when coverage drops below a threshold, and exclude code on purpose
  • Keep all of this in pyproject.toml
  • Say why 100% coverage does not mean the code is correct
  • Run tests in parallel with pytest-xdist, and explain why tests must not depend on each other
  • Write a GitHub Actions workflow that runs the suite on several Python versions

Prerequisites: configuration and warnings.


Before you write the test

The wrong way to use coverage is to open the report first and write whatever test turns the next red line green. That produces tests shaped like the code, not like the promise the code makes — and, as you will see, it misses bugs that have no line of their own. The right order is: plan from the contract, write the tests, and only then ask coverage what the plan forgot.

The contract. Read shipping_cost as a promise, not as code:

  • weight zero or below → ValueError whose message says "positive"
  • Bangladesh costs 60, India 90, anywhere else 500
  • above 5 kg, every extra kilogram adds 20; at exactly 5 kg nothing is added
  • no side effects: it returns a number and touches nothing else

What must be set up. A virtual environment with pytest and pytest-cov installed in it (and pytest-xdist later). The package must be importable from the tests — here shop, either installed or reached through pythonpath = ["src"] from the previous chapter. And you need to know the import name of the code you are measuring, because that is what goes after --cov=.

The plan. One row per behaviour, with boundaries and invalid input on purpose:

| Case | Input | Expected | |---|---|---| | happy path, home country | 1, "BD" | 60.0 | | second country | 1, "IN" | 90.0 | | anywhere else | 1, "US" | 500.0 | | boundary: exactly 5 kg | 5, "BD" | 60.0 (no surcharge) | | over the limit | 7, "BD" | 100.0 (60 + 2 × 20) | | invalid: zero weight | 0, "BD" | ValueError, "positive" |

What not to test. Code you did not write (Python's round, the coverage library itself). A __main__ block that only runs when a human executes the file. And never a test whose only job is to make a line run — a test with no meaningful assert raises the percentage and checks nothing.

Notice the two boundary rows. A test at 5 kg runs no line that the 1 kg test does not already run, so no coverage report will ever ask for it. And for the invalid case, a test with -1 would cover the raise line exactly as well as 0 does — but if someone typed weight_kg < 0 instead of <= 0, only the 0 test would fail. Coverage cannot tell those two tests apart; the contract can. That is the whole relationship: the plan decides what to test; coverage reports what the plan missed.

The two tests from the opening cover rows 1 and 5. The rest of the chapter measures that gap, closes it, and then automates the measuring.

Measuring with pytest-cov

Coverage is measured by the coverage library; pytest-cov is the plugin that switches it on from the pytest command line. Install it as a development dependency:

text
$ uv add --dev pytest-cov

Then tell it which package to watch with --cov=:

text
$ pytest -q --cov=shop
..                                                                       [100%]
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name               Stmts   Miss  Cover
--------------------------------------
shop/__init__.py       0      0   100%
shop/shipping.py      11      4    64%
--------------------------------------
TOTAL                 11      4    64%
2 passed in 0.09s

Stmts is the number of executable statements in the file. Miss is how many of them never ran. Cover is the share that did. Eleven statements, four never executed: 64%.

The value after --cov= is the code you want measured, not the tests. Your tests always run all of their own lines, so counting them would only flatter the total.

Reading the Missing column

A percentage tells you how much. It does not tell you where. Add --cov-report=term-missing:

text
$ pytest -q --cov=shop --cov-report=term-missing
..                                                                       [100%]
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name               Stmts   Miss  Cover   Missing
------------------------------------------------
shop/__init__.py       0      0   100%
shop/shipping.py      11      4    64%   3, 6-9
------------------------------------------------
TOTAL                 11      4    64%
2 passed in 0.09s

3, 6-9 are line numbers in shipping.py. Count down the file:

  • Line 3 is raise ValueError(...) — no test passed a weight of zero or less.
  • Lines 6 to 9 are elif country == "IN":, base = 90.0, else: and base = 500.0 — no test sent a parcel anywhere but Bangladesh.

Line 8, else:, is not a statement on its own, so it is not in the count of eleven; coverage just folds it into the range 6-9 to keep the list short.

That column is the to-do list. Each number is a line that could be completely wrong and your suite would still be green.

The HTML report

For a large file a list of numbers gets hard to follow. --cov-report=html writes a small website instead:

text
$ pytest -q --cov=shop --cov-report=html
..                                                                       [100%]
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Coverage HTML written to dir htmlcov
2 passed in 0.11s

Open htmlcov/index.html in a browser. Every file is listed with its percentage; click one and you see the source, with run lines marked green and missed lines marked red. It is the same information as Missing, laid over the code itself.

Both htmlcov/ and the .coverage data file are generated output. Put them in .gitignore.

Branch coverage — why line coverage lies

Here is a function where line coverage says everything is fine:

python
def final_price(total: float, coupon: str | None) -> float:
    if coupon == "SAVE10":
        total = total * 0.9
    return round(total, 2)
python
from shop.coupons import final_price


def test_coupon_takes_ten_percent_off():
    assert final_price(200.0, "SAVE10") == 180.0
text
$ pytest -q --cov=shop --cov-report=term-missing
.                                                                        [100%]
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name               Stmts   Miss  Cover   Missing
------------------------------------------------
shop/__init__.py       0      0   100%
shop/coupons.py        4      0   100%
------------------------------------------------
TOTAL                  4      0   100%
1 passed in 0.09s

100%. Every line ran. But nobody ever called final_price without a coupon. An if with no else has two ways out — into the body, or straight past it — and only one of them was tried. If someone later breaks the no-coupon path, this report will not notice.

Line coverage asks "did this line run?". Branch coverage asks "did each decision go both ways?". Turn it on with --cov-branch:

text
$ pytest -q --cov=shop --cov-branch --cov-report=term-missing
.                                                                        [100%]
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name               Stmts   Miss Branch BrPart  Cover   Missing
--------------------------------------------------------------
shop/__init__.py       0      0      0      0   100%
shop/coupons.py        4      0      2      1    83%   2->4
--------------------------------------------------------------
TOTAL                  4      0      2      1    83%
1 passed in 0.10s

Two new columns. Branch is the number of possible jumps (the if on line 2 can go to line 3 or to line 4: two). BrPart counts decisions that went only one way. And Missing now shows 2->4: the jump from line 2 straight to line 4 — the "no coupon" path — never happened.

Branch coverage is stricter and closer to the truth. There is little reason not to have it on all the time.

Making the number a gate: --cov-fail-under

A report that nobody reads changes nothing. --cov-fail-under turns the total into a pass/fail condition:

text
$ pytest -q --cov=shop --cov-branch --cov-report=term-missing --cov-fail-under=90
.
ERROR: Coverage failure: total of 83 is less than fail-under=90
                                                                         [100%]
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name               Stmts   Miss Branch BrPart  Cover   Missing
--------------------------------------------------------------
shop/__init__.py       0      0      0      0   100%
shop/coupons.py        4      0      2      1    83%   2->4
--------------------------------------------------------------
TOTAL                  4      0      2      1    83%
FAIL Required test coverage of 90% not reached. Total coverage: 83.33%
1 passed in 0.12s

Read the last line carefully: 1 passed. Every test passed, and yet the command exits with status 1. In CI that is a red build. The point is not that 90 is a magic number; it is that coverage can no longer quietly slide down when someone adds code without tests.

Choose a threshold a little below where you are today and raise it over time. Setting it to 100 on day one usually ends with people writing pointless tests to satisfy the number.

Excluding code on purpose: # pragma: no cover

Some lines are not worth testing — a block that only runs when you execute the file by hand, say. Add a second test, assert final_price(200.0, None) == 200.0, so final_price is fully covered — and then a __main__ block to the module:

python
def final_price(total: float, coupon: str | None) -> float:
    if coupon == "SAVE10":
        total = total * 0.9
    return round(total, 2)


if __name__ == "__main__":
    print(final_price(200.0, "SAVE10"))
text
$ pytest -q --cov=shop --cov-branch --cov-report=term-missing
..                                                                       [100%]
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name               Stmts   Miss Branch BrPart  Cover   Missing
--------------------------------------------------------------
shop/__init__.py       0      0      0      0   100%
shop/coupons.py        6      1      4      1    80%   8
--------------------------------------------------------------
TOTAL                  6      1      4      1    80%
2 passed in 0.11s

Line 8, the print, never runs under pytest — and it never will. Mark the if with a comment that coverage understands:

python
if __name__ == "__main__":  # pragma: no cover
    print(final_price(200.0, "SAVE10"))
text
$ pytest -q --cov=shop --cov-branch --cov-report=term-missing
..                                                                       [100%]
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name               Stmts   Miss Branch BrPart  Cover   Missing
--------------------------------------------------------------
shop/__init__.py       0      0      0      0   100%
shop/coupons.py        4      0      2      0   100%
--------------------------------------------------------------
TOTAL                  4      0      2      0   100%
2 passed in 0.08s

When the pragma sits on a line that opens a block, the whole block is excluded. Use it sparingly and honestly. Every pragma is a line you have decided not to check; if you find yourself adding one to make a hard branch "go away", that branch is probably the one that most needs a test.

Putting it in pyproject.toml

Typing four --cov flags every time is how they get forgotten. Coverage reads its own settings from pyproject.toml, and the addopts you met in the previous chapter makes pytest switch it on:

toml
[tool.pytest.ini_options]
pythonpath = ["src"]
testpaths = ["tests"]
addopts = "--cov --cov-report=term-missing"

[tool.coverage.run]
source = ["shop"]
branch = true

[tool.coverage.report]
fail_under = 90
exclude_also = [
    'if __name__ == "__main__":',
]
  • [tool.coverage.run] controls measuring: source is what to watch (so a bare --cov with no value is enough), branch = true is --cov-branch.
  • [tool.coverage.report] controls reporting: fail_under is --cov-fail-under, and exclude_also lists patterns to exclude in addition to # pragma: no cover — here every __main__ block, with no comment needed in the code.

Now a plain pytest measures, reports and enforces. You will see it run in the complete example below.

100% coverage does not mean correct

Coverage answers "did this line run?". It does not answer "did anyone check the result was right?". A small example:

python
def is_leap_year(year: int) -> bool:
    return year % 4 == 0
python
from shop.calendar_rules import is_leap_year


def test_2024_is_a_leap_year():
    assert is_leap_year(2024)


def test_2023_is_not():
    assert not is_leap_year(2023)
text
$ pytest -q --cov=shop --cov-branch --cov-report=term-missing
..                                                                       [100%]
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name                     Stmts   Miss Branch BrPart  Cover   Missing
--------------------------------------------------------------------
shop/__init__.py             0      0      0      0   100%
shop/calendar_rules.py       2      0      0      0   100%
--------------------------------------------------------------------
TOTAL                        2      0      0      0   100%
2 passed in 0.12s

Perfect score. And the function is wrong: a year divisible by 100 is not a leap year unless it is also divisible by 400. One more test, chosen by thinking about the rule rather than about lines:

python
def test_1900_is_not():
    assert not is_leap_year(1900)
text
$ pytest -q --cov=shop --cov-branch --cov-report=term-missing
..F                                                                      [100%]
=================================== FAILURES ===================================
_______________________________ test_1900_is_not _______________________________

    def test_1900_is_not():
>       assert not is_leap_year(1900)
E       assert not True
E        +  where True = is_leap_year(1900)

tests/test_calendar_rules.py:13: AssertionError
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name                     Stmts   Miss Branch BrPart  Cover   Missing
--------------------------------------------------------------------
shop/__init__.py             0      0      0      0   100%
shop/calendar_rules.py       2      0      0      0   100%
--------------------------------------------------------------------
TOTAL                        2      0      0      0   100%
=========================== short test summary info ============================
FAILED tests/test_calendar_rules.py::test_1900_is_not - assert not True
1 failed, 2 passed in 0.12s

Still 100% — and now a failure. Coverage could not have pointed at this bug, because there was no missing line; the missing thing was a case. Treat coverage as a map of where you have certainly not looked, never as proof that what you looked at is right.

Running tests in parallel with pytest-xdist

As a suite grows it gets slow, and a slow suite gets run less often. pytest-xdist spreads tests over several processes. Eight tests that each take half a second:

python
import time

import pytest


@pytest.mark.parametrize("n", range(8))
def test_slow_check(n):
    time.sleep(0.5)
    assert n >= 0
text
$ pytest -q test_slow.py
........                                                                 [100%]
8 passed in 4.10s

$ pytest -n auto test_slow.py
============================= test session starts ==============================
created: 4/4 workers
4 workers [8 items]

........                                                                 [100%]
============================== 8 passed in 1.52s ===============================

-n auto starts one worker per CPU core (four on this machine) and hands each one tests to run. It works together with --cov; the coverage data from all workers is combined into one report.

There is a price. Each worker is a separate process, and the order in which tests run is no longer the order in the file. Any test that quietly relied on another one having run first will break:

python
CART = []


def test_add_item():
    CART.append("pen")
    assert CART == ["pen"]


def test_cart_has_one_item():
    assert len(CART) == 1
text
$ pytest -q test_cart.py
..                                                                       [100%]
2 passed in 0.07s

$ pytest -n 2 test_cart.py
============================= test session starts ==============================
created: 2/2 workers
2 workers [2 items]

.F                                                                       [100%]
=================================== FAILURES ===================================
____________________________ test_cart_has_one_item ____________________________
[gw1] linux -- Python 3.12.3 /path/to/venv/bin/python

    def test_cart_has_one_item():
>       assert len(CART) == 1
E       assert 0 == 1
E        +  where 0 = len([])

test_cart.py:10: AssertionError
=========================== short test summary info ============================
FAILED test_cart.py::test_cart_has_one_item - assert 0 == 1
========================= 1 failed, 1 passed in 0.37s ==========================

The second test only passed because the first one had already filled the shared list. On a different worker, the list was empty. Running pytest test_cart.py::test_cart_has_one_item alone fails the same way — xdist did not create the bug, it revealed it.

The fix is the one from the fixtures chapters: give each test its own state.

python
import pytest


@pytest.fixture
def cart():
    return ["pen"]


def test_add_item(cart):
    cart.append("bag")
    assert cart == ["pen", "bag"]


def test_cart_has_one_item(cart):
    assert len(cart) == 1
text
$ pytest -n 2 test_cart_fixed.py
============================= test session starts ==============================
created: 2/2 workers
2 workers [2 items]

..                                                                       [100%]
============================== 2 passed in 0.39s ===============================

A test should pass alone, in any order, on any worker. There is a plugin, pytest-randomly, built entirely on that idea: it shuffles the order of tests on every run, so a hidden dependence shows up early, on your machine, rather than months later in CI. You do not need it to follow this course; it is enough to know that "the suite passes only in file order" is a bug.

Running it on every push: GitHub Actions

The last step is to stop relying on people remembering to run the suite. Declare the test tools as a development group in pyproject.toml, so uv sync installs them:

toml
[dependency-groups]
dev = [
    "pytest>=9",
    "pytest-cov>=7",
    "pytest-xdist>=3.8",
]

Then add .github/workflows/tests.yml:

yaml
name: tests

on:
  push:
    branches: [main]
  pull_request:

jobs:
  test:
    runs-on: ubuntu-latest
    strategy:
      fail-fast: false
      matrix:
        python-version: ["3.12", "3.13", "3.14"]

    steps:
      - uses: actions/checkout@v7

      - uses: astral-sh/setup-uv@v10
        with:
          python-version: ${{ matrix.python-version }}

      - name: Install dependencies
        run: uv sync --locked

      - name: Run tests
        run: uv run pytest -n auto --cov-report=xml

      - name: Keep the coverage report
        if: matrix.python-version == '3.12'
        uses: actions/upload-artifact@v7
        with:
          name: coverage-xml
          path: coverage.xml

What each part does:

  • on: — run on every push to main and on every pull request.
  • matrix: — the job runs three times, once per Python version, side by side. fail-fast: false lets all three finish, so you see whether a failure is specific to one version.
  • setup-uv installs uv and pins the Python version for that copy of the job.
  • uv sync --locked installs exactly what uv.lock says, and fails if the lock file is out of date — so CI tests what you committed.
  • uv run pytest -n auto --cov-report=xml runs in parallel; coverage settings, including fail_under, come from pyproject.toml, so a drop in coverage fails the build. The XML report sits alongside the terminal one, ready for a coverage service or for download.

Locally, after uv sync, the same command gives the same result (the last lines of the output):

text
$ uv run pytest -n auto --cov-report=xml
TOTAL                     15      0     10      1    96%
Coverage XML written to file coverage.xml
Required test coverage of 90.0% reached. Total coverage: 96.00%
============================== 7 passed in 0.43s ===============================

A complete example

A project with a src layout, both modules from this chapter, and every setting in one place.

text
shop-project/
├── pyproject.toml
├── src/
│   └── shop/
│       ├── __init__.py
│       ├── coupons.py
│       └── shipping.py
└── tests/
    ├── test_coupons.py
    └── test_shipping.py

pyproject.toml is the one from the configuration section above. coupons.py keeps its __main__ block, with no pragma — exclude_also handles it. test_coupons.py still has only the SAVE10 test. And test_shipping.py is now the test plan from the start of the chapter, row for row:

python
import pytest

from shop.shipping import shipping_cost


@pytest.mark.parametrize(
    ("weight", "country", "expected"),
    [
        (1, "BD", 60.0),
        (1, "IN", 90.0),
        (1, "US", 500.0),
        (5, "BD", 60.0),
        (7, "BD", 100.0),
    ],
)
def test_shipping_cost(weight, country, expected):
    assert shipping_cost(weight, country) == expected


def test_zero_weight_is_rejected():
    with pytest.raises(ValueError, match="positive"):
        shipping_cost(0, "BD")
text
$ pytest
============================= test session starts ==============================
configfile: pyproject.toml
testpaths: tests
collected 7 items

tests/test_coupons.py .                                                  [ 14%]
tests/test_shipping.py ......                                            [100%]

================================ tests coverage ================================
_______________ coverage: platform linux, python 3.12.3-final-0 ________________

Name                   Stmts   Miss Branch BrPart  Cover   Missing
------------------------------------------------------------------
src/shop/__init__.py       0      0      0      0   100%
src/shop/coupons.py        4      0      2      1    83%   2->4
src/shop/shipping.py      11      0      8      0   100%
------------------------------------------------------------------
TOTAL                     15      0     10      1    96%
Required test coverage of 90.0% reached. Total coverage: 96.00%
============================== 7 passed in 0.12s ===============================

Three things are worth noticing.

First, plain pytest did everything — no flags. Measuring, branch checking, the missing list and the threshold all came from pyproject.toml.

Second, shipping.py went from 64% to 100% because the plan was written down, not because we chased red lines. Coverage then confirms the plan reached every line and every branch. The (5, "BD", 60.0) row and the choice of 0 rather than -1 for the invalid weight add nothing to the percentage — they are there because the contract has edges, and edges are where typos like < for <= live.

Third, the build passes at 96% while 2->4 is still listed. The threshold is a floor, not a finish line. The report is still telling you exactly which test to write next: final_price with no coupon.


When it breaks

error: unrecognized arguments: --cov=shop pytest-cov is not installed in the environment pytest is running from. Install it (uv add --dev pytest-cov) and run through that environment, for example with uv run pytest. The same message with -n means pytest-xdist is missing.

CoverageWarning: Module shopp was never imported. (module-not-imported) followed by No data was collected. The name after --cov= does not match any package your tests imported — here a typo. Check the spelling, and with a src layout give the import name (shop), not the folder (src).

WARNING: Failed to generate report: No data to report. Same cause as above: nothing was measured, so there is nothing to report. The tests still passed, which is exactly why this one is easy to miss.

FAIL Required test coverage of 90% not reached. Total coverage: 83.33% Not a broken test — the threshold. Look at the Missing column and write the test it points at. Lowering fail_under to make the build green defeats the point of having it.

A test passes on its own but fails under -n auto (or the reverse) The tests share state: a module-level list or dict, a file at a fixed path, an environment variable, a database row. Find what the failing test assumes is already there, and give it to the test through a fixture, tmp_path or monkeypatch instead.

Coverage is 100% but users still hit bugs Expected. Coverage shows what ran, not what was checked. Look for cases — boundaries, empty inputs, the year 1900 — not lines.