Chapter 05

Floats and pytest.approx — testing "close enough"

Why 0.1 + 0.2 == 0.3 is false, what pytest.approx tolerates by default, when you need rel and abs, and why money wants Decimal rather than approx. With NaN, unordered results and timestamps.

35 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

Ask Python something a child could answer:

python
print(0.1 + 0.2)
print(0.1 + 0.2 == 0.3)
text
0.30000000000000004
False

That is not a bug in your machine, and it is not a bug in Python. It is how almost every programming language stores decimal fractions. And it walks straight into your tests.

test_total.py:

python
def total(prices):
    return sum(prices)


def test_total():
    assert total([0.1, 0.2]) == 0.3

Then pytest -q:

text
F                                                                        [100%]
=================================== FAILURES ===================================
__________________________________ test_total __________________________________

    def test_total():
>       assert total([0.1, 0.2]) == 0.3
E       assert 0.30000000000000004 == 0.3
E        +  where 0.30000000000000004 = total([0.1, 0.2])

test_total.py:6: AssertionError
=========================== short test summary info ============================
FAILED test_total.py::test_total - assert 0.30000000000000004 == 0.3
1 failed in 0.01s

The function is correct. The test is wrong. It asks for an exactness that floating-point numbers cannot give. This chapter is about asking the right question instead: is it close enough? It is also about the cases where "close enough" is the wrong question, and money is the main one.

By the end of this chapter you can

  • Explain in two sentences why 0.1 + 0.2 == 0.3 is False
  • Compare floats in a test with pytest.approx, and read its default tolerance
  • Set your own tolerance with rel= and abs=, and know which one you need near zero
  • Use approx on lists, tuples and dicts, and read the report it gives when they differ
  • Choose between pytest.approx, math.isclose and Decimal
  • Test NaN, results whose order is not promised, and timestamps

Prerequisites: testing exceptions.


Before you write the test

With numbers, the most important decision comes before any code: for each result, what kind of equality has the code promised? Pick wrong and the test is either flaky (too strict) or blind (too loose). Settle three things first.

1. The contract. Take the small stats.py module that this chapter ends with. In plain words, it promises:

  • mean(values) returns the average of a list of floats, and nan for an empty list.
  • shares(counts) turns counts into fractions that add up to one.
  • with_vat(price) adds 15% VAT to a Decimal price, rounded half-up to the cent. A float price is refused with a TypeError.
  • countries(orders) returns each country once. It does not promise any order.

Each line already tells you how to compare. "Average of floats" means a tolerance. "Rounded to the cent" means exact. "Refused" means pytest.raises. "Not in any order" means sort or use a set before comparing.

2. The setup. Nothing new to install. You need the venv from chapter one with pytest in it, the module importable from the test file (both in one folder, run pytest from there), and two standard library modules, math and decimal. No fixtures, no files, no network.

3. The plan. Write the cases down before writing the tests. For every row, decide the comparison too, not just the expected value:

| Case | Input | Expected | Compare with | | --- | --- | --- | --- | | Happy path, float arithmetic | mean([0.1, 0.2, 0.3]) | 0.2 | approx (default tolerance) | | Boundary: result near zero | mean([0.1, 0.2, -0.3]) | 0.0 | approx(0.0, abs=1e-9) | | Edge: empty input | mean([]) | nan | math.isnan | | Container of floats | shares({"tea": 1, "coffee": 2}) | {"tea": 1/3, "coffee": 2/3} | approx on the dict | | Money, happy path | with_vat(Decimal("19.99")) | Decimal("22.99") | exact == | | Money, rounding boundary | with_vat(Decimal("0.10")) | Decimal("0.12") (0.115 rounds up) | exact == | | Invalid input | with_vat(19.99) | TypeError | pytest.raises | | Order not promised | countries(...) with a duplicate | ["BD", "IN"] | sorted(...) == |

What not to test. Do not test that 0.1 + 0.2 is 0.30000000000000004. That is Python's behaviour, not yours, and pinning it makes your test a test of the float format. Do not test the order of a result that comes from a set, nor the exact microsecond of a timestamp. Neither is promised. And do not pick a tolerance by widening it until the test turns green. A tolerance is a statement about the problem ("half a degree", "one percent"), never about the test.

The sections below teach each tool in that table. The complete example at the end implements the plan row by row.


Why 0.1 + 0.2 is not 0.3

A float is stored in binary, in a fixed amount of space (64 bits). Some fractions have no exact binary form, just as 1/3 has no exact decimal form: 0.3333… goes on forever, and at some point you have to stop writing. 0.1 is one of those fractions in binary. Python keeps the nearest number it can hold.

You can see the stored values by asking for more digits than Python normally shows:

python
print(f"{0.1:.20f}")
print(f"{0.2:.20f}")
print(f"{0.3:.20f}")
print(f"{0.1 + 0.2:.20f}")
text
0.10000000000000000555
0.20000000000000001110
0.29999999999999998890
0.30000000000000004441

0.1 is stored a hair too large, and so is 0.2. Adding them adds the two errors. Meanwhile 0.3 is stored a hair too small. The two results land on neighbouring floats, and == compares floats bit by bit, so it says False.

That leads to the one rule of this chapter: never use == on a float that came out of arithmetic. A float you typed in yourself and passed through unchanged is fine. A float that was added, divided or multiplied should be compared with a tolerance.

pytest.approx — equal, within a tolerance

Change one line of the test:

python
import pytest


def total(prices):
    return sum(prices)


def test_total():
    assert total([0.1, 0.2]) == pytest.approx(0.3)
text
.                                                                        [100%]
1 passed in 0.01s

pytest.approx(0.3) builds an object that is "equal" to any number close enough to 0.3. How close? Print one and it tells you:

python
import pytest

print(pytest.approx(0.3))
print(pytest.approx(250.0))
print(pytest.approx(1_000_000.0))
print(pytest.approx(0.0))
text
0.3 ± 3.0e-07
250.0 ± 2.5e-04
1000000.0 ± 1
0.0 ± 1.0e-12

The defaults are two tolerances, and the larger of the two wins:

  • relative rel=1e-6: one part in a million of the expected value. For 0.3 that is 0.0000003; for a million it is 1.
  • absolute abs=1e-12: a fixed floor, so the tolerance never quite reaches zero.

The relative tolerance is the one that matters almost everywhere, because it grows and shrinks with the number. An error of 0.0000003 is noise next to 0.3; next to 1_000_000 it would be absurdly strict.

Near zero, the picture changes. One part in a million of zero is zero, so only the tiny absolute floor is left:

python
import pytest

print(0.000001 == pytest.approx(0.0))
print(0.000001 == pytest.approx(0.0, abs=1e-5))
text
False
True

A millionth looks like zero to a person, but not to approx(0.0). If a result is expected to be zero, or near it, give approx an absolute tolerance yourself.

Choosing your own tolerance: rel= and abs=

The defaults suit "the same calculation, done slightly differently". When your code rounds, measures or estimates, decide what "close enough" means and say it:

python
import pytest

print(100.4 == pytest.approx(100, abs=0.5))
print(100.6 == pytest.approx(100, abs=0.5))
print(100.6 == pytest.approx(100, rel=0.01))
print(pytest.approx(100, abs=0.5))
print(pytest.approx(100, rel=0.01))
print(pytest.approx(100, rel=0.01, abs=5))
text
True
False
True
100 ± 0.5
100 ± 1
100 ± 5
  • abs=0.5 means "within half a unit, whatever the size of the number".
  • rel=0.01 means "within one percent of the expected value". For 100 that is ± 1.
  • Give both and, as with the defaults, the larger tolerance wins: here abs=5 beats one percent of 100.

Choose with a sentence in mind. "The temperature reading may be off by half a degree" is abs=0.5. "The estimate may be off by one percent" is rel=0.01. And when the expected value is zero, only abs can help.

The wrong way is to start from a failing test and widen the tolerance until it goes green. A tolerance that was chosen to make the test pass will pass a bug just as happily:

python
import pytest

print(0.95 == pytest.approx(1.0, rel=0.1))
text
True

A result that is five percent wrong now counts as "equal". If the code should be accurate to rounding error, keep the default. A loose tolerance needs a reason you can say out loud, and that reason belongs in a comment next to it.

Lists, tuples and dicts

approx also accepts a container of numbers and compares it element by element, each with its own tolerance:

python
import pytest

shares = [0.1 + 0.2, 0.7]
point = (1 / 3, 2 / 3)
rates = {"tax": 0.1 + 0.05, "tip": 0.1}

print(shares == pytest.approx([0.3, 0.7]))
print(point == pytest.approx((0.333333, 0.666667), rel=1e-5))
print(rates == pytest.approx({"tax": 0.15, "tip": 0.1}))
print(pytest.approx([0.3, 0.7]))
text
True
True
True
approx([0.3 ± 3.0e-07, 0.7 ± 7.0e-07])

A list compares by position. A dict compares by key, so the order of the keys does not matter, but the keys themselves must match exactly. The lengths must match too.

The real gain shows up when a comparison fails. test_shares.py:

python
import pytest


def test_shares():
    assert [0.1 + 0.2, 0.25, 0.6] == pytest.approx([0.3, 0.2, 0.6])
text
F                                                                        [100%]
=================================== FAILURES ===================================
_________________________________ test_shares __________________________________

    def test_shares():
>       assert [0.1 + 0.2, 0.25, 0.6] == pytest.approx([0.3, 0.2, 0.6])
E       assert [0.3000000000...04, 0.25, 0.6] == approx([0.3 ±....6 ± 6.0e-07])
E
E         comparison failed. Mismatched elements: 1 / 3:
E         Max absolute difference: 0.04999999999999999
E         Max relative difference: 0.19999999999999996
E         Index | Obtained | Expected
E         1     | 0.25     | 0.2 ± 2.0e-07

test_shares.py:5: AssertionError
=========================== short test summary info ============================
FAILED test_shares.py::test_shares - assert [0.3000000000...04, 0.25, 0.6] ==...
1 failed in 0.01s

Read it from the bottom of the E lines up. One element of three did not match. The table says which one (index 1), what came out (0.25) and what was expected, with its tolerance. Element 0, the 0.30000000000000004, is not listed. It was close enough, so it is not the problem. With a dict, the Index column holds the key instead.

What approx is not for

approx exists for floats. Hand it something else and it quietly falls back to plain ==:

python
import pytest

print("abc" == pytest.approx("abc"))
print(pytest.approx("abc"))
print(10 == pytest.approx(10))
print(10.000001 == pytest.approx(10))
text
True
abc
True
True

For a string, approx adds nothing, and there is no tolerance to show. For a whole number, it adds something you probably did not want. If count_items() is supposed to return 10, then 10.000001 is a bug, and approx(10) lets it through. Integers, strings, booleans and None are compared with ==.

math.isclose — the standard library version

Python has its own tolerance check, math.isclose. Its defaults are different: rel_tol=1e-9 (stricter than approx) and abs_tol=0.0 (no floor at all):

python
import math

print(math.isclose(0.1 + 0.2, 0.3))
print(math.isclose(0.000001, 0.0))
print(math.isclose(0.000001, 0.0, abs_tol=1e-5))
print(math.isclose(100.6, 100, rel_tol=0.01))
text
True
False
True
True

With abs_tol=0.0, nothing except exactly 0.0 is ever close to zero. The same lesson as before, only sharper.

In application code, math.isclose is the right tool, because pytest is not installed in production. In a test, prefer approx, and the reason is the report. test_close.py:

python
import math

import pytest


def test_with_isclose():
    assert math.isclose(0.3001, 0.3)


def test_with_approx():
    assert 0.3001 == pytest.approx(0.3)
text
FF                                                                       [100%]
=================================== FAILURES ===================================
______________________________ test_with_isclose _______________________________

    def test_with_isclose():
>       assert math.isclose(0.3001, 0.3)
E       assert False
E        +  where False = <built-in function isclose>(0.3001, 0.3)
E        +    where <built-in function isclose> = math.isclose

test_close.py:7: AssertionError
_______________________________ test_with_approx _______________________________

    def test_with_approx():
>       assert 0.3001 == pytest.approx(0.3)
E       assert 0.3001 == 0.3 ± 3.0e-07
E
E         comparison failed
E         Obtained: 0.3001
E         Expected: 0.3 ± 3.0e-07

test_close.py:11: AssertionError
=========================== short test summary info ============================
FAILED test_close.py::test_with_isclose - assert False
FAILED test_close.py::test_with_approx - assert 0.3001 == 0.3 ± 3.0e-07
2 failed in 0.01s

assert False tells you that the numbers differ. approx also tells you by how much, and what tolerance was allowed. math.isclose also has no answer for lists or dicts.

Money is not a tolerance problem

When a test on prices fails with 0.30000000000000004, reaching for approx is tempting. Look at what that tolerance means for a large invoice:

python
import pytest

expected = 1_000_000.00
charged = 1_000_000.99

print(charged == pytest.approx(expected))
text
True

Ninety-nine cents of difference, and the test passes. One part in a million of a million is one whole unit. For money, "close enough" is not a correct answer. A customer charged one cent too much has been charged the wrong amount.

The fix is not in the test. The code should not hold money in a float to begin with. Python's decimal.Decimal stores decimal digits exactly as written:

python
from decimal import Decimal

print(Decimal("0.10") + Decimal("0.20"))
print(Decimal("0.10") + Decimal("0.20") == Decimal("0.30"))
print(Decimal("1000000.99") == Decimal("1000000.00"))
print(Decimal(0.1))
text
0.30
True
False
0.1000000000000000055511151231257827021181583404541015625

With Decimal, == is exact again, which is what a money test needs. The last line is the trap. Decimal(0.1) is built from a float, so it copies the float's error faithfully. Always build a Decimal from a string.

Rounding to cents is explicit, and you choose the rule:

python
from decimal import ROUND_HALF_UP, Decimal

price = Decimal("19.99")
vat = price * Decimal("0.15")

print(vat)
print(vat.quantize(Decimal("0.01"), rounding=ROUND_HALF_UP))
text
2.9985
3.00

So the split is simple. Measurements (lengths, averages, ratios, sensor readings) are floats, tested with approx. Amounts of money are Decimal, tested with ==.

NaN is not equal to itself

float("nan") means "not a number", the result of something like 0 * inf or a missing reading. By definition it is equal to nothing, itself included:

python
import math

import pytest

missing = float("nan")

print(missing == missing)
print(missing == pytest.approx(float("nan")))
print(missing == pytest.approx(float("nan"), nan_ok=True))
print(math.isnan(missing))
text
False
False
True
True

approx follows that rule unless you pass nan_ok=True, which says "a NaN here is expected, and matches a NaN". For a single value, assert math.isnan(result) reads more plainly. nan_ok=True earns its place on containers, where one NaN sits among ordinary numbers: [1.5, nan, 2.0] == pytest.approx([1.5, nan, 2.0], nan_ok=True) is True.

Order you did not promise, and time that will not hold still

Floats are not the only values that come out "nearly" right. Two more cases need the same thinking: decide what is promised, and test only that.

Order. This function collects the tags of some posts through a set:

tags.py:

python
def unique_tags(posts):
    tags = set()
    for post in posts:
        tags.update(post["tags"])
    return list(tags)

test_tags.py:

python
from tags import unique_tags

POSTS = [
    {"tags": ["python", "testing"]},
    {"tags": ["testing", "floats"]},
]


def test_order_is_a_guess():
    assert unique_tags(POSTS) == ["floats", "python", "testing"]


def test_sorted():
    assert sorted(unique_tags(POSTS)) == ["floats", "python", "testing"]


def test_as_a_set():
    assert set(unique_tags(POSTS)) == {"floats", "python", "testing"}

The order of strings in a set changes from run to run, because string hashing is randomised per process. Fixing the seed makes that visible. PYTHONHASHSEED=1 pytest -q:

text
F..                                                                      [100%]
=================================== FAILURES ===================================
____________________________ test_order_is_a_guess _____________________________

    def test_order_is_a_guess():
>       assert unique_tags(POSTS) == ["floats", "python", "testing"]
E       AssertionError: assert ['python', 't...ng', 'floats'] == ['floats', 'p...n', 'testing']
E
E         At index 0 diff: 'python' != 'floats'
E         Use -v to get more diff

test_tags.py:10: AssertionError
=========================== short test summary info ============================
FAILED test_tags.py::test_order_is_a_guess - AssertionError: assert ['python'...
1 failed, 2 passed in 0.01s

Run it again with PYTHONHASHSEED=2 and all three pass. Same code, same test, different result. That is a flaky test, and it is worse than a failing one. If the function does not promise an order, do not test one. Compare sorted(...), or compare as a set. One caution: a set also hides duplicates. If duplicates matter, use sorted, or collections.Counter, which counts them.

Time. A timestamp taken inside your code can never equal one taken in the test, because time passes between the two calls. Test a window instead:

orders.py:

python
from datetime import datetime, timezone


def make_order(item):
    return {"item": item, "created_at": datetime.now(timezone.utc)}

test_orders.py:

python
from datetime import datetime, timezone

from orders import make_order


def test_created_between_before_and_after():
    before = datetime.now(timezone.utc)
    order = make_order("pen")
    after = datetime.now(timezone.utc)

    assert before <= order["created_at"] <= after
text
.                                                                        [100%]
1 passed in 0.01s

The test is exact and needs no tolerance at all: the stamp must fall between two moments you recorded. If you prefer a tolerance, approx accepts a datetime when abs= is a timedelta: order["created_at"] == pytest.approx(now, abs=timedelta(seconds=1)). When a test has to pin time to one exact value, the answer is to control the clock, and that comes later in the course with monkeypatch.


A complete example

stats.py:

python
import math
from decimal import ROUND_HALF_UP, Decimal

CENT = Decimal("0.01")


def mean(values):
    if not values:
        return math.nan
    return sum(values) / len(values)


def shares(counts):
    total = sum(counts.values())
    return {name: count / total for name, count in counts.items()}


def with_vat(price, rate=Decimal("0.15")):
    return (price * (1 + rate)).quantize(CENT, rounding=ROUND_HALF_UP)


def countries(orders):
    return list({order["country"] for order in orders})

test_stats.py:

python
import math
from decimal import Decimal

import pytest

from stats import countries, mean, shares, with_vat


def test_mean_of_measurements():
    # 0.6000000000000001 / 3 -- close, never exact
    assert mean([0.1, 0.2, 0.3]) == pytest.approx(0.2)


def test_mean_near_zero():
    # (0.1 + 0.2 - 0.3) / 3 is about 1.9e-17, not 0.0
    assert mean([0.1, 0.2, -0.3]) == pytest.approx(0.0, abs=1e-9)


def test_mean_of_nothing_is_nan():
    assert math.isnan(mean([]))


def test_shares_are_fractions():
    result = shares({"tea": 1, "coffee": 2})
    assert result == pytest.approx({"tea": 1 / 3, "coffee": 2 / 3})
    assert sum(result.values()) == pytest.approx(1.0)


def test_shares_to_two_places():
    result = shares({"tea": 1, "coffee": 2})
    assert result == pytest.approx({"tea": 0.33, "coffee": 0.67}, abs=0.005)


def test_vat_is_exact_to_the_cent():
    assert with_vat(Decimal("19.99")) == Decimal("22.99")
    assert with_vat(Decimal("0.10")) == Decimal("0.12")


def test_vat_refuses_a_float():
    with pytest.raises(TypeError):
        with_vat(19.99)


def test_countries_ignores_order_and_duplicates():
    orders = [{"country": "BD"}, {"country": "IN"}, {"country": "BD"}]
    assert sorted(countries(orders)) == ["BD", "IN"]

Then pytest -v:

text
============================= test session starts ==============================
collecting ... collected 8 items

test_stats.py::test_mean_of_measurements PASSED                          [ 12%]
test_stats.py::test_mean_near_zero PASSED                                [ 25%]
test_stats.py::test_mean_of_nothing_is_nan PASSED                        [ 37%]
test_stats.py::test_shares_are_fractions PASSED                          [ 50%]
test_stats.py::test_shares_to_two_places PASSED                          [ 62%]
test_stats.py::test_vat_is_exact_to_the_cent PASSED                      [ 75%]
test_stats.py::test_vat_refuses_a_float PASSED                           [ 87%]
test_stats.py::test_countries_ignores_order_and_duplicates PASSED        [100%]

============================== 8 passed in 0.01s ===============================

Eight tests, one for each row of the plan in "Before you write the test". Each picks its comparison on purpose:

  • mean is arithmetic on floats, so approx with its defaults.
  • A mean that should be zero comes out as 1.9e-17. Relative tolerance is useless at zero, so the test gives an abs= it can justify: anything below a billionth is rounding noise here.
  • An empty mean is NaN by design, so math.isnan, because == could never pass.
  • shares returns a dict of floats, so approx on the whole dict. The second test states its tolerance (abs=0.005, half a hundredth) because it compares against numbers rounded to two places.
  • with_vat is money, so Decimal and exact ==. 0.10 × 1.15 = 0.115 is the case that proves the rounding rule: ROUND_HALF_UP makes it 0.12.
  • A float price is refused rather than silently converted. That refusal is part of the contract, so it gets its own test with pytest.raises, as in the previous chapter.
  • countries promises no order, so the test sorts before comparing.

When it breaks

assert 0.30000000000000004 == 0.3 A computed float compared with ==. Wrap the expected value: == pytest.approx(0.3). If the value is money, fix the code to use Decimal instead.

AssertionError: approx() is not supported in a boolean context. You wrote assert pytest.approx(x) with no comparison. approx only means something on one side of ==, as the message suggests: assert a == approx(b).

TypeError: '>' not supported between instances of 'ApproxScalar' and 'float' approx supports only == and !=. For "at least" or "at most", compare the plain numbers: assert result > 0.2.

Impossible to compare lists with different sizes. The list and the approx list have different lengths. A tolerance applies to values, never to how many there are. The report gives both lengths on the next line.

assert nan == nan ± ??? A NaN on both sides, and NaN never equals anything. Use math.isnan(result), or pass nan_ok=True if a NaN is the expected value.

*`TypeError: unsupported operand type(s) for : 'decimal.Decimal' and 'float'** Decimal refuses to mix with float, on purpose, so an inexact number cannot sneak in. Write the other operand as a Decimal too: Decimal("1.15")`.

Decimal('0.1000000000000000055511151231257827021181583404541015625') in a report A Decimal was built from a float, Decimal(0.1), and inherited its error. Build it from a string: Decimal("0.1").

A test that passes on one run and fails on the next Look for order coming from a set, or for a timestamp compared with ==. Compare sorted(...) or a set, and test time as a window.