Chapter 18

Comprehensions — a loop on one line

List, dictionary and set comprehensions, the difference between a filtering if and a choosing if/else, and when a comprehension is the wrong choice.

30 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

One shape of code has kept reappearing over the last six chapters:

python
marks = [72, 45, 90, 33]

doubled = []
for mark in marks:
    doubled.append(mark * 2)
print(doubled)

doubled = [mark * 2 for mark in marks]
print(doubled)
text
[144, 90, 180, 66]
[144, 90, 180, 66]

The four lines above and the one line below do exactly the same thing.

Three of those four lines are not the work, they are the bookkeeping — make an empty list, walk over the source, append each time. The only real decision is mark * 2. Everything else is the same every time.

A comprehension removes that repetition. It is not a new ability — it is what you could already write, in fewer words. Which is also where its limit comes from, and that is the subject of this chapter's last section: shorten what becomes clearer for being shortened, not everything.

By the end of this chapter you can

  • Write and read a list comprehension
  • Filter with if, and choose with if/else
  • Write dictionary and set comprehensions
  • Say when a comprehension is the wrong choice
  • Take an unreadable comprehension apart back into a loop

Prerequisites: Nested data — a list of dictionaries.


The shape

text
[ what I want   for  each thing   in  where it comes from ]

Read aloud it is an ordinary sentence: "`mark 2 for each mark in marks`"*.

One way to hold it: read the for part first, then the beginning. Knowing where things come from makes what is being built easy to see.

Filtering

An if on the end keeps only the things that match:

python
marks = [72, 45, 90, 33]

passed = [mark for mark in marks if mark >= 40]
print(passed)
text
[72, 45, 90]

Here the first part is just mark — no calculation, taken as it is. Filtering and transforming can be done together:

python
names = ["  rafi ", "AHMED", " dia"]

clean = [name.strip().lower() for name in names if name.strip()]
print(clean)
text
['rafi', 'ahmed', 'dia']

The condition is name.strip() with no comparison at all. Recall chapter eight's truthiness rule: empty text is false. So somebody who typed only spaces is dropped.

if at the end versus if at the start

Two different things, and the position is the difference:

python
marks = [72, 33, 90]

labels = ["pass" if m >= 40 else "fail" for m in marks]
print(labels)
text
['pass', 'fail', 'pass']

An if at the end filters — some things never reach the result, so the result can be shorter.

An if ... else at the start chooses — every thing produces something, so the length is unchanged. It is the one-line condition from chapter fifteen, inside a comprehension.

Does the length change? That question is the easiest way to tell the two apart.

Dictionaries and sets too

The same shape, different brackets:

python
marks = {"rafi": 72, "ahmed": 45, "bilal": 90}

passed = {name: mark for name, mark in marks.items() if mark >= 50}
print(passed)

doubled = {name: mark * 2 for name, mark in marks.items()}
print(doubled)
text
{'rafi': 72, 'bilal': 90}
{'rafi': 144, 'ahmed': 90, 'bilal': 180}

Curly brackets and a colon make a dictionary comprehension. marks.items() hands back a tuple each time round and name, mark unpacks it, exactly as in a loop.

The same curly brackets without a colon make a set:

python
orders = [
    {"item": "pen"},
    {"item": "bag"},
    {"item": "pen"},
]

items = {order["item"] for order in orders}
print(sorted(items))
text
['bag', 'pen']

Chapter sixteen's "how many different" question in one line. pen arrives twice and appears once.

Numbers into text — which you have already seen

This line appeared in chapter thirteen with a promise that the explanation was coming:

python
marks = [72, 45, 90]
print(", ".join(str(m) for m in marks))
print([str(m) for m in marks])
text
72, 45, 90
['72', '45', '90']

Now it can be read in full. str(m) for m in marks is a comprehension without the square brackets — that form is called a generator expression, and rather than building the whole list at once it hands join one item at a time. The brackets can be dropped when it is a function's only argument, and on large data it saves memory.

When not to write one

Being able to write a comprehension does not mean you should. Look at this:

python
rows = [["1", "2"], ["3", "x"]]

numbers = [int(c) for row in rows for c in row if c.isdigit()]
print(numbers)
text
[1, 2, 3]

The line is correct. But it has two fors and an if, and you have to remember the order: the fors read outermost to innermost, the way nested loops would be written. Six months later that line is work to read.

As a loop:

python
rows = [["1", "2"], ["3", "x"]]

numbers = []
for row in rows:
    for c in row:
        if c.isdigit():
            numbers.append(int(c))

print(numbers)
text
[1, 2, 3]

More lines, less thinking. The goal of code is not to be short but to be clear — and most of the time shortening makes it clearer, but not always.

As a working rule: one for and at most one if — a comprehension. More than that — a loop. And never put a print or an append inside a comprehension; it is a thing for building a value, not for doing work.


A complete example

views.py:

python
# One list of records in, three different views out
orders = [
    {"customer": "rafi", "item": "pen", "quantity": 3, "price": 15.0},
    {"customer": "ahmed", "item": "bag", "quantity": 1, "price": 850.0},
    {"customer": "rafi", "item": "ink", "quantity": 2, "price": 120.0},
    {"customer": "dia", "item": "pen", "quantity": 5, "price": 15.0},
]

totals = [order["quantity"] * order["price"] for order in orders]
print("Totals    :", totals)

big = [order["item"] for order in orders if order["quantity"] * order["price"] > 100]
print("Over 100  :", big)

items = {order["item"] for order in orders}
print("Distinct  :", sorted(items))

by_item = {order["item"]: order["price"] for order in orders}
print("Price list:", by_item)

print("Grand     :", sum(totals))
text
Totals    : [45.0, 850.0, 240.0, 75.0]
Over 100  : ['bag', 'ink']
Distinct  : ['bag', 'ink', 'pen']
Price list: {'pen': 15.0, 'bag': 850.0, 'ink': 120.0}
Grand     : 1210.0

The same data as chapter seventeen, with most of what was a loop there now one line.

Three things worth noticing.

totals and sum(totals) are separate lines. The total could have been sum(order["quantity"] * order["price"] for order in orders), but then the individual line values would not exist. An intermediate name is often worth more than a one-line trick.

The same calculation appears twice inside big — once in the condition, and it would be needed again in the result. Here the result only takes the name, so nothing is lost; but when you need both, that is a good reason to leave the comprehension and write a loop, where the value can be computed once and given a name.

by_item lost a record. Four orders, three pairs in the result — pen appears twice, and by chapter fifteen's rule the later one quietly replaced the earlier. Here both pens have the same price so no harm was done; with different prices it would be silently wrong. A comprehension does not create that problem, but being short gives it less chance to be noticed.


When it breaks

SyntaxError from putting if and else at the end [m for m in marks if m >= 40 else 0] is not valid. The filtering if goes at the end and has no else; the choosing if ... else goes at the front: [m if m >= 40 else 0 for m in marks].

The result is full of None Something inside the comprehension returns nothing — such as [items.append(x) for x in other]. append returns None, so the result is a list of None. A comprehension builds values; to do work, write an ordinary loop.

NameError — the comprehension's variable is gone outside it After [m * 2 for m in marks] you cannot use m. In chapter eleven a for loop's variable survived the loop; a comprehension's does not — it lives only inside.

A dictionary comprehension produced fewer pairs than expected The same key was produced more than once and the later one replaced the earlier. Compare with len().

I cannot read the line That is reason enough. Break it into a loop — the code will not run slower, and the next person to read it, probably you, will be grateful.