Sorting lists, and the methods that come with them
What separates sorted() from .sort(), the answers max/min/sum/count/index give, changing the rule with key=, and turning a list into readable text with join().
- 1Encounter
- 2Understand
- 3Worked
- 4Predict
- 5Apply
- 6Stretch
The problem we are solving
Seven people's marks are in a list. Three questions need answering: what is the highest, what is the average, and who are the top three?
With chapter twelve's tools the first two can be done with a loop — laborious, but possible. The third is different. "Top three" means putting them in order first, and we have no way to sort anything yet.
These jobs come up so often that Python ships them. Today's chapter is that list of tools — and one trap almost everyone falls into at least once: there are two forms of the same operation, one that changes the list and one that hands back a new one, and choosing the wrong one raises no error at all.
By the end of this chapter you can
- Say what separates
sorted()from.sort(), and which each job wants - Get answers out of a list with
max,min,sum,countandindex - Change the rule a sort uses with
key= - Turn a list into readable text with
join() - Explain why sorting a list of mixed types raises a
TypeError
Prerequisites: Lists — many values under one name.
sorted() and .sort() — the heart of this chapter
Both sort. The difference is what comes back.
marks = [72, 45, 90, 61]
ordered = sorted(marks)
print(ordered)
print(marks)
marks.sort()
print(marks)[45, 61, 72, 90]
[72, 45, 90, 61]
[45, 61, 72, 90]Read those three lines of output carefully.
sorted(marks) built a new sorted list and handed it back. The original marks is untouched — the second line is the proof.
marks.sort() built nothing. It sorted the original list, and the old order is gone for good.
It is chapter twelve's rule returning: a method that changes the list does not hand anything back.
marks = [72, 45, 90, 61]
marks = marks.sort()
print(marks)NoneExactly like append. With .sort() you must not keep the result; with sorted() you must, or the work is thrown away.
Which, when? If the original order is no longer needed, .sort() is simpler and cheaper. If the original list is used elsewhere, or you are not sure, use sorted() — it takes nothing away from anybody. When in doubt, sorted().
Reverse order, and numeric answers
marks = [72, 45, 90, 61]
print(sorted(marks, reverse=True))
print(max(marks), min(marks), sum(marks))
print(round(sum(marks) / len(marks), 2))[90, 72, 61, 45]
90 45 268
67.0max, min and sum are functions rather than methods, so it is max(marks) and not marks.max(). There is nothing special for an average: the total over the count is the average, and round(..., 2) trims it for display.
Notice that chapter ten's accumulating loop has collapsed into one line. Learning it was not wasted — what sum does is that loop, and one day you will need to accumulate something Python has no ready-made function for.
Counting and finding
marks = [72, 45, 90, 45, 61]
print(marks.count(45))
print(marks.index(45))
print(marks.index(90))2
1
2count(x) says how many times x appears. index(x) says where the first one is — with two 45s it reports index 1 and never mentions the second.
And looking for something that is not there stops the program:
marks = [72, 45, 90]
print(marks.index(50))ValueError: 50 is not in listJust like remove. Get into the habit of checking with if 50 in marks: first.
Sorting text, and key=
names = ["rafi", "Bilal", "ahmed"]
print(sorted(names))
print(sorted(names, key=str.lower))['Bilal', 'ahmed', 'rafi']
['ahmed', 'Bilal', 'rafi']People see the first result and assume Python got it wrong. It did not — it sorts by character code, and there every capital letter comes before every lowercase one. So Bilal jumped to the front.
key=str.lower says: while sorting, treat each name as though it were lowercase. The names themselves are unchanged — Bilal still prints with its capital B — only the comparison uses the lowercased form.
key can be any rule:
words = ["bag", "pen", "notebook", "ink"]
print(sorted(words, key=len))['bag', 'pen', 'ink', 'notebook']Sorted by length. Notice that bag, pen and ink are all three characters long and stayed in the order they were already in — Python does not disturb ties. That property is called a stable sort, and it is what makes sorting by several criteria possible.
Mixed types cannot be sorted
values = [3, "1", 2]
print(sorted(values))TypeError: '<' not supported between instances of 'str' and 'int'Sorting means comparing things in pairs, and chapter eight showed that a number and a piece of text cannot be compared with <. The message is word for word the same one. This happens routinely with lists read from a CSV file, where every number is really text.
From a list to text — join
Printing a list directly gives you brackets and quotes, which is not for showing to a person:
names = ["rafi", "ahmed", "bilal"]
print(", ".join(names))
print(" | ".join(names))rafi, ahmed, bilal
rafi | ahmed | bilalThe line reads backwards the first time: the joining text comes first, and the list goes in the brackets. The way to remember it — "join these with this".
There is one condition, and it is strict: every value has to be text.
marks = [72, 45, 90]
print(", ".join(marks))TypeError: sequence item 0: expected str instance, int foundTo join a list of numbers, turn each one into text first:
marks = [72, 45, 90]
print(", ".join(str(m) for m in marks))72, 45, 90That str(m) for m in marks is chapter eighteen's subject — for now just read it as "put str() around each value".
A complete example
report.py:
# A marks report: order them, summarise them, and show the top three
marks = [72, 45, 90, 61, 88, 45, 33]
ordered = sorted(marks, reverse=True)
print("Marks :", marks)
print("Ordered :", ordered)
print("Count :", len(marks))
print("Highest :", max(marks))
print("Lowest :", min(marks))
print("Total :", sum(marks))
print("Average :", round(sum(marks) / len(marks), 2))
print("How many 45s:", marks.count(45))
top_three = ordered[:3]
print("Top three:", top_three)
print("As text :", ", ".join(str(m) for m in top_three))
passed = []
for mark in marks:
if mark >= 40:
passed.append(mark)
print("Passed :", len(passed), "of", len(marks))Marks : [72, 45, 90, 61, 88, 45, 33]
Ordered : [90, 88, 72, 61, 45, 45, 33]
Count : 7
Highest : 90
Lowest : 33
Total : 434
Average : 62.0
How many 45s: 2
Top three: [90, 88, 72]
As text : 90, 88, 72
Passed : 6 of 7Two design decisions are worth noticing.
sorted() was chosen over .sort() because the original marks is printed again on the first line. With .sort(), the "Marks" and "Ordered" lines would have shown the same thing and nothing would have complained. That is precisely the silent mistake this chapter opened with.
The top three came out of a slice — ordered[:3]. No separate "top three" function was needed: the first three of a descending list are the top three. Once something is sorted, a surprising number of questions become one slice.
When it breaks
My list became None after sorting it The line said marks = marks.sort(). .sort() sorts the list itself and returns None — write just marks.sort(). For a new list, ordered = sorted(marks).
Sorting changed the original list too You used .sort(), which sorts in place. Use sorted() to leave the original alone.
TypeError: '<' not supported between instances of 'str' and 'int' The list mixes numbers and text. Run print(marks) and see which values have quotes around them; if it came from a file they are probably all text, and need int().
ValueError: 50 is not in list index() or remove() was given a value the list does not contain. Check with in first.
TypeError: sequence item 0: expected str instance, int found join() was handed a list of numbers. Wrap each in str().
After sorting names, all the capitalised ones came first That is correct behaviour, not a bug. To sort the way a person expects, sorted(names, key=str.lower).
Step 4 of 6 — Predict
Check your understanding
sorted() was used. What do the two lines print?
marks = [72, 45, 90]
ordered = sorted(marks)
print(ordered)
print(marks)- A[45, 72, 90] [72, 45, 90]
- B[45, 72, 90] [45, 72, 90]
- C[72, 45, 90] [45, 72, 90]
- D[45, 72, 90] None
The result of .sort() is stored back into marks. What is printed?
marks = [72, 45, 90]
marks = marks.sort()
print(marks)- ANone
- B[45, 72, 90]
- C[72, 45, 90]
- DAn `AttributeError`
An attempt to join a list of numbers with commas. What happens?
marks = [72, 45, 90]
print(", ".join(marks))- AA `TypeError` — `join` can only join text
- B72, 45, 90
- C724590
- D[72, 45, 90]
Answering needs an account
Sign in to check your answers
The questions are above, and working them out in your head is the part that matters. Sign in to see the answers, the explanations and the three-level hints.
Your turn
Write a file called inventory.py holding a list of at least seven product names and a list of their prices.
From the prices, print:
- The highest, the lowest, the total and the average (to two places)
- A new sorted list — and then prove the original is unchanged
- The three most expensive, using a slice
From the names, print:
- Alphabetical order, ignoring capitalisation
- Sorted by the length of the name
- All the names on one line, comma-separated, using
join
Then run one experiment: call .sort() on the prices, and print the lines from part 1 again. Which numbers changed and which did not? Explain to yourself why the ones that did not change could not have — hold on to that answer and this chapter's trap will never catch you.
Step 6 of 6
Stretch — the chapter quiz
Ten questions from easy to hard. The last ones are difficult on purpose.
Sign in to take the quiz