Chapter 23

Reading and writing files

Reading and writing with `with open`, what separates r/w/a/x, the newline at the end of every line, why `encoding` is always written, and pathlib for the small jobs.

35 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

Everything we have built across twenty-three chapters disappeared the moment the program ended.

python
basket = []
basket.append("pen")
print(basket)
text
['pen']

Run it again and the basket is empty again. Lists, dictionaries, functions — all of it lives in memory, and memory lasts exactly as long as the program.

The last chapter was about sharing code between programs. This one is about data — how one run leaves its work for the next run, or for somebody else.

python
with open("notes.txt", "w", encoding="utf-8") as fh:
    fh.write("first line\n")
    fh.write("second line\n")

with open("notes.txt", "r", encoding="utf-8") as fh:
    print(fh.read(), end="")
text
first line
second line

The program has ended, and notes.txt is still there.

By the end of this chapter you can

  • Read and write files with with open(...), and say why with
  • Say what separates "r", "w", "a" and "x" — especially what "w" destroys
  • Read a file one line at a time, and handle the \n at the end of each
  • Say why encoding="utf-8" is always written
  • Use pathlib.Path for the small jobs
  • Read a FileNotFoundError and a UnicodeDecodeError

Prerequisites: Modules and imports.


Why with

Opening a file borrows something from the operating system, and a loan has to be repaid. Without closing, what you wrote may never reach the disk, and enough unclosed files will stop the program.

A with block makes that repayment itself — the file is closed as the block ends, even if something inside it raised.

python
with open("notes.txt", "r", encoding="utf-8") as fh:
    text = fh.read()

print(fh.closed)
print(fh.read())
text
True
ValueError: I/O operation on closed file.

The name fh still exists after the block — scope works exactly as before — but the file is shut. In this course open() is always written with with.

Reading

The whole contents at once come from .read():

python
with open("notes.txt", "r", encoding="utf-8") as fh:
    text = fh.read()

print(repr(text))
print(len(text))
text
'first line\nsecond line\n'
23

repr is used deliberately, because a plain print hides the line breaks. A file is one long piece of text, with its line endings sitting inside it as \n — remembering that is half of this chapter.

One line at a time

.read() is a poor choice for a large file, because it lifts the whole thing into memory. Looping over the file itself makes Python read a line at a time:

python
with open("orders.txt", "r", encoding="utf-8") as fh:
    for line in fh:
        print(repr(line))
text
'pen,15.0,3\n'
'bag,850.0,1\n'
'ink,120.0,2\n'

Every line ends with \n. This is the most common confusion of all — int(line) does not work, line == "pen" does not match, and the reason is invisible.

So the first move is almost always .strip():

python
with open("orders.txt", "r", encoding="utf-8") as fh:
    for line in fh:
        name, price, qty = line.strip().split(",")
        print(f"{name:<5} {float(price) * int(qty):>8.2f}")
text
pen      45.00
bag     850.00
ink     240.00

.strip() then .split(",") then float() and int() — that sequence will come back again and again when reading files. Everything that comes out of a file is text, never numbers; chapter seven's conversion is needed here too.

When all the lines are wanted at once, there are two ways:

python
with open("orders.txt", "r", encoding="utf-8") as fh:
    lines = fh.read().splitlines()

print(lines)
text
['pen,15.0,3', 'bag,850.0,1', 'ink,120.0,2']

.splitlines() drops the line endings. There is also .readlines(), but it keeps the \n — which is where most people trip.

Writing, and one dangerous letter

python
with open("notes.txt", "w", encoding="utf-8") as fh:
    fh.write("new\n")

The file previously contained old content. Now:

text
new

"w" empties the file — at the moment of opening, before anything is written. There is no warning and no way back.

To keep what is there and add to the end, "a":

python
with open("notes.txt", "a", encoding="utf-8") as fh:
    fh.write("new\n")
text
old content
new

And to refuse to write at all if the file already exists, "x":

text
FileExistsError: [Errno 17] File exists: 'notes.txt'

The four modes in one line: "r" reads, "w" erases and writes, "a" adds to the end, "x" only creates.

Pause before typing "w". People lose real work to this one letter. If the file you are opening was not made by your program, think before writing "w".

The \n is yours to supply

print breaks the line for you. write does not.

python
rows = ["pen", "bag"]

with open("out.txt", "w", encoding="utf-8") as fh:
    for row in rows:
        fh.write(row)
text
'penbag'

So the cleanest way to write a list of lines is:

python
rows = ["pen", "bag"]

with open("out.txt", "w", encoding="utf-8") as fh:
    fh.write("\n".join(rows) + "\n")
text
'pen\nbag\n'

"\n".join(rows) puts breaks between the lines, and the trailing + "\n" ends the file on a newline — which is the convention for text files, and what many tools expect.

encoding="utf-8" — always

Without encoding, Python uses the operating system's default, which differs from machine to machine. The result is a program that works where you are and breaks where somebody else is.

Reading with the wrong one gives:

text
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 3: invalid start byte

So write encoding="utf-8" every time you write open() — for reading and for writing alike. The only exception is binary mode ("rb", "wb"), where the question does not arise because there is no encoding.

pathlib — a short path for small jobs

For a small file the whole with block can feel like a lot.

python
from pathlib import Path

path = Path("orders.txt")
print(path.exists())
print(len(path.read_text(encoding="utf-8").splitlines()))

out = Path("report.txt")
out.write_text("done\n", encoding="utf-8")
print(out.read_text(encoding="utf-8"), end="")
text
True
3
done

read_text and write_text open and close the file internally, so no with is needed. Path does more besides — .exists(), .name, .suffix, .parent — and folder paths are joined with /, as in Path("data") / "orders.txt".

Which when? Path when the whole file is wanted at once, with open(...) when reading line by line. write_text has one more advantage: it replaces the file by default, so the mistake of typing "w" is at least never accidental.


A complete example

orders.txt:

text
pen,15.0,3
bag,850.0,1
ink,120.0,2

main.py:

python
"""Read an order file, total it, and write a report beside it."""

from pathlib import Path

TAX_RATE = 0.15


def read_order(path):
    """Each line is name,price,quantity. Blank lines are skipped."""
    lines = []
    for raw in path.read_text(encoding="utf-8").splitlines():
        if not raw.strip():
            continue
        name, price, quantity = raw.split(",")
        lines.append((name, float(price), int(quantity)))
    return lines


def report(lines):
    rows = []
    total = 0.0
    for name, price, quantity in lines:
        amount = round(price * quantity * (1 + TAX_RATE), 2)
        total += amount
        rows.append(f"{name:<6} {quantity:>3} {amount:>9.2f}")
    rows.append("-" * 20)
    rows.append(f"{'total':<6} {'':>3} {total:>9.2f}")
    return rows


def main():
    order = read_order(Path("orders.txt"))
    rows = report(order)

    Path("report.txt").write_text("\n".join(rows) + "\n", encoding="utf-8")

    print(Path("report.txt").read_text(encoding="utf-8"), end="")
    print()
    print("lines read:", len(order))


if __name__ == "__main__":
    main()
text
pen      3     51.75
bag      1    977.50
ink      2    276.00
--------------------
total        1305.25

lines read: 3

Four things worth looking at.

Three functions with three separate jobs. read_order reads from disk, report only calculates, and main writes. report is given no file at all — it is given a list and returns a list. So testing it needs no file, and in chapter twenty-eight we will be grateful for exactly that.

The blank line is handled in advance. A stray newline at the end of a file is entirely normal, and without handling it raw.split(",") would not get three things and would raise a ValueError. if not raw.strip(): continue settles it in one line.

float(price) and int(quantity) are not optional, because everything out of a file arrives as text. Forget them and price * quantity would quietly repeat the text instead — the worst kind of bug, the kind that does not break.

The report goes to a separate file. orders.txt is never opened with "w" — the input file stays intact. That is a habit worth forming: do not write to what you are reading.


When it breaks

FileNotFoundError: [Errno 2] No such file or directory: 'orders.txt' The file is not there, or you are running the program from a different folder. The path is worked out from where you ran it, not from the script's folder. Print Path("orders.txt").resolve() to see where it is looking.

UnicodeDecodeError: 'utf-8' codec can't decode byte ... The file was written in another encoding, or it is not a text file at all. Fix the encoding, or open it with "rb" if it is binary.

ValueError: I/O operation on closed file. The file was used outside its with block. Read what you need inside the block.

ValueError: not enough values to unpack Some line does not have the commas you expected — often the empty line at the end. Skip blank lines, and when in doubt print(repr(raw)).

I wrote the file but it is empty It is being read before the with block has finished. Writes reach the disk when the file is closed.

All my lines ran together into one write does not break lines. The "\n" is yours to supply.

Everything in my file is gone It was opened with "w". Use "r" to read and "a" to add to the end.