Chapter 22

Modules and imports

Using a file you wrote from another file, what an import actually does, why `if __name__ == "__main__"` is needed, the standard library, and hiding a module behind a file of your own.

33 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

In the last chapter we wrote a line_total function. Now a second program needs the same thing.

The easy way is to copy it across. That works, on the first day.

Then the tax rate changes. You fix one file and forget the other. Now two programs show two different totals for the same order, and there is no way to say which is right — because both are right, according to their own file.

The problem is not the effort of copying. The problem is that the truth now lives in two places.

python
import pricing

print(pricing.TAX_RATE)
print(pricing.line_total(15.0, 3))
text
0.15
51.75

One file, and as many programs as you like using it. When the rate changes, it changes in one place.

This chapter is that one line — import — and what actually happens behind it.

By the end of this chapter you can

  • Import a file you wrote from another file
  • Say what separates import x, from x import y and import x as z
  • Say what happens at the moment of an import, and how many times it happens
  • Explain why if __name__ == "__main__": is written
  • Bring in modules from the standard library
  • Know where Python looks, and avoid hiding a module behind a file of your own

Prerequisites: Scope and mutable defaults.


A module is just a file

A file called pricing.py is a module called pricing. Nothing has to be declared and no special line has to be written — having the file is enough.

import pricing brings in the whole module, and things inside it are reached through a dot: pricing.TAX_RATE, pricing.line_total(...).

The dot can feel like extra typing, but it does a job — when reading, it says where the thing came from. In a large file that is worth a great deal.

Bringing in only what you need

python
from pricing import line_total

print(line_total(15.0, 3))
print(TAX_RATE)
text
51.75
NameError: name 'TAX_RATE' is not defined

from ... import ... brings in only the name that was asked for, and that name is then used directly — no dot required.

But the other names do not come. TAX_RATE was not asked for, so it is not here, even though it sits in the same file.

Which to use? Roughly: from when you need two or three names, the whole module when you need many. And one thing never to do — from pricing import *. It pulls in every name silently, and afterwards there is no way to find out where line_total came from.

Renaming on the way in

python
import pricing as p

print(p.line_total(15.0, 3))
text
51.75

as gives the module a shorter name inside this file.

Write it when the name really is long, or when the convention is established. A one-letter name out of laziness throws away the readability the dot was providing — nobody reading p.line_total(...) can tell where it comes from.

Importing means running the file

This is the part most people discover later, usually at a bad moment.

tools.py:

python
print("tools is being set up")


def ping():
    return "pong"

main.py:

python
print("main starts")
import tools
import tools

print(tools.ping())
print("main ends")
text
main starts
tools is being set up
pong
main ends

Two things to notice.

Importing means running the file, top to bottom. The def lines create functions, and everything else — such as that print — actually happens. Put a print or an input at the top level of a module and it lands on whoever imports it.

The second import tools did nothing. A module runs once per program; after that Python hands back the one it already has. So importing the same module in ten places is safe and cheap.

__name__ and that famous line

Every module has a name called __name__ inside it, which Python sets itself.

tools.py:

python
def ping():
    return "pong"


print("module name is", __name__)

if __name__ == "__main__":
    print("running directly")

Run as python tools.py:

text
module name is __main__
running directly

But import tools from another file:

text
module name is tools
imported, name was tools

The rule is simple: the file you are running has __name__ of "__main__"; files being imported have their own name.

So whatever is inside if __name__ == "__main__": runs only when the file is run directly, never when it is imported.

Why that is needed follows from the previous section. Importing means running the file — so if main() sits at the end of a file, anyone who imports it suddenly runs your program too. That line closes the door.

The standard library

A great many modules ship with Python itself. Nothing has to be installed.

python
import math
import statistics

print(math.sqrt(144))
print(math.floor(3.7), math.ceil(3.2))
print(statistics.mean([10, 20, 60]))
text
12.0
3 4
30

Another one that comes up often:

python
import random

random.seed(7)
print(random.randint(1, 100))
print(random.choice(["pen", "bag", "ink"]))
text
42
pen

Without random.seed(7) every run would give different numbers — which is, after all, the point. Setting a seed fixes the sequence, so the output above will be the same on your machine. For writing tests this is indispensable: to test something random, you first have to make it repeatable.

This is not "random", it is "random-looking". random produces numbers from a calculation, so the same seed always gives the same sequence. Do not use it for passwords or tokens — the secrets module exists for that.

Where Python looks

For import pricing, Python searches in order: first the modules built into the interpreter, then the folder of the script being run, then the standard library, then installed packages.

That order has a consequence, and it catches people out.

There is a file called random.py in your folder, and you write:

python
import random

print(random.choice([1, 2, 3]))
text
AttributeError: module 'random' has no attribute 'choice'

Your file came before the real random, so yours is what got imported — and it has no choice in it.

The message is confusing, because it says random has no choice, which sounds impossible. But it is talking about your random.

The rule: never name a file after a module. random.py, string.py, json.py, csv.py — they look like innocent names, which is exactly why the trap works.

One exception, which explains the rule. Make a file called math.py and import math still gives you the real one, not yours, because math is built into the interpreter and that comes first in the list. So the trap fires for some names and not others — which is harder to remember than simply renaming the file.

Arranging files in folders

As files multiply they can be grouped into folders. Put an __init__.py file inside the folder — empty is fine — and the folder becomes a package.

text
shop/
    __init__.py
    pricing.py
main.py

The path can then be written with dots:

python
from shop.pricing import line_total

print(line_total(15.0, 3))
text
51.75

That is enough for now. For the rest of this chapter we stay with flat files.


A complete example

Three files, with the dependencies running in one direction only.

pricing.py:

python
"""Prices and tax. Nothing here prints, and nothing here asks for input."""

TAX_RATE = 0.15


def line_total(price, quantity):
    """One line of an order, tax included."""
    return round(price * quantity * (1 + TAX_RATE), 2)


def order_total(lines):
    """`lines` is a list of (name, price, quantity) tuples."""
    return round(sum(line_total(price, qty) for _, price, qty in lines), 2)

report.py:

python
"""Turning numbers into lines of text. It imports pricing; pricing imports nothing."""

import pricing


def receipt(lines):
    rows = [f"{name:<8} {pricing.line_total(price, qty):>8.2f}"
            for name, price, qty in lines]
    rows.append("-" * 17)
    rows.append(f"{'total':<8} {pricing.order_total(lines):>8.2f}")
    return rows


if __name__ == "__main__":
    print("report.py has no data of its own to show")

main.py:

python
"""The only file that is meant to be run."""

import statistics

from report import receipt

ORDER = [
    ("pen", 15.0, 3),
    ("bag", 850.0, 1),
    ("ink", 120.0, 2),
]


def main():
    for row in receipt(ORDER):
        print(row)

    print()
    quantities = [qty for _, _, qty in ORDER]
    print("mean quantity:", statistics.mean(quantities))


if __name__ == "__main__":
    main()

Run as python main.py:

text
pen         51.75
bag        977.50
ink        276.00
-----------------
total     1305.25

mean quantity: 2

Four things worth looking at.

The dependencies run one way. main knows report, report knows pricing, and pricing knows nobody. That ordering makes a chain, and the file at the end of it — pricing — can be tested entirely on its own. Had pricing needed to import report back, neither of them could stand alone any more.

pricing.py prints nothing and asks for nothing. It takes values and returns values — chapter twenty-one's pure functions. Since importing means running the file, that restraint is not politeness, it is a requirement.

report.py has an if __name__ == "__main__": too, and it does nothing useful — the file has no data of its own, so it has nothing to show. It could have been left out. It is there so that if the file is ever run directly it at least says something, rather than doing nothing in silence.

No file except main.py does anything by itself. The whole program has exactly one entry point, and reading it tells you what will happen.


When it breaks

ModuleNotFoundError: No module named 'priceing' A misspelling, or the file is in another folder. Check that it sits beside the script you are running.

ImportError: cannot import name 'line_totl' from 'pricing' The module was found, the name inside it was not. The message often suggests the near miss — Did you mean: 'line_total'?

AttributeError: module 'random' has no attribute 'choice' Almost always a file of yours hiding the real module. Look for that name in your folder and rename it. Delete the __pycache__ folder as well.

ImportError: cannot import name 'A' from partially initialized module 'a' (most likely due to a circular import) Two files import each other. The usual fix is to move whatever both of them need into a third file that imports neither.

Something printed, or asked for input, the moment I imported it There is code at the module's top level. Put it in a function and call it from inside if __name__ == "__main__":.

I edited the file but I am still seeing the old behaviour Run the program again — a running program does not re-read the module. In a notebook, restart the kernel.