Python ships with a large collection of modules you never have to install. Knowing what is in there is one of the highest-value things a beginner can learn, because it turns “I need to write a function for this” into “there is already one”. This guide covers the fifteen modules that come up most often, with a short example of each.
Files and paths#
from pathlib import Path
p = Path("data") / "2026" / "report.csv"
p.parent.mkdir(parents=True, exist_ok=True)
print(p.name, p.stem, p.suffix)
print(p.exists(), p.is_file())
for csv_file in Path("data").rglob("*.csv"):
print(csv_file, csv_file.stat().st_size)
pathlib replaces almost all of os.path. Its companions are shutil for copying and moving, glob for pattern matching, and tempfile for scratch files that clean themselves up.
Data formats#
import json
data = {"name": "Ana", "scores": [1, 2, 3]}
text = json.dumps(data, indent=2)
back = json.loads(text)
with open("out.json", "w", encoding="utf-8") as f:
json.dump(data, f, indent=2)
import csv
with open("people.csv", newline="", encoding="utf-8") as f:
for row in csv.DictReader(f):
print(row["name"], row["age"])
with open("out.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["name", "age"])
writer.writeheader()
writer.writerows(rows)
The newline="" argument is not optional — without it you get blank lines between rows on Windows. Also here: sqlite3 for a real database in a single file, and configparser for INI-style settings.
Dates and times#
from datetime import datetime, date, timedelta, timezone
now = datetime.now(timezone.utc)
print(now.isoformat())
print(now.strftime("%d %B %Y"))
deadline = date.today() + timedelta(days=30)
print((deadline - date.today()).days)
parsed = datetime.strptime("2026-09-06", "%Y-%m-%d")
Better containers#
from collections import Counter, defaultdict, deque, namedtuple
words = "the cat sat on the mat the end".split()
print(Counter(words).most_common(2)) # [('the', 3), ('cat', 1)]
groups = defaultdict(list)
for word in words:
groups[len(word)].append(word)
queue = deque([1, 2, 3])
queue.appendleft(0)
queue.popleft()
Point = namedtuple("Point", "x y")
p = Point(3, 4)
print(p.x, p.y)
Counter alone replaces a surprising amount of hand-written counting code.
Iteration tools#
import itertools
print(list(itertools.chain([1, 2], [3, 4]))) # [1, 2, 3, 4]
print(list(itertools.combinations("abc", 2))) # pairs
print(list(itertools.product([1, 2], "ab"))) # every combination
print(list(itertools.islice(itertools.count(10), 3))) # [10, 11, 12]
data = sorted(people, key=lambda p: p["city"])
for city, group in itertools.groupby(data, key=lambda p: p["city"]):
print(city, len(list(group)))
Text and patterns#
import re
text = "Call 0161 496 0000 or 0207 946 0958"
print(re.findall(r"\d{4} \d{3} \d{4}", text))
print(re.sub(r"\s+", " ", " lots of space ").strip())
match = re.search(r"(\w+)@(\w+)\.com", "write to ana@example.com")
if match:
print(match.group(1), match.group(2))
Also worth knowing: textwrap for wrapping and indenting, string for character constants, difflib for “did you mean” suggestions.
Command-line programs#
import argparse
parser = argparse.ArgumentParser(description="Summarise a CSV file.")
parser.add_argument("path", help="the file to read")
parser.add_argument("--column", default="amount")
parser.add_argument("--verbose", action="store_true")
args = parser.parse_args()
print(args.path, args.column, args.verbose)
You get --help generated automatically, and clear errors for missing arguments. Reading sys.argv by hand is almost never worth it.
Logging instead of print#
import logging
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s %(levelname)s %(message)s",
)
log = logging.getLogger(__name__)
log.info("Started")
log.warning("Nothing found for %s", name)
log.error("Failed to open %s", path)
The advantage over print is that you can turn levels on and off, send output to a file, and add timestamps without touching any of the call sites.
Maths and randomness#
import math, random, statistics, decimal
print(math.sqrt(16), math.ceil(4.1), math.isclose(0.1 + 0.2, 0.3))
print(random.choice(["a", "b"]), random.sample(range(100), 3))
print(statistics.mean([1, 2, 3]), statistics.median([1, 2, 3, 4]))
print(decimal.Decimal("0.1") + decimal.Decimal("0.2")) # exactly 0.3
What these replace#
| Instead of installing | Use |
|---|---|
| A path helper library | pathlib |
| A CSV package | csv |
| A small key-value store | sqlite3 or shelve |
| An argument parser | argparse |
| A basic test runner | unittest |
| A simple web server for testing | http.server |
| A benchmarking helper | timeit |
| A memoisation decorator | functools.lru_cache |
Exploring a module you do not know#
import statistics
print(dir(statistics)) # every name in it
help(statistics.median) # the documentation for one
print(statistics.__file__) # where the source lives
Reading the actual source is more approachable than people expect — large parts of the standard library are plain, well-commented Python.
Questions people ask#
Do standard library modules need installing?
No. They come with Python. The exception is a handful of optional C extensions — lzma, sqlite3, ssl — which can be missing if Python was compiled without their system libraries.
Is the standard library slower than third-party packages?
Sometimes, for specialised work — orjson beats json, httpx beats urllib for ergonomics. For ordinary code the standard library is fast enough and has no installation cost.
Which modules should I learn first?
pathlib, json, csv, datetime and collections. Those five cover a large share of everyday scripting.
How do I know what version added a feature?
The documentation marks it, with notes like “New in version 3.9”. Worth checking if your code has to run on an older Python.
Where to go next#
- Python file handling — pathlib and open in practice.
- Python dictionaries explained — the base for collections.
- Building a complete Python project — where argparse and logging fit.