To get a file’s size in Python: Path("report.pdf").stat().st_size. That returns bytes as an integer. Everything else — formatting it as MB, summing a folder, handling missing files — builds on that one call. This guide covers all of it, including the difference between KB and KiB that trips people up when their number does not match what the operating system shows.
Getting the size#
from pathlib import Path
path = Path("report.pdf")
size = path.stat().st_size
print(size, "bytes")
The os equivalent, for older code:
import os
size = os.path.getsize("report.pdf")
size = os.stat("report.pdf").st_size # same thing
All of these raise FileNotFoundError when the file is missing, so check first if that is possible:
if path.is_file():
print(path.stat().st_size)
else:
print("No such file")
is_file() is better than exists() here — a directory exists but asking for its size gives you the size of the directory entry, not its contents, which is almost never what anyone means.
Formatting bytes for people#
def human_size(num_bytes, binary=False):
"""Format a byte count. binary=True gives KiB, MiB (base 1024)."""
step = 1024 if binary else 1000
units = ["KiB", "MiB", "GiB", "TiB"] if binary else ["kB", "MB", "GB", "TB"]
if num_bytes < step:
return f"{num_bytes} B"
value = float(num_bytes)
for unit in units:
value /= step
if value < step:
return f"{value:.1f} {unit}"
return f"{value:.1f} {units[-1]}"
print(human_size(1000)) # 1.0 kB
print(human_size(1024, True)) # 1.0 KiB
print(human_size(5_242_880)) # 5.2 MB
print(human_size(5_242_880, True)) # 5.0 MiB
Why your number does not match the file manager#
Two separate reasons, and they are often confused:
1. Base 1000 versus base 1024. A “kilobyte” officially means 1000 bytes, but many tools use 1024 and still call it KB. Windows shows 1024-based numbers labelled KB and MB; macOS uses 1000-based. So a 5,242,880-byte file is “5.24 MB” on macOS and “5.00 MB” on Windows, and both are telling the truth by their own definition.
2. Size on disk versus size. Files occupy whole allocation blocks, typically 4 KiB. A one-byte file uses 4096 bytes on disk. st_size reports the logical size; the file manager may show either.
stat = Path("tiny.txt").stat()
print(stat.st_size) # 1 - logical size
print(stat.st_blocks * 512) # 4096 - actual disk use (Unix only)
Totalling a folder#
from pathlib import Path
def folder_size(folder):
"""Total size of every file below folder, in bytes."""
total = 0
for path in Path(folder).rglob("*"):
if path.is_file() and not path.is_symlink():
try:
total += path.stat().st_size
except (PermissionError, FileNotFoundError):
continue # skip what we cannot read
return total
print(human_size(folder_size("documents")))
Three details that matter in real folders. Skipping symlinks stops the same file being counted twice, or an infinite loop if a link points at a parent. The try/except handles files that disappear or are locked mid-scan. And rglob("*") includes directories, hence the is_file() check.
For large trees, os.scandir is noticeably faster because it reads size information from the directory entry rather than calling stat separately:
import os
def folder_size_fast(folder):
total = 0
with os.scandir(folder) as entries:
for entry in entries:
try:
if entry.is_file(follow_symlinks=False):
total += entry.stat(follow_symlinks=False).st_size
elif entry.is_dir(follow_symlinks=False):
total += folder_size_fast(entry.path)
except (PermissionError, FileNotFoundError):
continue
return total
Finding the biggest files#
from pathlib import Path
files = [
(p.stat().st_size, p)
for p in Path("documents").rglob("*")
if p.is_file()
]
for size, path in sorted(files, reverse=True)[:10]:
print(f"{human_size(size):>10} {path}")
That is a usable disk-space report in six lines, and it is the sort of small script worth keeping around.
Checking a size without reading the file#
MAX_UPLOAD = 5 * 1024 * 1024 # 5 MiB
path = Path("upload.jpg")
if path.stat().st_size > MAX_UPLOAD:
raise ValueError(f"File is {human_size(path.stat().st_size)}, limit is 5 MiB")
stat() reads metadata only, so this costs nothing regardless of how large the file is. Checking the size before read() is what stops a huge file exhausting memory.
Other metadata from the same call#
from datetime import datetime
stat = Path("report.pdf").stat()
print(stat.st_size) # bytes
print(datetime.fromtimestamp(stat.st_mtime)) # last modified
print(datetime.fromtimestamp(stat.st_atime)) # last accessed
print(oct(stat.st_mode)[-3:]) # permissions, Unix
Call stat() once and reuse the result. Each call is a filesystem operation, and in a loop over thousands of files the difference is measurable.
Questions people ask#
Does st_size include the file name and metadata?
No. It is the size of the contents only. Names and metadata live in the directory entry and the inode.
Why is an empty folder not zero bytes?
A directory is itself a file on disk holding the list of its entries, so it has a small size of its own. Sum the files inside it rather than asking for the folder’s own size.
How do I get the size of a file on a remote server?
Depends on the protocol. Over HTTP, a HEAD request returns a Content-Length header. Over SFTP, paramiko exposes the same st_size attribute you already know.
Is os.path.getsize faster than Path.stat?
They call the same system function, so any difference is noise. Use whichever fits the rest of your code — pathlib for new work.
Where to go next#
- Getting the file name from a path — the rest of the pathlib basics.
- Moving files with shutil — acting on the files you have measured.
- Python file handling — reading and writing.