Skip to content
Happy Programming Guide
Start learning
Python

Python File Size: A Comprehensive Guide

Get a file size in Python with pathlib or os.path, format bytes as KB and MB properly, and total up a whole folder including subfolders.

A backlit keyboard in a dark room

To get a file’s size in Python: Path("report.pdf").stat().st_size. That returns bytes as an integer. Everything else — formatting it as MB, summing a folder, handling missing files — builds on that one call. This guide covers all of it, including the difference between KB and KiB that trips people up when their number does not match what the operating system shows.

Getting the size#

Python
from pathlib import Path

path = Path("report.pdf")
size = path.stat().st_size

print(size, "bytes")

The os equivalent, for older code:

Python
import os

size = os.path.getsize("report.pdf")
size = os.stat("report.pdf").st_size      # same thing

All of these raise FileNotFoundError when the file is missing, so check first if that is possible:

Python
if path.is_file():
    print(path.stat().st_size)
else:
    print("No such file")

is_file() is better than exists() here — a directory exists but asking for its size gives you the size of the directory entry, not its contents, which is almost never what anyone means.

Formatting bytes for people#

Python
def human_size(num_bytes, binary=False):
    """Format a byte count. binary=True gives KiB, MiB (base 1024)."""
    step = 1024 if binary else 1000
    units = ["KiB", "MiB", "GiB", "TiB"] if binary else ["kB", "MB", "GB", "TB"]

    if num_bytes < step:
        return f"{num_bytes} B"

    value = float(num_bytes)
    for unit in units:
        value /= step
        if value < step:
            return f"{value:.1f} {unit}"
    return f"{value:.1f} {units[-1]}"


print(human_size(1000))          # 1.0 kB
print(human_size(1024, True))    # 1.0 KiB
print(human_size(5_242_880))     # 5.2 MB
print(human_size(5_242_880, True))   # 5.0 MiB

Why your number does not match the file manager#

Two separate reasons, and they are often confused:

1. Base 1000 versus base 1024. A “kilobyte” officially means 1000 bytes, but many tools use 1024 and still call it KB. Windows shows 1024-based numbers labelled KB and MB; macOS uses 1000-based. So a 5,242,880-byte file is “5.24 MB” on macOS and “5.00 MB” on Windows, and both are telling the truth by their own definition.

2. Size on disk versus size. Files occupy whole allocation blocks, typically 4 KiB. A one-byte file uses 4096 bytes on disk. st_size reports the logical size; the file manager may show either.

Python
stat = Path("tiny.txt").stat()

print(stat.st_size)                  # 1 - logical size
print(stat.st_blocks * 512)          # 4096 - actual disk use (Unix only)

Totalling a folder#

Python
from pathlib import Path


def folder_size(folder):
    """Total size of every file below folder, in bytes."""
    total = 0
    for path in Path(folder).rglob("*"):
        if path.is_file() and not path.is_symlink():
            try:
                total += path.stat().st_size
            except (PermissionError, FileNotFoundError):
                continue          # skip what we cannot read
    return total


print(human_size(folder_size("documents")))

Three details that matter in real folders. Skipping symlinks stops the same file being counted twice, or an infinite loop if a link points at a parent. The try/except handles files that disappear or are locked mid-scan. And rglob("*") includes directories, hence the is_file() check.

For large trees, os.scandir is noticeably faster because it reads size information from the directory entry rather than calling stat separately:

Python
import os


def folder_size_fast(folder):
    total = 0
    with os.scandir(folder) as entries:
        for entry in entries:
            try:
                if entry.is_file(follow_symlinks=False):
                    total += entry.stat(follow_symlinks=False).st_size
                elif entry.is_dir(follow_symlinks=False):
                    total += folder_size_fast(entry.path)
            except (PermissionError, FileNotFoundError):
                continue
    return total

Finding the biggest files#

Python
from pathlib import Path

files = [
    (p.stat().st_size, p)
    for p in Path("documents").rglob("*")
    if p.is_file()
]

for size, path in sorted(files, reverse=True)[:10]:
    print(f"{human_size(size):>10}  {path}")

That is a usable disk-space report in six lines, and it is the sort of small script worth keeping around.

Checking a size without reading the file#

Python
MAX_UPLOAD = 5 * 1024 * 1024      # 5 MiB

path = Path("upload.jpg")

if path.stat().st_size > MAX_UPLOAD:
    raise ValueError(f"File is {human_size(path.stat().st_size)}, limit is 5 MiB")

stat() reads metadata only, so this costs nothing regardless of how large the file is. Checking the size before read() is what stops a huge file exhausting memory.

Other metadata from the same call#

Python
from datetime import datetime

stat = Path("report.pdf").stat()

print(stat.st_size)                                      # bytes
print(datetime.fromtimestamp(stat.st_mtime))             # last modified
print(datetime.fromtimestamp(stat.st_atime))             # last accessed
print(oct(stat.st_mode)[-3:])                            # permissions, Unix

Call stat() once and reuse the result. Each call is a filesystem operation, and in a loop over thousands of files the difference is measurable.

Questions people ask#

Does st_size include the file name and metadata?

No. It is the size of the contents only. Names and metadata live in the directory entry and the inode.

Why is an empty folder not zero bytes?

A directory is itself a file on disk holding the list of its entries, so it has a small size of its own. Sum the files inside it rather than asking for the folder’s own size.

How do I get the size of a file on a remote server?

Depends on the protocol. Over HTTP, a HEAD request returns a Content-Length header. Over SFTP, paramiko exposes the same st_size attribute you already know.

Is os.path.getsize faster than Path.stat?

They call the same system function, so any difference is noise. Use whichever fits the rest of your code — pathlib for new work.

Where to go next#

Moving files safely in Python with shutilRead next

Keep reading

Python

Python File Moving Using shutil

How shutil.move works, when it differs from os.rename, and how to move files safely across drives without overwriting anything by accident.

4 min read

Keep going — pick your next guide

The fastest way to improve is to read one guide, then build the thing it describes. Start with the basics, or jump straight to a project.

Ask a question or share what worked

Your email address will not be published. Required fields are marked *