Skip to content
Happy Programming Guide
Start learning
Python

Cleaning Python Strings: Removing Punctuation

Four ways to strip punctuation from a string in Python, which is fastest, and how to handle apostrophes, hyphens and non-English text properly.

An open laptop on a wooden desk

The fastest way to remove punctuation from a string in Python is str.translate() with a table built from string.punctuation. Regular expressions are more flexible, and a comprehension is the most readable. Which one you want depends on whether you need to keep apostrophes in words like “don’t” — a decision that matters more than the technique.

Method 1: str.translate (fastest)#

Python
import string

text = "Hello, world! It's a fine day... isn't it?"
table = str.maketrans("", "", string.punctuation)

print(text.translate(table))
Output
Hello world Its a fine day isnt it

str.maketrans("", "", chars) builds a table that deletes every character in the third argument. The translation itself happens in C, which is why this is several times faster than the alternatives on long text.

Build the table once if you are cleaning many strings:

Python
PUNCT_TABLE = str.maketrans("", "", string.punctuation)

cleaned = [line.translate(PUNCT_TABLE) for line in lines]

Method 2: a regular expression (most flexible)#

Python
import re

text = "Hello, world! It's a fine day."

print(re.sub(r"[^\w\s]", "", text))        # remove anything not word or space
print(re.sub(r"[^\w\s']", "", text))       # keep apostrophes
print(re.sub(r"[^a-zA-Z\s]", "", text))     # letters and spaces only

The escape for a word character means letters, digits and underscore; the whitespace escape means spaces, tabs and newlines. The caret inside the brackets negates the set, so the pattern matches everything else.

Compile the pattern when reusing it:

Python
PUNCT_RE = re.compile(r"[^\w\s]")
cleaned = [PUNCT_RE.sub("", line) for line in lines]

Method 3: a comprehension (most readable)#

Python
import string

text = "Hello, world!"
cleaned = "".join(ch for ch in text if ch not in string.punctuation)
print(cleaned)      # Hello world

Slower, because it runs a Python-level loop over every character, but obvious to anyone reading it. Perfectly fine for short strings.

Method 4: keep only what you want#

Python
text = "Hello, world! 123"

print("".join(ch for ch in text if ch.isalnum() or ch.isspace()))
# Hello world 123

print("".join(ch for ch in text if ch.isalpha() or ch.isspace()))
# Hello world

An allow-list is safer than a deny-list when the input might contain characters you did not anticipate — emoji, currency symbols, mathematical notation.

What string.punctuation actually contains#

Thirty-two ASCII characters: the brackets, quotes, slashes, arithmetic signs and the rest of the symbols on a standard keyboard. It does not include curly quotes, em dashes, ellipsis characters or any non-English punctuation — all of which appear constantly in text copied from documents and web pages.

Python
import string

text = "She said \u201chello\u201d \u2014 then left\u2026"
print(text.translate(str.maketrans("", "", string.punctuation)))
# She said "hello" - then left...   (the fancy characters survive)

For real-world text, use Unicode categories instead:

Python
import unicodedata


def strip_punctuation(text):
    """Remove punctuation of any language, not just ASCII."""
    return "".join(
        ch for ch in text
        if not unicodedata.category(ch).startswith("P")
    )


print(strip_punctuation("She said \u201chello\u201d \u2014 then left\u2026"))
# She said hello  then left

Categories beginning with P are punctuation. Add S to that check if you also want to remove symbols such as currency signs and mathematical operators.

Apostrophes and hyphens#

Stripping all punctuation turns “don’t” into “dont” and “well-known” into “wellknown”. For word counting that is usually acceptable. For anything a person will read, it is not.

Python
import re

text = "It's a well-known state-of-the-art system, isn't it?"

# Remove an apostrophe or hyphen unless it has a letter on both sides
cleaned = re.sub(r"(?<![A-Za-z])['-]|['-](?![A-Za-z])", "", text)
cleaned = re.sub(r"[^\w\s'-]", "", cleaned)

print(cleaned)
# It's a well-known state-of-the-art system isn't it

The lookarounds are what keep the marks inside words while removing quotation marks and dashes.

Curly apostrophes need normalising first, or they slip through:

Python
text = text.replace("\u2019", "'").replace("\u2018", "'")

A reusable function#

Python
import re
import unicodedata

QUOTES = {"\u2018": "'", "\u2019": "'", "\u201c": '"', "\u201d": '"'}


def clean_text(text, keep_apostrophes=True, lowercase=True):
    """Normalise and strip punctuation from text."""
    for fancy, plain in QUOTES.items():
        text = text.replace(fancy, plain)

    text = unicodedata.normalize("NFKC", text)

    keep = "'" if keep_apostrophes else ""
    text = "".join(
        ch for ch in text
        if not unicodedata.category(ch).startswith("P") or ch in keep
    )

    text = re.sub(r"\s+", " ", text).strip()
    return text.lower() if lowercase else text


print(clean_text("  She said \u201cIt's fine\u201d \u2014 really!  "))
# she said it's fine really

The NFKC normalisation is worth keeping: it converts compatibility characters — ligatures, full-width Latin letters, some superscripts — into their ordinary equivalents, which prevents visually identical strings comparing as different.

Speed comparison#

Python
import timeit, string, re

text = "Hello, world! It's a fine day. " * 1000
table = str.maketrans("", "", string.punctuation)
pattern = re.compile(r"[^\w\s]")

print(timeit.timeit(lambda: text.translate(table), number=1000))
print(timeit.timeit(lambda: pattern.sub("", text), number=1000))
print(timeit.timeit(
    lambda: "".join(c for c in text if c not in string.punctuation), number=1000))

translate is typically several times faster than the compiled regex, and the comprehension is slower again by roughly an order of magnitude. On short strings none of this matters; on a large corpus it decides whether the job takes seconds or minutes.

Questions people ask#

How do I remove numbers as well?

Add a digit check with if not ch.isdigit() in the comprehension, or use a regex that keeps only letters and whitespace.

Should I lowercase before or after removing punctuation?

Either — the two do not interact. Do it once, in a single function, so the whole pipeline is consistent.

How do I remove extra spaces left behind?

re.sub(r"\s+", " ", text).strip() collapses any run of whitespace into one space and trims the ends.

Does this handle other alphabets?

The unicodedata version does — it identifies punctuation by category rather than by a fixed ASCII list, so Greek, Cyrillic and CJK text are handled correctly.

Where to go next#

Python string methods worth knowing by heartRead next

Keep reading

Keep going — pick your next guide

The fastest way to improve is to read one guide, then build the thing it describes. Start with the basics, or jump straight to a project.

Ask a question or share what worked

Your email address will not be published. Required fields are marked *