The fastest way to remove punctuation from a string in Python is str.translate() with a table built from string.punctuation. Regular expressions are more flexible, and a comprehension is the most readable. Which one you want depends on whether you need to keep apostrophes in words like “don’t” — a decision that matters more than the technique.
Method 1: str.translate (fastest)#
import string
text = "Hello, world! It's a fine day... isn't it?"
table = str.maketrans("", "", string.punctuation)
print(text.translate(table))
Hello world Its a fine day isnt it
str.maketrans("", "", chars) builds a table that deletes every character in the third argument. The translation itself happens in C, which is why this is several times faster than the alternatives on long text.
Build the table once if you are cleaning many strings:
PUNCT_TABLE = str.maketrans("", "", string.punctuation)
cleaned = [line.translate(PUNCT_TABLE) for line in lines]
Method 2: a regular expression (most flexible)#
import re
text = "Hello, world! It's a fine day."
print(re.sub(r"[^\w\s]", "", text)) # remove anything not word or space
print(re.sub(r"[^\w\s']", "", text)) # keep apostrophes
print(re.sub(r"[^a-zA-Z\s]", "", text)) # letters and spaces only
The escape for a word character means letters, digits and underscore; the whitespace escape means spaces, tabs and newlines. The caret inside the brackets negates the set, so the pattern matches everything else.
Compile the pattern when reusing it:
PUNCT_RE = re.compile(r"[^\w\s]")
cleaned = [PUNCT_RE.sub("", line) for line in lines]
Method 3: a comprehension (most readable)#
import string
text = "Hello, world!"
cleaned = "".join(ch for ch in text if ch not in string.punctuation)
print(cleaned) # Hello world
Slower, because it runs a Python-level loop over every character, but obvious to anyone reading it. Perfectly fine for short strings.
Method 4: keep only what you want#
text = "Hello, world! 123"
print("".join(ch for ch in text if ch.isalnum() or ch.isspace()))
# Hello world 123
print("".join(ch for ch in text if ch.isalpha() or ch.isspace()))
# Hello world
An allow-list is safer than a deny-list when the input might contain characters you did not anticipate — emoji, currency symbols, mathematical notation.
What string.punctuation actually contains#
Thirty-two ASCII characters: the brackets, quotes, slashes, arithmetic signs and the rest of the symbols on a standard keyboard. It does not include curly quotes, em dashes, ellipsis characters or any non-English punctuation — all of which appear constantly in text copied from documents and web pages.
import string
text = "She said \u201chello\u201d \u2014 then left\u2026"
print(text.translate(str.maketrans("", "", string.punctuation)))
# She said "hello" - then left... (the fancy characters survive)
For real-world text, use Unicode categories instead:
import unicodedata
def strip_punctuation(text):
"""Remove punctuation of any language, not just ASCII."""
return "".join(
ch for ch in text
if not unicodedata.category(ch).startswith("P")
)
print(strip_punctuation("She said \u201chello\u201d \u2014 then left\u2026"))
# She said hello then left
Categories beginning with P are punctuation. Add S to that check if you also want to remove symbols such as currency signs and mathematical operators.
Apostrophes and hyphens#
Stripping all punctuation turns “don’t” into “dont” and “well-known” into “wellknown”. For word counting that is usually acceptable. For anything a person will read, it is not.
import re
text = "It's a well-known state-of-the-art system, isn't it?"
# Remove an apostrophe or hyphen unless it has a letter on both sides
cleaned = re.sub(r"(?<![A-Za-z])['-]|['-](?![A-Za-z])", "", text)
cleaned = re.sub(r"[^\w\s'-]", "", cleaned)
print(cleaned)
# It's a well-known state-of-the-art system isn't it
The lookarounds are what keep the marks inside words while removing quotation marks and dashes.
Curly apostrophes need normalising first, or they slip through:
text = text.replace("\u2019", "'").replace("\u2018", "'")
A reusable function#
import re
import unicodedata
QUOTES = {"\u2018": "'", "\u2019": "'", "\u201c": '"', "\u201d": '"'}
def clean_text(text, keep_apostrophes=True, lowercase=True):
"""Normalise and strip punctuation from text."""
for fancy, plain in QUOTES.items():
text = text.replace(fancy, plain)
text = unicodedata.normalize("NFKC", text)
keep = "'" if keep_apostrophes else ""
text = "".join(
ch for ch in text
if not unicodedata.category(ch).startswith("P") or ch in keep
)
text = re.sub(r"\s+", " ", text).strip()
return text.lower() if lowercase else text
print(clean_text(" She said \u201cIt's fine\u201d \u2014 really! "))
# she said it's fine really
The NFKC normalisation is worth keeping: it converts compatibility characters — ligatures, full-width Latin letters, some superscripts — into their ordinary equivalents, which prevents visually identical strings comparing as different.
Speed comparison#
import timeit, string, re
text = "Hello, world! It's a fine day. " * 1000
table = str.maketrans("", "", string.punctuation)
pattern = re.compile(r"[^\w\s]")
print(timeit.timeit(lambda: text.translate(table), number=1000))
print(timeit.timeit(lambda: pattern.sub("", text), number=1000))
print(timeit.timeit(
lambda: "".join(c for c in text if c not in string.punctuation), number=1000))
translate is typically several times faster than the compiled regex, and the comprehension is slower again by roughly an order of magnitude. On short strings none of this matters; on a large corpus it decides whether the job takes seconds or minutes.
Questions people ask#
How do I remove numbers as well?
Add a digit check with if not ch.isdigit() in the comprehension, or use a regex that keeps only letters and whitespace.
Should I lowercase before or after removing punctuation?
Either — the two do not interact. Do it once, in a single function, so the whole pipeline is consistent.
How do I remove extra spaces left behind?
re.sub(r"\s+", " ", text).strip() collapses any run of whitespace into one space and trims the ends.
Does this handle other alphabets?
The unicodedata version does — it identifies punctuation by category rather than by a fixed ASCII list, so Greek, Cyrillic and CJK text are handled correctly.
Where to go next#
- Python string methods — the wider toolkit these build on.
- Python lists explained — for what you do after splitting into words.
- Web scraping and automation — where messy text usually comes from.