How to Extract Email Addresses from Text with Regex in Python and JavaScript
Extracting email addresses from raw text is a common task: cleaning a support inbox export, scanning a log file, or collecting contacts from a page. A single regex pattern does most of the work in both Python and JavaScript. This guide explains the pattern piece by piece, shows boundary-safe matching, and gives complete examples you can adapt.
The core email pattern
This pattern matches the vast majority of real-world addresses:
[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,} It has three parts:
- Local part
[A-Za-z0-9._%+-]+— the part before the @. It allows letters, digits, dots, underscores, percent, plus, and hyphen, which covers the characters the RFC permits in practice. - Domain label
[A-Za-z0-9.-]+— hostnames, including subdomains like mail.example.com. - Top-level domain
\.[A-Za-z]{2,}— the dot plus a TLD of at least two letters, so a match like [email protected] is rejected.
If you want to reject addresses ending in a dot or hyphen, tighten the pattern slightly:
[A-Za-z0-9._%+-]+@[A-Za-z0-9-]+(?:\.[A-Za-z0-9-]+)*\.[A-Za-z]{2,} Extract emails in Python
Python’s re module handles the whole task in one call:
import re
pattern = r"[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}"
text = "Contact [email protected] or [email protected] for help."
emails = re.findall(pattern, text)
print(emails)
# ['[email protected]', '[email protected]'] Use re.findall when you only need the addresses. Use re.finditer when you also need their positions, for example to highlight them in the original text:
for match in re.finditer(pattern, text, re.IGNORECASE):
print(match.group(), match.start(), match.end()) Extract emails in JavaScript
JavaScript uses the same pattern with a g flag to collect every match:
const pattern = /[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}/g;
const text = "Contact [email protected] or [email protected] for help.";
const emails = text.match(pattern) || [];
console.log(emails);
// ['[email protected]', '[email protected]'] Add the i flag (/.../gi) when the text may contain uppercase addresses and you want case-insensitive matching.
Avoid matching inside longer words
In many real texts an address sits next to punctuation or belongs inside a URL. To keep matches clean, add word boundaries on both sides:
\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b The boundary prevents a fragment like [email protected] from being treated as a valid address when it is actually part of a longer token.
Filter duplicates and normalize case
Email domains are case-insensitive, and the same address often appears more than once. Normalize before deduplicating:
# Python
seen = set()
unique = []
for email in re.findall(pattern, text, re.IGNORECASE):
key = email.lower()
if key not in seen:
seen.add(key)
unique.append(email)
print(unique) // JavaScript
const unique = [...new Set(
(text.match(/[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}/gi) || [])
.map(email => email.toLowerCase())
)];
console.log(unique); Handle noisy input
Log files and web pages add noise: addresses inside links (mailto:), obfuscated forms (alice [at] example.com), or false positives like user@localhost. Two practical steps:
- Strip known prefixes such as
mailto:before matching. - Filter results with a small blocklist (for example, addresses whose domain ends in
.local,.test, or.example) when your source is internal.
bad = {"localhost", ".local", ".test", ".example"}
clean = [e for e in unique if not any(e.endswith(s) for s in bad)] When regex is not enough
No single pattern is perfect for every address. The RFC 5322 grammar allows quoted local parts and comments, which this practical pattern deliberately ignores. If you are validating user input for a signup form, pair the regex with a verification email instead of trusting the pattern alone. For extraction, the pattern above balances precision and recall well.
Common questions
Why does my pattern also match inside URLs? Because an email inside a URL is still text that matches. Use word boundaries and, when needed, strip URL contexts before extraction.
Should I use \w instead of [A-Za-z0-9]? In Python, \w also matches Unicode letters, which can create false positives in some locales. The explicit character class keeps behavior predictable.
How do I extract from a file? Read the file as text, then apply the same pattern. For very large files, process line by line and accumulate matches.
Next steps
Start with the core pattern, add word boundaries, then deduplicate with a case-normalized set. If your source is HTML, parse it first or at least strip tags so addresses inside markup do not get split. Keep a small test set of real addresses and run it after every change to the pattern.