Home >Blog

How to Extract Email Addresses from Text with Regex in Python and JavaScript

Published Updated

Extracting email addresses from raw text is a common task: cleaning a support inbox export, scanning a log file, or collecting contacts from a page. A single regex pattern does most of the work in both Python and JavaScript. This guide explains the pattern piece by piece, shows boundary-safe matching, and gives complete examples you can adapt.

The core email pattern

This pattern matches the vast majority of real-world addresses:

[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}

It has three parts:

  • Local part[A-Za-z0-9._%+-]+ — the part before the @. It allows letters, digits, dots, underscores, percent, plus, and hyphen, which covers the characters the RFC permits in practice.
  • Domain label[A-Za-z0-9.-]+ — hostnames, including subdomains like mail.example.com.
  • Top-level domain\.[A-Za-z]{2,} — the dot plus a TLD of at least two letters, so a match like [email protected] is rejected.

If you want to reject addresses ending in a dot or hyphen, tighten the pattern slightly:

[A-Za-z0-9._%+-]+@[A-Za-z0-9-]+(?:\.[A-Za-z0-9-]+)*\.[A-Za-z]{2,}

Extract emails in Python

Python’s re module handles the whole task in one call:

import re

pattern = r"[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}"
text = "Contact [email protected] or [email protected] for help."

emails = re.findall(pattern, text)
print(emails)
# ['[email protected]', '[email protected]']

Use re.findall when you only need the addresses. Use re.finditer when you also need their positions, for example to highlight them in the original text:

for match in re.finditer(pattern, text, re.IGNORECASE):
    print(match.group(), match.start(), match.end())

Extract emails in JavaScript

JavaScript uses the same pattern with a g flag to collect every match:

const pattern = /[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}/g;
const text = "Contact [email protected] or [email protected] for help.";

const emails = text.match(pattern) || [];
console.log(emails);
// ['[email protected]', '[email protected]']

Add the i flag (/.../gi) when the text may contain uppercase addresses and you want case-insensitive matching.

Avoid matching inside longer words

In many real texts an address sits next to punctuation or belongs inside a URL. To keep matches clean, add word boundaries on both sides:

\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b

The boundary prevents a fragment like [email protected] from being treated as a valid address when it is actually part of a longer token.

Filter duplicates and normalize case

Email domains are case-insensitive, and the same address often appears more than once. Normalize before deduplicating:

# Python
seen = set()
unique = []
for email in re.findall(pattern, text, re.IGNORECASE):
    key = email.lower()
    if key not in seen:
        seen.add(key)
        unique.append(email)
print(unique)
// JavaScript
const unique = [...new Set(
  (text.match(/[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}/gi) || [])
    .map(email => email.toLowerCase())
)];
console.log(unique);

Handle noisy input

Log files and web pages add noise: addresses inside links (mailto:), obfuscated forms (alice [at] example.com), or false positives like user@localhost. Two practical steps:

  • Strip known prefixes such as mailto: before matching.
  • Filter results with a small blocklist (for example, addresses whose domain ends in .local, .test, or .example) when your source is internal.
bad = {"localhost", ".local", ".test", ".example"}
clean = [e for e in unique if not any(e.endswith(s) for s in bad)]

When regex is not enough

No single pattern is perfect for every address. The RFC 5322 grammar allows quoted local parts and comments, which this practical pattern deliberately ignores. If you are validating user input for a signup form, pair the regex with a verification email instead of trusting the pattern alone. For extraction, the pattern above balances precision and recall well.

Common questions

Why does my pattern also match inside URLs? Because an email inside a URL is still text that matches. Use word boundaries and, when needed, strip URL contexts before extraction.

Should I use \w instead of [A-Za-z0-9]? In Python, \w also matches Unicode letters, which can create false positives in some locales. The explicit character class keeps behavior predictable.

How do I extract from a file? Read the file as text, then apply the same pattern. For very large files, process line by line and accumulate matches.

Next steps

Start with the core pattern, add word boundaries, then deduplicate with a case-normalized set. If your source is HTML, parse it first or at least strip tags so addresses inside markup do not get split. Keep a small test set of real addresses and run it after every change to the pattern.

FAQ

Why does my email regex also match inside URLs?
An email inside a URL is still text that matches the pattern. Use word boundaries and, when needed, strip URL contexts before extraction.
Should I use \w instead of [A-Za-z0-9] in the email pattern?
In Python, \w also matches Unicode letters, which can create false positives in some locales. The explicit character class keeps behavior predictable.
How do I extract emails from a large file?
Read the file as text and apply the same pattern, or process it line by line and accumulate matches to keep memory usage low.

Table of Contents

Information

  • Hits211
  • Published date2026/09/15
0/500
Share your thoughts respectfully.

More Posts

Explore more articles from this section
Regex for IP Addresses: Accurate IPv4 and IPv6 Validation Examples
Sep 09, 2026287views

Regex for IP Addresses: Accurate IPv4 and IPv6 Validation Examples

Use regex to find possible IP addresses, then use the right validation rule for the job. This guide covers loo...

#regex for IP address#IP address regex#IPv4 regex#IPv6 regex
Password Validation Regex: Examples for Length, Numbers, Symbols, and Strong Passwords
Sep 07, 2026268views

Password Validation Regex: Examples for Length, Numbers, Symbols, and Strong Passwords

Learn practical password validation regex patterns for length, digits, symbols, and mixed character classes. T...

#password validation regex#regex for password validation#strong password regex#regex lookahead
What Regex Should You Use for Phone Numbers? Patterns for US and International Formats
Aug 28, 2026256views

What Regex Should You Use for Phone Numbers? Patterns for US and International Formats

Compare practical regex patterns for US formatted numbers, normalized local numbers, plus-prefixed internation...

#phone number regex#regex for phone number#US phone regex#international phone regex
How to Validate a URL with Regex: HTTP, HTTPS, Domains, and Query Strings
Sep 03, 2026245views

How to Validate a URL with Regex: HTTP, HTTPS, Domains, and Query Strings

Build a practical HTTP and HTTPS URL regex by separating schemes, DNS domains, ports, paths, query strings, an...

#URL validation regex#regex for URL validation#HTTP URL regex#HTTPS URL regex
Regex for Numbers Only: Integers, Decimals, Negative Values, and Leading Zeros
Sep 12, 2026211views

Regex for Numbers Only: Integers, Decimals, Negative Values, and Leading Zeros

Choose the right regex for numbers-only input. This guide covers ASCII digits, integers, decimals, negative va...

#regex for numbers only#numbers only regex#integer regex#decimal regex
What Is the Best Regex for Email Validation? Practical Patterns and Limitations
Aug 31, 2026208views

What Is the Best Regex for Email Validation? Practical Patterns and Limitations

Choose a simple or practical email-validation regex based on your input policy, compare JavaScript and Python ...

#email validation regex#regex for email validation#email regex#JavaScript regex
Feedback email