Home >Blog

How to Replace Text with Python Regex: re.sub() Examples and Common Mistakes

Published Updated

Replacing part of a string is one of the most useful jobs for a regular expression. In Python, the standard re.sub() function applies a pattern, inserts replacement text, and returns a new string. This guide explains the basic call, capture-group replacements, callback functions, case-insensitive and multiline matching, and two practical workflows for log redaction and text formatting.

What re.sub() does

re.sub() searches the input string for every non-overlapping match and replaces each match. It does not modify the original string in place, so keep the returned value:

import re

text = "Order 1042 is ready. Order 1043 is delayed."
result = re.sub(r"Order \d+", "Order [hidden]", text)
print(result)
# Order [hidden] is ready. Order [hidden] is delayed.

The main form is re.sub(pattern, repl, string, count=0, flags=0). The pattern selects text, repl is either a string or a function, string is the source text, count limits the number of replacements, and flags changes how the pattern is matched.

Replace only the first few matches

Use count when the first match or first few matches are the only ones that should change:

text = "red, red, red"
result = re.sub(r"red", "blue", text, count=2)
print(result)
# blue, blue, red

With the default count=0, every matching occurrence is replaced.

Use capture groups in replacement text

Parentheses capture parts of the match. In a replacement string, refer to them with \1, \2, or the clearer named form \g<name>:

text = "Smith, Jane"
result = re.sub(r"(?P<last>[A-Za-z]+), (?P<first>[A-Za-z]+)",
                r"\g<first> \g<last>",
                text)
print(result)
# Jane Smith

Use raw strings for Python regex patterns and replacement strings whenever they contain backslashes. Named groups also avoid an ambiguity such as \10, which can be read as group 10 instead of group 1 followed by a zero.

Use a callback for conditional replacements

A replacement function receives a match object and returns the text that should replace it. This is useful when the replacement depends on the matched value:

def mask_number(match):
    value = match.group()
    return value[:2] + "*" * (len(value) - 4) + value[-2:]

text = "Call 4155550134 or 2125550199."
result = re.sub(r"\b\d{10}\b", mask_number, text)
print(result)
# Call 41******34 or 21******99.

The callback can inspect named groups, convert numbers, look up a replacement in a dictionary, or apply different rules to different matches. Keep the callback deterministic and return a string for every match.

Case-insensitive and multiline replacement

Ignore letter case

Pass re.IGNORECASE when the same word may appear with different capitalization:

text = "Error: disk full\nERROR: timeout\nerror: retry"
result = re.sub(r"error:", "Warning:", text, flags=re.IGNORECASE)
print(result)

This changes the matching rule, not the original capitalization of the replacement. If preserving capitalization matters, use a callback that examines match.group().

Process each line with multiline mode

re.MULTILINE makes ^ and $ refer to the start and end of each line instead of only the entire string:

text = "DEBUG ready\nDEBUG connected\nINFO complete"
result = re.sub(r"^DEBUG ", "", text, flags=re.MULTILINE)
print(result)
# ready
# connected
# INFO complete

re.DOTALL is different: it allows a dot to match newline characters. Use it only when a multi-line block is intentionally part of one match.

Example: redact sensitive values in logs

A log-cleaning step often needs to keep the key while hiding the value. A non-greedy pattern stops at the next comma or line break:

import re

log = "user=alice, token=sk_live_123456, action=login\nuser=bob, token=sk_test_abcdef, action=logout"
redacted = re.sub(
    r"(?i)(token=)[^,\n]+",
    r"\1[redacted]",
    log,
)
print(redacted)
# user=alice, token=[redacted], action=login
# user=bob, token=[redacted], action=logout

Test the expression against representative logs before using it in production. If values can contain commas or span lines, define the real delimiter first; a broad .* can hide too much text.

Example: format identifiers consistently

Capture groups can insert separators into an unformatted identifier. This example turns eight digits into a readable date:

text = "Release dates: 20260917 and 20261201"
formatted = re.sub(r"\b(\d{4})(\d{2})(\d{2})\b", r"\1-\2-\3", text)
print(formatted)
# Release dates: 2026-09-17 and 2026-12-01

For more complicated formatting, prefer a callback so that invalid values can be rejected or normalized before the final text is returned.

Common escaping mistakes and fixes

  • Backslashes disappear. Use raw strings such as r"\d+" for patterns instead of doubling every backslash in an ordinary string.
  • The replacement creates an unexpected group reference. Use r"\g<1>0" when you mean group 1 followed by zero, or use named groups.
  • Replacement text is treated as a pattern. The replacement string is not a second regex. For literal backslashes, use re.escape() only when you are escaping text for a pattern; for replacement text, construct the desired string directly or use a callback.
  • Only one line changes. Add re.MULTILINE when anchors should work per line, or process lines explicitly.
  • Too much text is replaced. Narrow the character class, use a non-greedy quantifier, or add boundaries such as \b. Always test both matching and non-matching examples.

Use RegexToolBox to prepare the pattern

RegexToolBox can generate a regex from a plain-language requirement, show the generated pattern, switch the code example to Python, and validate the pattern against test content. A practical workflow is to describe the text you need to find, choose Python in the language selector, copy the generated pattern, and then place it in a small re.sub() test like the examples above. The site helps with pattern generation and testing; the replacement logic, callback behavior, and final safety checks still belong in your Python code.

Common questions

Does re.sub() change the original string? No. Python strings are immutable, so assign the returned value to a variable.

When should I use a callback instead of a replacement string? Use a callback when the replacement depends on the matched value, needs conditional logic, or must preserve part of the original capitalization.

Why does ^ match only once in a multi-line string? Without re.MULTILINE, the anchors refer to the beginning and end of the complete string. Add the flag when each line should be handled separately.

How can I avoid replacing more text than intended? Use a precise pattern, define boundaries, test non-matches, and inspect the result on realistic samples before processing a whole file or log stream.

Can RegexToolBox run Python replacement code for me? The current site provides AI regex generation, generated code examples, and validation testing. Use those outputs to prepare a pattern, then run re.sub() in your own Python environment.

FAQ

Does re.sub() change the original Python string?
No. Python strings are immutable, so assign the value returned by re.sub() to a variable.
How do capture groups work in a replacement string?
Parenthesized groups can be reused with backreferences such as \1 or the named form \g; named groups are clearer when a replacement is complex.
When should a replacement callback be used?
Use a callback when the replacement depends on the matched value, needs conditional logic, or must preserve information from the match.
How do I replace matches on every line?
Use re.MULTILINE when ^ and $ should match the beginning and end of each line in a multi-line string.
Why do backslashes cause re.sub() errors?
Python string escaping and regex escaping are separate layers. Raw strings reduce accidental escapes, and named backreferences avoid ambiguous replacement syntax.

目次

情報

  • ヒット202
  • 公開日2026/09/18
0/500
敬意をもってご意見をお聞かせください。

もっと投稿

このセクションの記事をさらに探索
Regex for IP Addresses: Accurate IPv4 and IPv6 Validation Examples
Sep 09, 2026288ビュー

Regex for IP Addresses: Accurate IPv4 and IPv6 Validation Examples

Use regex to find possible IP addresses, then use the right validation rule for the job. This guide covers loo...

#regex for IP address#IP address regex#IPv4 regex#IPv6 regex
Password Validation Regex: Examples for Length, Numbers, Symbols, and Strong Passwords
Sep 07, 2026268ビュー

Password Validation Regex: Examples for Length, Numbers, Symbols, and Strong Passwords

Learn practical password validation regex patterns for length, digits, symbols, and mixed character classes. T...

#password validation regex#regex for password validation#strong password regex#regex lookahead
What Regex Should You Use for Phone Numbers? Patterns for US and International Formats
Aug 28, 2026256ビュー

What Regex Should You Use for Phone Numbers? Patterns for US and International Formats

Compare practical regex patterns for US formatted numbers, normalized local numbers, plus-prefixed internation...

#phone number regex#regex for phone number#US phone regex#international phone regex
How to Validate a URL with Regex: HTTP, HTTPS, Domains, and Query Strings
Sep 03, 2026245ビュー

How to Validate a URL with Regex: HTTP, HTTPS, Domains, and Query Strings

Build a practical HTTP and HTTPS URL regex by separating schemes, DNS domains, ports, paths, query strings, an...

#URL validation regex#regex for URL validation#HTTP URL regex#HTTPS URL regex
How to Extract Email Addresses from Text with Regex in Python and JavaScript
Sep 15, 2026211ビュー

How to Extract Email Addresses from Text with Regex in Python and JavaScript

Extract email addresses from any text with regex in Python and JavaScript: the pattern explained piece by piec...

#regex extract email#email regex python#email regex javascript#extract emails from text
Regex for Numbers Only: Integers, Decimals, Negative Values, and Leading Zeros
Sep 12, 2026211ビュー

Regex for Numbers Only: Integers, Decimals, Negative Values, and Leading Zeros

Choose the right regex for numbers-only input. This guide covers ASCII digits, integers, decimals, negative va...

#regex for numbers only#numbers only regex#integer regex#decimal regex
フィードバックメール