How to Replace Text with Python Regex: re.sub() Examples and Common Mistakes
Replacing part of a string is one of the most useful jobs for a regular expression. In Python, the standard re.sub() function applies a pattern, inserts replacement text, and returns a new string. This guide explains the basic call, capture-group replacements, callback functions, case-insensitive and multiline matching, and two practical workflows for log redaction and text formatting.
What re.sub() does
re.sub() searches the input string for every non-overlapping match and replaces each match. It does not modify the original string in place, so keep the returned value:
import re
text = "Order 1042 is ready. Order 1043 is delayed."
result = re.sub(r"Order \d+", "Order [hidden]", text)
print(result)
# Order [hidden] is ready. Order [hidden] is delayed. The main form is re.sub(pattern, repl, string, count=0, flags=0). The pattern selects text, repl is either a string or a function, string is the source text, count limits the number of replacements, and flags changes how the pattern is matched.
Replace only the first few matches
Use count when the first match or first few matches are the only ones that should change:
text = "red, red, red"
result = re.sub(r"red", "blue", text, count=2)
print(result)
# blue, blue, red With the default count=0, every matching occurrence is replaced.
Use capture groups in replacement text
Parentheses capture parts of the match. In a replacement string, refer to them with \1, \2, or the clearer named form \g<name>:
text = "Smith, Jane"
result = re.sub(r"(?P<last>[A-Za-z]+), (?P<first>[A-Za-z]+)",
r"\g<first> \g<last>",
text)
print(result)
# Jane Smith Use raw strings for Python regex patterns and replacement strings whenever they contain backslashes. Named groups also avoid an ambiguity such as \10, which can be read as group 10 instead of group 1 followed by a zero.
Use a callback for conditional replacements
A replacement function receives a match object and returns the text that should replace it. This is useful when the replacement depends on the matched value:
def mask_number(match):
value = match.group()
return value[:2] + "*" * (len(value) - 4) + value[-2:]
text = "Call 4155550134 or 2125550199."
result = re.sub(r"\b\d{10}\b", mask_number, text)
print(result)
# Call 41******34 or 21******99. The callback can inspect named groups, convert numbers, look up a replacement in a dictionary, or apply different rules to different matches. Keep the callback deterministic and return a string for every match.
Case-insensitive and multiline replacement
Ignore letter case
Pass re.IGNORECASE when the same word may appear with different capitalization:
text = "Error: disk full\nERROR: timeout\nerror: retry"
result = re.sub(r"error:", "Warning:", text, flags=re.IGNORECASE)
print(result) This changes the matching rule, not the original capitalization of the replacement. If preserving capitalization matters, use a callback that examines match.group().
Process each line with multiline mode
re.MULTILINE makes ^ and $ refer to the start and end of each line instead of only the entire string:
text = "DEBUG ready\nDEBUG connected\nINFO complete"
result = re.sub(r"^DEBUG ", "", text, flags=re.MULTILINE)
print(result)
# ready
# connected
# INFO complete re.DOTALL is different: it allows a dot to match newline characters. Use it only when a multi-line block is intentionally part of one match.
Example: redact sensitive values in logs
A log-cleaning step often needs to keep the key while hiding the value. A non-greedy pattern stops at the next comma or line break:
import re
log = "user=alice, token=sk_live_123456, action=login\nuser=bob, token=sk_test_abcdef, action=logout"
redacted = re.sub(
r"(?i)(token=)[^,\n]+",
r"\1[redacted]",
log,
)
print(redacted)
# user=alice, token=[redacted], action=login
# user=bob, token=[redacted], action=logout Test the expression against representative logs before using it in production. If values can contain commas or span lines, define the real delimiter first; a broad .* can hide too much text.
Example: format identifiers consistently
Capture groups can insert separators into an unformatted identifier. This example turns eight digits into a readable date:
text = "Release dates: 20260917 and 20261201"
formatted = re.sub(r"\b(\d{4})(\d{2})(\d{2})\b", r"\1-\2-\3", text)
print(formatted)
# Release dates: 2026-09-17 and 2026-12-01 For more complicated formatting, prefer a callback so that invalid values can be rejected or normalized before the final text is returned.
Common escaping mistakes and fixes
- Backslashes disappear. Use raw strings such as
r"\d+"for patterns instead of doubling every backslash in an ordinary string. - The replacement creates an unexpected group reference. Use
r"\g<1>0"when you mean group 1 followed by zero, or use named groups. - Replacement text is treated as a pattern. The replacement string is not a second regex. For literal backslashes, use
re.escape()only when you are escaping text for a pattern; for replacement text, construct the desired string directly or use a callback. - Only one line changes. Add
re.MULTILINEwhen anchors should work per line, or process lines explicitly. - Too much text is replaced. Narrow the character class, use a non-greedy quantifier, or add boundaries such as
\b. Always test both matching and non-matching examples.
Use RegexToolBox to prepare the pattern
RegexToolBox can generate a regex from a plain-language requirement, show the generated pattern, switch the code example to Python, and validate the pattern against test content. A practical workflow is to describe the text you need to find, choose Python in the language selector, copy the generated pattern, and then place it in a small re.sub() test like the examples above. The site helps with pattern generation and testing; the replacement logic, callback behavior, and final safety checks still belong in your Python code.
Common questions
Does re.sub() change the original string? No. Python strings are immutable, so assign the returned value to a variable.
When should I use a callback instead of a replacement string? Use a callback when the replacement depends on the matched value, needs conditional logic, or must preserve part of the original capitalization.
Why does ^ match only once in a multi-line string? Without re.MULTILINE, the anchors refer to the beginning and end of the complete string. Add the flag when each line should be handled separately.
How can I avoid replacing more text than intended? Use a precise pattern, define boundaries, test non-matches, and inspect the result on realistic samples before processing a whole file or log stream.
Can RegexToolBox run Python replacement code for me? The current site provides AI regex generation, generated code examples, and validation testing. Use those outputs to prepare a pattern, then run re.sub() in your own Python environment.