Classic Regex Interview Questions
I. Basic Concept Review
Before diving into specific questions, interviewers often start by checking your understanding of fundamental Regex concepts.
What is a Regular Expression?
A regular expression is a string that describes a text pattern. It can be used to match, search, and replace specific content within text using special characters called metacharacters.
What are common metacharacters and their meanings?
.(dot): Matches any single character except newline.*(asterisk): Matches the preceding character zero or more times.+(plus): Matches the preceding character one or more times.?(question mark): Matches the preceding character zero or one time.[](square brackets): Defines a character set, matching any one character inside.[^](negated square brackets): Defines a negated character set, matching any one character not inside.^(caret): Matches the start of the string.$(dollar sign): Matches the end of the string.\d: Matches any digit (0-9).\w: Matches any alphanumeric character (a-z, A-Z, 0-9, _).\s: Matches any whitespace character (space, tab, newline, etc.).()(parentheses): Used for grouping and capturing matched content.|(pipe): Represents alternation, matching either the left or right pattern.\(backslash): Escapes special characters, removing their special meaning.
What is greedy vs. non-greedy matching? How to achieve non-greedy matching?
- Greedy matching: By default, the Regex engine matches as many characters as possible.
- Non-greedy matching: The Regex engine matches as few characters as possible.
- Achieving non-greedy matching: Add a
?after quantifiers (*,+,?,{m,n}). For example,.*is greedy, while.*?is non-greedy.
II. Classic Interview Questions & Solutions
Below are common Regex interview questions with detailed explanations and Python code examples.
1. Validate Email Address
Question: Write a Regex to verify if an email address format is correct.
Approach: A simple email format usually includes: username, @ symbol, and domain.
Regex:
import re
def validate_email(email):
pattern = r"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$"
return bool(re.match(pattern, email))
# Tests
print(validate_email("[email protected]")) # True
print(validate_email("[email protected]")) # True
print(validate_email("invalid-email")) # False
print(validate_email("test@example")) # False
Explanation:^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$ matches a basic email format.
Note: This is a simple email validation Regex; stricter validation requires considering more cases like domain length and special characters.
2. Extract Content from HTML Tags
Question: Write a Regex to extract content from HTML tags, e.g., extract "This is a paragraph." from This is a paragraph..
Approach: Use grouping () to capture content inside tags.
Regex:
import re
def extract_content(html):
pattern = r"<[^>]+>(.*?)[^>]+>"
match = re.search(pattern, html)
if match:
return match.group(1) # Return first captured group
else:
return None
# Tests
html = "This is a paragraph.
"
print(extract_content(html)) # This is a paragraph.
html = "This is a heading
"
print(extract_content(html)) # This is a heading
html = "Some text"
print(extract_content(html)) # Some text
Explanation:<[^>]+>(.*?)[^>]+> matches HTML tags and uses non-greedy matching to extract content.
Note: This Regex only extracts content from simple HTML tags; for nested tags, more complex Regex or an HTML parser is needed.
3. Match IP Address
Question: Write a Regex to validate if an IP address format is correct.
Approach: An IP address consists of four numbers, each ranging 0-255, separated by dots.
Regex:
import re
def validate_ip(ip):
pattern = r"^((25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)$"
return bool(re.match(pattern, ip))
# Tests
print(validate_ip("192.168.1.1")) # True
print(validate_ip("10.0.0.255")) # True
print(validate_ip("256.0.0.1")) # False
print(validate_ip("192.168.1.256")) # False
print(validate_ip("192.168.1")) # False
Explanation:^((25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)$ matches IP address format and validates each number's range.
4. Replace Sensitive Words in String
Question: Write a function using Regex to replace sensitive words in a string with *.
Approach: Use the re.sub() function for replacement.
Code:
import re
def censor_words(text, sensitive_words):
pattern = "|".join(re.escape(word) for word in sensitive_words)
return re.sub(pattern, "*", text)
# Tests
text = "This is a bad word and another bad word."
sensitive_words = ["bad", "another"]
print(censor_words(text, sensitive_words)) # This is a * word and * * word.
text = "This is a test with special characters like . and *."
sensitive_words = [".", "*"]
print(censor_words(text, sensitive_words)) # This is a test with special characters like * and *.
Explanation: Uses re.sub() to replace matched sensitive words with *.
5. Extract Domain from URL
Question: Write a Regex to extract the domain name from a URL.
Approach: URL format is typically protocol://domain/path; we need to extract the domain part.
Regex:
import re
def extract_domain(url):
pattern = r"^(?:https?:\/\/)?(?:www\.)?([a-zA-Z0-9.-]+)\.([a-zA-Z]{2,6})(?:\/.*)?$"
match = re.match(pattern, url)
if match:
return match.group(1) + "." + match.group(2)
else:
return None
# Tests
print(extract_domain("http://www.example.com/path")) # example.com
print(extract_domain("https://example.co.uk")) # example.co.uk
print(extract_domain("example.com")) # example.com
print(extract_domain("sub.example.com/path")) # sub.example.com
Explanation:^(?:https?:\/\/)?(?:www\.)?([a-zA-Z0-9.-]+)\.([a-zA-Z]{2,6})(?:\/.*)?$ matches URL format and extracts the domain.
III. Summary & Tips
Mastering Regex requires continuous practice. Before your interview, we recommend:
- Review basic Regex concepts and metacharacters.
- Practice more exercises to familiarize yourself with common Regex patterns.
- Understand the difference between greedy and non-greedy matching.
- Get familiar with the Regex API in your programming language.
- During the interview, clearly explain your thought process and Regex patterns.
We hope this article helps you better prepare for Regex interviews—good luck!