Learn Regular Expressions from Scratch
1. What is a Regular Expression?
A Regular Expression (Regex) is a powerful text-processing tool that uses specific patterns to match, find, and manipulate strings. Whether you are a programmer, data analyst, or everyday user, mastering regular expressions can dramatically improve your text-processing efficiency.
Simple Examples:
\dmatches any digit (0-9)[a-z]matches any lowercase letter^hellomatches strings starting with "hello"
2. Basic Syntax Elements
1. Character Matching
- Literal characters: match themselves, e.g.
amatches the letter "a" - Dot
.: matches any single character (except newline) - Character class
[]: matches any one character inside, e.g.[aeiou]matches any vowel
2. Predefined Character Classes
\d: digit, equivalent to[0-9]\w: word character, equivalent to[a-zA-Z0-9_]\s: whitespace (space, tab, newline, etc.)- Uppercase forms negate the class, e.g.
\Dmatches non-digits
3. Quantifiers
*: 0 or more times+: 1 or more times?: 0 or 1 time{n}: exactly n times{n,}: at least n times{n,m}: between n and m times
4. Anchors
^: start of string$: end of string\b: word boundary
3. Practical Examples
1. Email Validation
^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$ Pattern Breakdown:
^- start of string[a-zA-Z0-9._%+-]+- username part, allows letters, digits and certain symbols@- required @ symbol[a-zA-Z0-9.-]+- domain part\.- required dot[a-zA-Z]{2,}$- TLD, at least 2 letters
Test Cases:
2. URL Extraction
(https?:\/\/(?:www\.)?[a-zA-Z0-9-]+\.[a-zA-Z]{2,}(?:\/[^\s]*)?) Pattern Breakdown:
https?- matches http or https:\/\/- matches ://(?:www\.)?- optional www.[a-zA-Z0-9-]+- domain body\.[a-zA-Z]{2,}- top-level domain(?:\/[^\s]*)?- optional path
Test Cases:
3. Date Format Matching (YYYY-MM-DD)
^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12][0-9]|3[01])$ Pattern Breakdown:
^\d{4}- 4-digit year(0[1-9]|1[0-2])- months 01-12(0[1-9]|[12][0-9]|3[01])- days 01-31$- end of string
Test Cases:
4. Extract HTML Tag Content
<([a-z]+)(?:\s+[^>]*)?>(.*?)<\/\1> Pattern Breakdown:
<([a-z]+)- matches opening tag name(?:\s+[^>]*)?- optional attributes>(.*?)<\/\1>- matches content until closing tag.*?- non-greedy match
Test Cases:
This is a paragraph
5. Password Strength Validation
Requirements: 8-20 chars, must include uppercase, lowercase, digit and special char
^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[@$!%*?&])[A-Za-z\d@$!%*?&]{8,20}$ Pattern Breakdown:
^- start of string(?=.*[a-z])- at least one lowercase(?=.*[A-Z])- at least one uppercase(?=.*\d)- at least one digit(?=.*[@$!%*?&])- at least one special[A-Za-z\d@$!%*?&]{8,20}- allowed chars & length$- end of string
Test Cases:
6. Extract Chinese Text
[\u4e00-\u9fa5]+ Pattern Breakdown:
[\u4e00-\u9fa5]- Unicode CJK range+- one or more Chinese characters
Test Cases:
4. Advanced Techniques
1. Capturing Groups
Use () to capture matched content
(\d{3})-(\d{4}) Matches "123-4567" and stores "123" and "4567"
2. Non-capturing Groups
Use (?:...) to group without capturing
(?:www\.)?example\.com 3. Positive/Negative Lookahead
(?=...)positive lookahead\d+(?=px)Matches digits followed by "px"
(?!...)negative lookahead\d+(?!px)Matches digits not followed by "px"
4. Greedy vs Lazy Matching
- Default is greedy (matches as much as possible)
<.*>Matches entire HTML tag and its content
- Add
?after quantifier for lazy (matches as little as possible)<.*?>Matches only the HTML tag itself
5. Learning Resources & Tools
1. Online Testing Tools
- 🔗
Regex101
Full-featured regex testing and learning platform
- 🔗
RegExr
User-friendly regex learning tool
- 🔗
RegexPal
Simple regex tester
2. Practice Platforms
- 🏆
RegexOne
Interactive regex tutorial
- 🏆
HackerRank Regex Challenges
Regex exercises on a competitive programming platform
3. Language Support
- 🐍
Python:
remodule - 🟨
JavaScript: built-in RegExp
- ☕
Java:
java.util.regexpackage - 🐘
PHP:
preg_functions