Regex for IP Addresses: Accurate IPv4 and IPv6 Validation Examples
An IP address looks simple until a regex has to decide what counts as one. The string 192.168.1.1 is familiar, but 999.1.1.1, 01.2.3.4, 2001:db8::1, and ::ffff:192.0.2.128 force you to make a policy choice. Are you scanning a log for things that look like addresses, or rejecting every value that a network library would reject? Those are different jobs, and they need different tools.
This guide keeps the distinction explicit. Use a forgiving regex when you need candidates from text. Use a strict IPv4 regex only when the input contract is truly IPv4. For IPv6, a small regex is often a trap; a language parser is usually clearer and more accurate.
Start by choosing extraction or validation
Extraction asks whether a stretch of text resembles an IP address. A web-server log may contain punctuation, a port, a surrounding URL, or a malformed value you still want to review. A loose pattern is useful here because missing a candidate can be worse than returning a few false positives.
Validation asks a stricter question: is the whole supplied value a valid address under this application's rules? That means anchors, numeric range checks, and a decision about leading zeroes. It may also mean accepting IPv6, an IPv4-mapped IPv6 address, or a zone identifier. Do not reuse a log-scanning pattern as a form validator just because it matched the happy-path examples.
Candidate IPv4 extraction:
\b(?:\d{1,3}\.){3}\d{1,3}\b
Strict whole-string IPv4 validation:
^(?:(?:25[0-5]|2[0-4]\d|1\d{2}|[1-9]?\d)\.){3}(?:25[0-5]|2[0-4]\d|1\d{2}|[1-9]?\d)$ The first pattern intentionally accepts 300.400.500.600. It is a candidate finder. The second checks four decimal octets and keeps every octet within 0 through 255. It also rejects multi-digit octets with a leading zero, such as 001. That is a sensible default for modern dotted-decimal input, though legacy systems sometimes choose a different rule.
A simple IPv4 regex for controlled input
If your UI already separates the four octets, or you only need a readable first-pass check, use this simpler form:
^(?:\d{1,3}\.){3}\d{1,3}$ It confirms the shape: four groups of one to three digits, separated by dots. It does not know that 256 is too large. That limitation is fine when a later numeric check handles the ranges, or when the pattern is only used to route input to another parser. It is not enough for a public API boundary on its own.
Keeping this pattern small has a practical benefit. People can read it six months later. A strict regex is compact enough to use, but it deserves a comment or a named constant so nobody assumes it accepts every dotted sequence.
How the strict IPv4 pattern works
The useful building block is one decimal octet:
25[0-5] # 250 through 255
2[0-4]\d # 200 through 249
1\d{2} # 100 through 199
[1-9]?\d # 0 through 99, with no 00 or 01 Join that alternative with dots, repeat it for the first three octets, and place the final octet after the repetition. The anchors ^ and $ say that the entire input must match. In JavaScript, prefer $/ for ordinary single-line input; if your value can carry a trailing newline from a file or form, trim it before validation rather than relying on regex flags to hide the problem.
The pattern treats 0.0.0.0, 127.0.0.1, 192.168.0.1, and 255.255.255.255 as syntactically valid. That does not mean they are valid for every business rule. A firewall screen might reject multicast, unspecified, loopback, private, or broadcast ranges after syntax validation. Regex answers only the grammar question.
Boundaries matter when extracting from prose
Word boundaries are handy for plain log lines, but they are not a universal token boundary. In many regex engines, underscores count as word characters. A candidate next to an underscore may therefore behave differently from one next to a comma. URLs and host:port strings add another wrinkle: 192.0.2.10:8080 contains a valid IPv4 candidate but is not itself an IPv4 address.
For plain text, this is a reasonable extraction pattern:
(?<!\d)(?:\d{1,3}\.){3}\d{1,3}(?!\d) It avoids taking a substring from a longer digit run. The lookbehind is unsupported in a few older runtimes, so use the word-boundary version or a capture-group approach when compatibility matters. Either way, run every candidate through a strict validator before storing it as an address.
Why IPv6 resists a tidy regex
IPv6 is flexible by design. It allows eight hexadecimal groups, omits leading zeroes, compresses one run of zero groups with ::, and can end with an embedded IPv4 address. These values may represent the same address:
2001:0db8:0000:0000:0000:ff00:0042:8329
2001:db8::ff00:42:8329 A full IPv6 regex must count groups before and after ::, prevent a second compression marker, handle the all-zero shorthand ::, and decide whether IPv4-mapped forms are allowed. It can be written, but it is difficult to audit and easy to change incorrectly. A short pattern such as ^[0-9A-Fa-f:]+$ only checks that the characters look plausible. It accepts many invalid addresses, including repeated colons in the wrong places.
For extraction, a deliberately loose IPv6 candidate pattern can be useful:
(?i)(?<![:0-9a-f])[0-9a-f]{0,4}(?::[0-9a-f]{0,4}){2,7}(?![:0-9a-f]) Treat it as a search aid, not proof. It may miss an unusual legal form or return text that fails strict parsing. That trade-off is acceptable for triage and unacceptable for a security decision.
Use a native parser when correctness matters
Most mainstream languages already include an IP parser. It understands IPv4 and IPv6 syntax, makes intent obvious, and is easier to test than a long regex. Pair it with a small regex only when you need to locate possible tokens inside a larger string.
// JavaScript (Node.js)
import net from 'node:net';
const family = net.isIP(value); // 0, 4, or 6
# Python
import ipaddress
ipaddress.ip_address(value) # raises ValueError when invalid
// PHP
filter_var($value, FILTER_VALIDATE_IP);
// .NET
IPAddress.TryParse(value, out var address); Native parsers still need policy around them. Decide whether whitespace is trimmed, whether bracketed IPv6 from a URL such as [2001:db8::1] is accepted, and whether a port belongs in the same input. Usually, parse the host and port separately. Do not pass a whole URL or host:port string to an address parser and then blame the parser when it rejects it.
Do not mix the address with its transport syntax
An address often arrives with context that belongs to another grammar. A proxy log might contain 203.0.113.7:443; a URL uses brackets for an IPv6 host, as in https://[2001:db8::7]:8443/health. Neither full string is a bare IP address. A strict address validator should reject both, because accepting several formats in one field makes later code guess what the user meant.
Split the work into stages. First parse the URL with a URL library when the input is a URL. Then obtain the hostname and port from that parsed result. For a connection endpoint, parse the port as an integer in the range 1 through 65535 and validate the host separately. IPv6 makes this separation especially important: colons are part of the address and also separate a port in an endpoint. Brackets resolve that ambiguity in URL syntax, but they are not part of the raw address value.
The same idea applies to CIDR notation. 10.0.0.0/8 describes a network, not one host address. Validate the IP portion and the prefix length as separate pieces, or use a network parser such as Python's ipaddress.ip_network(). A regex that tries to accept addresses, ports, brackets, and CIDR suffixes all at once tends to become broad enough to accept malformed combinations.
Pick the smallest rule that fits the input
| Input you have | Recommended approach |
|---|---|
| Free-form log or message text | Loose candidate regex, then parse each candidate |
| One IPv4-only form field | Strict anchored IPv4 regex or a native parser |
| IPv4 or IPv6 form field | Native IP parser |
| URL, endpoint, or CIDR range | Use a URL, endpoint, or network parser first |
This is less about avoiding regex and more about giving each parser a single, clear job. A readable validation path is easier to review during an incident, and the test cases map directly to the formats the application claims to accept.
Language-specific regex escaping
The regex itself is only part of the code. A backslash may be consumed by the string literal before the regex engine sees it. JavaScript regex literals are the least noisy:
const ipv4 = /^(?:(?:25[0-5]|2[0-4]\d|1\d{2}|[1-9]?\d)\.){3}(?:25[0-5]|2[0-4]\d|1\d{2}|[1-9]?\d)$/;
const valid = ipv4.test(value); In a JavaScript string passed to new RegExp(), each regex backslash needs another backslash: \\d in the string becomes \d for the regex engine. Python raw strings, such as r'...', avoid most double escaping. In PHP, a single-quoted pattern is often easier to read than a double-quoted one. Test the exact source form used by your application, not a copied display version.
Boundary tests for your test suite
Do not stop at a few normal office-network values. These cases catch the range, anchor, and leading-zero mistakes that turn up in production:
| Input | Strict IPv4 result | Reason |
|---|---|---|
0.0.0.0 | Accept | Lowest valid numeric address |
255.255.255.255 | Accept | Highest valid numeric address |
256.0.0.1 | Reject | First octet exceeds 255 |
192.168.1 | Reject | Only three octets |
192.168.1.1.5 | Reject | Anchors prevent a partial match |
192.168.001.1 | Reject by default | Leading-zero policy |
2001:db8::1 | Reject | Valid IPv6, not IPv4 |
Also test copied input with spaces, a trailing newline, tabs, and Unicode lookalike dots. Whether those should be normalized is an application decision. Security-sensitive endpoints are usually better off rejecting unexpected characters and reporting a clear input error.
A practical rule of thumb
Use a loose regex to pull possible addresses from text. Use the strict IPv4 regex when you specifically require an unbracketed dotted IPv4 value. Use a built-in IP parser for IPv6 or for any boundary where a false acceptance has consequences. That split keeps the regex understandable and leaves address semantics to code that was built for them.