Python Regular Expressions (re Module)
Regular expressions (regex or regexp) are powerful sequences of characters that define a search pattern. They are widely used for string searching, manipulation, and validation. Python’s built-in re module provides full support for regular expressions, allowing you to perform complex text processing tasks efficiently.
1. Using Word Boundaries (\b)
The \b special sequence matches an empty string, but only when it is at the beginning or end of a word. This allows you to match whole words and prevent partial matches.
import re
text = "The time has come. This is a time-sensitive issue."
time_pattern = r"\btime\b"
# re.findall finds all non-overlapping matches
print(re.findall(time_pattern, text)) # Output: ['time', 'time']
# re.search finds the first occurrence
match = re.search(time_pattern, text)
if match:
print(f"First match found: {match.group()}") # Output: First match found: time
# Explanation: '\btime\b' matches "time" but not "time-sensitive" because of the word boundary.2. Matching Network Ports (0-65535)
This pattern matches any valid TCP/UDP port number, which ranges from 0 to 65535.
port_pattern = r'\b((6553[0-5])|(655[0-2][0-9])|(65[0-4][0-9]{2})|(6[0-4][0-9]{3})|([1-5][0-9]{4})|([0-5]{1,4}))\b'
# A simpler, often sufficient pattern for common use, but less strict:
# port_pattern = r'\b(?:[0-9]{1,4}|[1-5][0-9]{4}|6[0-4][0-9]{3}|65[0-4][0-9]{2}|655[0-3][0-5])\b'
# For the original pattern from the document, which captures 0-65000 and 65535:
port_pattern_original = r'\b([0-9]|[1-5][0-9]{1,4}|6[0-4][0-9]{3}|65000)\b'
# The original pattern is slightly incomplete as it misses ports between 65001-65534.
# A more robust pattern for 0-65535:
port_pattern_robust = r'\b(?:[0-9]|[1-9][0-9]{1,3}|[1-5][0-9]{4}|6[0-4][0-9]{3}|65[0-4][0-9]{2}|655[0-2][0-9]|6553[0-5])\b'
# Example:
text_ports = "Server runs on port 80 and a special service on 65000."
print(re.findall(port_pattern_robust, text_ports)) # Output: ['80', '65000']Note: The original pattern in the document r'\b([0-9]|[1-5][0-9]{1,4}|6[0-4][0-9]{3}|65000)\b' correctly handles 0-65000, but would miss ports 65001-65534. The port_pattern_robust provides a more complete range for 0-65535.
3. Extracting Content Between Quotes
This pattern uses a non-greedy quantifier *? to match any characters between double quotes.
text_quotes = 'He said "Hello World" and then "Goodbye!".'
between_quotes_pattern = r'"(.*?)"'
print(re.findall(between_quotes_pattern, text_quotes)) # Output: ['Hello World', 'Goodbye!'](.*?): This is a capturing group..matches any character (except newline).*matches zero or more occurrences.?makes the*non-greedy, meaning it matches the shortest possible string.
4. Matching IP Addresses (IPv4)
A common pattern for recognizing standard IPv4 addresses.
text_ip = "Connect to 192.168.1.100 or 10.0.0.5."
ip_pattern = r"(\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3})"
print(re.findall(ip_pattern, text_ip)) # Output: ['192.168.1.100', '10.0.0.5']\d{1,3}: Matches a digit (\d) one to three times.\.: Matches a literal dot (escaped because.is a special regex character).
5. Extracting Content Around an Equal Sign
Content Before Equal Sign
text_equal = "name=Alice, age=30, city=New York"
pattern_before_equal = r'([^=]+)\s*='
print(re.findall(pattern_before_equal, text_equal)) # Output: ['name', ' age', ' city']([^=]+): Captures one or more characters that are not an equal sign.\s*: Matches zero or more whitespace characters.
Content After Equal Sign
pattern_after_equal = r'=\s*([^,]+)'
print(re.findall(pattern_after_equal, text_equal)) # Output: ['Alice', '30', 'New York']=\s*: Matches an equal sign followed by zero or more whitespace characters.([^,]+): Captures one or more characters that are not a comma.
6. Common re Module Functions
re.match(pattern, string): Matches only at the beginning of the string.re.search(pattern, string): Scans through the string looking for the first location where the pattern produces a match.re.findall(pattern, string): Finds all non-overlapping matches ofpatterninstring, returning them as a list of strings or tuples.re.finditer(pattern, string): Similar tofindall, but returns an iterator yielding match objects.re.sub(pattern, repl, string): Replaces occurrences ofpatterninstringwithrepl.re.split(pattern, string): Splitsstringby occurrences ofpattern.
Tip: Use online regex testers (e.g., regex101.com) to build and test your patterns interactively.