Python Wildcard Pattern Matching with fnmatch and glob
Learn Python's wildcard pattern matching with fnmatch, glob, and pathlib to filter strings and file paths using *, ?, and character sets.
Python's wildcard pattern matching is useful when you want to filter strings or file paths with simple patterns such as *.txt or data_?.csv. The standard library provides three main tools: fnmatch for plain string matching, glob for filesystem paths represented as strings, and pathlib.Path.glob() for object-oriented path matching. They share the same wildcard characters: * matches any sequence of characters, ? matches one character, and [seq] matches one character from a set or range. For path globbing, wildcards apply per path segment, so they do not cross directory boundaries.
Matching Strings with fnmatch
The fnmatch module translates a wildcard pattern into a regular expression internally and matches it against the whole string. The core function is fnmatch.fnmatch(name, pattern), which returns True when name matches pattern:
import fnmatch print(fnmatch.fnmatch('report_final.txt', 'report_*.txt')) # True print(fnmatch.fnmatch('data_1.csv', 'data_?.csv')) # True print(fnmatch.fnmatch('data_10.csv', 'data_?.csv')) # False print(fnmatch.fnmatch('log.txt', 'log[0-9].txt')) # False
The [seq] syntax supports ranges like [a-z] and negation with [!seq]. On Windows, fnmatch normalizes case because the underlying OS treats filenames case-insensitively. On other platforms, matching is case-sensitive by default. If you need consistent case-sensitive behavior across operating systems, use fnmatch.fnmatchcase instead.
fnmatch also provides fnmatch.filter(names, pattern) to filter an iterable of strings in one call:
files = ['a.py', 'b.py', 'c.txt', 'd.py'] py_files = fnmatch.filter(files, '*.py') print(py_files) # ['a.py', 'b.py', 'd.py']
fnmatch.filter translates the pattern once and applies it to each name, which is more concise than using fnmatch in a list comprehension.
Matching File Paths with glob
While fnmatch works on plain strings, glob applies wildcard patterns to filesystem paths. The glob.glob(pathname) function returns a list of path strings that match the pattern. Because glob matches each path segment separately, *, ?, and bracket expressions do not cross directory boundaries; *.txt matches only files in the current directory. To match recursively, use ** with recursive=True:
import glob # All .py files in current directory print(glob.glob('*.py')) # All .py files in current directory and subdirectories print(glob.glob('**/*.py', recursive=True))
glob also supports the ? and [seq] forms. The returned paths are not sorted by default, so sort them yourself if order matters. If no matches are found, an empty list is returned.
A common mistake is assuming that glob returns absolute paths. It returns paths exactly as constructed from the pattern, so relative patterns yield relative paths. Use os.path.abspath or pathlib.Path.resolve() if you need absolute paths.
Using pathlib for Wildcard Paths
The pathlib module offers an object-oriented interface for filesystem paths and includes wildcard support through the Path.glob() method. This method returns a generator of Path objects that match the given pattern. The pattern syntax is the same as glob, including ** for recursive matching:
from pathlib import Path for path in Path('.').glob('*.py'): print(path) for path in Path('.').glob('**/*.txt'): print(path)
Using Path.glob() is often more convenient than glob.glob() because you can chain path operations directly:
configs = [p for p in Path('config').glob('*.json') if p.stat().st_size > 1000]
pathlib also provides Path.rglob(pattern) for recursive matching. It is equivalent to Path.glob('**/' + pattern), and useful for finding files with a specific extension in a directory tree:
for path in Path('src').rglob('*.py'): print(path)
When Wildcards Are Not Enough: Use Regular Expressions
Wildcard syntax is intentionally limited. It cannot express patterns like a date in YYYY-MM-DD format or a name that does not contain a digit. When you need quantifiers, alternation, lookaheads, or more precise character classes, use the re module directly:
import re date_file = re.compile(r'\d{4}-\d{2}-\d{2}\.txt') for name in filenames: if date_file.fullmatch(name): print(name)
The fnmatch.translate() function exposes the regular expression that fnmatch generates, but in practice you rarely need to convert a wildcard to a regex by hand. Write the regular expression explicitly when the pattern needs more logic than wildcards provide.
Performance and Runtime Considerations
fnmatch caches the compiled form of each distinct pattern, so repeated calls with the same pattern are usually fast. If profiling shows that translation overhead matters because you are matching many different patterns, translate and compile the pattern yourself:
import fnmatch import re pattern = re.compile(fnmatch.translate('*.txt')) for name in many_names: if pattern.match(name): # process
This bypasses the OS-specific case normalization applied by fnmatch.fnmatch, so use it when you want consistent case-sensitive behavior. glob performs filesystem I/O, so its performance depends on the number of files and directories scanned. Using ** recursively can be expensive on large trees; if you only need a single level, avoid ** and use * instead.
Another memory consideration: glob.glob() returns a list, which can be large if many files match. pathlib.Path.glob() returns a generator, allowing you to process matches lazily and avoid holding all results in memory.
Common Mistakes and Edge Cases
Hidden files in glob
Unlike a plain fnmatch string match, glob treats a leading dot specially: a pattern such as *.txt will not match .hidden.txt. Use an explicit dot in the pattern, such as .*.txt, to include hidden files.
Wildcards and path separators
glob applies wildcard patterns to path segments, so *, ?, and [seq] do not cross path separators. fnmatch, by contrast, matches a whole string: fnmatch.fnmatch('a/b', 'a?b') returns True, while glob.glob('a?b') will not match a path named a/b. This difference matters when you test file paths with fnmatch and expect glob behavior.
Bracket expressions
[!a-z] matches any character not in the specified range; the ! must be the first character after [. fnmatch does not use backslashes to escape wildcard characters. To match a literal * or ?, use [*] or [?].
Choosing the Right Wildcard Tool
Use fnmatch when you need to match strings that are not necessarily file paths, such as filtering log messages or validating a simple pattern. Use glob when you need to find files on disk and want the results as strings. Use pathlib.Path.glob() when you are already working with Path objects, want to chain path operations, or need a generator to handle large directory trees.
For patterns that require more expressive power than wildcards provide, switch to regular expressions. The decision is not about which tool is better; it is about matching the tool to the complexity of the pattern and the data source. Wildcards are ideal for quick, human-readable filters; regex is necessary when the pattern has logical conditions that wildcards cannot encode.