Using Python String find() to Locate Substrings
How Python's str.find() locates substrings, including its return value semantics, start/end parameters, and comparisons with index(), in, and re.search().
To locate a substring inside a Python string, str.find() returns the lowest index where the substring occurs, or -1 if it is not found. This makes it useful for conditional checks where you need both existence and position without raising an exception.
line = 'the quick brown fox' position = line.find('quick') print(position) # 4
The method is available on all string objects and works with any substring, including empty strings. Its exact behavior, especially the return value and the optional parameters, matters in text parsing and validation code.
How str.find() Works
The signature of str.find() is str.find(sub[, start[, end]]). The sub argument is the substring you are looking for. The optional start and end parameters define a slice of the original string to search within, using the same semantics as slicing: start is inclusive and end is exclusive.
text = 'abcabcabc' print(text.find('abc')) # 0 print(text.find('abc', 1)) # 3 print(text.find('abc', 1, 6)) # 3
The method scans the string from left to right and returns the first match. If no match is found, it returns -1. A missing substring does not raise an error.
Return Value Semantics and Common Pitfalls
Because find() returns -1 when the substring is absent, check for >= 0 when you want to know whether a match exists. A common bug is if text.find(sub) > 0, which skips a match at index 0.
if text.find('start') >= 0: print('found')
Another pitfall is treating the return value as a boolean. -1 is truthy in Python, so if text.find(sub) is true when the substring is missing, and also true when the match starts after index 0; it is false only when the match starts at index 0. Always compare explicitly with >= 0 or == -1.
Using start and end to Restrict the Search
The start and end parameters search within a specific region of a string, such as after a known prefix or before a delimiter. They avoid building a slice in your own code, which would copy the string.
log_line = 'ERROR: disk full' if log_line.find('disk', 6) >= 0: print('disk-related error')
start and end follow slice semantics: negative values count from the end of the string, and out-of-range indices are adjusted to the nearest valid boundary. This is the same boundary handling used by slicing.
Differences Between find() and index()
The str.index() method behaves like find() but raises a ValueError when the substring is not found. This changes the error-handling strategy: find() is appropriate when absence is a normal condition, while index() is useful when a missing substring indicates malformed input that should fail loudly.
| Method | Missing substring behavior | Use case |
|---|---|---|
find() | Returns -1 | Conditional checks, optional substrings |
index() | Raises ValueError | Required substrings, input validation |
# find() for optional parsing pos = data.find('\n') if pos != -1: line = data[:pos] # index() for required format pos = data.index(':') # raises if missing
Choosing between them depends on whether the substring is expected to exist. If absence is a condition you handle explicitly, find() avoids a try/except block.
Performance and Runtime Behavior
The find() method searches for a literal substring directly and does not use a regex engine, so there is no pattern to compile or match. For simple literal substring searches, find() is generally the more direct choice than re.search().
However, find() does not support pattern matching. For a case-insensitive search, normalize the string with .lower() or .casefold() before calling find(), or use a regex with the re.IGNORECASE flag. Normalizing copies the string, which adds memory and time. For repeated searches on the same string, store the normalized version once.
# Case-insensitive search if text.lower().find('error') >= 0: print('error found')
Using start and end avoids creating a slice in your own code, and find() returns an integer rather than a new substring. This keeps extra memory use low for single searches.
When to Use find() vs in vs re.search
Python offers several ways to search for substrings, and the right choice depends on what you need beyond a simple index.
- Use
inwhen you only need a boolean result and do not care about the position. - Use
find()when you need the index of the first occurrence and want to handle absence without exceptions. - Use
re.search()when you need pattern matching, such as wildcards, character classes, or alternation.
# Boolean check if 'error' in text: pass # Index needed pos = text.find('error') # Pattern needed import re if re.search(r'error\s+\d+', text): pass
For a single literal substring, find() is more direct than re.search(). For multiple different substrings, a regex with alternation lets you search for all of them in one call. That can be worth it when the pattern is reused, despite regex compilation cost.
Edge Cases: Empty Substring and Overlapping Matches
find() with an empty substring returns the effective start position (or 0 if not specified), because an empty string is considered to exist at every position. This can lead to unexpected behavior if you do not guard against it.
print('abc'.find('')) # 0 print('abc'.find('', 2)) # 2
find() returns only the first occurrence. To find all occurrences, use a loop that advances the start position past the previous match; this can include overlapping matches if you advance by one character.
text = 'aaaa' start = 0 while True: pos = text.find('aa', start) if pos == -1: break print(pos) start = pos + 1 # allows overlapping
This loop prints 0, 1, and 2. If you set start = pos + 2, you get non-overlapping matches. The choice depends on whether overlapping matches are relevant to your use case.
Compatibility and Version Notes
str.find() has existed since early Python versions, and its behavior is stable across Python 2 and Python 3. In Python 3, strings are Unicode by default, so find() returns a character index, not a byte offset. For byte strings, bytes.find() works similarly.
If you need byte offsets, encode the string and search in the encoded bytes. The following example shows why byte offsets can differ from character offsets when earlier characters are multibyte:
s = 'éa' print(s.find('a')) # 1, character index b = s.encode('utf-8') print(b.find(b'a')) # 2, byte index: 'é' occupies 2 bytes
For most text processing, the character index is the correct choice. Only when interfacing with binary protocols or file offsets should you convert to bytes.