Back to Blog
Python

Using Python String find() to Locate Substrings

How Python's str.find() locates substrings, including its return value semantics, start/end parameters, and comparisons with index(), in, and re.search().

PythonString MethodsSubstring Searchstr.findText Processing
Python string find method locating a substring within a text string, shown as a magnifying glass over a line of code.

To locate a substring inside a Python string, str.find() returns the lowest index where the substring occurs, or -1 if it is not found. This makes it useful for conditional checks where you need both existence and position without raising an exception.

line = 'the quick brown fox' position = line.find('quick') print(position) # 4

The method is available on all string objects and works with any substring, including empty strings. Its exact behavior, especially the return value and the optional parameters, matters in text parsing and validation code.

How str.find() Works

The signature of str.find() is str.find(sub[, start[, end]]). The sub argument is the substring you are looking for. The optional start and end parameters define a slice of the original string to search within, using the same semantics as slicing: start is inclusive and end is exclusive.

text = 'abcabcabc' print(text.find('abc')) # 0 print(text.find('abc', 1)) # 3 print(text.find('abc', 1, 6)) # 3

The method scans the string from left to right and returns the first match. If no match is found, it returns -1. A missing substring does not raise an error.

Return Value Semantics and Common Pitfalls

Because find() returns -1 when the substring is absent, check for >= 0 when you want to know whether a match exists. A common bug is if text.find(sub) > 0, which skips a match at index 0.

if text.find('start') >= 0: print('found')

Another pitfall is treating the return value as a boolean. -1 is truthy in Python, so if text.find(sub) is true when the substring is missing, and also true when the match starts after index 0; it is false only when the match starts at index 0. Always compare explicitly with >= 0 or == -1.

Using start and end to Restrict the Search

The start and end parameters search within a specific region of a string, such as after a known prefix or before a delimiter. They avoid building a slice in your own code, which would copy the string.

log_line = 'ERROR: disk full' if log_line.find('disk', 6) >= 0: print('disk-related error')

start and end follow slice semantics: negative values count from the end of the string, and out-of-range indices are adjusted to the nearest valid boundary. This is the same boundary handling used by slicing.

Differences Between find() and index()

The str.index() method behaves like find() but raises a ValueError when the substring is not found. This changes the error-handling strategy: find() is appropriate when absence is a normal condition, while index() is useful when a missing substring indicates malformed input that should fail loudly.

MethodMissing substring behaviorUse case
find()Returns -1Conditional checks, optional substrings
index()Raises ValueErrorRequired substrings, input validation
# find() for optional parsing pos = data.find('\n') if pos != -1: line = data[:pos] # index() for required format pos = data.index(':') # raises if missing

Choosing between them depends on whether the substring is expected to exist. If absence is a condition you handle explicitly, find() avoids a try/except block.

Performance and Runtime Behavior

The find() method searches for a literal substring directly and does not use a regex engine, so there is no pattern to compile or match. For simple literal substring searches, find() is generally the more direct choice than re.search().

However, find() does not support pattern matching. For a case-insensitive search, normalize the string with .lower() or .casefold() before calling find(), or use a regex with the re.IGNORECASE flag. Normalizing copies the string, which adds memory and time. For repeated searches on the same string, store the normalized version once.

# Case-insensitive search if text.lower().find('error') >= 0: print('error found')

Using start and end avoids creating a slice in your own code, and find() returns an integer rather than a new substring. This keeps extra memory use low for single searches.

When to Use find() vs in vs re.search

Python offers several ways to search for substrings, and the right choice depends on what you need beyond a simple index.

  • Use in when you only need a boolean result and do not care about the position.
  • Use find() when you need the index of the first occurrence and want to handle absence without exceptions.
  • Use re.search() when you need pattern matching, such as wildcards, character classes, or alternation.
# Boolean check if 'error' in text: pass # Index needed pos = text.find('error') # Pattern needed import re if re.search(r'error\s+\d+', text): pass

For a single literal substring, find() is more direct than re.search(). For multiple different substrings, a regex with alternation lets you search for all of them in one call. That can be worth it when the pattern is reused, despite regex compilation cost.

Edge Cases: Empty Substring and Overlapping Matches

find() with an empty substring returns the effective start position (or 0 if not specified), because an empty string is considered to exist at every position. This can lead to unexpected behavior if you do not guard against it.

print('abc'.find('')) # 0 print('abc'.find('', 2)) # 2

find() returns only the first occurrence. To find all occurrences, use a loop that advances the start position past the previous match; this can include overlapping matches if you advance by one character.

text = 'aaaa' start = 0 while True: pos = text.find('aa', start) if pos == -1: break print(pos) start = pos + 1 # allows overlapping

This loop prints 0, 1, and 2. If you set start = pos + 2, you get non-overlapping matches. The choice depends on whether overlapping matches are relevant to your use case.

Compatibility and Version Notes

str.find() has existed since early Python versions, and its behavior is stable across Python 2 and Python 3. In Python 3, strings are Unicode by default, so find() returns a character index, not a byte offset. For byte strings, bytes.find() works similarly.

If you need byte offsets, encode the string and search in the encoded bytes. The following example shows why byte offsets can differ from character offsets when earlier characters are multibyte:

s = 'éa' print(s.find('a')) # 1, character index b = s.encode('utf-8') print(b.find(b'a')) # 2, byte index: 'é' occupies 2 bytes

For most text processing, the character index is the correct choice. Only when interfacing with binary protocols or file offsets should you convert to bytes.

Python String find(): Usage, Return Values, and Comparisons | RYUSLOG DEV