Python Contains: Membership Testing with the in Operator
Learn how Python's in operator checks membership across lists, strings, sets, and dictionaries, including performance tradeoffs and common pitfalls.
Python's in operator is the standard way to test membership. The expression value in container returns True if the value is present and False otherwise. This syntax works with lists, tuples, strings, sets, dictionaries, and any object that supports containment through __contains__ or iteration. How in behaves for each data structure matters for both correctness and performance.
The in Operator for Lists and Tuples
For lists and tuples, in performs a linear scan from the first element until it finds a match or reaches the end. The cost grows with the size of the collection. For a small list this is negligible, but for a list with thousands of elements, repeated membership checks can become a bottleneck.
fruits = ['apple', 'banana', 'cherry'] print('banana' in fruits) # True print('grape' in fruits) # False
The operator uses equality comparison (==) to determine whether an element matches. This matters when the list contains objects with custom equality logic: in will invoke __eq__ for each element until a match is found.
Tuples use the same membership behavior as lists: the check is a linear scan. The only structural difference is that tuples are immutable, but membership testing still walks every element until it finds a match.
Checking Substrings with in on Strings
When the container is a string, in checks substring membership rather than character equality. The expression substring in string returns True if the substring appears anywhere within the target string.
message = 'hello world' print('world' in message) # True print('xyz' in message) # False
This behavior is consistent with string methods like find() and index(), but in returns a boolean directly, making it the most readable choice for simple presence checks. The search cost depends on the lengths of both the substring and the target string.
For case-insensitive checks, normalize the strings first, for example by calling .lower() on both sides. The in operator does not perform case folding by itself.
Membership in Sets and Dictionaries
Sets and dictionary key lookups use hash-based storage, so in has an average-case time complexity of O(1) for those containers. When you write value in some_set, Python hashes the value and probes the underlying hash table. This makes membership testing on large sets dramatically faster than on lists.
unique_ids = {101, 202, 303} print(202 in unique_ids) # True
For dictionaries, in checks keys only, not values. This is a common point of confusion. To check whether a value appears in a dictionary, use value in config.values(); that is a linear scan over the values.
config = {'host': 'localhost', 'port': 8080} print('host' in config) # True print('localhost' in config) # False print('localhost' in config.values()) # True
Set elements and dictionary keys must be hashable, because both rely on hashing. Dictionary values do not need to be hashable. Mutable containers like lists and dictionaries are not hashable and cannot be used as set elements or dictionary keys. If you write some_list in some_set, Python raises a TypeError when it tries to hash the list.
How in Behaves with Custom Classes
Any class can define __contains__ to control how in works for its instances. When you write x in obj, Python calls obj.__contains__(x) if the method is defined. If it is not defined, Python falls back to iterating over the object's items and comparing each one for equality.
class Playlist: def __init__(self, songs): self.songs = songs def __contains__(self, title): return any(s.title == title for s in self.songs)
Now 'Song Title' in playlist returns True when any song in self.songs has that title. Membership is determined by song titles, not by comparing entire song objects. Defining __contains__ gives you precise control over what in means for your domain objects. Without it, Python would try to iterate over the instance, which may not be what you intend.
When implementing __contains__, return a boolean. Returning a truthy or falsy value works, but explicit True or False is clearer. Also note that __contains__ runs on every in check, so expensive work in the method is repeated for each check.
Performance Characteristics of in by Data Structure
The performance of in depends on the underlying data structure. Lists and tuples require a linear scan, so the worst-case cost is O(n). Sets and dictionary key lookups use hashing, which gives an average-case cost of O(1), but computing the hash adds a constant overhead. String membership uses a substring search; the cost depends on the lengths of the substring and the target string, and optimized implementations are much faster in practice than naive comparisons.
Choosing the right container is a common optimization. If you need to check membership frequently and the collection is large, converting a list to a set once and then using in on that set is usually worthwhile. The conversion costs O(n), but each later check is O(1).
# Repeated membership checks on a list items = [1, 2, 3, 4, 5] for x in range(1000): if x in items: pass # Convert to a set once when checks are frequent item_set = set(items) for x in range(1000): if x in item_set: pass
This pattern is most useful when the collection is large and the number of checks is high. For small collections, the cost of building a set may not be justified, so base the decision on expected size and check frequency.
Common Pitfalls and Edge Cases
One subtle issue involves floating-point values. The in operator uses equality, and float('nan') does not equal itself. As a result, float('nan') in [float('nan')] returns False. This follows from IEEE 754 semantics and Python's equality behavior.
Another pitfall is checking membership in a dictionary when you actually need to check values. As noted earlier, in on a dictionary only tests keys. Use value in dict.values() to test values; that is a linear scan and therefore O(n). If value lookups are frequent and the values are hashable, consider maintaining a separate set of values.
When working with custom objects, in relies on __eq__ unless you override __contains__. If two objects compare equal but have different identities, in still finds a match. Conversely, if a class does not define __eq__, object identity is used, so two distinct objects with the same attributes are not considered equal.
Finally, be aware that in on a generator or iterator advances it. The check consumes items until it finds a match or exhausts the iterator. If you need to preserve the values for later use, materialize them in a list or tuple first.
Using in correctly means understanding the data structure you are working with and the semantics of equality and hashing in Python. For most presence checks, value in collection is the clearest and most idiomatic choice.