Python __eq__ vs __hash__: The Contract You Must Follow
Understand how Python's __eq__ and __hash__ methods interact, why they must stay consistent, and how to implement them correctly for set membership and dictionary keys.
Python's __eq__ and __hash__ methods control how objects compare and whether they can be used as dictionary keys or set members. This article explains the contract between them, what happens when you define only one, and how to implement both correctly.
The Contract Between __eq__ and __hash__
The core rule is simple: if two objects compare as equal using __eq__, they must return the same value from __hash__. This invariant is what allows sets and dictionaries to locate objects efficiently. When you insert an object into a set, Python computes its hash to find the bucket, then uses __eq__ to resolve collisions. If equal objects had different hashes, they would be placed in different buckets, and membership checks would fail.
Python maintains this invariant for built-in types, but for custom classes you are responsible for it. The default __hash__ is based on id(), and the default __eq__ also uses identity, so they are consistent. As soon as you override __eq__, Python assumes equality semantics are changing and disables the default hash by setting __hash__ = None unless you explicitly define it.
What Happens When You Define Only __eq__
Consider a class that defines __eq__ but not __hash__:
class Point: def __init__(self, x, y): self.x = x self.y = y def __eq__(self, other): return isinstance(other, Point) and (self.x, self.y) == (other.x, other.y)
Instances of this class are unhashable:
p1 = Point(1, 2) p2 = Point(1, 2) print(p1 == p2) # True print(hash(p1)) # TypeError: unhashable type: 'Point'
This behavior is intentional. If Python kept the default id()-based hash, two equal points would have different hashes, violating the contract. By making the object unhashable, Python forces you to decide whether you want value-based equality and, if so, to provide a matching hash.
Implementing Both Methods Correctly
To make a class hashable with value-based equality, define both __eq__ and __hash__. The hash should be derived from the same attributes used in equality. A common pattern is to hash a tuple of those attributes:
class Point: def __init__(self, x, y): self.x = x self.y = y def __eq__(self, other): return isinstance(other, Point) and (self.x, self.y) == (other.x, other.y) def __hash__(self): return hash((self.x, self.y))
Now p1 and p2 have the same hash and compare equal, so they can be used in sets and as dictionary keys:
s = {p1, p2} print(len(s)) # 1
The tuple (self.x, self.y) is immutable, and hash() of a tuple combines the hashes of its elements in a deterministic way. This approach works as long as the attributes themselves are hashable.
The Role of Immutability
A hashable object must not change its hash value after it has been inserted into a set or dictionary. If an object's attributes are mutable, and those attributes contribute to the hash, then modifying the object changes its hash, breaking the data structure's internal invariants. For example:
class MutablePoint: def __init__(self, x, y): self.x = x self.y = y def __eq__(self, other): return isinstance(other, MutablePoint) and (self.x, self.y) == (other.x, other.y) def __hash__(self): return hash((self.x, self.y)) p = MutablePoint(1, 2) s = {p} p.x = 3 print(p in s) # False, because p's hash changed after insertion
After mutating p.x, its hash changes, so the set no longer finds it in the original bucket. This is why hashable objects are typically immutable. If you need mutable objects with value equality, keep them out of hash-based collections. If your data is conceptually immutable, a frozen dataclass is a convenient way to get value equality and a correct hash.
Common Mistakes That Break Hashing
A frequent error is using a mutable attribute in the hash function, such as a list. For instance:
class BadHash: def __init__(self, items): self.items = items # list is mutable def __eq__(self, other): return isinstance(other, BadHash) and self.items == other.items def __hash__(self): return hash(tuple(self.items))
The problem here is that self.items is mutable, not the conversion of a list of hashable items to a tuple inside __hash__. If the list changes after the object is inserted into a set, the hash changes and the object is no longer found in its original bucket. Store an immutable tuple from the start if the object needs stable hashing.
Another common mistake is defining __eq__ across unrelated types while using a hash that is not aligned with that equality. For example, if a Point compares equal to a tuple with the same coordinates, the hash must also be the same. If the hash includes the type, equal objects would have different hashes and violate the contract. It is acceptable for unequal objects to share a hash, but that creates extra collisions.
Performance Impact of a Poor Hash Function
The quality of your __hash__ implementation directly affects the performance of sets and dictionaries. A hash function that returns the same constant for all objects, such as return 0, is valid but causes every object to land in the same bucket. This turns set lookups from O(1) into O(n) in the worst case, because every insertion and lookup must compare against all existing objects via __eq__. While the contract is satisfied, the performance degrades significantly for large collections.
A good hash function should distribute objects uniformly across the hash space. Using hash() on a tuple of immutable attributes usually achieves this because Python's tuple hash combines element hashes with a well-tested algorithm. Avoid inventing your own hash arithmetic unless you have a specific reason; the built-in hash() is fast and well-distributed.
Note that Python's hash() for strings is salted per process for security reasons. That does not affect the contract, as long as you call hash() rather than hardcoding numeric hash values.
Using dataclass and NamedTuple for Automatic Hashing
Python's dataclasses module can generate __eq__ and __hash__ for you. By default, a dataclass has eq=True and frozen=False, so Python sets __hash__ to None, making instances unhashable. To get a safe hash, set frozen=True, which makes the dataclass immutable and generates a hash based on its fields:
from dataclasses import dataclass @dataclass(frozen=True) class Point: x: int y: int
Now Point instances are hashable and compare by value. The generated __hash__ combines the hashes of the fields in a deterministic order. (unsafe_hash=True also generates a hash for a mutable dataclass, but it is easy to break if fields change after the object is used in a set or dictionary.)
Similarly, NamedTuple subclasses are immutable and hashable by default, with equality based on field values. If you need a lightweight data container, a NamedTuple is often simpler than a custom class.
When to Deliberately Set __hash__ = None
There are cases where you want value equality but do not want objects to be usable as dictionary keys or set members. For example, a mutable object compared by value but whose state changes frequently should stay unhashable. Setting __hash__ = None explicitly makes the object unhashable while keeping __eq__:
class MutableRecord: def __init__(self, value): self.value = value def __eq__(self, other): return isinstance(other, MutableRecord) and self.value == other.value __hash__ = None
This is the same behavior Python applies automatically when you define __eq__ without __hash__. Being explicit can improve readability, especially when the class is part of a public API. You might also do this to prevent accidental use in sets when you know an object's equality is not stable over time.
In summary, the decision hinges on whether your objects are immutable and whether you need them in hash-based collections. If you need value equality and hashability, implement both methods consistently or use a frozen=True dataclass. If mutability is required, keep the object unhashable and rely on other data structures like lists for membership checks.