Back to Blog
Python

Python RapidFuzz: ratio, partial_ratio, and token_sort_ratio

RapidFuzz provides three main fuzzy matching scores in Python: ratio, partial_ratio, and token_sort_ratio. See how they differ and when to use each.

RapidFuzzfuzzy matchingstring similarityPythontext matching
A visual representation of fuzzy string matching using RapidFuzz functions, showing similarity scores between text samples.

RapidFuzz is a Python library that provides fast, C-accelerated fuzzy string matching. Three commonly used scoring functions are ratio, partial_ratio, and token_sort_ratio. Each one handles a different kind of mismatch, so choosing the right function depends on how the strings you are comparing differ.

Why fuzzy matching needs more than one similarity function

Exact string comparison fails in many real-world situations: a user types a product name with a typo, a database contains extra whitespace, or a list of names has inconsistent word order. RapidFuzz helps by scoring how similar two strings are. Its ratio, partial_ratio, and token_sort_ratio functions make different tradeoffs, and the best choice depends on the type of variation in your data.

The ratio function: whole-string similarity

ratio computes a normalized edit distance between two complete strings and returns a score from 0 to 100. A score of 100 means the strings are identical. The score is based on the edits needed to turn one string into the other.

from rapidfuzz import fuzz print(fuzz.ratio('hello world', 'hello world')) # 100.0 print(fuzz.ratio('hello world', 'hello worl')) # about 95.2 print(fuzz.ratio('hello world', 'world hello')) # well below 100

The third example shows a major limitation: ratio compares characters from beginning to end, so it is sensitive to word order. Two strings with the same words but a different order get a low score, even though they are semantically similar. That is where token_sort_ratio becomes useful.

The partial_ratio function: substring matching

partial_ratio is designed for cases where one string is a truncated or extended version of the other. It compares the shorter string against the best-matching substring of the longer string.

from rapidfuzz import fuzz print(fuzz.partial_ratio('hello world', 'hello')) # 100.0 print(fuzz.partial_ratio('hello world', 'world')) # 100.0

In both examples, the shorter string appears exactly as a substring of the longer string, so the score is 100. This function is useful when you expect one string to be a fragment of the other.

The token_sort_ratio function: word-order independence

token_sort_ratio normalizes and tokenizes the strings, sorts the tokens alphabetically, and then computes ratio on the sorted token sequences. As a result, the score is insensitive to word order.

from rapidfuzz import fuzz print(fuzz.token_sort_ratio('hello world', 'world hello')) # 100.0 print(fuzz.token_sort_ratio('python rapidfuzz', 'rapidfuzz python')) # 100.0 print(fuzz.token_sort_ratio('hello world', 'hello there world')) # below 100

The first two examples return 100 because the same tokens appear in both strings, just in a different order. The third example is below 100 because there is an extra token that changes the sorted sequence.

Comparing the three functions

The following table summarizes the general score patterns. It uses 100 for a perfect score and lower for a non-perfect score.

Input pairType of variationratiopartial_ratiotoken_sort_ratio
hello world vs world helloReversed word orderlowerlower100
hello world vs helloOne string is a substringlower100lower
hello world vs hello there worldExtra word insertedlowerlowerlower

No single function fits every scenario. ratio is the baseline, partial_ratio is best when one string is a substring of the other, and token_sort_ratio is best when word order varies.

Performance considerations

The cost of these functions depends on the length of the input strings. ratio needs an edit-distance calculation, commonly O(n*m), where n and m are the string lengths. partial_ratio can be more expensive for long strings because it evaluates candidate substrings of the longer string. token_sort_ratio adds tokenization and sorting; the sorting step is O(k log k), where k is the number of tokens.

In practice, RapidFuzz is implemented in C and is faster than a pure-Python implementation of the same algorithms. If you compare short strings such as usernames or product codes, ratio is usually sufficient. If word order can vary, token_sort_ratio is a better choice. For long strings where one string may be a fragment of the other, use partial_ratio carefully and consider a minimum length for the shorter string.

Practical example: matching product names

Consider a list of product names and a user query that may contain different word order. You can score the query against each product and pick the best match.

from rapidfuzz import fuzz products = [ 'Wireless Mouse', 'USB-C Hub', 'Laptop Stand', 'Mechanical Keyboard', ] query = 'keyboard mechanical' best_match = None best_score = 0 for product in products: score = fuzz.token_sort_ratio(query, product) if score > best_score: best_match = product best_score = score print(f'Best match: {best_match} (score: {best_score})')

Here, token_sort_ratio is appropriate because the user reversed the words. ratio would give a lower score and might miss the correct product.

Edge cases and limitations

ratio and partial_ratio compare the strings as passed, so they are case-sensitive and punctuation matters. token_sort_ratio, by default, uses RapidFuzz preprocessing, which lowercases strings and removes characters that are not letters or numbers before tokenizing. Pass processor=None to token_sort_ratio if you need raw token comparison, or preprocess the strings yourself before comparing.

partial_ratio can produce misleadingly high scores when the shorter string is a very common substring. For example, comparing the to a long text can return 100 even if the is not a meaningful match. In such cases, set a minimum length for the shorter string or combine scores from multiple functions.

RapidFuzz in Python: ratio vs partial_ratio vs token_sort_ratio | RYUSLOG DEV