Fuzzy matching decides whether two text values are close enough to treat as the same thing. The table below shows what that looks like on real data: names, addresses, product titles, search queries, and company names, with the similarity score and the accept/reject decision for each pair.
The scores assume a simple Levenshtein-based similarity with a threshold of 0.85. Real systems tune the comparator, preprocessing, and threshold to the specific field, so these numbers are a representative baseline, not a universal rule.
| String A | String B | Scenario | Similarity | Verdict |
|---|---|---|---|---|
Jonathan Smith |
Jonathon Smith |
Name typo: one character differs | 0.93 | ✅ Accept |
123 Main Street |
123 Main St |
Address: abbreviation expands to full word | 0.92 | ✅ Accept |
iPhone 14 Pro Max 256GB |
iPhone 14 Pro Max 256 GB |
Product title: extra whitespace | 0.96 | ✅ Accept |
pyhton list |
python list |
Search query: transposed letters in python |
0.82 | ❌ Reject* |
Acme Corp. |
Acme Corporation |
Company name: abbreviation vs full word | 0.94 | ✅ Accept |
John Smith |
Jane Smith |
Different person: first name differs meaningfully | 0.70 | ❌ Reject |
*pyhton vs python scores higher under Damerau-Levenshtein (0.91) because it counts the adjacent letter swap as one operation instead of two. The choice of comparator always changes borderline cases.
The first four pairs produce a high enough score to pass the threshold because the differences are local and small: a single substitution, a missing suffix, an extra space, or a transposition.
The Acme pair works well after removing the period from Corp. and stemming the suffix, because Corp and Corporation share most of their characters with only a suffix addition.
The John Smith / Jane Smith pair fails because three of the first four characters differ, and the edit cost is large relative to the short length of the string.
How One Score Is Produced
Take Jonathan Smith and Jonathon Smith.
After lowercasing and removing spaces, the two strings are jonathansmith and jonathonsmith, each 14 characters long.
The difference is one character: an a in the first string where the second has an o, at position 5.
A Levenshtein matcher considers every possible sequence of edits and finds the cheapest route.
Substituting a with o costs one edit.
Everything else matches, so the total distance is 1.
The normalized similarity formula turns that distance into a 0-to-1 score:
At a threshold of 0.85, the pair passes easily. One small edit in a 14-character string is a tiny fraction of the whole, and the similarity score reflects that.
Contrast this with John Smith vs Jane Smith.
The normalized strings johnsmith and janesmith are both 10 characters long.
The first four characters john must be transformed into jane, which requires three substitutions (o→a, h→n, n→e).
Three edits in a 10-character string give a similarity of , which lands below the threshold.
The same absolute edit count is more meaningful in a shorter field, and the score captures that naturally.
The interactive matrix below makes this process visible by showing the cheapest repair route for any pair of strings. Edit the source and target boxes to try your own examples, or click the example buttons to see how well-known pairs break down.
Each cell in the matrix represents the cheapest repair cost between two prefixes. The highlighted route is the sequence of operations with the lowest total cost: diagonal moves for matches and substitutions, horizontal moves to delete a character, vertical moves to insert one. The right-hand panel lists each repair so you can see exactly which characters changed, were removed, or were added.
Try replacing one of the preset examples with John as the source and Jane as the target.
Watch the repair sequence fill with three substitution operations at the start of the string, and notice how the final distance of 3 produces a lower similarity score than the single-edit name example above.
Preprocessing Changes the Game
Several of the examples in the table would fail without preprocessing.
The address pair 123 Main Street and 123 Main St differs by five characters (reet), but those five characters are a standard abbreviation suffix.
Lowercasing the strings and expanding St to Street or stripping the suffix entirely makes the difference invisible to the comparator.
Similarly, Acme Corp. and Acme Corporation need the period removed and the suffix handled.
Without that cleanup, the raw edit distance between acme corp. and acme corporation is large enough to push the similarity below 0.85.
With it, the match passes.
These preprocessing choices are not automatic cleanup. Removing periods might be correct for company suffixes but wrong for version numbers or domain names. Expanding abbreviations too aggressively can create false matches when a short form has multiple meanings. The full fuzzy matching tutorial covers preprocessing and threshold tuning in depth, including how to decide which normalizations are safe for your data and which risk hiding meaningful differences.
Choosing the Right Comparator for Each Example
The table above uses Levenshtein similarity because it is the most explainable baseline, but it is not always the best tool.
The pyhton / python example illustrates why: Levenshtein treats the transposed ht as two separate substitutions, producing a score below the threshold, while a comparator that allows adjacent swaps (such as Damerau-Levenshtein) scores the pair above it.
For search queries where transposition is a common typing error, the extra operation can mean the difference between showing a result and hiding it.
The Acme Corp. / Acme Corporation pair hints at a different need.
Even after removing the period, character-level comparators must account for the extra oration suffix.
A token-based comparator that splits on whitespace and compares word overlap would see acme matching acme and corp nearly matching corporation, producing a confident score regardless of exact character alignment.
For a survey of when to use each comparator, see Fuzzy String Matching Algorithms Explained. For the core concepts behind edit distance, comparators, and thresholds with the interactive repair matrix and the three-lane comparator race, the interactive fuzzy matching tutorial covers the full learning path.