TL;DR: "Duplicate" lines aren't always identical bytes — trailing spaces, different casing or Unicode normalization can hide exact duplicates or create false ones. Knowing how comparison works keeps your cleanup accurate instead of accidentally deleting rows you needed. Pair this guide with Remove Duplicate Lines, Text Difference, Word Frequency and Remove Extra Spaces.
Duplicate Finder on CharCount removes repeated lines from lists, logs, CSV exports and keyword lists. The tricky part isn't the removal — it's deciding what counts as "the same line" before you remove anything.
Background: MDN's JavaScript Set reference explains the data structure most duplicate-removal tools use internally, MDN's String.normalize() guide covers Unicode normalization, and the Unicode text segmentation report (UAX #29) defines how text is split into meaningful units in the first place.

How duplicate detection actually works
Under the hood, a duplicate-line tool walks each line, and checks whether it has already been seen — typically using a Set, which stores unique values and rejects exact repeats in constant time. That's fast, but "exact" is doing a lot of work in that sentence: a Set compares strings byte-for-byte, so "Apple" and "apple" are different entries, and so are "text " and "text" (trailing space).
This is why a naive dedupe on a CSV export or a scraped keyword list often "fails" — it's not broken, it's comparing exactly what you gave it, including invisible whitespace and case differences you didn't notice.
Order-preserving vs. sort-and-dedupe
There are two common strategies. Order-preserving dedupe keeps the first occurrence of each line and its original position — the right choice for logs, chat exports or anything where sequence carries meaning. Sort-and-dedupe (like the classic Unix sort | uniq pipeline) reorders everything alphabetically first, which is faster on huge files but destroys the original order — fine for a keyword list, wrong for a changelog.
Pick order-preserving whenever "what came first" matters; pick sort-and-dedupe only when the list's order was arbitrary to begin with.
Where normalization changes the count
Unicode allows the same visible character to exist as different byte sequences — an accented "é" can be one composed code point (NFC) or an "e" plus a separate combining accent mark (NFD). Two lines that look identical on screen can fail an exact-match dedupe if one came from a Mac export and the other from a Windows one, because they're in different normalization forms.
String.normalize('NFC') collapses both forms to the same representation before comparison, which is why a good duplicate finder normalizes text first — otherwise it silently under-counts duplicates that are visually, but not byte-for-byte, identical.
A safer cleanup workflow
- Decide first: does line order matter for this list? That picks your strategy.
- Trim trailing spaces and normalize case only if the list is meant to be case-insensitive (a keyword list often is; a password list never is).
- Run the dedupe, then spot-check a sample of removed lines to confirm they really were duplicates.
- Keep the original file until you've verified the cleaned version.
For anything ambiguous, run the before/after through Text Difference to see exactly what was removed, and Word Frequency to sanity-check that no legitimate repeated term got treated as noise.