A list with repeats is easy to clean until the repeats are not exact. "Apple", "apple", and "apple " with a space at the end look like one item to a person and three to a program. A name pasted from a web page may have a non-breaking space, and a letter with an accent may be stored as two characters. This page removes the repeats from a list and shows you the ones that look the same but are written differently, so you decide what counts as a repeat.
Paste the list, choose how it is separated, and choose what to do with repeats: keep the first or the last copy, keep only what appears once, keep only what repeats, or count each item. The page also writes the same steps as code. If you came here to remove duplicates from a list in Python, the code box gives you the exact lines for the choices you made. Everything is processed in your browser, nothing is uploaded, and the code behind the page makes no network requests.
How to use the Remove Duplicates from List: Lines, Words, and Items
- Paste your listPaste the items in the box, or open a text file. Choose how the items are separated: new lines, commas, semicolons, tabs, pipes, spaces, or a separator you type.
- Choose what to doKeep the first or the last copy of each item, keep only items that appear once, keep only items that repeat, or count how often each appears. You can sort the result.
- Decide what counts as the sameTrim spaces, ignore capital letters or accents, remove invisible characters, or join accents. If some items look the same but are written differently, the page tells you how many more would be merged by each choice.
- Check the resultThe tiles show how many items there were, how many are different, and how many were removed. Open the list of merged groups to see each spelling and why they were counted as one.
- Copy it or take the codeCopy the result, download it, or use it as the input for another pass. The code box shows the same steps in Python, JavaScript, shell, SQL, and Excel.
What counts as a duplicate
A program compares text character by character, so two items are the same only if every character is the same. Several kinds of difference are invisible. A space at the end of a line, a tab instead of a space, and a non-breaking space that looks like a normal one all make an item different. A zero-width space, a soft hyphen, and similar characters take no room at all. The page lets you trim spaces, turn runs of spaces into one, and remove invisible characters, which changes the items you keep.
Other differences are about letters. Capital letters make "Apple" and "apple" different, and the page can ignore them. It folds case the way Python does. The Python documentation says casefolding is similar to lowercasing but more aggressive, because it removes all case distinctions, and its example is that the German ß is equivalent to ss, which lower() would not change. An accent can also be ignored, so "café" and "cafe" are one item. When the page ignores case or accents, it keeps the spelling of the first copy and does not rewrite it.
The same character in two forms
Unicode can write one letter in more than one way. The Unicode normalization report calls characters canonically equivalent when they represent the same abstract character and should always look and behave the same. Its example is é, which can be one character or an e followed by a combining acute accent. The two look the same and are different to a program, so a list that was typed on two systems can hold both. The option to join a letter and its separate accent turns them into one, using the form called NFC.
A weaker kind is compatibility equivalence, where characters represent the same thing but may look or behave differently. The report's examples are the ligature fi, which is compatible with the two letters f and i, and the full-width forms used in East Asian text. The option to treat look-alike forms as the same uses the form called NFKC for the comparison only, and the item you keep is not changed.
Which copy to keep
Removing repeats usually means keeping the first copy of each item, in the order they first appear. The page can also keep the last copy, which gives the order of the last appearances. Two other choices answer different questions. Keeping only items that appear once drops every item that has a repeat, including the first copy, which finds what is unique. Keeping only items that repeat gives one copy of each item that appears more than once, which finds the duplicates. Counting gives each item with the number of times it appears, most frequent first.
Lists with commas, and quotes
A list separated by commas is harder than a list of lines, because an item can hold a comma. In the usual convention for such files, an item is put in double quotes when it holds the separator, and a quote inside is written twice. The page reads a comma, semicolon, tab, or pipe list that way, so "Smith, J" stays one item, and it puts the quotes back when it writes the result. Spaces after a separator are not part of the item. A line break always ends an item, so you can paste several lines of comma lists at once.
Removing duplicates from a list in Python
The shortest answer in Python is list(dict.fromkeys(items)). A dictionary cannot hold the same key twice, and what's new in Python 3.7 says the insertion-order preservation of dict objects has been declared an official part of the language specification, so the result keeps the order of first appearance. list(set(items)) also removes repeats but does not keep the order. For a list of lists, each inner list has to be turned into a tuple first, because a list cannot be a dictionary key.
When the comparison is looser, the idea is a key function: compute a key for each item, and keep the first item for each key. For a case-insensitive list the key is item.casefold(), and to ignore spaces it is item.strip() first. The page writes this for you with the choices you made, using the unicodedata module for accents and normalization, and the Counter class from collections for counts. One detail to know is that strip() in Python and trim() in JavaScript do not remove exactly the same characters, and the limitations below list them.
Excel, shell, and SQL
In Excel for Microsoft 365 and Excel 2021, the UNIQUE function returns a list of unique values in a range. Its optional third argument, exactly_once, returns only the values that occur once. The page writes formulas with UNIQUE, TRIM, SORT, and COUNTIF for a list in A2:A100. They were written from Microsoft's documentation and were not run in Excel.
In a shell, awk '!seen[$0]++' prints each line the first time it is seen, which removes repeats and keeps the order. In SQL, DISTINCT or GROUP BY gives one row per value, and HAVING COUNT(*) finds values that appear once or more than once. The page writes these for the choices it can express. A database does not promise an order unless you add ORDER BY.
How the page was tested
The page was compared with code that it writes, run by other programs. 2,500 random lists were built from words, padded items, accented and decomposed letters, ß and SS, full-width letters, ligatures, Greek letters with final sigma, emoji, and invisible characters, with random choices of separator, mode, order, and options. The Python code the page writes gave the same list as the page on all 2,500, and so did the JavaScript code. Where a list was plain ASCII, the shell commands were run with bash and awk on 108 lists, and the SQL with SQLite on 21 lists. They agreed on all of them.
The case folding and the normalization were also compared with Python for every one of the 1,112,064 code points. Python's casefold differs from JavaScript's upper-then-lower-case on 174 characters, such as the Turkish dotless ı and the Cherokee letters, so the page carries a small table of exceptions generated from Python. NFC and NFD agreed on every code point. The versions of Unicode in Node and Python differ, which accounts for one NFKC difference and 42 characters that are marks in one and not the other. These are characters added in the newest Unicode version.
Limits and accuracy
- Python's strip() also removes the characters U+0085 and U+001C to U+001F, which JavaScript's trim() does not, and trim() removes U+FEFF, which strip() does not. The page trims the way JavaScript does. The Python code uses strip(), so the two differ for these rare characters.
- Sorting A to Z is by code point, the same as Python's sorted(), not by the rules of a language. Natural order and the order of accented letters follow the browser and may differ from another program.
- Ignoring case follows Python's casefold for the Unicode version of Python 3.14. Letters added to Unicode later are folded the JavaScript way.
- Quotes are read only when the separator is a comma, semicolon, tab, or pipe. With new lines, spaces, or your own separator, a quote is an ordinary character.
- The shell and SQL code is written only for the choices it can express. SQL trims only spaces, and shell tolower works for ASCII letters in most systems.
- The Excel formulas were not run in Excel.
- A list of up to 5 million characters can be processed. A larger one should be split, or handled in code.
Frequently asked questions
How do I remove duplicates from a list?
Paste the list, choose how the items are separated, and choose to remove repeats and keep the first of each. The page shows the result, how many items were removed, and any items that look the same but are written differently.
How do I remove duplicates from a list in Python?
Use list(dict.fromkeys(items)). It removes repeats and keeps the order, which Python guarantees for dictionaries since version 3.7. The page also writes code for a looser comparison, such as ignoring case or spaces, with the choices you made.
Why were my duplicates not removed?
The items probably differ in a way you cannot see: a space at the end, a non-breaking space, a different capital letter, an accent written as a separate character, or an invisible character. The page lists what each choice would merge, so you can switch it on.
Does it keep the first or the last copy?
You choose. By default it keeps the first copy of each item in the order it first appears. You can also keep the last copy, keep only items that appear once, keep only items that repeat, or count how often each item appears.
Can it remove duplicates and sort the list?
Yes. Choose A to Z, Z to A, natural order, or shortest first. A to Z orders by code point, as Python's sorted() does. The default keeps the order of the list.
Does it treat Apple and apple as the same?
Only if you tick the option to ignore capital letters. The first spelling is kept. The option folds case the way Python's casefold does, so the German ß and ss are also counted as the same.
Is my list uploaded?
No. The list is processed in your browser, and the code behind the page makes no network requests. Nothing is saved, so copy the result before you close the page.
Research and references
This page was written and checked against the sources below.

