How to Remove Duplicate Lines From a List (5 Easy Ways)

Handwritten to-do list in a spiral notebook with a pen
Photo: “Todo List” by Freestocks.org · CC0 1.0 · via Openverse (stocksnap)

Short answer

Paste the list into an online duplicate remover and copy the cleaned result, or use the built-in tools: Data › Remove Duplicates in Excel, =UNIQUE(A2:A) in Google Sheets, Edit › Line Operations › Remove Duplicate Lines in Notepad++, or sort -u on the command line. Decide first whether “Apple” and “apple ” should count as the same line.

Duplicate lines creep into email lists, keyword lists, product codes, survey answers and log files whenever data is merged from several sources. Removing them by hand is tedious and easy to get wrong. Here are five quick methods, from a one-click online tool to spreadsheet and command-line options, plus the settings that decide what counts as a duplicate.

First, decide what counts as a duplicate

Computers compare text exactly, so these lines are all different unless you tell the tool otherwise:

In most real lists, such as email addresses or names, you want these treated as the same. That means trimming spaces and ignoring case before comparing. You should also decide whether to keep the original order (first occurrence wins) or sort the list, and whether to remove empty lines.

Method 1: Online duplicate line remover

  1. Open the remove duplicate lines tool.
  2. Paste your list, one item per line.
  3. Choose the options: trim spaces, ignore case, remove empty lines, and sort A–Z or Z–A (or keep the original order).
  4. Copy the cleaned list.

The tool shows how many lines were removed, and everything runs in your browser, so private lists such as customer emails are not uploaded. This is the quickest option for lists you copy from emails, documents or web pages.

Method 2: Excel

  1. Select the column or range.
  2. Go to Data › Remove Duplicates (in the Data Tools group).
  3. Tick the columns that must match for a row to count as a duplicate, then click OK.

Excel deletes the duplicate rows and tells you how many were removed. It keeps the first occurrence. Because this permanently changes your data, work on a copy. To see duplicates before deleting them, use Home › Conditional Formatting › Highlight Cells Rules › Duplicate Values. In newer versions of Excel, =UNIQUE(A2:A100) returns a de-duplicated list in a new place without touching the original.

Excel’s comparison ignores case, but not extra spaces. Clean those first with =TRIM(A2).

Method 3: Google Sheets

Sheets also has Data cleanup › Trim whitespace to remove stray spaces first.

Method 4: Notepad++ and code editors

In Notepad++, choose Edit › Line Operations › Remove Duplicate Lines. There is also “Remove Consecutive Duplicate Lines”, which only removes repeats that sit next to each other, so sort first if you use that option. In Visual Studio Code, you can sort lines with the “Sort Lines Ascending” command and then remove adjacent duplicates, or use an extension.

Method 5: The command line

On macOS, Linux or Windows Subsystem for Linux:

CommandWhat it does
sort -u list.txtSorts and removes duplicates
sort list.txt | uniqSame result; uniq only removes adjacent repeats, so sort first
sort -f -u list.txtIgnores case when sorting and comparing
awk '!seen[$0]++' list.txtRemoves duplicates and keeps the original order
sort list.txt | uniq -c | sort -rnCounts how often each line appears, most common first

Add > clean.txt to the end of a command to save the result in a new file.

Worked example

Input:

banana
Apple
apple

cherry
banana

With trim spaces, ignore case and remove empty lines turned on, and the original order kept, the output is banana, Apple, cherry: three lines instead of six. Without trimming and ignoring case, “Apple” and “apple ” would both survive.

Tips for email and contact lists

Email lists are the most common reason people remove duplicates, and they have a few quirks. The domain part of an address is not case-sensitive, and in practice most providers treat the whole address that way, so ignoring case is usually safe. Watch for copies of the same address with different labels, such as “Jane Doe <[email protected]>” and “[email protected]”: strip the names first so only the addresses remain. Finally, keep a record of people who unsubscribed, so that merging lists does not add them back by accident.

After removing duplicates

Common mistakes

Frequently asked questions

How do I remove duplicate lines online?

Paste your list into an online duplicate remover, choose whether to trim spaces and ignore case, and copy the cleaned result.

How do I remove duplicates in Excel but keep the first one?

Data › Remove Duplicates keeps the first occurrence and deletes later repeats. Work on a copy because the change is permanent.

How do I remove duplicates without changing the order?

Use a tool that keeps the original order, the UNIQUE function in Google Sheets or Excel, or awk '!seen[$0]++' on the command line.

Why are some duplicates not being removed?

They probably differ by a space or capital letter. Turn on trimming and ignore case, or clean the data with TRIM first.

What is the difference between sort -u and uniq?

sort -u sorts and removes all duplicates. uniq only removes repeated lines that are next to each other, so it is usually used after sort.

Sources