Finding All Duplicates in a List in Java

Azure Spring Apps is a fully managed service from Microsoft (built in collaboration with VMware), focused on building and deploying Spring Boot applications on Azure Cloud without worrying about Kubernetes.

And, the Enterprise plan comes with some interesting features, such as commercial Spring runtime support, a 99.95% SLA and some deep discounts (up to 47%) when you are ready for production.

>> Learn more and deploy your first Spring Boot app to Azure.

You can also ask questions and leave feedback on the Azure Spring Apps GitHub page.

Slow MySQL query performance is all too common. Of course it is. A good way to go is, naturally, a dedicated profiler that actually understands the ins and outs of MySQL.

The Jet Profiler was built for MySQL only, so it can do things like real-time query performance, focus on most used tables or most frequent queries, quickly identify performance issues and basically help you optimize your queries.

Critically, it has very minimal impact on your server's performance, with most of the profiling work done separately - so it needs no server changes, agents or separate services.

Basically, you install the desktop application, connect to your MySQL server, hit the record button, and you'll have results within minutes:

>> Try out the Profiler

Accelerate Your Jakarta EE Development with Payara Server!

With best-in-class guides and documentation, Payara essentially simplifies deployment to diverse infrastructures.

Beyond that, it provides intelligent insights and actions to optimize Jakarta EE applications.

The goal is to apply an opinionated approach to get to what's essential for mission-critical applications - really solid scalability, availability, security, and long-term support:

>> Download and Explore the Guide (to learn more)

The AI Assistant to boost Boost your productivity writing unit tests - Machinet AI.

AI is all the rage these days, but for very good reason. The highly practical coding companion, you'll get the power of AI-assisted coding and automated unit test generation.
Machinet's Unit Test AI Agent utilizes your own project context to create meaningful unit tests that intelligently aligns with the behavior of the code.
And, the AI Chat crafts code and fixes errors with ease, like a helpful sidekick.

Simplify Your Coding Journey with Machinet AI:

>> Install Machinet AI in your IntelliJ

Looking for the ideal Linux distro for running modern Spring apps in the cloud?

Meet Alpaquita Linux: lightweight, secure, and powerful enough to handle heavy workloads.

This distro is specifically designed for running Java apps. It builds upon Alpine and features significant enhancements to excel in high-density container environments while meeting enterprise-grade security standards.

Specifically, the container image size is ~30% smaller than standard options, and it consumes up to 30% less RAM:

>> Try Alpaquita Containers now.

DbSchema is a super-flexible database designer, which can take you from designing the DB with your team all the way to safely deploying the schema.

The way it does all of that is by using a design model, a database-independent image of the schema, which can be shared in a team using GIT and compared or deployed on to any database.

And, of course, it can be heavily visual, allowing you to interact with the database using diagrams, visually compose queries, explore the data, generate random data, import data or build HTML5 database reports.

>> Take a look at DBSchema

Slow MySQL query performance is all too common. Of course it is. A good way to go is, naturally, a dedicated profiler that actually understands the ins and outs of MySQL.

Critically, it has very minimal impact on your server's performance, with most of the profiling work done separately - so it needs no server changes, agents or separate services.

Basically, you install the desktop application, connect to your MySQL server, hit the record button, and you'll have results within minutes:

>> Try out the Profiler

1. Introduction

In this article, we’ll learn different approaches to finding duplicates in a List in Java.

Given a list of integers with duplicate elements, we’ll be finding the duplicate elements in it. For example, given the input list [1, 2, 3, 3, 4, 4, 5], the output List will be [3, 4].

2. Finding Duplicates Using Collections

In this section, we’ll discuss two ways of using Collections to extract duplicate elements present in a list.

2.1. Using the contains() Method of Set

Set in Java doesn’t contain duplicates. The contains() method in Set returns true only if the element is already present in it.

We’ll add elements to the Set if contains() returns false. Otherwise, we’ll add the element to the output list. The output list thus contains the duplicate elements:

List<Integer> listDuplicateUsingSet(List<Integer> list) {
    List<Integer> duplicates = new ArrayList<>();
    Set<Integer> set = new HashSet<>();
    for (Integer i : list) {
        if (set.contains(i)) {
            duplicates.add(i);
        } else {
            set.add(i);
        }
    }
    return duplicates;
}

Let’s write a test to check if the list duplicates contains only the duplicate elements:

@Test
void givenList_whenUsingSet_thenReturnDuplicateElements() {
    List<Integer> list = Arrays.asList(1, 2, 3, 3, 4, 4, 5);
    List<Integer> duplicates = listDuplicate.listDuplicateUsingSet(list);
    Assert.assertEquals(duplicates.size(), 2);
    Assert.assertEquals(duplicates.contains(3), true);
    Assert.assertEquals(duplicates.contains(4), true);
    Assert.assertEquals(duplicates.contains(1), false);
}

Here we see that the output list only contains two elements 3 and 4.

This approach takes O(n) time for n elements in a list and extra space of size n for the set.

2.2. Using a Map and Storing the Frequency of Elements

We can use a Map to store the frequency of each element and then add them to the output list only when the frequency of the element isn’t 1:

List<Integer> listDuplicateUsingMap(List<Integer> list) {
    List<Integer> duplicates = new ArrayList<>();
    Map<Integer, Integer> frequencyMap = new HashMap<>();
    for (Integer number : list) {
        frequencyMap.put(number, frequencyMap.getOrDefault(number, 0) + 1);
    }
    for (int number : frequencyMap.keySet()) {
        if (frequencyMap.get(number) != 1) {
            duplicates.add(number);
        }
    }
    return duplicates;
}

Let’s write a test to check if the list duplicates contain only duplicate elements:

@Test
void givenList_whenUsingFrequencyMap_thenReturnDuplicateElements() {
    List<Integer> list = Arrays.asList(1, 2, 3, 3, 4, 4, 5);
    List<Integer> duplicates = listDuplicate.listDuplicateUsingMap(list);
    Assert.assertEquals(duplicates.size(), 2);
    Assert.assertEquals(duplicates.contains(3), true);
    Assert.assertEquals(duplicates.contains(4), true);
    Assert.assertEquals(duplicates.contains(1), false);
}

Here we see that the output list contains only two elements, 3 and 4.

This approach takes O(n) time for n elements in a list and extra space of size n for the map.

3. Using Streams in Java 8

In this section, we’ll discuss three ways of using Streams to extract duplicate elements present in a list.

3.1. Using filter() and Set.add() Method

Set.add() adds the specified element to this set if it’s not already present. If this set already contains the element, the call leaves the set unchanged and returns false.

Here, we’ll use a Set and convert the list to a stream. The stream is added to the Set, and the duplicate elements are filtered and collected into List:

List<Integer> listDuplicateUsingFilterAndSetAdd(List<Integer> list) {
    Set<Integer> elements = new HashSet<Integer>();
    return list.stream()
      .filter(n -> !elements.add(n))
      .collect(Collectors.toList());
}

Let’s write a test to check if the list duplicates contain only duplicate elements:

@Test
void givenList_whenUsingFilterAndSetAdd_thenReturnDuplicateElements() {
    List<Integer> list = Arrays.asList(1, 2, 3, 3, 4, 4, 5);
    List<Integer> duplicates = listDuplicate.listDuplicateUsingFilterAndSetAdd(list);
    Assert.assertEquals(duplicates.size(), 2);
    Assert.assertEquals(duplicates.contains(3), true);
    Assert.assertEquals(duplicates.contains(4), true);
    Assert.assertEquals(duplicates.contains(1), false);
}

Here we see that the output elements contain only two elements, 3 and 4, as expected.

This approach using filter() with Set.add() is the fastest algorithm to find duplicate elements with O(n) time complexity and extra space of size n for the set.

3.2. Using Collections.frequency()

Collections.frequency() returns the number of elements in the specified collection, which is equal to a specified value. Here we’ll convert List to Stream and filter out only the elements that return a value greater than one from Collections.frequency().

We’ll collect these elements into Set to avoid repetitions and finally convert Set to List:

List<Integer> listDuplicateUsingCollectionsFrequency(List<Integer> list) {
    List<Integer> duplicates = new ArrayList<>();
    Set<Integer> set = list.stream()
      .filter(i -> Collections.frequency(list, i) > 1)
      .collect(Collectors.toSet());
    duplicates.addAll(set);
    return duplicates;
}

Let’s write a test to check if the duplicates contain only duplicate elements:

@Test
void givenList_whenUsingCollectionsFrequency_thenReturnDuplicateElements() {
    List<Integer> list = Arrays.asList(1, 2, 3, 3, 4, 4, 5);
    List<Integer> duplicates = listDuplicate.listDuplicateUsingCollectionsFrequency(list);
    Assert.assertEquals(duplicates.size(), 2);
    Assert.assertEquals(duplicates.contains(3), true);
    Assert.assertEquals(duplicates.contains(4), true);
    Assert.assertEquals(duplicates.contains(1), false);
}

As expected, the output list contains only two elements, 3 and 4.

This approach using Collections.frequency() is the slowest because it compares each element with a list – Collections.frequency(list, i) whose complexity is O(n). So the overall complexity is O(n*n). It also requires an extra space of size n for the set.

3.3. Using Map and Collectors.groupingBy()

Collectors.groupingBy() returns a collector implementing a cascaded “group by” operation on input elements.

It groups elements according to a classification function and then performs a reduction operation on the associated values with a given key using the specified downstream collector. The classification function maps elements to some key type K. The downstream collector operates on input elements and produces a result of type D. The resulting collector produces a Map<K, D>.

Here we’ll use Function.identity() as the classification function and Collectors.counting() as the downstream collector.

Function.identity() returns a function that always returns its input argument. Collectors.counting() returns a collector accepting elements that count the number of input elements. If no elements are present, the result is zero. Thus we’ll get a map of elements and their frequency using Collectors.groupingBy().

Then we convert the EntrySet of this Map into a Stream, filter out only the elements that have a value greater than 1, and collect them in a Set to avoid repetitions. Then the Set is converted into a List:

List<Integer> listDuplicateUsingMapAndCollectorsGroupingBy(List<Integer> list) {
    List<Integer> duplicates = new ArrayList<>();
    Set<Integer> set = list.stream()
      .collect(Collectors.groupingBy(Function.identity(), Collectors.counting()))
      .entrySet()
      .stream()
      .filter(m -> m.getValue() > 1)
      .map(Map.Entry::getKey)
      .collect(Collectors.toSet());
    duplicates.addAll(set);
    return duplicates;
}

Let’s write a test to check if the list duplicates contains only duplicate elements:

@Test
void givenList_whenUsingMapAndCollectorsGroupingBy_thenReturnDuplicateElements() {
    List<Integer> list = Arrays.asList(1, 2, 3, 3, 4, 4, 5);
    List<Integer> duplicates = listDuplicate.listDuplicateUsingCollectionsFrequency(list);
    Assert.assertEquals(duplicates.size(), 2);
    Assert.assertEquals(duplicates.contains(3), true);
    Assert.assertEquals(duplicates.contains(4), true);
    Assert.assertEquals(duplicates.contains(1), false);
}

Here we see that the output elements contain only two elements 3 and 4.

Collectors.groupingBy() takes O(n) time. A filter() operation is done on the resulting EntrySet, but the complexity remains O(n) as the map lookup time is O(1). It also requires an extra space of n for the set.

4. Conclusion

In this article, we learned about different ways of extracting duplicate elements from a List in Java.

We discussed approaches using Set and Map and their corresponding approaches using Stream. The code using Stream is far more declarative and conveys the intent of the code clearly without the need of external iterators.

As always, the complete code samples for this article can be found over on GitHub.

Finding All Duplicates in a List in Java

Get started with Spring and Spring Boot, through the Learn Spring course:

1. Introduction

2. Finding Duplicates Using Collections

2.1. Using the contains() Method of Set

2.2. Using a Map and Storing the Frequency of Elements

3. Using Streams in Java 8

3.1. Using filter() and Set.add() Method

3.2. Using Collections.frequency()

3.3. Using Map and Collectors.groupingBy()

4. Conclusion

Get started with Spring and Spring Boot, through the Learn Spring course:

REST with Spring

Learn Spring Security ▼▲

Learn Spring Security Core

Learn Spring Security OAuth

Learn Spring

Learn Spring Data JPA

Persistence

REST

Security

Full Archive

Baeldung Ebooks

About Baeldung

Write for Baeldung

Get started with Spring and Spring Boot, through the Learn Spring course:

1. Introduction

2. Finding Duplicates Using Collections

2.1. Using the contains() Method of Set

2.2. Using a Map and Storing the Frequency of Elements

3. Using Streams in Java 8

3.1. Using filter() and Set.add() Method

3.2. Using Collections.frequency()

3.3. Using Map and Collectors.groupingBy()

4. Conclusion

Get started with Spring and Spring Boot, through the Learn Spring course: