Python Sets: What They Are and How to Use Them

By Dr. Zubair Khalid, DVM, MS, PhD ·

Python Sets: What They Are and How to Use Them

A Python set is an unordered collection of unique elements. You create one with curly braces or the set() function, and you can combine sets with mathematical operations such as union, intersection and difference [1]. Sets are the fastest built-in way to test membership and to remove duplicate values from a collection.

Quick Answer

  • A set stores unordered, unique items. Duplicates are dropped automatically [1].
  • Create one with curly braces, as in {1, 2, 3}, or with set() from any iterable [1].
  • An empty set must be written set(). The literal {} creates an empty dictionary instead [1].
  • Sets support union |, intersection &, difference - and symmetric difference ^ [1].
  • Because sets are unordered, printing or iterating over one can give a different order than you expect [1].

What a Python Set Means

In plain terms, a set is a bag that refuses to hold the same item twice. If you add a value that is already there, nothing changes. Order does not matter, so there is no "first" or "last" element.

The precise definition from the Python documentation is that a set is an unordered collection with no duplicate elements [1]. Its basic uses include membership testing and eliminating duplicate entries, and set objects also support mathematical operations like union, intersection, difference and symmetric difference [1].

That combination of properties is what makes sets useful in data work. When you care about which values appear and not how many times or in what order, a set is the right structure. If you need order or counts, a Python list or another sequence type is a better fit.

How It Works

A set is built on a hash table. Each element is hashed to a location, which is why membership tests are fast and why elements must be hashable (immutable types such as numbers, strings and tuples of immutables). Lists and dictionaries cannot go inside a set.

The four core operations follow set theory. For two sets $A$ and $B$:

$$A \cup B = \{x : x \in A \text{ or } x \in B\}$$

$$A \cap B = \{x : x \in A \text{ and } x \in B\}$$

$$A - B = \{x : x \in A \text{ and } x \notin B\}$$

$$A \triangle B = \{x : x \in A \text{ or } x \in B, \text{ but not both}\}$$

Each symbol maps to a Python operator:

SymbolMeaningOperatorMethod
$\cup$Union, everything in either set`\`.union()
$\cap$Intersection, only what both share&.intersection()
$-$Difference, in the first but not the second-.difference()
$\triangle$Symmetric difference, in exactly one set^.symmetric_difference()

The operator forms require both operands to be sets. The method forms accept any iterable, so week1.union([999]) works even though a list is not a set.

Worked Example

The dataset is a list of survey respondent IDs collected in two consecutive weeks. Week 1 has 8 IDs and week 2 has 8 IDs, with some respondents appearing in both weeks.

weekrespondent_ids
week1101, 102, 103, 104, 105, 106, 107, 108
week2105, 106, 107, 108, 109, 110, 111, 112

Here is the code and its output.

week1 = {101, 102, 103, 104, 105, 106, 107, 108}
week2 = {105, 106, 107, 108, 109, 110, 111, 112}
print(week1 | week2)  # union
print(week1 & week2)  # intersection
print(week1 - week2)  # difference
{101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112}
{105, 106, 107, 108}
{104, 101, 102, 103}

Walking through each step:

  • week1 set: [101, 102, 103, 104, 105, 106, 107, 108] (n=8)
  • week2 set: [105, 106, 107, 108, 109, 110, 111, 112] (n=8)
  • week1 | week2 (union): [101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112] (n=12)
  • week1 & week2 (intersection): [105, 106, 107, 108] (n=4)
  • week1 - week2 (difference): [101, 102, 103, 104] (n=4)
  • week2 - week1 (difference): [109, 110, 111, 112] (n=4)
  • week1 ^ week2 (symmetric difference): [101, 102, 103, 104, 109, 110, 111, 112] (n=8)

The counts tell the story. Union is 12, intersection is 4, only week 1 is 4, only week 2 is 4, and the symmetric difference is 8. Note that $8 + 8 = 16$, and $16 - 4 = 12$, which matches the union count because the 4 shared IDs are counted twice in the raw totals.

How to Interpret It

The intersection is your repeat group. Four respondents, IDs 105 through 108, answered in both weeks. If you are measuring retention or panel fatigue, that is the number you report.

The difference tells you who is new or lost. week1 - week2 gives the 4 IDs that appeared only in week 1, so those respondents dropped out. week2 - week1 gives the 4 IDs that appeared only in week 2, so those are new respondents.

The union is your total unique reach. Twelve distinct people responded across the two weeks, even though there were 16 responses in total. That gap between 16 responses and 12 people is exactly the kind of duplicate counting that sets remove automatically [1].

The symmetric difference is the "changed" group. Eight IDs appear in one week but not the other, which is the union minus the intersection.

When to Use It (and when not to)

Use a set when you need to:

  • Remove duplicates from a list of IDs, emails or category labels.
  • Test membership quickly, as in 'orange' in basket [1].
  • Compare two groups and find overlap, additions or removals.
  • Build a vocabulary of unique tokens before counting them.

Do not use a set when order matters, when you need to store duplicates, or when you need to index by position. Sets have no index, so week1[0] raises an error. If you need key-value lookups, a Python dictionary is the right tool. If you need ordered, labeled columns for analysis, a pandas DataFrame is usually the better container.

Python Set vs Python List

The closest related idea is the list, since both hold a collection of values. The difference is in the guarantees each one makes.

FeatureSetList
OrderUnordered [1]Ordered
DuplicatesNot allowed [1]Allowed
IndexingNot supportedSupported
Membership testFast, hash-basedSlower, scans elements
Typical useUnique values, set mathSequences, ordered data

If you convert a list to a set and back, you lose both the order and the duplicates. That is often exactly what you want when cleaning data, and exactly what you do not want when the order carries meaning.

Common Mistakes

  • Writing {} for an empty set. That creates an empty dictionary. Use set() instead [1].
  • Expecting a stable print order. Sets are unordered, so printing or iterating can produce a different order than you expect [1]. Sort the result if you need a fixed order.
  • Putting a list inside a set. Lists are mutable and unhashable, so {[1, 2]} fails. Convert the inner list to a tuple first.
  • Using - when you meant ^. Difference is one-directional, so week1 - week2 and week2 - week1 give different results. Symmetric difference gives both sides at once.
  • Assuming | works with a list. The operator needs a set on both sides. Use .union() if one side is a list.
  • Forgetting that sets drop duplicates silently. If you needed the counts, you have already lost them. Count first, then convert.

Limitations

Sets cannot preserve order or frequency, so they are the wrong structure for any question that starts with "how many times" or "in what order." Once you convert a list to a set, the duplicate information is gone and you cannot recover it from the set alone.

Sets also require hashable elements. That rules out lists, dictionaries and other mutable objects, which means nested or structured records need to be converted first. Finally, because sets are unordered, any output that depends on iteration order is not reproducible across runs unless you sort it explicitly. For a full picture of how sets fit with the other built-in containers, see Python data types explained.

Frequently Asked Questions

How do I create an empty Python set?

Call set() with no arguments. The literal {} creates an empty dictionary, not an empty set, so it will not behave the way you expect [1]. You can also build a set from any iterable, as in set([1, 2, 2, 3]), which returns {1, 2, 3}.

Does a Python set keep the order I add items in?

No. A set is an unordered collection, so iterating over it or printing it can produce the elements in a different order than you expect [1]. If you need a consistent order, sort the set into a list before displaying it.

What is the difference between union and intersection?

Union returns everything that appears in either set. Intersection returns only the elements that appear in both. In the survey example, the union had 12 IDs and the intersection had 4.

Can I put a list inside a set?

No. Set elements must be hashable, and lists are mutable so they are not hashable. Convert the list to a tuple, which is immutable, and then add it to the set.

How do I remove duplicates from a list with a set?

Pass the list to set(), which drops every repeated value, then convert back with list() if you need a list. Remember that the original order is not preserved. If order matters, use a loop or a dictionary to keep first occurrences instead, as shown in the Python for loop guide.

References

  1. 5. Data Structures, Python 3.14.8 documentation

Further Reading

Related Articles