State of M&A Data Rooms — Q2 2026 Read the report →

The Biggest Data Breaches in History (Correctly Classified) in 2026

Co-founder and CEO at Peony. I built the data room platform with a background in document security, file systems, and AI. Founded Peony in 2021 in San Francisco.

The Biggest Data Breaches in History (Correctly Classified)

I'm Deqian Jia, co-founder of Peony, a data room company. I spend my days thinking about who can see a document and who cannot, so I read a lot of breach coverage. Most "biggest data breaches ever" lists have the same three problems: they quote the first headline number and never update it after the count is revised, they call scrapes and misconfigurations "hacks," and they add unrelated incidents together to manufacture a scary total. All three are avoidable.

This post fixes them. Every incident below carries the current authoritative figure, the classification it actually belongs to, and the source named in the sentence. Where a number was revised, I lead with the current one and note the history. Where a "record" is not a "person," I say so. And where a famous number is really a scrape, an exposure, or a compilation of old leaks rather than a fresh breach, it gets labeled that way, because that distinction is the entire point.

What is the biggest data breach in history?

The biggest data breach in history, by account count, is the 2013 Yahoo breach: all 3 billion Yahoo user accounts, per the company's October 3, 2017 disclosure. It is a genuine breach, meaning attackers accessed and took the data from Yahoo's systems.

The caveat is what most lists skip. Several numbers quoted as bigger than Yahoo are not breaches at all. First American Financial's roughly 885 million documents were an exposure left publicly readable, not a confirmed theft. Facebook's 533 million-record dataset was scraped from public profiles, not hacked out of Meta's servers. The "16 billion passwords" of June 2025 were a compilation of old leaks, not a single new incident. Once you sort incidents by what actually happened, Yahoo's 3 billion accounts is the largest true breach on the public record, and the rest of this list is ranked and tagged so you can tell one category from another.

The biggest data breaches, ranked and classified

Here are the incidents covered below, with the current authoritative count and the correct classification for each. The classification column is the one to read carefully: a bigger number in an exposure or scrape row does not mean a bigger breach. Do not add these figures together, and remember that a record is not always a person.

IncidentYearCurrent authoritative countClassificationSource (named in-sentence below)
Yahoo2013~3 billion accountsBreachYahoo SEC 8-K, Oct 2017
First American Financial2019~885 million documents exposedExposureKrebsOnSecurity / NYDFS
Facebooksurfaced 2021~533 million usersScrapeMeta newsroom, Apr 2021
Marriott / Starwood2018~383 million recordsBreachMarriott press release, Jan 2019
Change Healthcare2024~192.7 million peopleBreachOCR filing, July 2025
Equifax2017~147 million US consumersBreachCFPB / FTC settlement
MOVEit / Cl0p202395,788,491 people across 2,773 orgsCampaignEmsisoft tracker, as of Jun 28 2024
National Public Data2024contested; ~272M unique SSNs (est.)Breach (contested)TechCrunch / researcher estimates

A note on how this is ordered: the rows are sorted by the size of the current authoritative count, not by severity, with one deliberate exception: National Public Data sits last because its figures are contested estimates rather than a confirmed count. First American's 885 million sits high because that is how many documents were accessible, but it is an exposure, and regulators documented unauthorized access to far fewer. That is exactly the trap this list is built to avoid. Two frequently-cited 2025 to 2026 events, the Salesforce-ecosystem extortion campaign and the 2026 Canvas incident, are discussed below but deliberately left out of the ranked table because neither has a company-confirmed record count.

Yahoo, 2013: 3 billion accounts (breach)

The 2013 Yahoo breach affected all 3 billion Yahoo user accounts, per Yahoo's statement filed with the SEC as an 8-K on October 3, 2017. This is the current authoritative figure, and it is the result of two upward revisions, which is why so many older lists still quote the wrong number.

The revision history is worth getting right, because two separate incidents get conflated constantly. In September 2016, Yahoo disclosed a breach of at least 500 million accounts, an event dated to 2014. In December 2016, it disclosed a different breach, from August 2013, affecting more than 1 billion accounts. Then on October 3, 2017, during the Verizon acquisition, Yahoo raised the 2013 figure to all roughly 3 billion accounts. So the 3 billion number belongs to the 2013 event and the 500 million number to the 2014 event. They are not the same breach and must not be summed.

As for what was taken, Yahoo stated the stolen data did not include passwords in clear text, payment card data, or bank account information. It was account information, later tied to US Department of Justice indictments of state-sponsored and for-hire actors. This is a textbook breach: attackers stole account databases directly from Yahoo's systems.

Change Healthcare, 2024: about 192.7 million people (breach)

The 2024 Change Healthcare breach affected approximately 192.7 million individuals — the estimate Change reported to the HHS Office for Civil Rights as of July 31, 2025. It is the largest breach of US medical data in history, touching nearly two-thirds of the US population.

This figure kept being revised upward — which is exactly how mega-breach numbers behave. Change first reported the regulatory minimum of 500 individuals, then about 100 million to OCR in October 2024, then roughly 190 million per a UnitedHealth Group spokesperson in January 2025, and finally 192.7 million in the July 2025 OCR update. Any source still citing 100 million — or even 190 million — is out of date.

The data involved is unusually sensitive: names, addresses, dates of birth, social security numbers, driver's license or passport numbers, diagnoses, medications, test results, and health-insurance and financial information. How it happened is the part that matters for prevention: attackers, a BlackCat/ALPHV affiliate, used stolen credentials on a Citrix remote-access portal that lacked multi-factor authentication to breach Change Healthcare in February 2024. That single detail, valid credentials plus no MFA, is the thread running through most of this list.

First American Financial, 2019: about 885 million documents (exposure)

In 2019, First American Financial left approximately 885 million documents publicly accessible, dating back to 2003. KrebsOnSecurity, reporting the New York Department of Financial Services charges, described approximately 885 million records related to mortgage deals going back to 2003, and noted the records were available without authentication to anyone with a Web browser.

This is why classification matters. First American is an exposure, not a confirmed exfiltration. The 885 million documents were accessible, but that is not the same as saying 885 million records were stolen. NYDFS documented unauthorized access to over roughly 350,000 documents during its review window. So the honest framing is that about 885 million documents were left publicly accessible, and a much smaller number is known to have been accessed.

The data on offer was severe: bank account numbers and statements, mortgage and tax records, social security numbers, wire transaction receipts, and drivers license images. The cause was an insecure direct object reference flaw in the EaglePro application, which let anyone view 16 years of title and mortgage documents by changing a number in the URL. The flaw was introduced in May 2014, found internally in December 2018, and not fixed until public disclosure in May 2019.

This was NYDFS's first-ever cybersecurity enforcement action, with charges brought in July 2020. Two distinct penalties often get muddled, so precisely: the SEC imposed a civil penalty of $487,616 in 2021, and the NYDFS matter was resolved with a $1,000,000 settlement in 2023.

MOVEit, 2023: about 95.8 million people across 2,773 organizations (campaign)

The 2023 MOVEit event was a campaign, not one company's breach, and treating it as a single incident with one record count is a mistake. The ransomware and extortion group Cl0p exploited a SQL-injection zero-day (CVE-2023-34362) in Progress Software's MOVEit Transfer file-transfer tool and used it to steal data from thousands of downstream organizations that all happened to use the same software.

The most-cited tally comes from Emsisoft's tracker, which put the impact at 2,773 organizations and 95,788,491 individuals, current as of June 28, 2024. Two caveats travel with that number and should never be dropped. First, it is an aggregate across roughly 2,773 organizations, not a single company's loss. Second, it is a tracker estimate as of a specific date, not a final official count. Emsisoft's sector split shows the spread: education at 39.1%, health at 20.1%, and finance and professional services at 13.3%.

MOVEit is the canonical example of supply-chain risk, where one vulnerability in one widely-used tool cascades into hundreds of separate disclosures. It pairs directly with the Verizon data covered later: third-party involvement is now a factor in nearly half of all breaches.

Facebook, surfaced 2021: 533 million users (scrape)

The dataset of about 533 million Facebook users that circulated in 2021 was scraped, not hacked, and this is one of the most abused misclassifications in every breach list. Meta's own newsroom statement on April 6, 2021 said malicious actors obtained this data not through hacking its systems but by scraping it from the platform, and that the data was scraped from people's Facebook profiles by malicious actors using the contact importer prior to September 2019.

So the mechanics are the opposite of an intrusion: attackers abused the legitimate contact-importer feature to harvest data from public profiles at scale, before September 2019. No one broke into Facebook's servers. Frame it as a roughly 533 million-record scraped dataset, never as "Facebook was hacked."

One more precision point: the 533 million figure itself traces to Business Insider's reporting, not to a Meta count. Meta's post does not cite that number. That is a good habit for reading any breach story: notice whether the headline figure comes from the company or from a third party, because the two are not equally reliable.

Marriott, 2018: about 383 million records (breach)

Marriott's Starwood breach is the cleanest example of why a record is not a person. Marriott's stated upper limit is approximately 383 million records, not 383 million unique guests, because many guests had multiple records in the database.

The revision runs the other direction from Yahoo and Change Healthcare, downward rather than up. Marriott's initial disclosure on November 30, 2018 said up to about 500 million guests. Its January 4, 2019 update revised the upper limit down to approximately 383 million records and explicitly cautioned that this was records, not unique guests. The UK Information Commissioner's Office later worked with a figure of about 339 million guest records globally. So the correct phrasing is roughly 383 million records, revised down from an initial roughly 500 million, and always records rather than people.

The exposed data included guest information, about 20.3 million encrypted passport numbers, about 5.25 million unencrypted passport numbers, and some encrypted payment card data, with no evidence that the passport master encryption key was taken. The cause was a long-running intrusion: attackers were inside the legacy Starwood reservation database from around 2014, undetected until September 2018. Marriott inherited that compromise when it acquired Starwood, which makes this a due-diligence lesson as much as a security one.

Equifax, 2017: about 147 million US consumers (breach)

The 2017 Equifax breach exposed the sensitive personal information of approximately 147 million US consumers. The joint CFPB, FTC, and states settlement announcement stated that a data breach at the company resulted in the exposure of approximately 147 million US consumers' sensitive personal information, and specified the data as names, addresses, social security numbers, and dates of birth.

The cause was a known, unpatched vulnerability: attackers exploited an Apache Struts web-application flaw for which a fix was already available. That places Equifax squarely in the category Verizon now calls the top initial-access vector, vulnerability exploitation.

The settlement is frequently misquoted, so the figures from the same source are: a global settlement of up to $700 million in relief and penalties, including a consumer-relief fund of up to $425 million and a $100 million CFPB civil money penalty. The payout has a long tail. As of August 2025, the court-appointed settlement administrator was distributing final payments from the $425 million restitution fund and had added funds to many prepaid cards. The claim-filing deadline of January 22, 2024 has closed, so the settlement is in its final distribution phase rather than accepting new claims.

National Public Data, 2024: contested figures (breach)

National Public Data (NPD), a data broker, had its files breached in 2024, and it belongs on this list mainly as a warning about how record counts get inflated. The threat actor's marketing claimed 2.9 billion records, and that number spread as if it meant 2.9 billion people. It does not.

The 2.9 billion figure is a record count heavily padded with duplicates and dead or outdated data, not a count of individuals. Researchers analyzing the data put the number of unique social security numbers on the order of about 272 million, itself an estimate, and noted that many records were stale. So the honest version is: a data broker whose files were breached in 2024, where the widely-cited 2.9 billion records is inflated by duplicates and outdated data, and researchers estimate on the order of about 272 million unique social security numbers. Do not headline 2.9 billion people, and do not sum this with any other incident.

The corporate aftermath is confirmed and telling: NPD's parent, Jerico Pictures, Inc., filed for Chapter 11 bankruptcy in October 2024, citing breach fallout, more than 20 class actions, and state attorney general inquiries, per TechCrunch. The site itself later returned, per Malwarebytes reporting in August 2025.

What about the 2026 Canvas incident?

The 2026 Canvas incident at Instructure is recent enough to appear on freshly-written lists, so it is worth stating carefully rather than omitting. It is left out of the ranked table on purpose, because the big number attached to it is an attacker claim, not a company-confirmed count.

What Instructure officially confirmed: a cybersecurity attack in late April to May 2026 involving names, email addresses, student ID numbers, and messages among users, and the company said it found no evidence that passwords, birth dates, government IDs, or financial information were involved. On May 11, 2026, Instructure said it had reached an agreement with the actor and that the compromised data was destroyed.

What is not company-confirmed: the group ShinyHunters claimed roughly 275 million users, about 9,000 schools, and 3.65 TB of data, and a roughly $10 million ransom figure is an unconfirmed rumor. As of late August 2026 that remains true: Instructure has not confirmed the attacker's claimed counts, no public leak of the Canvas data has surfaced since the company's May 11 statement that the data was destroyed, and putative class actions filed from May 2026 onward are working through the courts. So if you see Canvas ranked by "275 million users" as established, that figure is an attacker or third-party claim, not confirmed by Instructure. The same discipline applies to the 2025 Salesforce-ecosystem extortion campaign by the group tracked as ShinyHunters and UNC6040: it stole data from dozens of companies' Salesforce environments through voice-phishing of their users, but Salesforce stressed the platform had not been compromised and the issue was not due to any known vulnerability in its technology. There is no approved single record count for that campaign, so it is a trend example here, not a ranked entry.

Breach vs exposure vs scrape vs compilation vs campaign

This is the section that separates a rigorous list from a scary one. Five things get called "the biggest breach" in casual coverage, and only one of them is actually a breach. Here is how to tell them apart.

CategoryWhat it meansWas there an intrusion?Example on this list
BreachUnauthorized access to or acquisition of data from a systemYesYahoo, Equifax, Marriott, Change Healthcare
ExposureData left publicly accessible via misconfiguration or a flaw such as an IDORNot necessarily; accessible is not the same as stolenFirst American Financial
ScrapeAutomated collection of data from public or semi-public profiles, often via an abused featureNo; the source data was already reachableFacebook
CompilationAn aggregated dump of usernames and passwords from many prior leaks and infostealer logsNo; it repackages old incidentsThe "16 billion passwords" story
CampaignOne threat actor exploiting one tool or vector to hit many organizations at onceYes, but distributed across many victims, with no single countMOVEit, the Salesforce campaign

Four rules fall out of that table, and they are the rules most lists break:

  • A record is not a person. Marriott is about 383 million records, not guests. NPD's 2.9 billion records is not 2.9 billion people.
  • Never sum across incidents. Adding Yahoo, MOVEit, and NPD together produces a meaningless number. Each has its own scope, its own de-duplication, and its own definition of what was counted.
  • A tracker is an estimate as of a date, not a final count. MOVEit's 95,788,491 is current as of June 28, 2024, and it may move.
  • An attacker's claim is not a company-confirmed figure. Canvas's 275 million and NPD's 2.9 billion are attacker or marketing claims. Label them as such.

Was the "16 billion passwords" leak really a breach?

No. The "16 billion passwords" headlines of June 2025 did not describe a new breach. They described a compilation of previously-leaked credentials and infostealer logs, repackaged from many prior incidents, not a fresh intrusion at any single company.

The story was first reported by Cybernews on June 18, 2025 and was rapidly walked back by security press. BleepingComputer's coverage stated plainly that the 16 billion credentials leak is not a new data breach. CyberScoop called the "16 billion password breach" story a farce, and BankInfoSecurity headlined its analysis as a hype alert about "the largest data breach in history that wasn't."

The practical takeaway: there was no centralized breach of Apple, Google, or Facebook behind those numbers. Aggregated credential dumps are assembled from data already stolen in earlier, unrelated events, sometimes years apart. They are a real reason to use unique passwords and a password manager, but they are not a single incident and do not belong in a ranking of breaches. Any list that puts "16 billion" at the top is counting a compilation as if it were one hack.

What do the biggest breaches have in common?

Access is the common thread. Strip away the differences in scale and sector, and most of these incidents come down to someone getting in through a credential, an unpatched vulnerability, or a third party's connection.

The current industry data backs this up. Verizon's 2026 Data Breach Investigations Report found that nearly a third, 31%, of all breaches start with vulnerability exploitation, the first time in about 19 years that vulnerability exploitation surpassed stolen credentials as the top initial-access vector. The same report found that breaches involving a third party now account for 48% of all breaches.

Both numbers map onto the incidents above. Equifax and MOVEit were vulnerability exploitation: a known Apache Struts flaw in one case, a MOVEit Transfer zero-day in the other. Change Healthcare was the credential story, valid logins on a remote-access portal with no multi-factor authentication. And MOVEit is the archetype of the 48% third-party figure, since the victims were breached through software they depended on. The lesson is not that any one company was careless. It is that access control and third-party access are where failures concentrate, so that is where defenses pay off most.

What does a data breach cost?

Per IBM's 2026 Cost of a Data Breach Report, published July 29, 2026, the global average cost of a data breach rose about 12% to $4.99 million — reversing the prior year's decline (the 2025 report had recorded the first drop in five years, to $4.44 million). In the United States, the average reached a record $11.5 million, more than double the global figure. The 2026 report's sharpest finding: one in four malicious breaches was AI-enabled — a 56% jump year over year — and those breaches averaged about $6 million each.

Those averages sit well below the headline settlements in this post, and that gap is instructive. Equifax's up-to-$700 million settlement and First American's penalties are outliers driven by scale, regulatory exposure, and litigation, not typical breach economics. For most organizations the real cost is the unglamorous average: detection and response, notification, credit monitoring, legal fees, lost business, and operational drag. The cheapest breach is the one that access controls prevent from happening.

How businesses reduce document-breach risk

If access is where the biggest breaches begin, the practical question for any business is narrower than "how do we stop hackers." It is: when we send a sensitive document outside the company, who can open it, for how long, and can we prove it and cut it off?

Most document leaks are not sophisticated intrusions. They are a link that stayed live too long, a file forwarded to the wrong person, an attachment sitting in a later-compromised inbox, or an over-broad "anyone with the link" share. Standard tools make this worse: as I've written in Is OneDrive secure?, the security depends heavily on how sharing is configured, which is easy to get wrong. Email is worse still, because once an attachment leaves you cannot expire it, watermark it, or see who opened it, the gap covered in how to send confidential documents via email.

A permissioned data room closes that gap by binding access to identity instead of to a link. The controls that map to the failure modes above:

  • Link expiry and revoke so access is time-bound and can be cut instantly when a deal falls through or a recipient leaves. At Peony, this is on every tier, including the free plan.
  • Dynamic per-viewer watermarks that stamp each page view with the viewer's identity, so a leaked screenshot points back to its source. On Peony this is on the Data Room plan at $52/admin/month.
  • Audit trails and page-level analytics so you can see who opened what, when, and for how long. Page-level analytics is free on Peony.
  • Screenshot protection that blocks and logs capture attempts, on Peony's Business plan at $30/admin/month.

I run Peony, a data room company used by 6,800+ customers, and Peony is SOC 2 Type II. A data room would not have prevented Equifax or MOVEit, which were platform and supply-chain failures at enormous scale. But for the ordinary act of sharing a contract, a diligence file, or a board deck with an outside party, the difference between an open link and a permissioned, revocable, watermarked one is the difference between hoping nothing leaks and being able to see, and stop, when it does. If you are evaluating options, I compared the category in my ranking of document security platforms; if you are assessing a company's exposure rather than your own, cybersecurity due diligence walks through how to read a target's breach history and controls.

The honest starting point is the free tier: link expiry, revoke, and page-level analytics at no cost, covering the two failure modes, stale access and no visibility, behind most everyday leaks. The full breakdown is on the pricing page. It is why 6,800+ customers keep their sensitive documents in a room rather than an inbox.

Frequently asked questions

What is the biggest data breach in history?

The biggest data breach in history by account count is the 2013 Yahoo breach, which the company disclosed on October 3, 2017 affected all 3 billion Yahoo user accounts. It is a genuine breach: attackers stole account databases. Yahoo stated the stolen data did not include passwords in clear text, payment card data, or bank account information. Note that many larger-sounding figures you will see quoted are exposures, scrapes, or credential compilations, not breaches, so a raw ranking that mixes those categories is misleading.

How many accounts were affected in the Yahoo breach?

All 3 billion Yahoo user accounts were affected in the 2013 breach, per Yahoo's October 3, 2017 statement filed with the SEC. The figure was revised twice: Yahoo first disclosed more than 1 billion accounts in December 2016, then raised the 2013 total to all 3 billion accounts in October 2017. A separate 2014 Yahoo breach affected at least 500 million accounts. These are two different incidents and should not be combined into one number.

How many people were affected by the Change Healthcare breach?

Approximately 192.7 million people were affected by the 2024 Change Healthcare breach — the figure Change reported to the HHS Office for Civil Rights as of July 31, 2025. The estimate climbed in stages: about 100 million filed to HHS in October 2024, roughly 190 million per UnitedHealth Group in January 2025, then 192.7 million. It is the largest breach of US medical data in history. Attackers used stolen credentials on a Citrix remote-access portal that lacked multi-factor authentication.

Was the 16 billion password leak real?

The June 2025 '16 billion passwords' headlines did not describe a new breach. What was actually found, first reported by Cybernews, was a compilation of previously-leaked credentials and infostealer logs aggregated from many prior incidents, not a fresh intrusion at any company. Security press including BleepingComputer and CyberScoop debunked the framing, with BleepingComputer stating plainly that the 16 billion credentials leak is not a new data breach. There was no centralized breach of Apple, Google, or Facebook behind those numbers.

What is the difference between a data breach and a data leak?

A data breach is unauthorized access to or acquisition of data from a system, meaning an attacker got in. A data leak, often called an exposure, is data left publicly accessible through a misconfiguration or flaw such as an insecure direct object reference, which does not by itself prove the data was stolen. Two related categories are a scrape, the automated collection of data from public or semi-public profiles, and a compilation, an aggregated dump of credentials from many prior leaks. Mixing these categories inflates rankings.

What was the Equifax breach?

The 2017 Equifax breach exposed the sensitive personal information of approximately 147 million US consumers, per the joint CFPB, FTC, and states settlement, which described names, addresses, social security numbers, and dates of birth. Attackers exploited an unpatched Apache Struts web-application vulnerability. The global settlement reached up to $700 million, including a consumer-relief fund of up to $425 million and a $100 million CFPB civil penalty. As of August 2025, the court-appointed administrator was distributing final payments from the restitution fund.

How many people were affected by the Marriott breach?

Marriott's stated upper limit is approximately 383 million records, not 383 million unique guests. The distinction matters: many guests had multiple records. Marriott's initial November 2018 disclosure said up to about 500 million guests, then a January 4, 2019 update revised the upper limit down to approximately 383 million records and cautioned that this was records, not unique people. The compromise sat inside the legacy Starwood reservation database from around 2014 until it was found in September 2018.

Was the Facebook breach actually a hack?

No. The dataset of about 533 million users that surfaced in 2021 was scraped, not hacked. Meta's April 6, 2021 statement said malicious actors obtained the data not through hacking Meta's systems but by scraping it from the platform using the contact importer feature prior to September 2019. Calling it a hack is one of the most common misclassifications in breach lists. The 533 million figure itself traces to Business Insider's reporting, not to a Meta count.

How many people did the MOVEit attacks affect?

The MOVEit attacks were a campaign, not a single company's breach. Cl0p exploited a zero-day in Progress Software's MOVEit Transfer tool and hit thousands of downstream organizations at once. Emsisoft's tracker put the total at 2,773 organizations and 95,788,491 individuals, current as of June 28, 2024. Because it is an aggregate across many organizations and a tracker estimate as of a date rather than a final official count, it should never be presented as one company being breached for one number.

What is the safest way to share sensitive business documents externally?

The safest way is a permissioned data room rather than email attachments or open links, because most mega-incidents trace to credentials and over-broad access. I run Peony, a data room used by 6,800+ customers. On every tier including the free plan, you get link expiry and revoke so access can be cut instantly, plus page-level analytics. The Data Room plan at $52/admin/month adds dynamic per-viewer watermarks that bind each page view to a specific person, and Peony is SOC 2 Type II. Viewers are always free.