Italy’s data protection authority has fined IQVIA Solutions Italy €7 million, approximately $7.8 million, after concluding that a health-data database covering roughly one million patients from about 800 general practitioners had not been properly anonymized. IQVIA had treated the records as anonymous because direct identifiers were generally replaced with patient-specific codes, but the Italian Garante found that those codes allowed individual patients to be followed longitudinally and, when combined with rich clinical and demographic information, made it possible to single out and potentially re-identify people using reasonable means. The case is important far beyond IQVIA because it challenges a common assumption in healthcare analytics: removing a patient’s name does not necessarily remove the patient’s identity.

The database contained information including year of birth, sex, diagnoses, symptoms, prescriptions, medical examinations, vaccinations and location-related information. Because each patient retained a consistent code across records, a person’s medical history could effectively be reconstructed over time. The regulator’s position was that the combination of persistent identifiers and highly detailed medical characteristics created enough uniqueness to isolate individuals and make re-identification realistically possible. In other words, IQVIA had removed obvious identifiers but retained enough structure around each person that the records remained personal data under GDPR rather than becoming truly anonymous statistical information.

That distinction matters enormously because anonymous data and pseudonymized data have very different legal and security consequences. Properly anonymous information falls outside many GDPR requirements because the individual can no longer reasonably be identified. Pseudonymized data, by contrast, remains personal data when a person can still be singled out, linked across records or reidentified using additional information. Replacing “John Smith” with “Patient 47A92” may reduce immediate visibility, but if Patient 47A92 appears repeatedly alongside the same age, location, diagnoses, prescriptions and treatment history, the underlying individual may still be distinguishable from everyone else in the dataset.

The practical test is therefore not simply:

“Did we remove the name?”

It is:

“Can someone still work out who this person is?”

That is a much higher bar.

Longitudinal health datasets are particularly difficult to anonymize because medical histories are inherently distinctive. A person’s combination of age, geographical area, disease history, specialist visits, medication sequence, vaccination history and diagnostic procedures can function almost like a fingerprint. Each individual attribute may be shared by thousands of people, yet the combination can narrow the population dramatically. Add enough fields and repeated observations over time, and “anonymous” records can become surprisingly recognizable.

This is why the use of a persistent patient code is central to the Italian regulator’s reasoning. The code did not simply conceal a name; it preserved continuity. Every time the same person appeared again, IQVIA could connect the new information to the existing health history. That continuity is extremely valuable for analytics because longitudinal research depends on understanding how one patient changes over time. Unfortunately, the feature that makes the dataset useful for longitudinal research is also what makes genuine anonymization significantly harder.

The underlying tension can be summarized rather neatly:

the more useful longitudinal patient data becomes for analytics, the harder it becomes to argue that the patient has disappeared from the data.

That does not mean longitudinal medical research is impossible. It means organizations need an appropriate legal basis, strong governance, privacy safeguards and realistic assumptions about whether the information is anonymous, pseudonymous or directly identifiable.

The Garante also found a more obvious problem within a subset of the database. Information relating to more than 3,300 patients contained direct identifiers including names, Italian tax identification numbers, addresses and contact details, and more than 3,000 of those individuals also had health information associated with their identities. Those records require little philosophical debate about anonymization because the individuals were directly identifiable.

The regulator’s findings went substantially beyond inadequate anonymization. It concluded that IQVIA lacked an appropriate legal basis for processing the health information, had not adequately informed patients, had not established defined retention periods, had not performed the required data-protection impact assessment, and had not implemented sufficient security measures. Some records in the database dated back as far as 2001, meaning sensitive health information had accumulated over more than two decades without defined retention limits.

That retention issue deserves almost as much attention as the anonymization finding. Health information does not become less sensitive merely because it becomes old. A diagnosis from twenty years ago may remain deeply private, while historical treatment and prescription information can reveal chronic illnesses, psychiatric conditions, reproductive health, substance-use issues and other aspects of a person’s life that remain sensitive indefinitely. If an organization retains health information for 20 or 25 years, it assumes responsibility for protecting that information for just as long.

The security principle is straightforward:

data that no longer exists cannot be breached, reidentified, leaked or misused.

Organizations therefore need a defensible answer to why each category of sensitive data is retained and for how long. “Storage is cheap” is technically true and, as security strategies go, magnificently unhelpful.

The IQVIA decision also illustrates why data minimization and retention are connected to cybersecurity, not merely regulatory compliance. Every additional attribute increases the possibility of re-identification. Every additional year of history increases the richness of the longitudinal profile. Every duplicate dataset creates another place from which the information can escape. Reducing data fields, geographical precision and retention periods can therefore reduce both privacy risk and the attractiveness of the dataset to attackers.

The regulator’s approach is also important for organizations that rely heavily on hashing, tokenization or coded identifiers. A hash or random identifier may prevent an analyst from immediately seeing a person’s name, but it does not necessarily anonymize the underlying record if the identifier remains stable. Stable identifiers make it possible to correlate the same person across tables, datasets and time periods. That capability can be operationally valuable, but it also means the person remains distinguishable.

This creates a broader rule for privacy engineering:

pseudonymization reduces exposure; anonymization removes identifiability. They are not interchangeable terms.

Pseudonymization is still valuable. It can reduce accidental disclosure, prevent ordinary analysts from seeing direct identifiers, separate identity data from clinical information and limit the consequences of some types of breach. GDPR itself recognizes pseudonymization as an important security measure. The problem begins when organizations treat pseudonymization as proof that the resulting information is no longer personal data.

The IQVIA ruling effectively says that the label applied by the organization does not determine the legal status of the data.

Calling a dataset anonymous does not make it anonymous.

The regulator looks at whether people can actually be identified.

The Garante also concluded that IQVIA was a data controller from the point the information was collected from general practitioners, rather than merely a passive recipient of already anonymous information. That finding matters because the controller determines the purposes and essential means of processing and therefore carries much more substantial GDPR responsibility. IQVIA could not avoid those obligations simply by describing the upstream coding process as anonymization.

That is an important lesson for data brokers, analytics companies, research organizations and AI companies that obtain datasets from external providers. Responsibility does not automatically disappear because another organization removed names before the information arrived. If the recipient helps design the coding mechanism, specifies the data fields, controls longitudinal identifiers or can reasonably reidentify individuals, regulators may still consider the recipient responsible for personal-data processing.

Healthcare organizations should therefore evaluate anonymization from the attacker’s or recipient’s perspective, not only from the perspective of the person performing the initial transformation. The relevant question is not whether one isolated file contains a name. It is whether the available information, combined with other reasonably accessible sources, can identify the person.

External datasets make this increasingly difficult.

Public registers, leaked databases, social media, commercial data brokers and location information can all provide auxiliary data useful for re-identification. A dataset that appeared anonymous ten years ago may be much easier to reidentify today because so much additional information about individuals exists elsewhere.

This is why modern anonymization requires thinking about linkage attacks. Suppose a health dataset contains year of birth, sex, approximate location and several medical events. Another dataset may contain age, location and employment information. A third may contain public announcements about a particular illness or hospital visit. Individually, none may identify the person. Combined, they may narrow the possibilities to one individual.

Healthcare information is especially vulnerable to this because rare diseases, unusual treatment sequences and specific combinations of medical events can be highly distinctive.

The regulator’s decision therefore reflects a shift away from simplistic de-identification toward a more realistic risk-based model: could someone reasonably single out or reidentify a patient using the available data and reasonably obtainable auxiliary information? If the answer is yes, the organization still possesses personal data and needs the corresponding legal basis, governance and security controls.

The fine also demonstrates that privacy failures increasingly overlap with cybersecurity failures. IQVIA was not sanctioned merely because someone had broken into its systems. The issue was how the information was designed, collected, retained and protected in the first place. Security teams often concentrate on perimeter controls, malware, vulnerabilities and access management while assuming privacy architecture belongs to legal or compliance teams. Large-scale health-data platforms make that organizational separation increasingly artificial.

A dataset can be perfectly protected from external attackers and still be processed unlawfully.

Conversely, a privacy-preserving architecture can substantially reduce the consequences when technical security eventually fails.

The strongest approach combines both disciplines.

For health-data analytics, that means minimizing identifiers, limiting longitudinal linkage where it is unnecessary, restricting geographical granularity, separating identity data from clinical data, encrypting sensitive fields, limiting employee access, maintaining detailed audit logs, enforcing retention periods, performing formal re-identification testing and regularly reassessing whether previously anonymized datasets remain genuinely anonymous.

The data protection impact assessment, which the Italian authority says IQVIA failed to perform, is particularly important for a dataset of this type. A DPIA is not supposed to be paperwork produced after the architecture has already been decided. Its value comes from forcing the organization to consider questions such as: How could patients be reidentified? What happens if datasets are linked? Which attributes are actually necessary? How long should they remain? Who can access them? What harm could occur if the information becomes attributable to an individual?

Had those questions been addressed rigorously at design time, the distinction between persistent pseudonyms and genuine anonymity should have been difficult to ignore.

The case also has implications for AI and machine-learning projects using health datasets. AI teams often seek very large longitudinal datasets because model performance improves with breadth and historical context. But the characteristics that make those datasets valuable to models also increase privacy risk. Long time series, detailed medical features and stable identifiers create excellent training signals and excellent re-identification signals simultaneously.

Removing names before training is therefore not necessarily sufficient.

Organizations need to consider whether the training dataset itself remains personal data, whether patient consent or another legal basis covers the intended processing, whether models could memorize sensitive records, and whether the outputs could reveal information about individuals.

The regulator’s reasoning is likely to resonate well beyond healthcare. Advertising platforms, location analytics firms, financial-data providers, telecommunications companies and data brokers frequently use persistent pseudonymous identifiers to follow individuals across time. The same conceptual question applies: if a stable identifier enables a person to be singled out and correlated across a rich dataset, can the organization genuinely claim the information is anonymous?

The answer increasingly appears to depend on what can be inferred, not merely which obvious fields were deleted.

IQVIA has disputed aspects of the regulator’s position and says it reserves the right to appeal. The company told BleepingComputer that it remains committed to responsible data use and uses safeguards including pseudonymization and encryption. IQVIA also emphasized that the dataset involved in the Italian decision is not used for its clinical-research services and does not concern clinical trials conducted for sponsors. The company says it has already begun implementing measures needed to align its processing with the regulator’s guidance.

That distinction around clinical trials is worth preserving because IQVIA operates across numerous healthcare and research businesses. The ruling concerns a specific Italian patient-data database rather than establishing that all IQVIA clinical research datasets or services suffer the same problem.

The regulator has given IQVIA 120 days to bring the relevant processing into compliance if it intends to continue the activity. Alternatively, the anonymization must be performed independently by the general practitioners under safeguards defined by the authority. The Garante says it considered the scale of the dataset, the sensitivity of the health information, the fact that doctors stopped transmitting information from 2023 onward, and IQVIA’s cooperation during the investigation when determining the €7 million penalty.

The investigation itself began after inspections in April 2025, and the final decision, No. 710, was adopted on September 23, 2026 and announced publicly on October 2. The regulator also incorporated a separate personal-data breach notification from IQVIA into the broader proceeding.

Interestingly, this is not IQVIA’s only substantial European health-data enforcement action in 2026. France’s CNIL fined IQVIA Operations France €5 million in May 2026 over separate health-data warehouse compliance issues involving safeguards imposed when the warehouses were authorized. That is a different case with different facts and should not be conflated with the Italian anonymization decision, but together the actions show European regulators placing increasing scrutiny on large-scale health-data analytics.

The broader lesson from the Italian case is therefore not simply that organizations need “better anonymization.”

It is that anonymization is a property of the resulting dataset, not a processing step or a label.

A company can hash every name, replace every identifier and remove every email address, yet still retain a dataset in which individual people are distinguishable and reasonably reidentifiable.

The correct test should look something like:

direct identifiers removed → persistent identifiers assessed → quasi-identifiers measured → longitudinal uniqueness tested → external data linkage considered → re-identification risk evaluated → only then determine whether the dataset is genuinely anonymous

Anything less risks confusing concealment with anonymity.

The case can be summarized as: health records collected from roughly 800 general practitioners → approximately one million patients assigned persistent codes → detailed longitudinal clinical and location information retained → IQVIA treats dataset as anonymous → Italian regulator determines patients remain singlable and reasonably reidentifiable → additional direct identifiers found for over 3,300 patients → legal basis, notice, retention, DPIA and security shortcomings identified → €7 million penalty and 120-day remediation order.

The most useful security and privacy lesson is surprisingly simple:

Removing someone’s name does not remove the person.

If the remaining information still tells a sufficiently detailed story about one individual, the organization is still handling that individual’s data, however anonymous the database happens to be called.


Italy's Data Protection Authority (GPDP) has fined IQVIA €7 million ($7.8M) over poor data-processing practices that the agency says could have put roughly one million patients at risk of data exposure and de-anonymization. [...]

Source: IQVIA fined $7.8 million for failing to properly anonymize health data via Bleeping Computer — published 05 Oct 2026.