Article 1
When Is Genomic Data Truly Anonymous?
By Eshara Chotoo
One of the most important questions in genomic data law is deceptively simple: when does data stop being “personal information”?
South Africa’s courts are increasingly being asked to grapple with the meaning of identifiability under POPIA — including whether information that does not contain a person’s name can nevertheless relate to an identifiable person.
The underlying legal disputes may not concern genomics directly. But the principle matters enormously for genomics.
Genomic information is fundamentally different from most ordinary datasets. Removing a name, identity number or contact detail does not necessarily render genomic information anonymous. A genome may still be capable of being associated with an individual when combined with family information, phenotype, geography, clinical records or other datasets.
That distinction between pseudonymisation and true anonymisation will become increasingly important as African genomic programmes scale.
From contractual language to legal frameworks
From a contractual perspective, this means we need to move beyond simplistic wording such as:
- “the data has been de-identified and therefore no longer constitutes personal information.”
A stronger legal framework distinguishes clearly between:
- identifiable data;
- pseudonymised data; and
- genuinely de-identified data.
Each category carries different legal and practical risks.
Genomic-data sovereignty beyond privacy classification
But there is an even more important point. Genomic-data sovereignty should not depend entirely on whether a court ultimately classifies a particular dataset as personal information under privacy legislation.
Contracts can — and arguably should — independently regulate issues such as:
- re-identification;
- linkage with external datasets;
- cross-border transfer;
- onward sharing;
- permitted uses;
- retention and deletion;
- AI and model training;
- commercialisation; and
- derivatives generated from the data.
For African genomic projects, this matters because the scientific and economic value of a dataset does not disappear simply because a participant’s name has been removed.
Why this matters to the 108,000 Africa Genomes Project
This is not merely an academic question for Antares. Through the 108,000 Africa Genomes Project, we are confronting these questions in practice as we work toward building a large-scale African genomic resource with participation across the continent.
One of the principles shaping the Project is that privacy, scientific utility and African data sovereignty cannot be treated as separate conversations.
Genomic research requires collaboration. Precision medicine requires access to high-quality data. Scientific discovery requires scale. But none of those objectives necessarily requires Africa to relinquish agency over the genomic resources generated from African populations.
As genomic datasets increase in scale and value, the legal frameworks surrounding them have to evolve just as rapidly. The question is therefore becoming bigger than privacy.
It is increasingly about control, permitted use, accountability and sovereignty over genomic information itself.
Africa should not simply participate in that conversation. Africa should help define it.
References & Further Reading
- Protection of Personal Information Act 4 of 2013 (POPIA), Republic of South Africa
- Information Regulator South Africa — regulatory guidance and enforcement activity
- Regulations relating to the Processing of Data Subjects’ Health Information by Certain Responsible Parties, 2026, Government Gazette No. 54268, 6 March 2026
- South African Government Gazette No. 54268, 6 March 2026
Note: The discussion in this article regarding genomic identifiability reflects the application of existing privacy-law principles to genomic information. Genomic data presents distinctive re-identification risks because it may remain identifying or become identifiable when combined with other datasets.
