Working with research data at the Desmond Tutu TB Centre (DTTC) within Stellenbosch University, and later developing DroitPlus, taught me the same lesson in two very different environments: data governance is not primarily about databases, policies or compliance forms. It is about protecting people while enabling institutions to use information for legitimate public benefit.
In clinical research, a poorly governed record can expose a participant’s health status, compromise a study or undermine trust in an institution. On a digital platform, weak controls can disclose a person’s legal problem, identity, financial circumstances or private communications. When artificial intelligence is introduced, the consequences extend further: sensitive data may influence automated recommendations, train models, reproduce historical inequalities or travel through systems that neither the individual nor the original data collector fully understands.
Africa therefore needs more than generic calls for “responsible data use.” We need a practical, risk-based governance model that applies across clinical trials, public institutions and digital platforms.
The central policy question should not be: Can we collect this data?
It should be: What harm could arise from collecting, linking, analysing, sharing or automating decisions with it—and what level of protection does that risk require?
What clinical research teaches us
Clinical trials operate within demanding ethical and regulatory environments for good reason. A research participant is not merely a row in a database. That row may contain a diagnosis, laboratory result, treatment history, residential location, demographic profile and study identifier. Individually or in combination, these fields can reveal a person’s identity and circumstances.
At the DTTC—an academic research centre focused on TB and HIV within Stellenbosch University’s Faculty of Medicine and Health Sciences—data teams support the full research-data lifecycle, including collection, management, storage, restructuring and analysis. The Centre also works with large routine health datasets and clinical research networks. This environment demonstrates why quality, security, traceability and ethical oversight cannot be separated. DTTC describes these capabilities as spanning the full lifecycle of research data.
Several principles from clinical research should inform the governance of every sensitive digital system:
- Collect information for a defined purpose.
- Limit access according to responsibility.
- document how records are created and changed.
- Separate identifying information where possible.
- Validate data before relying on it.
- investigate deviations and security incidents.
- retain information only as long as justified.
- maintain human accountability for consequential decisions.
These practices are sometimes treated as obstacles to innovation. In reality, they are what make trustworthy innovation possible.
A clinical dataset that cannot be validated cannot support reliable evidence. An AI system built on poorly governed data suffers from the same weakness—only at greater speed and scale.
Digital platforms create a different risk environment
Building DroitPlus moved these questions into another domain. Digital platforms can broaden access to information and services, particularly where professional assistance is costly, geographically distant or difficult to navigate. But they also concentrate sensitive information and introduce new relationships among users, platform operators, cloud providers, analytics services and AI models.
This changes the risk profile.
A traditional database may have a limited set of authorised users and predictable reporting functions. A modern platform may process free-text narratives, uploaded documents, behavioural logs and inferred attributes. It may connect to third-party services, operate across borders and use algorithms whose outputs change over time.
Free text is especially dangerous. People rarely disclose information in neat, predefined categories. A description of a health, employment or legal problem may simultaneously reveal family relationships, migration status, disability, income, location or exposure to violence.
Anonymisation is also not a universal solution. Once datasets are linked, records thought to be anonymous may become identifiable through combinations of age, location, dates and uncommon events. In communities with small populations or distinctive characteristics, re-identification may be easier still.
Governance designed only around whether a name or identity number is present is therefore obsolete.
Africa should regulate by risk, not by sector alone
Africa’s data-policy landscape is developing, but implementation remains uneven. The African Union’s Data Policy Framework calls for harmonised governance, trustworthy digital environments, protection of digital rights and the equitable use of data for development. That continental direction is important. It now needs to be translated into operational rules institutions can apply. The AU framework explicitly connects data governance with rights, trust and inclusive development.
A risk-based framework would classify data activities according to their potential impact.
Tier 1: Low risk
This could include public information, properly aggregated statistics or genuinely anonymous operational data.
Controls should remain proportionate: baseline security, quality checks, clear ownership and documented retention rules.
Tier 2: Moderate risk
This category could cover identifiable customer records, internal employee information, platform usage data or pseudonymised research records.
It should require role-based access, encryption, processing registers, contractual controls for service providers, incident-response procedures and periodic audits.
Tier 3: High risk
This should include health information, biometric data, children’s data, detailed legal matters, precise location histories, linked administrative datasets and information about vulnerable populations.
High-risk processing should trigger mandatory impact assessments, stronger segregation, documented necessity and proportionality, independent oversight, auditable access logs and enforceable restrictions on secondary use.
Tier 4: Critical or consequential use
The highest tier should apply when sensitive data informs decisions about diagnosis, treatment, employment, credit, insurance, policing, migration, social benefits or access to justice.
It should also cover AI systems that rank people, predict behaviour or generate recommendations capable of materially affecting their rights.
These systems should face pre-deployment assessment, independent validation, bias and performance testing, meaningful human review, accessible explanations, appeal mechanisms and ongoing monitoring after deployment. Some applications may be too dangerous to permit, regardless of technical safeguards.
This approach would allow governments to protect people without imposing the same administrative burden on every spreadsheet, research database and AI system.
Consent is necessary, but it is not enough
African institutions too often treat consent as the entire ethical foundation of data processing. Once a user clicks “accept,” responsibility is considered transferred to the individual.
That is neither realistic nor fair.
People may need healthcare, employment, education or legal support. They often cannot negotiate platform terms or understand complex chains of secondary processing. Consent obtained through long notices, default settings or unequal relationships offers limited protection.
Institutions must remain accountable even when consent has been obtained. They should have to demonstrate:
- a legitimate and clearly defined purpose;
- that each category of information is necessary;
- that less intrusive alternatives were considered;
- that secondary uses are compatible with the original context;
- that people can exercise meaningful rights;
- and that expected public benefits outweigh foreseeable harms.
South Africa’s Protection of Personal Information Act provides an important legal foundation, including specific safeguards for health information and other categories of special personal information. But legal compliance should be treated as the minimum threshold, not the final objective. POPIA’s provisions on special personal information are contained in sections 26–33.
AI governance begins with data governance
My experience across data engineering, business intelligence and AI has reinforced a simple point: an algorithm inherits the strengths and weaknesses of the system around it.
A model may be technically sophisticated and still be unsafe because:
- its training data excludes rural or marginalised populations;
- historical records reflect unequal access to services;
- source data were collected for a different purpose;
- important fields have inconsistent definitions;
- labels encode subjective or discriminatory judgments;
- performance is measured only at an overall population level;
- or no one monitors whether accuracy deteriorates after deployment.
Business intelligence systems can create similar problems without being labelled “AI.” A dashboard determines which indicators leaders see, how categories are defined and which communities appear to be succeeding or failing. Data pipelines embed policy choices long before a visualisation reaches a decision-maker.
Governance must therefore cover the full chain:
collection → validation → integration → analysis → presentation → decision → review
Every consequential metric and model should have a documented lineage: where the data came from, who transformed it, what assumptions were made, which groups may be underrepresented and who is accountable for the resulting decision.
The World Health Organization recommends that AI for health protect autonomy, promote safety, remain transparent, assign accountability, advance inclusion and undergo continuous assessment. These principles should inform African AI policy well beyond healthcare. WHO’s guidance places ethics and human rights at the centre of AI design and deployment.
Five policy actions Africa can take now
First, governments should require data-protection and algorithmic-impact assessments for high-risk systems. These should be living documents, revisited when datasets, models, vendors or purposes change—not forms completed once and filed away.
Second, public procurement must become a governance checkpoint. No public body should purchase a sensitive data or AI system without contractual rights to audit it, test performance, investigate incidents, retrieve institutional data and exit the service without becoming operationally trapped.
Third, regulators should require public registers of consequential automated systems. Citizens should be able to discover where algorithms are used in public services, what they influence, who operates them and how an affected decision can be challenged. Transparency need not expose source code or legitimate security controls, but secrecy cannot be the default.
Fourth, Africa must invest in institutional capacity. Data protection authorities, research ethics committees and sector regulators need data engineers, security specialists, statisticians and AI auditors—not only lawyers. Likewise, technical teams need training in ethics, human rights and public administration. Effective governance is multidisciplinary.
Fifth, affected communities must participate in design and oversight. Africa cannot achieve data sovereignty merely by storing information within national borders. Sovereignty also requires African institutions and communities to shape the purposes, standards and value derived from their data.
From compliance to public trust
The debate about sensitive data is often framed as a contest between protection and innovation. That framing is mistaken.
Weak governance may enable rapid deployment, but it also produces breaches, unreliable analytics, discriminatory outcomes and public resistance. Once trust is lost, even socially valuable data initiatives become harder to sustain.
Good governance does not mean eliminating all risk. Clinical research itself would be impossible under such a standard. It means identifying risk, reducing it, assigning responsibility and ensuring that people have recourse when systems fail.
Across clinical trials and digital platforms, the underlying duty remains the same: institutions must be able to explain why they hold sensitive data, how it serves the people represented in it, who may use it, how harm is prevented and who will answer when something goes wrong.
Africa has an opportunity to build a governance model suited to its own realities—one that recognises infrastructure constraints and urgent development needs without accepting weaker protection for African people.
The most successful digital future will not belong to the institutions that collect the most data. It will belong to those that earn the greatest trust—and can prove that the intelligence extracted from data produces public value without sacrificing human dignity.
From Clinical Trials to Digital Platforms: A Risk-Based Approach to Governing Sensitive Data in Africa.
Drawing on experience in clinical research, data engineering and digital platform development, this article proposes a risk-based framework for governing sensitive data in Africa—one that protects individual rights, strengthens public trust and enables responsible innovation.

Share this article


Conversation
0 comments
Be the first to contribute to this conversation.