At the heart of digital mental health is trust.
Most people would not consider it acceptable to show a recording of a therapy session at a conference simply because a face had been blurred and a name removed. They would want to know whether the person had agreed to that use of the recording in the first place. The concern isn’t only whether the person could be identified. It’s whether something disclosed within a therapeutic relationship is being used beyond the purpose for which it was shared.
The person is still there. So is the therapeutic relationship.
Something similar is now happening with data. A de-identified record may no longer identify a person directly, but it can still carry what was learned inside a therapeutic relationship. The issue is not only whether the individual remains identifiable. It is whether the obligations created by that relationship continue to travel with disclosures once they are linked, analysed, transformed, and reused.
For those implementing digital mental health services, platforms, clinical programmes, and national initiatives, this is not an abstract concern. Continuity across settings, records that follow a person, and analytics that support better care are all worth building. The argument here is not against that direction of travel. It is that one governance question remains underdeveloped, and it will be easier to answer now than after these systems are embedded.
The Bargain We Struck
Think about a note written last Tuesday. It sits in a platform that is encrypted, access-logged, hosted where the service requires, and audited. Nobody has breached anything.
A de-identified version of that note may still be released, linked, reused, or used in model training under terms the organisation agreed to. Removing the identifiers can lift many of the rules that protected it.
De-identification made sense as a legal and governance threshold when records were relatively bounded. A form, a discharge summary, or a table of values could be stripped of direct identifiers and treated differently because the remaining object was still close to the original record and purpose.
That is not a failure of those laws, or of the organisations working within them. The statutes govern personal health information while a person remains identifiable, and they were built to do that well. The question begins where that condition ends: not with compliance, but with public interest and public trust.
When Records Become a Representation
Consider one person in crisis over a single week. A police wellness check, an emergency presentation, a crisis team contact, a pharmacy record, a therapy session, and a benefits claim.
Six organisations. Multiple custodians, legal authorities, and consent processes.
Two things then happen to those records. Interoperability links them: the standards, platforms, and agreements that allow organisations to recognise that separate encounters belong to the same person and join them into one continuous history. Artificial intelligence infers from them: it reads across what has been linked, turns narrative into structured summary, and produces statements and risk scores that appear nowhere in the source material.
Together they produce something else: a substrate from which claims about a person can be generated that no individual event supports and no person authored.
A file records what happened; you can read out of it what was put in. But once records are linked and interpreted across settings, they begin to constitute a different kind of object: not only a file about a person, but a representation of one.
For the person’s care, this is largely how the system should work. Continuity across services is the point of connecting records, and each organisation had a lawful basis for what it held. At this stage, nothing has gone wrong.
The question arises when what has been learned about a person is later used for something else.
The position paper When health records become representations of people calls this emerging object a Digital Subject Representation, or DSR: a de-identified but composed representation of a person, assembled across time, setting, or source, and containing structured, unstructured, and derived content from which behavioural or predictive claims can be generated.
Where the Concern Lies
The concern isn’t that such representations exist. They’re often the product of exactly the kinds of continuity, coordination, and insight health systems are trying to achieve. The question is the same as the one posed by the therapy recording: what obligations remain when something disclosed within a therapeutic relationship is used beyond the purpose for which it was shared?
De-identification is what can release the representation from the rules that applied to identifiable records. Composition is what makes the representation valuable enough to reuse.
Not all secondary use raises this. De-identification is not anonymisation: removing the identifiers leaves one person’s record intact, while anonymisation means no individual can be picked out at all. Where representations are combined so that no individual trajectory can be retrieved, the result is knowledge about a population, and most of these concerns fall away.
The concern here is the other case. The identifiers are removed, but each record remains one person’s composed history, and models are trained on them one whole history at a time. A composite is knowledge derived across people. A training corpus of Digital Subject Representations is a collection of people.
A Digital Subject Representation can be used, trained upon, or commercialised in ways that affect the person it describes without their knowledge, consent, or recourse.
And at that point, which consent governs? None of them reaches the composed object.
Secondary Confidentiality
Ordinary confidentiality asks one question: can anyone work out who this person is?
Secondary confidentiality asks another: what can legitimately be done with what has been learned about them?
Once the object in question is a DSR rather than a bounded record, that second question is the one that matters, and existing frameworks are not built to answer it.
De-identification removes the identifier. It doesn’t necessarily remove the effects of what has been learned.
What the Field Is Already Doing
Serious governance work is already underway. The European Health Data Space, OECD guidance on secondary use, WHO guidance on AI governance, and Indigenous data sovereignty frameworks all address important parts of the problem.
Yet these frameworks are generally triggered by access, use, purpose, governance arrangements, or AI systems themselves. Few are explicitly triggered by the formation of a Digital Subject Representation.
As a result, obligations relating to composition, derivation, and longitudinal representation can remain unclear. The expectations attached to what people disclosed become harder to trace once information is linked, transformed, and reused. That is the governance gap explored here.
What Would Need to Change in Governance
If the obligations created by disclosure are meant to survive beyond the original encounter, governance needs a way to recognise them, carry them forward, and assign responsibility for them.
The DSR position paper proposes four starting points for discussion:
Triggers: governance review when records are linked, purposes change, or new assertions are derived.
Inheritance: derived outputs carry the obligations attached to the records from which they were generated. A risk score is a new assertion, not a copy of the record it came from, and without this rule the obligations simply do not follow it forward.
Carriage: governance conditions travel with the representation as it moves between systems.
Accountability: naming who answers for a representation that emerges across multiple custodians and platforms.
The implication is proportionate rather than prohibitive. Research, secondary use, and model development all remain possible. What changes is that de-identification alone would no longer settle whether a holding is low risk.
What Happens If We Get This Wrong
The risk here is not primarily legal. It is that people discover, after the fact, that deeply personal disclosures can be carried into linked, derived, or de-identified representations in ways they did not expect.
People disclose intensely personal information in mental health care because they believe the therapeutic relationship creates obligations around what happens to those disclosures. When those expectations turn out not to extend to how their digital subject representations are used, trust becomes harder to sustain.
England’s care.data programme illustrates the broader risk. More than a million people opted out after sustained public concern. The transfers may have been lawful and the data de-identified, but legality did not resolve the trust question.
Innovation in mental health depends on disclosure. Disclosure depends on trust. Protecting that trust is what makes the rest of the work possible.
What a Service Can Do Now
None of this requires waiting for policy.
Services already making decisions about interoperability, analytics, AI, and data sharing can begin with a practical question: would these decisions align with what service users would reasonably expect to happen with the information they disclose?
Holistic Research Canada’s Clinical Data Governance Checklist provides a practical starting point and is freely available to use and adapt. The position paper When health records become representations of people develops the full argument. Related governance tools, position papers, and research projects are openly shared through the Clinical Data Governance community on Zenodo.
Current work includes Illustrating Emerging Personal Health Data Governance Challenges in AI-Enabled Health Data Ecosystems: An International Scenario-Based Study Using the Clinical Data Governance Checklist. The proposal would benefit from replication across a wider range of health systems, particularly in the Global South and other jurisdictions beyond Western high-income democracies.
The governance questions discussed here are unlikely to have a single answer. We will need to work them through across different cultures, legal systems, technologies, and models of care.
Conclusion
De-identification can remove the identifier while leaving a representation of the person intact.
Mental health care depends on people sharing information they may tell no one else. Practitioners treat informed consent as a sacred trust because disclosure is only possible when people understand and accept what will happen next. The question is not whether de-identification can remove a name. It is whether the expectations created in care disappear when information is de-identified, linked, and reused. That is where trust is built, and where it is lost.

