Wednesday, September 16, 2026

The Researcher’s Promise

 

The Researcher’s Promise

On Anonymity, Trust, and the Limits of Open Science

Imagine a researcher arriving in a small Norwegian rural community to interview five people with drug-related problems.

Or a researcher approaching ten homeless people in a small town.

They are being asked to speak about their lives. About addiction. Family. Illness. Perhaps crime, shame, violence, or betrayal. Some of what they disclose may never previously have been shared with anyone.

The researcher explains the project, provides an information sheet, and asks for consent.

Their names will not be published. The information they provide will be treated confidentially. No one should be able to recognise them in what is later published.

It is easy to write this in an information sheet.

It is much harder to keep the promise.

When Everyone Knows Everyone

Norway is a small country. Many villages and smaller towns are socially transparent communities in which people know one another, or know someone who knows someone.

Suppose a researcher writes about “a man in his fifties, formerly employed at the local factory, with three adult children, who developed a serious substance-use problem following a divorce.”

His name has been removed.

Yet half the village may still know who he is.

If the researcher removes his occupation, age, family situation, and place of residence, he becomes more difficult to identify. But at the same time, parts of the context that made the interview meaningful disappear. Perhaps the loss of his job, the divorce, and life in a small community are precisely what must be understood if his story is to make sense.

This is a well-known challenge in qualitative research. The details that give a narrative meaning may simultaneously make the participant identifiable, and the problem becomes particularly acute in small and closely interconnected communities (Kaiser, 2009; Teufel-Shone & Williams, 2010).

A fundamental problem therefore arises:

The more context we preserve, the better we may understand the person’s experience. Yet the more context we preserve, the greater the risk that the person will be recognised.

Anonymisation is therefore not simply a matter of replacing “Ole Hansen” with “Participant 4.”

The Norwegian National Research Ethics Committee for the Social Sciences and the Humanities (NESH) uses the concept of reverse identification to describe situations in which individuals can be identified through combinations of information even when direct identifiers have been removed. Where research participants have been promised anonymity, NESH emphasises that this constitutes a promise that individuals will not be identifiable in the research or its dissemination (NESH, 2023).

This has long been a recognised problem in qualitative research conducted in small and socially transparent communities.

Recent research, however, shows that the problem extends far beyond the Norwegian village.

327 Research Projects

In 2026, Abel Brodeur and David Valenta presented a study of personally identifiable information in social science research data.

They examined 327 publicly available replication packages associated with articles published between 2020 and 2024 in eleven leading journals in economics, political science, and psychology. The studies were based on individual-level data collected by the researchers themselves through field experiments, laboratory experiments, online experiments, surveys, and related methods (Brodeur & Valenta, 2026).

In 69 of the 327 packages, they found information that could contribute to the identification of research participants. This corresponds to 21.1 per cent.

The authors report three partly overlapping categories: 14.4 per cent of the packages contained direct identifiers, 7.0 per cent contained indirectly identifying information, and 7.0 per cent contained IP addresses. Because the categories overlap, these percentages should not be added together (Brodeur & Valenta, 2026).

The information identified included names, email addresses, telephone numbers, geographic coordinates, and identifiers associated with research platforms. Identifying information was also found in open-text fields in which participants themselves had entered names or contact details.

In one particularly serious case, research participants could be linked to information concerning criminal acts that they had disclosed to researchers in confidence (Brodeur & Valenta, 2026).

These are not minor imperfections in a dataset.

They concern human beings who may become identifiable through information they entrusted to researchers.

At the same time, precision matters. Brodeur and Valenta’s study is a discussion paper, and it is not a study of qualitative interview research.

We therefore cannot conclude that 21 per cent of qualitative research breaches anonymity.

But the study demonstrates something else that is at least as important.

The problem of identification is not confined to small Norwegian communities in which everyone knows everyone else.

It also exists at the centre of international, digital, and open science.

A Name Does Not Have to Be a Name

The problem becomes even clearer when research material consists of language.

A qualitative interview is not a table.

It is a narrative.

People speak about their workplace, their spouse, their children, the hospital in which they were treated, the school they attended, the municipality in which they live, or the event that changed their life.

Each individual characteristic may appear harmless.

Together, they may become a fingerprint.

Kaiser (2009) describes this as a fundamental tension in qualitative research: researchers seek to communicate rich and detailed accounts of human lives, yet the very details that make those accounts meaningful may also make participants recognisable.

The researcher is therefore confronted with a paradox. Remove too much, and we may destroy the context that makes the narrative intelligible. Remove too little, and we may expose the person who told the story.

This is not merely a technical problem.

It is a problem of practical philosophy.

Anonymity and Confidentiality

Our language also requires precision.

In a qualitative research interview, the researcher usually knows who the participant is. The participant is therefore not anonymous to the researcher.

What is often actually promised is confidentiality.

I know who you are, but I will protect what you tell me.

The distinction may appear linguistic. Ethically, it is fundamental.

NESH explicitly distinguishes between anonymity and confidentiality. If a researcher promises confidentiality, this is a commitment that information will be treated in confidence and will not be disclosed in ways that exceed what has been agreed. NESH directly connects the credibility of the researcher and participants’ trust in research to this obligation (NESH, 2023).

For when a researcher promises confidentiality, the issue is not merely how data are processed.

It concerns a relationship between human beings.

And I keep returning to the word promise.

A promise is not simply a clause in an information sheet.

It is something one human being gives to another.

Trust Comes First

K. E. Løgstrup begins The Ethical Demand with trust. Ordinarily, he argues, we encounter one another with a fundamental trust. Precisely because human beings are mutually dependent and vulnerable, we inevitably acquire some degree of power over another person’s life (Løgstrup, 1956).

Løgstrup later developed the idea of the sovereign expressions of life. Trust belongs among them, together with phenomena such as mercy and openness in speech. They are not primarily rules that we decide to follow. They arise spontaneously in human life and make coexistence possible (Løgstrup, 1972).

The qualitative research interview illustrates this in an almost concentrated form.

The participant does not merely provide information.

He confides.

She tells.

Something of a life is placed in another person’s hands.

It is here that Løgstrup’s ethical demand becomes particularly relevant. The demand arises because, through addressing us, the other person has entrusted something to us. We may use it well or badly. We may protect it or expose it.

The participant may know less than the researcher about data repositories, replication packages, metadata, and future possibilities for linking different sources of data.

Yet the participant must dare to trust the researcher.

There is an asymmetry in this relationship.

The researcher gains knowledge that the other person makes available.

And for precisely that reason, the researcher acquires responsibility.

Trust is not an addition to research.

It is often a precondition for the production of knowledge itself.

I Have Also Experienced This from the Other Side of the Table

Through many years of work as a social worker and therapist, I have sat in conversations in which people have spoken about matters they could not readily disclose to others.

Such encounters teach one something about confidentiality that no legal provision, by itself, can teach.

Confidentiality is, of course, concerned with law, rules, and professional obligations. But for the person on the other side of the table, it also concerns something more elementary:

Can I tell you this without it later being used against me?

That is a question the researcher must also be able to answer.

Not only on the day the consent form is signed.

But when the interview is transcribed.

When the material is analysed.

When the article is written.

When the data are archived.

And when, perhaps many years later, someone wishes to use them again.

The Human Being Must Not Become Raw Material

Here, Kant’s formulation of the human being as an end in itself acquires a very concrete meaning.

In the Groundwork of the Metaphysics of Morals, Kant formulates his humanity principle: a human being must never be treated merely as a means, but always at the same time as an end in themselves (Kant, 1785/2012).

Research participants contribute data.

But research participants are not data.

This can be easy to forget once an interview has been transcribed. The voice becomes text. The person becomes “Participant 7.” The story is coded and distributed across analytical categories.

The man from the village who told the researcher about his divorce and substance-use problems may, several years later, have become a file on a server.

But he is still the man who told the story.

That does not disappear because the researcher has assigned him a number.

The question “Can this dataset be shared?” should therefore not come first.

Another question must precede it:

What did I promise the person who told me this?

Rules Are Necessary, but Not Sufficient

Aristotle called practical wisdom phronesis. It concerns the capacity to judge how one ought to act in concrete situations in which general rules alone cannot provide the answer (Aristotle, 2009).

This is particularly evident in qualitative research.

No general rule can tell a researcher in advance exactly how much must be removed from the story of the man with substance-use problems in the small village.

It requires knowledge of the place.

Of the social environment.

Of the people.

Of the material.

And of what may happen if the person is recognised.

The same description may be entirely unproblematic in a large city and identifying in a village with only a few hundred inhabitants.

Practical wisdom consists precisely in being able to perceive such differences.

Research ethics can therefore never be reduced to the question:

Have I followed the rules?

We must also ask:

Have I understood the situation?

As Open as Possible

The problem has acquired a new dimension through the movement towards open science.

There are strong arguments for sharing research data.

Other researchers should be able to examine analyses. Errors should be detectable. Results should be replicable. Existing material may contribute to new knowledge.

These are important scientific values.

But “open” is not synonymous with “ethically right.”

Brodeur and Valenta’s findings demonstrate precisely this tension. Replication packages were made public in order to strengthen openness and the verifiability of science. Yet the same openness simultaneously made identifying information about research participants publicly available (Brodeur & Valenta, 2026).

For qualitative data, the tension becomes particularly pronounced. Kvale, Pharo, and Darch (2023) show how the sharing of qualitative interview data involves a difficult balance between reuse and openness on the one hand, and privacy, context, and participants’ own understandings of what should be shared on the other. Strong anonymisation may protect participants, but it may also remove so much context that the scientific value of the material is diminished.

Two legitimate values may therefore come into conflict.

Openness protects the verifiability of science.

Confidentiality protects the person who made the research possible.

There is no universal rule according to which one must always take precedence over the other.

Open science, too, therefore requires practical wisdom.

Who Receives the Benefit—and Who Bears the Risk?

The benefits of data sharing largely accrue to research.

Results can be checked. New analyses can be conducted. Studies can be replicated. Knowledge can grow.

But who bears the consequences if anonymisation fails?

It may be the man with the substance-use problems.

The homeless woman.

The young person who disclosed violence at home.

The employee who criticised the workplace.

The patient who spoke about an illness no one else knows about.

Here John Rawls can help us through a thought experiment.

Imagine that we were to design the rules governing data sharing from behind his veil of ignorance. We would not know which position we ourselves would occupy. We would not know whether we would be the researcher seeking access to the material or the vulnerable research participant whose story is contained within it (Rawls, 1999).

What rules would we choose?

Rawls did not develop the veil of ignorance as a theory of research data. But the thought experiment forces us to consider the distribution of benefit and risk from a position in which we do not already know who we are.

The question then becomes more demanding than “Could these data be useful to other researchers?”

We must also ask:

Would I accept this risk if the story were mine?

A Promise That Endures

There is a temptation to regard consent as the conclusion of the ethical problem.

The participant has signed.

The ethical requirements have been satisfied.

But consent does not bring responsibility to an end.

It begins it.

A person consenting today cannot possibly foresee all the technologies, databases, and possibilities for data linkage that may exist ten or twenty years from now.

Information that appears minimally identifying in isolation may, when combined, make individuals recognisable. Both research on confidentiality in qualitative studies and NESH’s discussion of reverse identification demonstrate why anonymisation must therefore be understood as an ongoing ethical judgement rather than as a one-off technical procedure (Kaiser, 2009; NESH, 2023).

Brodeur and Valenta (2026) also emphasise that their estimate of indirect identification is likely to be conservative. It is difficult to investigate in advance every possible combination of variables against every external source of information to which a future user may have access.

A promise of confidentiality cannot therefore be understood as a single act performed at the beginning of a research project.

It must be honoured throughout the life of the material.

From Thou to It

Martin Buber distinguishes between two fundamental ways of relating to the world: I–Thou and I–It (Buber, 1923/1970).

Research requires the I–It relation.

We must be able to analyse, categorise, compare, and systematise. Without a degree of objectification, there can be no science.

The problem arises when the It becomes all that remains.

When the person who once sat before us has become nothing more than a dataset.

Buber’s point is not that we must never relate to another person as an object of knowledge. That would make both science and everyday life impossible. Rather, a human being cannot be exhausted by this relation. Behind the category, the variable, and the interview excerpt there remains a Thou—a person who cannot be reduced to what we know about them.

Three different philosophical traditions converge here: Løgstrup’s ethical demand, Kant’s requirement that the human being never be treated merely as a means, and Buber’s distinction between Thou and It. Each points towards the same insight: a human being can never be exhausted by the information we collect about them.

This may be the deepest ethical reminder of all.

The researcher did not first encounter “Participant 7.”

The researcher encountered a human being.

The dataset came later.

The Researcher’s Promise

Brodeur and Valenta’s study shows that personally identifiable information has found its way into publicly available research data far more often than we might have expected.

But it would be too simple to turn this into a story about careless researchers.

The problem is larger.

Modern research has made the relationship between openness and confidentiality more difficult.

More can be stored.

More can be linked.

More can be shared.

And therefore more can also be exposed.

We need better technical solutions. Better checking of datasets before publication. More secure repositories. Restricted access where open data create unreasonable risks. Better information for research participants. Research on the sharing of qualitative data points precisely towards the need for arrangements that take account both of scientific reuse and of participants’ understanding of what they have entrusted to researchers (Kvale et al., 2023).

All of this is necessary.

But no technical solution can replace the ethical question that comes first.

At some point, a human being sat before the researcher and spoke.

Perhaps hesitantly.

Perhaps about something very few other people knew.

That person was not first and foremost an informant, a unit of analysis, or a data point.

That person was a Thou.

The researcher came to know something because another human being dared to trust that the story would be received responsibly.

That is where research ethics begins.

Not in the dataset.

Not in the repository.

Not when the article is submitted to a journal.

It begins in the encounter between human beings.

And it does not end when the encounter is over.

For the researcher’s promise follows the story onwards.

Perhaps, then, the first requirement of good research is not that we share everything we know.

It is that we prove ourselves worthy of the trust that made the knowledge possible.

References

Aristotle. (2009). The Nicomachean ethics (D. Ross, Trans.; L. Brown, Rev.). Oxford University Press. https://doi.org/10.1093/actrade/9780199213610.book.1

Brodeur, A., & Valenta, D. (2026). On the prevalence of personally identifiable information (PII) in the social sciences (I4R Discussion Paper Series No. 320). Institute for Replication. https://hdl.handle.net/10419/343373

Buber, M. (1970). I and thou (W. Kaufmann, Trans.). Charles Scribner’s Sons. (Original work published 1923)

Den nasjonale forskningsetiske komité for samfunnsvitenskap og humaniora. (2023). Forskningsetiske retningslinjer for samfunnsvitenskap og humaniora [Research ethics guidelines for the social sciences and the humanities] (5th ed.; originally published 2021, updated 2023). https://www.forskningsetikk.no/retningslinjer/hum-sam/forskningsetiske-retningslinjer-for-samfunnsvitenskap-og-humaniora/

Kaiser, K. (2009). Protecting respondent confidentiality in qualitative research. Qualitative Health Research, 19(11), 1632–1641. https://doi.org/10.1177/1049732309350879

Kant, I. (2012). Groundwork of the metaphysics of morals (M. Gregor & J. Timmermann, Trans.; 2nd ed.). Cambridge University Press. (Original work published 1785)

Kvale, L. H., Pharo, N., & Darch, P. (2023). Sharing qualitative interview data in dialogue with research participants. Proceedings of the Association for Information Science and Technology, 60(1), 223–232. https://doi.org/10.1002/pra2.783

Løgstrup, K. E. (1956). Den etiske fordring [The ethical demand]. Gyldendal.

Løgstrup, K. E. (1972). Norm og spontanitet: Etik og politik mellem teknokrati og dilettantokrati [Norm and spontaneity: Ethics and politics between technocracy and dilettantocracy]. Gyldendal.

Rawls, J. (1999). A theory of justice (Rev. ed.). Belknap Press of Harvard University Press.

Teufel-Shone, N. I., & Williams, S. (2010). Focus groups in small communities. Preventing Chronic Disease, 7(3), Article A67.


Perhaps, then, the first requirement of good research is not that we share everything we know.

It is that we prove ourselves worthy of the trust that made the knowledge possible.



This essay was written in a conversation with Claude/Anthropic and OpenAI/ChatGPT 


No comments:

Post a Comment