Measuring Shame
When Numbers Met Lived Experience
Many years before I tried to measure shame, I had attempted to measure the views of human nature and society held by fifty employees in the child welfare service where I worked.
I had devised questions intended to place them along certain fundamental oppositions. Were they voluntarists or determinists? Did they primarily regard the human being as free and capable of action, or as shaped by forces beyond individual control? Was their view of society characterised by confidence in laws and regulations, or were they revolutionaries who believed that society needed more fundamental change?
The answers were collected, organised and calculated.
The result was a view of human nature measuring 1.2 and a view of society measuring 2.3.
I showed the figures to my supervisor, John Lundstøl. He laughed heartily.
“These are numbers without meaning,” he said.
That was his brief reply.
I had carried out the calculations correctly. The numbers were not arbitrary. They could be traced back to the answers given by the fifty employees. Yet Lundstøl had seen something I had been unable to resolve through the questionnaire: What, exactly, was a view of human nature measuring 1.2? How did a child welfare worker meet a child differently because his or her view of society had been calculated as 2.3?
The numbers were precise. But what did they mean?
A person may believe that human beings possess freedom and responsibility while also recognising that childhood circumstances, poverty, violence and neglect restrict their possibilities. A child welfare worker may be bound by laws and regulations while at the same time believing that society must be changed. Practical work does not take place at the extreme points of a scale, but in the tension between them.
I had measured the answers. It was less certain that I had measured their view of the human being.
Lundstøl’s laughter stayed with me. Many years later, when I found myself surrounded by extensive statistical material on shame, his words returned:
Numbers without meaning.
That Which Hides Itself
Shame is an emotion that makes the human being hide. The eyes are lowered, the words become fewer, the voice softer. Some cover their face with their hands. Others smile in a way that does not quite correspond to what they are telling. Shame often reveals itself precisely by trying to make itself invisible.
Yet I decided to measure it.
This may sound contradictory. How can we measure something that withdraws the moment it is seen? How can we assign a number to an experience that often lacks words?
Height can be measured with a measuring tape and weight with a scale. Shame has no comparable extension or weight. It cannot be observed directly. It must be approached through what a person says, does or imagines that he or she would do in particular situations.
When I began my exploration of shame, measurement nevertheless seemed a necessary part of the work. I did not want to approach shame only philosophically or through people’s stories. I also wanted to investigate whether it could be identified as a distinct emotional disposition.
Were some people more inclined than others to react with shame? Could shame be distinguished from guilt, detachment, externalisation and pride? Could such differences be expressed in numbers?
I used the TOSCA-3, the Test of Self-Conscious Affect, developed by June Price Tangney and her colleagues. The test consists of sixteen everyday scenarios. Respondents are asked to imagine that they have forgotten an appointment, broken something at work, made a mistake for which a colleague is blamed, or spilled red wine on a light-coloured carpet without anyone noticing.
Each scenario is followed by several possible reactions. Some are intended to express shame, others guilt, detachment, externalisation or pride.
Respondents do not choose only one reaction. Each possibility is rated on a scale from one to five. The test therefore allows for the possibility that a person may react in several ways at once. One may feel both guilt and shame, wish to make amends and at the same time want to disappear.
At first, this seemed reasonable. Emotional life is complex, and the test attempted to take this complexity into account. Yet the problems began as soon as it was translated from English into Norwegian.
Translating an Emotion
Translating a psychological measurement instrument is not simply a matter of finding Norwegian words that roughly correspond to the English ones. One must also try to preserve the emotional meaning of each response.
A statement that appears self-evident in one cultural setting may seem strange in another.
In one scenario, the person had broken something at work and then tried to hide it. Several of those who tested the Norwegian translation reacted to the situation itself. Breaking something by accident might be common enough, but was it equally common to hide it? A more likely Norwegian response might be to repair it or report what had happened.
In another scenario, someone had forgotten a lunch appointment with a friend and did not realise it until five o’clock. One response was: “I am inconsiderate.” But what did the English word inconsiderate mean in this context? Thoughtless, insensitive or lacking in consideration? The choice of one Norwegian word rather than another could influence which emotion the respondent believed he or she was rating.
The problem became even clearer in the phrase:
“I would feel small—like a rat.”
This is not a common Norwegian expression. Should the rat be retained because it was part of the original test? Or should only the feeling of being small be expressed? And if so, would it still be the same question?
The work of translation thus revealed something fundamental: Emotions do not exist independently of language and culture. They are expressed through images, metaphors and ideas that have developed within particular forms of life.
When a word is translated, the experience itself may also shift.
What I was attempting to translate, therefore, was not merely a test. I was trying to translate possible human reactions.
The Numbers
I conducted two surveys. The first included 221 students enrolled in health and social work programmes. The second included 180 adult users of Norwegian centres for people who had experienced incest and sexual abuse. Of these, 165 reported that they had been sexually abused as children.
I expected clear differences between the two groups.
The sexual abuse of children is surrounded by secrecy, violation and shame. It therefore seemed reasonable to expect people with such experiences to score considerably higher in proneness to shame and guilt than the students.
The results were not so clear-cut.
The women attending the centres did have a higher average shame score than the female students. Overall, however, the differences between the groups were smaller than I had expected. The results also resembled findings from American studies of university students.
This could be interpreted as confirmation that the test worked. The statistical analyses also showed reasonably good reliability for the measurement of shame and guilt. Because the test produced consistent results, it appeared to have captured something stable.
But reliability is not the same as meaning.
A measurement instrument may produce stable results without it being clear what those results express. A clock that is always ten minutes slow is reliable in one sense: it consistently displays the same error. But it does not show the correct time.
The decisive question was therefore not only whether the test measured consistently. The question was whether it measured what it claimed to measure.
Did the shame scale really measure shame?
Once again, I heard Lundstøl’s voice:
What do the numbers mean?
When the Participants Answered in Their Own Words
To investigate the question more closely, I brought the test into the focus groups. The participants had already completed the TOSCA-3. I now invited them to discuss the scenarios.
I wanted to know whether the situations seemed credible and whether the response alternatives expressed reactions they recognised.
The test was thereby moved from the questionnaire into the conversation.
It soon became clear that the participants did not simply accept the emotional categories they had been given. They did not respond only in terms of shame, guilt, detachment or externalisation. They also spoke of sorrow, anger, respect, disappointment, responsibility and the need to put things right.
About one situation, someone said:
It is sad to be forgotten.
Another reacted to the distinction between responsibility and guilt:
Responsibility is not the same as guilt.
Several pointed out that their reaction depended upon the situation. How serious was the mistake? Who was the other person? Had the act been intentional? What kind of relationship did they have with the person who had been affected?
Whenever the test attempted to isolate a single emotional reaction, the participants returned the situation to the context from which it had been removed.
They also said that shame and guilt were difficult to separate:
Feelings of guilt and shame overlap.
And even more clearly:
Guilt and shame are knitted together.
The participants used the six emotions included in the test in ways that were intertwined with one another and with reactions for which the test had no separate categories. The scenarios gave them a safe and indirect way of approaching difficult emotions, yet they found some situations unclear and questioned whether the test really distinguished shame from guilt in the way it claimed.
This did not necessarily mean that the participants were unable to distinguish between emotions. It could just as easily have been the measurement instrument that demanded a clearer distinction than lived experience allowed.
In theory, guilt may be connected to the action:
I have done something wrong.
Shame is more often connected to the self:
There is something wrong with me.
But in a lived life, this boundary is not always clear. Someone who has done something wrong may experience himself or herself as a bad person. Someone who has been subjected to a violation may carry shame for an act committed by another person.
Guilt may move into shame. Shame may clothe itself in the language of guilt. Both may remain hidden beneath anger, silence or withdrawal.
What appeared in the questionnaire as two separate categories could, in experience, be woven together.
When the Numbers Began to Resist
I tried to examine the structure of the test more closely by means of factor analysis. The purpose was to see whether the responses actually clustered around the six emotional factors on which the test was based: shame, guilt, externalisation, detachment and two forms of pride.
I had expected six reasonably clear clusters.
I did not find them.
Among the students, guilt appeared together with externalisation, although with opposite values. Shame appeared together with detachment. The two forms of pride clustered as one factor rather than two.
Among the users of the centres, shame and guilt appeared within the same factor. Here too, the two forms of pride were difficult to distinguish.
In the combined material, shame was likewise not a clearly defined and independent factor. It appeared in connection with other reactions.
The statistical material therefore seemed to confirm what the participants had already said in the conversations: The emotions did not necessarily occur separately. They were entangled.
June Price Tangney, who had helped develop the test, pointed out in an email that traditional factor analysis was not well suited to the TOSCA-3. The response alternatives were nested within the individual scenarios. A more complex confirmatory analysis would therefore be required.
This was an important methodological objection. The structure of the test could explain why the factor analysis had not produced the clear results I had expected. My analysis could therefore not, by itself, determine whether the test measured what it claimed to measure.
But the objection did not remove the fundamental question.
The fact that a test is difficult to examine statistically does not in itself mean that its categories correspond to lived experience. In the dissertation, I therefore had to leave the question of the test’s construct validity open and continue the inquiry through interviews with people who had their own experiences of shame.
What Had I Measured?
I had numbers for shame. I could calculate means, standard deviations, correlations and reliability. I could compare women and men, students and people with experiences of sexual abuse. I could examine how shame varied together with guilt, detachment and pride.
Yet the more analyses I conducted, the more insistent the question became:
What did the numbers actually show?
Perhaps they showed a tendency towards self-criticism. Perhaps they measured a wish to hide, a feeling of inadequacy or fear of other people’s judgement. Perhaps they measured several of these experiences at once.
Just as I had once arrived at a view of human nature measuring 1.2, I could now calculate a person’s proneness to shame. The calculation was more sophisticated, the instrument more thoroughly tested and the statistical analyses considerably more advanced. Yet the philosophical question remained the same:
What kind of reality lies behind the number?
An average cannot tell us how a person experiences himself or herself when shame strikes. A correlation may show that shame and guilt vary together, but not how they are woven together within a particular life. A reliability coefficient may show that the responses possess a degree of internal consistency, but not whether the person recognises himself or herself in the category the researcher has assigned to the answer.
This was what Lundstøl had seen so many years earlier.
A number does not acquire meaning by itself. It must be returned to the experience from which it was derived.
A Measurement Instrument as an Entrance
I did not conclude that the measurement was worthless. The numbers revealed patterns that deserved investigation. They showed, among other things, a stronger relationship between shame and guilt among the users of the centres than among the students. They made it possible to formulate questions that might otherwise never have been asked.
Measurement was therefore not the end of the inquiry. It became an entrance.
When the participants discussed the scenarios, they could approach difficult emotions indirectly. They did not immediately have to tell others about their most painful experiences. They could begin by talking about a forgotten appointment, a broken object or a mistake at work.
The ordinary scenario created distance. That distance offered a degree of safety.
At the same time, the participants began to fill the constructed situations with their own experiences. They corrected the questions. They contradicted the response alternatives. They added new emotions and new moral meanings.
In this way, the test itself became part of the conversation it had been intended to measure.
It did not merely show how the participants reacted. It also revealed how difficult it is to reduce a person’s emotional life to predetermined categories.
The measurement had therefore not failed. But perhaps its value lay somewhere other than where I had first expected. It did not provide a final answer to what shame was. It led me towards the questions that had to be explored through conversation.
From Measurement to Encounter
Perhaps the greatest value of measurement does not lie in its ability to give us a final answer. Perhaps its value lies in revealing the limits of what can be counted.
When the numbers no longer correspond to people’s descriptions, the researcher faces a choice. He can defend the instrument and regard the discrepancies as noise. Or he can listen to the resistance and ask whether the categories themselves need to be reconsidered.
For me, this resistance led onwards to the interviews.
I had to move from the question:
How much shame does this person have?
to the question:
How is shame experienced, and what meaning has it acquired in this person’s life?
This was not a movement from science to ignorance. It was a movement from one form of knowledge to another.
The numbers provided breadth. The conversation provided depth.
Measurement sought differences between shame and guilt. The participants showed how they could be woven together.
The test located the emotions within the individual. The stories showed how shame was connected to relationships, violations, silence and the gaze of others.
The questionnaire asked what the person would probably feel in an imagined situation. The conversation opened towards what the person had actually experienced.
Numbers Without People
Looking back, I believe that both the study of views of human nature and the attempt to measure shame taught me something about the necessary humility of research.
Science needs concepts, categories and methods. Without distinctions, we cannot investigate anything systematically. But the categories must never be confused with the experiences they attempt to describe.
The word shame is not shame itself.
The number is not the feeling.
The score is not the person.
This does not mean that we should stop measuring. But we must understand what measurement can and cannot do. It can reveal relationships, differences and patterns. It can help us formulate better questions. But by itself, it cannot tell us what it is like to live with a shame that has taken root in the body, language and relationships with other people.
John Lundstøl did not dismiss the numbers because numbers are meaningless in themselves. He reminded me that they become meaningless when their connection to human lives disappears.
A view of human nature measuring 1.2 meant little unless I knew how that view was expressed in the encounter with a child.
A shame score meant little unless I knew what the person was ashamed of, whom the shame was connected to and how it had shaped that person’s life.
I began with a wish to make shame visible. I tried to place it in a questionnaire, assign it a value from one to five and examine how it was related to other emotions.
But shame could not be held within a single column.
It appeared in the answers assigned to guilt, in the wish to hide, in the rejection of responsibility, in silence and in the need to put right what had happened. It also revealed itself in what did not fit: in the reservation, the misunderstanding, the protest and the pause.
The numbers had not been useless. They had led me to the boundary of what they could tell me.
On the other side of that boundary sat the human beings.
Only when measurement was interrupted by conversation did the numbers begin to acquire meaning.
The numbers had not been useless.
They had led me to the boundary of what they could tell me.
On the other side of that boundary sat the human beings.
Only when measurement was interrupted by conversation
did the numbers begin to acquire meaning.
This essay was written in a conversation with OpenAI/ChatGPT
No comments:
Post a Comment