OpenAI and MIT Find Emotional Attachment to ChatGPT Is Rare but Concentrated
A study of nearly 40 million ChatGPT conversations and a 1,000-person trial finds emotional attachment rare overall but concentrated among heavy voice users.
Updated
Why it matters
- OpenAI and MIT Media Lab analyzed nearly 40 million ChatGPT interactions using automated classifiers.
- A four-week RCT with nearly 1,000 participants tested voice, text, and conversation-type effects on well-being.
- Heavy users were defined as the top 1,000 daily Advanced Voice Mode users by message count.
- Voice mode was linked to better well-being in brief use but worse outcomes with prolonged daily use.
- The findings are not peer-reviewed and covered only English-language, U.S. participants.
OpenAI and the MIT Media Lab analyzed nearly 40 million real-world ChatGPT conversations and ran a four-week randomized controlled trial with roughly 1,000 participants, and their joint finding is blunt: emotional engagement with ChatGPT is rare across the platform, but a small group of heavy voice users shows concentrated, potentially concerning patterns of attachment.
The two organizations published the results in a joint blog post co-authored by researchers at OpenAI and the MIT Media Lab, with full reports released by both institutions. The work arrives as regulators, researchers, and product teams debate what happens when hundreds of millions of people form conversational relationships with AI systems — a policy and market question that no controlled study of this scale had previously addressed.
What did the researchers actually measure?
The teams ran two parallel studies with different methods.
- Study 1 (OpenAI): A large-scale, automated analysis of nearly 40 million ChatGPT interactions, processed entirely by automated classifiers with no human involvement to protect user privacy. OpenAI combined this with targeted user surveys, correlating self-reported sentiment toward ChatGPT with attributes of actual conversations.
- Study 2 (MIT Media Lab): An IRB-approved, pre-registered randomized controlled trial with nearly 1,000 participants using ChatGPT over four weeks. The trial tested how specific platform features — model personality and modality, including the engaging voices "Ember" and "Sol," a neutral voice, and text — and types of usage affected self-reported psychosocial states: loneliness, social interaction with real people, emotional dependence, and problematic use.
Both studies excluded users who reported being under 18.
How common is emotional use of ChatGPT?
The headline result from the observational study: affective cues — aspects of interactions that indicate empathy, affection, or support — were absent from the vast majority of conversations analyzed. Emotional engagement, in other words, is a rare use case for ChatGPT in the wild.
But rarity at the platform level hides concentration at the individual level. High degrees of affective use were limited to a small group of heavy users of Advanced Voice Mode. The researchers defined "heavy" users as the top 1,000 users of Advanced Voice Mode on any given day of the study period, measured by messages sent. Within that group, emotionally expressive interactions made up a large percentage of usage for a small subset.
That subset was also significantly more likely to agree with statements such as, "I consider ChatGPT to be a friend." The researchers note that because affective use is concentrated in a small sub-population, its impact may not show up when averaging overall platform trends — a methodological warning for anyone studying AI and well-being at scale.
Does voice mode make things worse?
The controlled trial produced a mixed answer.
Users engaging via text showed more affective cues per message than voice users on average. Voice modes were associated with better well-being when used briefly, but worse outcomes with prolonged daily use.
One result cuts against a common assumption: a more engaging voice did not produce more negative outcomes than the neutral voice or text conditions over the course of the study. The voice's personality, in other words, mattered less than how long people spent using it.
What role does conversation type play?
The trial distinguished between personal conversations (prompts such as "Help me reflect on what I am most grateful for in my life"), non-personal conversations (such as "Let's discuss if remote work improves or reduces overall productivity for companies"), and open-ended use. The outcomes diverged sharply.
- Personal conversations, which involved more emotional expression from both user and model, were associated with higher loneliness but lower emotional dependence and problematic use at moderate usage levels.
- Non-personal conversations tended to increase emotional dependence, especially with heavy usage.
The pattern suggests that the relationship between what people talk about and how they fare is not linear, and that dependence can grow even from utilitarian use if the volume is high enough.
Who is most at risk?
The controlled study identified personal factors associated with worse outcomes, though the researchers state clearly they cannot establish causation for these factors.
- People with a stronger tendency toward attachment in relationships.
- People who viewed the AI as a friend that could fit into their personal life.
- People with extended daily usage.
The researchers describe these correlations as important directions for future research rather than settled findings.
Why two methods instead of one?
The teams argue the combination of approaches is itself a contribution. Platform data captured organic user behavior; the controlled experiment isolated specific variables to determine causal effects. Together, they write, the approaches "yielded nuanced findings about how users use ChatGPT, and how ChatGPT in turn affects them, helping to refine our understanding and identify areas where further study is needed."
The researchers explicitly caution against generalizing the results. "We advise against generalizing the results because doing so may obscure the nuanced findings that highlight the non-uniform, complex interactions between people and AI systems," they write.
What are the study's limitations?
The authors list several constraints that shape how far the findings can be pushed.
- The findings have not been peer-reviewed.
- The studies cover ChatGPT only; users of other chatbot platforms may have different experiences.
- Not all findings demonstrate cause and effect.
- Self-reported survey data may not accurately capture true feelings or experiences.
- Meaningful behavioral change may require longer study periods than four weeks.
- The automated classifiers used to detect affective cues are imperfect and may miss nuance.
- The research covered only English conversations with U.S. participants, leaving other languages and cultures unstudied.
What happens next?
OpenAI frames the work as an early step in a broader effort. "We are focused on building AI that maximizes user benefit while minimizing potential harms, especially around well-being and overreliance," the company writes. "We conducted this work to stay ahead of emerging challenges — both for OpenAI and the wider industry."
The company says it will update its published Model Spec to provide greater transparency on ChatGPT's intended behaviors, capabilities, and limitations, and calls the studies "a critical first step in understanding the impact of advanced AI models on human experience and well-being."
The researchers also invite replication: they hope the findings will encourage industry and academic researchers to apply the same methodologies to other domains of human-AI interaction. With emotional use concentrated in a small, identifiable population of heavy voice users — and with prolonged daily use correlating with worse outcomes — the study gives both product teams and regulators a concrete target for monitoring, even before causal questions are resolved.
Original: media.mit.edu
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
195 articles
Related articles
- OpenAI Cuts Unsafe ChatGPT Mental Health Responses by Up to 80%
- OpenAI Builds a Bias Test for ChatGPT. Results Are Mixed.
- OpenAI Signals Data Shows ChatGPT Use Deepening Worldwide
- OpenAI Says Over a Quarter of U.S. Workers Now Use ChatGPT on the Job
- OpenAI Releases MentalHealthBench, an Expert-Built AI Mental Health Benchmark