Research

AI-generated text

Independent dataset shows big AI firms' usage reports omit many sensitive uses

A team led by Stanford and MIT researchers built the AI Observatory, an independent platform aggregating 24,521 real user–AI conversations from 2023–2025 to analyze how people use models like ChatGPT, Claude, Gemini and Grok.

Independent dataset shows big AI firms' usage reports omit many sensitive uses

The AI Observatory, a new research effort led by scholars from Stanford and MIT and involving the Data Provenance Initiative and others, assembled an independent dataset of real user–AI conversations to analyse how people interact with popular models such as ChatGPT, Claude, Gemini and Grok. The platform aggregates conversations collected with users’ consent from seven existing datasets.

The project’s stated aim is to give researchers and policymakers independent information about generative AI usage, since decisions about AI’s benefits and risks are often made using limited or proprietary company statistics, says Anka Reuel, a Computer Science PhD candidate at Stanford’s Trustworthy AI Research (STAIR) Lab and co‑lead of the Observatory.

Key findings

The AI Observatory analysed 24,521 conversations comprising 85,633 conversational turns (prompt + response) from about 5,000 users interacting with 52 different models between 2023 and 2025. Main findings include:

  • Usage varies substantially across models in topics, interaction styles and the prevalence and types of sensitive use cases.
  • Over time, some datasets (notably WildChat) showed longer and more elaborate conversations — increases in prompt tokens, response tokens and conversation turns — and a rise in small talk alongside reduced AI self‑disclosure (the assistant stating it is a chatbot), suggesting growing ‘‘AI companionship.’’
  • The share of exchanges labelled as sensitive (potentially harmful or restricted content, including sexual harassment and hate speech) declined over the study period, which may reflect more effective platform safeguards.

Differences versus company reports

The Observatory points out that prominent corporate reports can omit large portions of non‑work use. The Anthropic Economic Index, for example, focuses on work‑ and productivity‑related uses of Claude and filters out unrelated conversations. When the Observatory applied Anthropic’s filtering method to its own dataset, it found that 48% of the conversations would have been excluded.

Conversations that would have been filtered out by Anthropic’s approach were more likely to contain health and relationship topics (44.2% vs. 31.2% in Anthropic’s analysis), adult or illicit topics (7.9% vs. 2.1%), harassment and hate (27.5% vs. 5.66%) and sexual content (16.7% vs. 2.4%). OpenAI’s 2025 report similarly found that only about 30% of consumer ChatGPT use was work‑related.

David Widder, assistant professor at the University of Texas at Austin’s School of Information (not part of the Observatory), noted that while Anthropic has published separate posts on use for support, companionship and even CSAM generation, the Observatory’s consolidated, ‘‘bird’s‑eye’’ analysis helps researchers see the variety of uses more consistently.

Model‑specific patterns

The Observatory found model‑specific concentrations of use. Grok and Gemini were more often used for information retrieval; Grok in particular was frequently used for news and political information and was also a locus of misinformation, consistent with other research. (xAI did not respond to a request for comment.)

Other patterns emerged: Anthropic models were relatively common for coding tasks; Gemini for social and roleplay interactions; and ChatGPT for homework assistance. The study also found intra‑model differences: ChatGPT conversations powered by GPT‑3.5 tended to be shorter, while GPT‑4o produced longer, more iterative conversations — a pattern that aligns with previous observations about GPT‑4o’s role in fostering stronger emotional engagement.

Limits and implications

The Observatory’s dataset is still limited. It comprises voluntarily contributed conversations and therefore likely underrepresents especially sensitive uses that people are reluctant to share. Moreover, the 24,521 conversations are small compared with the proprietary datasets available to large labs: Anthropic’s Economic AI Index is based on 1 million Claude conversations, and OpenAI’s ChatGPT usage report analysed 1.5 million conversations.

Anthropic said its published research reflects its teams’ specific questions and interests and emphasized the importance of supporting external independent research. OpenAI did not respond to requests for comment.

Reuel and her co‑researchers plan to keep expanding the Observatory’s dataset and make it available to researchers for independent analysis. Ideally, they say, large AI firms would share data with independent researchers in privacy‑preserving ways. Until then, the team warns, policymakers and other stakeholders risk making consequential decisions relying on incomplete company narratives rather than a fuller picture of how generative AI is used.