Ever wonder what people really do with their AI chatbots? Not just the work stuff, but the late-night confessions, the relationship advice, the… other things?
Turns out, the big AI companies like Anthropic and OpenAI are pretty selective about what they share. They release reports on how people use Claude and ChatGPT, but researchers have been side-eyeing the data, noting there's no independent way to verify it. Because, naturally, companies tend to highlight the most flattering bits.

The AI Observatory Lifts the Veil
Enter the AI Observatory, a new project co-led by Stanford's Anka Reuel. This public platform is doing the heavy lifting, gathering and analyzing actual conversations people are having with popular AI models like Claude and Gemini. They're not just guessing; they're working with user-consented data from seven existing datasets.
We're a new kind of news feed.
Regular news is designed to drain you. We're a non-profit built to restore you. Every story we publish is scored for impact, progress, and hope.
Start Your News DetoxWhy does this matter? Because crucial decisions about AI's future — its benefits, its risks, its potential for utter chaos — are being made with what Reuel calls "very limited data." Basically, we're flying blind, guided by corporate press releases.
And what the Observatory found is, shall we say, enlightening. People are using AI for far more sensitive behaviors than what the official company reports let on. Those reports, it seems, are a bit too focused on work and productivity, conveniently filtering out the messier, more human stuff.
Take Anthropic's Economic Index, a widely cited source on Claude AI usage. It focuses squarely on work-related tasks. But when the AI Observatory applied Anthropic's own filters to their independent dataset, nearly half (48%) of the conversations vanished. Poof.
What was in that missing 48%? A whole lot of:
- Health and relationships: 44.2% (vs. Anthropic's 31.2%)
- Adult or illicit topics: 7.9% (vs. Anthropic's 2.1%)
- Harassment and hate: 27.5% (vs. Anthropic's 5.66%)
- Sexual content: 16.7% (vs. Anthropic's 2.4%)
Let that sink in. The gap between what's reported and what's actually happening is… significant. OpenAI's own report on ChatGPT found only 30% of consumer use was work-related, which still leaves a lot of room for the weird and wonderful.
Chatbots Get Chatty (and Grok Gets Political)
The Observatory also tracked how AI use evolves over time and across different models from 2023 to 2025. Conversations are getting longer, more detailed, and, perhaps most interestingly, there's been a noticeable increase in "small talk." People are turning to AI for companionship, and the bots are getting better at pretending not to be bots.
Good news: "sensitive use" (think hate speech, harassment) has decreased, suggesting those safeguards might actually be working. Which is, you know, a relief.
But the models themselves have personalities:
- Grok and Gemini are the info-seekers. Grok, in particular, is the go-to for news and politics — and, naturally, a higher concentration of misinformation. Because apparently that's where we are now.
- Anthropic is for the coders.
- Gemini is for social interactions and roleplay. Your AI Dungeon Master, if you will.
- ChatGPT is the homework helper. (Sorry, teachers.)
Even within the same brand, there are differences. GPT-3.5 gets you short chats, while GPT-4o leads to those long, emotionally attached conversations. As Shayne Longpre, co-lead of the research, dryly notes, "No single company report tells the whole story." You don't say.
The Great Data Divide
The AI Observatory's dataset is impressive: 24,521 conversations, 85,633 user-AI turns, involving 5,000 users and 52 different models. But it's a drop in the bucket compared to the millions of conversations the big AI labs are sitting on. Anthropic analyzed 1 million Claude chats; OpenAI, 1.5 million ChatGPT ones. And they're not exactly keen on sharing.
An Anthropic rep said their research reflects their interests, and they support external research. OpenAI, meanwhile, ghosted requests for comment. Which, again, tells a story of its own.
This data gap means independent researchers are essentially trying to piece together a puzzle with half the pieces missing. Without a clear, unbiased picture of how AI is actually being used, policymakers and researchers are, as Reuel puts it, "completely operating in the wild." And that's a pretty wild place to be when you're shaping the future.










