In September 2025, OpenAI published research with a team of Harvard-affiliated economists that analyzed 1.5 million anonymized ChatGPT conversations — the largest attempt yet to measure, empirically, what people actually do with a chatbot rather than what they claim to do in a survey. The finding that got the most attention: everyday use, like personal advice, writing help, and general questions, made up a larger share of conversations than office productivity tasks, complicating the assumption that chatbots are primarily workplace tools.

Anthropic ran a parallel effort the same year, mapping millions of anonymized Claude conversations onto the U.S. Department of Labor's occupational task database to see which jobs and tasks actually show up in real usage. Both projects were genuine upgrades over the survey data that had dominated the conversation until then — pollsters asking people whether they've tried ChatGPT, employers asking staff to self-report AI adoption, vendors citing weekly active user counts with no breakdown of what those users were doing with their time.

According to MIT Tech Review, even these usage-log studies leave the central question unanswered: despite two of the best-resourced AI labs in the world running large-scale analyses of their own products, nobody — not researchers, not the labs, not the companies paying for AI subscriptions across their workforce — has a reliable, complete picture of how AI is actually used across the economy. The richest data stays locked inside the companies that generate it.

Why "how people use AI" is so hard to pin down

Three problems compound each other. First, the boundary of "AI use" keeps moving. Is an AI-generated summary at the top of a search results page a case of AI use if the person never typed a prompt? Is autocomplete? Most research only counts deliberate interactions with a named chatbot product, which already misses a growing share of AI-mediated activity that users never consciously register as "using AI."

Second, the two dominant data sources carry opposite biases. Surveys rely on self-report, and self-report is unreliable for anything people feel proud or embarrassed about — employees under-report using AI to cut corners on work, while managers over-report adoption rates to look competitive to their boards. Usage logs solve the honesty problem but introduce a sampling problem: they only capture the slice of behavior that happens to run through one company's product, with no visibility into what the same person does on a competing tool minutes later.

Third, most of the granular data never reaches outside researchers at all. What the public sees is whatever slice a lab chooses to publish, shaped by that lab's own incentive to tell a particular story about adoption.

What the landmark studies cover — and what they skip

Even the best current research has structural blind spots:

What this means for people building AI products

For teams actually shipping AI features, the practical lesson isn't that usage data is worthless — it's that no one else's usage data describes your users:

AiiN's takeaway

The industry is making enormous product and investment bets on adoption curves that nobody can fully verify end to end. Surveys undercount what people are ashamed of; usage logs overcount what one company happens to see; embedded AI features go uncounted by both. In our estimation, this gap likely persists for years rather than closing with the next big study, simply because the incentive to publish a flattering number will always outweigh the incentive to publish a complete one. Until that changes, the most reliable data any builder has is the instrumentation running inside their own product.