Anthropic opens real Claude usage data to outside researchers for first time

Three independent teams analyzed around 250,000 privacy-preserved Claude conversations, with Stanford researchers finding more than half involved work that could affect other people or be difficult to undo

A person seated at a desk, focused on a laptop displaying a colorful graphic on the screen.

Anthropic has opened privacy-preserved Claude usage data to external researchers for the first time

Anthropic has given external researchers access to real-world Claude usage data for the first time, opening a part of AI research that has largely remained inside the companies building the models.

Three independent groups at Stanford University, the University of Oxford and nonprofit AI evaluator METR designed studies using aggregated data from roughly 250,000 Claude.ai and Claude Code conversations recorded between April and May 2026.

The researchers never saw the underlying conversations. Anthropic ran their questions through Anthropic Insights, its privacy-preserving analysis system, and supplied aggregated outputs for the teams to analyze independently.

Anthropic says its contractual review rights were limited to user privacy, information that could enable violations of its usage policies, confidential company information and research accuracy. The researchers were otherwise free to publish findings even if they were unfavorable to the company.

The first completed study has already produced a notable result. Stanford University’s Social and Language Technologies Lab found that more than half of the Claude conversations it analyzed involved people delegating “consequential” work, defined as tasks that affect other people or are difficult to reverse.

Professional guidance stood out, particularly legal and financial questions.

People usually remained in charge, but not always in the same way

The Stanford SALT Lab examined how people divide work with AI, the roles users retain and where collaboration becomes difficult.

In nearly three-quarters of conversations, people set the direction and Claude assisted. Users also generally adapted Claude’s output instead of taking it verbatim.

That does not necessarily mean everyone understood the work being produced to the same degree. Researchers found variation in how much users appeared to understand and learn from Claude’s responses even when they were directing the system.

Friction was also common, but it was not automatically a sign that collaboration had failed. The study found that working through misunderstandings could push users to clarify what they wanted, refine the output and remain engaged with the problem.

Two further studies are still underway.

Oxford University’s Human Information Processing Lab is examining how people appear to feel while using Claude and how those states relate to the model’s behavior.

Its early findings suggest patterns appear on both sides of the conversation. Warmer Claude responses appeared alongside more positive user behavior, refusals and disagreements were associated with users pushing back, while more eccentric responses appeared alongside greater intellectual engagement.

Researchers also found that states such as absorption, frustration and enjoyment resembled patterns observed in separate research into everyday internet browsing.

METR, meanwhile, is using Claude Code conversations to investigate productivity gains from coding agents and whether those gains change as models become more capable.

Its preliminary analysis suggests newer models may save users more time than earlier generations, although the study is still being completed.

Independent access comes with methodological limits

Opening real usage data solves one research problem but creates others.

Public AI datasets give outside researchers freedom to ask their own questions, but Anthropic argues they can skew toward casual or creative interactions and may not represent everyday usage accurately.

Anthropic Insights offers access to actual Claude usage patterns without exposing the conversations themselves.

Researchers submit questions such as what type of guidance a person is seeking. Claude evaluates conversations against that question, then groups the results into categories and percentages.

That means the analysis itself still depends on an AI model making judgments.

Anthropic found that small changes in how questions were phrased could alter the categories produced. Because researchers cannot inspect the original conversations, misleading classifications can also be difficult to spot.

External teams initially tested their questions using WildChat, a public collection of human-AI conversations where they could compare the classifications with the original material. Some questions that worked on WildChat performed less reliably on Claude conversations because the datasets contained different patterns of use.

Anthropic says it is now exploring ways for researchers to develop and validate those questions more effectively before future studies begin.

There were also limits on what could be released. Some aggregated categories exposed attempts to violate Anthropic’s policies. Most were retained so researchers could examine misuse, but categories revealing how users bypassed safeguards were removed or altered.

Anthropic says fewer than 5% of categories and conversations were affected in each study, and researchers were told what had been changed and why.

A separate third-party privacy audit of the pilot was carried out by Imperial College London.

Anthropic now wants more researchers to apply

The company describes the pilot as resource-intensive and slower than its normal internal research process, which means broader access is not yet guaranteed at scale.

For now, Anthropic is collecting expressions of interest from researchers who want to conduct studies that would otherwise require access to data held inside an AI company.

The aggregated datasets from the three pilot projects are also being released publicly.

Anthropic says the next phase will focus on determining how many external projects the system can support while maintaining privacy, safety and research quality, with the Stanford study complete and further findings from Oxford and METR still to come.

Previous
Previous

Soroush Saghafian joins University of Maryland after 11 years at Harvard

Next
Next

OpenAI rolls out GPT-6 Astra for complex research, coding and computer-use tasks