AI developers including OpenAI and Anthropic are purchasing internal corporate data such as employee chats, emails, video conference recordings, and code change histories. The shift comes as the supply of publicly available internet text approaches exhaustion, according to a report by The Information published on August 13 that cited roughly 20 sources, among them startup founders and data industry insiders.

Rising Demand for Enterprise Communication Records

Demand for this type of data has intensified over the past several months. The driver is the push to build AI agents that can handle real-world tasks, including customer service and invoice verification. Bobby Samuels, CEO of data brokerage startup Protege, said his company’s gross transaction volume climbed from $30 million last year to at least $100 million this year.

Buyers focus mainly on startups that face bankruptcy or acquisition. Warmly, an AI agent startup recently acquired by HubSpot, received four separate offers of up to $300,000 for its internal meeting minutes and emails after signing the acquisition agreement. The company declined every offer. Records of live human interactions—developers working through problems over email, CFOs reviewing financial performance, software demonstration videos—are viewed as essential training material for the next generation of AI systems.

The Approaching Limit on Public Text

The purchases reflect a structural constraint on public data. Epoch AI estimates that the effective stock of quality public human-written text totals roughly 300 trillion tokens. Frontier models could fully consume that stock between 2026 and 2032. Under more aggressive training assumptions, the timeline could move as early as 2027. Fresh human text still appears online, yet its growth rate lags far behind the expansion of AI training datasets.

Earlier Controversies Over Training Material

The move toward private data follows earlier disputes about how AI companies obtained material. Anthropic’s internal Project Panama initiative involved buying copyrighted books, removing their bindings for scanning, and destroying the physical copies afterward. A federal judge in San Francisco approved a $1.5 billion class-action settlement in July related to the program. An unsealed internal planning document described the effort as an attempt “to destructively scan all the books in the world.”

Anonymization Challenges and Rising Training Costs

De-identification of corporate datasets has become a major hurdle. Shub Sinha, CEO of data processing firm Integral, said his company runs a rigorous anonymization process designed to keep the data useful while meeting privacy regulations. The growth of the private-text market shows that training costs, once driven mainly by compute, are turning into a more significant expense line for AI developers.