
For years, AI models have gotten better by training on more data. GPT-3 used 45 terabytes of text. GPT-4 used even more. But there’s a problem: we’re running out of new, high-quality text to feed these models.
Researchers estimate that publicly available text data—books, websites, articles, forums—will be largely exhausted by 2027 or 2028. The internet isn’t infinite, and AI companies have already scraped most of it. So what happens next?
Why Running Out of Data Matters
Large language models learn patterns from massive datasets. More diverse, high-quality data generally means better performance. But the well is running dry.
Most accessible text has already been used. Books, Wikipedia, Reddit, news archives, academic papers—all scraped and tokenized. What’s left is either lower quality, duplicative, or locked behind paywalls and copyright restrictions.
This creates a bottleneck. If models can’t train on fresh data, improvement slows. The easy gains from “bigger datasets” are over.
Synthetic Data: Training AI on AI-Generated Text
One solution is synthetic data—text generated by AI models themselves. OpenAI, Google, and Anthropic are all experimenting with this approach.
The idea: use a strong model to generate new training examples. For instance, Claude could write thousands of practice coding problems, which are then used to train the next version of Claude.
But synthetic data has risks. Models trained heavily on AI-generated text can suffer from “model collapse”—they start repeating the same patterns, losing diversity and originality. It’s like photocopying a photocopy: each generation degrades.
Researchers are working on hybrid approaches: combining synthetic data with carefully curated human-written text, and using AI to filter out low-quality synthetic examples before training.
Multimodal Data and Private Datasets
Another path: go beyond text. Video, audio, images, and sensor data offer massive new training opportunities. Models like GPT-4 and Gemini already handle images. Future models may learn from video, simulations, or even robotics data.
Companies are also turning to private datasets—proprietary archives, licensed content, and partnerships with publishers. OpenAI has deals with news organizations. Google has YouTube. These datasets aren’t available to smaller competitors, which could concentrate power among a few large players.
What This Means for You
In practice, you may notice AI progress shifting. Instead of dramatic leaps every six months, improvements might come from better reasoning algorithms, fine-tuning, and efficiency—not just larger datasets.
Tools you use daily (ChatGPT, Claude, Gemini) will keep getting better, but the pace and nature of progress may change. Expect more focus on specialized models, domain-specific fine-tuning, and smarter use of existing data.
For businesses and developers, this shift underscores the value of proprietary data. If you have unique datasets—customer interactions, domain expertise, internal documents—they become more valuable for training custom models.
The data crunch is real, but it’s not the end of AI progress. It’s a forcing function for smarter, more efficient approaches. And that might lead to better AI in the long run.
Want one useful AI insight like this in your inbox every day? Subscribe to the One Two Three AI newsletter and stay ahead of what’s changing.
