
Throughout 2026, there has been a steady stream of news stories about the lengths some AI companies are going to in order to harvest information. In June of this year, court filings revealed that Anthropic, owner of the Claude chatbot, had purchased, scanned and logged millions of physical books to better train its model. More recently, second hand booksellers have noted scores of their novels being purchased and shipped to far-flung warehouses, with the suspicion being that AI trainers are behind it. This is where specialisation can become an advantage.
There has long been an assumption that the amount of data and information generated and available in the internet age is almost limitless. Add in the quantity being newly generated each day and the sum total of available data would seem to be on a scale beyond imagining. Yet the emergence of LLMs and their trainers insatiable need for information has exposed a stark limit that was not previously obvious.
When people think of the rise of AI and the central role LLMs are already playing in our lives, they tend to picture technology at its most cutting edge, with algorithms and code whizzing around behind the scenes. What they do not envisage is teams of people and shell companies buying up some of the world’s oldest, dustiest books. Yet having seemingly ‘tapped out’ the internet, that is exactly what they are doing.
What does this mean, at least in the short-to-medium term, for general-use LLMs? It points towards stagnation, with an eventual drying up of the information wells on which they draw upon. No doubt their development, in time, will take them down different paths that allow them to improve, but that looks likely to have a longer time horizon. For now, it represents a potentially significant slowdown.
That creates an opportunity for highly specialised AI tools designed with a specific purpose in mind. Instead of fixating on LLMs that try and capture, evaluate, summarise and deliver the total of human knowledge, those company’s which produce tools with a specific, highly impactful purpose in mind stand to arguably gain the most.
Take OpenEvidence. The company was created by former Kensho Technologies founder, Daniel Nadler, and has been billed as ‘ChatGPT for doctors’. However, this description is misleading and disguises the true nature of how it operates. OpenEvidence does not seek to harvest as much information as possible. Indeed, it does not even seek to harvest all known medical information, recognising that much of this information is poor quality and potentially harmful.
Instead, Nadler has built a tool that depends on a targeted corpus of high quality, accurate and up-to-date medical information. OpenEvidence does not voraciously seek out new material for new material’s sake. It seeks only to update the bank of data it depends on with information that will materially add to the tools ability to deliver high quality, medically accurate responses.
In the world of mass medicine, particularly at the initial consultation stage, this is crucial. Quality vastly outweighs quantity in terms of importance. The doctors using this tool, 45% of all physicians in America according to the company’s latest data, must be certain that the information it provides is accurate. That is why OpenEvidence is trained on medical libraries, peer reviewed journals and verified research studies.
That isn’t to say quantity doesn’t matter at all. The purpose behind the tool is to provide doctors with rapid, distilled access to an amount of medical knowledge that they could not possibly hold inside their head. The amount of good quality research available that can inform medical decisions could not possibly be delivered to frontline medical professionals in any other way.
However, what the example of OpenEvidence shows is that those AI tools designed with a specific, clear focus in mind are well placed to thrive in the competitive marketplace seeking to harness the broader technology. Whilst the generalised tools are scrambling around for any data they can get their hands on, specialised tools which methodically, purposefully gather higher quality information for a dedicated cause look set to thrive. In the race for AI information, less can very much be more, as Daniel Nadler and his team are fast learning.