AI is taking over the internet. About 10% of all web pages in 2026 were created, at least in part, with AI, according to a study by Pew Research.
That study includes webpages created before ChatGPT was created in 2022. Removing the older web pages from the study shows that a full 35%, or more than one in three pages, were created by an AI model.
The Large Language Models (LLMs) used to train generative AI assistants "make vast bodies of text more accessible through summaries, explanations, and translations," writes attorney Joanna Olin in The Chronicle of Higher Education. However, the biggest risk comes when the LLMs are used to train the later chatbots. Olin writes, "The concern, then, is not only whether a particular answer is accurate, but whether the human-authored record beneath it remains identifiable and accessible."
Olin uses French jurisprudence as an example. Based on Roman law, French scholars in the 16th century pored over Roman records, finding much that was obscure, contradicted by later scholarship, and corrupted and corrected. Using the principle of ad fontes, or “back to the sources,” the scholars were able to construct a legal framework based on the past and present law.
With so much AI slop being put out and the LLMs gobbling it up to create new models, at what point does the human disappear and research devolve into an AI smorgasbord?
The Chronicle of Higher Education:
The technical debate over “model collapse” illustrates one part of the problem. The term refers to the degradation that can occur when large language models are trained recursively on data generated by earlier models, rather than primarily on human-authored material. Repeated cycles can reinforce errors and simplifications while diminishing rare or complex features of the original data. AI researchers argue that preserving access to the original human-generated distribution is crucial if models are to avoid that drift. Others warn that recursive synthetic training can weaken a system’s ability to generalize to real-world data.
Not everyone accepts the alarm around the “model collapse” forecast. But the debate still exposes a deeper question, one that would have been familiar to the French legal humanists. What happens when a system of knowledge becomes increasingly detached from the human-authored sources on which it ultimately depends?
Higher education holds a vital responsibility to preserve authentic, human-authored data. As generative AI models synthesize vast quantities of web text, risking "model collapse" and blurring the lines between original sources and machine interpretations, colleges serve as essential guardians of the original source information.
"Colleges must sustain their libraries, archives, special collections, digital repositories, and systems of provenance, as well as the librarians, archivists, and scholars of library and information science whose expertise makes those resources trustworthy and usable" Olin writes, "AI-generated summaries may help researchers orient themselves with unfamiliar material, but they cannot replace the scholarship that interprets, the evidence that substantiates, or the human judgment that weighs competing claims."
As AI models become faster, more sophisticated, and capable of performing more complex tasks, universities must foster a "return to sources" (ad fontes). Beyond teaching students digital literacy, colleges must preserve the capacity to look beneath AI-generated summaries and test AI claims against original, human-created source material.
AI is the most powerful system of textual synthesis ever created, which has made source-consciousness more urgent than ever. The response should not be to reject the valuable synthesis that AI can provide, any more than the legal humanists rejected the gloss. Rather, it should be to preserve the discipline and the institutional capacity to look beneath it.
For ad fontes to take institutional form, colleges must protect the collections, professional expertise, and systems of provenance that allow synthetic claims to be tested against original sources and historical context. The call to return to the sources is a condition for preserving the integrity of knowledge and our ability to distinguish well-grounded truth from unverified synthesis.
I'm skeptical. We are lazy creatures, and AI is just too tempting to use in lieu of actual research and writing. With 35% of post 2022 webpages generated in part or wholly by AI, we can expect that percentage to rise.
As for scholars maintaining contact with human-generated sources, I wish them luck. Until AI can come up with a sure-fire way to determine if writing is manufactured by a machine or in the human brain, scholars will have to be extra careful in accepting a product as human-generated.
Recommended: WaPo Writer, Fired for Anti-Charlie Kirk Comments, Found to Have the Right to Be Awful






