Collapsing Distance: The Curse of Ground Truth in Computational Narrative Understanding
Abstract: As AI model size grows, neural scaling laws have become a crucial tool to predict the improvements of large models when increasing capacity and the size of original (human or natural) training data. Yet, the widespread use of popular models means that the ecosystem of online data and text will co-evolve to progressively contain increased amounts of synthesized data. In this paper we ask: How will the scaling laws change in the inevitable regime where synthetic data makes its way into the training corpus? Will future models, still improve, or be doomed to degenerate up to total (model) collapse? We develop a theoretical framework of model collapse through the lens of scaling laws. We discover a wide range of decay phenomena, analyzing loss of scaling, shifted scaling with number of generations, the ‘‘un-learning” of skills, and grokking when mixing human and synthesized data. Our theory is validated by large-scale experiments with a transformer on an arithmetic task and text generation using the large language model Llama2.
Show BibTeX
@inproceedings{DBLP:conf/text2story/HuangU26,
author = {Junbo Huang and
Ricardo Usbeck},
editor = {Ricardo Campos and
Al{\'{\i}}pio M{\'{a}}rio Jorge and
Adam Jatowt and
Sumit Bhatia and
Marina Litvak},
title = {Collapsing Distance: The Curse of Ground Truth in Computational Narrative
Understanding},
booktitle = {Proceedings of Text2Story - Ninth Workshop on Narrative Extraction
From Texts held in conjunction with the 48th European Conference on
Information Retrieval {(ECIR} 2026), Delft, The Netherlands, March
29, 2026},
series = {{CEUR} Workshop Proceedings},
volume = {4202},
publisher = {CEUR-WS.org},
year = {2026},
url = {https://ceur-ws.org/Vol-4202/paper13.pdf},
timestamp = {Thu, 07 May 2026 15:41:11 +0200},
biburl = {https://dblp.org/rec/conf/text2story/HuangU26.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}