Since conversational generative AI chatbots became available a few years ago, the internet has rapidly become populated with AI-generated texts. Many worry that, as these artificial texts are reused as training data for AI chatbots, they could impoverish our language and amplify existing biases.
Over the past few years, this question has become a major research topic, with hundreds of articles published so far. “There have been many studies lately, with different results. Most of them show that the diversity of the data obtained with models retrained on data generated by the same model decreases until it collapses. It can get to creating what they call “garbage,” texts that are just a random collection of words,” explains Matteo Marsili, a Senior Research Scientist in ICTP’s Quantitative Life Sciences section. Marsili recently collaborated with former ICTP Diploma student Fariba Jangjoo from the Kavli Institute for Systems Neuroscience in Norway, to understand the problem from a deeper and more general perspective. Their paper was published in Physical Review Letters.
More news
ICTP: A Safe Harbour
Befriending Earthquakes
Inside Matter's Secrets
India Lecture Tour
The Cost of Surviving
Launching Salam's Centennial Year
The ICTP SciFabLab Meets Kuwait
Science is Our Common Language
Bright Imperfections?
Eccellence in Medical Physics
Ramanujan Prize 2025 Announced
Public film screening at ICTP
How Bacteria Predict the Future
Fluorescence in Amino Acid Crystals
ICTP Welcomes Delegation from India
Assessing Climate Change
ICTP Success Story