VLDB 2026 Research / reviewers in the wild / expert
Jérémy Perez
dblp:372/2869
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 62% Generative modeling · 19% Trustworthy machine learning · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification |
0.9 | 1 | 2025 | Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? · EMNLP 2025 |
Natural language and speech › Language models and text generation
evaluation of language models |
0.9 | 1 | 2025 | When LLMs Play the Telephone Game: Cultural Attractors as Conceptual Tools to Evaluate LLMs in Multi-turn Settings · ICLR 2025 |
Machine learning › Generative modeling
model collapse |
0.9 | 1 | 2025 | Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? · EMNLP 2025 |
Natural language and speech › Language models and text generation
synthetic data |
0.9 | 1 | 2025 | Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
transmission chain design · 1.7telephone game experiments · 1.7regression analysis · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?abstractLarge language models (LLMs) are increasingly used in the creation of online content, creating feedback loops as subsequent generations of models will be trained on this synthetic data.Such loops were shown to lead to distribution shifts -models misrepresenting the true underlying distributions of human data (also called model collapse).However, how human data properties affect such shifts remains poorly understood.In this paper, we provide the first empirical examination of the effect of such properties on the outcome of recursive training.We first confirm that using different human datasets leads to distribution shifts of different magnitudes.Through exhaustive manipulation of dataset properties combined with regression analyses, we then identify a set of properties associated with distribution shift magnitudes.Lexical diversity is found to amplify these shifts, while semantic diversity and data quality mitigate them.Furthermore, we find that these influences are highly modular: data scrapped from a given internet domain has little influence on the content generated for another domain.Finally, experiments on political bias reveal that human data properties affect whether the initial bias will be amplified or reduced.Overall, our results portray a novel view, where different parts of internet may undergo different types of distribution shift. Grgur Kovac, Jérémy Perez, Rémy Portelas, Peter Ford Dominey, Pierre-Yves Oudeyer |
EMNLP | 2 |
| 2025 | When LLMs Play the Telephone Game: Cultural Attractors as Conceptual Tools to Evaluate LLMs in Multi-turn SettingsabstractAs large language models (LLMs) start interacting with each other and generating an increasing amount of text online, it becomes crucial to better understand how information is transformed as it passes from one LLM to the next. While significant research has examined individual LLM behaviors, existing studies have largely overlooked the collective behaviors and information distortions arising from iterated LLM interactions. Small biases, negligible at the single output level, risk being amplified in iterated interactions, potentially leading the content to evolve towards attractor states. In a series of _telephone game experiments_, we apply a transmission chain design borrowed from the human cultural evolution literature: LLM agents iteratively receive, produce, and transmit texts from the previous to the next agent in the chain. By tracking the evolution of text _toxicity_, _positivity_, _difficulty_, and _length_ across transmission chains, we uncover the existence of biases and attractors, and study their dependence on the initial text, the instructions, language model, and model size. For instance, we find that more open-ended instructions lead to stronger attraction effects compared to more constrained tasks. We also find that different text properties display different sensitivity to attraction effects, with _toxicity_ leading to stronger attractors than _length_. These findings highlight the importance of accounting for multi-step transmission dynamics and represent a first step towards a more comprehensive understanding of LLM cultural dynamics. Jérémy Perez, Grgur Kovac, Corentin Léger, Cédric Colas, Gaia Molinaro, Maxime Derex, Pierre-Yves Oudeyer, Clément Moulin-Frier |
ICLR | 1 |