EDBT 2026 Demo / reviewers in the wild / expert
Apratim Dey
dblp:374/6010
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 52% Trustworthy machine learning · 16% Generative modeling · 16% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity · ICLR 2025 |
Natural language and speech › Language models and text generation
knowledge editing |
0.9 | 1 | 2025 | Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity · ICLR 2025 |
Machine learning › Generative modeling
model collapse |
0.9 | 1 | 2025 | Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating World · ICML 2025 |
Machine learning › Representation and self-supervised learning › pre-training
pretraining data |
0.9 | 1 | 2025 | Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating World · ICML 2025 |
Natural language and speech › Language models and text generation
synthetic data |
0.9 | 1 | 2025 | Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating World · ICML 2025 |
Machine learning › Trustworthy machine learning
toxicity reduction |
0.9 | 1 | 2025 | Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity · ICLR 2025 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.3 | 1 | 2025 | Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
subspace identification · 0.9projection filter · 0.9multivariate gaussian estimation · 0.9language model fine-tuning · 0.9kernel density estimation · 0.9factor analysis · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Model Editing as a Robust and Denoised variant of DPO: A Case Study on ToxicityabstractRecent alignment algorithms such as direct preference optimization (DPO) have been developed to improve the safety of large language models (LLMs) by training these models to match human behaviors exemplified by preference data. However, these methods are
both computationally intensive and lacking in controllability and transparency, inhibiting their widespread use. Furthermore, these tuning-based methods require large-scale preference data for training and are susceptible to noisy preference data. In this paper, we introduce a tuning-free alignment alternative, ProFS (Projection Filter for Subspaces), and demonstrate its effectiveness under the use case of toxicity reduction. Grounded on theory from factor analysis, ProFS is a sample-efficient model editing approach that identifies a toxic subspace in the model parameter space and reduces model toxicity by projecting away the detected subspace. The toxic subspace is identified by extracting preference data embeddings from the language model, and removing non-toxic information from these embeddings. We show that ProFS is more sample-efficient than DPO, further showcasing greater robustness to noisy data. Finally, we attempt to connect tuning based alignment with editing, by establishing both theoretical and empirical connections between ProFS and DPO, showing that ProFS can be interpreted as a denoised version of a single DPO step. Rheeya Uppaal, Apratim Dey, Yiting He, Yiqiao Zhong, Junjie Hu 0001 |
ICLR | 2 |
| 2025 | Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating WorldabstractWhat happens when generative machine learning models are pretrained on web-scale datasets containing data generated by earlier models? Some prior work warns of “model collapse” as the web is overwhelmed by synthetic data; other work suggests the problem can be contained (i.e. collapse can be avoided) by managing how available data are used in pretraining. In this paper, we report experiments on three ways of using data (training-workflows), across three generative model task-settings (multivariate Gaussian estimation, kernel density estimation, and language-model fine-tuning) to further confirm the possibility of containment: (a) we confirm that the training-workflow of {\it replacing} all real data by successive generations of purely synthetic data indeed suffers model collapse in all task-settings studied; (b) we consider the training-workflow of {\it accumulating} synthetic data alongside real data and training on all data combined and confirming that, although the proportion of real data eventually becomes zero, models remain stable and their test losses do not diverge under this training-workflow; (c) we consider a training-workflow where real and synthetic data accumulate together but successive generations of pretraining are constrained to use fixed-size data subsets each generation. In this workflow, we observe slow and gradual rather than explosive degradation of test loss performance across generations. Our insights are particularly important when forecasting whether future frontier generative models will collapse or thrive, and our results open avenues for empirically and mathematically studying the context-dependent value of synthetic data. Joshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser, Rafael Rafailov, David L. Donoho, Oluwasanmi Koyejo |
ICML | 3 |