VLDB 2026 Research / reviewers in the wild / expert
Rui Wang 0119
dblp:06/2293-119
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0008-5461-0353ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 50% Language models and text generation · 25% Trustworthy machine learning · 25% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient pretraining |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Machine learning › Trustworthy machine learning › Data-centric AI
data influence |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Machine learning › Efficient and distributed learning
data selection |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
kronecker product · 0.9influence functions · 0.9clustering · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language ModelsabstractData selection is of great significance in pretraining large language models, given the variation in quality within the large-scale available training corpora.
To achieve this, researchers are currently investigating the use of data influence to measure the importance of data instances, $i.e.,$ a high influence score indicates that incorporating this instance to the training set is likely to enhance the model performance. Consequently, they select the top-$k$ instances with the highest scores. However, this approach has several limitations.
(1) Calculating the accurate influence of all available data is time-consuming.
(2) The selected data instances are not diverse enough, which may hinder the pretrained model's ability to generalize effectively to various downstream tasks.
In this paper, we introduce $\texttt{Quad}$, a data selection approach that considers both quality and diversity by using data influence to achieve state-of-the-art pretraining results.
To compute the influence ($i.e.,$ the quality) more accurately and efficiently, we incorporate the attention layers to capture more semantic details, which can be accelerated through the Kronecker product.
For the diversity, $\texttt{Quad}$ clusters the dataset into similar data instances within each cluster and diverse instances across different clusters. For each cluster, if we opt to select data from it, we take some samples to evaluate the influence to prevent processing all instances. Overall, we favor clusters with highly influential instances (ensuring high quality) or clusters that have been selected less frequently (ensuring diversity), thereby well balancing between quality and diversity. Experiments on Slimpajama and FineWeb over 7B large language models demonstrate that $\texttt{Quad}$ significantly outperforms other data selection methods with a low FLOPs consumption. Further analysis also validates the effectiveness of our influence calculation. Chi Zhang 0102, Huaping Zhong, Chengliang Chai, Rui Wang 0119, Xinlin Zhuang, Tianyi Bai, Jiantao Qiu, Lei Cao 0004, Ju Fan, Ye Yuan 0001, Guoren Wang, Conghui He |
ICLR | 5 |