VLDB 2026 Research / reviewers in the wild / expert
Xinlin Zhuang
dblp:366/3295
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0006-1822-8224ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Efficient and distributed learning · 54% Language models and text generation · 29% Trustworthy machine learning · 10% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
data-efficient learning |
1.7 | 2 | 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025 Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
data selection |
1.7 | 2 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
1.7 | 2 | 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025 Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
federated learning |
1.0 | 1 | 2026 | FedCD: Towards Consolidated Distillation for Heterogeneous Federated Learning · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.0 | 1 | 2026 | FedCD: Towards Consolidated Distillation for Heterogeneous Federated Learning · AAAI 2026 |
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient pretraining |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Machine learning › Trustworthy machine learning › Data-centric AI
data influence |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Machine learning › Representation and self-supervised learning › pre-training › data-centric pre-training
data selection for pre-training |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.9 | 1 | 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model training
pretraining data selection |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
logit-based distillation · 1.0gaussian mixture model · 1.0feature-based distillation · 1.0cross-layer attention · 1.0multi-dimensional data selection · 0.9multi-actor collaboration · 0.9kronecker product · 0.9influence functions · 0.9curriculum learning · 0.9clustering · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedCD: Towards Consolidated Distillation for Heterogeneous Federated LearningabstractKnowledge Distillation (KD) serves as an effective approach to addressing heterogeneity issues in Federated Learning (FL), leveraging additional datasets to align local and global models better. There are two primary distillation paradigms: feature-based distillation, which utilizes intermediate-layer features of the network, and logit-based distillation, which employs the final layer's logit outputs. However, existing studies often select distillation methods based on intuitive and empirical evidence when facing different heterogeneous settings, neglecting the intrinsic relationship between distillation paradigms and heterogeneity. This oversight may result in suboptimal federated knowledge distillation performance under heterogeneous conditions. In this paper, we propose the Consolidated Distillation for Heterogeneous Federated Learning - FedCD that balances knowledge representations from both feature-based and logit-based distillation to enhance performance. Specifically, to address the misalignment between knowledge conveyed by features and logits, we aggregate features from different layers via cross-layer attention to preserve semantic knowledge, followed by distribution modeling using Gaussian Mixture Models. This process strengthens knowledge distillation by constraining the transformation of different network layers' features under a consolidated distribution, thereby mitigating impacts from both data and model heterogeneity. Extensive experiments demonstrate that FedCD outperforms state-of-the-art methods by over 10.72% and validate the effectiveness of our approach. Yichen Li 0006, Huifa Li, Xinlin Zhuang, Haochen Xue, Haozhao Wang, Muhammad Imran Razzak |
AAAI | 5 |
| 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor CollaborationabstractEfficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to 10.5% across multiple language model benchmarks compared to the state-of-the-art methods. Code and checkpoints are publicly released at https://github.com/Relaxed-System-Lab/multi-actor-data-selection. Tianyi Bai, Ling Yang 0006, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Lijun Wu 0003, Jiantao Qiu, Wentao Zhang 0001, Binhang Yuan, Conghui He |
ACL (1) | 5 |
| 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language ModelsabstractXinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang, Tianyi Bai, Xingjian Wei, Qiu Jiantao, Chi Zhang, Ying Qian, Conghui He. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang, Tianyi Bai, Xingjian Wei, Jiantao Qiu, Conghui He |
ACL (1) | 1 |
| 2025 | FinDABench: Benchmarking Financial Data Analysis Ability of Large Language ModelsabstractLarge Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks. However, their proficiency and reliability in the specialized domain of financial data analysis, particularly focusing on data-driven thinking, remain uncertain. To bridge this gap, we introduce FinDABench, a comprehensive benchmark designed to evaluate the financial data analysis capabilities of LLMs within this context. The benchmark comprises 15,200 training instances and 8,900 test instances, all meticulously crafted by human experts. FinDABench assesses LLMs across three dimensions: 1) Core Ability, evaluating the models’ ability to perform financial indicator calculation and corporate sentiment risk assessment; 2) Analytical Ability, determining the models’ ability to quickly comprehend textual information and analyze abnormal financial reports; and 3) Technical Ability, examining the models’ use of technical knowledge to address real-world data analysis challenges involving analysis generation and charts visualization from multiple perspectives. We will release FinDABench, and the evaluation scripts at https://github.com/xxx. FinDABench aims to provide a measure for in-depth analysis of LLM abilities and foster the advancement of LLMs in the field of financial data analysis. Shangqing Zhao, Chenghao Jia, Xinlin Zhuang, Zhaoguang Long, Aimin Zhou, Man Lan, Yang Chong |
COLING | 4 |
| 2025 | Harnessing Diversity for Important Data Selection in Pretraining Large Language ModelsabstractData selection is of great significance in pretraining large language models, given the variation in quality within the large-scale available training corpora.
To achieve this, researchers are currently investigating the use of data influence to measure the importance of data instances, $i.e.,$ a high influence score indicates that incorporating this instance to the training set is likely to enhance the model performance. Consequently, they select the top-$k$ instances with the highest scores. However, this approach has several limitations.
(1) Calculating the accurate influence of all available data is time-consuming.
(2) The selected data instances are not diverse enough, which may hinder the pretrained model's ability to generalize effectively to various downstream tasks.
In this paper, we introduce $\texttt{Quad}$, a data selection approach that considers both quality and diversity by using data influence to achieve state-of-the-art pretraining results.
To compute the influence ($i.e.,$ the quality) more accurately and efficiently, we incorporate the attention layers to capture more semantic details, which can be accelerated through the Kronecker product.
For the diversity, $\texttt{Quad}$ clusters the dataset into similar data instances within each cluster and diverse instances across different clusters. For each cluster, if we opt to select data from it, we take some samples to evaluate the influence to prevent processing all instances. Overall, we favor clusters with highly influential instances (ensuring high quality) or clusters that have been selected less frequently (ensuring diversity), thereby well balancing between quality and diversity. Experiments on Slimpajama and FineWeb over 7B large language models demonstrate that $\texttt{Quad}$ significantly outperforms other data selection methods with a low FLOPs consumption. Further analysis also validates the effectiveness of our influence calculation. Chi Zhang 0102, Huaping Zhong, Chengliang Chai, Rui Wang 0119, Xinlin Zhuang, Tianyi Bai, Jiantao Qiu, Lei Cao 0004, Ju Fan, Ye Yuan 0001, Guoren Wang, Conghui He |
ICLR | 6 |
| 2024 | A Lightweight and Effective Multi-View Knowledge Distillation Framework for Text-Image RetrievalabstractLarge-scale dual-stream Vision-Language Pre-training (VLP) models provide an efficient solution for text-image retrieval tasks. Despite this, their performance often falls short of the most current single-stream models, primarily due to limited fine-grained text-image interactions. Recent trends indicate a union of these two types of networks. Some methods adopt a retrieve and rerank strategy, their performance improvements largely hinge on the single-stream encoder during inference. Other approaches utilize knowledge distillation to strengthen either the single-stream encoder or the dual-stream encoder, surpassing their previous capabilities. However, existing distillation techniques typically focus on a single knowledge type, neglecting the richer insights available in the teacher model. To bridge this gap, we introduce a Lightweight and Effective Multi-View Knowledge Distillation approach, named LEMKD, for text-image retrieval. This method effectively utilizes response-based, feature-based and relation-based knowledge, transferring the knowledge from the single-stream encoder to the dual-stream encoder. Our approach is executed on the widely used MS-COCO and Flickr30K datasets. Results demonstrate that LEMKD not only matches the exceptional performance of the most advanced single-stream models but also excels in dual-stream encoder performance amidst the recent integration of single-stream and dual-stream models. Yuxiang Song, Yuxuan Zheng, Shangqing Zhao, Xinlin Zhuang, Zhaoguang Long, Changzhi Sun, Aimin Zhou, Man Lan |
IJCNN | 5 |
| 2024 | Bread: A Hybrid Approach for Instruction Data Mining Through Balanced Retrieval and Dynamic Data Sampling
Xinlin Zhuang, Xin Mao 0002, Hongyi Wu, Shangqing Zhao, Yuxiang Song, Chenghao Jia, Man Lan |
NLPCC (2) | 1 |