VLDB 2026 Research / reviewers in the wild / expert
Xinyi Shang
dblp:300/0145
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 63% Deep learning architectures and training · 16% Language models and text generation · 11% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
2.8 | 4 | 2025 | Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatch · CVPR 2025 Revisiting Weighted Aggregation in Federated Learning with Neural Networks · ICML 2023 No Fear of Classifier Biases: Neural Collapse Inspired Federated Learning with Synthetic and Fixed Classifier · ICCV 2023 |
Natural language and speech › Language models and text generation
large language model training |
1.0 | 1 | 2026 | LLMSurgeon: Diagnosing Data Mixture of Large Language Models · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
dataset distillation |
0.9 | 1 | 2025 | GIFT: Unlocking Full Potential of Labels in Distilled Dataset at Near-zero Cost · ICLR 2025 |
Machine learning › Efficient and distributed learning › federated learning › label-efficient federated learning
federated semi-supervised learning |
0.9 | 1 | 2025 | Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatch · CVPR 2025 |
Computer vision › Segmentation and scene understanding › pseudo-label learning
pseudo-label refinement |
0.9 | 1 | 2025 | Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatch · CVPR 2025 |
Machine learning › Efficient and distributed learning › federated learning
heterogeneous federated learning |
0.7 | 1 | 2023 | No Fear of Classifier Biases: Neural Collapse Inspired Federated Learning with Synthetic and Fixed Classifier · ICCV 2023 |
Machine learning › Deep learning architectures and training
weight averaging |
0.7 | 1 | 2023 | Revisiting Weighted Aggregation in Federated Learning with Neural Networks · ICML 2023 |
Machine learning › Efficient and distributed learning › data-centric learning
data-centric training |
0.3 | 1 | 2026 | LLMSurgeon: Diagnosing Data Mixture of Large Language Models · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › federated learning
data heterogeneity |
0.3 | 1 | 2025 | Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatch · CVPR 2025 |
Machine learning › Deep learning architectures and training
neural collapse |
0.2 | 1 | 2023 | No Fear of Classifier Biases: Neural Collapse Inspired Federated Learning with Synthetic and Fixed Classifier · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
mixture analysis · 1.0data attribution · 1.0soft label refinement · 0.9ensemble aggregation · 0.9cosine similarity loss · 0.9consistency regularization · 0.9confidence discrepancy · 0.9synthetic fixed classifier · 0.7simplex equiangular tight frame · 0.7neural collapse · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMSurgeon: Diagnosing Data Mixture of Large Language ModelsabstractYaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang, Jiacheng Liu, Xinyue Bi, Zhaoyi Li, Zhiqiang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang, Xinyue Bi |
ACL (1) | 4 |
| 2026 | Dawid-Skene-model-based label-noise mitigation for federated learningabstractFederated learning (FL) enables collaborative model training without centralising raw data, but its performance is susceptible to label noise from clients. A common mitigation strategy involves using a clean, labelled public dataset at the server to assess client reliability. However, this approach is impractical due to the unrealistic assumption of availability of a clean, labelled public dataset. To address this issue, we propose FedDS, a novel approach that brings the Dawid-Skene model from statistical analysis to FL, which enables the estimation of the reliability of each client in FL without requiring any labelled data at the server. This approach effectively mitigates the adverse impact of heterogeneous label noise under a weaker and more practical assumption, offering a robust aggregation strategy for real-world FL scenarios with label noise. The code is available at https://github.com/Gia99999/FedDS . Jia Dong, Rui Zhu 0006, Xinyi Shang, Jing-Hao Xue |
Inf. Sci. | 3 |
| 2026 | Federated learning with noisy labels: A comprehensive and concise review of current methodologies and future directionsabstractFederated learning, a vital paradigm in modern machine learning, enables private and decentralised training of models that is crucial for learning from sensitive data. Noisy label learning, another vital paradigm in modern machine learning, addresses the training of models from the data with potentially incorrect labels. Their integration, namely federated learning with noisy labels (FLNL), is an emerging but challenging topic arising from the practice of machine learning, which, however, still lacks a review of its research progress. The aim of this paper is to fill in this gap. We first summarise four core challenges to FLNL: localised label noise, across-client heterogeneity of label noise, localised overfitting to label noise, and inadequate benchmarking. We then propose a taxonomy to categorise current FLNL studies into four types that address the four challenges correspondingly: sample-wise methods, client-wise methods, model-wise methods, and benchmark-wise studies. This work offers the first comprehensive and concise review dedicated to FLNL; moreover, we also provide future research directions for this rapidly evolving and practically significant field. Jia Dong, Rui Zhu 0006, Xinyi Shang, Jing-Hao Xue |
Neural Networks | 3 |
| 2025 | Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-MismatchabstractFederated Semi-Supervised Learning (FSSL) aims to leverage unlabeled data across clients with limited labeled data to train a global model with strong generalization ability. Most FSSL methods rely on consistency regularization with pseudo-labels, converting predictions from local or global models into hard pseudo-labels as supervisory signals. However, we discover that the quality of pseudo-label is largely deteriorated by data heterogeneity, an intrinsic facet of federated learning. In this paper, we study the problem of FSSL in-depth and show that (1) heterogeneity exacerbates pseudo-label mismatches, further degrading model performance and convergence, and (2) local and global models’ predictive tendencies diverge as heterogeneity increases. Motivated by these findings, we propose a simple and effective method called Semi-supervised Aggregation for Globally-Enhanced Ensemble (SAGE), that can flexibly correct pseudo-labels based on confidence discrepancies. This strategy effectively mitigates performance degradation caused by incorrect pseudo-labels and enhances consensus between local and global models. Experimental results demonstrate that SAGE outperforms existing FSSL methods in both performance and convergence. Our code is available at https://github.com/Jay-Codeman/SAGE. Xinyi Shang, Yiqun Zhang 0006, Yang Lu 0009, Chen Gong 0002, Jing-Hao Xue, Hanzi Wang |
CVPR | 2 |
| 2025 | GIFT: Unlocking Full Potential of Labels in Distilled Dataset at Near-zero CostabstractRecent advancements in dataset distillation have demonstrated the significant benefits of employing soft labels generated by pre-trained teacher models.
In this paper, we introduce a novel perspective by emphasizing the full utilization of labels.
We first conduct a comprehensive comparison of various loss functions for soft label utilization in dataset distillation, revealing that the model trained on the synthetic dataset exhibits high sensitivity to the choice of loss function for soft label utilization.
This finding highlights the necessity of a universal loss function for training models on synthetic datasets.
Building on these insights, we introduce an extremely simple yet surprisingly effective plug-and-play approach, GIFT, which encompasses soft label refinement and a cosine similarity-based loss function to efficiently leverage full label information.
Extensive experiments indicate that GIFT consistently enhances state-of-the-art dataset distillation methods across various dataset scales without incurring additional computational costs.
Importantly, GIFT significantly enhances cross-optimizer generalization, an area previously overlooked.
For instance, on ImageNet-1K with IPC = 10, GIFT enhances the state-of-the-art method RDED by 30.8% in cross-optimizer generalization. Our code is available at https://github.com/LINs-lab/GIFT. Xinyi Shang |
ICLR | 1 |
| 2025 | Benchmarking large language models for genomic knowledge with GeneTuringabstractLarge language models (LLMs) show promise in biomedical research, but their effectiveness for genomic inquiry remains unclear. We developed GeneTuring, a benchmark consisting of 16 genomics tasks with 1600 curated questions, and manually evaluated 48 000 answers from 10 LLM configurations, including GPT-4o (via API, ChatGPT with web access, and a custom Generative Pretrained Transformer (GPT) setup), GPT-3.5, Claude 3.5, Gemini Advanced, GeneGPT (both slim and full), BioGPT, and BioMedLM. A custom GPT-4o configuration integrated with National Center for Biotechnology Information (NCBI) Application Programming Interfaces (APIs), developed in this study as SeqSnap, achieved the best overall performance. GPT-4o with web access and GeneGPT demonstrated complementary strengths. Our findings highlight both the promise and current limitations of LLMs in genomics, and emphasize the value of combining LLMs with domain-specific tools for robust genomic intelligence. GeneTuring offers a key resource for benchmarking and improving LLMs in biomedical research. Xinyi Shang, Wenpin Hou |
Briefings Bioinform. | 1 |
| 2025 | Secure color image encryption algorithm for face recognition using Zaslavsky and Arnold cat maps with binary bit-plane decomposition
Lei Ding 0010, Xinyi Shang, Lianhai Wang |
Inf. Sci. | 5 |
| 2025 | FediOS: decoupling orthogonal subspaces for personalization in feature-skew federated learning
Lingzhi Gao, Zexi Li 0001, Xinyi Shang, Yang Lu 0009, Chao Wu 0001 |
Mach. Learn. | 3 |
| 2023 | No Fear of Classifier Biases: Neural Collapse Inspired Federated Learning with Synthetic and Fixed ClassifierabstractData heterogeneity is an inherent challenge that hinders the performance of federated learning (FL). Recent studies have identified the biased classifiers of local models as the key bottleneck. Previous attempts have used classifier calibration after FL training, but this approach falls short in improving the poor feature representations caused by training-time classifier biases. Resolving the classifier bias dilemma in FL requires a full understanding of the mechanisms behind the classifier. Recent advances in neural collapse have shown that the classifiers and feature prototypes under perfect training scenarios collapse into an optimal structure called simplex equiangular tight frame (ETF). Building on this neural collapse insight, we propose a solution to the FL's classifier bias problem by utilizing a synthetic and fixed ETF classifier during training. The optimal classifier structure enables all clients to learn unified and optimal feature representations even under extremely heterogeneous data. We devise several effective modules to better adapt the ETF structure in FL, achieving both high generalization and personalization. Extensive experiments demonstrate that our method achieves state-of-the-art performances on CIFAR-10, CIFAR-100, and Tiny-ImageNet. The code is available at https://github.com/ZexiLee/ICCV-2023-FedETF. Zexi Li 0001, Xinyi Shang, Tao Lin 0004, Chao Wu 0001 |
ICCV | 2 |
| 2023 | Revisiting Weighted Aggregation in Federated Learning with Neural NetworksabstractIn federated learning (FL), weighted aggregation of local models is conducted to generate a global model, and the aggregation weights are normalized (the sum of weights is 1) and proportional to the local data sizes. In this paper, we revisit the weighted aggregation process and gain new insights into the training dynamics of FL. First, we find that the sum of weights can be smaller than 1, causing global weight shrinking effect (analogous to weight decay) and improving generalization. We explore how the optimal shrinking factor is affected by clients' data heterogeneity and local epochs. Second, we dive into the relative aggregation weights among clients to depict the clients' importance. We develop client coherence to study the learning dynamics and find a critical point that exists. Before entering the critical point, more coherent clients play more essential roles in generalization. Based on the above insights, we propose an effective method for Federated Learning with Learnable Aggregation Weights, named as FedLAW. Extensive experiments verify that our method can improve the generalization of the global model by a large margin on different datasets and models. Zexi Li 0001, Tao Lin 0004, Xinyi Shang, Chao Wu 0001 |
ICML | 3 |
| 2022 | FEDIC: Federated Learning on Non-IID and Long-Tailed Data via Calibrated DistillationabstractFederated learning provides a privacy guarantee for generating good deep learning models on distributed clients with different kinds of data. Nevertheless, dealing with non-IID data is one of the most challenging problems for federated learning. Researchers have proposed a variety of methods to eliminate the negative influence of non-IIDness. However, they only focus on the non-IID data provided that the universal class distribution is balanced. In many real-world applications, the universal class distribution is long-tailed, which causes the model seriously biased. Therefore, this paper studies the joint problem of non-IID and long-tailed data in federated learning and proposes a corresponding solution called Federated Ensemble Distillation with Imbalance Calibration (FEDIC). To deal with non-IID data, FEDIC uses model ensemble to take advantage of the diversity of models trained on non-IID data. Then, a new distillation method with logit adjustment and calibration gating network is proposed to solve the long-tail problem effectively. We evaluate FEDIC on CIFAR-10-LT, CIFAR-100-LT, and ImageNet-LT with a highly non-IID experimental setting, in comparison with the state-of-the-art methods of federated learning and long-tail learning. Our code is available at https://github.com/shangxinyi/FEDIC. Xinyi Shang, Yang Lu 0009, Yiu-Ming Cheung, Hanzi Wang |
ICME | 1 |
| 2022 | Federated Learning on Heterogeneous and Long-Tailed Data via Classifier Re-Training with Federated FeaturesabstractFederated learning (FL) provides a privacy-preserving solution for distributed machine learning tasks. One challenging problem that severely damages the performance of FL models is the co-occurrence of data heterogeneity and long-tail distribution, which frequently appears in real FL applications. In this paper, we reveal an intriguing fact that the biased classifier is the primary factor leading to the poor performance of the global model. Motivated by the above finding, we propose a novel and privacy-preserving FL method for heterogeneous and long-tailed data via Classifier Re-training with Federated Features (CReFF). The classifier re-trained on federated features can produce comparable performance as the one re-trained on real data in a privacy-preserving manner without information leakage of local data or class distribution. Experiments on several benchmark datasets show that the proposed CReFF is an effective solution to obtain a promising FL model under heterogeneous and long-tailed data. Comparative results with the state-of-the-art FL methods also validate the superiority of CReFF. Our code is available at https://github.com/shangxinyi/CReFF-FL. Xinyi Shang, Yang Lu 0009, Gang Huang 0004, Hanzi Wang |
IJCAI | 1 |