VLDB 2026 Research / reviewers in the wild / expert
Yucong Dai
dblp:312/7771
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0000-3729-3650ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Trustworthy machine learning · 29% Transfer learning and domain adaptation · 19% Language models and text generation · 17% | |
| Databases, data mining, and information retrieval
2 papers |
Machine learning and data management · 40% Data stream processing · 21% Data mining · 21% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › stochastic gradient descent
adaptive stochastic gradient descent |
0.9 | 1 | 2025 | Adaptive Learning in Imbalanced Data Streams With Unpredictable Feature Evolution · IEEE Trans. Knowl. Data Eng. 2025 |
Machine learning › Trustworthy machine learning › fairness
causal fairness |
0.9 | 1 | 2025 | Towards counterfactual fairness through auxiliary variables · ICLR 2025 |
Machine learning › Trustworthy machine learning › fairness › causal fairness
counterfactual fairness |
0.9 | 1 | 2025 | Towards counterfactual fairness through auxiliary variables · ICLR 2025 |
Machine learning › Reinforcement learning › regret minimization
dynamic regret minimization |
0.9 | 1 | 2025 | Label Shift Meets Online Learning: Ensuring Consistent Adaptation with Universal Dynamic Regret · CVPR 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.9 | 1 | 2025 | Towards counterfactual fairness through auxiliary variables · ICLR 2025 |
Machine learning › Transfer learning and domain adaptation
label shift |
0.9 | 1 | 2025 | Label Shift Meets Online Learning: Ensuring Consistent Adaptation with Universal Dynamic Regret · CVPR 2025 |
Machine learning › Transfer learning and domain adaptation › label shift
online label shift |
0.9 | 1 | 2025 | Label Shift Meets Online Learning: Ensuring Consistent Adaptation with Universal Dynamic Regret · CVPR 2025 |
Machine learning › Learning theory
online learning |
0.9 | 1 | 2025 | Label Shift Meets Online Learning: Ensuring Consistent Adaptation with Universal Dynamic Regret · CVPR 2025 |
Data stream processing › evolving data
concept drift |
0.9 | 1 | 2025 | Adaptive Learning in Imbalanced Data Streams With Unpredictable Feature Evolution · IEEE Trans. Knowl. Data Eng. 2025 |
Machine learning and data management
feature space evolution |
0.9 | 1 | 2025 | Adaptive Learning in Imbalanced Data Streams With Unpredictable Feature Evolution · IEEE Trans. Knowl. Data Eng. 2025 |
Natural language and speech › Language models and text generation
dataset refinement |
0.8 | 1 | 2024 | SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024 |
Natural language and speech › Language models and text generation
instruction tuning |
0.8 | 1 | 2024 | SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024 |
Data integration and cleaning
data curation |
0.8 | 1 | 2024 | SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024 |
Machine learning and data management
data selection |
0.8 | 1 | 2024 | SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
data-efficient learning |
0.2 | 1 | 2024 | SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
ensemble methods · 2.6projected technique · 1.7shapley value · 1.5dataset refinement · 1.5reweighting · 0.9re-weighting · 0.9optimistic online algorithm · 0.9exogenous variable · 0.9causal reasoning · 0.9auxiliary variables · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Statistical and computational trade-offs in imbalanced kernel clustering
Jing Zhang 0064, Yucong Dai, Chenping Hou |
Frontiers Comput. Sci. | 2 |
| 2025 | Label Shift Meets Online Learning: Ensuring Consistent Adaptation with Universal Dynamic RegretabstractLabel shift, which investigates the adaptation of label distributions between the fixed source and target domains, has attracted significant research interests and broad applications in offline settings. In real-world scenarios, however, data often arrives as a continuous stream. Addressing label shift in online learning settings is paramount. Existing strategies, which tailor traditional offline label shift techniques to online settings, have degraded performance due to the inconsistent estimation of label distributions and violation of convex assumption for theoretical guarantee. In this paper, we propose a novel method to ensure consistent adaptation to online label shift. We construct a new convex risk estimator that is pivotal for both online optimization and theoretical analysis. Furthermore, we enhance an optimistic online algorithm as the base learner and refine the classifier using an ensemble method. Theoretically, we derive a universal dynamic regret which achieves minimax optimal. Extensive experiments on both real-world datasets and human motion task demonstrate the superiority of our method comparing existing methods. Yucong Dai, Shilin Gu, Ruidong Fan, Chao Xu 0008, Chenping Hou |
CVPR | 1 |
| 2025 | Towards counterfactual fairness through auxiliary variablesabstractThe challenge of balancing fairness and predictive accuracy in machine learning models, especially when sensitive attributes such as race, gender, or age are considered, has motivated substantial research in recent years. Counterfactual fairness ensures that predictions remain consistent across counterfactual variations of sensitive attributes, which is a crucial concept in addressing societal biases.
However, existing counterfactual fairness approaches usually overlook intrinsic information about sensitive features, limiting their ability to achieve fairness while simultaneously maintaining performance. To tackle this challenge, we introduce EXOgenous Causal reasoning (EXOC), a novel causal reasoning framework motivated by exogenous variables. It leverages auxiliary variables to uncover intrinsic properties that give rise to sensitive attributes. Our framework explicitly defines an auxiliary node and a control node that contribute to counterfactual fairness and control the information flow within the model. Our evaluation, conducted on synthetic and real-world datasets, validates EXOC's superiority, showing that it outperforms state-of-the-art approaches in achieving counterfactual fairness without sacrificing accuracy. Our code is available at https://github.com/CASE-Lab-UMD/counterfactual_fairness_2025. Bowei Tian, Shwai He, Wanghao Ye, Guoheng Sun, Yucong Dai, Yongkai Wu, Ang Li 0005 |
ICLR | 6 |
| 2025 | Adaptive Learning in Imbalanced Data Streams With Unpredictable Feature EvolutionabstractLearning from data streams collected sequentially over time are widely spread in real-world applications. Previous methods typically assume that the data stream has a feature space with a fixed or clearly defined evolution pattern, as well as a balanced class distribution. However, in many practical scenarios, such as environmental monitoring systems, the frequency of anomalous events is significantly imbalanced compared to normal ones and the feature space dynamically changes due to ecological evolution and sensor lifespan. To alleviate this important but rarely studied problem, we propose the Adaptive Learning in Imbalace data streams with Unpredictable feature evolution (ALIU) algorithm. As data streams with imbalanced class distribution arrive, ALIU first mitigates the model's bias for the majority class by reweighting the adaptive gradient descent magnitudes between different classes. Then, a new loss function is proposed that simultaneously focuses on misclassifications and maintains model robustness. Further, when imbalanced data streams arrive with feature evolutions, we reuse the previously learned model and update the incomplete and augmented features by adopting the adaptive gradient strategy and ensemble method, respectively. Finally, we utilize the projected technique to build a sparse yet efficient model. Based on a few common and mild assumptions, we theoretically analyze that the ALIU satisfies a sub-linear regret bound under both convex and strong convex loss functions and the performance of model can be improved with the assistance of old features. Besides, extensive experimental results further demonstrate the effectiveness of our proposed algorithm. Jiahang Tu, Xijia Tang, Shilin Gu, Yucong Dai, Ruidong Fan, Chenping Hou |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Fair Weak-Supervised Learning: A Multiple-Instance Learning ApproachabstractWith the prevalence of machine learning in many high-stakes decision-making processes, e.g., hiring and admission, it is important to take fairness into account when practitioners design and deploy machine learning models, especially in scenarios with imperfectly labeled data. Multiple-Instance Learning (MIL) is a weakly supervised approach where instances are grouped in labeled bags, each containing several instances sharing the same label. However, current fairness-centric methods in machine learning often fall short when applied to MIL due to their reliance on instance-level labels. In this work, we introduce a Fair Multiple-Instance Learning (FMIL) framework to ensure fairness in weakly supervised learning. In particular, our method bridges the gap between bag-level and instance-level labeling by leveraging the bag labels, inferring high-confidence instance labels to improve both accuracy and fairness in MIL classifiers. Comprehensive experiments underscore that our FMIL framework substantially reduces biases in MIL without compromising accuracy. Yucong Dai, Xiangyu Jiang, Yaowei Hu 0001, Lu Zhang 0021, Yongkai Wu |
IJCNN | 1 |
| 2024 | SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-TuningabstractThe pre-trained Large Language Models (LLMs) can be adapted for many downstream tasks and tailored to align with human preferences through fine-tuning. Recent studies have discovered that LLMs can achieve desirable performance with only a small amount of high-quality data, suggesting that a large portion of the data in these extensive datasets is redundant or even harmful. Identifying high-quality data from vast datasets to curate small yet effective datasets has emerged as a critical challenge. In this paper, we introduce SHED, an automated dataset refinement framework based on Shapley value for instruction fine-tuning. SHED eliminates the need for human intervention or the use of commercial LLMs. Moreover, the datasets curated through SHED exhibit transferability, indicating they can be reused across different LLMs with consistently high performance. We conduct extensive experiments to evaluate the datasets curated by SHED. The results demonstrate SHED's superiority over state-of-the-art methods across various tasks and LLMs; notably, datasets comprising only 10% of the original data selected by SHED achieve performance comparable to or surpassing that of the full datasets. Yexiao He, Zheyu Shen, Guoheng Sun, Yucong Dai, Yongkai Wu, Hongyi Wang 0001, Ang Li 0005 |
NeurIPS | 5 |
| 2023 | Fair Selection through Kernel Density EstimationabstractWith the prevalence of machine learning in many high-stakes decision-making processes, e.g., hiring and admission, it is important to take fairness into consideration when practitioners design and deploy machine learning models. Although many approaches have been developed for fair machine learning, most of them focus on classification. In this paper, we target a notable but under-explored task, selection, where the number of selected individuals cannot exceed a pre-defined budget, such as employee hiring or university admission with limited positions or capabilities. In the selection task, existing fairness notions designed for classification are not suitable. In particular, our experimental results show that the selection models subject to common fairness notions may still make biased predictions against the underrepresented group. Hence, we propose a novel fairness notion, Selection Parity, which captures the demographic diversity among the selected groups in this restricted selection problem. Since the selection of qualified individuals with a fixed budget is non-differentiable, existing fairness regularization terms cannot be directly integrated with the selection task. To close the gap, we develop a novel in-processing framework named Fair Selection with the Differentiable Distribution Difference constraint (FS-DD), which incorporates a differentiable constraint into the training process and produces fair decisions for selection problems. Our theoretical analysis shows that common fairness metrics are bounded by the proposed Distribution Difference measurement. In other words, the FS-DD framework can guarantee fairness with regard to the common existing fairness metrics. We evaluate the performance of our method as well as several baselines on four real-world datasets. The experimental results demonstrate that the proposed method achieves fairness in various selection settings. In addition, the proposed method has a better fairness-accuracy trade-off compared with existing baseline methods. Xiangyu Jiang, Yucong Dai, Yongkai Wu |
IJCNN | 2 |
| 2021 | Stereo Superpixel Segmentation Via Dual-Attention Fusion NetworksabstractStereo image pairs can improve performance of many tasks benefiting from the additional information obtained from a second viewpoint when compared with single images. Existing superpixel segmentation algorithms for stereo images mostly adopt single images as input, and neglect the correspondence between the left and right views. In this work, we consider to exploit the depth information between stereo image pairs, and propose an end-to-end dual-attention fusion network for stereo images to generate parallax-consistency superpixels. We first utilize a deep convolution network to extract the deep features of stereo images. Then, to effectively utilize the additional information from the other view, features of the left and right views is integrated by a parallax attention and channel attention mechanism. Finally, the stereo superpixels are generated by a differentiable clustering algorithm, which is end-to-end trainable with deep learning networks. Comprehensive experimental results demonstrate that our method can outperform the state-of-the-art performance on the KITTI2015 and Cityscapes dataset. Yajuan Du, Hua Li 0012, Yucong Dai |
ICME | 4 |