VLDB 2026 Research / reviewers in the wild / expert
Yihao Xue
dblp:271/2194
· DBLP profile ↗
13ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-3310-4864ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Representation and self-supervised learning · 40% Trustworthy machine learning · 34% Transfer learning and domain adaptation · 9% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
contrastive learning |
2.0 | 3 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023 Investigating Why Contrastive Learning Benefits Robustness against Label Noise · ICML 2022 |
Machine learning › Trustworthy machine learning
robustness |
1.6 | 3 | 2024 | Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024 Investigating Why Contrastive Learning Benefits Robustness against Label Noise · ICML 2022 Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Natural language and speech › Language models and text generation › alignment › scalable oversight
weak-to-strong generalization |
0.9 | 1 | 2025 | Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions · ICML 2025 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.8 | 1 | 2024 | Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings · ICML 2024 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot adaptation |
0.8 | 1 | 2024 | Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings · ICML 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning |
0.8 | 1 | 2024 | Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
projection head |
0.8 | 1 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation analysis
representation learning theory |
0.8 | 1 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.8 | 1 | 2024 | Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.8 | 1 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
feature suppression |
0.7 | 1 | 2023 | Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023 |
Machine learning › Learning theory › implicit bias
implicit bias of gradient descent |
0.7 | 1 | 2023 | Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023 |
Machine learning › Learning theory › inductive bias
simplicity bias |
0.7 | 1 | 2023 | Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023 |
Machine learning › Trustworthy machine learning › robustness › robust learning
fine-tuning robustness |
0.6 | 1 | 2022 | Investigating Why Contrastive Learning Benefits Robustness against Label Noise · ICML 2022 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
label noise robustness |
0.6 | 1 | 2022 | Investigating Why Contrastive Learning Benefits Robustness against Label Noise · ICML 2022 |
Machine learning › Efficient and distributed learning
federated learning |
0.5 | 1 | 2021 | Toward Understanding the Influence of Individual Clients in Federated Learning · AAAI 2021 |
Machine learning › Trustworthy machine learning › interpretability
model debugging |
0.5 | 1 | 2021 | Toward Understanding the Influence of Individual Clients in Federated Learning · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
theoretical analysis · 2.0representation analysis · 0.9kernel methods · 0.9linear classifier · 0.8layer-wise feature weighting analysis · 0.8intra-class contrasting · 0.8inter-class feature sharing · 0.8embedding mixing · 0.8contrastive loss · 0.8stochastic gradient descent · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TPPPA: A Triangular Partition Path Planning Algorithm for UAV Coverage in Irregular Areas
Gordon Owusu Boateng, Yihao Xue, Bintao Hu, Xingzhen Duan |
IWCMC | 4 |
| 2026 | Dual Attention-Guided Ensemble Framework for High-Speed Train Fault Diagnosis: Optimizing Multiscale Features From Multiple SensorsabstractHealth monitoring and fault diagnosis of high-speed train traction systems are essential for maintaining reliable operation. To effectively process multisensor signals and avoid overfitting, ensemble learning methods are employed, leveraging multiple base models to integrate data from various sensors and enhance fault diagnosis performance. However, conventional ensemble frameworks are often burdened by excessive model parameters, limiting their applicability on edge computing processors. To address these challenges, this study proposes a novel dual attention-guided ensemble framework. This framework incorporates multiple multiscale feature attention (MFA) modules and a decision fusion attention (DFA) module, designed to capture critical features from multisensor signals, optimize the capacity of prominent feature extraction, and simultaneously reducing trainable parameters. The proposed ensemble framework is validated on the hardware-in-the-loop (HIL) simulation platform for high-speed train traction control systems, with experimental results demonstrating its superior effectiveness over several recently published ensemble learning methods. Yihao Xue, Rui Yang 0007, Xiaohan Chen 0003, Yifan Zhan, Baoye Song, Zidong Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical PredictionsabstractWeak-to-Strong Generalization (W2SG), where a weak model supervises a stronger one, serves as an important analogy for understanding how humans might guide superhuman intelligence in the future. Promising empirical results revealed that a strong model can surpass its weak supervisor. While recent work has offered theoretical insights into this phenomenon, a clear understanding of the interactions between weak and strong models that drive W2SG remains elusive. We investigate W2SG through a theoretical lens and show that it can be characterized using kernels derived from the principal components of weak and strong models' internal representations. These kernels can be used to define a space that, at a high level, captures what the weak model is unable to learn but is learnable by the strong model. The projection of labels onto this space quantifies how much the strong model falls short of its full potential due to weak supervision. This characterization also provides insights into how certain errors in weak supervision can be corrected by the strong model, regardless of overfitting. Our theory has significant practical implications, providing a representation-based metric that predicts W2SG performance trends without requiring labels, as shown in experiments on molecular predictions with transformers and 5 NLP tasks involving 52 LLMs. Yihao Xue, Jiping Li, Baharan Mirzasoleiman |
ICML | 1 |
| 2025 | Separable Convolutional Network-Based Fault Diagnosis for High-Speed Train: A Gossip Strategy-Based Optimization ApproachabstractWith the rapid development of high-speed train, health monitoring of high-speed train traction power system has gradually become a popular research topic. The traction asynchronous motor, as a key component in the traction power systems, greatly affects the reliability, stability, and safety of high-speed train operation. Normally, when faults occur, the train needs to immediately slow down or even stop to avoid unimaginable losses, resulting in limited fault data. Traditional data-driven fault diagnosis methods may face the local optimum problem during the optimization process when training samples are insufficient. In this study, a novel gossip strategy-based fault diagnosis method is proposed to prevent the local optimum problem, thus improving fault diagnosis performance. The proposed gossip strategy-based fault diagnosis method is validated on the hardware-in-the-loop high-speed train traction control system simulation platform, and the experimental results unequivocally show that the proposed method outperforms other well-known methods. Yihao Xue, Rui Yang 0007, Xiaohan Chen 0003, Baoye Song, Zidong Wang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Investigating the Benefits of Projection Head for Representation LearningabstractAn effective technique for obtaining high-quality representations is adding a projection head on top of the encoder during training, then discarding it and using the pre-projection representations. Despite its proven practical effectiveness, the reason behind the success of this technique is poorly understood. The pre-projection representations are not directly optimized by the loss function, raising the question: what makes them better? In this work, we provide a rigorous theoretical answer to this question. We start by examining linear models trained with self-supervised contrastive loss. We reveal that the implicit bias of training algorithms leads to layer-wise progressive feature weighting, where features become increasingly unequal as we go deeper into the layers. Consequently, lower layers tend to have more normalized and less specialized representations. We theoretically characterize scenarios where such representations are more beneficial, highlighting the intricate interplay between data augmentation and input features. Additionally, we demonstrate that introducing non-linearity into the network allows lower layers to learn features that are completely absent in higher layers. Finally, we show how this mechanism improves the robustness in supervised contrastive learning and supervised learning. We empirically validate our results through various experiments on CIFAR-10/100, UrbanCars and shifted versions of ImageNet. We also introduce a potential alternative to projection head, which offers a more interpretable and controllable design. Yihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi 0004, Baharan Mirzasoleiman |
ICLR | 1 |
| 2024 | Understanding the Robustness of Multi-modal Contrastive Learning to Distribution ShiftabstractRecently, multimodal contrastive learning (MMCL) approaches, such as CLIP, have achieved a remarkable success in learning representations that are robust against distribution shift and generalize to new domains. Despite the empirical success, the mechanism behind learning such generalizable representations is not understood. In this work, we rigorously analyze this problem and
uncover two mechanisms behind MMCL's robustness: \emph{intra-class contrasting}, which allows the model to learn features with a high variance, and \emph{inter-class feature sharing}, where annotated details in one class help learning other classes better. Both mechanisms prevent spurious features that are over-represented in the training data to overshadow the generalizable core features. This yields superior zero-shot classification accuracy under distribution shift. Furthermore, we theoretically demonstrate the benefits of using rich captions on robustness and explore the effect of annotating different types of details in the captions. We validate our theoretical findings through experiments, including a well-designed synthetic experiment and an experiment involving training CLIP models on MSCOCO/Conceptual Captions and evaluating them on shifted ImageNets. Yihao Xue, Siddharth Joshi 0004, Baharan Mirzasoleiman |
ICLR | 1 |
| 2024 | Few-shot Adaptation to Distribution Shifts By Mixing Source and Target EmbeddingsabstractPretrained machine learning models need to be adapted to distribution shifts when deployed in new target environments. When obtaining labeled data from the target distribution is expensive, few-shot adaptation with only a few examples from the target distribution becomes essential. In this work, we propose MixPro, a lightweight and highly data-efficient approach for few-shot adaptation. MixPro first generates a relatively large dataset by mixing (linearly combining) pre-trained embeddings of large source data with those of the few target examples. This process preserves important features of both source and target distributions, while mitigating the specific noise in the small target data. Then, it trains a linear classifier on the mixed embeddings to effectively adapts the model to the target distribution without overfitting the small target data. Theoretically, we demonstrate the advantages of MixPro over previous methods. Our experiments, conducted across various model architectures on 8 datasets featuring different types of distribution shifts, reveal that MixPro can outperform baselines by as much as 7%, with only 2-4 target examples. Yihao Xue, Ali Payani, Yu Yang 0007, Baharan Mirzasoleiman |
ICML | 1 |
| 2024 | Investigating the Impact of Model Width and Density on Generalization in Presence of Label NoiseabstractIncreasing the size of overparameterized neural networks has been a key in achieving state-of-the-art performance. This is captured by the double descent phenomenon, where the test loss follows a decreasing-increasing-decreasing pattern (or sometimes monotonically decreasing) as model width increases. However, the effect of label noise on the test loss curve has not been fully explored. In this work, we uncover an intriguing phenomenon where label noise leads to a final ascent in the originally observed double descent curve. Specifically, under a sufficiently large noise-to-sample-size ratio, optimal generalization is achieved at intermediate widths. Through theoretical analysis, we attribute this phenomenon to the shape transition of test loss variance induced by label noise. Furthermore, we extend the final ascent phenomenon to model density and provide the first theoretical characterization showing that reducing density by randomly dropping trainable parameters improves generalization under label noise. We also thoroughly examine the roles of regularization and sample size. Surprisingly, we find that larger $\ell_2$ regularization and robust learning methods against label noise exacerbate the final ascent. We confirm the validity of our findings through extensive experiments on ReLu networks trained on MNIST, ResNets/ViT trained on CIFAR-10/100, and InceptionResNet-v2 trained on Stanford Cars with real-world noisy labels. Yihao Xue, Kyle Whitecross, Baharan Mirzasoleiman |
UAI | 1 |
| 2023 | Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature SuppressionabstractContrastive learning (CL) has emerged as a powerful technique for representation learning, with or without label supervision. However, supervised CL is prone to collapsing representations of subclasses within a class by not capturing all their features, and unsupervised CL may suppress harder class-relevant features by focusing on learning easy class-irrelevant features; both significantly compromise representation quality. Yet, there is no theoretical understanding of class collapse or feature suppression at test time. We provide the first unified theoretically rigorous framework to determine which features are learnt by CL. Our analysis indicate that, perhaps surprisingly, bias of (stochastic) gradient descent towards finding simpler solutions is a key factor in collapsing subclass representations and suppressing harder class-relevant features. Moreover, we present increasing embedding dimensionality and improving the quality of data augmentations as two theoretically motivated solutions to feature suppression. We also provide the first theoretical explanation for why employing supervised and unsupervised CL together yields higher-quality representations, even when using commonly-used stochastic gradient methods. Yihao Xue, Siddharth Joshi 0004, Eric Gan, Baharan Mirzasoleiman |
ICML | 1 |
| 2023 | A novel momentum prototypical neural network to cross-domain fault diagnosis for rotating machinery subject to cold-start
Xiaohan Chen 0003, Rui Yang 0007, Yihao Xue, Baoye Song, Maiying Zhong |
Neurocomputing | 3 |
| 2022 | Investigating Why Contrastive Learning Benefits Robustness against Label NoiseabstractSelf-supervised Contrastive Learning (CL) has been recently shown to be very effective in preventing deep networks from overfitting noisy labels. Despite its empirical success, the theoretical understanding of the effect of contrastive learning on boosting robustness is very limited. In this work, we rigorously prove that the representation matrix learned by contrastive learning boosts robustness, by having: (i) one prominent singular value corresponding to each sub-class in the data, and significantly smaller remaining singular values; and (ii) a large alignment between the prominent singular vectors and the clean labels of each sub-class. The above properties enable a linear layer trained on such representations to effectively learn the clean labels without overfitting the noise. We further show that the low-rank structure of the Jacobian of deep networks pre-trained with contrastive learning allows them to achieve a superior performance initially, when fine-tuned on noisy labels. Finally, we demonstrate that the initial robustness provided by contrastive learning enables robust training methods to achieve state-of-the-art performance under extreme noise levels, e.g., an average of 27.18% and 15.58% increase in accuracy on CIFAR-10 and CIFAR-100 with 80% symmetric noisy labels, and 4.11% increase in accuracy on WebVision. Yihao Xue, Kyle Whitecross, Baharan Mirzasoleiman |
ICML | 1 |
| 2022 | Polyphonic music generation generative adversarial network with Markov decision process
Wenkai Huang 0001, Yihao Xue, Zefeng Xu, Guanglong Peng |
Multim. Tools Appl. | 2 |
| 2021 | Toward Understanding the Influence of Individual Clients in Federated LearningabstractFederated learning allows mobile clients to jointly train a global model without sending their private data to a central server. Extensive works have studied the performance guarantee of the global model, however, it is still unclear how each individual client influences the collaborative training process. In this work, we defined a new notion, called {\em Fed-Influence}, to quantify this influence over the model parameters, and proposed an effective and efficient algorithm to estimate this metric. In particular, our design satisfies several desirable properties: (1) it requires neither retraining nor retracing, adding only linear computational overhead to clients and the server; (2) it strictly maintains the tenets of federated learning, without revealing any client's local private data; and (3) it works well on both convex and non-convex loss functions, and does not require the final model to be optimal. Empirical results on a synthetic dataset and the FEMNIST dataset demonstrate that our estimation method can approximate Fed-Influence with small bias. Further, we show an application of Fed-Influence in model debugging. Yihao Xue, Chaoyue Niu, Zhenzhe Zheng 0001, Shaojie Tang 0001, Chengfei Lyu, Fan Wu 0006, Guihai Chen |
AAAI | 1 |