Yihao Xue

dblp:271/2194 · DBLP profile ↗
← Back
13ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-3310-4864ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Representation and self-supervised learning · 40% Trustworthy machine learning · 34% Transfer learning and domain adaptation · 9%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
2.032024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023
Investigating Why Contrastive Learning Benefits Robustness against Label Noise · ICML 2022
Machine learning › Trustworthy machine learning
robustness
1.632024
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024
Investigating Why Contrastive Learning Benefits Robustness against Label Noise · ICML 2022
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Natural language and speech › Language models and text generation › alignment › scalable oversight
weak-to-strong generalization
0.912025
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions · ICML 2025
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.812024
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings · ICML 2024
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot adaptation
0.812024
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings · ICML 2024
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning
0.812024
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024
Machine learning › Representation and self-supervised learning › contrastive learning
projection head
0.812024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Representation and self-supervised learning › representation analysis
representation learning theory
0.812024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.812024
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.812024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Representation and self-supervised learning › contrastive learning
feature suppression
0.712023
Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023
Machine learning › Learning theory › implicit bias
implicit bias of gradient descent
0.712023
Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023
Machine learning › Learning theory › inductive bias
simplicity bias
0.712023
Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023
Machine learning › Trustworthy machine learning › robustness › robust learning
fine-tuning robustness
0.612022
Investigating Why Contrastive Learning Benefits Robustness against Label Noise · ICML 2022
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
label noise robustness
0.612022
Investigating Why Contrastive Learning Benefits Robustness against Label Noise · ICML 2022
Machine learning › Efficient and distributed learning
federated learning
0.512021
Toward Understanding the Influence of Individual Clients in Federated Learning · AAAI 2021
Machine learning › Trustworthy machine learning › interpretability
model debugging
0.512021
Toward Understanding the Influence of Individual Clients in Federated Learning · AAAI 2021

Methods — techniques the papers use, named apart from their topics

theoretical analysis · 2.0representation analysis · 0.9kernel methods · 0.9linear classifier · 0.8layer-wise feature weighting analysis · 0.8intra-class contrasting · 0.8inter-class feature sharing · 0.8embedding mixing · 0.8contrastive loss · 0.8stochastic gradient descent · 0.7
YearPublicationVenuePosition
2026 TPPPA: A Triangular Partition Path Planning Algorithm for UAV Coverage in Irregular Areas
Gordon Owusu Boateng, Yihao Xue, Bintao Hu, Xingzhen Duan
IWCMC4
2026 Dual Attention-Guided Ensemble Framework for High-Speed Train Fault Diagnosis: Optimizing Multiscale Features From Multiple Sensors
abstract
Health monitoring and fault diagnosis of high-speed train traction systems are essential for maintaining reliable operation. To effectively process multisensor signals and avoid overfitting, ensemble learning methods are employed, leveraging multiple base models to integrate data from various sensors and enhance fault diagnosis performance. However, conventional ensemble frameworks are often burdened by excessive model parameters, limiting their applicability on edge computing processors. To address these challenges, this study proposes a novel dual attention-guided ensemble framework. This framework incorporates multiple multiscale feature attention (MFA) modules and a decision fusion attention (DFA) module, designed to capture critical features from multisensor signals, optimize the capacity of prominent feature extraction, and simultaneously reducing trainable parameters. The proposed ensemble framework is validated on the hardware-in-the-loop (HIL) simulation platform for high-speed train traction control systems, with experimental results demonstrating its superior effectiveness over several recently published ensemble learning methods.
Yihao Xue, Rui Yang 0007, Xiaohan Chen 0003, Yifan Zhan, Baoye Song, Zidong Wang 0001
IEEE Trans. Intell. Transp. Syst.1
2025 Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
abstract
Weak-to-Strong Generalization (W2SG), where a weak model supervises a stronger one, serves as an important analogy for understanding how humans might guide superhuman intelligence in the future. Promising empirical results revealed that a strong model can surpass its weak supervisor. While recent work has offered theoretical insights into this phenomenon, a clear understanding of the interactions between weak and strong models that drive W2SG remains elusive. We investigate W2SG through a theoretical lens and show that it can be characterized using kernels derived from the principal components of weak and strong models' internal representations. These kernels can be used to define a space that, at a high level, captures what the weak model is unable to learn but is learnable by the strong model. The projection of labels onto this space quantifies how much the strong model falls short of its full potential due to weak supervision. This characterization also provides insights into how certain errors in weak supervision can be corrected by the strong model, regardless of overfitting. Our theory has significant practical implications, providing a representation-based metric that predicts W2SG performance trends without requiring labels, as shown in experiments on molecular predictions with transformers and 5 NLP tasks involving 52 LLMs.
Yihao Xue, Jiping Li, Baharan Mirzasoleiman
ICML1
2025 Separable Convolutional Network-Based Fault Diagnosis for High-Speed Train: A Gossip Strategy-Based Optimization Approach
abstract
With the rapid development of high-speed train, health monitoring of high-speed train traction power system has gradually become a popular research topic. The traction asynchronous motor, as a key component in the traction power systems, greatly affects the reliability, stability, and safety of high-speed train operation. Normally, when faults occur, the train needs to immediately slow down or even stop to avoid unimaginable losses, resulting in limited fault data. Traditional data-driven fault diagnosis methods may face the local optimum problem during the optimization process when training samples are insufficient. In this study, a novel gossip strategy-based fault diagnosis method is proposed to prevent the local optimum problem, thus improving fault diagnosis performance. The proposed gossip strategy-based fault diagnosis method is validated on the hardware-in-the-loop high-speed train traction control system simulation platform, and the experimental results unequivocally show that the proposed method outperforms other well-known methods.
Yihao Xue, Rui Yang 0007, Xiaohan Chen 0003, Baoye Song, Zidong Wang 0001
IEEE Trans. Ind. Informatics1
2024 Investigating the Benefits of Projection Head for Representation Learning
abstract
An effective technique for obtaining high-quality representations is adding a projection head on top of the encoder during training, then discarding it and using the pre-projection representations. Despite its proven practical effectiveness, the reason behind the success of this technique is poorly understood. The pre-projection representations are not directly optimized by the loss function, raising the question: what makes them better? In this work, we provide a rigorous theoretical answer to this question. We start by examining linear models trained with self-supervised contrastive loss. We reveal that the implicit bias of training algorithms leads to layer-wise progressive feature weighting, where features become increasingly unequal as we go deeper into the layers. Consequently, lower layers tend to have more normalized and less specialized representations. We theoretically characterize scenarios where such representations are more beneficial, highlighting the intricate interplay between data augmentation and input features. Additionally, we demonstrate that introducing non-linearity into the network allows lower layers to learn features that are completely absent in higher layers. Finally, we show how this mechanism improves the robustness in supervised contrastive learning and supervised learning. We empirically validate our results through various experiments on CIFAR-10/100, UrbanCars and shifted versions of ImageNet. We also introduce a potential alternative to projection head, which offers a more interpretable and controllable design.
Yihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi 0004, Baharan Mirzasoleiman
ICLR1
2024 Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift
abstract
Recently, multimodal contrastive learning (MMCL) approaches, such as CLIP, have achieved a remarkable success in learning representations that are robust against distribution shift and generalize to new domains. Despite the empirical success, the mechanism behind learning such generalizable representations is not understood. In this work, we rigorously analyze this problem and uncover two mechanisms behind MMCL's robustness: \emph{intra-class contrasting}, which allows the model to learn features with a high variance, and \emph{inter-class feature sharing}, where annotated details in one class help learning other classes better. Both mechanisms prevent spurious features that are over-represented in the training data to overshadow the generalizable core features. This yields superior zero-shot classification accuracy under distribution shift. Furthermore, we theoretically demonstrate the benefits of using rich captions on robustness and explore the effect of annotating different types of details in the captions. We validate our theoretical findings through experiments, including a well-designed synthetic experiment and an experiment involving training CLIP models on MSCOCO/Conceptual Captions and evaluating them on shifted ImageNets.
Yihao Xue, Siddharth Joshi 0004, Baharan Mirzasoleiman
ICLR1
2024 Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
abstract
Pretrained machine learning models need to be adapted to distribution shifts when deployed in new target environments. When obtaining labeled data from the target distribution is expensive, few-shot adaptation with only a few examples from the target distribution becomes essential. In this work, we propose MixPro, a lightweight and highly data-efficient approach for few-shot adaptation. MixPro first generates a relatively large dataset by mixing (linearly combining) pre-trained embeddings of large source data with those of the few target examples. This process preserves important features of both source and target distributions, while mitigating the specific noise in the small target data. Then, it trains a linear classifier on the mixed embeddings to effectively adapts the model to the target distribution without overfitting the small target data. Theoretically, we demonstrate the advantages of MixPro over previous methods. Our experiments, conducted across various model architectures on 8 datasets featuring different types of distribution shifts, reveal that MixPro can outperform baselines by as much as 7%, with only 2-4 target examples.
Yihao Xue, Ali Payani, Yu Yang 0007, Baharan Mirzasoleiman
ICML1
2024 Investigating the Impact of Model Width and Density on Generalization in Presence of Label Noise
abstract
Increasing the size of overparameterized neural networks has been a key in achieving state-of-the-art performance. This is captured by the double descent phenomenon, where the test loss follows a decreasing-increasing-decreasing pattern (or sometimes monotonically decreasing) as model width increases. However, the effect of label noise on the test loss curve has not been fully explored. In this work, we uncover an intriguing phenomenon where label noise leads to a final ascent in the originally observed double descent curve. Specifically, under a sufficiently large noise-to-sample-size ratio, optimal generalization is achieved at intermediate widths. Through theoretical analysis, we attribute this phenomenon to the shape transition of test loss variance induced by label noise. Furthermore, we extend the final ascent phenomenon to model density and provide the first theoretical characterization showing that reducing density by randomly dropping trainable parameters improves generalization under label noise. We also thoroughly examine the roles of regularization and sample size. Surprisingly, we find that larger $\ell_2$ regularization and robust learning methods against label noise exacerbate the final ascent. We confirm the validity of our findings through extensive experiments on ReLu networks trained on MNIST, ResNets/ViT trained on CIFAR-10/100, and InceptionResNet-v2 trained on Stanford Cars with real-world noisy labels.
Yihao Xue, Kyle Whitecross, Baharan Mirzasoleiman
UAI1
2023 Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression
abstract
Contrastive learning (CL) has emerged as a powerful technique for representation learning, with or without label supervision. However, supervised CL is prone to collapsing representations of subclasses within a class by not capturing all their features, and unsupervised CL may suppress harder class-relevant features by focusing on learning easy class-irrelevant features; both significantly compromise representation quality. Yet, there is no theoretical understanding of class collapse or feature suppression at test time. We provide the first unified theoretically rigorous framework to determine which features are learnt by CL. Our analysis indicate that, perhaps surprisingly, bias of (stochastic) gradient descent towards finding simpler solutions is a key factor in collapsing subclass representations and suppressing harder class-relevant features. Moreover, we present increasing embedding dimensionality and improving the quality of data augmentations as two theoretically motivated solutions to feature suppression. We also provide the first theoretical explanation for why employing supervised and unsupervised CL together yields higher-quality representations, even when using commonly-used stochastic gradient methods.
Yihao Xue, Siddharth Joshi 0004, Eric Gan, Baharan Mirzasoleiman
ICML1
2023 A novel momentum prototypical neural network to cross-domain fault diagnosis for rotating machinery subject to cold-start
Xiaohan Chen 0003, Rui Yang 0007, Yihao Xue, Baoye Song, Maiying Zhong
Neurocomputing3
2022 Investigating Why Contrastive Learning Benefits Robustness against Label Noise
abstract
Self-supervised Contrastive Learning (CL) has been recently shown to be very effective in preventing deep networks from overfitting noisy labels. Despite its empirical success, the theoretical understanding of the effect of contrastive learning on boosting robustness is very limited. In this work, we rigorously prove that the representation matrix learned by contrastive learning boosts robustness, by having: (i) one prominent singular value corresponding to each sub-class in the data, and significantly smaller remaining singular values; and (ii) a large alignment between the prominent singular vectors and the clean labels of each sub-class. The above properties enable a linear layer trained on such representations to effectively learn the clean labels without overfitting the noise. We further show that the low-rank structure of the Jacobian of deep networks pre-trained with contrastive learning allows them to achieve a superior performance initially, when fine-tuned on noisy labels. Finally, we demonstrate that the initial robustness provided by contrastive learning enables robust training methods to achieve state-of-the-art performance under extreme noise levels, e.g., an average of 27.18% and 15.58% increase in accuracy on CIFAR-10 and CIFAR-100 with 80% symmetric noisy labels, and 4.11% increase in accuracy on WebVision.
Yihao Xue, Kyle Whitecross, Baharan Mirzasoleiman
ICML1
2022 Polyphonic music generation generative adversarial network with Markov decision process
Wenkai Huang 0001, Yihao Xue, Zefeng Xu, Guanglong Peng
Multim. Tools Appl.2
2021 Toward Understanding the Influence of Individual Clients in Federated Learning
abstract
Federated learning allows mobile clients to jointly train a global model without sending their private data to a central server. Extensive works have studied the performance guarantee of the global model, however, it is still unclear how each individual client influences the collaborative training process. In this work, we defined a new notion, called {\em Fed-Influence}, to quantify this influence over the model parameters, and proposed an effective and efficient algorithm to estimate this metric. In particular, our design satisfies several desirable properties: (1) it requires neither retraining nor retracing, adding only linear computational overhead to clients and the server; (2) it strictly maintains the tenets of federated learning, without revealing any client's local private data; and (3) it works well on both convex and non-convex loss functions, and does not require the final model to be optimal. Empirical results on a synthetic dataset and the FEMNIST dataset demonstrate that our estimation method can approximate Fed-Influence with small bias. Further, we show an application of Fed-Influence in model debugging.
Yihao Xue, Chaoyue Niu, Zhenzhe Zheng 0001, Shaojie Tang 0001, Chengfei Lyu, Fan Wu 0006, Guihai Chen
AAAI1