Yu Pan 0005

dblp:76/1503-5 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-7515-8492ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing progressive ensemble learning via normalized extra-Gradient initialization
Zheshun Wu, Yu Pan 0005, Dun Zeng, Qifan Wang 0001, Zenglin Xu, Jie Liu 0001
Neural Networks2
2025 On the Power of Adaptive Weighted Aggregation in Heterogeneous Federated Learning and Beyond
abstract
Federated averaging (FedAvg) is the most fundamental algorithm in Federated learning (FL). Previous theoretical results assert that FedAvg convergence and generalization degenerate under heterogeneous clients. However, recent empirical results show that FedAvg can perform well in many real-world heterogeneous tasks. These results reveal an inconsistency between FL theory and practice that is not fully explained. In this paper, we show that common heterogeneity measures contribute to this inconsistency based on rigorous convergence analysis. Furthermore, we introduce a new measure \textit{client consensus dynamics} and prove that \textit{FedAvg can effectively handle client heterogeneity when an appropriate aggregation strategy is used}. Building on this theoretical insight, we present a simple and effective FedAvg variant termed FedAWARE. Extensive experiments on three datasets and two modern neural network architectures demonstrate that FedAWARE ensures faster convergence and better generalization in heterogeneous client settings. Moreover, our results show that FedAWARE can significantly enhance the generalization performance of advanced FL algorithms when used as a plug-in module.
Dun Zeng, Zenglin Xu, Yu Pan 0005, Qifan Wang 0001, Xiaoying Tang 0002
AISTATS4
2025 IDInit: A Universal and Stable Initialization Method for Neural Network Training
abstract
Deep neural networks have achieved remarkable accomplishments in practice. The success of these networks hinges on effective initialization methods, which are vital for ensuring stable and rapid convergence during training. Recently, initialization methods that maintain identity transition within layers have shown good efficiency in network training. These techniques (e.g., Fixup) set specific weights to zero to achieve identity control. However, settings of remaining weight (e.g., Fixup uses random values to initialize non-zero weights) will affect the inductive bias that is achieved only by a zero weight, which may be harmful to training. Addressing this concern, we introduce fully identical initialization (IDInit), a novel method that preserves identity in both the main and sub-stem layers of residual networks. IDInit employs a padded identity-like matrix to overcome rank constraints in non-square weight matrices. Furthermore, we show the convergence problem of an identity matrix can be solved by stochastic gradient descent. Additionally, we enhance the universality of IDInit by processing higher-order weights and addressing dead neuron problems. IDInit is a straightforward yet effective initialization method, with improved convergence, stability, and performance across various settings, including large-scale datasets and deep models.
Yu Pan 0005, Chaozheng Wang, Zekai Wu, Qifan Wang 0001, Min Zhang 0014, Zenglin Xu
ICLR1
2024 Preparing Lessons for Progressive Training on Language Models
abstract
The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prepares lessons for expanding operations by learning high-layer functionality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs.
Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Jiaxin Shi, Zenglin Xu, Ming Zhang 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
AAAI1
2024 Tensorized Hypergraph Neural Networks
abstract
Hypergraph neural networks (HGNN) have recently become attractive and received significant attention due to their excellent performance in various domains. However, most existing HGNNs rely on first-order approximations of hypergraph connectivity patterns, which ignores important high-order information. To address this issue, we propose a novel adjacency-tensor-based Tensorized Hypergraph Neural Network (THNN). THNN is a faithful hypergraph modeling framework through high-order outer product feature message passing and is a natural tensor extension of the adjacency-matrix-based graph neural networks. The proposed THNN is equivalent to a high-order polynomial regression scheme, which enables THNN with the ability to efficiently extract high-order information from uniform hypergraphs. Moreover, in consideration of the exponential complexity of directly processing high-order outer product features, we propose using a partially symmetric CP decomposition approach to reduce model complexity to a linear degree. Additionally, we propose two simple yet effective extensions of our method for non-uniform hypergraphs commonly found in real-world applications. Results from experiments on two widely used hypergraph datasets for 3-D visual object classification show the model's promising performance.
Maolin Wang 0001, Yaoming Zhen, Yu Pan 0005, Yao Zhao 0011, Chenyi Zhuang, Zenglin Xu, Ruocheng Guo, Xiangyu Zhao 0001
SDM3
2023 Reusing Pretrained Models by Multi-linear Operators for Efficient Training
abstract
Training large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despite the successes of these previous studies, they grew pretrained models by mapping partial weights only, ignoring potential correlations across the entire model. As we show in this paper, there are inter- and intra-interactions among the weights of both the pretrained and the target models. As a result, the partial mapping may not capture the complete information and lead to inadequate growth. In this paper, we propose a method that linearly correlates each weight of the target model to all the weights of the pretrained model to further enhance acceleration ability. We utilize multi-linear operators to reduce computational and spacial complexity, enabling acceptable resource requirements. Experiments demonstrate that our method can save 76\% computational costs on DeiT-base transferred from DeiT-small, which outperforms bert2BERT by +12\% and LiGO by +21\%, respectively.
Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Zenglin Xu, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
NeurIPS1
2023 RegNet: Self-Regulated Network for Image Classification
abstract
The ResNet and its variants have achieved remarkable successes in various computer vision tasks. Despite its success in making gradient flow through building blocks, the information communication of intermediate layers of blocks is ignored. To address this issue, in this brief, we propose to introduce a regulator module as a memory mechanism to extract complementary features of the intermediate layers, which are further fed to the ResNet. In particular, the regulator module is composed of convolutional recurrent neural networks (RNNs) [e.g., convolutional long short-term memories (LSTMs) or convolutional gated recurrent units (GRUs)], which are shown to be good at extracting spatio-temporal information. We named the new regulated network as regulated residual network (RegNet). The regulator module can be easily implemented and appended to any ResNet architecture. Experimental results on three image classification datasets have demonstrated the promising performance of the proposed architecture compared with the standard ResNet, squeeze-and-excitation ResNet, and other state-of-the-art architectures.
Yu Pan 0005, Xinglin Pan, Steven C. H. Hoi, Zhang Yi 0001, Zenglin Xu
IEEE Trans. Neural Networks Learn. Syst.2
2022 Semantically Proportional Patchmix for Few-Shot Learning
abstract
Few-shot learning aims to classify unseen classes with only a limited number of labeled data. Recent works have demonstrated that training models with a simple transfer learning strategy can achieve competitive results in few-shot classification. Although excelling at distinguishing training data, these models are not well generalized to unseen data, probably due to insufficient feature representations on evaluation. To tackle this issue, we propose Semantically Proportional Patchmix (SePPMix), in which patches are cut and pasted among training images and the ground truth labels are mixed proportionally to the semantic information of the patches. In this way, we can improve the generalization ability of the model by regional dropout effect without introducing severe label noise. To learn more robust representations of data, we further take rotate transformation on the mixed images and predict rotations as a rule-based regularizer. Extensive experiments on prevalent few-shot benchmarks have shown the effectiveness of our proposed method.
Jingquan Wang, Yu Pan 0005, Zenglin Xu
ICASSP3
2022 A Unified Weight Initialization Paradigm for Tensorial Convolutional Neural Networks
abstract
Tensorial Convolutional Neural Networks (TCNNs) have attracted much research attention for their power in reducing model parameters or enhancing the generalization ability. However, exploration of TCNNs is hindered even from weight initialization methods. To be specific, general initialization methods, such as Xavier or Kaiming initialization, usually fail to generate appropriate weights for TCNNs. Meanwhile, although there are ad-hoc approaches for specific architectures (e.g., Tensor Ring Nets), they are not applicable to TCNNs with other tensor decomposition methods (e.g., CP or Tucker decomposition). To address this problem, we propose a universal weight initialization paradigm, which generalizes Xavier and Kaiming methods and can be widely applicable to arbitrary TCNNs. Specifically, we first present the Reproducing Transformation to convert the backward process in TCNNs to an equivalent convolution process. Then, based on the convolution operators in the forward and backward processes, we build a unified paradigm to control the variance of features and gradients in TCNNs. Thus, we can derive fan-in and fan-out initialization for various TCNNs. We demonstrate that our paradigm can stabilize the training of TCNNs, leading to faster convergence and better results.
Yu Pan 0005, Zeyong Su, Ao Liu 0008, Jingquan Wang, Nannan Li 0001, Zenglin Xu
ICML1
2022 TedNet: A Pytorch toolkit for tensor decomposition networks
Yu Pan 0005, Maolin Wang 0001, Zenglin Xu
Neurocomputing1
2022 AFINet: Attentive Feature Integration Networks for image classification
Xinglin Pan, Yu Pan 0005, Liangjian Wen, Wenxiang Lin, Hongguang Fu, Zenglin Xu
Neural Networks3
2020 Concatenated Tensor Networks for Deep Multi-Task Learning
Maolin Wang 0001, Zeyong Su, Xu Luo 0003, Yu Pan 0005, Shenggen Zheng, Zenglin Xu
ICONIP (5)4
2019 Compressing Recurrent Neural Networks with Tensor Ring for Action Recognition
abstract
Recurrent Neural Networks (RNNs) and their variants, such as Long-Short Term Memory (LSTM) networks, and Gated Recurrent Unit (GRU) networks, have achieved promising performance in sequential data modeling. The hidden layers in RNNs can be regarded as the memory units, which are helpful in storing information in sequential contexts. However, when dealing with high dimensional input data, such as video and text, the input-to-hidden linear transformation in RNNs brings high memory usage and huge computational cost. This makes the training of RNNs very difficult. To address this challenge, we propose a novel compact LSTM model, named as TR-LSTM, by utilizing the low-rank tensor ring decomposition (TRD) to reformulate the input-to-hidden transformation. Compared with other tensor decomposition methods, TR-LSTM is more stable. In addition, TR-LSTM can complete an end-to-end training and also provide a fundamental building block for RNNs in handling large input data. Experiments on real-world action recognition datasets have demonstrated the promising performance of the proposed TR-LSTM compared with the tensor-train LSTM and other state-of-the-art competitors.
Yu Pan 0005, Maolin Wang 0001, Jinmian Ye, Fei Wang 0001, Zenglin Xu
AAAI1
2019 Tensor Ring Restricted Boltzmann Machines
abstract
Restricted Boltzmann Machines are important and useful generative models which learn a probability distribution from a set of vector inputs. Despite their success in a number of applications, standard RBMs designed for vectorized inputs are incapable of dealing with high-order data, since vectorization of high-order data may cause both modes collapsing and explosive parameter growth. To address this issue, we formulate a new tensor-input RBM model, which employs the tensor-ring (TR) decomposition structure to naturally represent the high-order relationship between the visual layer and the hidden layer. For convenience, we name the proposed model as TR-RBM. In particular, the tensor ring decomposition enjoys many good properties, such as the rank stableness, leading to better generalization performance compared with other low-rank decomposition methods. Moreover, TR-RBM can also reduce the complexity of RBM by reshaping of both visible and hidden layers into the tensor forms, leading a significant drop of parameter size. Experimental results in comparison with the classical RBMs and the Matrix-Product-Operator RBM have shown the promising performance of the proposed method in the tasks of feature extraction and denoising.
Maolin Wang 0001, Chenbin Zhang, Yu Pan 0005, Zenglin Xu
IJCNN3