Xuwei Xu

dblp:332/1103 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-3434-7451ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 80% Deep learning architectures and training · 9% 3D vision · 6%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
2.432026
Understanding the Effects of Projectors in Knowledge Distillation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Medium-Difficulty Samples Constitute Smoothed Decision Boundary for Knowledge Distillation on Pruned Datasets · ICLR 2025
Improved Feature Distillation via Projector Ensemble · NeurIPS 2022
Machine learning › Efficient and distributed learning
model compression
1.922026
Understanding the Effects of Projectors in Knowledge Distillation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers · ICML 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
feature distillation
1.622026
Understanding the Effects of Projectors in Knowledge Distillation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Improved Feature Distillation via Projector Ensemble · NeurIPS 2022
Machine learning › Efficient and distributed learning › data selection
data pruning
0.912025
Medium-Difficulty Samples Constitute Smoothed Decision Boundary for Knowledge Distillation on Pruned Datasets · ICLR 2025
Machine learning › Efficient and distributed learning › model compression
structural re-parameterization
0.912025
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers · ICML 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.912025
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers · ICML 2025
Recommender systems › reinforcement-learning-based recommendation
offline reinforcement learning for recommendation
0.912025
DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic Reward · SIGIR 2025
Recommender systems
sequential recommendation
0.912025
DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic Reward · SIGIR 2025
Computer vision › 3D vision
feature matching
0.612022
Improved Feature Distillation via Projector Ensemble · NeurIPS 2022
Machine learning › Efficient and distributed learning
inference efficiency
0.312025
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers · ICML 2025
Machine learning › Reinforcement learning › offline reinforcement learning
model-based offline reinforcement learning
0.312025
DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic Reward · SIGIR 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.312025
DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic Reward · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

world model · 1.7uncertainty penalty · 1.7dual-agent framework · 1.7projector ensemble · 1.6model calibration · 1.0centered kernel alignment · 1.0structural re-parameterization · 0.9knowledge distillation · 0.9channel idle mechanism · 0.9multi-task learning · 0.6
YearPublicationVenuePosition
2026 Understanding the Effects of Projectors in Knowledge Distillation
abstract
Conventionally, during the knowledge distillation process (e.g., feature distillation), an additional projector is often required to perform feature transformation due to the dimension mismatch between the teacher and the student networks. Interestingly, we discovered that even if the student and the teacher have the same feature dimensions, adding a projector still helps to improve the distillation performance. In addition, projectors even improve logit distillation if we add them to the architecture too. Inspired by these surprising findings and the general lack of understanding of the projectors in the knowledge distillation process from existing literature, this paper investigates the implicit role that projectors play, but so far been overlooked. Our empirical study shows that the student with a projector 1) obtains a better trade-off between the training accuracy and the testing accuracy compared to the student without a projector when it has the same feature dimensions as the teacher, 2) better preserves its similarity to the teacher beyond shallow and numeric resemblance, from the view of Centered Kernel Alignment (CKA) (Kornblith et al., 2019), and 3) avoids being over-confident (Guo et al., 2017) as the teacher does at the testing phase. Motivated by the positive effects of projectors, we propose a projector ensemble-based feature distillation method to further improve distillation performance. Despite the simplicity of the proposed strategy, empirical results from the evaluation of classification tasks on benchmark datasets demonstrate the superior classification performance of our method on a broad range of teacher-student pairs and verify, from the aspects of CKA and model calibration that the student's features are of improved quality with the projector ensemble design.
Yudong Chen 0002, Sen Wang 0001, Jiajun Liu 0004, Xuwei Xu, Frank de Hoog, Branislav Kusy, Zi Huang
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Medium-Difficulty Samples Constitute Smoothed Decision Boundary for Knowledge Distillation on Pruned Datasets
abstract
This paper tackles a new problem of dataset pruning for Knowledge Distillation (KD), from a fresh perspective of Decision Boundary (DB) preservation and drifts. Existing dataset pruning methods generally assume that the post-pruning DB formed by the selected samples can be well-captured by future networks that use those samples for training. Therefore, they tend to preserve hard samples since hard samples are closer to the DB and better characterize the nuances in the distribution of the entire dataset. However, in KD, the limited learning capacity from the student network leads to imperfect preservation of the teacher's feature distribution, resulting in the drift of DB in the student space. Specifically, hard samples worsen such drifts as they are difficult for the student to learn, creating a situation where the student's DB can drift deeper into other classes and make incorrect classifications. Motivated by these findings, our method selects medium-difficulty samples for KD-based dataset pruning. We show that these samples constitute a smoothed version of the teacher's DB and are easier for the student to learn, obtaining a general feature distribution preservation for a class of samples and reasonable DB between different classes for the student. In addition, to reduce the distributional shift due to dataset pruning, we leverage the class-wise distributional information of the teacher's outputs to reshape the logits of the preserved samples. Experiments show that the proposed static pruning method can even perform better than the state-of-the-art dynamic pruning method which needs access to the entire dataset. In addition, our method halves the training times of KD and improves the student's accuracy by 0.4% on ImageNet with a 50% keep ratio. When the ratio further increases to 70%, our method achieves higher accuracy over the vanilla KD while reducing the training times by 30%. Code is available at https://github.com/chenyd7/MDSLR.
Yudong Chen 0002, Xuwei Xu, Frank de Hoog, Jiajun Liu 0004, Sen Wang 0001
ICLR2
2025 RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
abstract
We reveal that feedforward network (FFN) layers, rather than attention layers, are the primary contributors to Vision Transformer (ViT) inference latency, with their impact signifying as model size increases. This finding highlights a critical opportunity for optimizing the efficiency of large-scale ViTs by focusing on FFN layers. In this work, we propose a novel channel idle mechanism that facilitates post-training structural reparameterization for efficient FFN layers during testing. Specifically, a set of feature channels remains idle and bypasses the nonlinear activation function in each FFN layer, thereby forming a linear pathway that enables structural reparameterization during inference. This mechanism results in a family of **RePa**rameterizable **Vi**sion **T**ransformers (RePaViTs), which achieve remarkable latency reductions with acceptable sacrifices (sometimes gains) in accuracy across various ViTs. The benefits of our method scale consistently with model sizes, demonstrating greater speed improvements and progressively narrowing accuracy gaps or even higher accuracies on larger models. In particular, RePa-ViT-Large and RePa-ViT-Huge enjoy **66.8%** and **68.7%** speed-ups with **+1.7%** and **+1.1%** higher top-1 accuracies under the same training strategy, respectively. RePaViT is the first to employ structural reparameterization on FFN layers to expedite ViTs to our best knowledge, and we believe that it represents an auspicious direction for efficient ViTs. Source code is available at https://github.com/Ackesnal/RePaViT.
Xuwei Xu, Yang Li 0184, Yudong Chen 0002, Jiajun Liu 0004, Sen Wang 0001
ICML1
2025 DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic Reward
abstract
Model-based offline reinforcement learning (RL) has emerged as a promising approach for recommender systems, enabling effective policy learning by interacting with frozen world models. However, the reward functions in these world models, trained on sparse offline logs, often suffer from inaccuracies. Specifically, existing methods face two major limitations in addressing this challenge: (1) deterministic use of reward functions as static look-up tables, which propagates inaccuracies during policy learning, and (2) static uncertainty designs that fail to effectively capture decision risks and mitigate the impact of these inaccuracies. In this work, a dual-agent framework, DARLR, is proposed to dynamically update world models to enhance recommendation policies. To achieve this, a selector is introduced to identify reference users by balancing similarity and diversity so that the recommender can aggregate information from these users and iteratively refine reward estimations for dynamic reward shaping. Further, the statistical features of the selected users guide the dynamic adaptation of an uncertainty penalty to better align with evolving recommendation requirements. Extensive experiments on four benchmark datasets demonstrate the superior performance of DARLR, validating its effectiveness. The code is available at this address.
Yi Zhang 0105, Ruihong Qiu, Xuwei Xu, Jiajun Liu 0004, Sen Wang 0001
SIGIR3
2024 Towards Cost-Efficient Federated Multi-agent RL with Learnable Aggregation
Yi Zhang 0105, Sen Wang 0001, Zhi Chen 0010, Xuwei Xu, Stanislav Funiak, Jiajun Liu 0004
PAKDD (2)4
2024 GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
abstract
Vision Transformers (ViTs) have revolutionized the field of computer vision, yet their deployments on resource-constrained devices remain challenging due to high computational demands. To expedite pre-trained ViTs, token pruning and token merging approaches have been developed, which aim at reducing the number of tokens involved in the computation. However, these methods still have some limitations, such as image information loss from pruned tokens and inefficiency in the token-matching process. In this paper, we introduce a novel Graph-based Token Propagation (GTP) method to resolve the challenge of balancing model efficiency and information preservation for efficient ViTs. Inspired by graph summarization algorithms, GTP meticulously propagates less significant tokens’ information to spatially and semantically connected tokens that are of greater importance. Consequently, the remaining few tokens serve as a summarization of the entire token graph, allowing the method to reduce computational complexity while preserving essential information of eliminated tokens. Combined with an innovative token selection strategy, GTP can efficiently identify image tokens to be propagated. Extensive experiments have validated GTP’s effectiveness, demonstrating both efficiency and performance improvements. Specifically, GTP decreases the computational complexity of both DeiT-S and DeiT-B by up to 26% with only a minimal 0.3% accuracy drop on ImageNet-1K without finetuning, and remarkably surpasses the state-of-the-art token merging method on various backbones at an even faster inference speed. The source code is available at https://github.com/Ackesnal/GTP-ViT.
Xuwei Xu, Sen Wang 0001, Yudong Chen 0002, Yanping Zheng, Zhewei Wei, Jiajun Liu 0004
WACV1
2022 Improved Feature Distillation via Projector Ensemble
abstract
In knowledge distillation, previous feature distillation methods mainly focus on the design of loss functions and the selection of the distilled layers, while the effect of the feature projector between the student and the teacher remains under-explored. In this paper, we first discuss a plausible mechanism of the projector with empirical evidence and then propose a new feature distillation method based on a projector ensemble for further performance improvement. We observe that the student network benefits from a projector even if the feature dimensions of the student and the teacher are the same. Training a student backbone without a projector can be considered as a multi-task learning process, namely achieving discriminative feature extraction for classification and feature matching between the student and the teacher for distillation at the same time. We hypothesize and empirically verify that without a projector, the student network tends to overfit the teacher's feature distributions despite having different architecture and weights initialization. This leads to degradation on the quality of the student's deep features that are eventually used in classification. Adding a projector, on the other hand, disentangles the two learning tasks and helps the student network to focus better on the main feature extraction task while still being able to utilize teacher features as a guidance through the projector. Motivated by the positive effect of the projector in feature distillation, we propose an ensemble of projectors to further improve the quality of student features. Experimental results on different datasets with a series of teacher-student pairs illustrate the effectiveness of the proposed method. Code is available at https://github.com/chenyd7/PEFD.
Yudong Chen 0002, Sen Wang 0001, Jiajun Liu 0004, Xuwei Xu, Frank de Hoog, Zi Huang
NeurIPS4