Wei Chen 0124

dblp:181/2832-124 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0001-6722-4322ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Efficient and distributed learning · 33% Deep learning architectures and training · 33% Generative modeling · 15%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
2.632025
Large Convolutional Model Tuning via Filter Subspace · ICLR 2025
Sparse Fine-Tuning of Transformers for Generative Tasks · ICCV 2025
Coeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large Models · CVPR 2025
Machine learning › Deep learning architectures and training
convolutional neural network
1.522025
Large Convolutional Model Tuning via Filter Subspace · ICLR 2025
Inner Product-based Neural Network Similarity · NeurIPS 2023
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
In-Context Compositional Learning vis Sparse Coding Transformer · NeurIPS 2025
Natural language and speech › Language models and text generation
compositional generalization
0.912025
In-Context Compositional Learning vis Sparse Coding Transformer · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › personalized image generation
concept customization
0.912025
Sparse Fine-Tuning of Transformers for Generative Tasks · ICCV 2025
Machine learning › Deep learning architectures and training › convolutional neural network
filter decomposition
0.912025
Large Convolutional Model Tuning via Filter Subspace · ICLR 2025
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
sparse fine-tuning
0.912025
Sparse Fine-Tuning of Transformers for Generative Tasks · ICCV 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Sparse Fine-Tuning of Transformers for Generative Tasks · ICCV 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
In-Context Compositional Learning vis Sparse Coding Transformer · NeurIPS 2025
Machine learning › Deep learning architectures and training › attention mechanism
transformer attention
0.912025
Coeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large Models · CVPR 2025
Visual content generation and editing
image editing
0.912025
Sparse Fine-Tuning of Transformers for Generative Tasks · ICCV 2025
Visual content generation and editing › image editing
text-guided image editing
0.912025
Sparse Fine-Tuning of Transformers for Generative Tasks · ICCV 2025
Machine learning › Efficient and distributed learning
model compression
0.822025
Continual Learning with Filter Atom Swapping · ICLR 2022
Large Convolutional Model Tuning via Filter Subspace · ICLR 2025
Machine learning › Learning paradigms
continual learning
0.822023
Continual Learning with Filter Atom Swapping · ICLR 2022
Inner Product-based Neural Network Similarity · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation analysis
representation similarity
0.712023
Inner Product-based Neural Network Similarity · NeurIPS 2023
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning
0.612022
Continual Learning with Filter Atom Swapping · ICLR 2022
Machine learning › Generative modeling
implicit generative model
0.512021
Run-Sort-ReRun: Escaping Batch Size Limitations in Sliced Wasserstein Generative Models · ICML 2021
Machine learning › Optimization for machine learning › optimal transport
sliced wasserstein distance
0.512021
Run-Sort-ReRun: Escaping Batch Size Limitations in Sliced Wasserstein Generative Models · ICML 2021
Edge and fog computing
edge intelligence
0.412020
INVITED: New Directions in Distributed Deep Learning: Bringing the Network at Forefront of IoT Design · DAC 2020
Distributed systems › distributed machine learning
federated learning
0.412020
INVITED: New Directions in Distributed Deep Learning: Bringing the Network at Forefront of IoT Design · DAC 2020
Machine learning › Efficient and distributed learning
federated learning
0.212023
Inner Product-based Neural Network Similarity · NeurIPS 2023
Machine learning › Optimization for machine learning
optimal transport
0.112021
Run-Sort-ReRun: Escaping Batch Size Limitations in Sliced Wasserstein Generative Models · ICML 2021

Methods — techniques the papers use, named apart from their topics

sparse coding · 2.6transformer fine-tuning · 1.7feature dictionary · 1.7filter atom decomposition · 1.5recursive decomposition · 0.9graph convolution · 0.9dropout regularization · 0.9dictionary learning · 0.9federated learning · 0.9deep learning · 0.9cosine distance · 0.7sliced wasserstein distance · 0.5
YearPublicationVenuePosition
2025 Coeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large Models
abstract
Transformer-based large pre-trained models have shown remarkable generalization ability, and various parameter-efficient fine-tuning (PEFT) methods have been proposed to customize these models on downstream tasks with minimal computational and memory budgets. Previous PEFT methods are primarily designed from a tensor-decomposition perspective that tries to effectively tune the linear transformation by finding the smallest subset of parameters to train. Our study adopts an orthogonal view by representing the attention operation as a graph convolution and formulating the multi-head attention maps as a convolutional filter subspace, with each attention map as a subspace element. In this paper, we propose to tune the large pre-trained transformers by learning a small set of combination coefficients that construct a more expressive filter subspace from the original multi-head attention maps. We show analytically and experimentally that the tuned filter subspace can effectively expand the feature space of the multi-head attention and further enhance the capacity of transformers. We further stabilize the fine-tuning with a residual parameterization of the tunable subspace coefficients, and enhance the generalization with a regularization design by directly applying dropout on the tunable coefficient during training. The tunable coefficients take a tiny number of parameters and can be combined with previous PEFT methods in a plug-and-play manner. Extensive experiments show that our approach achieves superior performances than PEFT baselines with neglectable additional parameters.1
Zichen Miao, Wei Chen 0124, Qiang Qiu 0001
CVPR2
2025 Sparse Fine-Tuning of Transformers for Generative Tasks
abstract
Large pre-trained transformers have revolutionized artificial intelligence across various domains, and fine-tuning remains the dominant approach for adapting these models to downstream tasks due to the cost of training from scratch. However, in existing fine-tuning methods, the updated representations are formed as a dense combination of modified parameters, making it challenging to interpret their contributions and understand how the model adapts to new tasks. In this work, we introduce a fine-tuning framework inspired by sparse coding, where fine-tuned features are represented as a sparse combination of basic elements, i.e., feature dictionary atoms. The feature dictionary atoms function as fundamental building blocks of the representation, and tuning atoms allows for seamless adaptation to downstream tasks. Sparse coefficients then serve as indicators of atom importance, identifying the contribution of each atom to the updated representation. Leveraging the atom selection capability of sparse coefficients, we first demonstrate that our method enhances image editing performance by improving text alignment through the removal of unimportant feature dictionary atoms. Additionally, we validate the effectiveness of our approach in the text-to-image concept customization task, where our method efficiently constructs the target concept using a sparse combination of feature dictionary atoms, outperforming various baseline fine-tuning methods.
Wei Chen 0124, Jingxi Yu, Zichen Miao, Qiang Qiu 0001
ICCV1
2025 Large Convolutional Model Tuning via Filter Subspace
abstract
Efficient fine-tuning methods are critical to address the high computational and parameter complexity while adapting large pre-trained models to downstream tasks. Our study is inspired by prior research that represents each convolution filter as a linear combination of a small set of filter subspace elements, referred to as filter atoms. In this paper, we propose to fine-tune pre-trained models by adjusting only filter atoms, which are responsible for spatial-only convolution, while preserving spatially-invariant channel combination knowledge in atom coefficients. In this way, we bring a new filter subspace view for model tuning. Furthermore, each filter atom can be recursively decomposed as a combination of another set of atoms, which naturally expands the number of tunable parameters in the filter subspace. By only adapting filter atoms constructed by a small number of parameters, while maintaining the rest of model parameters constant, the proposed approach is highly parameter-efficient. It effectively preserves the capabilities of pre-trained models and prevents overfitting to downstream tasks. Extensive experiments show that such a simple scheme surpasses previous tuning baselines for both discriminate and generative tasks.
Wei Chen 0124, Zichen Miao, Qiang Qiu 0001
ICLR1
2025 In-Context Compositional Learning vis Sparse Coding Transformer
abstract
Recent advances in AI, driven by Transformer architectures, have achieved remarkable success in language, vision, and multimodal reasoning, and there is growing demand for them to address in-context compositional learning tasks. In these tasks, models solve the target problems by inferring compositional rules from context examples, which are composed of basic components structured by underlying rules. However, some of these tasks remain challenging for Transformers, which are not inherently designed to handle compositional tasks and offer limited structural inductive bias. Inspired by sparse coding, we propose a reformulation of the attention to enhance its capability for compositional tasks. In sparse coding, data are represented as sparse combinations of basic elements, with the resulting coefficients capturing the underlying compositional structure of the input. Specifically, we reinterpret the standard attention block as projecting inputs into outputs through projections onto two sets of learned dictionary atoms: an *encoding dictionary* and a *decoding dictionary*. The encoding dictionary decomposes the input into a set of coefficients, which represent the compositional structure of the input. To enhance structured representations, we impose sparsity on these coefficients. The sparse coefficients are then used to linearly combine the decoding dictionary atoms to generate the output. Furthermore, to assist compositional generalization tasks, we propose estimating the coefficients of the target problem as a linear combination of the coefficients obtained from the context examples. We demonstrate the effectiveness of our approach on the S-RAVEN and RAVEN datasets. For certain compositional generalization tasks, our method maintains performance even when standard Transformers fail, owing to its ability to learn and apply compositional rules.
Wei Chen 0124, Jingxi Yu, Zichen Miao, Qiang Qiu 0001
NeurIPS1
2023 FLAIR: Defense against Model Poisoning Attack in Federated Learning
abstract
Federated learning—multi-party, distributed learning in a decentralized environment—is vulnerable to model poisoning attacks, more so than centralized learning. This is because malicious clients can collude and send in carefully tailored model updates to make the global model inaccurate. This motivated the development of Byzantine-resilient federated learning algorithms, such as Krum, Bulyan, FABA, and FoolsGold. However, a recently developed untargeted model poisoning attack showed that all prior defenses can be bypassed. The attack uses the intuition that simply by changing the sign of the gradient updates that the optimizer is computing, for a set of malicious clients, a model can be diverted from the optima to increase the test error rate. In this work, we develop FLAIR—a defense against this directed deviation attack (DDA), a state-of-the-art model poisoning attack. FLAIR is based on our intuition that in federated learning, certain patterns of gradient flips are indicative of an attack. This intuition is remarkably stable across different learning algorithms, models, and datasets. FLAIR assigns reputation scores to the participating clients based on their behavior during the training phase and then takes a weighted contribution of the clients. We show that where the existing defense baselines of FABA [IJCAI ’19], FoolsGold [Usenix ’20], and FLTrust [NDSS ’21] fail when 20-30% of the clients are malicious, FLAIR provides byzantine-robustness upto a malicious client percentage of 45%. We also show that FLAIR provides robustness against even a white-box version of DDA.
Wei Chen 0124, Joshua Zhao 0001, Qiang Qiu 0001, Saurabh Bagchi, Somali Chaterji
AsiaCCS2
2023 Inner Product-based Neural Network Similarity
abstract
Analyzing representational similarity among neural networks (NNs) is essential for interpreting or transferring deep models. In application scenarios where numerous NN models are learned, it becomes crucial to assess model similarities in computationally efficient ways. In this paper, we propose a new paradigm for reducing NN representational similarity to filter subspace distance. Specifically, when convolutional filters are decomposed as a linear combination of a set of filter subspace elements, denoted as filter atoms, and have those decomposed atom coefficients shared across networks, NN representational similarity can be significantly simplified as calculating the cosine distance among respective filter atoms, to achieve millions of times computation reduction over popular probing-based methods. We provide both theoretical and empirical evidence that such simplified filter subspace-based similarity preserves a strong linear correlation with other popular probing-based metrics, while being significantly more efficient to obtain and robust to probing data. We further validate the effectiveness of the proposed method in various application scenarios where numerous models exist, such as federated and continual learning as well as analyzing training dynamics. We hope our findings can help further explorations of real-time large-scale representational similarity analysis in neural networks.
Wei Chen 0124, Zichen Miao, Qiang Qiu 0001
NeurIPS1
2022 Continual Learning with Filter Atom Swapping
Zichen Miao, Ze Wang 0008, Wei Chen 0124, Qiang Qiu 0001
ICLR3
2022 FT-DeepNets: Fault-Tolerant Convolutional Neural Networks with Kernel-based Duplication
abstract
Deep neural network (deepnet) applications play a crucial role in safety-critical systems such as autonomous vehicles (AVs). An AV must drive safely towards its destination, avoiding obstacles, and respond quickly when the vehicle must stop. Any transient errors in software calculations or hardware memory in these deepnet applications can potentially lead to dramatically incorrect results. Therefore, assessing and mitigating any transient errors and providing robust results are important for safety-critical systems. Previous research on this subject focused on detecting errors and then recovering from the errors by re-running the network. Other approaches were based on the extent of full network duplication such as the ensemble learning-based approach to boost system fault-tolerance by leveraging each model’s advantages. However, it is hard to detect errors in a deep neural network, and the computational overhead of full redundancy can be substantial.We first study the impact of the error types and locations in deepnets. We next focus on selecting which part should be duplicated using multiple ranking methods to measure the order of importance among neurons. We find that the duplication overhead for computation and memory is a trade-off between algorithmic performance and robustness. To achieve higher robustness with less system overhead, we present two error protection mechanisms that only duplicate parts of the network from critical neurons. Finally, we substantiate the practical feasibility of our approach and evaluate the improvement in the accuracy of a deepnet in the presence of errors. We demonstrate these results using a case study with real-world applications on an Nvidia GeForce RTX 2070Ti GPU and an Nvidia Xavier embedded platform used by automotive OEMs.
Iljoo Baek, Wei Chen 0124, Soheil Samii, Ragunathan Rajkumar
WACV2
2021 Run-Sort-ReRun: Escaping Batch Size Limitations in Sliced Wasserstein Generative Models
abstract
When training an implicit generative model, ideally one would like the generator to reproduce all the different modes and subtleties of the target distribution. Naturally, when comparing two empirical distributions, the larger the sample population, the more these statistical nuances can be captured. However, existing objective functions are computationally constrained in the amount of samples they can consider by the memory required to process a batch of samples. In this paper, we build upon recent progress in sliced Wasserstein distances, a family of differentiable metrics for distribution discrepancy based on the Optimal Transport paradigm. We introduce a procedure to train these distances with virtually any batch size, allowing the discrepancy measure to capture richer statistics and better approximating the distance between the underlying continuous distributions. As an example, we demonstrate the matching of the distribution of Inception features with batches of tens of thousands of samples, achieving FID scores that outperform state-of-the-art implicit generative models.
José Lezama, Wei Chen 0124, Qiang Qiu 0001
ICML2
2020 INVITED: New Directions in Distributed Deep Learning: Bringing the Network at Forefront of IoT Design
abstract
In this paper, we first highlight three major challenges to large-scale adoption of deep learning at the edge: (i) Hardware-constrained IoT devices, (ii) Data security and privacy in the IoT era, and (iii) Lack of network-aware deep learning algorithms for distributed inference across multiple IoT devices. We then provide a unified view targeting three research directions that naturally emerge from the above challenges: (1) Federated learning for training deep networks, (2) Data-independent deployment of learning algorithms, and (3) Communication-aware distributed inference. We believe that the above research directions need a network-centric approach to enable the edge intelligence and, therefore, fully exploit the true potential of IoT.
Kartikeya Bhardwaj, Wei Chen 0124, Radu Marculescu
DAC2
2020 FedMAX: Mitigating Activation Divergence for Accurate and Communication-Efficient Federated Learning
Wei Chen 0124, Kartikeya Bhardwaj, Radu Marculescu
ECML/PKDD (2)1