Guanchu Wang

dblp:213/0985 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0003-3258-762XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Catastrophic Forgetting in Kolmogorov-Arnold Networks
abstract
Catastrophic forgetting is a longstanding challenge in continual learning, where models lose knowledge from earlier tasks when learning new ones. While various mitigation strategies have been proposed for Multi-Layer Perceptrons (MLPs), recent architectural advances like Kolmogorov-Arnold Networks (KANs) have been suggested to offer intrinsic resistance to forgetting by leveraging localized spline-based activations. However, the practical behavior of KANs under continual learning remains unclear, and their limitations are not well understood. To address this, we present a comprehensive study of catastrophic forgetting in KANs and develop a theoretical framework that links forgetting to activation support overlap and intrinsic data dimension. We validate these analyses through systematic experiments on synthetic and vision tasks, measuring forgetting dynamics under varying model configurations and data complexity. Further, we introduce KAN-LoRA, a novel adapter design for parameter-efficient continual fine-tuning of language models, and evaluate its effectiveness in knowledge editing tasks. Our findings reveal that while KANs exhibit promising retention in low-dimensional algorithmic settings, they remain vulnerable to forgetting in high-dimensional domains such as image classification and language modeling. These results advance the understanding of KANs’ strengths and limitations, offering practical insights for continual learning system design.
Mohammad Marufur Rahman, Guanchu Wang, Kaixiong Zhou, Minghan Chen 0001, Fan Yang 0023
AAAI2
2025 MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation
abstract
Chia-Yuan Chang, Zhimeng Jiang, Vineeth Rakesh, Menghai Pan, Chin-Chia Michael Yeh, Guanchu Wang, Mingzhi Hu, Zhichao Xu, Yan Zheng, Mahashweta Das, Na Zou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chia-Yuan Chang 0002, Zhimeng Jiang, Vineeth Rakesh, Menghai Pan, Chin-Chia Michael Yeh, Guanchu Wang, Mingzhi Hu, Yan Zheng 0001, Mahashweta Das, Na Zou 0001
ACL (1)6
2025 Quantized Can Still Be Calibrated: A Unified Framework to Calibration in Quantized Large Language Models
abstract
Although weight quantization helps large language models (LLMs) in resource-constrained environments, its influence on the uncertainty calibration remains unexplored.To bridge this gap, we present a comprehensive investigation of uncertainty calibration for quantized LLMs in this work.Specifically, we propose an analytic method to estimate the upper bound of calibration error (UBCE) for LLMs.Our method separately discusses the calibration error of the model's correct and incorrect predictions, indicating a theoretical improvement of calibration error caused by weight quantization.Our study demonstrates that quantized models consistently exhibit worse calibration performance than full-precision models, supported by consistent analysis across multiple LLMs and datasets.To address the calibration issues of quantized models, we propose a novel post-calibration method to recover the calibration performance of quantized models through soft-prompt tuning.Specifically, we inject soft tokens into quantized models after the embedding layers and optimize these tokens to recover the calibration error caused by weight quantization.Experimental results on multiple datasets demonstrate its effectiveness in improving the uncertainty calibration of quantized LLMs, facilitating more reliable weight quantization in resource-constrained environments.
Mingyu Zhong, Guanchu Wang, Yu-Neng Chuang, Na Zou 0001
ACL (1)2
2025 Personalizing Low-Rank Bayesian Neural Networks Via Federated Learning
abstract
To support real-world decision-making, it is crucial for models to be well-calibrated, i.e., to assign reliable confidence estimates to their predictions. Uncertainty quantification is particularly important in personalized federated learning (PFL), as participating clients typically have small local datasets, making it difficult to unambiguously determine optimal model parameters. Bayesian PFL (BPFL) methods can potentially enhance calibration, but they often come with considerable computational and memory requirements due to the need to track the variances of all the individual model parameters. Furthermore, different clients may exhibit heterogeneous uncertainty levels owing to varying local dataset sizes and distributions. To address these challenges, we propose LR-BPFL, a novel BPFL method that learns a global deterministic model along with personalized low-rank Bayesian corrections. To tailor the local model to each client’s inherent uncertainty level, LR-BPFL incorporates an adaptive rank selection mechanism. We evaluate LR-BPFL across a variety of datasets, demonstrating its advantages in terms of calibration, accuracy, as well as computational and memory requirements. The code is available at \url{https://github.com/Bernie0115/LR-BPFL.}
Dongzhu Liu, Osvaldo Simeone, Guanchu Wang, Dimitrios P. Pezaros, Guangxu Zhu
AISTATS4
2025 Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining
abstract
Low-rank optimization has emerged as a promising approach to enabling memory-efficient training of large language models (LLMs). Existing low-rank optimization methods typically project gradients onto a low-rank subspace, reducing the memory cost of storing optimizer states. A key challenge in these methods is selecting suitable subspaces to ensure an effective optimization trajectory. Most existing approaches select the dominant subspace to preserve gradient information, as this intuitively provides the best approximation. However, we find that in practice, the dominant subspace stops changing during pretraining, thereby constraining weight updates to similar subspaces. In this paper, we propose importance sampling for low-rank optimization in LLM pretraining with a provable convergence guarantee, which the dominant subspace approach does not have. Empirically, we demonstrate that our method significantly outperforms previous methods in LLM pretraining tasks.
Junze Yin, Guanchu Wang, Zirui Liu 0001, Lin Yang 0011, Tianyi Zhang 0011, Anshumali Shrivastava, Vladimir Braverman
NeurIPS3
2025 Efficient GNN Explanation via Learning Removal-based Attribution
abstract
As Graph Neural Networks (GNNs) have been widely used in real-world applications, model explanations are required not only by users but also by legal regulations. However, simultaneously achieving high fidelity and low computational costs in generating explanations has been a challenge for current methods. In this work, we propose a framework of GNN explanation named L e A rn R emoval-based A ttribution (LARA) to address this problem. Specifically, we introduce removal-based attribution and demonstrate its substantiated link to interpretability fidelity theoretically and experimentally. The explainer in LARA learns to generate removal-based attribution which enables providing explanations with high fidelity. A strategy of subgraph sampling is designed in LARA to improve the scalability of the training process. In the deployment, LARA can efficiently generate the explanation through a feed-forward pass. We benchmark our approach with other state-of-the-art GNN explanation methods on six datasets. Results highlight the effectiveness of our framework regarding both efficiency and fidelity. In particular, LARA is 3.1 \(\times\) faster and achieves higher fidelity than the state-of-the-art method on the large dataset ogbn-arxiv (more than 160K nodes and 1M edges), showing its great potential in real-world applications. Our source code is available at https://github.com/yaorong0921/LARA .
Yao Rong 0001, Guanchu Wang, Qizhang Feng, Ninghao Liu 0001, Zirui Liu 0001, Enkelejda Kasneci, Xia Ben Hu
ACM Trans. Knowl. Discov. Data2
2024 Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
abstract
Guanchu Wang, Yu-Neng Chuang, Ruixiang Tang, Shaochen Zhong, Jiayi Yuan, Hongye Jin, Zirui Liu, Vipin Chaudhary, Shuai Xu, James Caverlee, Xia Hu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Guanchu Wang, Yu-Neng Chuang, Ruixiang Tang, Shaochen Zhong, Jiayi Yuan 0001, Hongye Jin, Zirui Liu 0001, Vipin Chaudhary, James Caverlee, Xia Ben Hu
EMNLP1
2024 TVE: Learning Meta-attribution for Transferable Vision Explainer
abstract
Explainable machine learning significantly improves the transparency of deep neural networks. However, existing work is constrained to explaining the behavior of individual model predictions, and lacks the ability to transfer the explanation across various models and tasks. This limitation results in explaining various tasks being time- and resource-consuming. To address this problem, we introduce a Transferable Vision Explainer (TVE) that can effectively explain various vision models in downstream tasks. Specifically, the transferability of TVE is realized through a pre-training process on large-scale datasets towards learning the meta-attribution. This meta-attribution leverages the versatility of generic backbone encoders to comprehensively encode the attribution knowledge for the input instance, which enables TVE to seamlessly transfer to explaining various downstream tasks, without the need for training on task-specific data. Empirical studies involve explaining three different architectures of vision models across three diverse downstream datasets. The experiment results indicate TVE is effective in explaining these tasks without the need for additional training on downstream data.
Guanchu Wang, Yu-Neng Chuang, Fan Yang 0023, Mengnan Du, Chia-Yuan Chang 0002, Shaochen Zhong, Zirui Liu 0001, Zhaozhuo Xu, Kaixiong Zhou, Xuanting Cai, Xia Ben Hu
ICML1
2023 DiscoverPath: A Knowledge Refinement and Retrieval System for Interdisciplinarity on Biomedical Research
abstract
The exponential growth in scholarly publications necessitates advanced tools for efficient article retrieval, especially in interdisciplinary fields where diverse terminologies are used to describe similar research. Traditional keyword-based search engines often fall short in assisting users who may not be familiar with specific terminologies. To address this, we present a knowledge graph based paper search engine for biomedical research to enhance the user experience in discovering relevant queries and articles. The system, dubbed DiscoverPath, employs Named Entity Recognition (NER) and part-of-speech (POS) tagging to extract terminologies and relationships from article abstracts to create a KG. To reduce information overload, DiscoverPath presents users with a focused subgraph containing the queried entity and its neighboring nodes and incorporates a query recommendation system enabling users to iteratively refine their queries. The system is equipped with an accessible Graphical User Interface that provides an intuitive visualization of the KG, query recommendations, and detailed article information, enabling efficient article retrieval, thus fostering interdisciplinary knowledge exploration. DiscoverPath is open-sourced at https://github.com/ynchuang/DiscoverPath with a demo video at Youtube.
Yu-Neng Chuang, Guanchu Wang, Chia-Yuan Chang 0002, Kwei-Herng Lai, Daochen Zha, Ruixiang Tang, Fan Yang 0023, Alfredo Costilla-Reyes, Kaixiong Zhou, Xiaoqian Jiang, Xia Ben Hu
CIKM2
2023 CoRTX: Contrastive Framework for Real-time Explanation
Yu-Neng Chuang, Guanchu Wang, Fan Yang 0023, Pushkar Tripathi, Xuanting Cai, Xia Ben Hu
ICLR2
2023 DIVISION: Memory Efficient Training via Dual Activation Precision
abstract
Activation compressed training provides a solution towards reducing the memory cost of training deep neural networks (DNNs). However, state-of-the-art work combines a search of quantization bit-width with the training, which makes the procedure complicated and less transparent. To this end, we propose a simple and effective method to compress DNN training. Our method is motivated by an instructive observation: DNN backward propagation mainly utilizes the low-frequency component (LFC) of the activation maps, while the majority of memory is for caching the high-frequency component (HFC) during the training. This indicates the HFC of activation maps is highly redundant and compressible, which inspires our proposed Dual Activation Precision (DIVISION). During the training, DIVISION preserves a high-precision copy of LFC and compresses the HFC into a light-weight copy with low numerical precision. This can significantly reduce the memory cost while maintaining a competitive model accuracy. Experiment results show DIVISION has better comprehensive performance than state-of-the-art methods, including over 10x compression of activation maps and competitive training throughput, without loss of model accuracy. The source code is available at https://github.com/guanchuwang/division.
Guanchu Wang, Zirui Liu 0001, Zhimeng Jiang, Ninghao Liu 0001, Na Zou 0001, Xia Ben Hu
ICML1
2023 Chasing Fairness Under Distribution Shift: A Model Weight Perturbation Approach
abstract
Fairness in machine learning has attracted increasing attention in recent years. The fairness methods improving algorithmic fairness for in-distribution data may not perform well under distribution shifts. In this paper, we first theoretically demonstrate the inherent connection between distribution shift, data perturbation, and model weight perturbation. Subsequently, we analyze the sufficient conditions to guarantee fairness (i.e., low demographic parity) for the target dataset, including fairness for the source dataset, and low prediction difference between the source and target datasets for each sensitive attribute group. Motivated by these sufficient conditions, we propose robust fairness regularization (RFR) by considering the worst case within the model weight perturbation ball for each sensitive attribute group. We evaluate the effectiveness of our proposed RFR algorithm on synthetic and real distribution shifts across various datasets. Experimental results demonstrate that RFR achieves better fairness-accuracy trade-off performance compared with several baselines. The source code is available at \url{https://github.com/zhimengj0326/RFR_NeurIPS23}.
Zhimeng Jiang, Hongye Jin, Guanchu Wang, Rui Chen 0012, Na Zou 0001, Xia Ben Hu
NeurIPS4
2023 Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model
abstract
As the model size grows rapidly, fine-tuning the large pre-trained language model has become increasingly difficult due to its extensive memory usage. Previous works usually focus on reducing the number of trainable parameters in the network. While the model parameters do contribute to memory usage, the primary memory bottleneck during training arises from storing feature maps, also known as activations, as they are crucial for gradient calculation. Notably, machine learning models are typically trained using stochastic gradient descent. We argue that in stochastic optimization, models can handle noisy gradients as long as the gradient estimator is unbiased with reasonable variance. Following this motivation, we propose a new family of unbiased estimators called \sas, for matrix production with reduced variance, which only requires storing the sub-sampled activations for calculating the gradient. Our work provides both theoretical and experimental evidence that, in the context of tuning transformers, our proposed estimators exhibit lower variance compared to existing ones. By replacing the linear operation with our approximated one in transformers, we can achieve up to 2.7X peak memory reduction with almost no accuracy drop and enables up to $6.4\times$ larger batch size. Under the same hardware, \sas enables better down-streaming task performance by applying larger models and/or faster training speed with larger batch sizes. The code is available at https://anonymous.4open.science/r/WTACRS-A5C5/.
Zirui Liu 0001, Guanchu Wang, Shaochen Zhong, Zhaozhuo Xu, Daochen Zha, Ruixiang Tang, Zhimeng Jiang, Kaixiong Zhou, Vipin Chaudhary, Xia Ben Hu
NeurIPS2
2023 Mitigating Algorithmic Bias with Limited Annotations
Guanchu Wang, Mengnan Du, Ninghao Liu 0001, Na Zou 0001, Xia Ben Hu
ECML/PKDD (2)1
2022 BED: A Real-Time Object Detection System for Edge Devices
abstract
Deploying deep neural networks (DNNs) on edge devices provides efficient and effective solutions for the real-world tasks. Edge devices have been used for collecting a large volume of data efficiently in different domains. DNNs have been an effective tool for data processing and analysis. However, designing DNNs on edge devices is challenging due to the limited computational resources and memory. To tackle this challenge, we demonstrate oBject detection system for Edge Devices (BED) on the MAX78000 DNN accelerator. It integrates on-device DNN inference with a camera and an LCD display for image acquisition and detection exhibition, respectively. BED is a concise, effective and detailed solution, including model training, quantization, synthesis and deployment. The entire repository is open-sourced on Github1, including a Graphical User Interface (GUI) for on-chip debugging. Experiment results indicate that BED can produce accurate detection with a 300-KB tiny DNN model, which takes only 91.9 ms of inference time and 1.845 mJ of energy. The real-time detection is available at YouTube.
Guanchu Wang, Zaid Pervaiz Bhat, Zhimeng Jiang, Yi-Wei Chen, Daochen Zha, Alfredo Costilla-Reyes, Afshin Niktash, Mehmet Görkem Ulkar, Osman Erman Okman, Xuanting Cai, Xia Ben Hu
CIKM1
2022 Accelerating Shapley Explanation via Contributive Cooperator Selection
abstract
Even though Shapley value provides an effective explanation for a DNN model prediction, the computation relies on the enumeration of all possible input feature coalitions, which leads to the exponentially growing complexity. To address this problem, we propose a novel method SHEAR to significantly accelerate the Shapley explanation for DNN models, where only a few coalitions of input features are involved in the computation. The selection of the feature coalitions follows our proposed Shapley chain rule to minimize the absolute error from the ground-truth Shapley values, such that the computation can be both efficient and accurate. To demonstrate the effectiveness, we comprehensively evaluate SHEAR across multiple metrics including the absolute error from the ground-truth Shapley value, the faithfulness of the explanations, and running speed. The experimental results indicate SHEAR consistently outperforms state-of-the-art baseline methods across different evaluation metrics, which demonstrates its potentials in real-world applications where the computational resource is limited.
Guanchu Wang, Yu-Neng Chuang, Mengnan Du, Fan Yang 0023, Pushkar Tripathi, Xuanting Cai, Xia Ben Hu
ICML1
2021 TODS: An Automated Time Series Outlier Detection System
abstract
We present TODS, an automated Time Series Outlier Detection System for research and industrial applications. TODS is a highly modular system that supports easy pipeline construction. The basic building block of TODS is primitive, which is an implementation of a function with hyperparameters. TODS currently supports 70 primitives, including data processing, time series processing, feature analysis, detection algorithms, and a reinforcement module. Users can freely construct a pipeline using these primitives and perform end- to-end outlier detection with the constructed pipeline. TODS provides a Graphical User Interface (GUI), where users can flexibly design a pipeline with drag-and-drop. Moreover, a data-driven searcher is provided to automatically discover the most suitable pipelines given a dataset. TODS is released under Apache 2.0 license at https://github.com/datamllab/tods. A video is available on YouTube (https://youtu.be/JOtYxTclZgQ)
Kwei-Herng Lai, Daochen Zha, Guanchu Wang, Yue Zhao 0016, Yile Chen 0002, Purav Zumkhawaka, Minyang Wan, Diego Martinez, Xia Ben Hu
AAAI3
2021 Unsupervised Discovery of Transitional Skills for Deep Reinforcement Learning
abstract
By maximizing an information theoretic objective, a few recent methods empower the agent to explore the environment and learn skills without extrinsic reward. However, when considering using multiple consecutive skills to complete a specific task, the transition from one to another cannot guarantee the success of the process due to the evident gap between skills. In this paper, we propose a novel unsupervised reinforcement learning approach to learn transitional skills in addition to pursuing diverse primitive skills. By introducing an extra latent variable for exploring the dependence between skills, our method discovers both primitive and transitional skills by optimizing a novel information theoretic objective. Considering various robotic tasks, our results demonstrate the effectiveness on learning both diverse primitive skills and transitional skills, and further exhibit the superiority of our method in smooth transition of skills over the baselines. Videos of transitional skills can be found on the project website: https://sites.google.com/view/udts-skill.
Qiangxing Tian, Guanchu Wang
IJCNN3
2021 Fairness via Representation Neutralization
abstract
Existing bias mitigation methods for DNN models primarily work on learning debiased encoders. This process not only requires a lot of instance-level annotations for sensitive attributes, it also does not guarantee that all fairness sensitive information has been removed from the encoder. To address these limitations, we explore the following research question: Can we reduce the discrimination of DNN models by only debiasing the classification head, even with biased representations as inputs? To this end, we propose a new mitigation technique, namely, Representation Neutralization for Fairness (RNF) that achieves fairness by debiasing only the task-specific classification head of DNN models. To this end, we leverage samples with the same ground-truth label but different sensitive attributes, and use their neutralized representations to train the classification head of the DNN model. The key idea of RNF is to discourage the classification head from capturing spurious correlation between fairness sensitive information in encoder representations with specific class labels. To address low-resource settings with no access to sensitive attribute annotations, we leverage a bias-amplified model to generate proxy annotations for sensitive attributes. Experimental results over several benchmark datasets demonstrate our RNF framework to effectively reduce discrimination of DNN models with minimal degradation in task-specific performance.
Mengnan Du, Subhabrata Mukherjee, Guanchu Wang, Ruixiang Tang, Ahmed Awadallah 0001, Xia Ben Hu
NeurIPS3
2021 On the Achievable Rate and Capacity for a Sample-Based Practical Photon-Counting Receiver
abstract
We investigate the achievable rate and capacity of a non-perfect photon-counting receiver assuming that the shot and thermal noise is negligible, and that the sampling interval is not shorter than the dead time. For long and fixed symbol duration, the achievable rate under on-off keying modulation is investigated based on Kullback-Leibler divergence and Chernoff divergence. We prove the tightness of the derived bounds for large peak power with zero background radiation at an exponential convergence rates, and for low peak power at an order-two convergence rates. Moreover, we propose an approximation on the achievable rate, which is more accurate compared with the derived bounds in the medium signal to noise ratio (SNR) regime. We also verify the accuracy of the proposed model on the achievable rate analysis via the comparison with the model with non-negligible shot and thermal noise. As for the capacity analysis, we assume that the symbol duration can be arbitrarily short, and demonstrate that the capacity approaches that of the continuous-time Poisson channel as both the sampling interval and the symbol duration approach zero, and the sampling interval equals the symbol duration. For large peak power, the capacity with a non-perfect receiver converges, while that of continuous Poisson capacity channel linearly increases.
Zhimeng Jiang, Chen Gong 0001, Guanchu Wang, Zhengyuan Xu
IEEE Trans. Commun.3
2020 Independent Skill Transfer for Deep Reinforcement Learning
abstract
Recently, diverse primitive skills have been learned by adopting the entropy as intrinsic reward, which further shows that new practical skills can be produced by combining a variety of primitive skills. This is essentially skill transfer, very useful for learning high-level skills but quite challenging due to the low efficiency of transferring primitive skills. In this paper, we propose a novel efficient skill transfer method, where we learn independent skills and only independent components of skills are transferred instead of the whole set of skills. More concretely, independent components of skills are obtained through independent component analysis (ICA), which always have a smaller amount (or lower dimension) compared with their mixtures. With a lower dimension, independent skill transfer (IST) exhibits a higher efficiency on learning a given task. Extensive experiments including three robotic tasks demonstrate the effectiveness and high efficiency of our proposed IST method in comparison to direct primitive-skill transfer and conventional reinforcement learning.
Qiangxing Tian, Guanchu Wang, Yachen Kang
IJCAI2
2018 Signal Characterization for Multiple Access Non-Line of Sight Scattering Communication
abstract
Due to the extremely large path loss of non-line of sight optical wireless communication, the received signal exhibits the characteristics of discrete photoelectrons. In this paper, we investigate the achievable rates and signal detection of on-off keying modulation for discrete-time Poisson multiple access channel . Both non-orthogonal multiple access and code division multiple access are considered, for both single-rate and multi-rate multiple access channels. Maximum likelihood (ML) detections are adopted as the optimal criterion for signal detection, whose complexity grows exponentially with the number of users. Low computational complexity detections are proposed with M-order correction that can perform closely to the ML detection. Numerical and simulation results are presented to show the achievable rate regions, optimal power allocations, the performance of ML, and proposed low-complexity detections. Finally, the experimental results of the proposed detection approaches agree with the simulation results.
Guanchu Wang, Chen Gong 0001, Zhengyuan Xu
IEEE Trans. Commun.1
2017 Signal Detection and Achievable Rates for Multiple Access Optical Wireless Scattering Communication
abstract
Due to the extremely large path loss of non-line of sight optical wireless communication, the received signal exhibits the characteristics of discrete photons. In this work we investigate the signal detection and achievable rates of OOK modulation for discrete Poisson multiple access channel (MAC), for both code division multiple access (CDMA) and non-orthogonal multiple access (NOMA). For the signal detection, we adopt the maximum likelihood (ML) criterion for the optimal detection, whose complexity grows exponentially with the number of users. Low computational complexity detection is proposed with ℳ-order correction algorithm that can perform closely to ML detection. Moreover, we address the achievable rates and rationalize the power allocation for discrete Poisson MAC for both NOMA and CDMA. Numerical and simulation results are presented to show the performance of the proposed detection, the achievable rate region as well as the optimal power allocation of two-user discrete Poisson NOMA and CDMA.
Guanchu Wang, Chen Gong 0001, Zhengyuan Xu
GLOBECOM1