Chen Liang 0006

dblp:35/3221-6 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
13since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Taxonomy-Based Negative Sampling in Personalized Semantic Search for E-Commerce
Uthman Jinadu, Siawpeng Er, Chen Liang 0006, Aleksandar Velkoski
IEEE Big Data4
2025 Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
abstract
Efficiently modeling sequences with infinite context length has long been a challenging problem. Previous approaches have either suffered from quadratic computational complexity or limited extrapolation ability in length generalization. In this work, we present Samba, a simple hybrid architecture that layer-wise combines Mamba, a selective State Space Model (SSM), with Sliding Window Attention (SWA). Samba selectively compresses a given sequence into recurrent hidden states while still maintaining the ability to precisely recall recent memories with the attention mechanism. We scale Samba up to 3.8B parameters with 3.2T training tokens and demonstrate that it significantly outperforms state-of-the-art models across a variety of benchmarks. Pretrained on sequences of 4K length, Samba shows improved perplexity in context lengths of up to 1M in zero-shot. When finetuned on 4K-length sequences, Samba efficiently extrapolates to a 256K context length with perfect memory recall on the Passkey Retrieval task, and exhibits superior retrieval extrapolation on the challenging Phonebook task compared to full-attention models. As a linear-time sequence model, Samba achieves a 3.73× higher throughput compared to Transformers with grouped-query attention for user prompts of 128K length, and a 3.64× speedup when generating 64K tokens with unlimited streaming.
Liliang Ren, Yang Liu 0003, Yadong Lu, Yelong Shen, Chen Liang 0006, Weizhu Chen
ICLR5
2024 LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language Models
abstract
Quantization is an indispensable technique for serving Large Language Models (LLMs) and has recently found its way into LoRA fine-tuning (Dettmers et al., 2023). In this work we focus on the scenario where quantization and LoRA fine- tuning are applied together on a pre-trained model. In such cases it is common to observe a consistent gap in the performance on downstream tasks between full fine-tuning and quantization plus LoRA fine-tuning approach. In response, we propose LoftQ (LoRA-Fine-Tuning-aware Quantization), a novel quantization framework that simultaneously quantizes an LLM and finds a proper low-rank initialization for LoRA fine-tuning. Such an initialization alleviates the discrep- ancy between the quantized and full-precision model and significantly improves the generalization in downstream tasks. We evaluate our method on natural lan- guage understanding, question answering, summarization, and natural language generation tasks. Experiments show that our method is highly effective and out- performs existing quantization methods, especially in the challenging 2-bit and 2/4-bit mixed precision regimes. We will release our code.
Yifan Yu 0008, Chen Liang 0006, Nikos Karampatziakis, Weizhu Chen, Tuo Zhao
ICLR3
2023 HomoDistil: Homotopic Task-Agnostic Distillation of Pre-trained Transformers
Chen Liang 0006, Haoming Jiang, Zheng Li 0018, Xianfeng Tang, Tuo Zhao
ICLR1
2023 LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation
abstract
Transformer models have achieved remarkable results in various natural language tasks, but they are often prohibitively large, requiring massive memories and computational resources. To re- duce the size and complexity of these models, we propose LoSparse (Low-Rank and Sparse ap- proximation), a novel model compression tech- nique that approximates a weight matrix by the sum of a low-rank matrix and a sparse matrix. Our method combines the advantages of both low- rank approximations and pruning, while avoid- ing their limitations. Low-rank approximation compresses the coherent and expressive parts in neurons, while pruning removes the incoherent and non-expressive parts in neurons. Pruning enhances the diversity of low-rank approxima- tions, and low-rank approximation prevents prun- ing from losing too many expressive neurons. We evaluate our method on natural language under- standing, question answering, and natural lan- guage generation tasks. We show that it signif- icantly outperforms existing compression meth- ods. Our code is publicly available at https: //github.com/yxli2123/LoSparse
Yifan Yu 0008, Qingru Zhang, Chen Liang 0006, Weizhu Chen, Tuo Zhao
ICML4
2023 Less is More: Task-aware Layer-wise Distillation for Language Model Compression
abstract
Layer-wise distillation is a powerful tool to compress large models (i.e. teacher models) into small ones (i.e., student models). The student distills knowledge from the teacher by mimicking the hidden representations of the teacher at every intermediate layer. However, layer-wise distillation is difficult. Since the student has a smaller model capacity than the teacher, it is often under-fitted. Furthermore, the hidden representations of the teacher contain redundant information that the student does not necessarily need for the target task's learning. To address these challenges, we propose a novel Task-aware layEr-wise Distillation (TED). TED designs task-aware filters to align the hidden representations of the student and the teacher at each layer. The filters select the knowledge that is useful for the target task from the hidden representations. As such, TED reduces the knowledge gap between the two models and helps the student to fit better on the target task. We evaluate TED in two scenarios: continual pre-training and fine-tuning. TED demonstrates significant and consistent improvements over existing distillation methods in both scenarios. Code is available at https://github.com/cliang1453/task-aware-distillation.
Chen Liang 0006, Simiao Zuo, Qingru Zhang, Weizhu Chen, Tuo Zhao
ICML1
2023 Module-wise Adaptive Distillation for Multimodality Foundation Models
abstract
Pre-trained multimodal foundation models have demonstrated remarkable generalizability but pose challenges for deployment due to their large sizes. One effective approach to reducing their sizes is layerwise distillation, wherein small student models are trained to match the hidden representations of large teacher models at each layer. Motivated by our observation that certain architecture components, referred to as modules, contribute more significantly to the student's performance than others, we propose to track the contributions of individual modules by recording the loss decrement after distillation each module and choose the module with a greater contribution to distill more frequently. Such an approach can be naturally formulated as a multi-armed bandit (MAB) problem, where modules and loss decrements are considered as arms and rewards, respectively. We then develop a modified-Thompson sampling algorithm named OPTIMA to address the nonstationarity of module contributions resulting from model updating. Specifically, we leverage the observed contributions in recent history to estimate the changing contribution of each module and select modules based on these estimations to maximize the cumulative contribution. We evaluate the effectiveness of OPTIMA through distillation experiments on various multimodal understanding and image captioning tasks, using the CoCa-Large model \citep{yu2022coca} as the teacher model.
Chen Liang 0006, Ming-Hsuan Yang 0001, Matthew Brown 0001, Yin Cui, Tuo Zhao, Boqing Gong, Tianyi Zhou 0002
NeurIPS1
2022 CAMERO: Consistency Regularized Ensemble of Perturbed Language Models with Weight Sharing
abstract
Model ensemble is a popular approach to produce a low-variance and well-generalized model.However, it induces large memory and inference costs, which are often not affordable for real-world deployment.Existing work has resorted to sharing weights among models.However, when increasing the proportion of the shared weights, the resulting models tend to be similar, and the benefits of using model ensemble diminish.To retain ensemble benefits while maintaining a low memory cost, we propose a consistency-regularized ensemble learning approach based on perturbed models, named CAMERO.Specifically, we share the weights of bottom layers across all models and apply different perturbations to the hidden representations for different models, which can effectively promote the model diversity.Meanwhile, we apply a prediction consistency regularizer across the perturbed models to control the variance due to the model diversity.Our experiments using large language models demonstrate that CAMERO significantly improves the generalization performance of the ensemble model.Specifically, CAMERO outperforms the standard ensemble of 8 BERT-base models on the GLUE benchmark by 0.7 with a significantly smaller model size (114.2Mvs. 880.6M).* Work was done during an internship at Microsoft Azure AI.
Chen Liang 0006, Yelong Shen, Weizhu Chen, Tuo Zhao
ACL (1)1
2022 No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer Models
Chen Liang 0006, Haoming Jiang, Simiao Zuo, Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen, Tuo Zhao
ICLR1
2022 PLATON: Pruning Large Transformer Models with Upper Confidence Bound of Weight Importance
abstract
Large Transformer-based models have exhibited superior performance in various natural language processing and computer vision tasks. However, these models contain enormous amounts of parameters, which restrict their deployment to real-world applications. To reduce the model size, researchers prune these models based on the weights’ importance scores. However, such scores are usually estimated on mini-batches during training, which incurs large variability/uncertainty due to mini-batch sampling and complicated training dynamics. As a result, some crucial weights could be pruned by commonly used pruning methods because of such uncertainty, which makes training unstable and hurts generalization. To resolve this issue, we propose PLATON, which captures the uncertainty of importance scores by upper confidence bound of importance estimation. In particular, for the weights with low importance scores but high uncertainty, PLATON tends to retain them and explores their capacity. We conduct extensive experiments with several Transformer-based models on natural language understanding, question answering and image classification to validate the effectiveness of PLATON. Results demonstrate that PLATON manifests notable improvement under different sparsity levels. Our code is publicly available at https://github.com/QingruZhang/PLATON.
Qingru Zhang, Simiao Zuo, Chen Liang 0006, Alexander Bukharin, Weizhu Chen, Tuo Zhao
ICML3
2022 MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation
abstract
Simiao Zuo, Qingru Zhang, Chen Liang, Pengcheng He, Tuo Zhao, Weizhu Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Simiao Zuo, Qingru Zhang, Chen Liang 0006, Tuo Zhao, Weizhu Chen
NAACL-HLT3
2021 Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
abstract
Chen Liang, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu, Pengcheng He, Tuo Zhao, Weizhu Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chen Liang 0006, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu 0003, Tuo Zhao, Weizhu Chen
ACL/IJCNLP (1)1
2021 Adversarial Regularization as Stackelberg Game: An Unrolled Optimization Approach
abstract
Adversarial regularization has been shown to improve the generalization performance of deep learning models in various natural language processing tasks.Existing works usually formulate the method as a zero-sum game, which is solved by alternating gradient descent/ascent algorithms.Such a formulation treats the adversarial and the defending players equally, which is undesirable because only the defending player contributes to the generalization performance.To address this issue, we propose Stackelberg Adversarial Regularization (SALT), which formulates adversarial regularization as a Stackelberg game.This formulation induces a competition between a leader and a follower, where the follower generates perturbations, and the leader trains the model subject to the perturbations.Different from conventional approaches, in SALT, the leader is in an advantageous position.When the leader moves, it recognizes the strategy of the follower and takes the anticipated follower's outcomes into consideration.Such a leader's advantage enables us to improve the model fitting to the unperturbed data.The leader's strategic information is captured by the Stackelberg gradient, which is obtained using an unrolling algorithm.Our experimental results on a set of machine translation and natural language understanding tasks show that SALT outperforms existing adversarial regularization baselines across all tasks.Our code is publicly available.
Simiao Zuo, Chen Liang 0006, Haoming Jiang, Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen, Tuo Zhao
EMNLP (1)2
2020 Multi-Domain Neural Machine Translation with Word-Level Adaptive Layer-wise Domain Mixing
abstract
Many multi-domain neural machine translation (NMT) models achieve knowledge transfer by enforcing one encoder to learn shared embedding across domains.However, this design lacks adaptation to individual domains.To overcome this limitation, we propose a novel multi-domain NMT model using individual modules for each domain, on which we apply word-level, adaptive and layer-wise domain mixing.We first observe that words in a sentence are often related to multiple domains.Hence, we assume each word has a domain proportion, which indicates its domain preference.Then word representations are obtained by mixing their embedding in individual domains based on their domain proportions.We show this can be achieved by carefully designing multi-head dot-product attention modules for different domains, and eventually taking weighted averages of their parameters by word-level layer-wise domain proportions.Through this, we can achieve effective domain knowledge sharing, and capture fine-grained domain-specific knowledge as well.Our experiments show that our proposed model outperforms existing ones in several NMT tasks.
Haoming Jiang, Chen Liang 0006, Tuo Zhao
ACL2
2020 BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant Supervision
abstract
We study the open-domain named entity recognition (NER) problem under distant supervision. The distant supervision, though does not require large amounts of manual annotations, yields highly incomplete and noisy distant labels via external knowledge bases. To address this challenge, we propose a new computational framework -- BOND, which leverages the power of pre-trained language models (e.g., BERT and RoBERTa) to improve the prediction performance of NER models. Specifically, we propose a two-stage training algorithm: In the first stage, we adapt the pre-trained language model to the NER tasks using the distant labels, which can significantly improve the recall and precision; In the second stage, we drop the distant labels, and propose a self-training approach to further improve the model performance. Thorough experiments on 5 benchmark datasets demonstrate the superiority of BOND over existing distantly supervised NER methods. The code and distantly labeled data have been released in https://github.com/cliang1453/BOND.
Chen Liang 0006, Yue Yu 0001, Haoming Jiang, Siawpeng Er, Tuo Zhao, Chao Zhang 0014
KDD1