Wei Zhu 0016

dblp:83/4805-16 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-6389-6866ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TS-DiT: A diffusion transformer model for time series forecasting
Wei Zhu 0016, Jinhui Song
Expert Syst. Appl.2
2024 IAPT: Instance-Aware Prompt Tuning for Large Language Models
abstract
Soft prompt tuning is a widely studied parameter-efficient fine-tuning method.However, it has a clear drawback: many soft tokens must be inserted into the input sequences to guarantee downstream performance.As a result, soft prompt tuning is less considered than Low-rank adaptation (LoRA) in the large language modeling (LLM) era.In this work, we propose a novel prompt tuning method, Instruction-Aware Prompt Tuning (IAPT), that requires only four soft tokens.First, we install a parameter-efficient soft prompt generator at each Transformer layer to generate idiosyncratic soft prompts for each input instruction.The generated soft prompts can be seen as a semantic summary of the input instructions and can effectively guide the output generation.Second, the soft prompt generators are modules with a bottleneck architecture consisting of a self-attention pooling operation, two linear projections, and an activation function.Pilot experiments show that prompt generators at different Transformer layers require different activation functions.Thus, we propose to learn the idiosyncratic activation functions for prompt generators automatically with the help of rational functions.We have conducted experiments on various tasks, and the experimental results demonstrate that (a) our IAPT method can outperform the recent baselines with comparable tunable parameters.(b) Our IAPT method is more efficient than LoRA under the singlebackbone multi-tenant setting.
Wei Zhu 0016, Aaron Xuxiang Tian, Congrui Yin, Yuan Ni, Xiaoling Wang 0004, Guo Tong Xie
ACL (1)1
2024 Chimera Model of Candidate Soups for Non-Autoregressive Translation
Huanran Zheng, Wei Zhu 0016, Xiaoling Wang 0004
DASFAA (2)2
2024 ULTRAFEEDBACK: Boosting Language Models with Scaled AI Feedback
abstract
Learning from human feedback has become a pivot technique in aligning large language models (LLMs) with human preferences. However, acquiring vast and premium human feedback is bottlenecked by time, labor, and human capability, resulting in small sizes or limited topics of current datasets. This further hinders feedback learning as well as alignment research within the open-source community. To address this issue, we explore how to go beyond human feedback and collect high-quality AI feedback automatically for a scalable alternative. Specifically, we identify scale and diversity as the key factors for feedback data to take effect. Accordingly, we first broaden instructions and responses in both amount and breadth to encompass a wider range of user-assistant interactions. Then, we meticulously apply a series of techniques to mitigate annotation biases for more reliable AI feedback. We finally present UltraFeedback, a large-scale, high-quality, and diversified AI feedback dataset, which contains over 1 million GPT-4 feedback for 250k user-assistant conversations from various aspects. Built upon UltraFeedback, we align a LLaMA-based model by best-of-$n$ sampling and reinforcement learning, demonstrating its exceptional performance on chat benchmarks. Our work validates the effectiveness of scaled AI feedback data in constructing strong open-source chat language models, serving as a solid foundation for future feedback learning research.
Ganqu Cui, Lifan Yuan, Ning Ding 0002, Guanming Yao, Bingxiang He, Wei Zhu 0016, Yuan Ni, Guo Tong Xie, Ruobing Xie, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001
ICML6
2024 NAT4AT: Using Non-Autoregressive Translation Makes Autoregressive Translation Faster and Better
abstract
With the increasing number of web documents, the demand for translation has increased dramatically. Non-autoregressive translation (NAT) models can significantly reduce decoding latency to meet the growing translation needs, but they sacrifice translation quality. And there is still an irreparable performance gap between NAT models and strong autoregressive translation (AT) models at the corpus level. However, more fine-grained comparative experiments on AT and NAT are currently lacking. Therefore, in this paper, we first conducted analysis experiments at the sentence level and found complementarity and high similarity between the translations generated by AT and NAT. Then, based on this observation, we propose a general and effective method called NAT4AT, which can not only use NAT to speed up the inference speed of AT significantly but also improve its final translation quality. Specifically, NAT4AT first uses a NAT model to generate an original translation in parallel and then uses an AT model as a correction model to revise errors in the original translation. In this way, the AT model no longer needs to predict the entire translation but only needs to predict a small number of error parts in the NAT result. Extensive experimental results on major WMT benchmarks verify the generality and effectiveness of our method, whose translation quality is superior to the strong AT model and achieves a 5.0x speedup.
Huanran Zheng, Wei Zhu 0016, Xiaoling Wang 0004
WWW2
2023 Unified Demonstration Retriever for In-Context Learning
abstract
Xiaonan Li, Kai Lv, Hang Yan, Tianyang Lin, Wei Zhu, Yuan Ni, Guotong Xie, Xiaoling Wang, Xipeng Qiu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Kai Lv 0001, Hang Yan 0001, Tianyang Lin, Wei Zhu 0016, Yuan Ni, Guo Tong Xie, Xiaoling Wang 0004, Xipeng Qiu
ACL (1)5
2023 F-PABEE: Flexible-Patience-Based Early Exiting For Single-Label and Multi-Label Text Classification Tasks
abstract
Computational complexity and overthinking problems have become the bottlenecks for pre-training language models (PLMs) with millions or even trillions of parameters. A Flexible-Patience-Based Early Exiting method (F-PABEE) has been proposed to alleviate the problems mentioned above for single-label classification (SLC) and multi-label classification (MLC) tasks. F-PABEE makes predictions at the classifier and will exit early if predicted distributions of cross-layer are consecutively similar. It is more flexible than the previous state-of-the-art (SOTA) early exiting method PABEE because it can simultaneously adjust the similarity score thresholds and the patience parameters. Extensive experiments show that: (1) F-PABEE makes a better speedup-accuracy balance than existing early exiting strategies on both SLC and MLC tasks. (2) F-PABEE achieves faster inference and better performances on different PLMs such as BERT and ALBERT. (3) F-PABEE-JSKD performs best for F-PABEE with different similarity measures.
Xiangxiang Gao, Wei Zhu 0016, Jiasheng Gao, Congrui Yin
ICASSP2
2023 ACF: Aligned Contrastive Finetuning For Language and Vision Tasks
abstract
Contrastive learning (CL) has achieved great success in various fields with self-supervised learning. However, CL under the supervised setting is not fully explored, especially how to utilize the class labels in CL. We propose a novel aligned contrastive finetuning (ACF) approach in this work. Specifically, we consider the label embeddings as labeled instances and put them in an InfoNCE loss objective together with the instance representations, thus aligning the label embeddings and instance representation in the same semantic space. In addition, we design a correlation-based regularization term to alleviate the anisotropy problem. Extensive experiments are conducted on language understanding and image classification tasks, demonstrating our ACF method’s competitiveness. ACF is off-the-shelf and can be plugged into any pre-trained models without additional network architectures or computation overhead.
Wei Zhu 0016, Xiaoling Wang 0004, Yuan Ni, Guo Tong Xie
ICASSP1
2023 Multi-task entity linking with supervision from a taxonomy
Xuwu Wang, Wei Zhu 0016, Yuan Ni, Guo Tong Xie, Deqing Yang, Yanghua Xiao
Knowl. Inf. Syst.3
2022 Candidate Soups: Fusing Candidate Results Improves Translation Quality for Non-Autoregressive Translation
abstract
Non-autoregressive translation (NAT) model achieves a much faster inference speed than the autoregressive translation (AT) model because it can simultaneously predict all tokens during inference.However, its translation quality suffers from degradation compared to AT.And existing NAT methods only focus on improving the NAT model's performance but do not fully utilize it.In this paper, we propose a simple but effective method called "Candidate Soups," which can obtain high-quality translations while maintaining the inference speed of NAT models.Unlike previous approaches that pick the individual result and discard the remainders, Candidate Soups (CDS) can fully use the valuable information in the different candidate translations through model uncertainty.Extensive experiments on two benchmarks (WMT'14 EN-DE and WMT'16 EN-RO) demonstrate the effectiveness and generality of our proposed method, which can significantly improve the translation quality of various base models.More notably, our best variant outperforms the AT model on three translation tasks with 7.6× speedup. 1
Huanran Zheng, Wei Zhu 0016, Pengfei Wang 0009, Xiaoling Wang 0004
EMNLP2
2021 LeeBERT: Learned Early Exit for BERT with cross-level optimization
abstract
Wei Zhu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei Zhu 0016
ACL/IJCNLP (1)1
2021 GAML-BERT: Improving BERT Early Exiting by Gradient Aligned Mutual Learning
abstract
In this work, we propose a novel framework, Gradient Aligned Mutual Learning BERT (GAML-BERT), for improving the early exiting of BERT.GAML-BERT's contributions are two-fold.We conduct a set of pilot experiments, which shows that mutual knowledge distillation between a shallow exit and a deep exit leads to better performances for both.From this observation, we use mutual learning to improve BERT's early exiting performances, that is, we ask each exit of a multi-exit BERT to distill knowledge from each other.Second, we propose GA, a novel training method that aligns the gradients from knowledge distillation to cross-entropy losses.Extensive experiments are conducted on the GLUE benchmark, which shows that our GAML-BERT can significantly outperform the state-of-the-art (SOTA) BERT early exiting methods.
Wei Zhu 0016, Xiaoling Wang 0004, Yuan Ni, Guo Tong Xie
EMNLP (1)1
2021 AutoNLU: Architecture Search for Sentence and Cross-sentence Attention Modeling with Re-designed Search Space
Wei Zhu 0016
NLPCC (1)1
2021 AutoTrans: Automating Transformer Design via Reinforced Architecture Search
Wei Zhu 0016, Xiaoling Wang 0004, Yuan Ni, Guo Tong Xie
NLPCC (1)1
2020 Mining Infrequent High-Quality Phrases from Domain-Specific Corpora
abstract
Phrase mining is a fundamental task for text analysis and has various downstream applications such as named entity recognition, topic modeling, and relation extraction. In this paper, we focus on mining high-quality phrases from domain-specific corpora with special consideration of infrequent ones. Previous methods might miss infrequent high-quality phrases in the candidate selection stage. And these methods rely on explicit features to mine phrases while rarely considering the implicit features. In addition, completeness is rarely explicitly considered in the evaluation of a high-quality phrase. In this paper, we propose a novel approach that exploits a sequence labeling model to capture infrequent phrases. And we employ implicit semantic features and contextual POS tag statistics to measure meaningfulness and completeness, respectively. Experiments over four real-world corpora demonstrate that our method achieves significant improvements over previous state-of-the-art methods across different domains and languages.
Wei Zhu 0016, Sihang Jiang 0001, Sheng Zhang 0027, Yuan Ni, Guo Tong Xie, Yanghua Xiao
CIKM2
2020 Automatic Student Network Search for Knowledge Distillation
abstract
Pre-trained language models (PLMs), such as BERT, have achieved outstanding performance on multiple natural language processing (NLP) tasks. However, such pre-trained models usually contain a huge number of parameters and are computationally expensive. The high resource demand hinders their application on resource-restricted devices like mobile phones. Knowledge distillation (KD) is an effective compression approach, aiming at encouraging a light-weight student network to imitate the teacher network, and accordingly latent knowledge is transferred from the teacher to student. However, the great majority of student networks in previous KD methods are manually designed, normally a subnetwork of the teacher network. Transformer is generally utilized as the student for compressing BERT but still contains masses of parameters. Motivated by this, we propose a novel approach named NAS-KD, which automatically generates an optimal student network using neural architecture search (NAS) to enhance the distillation for BERT. Experiment on 7 classification tasks in NLP domain demonstrates that NAS-KD can substantially reduce the size of BERT without much performance sacrifice.
Zhexi Zhang, Wei Zhu 0016, Junchi Yan, Peng Gao 0015, Guo Tong Xie
ICPR2