EDBT 2026 Demo / reviewers in the wild / expert
Canwen Xu
dblp:234/8009
· DBLP profile ↗
20ranked-venue papers
9as first author
14since 2021 · last 2024
0000-0002-1552-999XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsabstractLarge Language Models (LLMs) have greatly advanced code auto-completion systems, with a potential for substantial productivity enhancements for developers. However, current benchmarks mainly focus on single-file tasks, leaving an assessment gap for more complex, real-world, multi-file programming scenarios. To fill this gap, we introduce RepoBench, a new benchmark specifically designed for evaluating repository-level code auto-completion systems. RepoBench consists of three interconnected evaluation tasks: RepoBench-R (Retrieval), RepoBench-C (Code Completion), and RepoBench-P (Pipeline). Each task respectively measures the system's ability to retrieve the most relevant code snippets from other files as cross-file context, predict the next line of code with cross-file and in-file context, and handle complex tasks that require a combination of both retrieval and next-line prediction. RepoBench aims to facilitate a more complete comparison of performance and encouraging continuous improvement in auto-completion systems. RepoBench is actively maintained with the latest code, serving as a live benchmark publicly available at https://github.com/Leolty/repobench. Tianyang Liu 0003, Canwen Xu, Julian J. McAuley |
ICLR | 2 |
| 2023 | A Survey on Model Compression and Acceleration for Pretrained Language ModelsabstractDespite achieving state-of-the-art performance on many NLP tasks, the high energy cost and long inference delay prevent Transformer-based pretrained language models (PLMs) from seeing broader adoption including for edge and mobile computing. Efficient NLP research aims to comprehensively consider computation, time and carbon emission for the entire life-cycle of NLP, including data preparation, model training and inference. In this survey, we focus on the inference stage and review the current state of model compression and acceleration for pretrained language models, including benchmarks, metrics and methodology. Canwen Xu, Julian J. McAuley |
AAAI | 1 |
| 2023 | Spoiler Detection as Semantic Text MatchingabstractEngaging with discussion of TV shows online often requires individuals to refrain from consuming show-related content for extended periods to avoid spoilers.While existing research on spoiler detection shows promising results in safeguarding viewers from general spoilers, it fails to address the issue of users abstaining from show-related content during their watch.This is primarily because the definition of a spoiler varies depending on the viewer's progress in the show, and conventional spoiler detection methods lack the granularity to capture this complexity.To tackle this challenge, we propose the task of spoiler matching, which involves assigning an episode number to a spoiler given a specific TV show.We frame this task as semantic text matching and introduce a dataset comprised of comments and episode summaries to evaluate model performance.Given the length of each example, our dataset can also serve as a benchmark for longrange language models. 1 2 Ryan Tran, Canwen Xu, Julian J. McAuley |
EMNLP | 2 |
| 2023 | Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataabstractChat models, such as ChatGPT, have shown impressive capabilities and have been rapidly adopted across numerous domains.However, these models are only accessible through a restricted API, creating barriers for new research and progress in the field.We propose a pipeline that can automatically generate a highquality multi-turn chat corpus by leveraging ChatGPT to engage in a conversation with itself.Subsequently, we employ parameter-efficient tuning to enhance LLaMA, an open-source large language model.The resulting model, named Baize, demonstrates good performance in multi-turn dialogues with guardrails that minimize potential risks.Additionally, we propose a new technique called Self-Distill with Feedback, to further improve the performance of the Baize models with feedback from ChatGPT.The Baize models and data are released for research purposes only. 1 Canwen Xu, Daya Guo, Nan Duan 0001, Julian J. McAuley |
EMNLP | 1 |
| 2023 | LongCoder: A Long-Range Pre-trained Language Model for Code CompletionabstractIn this paper, we introduce a new task for code completion that focuses on handling long code input and propose a sparse Transformer model, called LongCoder, to address this task. LongCoder employs a sliding window mechanism for self-attention and introduces two types of globally accessible tokens - bridge tokens and memory tokens - to improve performance and efficiency. Bridge tokens are inserted throughout the input sequence to aggregate local information and facilitate global interaction, while memory tokens are included to highlight important statements that may be invoked later and need to be memorized, such as package imports and definitions of classes, functions, or structures. We conduct experiments on a newly constructed dataset that contains longer code context and the publicly available CodeXGLUE benchmark. Experimental results demonstrate that LongCoder achieves superior performance on code completion tasks compared to previous models while maintaining comparable efficiency in terms of computational resources during inference. Daya Guo, Canwen Xu, Nan Duan 0001, Jian Yin 0001, Julian J. McAuley |
ICML | 2 |
| 2022 | Leashing the Inner Demons: Self-Detoxification for Language ModelsabstractLanguage models (LMs) can reproduce (or amplify) toxic language seen during training, which poses a risk to their practical application. In this paper, we conduct extensive experiments to study this phenomenon. We analyze the impact of prompts, decoding strategies and training corpora on the output toxicity. Based on our findings, we propose a simple yet effective unsupervised method for language models to ``detoxify'' themselves without an additional large corpus or external discriminator. Compared to a supervised baseline, our proposed method shows better toxicity reduction with good generation quality in the generated content under multiple settings. Warning: some examples shown in the paper may contain uncensored offensive content. Canwen Xu, Zexue He, Zhankui He, Julian J. McAuley |
AAAI | 1 |
| 2022 | BERT Learns to Teach: Knowledge Distillation with Meta LearningabstractWe present Knowledge Distillation with Meta Learning (MetaDistil), a simple yet effective alternative to traditional knowledge distillation (KD) methods where the teacher model is fixed during training.We show the teacher network can learn to better transfer knowledge to the student network (i.e., learning to teach) with the feedback from the performance of the distilled student network in a meta learning framework.Moreover, we introduce a pilot update mechanism to improve the alignment between the inner-learner and meta-learner in meta learning algorithms that focus on an improved inner-learner.Experiments on various benchmarks show that MetaDistil can yield significant improvements compared with traditional KD algorithms and is less sensitive to the choice of different student capacity and hyperparameters, facilitating the use of KD on different tasks and models. 1 Wangchunshu Zhou, Canwen Xu, Julian J. McAuley |
ACL (1) | 2 |
| 2022 | InforMask: Unsupervised Informative Masking for Language Model PretrainingabstractMasked language modeling is widely used for pretraining large language models for natural language understanding (NLU).However, random masking is suboptimal, allocating an equal masking rate for all tokens.In this paper, we propose InforMask, a new unsupervised masking strategy for training masked language models.InforMask exploits Pointwise Mutual Information (PMI) to select the most informative tokens to mask.We further propose two optimizations for InforMask to improve its efficiency.With a one-off preprocessing step, InforMask outperforms random masking and previously proposed masking strategies on the factual recall benchmark LAMA and the question answering benchmark SQuAD v1 and v2. 1 Nafis Sadeq, Canwen Xu, Julian J. McAuley |
EMNLP | 2 |
| 2022 | Efficiently Tuned Parameters Are Task EmbeddingsabstractIntermediate-task transfer can benefit a wide range of NLP tasks with properly selected source datasets.However, it is computationally infeasible to experiment with all intermediate transfer combinations, making choosing a useful source task a challenging problem.In this paper, we anticipate that task-specific parameters updated in parameter-efficient tuning methods are likely to encode task-specific information.Therefore, such parameters can be predictive for inter-task transferability.Thus, we propose to exploit these efficiently tuned parameters as off-the-shelf task embeddings for the efficient selection of source datasets for intermediate-task transfer.We experiment with 11 text classification tasks and 11 question answering tasks.Experimental results show that our approach can consistently outperform existing inter-task transferability prediction methods while being conceptually simple and computationally efficient.Our analysis also reveals that the ability of efficiently tuned parameters on transferability prediction is disentangled with their in-task performance.This allows us to use parameters from early checkpoints as task embeddings to further improve efficiency. 1 Wangchunshu Zhou, Canwen Xu, Julian J. McAuley |
EMNLP | 2 |
| 2022 | Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim 0002, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Mike Tian-Jian Jiang, Matteo Manica, Sheng Shen 0001, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf 0008, Alexander M. Rush |
ICLR | 12 |
| 2022 | Automatic Multi-Label Prompting: Simple and Interpretable Few-Shot ClassificationabstractPrompt-based learning (i.e., prompting) is an emerging paradigm for exploiting knowledge learned by a pretrained language model.In this paper, we propose Automatic Multi-Label Prompting (AMuLaP), a simple yet effective method to automatically select label mappings for few-shot text classification with prompting.Our method exploits one-to-many label mappings and a statistics-based algorithm to select label mappings given a prompt template.Our experiments demonstrate that AMu-LaP achieves competitive performance on the GLUE benchmark without human effort or external resources.1 Canwen Xu, Julian J. McAuley |
NAACL-HLT | 2 |
| 2021 | Beyond Preserved Accuracy: Evaluating Loyalty and Robustness of BERT CompressionabstractRecent studies on compression of pretrained language models (e.g., BERT) usually use preserved accuracy as the metric for evaluation.In this paper, we propose two new metrics, label loyalty and probability loyalty that measure how closely a compressed model (i.e., student) mimics the original model (i.e., teacher).We also explore the effect of compression with regard to robustness under adversarial attacks.We benchmark quantization, pruning, knowledge distillation and progressive module replacing with loyalty and robustness.By combining multiple compression techniques, we provide a practical strategy to achieve better accuracy, loyalty and robustness. 1 Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Julian J. McAuley, Furu Wei |
EMNLP (1) | 1 |
| 2021 | Improving Sequence-to-Sequence Pre-training via Sequence Span RewritingabstractIn this paper, we propose Sequence Span Rewriting (SSR), a self-supervised task for sequence-to-sequence (Seq2Seq) pre-training.SSR learns to refine the machine-generated imperfect text spans into ground truth text.SSR provides more fine-grained and informative supervision in addition to the original textinfilling objective.Compared to the prevalent text infilling objectives for Seq2Seq pretraining, SSR is naturally more consistent with many downstream generation tasks that require sentence rewriting (e.g., text summarization, question generation, grammatical error correction, and paraphrase generation).We conduct extensive experiments by using SSR to improve the typical Seq2Seq pre-trained model T5 in a continual pre-training setting and show substantial improvements over T5 on various natural language generation tasks. 1 Wangchunshu Zhou, Tao Ge 0001, Canwen Xu, Ke Xu 0001, Furu Wei |
EMNLP (1) | 3 |
| 2021 | Blow the Dog Whistle: A Chinese Dataset for Cant Understanding with Common Sense and World KnowledgeabstractCanwen Xu, Wangchunshu Zhou, Tao Ge, Ke Xu, Julian McAuley, Furu Wei. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Julian J. McAuley, Furu Wei |
NAACL-HLT | 1 |
| 2020 | Pre-train and Plug-in: Flexible Conditional Text Generation with Variational Auto-EncodersabstractConditional Text Generation has drawn much attention as a topic of Natural Language Generation (NLG) which provides the possibility for humans to control the properties of generated contents.Current conditional generation models cannot handle emerging conditions due to their joint end-to-end learning fashion.When a new condition added, these techniques require full retraining.In this paper, we present a new framework named Pre-train and Plug-in Variational Auto-Encoder (PPVAE) towards flexible conditional text generation.PPVAE decouples the text generation module from the condition representation module to allow "one-to-many" conditional generation.When a fresh condition emerges, only a lightweight network needs to be trained and works as a plug-in for PPVAE, which is efficient and desirable for real-world applications.Extensive experiments demonstrate the superiority of PPVAE against the existing alternatives with better conditionality and diversity but less training effort. 1 Canwen Xu, Jiaxin Pei, Jialong Han |
ACL | 2 |
| 2020 | MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and SummarizationabstractRecently, large-scale datasets have vastly facilitated the development in nearly all domains of Natural Language Processing.However, there is currently no cross-task dataset in NLP, which hinders the development of multi-task learning.We propose MATINF, the first jointly labeled large-scale dataset for classification, question answering and summarization.MAT-INF contains 1.07 million question-answer pairs with human-labeled categories and usergenerated question descriptions.Based on such rich information, MATINF is applicable for three major NLP tasks, including classification, question answering, and summarization.We benchmark existing methods and a novel multi-task baseline over MATINF to inspire further research.Our comprehensive comparison and experiments over MATINF and other datasets demonstrate the merits held by MAT-INF. 1 Canwen Xu, Jiaxin Pei, Yiyu Liu |
ACL | 1 |
| 2020 | BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingabstractIn this paper, we propose a novel model compression approach to effectively compress BERT by progressive module replacing.Our approach first divides the original BERT into several modules and builds their compact substitutes.Then, we randomly replace the original modules with their substitutes to train the compact modules to mimic the behavior of the original modules.We progressively increase the probability of replacement through the training.In this way, our approach brings a deeper level of interaction between the original and compact models.Compared to the previous knowledge distillation approaches for BERT compression, our approach does not introduce any additional loss function.Our approach outperforms existing knowledge distillation approaches on GLUE benchmark, showing a new perspective of model compression.1 Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Furu Wei, Ming Zhou 0001 |
EMNLP (1) | 1 |
| 2020 | BERT Loses Patience: Fast and Robust Inference with Early ExitabstractIn this paper, we propose Patience-based Early Exit, a straightforward yet effective inference method that can be used as a plug-and-play technique to simultaneously improve the efficiency and robustness of a pretrained language model (PLM). To achieve this, our approach couples an internal-classifier with each layer of a PLM and dynamically stops inference when the intermediate predictions of the internal classifiers do not change for a pre-defined number of steps. Our approach improves inference efficiency as it allows the model to make a prediction with fewer layers. Meanwhile, experimental results with an ALBERT model show that our method can improve the accuracy and robustness of the model by preventing it from overthinking and exploiting multiple classifiers for prediction, yielding a better accuracy-speed trade-off compared to existing early exit methods. Wangchunshu Zhou, Canwen Xu, Tao Ge 0001, Julian J. McAuley, Ke Xu 0001, Furu Wei |
NeurIPS | 2 |
| 2019 | Exploiting Multiple Embeddings for Chinese Named Entity RecognitionabstractIdentifying the named entities mentioned in text would enrich many semantic applications at the downstream level. However, due to the predominant usage of colloquial language in microblogs, the named entity recognition (NER) in Chinese microblogs experience significant performance deterioration, compared with performing NER in formal Chinese corpus. In this paper, we propose a simple yet effective neural framework to derive the character-level embeddings for NER in Chinese text, named ME-CNER. A character embedding is derived with rich semantic information harnessed at multiple granularities, ranging from radical, character to word levels. The experimental results demonstrate that the proposed approach achieves a large performance improvement on Weibo dataset and comparable performance on MSRA news dataset with lower computational cost against the existing state-of-the-art alternatives. Canwen Xu, Jialong Han |
CIKM | 1 |
| 2019 | DLocRL: A Deep Learning Pipeline for Fine-Grained Location Recognition and Linking in TweetsabstractIn recent years, with the prevalence of social media and smart devices, people causally reveal their locations such as shops, hotels, and restaurants in their tweets. Recognizing and linking such fine-grained location mentions to well-defined location profiles are beneficial for retrieval and recommendation systems. In this paper, we propose DLocRL, a new deep learning pipeline for fine-grained location recognition and linking in tweets, and verify its effectiveness on a real-world Twitter dataset. Canwen Xu, Jing Li 0034, Xiangyang Luo 0001, Jiaxin Pei, Chenliang Li 0005, Donghong Ji |
WWW | 1 |