VLDB 2026 Research / reviewers in the wild / expert
Lei Shu 0004
dblp:19/2932-4
· DBLP profile ↗
22ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating Inference of Retrieval-Augmented Generation via Sparse Context SelectionabstractLarge language models (LLMs) augmented with retrieval exhibit robust performance and extensive versatility by incorporating external contexts. However, the input length grows linearly in the number of retrieved documents, causing a dramatic increase in latency. In this paper, we propose a novel paradigm named Sparse RAG, which seeks to cut computation costs through sparsity. Specifically, Sparse RAG encodes retrieved documents in parallel, which eliminates latency introduced by long-range attention of retrieved documents. Then, LLMs selectively decode the output by only attending to highly relevant caches auto-regressively, which are chosen via prompting LLMs with special control tokens. It is notable that Sparse RAG combines the assessment of each individual document and the generation of the response into a single process. The designed sparse mechanism in a RAG system can facilitate the reduction of the number of documents loaded during decoding for accelerating the inference of the RAG system. Additionally, filtering out undesirable contexts enhances the model’s focus on relevant context, inherently improving its generation quality. Evaluation results on four datasets show that Sparse RAG can be used to strike an optimal balance between generation quality and computational efficiency, demonstrating its generalizability across tasks. Jia-Chen Gu, Caitlin Sikora, Ho Ko, Yinxiao Liu, Chu-Cheng Lin, Lei Shu 0004, Liangchen Luo, Lei Meng 0008, Jindong Chen |
ICLR | 7 |
| 2024 | RewriteLM: An Instruction-Tuned Large Language Model for Text RewritingabstractLarge Language Models (LLMs) have demonstrated impressive capabilities in creative tasks such as storytelling and E-mail generation. However, as LLMs are primarily trained on final text results rather than intermediate revisions, it might be challenging for them to perform text rewriting tasks. Most studies in the rewriting tasks focus on a particular transformation type within the boundaries of single sentences. In this work, we develop new strategies for instruction tuning and reinforcement learning to better align LLMs for cross-sentence rewriting tasks using diverse wording and structures expressed through natural languages including 1) generating rewriting instruction data from Wiki edits and public corpus through instruction generation and chain-of-thought prompting; 2) collecting comparison data for reward model training through a new ranking function. To facilitate this research, we introduce OpenRewriteEval, a novel benchmark covers a wide variety of rewriting types expressed through natural language instructions. Our results show significant improvements over a variety of baselines. Lei Shu 0004, Liangchen Luo, Jayakumar Hoskere, Yinxiao Liu, Simon Tong, Jindong Chen, Lei Meng 0008 |
AAAI | 1 |
| 2024 | Enhancing Reinforcement Learning with Dense Rewards from Language Model CriticabstractReinforcement learning (RL) can align language models with non-differentiable reward signals, such as human preferences.However, a major challenge arises from the sparsity of these reward signals -typically, there is only a single reward for an entire output.This sparsity of rewards can lead to inefficient and unstable learning.To address this challenge, our paper introduces an novel framework that utilizes the critique capability of Large Language Models (LLMs) to produce intermediate-step rewards during RL training.Our approach pairs a policy model with a critic language model that provides feedback on each part of the policy's output.This feedback is then translated into token or span-level rewards that can be used to guide the RL training process.We investigate this approach under two different settings: one where the policy model is smaller and is paired with a more powerful critic model, and another where a single language model fulfills both roles.We assess our approach on three text generation tasks: sentiment control, language model detoxification, and summarization.Experimental results show that incorporating artificial intrinsic rewards significantly improve both sample efficiency and the overall performance of the policy model, supported by both automatic and human evaluation.The code is available under Google Research github * . Lei Shu 0004, Nevan Wichers, Yinxiao Liu, Lei Meng 0008 |
EMNLP | 2 |
| 2022 | Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIPabstractIn an out-of-distribution (OOD) detection problem, samples of known classes (also called in-distribution classes) are used to train a special classifier. In testing, the classifier can (1) classify the test samples of known classes to their respective classes and also (2) detect samples that do not belong to any of the known classes (i.e., they belong to some unknown or OOD classes). This paper studies the problem of zero-shot out-of-distribution (OOD) detection, which still performs the same two tasks in testing but has no training except using the given known class names. This paper proposes a novel and yet simple method (called ZOC) to solve the problem. ZOC builds on top of the recent advances in zero-shot classification through multi-modal representation learning. It first extends the pre-trained language-vision model CLIP by training a text-based image description generator on top of CLIP. In testing, it uses the extended model to generate candidate unknown class names for each test sample and computes a confidence score based on both the known class names and candidate unknown class names for zero-shot OOD detection. Experimental results on 5 benchmark datasets for OOD detection demonstrate that ZOC outperforms the baselines by a large margin. Sepideh Esmaeilpour, Bing Liu 0001, Eric Robertson 0001, Lei Shu 0004 |
AAAI | 4 |
| 2022 | Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue SystemabstractPre-trained language models have been recently shown to benefit task-oriented dialogue (TOD) systems.Despite their success, existing methods often formulate this task as a cascaded generation problem which can lead to error accumulation across different sub-tasks and greater data annotation overhead.In this study, we present PPTOD, a unified plug-andplay model for task-oriented dialogue.In addition, we introduce a new dialogue multi-task pre-training strategy that allows the model to learn the primary TOD task completion skills from heterogeneous dialog corpora.We extensively test our model on three benchmark TOD tasks, including end-to-end dialogue modelling, dialogue state tracking, and intent classification.Experimental results show that PPTOD achieves new state of the art on all evaluated tasks in both high-resource and lowresource scenarios.Furthermore, comparisons against previous SOTA methods show that the responses generated by PPTOD are more factually correct and semantically coherent as judged by human annotators. 1 Yixuan Su, Lei Shu 0004, Elman Mansimov, Arshit Gupta, Deng Cai 0002, Yi-An Lai |
ACL (1) | 2 |
| 2022 | Continual Training of Language Models for Few-Shot LearningabstractRecent work on applying large language models (LMs) achieves impressive performance in many NLP applications.Adapting or posttraining an LM using an unlabeled domain corpus can produce even better performance for end-tasks in the domain.This paper proposes the problem of continually extending an LM by incrementally post-train the LM with a sequence of unlabeled domain corpora to expand its knowledge without forgetting its previous skills.The goal is to improve the few-shot end-task learning in these domains.The resulting system is called CPT (Continual Post-Training), which to our knowledge, is the first continual post-training system.Experimental results verify its effectiveness. Zixuan Ke, Haowei Lin, Yijia Shao, Hu Xu 0001, Lei Shu 0004, Bing Liu 0001 |
EMNLP | 5 |
| 2022 | Adapting a Language Model While Preserving its General KnowledgeabstractDomain-adaptive pre-training (or DA-training for short), also known as post-training, aims to train a pre-trained general-purpose language model (LM) using an unlabeled corpus of a particular domain to adapt the LM so that endtasks in the domain can give improved performances.However, existing DA-training methods are in some sense blind as they do not explicitly identify what knowledge in the LM should be preserved and what should be changed by the domain corpus.This paper shows that the existing methods are suboptimal and proposes a novel method to perform a more informed adaptation of the knowledge in the LM by (1) soft-masking the attention heads based on their importance to best preserve the general knowledge in the LM and (2) contrasting the representations of the general and the full (both general and domain knowledge) to learn an integrated representation with both general and domain-specific knowledge.Experimental results will demonstrate the effectiveness of the proposed approach.1 Zixuan Ke, Yijia Shao, Haowei Lin, Hu Xu 0001, Lei Shu 0004, Bing Liu 0001 |
EMNLP | 5 |
| 2022 | Measuring and Reducing Model Update Regression in Structured Prediction for NLPabstractRecent advance in deep learning has led to rapid adoption of machine learning based NLP models in a wide range of applications. Despite the continuous gain in accuracy, backward compatibility is also an important aspect for industrial applications, yet it received little research attention. Backward compatibility requires that the new model does not regress on cases that were correctly handled by its predecessor. This work studies model update regression in structured prediction tasks. We choose syntactic dependency parsing and conversational semantic parsing as representative examples of structured prediction tasks in NLP. First, we measure and analyze model update regression in different model update settings. Next, we explore and benchmark existing techniques for reducing model update regression including model ensemble and knowledge distillation. We further propose a simple and effective method, Backward-Congruent Re-ranking (BCR), by taking into account the characteristics of structured output. Experiments show that BCR can better mitigate model update regression than model ensemble and knowledge distillation approaches. Deng Cai 0002, Elman Mansimov, Yi-An Lai, Yixuan Su, Lei Shu 0004 |
NeurIPS | 5 |
| 2021 | CLASSIC: Continual and Contrastive Learning of Aspect Sentiment Classification TasksabstractThis paper studies continual learning (CL) of a sequence of aspect sentiment classification (ASC) tasks in a particular CL setting called domain incremental learning (DIL). Each task is from a different domain or product. The DIL setting is particularly suited to ASC because in testing the system needs not know the task/domain to which the test data belongs. To our knowledge, this setting has not been studied before for ASC. This paper proposes a novel model called CLASSIC. The key novelty is a contrastive continual learning method that enables both knowledge transfer across tasks and knowledge distillation from old tasks to the new task, which eliminates the need for task ids in testing. Experimental results show the high effectiveness of CLASSIC. Zixuan Ke, Bing Liu 0001, Hu Xu 0001, Lei Shu 0004 |
EMNLP (1) | 4 |
| 2021 | Achieving Forgetting Prevention and Knowledge Transfer in Continual LearningabstractContinual learning (CL) learns a sequence of tasks incrementally with the goal of achieving two main objectives: overcoming catastrophic forgetting (CF) and encouraging knowledge transfer (KT) across tasks. However, most existing techniques focus only on overcoming CF and have no mechanism to encourage KT, and thus do not do well in KT. Although several papers have tried to deal with both CF and KT, our experiments show that they suffer from serious CF when the tasks do not have much shared knowledge. Another observation is that most current CL methods do not use pre-trained models, but it has been shown that such models can significantly improve the end task performance. For example, in natural language processing, fine-tuning a BERT-like pre-trained language model is one of the most effective approaches. However, for CL, this approach suffers from serious CF. An interesting question is how to make the best use of pre-trained models for CL. This paper proposes a novel model called CTR to solve these problems. Our experimental results demonstrate the effectiveness of CTR Zixuan Ke, Bing Liu 0001, Nianzu Ma, Hu Xu 0001, Lei Shu 0004 |
NeurIPS | 5 |
| 2020 | Understanding Pre-trained BERT for Aspect-based Sentiment AnalysisabstractThis paper analyzes the pre-trained hidden representations learned from reviews on BERT for tasks in aspect-based sentiment analysis (ABSA).Our work is motivated by the recent progress in BERT-based language models for ABSA.However, it is not clear how the general proxy task of (masked) language model trained on unlabeled corpus without annotations of aspects or opinions can provide important features for downstream tasks in ABSA.By leveraging the annotated datasets in ABSA, we investigate both the attentions and the learned representations of BERT pre-trained on reviews.We found that BERT uses very few self-attention heads to encode context words (such as prepositions or pronouns that indicating an aspect) and opinion words for an aspect.Most features in the representation of an aspect are dedicated to the finegrained semantics of the domain (or product category) and the aspect itself, instead of carrying summarized opinions from its context.We hope this investigation can help future research in improving self-supervised learning, unsupervised learning and fine-tuning for ABSA. 1 Hu Xu 0001, Lei Shu 0004, Philip S. Yu, Bing Liu 0001 |
COLING | 2 |
| 2020 | Continual Learning with Knowledge Transfer for Sentiment Classification
Zixuan Ke, Bing Liu 0001, Hao Wang 0008, Lei Shu 0004 |
ECML/PKDD (3) | 4 |
| 2019 | Modeling Multi-Action Policy for Task-Oriented DialoguesabstractLei Shu, Hu Xu, Bing Liu, Piero Molino. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lei Shu 0004, Hu Xu 0001, Bing Liu 0001, Piero Molino |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Flexibly-Structured Model for Task-Oriented DialoguesabstractThis paper proposes a novel end-to-end architecture for task-oriented dialogue systems.It is based on a simple and practical yet very effective sequence-to-sequence approach, where language understanding and state tracking tasks are modeled jointly with a structured copy-augmented sequential decoder and a multi-label decoder for each slot.The policy engine and language generation tasks are modeled jointly following that.The copyaugmented sequential decoder deals with new or unknown values in the conversation, while the multi-label decoder combined with the sequential decoder ensures the explicit assignment of values to slots.On the generation part, slot binary classifiers are used to improve performance.This architecture is scalable to real-world scenarios and is shown through an empirical evaluation to achieve state-of-the-art performance on both the Cambridge Restaurant dataset and the Stanford in-car assistant dataset 1 . Lei Shu 0004, Piero Molino, Mahdi Namazifar, Hu Xu 0001, Bing Liu 0001, Huaixiu Zheng, Gökhan Tür |
SIGdial | 1 |
| 2019 | Open-world Learning and Application to Product ClassificationabstractClassic supervised learning makes the closed-world assumption that the classes seen in testing must have appeared in training. However, this assumption is often violated in real-world applications. For example, in a social media site, new topics emerge constantly and in e-commerce, new categories of products appear daily. A model that cannot detect new/unseen topics or products is hard to function well in such open environments. A desirable model working in such environments must be able to (1) reject examples from unseen classes (not appeared in training) and (2) incrementally learn the new/unseen classes to expand the existing model. This is called open-world learning (OWL). This paper proposes a new OWL method based on meta-learning. The key novelty is that the model maintains only a dynamic set of seen classes that allows new classes to be added or deleted with no need for model re-training. Each class is represented by a small set of training examples. In testing, the meta-classifier only uses the examples of the maintained seen classes (including the newly added classes) on-the-fly for classification and rejection. Experimental results with e-commerce product classification show that the proposed method is highly effective1. Hu Xu 0001, Bing Liu 0001, Lei Shu 0004, Philip S. Yu |
WWW | 3 |
| 2018 | Dual Attention Network for Product Compatibility and Function Satisfiability AnalysisabstractProduct compatibility and functionality are of utmost importance to customers when they purchase products, and to sellers and manufacturers when they sell products. Due to the huge number of products available online, it is infeasible to enumerate and test the compatibility and functionality of every product. In this paper, we address two closely related problems: product compatibility analysis and function satisfiability analysis, where the second problem is a generalization of the first problem (e.g., whether a product works with another product can be considered as a special function). We first identify a novel question and answering corpus that is up-to-date regarding product compatibility and functionality information. To allow automatic discovery product compatibility and functionality, we then propose a deep learning model called Dual Attention Network (DAN). Given a QA pair for a to-be-purchased product, DAN learns to 1) discover complementary products (or functions), and 2) accurately predict the actual compatibility (or satisfiability) of the discovered products (or functions). The challenges addressed by the model include the briefness of QAs, linguistic patterns indicating compatibility, and the appropriate fusion of questions and answers. We conduct experiments to quantitatively and qualitatively show that the identified products and functions have both high coverage and accuracy, compared with a wide spectrum of baselines. Hu Xu 0001, Sihong Xie, Lei Shu 0004, Philip S. Yu |
AAAI | 3 |
| 2018 | Lifelong Domain Word Embedding via Meta-LearningabstractLearning high-quality domain word embeddings is important for achieving good performance in many NLP tasks. General-purpose embeddings trained on large-scale corpora are often sub-optimal for domain-specific applications. However, domain-specific tasks often do not have large in-domain corpora for training high-quality domain embeddings. In this paper, we propose a novel lifelong learning setting for domain embedding. That is, when performing the new domain embedding, the system has seen many past domains, and it tries to expand the new in-domain corpus by exploiting the corpora from the past domains via meta-learning. The proposed meta-learner characterizes the similarities of the contexts of the same word in many domain corpora, which helps retrieve relevant data from the past domains to expand the new domain corpus. Experimental results show that domain embeddings produced from such a process improve the performance of the downstream tasks. Hu Xu 0001, Bing Liu 0001, Lei Shu 0004, Philip S. Yu |
IJCAI | 3 |
| 2017 | Product function need recognition via semi-supervised attention networkabstractFunctionality is of utmost importance to customers when they purchase products. However, it is unclear to customers whether a product can really satisfy their needs on functions. Further, missing functions may be intentionally hidden by the manufacturers or the sellers. As a result, a customer needs to spend a fair amount of time before purchasing or just purchase the product on his/her own risk. In this paper, we first identify a novel QA corpus that is dense on product functionality information1. We then design a neural network called Semi-supervised Attention Network (SAN) to discover product functions from questions. This model leverages unlabeled data as contextual information to perform semi-supervised sequence labeling. We conduct experiments to show that the extracted function have both high coverage and accuracy, compared with a wide spectrum of baselines. Hu Xu 0001, Sihong Xie, Lei Shu 0004, Philip S. Yu |
IEEE BigData | 3 |
| 2017 | DOC: Deep Open Classification of Text DocumentsabstractTraditional supervised learning makes the closed-world assumption that the classes appeared in the test data must have appeared in training.This also applies to text learning or text classification.As learning is used increasingly in dynamic open environments where some new/test documents may not belong to any of the training classes, identifying these novel documents during classification presents an important problem.This problem is called openworld classification or open classification.This paper proposes a novel deep learning based approach.It outperforms existing state-of-the-art techniques dramatically. Lei Shu 0004, Hu Xu 0001, Bing Liu 0001 |
EMNLP | 1 |
| 2016 | CER: Complementary entity recognition via knowledge expansion on large unlabeled product reviewsabstractProduct reviews contain a lot of useful information about product features and customer opinions. One important product feature is the complementary entity (products) that may potentially work together with the reviewed product. Knowing complementary entities of the reviewed product is very important because customers want to buy compatible products and avoid incompatible ones. In this paper, we address the problem of Complementary Entity Recognition (CER). Since no existing method can solve this problem, we first propose a novel unsupervised method to utilize syntactic dependency paths to recognize complementary entities. Then we expand category-level domain knowledge about complementary entities using only a few general seed verbs on a large amount of unlabeled reviews. The domain knowledge helps the unsupervised method to adapt to different products and greatly improves the precision of the CER task. The advantage of the proposed method is that it does not require any labeled data for training. We conducted experiments on 7 popular products with about 1200 reviews in total to demonstrate that the proposed approach is effective. Hu Xu 0001, Sihong Xie, Lei Shu 0004, Philip S. Yu |
IEEE BigData | 3 |
| 2016 | Lifelong-RL: Lifelong Relaxation Labeling for Separating Entities and Aspects in Opinion Targetsabstract. Extensive experiments show that the proposed algorithm Lifelong-RL outperforms baseline methods markedly. Lei Shu 0004, Bing Liu 0001, Hu Xu 0001, Annice Kim |
EMNLP | 1 |
| 2013 | Planning Paths with Fewer Turns on Grid MapsabstractIn this paper, we consider the problem of planning any-angle paths with small numbers of turns on grid maps. We propose a novel heuristic search algorithm called Link* that returns paths containing fewer turns at the cost of slightly longer path lengths. Experimental results demonstrate that Link* can produce paths with fewer turns than other any-angle path planning algorithms while still maintaining comparable path lengths. Because it produces this type of path, artificial agents can take advantage of Link* when the cost of turns is expensive. Hu Xu 0001, Lei Shu 0004, May Huang |
SOCS | 2 |