VLDB 2026 Research / reviewers in the wild / expert
Lei Shen 0001
dblp:63/2022-1
· DBLP profile ↗
22ranked-venue papers
6as first author
16since 2021 · last 2024
0000-0003-4782-3779ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MuJo-SF: Multimodal Joint Slot Filling for Attribute Value Prediction of E-Commerce CommoditiesabstractSupplementing product attribute information is a critical step for E-commerce platforms, which further benefits various downstream tasks, including product recommendation, product search, and product knowledge graph construction. Intuitively, the visual information available on e-commerce platforms can effectively function as a primary source for certain product attributes. However, existing works either extract attribute values solely from textual product descriptions or leverage limited visual information (e.g., image features or optical character recognition tokens) to assist extraction, without mining the fine-grained visual cues linked with the products effectively. In this paper, we propose a novel task -Multimodal Joint Slot Filling(MuJo-SF) - that aims to combine multimodal information from both product descriptions and their corresponding product images to jointly fill values into the pre-defined product attribute set. To this end, we develop MAVP, a new dataset with 79 k instances of product description-image pairs. Specifically, we present a strategy to fulfill visualized saliency ascription, which aims to distinguish between text-dependent and image-dependent attributes. For those image-dependent attributes, we annotate the corresponding values from images using distant supervision. Then, we design a model for MuJo-SF, which combines multimodal representations and fills image-dependent and text-dependent attributes separately. Finally, we conduct extensive experiments on MAVP and provide rich results for MuJo-SF, which can be used as baselines to facilitate future research. Meihuizi Jia, Lei Shen 0001, Anh Tuan Luu, Meng Chen 0006, Lejian Liao, Shaozu Yuan, Xiaodong He 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query GroundingabstractMultimodal named entity recognition (MNER) is a critical step in information extraction, which aims to detect entity spans and classify them to corresponding entity types given a sentence-image pair. Existing methods either (1) obtain named entities with coarse-grained visual clues from attention mechanisms, or (2) first detect fine-grained visual regions with toolkits and then recognize named entities. However, they suffer from improper alignment between entity types and visual regions or error propagation in the two-stage manner, which finally imports irrelevant visual information into texts. In this paper, we propose a novel end-to-end framework named MNER-QG that can simultaneously perform MRC-based multimodal named entity recognition and query grounding. Specifically, with the assistance of queries, MNER-QG can provide prior knowledge of entity types and visual regions, and further enhance representations of both text and image. To conduct the query grounding task, we provide manual annotations and weak supervisions that are obtained via training a highly flexible visual grounding model with transfer learning. We conduct extensive experiments on two public MNER datasets, Twitter2015 and Twitter2017. Experimental results show that MNER-QG outperforms the current state-of-the-art models on the MNER task, and also improves the query grounding performance. Meihuizi Jia, Lei Shen 0001, Lejian Liao, Meng Chen 0006, Xiaodong He 0001 |
AAAI | 2 |
| 2023 | DiffusEmp: A Diffusion Model-Based Framework with Multi-Grained Control for Empathetic Response GenerationabstractEmpathy is a crucial factor in open-domain conversations, which naturally shows one's caring and understanding to others.Though several methods have been proposed to generate empathetic responses, existing works often lead to monotonous empathy that refers to generic and safe expressions.In this paper, we propose to use explicit control to guide the empathy expression and design a framework DIFFUSEMP based on conditional diffusion language model to unify the utilization of dialogue context and attribute-oriented control signals.Specifically, communication mechanism, intent, and semantic frame are imported as multi-grained signals that control the empathy realization from coarse to fine levels.We then design a specific masking strategy to reflect the relationship between multi-grained signals and response tokens, and integrate it into the diffusion model to influence the generative process.Experimental results on a benchmark dataset EMPA-THETICDIALOGUE show that our framework outperforms competitive baselines in terms of controllability, informativeness, and diversity without the loss of context-relatedness. Guanqun Bi, Lei Shen 0001, Yanan Cao 0001, Meng Chen 0006, Yuqiang Xie, Zheng Lin 0001, Xiaodong He 0001 |
ACL (1) | 2 |
| 2023 | Tackling Modality Heterogeneity with Multi-View Calibration Network for Multimodal Sentiment DetectionabstractYiwei Wei, Shaozu Yuan, Ruosong Yang, Lei Shen, Zhangmeizhi Li, Longbiao Wang, Meng Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shaozu Yuan, Ruosong Yang, Lei Shen 0001, Zhangmeizhi Li, Longbiao Wang, Meng Chen 0006 |
ACL (1) | 4 |
| 2023 | Dialog-Post: Multi-Level Self-Supervised Objectives and Hierarchical Model for Dialogue Post-TrainingabstractDialogue representation and understanding aim to convert conversational inputs into embeddings and fulfill discriminative tasks.Compared with free-form text, dialogue has two important characteristics, hierarchical semantic structure and multi-facet attributes.Therefore, directly applying the pretrained language models (PLMs) might result in unsatisfactory performance.Recently, several work focused on the dialogue-adaptive post-training (Dial-Post) that further trains PLMs to fit dialogues.To model dialogues more comprehensively, we propose a DialPost method, DIALOG-POST, with multi-level self-supervised objectives and a hierarchical model.These objectives leverage dialogue-specific attributes and use selfsupervised signals to fully facilitate the representation and understanding of dialogues.The novel model is a hierarchical segment-wise self-attention network, which contains innersegment and inter-segment self-attention sublayers followed by an aggregation and updating module.To evaluate the effectiveness of our methods, we first apply two public datasets for the verification of representation ability.Then we conduct experiments on a newly-labelled dataset that is annotated with 4 dialogue understanding tasks.Experimental results show that our method outperforms existing SOTA models and achieves a 3.3% improvement on average. Token-level SSOs Utterance-level SSO Dialogue-level SSOs𝑄: 这个手机支持5G吗?𝑄: Zhenyu Zhang 0029, Lei Shen 0001, Meng Chen 0006, Xiaodong He 0001 |
ACL (1) | 2 |
| 2023 | POSPAN: Position-Constrained Span Masking for Language Model Pre-trainingabstractSpan-level masked language modeling (MLM) has shown to be advantageous to pre-trained language models over the original single-token MLM, as entities/phrases and their dependencies are critical to language understanding. Previous works only consider span length with some discrete distributions, while the dependencies among spans are ignored, i.e., assuming that the positions of masked spans are uniformly distributed. In this paper, we present POSPAN, a general framework to allow diverse position-constrained span masking strategies via the combination of span length distribution and position constraint distribution, which unifies all existing span-level masking methods. To verify the effectiveness of POSPAN in pre-training, we evaluate it on the datasets from several NLU benchmarks. Experimental results indicate that the position constraint is capable of enhancing span-level masking broadly, and our best POSPAN setting consistently outperforms its span-length-only counterparts and vanilla MLM. We also conduct theoretical analysis for the position constraint in masked language models to shed light on the reason why POSPAN works well, demonstrating the rationality and necessity of POSPAN. Zhenyu Zhang 0029, Lei Shen 0001, Meng Chen 0006, Xiaodong He 0001 |
CIKM | 2 |
| 2023 | MPP-net: Multi-perspective perception network for dense video captioning
Shaozu Yuan, Meng Chen 0006, Longbiao Wang, Lei Shen 0001, Zhiling Yan |
Neurocomputing | 6 |
| 2022 | Few-Shot Table Understanding: A Benchmark Dataset and Pre-Training BaselineabstractFew-shot table understanding is a critical and challenging problem in real-world scenario as annotations over large amount of tables are usually costly. Pre-trained language models (PLMs), which have recently flourished on tabular data, have demonstrated their effectiveness for table understanding tasks. However, few-shot table understanding is rarely explored due to the deficiency of public table pre-training corpus and well-defined downstream benchmark tasks, especially in Chinese. In this paper, we establish a benchmark dataset, FewTUD, which consists of 5 different tasks with human annotations to systematically explore the few-shot table understanding in depth. Since there is no large number of public Chinese tables, we also collect a large-scale, multi-domain tabular corpus to facilitate future Chinese table pre-training, which includes one million tables and related natural language text with auxiliary supervised interaction signals. Finally, we present FewTPT, a novel table PLM with rich interactions over tabular data, and evaluate its performance comprehensively on the benchmark. Our dataset and model will be released to the public soon. Ruixue Liu, Shaozu Yuan, Aijun Dai, Lei Shen 0001, Tiangang Zhu, Meng Chen 0006, Xiaodong He 0001 |
COLING | 4 |
| 2022 | Query Prior Matters: A MRC Framework for Multimodal Named Entity RecognitionabstractMultimodal named entity recognition (MNER) is a vision-language task where the system is required to detect entity spans and corresponding entity types given a sentence-image pair. Existing methods capture text-image relations with various attention mechanisms that only obtain implicit alignments between entity types and image regions. To locate regions more accurately and better model cross-/within-modal relations, we propose a machine reading comprehension based framework for MNER, namely MRC-MNER. By utilizing queries in MRC, our framework can provide prior information about entity types and image regions. Specifically, we design two stages, Query-Guided Visual Grounding and Multi-Level Modal Interaction, to align fine-grained type-region information and simulate text-image/inner-text interactions respectively. For the former, we train a visual grounding model via transfer learning to extract region candidates that can be further integrated into the second stage to enhance token representations. For the latter, we design text-image and inner-text interaction modules along with three sub-tasks for MRC-MNER. To verify the effectiveness of our model, we conduct extensive experiments on two public MNER datasets, Twitter2015 and Twitter2017. Experimental results show that MRC-MNER outperforms the current state-of-the-art models on Twitter2017, and yields competitive results on Twitter2015. Meihuizi Jia, Lei Shen 0001, Jinhui Pang, Lejian Liao, Yang Song 0008, Meng Chen 0006, Xiaodong He 0001 |
ACM Multimedia | 3 |
| 2022 | Modeling semantic and emotional relationship in multi-turn emotional conversations using multi-task learning
Fuwei Cui, Hui Di, Lei Shen 0001, Kazushige Ouchi, Jin An Xu |
Appl. Intell. | 3 |
| 2021 | Probing Product Description Generation via Posterior DistillationabstractIn product description generation (PDG), the user-cared aspect is critical for the recommendation system, which can not only improve user's experiences but also obtain more clicks. High-quality customer reviews can be considered as an ideal source to mine user-cared aspects. However, in reality, a large number of new products (known as long-tailed commodities) cannot gather sufficient amount of customer reviews, which brings a big challenge in the product description generation task. Existing works tend to generate the product description solely based on item information, i.e., product attributes or title words, which leads to tedious contents and cannot attract customers effectively. To tackle this problem, we propose an adaptive posterior network based on Transformer architecture that can utilize user-cared information from customer reviews. Specifically, we first extend the self-attentive Transformer encoder to encode product titles and attributes. Then, we apply an adaptive posterior distillation module to utilize useful review information, which integrates user-cared aspects to the generation process. Finally, we apply a Transformer-based decoding phase with copy mechanism to automatically generate the product description. Besides, we also collect a large-scare Chinese product description dataset to support our work and further research in this field. Experimental results show that our model is superior to traditional generative models in both automatic indicators and human evaluation. Haolan Zhan, Hainan Zhang 0001, Hongshen Chen, Lei Shen 0001, Zhuoye Ding, Yongjun Bao, Weipeng Yan, Yanyan Lan |
AAAI | 4 |
| 2021 | GTM: A Generative Triple-wise Model for Conversational Question GenerationabstractLei Shen, Fandong Meng, Jinchao Zhang, Yang Feng, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Lei Shen 0001, Fandong Meng, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016 |
ACL/IJCNLP (1) | 1 |
| 2021 | Identifying Untrustworthy Samples: Data Filtering for Open-domain Dialogues with Bayesian OptimizationabstractBeing able to reply with a related, fluent, and informative response is an indispensable requirement for building high-quality conversational agents. In order to generate better responses, some approaches have been proposed, such as feeding extra information by collecting large-scale datasets with human annotations, designing neural conversational models (NCMs) with complex architecture and loss functions, or filtering out untrustworthy samples based on a dialogue attribute, e.g., Relatedness or Genericness. In this paper, we follow the third research branch and present a data filtering method for open-domain dialogues, which identifies untrustworthy samples from training data with a quality measure that linearly combines seven dialogue attributes. The attribute weights are obtained via Bayesian Optimization (BayesOpt) that aims to optimize an objective function for dialogue generation iteratively on the validation set. Then we score training samples with the quality measure, sort them in descending order, and filter out those at the bottom. Furthermore, to accelerate the "filter-train-evaluate'' iterations involved in BayesOpt on large-scale datasets, we propose a training framework that integrates maximum likelihood estimation (MLE) and negative training method (NEG). The training method updates parameters of a trained NCMs on two small sets with newly maintained and removed samples, respectively. Specifically, MLE is applied to maximize the log-likelihood of newly maintained samples, while NEG is used to minimize the log-likelihood of newly removed ones. Experimental results on two datasets show that our method can effectively identify untrustworthy samples, and NCMs trained on the filtered datasets achieve better performance. Lei Shen 0001, Haolan Zhan, Hongshen Chen, Xiaodan Zhu 0001 |
CIKM | 1 |
| 2021 | CoLV: A Collaborative Latent Variable Model for Knowledge-Grounded Dialogue GenerationabstractKnowledge-grounded dialogue generation has achieved promising performance with the engagement of external knowledge sources.Typical approaches towards this task usually perform relatively independent two sub-tasks, i.e., knowledge selection and knowledge-aware response generation.In this paper, in order to improve the diversity of both knowledge selection and knowledge-aware response generation, we propose a collaborative latent variable (CoLV) model to integrate these two aspects simultaneously in separate yet collaborative latent spaces, so as to capture the inherent correlation between knowledge selection and response generation.During generation, our proposed model firstly draws knowledge candidate from the latent space conditioned on the dialogue context, and then samples a response from another collaborative latent space conditioned on both the context and the selected knowledge.Experimental results on two widely-used knowledge-grounded dialogue datasets show that our model outperforms previous methods on both knowledge selection and response generation. Haolan Zhan, Lei Shen 0001, Hongshen Chen, Hainan Zhang 0001 |
EMNLP (1) | 2 |
| 2021 | Learning to Select Context in a Hierarchical and Global Perspective for Open-Domain Dialogue GenerationabstractOpen-domain multi-turn conversations mainly have three features, which are hierarchical semantic structure, redundant information, and long-term dependency. Grounded on these, selecting relevant context becomes a challenge step for multiturn dialogue generation. However, existing methods cannot differentiate both useful words and utterances in long distances from a response. Besides, previous work just performs context selection based on a state in the decoder, which lacks a global guidance and could lead some focuses on irrelevant or unnecessary information. In this paper, we propose a novel model with hierarchical self-attention mechanism and distant supervision to not only detect relevant words and utterances in short and long distances, but also discern related information globally when decoding. Experimental results on two public datasets of both automatic and human evaluations show that our model significantly outperforms other baselines in terms of fluency, coherence, and informativeness. Lei Shen 0001, Haolan Zhan, Yang Feng 0004 |
ICASSP | 1 |
| 2021 | Text is NOT Enough: Integrating Visual Impressions into Open-domain Dialogue GenerationabstractOpen-domain dialogue generation in natural language processing (NLP) is by default a pure-language task, which aims to satisfy human need for daily communication on open-ended topics by producing related and informative responses. In this paper, we point out that hidden images, named as visual impressions (VIs), can be explored from the text-only data to enhance dialogue understanding and help generate better responses. Besides, the semantic dependency between an dialogue post and its response is complicated, e.g., few word alignments and some topic transitions. Therefore, the visual impressions of them are not shared, and it is more reasonable to integrate the response visual impressions (RVIs) into the decoder, rather than the post visual impressions (PVIs). However, both the response and its RVIs are not given directly in the test process. To handle the above issues, we propose a framework to explicitly construct VIs based on pure-language dialogue datasets and utilize them for better dialogue understanding and generation. Specifically, we obtain a group of images (PVIs) for each post based on a pre-trained word-image mapping model. These PVIs are used in a co-attention encoder to get a post representation with both visual and textual information. Since the RVIs are not provided during testing, we design a cascade decoder that consists of two sub-decoders. The first sub-decoder predicts the content words in response, and applies the word-image mapping model to get corresponding RVIs. Then, the second sub-decoder generates the response based on the post and RVIs. Experimental results on two open-domain dialogue datasets show that our proposed approach achieves superior performance over competitive baselines in terms of fluency, relatedness, and diversity. Lei Shen 0001, Haolan Zhan, Yonghao Song |
ACM Multimedia | 1 |
| 2020 | CDL: Curriculum Dual Learning for Emotion-Controllable Response GenerationabstractEmotion-controllable response generation is an attractive and valuable task that aims to make open-domain conversations more empathetic and engaging. Existing methods mainly enhance the emotion expression by adding regularization terms to standard cross-entropy loss and thus influence the training process. However, due to the lack of further consideration of content consistency, the common problem of response generation tasks, safe response, is intensified. Besides, query emotions that can help model the relationship between query and response are simply ignored in previous models, which would further hurt the coherence. To alleviate these problems, we propose a novel framework named Curriculum Dual Learning (CDL) which extends the emotion-controllable response generation to a dual task to generate emotional responses and emotional queries alternatively. CDL utilizes two rewards focusing on emotion and content to improve the duality. Additionally, it applies curriculum learning to gradually generate high-quality responses based on the difficulties of expressing various emotions. Experimental results show that CDL significantly outperforms the baselines in terms of coherence, diversity, and relation to emotion factors. Lei Shen 0001, Yang Feng 0004 |
ACL | 1 |
| 2020 | Compose Like Humans: Jointly Improving the Coherence and Novelty for Modern Chinese Poetry GenerationabstractChinese poetry is an important part of worldwide culture, and classical and modern sub-branches are quite different. The former is a unique genre and has strict constraints, while the latter is very flexible in length, optional to have rhymes, and similar to modern poetry in other languages. Thus, it requires more to control the coherence and improve the novelty. In this paper, we propose a generate-retrieve-then-refine paradigm to jointly improve the coherence and novelty. In the first stage, a draft is generated given keywords (i.e., topics) only. The second stage produces a "refining vector" from retrieval lines. At last, we take into consideration both the draft and the "refining vector" to generate a new poem. The draft provides future sentence-level information for a line to be generated. Meanwhile, the "refining vector" points out the direction of refinement based on impressive words detection mechanism which can learn good patterns from references and then create new ones via insertion operation. Experimental results on a collected large-scale modern Chinese poetry dataset show that our proposed approach can not only generate more coherent poems, but also improve the diversity and novelty. Lei Shen 0001, Meng Chen 0006 |
IJCNN | 1 |
| 2020 | The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer ServiceabstractHuman conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge amount of real conversation data. In this paper, we construct a large-scale real scenario Chinese E-commerce conversation corpus, JDDC, with more than 1 million multi-turn dialogues, 20 million utterances, and 150 million words. The dataset reflects several characteristics of human-human conversations, e.g., goal-driven, and long-term dependency among the context. It also covers various dialogue types including task-oriented, chitchat and question-answering. Extra intent information and three well-annotated challenge sets are also provided. Then, we evaluate several retrieval-based and generative models to provide basic benchmark performance on the JDDC corpus. And we hope JDDC can serve as an effective testbed and benefit the development of fundamental research in dialogue task. Meng Chen 0006, Ruixue Liu, Lei Shen 0001, Shaozu Yuan, Jingyan Zhou, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
LREC | 3 |
| 2020 | Enhancing Multi-turn Dialogue Modeling with Intent Information for E-Commerce Customer Service
Ruixue Liu, Meng Chen 0006, Hang Liu 0005, Lei Shen 0001, Yang Song 0008, Xiaodong He 0001 |
NLPCC (1) | 4 |
| 2020 | User-Inspired Posterior Network for Recommendation Reason GenerationabstractRecommendation reason generation, aiming at showing the selling points of products for customers, plays a vital role in attracting customers' attention as well as improving user experience. A simple and effective way is to extract keywords directly from the knowledge-base of products, i.e., attributes or title, as the recommendation reason. However, generating recommendation reason from product knowledge doesn't naturally respond to users' interests. Fortunately, on some E-commerce websites, there exists more and more user-generated content (user-content for short), i.e., product question-answering (QA) discussions, which reflect user-cared aspects. Therefore, in this paper, we consider generating the recommendation reason by taking into account not only the product attributes but also the customer-generated product QA discussions. In reality, adequate user-content is only possible for the most popular commodities, whereas large sums of long-tail products or new products cannot gather a sufficient number of user-content. To tackle this problem, we propose a user-inspired multi-source posterior transformer (MSPT), which induces the model reflecting the users' interests with a posterior multiple QA discussions module, and generating recommendation reasons containing the product attributes as well as the user-cared aspects. Experimental results show that our model is superior to traditional generative models. Additionally, the analysis also shows that our model can focus more on the user-cared aspects than baselines. Haolan Zhan, Hainan Zhang 0001, Hongshen Chen, Lei Shen 0001, Yanyan Lan, Zhuoye Ding, Dawei Yin 0001 |
SIGIR | 4 |
| 2018 | Speeding Up Neural Machine Translation Decoding by Cube PruningabstractAlthough neural machine translation has achieved promising results, it suffers from slow translation speed.The direct consequence is that a trade-off has to be made between translation quality and speed, thus its performance can not come into full play.We apply cube pruning, a popular technique to speed up dynamic programming, into neural machine translation to speed up the translation.To construct the equivalence class, similar target hidden states are combined, leading to less RNN expansion operations on the target side and less softmax operations over the large target vocabulary.The experiments show that, at the same or even better translation quality, our method can translate faster compared with naive beam search by 3.3× on GPUs and 3.5× on CPUs. Wen Zhang 0009, Liang Huang 0001, Yang Feng 0004, Lei Shen 0001, Qun Liu 0001 |
EMNLP | 4 |