EDBT 2026 Demo / reviewers in the wild / expert
Meng Chen 0006
dblp:25/687-6
· DBLP profile ↗
41ranked-venue papers
1as first author
33since 2021 · last 2025
0009-0006-9908-4524ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 1 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 22 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AutoMV: An Autonomous Agent Framework for Real Estate Marketing Video GenerationabstractIn this paper, we introduce AutoMV, an autonomous agent framework designed for generating real estate marketing videos. The framework integrates a diverse set of existing models into a tool library, allowing the agent to intelligently select and execute the appropriate tools. Given property images and text, the agent decomposes the task into manageable subtasks, generating storyline directives and corresponding camera movement trajectories to guide the video production process. By automatically applying video synthesis techniques and incorporating multimedia elements such as subtitles and background music, the agent transforms static real estate images into dynamic, visually appealing videos, thereby optimizing their impact for digital marketing purposes. Kuizong Wu, Shaozu Yuan, Chang Shen, Meng Chen 0006 |
AAAI | 5 |
| 2025 | DeepMSD: Advancing Multimodal Sarcasm Detection Through Knowledge-Augmented Graph ReasoningabstractMultimodal sarcasm detection (MSD) requires predicting the sarcastic sentiment by understanding diverse modalities of data (e.g., text, image). Beyond the surface-level information conveyed in the post data, understanding the underlying deep-level knowledge-such as the background and intent behind the data-is crucial for understanding the sarcastic sentiment. However, previous works have often overlooked this aspect, limiting their potential to achieve superior performance. To tackle this challenge, we propose DeepMSD, a novel framework that generates supplemental deep-level knowledge to enhance the understanding of sarcastic content. Specifically, we first devise a Deep-level Knowledge Extraction Module that leverages large vision-language models to generate deep-level information behind the text-image pairs. Additionally, we devise a Cross-knowledge Graph Reasoning Module to model how humans use prior knowledge to identify sarcastic cues in multimodal posts. This module constructs cross-knowledge graphs that connect deep-level knowledge with surface-level knowledge. As such, it enables a more profound exploration of the cues underlying sarcasm. Experiments on the public MSD dataset demonstrate that our approach significantly surpasses previous state-of-the-art methods. Hengyang Zhou, Shaozu Yuan, Meng Chen 0006, Zhiyang Jia, Longbiao Wang, Xiaodong He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Enhancing Semantic Awareness by Sentimental Constraint With Automatic Outlier Masking for Multimodal Sarcasm DetectionabstractMultimodal sarcasm detection, aiming to uncover sarcastic sentiment behind multimodal data, has gained substantial attention in multimodal communities. Recent advancements in multimodal sarcasm detection (MSD) methods have primarily focused on modality alignment with pre-trained vision-language (V-L) model. However, text-image pairs often exhibit weak or even opposite semantic correlations in MSD tasks. Consequently, directly aligning these modalities can potentially result in feature shift and inter-class confusion, ultimately hindering the model's ability. To alleviate this issue, we propose the Enhancing Semantic Awareness Model (ESAM) for multimodal sarcasm detection. Specifically, we first devise a Modality-decoupled Framework (MDF) to separate the textual and visual features from the fused multimodal representation. This decoupling enables the parallel integration of the Sentimental Congruity Constraint (SCC) within both visual and textual latent spaces, thereby enhancing the semantic awareness of different modalities. Furthermore, given that certain outlier samples with ambiguous sentiments can mislead the training and weaken the performance of SCC, we further incorporate Automatic Outlier Masking. This mechanism automatically detects and masks the outliers, guiding the model to focus on more informative samples during training. Experimental results on two public MSD datasets validate the robustness and superiority of our proposed ESAM model. Shaozu Yuan, Hengyang Zhou, Qinfu Xu, Meng Chen 0006, Xiaodong He 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | G^2SAM: Graph-Based Global Semantic Awareness Method for Multimodal Sarcasm DetectionabstractMultimodal sarcasm detection, aiming to detect the ironic sentiment within multimodal social data, has gained substantial popularity in both the natural language processing and computer vision communities. Recently, graph-based studies by drawing sentimental relations to detect multimodal sarcasm have made notable advancements. However, they have neglected exploiting graph-based global semantic congruity from existing instances to facilitate the prediction, which ultimately hinders the model's performance. In this paper, we introduce a new inference paradigm that leverages global graph-based semantic awareness to handle this task. Firstly, we construct fine-grained multimodal graphs for each instance and integrate them into semantic space to draw graph-based relations. During inference, we leverage global semantic congruity to retrieve k-nearest neighbor instances in semantic space as references for voting on the final prediction. To enhance the semantic correlation of representation in semantic space, we also introduce label-aware graph contrastive learning to further improve the performance. Experimental results demonstrate that our model achieves state-of-the-art (SOTA) performance in multimodal sarcasm detection. The code will be available at https://github.com/upccpu/G2SAM. Shaozu Yuan, Hengyang Zhou, Longbiao Wang, Zhiling Yan, Ruosong Yang, Meng Chen 0006 |
AAAI | 7 |
| 2024 | MuJo-SF: Multimodal Joint Slot Filling for Attribute Value Prediction of E-Commerce CommoditiesabstractSupplementing product attribute information is a critical step for E-commerce platforms, which further benefits various downstream tasks, including product recommendation, product search, and product knowledge graph construction. Intuitively, the visual information available on e-commerce platforms can effectively function as a primary source for certain product attributes. However, existing works either extract attribute values solely from textual product descriptions or leverage limited visual information (e.g., image features or optical character recognition tokens) to assist extraction, without mining the fine-grained visual cues linked with the products effectively. In this paper, we propose a novel task -Multimodal Joint Slot Filling(MuJo-SF) - that aims to combine multimodal information from both product descriptions and their corresponding product images to jointly fill values into the pre-defined product attribute set. To this end, we develop MAVP, a new dataset with 79 k instances of product description-image pairs. Specifically, we present a strategy to fulfill visualized saliency ascription, which aims to distinguish between text-dependent and image-dependent attributes. For those image-dependent attributes, we annotate the corresponding values from images using distant supervision. Then, we design a model for MuJo-SF, which combines multimodal representations and fills image-dependent and text-dependent attributes separately. Finally, we conduct extensive experiments on MAVP and provide rich results for MuJo-SF, which can be used as baselines to facilitate future research. Meihuizi Jia, Lei Shen 0001, Anh Tuan Luu, Meng Chen 0006, Lejian Liao, Shaozu Yuan, Xiaodong He 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query GroundingabstractMultimodal named entity recognition (MNER) is a critical step in information extraction, which aims to detect entity spans and classify them to corresponding entity types given a sentence-image pair. Existing methods either (1) obtain named entities with coarse-grained visual clues from attention mechanisms, or (2) first detect fine-grained visual regions with toolkits and then recognize named entities. However, they suffer from improper alignment between entity types and visual regions or error propagation in the two-stage manner, which finally imports irrelevant visual information into texts. In this paper, we propose a novel end-to-end framework named MNER-QG that can simultaneously perform MRC-based multimodal named entity recognition and query grounding. Specifically, with the assistance of queries, MNER-QG can provide prior knowledge of entity types and visual regions, and further enhance representations of both text and image. To conduct the query grounding task, we provide manual annotations and weak supervisions that are obtained via training a highly flexible visual grounding model with transfer learning. We conduct extensive experiments on two public MNER datasets, Twitter2015 and Twitter2017. Experimental results show that MNER-QG outperforms the current state-of-the-art models on the MNER task, and also improves the query grounding performance. Meihuizi Jia, Lei Shen 0001, Lejian Liao, Meng Chen 0006, Xiaodong He 0001 |
AAAI | 5 |
| 2023 | DiffusEmp: A Diffusion Model-Based Framework with Multi-Grained Control for Empathetic Response GenerationabstractEmpathy is a crucial factor in open-domain conversations, which naturally shows one's caring and understanding to others.Though several methods have been proposed to generate empathetic responses, existing works often lead to monotonous empathy that refers to generic and safe expressions.In this paper, we propose to use explicit control to guide the empathy expression and design a framework DIFFUSEMP based on conditional diffusion language model to unify the utilization of dialogue context and attribute-oriented control signals.Specifically, communication mechanism, intent, and semantic frame are imported as multi-grained signals that control the empathy realization from coarse to fine levels.We then design a specific masking strategy to reflect the relationship between multi-grained signals and response tokens, and integrate it into the diffusion model to influence the generative process.Experimental results on a benchmark dataset EMPA-THETICDIALOGUE show that our framework outperforms competitive baselines in terms of controllability, informativeness, and diversity without the loss of context-relatedness. Guanqun Bi, Lei Shen 0001, Yanan Cao 0001, Meng Chen 0006, Yuqiang Xie, Zheng Lin 0001, Xiaodong He 0001 |
ACL (1) | 4 |
| 2023 | Tackling Modality Heterogeneity with Multi-View Calibration Network for Multimodal Sentiment DetectionabstractYiwei Wei, Shaozu Yuan, Ruosong Yang, Lei Shen, Zhangmeizhi Li, Longbiao Wang, Meng Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shaozu Yuan, Ruosong Yang, Lei Shen 0001, Zhangmeizhi Li, Longbiao Wang, Meng Chen 0006 |
ACL (1) | 7 |
| 2023 | Dialog-Post: Multi-Level Self-Supervised Objectives and Hierarchical Model for Dialogue Post-TrainingabstractDialogue representation and understanding aim to convert conversational inputs into embeddings and fulfill discriminative tasks.Compared with free-form text, dialogue has two important characteristics, hierarchical semantic structure and multi-facet attributes.Therefore, directly applying the pretrained language models (PLMs) might result in unsatisfactory performance.Recently, several work focused on the dialogue-adaptive post-training (Dial-Post) that further trains PLMs to fit dialogues.To model dialogues more comprehensively, we propose a DialPost method, DIALOG-POST, with multi-level self-supervised objectives and a hierarchical model.These objectives leverage dialogue-specific attributes and use selfsupervised signals to fully facilitate the representation and understanding of dialogues.The novel model is a hierarchical segment-wise self-attention network, which contains innersegment and inter-segment self-attention sublayers followed by an aggregation and updating module.To evaluate the effectiveness of our methods, we first apply two public datasets for the verification of representation ability.Then we conduct experiments on a newly-labelled dataset that is annotated with 4 dialogue understanding tasks.Experimental results show that our method outperforms existing SOTA models and achieves a 3.3% improvement on average. Token-level SSOs Utterance-level SSO Dialogue-level SSOs𝑄: 这个手机支持5G吗?𝑄: Zhenyu Zhang 0029, Lei Shen 0001, Meng Chen 0006, Xiaodong He 0001 |
ACL (1) | 4 |
| 2023 | POSPAN: Position-Constrained Span Masking for Language Model Pre-trainingabstractSpan-level masked language modeling (MLM) has shown to be advantageous to pre-trained language models over the original single-token MLM, as entities/phrases and their dependencies are critical to language understanding. Previous works only consider span length with some discrete distributions, while the dependencies among spans are ignored, i.e., assuming that the positions of masked spans are uniformly distributed. In this paper, we present POSPAN, a general framework to allow diverse position-constrained span masking strategies via the combination of span length distribution and position constraint distribution, which unifies all existing span-level masking methods. To verify the effectiveness of POSPAN in pre-training, we evaluate it on the datasets from several NLU benchmarks. Experimental results indicate that the position constraint is capable of enhancing span-level masking broadly, and our best POSPAN setting consistently outperforms its span-length-only counterparts and vanilla MLM. We also conduct theoretical analysis for the position constraint in masked language models to shed light on the reason why POSPAN works well, demonstrating the rationality and necessity of POSPAN. Zhenyu Zhang 0029, Lei Shen 0001, Meng Chen 0006, Xiaodong He 0001 |
CIKM | 4 |
| 2023 | UFO2: A Unified Pre-Training Framework for Online and Offline Speech RecognitionabstractIn this paper, we propose a Unified pre-training Framework for Online and Offline (UFO2) Automatic Speech Recognition (ASR), which 1) simplifies the two separate training workflows for online and offline modes into one process, and 2) improves the Word Error Rate (WER) performance with limited utterance annotating. Specifically, we extend the conventional offline-mode Self-Supervised Learning (SSL)-based ASR approach to a unified manner, where the model training is conditioned on both the full-context and dynamic- chunked inputs. To enhance the pre-trained representation model, stop-gradient operation is applied to decouple the online-mode objectives to the quantizer. Moreover, in both the pre-training and the downstream fine-tuning stages, joint losses are proposed to train the unified model with full-weight sharing for the two modes. Experimental results on the LibriSpeech dataset show that UFO2 outperforms the SSL-based baseline method by 29.7% and 18.2% relative WER reduction in offline and online modes, respectively. Qingtao Li, Fangzhu Li, Fan Lu 0003, Meng Chen 0006, Xiaodong He 0001 |
ICASSP | 7 |
| 2023 | Improving Disfluency Detection with Multi-Scale Self Attention and Contrastive LearningabstractDisfluency detection aims to recognize disfluencies in sentences. Existing works usually adopt a sequence labeling model to tackle this task. They also attempt to integrate into models the feature that the disfluencies are similar to the correct phrase, the so-called "rough copy". However, they heavily rely on hand-craft features or word-to-word match patterns, which are insufficient to precisely capture such rough copy and cause under-tagging and over-tagging problems. To alleviate these problems, we propose a multi-scale self-attention mechanism (MSAT) and design contrastive learning (CL) loss for this task. Specifically, the MSAT leverages token representations to learn representations for different scales of phrases, and then compute similarity among them. The CL adopts the fluent version of the input to build the positive and negative samples and encourages the model to keep the fluent version consistent with the input in semantics. We conduct experiments on a public English dataset Switchboard, and an in-house Chinese dataset Waihu, which is derived from an online conversation bot. Results show that our method outperforms the baselines and achieves superior performance on both datasets. Peiying Wang, Chaoqun Duan, Meng Chen 0006, Xiaodong He 0001 |
ICASSP | 3 |
| 2023 | Enhancing Multimodal Alignment with Momentum Augmentation for Dense Video CaptioningabstractDense video captioning aims to localize multiple events from an untrimmed video and generate corresponding captions for each event. Fusing different modalities(e.g. rgb, flow, audio) via transformer structure is a promising way to improve the caption performance. However, it is challenging for the cross-modal encoder to learn multimodal interactions due to their inherent disparities of distribution. In this paper, we propose a novel transformer structure with contrastive learning to align different modalities. Specifically, to avoid the limitation of small batch size and false contrastive targets, we design an event-aligned momentum augmentation strategy to apply contrast learning for dense video captioning. The experimental result shows that our proposals outperform all existing multimodal fusion methods for dense video captioning. Shaozu Yuan, Meng Chen 0006, Longbiao Wang |
ICASSP | 3 |
| 2023 | OTF: Optimal Transport based Fusion of Supervised and Self-Supervised Learning Models for Automatic Speech Recognition
Qingtao Li, Fangzhu Li, Fan Lu 0003, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001 |
INTERSPEECH | 7 |
| 2023 | Leveraging Label Information for Multimodal Emotion Recognition
Peiying Wang, Sunlu Zeng, Fan Lu 0003, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001 |
INTERSPEECH | 5 |
| 2023 | Enhancing New Intent Discovery via Robust Neighbor-based Contrastive Learning
Zhenhe Wu, Xiaoguang Yu, Meng Chen 0006, Liangqing Wu, Jiahao Ji, Zhoujun Li 0001 |
INTERSPEECH | 3 |
| 2023 | MPP-net: Multi-perspective perception network for dense video captioning
Shaozu Yuan, Meng Chen 0006, Longbiao Wang, Lei Shen 0001, Zhiling Yan |
Neurocomputing | 3 |
| 2022 | Legal Charge Prediction via Bilinear Attention NetworkabstractThe legal charge prediction task aims to judge appropriate charges according to the given fact description in cases. Most existing methods formulate it as a multi-class text classification problem and have achieved tremendous progress. However, the performance on low-frequency charges is still unsatisfactory. Previous studies indicate leveraging the charge label information can facilitate this task, but the approaches to utilizing the label information are not fully explored. In this paper, inspired by the vision-language information fusion techniques in the multi-modal field, we propose a novel model (denoted as LeapBank) by fusing the representations of text and labels to enhance the legal charge prediction task. Specifically, we devise a representation fusion block based on the bilinear attention network to interact the labels and text tokens seamlessly. Extensive experiments are conducted on three real-world datasets to compare our proposed method with state-of-the-art models. Experimental results show that LeapBank obtains up to 8.5% Macro-F1 improvements on the low-frequency charges, demonstrating our model's superiority and competitiveness. Yuquan Le, Meng Chen 0006, Zhe Quan, Xiaodong He 0001, Kenli Li 0001 |
CIKM | 3 |
| 2022 | Few-Shot Table Understanding: A Benchmark Dataset and Pre-Training BaselineabstractFew-shot table understanding is a critical and challenging problem in real-world scenario as annotations over large amount of tables are usually costly. Pre-trained language models (PLMs), which have recently flourished on tabular data, have demonstrated their effectiveness for table understanding tasks. However, few-shot table understanding is rarely explored due to the deficiency of public table pre-training corpus and well-defined downstream benchmark tasks, especially in Chinese. In this paper, we establish a benchmark dataset, FewTUD, which consists of 5 different tasks with human annotations to systematically explore the few-shot table understanding in depth. Since there is no large number of public Chinese tables, we also collect a large-scale, multi-domain tabular corpus to facilitate future Chinese table pre-training, which includes one million tables and related natural language text with auxiliary supervised interaction signals. Finally, we present FewTPT, a novel table PLM with rich interactions over tabular data, and evaluate its performance comprehensively on the benchmark. Our dataset and model will be released to the public soon. Ruixue Liu, Shaozu Yuan, Aijun Dai, Lei Shen 0001, Tiangang Zhu, Meng Chen 0006, Xiaodong He 0001 |
COLING | 6 |
| 2022 | Building Robust Spoken Language Understanding by Cross Attention Between Phoneme Sequence and ASR HypothesisabstractBuilding Spoken Language Understanding (SLU) robust to Automatic Speech Recognition (ASR) errors is an essential issue for various voice-enabled virtual assistants. Considering that most ASR errors are caused by phonetic confusion between similar-sounding expressions, intuitively, leveraging the phoneme sequence of speech can complement ASR hypothesis and enhance the robustness of SLU. This paper proposes a novel model with Cross Attention for SLU (denoted as CASLU). The cross attention block is devised to catch the fine-grained interactions between phoneme and word embeddings in order to make the joint representations catch the phonetic and semantic features of input simultaneously and for overcoming the ASR errors in downstream natural language understanding (NLU) tasks. Extensive experiments are conducted on three datasets, showing the effectiveness and competitiveness of our approach. Additionally, We also validate the universality of CASLU and prove its complementarity when combining with other robust SLU techniques. Zexun Wang, Yuquan Le, Mingchao Feng, Meng Chen 0006, Xiaodong He 0001 |
ICASSP | 6 |
| 2022 | Gated Multimodal Fusion with Contrastive Learning for Turn-Taking Prediction in Human-Robot DialogueabstractTurn-taking, aiming to decide when the next speaker can start talking, is an essential component in building human-robot spoken dialogue systems. Previous studies indicate that multi-modal cues can facilitate this challenging task. However, due to the paucity of public multimodal datasets, current methods are mostly limited to either utilizing unimodal features or simplistic multimodal ensemble models. Besides, the inherent class imbalance in real scenario, e.g. sentence ending with short pause will be mostly regarded as the end of turn, also poses great challenge to the turn-taking decision. In this paper, we first collect a large-scale annotated corpus for turn-taking with over 5,000 real human-robot dialogues in speech and text modalities. Then, a novel gated multimodal fusion mechanism is devised to utilize various information seamlessly for turn-taking prediction. More importantly, to tackle the data imbalance issue, we design a simple yet effective data augmentation method to construct negative instances without supervision and apply contrastive learning to obtain better feature representations. Extensive experiments are conducted and the results demonstrate the superiority and competitiveness of our model over several state-of-the-art baselines. Jiudong Yang, Peiying Wang, Mingchao Feng, Meng Chen 0006, Xiaodong He 0001 |
ICASSP | 5 |
| 2022 | SE-GAN: Skeleton Enhanced Gan-Based Model for Brush Handwriting Font GenerationabstractPrevious works on font generation mainly focus on the standard print fonts where character's shape is stable and strokes are clearly separated. There is rare research on brush hand-writing font generation, which involves holistic structure changes and complex strokes transfer. To address this issue, we propose a novel GAN-based image translation model by integrating the skeleton information. We first extract the skeleton from training images, then design an image encoder and a skeleton encoder to extract corresponding features. A self-attentive refined attention module is devised to guide the model to learn distinctive features between different domains. A skeleton discriminator is involved to first synthesize the skeleton image from the generated image with a pre-trained generator, then to judge its realness to the target one. We also contribute a large-scale brush handwriting font image dataset with six styles and 15,000 high-resolution images. Both quantitative and qualitative experimental results demonstrate the competitiveness of our proposed model. Shaozu Yuan, Ruixue Liu, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001 |
ICME | 3 |
| 2022 | Learning to Generate Poetic Chinese Landscape Painting with CalligraphyabstractIn this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generation, Polaca takes the classic poetry as input and outputs the artistic landscape painting image with the corresponding calligraphy. It is equipped with three different modules to complete the whole piece of landscape painting artwork: the first one is a text-to-image module to generate landscape painting image, the second one is an image-to-image module to generate stylistic calligraphy image, and the third one is an image fusion module to fuse the two images into a whole piece of aesthetic artwork. Shaozu Yuan, Aijun Dai, Zhiling Yan, Ruixue Liu, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001 |
IJCAI | 5 |
| 2022 | SCaLa: Supervised Contrastive Learning for End-to-End Speech RecognitionabstractEnd-to-end Automatic Speech Recognition (ASR) models are usually trained to optimize the loss of the whole token sequence, while neglecting explicit phonemic-granularity supervision.This could result in recognition errors due to similarphoneme confusion or phoneme reduction.To alleviate this problem, we propose a novel framework based on Supervised Contrastive Learning (SCaLa) to enhance phonemic representation learning for end-to-end ASR systems.Specifically, we extend the self-supervised Masked Contrastive Predictive Coding (MCPC) to a fully-supervised setting, where the supervision is applied in the following way.First, SCaLa masks variablelength encoder features according to phoneme boundaries given phoneme forced-alignment extracted from a pre-trained acoustic model; it then predicts the masked features via contrastive learning.The forced-alignment can provide phoneme labels to mitigate the noise introduced by positive-negative pairs in selfsupervised MCPC.Experiments on reading and spontaneous speech datasets show that our proposed approach achieves 2.8 and 1.4 points Character Error Rate (CER) absolute reductions compared to the baseline, respectively. Runyu Wang, Fan Lu 0003, Zhengchen Zhang, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001 |
INTERSPEECH | 6 |
| 2022 | Cross-modal Transfer Learning via Multi-grained Alignment for End-to-End Spoken Language Understanding
Zexun Wang, Hang Liu 0005, Peiying Wang, Mingchao Feng, Meng Chen 0006, Xiaodong He 0001 |
INTERSPEECH | 6 |
| 2022 | E-ConvRec: A Large-Scale Conversational Recommendation Dataset for E-Commerce Customer ServiceabstractThere has been a growing interest in developing conversational recommendation system (CRS), which provides valuable recommendations to users through conversations. Compared to the traditional recommendation, it advocates wealthier interactions and provides possibilities to obtain users’ exact preferences explicitly. Nevertheless, the corresponding research on this topic is limited due to the lack of broad-coverage dialogue corpus, especially real-world dialogue corpus. To handle this issue and facilitate our exploration, we construct E-ConvRec, an authentic Chinese dialogue dataset consisting of over 25k dialogues and 770k utterances, which contains user profile, product knowledge base (KB), and multiple sequential real conversations between users and recommenders. Next, we explore conversational recommendation in a real scene from multiple facets based on the dataset. Therefore, we particularly design three tasks: user preference recognition, dialogue management, and personalized recommendation. In the light of the three tasks, we establish baseline results on E-ConvRec to facilitate future studies. Meihuizi Jia, Ruixue Liu, Peiying Wang, Yang Song 0008, Zexi Xi, Haobin Li, Meng Chen 0006, Jinhui Pang, Xiaodong He 0001 |
LREC | 8 |
| 2022 | Query Prior Matters: A MRC Framework for Multimodal Named Entity RecognitionabstractMultimodal named entity recognition (MNER) is a vision-language task where the system is required to detect entity spans and corresponding entity types given a sentence-image pair. Existing methods capture text-image relations with various attention mechanisms that only obtain implicit alignments between entity types and image regions. To locate regions more accurately and better model cross-/within-modal relations, we propose a machine reading comprehension based framework for MNER, namely MRC-MNER. By utilizing queries in MRC, our framework can provide prior information about entity types and image regions. Specifically, we design two stages, Query-Guided Visual Grounding and Multi-Level Modal Interaction, to align fine-grained type-region information and simulate text-image/inner-text interactions respectively. For the former, we train a visual grounding model via transfer learning to extract region candidates that can be further integrated into the second stage to enhance token representations. For the latter, we design text-image and inner-text interaction modules along with three sub-tasks for MRC-MNER. To verify the effectiveness of our model, we conduct extensive experiments on two public MNER datasets, Twitter2015 and Twitter2017. Experimental results show that MRC-MNER outperforms the current state-of-the-art models on Twitter2017, and yields competitive results on Twitter2015. Meihuizi Jia, Lei Shen 0001, Jinhui Pang, Lejian Liao, Yang Song 0008, Meng Chen 0006, Xiaodong He 0001 |
ACM Multimedia | 7 |
| 2022 | Label Anchored Contrastive Learning for Language UnderstandingabstractContrastive learning (CL) has achieved astonishing progress in computer vision, speech, and natural language processing fields recently with self-supervised learning.However, CL approach to the supervised setting is not fully explored, especially for the natural language understanding classification task.Intuitively, the class label itself has the intrinsic ability to perform hard positive/negative mining, which is crucial for CL.Motivated by this, we propose a novel label anchored contrastive learning approach (denoted as LaCon) for language understanding.Specifically, three contrastive objectives are devised, including a multi-head instance-centered contrastive loss (ICL), a label-centered contrastive loss (LCL), and a label embedding regularizer (LER).Our approach does not require any specialized network architecture or any extra data augmentation, thus it can be easily plugged into existing powerful pre-trained language models.Compared to the state-of-the-art baselines, LaCon obtains up to 4.1% improvement on the popular datasets of GLUE and CLUE benchmarks.Besides, LaCon also demonstrates significant advantages under the few-shot and data imbalance settings, which obtains up to 9.4% improvement on the FewGLUE and FewCLUE benchmarking tasks. Zhenyu Zhang 0029, Meng Chen 0006, Xiaodong He 0001 |
NAACL-HLT | 3 |
| 2022 | MCIC: Multimodal Conversational Intent Classification for E-commerce Customer Service
Shaozu Yuan, Hang Liu 0005, Zhiling Yan, Ruixue Liu, Meng Chen 0006 |
NLPCC (1) | 7 |
| 2021 | DialogueBERT: A Self-Supervised Learning based Dialogue Pre-training EncoderabstractWith the rapid development of artificial intelligence, conversational bots have became prevalent in mainstream E-commerce platforms, which can provide convenient customer service timely. To satisfy the user, the conversational bots need to understand the user's intention, detect the user's emotion, and extract the key entities from the conversational utterances. However, understanding dialogues is regarded as a very challenging task. Different from common language understanding, utterances in dialogues appear alternately from different roles and are usually organized as hierarchical structures. To facilitate the understanding of dialogues, in this paper, we propose a novel contextual dialogue encoder (i.e. DialogueBERT) based on the popular pre-trained language model BERT. Five self-supervised learning pre-training tasks are devised for learning the particularity of dialouge utterances. Four different input embeddings are integrated to catch the relationship between utterances, including turn embedding, role embedding, token embedding and position embedding. DialogueBERT was pre-trained with 70 million dialogues in real scenario, and then fine-tuned in three different downstream dialogue understanding tasks. Experimental results show that DialogueBERT achieves exciting results with 88.63% accuracy for intent recognition, 94.25% accuracy for emotion recognition and 97.04% F1 score for named entity recognition, which outperforms several strong baselines by a large margin. Zhenyu Zhang 0029, Meng Chen 0006 |
CIKM | 3 |
| 2021 | Conversational Query Rewriting with Self-Supervised LearningabstractContext modeling plays a critical role in building multi-turn dialogue systems. Conversational Query Rewriting (CQR) aims to simplify the multi-turn dialogue modeling into a single-turn problem by explicitly rewriting the conversational query into a self-contained utterance. However, existing approaches rely on massive supervised training data, which is labor-intensive to annotate. And the detection of the omitted important information from context can be further improved. Besides, intent consistency constraint between contextual query and rewritten query is also ignored. To tackle these issues, we first propose to construct a large-scale CQR dataset automatically via self-supervised learning, which does not need human annotation. Then we introduce a novel CQR model Teresa based on Transformer, which is enhanced by self-attentive keywords detection and intent consistency constraint. Finally, we conduct extensive experiments on two public datasets. Experimental results demonstrate that our proposed model outperforms existing CQR baselines significantly, and also prove the effectiveness of self-supervised learning on improving the CQR performance. Hang Liu 0005, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 2 |
| 2021 | ViDA-MAN: Visual Dialog with Digital HumansabstractWe demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text or voice-based system, ViDA-MAN offers human-like interactions (e.g, vivid voice, natural facial expression and body gestures). Given a speech request, the demonstration is able to response with high quality videos in sub-second latency. To deliver immersive user experience, ViDA-MAN seamlessly integrates multi-modal techniques including Acoustic Speech Recognition (ASR), multi-turn dialog, Text To Speech (TTS), talking heads video generation. Backed with large knowledge base, ViDA-MAN is able to chat with users on a number of topics including chit-chat, weather, device control, News recommendations, booking hotels, as well as answering questions via structured knowledge. Jiawei Zuo, Liqin Jiang, Meng Chen 0006, Zhengchen Zhang, Wei Zhang 0031, Xiaodong He 0001, Tao Mei 0001 |
ACM Multimedia | 6 |
| 2021 | Learning to Compose Stylistic Calligraphy Artwork with EmotionsabstractEmotion plays a critical role in calligraphy composition, which makes the calligraphy artwork impressive and have a soul. However, previous research on calligraphy generation all neglected the emotion as a major contributor to the artistry of calligraphy. Such defects prevent them from generating aesthetic, stylistic, and diverse calligraphy artworks, but only static handwriting font library instead. To address this problem, we propose a novel cross-modal approach to generate stylistic and diverse Chinese calligraphy artwork driven by different emotions automatically. We firstly detect the emotions in the text by a classifier, then generate the emotional Chinese character images via a novel modified Generative Adversarial Network (GAN) structure, finally we predict the layout for all character images with a recurrent neural network. We also collect a large-scale stylistic Chinese calligraphy image dataset with rich emotions. Experimental results demonstrate that our model outperforms all baseline image translation models significantly for different emotional styles in terms of content accuracy and style discrepancy. Besides, our layout algorithm can also learn the patterns and habits of calligrapher, and makes the generated calligraphy more artistic. To the best of our knowledge, we are the first to work on emotion-driven discourse-level Chinese calligraphy artwork composition. Shaozu Yuan, Ruixue Liu, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001 |
ACM Multimedia | 3 |
| 2020 | Learning to Predict Charges for Legal Judgment via Self-Attentive Capsule NetworkabstractWith the rapid development of deep learning technology, more and more traditional industries are changed by Artificial Intelligence. The legal industry is such a popular scenario which attracts lots of researchers' interests. In this work, we focus on automatic charge prediction, which predicts the final charges according to the given fact descriptions in criminal cases. It is crucial for legal assistant systems and can help the judges improve work efficiency greatly. However, extremely imbalanced data distribution and lengthy fact descriptions make this task especially challenging. To tackle these two issues, we propose a novel model, namely Self-Attentive Capsule Network (dubbed as SAttCaps). In particular, we devise a self-attentive dynamic routing, which can not only capture long-range dependency more directly than vanilla dynamic routing, but also learn the high-level generalized features better. The experimental results on three real-world datasets demonstrate that our model significantly outperforms the baselines and creates new state-of-the-art performance. Moreover, our model performs much better than the baselines especially in the low-frequency charges and can bring 5.7% absolute improvement under F1 score. Yuquan Le, Congqing He, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ECAI | 3 |
| 2020 | Compose Like Humans: Jointly Improving the Coherence and Novelty for Modern Chinese Poetry GenerationabstractChinese poetry is an important part of worldwide culture, and classical and modern sub-branches are quite different. The former is a unique genre and has strict constraints, while the latter is very flexible in length, optional to have rhymes, and similar to modern poetry in other languages. Thus, it requires more to control the coherence and improve the novelty. In this paper, we propose a generate-retrieve-then-refine paradigm to jointly improve the coherence and novelty. In the first stage, a draft is generated given keywords (i.e., topics) only. The second stage produces a "refining vector" from retrieval lines. At last, we take into consideration both the draft and the "refining vector" to generate a new poem. The draft provides future sentence-level information for a line to be generated. Meanwhile, the "refining vector" points out the direction of refinement based on impressive words detection mechanism which can learn good patterns from references and then create new ones via insertion operation. Experimental results on a collected large-scale modern Chinese poetry dataset show that our proposed approach can not only generate more coherent poems, but also improve the diversity and novelty. Lei Shen 0001, Meng Chen 0006 |
IJCNN | 3 |
| 2020 | The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer ServiceabstractHuman conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge amount of real conversation data. In this paper, we construct a large-scale real scenario Chinese E-commerce conversation corpus, JDDC, with more than 1 million multi-turn dialogues, 20 million utterances, and 150 million words. The dataset reflects several characteristics of human-human conversations, e.g., goal-driven, and long-term dependency among the context. It also covers various dialogue types including task-oriented, chitchat and question-answering. Extra intent information and three well-annotated challenge sets are also provided. Then, we evaluate several retrieval-based and generative models to provide basic benchmark performance on the JDDC corpus. And we hope JDDC can serve as an effective testbed and benefit the development of fundamental research in dialogue task. Meng Chen 0006, Ruixue Liu, Lei Shen 0001, Shaozu Yuan, Jingyan Zhou, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
LREC | 1 |
| 2020 | MaLiang: An Emotion-driven Chinese Calligraphy Artwork Composition SystemabstractWe present a novel Chinese calligraphy artwork composition system (MaLiang) which can generate aesthetic, stylistic and diverse calligraphy images based on the emotion status from the input text. Different from previous research, it's the first work to endow the calligraphy synthesis with the ability to express fickle emotions and composite a whole piece of discourse-level calligraphy artwork instead of single character images. The system consists of three modules: emotion detection, character image generation, and layout prediction. As a creative form of interactive art, MaLiang has been exhibited in several famous international art festivals. Ruixue Liu, Shaozu Yuan, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001 |
ACM Multimedia | 3 |
| 2020 | Enhancing Multi-turn Dialogue Modeling with Intent Information for E-Commerce Customer Service
Ruixue Liu, Meng Chen 0006, Hang Liu 0005, Lei Shen 0001, Yang Song 0008, Xiaodong He 0001 |
NLPCC (1) | 2 |
| 2019 | Mappa Mundi: An Interactive Artistic Mind Map Generator with Artificial ImaginationabstractWe present a novel real-time, collaborative, and interactive AI painting system, Mappa Mundi, for artistic Mind Map creation. The system consists of a voice-based input interface, an automatic topic expansion module, and an image projection module. The key innovation is to inject Artificial Imagination into painting creation by considering lexical and phonological similarities of language, learning and inheriting artist’s original painting style, and applying the principles of Dadaism and impossibility of improvisation. Our system indicates that AI and artist can collaborate seamlessly to create imaginative artistic painting and Mappa Mundi has been applied in art exhibition in UCCA, Beijing. Ruixue Liu, Baoyang Chen, Meng Chen 0006, Youzheng Wu, Zhijie Qiu, Xiaodong He 0001 |
IJCAI | 3 |
| 2019 | Automated Thematic and Emotional Modern Chinese Poetry Composition
Meng Chen 0006, Yang Song 0008, Xiaodong He 0001, Bowen Zhou 0001 |
NLPCC (1) | 2 |
| 2019 | A Sequence-to-Action Architecture for Character-Based Chinese Dependency Parsing with Status History
Hang Liu 0005, Meng Chen 0006, Jin An Xu, Yufeng Chen 0005 |
NLPCC (2) | 3 |