EDBT 2026 Demo / reviewers in the wild / expert
Youzheng Wu
dblp:01/3620
· DBLP profile ↗
48ranked-venue papers
11as first author
26since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 10 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Comet: Dialog Context Fusion Mechanism for End-to-End Task-Oriented Dialog with Multi-task LearningabstractExisting end-to-end task-oriented dialog systems often encounter challenges arising from implicit information, coreference, and the presence of noisy and irrelevant data within the dialog context. These issues hinder the system’s ability to fully comprehend critical information and lead to inaccurate responses. To address these concerns, we propose Comet, a dialog context fusion mechanism for end-to-end task-oriented dialog, augmented with three supplementary tasks: dialog summarization, domain prediction, and slot detection. Dialog summarization facilitates a more comprehensive understanding of important dialog context information by Comet. Domain prediction enables Comet to concentrate on domain-specific information, thus reducing interference from irrelevant information. Slot detection empowers Comet to accurately identify and comprehend essential dialog context information. Additionally, we introduce a data refinement strategy to enhance the comprehensiveness and recommendability of the generated responses. Experimental results demonstrate the superior performance of our proposed methods compared to existing end-to-end task-oriented dialog systems, achieving state-of-the-art results on the MultiWOZ and CrossWOZ datasets. Haipeng Sun, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001 |
COLING | 3 |
| 2025 | UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech RecognitionabstractRecent advancements in scaling up models have significantly improved performance in Automatic Speech Recognition (ASR) tasks. However, training large ASR models from scratch remains costly. To address this issue, we introduce UME, a novel method that efficiently Upcycles pretrained dense ASR checkpoints into larger Mixture-of-Eperts (MoE) architectures. Initially, feed-forward networks are converted into MoE layers. By reusing the pretrained weights, we establish a robust foundation for the expanded model, significantly reducing optimization time. Then, layer freezing and expert balancing strategies are employed to continue training the model, further enhancing performance. Experiments on a mixture of 170k-hour Mandarin and English datasets show that UME: 1) surpasses the pretrained baseline by a margin of 11.9% relative error rate reduction while maintaining comparable latency; 2) reduces training time by up to 86.7% and achieves superior accuracy compared to training models of the same size from scratch. Shanyong Yu, Fan Lu 0003, Youzheng Wu, Xiaodong He 0001 |
ICASSP | 5 |
| 2024 | An efficient confusing choices decoupling framework for multi-choice tasks over texts
Yingyao Wang, Junwei Bao 0001, Chaoqun Duan, Youzheng Wu, Xiaodong He 0001, Conghui Zhu, Tiejun Zhao |
Neural Comput. Appl. | 4 |
| 2024 | Operation-Augmented Numerical Reasoning for Question AnsweringabstractQuestion answering requiring numerical reasoning, which generally involves symbolic operations such as sorting, counting, and addition, is a challenging task. To address such a problem, existing mixture-of-experts (MoE)-based methods design several specific answer predictors to handle different types of questions and achieve promising performance. However, they ignore the modeling and exploitation of fine-grained reasoning-related operations to support numerical reasoning, encountering the inadequacy in reasoning capability and interpretability. To alleviate this issue, we propose OPERA, an operation-augmented numerical reasoning framework. Concretely, we systematically define a scalable operation set to model numerical reasoning. We first identify reasoning-related operations based on context and then softly execute them to imitate the answer reasoning procedure via an operation-aware cross-attention mechanism. Finally, we utilize the operation-augmented semantic representation of execution results to support answer prediction. We verify the effectiveness and generalization of OPERA in two scenarios with different knowledge sources and reasoning capabilities. Specifically, we conduct extensive experiments on two textual datasets, DROP and RACENum, and a table-text hybrid dataset TAT-QA. Experiment results show that OPERA outperforms previous strong methods on the DROP, RACENum, and TAT-QA datasets. Further, we statistically and visually analyze its interpretability. Yongwei Zhou, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationabstractRecently, the contrastive language-image pre-training, e.g., CLIP, has demonstrated promising results on various downstream tasks. The pre-trained model can capture enriched visual concepts for images by learning from a large scale of text-image data. However, transferring the learned visual knowledge to open-vocabulary semantic segmentation is still under-explored. In this paper, we propose a CLIP-based model named SegCLIP for the topic of open-vocabulary segmentation in an annotation-free manner. The SegCLIP achieves segmentation based on ViT and the main idea is to gather patches with learnable centers to semantic regions through training on text-image pairs. The gathering operation can dynamically capture the semantic groups, which can be used to generate the final segmentation results. We further propose a reconstruction loss on masked patches and a superpixel-based KL loss with pseudo-labels to enhance the visual representation. Experimental results show that our model achieves comparable or superior segmentation accuracy on the PASCAL VOC 2012 (+0.3% mIoU), PASCAL Context (+2.3% mIoU), and COCO (+2.2% mIoU) compared with baselines. We release the code at https://github.com/ArrowLuo/SegCLIP. Huaishao Luo, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001, Tianrui Li 0001 |
ICML | 3 |
| 2023 | OTF: Optimal Transport based Fusion of Supervised and Self-Supervised Learning Models for Automatic Speech Recognition
Qingtao Li, Fangzhu Li, Fan Lu 0003, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001 |
INTERSPEECH | 8 |
| 2023 | Leveraging Label Information for Multimodal Emotion Recognition
Peiying Wang, Sunlu Zeng, Fan Lu 0003, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001 |
INTERSPEECH | 6 |
| 2023 | MaskedSpeech: Context-aware Speech Synthesis with Masking Strategy
Ya-Jie Zhang, Yanghao Yue, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001 |
INTERSPEECH | 5 |
| 2023 | Prosody Modelling With Pre-Trained Cross-Utterance Representations for Improved Speech SynthesisabstractWhen humans speak multiple utterances in a continuous manner, the prosodic features generated in each utterance are related to those in its neighbouring utterances. Such cross-utterance (CU) dependencies are often ignored by the current neural text-to-speech (TTS) systems, which reduces the naturalness and expressiveness of the synthesized speeches. In this paper, we propose to improve the prosody modelling ability of neural TTS systems using pre-trained CU acoustic and text representations. Such CU acoustic representations are derived using the Wav2Vec 2.0 model (W2V2) from the synthesized audios of the past utterances, while the CU text representations are extracted using the Bidirectional Encoder Representation from Transformers (BERT) model from the scripts of the future utterances. Experimental results on a Mandarin audiobook and an English audiobook showed the naturalness and expressiveness of the synthesized audios were significantly improved by incorporating such pre-trained W2V2 and BERT CU representations into the Fastspeech2 TTS framework. Ya-Jie Zhang, Chao Zhang 0031, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | A Multi-Factor Classification Framework for Completing Users' Fuzzy Queries (Student Abstract)abstractIntent identification is the key technology in dialogue system. However, not all online queries are clear or complete. To identify users' intents from those fuzzy queries accurately, this paper proposes a multi-factor classification framework on the query level. Experimental results on our online serving system JIMI demonstrate the effectiveness of our proposed framework. Liangqing Wu, Xiaoguang Yu, Shuangyong Song, Youzheng Wu, Xiaodong He 0001 |
AAAI | 7 |
| 2022 | Fine- and Coarse-Granularity Hybrid Self-Attention for Efficient BERTabstractTransformer-based pre-trained models, such as BERT, have shown extraordinary success in achieving state-of-the-art results in many natural language processing applications.However, deploying these models can be prohibitively costly, as the standard self-attention mechanism of the Transformer suffers from quadratic computational cost in the input sequence length.To confront this, we propose FCA, a fine-and coarse-granularity hybrid self-attention that reduces the computation cost through progressively shortening the computational sequence length in self-attention.Specifically, FCA conducts an attention-based scoring strategy to determine the informativeness of tokens at each layer.Then, the informative tokens serve as the fine-granularity computing units in selfattention and the uninformative tokens are replaced with one or several clusters as the coarsegranularity computing units in self-attention.Experiments on GLUE and RACE datasets show that BERT with FCA achieves 2x reduction in FLOPs over original BERT with <1% loss in accuracy.We show that FCA offers significantly better trade-off between accuracy and FLOPs compared to prior methods 1 . Yifan Wang 0016, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001 |
ACL (1) | 4 |
| 2022 | AutoQGS: Auto-Prompt for Low-Resource Knowledge-based Question Generation from SPARQLabstractThis study investigates the task of knowledge-based question generation (KBQG). Conventional KBQG works generated questions from fact triples in the knowledge graph, which could not express complex operations like aggregation and comparison in SPARQL. Moreover, due to the costly annotation of large-scale SPARQL-question pairs, KBQG from SPARQL under low-resource scenarios urgently needs to be explored. Recently, since the generative pre-trained language models (PLMs) typically trained in natural language (NL)-to-NL paradigm have been proven effective for low-resource generation, e.g., T5 and BART, how to effectively utilize them to generate NL-question from non-NL SPARQL is challenging. To address these challenges, AutoQGS, an auto-prompt approach for low-resource KBQG from SPARQL, is proposed. Firstly, we put forward to generate questions directly from SPARQL for KBQG task to handle complex operations. Secondly, we propose an auto-prompter trained on large-scale unsupervised data to rephrase SPARQL into NL description, smoothing the low-resource transformation from non-NL SPARQL to NL question with PLMs. Experimental results on the WebQuestionsSP, ComlexWebQuestions 1.1, and PathQuestions show that our model achieves state-of-the-art performance, especially in low-resource settings. Furthermore, a corpora of 330k factoid complex question-SPARQL pairs is generated for further KBQG research. Guanming Xiong, Junwei Bao 0001, Wen Zhao 0008, Youzheng Wu, Xiaodong He 0001 |
CIKM | 4 |
| 2022 | PRINCE: Prefix-Masked Decoding for Knowledge Enhanced Sequence-to-Sequence Pre-TrainingabstractPre-trained Language Models (PLMs) have shown effectiveness in various Natural Language Processing (NLP) tasks.Denoising autoencoder is one of the most successful pretraining frameworks, learning to recompose the original text given a noise-corrupted one.The existing studies mainly focus on injecting noises into the input.This paper introduces a simple yet effective pre-training paradigm, equipped with a knowledge-enhanced decoder that predicts the next entity token with noises in the prefix, explicitly strengthening the representation learning of entities that span over multiple input tokens.Specifically, when predicting the next token within an entity, we feed masks into the prefix in place of some of the previous ground-truth tokens that constitute the entity.Our model achieves new state-of-the-art results on two knowledge-driven data-to-text generation tasks with up to 2% BLEU gains. Song Xu 0002, Haoran Li 0001, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001 |
EMNLP | 4 |
| 2022 | JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and SummarizationabstractThe popularity of multimodal dialogue has stimulated the need for a new generation of dialogue agents with multimodal interactivity.When users communicate with customer service, they may express their requirements by means of text, images, or even videos.Visual information usually acts as discriminators for product models, or indicators of product failures, which play an important role in the Ecommerce scenario.On the other hand, detailed information provided by the images is limited, and typically, customer service systems cannot understand the intent of users without the input text.Thus, bridging the gap between the image and text is crucial for communicating with customers.In this paper, we construct JDDC 2.1, a large-scale multimodal multi-turn dialogue dataset collected from a mainstream Chinese E-commerce platform 1 , containing about 246K dialogue sessions, 3M utterances, and 507K images, along with product knowledge bases and image category annotations.Over our dataset, we jointly define four tasks: the multimodal dialogue response generation task, the multimodal query rewriting task, the multimodal dialogue discourse parsing task, and the multimodal dialogue summarization task.JDDC 2.1 is the first corpus with annotations for all the above tasks over the same dialogue sessions, which facilitates the comprehensive research around the dialogue.In addition, we present several text-only and multimodal baselines and show the importance of visual information for these tasks.Our dataset and implements will be publicly available. Haoran Li 0001, Youzheng Wu, Xiaodong He 0001 |
EMNLP | 3 |
| 2022 | UniRPG: Unified Discrete Reasoning over Table and Text as Program GenerationabstractQuestion answering requiring discrete reasoning, e.g., arithmetic computing, comparison, and counting, over knowledge is a challenging task.In this paper, we propose UniRPG, a semantic-parsing-based approach advanced in interpretability and scalability, to perform Unified discrete Reasoning over heterogeneous knowledge resources, i.e., table and text, as Program Generation.Concretely, UniRPG consists of a neural programmer and a symbolic program executor, where a program is the composition of a set of pre-defined general atomic and higher-order operations and arguments extracted from table and text.First, the programmer parses a question into a program by generating operations and copying arguments, and then, the executor derives answers from table and text based on the program.To alleviate the costly program annotation issue, we design a distant supervision approach for programmer learning, where pseudo programs are automatically constructed without annotated derivations.Extensive experiments on the TAT-QA dataset show that UniRPG achieves tremendous improvements and enhances interpretability and scalability compared with previous state-of-theart methods, even without derivation annotation.Moreover, it achieves promising performance on the textual dataset DROP without derivation annotation. 1 Yongwei Zhou, Junwei Bao 0001, Chaoqun Duan, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao |
EMNLP | 4 |
| 2022 | SCaLa: Supervised Contrastive Learning for End-to-End Speech RecognitionabstractEnd-to-end Automatic Speech Recognition (ASR) models are usually trained to optimize the loss of the whole token sequence, while neglecting explicit phonemic-granularity supervision.This could result in recognition errors due to similarphoneme confusion or phoneme reduction.To alleviate this problem, we propose a novel framework based on Supervised Contrastive Learning (SCaLa) to enhance phonemic representation learning for end-to-end ASR systems.Specifically, we extend the self-supervised Masked Contrastive Predictive Coding (MCPC) to a fully-supervised setting, where the supervision is applied in the following way.First, SCaLa masks variablelength encoder features according to phoneme boundaries given phoneme forced-alignment extracted from a pre-trained acoustic model; it then predicts the masked features via contrastive learning.The forced-alignment can provide phoneme labels to mitigate the noise introduced by positive-negative pairs in selfsupervised MCPC.Experiments on reading and spontaneous speech datasets show that our proposed approach achieves 2.8 and 1.4 points Character Error Rate (CER) absolute reductions compared to the baseline, respectively. Runyu Wang, Fan Lu 0003, Zhengchen Zhang, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001 |
INTERSPEECH | 7 |
| 2022 | LUNA: Learning Slot-Turn Alignment for Dialogue State TrackingabstractYifan Wang, Jing Zhao, Junwei Bao, Chaoqun Duan, Youzheng Wu, Xiaodong He. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yifan Wang 0016, Junwei Bao 0001, Chaoqun Duan, Youzheng Wu, Xiaodong He 0001 |
NAACL-HLT | 5 |
| 2022 | OPERA: Operation-Pivoted Discrete Reasoning over TextabstractYongwei Zhou, Junwei Bao, Chaoqun Duan, Haipeng Sun, Jiahui Liang, Yifan Wang, Jing Zhao, Youzheng Wu, Xiaodong He, Tiejun Zhao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yongwei Zhou, Junwei Bao 0001, Chaoqun Duan, Haipeng Sun, Jiahui Liang, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao |
NAACL-HLT | 8 |
| 2022 | Overview of the NLPCC 2022 Shared Task on Multimodal Product Summarization
Haoran Li 0001, Peng Yuan 0002, Haoning Zhang, Weikang Li, Song Xu 0002, Youzheng Wu, Xiaodong He 0001 |
NLPCC (2) | 6 |
| 2021 | Incremental Learning for End-to-End Automatic Speech RecognitionabstractIn this paper, we propose an incremental learning method for end-to-end Automatic Speech Recognition (ASR) which enables an ASR system to perform well on new tasks while maintaining the performance on its originally learned ones. To mitigate catastrophic forgetting during incremental learning, we design a novel explainability-based knowledge distillation for ASR models, which is combined with a response-based knowledge distillation to maintain the original model's predictions and the “reason” for the predictions. Our method works without access to the training data of original tasks, which addresses the cases where the previous data is no longer available or joint training is costly. Results on a multi-stage sequential training task show that our method outperforms existing ones in mitigating forgetting. Furthermore, in two practical scenarios, compared to the target-reference joint training method, the performance drop of our method is 0.02% Character Error Rate (CER), which is 97% smaller than the drops of the baseline methods. Libo Zi, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ASRU | 5 |
| 2021 | Learn to Copy from the Copying History: Correlational Copy Network for Abstractive SummarizationabstractThe copying mechanism has had considerable success in abstractive summarization, facilitating models to directly copy words from the input text to the output summary.Existing works mostly employ encoder-decoder attention, which applies copying at each time step independently of the former ones.However, this may sometimes lead to incomplete copying.In this paper, we propose a novel copying scheme named Correlational Copying Network (CoCoNet) that enhances the standard copying mechanism by keeping track of the copying history.It thereby takes advantage of prior copying distributions and, at each time step, explicitly encourages the model to copy the input word that is relevant to the previously copied one.In addition, we strengthen CoCoNet through pretraining with suitable corpora that simulate the copying behaviors.Experimental results show that CoCoNet can copy more accurately and achieves new state-of-the-art performances on summarization benchmarks, including CNN/DailyMail for news summarization and SAMSum for dialogue summarization.Our code is available at https:// github.com/hrlinlp/coconet. Haoran Li 0001, Song Xu 0002, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
EMNLP (1) | 5 |
| 2021 | Conversational Query Rewriting with Self-Supervised LearningabstractContext modeling plays a critical role in building multi-turn dialogue systems. Conversational Query Rewriting (CQR) aims to simplify the multi-turn dialogue modeling into a single-turn problem by explicitly rewriting the conversational query into a self-contained utterance. However, existing approaches rely on massive supervised training data, which is labor-intensive to annotate. And the detection of the omitted important information from context can be further improved. Besides, intent consistency constraint between contextual query and rewritten query is also ignored. To tackle these issues, we first propose to construct a large-scale CQR dataset automatically via self-supervised learning, which does not need human annotation. Then we introduce a novel CQR model Teresa based on Transformer, which is enhanced by self-attentive keywords detection and intent consistency constraint. Finally, we conduct extensive experiments on two public datasets. Experimental results demonstrate that our proposed model outperforms existing CQR baselines significantly, and also prove the effectiveness of self-supervised learning on improving the CQR performance. Hang Liu 0005, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 3 |
| 2021 | Dian: Duration Informed Auto-Regressive Network for Voice CloningabstractIn this paper, we propose a novel end-to-end speech synthesis approach, Duration Informed Auto-regressive Network (DIAN), which consists of an acoustic model and a separate duration model. Un-like other auto-regressive TTS methods, the duration information of phonemes is provided as part of the input to the acoustic model, which enables the removal of the attention mechanism between its encoder and decoder parts. This eliminates the common seen skipping and repeating issues and improves speech intelligibility while ensuring high speech quality. A Transformer-based duration model is used to predict the duration of each phoneme for the attention-free acoustic model. We developed our TTS systems for the multi-speaker multi-style voice cloning challenge (M2VoC) using the proposed DIAN approach. In our procedure, a multi-speaker attention-free acoustic model and its Transformer-based duration model are first separately trained based on the training data released by M2VoC. Next, the multi-speaker models are adapted to form the speaker-specific models with the speaker-dependent data and transfer learning. At last, a speaker-specific LPCNet is estimated and used to synthesize the speech of the corresponding speaker. The M2VoC results showed that our proposed approach achieved the 3rd-place in the speech quality ranking and the 4th-place in the speaker similarity and style similarity ranking in the Track1-a task. Zhengchen Zhang, Chao Zhang 0031, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 5 |
| 2021 | SGG: Learning to Select, Guide, and Generate for Keyphrase GenerationabstractJing Zhao, Junwei Bao, Yifan Wang, Youzheng Wu, Xiaodong He, Bowen Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Junwei Bao 0001, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
NAACL-HLT | 4 |
| 2021 | CUSTOM: Aspect-Oriented Product Summarization for E-Commerce
Jiahui Liang, Junwei Bao 0001, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
NLPCC (2) | 4 |
| 2021 | EviDR: Evidence-Emphasized Discrete Reasoning for Reasoning Machine Reading Comprehension
Yongwei Zhou, Junwei Bao 0001, Haipeng Sun, Jiahui Liang, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001, Tiejun Zhao |
NLPCC (1) | 5 |
| 2020 | Aspect-Aware Multimodal Summarization for Chinese E-Commerce ProductsabstractWe present an abstractive summarization system that produces summary for Chinese e-commerce products. This task is more challenging than general text summarization. First, the appearance of a product typically plays a significant role in customers' decisions to buy the product or not, which requires that the summarization model effectively use the visual information of the product. Furthermore, different products have remarkable features in various aspects, such as “energy efficiency” and “large capacity” for refrigerators. Meanwhile, different customers may care about different aspects. Thus, the summarizer needs to capture the most attractive aspects of a product that resonate with potential purchasers. We propose an aspect-aware multimodal summarization model that can effectively incorporate the visual information and also determine the most salient aspects of a product. We construct a large-scale Chinese e-commerce product summarization dataset that contains approximately 1.4 million manually created product summaries that are paired with detailed product information, including an image, a title, and other textual descriptions for each product. The experimental results on this dataset demonstrate that our models significantly outperform the comparative methods in terms of both the ROUGE score and manual evaluations. Haoran Li 0001, Peng Yuan 0002, Song Xu 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
AAAI | 4 |
| 2020 | Self-Attention Guided Copy Mechanism for Abstractive SummarizationabstractCopy module has been widely equipped in the recent abstractive summarization models, which facilitates the decoder to extract words from the source into the summary.Generally, the encoder-decoder attention is served as the copy distribution, while how to guarantee that important words in the source are copied remains a challenge.In this work, we propose a Transformer-based model to enhance the copy mechanism.Specifically, we identify the importance of each source word based on the degree centrality with a directed graph built by the self-attention layer in the Transformer.We use the centrality of each source word to guide the copy process explicitly.Experimental results show that the self-attention graph provides useful guidance for the copy distribution.Our proposed models significantly outperform the baseline methods on the CNN/Daily Mail dataset and the Gigaword dataset. Song Xu 0002, Haoran Li 0001, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ACL | 4 |
| 2020 | Learning to Decouple Relations: Few-Shot Relation Classification with Entity-Guided Attention and Confusion-Aware TrainingabstractThis paper aims to enhance the few-shot relation classification especially for sentences that jointly describe multiple relations.Due to the fact that some relations usually keep high cooccurrence in the same context, previous few-shot relation classifiers struggle to distinguish them with few annotated instances.To alleviate the above relation confusion problem, we propose CTEG, a model equipped with two mechanisms to learn to decouple these easily-confused relations.On the one hand, an Entity-Guided Attention (EGA) mechanism, which leverages the syntactic relations and relative positions between each word and the specified entity pair, is introduced to guide the attention to filter out information causing confusion.On the other hand, a Confusion-Aware Training (CAT) method is proposed to explicitly learn to distinguish relations by playing a pushing-away game between classifying a sentence into a true relation and its confusing relation.Extensive experiments are conducted on the FewRel dataset, and the results show that our proposed model achieves comparable and even much better results to strong baselines in terms of accuracy.Furthermore, the ablation test and case study verify the effectiveness of our proposed EGA and CAT, especially in addressing the relation confusion problem. Yingyao Wang, Junwei Bao 0001, Guangyi Liu 0005, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001, Tiejun Zhao |
COLING | 4 |
| 2020 | On the Faithfulness for E-commerce Product SummarizationabstractIn this work, we present a model to generate e-commerce product summaries.The consistency between the generated summary and the product attributes is an essential criterion for the ecommerce product summarization task.To enhance the consistency, first, we encode the product attribute table to guide the process of summary generation.Second, we identify the attribute words from the vocabulary, and we constrain these attribute words can be presented in the summaries only through copying from the source, i.e., the attribute words not in the source cannot be generated.We construct a Chinese e-commerce product summarization dataset, and the experimental results on this dataset demonstrate that our models significantly improve the faithfulness. Peng Yuan 0002, Haoran Li 0001, Song Xu 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
COLING | 4 |
| 2020 | Learning to Predict Charges for Legal Judgment via Self-Attentive Capsule NetworkabstractWith the rapid development of deep learning technology, more and more traditional industries are changed by Artificial Intelligence. The legal industry is such a popular scenario which attracts lots of researchers' interests. In this work, we focus on automatic charge prediction, which predicts the final charges according to the given fact descriptions in criminal cases. It is crucial for legal assistant systems and can help the judges improve work efficiency greatly. However, extremely imbalanced data distribution and lengthy fact descriptions make this task especially challenging. To tackle these two issues, we propose a novel model, namely Self-Attentive Capsule Network (dubbed as SAttCaps). In particular, we devise a self-attentive dynamic routing, which can not only capture long-range dependency more directly than vanilla dynamic routing, but also learn the high-level generalized features better. The experimental results on three real-world datasets demonstrate that our model significantly outperforms the baselines and creates new state-of-the-art performance. Moreover, our model performs much better than the baselines especially in the low-frequency charges and can bring 5.7% absolute improvement under F1 score. Yuquan Le, Congqing He, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ECAI | 4 |
| 2020 | Multimodal Joint Attribute Prediction and Value Extraction for E-commerce ProductabstractProduct attribute values are essential in many e-commerce scenarios, such as customer service robots, product recommendations, and product retrieval.While in the real world, the attribute values of a product are usually incomplete and vary over time, which greatly hinders the practical applications.In this paper, we propose a multimodal method to jointly predict product attributes and extract values from textual product descriptions with the help of the product images.We argue that product attributes and values are highly correlated, e.g., it will be easier to extract the values on condition that the product attributes are given.Thus, we jointly model the attribute prediction and value extraction tasks from multiple aspects towards the interactions between attributes and values.Moreover, product images have distinct effects on our tasks for different product attributes and values.Thus, we selectively draw useful visual information from product images to enhance our model.We annotate a multimodal product attribute value dataset that contains 87,194 instances, and the experimental results on this dataset demonstrate that explicitly modeling the relationship between attributes and values facilitates our method to establish the correspondence between them, and selectively utilizing visual product information is necessary for the task.Our code and dataset are available 1 . Tiangang Zhu, Haoran Li 0001, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
EMNLP (1) | 4 |
| 2020 | The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer ServiceabstractHuman conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge amount of real conversation data. In this paper, we construct a large-scale real scenario Chinese E-commerce conversation corpus, JDDC, with more than 1 million multi-turn dialogues, 20 million utterances, and 150 million words. The dataset reflects several characteristics of human-human conversations, e.g., goal-driven, and long-term dependency among the context. It also covers various dialogue types including task-oriented, chitchat and question-answering. Extra intent information and three well-annotated challenge sets are also provided. Then, we evaluate several retrieval-based and generative models to provide basic benchmark performance on the JDDC corpus. And we hope JDDC can serve as an effective testbed and benefit the development of fundamental research in dialogue task. Meng Chen 0006, Ruixue Liu, Lei Shen 0001, Shaozu Yuan, Jingyan Zhou, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
LREC | 6 |
| 2019 | Mappa Mundi: An Interactive Artistic Mind Map Generator with Artificial ImaginationabstractWe present a novel real-time, collaborative, and interactive AI painting system, Mappa Mundi, for artistic Mind Map creation. The system consists of a voice-based input interface, an automatic topic expansion module, and an image projection module. The key innovation is to inject Artificial Imagination into painting creation by considering lexical and phonological similarities of language, learning and inheriting artist’s original painting style, and applying the principles of Dadaism and impossibility of improvisation. Our system indicates that AI and artist can collaborate seamlessly to create imaginative artistic painting and Mappa Mundi has been applied in art exhibition in UCCA, Beijing. Ruixue Liu, Baoyang Chen, Meng Chen 0006, Youzheng Wu, Zhijie Qiu, Xiaodong He 0001 |
IJCAI | 4 |
| 2016 | Attention-Based Convolutional Neural Networks for Sentence Classification
Youzheng Wu |
INTERSPEECH | 2 |
| 2015 | Leveraging social Q&A collections for improving complex question answeringabstractThis paper regards social question-and-answer (Q&A) collections such as Yahoo! Answers as knowledge repositories and investigates techniques to mine knowledge from them to improve sentence-based complex question answering (QA) systems. Specifically, we present a question-type-specific method (QTSM) that extracts question-type-dependent cue expressions from social Q&A pairs in which the question types are the same as the submitted questions. We compare our approach with the question-specific and monolingual translation-based methods presented in previous works. The question-specific method (QSM) extracts question-dependent answer words from social Q&A pairs in which the questions resemble the submitted question. The monolingual translation-based method (MTM) learns word-to-word translation probabilities from all of the social Q&A pairs without considering the question or its type. Experiments on the extension of the NTCIR 2008 Chinese test data set demonstrate that our models that exploit social Q&A collections are significantly more effective than baseline methods such as LexRank. The performance ranking of these methods is QTSM > {QSM, MTM}. The largest F3 improvements in our proposed QTSM over QSM and MTM reach 6.0% and 5.8%, respectively. Youzheng Wu, Chiori Hori, Hideki Kashioka, Hisashi Kawai |
Comput. Speech Lang. | 1 |
| 2014 | Recurrent Neural Network-based Tuple Sequence Model for Machine Translation
Youzheng Wu, Taro Watanabe, Chiori Hori |
COLING | 1 |
| 2014 | Translating TED speeches by recurrent neural network based translation modelabstractThis paper presents our recent progress on translating TED speeches1, a collection of public lectures covering a variety of topics. Specially, we use word-to-word alignment to compose translation units of bilingual tuples and present a recurrent neural network-based translation model (RNNTM) to capture long-span context during estimating translation probabilities of bilingual tuples. However, this RNNTM has severe data sparsity problem due to large tuple vocabulary and limited training data. Therefore, a factored RNNTM, which takes bilingual tuples in addition to source and target phrases of the tuples as input features, is proposed to partially address the problem. Our experimental results on the IWSLT2012 test sets show that the proposed models significantly improve the translation quality over state-of-the-art phrase-based translation systems. Youzheng Wu, Xinhui Hu, Chiori Hori |
ICASSP | 1 |
| 2013 | A lecture transcription system combining neural network acoustic and language modelsabstractThis paper presents a new system for automatic transcription of lectures. The system combines a number of novel features, including deep neural network acoustic models using multi-level adaptive networks to incorporate out-of-domain information, and factored recurrent neural network language models. We demonstrate that the system achieves large improvements on the TED lecture transcription task from the 2012 IWSLT evaluation - our results are currently the best reported on this task, showing an relative WER reduction of more than 16% compared to the closest competing system from the evaluation. Peter Bell 0001, Hitoshi Yamamoto, Pawel Swietojanski, Youzheng Wu, Fergus R. McInnes, Chiori Hori, Steve Renals |
INTERSPEECH | 4 |
| 2012 | Factored Language Model based on Recurrent Neural Network
Youzheng Wu, Xugang Lu, Hitoshi Yamamoto, Shigeki Matsuda, Chiori Hori, Hideki Kashioka |
COLING | 1 |
| 2012 | Leveraging Social Annotation for Topic Language Model Adaptation
Youzheng Wu, Kazuhiko Abe, Paul R. Dixon, Chiori Hori, Hideki Kashioka |
INTERSPEECH | 1 |
| 2011 | Improving Related Entity Finding via Incorporating Homepages and Recognizing Fine-grained Entities
Youzheng Wu, Chiori Hori, Hisashi Kawai, Hideki Kashioka |
IJCNLP | 1 |
| 2011 | Answering Complex Questions via Exploiting Social Q&A Collection
Youzheng Wu, Chiori Hori, Hisashi Kawai, Hideki Kashioka |
IJCNLP | 1 |
| 2009 | An Unsupervised Model of Exploiting the Web to Answer Definitional QuestionsabstractIn order to build accurate target profiles, most definition question answering (QA) systems primarily involve utilizing various external resources, such as WordNet, Wikipedia, Biograpy.com, etc. However, these external resources are not always available or helpful when answering definition questions. In contrast, this paper proposes an unsupervised classification model, called the U-Model, which can liberate definitional QA systems from heavily depending on a variety of external resources via applying sentence expansion ($SE$) and SVM classifier. Experimental results from testing on English TREC test sets reveal that the proposed U-Model can not only significantly outperform baseline system but also require no specific external resources. Youzheng Wu, Hideki Kashioka |
Web Intelligence | 1 |
| 2008 | Learning Reliable Information for Dependency Parsing Adaptation
Wenliang Chen, Youzheng Wu, Hitoshi Isahara |
COLING | 2 |
| 2007 | Using Clustering Approaches to Open-Domain Question Answering
Youzheng Wu, Hideki Kashioka, Jun Zhao 0001 |
CICLing | 1 |
| 2007 | Mining redundancy in candidate-bearing snippets to improve web question answeringabstractConventional question answering (QA) techniques independently process candidate-bearing snippets to select an exact answer to a question from candidate answers. This paper presents two novel ways of utilizing redundancy in candidate-bearing snippets to help select an exact answer to a question in our Web QA system, i.e., cluster-based language model (CLM-M) and unsupervised SVM classifier (U-SVM) techniques. The comparative experiments demonstrate that the proposed methods significantly outperform the language model-based (LM-M) and supervised SVM-based (S-SVM) techniques that do not utilize this redundancy in the candidate-bearing snippets. Using the CLM-M, the top_1 score is increased from 36.03% (LM-M) to 46.96%; and the top_1 improvement in the U-SVM over the S-SVM is about 23%. Moreover, a cross-model comparison shows that the performance ranking of these models is: U-SVM > CLM-LM > LM-M > S-SVM > R-M (the retrieval-based model). Youzheng Wu, Xinhui Hu, Hideki Kashioka |
CIKM | 1 |
| 2007 | Learning Unsupervised SVM Classifier for Answer Selection in Web Question Answering
Youzheng Wu, Ruiqiang Zhang, Xinhui Hu, Hideki Kashioka |
EMNLP-CoNLL | 1 |