Xiaodong He 0001

dblp:03/3923-1 · DBLP profile ↗
← Back
199ranked-venue papers
22as first author
64since 2021 · last 2025
0000-0002-9463-9168ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 149 · 13 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 100 · 13 first-author · 32 since 2021Databases, data management, data science and information retrieval · 18 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author
YearPublicationVenuePosition
2025 Comet: Dialog Context Fusion Mechanism for End-to-End Task-Oriented Dialog with Multi-task Learning
abstract
Existing end-to-end task-oriented dialog systems often encounter challenges arising from implicit information, coreference, and the presence of noisy and irrelevant data within the dialog context. These issues hinder the system’s ability to fully comprehend critical information and lead to inaccurate responses. To address these concerns, we propose Comet, a dialog context fusion mechanism for end-to-end task-oriented dialog, augmented with three supplementary tasks: dialog summarization, domain prediction, and slot detection. Dialog summarization facilitates a more comprehensive understanding of important dialog context information by Comet. Domain prediction enables Comet to concentrate on domain-specific information, thus reducing interference from irrelevant information. Slot detection empowers Comet to accurately identify and comprehend essential dialog context information. Additionally, we introduce a data refinement strategy to enhance the comprehensiveness and recommendability of the generated responses. Experimental results demonstrate the superior performance of our proposed methods compared to existing end-to-end task-oriented dialog systems, achieving state-of-the-art results on the MultiWOZ and CrossWOZ datasets.
Haipeng Sun, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001
COLING4
2025 HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation
abstract
Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale videos with accurate captions for HOI. To address this issue, we introduce HOIGen-1M, the first large-scale dataset for HOI Generation, consisting of over one million high-quality videos collected from diverse sources. In particular, to guarantee the high quality of videos, we first design an efficient framework to automatically curate HOI videos using the powerful multimodal large language models (MLLMs), and then the videos are further cleaned by human annotators. Moreover, to obtain accurate textual captions for HOI videos, we design a novel video description method based on a Mixture-of-Multimodal-Experts (MoME) strategy that not only generates expressive captions but also eliminates the hallucination by individual MLLM. Further-more, due to the lack of an evaluation framework for gen-erated HOI videos, we propose two new metrics to assess the quality of generated videos in a coarse-to-fine manner. Extensive experiments reveal that current T2V models struggle to generate high-quality HOI videos and confirm that our HOIGen-1M dataset is instrumental for improving HOI video generation.
Kun Liu 0016, Qi Liu 0081, Xinchen Liu, Yongdong Zhang 0001, Jiebo Luo 0001, Xiaodong He 0001, Wu Liu 0005
CVPR7
2025 Scaling Down Text Encoders of Text-to-Image Diffusion Models
abstract
Text encoders in diffusion models have rapidly evolved, transitioning from CLIP to T5-XXL. Although this evolution has significantly enhanced the models’ ability to understand complex prompts and generate text, it also leads to a substantial increase in the number of parameters. Despite T5 series encoders being trained on the C4 natural language corpus, which includes a significant amount of non-visual data, diffusion models with T5 encoder do not respond to those non-visual prompts, indicating redundancy in representational power. Therefore, it raises an important question: "Do we really need such a large text encoder?" In pursuit of an answer, we employ vision-based knowledge distillation to train a series of T5 encoder models. To fully inherit T5-XXL’s capabilities, we constructed our dataset based on three criteria: image quality, semantic understanding, and text-rendering. Our results demonstrate the scaling down pattern that the distilled T5-base model can generate images of comparable quality to those produced by T5-XXL, while being 50 times smaller in size. This reduction in model size significantly lowers the GPU requirements for running state-of-the-art models such as FLUX and SD3, making high-quality text-to-image generation more accessible. Our code is available at https: //lifuwang-66.github.io/ScalingDownTE/.
Daqing Liu, Xinchen Liu, Xiaodong He 0001
CVPR4
2025 UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition
abstract
Recent advancements in scaling up models have significantly improved performance in Automatic Speech Recognition (ASR) tasks. However, training large ASR models from scratch remains costly. To address this issue, we introduce UME, a novel method that efficiently Upcycles pretrained dense ASR checkpoints into larger Mixture-of-Eperts (MoE) architectures. Initially, feed-forward networks are converted into MoE layers. By reusing the pretrained weights, we establish a robust foundation for the expanded model, significantly reducing optimization time. Then, layer freezing and expert balancing strategies are employed to continue training the model, further enhancing performance. Experiments on a mixture of 170k-hour Mandarin and English datasets show that UME: 1) surpasses the pretrained baseline by a margin of 11.9% relative error rate reduction while maintaining comparable latency; 2) reduces training time by up to 86.7% and achieves superior accuracy compared to training models of the same size from scratch.
Shanyong Yu, Fan Lu 0003, Youzheng Wu, Xiaodong He 0001
ICASSP6
2025 DeepMSD: Advancing Multimodal Sarcasm Detection Through Knowledge-Augmented Graph Reasoning
abstract
Multimodal sarcasm detection (MSD) requires predicting the sarcastic sentiment by understanding diverse modalities of data (e.g., text, image). Beyond the surface-level information conveyed in the post data, understanding the underlying deep-level knowledge-such as the background and intent behind the data-is crucial for understanding the sarcastic sentiment. However, previous works have often overlooked this aspect, limiting their potential to achieve superior performance. To tackle this challenge, we propose DeepMSD, a novel framework that generates supplemental deep-level knowledge to enhance the understanding of sarcastic content. Specifically, we first devise a Deep-level Knowledge Extraction Module that leverages large vision-language models to generate deep-level information behind the text-image pairs. Additionally, we devise a Cross-knowledge Graph Reasoning Module to model how humans use prior knowledge to identify sarcastic cues in multimodal posts. This module constructs cross-knowledge graphs that connect deep-level knowledge with surface-level knowledge. As such, it enables a more profound exploration of the cues underlying sarcasm. Experiments on the public MSD dataset demonstrate that our approach significantly surpasses previous state-of-the-art methods.
Hengyang Zhou, Shaozu Yuan, Meng Chen 0006, Zhiyang Jia, Longbiao Wang, Xiaodong He 0001
IEEE Trans. Circuits Syst. Video Technol.8
2025 Enhancing Semantic Awareness by Sentimental Constraint With Automatic Outlier Masking for Multimodal Sarcasm Detection
abstract
Multimodal sarcasm detection, aiming to uncover sarcastic sentiment behind multimodal data, has gained substantial attention in multimodal communities. Recent advancements in multimodal sarcasm detection (MSD) methods have primarily focused on modality alignment with pre-trained vision-language (V-L) model. However, text-image pairs often exhibit weak or even opposite semantic correlations in MSD tasks. Consequently, directly aligning these modalities can potentially result in feature shift and inter-class confusion, ultimately hindering the model's ability. To alleviate this issue, we propose the Enhancing Semantic Awareness Model (ESAM) for multimodal sarcasm detection. Specifically, we first devise a Modality-decoupled Framework (MDF) to separate the textual and visual features from the fused multimodal representation. This decoupling enables the parallel integration of the Sentimental Congruity Constraint (SCC) within both visual and textual latent spaces, thereby enhancing the semantic awareness of different modalities. Furthermore, given that certain outlier samples with ambiguous sentiments can mislead the training and weaken the performance of SCC, we further incorporate Automatic Outlier Masking. This mechanism automatically detects and masks the outliers, guiding the model to focus on more informative samples during training. Experimental results on two public MSD datasets validate the robustness and superiority of our proposed ESAM model.
Shaozu Yuan, Hengyang Zhou, Qinfu Xu, Meng Chen 0006, Xiaodong He 0001
IEEE Trans. Multim.6
2024 POCE: Primal Policy Optimization with Conservative Estimation for Multi-constraint Offline Reinforcement Learning
abstract
Multi-constraint offline reinforcement learning (RL) promises to learn policies that satisfy both cumulative and state- wise costs from offline datasets. This arrangement provides an effective approach for the widespread appli-cation of RL in high-risk scenarios where both cumulative and state-wise costs need to be considered simulta-neously. However, previously constrained offline RL algorithms are primarily designed to handle single-constraint problems related to cumulative cost, which faces challenges when addressing multi-constraint tasks that involve both cumulative and state-wise costs. In this work, we pro-pose a novel Primal policy Optimization with Conservative Estimation algorithm (POCE) to address the problem of multi-constraint offline RL. Concretely, we reframe the ob-jective of multi-constraint offline RL by introducing the con-cept of Maximum Markov Decision Processes (MMDP). Subsequently, we present a primal policy optimization al-gorithm to confront the multi-constraint problems, which improves the stability and convergence speed of model training. Furthermore, we propose a conditional Bell-man operator to estimate cumulative and state-wise Q-values, reducing the extrapolation error caused by out-of-distribution (OOD) actions. Finally, extensive experiments demonstrate that the POCE algorithm achieves competitive performance across multiple experimental tasks, particu-larly outperforming baseline algorithms in terms of safety. Our code is available at github. POCE.
Jiayi Guan, Li Shen 0008, Ao Zhou 0005, Lusong Li, Han Hu 0003, Xiaodong He 0001, Guang Chen 0001, Changjun Jiang 0002
CVPR6
2024 Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld
abstract
While large language models (LLMs) excel in a simulated world of texts, they struggle to interact with the more realistic world without perceptions of other modalities such as visual or audio signals. Although vision-language models (VLMs) integrate LLM modules (1) aligned with static image features, and (2) may possess prior knowledge of world dynamics (as demonstrated in the text world), they have not been trained in an embodied visual world and thus cannot align with its dynamics. On the other hand, training an embodied agent in a noisy visual world without expert guidance is often chal-lenging and inefficient. In this paper, we train a VLM agent living in a visual world using an LLM agent excelling in a parallel text world. Specifically, we distill LLM's reflection outcomes (improved actions by analyzing mistakes) in a text world's tasks to finetune the VLM on the same tasks of the visual world, resulting in an Embodied Multi-Modal Agent (EMMA) quickly adapting to the visual world dy-namics. Such cross-modality imitation learning between the two parallel worlds is achieved by a novel DAgger-DPO algorithm, enabling EMMA to generalize to a broad scope of new tasks without any further guidance from the LLM expert. Extensive evaluations on the ALFWorld benchmark's diverse tasks highlight EMMA's superior performance to SOTA VLM-based agents, e.g., 20%-70% improvement in the success rate.
Tianyi Zhou 0001, Kanxue Li, Dapeng Tao, Lusong Li, Li Shen 0008, Xiaodong He 0001, Jing Jiang 0002, Yuhui Shi 0001
CVPR7
2024 MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models
Kanxue Li, Baosheng Yu, Yibing Zhan, Qiong Cao, Li Shen 0008, Lusong Li, Dapeng Tao, Xiaodong He 0001
IJCAI14
2024 An efficient confusing choices decoupling framework for multi-choice tasks over texts
Yingyao Wang, Junwei Bao 0001, Chaoqun Duan, Youzheng Wu, Xiaodong He 0001, Conghui Zhu, Tiejun Zhao
Neural Comput. Appl.5
2024 Operation-Augmented Numerical Reasoning for Question Answering
abstract
Question answering requiring numerical reasoning, which generally involves symbolic operations such as sorting, counting, and addition, is a challenging task. To address such a problem, existing mixture-of-experts (MoE)-based methods design several specific answer predictors to handle different types of questions and achieve promising performance. However, they ignore the modeling and exploitation of fine-grained reasoning-related operations to support numerical reasoning, encountering the inadequacy in reasoning capability and interpretability. To alleviate this issue, we propose OPERA, an operation-augmented numerical reasoning framework. Concretely, we systematically define a scalable operation set to model numerical reasoning. We first identify reasoning-related operations based on context and then softly execute them to imitate the answer reasoning procedure via an operation-aware cross-attention mechanism. Finally, we utilize the operation-augmented semantic representation of execution results to support answer prediction. We verify the effectiveness and generalization of OPERA in two scenarios with different knowledge sources and reasoning capabilities. Specifically, we conduct extensive experiments on two textual datasets, DROP and RACENum, and a table-text hybrid dataset TAT-QA. Experiment results show that OPERA outperforms previous strong methods on the DROP, RACENum, and TAT-QA datasets. Further, we statistically and visually analyze its interpretability.
Yongwei Zhou, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 MuJo-SF: Multimodal Joint Slot Filling for Attribute Value Prediction of E-Commerce Commodities
abstract
Supplementing product attribute information is a critical step for E-commerce platforms, which further benefits various downstream tasks, including product recommendation, product search, and product knowledge graph construction. Intuitively, the visual information available on e-commerce platforms can effectively function as a primary source for certain product attributes. However, existing works either extract attribute values solely from textual product descriptions or leverage limited visual information (e.g., image features or optical character recognition tokens) to assist extraction, without mining the fine-grained visual cues linked with the products effectively. In this paper, we propose a novel task -Multimodal Joint Slot Filling(MuJo-SF) - that aims to combine multimodal information from both product descriptions and their corresponding product images to jointly fill values into the pre-defined product attribute set. To this end, we develop MAVP, a new dataset with 79 k instances of product description-image pairs. Specifically, we present a strategy to fulfill visualized saliency ascription, which aims to distinguish between text-dependent and image-dependent attributes. For those image-dependent attributes, we annotate the corresponding values from images using distant supervision. Then, we design a model for MuJo-SF, which combines multimodal representations and fills image-dependent and text-dependent attributes separately. Finally, we conduct extensive experiments on MAVP and provide rich results for MuJo-SF, which can be used as baselines to facilitate future research.
Meihuizi Jia, Lei Shen 0001, Anh Tuan Luu, Meng Chen 0006, Lejian Liao, Shaozu Yuan, Xiaodong He 0001
IEEE Trans. Multim.8
2023 MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query Grounding
abstract
Multimodal named entity recognition (MNER) is a critical step in information extraction, which aims to detect entity spans and classify them to corresponding entity types given a sentence-image pair. Existing methods either (1) obtain named entities with coarse-grained visual clues from attention mechanisms, or (2) first detect fine-grained visual regions with toolkits and then recognize named entities. However, they suffer from improper alignment between entity types and visual regions or error propagation in the two-stage manner, which finally imports irrelevant visual information into texts. In this paper, we propose a novel end-to-end framework named MNER-QG that can simultaneously perform MRC-based multimodal named entity recognition and query grounding. Specifically, with the assistance of queries, MNER-QG can provide prior knowledge of entity types and visual regions, and further enhance representations of both text and image. To conduct the query grounding task, we provide manual annotations and weak supervisions that are obtained via training a highly flexible visual grounding model with transfer learning. We conduct extensive experiments on two public MNER datasets, Twitter2015 and Twitter2017. Experimental results show that MNER-QG outperforms the current state-of-the-art models on the MNER task, and also improves the query grounding performance.
Meihuizi Jia, Lei Shen 0001, Lejian Liao, Meng Chen 0006, Xiaodong He 0001
AAAI6
2023 DiffusEmp: A Diffusion Model-Based Framework with Multi-Grained Control for Empathetic Response Generation
abstract
Empathy is a crucial factor in open-domain conversations, which naturally shows one's caring and understanding to others.Though several methods have been proposed to generate empathetic responses, existing works often lead to monotonous empathy that refers to generic and safe expressions.In this paper, we propose to use explicit control to guide the empathy expression and design a framework DIFFUSEMP based on conditional diffusion language model to unify the utilization of dialogue context and attribute-oriented control signals.Specifically, communication mechanism, intent, and semantic frame are imported as multi-grained signals that control the empathy realization from coarse to fine levels.We then design a specific masking strategy to reflect the relationship between multi-grained signals and response tokens, and integrate it into the diffusion model to influence the generative process.Experimental results on a benchmark dataset EMPA-THETICDIALOGUE show that our framework outperforms competitive baselines in terms of controllability, informativeness, and diversity without the loss of context-relatedness.
Guanqun Bi, Lei Shen 0001, Yanan Cao 0001, Meng Chen 0006, Yuqiang Xie, Zheng Lin 0001, Xiaodong He 0001
ACL (1)7
2023 Dialog-Post: Multi-Level Self-Supervised Objectives and Hierarchical Model for Dialogue Post-Training
abstract
Dialogue representation and understanding aim to convert conversational inputs into embeddings and fulfill discriminative tasks.Compared with free-form text, dialogue has two important characteristics, hierarchical semantic structure and multi-facet attributes.Therefore, directly applying the pretrained language models (PLMs) might result in unsatisfactory performance.Recently, several work focused on the dialogue-adaptive post-training (Dial-Post) that further trains PLMs to fit dialogues.To model dialogues more comprehensively, we propose a DialPost method, DIALOG-POST, with multi-level self-supervised objectives and a hierarchical model.These objectives leverage dialogue-specific attributes and use selfsupervised signals to fully facilitate the representation and understanding of dialogues.The novel model is a hierarchical segment-wise self-attention network, which contains innersegment and inter-segment self-attention sublayers followed by an aggregation and updating module.To evaluate the effectiveness of our methods, we first apply two public datasets for the verification of representation ability.Then we conduct experiments on a newly-labelled dataset that is annotated with 4 dialogue understanding tasks.Experimental results show that our method outperforms existing SOTA models and achieves a 3.3% improvement on average. Token-level SSOs Utterance-level SSO Dialogue-level SSOs𝑄: 这个手机支持5G吗?𝑄:
Zhenyu Zhang 0029, Lei Shen 0001, Meng Chen 0006, Xiaodong He 0001
ACL (1)5
2023 POSPAN: Position-Constrained Span Masking for Language Model Pre-training
abstract
Span-level masked language modeling (MLM) has shown to be advantageous to pre-trained language models over the original single-token MLM, as entities/phrases and their dependencies are critical to language understanding. Previous works only consider span length with some discrete distributions, while the dependencies among spans are ignored, i.e., assuming that the positions of masked spans are uniformly distributed. In this paper, we present POSPAN, a general framework to allow diverse position-constrained span masking strategies via the combination of span length distribution and position constraint distribution, which unifies all existing span-level masking methods. To verify the effectiveness of POSPAN in pre-training, we evaluate it on the datasets from several NLU benchmarks. Experimental results indicate that the position constraint is capable of enhancing span-level masking broadly, and our best POSPAN setting consistently outperforms its span-length-only counterparts and vanilla MLM. We also conduct theoretical analysis for the position constraint in masked language models to shed light on the reason why POSPAN works well, demonstrating the rationality and necessity of POSPAN.
Zhenyu Zhang 0029, Lei Shen 0001, Meng Chen 0006, Xiaodong He 0001
CIKM5
2023 Composable Text Controls in Latent Space with ODEs
abstract
Guangyi Liu, Zeyu Feng, Yuan Gao, Zichao Yang, Xiaodan Liang, Junwei Bao, Xiaodong He, Shuguang Cui, Zhen Li, Zhiting Hu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Guangyi Liu 0005, Zeyu Feng, Xiaodan Liang, Junwei Bao 0001, Xiaodong He 0001, Shuguang Cui, Zhen Li 0026, Zhiting Hu
EMNLP7
2023 UFO2: A Unified Pre-Training Framework for Online and Offline Speech Recognition
abstract
In this paper, we propose a Unified pre-training Framework for Online and Offline (UFO2) Automatic Speech Recognition (ASR), which 1) simplifies the two separate training workflows for online and offline modes into one process, and 2) improves the Word Error Rate (WER) performance with limited utterance annotating. Specifically, we extend the conventional offline-mode Self-Supervised Learning (SSL)-based ASR approach to a unified manner, where the model training is conditioned on both the full-context and dynamic- chunked inputs. To enhance the pre-trained representation model, stop-gradient operation is applied to decouple the online-mode objectives to the quantizer. Moreover, in both the pre-training and the downstream fine-tuning stages, joint losses are proposed to train the unified model with full-weight sharing for the two modes. Experimental results on the LibriSpeech dataset show that UFO2 outperforms the SSL-based baseline method by 29.7% and 18.2% relative WER reduction in offline and online modes, respectively.
Qingtao Li, Fangzhu Li, Fan Lu 0003, Meng Chen 0006, Xiaodong He 0001
ICASSP8
2023 Improving Disfluency Detection with Multi-Scale Self Attention and Contrastive Learning
abstract
Disfluency detection aims to recognize disfluencies in sentences. Existing works usually adopt a sequence labeling model to tackle this task. They also attempt to integrate into models the feature that the disfluencies are similar to the correct phrase, the so-called "rough copy". However, they heavily rely on hand-craft features or word-to-word match patterns, which are insufficient to precisely capture such rough copy and cause under-tagging and over-tagging problems. To alleviate these problems, we propose a multi-scale self-attention mechanism (MSAT) and design contrastive learning (CL) loss for this task. Specifically, the MSAT leverages token representations to learn representations for different scales of phrases, and then compute similarity among them. The CL adopts the fluent version of the input to build the positive and negative samples and encourages the model to keep the fluent version consistent with the input in semantics. We conduct experiments on a public English dataset Switchboard, and an in-house Chinese dataset Waihu, which is derived from an online conversation bot. Results show that our method outperforms the baselines and achieves superior performance on both datasets.
Peiying Wang, Chaoqun Duan, Meng Chen 0006, Xiaodong He 0001
ICASSP4
2023 SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation
abstract
Recently, the contrastive language-image pre-training, e.g., CLIP, has demonstrated promising results on various downstream tasks. The pre-trained model can capture enriched visual concepts for images by learning from a large scale of text-image data. However, transferring the learned visual knowledge to open-vocabulary semantic segmentation is still under-explored. In this paper, we propose a CLIP-based model named SegCLIP for the topic of open-vocabulary segmentation in an annotation-free manner. The SegCLIP achieves segmentation based on ViT and the main idea is to gather patches with learnable centers to semantic regions through training on text-image pairs. The gathering operation can dynamically capture the semantic groups, which can be used to generate the final segmentation results. We further propose a reconstruction loss on masked patches and a superpixel-based KL loss with pseudo-labels to enhance the visual representation. Experimental results show that our model achieves comparable or superior segmentation accuracy on the PASCAL VOC 2012 (+0.3% mIoU), PASCAL Context (+2.3% mIoU), and COCO (+2.2% mIoU) compared with baselines. We release the code at https://github.com/ArrowLuo/SegCLIP.
Huaishao Luo, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001, Tianrui Li 0001
ICML4
2023 OTF: Optimal Transport based Fusion of Supervised and Self-Supervised Learning Models for Automatic Speech Recognition
Qingtao Li, Fangzhu Li, Fan Lu 0003, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001
INTERSPEECH9
2023 Leveraging Label Information for Multimodal Emotion Recognition
Peiying Wang, Sunlu Zeng, Fan Lu 0003, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001
INTERSPEECH7
2023 MaskedSpeech: Context-aware Speech Synthesis with Masking Strategy
Ya-Jie Zhang, Yanghao Yue, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001
INTERSPEECH6
2023 Prosody Modelling With Pre-Trained Cross-Utterance Representations for Improved Speech Synthesis
abstract
When humans speak multiple utterances in a continuous manner, the prosodic features generated in each utterance are related to those in its neighbouring utterances. Such cross-utterance (CU) dependencies are often ignored by the current neural text-to-speech (TTS) systems, which reduces the naturalness and expressiveness of the synthesized speeches. In this paper, we propose to improve the prosody modelling ability of neural TTS systems using pre-trained CU acoustic and text representations. Such CU acoustic representations are derived using the Wav2Vec 2.0 model (W2V2) from the synthesized audios of the past utterances, while the CU text representations are extracted using the Bidirectional Encoder Representation from Transformers (BERT) model from the scripts of the future utterances. Experimental results on a Mandarin audiobook and an English audiobook showed the naturalness and expressiveness of the synthesized audios were significantly improved by incorporating such pre-trained W2V2 and BERT CU representations into the Fastspeech2 TTS framework.
Ya-Jie Zhang, Chao Zhang 0031, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2022 SimCTC: A Simple Contrast Learning Method of Text Clustering (Student Abstract)
abstract
This paper presents SimCTC, a simple contrastive learning (CL) framework that greatly advances the state-of-the-art text clustering models. In SimCTC, a pre-trained BERT model first maps the input sequence to the representation space, which is then followed by three different loss function heads: Clustering head, Instance-CL head and Cluster-CL head. Experimental results on multiple benchmark datasets demonstrate that SimCTC remarkably outperforms 6 competitive text clustering methods with 1%-6% improvement on Accuracy (ACC) and 1%-4% improvement on Normalized Mutual Information (NMI). Moreover, our results also show that the clustering performance can be further improved by setting an appropriate number of clusters in the cluster-level objective.
Chen Li 0021, Xiaoguang Yu, Shuangyong Song, Xiaodong He 0001
AAAI6
2022 A Multi-Factor Classification Framework for Completing Users' Fuzzy Queries (Student Abstract)
abstract
Intent identification is the key technology in dialogue system. However, not all online queries are clear or complete. To identify users' intents from those fuzzy queries accurately, this paper proposes a multi-factor classification framework on the query level. Experimental results on our online serving system JIMI demonstrate the effectiveness of our proposed framework.
Liangqing Wu, Xiaoguang Yu, Shuangyong Song, Youzheng Wu, Xiaodong He 0001
AAAI8
2022 Fine- and Coarse-Granularity Hybrid Self-Attention for Efficient BERT
abstract
Transformer-based pre-trained models, such as BERT, have shown extraordinary success in achieving state-of-the-art results in many natural language processing applications.However, deploying these models can be prohibitively costly, as the standard self-attention mechanism of the Transformer suffers from quadratic computational cost in the input sequence length.To confront this, we propose FCA, a fine-and coarse-granularity hybrid self-attention that reduces the computation cost through progressively shortening the computational sequence length in self-attention.Specifically, FCA conducts an attention-based scoring strategy to determine the informativeness of tokens at each layer.Then, the informative tokens serve as the fine-granularity computing units in selfattention and the uninformative tokens are replaced with one or several clusters as the coarsegranularity computing units in self-attention.Experiments on GLUE and RACE datasets show that BERT with FCA achieves 2x reduction in FLOPs over original BERT with <1% loss in accuracy.We show that FCA offers significantly better trade-off between accuracy and FLOPs compared to prior methods 1 .
Yifan Wang 0016, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001
ACL (1)5
2022 Legal Charge Prediction via Bilinear Attention Network
abstract
The legal charge prediction task aims to judge appropriate charges according to the given fact description in cases. Most existing methods formulate it as a multi-class text classification problem and have achieved tremendous progress. However, the performance on low-frequency charges is still unsatisfactory. Previous studies indicate leveraging the charge label information can facilitate this task, but the approaches to utilizing the label information are not fully explored. In this paper, inspired by the vision-language information fusion techniques in the multi-modal field, we propose a novel model (denoted as LeapBank) by fusing the representations of text and labels to enhance the legal charge prediction task. Specifically, we devise a representation fusion block based on the bilinear attention network to interact the labels and text tokens seamlessly. Extensive experiments are conducted on three real-world datasets to compare our proposed method with state-of-the-art models. Experimental results show that LeapBank obtains up to 8.5% Macro-F1 improvements on the low-frequency charges, demonstrating our model's superiority and competitiveness.
Yuquan Le, Meng Chen 0006, Zhe Quan, Xiaodong He 0001, Kenli Li 0001
CIKM5
2022 AutoQGS: Auto-Prompt for Low-Resource Knowledge-based Question Generation from SPARQL
abstract
This study investigates the task of knowledge-based question generation (KBQG). Conventional KBQG works generated questions from fact triples in the knowledge graph, which could not express complex operations like aggregation and comparison in SPARQL. Moreover, due to the costly annotation of large-scale SPARQL-question pairs, KBQG from SPARQL under low-resource scenarios urgently needs to be explored. Recently, since the generative pre-trained language models (PLMs) typically trained in natural language (NL)-to-NL paradigm have been proven effective for low-resource generation, e.g., T5 and BART, how to effectively utilize them to generate NL-question from non-NL SPARQL is challenging. To address these challenges, AutoQGS, an auto-prompt approach for low-resource KBQG from SPARQL, is proposed. Firstly, we put forward to generate questions directly from SPARQL for KBQG task to handle complex operations. Secondly, we propose an auto-prompter trained on large-scale unsupervised data to rephrase SPARQL into NL description, smoothing the low-resource transformation from non-NL SPARQL to NL question with PLMs. Experimental results on the WebQuestionsSP, ComlexWebQuestions 1.1, and PathQuestions show that our model achieves state-of-the-art performance, especially in low-resource settings. Furthermore, a corpora of 330k factoid complex question-SPARQL pairs is generated for further KBQG research.
Guanming Xiong, Junwei Bao 0001, Wen Zhao 0008, Youzheng Wu, Xiaodong He 0001
CIKM5
2022 Few-Shot Table Understanding: A Benchmark Dataset and Pre-Training Baseline
abstract
Few-shot table understanding is a critical and challenging problem in real-world scenario as annotations over large amount of tables are usually costly. Pre-trained language models (PLMs), which have recently flourished on tabular data, have demonstrated their effectiveness for table understanding tasks. However, few-shot table understanding is rarely explored due to the deficiency of public table pre-training corpus and well-defined downstream benchmark tasks, especially in Chinese. In this paper, we establish a benchmark dataset, FewTUD, which consists of 5 different tasks with human annotations to systematically explore the few-shot table understanding in depth. Since there is no large number of public Chinese tables, we also collect a large-scale, multi-domain tabular corpus to facilitate future Chinese table pre-training, which includes one million tables and related natural language text with auxiliary supervised interaction signals. Finally, we present FewTPT, a novel table PLM with rich interactions over tabular data, and evaluate its performance comprehensively on the benchmark. Our dataset and model will be released to the public soon.
Ruixue Liu, Shaozu Yuan, Aijun Dai, Lei Shen 0001, Tiangang Zhu, Meng Chen 0006, Xiaodong He 0001
COLING7
2022 Tracking Satisfaction States for Customer Satisfaction Prediction in E-commerce Service Chatbots
abstract
Due to the increasing use of service chatbots in E-commerce platforms in recent years, customer satisfaction prediction (CSP) is gaining more and more attention. CSP is dedicated to evaluating subjective customer satisfaction in conversational service and thus helps improve customer service experience. However, previous methods focus on modeling customer-chatbot interaction across different turns, which are hard to represent the important dynamic satisfaction states throughout the customer journey. In this work, we investigate the problem of satisfaction states tracking and its effects on CSP in E-commerce service chatbots. To this end, we propose a dialogue-level classification model named DialogueCSP to track satisfaction states for CSP. In particular, we explore a novel two-step interaction module to represent the dynamic satisfaction states at each turn. In order to capture dialogue-level satisfaction states for CSP, we further introduce dialogue-aware attentions to integrate historical informative cues into the interaction module. To evaluate the proposed approach, we also build a Chinese E-commerce dataset for CSP. Experiment results demonstrate that our model significantly outperforms multiple baselines, illustrating the benefits of satisfaction states tracking on CSP.
Liangqing Wu, Shuangyong Song, Xiaoguang Yu, Xiaodong He 0001, Guohong Fu
COLING5
2022 Beyond QA: 'Heuristic QA' Strategies in JIMI
Shuangyong Song, Jianghua Lin, Xiaoguang Yu, Xiaodong He 0001
DASFAA (3)5
2022 PRINCE: Prefix-Masked Decoding for Knowledge Enhanced Sequence-to-Sequence Pre-Training
abstract
Pre-trained Language Models (PLMs) have shown effectiveness in various Natural Language Processing (NLP) tasks.Denoising autoencoder is one of the most successful pretraining frameworks, learning to recompose the original text given a noise-corrupted one.The existing studies mainly focus on injecting noises into the input.This paper introduces a simple yet effective pre-training paradigm, equipped with a knowledge-enhanced decoder that predicts the next entity token with noises in the prefix, explicitly strengthening the representation learning of entities that span over multiple input tokens.Specifically, when predicting the next token within an entity, we feed masks into the prefix in place of some of the previous ground-truth tokens that constitute the entity.Our model achieves new state-of-the-art results on two knowledge-driven data-to-text generation tasks with up to 2% BLEU gains.
Song Xu 0002, Haoran Li 0001, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001
EMNLP5
2022 Correctable-DST: Mitigating Historical Context Mismatch between Training and Inference for Improved Dialogue State Tracking
abstract
Hongyan Xie, Haoxiang Su, Shuangyong Song, Hao Huang, Bo Zou, Kun Deng, Jianghua Lin, Zhihui Zhang, Xiaodong He. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Hongyan Xie, Haoxiang Su, Shuangyong Song, Jianghua Lin, Xiaodong He 0001
EMNLP9
2022 JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and Summarization
abstract
The popularity of multimodal dialogue has stimulated the need for a new generation of dialogue agents with multimodal interactivity.When users communicate with customer service, they may express their requirements by means of text, images, or even videos.Visual information usually acts as discriminators for product models, or indicators of product failures, which play an important role in the Ecommerce scenario.On the other hand, detailed information provided by the images is limited, and typically, customer service systems cannot understand the intent of users without the input text.Thus, bridging the gap between the image and text is crucial for communicating with customers.In this paper, we construct JDDC 2.1, a large-scale multimodal multi-turn dialogue dataset collected from a mainstream Chinese E-commerce platform 1 , containing about 246K dialogue sessions, 3M utterances, and 507K images, along with product knowledge bases and image category annotations.Over our dataset, we jointly define four tasks: the multimodal dialogue response generation task, the multimodal query rewriting task, the multimodal dialogue discourse parsing task, and the multimodal dialogue summarization task.JDDC 2.1 is the first corpus with annotations for all the above tasks over the same dialogue sessions, which facilitates the comprehensive research around the dialogue.In addition, we present several text-only and multimodal baselines and show the importance of visual information for these tasks.Our dataset and implements will be publicly available.
Haoran Li 0001, Youzheng Wu, Xiaodong He 0001
EMNLP4
2022 UniRPG: Unified Discrete Reasoning over Table and Text as Program Generation
abstract
Question answering requiring discrete reasoning, e.g., arithmetic computing, comparison, and counting, over knowledge is a challenging task.In this paper, we propose UniRPG, a semantic-parsing-based approach advanced in interpretability and scalability, to perform Unified discrete Reasoning over heterogeneous knowledge resources, i.e., table and text, as Program Generation.Concretely, UniRPG consists of a neural programmer and a symbolic program executor, where a program is the composition of a set of pre-defined general atomic and higher-order operations and arguments extracted from table and text.First, the programmer parses a question into a program by generating operations and copying arguments, and then, the executor derives answers from table and text based on the program.To alleviate the costly program annotation issue, we design a distant supervision approach for programmer learning, where pseudo programs are automatically constructed without annotated derivations.Extensive experiments on the TAT-QA dataset show that UniRPG achieves tremendous improvements and enhances interpretability and scalability compared with previous state-of-theart methods, even without derivation annotation.Moreover, it achieves promising performance on the textual dataset DROP without derivation annotation. 1
Yongwei Zhou, Junwei Bao 0001, Chaoqun Duan, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao
EMNLP5
2022 Building Robust Spoken Language Understanding by Cross Attention Between Phoneme Sequence and ASR Hypothesis
abstract
Building Spoken Language Understanding (SLU) robust to Automatic Speech Recognition (ASR) errors is an essential issue for various voice-enabled virtual assistants. Considering that most ASR errors are caused by phonetic confusion between similar-sounding expressions, intuitively, leveraging the phoneme sequence of speech can complement ASR hypothesis and enhance the robustness of SLU. This paper proposes a novel model with Cross Attention for SLU (denoted as CASLU). The cross attention block is devised to catch the fine-grained interactions between phoneme and word embeddings in order to make the joint representations catch the phonetic and semantic features of input simultaneously and for overcoming the ASR errors in downstream natural language understanding (NLU) tasks. Extensive experiments are conducted on three datasets, showing the effectiveness and competitiveness of our approach. Additionally, We also validate the universality of CASLU and prove its complementarity when combining with other robust SLU techniques.
Zexun Wang, Yuquan Le, Mingchao Feng, Meng Chen 0006, Xiaodong He 0001
ICASSP7
2022 Gated Multimodal Fusion with Contrastive Learning for Turn-Taking Prediction in Human-Robot Dialogue
abstract
Turn-taking, aiming to decide when the next speaker can start talking, is an essential component in building human-robot spoken dialogue systems. Previous studies indicate that multi-modal cues can facilitate this challenging task. However, due to the paucity of public multimodal datasets, current methods are mostly limited to either utilizing unimodal features or simplistic multimodal ensemble models. Besides, the inherent class imbalance in real scenario, e.g. sentence ending with short pause will be mostly regarded as the end of turn, also poses great challenge to the turn-taking decision. In this paper, we first collect a large-scale annotated corpus for turn-taking with over 5,000 real human-robot dialogues in speech and text modalities. Then, a novel gated multimodal fusion mechanism is devised to utilize various information seamlessly for turn-taking prediction. More importantly, to tackle the data imbalance issue, we design a simple yet effective data augmentation method to construct negative instances without supervision and apply contrastive learning to obtain better feature representations. Extensive experiments are conducted and the results demonstrate the superiority and competitiveness of our model over several state-of-the-art baselines.
Jiudong Yang, Peiying Wang, Mingchao Feng, Meng Chen 0006, Xiaodong He 0001
ICASSP6
2022 SE-GAN: Skeleton Enhanced Gan-Based Model for Brush Handwriting Font Generation
abstract
Previous works on font generation mainly focus on the standard print fonts where character's shape is stable and strokes are clearly separated. There is rare research on brush hand-writing font generation, which involves holistic structure changes and complex strokes transfer. To address this issue, we propose a novel GAN-based image translation model by integrating the skeleton information. We first extract the skeleton from training images, then design an image encoder and a skeleton encoder to extract corresponding features. A self-attentive refined attention module is devised to guide the model to learn distinctive features between different domains. A skeleton discriminator is involved to first synthesize the skeleton image from the generated image with a pre-trained generator, then to judge its realness to the target one. We also contribute a large-scale brush handwriting font image dataset with six styles and 15,000 high-resolution images. Both quantitative and qualitative experimental results demonstrate the competitiveness of our proposed model.
Shaozu Yuan, Ruixue Liu, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001
ICME6
2022 Cross-modal Contrastive Distillation for Instructional Activity Anticipation
abstract
In this study, we aim to predict the plausible future action steps given an observation of the past and study the task of instructional activity anticipation. Unlike previous anticipation tasks that aim at action label prediction, our work targets at generating natural language outputs that provide interpretable and accurate descriptions of future action steps. It is a challenging task due to the lack of semantic information extracted from the instructional videos. To overcome this challenge, we propose a novel knowledge distillation framework to exploit the related external textual knowledge to assist the visual anticipation task. However, previous knowledge distillation techniques generally transfer information within the same modality. To bridge the gap between the visual and text modalities during the distillation process, we devise a novel cross-modal contrastive distillation (CCD) scheme, which facilitates knowledge distillation between teacher and student in heterogeneous modalities with the proposed cross-modal distillation loss. We evaluate our method on the Tasty Videos dataset. CCD improves the anticipation performance of the visual-alone student model by a large margin of 40.2% relatively in BLEU4. Our approach also outperforms the state-of-the-art approaches by a large margin.
Zhengyuan Yang, Jingen Liu, Jing Huang 0019, Xiaodong He 0001, Tao Mei 0001, Chenliang Xu, Jiebo Luo 0001
ICPR4
2022 Learning to Generate Poetic Chinese Landscape Painting with Calligraphy
abstract
In this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generation, Polaca takes the classic poetry as input and outputs the artistic landscape painting image with the corresponding calligraphy. It is equipped with three different modules to complete the whole piece of landscape painting artwork: the first one is a text-to-image module to generate landscape painting image, the second one is an image-to-image module to generate stylistic calligraphy image, and the third one is an image fusion module to fuse the two images into a whole piece of aesthetic artwork.
Shaozu Yuan, Aijun Dai, Zhiling Yan, Ruixue Liu, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001
IJCAI8
2022 SCaLa: Supervised Contrastive Learning for End-to-End Speech Recognition
abstract
End-to-end Automatic Speech Recognition (ASR) models are usually trained to optimize the loss of the whole token sequence, while neglecting explicit phonemic-granularity supervision.This could result in recognition errors due to similarphoneme confusion or phoneme reduction.To alleviate this problem, we propose a novel framework based on Supervised Contrastive Learning (SCaLa) to enhance phonemic representation learning for end-to-end ASR systems.Specifically, we extend the self-supervised Masked Contrastive Predictive Coding (MCPC) to a fully-supervised setting, where the supervision is applied in the following way.First, SCaLa masks variablelength encoder features according to phoneme boundaries given phoneme forced-alignment extracted from a pre-trained acoustic model; it then predicts the masked features via contrastive learning.The forced-alignment can provide phoneme labels to mitigate the noise introduced by positive-negative pairs in selfsupervised MCPC.Experiments on reading and spontaneous speech datasets show that our proposed approach achieves 2.8 and 1.4 points Character Error Rate (CER) absolute reductions compared to the baseline, respectively.
Runyu Wang, Fan Lu 0003, Zhengchen Zhang, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001
INTERSPEECH8
2022 Cross-modal Transfer Learning via Multi-grained Alignment for End-to-End Spoken Language Understanding
Zexun Wang, Hang Liu 0005, Peiying Wang, Mingchao Feng, Meng Chen 0006, Xiaodong He 0001
INTERSPEECH7
2022 E-ConvRec: A Large-Scale Conversational Recommendation Dataset for E-Commerce Customer Service
abstract
There has been a growing interest in developing conversational recommendation system (CRS), which provides valuable recommendations to users through conversations. Compared to the traditional recommendation, it advocates wealthier interactions and provides possibilities to obtain users’ exact preferences explicitly. Nevertheless, the corresponding research on this topic is limited due to the lack of broad-coverage dialogue corpus, especially real-world dialogue corpus. To handle this issue and facilitate our exploration, we construct E-ConvRec, an authentic Chinese dialogue dataset consisting of over 25k dialogues and 770k utterances, which contains user profile, product knowledge base (KB), and multiple sequential real conversations between users and recommenders. Next, we explore conversational recommendation in a real scene from multiple facets based on the dataset. Therefore, we particularly design three tasks: user preference recognition, dialogue management, and personalized recommendation. In the light of the three tasks, we establish baseline results on E-ConvRec to facilitate future studies.
Meihuizi Jia, Ruixue Liu, Peiying Wang, Yang Song 0008, Zexi Xi, Haobin Li, Meng Chen 0006, Jinhui Pang, Xiaodong He 0001
LREC10
2022 Query Prior Matters: A MRC Framework for Multimodal Named Entity Recognition
abstract
Multimodal named entity recognition (MNER) is a vision-language task where the system is required to detect entity spans and corresponding entity types given a sentence-image pair. Existing methods capture text-image relations with various attention mechanisms that only obtain implicit alignments between entity types and image regions. To locate regions more accurately and better model cross-/within-modal relations, we propose a machine reading comprehension based framework for MNER, namely MRC-MNER. By utilizing queries in MRC, our framework can provide prior information about entity types and image regions. Specifically, we design two stages, Query-Guided Visual Grounding and Multi-Level Modal Interaction, to align fine-grained type-region information and simulate text-image/inner-text interactions respectively. For the former, we train a visual grounding model via transfer learning to extract region candidates that can be further integrated into the second stage to enhance token representations. For the latter, we design text-image and inner-text interaction modules along with three sub-tasks for MRC-MNER. To verify the effectiveness of our model, we conduct extensive experiments on two public MNER datasets, Twitter2015 and Twitter2017. Experimental results show that MRC-MNER outperforms the current state-of-the-art models on Twitter2017, and yields competitive results on Twitter2015.
Meihuizi Jia, Lei Shen 0001, Jinhui Pang, Lejian Liao, Yang Song 0008, Meng Chen 0006, Xiaodong He 0001
ACM Multimedia8
2022 Don't Take It Literally: An Edit-Invariant Sequence Loss for Text Generation
abstract
Guangyi Liu, Zichao Yang, Tianhua Tao, Xiaodan Liang, Junwei Bao, Zhen Li, Xiaodong He, Shuguang Cui, Zhiting Hu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Guangyi Liu 0005, Tianhua Tao, Xiaodan Liang, Junwei Bao 0001, Zhen Li 0026, Xiaodong He 0001, Shuguang Cui, Zhiting Hu
NAACL-HLT7
2022 LUNA: Learning Slot-Turn Alignment for Dialogue State Tracking
abstract
Yifan Wang, Jing Zhao, Junwei Bao, Chaoqun Duan, Youzheng Wu, Xiaodong He. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yifan Wang 0016, Junwei Bao 0001, Chaoqun Duan, Youzheng Wu, Xiaodong He 0001
NAACL-HLT6
2022 Label Anchored Contrastive Learning for Language Understanding
abstract
Contrastive learning (CL) has achieved astonishing progress in computer vision, speech, and natural language processing fields recently with self-supervised learning.However, CL approach to the supervised setting is not fully explored, especially for the natural language understanding classification task.Intuitively, the class label itself has the intrinsic ability to perform hard positive/negative mining, which is crucial for CL.Motivated by this, we propose a novel label anchored contrastive learning approach (denoted as LaCon) for language understanding.Specifically, three contrastive objectives are devised, including a multi-head instance-centered contrastive loss (ICL), a label-centered contrastive loss (LCL), and a label embedding regularizer (LER).Our approach does not require any specialized network architecture or any extra data augmentation, thus it can be easily plugged into existing powerful pre-trained language models.Compared to the state-of-the-art baselines, LaCon obtains up to 4.1% improvement on the popular datasets of GLUE and CLUE benchmarks.Besides, LaCon also demonstrates significant advantages under the few-shot and data imbalance settings, which obtains up to 9.4% improvement on the FewGLUE and FewCLUE benchmarking tasks.
Zhenyu Zhang 0029, Meng Chen 0006, Xiaodong He 0001
NAACL-HLT4
2022 OPERA: Operation-Pivoted Discrete Reasoning over Text
abstract
Yongwei Zhou, Junwei Bao, Chaoqun Duan, Haipeng Sun, Jiahui Liang, Yifan Wang, Jing Zhao, Youzheng Wu, Xiaodong He, Tiejun Zhao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yongwei Zhou, Junwei Bao 0001, Chaoqun Duan, Haipeng Sun, Jiahui Liang, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao
NAACL-HLT9
2022 Overview of the NLPCC 2022 Shared Task on Multimodal Product Summarization
Haoran Li 0001, Peng Yuan 0002, Haoning Zhang, Weikang Li, Song Xu 0002, Youzheng Wu, Xiaodong He 0001
NLPCC (2)7
2022 MFDG: A Multi-Factor Dialogue Graph Model for Dialogue Intent Classification
Jinhui Pang, Huinan Xu, Shuangyong Song, Xiaodong He 0001
ECML/PKDD (2)5
2022 DialCSP: A Two-Stage Attention-Based Model for Customer Satisfaction Prediction in E-commerce Customer Service
Zhenhe Wu, Liangqing Wu, Shuangyong Song, Jiahao Ji, Zhoujun Li 0001, Xiaodong He 0001
ECML/PKDD (3)7
2021 Incremental Learning for End-to-End Automatic Speech Recognition
abstract
In this paper, we propose an incremental learning method for end-to-end Automatic Speech Recognition (ASR) which enables an ASR system to perform well on new tasks while maintaining the performance on its originally learned ones. To mitigate catastrophic forgetting during incremental learning, we design a novel explainability-based knowledge distillation for ASR models, which is combined with a response-based knowledge distillation to maintain the original model's predictions and the “reason” for the predictions. Our method works without access to the training data of original tasks, which addresses the cases where the previous data is no longer available or joint training is costly. Results on a multi-stage sequential training task show that our method outperforms existing ones in mitigating forgetting. Furthermore, in two practical scenarios, compared to the target-reference joint training method, the performance drop of our method is 0.02% Character Error Rate (CER), which is 97% smaller than the drops of the baseline methods.
Libo Zi, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
ASRU6
2021 Learn to Copy from the Copying History: Correlational Copy Network for Abstractive Summarization
abstract
The copying mechanism has had considerable success in abstractive summarization, facilitating models to directly copy words from the input text to the output summary.Existing works mostly employ encoder-decoder attention, which applies copying at each time step independently of the former ones.However, this may sometimes lead to incomplete copying.In this paper, we propose a novel copying scheme named Correlational Copying Network (CoCoNet) that enhances the standard copying mechanism by keeping track of the copying history.It thereby takes advantage of prior copying distributions and, at each time step, explicitly encourages the model to copy the input word that is relevant to the previously copied one.In addition, we strengthen CoCoNet through pretraining with suitable corpora that simulate the copying behaviors.Experimental results show that CoCoNet can copy more accurately and achieves new state-of-the-art performances on summarization benchmarks, including CNN/DailyMail for news summarization and SAMSum for dialogue summarization.Our code is available at https:// github.com/hrlinlp/coconet.
Haoran Li 0001, Song Xu 0002, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
EMNLP (1)6
2021 Conversational Query Rewriting with Self-Supervised Learning
abstract
Context modeling plays a critical role in building multi-turn dialogue systems. Conversational Query Rewriting (CQR) aims to simplify the multi-turn dialogue modeling into a single-turn problem by explicitly rewriting the conversational query into a self-contained utterance. However, existing approaches rely on massive supervised training data, which is labor-intensive to annotate. And the detection of the omitted important information from context can be further improved. Besides, intent consistency constraint between contextual query and rewritten query is also ignored. To tackle these issues, we first propose to construct a large-scale CQR dataset automatically via self-supervised learning, which does not need human annotation. Then we introduce a novel CQR model Teresa based on Transformer, which is enhanced by self-attentive keywords detection and intent consistency constraint. Finally, we conduct extensive experiments on two public datasets. Experimental results demonstrate that our proposed model outperforms existing CQR baselines significantly, and also prove the effectiveness of self-supervised learning on improving the CQR performance.
Hang Liu 0005, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
ICASSP4
2021 Dian: Duration Informed Auto-Regressive Network for Voice Cloning
abstract
In this paper, we propose a novel end-to-end speech synthesis approach, Duration Informed Auto-regressive Network (DIAN), which consists of an acoustic model and a separate duration model. Un-like other auto-regressive TTS methods, the duration information of phonemes is provided as part of the input to the acoustic model, which enables the removal of the attention mechanism between its encoder and decoder parts. This eliminates the common seen skipping and repeating issues and improves speech intelligibility while ensuring high speech quality. A Transformer-based duration model is used to predict the duration of each phoneme for the attention-free acoustic model. We developed our TTS systems for the multi-speaker multi-style voice cloning challenge (M2VoC) using the proposed DIAN approach. In our procedure, a multi-speaker attention-free acoustic model and its Transformer-based duration model are first separately trained based on the training data released by M2VoC. Next, the multi-speaker models are adapted to form the speaker-specific models with the speaker-dependent data and transfer learning. At last, a speaker-specific LPCNet is estimated and used to synthesize the speech of the corresponding speaker. The M2VoC results showed that our proposed approach achieved the 3rd-place in the speech quality ranking and the 4th-place in the speaker similarity and style similarity ranking in the Track1-a task.
Zhengchen Zhang, Chao Zhang 0031, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
ICASSP6
2021 Improving Prosody Modelling with Cross-Utterance Bert Embeddings for End-to-End Speech Synthesis
abstract
Although speech prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account the information within each sentence. This makes it challenging when converting a paragraph of text into natural and expressive speech. In this paper, we propose to use the text embeddings of the neighboring sentences to improve the prosody generation for each utterance of a paragraph in an end-to-end fashion without using any explicit prosody features. More specifically, cross-utterance (CU) context vectors, which are produced by an additional CU encoder based on the sentence embeddings extracted by a pretrained BERT model, are used to augment the input of the Tacotron2 decoder. Two types of BERT embeddings are investigated, which leads to the use of different CU encoder structures. Experimental results on a Mandarin audiobook dataset and the LJ-Speech English audiobook dataset demonstrate the use of CU information can improve the naturalness and expressiveness of the synthesized speech. Subjective listening testing shows most of the participants prefer the voice generated using the CU encoder over that generated using standard Tacotron2. It is also found that the prosody can be controlled indirectly by changing the neighbouring sentences.
Zhengchen Zhang, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001
ICASSP5
2021 Neural Kalman Filtering for Speech Enhancement
abstract
Conventional learning-based speech enhancement methods usually utilize existing building blocks to design the deep neural networks (DNNs), while how to effectively integrate the statistical signal processing based schemes, which are expert-knowledge driven and could ameliorate the over-fitting problem, into the network design remains an open issue. In this paper, we extend the conventional Kalman filtering (KF) and propose a supervised-learning based neural Kalman filter (NKF) for speech enhancement. Similar to KF, the proposed method first obtains a prediction from the speech evolution model and then integrates the short-term instantaneous observation by linear weighting, and the weights are calculated by comparing between the speech prediction residual error and the environmental noise level. An end-to-end network is designed to convert the speech linear prediction model in KF to non-linear, and to compact all other conventional linear filtering operations. Different with other DNN based methods, the proposed method provides a specialized network design inspired from the conventional signal processing, the backpropagation can be directly applied on the linear filtering operations integrated from KF. We conduct experiments in different noisy conditions, and the results demonstrate that the proposed method outperforms the baseline methods which are based on either signal processing or DNNs.
Wei Xue 0002, Gang Quan, Chao Zhang 0031, Guo-Hong Ding, Xiaodong He 0001, Bowen Zhou 0001
ICASSP5
2021 ViDA-MAN: Visual Dialog with Digital Humans
abstract
We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text or voice-based system, ViDA-MAN offers human-like interactions (e.g, vivid voice, natural facial expression and body gestures). Given a speech request, the demonstration is able to response with high quality videos in sub-second latency. To deliver immersive user experience, ViDA-MAN seamlessly integrates multi-modal techniques including Acoustic Speech Recognition (ASR), multi-turn dialog, Text To Speech (TTS), talking heads video generation. Backed with large knowledge base, ViDA-MAN is able to chat with users on a number of topics including chit-chat, weather, device control, News recommendations, booking hotels, as well as answering questions via structured knowledge.
Jiawei Zuo, Liqin Jiang, Meng Chen 0006, Zhengchen Zhang, Wei Zhang 0031, Xiaodong He 0001, Tao Mei 0001
ACM Multimedia9
2021 Learning to Compose Stylistic Calligraphy Artwork with Emotions
abstract
Emotion plays a critical role in calligraphy composition, which makes the calligraphy artwork impressive and have a soul. However, previous research on calligraphy generation all neglected the emotion as a major contributor to the artistry of calligraphy. Such defects prevent them from generating aesthetic, stylistic, and diverse calligraphy artworks, but only static handwriting font library instead. To address this problem, we propose a novel cross-modal approach to generate stylistic and diverse Chinese calligraphy artwork driven by different emotions automatically. We firstly detect the emotions in the text by a classifier, then generate the emotional Chinese character images via a novel modified Generative Adversarial Network (GAN) structure, finally we predict the layout for all character images with a recurrent neural network. We also collect a large-scale stylistic Chinese calligraphy image dataset with rich emotions. Experimental results demonstrate that our model outperforms all baseline image translation models significantly for different emotional styles in terms of content accuracy and style discrepancy. Besides, our layout algorithm can also learn the patterns and habits of calligrapher, and makes the generated calligraphy more artistic. To the best of our knowledge, we are the first to work on emotion-driven discourse-level Chinese calligraphy artwork composition.
Shaozu Yuan, Ruixue Liu, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001
ACM Multimedia6
2021 Graph Ensemble Learning over Multiple Dependency Trees for Aspect-level Sentiment Classification
abstract
Xiaochen Hou, Peng Qi, Guangtao Wang, Rex Ying, Jing Huang, Xiaodong He, Bowen Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xiaochen Hou, Peng Qi 0003, Guangtao Wang, Rex Ying, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001
NAACL-HLT6
2021 SGG: Learning to Select, Guide, and Generate for Keyphrase Generation
abstract
Jing Zhao, Junwei Bao, Yifan Wang, Youzheng Wu, Xiaodong He, Bowen Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Junwei Bao 0001, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
NAACL-HLT5
2021 CUSTOM: Aspect-Oriented Product Summarization for E-Commerce
Jiahui Liang, Junwei Bao 0001, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
NLPCC (2)5
2021 EviDR: Evidence-Emphasized Discrete Reasoning for Reasoning Machine Reading Comprehension
Yongwei Zhou, Junwei Bao 0001, Haipeng Sun, Jiahui Liang, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001, Tiejun Zhao
NLPCC (1)6
2020 Zero-Shot Text-to-SQL Learning with Auxiliary Task
abstract
Recent years have seen great success in the use of neural seq2seq models on the text-to-SQL task. However, little work has paid attention to how these models generalize to realistic unseen data, which naturally raises a question: does this impressive performance signify a perfect generalization model, or are there still some limitations?In this paper, we first diagnose the bottleneck of the text-to-SQL task by providing a new testbed, in which we observe that existing models present poor generalization ability on rarely-seen data. The above analysis encourages us to design a simple but effective auxiliary task, which serves as a supportive model as well as a regularization term to the generation task to increase the models' generalization. Experimentally, We evaluate our models on a large text-to-SQL dataset WikiSQL. Compared to a strong baseline coarse-to-fine model, our models improve over the baseline by more than 3% absolute in accuracy on the whole dataset. More interestingly, on a zero-shot subset test of WikiSQL, our models achieve 5% absolute accuracy gain over the baseline, clearly demonstrating its superior generalizability.
Shuaichen Chang, Yun Tang 0002, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001
AAAI5
2020 Aspect-Aware Multimodal Summarization for Chinese E-Commerce Products
abstract
We present an abstractive summarization system that produces summary for Chinese e-commerce products. This task is more challenging than general text summarization. First, the appearance of a product typically plays a significant role in customers' decisions to buy the product or not, which requires that the summarization model effectively use the visual information of the product. Furthermore, different products have remarkable features in various aspects, such as “energy efficiency” and “large capacity” for refrigerators. Meanwhile, different customers may care about different aspects. Thus, the summarizer needs to capture the most attractive aspects of a product that resonate with potential purchasers. We propose an aspect-aware multimodal summarization model that can effectively incorporate the visual information and also determine the most salient aspects of a product. We construct a large-scale Chinese e-commerce product summarization dataset that contains approximately 1.4 million manually created product summaries that are paired with detailed product information, including an image, a title, and other textual descriptions for each product. The experimental results on this dataset demonstrate that our models significantly outperform the comparative methods in terms of both the ROUGE score and manual evaluations.
Haoran Li 0001, Peng Yuan 0002, Song Xu 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
AAAI5
2020 Keywords-Guided Abstractive Sentence Summarization
abstract
We study the problem of generating a summary for a given sentence. Existing researches on abstractive sentence summarization ignore that keywords in the input sentence provide significant clues for valuable content, and humans tend to write summaries covering these keywords. In this paper, we propose an abstractive sentence summarization method by applying guidance signals of keywords to both the encoder and the decoder in the sequence-to-sequence model. A multi-task learning framework is adopted to jointly learn to extract keywords and generate a summary for the input sentence. We apply keywords-guided selective encoding strategies to filter source information by investigating the interactions between the input sentence and the keywords. We extend pointer-generator network by a dual-attention and a dual-copy mechanism, which can integrate the semantics of the input sentence and the keywords, and copy words from both the input sentence and the keywords. We demonstrate that multi-task learning and keywords-oriented guidance facilitate sentence summarization task, achieving better performance than the competitive models on the English Gigaword sentence summarization dataset.
Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Chengqing Zong, Xiaodong He 0001
AAAI5
2020 Select, Answer and Explain: Interpretable Multi-Hop Reading Comprehension over Multiple Documents
abstract
Interpretable multi-hop reading comprehension (RC) over multiple documents is a challenging problem because it demands reasoning over multiple information sources and explaining the answer prediction by providing supporting evidences. In this paper, we propose an effective and interpretable Select, Answer and Explain (SAE) system to solve the multi-document RC problem. Our system first filters out answer-unrelated documents and thus reduce the amount of distraction information. This is achieved by a document classifier trained with a novel pairwise learning-to-rank loss. The selected answer-related documents are then input to a model to jointly predict the answer and supporting sentences. The model is optimized with a multi-task learning objective on both token level for answer prediction and sentence level for supporting sentences prediction, together with an attention-based interaction between these two tasks. Evaluated on HotpotQA, a challenging multi-hop RC data set, the proposed SAE system achieves top competitive performance in distractor setting compared to other existing systems on the leaderboard.
Kevin Huang 0002, Guangtao Wang, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001
AAAI5
2020 Orthogonal Relation Transforms with Graph Context Modeling for Knowledge Graph Embedding
abstract
Distance-based knowledge graph embeddings have shown substantial improvement on the knowledge graph link prediction task, from TransE to the latest state-of-the-art RotatE.However, complex relations such as N-to-1, 1-to-N and N-to-N still remain challenging to predict.In this work, we propose a novel distance-based approach for knowledge graph link prediction.First we extend the RotatE from 2D complex domain to high dimensional space with orthogonal transforms to model relations.The orthogonal transform embedding for relations keeps the capability for modeling symmetric/anti-symmetric, inverse and compositional relations while achieves better modeling capacity.Second, the graph context is integrated into distance scoring functions directly.Specifically, graph context is explicitly modeled via two directed context representations.Each node embedding in knowledge graph is augmented with two context representations, which are computed from the neighboring outgoing and incoming nodes/edges respectively.The proposed approach improves prediction accuracy on the difficult N-to-1, 1-to-N and N-to-N cases.Our experimental results show that it achieves state-of-the-art results on two common benchmarks FB15k-237 and WNRR-18, especially on FB15k-237 which has many high in-degree nodes.Code available at https://github. com/JD-AI-Research-Silicon-Valley/ KGEmbedding-OTE.
Yun Tang 0002, Jing Huang 0019, Guangtao Wang, Xiaodong He 0001, Bowen Zhou 0001
ACL4
2020 Self-Attention Guided Copy Mechanism for Abstractive Summarization
abstract
Copy module has been widely equipped in the recent abstractive summarization models, which facilitates the decoder to extract words from the source into the summary.Generally, the encoder-decoder attention is served as the copy distribution, while how to guarantee that important words in the source are copied remains a challenge.In this work, we propose a Transformer-based model to enhance the copy mechanism.Specifically, we identify the importance of each source word based on the degree centrality with a directed graph built by the self-attention layer in the Transformer.We use the centrality of each source word to guide the copy process explicitly.Experimental results show that the self-attention graph provides useful guidance for the copy distribution.Our proposed models significantly outperform the baseline methods on the CNN/Daily Mail dataset and the Gigaword dataset.
Song Xu 0002, Haoran Li 0001, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
ACL5
2020 Multimodal Sentence Summarization via Multimodal Selective Encoding
abstract
This paper studies the problem of generating a summary for a given sentence-image pair.Existing multimodal sequence-to-sequence approaches mainly focus on enhancing the decoder by visual signals, while ignoring that the image can improve the ability of the encoder to identify highlights of a news event or a document.Thus, we propose a multimodal selective gate network that considers reciprocal relationships between textual and multi-level visual features, including global image descriptor, activation grids, and object proposals, to select highlights of the event when encoding the source sentence.In addition, we introduce a modality regularization to encourage the summary to capture the highlights embedded in the image more accurately.To verify the generalization of our model, we adopt the multimodal selective gate to the text-based decoder and multimodal-based decoder.Experimental results on a public multimodal sentence summarization dataset demonstrate the advantage of our models over baselines.Further analysis suggests that our proposed multimodal selective gate network can effectively select important information in the input sentence.
Haoran Li 0001, Junnan Zhu, Jiajun Zhang 0001, Xiaodong He 0001, Chengqing Zong
COLING4
2020 Learning to Decouple Relations: Few-Shot Relation Classification with Entity-Guided Attention and Confusion-Aware Training
abstract
This paper aims to enhance the few-shot relation classification especially for sentences that jointly describe multiple relations.Due to the fact that some relations usually keep high cooccurrence in the same context, previous few-shot relation classifiers struggle to distinguish them with few annotated instances.To alleviate the above relation confusion problem, we propose CTEG, a model equipped with two mechanisms to learn to decouple these easily-confused relations.On the one hand, an Entity-Guided Attention (EGA) mechanism, which leverages the syntactic relations and relative positions between each word and the specified entity pair, is introduced to guide the attention to filter out information causing confusion.On the other hand, a Confusion-Aware Training (CAT) method is proposed to explicitly learn to distinguish relations by playing a pushing-away game between classifying a sentence into a true relation and its confusing relation.Extensive experiments are conducted on the FewRel dataset, and the results show that our proposed model achieves comparable and even much better results to strong baselines in terms of accuracy.Furthermore, the ablation test and case study verify the effectiveness of our proposed EGA and CAT, especially in addressing the relation confusion problem.
Yingyao Wang, Junwei Bao 0001, Guangyi Liu 0005, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001, Tiejun Zhao
COLING5
2020 On the Faithfulness for E-commerce Product Summarization
abstract
In this work, we present a model to generate e-commerce product summaries.The consistency between the generated summary and the product attributes is an essential criterion for the ecommerce product summarization task.To enhance the consistency, first, we encode the product attribute table to guide the process of summary generation.Second, we identify the attribute words from the vocabulary, and we constrain these attribute words can be presented in the summaries only through copying from the source, i.e., the attribute words not in the source cannot be generated.We construct a Chinese e-commerce product summarization dataset, and the experimental results on this dataset demonstrate that our models significantly improve the faithfulness.
Peng Yuan 0002, Haoran Li 0001, Song Xu 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
COLING5
2020 Learning to Predict Charges for Legal Judgment via Self-Attentive Capsule Network
abstract
With the rapid development of deep learning technology, more and more traditional industries are changed by Artificial Intelligence. The legal industry is such a popular scenario which attracts lots of researchers' interests. In this work, we focus on automatic charge prediction, which predicts the final charges according to the given fact descriptions in criminal cases. It is crucial for legal assistant systems and can help the judges improve work efficiency greatly. However, extremely imbalanced data distribution and lengthy fact descriptions make this task especially challenging. To tackle these two issues, we propose a novel model, namely Self-Attentive Capsule Network (dubbed as SAttCaps). In particular, we devise a self-attentive dynamic routing, which can not only capture long-range dependency more directly than vanilla dynamic routing, but also learn the high-level generalized features better. The experimental results on three real-world datasets demonstrate that our model significantly outperforms the baselines and creates new state-of-the-art performance. Moreover, our model performs much better than the baselines especially in the low-frequency charges and can bring 5.7% absolute improvement under F1 score.
Yuquan Le, Congqing He, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
ECAI5
2020 Multimodal Joint Attribute Prediction and Value Extraction for E-commerce Product
abstract
Product attribute values are essential in many e-commerce scenarios, such as customer service robots, product recommendations, and product retrieval.While in the real world, the attribute values of a product are usually incomplete and vary over time, which greatly hinders the practical applications.In this paper, we propose a multimodal method to jointly predict product attributes and extract values from textual product descriptions with the help of the product images.We argue that product attributes and values are highly correlated, e.g., it will be easier to extract the values on condition that the product attributes are given.Thus, we jointly model the attribute prediction and value extraction tasks from multiple aspects towards the interactions between attributes and values.Moreover, product images have distinct effects on our tasks for different product attributes and values.Thus, we selectively draw useful visual information from product images to enhance our model.We annotate a multimodal product attribute value dataset that contains 87,194 instances, and the experimental results on this dataset demonstrate that explicitly modeling the relationship between attributes and values facilitates our method to establish the correspondence between them, and selectively utilizing visual product information is necessary for the task.Our code and dataset are available 1 .
Tiangang Zhu, Haoran Li 0001, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
EMNLP (1)5
2020 Efficient WaveGlow: An Improved WaveGlow Vocoder with Enhanced Speed
Zhengchen Zhang, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001
INTERSPEECH5
2020 The JD AI Speaker Verification System for the FFSVC 2020 Challenge
abstract
This paper presents the development of our systems for the Interspeech 2020 Far-Field Speaker Verification Challenge (FFSVC). Our focus is the task 2 of the challenge, which is to perform far-field text-independent speaker verification using a single microphone array. The FFSVC training set provided by the challenge is augmented by pre-processing the far-field data with both beamforming, voice channel switching, and a combination of weighted prediction error (WPE) and beamforming. Two open-access corpora, CHData in Mandarin and VoxCeleb2 in English, are augmented using multiple methods and mixed with the augmented FFSVC data to form the final training data. Four different model structures are used to model speaker characteristics: ResNet, extended time-delay neural network (ETDNN), Transformer, and factorized TDNN (FTDNN), whose output values are pooled across time using the self-attentive structure, the statistic pooling structure, and the GVLAD structure. The final results are derived by fusing the adaptively normalized scores of the four systems with a two-stage fusion method, which achieves a minimum of the detection cost function (minDCF) of 0.3407 and an equal error rate (EER) of 2.67% on the development set of the challenge.
Ying Tong, Wei Xue 0002, Shanluo Huang, Fan Lu 0003, Chao Zhang 0031, Guo-Hong Ding, Xiaodong He 0001
INTERSPEECH7
2020 Sound Event Localization and Detection Based on Multiple DOA Beamforming and Multi-Task Learning
abstract
The performance of sound event localization and detection (SELD) degrades in source-overlapping cases since features of different sources collapse with each other, and the network tends to fail to learn to separate these features effectively. In this paper, by leveraging the conventional microphone array signal processing to generate comprehensive representations for SELD, we propose a new SELD method based on multiple direction of arrival (DOA) beamforming and multi-task learning. By using multiple beamformers to extract the signals from different DOAs, the sound field is more diversely described, and specialised representations of target source and noises can be obtained. With labelled training data, the steering vector is estimated based on the cross-power spectra (CPS) and the signal presence probability (SPP), which eliminates the need of knowing the array geometry. We design two networks for sound event localization (SED) and sound source localization (SSL) and use a multi-task learning scheme for SED, in which the SSL-related task act as a regularization. Experimental results using the database of DCASE2019 SELD task show that the proposed method achieves the state-of-art performance.
Wei Xue 0002, Ying Tong, Chao Zhang 0031, Guo-Hong Ding, Xiaodong He 0001, Bowen Zhou 0001
INTERSPEECH5
2020 The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer Service
abstract
Human conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge amount of real conversation data. In this paper, we construct a large-scale real scenario Chinese E-commerce conversation corpus, JDDC, with more than 1 million multi-turn dialogues, 20 million utterances, and 150 million words. The dataset reflects several characteristics of human-human conversations, e.g., goal-driven, and long-term dependency among the context. It also covers various dialogue types including task-oriented, chitchat and question-answering. Extra intent information and three well-annotated challenge sets are also provided. Then, we evaluate several retrieval-based and generative models to provide basic benchmark performance on the JDDC corpus. And we hope JDDC can serve as an effective testbed and benefit the development of fundamental research in dialogue task.
Meng Chen 0006, Ruixue Liu, Lei Shen 0001, Shaozu Yuan, Jingyan Zhou, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
LREC7
2020 MaLiang: An Emotion-driven Chinese Calligraphy Artwork Composition System
abstract
We present a novel Chinese calligraphy artwork composition system (MaLiang) which can generate aesthetic, stylistic and diverse calligraphy images based on the emotion status from the input text. Different from previous research, it's the first work to endow the calligraphy synthesis with the ability to express fickle emotions and composite a whole piece of discourse-level calligraphy artwork instead of single character images. The system consists of three modules: emotion detection, character image generation, and layout prediction. As a creative form of interactive art, MaLiang has been exhibited in several famous international art festivals.
Ruixue Liu, Shaozu Yuan, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001
ACM Multimedia6
2020 Group Contextual Encoding for 3D Point Clouds
abstract
Global context is crucial for 3D point cloud scene understanding tasks. In this work, we extended the contextual encoding layer that was originally designed for 2D tasks to 3D Point Cloud scenarios. The encoding layer learns a set of code words in the feature space of the 3D point cloud to characterize the global semantic context, and then based on these code words, the method learns a global contextual descriptor to reweight the featuremaps accordingly. Moreover, compared to 2D scenarios, data sparsity becomes a major issue in 3D point cloud scenarios, and the performance of contextual encoding quickly saturates when the number of code words increases. To mitigate this problem, we further proposed a group contextual encoding method, which divides the channel into groups and then performs encoding on group-divided feature vectors. This method facilitates learning of global context in grouped subspace for 3D point clouds. We evaluate the effectiveness and generalizability of our method on three widely-studied 3D point cloud tasks. Experimental results have shown that the proposed method outperformed the VoteNet remarkably with 3 mAP on the benchmark of SUN-RGBD, with the metrics of mAP@ 0.25, and a much greater margin of 6.57 mAP on ScanNet with the metrics of mAP@ 0.5. Compared to the baseline of PointNet++, the proposed method leads to an accuracy of 86 %, outperforming the baseline by 1.5 %. Our proposed method have outperformed the non-grouping baseline methods across the board and establishes new state-of-the-art on these benchmarks.
Xu Liu 0017, Chengtao Li, Jian Wang 0100, Boxin Shi, Xiaodong He 0001
NeurIPS6
2020 Enhancing Multi-turn Dialogue Modeling with Intent Information for E-Commerce Customer Service
Ruixue Liu, Meng Chen 0006, Hang Liu 0005, Lei Shen 0001, Yang Song 0008, Xiaodong He 0001
NLPCC (1)6
2020 AIIS: The SIGIR 2020 Workshop on Applied Interactive Information Systems
abstract
Nowadays, intelligent information systems, especially the interactive information systems (e.g., conversational interaction systems like Siri, and Cortana; news feed recommender systems, and interactive search engines, etc.), are ubiquitous in real-world applications. These systems either converse with users explicitly through natural languages, or mine users interests and respond to users requests implicitly. Interactivity has become a crucial element towards intelligent information systems. Despite the fact that interactive information systems have gained significant progress, there are still many challenges to be addressed when applying these models to real-world scenarios. This half day workshop explores challenges and potential research, development, and application directions in applied interactive information systems. We aim to discuss the issues of applying interactive information models to production systems, as well as to shed some light on the fundamental characteristics, i.e., interactivity and applicability, of different interactive tasks. We welcome practical, theoretical, experimental, and methodological studies that advances the interactivity towards intelligent information systems. The workshop aims to bring together a diverse set of practitioners and researchers interested in investigating the interaction between human and information systems to develop more intelligent information systems.
Hongshen Chen, Zhaochun Ren, Pengjie Ren, Dawei Yin 0001, Xiaodong He 0001
SIGIR5
2019 Attentive Tensor Product Learning
abstract
This paper proposes a novel neural architecture — Attentive Tensor Product Learning (ATPL) — to represent grammatical structures of natural language in deep learning models. ATPL exploits Tensor Product Representations (TPR), a structured neural-symbolic model developed in cognitive science, to integrate deep learning with explicit natural language structures and rules. The key ideas of ATPL are: 1) unsupervised learning of role-unbinding vectors of words via the TPR-based deep neural network; 2) the use of attention modules to compute TPR; and 3) the integration of TPR with typical deep learning architectures including long short-term memory and feedforward neural networks. The novelty of our approach lies in its ability to extract the grammatical structure of a sentence by using role-unbinding vectors, which are obtained in an unsupervised manner. Our ATPL approach is applied to 1) image captioning, 2) part of speech (POS) tagging, and 3) constituency parsing of a natural language sentence. The experimental results demonstrate the effectiveness of the proposed approach in all these three natural language processing tasks.
Qiuyuan Huang, Li Deng 0001, Dapeng Oliver Wu, Chang Liu 0021, Xiaodong He 0001
AAAI5
2019 Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
abstract
We propose a hierarchically structured reinforcement learning approach to address the challenges of planning for generating coherent multi-sentence stories for the visual storytelling task. Within our framework, the task of generating a story given a sequence of images is divided across a two-level hierarchical decoder. The high-level decoder constructs a plan by generating a semantic concept (i.e., topic) for each image in sequence. The low-level decoder generates a sentence for each image using a semantic compositional network, which effectively grounds the sentence generation conditioned on the topic. The two decoders are jointly trained end-to-end using reinforcement learning. We evaluate our model on the visual storytelling (VIST) dataset. Empirical results from both automatic and human evaluations demonstrate that the proposed hierarchically structured reinforced training achieves significantly better performance compared to a strong flat deep reinforcement learning baseline.
Qiuyuan Huang, Zhe Gan, Asli Celikyilmaz, Dapeng Oliver Wu, Xiaodong He 0001
AAAI6
2019 End-to-End Structure-Aware Convolutional Networks for Knowledge Base Completion
abstract
Knowledge graph embedding has been an active research topic for knowledge base completion, with progressive improvement from the initial TransE, TransH, DistMult et al to the current state-of-the-art ConvE. ConvE uses 2D convolution over embeddings and multiple layers of nonlinear features to model knowledge graphs. The model can be efficiently trained and scalable to large knowledge graphs. However, there is no structure enforcement in the embedding space of ConvE. The recent graph convolutional network (GCN) provides another way of learning graph node embedding by successfully utilizing graph connectivity structure. In this work, we propose a novel end-to-end StructureAware Convolutional Network (SACN) that takes the benefit of GCN and ConvE together. SACN consists of an encoder of a weighted graph convolutional network (WGCN), and a decoder of a convolutional network called Conv-TransE. WGCN utilizes knowledge graph node structure, node attributes and edge relation types. It has learnable weights that adapt the amount of information from neighbors used in local aggregation, leading to more accurate embeddings of graph nodes. Node attributes in the graph are represented as additional nodes in the WGCN. The decoder Conv-TransE enables the state-of-the-art ConvE to be translational between entities and relations while keeps the same link prediction performance as ConvE. We demonstrate the effectiveness of the proposed SACN on standard FB15k-237 and WN18RR datasets, and it gives about 10% relative improvement over the state-of-theart ConvE in terms of HITS@1, HITS@3 and HITS@10.
Yun Tang 0002, Jing Huang 0019, Jinbo Bi, Xiaodong He 0001, Bowen Zhou 0001
AAAI5
2019 Multi-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous Graphs
abstract
Multi-hop reading comprehension (RC) across documents poses new challenge over single-document RC because it requires reasoning over multiple documents to reach the final answer. In this paper, we propose a new model to tackle the multi-hop RC problem. We introduce a heterogeneous graph with different types of nodes and edges, which is named as Heterogeneous Document-Entity (HDE) graph. The advantage of HDE graph is that it contains different granularity levels of information including candidates, documents and entities in specific document contexts. Our proposed model can do reasoning over the HDE graph with nodes representation initialized with co-attention and self-attention based context encoders. We employ Graph Neural Networks (GNN) based message passing algorithms to accumulate evidences on the proposed HDE graph. Evaluated on the blind test set of the Qangaroo WikiHop data set, our HDE graph based single model delivers competitive result, and the ensemble model achieves the state-of-the-art performance.
Guangtao Wang, Jing Huang 0019, Yun Tang 0002, Xiaodong He 0001, Bowen Zhou 0001
ACL (1)5
2019 Relation Module for Non-Answerable Predictions on Reading Comprehension
abstract
Machine reading comprehension (MRC) has attracted significant amounts of research attention recently, due to an increase of challenging reading comprehension datasets.In this paper, we aim to improve a MRC model's ability to determine whether a question has an answer in a given context (e.g. the recently proposed SQuAD 2.0 task).Our solution is a relation module that is adaptable to any MRC model.The relation module consists of both semantic extraction and relational information.We first extract high level semantics as objects from both question and context with multihead self-attentive pooling.These semantic objects are then passed to a relation network, which generates relationship scores for each object pair in a sentence.These scores are used to determine whether a question is nonanswerable.We test the relation module on the SQuAD 2.0 dataset using both the BiDAF and BERT models as baseline readers.We obtain 1.8% gain of F1 accuracy on top of the BiDAF reader, and 1.0% on top of the BERT base model.These results show the effectiveness of our relation module on MRC.
Kevin Huang 0002, Yun Tang 0002, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001
CoNLL4
2019 Object-Driven Text-To-Image Synthesis via Adversarial Training
abstract
In this paper, we propose Object-driven Attentive Generative Adversarial Newtorks (Obj-GANs) that allow attention-driven, multi-stage refinement for synthesizing complex images from text descriptions. With a novel object-driven attentive generative network, the Obj-GAN can synthesize salient objects by paying attention to their most relevant words in the text descriptions and their pre-generated class label. In addition, a novel object-wise discriminator based on the Fast R-CNN model is proposed to provide rich object-wise discrimination signals on whether the synthesized object matches the text description and the pre-generated class label. The proposed Obj-GAN significantly outperforms the previous state of the art in various metrics on the large-scale MS-COCO benchmark, increasing the inception score by 27% and decreasing the FID score by 11%. A thorough comparison between the classic grid attention and the new object-driven attention is provided through analyzing their mechanisms and visualizing their attention layers, showing insights of how the proposed model generates complex scenes in high quality.
Wenbo Li 0001, Pengchuan Zhang, Lei Zhang 0001, Qiuyuan Huang, Xiaodong He 0001, Siwei Lyu, Jianfeng Gao 0001
CVPR5
2019 Deep Speaker Embedding Learning with Multi-level Pooling for Text-independent Speaker Verification
abstract
This paper aims to improve the widely used deep speaker embedding x-vector model. We propose the following improvements: (1) a hybrid neural network structure using both time delay neural network (TDNN) and long short-term memory neural networks (LSTM) to generate complementary speaker information at different levels; (2) a multi-level pooling strategy to collect speaker information from both TDNN and LSTM layers; (3) a regularization scheme on the speaker embedding extraction layer to make the extracted embeddings suitable for the following fusion step. The synergy of these improvements are shown on the NIST SRE 2016 eval test (with a 19% EER reduction) and SRE 2018 dev test (with a 9% EER reduction), as well as more than 10% DCF scores reduction on these two test sets over the x-vector baseline.
Yun Tang 0002, Guo-Hong Ding, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001
ICASSP4
2019 Dynamic Item Block and Prediction Enhancing Block for Sequential Recommendation
abstract
Sequential recommendation systems have become a research hotpot recently to suggest users with the next item of interest (to interact with). However, existing approaches suffer from two limitations: (1) The representation of an item is relatively static and fixed for all users. We argue that even a same item should be represented distinctively with respect to different users and time steps. (2) The generation of a prediction for a user over an item is computed in a single scale (e.g., by their inner product), ignoring the nature of multi-scale user preferences. To resolve these issues, in this paper we propose two enhancing building blocks for sequential recommendation. Specifically, we devise a Dynamic Item Block (DIB) to learn dynamic item representation by aggregating the embeddings of those who rated the same item before that time step. Then, we come up with a Prediction Enhancing Block (PEB) to project user representation into multiple scales, based on which many predictions can be made and attentively aggregated for enhanced learning. Each prediction is generated by a softmax over a sampled itemset rather than the whole item space for efficiency. We conduct a series of experiments on four real datasets, and show that even a basic model can be greatly enhanced with the involvement of DIB and PEB in terms of ranking accuracy. The code and datasets can be obtained from https://github.com/ouououououou/DIB-PEB-Sequential-RS
Guibing Guo, Shichang Ouyang, Xiaodong He 0001, Fajie Yuan
IJCAI3
2019 Discrete Trust-aware Matrix Factorization for Fast Recommendation
abstract
Trust-aware recommender systems have received much attention recently for their abilities to capture the influence among connected users. However, they suffer from the efficiency issue due to large amount of data and time-consuming real-valued operations. Although existing discrete collaborative filtering may alleviate this issue to some extent, it is unable to accommodate social influence. In this paper we propose a discrete trust-aware matrix factorization (DTMF) model to take dual advantages of both social relations and discrete technique for fast recommendation. Specifically, we map the latent representation of users and items into a joint hamming space by recovering the rating and trust interactions between users and items. We adopt a sophisticated discrete coordinate descent (DCD) approach to optimize our proposed model. In addition, experiments on two real-world datasets demonstrate the superiority of our approach against other state-of-the-art approaches in terms of ranking accuracy and efficiency.
Guibing Guo, Enneng Yang, Li Shen 0008, Xiaochun Yang 0001, Xiaodong He 0001
IJCAI5
2019 Mappa Mundi: An Interactive Artistic Mind Map Generator with Artificial Imagination
abstract
We present a novel real-time, collaborative, and interactive AI painting system, Mappa Mundi, for artistic Mind Map creation. The system consists of a voice-based input interface, an automatic topic expansion module, and an image projection module. The key innovation is to inject Artificial Imagination into painting creation by considering lexical and phonological similarities of language, learning and inheriting artist’s original painting style, and applying the principles of Dadaism and impossibility of improvisation. Our system indicates that AI and artist can collaborate seamlessly to create imaginative artistic painting and Mappa Mundi has been applied in art exhibition in UCCA, Beijing.
Ruixue Liu, Baoyang Chen, Meng Chen 0006, Youzheng Wu, Zhijie Qiu, Xiaodong He 0001
IJCAI6
2019 Knowledgeable Storyteller: A Commonsense-Driven Generative Model for Visual Storytelling
abstract
The visual storytelling (VST) task aims at generating a reasonable and coherent paragraph-level story with the image stream as input. Different from caption that is a direct and literal description of image content, the story in the VST task tends to contain plenty of imaginary concepts that do not appear in the image. This requires the AI agent to reason and associate with the imaginary concepts based on implicit commonsense knowledge to generate a reasonable story describing the image stream. Therefore, in this work, we present a commonsense-driven generative model, which aims to introduce crucial commonsense from the external knowledge base for visual storytelling. Our approach first extracts a set of candidate knowledge graphs from the knowledge base. Then, an elaborately designed vision-aware directional encoding schema is adopted to effectively integrate the most informative commonsense. Besides, we strive to maximize the semantic similarity within the output during decoding to enhance the coherence of the generated text. Results show that our approach can outperform the state-of-the-art systems by a large margin, which achieves a 29\% relative improvement of CIDEr score. With additional commonsense and semantic-relevance based objective, the generated stories are more diverse and coherent.
Fuli Luo, Lei Li 0039, Zhiyi Yin, Xiaodong He 0001, Xu Sun 0001
IJCAI6
2019 Multi-Stride Self-Attention for Speech Recognition
Kyu Jeong Han, Jing Huang 0019, Yun Tang 0002, Xiaodong He 0001, Bowen Zhou 0001
INTERSPEECH4
2019 Speaker Diarization with Lexical Information
abstract
This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn probabilities with speaker embeddings into a speaker clustering process to improve the overall diarization accuracy. To integrate lexical and acoustic information in a comprehensive way during clustering, we introduce an adjacency matrix integration for spectral clustering. Since words and word boundary information for word-level speaker turn probability estimation are provided by a speech recognition system, our proposed method works without any human intervention for manual transcriptions. We show that the proposed method improves diarization performance on various evaluation datasets compared to the baseline diarization system using acoustic information only in speaker embeddings.
Tae Jin Park, Kyu Jeong Han, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH4
2019 Direct-Path Signal Cross-Correlation Estimation for Sound Source Localization in Reverberation
abstract
Sound source localization (SSL) is challenging in presence of reverberation since the cross-correlation between the direct-path signals in different microphones, which indicates the spatial information of the sound source, is interfered by the reverberation signal components. A novel algorithm is proposed in this paper to estimate the cross-correlation of the direct-path speech signals, such that the robustness of SSL to reverberation can be improved. The proposed method follows a similar scheme to the multichannel linear prediction (MCLP), which is commonly used for speech dereverberation, while avoids the explicit estimation of the direct-path signal of each channel. This is achieved by revealing the relationship between the direct-path signal cross-correlation (DPCC) and the MCLP coefficient vector, and finally deriving the DPCC by using only the multichannel reverberant signals. It is also shown that the pre-whitening operation, which is widely used for SSL, can be inherently integrated into the estimated DPCC. An adaptive method is further derived to facilitate online frame-level SSL. The proposed method can be easily applied to conventional cross-correlation based SSL methods by using the DPCC rather than the full cross-correlation. Experiments conducted in various reverberant conditions demonstrate the effectiveness of the proposed method.
Wei Xue 0002, Ying Tong, Guo-Hong Ding, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001
INTERSPEECH6
2019 Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image Representations
abstract
In vision-and-language grounding problems, fine-grained representations of the image are considered to be of paramount importance. Most of the current systems incorporate visual features and textual concepts as a sketch of an image. However, plainly inferred representations are usually undesirable in that they are composed of separate components, the relations of which are elusive. In this work, we aim at representing an image with a set of integrated visual regions and corresponding textual concepts, reflecting certain semantics. To this end, we build the Mutual Iterative Attention (MIA) module, which integrates correlated visual features and textual concepts, respectively, by aligning the two modalities. We evaluate the proposed approach on two representative vision-and-language grounding tasks, i.e., image captioning and visual question answering. In both tasks, the semantic-grounded image representations consistently boost the performance of the baseline models under all metrics across the board. The results demonstrate that our approach is effective and generalizes well to a wide range of models for image-related applications. (The code is available at \url{https://github.com/fenglinliu98/MIA)
Yuanxin Liu, Xuancheng Ren, Xiaodong He 0001, Xu Sun 0001
NeurIPS4
2019 Automated Thematic and Emotional Modern Chinese Poetry Composition
Meng Chen 0006, Yang Song 0008, Xiaodong He 0001, Bowen Zhou 0001
NLPCC (1)4
2018 Question-Answering with Grammatically-Interpretable Representations
abstract
We introduce an architecture, the Tensor Product RecurrentNetwork (TPRN). In our application of TPRN, internal representations—learned by end-to-end optimization in a deep neural network performing a textual question-answering(QA) task—can be interpreted using basic concepts from linguistic theory. No performance penalty need be paid for this increased interpretability: the proposed model performs comparably to a state-of-the-art system on the SQuAD QA task.The internal representation which is interpreted is a Tensor Product Representation: for each input word, the model selects a symbol to encode the word, and a role in which to place the symbol, and binds the two together. The selection is via soft attention. The overall interpretation is built from interpretations of the symbols, as recruited by the trained model, and interpretations of the roles as used by the model. We find support for our initial hypothesis that symbols can be interpreted as lexical-semantic word meanings, while roles can be interpreted as approximations of grammatical roles (or categories)such as subject, wh-word, determiner, etc. Fine-grained analysis reveals specific correspondences between the learned roles and parts of speech as assigned by a standard tagger(Toutanova et al. 2003), and finds several discrepancies in the model’s favor. In this sense, the model learns significant aspectsof grammar, after having been exposed solely to linguistically unannotated text, questions, and answers: no prior linguistic knowledge is given to the model. What is given is the means to build representations using symbols and roles, with an inductive bias favoring use of these in an approximately discrete manner.
Hamid Palangi, Paul Smolensky, Xiaodong He 0001, Li Deng 0001
AAAI3
2018 Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
abstract
Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reasoning. In this work, we propose a combined bottom-up and top-down attention mechanism that enables attention to be calculated at the level of objects and other salient image regions. This is the natural basis for attention to be considered. Within our approach, the bottom-up mechanism (based on Faster R-CNN) proposes image regions, each with an associated feature vector, while the top-down mechanism determines feature weightings. Applying this approach to image captioning, our results on the MSCOCO test server establish a new state-of-the-art for the task, achieving CIDEr / SPICE / BLEU-4 scores of 117.9, 21.5 and 36.9, respectively. Demonstrating the broad applicability of the method, applying the same approach to VQA we obtain first place in the 2017 VQA Challenge.
Peter Anderson 0001, Xiaodong He 0001, Chris Buehler, Damien Teney, Mark Johnson 0001, Stephen Gould, Lei Zhang 0001
CVPR2
2018 CleanNet: Transfer Learning for Scalable Image Classifier Training With Label Noise
abstract
In this paper, we study the problem of learning image classification models with label noise. Existing approaches depending on human supervision are generally not scalable as manually identifying correct or incorrect labels is time-consuming, whereas approaches not relying on human supervision are scalable but less effective. To reduce the amount of human supervision for label noise cleaning, we introduce CleanNet, a joint neural embedding network, which only requires a fraction of the classes being manually verified to provide the knowledge of label noise that can be transferred to other classes. We further integrate CleanNet and conventional convolutional neural network classifier into one framework for image classification learning. We demonstrate the effectiveness of the proposed algorithm on both of the label noise detection task and the image classification on noisy data task on several large-scale datasets. Experimental results show that CleanNet can reduce label noise detection error rate on held-out classes where no human supervision available by 41.5% compared to current weakly supervised methods. It also achieves 47% of the performance gain of verifying all images with only 3.2% images verified on an image classification task. Source code and dataset will be available at kuanghuei.github.io/CleanNetProject.
Kuang-Huei Lee, Xiaodong He 0001, Lei Zhang 0001, Linjun Yang
CVPR2
2018 Tips and Tricks for Visual Question Answering: Learnings From the 2017 Challenge
abstract
Deep Learning has had a transformative impact on Computer Vision, but for all of the success there is also a significant cost. This is that the models and procedures used are so complex and intertwined that it is often impossible to distinguish the impact of the individual design and engineering choices each model embodies. This ambiguity diverts progress in the field, and leads to a situation where developing a state-of-the-art model is as much an art as a science. As a step towards addressing this problem we present a massive exploration of the effects of the myriad architectural and hyperparameter choices that must be made in generating a state-of-the-art model. The model is of particular interest because it won the 2017 Visual Question Answering Challenge. We provide a detailed analysis of the impact of each choice on model performance, in the hope that it will inform others in developing models, but also that it might set a precedent that will accelerate scientific progress in the field.
Damien Teney, Peter Anderson 0001, Xiaodong He 0001, Anton van den Hengel
CVPR3
2018 AttnGAN: Fine-Grained Text to Image Generation With Attentional Generative Adversarial Networks
abstract
In this paper, we propose an Attentional Generative Adversarial Network (AttnGAN) that allows attention-driven, multi-stage refinement for fine-grained text-to-image generation. With a novel attentional generative network, the AttnGAN can synthesize fine-grained details at different sub-regions of the image by paying attentions to the relevant words in the natural language description. In addition, a deep attentional multimodal similarity model is proposed to compute a fine-grained image-text matching loss for training the generator. The proposed AttnGAN significantly outperforms the previous state of the art, boosting the best reported inception score by 14.14% on the CUB dataset and 170.25% on the more challenging COCO dataset. A detailed analysis is also performed by visualizing the attention layers of the AttnGAN. It for the first time shows that the layered attentional GAN is able to automatically select the condition at the word level for generating different parts of the image.
Tao Xu 0029, Pengchuan Zhang, Qiuyuan Huang, Han Zhang 0010, Zhe Gan, Sharon X. Huang, Xiaodong He 0001
CVPR7
2018 Stacked Cross Attention for Image-Text Matching
Kuang-Huei Lee, Gang Hua 0001, Houdong Hu, Xiaodong He 0001
ECCV (4)5
2018 Policy Shaping and Generalized Update Equations for Semantic Parsing from Denotations
abstract
Semantic parsing from denotations faces two key challenges in model training: (1) given only the denotations (e.g., answers), search for good candidate semantic parses, and (2) choose the best model update algorithm.We propose effective and general solutions to each of them.Using policy shaping, we bias the search procedure towards semantic parses that are more compatible to the text, which provide better supervision signals for training.In addition, we propose an update equation that generalizes three different families of learning algorithms, which enables fast model exploration.When experimented on a recently proposed sequential question answering dataset, our framework leads to a new state-of-theart model that outperforms previous work by 5.0% absolute on exact match accuracy.Question: what nation scored the most points
Dipendra Misra, Ming-Wei Chang, Xiaodong He 0001, Scott Yih
EMNLP3
2018 Constrained Convolutional-Recurrent Networks to Improve Speech Quality with Low Impact on Recognition Accuracy
abstract
For a speech-enhancement algorithm, it is highly desirable to simultaneously improve perceptual quality and recognition rate. Thanks to computational costs and model complexities, it is challenging to train a model that effectively optimizes both metrics at the same time. In this paper, we propose a method for speech enhancement that combines local and global contextual structures information through convolutional-recurrent neural networks that improves perceptual quality. At the same time, we introduce a new constraint on the objective function using a language model/decoder that limits the impact on recognition rate. Based on experiments conducted with real user data, we demonstrate that our new context-augmented machine-learning approach for speech enhancement improves PESQ and WER by an additional 24.5% and 51.3%, respectively, when compared to the best-performing methods in the literature.
Rasool Fakoor, Xiaodong He 0001, Ivan Tashev, Shuayb Zarar
ICASSP2
2018 On the Discrimination-Generalization Tradeoff in GANs
Pengchuan Zhang, Qiang Liu 0001, Dengyong Zhou, Tao Xu 0029, Xiaodong He 0001
ICLR (Poster)5
2018 Discourse-Aware Neural Rewards for Coherent Text Generation
abstract
Antoine Bosselut, Asli Celikyilmaz, Xiaodong He, Jianfeng Gao, Po-Sen Huang, Yejin Choi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Antoine Bosselut, Asli Celikyilmaz, Xiaodong He 0001, Jianfeng Gao 0001, Po-Sen Huang, Yejin Choi 0001
NAACL-HLT3
2018 Deep Communicating Agents for Abstractive Summarization
abstract
Asli Celikyilmaz, Antoine Bosselut, Xiaodong He, Yejin Choi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Asli Celikyilmaz, Antoine Bosselut, Xiaodong He 0001, Yejin Choi 0001
NAACL-HLT3
2018 Tensor Product Generation Networks for Deep NLP Modeling
abstract
Qiuyuan Huang, Paul Smolensky, Xiaodong He, Li Deng, Dapeng Wu. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Qiuyuan Huang, Paul Smolensky, Xiaodong He 0001, Li Deng 0001, Dapeng Oliver Wu
NAACL-HLT3
2018 From Eliza to XiaoIce: challenges and opportunities with social chatbots
abstract
Conversational systems have come a long way since their inception in the 1960s. After decades of research and development, we have seen progress from Eliza and Parry in the 1960s and 1970s, to task-completion systems as in the Defense Advanced Research Projects Agency (DARPA) communicator program in the 2000s, to intelligent personal assistants such as Siri, in the 2010s, to today’s social chatbots like XiaoIce. Social chatbots’ appeal lies not only in their ability to respond to users’ diverse requests, but also in being able to establish an emotional connection with users. The latter is done by satisfying users’ need for communication, affection, as well as social belonging. To further the advancement and adoption of social chatbots, their design must focus on user engagement and take both intellectual quotient (IQ) and emotional quotient (EQ) into account. Users should want to engage with a social chatbot; as such, we define the success metric for social chatbots as conversation-turns per session (CPS). Using XiaoIce as an illustrative example, we discuss key technologies in building social chatbots from core chat to visual awareness to skills. We also show how XiaoIce can dynamically recognize emotion and engage the user throughout long conversations with appropriate interpersonal responses. As we become the first generation of humans ever living with artificial intelligenc (AI), we have a responsibility to design social chatbots to be both useful and empathetic, so they will become ubiquitous and help society as a whole.
Harry Shum, Xiaodong He 0001
Frontiers Inf. Technol. Electron. Eng.2
2017 Deep Learning with Low Precision by Half-Wave Gaussian Quantization
abstract
The problem of quantizing the activations of a deep neural network is considered. An examination of the popular binary quantization approach shows that this consists of approximating a classical non-linearity, the hyperbolic tangent, by two functions: a piecewise constant sign function, which is used in feedforward network computations, and a piecewise linear hard tanh function, used in the backpropagation step during network learning. The problem of approximating the widely used ReLU non-linearity is then considered. An half-wave Gaussian quantizer (HWGQ) is proposed for forward approximation and shown to have efficient implementation, by exploiting the statistics of of network activations and batch normalization operations. To overcome the problem of gradient mismatch, due to the use of different forward and backward approximations, several piece-wise backward approximators are then investigated. The implementation of the resulting quantized network, denoted as HWGQ-Net, is shown to achieve much closer performance to full precision networks, such as AlexNet, ResNet, GoogLeNet and VGG-Net, than previously available low-precision networks, with 1-bit binary weights and 2-bit quantized activations.
Zhaowei Cai, Xiaodong He 0001, Jian Sun 0001, Nuno Vasconcelos
CVPR2
2017 StyleNet: Generating Attractive Visual Captions with Styles
abstract
We propose a novel framework named StyleNet to address the task of generating attractive captions for images and videos with different styles. To this end, we devise a novel model component, named factored LSTM, which automatically distills the style factors in the monolingual text corpus. Then at runtime, we can explicitly control the style in the caption generation process so as to produce attractive visual captions with the desired style. Our approach achieves this goal by leveraging two sets of data: 1) factual image/video-caption paired data, and 2) stylized monolingual text data (e.g., romantic and humorous sentences). We show experimentally that StyleNet outperforms existing approaches for generating visual captions with different styles, measured in both automatic and human evaluation metrics on the newly collected FlickrStyle10K image caption dataset, which contains 10K Flickr images with corresponding humorous and romantic captions.
Chuang Gan 0001, Zhe Gan, Xiaodong He 0001, Jianfeng Gao 0001, Li Deng 0001
CVPR3
2017 Semantic Compositional Networks for Visual Captioning
abstract
A Semantic Compositional Network (SCN) is developed for image captioning, in which semantic concepts (i.e., tags) are detected from the image, and the probability of each tag is used to compose the parameters in a long short-term memory (LSTM) network. The SCN extends each weight matrix of the LSTM to an ensemble of tag-dependent weight matrices. The degree to which each member of the ensemble is used to generate an image caption is tied to the image-dependent probability of the corresponding tag. In addition to captioning images, we also extend the SCN to generate captions for video clips. We qualitatively analyze semantic composition in SCNs, and quantitatively evaluate the algorithm on three benchmark datasets: COCO, Flickr30k, and Youtube2Text. Experimental results show that the proposed method significantly outperforms prior state-of-the-art approaches, across multiple evaluation metrics.
Zhe Gan, Chuang Gan 0001, Xiaodong He 0001, Yunchen Pu, Kenneth Tran, Jianfeng Gao 0001, Lawrence Carin, Li Deng 0001
CVPR3
2017 Learning Generic Sentence Representations Using Convolutional Neural Networks
abstract
We propose a new encoder-decoder approach to learn distributed sentence representations that are applicable to multiple purposes.The model is learned by using a convolutional neural network as an encoder to map an input sentence into a continuous vector, and using a long short-term memory recurrent neural network as a decoder.Several tasks are considered, including sentence reconstruction and future sentence prediction.Further, a hierarchical encoderdecoder model is proposed to encode a sentence to predict multiple future sentences.By training our models on a large collection of novels, we obtain a highly generic convolutional sentence encoder that performs well in practice.Experimental results on several benchmark datasets, and across a broad range of applications, demonstrate the superiority of the proposed model over competing methods.
Zhe Gan, Yunchen Pu, Ricardo Henao, Chunyuan Li, Xiaodong He 0001, Lawrence Carin
EMNLP5
2017 Two-Stage Synthesis Networks for Transfer Learning in Machine Comprehension
abstract
We develop a technique for transfer learning in machine comprehension (MC) using a novel two-stage synthesis network (SynNet).Given a high-performing MC model in one domain, our technique aims to answer questions about documents in another domain, where we use no labeled data of question-answer pairs.Using the proposed SynNet with a pretrained model on the SQuAD dataset, we achieve an F1 measure of 46.6% on the challenging NewsQA dataset, approaching performance of in-domain models (F1 measure of 50.0%) and outperforming the out-ofdomain baseline by 7.6%, without use of provided annotations. 1
David Golub, Po-Sen Huang, Xiaodong He 0001, Li Deng 0001
EMNLP3
2017 Character-level deep conflation for business data analytics
abstract
Connecting different text attributes associated with the same entity (conflation) is important in business data analytics since it could help merge two different tables in a database to provide a more comprehensive profile of an entity. However, the conflation task is challenging because two text strings that describe the same entity could be quite different from each other for reasons such as misspelling. It is therefore critical to develop a conflation model that is able to truly understand the semantic meaning of the strings and match them at the semantic level. To this end, we develop a character-level deep conflation model that encodes the input text strings from character level into finite dimension feature vectors, which are then used to compute the cosine similarity between the text strings. The model is trained in an end-to-end manner using back propagation and stochastic gradient descent to maximize the likelihood of the correct association. Specifically, we propose two variants of the deep conflation model, based on long-short-term memory (LSTM) recurrent neural network (RNN) and convolutional neural network (CNN), respectively. Both models perform well on a real-world business analytics dataset and significantly outperform the baseline bag-of-character (BoC) model.
Zhe Gan, P. D. Singh, Ameet Joshi, Xiaodong He 0001, Jianshu Chen, Jianfeng Gao 0001, Li Deng 0001
ICASSP4
2017 Adversarial Ranking for Language Generation
abstract
Generative adversarial networks (GANs) have great successes on synthesizing data. However, the existing GANs restrict the discriminator to be a binary classifier, and thus limit their learning capacity for tasks that need to synthesize output with rich structures such as natural language descriptions. In this paper, we propose a novel generative adversarial network, RankGAN, for generating high-quality language descriptions. Rather than training the discriminator to learn and assign absolute binary predicate for individual data sample, the proposed RankGAN is able to analyze and rank a collection of human-written and machine-written sentences by giving a reference group. By viewing a set of data samples collectively and evaluating their quality through relative ranking scores, the discriminator is able to make better assessment which in turn helps to learn a better generator. The proposed RankGAN is optimized through the policy gradient technique. Experimental results on multiple public datasets clearly demonstrate the effectiveness of the proposed approach.
Dianqi Li, Xiaodong He 0001, Ming-Ting Sun, Zhengyou Zhang
NIPS3
2017 Editorial
Qun Liu 0001, Xiaodong He 0001, Hermann Ney
Mach. Transl.2
2016 Deep Reinforcement Learning with a Natural Language Action Space
abstract
Ji He, Jianshu Chen, Xiaodong He, Jianfeng Gao, Lihong Li, Li Deng, Mari Ostendorf. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Jianshu Chen, Xiaodong He 0001, Jianfeng Gao 0001, Lihong Li 0001, Li Deng 0001, Mari Ostendorf
ACL (1)3
2016 Generating Natural Questions About an Image
abstract
Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Margaret Mitchell, Xiaodong He, Lucy Vanderwende. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Margaret Mitchell, Xiaodong He 0001, Lucy Vanderwende
ACL (1)5
2016 Stacked Attention Networks for Image Question Answering
abstract
This paper presents stacked attention networks (SANs) that learn to answer natural language questions from images. SANs use semantic representation of a question as query to search for the regions in an image that are related to the answer. We argue that image question answering (QA) often requires multiple steps of reasoning. Thus, we develop a multiple-layer SAN in which we query an image multiple times to infer the answer progressively. Experiments conducted on four image QA data sets demonstrate that the proposed SANs significantly outperform previous state-of-the-art approaches. The visualization of the attention layers illustrates the progress that the SAN locates the relevant visual clues that lead to the answer of the question layer-by-layer.
Xiaodong He 0001, Jianfeng Gao 0001, Li Deng 0001, Alexander J. Smola
CVPR2
2016 MS-Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition
Yandong Guo, Lei Zhang 0001, Yuxiao Hu 0001, Xiaodong He 0001, Jianfeng Gao 0001
ECCV (3)4
2016 Bi-directional Attention with Agreement for Dependency Parsing
abstract
We develop a novel bi-directional attention model for dependency parsing, which learns to agree on headword predictions from the forward and backward parsing directions. The parsing procedure for each direction is formulated as sequentially querying the memory component that stores continuous headword embeddings. The proposed parser makes use of {\it soft} headword embeddings, allowing the model to implicitly capture high-order parsing history without dramatically increasing the computational complexity. We conduct experiments on English, Chinese, and 12 other languages from the CoNLL 2006 shared task, showing that the proposed model achieves state-of-the-art unlabeled attachment scores on 6 languages.
Hao Cheng 0002, Hao Fang 0002, Xiaodong He 0001, Jianfeng Gao 0001, Li Deng 0001
EMNLP3
2016 Character-Level Question Answering with Attention
abstract
We show that a character-level encoderdecoder framework can be successfully applied to question answering with a structured knowledge base.We use our model for singlerelation question answering and demonstrate the effectiveness of our approach on the Sim-pleQuestions dataset (Bordes et al., 2015), where we improve state-of-the-art accuracy from 63.9% to 70.9%, without use of ensembles.Importantly, our character-level model has 16x fewer parameters than an equivalent word-level model, can be learned with significantly less data compared to previous work, which relies on data augmentation, and is robust to new entities in testing. 1
Xiaodong He 0001, David Golub
EMNLP1
2016 Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads
abstract
We introduce an online popularity prediction and tracking task as a benchmark task for reinforcement learning with a combinatorial, natural language action space.A specified number of discussion threads predicted to be popular are recommended, chosen from a fixed window of recent comments to track.Novel deep reinforcement learning architectures are studied for effective modeling of the value function associated with actions comprised of interdependent sub-actions.The proposed model, which represents dependence between sub-actions through a bi-directional LSTM, gives the best performance across different experimental configurations and domains, and it also generalizes well with varying numbers of recommendation requests.
Mari Ostendorf, Xiaodong He 0001, Jianshu Chen, Jianfeng Gao 0001, Lihong Li 0001, Li Deng 0001
EMNLP3
2016 Zero-shot learning of intent embeddings for expansion by convolutional deep structured semantic models
abstract
The recent surge of intelligent personal assistants motivates spoken language understanding of dialogue systems. However, the domain constraint along with the inflexible intent schema remains a big issue. This paper focuses on the task of intent expansion, which helps remove the domain limit and make an intent schema flexible. A con-volutional deep structured semantic model (CDSSM) is applied to jointly learn the representations for human intents and associated utterances. Then it can flexibly generate new intent embeddings without the need of training samples and model-retraining, which bridges the semantic relation between seen and unseen intents and further performs more robust results. Experiments show that CDSSM is capable of performing zero-shot learning effectively, e.g. generating embeddings of previously unseen intents, and therefore expand to new intents without re-training, and outperforms other semantic embeddings. The discussion and analysis of experiments provide a future direction for reducing human effort about annotating data and removing the domain constraint in spoken dialogue systems.
Yun-Nung Chen, Dilek Hakkani-Tür, Xiaodong He 0001
ICASSP3
2016 Interpreting the prediction process of a deep network constructed from supervised topic models
abstract
In this paper, we propose an approach to interpret the prediction process of the BP-sLDA model, which is a supervised Latent Dirichlet Allocation model trained by Back Propagation over a deep architecture. The model is shown to achieve state-of-the-art prediction performance on several large-scale text analysis tasks. To interpret the prediction process of the model, often demanded by business data analytics applications, we perform evidence analysis on each pair-wise decision boundary over the topic distribution space, which is decomposed into a positive and a negative components. Then, for each element in the current document, a novel evidence score is defined by exploiting this topic decomposition and the generative nature of LDA. Then the score is used to rank the relative evidence of each element for the effectiveness of model prediction. We demonstrate the effectiveness of the method on a large-scale binary classification task on a corporate proprietary dataset with business-centric applications.
Jianshu Chen, Xiaodong He 0001, Jianfeng Gao 0001, Li Deng 0001
ICASSP3
2016 Visual Storytelling
abstract
Ting-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, Margaret Mitchell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Ting-Hao 'Kenneth' Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross B. Girshick, Xiaodong He 0001, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, Margaret Mitchell
HLT-NAACL8
2016 A Corpus and Cloze Evaluation for Deeper Understanding of Commonsense Stories
abstract
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, James Allen. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He 0001, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, James F. Allen
HLT-NAACL3
2016 Hierarchical Attention Networks for Document Classification
abstract
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, Eduard Hovy. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Diyi Yang, Chris Dyer, Xiaodong He 0001, Alexander J. Smola, Eduard H. Hovy
HLT-NAACL4
2016 Multi-Rate Deep Learning for Temporal Recommendation
abstract
Modeling temporal behavior in recommendation systems is an important and challenging problem. Its challenges come from the fact that temporal modeling increases the cost of parameter estimation and inference, while requiring large amount of data to reliably learn the model with the additional time dimensions. Therefore, it is often difficult to model temporal behavior in large-scale real-world recommendation systems. In this work, we propose a novel deep neural network based architecture that models the combination of long-term static and short-term temporal user preferences to improve the recommendation performance. To train the model efficiently for large-scale applications, we propose a novel pre-train method to reduce the number of free parameters significantly. The resulted model is applied to a real-world data set from a commercial News recommendation system. We compare to a set of established baselines and the experimental results show that our method outperforms the state-of-the-art significantly.
Yang Song 0008, Ali Mamdouh Elkahky, Xiaodong He 0001
SIGIR3
2016 Table Cell Search for Question Answering
abstract
Tables are pervasive on the Web. Informative web tables range across a large variety of topics, which can naturally serve as a significant resource to satisfy user information needs. Driven by such observations, in this paper, we investigate an important yet largely under-addressed problem: Given millions of tables, how to precisely retrieve table cells to answer a user question. This work proposes a novel table cell search framework to attack this problem. We first formulate the concept of a relational chain which connects two cells in a table and represents the semantic relation between them. With the help of search engine snippets, our framework generates a set of relational chains pointing to potentially correct answer cells. We further employ deep neural networks to conduct more fine-grained inference on which relational chains best match the input question and finally extract the corresponding answer cells. Based on millions of tables crawled from the Web, we evaluate our framework in the open-domain question answering (QA) setting, using both the well-known WebQuestions dataset and user queries mined from Bing search engine logs. On WebQuestions, our framework is comparable to state-of-the-art QA systems based on knowledge bases (KBs), while on Bing queries, it outperforms other systems with a 56.7% relative gain. Moreover, when combined with results from our framework, KB-based QA performance can obtain a relative improvement of 28.1% to 66.7%, demonstrating that web tables supply rich knowledge that might not exist or is difficult to be identified in existing KBs.
Huan Sun 0001, Hao Ma 0001, Xiaodong He 0001, Scott Yih, Yu Su 0001, Xifeng Yan
WWW3
2016 Deep Sentence Embedding Using Long Short-Term Memory Networks: Analysis and Application to Information Retrieval
abstract
This paper develops a model that addresses sentence embedding, a hot topic in current natural language processing research, using recurrent neural networks (RNN) with Long Short-Term Memory (LSTM) cells. The proposed LSTM-RNN model sequentially takes each word in a sentence, extracts its information, and embeds it into a semantic vector. Due to its ability to capture long term memory, the LSTM-RNN accumulates increasingly richer information as it goes through the sentence, and when it reaches the last word, the hidden layer of the network provides a semantic representation of the whole sentence. In this paper, the LSTM-RNN is trained in a weakly supervised manner on user click-through data logged by a commercial web search engine. Visualization and analysis are performed to understand how the embedding process works. The model is found to automatically attenuate the unimportant words and detect the salient keywords in the sentence. Furthermore, these detected keywords are found to automatically activate different cells of the LSTM-RNN, where words belonging to a similar topic activate the same cell. As a semantic representation of the sentence, the embedding vector can be used in many different applications. These automatic keyword detection and topic allocation abilities enabled by the LSTM-RNN allow the network to perform document retrieval, a difficult language processing task, where the similarity between the query and documents can be measured by the distance between their corresponding sentence embedding vectors computed by the LSTM-RNN. On a web search task, the LSTM-RNN embedding is shown to significantly outperform several existing state of the art methods. We emphasize that the proposed model generates sentence embedding vectors that are specially useful for web document retrieval tasks. A comparison with a well known general sentence embedding method, the Paragraph Vector, is performed. The results show that the proposed method in this paper significantly outperforms Paragraph Vector method for web document retrieval task.
Hamid Palangi, Li Deng 0001, Yelong Shen, Jianfeng Gao 0001, Xiaodong He 0001, Jianshu Chen, Xinying Song, Rabab K. Ward
IEEE ACM Trans. Audio Speech Lang. Process.5
2015 Semantic Parsing via Staged Query Graph Generation: Question Answering with Knowledge Base
abstract
Wen-tau Yih, Ming-Wei Chang, Xiaodong He, Jianfeng Gao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Scott Yih, Ming-Wei Chang, Xiaodong He 0001, Jianfeng Gao 0001
ACL (1)3
2015 Detecting actionable items in meetings by convolutional deep structured semantic models
abstract
The recent success of voice interaction with smart devices (human-machine genre) and improvements in speech recognition for conversational speech show the possibility of conversation-related applications. This paper investigates the task of actionable item detection in meetings (human-human genre), where the intelligent assistant dynamically provides the participants access to information (e.g. scheduling a meeting, taking notes) without interrupting the meetings. A convolutional deep structured semantic model (CDSSM) is applied to learn the latent semantics for human actions and utterances from human-machine (source genre) and human-human (target) interactions. Furthermore, considering the mismatch between source and target genre and scarcity of annotated data sets for the target genre, we develop adaptation techniques that adjust the learned embeddings to better fit the target genre. Experiments show that CDSSM performs better for actionable item detection compared to baselines using lexical features (27.5% relative) and other semantic features (15.9% relative) when the source genre and target genre match with each other. When the target genre mismatches with the source genre, our proposed adaptation techniques further improve the performance. The discussion and analysis of the experiments provide a reasonable direction for such an actionable item detection task1.
Yun-Nung Chen, Dilek Hakkani-Tür, Xiaodong He 0001
ASRU3
2015 From captions to visual concepts and back
abstract
This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time.
Hao Fang 0002, Saurabh Gupta 0001, Forrest N. Iandola, Rupesh Kumar Srivastava, Li Deng 0001, Piotr Dollár, Jianfeng Gao 0001, Xiaodong He 0001, Margaret Mitchell, John C. Platt, C. Lawrence Zitnick, Geoffrey Zweig
CVPR8
2015 Representation Learning Using Multi-Task Deep Neural Networks for Semantic Classification and Information Retrieval
abstract
Xiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, Ye-yi Wang. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Xiaodong Liu 0003, Jianfeng Gao 0001, Xiaodong He 0001, Li Deng 0001, Kevin Duh, Ye-Yi Wang
HLT-NAACL3
2015 Deep Learning and Continuous Representations for Natural Language Processing
abstract
Deep learning techniques have demonstrated tremendous success in the speech and language processing community in recent years, establishing new state-ofthe-art performance in speech recognition, language modeling, and have shown great potential for many other natural language processing tasks. The focus of this tutorial is to provide an extensive overview on recent deep learning approaches to problems in language or text processing, with particular emphasis on important real-world applications including language understanding, semantic representation modeling, question answering and semantic parsing, etc.
Scott Yih, Xiaodong He 0001, Jianfeng Gao 0001
HLT-NAACL2
2015 End-to-end Learning of LDA by Mirror-Descent Back Propagation over a Deep Architecture
abstract
We develop a fully discriminative learning approach for supervised Latent Dirichlet Allocation (LDA) model using Back Propagation (i.e., BP-sLDA), which maximizes the posterior probability of the prediction variable given the input document. Different from traditional variational learning or Gibbs sampling approaches, the proposed learning method applies (i) the mirror descent algorithm for maximum a posterior inference and (ii) back propagation over a deep architecture together with stochastic gradient/mirror descent for model parameter estimation, leading to scalable and end-to-end discriminative learning of the model. As a byproduct, we also apply this technique to develop a new learning method for the traditional unsupervised LDA model (i.e., BP-LDA). Experimental results on three real-world regression and classification tasks show that the proposed methods significantly outperform the previous supervised topic models, neural networks, and is on par with deep neural networks.
Jianshu Chen, Yelong Shen, Xiaodong He 0001, Jianfeng Gao 0001, Xinying Song, Li Deng 0001
NIPS5
2015 A Multi-View Deep Learning Approach for Cross Domain User Modeling in Recommendation Systems
abstract
Recent online services rely heavily on automatic personalization to recommend relevant content to a large number of users. This requires systems to scale promptly to accommodate the stream of new users visiting the online services for the first time. In this work, we propose a content-based recommendation system to address both the recommendation quality and the system scalability. We propose to use a rich feature set to represent users, according to their web browsing history and search queries. We use a Deep Learning approach to map users and items to a latent space where the similarity between users and their preferred items is maximized. We extend the model to jointly learn from features of items from different domains and user features by introducing a multi-view Deep Learning model. We show how to make this rich-feature based user representation scalable by reducing the dimension of the inputs and the amount of training data. The rich user feature representation allows the model to learn relevant user behavior patterns and give useful recommendations for users who do not have any interaction with the service, given that they have adequate search and browsing history. The combination of different domains into a single model for learning helps improve the recommendation quality across all the domains, as well as having a more compact and a semantically richer user latent feature vector. We experiment with our approach on three real-world recommendation systems acquired from different sources of Microsoft products: Windows Apps recommendation, News recommendation, and Movie/TV recommendation. Results indicate that our approach is significantly better than the state-of-the-art algorithms (up to 49% enhancement on existing users and 115% enhancement on new users). In addition, experiments on a publicly open data set also indicate the superiority of our method in comparison with transitional generative topic models, for modeling cross-domain recommender systems. Scalability analysis show that our multi-view DNN model can easily scale to encompass millions of users and billions of item entries. Experimental results also confirm that combining features from all domains produces much better performance than building separate models for each domain.
Ali Mamdouh Elkahky, Yang Song 0008, Xiaodong He 0001
WWW3
2015 Introduction to the Special Section on Continuous Space and Related Methods in Natural Language Processing
abstract
The articles in this special section discuss some latest findings on research problems related to the application of continuous space and related models in Natural Language Processing (NLP).
Haizhou Li 0001, Marcello Federico, Xiaodong He 0001, Helen M. Meng, Isabel Trancoso
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Using Recurrent Neural Networks for Slot Filling in Spoken Language Understanding
abstract
Semantic slot filling is one of the most challenging problems in spoken language understanding (SLU). In this paper, we propose to use recurrent neural networks (RNNs) for this task, and present several novel architectures designed to efficiently model past and future temporal dependencies. Specifically, we implemented and compared several important RNN architectures, including Elman, Jordan, and hybrid variants. To facilitate reproducibility, we implemented these networks with the publicly available Theano neural network toolkit and completed experiments on the well-known airline travel information system (ATIS) benchmark. In addition, we compared the approaches on two custom SLU data sets from the entertainment and movies domains. Our results show that the RNN-based models outperform the conditional random field (CRF) baseline by 2% in absolute error reduction on the ATIS benchmark. We improve the state-of-the-art by 0.5% in the Entertainment domain, and 6.7% for the movies domain.
Grégoire Mesnil, Yann N. Dauphin, Kaisheng Yao, Yoshua Bengio, Li Deng 0001, Dilek Hakkani-Tür, Xiaodong He 0001, Larry Heck, Gökhan Tür, Dong Yu 0001, Geoffrey Zweig
IEEE ACM Trans. Audio Speech Lang. Process.7
2014 Learning Continuous Phrase Representations for Translation Modeling
abstract
This paper tackles the sparsity problem in estimating phrase translation probabilities by learning continuous phrase representations, whose distributed nature enables the sharing of related phrases in their representations.A pair of source and target phrases are projected into continuous-valued vector representations in a low-dimensional latent space, where their translation score is computed by the distance between the pair in this new space.The projection is performed by a neural network whose weights are learned on parallel training data.Experimental evaluation has been performed on two WMT translation tasks.Our best result improves the performance of a state-of-the-art phrase-based statistical machine translation system trained on WMT 2012 French-English data by up to 1.3 BLEU points.
Jianfeng Gao 0001, Xiaodong He 0001, Scott Yih, Li Deng 0001
ACL (1)2
2014 A Latent Semantic Model with Convolutional-Pooling Structure for Information Retrieval
abstract
In this paper, we propose a new latent semantic model that incorporates a convolutional-pooling structure over word sequences to learn low-dimensional, semantic vector representations for search queries and Web documents. In order to capture the rich contextual structures in a query or a document, we start with each word within a temporal context window in a word sequence to directly capture contextual features at the word n-gram level. Next, the salient word n-gram features in the word sequence are discovered by the model and are then aggregated to form a sentence-level feature vector. Finally, a non-linear transformation is applied to extract high-level semantic information to generate a continuous vector representation for the full text string. The proposed convolutional latent semantic model (CLSM) is trained on clickthrough data and is evaluated on a Web document ranking task using a large-scale, real-world data set. Results show that the proposed model effectively captures salient semantic information in queries and documents for the task while significantly outperforming previous state-of-the-art semantic models.
Yelong Shen, Xiaodong He 0001, Jianfeng Gao 0001, Li Deng 0001, Grégoire Mesnil
CIKM2
2014 Modeling Interestingness with Deep Neural Networks
abstract
This paper presents a deep semantic similarity model (DSSM), a special type of deep neural networks designed for text analysis, for recommending target documents to be of interest to a user based on a source document that she is reading.We observe, identify, and detect naturally occurring signals of interestingness in click transitions on the Web between source and target documents, which we collect from commercial Web browser logs.The DSSM is trained on millions of Web transitions, and maps source-target document pairs to feature vectors in a latent space in such a way that the distance between source documents and their corresponding interesting targets in that space is minimized.The effectiveness of the DSSM is demonstrated using two interestingness tasks: automatic highlighting and contextual entity search.The results on large-scale, real-world datasets show that the semantics of documents are important for modeling interestingness and that the DSSM leads to significant quality improvement on both tasks, outperforming not only the classic document models that do not use semantics but also state-of-the-art topic models.
Jianfeng Gao 0001, Patrick Pantel, Michael Gamon, Xiaodong He 0001, Li Deng 0001
EMNLP4
2014 Modeling action-level satisfaction for search task satisfaction prediction
abstract
Search satisfaction is a property of a user's search process. Understanding it is critical for search providers to evaluate the performance and improve the effectiveness of search engines. Existing methods model search satisfaction holistically at the search-task level, ignoring important dependencies between action-level satisfaction and overall task satisfaction. We hypothesize that searchers' latent action-level satisfaction (i.e., whether they believe they were satisfied with the results of a query or click) influences their observed search behaviors and contributes to overall search satisfaction. We conjecture that by modeling search satisfaction at the action level, we can build more complete and more accurate predictors of search-task satisfaction. To do this, we develop a latent structural learning method, whereby rich structured features and dependency relations unique to search satisfaction prediction are explored. Using in-situ search satisfaction judgments provided by searchers, we show that there is significant value in modeling action-level satisfaction in search-task satisfaction prediction. In addition, experimental results on large-scale logs from Bing.com demonstrate clear benefit from using inferred action satisfaction labels for other applications such as document relevance estimation and query suggestion.
Hongning Wang, Yang Song 0008, Ming-Wei Chang, Xiaodong He 0001, Ahmed Awadallah 0001, Ryen W. White
SIGIR4
2014 Adapting deep RankNet for personalized search
abstract
RankNet is one of the widely adopted ranking models for web search tasks. However, adapting a generic RankNet for personalized search is little studied. In this paper, we first continue-trained a variety of RankNets with different number of hidden layers and network structures over a previously trained global RankNet model, and observed that a deep neural network with five hidden layers gives the best performance. To further improve the performance of adaptation, we propose a set of novel methods categorized into two groups. In the first group, three methods are proposed to properly assess the usefulness of each adaptation instance and only leverage the most informative instances to adapt a user-specific RankNet model. These assessments are based on KL-divergence, click entropy or a heuristic to ignore top clicks in adaptation queries. In the second group, two methods are proposed to regularize the training of the neural network in RankNet: one of these methods regularize the error back-propagation via a truncated gradient approach, while the other method limits the depth of the back propagation when adapting the neural network. We empirically evaluate our approaches using a large-scale real-world data set. Experimental results exhibit that our methods all give significant improvements over a strong baseline ranking system, and the truncated gradient approach gives the best performance, significantly better than all others.
Yang Song 0008, Hongning Wang, Xiaodong He 0001
WSDM3
2013 Learning deep structured semantic models for web search using clickthrough data
abstract
Latent semantic models, such as LSA, intend to map a query to its relevant documents at the semantic level where keyword-based matching often fails. In this study we strive to develop a series of new latent semantic models with a deep structure that project queries and documents into a common low-dimensional space where the relevance of a document given a query is readily computed as the distance between them. The proposed deep structured semantic models are discriminatively trained by maximizing the conditional likelihood of the clicked documents given a query using the clickthrough data. To make our models applicable to large-scale Web search applications, we also use a technique called word hashing, which is shown to effectively scale up our semantic models to handle large vocabularies which are common in such tasks. The new models are evaluated on a Web document ranking task using a real-world data set. Results show that our best model significantly outperforms other latent semantic models, which were considered state-of-the-art in the performance prior to the work presented in this paper.
Po-Sen Huang, Xiaodong He 0001, Jianfeng Gao 0001, Li Deng 0001, Alex Acero, Larry Heck
CIKM2
2013 Deep stacking networks for information retrieval
abstract
Deep stacking networks (DSN) are a special type of deep model equipped with parallel and scalable learning. We report successful applications of DSN to an information retrieval (IR) task pertaining to relevance prediction for sponsor search after careful regularization methods are incorporated to the previous DSN methods developed for speech and image classification tasks. The DSN-based system significantly outperforms the LambdaRank-based system which represents a recent state-of-the-art for IR in normalized discounted cumulative gain (NDCG) measures, despite the use of mean square error as DSN's training objective. We demonstrate desirable monotonic correlation between NDCG and classification rate in a wide range of IR quality. The weaker correlation and more flat relationship in the high IR-quality region suggest the need for developing new learning objectives and optimization methods.
Li Deng 0001, Xiaodong He 0001, Jianfeng Gao 0001
ICASSP2
2013 Recent advances in deep learning for speech research at Microsoft
abstract
Deep learning is becoming a mainstream technology for speech recognition at industrial scale. In this paper, we provide an overview of the work by Microsoft speech researchers since 2009 in this area, focusing on more recent advances which shed light to the basic capabilities and limitations of the current deep learning technology. We organize this overview along the feature-domain and model-domain dimensions according to the conventional approach to analyzing speech systems. Selected experimental results, including speech recognition and related applications such as spoken dialogue and language modeling, are presented to demonstrate and analyze the strengths and weaknesses of the techniques described in the paper. Potential improvement of these techniques and future research directions are discussed.
Li Deng 0001, Jinyu Li 0001, Jui-Ting Huang, Kaisheng Yao, Dong Yu 0001, Frank Seide, Michael L. Seltzer, Geoffrey Zweig, Xiaodong He 0001, Jason D. Williams, Yifan Gong 0001, Alex Acero
ICASSP9
2013 End-to-end learning of parsing models for information retrieval
abstract
Parsers have been shown to be helpful in information retrieval tasks because they are able to model long-span word dependencies efficiently. While previous work focused on using traditional syntactic parse trees, this paper proposes a new approach where, unlike previous work, the parser parameters are discriminatively trained to directly optimize a non-convex and non-smooth IR measure. The relevance between a document and a query is then modeled by the weighted tree edit distance between their parses. We evaluate our method on a large scale web search task consisting of a real world query set. Results show that the new parser is more effective for document retrieval than using traditional syntactic parse trees. It gives significant improvement, especially for long queries where proper modeling of long-span dependencies is crucial.
Jennifer Gillenwater, Xiaodong He 0001, Jianfeng Gao 0001, Li Deng 0001
ICASSP2
2013 Multi-style adaptive training for robust cross-lingual spoken language understanding
abstract
Given the increasingly available machine translation (MT) services nowadays, one efficient strategy for cross-lingual spoken language understanding (SLU) is to first translate the input utterance from the second language into the primary language, and then call the primary language SLU system to decode the semantic knowledge. However, errors introduced in the MT process create a condition similar to the “mismatch” condition encountered in robust speech recognition. Such mismatch makes the performance of cross-lingual SLU far from acceptable. Motivated by successful solutions developed in robust speech recognition, we in this paper propose a multi-style adaptive training method to improve the robustness of the SLU system for cross-lingual SLU tasks. For evaluation, we created an English-Chinese bilingual ATIS database, and then carried out a series of experiments on that database to experimentally assess the proposed methods. Experimental results show that, without relying on any data in the second language, the proposed method significantly improves the performance on a cross-lingual SLU task while producing no degradation for input in the primary language. This greatly facilitates porting SLU to as many languages as there are MT systems without any human effort. We further study the robustness of this approach to another type of mismatch condition, caused by speech recognition errors, and demonstrate its success also.
Xiaodong He 0001, Li Deng 0001, Dilek Hakkani-Tür, Gökhan Tür
ICASSP1
2013 Random features for Kernel Deep Convex Network
abstract
The recently developed deep learning architecture, a kernel version of the deep convex network (K-DCN), is improved to address the scalability problem when the training and testing samples become very large. We have developed a solution based on the use of random Fourier features, which possess the strong theoretical property of approximating the Gaussian kernel while rendering efficient computation in both training and evaluation of the K-DCN with large training samples. We empirically demonstrate that just like the conventional K-DCN exploiting rigorous Gaussian kernels, the use of random Fourier features also enables successful stacking of kernel modules to form a deep architecture. Our evaluation experiments on phone recognition and speech understanding tasks both show the computational efficiency of the K-DCN which makes use of random features. With sufficient depth in the K-DCN, the phone recognition accuracy and slot-filling accuracy are shown to be comparable or slightly higher than the K-DCN with Gaussian kernels while significant computational saving has been achieved.
Po-Sen Huang, Li Deng 0001, Mark Hasegawa-Johnson, Xiaodong He 0001
ICASSP4
2013 Investigation of recurrent-neural-network architectures and learning methods for spoken language understanding
abstract
One of the key problems in spoken language understanding (SLU) is the task of slot filling. In light of the recent success of applying deep neural network technologies in domain detection and intent identification, we carried out an in-depth investigation on the use of recurrent neural networks for the more difficult task of slot filling involving sequence discrimination. In this work, we implemented and compared several important recurrent-neural-network architectures, including the Elman-type and Jordan-type recurrent networks and their variants. To make the results easy to reproduce and compare, we implemented these networks on the common Theano neural network toolkit, and evaluated them on the ATIS benchmark. We also compared our results to a conditional random fields (CRF) baseline. Our results show that on this task, both types of recurrent networks outperform the CRF baseline substantially, and a bi-directional Jordantype network that takes into account both past and future dependencies among slots works best, outperforming a CRFbased baseline by 14% in relative error reduction.
Grégoire Mesnil, Xiaodong He 0001, Li Deng 0001, Yoshua Bengio
INTERSPEECH2
2013 Training MRF-Based Phrase Translation Models using Gradient Ascent
Jianfeng Gao 0001, Xiaodong He 0001
HLT-NAACL2
2013 Personalized ranking model adaptation for web search
abstract
Search engines train and apply a single ranking model across all users, but searchers' information needs are diverse and cover a broad range of topics. Hence, a single user-independent ranking model is insufficient to satisfy different users' result preferences. Conventional personalization methods learn separate models of user interests and use those to re-rank the results from the generic model. Those methods require significant user history information to learn user preferences, have low coverage in the case of memory-based methods that learn direct associations between query-URL pairs, and have limited opportunity to markedly affect the ranking given that they only re-order top-ranked items.
Hongning Wang, Xiaodong He 0001, Ming-Wei Chang, Yang Song 0008, Ryen W. White
SIGIR2
2013 Learning to extract cross-session search tasks
abstract
Search tasks, comprising a series of search queries serving the same information need, have recently been recognized as an accurate atomic unit for modeling user search intent. Most prior research in this area has focused on short-term search tasks within a single search session, and heavily depend on human annotations for supervised classification model learning. In this work, we target the identification of long-term, or cross-session, search tasks (transcending session boundaries) by investigating inter-query dependencies learned from users' searching behaviors. A semi-supervised clustering model is proposed based on the latent structural SVM framework, and a set of effective automatic annotation rules are proposed as weak supervision to release the burden of manual annotation. Experimental results based on a large-scale search log collected from Bing.com confirms the effectiveness of the proposed model in identifying cross-session search tasks and the utility of the introduced weak supervision signals. Our learned model enables a more comprehensive understanding of users' search behaviors via search logs and facilitates the development of dedicated search-engine support for long-term tasks.
Hongning Wang, Yang Song 0008, Ming-Wei Chang, Xiaodong He 0001, Ryen W. White
WWW4
2013 Enhancing personalized search by mining and modeling task behavior
abstract
Personalized search systems tailor search results to the current user intent using historic search interactions. This relies on being able to find pertinent information in that user's search history, which can be challenging for unseen queries and for new search scenarios. Building richer models of users' current and historic search tasks can help improve the likelihood of finding relevant content and enhance the relevance and coverage of personalization methods. The task-based approach can be applied to the current user's search history, or as we focus on here, all users' search histories as so-called "groupization" (a variant of personalization whereby other users' profiles can be used to personalize the search experience). We describe a method whereby we mine historic search-engine logs to find other users performing similar tasks to the current user and leverage their on-task behavior to identify Web pages to promote in the current ranking. We investigate the effectiveness of this approach versus query-based matching and finding related historic activity from the current user (i.e., group versus individual). As part of our studies we also explore the use of the on-task behavior of particular user cohorts, such as people who are expert in the topic currently being searched, rather than all other users. Our approach yields promising gains in retrieval performance, and has direct implications for improving personalization in search systems.
Ryen W. White, Ahmed Awadallah 0001, Xiaodong He 0001, Yang Song 0008, Hongning Wang
WWW4
2013 Speech-Centric Information Processing: An Optimization-Oriented Approach
abstract
Automatic speech recognition (ASR) is a central and common component of voice-driven information processing systems in human language technology, including spoken language translation (SLT), spoken language understanding (SLU), voice search, spoken document retrieval, and so on. Interfacing ASR with its downstream text-based processing tasks of translation, understanding, and information retrieval (IR) creates both challenges and opportunities in optimal design of the combined, speech-enabled systems. We present an optimization-oriented statistical framework for the overall system design where the interactions between the subsystems in tandem are fully incorporated and where design consistency is established between the optimization objectives and the end-to-end system performance metrics. Techniques for optimizing such objectives in both the decoding and learning phases of the speech-centric information processing (SCIP) system design are described, in which the uncertainty in speech recognition subsystem's outputs is fully considered and marginalized. This paper provides an overview of the past and current work in this area. Future challenges and new opportunities are also discussed and analyzed.
Xiaodong He 0001, Li Deng 0001
Proc. IEEE1
2013 Optimization Algorithms and Applications for Speech and Language Processing
abstract
Optimization techniques have been used for many years in the formulation and solution of computational problems arising in speech and language processing. Such techniques are found in the Baum-Welch, extended Baum-Welch (EBW), Rprop, and GIS algorithms, for example. Additionally, the use of regularization terms has been seen in other applications of sparse optimization. This paper outlines a range of problems in which optimization formulations and algorithms play a role, giving some additional details on certain application problems in machine translation, speaker/language recognition, and automatic speech recognition. Several approaches developed in the speech and language processing communities are described in a way that makes them more recognizable as optimization procedures. Our survey is not exhaustive and is complemented by other papers in this volume.
Stephen J. Wright 0001, Dimitri Kanevsky, Li Deng 0001, Xiaodong He 0001, Georg Heigold, Haizhou Li 0001
IEEE Trans. Speech Audio Process.4
2012 Maximum Expected BLEU Training of Phrase and Lexicon Translation Models
Xiaodong He 0001, Li Deng 0001
ACL (1)1
2012 Learning Lexicon Models from Search Logs for Query Expansion
Jianfeng Gao 0001, Shasha Xie, Xiaodong He 0001, Alnur Ali
EMNLP-CoNLL3
2012 New methods and evaluation experiments on translating TED talks in the IWSLT benchmark
abstract
The IWSLT benchmark task is an annual evaluation campaign on spoken language translation held by the International Workshop on Spoken Language Processing (IWSLT). The task is to translate TED talks (www.ted.com). This task presents two unique challenges: Firstly, the underlying topic switches sharply from talk to talk, and each one contains only tens to hundreds of utterances. The translation system therefore needs to adapt to the current topic quickly and dynamically. Secondly, unlike other machine translation benchmark tasks, only a very small relevant parallel corpus (transcripts of TED talks) is available. Therefore, it is necessary to perform accurate translation model estimation with limited data. In this paper, we present our recent progress and two new methods on the IWSLT TED talk translation task from Chinese into English. In particular, to address the first problem, we use unsupervised topic modeling to select additional topic-dependent parallel data from a globally irrelevant corpus. These additional data slices can then be used to build an unsupervised topic-adapted machine translation system. For the second problem, we develop a discriminative training method to estimate the translation models more accurately. Our experimental evaluation results show that both methods improve the translation quality over a state-of-the-art baseline.
Amittai Axelrod, Xiaodong He 0001, Li Deng 0001, Alex Acero, Mei-Yuh Hwang
ICASSP2
2012 Optimization in speech-centric information processing: Criteria and techniques
abstract
Automatic speech recognition (ASR) is an enabling technology for a wide range of information processing applications including speech translation, voice search (i.e., information retrieval with speech input), and conversational understanding. In these speech-centric applications, the output of ASR as “noisy” text is fed into down-stream processing systems to accomplish the designated tasks of translation, information retrieval, or natural language understanding, etc. In conventional applications, the ASR model as a sub-system is usually trained without considering the down-stream systems. This often leads to sub-optimal end-to-end performance. In this paper, we propose a unifying end-to-end optimization framework in which the model parameters in all sub-systems including ASR are learned by Extended Baum-Welch (EBW) algorithms via optimizing the criteria directly tied to the end-to-end performance measure. We demonstrate the effectiveness of the proposed approach on a speech translation task using the spoken language translation benchmark test of IWSLT. Our experimental results show that the proposed method leads to significant improvement of translation quality over the conventional techniques based on separate modular sub-system design. We also analyze the EBW-based optimization algorithms employed in our work and discuss its relationship with other popular optimization techniques.
Xiaodong He 0001, Li Deng 0001
ICASSP1
2012 Towards deeper understanding: Deep convex networks for semantic utterance classification
abstract
Following the recent advances in deep learning techniques, in this paper, we present the application of special type of deep architecture - deep convex networks (DCNs) - for semantic utterance classification (SUC). DCNs are shown to have several advantages over deep belief networks (DBNs) including classification accuracy and training scalability. However, adoption of DCNs for SUC comes with non-trivial issues. Specifically, SUC has an extremely sparse input feature space encompassing a very large number of lexical and semantic features. This is about a few thousand times larger than the feature space for acoustic modeling, yet with a much smaller number of training samples. Experimental results we obtained on a domain classification task for spoken language understanding demonstrate the effectiveness of DCNs. The DCN-based method produces higher SUC accuracy than the Boosting-based discriminative classifier with word trigrams.
Gökhan Tür, Li Deng 0001, Dilek Hakkani-Tür, Xiaodong He 0001
ICASSP4
2012 Use of kernel deep convex networks and end-to-end learning for spoken language understanding
abstract
We present our recent and ongoing work on applying deep learning techniques to spoken language understanding (SLU) problems. The previously developed deep convex network (DCN) is extended to its kernel version (K-DCN) where the number of hidden units in each DCN layer approaches infinity using the kernel trick. We report experimental results demonstrating dramatic error reduction achieved by the K-DCN over both the Boosting-based baseline and the DCN on a domain classification task of SLU, especially when a highly correlated set of features extracted from search query click logs are used. Not only can DCN and K-DCN be used as a domain or intent classifier for SLU, they can also be used as local, discriminative feature extractors for the slot filling task of SLU. The interface of K-DCN to slot filling systems via the softmax function is presented. Finally, we outline an end-to-end learning strategy for training the softmax parameters (and potentially all DCN and K-DCN parameters) where the learning objective can take any performance measure (e.g. the F-measure) for the full SLU system.
Li Deng 0001, Gökhan Tür, Xiaodong He 0001, Dilek Hakkani-Tür
SLT3
2011 Domain Adaptation via Pseudo In-Domain Data Selection
Amittai Axelrod, Xiaodong He 0001, Jianfeng Gao 0001
EMNLP2
2011 Why word error rate is not a good metric for speech recognizer training for the speech translation task?
abstract
Speech translation (ST) is an enabling technology for cross-lingual oral communication. A ST system consists of two major components: an automatic speech recognizer (ASR) and a machine translator (MT). Nowadays, most ASR systems are trained and tuned by minimizing word error rate (WER). However, WER counts word errors at the surface level. It does not consider the contextual and syntactic roles of a word, which are often critical for MT. In the end-to-end ST scenarios, whether WER is a good metric for the ASR component of the full ST system is an open issue and lacks systematic studies. In this paper, we report our recent investigation on this issue, focusing on the interactions of ASR and MT in a ST system. We show that BLEU-oriented global optimization of ASR system parameters improves the translation quality by an absolute 1.5% BLEU score, while sacrificing WER over the conventional, WER-optimized ASR system. We also conducted an in-depth study on the impact of ASR errors on the final ST output. Our findings suggest that the speech recognizer component of the full ST system should be optimized by translation metrics instead of the traditional WER.
Xiaodong He 0001, Li Deng 0001, Alex Acero
ICASSP1
2011 A novel decision function and the associated decision-feedback learning for speech translation
abstract
In this paper we report our recent development of an end-to-end integrative design methodology for speech translation. Specifically, a novel decision function is proposed based on the Bayesian analysis, and the associated discriminative learning technique is presented based on the decision-feedback principle. The decision function in our end-to-end design methodology integrates acoustic scores, language model scores and translation scores to refine the translation hypotheses and to determine the best translation candidate. This Bayesian-guided decision function is then embedded into the training process that jointly learns the parameters in speech recognition and machine translation sub-systems in the overall speech translation system. The resulting decision-feedback learning takes a functional form similar to the minimum classification error training. Experimental results obtained on the IWSLT DIALOG 2010 database showed that the proposed system outperformed the baseline system in terms of BLEU score by 2.3 points.
Li Deng 0001, Xiaodong He 0001, Alex Acero
ICASSP3
2011 Robust Speech Translation by Domain Adaptation
abstract
Speech translation tasks usually are different from text-based machine translation tasks, and the training data for speech translation tasks are usually very limited. Therefore, domain adaptation is crucial to achieve robust performance across different conditions in speech translation. In this paper, we study the problem of adapting a general-domain, writing-textstyle machine translation system to a travel-domain, speech translation task. We study a variety of domain adaptation techniques, including data selection and incorporation of multiple translation models, in a unified decoding process. The experimental results demonstrate significant BLEU score improvement on the targeting scenario after domain adaptation. The results also demonstrate robust translation performance achieved across multiple conditions via joint data selection and model combination. We finally analyze and compare the robust techniques developed for speech recognition and speech translation, and point out further directions for robust translation via variability-adaptive and discriminatively-adaptive learning.
Xiaodong He 0001, Li Deng 0001
INTERSPEECH1
2010 Clickthrough-based translation models for web search: from word models to phrase models
abstract
Web search is challenging partly due to the fact that search queries and Web documents use different language styles and vocabularies. This paper provides a quantitative analysis of the language discrepancy issue, and explores the use of clickthrough data to bridge documents and queries. We assume that a query is parallel to the titles of documents clicked on for that query. Two translation models are trained and integrated into retrieval models: A word-based translation model that learns the translation probability between single words, and a phrase-based translation model that learns the translation probability between multi-term phrases. Experiments are carried out on a real world data set. The results show that the retrieval systems that use the translation models outperform significantly the systems that do not. The paper also demonstrates that standard statistical machine translation techniques such as word alignment, bilingual phrase extraction, and phrase-based decoding, can be adapted for building a better Web document retrieval system.
Jianfeng Gao 0001, Xiaodong He 0001, Jian-Yun Nie
CIKM2
2009 Incremental HMM Alignment for MT System Combination
Chi-Ho Li, Xiaodong He 0001
ACL/IJCNLP2
2009 Joint Optimization for Machine Translation System Combination
Xiaodong He 0001, Kristina Toutanova
EMNLP1
2009 Improved Monolingual Hypothesis Alignment for Machine Translation System Combination
abstract
This article presents a new hypothesis alignment method for combining outputs of multiple machine translation (MT) systems. An indirect hidden Markov model (IHMM) is proposed to address the synonym matching and word ordering issues in hypothesis alignment. Unlike traditional HMMs whose parameters are trained via maximum likelihood estimation (MLE), the parameters of the IHMM are estimated indirectly from a variety of sources including word semantic similarity, word surface similarity, and a distance-based distortion penalty. The IHMM-based method significantly outperforms the state-of-the-art, TER-based alignment model in our experiments on NIST benchmark datasets. Our combined SMT system using the proposed method achieved the best Chinese-to-English translation result in the constrained training track of the 2008 NIST Open MT Evaluation.
Xiaodong He 0001, Jianfeng Gao 0001, Patrick Nguyen
ACM Trans. Asian Lang. Inf. Process.1
2008 Indirect-HMM-based Hypothesis Alignment for Combining Outputs from Machine Translation Systems
Xiaodong He 0001, Jianfeng Gao 0001, Patrick Nguyen
EMNLP1
2008 Large-margin minimum classification error training: A theoretical risk minimization perspective
Dong Yu 0001, Li Deng 0001, Xiaodong He 0001, Alex Acero
Comput. Speech Lang.3
2007 Large-Margin Minimum Classification Error Training for Large-Scale Speech Recognition Tasks
abstract
Recently, we have developed a novel discriminative training method named large-margin minimum classification error (LM-MCE) training that incorporates the idea of discriminative margin into the conventional minimum classification error (MCE) training method. In our previous work, this novel approach was formulated specifically for the MCE training using the sigmoid loss function and its effectiveness was demonstrated on the TIDIGITS task alone. In this paper two additional contributions are made. First, we formulate LM-MCE as a Bayes risk minimization problem whose loss function not only includes empirical error rates but also a margin-bound risk. This new formulation allows us to extend the same technique to a wide variety of MCE based training. Second, we have successfully applied LM-MCE training approach to the Microsoft internal large vocabulary telephony speech recognition task (with 2000 hours of training data and 120K of vocabulary) and achieved significant recognition accuracy improvement across-the-board. To our best knowledge, this is the first time that the large-margin approach is demonstrated to be successful in large-scale speech recognition tasks.
Dong Yu 0001, Li Deng 0001, Xiaodong He 0001, Alex Acero
ICASSP (4)3
2007 Phone-discriminating minimum classification error (p-MCE) training for phonetic recognition
abstract
In this paper, we report a study on performance comparisons of discriminative training methods for phone recognition using the TIMIT database. We propose a new method of phonediscriminating minimum classification error (P-MCE), which performs MCE training at the sub-string or phone level instead of at the traditional string level. Aiming at minimizing the phone recognition error rate, P-MCE nevertheless takes advantage of the well-known, efficient training routine derived from the conventional string-based MCE, using specially constructed one-best lists selected from phone lattices. Extensive investigations and comparisons are conducted between the PMCE and other discriminative training methods including maximum mutual information (MMI), minimum phone or word error (MPE/MWE), and the other two MCE methods. The P-MCE outperforms most of experimented approaches on the standard TIMIT database in terms of the continuous phonetic recognition accuracy. P-MCE achieves comparable results with the MPE method which also aims at reducing phone-level recognition errors.
Xiaodong He 0001, Li Deng 0001
INTERSPEECH2
2007 Automatic validation of terminology translation consistenscy with statistical method
Masaki Itagaki, Takako Aikawa, Xiaodong He 0001
MTSummit3
2007 Prior knowledge guided maximum expected likelihood based model selection and adaptation for nonnative speech recognition
Xiaodong He 0001, Yunxin Zhao
Comput. Speech Lang.1
2007 A new look at discriminative training for hidden Markov models
Xiaodong He 0001, Li Deng 0001
Pattern Recognit. Lett.1
2006 Robust feature space adaptation for telephony speech recognition
abstract
Speaker adaptation is critical for modern speech recognition systems. Due to the computational and multi-channel model sharing considerations, the use of model adaptation techniques is limited in telephony speech recognition systems. On the other hand, feature space adaptation methods such as feature space maximum likelihood linear regression (fMLLR) are efficient approaches suitable for telephony systems. In this work, we first describe techniques for efficient implementation of online fMLLR adaptation. Then feature space maximum a posteriori linear regression (fMAPLR) is proposed to incorporate prior knowledge for the feature transform estimation and improve the robustness of the conventional fMLLR approach. Experiments on telephony data indicate that fMAPLR is significantly more robust than fMLLR, and outperforms fMLLR especially when the adaptation data is very limited. Index Terms: Speaker adaptation, telephony, speech recognition. 1.
Jon Hamaker, Xiaodong He 0001
INTERSPEECH3
2006 Use of incrementally regulated discriminative margins in MCE training for speech recognition
abstract
In this paper, we report our recent development of a novel discriminative learning technique which embeds the concept of discriminative margin into the well established minimum classification error (MCE) method. The idea is to impose an incrementally adjusted “margin ” in the loss function of MCE algorithm so that not only error rates are minimized but also discrimination “robustness ” between training and test sets is maintained. Experimental evaluation shows that the use of the margin improves a state-of-the-art MCE method by reducing 17 % digit errors and 19 % string errors in the TIDigits recognition task. The string error rate of 0.55 % and digit error rate of 0.19 % we have obtained are the best-ever results reported on this task in the literature. Index Terms: discriminative training, margin, minimum error 1.
Dong Yu 0001, Li Deng 0001, Xiaodong He 0001, Alex Acero
INTERSPEECH3
2006 A Novel Learning Method for Hidden Markov Models in Speech and Audio Processing
abstract
In recent years, various discriminative learning techniques for HMMs have consistently yielded significant benefits in speech recognition. In this paper, we present a novel optimization technique using the minimum classification error (MCE) criterion to optimize the HMM parameters. Unlike maximum mutual information training where an extended Baum-Welch (EBW) algorithm exists to optimize its objective function, for MCE training the original EBW algorithm cannot be directly applied. In this work, we extend the original EBW algorithm and derive a novel method for MCE-based model parameter estimation. Compared with conventional gradient descent methods for MCE learning, the proposed method gives a solid theoretical basis, stable convergence, and it is well suited for the large-scale batch-mode training process essential in large-scale speech recognition and other pattern recognition applications. Evaluation experiments, including model training and speech recognition, are reported on both a small vocabulary task (TI-digits) and a large vocabulary task (WSJ), where the effectiveness of the proposed method is demonstrated. We expect new future applications and success of this novel learning method in general pattern recognition and multimedia processing, in addition to speech and audio processing applications we present in this paper
Xiaodong He 0001, Li Deng 0001, Wu Chou
MMSP1
2004 Prior knowledge guided MEL based model selection and adaptation for nonnative speech recognition
abstract
An improved method of model complexity selection for nonnative speech recognition is proposed by using maximum a posteriori estimation of bias distributions. An algorithm is described for estimating the hyper-parameters of the prior distributions, and an automatic accent detection algorithm is also proposed for integration with dynamic model selection and adaptation. Experiments were performed on the WSJ1 task with American English speech, British accent speech, and Mandarin Chinese accent speech. Results show that the use of prior knowledge of accents enabled reliable estimation of bias distributions in the case of a very small amount of adaptation speech, or without adaptation speech. Recognition results show that the new approach is superior to the previous MEL (maximum expected likelihood) method, especially when the adaptation data are extremely limited.
Xiaodong He 0001, Yunxin Zhao
ICASSP (1)1
2003 Minimum classification error linear regression for acoustic model adaptation of continuous density HMMs
abstract
In this paper, a concatenated "super" string model based minimum classification error (MCE) model adaptation approach is described. We show that the error rate minimization in the proposed approach can be formulated into maximizing a special ratio of two positive functions. The proposed string model is used to derive the growth transform based error rate minimization for MCE linear regression (MCELR). It provides an effective solution to apply MCE approach to acoustic model adaptation with sparse data. The proposed MCELR approach is studied and compared with the maximum likelihood linear regression (MLLR) based model adaptation. Experiments on large vocabulary speech recognition tasks are performed. Experimental results indicate that the proposed MCELR model adaptation can lead to significant speech recognition performance improvement and its performance advantage over the MLLR based approach is observed even when the amount of adaptation data is sparse.
Xiaodong He 0001, Wu Chou
ICASSP (1)1
2003 minimum classification error linear regression for acoustic model adaptation of continuous density HMMS
abstract
In this paper, a concatenated "super" string model based minimum classification error (MCE) model adaptation approach is described. We show that the error rate minimization in the proposed approach can be formulated into maximizing a special ratio of two positive functions. The proposed string model is used to derive the growth transform based error rate minimization for MCE linear regression (MCELR). It provides an effective solution to apply MCE approach to acoustic model adaptation with sparse data. The proposed MCELR approach is studied and compared with the maximum likelihood linear regression (MLLR) based model adaptation. Experiments on large vocabulary speech recognition tasks are performed. Experimental results indicate that the proposed MCELR model adaptation can lead to significant speech recognition performance improvement and its performance advantage over the MLLR based approach is observed even when the amount of adaptation data is sparse.
Xiaodong He 0001, Wu Chou
ICME1
2003 Maximum a posteriori linear regression (MAPLR) variance adaptation for continuous density HMMS
abstract
In this paper, the theoretical framework of maximum a posteriori linear regression (MAPLR) based variance adaptation for continuous density HMMs is described. In our approach, a class of informative prior distribution for MAPLR based variance adaptation is identified, from which the close form solution of MAPLR based variance adaptation is obtained under its EM formulation. Effects of the proposed prior distribution in MAPLR based variance adaptation are characterized and compared with conventional maximum likelihood linear regression (MLLR) based variance adaptation. These findings provide a consistent Bayesian theoretical framework to incorporate prior knowledge in linear regression based variance adaptation. Experiments on large vocabulary speech recognition tasks were performed. The experimental results indicate that significant performance gain over the MLLR based variance adaptation can be obtained based on the proposed approach.
Wu Chou, Xiaodong He 0001
INTERSPEECH2
2003 Minimum classification error (MCE) model adaptation of continuous density HMMS
abstract
In this paper, a framework of minimum classification error (MCE) model adaptation for continuous density HMMs is proposed based on the approach of "super " string model. We show that the error rate minimization in the proposed approach can be formulated into maximizing a special ratio of two positive functions, and from that a general growth transform algorithm is derived for MCE based model adaptation. This algorithm departs from the generalized probability descent (GPD) algorithm, and it is well suited for model adaptation with a small amount of training data. The proposed approach is applied to linear regression based variance adaptation, and the close form solution for variance adaptation using MCE linear regression (MCELR) is derived. The MCELR approach is evaluated on large vocabulary speech recognition tasks. The relative performance gain is more than doubled on the standard (WSJ Spoke 3) database, comparing to maximum likelihood linear regression (MLLR) based variance adaptation for the same amount of adaptation data. 1.
Xiaodong He 0001, Wu Chou
INTERSPEECH1
2003 Fast model selection based speaker adaptation for nonnative speech
abstract
The problem of adapting acoustic models of native English speech to nonnative speakers is addressed from a perspective of adaptive model complexity selection. The goal is to select model complexity dynamically for each nonnative talker so as to optimize the balance between model robustness to pronunciation variations and model detailedness for discrimination of speech sounds. A maximum expected likelihood (MEL) based technique is proposed to enable reliable complexity selection when adaptation data are sparse, where expectation of log-likelihood (EL) of adaptation data is computed based on distributions of mismatch biases between model and data, and model complexity is selected to maximize EL. The MEL based complexity selection is further combined with MLLR (maximum likelihood linear regression) to enable adaptation of both complexity and parameters of acoustic models. Experiments were performed on WSJ1 data of speakers with a wide range of foreign accents. Results show that the MEL based complexity selection is feasible when using as little as one adaptation utterance, and it is able to select dynamically the proper model complexity as the adaptation data increases. Compared with the standard MLLR, the MEL+MLLR method leads to consistent and significant improvement to recognition accuracy on nonnative speakers, without performance degradation on native speakers.
Xiaodong He 0001, Yunxin Zhao
IEEE Trans. Speech Audio Process.1
2002 Fast model adaptation and complexity selection for nonnative English speakers
abstract
In this paper, the problem of fast model adaptation and complexity selection for nonnative speaker is investigated. The key challenge lies in reliable complexity selection when only a small amount of adaptation data is available. A novel technique of combining a maximum likelihood (ML) based state-tying with a pseudo likelihood (PL) based state-tying is proposed to enable model complexity selection from using as little as three adaptation speech sentences. In MUPL, ML model complexity selection is performed on nodes with sufficient adaptation data, and PL based state tying is performed on nodes with insufficient adaptation data. Experiments were performed on WSJ data of six nonnative speakers. The combined model adaptation and complexity selection method led to consistent and significant improvement on recognition accuracy over MLLR, with an average error reduction of 13% when a varying number of adaptation speech sentences were taken from each speaker.
Xiaodong He 0001, Yunxin Zhao
ICASSP1
2002 Maximum expected likelihood based model selection and adaptation for nonnative English speakers
abstract
In this paper, the problem of fast model adaptation for nonnative speakers is addressed from a perspective of model complexity selection. The key challenge lies in reliable complexity selection when only a small amount of adaptation data is available. A novel maximum expected likelihood (MEL) based technique is proposed to enable model complexity selection from using as little as one adaptation sentence. In MEL, the expectation of loglikelihood is computed based on the mismatch bias between model and data which is measured by a small amount of adaptation data, and model complexity is selected to maximize EL. Experiments were performed on WSJ data of speakers with a wide range of foreign accents. The proposed method led to consistent and significant improvement on recognition accuracy over MLLR for nonnative speakers, without performance degradation on native speakers. The proposed method was able to dynamically select optimal model complexity as the available adaptation data increased. 1.
Xiaodong He 0001, Yunxin Zhao
INTERSPEECH1
2001 Model complexity optimization for nonnative English speakers
abstract
In this paper, a study is made on selecting existing acoustic models that are trained from native English speech for improving recognition of nonnative English talkers ’ speech. The problem is addressed from the perspective that foreign accents prevent detailed triphone models that are commonly used in highperformance speech recognition systems to match well with these talkers ’ speech, and therefore an appropriate level of context-dependent acoustic modeling is needed for foreign accent speakers. In this work, model complexity selection is accomplished by empirically choosing a set of model tying thresholds and by using the principle of MDL. An experiment was performed on the Wall Street Journal task on three nonnative English talkers with Chinese accent (276 sentences). Compared to the result obtained from using the models optimized to native English speakers, the best model tying threshold and MDL yielded similar and significant reduction to recognition word errors by 23%. 1.
Xiaodong He 0001, Yunxin Zhao
INTERSPEECH1
2000 A combined adaptive and decision tree based speech separation technique for telemedicine applications
abstract
We present a novel technique for separation of doctor and patient’s speech in conversations over a telemedicine network. The mixed speech signals acquired at doctor’s site is first broken into single talkers ’ speech segments and background by using thresholds of energy and duration. The speech segments are then identified as spoken by doctor or patient in two steps. In the first step, Gaussian mixture models (GMM) of doctor and patient are used, where the doctor’s model is obtained from his/her training speech, and the patient’s model is initialized by a general speaker model and then adapted by the patient’s speech. In the second step, a decision tree that uses contextual and confidence features is applied to refine the identification results. Preliminary experiments were performed on three data sets collected in telemedicine. Without adaptation and decision tree, error rates at the segment-level and frame-level were 25.44 % and 16.53%, respectively. With adaptation, segment and frame error rates were reduced to 13.11 % and 7.85%, and with decision tree, the error rates were further reduced to 10.48 % and 6.73%, respectively. 1.
Yunxin Zhao, Xiaodong He 0001, Laura Schopp
INTERSPEECH3
1999 Research on speech units modeling in continuous speech recognition
abstract
It is often expedient to consider using more than one single HMM to characterize a speech unit. In this paper, we suggest a new speech units modeling method based on analysis of parameters of HMMs obtained by preliminary training. By analyzing the emission probability function of a state of a HMM obtained by segmental k-means training, we can obtain the distribution of the source data and determine the splitting of that model. The experimental results, based on totally 264,500 phoneme occurring in the 9180 sentences from 60 speakers, showed that approximate 10% improvement of the recognition rate of the basic phoneme was achieved.
Xiaodong He 0001, Jian-Lai Zhou, Tiecheng Yu
EUROSPEECH1
1999 Study on tone classification of Chinese continuous speech in speech recognition system
Xiaodong He 0001, Fuyuan Mo, Tiecheng Yu
EUROSPEECH2
1999 A new hybrid structure of speech recognizer based on HMM and neural network
abstract
In this paper, we introduced a new framework of speech recognizer based on HMM and neural net. Unlike the traditional hybrid system, the neural net was used as a post processor, which classify the speech data segmented by HMM recognizer. The purpose of this method is to improve the top-choice accuracy of HMM based speech recognition system in our lab. Major issues such as how to use the segmentation information of HMM in neural net, the structure of the neural net, the choice of the error metric for training neural net, and the determination of the training procedure are investigated within a set of experiments. In these experiments, we attempt to recognize 68 phoneme like units in continuous speech. Our results indicate that this is a potential method. About 20% can be obtained to improve the recognition accuracy for multi-speaker system in syllable level, and 10% for speaker independent system.
Jian-Lai Zhou, Xiaodong He 0001, Tiecheng Yu, Fuyuan Mo
EUROSPEECH2