VLDB 2026 Research / reviewers in the wild / expert
Ming Zhou 0001
dblp:16/1161-1
· DBLP profile ↗
261ranked-venue papers
2as first author
27since 2021 · last 2025
0000-0002-2551-2964ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 238 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 8 since 2021Databases, data management, data science and information retrieval · 28 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Task-level Distributionally Robust Optimization for Large Language Model-based Dense RetrievalabstractLarge Language Model-based Dense Retrieval (LLM-DR) optimizes over numerous heterogeneous fine-tuning collections from different domains. However, the discussion about its training data distribution is still minimal. Previous studies rely on empirically assigned dataset choices or sampling ratios, which inevitably lead to sub-optimal retrieval performances. In this paper, we propose a new task-level Distributionally Robust Optimization (tDRO) algorithm for LLM-DR fine-tuning, targeted at improving the universal domain generalization ability by end-to-end reweighting the data distribution of each task. The tDRO parameterizes the domain weights and updates them with scaled domain gradients. The optimized weights are then transferred to the LLM-DR fine-tuning to train more robust retrievers. Experiments show optimal improvements in large-scale retrieval benchmarks and reduce up to 30% dataset usage after applying our optimization algorithm with a series of different-sized LLM-DR models. Guangyuan Ma, Yongliang Ma, Xing Wu 0002, Zhenpeng Su, Ming Zhou 0001, Songlin Hu 0001 |
AAAI | 5 |
| 2025 | Pre-Training a Graph Recurrent Network for Text UnderstandingabstractTransformer-based pre-trained models have gained much advance in recent years, Transformer architecture also becomes one of the most important backbones in natural language processing. Recent works show that the attention mechanism inside Transformer may not be necessary, and Transformer alternatives such as convolutional neural networks, multi-layer perceptron, and state space model have also been investigated. Transformer-based models have two main limitations: First, they have quadratic time complexity due to the full attention mechanism, which leads to high computational costs. Second, they rely on representation of a special token such as [CLS] to encode entire text, which limits its sentence-level expressiveness. In this paper, we consider a graph recurrent network with linear time complexity for language model pre-training, which builds a graph structure for each sequence with local token-level communications, together with a sentence-level representation detached from other normal tokens. On both English and Chinese text understanding tasks, our model can achieve comparable performance to existing pre-trained models while also achieving higher inference efficiency. Furthermore, we discovered that the representations generated by our model are more diverse and uniform compared to that of Transformer, which alleviates the problems in existing pre-trained models such as representation degradation. Yile Wang 0001, Linyi Yang, Zhiyang Teng, Ming Zhou 0001, Yue Zhang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | HORIZON: High-Resolution Semantically Controlled Panorama SynthesisabstractPanorama synthesis endeavors to craft captivating 360-degree visual landscapes, immersing users in the heart of virtual worlds. Nevertheless, contemporary panoramic synthesis techniques grapple with the challenge of semantically guiding the content generation process. Although recent breakthroughs in visual synthesis have unlocked the potential for semantic control in 2D flat images, a direct application of these methods to panorama synthesis yields distorted content. In this study, we unveil an innovative framework for generating high-resolution panoramas, adeptly addressing the issues of spherical distortion and edge discontinuity through sophisticated spherical modeling. Our pioneering approach empowers users with semantic control, harnessing both image and text inputs, while concurrently streamlining the generation of high-resolution panoramas using parallel decoding. We rigorously evaluate our methodology on a diverse array of indoor and outdoor datasets, establishing its superiority over recent related work, in terms of both quantitative and qualitative performance metrics. Our research elevates the controllability, efficiency, and fidelity of panorama synthesis to new levels. Kun Yan 0004, Lei Ji 0001, Chenfei Wu, Ming Zhou 0001, Nan Duan 0001, Shuai Ma 0001 |
AAAI | 5 |
| 2024 | Quantified emotion analysis based on design principles of color feature recognition in pictures
Chin-Ling Chen, Qing-Yang Huang, Ming Zhou 0001, Der-Chen Huang, Ling-Chun Liu, Yong-Yuan Deng |
Multim. Tools Appl. | 3 |
| 2023 | LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language UnderstandingabstractNLP research on logical reasoning regains momentum with the recent releases of a handful of datasets, notably LogiQA and Reclor. Logical reasoning is exploited in many probing tasks over large Pre-trained Language Models (PLMs) and downstream tasks like question-answering and dialogue systems. In this paper, we release LogiQA 2.0. The dataset is an amendment and re-annotation of LogiQA in 2020, a large-scale logical reasoning reading comprehension dataset adapted from the Chinese Civil Service Examination. We increase the data size, refine the texts with manual translation by professionals, and improve the quality by removing items with distinctive cultural features like Chinese idioms. Furthermore, we conduct a fine-grained annotation on the dataset and turn it into a two-way natural language inference (NLI) task, resulting in 35k premise-hypothesis pairs with gold labels, making it the first large-scale NLI dataset for complex logical reasoning. Compared to Question Answering, Natural Language Inference excels in generalizability and helps downstream tasks better. We establish a baseline for logical reasoning in NLI and incite further research. Hanmeng Liu, Jian Liu 0030, Leyang Cui, Zhiyang Teng, Nan Duan 0001, Ming Zhou 0001, Yue Zhang 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Robust principal component analysis via weighted nuclear norm with modified second-order total variation regularization
Yi Dou, Xinling Liu, Ming Zhou 0001 |
Vis. Comput. | 3 |
| 2022 | UniXcoder: Unified Cross-Modal Pre-training for Code RepresentationabstractPre-trained models for programming languages have recently demonstrated great success on code intelligence.To support both code-related understanding and generation tasks, recent works attempt to pre-train unified encoder-decoder models.However, such encoder-decoder framework is sub-optimal for auto-regressive tasks, especially code completion that requires a decoder-only manner for efficient inference.In this paper, we present UniXcoder, a unified cross-modal pre-trained model for programming language.The model utilizes mask attention matrices with prefix adapters to control the behavior of the model and leverages cross-modal contents like AST and code comment to enhance code representation.To encode AST that is represented as a tree in parallel, we propose a one-to-one mapping method to transform AST in a sequence structure that retains all structural information from the tree.Furthermore, we propose to utilize multi-modal contents to learn representation of code fragment with contrastive learning, and then align representations among programming languages using a cross-modal generation task.We evaluate UniXcoder on five code-related tasks over nine datasets.To further evaluate the performance of code fragment representation, we also construct a dataset for a new task, called zero-shot code-to-code search.Results show that our model achieves state-of-the-art performance on most tasks and analysis reveals that comment and AST can both enhance UniXcoder. Daya Guo, Nan Duan 0001, Yanlin Wang 0001, Ming Zhou 0001, Jian Yin 0001 |
ACL (1) | 5 |
| 2022 | Trace Controlled Text to Image Generation
Kun Yan 0004, Lei Ji 0001, Chenfei Wu, Jianmin Bao, Ming Zhou 0001, Nan Duan 0001, Shuai Ma 0001 |
ECCV (36) | 5 |
| 2022 | Reasoning over Hybrid Chain for Table-and-Text Open Domain Question AnsweringabstractTabular and textual question answering requires systems to perform reasoning over heterogeneous information, considering table structure, and the connections among table and text. In this paper, we propose a ChAin-centric Reasoning and Pre-training framework (CARP). CARP utilizes hybrid chain to model the explicit intermediate reasoning process across table and text for question answering. We also propose a novel chain-centric pre-training method, to enhance the pre-trained model in identifying the cross-modality reasoning process and alleviating the data sparsity problem. This method constructs the large-scale reasoning corpus by synthesizing pseudo heterogeneous reasoning paths from Wikipedia and generating corresponding questions. We evaluate our system on OTT-QA, a large-scale table-and-text open-domain question answering benchmark, and our system achieves the state-of-the-art performance. Further analyses illustrate that the explicit hybrid chain offers substantial performance improvement and interpretablity of the intermediate reasoning process, and the chain-centric pre-training boosts the performance on the chain extraction. Wanjun Zhong, Junjie Huang 0008, Qian Liu 0033, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001, Nan Duan 0001 |
IJCAI | 4 |
| 2022 | BlonDe: An Automatic Evaluation Metric for Document-level Machine TranslationabstractYuchen Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tianyu Liu 0004, Shuming Ma, Dongdong Zhang 0001, Jian Yang 0030, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou 0001 |
NAACL-HLT | 10 |
| 2022 | ProQA: Structural Prompt-based Pre-training for Unified Question AnsweringabstractWanjun Zhong, Yifan Gao, Ning Ding, Yujia Qin, Zhiyuan Liu, Ming Zhou, Jiahai Wang, Jian Yin, Nan Duan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wanjun Zhong, Yifan Gao 0001, Ning Ding 0002, Yujia Qin, Zhiyuan Liu 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001, Nan Duan 0001 |
NAACL-HLT | 6 |
| 2022 | From LSAT: The Progress and Challenges of Complex ReasoningabstractComplex reasoning aims to draw a correct inference based on complex rules. As a hallmark of human intelligence, it involves a degree of explicit reading comprehension, interpretation of logical knowledge and complex rule application. In this paper, we take a step forward in complex reasoning by systematically studying the three challenging and domain-general tasks of the Law School Admission Test (LSAT), including analytical reasoning, logical reasoning and reading comprehension. We propose a hybrid reasoning system to integrate these three tasks and achieve impressive overall performance on the LSAT tests. The experimental results demonstrate that our system endows itself a certain complex reasoning ability, especially the fundamental reading comprehension and challenging logical reasoning capacities. Further analysis also shows the effectiveness of combining the pre-trained models with the task-specific reasoning module, and integrating symbolic knowledge into discrete interpretable reasoning steps in complex reasoning. We further shed a light on the potential future directions, like unsupervised symbolic knowledge extraction, model interpretability, few-shot learning and comprehensive benchmark for complex reasoning. Siyuan Wang 0025, Zhongkun Liu, Wanjun Zhong, Ming Zhou 0001, Zhongyu Wei, Zhumin Chen, Nan Duan 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Conditional Sentence Generation and Cross-Modal Reranking for Sign Language TranslationabstractSign Language Translation (SLT) aims to generate spoken language translations from sign language videos. Currently, the available sign language datasets are relatively too small to learn the linguistic properties of spoken language. In this paper, towards effective SLT, we propose a novel framework which takes the advantage of the spoken language grammar learnt from a large corpus of text sentences. Our framework consists of three key modules: word existence verification, conditional sentence generation and cross-modal re-ranking. We first check the existence of words in the vocabulary by a series of binary classification in parallel. After that, the appearing words are assembled and guided by a pretrained spoken language generator to produce multiple candidate sentences in spoken language manner. Last but not least, we select the sentence most semantically similar to the input sign video as the translation result with a crossmodal re-ranking model. We evaluate our framework on two large scale continuous SLT benchmarks,i.e., CSL and RWTHPHOENIX-Weather 2014 T. Experimental results demonstrate that the proposed framework achieves promising performance on both datasets. Jian Zhao 0018, Weizhen Qi, Wengang Zhou 0001, Nan Duan 0001, Ming Zhou 0001, Houqiang Li |
IEEE Trans. Multim. | 5 |
| 2021 | Compare to The Knowledge: Graph Neural Fake News Detection with External KnowledgeabstractLinmei Hu, Tianchi Yang, Luhao Zhang, Wanjun Zhong, Duyu Tang, Chuan Shi, Nan Duan, Ming Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Linmei Hu, Tianchi Yang, Luhao Zhang, Wanjun Zhong, Duyu Tang, Chuan Shi 0001, Nan Duan 0001, Ming Zhou 0001 |
ACL/IJCNLP (1) | 8 |
| 2021 | CoSQA: 20, 000+ Web Queries for Code Search and Question AnsweringabstractJunjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, Nan Duan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Junjie Huang 0008, Duyu Tang, Linjun Shou, Ming Gong 0001, Ke Xu 0001, Daxin Jiang, Ming Zhou 0001, Nan Duan 0001 |
ACL/IJCNLP (1) | 7 |
| 2021 | Learning to Ask Conversational Questions by Optimizing Levenshtein DistanceabstractZhongkun Liu, Pengjie Ren, Zhumin Chen, Zhaochun Ren, Maarten de Rijke, Ming Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhongkun Liu, Pengjie Ren, Zhumin Chen, Zhaochun Ren, Maarten de Rijke, Ming Zhou 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | SemFace: Pre-training Encoder and Decoder with a Semantic Interface for Neural Machine TranslationabstractShuo Ren, Long Zhou, Shujie Liu, Furu Wei, Ming Zhou, Shuai Ma. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shuo Ren 0002, Shujie Liu 0001, Furu Wei, Ming Zhou 0001, Shuai Ma 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | Control Image Captioning Spatially and TemporallyabstractKun Yan, Lei Ji, Huaishao Luo, Ming Zhou, Nan Duan, Shuai Ma. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kun Yan 0004, Lei Ji 0001, Huaishao Luo, Ming Zhou 0001, Nan Duan 0001, Shuai Ma 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Jointly Learning to Repair Code and Generate Commit MessageabstractWe propose a novel task of jointly repairing program codes and generating commit messages.Code repair and commit message generation are two essential and related tasks for software development.However, existing work usually performs the two tasks independently.We construct a multilingual triple dataset including buggy code, fixed code, and commit messages for this novel task.We provide the cascaded models as baseline, which are enhanced with different training approaches, including the teacher-student method, the multi-task method, and the backtranslation method.To deal with the error propagation problem of the cascaded method, the joint model is proposed that can both repair the code and generate the commit message in a unified framework.Experimental results show that the enhanced cascaded model with teacher-student method and multitask-learning method achieves the best score on different metrics of automated code repair, and the joint model behaves better than the cascaded model on commit message generation. Jiaqi Bai 0001, Ambrosio Blanco, Shujie Liu 0001, Furu Wei, Ming Zhou 0001, Zhoujun Li 0001 |
EMNLP (1) | 6 |
| 2021 | Continuous Speech Separation with ConformerabstractContinuous speech separation was recently proposed to deal with the overlapped speech in natural conversations. While it was shown to significantly improve the speech recognition performance for multichannel conversation transcription, its effectiveness has yet to be proven for a single-channel recording scenario. This paper examines the use of Conformer architecture in lieu of recurrent neural networks for the separation model. Conformer allows the separation model to efficiently capture both local and global context information, which is helpful for speech separation. Experimental results using the LibriCSS dataset show that the Conformer separation model achieves the state of the art results for both single-channel and multi-channel settings. Results for real meeting recordings are also presented, showing significant performance gains in both word error rate (WER) and speaker-attributed WER. Sanyuan Chen, Yu Wu 0012, Zhuo Chen 0006, Jian Wu 0027, Jinyu Li 0001, Takuya Yoshioka, Chengyi Wang 0002, Shujie Liu 0001, Ming Zhou 0001 |
ICASSP | 9 |
| 2021 | GraphCodeBERT: Pre-training Code Representations with Data Flow
Daya Guo, Shuo Ren 0002, Zhangyin Feng, Duyu Tang, Shujie Liu 0001, Nan Duan 0001, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin 0001, Daxin Jiang, Ming Zhou 0001 |
ICLR | 18 |
| 2021 | BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale PretrainingabstractIn this paper, we propose BANG, a new pretraining model to Bridge the gap between Autoregressive (AR) and Non-autoregressive (NAR) Generation. AR and NAR generation can be uniformly regarded as to what extent previous tokens can be attended, and BANG bridges AR and NAR generation through designing a novel model structure for large-scale pre-training. A pretrained BANG model can simultaneously support AR, NAR, and semi-NAR generation to meet different requirements. Experiments on question generation (SQuAD 1.1), summarization (XSum), and dialogue generation (PersonaChat) show that BANG improves NAR and semi-NAR performance significantly as well as attaining comparable performance with strong AR pretrained models. Compared with the semi-NAR strong baselines, BANG achieves absolute improvements of 14.01 and 5.24 in the overall scores of SQuAD 1.1 and XSum, respectively. In addition, BANG achieves absolute improvements of 10.73, 6.39, and 5.90 in the overall scores of SQuAD, XSUM, and PersonaChat compared with the NAR strong baselines, respectively. Our code will be made publicly available. Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Weizhu Chen, Dayiheng Liu, Kewen Tang, Houqiang Li, Jiusheng Chen, Ruofei Zhang, Ming Zhou 0001, Nan Duan 0001 |
ICML | 11 |
| 2021 | InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-TrainingabstractZewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, Ming Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zewen Chi, Li Dong 0004, Furu Wei, Nan Yang 0002, Saksham Singhal, Wenhui Wang 0003, Xianling Mao, Heyan Huang, Ming Zhou 0001 |
NAACL-HLT | 10 |
| 2021 | Smart-Start Decoding for Neural Machine TranslationabstractJian Yang, Shuming Ma, Dongdong Zhang, Juncheng Wan, Zhoujun Li, Ming Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Juncheng Wan, Zhoujun Li 0001, Ming Zhou 0001 |
NAACL-HLT | 6 |
| 2021 | XGPT: Cross-modal Generative Pre-Training for Image Captioning
Qiaolin Xia, Haoyang Huang, Nan Duan 0001, Dongdong Zhang 0001, Lei Ji 0001, Zhifang Sui, Edward Dong Bo Cui, Taroon Bharti, Ming Zhou 0001 |
NLPCC (1) | 9 |
| 2021 | Tree-Capsule: Tree-Structured Capsule Network for Improving Relation Extraction
Tianchi Yang, Linmei Hu, Luhao Zhang, Chuan Shi 0001, Cheng Yang 0002, Nan Duan 0001, Ming Zhou 0001 |
PAKDD (3) | 7 |
| 2021 | A blockchain-based intelligent anti-switch package in tracing logistics system
Chin-Ling Chen, Yong-Yuan Deng, Wei Weng 0002, Ming Zhou 0001, Hongyu Sun 0005 |
J. Supercomput. | 4 |
| 2020 | Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech TranslationabstractEnd-to-end speech translation, a hot topic in recent years, aims to translate a segment of audio into a specific language with an end-to-end model. Conventional approaches employ multi-task learning and pre-training methods for this task, but they suffer from the huge gap between pre-training and fine-tuning. To address these issues, we propose a Tandem Connectionist Encoding Network (TCEN) which bridges the gap by reusing all subnets in fine-tuning, keeping the roles of subnets consistent, and pre-training the attention module. Furthermore, we propose two simple but effective methods to guarantee the speech encoder outputs and the MT encoder inputs are consistent in terms of semantic representation and sequence length. Experimental results show that our model leads to significant improvements in En-De and En-Fr translation irrespective of the backbones. Chengyi Wang 0002, Yu Wu 0012, Shujie Liu 0001, Zhenglu Yang, Ming Zhou 0001 |
AAAI | 5 |
| 2020 | Alternating Language Modeling for Cross-Lingual Pre-TrainingabstractLanguage model pre-training has achieved success in many natural language processing tasks. Existing methods for cross-lingual pre-training adopt Translation Language Model to predict masked words with the concatenation of the source sentence and its target equivalent. In this work, we introduce a novel cross-lingual pre-training method, called Alternating Language Modeling (ALM). It code-switches sentences of different languages rather than simple concatenation, hoping to capture the rich cross-lingual context of words and phrases. More specifically, we randomly substitute source phrases with target translations to create code-switched sentences. Then, we use these code-switched data to train ALM model to learn to predict words of different languages. We evaluate our pre-training ALM on the downstream tasks of machine translation and cross-lingual classification. Experiments show that ALM can outperform the previous pre-training methods on three benchmarks.1 Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Shuangzhi Wu, Zhoujun Li 0001, Ming Zhou 0001 |
AAAI | 6 |
| 2020 | MuTual: A Dataset for Multi-Turn Dialogue ReasoningabstractNon-task oriented dialogue systems have achieved great success in recent years due to largely accessible conversation data and the development of deep learning techniques.Given a context, current systems are able to yield a relevant and fluent response, but sometimes make logical mistakes because of weak reasoning capabilities.To facilitate the conversation reasoning research, we introduce Mu-Tual, a novel dataset for Multi-Turn dialogue Reasoning, consisting of 8,860 manually annotated dialogues based on Chinese student English listening comprehension exams.Compared to previous benchmarks for non-task oriented dialogue systems, MuTual is much more challenging since it requires a model that can handle various reasoning problems.Empirical results show that state-of-the-art methods only reach 71%, which is far behind the human performance of 94%, indicating that there is ample room for improving reasoning ability.MuTual is available at https://github. com/Nealcly/MuTual. * Contribution during internship at MSRA.M: Ma'am Leyang Cui, Yu Wu 0012, Shujie Liu 0001, Yue Zhang 0004, Ming Zhou 0001 |
ACL | 5 |
| 2020 | Evidence-Aware Inferential Text Generation with Vector Quantised Variational AutoEncoderabstractGenerating inferential texts about an event in different perspectives requires reasoning over different contexts that the event occurs.Existing works usually ignore the context that is not explicitly provided, resulting in a context-independent semantic representation that struggles to support the generation.To address this, we propose an approach that automatically finds evidence for an event from a large text corpus, and leverages the evidence to guide the generation of inferential texts.Our approach works in an encoderdecoder manner and is equipped with a Vector Quantised-Variational Autoencoder, where the encoder outputs representations from a distribution over discrete variables.Such discrete representations enable automatically selecting relevant evidence, which not only facilitates evidence-aware generation, but also provides a natural way to uncover rationales behind the generation.Our approach provides state-ofthe-art performance on both Event2Mind and ATOMIC datasets.More importantly, we find that with discrete representations, our model selectively uses evidence to generate different inferential texts. Daya Guo, Duyu Tang, Nan Duan 0001, Jian Yin 0001, Daxin Jiang, Ming Zhou 0001 |
ACL | 6 |
| 2020 | Graph Neural News Recommendation with Unsupervised Preference DisentanglementabstractWith the explosion of news information, personalized news recommendation has become very important for users to quickly find their interested contents. Most existing methods usually learn the representations of users and news from news contents for recommendation. However, they seldom consider high-order connectivity underlying the user-news interactions. Moreover, existing methods failed to disentangle a user’s latent preference factors which cause her clicks on different news. In this paper, we model the user-news interactions as a bipartite graph and propose a novel Graph Neural News Recommendation model with Unsupervised Preference Disentanglement, named GNUD. Our model can encode high-order relationships into user and news representations by information propagation along the graph. Furthermore, the learned representations are disentangled with latent preference factors by a neighborhood routing algorithm, which can enhance expressiveness and interpretability. A preference regularizer is also designed to force each disentangled subspace to independently reflect an isolated preference, improving the quality of the disentangled representations. Experimental results on real-world news datasets demonstrate that our proposed model can effectively improve the performance of news recommendation and outperform state-of-the-art news recommendation methods. Linmei Hu, Siyong Xu, Cheng Yang 0002, Chuan Shi 0001, Nan Duan 0001, Xing Xie 0001, Ming Zhou 0001 |
ACL | 8 |
| 2020 | A Simple and Effective Unified Encoder for Document-Level Machine TranslationabstractMost of the existing models for documentlevel machine translation adopt dual-encoder structures.The representation of the source sentences and the document-level contexts 1 are modeled with two separate encoders.Although these models can make use of the document-level contexts, they do not fully model the interaction between the contexts and the source sentences, and can not directly adapt to the recent pre-training models (e.g., BERT) which encodes multiple sentences with a single encoder.In this work, we propose a simple and effective unified encoder that can outperform the baseline models of dualencoder models in terms of BLEU and ME-TEOR scores.Moreover, the pre-training models can further boost the performance of our proposed model. Shuming Ma, Dongdong Zhang 0001, Ming Zhou 0001 |
ACL | 3 |
| 2020 | A Graph-based Coarse-to-fine Method for Unsupervised Bilingual Lexicon InductionabstractUnsupervised bilingual lexicon induction is the task of inducing word translations from monolingual corpora of two languages.Recent methods are mostly based on unsupervised cross-lingual word embeddings, the key to which is to find initial solutions of word translations, followed by the learning and refinement of mappings between the embedding spaces of two languages.However, previous methods find initial solutions just based on word-level information, which may be (1) limited and inaccurate, and (2) prone to contain some noise introduced by the insufficiently pre-trained embeddings of some words.To deal with those issues, in this paper, we propose a novel graph-based paradigm to induce bilingual lexicons in a coarse-to-fine way.We first build a graph for each language with its vertices representing different words.Then we extract word cliques from the graphs and map the cliques of two languages.Based on that, we induce the initial word translation solution with the central words of the aligned cliques.This coarse-to-fine approach not only leverages clique-level information, which is richer and more accurate, but also effectively reduces the bad effect of the noise in the pre-trained embeddings.Finally, we take the initial solution as the seed to learn cross-lingual embeddings, from which we induce bilingual lexicons.Experiments show that our approach improves the performance of bilingual lexicon induction compared with previous methods. Shuo Ren 0002, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001 |
ACL | 3 |
| 2020 | A Retrieve-and-Rewrite Initialization Method for Unsupervised Machine TranslationabstractThe commonly used framework for unsupervised machine translation builds initial translation models of both translation directions, and then performs iterative back-translation to jointly boost their translation performance.The initialization stage is very important since bad initialization may wrongly squeeze the search space, and too much noise introduced in this stage may hurt the final performance.In this paper, we propose a novel retrieval and rewriting based method to better initialize unsupervised translation models.We first retrieve semantically comparable sentences from monolingual corpora of two languages and then rewrite the target side to minimize the semantic gap between the source and retrieved targets with a designed rewriting model.The rewritten sentence pairs are used to initialize SMT models which are used to generate pseudo data for two NMT models, followed by the iterative back-translation.Experiments show that our method can build better initial unsupervised translation models and improve the final translation performance by over 4 BLEU scores. Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001 |
ACL | 4 |
| 2020 | Curriculum Pre-training for End-to-End Speech TranslationabstractEnd-to-end speech translation poses a heavy burden on the encoder because it has to transcribe, understand, and learn cross-lingual semantics simultaneously.To obtain a powerful encoder, traditional methods pre-train it on ASR data to capture speech features.However, we argue that pre-training the encoder only through simple speech recognition is not enough, and high-level linguistic knowledge should be considered.Inspired by this, we propose a curriculum pre-training method that includes an elementary course for transcription learning and two advanced courses for understanding the utterance and mapping words in two languages.The difficulty of these courses is gradually increasing.Experiments show that our curriculum pre-training method leads to significant improvements on En-De and En-Fr speech translation benchmarks. Chengyi Wang 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Zhenglu Yang |
ACL | 4 |
| 2020 | MIND: A Large-scale Dataset for News RecommendationabstractFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, Ming Zhou. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Fangzhao Wu, Jiun-Hung Chen, Chuhan Wu, Tao Qi 0001, Jianxun Lian, Xing Xie 0001, Jianfeng Gao 0001, Winnie Wu, Ming Zhou 0001 |
ACL | 11 |
| 2020 | Improving Neural Machine Translation with Soft Template PredictionabstractAlthough neural machine translation (NMT) has achieved significant progress in recent years, most previous NMT models only depend on the source text to generate translation.Inspired by the success of template-based and syntax-based approaches in other fields, we propose to use extracted templates from tree structures as soft target templates to guide the translation procedure.In order to learn the syntactic structure of the target sentences, we adopt the constituency-based parse tree to generate candidate templates.We incorporate the template information into the encoder-decoder framework to jointly utilize the templates and source text.Experiments show that our model significantly outperforms the baseline models on four benchmarks and demonstrate the effectiveness of soft target templates. Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Zhoujun Li 0001, Ming Zhou 0001 |
ACL | 5 |
| 2020 | Document Modeling with Graph Attention Networks for Multi-grained Machine Reading ComprehensionabstractNatural Questions is a new challenging machine reading comprehension benchmark with two-grained answers, which are a long answer (typically a paragraph) and a short answer (one or more entities inside the long answer).Despite the effectiveness of existing methods on this benchmark, they treat these two sub-tasks individually during training while ignoring their dependencies.To address this issue, we present a novel multi-grained machine reading comprehension framework that focuses on modeling documents at their hierarchical nature, which are different levels of granularity: documents, paragraphs, sentences, and tokens.We utilize graph attention networks to obtain different levels of representations so that they can be learned simultaneously.The long and short answers can be extracted from paragraphlevel representation and token-level representation, respectively.In this way, we can model the dependencies between the two-grained answers to provide evidence for each other.We jointly train the two sub-tasks, and our experiments show that our approach significantly outperforms previous systems at both long and short answer criteria. Bo Zheng 0010, Haoyang Wen, Yaobo Liang, Nan Duan 0001, Wanxiang Che, Daxin Jiang, Ming Zhou 0001, Ting Liu 0001 |
ACL | 7 |
| 2020 | LogicalFactChecker: Leveraging Logical Operations for Fact Checking with Graph Module NetworkabstractWanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan, Ming Zhou, Ming Gong, Linjun Shou, Daxin Jiang, Jiahai Wang, Jian Yin. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Wanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan 0001, Ming Zhou 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Jiahai Wang, Jian Yin 0001 |
ACL | 5 |
| 2020 | Reasoning Over Semantic-Level Graph for Fact CheckingabstractFact checking is a challenging task because verifying the truthfulness of a claim requires reasoning about multiple retrievable evidence.In this work, we present a method suitable for reasoning about the semantic-level structure of evidence.Unlike most previous works, which typically represent evidence sentences with either string concatenation or fusing the features of isolated evidence sentences, our approach operates on rich semantic structures of evidence obtained by semantic role labeling.We propose two mechanisms to exploit the structure of evidence while leveraging the advances of pre-trained models like BERT, GPT or XLNet.Specifically, using XLNet as the backbone, we first utilize the graph structure to re-define the relative distances of words, with the intuition that semantically related words should have short distances.Then, we adopt graph convolutional network and graph attention network to propagate and aggregate information from neighboring nodes on the graph.We evaluate our system on FEVER, a benchmark dataset for fact checking, and find that rich structural information is helpful and both our graph-based mechanisms improve the accuracy.Our model is the state-of-the-art system in terms of both official evaluation metrics, namely claim verification accuracy and FEVER score. Wanjun Zhong, Jingjing Xu 0001, Duyu Tang, Zenan Xu, Nan Duan 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001 |
ACL | 6 |
| 2020 | Unsupervised Fine-tuning for Text ClusteringabstractFine-tuning with pre-trained language models (e.g.BERT) has achieved great success in many language understanding tasks in supervised settings (e.g.text classification).However, relatively little work has been focused on applying pre-trained models in unsupervised settings, such as text clustering.In this paper, we propose a novel method to fine-tune pre-trained models unsupervisedly for text clustering, which simultaneously learns text representations and cluster assignments using a clustering oriented loss.Experiments on three text clustering datasets (namely TREC-6, Yelp, and DBpedia) show that our model outperforms the baseline methods and achieves stateof-the-art results. Shaohan Huang, Furu Wei, Lei Cui 0001, Xingxing Zhang 0002, Ming Zhou 0001 |
COLING | 5 |
| 2020 | DocBank: A Benchmark Dataset for Document Layout AnalysisabstractDocument layout analysis usually relies on computer vision models to understand documents while ignoring textual information that is vital to capture.Meanwhile, high quality labeled datasets with both visual and textual information are still insufficient.In this paper, we present DocBank, a benchmark dataset that contains 500K document pages with fine-grained tokenlevel annotations for document layout analysis.DocBank is constructed using a simple yet effective way with weak supervision from the L A T E X documents available on the arXiv.com.With DocBank, models from different modalities can be compared fairly and multi-modal approaches will be further investigated and boost the performance of document layout analysis.We build several strong baselines and manually split train/dev/test sets for evaluation.Experiment results show that models trained on DocBank accurately recognize the layout information for a variety of documents.The DocBank dataset is publicly available at https: //github.com/doc-analysis/DocBank. Minghao Li 0004, Yiheng Xu, Lei Cui 0001, Shaohan Huang, Furu Wei, Zhoujun Li 0001, Ming Zhou 0001 |
COLING | 7 |
| 2020 | At Which Level Should We Extract? An Empirical Analysis on Extractive Document SummarizationabstractExtractive methods have been proven effective in automatic document summarization.Previous works perform this task by identifying informative contents at sentence level.However, it is unclear whether performing extraction at sentence level is the best solution.In this work, we show that unnecessity and redundancy issues exist when extracting full sentences, and extracting sub-sentential units is a promising alternative.Specifically, we propose extracting sub-sentential units based on the constituency parsing tree.A neural extractive model which leverages the subsentential information and extracts them is presented.Extensive experiments and analyses show that extracting sub-sentential units performs competitively comparing to full sentence extraction under the evaluation of both automatic and human evaluations.Hopefully, our work could provide some inspiration of the basic extraction units in extractive summarization for future research. Qingyu Zhou, Furu Wei, Ming Zhou 0001 |
COLING | 3 |
| 2020 | Improving the Efficiency of Grammatical Error Correction with Erroneous Span Detection and CorrectionabstractWe propose a novel language-independent approach to improve the efficiency for Grammatical Error Correction (GEC) by dividing the task into two subtasks: Erroneous Span Detection (ESD) and Erroneous Span Correction (ESC).ESD identifies grammatically incorrect text spans with an efficient sequence tagging model.Then, ESC leverages a seq2seq model to take the sentence with annotated erroneous spans as input and only outputs the corrected text for these spans.Experiments show our approach performs comparably to conventional seq2seq approaches in both English and Chinese GEC benchmarks with less than 50% time cost for inference. Mengyun Chen, Tao Ge 0001, Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001 |
EMNLP (1) | 5 |
| 2020 | XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationabstractYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yaobo Liang, Nan Duan 0001, Yeyun Gong, Ning Wu 0013, Fenfei Guo, Weizhen Qi, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Dong Bo Cui, Sining Wei, Taroon Bharti, Jiun-Hung Chen, Winnie Wu, Fan Yang 0024, Daniel Campos, Rangan Majumder, Ming Zhou 0001 |
EMNLP (1) | 24 |
| 2020 | Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous SpaceabstractIn this paper, we propose a novel data augmentation method, referred to as Controllable Rewriting based Question Data Augmentation (CRQDA), for machine reading comprehension (MRC), question generation, and question-answering natural language inference tasks.We treat the question data augmentation task as a constrained question rewriting problem to generate context-relevant, high-quality, and diverse question data samples.CRQDA utilizes a Transformer autoencoder to map the original discrete question into a continuous embedding space.It then uses a pre-trained MRC model to revise the question representation iteratively with gradientbased optimization.Finally, the revised question representations are mapped back into the discrete space, which serve as additional question data.Comprehensive experiments on SQuAD 2.0, SQuAD 1.1 question generation, and QNLI tasks demonstrate the effectiveness of CRQDA 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Jiusheng Chen, Jiancheng Lv 0001, Nan Duan 0001, Ming Zhou 0001 |
EMNLP (1) | 8 |
| 2020 | Leveraging Declarative Knowledge in Text and First-Order Logic for Fine-Grained Propaganda DetectionabstractRuize Wang, Duyu Tang, Nan Duan, Wanjun Zhong, Zhongyu Wei, Xuanjing Huang, Daxin Jiang, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Duyu Tang, Nan Duan 0001, Wanjun Zhong, Zhongyu Wei, Xuanjing Huang 0001, Daxin Jiang, Ming Zhou 0001 |
EMNLP (1) | 8 |
| 2020 | BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingabstractIn this paper, we propose a novel model compression approach to effectively compress BERT by progressive module replacing.Our approach first divides the original BERT into several modules and builds their compact substitutes.Then, we randomly replace the original modules with their substitutes to train the compact modules to mimic the behavior of the original modules.We progressively increase the probability of replacement through the training.In this way, our approach brings a deeper level of interaction between the original and compact models.Compared to the previous knowledge distillation approaches for BERT compression, our approach does not introduce any additional loss function.Our approach outperforms existing knowledge distillation approaches on GLUE benchmark, showing a new perspective of model compression.1 Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Furu Wei, Ming Zhou 0001 |
EMNLP (1) | 5 |
| 2020 | Neural Deepfake Detection with Factual Structure of TextabstractDeepfake detection, the task of automatically discriminating machine-generated text, is increasingly critical with recent advances in natural language generative models.Existing approaches to deepfake detection typically represent documents with coarse-grained representations.However, they struggle to capture factual structures of documents, which is a discriminative factor between machinegenerated and human-written text according to our statistical analysis.To address this, we propose a graph-based model that utilizes the factual structure of a document for deepfake detection of text.Our approach represents the factual structure of a given document as an entity graph, which is further utilized to learn sentence representations with a graph neural network.Sentence representations are then composed to a document representation for making predictions, where consistent relations between neighboring sentences are sequentially modeled.Results of experiments on two public deepfake datasets show that our approach significantly improves strong base models built with RoBERTa.Model analysis further indicates that our model can distinguish the difference in the factual structure between machine-generated text and humanwritten text. Wanjun Zhong, Duyu Tang, Zenan Xu, Nan Duan 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001 |
EMNLP (1) | 6 |
| 2020 | Pre-training for Abstractive Document Summarization by Reinstating Source TextabstractAbstractive document summarization is usually modeled as a sequence-to-sequence (SEQ2SEQ) learning problem.Unfortunately, training large SEQ2SEQ based summarization models on limited supervised summarization data is challenging.This paper presents three sequence-to-sequence pre-training (in shorthand, STEP) objectives which allow us to pre-train a SEQ2SEQ based abstractive summarization model on unlabeled text.The main idea is that, given an input text artificially constructed from a document, a model is pre-trained to reinstate the original document.These objectives include sentence reordering, next sentence generation and masked document generation, which have close relations with the abstractive document summarization task.Experiments on two benchmark summarization datasets (i.e., CNN/DailyMail and New York Times) show that all three objectives can improve performance upon baselines.Compared to models pre-trained on large-scale data (≥160GB), our method, with only 19GB text for pre-training, achieves comparable results, which demonstrates its effectiveness.Code and models are public available at https://github.com/ zoezou2015/abs_pretraining. Xingxing Zhang 0002, Wei Lu 0011, Furu Wei, Ming Zhou 0001 |
EMNLP (1) | 5 |
| 2020 | Self-Adversarial Learning with Comparative Discrimination for Text Generation
Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Furu Wei, Ming Zhou 0001 |
ICLR | 5 |
| 2020 | UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingabstractWe propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseudo-masked language model (PMLM). Given an input text with masked tokens, we rely on conventional masks to learn inter-relations between corrupted tokens and context via autoencoding, and pseudo masks to learn intra-relations between masked spans via partially autoregressive modeling. With well-designed position embeddings and self-attention masks, the context encodings are reused to avoid redundant computation. Moreover, conventional masks used for autoencoding provide global masking information, so that all the position embeddings are accessible in partially autoregressive language modeling. In addition, the two tasks pre-train a unified language model as a bidirectional encoder and a sequence-to-sequence decoder, respectively. Our experiments show that the unified language models pre-trained using PMLM achieve new state-of-the-art results on a wide range of language understanding and generation tasks across several widely used benchmarks. The code and pre-trained models are available at https://github.com/microsoft/unilm. Hangbo Bao, Li Dong 0004, Furu Wei, Wenhui Wang 0003, Nan Yang 0002, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
ICML | 10 |
| 2020 | Semantic Mask for Transformer Based End-to-End Speech RecognitionabstractAttention-based encoder-decoder model has achieved impressive results for both automatic speech recognition (ASR) and text-to-speech (TTS) tasks.This approach takes advantage of the memorization capacity of neural networks to learn the mapping from the input sequence to the output sequence from scratch, without the assumption of prior knowledge such as the alignments.However, this model is prone to overfitting, especially when the amount of training data is limited.Inspired by SpecAugment and BERT, in this paper, we propose a semantic mask based regularization for training such kind of end-toend (E2E) model.The idea is to mask the input features corresponding to a particular output token, e.g., a word or a wordpiece, in order to encourage the model to fill the token based on the contextual information.While this approach is applicable to the encoder-decoder framework with any type of neural network architecture, we study the transformer-based model for ASR in this work.We perform experiments on Librispeech 960h and TedLium2 data sets, and achieve the state-of-the-art performance on the test set in the scope of E2E models. Chengyi Wang 0002, Yu Wu 0012, Yujiao Du, Jinyu Li 0001, Shujie Liu 0001, Liang Lu 0001, Shuo Ren 0002, Guoli Ye, Sheng Zhao 0002, Ming Zhou 0001 |
INTERSPEECH | 10 |
| 2020 | Low Latency End-to-End Streaming Speech Recognition with a Scout NetworkabstractThe attention-based Transformer model has achieved promising results for speech recognition (SR) in the offline mode.However, in the streaming mode, the Transformer model usually incurs significant latency to maintain its recognition accuracy when applying a fixed-length look-ahead window in each encoder layer.In this paper, we propose a novel low-latency streaming approach for Transformer models, which consists of a scout network and a recognition network.The scout network detects the whole word boundary without seeing any future frames, while the recognition network predicts the next subword by utilizing the information from all the frames before the predicted boundary.Our model achieves the best performance (2.7/6.4WER) with only 639 ms latency on the test-clean and test-other data sets of Librispeech. Chengyi Wang 0002, Yu Wu 0012, Liang Lu 0001, Shujie Liu 0001, Jinyu Li 0001, Guoli Ye, Ming Zhou 0001 |
INTERSPEECH | 7 |
| 2020 | MoBoAligner: A Neural Alignment Model for Non-Autoregressive TTS with Monotonic Boundary SearchabstractTo speed up the inference of neural speech synthesis, nonautoregressive models receive increasing attention recently.In non-autoregressive models, additional durations of text tokens are required to make a hard alignment between the encoder and the decoder.The duration-based alignment plays a crucial role since it controls the correspondence between text tokens and spectrum frames and determines the rhythm and speed of synthesized audio.To get better duration-based alignment and improve the quality of non-autoregressive speech synthesis, in this paper, we propose a novel neural alignment model named MoBoAligner.Given the pairs of the text and mel spectrum, MoBoAligner tries to identify the boundaries of text tokens in the given mel spectrum frames based on the tokenframe similarity in the neural semantic space with an end-toend framework.With these boundaries, durations can be extracted and used in the training of non-autoregressive TTS models.Compared with the duration extracted by TransformerTTS, MoBoAligner brings improvement for the non-autoregressive TTS model on MOS (our 3.74 comparing to baseline's 3.44).Besides, MoBoAligner is task-specified and lightweight, which reduces the parameter number by 45% and the training time consuming by 30%. Naihan Li, Shujie Liu 0001, Sheng Zhao 0002, Ming Zhou 0001 |
INTERSPEECH | 6 |
| 2020 | LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingabstractPre-training techniques have been verified successfully in a variety of NLP tasks in recent years. Despite the widespread use of pre-training models for NLP applications, they almost exclusively focus on text-level manipulation, while neglecting layout and style information that is vital for document image understanding. In this paper, we propose the LayoutLM to jointly model interactions between text and layout information across scanned document images, which is beneficial for a great number of real-world document image understanding tasks such as information extraction from scanned documents. Furthermore, we also leverage image features to incorporate words' visual information into LayoutLM. To the best of our knowledge, this is the first time that text and layout are jointly learned in a single framework for document-level pre-training. It achieves new state-of-the-art results in several downstream tasks, including form understanding (from 70.72 to 79.27), receipt understanding (from 94.02 to 95.24) and document image classification (from 93.07 to 94.42). The code and pre-trained LayoutLM models are publicly available at https://aka.ms/layoutlm. Yiheng Xu, Minghao Li 0004, Lei Cui 0001, Shaohan Huang, Furu Wei, Ming Zhou 0001 |
KDD | 6 |
| 2020 | TableBank: Table Benchmark for Image-based Table Detection and RecognitionabstractWe present TableBank, a new image-based table detection and recognition dataset built with novel weak supervision from Word and Latex documents on the internet. Existing research for image-based table detection and recognition usually fine-tunes pre-trained models on out-of-domain data with a few thousand human-labeled examples, which is difficult to generalize on real-world applications. With TableBank that contains 417K high quality labeled tables, we build several strong baselines using state-of-the-art models with deep neural networks. We make TableBank publicly available and hope it will empower more deep learning approaches in the table detection and recognition task. The dataset and models can be downloaded from https://github.com/doc-analysis/TableBank. Minghao Li 0004, Lei Cui 0001, Shaohan Huang, Furu Wei, Ming Zhou 0001, Zhoujun Li 0001 |
LREC | 5 |
| 2020 | Learning Semantic Concepts and Temporal Alignment for Narrated Video Procedural CaptioningabstractVideo captioning is a fundamental task for visual understanding. Previous works employ end-to-end networks to learn from the low-level vision feature and generate descriptive captions, which are hard to recognize fine-grained objects and lacks the understanding of crucial semantic concepts. According to DPC [19], these concepts generally present in the narrative transcripts of the instructional videos. The incorporation of transcript and video can improve the captioning performance. However, DPC directly concatenates the embedding of transcript with video features, which is incapable of fusing language and vision features effectively and leads to the temporal mis-alignment between transcript and video. This motivates us to 1) learn the semantic concepts explicitly and 2) design a temporal alignment mechanism to better align the video and transcript for the captioning task. In this paper, we start with an encoder-decoder backbone using transformer models. Firstly, we design a semantic concept prediction module as a multi-task to train the encoder in a supervised way. Then, we develop an attention based cross-modality temporal alignment method that combines the sequential video frames and transcript sentences. Finally, we adopt a copy mechanism to enable the decoder(generation) module to copy important concepts from source transcript directly. The extensive experimental results demonstrate the effectiveness of our model, which achieves state-of-the-art results on YouCookII dataset. Botian Shi, Lei Ji 0001, Zhendong Niu, Nan Duan 0001, Ming Zhou 0001, Xilin Chen 0001 |
ACM Multimedia | 5 |
| 2020 | MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersabstractPre-trained language models (e.g., BERT (Devlin et al., 2018) and its variants) have achieved remarkable success in varieties of NLP tasks. However, these models usually consist of hundreds of millions of parameters which brings challenges for fine-tuning and online serving in real-life applications due to latency and capacity constraints. In this work, we present a simple and effective approach to compress large Transformer (Vaswani et al., 2017) based pre-trained models, termed as deep self-attention distillation. The small model (student) is trained by deeply mimicking the self-attention module, which plays a vital role in Transformer networks, of the large model (teacher). Specifically, we propose distilling the self-attention module of the last Transformer layer of the teacher, which is effective and flexible for the student. Furthermore, we introduce the scaled dot-product between values in the self-attention module as the new deep self-attention knowledge, in addition to the attention distributions (i.e., the scaled dot-product of queries and keys) that have been used in existing works. Moreover, we show that introducing a teacher assistant (Mirzadeh et al., 2019) also helps the distillation of large pre-trained Transformer models. Experimental results demonstrate that our monolingual model outperforms state-of-the-art baselines in different parameter size of student models. In particular, it retains more than 99% accuracy on SQuAD 2.0 and several GLUE benchmark tasks using 50% of the Transformer parameters and computations of the teacher model. We also obtain competitive results in applying deep self-attention distillation to multilingual pre-trained models. Wenhui Wang 0003, Furu Wei, Li Dong 0004, Hangbo Bao, Nan Yang 0002, Ming Zhou 0001 |
NeurIPS | 6 |
| 2020 | ProphetNet-Ads: A Looking Ahead Strategy for Generative Retrieval Models in Sponsored Search Engine
Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Ruofei Zhang, Houqiang Li, Nan Duan 0001, Ming Zhou 0001 |
NLPCC (2) | 9 |
| 2020 | A Joint Sentence Scoring and Selection Framework for Neural Extractive Document SummarizationabstractExtractive document summarization methods aim to extract important sentences to form a summary. Previous works perform this task by first scoring all sentences in the document then selecting most informative ones; while we propose to jointly learn the two steps with a novel end-to-end neural network framework. Specifically, the sentences in the input document are represented as real-valued vectors through a neural document encoder. Then the method builds the output summary by extracting important sentences one by one. Different from previous works, the proposed joint sentence scoring and selection framework directly predicts the relative sentence importance score according to both sentence content and previously selected sentences. We evaluate the proposed framework with two realizations: a hierarchical recurrent neural network based model; and a pre-training based model that uses BERT as the document encoder. Experiments on two datasets show that the proposed joint framework outperforms the state-of-the-art extractive summarization models which treat sentence scoring and selection as two subtasks. Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Joint Learning of Question Answering and Question GenerationabstractQuestion answering (QA) and question generation (QG) are closely related tasks that could improve each other; however, the connection of these two tasks is not well explored in the literature. In this paper, we present two training algorithms for learning better QA and QG models through leveraging one another. The first algorithm extends Generative Adversarial Network (GAN), which selectively incorporates artificially generated instances as additional QA training data. The second algorithm is an extension of dual learning, which incorporates the probabilistic correlation of QA and QG as additional regularization in training objectives. To test the scalability of our algorithms, we conduct experiments on both document based and table based question answering tasks. Results show that both algorithms improve a QA model in terms of accuracy and QG model in terms of BLEU score. Moreover, we find that the performance of a QG model could be easily improved by a QA model via policy gradient, however, directly applying GAN that regards all the generated questions as negative instances could not improve the accuracy of the QA model. Our algorithm that selectively assigns labels to generated questions would bring a performance boost. Duyu Tang, Nan Duan 0001, Tao Qin 0001, Shujie Liu 0001, Ming Zhou 0001, Yuanhua Lv, Wenpeng Yin 0001, Bing Qin 0001, Ting Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2019 | Response Generation by Context-Aware Prototype EditingabstractOpen domain response generation has achieved remarkable progress in recent years, but sometimes yields short and uninformative responses. We propose a new paradigm, prototypethen-edit for response generation, that first retrieves a prototype response from a pre-defined index and then edits the prototype response according to the differences between the prototype context and current context. Our motivation is that the retrieved prototype provides a good start-point for generation because it is grammatical and informative, and the post-editing process further improves the relevance and coherence of the prototype. In practice, we design a contextaware editing model that is built upon an encoder-decoder framework augmented with an editing vector. We first generate an edit vector by considering lexical differences between a prototype context and current context. After that, the edit vector and the prototype response representation are fed to a decoder to generate a new response. Experiment results on a large scale dataset demonstrate that our new paradigm significantly increases the relevance, diversity and originality of generation results, compared to traditional generative models. Furthermore, our model outperforms retrieval-based methods in terms of relevance and originality. Yu Wu 0012, Furu Wei, Shaohan Huang, Yunli Wang, Zhoujun Li 0001, Ming Zhou 0001 |
AAAI | 6 |
| 2019 | Unsupervised Neural Machine Translation with SMT as Posterior RegularizationabstractWithout real bilingual corpus available, unsupervised Neural Machine Translation (NMT) typically requires pseudo parallel data generated with the back-translation method for the model training. However, due to weak supervision, the pseudo data inevitably contain noises and errors that will be accumulated and reinforced in the subsequent training process, leading to bad translation performance. To address this issue, we introduce phrase based Statistic Machine Translation (SMT) models which are robust to noisy data, as posterior regularizations to guide the training of unsupervised NMT models in the iterative back-translation process. Our method starts from SMT models built with pre-trained language models and word-level translation tables inferred from cross-lingual embeddings. Then SMT and NMT models are optimized jointly and boost each other incrementally in a unified EM framework. In this way, (1) the negative effect caused by errors in the iterative back-translation process can be alleviated timely by SMT filtering noises from its phrase tables; meanwhile, (2) NMT can compensate for the deficiency of fluency inherent in SMT. Experiments conducted on en-fr and en-de translation tasks show that our method outperforms the strong baseline and achieves new state-of-the-art unsupervised machine translation performance. Shuo Ren 0002, Zhirui Zhang, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001 |
AAAI | 4 |
| 2019 | Regularizing Neural Machine Translation by Target-Bidirectional AgreementabstractAlthough Neural Machine Translation (NMT) has achieved remarkable progress in the past several years, most NMT systems still suffer from a fundamental shortcoming as in other sequence generation tasks: errors made early in generation process are fed as inputs to the model and can be quickly amplified, harming subsequent sequence generation. To address this issue, we propose a novel model regularization method for NMT training, which aims to improve the agreement between translations generated by left-to-right (L2R) and right-to-left (R2L) NMT decoders. This goal is achieved by introducing two Kullback-Leibler divergence regularization terms into the NMT training objective to reduce the mismatch between output probabilities of L2R and R2L models. In addition, we also employ a joint training strategy to allow L2R and R2L models to improve each other in an interactive update process. Experimental results show that our proposed method significantly outperforms state-of-the-art baselines on Chinese-English and English-German translation tasks. Zhirui Zhang, Shuangzhi Wu, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Tong Xu 0001 |
AAAI | 5 |
| 2019 | Automatic Grammatical Error Correction for Sequence-to-sequence Text Generation: An Empirical StudyabstractSequence-to-sequence (seq2seq) models have achieved tremendous success in text generation tasks.However, there is no guarantee that they can always generate sentences without grammatical errors.In this paper, we present a preliminary empirical study on whether and how much automatic grammatical error correction can help improve seq2seq text generation.We conduct experiments across various seq2seq text generation tasks including machine translation, formality style transfer, sentence compression and simplification.Experiments show the state-of-the-art grammatical error correction system can improve the grammaticality of generated text and can bring taskoriented improvements in the tasks where target sentences are in a formal style. Tao Ge 0001, Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001 |
ACL (1) | 4 |
| 2019 | Coupling Retrieval and Meta-Learning for Context-Dependent Semantic Parsingabstract5 9 469 9 65 67 1 016 9 93969696 5 5 67 21 22 67 !"626 ! 5 67#$12 !6%5 6&5 6 '()*+(,,-,./0100234,5671 )8*+9 )2+):*72;01 < = )2>29 ?)8 = 9 1 @A B702C3,2CD)@E0F,8 01 ,8 @,.G9 C/01 0H20-@= 9 =023I8 ,+)= = 9 2C:B702CJ(,7:IA KA 4(9 20 !L9 +8 ,= ,. 1K)= )08 +(H= 9 0:G)9 M 9 2C:4(9 20 NOPQRSTUVWXYZ[X\\]SX^UVWXY_`\S\P`aRP`b^ NRPcW^O[^W^RPW^[VX^OdeQP_UVXbfQ\Qgc`bQV hi j 21 (9 =606)8 :k)68 )= )21020668 ,0+(1 ,9 2+,8 < 6,8 01 )8 )1 8 9 )?)3301 06,9 21 =0== 766,8 1 9 2C)?9 < 3)2+).,8+,21 )l1 < 3)6)23)21= )5021 9 +608 = 9 2C: = 7+(0=C)2)8 01 9 2C= ,78 +)+,3)+,239 1 9 ,2)3,2 1 ()+-0= =)2?9 8 ,25)21 Am780668 ,0+(201 78 0--@ +,5F9 2)=08 )1 8 9 )?0-5,3)-02305)1 0< -)08 2)8 : k()8 )1 ().,8 5)8-)08 2=1 ,n23= 9 59 -08301 < 06,9 21 =. 8 ,51 ()1 8 09 29 2C301 0:0231 ()-01 1 )8 +,2= 9 3)8 =8 )1 8 9 )?)3301 06,9 21 =0=06= )73,1 0= o .,8. 0= 103061 01 9 ,2A*6)+9 n+0--@:,788 )1 8 9 )?< )89 =0+,21 )l1 < 0k08 ))2+,3)8 < 3)+,3)85,3)-k9 1 (0-01 )21?08 9 0F-)k(9 +(1 0o)=+,21 )l1)2< ?9 8 ,25)219 21 ,+,2= 9 3)8 01 9 ,2:023,785)1 0< -)08 2)8-)08 2=1 ,71 9 -9 J)8 )1 8 9 )?)3301 06,9 21 = 9 205,3)-< 0C2,= 1 9 +5)1 0< -)08 29 2C608 039 C5 .,8. 0= 103061 01 9 ,2Ap)+,237+1)l6)8 9 5)21 = ,24mq4m/r0234*sH 301 0= )1 = :k()8 ) 1 ()+,21 )l18 ).)8 =1 ,+-0= =)2?9 8 ,25)219 2t H< uH+,3)=023+,2?)8 = 01 9 ,20-(9 = 1 ,8 @:8 )= 6)+< 1 9 ?)-@Ap)7=)=)v7)2+)< 1 ,< 0+1 9 ,25,3)-0= 1 ()F0= )=)5021 9 +608 = )8 :k(9 +(6)8 .,8 5=1 () = 1 01 )< ,. < 1 ()< 08 10++78 0+@,2F,1 (301 0= )1 = AK)< = 7-1 == (,k1 (01F,1 (1 ()+,21 )l1 < 0k08 )8 )1 8 9 )?< )80231 ()5)1 0< -)08 29 2C=1 8 01 )C@9 568 ,?)0+< +78 0+@:023,780668 ,0+(6)8 .,8 5=F)1 1 )81 (02 8 )1 8 9 )?)< 023< )39 1F0= )-9 2)= A w x 6 12 5 16 4,21 )l1 < 3)6)23)21= )5021 9 +608 = 9 2C09 5=1 ,506 0201 78 0--02C70C)71 1 )8 02+)1 ,0=1 8 7+1 78 0--,C9 < +0-.,8 5y )A CA= ,78 +)+,3)z+,239 1 9 ,2)3,20C9 ?< )2+,21 )l1y )A CA+-0= =)2?9 8 ,25)21 zy E9 2C)10-A : {|}~E,2C)10-A :{|}~j @@)8)10-A :{|}j @< )8)10-A :{|}*7(8)10-A :{|}*7(8023H8 1 J9 : {|}z A*1 02308 30668 ,0+()=1 @69 +0--@-)08 20,2)< = 9 J)< n1 = < 0--5,3)-,21 ())21 9 8 )1 8 09 29 2C301 0= )1 : k(9 +(9 =. )3k9 1 ()0+()l056-)9 239 ?9 370--@9 21 () 1 8 09 29 2C6(0= )02350o)=68 )39 +1 9 ,2=.,8)0+(1 )= 1 )l056-)9 21 ()9 2. )8 )2+)6(0= )A,k)?)8 :1 0o9 2C p,8 o3,2)k(9 -)1 (9 =071 (,8k0=029 21 )8 201L9 +8 ,= ,. 1 K)= )08 +(A +,3)C)2)8 01 9 ,20=02)l056-):68 ,C8 055)8 =7= 7< 0--@3,2,1k8 9 1 )+,3)=.8 ,5= +8 01 +(9 21 ()8 )0k,8 -3Ap()21 ()@k8 9 1 )069 )+),.+,3)920608 < 1 9 +7-08)2?9 8 ,25)21 :1 ()@1 @69 +0--@-)?)8 0C)60= 1 )l6)8 9 )2+),2k8 9 1 9 2C,88 )039 2C+,3)=9 21 ()= 9 59 < -08= 9 1 701 9 ,20=0C79 302+)AL)02k(9 -):301 06,9 21 = .,801 0= o50@?08 @k9 3)-@y 702C)10-A :{|}0z : 1 (7=9 19 =3)= 9 8 0F-)1 ,-)08 206)8 = ,20-9 J)35,3< )-.,81 ()1 08 C)1301 06,9 21 Aj 21 (9 =k,8 o:k)= 1 73@ (,k1 ,071 ,501 9 +0--@8 )1 8 9 )?)= 9 59 -08301 06,9 21 =9 2 0+,21 )l1 < 3)6)23)21= +)208 9 ,0237= )1 ()50=1 () = 766,8 1 9 2C)?9 3)2+)1 ,. 0+9 -9 1 01 )= )5021 9 +608 = 9 2CA '()8 )08 )8 )+)2101 1 )561 =01)l6-,9 1 9 2C8 )< 1 8 9 )?)3)l056-)=1 ,9 568 ,?)1 ()C)2)8 01 9 ,2,.-,C< 9 +0-.,8 50231 )l1 AK)1 8 9 )?)< 023< )39 10668 ,0+()= y 0= (9 5,1 ,)10-A :{|}702C)10-A :{|}Fp7 )10-A :{|}B7)10-A :{|}z1 @69 +0--@n8 = 17= )0 +,21 )l1 < 9 23)6)23)218 )1 8 9 )?)81 ,n231 ()5,= 18 )-< )?021301 06,9 21 :0231 ()27= )9 10=020339 1 9 ,20-9 2671,.1 ())39 1 9 2C5,3)-A,k)?)8:0+,21 )l1 < 0k08 )8 )1 8 9 )?)89 =?)8 @9 56,8 1 021.,81 ()1 0= o,.+,21 )l1 < 3)6)23)21= )5021 9 +608 = 9 2CA,8)l05< 6-)= :0== (,k29 29 C78 )}:+-0= =)2?9 8 ,25)21+02 ()-61 ()8 )1 8 9 )?)83)+9 3)k()1 ()81 ()3)= 9 8 )3+,3) ,. 9 =C)2)8 01 )3F@39 8 )+1 -@ +0--9 2C ,89 1 )8 01 9 2C1 () ¡08 8 0@ 1 ,9 2+8 )5)21)0+()-)5)21 A78 1 ()8 5,8 ):8 )1 8 9 )?)< 023< )39 10668 ,0+()=1 @69 +0--@+,2= 9 3)8,2-@,2) = 9 59 -08)l056-)1 ,)39 1 Aj 2=)5021 9 +608 = 9 2C:1 () 601 1 )8 2,.0= 1 8 7+1 78 0-,71 67150@+,5).8 ,539 .< .)8 )218 )1 8 9 )?)3)l056-)= A'()8 )0-= ,)l9 = 1k,8 o= 1 ,71 9 -9 J)57-1 9 6-))l056-)=1 ,C79 3)1 ()= )5021 9 + 608 = )8y 0@01 9)10-A :{|}702C)10-A :{|}0z : (,k)?)8 :1 ()= )0668 ,0+()=)9 1 ()87= )0()78 9 = 1 9 + k0@1 ,)l6-,9 11 ()8 )1 8 9 )?)3-,C9 +0-.,8 5= 7+(0= 9 2+8 )0= 9 2C1 ()68 ,F0F9 -9 1 @,.0+1 9 ,2=y 0@01 9)10-A : {|}z,87= )08 )-)?02+).72+1 9 ,23)= 9 C2)3023 -)08 2)3F0= )3,2)l6)8 1 9 = )0F,711 ()1 08 C)1-,C9 < +0-.,8 5y 702C)10-A :{|}0z Ap()2k)+,2= 9 3)8 1 ()+,21 )l1)2?9 8 ,25)21 :9 1 ¢ =2,21 8 9 ?9 0-1 ,3)= 9 C2 Daya Guo, Duyu Tang, Nan Duan 0001, Ming Zhou 0001, Jian Yin 0001 |
ACL (1) | 4 |
| 2019 | Dense Procedure Captioning in Narrated Instructional VideosabstractUnderstanding narrated instructional videos is important for both research and real-world web applications.Motivated by video dense captioning, we propose a model to generate procedure captions from narrated instructional videos which are a sequence of stepwise clips with description.Previous works on video dense captioning learn video segments and generate captions without considering transcripts.We argue that transcripts in narrated instructional videos can enhance video representation by providing fine-grained complimentary and semantic textual information.In this paper, we introduce a framework to ( 1) extract procedures by a cross-modality module, which fuses video content with the entire transcript; and (2) generate captions by encoding video frames as well as a snippet of transcripts within each extracted procedure.Experiments show that our model can achieve state-of-the-art performance in procedure extraction and captioning, and the ablation studies demonstrate that both the video frames and the transcripts are important for the task. Botian Shi, Lei Ji 0001, Yaobo Liang, Nan Duan 0001, Peng Chen 0029, Zhendong Niu, Ming Zhou 0001 |
ACL (1) | 7 |
| 2019 | HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document SummarizationabstractNeural extractive summarization models usually employ a hierarchical encoder for document encoding and they are trained using sentence-level labels, which are created heuristically using rule-based methods.Training the hierarchical encoder with these inaccurate labels is challenging.Inspired by the recent work on pre-training transformer sentence encoders (Devlin et al., 2018), we propose HIBERT (as shorthand for HIerachical Bidirectional Encoder Representations from Transformers) for document encoding and a method to pre-train it using unlabeled data.We apply the pre-trained HIBERT to our summarization model and it outperforms its randomly initialized counterpart by 1.25 ROUGE on the CNN/Dailymail dataset and by 2.0 ROUGE on a version of New York Times dataset.We also achieve the state-of-the-art performance on these two datasets. Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001 |
ACL (1) | 3 |
| 2019 | BERT-based Lexical SubstitutionabstractPrevious studies on lexical substitution tend to obtain substitute candidates by finding the target word's synonyms from lexical resources (e.g., WordNet) and then rank the candidates based on its contexts.These approaches have two limitations: (1) They are likely to overlook good substitute candidates that are not the synonyms of the target words in the lexical resources;(2) They fail to take into account the substitution's influence on the global context of the sentence.To address these issues, we propose an end-toend BERT-based lexical substitution approach which can propose and validate substitute candidates without using any annotated data or manually curated resources.Our approach first applies dropout to the target word's embedding for partially masking the word, allowing BERT to take balanced consideration of the target word's semantics and contexts for proposing substitute candidates, and then validates the candidates based on their substitution's influence on the global contextualized representation of the sentence.Experiments show our approach performs well in both proposing and ranking substitute candidates, achieving the state-of-the-art results in both LS07 and LS14 benchmarks. Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Furu Wei, Ming Zhou 0001 |
ACL (1) | 5 |
| 2019 | Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual TasksabstractHaoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, Ming Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Haoyang Huang, Yaobo Liang, Nan Duan 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Ming Zhou 0001 |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Explicit Cross-lingual Pre-training for Unsupervised Machine TranslationabstractShuo Ren, Yu Wu, Shujie Liu, Ming Zhou, Shuai Ma. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Asking Clarification Questions in Knowledge-Based Question AnsweringabstractJingjing Xu, Yuechen Wang, Duyu Tang, Nan Duan, Pengcheng Yang, Qi Zeng, Ming Zhou, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jingjing Xu 0001, Yuechen Wang, Duyu Tang, Nan Duan 0001, Qi Zeng 0001, Ming Zhou 0001, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Unified Language Model Pre-training for Natural Language Understanding and GenerationabstractThis paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction. The unified modeling is achieved by employing a shared Transformer network and utilizing specific self-attention masks to control what context the prediction conditions on. UniLM compares favorably with BERT on the GLUE benchmark, and the SQuAD 2.0 and CoQA question answering tasks. Moreover, UniLM achieves new state-of-the-art results on five natural language generation datasets, including improving the CNN/DailyMail abstractive summarization ROUGE-L to 40.51 (2.04 absolute improvement), the Gigaword abstractive summarization ROUGE-L to 35.75 (0.86 absolute improvement), the CoQA generative question answering F1 score to 82.5 (37.1 absolute improvement), the SQuAD question generation BLEU-4 to 22.12 (3.75 absolute improvement), and the DSTC7 document-grounded dialog response generation NIST-4 to 2.67 (human performance is 2.65). The code and pre-trained models are available at https://github.com/microsoft/unilm. Li Dong 0004, Nan Yang 0002, Wenhui Wang 0003, Furu Wei, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
NeurIPS | 8 |
| 2019 | A Tensorized Transformer for Language ModelingabstractLatest development of neural models has connected the encoder and decoder through a self-attention mechanism. In particular, Transformer, which is solely based on self-attention, has led to breakthroughs in Natural Language Processing (NLP) tasks. However, the multi-head attention mechanism, as a key component of Transformer, limits the effective deployment of the model to a resource-limited setting. In this paper, based on the ideas of tensor decomposition and parameters sharing, we propose a novel self-attention model (namely Multi-linear attention) with Block-Term Tensor Decomposition (BTD). We test and verify the proposed attention method on three language modeling tasks (i.e., PTB, WikiText-103 and One-billion) and a neural machine translation task (i.e., WMT-2016 English-German). Multi-linear attention can not only largely compress the model parameters but also obtain performance improvements, compared with a number of language modeling approaches, such as Transformer, Transformer-XL, and Transformer with tensor train decomposition. Xindian Ma, Peng Zhang 0002, Nan Duan 0001, Yuexian Hou, Ming Zhou 0001, Dawei Song 0001 |
NeurIPS | 6 |
| 2019 | Neural Melody Composition from Lyrics
Hangbo Bao, Shaohan Huang, Furu Wei, Lei Cui 0001, Yu Wu 0012, Chuanqi Tan, Ming Zhou 0001 |
NLPCC (1) | 8 |
| 2019 | Document-Based Question Answering Improves Query-Focused Multi-document Summarization
Weikang Li, Xingxing Zhang 0002, Yunfang Wu, Furu Wei, Ming Zhou 0001 |
NLPCC (2) | 5 |
| 2019 | Knowledge-Aware Conversational Semantic Parsing over Web Tables
Duyu Tang, Jingjing Xu 0001, Nan Duan 0001, Bing Qin 0001, Ting Liu 0001, Ming Zhou 0001 |
NLPCC (1) | 8 |
| 2019 | Effective Soft-Adaptation for Neural Machine Translation
Shuangzhi Wu, Dongdong Zhang 0001, Ming Zhou 0001 |
NLPCC (2) | 3 |
| 2019 | Improving Question Answering by Commonsense-Based Pre-training
Wanjun Zhong, Duyu Tang, Nan Duan 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001 |
NLPCC (1) | 4 |
| 2019 | A Sequential Matching Framework for Multi-Turn Response Selection in Retrieval-Based ChatbotsabstractWe study the problem of response selection for multi-turn conversation in retrieval-based chatbots. The task involves matching a response candidate with a conversation context, the challenges for which include how to recognize important parts of the context, and how to model the relationships among utterances in the context. Existing matching methods may lose important information in contexts as we can interpret them with a unified framework in which contexts are transformed to fixed-length vectors without any interaction with responses before matching. This motivates us to propose a new matching framework that can sufficiently carry important information in contexts to matching and model relationships among utterances at the same time. The new framework, which we call a sequential matching framework (SMF), lets each utterance in a context interact with a response candidate at the first step and transforms the pair to a matching vector. The matching vectors are then accumulated following the order of the utterances in the context with a recurrent neural network (RNN) that models relationships among utterances. Context-response matching is then calculated with the hidden states of the RNN. Under SMF, we propose a sequential convolutional network and sequential attention network and conduct experiments on two public data sets to test their performance. Experiment results show that both models can significantly outperform state-of-the-art matching methods. We also show that the models are interpretable with visualizations that provide us insights on how they capture and leverage important information in contexts for matching. Yu Wu 0012, Wei Wu 0014, Chen Xing, Can Xu 0002, Zhoujun Li 0001, Ming Zhou 0001 |
Comput. Linguistics | 6 |
| 2019 | Text Generation From TablesabstractThis paper proposes a neural generative model, namely Table2Seq, to generate a natural language sentence based on a table. Specifically, the model maps a table to continuous vectors and then generates a natural language sentence by leveraging the semantics of a table. Since rare words, e.g., entities and values, usually appear in a table, we develop a flexible copying mechanism that selectively replicates contents from the table to the output sequence. We conduct extensive experiments to demonstrate the effectiveness of our Table2Seq model and the utility of the designed copying mechanism. On the WIKIBIO and SIMPLEQUESTIONS datasets, the Table2Seq model improves the state-of-the-art results from 34.70 to 40.26 and from 33.32 to 39.12 in terms of BLEU-4 scores, respectively. Moreover, we construct an open-domain dataset WIKITABLETEXT that includes 13 318 descriptive sentences for 4962 tables. Our Table2Seq model achieves a BLEU-4 score of 38.23 on WIKITABLETEXT outperforming template-based and language model based approaches. Furthermore, through experiments on 1 M table-query pairs from a search engine, our Table2Seq model considering the structured part of a table, i.e., table attributes and table cells, as additional information outperforms a sequence-to-sequence model considering only the sequential part of a table, i.e., table caption. Junwei Bao 0001, Duyu Tang, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Table-to-Text: Describing Table Region With Natural LanguageabstractIn this paper, we present a generative model to generate a natural language sentence describing a table region, e.g., a row. The model maps a row from a table to a continuous vector and then generates a natural language sentence by leveraging the semantics of a table. To deal with rare words appearing in a table, we develop a flexible copying mechanism that selectively replicates contents from the table in the output sequence. Extensive experiments demonstrate the accuracy of the model and the power of the copying mechanism. On two synthetic datasets, WIKIBIO and SIMPLEQUESTIONS, our model improves the current state-of-the-art BLEU-4 score from 34.70 to 40.26 and from 33.32 to 39.12, respectively. Furthermore, we introduce an open-domain dataset WIKITABLETEXT including 13,318 explanatory sentences for 4,962 tables. Our model achieves a BLEU-4 score of 38.23, which outperforms template based and language model based approaches. Junwei Bao 0001, Duyu Tang, Nan Duan 0001, Yuanhua Lv, Ming Zhou 0001, Tiejun Zhao |
AAAI | 6 |
| 2018 | S-Net: From Answer Extraction to Answer Synthesis for Machine Reading ComprehensionabstractIn this paper, we present a novel approach to machine reading comprehension for the MS-MARCO dataset. Unlike the SQuAD dataset that aims to answer a question with exact text spans in a passage, the MS-MARCO dataset defines the task as answering a question from multiple passages and the words in the answer are not necessary in the passages. We therefore develop an extraction-then-synthesis framework to synthesize answers from extraction results. Specifically, the answer extraction model is first employed to predict the most important sub-spans from the passage as evidence, and the answer synthesis model takes the evidence as additional features along with the question and passage to further elaborate the final answers. We build the answer extraction model with state-of-the-art neural networks for single passage reading comprehension, and propose an additional task of passage ranking to help answer extraction in multiple passages. The answer synthesis model is based on the sequence-to-sequence neural networks with extracted evidences as features. Experiments show that our extraction-then-synthesis method outperforms state-of-the-art methods. Chuanqi Tan, Furu Wei, Nan Yang 0002, Bowen Du 0001, Weifeng Lv, Ming Zhou 0001 |
AAAI | 6 |
| 2018 | Hierarchical Recurrent Attention Network for Response GenerationabstractWe study multi-turn response generation in chatbots where a response is generated according to a conversation context. Existing work has modeled the hierarchy of the context, but does not pay enough attention to the fact that words and utterances in the context are differentially important. As a result, they may lose important information in context and generate irrelevant responses. We propose a hierarchical recurrent attention network (HRAN) to model both the hierarchy and the importance variance in a unified framework. In HRAN, a hierarchical attention mechanism attends to important parts within and among utterances with word level attention and utterance level attention respectively. Chen Xing, Yu Wu 0012, Wei Wu 0014, Yalou Huang, Ming Zhou 0001 |
AAAI | 5 |
| 2018 | Assertion-Based QA With Question-Aware Open Information ExtractionabstractWe present assertion based question answering (ABQA), an open domain question answering task that takes a question and a passage as inputs, and outputs a semi-structured assertion consisting of a subject, a predicate and a list of arguments. An assertion conveys more evidences than a short answer span in reading comprehension, and it is more concise than a tedious passage in passage-based QA. These advantages make ABQA more suitable for human-computer interaction scenarios such as voice-controlled speakers. Further progress towards improving ABQA requires richer supervised dataset and powerful models of text understanding. To remedy this, we introduce a new dataset called WebAssertions, which includes hand-annotated QA labels for 358,427 assertions in 55,960 web passages. To address ABQA, we develop both generative and extractive approaches. The backbone of our generative approach is sequence to sequence learning. In order to capture the structure of the output assertion, we introduce a hierarchical decoder that first generates the structure of the assertion and then generates the words of each field. The extractive approach is based on learning to rank. Features at different levels of granularity are designed to measure the semantic relevance between a question and an assertion. Experimental results show that our approaches have the ability to infer question-aware assertions from a passage. We further evaluate our approaches by incorporating the ABQA results as additional features in passage-based QA. Results on two datasets show that ABQA features significantly improve the accuracy on passage-based QA. Duyu Tang, Nan Duan 0001, Shujie Liu 0001, Daxin Jiang, Ming Zhou 0001, Zhoujun Li 0001 |
AAAI | 7 |
| 2018 | Joint Training for Neural Machine Translation Models with Monolingual DataabstractMonolingual data have been demonstrated to be helpful in improving translation quality of both statistical machine translation (SMT) systems and neural machine translation (NMT) systems, especially in resource-poor or domain adaptation tasks where parallel data are not rich enough. In this paper, we propose a novel approach to better leveraging monolingual data for neural machine translation by jointly learning source-to-target and target-to-source NMT models for a language pair with a joint EM optimization method. The training process starts with two initial NMT models pre-trained on parallel data for each direction, and these two models are iteratively updated by incrementally decreasing translation losses on training data.In each iteration step, both NMT models are first used to translate monolingual data from one language to the other, forming pseudo-training data of the other NMT model. Then two new NMT models are learnt from parallel data together with the pseudo training data. Both NMT models are expected to be improved and better pseudo-training data can be generated in next step. Experiment results on Chinese-English and English-German translation tasks show that our approach can simultaneously improve translation quality of source-to-target and target-to-source models, significantly outperforming strong baseline systems which are enhanced with monolingual data for model training including back-translation. Zhirui Zhang, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Enhong Chen |
AAAI | 4 |
| 2018 | Sequential Copying NetworksabstractCopying mechanism shows effectiveness in sequence-to-sequence based neural network models for text generation tasks, such as abstractive sentence summarization and question generation. However, existing works on modeling copying or pointing mechanism only considers single word copying from the source sentences. In this paper, we propose a novel copying framework, named Sequential Copying Networks (SeqCopyNet), which not only learns to copy single words, but also copies sequences from the input sentence. It leverages the pointer networks to explicitly select a sub-span from the source side to target side, and integrates this sequential copying mechanism to the generation process in the encoder-decoder paradigm. Experiments on abstractive sentence summarization and question generation tasks show that the proposed SeqCopyNet can copy meaningful spans and outperforms the baseline models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Ming Zhou 0001 |
AAAI | 4 |
| 2018 | Neural Document Summarization by Jointly Learning to Score and Select SentencesabstractSentence scoring and sentence selection are two main steps in extractive document summarization systems.However, previous works treat them as two separated subtasks.In this paper, we present a novel end-to-end neural network framework for extractive document summarization by jointly learning to score and select sentences.It first reads the document sentences with a hierarchical encoder to obtain the representation of sentences.Then it builds the output summary by extracting sentences one by one.Different from previous methods, our approach integrates the selection strategy into the scoring model, which directly predicts the relative importance given previously selected sentences.Experiments on the CNN/Daily Mail dataset show that the proposed framework significantly outperforms the state-of-the-art extractive summarization models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao |
ACL (1) | 5 |
| 2018 | Semantic Parsing with Syntax- and Table-Aware SQL GenerationabstractYibo Sun, Duyu Tang, Nan Duan, Jianshu Ji, Guihong Cao, Xiaocheng Feng, Bing Qin, Ting Liu, Ming Zhou. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Duyu Tang, Nan Duan 0001, Jianshu Ji, Guihong Cao, Bing Qin 0001, Ting Liu 0001, Ming Zhou 0001 |
ACL (1) | 9 |
| 2018 | Triangular Architecture for Rare Language TranslationabstractNeural Machine Translation (NMT) performs poor on the low-resource language pair (X, Z), especially when Z is a rare language.By introducing another rich language Y , we propose a novel triangular training architecture (TA-NMT) to leverage bilingual data (Y, Z) (may be small) and (X, Y ) (can be rich) to improve the translation performance of lowresource pairs.In this triangular architecture, Z is taken as the intermediate latent variable, and translation models of Z are jointly optimized with a unified bidirectional EM algorithm under the goal of maximizing the translation likelihood of (X, Y ).Empirical results demonstrate that our method significantly improves the translation quality of rare languages on MultiUN and IWSLT2012 datasets, and achieves even better performance combining back-translation methods. Shuo Ren 0002, Wenhu Chen, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Shuai Ma 0001 |
ACL (1) | 5 |
| 2018 | Fluency Boost Learning and Inference for Neural Grammatical Error CorrectionabstractMost of the neural sequence-to-sequence (seq2seq) models for grammatical error correction (GEC) have two limitations: (1) a seq2seq model may not be well generalized with only limited error-corrected data; (2) a seq2seq model may fail to completely correct a sentence with multiple errors through normal seq2seq inference.We attempt to address these limitations by proposing a fluency boost learning and inference mechanism.Fluency boosting learning generates fluency-boost sentence pairs during training, enabling the error correction model to learn how to improve a sentence's fluency from more instances, while fluency boosting inference allows the model to correct a sentence incrementally through multi-round seq2seq inference until the sentence's fluency stops increasing.Experiments show our approaches improve the performance of seq2seq models for GEC, achieving state-of-the-art results on both CoNLL-2014 and JFLEG benchmark datasets. Tao Ge 0001, Furu Wei, Ming Zhou 0001 |
ACL (1) | 3 |
| 2018 | Bidirectional Generative Adversarial Networks for Neural Machine TranslationabstractGenerative Adversarial Network (GAN) has been proposed to tackle the exposure bias problem of Neural Machine Translation (NMT).However, the discriminator typically results in the instability of the GAN training due to the inadequate training problem: the search space is so huge that sampled translations are not sufficient for discriminator training.To address this issue and stabilize the GAN training, in this paper, we propose a novel Bidirectional Generative Adversarial Network for Neural Machine Translation (BGAN-NMT), which aims to introduce a generator model to act as the discriminator, whereby the discriminator naturally considers the entire translation space so that the inadequate training problem can be alleviated.To satisfy this property, generator and discriminator are both designed to model the joint probability of sentence pairs, with the difference that, the generator decomposes the joint probability with a source language model and a source-to-target translation model, while the discriminator is formulated as a target language model and a target-to-source translation model.To further leverage the symmetry of them, an auxiliary GAN is introduced and adopts generator and discriminator models of original one as its own discriminator and generator respectively.Two GANs are alternately trained to update the parameters.Experiment results on German-English and Chinese-English translation tasks demonstrate that our method not only stabilizes GAN training but also achieves significant improvements over baseline systems. Zhirui Zhang, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Enhong Chen |
CoNLL | 4 |
| 2018 | Visual Question Generation as Dual Task of Visual Question AnsweringabstractVisual question answering (VQA) and visual question generation (VQG) are two trending topics in the computer vision, but they are usually explored separately despite their intrinsic complementary relationship. In this paper, we propose an end-to-end unified model, the Invertible Question Answering Network (iQAN), to introduce question generation as a dual task of question answering to improve the VQA performance. With our proposed invertible bilinear fusion module and parameter sharing scheme, our iQAN can accomplish VQA and its dual task VQG simultaneously. By jointly trained on two tasks with our proposed dual regularizes (termed as Dual Training), our model has a better understanding of the interactions among images, questions and answers. After training, iQAN can take either question or answer as input, and output the counterpart. Evaluated on the CLEVR and VQA2 datasets, our iQAN improves the top-1 accuracy of the prior art MUTAN VQA method by 1.33% and 0.88% (absolute increase) respectiely. We also show that our proposed dual training framework can consistently improve model performances of many popular VQA architectures1. Yikang Li 0002, Nan Duan 0001, Bolei Zhou, Xiao Chu, Wanli Ouyang, Xiaogang Wang 0001, Ming Zhou 0001 |
CVPR | 7 |
| 2018 | Fine-grained Coordinated Cross-lingual Text Stream Alignment for Endless Language Knowledge AcquisitionabstractThis paper proposes to study fine-grained coordinated cross-lingual text stream alignment through a novel information network decipherment paradigm.We use Burst Information Networks as media to represent text streams and present a simple yet effective network decipherment algorithm with diverse clues to decipher the networks for accurate text stream alignment.Experiments on Chinese-English news streams show our approach not only outperforms previous approaches on bilingual lexicon extraction from coordinated text streams but also can harvest high-quality alignments from large amounts of streaming data for endless language knowledge mining, which makes it promising to be a new paradigm for automatic language knowledge acquisition. Tao Ge 0001, Qing Dou, Heng Ji 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
EMNLP | 8 |
| 2018 | Question Generation from SQL Queries Improves Neural Semantic ParsingabstractWe study how to learn a semantic parser of state-of-the-art accuracy with less supervised training data.We conduct our study on WikiSQL, the largest hand-annotated semantic parsing dataset to date.First, we demonstrate that question generation is an effective method that empowers us to learn a state-ofthe-art neural network based semantic parser with thirty percent of the supervised training data.Second, we show that applying question generation to the full supervised training data further improves the state-of-the-art model.In addition, we observe that there is a logarithmic relationship between the accuracy of a semantic parser and the amount of training data. Daya Guo, Duyu Tang, Nan Duan 0001, Jian Yin 0001, Hong Chi, James Cao, Peng Chen 0029, Ming Zhou 0001 |
EMNLP | 9 |
| 2018 | Attention-Guided Answer Distillation for Machine Reading ComprehensionabstractDespite that current reading comprehension systems have achieved significant advancements, their promising performances are often obtained at the cost of making an ensemble of numerous models. Besides, existing approaches are also vulnerable to adversarial attacks. This paper tackles these problems by leveraging knowledge distillation, which aims to transfer knowledge from an ensemble model to a single model. We first demonstrate that vanilla knowledge distillation applied to answer span prediction is effective for reading comprehension systems. We then propose two novel approaches that not only penalize the prediction on confusing answers but also guide the training with alignment information distilled from the ensemble. Experiments show that our best student model has only a slight drop of 0.4% F1 on the SQuAD test set compared to the ensemble teacher, while running 12x faster during inference. It even outperforms the teacher on adversarial SQuAD datasets and NarrativeQA benchmark. Yuxing Peng 0001, Furu Wei, Zhen Huang 0006, Dongsheng Li 0001, Nan Yang 0002, Ming Zhou 0001 |
EMNLP | 7 |
| 2018 | Neural Latent Extractive Document SummarizationabstractExtractive summarization models require sentence-level labels, which are usually created heuristically (e.g., with rule-based methods) given that most summarization datasets only have document-summary pairs.Since these labels might be suboptimal, we propose a latent variable extractive model where sentences are viewed as latent variables and sentences with activated variables are used to infer gold summaries.During training the loss comes directly from gold summaries.Experiments on the CNN/Dailymail dataset show that our model improves over a strong extractive baseline trained on heuristically approximated labels and also performs competitively to several recent models. Xingxing Zhang 0002, Mirella Lapata, Furu Wei, Ming Zhou 0001 |
EMNLP | 4 |
| 2018 | Reinforced Mnemonic Reader for Machine Reading ComprehensionabstractIn this paper, we introduce the Reinforced Mnemonic Reader for machine reading comprehension tasks, which enhances previous attentive readers in two aspects. First, a reattention mechanism is proposed to refine current attentions by directly accessing to past attentions that are temporally memorized in a multi-round alignment architecture, so as to avoid the problems of attention redundancy and attention deficiency. Second, a new optimization approach, called dynamic-critical reinforcement learning, is introduced to extend the standard supervised method. It always encourages to predict a more acceptable answer so as to address the convergence suppression problem occurred in traditional reinforcement learning algorithms. Extensive experiments on the Stanford Question Answering Dataset (SQuAD) show that our model achieves state-of-the-art results. Meanwhile, our model outperforms previous systems by over 6% in terms of both Exact Match and F1 metrics on two adversarial SQuAD datasets. Yuxing Peng 0001, Zhen Huang 0006, Xipeng Qiu, Furu Wei, Ming Zhou 0001 |
IJCAI | 6 |
| 2018 | Multiway Attention Networks for Modeling Sentence PairsabstractModeling sentence pairs plays the vital role for judging the relationship between two sentences, such as paraphrase identification, natural language inference, and answer sentence selection. Previous work achieves very promising results using neural networks with attention mechanism. In this paper, we propose the multiway attention networks which employ multiple attention functions to match sentence pairs under the matching-aggregation framework. Specifically, we design four attention functions to match words in corresponding sentences. Then, we aggregate the matching information from each function, and combine the information from all functions to obtain the final representation. Experimental results demonstrate that the proposed multiway attention networks improve the result on the Quora Question Pairs, SNLI, MultiNLI, and answer sentence selection task on the SQuAD dataset. Chuanqi Tan, Furu Wei, Wenhui Wang 0003, Weifeng Lv, Ming Zhou 0001 |
IJCAI | 5 |
| 2018 | R-VQA: Learning Visual Relation Facts with Semantic Attention for Visual Question AnsweringabstractRecently, Visual Question Answering (VQA) has emerged as one of the most significant tasks in multimodal learning as it requires understanding both visual and textual modalities. Existing methods mainly rely on extracting image and question features to learn their joint feature embedding via multimodal fusion or attention mechanism. Some recent studies utilize external VQA-independent models to detect candidate entities or attributes in images, which serve as semantic knowledge complementary to the VQA task. However, these candidate entities or attributes might be unrelated to the VQA task and have limited semantic capacities. To better utilize semantic knowledge in images, we propose a novel framework to learn visual relation facts for VQA. Specifically, we build up a Relation-VQA (R-VQA) dataset based on the Visual Genome dataset via a semantic similarity module, in which each data consists of an image, a corresponding question, a correct answer and a supporting relation fact. A well-defined relation detector is then adopted to predict visual question-related relation facts. We further propose a multi-step attention model composed of visual attention and semantic attention sequentially to extract related visual knowledge and semantic knowledge. We conduct comprehensive experiments on the two benchmark datasets, demonstrating that our model achieves state-of-the-art performance and verifying the benefit of considering visual relation facts. Pan Lu, Lei Ji 0001, Wei Zhang 0056, Nan Duan 0001, Ming Zhou 0001, Jianyong Wang 0001 |
KDD | 5 |
| 2018 | EventWiki: A Knowledge Base of Major Events
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
LREC | 6 |
| 2018 | Generative Bridging Network for Neural Sequence PredictionabstractWenhu Chen, Guanlin Li, Shuo Ren, Shujie Liu, Zhirui Zhang, Mu Li, Ming Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Wenhu Chen, Shuo Ren 0002, Shujie Liu 0001, Zhirui Zhang, Mu Li 0001, Ming Zhou 0001 |
NAACL-HLT | 7 |
| 2018 | Learning to Collaborate for Question Answering and AskingabstractDuyu Tang, Nan Duan, Zhao Yan, Zhirui Zhang, Yibo Sun, Shujie Liu, Yuanhua Lv, Ming Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Duyu Tang, Nan Duan 0001, Zhirui Zhang, Shujie Liu 0001, Yuanhua Lv, Ming Zhou 0001 |
NAACL-HLT | 8 |
| 2018 | Dialog-to-Action: Conversational Question Answering Over a Large-Scale Knowledge BaseabstractWe present an approach to map utterances in conversation to logical forms, which will be executed on a large-scale knowledge base. To handle enormous ellipsis phenomena in conversation, we introduce dialog memory management to manipulate historical entities, predicates, and logical forms when inferring the logical form of current utterances. Dialog memory management is embodied in a generative model, in which a logical form is interpreted in a top-down manner following a small and flexible grammar. We learn the model from denotations without explicit annotation of logical forms, and evaluate it on a large-scale dataset consisting of 200K dialogs over 12.8M entities. Results verify the benefits of modeling dialog memory, and show that our semantic parsing-based approach outperforms a memory network based encoder-decoder model by a huge margin. Daya Guo, Duyu Tang, Nan Duan 0001, Ming Zhou 0001, Jian Yin 0001 |
NeurIPS | 4 |
| 2018 | SeRI: A Dataset for Sub-event Relation Inference from an Encyclopedia
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
NLPCC (2) | 6 |
| 2018 | I Know There Is No Answer: Modeling Answer Validation for Machine Reading Comprehension
Chuanqi Tan, Furu Wei, Qingyu Zhou, Nan Yang 0002, Weifeng Lv, Ming Zhou 0001 |
NLPCC (1) | 6 |
| 2018 | Improved Neural Machine Translation with Chinese Phonologic Features
Jian Yang 0030, Shuangzhi Wu, Dongdong Zhang 0001, Zhoujun Li 0001, Ming Zhou 0001 |
NLPCC (1) | 5 |
| 2018 | Coarse-To-Fine Learning for Neural Machine Translation
Zhirui Zhang, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Enhong Chen |
NLPCC (1) | 4 |
| 2018 | Response selection with topic clues for retrieval-based chatbots
Yu Wu 0012, Zhoujun Li 0001, Wei Wu 0014, Ming Zhou 0001 |
Neurocomputing | 4 |
| 2018 | Response selection from unstructured documents for human-computer conversation systems
Nan Duan 0001, Junwei Bao 0001, Peng Chen 0029, Ming Zhou 0001, Zhoujun Li 0001 |
Knowl. Based Syst. | 5 |
| 2018 | Question Generation With Doubly Adversarial NetsabstractWe study the problem of question generation on a specific domain, where there are no labeled data. To address this problem, we propose a novel neural question generation approach called DoubAN, or doubly adversarial nets, which fully utilizes labeled data from other domains (source domains) and unlabeled data from the target domain. Learning a DoubAN involves two adversarial procedures between a question generator and two adversaries. One adversary is a domain-classification discriminator (DC-Dis), which is designed to help the generator learn domain-general representations of the input text. The other is a question-answering discriminator (QA-Dis), which provides more training data with estimated reward scores for generated text-question pairs. We conduct experiments on the SQuAD dataset as target-domain unlabeled data and the NewsQA dataset as source-domain labeled data. Experiment results show that our DoubAN achieves better results than baselines. Compared to model variants, which adopt only DC-Dis or QA-Dis, we find that the DC-Dis and QA-Dis indirectly interact with each other and jointly improve the quality of generated questions on the target domain. Moreover, extensive analysis and discussion prove the reasonableness and effectiveness of our proposed approach. Junwei Bao 0001, Yeyun Gong, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Context-Aware Answer Sentence Selection With Hierarchical Gated Recurrent Neural NetworksabstractIn this paper, we study the task of reading comprehension style answer sentence selection that aims to select the best sentence from a given passage to answer a question. Unlike most previous works that match the question and each candidate sentence separately, we observe that the context information among sentences in the same passage plays a vital role in this task. We propose modeling context information with hierarchical gated recurrent neural networks. Specifically, we first apply a word level recurrent neural network to model the context independent matching between the question and each candidate sentence. We then employ a sentence level recurrent neural network to incorporate the context information among all candidate sentences. Moreover, we introduce the gate mechanism to select matching information before feeding into recurrent neural networks at both word and sentence level. Experiments on the WikiQA and SQuAD datasets show that our model outperforms state-of-the-art methods. Chuanqi Tan, Furu Wei, Qingyu Zhou, Nan Yang 0002, Bowen Du 0001, Weifeng Lv, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2018 | Dependency-to-Dependency Neural Machine TranslationabstractRecent research has proven that syntactic knowledge is effective to improve the performance of neural machine translation (NMT). Most previous work focuses on leveraging either source or target syntax in the recurrent neural network (RNN) based encoder–decoder model. In this paper, we simultaneously use both source and target dependency tree to improve the NMT model. First, we propose a simple but effective syntax-aware encoder to incorporate source dependency tree into NMT. The new encoder enriches each source state with dependence relations from the tree. Then, we propose a novel sequence-to-dependence framework. In this framework, the target translation and its corresponding dependence tree are jointly constructed and modeled. During decoding, the tree structure is used as context to facilitate word generations. Finally, we extend the sequence-to-dependence framework with the syntax-aware encoder to build a dependence-NMT model and apply the dependence-based framework to the Transformer. Experimental results on several translation tasks show that both source and target dependence structures can improve the translation quality and their effects can be accumulated. Shuangzhi Wu, Dongdong Zhang 0001, Zhirui Zhang, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2017 | Topic Aware Neural Response GenerationabstractWe consider incorporating topic information into a sequence-to-sequence framework to generate informative and interesting responses for chatbots. To this end, we propose a topic aware sequence-to-sequence (TA-Seq2Seq) model. The model utilizes topics to simulate prior human knowledge that guides them to form informative and interesting responses in conversation, and leverages topic information in generation by a joint attention mechanism and a biased generation probability. The joint attention mechanism summarizes the hidden vectors of an input message as context vectors by message attention and synthesizes topic vectors by topic attention from the topic words of the message obtained from a pre-trained LDA model, with these vectors jointly affecting the generation of words in decoding. To increase the possibility of topic words appearing in responses, the model modifies the generation probability of topic words by adding an extra probability item to bias the overall distribution. Empirical studies on both automatic evaluation metrics and human annotations show that TA-Seq2Seq can generate more informative and interesting responses, significantly outperforming state-of-the-art response generation models. Chen Xing, Wei Wu 0014, Yu Wu 0012, Jie Liu 0007, Yalou Huang, Ming Zhou 0001, Wei-Ying Ma |
AAAI | 6 |
| 2017 | Building Task-Oriented Dialogue Systems for Online ShoppingabstractWe present a general solution towards building task-oriented dialogue systems for online shopping, aiming to assist online customers in completing various purchase-related tasks, such as searching products and answering questions, in a natural language conversation manner. As a pioneering work, we show what & how existing NLP techniques, data resources, and crowdsourcing can be leveraged to build such task-oriented dialogue systems for E-commerce usage. To demonstrate its effectiveness, we integrate our system into a mobile online shopping app. To the best of our knowledge, this is the first time that an AI bot in Chinese is practically used in online shopping scenario with millions of real consumers. Interesting and insightful observations are shown in the experimental part, based on the analysis of human-bot conversation log. Several current challenges are also pointed out as our future directions. Nan Duan 0001, Peng Chen 0029, Ming Zhou 0001, Jianshe Zhou, Zhoujun Li 0001 |
AAAI | 4 |
| 2017 | Chunk-based Decoder for Neural Machine TranslationabstractShonosuke Ishiwatari, Jingtao Yao, Shujie Liu, Mu Li, Ming Zhou, Naoki Yoshinaga, Masaru Kitsuregawa, Weijia Jia. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Shonosuke Ishiwatari, JingTao Yao 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Naoki Yoshinaga 0001, Masaru Kitsuregawa, Weijia Jia 0001 |
ACL (1) | 5 |
| 2017 | Gated Self-Matching Networks for Reading Comprehension and Question AnsweringabstractIn this paper, we present the gated selfmatching networks for reading comprehension style question answering, which aims to answer questions from a given passage.We first match the question and passage with gated attention-based recurrent networks to obtain the question-aware passage representation.Then we propose a self-matching attention mechanism to refine the representation by matching the passage against itself, which effectively encodes information from the whole passage.We finally employ the pointer networks to locate the positions of answers from the passages.We conduct extensive experiments on the SQuAD dataset.The single model achieves 71.3% on the evaluation metrics of exact match on the hidden test set, while the ensemble model further boosts the results to 75.9%.At the time of submission of the paper, our model holds the first place on the SQuAD leaderboard for both single and ensemble model. Wenhui Wang 0003, Nan Yang 0002, Furu Wei, Baobao Chang, Ming Zhou 0001 |
ACL (1) | 5 |
| 2017 | Sequential Matching Network: A New Architecture for Multi-turn Response Selection in Retrieval-Based ChatbotsabstractWe study response selection for multiturn conversation in retrieval-based chatbots.Existing work either concatenates utterances in context or matches a response with a highly abstract context vector finally, which may lose relationships among utterances or important contextual information.We propose a sequential matching network (SMN) to address both problems.SMN first matches a response with each utterance in the context on multiple levels of granularity, and distills important matching information from each pair as a vector with convolution and pooling operations.The vectors are then accumulated in a chronological order through a recurrent neural network (RNN) which models relationships among utterances.The final matching score is calculated with the hidden states of the RNN.An empirical study on two public data sets shows that SMN can significantly outperform stateof-the-art methods for response selection in multi-turn conversation. Yu Wu 0012, Wei Wu 0014, Chen Xing, Ming Zhou 0001, Zhoujun Li 0001 |
ACL (1) | 4 |
| 2017 | Sequence-to-Dependency Neural Machine TranslationabstractNowadays a typical Neural Machine Translation (NMT) model generates translations from left to right as a linear sequence, during which latent syntactic structures of the target sentences are not explicitly concerned.Inspired by the success of using syntactic knowledge of target language for improving statistical machine translation, in this paper we propose a novel Sequence-to-Dependency Neural Machine Translation (SD-NMT) method, in which the target word sequence and its corresponding dependency structure are jointly constructed and modeled, and this structure is used as context to facilitate word generations.Experimental results show that the proposed method significantly outperforms state-of-the-art baselines on Chinese-English and Japanese-English translation tasks. Shuangzhi Wu, Dongdong Zhang 0001, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001 |
ACL (1) | 5 |
| 2017 | Selective Encoding for Abstractive Sentence SummarizationabstractWe propose a selective encoding model to extend the sequence-to-sequence framework for abstractive sentence summarization.It consists of a sentence encoder, a selective gate network, and an attention equipped decoder.The sentence encoder and decoder are built with recurrent neural networks.The selective gate network constructs a second level sentence representation by controlling the information flow from encoder to decoder.The second level representation is tailored for sentence summarization task, which leads to better performance.We evaluate our model on the English Gigaword, DUC 2004 and MSR abstractive sentence summarization datasets.The experimental results show that the proposed selective encoding model outperforms the state-ofthe-art baseline models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Ming Zhou 0001 |
ACL (1) | 4 |
| 2017 | Learning to Generate Product Reviews from AttributesabstractLi Dong, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou, Ke Xu. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Li Dong 0004, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou 0001, Ke Xu 0001 |
EACL (1) | 5 |
| 2017 | Question Generation for Question AnsweringabstractThis paper presents how to generate questions from given passages using neural networks, where large scale QA pairs are automatically crawled and processed from Community-QA website, and used as training data.The contribution of the paper is 2-fold: First, two types of question generation approaches are proposed, one is a retrieval-based method using convolution neural network (CNN), the other is a generation-based method using recurrent neural network (RNN); Second, we show how to leverage the generated questions to improve existing question answering systems.We evaluate our question generation method for the answer sentence selection task on three benchmark datasets, including SQuAD, MS MARCO, and WikiQA.Experimental results show that, by using generated questions as an extra signal, significant QA improvement can be achieved. Nan Duan 0001, Duyu Tang, Peng Chen 0029, Ming Zhou 0001 |
EMNLP | 4 |
| 2017 | Entity Linking for Queries by Searching Wikipedia SentencesabstractWe present a simple yet effective approach for linking entities in queries.The key idea is to search sentences similar to a query from Wikipedia articles and directly use the human-annotated entities in the similar sentences as candidate entities for the query.Then, we employ a rich set of features, such as link-probability, contextmatching, word embeddings, and relatedness among candidate entities as well as their related entities, to rank the candidates under a regression based framework.The advantages of our approach lie in two aspects, which contribute to the ranking process and final linking result.First, it can greatly reduce the number of candidate entities by filtering out irrelevant entities with the words in the query.Second, we can obtain the query sensitive prior probability in addition to the static linkprobability derived from all Wikipedia articles.We conduct experiments on two benchmark datasets on entity linking for queries, namely the ERD14 dataset and the GERDAQ dataset.Experimental results show that our method outperforms state-of-the-art systems and yields 75.0% in F1 on the ERD14 dataset and 56.9% on the GERDAQ dataset. Chuanqi Tan, Furu Wei, Pengjie Ren, Weifeng Lv, Ming Zhou 0001 |
EMNLP | 5 |
| 2017 | Stack-based Multi-layer Attention for Transition-based Dependency ParsingabstractAlthough sequence-to-sequence (seq2seq) network has achieved significant success in many NLP tasks such as machine translation and text summarization, simply applying this approach to transition-based dependency parsing cannot yield a comparable performance gain as in other stateof-the-art methods, such as stack-LSTM and head selection.In this paper, we propose a stack-based multi-layer attention model for seq2seq learning to better leverage structural linguistics information.In our method, two binary vectors are used to track the decoding stack in transition-based parsing, and multi-layer attention is introduced to capture multiple word dependencies in partial trees.We conduct experiments on PTB and CTB datasets, and the results show that our proposed model achieves state-of-the-art accuracy and significant improvement in labeled precision with respect to the baseline seq2seq model. Zhirui Zhang, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Enhong Chen |
EMNLP | 4 |
| 2017 | Improved Neural Machine Translation with Source SyntaxabstractNeural Machine Translation (NMT) based on the encoder-decoder architecture has recently achieved the state-of-the-art performance. Researchers have proven that extending word level attention to phrase level attention by incorporating source-side phrase structure can enhance the attention model and achieve promising improvement. However, word dependencies that can be crucial to correctly understand a source sentence are not always in a consecutive fashion (i.e. phrase structure), sometimes they can be in long distance. Phrase structures are not the best way to explicitly model long distance dependencies. In this paper we propose a simple but effective method to incorporate source-side long distance dependencies into NMT. Our method based on dependency trees enriches each source state with global dependency structures, which can better capture the inherent syntactic structure of source sentences. Experiments on Chinese-English and English-Japanese translation tasks show that our proposed method outperforms state-of-the-art SMT and NMT baselines. Shuangzhi Wu, Ming Zhou 0001, Dongdong Zhang 0001 |
IJCAI | 2 |
| 2017 | An Information Retrieval-Based Approach to Table-Based Question Answering
Junwei Bao 0001, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
NLPCC | 3 |
| 2017 | Modeling Indicative Context for Statistical Machine Translation
Shuangzhi Wu, Dongdong Zhang 0001, Shujie Liu 0001, Ming Zhou 0001 |
NLPCC | 4 |
| 2017 | Neural Question Generation from Text: A Preliminary Study
Qingyu Zhou, Nan Yang 0002, Furu Wei, Chuanqi Tan, Hangbo Bao, Ming Zhou 0001 |
NLPCC | 6 |
| 2017 | Named entity disambiguation for questions in community question answering
Fang Wang 0019, Wei Wu 0014, Zhoujun Li 0001, Ming Zhou 0001 |
Knowl. Based Syst. | 4 |
| 2016 | TGSum: Build Tweet Guided Multi-Document Summarization DatasetabstractThe development of summarization research has been significantly hampered by the costly acquisition of reference summaries. This paper proposes an effective way to automatically collect large scales of news-related multi-document summaries with reference to social media's reactions. We utilize two types of social labels in tweets, i.e., hashtags and hyper-links. Hashtags are used to cluster documents into different topic sets. Also, a tweet with a hyper-link often highlights certain key points of the corresponding document. We synthesize a linked document cluster to form a reference summary which can cover most key points. To this aim, we adopt the ROUGE metrics to measure the coverage ratio, and develop an Integer Linear Programming solution to discover the sentence set reaching the upper bound of ROUGE. Since we allow summary sentences to be selected from both documents and high-quality tweets, the generated reference summaries could be abstractive. Both informativeness and readability of the collected summaries are verified by manual judgment. In addition, we train a Support Vector Regression summarizer on DUC generic multi-document summarization benchmarks. With the collected data as extra training resource, the performance of the summarizer improves a lot on all the test sets. We release this dataset for further research. Ziqiang Cao, Chengyao Chen, Wenjie Li 0002, Sujian Li, Furu Wei, Ming Zhou 0001 |
AAAI | 6 |
| 2016 | Jointly Modeling Topics and Intents with Global Order StructureabstractModeling document structure is of great importance for discourse analysis and related applications. The goal of this research is to capture the document intent structure by modeling documents as a mixture of topic words and rhetorical words. While the topics are relatively unchanged through one document, the rhetorical functions of sentences usually change following certain orders in discourse. We propose GMM-LDA, a topic modeling based Bayesian unsupervised model, to analyze the document intent structure cooperated with order information. Our model is flexible that has the ability to combine the annotations and do supervised learning. Additionally, entropic regularization can be introduced to model the significant divergence between topics and intents. We perform experiments in both unsupervised and supervised settings, results show the superiority of our model over several state-of-the-art baselines. Jun Zhu 0001, Nan Yang 0002, Tian Tian 0001, Ming Zhou 0001, Bo Zhang 0010 |
AAAI | 5 |
| 2016 | Improving Recommendation of Tail Tags for Questions in Community Question AnsweringabstractWe study tag recommendation for questions in community question answering (CQA). Tags represent the semantic summarization of questions are useful for navigation and expert finding in CQA and can facilitate content consumption such as searching and mining in these web sites. The task is challenging, as both questions and tags are short and a large fraction of tags are tail tags which occur very infrequently. To solve these problems, we propose matching questions and tags not only by themselves, but also by similar questions and similar tags. The idea is then formalized as a model in which we calculate question-tag similarity using a linear combination of similarity with similar questions and tags weighted by tag importance.Question similarity, tag similarity, and tag importance are learned in a supervised random walk framework by fusing multiple features. Our model thus can not only accurately identify question-tag similarity for head tags, but also improve the accuracy of recommendation of tail tags. Experimental results show that the proposed method significantly outperforms state-of-the-art methods on tag recommendation for questions. Particularly, it improves tail tag recommendation accuracy by a large margin. Yu Wu 0012, Wei Wu 0014, Zhoujun Li 0001, Ming Zhou 0001 |
AAAI | 4 |
| 2016 | Knowledge-Based Semantic Embedding for Machine TranslationabstractIn this paper, with the help of knowledge base, we build and formulate a semantic space to connect the source and target languages, and apply it to the sequence-to-sequence framework to propose a Knowledge-Based Semantic Embedding (KBSE) method.In our KB-SE method, the source sentence is firstly mapped into a knowledge based semantic space, and the target sentence is generated using a recurrent neural network with the internal meaning preserved.Experiments are conducted on two translation tasks, the electric business data and movie data, and the results show that our proposed method can achieve outstanding performance, compared with both the traditional SMT methods and the existing encoder-decoder models. Shujie Liu 0001, Shuo Ren 0002, Mu Li 0001, Ming Zhou 0001, Xu Sun 0001, Houfeng Wang |
ACL (1) | 6 |
| 2016 | DocChat: An Information Retrieval Approach for Chatbot Engines Using Unstructured DocumentsabstractMost current chatbot engines are designed to reply to user utterances based on existing utterance-response (or Q-R) 1 pairs.In this paper, we present DocChat, a novel information retrieval approach for chatbot engines that can leverage unstructured documents, instead of Q-R pairs, to respond to utterances.A learning to rank model with features designed at different levels of granularity is proposed to measure the relevance between utterances and responses directly.We evaluate our proposed approach in both English and Chinese: (i) For English, we evaluate Doc-Chat on WikiQA and QASent, two answer sentence selection tasks, and compare it with state-of-the-art methods.Reasonable improvements and good adaptability are observed.(ii) For Chinese, we compare DocChat with XiaoIce 2 , a famous chitchat engine in China, and side-by-side evaluation shows that DocChat is a perfect complement for chatbot engines using Q-R pairs as main source of responses. Nan Duan 0001, Junwei Bao 0001, Peng Chen 0029, Ming Zhou 0001, Zhoujun Li 0001, Jianshe Zhou |
ACL (1) | 5 |
| 2016 | Constraint-Based Question Answering with Knowledge GraphabstractWebQuestions and SimpleQuestions are two benchmark data-sets commonly used in recent knowledge-based question answering (KBQA) work. Most questions in them are ‘simple’ questions which can be answered based on a single relation in the knowledge base. Such data-sets lack the capability of evaluating KBQA systems on complicated questions. Motivated by this issue, we release a new data-set, namely ComplexQuestions, aiming to measure the quality of KBQA systems on ‘multi-constraint’ questions which require multiple knowledge base relations to get the answer. Beside, we propose a novel systematic KBQA approach to solve multi-constraint questions. Compared to state-of-the-art methods, our approach not only obtains comparable results on the two existing benchmark data-sets, but also achieves significant improvements on the ComplexQuestions. Junwei Bao 0001, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
COLING | 4 |
| 2016 | Improving Attention Modeling with Implicit Distortion and Fertility for Machine TranslationabstractIn neural machine translation, the attention mechanism facilitates the translation process by producing a soft alignment between the source sentence and the target sentence. However, without dedicated distortion and fertility models seen in traditional SMT systems, the learned alignment may not be accurate, which can lead to low translation quality. In this paper, we propose two novel models to improve attention-based neural machine translation. We propose a recurrent attention mechanism as an implicit distortion model, and a fertility conditioned decoder as an implicit fertility model. We conduct experiments on large-scale Chinese–English translation tasks. The results show that our models significantly improve both the alignment and translation quality compared to the original attention mechanism and several other variations. Shujie Liu 0001, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001, Kenny Q. Zhu |
COLING | 5 |
| 2016 | Event Detection with Burst Information NetworksabstractRetrospective event detection is an important task for discovering previously unidentified events in a text stream. In this paper, we propose two fast centroid-aware event detection models based on a novel text stream representation – Burst Information Networks (BINets) for addressing the challenge. The BINets are time-aware, efficient and can be easily analyzed for identifying key information (centroids). These advantages allow the BINet-based approaches to achieve the state-of-the-art performance on multiple datasets, demonstrating the efficacy of BINets for the task of event detection. Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Ming Zhou 0001 |
COLING | 5 |
| 2016 | Detecting Context Dependent Messages in a Conversational EnvironmentabstractWhile automatic response generation for building chatbot systems has drawn a lot of attention recently, there is limited understanding on when we need to consider the linguistic context of an input text in the generation process. The task is challenging, as messages in a conversational environment are short and informal, and evidence that can indicate a message is context dependent is scarce. After a study of social conversation data crawled from the web, we observed that some characteristics estimated from the responses of messages are discriminative for identifying context dependent messages. With the characteristics as weak supervision, we propose using a Long Short Term Memory (LSTM) network to learn a classifier. Our method carries out text representation and classifier learning in a unified framework. Experimental results show that the proposed method can significantly outperform baseline methods on accuracy of classification. Chaozhuo Li, Yu Wu 0012, Wei Wu 0014, Chen Xing, Zhoujun Li 0001, Ming Zhou 0001 |
COLING | 6 |
| 2016 | A Redundancy-Aware Sentence Regression Framework for Extractive SummarizationabstractExisting sentence regression methods for extractive summarization usually model sentence importance and redundancy in two separate processes. They first evaluate the importance f(s) of each sentence s and then select sentences to generate a summary based on both the importance scores and redundancy among sentences. In this paper, we propose to model importance and redundancy simultaneously by directly evaluating the relative importance f(s|S) of a sentence s given a set of selected sentences S. Specifically, we present a new framework to conduct regression with respect to the relative gain of s given S calculated by the ROUGE metric. Besides the single sentence features, additional features derived from the sentence relations are incorporated. Experiments on the DUC 2001, 2002 and 2004 multi-document summarization datasets show that the proposed method outperforms state-of-the-art extractive summarization approaches. Pengjie Ren, Furu Wei, Zhumin Chen, Jun Ma 0001, Ming Zhou 0001 |
COLING | 5 |
| 2016 | News Stream Summarization using Burst Information NetworksabstractThis paper studies summarizing key information from news streams. We propose simple yet effective models to solve the problem based on a novel and promising representation of text streams – Burst Information Networks (BINets). A BINet can be aware of redundant information, allows global analysis of a text stream, and can be efficiently built and dynamically updated, which perfectly fits the demands of text stream summarization. Extensive experiments show that the BINet-based approaches are not only efficient and can be used in a real-time online summarization setting, but also can generate high-quality summaries, outperforming the state-of-the-art approach. Tao Ge 0001, Lei Cui 0001, Baobao Chang, Sujian Li, Ming Zhou 0001, Zhifang Sui |
EMNLP | 5 |
| 2016 | Solving and Generating Chinese Character RiddlesabstractChinese character riddle is a riddle game in which the riddle solution is a single Chinese character.It is closely connected with the shape, pronunciation or meaning of Chinese characters.The riddle description (sentence) is usually composed of phrases with rich linguistic phenomena (such as pun, simile, and metaphor), which are associated to different parts (namely radicals) of the solution character.In this paper, we propose a statistical framework to solve and generate Chinese character riddles.Specifically, we learn the alignments and rules to identify the metaphors between phrases in riddles and radicals in characters.Then, in the solving phase, we utilize a dynamic programming method to combine the identified metaphors to obtain candidate solutions.In the riddle generation phase, we use a template-based method and a replacement-based method to obtain candidate riddle descriptions.We then use Ranking SVM to rerank the candidates both in the solving and generation process.Experimental results in the solving task show that the proposed method outperforms baseline methods.We also get very promising results in the generation task according to human judges. Chuanqi Tan, Furu Wei, Li Dong 0004, Weifeng Lv, Ming Zhou 0001 |
EMNLP | 5 |
| 2016 | Unsupervised Word and Dependency Path Embeddings for Aspect Term Extraction
Yichun Yin, Furu Wei, Li Dong 0004, Kaimeng Xu, Ming Zhang 0004, Ming Zhou 0001 |
IJCAI | 6 |
| 2016 | Learning Distributed Representations of Data in Community Question Answering for Question RetrievalabstractWe study the problem of question retrieval in community question answering (CQA). The biggest challenge within this task is lexical gaps between questions since similar questions are usually expressed with different but semantically related words. To bridge the gaps, state-of-the-art methods incorporate extra information such as word-to-word translation and categories of questions into the traditional language models. We find that the existing language model based methods can be interpreted using a new framework, that is they represent words and question categories in a vector space and calculate question-question similarities with a linear combination of dot products of the vectors. The problem is that these methods are either heuristic on data representation or difficult to scale up. We propose a principled and efficient approach to learning representations of data in CQA. In our method, we simultaneously learn vectors of words and vectors of question categories by optimizing an objective function naturally derived from the framework. In question retrieval, we incorporate learnt representations into traditional language models in an effective and efficient way. We conduct experiments on large scale data from Yahoo! Answers and Baidu Knows, and compared our method with state-of-the-art methods on two public data sets. Experimental results show that our method can significantly improve on baseline methods for retrieval relevance. On 1 million training data, our method takes less than 50 minutes to learn a model on a single multicore machine, while the translation based language model needs more than 2 days to learn a translation table on the same machine. Kai Zhang 0038, Wei Wu 0014, Fang Wang 0019, Ming Zhou 0001, Zhoujun Li 0001 |
WSDM | 4 |
| 2016 | Adaptive Multi-Compositionality for Recursive Neural Network ModelsabstractRecursive neural network models have achieved promising results in many natural language processing tasks. The main difference among these models lies in the composition function, i.e., how to obtain the vector representation for a phrase or sentence using the representations of words it contains. This paper introduces a novel Adaptive Multi-Compositionality (AdaMC) layer to recursive neural network models. The basic idea is to use more than one composition function and adaptively select them depending on input vectors. We develop a general framework to model the semantic composition as a distribution of these composition functions. The composition functions and parameters used for adaptive selection are jointly learnt from the supervision of specific tasks. We integrate AdaMC into existing recursive neural network models and conduct extensive experiments on the Stanford Sentiment Treebank and semantic relation classification task. The experimental results demonstrate that AdaMC improves the performance of recursive neural network models and outperforms the baseline methods. Li Dong 0004, Furu Wei, Ke Xu 0001, Shixia Liu, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | Sentiment Embeddings with Applications to Sentiment AnalysisabstractWe propose learning sentiment-specific word embeddings dubbed sentiment embeddings in this paper. Existing word embedding learning algorithms typically only use the contexts of words but ignore the sentiment of texts. It is problematic for sentiment analysis because the words with similar contexts but opposite sentiment polarity, such asgoodandbad, are mapped to neighboring word vectors. We address this issue by encoding sentiment information of texts (e.g., sentences and words) together with contexts of words in sentiment embeddings. By combining context and sentiment level evidences, the nearest neighbors in sentiment embedding space are semantically similar and it favors words with the same sentiment polarity. In order to learn sentiment embeddings effectively, we develop a number of neural networks with tailoring loss functions, and collect massive texts automatically with sentiment signals like emoticons as the training data. Sentiment embeddings can be naturally used as word features for a variety of sentiment analysis tasks without feature engineering. We apply sentiment embeddings to word-level sentiment analysis, sentence level sentiment classification, and building sentiment lexicons. Experimental results show that sentiment embeddings consistently outperform context-based embeddings on several benchmark datasets of these tasks. This work provides insights on the design of neural networks for learning task-specific word embeddings in other natural language processing tasks. Duyu Tang, Furu Wei, Bing Qin 0001, Nan Yang 0002, Ting Liu 0001, Ming Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2015 | Ranking with Recursive Neural Networks and Its Application to Multi-Document SummarizationabstractWe develop a Ranking framework upon Recursive Neural Networks (R2N2) to rank sentences for multi-document summarization. It formulates the sentence ranking task as a hierarchical regression process, which simultaneously measures the salience of a sentence and its constituents (e.g., phrases) in the parsing tree. This enables us to draw on word-level to sentence-level supervisions derived from reference summaries.In addition, recursive neural networks are used to automatically learn ranking features over the tree, with hand-crafted feature vectors of words as inputs. Hierarchical regressions are then conducted with learned features concatenating raw features.Ranking scores of sentences and words are utilized to effectively select informative and non-redundant sentences to generate summaries.Experiments on the DUC 2001, 2002 and 2004 multi-document summarization datasets show that R2N2 outperforms state-of-the-art extractive summarization approaches. Ziqiang Cao, Furu Wei, Li Dong 0004, Sujian Li, Ming Zhou 0001 |
AAAI | 5 |
| 2015 | Mining Query Subtopics from Questions in Community Question AnsweringabstractThis paper proposes mining query subtopics from questions in community question answering (CQA). The subtopics are represented as a number of clusters of questions with keywords summarizing the clusters. The task is unique in that the subtopics from questions can not only facilitate user browsing in CQA search, but also describe aspects of queries from a question-answering perspective. The challenges of the task include how to group semantically similar questions and how to find keywords capable of summarizing the clusters. We formulate the subtopic mining task as a non-negative matrix factorization (NMF) problem and further extend the model of NMF to incorporate question similarity estimated from metadata of CQA into learning. Compared with existing methods, our method can jointly optimize question clustering and keyword extraction and encourage the former task to enhance the latter. Experimental results on large scale real world CQA datasets show that the proposed method significantly outperforms the existing methods in terms of keyword extraction, while achieving a comparable performance to the state-of-the-art methods for question clustering. Yu Wu 0012, Wei Wu 0014, Zhoujun Li 0001, Ming Zhou 0001 |
AAAI | 4 |
| 2015 | Question Answering over Freebase with Multi-Column Convolutional Neural NetworksabstractLi Dong, Furu Wei, Ming Zhou, Ke Xu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Li Dong 0004, Furu Wei, Ming Zhou 0001, Ke Xu 0001 |
ACL (1) | 3 |
| 2015 | Efficient Disfluency Detection with Transition-based ParsingabstractShuangzhi Wu, Dongdong Zhang, Ming Zhou, Tiejun Zhao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Shuangzhi Wu, Dongdong Zhang 0001, Ming Zhou 0001, Tiejun Zhao |
ACL (1) | 3 |
| 2015 | Answering Questions with Complex Semantic Constraints on Open Knowledge BasesabstractA knowledge-based question-answering system (KB-QA) is one that answers natural language questions with information stored in a large-scale knowledge base (KB). Existing KB-QA systems are either powered by curated KBs in which factual knowledge is encoded in entities and relations with well-structured schemas, or by open KBs, which contain assertions represented in the form of triples (e.g., subject; relation phrase; argument). We show that both approaches fall short in answering questions with complex prepositional or adverbial constraints. We propose using n-tuple assertions, which are assertions with an arbitrary number of arguments, and n-tuple open KB (nOKB), which is an open knowledge base of n-tuple assertions. We present TAQA, a novel KB-QA system that is based on an nOKB and illustrate via experiments how TAQA can effectively answer complex questions with rich semantic constraints. Our work also results in a new open KB containing 120M n-tuple assertions and a collection of 300 labeled complex questions, which is made publicly available for further research. Nan Duan 0001, Ben Kao, Junwei Bao 0001, Ming Zhou 0001 |
CIKM | 5 |
| 2015 | Cold-Start Expert Finding in Community Question Answering via Graph Regularization
Zhou Zhao 0001, Furu Wei, Ming Zhou 0001, Wilfred Ng |
DASFAA (1) | 3 |
| 2015 | Crowd-Selection Query Processing in Crowdsourcing Databases: A Task-Driven ApproachabstractCrowd-selection is essential to crowdsourcing applications, since choosing the right workers with particular expertise to carry out specific crowdsourced tasks is extremely important. The central problem is simple but tricky: given a crowdsourced task, who is the right worker to ask? Currently, most existing work has mainly studied the problem of crowd-selection for simple crowdsourced tasks such as decision making and sentiment analysis. Their crowd-selection procedures are based on the trustworthiness of workers. However, for some complex tasks such as document review and question answering, selecting workers based on the latent category of tasks is a better solution. In this paper, we formulate a new problem of task-driven crowd-selection for complex tasks. We first develop a Bayesian generative model to exploit "who knows what" for the workers in the crowdsourcing environment. The model provides a principle and natural framework for capturing the latent skills of workers as well as the latent categories of crowdsourced tasks. The inference of the latent skills of workers is based on past resolved crowdsourced tasks with feedback scores. We assume that the feedback scores can illustrate the performance of the workers for the tasks. We then devise a variational algorithm that transforms the latent skill inference with the proposed model into a standard optimization problem, which can be solved efficiently. We verify the performance of our method through extensive experiments on the data collected from three well-known crowdsourcing platforms for question answering tasks such as Quora, Yahoo! Answer and Stack Overflow. Zhou Zhao 0001, Furu Wei, Ming Zhou 0001, Weikeng Chen, Wilfred Ng |
EDBT | 3 |
| 2015 | Hierarchical Recurrent Neural Network for Document ModelingabstractThis paper proposes a novel hierarchical recurrent neural network language model (HRNNLM) for document modeling.After establishing a RNN to capture the coherence between sentences in a document, HRNNLM integrates it as the sentence history information into the word level RNN to predict the word sequence with cross-sentence contextual information.A two-step training approach is designed, in which sentence-level and word-level language models are approximated for the convergence in a pipeline style.Examined by the standard sentence reordering scenario, HRNNLM is proved for its better accuracy in modeling the sentence coherence.And at the word level, experimental results also indicate a significant lower model perplexity, followed by a practical better translation result when applied to a Chinese-English document translation reranking task. Shujie Liu 0001, Muyun Yang, Mu Li 0001, Ming Zhou 0001, Sheng Li 0003 |
EMNLP | 5 |
| 2015 | A Hybrid Neural Model for Type Classification of Entity Mentions
Li Dong 0004, Furu Wei, Ming Zhou 0001, Ke Xu 0001 |
IJCAI | 4 |
| 2015 | Entity Translation with Collective Inference in Knowledge Graph
Qinglin Li, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001 |
NLPCC | 5 |
| 2015 | A Statistical Parsing Framework for Sentiment ClassificationabstractWe present a statistical parsing framework for sentence-level sentiment classification in this article. Unlike previous works that use syntactic parsing results for sentiment analysis, we develop a statistical parser to directly analyze the sentiment structure of a sentence. We show that complicated phenomena in sentiment analysis (e.g., negation, intensification, and contrast) can be handled the same way as simple and straightforward sentiment expressions in a unified and probabilistic way. We formulate the sentiment grammar upon Context-Free Grammars (CFGs), and provide a formal description of the sentiment parsing framework. We develop the parsing model to obtain possible sentiment parse trees for a sentence, from which the polarity model is proposed to derive the sentiment strength and polarity, and the ranking model is dedicated to selecting the best sentiment tree. We train the parser directly from examples of sentences annotated only with sentiment polarity labels but without any syntactic annotations or polarity annotations of constituents within sentences. Therefore we can obtain training data easily. In particular, we train a sentiment parser, s.parser, from a large amount of review sentences with users' ratings as rough sentiment polarity labels. Extensive experiments on existing benchmark data sets show significant improvements over baseline sentiment classification approaches. Li Dong 0004, Furu Wei, Shujie Liu 0001, Ming Zhou 0001, Ke Xu 0001 |
Comput. Linguistics | 4 |
| 2015 | Cross-lingual Sentiment Lexicon Learning With Bilingual Word Graph Label PropagationabstractIn this article we address the task of cross-lingual sentiment lexicon learning, which aims to automatically generate sentiment lexicons for the target languages with available English sentiment lexicons. We formalize the task as a learning problem on a bilingual word graph, in which the intra-language relations among the words in the same language and the inter-language relations among the words between different languages are properly represented. With the words in the English sentiment lexicon as seeds, we propose a bilingual word graph label propagation approach to induce sentiment polarities of the unlabeled words in the target language. Particularly, we show that both synonym and antonym word relations can be used to build the intra-language relation, and that the word alignment information derived from bilingual parallel sentences can be effectively leveraged to build the inter-language relation. The evaluation of Chinese sentiment lexicon learning shows that the proposed approach outperforms existing approaches in both precision and recall. Experiments conducted on the NTCIR data set further demonstrate the effectiveness of the learned sentiment lexicon in sentence-level sentiment classification. Dehong Gao, Furu Wei, Wenjie Li 0002, Ming Zhou 0001 |
Comput. Linguistics | 5 |
| 2015 | Towards Machine Translation in Semantic Vector SpaceabstractMeasuring the quality of the translation rules and their composition is an essential issue in the conventional statistical machine translation (SMT) framework. To express the translation quality, the previous lexical and phrasal probabilities are calculated only according to the co-occurrence statistics in the bilingual corpus and may be not reliable due to the data sparseness problem. To address this issue, we propose measuring the quality of the translation rules and their composition in the semantic vector embedding space (VES). We present a recursive neural network (RNN)-based translation framework, which includes two submodels. One is the bilingually-constrained recursive auto-encoder, which is proposed to convert the lexical translation rules into compact real-valued vectors in the semantic VES. The other is a type-dependent recursive neural network, which is proposed to perform the decoding process by minimizing the semantic gap (meaning distance) between the source language string and its translation candidates at each state in a bottom-up structure. The RNN-based translation model is trained using a max-margin objective function that maximizes the margin between the reference translation and the n-best translations in forced decoding. In the experiments, we first show that the proposed vector representations for the translation rules are very reliable for application in translation modeling. We further show that the proposed type-dependent, RNN-based model can significantly improve the translation quality in the large-scale, end-to-end Chinese-to-English translation evaluation. Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2015 | A Joint Segmentation and Classification Framework for Sentence Level Sentiment ClassificationabstractIn this paper, we propose a joint segmentation and classification framework for sentence-level sentiment classification. It is widely recognized that phrasal information is crucial for sentiment classification. However, existing sentiment classification algorithms typically split a sentence as a word sequence, which does not effectively handle the inconsistent sentiment polarity between a phrase and the words it contains, such as {“not bad,” “bad”} and {“a great deal of,” “great”}. We address this issue by developing a joint framework for sentence-level sentiment classification. It simultaneously generates useful segmentations and predicts sentence-level polarity based on the segmentation results. Specifically, we develop a candidate generation model to produce segmentation candidates of a sentence; a segmentation ranking model to score the usefulness of a segmentation candidate for sentiment classification; and a classification model for predicting the sentiment polarity of a segmentation. We train the joint framework directly from sentences annotated with only sentiment polarity, without using any syntactic or sentiment annotations in segmentation level. We conduct experiments for sentiment classification on two benchmark datasets: a tweet dataset and a review dataset. Experimental results show that: 1) our method performs comparably with state-of-the-art methods on both datasets; 2) joint modeling segmentation and classification outperforms pipelined baseline methods in various experimental settings. Duyu Tang, Bing Qin 0001, Furu Wei, Li Dong 0004, Ting Liu 0001, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2015 | Can You Trust Online Ratings? A Mutual Reinforcement Model for Trustworthy Online Rating SystemsabstractThe average of customer ratings on a product, which we call a reputation, is one of the key factors in online purchasing decisions. There is, however, no guarantee of the trustworthiness of a reputation since it can be manipulated rather easily. In this paper, we define false reputation as the problem of a reputation being manipulated by unfair ratings and design a general framework that provides trustworthy reputations. For this purpose, we propose TRUE-REPUTATION, an algorithm that iteratively adjusts a reputation based on the confidence of customer ratings. We also show the effectiveness of TRUE-REPUTATION through extensive experiments in comparisons to state-of-the-art approaches. Hyun-Kyo Oh, Sang-Wook Kim, Sunju Park, Ming Zhou 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2014 | Machine Translation with Real-Time Web Search
Lei Cui 0001, Ming Zhou 0001, Dongdong Zhang 0001, Mu Li 0001 |
AAAI | 2 |
| 2014 | Adaptive Multi-Compositionality for Recursive Neural Models with Applications to Sentiment AnalysisabstractRecursive neural models have achieved promising results in many natural language processing tasks. The main difference among these models lies in the composition function, i.e., how to obtain the vector representation for a phrase or sentence using the representations of words it contains. This paper introduces a novel Adaptive Multi-Compositionality (AdaMC) layer to recursive neural models. The basic idea is to use more than one composition functions and adaptively select them depending on the input vectors. We present a general framework to model each semantic composition as a distribution over these composition functions. The composition functions and parameters used for adaptive selection are learned jointly from data. We integrate AdaMC into existing recursive neural models and conduct extensive experiments on the Stanford Sentiment Treebank. The results illustrate that AdaMC significantly outperforms state-of-the-art sentiment classification methods. It helps push the best accuracy of sentence-level negative/positive classification from 85.4% up to 88.5%. Li Dong 0004, Furu Wei, Ming Zhou 0001, Ke Xu 0001 |
AAAI | 3 |
| 2014 | Mind the Gap: Machine Translation by Minimizing the Semantic Gap in Embedding SpaceabstractThe conventional statistical machine translation (SMT) methods perform the decoding process by compositing a set of the translation rules which are associated with high probabilities. However, the probabilities of the translation rules are calculated only according to the cooccurrence statistics in the bilingual corpus rather than the semantic meaning similarity. In this paper, we propose a Recursive Neural Network (RNN) based model that converts each translation rule into a compact real-valued vector in the semantic embedding space and performs the decoding process by minimizing the semantic gap between the source language string and its translation candidates at each state in a bottom-up structure. The RNN-based translation model is trained using a max-margin objective function. Extensive experiments on Chinese-to-English translation show that our RNN-based model can significantly improve the translation quality by up to 1.68 BLEU score. Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong |
AAAI | 4 |
| 2014 | Knowledge-Based Question Answering as Machine TranslationabstractA typical knowledge-based question answering (KB-QA) system faces two challenges: one is to transform natural language questions into their meaning representations (MRs); the other is to retrieve answers from knowledge bases (KBs) using generated MRs.Unlike previous methods which treat them in a cascaded manner, we present a translation-based approach to solve these two tasks in one unified framework.We translate questions to answers based on CYK parsing.Answers as translations of the span covered by each CYK cell are obtained by a question translation method, which first generates formal triple queries as MRs for the span based on question patterns and relation expressions, and then retrieves answers from a given KB based on triple queries generated.A linear model is defined over derivations, and minimum error rate training is used to tune feature weights based on a set of question-answer pairs.Compared to a KB-QA system using a state-of-the-art semantic parser, our method achieves better results. Junwei Bao 0001, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
ACL (1) | 3 |
| 2014 | Learning Topic Representation for SMT with Neural NetworksabstractStatistical Machine Translation (SMT) usually utilizes contextual information to disambiguate translation candidates. However, it is often limited to contexts within sentence boundaries, hence broader topical information cannot be leveraged. In this paper, we propose a novel approach to learning topic representation for paral-lel data using a neural network architec-ture, where abundant topical contexts are embedded via topic relevant monolingual data. By associating each translation rule with the topic representation, topic rele-vant rules are selected according to the dis-tributional similarity with the source text during SMT decoding. Experimental re-sults show that our method significantly improves translation accuracy in the NIST Chinese-to-English translation task com-pared to a state-of-the-art baseline. 1 Lei Cui 0001, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Muyun Yang |
ACL (1) | 6 |
| 2014 | A Recursive Recurrent Neural Network for Statistical Machine TranslationabstractIn this paper, we propose a novel recursive recurrent neural network (R 2 NN) to model the end-to-end decoding process for statistical machine translation.R 2 NN is a combination of recursive neural network and recurrent neural network, and in turn integrates their respective capabilities: (1) new information can be used to generate the next hidden state, like recurrent neural networks, so that language model and translation model can be integrated naturally; (2) a tree structure can be built, as recursive neural networks, so as to generate the translation candidates in a bottom up manner.A semi-supervised training approach is proposed to train the parameters, and the phrase pair embedding is explored to model translation confidence directly.Experiments on a Chinese to English translation task show that our proposed R 2 NN can outperform the stateof-the-art baseline by about 1.5 points in BLEU. Shujie Liu 0001, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001 |
ACL (1) | 4 |
| 2014 | Learning Sentiment-Specific Word Embedding for Twitter Sentiment ClassificationabstractWe present a method that learns word embedding for Twitter sentiment classification in this paper.Most existing algorithms for learning continuous word representations typically only model the syntactic context of words but ignore the sentiment of text.This is problematic for sentiment analysis as they usually map words with similar syntactic context but opposite sentiment polarity, such as good and bad, to neighboring word vectors.We address this issue by learning sentimentspecific word embedding (SSWE), which encodes sentiment information in the continuous representation of words.Specifically, we develop three neural networks to effectively incorporate the supervision from sentiment polarity of text (e.g.sentences or tweets) in their loss functions.To obtain large scale training corpora, we learn the sentiment-specific word embedding from massive distant-supervised tweets collected by positive and negative emoticons.Experiments on applying SS-WE to a benchmark Twitter sentiment classification dataset in SemEval 2013 show that (1) the SSWE feature performs comparably with hand-crafted features in the top-performed system; (2) the performance is further improved by concatenating SSWE with existing feature set. Duyu Tang, Furu Wei, Nan Yang 0002, Ming Zhou 0001, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 4 |
| 2014 | Bilingually-constrained Phrase Embeddings for Machine TranslationabstractWe propose Bilingually-constrained Recursive Auto-encoders (BRAE) to learn semantic phrase embeddings (compact vector representations for phrases), which can distinguish the phrases with different semantic meanings.The BRAE is trained in a way that minimizes the semantic distance of translation equivalents and maximizes the semantic distance of nontranslation pairs simultaneously.After training, the model learns how to embed each phrase semantically in two languages and also learns how to transform semantic embedding space in one language to the other.We evaluate our proposed method on two end-to-end SMT tasks (phrase table pruning and decoding with phrasal semantic similarities) which need to measure semantic similarity between a source phrase and its translation candidates.Extensive experiments show that the BRAE is remarkably effective in these two tasks. Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong |
ACL (1) | 4 |
| 2014 | Question Retrieval with High Quality Answers in Community Question AnsweringabstractThis paper studies the problem of question retrieval in community question answering (CQA). To bridge lexical gaps in questions, which is regarded as the biggest challenge in retrieval, state-of-the-art methods learn translation models using answers under an assumption that they are parallel texts. In practice, however, questions and answers are far from "parallel". Indeed, they are heterogeneous for both the literal level and user behaviors. There are a particularly large number of low quality answers, to which the performance of translation models is vulnerable. To address these problems, we propose a supervised question-answer topic modeling approach. The approach assumes that questions and answers share some common latent topics and are generated in a "question language" and "answer language" respectively following the topics. The topics also determine an answer quality signal. Compared with translation models, our approach not only comprehensively models user behaviors on CQA portals, but also highlights the instinctive heterogeneity of questions and answers. More importantly, it takes answer quality into account and performs robustly against noise in answers. With the topic modeling approach, we propose a topic-based language model, which matches questions not only on a term level but also on a topic level. We conducted experiments on large scale data from Yahoo! Answers and Baidu Knows. Experimental results show that the proposed model can significantly outperform state-of-the-art retrieval models in CQA. Kai Zhang 0038, Wei Wu 0014, Haocheng Wu, Zhoujun Li 0001, Ming Zhou 0001 |
CIKM | 5 |
| 2014 | SocialTransfer: Transferring Social Knowledge for Cold-Start CowdsourcingabstractAn essential component of building a successful crowdsourcing market is effective task matching, which matches a given task to the right crowdworkers. In order to provide high- quality task matching, crowdsourcing systems rely on past task-solving activities of crowdworkers. However, the average number of past activities of crowdworkers in most crowd- sourcing systems is very small. We call the workers who have only solved a small number of tasks cold-start crowdworkers. We observe that most of the workers in crowdsourcing systems are cold-start crowdworkers, and crowdsourcing systems actually enjoy great benefits from cold-start crowd-workers. However, the problem of task matching with the presence of many cold-start crowdworkers has not been well studied. We propose a new approach to address this issue. Our main idea, motivated by the prevalence of online social networks, is to transfer the knowledge about crowdworkers in their social networks to crowdsourcing systems for task matching. We propose a SocialTransfer model for cold-start crowdsourcing, which not only infers the expertise of warm- start crowdworkers from their past activities, but also transfers the expertise knowledge to cold-start crowdworkers via social connections. We evaluate the SocialTransfer model on the well-known crowdsourcing system Quora, using knowledge from the popular social network Twitter. Experimental results show that, by transferring social knowledge, our method achieves significant improvements over the state-of-the-art methods. Zhou Zhao 0001, James Cheng, Furu Wei, Ming Zhou 0001, Wilfred Ng, Yingjun Wu |
CIKM | 4 |
| 2014 | A Lexicalized Reordering Model for Hierarchical Phrase-based Translation
Hailong Cao, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001, Tiejun Zhao |
COLING | 4 |
| 2014 | Soft Dependency Matching for Hierarchical Phrase-based Machine Translation
Hailong Cao, Dongdong Zhang 0001, Ming Zhou 0001, Tiejun Zhao |
COLING | 3 |
| 2014 | Building Large-Scale Twitter-Specific Sentiment Lexicon : A Representation Learning Approach
Duyu Tang, Furu Wei, Bing Qin 0001, Ming Zhou 0001, Ting Liu 0001 |
COLING | 4 |
| 2014 | A Joint Segmentation and Classification Framework for Sentiment AnalysisabstractIn this paper, we propose a joint segmentation and classification framework for sentiment analysis.Existing sentiment classification algorithms typically split a sentence as a word sequence, which does not effectively handle the inconsistent sentiment polarity between a phrase and the words it contains, such as "not bad" and "a great deal of ".We address this issue by developing a joint segmentation and classification framework (JSC), which simultaneously conducts sentence segmentation and sentence-level sentiment classification.Specifically, we use a log-linear model to score each segmentation candidate, and exploit the phrasal information of top-ranked segmentations as features to build the sentiment classifier.A marginal log-likelihood objective function is devised for the segmentation model, which is optimized for enhancing the sentiment classification performance.The joint model is trained only based on the annotated sentiment polarity of sentences, without any segmentation annotations.Experiments on a benchmark Twitter sentiment classification dataset in SemEval 2013 show that, our joint model performs comparably with the state-of-the-art methods. Duyu Tang, Furu Wei, Bing Qin 0001, Li Dong 0004, Ting Liu 0001, Ming Zhou 0001 |
EMNLP | 6 |
| 2014 | Joint Relational Embeddings for Knowledge-based Question AnsweringabstractTransforming a natural language (NL) question into a corresponding logical form (LF) is central to the knowledge-based question answering (KB-QA) task.Unlike most previous methods that achieve this goal based on mappings between lexicalized phrases and logical predicates, this paper goes one step further and proposes a novel embedding-based approach that maps NL-questions into LFs for KB-QA by leveraging semantic associations between lexical representations and KBproperties in the latent space.Experimental results demonstrate that our proposed method outperforms three KB-QA baseline methods on two publicly released QA data sets. Nan Duan 0001, Ming Zhou 0001, Hae-Chang Rim |
EMNLP | 3 |
| 2014 | Answer Extraction with Multiple Extraction Engines for Web-Based Question Answering
Furu Wei, Ming Zhou 0001 |
NLPCC | 3 |
| 2014 | Improving search relevance for short queries in community question answeringabstractRelevant question retrieval and ranking is a typical task in community question answering (CQA). Existing methods mainly focus on long and syntactically structured queries. However, when an input query is short, the task becomes challenging, due to a lack information regarding user intent. In this paper, we mine different types of user intent from various sources for short queries. With these intent signals, we propose a new intent-based language model. The model takes advantage of both state-of-the-art relevance models and the extra intent information mined from multiple sources. We further employ a state-of-the-art learning-to-rank approach to estimate parameters in the model from training data. Experiments show that by leveraging user intent prediction, our model significantly outperforms the state-of-the-art relevance models in question search. Haocheng Wu, Wei Wu 0014, Ming Zhou 0001, Enhong Chen, Lei Duan, Harry Shum |
WSDM | 3 |
| 2013 | The Automated Acquisition of Suggestions from TweetsabstractThis paper targets at automatically detecting and classifying user's suggestions from tweets. The short and informal nature of tweets, along with the imbalanced characteristics of suggestion tweets, makes the task extremely challenging. To this end, we develop a classification framework on Factorization Machines, which is effective and efficient especially in classification tasks with feature sparsity settings. Moreover, we tackle the imbalance problem by introducing cost-sensitive learning techniques in Factorization Machines. Extensively experimental studies on a manually annotated real-life data set show that the proposed approach significantly improves the baseline approach, and yields the precision of 71.06% and recall of 67.86%. We also investigate the reason why Factorization Machines perform better. Finally, we introduce the first manually annotated dataset for suggestion classification. Li Dong 0004, Furu Wei, Yajuan Duan, Ming Zhou 0001, Ke Xu 0001 |
AAAI | 5 |
| 2013 | Entity Linking for Tweets
Haocheng Wu, Ming Zhou 0001, Furu Wei |
ACL (1) | 4 |
| 2013 | Word Alignment Modeling with Context Dependent Deep Neural Network
Nan Yang 0002, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Nenghai Yu |
ACL (1) | 4 |
| 2013 | VisualComm: a tool to support communication between deaf and hearing persons with the KinectabstractWith the quickly increasing of the deaf community, how to communicate with the hearing persons is becoming a serious social problem. Furthermore, the investigation indicates that the deaf community is more self-enclosed and won't exchange ideas with the hearing. To address this challenge, we develop VisualComm, a tool to support communication between deaf and hearing persons with sign language recognition technology by using the Kinect. The main contribution of the system is a holistic solution of a two-way communication between deaf and hearings, and furthermore it is a seamless experience tailored for this particular activity. Currently we have implemented the basic communication based on 370 daily Chinese words for signer. Xiujuan Chai, Xilin Chen 0001, Ming Zhou 0001, Hanjing Li |
ASSETS | 4 |
| 2013 | Trustable aggregation of online ratingsabstractThe average of the customer ratings on the product, which we call reputation, is one of the key factors in online purchasing decision of a product. There is, however, no guarantee in the trustworthiness of the reputation since it can be manipulated rather easily. In this paper, we define false reputation as the problem of the reputation to be manipulated by unfair ratings, and design a general framework that provides trustable reputation. For this purpose, we propose TRUEREPUTATION, an algorithm that iteratively adjusts the reputation based on the confidence of customer ratings. Hyun-Kyo Oh, Sang-Wook Kim, Sunju Park, Ming Zhou 0001 |
CIKM | 4 |
| 2013 | Multi-Domain Adaptation for SMT Using Multi-Task LearningabstractDomain adaptation for SMT usually adapts models to an individual specific domain.However, it often lacks some correlation among different domains where common knowledge could be shared to improve the overall translation quality.In this paper, we propose a novel multi-domain adaptation approach for SMT using Multi-Task Learning (MTL), with in-domain models tailored for each specific domain and a general-domain model shared by different domains.The parameters of these models are tuned jointly via MTL so that they can learn general knowledge more accurately and exploit domain knowledge better.Our experiments on a largescale English-to-Chinese translation task validate that the MTL-based adaptation approach significantly and consistently improves the translation quality compared to a non-adapted baseline.Furthermore, it also outperforms the individual adaptation of each specific domain. Lei Cui 0001, Xilun Chen 0002, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001 |
EMNLP | 6 |
| 2013 | Efficient Collective Entity Linking with StackingabstractEntity disambiguation works by linking ambiguous mentions in text to their corresponding real-world entities in knowledge base.Recent collective disambiguation methods enforce coherence among contextual decisions at the cost of non-trivial inference processes.We propose a fast collective disambiguation approach based on stacking.First, we train a local predictor g 0 with learning to rank as base learner, to generate initial ranking list of candidates.Second, top k candidates of related instances are searched for constructing expressive global coherence features.A global predictor g 1 is trained in the augmented feature space and stacking is employed to tackle the train/test mismatch problem.The proposed method is fast and easy to implement.Experiments show its effectiveness over various algorithms on several public datasets.By learning a rich semantic relatedness measure between entity categories and context document, performance is further improved. Zhengyan He, Shujie Liu 0001, Yang Song 0021, Mu Li 0001, Ming Zhou 0001, Houfeng Wang |
EMNLP | 5 |
| 2013 | Answer Extraction from Passage Graph for Question Answering
Nan Duan 0001, Yajuan Duan, Ming Zhou 0001 |
IJCAI | 4 |
| 2013 | Collective Corpus Weighting and Phrase Scoring for SMT Using Graph-Based Random Walk
Lei Cui 0001, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001 |
NLPCC | 5 |
| 2013 | The Spoken/Written Language Classification of English Sentences with Bilingual Information
Kuan Li, Zhongyang Xiong, Ming Zhou 0001 |
NLPCC | 5 |
| 2013 | Two-stage NER for tweets with clustering
Ming Zhou 0001 |
Inf. Process. Manag. | 2 |
| 2013 | Named entity recognition for tweetsabstractTwo main challenges of Named Entity Recognition (NER) for tweets are the insufficient information in a tweet and the lack of training data. We propose a novel method consisting of three core elements: (1) normalization of tweets; (2) combination of a K-Nearest Neighbors (KNN) classifier with a linear Conditional Random Fields (CRF) model; and (3) semisupervised learning framework. The tweet normalization preprocessing corrects common ill-formed words using a global linear model. The KNN-based classifier conducts prelabeling to collect global coarse evidence across tweets while the CRF model conducts sequential labeling to capture fine-grained information encoded in a tweet. The semisupervised learning plus the gazetteers alleviate the lack of training data. Extensive experiments show the advantages of our method over the baselines as well as the effectiveness of normalization, KNN, and semisupervised learning. Furu Wei, Shaodian Zhang, Ming Zhou 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2013 | Entity Translation Mining from Comparable Corpora: Combining Graph Mapping with Corpus Latent FeaturesabstractThis paper addresses the problem of mining named entity translations from comparable corpora, specifically, mining English and Chinese named entity translation. We first observe that existing approaches use one or more of the following named entity similarity metrics: entity, entity context, and relationship. Motivated by this observation, we propose a new holistic approach by 1) combining all similarity types used and 2) additionally considering relationship context similarity between pairs of named entities, a missing quadrant in the taxonomy of similarity metrics. We abstract the named entity translation problem as the matching of two named entity graphs extracted from the comparable corpora. Specifically, named entity graphs are first constructed from comparable corpora to extract relationship between named entities. Entity similarity and entity context similarity are then calculated from every pair of bilingual named entities. A reinforcing method is utilized to reflect relationship similarity and relationship context similarity between named entities. We also discover "latent" features lost in the graph extraction process and integrate this into our framework. According to our experimental results, our holistic graph-based approach and its enhancement using corpus latent features are highly effective and our framework significantly outperforms previous approaches. Jinhan Kim, Seung-won Hwang, Long Jiang, Young-In Song, Ming Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2012 | Generating Chinese Classical Poems with Statistical Machine Translation ModelsabstractThis paper describes a statistical approach to generation of Chinese classical poetry and proposes a novel method to automatically evaluate poems. The system accepts a set of keywords representing the writing intents from a writer and generates sentences one by one to form a completed poem. A statistical machine translation (SMT) system is applied to generate new sentences, given the sentences generated previously. For each line of sentence a specific model specially trained for that line is used, as opposed to using a single model for all sentences. To enhance the coherence of sentences on every line, a coherence model using mutual information is applied to select candidates with better consistency with previous sentences. In addition, we demonstrate the effectiveness of the BLEU metric for evaluation with a novel method of generating diverse references. Ming Zhou 0001, Long Jiang |
AAAI | 2 |
| 2012 | Collective Nominal Semantic Role Labeling for TweetsabstractTweets have become an increasingly popular source of fresh information. We investigate the task of Nominal Semantic Role Labeling (NSRL) for tweets, which aims to identify predicate-argument structures defined by nominals in tweets. Studies of this task can help fine-grained information extraction and retrieval from tweets. There are two main challenges in this task: 1) The lack of information in a single tweet, rooted in the short and noisy nature of tweets; and 2) recovery of implicit arguments. We propose jointly conducting NSRL on multiple similar tweets using a graphical model, leveraging the redundancy in tweets to tackle these challenges. Extensive evaluations on a human annotated data set demonstrate that our method outperforms two baselines with an absolute gain of 2.7% in F1. Zhongyang Fu, Furu Wei, Ming Zhou 0001 |
AAAI | 4 |
| 2012 | Exacting Social Events for Tweets Using a Factor GraphabstractSocial events are events that occur between people where at least one person is aware of the other and of the event taking place. Extracting social events can play an important role in a wide range of applications, such as the construction of social network. In this paper, we introduce the task of social event extraction for tweets, an important source of fresh events. One main challenge is the lack of information in a single tweet, which is rooted in the short and noise-prone nature of tweets. We propose to collectively extract social events from multiple similar tweets using a novel factor graph, to harvest the redundance in tweets, i.e., the repeated occurrences of a social event in several tweets. We evaluate our method on a human annotated data set, and show that it outperforms all baselines, with an absolute gain of 21% in F1. Xiangyang Zhou, Zhongyang Fu, Furu Wei, Ming Zhou 0001 |
AAAI | 5 |
| 2012 | Learning Translation Consensus with Structured Label Propagation
Shujie Liu 0001, Chi-Ho Li, Mu Li 0001, Ming Zhou 0001 |
ACL (1) | 4 |
| 2012 | Joint Inference of Named Entity Recognition and Normalization for Tweets
Ming Zhou 0001, Xiangyang Zhou, Zhongyang Fu, Furu Wei |
ACL (1) | 2 |
| 2012 | Cross-Lingual Mixture Model for Sentiment Classification
Xinfan Meng, Furu Wei, Ming Zhou 0001, Houfeng Wang |
ACL (1) | 4 |
| 2012 | Graph-based collective classification for tweetsabstractIn this paper, we address the problem of classifying tweets into topical categories. Because of the short, noisy and ambiguous nature of tweets, we propose to collectively conduct the classification by exploiting the context information (i.e. related tweets) other than individually as in conventional text classification methods. In particular, we augment the content-based representation of text with tweets sharing same #hashtag or URL, which results in a tweet graph. We then formulate the tweet classification task under a graph optimization framework. We investigate three popular approaches, namely, Loopy Belief Propagation (LBP), Relaxation Labeling (RL), and Iterative Classification Algorithm (ICA). Extensive experiment results show that the graph-based tweet classification approach remarkably improves the performance, while the ICA model with relationship of sharing the same #hashtag gives the best result on separate tweet graph. Yajuan Duan, Furu Wei, Ming Zhou 0001, Harry Shum |
CIKM | 3 |
| 2012 | Twitter Topic Summarization by Ranking Tweets using Social Influence and Content Quality
Yajuan Duan, Furu Wei, Ming Zhou 0001, Harry Shum |
COLING | 4 |
| 2012 | Graph-Based Multi-Tweet Summarization using Social Signals
Furu Wei, Ming Zhou 0001 |
COLING | 4 |
| 2012 | Forced Derivation Tree based Model Training to Statistical Machine Translation
Nan Duan 0001, Mu Li 0001, Ming Zhou 0001 |
EMNLP-CoNLL | 3 |
| 2012 | Re-training Monolingual Parser Bilingually for Syntactic SMT
Shujie Liu 0001, Chi-Ho Li, Mu Li 0001, Ming Zhou 0001 |
EMNLP-CoNLL | 4 |
| 2012 | Entity-centric topic-oriented opinion summarization in twitterabstractMicroblogging services, such as Twitter, have become popular channels for people to express their opinions towards a broad range of topics. Twitter generates a huge volume of instant messages (i.e. tweets) carrying users' sentiments and attitudes every minute, which both necessitates automatic opinion summarization and poses great challenges to the summarization system. In this paper, we study the problem of opinion summarization for entities, such as celebrities and brands, in Twitter. We propose an entity-centric topic-based opinion summarization framework, which aims to produce opinion summaries in accordance with topics and remarkably emphasizing the insight behind the opinions. To this end, we first mine topics from #hashtags, the human-annotated semantic tags in tweets. We integrate the #hashtags as weakly supervised information into topic modeling algorithms to obtain better interpretation and representation for calculating the similarity among them, and adopt Affinity Propagation algorithm to group #hashtags into coherent topics. Subsequently, we use templates generalized from paraphrasing to identify tweets with deep insights, which reveal reasons, express demands or reflect viewpoints. Afterwards, we develop a target (i.e. entity) dependent sentiment classification approach to identifying the opinion towards a given target (i.e. entity) of tweets. Finally, the opinion summary is generated through integrating information from dimensions of topic, opinion and insight, as well as other factors (e.g. topic relevancy, redundancy and language styles) in an unified optimization framework. We conduct extensive experiments on a real-life data set to evaluate the performance of individual opinion summarization modules as well as the quality of the produced summary. The promising experiment results show the effectiveness of the proposed framework and algorithms. Xinfan Meng, Furu Wei, Ming Zhou 0001, Sujian Li, Houfeng Wang |
KDD | 4 |
| 2011 | Enhancing Semantic Role Labeling for Tweets Using Self-TrainingabstractSemantic Role Labeling (SRL) for tweets is a meaningful task that can benefit a wide range of applications such as fine-grained information extraction and retrieval from tweets. One main challenge of the task is the lack of annotated tweets, which is required to train a statistical model. We introduce self-training to SRL, leveraging abundant unlabeled tweets to alleviate its depending on annotated tweets. A novel strategy of tweet selection is presented, ensuring the chosen tweets are both correct and informative. More specifically, the correctness is estimated according to the labeling confidences and agreement of two Conditional Random Fields based labelers, which are trained on the randomly evenly spitted labeled data; while the informativeness is in proportion to the maximum distance between the tweet and the already selected tweets. We evaluate our method on a human annotated data set and show that bootstrapping improve a baseline by 3.4% F1. Kuan Li, Ming Zhou 0001, Zhongyang Xiong |
AAAI | 3 |
| 2011 | Hypothesis Mixture Decoding for Statistical Machine Translation
Nan Duan 0001, Mu Li 0001, Ming Zhou 0001 |
ACL | 3 |
| 2011 | Target-dependent Twitter Sentiment Classification
Long Jiang, Mo Yu, Ming Zhou 0001, Tiejun Zhao |
ACL | 3 |
| 2011 | Recognizing Named Entities in Tweets
Shaodian Zhang, Furu Wei, Ming Zhou 0001 |
ACL | 4 |
| 2011 | Correcting Verb Selection Errors for ESL with the Perceptron
Ming Zhou 0001 |
CICLing (2) | 3 |
| 2011 | Mining entity translations from comparable corpora: a holistic graph mapping approachabstractThis paper addresses the problem of mining named entity translations from comparable corpora, specifically, mining English and Chinese named entity translation. We first observe that existing approaches use one or more of the following named entity similarity metrics: entity, entity context, and relationship. Inspired by this observation, in this paper, we propose a new holistic approach, by (1) combining all similarity types used and (2) additionally considering relationship context similarity between pairs of named entities, a missing quadrant in the taxonomy of similarity metrics. We abstract the named entity translation problem as the matching of two named entity graphs extracted from the comparable corpora. Specifically, named entity graphs are first constructed from comparable corpora to extract relationship between named entities. Entity similarity and entity context similarity are then calculated from every pair of bilingual named entities. A reinforcing method is utilized to reflect relationship similarity and relationship context similarity between named entities. According to our experimental results, our holistic graph-based approach significantly outperforms previous approaches. Jinhan Kim, Long Jiang, Seung-won Hwang, Young-In Song, Ming Zhou 0001 |
CIKM | 5 |
| 2011 | Topic sentiment analysis in twitter: a graph-based hashtag sentiment classification approachabstractTwitter is one of the biggest platforms where massive instant messages (i.e. tweets) are published every day. Users tend to express their real feelings freely in Twitter, which makes it an ideal source for capturing the opinions towards various interesting topics, such as brands, products or celebrities, etc. Naturally, people may anticipate an approach to receiving the common sentiment tendency towards these topics directly rather than through reading the huge amount of tweets about them. On the other side, Hashtags, starting with a symbol "#" ahead of keywords or phrases, are widely used in tweets as coarse-grained topics. In this paper, instead of presenting the sentiment polarity of each tweet relevant to the topic, we focus our study on hashtag-level sentiment classification. This task aims to automatically generate the overall sentiment polarity for a given hashtag in a certain time period, which markedly differs from the conventional sentence-level and document-level sentiment analysis. Our investigation illustrates that three types of information is useful to address the task, including (1) sentiment polarity of tweets containing the hashtag; (2) hashtags co-occurrence relationship and (3) the literal meaning of hashtags. Consequently, in order to incorporate the first two types of information into a classification framework where hashtags can be classified collectively, we propose a novel graph model and investigate three approximate collective classification algorithms for inference. Going one step further, we show that the performance can be remarkably improved using an enhanced boosting classification setting in which we employ the literal meaning of hashtags as semi-supervised information. Experimental results on a real-life data set consisting of 29,195 tweets and 2,181 hashtags show the effectiveness of the proposed model and algorithms. Furu Wei, Ming Zhou 0001, Ming Zhang 0004 |
CIKM | 4 |
| 2011 | Collective Semantic Role Labeling for Tweets with ClusteringabstractAs tweets have become a comprehensive repository of fresh information, Semantic Role Labeling (SRL) for tweets has aroused great research interests because of its central role in a wide range of tweet related studies such as fine-grained information extraction, sentiment analysis and summarization. However, the fact that a tweet is often too short and informal to provide sufficient information poses a major challenge. To tackle this challenge, we propose a new method to collectively label similar tweets. The underlying idea is to exploit similar tweets to make up for the lack of information in a tweet. Specifically, similar tweets are first grouped together by clustering. Then for each cluster a two-stage labeling is conducted: One labeler conducts SRL to get statistical information, such as the predicate/argument/role triples that occur frequently, from its highly confidently labeled results; then in the second stage, another labeler performs SRL with such statistical information to refine the results. Experimental results on a human annotated dataset show that our approach remarkably improves SRL by 3.1 % F1. Kuan Li, Ming Zhou 0001, Zhongyang Xiong |
IJCAI | 3 |
| 2011 | User-level sentiment analysis incorporating social networksabstractWe show that information about social relationships can be used to improve user-level sentiment analysis. The main motivation behind our approach is that users that are somehow "connected" may be more likely to hold similar opinions; therefore, relationship information can complement what we can extract about a user's viewpoints from their utterances. Employing Twitter as a source for our experimental data, and working within a semi-supervised framework, we propose models that are induced either from the Twitter follower/followee network or from the network in Twitter formed by users referring to each other using "@" mentions. Our transductive learning results reveal that incorporating social-network information can indeed lead to statistically significant sentiment classification improvements over the performance of an approach based on Support Vector Machines having access only to textual features. Chenhao Tan, Lillian Lee, Jie Tang 0001, Long Jiang, Ming Zhou 0001, Ping Li 0001 |
KDD | 5 |
| 2011 | Statistic Machine Translation Boosted with Spurious Word Deletion
Shujie Liu 0001, Chi-Ho Li, Ming Zhou 0001 |
MTSummit | 3 |
| 2011 | A Unified SMT Framework Combining MIRA and MERT
Shujie Liu 0001, Chi-Ho Li, Ming Zhou 0001 |
MTSummit | 3 |
| 2011 | Function Word Generation in Statistical Machine Translation Systems
Lei Cui 0001, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001 |
MTSummit | 4 |
| 2011 | Improving Phrase Extraction via MBR Phrase Scoring and Pruning
Nan Duan 0001, Mu Li 0001, Ming Zhou 0001, Lei Cui 0001 |
MTSummit | 3 |
| 2011 | Tulsa: web search for writing assistanceabstractNo abstract available. Duo Ding, Xingping Jiang, Matthew R. Scott, Ming Zhou 0001, Yong Yu 0001 |
SIGIR | 4 |
| 2011 | QuickView: advanced search of tweetsabstractTweets have become a comprehensive repository for real-time information. However, it is often hard for users to quickly get information they are interested in from tweets, owing to the sheer volume of tweets as well as their noisy and informal nature. We present QuickView, an NLP-based tweet search platform to tackle this issue. Specifically, it exploits a series of natural language processing technologies, such as tweet normalization, named entity recognition, semantic role labeling, sentiment analysis, tweet classification, to extract useful information, i.e., named entities, events, opinions, etc., from a large volume of tweets. Then, non-noisy tweets, together with the mined information, are indexed, on top of which two brand new scenarios are enabled, i.e., categorized browsing and advanced search, allowing users to effectively access either the tweets or fine-grained information they are interested in. Long Jiang, Furu Wei, Ming Zhou 0001 |
SIGIR | 4 |
| 2010 | Discriminative Pruning for Discriminative ITG Alignment
Shujie Liu 0001, Chi-Ho Li, Ming Zhou 0001 |
ACL | 3 |
| 2010 | An Empirical Study on Learning to Rank of Tweets
Yajuan Duan, Long Jiang, Tao Qin 0001, Ming Zhou 0001, Harry Shum |
COLING | 4 |
| 2010 | Mixture Model-based Minimum Bayes Risk Decoding using Multiple Machine Translation Systems
Nan Duan 0001, Mu Li 0001, Dongdong Zhang 0001, Ming Zhou 0001 |
COLING | 4 |
| 2010 | Translation Model Generalization using Probability Averaging for Machine Translation
Nan Duan 0001, Ming Zhou 0001 |
COLING | 3 |
| 2010 | An Empirical Study on Web Mining of Parallel Data
Gumwon Hong, Chi-Ho Li, Ming Zhou 0001, Hae-Chang Rim |
COLING | 3 |
| 2010 | Adaptive Development Data Selection for Log-linear Model in Statistical Machine Translation
Mu Li 0001, Yinggong Zhao, Dongdong Zhang 0001, Ming Zhou 0001 |
COLING | 4 |
| 2010 | Semantic Role Labeling for News Tweets
Kuan Li, Ming Zhou 0001, Long Jiang, Zhongyang Xiong, Changning Huang |
COLING | 4 |
| 2010 | SRL-Based Verb Selection for ESL
Kuan Li, Stephan Hyeonjun Stiller, Ming Zhou 0001 |
EMNLP | 5 |
| 2010 | Exploiting query logs for cross-lingual query suggestionsabstractQuery suggestion aims to suggest relevant queries for a given query, which helps users better specify their information needs. Previous work on query suggestion has been limited to the same language. In this article, we extend it to cross-lingual query suggestion (CLQS): for a query in one language, we suggest similar or relevant queries in other languages. This is very important to the scenarios of cross-language information retrieval (CLIR) and other related cross-lingual applications. Instead of relying on existing query translation technologies for CLQS, we present an effective means to map the input query of one language to queries of the other language in the query log. Important monolingual and cross-lingual information such as word translation relations and word co-occurrence statistics, and so on, are used to estimate the cross-lingual query similarity with a discriminative model. Benchmarks show that the resulting CLQS system significantly outperforms a baseline system that uses dictionary-based query translation. Besides, we evaluate CLQS with French-English and Chinese-English CLIR tasks on TREC-6 and NTCIR-4 collections, respectively. The CLIR experiments using typical retrieval models demonstrate that the CLQS-based approach has significantly higher effectiveness than several traditional query translation methods. We find that when combined with pseudo-relevance feedback, the effectiveness of CLIR using CLQS is enhanced for different pairs of languages. Wei Gao 0001, Cheng Niu, Jian-Yun Nie, Ming Zhou 0001, Kam-Fai Wong, Hsiao-Wuen Hon |
ACM Trans. Inf. Syst. | 4 |
| 2009 | Exploiting Bilingual Information to Improve Web Search
Wei Gao 0001, John Blitzer, Ming Zhou 0001, Kam-Fai Wong |
ACL/IJCNLP | 3 |
| 2009 | Mining Bilingual Data from the Web with Adaptively Learnt Patterns
Long Jiang, Shiquan Yang, Ming Zhou 0001, Qingsheng Zhu |
ACL/IJCNLP | 3 |
| 2009 | Collaborative Decoding: Partial Hypothesis Re-ranking Using Translation Consensus between Decoders
Mu Li 0001, Nan Duan 0001, Dongdong Zhang 0001, Chi-Ho Li, Ming Zhou 0001 |
ACL/IJCNLP | 5 |
| 2009 | Joint Ranking for Multilingual Web Search
Wei Gao 0001, Cheng Niu, Ming Zhou 0001, Kam-Fai Wong |
ECIR | 3 |
| 2009 | The Feature Subspace Method for SMT System Combination
Nan Duan 0001, Mu Li 0001, Tong Xiao 0001, Ming Zhou 0001 |
EMNLP | 4 |
| 2009 | Better Synchronous Binarization for Machine Translation
Tong Xiao 0001, Mu Li 0001, Dongdong Zhang 0001, Ming Zhou 0001 |
EMNLP | 5 |
| 2009 | Generating Chinese Couplets and Quatrain Using a Statistical Approach
Ming Zhou 0001, Long Jiang |
PACLIC | 1 |
| 2008 | Measure Word Generation for English-Chinese SMT Systems
Dongdong Zhang 0001, Mu Li 0001, Nan Duan 0001, Chi-Ho Li, Ming Zhou 0001 |
ACL | 5 |
| 2008 | Combining Multiple Resources to Improve SMT-based Paraphrasing Model
Cheng Niu, Ming Zhou 0001, Ting Liu 0001, Sheng Li 0003 |
ACL | 3 |
| 2008 | Generating Chinese Couplets using a Statistical MT Approach
Long Jiang, Ming Zhou 0001 |
COLING | 2 |
| 2008 | Diagnostic Evaluation of Machine Translation Systems Using Automatically Constructed Linguistic Check-Points
Ming Zhou 0001, Shujie Liu 0001, Mu Li 0001, Dongdong Zhang 0001, Tiejun Zhao |
COLING | 1 |
| 2007 | Mining Sequential Patterns and Tree Patterns to Detect Erroneous Sentences
Guihua Sun, Gao Cong, Chin-Yew Lin, Ming Zhou 0001 |
AAAI | 5 |
| 2007 | A Probabilistic Approach to Syntax-based Reordering for Statistical Machine Translation
Chi-Ho Li, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001, Yi Guan |
ACL | 5 |
| 2007 | Detecting Erroneous Sentences using Automatically Mined Sequential Patterns
Guihua Sun, Gao Cong, Ming Zhou 0001, Zhongyang Xiong, John Lee 0001, Chin-Yew Lin |
ACL | 4 |
| 2007 | Improving Query Spelling Correction Using Web Search Results
Mu Li 0001, Ming Zhou 0001 |
EMNLP-CoNLL | 3 |
| 2007 | Phrase Reordering Model Integrating Syntactic Knowledge for SMT
Dongdong Zhang 0001, Mu Li 0001, Chi-Ho Li, Ming Zhou 0001 |
EMNLP-CoNLL | 4 |
| 2007 | Named Entity Translation with Web Mining and Transliteration
Long Jiang, Ming Zhou 0001, Lee-Feng Chien, Cheng Niu |
IJCAI | 2 |
| 2007 | Learning Question Paraphrases for QA from Encarta Logs
Ming Zhou 0001, Ting Liu 0001 |
IJCAI | 2 |
| 2007 | Cross-lingual query suggestion using query logs of different languagesabstractQuery suggestion aims to suggest relevant queries for a given query, which help users better specify their information needs. Previously, the suggested terms are mostly in the same language of the input query. In this paper, we extend it to cross-lingual query suggestion (CLQS): for a query in one language, we suggest similar or relevant queries in other languages. This is very important to scenarios of cross-language information retrieval (CLIR) and cross-lingual keyword bidding for search engine advertisement. Instead of relying on existing query translation technologies for CLQS, we present an effective means to map the input query of one language to queries of the other language in the query log. Important monolingual and cross-lingual information such as word translation relations and word co-occurrence statistics, etc. are used to estimate the cross-lingual query similarity with a discriminative model. Benchmarks show that the resulting CLQS system significantly out performs a baseline system based on dictionary-based query translation. Besides, the resulting CLQS is tested with French to English CLIR tasks on TREC collections. The results demonstrate higher effectiveness than the traditional query translation methods. Wei Gao 0001, Cheng Niu, Jian-Yun Nie, Ming Zhou 0001, Kam-Fai Wong, Hsiao-Wuen Hon |
SIGIR | 4 |
| 2006 | Reranking Answers for Definitional QA Using Language ModelingabstractStatistical ranking methods based on centroid vector (profile) extracted from external knowledge have become widely adopted in the top definitional QA systems in TREC 2003 and 2004. In these approaches, terms in the centroid vector are treated as a bag of words based on the independent assumption. To relax this assumption, this paper proposes a novel language model-based answer reranking method to improve the existing bag-of-words model approach by considering the dependence of the words in the centroid vector. Experiments have been conducted to evaluate the different dependence models. The results on the TREC 2003 test set show that the reranking approach with biterm language model, significantly outperforms the one with the bag-of-words model and unigram language model by 14.9% and 12.5% respectively in F-Measure(5). Ming Zhou 0001 |
ACL | 2 |
| 2006 | Exploring Distributional Similarity Based Models for Query Spelling CorrectionabstractA query speller is crucial to search engine in improving web search relevance. This paper describes novel methods for use of distributional similarity estimated from query logs in learning improved query spelling correction models. The key to our methods is the property of distributional similarity between two terms: it is high between a frequently occurring misspelling and its correction, and low between two irrelevant terms only with similar spellings. We present two models that are able to take advantage of this property. Experimental results demonstrate that the distributional similarity based models can significantly outperform their baseline systems in the web query spelling correction task. Mu Li 0001, Muhua Zhu, Ming Zhou 0001 |
ACL | 4 |
| 2006 | A DOM Tree Alignment Model for Mining Parallel Data from the WebabstractThis paper presents a new web mining scheme for parallel data acquisition.Based on the Document Object Model (DOM), a web page is represented as a DOM tree.Then a DOM tree alignment model is proposed to identify the translationally equivalent texts and hyperlinks between two parallel DOM trees.By tracing the identified parallel hyperlinks, parallel web documents are recursively mined.Compared with previous mining schemes, the benchmarks show that this new mining scheme improves the mining coverage, reduces mining bandwidth, and enhances the quality of mined parallel sentences. Cheng Niu, Ming Zhou 0001, Jianfeng Gao 0001 |
ACL | 3 |
| 2006 | Statistical query translation models for cross-language information retrievalabstractQuery translation is an important task in cross-language information retrieval (CLIR), which aims to determine the best translation words and weights for a query. This article presents three statistical query translation models that focus on the resolution of query translation ambiguities. All the models assume that the selection of the translation of a query term depends on the translations of other terms in the query. They differ in the way linguistic structures are detected and exploited. The co-occurrence model treats a query as a bag of words and uses all the other terms in the query as the context for translation disambiguation. The other two models exploit linguistic dependencies among terms. The noun phrase (NP) translation model detects NPs in a query, and translates each NP as a unit by assuming that the translation of a term only depends on other terms within the same NP. Similarly, the dependency translation model detects and translates dependency triples, such as verb-object, as units. The evaluations show that linguistic structures always lead to more precise translations. The experiments of CLIR on TREC Chinese collections show that all three models have a positive impact on query translation and lead to significant improvements of CLIR performance over the simple dictionary-based translation method. The best results are obtained by combining the three models. Jianfeng Gao 0001, Jian-Yun Nie, Ming Zhou 0001 |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2005 | Resume Information Extraction with Cascaded Hybrid ModelabstractThis paper presents an effective approach for resume information extraction to support automatic resume management and routing. A cascaded information extraction (IE) framework is designed. In the first pass, a resume is segmented into a consecutive blocks attached with labels indicating the information types. Then in the second pass, the detailed information, such as Name and Address, are identified in certain blocks (e.g. blocks labelled with Personal Information), instead of searching globally in the entire resume. The most appropriate model is selected through experiments for each IE task in different passes. The experimental results show that this cascaded hybrid model achieves better F-score than flat models that do not apply the hierarchical structure of resumes. It also shows that applying different IE models in different passes according to the contextual structure is effective. Gang Guan, Ming Zhou 0001 |
ACL | 3 |
| 2004 | Collocation Translation Acquisition Using Monolingual CorporaabstractCollocation translation is important for machine translation and many other NLP tasks. Unlike previous methods using bilingual parallel corpora, this paper presents a new method for acquiring collocation translations by making use of monolingual corpora and linguistic knowledge. First, dependency triples are extracted from Chinese and English corpora with dependency parsers. Then, a dependency triple translation model is estimated using the EM algorithm based on a dependency correspondence assumption. The generated triple translation model is used to extract collocation translations from two monolingual corpora. Experiments show that our approach outperforms the existing monolingual corpus based methods in dependency triple translation and achieves promising results in collocation translation extraction. Yajuan Lü, Ming Zhou 0001 |
ACL | 2 |
| 2004 | A New Approach for English-Chinese Named Entity Alignment
Donghui Feng 0001, Yajuan Lü, Ming Zhou 0001 |
EMNLP | 3 |
| 2003 | Synonymous Collocation Extraction Using Translation InformationabstractAutomatically acquiring synonymous collocation pairs such as and from corpora is a challenging task. For this task, we can, in general, have a large monolingual corpus and/or a very limited bilingual corpus. Methods that use monolingual corpora alone or use bilingual corpora alone are apparently inadequate because of low precision or low coverage. In this paper, we propose a method that uses both these resources to get an optimal compromise of precision and coverage. This method first gets candidates of synonymous collocation pairs based on a monolingual corpus and a word thesaurus, and then selects the appropriate pairs from the candidates using their translations in a second language. The translations of the candidates are obtained with a statistical translation model which is trained with a small bilingual corpus and a large monolingual corpus. The translation information is proved as effective to select synonymous collocation pairs. Experimental results indicate that the average precision and recall of our approach are 74% and 64% respectively, which outperform those methods that only use monolingual corpora and those that only use bilingual corpora. Ming Zhou 0001 |
ACL | 2 |
| 2002 | Chinese Named Entity Identification Using Class-based Language Model
Jian Sun 0001, Jianfeng Gao 0001, Lei Zhang 0001, Ming Zhou 0001, Changning Huang |
COLING | 4 |
| 2002 | An Automatic Evaluation Method for Localization Oriented Lexicalised EBMT System
Ming Zhou 0001, Tiejun Zhao, Hao Yu 0005, Sheng Li 0003 |
COLING | 2 |
| 2002 | Resolving query translation ambiguity using a decaying co-occurrence model and syntactic dependence relationsabstractBilingual dictionaries have been commonly used for query translation in cross-language information retrieval (CLIR). However, we are faced with the problem of translation selection. Several recent studies suggested the utilization of term co-occurrences in this selection. This paper presents two extensions to improve them. First, we extend the basic co-occurrence model by adding a decaying factor that decreases the mutual information when the distance between the terms increases. Second, we incorporate a triple translation model, in which syntactic dependence relations (represented as triples) are integrated. Our evaluation on translation accuracy shows that translating triples as units is more precise than a word-by-word translation. Our CLIR experiments show that the addition of the decaying factor leads to substantial improvements of the basic co-occurrence model; and the triple translation model brings further improvements. Jianfeng Gao 0001, Ming Zhou 0001, Jian-Yun Nie, Hongzhao He |
SIGIR | 2 |
| 2001 | Improving Query Translation for Cross-Language Information Retrieval Using Statistical ModelsabstractDictionaries have often been used for query translation in cross-language information retrieval (CLIR). However, we are faced with the problem of translation ambiguity, i.e. multiple translations are stored in a dictionary for a word. In addition, a word-by-word query translation is not precise enough. In this paper, we explore several methods to improve the previous dictionary-based query translation. First, as many as possible, noun phrases are recognized and translated as a whole by using statistical models and phrase translation patterns. Second, the best word translations are selected based on the cohesion of the translation words. Our experimental results on TREC English-Chinese CLIR collection show that these techniques result in significant improvements over the simple dictionary approaches, and achieve even better performance than a high-quality machine translation system. Jianfeng Gao 0001, Endong Xun, Ming Zhou 0001, Changning Huang, Jian-Yun Nie |
SIGIR | 3 |
| 2000 | PENS: A Machine-aided English Writing System for Chinese UsersabstractWriting English is a big barrier for most Chinese users. To build a computer-aided system that helps Chinese users not only on spelling checking and grammar checking but also on writing in the way of native-English is a challenging task. Although machine translation is widely used for this purpose, how to find an efficient way in which human collaborates with computers remains an open issue. In this paper, based on the comprehensive study of Chinese users requirements, we propose an approach to machine aided English writing system, which consists of two components: 1) a statistical approach to word spelling help, and 2) an information retrieval based approach to intelligent recommendation by providing suggestive example sentences. Both components work together in a unified way, and highly improve the productivity of English writing. We also developed a pilot system, namely PENS (Perfect ENglish System). Preliminary experiments show very promising results. Ting Liu 0001, Ming Zhou 0001, Jianfeng Gao 0001, Endong Xun, Changning Huang |
ACL | 2 |
| 2000 | A Unified Statistical Model for the Identification of English BaseNPabstractThis paper presents a novel statistical model for automatic identification of English baseNP. It uses two steps: the N-best Part-Of-Speech (POS) tagging and baseNP identification given the N-best POS-sequences. Unlike the other approaches where the two steps are separated, we integrate them into a unified statistical framework. Our model also integrates lexical information. Finally, Viterbi algorithm is applied to make global search in the entire sentence, allowing us to obtain linear complexity for the entire process. Compared with other methods using the same testing set, our approach achieves 92.3% in precision and 93.2% in recall. The result is comparable with or better than the previously reported results. Endong Xun, Changning Huang, Ming Zhou 0001 |
ACL | 3 |