Ming Gong 0001

dblp:34/4521-1 · DBLP profile ↗
← Back
44ranked-venue papers
3as first author
29since 2021 · last 2026
0000-0001-6140-7187ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 2 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
abstract
Shiping Yang, Jie Wu, Wenbiao Ding, Ning Wu, Shining Liang, Ming Gong, Hongzhi Li, Hengyuan Zhang, Angel X. Chang, Dongmei Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jie Wu 0018, Wenbiao Ding, Ning Wu 0013, Shining Liang, Ming Gong 0001, Angel X. Chang, Dongmei Zhang 0001
ACL (1)6
2025 Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
abstract
Truthfulness stands out as an essential challenge for Large Language Models (LLMs). Although many works have developed various ways for truthfulness enhancement, they seldom focus on truthfulness in multilingual scenarios. Meanwhile, contemporary multilingual aligning technologies struggle to balance massive languages and often exhibit serious truthfulness gaps across different languages, especially those that differ greatly from English. In our work, we extend truthfulness evaluation to multilingual contexts and propose a practical method for cross-lingual truthfulness transfer called Fact-aware Multilingual Selective Synergy (FaMSS). FaMSS is able to select an optimal subset of all tested languages by language bias and transfer contributions, and then employ translation instruction tuning for cross-lingual truthfulness transfer. Experimental results demonstrate that our approach can effectively reduce the multilingual representation disparity and boost cross-lingual truthfulness transfer of LLMs.
Ning Wu 0013, Wenbiao Ding, Shining Liang, Ming Gong 0001, Dongmei Zhang 0001
COLING5
2024 Grounding Language Models for Visual Entity Recognition
Zilin Xiao, Ming Gong 0001, Paola Cascante-Bonilla, Xingyao Zhang 0003, Jie Wu 0018, Vicente Ordonez
ECCV (11)2
2023 WIERT: Web Information Extraction via Render Tree
abstract
Web information extraction (WIE) is a fundamental problem in web document understanding, with a significant impact on various applications. Visual information plays a crucial role in WIE tasks as the nodes containing relevant information are often visually distinct, such as being in a larger font size or having a brighter color, from the other nodes. However, rendering visual information of a web page can be computationally expensive. Previous works have mainly focused on the Document Object Model (DOM) tree, which lacks visual information. To efficiently exploit visual information, we propose leveraging the render tree, which combines the DOM tree and Cascading Style Sheets Object Model (CSSOM) tree, and contains not only content and layout information but also rich visual information at a little additional acquisition cost compared to the DOM tree. In this paper, we present WIERT, a method that effectively utilizes the render tree of a web page based on a pretrained language model. We evaluate WIERT on the Klarna product page dataset, a manually labeled dataset of renderable e-commerce web pages, demonstrating its effectiveness and robustness.
Zimeng Li 0002, Linjun Shou, Ming Gong 0001, Daxin Jiang
AAAI4
2023 A Graph Fusion Approach for Cross-Lingual Machine Reading Comprehension
abstract
Although great progress has been made for Machine Reading Comprehension (MRC) in English, scaling out to a large number of languages remains a huge challenge due to the lack of large amounts of annotated training data in non-English languages. To address this challenge, some recent efforts of cross-lingual MRC employ machine translation to transfer knowledge from English to other languages, through either explicit alignment or implicit attention. For effective knowledge transition, it is beneficial to leverage both semantic and syntactic information. However, the existing methods fail to explicitly incorporate syntax information in model learning. Consequently, the models are not robust to errors in alignment and noises in attention. In this work, we propose a novel approach, which jointly models the cross-lingual alignment information and the mono-lingual syntax information using a graph. We develop a series of algorithms, including graph construction, learning, and pre-training. The experiments on two benchmark datasets for cross-lingual MRC show that our approach outperforms all strong baselines, which verifies the effectiveness of syntax information for cross-lingual MRC.
Zenan Xu, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Qinliang Su, Xiaojun Quan, Daxin Jiang
AAAI4
2023 Alleviating Over-smoothing for Unsupervised Sentence Representation
abstract
Nuo Chen, Linjun Shou, Jian Pei, Ming Gong, Bowen Cao, Jianhui Chang, Jia Li, Daxin Jiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Nuo Chen 0001, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Bowen Cao, Jianhui Chang, Jia Li 0009, Daxin Jiang
ACL (1)4
2023 NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation
abstract
Shengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, Jianlong Fu, Ming Gong, Lijuan Wang, Zicheng Liu, Houqiang Li, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shengming Yin, Chenfei Wu, Huan Yang 0005, Xiaodong Wang 0023, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Jianlong Fu, Ming Gong 0001, Zicheng Liu 0001, Houqiang Li, Nan Duan 0001
ACL (1)12
2023 RUEL: Retrieval-Augmented User Representation with Edge Browser Logs for Sequential Recommendation
abstract
Online recommender systems (RS) aim to match user needs with the vast amount of resources available on various platforms. A key challenge is to model user preferences accurately under the condition of data sparsity. To address this challenge, some methods have leveraged external user behavior data from multiple platforms to enrich user representation. However, all of these methods require a consistent user ID across platforms and ignore the information from similar users. In this study, we propose RUEL, a novel retrieval-based sequential recommender that can effectively incorporate external anonymous user behavior data from Edge browser logs to enhance recommendation. We first collect and preprocess a large volume of Edge browser logs over a one-year period and link them to target entities that correspond to candidate items in recommendation datasets. We then design a contrastive learning framework with a momentum encoder and a memory bank to retrieve the most relevant and diverse browsing sequences from the full browsing log based on the semantic similarity between user representations. After retrieval, we apply an item-level attentive selector to filter out noisy items and generate refined sequence embeddings for the final predictor. RUEL is the first method that connects user browsing data with typical recommendation datasets and can be generalized to various recommendation scenarios and datasets. We conduct extensive experiments on four real datasets for sequential recommendation tasks and demonstrate that RUEL significantly outperforms state-of-the-art baselines. We also conduct ablation studies and qualitative analysis to validate the effectiveness of each component of RUEL and provide additional insights into our method.
Ning Wu 0013, Ming Gong 0001, Linjun Shou, Jian Pei 0001, Daxin Jiang
CIKM2
2023 Instructed Language Models with Retrievers Are Powerful Entity Linkers
abstract
Generative approaches powered by large language models (LLMs) have demonstrated emergent abilities in tasks that require complex reasoning abilities.Yet the generative nature still makes the generated content suffer from hallucinations, thus unsuitable for entity-centric tasks like entity linking (EL) requiring precise entity predictions over a large knowledge base.We present Instructed Generative Entity Linker (INSGENEL), the first approach that enables casual language models to perform entity linking over knowledge bases.Several methods to equip language models with EL capability were proposed in this work, including (i) a sequence-to-sequence training EL objective with instruction-tuning, (ii) a novel generative EL framework based on a light-weight potential mention retriever that frees the model from heavy and non-parallelizable decoding, achieving 4× speedup without compromise on linking metrics.INSGENEL outperforms previous generative alternatives with +6.8 F1 points gain on average, also with a huge advantage in training data efficiency and training compute consumption.In addition, our skillfully engineered incontext learning (ICL) framework for EL still lags behind INSGENEL significantly, reaffirming that the EL task remains a persistent hurdle for general LLMs.
Zilin Xiao, Ming Gong 0001, Jie Wu 0018, Xingyao Zhang 0003, Linjun Shou, Daxin Jiang
EMNLP2
2023 Modeling Sequential Sentence Relation to Improve Cross-lingual Dense Retrieval
Shunyu Zhang, Yaobo Liang, Ming Gong 0001, Daxin Jiang, Nan Duan 0001
ICLR3
2023 Large Language Models are Diverse Role-Players for Summarization Evaluation
Ning Wu 0013, Ming Gong 0001, Linjun Shou, Shining Liang, Daxin Jiang
NLPCC (1)2
2023 Improving Readability for Automatic Speech Recognition Transcription
abstract
Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to grammatical errors, disfluency, and other noises common in spoken communication. These readable issues introduced by speakers and ASR systems will impair the performance of downstream tasks and the understanding of human readers. In this work, we present a task called ASR post-processing for readability (APR) and formulate it as a sequence-to-sequence text generation problem. The APR task aims to transform the noisy ASR output into a readable text for humans and downstream tasks while maintaining the semantic meaning of speakers. We further study the APR task from the benchmark dataset, evaluation metrics, and baseline models: First, to address the lack of task-specific data, we propose a method to construct a dataset for the APR task by using the data collected for grammatical error correction. Second, we utilize metrics adapted or borrowed from similar tasks to evaluate model performance on the APR task. Lastly, we use several typical or adapted pre-trained models as the baseline models for the APR task. Furthermore, we fine-tune the baseline models on the constructed dataset and compare their performance with a traditional pipeline method in terms of proposed evaluation metrics. Experimental results show that all the fine-tuned baseline models perform better than the traditional pipeline method, and our adapted RoBERTa model outperforms the pipeline method by 4.95 and 6.63 BLEU points on two test sets, respectively. The human evaluation and case study further reveal the ability of the proposed model to improve the readability of ASR transcripts.
Junwei Liao, Sefik Emre Eskimez, Liyang Lu, Yu Shi 0001, Ming Gong 0001, Linjun Shou, Hong Qu 0002, Michael Zeng 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2022 From Good to Best: Two-Stage Training for Cross-Lingual Machine Reading Comprehension
abstract
Cross-lingual Machine Reading Comprehension (xMRC) is a challenging task due to the lack of training data in low-resource languages. Recent approaches use training data only in a resource-rich language (such as English) to fine-tune large-scale cross-lingual pre-trained language models, which transfer knowledge from resource-rich languages (source) to low-resource languages (target). Due to the big difference between languages, the model fine-tuned only by the source language may not perform well for target languages. In our study, we make an interesting observation that while the top 1 result predicted by the previous approaches may often fail to hit the ground-truth answer, there are still good chances for the correct answer to be contained in the set of top k predicted results. Intuitively, the previous approaches have empowered the model certain level of capability to roughly distinguish good answers from bad ones. However, without sufficient training data, it is not powerful enough to capture the nuances between the accurate answer and those approximate ones. Based on this observation, we develop a two-stage approach to enhance the model performance. The first stage targets at recall; we design a hard-learning (HL) algorithm to maximize the likelihood that the top k predictions contain the accurate answer. The second stage focuses on precision, where an answer-aware contrastive learning (AA-CL) mechanism is developed to learn the minute difference between the accurate answer and other candidates. Extensive experiments show that our model significantly outperforms strong baselines on two cross-lingual MRC benchmark datasets.
Nuo Chen 0001, Linjun Shou, Ming Gong 0001, Jian Pei 0001
AAAI3
2022 Multi-View Document Representation Learning for Open-Domain Dense Retrieval
abstract
Dense retrieval has achieved impressive advances in first-stage retrieval from a largescale document collection, which is built on bi-encoder architecture to produce single vector representation of query and document.However, a document can usually answer multiple potential queries from different views.So the single vector representation of a document is hard to match with multi-view queries, and faces a semantic mismatch problem.This paper proposes a multi-view document representation learning framework, aiming to produce multiview embeddings to represent documents and enforce them to align with different queries.First, we propose a simple yet effective method of generating multiple embeddings through viewers.Second, to prevent multi-view embeddings from collapsing to the same one, we further propose a global-local loss with annealed temperature to encourage the multiple viewers to better align with different potential queries.Experiments show our method outperforms recent works and achieves state-of-the-art results. * Work done during internship at Microsoft Research Asia.Q1: Where can people using iPods on planes view the device's interface?A1: Individual seat-back displays.Q2: What are two airlines that considered implementing iPod connections but did not join the 2007 agreement?A2: KLM and Air France.
Shunyu Zhang, Yaobo Liang, Ming Gong 0001, Daxin Jiang, Nan Duan 0001
ACL (1)3
2022 Tiger: Transferable Interest Graph Embedding for Domain-Level Zero-Shot Recommendation
abstract
Recommender systems play a significant role in online services and have attracted wide attention from both academia and industry. In this paper, we focus on an important, practical, but often overlooked task: domain-level zero-shot recommendation (DZSR). The challenge of DZSR mainly lies in the absence of collaborative behaviors in the target domain, which may be caused by various reasons, such as the domain being newly launched without existing user-item interactions, or users' behaviors being too sensitive to collect for training. To address this challenge, we propose a Transferable Interest Graph Embedding technique for Recommendations (Tiger). The key idea is to connect isolated collaborative filtering datasets with a knowledge graph tailored to recommendations, then propagate collaborative signals from public domains to the zero-shot target domain. The backbone of Tiger is the transferable interest extractor, which is a simple yet effective graph convolutional network (GCN) aggregating multiple hops of neighbors on a shared interest graph. We find that the bottom layers of GCN preserve more domain-specific information while the upper layers represent universal interest better. Thus, in Tiger, we discard the bottom layers of GCN to reconstruct user interest so that collaborative signals can be successfully propagated to other domains, and retain the bottom layers of GCN to include domain-specific information for items. Extensive experiments with four public datasets demonstrate that Tiger can effectively make recommendations for a zero-shot domain and outperform several alternative baselines.
Jianhuan Zhuo, Jianxun Lian, Lanling Xu, Ming Gong 0001, Linjun Shou, Daxin Jiang, Xing Xie 0001, Yinliang Yue
CIKM4
2022 Label-aware Multi-level Contrastive Learning for Cross-lingual Spoken Language Understanding
abstract
Despite the great success of spoken language understanding (SLU) in high-resource languages, it remains challenging in low-resource languages mainly due to the lack of labeled training data.The recent multilingual codeswitching approach achieves better alignments of model representations across languages by constructing a mixed-language context in zeroshot cross-lingual SLU.However, current codeswitching methods are limited to implicit alignment and disregard the inherent semantic structure in SLU, i.e., the hierarchical inclusion of utterances, slots, and words.In this paper, we propose to model the utterance-slot-word structure by a multi-level contrastive learning framework at the utterance, slot, and word levels to facilitate explicit alignment.Novel codeswitching schemes are introduced to generate hard negative examples for our contrastive learning framework.Furthermore, we develop a label-aware joint model leveraging label semantics to enhance the implicit alignment and feed to contrastive learning.Our experimental results show that our proposed methods significantly improve the performance compared with the strong baselines on two zero-shot crosslingual SLU benchmark datasets.
Shining Liang, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Wanli Zuo, Xianglin Zuo, Daxin Jiang
EMNLP4
2022 Empowering Dual-Encoder with Query Generator for Cross-Lingual Dense Retrieval
abstract
In monolingual dense retrieval, lots of works focus on how to distill knowledge from crossencoder re-ranker to dual-encoder retriever and these methods achieve better performance due to the effectiveness of cross-encoder re-ranker.However, we find that the performance of the cross-encoder re-ranker is heavily influenced by the number of training samples and the quality of negative samples, which is hard to obtain in the cross-lingual setting.In this paper, we propose to use a query generator as the teacher in the cross-lingual setting, which is less dependent on enough training samples and highquality negative samples.In addition to traditional knowledge distillation, we further propose a novel enhancement method, which uses the query generator to help the dual-encoder align queries from different languages, but does not need any additional parallel sentences.The experimental results show that our method outperforms the state-of-the-art methods on two benchmark datasets.
Houxing Ren, Linjun Shou, Ning Wu 0013, Ming Gong 0001, Daxin Jiang
EMNLP4
2022 Unsupervised Context Aware Sentence Representation Pretraining for Multi-lingual Dense Retrieval
abstract
Recent research demonstrates the effectiveness of using pretrained language models (PLM) to improve dense retrieval and multilingual dense retrieval. In this work, we present a simple but effective monolingual pretraining task called contrastive context prediction (CCP) to learn sentence representation by modeling sentence level contextual relation. By pushing the embedding of sentences in a local context closer and pushing random negative samples away, different languages could form isomorphic structure, then sentence pairs in two different languages will be automatically aligned. Our experiments show that model collapse and information leakage are very easy to happen during contrastive training of language model, but language-specific memory bank and asymmetric batch normalization operation play an essential role in preventing collapsing and information leakage, respectively. Besides, a post-processing for sentence embedding is also very effective to achieve better retrieval performance. On the multilingual sentence retrieval task Tatoeba, our model achieves new SOTA results among methods without using bilingual data. Our model also shows larger gain on Tatoeba when transferring between non-English pairs. On two multi-lingual query-passage retrieval tasks, XOR Retrieve and Mr.TYDI, our model even achieves two SOTA results in both zero-shot and supervised setting among all pretraining models using bilingual data.
Ning Wu 0013, Yaobo Liang, Houxing Ren, Linjun Shou, Nan Duan 0001, Ming Gong 0001, Daxin Jiang
IJCAI6
2022 Bridging the Gap between Language Models and Cross-Lingual Sequence Labeling
abstract
Nuo Chen, Linjun Shou, Ming Gong, Jian Pei, Daxin Jiang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Nuo Chen 0001, Linjun Shou, Ming Gong 0001, Jian Pei 0001, Daxin Jiang
NAACL-HLT3
2022 Graph Fusion Network for Text Classification
Yong Dai 0001, Linjun Shou, Ming Gong 0001, Xiaolin Xia, Zhao Kang 0001, Zenglin Xu, Daxin Jiang
Knowl. Based Syst.3
2021 Reinforced Multi-Teacher Selection for Knowledge Distillation
abstract
In natural language processing (NLP) tasks, slow inference speed and huge footprints in GPU usage remain the bottleneck of applying pre-trained deep models in production. As a popular method for model compression, knowledge distillation transfers knowledge from one or multiple large (teacher) models to a small (student) model. When multiple teacher models are available in distillation, the state-of-the-art methods assign a fixed weight to a teacher model in the whole distillation. Furthermore, most of the existing methods allocate an equal weight to every teacher model. In this paper, we observe that, due to the complexity of training examples and the differences in student model capability, learning differentially from teacher models can lead to better performance of student models distilled. We systematically develop a reinforced method to dynamically assign weights to teacher models for different training instances and optimize the performance of student model. Our extensive experimental results on several NLP tasks clearly verify the feasibility and effectiveness of our approach.
Fei Yuan 0010, Linjun Shou, Jian Pei 0001, Wutao Lin, Ming Gong 0001, Daxin Jiang
AAAI5
2021 CoSQA: 20, 000+ Web Queries for Code Search and Question Answering
abstract
Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, Nan Duan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Junjie Huang 0008, Duyu Tang, Linjun Shou, Ming Gong 0001, Ke Xu 0001, Daxin Jiang, Ming Zhou 0001, Nan Duan 0001
ACL/IJCNLP (1)4
2021 Syntax-Enhanced Pre-trained Model
abstract
Zenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, Nan Duan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong 0001, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, Nan Duan 0001
ACL/IJCNLP (1)6
2021 Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language Understanding
abstract
Lack of training data presents a grand challenge to scaling out spoken language understanding (SLU) to low-resource languages.Although various data augmentation approaches have been proposed to synthesize training data in low-resource target languages, the augmented data sets are often noisy, and thus impede the performance of SLU models.In this paper we focus on mitigating noise in augmented data.We develop a denoising training approach.Multiple models are trained with data produced by various augmented methods.Those models provide supervision signals to each other.The experimental results show that our method outperforms the existing state of the art by 3.05 and 4.24 percentage points on two benchmark datasets, respectively.The code will be made open sourced on github.
Yingmei Guo, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Mingxing Xu, Zhiyong Wu 0001, Daxin Jiang
EMNLP (1)4
2021 Generating Human Readable Transcript for Automatic Speech Recognition with Pre-Trained Language Model
abstract
Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to disfluency, filter words, and other errata common in spoken communication. Many downstream tasks and human readers rely on the output of the ASR system; therefore, errors introduced by the speaker and ASR system alike will be propagated to the next task in the pipeline. In this work, we propose an ASR post-processing model that aims to transform the incorrect and noisy ASR output into a readable text for humans and downstream tasks. We leverage the Metadata Extraction (MDE) corpus to construct a task-specific dataset for our study. Since the dataset is small, we propose a novel data augmentation method and use a two-stage training strategy to fine-tune the RoBERTa pre-trained model. On the constructed test set, our model outperforms a production two-step pipeline-based post-processing method by a large margin of 13.26 on readability-aware WER (RA-WER) and 17.53 on BLEU metrics. Human evaluation also demonstrates that our method can generate more human-readable transcripts than the baseline method.
Junwei Liao, Yu Shi 0001, Ming Gong 0001, Linjun Shou, Sefik Emre Eskimez, Liyang Lu, Hong Qu 0002, Michael Zeng 0001
ICASSP3
2021 Improving Zero-shot Neural Machine Translation on Language-specific Encoders- Decoders
abstract
Recently, universal neural machine translation (NMT) with shared encoder-decoder gained good performance on zero-shot translation. Unlike universal NMT, jointly trained language-specific encoders-decoders aim to achieve universal representation across non-shared modules, each of which is for a language or language family. The non-shared architecture has the advantage of mitigating internal language competition, especially when the shared vocabulary and model parameters are restricted in their size. However, the performance of using multiple encoders and decoders on zero-shot translation still lags behind universal NMT. In this work, we study zero-shot translation using language-specific encoders-decoders. We propose to generalize the non-shared architecture and universal NMT by differentiating the Transformer layers between language-specific and interlingua. By selectively sharing parameters and applying cross-attentions, we explore maximizing the representation universality and realizing the best alignment of language-agnostic information. We also introduce a denoising auto-encoding (DAE) objective to jointly train the model with the translation task in a multi-task manner. Experiments on two public multilingual parallel datasets show that our proposed model achieves competitive or better results than universal NMT and the strong pivot baseline. Moreover, we experiment incrementally adding new language to the trained model by only updating the new model parameters. With this little effort, the zero-shot translation between this newly added language and existing languages achieves a comparable result with the model trained jointly from scratch on all languages.
Junwei Liao, Yu Shi 0001, Ming Gong 0001, Linjun Shou, Hong Qu 0002, Michael Zeng 0001
IJCNN3
2021 Reinforced Iterative Knowledge Distillation for Cross-Lingual Named Entity Recognition
abstract
Named entity recognition (NER) is a fundamental component in many applications, such as Web Search and Voice Assistants. Although deep neural networks greatly improve the performance of NER, due to the requirement of large amounts of training data, deep neural networks can hardly scale out to many languages in an industry setting. To tackle this challenge, cross-lingual NER transfers knowledge from a rich-resource language to languages with low resources through pre-trained multilingual language models. Instead of using training data in target languages, cross-lingual NER has to rely on only training data in source languages, and optionally adds the translated training data derived from source languages. However, the existing cross-lingual NER methods do not make good use of rich unlabeled data in target languages, which is relatively easy to collect in industry applications. To address the opportunities and challenges, in this paper we describe our novel practice in Microsoft to leverage such large amounts of unlabeled data in target languages in real production settings. To effectively extract weak supervision signals from the unlabeled data, we develop a novel approach based on the ideas of semi-supervised learning and reinforcement learning. The empirical study on three benchmark data sets verifies that our approach establishes the new state-of-the-art performance with clear edges. Now, the NER techniques reported in this paper are on their way to become a fundamental component for Web ranking, Entity Pane, Answers Triggering, and Question Answering in the Microsoft Bing search engine. Moreover, our techniques will also serve as part of the Spoken Language Understanding module for a commercial voice assistant. We plan to open source the code of the prototype framework after deployment.
Shining Liang, Ming Gong 0001, Jian Pei 0001, Linjun Shou, Wanli Zuo, Xianglin Zuo, Daxin Jiang
KDD2
2021 Language Scaling: Applications, Challenges and Approaches
abstract
Language scaling aims to deploy Natural Language Processing (NLP) applications economically across many countries/regions with different languages. Language scaling has been heavily invested by industry since many parties want to deploy their applications/services to global markets. At the same time, scaling out NLP applications to various languages, essentially a data science problem, remains a grand challenge due to the huge differences in the morphology, syntaxes, and pragmatics among different languages. We present a comprehensive survey and tutorial on language scaling. We start with a clear problem description for language scaling and an intuitive discussion on the overall challenges. Then, we outline two major categories of approaches to language scaling, namely, model transfer and data transfer. We present a taxonomy to summarize various methods in literature. A large part of the tutorial is organized to address various types of NLP applications. Finally, we discuss several important challenges in this area and future directions.
Linjun Shou, Ming Gong 0001, Jian Pei 0001, Xiubo Geng, Xingjie Zhou, Daxin Jiang
KDD2
2021 CalibreNet: Calibration Networks for Multilingual Sequence Labeling
abstract
Lack of training data in low-resource languages presents huge challenges to sequence labeling tasks such as named entity recognition (NER) and machine reading comprehension (MRC). One major obstacle is the errors on the boundary of predicted answers. To tackle this problem, we propose CalibreNet, which predicts answers in two steps. In the first step, any existing sequence labeling method can be adopted as a base model to generate an initial answer. In the second step, CalibreNet refines the boundary of the initial answer. To tackle the challenge of lack of training data in low-resource languages, we dedicatedly develop a novel unsupervised phrase boundary recovery pre-training task to enhance the multilingual boundary detection capability of CalibreNet. Experiments on two cross-lingual benchmark datasets show that the proposed approach achieves SOTA results on zero-shot cross-lingual NER and MRC tasks.
Shining Liang, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Wanli Zuo, Daxin Jiang
WSDM4
2020 Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training
abstract
We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM (Lample and Conneau 2019) and Unicoder (Huang et al. 2019), both visual and linguistic contents are fed into a multi-layer Transformer (Vaswani et al. 2017) for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling(MLM), Masked Object Classification(MOC) and Visual-linguistic Matching(VLM). The first two tasks learn context-aware representations for input tokens based on linguistic and visual contents jointly. The last task tries to predict whether an image and a text describe each other. After pretraining on large-scale image-caption pairs, we transfer Unicoder-VL to caption-based image-text retrieval and visual commonsense reasoning, with just one additional output layer. We achieve state-of-the-art or comparable results on both two tasks and show the powerful ability of the cross-modal pre-training.
Nan Duan 0001, Yuejian Fang, Ming Gong 0001, Daxin Jiang
AAAI4
2020 Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering
abstract
Commonsense question answering aims to answer questions which require background knowledge that is not explicitly expressed in the question. The key challenge is how to obtain evidence from external knowledge and make predictions based on the evidence. Recent studies either learn to generate evidence from human-annotated evidence which is expensive to collect, or extract evidence from either structured or unstructured knowledge bases which fails to take advantages of both sources simultaneously. In this work, we propose to automatically extract evidence from heterogeneous knowledge sources, and answer questions based on the extracted evidence. Specifically, we extract evidence from both structured knowledge base (i.e. ConceptNet) and Wikipedia plain texts. We construct graphs for both sources to obtain the relational structures of evidence. Based on these graphs, we propose a graph-based approach consisting of a graph-based contextual word representation learning module and a graph-based inference module. The first module utilizes graph structural information to re-define the distance between words for learning better contextual word representations. The second module adopts graph convolutional network to encode neighbor information into the representations of nodes, and aggregates evidence with graph attention mechanism for predicting the final answer. Experimental results on CommonsenseQA dataset illustrate that our graph-based approach over both knowledge sources brings improvement over strong baselines. Our approach achieves the state-of-the-art accuracy (75.3%) on the CommonsenseQA dataset.
Shangwen Lv, Daya Guo, Jingjing Xu 0001, Duyu Tang, Nan Duan 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Songlin Hu 0001
AAAI6
2020 Enhancing Answer Boundary Detection for Multilingual Machine Reading Comprehension
abstract
Multilingual pre-trained models could leverage the training data from a rich source language (such as English) to improve the performance on low resource languages.However, the transfer effectiveness on the multilingual Machine Reading Comprehension (MRC) task is substantially poorer than that for sentence classification tasks, mainly due to the requirement of MRC to detect the word level answer boundary.In this paper, we propose two auxiliary tasks to introduce additional phrase boundary supervision in the fine-tuning stage:(1) a mixed MRC task, which translates the question or passage to other languages and builds cross-lingual question-passage pairs; and (2) a language-agnostic knowledge masking task by leveraging knowledge phrases mined from the Web.Extensive experiments on two cross-lingual MRC datasets show the effectiveness of our proposed approach.† Random N-gram Masking shows gains in English SQuAD.
Fei Yuan 0010, Linjun Shou, Xuanyu Bai, Ming Gong 0001, Yaobo Liang, Nan Duan 0001, Daxin Jiang
ACL4
2020 LogicalFactChecker: Leveraging Logical Operations for Fact Checking with Graph Module Network
abstract
Wanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan, Ming Zhou, Ming Gong, Linjun Shou, Daxin Jiang, Jiahai Wang, Jian Yin. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Wanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan 0001, Ming Zhou 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Jiahai Wang, Jian Yin 0001
ACL6
2020 Cross-lingual Machine Reading Comprehension with Language Branch Knowledge Distillation
abstract
Cross-lingual Machine Reading Comprehension (CLMRC) remains a challenging problem due to the lack of large-scale annotated datasets in low-source languages, such as Arabic, Hindi, and Vietnamese.Many previous approaches use translation data by translating from a rich-source language, such as English, to low-source languages as auxiliary supervision.However, how to effectively leverage translation data and reduce the impact of noise introduced by translation remains onerous.In this paper, we tackle this challenge and enhance the cross-lingual transferring performance by a novel augmentation approach named Language Branch Machine Reading Comprehension (LBMRC).A language branch is a group of passages in one single language paired with questions in all target languages.We train multiple machine reading comprehension (MRC) models proficient in individual language based on LBMRC.Then, we devise a multilingual distillation approach to amalgamate knowledge from multiple language branch models to a single model for all target languages.Combining the LBMRC and multilingual distillation can be more robust to the data noises, therefore, improving the model's cross-lingual ability.Meanwhile, the produced single multilingual model can apply to all target languages, which saves the cost of training, inference, and maintenance for multiple models.Extensive experiments on two CLMRC benchmarks clearly show the effectiveness of our proposed method.
Junhao Liu 0001, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Min Yang 0007, Daxin Jiang
COLING4
2020 A Graph Representation of Semi-structured Data for Web Question Answering
abstract
The abundant semi-structured data on the Web, such as HTML-based tables and lists, provide commercial search engines a rich information source for question answering (QA).Different from plain text passages in Web documents, Web tables and lists have inherent structures, which carry semantic correlations among various elements in tables and lists.Many existing studies treat tables and lists as flat documents with pieces of text and do not make good use of semantic information hidden in structures.In this paper, we propose a novel graph representation of Web tables and lists based on a systematic categorization of the components in semi-structured data as well as their relations.We also develop pre-training and reasoning techniques on the graph model for the QA task.Extensive experiments on several real datasets collected from a commercial engine verify the effectiveness of our approach.Our method improves F1 score by 3.90 points over the state-of-the-art baselines.
Xingyao Zhang 0003, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Lijie Wen 0001, Daxin Jiang
COLING4
2020 XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and Generation
abstract
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Yaobo Liang, Nan Duan 0001, Yeyun Gong, Ning Wu 0013, Fenfei Guo, Weizhen Qi, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Dong Bo Cui, Sining Wei, Taroon Bharti, Jiun-Hung Chen, Winnie Wu, Fan Yang 0024, Daniel Campos, Rangan Majumder, Ming Zhou 0001
EMNLP (1)7
2020 Mining Implicit Relevance Feedback from User Behavior for Web Question Answering
abstract
Training and refreshing a web-scale Question Answering (QA) system for a multi-lingual commercial search engine often requires a huge amount of training examples. One principled idea is to mine implicit relevance feedback from user behavior recorded in search engine logs. All previous works on mining implicit relevance feedback target at relevance of web documents rather than passages. Due to several unique characteristics of QA tasks, the existing user behavior models for web documents cannot be applied to infer passage relevance. In this paper, we make the first study to explore the correlation between user behavior and passage relevance, and propose a novel approach for mining training data for Web QA. We conduct extensive experiments on four test datasets and the results show our approach significantly improves the accuracy of passage ranking without extra human labeled data. In practice, this work has proved effective to substantially reduce the human labeling cost for the QA service in a global commercial search engine, especially for languages with low resources. Our techniques have been deployed in multi-language services.
Linjun Shou, Shining Bo, Feixiang Cheng, Ming Gong 0001, Jian Pei 0001, Daxin Jiang
KDD4
2020 Model Compression with Two-stage Multi-teacher Knowledge Distillation for Web Question Answering System
abstract
Deep pre-training and fine-tuning models (such as BERT and OpenAI GPT) have demonstrated excellent results in question answering areas. However, due to the sheer amount of model parameters, the inference speed of these models is very slow. How to apply these complex models to real business scenarios becomes a challenging but practical problem. Previous model compression methods usually suffer from information loss during the model compression procedure, leading to inferior models compared with the original one. To tackle this challenge, we propose a Two-stage Multi-teacher Knowledge Distillation (TMKD for short) method for web Question Answering system. We first develop a general Q&A distillation task for student model pre-training, and further fine-tune this pre-trained student model with multi-teacher knowledge distillation on downstream tasks (like Web Q&A task, MNLI, SNLI, RTE tasks from GLUE), which effectively reduces the overfitting bias in individual teacher models, and transfers more general knowledge to the student model. The experiment results show that our method can significantly outperform the baseline methods and even achieve comparable results with the original teacher models, along with substantial speedup of model inference.
Ze Yang 0005, Linjun Shou, Ming Gong 0001, Wutao Lin, Daxin Jiang
WSDM3
2019 Joint Type Inference on Entities and Relations via Graph Convolutional Networks
abstract
We develop a new paradigm for the task of joint entity relation extraction.It first identifies entity spans, then performs a joint inference on entity types and relation types.To tackle the joint type inference task, we propose a novel graph convolutional network (GCN) running on an entity-relation bipartite graph.By introducing a binary relation classification task, we are able to utilize the structure of entity-relation bipartite graph in a more efficient and interpretable way.Experiments on ACE05 show that our model outperforms existing joint models in entity performance and is competitive with the state-of-the-art in relation performance.
Changzhi Sun, Yeyun Gong, Yuanbin Wu, Ming Gong 0001, Daxin Jiang, Man Lan, Shiliang Sun, Nan Duan 0001
ACL (1)4
2019 Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks
abstract
Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, Ming Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Haoyang Huang, Yaobo Liang, Nan Duan 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Ming Zhou 0001
EMNLP/IJCNLP (1)4
2017 Naturally combined shape-color moment invariants under affine transformations
Ming Gong 0001, You Hao, Hanlin Mo, Hua Li 0009
Comput. Vis. Image Underst.1
2013 Gaussian-curvature-derived invariants for isometry
Weiguo Cao, Ming Gong 0001, Hua Li 0009
Sci. China Inf. Sci.4
2013 Moment invariants to affine transformation of colours
Ming Gong 0001, Hua Li 0009, Weiguo Cao
Pattern Recognit. Lett.1
2011 A Kind of Shape-Color Moment Invariants
abstract
Shape or color based moment invariants are conventional pattern sensitive features in the object recognition and image description. However, the existing moment invariants are not robust because they handle simplified cases, such as single impact of shape transformation, single color transformation, and a simple combination of them. In this paper, we propose a kind of shape-color moment invariants (SCMIs), taking complicated shape and color transformation into account. It is applicable to more complicated cases, such as changing viewpoints, different illuminations, different camera white balance modes, and different color spaces. The integral-based SCMI construction framework is extensible and general. We evaluate the technique through comprehensive experiments.
Ming Gong 0001, Weiguo Cao, Hua Li 0009
CAD/Graphics1