Hai Wang 0013

dblp:59/3767-13 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Question answering and dialogue systems · 87% Information extraction and text analysis · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
cloze-style question answering
0.212016
Who did What: A Large-Scale Person-Centered Cloze Dataset · EMNLP 2016
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.212016
Who did What: A Large-Scale Person-Centered Cloze Dataset · EMNLP 2016
Natural language and speech › Information extraction and text analysis
named entity recognition
0.112016
Who did What: A Large-Scale Person-Centered Cloze Dataset · EMNLP 2016

Methods — techniques the papers use, named apart from their topics

multiple-choice question generation · 0.2dataset construction · 0.2
YearPublicationVenuePosition
2019 Improving Pre-Trained Multilingual Model with Vocabulary Expansion
abstract
Recently, pre-trained language models have achieved remarkable success in a broad range of natural language processing tasks.However, in multilingual setting, it is extremely resource-consuming to pre-train a deep language model over large-scale corpora for each language.Instead of exhaustively pre-training monolingual language models independently, an alternative solution is to pre-train a powerful multilingual deep language model over large-scale corpora in hundreds of languages.However, the vocabulary size for each language in such a model is relatively small, especially for low-resource languages.This limitation inevitably hinders the performance of these multilingual models on tasks such as sequence labeling, wherein in-depth token-level or sentence-level understanding is essential.In this paper, inspired by previous methods designed for monolingual settings, we investigate two approaches (i.e., joint mapping and mixture mapping) based on a pre-trained multilingual model BERT for addressing the out-of-vocabulary (OOV) problem on a variety of tasks, including part-of-speech tagging, named entity recognition, machine translation quality estimation, and machine reading comprehension.Experimental results show that using mixture mapping is more promising.To the best of our knowledge, this is the first work that attempts to address and discuss the OOV issue in multilingual settings.
Hai Wang 0013, Dian Yu 0001, Kai Sun 0006, Jianshu Chen, Dong Yu 0001
CoNLL1
2019 Evidence Sentence Extraction for Machine Reading Comprehension
abstract
Remarkable success has been achieved in the last few years on some limited machine reading comprehension (MRC) tasks.However, it is still difficult to interpret the predictions of existing MRC models.In this paper, we focus on extracting evidence sentences that can explain or support the answers of multiplechoice MRC tasks, where the majority of answer options cannot be directly extracted from reference documents.Due to the lack of ground truth evidence sentence labels in most cases, we apply distant supervision to generate imperfect labels and then use them to train an evidence sentence extractor.To denoise the noisy labels, we apply a recently proposed deep probabilistic logic learning framework to incorporate both sentence-level and cross-sentence linguistic indicators for indirect supervision.We feed the extracted evidence sentences into existing MRC models and evaluate the end-to-end performance on three challenging multiplechoice MRC datasets: MultiRC, RACE, and DREAM, achieving comparable or better performance than the same models that take as input the full reference document.To the best of our knowledge, this is the first work extracting evidence sentences for multiple-choice MRC.
Hai Wang 0013, Dian Yu 0001, Kai Sun 0006, Jianshu Chen, Dong Yu 0001, David A. McAllester, Dan Roth 0001
CoNLL1
2016 Who did What: A Large-Scale Person-Centered Cloze Dataset
abstract
We have constructed a new "Who-did-What" dataset of over 200,000 fill-in-the-gap (cloze) multiple choice reading comprehension problems constructed from the LDC English Gigaword newswire corpus.The WDW dataset has a variety of novel features.First, in contrast with the CNN and Daily Mail datasets (Hermann et al., 2015) we avoid using article summaries for question formation.Instead, each problem is formed from two independent articles -an article given as the passage to be read and a separate article on the same events used to form the question.Second, we avoid anonymization -each choice is a person named entity.Third, the problems have been filtered to remove a fraction that are easily solved by simple baselines, while remaining 84% solvable by humans.We report performance benchmarks of standard systems and propose the WDW dataset as a challenge task for the community.1
Hai Wang 0013, Mohit Bansal, Kevin Gimpel, David A. McAllester
EMNLP2