VLDB 2026 Research / reviewers in the wild / expert
Hai Wang 0013
dblp:59/3767-13
· DBLP profile ↗
3ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Question answering and dialogue systems · 87% Information extraction and text analysis · 13% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
cloze-style question answering |
0.2 | 1 | 2016 | Who did What: A Large-Scale Person-Centered Cloze Dataset · EMNLP 2016 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.2 | 1 | 2016 | Who did What: A Large-Scale Person-Centered Cloze Dataset · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 1 | 2016 | Who did What: A Large-Scale Person-Centered Cloze Dataset · EMNLP 2016 |
Methods — techniques the papers use, named apart from their topics
multiple-choice question generation · 0.2dataset construction · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Improving Pre-Trained Multilingual Model with Vocabulary ExpansionabstractRecently, pre-trained language models have achieved remarkable success in a broad range of natural language processing tasks.However, in multilingual setting, it is extremely resource-consuming to pre-train a deep language model over large-scale corpora for each language.Instead of exhaustively pre-training monolingual language models independently, an alternative solution is to pre-train a powerful multilingual deep language model over large-scale corpora in hundreds of languages.However, the vocabulary size for each language in such a model is relatively small, especially for low-resource languages.This limitation inevitably hinders the performance of these multilingual models on tasks such as sequence labeling, wherein in-depth token-level or sentence-level understanding is essential.In this paper, inspired by previous methods designed for monolingual settings, we investigate two approaches (i.e., joint mapping and mixture mapping) based on a pre-trained multilingual model BERT for addressing the out-of-vocabulary (OOV) problem on a variety of tasks, including part-of-speech tagging, named entity recognition, machine translation quality estimation, and machine reading comprehension.Experimental results show that using mixture mapping is more promising.To the best of our knowledge, this is the first work that attempts to address and discuss the OOV issue in multilingual settings. Hai Wang 0013, Dian Yu 0001, Kai Sun 0006, Jianshu Chen, Dong Yu 0001 |
CoNLL | 1 |
| 2019 | Evidence Sentence Extraction for Machine Reading ComprehensionabstractRemarkable success has been achieved in the last few years on some limited machine reading comprehension (MRC) tasks.However, it is still difficult to interpret the predictions of existing MRC models.In this paper, we focus on extracting evidence sentences that can explain or support the answers of multiplechoice MRC tasks, where the majority of answer options cannot be directly extracted from reference documents.Due to the lack of ground truth evidence sentence labels in most cases, we apply distant supervision to generate imperfect labels and then use them to train an evidence sentence extractor.To denoise the noisy labels, we apply a recently proposed deep probabilistic logic learning framework to incorporate both sentence-level and cross-sentence linguistic indicators for indirect supervision.We feed the extracted evidence sentences into existing MRC models and evaluate the end-to-end performance on three challenging multiplechoice MRC datasets: MultiRC, RACE, and DREAM, achieving comparable or better performance than the same models that take as input the full reference document.To the best of our knowledge, this is the first work extracting evidence sentences for multiple-choice MRC. Hai Wang 0013, Dian Yu 0001, Kai Sun 0006, Jianshu Chen, Dong Yu 0001, David A. McAllester, Dan Roth 0001 |
CoNLL | 1 |
| 2016 | Who did What: A Large-Scale Person-Centered Cloze DatasetabstractWe have constructed a new "Who-did-What" dataset of over 200,000 fill-in-the-gap (cloze) multiple choice reading comprehension problems constructed from the LDC English Gigaword newswire corpus.The WDW dataset has a variety of novel features.First, in contrast with the CNN and Daily Mail datasets (Hermann et al., 2015) we avoid using article summaries for question formation.Instead, each problem is formed from two independent articles -an article given as the passage to be read and a separate article on the same events used to form the question.Second, we avoid anonymization -each choice is a person named entity.Third, the problems have been filtered to remove a fraction that are easily solved by simple baselines, while remaining 84% solvable by humans.We report performance benchmarks of standard systems and propose the WDW dataset as a challenge task for the community.1 Hai Wang 0013, Mohit Bansal, Kevin Gimpel, David A. McAllester |
EMNLP | 2 |