EDBT 2026 Demo / reviewers in the wild / expert
Wenhao Liu 0003
dblp:86/8117-3
· DBLP profile ↗
11ranked-venue papers
0as first author
9since 2021 · last 2023
0009-0004-9828-6736ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Question answering and dialogue systems · 25% Language models and text generation · 23% Information extraction and text analysis · 19% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 50% Learning and educational technologies · 50% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 100% |
Topics — the 25 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
text summarization |
0.8 | 2 | 2023 | HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder Models · EMNLP 2022 Marvista: Exploring the Design of a Human-AI Collaborative News Reading Tool · ACM Trans. Comput. Hum. Interact. 2023 |
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering |
0.6 | 1 | 2022 | QAConv: Question Answering on Informative Conversations · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis › fact-checking
evidence retrieval |
0.6 | 1 | 2022 | DialFact: A Benchmark for Fact-Checking in Dialogue · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis
fact-checking |
0.6 | 1 | 2022 | DialFact: A Benchmark for Fact-Checking in Dialogue · ACL (1) 2022 |
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection |
0.6 | 1 | 2022 | Open Vocabulary Object Detection with Pseudo Bounding-Box Labels · ECCV (10) 2022 |
Natural language and speech › Information extraction and text analysis
text classification |
0.6 | 1 | 2022 | Conformal Predictor for Improving Zero-Shot Text Classification Efficiency · EMNLP 2022 |
Natural language and speech › Language models and text generation
text generation evaluation |
0.6 | 1 | 2022 | Near-Negative Distinction: Giving a Second Life to Human Evaluation Datasets · EMNLP 2022 |
Information retrieval › query suggestion
query auto-completion |
0.5 | 1 | 2021 | QueryBlazer: Efficient Query Autocompletion Framework · WSDM 2021 |
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization |
0.4 | 1 | 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020 |
Machine learning › Transfer learning and domain adaptation
cross-task transfer |
0.4 | 1 | 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › dialogue understanding
dialogue act classification |
0.4 | 1 | 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.4 | 1 | 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › intent detection
few-shot intent detection |
0.4 | 1 | 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems
intent detection |
0.4 | 1 | 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.4 | 1 | 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.4 | 1 | 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text summarization
explainable summarization |
0.2 | 1 | 2023 | Marvista: Exploring the Design of a Human-AI Collaborative News Reading Tool · ACM Trans. Comput. Hum. Interact. 2023 |
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
0.2 | 1 | 2022 | Conformal Predictor for Improving Zero-Shot Text Classification Efficiency · EMNLP 2022 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.2 | 1 | 2022 | HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder Models · EMNLP 2022 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.2 | 1 | 2022 | Conformal Predictor for Improving Zero-Shot Text Classification Efficiency · EMNLP 2022 |
Information retrieval › evaluation
benchmark |
0.2 | 1 | 2022 | QAConv: Question Answering on Informative Conversations · ACL (1) 2022 |
Information retrieval
evaluation |
0.2 | 1 | 2022 | QAConv: Question Answering on Informative Conversations · ACL (1) 2022 |
Information retrieval › text summarization
summarization evaluation |
0.2 | 1 | 2022 | Near-Negative Distinction: Giving a Second Life to Human Evaluation Datasets · EMNLP 2022 |
Information retrieval
text summarization |
0.2 | 1 | 2022 | Near-Negative Distinction: Giving a Second Life to Human Evaluation Datasets · EMNLP 2022 |
Machine learning › Learning theory › classification › nonparametric classification
nearest neighbor classification |
0.1 | 1 | 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
natural language processing · 1.3large language model · 1.3question generation · 1.1dialogue summarization · 1.1pseudo-labeling · 0.6next sentence prediction · 0.6natural language inference · 0.6multi-decoder models · 0.6mixture of experts · 0.6conformal predictor · 0.6subword tokenization · 0.5n-gram language model · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Marvista: Exploring the Design of a Human-AI Collaborative News Reading ToolabstractWe explore the design of Marvista—a human-AI collaborative tool that employs a suite of natural language processing models to provide end-to-end support for reading online news articles. Before reading an article, Marvista helps a user plan what to read by filtering text based on how much time one can spend and what questions one is interested to find out from the article. During reading, Marvista helps the user reflect on their understanding of each paragraph with AI-generated questions. After reading, Marvista generates an explainable human-AI summary that combines AI’s processing of the text, the user’s reading behavior, and user-generated data in the reading process. In contrast to prior work that offered (content-independent) interaction techniques or devices for reading, Marvista takes a human-AI collaborative approach that contributes text-specific guidance (content-aware) to support the entire reading process. Xiang 'Anthony' Chen, Chien-Sheng Wu, Lidiya Murakhovs'ka, Philippe Laban, Wenhao Liu 0003, Caiming Xiong |
ACM Trans. Comput. Hum. Interact. | 6 |
| 2022 | DialFact: A Benchmark for Fact-Checking in DialogueabstractFact-checking is an essential tool to mitigate the spread of misinformation and disinformation.We introduce the task of fact-checking in dialogue, which is a relatively unexplored area.We construct DIALFACT, a testing benchmark dataset of 22,245 annotated conversational claims, paired with pieces of evidence from Wikipedia.There are three sub-tasks in DIALFACT: 1) Verifiable claim detection task distinguishes whether a response carries verifiable factual information; 2) Evidence retrieval task retrieves the most relevant Wikipedia snippets as evidence; 3) Claim verification task predicts a dialogue response to be supported, refuted, or not enough information.We found that existing fact-checking models trained on non-dialogue data like FEVER (Thorne et al., 2018) fail to perform well on our task, and thus, we propose a simple yet data-efficient solution to effectively improve fact-checking performance in dialogue.We point out unique challenges in DIALFACT such as handling the colloquialisms, coreferences and retrieval ambiguities in the error analysis to shed light on future research in this direction 1 .Dialogue Context: I have family in Ireland!Have you ever been there?Evidence: Ireland is an island in the North Atlantic.Non-Verifiable Response: I haven't been but want to!Verifiable Supported Response: I haven't.It is an island in the north Atlantic right?Verifiable Refuted Response: I haven't been.Isn't it somewhere in north Pacific?Verifiable NEI Response: I haven't been.I heard it's the most popular tourist location in Europe! Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong |
ACL (1) | 3 |
| 2022 | QAConv: Question Answering on Informative ConversationsabstractThis paper introduces QAConv, 1 , a new question answering (QA) dataset that uses conversations as a knowledge source.We focus on informative conversations, including business emails, panel discussions, and work channels.Unlike open-domain and task-oriented dialogues, these conversations are usually long, complex, asynchronous, and involve strong domain knowledge.In total, we collect 34,608 QA pairs from 10,259 selected conversations with both human-written and machinegenerated questions.We use a question generator and a dialogue summarizer as auxiliary tools to collect and recommend questions.The dataset has two testing scenarios: chunk mode and full mode, depending on whether the grounded partial conversation is provided or retrieved.Experimental results show that stateof-the-art pretrained QA systems have limited zero-shot performance and tend to predict our questions as unanswerable.Our dataset provides a new training and evaluation testbed to facilitate QA on conversations research. Chien-Sheng Wu, Andrea Madotto, Wenhao Liu 0003, Pascale Fung, Caiming Xiong |
ACL (1) | 3 |
| 2022 | Open Vocabulary Object Detection with Pseudo Bounding-Box Labels
Mingfei Gao, Chen Xing, Juan Carlos Niebles, Junnan Li 0001, Ran Xu 0001, Wenhao Liu 0003, Caiming Xiong |
ECCV (10) | 6 |
| 2022 | Conformal Predictor for Improving Zero-Shot Text Classification EfficiencyabstractPre-trained language models (PLMs) have been shown effective for zero-shot (0shot) text classification.0shot models based on natural language inference (NLI) and next sentence prediction (NSP) employ cross-encoder architecture and infer by making a forward pass through the model for each label-text pair separately.This increases the computational cost to make inferences linearly in the number of labels.In this work, we improve the efficiency of such cross-encoder-based 0shot models by restricting the number of likely labels using another fast base classifier-based conformal predictor (CP) calibrated on samples labeled by the 0shot model.Since a CP generates prediction sets with coverage guarantees, it reduces the number of target labels without excluding the most probable label based on the 0shot model.We experiment with three intent and two topic classification datasets.With a suitable CP for each dataset, we reduce the average inference time for NLI-and NSP-based models by 25.6% and 22.2% respectively, without dropping performance below the predefined error rate of 1%. Prafulla Kumar Choubey, Yu Bai 0017, Chien-Sheng Wu, Wenhao Liu 0003, Nazneen Fatema Rajani |
EMNLP | 4 |
| 2022 | HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder ModelsabstractSummarization systems make numerous "decisions" about summary properties during inference, e.g.degree of copying, specificity and length of outputs, etc.However, these are implicitly encoded within model parameters and specific styles cannot be enforced.To address this, we introduce HYDRASUM, a new summarization architecture that extends the single decoder framework of current models to a mixture-of-experts version with multiple decoders.We show that HYDRASUM's multiple decoders automatically learn contrasting summary styles when trained under the standard training objective without any extra supervision.Through experiments on three summarization datasets (CNN, NEWSROOM and XSUM), we show that HYDRASUM provides a simple mechanism to obtain stylistically-diverse summaries by sampling from either individual decoders or their mixtures, outperforming baseline models.Finally, we demonstrate that a small modification to the gating strategy during training can enforce an even stricter style partitioning, e.g.high-vs low-abstractiveness or high-vs low-specificity, allowing users to sample from a larger area in the generation space and vary summary styles along multiple dimensions. 1Input Article: Insights into the workings of the human body that Leonardo da Vinci could only obtain by dissecting scores of corpses and recording the results in exquisite drawings will be displayed for the first time beside modern 3D films, CT and MRI scans, which show how close the Renaissance genius got to the truth of what lies under the skin.[…] the Edinburgh show will be the first to compare Leonardo's results with scalpel and pen with the best results of modern technology.[…] The exhibition will show how close Leonardo got in some of his last medical experiments to discovering the role of the beating heart in the circulation of the blood, a century before William Harvey worked it out.[…] Edinburgh show will be first to compare Renaissance genius's results with best results of modern technology.Edinburgh show will be first to compare Renaissance genius's results with the best results of modern technology.Edinburgh show will be first to compare Leonardo's results with the best results of modern technology. Low diversity Baseline BARTEdinburgh show will be first to compare Leonardo's results with best results of modern technology.Modern imaging techniques will be displayed alongside Leonardo da Vinci's anatomical drawings in Edinburgh exhibition. Tanya Goyal, Nazneen Fatema Rajani, Wenhao Liu 0003, Wojciech Kryscinski |
EMNLP | 3 |
| 2022 | Near-Negative Distinction: Giving a Second Life to Human Evaluation DatasetsabstractPrecisely assessing the progress in natural language generation (NLG) tasks is challenging, and human evaluation to establish a preference in a model's output over another is often necessary.However, human evaluation is usually costly, difficult to reproduce, and non-reusable.In this paper, we propose a new and simple automatic evaluation method for NLG called Near-Negative Distinction (NND) that repurposes prior human annotations into NND tests.In an NND test, an NLG model must place a higher likelihood on a high-quality output candidate than on a near-negative candidate with a known error.Model performance is established by the number of NND tests a model passes, as well as the distribution over task-specific errors the model fails on.Through experiments on three NLG tasks (question generation, question answering, and summarization), we show that NND achieves a higher correlation with human judgments than standard NLG evaluation metrics.We then illustrate NND evaluation in four practical scenarios, for example performing fine-grain model analysis, or studying model training dynamics.Our findings suggest that NND can give a second life to human annotations and provide low-cost NLG evaluation. Philippe Laban, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong |
EMNLP | 3 |
| 2022 | QAFactEval: Improved QA-Based Factual Consistency Evaluation for SummarizationabstractAlexander Fabbri, Chien-Sheng Wu, Wenhao Liu, Caiming Xiong. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Alexander R. Fabbri, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong |
NAACL-HLT | 3 |
| 2021 | QueryBlazer: Efficient Query Autocompletion FrameworkabstractQuery autocompletion is an essential feature in search engines that predicts and suggests query completions to a user's incomplete prefix input, a critical feature to enhance the user experience. While a generic lookup-based system can provide completions with great efficiency, it is unable to address prefixes not seen in the past. On the other hand, a generative system can complete unseen queries with superior accuracy but requires substantial computational overhead at runtime, making it costly for a large-scale system. Here, we present an efficient, fully-generative query autocompletion framework. Our framework employs an n-gram language model at a subword-level and exploits the n-gram model's inherent data structure to precompute completions prior to runtime. Evaluation results on public dataset show that our framework is not only as effective as previous systems with neural language models, but also reduces computational overhead at runtime, expediting the speed by more than two orders of magnitude. The goal of this work is to showcase a generative query completion system that is an attractive choice for large-scale deployments. Young Mo Kang, Wenhao Liu 0003, Yingbo Zhou 0002 |
WSDM | 2 |
| 2020 | Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act TaggingabstractThe concept of Dialogue Act (DA) is universal across different task-oriented dialogue domains -the act of "request" carries the same speaker intention whether it is for restaurant reservation or flight booking.However, DA taggers trained on one domain do not generalize well to other domains, which leaves us with the expensive need for a large amount of annotated data in the target domain.In this work, we investigate how to better adapt DA taggers to desired target domains with only unlabeled data.We propose MASKAUGMENT, a controllable mechanism that augments text input by leveraging the pre-trained MASK token from BERT model.Inspired by consistency regularization, we use MASKAUGMENT to introduce an unsupervised teacher-student learning scheme to examine the domain adaptation of DA taggers.Our extensive experiments on the Simulated Dialogue (GSim) and Schema-Guided Dialogue (SGD) datasets show that MASKAUGMENT is useful in improving the cross-domain generalization for DA tagging. Semih Yavuz, Kazuma Hashimoto, Wenhao Liu 0003, Nitish Shirish Keskar, Richard Socher, Caiming Xiong |
EMNLP (1) | 3 |
| 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language InferenceabstractJianguo Zhang, Kazuma Hashimoto, Wenhao Liu, Chien-Sheng Wu, Yao Wan, Philip Yu, Richard Socher, Caiming Xiong. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Jianguo Zhang 0005, Kazuma Hashimoto, Wenhao Liu 0003, Chien-Sheng Wu, Yao Wan 0001, Philip S. Yu, Richard Socher, Caiming Xiong |
EMNLP (1) | 3 |