Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wenhao Liu 0003

dblp:86/8117-3 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
9since 2021 · last 2023
0009-0004-9828-6736ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Question answering and dialogue systems · 25% Language models and text generation · 23% Information extraction and text analysis · 19%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 50% Learning and educational technologies · 50%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%

Topics — the 25 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
text summarization
0.822023
HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder Models · EMNLP 2022
Marvista: Exploring the Design of a Human-AI Collaborative News Reading Tool · ACM Trans. Comput. Hum. Interact. 2023
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering
0.612022
QAConv: Question Answering on Informative Conversations · ACL (1) 2022
Natural language and speech › Information extraction and text analysis › fact-checking
evidence retrieval
0.612022
DialFact: A Benchmark for Fact-Checking in Dialogue · ACL (1) 2022
Natural language and speech › Information extraction and text analysis
fact-checking
0.612022
DialFact: A Benchmark for Fact-Checking in Dialogue · ACL (1) 2022
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
0.612022
Open Vocabulary Object Detection with Pseudo Bounding-Box Labels · ECCV (10) 2022
Natural language and speech › Information extraction and text analysis
text classification
0.612022
Conformal Predictor for Improving Zero-Shot Text Classification Efficiency · EMNLP 2022
Natural language and speech › Language models and text generation
text generation evaluation
0.612022
Near-Negative Distinction: Giving a Second Life to Human Evaluation Datasets · EMNLP 2022
Information retrieval › query suggestion
query auto-completion
0.512021
QueryBlazer: Efficient Query Autocompletion Framework · WSDM 2021
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization
0.412020
Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020
Machine learning › Transfer learning and domain adaptation
cross-task transfer
0.412020
Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems › dialogue understanding
dialogue act classification
0.412020
Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.412020
Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems › intent detection
few-shot intent detection
0.412020
Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems
intent detection
0.412020
Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.412020
Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.412020
Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020
Natural language and speech › Language models and text generation › text summarization
explainable summarization
0.212023
Marvista: Exploring the Design of a Human-AI Collaborative News Reading Tool · ACM Trans. Comput. Hum. Interact. 2023
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction
0.212022
Conformal Predictor for Improving Zero-Shot Text Classification Efficiency · EMNLP 2022
Machine learning › Deep learning architectures and training
mixture of experts
0.212022
HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder Models · EMNLP 2022
Machine learning › Trustworthy machine learning
uncertainty estimation
0.212022
Conformal Predictor for Improving Zero-Shot Text Classification Efficiency · EMNLP 2022
Information retrieval › evaluation
benchmark
0.212022
QAConv: Question Answering on Informative Conversations · ACL (1) 2022
Information retrieval
evaluation
0.212022
QAConv: Question Answering on Informative Conversations · ACL (1) 2022
Information retrieval › text summarization
summarization evaluation
0.212022
Near-Negative Distinction: Giving a Second Life to Human Evaluation Datasets · EMNLP 2022
Information retrieval
text summarization
0.212022
Near-Negative Distinction: Giving a Second Life to Human Evaluation Datasets · EMNLP 2022
Machine learning › Learning theory › classification › nonparametric classification
nearest neighbor classification
0.112020
Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

natural language processing · 1.3large language model · 1.3question generation · 1.1dialogue summarization · 1.1pseudo-labeling · 0.6next sentence prediction · 0.6natural language inference · 0.6multi-decoder models · 0.6mixture of experts · 0.6conformal predictor · 0.6subword tokenization · 0.5n-gram language model · 0.5
YearPublicationVenuePosition
2023 Marvista: Exploring the Design of a Human-AI Collaborative News Reading Tool
abstract
We explore the design of Marvista—a human-AI collaborative tool that employs a suite of natural language processing models to provide end-to-end support for reading online news articles. Before reading an article, Marvista helps a user plan what to read by filtering text based on how much time one can spend and what questions one is interested to find out from the article. During reading, Marvista helps the user reflect on their understanding of each paragraph with AI-generated questions. After reading, Marvista generates an explainable human-AI summary that combines AI’s processing of the text, the user’s reading behavior, and user-generated data in the reading process. In contrast to prior work that offered (content-independent) interaction techniques or devices for reading, Marvista takes a human-AI collaborative approach that contributes text-specific guidance (content-aware) to support the entire reading process.
Xiang 'Anthony' Chen, Chien-Sheng Wu, Lidiya Murakhovs'ka, Philippe Laban, Wenhao Liu 0003, Caiming Xiong
ACM Trans. Comput. Hum. Interact.6
2022 DialFact: A Benchmark for Fact-Checking in Dialogue
abstract
Fact-checking is an essential tool to mitigate the spread of misinformation and disinformation.We introduce the task of fact-checking in dialogue, which is a relatively unexplored area.We construct DIALFACT, a testing benchmark dataset of 22,245 annotated conversational claims, paired with pieces of evidence from Wikipedia.There are three sub-tasks in DIALFACT: 1) Verifiable claim detection task distinguishes whether a response carries verifiable factual information; 2) Evidence retrieval task retrieves the most relevant Wikipedia snippets as evidence; 3) Claim verification task predicts a dialogue response to be supported, refuted, or not enough information.We found that existing fact-checking models trained on non-dialogue data like FEVER (Thorne et al., 2018) fail to perform well on our task, and thus, we propose a simple yet data-efficient solution to effectively improve fact-checking performance in dialogue.We point out unique challenges in DIALFACT such as handling the colloquialisms, coreferences and retrieval ambiguities in the error analysis to shed light on future research in this direction 1 .Dialogue Context: I have family in Ireland!Have you ever been there?Evidence: Ireland is an island in the North Atlantic.Non-Verifiable Response: I haven't been but want to!Verifiable Supported Response: I haven't.It is an island in the north Atlantic right?Verifiable Refuted Response: I haven't been.Isn't it somewhere in north Pacific?Verifiable NEI Response: I haven't been.I heard it's the most popular tourist location in Europe!
Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong
ACL (1)3
2022 QAConv: Question Answering on Informative Conversations
abstract
This paper introduces QAConv, 1 , a new question answering (QA) dataset that uses conversations as a knowledge source.We focus on informative conversations, including business emails, panel discussions, and work channels.Unlike open-domain and task-oriented dialogues, these conversations are usually long, complex, asynchronous, and involve strong domain knowledge.In total, we collect 34,608 QA pairs from 10,259 selected conversations with both human-written and machinegenerated questions.We use a question generator and a dialogue summarizer as auxiliary tools to collect and recommend questions.The dataset has two testing scenarios: chunk mode and full mode, depending on whether the grounded partial conversation is provided or retrieved.Experimental results show that stateof-the-art pretrained QA systems have limited zero-shot performance and tend to predict our questions as unanswerable.Our dataset provides a new training and evaluation testbed to facilitate QA on conversations research.
Chien-Sheng Wu, Andrea Madotto, Wenhao Liu 0003, Pascale Fung, Caiming Xiong
ACL (1)3
2022 Open Vocabulary Object Detection with Pseudo Bounding-Box Labels
Mingfei Gao, Chen Xing, Juan Carlos Niebles, Junnan Li 0001, Ran Xu 0001, Wenhao Liu 0003, Caiming Xiong
ECCV (10)6
2022 Conformal Predictor for Improving Zero-Shot Text Classification Efficiency
abstract
Pre-trained language models (PLMs) have been shown effective for zero-shot (0shot) text classification.0shot models based on natural language inference (NLI) and next sentence prediction (NSP) employ cross-encoder architecture and infer by making a forward pass through the model for each label-text pair separately.This increases the computational cost to make inferences linearly in the number of labels.In this work, we improve the efficiency of such cross-encoder-based 0shot models by restricting the number of likely labels using another fast base classifier-based conformal predictor (CP) calibrated on samples labeled by the 0shot model.Since a CP generates prediction sets with coverage guarantees, it reduces the number of target labels without excluding the most probable label based on the 0shot model.We experiment with three intent and two topic classification datasets.With a suitable CP for each dataset, we reduce the average inference time for NLI-and NSP-based models by 25.6% and 22.2% respectively, without dropping performance below the predefined error rate of 1%.
Prafulla Kumar Choubey, Yu Bai 0017, Chien-Sheng Wu, Wenhao Liu 0003, Nazneen Fatema Rajani
EMNLP4
2022 HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder Models
abstract
Summarization systems make numerous "decisions" about summary properties during inference, e.g.degree of copying, specificity and length of outputs, etc.However, these are implicitly encoded within model parameters and specific styles cannot be enforced.To address this, we introduce HYDRASUM, a new summarization architecture that extends the single decoder framework of current models to a mixture-of-experts version with multiple decoders.We show that HYDRASUM's multiple decoders automatically learn contrasting summary styles when trained under the standard training objective without any extra supervision.Through experiments on three summarization datasets (CNN, NEWSROOM and XSUM), we show that HYDRASUM provides a simple mechanism to obtain stylistically-diverse summaries by sampling from either individual decoders or their mixtures, outperforming baseline models.Finally, we demonstrate that a small modification to the gating strategy during training can enforce an even stricter style partitioning, e.g.high-vs low-abstractiveness or high-vs low-specificity, allowing users to sample from a larger area in the generation space and vary summary styles along multiple dimensions. 1Input Article: Insights into the workings of the human body that Leonardo da Vinci could only obtain by dissecting scores of corpses and recording the results in exquisite drawings will be displayed for the first time beside modern 3D films, CT and MRI scans, which show how close the Renaissance genius got to the truth of what lies under the skin.[…] the Edinburgh show will be the first to compare Leonardo's results with scalpel and pen with the best results of modern technology.[…] The exhibition will show how close Leonardo got in some of his last medical experiments to discovering the role of the beating heart in the circulation of the blood, a century before William Harvey worked it out.[…] Edinburgh show will be first to compare Renaissance genius's results with best results of modern technology.Edinburgh show will be first to compare Renaissance genius's results with the best results of modern technology.Edinburgh show will be first to compare Leonardo's results with the best results of modern technology. Low diversity Baseline BARTEdinburgh show will be first to compare Leonardo's results with best results of modern technology.Modern imaging techniques will be displayed alongside Leonardo da Vinci's anatomical drawings in Edinburgh exhibition.
Tanya Goyal, Nazneen Fatema Rajani, Wenhao Liu 0003, Wojciech Kryscinski
EMNLP3
2022 Near-Negative Distinction: Giving a Second Life to Human Evaluation Datasets
abstract
Precisely assessing the progress in natural language generation (NLG) tasks is challenging, and human evaluation to establish a preference in a model's output over another is often necessary.However, human evaluation is usually costly, difficult to reproduce, and non-reusable.In this paper, we propose a new and simple automatic evaluation method for NLG called Near-Negative Distinction (NND) that repurposes prior human annotations into NND tests.In an NND test, an NLG model must place a higher likelihood on a high-quality output candidate than on a near-negative candidate with a known error.Model performance is established by the number of NND tests a model passes, as well as the distribution over task-specific errors the model fails on.Through experiments on three NLG tasks (question generation, question answering, and summarization), we show that NND achieves a higher correlation with human judgments than standard NLG evaluation metrics.We then illustrate NND evaluation in four practical scenarios, for example performing fine-grain model analysis, or studying model training dynamics.Our findings suggest that NND can give a second life to human annotations and provide low-cost NLG evaluation.
Philippe Laban, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong
EMNLP3
2022 QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization
abstract
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, Caiming Xiong. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Alexander R. Fabbri, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong
NAACL-HLT3
2021 QueryBlazer: Efficient Query Autocompletion Framework
abstract
Query autocompletion is an essential feature in search engines that predicts and suggests query completions to a user's incomplete prefix input, a critical feature to enhance the user experience. While a generic lookup-based system can provide completions with great efficiency, it is unable to address prefixes not seen in the past. On the other hand, a generative system can complete unseen queries with superior accuracy but requires substantial computational overhead at runtime, making it costly for a large-scale system. Here, we present an efficient, fully-generative query autocompletion framework. Our framework employs an n-gram language model at a subword-level and exploits the n-gram model's inherent data structure to precompute completions prior to runtime. Evaluation results on public dataset show that our framework is not only as effective as previous systems with neural language models, but also reduces computational overhead at runtime, expediting the speed by more than two orders of magnitude. The goal of this work is to showcase a generative query completion system that is an attractive choice for large-scale deployments.
Young Mo Kang, Wenhao Liu 0003, Yingbo Zhou 0002
WSDM2
2020 Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging
abstract
The concept of Dialogue Act (DA) is universal across different task-oriented dialogue domains -the act of "request" carries the same speaker intention whether it is for restaurant reservation or flight booking.However, DA taggers trained on one domain do not generalize well to other domains, which leaves us with the expensive need for a large amount of annotated data in the target domain.In this work, we investigate how to better adapt DA taggers to desired target domains with only unlabeled data.We propose MASKAUGMENT, a controllable mechanism that augments text input by leveraging the pre-trained MASK token from BERT model.Inspired by consistency regularization, we use MASKAUGMENT to introduce an unsupervised teacher-student learning scheme to examine the domain adaptation of DA taggers.Our extensive experiments on the Simulated Dialogue (GSim) and Schema-Guided Dialogue (SGD) datasets show that MASKAUGMENT is useful in improving the cross-domain generalization for DA tagging.
Semih Yavuz, Kazuma Hashimoto, Wenhao Liu 0003, Nitish Shirish Keskar, Richard Socher, Caiming Xiong
EMNLP (1)3
2020 Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference
abstract
Jianguo Zhang, Kazuma Hashimoto, Wenhao Liu, Chien-Sheng Wu, Yao Wan, Philip Yu, Richard Socher, Caiming Xiong. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jianguo Zhang 0005, Kazuma Hashimoto, Wenhao Liu 0003, Chien-Sheng Wu, Yao Wan 0001, Philip S. Yu, Richard Socher, Caiming Xiong
EMNLP (1)3