VLDB 2026 Research / reviewers in the wild / expert
Yizhu Liu
dblp:219/0670
· DBLP profile ↗
9ranked-venue papers
6as first author
7since 2021 · last 2023
0000-0002-7241-120XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 74% Information extraction and text analysis · 12% Learning paradigms · 9% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
text summarization |
1.8 | 3 | 2023 | Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language Model · EMNLP 2023 Opinion Summarization by Weak-Supervision from Mix-structured Data · EMNLP 2022 Length Control in Abstractive Summarization by Pretraining Information Selection · ACL (1) 2022 |
Natural language and speech › Language models and text generation › text summarization
low-resource summarization |
0.9 | 2 | 2022 | Length Control in Abstractive Summarization by Pretraining Information Selection · ACL (1) 2022 Controlling Length in Abstractive Summarization Using a Convolutional Neural Network · EMNLP 2018 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.8 | 2 | 2021 | Keyword-aware Abstractive Summarization by Extracting Set-level Intermediate Summaries · WWW 2021 Controlling Length in Abstractive Summarization Using a Convolutional Neural Network · EMNLP 2018 |
Machine learning › Learning paradigms
curriculum learning |
0.7 | 1 | 2023 | In-sample Curriculum Learning by Sequence Completion for Natural Language Generation · ACL (1) 2023 |
Natural language and speech › Language models and text generation › evaluation of language models
faithfulness evaluation |
0.7 | 1 | 2023 | Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language Model · EMNLP 2023 |
Natural language and speech › Language models and text generation
text generation |
0.7 | 1 | 2023 | In-sample Curriculum Learning by Sequence Completion for Natural Language Generation · ACL (1) 2023 |
Natural language and speech › Language models and text generation › text summarization
opinion summarization |
0.6 | 1 | 2022 | Opinion Summarization by Weak-Supervision from Mix-structured Data · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis
keyphrase extraction |
0.5 | 1 | 2021 | Keyword-aware Abstractive Summarization by Extracting Set-level Intermediate Summaries · WWW 2021 |
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse relation recognition |
0.4 | 1 | 2020 | Multi-turn Response Selection using Dialogue Dependency Relations · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems
response selection |
0.4 | 1 | 2020 | Multi-turn Response Selection using Dialogue Dependency Relations · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
pre-training · 1.1sequence completion · 0.7probability change measurement · 0.7weak supervision · 0.6length-aware attention mechanism · 0.6aspect-based sentiment analysis · 0.6reinforcement learning · 0.5extractor-abstractor framework · 0.5pre-trained transformer · 0.4attention · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | In-sample Curriculum Learning by Sequence Completion for Natural Language GenerationabstractCurriculum learning has shown promising improvements in multiple domains by training machine learning models from easy samples to hard ones.Previous works which either design rules or train models for scoring the difficulty highly rely on task-specific expertise, and cannot generalize.Inspired by the "easy-to-hard" intuition, we propose to do in-sample curriculum learning for natural language generation tasks.Our learning strategy starts training the model to generate the last few words, i.e., do sequence completion, and gradually extends to generate the whole output sequence.Comprehensive experiments show that it generalizes well to different tasks and achieves significant improvements over strong baselines. Qi Jia 0003, Yizhu Liu, Haifeng Tang, Kenny Q. Zhu |
ACL (1) | 2 |
| 2023 | Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language ModelabstractDespite tremendous improvements in natural language generation, summarization models still suffer from the unfaithfulness issue.Previous work evaluates faithfulness either using models trained on the other tasks or in-domain synthetic data, or prompting a large model such as ChatGPT.This paper proposes to do zero-shot faithfulness evaluation simply with a moderately-sized foundation language model.We introduce a new metric FFLM, which is a combination of probability changes based on the intuition that prefixing a piece of text that is consistent with the output will increase the probability of predicting the output.Experiments show that FFLM performs competitively with or even outperforms ChatGPT on both inconsistency detection and faithfulness rating with 24x fewer parameters.FFLM also achieves improvements over other strong baselines. Qi Jia 0003, Yizhu Liu, Kenny Q. Zhu |
EMNLP | 3 |
| 2023 | Reducing repetition in convolutional abstractive summarizationabstractAbstract Convolutional sequence to sequence (CNN seq2seq) models have met success in abstractive summarization. However, their outputs often contain repetitive word sequences and logical inconsistencies, limiting the practicality of their application. In this paper, we find the reasons behind the repetition problem in CNN-based abstractive summarization through observing the attention map between the summaries with repetition and their corresponding source documents and mitigate the repetition problem. We propose to reduce the repetition in summaries by attention filter mechanism (ATTF) and sentence-level backtracking decoder (SBD), which dynamically redistributes attention over the input sequence as the output sentences are generated. The ATTF can record previously attended locations in the source document directly and prevent the decoder from attending to these locations. The SBD prevents the decoder from generating similar sentences more than once via backtracking at test. The proposed model outperforms the baselines in terms of ROUGE score, repeatedness, and readability. The results show that this approach generates high-quality summaries with minimal repetition and makes the reading experience better. Yizhu Liu, Xusheng Luo, Kenny Q. Zhu |
Nat. Lang. Eng. | 1 |
| 2022 | Length Control in Abstractive Summarization by Pretraining Information SelectionabstractPrevious length-controllable summarization models mostly control lengths at the decoding stage, whereas the encoding or the selection of information from the source document is not sensitive to the designed length.They also tend to generate summaries as long as those in the training data.In this paper, we propose a length-aware attention mechanism (LAAM) to adapt the encoding of the source based on the desired length.Our approach works by training LAAM on a summary length balanced dataset built from the original training data, and then fine-tuning as usual.Results show that this approach is effective in generating high-quality summaries with desired lengths and even those short lengths never seen in the original training set. Yizhu Liu, Qi Jia 0003, Kenny Q. Zhu |
ACL (1) | 1 |
| 2022 | Opinion Summarization by Weak-Supervision from Mix-structured DataabstractOpinion summarization of multiple reviews suffers from the lack of reference summaries for training.Most previous approaches construct multiple reviews and their summary based on textual similarities between reviews, resulting in information mismatch between the review input and the summary.In this paper, we convert each review into a mix of structured and unstructured data, which we call opinion-aspect pairs (OAs) and implicit sentences (ISs).We propose a new method to synthesize training pairs of such mix-structured data as input and the textual summary as output, and design a summarization model with OA encoder and IS encoder.Experiments show that our approach outperforms previous methods on Yelp, Amazon and RottenTomatos datasets. Yizhu Liu, Qi Jia 0003, Kenny Q. Zhu |
EMNLP | 1 |
| 2022 | Reference-free Summarization Evaluation via Semantic Correlation and Compression RatioabstractA document can be summarized in a number of ways.Reference-based evaluation of summarization has been criticized for its inflexibility.In this paper, we propose a new automatic reference-free evaluation metric that compares semantic distribution between source document and summary by pretrained language models and considers summary compression ratio.The experiments show that this metric is more consistent with human evaluation in terms of coherence, consistency, relevance, fluency. Yizhu Liu, Qi Jia 0003, Kenny Q. Zhu |
NAACL-HLT | 1 |
| 2021 | Keyword-aware Abstractive Summarization by Extracting Set-level Intermediate SummariesabstractAbstractive summarization is useful in providing a summary or a digest of news or other web texts and enhancing users reading experience, especially when they are reading on small displays such as mobile phones. However, existing encoder-decoder summarization models have difficulty learning the latent alignment between source documents and summaries because of their vast disparity in length. In this paper, we propose a extractor-abstractor framework in which the keyword-based extractor selects a few sets of salient sentences from the input document and then the abstractor paraphrases these sets of sentences in parallel, which are more aligned to the summary, to generate the final summary. The new extractor and abstractor are pretrained from a set of “pseudo summaries” extracted by specially designed heuristics, and then further trained together in a reinforcement learning framework. The results show that the proposed model generates high-quality summaries with faster training speed and less training memory footprint, and outperforms the state-of-the-art models on CNN/Daily Mail, Webis-TLDR-17, Webis-Snippet-20, WikiHow and DUC-2002 datasets. Yizhu Liu, Qi Jia 0003, Kenny Q. Zhu |
WWW | 1 |
| 2020 | Multi-turn Response Selection using Dialogue Dependency RelationsabstractMulti-turn response selection is a task designed for developing dialogue agents.The performance on this task has a remarkable improvement with pre-trained language models.However, these models simply concatenate the turns in dialogue history as the input and largely ignore the dependencies between the turns.In this paper, we propose a dialogue extraction algorithm to transform a dialogue history into threads based on their dependency relations.Each thread can be regarded as a self-contained sub-dialogue.We also propose Thread-Encoder model to encode threads and candidates into compact representations by pre-trained Transformers and finally get the matching score through an attention layer.The experiments show that dependency relations are helpful for dialogue context understanding, and our model outperforms the state-of-the-art baselines on both DSTC7 and DSTC8*, with competitive results on UbuntuV2. Qi Jia 0003, Yizhu Liu, Kenny Q. Zhu, Haifeng Tang |
EMNLP (1) | 2 |
| 2018 | Controlling Length in Abstractive Summarization Using a Convolutional Neural NetworkabstractConvolutional neural networks (CNNs) have met great success in abstractive summarization, but they cannot effectively generate summaries of desired lengths.Because generated summaries are used in difference scenarios which may have space or length constraints, the ability to control the summary length in abstractive summarization is an important problem.In this paper, we propose an approach to constrain the summary length by extending a convolutional sequence to sequence model.The results show that this approach generates high-quality summaries with user defined length, and outperforms the baselines consistently in terms of ROUGE score, length variations and semantic similarity. Yizhu Liu, Zhiyi Luo, Kenny Q. Zhu |
EMNLP | 1 |