VLDB 2026 Research / reviewers in the wild / expert
Zichao Li 0001
dblp:95/147-1
· DBLP profile ↗
4ranked-venue papers
2as first author
1since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 52% Efficient and distributed learning · 22% Optimization for machine learning · 14% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
1.1 | 3 | 2020 | Unsupervised Text Generation by Learning from Search · NeurIPS 2020 Decomposable Neural Paraphrase Generation · ACL (1) 2019 Paraphrase Generation with Deep Reinforcement Learning · EMNLP 2018 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.7 | 1 | 2023 | f-Divergence Minimization for Sequence-Level Knowledge Distillation · ACL (1) 2023 |
Machine learning › Optimization for machine learning › black-box optimization › zeroth-order optimization
simulated annealing |
0.4 | 1 | 2020 | Unsupervised Text Generation by Learning from Search · NeurIPS 2020 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 1 | 2020 | Unsupervised Text Generation by Learning from Search · NeurIPS 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.4 | 1 | 2019 | Decomposable Neural Paraphrase Generation · ACL (1) 2019 |
Methods — techniques the papers use, named apart from their topics
f-divergence minimization · 0.7simulated annealing · 0.4search-based learning · 0.4conditional generative model · 0.4transformer · 0.4encoder-decoder · 0.4disentangled representation · 0.4inverse reinforcement learning · 0.3deep reinforcement learning · 0.3deep matching · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | f-Divergence Minimization for Sequence-Level Knowledge DistillationabstractKnowledge distillation (KD) is the process of transferring knowledge from a large model to a small one.It has gained increasing attention in the natural language processing community, driven by the demands of compressing evergrowing language models.In this work, we propose an f -DISTILL framework, which formulates sequence-level knowledge distillation as minimizing a generalized f -divergence function.We propose four distilling variants under our framework and show that existing SeqKD and ENGINE approaches are approximations of our f -DISTILL methods.We further derive step-wise decomposition for our f -DISTILL, reducing intractable sequence-level divergence to word-level losses that can be computed in a tractable manner.Experiments across four datasets show that our methods outperform existing KD approaches, and that our symmetric distilling losses can better force the student to learn from the teacher distribution.1 Yuqiao Wen, Zichao Li 0001, Wenyu Du, Lili Mou |
ACL (1) | 2 |
| 2020 | Unsupervised Text Generation by Learning from SearchabstractIn this work, we propose TGLS, a novel framework for unsupervised Text Generation by Learning from Search. We start by applying a strong search algorithm (in particular, simulated annealing) towards a heuristically defined objective that (roughly) estimates the quality of sentences. Then, a conditional generative model learns from the search results, and meanwhile smooth out the noise of search. The alternation between search and learning can be repeated for performance bootstrapping. We demonstrate the effectiveness of TGLS on two real-world natural language generation tasks, unsupervised paraphrasing and text formalization. Our model significantly outperforms unsupervised baseline methods in both tasks. Especially, it achieves comparable performance to strong supervised methods for paraphrase generation. Jingjing Li 0007, Zichao Li 0001, Lili Mou, Xin Jiang 0002, Michael R. Lyu, Irwin King |
NeurIPS | 2 |
| 2019 | Decomposable Neural Paraphrase GenerationabstractParaphrasing exists at different granularity levels, such as lexical level, phrasal level and sentential level.This paper presents Decomposable Neural Paraphrase Generator (DNPG), a Transformer-based model that can learn and generate paraphrases of a sentence at different levels of granularity in a disentangled way.Specifically, the model is composed of multiple encoders and decoders with different structures, each of which corresponds to a specific granularity.The empirical study shows that the decomposition mechanism of DNPG makes paraphrase generation more interpretable and controllable.Based on DNPG, we further develop an unsupervised domain adaptation method for paraphrase generation.Experimental results show that the proposed model achieves competitive in-domain performance compared to the state-of-the-art neural models, and significantly better performance when adapting to a new domain.What is the population of New York?How many people is there in NYC?Who wrote the Winnie the Pooh books?Who is the author of winnie the pooh?What is the best phone to buy below 15k?Which are best mobile phones to buy under 15000?How can I be a good geologist?What should I do to be a great geologist?How do I reword a sentence to avoid plagiarism?How can I paraphrase my essay and avoid plagiarism? Zichao Li 0001, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001 |
ACL (1) | 1 |
| 2018 | Paraphrase Generation with Deep Reinforcement LearningabstractAutomatic generation of paraphrases from a given sentence is an important yet challenging task in natural language processing (NLP).In this paper, we present a deep reinforcement learning approach to paraphrase generation.Specifically, we propose a new framework for the task, which consists of a generator and an evaluator, both of which are learned from data.The generator, built as a sequenceto-sequence learning model, can produce paraphrases given a sentence.The evaluator, constructed as a deep matching model, can judge whether two sentences are paraphrases of each other.The generator is first trained by deep learning and then further fine-tuned by reinforcement learning in which the reward is given by the evaluator.For the learning of the evaluator, we propose two methods based on supervised learning and inverse reinforcement learning respectively, depending on the type of available training data.Experimental results on two datasets demonstrate the proposed models (the generators) can produce more accurate paraphrases and outperform the stateof-the-art methods in paraphrase generation in both automatic evaluation and human evaluation. Zichao Li 0001, Xin Jiang 0002, Lifeng Shang, Hang Li 0001 |
EMNLP | 1 |