VLDB 2026 Research / reviewers in the wild / expert
Yufei Tian
dblp:222/7636
· DBLP profile ↗
12ranked-venue papers
7as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SkillVerse : Assessing and Enhancing LLMs with Tree EvaluationabstractAs language models evolve to tackle complex and multifaceted tasks, their evaluation must adapt to capture this intricacy.A granular, skillspecific understanding of model capabilities can empower researchers to make informed model development plans.In this paper, we introduce SKILLVERSE, an unsupervised treestructured diagnosis framework for understanding model proficiency in specific abilities.With LLM as a judge, SKILLVERSE first critiques the model responses, and then organizes them into a hierarchical structure termed dendrogram.Given proficiency at arbitrary levels of granularity, SKILLVERSE is flexible to produce insights of behaviors of modern large models.We also demonstrate its efficacy in two downstream tasks: 1) improving model in-context learning by 25% using a tree-search algorithm to select more informative few shots, and 2) accurately predicting new model weaknesses with a 55% success rate, 22% higher than the baseline. Yufei Tian, Jiao Sun, Nanyun Peng 0001 |
ACL (1) | 1 |
| 2025 | REFFLY: Melody-Constrained Lyrics Editing ModelabstractSongyan Zhao, Bingxuan Li, Yufei Tian, Nanyun Peng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Songyan Zhao, Yufei Tian, Nanyun Peng 0001 |
NAACL (Long Papers) | 3 |
| 2024 | Are Large Language Models Capable of Generating Human-Level Narratives?abstractYufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, Nanyun Peng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen 0001, Jonathan May, Nanyun Peng 0001 |
EMNLP | 1 |
| 2024 | MacGyver: Are Large Language Models Creative Problem Solvers?abstractYufei Tian, Abhilasha Ravichander, Lianhui Qin, Ronan Le Bras, Raja Marjieh, Nanyun Peng, Yejin Choi, Thomas Griffiths, Faeze Brahman. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yufei Tian, Abhilasha Ravichander, Lianhui Qin, Ronan Le Bras 0001, Raja Marjieh, Nanyun Peng 0001, Yejin Choi 0001, Thomas L. Griffiths 0001, Faeze Brahman |
NAACL-HLT | 1 |
| 2023 | Unsupervised Melody-to-Lyrics GenerationabstractYufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar Sigurdsson, Chenyang Tao, Wenbo Zhao, Tagyoung Chung, Jing Huang, Nanyun Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar A. Sigurdsson, Chenyang Tao, Wenbo Zhao 0006, Tagyoung Chung, Jing Huang 0020, Nanyun Peng 0001 |
ACL (1) | 1 |
| 2023 | Evaluating Large Language Models on Controlled Generation TasksabstractJiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu, Qian Hu, Rahul Gupta, John Wieting, Nanyun Peng, Xuezhe Ma. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jiao Sun, Yufei Tian, Wangchunshu Zhou, Rahul Gupta 0001, John Wieting, Nanyun Peng 0001, Xuezhe Ma |
EMNLP | 2 |
| 2023 | Harnessing Black-Box Control to Boost Commonsense in LM's GenerationabstractLarge language models (LLMs) such as GPT-3 have demonstrated a strong capability to generate coherent and contextually relevant text.However, amidst their successes, a crucial issue persists: their generated outputs still lack commonsense at times.Yet fine-tuning the entire LLM towards more commonsensical outputs is computationally expensive if not infeasible.In this paper, we present a computation-efficient framework that steers a frozen Pre-Trained Language Model (PTLM) towards more commonsensical generation (i.e., producing a meaningful and plausible output that incorporates a list of concepts).Specifically, we first construct a reference-free evaluator that assigns a sentence with a commonsensical score by grounding the sentence to a dynamic commonsense knowledge base from four different relational aspects.We then use the scorer as the oracle for commonsense knowledge, and extend the controllable generation method called NADO to train an auxiliary head that guides a fixed PTLM to better satisfy the oracle.We test our framework on a series of GPT-2-, FLAN-T5-and Alpaca-based language models (LMs) on two constrained concept-tosentence benchmarks.Human evaluation results demonstrate that our method consistently leads to the most commonsensical outputs. 1 Yufei Tian, Felix Zhang, Nanyun Peng 0001 |
EMNLP | 1 |
| 2022 | Paraphrase Generation as Unsupervised Machine TranslationabstractIn this paper, we propose a new paradigm for paraphrase generation by treating the task as unsupervised machine translation (UMT) based on the assumption that there must be pairs of sentences expressing the same meaning in a large-scale unlabeled monolingual corpus. The proposed paradigm first splits a large unlabeled corpus into multiple clusters, and trains multiple UMT models using pairs of these clusters. Then based on the paraphrase pairs produced by these UMT models, a unified surrogate model can be trained to serve as the final model to generate paraphrases, which can be directly used for test in the unsupervised setup, or be finetuned on labeled datasets in the supervised setup. The proposed method offers merits over machine-translation-based paraphrase generation methods, as it avoids reliance on bilingual sentence pairs. It also allows human intervene with the model so that more diverse paraphrases can be generated using different filtering criteria. Extensive experiments on existing paraphrase dataset for both the supervised and unsupervised setups demonstrate the effectiveness the proposed paradigm. Xiaofei Sun 0001, Yufei Tian, Yuxian Meng, Nanyun Peng 0001, Fei Wu 0001, Jiwei Li 0001, Chun Fan 0001 |
COLING | 2 |
| 2022 | Go Back in Time: Generating Flashbacks in Stories with Event Temporal PromptsabstractStories or narratives are comprised of a sequence of events.To compose interesting stories, professional writers often leverage a creative writing technique called flashback that inserts past events into current storylines as we commonly observe in novels and plays.However, it is challenging for machines to generate flashbacks as it requires solid understanding of event temporal order (e.g.feeling hungry before eat, not vice versa), and the creativity to arrange storylines so that earlier events do not always appear first in narrative order.Two major issues in existing systems exacerbate the challenges: 1) temporal bias in pretraining and story datasets that leads to monotonic event temporal orders; 2) lack of explicit guidance that helps machines decide where to insert flashbacks.We propose to address these issues using structured storylines to encode events and their pair-wise temporal relations ( before , after and vague ) as temporal prompts that guide how stories should unfold temporally.We leverage a Plan-and-Write framework enhanced by reinforcement learning to generate storylines and stories end-toend.Evaluation results show that the proposed method can generate more interesting stories with flashbacks while maintaining textual diversity, fluency and temporal coherence.1 Rujun Han, Hong Chen 0017, Yufei Tian, Nanyun Peng 0001 |
NAACL-HLT | 3 |
| 2022 | AmbiPun: Generating Humorous Puns with Ambiguous ContextabstractWe propose a simple yet effective way to generate pun sentences that does not require any training on existing puns.Our approach is inspired by humor theories that ambiguity comes from the context rather than the pun word itself.Given a pair of definitions of a pun word, 1 our model first produces a list of related concepts through a reverse dictionary to identify unambiguous words to represent the pun and the alternative senses.We then utilize one-shot GPT3 to generate context words and then generate puns incorporating context words from both senses.Human evaluation shows that our method successfully generates puns 52% of the time, outperforming well crafted baselines and the state-of-the-art models by a large margin. * Equal contribution.† Work done when the author is interning at UCLA. 1 We focus on generating homographic puns where two or more meanings of a word form an intended humorous effect. Anirudh Mittal, Yufei Tian, Nanyun Peng 0001 |
NAACL-HLT | 2 |
| 2022 | Zero-shot Sonnet Generation with Discourse-level Planning and Aesthetics FeaturesabstractPoetry generation, and creative language generation in general, usually suffers from the lack of large training data.In this paper, we present a novel framework to generate sonnets that does not require training on poems.We design a hierarchical framework which plans the poem sketch before decoding.Specifically, a content planning module is trained on non-poetic texts to obtain discourse-level coherence; then a rhyme module generates rhyme words and a polishing module introduces imagery and similes for aesthetics purposes.Finally, we design a constrained decoding algorithm to impose the meter-and-rhyme constraint of the generated sonnets.Automatic and human evaluation show that our multi-stage approach without training on poem corpora generates more coherent, poetic, and creative sonnets than several strong baselines.1 Yufei Tian, Nanyun Peng 0001 |
NAACL-HLT | 1 |
| 2019 | Aspect and Opinion Aware Abstractive Review Summarization with Reinforced Hard Typed DecoderabstractIn this paper, we study abstractive review summarization. Observing that review summaries often consist of aspect words, opinion words and context words, we propose a two-stage reinforcement learning approach, which first predicts the output word type from the three types, and then leverages the predicted word type to generate the final word distribution. Experimental results on two Amazon product review datasets demonstrate that our method can consistently outperform several strong baseline approaches based on ROUGE scores. Yufei Tian, Jianfei Yu, Jing Jiang 0001 |
CIKM | 1 |