EDBT 2026 Demo / reviewers in the wild / expert
Yixin Liu 0003
dblp:140/7348-3
· DBLP profile ↗
23ranked-venue papers
10as first author
23since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 10 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Mathematical Reasoning Beyond AccuracyabstractThe leaderboard of Large Language Models (LLMs) in mathematical tasks has been continuously updated. However, the majority of evaluations focus solely on the final results, neglecting the quality of the intermediate steps. This oversight can mask underlying problems, such as logical errors or unnecessary steps in the reasoning process. To measure reasoning beyond final-answer accuracy, we introduce ReasonEval, a new methodology for evaluating the quality of reasoning steps. ReasonEval employs validity and redundancy to characterize the reasoning quality, as well as accompanying LLMs to assess them automatically. We explore different design options for the LLM-based evaluators and empirically demonstrate that ReasonEval, when instantiated with base models possessing strong mathematical knowledge and trained with high-quality labeled data, consistently outperforms baseline methods in the meta-evaluation datasets. We also highlight the strong generalization capabilities of ReasonEval. By utilizing ReasonEval to evaluate LLMs specialized in math, we find that an increase in final-answer accuracy does not necessarily guarantee an improvement in the overall quality of the reasoning steps for challenging mathematical problems. Additionally, we observe that ReasonEval can play a significant role in data selection. We open-source the best-performing model, meta-evaluation script, and all evaluation results to facilitate future research. Shijie Xia, Xuefeng Li 0003, Yixin Liu 0003, Sherry Tongshuang Wu, Pengfei Liu 0003 |
AAAI | 3 |
| 2025 | AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific ResearchabstractYilun Zhao, Weiyuan Chen, Zhijian Xu, Manasi Patwardhan, Chengye Wang, Yixin Liu, Lovekesh Vig, Arman Cohan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yilun Zhao 0001, Weiyuan Chen, Manasi Patwardhan 0001, Chengye Wang, Yixin Liu 0003, Lovekesh Vig, Arman Cohan |
ACL (1) | 6 |
| 2025 | MMVU: Measuring Expert-Level Multi-Discipline Video UnderstandingabstractWe introduce $\color{Blue}{\text{MMVU}}$, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. $\color{Blue}{\text{MMVU}}$ includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. Compared to prior benchmarks, $\color{Blue}{\text{MMVU}}$ features three key advancements. First, it challenges models to apply domain-specific knowledge and perform expert-level reasoning to analyze specialized-domain videos, moving beyond the basic visual perception typically assessed in current video benchmarks. Second, each example is annotated by human experts from scratch. We implement strict data quality controls to ensure the high quality of the dataset. Finally, each example is enriched with expert-annotated reasoning rationals and relevant domain knowledge, facilitating in-depth analysis. We conduct an extensive evaluation of 36 frontier multimodal foundation models on $\color{Blue}{\text{MMVU}}$. The latest System-2-capable models, o1 and Gemini 2.0 Flash Thinking, achieve the highest performance among the tested models. However, they still fall short of matching human expertise. Through in-depth error analyses and case studies, we offer actionable insights for future advancements in expert-level, knowledge-intensive video understanding for specialized domains. Yilun Zhao 0001, Haowei Zhang 0002, Lujing Xie, Tongyan Hu, Guo Gan, Yitao Long, Weiyuan Chen, Chuhan Li, Chengye Wang, Ziyao Shangguan, Zhenwen Liang, Yixin Liu 0003, Chen Zhao 0013, Arman Cohan |
CVPR | 14 |
| 2025 | CourtReasoner: Can LLM Agents Reason Like Judges?abstractSophia Simeng Han, Yoshiki Takashima, Shannon Zejiang Shen, Chen Liu, Yixin Liu, Roque K. Thuo, Sonia Knowlton, Ruzica Piskac, Scott J Shapiro, Arman Cohan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Simeng Han, Yoshiki Takashima, Shannon Shen 0001, Chen Liu 0020, Yixin Liu 0003, Roque K. Thuo, Sonia Knowlton, Ruzica Piskac, Scott J. Shapiro, Arman Cohan |
EMNLP | 5 |
| 2025 | ReIFE: Re-evaluating Instruction-Following EvaluationabstractYixin Liu, Kejian Shi, Alexander Fabbri, Yilun Zhao, PeiFeng Wang, Chien-Sheng Wu, Shafiq Joty, Arman Cohan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yixin Liu 0003, Kejian Shi, Alexander R. Fabbri, Yilun Zhao 0001, Peifeng Wang, Chien-Sheng Wu, Shafiq R. Joty, Arman Cohan |
NAACL (Long Papers) | 1 |
| 2025 | SCIURus: Shared Circuits for Interpretable Uncertainty Representations in Language ModelsabstractCarter Teplica, Yixin Liu, Arman Cohan, Tim G. J. Rudner. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Carter Teplica, Yixin Liu 0003, Arman Cohan, Tim G. J. Rudner |
NAACL (Long Papers) | 2 |
| 2025 | On Evaluating LLM Alignment by Evaluating LLMs as JudgesabstractAlignment with human preferences is an important evaluation aspect of LLMs, requiring them to be helpful, honest, safe, and to precisely follow human instructions. Evaluating large language models' (LLMs) alignment typically involves directly assessing their open-ended responses, requiring human annotators or strong LLM judges. Conversely, LLMs themselves have also been extensively evaluated as judges for assessing alignment. In this work, we examine the relationship between LLMs' generation and evaluation capabilities in aligning with human preferences. To this end, we first conduct a comprehensive analysis of the generation-evaluation consistency (GE-consistency) among various LLMs, revealing a strong correlation between their generation and evaluation capabilities when evaluated by a strong LLM preference oracle (GPT-4o). Utilizing this finding, we propose a benchmarking paradigm that measures LLM alignment with human preferences without directly evaluating their generated outputs, instead assessing LLMs in their role as evaluators. Our evaluation shows that our proposed benchmark, AlignEval, matches or surpasses widely used automatic LLM evaluation benchmarks, such as AlpacaEval and Arena-Hard, in capturing human preferences when ranking LLMs. Our study offers valuable insights into the connection between LLMs' generation and evaluation capabilities, and introduces a benchmark that assesses alignment without directly evaluating model outputs. Yixin Liu 0003, Pengfei Liu 0003, Arman Cohan |
NeurIPS | 1 |
| 2025 | SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded TasksabstractWe present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons.By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses.The platform currently supports 44 open-source and proprietary foundation models and has collected over 19,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality.We discuss the results and insights based on the model ranking leaderboard.To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on our collected preference data. The benchmark measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark’s challenges and emphasize the need for more reliable automated evaluation methods. Yilun Zhao 0001, Tiansheng Hu, Sihong Wu, Ronan Le Bras 0001, Yixin Liu 0003, Robert Tang, Joseph Chee Chang, Jesse Dodge, Jonathan Bragg, Chen Zhao 0013, Hannaneh Hajishirzi, Doug Downey, Arman Cohan |
NeurIPS | 6 |
| 2024 | DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Financial DocumentsabstractYilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu, Xiangru Tang, Rui Zhang, Arman Cohan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yilun Zhao 0001, Yitao Long, Hongjun Liu 0001, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu 0003, Xiangru Tang, Rui Zhang 0037, Arman Cohan |
ACL (1) | 7 |
| 2024 | FOLIO: Natural Language Reasoning with First-Order LogicabstractSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev |
EMNLP | 17 |
| 2024 | On Learning to Summarize with Large Language Models as ReferencesabstractYixin Liu, Kejian Shi, Katherine He, Longtian Ye, Alexander Fabbri, Pengfei Liu, Dragomir Radev, Arman Cohan. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yixin Liu 0003, Kejian Shi, Katherine He, Longtian Ye, Alexander R. Fabbri, Pengfei Liu 0003, Dragomir R. Radev, Arman Cohan |
NAACL-HLT | 1 |
| 2024 | Fair Abstractive Summarization of Diverse PerspectivesabstractYusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, Rui Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yusen Zhang 0001, Yixin Liu 0003, Alexander R. Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao 0001, Dragomir R. Radev, Kathy McKeown, Rui Zhang 0037 |
NAACL-HLT | 3 |
| 2023 | On Improving Summarization Factual Consistency from Natural Language FeedbackabstractYixin Liu, Budhaditya Deb, Milagro Teruel, Aaron Halfaker, Dragomir Radev, Ahmed Hassan Awadallah. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yixin Liu 0003, Budhaditya Deb, Milagro Teruel, Aaron Halfaker, Dragomir R. Radev, Ahmed Awadallah 0001 |
ACL (1) | 1 |
| 2023 | Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationabstractYixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yixin Liu 0003, Alexander R. Fabbri, Pengfei Liu 0003, Yilun Zhao 0001, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
ACL (1) | 1 |
| 2023 | A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for SummarizationabstractLining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Lining Zhang, Simon Mille, Yufang Hou 0001, Daniel Deutsch, Elizabeth Clark, Yixin Liu 0003, Saad Mahamood, Sebastian Gehrmann, Miruna-Adriana Clinciu, Khyathi Raghavi Chandu, João Sedoc |
ACL (1) | 6 |
| 2023 | QTSumm: Query-Focused Summarization over Tabular DataabstractYilun Zhao, Zhenting Qi, Linyong Nan, Boyu Mi, Yixin Liu, Weijin Zou, Simeng Han, Ruizhe Chen, Xiangru Tang, Yumo Xu, Dragomir Radev, Arman Cohan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yilun Zhao 0001, Zhenting Qi, Linyong Nan, Boyu Mi, Yixin Liu 0003, Weijin Zou, Simeng Han, Ruizhe Chen, Xiangru Tang, Yumo Xu, Dragomir R. Radev, Arman Cohan |
EMNLP | 5 |
| 2023 | Towards Interpretable and Efficient Automatic Reference-Based Summarization EvaluationabstractYixin Liu, Alexander Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yixin Liu 0003, Alexander R. Fabbri, Yilun Zhao 0001, Pengfei Liu 0003, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
EMNLP | 1 |
| 2022 | BRIO: Bringing Order to Abstractive SummarizationabstractAbstractive summarization models are commonly trained using maximum likelihood estimation, which assumes a deterministic (onepoint) target distribution in which an ideal model will assign all the probability mass to the reference summary.This assumption may lead to performance degradation during inference, where the model needs to compare several system-generated (candidate) summaries that have deviated from the reference summary.To address this problem, we propose a novel training paradigm which assumes a non-deterministic distribution so that different candidate summaries are assigned probability mass according to their quality.Our method achieves a new state-of-the-art result on the CNN/DailyMail (47.78 ROUGE-1) and XSum (49.07 ROUGE-1) datasets.Further analysis also shows that our model can estimate probabilities of candidate summaries that are more correlated with their level of quality. 1 Yixin Liu 0003, Pengfei Liu 0003, Dragomir R. Radev, Graham Neubig |
ACL (1) | 1 |
| 2022 | Leveraging Locality in Abstractive Text SummarizationabstractNeural attention models have achieved significant improvements on many natural language processing tasks.However, the quadratic memory complexity of the self-attention module with respect to the input length hinders their applications in long text summarization.Instead of designing more efficient attention modules, we approach this problem by investigating if models with a restricted context can have competitive performance compared with the memory-efficient attention models that maintain a global context by treating the input as a single sequence.Our model is applied to individual pages, which contain parts of inputs grouped by the principle of locality, during both the encoding and decoding stages.We empirically investigated three kinds of locality in text summarization at different levels of granularity, ranging from sentences to documents.Our experimental results show that our model has a better performance compared with strong baseline models with efficient attention modules, and our analysis provides further insights into our locality-aware modeling strategy.1 Yixin Liu 0003, Ansong Ni, Linyong Nan, Budhaditya Deb, Chenguang Zhu 0001, Ahmed Awadallah 0001, Dragomir R. Radev |
EMNLP | 1 |
| 2022 | R2D2: Robust Data-to-Text with Replacement DetectionabstractUnfaithful text generation is a common problem for text generation systems.In the case of Data-to-Text (D2T) systems, the factuality of the generated text is particularly crucial for any real-world applications.We introduce R2D2, a training framework that addresses unfaithful Data-to-Text generation by training a system both as a generator and a faithfulness discriminator with additional replacement detection and unlikelihood learning tasks.To facilitate such training, we propose two methods for sampling unfaithful sentences.We argue that the poor entity retrieval capability of D2T systems is one of the primary sources of unfaithfulness, so in addition to the existing metrics, we further propose named entity based metrics to evaluate the fidelity of D2T generations.Our experimental results show that R2D2 systems could effectively mitigate the unfaithful text generation, and they achieve new state-of-the-art results on FeTaQA, LogicNLG, and ToTTo, all with significant improvements. Linyong Nan, Lorenzo Jaime Yu Flores, Yilun Zhao 0001, Yixin Liu 0003, Luke Benson, Weijin Zou, Dragomir R. Radev |
EMNLP | 4 |
| 2022 | Surfer100: Generating Surveys From Web Resources, Wikipedia-styleabstractFast-developing fields such as Artificial Intelligence (AI) often outpace the efforts of encyclopedic sources such as Wikipedia, which either do not completely cover recently-introduced topics or lack such content entirely. As a result, methods for automatically producing content are valuable tools to address this information overload. We show that recent advances in pretrained language modeling can be combined for a two-stage extractive and abstractive approach for Wikipedia lead paragraph generation. We extend this approach to generate longer Wikipedia-style summaries with sections and examine how such methods struggle in this application through detailed studies with 100 reference human-collected surveys. This is the first study on utilizing web resources for long Wikipedia-style summaries to the best of our knowledge. Irene Li, Alexander R. Fabbri, Rina Kawamura, Yixin Liu 0003, Xiangru Tang, Jaesung Tae, Chang Shen, Sally Ma, Tomoe Mizutani, Dragomir R. Radev |
LREC | 4 |
| 2021 | RefSum: Refactoring Neural SummarizationabstractAlthough some recent works show potential complementarity among different state-of-theart systems, few works try to investigate this problem in text summarization.Researchers in other areas commonly refer to the techniques of reranking or stacking to approach this problem.In this work, we highlight several limitations of previous methods, which motivates us to present a new framework Refactor that provides a unified view of text summarization and summaries combination.Experimentally, we perform a comprehensive evaluation that involves twenty-two base systems, four datasets, and three different application scenarios.Besides new state-of-the-art results on CNN/DailyMail dataset (46.18 ROUGE-1), we also elaborate on how our proposed method addresses the limitations of the traditional methods and the effectiveness of the Refactor model sheds light on insight for performance improvement.Our system can be directly used by other researchers as an offthe-shelf tool to achieve further performance improvements.We open-source all the code and provide a convenient interface to use it:https://github.com/yixinL7/ Refactoring-Summarization. Yixin Liu 0003, Zi-Yi Dou, Pengfei Liu 0003 |
NAACL-HLT | 1 |
| 2021 | On Learning Text Style Transfer with Direct RewardsabstractIn most cases, the lack of parallel corpora makes it impossible to directly train supervised models for the text style transfer task.In this paper, we explore training algorithms that instead optimize reward functions that explicitly consider different aspects of the styletransferred outputs.In particular, we leverage semantic similarity metrics originally used for fine-tuning neural machine translation models to explicitly assess the preservation of content between system outputs and input texts.We also investigate the potential weaknesses of the existing automatic metrics and propose efficient strategies of using these metrics for training.The experimental results show that our model provides significant gains in both automatic and human evaluation over strong baselines, indicating the effectiveness of our proposed methods and training strategies. 1 Yixin Liu 0003, Graham Neubig, John Wieting |
NAACL-HLT | 1 |