Yixin Liu 0003

dblp:140/7348-3 · DBLP profile ↗
← Back
23ranked-venue papers
10as first author
23since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 10 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Evaluating Mathematical Reasoning Beyond Accuracy
abstract
The leaderboard of Large Language Models (LLMs) in mathematical tasks has been continuously updated. However, the majority of evaluations focus solely on the final results, neglecting the quality of the intermediate steps. This oversight can mask underlying problems, such as logical errors or unnecessary steps in the reasoning process. To measure reasoning beyond final-answer accuracy, we introduce ReasonEval, a new methodology for evaluating the quality of reasoning steps. ReasonEval employs validity and redundancy to characterize the reasoning quality, as well as accompanying LLMs to assess them automatically. We explore different design options for the LLM-based evaluators and empirically demonstrate that ReasonEval, when instantiated with base models possessing strong mathematical knowledge and trained with high-quality labeled data, consistently outperforms baseline methods in the meta-evaluation datasets. We also highlight the strong generalization capabilities of ReasonEval. By utilizing ReasonEval to evaluate LLMs specialized in math, we find that an increase in final-answer accuracy does not necessarily guarantee an improvement in the overall quality of the reasoning steps for challenging mathematical problems. Additionally, we observe that ReasonEval can play a significant role in data selection. We open-source the best-performing model, meta-evaluation script, and all evaluation results to facilitate future research.
Shijie Xia, Xuefeng Li 0003, Yixin Liu 0003, Sherry Tongshuang Wu, Pengfei Liu 0003
AAAI3
2025 AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
abstract
Yilun Zhao, Weiyuan Chen, Zhijian Xu, Manasi Patwardhan, Chengye Wang, Yixin Liu, Lovekesh Vig, Arman Cohan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yilun Zhao 0001, Weiyuan Chen, Manasi Patwardhan 0001, Chengye Wang, Yixin Liu 0003, Lovekesh Vig, Arman Cohan
ACL (1)6
2025 MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
abstract
We introduce $\color{Blue}{\text{MMVU}}$, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. $\color{Blue}{\text{MMVU}}$ includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. Compared to prior benchmarks, $\color{Blue}{\text{MMVU}}$ features three key advancements. First, it challenges models to apply domain-specific knowledge and perform expert-level reasoning to analyze specialized-domain videos, moving beyond the basic visual perception typically assessed in current video benchmarks. Second, each example is annotated by human experts from scratch. We implement strict data quality controls to ensure the high quality of the dataset. Finally, each example is enriched with expert-annotated reasoning rationals and relevant domain knowledge, facilitating in-depth analysis. We conduct an extensive evaluation of 36 frontier multimodal foundation models on $\color{Blue}{\text{MMVU}}$. The latest System-2-capable models, o1 and Gemini 2.0 Flash Thinking, achieve the highest performance among the tested models. However, they still fall short of matching human expertise. Through in-depth error analyses and case studies, we offer actionable insights for future advancements in expert-level, knowledge-intensive video understanding for specialized domains.
Yilun Zhao 0001, Haowei Zhang 0002, Lujing Xie, Tongyan Hu, Guo Gan, Yitao Long, Weiyuan Chen, Chuhan Li, Chengye Wang, Ziyao Shangguan, Zhenwen Liang, Yixin Liu 0003, Chen Zhao 0013, Arman Cohan
CVPR14
2025 CourtReasoner: Can LLM Agents Reason Like Judges?
abstract
Sophia Simeng Han, Yoshiki Takashima, Shannon Zejiang Shen, Chen Liu, Yixin Liu, Roque K. Thuo, Sonia Knowlton, Ruzica Piskac, Scott J Shapiro, Arman Cohan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Simeng Han, Yoshiki Takashima, Shannon Shen 0001, Chen Liu 0020, Yixin Liu 0003, Roque K. Thuo, Sonia Knowlton, Ruzica Piskac, Scott J. Shapiro, Arman Cohan
EMNLP5
2025 ReIFE: Re-evaluating Instruction-Following Evaluation
abstract
Yixin Liu, Kejian Shi, Alexander Fabbri, Yilun Zhao, PeiFeng Wang, Chien-Sheng Wu, Shafiq Joty, Arman Cohan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yixin Liu 0003, Kejian Shi, Alexander R. Fabbri, Yilun Zhao 0001, Peifeng Wang, Chien-Sheng Wu, Shafiq R. Joty, Arman Cohan
NAACL (Long Papers)1
2025 SCIURus: Shared Circuits for Interpretable Uncertainty Representations in Language Models
abstract
Carter Teplica, Yixin Liu, Arman Cohan, Tim G. J. Rudner. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Carter Teplica, Yixin Liu 0003, Arman Cohan, Tim G. J. Rudner
NAACL (Long Papers)2
2025 On Evaluating LLM Alignment by Evaluating LLMs as Judges
abstract
Alignment with human preferences is an important evaluation aspect of LLMs, requiring them to be helpful, honest, safe, and to precisely follow human instructions. Evaluating large language models' (LLMs) alignment typically involves directly assessing their open-ended responses, requiring human annotators or strong LLM judges. Conversely, LLMs themselves have also been extensively evaluated as judges for assessing alignment. In this work, we examine the relationship between LLMs' generation and evaluation capabilities in aligning with human preferences. To this end, we first conduct a comprehensive analysis of the generation-evaluation consistency (GE-consistency) among various LLMs, revealing a strong correlation between their generation and evaluation capabilities when evaluated by a strong LLM preference oracle (GPT-4o). Utilizing this finding, we propose a benchmarking paradigm that measures LLM alignment with human preferences without directly evaluating their generated outputs, instead assessing LLMs in their role as evaluators. Our evaluation shows that our proposed benchmark, AlignEval, matches or surpasses widely used automatic LLM evaluation benchmarks, such as AlpacaEval and Arena-Hard, in capturing human preferences when ranking LLMs. Our study offers valuable insights into the connection between LLMs' generation and evaluation capabilities, and introduces a benchmark that assesses alignment without directly evaluating model outputs.
Yixin Liu 0003, Pengfei Liu 0003, Arman Cohan
NeurIPS1
2025 SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
abstract
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons.By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses.The platform currently supports 44 open-source and proprietary foundation models and has collected over 19,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality.We discuss the results and insights based on the model ranking leaderboard.To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on our collected preference data. The benchmark measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark’s challenges and emphasize the need for more reliable automated evaluation methods.
Yilun Zhao 0001, Tiansheng Hu, Sihong Wu, Ronan Le Bras 0001, Yixin Liu 0003, Robert Tang, Joseph Chee Chang, Jesse Dodge, Jonathan Bragg, Chen Zhao 0013, Hannaneh Hajishirzi, Doug Downey, Arman Cohan
NeurIPS6
2024 DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Financial Documents
abstract
Yilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu, Xiangru Tang, Rui Zhang, Arman Cohan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yilun Zhao 0001, Yitao Long, Hongjun Liu 0001, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu 0003, Xiangru Tang, Rui Zhang 0037, Arman Cohan
ACL (1)7
2024 FOLIO: Natural Language Reasoning with First-Order Logic
abstract
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev
EMNLP17
2024 On Learning to Summarize with Large Language Models as References
abstract
Yixin Liu, Kejian Shi, Katherine He, Longtian Ye, Alexander Fabbri, Pengfei Liu, Dragomir Radev, Arman Cohan. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yixin Liu 0003, Kejian Shi, Katherine He, Longtian Ye, Alexander R. Fabbri, Pengfei Liu 0003, Dragomir R. Radev, Arman Cohan
NAACL-HLT1
2024 Fair Abstractive Summarization of Diverse Perspectives
abstract
Yusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, Rui Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yusen Zhang 0001, Yixin Liu 0003, Alexander R. Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao 0001, Dragomir R. Radev, Kathy McKeown, Rui Zhang 0037
NAACL-HLT3
2023 On Improving Summarization Factual Consistency from Natural Language Feedback
abstract
Yixin Liu, Budhaditya Deb, Milagro Teruel, Aaron Halfaker, Dragomir Radev, Ahmed Hassan Awadallah. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yixin Liu 0003, Budhaditya Deb, Milagro Teruel, Aaron Halfaker, Dragomir R. Radev, Ahmed Awadallah 0001
ACL (1)1
2023 Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation
abstract
Yixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yixin Liu 0003, Alexander R. Fabbri, Pengfei Liu 0003, Yilun Zhao 0001, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev
ACL (1)1
2023 A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization
abstract
Lining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Lining Zhang, Simon Mille, Yufang Hou 0001, Daniel Deutsch, Elizabeth Clark, Yixin Liu 0003, Saad Mahamood, Sebastian Gehrmann, Miruna-Adriana Clinciu, Khyathi Raghavi Chandu, João Sedoc
ACL (1)6
2023 QTSumm: Query-Focused Summarization over Tabular Data
abstract
Yilun Zhao, Zhenting Qi, Linyong Nan, Boyu Mi, Yixin Liu, Weijin Zou, Simeng Han, Ruizhe Chen, Xiangru Tang, Yumo Xu, Dragomir Radev, Arman Cohan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Yilun Zhao 0001, Zhenting Qi, Linyong Nan, Boyu Mi, Yixin Liu 0003, Weijin Zou, Simeng Han, Ruizhe Chen, Xiangru Tang, Yumo Xu, Dragomir R. Radev, Arman Cohan
EMNLP5
2023 Towards Interpretable and Efficient Automatic Reference-Based Summarization Evaluation
abstract
Yixin Liu, Alexander Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Yixin Liu 0003, Alexander R. Fabbri, Yilun Zhao 0001, Pengfei Liu 0003, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev
EMNLP1
2022 BRIO: Bringing Order to Abstractive Summarization
abstract
Abstractive summarization models are commonly trained using maximum likelihood estimation, which assumes a deterministic (onepoint) target distribution in which an ideal model will assign all the probability mass to the reference summary.This assumption may lead to performance degradation during inference, where the model needs to compare several system-generated (candidate) summaries that have deviated from the reference summary.To address this problem, we propose a novel training paradigm which assumes a non-deterministic distribution so that different candidate summaries are assigned probability mass according to their quality.Our method achieves a new state-of-the-art result on the CNN/DailyMail (47.78 ROUGE-1) and XSum (49.07 ROUGE-1) datasets.Further analysis also shows that our model can estimate probabilities of candidate summaries that are more correlated with their level of quality. 1
Yixin Liu 0003, Pengfei Liu 0003, Dragomir R. Radev, Graham Neubig
ACL (1)1
2022 Leveraging Locality in Abstractive Text Summarization
abstract
Neural attention models have achieved significant improvements on many natural language processing tasks.However, the quadratic memory complexity of the self-attention module with respect to the input length hinders their applications in long text summarization.Instead of designing more efficient attention modules, we approach this problem by investigating if models with a restricted context can have competitive performance compared with the memory-efficient attention models that maintain a global context by treating the input as a single sequence.Our model is applied to individual pages, which contain parts of inputs grouped by the principle of locality, during both the encoding and decoding stages.We empirically investigated three kinds of locality in text summarization at different levels of granularity, ranging from sentences to documents.Our experimental results show that our model has a better performance compared with strong baseline models with efficient attention modules, and our analysis provides further insights into our locality-aware modeling strategy.1
Yixin Liu 0003, Ansong Ni, Linyong Nan, Budhaditya Deb, Chenguang Zhu 0001, Ahmed Awadallah 0001, Dragomir R. Radev
EMNLP1
2022 R2D2: Robust Data-to-Text with Replacement Detection
abstract
Unfaithful text generation is a common problem for text generation systems.In the case of Data-to-Text (D2T) systems, the factuality of the generated text is particularly crucial for any real-world applications.We introduce R2D2, a training framework that addresses unfaithful Data-to-Text generation by training a system both as a generator and a faithfulness discriminator with additional replacement detection and unlikelihood learning tasks.To facilitate such training, we propose two methods for sampling unfaithful sentences.We argue that the poor entity retrieval capability of D2T systems is one of the primary sources of unfaithfulness, so in addition to the existing metrics, we further propose named entity based metrics to evaluate the fidelity of D2T generations.Our experimental results show that R2D2 systems could effectively mitigate the unfaithful text generation, and they achieve new state-of-the-art results on FeTaQA, LogicNLG, and ToTTo, all with significant improvements.
Linyong Nan, Lorenzo Jaime Yu Flores, Yilun Zhao 0001, Yixin Liu 0003, Luke Benson, Weijin Zou, Dragomir R. Radev
EMNLP4
2022 Surfer100: Generating Surveys From Web Resources, Wikipedia-style
abstract
Fast-developing fields such as Artificial Intelligence (AI) often outpace the efforts of encyclopedic sources such as Wikipedia, which either do not completely cover recently-introduced topics or lack such content entirely. As a result, methods for automatically producing content are valuable tools to address this information overload. We show that recent advances in pretrained language modeling can be combined for a two-stage extractive and abstractive approach for Wikipedia lead paragraph generation. We extend this approach to generate longer Wikipedia-style summaries with sections and examine how such methods struggle in this application through detailed studies with 100 reference human-collected surveys. This is the first study on utilizing web resources for long Wikipedia-style summaries to the best of our knowledge.
Irene Li, Alexander R. Fabbri, Rina Kawamura, Yixin Liu 0003, Xiangru Tang, Jaesung Tae, Chang Shen, Sally Ma, Tomoe Mizutani, Dragomir R. Radev
LREC4
2021 RefSum: Refactoring Neural Summarization
abstract
Although some recent works show potential complementarity among different state-of-theart systems, few works try to investigate this problem in text summarization.Researchers in other areas commonly refer to the techniques of reranking or stacking to approach this problem.In this work, we highlight several limitations of previous methods, which motivates us to present a new framework Refactor that provides a unified view of text summarization and summaries combination.Experimentally, we perform a comprehensive evaluation that involves twenty-two base systems, four datasets, and three different application scenarios.Besides new state-of-the-art results on CNN/DailyMail dataset (46.18 ROUGE-1), we also elaborate on how our proposed method addresses the limitations of the traditional methods and the effectiveness of the Refactor model sheds light on insight for performance improvement.Our system can be directly used by other researchers as an offthe-shelf tool to achieve further performance improvements.We open-source all the code and provide a convenient interface to use it:https://github.com/yixinL7/ Refactoring-Summarization.
Yixin Liu 0003, Zi-Yi Dou, Pengfei Liu 0003
NAACL-HLT1
2021 On Learning Text Style Transfer with Direct Rewards
abstract
In most cases, the lack of parallel corpora makes it impossible to directly train supervised models for the text style transfer task.In this paper, we explore training algorithms that instead optimize reward functions that explicitly consider different aspects of the styletransferred outputs.In particular, we leverage semantic similarity metrics originally used for fine-tuning neural machine translation models to explicitly assess the preservation of content between system outputs and input texts.We also investigate the potential weaknesses of the existing automatic metrics and propose efficient strategies of using these metrics for training.The experimental results show that our model provides significant gains in both automatic and human evaluation over strong baselines, indicating the effectiveness of our proposed methods and training strategies. 1
Yixin Liu 0003, Graham Neubig, John Wieting
NAACL-HLT1