VLDB 2026 Research / reviewers in the wild / expert
Pengfei Liu 0003
dblp:34/3381-3
· DBLP profile ↗
68ranked-venue papers
12as first author
42since 2021 · last 2026
0000-0001-9030-1875ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 68 · 12 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time ScalingabstractTest-time compute scaling has emerged as a powerful paradigm for enhancing mathematical reasoning in large language models (LLMs) by allocating additional computational resources during inference. However, current methods employ uniform resource distribution across all reasoning sub-problems, creating fundamental bottlenecks where challenging sub-problems receive insufficient attention while routine operations consume disproportionate resources. This uniform allocation creates performance bottlenecks where additional computational resources yield diminishing returns. Inspired by dual-process theory, we propose SCALE (Selective Resource Allocation), a framework that selectively allocates computational resources based on sub-problem difficulty. SCALE operates through four stages: (1) problem decomposition into sequential reasoning sub-problems, (2) difficulty assessment of each sub-problem to distinguish between routine operations and computationally challenging sub-problems, (3) selective processing mode assignment between System 1 for simple sub-problems and System 2 for complex ones, and (4) sequential execution with context propagation. By concentrating resources on challenging sub-problems while processing routine operations efficiently, SCALE achieves substantial performance improvements with superior resource utilization. Extensive experiments demonstrate that SCALE significantly outperforms uniform scaling baselines, achieving accuracy improvements of up to 13.75 percentage points (57.50% to 71.25% on AIME25) while reducing computational costs by 33-53%, representing a major advance in test-time scaling that addresses fundamental limitations of current approaches. Chunpu Xu, Ruifeng Yuan, Jessie Wang 0004, Wenjie Li 0002, Pengfei Liu 0003 |
AAAI | 6 |
| 2026 | AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsabstractKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang, Jie Sun, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Tianze Xu, Weiye Si, Wenjie Li, Dequan Wang, Pengfei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhao Shi, Mohan Jiang, Jie Sun 0030, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Weiye Si, Wenjie Li 0002, Dequan Wang, Pengfei Liu 0003 |
ACL (1) | 14 |
| 2025 | Evaluating Mathematical Reasoning Beyond AccuracyabstractThe leaderboard of Large Language Models (LLMs) in mathematical tasks has been continuously updated. However, the majority of evaluations focus solely on the final results, neglecting the quality of the intermediate steps. This oversight can mask underlying problems, such as logical errors or unnecessary steps in the reasoning process. To measure reasoning beyond final-answer accuracy, we introduce ReasonEval, a new methodology for evaluating the quality of reasoning steps. ReasonEval employs validity and redundancy to characterize the reasoning quality, as well as accompanying LLMs to assess them automatically. We explore different design options for the LLM-based evaluators and empirically demonstrate that ReasonEval, when instantiated with base models possessing strong mathematical knowledge and trained with high-quality labeled data, consistently outperforms baseline methods in the meta-evaluation datasets. We also highlight the strong generalization capabilities of ReasonEval. By utilizing ReasonEval to evaluate LLMs specialized in math, we find that an increase in final-answer accuracy does not necessarily guarantee an improvement in the overall quality of the reasoning steps for challenging mathematical problems. Additionally, we observe that ReasonEval can play a significant role in data selection. We open-source the best-performing model, meta-evaluation script, and all evaluation results to facilitate future research. Shijie Xia, Xuefeng Li 0003, Yixin Liu 0003, Sherry Tongshuang Wu, Pengfei Liu 0003 |
AAAI | 5 |
| 2025 | Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesabstractAs Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial. While existing benchmarks assess basic ToM abilities, they predominantly focus on static snapshots of mental states, overlooking the temporal evolution that characterizes real-world social interactions. We present DynToM, a novel benchmark specifically designed to evaluate LLMs’ ability to understand and track the temporal progression of mental states across interconnected scenarios. Through a systematic four-step framework, we generate 1,100 social contexts encompassing 5,500 scenarios and 78,100 questions, each validated for realism and quality. Our comprehensive evaluation of ten state-of-the-art LLMs reveals that their average performance underperforms humans by 44.7%, with performance degrading significantly when tracking and reasoning about the shift of mental states. This performance gap highlights fundamental limitations in current LLMs’ ability to model the dynamic nature of human mental states. Jiashuo Wang, Qiancheng Xu, Changhe Song, Chunpu Xu, Wenjie Li 0002, Pengfei Liu 0003 |
ACL (1) | 8 |
| 2025 | DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world EnvironmentsabstractLarge Language Models (LLMs) with web search capabilities show significant potential for deep research, yet current methods-brittle prompt engineering or RAG-based reinforcement learning in controlled environments-fail to capture real-world complexities.In this paper, we introduce DeepResearcher, the first comprehensive framework for end-to-end training of LLM-based deep research agents through scaling reinforcement learning (RL) in real-world environments with authentic web search interactions.Unlike RAG approaches reliant on fixed corpora, DeepResearcher trains agents to navigate the noisy, dynamic open web.We implement a specialized multi-agent architecture where browsing agents extract relevant information from various webpage structures and overcoming significant technical challenges.Extensive experiments on open-domain research tasks demonstrate that DeepResearcher achieves substantial improvements of up to 28.9 points over prompt engineering-based baselines and up to 7.2 points over RAG-based RL agents.Our qualitative analysis reveals emergent cognitive behaviors from end-to-end RL training, such as planning, cross-validation, self-reflection for research redirection, and maintain honesty when unable to find definitive answers.Our results highlight that end-to-end training in realworld web environments is fundamental for developing robust research capabilities aligned with real-world applications.The source code for DeepResearcher is released at: https:// github.com/GAIR-NLP/DeepResearcher. Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, Pengfei Liu 0003 |
EMNLP | 7 |
| 2025 | Progress or Regress? Self-Improvement Reversal in Post-trainingabstractSelf-improvement through post-training methods such as iterative preference learning has been acclaimed for enhancing the problem-solving capabilities (e.g., mathematical reasoning) of Large Language Models (LLMs) without human intervention. However, as our exploration deepens, it is crucial to critically assess whether these enhancements indeed signify comprehensive progress or if they could lead to unintended regressions. Through rigorous experimentation and analysis across diverse problem-solving tasks, we uncover nuances in the self-improvement trajectories of LLMs. Our study introduces the concept of \emph{self-improvement reversal}, where models showing improved overall accuracy metrics might paradoxically exhibit declines in broader, essential capabilities. We propose a comprehensive evaluative framework to scrutinize the underlying mechanisms and outcomes of post-training self-improvement, aiming to discern between superficial metric improvements and genuine enhancements in model functionality. The findings emphasize the complexity of technological advancements in LLMs, underscoring the need for a nuanced understanding of the \textit{progress or regress} dichotomy in their development. Xuefeng Li 0003, Pengfei Liu 0003 |
ICLR | 3 |
| 2025 | OmniBal: Towards Fast Instruction-Tuning for Vision-Language Models via Omniverse Computation BalanceabstractVision-language instruction-tuning models have recently achieved significant performance improvements. In this work, we discover that large-scale 3D parallel training on those models leads to an imbalanced computation load across different devices. The vision and language parts are inherently heterogeneous: their data distribution and model architecture differ significantly, which affects distributed training efficiency. To address this issue, we rebalance the computational load from data, model, and memory perspectives, achieving more balanced computation across devices. Specifically, for the data, instances are grouped into new balanced mini-batches within and across devices. A search-based method is employed for the model to achieve a more balanced partitioning. For memory optimization, we adaptively adjust the re-computation strategy for each partition to utilize the available memory fully. These three perspectives are not independent but are closely connected, forming an omniverse balanced training framework. Extensive experiments are conducted to validate the effectiveness of our method. Compared with the open-source training code of InternVL-Chat, training time is reduced greatly, achieving about 1.8$\times$ speed-up. Our method’s efficacy and generalizability are further validated across various models and datasets. Codes will be released at https://github.com/ModelTC/OmniBal. Yongqiang Yao, Jingru Tan, Feizhao Zhang, Yazhe Niu, Xin Jin 0008, Bo Li 0126, Pengfei Liu 0003, Ruihao Gong, Dahua Lin, Ningyi Xu |
ICML | 8 |
| 2025 | Programming Every Example: Lifting Pre-training Data Quality Like Experts at ScaleabstractLarge language model pre-training has traditionally relied on human experts to craft heuristics for improving the corpora quality, resulting in numerous rules developed to date. However, these fixed rules lack the flexibility to address the unique characteristics of individual examples, yet crafting sample-wise rules is impractical for human experts. In this paper, we show that even small language models, with only 0.3B parameters, can exhibit substantial data refining capabilities. We propose Programming Every Example (ProX), a novel framework that treats data refinement as a programming task, and enables the model to refine corpora by generating and executing fine-grained operations, such as string normalization, for each individual example at scale. Experiments show that models trained on ProX-refined data consistently outperform other baselines across 10 benchmarks, demonstrating effectiveness across model sizes (up to 1.7B) and pre-training corpora (C4, RedPajama-V2, FineWeb, FineWeb-Edu, and DCLM). ProX also shows great potential in continual pre-training: on math domain, ProX boosts 7B models by up to 20% within 10B tokens—results typically achieved with much larger scale training (e.g., 200B tokens). We believe ProX offers a way to curate high-quality pre-training data, and finally contributes to efficient LLM development. Zengzhi Wang, Qian Liu 0033, Pengfei Liu 0003 |
ICML | 5 |
| 2025 | On Evaluating LLM Alignment by Evaluating LLMs as JudgesabstractAlignment with human preferences is an important evaluation aspect of LLMs, requiring them to be helpful, honest, safe, and to precisely follow human instructions. Evaluating large language models' (LLMs) alignment typically involves directly assessing their open-ended responses, requiring human annotators or strong LLM judges. Conversely, LLMs themselves have also been extensively evaluated as judges for assessing alignment. In this work, we examine the relationship between LLMs' generation and evaluation capabilities in aligning with human preferences. To this end, we first conduct a comprehensive analysis of the generation-evaluation consistency (GE-consistency) among various LLMs, revealing a strong correlation between their generation and evaluation capabilities when evaluated by a strong LLM preference oracle (GPT-4o). Utilizing this finding, we propose a benchmarking paradigm that measures LLM alignment with human preferences without directly evaluating their generated outputs, instead assessing LLMs in their role as evaluators. Our evaluation shows that our proposed benchmark, AlignEval, matches or surpasses widely used automatic LLM evaluation benchmarks, such as AlpacaEval and Arena-Hard, in capturing human preferences when ranking LLMs. Our study offers valuable insights into the connection between LLMs' generation and evaluation capabilities, and introduces a benchmark that assesses alignment without directly evaluating model outputs. Yixin Liu 0003, Pengfei Liu 0003, Arman Cohan |
NeurIPS | 2 |
| 2025 | LIMOPro: Reasoning Refinement for Efficient and Effective Test-time ScalingabstractLarge language models (LLMs) have demonstrated remarkable reasoning capabilities through test-time scaling approaches, particularly when fine-tuned with chain-of-thought (CoT) data distilled from more powerful large reasoning models (LRMs). However, these reasoning chains often contain verbose elements that mirror human problem-solving, categorized as progressive reasoning (the essential solution development path) and functional elements (verification processes, alternative solution approaches, and error corrections). While progressive reasoning is crucial, the functional elements significantly increase computational demands during test-time inference. We introduce PIR (Perplexity-based Importance Refinement), a principled framework that quantitatively evaluates the importance of each reasoning step based on its impact on answer prediction confidence. PIR systematically identifies and selectively prunes only low-importance functional steps while preserving all progressive reasoning components, creating optimized training data that maintains the integrity of the core solution path while reducing verbosity. Models fine-tuned on PIR-optimized data exhibit superior test-time scaling properties, generating more concise reasoning chains while achieving improved accuracy (+0.9\% to +6.6\%) with significantly reduced token usage (-3\% to -41\%) across challenging reasoning benchmarks (AIME, AMC, and GPQA Diamond). Our approach demonstrates strong generalizability across different model sizes, data sources, and token budgets, offering a practical solution for deploying reasoning-capable LLMs in scenarios where efficient test-time scaling, response time, and computational efficiency are valuable constraints. Code and dataset are available at the [LIMOPro GitHub repository.](https://github.com/GAIR-NLP/LIMOPro) Jiashuo Wang, Ruifeng Yuan, Chunpu Xu, Kaishuai Xu, Wenjie Li 0002, Pengfei Liu 0003 |
NeurIPS | 7 |
| 2024 | Dissecting Human and LLM PreferencesabstractAs a relative quality comparison of model responses, human and Large Language Model (LLM) preferences serve as common alignment goals in model fine-tuning and criteria in evaluation.Yet, these preferences merely reflect broad tendencies, resulting in less explainable and controllable models with potential safety risks.In this work, we dissect the preferences of human and 32 different LLMs to understand their quantitative composition, using annotations from real-world user-model conversations for a fine-grained, scenario-wise analysis.We find that humans are less sensitive to errors, favor responses that support their stances, and show clear dislike when models admit their limits.On the contrary, advanced LLMs like GPT-4-Turbo emphasize correctness, clarity, and harmlessness more.Additionally, LLMs of similar sizes tend to exhibit similar preferences, regardless of their training methods, and fine-tuning for alignment does not significantly alter the preferences of pretrained-only LLMs.Finally, we show that preference-based evaluation can be intentionally manipulated.In both training-free and training-based settings, aligning a model with the preferences of judges boosts scores, while injecting the least preferred properties lowers them.This results in notable score shifts: up to 0.59 on MT-Bench (1-10 scale) and 31.94 on AlpacaEval 2.0 (0-100 scale), highlighting the significant impact of this strategic adaptation.We have made all resources of this project publicly available. Shichao Sun, Yikai Zhang 0003, Hai Zhao 0001, Pengfei Liu 0003 |
ACL (1) | 6 |
| 2024 | MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story GenerationabstractA story premise succinctly defines a story's main idea, foundation, and trajectory.It serves as the initial trigger in automatic story generation.Existing sources of story premises are limited by a lack of diversity, uneven quality, and high costs that make them difficult to scale.In response, we introduce Modular Story Premise Synthesis (MoPS) which breaks down story premises into modules like background and persona for automated design and generation.MoPS consists of three phases: (1) Precollect a consistent set of candidates for each module to form a nested dictionary.(2) Extract a key path from the nested dictionary as the premise design.(3) Instruct an LLM to integrate the design into a coherent premise sentence.Thorough evaluations demonstrate that our synthesized premises excel in diversity, fascination, completeness, and originality compared to those induced from large language models and captured from public story datasets.Similarly, the extended novels and scripts generated from our premises also exhibit higher quality.In supplementary materials, we provide the MoPS code suite, along with 7.6k generated premises and 1k extended stories. Yu Qiao 0001, Pengfei Liu 0003 |
ACL (1) | 3 |
| 2024 | ECON: On the Detection and Resolution of Evidence ConflictsabstractCheng Jiayang, Chunkit Chan, Qianqian Zhuang, Lin Qiu, Tianhang Zhang, Tengxiao Liu, Yangqiu Song, Yue Zhang, Pengfei Liu, Zheng Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Cheng Jiayang, Chunkit Chan, Qianqian Zhuang, Tianhang Zhang, Tengxiao Liu, Yangqiu Song, Yue Zhang 0004, Pengfei Liu 0003, Zheng Zhang 0001 |
EMNLP | 9 |
| 2024 | FRoG: Evaluating Fuzzy Reasoning of Generalized Quantifiers in LLMsabstractFuzzy reasoning is vital due to the frequent use of imprecise information in daily contexts.However, the ability of current large language models (LLMs) to handle such reasoning remains largely uncharted.In this paper, we introduce a new benchmark, FROG, for fuzzy reasoning, featuring real-world mathematical word problems that incorporate generalized quantifiers.Our experimental findings reveal that fuzzy reasoning continues to pose significant challenges for LLMs.Moreover, we find that existing methods designed to enhance reasoning do not consistently improve performance in tasks involving fuzzy logic.Additionally, our results show an inverse scaling effect in the performance of LLMs on FROG.Interestingly, we also demonstrate that strong mathematical reasoning skills are not necessarily indicative of success on our benchmark 1 . Yiyuan Li, Shichao Sun, Pengfei Liu 0003 |
EMNLP | 3 |
| 2024 | Generative Judge for Evaluating AlignmentabstractThe rapid development of Large Language Models (LLMs) has substantially expanded the range of tasks they can address. In the field of Natural Language Processing (NLP), researchers have shifted their focus from conventional NLP tasks (e.g., sequence tagging and parsing) towards tasks that revolve around aligning with human needs (e.g., brainstorming and email writing). This shift in task distribution imposes new requirements on evaluating these aligned models regarding *generality* (i.e., assessing performance across diverse scenarios), *flexibility* (i.e., examining under different protocols), and *interpretability* (i.e., scrutinizing models with explanations). In this paper, we propose a generative judge with 13B parameters, **Auto-J**, designed to address these challenges. Our model is trained on user queries and LLM-generated responses under massive real-world scenarios and accommodates diverse evaluation protocols (e.g., pairwise response comparison and single-response evaluation) with well-structured natural language critiques. To demonstrate the efficacy of our approach, we construct a new testbed covering 58 different scenarios. Experimentally, **Auto-J** outperforms a series of strong competitors, including both open-source and closed-source models, by a large margin. We also provide detailed analysis and case studies to further reveal the potential of our method and make a variety of resources public at https://github.com/GAIR-NLP/auto-j. Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao 0001, Pengfei Liu 0003 |
ICLR | 6 |
| 2024 | GPTScore: Evaluate as You DesireabstractJinlan Fu, See-Kiong Ng, Zhengbao Jiang, Pengfei Liu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, Pengfei Liu 0003 |
NAACL-HLT | 4 |
| 2024 | On Learning to Summarize with Large Language Models as ReferencesabstractYixin Liu, Kejian Shi, Katherine He, Longtian Ye, Alexander Fabbri, Pengfei Liu, Dragomir Radev, Arman Cohan. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yixin Liu 0003, Kejian Shi, Katherine He, Longtian Ye, Alexander R. Fabbri, Pengfei Liu 0003, Dragomir R. Radev, Arman Cohan |
NAACL-HLT | 6 |
| 2024 | Alignment for HonestyabstractRecent research has made significant strides in aligning large language models (LLMs) with helpfulness and harmlessness. In this paper, we argue for the importance of alignment for \emph{honesty}, ensuring that LLMs proactively refuse to answer questions when they lack knowledge, while still not being overly conservative. However, a pivotal aspect of alignment for honesty involves discerning an LLM's knowledge boundaries, which demands comprehensive solutions in terms of metric development, benchmark creation, and training methodologies. We address these challenges by first establishing a precise problem definition and defining ``honesty'' inspired by the Analects of Confucius. This serves as a cornerstone for developing metrics that effectively measure an LLM's honesty by quantifying its progress post-alignment. Furthermore, we introduce a flexible training framework which is further instantiated by several efficient fine-tuning techniques that emphasize honesty without sacrificing performance on other tasks. Our extensive experiments reveal that these aligned models show a marked increase in honesty, as indicated by our proposed metrics. We open-source all relevant resources to facilitate future research at \url{https://github.com/GAIR-NLP/alignment-for-honesty}. Yuqing Yang 0004, Ethan Chern, Xipeng Qiu, Graham Neubig, Pengfei Liu 0003 |
NeurIPS | 5 |
| 2024 | OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AIabstractThe evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abilities in problem-solving and scientific discovery (i.e., AI4Science) once exclusive to human intellect. To comprehensively evaluate current models' performance in cognitive reasoning abilities, we introduce OlympicArena, which includes 11,163 bilingual problems across both text-only and interleaved text-image modalities. These challenges encompass a wide range of disciplines spanning seven fields and 62 international Olympic competitions, rigorously examined for data leakage. We argue that the challenges in Olympic competition problems are ideal for evaluating AI's cognitive reasoning due to their complexity and interdisciplinary nature, which are essential for tackling complex scientific challenges and facilitating discoveries. Beyond evaluating performance across various disciplines using answer-only criteria, we conduct detailed experiments and analyses from multiple perspectives. We delve into the models' cognitive reasoning abilities, their performance across different modalities, and their outcomes in process-level evaluations, which are vital for tasks requiring complex reasoning with lengthy solutions. Our extensive evaluations reveal that even advanced models like GPT-4o only achieve a 39.97\% overall accuracy (28.67\% for mathematics and 29.71\% for physics), illustrating current AI limitations in complex reasoning and multimodal integration. Through the OlympicArena, we aim to advance AI towards superintelligence, equipping it to address more complex challenges in science and beyond. We also provide a comprehensive set of resources to support AI research, including a benchmark dataset, an open-source annotation platform, a detailed evaluation tool, and a leaderboard with automatic submission features. Zengzhi Wang, Shijie Xia, Xuefeng Li 0003, Haoyang Zou, Ruijie Xu 0005, Run-Ze Fan, Lyumanshan Ye, Ethan Chern, Yixin Ye, Yikai Zhang 0003, Yuqing Yang 0004, Binjie Wang, Shichao Sun, Yiyuan Li, Steffi Chern, Yiwei Qin, Jiadi Su, Yixiu Liu, Shaoting Zhang 0001, Dahua Lin, Yu Qiao 0001, Pengfei Liu 0003 |
NeurIPS | 28 |
| 2024 | RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented GenerationabstractDespite Retrieval-Augmented Generation (RAG) has shown promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose a fine-grained evaluation framework, RAGChecker, that incorporates a suite of diagnostic metrics for both the retrieval and generation modules. Meta evaluation verifies that RAGChecker has significantly better correlations with human judgments than other evaluation metrics. Using RAGChecker, we evaluate 8 RAG systems and conduct an in-depth analysis of their performance, revealing insightful patterns and trade-offs in the design choices of RAG architectures. The metrics of RAGChecker can guide researchers and practitioners in developing more effective RAG systems. Dongyu Ru, Xiangkun Hu, Tianhang Zhang, Peng Shi 0010, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li 0010, Binjie Wang, Jiarong Jiang, Tong He 0002, Zhiguo Wang 0006, Pengfei Liu 0003, Yue Zhang 0004, Zheng Zhang 0001 |
NeurIPS | 16 |
| 2024 | MathPile: A Billion-Token-Scale Pretraining Corpus for MathabstractHigh-quality, large-scale corpora are the cornerstone of building foundation models. In this work, we introduce MathPile, a diverse and high-quality math-centric corpus comprising about 9.5 billion tokens. Throughout its creation, we adhered to the principle of “less is more”, firmly believing in the supremacy of data quality over quantity, even in the pre-training phase. Our meticulous data collection and processing efforts included a complex suite of preprocessing, prefiltering, language identification, cleaning, filtering, and deduplication, ensuring the high quality of our corpus. Furthermore, we performed data contamination detection on downstream benchmark test sets to eliminate duplicates and conducted continual pre-training experiments, booting the performance on common mathematical reasoning benchmarks. We aim for our MathPile to boost language models’ mathematical reasoning abilities and open-source its different versions and processing scripts to advance the field. Zengzhi Wang, Xuefeng Li 0003, Pengfei Liu 0003 |
NeurIPS | 4 |
| 2023 | DataFinder: Scientific Dataset Recommendation from Natural Language DescriptionsabstractModern machine learning relies on datasets to develop and validate research ideas.Given the growth of publicly available data, finding the right dataset to use is increasingly difficult.Any research question imposes explicit and implicit constraints on how well a given dataset will enable researchers to answer this question, such as dataset size, modality, and domain.We operationalize the task of recommending datasets given a short natural language description of a research idea, to help people find relevant datasets for their needs.Dataset recommendation poses unique challenges as an information retrieval problem; datasets are hard to directly index for search and there are no corpora readily available for this task.To facilitate this task, we build the DataFinder Dataset which consists of a larger automatically-constructed training set (17.5K queries) and a smaller expertannotated evaluation set (392 queries).Using this data, we compare various information retrieval algorithms on our test set and present a superior bi-encoder retriever for text-based dataset recommendation.This system, trained on the DataFinder Dataset, finds more relevant search results than existing third-party dataset search engines.To encourage progress on dataset recommendation, we release our dataset and models to the public.1 Vijay Viswanathan 0002, Luyu Gao, Sherry Tongshuang Wu, Pengfei Liu 0003, Graham Neubig |
ACL (1) | 4 |
| 2023 | Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationabstractYixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yixin Liu 0003, Alexander R. Fabbri, Pengfei Liu 0003, Yilun Zhao 0001, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
ACL (1) | 3 |
| 2023 | Towards Interpretable and Efficient Automatic Reference-Based Summarization EvaluationabstractYixin Liu, Alexander Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yixin Liu 0003, Alexander R. Fabbri, Yilun Zhao 0001, Pengfei Liu 0003, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
EMNLP | 4 |
| 2023 | GlobalBench: A Benchmark for Global Progress in Natural Language ProcessingabstractYueqi Song, Simran Khanuja, Pengfei Liu, Fahim Faisal, Alissa Ostapenko, Genta Winata, Alham Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yueqi Song, Simran Khanuja, Pengfei Liu 0003, Fahim Faisal, Alissa Ostapenko, Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig |
EMNLP | 3 |
| 2023 | PAL: Program-aided Language ModelsabstractLarge language models (LLMs) have demonstrated an impressive ability to perform arithmetic and symbolic reasoning tasks, when provided with a few examples at test time ("few-shot prompting"). Much of this success can be attributed to prompting methods such as "chain-of-thought", which employ LLMs for both understanding the problem description by decomposing it into steps, as well as solving each step of the problem. While LLMs seem to be adept at this sort of step-by-step decomposition, LLMs often make logical and arithmetic mistakes in the solution part, even when the problem is decomposed correctly. In this paper, we present Program-Aided Language models (PAL): a novel approach that uses the LLM to read natural language problems and generate programs as the intermediate reasoning steps, but offloads the solution step to a runtime such as a Python interpreter. With PAL, decomposing the natural language problem into runnable steps remains the only learning task for the LLM, while solving is delegated to the interpreter. We demonstrate this synergy between a neural LLM and a symbolic interpreter across 13 mathematical, symbolic, and algorithmic reasoning tasks from BIG-Bench Hard and others. In all these natural language reasoning tasks, generating code using an LLM and reasoning using a Python interpreter leads to more accurate results than much larger models. For example, PAL using Codex achieves state-of-the-art few-shot accuracy on GSM8K, surpassing PaLM which uses chain-of-thought by absolute 15% top-1. Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 0002, Pengfei Liu 0003, Yiming Yang 0002, Jamie Callan, Graham Neubig |
ICML | 5 |
| 2023 | FELM: Benchmarking Factuality Evaluation of Large Language ModelsabstractAssessing factuality of text generated by large language models (LLMs) is an emerging yet crucial research area, aimed at alerting users to potential errors and guiding the development of more reliable LLMs. Nonetheless, the evaluators assessing factuality necessitate suitable evaluation themselves to gauge progress and foster advancements. This direction remains under-explored, resulting in substantial impediments to the progress of factuality evaluators. To mitigate this issue, we introduce a benchmark for Factuality Evaluation of large Language Models, referred to as FELM. In this benchmark, we collect responses generated from LLMs and annotate factuality labels in a fine-grained manner. Contrary to previous studies that primarily concentrate on the factuality of world knowledge (e.g. information from Wikipedia), FELM focuses on factuality across diverse domains, spanning from world knowledge to math and reasoning. Our annotation is based on text segments, which can help pinpoint specific factual errors. The factuality annotations are further supplemented by predefined error types and reference links that either support or contradict the statement. In our experiments, we investigate the performance of several LLM-based factuality evaluators on FELM, including both vanilla LLMs and those augmented with retrieval mechanisms and chain-of-thought processes. Our findings reveal that while retrieval aids factuality evaluation, current LLMs are far from satisfactory to faithfully detect factual errors. Shiqi Chen 0002, Yiran Zhao 0006, Jinghan Zhang 0006, I-Chun Chern, Siyang Gao, Pengfei Liu 0003, Junxian He |
NeurIPS | 6 |
| 2022 | KID-Review: Knowledge-Guided Scientific Review Generation with Oracle Pre-trainingabstractThe surge in the number of scientific submissions has brought challenges to the work of peer review. In this paper, as a first step, we explore the possibility of designing an automated system, which is not meant to replace humans, but rather providing a first-pass draft for a machine-assisted human review process. Specifically, we present an end-to-end knowledge-guided review generation framework for scientific papers grounded in cognitive psychology research that a better understanding of text requires different types of knowledge. In practice, we found that this seemingly intuitive idea suffered from training difficulties. In order to solve this problem, we put forward an oracle pre-training strategy, which can not only make the Kid-Review better educated but also make the generated review cover more aspects. Experimentally, we perform a comprehensive evaluation (human and automatic) from different perspectives. Empirical results have shown the effectiveness of different types of knowledge as well as oracle pre-training. We make all code, relevant dataset available: https://github.com/Anonymous4nlp233/KIDReview as well as the Kid-Review system: http://nlpeer.reviews. Weizhe Yuan, Pengfei Liu 0003 |
AAAI | 2 |
| 2022 | BRIO: Bringing Order to Abstractive SummarizationabstractAbstractive summarization models are commonly trained using maximum likelihood estimation, which assumes a deterministic (onepoint) target distribution in which an ideal model will assign all the probability mass to the reference summary.This assumption may lead to performance degradation during inference, where the model needs to compare several system-generated (candidate) summaries that have deviated from the reference summary.To address this problem, we propose a novel training paradigm which assumes a non-deterministic distribution so that different candidate summaries are assigned probability mass according to their quality.Our method achieves a new state-of-the-art result on the CNN/DailyMail (47.78 ROUGE-1) and XSum (49.07 ROUGE-1) datasets.Further analysis also shows that our model can estimate probabilities of candidate summaries that are more correlated with their level of quality. 1 Yixin Liu 0003, Pengfei Liu 0003, Dragomir R. Radev, Graham Neubig |
ACL (1) | 2 |
| 2022 | Polyglot Prompt: Multilingual Multitask Prompt TrainingabstractThis paper aims for a potential architectural improvement for multilingual learning and asks: Can different tasks from different languages be modeled in a monolithic framework, i.e. without any task/language-specific module?The benefit of achieving this could open new doors for future multilingual research, including allowing systems trained on low resources to be further assisted by other languages as well as other tasks.We approach this goal by developing a learning framework named Polyglot Prompting to exploit prompting methods for learning a unified semantic space for different languages and tasks with multilingual prompt engineering.We performed a comprehensive evaluation of 6 tasks, namely topic classification, sentiment classification, named entity recognition, question answering, natural language inference, and summarization, covering 24 datasets and 49 languages.The experimental results demonstrated the efficacy of multilingual multitask prompt-based learning and led to inspiring observations.We also present an interpretable multilingual evaluation methodology and show how the proposed framework, multilingual multitask prompt training, works.We release all datasets prompted in the best setting and code. 1 Jinlan Fu, See-Kiong Ng, Pengfei Liu 0003 |
EMNLP | 3 |
| 2022 | Towards a Unified Multi-Dimensional Evaluator for Text GenerationabstractMulti-dimensional evaluation is the dominant paradigm for human evaluation in Natural Language Generation (NLG), i.e., evaluating the generated text from multiple explainable dimensions, such as coherence and fluency.However, automatic evaluation in NLG is still dominated by similarity-based metrics, and we lack a reliable framework for a more comprehensive evaluation of advanced models.In this paper, we propose a unified multi-dimensional evaluator UNIEVAL for NLG.We re-frame NLG evaluation as a Boolean Question Answering (QA) task, and by guiding the model with different questions, we can use one evaluator to evaluate from multiple dimensions.Furthermore, thanks to the unified Boolean QA format, we are able to introduce an intermediate learning phase that enables UNIEVAL to incorporate external knowledge from multiple related tasks and gain further improvement.Experiments on three typical NLG tasks show that UNIEVAL correlates substantially better with human judgments than existing metrics.Specifically, compared to the top-performing unified evaluators, UNIEVAL achieves a 23% higher correlation on text summarization, and over 43% on dialogue response generation.Also, UNIEVAL demonstrates a strong zero-shot learning ability for unseen evaluation dimensions and tasks.Source code, data and all pre-trained evaluators are available on our GitHub repository 1 . Generated Summary:Harry Kane is nominated for both the PFA player and young player of the season.The Spurs striker has been released from the awards ceremony on Sunday.The Tottenham striker features in a new animation.Reference Summary: Harry Kane has been in superb form for Tottenham this season.The 21-year-old has scored 30 goals in all competitions for Spurs.Kane also made his England debut and scored within two minutes.Document: Harry Kane's celebrations this season have always shown him to be an animated young man . . .Similarity-based Evaluators ROUGE-1: 0.44 ROUGE-2: 0.25 ROUGE-L: 0.42 BERTScore: 0.24 Single-dimensional Evaluators (predicted by two different evaluators (Deng et al., 2021)) Consistency: 0.87 Relevance: 0.74 Unified Evaluator (predicted by BARTScore, and the scoring range is negative infinity to 0) Precision: -5.45 Recall: -4.93 F1: -5.19 Ming Zhong 0005, Yang Liu 0005, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu 0003, Chenguang Zhu 0001, Heng Ji 0001, Jiawei Han 0001 |
EMNLP | 6 |
| 2022 | Are All the Datasets in Benchmark Necessary? A Pilot Study of Dataset Evaluation for Text ClassificationabstractIn this paper, we ask the research question of whether all the datasets in the benchmark are necessary.We approach this by first characterizing the distinguishability of datasets when comparing different systems.Experiments on 9 datasets and 36 systems show that several existing benchmark datasets contribute little to discriminating top-scoring systems, while those less used datasets exhibit impressive discriminative power.We further, taking the text classification task as a case study, investigate the possibility of predicting dataset discrimination based on its properties (e.g., average sentence length).Our preliminary experiments promisingly show that given a sufficient number of training experimental records, a meaningful predictor can be learned to estimate dataset discrimination over unseen datasets.We released all datasets with features explored in this work on DataLab. Jinlan Fu, See-Kiong Ng, Pengfei Liu 0003 |
NAACL-HLT | 4 |
| 2022 | Can We Automate Scientific Reviewing?abstractThe rapid development of science and technology has been accompanied by an exponential growth in peer-reviewed scientific publications. At the same time, the review of each paper is a laborious process that must be carried out by subject matter experts. Thus, providing high-quality reviews of this growing number of papers is a significant challenge. In this work, we ask the question “can we automate scientific reviewing? ”, discussing the possibility of using natural language processing (NLP) models to generate peer reviews for scientific papers. Because it is non-trivial to define what a “good” review is in the first place, we first discuss possible evaluation metrics that could be used to judge success in this task. We then focus on the machine learning domain and collect a dataset of papers in the domain, annotate them with different aspects of content covered in each review, and train targeted summarization models that take in papers as input and generate reviews as output. Comprehensive experimental results on the test set show that while system-generated reviews are comprehensive, touching upon more aspects of the paper than human-written reviews, the generated texts are less constructive and less factual than human-written reviews for all aspects except the explanation of the core ideas of the papers, which are largely factually correct. Given these results, we pose eight challenges in the pursuit of a good review generation system together with potential solutions, which, hopefully, will inspire more future research in this direction. We make relevant resource publicly available for use by future research: https://github. com/neulab/ReviewAdvisor. In addition, while our conclusion is that the technology is not yet ready for use in high-stakes review settings we provide a system demo, ReviewAdvisor (http://review.nlpedia.ai/), showing the current capabilities and failings of state-of-the-art NLP models at this task (see demo screenshot in A.2). A review of this paper written by the system proposed in this paper can be found in A.1. Weizhe Yuan, Pengfei Liu 0003, Graham Neubig |
J. Artif. Intell. Res. | 2 |
| 2021 | SpanNER: Named Entity Re-/Recognition as Span PredictionabstractJinlan Fu, Xuanjing Huang, Pengfei Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jinlan Fu, Xuanjing Huang 0001, Pengfei Liu 0003 |
ACL/IJCNLP (1) | 3 |
| 2021 | CitationIE: Leveraging the Citation Graph for Scientific Information ExtractionabstractVijay Viswanathan, Graham Neubig, Pengfei Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Vijay Viswanathan 0002, Graham Neubig, Pengfei Liu 0003 |
ACL/IJCNLP (1) | 3 |
| 2021 | Towards More Fine-grained and Reliable NLP Performance PredictionabstractPerformance prediction, the task of estimating a system's performance without performing experiments, allows us to reduce the experimental burden caused by the combinatorial explosion of different datasets, languages, tasks, and models.In this paper, we make two contributions to improving performance prediction for NLP tasks.First, we examine performance predictors not only for holistic measures of accuracy like F1 or BLEU, but also fine-grained performance measures such as accuracy over individual classes of examples.Second, we propose methods to understand the reliability of a performance prediction model from two angles: confidence intervals and calibration.We perform an analysis of four types of NLP tasks, and both demonstrate the feasibility of fine-grained performance prediction and the necessity to perform reliability analysis for performance prediction methods in the future.We make our code publicly available Zihuiwen Ye, Pengfei Liu 0003, Jinlan Fu, Graham Neubig |
EACL | 2 |
| 2021 | XTREME-R: Towards More Challenging and Nuanced Multilingual EvaluationabstractSebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, Melvin Johnson. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Sebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu 0003, Junjie Hu 0001, Dan Garrette, Graham Neubig, Melvin Johnson |
EMNLP (1) | 7 |
| 2021 | Does syntax matter? A strong baseline for Aspect-based Sentiment Analysis with RoBERTaabstractJunqi Dai, Hang Yan, Tianxiang Sun, Pengfei Liu, Xipeng Qiu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Junqi Dai, Hang Yan 0001, Tianxiang Sun, Pengfei Liu 0003, Xipeng Qiu |
NAACL-HLT | 4 |
| 2021 | GSum: A General Framework for Guided Neural Abstractive SummarizationabstractZi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, Graham Neubig. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zi-Yi Dou, Pengfei Liu 0003, Hiroaki Hayashi, Zhengbao Jiang, Graham Neubig |
NAACL-HLT | 2 |
| 2021 | Larger-Context Tagging: When and Why Does It Work?abstractJinlan Fu, Liangjing Feng, Qi Zhang, Xuanjing Huang, Pengfei Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jinlan Fu, Liangjing Feng, Qi Zhang 0001, Xuanjing Huang 0001, Pengfei Liu 0003 |
NAACL-HLT | 5 |
| 2021 | RefSum: Refactoring Neural SummarizationabstractAlthough some recent works show potential complementarity among different state-of-theart systems, few works try to investigate this problem in text summarization.Researchers in other areas commonly refer to the techniques of reranking or stacking to approach this problem.In this work, we highlight several limitations of previous methods, which motivates us to present a new framework Refactor that provides a unified view of text summarization and summaries combination.Experimentally, we perform a comprehensive evaluation that involves twenty-two base systems, four datasets, and three different application scenarios.Besides new state-of-the-art results on CNN/DailyMail dataset (46.18 ROUGE-1), we also elaborate on how our proposed method addresses the limitations of the traditional methods and the effectiveness of the Refactor model sheds light on insight for performance improvement.Our system can be directly used by other researchers as an offthe-shelf tool to achieve further performance improvements.We open-source all the code and provide a convenient interface to use it:https://github.com/yixinL7/ Refactoring-Summarization. Yixin Liu 0003, Zi-Yi Dou, Pengfei Liu 0003 |
NAACL-HLT | 3 |
| 2021 | BARTScore: Evaluating Generated Text as Text GenerationabstractA wide variety of NLP applications, such as machine translation, summarization, and dialog, involve text generation. One major challenge for these applications is how to evaluate whether such generated texts are actually fluent, accurate, or effective. In this work, we conceptualize the evaluation of generated text as a text generation problem, modeled using pre-trained sequence-to-sequence models. The general idea is that models trained to convert the generated text to/from a reference output or the source text will achieve higher scores when the generated text is better. We operationalize this idea using BART, an encoder-decoder based pre-trained model, and propose a metric BARTScore with a number of variants that can be flexibly applied in an unsupervised fashion to evaluation of text from different perspectives (e.g. informativeness, fluency, or factuality). BARTScore is conceptually simple and empirically effective. It can outperform existing top-scoring metrics in 16 of 22 test settings, covering evaluation of 16 datasets (e.g., machine translation, text summarization) and 7 different perspectives (e.g., informativeness, factuality). Code to calculate BARTScore is available at https://github.com/neulab/BARTScore, and we have released an interactive leaderboard for meta-evaluation at http://explainaboard.nlpedia.ai/leaderboard/task-meval/ on the ExplainaBoard platform, which allows us to interactively understand the strengths, weaknesses, and complementarity of each metric. Weizhe Yuan, Graham Neubig, Pengfei Liu 0003 |
NeurIPS | 3 |
| 2020 | Rethinking Generalization of Neural Models: A Named Entity Recognition Case StudyabstractWhile neural network-based models have achieved impressive performance on a large body of NLP tasks, the generalization behavior of different models remains poorly understood: Does this excellent performance imply a perfect generalization model, or are there still some limitations? In this paper, we take the NER task as a testbed to analyze the generalization behavior of existing models from different perspectives and characterize the differences of their generalization abilities through the lens of our proposed measures, which guides us to better design models and training methods. Experiments with in-depth analyses diagnose the bottleneck of existing neural NER models in terms of breakdown performance analysis, annotation errors, dataset bias, and category relationships, which suggest directions for improvement. We have released the datasets: (ReCoNLL, PLONER) for the future research at our project page: http://pfliu.com/InterpretNER/. Jinlan Fu, Pengfei Liu 0003, Qi Zhang 0001 |
AAAI | 2 |
| 2020 | Multi-Scale Self-Attention for Text ClassificationabstractIn this paper, we introduce the prior knowledge, multi-scale structure, into self-attention modules. We propose a Multi-Scale Transformer which uses multi-scale multi-head self-attention to capture features from different scales. Based on the linguistic perspective and the analysis of pre-trained Transformer (BERT) on a huge corpus, we further design a strategy to control the scale distribution for each layer. Results of three different kinds of tasks (21 datasets) show our Multi-Scale Transformer outperforms the standard Transformer consistently and significantly on small and moderate size datasets. Qipeng Guo, Xipeng Qiu, Pengfei Liu 0003, Xiangyang Xue 0001, Zheng Zhang 0001 |
AAAI | 3 |
| 2020 | Learning Sparse Sharing Architectures for Multiple TasksabstractMost existing deep multi-task learning models are based on parameter sharing, such as hard sharing, hierarchical sharing, and soft sharing. How choosing a suitable sharing mechanism depends on the relations among the tasks, which is not easy since it is difficult to understand the underlying shared factors among these tasks. In this paper, we propose a novel parameter sharing mechanism, named Sparse Sharing. Given multiple tasks, our approach automatically finds a sparse sharing structure. We start with an over-parameterized base network, from which each task extracts a subnetwork. The subnetworks of multiple tasks are partially overlapped and trained in parallel. We show that both hard sharing and hierarchical sharing can be formulated as particular instances of the sparse sharing framework. We conduct extensive experiments on three sequence labeling tasks. Compared with single-task models and three typical multi-task learning baselines, our proposed approach achieves consistent improvement while requiring fewer parameters. Tianxiang Sun, Yunfan Shao, Pengfei Liu 0003, Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001 |
AAAI | 4 |
| 2020 | Heterogeneous Graph Neural Networks for Extractive Document SummarizationabstractAs a crucial step in extractive document summarization, learning cross-sentence relations has been explored by a plethora of approaches.An intuitive way is to put them in the graphbased neural network, which has a more complex structure for capturing inter-sentence relationships.In this paper, we present a heterogeneous graph-based neural network for extractive summarization (HETERSUMGRAPH), which contains semantic nodes of different granularity levels apart from sentences.These additional nodes act as the intermediary between sentences and enrich the cross-sentence relations.Besides, our graph structure is flexible in natural extension from a singledocument setting to multi-document via introducing document nodes.To our knowledge, we are the first one to introduce different types of nodes into graph-based neural networks for extractive document summarization and perform a comprehensive qualitative analysis to investigate their benefits.The code will be released on Github 1 . Danqing Wang, Pengfei Liu 0003, Yining Zheng, Xipeng Qiu, Xuanjing Huang 0001 |
ACL | 2 |
| 2020 | Extractive Summarization as Text MatchingabstractThis paper creates a paradigm shift with regard to the way we build neural extractive summarization systems.Instead of following the commonly used framework of extracting sentences individually and modeling the relationship between sentences, we formulate the extractive summarization task as a semantic text matching problem, in which a source document and candidate summaries will be (extracted from the original text) matched in a semantic space.Notably, this paradigm shift to semantic matching framework is well-grounded in our comprehensive analysis of the inherent gap between sentence-level and summary-level extractors based on the property of the dataset.Besides, even instantiating the framework with a simple form of a matching model, we have driven the state-of-the-art extractive result on CNN/DailyMail to a new level (44.41 in ROUGE-1).Experiments on the other five datasets also show the effectiveness of the matching framework.We believe the power of this matching-based summarization framework has not been fully exploited.To encourage more instantiations in the future, we have released our codes, processed dataset, as well as generated summaries in https://github. com/maszhongming/MatchSum. Ming Zhong 0005, Pengfei Liu 0003, Yiran Chen 0013, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001 |
ACL | 2 |
| 2020 | Re-evaluating Evaluation in Text SummarizationabstractAutomated evaluation metrics as a stand-in for manual evaluation are an essential part of the development of text-generation tasks such as text summarization.However, while the field has progressed, our standard metrics have not -for nearly 20 years ROUGE has been the standard evaluation in most summarization papers.In this paper, we make an attempt to re-evaluate the evaluation method for text summarization: assessing the reliability of automatic metrics using top-scoring system outputs, both abstractive and extractive, on recently popular datasets for both systemlevel and summary-level evaluation settings.We find that conclusions about evaluation metrics on older datasets do not necessarily hold on modern datasets and systems.We release a dataset of human judgments that are collected from 25 top-scoring neural summarization systems (14 abstractive and 11 extractive): Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu 0003, Graham Neubig |
EMNLP (1) | 4 |
| 2020 | Interpretable Multi-dataset Evaluation for Named Entity RecognitionabstractWith the proliferation of models for natural language processing tasks, it is even harder to understand the differences between models and their relative merits.Simply looking at differences between holistic metrics such as accuracy, BLEU, or F1 does not tell us why or how particular methods perform differently and how diverse datasets influence the model design choices.In this paper, we present a general methodology for interpretable evaluation for the named entity recognition (NER) task.The proposed evaluation method enables us to interpret the differences in models and datasets, as well as the interplay between them, identifying the strengths and weaknesses of current systems.By making our analysis tool available, we make it easy for future researchers to run similar analyses and drive progress in this area: https: //github.com/neulab/InterpretEval. Jinlan Fu, Pengfei Liu 0003, Graham Neubig |
EMNLP (1) | 2 |
| 2020 | RethinkCWS: Is Chinese Word Segmentation a Solved Task?abstractThe performance of the Chinese Word Segmentation (CWS) systems has gradually reached a plateau with the rapid development of deep neural networks, especially the successful use of large pre-trained models.In this paper, we take stock of what we have achieved and rethink what's left in the CWS task.Methodologically, we propose a finegrained evaluation for existing CWS systems, which not only allows us to diagnose the strengths and weaknesses of existing models (under the in-dataset setting), but enables us to quantify the discrepancy between different criterion and alleviate the negative transfer problem when doing multi-criteria learning.Strategically, despite not aiming to propose a novel model in this paper, our comprehensive experiments on eight models and seven datasets, as well as thorough analysis, could search for some promising direction for future research.We make all codes publicly available and release an interface that can quickly evaluate and diagnose user's models: https://github. com/neulab/InterpretEval. Jinlan Fu, Pengfei Liu 0003, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP (1) | 2 |
| 2019 | Contextualized Non-Local Neural Networks for Sequence LearningabstractRecently, a large number of neural mechanisms and models have been proposed for sequence learning, of which selfattention, as exemplified by the Transformer model, and graph neural networks (GNNs) have attracted much attention. In this paper, we propose an approach that combines and draws on the complementary strengths of these two methods. Specifically, we propose contextualized non-local neural networks (CN3), which can both dynamically construct a task-specific structure of a sentence and leverage rich local dependencies within a particular neighbourhood.Experimental results on ten NLP tasks in text classification, semantic matching, and sequence labelling show that our proposed model outperforms competitive baselines and discovers task-specific dependency structures, thus providing better interpretability to users. Pengfei Liu 0003, Shuaichen Chang, Xuanjing Huang 0001, Jackie Chi Kit Cheung |
AAAI | 1 |
| 2019 | Learning Multi-Task Communication with Message Passing for Sequence LearningabstractWe present two architectures for multi-task learning with neural sequence models. Our approach allows the relationships between different tasks to be learned dynamically, rather than using an ad-hoc pre-defined structure as in previous work. We adopt the idea from message-passing graph neural networks, and propose a general graph multi-task learning framework in which different tasks can communicate with each other in an effective and interpretable way. We conduct extensive experiments in text classification and sequence labelling to evaluate our approach on multi-task learning and transfer learning. The empirical results show that our models not only outperform competitive baselines, but also learn interpretable and transferable patterns across tasks. Pengfei Liu 0003, Jie Fu 0001, Yue Dong 0002, Xipeng Qiu, Jackie Chi Kit Cheung |
AAAI | 1 |
| 2019 | TIGS: An Inference Algorithm for Text Infilling with Gradient SearchabstractText infilling is defined as a task for filling in the missing part of a sentence or paragraph, which is suitable for many real-world natural language generation scenarios.However, given a well-trained sequential generative model, generating missing symbols conditioned on the context is challenging for existing greedy approximate inference algorithms.In this paper, we propose an iterative inference algorithm based on gradient search, which is the first inference algorithm that can be broadly applied to any neural sequence generative models for text infilling tasks.We compare the proposed method with strong baselines on three text infilling tasks with various mask ratios and different mask strategies.The results show that our proposed method is effective and efficient for fill-in-the-blank tasks, consistently outperforming all baselines.1 Dayiheng Liu, Jie Fu 0001, Pengfei Liu 0003, Jiancheng Lv 0001 |
ACL (1) | 3 |
| 2019 | Searching for Effective Neural Extractive Summarization: What Works and What's NextabstractThe recent years have seen remarkable success in the use of deep neural networks on text summarization.However, there is no clear understanding of why they perform so well, or how they might be improved.In this paper, we seek to better understand how neural extractive summarization systems could benefit from different types of model architectures, transferable knowledge and learning schemas.Additionally, we find an effective way to improve current frameworks and achieve the state-ofthe-art result on CNN/DailyMail by a large margin based on our observations and analyses.Hopefully, our work could provide more clues for future research on extractive summarization.Source code will be available on Github 1 . Ming Zhong 0005, Pengfei Liu 0003, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 2 |
| 2018 | Meta Multi-Task Learning for Sequence ModelingabstractSemantic composition functions have been playing a pivotal role in neural representation learning of text sequences. In spite of their success, most existing models suffer from the underfitting problem: they use the same shared compositional function on all the positions in the sequence, thereby lacking expressive power due to incapacity to capture the richness of compositionality. Besides, the composition functions of different tasks are independent and learned from scratch. In this paper, we propose a new sharing scheme of composition function across multiple tasks. Specifically, we use a shared meta-network to capture the meta-knowledge of semantic composition and generate the parameters of the task-specific semantic composition models. We conduct extensive experiments on two types of tasks, text classification and sequence tagging, which demonstrate the benefits of our approach. Besides, we show that the shared meta-knowledge learned by our proposed model can be regarded as off-the-shelf knowledge and easily transferred to new tasks. Jun-Kun Chen, Xipeng Qiu, Pengfei Liu 0003, Xuanjing Huang 0001 |
AAAI | 3 |
| 2017 | Adversarial Multi-task Learning for Text ClassificationabstractNeural network models have shown their promising opportunities for multi-task learning, which focus on learning the shared layers to extract the common and task-invariant features.However, in most existing approaches, the extracted shared features are prone to be contaminated by task-specific features or the noise brought by other tasks.In this paper, we propose an adversarial multi-task learning framework, alleviating the shared and private latent feature spaces from interfering with each other.We conduct extensive experiments on 16 different text classification tasks, which demonstrates the benefits of our approach.Besides, we show that the shared knowledge learned by our proposed model can be regarded as off-the-shelf knowledge and easily transferred to new tasks.The datasets of all 16 tasks are publicly available at Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 1 |
| 2017 | Idiom-Aware Compositional Distributed SemanticsabstractIdioms are peculiar linguistic constructions that impose great challenges for representing the semantics of language, especially in current prevailing end-to-end neural models, which assume that the semantics of a phrase or sentence can be literally composed from its constitutive words.In this paper, we propose an idiomaware distributed semantic model to build representation of sentences on the basis of understanding their contained idioms.Our models are grounded in the literalfirst psycholinguistic hypothesis, which can adaptively learn semantic compositionality of a phrase literally or idiomatically.To better evaluate our models, we also construct an idiom-enriched sentiment classification dataset with considerable scale and abundant peculiarities of idioms.The qualitative and quantitative experimental analyses demonstrate the efficacy of our models.The newly-introduced datasets are publicly available at Pengfei Liu 0003, Kaiyu Qian, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 1 |
| 2017 | Dynamic Compositional Neural Networks over Tree StructureabstractTree-structured neural networks have proven to be effective in learning semantic representations by exploitingsyntactic information. In spite of their success, most existing models suffer from the underfitting problem: they recursively use the same shared compositional function throughout the whole compositional process and lack expressive power due to inability to capture the richness of compositionality.In this paper, we address this issue by introducing the dynamic compositional neural networks over tree structure (DC-TreeNN), in which the compositional function is dynamically generated by a meta network.The role of meta-network is to capture the metaknowledge across the different compositional rules and formulate them. Experimental results on two typical tasks show the effectiveness of the proposed models. Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 1 |
| 2017 | Adaptive Semantic Compositionality for Sentence ModellingabstractRepresenting a sentence with a fixed vector has shown its effectiveness in various NLP tasks. Most of the existing methods are based on neural network, which recursively apply different composition functions to a sequence of word vectors thereby obtaining a sentence vector.A hypothesis behind these approaches is that the meaning of any phrase can be composed of the meanings of its constituents.However, many phrases, such as idioms, are apparently non-compositional.To address this problem, we introduce a parameterized compositional switch, which outputs a scalar to adaptively determine whether the meaning of a phrase should be composed of its two constituents.We evaluate our model on five datasets of sentiment classification and demonstrate its efficacy with qualitative and quantitative experimental analysis . Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 1 |
| 2016 | Discourse Relations Detection via a Mixed Generative-Discriminative FrameworkabstractWord embeddings, which can better capture the fine-grained semantics of words, have proven to be useful for a variety of natural language processing tasks. However, because discourse structures describe the relationships between segments of discourse, word embeddings cannot be directly integrated to perform the task. In this paper, we introduce a mixed generative-discriminative framework, in which we use vector offsets between embeddings of words to represent the semantic relations between text segments and Fisher kernel framework to convert a variable number of vector offsets into a fixed length vector. In order to incorporate the weights of these offsets into the vector, we also propose the Weighted Fisher Vector. Experimental results on two different datasets show that the proposed method without using manually designed features can achieve better performance on recognizing the discourse level relations in most cases. Jifan Chen, Qi Zhang 0001, Pengfei Liu 0003, Xuanjing Huang 0001 |
AAAI | 3 |
| 2016 | Implicit Discourse Relation Detection via a Deep Architecture with Gated Relevance NetworkabstractWord pairs, which are one of the most easily accessible features between two text segments, have been proven to be very useful for detecting the discourse relations held between text segments.However, because of the data sparsity problem, the performance achieved by using word pair features is limited.In this paper, in order to overcome the data sparsity problem, we propose the use of word embeddings to replace the original words.Moreover, we adopt a gated relevance network to capture the semantic interaction between word pairs, and then aggregate those semantic interactions using a pooling layer to select the most informative interactions.Experimental results on Penn Discourse Tree Bank show that the proposed method without using manually designed features can achieve better performance on recognizing the discourse level relations in all of the relations. Jifan Chen, Qi Zhang 0001, Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 3 |
| 2016 | Deep Fusion LSTMs for Text Semantic MatchingabstractRecently, there is rising interest in modelling the interactions of text pair with deep neural networks.In this paper, we propose a model of deep fusion LSTMs (DF-LSTMs) to model the strong interaction of text pair in a recursive matching way.Specifically, DF-LSTMs consist of two interdependent LSTMs, each of which models a sequence under the influence of another.We also use external memory to increase the capacity of LSTMs, thereby possibly capturing more complicated matching patterns.Experiments on two very large datasets demonstrate the efficacy of our proposed architecture.Furthermore, we present an elaborate qualitative analysis of our models, giving an intuitive understanding how our model worked. Pengfei Liu 0003, Xipeng Qiu, Jifan Chen, Xuanjing Huang 0001 |
ACL (1) | 1 |
| 2016 | Deep Multi-Task Learning with Shared Memory for Text Classification
Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 1 |
| 2016 | Modelling Interaction of Sentence Pair with Coupled-LSTMsabstractRecently, there is rising interest in modelling the interactions of two sentences with deep neural networks.However, most of the existing methods encode two sequences with separate encoders, in which a sentence is encoded with little or no information from the other sentence.In this paper, we propose a deep architecture to model the strong interaction of sentence pair with two coupled-LSTMs.Specifically, we introduce two coupled ways to model the interdependences of two LSTMs, coupling the local contextualized interactions of two sentences.We then aggregate these interactions and use a dynamic pooling to select the most informative features.Experiments on two very large datasets demonstrate the efficacy of our proposed architectures. Pengfei Liu 0003, Xipeng Qiu, Yaqian Zhou 0001, Jifan Chen, Xuanjing Huang 0001 |
EMNLP | 1 |
| 2016 | Recurrent Neural Network for Text Classification with Multi-Task Learning
Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 1 |
| 2015 | Long Short-Term Memory Neural Networks for Chinese Word SegmentationabstractCurrently most of state-of-the-art methods for Chinese word segmentation are based on supervised learning, whose features are mostly extracted from a local context.These methods cannot utilize the long distance information which is also crucial for word segmentation.In this paper, we propose a novel neural network model for Chinese word segmentation, which adopts the long short-term memory (LSTM) neural network to keep the previous important information in memory cell and avoids the limit of window size of local context.Experiments on PKU, MSRA and CTB6 benchmark datasets show that our model outperforms the previous neural network models and state-of-the-art methods. Xinchi Chen, Xipeng Qiu, Pengfei Liu 0003, Xuanjing Huang 0001 |
EMNLP | 4 |
| 2015 | Multi-Timescale Long Short-Term Memory Neural Network for Modelling Sentences and DocumentsabstractNeural network based methods have obtained great progress on a variety of natural language processing tasks.However, it is still a challenge task to model long texts, such as sentences and documents.In this paper, we propose a multi-timescale long short-term memory (MT-LSTM) neural network to model long texts.MT-LSTM partitions the hidden states of the standard LSTM into several groups.Each group is activated at different time periods.Thus, MT-LSTM can model very long documents as well as short sentences.Experiments on four benchmark datasets show that our model outperforms the other neural models in text classification task. Pengfei Liu 0003, Xipeng Qiu, Xinchi Chen, Shiyu Wu, Xuanjing Huang 0001 |
EMNLP | 1 |
| 2015 | Learning Context-Sensitive Word Embeddings with Neural Tensor Skip-Gram Model
Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 1 |