EDBT 2026 Demo / reviewers in the wild / expert
Ming Zhong 0005
dblp:92/2292-5
· DBLP profile ↗
24ranked-venue papers
7as first author
22since 2021 · last 2025
0000-0001-5728-0224ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 20 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Long Chain-of-Thought Fine-tuning via Understanding-to-Reasoning TransitionabstractChenxin An, Zhihui Xie, Xiaonan Li, Ming Zhong, Shansan Gong, Lei Li, Jun Zhang, Jingjing Xu, Lingpeng Kong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chenxin An, Zhihui Xie 0002, Ming Zhong 0005, Shansan Gong, Lei Li 0039, Jun Zhang 0003, Jingjing Xu 0001, Lingpeng Kong |
EMNLP | 4 |
| 2025 | Law of the Weakest Link: Cross Capabilities of Large Language ModelsabstractThe development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term **cross capabilities**. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce *CrossEval*, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that current LLMs consistently exhibit the ``Law of the Weakest Link,'' where cross-capability performance is significantly constrained by the weakest component. Across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight LLMs' underperformance in cross-capability tasks, emphasizing the need to identify and improve their weakest capabilities as a key research priority. The code, benchmarks, and evaluations are available on our [project website](https://www.llm-cross-capabilities.org). Ming Zhong 0005, Aston Zhang, Wenhan Xiong, Chenguang Zhu 0001, Zhengxing Chen, Chloe Bi, Mike Lewis, Sravya Popuri, Sharan Narang, Melanie Kambadur, Dhruv Mahajan 0001, Sergey Edunov, Jiawei Han 0001, Laurens van der Maaten |
ICLR | 1 |
| 2025 | Why Does the Effective Context Length of LLMs Fall Short?abstractAdvancements in distributed training and efficient attention mechanisms have significantly expanded the context window sizes of large language models (LLMs). However, recent work reveals that the effective context lengths of open-source LLMs often fall short, typically not exceeding half of their training lengths. In this work, we attribute this limitation to the left-skewed frequency distribution of relative positions formed in LLMs pretraining and post-training stages, which impedes their ability to effectively gather distant information.
To address this challenge, we introduce Shifted Rotray Position Embedding (STRING). STRING shifts well-trained positions to overwrite the original ineffective positions during inference, enhancing performance within their existing training lengths.
Experimental results show that without additional training, STRING dramatically improves the performance of the latest large-scale models, such as Llama3.1 70B and Qwen2 72B, by over 10 points on popular long-context benchmarks RULER and InfiniteBench, establishing new state-of-the-art results for open-source LLMs. Compared to commercial models, Llama 3.1 70B with STRING even achieves better performance than GPT-4-128K and clearly surpasses Claude 2 and Kimi-chat. Chenxin An, Jun Zhang 0003, Ming Zhong 0005, Lei Li 0039, Shansan Gong, Yao Luo, Jingjing Xu 0001, Lingpeng Kong |
ICLR | 3 |
| 2025 | Retrieval And Structuring Augmented Generation with Large Language ModelsabstractLarge Language Models (LLMs) have revolutionized natural language processing with their remarkable capabilities in text generation and reasoning. However, these models face critical challenges when deployed in real-world applications, including hallucination generation, outdated knowledge, and limited domain expertise. Retrieval And Structuring (RAS) Augmented Generation addresses these limitations by integrating dynamic information retrieval with structured knowledge representations. This survey (1) examines retrieval mechanisms including sparse, dense, and hybrid approaches for accessing external knowledge; (2) explore text structuring techniques such as taxonomy construction, hierarchical classification, and information extraction that transform unstructured text into organized representations; and (3) investigate how these structured representations integrate with LLMs through prompt-based methods, reasoning frameworks, and knowledge embedding techniques. It also identifies technical challenges in retrieval efficiency, structure quality, and knowledge integration, while highlighting research opportunities in multimodal retrieval, cross-lingual structures, and interactive systems. This comprehensive overview provides researchers and practitioners with insights into RAS methods, applications, and future directions. Pengcheng Jiang, Siru Ouyang, Yizhu Jiao, Ming Zhong 0005, Runchu Tian, Jiawei Han 0001 |
KDD (2) | 4 |
| 2025 | Multimodal Search in Chemical Documents and ReactionsabstractWe present a multimodal search tool for retrieval of chemical reactions, molecular structures, and associated text from scientific literature.Queries may combine molecular diagrams, textual descriptions, and reaction data, allowing users to connect different chemical information representations.Indexing includes chemical diagram extraction and parsing, extraction of reaction data from text in tabular form, and cross-modal linking of diagrams with their mentions in text.We describe the system's architecture and retrieval features, along with expert assessments of the system.Our demo highlights the workflow and search components.Online demo: https://www.cs.rit.edu/ Ayush Kumar Shah, Abhisek Dey, Leo Luo, Bryan Amador, Patrick Philippy, Ming Zhong 0005, Siru Ouyang, David Mark Friday, David Bianchi, Nick Jackson, Richard Zanibbi, Jiawei Han 0001 |
SIGIR | 6 |
| 2024 | L-Eval: Instituting Standardized Evaluation for Long Context Language ModelsabstractChenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao, Mukai Li, Jun Zhang, Lingpeng Kong, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chenxin An, Shansan Gong, Ming Zhong 0005, Xingjian Zhao, Mukai Li, Jun Zhang 0003, Lingpeng Kong, Xipeng Qiu |
ACL (1) | 3 |
| 2024 | ActionIE: Action Extraction from Scientific Literature with Programming LanguagesabstractXianrui Zhong, Yufeng Du, Siru Ouyang, Ming Zhong, Tingfeng Luo, Qirong Ho, Hao Peng, Heng Ji, Jiawei Han. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xianrui Zhong, Yufeng Du, Siru Ouyang, Ming Zhong 0005, Tingfeng Luo, Qirong Ho, Hao Peng 0009, Heng Ji 0001, Jiawei Han 0001 |
ACL (1) | 4 |
| 2024 | Seeking Neural Nuggets: Knowledge Transfer in Large Language Models from a Parametric PerspectiveabstractLarge Language Models (LLMs) inherently encode a wealth of knowledge within their parameters through pre-training on extensive corpora. While prior research has delved into operations on these parameters to manipulate the underlying implicit knowledge — encompassing detection, editing, and merging — there remains an ambiguous understanding regarding their transferability across models with varying scales. In this paper, we seek to empirically investigate knowledge transfer from larger to smaller models through a parametric perspective. To achieve this, we employ sensitivity-based techniques to extract and align knowledge-specific parameters between different LLMs. Moreover, the LoRA module is used as the intermediary mechanism for injecting the extracted knowledge into smaller models. Evaluations across four benchmarks validate the efficacy of our proposed method. Our findings highlight the critical factors contributing to the process of parametric knowledge transfer, underscoring the transferability of model parameters across LLMs of different scales. Project website: https://maszhongming.github.io/ParaKnowTransfer. Ming Zhong 0005, Chenxin An, Weizhu Chen, Jiawei Han 0001 |
ICLR | 1 |
| 2024 | Automated Mining of Structured Knowledge from Text in the Era of Large Language ModelsabstractMassive amount of unstructured text data are generated daily, ranging from news articles to scientific papers. How to mine structured knowledge from the text data remains a crucial research question. Recently, large language models (LLMs) have shed light on the text mining field with their superior text understanding and instruction-following ability. There are typically two ways of utilizing LLMs: fine-tune the LLMs with human-annotated training data, which is labor intensive and hard to scale; prompt the LLMs in a zero-shot or few-shot way, which cannot take advantage of the useful information in the massive text data. Therefore, it remains a challenge on automated mining of structured knowledge from massive text data in the era of large language models. Yunyi Zhang 0001, Ming Zhong 0005, Siru Ouyang, Yizhu Jiao, Sizhe Zhou, Linyi Ding, Jiawei Han 0001 |
KDD | 2 |
| 2023 | Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved AnnotationabstractYulong Chen, Huajian Zhang, Yijie Zhou, Xuefeng Bai, Yueguan Wang, Ming Zhong, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yulong Chen 0001, Xuefeng Bai 0001, Yueguan Wang, Ming Zhong 0005, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang 0004 |
ACL (1) | 6 |
| 2023 | Instruct and Extract: Instruction Tuning for On-Demand Information ExtractionabstractLarge language models with instructionfollowing capabilities open the door to a wider group of users.However, when it comes to information extraction -a classic task in natural language processing -most task-specific systems cannot align well with long-tail ad hoc extraction use cases for non-expert users.To address this, we propose a novel paradigm, termed On-Demand Information Extraction, to fulfill the personalized demands of real-world users.Our task aims to follow the instructions to extract the desired content from the associated text and present it in a structured tabular format.The table headers can either be userspecified or inferred contextually by the model.To facilitate research in this emerging area, we present a benchmark named INSTRUCTIE, inclusive of both automatically generated training data, as well as the human-annotated test set.Building on INSTRUCTIE, we further develop an On-Demand Information Extractor, ODIE.Comprehensive evaluations on our benchmark reveal that ODIE substantially outperforms the existing open-source models of similar size.Our code and dataset are released on https://github.com/yzjiao/On-Demand-IE. Yizhu Jiao, Ming Zhong 0005, Ruining Zhao, Siru Ouyang, Heng Ji 0001, Jiawei Han 0001 |
EMNLP | 2 |
| 2023 | The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT InteractionsabstractSiru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, Jiawei Han. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Siru Ouyang, Shuohang Wang, Ming Zhong 0005, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu 0001, Heng Ji 0001, Jiawei Han 0001 |
EMNLP | 4 |
| 2023 | Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data CurationabstractInstruction tuning has emerged to enhance the capabilities of large language models (LLMs) to comprehend instructions and generate appropriate responses.Existing methods either manually annotate or employ LLM (e.g., GPTseries) to generate data for instruction tuning.However, they often overlook associating instructions with existing annotated datasets.In this paper, we propose DYNOSAUR, a dynamic growth paradigm for the automatic curation of instruction-tuning data.Based on the metadata of existing datasets, we use LLMs to automatically construct instruction-tuning data by identifying relevant data fields and generating appropriate instructions.By leveraging the existing annotated datasets, DYNOSAUR offers several advantages: 1) it reduces the API cost for generating instructions (e.g., it costs less than $12 USD by calling GPT-3.5-turbo for generating 800K instruction tuning samples; 2) it provides high-quality data for instruction tuning (e.g., it performs better than ALPACA and FLAN on SUPER-NI and LONGFORM with comparable data sizes); and 3) it supports the continuous improvement of models by generating instruction-tuning data when a new annotated dataset becomes available.We further investigate a continual learning scheme for learning with the ever-growing instruction-tuning dataset, and demonstrate that replaying tasks with diverse instruction embeddings not only helps mitigate forgetting issues but generalizes to unseen tasks better. Da Yin, Xiao Liu 0032, Fan Yin, Ming Zhong 0005, Hritik Bansal, Jiawei Han 0001, Kai-Wei Chang 0001 |
EMNLP | 4 |
| 2023 | Unsupervised Event Chain Mining from Multiple DocumentsabstractMassive and fast-evolving news articles keep emerging on the web. To effectively summarize and provide concise insights into real-world events, we propose a new event knowledge extraction task Event Chain Mining in this paper. Given multiple documents about a super event, it aims to mine a series of salient events in temporal order. For example, the event chain of super event Mexico Earthquake in 2017 is {earthquake hit Mexico, destroy houses, kill people, block roads}. This task can help readers capture the gist of texts quickly, thereby improving reading efficiency and deepening text comprehension. To address this task, we regard an event as a cluster of different mentions of similar meanings. In this way, we can identify the different expressions of events, enrich their semantic knowledge and replenish relation information among them. Taking events as the basic unit, we present a novel unsupervised framework, EMiner. Specifically, we extract event mentions from texts and merge them with similar meanings into a cluster as a single event. By jointly incorporating both content and commonsense, essential events are then selected and arranged chronologically to form an event chain. Meanwhile, we annotate a multi-document benchmark to build a comprehensive testbed for the proposed task. Extensive experiments are conducted to verify the effectiveness of EMiner in terms of both automatic and human evaluations. Yizhu Jiao, Ming Zhong 0005, Yunyi Zhang 0001, Chao Zhang 0014, Jiawei Han 0001 |
WWW | 2 |
| 2022 | DialogLM: Pre-trained Model for Long Dialogue Understanding and SummarizationabstractDialogue is an essential part of human communication and cooperation. Existing research mainly focuses on short dialogue scenarios in a one-on-one fashion. However, multi-person interactions in the real world, such as meetings or interviews, are frequently over a few thousand words. There is still a lack of corresponding research and powerful tools to understand and process such long dialogues. Therefore, in this work, we present a pre-training framework for long dialogue understanding and summarization. Considering the nature of long conversations, we propose a window-based denoising approach for generative pre-training. For a dialogue, it corrupts a window of text with dialogue-inspired noise, and guides the model to reconstruct this window based on the content of the remaining conversation. Furthermore, to process longer input, we augment the model with sparse attention which is combined with conventional attention in a hybrid manner. We conduct extensive experiments on five datasets of long dialogues, covering tasks of dialogue summarization, abstractive question answering and topic segmentation. Experimentally, we show that our pre-trained model DialogLM significantly surpasses the state-of-the-art models across datasets and tasks. Source code and all the pre-trained models are available on our GitHub repository (https://github.com/microsoft/DialogLM). Ming Zhong 0005, Yang Liu 0124, Yichong Xu, Chenguang Zhu 0001, Michael Zeng 0001 |
AAAI | 1 |
| 2022 | CoLo: A Contrastive Learning Based Re-ranking Framework for One-Stage SummarizationabstractTraditional training paradigms for extractive and abstractive summarization systems always only use token-level or sentence-level training objectives. However, the output summary is always evaluated from summary-level which leads to the inconsistency in training and evaluation. In this paper, we propose a Contrastive Learning based re-ranking framework for one-stage summarization called CoLo. By modeling a contrastive objective, we show that the summarization model is able to directly generate summaries according to the summary-level score without additional modules and parameters. Extensive experiments demonstrate that CoLo boosts the extractive and abstractive results of one-stage systems on CNN/DailyMail benchmark to 44.58 and 46.33 ROUGE-1 score while preserving the parameter efficiency and inference efficiency. Compared with state-of-the-art multi-stage systems, we save more than 100 GPU training hours and obtaining 3x 8x speed-up ratio during inference while maintaining comparable results. Chenxin An, Ming Zhong 0005, Zhiyong Wu 0003, Xuanjing Huang 0001, Xipeng Qiu |
COLING | 2 |
| 2022 | Improving Abstractive Dialogue Summarization with Speaker-Aware Supervised Contrastive LearningabstractPre-trained models have brought remarkable success on the text summarization task. For dialogue summarization, the subdomain of text summarization, utterances are concatenated to flat text before being processed. As a result, existing summarization systems based on pre-trained models are unable to recognize the unique format of the speaker-utterance pair well in the dialogue. To investigate this issue, we conduct probing tests and manual analysis, and find that the powerful pre-trained model can not identify different speakers well in the conversation, which leads to various factual errors. Moreover, we propose three speaker-aware supervised contrastive learning (SCL) tasks: Token-level SCL, Turn-level SCL, and Global-level SCL. Comprehensive experiments demonstrate that our methods achieve significant performance improvement on two mainstream dialogue summarization datasets. According to detailed human evaluations, pre-trained models equipped with SCL tasks effectively generate summaries with better factual consistency. Zhichao Geng, Ming Zhong 0005, Zhangyue Yin, Xipeng Qiu, Xuanjing Huang 0001 |
COLING | 2 |
| 2022 | CiteSum: Citation Text-guided Scientific Extreme Summarization and Domain Adaptation with Limited SupervisionabstractScientific extreme summarization (TLDR) aims to form ultra-short summaries of scientific papers.Previous efforts on curating scientific TLDR datasets failed to scale up due to the heavy human annotation and domain expertise required.In this paper, we propose a simple yet effective approach to automatically extracting TLDR summaries for scientific papers from their citation texts.Based on the proposed approach, we create a new benchmark CiteSum without human annotation, which is around 30 times larger than the previous human-curated dataset SciTLDR.We conduct a comprehensive analysis of CiteSum, examining its data characteristics and establishing strong baselines.We further demonstrate the usefulness of CiteSum by adapting models pre-trained on CiteSum (named CITES) to new tasks and domains with limited supervision.For scientific extreme summarization, CITES outperforms most fully-supervised methods on SciTLDR without any fine-tuning and obtains state-of-theart results with only 128 examples.For news extreme summarization, CITES achieves significant gains on XSum over its base model (not pre-trained on CiteSum), e.g., +7.2 ROUGE-1 zero-shot performance and state-of-the-art few-shot performance.For news headline generation, CITES performs the best among unsupervised and zero-shot methods on Gigaword. 1 Yuning Mao, Ming Zhong 0005, Jiawei Han 0001 |
EMNLP | 2 |
| 2022 | UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsabstractTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009 |
EMNLP | 8 |
| 2022 | Towards a Unified Multi-Dimensional Evaluator for Text GenerationabstractMulti-dimensional evaluation is the dominant paradigm for human evaluation in Natural Language Generation (NLG), i.e., evaluating the generated text from multiple explainable dimensions, such as coherence and fluency.However, automatic evaluation in NLG is still dominated by similarity-based metrics, and we lack a reliable framework for a more comprehensive evaluation of advanced models.In this paper, we propose a unified multi-dimensional evaluator UNIEVAL for NLG.We re-frame NLG evaluation as a Boolean Question Answering (QA) task, and by guiding the model with different questions, we can use one evaluator to evaluate from multiple dimensions.Furthermore, thanks to the unified Boolean QA format, we are able to introduce an intermediate learning phase that enables UNIEVAL to incorporate external knowledge from multiple related tasks and gain further improvement.Experiments on three typical NLG tasks show that UNIEVAL correlates substantially better with human judgments than existing metrics.Specifically, compared to the top-performing unified evaluators, UNIEVAL achieves a 23% higher correlation on text summarization, and over 43% on dialogue response generation.Also, UNIEVAL demonstrates a strong zero-shot learning ability for unseen evaluation dimensions and tasks.Source code, data and all pre-trained evaluators are available on our GitHub repository 1 . Generated Summary:Harry Kane is nominated for both the PFA player and young player of the season.The Spurs striker has been released from the awards ceremony on Sunday.The Tottenham striker features in a new animation.Reference Summary: Harry Kane has been in superb form for Tottenham this season.The 21-year-old has scored 30 goals in all competitions for Spurs.Kane also made his England debut and scored within two minutes.Document: Harry Kane's celebrations this season have always shown him to be an animated young man . . .Similarity-based Evaluators ROUGE-1: 0.44 ROUGE-2: 0.25 ROUGE-L: 0.42 BERTScore: 0.24 Single-dimensional Evaluators (predicted by two different evaluators (Deng et al., 2021)) Consistency: 0.87 Relevance: 0.74 Unified Evaluator (predicted by BARTScore, and the scoring range is negative infinity to 0) Precision: -5.45 Recall: -4.93 F1: -5.19 Ming Zhong 0005, Yang Liu 0005, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu 0003, Chenguang Zhu 0001, Heng Ji 0001, Jiawei Han 0001 |
EMNLP | 1 |
| 2021 | Enhancing Scientific Papers Summarization with Citation GraphabstractPrevious work for text summarization in scientific domain mainly focused on the content of the input document, but seldom considering its citation network. However, scientific papers are full of uncommon domain-specific terms, making it almost impossible for the model to understand its true meaning without the help of the relevant research community. In this paper, we redefine the task of scientific papers summarization by utilizing their citation graph and propose a citation graph-based summarization model CGSum which can incorporate the information of both the source paper and its references. In addition, we construct a novel scientific papers summarization dataset Semantic Scholar Network (SSN) which contains 141K research papers in different domains and 661K citation relationships. The entire dataset constitutes a large connected citation graph. Extensive experiments show that our model can achieve competitive performance when compared with the pretrained models even with a simple architecture. The results also indicates the citation graph is crucial to better understand the content of papers and generate high-quality summaries. Chenxin An, Ming Zhong 0005, Yiran Chen 0013, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001 |
AAAI | 2 |
| 2021 | QMSum: A New Benchmark for Query-based Multi-domain Meeting SummarizationabstractMing Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, Dragomir Radev. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Ming Zhong 0005, Da Yin, Tao Yu 0009, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Awadallah 0001, Asli Celikyilmaz, Yang Liu 0124, Xipeng Qiu, Dragomir R. Radev |
NAACL-HLT | 1 |
| 2020 | Extractive Summarization as Text MatchingabstractThis paper creates a paradigm shift with regard to the way we build neural extractive summarization systems.Instead of following the commonly used framework of extracting sentences individually and modeling the relationship between sentences, we formulate the extractive summarization task as a semantic text matching problem, in which a source document and candidate summaries will be (extracted from the original text) matched in a semantic space.Notably, this paradigm shift to semantic matching framework is well-grounded in our comprehensive analysis of the inherent gap between sentence-level and summary-level extractors based on the property of the dataset.Besides, even instantiating the framework with a simple form of a matching model, we have driven the state-of-the-art extractive result on CNN/DailyMail to a new level (44.41 in ROUGE-1).Experiments on the other five datasets also show the effectiveness of the matching framework.We believe the power of this matching-based summarization framework has not been fully exploited.To encourage more instantiations in the future, we have released our codes, processed dataset, as well as generated summaries in https://github. com/maszhongming/MatchSum. Ming Zhong 0005, Pengfei Liu 0003, Yiran Chen 0013, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001 |
ACL | 1 |
| 2019 | Searching for Effective Neural Extractive Summarization: What Works and What's NextabstractThe recent years have seen remarkable success in the use of deep neural networks on text summarization.However, there is no clear understanding of why they perform so well, or how they might be improved.In this paper, we seek to better understand how neural extractive summarization systems could benefit from different types of model architectures, transferable knowledge and learning schemas.Additionally, we find an effective way to improve current frameworks and achieve the state-ofthe-art result on CNN/DailyMail by a large margin based on our observations and analyses.Hopefully, our work could provide more clues for future research on extractive summarization.Source code will be available on Github 1 . Ming Zhong 0005, Pengfei Liu 0003, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 1 |