Minlong Peng

dblp:205/9010 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
abstract
Existing reinforcement learning methods for Chain-of-Thought reasoning suffer from two critical limitations. First, they operate as monolithic black boxes that provide undifferentiated reward signals, obscuring individual step contributions and hindering error diagnosis. Second, sequential decoding has O(n) time complexity. This makes real-time deployment impractical for complex reasoning tasks. We present DeCoRL (Decoupled Reasoning Chains via Coordinated Reinforcement Learning), a novel framework that transforms reasoning from sequential processing into collaborative modular orchestration. DeCoRL trains lightweight specialized models to generate reasoning sub-steps concurrently, eliminating sequential bottlenecks through parallel processing. To enable precise error attribution, the framework designs modular reward functions that score each sub-step independently. Cascaded DRPO optimization then coordinates these rewards while preserving inter-step dependencies. Comprehensive evaluation demonstrates state-of-the-art results across RM-Bench, RMB, and RewardBench, outperforming existing methods including large-scale models. DeCoRL delivers 3.8 times faster inference while maintaining superior solution quality and offers a 22.7% improvement in interpretability through explicit reward attribution. These advancements, combined with a 72.4% reduction in energy consumption and a 68% increase in throughput, make real-time deployment of complex reasoning systems a reality.
Ziyuan Gao, Di Liang, Xianjie Wu, Philippe Morel, Minlong Peng
AAAI5
2026 Reinforcement Learning Enhanced Muti-hop Reasoning for Temporal Knowledge Question Answering
abstract
Temporal knowledge graph question answering (TKGQA) involves multi-hop reasoning over temporally constrained entity relationships in the knowledge graph to answer a given question. However, at each hop, large language models (LLMs) retrieve subgraphs with numerous temporally similar and semantically complex relations, increasing the risk of suboptimal decisions and error propagation. To address these challenges, we propose the multi-hop reasoning enhanced (MRE) framework, which enhances both forward and backward reasoning to improve the identification of globally optimal reasoning trajectories. Specifically, MRE begins with prompt engineering to guide LLM in generating diverse reasoning trajectories for the given question. Valid reasoning trajectories are then selected for supervised fine-tuning, serving as a cold-start strategy. Finally, we introduce Tree-Group Relative Policy Optimization (T-GRPO)—a recursive, tree-structured learning-by-exploration approach. At each hop, exploration establishes strong causal dependencies on the previous hop, while evaluation is informed by multi-path exploration feedback from subsequent hops. Experimental results on two TKGQA benchmarks indicate that the proposed MRE-based model consistently surpasses state-of-the-art (SOTA) approaches in handling complex multi-hop queries. Further analysis highlights improved interpretability and robustness to noisy temporal annotations.
Wuzhenghong Wen, Yuwei Sun, Minlong Peng
AAAI5
2026 Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning
abstract
Zekai Lin, Chao Xue, Di Liang, Xingsheng Han, Peiyang Liu, Xianjie Wu, Lei Jiang, Yu Lu, Bob Simons, Shuang Liang, Minlong Peng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zekai Lin, Di Liang, Xingsheng Han, Peiyang Liu, Xianjie Wu, Bob Simons, Minlong Peng
ACL (1)11
2026 Lingua-Graph: A Unified Representation of Cross-Task Common Substructures for Analytic Language Processing
abstract
Structural understanding of natural language requires explicit recovery of internal meaning structures (entities, facts, nested relations), yet current structural-analytic tasks are fragmented by inconsistent task requirements across datasets.We investigate the problem of robust cross-task structural understanding under heterogeneous requirements across structural-analytic tasks and outline a perspective called Analytic NLP in which tasks can be reformulated into a representation-thendecision paradigm.In this paper, we suggest a solution for the representation layer, called Lingua-Graph, which explicitly captures entities, facts, and relations.By representing predictions as explicit graphs with labeled nodes and edges, Lingua-Graph also improves interpretability, enabling transparent inspection and error analysis of intermediate meaning structures.We construct a labeled Lingua-Graph dataset and train a baseline parser.Experiments show that Lingua-Graph provides substantially higher entity-structure hostability than alternative representations on average, and OpenIE systems based on Lingua-Graph achieve superior performance on three benchmarks, demonstrating that better intermediate structures translate into downstream gains.The data, code and the trained model are publicly released at https://github.com/rudaoshi/Lingua.
Mingming Sun 0001, Runze Jiang, Zhu Zhangchenxi, Minlong Peng, Yunfeng Cai
ACL (1)4
2026 Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models
abstract
Chao Xue, Yao Wang, Mengqiao Liu, Di Liang, Xingsheng Han, Peiyang Liu, Xianjie Wu, Chenyao Lu, Lei Jiang, Yu Lu, Haibo Shi, Shuang Liang, Minlong Peng, Flora D. Salim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mengqiao Liu, Di Liang, Xingsheng Han, Peiyang Liu, Xianjie Wu, Chenyao Lu, Haibo Shi, Minlong Peng, Flora D. Salim
ACL (1)13
2026 DocTER: Evaluating document-based knowledge editing
Suhang Wu, Ante Wang, Minlong Peng, Yujie Lin 0003, Mingming Sun 0001, Jinsong Su
Inf. Process. Manag.3
2025 Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
abstract
Supervised fine-tuning (SFT) is a pivotal approach to adapting large language models (LLMs) for downstream tasks; however, performance often suffers from the "seesaw phenomenon", where indiscriminate parameter updates yield progress on certain tasks at the expense of others.To address this challenge, we propose a novel Core Parameter Isolation Fine-Tuning (CPI-FT) framework.Specifically, we first independently fine-tune the LLM on each task to identify its core parameter regions by quantifying parameter update magnitudes.Tasks with similar core regions are then grouped based on region overlap, forming clusters for joint modeling.We further introduce a parameter fusion technique: for each task, core parameters from its individually finetuned model are directly transplanted into a unified backbone, while non-core parameters from different tasks are smoothly integrated via Spherical Linear Interpolation (SLERP), mitigating destructive interference.A lightweight, pipelined SFT training phase using mixed-task data is subsequently employed, while freezing core regions from prior tasks to prevent catastrophic forgetting.Extensive experiments on multiple public benchmarks demonstrate that our approach significantly alleviates task interference and forgetting, consistently outperforming vanilla multi-task and multi-stage finetuning baselines.
Di Liang, Minlong Peng
EMNLP3
2024 One2Set + Large Language Model: Best Partners for Keyphrase Generation
abstract
Keyphrase generation (KPG) aims to automatically generate a collection of phrases representing the core concepts of a given document.The dominant paradigms in KPG include ONE2SEQ and ONE2SET.Recently, there has been increasing interest in applying large language models (LLMs) to KPG.Our preliminary experiments reveal that it is challenging for a single model to excel in both recall and precision.Further analysis shows that: 1) the ONE2SET paradigm owns the advantage of high recall, but suffers from improper assignments of supervision signals during training; 2) LLMs are powerful in keyphrase selection, but existing selection methods often make redundant selections.Given these observations, we introduce a generate-then-select framework decomposing KPG into two steps, where we adopt a ONE2SET-based model as generator to produce candidates and then use an LLM as selector to select keyphrases from these candidates.Particularly, we make two important improvements on our generator and selector: 1) we design an Optimal Transport-based assignment strategy to address the above improper assignments; 2) we model the keyphrase selection as a sequence labeling task to alleviate redundant selections.Experimental results on multiple benchmark datasets show that our framework significantly surpasses state-of-theart models, especially in absent keyphrase prediction.We release our code at https:// github.
Liangying Shao, Minlong Peng, Guoqi Ma, Mingming Sun 0001, Jinsong Su
EMNLP3
2023 Actively Supervised Clustering for Open Relation Extraction
abstract
Jun Zhao, Yongxin Zhang, Qi Zhang, Tao Gui, Zhongyu Wei, Minlong Peng, Mingming Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jun Zhao 0019, Qi Zhang 0001, Tao Gui, Zhongyu Wei, Minlong Peng, Mingming Sun 0001
ACL (1)6
2023 RE-Matching: A Fine-Grained Semantic Matching Method for Zero-Shot Relation Extraction
abstract
Jun Zhao, WenYu Zhan, Xin Zhao, Qi Zhang, Tao Gui, Zhongyu Wei, Junzhe Wang, Minlong Peng, Mingming Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jun Zhao 0019, Wenyu Zhan, Qi Zhang 0001, Tao Gui, Zhongyu Wei, Junzhe Wang 0001, Minlong Peng, Mingming Sun 0001
ACL (1)8
2022 OIE@OIA: an Adaptable and Efficient Open Information Extraction Framework
abstract
Different Open Information Extraction (OIE)tasks require different types of information, so the OIE field requires strong adaptability of OIE algorithms to meet different task requirements.This paper discusses the adaptability problem in existing OIE systems and designs a new adaptable and efficient OIE system -OIE@OIA as a solution.OIE@OIA follows the methodology of Open Information eXpression (OIX): parsing a sentence to an Open Information Annotation (OIA) Graph and then adapting the OIA graph to different OIE tasks with simple rules.As the core of our OIE@OIA system, we implement an endto-end OIA generator by annotating a dataset (we make it open available) and designing an efficient learning algorithm for the complex OIA graph.We easily adapt the OIE@OIA system to accomplish three popular OIE tasks.The experimental show that our OIE@OIA achieves new SOTA performances on these tasks, showing the great adaptability of our OIE@OIA system.Furthermore, compared to other end-to-end OIE baselines that need millions of samples for training, our OIE@OIA needs much fewer training samples (12K), showing a significant advantage in terms of efficiency.
Xin Wang 0017, Minlong Peng, Mingming Sun 0001, Ping Li 0001
ACL (1)2
2021 Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic Modeling
abstract
In a customer service system, dialogue summarization can boost service efficiency by automatically creating summaries for long spoken dialogues in which customers and agents try to address issues about specific topics. In this work, we focus on topic-oriented dialogue summarization, which generates highly abstractive summaries that preserve the main ideas from dialogues. In spoken dialogues, abundant dialogue noise and common semantics could obscure the underlying informative content, making the general topic modeling approaches difficult to apply. In addition, for customer service, role-specific information matters and is an indispensable part of a summary. To effectively perform topic modeling on dialogues and capture multi-role information, in this work we propose a novel topic-augmented two-stage dialogue summarizer (TDS) jointly with a saliency-aware neural topic model (SATM) for topic-oriented summarization of customer service dialogues. Comprehensive studies on a real-world Chinese customer service dataset demonstrated the superiority of our method against several strong baselines.
Yicheng Zou, Lujun Zhao, Yangyang Kang, Minlong Peng, Zhuoren Jiang, Changlong Sun, Qi Zhang 0001, Xuanjing Huang 0001, Xiaozhong Liu 0001
AAAI5
2020 Simplify the Usage of Lexicon in Chinese NER
abstract
Recently, many works have tried to augment the performance of Chinese named entity recognition (NER) using word lexicons.As a representative, Lattice-LSTM (Zhang and Yang, 2018) has achieved new benchmark results on several public Chinese NER datasets.However, Lattice-LSTM has a complex model architecture.This limits its application in many industrial areas where real-time NER responses are needed.In this work, we propose a simple but effective method for incorporating the word lexicon into the character representations.This method avoids designing a complicated sequence modeling architecture, and for any neural NER model, it requires only subtle adjustment of the character representation layer to introduce the lexicon information.Experimental studies on four benchmark Chinese NER datasets show that our method achieves an inference speed up to 6.15 times faster than those of state-ofthe-art methods, along with a better performance.The experimental results also show that the proposed method can be easily incorporated with pre-trained models like BERT. 1 * Equal contribution.
Ruotian Ma, Minlong Peng, Qi Zhang 0001, Zhongyu Wei, Xuanjing Huang 0001
ACL2
2020 Weighed Domain-Invariant Representation Learning for Cross-domain Sentiment Analysis
abstract
Cross-domain sentiment analysis is currently a hot topic in both the research and industrial areas.One of the most popular framework for the task is domain-invariant representation learning (DIRL), which aims to learn a distribution-invariant feature representation across domains.However, in this work, we find out that applying DIRL may degrade the cross-domain performance when the label distribution P(Y) changes across domains.To address this problem, we propose a modification to DIRL, obtaining a novel weighted domain-invariant representation learning (WDIRL) framework.We show that it is easy to transfer existing models from the DIRL framework to the WDIRL framework.Empirical studies on extensive cross-domain sentiment analysis tasks verified our statements and showed the effectiveness of our proposed solution.
Minlong Peng, Qi Zhang 0001
COLING1
2020 Learning to Generate Representations for Novel Words: Mimic the OOV Situation in Training
Minlong Peng, Qi Zhang 0001, Qin Liu 0010, Xuanjing Huang 0001
NLPCC (1)2
2019 Cooperative Multimodal Approach to Depression Detection in Twitter
abstract
The advent of social media has presented a promising new opportunity for the early detection of depression. To do so effectively, there are two challenges to overcome. The first is that textual and visual information must be jointly considered to make accurate inferences about depression. The second challenge is that due to the variety of content types posted by users, it is difficult to extract many of the relevant indicator texts and images. In this work, we propose the use of a novel cooperative multi-agent model to address these challenges. From the historical posts of users, the proposed method can automatically select related indicator texts and images. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods by a large margin (over 30% error reduction). In several experiments and examples, we also verify that the selected posts can successfully indicate user depression, and our model can obtained a robust performance in realistic scenarios.
Tao Gui, Qi Zhang 0001, Minlong Peng, Keyu Ding
AAAI4
2019 Long Short-Term Memory with Dynamic Skip Connections
abstract
In recent years, long short-term memory (LSTM) has been successfully used to model sequential data of variable length. However, LSTM can still experience difficulty in capturing long-term dependencies. In this work, we tried to alleviate this problem by introducing a dynamic skip connection, which can learn to directly connect two dependent words. Since there is no dependency information in the training data, we propose a novel reinforcement learning-based method to model the dependency relationship and connect dependent words. The proposed model computes the recurrent transition functions based on the skip connections, which provides a dynamic skipping advantage over RNNs that always tackle entire sentences sequentially. Our experimental results on three natural language processing tasks demonstrate that the proposed method can achieve better performance than existing methods. In the number prediction experiment, the proposed model outperformed LSTM with respect to accuracy by nearly 20%.
Tao Gui, Qi Zhang 0001, Lujun Zhao, Yaosong Lin, Minlong Peng, Jingjing Gong, Xuanjing Huang 0001
AAAI5
2019 Trainable Undersampling for Class-Imbalance Learning
abstract
Undersampling has been widely used in the class-imbalance learning area. The main deficiency of most existing undersampling methods is that their data sampling strategies are heuristic-based and independent of the used classifier and evaluation metric. Thus, they may discard informative instances for the classifier during the data sampling. In this work, we propose a meta-learning method built on the undersampling to address this issue. The key idea of this method is to parametrize the data sampler and train it to optimize the classification performance over the evaluation metric. We solve the non-differentiable optimization problem for training the data sampler via reinforcement learning. By incorporating evaluation metric optimization into the data sampling process, the proposed method can learn which instance should be discarded for the given classifier and evaluation metric. In addition, as a data level operation, this method can be easily applied to arbitrary evaluation metric and classifier, including non-parametric ones (e.g., C4.5 and KNN). Experimental results on both synthetic and realistic datasets demonstrate the effectiveness of the proposed method.
Minlong Peng, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001, Yu-Gang Jiang 0001, Keyu Ding
AAAI1
2019 Distantly Supervised Named Entity Recognition using Positive-Unlabeled Learning
abstract
In this work, we explore the way to perform named entity recognition (NER) using only unlabeled data and named entity dictionaries.To this end, we formulate the task as a positive-unlabeled (PU) learning problem and accordingly propose a novel PU learning algorithm to perform the task.We prove that the proposed algorithm can unbiasedly and consistently estimate the task loss as if there is fully labeled data.A key feature of the proposed method is that it does not require the dictionaries to label every entity within a sentence, and it even does not require the dictionaries to label all of the words constituting an entity.This greatly reduces the requirement on the quality of the dictionaries and makes our method generalize well with quite simple dictionaries.Empirical studies on four public NER datasets demonstrate the effectiveness of our proposed method.We have published the source code at https:// github.com/v-mipeng/LexiconNER.
Minlong Peng, Qi Zhang 0001, Jinlan Fu, Xuanjing Huang 0001
ACL (1)1
2019 A Lexicon-Based Graph Neural Network for Chinese NER
abstract
Tao Gui, Yicheng Zou, Qi Zhang, Minlong Peng, Jinlan Fu, Zhongyu Wei, Xuanjing Huang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tao Gui, Yicheng Zou, Qi Zhang 0001, Minlong Peng, Jinlan Fu, Zhongyu Wei, Xuanjing Huang 0001
EMNLP/IJCNLP (1)4
2019 Learning Task-Specific Representation for Novel Words in Sequence Labeling
abstract
Word representation is a key component in neural-network-based sequence labeling systems. However, representations of unseen or rare words trained on the end task are usually poor for appreciable performance. This is commonly referred to as the out-of-vocabulary (OOV) problem. In this work, we address the OOV problem in sequence labeling using only training data of the task. To this end, we propose a novel method to predict representations for OOV words from their surface-forms (e.g., character sequence) and contexts. The method is specifically designed to avoid the error propagation problem suffered by existing approaches in the same paradigm. To evaluate its effectiveness, we performed extensive empirical studies on four part-of-speech tagging (POS) tasks and four named entity recognition (NER) tasks. Experimental results show that the proposed method can achieve better or competitive performance on the OOV problem compared with existing state-of-the-art methods.
Minlong Peng, Qi Zhang 0001, Tao Gui, Jinlan Fu, Xuanjing Huang 0001
IJCAI1
2019 Model the Long-Term Post History for Hashtag Recommendation
Minlong Peng, Qiyuan Bian, Qi Zhang 0001, Tao Gui, Jinlan Fu, Lanjun Zeng, Xuanjing Huang 0001
NLPCC (1)1
2019 Mention Recommendation in Twitter with Cooperative Multi-Agent Reinforcement Learning
abstract
In Twitter-like social networking services, the "@'' symbol can be used with the tweet to mention users whom the user wants to alert regarding the message. An automatic suggestion to the user of a small list of candidate names can improve communication efficiency. Previous work usually used several most recent tweets or randomly select historical tweets to make an inference about this preferred list of names. However, because there are too many historical tweets by users and a wide variety of content types, the use of several tweets cannot guarantee the desired results. In this work, we propose the use of a novel cooperative multi-agent approach to mention recommendation, which incorporates dozens of more historical tweets than earlier approaches. The proposed method can effectively select a small set of historical tweets and cooperatively extract relevant indicator tweets from both the user and mentioned users. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods.
Tao Gui, Qi Zhang 0001, Minlong Peng, Yunhua Zhou, Xuanjing Huang 0001
SIGIR5
2019 Adaptive Multi-Attention Network Incorporating Answer Information for Duplicate Question Detection
abstract
Community-based question answering (CQA), which provides a platform for people with diverse backgrounds to share information and knowledge, has become increasingly popular. With the accumulation of site data, methods to detect duplicate questions in CQA sites have attracted considerable attention. Existing methods typically use only questions to complete the task. However, the paired answers may also provide valuable information. In this paper, we propose an answer information- enhanced adaptive multi-attention network (AMAN) to perform this task. AMAN takes full advantage of the semantic information in the paired answers while alleviating the noise problem caused by adding the answers. To evaluate the proposed method, we use a CQADupStack set and the Quora question-pair dataset expanded with paired answers. Experimental results demonstrate that the proposed model can achieve state-of-the-art performance on the above two data sets.
Di Liang, Fubao Zhang, Qi Zhang 0001, Jinlan Fu, Minlong Peng, Tao Gui, Xuanjing Huang 0001
SIGIR6
2019 Implicit discourse relation detection using concatenated word embeddings and a gated relevance network
Jinlan Fu, Qi Zhang 0001, Jifan Chen, Minlong Peng, Tao Gui, Xipeng Qiu, Xuanjing Huang 0001
Sci. China Inf. Sci.4
2018 Cross-Domain Sentiment Classification with Target Domain Specific Information
abstract
The task of adopting a model with good performance to a target domain that is different from the source domain used for training has received considerable attention in sentiment analysis.Most existing approaches mainly focus on learning representations that are domain-invariant in both the source and target domains.Few of them pay attention to domain specific information, which should also be informative.In this work, we propose a method to simultaneously extract domain specific and invariant representations and train a classifier on each of the representation, respectively.And we introduce a few target domain labeled data for learning domain-specific information.To effectively utilize the target domain labeled data, we train the domain-invariant representation based classifier with both the source and target domain labeled data and train the domain-specific representation based classifier with only the target domain labeled data.These two classifiers then boost each other in a co-training style.Extensive sentiment analysis experiments demonstrated that the proposed method could achieve better performance than state-of-the-art methods.
Minlong Peng, Qi Zhang 0001, Yu-Gang Jiang 0001, Xuanjing Huang 0001
ACL (1)1
2018 Transferring from Formal Newswire Domain with Hypernet for Twitter POS Tagging
abstract
Part-of-Speech (POS) tagging for Twitter has received considerable attention in recent years.Because most POS tagging methods are based on supervised models, they usually require a large amount of labeled data for training.However, the existing labeled datasets for Twitter are much smaller than those for newswire text.Hence, to help POS tagging for Twitter, most domain adaptation methods try to leverage newswire datasets by learning the shared features between the two domains.However, from a linguistic perspective, Twitter users not only tend to mimic the formal expressions of traditional media, like news, but they also appear to be developing linguistically informal styles.Therefore, POS tagging for the formal Twitter context can be learned together with the newswire dataset, while POS tagging for the informal Twitter context should be learned separately.To achieve this task, in this work, we propose a hypernetworkbased method to generate different parameters to separately model contexts with different expression styles.Experimental results on three different datasets show that our approach achieves better performance than state-of-theart methods in most cases.
Tao Gui, Qi Zhang 0001, Jingjing Gong, Minlong Peng, Di Liang, Keyu Ding, Xuanjing Huang 0001
EMNLP4
2017 Part-of-Speech Tagging for Twitter with Adversarial Neural Networks
abstract
In this work, we study the problem of partof-speech tagging for Tweets.In contrast to newswire articles, Tweets are usually informal and contain numerous out-ofvocabulary words.Moreover, there is a lack of large scale labeled datasets for this domain.To tackle these challenges, we propose a novel neural network to make use of out-of-domain labeled data, unlabeled in-domain data, and labeled indomain data.Inspired by adversarial neural networks, the proposed method tries to learn common features through adversarial discriminator.In addition, we hypothesize that domain-specific features of target domain should be preserved in some degree.Hence, the proposed method adopts a sequence-to-sequence autoencoder to perform this task.Experimental results on three different datasets show that our method achieves better performance than state-of-the-art methods.
Tao Gui, Qi Zhang 0001, Haoran Huang, Minlong Peng, Xuanjing Huang 0001
EMNLP4