Wenpeng Hu

dblp:191/6009 · DBLP profile ↗
← Back
36ranked-venue papers
8as first author
22since 2021 · last 2026
0009-0007-8554-8778ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 8 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CGMIS: Concept-Graph Based Multi-Hop Instructions Synthesis for Enhancing Long-Context Reasoning
abstract
High-quality multi-hop instruction data is critical for enhancing the reasoning capabilities of large language models (LLMs) in complex long-context scenarios, e.g., long-form reasoning. Nevertheless, there is currently a notable scarcity of such datasets within the community, and existing data synthesis approaches typically fail to provide explicit modeling of intermediate reasoning steps, resulting in unverifiable and potentially erroneous samples. To mitigate above issue, we design the Concept-Graph based Multi-hop Instructions Synthesis (CGMIS) framework, which constructs long-form reasoning paths via concept graph traversal and automatically generates verifiable multi-hop data. The CGMIS framework not only guarantees the accuracy and verifiability of the synthesized data but also enables the construction of high-quality multi-hop instruction datasets from arbitrary corpora. Experiments show that fine-tuning with CGMIS-generated data achieves state-of-the-art performance across 13 long-context reasoning tasks on various models, using only 10% of the data volume required by existing methods.
Zechen Sun, Zecheng Tang, Juntao Li 0005, Wenpeng Hu, Wenliang Chen, Zhunchen Luo, Qiaoming Zhu
AAAI4
2026 IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking
abstract
Zechen Sun, Yuyang Sun, Zecheng Tang, Juntao Li, Wenpeng Hu, Wenliang Chen, Zhunchen Luo, Guotong Geng, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zechen Sun, Zecheng Tang, Juntao Li 0005, Wenpeng Hu, Wenliang Chen, Zhunchen Luo, Guotong Geng, Min Zhang 0005
ACL (1)5
2026 CAGE-MoE: Cross-Projection Alignment-Guided Compression for MoE Models
Zhunchen Luo, Long Sheng, Guotong Geng, Wenpeng Hu
ICIC (9)6
2026 LLaMA-MoT: A cost-effective framework for visual-linguistic instruction tuning based on multi-head adapters and chain-of-thought
Turdi Tohti, Wenpeng Hu, Tianwei Yan 0001, Shaohuang Wang, Askar Hamdulla
Expert Syst. Appl.3
2026 WizardEvent: Empowering Event Reasoning by Hybrid Event-Aware Data Synthesizing
abstract
Event reasoning is to reason with events and certain inter-event relations. These cutting-edge techniques possess crucial and fundamental capabilities that underlie various applications. Large language models (LLMs) have made advances in event reasoning owing to their wealth of training. However, the LLMs commonly used today still do not consistently demonstrate proficiency in managing event reasoning as humans. This discrepancy arises from not explicitly modeling events and their relations and insufficient knowledge of event relations. In addition, the different reasoning paradigms of the LLMs are trained in an imbalanced way. In this paper, we propose WIZARDEVENT, to synthesize data from the unlabeled corpus with the proposed hybrid event-aware instruction tuning. Specifically, we first represent the events and their relation in a novel structure and then extract the knowledge from the raw text. Second, we introduce hybrid event reasoning paradigms with four reasoning formats. Lastly, we wrap our constructed WIZARDEVENT with the paradigms to create the instruction tuning dataset. We fine-tune the model with this enriched dataset, significantly improving the event reasoning. The performance of WIZARDEVENT is rigorously evaluated through extensive experiments. The results demonstrate that WIZARDEVENT substantially outperforms baselines, indicating the effectiveness of our approach.
Zhengwei Tao, Xiancai Chen, Zhi Jin 0001, Xiaoying Bai, Haiyan Zhao 0001, Wenpeng Hu, Chongyang Tao, Shuai Ma 0001
IEEE Trans. Knowl. Data Eng.6
2025 Unveiling the Potential of BERT-family: A New Recipe for Building Scalable, General and Competitive Large Language Models
abstract
BERT-family have been increasingly explored for adaptation to scenarios beyond language understanding tasks, with more recent efforts focused on enabling them to become good instruction followers.These explorations have endowed BERT-family with new roles and human expectations, showcasing their potential on par with current state-of-the-art (SOTA) large language models (LLMs).However, several certain shortcomings in previous BERT-family, such as the relatively sub-optimal training corpora, learning procedure, and model architecture, all impede the further advancement of these models for serving as general and competitive LLMs.Therefore, we aim to address these deficiencies in this paper.Our study not only introduces a more suitable pre-training task that helps BERT-family excel in wider applications to realize generality but also explores the integration of cutting-edge technologies into our model to further enhance their capabilities.Our final models, termed Bidirectional General Language Models (BiGLM), exhibit performance levels comparable to current SOTA LLMs across a spectrum of tasks.Moreover, we conduct detailed analyses to study the effects of scaling and training corpora for BiGLM.To the best of our knowledge, our work represents the early attempt to offer a recipe for building novel types of scalable, general, and competitive LLMs that diverge from current autoregressive modeling methodology.Our codes and models are available on Github 1 .
Yisheng Xiao, Juntao Li 0005, Wenpeng Hu, Zhunchen Luo, Min Zhang 0005
ACL (1)3
2025 AdaDARE-gamma: Balancing Stability and Plasticity in Multi-modal LLMs through Efficient Adaptation
abstract
Adapting Multi-modal Large Language Models (MLLMs) to target tasks often suffers from catastrophic forgetting, where acquiring new task-specific knowledge compromises performance on pre-trained tasks. In this paper, we introduce AdaDARE-γ, an efficient approach that alleviates catastrophic forgetting by controllably injecting new task-specific knowledge through adaptive parameter selection from fine-tuned models without requiring retraining procedures. This approach consists two key innovations: (1) an adaptive parameter selection mechanism that identifies and retains the most task-relevant parameters from fine-tuned models, and (2) a controlled task-specific information injection strategy that precisely balances the preservation of pre-trained knowledge with the acquisition of new capabilities. Theoretical analysis proves the optimality of our parameter selection strategy and establishes bounds for the task-specific information injection factor. Extensive experiments on InstructBLIP and LLaVA-1.5 across image captioning and visual question answering tasks demonstrate that AdaDARE-γ establishes new state-of-the-art results in balancing model performance. Specifically, it maintains 98.2% of pre-training effectiveness on original tasks while achieving 98.7% of standard fine-tuning performance on target tasks.
Jintao Yang, Zhunchen Luo, Yunbo Cao, Wenpeng Hu
CVPR7
2025 Leveraging Statistical Machine Learning to Boost Large Language Models
abstract
This paper addresses the challenge of insufficient sampling diversity in large language models (LLMs) under the conventional autoregressive decoding framework. Our research reveals that, in some tasks where LLMs underperform, they actually possess the capability to provide correct answers. However, this potential is limited by inadequate sampling diversity in the decoding process, which prevents the models from effectively and fully leveraging their reasoning capabilities. To address this issue, we propose the Adaptive Trigram Model-Assisted Decoding Strategy (ATM-ADS), a method designed to enhance the reasoning abilities of large models without requiring fine-tuning. By dynamically providing diverse candidate vocabularies, combined with a decoding dynamic decision-making mechanism and an adaptive smoothing strategy, our approach substantially enhances sampling diversity during the decoding process. Experimental results demonstrate that our ATM-ADS significantly improves model performance across multiple task scenarios, highlighting its broad potential for cross-task and cross-linguistic applications.
Zhiyu Ding, Wenpeng Hu, Jianyong Duan, Jiaxin Bai, Zhunchen Luo
IJCNN2
2025 Knowledge-Optimized Multi-Agent Dynamics Framework for Zero-Shot Relation Triplet Extraction
abstract
Relation triplet extraction aims to identify entity pairs and their relations from unstructured text. Traditional zero-shot learning methods are hindered by predefined relation types and the lack of large-scale annotated data, limiting generalization to unseen relations. To address these challenges, we propose the KOMADF (Knowledge-Optimized Multi-Agent Dynamics Framework). KOMADF constructs task-specific knowledge graphs and incorporates a distillation module to refine relation labels, filtering out irrelevant candidates while preserving semantically similar or multi-type labels. These refined labels generate problem representations that guide a multi-agent mechanism to dynamically create context-aware prompt templates. By aligning task objectives with pre-trained language models, these templates enable generalization to unseen relations.Experiments on two public datasets, FewRel and Wiki-ZSL, show that the proposed method substantially improves precision, recall, and F1 score. These findings confirm KOMADF’s adaptability and robustness, establishing it as an effective solution to the challenges of zero-shot relation triplet extraction. Furthermore, the proposed framework lays a solid theoretical foundation for the automated construction of dynamic knowledge graphs.
Jianyong Duan, Wenpeng Hu, Zhunchen Luo
IJCNN4
2025 Are Task-Specific Datasets the Be-All and End-All for LLM Performance?
abstract
This study explores the efficacy of fine-tuning large language models (LLMs) using non-task-specific datasets, challenging the traditional reliance on task-specific datasets. Based on the SUP-NATINST and Stanford Alpaca dataset, this research explores some schemes to enhance the application capabilities of LLMs in specific domains through strategic data selection. The main contributions include an innovative fine-tuning data selection strategy that emphasizes the importance of dataset selection by comparing non-task-specific datasets with task-specific datasets, demonstrating the potential advantages of unconventional datasets in improving model task performance. The study also reveals the model’s performance sensitivity to data ratios, challenges the concept of optimal data ratios, and explores the impact of data volume changes on model accuracy. Furthermore, it deepens the understanding of how different task types affect model performance fine-tuning. By analyzing the performance across different tasks and dimensions and dissecting the impact of data ratio changes, the research indicates that non-specific task data allocation can achieve better performance than specific task datasets.
Jintao Yang, Yushan Tan, Wenpeng Hu, Junyao Zhou, Zonghao Yang, Zhunchen Luo
IJCNN3
2025 TRRD: Enabling Accurate and Reliable Retrieval of Time-sensitive Information
abstract
In the era of information explosion, locating and retrieving data or events within specific temporal ranges has become an inevitable challenge. Conventional retrieval models, exhibit limited sensitivity to temporal information, while large language models primarily focus on temporal reasoning, often in the context of question answering or fact verification tasks. Moreover, these models demonstrate vulnerabilities in temporal reasoning tasks, making them unsuitable for retrieval tasks based on temporal ranges. To address this issue, this study constructs a dataset specifically designed for temporal range retrieval and ranking. The dataset includes a variety of temporal range samples to fine-tune models, enhancing their capabilities in temporal retrieval and ranking. To further improve the generalization ability of the dataset, we propose a data augmentation method. We fine-tuned the current mainstream retrieval models using our dataset, and the results show significant improvements in their temporal retrieval performance, even in zero-shot scenarios. Experiments conducted on three test datasets demonstrate that our approach substantially enhances the models’ ability to handle temporal range retrieval tasks. Therefore, we believe that our dataset can provide a valuable reference for future research work.
Junyao Zhou, Yushan Tan, ZiBo Yi, Ruiqing Du, Wenpeng Hu
IJCNN6
2025 DPIMerge: An Efficient Dynamic Parameter Interpolation Framework for Alleviating Pure Text Forgetting in Multimodal Large Models
Haoguang Wen, Wenpeng Hu, Zhunchen Luo, Lingqiang Chen, Shuyi Wu
NLPCC (2)3
2025 D-PathVer: Dynamic Reasoning Pathways for Complex Claim Verification
Lingxiao Zheng, Zhunchen Luo, Wenpeng Hu, Yunbo Cao
NLPCC (3)3
2024 REM: A Ranking-Based Automatic Evaluation Method for LLMs
Jintao Yang, Yushan Tan, Wenpeng Hu, Zonghao Yang, Zhunchen Luo
ICANN (5)3
2024 Exploring Instruction Feature Adaptation for Event Argument Extraction in Large Language Models
abstract
Event argument extraction (EAE) is an important and critical task in natural language processing, whose goal is to transform unstructured text into structured information. However, previous methods are costly for annotation, which limits their application. With the emergence of large language models (LLMs), many approaches use prompt learning or instruction learning to solve the EAE task. However, due to the complexity of the EAE, the defined instructions are often abstract and difficult for the model to accurately understand the task objective, while the cost of gradient propagation based methods is too high. Therefore, we propose an Instruction Feature Adaptation event argument extraction (IFA) that uses very few shots to tune the task instruction features, while not requiring a highcost gradient optimisation method to enhance the model’s performance on the event argument extraction task. Experiments demonstrate that the method is able to optimise the instruction features, enabling the model to perform the event argument task more consistently based on the instructions, ultimately improving performance for low-resource. Code is available at https://github.com/yangzonghao1024/IFA.
Zonghao Yang, Jintao Yang, Yushan Tan, Changhai Tian, Junyao Zhou, Wenpeng Hu, Zhunchen Luo
IJCNN7
2024 Instruct Large Language Models to Generate Scientific Literature Survey Step by Step
Yuxuan Lai, Wenpeng Hu
NLPCC (5)4
2024 LLM Assists Hypothesis Generation and Testing for Deliberative Questions
Fuchun Wang, Xian Zhou 0003, Wenpeng Hu, Zhunchen Luo, Xiaoying Bai
NLPCC (2)3
2022 Adaptive Orthogonal Projection for Batch and Online Continual Learning
abstract
Catastrophic forgetting is a key obstacle to continual learning. One of the state-of-the-art approaches is orthogonal projection. The idea of this approach is to learn each task by updating the network parameters or weights only in the direction orthogonal to the subspace spanned by all previous task inputs. This ensures no interference with tasks that have been learned. The system OWM that uses the idea performs very well against other state-of-the-art systems. In this paper, we first discuss an issue that we discovered in the mathematical derivation of this approach and then propose a novel method, called AOP (Adaptive Orthogonal Projection), to resolve it, which results in significant accuracy gains in empirical evaluations in both the batch and online continual learning settings without saving any previous training data as in replay-based methods.
Yiduo Guo, Wenpeng Hu, Dongyan Zhao 0001, Bing Liu 0001
AAAI2
2022 CMG: A Class-Mixed Generation Approach to Out-of-Distribution Detection
Mengyu Wang 0002, Yijia Shao, Haowei Lin, Wenpeng Hu, Bing Liu 0001
ECML/PKDD (4)4
2021 Predictive Adversarial Learning from Positive and Unlabeled Data
abstract
This paper studies learning from positive and unlabeled examples, known as PU learning. It proposes a novel PU learning method called Predictive Adversarial Networks (PAN) based on GAN (Generative Adversarial Networks). GAN learns a generator to generate data (e.g., images) to fool a discriminator which tries to determine whether the generated data belong to a (positive) training class. PU learning can be casted as trying to identify (not generate) likely positive instances from the unlabeled set to fool a discriminator that determines whether the identified likely positive instances from the unlabeled set are indeed positive. However, directly applying GAN is problematic because GAN focuses on only the positive data. The resulting PU learning method will have high precision but low recall. We propose a new objective function based on KL-divergence. Evaluation using both image and text data shows that PAN outperforms state-of-the-art PU learning methods and also a direct adaptation of GAN for PU learning.
Wenpeng Hu, Ran Le, Bing Liu 0001, Jinwen Ma, Dongyan Zhao 0001, Rui Yan 0001
AAAI1
2021 Continual Learning by Using Information of Each Class Holistically
abstract
Continual learning (CL) incrementally learns a sequence of tasks while solving the catastrophic forgetting (CF) problem. Existing methods mainly try to deal with CF directly. In this paper, we propose to avoid CF by considering the features of each class holistically rather than only the discriminative information for classifying the classes seen so far. This latter approach is prone to CF because the discriminative information for old classes may not be sufficiently discriminative for the new class to be learned. Consequently, in learning each new task, the network parameters for previous tasks have to be revised, which causes CF. With the holistic consideration, after adding new tasks, the system can still do well for previous tasks. The proposed technique is called Per-class Continual Learning (PCL). PCL has two key novelties. (1) It proposes a one-class learning based technique for CL, which considers features of each class holistically and represents a new approach to solving the CL problem. (2) It proposes a method to extract discriminative information after training to further improve the accuracy. Empirical evaluation shows that PCL markedly outperforms the state-of-the-art baselines for one or more classes per task. More tasks also result in more gains.
Wenpeng Hu, Mengyu Wang 0002, Jinwen Ma, Bing Liu 0001
AAAI1
2021 BNS: Building Network Structures Dynamically for Continual Learning
abstract
Continual learning (CL) of a sequence of tasks is often accompanied with the catastrophic forgetting(CF) problem. Existing research has achieved remarkable results in overcoming CF, especially for task continual learning. However, limited work has been done to achieve another important goal of CL,knowledge transfer.In this paper, we propose a technique (called BNS) to do both. The novelty of BNS is that it dynamically builds a network to learn each new task to overcome CF and to transfer knowledge across tasks at the same time. Experimental results show that when the tasks are different (with little shared knowledge), BNS can already outperform the state-of-the-art baselines. When the tasks are similar and have shared knowledge, BNS outperforms the baselines substantially by a large margin due to its knowledge transfer capability.
Wenpeng Hu, Dongyan Zhao 0001, Bing Liu 0001
NeurIPS2
2020 Feature Projection for Improved Text Classification
abstract
In classification, there are usually some good features that are indicative of class labels.For example, in sentiment classification, words like good and nice are indicative of the positive sentiment and words like bad and terrible are indicative of the negative sentiment.However, there are also many common features (e.g., words) that are not indicative of any specific class (e.g., voice and screen, which are common to both sentiment classes and are not discriminative for classification).Although deep learning has made significant progresses in generating discriminative features through its powerful representation learning, we believe there is still room for improvement.In this paper, we propose a novel angle to further improve this representation learning, i.e., feature projection.This method projects existing features into the orthogonal space of the common features.The resulting projection is thus perpendicular to the common features and more discriminative for classification.We apply this new method to improve CNN, RNN, Transformer, and Bert based text classification and obtain markedly better results.
Wenpeng Hu, Bing Liu 0001
ACL2
2020 Translation vs. Dialogue: A Comparative Analysis of Sequence-to-Sequence Modeling
abstract
Understanding neural models is a major topic of interest in the deep learning community. In this paper, we propose to interpret a general neural model comparatively. Specifically, we study the sequence-to-sequence (Seq2Seq) model in the contexts of two mainstream NLP tasks–machine translation and dialogue response generation–as they both use the seq2seq model. We investigate how the two tasks are different and how their task difference results in major differences in the behaviors of the resulting translation and dialogue generation systems. This study allows us to make several interesting observations and gain valuable insights, which can be used to help develop better translation and dialogue generation models. To our knowledge, no such comparative study has been done so far.
Wenpeng Hu, Ran Le, Bing Liu 0001, Jinwen Ma, Dongyan Zhao 0001, Rui Yan 0001
COLING1
2020 Transformation of Dense and Sparse Text Representations
abstract
Sparsity is regarded as a desirable property of representations, especially in terms of explanation.However, its usage has been limited due to the gap with dense representations.Most research progresses in NLP in recent years are based on dense representations.Thus the desirable property of sparsity cannot be leveraged.Inspired by Fourier Transformation, in this paper, we propose a novel Semantic Transformation method to bridge the dense and sparse spaces, which can facilitate the NLP research to shift from dense spaces to sparse spaces or to jointly use both spaces.Experiments using classification tasks and natural language inference task show that the proposed Semantic Transformation is effective.
Wenpeng Hu, Mengyu Wang 0002, Bing Liu 0001, Jinwen Ma, Dongyan Zhao 0001
COLING1
2020 HRN: A Holistic Approach to One Class Learning
abstract
Existing neural network based one-class learning methods mainly use various forms of auto-encoders or GAN style adversarial training to learn a latent representation of the given one class of data. This paper proposes an entirely different approach based on a novel regularization, called holistic regularization (or H-regularization), which enables the system to consider the data holistically, not to produce a model that biases towards some features. Combined with a proposed 2-norm instance-level data normalization, we obtain an effective one-class learning method, called HRN. To our knowledge, the proposed regularization and the normalization method have not been reported before. Experimental evaluation using both benchmark image classification and traditional anomaly detection datasets show that HRN markedly outperforms the state-of-the-art existing deep/non-deep learning models.
Wenpeng Hu, Mengyu Wang 0002, Jinwen Ma, Bing Liu 0001
NeurIPS1
2019 One Time of Interaction May Not Be Enough: Go Deep with an Interaction-over-Interaction Network for Response Selection in Dialogues
abstract
Currently, researchers have paid great attention to retrieval-based dialogues in opendomain.In particular, people study the problem by investigating context-response matching for multi-turn response selection based on publicly recognized benchmark data sets.State-of-the-art methods require a response to interact with each utterance in a context from the beginning, but the interaction is performed in a shallow way.In this work, we let utterance-response interaction go deep by proposing an interaction-over-interaction network (IoI).The model performs matching by stacking multiple interaction blocks in which residual information from one time of interaction initiates the interaction process again.Thus, matching information within an utterance-response pair is extracted from the interaction of the pair in an iterative fashion, and the information flows along the chain of the blocks via representations.Evaluation results on three benchmark data sets indicate that IoI can significantly outperform state-of-theart methods in terms of various matching metrics.Through further analysis, we also unveil how the depth of interaction affects the performance of IoI.
Chongyang Tao, Wei Wu 0014, Can Xu 0002, Wenpeng Hu, Dongyan Zhao 0001, Rui Yan 0001
ACL (1)4
2019 Query-bag Matching with Mutual Coverage for Information-seeking Conversations in E-commerce
abstract
Information-seeking conversation system aims at satisfying the information needs of users through conversations. Text matching between a user query and a pre-collected question is an important part of the information-seeking conversation in E-commerce. In the practical scenario, a sort of questions always correspond to a same answer. Naturally, these questions can form a bag. Learning the matching between user query and bag directly may improve the conversation performance, denoted as query-bag matching. Inspired by such opinion, we propose a query-bag matching model which mainly utilizes the mutual coverage between query and bag and measures the degree of the content in the query mentioned by the bag, and vice verse. In addition, the learned bag representation in word level helps find the main points of a bag in a fine grade and promotes the query-bag matching performance. Experiments on two datasets show the effectiveness of our model.
Zhenxin Fu, Wenpeng Hu, Dongyan Zhao 0001, Haiqing Chen, Rui Yan 0001
CIKM3
2019 Towards Effective and Interpretable Person-Job Fitting
abstract
The diversity of job requirements and the complexity of job seekers' abilities put forward higher requirements for the accuracy and interpretability of Person-Job Fit system. Interpretable Person-Job Fit system can show reasons for giving recommendations or not recommending specific jobs to some people, and vice versa. Such reasons help us understand according to what the final decision is made by the system and guarantee a high recommending accuracy. Existing studies on Person-Job Fit have focused on 1) one perspective, without considering the variances of role and psychological motivation between interviewer and job seeker; 2) modeling the matching degree between resume and job requirements directly through a deep neural network without interaction matching modules, which leads to shortage on interpretation. To this end, we propose an Interpretable Person-Job Fit (IPJF) model, which 1) models the Person-Job Fit problem from the perspectives/intentions of employer and job seeker in a multi-tasks optimization fashion to interpretively formulate the Person-Job Fit process; 2) leverages deep interactive representation learning to automatically learn the interdependence between a resume and job requirements without relying on a clear list of job seeker's abilities, and deploys the optimizing problem as a learning to rank problem. Experiments on large real dataset show that the proposed IPJF model outperforms state-of-the-art baselines and also gives promising interpretable recommending reasons.
Ran Le, Wenpeng Hu, Yang Song 0021, Tao Zhang 0070, Dongyan Zhao 0001, Rui Yan 0001
CIKM2
2019 Modeling Personalization in Continuous Space for Response Generation via Augmented Wasserstein Autoencoders
abstract
Zhangming Chan, Juntao Li, Xiaopeng Yang, Xiuying Chen, Wenpeng Hu, Dongyan Zhao, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zhangming Chan, Juntao Li 0005, Xiaopeng Yang 0002, Xiuying Chen, Wenpeng Hu, Dongyan Zhao 0001, Rui Yan 0001
EMNLP/IJCNLP (1)5
2019 Who Is Speaking to Whom? Learning to Identify Utterance Addressee in Multi-Party Conversations
abstract
Ran Le, Wenpeng Hu, Mingyue Shang, Zhenjun You, Lidong Bing, Dongyan Zhao, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ran Le, Wenpeng Hu, Mingyue Shang, Zhenjun You, Lidong Bing, Dongyan Zhao 0001, Rui Yan 0001
EMNLP/IJCNLP (1)2
2019 Overcoming Catastrophic Forgetting for Continual Learning via Model Adaptation
Wenpeng Hu, Bing Liu 0001, Chongyang Tao, Zhengwei Tao, Jinwen Ma, Dongyan Zhao 0001, Rui Yan 0001
ICLR (Poster)1
2019 GSN: A Graph-Structured Network for Multi-Party Dialogues
abstract
Existing neural models for dialogue response generation assume that utterances are sequentially organized. However, many real-world dialogues involve multiple interlocutors (i.e., multi-party dialogues), where the assumption does not hold as utterances from different interlocutors can occur ``in parallel.'' This paper generalizes existing sequence-based models to a Graph-Structured neural Network (GSN) for dialogue modeling. The core of GSN is a graph-based encoder that can model the information flow along the graph-structured dialogues (two-party sequential dialogues are a special case). Experimental results show that GSN significantly outperforms existing sequence-based models.
Wenpeng Hu, Zhangming Chan, Bing Liu 0001, Dongyan Zhao 0001, Jinwen Ma, Rui Yan 0001
IJCAI1
2019 FlexNER: A Flexible LSTM-CNN Stack Framework for Named Entity Recognition
Hongyin Zhu, Wenpeng Hu, Yi Zeng 0001
NLPCC (2)2
2019 Multi-Representation Fusion Network for Multi-Turn Response Selection in Retrieval-Based Chatbots
abstract
We consider context-response matching with multiple types of representations for multi-turn response selection in retrieval-based chatbots. The representations encode semantics of contexts and responses on words, n-grams, and sub-sequences of utterances, and capture both short-term and long-term dependencies among words. With such a number of representations in hand, we study how to fuse them in a deep neural architecture for matching and how each of them contributes to matching. To this end, we propose a multi-representation fusion network where the representations can be fused into matching at an early stage, at an intermediate stage, or at the last stage. We empirically compare different representations and fusing strategies on two benchmark data sets. Evaluation results indicate that late fusion is always better than early fusion, and by fusing the representations at the last stage, our model significantly outperforms the existing methods, and achieves new state-of-the-art performance on both data sets. Through a thorough ablation study, we demonstrate the effect of each representation to matching, which sheds light on how to select them in practical systems.
Chongyang Tao, Wei Wu 0014, Can Xu 0002, Wenpeng Hu, Dongyan Zhao 0001, Rui Yan 0001
WSDM4
2016 Different Contexts Lead to Different Word Embeddings
abstract
Recent work for learning word representations has applied successfully to many NLP applications, such as sentiment analysis and question answering. However, most of these models assume a single vector per word type without considering polysemy and homonymy. In this paper, we present an extension to the CBOW model which not only improves the quality of embeddings but also makes embeddings suitable for polysemy. It differs from most of the related work in that it learns one semantic center embedding and one context bias instead of training multiple embeddings per word type. Different context leads to different bias which is defined as the weighted average embeddings of local context. Experimental results on similarity task and analogy task show that the word representations learned by the proposed method outperform the competitive baselines.
Wenpeng Hu
COLING1