EDBT 2026 Demo / reviewers in the wild / expert
Pei Ke
dblp:10/2179
· DBLP profile ↗
31ranked-venue papers
7as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 7 first-author · 20 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following EvaluationabstractBosi Wen, Yilin Niu, Cunxiang Wang, Pei Ke, Xiaoying Ling, Ying Zhang, Aohan Zeng, Hongning Wang, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bosi Wen, Yilin Niu, Cunxiang Wang, Pei Ke, Xiaoying Ling, Aohan Zeng, Hongning Wang, Minlie Huang |
ACL (1) | 4 |
| 2026 | IF-RewardBench: Benchmarking Judge Models for Instruction-Following EvaluationabstractBosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling, Ying Zhang, Pei Ke, Hongning Wang, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling, Pei Ke, Hongning Wang, Minlie Huang |
ACL (1) | 6 |
| 2026 | From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language ModelsabstractEvent analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks. To address these limitations, we introduce MiGUE-Bench, a systematic benchmark for assessing the performance of LLMs in multi-granularity event analysis. To support large-scale evaluation, we first develop an LLM-driven self-correcting annotation framework called MiGUE-Pipeline, enabling scalable acquisition of high-quality source data of events with automatic labels. Then, we design four core tasks in our benchmark, i.e., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross-document narratives. Extensive experiments on state-of-the-art LLMs and retrieval-augmented generation (RAG) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks. Tao Wen 0011, Shuai Shao 0015, Pei Ke, Xu Han 0007, Jie Zou 0001, Tao Tian, Jinjie Qiu, Ke Qin |
SIGIR | 3 |
| 2026 | Benchmarking and enhancing the ability of large language models on event generalization
Shuai Shao 0015, Tao Wen 0011, Yuezhou Dong, Pei Ke, Ke Qin |
Inf. Process. Manag. | 4 |
| 2025 | CharacterBench: Benchmarking Character Customization of Large Language ModelsabstractCharacter-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs’ character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters’ responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark’s potential to optimize LLMs’ character customization. Jinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi, Pei Ke, Zhuang Chen 0002, Xiyao Xiao, Libiao Peng, Kuntian Tang, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang |
AAAI | 6 |
| 2025 | The superalignment of superhuman intelligence with large language models
Minlie Huang, Yingkang Wang, Shiyao Cui, Pei Ke, Jie Tang 0001 |
Sci. China Inf. Sci. | 4 |
| 2025 | Cross-Graph Knowledge Exchange for Personalized Response Generation in Dialogue SystemsabstractRecent advancements in language models have greatly improved dialogue systems, but they still face challenges in generating personalized responses that are consistent with the user’s persona and dialogue context. Existing approaches typically model dialogue context and persona information together in a unified manner, but they lack fine-grained differentiation between the two, leading to inaccurate user modeling. This misalignment hinders the ability of dialogue systems to produce personalized responses. In this work, we propose cross-graph knowledge exchange (CKE), a novel algorithm designed to enhance personalized response generation in dialogue systems. CKE constructs separate dialogue user graphs for each party in the dialogue, representing both their dialogue context and persona information. These graphs are then utilized to perform cross-graph structured knowledge aggregation, where the aggregation is under supervised by both its own and the other party’s persona and dialogue context, providing richer, more accurate representations. Furthermore, CKE introduces a hybrid prompt template that combines both discrete and continuous elements, improving the language model’s ability to leverage graph-structured information. The experimental results demonstrate that CKE significantly outperforms existing baseline methods in generating more coherent, contextually appropriate, and personalized responses in dialogue systems. Yuezhou Dong, Ke Qin, Pei Ke, Shuang Liang 0002, Guangchun Luo |
IEEE Internet Things J. | 3 |
| 2025 | Safe and effective post-fine-tuning alignment in large language models
Minrui Jiang, Xiurui Xie, Pei Ke, Guisong Liu |
Knowl. Based Syst. | 4 |
| 2024 | Black-Box Prompt Optimization: Aligning Large Language Models without Model TrainingabstractJiale Cheng, Xiao Liu, Kehan Zheng, Pei Ke, Hongning Wang, Yuxiao Dong, Jie Tang, Minlie Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xiao Liu 0036, Kehan Zheng, Pei Ke, Hongning Wang, Yuxiao Dong, Jie Tang 0001, Minlie Huang |
ACL (1) | 4 |
| 2024 | CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model GenerationabstractPei Ke, Bosi Wen, Andrew Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, Jie Tang, Minlie Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Pei Ke, Bosi Wen, Andrew Feng, Xiao Liu 0036, Xuanyu Lei, Shengyuan Wang 0002, Aohan Zeng, Yuxiao Dong, Hongning Wang, Jie Tang 0001, Minlie Huang |
ACL (1) | 1 |
| 2024 | AlignBench: Benchmarking Chinese Alignment of Large Language ModelsabstractXiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang, Andrew Feng, Bosi Wen, Jiale Cheng, Pei Ke, Yifan Xu, Weng Lam Tam, Xiaohan Zhang, Lichao Sun, Xiaotao Gu, Hongning Wang, Jing Zhang, Minlie Huang, Yuxiao Dong, Jie Tang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xiao Liu 0036, Xuanyu Lei, Shengyuan Wang 0002, Yue Huang 0001, Andrew Feng, Bosi Wen, Pei Ke, Yifan Xu 0014, Weng Lam Tam, Lichao Sun 0001, Xiaotao Gu, Hongning Wang, Jing Zhang 0001, Minlie Huang, Yuxiao Dong, Jie Tang 0001 |
ACL (1) | 8 |
| 2024 | Learning Task Decomposition to Assist Humans in Competitive ProgrammingabstractWhen using language models (LMs) to solve complex problems, humans might struggle to understand the LM-generated solutions and repair the flawed ones.To assist humans in repairing them, we propose to automatically decompose complex solutions into multiple simpler pieces that correspond to specific subtasks.We introduce a novel objective for learning task decomposition, termed assistive value (AssistV), which measures the feasibility and speed for humans to repair the decomposed solution.We collect a dataset of human repair experiences on different decomposed solutions.Utilizing the collected data as in-context examples, we then learn to critique, refine, and rank decomposed solutions to improve AssistV.We validate our method under competitive programming problems: under 177 hours of human study, our method enables non-experts to solve 33.3% more problems, speeds them up by 3.3x, and empowers them to match unassisted experts. Jiaxin Wen, Ruiqi Zhong, Pei Ke, Zhihong Shao, Hongning Wang, Minlie Huang |
ACL (1) | 3 |
| 2024 | Defending Large Language Models Against Jailbreaking Attacks Through Goal PrioritizationabstractWhile significant attention has been dedicated to exploiting weaknesses in LLMs through jailbreaking attacks, there remains a paucity of effort in defending against these attacks.We point out a pivotal factor contributing to the success of jailbreaks: the intrinsic conflict between the goals of being helpful and ensuring safety.Accordingly, we propose to integrate goal prioritization at both training and inference stages to counteract.Implementing goal prioritization during inference substantially diminishes the Attack Success Rate (ASR) of jailbreaking from 66.4% to 3.6% for ChatGPT.And integrating goal prioritization into model training reduces the ASR from 71.0% to 6.6% for Llama2-13B.Remarkably, even in scenarios where no jailbreaking samples are included during training, our approach slashes the ASR by half.Additionally, our findings reveal that while stronger LLMs face greater safety risks, they also possess a greater capacity to be steered towards defending against such attacks, both because of their stronger ability in instruction following.Our work thus contributes to the comprehension of jailbreaking attacks and defenses, and sheds light on the relationship between LLMs' capability and safety.Our code is available at https://github.com/thu-coai/ JailbreakDefense_GoalPriority. Zhexin Zhang, Junxiao Yang, Pei Ke, Fei Mi, Hongning Wang, Minlie Huang |
ACL (1) | 3 |
| 2024 | Language Model Decoding as Direct Metrics OptimizationabstractDespite the remarkable advances in language modeling, current mainstream decoding methods still struggle to generate texts that align with human texts across different aspects. In particular, sampling-based methods produce less-repetitive texts which are often disjunctive in discourse, while search-based methods maintain topic coherence at the cost of increased repetition. Overall, these methods fall short in achieving holistic alignment across a broad range of aspects. In this work, we frame decoding from a language model as an optimization problem with the goal of strictly matching the expected performance with human texts measured by multiple metrics of desired aspects simultaneously. The resulting decoding distribution enjoys an analytical solution that scales the input language model distribution via a sequence-level energy function defined by these metrics. And most importantly, we prove that this induced distribution is guaranteed to improve the perplexity on human texts, which suggests a better approximation to the underlying distribution of human texts. To facilitate tractable sampling from this globally normalized distribution, we adopt the Sampling-Importance-Resampling technique. Experiments on various domains and model scales demonstrate the superiority of our method in metrics alignment with human texts and human evaluation over strong baselines. Haozhe Ji, Pei Ke, Hongning Wang, Minlie Huang |
ICLR | 2 |
| 2024 | Towards Efficient Exact Optimization of Language Model AlignmentabstractThe alignment of language models with human preferences is vital for their application in real-world tasks. The problem is formulated as optimizing the model's policy to maximize the expected reward that reflects human preferences with minimal deviation from the initial policy. While considered as a straightforward solution, reinforcement learning (RL) suffers from high variance in policy updates, which impedes efficient policy improvement. Recently, direct preference optimization (DPO) was proposed to directly optimize the policy from preference data. However, we show that DPO derived based on the optimal solution of the problem leads to a compromised mean-seeking approximation of the optimal solution in practice. In this paper, we propose efficient exact optimization (EXO) of the alignment objective. EXO is guaranteed to optimize in the same direction as RL algorithms asymptotically for arbitrary policy parametrization. This leads to the same mode-seeking solution, while enables efficient optimization by circumventing the complexities of RL. We also compare our method to DPO with both theoretical and empirical analyses, and further demonstrate the advantages of our method over existing approaches on realistic human preference data. Code is available at https://github.com/haozheji/exact-optimization. Haozhe Ji, Cheng Lu 0011, Yilin Niu, Pei Ke, Hongning Wang, Jun Zhu 0001, Jie Tang 0001, Minlie Huang |
ICML | 4 |
| 2024 | Benchmarking Complex Instruction-Following with Multiple Constraints CompositionabstractInstruction following is one of the fundamental capabilities of large language models (LLMs). As the ability of LLMs is constantly improving, they have been increasingly applied to deal with complex human instructions in real-world scenarios. Therefore, how to evaluate the ability of complex instruction-following of LLMs has become a critical research problem. Existing benchmarks mainly focus on modeling different types of constraints in human instructions while neglecting the composition of different constraints, which is an indispensable constituent in complex instructions. To this end, we propose ComplexBench, a benchmark for comprehensively evaluating the ability of LLMs to follow complex instructions composed of multiple constraints. We propose a hierarchical taxonomy for complex instructions, including 4 constraint types, 19 constraint dimensions, and 4 composition types, and manually collect a high-quality dataset accordingly. To make the evaluation reliable, we augment LLM-based evaluators with rules to effectively verify whether generated texts can satisfy each constraint and composition. Furthermore, we obtain the final evaluation score based on the dependency structure determined by different composition types. ComplexBench identifies significant deficiencies in existing LLMs when dealing with complex instructions with multiple constraints composition. Bosi Wen, Pei Ke, Xiaotao Gu, Lindong Wu, Jinfeng Zhou, Wenchuang Li, Binxin Hu, Wendy Gao, Jiaxing Xu, Jie Tang 0001, Hongning Wang, Minlie Huang |
NeurIPS | 2 |
| 2024 | Perception of Knowledge Boundary for Large Language Models through Semi-open-ended Question AnsweringabstractLarge Language Models (LLMs) are widely used for knowledge-seeking purposes yet suffer from hallucinations. The knowledge boundary of an LLM limits its factual understanding, beyond which it may begin to hallucinate. Investigating the perception of LLMs' knowledge boundary is crucial for detecting hallucinations and LLMs' reliable generation. Current studies perceive LLMs' knowledge boundary on questions with concrete answers (close-ended questions) while paying limited attention to semi-open-ended questions that correspond to many potential answers. Some researchers achieve it by judging whether the question is answerable or not. However, this paradigm is not so suitable for semi-open-ended questions, which are usually ``partially answerable questions'' containing both answerable answers and ambiguous (unanswerable) answers. Ambiguous answers are essential for knowledge-seeking, but it may go beyond the knowledge boundary of LLMs. In this paper, we perceive the LLMs' knowledge boundary with semi-open-ended questions by discovering more ambiguous answers. First, we apply an LLM-based approach to construct semi-open-ended questions and obtain answers from a target LLM. Unfortunately, the output probabilities of mainstream black-box LLMs are inaccessible to sample more low-probability ambiguous answers. Therefore, we apply an open-sourced auxiliary model to explore ambiguous answers for the target LLM. We calculate the nearest semantic representation for existing answers to estimate their probabilities, with which we reduce the generation probability of high-probability existing answers to achieve a more effective generation. Finally, we compare the results from the RAG-based evaluation and LLM self-evaluation to categorize four types of ambiguous answers that are beyond the knowledge boundary of the target LLM. Following our method, we construct a dataset to perceive the knowledge boundary for GPT-4. We find that GPT-4 performs poorly on semi-open-ended questions and is often unaware of its knowledge boundary. Besides, our auxiliary model, LLaMA-2-13B, is effective in discovering many ambiguous answers, including correct answers neglected by GPT-4 and delusive wrong answers GPT-4 struggles to identify. Zhihua Wen, Zhiliang Tian, Zexin Jian, Zhen Huang 0006, Pei Ke, Yifu Gao, Minlie Huang, Dongsheng Li 0001 |
NeurIPS | 5 |
| 2024 | ChatGPT: potential, prospects, and limitations
Jie Zhou 0015, Pei Ke, Xipeng Qiu, Minlie Huang, Junping Zhang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2023 | DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringabstractExisting evaluation metrics for natural language generation (NLG) tasks face the challenges on generalization ability and interpretability.Specifically, most of the wellperformed metrics are required to train on evaluation datasets of specific NLG tasks and evaluation dimensions, which may cause over-fitting to task-specific datasets.Furthermore, existing metrics only provide an evaluation score for each dimension without revealing the evidence to interpret how this score is obtained.To deal with these challenges, we propose a simple yet effective metric called DecompEval.This metric formulates NLG evaluation as an instruction-style question answering task and utilizes instruction-tuned pre-trained language models (PLMs) without training on evaluation datasets, aiming to enhance the generalization ability.To make the evaluation process more interpretable, we decompose our devised instruction-style question about the quality of generated texts into the subquestions that measure the quality of each sentence.The subquestions with their answers generated by PLMs are then recomposed as evidence to obtain the evaluation result.Experimental results show that DecompEval achieves state-of-the-art performance in untrained metrics for evaluating text summarization and dialogue generation, which also exhibits strong dimension-level / task-level generalization ability and interpretability 1 . Pei Ke, Fei Huang 0005, Fei Mi, Yasheng Wang, Qun Liu 0001, Xiaoyan Zhu 0001, Minlie Huang |
ACL (1) | 1 |
| 2023 | Unveiling the Implicit Toxicity in Large Language ModelsabstractThe open-endedness of large language models (LLMs) combined with their impressive capabilities may lead to new safety issues when being exploited for malicious use.While recent studies primarily focus on probing toxic outputs that can be easily detected with existing toxicity classifiers, we show that LLMs can generate diverse implicit toxic outputs that are exceptionally difficult to detect via simply zero-shot prompting.Moreover, we propose a reinforcement learning (RL) based attacking method to further induce the implicit toxicity in LLMs.Specifically, we optimize the language model with a reward that prefers implicit toxic outputs to explicit toxic and non-toxic ones.Experiments on five widely-adopted toxicity classifiers demonstrate that the attack success rate can be significantly improved through RL fine-tuning.For instance, the RL-finetuned LLaMA-13B model achieves an attack success rate of 90.04% on BAD and 62.85% on Davinci003.Our findings suggest that LLMs pose a significant threat in generating undetectable implicit toxic outputs.We further show that fine-tuning toxicity classifiers on the annotated examples from our attacking method can effectively enhance their ability to detect LLM-generated implicit toxic language.The code is publicly available at https://github. com/thu-coai/Implicit-Toxicity. Jiaxin Wen, Pei Ke, Hao Sun 0012, Zhexin Zhang, Chengfei Li, Jinfeng Bai, Minlie Huang |
EMNLP | 2 |
| 2023 | Automating Vehicle SOA Threat Analysis Using a Model-Based Methodology
Yuri Gil Dantas, Simon Barner, Pei Ke, Vivek Nigam, Ulrich Schöpp |
ICISSP | 3 |
| 2023 | Tailoring Language Generation Models under Total Variation Distance
Haozhe Ji, Pei Ke, Zhipeng Hu, Minlie Huang |
ICLR | 2 |
| 2023 | Directed Acyclic Transformer Pre-training for High-quality Non-autoregressive Text GenerationabstractAbstract Non-AutoRegressive (NAR) text generation models have drawn much attention because of their significantly faster decoding speed and good generation quality in machine translation. However, in a wider range of text generation tasks, existing NAR models lack proper pre-training, making them still far behind the pre-trained autoregressive models. In this paper, we propose Pre-trained Directed Acyclic Transformer (PreDAT) and a novel pre-training task to promote prediction consistency in NAR generation. Experiments on five text generation tasks show that our PreDAT remarkably outperforms existing pre-trained NAR models (+4.2 score on average) and even achieves better results than pre-trained autoregressive baselines in n-gram-based metrics, along with 17 times speedup in throughput. Further analysis shows that PreDAT benefits from the unbiased prediction order that alleviates the error accumulation problem in autoregressive generation, which provides new insights into the advantages of NAR generation.1 Fei Huang 0005, Pei Ke, Minlie Huang |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text GenerationabstractExisting reference-free metrics have obvious limitations for evaluating controlled text generation models.Unsupervised metrics can only provide a task-agnostic evaluation result which correlates weakly with human judgments, whereas supervised ones may overfit task-specific data with poor generalization ability to other datasets.In this paper, we propose an unsupervised reference-free metric called CTRLEval, which evaluates controlled text generation from different aspects by formulating each aspect into multiple text infilling tasks.On top of these tasks, the metric assembles the generation probabilities from a pre-trained language model without any model training.Experimental results show that our metric has higher correlations with human judgments than other baselines, while obtaining better generalization of evaluating generated texts from different models and with different qualities 1 . Pei Ke, Hao Zhou 0012, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Xiaoyan Zhu 0001, Minlie Huang |
ACL (1) | 1 |
| 2022 | Learning Instructions with Unlabeled Data for Zero-Shot Cross-Task GeneralizationabstractTraining language models to learn from human instructions for zero-shot cross-task generalization has attracted much attention in NLP communities.Recently, instruction tuning (IT), which fine-tunes a pre-trained language model on a massive collection of tasks described via human-craft instructions, has been shown effective in instruction learning for unseen tasks.However, IT relies on a large amount of humanannotated samples, which restricts its generalization.Unlike labeled data, unlabeled data are often massive and cheap to obtain.In this work, we study how IT can be improved with unlabeled data.We first empirically explore the IT performance trends versus the number of labeled data, instructions, and training tasks.We find it critical to enlarge the number of training instructions, and the instructions can be underutilized due to the scarcity of labeled data.Then, we propose Unlabeled Data Augmented Instruction Tuning (UDIT) to take better advantage of the instructions during IT by constructing pseudo-labeled data from unlabeled plain texts.We conduct extensive experiments to show UDIT's effectiveness in various scenarios of tasks and datasets.We also comprehensively analyze the key factors of UDIT to investigate how to better improve IT with unlabeled data. Yuxian Gu, Pei Ke, Xiaoyan Zhu 0001, Minlie Huang |
EMNLP | 2 |
| 2022 | Curriculum-Based Self-Training Makes Better Few-Shot Learners for Data-to-Text GenerationabstractDespite the success of text-to-text pre-trained models in various natural language generation (NLG) tasks, the generation performance is largely restricted by the number of labeled data in downstream tasks, particularly in data-to-text generation tasks. Existing works mostly utilize abundant unlabeled structured data to conduct unsupervised pre-training for task adaption, which fail to model the complex relationship between source structured data and target texts. Thus, we introduce self-training as a better few-shot learner than task-adaptive pre-training, which explicitly captures this relationship via pseudo-labeled data generated by the pre-trained model. To alleviate the side-effect of low-quality pseudo-labeled data during self-training, we propose a novel method called Curriculum-Based Self-Training (CBST) to effectively leverage unlabeled data in a rearranged order determined by the difficulty of text generation. Experimental results show that our method can outperform fine-tuning and task-adaptive pre-training methods, and achieve state-of-the-art performance in the few-shot setting of data-to-text generation. Pei Ke, Haozhe Ji, Yi Huang 0017, Junlan Feng, Xiaoyan Zhu 0001, Minlie Huang |
IJCAI | 1 |
| 2020 | Language Generation with Multi-Hop Reasoning on Commonsense Knowledge GraphabstractDespite the success of generative pre-trained language models on a series of text generation tasks, they still suffer in cases where reasoning over underlying commonsense knowledge is required during generation.Existing approaches that integrate commonsense knowledge into generative pre-trained language models simply transfer relational knowledge by post-training on individual knowledge triples while ignoring rich connections within the knowledge graph.We argue that exploiting both the structural and semantic information of the knowledge graph facilitates commonsenseaware text generation.In this paper, we propose Generation with Multi-Hop Reasoning Flow (GRF) that enables pre-trained models with dynamic multi-hop reasoning on multirelational paths extracted from the external commonsense knowledge graph.We empirically show that our model outperforms existing baselines on three text generation tasks that require reasoning over commonsense knowledge.We also demonstrate the effectiveness of the dynamic multi-hop reasoning module with reasoning paths inferred by the model that provide rationale to the generation. 1 Haozhe Ji, Pei Ke, Shaohan Huang, Furu Wei, Xiaoyan Zhu 0001, Minlie Huang |
EMNLP (1) | 2 |
| 2020 | SentiLARE: Sentiment-Aware Language Representation Learning with Linguistic KnowledgeabstractMost of the existing pre-trained language representation models neglect to consider the linguistic knowledge of texts, which can promote language understanding in NLP tasks.To benefit the downstream tasks in sentiment analysis, we propose a novel language representation model called SentiLARE, which introduces word-level linguistic knowledge including part-of-speech tag and sentiment polarity (inferred from SentiWordNet) into pretrained models.We first propose a contextaware sentiment attention mechanism to acquire the sentiment polarity of each word with its part-of-speech tag by querying SentiWord-Net.Then, we devise a new pre-training task called label-aware masked language model to construct knowledge-aware language representation.Experiments show that SentiLARE obtains new state-of-the-art performance on a variety of sentiment analysis tasks 1 . Pei Ke, Haozhe Ji, Siyang Liu 0003, Xiaoyan Zhu 0001, Minlie Huang |
EMNLP (1) | 1 |
| 2020 | A Large-Scale Chinese Short-Text Conversation Dataset
Yida Wang 0009, Pei Ke, Yinhe Zheng, Kaili Huang, Yong Jiang 0001, Xiaoyan Zhu 0001, Minlie Huang |
NLPCC (1) | 2 |
| 2019 | ARAML: A Stable Adversarial Training Framework for Text GenerationabstractPei Ke, Fei Huang, Minlie Huang, Xiaoyan Zhu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Pei Ke, Fei Huang 0005, Minlie Huang, Xiaoyan Zhu 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Generating Informative Responses with Controlled Sentence FunctionabstractSentence function is a significant factor to achieve the purpose of the speaker, which, however, has not been touched in largescale conversation generation so far.In this paper, we present a model to generate informative responses with controlled sentence function.Our model utilizes a continuous latent variable to capture various word patterns that realize the expected sentence function, and introduces a type controller to deal with the compatibility of controlling sentence function and generating informative content.Conditioned on the latent variable, the type controller determines the type (i.e., function-related, topic, and ordinary word) of a word to be generated at each decoding position.Experiments show that our model outperforms state-of-the-art baselines, and it has the ability to generate responses with both controlled sentence function and informative content. Pei Ke, Jian Guan 0002, Minlie Huang, Xiaoyan Zhu 0001 |
ACL (1) | 1 |