EDBT 2026 Demo / reviewers in the wild / expert
Fei Mi
dblp:161/0068
· DBLP profile ↗
37ranked-venue papers
8as first author
30since 2021 · last 2026
0000-0001-6358-9922ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 6 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay WritingabstractPrompt-based essay writing is an effective and common way to assess students' critical thinking skills. Recent work has evaluated the impressive capabilities of Large Language Models (LLMs) on this task. However, most studies focus primarily on English. Those examining LLMs' performance in Chinese often rely on coarse-grained text quality metrics, overlooking the structural and rhetorical complexities of Chinese essays, particularly across diverse genres. We therefore propose EssayBench, a multi-genre benchmark specifically designed for Chinese essay writing, along with a fine-grained, genre-specific scoring framework that hierarchically aggregates scores to better align with human preferences. The dataset comprises 728 real-world prompts across four major genres (Argumentative, Narrative, Descriptive, and Expository), and includes both Open-Ended and Constrained types. Our evaluation protocol is validated through a comprehensive human agreement study. The results show that our protocol aligns well with human judgments, achieving a highest Spearman's correlation of 0.816 and outperforming coarse-grained evaluation methods by an average of 8.6\%. Finally, we benchmark 15 large LLMs, analyzing their strengths and limitations across genres and instruction types. We believe EssayBench offers a more reliable framework for evaluating Chinese essay generation and provides valuable insights for improving LLMs in this domain. Dongyuan Li, Ding Xia, Fei Mi, Yasheng Wang, Lifeng Shang, Baojun Wang |
AAAI | 4 |
| 2026 | Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical ReasoningabstractBowen Ding, Yuhan Chen, Jiayang Lyu, Jiyao Yuan, Qi Zhu, Shuangshuang Tian, Dantong Zhu, Futing Wang, Heyuan Deng, Fei Mi, Lifeng Shang, Tao Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiayang Lyu, Jiyao Yuan, Qi Zhu 0011, Shuangshuang Tian, Dantong Zhu, Futing Wang, Heyuan Deng, Fei Mi, Lifeng Shang |
ACL (1) | 10 |
| 2026 | How Should We Enhance the Safety of Large Reasoning Models: An Empirical StudyabstractZhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang, Qi Zhu, Shiyao Cui, Fei Mi, Lifeng Shang, Yingkang Wang, Hongning Wang, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang, Qi Zhu 0011, Shiyao Cui, Fei Mi, Lifeng Shang, Yingkang Wang, Hongning Wang, Minlie Huang |
ACL (1) | 7 |
| 2025 | Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-AlignmentabstractZhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong, Jiahui Gao, Fei Mi, Yu Zhang, Zhenguo Li, Xin Jiang, Qun Liu, James Kwok. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhili Liu, Yunhao Gou, Kai Chen 0023, Lanqing Hong, Jiahui Gao 0002, Fei Mi, Yu Zhang 0006, Zhenguo Li, Xin Jiang 0002, Qun Liu 0001, James T. Kwok |
ACL (1) | 6 |
| 2025 | UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language ModelsabstractBoyang Xue, Fei Mi, Qi Zhu, Hongru Wang, Rui Wang, Sheng Wang, Erxin Yu, Xuming Hu, Kam-Fai Wong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Boyang Xue, Fei Mi, Qi Zhu 0007, Hongru Wang 0003, Rui Wang 0092, Erxin Yu, Xuming Hu, Kam-Fai Wong |
ACL (1) | 2 |
| 2025 | Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical ReasoningabstractErxin Yu, Jing Li, Ming Liao, Qi Zhu, Boyang Xue, Minghui Xu, Baojun Wang, Lanqing Hong, Fei Mi, Lifeng Shang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Erxin Yu, Jing Li 0049, Ming Liao, Qi Zhu 0011, Boyang Xue, Baojun Wang, Lanqing Hong, Fei Mi, Lifeng Shang |
ACL (1) | 9 |
| 2025 | Dynamic data selection with normalized gradient-based influence approximation for targeted fine-tuning of LLMs
Zige Wang, Qi Zhu 0011, Fei Mi, Yasheng Wang, Lifeng Shang |
Knowl. Based Syst. | 3 |
| 2024 | FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsabstractYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yufei Wang 0005, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Wei Wang 0011 |
ACL (1) | 6 |
| 2024 | Defending Large Language Models Against Jailbreaking Attacks Through Goal PrioritizationabstractWhile significant attention has been dedicated to exploiting weaknesses in LLMs through jailbreaking attacks, there remains a paucity of effort in defending against these attacks.We point out a pivotal factor contributing to the success of jailbreaks: the intrinsic conflict between the goals of being helpful and ensuring safety.Accordingly, we propose to integrate goal prioritization at both training and inference stages to counteract.Implementing goal prioritization during inference substantially diminishes the Attack Success Rate (ASR) of jailbreaking from 66.4% to 3.6% for ChatGPT.And integrating goal prioritization into model training reduces the ASR from 71.0% to 6.6% for Llama2-13B.Remarkably, even in scenarios where no jailbreaking samples are included during training, our approach slashes the ASR by half.Additionally, our findings reveal that while stronger LLMs face greater safety risks, they also possess a greater capacity to be steered towards defending against such attacks, both because of their stronger ability in instruction following.Our work thus contributes to the comprehension of jailbreaking attacks and defenses, and sheds light on the relationship between LLMs' capability and safety.Our code is available at https://github.com/thu-coai/ JailbreakDefense_GoalPriority. Zhexin Zhang, Junxiao Yang, Pei Ke, Fei Mi, Hongning Wang, Minlie Huang |
ACL (1) | 4 |
| 2024 | UniRetriever: Multi-task Candidates Selection for Various Context-Adaptive Conversational RetrievalabstractConversational retrieval refers to an information retrieval system that operates in an iterative and interactive manner, requiring the retrieval of various external resources, such as persona, knowledge, and even response, to effectively engage with the user and successfully complete the dialogue. However, most previous work trained independent retrievers for each specific resource, resulting in sub-optimal performance and low efficiency. Thus, we propose a multi-task framework function as a universal retriever for three dominant retrieval tasks during the conversation: persona selection, knowledge selection, and response selection. To this end, we design a dual-encoder architecture consisting of a context-adaptive dialogue encoder and a candidate encoder, aiming to attention to the relevant context from the long dialogue and retrieve suitable candidates by simply a dot product. Furthermore, we introduce two loss constraints to capture the subtle relationship between dialogue context and different candidates by regarding historically selected candidates as hard negatives. Extensive experiments and analysis establish state-of-the-art retrieval quality both within and outside its training domain, revealing the promising potential and generalization capability of our model to serve as a universal retriever for different candidate selection tasks simultaneously. Hongru Wang 0003, Boyang Xue, Baohang Zhou, Rui Wang 0092, Fei Mi, Weichao Wang, Yasheng Wang, Kam-Fai Wong |
LREC/COLING | 5 |
| 2024 | CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue CoreferenceabstractAs large language models (LLMs) constantly evolve, ensuring their safety remains a critical research issue.Previous red teaming approaches for LLM safety have primarily focused on single prompt attack or goal hijacking.To the best of our knowledge, we are the first to study LLM safety in multi-turn dialogue coreference.We created a dataset of 1, 400 questions across 14 categories, each featuring multi-turn coreference safety attacks.We then conducted detailed evaluations on five widely used open-source LLMs.The results indicated that under multi-turn coreference safety attacks, the highest attack successful rate was 56% with the LLaMA2-Chat-7b model, while the lowest was 13.9% with the Mistral-7B-Instruct model.These findings highlight the safety vulnerabilities in LLMs during dialogue coreference interactions.Warning: This paper may contain offensive language or harmful content. 1 Erxin Yu, Jing Li 0049, Ming Liao, Zuchen Gao, Fei Mi, Lanqing Hong |
EMNLP | 6 |
| 2024 | Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake AnalysisabstractThe rapid development of large language models (LLMs) has not only provided numerous opportunities but also presented significant challenges. This becomes particularly evident when LLMs inadvertently generate harmful or toxic content, either unintentionally or because of intentional inducement. Existing alignment methods usually direct LLMs toward the favorable outcomes by utilizing human-annotated, flawless instruction-response pairs. Conversely, this study proposes a novel alignment technique based on mistake analysis, which deliberately exposes LLMs to erroneous content to learn the reasons for mistakes and how to avoid them. In this case, mistakes are repurposed into valuable data for alignment, effectively helping to avoid the production of erroneous responses. Without external models or human annotations, our method leverages a model's intrinsic ability to discern undesirable mistakes and improves the safety of its generated responses. Experimental results reveal that our method outperforms existing alignment approaches in enhancing model safety while maintaining the overall utility. Kai Chen 0023, Chunwei Wang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu 0004, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
ICLR | 6 |
| 2024 | Enhancing Large Language Models Against Inductive Instructions with Dual-critique PromptingabstractRui Wang, Hongru Wang, Fei Mi, Boyang Xue, Yi Chen, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Rui Wang 0092, Hongru Wang 0003, Fei Mi, Boyang Xue, Yi Chen 0007, Kam-Fai Wong, Ruifeng Xu 0001 |
NAACL-HLT | 3 |
| 2023 | KPT: Keyword-Guided Pre-training for Grounded Dialog GenerationabstractIncorporating external knowledge into the response generation process is essential to building more helpful and reliable dialog agents. However, collecting knowledge-grounded conversations is often costly, calling for a better pre-trained model for grounded dialog generation that generalizes well w.r.t. different types of knowledge. In this work, we propose KPT (Keyword-guided Pre-Training), a novel self-supervised pre-training method for grounded dialog generation without relying on extra knowledge annotation. Specifically, we use a pre-trained language model to extract the most uncertain tokens in the dialog as keywords. With these keywords, we construct two kinds of knowledge and pre-train a knowledge-grounded response generation model, aiming at handling two different scenarios: (1) the knowledge should be faithfully grounded; (2) it can be selectively used. For the former, the grounding knowledge consists of keywords extracted from the response. For the latter, the grounding knowledge is additionally augmented with keywords extracted from other utterances in the same dialog. Since the knowledge is extracted from the dialog itself, KPT can be easily performed on a large volume and variety of dialogue data. We considered three data sources (open-domain, task-oriented, conversational QA) with a total of 2.5M dialogues. We conduct extensive experiments on various few-shot knowledge-grounded generation tasks, including grounding on dialog acts, knowledge graphs, persona descriptions, and Wikipedia passages. Our comprehensive experiments and analyses demonstrate that KPT consistently outperforms state-of-the-art methods on these tasks with diverse grounding knowledge. Qi Zhu 0007, Fei Mi, Zheng Zhang 0020, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Xiaoyan Zhu 0001, Minlie Huang |
AAAI | 2 |
| 2023 | Towards Diverse, Relevant and Coherent Open-Domain Dialogue Generation via Hybrid Latent VariablesabstractConditional variational models, using either continuous or discrete latent variables, are powerful for open-domain dialogue response generation. However, previous works show that continuous latent variables tend to reduce the coherence of generated responses. In this paper, we also found that discrete latent variables have difficulty capturing more diverse expressions. To tackle these problems, we combine the merits of both continuous and discrete latent variables and propose a Hybrid Latent Variable (HLV) method. Specifically, HLV constrains the global semantics of responses through discrete latent variables and enriches responses with continuous latent variables. Thus, we diversify the generated responses while maintaining relevance and coherence. In addition, we propose Conditional Hybrid Variational Transformer (CHVT) to construct and to utilize HLV with transformers for dialogue generation. Through fine-grained symbolic-level semantic information and additive Gaussian mixing, we construct the distribution of continuous variables, prompting the generation of diverse expressions. Meanwhile, to maintain the relevance and coherence, the discrete latent variable is optimized by self-separation training. Experimental results on two dialogue generation datasets (DailyDialog and Opensubtitles) show that CHVT is superior to traditional transformer-based variational mechanism w.r.t. diversity, relevance and coherence metrics. Moreover, we also demonstrate the benefit of applying HLV to fine-tuning two pre-trained dialogue models (PLATO and BART-base). Bin Sun 0004, Fei Mi, Weichao Wang, Yiwei Li 0001, Kan Li 0001 |
AAAI | 3 |
| 2023 | MoralDial: A Framework to Train and Evaluate Moral Dialogue Systems via Moral DiscussionsabstractHao Sun, Zhexin Zhang, Fei Mi, Yasheng Wang, Wei Liu, Jianwei Cui, Bin Wang, Qun Liu, Minlie Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hao Sun 0012, Zhexin Zhang, Fei Mi, Yasheng Wang, Wei Liu 0005, Jianwei Cui 0002, Bin Wang 0004, Qun Liu 0001, Minlie Huang |
ACL (1) | 3 |
| 2023 | A Synthetic Data Generation Framework for Grounded DialoguesabstractTraining grounded response generation models often requires a large collection of grounded dialogues.However, it is costly to build such dialogues.In this paper, we present a synthetic data generation framework (SynDG) for grounded dialogues.The generation process utilizes large pre-trained language models and freely available knowledge data (e.g., Wikipedia pages, persona profiles, etc.).The key idea of designing SynDG is to consider dialogue flow and coherence in the generation process.Specifically, given knowledge data, we first heuristically determine a dialogue flow, which is a series of knowledge pieces.Then, we employ T5 to incrementally turn the dialogue flow into a dialogue.To ensure coherence of both the dialogue flow and the synthetic dialogue, we design a two-level filtering strategy, at the flow-level and the utterance-level respectively.Experiments on two public benchmarks show that the synthetic grounded dialogue data produced by our framework is able to significantly boost model performance in both full training data and low-resource scenarios. Jianzhu Bao, Rui Wang 0092, Yasheng Wang, Aixin Sun, Fei Mi, Ruifeng Xu 0001 |
ACL (1) | 6 |
| 2023 | DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringabstractExisting evaluation metrics for natural language generation (NLG) tasks face the challenges on generalization ability and interpretability.Specifically, most of the wellperformed metrics are required to train on evaluation datasets of specific NLG tasks and evaluation dimensions, which may cause over-fitting to task-specific datasets.Furthermore, existing metrics only provide an evaluation score for each dimension without revealing the evidence to interpret how this score is obtained.To deal with these challenges, we propose a simple yet effective metric called DecompEval.This metric formulates NLG evaluation as an instruction-style question answering task and utilizes instruction-tuned pre-trained language models (PLMs) without training on evaluation datasets, aiming to enhance the generalization ability.To make the evaluation process more interpretable, we decompose our devised instruction-style question about the quality of generated texts into the subquestions that measure the quality of each sentence.The subquestions with their answers generated by PLMs are then recomposed as evidence to obtain the evaluation result.Experimental results show that DecompEval achieves state-of-the-art performance in untrained metrics for evaluating text summarization and dialogue generation, which also exhibits strong dimension-level / task-level generalization ability and interpretability 1 . Pei Ke, Fei Huang 0005, Fei Mi, Yasheng Wang, Qun Liu 0001, Xiaoyan Zhu 0001, Minlie Huang |
ACL (1) | 3 |
| 2023 | One Cannot Stand for Everyone! Leveraging Multiple User Simulators to train Task-oriented Dialogue SystemsabstractYajiao Liu, Xin Jiang, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu, Xiang Wan, Benyou Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yajiao Liu, Xin Jiang 0002, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu 0001, Benyou Wang |
ACL (1) | 5 |
| 2023 | Retrieval-free Knowledge Injection through Multi-Document Traversal for Dialogue ModelsabstractRui Wang, Jianzhu Bao, Fei Mi, Yi Chen, Hongru Wang, Yasheng Wang, Yitong Li, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Rui Wang 0092, Jianzhu Bao, Fei Mi, Yi Chen 0007, Hongru Wang 0003, Yasheng Wang, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu 0001 |
ACL (1) | 3 |
| 2023 | ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain DialogueabstractIncorporating visual knowledge into text-only dialogue systems has become a potential direction to imitate the way humans think, imagine, and communicate.However, existing multimodal dialogue systems are either confined by the scale and quality of available datasets or the coarse concept of visual knowledge.To address these issues, we provide a new paradigm of constructing multimodal dialogues as well as two datasets extended from text-only dialogues under such paradigm (RESEE-WoW, RESEE-DD).We propose to explicitly split the visual knowledge into finer granularity ("turn-level" and "entity-level").To further boost the accuracy and diversity of augmented visual information, we retrieve them from the Internet or a large image dataset.To demonstrate the superiority and universality of the provided visual knowledge, we propose a simple but effective framework RESEE to add visual representation into vanilla dialogue models by modality concatenations.We also conduct extensive experiments and ablations w.r.t.different model configurations and visual knowledge settings.Empirically, encouraging results not only demonstrate the effectiveness of introducing visual knowledge at both entity and turn level but also verify the proposed model RESEE outperforms several state-of-the-art methods on automatic and human evaluations.By leveraging text and vision knowledge, RESEE can produce informative responses with real-world visual concepts.Our code is available at https: //github.com/ImKeTT/ReSee. Haoqin Tu, Fei Mi, Zhongliang Yang |
EMNLP | 3 |
| 2023 | History, Present and Future: Enhancing Dialogue Generation with Few-Shot History-Future PromptabstractDialogue history and response in open-domain dialogue are loosely coupled. Generating informative responses solely based on the original dialogue history is not easy, as dialogue history may not contain enough information or it may contain irrelevant noises. Intuitively, if a generation model can foresee possible dialogue future, or obtain real useful histories, it could generate more informative responses. In this paper, we propose a novel lightweight dialogue generation framework named few-shot history-future prompt that utilizes useful histories and simulated futures to help generate informative responses, without the need for fine-tuning or adding extra parameters. To obtain useful histories, we retrieve and combine relevant utterances from noisy multi- turn histories. Then we adopt a retrieval-generation hybrid approach to obtain diversified simulated futures. Such that our model could learn to condition on history combinations and simulated futures via few-shot learning. Experiments over publicly available datasets demonstrate that our method can help models generate better responses. Yasheng Wang, Fei Mi, Pingyi Zhou, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001 |
ICASSP | 4 |
| 2022 | CINS: Comprehensive Instruction for Few-Shot Learning in Task-Oriented Dialog SystemsabstractAs the labeling cost for different modules in task-oriented dialog (ToD) systems is high, a major challenge is to learn different tasks with the least amount of labeled data. Recently, pre-trained language models (PLMs) have shown promising results for few-shot learning in ToD. To better utilize the power of PLMs, this paper proposes Comprehensive Instruction (CINS) that exploits PLMs with extra task-specific instructions. We design a schema (definition, constraint, prompt) of instructions and their customized realizations for three important downstream tasks in ToD, ie. intent classification, dialog state tracking, and natural language generation. A sequence-to-sequence model (T5) is adopted to solve these three tasks in a unified framework. Extensive experiments are conducted on these ToD tasks in realistic few-shot learning scenarios with small validation data. Empirical results demonstrate that the proposed CINS approach consistently improves techniques that finetune PLMs with raw input or short prompt. Fei Mi, Yasheng Wang |
AAAI | 1 |
| 2022 | Continual Prompt Tuning for Dialog State TrackingabstractA desirable dialog system should be able to continually learn new skills without forgetting old ones, and thereby adapt to new domains or tasks in its life cycle.However, continually training a model often leads to a well-known catastrophic forgetting issue.In this paper, we present Continual Prompt Tuning, a parameterefficient framework that not only avoids forgetting but also enables knowledge transfer between tasks.To avoid forgetting, we only learn and store a few prompt tokens' embeddings for each task while freezing the backbone pre-trained model.To achieve bi-directional knowledge transfer among tasks, we propose several techniques (continual prompt initialization, query fusion, and memory replay) to transfer knowledge from preceding tasks and a memory-guided technique to transfer knowledge from subsequent tasks.Extensive experiments demonstrate the effectiveness and efficiency of our proposed method on continual learning for dialog state tracking, compared with state-of-the-art baselines. Qi Zhu 0007, Fei Mi, Xiaoyan Zhu 0001, Minlie Huang |
ACL (1) | 3 |
| 2022 | Pan More Gold from the Sand: Refining Open-domain Dialogue Training with Noisy Self-Retrieval GenerationabstractReal human conversation data are complicated, heterogeneous, and noisy, from which building open-domain dialogue systems remains a challenging task. In fact, such dialogue data still contains a wealth of information and knowledge, however, they are not fully explored. In this paper, we show existing open-domain dialogue generation methods that memorize context-response paired data with autoregressive or encode-decode language models underutilize the training data. Different from current approaches, using external knowledge, we explore a retrieval-generation training framework that can take advantage of the heterogeneous and noisy training data by considering them as “evidence”. In particular, we use BERTScore for retrieval, which gives better qualities of the evidence and generation. Experiments over publicly available datasets demonstrate that our method can help models generate better responses, even such training data are usually impressed as low-quality data. Such performance gain is comparable with those improved by enlarging the training set, even better. We also found that the model performance has a positive correlation with the relevance of the retrieved evidence. Moreover, our method performed well on zero-shot experiments, which indicates that our method can be more robust to real-world data. Yasheng Wang, Fei Mi, Pingyi Zhou, Xin Wang 0114, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001 |
COLING | 4 |
| 2022 | AEG: Argumentative Essay Generation via A Dual-Decoder Model with Content PlanningabstractArgument generation is an important but challenging task in computational argumentation.Existing studies have mainly focused on generating individual short arguments, while research on generating long and coherent argumentative essays is still under-explored.In this paper, we propose a new task, Argumentative Essay Generation (AEG).Given a writing prompt, the goal of AEG is to automatically generate an argumentative essay with strong persuasiveness.We construct a large-scale dataset, ArgEssay, for this new task and establish a strong model based on a dual-decoder Transformer architecture.Our proposed model contains two decoders, a planning decoder (PD) and a writing decoder (WD), where PD is used to generate a sequence for essay content planning and WD incorporates the planning information to write an essay.Further, we pre-train this model on a large news dataset to enhance the plan-and-write paradigm.Automatic and human evaluation results show that our model can generate more coherent and persuasive essays with higher diversity and less repetition compared to several baselines.1 Jianzhu Bao, Yasheng Wang, Fei Mi, Ruifeng Xu 0001 |
EMNLP | 4 |
| 2022 | COLD: A Benchmark for Chinese Offensive Language DetectionabstractOffensive language detection is increasingly crucial for maintaining a civilized social media platform and deploying pre-trained language models.However, this task in Chinese is still under exploration due to the scarcity of reliable datasets.To this end, we propose a benchmark -COLD for Chinese offensive language analysis, including a Chinese Offensive Language Dataset -COLDATASET and a baseline detector -COLDETECTOR which is trained on the dataset.We show that the COLD benchmark contributes to Chinese offensive language detection which is challenging for existing resources.We then deploy the COLDETECTOR and conduct detailed analyses on popular Chinese pre-trained language models.We first analyze the offensiveness of existing generative models and show that these models inevitably expose varying degrees of offensive issues.Furthermore, we investigate the factors that influence the offensive generations, and we find that anti-bias contents and keywords referring to certain groups or revealing negative attitudes trigger offensive outputs easier. Jiawen Deng 0006, Jingyan Zhou, Hao Sun 0012, Chujie Zheng, Fei Mi, Helen M. Meng, Minlie Huang |
EMNLP | 5 |
| 2022 | Slim: Explicit Slot-Intent Mapping with Bert for Joint Multi-Intent Detection and Slot FillingabstractUtterance-level intent detection and token-level slot filling are two key tasks for spoken language understanding (SLU) in task-oriented systems. Most existing approaches assume that only a single intent exists in an utterance. However, there are often multiple intents within an utterance in real-life scenarios. In this paper, we propose a multi-intent SLU framework, called SLIM, to jointly learn multi-intent detection and slot filling based on BERT. To fully exploit the existing annotation data and capture the interactions between slots and intents, SLIM introduces an explicit slot-intent classifier to learn the many-to-one mapping between slots and intents. Empirical results on three public multi-intent datasets demonstrate (1) the superior performance of SLIM compared to the current state-of-the-art for SLU with multiple intents and (2) the benefits obtained from the slot-intent classifier. Fengyu Cai, Wanhao Zhou, Fei Mi, Boi Faltings |
ICASSP | 3 |
| 2022 | Overview of NLPCC 2022 Shared Task 7: Fine-Grained Dialogue Social Bias Measurement
Jingyan Zhou, Fei Mi, Helen M. Meng, Jiawen Deng 0006 |
NLPCC (2) | 2 |
| 2021 | Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog SystemsabstractAs the labeling cost for different modules in task-oriented dialog (ToD) systems is expensive, a major challenge is to train different modules with the least amount of labeled data.Recently, large-scale pre-trained language models, have shown promising results for few-shot learning in ToD.In this paper, we devise a selftraining approach to utilize the abundant unlabeled dialog data to further improve state-ofthe-art pre-trained models in few-shot learning scenarios for ToD systems.Specifically, we propose a self-training approach that iteratively labels the most confident unlabeled data to train a stronger Student model.Moreover, a new text augmentation technique (GradAug) is proposed to better train the Student by replacing non-crucial tokens using a masked language model.We conduct extensive experiments and present analyses on four downstream tasks in ToD, including intent classification, dialog state tracking, dialog act prediction, and response selection.Empirical results demonstrate that the proposed self-training approach consistently improves state-of-the-art pre-trained models (BERT, ToD-BERT) when only a small number of labeled data are available. Fei Mi, Wanhao Zhou, Lingjing Kong 0001, Fengyu Cai, Minlie Huang, Boi Faltings |
EMNLP (1) | 1 |
| 2020 | Masking as an Efficient Alternative to Finetuning for Pretrained Language ModelsabstractWe present an efficient method of utilizing pretrained language models, where we learn selective binary masks for pretrained weights in lieu of modifying them through finetuning.Extensive evaluations of masking BERT, RoBERTa, and DistilBERT on eleven diverse NLP tasks show that our masking scheme yields performance comparable to finetuning, yet has a much smaller memory footprint when several tasks need to be inferred.Intrinsic evaluations show that representations computed by our binary masked language models encode information necessary for solving downstream tasks.Analyzing the loss landscape, we show that masking and finetuning produce models that reside in minima that can be connected by a line segment with nearly constant test accuracy.This confirms that masking can be utilized as an efficient alternative to finetuning. Tao Lin 0004, Fei Mi, Martin Jaggi, Hinrich Schütze |
EMNLP (1) | 3 |
| 2020 | Memory Augmented Neural Model for Incremental Session-based RecommendationabstractIncreasing concerns with privacy have stimulated interests in Session-based Recommendation (SR) using no personal data other than what is observed in the current browser session. Existing methods are evaluated in static settings which rarely occur in real-world applications. To better address the dynamic nature of SR tasks, we study an incremental SR scenario, where new items and preferences appear continuously. We show that existing neural recommenders can be used in incremental SR scenarios with small incremental updates to alleviate computation overhead and catastrophic forgetting. More importantly, we propose a general framework called Memory Augmented Neural model (MAN). MAN augments a base neural recommender with a continuously queried and updated nonparametric memory, and the predictions from the neural and the memory components are combined through another lightweight gating network. We empirically show that MAN is well-suited for the incremental SR task, and it consistently outperforms state-oft-he-art neural and nonparametric methods. We analyze the results and demonstrate that it is particularly good at incrementally learning preferences on new and infrequent items. Fei Mi, Boi Faltings |
IJCAI | 1 |
| 2020 | ADER: Adaptively Distilled Exemplar Replay Towards Continual Learning for Session-based RecommendationabstractSession-based recommendation has received growing attention recently due to the increasing privacy concern. Despite the recent success of neural session-based recommenders, they are typically developed in an offline manner using a static dataset. However, recommendation requires continual adaptation to take into account new and obsolete items and users, and requires “continual learning” in real-life applications. In this case, the recommender is updated continually and periodically with new data that arrives in each update cycle, and the updated model needs to provide recommendations for user activities before the next model update. A major challenge for continual learning with neural models is catastrophic forgetting, in which a continually trained model forgets user preference patterns it has learned before. To deal with this challenge, we propose a method called Adaptively Distilled Exemplar Replay (ADER) by periodically replaying previous training samples (i.e., exemplars) to the current model with an adaptive distillation loss. Experiments are conducted based on the state-of-the-art SASRec model using two widely used datasets to benchmark ADER with several well-known continual learning techniques. We empirically demonstrate that ADER consistently outperforms other baselines, and it even outperforms the method using all historical data at every update cycle. This result reveals that ADER is a promising solution to mitigate the catastrophic forgetting issue towards building more realistic and scalable session-based recommenders. Fei Mi, Boi Faltings |
RecSys | 1 |
| 2019 | Meta-Learning for Low-resource Natural Language Generation in Task-oriented Dialogue SystemsabstractNatural language generation (NLG) is an essential component of task-oriented dialogue systems. Despite the recent success of neural approaches for NLG, they are typically developed for particular domains with rich annotated training examples. In this paper, we study NLG in a low-resource setting to generate sentences in new scenarios with handful training examples. We formulate the problem from a meta-learning perspective, and propose a generalized optimization-based approach (Meta-NLG) based on the well-recognized model-agnostic meta-learning (MAML) algorithm. Meta-NLG defines a set of meta tasks, and directly incorporates the objective of adapting to new low-resource NLG tasks into the meta-learning optimization process. Extensive experiments are conducted on a large multi-domain dataset (MultiWoz) with diverse linguistic variations. We show that Meta-NLG significantly outperforms other training procedures in various low-resource configurations. We analyze the results, and demonstrate that Meta-NLG adapts extremely fast and well to low-resource situations. Fei Mi, Minlie Huang, Jiyong Zhang 0001, Boi Faltings |
IJCAI | 1 |
| 2017 | Adaptive Sequential Recommendation for Discussion Forums on MOOCs using Context Trees
Fei Mi, Boi Faltings |
EDM | 1 |
| 2016 | Adaptive Sequential Recommendation Using Context Trees
Fei Mi, Boi Faltings |
IJCAI | 1 |
| 2015 | Probabilistic Graphical Models for Boosting Cardinal and Ordinal Peer Grading in MOOCsabstractWith the enormous scale of massive open online courses (MOOCs), peer grading is vital for addressing the assessment challenge for open-ended assignments or exams while at the same time providing students with an effective learning experience through involvement in the grading process. Most existing MOOC platforms use simple schemes for aggregating peer grades, e.g., taking the median or mean. To enhance these schemes, some recent research attempts have developed machine learning methods under either the cardinal setting (for absolute judgment) or the ordinal setting (for relative judgment). In this paper, we seek to study both cardinal and ordinal aspects of peer grading within a common framework. First, we propose novel extensions to some existing probabilistic graphical models for cardi- nal peer grading. Not only do these extensions give su- perior performance in cardinal evaluation, but they also outperform conventional ordinal models in ordinal eval- uation. Next, we combine cardinal and ordinal models by augmenting ordinal models with cardinal predictions as prior. Such combination can achieve further performance boosts in both cardinal and ordinal evaluations, suggesting a new research direction to pursue for peer grading on MOOCs. Extensive experiments have been conducted using real peer grading data from a course called “Science, Technology, and Society in China I” offered by HKUST on the Coursera platform. Fei Mi, Dit-Yan Yeung |
AAAI | 1 |