EDBT 2026 Demo / reviewers in the wild / expert
Lei Wang 0185
dblp:181/2817-185
· DBLP profile ↗
33ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0003-1228-6758ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On Reasoning Behind Next Occupation Recommendation
Shan Dong, Palakorn Achananuparp, Hieu Hien Mai, Lei Wang 0185, Ee-Peng Lim |
PAKDD (4) | 4 |
| 2026 | Explainable and Interactive LLMs-Augmented Depression Detection in Social MediaabstractDepression detection based on social media content has received increasing attention in recent years, as it allows for early diagnosis before the user’s psychological state deteriorates. Although traditional methods of depression detection can provide a classification of whether the user is depressed or not, they cannot provide human-like explanations and interactions. In this article, we propose a next-generation paradigm for depression detection, namely an interpretable and interactive depression detection system based on large language models (LLMs). The proposed system not only yields a final diagnosis result, but also offers diagnostic evidence grounded in established diagnostic criteria. Furthermore, it enables users to engage in natural language dialogue with the system, facilitating a more personalized understanding of their mental state based on their social media content. The interactive dialogue allows for the provision of tailored recommendations, which users can utilize to enhance their well-being. In constructing the entire system, we also addressed some nontrivial challenges. First, we introduced the chain of thoughts technique and professional depression diagnostic criteria when constructing the prompts, enabling our system to make decisions based on professional diagnosis criteria and provide explanations. Second, LLMs are incapable of processing excessively long contextual texts, and the accumulated posts of a single user may amount to tens of thousands of words. To overcome this limitation, we integrated a tweet selector that selects the part of posts for diagnosis. The experiments demonstrate that our depression detection system achieves the best performance across various settings, including full data setting, few-shot setting, zero-shot setting, independent-identical-distribution (IID) setting, and out-of-distribution (OOD) setting. Additionally, case studies reveal the explanation and interactivity of our system. Zetong Chen, Xun Yang 0001, Lei Wang 0185, Yunshi Lan, Weijieying Ren, Richang Hong |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient TuningabstractExisting continual learning works explored strategies like memory replay, regularization, and parameter isolation, but little analysis was conducted on the optimization behavior of LLMs’ continual fine-tuning. In this work, we investigate the geometric connections of different minima along the continual LLM fine-tuning trajectories. We validate this phenomenon on LLMs and propose a new method called Interpolation-based LoRA (I-LoRA). I-LoRA can strike a balance between plasticity and stability through parameter interpolation, which constructs a dual-memory experience replay framework based on LoRA. Experiments on eight domain-specific benchmarks demonstrate that I-LoRA consistently shows significant improvement over previous approaches with up to 11% performance gains. Our code is available at https://anonymous.4open.science/r/LLMCL-3823. Xinlong Li, Weijieying Ren, Lei Wang 0185, Tianxiang Zhao 0001, Richang Hong |
ICASSP | 4 |
| 2025 | DVM: Towards Controllable LLM Agents in Social Deduction GamesabstractLarge Language Models (LLMs) have advanced the capability of game agents in social deduction games (SDGs). These games rely heavily on conversation-driven interactions and require agents to infer, make decisions, and express based on such information. While this progress leads to more sophisticated and strategic non-player characters (NPCs) in SDGs, there exists a need to control the proficiency of these agents. This control not only ensures that NPCs can adapt to varying difficulty levels during gameplay, but also provides insights into the safety and fairness of LLM agents. In this paper, we present DVM, a novel framework for developing controllable LLM agents for SDGs, and demonstrate its implementation on one of the most popular SDGs, Werewolf. DVM comprises three main components: Predictor, Decider, and Discussor. By integrating reinforcement learning with a win rate-constrained decision chain reward mechanism, we enable agents to dynamically adjust their gameplay proficiency to achieve specified win rates. Experiments show that DVM not only outperforms existing methods in the Werewolf game, but also successfully modulates its performance levels to meet predefined win rate targets. These results pave the way for LLM agents’ adaptive and balanced gameplay in SDGs, opening new avenues for research in controllable game agents. Zheng Zhang 0064, Yihuai Lan, Yangsen Chen, Lei Wang 0185, Hao Wang 0094 |
ICASSP | 4 |
| 2025 | ThinK: Thinner Key Cache by Query-Driven PruningabstractLarge Language Models (LLMs) have revolutionized the field of natural language processing, achieving unprecedented performance across a variety of applications.
However, their increased computational and memory demands present significant challenges, especially when handling long sequences.
This paper focuses on the long-context scenario, addressing the inefficiencies in KV cache memory consumption during inference.
Unlike existing approaches that optimize the memory based on the sequence length, we identify substantial redundancy in the channel dimension of the KV cache, as indicated by an uneven magnitude distribution and a low-rank structure in the attention weights.
In response, we propose ThinK, a novel query-dependent KV cache pruning method designed to minimize attention weight loss while selectively pruning the least significant channels. Our approach not only maintains or enhances model accuracy but also achieves a reduction in KV cache memory costs by over 20\% compared with vanilla KV cache eviction and quantization methods. For instance, ThinK integrated with KIVI can achieve $2.8\times$ peak memory reduction while maintaining nearly the same quality, enabling a batch size increase from 4$\times$ (with KIVI alone) to 5$\times$ when using a single GPU. Extensive evaluations on the LLaMA and Mistral models across various long-sequence datasets verified the efficiency of ThinK. Our code has been made available at https://github.com/SalesforceAIResearch/ThinK. Zhanming Jie, Hanze Dong, Lei Wang 0185, Aojun Zhou, Amrita Saha, Caiming Xiong, Doyen Sahoo |
ICLR | 4 |
| 2025 | Causal Intervention with Active Learning for Large Vision-Language Models in Egocentric ContextsabstractRecent advancements in Large Vision-Language Models (LVLMs) have attracted considerable attention due to their impressive performance across various downstream tasks. However, these tasks predominantly emphasize third-person perspectives and LVLMs demonstrate inadequate capability in reasoning from a first-person perspective. To mitigate these limitations, we introduce Causal Intervention with Active Learning (CIAL), an innovative approach designed to augment the first-person reasoning capabilities of LVLMs. Specifically, our method first incorporates an Active Learning-driven Knowledge Extraction (ALKE) scheme, which utilizes LVLMs themselves to automatically and autonomously acquire knowledge related to egocentric perspectives. Then, to optimize the model’s response in conjunction with the scene, a Knowledge-guided Causal Intervention (KCI) module is employed, thereby LVLMs can integrate both knowledge and certainty scores for inference. Comprehensive experiments conducted on the EgoThink benchmark demonstrate that our CIAL method significantly improves the models’ ability to understand and reason in egocentric contexts. Our anonymous code is available at https://github.com/running-alpaca/CIAL/. Wenxin Meng, Shenshen Li, Lei Wang 0185, Hao Yang 0015, Xing Xu 0001 |
ICME | 3 |
| 2024 | T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question AnsweringabstractLarge Language Models (LLMs) have recently demonstrated exceptional performance in various Natural Language Processing (NLP) tasks. They have also shown the ability to perform chain-of-thought (CoT) reasoning to solve complex problems. Recent studies have explored CoT reasoning in complex multimodal scenarios, such as the science question answering task, by fine-tuning multimodal models with high-quality human-annotated CoT rationales. However, collecting high-quality COT rationales is usually time-consuming and costly. Besides, the annotated rationales are hardly accurate due to the external essential information missed. To address these issues, we propose a novel method termed T-SciQ that aims at teaching science question answering with LLM signals. The T-SciQ approach generates high-quality CoT rationales as teaching signals and is advanced to train much smaller models to perform CoT reasoning in complex modalities. Additionally, we introduce a novel data mixing strategy to produce more effective teaching data samples for simple and complex science question answer problems. Extensive experimental results show that our T-SciQ method achieves a new state-of-the-art performance on the ScienceQA benchmark, with an accuracy of 96.18%. Moreover, our approach outperforms the most powerful fine-tuned baseline by 4.5%. The code is publicly available at https://github.com/T-SciQ/T-SciQ. Lei Wang 0185, Jiabang He, Xing Xu 0001, Heng Tao Shen |
AAAI | 1 |
| 2024 | LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon GameplayabstractThis paper explores the open research problem of understanding the social behaviors of LLM-based agents.Using Avalon as a testbed, we employ system prompts to guide LLM agents in gameplay.While previous studies have touched on gameplay with LLM agents, research on their social behaviors is lacking.We propose a novel framework, tailored for Avalon, features a multi-agent system facilitating efficient communication and interaction.We evaluate its performance based on game success and analyze LLM agents' social behaviors.Results affirm the framework's effectiveness in creating adaptive agents and suggest LLM-based agents' potential in navigating dynamic social interactions.By examining collaboration and confrontation behaviors, we offer insights into this field's research and applications.Our code is publicly available at https://github.com/ 3DAgentWorld/LLM-Game-Agent. Yihuai Lan, Lei Wang 0185, Yang Wang 0015, Deheng Ye, Peilin Zhao, Ee-Peng Lim, Hui Xiong 0001, Hao Wang 0094 |
EMNLP | 3 |
| 2024 | Gradient-Aware Logit Adjustment Loss for Long-Tailed ClassifierabstractIn the real-world setting, data often follows a long-tailed distribution, where head classes contain significantly more training samples than tail classes. Consequently, models trained on such data tend to be biased toward head classes. The medium of this bias is imbalanced gradients, which include not only the ratio of scale between positive and negative gradients but also imbalanced gradients from different negative classes. Therefore, we propose the Gradient-Aware Logit Adjustment (GALA) loss, which adjusts the logits based on accumulated gradients to balance the optimization process. Additionally, We find that most of the solutions to long-tailed problems are still biased towards head classes in the end, and we propose a simple and post hoc prediction re-balancing strategy to further mitigate the basis toward head class. Extensive experiments are conducted on multiple popular long-tailed recognition benchmark datasets to evaluate the effectiveness of these two designs. Our approach achieves top-1 accuracy of 48.5%, 41.4%, and 73.3% on CIFAR100-LT, Places-LT, and iNaturalist, outperforming the state-of-the-art method GCL by a significant margin of 3.62%, 0.76% and 1.2%, respectively. Code is available at https://github.com/lt-project-repository/lt-project. Weijieying Ren, Lei Wang 0185, Zetong Chen, Richang Hong |
ICASSP | 4 |
| 2024 | MoCoSA: Momentum Contrast for Knowledge Graph Completion with Structure-Augmented Pre-trained Language ModelsabstractKnowledge Graph Completion (KGC) aims to conduct reasoning on the facts within knowledge graphs and automatically infer missing links. Existing methods can mainly be categorized into structure-based or description-based. Structure-based methods effectively represent relational facts in knowledge graphs using entity embeddings and description-based methods leverage pre-trained language models (PLMs) to understand textual information. In this paper, we propose Momentum Contrast for knowledge graph completion with Structure-Augmented pre-trained language models (MoCoSA), which allows the PLM to perceive the structural information by the adaptable structure encoder. We proposed momentum hard negative and intra-relation negative sampling to improve learning efficiency. Experimental results demonstrate that our approach achieves state-of-the-art performance in terms of mean reciprocal rank (MRR), with improvements of 2.5% on WN18RR and 21% on OpenBG500. Jiabang He, Lei Wang 0185, Xiyao Li, Xing Xu 0001 |
ICME | 3 |
| 2024 | Mitigating Fine-Grained Hallucination by Fine-Tuning Large Vision-Language Models with Caption Rewrites
Lei Wang 0185, Jiabang He, Shenshen Li, Ee-Peng Lim |
MMM (4) | 1 |
| 2024 | A Prompt-Based Topic-Modeling Method for Depression Detection on Low-Resource DataabstractDepression has a large impact on one’s personal life, especially during the COVID-19 pandemic. People have been trying to develop reliable methods for the depression detection task. Recently, methods based on deep learning have attracted much attention from the research community. However, they still face the challenge that data collection and annotation are difficult and expensive. In many real-world applications, only a small number of or even no training data are available. In this context, we propose a Prompt-based Topic-modeling method for Depression Detection (PTDD) on low-resource data, aiming to establish an effective way of depression detection under the above challenging situation. Instead of learning discriminating features from a small amount of labeled data, the proposed framework turns to leverage the generalization power of pretrained language models. Specifically, based on the question-and-answer routine during the interview, we first reorganize the text data according to the predefined topics for each interviewee. Via the prompt-based framework, we then predict whether the next-sentence prompt is emotionally positive or not. Finally, the depression detection task can be achieved based on the obtained topicwise predictions through a simple voting process. In the experiments, we validate the effectiveness of our model under several low-resource data settings. The results and analysis demonstrate that our PTDD achieves acceptable performance when only a few training samples or even no training samples are available. Yanrong Guo, Lei Wang 0185, Shijie Hao, Richang Hong |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Math Word Problem Generation via Disentangled Memory RetrievalabstractThe task of math word problem (MWP) generation, which generates an MWP given an equation and relevant topic words, has increasingly attracted researchers’ attention. In this work, we introduce a simple memory retrieval module to search related training MWPs, which are used to augment the generation. To retrieve more relevant training data, we also propose a disentangled memory retrieval module based on the simple memory retrieval module. To this end, we first disentangle the training MWPs into logical description and scenario description and then record them in respective memory modules. Later, we use the given equation and topic words as queries to retrieve relevant logical descriptions and scenario descriptions from the corresponding memory modules, respectively. The retrieved results are then used to complement the process of the MWP generation. Extensive experiments and ablation studies verify the superior performance of our method and the effectiveness of each proposed module. The code is available at https://github.com/mwp-g/MWPG-DMR . Zhenzhen Hu 0004, Lei Wang 0185, Yunshi Lan, Richang Hong |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | Generalizing Math Word Problem Solvers via Solution DiversificationabstractCurrent math word problem (MWP) solvers are usually Seq2Seq models trained by the (one-problem; one-solution) pairs, each of which is made of a problem description and a solution showing reasoning flow to get the correct answer. However, one MWP problem naturally has multiple solution equations. The training of an MWP solver with (one-problem; one-solution) pairs excludes other correct solutions, and thus limits the generalizability of the MWP solver. One feasible solution to this limitation is to augment multiple solutions to a given problem. However, it is difficult to collect diverse and accurate augment solutions through human efforts. In this paper, we design a new training framework for an MWP solver by introducing a solution buffer and a solution discriminator. The buffer includes solutions generated by an MWP solver to encourage the training data diversity. The discriminator controls the quality of buffered solutions to participate in training. Our framework is flexibly applicable to a wide setting of fully, semi-weakly and weakly supervised training for all Seq2Seq MWP solvers. We conduct extensive experiments on a benchmark dataset Math23k and a new dataset named Weak12k, and show that our framework improves the performance of various MWP solvers under different settings by generating correct and diverse solutions. Zhenwen Liang, Lei Wang 0185, Yan Wang 0060, Jie Shao 0001, Xiangliang Zhang 0001 |
AAAI | 3 |
| 2023 | Alignment-Enriched Tuning for Patch-Level Pre-trained Document Image ModelsabstractAlignment between image and text has shown promising improvements on patch-level pre-trained document image models. However, investigating more effective or finer-grained alignment techniques during pre-training requires a large amount of computation cost and time. Thus, a question naturally arises: Could we fine-tune the pre-trained models adaptive to downstream tasks with alignment objectives and achieve comparable or better performance? In this paper, we propose a new model architecture with alignment-enriched tuning (dubbed AETNet) upon pre-trained document image models, to adapt downstream tasks with the joint task-specific supervised and alignment-aware contrastive objective. Specifically, we introduce an extra visual transformer as the alignment-ware image encoder and an extra text transformer as the alignment-ware text encoder before multimodal fusion. We consider alignment in the following three aspects: 1) document-level alignment by leveraging the cross-modal and intra-modal contrastive loss; 2) global-local alignment for modeling localized and structural information in document images; and 3) local-level alignment for more accurate patch-level information. Experiments on various downstream tasks show that AETNet can achieve state-of-the-art performance on various downstream tasks. Notably, AETNet consistently outperforms state-of-the-art pre-trained models, such as LayoutLMv3 with fine-tuning techniques, on three different downstream tasks. Code is available at https://github.com/MAEHCM/AET. Lei Wang 0185, Jiabang He, Xing Xu 0001 |
AAAI | 1 |
| 2023 | Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language ModelsabstractLei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, Ee-Peng Lim. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Lei Wang 0185, Wanyu Xu, Yihuai Lan, Yunshi Lan, Roy Ka-Wei Lee, Ee-Peng Lim |
ACL (1) | 1 |
| 2023 | FlaCGEC: A Chinese Grammatical Error Correction Dataset with Fine-grained Linguistic AnnotationabstractChinese Grammatical Error Correction (CGEC) has been attracting growing attention from researchers recently. In spite of the fact that multiple CGEC datasets have been developed to support the research, these datasets lack the ability to provide a deep linguistic topology of grammar errors, which is critical for interpreting and diagnosing CGEC approaches. To address this limitation, we introduce FlaCGEC, which is a new CGEC dataset featured with fine-grained linguistic annotation. Specifically, we collect raw corpus from the linguistic schema defined by Chinese language experts, conduct edits on sentences via rules, and refine generated samples manually, which results in 10k sentences with 78 instantiated grammar points and 3 types of edits. We evaluate various cutting-edge CGEC methods on the proposed FlaCGEC dataset and their unremarkable results indicate that this dataset is challenging in covering a large range of grammatical errors. In addition, we also treat FlaCGEC as a diagnostic dataset for testing generalization skills and conduct a thorough evaluation of existing CGEC models. Hanyue Du, Yike Zhao, Qingyuan Tian, Lei Wang 0185, Yunshi Lan |
CIKM | 5 |
| 2023 | Non-Autoregressive Math Word Problem Solver with Unified Tree StructureabstractExisting MWP solvers employ sequence or binary tree to present the solution expression and decode it from given problem description.However, such structures fail to handle the variants that can be derived via mathematical manipulation, e.g., (a 1 + a 2 ) * a 3 and a 1 * a 3 +a 2 * a 3 can both be possible valid solutions for a same problem but formulated as different expression sequences or trees.The multiple solution variants depicting different possible solving procedures for the same input problem would raise two issues: 1) making it hard for the model to learn the mapping function between the input and output spaces effectively, and 2) wrongly indicating wrong when evaluating a valid expression variant.To address these issues, we introduce a unified tree structure to present a solution expression, where the elements are permutable and identical for all the expression variants.We propose a novel non-autoregressive solver, named MWP-NAS, to parse the problem and deduce the solution expression based on the unified tree.For evaluating the possible expression variants, we design a path-based metric to evaluate the partial accuracy of expressions of a unified tree.The results from extensive experiments conducted on Math23K and MAWPS demonstrate the effectiveness of our proposed MWP-NAS.The codes and checkpoints are available at: https: //github.com/mengqunhan/MWP-NAS. Yi Bin, Mengqun Han, Lei Wang 0185, Yang Yang 0002, See-Kiong Ng, Heng Tao Shen |
EMNLP | 4 |
| 2023 | LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language ModelsabstractThe success of large language models (LLMs), like GPT-4 and ChatGPT, has led to the development of numerous cost-effective and accessible alternatives that are created by finetuning open-access LLMs with task-specific data (e.g., ChatDoctor) or instruction data (e.g., Alpaca).Among the various fine-tuning methods, adapter-based parameter-efficient fine-tuning (PEFT) is undoubtedly one of the most attractive topics, as it only requires fine-tuning a few external parameters instead of the entire LLMs while achieving comparable or even better performance.To enable further research on PEFT methods of LLMs, this paper presents LLM-Adapters, an easy-to-use framework that integrates various adapters into LLMs and can execute these adapter-based PEFT methods of LLMs for different tasks.The framework includes state-of-the-art open-access LLMs such as LLaMA, BLOOM, and GPT-J, as well as widely used adapters such as Series adapters, Parallel adapter, Prompt-based learning and Reparametrization-based methods.Moreover, we conduct extensive empirical studies on the impact of adapter types, placement locations, and hyper-parameters to the best design for each adapter-based methods.We evaluate the effectiveness of the adapters on fourteen datasets from two different reasoning tasks, Arithmetic Reasoning and Commonsense Reasoning.The results demonstrate that using adapter-based PEFT in smaller-scale LLMs (7B) with few extra trainable parameters yields comparable, and in some cases superior, performance to powerful LLMs (175B) in zero-shot inference on both reasoning tasks.The code and datasets can be found in https://github. com/AGI-Edgerunners/LLM-Adapters. Lei Wang 0185, Yihuai Lan, Wanyu Xu, Ee-Peng Lim, Lidong Bing, Xing Xu 0001, Soujanya Poria, Roy Ka-Wei Lee |
EMNLP | 2 |
| 2023 | ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information ExtractionabstractLarge language models (LLMs), such as GPT-3 and ChatGPT, have demonstrated remarkable results in various natural language processing (NLP) tasks with in-context learning, which involves inference based on a few demonstration examples. Despite their successes in NLP tasks, no investigation has been conducted to assess the ability of LLMs to perform document information extraction (DIE) using in-context learning. Applying LLMs to DIE poses two challenges: the modality and task gap. To this end, we propose a simple but effective in-context learning framework called ICL-D3IE, which enables LLMs to perform DIE with different types of demonstration examples. Specifically, we extract the most difficult and distinct segments from hard training documents as hard demonstrations for benefiting all test instances. We design demonstrations describing relationships that enable LLMs to understand positional relationships. We introduce formatting demonstrations for easy answer extraction. Additionally, the framework improves diverse demonstrations by updating them iteratively. Our experiments on three widely used benchmark datasets demonstrate that the ICL-D3IE framework enables Davinci-003/ChatGPT to achieve superior performance when compared to previous pre-trained methods fine-tuned with full training in both the in-distribution (ID) setting and in the out-of-distribution (OOD) setting. Code is available at https://github.com/MAEHCM/ICL-D3IE. Jiabang He, Lei Wang 0185, Xing Xu 0001, Heng Tao Shen |
ICCV | 2 |
| 2023 | Do-GOOD: Towards Distribution Shift Evaluation for Pre-Trained Visual Document Understanding ModelsabstractNumerous pre-training techniques for visual document understanding (VDU) have recently shown substantial improvements in performance across a wide range of document tasks. However, these pre-trained VDU models cannot guarantee continued success when the distribution of test data differs from the distribution of training data. In this paper, to investigate how robust existing pre-trained VDU models are to various distribution shifts, we first develop an out-of-distribution (OOD) benchmark termed Do-GOOD for the fine-Grained analysis on Document image-related tasks specifically. The Do-GOOD benchmark defines the underlying mechanisms that result in different distribution shifts and contains 9 OOD datasets covering 3 VDU related tasks, e.g., document information extraction, classification and question answering. We then evaluate the robustness and perform a fine-grained analysis of 5 latest VDU pre-trained models and 2 typical OOD generalization algorithms on these OOD datasets. Results from the experiments demonstrate that there is a significant performance gap between the in-distribution (ID) and OOD settings for document images, and that fine-grained analysis of distribution shifts can reveal the brittle nature of existing pre-trained VDU models and OOD generalization algorithms. The code and datasets for our Do-GOOD benchmark can be found at https://github.com/MAEHCM/Do-GOOD. Jiabang He, Lei Wang 0185, Xing Xu 0001, Heng Tao Shen |
SIGIR | 3 |
| 2022 | MWPToolkit: An Open-Source Framework for Deep Learning-Based Math Word Problem SolversabstractWhile Math Word Problem (MWP) solving has emerged as a popular field of study and made great progress in recent years, most existing methods are benchmarked solely on one or two datasets and implemented with different configurations. In this paper, we introduce the first open-source library for solving MWPs called MWPToolkit, which provides a unified, comprehensive, and extensible framework for the research purpose. Specifically, we deploy 17 deep learning-based MWP solvers and 6 MWP datasets in our toolkit. These MWP solvers are advanced models for MWP solving, covering the categories of Seq2seq, Seq2Tree, Graph2Tree, and Pre-trained Language Models. And these MWP datasets are popular datasets that are commonly used as benchmarks in existing work. Our toolkit is featured with highly modularized and reusable components, which can help researchers quickly get started and develop their own models. We have released the code and documentation of MWPToolkit in https://github.com/LYH-YF/MWPToolkit. Yihuai Lan, Lei Wang 0185, Yunshi Lan, Bing Tian Dai, Yan Wang 0060, Dongxiang Zhang, Ee-Peng Lim |
AAAI | 2 |
| 2022 | Explanation Guided Contrastive Learning for Sequential RecommendationabstractRecently, contrastive learning has been applied to the sequential recommendation task to address data sparsity caused by users with few item interactions and items with few user adoptions. Nevertheless, the existing contrastive learning-based methods fail to ensure that the positive (or negative) sequence obtained by some random augmentation (or sequence sampling) on a given anchor user sequence remains to be semantically similar (or different). When the positive and negative sequences turn out to be false positive and false negative respectively, it may lead to degraded recommendation performance. In this work, we address the above problem by proposing Explanation Guided Augmentations (EGA) and Explanation Guided Contrastive Learning for Sequential Recommendation (EC4SRec) model framework. The key idea behind EGA is to utilize explanation method(s) to determine items' importance in a user sequence and derive the positive and negative sequences accordingly. EC4SRec then combines both self-supervised and supervised contrastive learning over the positive and negative sequences generated by EGA operations to improve sequence representation learning for more accurate recommendation results. Extensive experiments on four real-world benchmark datasets demonstrate that EC4SRec outperforms the state-of-the-art sequential recommendation methods and two recent contrastive learning-based sequential recommendation methods, CL4SRec and DuoRec. Our experiments also show that EC4SRec can be easily adapted for different sequence encoder backbones (e.g., GRU4Rec and Caser), and improve their recommendation performance. Lei Wang 0185, Ee-Peng Lim, Zhiwei Liu 0001, Tianxiang Zhao 0001 |
CIKM | 1 |
| 2022 | Mitigating Popularity Bias in Recommendation with Unbalanced Interactions: A Gradient PerspectiveabstractRecommender systems learn from historical user-item interactions to identify preferred items for target users. These observed interactions are usually unbalanced following a long-tailed distribution. Such long-tailed data lead to popularity bias to recommend popular but not personalized items to users. We present a gradient perspective to understand two negative impacts of popularity bias in recommendation model optimization: (i) the gradient direction of popular item embeddings is closer to that of positive interactions, and (ii) the magnitude of positive gradient for popular items are much greater than that of unpopular items. To address these issues, we propose a simple yet efficient framework to mitigate popularity bias from a gradient perspective. Specifically, we first normalize each user embedding and record accumulated gradients of users and items via popularity bias measures in model training. To address the popularity bias issues, we develop a gradient-based embedding adjustment approach used in model testing. This strategy is generic, model-agnostic, and can be seamlessly integrated into most existing recommender systems. Our extensive experiments on two classic recommendation models and four real-world datasets demonstrate the effectiveness of our method over state-of-the-art debiasing baselines. Weijieying Ren, Lei Wang 0185, Kunpeng Liu 0001, Ruocheng Guo, Ee-Peng Lim, Yanjie Fu |
ICDM | 2 |
| 2022 | Math Word Problem Generation with Memory Retrieval
Zhenzhen Hu 0004, Lei Wang 0185, Yunshi Lan, Richang Hong |
PRCV (3) | 4 |
| 2020 | Graph-to-Tree Learning for Solving Math Word ProblemsabstractWhile the recent tree-based neural models have demonstrated promising results in generating solution expression for the math word problem (MWP), most of these models do not capture the relationships and order information among the quantities well.This results in poor quantity representations and incorrect solution expressions.In this paper, we propose Graph2Tree, a novel deep learning architecture that combines the merits of the graph-based encoder and tree-based decoder to generate better solution expressions.Included in our Graph2Tree framework are two graphs, namely the Quantity Cell Graph and Quantity Comparison Graph, which are designed to address limitations of existing methods by effectively representing the relationships and order information among the quantities in MWPs.We conduct extensive experiments on two available datasets.Our experiment results show that Graph2Tree outperforms the state-of-the-art baselines on two benchmark datasets significantly.We also discuss case studies and empirically examine Graph2Tree's effectiveness in translating the MWP text into solution expressions 1 . Lei Wang 0185, Roy Ka-Wei Lee, Yi Bin, Yan Wang 0060, Jie Shao 0001, Ee-Peng Lim |
ACL | 2 |
| 2020 | Next-Term Grade Prediction: A Machine Learning Approach
Audrey Tedja Widjaja, Lei Wang 0185, Nghia Trong Truong, Aldy Gunawan, Ee-Peng Lim |
EDM | 2 |
| 2020 | Teacher-Student Networks with Multiple Decoders for Solving Math Word ProblemabstractMath word problem (MWP) is challenging due to the limitation in training data where only one “standard” solution is available. MWP models often simply fit this solution rather than truly understand or solve the problem. The generalization of models (to diverse word scenarios) is thus limited. To address this problem, this paper proposes a novel approach, TSN-MD, by leveraging the teacher network to integrate the knowledge of equivalent solution expressions and then to regularize the learning behavior of the student network. In addition, we introduce the multiple-decoder student network to generate multiple candidate solution expressions by which the final answer is voted. In experiments, we conduct extensive comparisons and ablative studies on two large-scale MWP benchmarks, and show that using TSN-MD can surpass the state-of-the-art works by a large margin. More intriguingly, the visualization results demonstrate that TSN-MD not only produces correct final answers but also generates diverse equivalent expressions of the solution. Roy Ka-Wei Lee, Ee-Peng Lim, Lei Wang 0185, Jie Shao 0001, Qianru Sun |
IJCAI | 5 |
| 2020 | The Gap of Semantic Parsing: A Survey on Automatic Math Word Problem SolversabstractSolving mathematical word problems (MWPs) automatically is challenging, primarily due to the semantic gap between human-readable words and machine-understandable logics. Despite the long history dated back to the 1960s, MWPs have regained intensive attention in the past few years with the advancement of Artificial Intelligence (AI). Solving MWPs successfully is considered as a milestone towards general AI. Many systems have claimed promising results in self-crafted and small-scale datasets. However, when applied on large and diverse datasets, none of the proposed methods in the literature achieves high precision, revealing that current MWP solvers still have much room for improvement. This motivated us to present a comprehensive survey to deliver a clear and complete picture of automatic math problem solvers. In this survey, we emphasize on algebraic word problems, summarize their extracted features and proposed techniques to bridge the semantic gap, and compare their performance in the publicly accessible datasets. We also cover automatic solvers for other types of math problems such as geometric problems that require the understanding of diagrams. Finally, we identify several emerging research directions for the readers with interests in MWPs. Dongxiang Zhang, Lei Wang 0185, Bing Tian Dai, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Template-Based Math Word Problem Solvers with Recursive Neural NetworksabstractThe design of automatic solvers to arithmetic math word problems has attracted considerable attention in recent years and a large number of datasets and methods have been published. Among them, Math23K is the largest data corpus that is very helpful to evaluate the generality and robustness of a proposed solution. The best performer in Math23K is a seq2seq model based on LSTM to generate the math expression. However, the model suffers from performance degradation in large space of target expressions. In this paper, we propose a template-based solution based on recursive neural network for math expression construction. More specifically, we first apply a seq2seq model to predict a tree-structure template, with inferred numbers as leaf nodes and unknown operators as inner nodes. Then, we design a recursive neural network to encode the quantity with Bi-LSTM and self attention, and infer the unknown operator nodes in a bottom-up manner. The experimental results clearly establish the superiority of our new framework as we improve the accuracy by a wide margin in two of the largest datasets, i.e., from 58.1% to 66.9% in Math23K and from 62.8% to 66.8% in MAWPS. Lei Wang 0185, Dongxiang Zhang, Xing Xu 0001, Lianli Gao, Bing Tian Dai, Heng Tao Shen |
AAAI | 1 |
| 2019 | Modeling Intra-Relation in Math Word Problems with Different Functional Multi-Head AttentionsabstractSeveral deep learning models have been proposed for solving math word problems (MWPs) automatically.Although these models have the ability to capture features without manual efforts, their approaches to capturing features are not specifically designed for MWPs.To utilize the merits of deep learning models with simultaneous consideration of MWPs' specific features, we propose a group attention mechanism to extract global features, quantity-related features, quantitypair features and question-related features in MWPs respectively.The experimental results show that the proposed approach performs significantly better than previous state-of-the-art methods, and boost performance from 66.9% to 69.5% on Math23K with training-test split, from 65.8% to 66.9% on Math23K with 5-fold cross-validation and from 69.2% to 76.1% on MAWPS. Jierui Li, Lei Wang 0185, Yan Wang 0060, Bing Tian Dai, Dongxiang Zhang |
ACL (1) | 2 |
| 2018 | MathDQN: Solving Arithmetic Word Problems via Deep Reinforcement LearningabstractDesigning an automatic solver for math word problems has been considered as a crucial step towards general AI, with the ability of natural language understanding and logical inference. The state-of-the-art performance was achieved by enumerating all the possible expressions from the quantities in the text and customizing a scoring function to identify the one with the maximum probability. However, it incurs exponential search space with the number of quantities and beam search has to be applied to trade accuracy for efficiency. In this paper, we make the first attempt of applying deep reinforcement learning to solve arithmetic word problems. The motivation is that deep Q-network has witnessed success in solving various problems with big search space and achieves promising performance in terms of both accuracy and running time. To fit the math problem scenario, we propose our MathDQN that is customized from the general deep reinforcement learning framework. Technically, we design the states, actions, reward function, together with a feed-forward neural network as the deep Q-network. Extensive experimental results validate our superiority over state-of-the-art methods. Our MathDQN yields remarkable improvement on most of datasets and boosts the average precision among all the benchmark datasets by 15\%. Lei Wang 0185, Dongxiang Zhang, Lianli Gao, Jingkuan Song, Long Guo, Heng Tao Shen |
AAAI | 1 |
| 2018 | Translating Math Word Problem to Expression TreeabstractSequence-to-sequence (SEQ2SEQ) models have been successfully applied to automatic math word problem solving.Despite its simplicity, a drawback still remains: a math word problem can be correctly solved by more than one equations.This non-deterministic transduction harms the performance of maximum likelihood estimation.In this paper, by considering the uniqueness of expression tree, we propose an equation normalization method to normalize the duplicated equations.Moreover, we analyze the performance of three popular SEQ2SEQ models on the math word problem solving.We find that each model has its own specialty in solving problems, consequently an ensemble model is then proposed to combine their advantages.Experiments on dataset Math23K show that the ensemble model with equation normalization significantly outperforms the previous state-of-the-art methods. Lei Wang 0185, Yan Wang 0060, Deng Cai 0002, Dongxiang Zhang, Xiaojiang Liu |
EMNLP | 1 |