Wei Liu 0144

dblp:49/3283-144 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0002-3871-9454ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Time-Interval-Aware Disentangled Expert Modeling for Next-Basket Recommendation
abstract
Next-basket recommendation (NBR) is a type of recommendation that aims to predict a set of items a user will purchase based on their historical transaction basket sequences. It is governed by a dynamic interplay between two distinct user intents: habitual repurchase, which involves repeating past behaviors, and exploratory interest, which involves discovering new items. However, existing NBR methods generally suffer from two limitations: (1) they often entangle these conflicting motives within a single representation, causing habits to overshadow discovery, and (2) they rely on discrete sequential modeling that ignores continuous-time intervals and item-specific periodicities. In this paper, we propose a novel solution named Time-Interval Disentangled Experts (TIDE) to address these challenges. TIDE incorporates a Hawkes-enhanced Fourier Time Encoding to capture item-specific temporal periodicities and dynamic decay. To decouple user intentions, TIDE utilizes a dual-expert architecture that integrates a Habit Expert for recurring needs and a Pattern-Guided Exploration Expert for discovery. Combined with an item-aware gating mechanism, TIDE adaptively balances repurchase and exploration. Extensive experiments on four diverse real-world datasets demonstrate that TIDE consistently outperforms representative state-of-the-art NBR methods.
Zhiying Deng, Usman Farooq, Wei Liu 0144, Jianjun Li 0010
SIGIR5
2026 Confounder Balance in Next Basket Prediction
abstract
Next basket prediction (NBP) is essential for online businesses such as grocery shopping and retail, as it aims to learn users' interests based on historical basket sequences and predict a set of items for the next purchase. Existing methods tend to prioritize items with high interaction frequencies in the dataset. In this article, we use causal analysis to theoretically demonstrate that such methods can lead to biased results. Specifically, item interactions in the dataset represent item-specific weights that reflect average user preferences. However, these weights are imprecise due to varying user interest levels, leading to biased learning results and suboptimal predictions. We find that repeated interactions between a user and items can represent the user's level of interest, which can be leveraged to improve user interest modeling by assigning user-specific weights. In addition, different users exhibit varying preferences for items with high interaction frequencies, highlighting the necessity of different degrees of bias mitigation. Making personalized adjustments based on these differences can further refine existing methods. Consequently, we propose a simple yet effective confounder balance prediction (CBP) model to mitigate bias while preserving individual user interests. Specifically, we employ counterfactual inference to construct a counterfactual world in which predictions are influenced solely by either high interaction frequencies or user interests, both of which are potential confounders. This approach enables us to individually assess the impact of these confounders on interaction probabilities. The goal of CBP is to balance the confounders and thereby refine the learning process without resorting to oversimplified unbiased learning. Experiments on four widely used real-world datasets demonstrate the significant advantages of CBP over existing state-of-the-art methods. Our code is available at https://github.com/Hiiizhy/CBP.
Zhiying Deng, Jianjun Li 0010, Wei Liu 0144
IEEE Trans. Cybern.3
2025 Breaking Free from MMI: A New Frontier in Rationalization by Probing Input Utilization
abstract
Extracting a small subset of crucial rationales from the full input is a key problem in explainability research. The most widely used fundamental criterion for rationale extraction is the maximum mutual information (MMI) criterion. In this paper, we first demonstrate that MMI suffers from diminishing marginal returns. Once part of the rationale has been identified, finding the remaining portions contributes only marginally to increasing the mutual information, making it difficult to use MMI to locate the rest. In contrast to MMI that aims to reproduce the prediction, we seek to identify the parts of the input that the network can actually utilize. This is achieved by comparing how different rationale candidates match the capability space of the weight matrix. The weight matrix of a neural network is typically low-rank, meaning that the linear combinations of its column vectors can only cover part of the directions in a high-dimensional space (high-dimension: the dimensions of an input vector). If an input is fully utilized by the network, it generally matches these directions (e.g., a portion of a hypersphere), resulting in a representation with a high norm. Conversely, if an input primarily falls outside (orthogonal to) these directions, its representation norm will approach zero, behaving like noise that the network cannot effectively utilize. Building on this, we propose using the norms of rationale candidates as an alternative objective to MMI. Through experiments on four text classification datasets and one graph classification dataset using three network architectures (GRUs, BERT, and GCN), we show that our method outperforms MMI and its improved variants in identifying better rationales. We also compare our method with a representative LLM (llama-3.1-8b-instruct) and find that our simple method gets comparable results to it and can sometimes even outperform it.
Wei Liu 0144, Zhiying Deng, Zhongyu Niu, Haozhao Wang, Zhigang Zeng, Ruixuan Li 0001
ICLR1
2025 Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean Datasets
abstract
This study investigates the self-rationalization framework constructed with a cooperative game, where a generator initially extracts the most informative segment from raw input, and a subsequent predictor utilizes the selected subset for its input. The generator and predictor are trained collaboratively to maximize prediction accuracy. In this paper, we first uncover a potential caveat: such a cooperative game could unintentionally introduce a sampling bias during rationale extraction. Specifically, the generator might inadvertently create an incorrect correlation between the selected rationale candidate and the label, even when they are semantically unrelated in the original dataset. Subsequently, we elucidate the origins of this bias using both detailed theoretical analysis and empirical evidence. Our findings suggest a direction for inspecting these correlations through attacks, based on which we further introduce an instruction to prevent the predictor from learning the correlations. Through experiments on six text classification datasets and two graph classification datasets using three network architectures (GRUs, BERT, and GCN), we show that our method significantly outperforms recent rationalization methods.
Wei Liu 0144, Zhongyu Niu, Lang Gao, Zhiying Deng, Jun Wang 0018, Haozhao Wang, Ruixuan Li 0001
ICML1
2025 Quantifying Distributional Invariance in Causal Subgraph for IRM-Free Graph Generalization
abstract
Out-of-distribution generalization under distributional shifts remains a critical challenge for graph neural networks. Existing methods generally adopt the Invariant Risk Minimization (IRM) framework, requiring costly environment annotations or heuristically generated synthetic splits. To circumvent these limitations, in this work, we aim to develop an IRM-free method for capturing causal subgraphs. We first identify that causal subgraphs exhibit substantially smaller distributional variations than non-causal components across diverse environments, which we formalize as the Invariant Distribution Criterion and theoretically prove in this paper. Building on this criterion, we systematically uncover the quantitative relationship between distributional shift and representation norm for identifying the causal subgraph, and investigate its underlying mechanisms in depth. Finally, we propose an IRM-free method by introducing a norm-guided invariant distribution objective for causal subgraph discovery and prediction. Extensive experiments on two widely used benchmarks demonstrate that our method consistently outperforms state-of-the-art methods in graph generalization. Code is available at https://github.com/anders1123/IDG.
Yixiong Zou, Jun Wang 0018, Wei Liu 0144, Xiangyu Fu, Ruixuan Li 0001
NeurIPS4
2025 Privacy-Friendly Cross-Domain Recommendation via Distilling User-irrelevant Information
abstract
Privacy-preserving Cross-Domain Recommendation (CDR) has been extensively studied to address the cold-start problem using auxiliary source domains while simultaneously protecting sensitive information. However, existing privacy-preserving CDR methods rely heavily on transferring sensitive user embeddings or behaviour logs, which leads to adopt privacy methods to distort the data patterns before transferring it to the target domain. The distorted information can compromise overall performance during the knowledge transfer process. To overcome these challenges, our approach differs from existing privacy-preserving methods that focus on safeguarding user-sensitive information. Instead, we concentrate on distilling transferable knowledge from insensitive item embeddings, which we refer to as prototypes. Specifically, we propose a conditional model inversion mechanism to accurately distill prototypes for individual users. We have designed a new data format and corresponding learning paradigm for distilling transferable prototypes from traditional recommendation models using model inversion. These prototypes facilitate bridging the domain shift between distinct source and target domains in a privacy-friendly manner. Additionally, they enable the identification of top-k users in the target domain to substitute for cold-start users prediction. We conduct extensive experiments across large real-world datasets, and the results substantiate the effectiveness of PFCDR https://github.com/walcheng/PFCDR.
Cheng Wang 0025, Wenchao Xu 0001, Haozhao Wang, Wei Liu 0144, Ruixuan Li 0001
WWW4
2025 Re-Fed+: A Better Replay Strategy for Federated Incremental Learning
abstract
Federated learning (FL) has emerged as a significant distributed machine learning paradigm. It allows the training of a global model through user collaboration without the necessity of sharing their original data. Traditional FL generally assumes that each client's data remains fixed or static. However, in real-world scenarios, data typically arrives incrementally, leading to a dynamically expanding data domain. In this study, we examine catastrophic forgetting within Federated Incremental Learning (FIL) and focus on the training resources, where edge clients may not have sufficient storage to keep all data or computational budget to implement complex algorithms designed for the server-based environment. We propose a general and low-cost framework for FIL named Re-Fed+, which is designed to help clients cache important samples for replay. Specifically, when a new task arrives, each client initially caches selected previous samples based on their global and local significance. The client then trains the local model using both the cached samples and the new task samples. From a theoretical perspective, we analyze how effectively Re-Fed+ can identify significant samples for replay to alleviate the catastrophic forgetting issue. Empirically, we show that Re-Fed+ achieves competitive performance compared to state-of-the-art methods.
Yichen Li 0006, Haozhao Wang, Yining Qi, Wei Liu 0144, Ruixuan Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Exploring Practical Gaps in Using Cross Entropy to Implement Maximum Mutual Information Criterion for Rationalization
abstract
Abstract Rationalization is a framework that aims to build self-explanatory NLP models by extracting a subset of human-intelligible pieces of their inputting texts. It involves a cooperative game where a selector selects the most human-intelligible parts of the input as the rationale, followed by a predictor that makes predictions based on these selected rationales. Existing literature uses the cross-entropy between the model’s predictions and the ground-truth labels to measure the informativeness of the selected rationales, guiding the selector to choose better ones. In this study, we first theoretically analyze the objective of rationalization by decomposing it into two parts: the model-agnostic informativeness of the rationale candidates and the predictor’s degree of fit. We then provide various empirical evidence to support that, under this framework, the selector tends to sample from a limited small region, causing the predictor to overfit these localized areas. This results in a significant mismatch between the cross-entropy objective and the informativeness of the rationale candidates, leading to suboptimal solutions. To address this issue, we propose a simple yet effective method that introduces random vicinal1 perturbations to the selected rationale candidates. This approach broadens the predictor’s assessment to a vicinity around the selected rationale candidate. Compared to recent competitive methods, our method significantly improves rationale quality (by up to 6.6%) across six widely used classification datasets. The term “vicinal” is borrowed from vicinal risk minimization (Chapelle et al., 2000); “vicinal” means neighboring or adjacent.
Wei Liu 0144, Zhiying Deng, Zhongyu Niu, Jun Wang 0018, Haozhao Wang, Ruixuan Li 0001
Trans. Assoc. Comput. Linguistics1
2025 Unbiased Interest Modeling in Sequential Basket Analysis: Addressing Repetition Bias with Multi-Factor Estimation
abstract
Sequential basket analysis is a challenging task that focuses on modeling user interests through their shopping basket records. This study focuses on a newly identified bias: the repetition bias , which typically arises due to repurchase behavior . Existing methods typically oversimplify the relationship between repetitions and predictions. They assume that frequent repetition of an item by a user indicates a strong preference of the user. However, this assumption is flawed as it fails to consider that repetitions are not driven solely by user interests, as they can also be influenced by external factors, resulting in a biased understanding of user interests. In this article, we propose the CA usal intervention for R epetition D e-biasing ( CARD ), a novel solution to comprehensively estimate various influencing factors and address the repetition bias, thereby ensuring a more accurate learning of user interests. Specifically, we design a multi-factor estimation debiasing framework with constructed causal graphs to formalize the data generation process within the recommendation. We then analyze the variables that influence the recommendation, with the goal of identifying confounding variables that affect repurchase behavior and thereby locating the source of repetition bias. Since repetition bias originates from the influence of confounding variables on repurchase behavior, we resort to causal intervention methods to prevent its impacts and thus eliminate repetition bias at its source for unbiased user interest modeling. We evaluate CARD by conducting extensive experiments over three real-world datasets. The results demonstrate our approach’s competitiveness over the representative state-of-the-art baselines in achieving unbiased user interest modeling.
Zhiying Deng, Jianjun Li 0010, Wei Liu 0144
Trans. Recomm. Syst.3
2024 Decoupling Representation and Knowledge for Few-Shot Intent Classification and Slot Filling
abstract
Few-shot intent classification and slot filling are important but challenging tasks due to the scarcity of finely labeled data. Therefore, current works first train a model on source domains with sufficiently labeled data, and then transfer the model to target domains where only rarely labeled data is available. However, experience transferring as a whole usually suffers from gaps that exist among source domains and target domains. For instance, transferring domain-specific-knowledge-related experience is difficult. To tackle this problem, we propose a new method that explicitly decouples the transferring of general-semantic-representation-related experience and the domain-specific-knowledge-related experience. Specifically, for domain-specific-knowledge-related experience, we design two modules to capture intent-slot relation and slot-slot relation respectively. Extensive experiments on Snips and FewJoint datasets show that our method achieves state-of-the-art performance. The method improves the joint accuracy metric from 27.72% to 42.20% in the 1-shot setting, and from 46.54% to 60.79% in the 5-shot setting.
Yixiong Zou, Haozhao Wang, Jun Wang 0018, Wei Liu 0144, Ruixuan Li 0001
AAAI5
2024 Multi-scale Context-aware User Interest Learning for Behavior Pattern Modeling
Zhiying Deng, Jianjun Li 0010, Wei Liu 0144, Guohui Li 0001
DASFAA (3)4
2024 Adversarial Attack for Explanation Robustness of Rationalization Models
abstract
Rationalization models, which select a subset of input text as rationale—crucial for humans to understand and trust predictions—have recently emerged as a prominent research area in eXplainable Artificial Intelligence (XAI). However, most of previous studies mainly focus on improving the quality of the rationale, ignoring its robustness to malicious attack. Specifically, whether the rationalization models can still generate high-quality rationale under the adversarial attack remains unknown. To explore this, this paper proposes UAT2E, which aims to undermine the explainability of rationalization models without altering their predictions, thereby eliciting distrust in these models from human users. UAT2E employs the gradient-based search on triggers and then inserts them into the original input to conduct both the non-target and target attack. Experimental results on five datasets reveal the vulnerability of rationalization models in terms of explanation, where they tend to select more meaningless tokens under attacks. Based on this, we make a series of recommendations for improving rationalization models in terms of explanation.
Yuankai Zhang 0002, Lingxiao Kong, Haozhao Wang, Ruixuan Li 0001, Jun Wang 0018, Yuhua Li 0003, Wei Liu 0144
ECAI7
2024 PAGE: Parametric Generative Explainer for Graph Neural Network
abstract
This article introduces PAGE, a parameterized generative interpretive framework. PAGE is capable of providing faithful explanations for any graph neural network without necessitating prior knowledge or internal details. Specifically, we train the autoencoder to generate explanatory substructures by designing appropriate training strategy. Due to the dimensionality reduction of features in the latent space of the autoencoder, it becomes easier to extract causal features leading to the model’s output, which can be easily employed to generate explanations. To accomplish this, we introduce an additional discriminator to capture the causality between latent causal features and the model’s output. By designing appropriate optimization objectives, the well-trained discriminator can be employed to constrain the encoder in generating enhanced causal features. Finally, these features are mapped to substructures of the input graph through the decoder to serve as explanations. Compared to existing methods, PAGE operates at the sample scale rather than nodes or edges, eliminating the need for perturbation or encoding processes as seen in previous methods. Experimental results on both artificially synthesized and real-world datasets demonstrate that our approach not only exhibits the highest faithfulness and accuracy but also significantly outperforms baseline models in terms of efficiency.
Wei Liu 0144, Jun Wang 0018, Ruixuan Li 0001
ECAI2
2024 AWSSS: Adaptive Weighted Statistical Space Smoothing for Regression with Imbalance Data
abstract
Deep learning models, usually trained on large datasets, have witnessed great achievements for both classification and regression tasks during past years. However, in practical settings, the datasets are commonly imbalanced where the number of data samples differs among labels (i.e.,categories or target values), leading to serious performance degradation of the trained models. Although many prior works have been proposed to solve the data imbalance problem for the classification task, there are still limited works considering the regression task. In this work, we identify the distinct challenges in regression problems compared to traditional classification problems, such as managing the continuous nature of the target variable and navigating fuzzy decision-making boundaries. To address these challenges, we propose an Adaptive Weighted Statistical Space Smoothing method (AWSSS), which alleviates the impact of imbalanced data by learning from continuously valued, imbalanced data and extending local similarity across the entire target range. AWSSS extracts statistical values from the original dataset and smoothes statistical features. Subsequently, it aligns features across different regions and applies adaptive weights to facilitate knowledge transfer between adjacent areas. Finally, AWSSS calculates the loss values, which are used to refine the model’s parameters through iterative updates. Extensive experiments conducted on various datasets demonstrate the effectiveness of the proposed method as compared to state-of-the-art methods.
Xiaoquan Yi, Haozhao Wang, Zhenlong Zhu, Wei Liu 0144, Wenchao Xu 0001, Ruixuan Li 0001
HPCC4
2024 Enhancing the Rationale-Input Alignment for Self-explaining Rationalization
abstract
Rationalization empowers deep learning models with self-explaining capabilities through a cooperative game, where a generator selects a semantically consistent subset of the input as a rationale, and a subsequent predictor makes predictions based on the selected rationale. In this paper, we discover that rationalization is prone to a problem named rationale shift, which arises from the algorithmic bias of the cooperative game. Rationale shift refers to a situation where the semantics of the selected rationale may deviate from the original input, but the predictor still produces accurate predictions based on the deviation, resulting in a compromised generator with misleading feedback. To address this issue, we first demonstrate the importance of the alignment between the rationale and the full input through both empirical observations and theoretical analysis. Subsequently, we introduce a novel approach called DAR (Discriminatively Aligned Rationalization), which utilizes an auxiliary module pretrained on the full input to discriminatively align the selected rationale and the original input. We theoretically illustrate how DAR accomplishes the desired alignment, thereby overcoming the rationale shift problem. The experiments on two widely used real-world benchmarks show that the proposed method significantly improves the explanation quality (measured by the overlap between the model-selected explanation and the human-annotated rationale) as compared to state-of-the-art techniques. Additionally, results on two synthetic settings further validate the effectiveness of DAR in addressing the rationale shift problem.
Wei Liu 0144, Haozhao Wang, Jun Wang 0018, Zhiying Deng, Yuankai Zhang 0002, Cheng Wang 0025, Ruixuan Li 0001
ICDE1
2024 Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-Rationalization
abstract
An important line of research in the field of explainability is to extract a small subset of crucial rationales from the full input. The most widely used criterion for rationale extraction is the maximum mutual information (MMI) criterion. However, in certain datasets, there are spurious features non-causally correlated with the label and also get high mutual information, complicating the loss landscape of MMI. Although some penalty-based methods have been developed to penalize the spurious features (e.g., invariance penalty, intervention penalty, etc) to help MMI work better, these are merely remedial measures. In the optimization objectives of these methods, spurious features are still distinguished from plain noise, which hinders the discovery of causal rationales. This paper aims to develop a new criterion that treats spurious features as plain noise, allowing the model to work on datasets rich in spurious features as if it were working on clean datasets, thereby making rationale extraction easier. We theoretically observe that removing either plain noise or spurious features from the input does not alter the conditional distribution of the remaining components relative to the task label. However, significant changes in the conditional distribution occur only when causal features are eliminated. Based on this discovery, the paper proposes a criterion for \textbf{M}aximizing the \textbf{R}emaining \textbf{D}iscrepancy (MRD). Experiments on six widely used datasets show that our MRD criterion improves rationale quality (measured by the overlap with human-annotated rationales) by up to $10.4\%$ as compared to several recent competitive MMI variants. Code: \url{https://github.com/jugechengzi/Rationalization-MRD}.
Wei Liu 0144, Zhiying Deng, Zhongyu Niu, Jun Wang 0018, Haozhao Wang, YuanKai Zhang, Ruixuan Li 0001
NeurIPS1
2023 MGR: Multi-generator Based Rationalization
abstract
Rationalization is to employ a generator and a predictor to construct a self-explaining NLP model in which the generator selects a subset of human-intelligible pieces of the input text to the following predictor.However, rationalization suffers from two key challenges, i.e., spurious correlation and degeneration, where the predictor overfits the spurious or meaningless pieces solely selected by the not-yet well-trained generator and in turn deteriorates the generator.Although many studies have been proposed to address the two challenges, they are usually designed separately and do not take both of them into account.In this paper, we propose a simple yet effective method named MGR to simultaneously solve the two problems.The key idea of MGR is to employ multiple generators such that the occurrence stability of real pieces is improved and more meaningful pieces are delivered to the predictor.Empirically 1 , we show that MGR improves the F1 score by up to 20.9% as compared to state-of-the-art methods.
Wei Liu 0144, Haozhao Wang, Jun Wang 0018, Ruixuan Li 0001, Yuankai Zhang 0002
ACL (1)1
2023 Decoupled Rationalization with Asymmetric Learning Rates: A Flexible Lipschitz Restraint
abstract
A self-explaining rationalization model is generally constructed by a cooperative game where a generator selects the most human-intelligible pieces from the input text as rationales, followed by a predictor that makes predictions based on the selected rationales. However, such a cooperative game may incur the degeneration problem where the predictor overfits to the uninformative pieces generated by a not yet well-trained generator and in turn, leads the generator to converge to a sub-optimal model that tends to select senseless pieces. In this paper, we theoretically bridge degeneration with the predictor's Lipschitz continuity. Then, we empirically propose a simple but effective method named DR, which can naturally and flexibly restrain the Lipschitz constant of the predictor, to address the problem of degeneration. The main idea of DR is to decouple the generator and predictor to allocate them with asymmetric learning rates. A series of experiments conducted on two widely used benchmarks have verified the effectiveness of the proposed method. Codes: https://github.com/jugechengzi/Rationalization-DR.
Wei Liu 0144, Jun Wang 0018, Haozhao Wang, Ruixuan Li 0001, Yuankai Zhang 0002, Yixiong Zou
KDD1
2023 D-Separation for Causal Self-Explanation
abstract
Rationalization aims to strengthen the interpretability of NLP models by extracting a subset of human-intelligible pieces of their inputting texts. Conventional works generally employ the maximum mutual information (MMI) criterion to find the rationale that is most indicative of the target label. However, this criterion can be influenced by spurious features that correlate with the causal rationale or the target label. Instead of attempting to rectify the issues of the MMI criterion, we propose a novel criterion to uncover the causal rationale, termed the Minimum Conditional Dependence (MCD) criterion, which is grounded on our finding that the non-causal features and the target label are \emph{d-separated} by the causal rationale. By minimizing the dependence between the non-selected parts of the input and the target label conditioned on the selected rationale candidate, all the causes of the label are compelled to be selected. In this study, we employ a simple and practical measure for dependence, specifically the KL-divergence, to validate our proposed MCD criterion. Empirically, we demonstrate that MCD improves the F1 score by up to 13.7% compared to previous state-of-the-art MMI-based methods. Our code is in an anonymous repository: https://anonymous.4open.science/r/MCD-CE88.
Wei Liu 0144, Jun Wang 0018, Haozhao Wang, Ruixuan Li 0001, Zhiying Deng, Yuankai Zhang 0002
NeurIPS1
2023 Multi-view Multi-aspect Neural Networks for Next-basket Recommendation
abstract
Next-basket recommendation (NBR) is a type of recommendation that aims to recommend a set of items to users according to their historical basket sequences. Existing NBR methods suffer from two limitations: (1) overlooking low-level item correlations, which results in coarse-grained item representation; and (2) failing to consider spurious interests in repeated behaviors, leading to suboptimal user interest learning. To address these limitations, we propose a novel solution named Multi-view Multi-aspect Neural Recommendation (MMNR) for NBR, which first normalizes the interactions from both the user-side and item-side, respectively, aiming to remove the spurious interests, and utilizes them as weights for items from different views to construct differentiated representations for each interaction item, enabling comprehensive user interest learning. Then, to capture low-level item correlations, MMNR models different aspects of items to obtain disentangled representations of items, thereby fully capturing multiple user interests. Extensive experiments on real-world datasets demonstrate the effectiveness of MMNR, showing that it consistently outperforms several state-of-the-art NBR methods.
Zhiying Deng, Jianjun Li 0010, Zhiqiang Guo, Wei Liu 0144, Guohui Li 0001
SIGIR4
2022 FR: Folded Rationalization with a Unified Encoder
abstract
Rationalization aims to strengthen the interpretability of NLP models by extracting a subset of human-intelligible pieces of their inputting texts. Conventional works generally employ a two-phase model in which a generator selects the most important pieces, followed by a predictor that makes predictions based on the selected pieces. However, such a two-phase model may incur the degeneration problem where the predictor overfits to the noise generated by a not yet well-trained generator and in turn, leads the generator to converge to a suboptimal model that tends to select senseless pieces. To tackle this challenge, we propose Folded Rationalization (FR) that folds the two phases of the rationale model into one from the perspective of text semantic extraction. The key idea of FR is to employ a unified encoder between the generator and predictor, based on which FR can facilitate a better predictor by access to valuable information blocked by the generator in the traditional two-phase model and thus bring a better generator. Empirically, we show that FR improves the F1 score by up to 10.3% as compared to state-of-the-art methods.
Wei Liu 0144, Haozhao Wang, Jun Wang 0018, Ruixuan Li 0001, Yuankai Zhang 0002
NeurIPS1