Xin Wen 0017

dblp:42/4185-17 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-8379-4149ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Disentangled Denoising and Counterfactual Balance for Multimodal Recommendation
abstract
Recently, graph convolutional network-based dual-view multimodal recommendation methods have achieved great success. They extract multimodal and behavior features based on item-item and user-item graphs, respectively. However, they still have two- fold limitations. First, the relevance between multimodal semantics and user preferences is ignored, resulting in the propagation and coupling of preference-irrelevant noise. Second, the direct use of uneven factual user-item graphs is suboptimal, as both redundant noisy edges and missing positive interaction edges impair recommendations. To solve the above issues, we propose aDisentAngled deNoising andCounterfactual balancEmethod for multimodal recommendation, dubbed asDANCE. Specifically, for multimodal features, we explicitly disentangle them into preference-relevant and preference-irrelevant representations, to absorb and discard irrelevant noise via the latter. An orthogonal regularization and a contrastive learning task on preference relevance score prediction are proposed as the dual safeguard to prevent preference-relevant representations from encoding irrelevant noise. For behavior feature extraction, we construct a balanced user-item graph by integrating factual and counterfactual graphs. In this process, we pre-train a behavior simulator to build the counterfactual graph with full interactions. Top-$K$sampling is adopted to omit noisy edges and add missing edges in the graph. The final recommendation is performed upon the fused representation of preference-relevant multimodal and behavior representations. Extensive experiments on three public datasets verify the power of our DANCE.
Xin Wen 0017, Weizhi Nie, Jing Liu 0002, Yuting Su 0001, Anan Liu
IEEE Trans. Multim.1
2024 Knowledge-Enhanced Causal Reinforcement Learning Model for Interactive Recommendation
abstract
Owing to its inherently dynamic nature and economical training cost, offline reinforcement learning (RL) is typically employed to implement an interactive recommender system (IRS). A crucial challenge in offline RL-based IRSs is the data sparsity issue, i.e., it is hard to mine user preferences well from the limited number of user-item interactions. In this article, we propose a knowledge-enhanced causal reinforcement learning model (KCRL) to mitigate data sparsity in IRSs. We make technical extensions to the offline RL framework in terms of the reward function and state representation. Specifically, we first propose a group preference-injected causal user model (GCUM) to learn user satisfaction (i.e., reward) estimation. We introduce beneficial group preference information, namely, the group effect, via causal inference to compensate for incomplete user interests extracted from sparse data. Then, we learn the RL recommendation policy with the reward given by the GCUM. We propose a knowledge-enhanced state encoder (KSE) to generate knowledge-enriched user state representations at each time step, which is assisted by a self-constructed user-item knowledge graph. Extensive experimental results on real-world datasets demonstrate that our model significantly outperforms the baselines.
Weizhi Nie, Xin Wen 0017, Jing Liu 0002, Jiawei Chen 0007, Jiancan Wu, Guoqing Jin, Anan Liu
IEEE Trans. Multim.2
2024 CDCM: ChatGPT-Aided Diversity-Aware Causal Model for Interactive Recommendation
abstract
In recent years, interactive recommender systems (IRSs) have attracted extensive interest. Existing IRSs are typically implemented with offline reinforcement learning (RL). They are devoted to improving recommendation accuracy by optimizing the extraction of users' inherent preferences. However, there hasn't been much attention on recommendation diversity, which could result in the monotony effect,i.e., categories of recommended items are consistently fixed and unchanging. In this paper, we center on category diversification in IRSs while largely preserving or even boosting recommendation accuracy. To this end, we propose a ChatGPT-aided diversity-aware causal model (CDCM) to enhance the offline RL framework with causal inference and ChatGPT. Specifically, we first propose a diversity-aware causal user model (DCUM) to estimate user satisfaction. This model disentangles the causal effect of users' inherent preferences and the monotony effect to obtain user satisfaction with both accuracy and diversity. Then, DCUM is used to assist the RL agent in recommendation policy learning. A ChatGPT-aided state encoder (CSE) is proposed to provide user state representation for each time step of policy learning. With the help of ChatGPT, CSE incorporates multi-category information in line with users' potential preferences to promote diverse and relevant category recommendations. Extensive experiment results on two real-world datasets validate the superiority of our CDCM regarding both accuracy and diversity.
Xin Wen 0017, Weizhi Nie, Jing Liu 0002, Yuting Su 0001, Yongdong Zhang 0001, Anan Liu
IEEE Trans. Multim.1
2024 Privacy-preserving Multi-source Cross-domain Recommendation Based on Knowledge Graph
abstract
The cross-domain recommender systems aim to alleviate the data sparsity problem in the target domain by transferring knowledge from the auxiliary domain. However, existing works ignore the fact that the data sparsity problem may also exist in the single auxiliary domain, and sharing user behavior data is restricted by the privacy policy. In addition, their cross-domain models lack interpretability. To address these concerns, we propose a novel multi-source cross-domain model based on knowledge graph. Specifically, to avoid the insufficiency of single auxiliary domain, we construct a knowledge graph comprehensively leveraging items from multiple auxiliary domains. To avoid the leakage of user privacy when user information is transferred to multiple domains, we construct graph for information transfer between items to effectively avoid the propagation of users’ private information between different domains. We implicitly integrate the user–item interaction by transferring the learned item embeddings. To improve the interpretability of cross-domain knowledge transfer, we propose a knowledge graph-based retrieval and fusion method to transfer knowledge derived from multiple auxiliary domains. An attention-based fusion network is designed to enhance the representation of the targeted user and items with the transferred item embedding. We perform extensive experiments on three real-world datasets, demonstrating that our model outperforms the states of the art.
Jing Liu 0002, Litao Shang, Yuting Su 0001, Weizhi Nie, Xin Wen 0017, Anan Liu
ACM Trans. Multim. Comput. Commun. Appl.5
2023 MRFT: Multiscale Recurrent Fusion Transformer Based Prior Knowledge for Bit-Depth Enhancement
abstract
Bit-depth enhancement (BDE) plays an important role in providing high bit-depth data support for high-dynamic range (HDR) display. Although convolutional neural network (CNN) based BDE methods have achieved top performance, multiscale feature extraction and fusion still suffer from some inherent architectural flaws. Moreover, the training-data-scarce scene has not been effectively explored. To this end, this paper proposes an innovative multiscale recurrent fusion transformer (MRFT) framework, which contains three key components, i.e. multiscale transformer feature encoder, recurrent feature fusion module, and prior knowledge injection. Specifically, the multiscale transformer feature encoder consists of a prior-injected context encoder (PICE) and a multiscale local feature encoder (MLFE). PICE leverages the vanilla self-attention mechanism to extract the global context correlating spatially-distant contents for distinguishing long-distance false contours. MLFE exploits the local self-attention mechanism with varied window sizes to capture different-scale detail features. Then, a hierarchical recurrent decoder (HRD) is proposed as the recurrent feature fusion module to fuse multiscale visual information with global guidance. Via the circular query-key mechanism, global-to-local information is progressively fused. Furthermore, we propose a two-stage alternating optimization strategy for prior knowledge injection. By pre-parameterizing the global auxiliary priors, the training dilemma on the data-scarce domain is significantly alleviated. Extensive analyses on multiple benchmark datasets demonstrate the superiority of our MRFT in terms of quantitative measures and aesthetic effects.
Xin Wen 0017, Weizhi Nie, Jing Liu 0002, Yuting Su 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Iterative Residual Feature Refinement Network for Bit-Depth Enhancement
abstract
Bit-depth enhancement (BDE) restores high bit-depth (HBD) images from low bit-depth (LBD) ones, which has important applications. Recently, residual-optimized BDE algorithms based on convolutional neural networks (CNNs) have achieved top performance. However, they fail to use a single model to accurately recover all frequency information encoded by missing significant bits at one time on challenging large bit-depth recovery tasks. In this paper, we redefine BDE residual recovery from the perspective of image frequency characteristics. On this basis, we propose an iterative residual feature optimization strategy, which provides an implicit error correction mechanism and improves training and inference efficiency. Furthermore, we design a simple but effective iterative residual feature refinement network (IRFRN). By linking model complexity with the recovery of different frequency information, IRFRN enables a single model to simultaneously recover the missing low and high frequency information. Extensive experiments indicate that our method achieves the state-of-the-art quantitative and qualitative performance on large bit-depth recovery tasks.
Weizhi Nie, Xin Wen 0017, Jing Liu 0002, Yuting Su 0001
IEEE Signal Process. Lett.2
2022 Residual-Guided Multiscale Fusion Network for Bit-Depth Enhancement
abstract
Bit-depth enhancement (BDE) is a challenging task due to stubborn false contour artifacts and disappeared detailed information. Given the mixture of structural distortions and real edges in low bit-depth (LBD) images, both large and small receptive fields (RFs) are critical for BDE tasks. However, even powerful state-of-the-art CNN-based methods can hardly capture sufficient LBD features under multiple RFs. This paper proposes a residual-guided multiscale fusion network (RMFNet) to explore multiscale features in a residual manner. We find that the shuffling operation provides desired multiscale inputs for effectively distinguishing false contours from real edges without any loss of information. Therefore, we shuffle LBD images to multiple scales and then fully extract residual features under different RFs with corresponding subnets. To facilitate interscale guidance from the global context to the local context, we progressively transfer the encoded residual features between adjacent subnets from top to bottom. We further propose a dual-branch depthwise group fusion (DDGF) module to fully capture inter- and inner correlations of multiscale features with fewer parameters. Finally, extensive experiments show that our algorithm achieves excellent performance improvement both quantitatively and qualitatively, verifying its effectiveness.
Jing Liu 0002, Xin Wen 0017, Weizhi Nie, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2