Fengxin Li

dblp:216/1573 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-8828-2767ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis
abstract
Audio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce coarse gestures, lack expressiveness, and fail to fully align with audio semantics. To address these challenges, we propose ExGes, a novel retrieval-enhanced diffusion framework with three key designs: (1) a Motion Base Construction, which builds a gesture library from the training dataset; (2) a Motion Retrieval Module, employing contrastive learning and momentum distillation for retrieving fine-grained reference poses; and (3) a Precise Control Module, integrating partial masking and stochastic masking to enable flexible and fine-grained control. Experimental evaluations on BEAT2 demonstrate that ExGes reduces Fréchet Gesture Distance by 4.55%and improves motion diversity by 5.3% over EMAGE, with user studies revealing a 71.3% preference for its naturalness and semantic relevance.
Xukun Zhou, Fengxin Li, Yan Zhou 0003, Pengfei Wan 0001, Yeying Jin, Hongyuan Zhang 0001, Hongyan Liu 0002, Zhaoxin Fan, Jun He 0008, Xuelong Li 0001
IEEE Trans. Vis. Comput. Graph.2
2025 KAN v.s. MLP for Offline Reinforcement Learning
abstract
Kolmogorov-Arnold Networks (KAN) is an emerging neural network architecture in machine learning. It has greatly interested the research community about whether KAN can be a promising alternative to the commonly used Multi-Layer Perceptions (MLP). Experiments in various fields demonstrated that KAN-based machine learning can achieve comparable if not better performance than MLP-based methods, but with much smaller parameter scales and are more explainable. In this paper, we explore the incorporation of KAN into the actor and critic networks for offline reinforcement learning (RL). We evaluated the performance, parameter scales, and training efficiency of various KAN and MLP-based conservative Q-learning (CQL) on the classical D4RL benchmark for offline RL. Our study demonstrates that KAN can achieve performance close to the commonly used MLP with significantly fewer parameters. This allows us to choose the base networks according to the offline RL task requirements.
Haihong Guo, Fengxin Li, Jiao Li 0001, Hongyan Liu 0002
ICASSP2
2025 Offline Reinforcement Learning via Conservative Smoothing and Dynamics Controlling
abstract
Offline Reinforcement Learning (RL) optimizes policy using pre-collected data instead of direct environment interaction, offering a safe and cost-effective solution for sequential decision-making in the real world. However, it faces challenges such as distribution shift issues and vulnerability under perturbations. Researchers have developed various conservative methods to improve the robustness of offline RL. Nevertheless, model-based methods can result in transition distribution shift issues, while model-free value-based uncertainty penalty methods may not be sufficiently robust. To address these problems, we propose a new method called Robust Offline RL via Conservative Smoothing and Dynamics Controlling (RCSD). To achieve reliable value estimation of out-of-distribution (OOD) actions, RCSD uses both model-free uncertainty penalty and model-based simulation methods. It introduces a new one-step simulation method with conservative dynamics controlling to avoid value overestimation caused by transition distribution shifts. Moreover, RCSD considers both current and next states when generating OOD states to ensure cautious value estimation and efficient data utilization. RCSD uses conservative Q-smoothing and policy smoothing to strengthen the policy against sudden changes under perturbations. Experiments on D4RL benchmark demonstrate that RCSD can achieve state-of-the-art performance compared to baselines in either benchmark or adversarial attack tests.
Haihong Guo, Fengxin Li, Jiao Li 0001, Hongyan Liu 0002
ICASSP2
2025 Debiased Estimation for Cross-Domain Cold Start Recommendation
abstract
Mapping-based methods are critical solutions for Cross-Domain Cold Start Recommendation problem, which learns a mapping function to transfer knowledge from the source domain to the target domain based on overlap users between domains. However, in many CDCSR scenarios, there exists selection bias in observed overlap users due to the users being free to choose whether to rate or interact with items of a specific domain. This selection bias without proper treatment can introduce bias to the model learning process.To address the selection bias in observed overlap users, we propose two estimators for CDCSR: Inverse Propensity Weighting CDCSR Estimator (IPW-CDCSRE) and Doubly Robust CDCSR Estimator (DR-CDCSRE). IPW-CDCSRE learns propensity scores to adjust for selection bias and enhances user representations through the learned propensity model. DR-CDCSRE further employs an imputation model to consider prediction loss for non-overlap users. Additionally, we introduce the Recurrent Learning Process (RLP) to enhance the stability and effectiveness of DR-CDCSRE. To validate the effectiveness and generalization ability of the proposed estimators, we conducted extensive experiments on three real-world CDCSR scenarios, utilizing four base CDCSR models and two types of loss functions.
Fengxin Li, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001
ICASSP1
2025 Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
abstract
Audio-driven 3D face animation is crucial for live streaming and augmented reality, yet most existing methods focus on specific individuals with predefined speaking styles, limiting adaptability to varied styles. To address this, we introduce MetaFace, a novel methodology for speaking style adaptation based on meta-learning. MetaFace comprises three key components: the Robust Meta Initialization Stage (RMIS) for foundational style adaptation, the Dynamic Relation Mining Neural Process (DRMN) to connect observed and unobserved speaking styles, and a Low-rank Matrix Memory Reduction Approach to optimize model efficiency and style detail learning. These innovations enable MetaFace to significantly outperform existing baselines and set a new state-of-the-art, as demonstrated by our experimental results.
Xukun Zhou, Fengxin Li, Ziqiao Peng, Hongyan Liu 0002, Zhaoxin Fan, Jun He 0008
ICME2
2025 Diffusion Alignment for Cross Domain Recommendation
abstract
Cross-domain recommendation (CDR) is a critical solution to address the sparsity issue of conventional recommendation systems. In most scenarios, only a subset of users have interactions on both domains, which is referred to as partial user overlap CDR. Despite this, most CDR methods tend to design alignment mechanisms only for overlapping users, resulting in partial alignment issue. In this paper, we analyze the cause of partial alignment, including the limited expressive capacity of the models and the lack of alignment signal for non-overlapping users. To address this issue, we propose a new model DACDR (Diffusion Alignment Cross-domain recommendation) with two specially designed modules: a user latent diffusion module and a user alignment module. The user latent diffusion module integrates a diffusion model into the CDR process to enhance the expressive capacity of the model. The user alignment module introduces novel alignment mechanisms that consider both overlapping and non-overlapping users. We conduct extensive experiments on three CDR scenarios to evaluate the performance of DACDR. Our results demonstrate that DACDR outperforms the baselines.
Fengxin Li, Hongyan Liu 0002, Jun He 0008
ICMR1
2025 Topic Guided Multi-faceted Semantic Disentanglement for CTR prediction
abstract
Click-through rate (CTR) prediction lies at the heart of the online advertising ecosystem and recommendation systems, helping to improve user engagement and platform revenue. With the advent of Pretrained Language Models (PLMs), researchers focus on incorporating textual features to enhance semantic understanding in CTR prediction. However, existing methods typically aggregate a wealth of textual features and encode the informative text into a single semantic embedding. This mechanism leads to entangled embedding that fails to capture fine-grained feature interactions, ultimately limiting CTR prediction performance. To address this issue, we propose Multi-faceted Semantic Disentanglement for CTR prediction (MSD-CTR), a novel framework designed to disentangle and leverage multi-faceted textual information. MSD-CTR consists of two key components: Disentangled Semantic Topic Model (DSTopic) and Topic Guided Disentangled Representation Learning (TopicDRL). DSTopic employs a disentangled generative process to extract multi-faceted knowledge from the entangled textual information. Meanwhile, TopicDRL integrates the extracted multi-faceted knowledge into CTR prediction and introduces two alignment losses to guide disentangled semantic embedding learning. Extensive experiments on four real-world datasets demonstrate that MSD-CTR outperforms existing CTR models, highlighting the effectiveness of disentangling textual information for better click-through rate prediction.
Fengxin Li, Zhiqian Yin, Hongyan Liu 0002, Jingcai Guo, Jun He 0008, Haijie Gu
ACM Multimedia1
2025 LEADRE: Multi-Faceted Knowledge Enhanced LLM Empowered Display Advertisement Recommender System
abstract
Display advertising plays a crucial role in benefiting advertisers, publishers, and users. Traditional display advertising systems employ a multi-stage architecture comprising retrieval, coarse ranking, ranking, and re-ranking. However, conventional retrieval methods primarily rely on ID-based learning-to-rank mechanisms, often underutilizing the content information of ads, like ads' title, and description. This limitation reduces the ability to generate diverse and relevant recommendation lists. To address this challenge, we propose leveraging the extensive world knowledge of large language models (LLMs). However, effectively integrating LLMs into advertising systems presents three key challenges: ( i) How to accurately capture user interests, (ii) How to bridge the knowledge gap between LLMs and advertising systems , and ( iii) How to efficiently deploy LLMs at scale. To overcome these challenges, we introduce LEADRE —the L LM E mpowered Display AD vertisement RE commender system. LEADRE consists of three core components. The Intent-Aware Prompt Engineering module introduces multi-faceted knowledge and constructs intent-aware pairs, fine-tuning LLMs to generate ads tailored to users' personal interests. The Advertising-Specific Knowledge Alignment module incorporates auxiliary fine-tuning tasks and Direct Preference Optimization (DPO) to align LLMs with advertising semantics and business objectives. The Latency-Aware Model Deployment module integrates a hybrid service framework that balances latency-tolerant and latency-sensitive service, ensuring seamless online deployment. Extensive offline experiments validate the effectiveness of LEADRE, demonstrating significant improvements across multiple evaluation metrics. Furthermore, online A/B tests reveal a 1.57% and 1.17% increase in Gross Merchandise Value (GMV) for serviced users on WeChat Channels and Moments, respectively. LEADRE has been successfully deployed on both platforms, handling tens of billions of requests daily.
Fengxin Li, Xiaoxiang Deng, Haijie Gu, Biao Qin
Proc. VLDB Endow.1
2024 PoseRec: 3D Human Pose Driven Online Advertisement Recommendation for Micro-videos
abstract
In this paper, we present PoseRec, an innovative approach aimed at enhancing online advertisement recommendations for micro-videos to boost click-through rates. Addressing the inherent background bias introduced via direct video content learning from image frames, we exploit rich data within the 3D human pose. PoseRec capitalizes on the merits of 3D human pose detection and multi-frame pose data, resulting in superior advertisement recommendation performance. Additionally, we introduce a unique item-aware implicit prototype learning module and a pose-aware transductive hard-negative mining module to tackle the issues of ambiguity and sparsity in advertisement recommendation. Upon evaluation on our novel dataset, Pose-OBE, our method exhibits robust performance surpassing strong baselines, corroborating its effectiveness in resolving the complex challenges of micro-video advertisement recommendation.
Zhaoxin Fan, Fengxin Li, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001
ICMR2
2024 CausalCDR: Causal Embedding Learning for Cross-domain Recommendation
abstract
Cross-domain recommendation (CDR) methods achieve success in disentangling user preferences into domain-specific and domain-shared parts. However, recent research has shown that isolated domain-specific preference limits performance improvements. In this paper, we propose a new CDR framework, called CausalCDR, which identifies the limitations of existing methods and addresses existing issues. CausalCDR consists of two views: the causal view and the generative view. The causal view incorporates causality of variables into the CDR scenario, while the generative view implements the causal view by modeling the joint distribution of user interaction via encoding, causal, and generation stage. To optimize CausalCDR, we re-derive the Evidence Lower Bound (ELBO) and introduce a mutual information regularizer and an adversarial classifier. We evaluate CausalCDR on four real-world CDR scenarios and demonstrate its effectiveness in improving CDR performance.
Fengxin Li, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001
SDM1
2022 Why do Semantically Unrelated Categories Appear in the Same Session?: A Demand-aware Method
abstract
Session-based recommendation has recently attracted more and more research efforts. Most existing approaches are intuitively proposed to discover users' potential preferences or interests from the anonymous session data. This apparently ignores the fact that these sequential behavior data usually reflect session user's potential demand, i.e., a semantic level factor, and therefore how to estimate underlying demands from a session has become a challenging task. To tackle the aforementioned issue, this paper proposes a novel demand-aware graph neural network model. Particularly, a demand modeling component is designed to extract the underlying multiple demands of each session. Then, the demand-aware graph neural network is designed to first construct session demand graphs and then learn the demand-aware item embeddings to make the recommendation. The mutual information loss is further designed to enhance the quality of the learnt embeddings. Extensive experiments have been performed on two real-world datasets and the proposed model achieves the SOTA model performance.
Liqi Yang, Linhao Luo, Xiaofeng Zhang 0002, Fengxin Li, Xinni Zhang, Zelin Jiang
SIGIR4
2020 AlTwo: Vehicle Recognition in Foggy Weather Based on Two-Step Recognition Algorithm
Fengxin Li, Ziye Luo, Jingyu Huang, Lingzhan Wang, Jinpu Cai
ISNN1