Yan Gao 0029

dblp:46/3479-29 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-8012-1392ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Behavioral Feature Boosting via Substitute Relationships for E-commerce Search
abstract
On E-commerce platforms, new products often suffer from the cold-start problem: limited interaction data reduces their search visibility and hurts relevance ranking. To address this, we propose a simple yet effective behavior feature boosting method that leverages substitute relationships among products (BFS). BFS identifies substitutes—products that satisfy similar user needs—and aggregates their behavioral signals (e.g., clicks, add-to-carts, purchases, and ratings) to provide a warm start for new items. Incorporating these enriched signals into ranking models mitigates cold-start effects and improves relevance and competitiveness. Experiments on a large E-commerce platform, both offline and online, show that BFS significantly improves search relevance and product discovery for cold-start products. BFS is scalable and practical, improving user experience while increasing exposure for newly launched items in E-commerce search. The BFS-enhanced ranking model has been launched in production and has served customers since 2025.
Chaosheng Dong, Michinari Momma, Yan Gao 0029
SIGIR4
2026 Multi-modal Relational Item Representation Learning for Inferring Substitutable and Complementary Items
abstract
We study the problem of inferring substitutable and complementary items, which underpins applications such as alternative and follow-up purchase suggestions. Existing approaches typically learn from behavior-derived item-item associations using GNNs or leverage item content alone. However, these methods often overlook two key challenges: (i) user behaviors (e.g., co-view/co-purchase) only provide noisy weak supervision, and (ii) behavior signals are long-tailed, leaving many items with sparse associations. We propose MMSC, a self-supervised multi-modal relational representation learning framework that combines a multi-modal foundation model adapted to encode item metadata and a self-supervised denoising module that learns relationship-aware representations from noisy user behaviors, unified by a hierarchical aggregation mechanism. We further use LLM-assisted supervision to mitigate noise in behavior-derived supervision during training. Experiments on five real-world datasets show that MMSC consistently outperforms existing baselines by 26.1% for substitutable and 39.2% for complementary item inference, while remaining effective for cold-start items.
Junting Wang 0001, Chenghuan Guo, Hari Sundaram, Yan Gao 0029
SIGIR6
2025 Learning Attribute as Explicit Relation for Sequential Recommendation
abstract
The data on user behaviors is sparse given the vast array of user-item combinations. Attributes related to users (e.g., age), items (e.g., brand), and behaviors (e.g., co-purchase) serve as crucial input sources for item-item transitions of user's behavior prediction. While recent Transformer-based sequential recommender systems learn the attention matrix for each attribute to update item representations, the attention of a specific attribute is optimized by gradients from all input sources, leading to potential information mixture. Besides, Transformers mainly focus on intra-sequence attention for item attributes, neglecting cross-sequence relations and user attributes. Addressing these challenges, we propose the Attribute Transformer (AttrFormer) to learn attributes as explicit relations. This model transforms each type of attribute into an explicit relation defined in the feature space, and it ensures no information mixing among different input sources. Explicit relations introduce cross-sequence and intra-sequence relations. AttrFormer has novel relation-augmented heads to handle them at both the item and behavioral levels, seamlessly integrating the augmented heads into the multi-head attention mechanism. Furthermore, we employ position-to-position aggregation to refine behavior representation for users with similar patterns at the sequence level. To capture the subjective nature of user preferences, AttrFormer is trained using posterior targets where upcoming user behaviors follow a multinomial distribution with a Dirichlet prior. Our evaluations on four popular datasets, including Amazon (Toys & Games and Beauty) and MovieLens (1M and 25M versions), reveal that AttrFormer outperforms leading Transformer baselines, achieving around 20% improvement in NDCG@20 scores. Extensive ablation studies also demonstrate the efficiency of AttrFormer in managing long behavior sequences and inter-sequence relations.
Gang Liu 0025, Fan Yang 0084, Alireza Bagheri Garakani, Tian Tong, Yan Gao 0029, Meng Jiang 0001
KDD (1)6
2025 STIMULUS: Achieving Fast Convergence and Low Sample Complexity in Stochastic Multi-Objective Learning
abstract
Recently, multi-objective optimization (MOO) has gained attention for its broad applications in ML, operations research, and engineering. However, MOO algorithm design remains in its infancy and many existing MOO methods suffer from unsatisfactory convergence rate and sample complexity performance. To address this challenge, in this paper, we propose an algorithm called STIMULUS (**st**ochastic path-**i**ntegrated **mul**ti-gradient rec**u**rsive e**s**timator), a new and robust approach for solving MOO problems. Different from the traditional methods, STIMULUS introduces a simple yet powerful recursive framework for updating stochastic gradient estimates to improve convergence performance with low sample complexity. In addition, we introduce an enhanced version of \algns, termed \algmns, which incorporates a momentum term to further expedite convergence. We establish $\mathcal{O}(1/T)$ convergence rates of the proposed methods for non-convex settings and $\mathcal{O}(\exp{-\mu T})$ for strongly convex settings, where $T$ is the total number of iteration rounds. Additionally, we achieve the state-of-the-art $O\left(n+\sqrt{n}\epsilon^{-1}\right)$ sample complexities for non-convex settings and $\mathcal{O}\left(n+ \sqrt{n} \ln ({\mu/\epsilon})\right)$ for strongly convex settings, where $\epsilon>0$ is a desired stationarity error. Moreover, to alleviate the periodic full gradient evaluation requirement in STIMULUS and STIMULUS-M, we further propose enhanced versions with adaptive batching called STIMULUS$^+$/ STIMULUS-M$^+$ and provide their theoretical analysis.
Zhuqing Liu, Chaosheng Dong, Michinari Momma, Simone Shao, Shaoyuan Xu, Yan Gao 0029, Haibo Yang 0001, Jia Liu 0002
UAI6
2025 Divide and Orthogonalize: Efficient Continual Learning with Local Model Space Projection
abstract
Continual learning (CL) has gained increasing interest in recent years due to the need for models that can continuously learn new tasks while retaining knowledge from previous ones. However, existing CL methods often require either computationally expensive layer-wise gradient projections or large-scale storage of past task data, making them impractical for resource-constrained scenarios. To address these challenges, we propose a local model space projection (LMSP)-based continual learning framework that significantly reduces computational complexity from $\mathcal{O}(n^3)$ to $\mathcal{O}(n^2)$ while preserving both forward and backward knowledge transfer with minimal performance trade-offs. We establish a theoretical analysis of the error and convergence properties of LMSP compared to conventional global approaches. Extensive experiments on multiple public datasets demonstrate that our method achieves competitive performance while offering substantial efficiency gains, making it a promising solution for scalable continual learning.
Simone Shao, Tian Tong, Fan Yang 0084, Yetian Chen, Jia Liu 0002, Yan Gao 0029
UAI8
2024 Transitivity-Encoded Graph Attention Networks for Complementary Item Recommendations
abstract
In e-commerce recommender systems, providing product suggestions to customers that are often bought together, which is called “complementary recommendation,” not only improves customer experience but also boosts business impact. However, in practice, it is highly challenging to efficiently extract the complementary relations between the items due to noisy and low coverage of the co-purchased records in transaction datasets. To address these challenges, graph neural networks (GNN) have been increasingly adopted in complementary item recommendations thanks to their capabilities in integrating side-information and topological structures to extract these complex item relationships. However, most existing GNN-based methods fall short in learning better product complementary representation since they often utilize a simple one-to-one product-to-vector mapping strategy, which fails to describe the transitive logic of complementary items. To overcome this challenge, we propose a new GNN model called transitivity-encoded graph attention networks (TransGAT). To our knowledge, TransGAT is the first method that extends representation space by encoding the behavioral direction into embedding space in GNN and enabling mutual relationship extraction between complementary items. In order to better extract customer's intrinsic behavioral information, we further adopt the substitute information as the guidance by jointly learning complements and substitutes graphs and coupling them together. Moreover, several self-supervised data augmentation strategies are incorporated in our approach. Through evaluations on three real-world datasets, TransGAT consistently surpasses contemporary benchmarks, showcasing its prowess in complementary item recommendations.
Chenghuan Guo, Minghao Sun, Yan Gao 0029, Jia Liu 0002, Michinari Momma, Itetsu Taru
ICDM5
2024 Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning
abstract
Reinforcement learning with multiple, potentially conflicting objectives is pervasive in real-world applications, while this problem remains theoretically under-explored. This paper tackles the multi-objective reinforcement learning (MORL) problem and introduces an innovative actor-critic algorithm named MOAC which finds a policy by iteratively making trade-offs among conflicting reward signals. Notably, we provide the first analysis of finite-time Pareto-stationary convergence and corresponding sample complexity in both discounted and average reward settings. Our approach has two salient features: (a) MOAC mitigates the cumulative estimation bias resulting from finding an optimal common gradient descent direction out of stochastic samples. This enables provable convergence rate and sample complexity guarantees independent of the number of objectives; (b) With proper momentum coefficient, MOAC initializes the weights of individual policy gradients using samples from the environment, instead of manual initialization. This enhances the practicality and robustness of our algorithm. Finally, experiments conducted on a real-world dataset validate the effectiveness of our proposed method.
Hairi, Haibo Yang 0001, Jia Liu 0002, Tian Tong, Fan Yang 0084, Michinari Momma, Yan Gao 0029
ICML8
2024 Explainable Uncertainty Attribution for Sequential Recommendation
Carles Balsells Rodas, Fan Yang 0084, Zhishen Huang, Yan Gao 0029
SIGIR4
2023 Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion Analysis
abstract
Fashion representation learning involves the analysis and understanding of various visual elements at different granularities and the interactions among them. Existing works often learn fine-grained fashion representations at the attribute level without considering their relationships and inter-dependencies across different classes. In this work, we propose to learn an attribute and class-specific fashion representation duet to better model such attribute relationships and inter-dependencies by leveraging prior knowledge about the taxonomy of fashion attributes and classes. Through two sub-networks for the attributes and classes, respectively, our proposed an embedding network progressively learns and refines the visual representation of a fashion image to improve its robustness for fashion retrieval. A multi-granularity loss consisting of attribute-level and class-level losses is proposed to introduce appropriate inductive bias to learn across different granularities of the fashion representations. Experimental results on three benchmark datasets demonstrate the effectiveness of our method, which outperforms the state-of-the-art methods by a large margin.
Yan Gao 0029, Jingjing Meng
CVPR2
2022 Fine-Grained Fashion Representation Learning by Online Deep Clustering
Ning Xie 0009, Yan Gao 0029, Chien-Chih Wang
ECCV (27)3