VLDB 2026 Research / reviewers in the wild / expert
Yun He 0001
dblp:02/6731-1
· DBLP profile ↗
13ranked-venue papers
6as first author
5since 2021 · last 2022
0000-0001-9462-4583ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | MetaBalance: Improving Multi-Task Recommendations via Adapting Gradient Magnitudes of Auxiliary TasksabstractIn many personalized recommendation scenarios, the generalization ability of a target task can be improved via learning with additional auxiliary tasks alongside this target task on a multi-task network. However, this method often suffers from a serious optimization imbalance problem. On the one hand, one or more auxiliary tasks might have a larger influence than the target task and even dominate the network weights, resulting in worse recommendation accuracy for the target task. On the other hand, the influence of one or more auxiliary tasks might be too weak to assist the target task. More challenging is that this imbalance dynamically changes throughout the training process and varies across the parts of the same network. We propose a new method: MetaBalance to balance auxiliary losses via directly manipulating their gradients w.r.t the shared parameters in the multi-task network. Specifically, in each training iteration and adaptively for each part of the network, the gradient of an auxiliary loss is carefully reduced or enlarged to have a closer magnitude to the gradient of the target loss, preventing auxiliary tasks from being so strong that dominate the target task or too weak to help the target task. Moreover, the proximity between the gradient magnitudes can be flexibly adjusted to adapt MetaBalance to different scenarios. The experiments show that our proposed method achieves a significant improvement of 8.34% in terms of [email protected] upon the strongest baseline on two real-world datasets. The code of our approach can be found at here.1 Yun He 0001, Geng Ji 0001, Yunsong Guo, James Caverlee |
WWW | 1 |
| 2022 | Item Relationship Graph Neural Networks for E-CommerceabstractIn a modern e-commerce recommender system, it is important to understand the relationships among products. Recognizing product relationships-such as complements or substitutes-accurately is an essential task for generating better recommendation results, as well as improving explainability in recommendation. Products and their associated relationships naturally form a product graph, yet existing efforts do not fully exploit the product graph's topological structure. They usually only consider the information from directly connected products. In fact, the connectivity of products a few hops away also contains rich semantics and could be utilized for improved relationship prediction. In this work, we formulate the problem as a multilabel link prediction task and propose a novel graph neural network-based framework, item relationship graph neural network (IRGNN), for discovering multiple complex relationships simultaneously. We incorporate multihop relationships of products by recursively updating node embeddings using the messages from their neighbors. An edge relational network is designed to effectively capture relational information between products. Extensive experiments are conducted on real-world product data, validating the effectiveness of IRGNN, especially on large and sparse product graphs. Weiwen Liu, Yin Zhang 0011, Jianling Wang, Yun He 0001, James Caverlee, Patrick P. K. Chan, Daniel S. Yeung, Pheng-Ann Heng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Vibe check: social resonance learning for enhanced recommendationabstractSocial Resonance is a common socio-behavioral phenomenon in which users are more influenced by opinions that have similar vibes. That is, opinions from two different groups of users can mutually reinforce (or resonate with) each other to have an even stronger impact on the user. In this paper, we explore the powerful social resonance effect between social connections and other users in an eCommerce platform to improve recommendation. Specifically, we first formulate an item-aware user influence network that connects users who rate the same item. With the social network and item-aware user influence network, a novel graph-based mutual learning framework is proposed, which captures the resonance influence from both user local correlations and global connections. We then fuse these influence paths to predict the resonance-enhanced user preference towards items. Experiments on public benchmarks show the proposed approach outperforms state-of-the-art social recommendation methods. Yin Zhang 0011, Yun He 0001, James Caverlee |
ASONAM | 2 |
| 2021 | Popularity Bias in Dynamic RecommendationabstractPopularity bias is a long-standing challenge in recommender systems: popular items are overly recommended at the expense of less popular items that users may be interested in being under-recommended. Such a bias exerts detrimental impact on both users and item providers, and many efforts have been dedicated to studying and solving such a bias. However, most existing works situate the popularity bias in a static setting, where the bias is analyzed only for a single round of recommendation with logged data. These works fail to take account of the dynamic nature of real-world recommendation process, leaving several important research questions unanswered: how does the popularity bias evolve in a dynamic scenario? what are the impacts of unique factors in a dynamic recommendation process on the bias? and how to debias in this long-term dynamic process? In this work, we investigate the popularity bias in dynamic recommendation and aim to tackle these research gaps. Concretely, we conduct an empirical study by simulation experiments to analyze popularity bias in the dynamic scenario and propose a dynamic debiasing strategy and a novel False Positive Correction method utilizing false positive signals to debias, which show effective performance in extensive experiments. Ziwei Zhu 0001, Yun He 0001, Xing Zhao 0003, James Caverlee |
KDD | 2 |
| 2021 | Popularity-Opportunity Bias in Collaborative FilteringabstractThis paper connects equal opportunity to popularity bias in implicit recommenders to introduce the problem of popularity-opportunity bias. That is, conditioned on user preferences that a user likes both items, the more popular item is more likely to be recommended (or ranked higher) to the user than the less popular one. This type of bias is harmful, exerting negative effects on the engagement of both users and item providers. Thus, we conduct a three-part study: (i) By a comprehensive empirical study, we identify the existence of the popularity-opportunity bias in fundamental matrix factorization models on four datasets; (ii) coupled with this empirical study, our theoretical study shows that matrix factorization models inherently produce the bias; and (iii) we demonstrate the potential of alleviating this bias by both in-processing and post-processing algorithms. Extensive experiments on four datasets show the effective debiasing performance of these proposed methods compared with baselines designed for conventional popularity bias. Ziwei Zhu 0001, Yun He 0001, Xing Zhao 0003, Yin Zhang 0011, Jianling Wang, James Caverlee |
WSDM | 2 |
| 2020 | PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain KnowledgeabstractWe present a new benchmark dataset called PARADE for paraphrase identification that requires specialized domain knowledge.PA-RADE contains paraphrases that overlap very little at the lexical and syntactic level but are semantically equivalent based on computer science domain knowledge, as well as nonparaphrases that overlap greatly at the lexical and syntactic level but are not semantically equivalent based on this domain knowledge.Experiments show that both state-of-the-art neural models and non-expert human annotators have poor performance on PARADE.For example, BERT after fine-tuning achieves an F1 score of 0.709, which is much lower than its performance on other paraphrase identification datasets.PARADE can serve as a resource for researchers interested in testing models that incorporate domain knowledge.We make our data and code freely available.1 Yun He 0001, Zhuoer Wang, Yin Zhang 0011, Ruihong Huang, James Caverlee |
EMNLP (1) | 1 |
| 2020 | Infusing Disease Knowledge into BERT for Health Question Answering, Medical Inference and Disease Name RecognitionabstractKnowledge of a disease includes information of various aspects of the disease, such as signs and symptoms, diagnosis and treatment.This disease knowledge is critical for many healthrelated and biomedical tasks, including consumer health question answering, medical language inference and disease name recognition.While pre-trained language models like BERT have shown success in capturing syntactic, semantic, and world knowledge from text, we find they can be further complemented by specific information like knowledge of symptoms, diagnoses, treatments, and other disease aspects.Hence, we integrate BERT with disease knowledge for improving these important tasks.Specifically, we propose a new disease knowledge infusion training procedure and evaluate it on a suite of BERT models including BERT, BioBERT, SciBERT, Clinical-BERT, BlueBERT, and ALBERT.Experiments over the three tasks show that these models can be enhanced in nearly all cases, demonstrating the viability of disease knowledge infusion.For example, accuracy of BioBERT on consumer health question answering is improved from 68.29% to 72.09%, while new SOTA results are observed in two datasets.We make our data and code freely available. Yun He 0001, Ziwei Zhu 0001, Yin Zhang 0011, Qin Chen 0001, James Caverlee |
EMNLP (1) | 1 |
| 2020 | Content-Collaborative Disentanglement Representation Learning for Enhanced RecommendationabstractModern recommenders usually consider both collaborative features from user behavior data (e.g., clicks) and content information about the users and items (e.g., user ages or item images) for improved recommendations. While encouraging, the uncovered user preference representations derived from these collaborative and content-based perspectives can be entangled by intermixing the influence from each other, leading to sub-optimal performance and unstable recommendations. Hence, we propose to disentangle representations learned from user behavior data and content information. Specifically, we propose a novel two-level disentanglement generative recommendation model (DICER) that supports both content-collaborative disentanglement and feature disentanglement: for the content-collaborative disentanglement, DICER decomposes the features by their marginal distributions based on content and user-item interactions, to ensure the learned features from each type are statistically independent. For feature disentanglement, by decomposing the Kullback-Leibler divergence, we theoretically show that extracted features within each type are disentangled at a granular level. Furthermore, DICER utilizes a co-decoder that simultaneously decodes the content and user-item interactions to ensure the high-quality of learned features. Through extensive experiments on three real-world datasets, results show that DICER significantly outperforms other state-of-the-art methods by 13.5% in NDCG and 14.4% in hit ratio on average. Yin Zhang 0011, Ziwei Zhu 0001, Yun He 0001, James Caverlee |
RecSys | 3 |
| 2020 | Unbiased Implicit Recommendation and Propensity Estimation via Combinational Joint LearningabstractThis paper focuses on how to generate unbiased recommendations based on biased implicit user-item interactions. We propose a combinational joint learning framework to simultaneously learn unbiased user-item relevance and unbiased propensity. More specifically, we first present a new unbiased objective function for estimating propensity. We then show how a naïve joint learning approach faces an estimation-training overlap problem. Hence, we propose to jointly train multiple sub-models from different parts of the training dataset to avoid this problem. Finally, we show how to incorporate residual components trained by the complete training data to complement the relevance and propensity sub-models. Extensive experiments on two public datasets demonstrate the effectiveness of the proposed model with an improvement of 4% on average over the best alternatives. Ziwei Zhu 0001, Yun He 0001, Yin Zhang 0011, James Caverlee |
RecSys | 2 |
| 2020 | Consistency-Aware Recommendation for User-Generated Item List ContinuationabstractUser-generated item lists are popular on many platforms. Examples include video-based playlists on YouTube, image-based lists (or "boards") on Pinterest, book-based lists on Goodreads, and answer-based lists on question-answer forums like Zhihu. As users create these lists, a common challenge is in identifying what items to curate next. Some lists are organized around particular genres or topics, while others are seemingly incoherent, reflecting individual preferences for what items belong together. Furthermore, this heterogeneity in item consistency may vary from platform to platform, and from sub-community to sub-community. Hence, this paper proposes a generalizable approach for user-generated item list continuation. Complementary to methods that exploit specific content patterns (e.g., as in song-based playlists that rely on audio features), the proposed approach models the consistency of item lists based on human curation patterns, and so can be deployed across a wide range of varying item types (e.g., videos, images, books). A key contribution is in intelligently combining two preference models via a novel consistency-aware gating network -- a general user preference model that captures a user's overall interests, and a current preference priority model that captures a user's current (as of the most recent item) interests. In this way, the proposed consistency-aware recommender can dynamically adapt as user preferences evolve. Evaluation over four datasets (of songs, books, and answers) confirms these observations and demonstrates the effectiveness of the proposed model versus state-of-the-art alternatives. Further, all code and data are available at https://github.com/heyunh2015/ListContinuation_WSDM2020. Yun He 0001, Yin Zhang 0011, Weiwen Liu, James Caverlee |
WSDM | 1 |
| 2020 | Adaptive Hierarchical Translation-based Sequential RecommendationabstractWe propose an adaptive hierarchical translation-based sequential recommendation called HierTrans that first extends traditional item-level relations to the category-level, to help capture dynamic sequence patterns that can generalize across users and time. Then unlike item-level based methods, we build a novel hierarchical temporal graph that contains item multi-relations at the category-level and user dynamic sequences at the item-level. Based on the graph, HierTrans adaptively aggregates the high-order multi-relations among items and dynamic user preferences to capture the dynamic joint influence for next-item recommendation. Specifically, the user translation vector in HierTrans can adaptively change based on both a user’s previous interacted items and the item relations inside the user’s sequences, as well as the user’s personal dynamic preference. Experiments on public datasets demonstrate the proposed model HierTrans consistently outperforms state-of-the-art sequential recommendation methods. Yin Zhang 0011, Yun He 0001, Jianling Wang, James Caverlee |
WWW | 2 |
| 2019 | A Hierarchical Self-Attentive Model for Recommending User-Generated Item ListsabstractUser-generated item lists are a popular feature of many different platforms. Examples include lists of books on Goodreads, playlists on Spotify and YouTube, collections of images on Pinterest, and lists of answers on question-answer sites like Zhihu. Recommending item lists is critical for increasing user engagement and connecting users to new items, but many approaches are designed for the item-based recommendation, without careful consideration of the complex relationships between items and lists. Hence, in this paper, we propose a novel user-generated list recommendation model called AttList. Two unique features of AttList are careful modeling of (i) hierarchical user preference, which aggregates items to characterize the list that they belong to, and then aggregates these lists to estimate the user preference, naturally fitting into the hierarchical structure of item lists; and (ii) item and list consistency, through a novel self-attentive aggregation layer designed for capturing the consistency of neighboring items and lists to better model user preference. Through experiments over three real-world datasets reflecting different kinds of user-generated item lists, we find that AttList results in significant improvements in NDCG, [email protected], and [email protected] versus a suite of state-of-the-art baselines. Furthermore, all code and data are available at https://github.com/heyunh2015/AttList. Yun He 0001, Jianling Wang, Wei Niu 0003, James Caverlee |
CIKM | 1 |
| 2018 | Pseudo-Implicit Feedback for Alleviating Data Sparsity in Top-K RecommendationabstractWe propose PsiRec, a novel user preference propagation recommender that incorporates pseudo-implicit feedback for enriching the original sparse implicit feedback dataset. Three of the unique characteristics of PsiRec are: (i) it views user-item interactions as a bipartite graph and models pseudo-implicit feedback from this perspective; (ii) its random walks-based approach extracts graph structure information from this bipartite graph, toward estimating pseudo-implicit feedback; and (iii) it adopts a Skip-gram inspired measure of confidence in pseudo-implicit feedback that captures the pointwise mutual information between users and items. This pseudo-implicit feedback is ultimately incorporated into a new latent factor model to estimate user preference in cases of extreme sparsity. PsiRec results in improvements of 21.5% and 22.7% in terms of Precision@10 and Recall@10 over state-of-the-art Collaborative Denoising Auto-Encoders. Our implementation is available at https://github.com/heyunh2015/PsiRecICDM2018. Yun He 0001, Haochen Chen, Ziwei Zhu 0001, James Caverlee |
ICDM | 1 |