Jing Yi

dblp:06/8462 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Disentangling User Interest and Geographical Context for POI Recommendations
abstract
POI recommendation plays an important role in many applications, such as mobility prediction and location-based advertisements. Existing POI recommendation methods mainly capture the observed patterns in user visits for recommendations, without a comprehensive consideration of the underlying reasons behind the visits. Therefore, different causes of a visit, i.e., users’ interest and geographical context, are entangled. When the underlying causes change (e.g., when a user moves to a new place), the robustness of the recommendations cannot be guaranteed. To address the above challenges, we propose DUIG, a novel user interest and geographical influences disentanglement framework for POI recommendations. We first design a personalized disentanglement strategy to divide check-ins through geographical influence. Specifically, the colliding effect of causality is leveraged to the divide cause-specific check-ins, such that user interest and geographical influence can be properly disentangled in user and POI embeddings. Through this mechanism, even if the underlying reasons that affect a user’s preference change, intervention can be conducted upon the causes to make recommendations generalized to the new scenario. In addition, a geographical-aware negative sampling strategy is proposed to utilize hard negatives to regularize the embedding and disentanglement in the latent space, where a larger sampling probability is introduced for negative samples containing more geographic information. Extensive experiments on two real-world POI recommendation datasets demonstrate the superior performance of DUIG.
Wenhui Meng, Jiayi Xie, Jing Yi, Yaochen Zhu, Zhenzhong Chen 0001
ACM Trans. Intell. Syst. Technol.3
2024 Meta-path aware dynamic graph learning for friend recommendation with user mobility
Ding Ding 0004, Jing Yi, Jiayi Xie, Zhenzhong Chen 0001
Inf. Sci.2
2024 Exploring Universal Intrinsic Task Subspace for Few-Shot Learning via Prompt Tuning
abstract
Why can pre-trained language models (PLMs) learn universal representations and effectively adapt to broad NLP tasks differing a lot superficially? In this work, we empirically find evidence indicating that the adaptations of PLMs to various few-shot tasks can be reparameterized as optimizing only a few free parameters in a unified low-dimensionalintrinsic task subspace, which may help us understand why PLMs could easily adapt to various NLP tasks with small-scale data. To find such a subspace and examine its universality, we propose an analysis pipeline calledintrinsic prompt tuning(IPT). Specifically, we resort to the recent success of prompt tuning and decompose the soft prompts of multiple NLP tasks into the same low-dimensional nonlinear subspace, then we learn to adapt the PLM to unseen data or tasks by only tuning parameters in this subspace. In the experiments, we study diverse few-shot NLP tasks and surprisingly find that in a 250-dimensional subspace found with 100 tasks, by only tuning 250 free parameters, we can recover 97% and 83% of the full prompt tuning performance for 100 seen tasks (using different training data) and 20 unseen tasks, respectively, showing great generalization ability of the found intrinsic task subspace. Besides being an analysis tool, IPTcould further help us improve the prompt tuning stability.
Yujia Qin, Xiaozhi Wang, Yusheng Su, Yankai Lin 0001, Ning Ding 0002, Jing Yi, Weize Chen, Zhiyuan Liu 0001, Juan-Zi Li, Lei Hou 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 Deconfounded Cross-modal Matching for Content-based Micro-video Background Music Recommendation
abstract
Object-oriented micro-video background music recommendation is a complicated task where the matching degree between videos and background music is a major issue. However, music selections in user-generated content (UGC) are prone to selection bias caused by historical preferences of uploaders. Since historical preferences are not fully reliable and may reflect obsolete behaviors, over-reliance on them should be avoided as knowledge and interests dynamically evolve. In this article, we propose a Deconfounded Cross-Modal matching model to mitigate such bias. Specifically, uploaders’ personal preferences of music genres are identified as confounders that spuriously correlate music embeddings and background music selections, causing the learned system to over-recommend music from majority groups. To resolve such confounders, backdoor adjustment is utilized to deconfound the spurious correlation between music embeddings and prediction scores. We further utilize Monte Carlo estimator with batch-level average as the approximations to avoid integrating the entire confounder space calculated by the adjustment. Furthermore, we design a teacher–student network to utilize the matching of music videos, which is professionally generated content (PGC) with specialized matching, to better recommend content-matching background music. The PGC data are modeled by a teacher network to guide the matching of uploader-selected UGC data of student network by Kullback–Leibler–based knowledge transfer. Extensive experiments on the TT-150k-genre dataset demonstrate the effectiveness of the proposed method. The code is publicly available on https://github.com/jing-1/DecCM
Jing Yi, Zhenzhong Chen 0001
ACM Trans. Intell. Syst. Technol.1
2024 Deep Causal Reasoning for Recommendations
abstract
Traditional recommender systems aim to estimate a user’s rating to an item based on observed ratings from the population. As with all observational studies, hidden confounders, which are factors that affect both item exposures and user ratings, lead to a systematic bias in the estimation. Consequently, causal inference has been introduced in recommendations to address the influence of unobserved confounders. Observing that confounders in recommendations are usually shared among items and are therefore multi-cause confounders, we model the recommendation as a multi-cause multi-outcome (MCMO) inference problem. Specifically, to remedy the confounding bias, we estimate user-specific latent variables that render the item exposures independent Bernoulli trials. The generative distribution is parameterized by a DNN with factorized logistic likelihood and the intractable posteriors are estimated by variational inference. Controlling these factors as substitute confounders, under mild assumptions, can eliminate the bias incurred by multi-cause confounders. Furthermore, we show that MCMO modeling may lead to high variance due to scarce observations associated with the high-dimensional treatment space. Therefore, we theoretically demonstrate that controlling user features as pre-treatment variables can substantially improve sample efficiency and alleviate overfitting. Empirical studies on both simulated and real-world datasets demonstrate that the proposed deep causal recommender shows more robustness to unobserved confounders than state-of-the-art causal recommenders. Codes and datasets are released at https://github.com/yaochenzhu/Deep-Deconf.
Yaochen Zhu, Jing Yi, Jiayi Xie, Zhenzhong Chen 0001
ACM Trans. Intell. Syst. Technol.2
2024 Variational Mixture of Stochastic Experts Auto-Encoder for Multi-Modal Recommendation
abstract
Multi-modal data presents a promising opportunity for improving multimedia recommendation models, but it also introduces task-irrelevant noise that can reduce model robustness. In this paper, we propose a robust multi-modal recommendation approach that accounts for different levels of task-irrelevant noise across modalities. We explicitly consider the uncertainty associated with each modality and perform stochastic sampling-based fusion according to the precision of different modalsities, which serves as a measure of uncertainty. The influence of noisy modalities with high uncertainty is removed, filtering out task-irrelevant noise, and therefore a noise-robust multi-modal recommendation is achieved. Moreover, the stochastic sampling strategy intrinsically considers and simulates scenarios with absent modalities during multi-modal fusion. Consequently, it incorporates additional randomness into the training process, which enables the model to handle the problem of modality missing. Furthermore, the proposed fusion approach integrates the noise robustness of the Product-of-Experts (PoE) framework when modeling with Gaussian distributions, along with the flexibility of the Mixture-of-Experts (MoE) technique to represent diverse distributions of latent variables. This integration allows the proposed approach to achieve noise-robust modeling with non-Gaussian variables. Specifically, we derive a solvable evidence lower bound for the proposed variational mixture of stochastic experts (VMoSE) auto-encoder, where both Gaussian and Student-T distributions are used to model the latent variables. Constraints are added to match the similarities between the ID embeddings and the multi-modal joint embeddings by utilizing an Expectation maximization (EM)-style algorithm for better model optimization. Extensive experiments demonstrate the effectiveness of the proposed method in multi-modal fusion and the robustness to modality noise and modality missing.
Jing Yi, Zhenzhong Chen 0001
IEEE Trans. Multim.1
2023 A Coordination Control Strategy of HVDC with Interline DCPFC: Hierarchical Structure and Parameter Design
abstract
Multi-terminal DC transmission technology has complex power flow control problems, which require the introduction of a DC power flow controller(DCPFC) to improve the freedom of control. In order to solve the insufficient power flow control ability of existing DCPFC, this article presents a hierarchical control strategy for a system consisting of a distributed control layer and a centralized control layer, to avoid problems caused by overload and uncontrolled power flow. For the distributed control layer, the small-signal model of multi-line capacitance full bridge DCPFC is established, which provides a design basis for the selection of closed-loop proportional and integral (PI) (proportional and integral) control parameters. For the centralized control layer, the deviation of DC voltage in virtual synchronous generator (VSC) terminal and the operating line current ratings are two optimization objectives. The optimal power flow solution obtained from the optimization problem are achieved by DCPFC which enables to regulate VSC terminal voltage. Therefore, line losses in the DC system can be minimized. This article creatively proposes a hierarchical control framework about DCPFC and VSC, further enhances the development potential of DCPFC.
Yurui Zhou, Pengfeng Lin, Jing Yi, Hongyi Zhang 0013, Chuanchuan Hou
IECON4
2023 Cross-Modal Variational Auto-Encoder for Content-Based Micro-Video Background Music Recommendation
abstract
In this paper, we propose a cross-modal variational auto-encoder (CMVAE) for content-based micro-video background music recommendation. CMVAE is a hierarchical Bayesian generative model that matches relevant background music to a micro-video by projecting these two multimodal inputs into a shared low-dimensional latent space, where the alignment of two corresponding embeddings of a matched video-music pair is achieved by cross-generation. Moreover, the multimodal information is fused by the product-of-experts (PoE) principle, where the semantic information in visual and textual modalities of the micro-video are weighted according to their variance estimations such that the modality with a lower noise level is given more weights. Therefore, the micro-video latent variables contain less irrelevant information that results in a more robust model generalization. Furthermore, we establish a large-scale content-based micro-video background music recommendation dataset, TT-150k, composed of extracted features from approximately 3,000 different background music clips associated to 150,000 micro-videos from different users. Extensive experiments on the established TT-150k dataset demonstrate the effectiveness of the proposed method. A qualitative assessment of CMVAE by visualizing some recommendation results is also included.
Jing Yi, Yaochen Zhu, Jiayi Xie, Zhenzhong Chen 0001
IEEE Trans. Multim.1
2023 Multi-auxiliary Augmented Collaborative Variational Auto-encoder for Tag Recommendation
abstract
Recommending appropriate tags to items can facilitate content organization, retrieval, consumption, and other applications, where hybrid tag recommender systems have been utilized to integrate collaborative information and content information for better recommendations. In this article, we propose a multi-auxiliary augmented collaborative variational auto-encoder (MA-CVAE) for tag recommendation, which couples item collaborative information and item multi-auxiliary information, i.e., content and social graph, by defining a generative process. Specifically, the model learns deep latent embeddings from different item auxiliary information using variational auto-encoders (VAE), which could form a generative distribution over each auxiliary information by introducing a latent variable parameterized by deep neural network. Moreover, to recommend tags for new items, item multi-auxiliary latent embeddings are utilized as a surrogate through the item decoder for predicting recommendation probabilities of each tag, where reconstruction losses are added in the training phase to constrain the generation for feedback predictions via different auxiliary embeddings. In addition, an inductive variational graph auto-encoder is designed to infer latent embeddings of new items in the test phase, such that item social information could be exploited for new items. Extensive experiments on MovieLens and citeulike datasets demonstrate the effectiveness of our method.
Jing Yi, Xubin Ren, Zhenzhong Chen 0001
ACM Trans. Inf. Syst.1
2022 QuoteR: A Benchmark of Quote Recommendation for Writing
abstract
It is very common to use quotations (quotes) to make our writings more elegant or convincing.To help people find appropriate quotes efficiently, the task of quote recommendation is presented, aiming to recommend quotes that fit the current context of writing.There have been various quote recommendation approaches, but they are evaluated on different unpublished datasets.To facilitate the research on this task, we build a large and fully open quote recommendation dataset called QuoteR, which comprises three parts including English, standard Chinese and classical Chinese.Any part of it is larger than previous unpublished counterparts.We conduct an extensive evaluation of existing quote recommendation methods on QuoteR.Furthermore, we propose a new quote recommendation model that significantly outperforms previous methods on all three parts of QuoteR.All the code and data of this paper can be obtained at https://github.com/ thunlp/QuoteR.
Fanchao Qi, Yanhui Yang, Jing Yi, Zhili Cheng, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)3
2022 Exploring Mode Connectivity for Pre-trained Language Models
abstract
Recent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP.From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found.Although plenty of works have studied how to effectively and efficiently adapt PLMs to high-performance minima, little is known about the connection of various minima reached under different adaptation configurations.In this paper, we investigate the geometric connections of different minima through the lens of mode connectivity, which measures whether two minima can be connected with a low-loss path.We conduct empirical analyses to investigate three questions: (1) how could hyperparameters, specific tuning methods, and training data affect PLM's mode connectivity?(2) How does mode connectivity change during pretraining?(3) How does the PLM's task knowledge change along the path connecting two minima?In general, exploring the mode connectivity of PLMs conduces to understanding the geometric connection of different minima, which may help us fathom the inner workings of PLM downstream adaptation.The codes are publicly available at https://github.com/ thunlp/Mode-Connectivity-PLM.
Yujia Qin, Cheng Qian 0008, Jing Yi, Weize Chen, Yankai Lin 0001, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016
EMNLP3
2022 Knowledge Inheritance for Pre-trained Language Models
abstract
Yujia Qin, Yankai Lin, Jing Yi, Jiajie Zhang, Xu Han, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yujia Qin, Yankai Lin 0001, Jing Yi, Xu Han 0007, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
NAACL-HLT3
2022 Multi-Modal Variational Graph Auto-Encoder for Recommendation Systems
abstract
Graph embedding based methods have been used in recommendation systems recently, owing to their advances in modeling nodes as embeddings in a low-dimensional space. By effective neighborhood aggregation, graph convolutional networks can exploit high-order connections of neighbors such that the learned embeddings could be more informative thus improve the recommendation performance. However, user and item representations learned by graph aggregation inherently contain uncertainty due to sparsity of user-item interactions and noise of item features. To address these challenges, we propose a multi-modal variational graph auto-encoder (MVGAE) method. Specifically, we design modality-specific variational encoders that learn a Gaussian variable for each node whereas the mean vector represents semantic information and the variance vector denotes the noise level of the corresponding modality. Moreover, with the conditional independence assumption, the modality-specific Gaussian node embeddings are fused according to the product-of-experts principle, where the semantic information in each modality is weighted based on the estimated uncertainty level. Extensive experiments on three public datasets, Amazon Movies, Amazon Electronics and AliShop-7 C, demonstrate that our proposed method achieves competitive performance when compared with the state-of-the-art algorithms.
Jing Yi, Zhenzhong Chen 0001
IEEE Trans. Multim.1
2020 An Improved Triple Interline DC Power Flow Controller for Bidirectional Power Control
abstract
For the operational safety and stability of electrical system, effective power flow control in meshed high voltage direct current (HVDC) grids makes sense. In order to realize the objective of controlling several lines simultaneously and independently in a convenient way, interline dc power flow controller (IDCPFC) has been introduced. To cover some shortages of present IDCPFC topologies, an improved triple interline dc power flow controller (TI-DCPFC), capable of controlling power flow of two lines actively, has been proposed in this paper. To explain it thoroughly and explicitly, its operational principle is illustrated combined with the mathematic equations firstly. Then, to validate its capacity of dealing with occasions of bidirectional flows and the dual-freedom control function, the theoretical control strategy and simulation results in a typical working mode are referred to comprehensively.
Jing Yi, Xu Zhong
TENCON1
2020 A Multimodal Variational Encoder-Decoder Framework for Micro-video Popularity Prediction
abstract
Predicting the popularity of a micro-video is a challenging task, due to a number of factors impacting the distribution such as the diversity of the video content and user interests, complex online interactions, etc. In this paper, we propose a multimodal variational encoder-decoder (MMVED) framework that considers the uncertain factors as the randomness for the mapping from the multimodal features to the popularity. Specifically, the MMVED first encodes features from multiple modalities in the observation space into latent representations and learns their probability distributions based on variational inference, where only relevant features in the input modalities can be extracted into the latent representations. Then, the modality-specific hidden representations are fused through Bayesian reasoning such that the complementary information from all modalities is well utilized. Finally, a temporal decoder implemented as a recurrent neural network is designed to predict the popularity sequence of a certain micro-video. Experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed model in the micro-video popularity prediction task.
Jiayi Xie, Yaochen Zhu, Jing Yi, Yaosi Hu, Hongyi Liu 0003, Zhenzhong Chen 0001
WWW5
2019 Predicting gene-disease associations from the heterogeneous network using graph embedding
abstract
The discovery of gene-disease associations is important for the prevention, diagnosis and treatment of diseases. The studies on gene-disease associations have produced diverse data, which can facilitate the gene-disease association prediction. Integrating diverse information is critical for developing high-accuracy prediction models. In this paper, we propose a heterogeneous network-based method that enhances gene-disease association prediction by using graph embedding and ensemble learning, abbreviated as “HNEEM”. A heterogeneous network is constructed based on gene-disease associations, gene-chemical associations and disease-chemical associations, to combine diverse information. The network uses genes, diseases and chemicals as nodes, and uses their associations as edges. The graph embedding methods are utilized to extract representation vectors of nodes in the heterogeneous network, and the feature vectors of genes and diseases are merged to represent gene-disease pairs, and the random forest is employed to build the prediction model based on gene-disease pairs. We consider six types of graph embedding methods, and take the individual graph embedding method-generated features to build prediction models and use them as base predictors, and then combine base predictors to develop the ensemble learning model HNEEM. We comprehensively compare different graph embedding methods, and results demonstrate that the graph embedding methods produce satisfying results in the gene-disease association prediction, and integrating different graph embedding methods can make further improvements. In computational experiments, HNEEM produces better results compared to the state-of-the-art gene-disease perdition methods, and HNEEM is robust to the data richness as well. Moreover, the usefulness of the proposed method HNEEM is validated by the case studies. In conclusion, HNEEM is a promising method for predicting gene-disease associations.
Xiaochan Wang, Yuchong Gong, Jing Yi, Wen Zhang 0008
BIBM3
2019 A embedding model for text classification
abstract
Abstract Existing word embeddings learning algorithms only employ the contexts of words, but different text documents use words and their relevant parts of speech very differently. Based on the preceding assumption, in order to obtain appropriate word embeddings and further improve the effect of text classification, this paper studies in depth a representation of words combined with their parts of speech. First, using the parts of speech and context of words, a more expressive word embeddings can be obtained. Further, to improve the efficiency of look‐up tables, we construct a two‐dimensional table that is in the format to represent words in text documents. Finally, the two‐dimensional table and a Bayesian theorem are used for text classification. Experimental results show that our model has achieved more desirable results on standard data sets. And it has more preferable versatility and portability than alternative models.
Peiyu Liu 0001, Yuzhen Yang, Jing Yi, Zhenfang Zhu
Expert Syst. J. Knowl. Eng.4
2018 A Sentence Similarity Model Based on Word Embeddings and Dependency Syntax-Tree
Peiyu Liu 0001, Jing Yi, Yuzhen Yang, Weitong Liu
ICONIP (3)3
2018 A New Square Grid-Based Method for Identifying Solar Activities
abstract
The FITS is widely used as a common file format in astronomy. However, how to detect an effective celestial event from large-scale astronomical images with complex background is a thorny issue. This paper proposes a square grid structure for target detection (GSTD) of solar activities, inncluding the separation of the target area and the deletion of the background area. This kind of processing can greatly reduce the storage cost of the data set, while the solar activity area can be well preserved. The solar image is operated through an equally sized square grid to separate the solar activity target area and the background area to speed up the processing of images, improve processing accuracy, and effectively prevent image noise interference. The experiment results have shown that this method delivers satisfactory performance in accuracy and time-cost. Through the anti-jamming processing of image noise, the accurate positioning and effective recognition of solar activities are realized. This method has been proved to achieve good image segmentation recognition of solar phenomenon in the research of various types of solar activities. Moreover, it can provide a feasible way to reduce the storage occupancy of Content-Based Image Retrieval (CBIR) systems.
Weijiang Li, Jing Yi
Int. J. Pattern Recognit. Artif. Intell.2
2010 Surface area estimation of digitized 3D objects using quasi-Monte Carlo methods
Yu-Shen Liu, Jing Yi, Guo-Qin Zheng, Jean-Claude Paul
Pattern Recognit.2