Cairong Yan

dblp:96/6397 · DBLP profile ↗
← Back
55ranked-venue papers
27as first author
45since 2021 · last 2026
0000-0003-0313-8833ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 14 first-author · 26 since 2021Databases, data management, data science and information retrieval · 12 · 9 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Multi-armed bandits in recommender systems: advances, challenges, and future prospects
Cairong Yan, Jiaxin Nan, Zijian Wang 0010, Yongquan Wan
Knowl. Inf. Syst.1
2026 HASNN: Hierarchical attention spiking neural network for dynamic graph representation learning
Yanglan Gan, Yanzu Dong, Cairong Yan, Guobing Zou
Knowl. Based Syst.4
2026 HCMAF: Hierarchical Feature Aggregation and Cross-Modal Attention Fusion Framework for Multi-Omics Patient Classification
abstract
The accumulation of large-scale multi-omics datasets has brought new opportunities for precise disease treatment. However, the inherent complexity of inter- and intra-omics relationships presents considerable obstacles to the precise integration of multi-omics data. Here, we propose a Hierarchical Feature Aggregation and Cross-Modal Attention Fusion (HCMAF) framework to integrate multi-omics data for patient classification and biomarker identification. Specifically, to capture both the specific information inherent in each omics data and complex cross-omics interactions, HCMAF incorporates three innovative modules. The hierarchical feature aggregation graph attention (HGAT) module captures intra-omics topological features through adaptive neighborhood aggregation. The cross-modal attention (CMA) module pinpoints inter-omics complementarity by modeling cross-omics dependencies. Finally, the confidence-driven multi-omics fusion (CMF) module dynamically integrates omics-specific predictions through learnable reliability weights. Comprehensive experiments on four public benchmark datasets show that HCMAF achieves better classification performance and consistently surpasses leading existing methods. Further component analysis confirms the crucial role of the HGAT, CMA and CMF modules in ensuring overall model effectiveness.
Yanglan Gan, Hangkai Zhao, Cairong Yan, Guobing Zou
IEEE J. Biomed. Health Informatics4
2025 Compensating Information and Capturing Modal Preferences in Multimodal Recommendation: A Dual-Path Representation Learning Framework
abstract
In the context of information explosion, multimodal recommender systems (MMRS) have demonstrated great potential in capturing users' complex preferences and enhancing recommendation performance by integrating multimodal data such as images and text. However, multimodal data inherently suffers from semantic inconsistency, which can introduce information conflicts or noise. Moreover, users' reliance on different modalities varies dynamically with context and time (multimodal dynamic preferences). These challenges may lead to truth deviation and deep semantic mismatch, ultimately degrading recommendation performance. To address these issues, we propose Dual-Path Multimodal Recommendation (DPRec), a novel model that improves precision and robustness of the recommendation through cross-modal information compensation and dynamic modal preference learning. Specifically, DPRec first employs a cross-modal attention mechanism to dynamically model inter-modal correlations, effectively exploring complementary and shared features for robust user and item representations. Second, it integrates feature projection, modality alignment, and dynamic weighting mechanisms to adaptively adjust modality importance based on user context, ensuring flexibility in handling preference dynamics. Lastly, a modality contrastive loss is utilized to maximize mutual information between modalities, mitigating semantic mismatch by enhancing deep collaborative representations. Extensive experiments on three public datasets show that DPRec consistently outperforms state-of-the-art (SOTA) methods, achieving average improvements of 3.94% in Recall@20 and 3.84% in NDCG@20. Our code is publicly available at: https://anonymous.4open.science/r/DPRec-4D15.
Cairong Yan, Xubin Mao, Zijian Wang 0010, Xicheng Zhao, Linlin Meng
CIKM1
2025 CoCoB: Adaptive Collaborative Combinatorial Bandits for Online Recommendation
Cairong Yan, Jinyi Han, Jin Ju, Yanting Zhang 0001, Zijian Wang 0010, Xuan Shao
DASFAA (5)1
2025 KG-TS: Knowledge Graph-Driven Thompson Sampling for Online Recommendation
Cairong Yan, Hualu Xu, Yanting Zhang 0001, Zijian Wang 0010, Xuan Shao
DASFAA (5)1
2025 Dynamic Bidirectional Attentional Mamba Model for EEG-Based Motor Imagery Classification
Qianzi Shen, Zijian Wang 0010, Yanting Zhang 0001, Cairong Yan
ICIC (27)5
2025 A Virtual Camera Assisted Surround-View SLAM System for Robust Parking
Xuan Shao, Feiyang Lu, Cairong Yan
ICIC (14)3
2025 MCCVM: Multi-Scale Cross-axes Conv-VMamba for Medical Image Classification
abstract
The performance of medical image classification relies on the effective capture and balance of local and global features. Although hybrid models combining Convolutional Neural Networks (CNNs) and Transformers have achieved notable success, they face two critical challenges: insufficient cross-axes information modeling, which hampers spatial dependency capture, and difficulty dynamically balancing within multi-scale features essential for complex medical images. In addition, the quadratic complexity of the self-attention mechanism in Transformers limits efficiency on high-resolution imaging data. To tackle these issues, this paper proposes the Multi-scale Cross-axes Convolutional VMamba Model (MCCVM), which combines CNNs for local feature extraction, the Mamba model for efficient global feature processing, and novel mechanisms for cross-axes information modeling and feature fusion. The MCCVM incorporates a Convolutional VMamba Fusion Block (CMF), which replaces Transformers with the Mamba framework to enhance computational efficiency while maintaining global feature extraction. A Cross-axes Attention SSM Block (CASSM) is also introduced within the Mamba structure to better model cross-axes spatial dependencies. Finally, a channel-based deep convolutional gated feature fusion network (CGFFN) is employed to dynamically balance local and global features, ensuring a comprehensive representation of medical images. Extensive experiments on medical image datasets demonstrate the superiority and effectiveness of our MCCVM.
Zijian Wang 0010, Qianzi Shen, Yanting Zhang 0001, Cairong Yan
IJCNN5
2025 From spatial to semantic: attribute-aware fashion similarity learning via iterative positioning and attribute diverging
Yongquan Wan, Jianfei Zheng, Cairong Yan, Guobing Zou
Appl. Intell.3
2025 Precise spiking neurons for fitting any activation function in ANN-to-SNN Conversion
Qianzi Shen, Xuhang Li, Yanting Zhang 0001, Zijian Wang 0010, Cairong Yan
Appl. Intell.6
2025 cd-MBRec: Enhancing multi-behavior recommendation by explicitly modeling commonality and diversity
abstract
Multi-behavior recommendation models excel in extracting abundant information from user-item interactions to enhance performance; however, they encounter challenges in accuracy due to noise disturbance and ambiguous weight allocation. In this paper, we propose cd-MBRec, a novel model designed to amplify commonality among various behaviors, thereby minimizing noise interference while preserving behavior diversity to highlight semantic variations in feedback across distinct scenarios. Specifically, the model begins by constructing behavior matrices that models separate behaviors, along with an interaction matrix offering a broad overview of user behaviors. It employs graph neural networks to extract higher-order semantic and structural information from input data. Concurrently, the model integrates principles of Weber-Fechner Law for the adaptive allocation of initial weights to the multiple behaviors and utilizes matrix factorization techniques for efficient behavior embedding. Extensive experiments on two real-world datasets demonstrate that cd-MBRec surpasses existing state-of-the-art models in recommendation performance, achieving notable average improvements of 4.96% in HR@10 and 7.75% in NDCG@10.
Cairong Yan, Ziyang Zhu, Xiaopeng Guan, Yongquan Wan
Intell. Data Anal.1
2025 Inferring single-cell trajectories via critical cell identification using graph centrality algorithm
Yanglan Gan, Jiaqi Chu, Cairong Yan, Guobing Zou
Neurocomputing4
2025 Pose-Guided Transformer for Fine-Grained Action Quality Assessment
abstract
Action Quality Assessment (AQA) is a task aimed at automatically and fairly evaluating the level of movement execution, which holds significant importance for action understanding. Previous methods, while adept at extracting video features, often neglect human regions. This leads to a limited capability to discern subtle action differences and results in a lack of interpretative depth. In this work, we propose a Pose-Guided Transformer framework, termed PGT, for assessing action quality more accurately. Essentially, this framework incorporates pose information to augment human region features during video feature extraction. The PGT framework incorporates two critical modules: a pose-guided attention layer and a global-local feature extractor. The former is designed to isolate body-specific features, effectively minimizing background noise, while the latter further delineates fine-grained features by utilizing decomposed information from various human body parts. The proposed PGT achieves significant results on various challenging AQA benchmarks. Notably, on MTL-AQA dataset, with a Spearman’s rank correlation of 0.9630. Additionally, on the AQA-7 dataset, our approach achieves an average Spearman’s rank correlation of 0.8673, further validating the effectiveness of our method. These findings demonstrate that our framework excels in the task of action quality assessment, providing a viable solution for accurate and fair evaluation of movement execution.
Yanting Zhang 0001, Wenhao Chai, Cairong Yan, Wenhai Wang, Gaoang Wang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Dual-Path Multimodal Optimal Transport for Composed Image Retrieval
Cairong Yan, Yanting Zhang 0001, Yongquan Wan
ACCV (6)1
2024 TAN: A Tripartite Alignment Network Enhancing Composed Image Retrieval with Momentum Distillation
abstract
Composed image retrieval is designed to more accurately retrieve target images that align with user intentions by using a combination of reference images and descriptive modification texts. However, existing methods primarily focus on designing complex feature fusion networks while neglecting the prevalent issues of noise and inconsistent sample quality in training data, leading to insufficient cross-modal semantic alignment and sample relevance modeling. To address this, we propose an innovative Tripartite Alignment Network (TAN) that introduces a momentum distillation mechanism, leveraging the historical knowledge of a teacher network as additional super-vision to guide the optimization of the student network. During the feature encoder fine-tuning stage, we design response-based knowledge distillation and feature-based knowledge distillation techniques, explicitly strengthening modal alignment through composed-target contrastive learning and implicitly promoting modal fusion via composed-target matching learning. In the combiner training stage, we incorporate a lightweight combiner network and employ a cross-entropy-based matching loss function, encouraging high matching scores for relevant image-text pairs and low scores for irrelevant pairs. Extensive experiments on the FashionIQ and Shoes datasets demonstrate that TAN exhibits superior performance compared to existing state-of-the-art methods, with notable improvements in R@10 of +14.03% and +13.09%, respectively. These results affirm the effectiveness of momentum distillation in multimodal learning. Access the source code at https://github.com/Maserhe/TAN.
Yongquan Wan, Erhe Yang, Cairong Yan, Guobing Zou, Bofeng Zhang
ICDM3
2024 Taming Diffusion for Fashion Clothing Generation with Versatile Condition
Yanting Zhang 0001, Jingyi Guo, Cairong Yan, Zhijun Fang 0001
PRCV (5)3
2024 Inferring gene regulatory networks from single-cell transcriptomics based on graph embedding
abstract
MOTIVATION: Gene regulatory networks (GRNs) encode gene regulation in living organisms, and have become a critical tool to understand complex biological processes. However, due to the dynamic and complex nature of gene regulation, inferring GRNs from scRNA-seq data is still a challenging task. Existing computational methods usually focus on the close connections between genes, and ignore the global structure and distal regulatory relationships. RESULTS: In this study, we develop a supervised deep learning framework, IGEGRNS, to infer GRNs from scRNA-seq data based on graph embedding. In the framework, contextual information of genes is captured by GraphSAGE, which aggregates gene features and neighborhood structures to generate low-dimensional embedding for genes. Then, the k most influential nodes in the whole graph are filtered through Top-k pooling. Finally, potential regulatory relationships between genes are predicted by stacking CNNs. Compared with nine competing supervised and unsupervised methods, our method achieves better performance on six time-series scRNA-seq datasets. AVAILABILITY AND IMPLEMENTATION: Our method IGEGRNS is implemented in Python using the Pytorch machine learning library, and it is freely available at https://github.com/DHUDBlab/IGEGRNS.
Yanglan Gan, Jiacheng Yu, Cairong Yan, Guobing Zou
Bioinform.4
2024 MeFiNet: Modeling multi-semantic convolution-based feature interactions for CTR prediction
abstract
Extracting more information from feature interactions is essential to improve click-through rate (CTR) prediction accuracy. Although deep learning technology can help capture high-order feature interactions, the combination of features lacks interpretability. In this paper, we propose a multi-semantic feature interaction learning network (MeFiNet), which utilizes convolution operations to map feature interactions to multi-semantic spaces to improve their expressive ability and uses an improved Squeeze & Excitation method based on SENet to learn the importance of these interactions in different semantic spaces. The Squeeze operation helps to obtain the global importance distribution of semantic spaces, and the Excitation operation helps to dynamically re-assign the weights of semantic features so that both semantic diversity and feature diversity are considered in the model. The generated multi-semantic feature interactions are concatenated with the original feature embeddings and input into a deep learning network. Experiments on three public datasets demonstrate the effectiveness of the proposed model. Compared with state-of-the-art methods, the model achieves excellent performance (+0.18% in AUC and -0.34% in LogLoss VS DeepFM; +0.19% in AUC and -0.33% in LogLoss VS FiBiNet).
Cairong Yan, Xiaoke Li, Ran Tao 0005, Zhaohui Zhang 0001, Yongquan Wan
Intell. Data Anal.1
2024 Learning Attribute-guided Fashion Similarity with Spatial and Channel Attention
abstract
Fashion image retrieval is one of the important services of e-commerce platforms, and it is also the basis of various fashion-related AI applications. Studies have shown that in a multi-modal environment (images + attribute labels), embedding items into specific attribute spaces can support more fine-grained similarity measures, which is especially suitable for fashion retrieval tasks. In this paper, we propose an attention-based attribute-guided similarity learning network (AttnFashion) for fashion image retrieval. The core of this network is an attribute-guided spatial attention module and an attribute-guided channel attention module, which correspond to the mapping between attributes and image regions, and the mapping between attributes and high-level image semantics, respectively. To make these two modules interact deeply, we design a parallel structure that allows them to share attribute embeddings and guide each other to extract specific features, which also helps to reduce the network parameters of the attention modules. An adaptive feature fusion strategy is proposed to synthesise the features extracted by the two modules. Extensive experiments show that the proposed AttnFashion performs better than current competitive networks in the field of fine-grained attribute-based fashion retrieval.
Yongquan Wan, Cairong Yan, Bofeng Zhang
J. Exp. Theor. Artif. Intell.3
2023 Enhancing Session-Based Recommendation with Multi-granularity User Interest-Aware Graph Neural Networks
Cairong Yan, Xiangyang Feng, Yanglan Gan
CollaborateCom (3)1
2023 Thompson Sampling with Time-Varying Reward for Contextual Bandits
Cairong Yan, Hualu Xu, Haixia Han, Yanting Zhang 0001, Zijian Wang 0010
DASFAA (2)1
2023 TransLink: Transformer-Based Embedding for Tracklets' Global Link
abstract
Multi-object tracking (MOT) is essential to many tasks related to the smart transportation. Detecting and tracking humans on the road can give a vital feedback for either the moving vehicle or traffic control to ensure better driving safety and traffic flow. However, most trackers face a common problem of identity (ID) switch, resulting in an incomplete human trajectory prediction. In this paper, we propose a Transformer-based tracklet linking method called TransLink to mitigate the association failures. Specifically, the self-attention mechanism is well exploited to get the feature representation for tracklets, followed by a multilayer perceptron to predict the association likelihood, which can be further used in determining the tracklet association. Experiments on the MOT dataset demonstrate the effectiveness of the proposed module in lifting the tracking performances.
Yanting Zhang 0001, Shuanghong Wang, Yuxuan Fan, Gaoang Wang, Cairong Yan
ICASSP5
2023 MB-DP: A Multi-behavior Recommendation Model Integrating Dynamic Preferences
abstract
Multi-behavior recommendation has gained significant attention in recent years for its ability to outperform singlebehavior models.Current research related to multi-behavior models leaves room for improvement in the following two areas.First, the noise carried by individual behaviors and the additional noise generated during behavior processing is often overlooked, and these can ultimately degrade recommendation performance.Second, the specific time period of behavioral interactions and the frequency of interactions within that time period are also not taken into account.To address the above limitations, we propose a multibehavior recommendation model integrating dynamic preferences (MB-DP) that captures dynamic interests while smoothing and denoising multi-behavior information.MB-DP extracts low and high-order semantics from various behaviors and unifies the measurements to generate interaction predictions.Additionally, it analyzes the interaction time and frequency of each behavior using gated recurrent units to capture the dynamic preferences of users and improve the prediction values.Extensive experimental results on two real-world datasets show that MB-DP significantly improves recommendation performance compared to the state-ofthe-art baselines.
Cairong Yan, Xiaopeng Guan, Haixia Han, Zhaohui Zhang 0001
SEKE1
2023 DMFDDI: deep multimodal fusion for drug-drug interaction prediction
abstract
Drug combination therapy has gradually become a promising treatment strategy for complex or co-existing diseases. As drug-drug interactions (DDIs) may cause unexpected adverse drug reactions, DDI prediction is an important task in pharmacology and clinical applications. Recently, researchers have proposed several deep learning methods to predict DDIs. However, these methods mainly exploit the chemical or biological features of drugs, which is insufficient and limits the performances of DDI prediction. Here, we propose a new deep multimodal feature fusion framework for DDI prediction, DMFDDI, which fuses drug molecular graph, DDI network and the biochemical similarity features of drugs to predict DDIs. To fully extract drug molecular structure, we introduce an attention-gated graph neural network for capturing the global features of the molecular graph and the local features of each atom. A sparse graph convolution network is introduced to learn the topological structure information of the DDI network. In the multimodal feature fusion module, an attention mechanism is used to efficiently fuse different features. To validate the performance of DMFDDI, we compare it with 10 state-of-the-art methods. The comparison results demonstrate that DMFDDI achieves better performance in DDI prediction. Our method DMFDDI is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/DHUDEBLab/DMFDDI.git.
Yanglan Gan, Wenxiao Liu, Cairong Yan, Guobing Zou
Briefings Bioinform.4
2023 Predicting synergistic anticancer drug combination based on low-rank global attention mechanism and bilinear predictor
abstract
MOTIVATION: Drug combination therapy has exhibited remarkable therapeutic efficacy and has gradually become a promising clinical treatment strategy of complex diseases such as cancers. As the related databases keep expanding, computational methods based on deep learning model have become powerful tools to predict synergistic drug combinations. However, predicting effective synergistic drug combinations is still a challenge due to the high complexity of drug combinations, the lack of biological interpretability, and the large discrepancy in the response of drug combinations in vivo and in vitro biological systems. RESULTS: Here, we propose DGSSynADR, a new deep learning method based on global structured features of drugs and targets for predicting synergistic anticancer drug combinations. DGSSynADR constructs a heterogeneous graph by integrating the drug-drug, drug-target, protein-protein interactions and multi-omics data, utilizes a low-rank global attention (LRGA) model to perform global weighted aggregation of graph nodes and learn the global structured features of drugs and targets, and then feeds the embedded features into a bilinear predictor to predict the synergy scores of drug combinations in different cancer cell lines. Specifically, LRGA network brings better model generalization ability, and effectively reduces the complexity of graph computation. The bilinear predictor facilitates the dimension transformation of the features and fuses the feature representation of the two drugs to improve the prediction performance. The loss function Smooth L1 effectively avoids gradient explosion, contributing to better model convergence. To validate the performance of DGSSynADR, we compare it with seven competitive methods. The comparison results demonstrate that DGSSynADR achieves better performance. Meanwhile, the prediction of DGSSynADR is validated by previous findings in case studies. Furthermore, detailed ablation studies indicate that the one-hot coding drug feature, LRGA model and bilinear predictor play a key role in improving the prediction performance. AVAILABILITY AND IMPLEMENTATION: DGSSynADR is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/DHUDBlab/DGSSynADR.
Yanglan Gan, Cairong Yan, Guobing Zou
Bioinform.4
2023 Attribute-guided and attribute-manipulated similarity learning network for fashion image retrieval
abstract
Learning the similarity between fashion items is essential for many fashion-related tasks. Most methods based on global or local image similarity cannot meet the fine-grained retrieval requirements related to attributes. We are the first to clearly distinguish the concepts of attribute name and their values and divide fashion retrieval tasks that combine images and text into: attribute-guided retrieval and attribute-manipulated retrieval. We propose a hierarchical attribute-aware embedding network (HAEN) that takes images and attributes as input, learns multiple attribute-specific embedding spaces, and measures fine-grained similarity in the corresponding spaces. It can accurately map different attributes to the corresponding areas of the image, thereby facilitating the feature fusion of two different modalities of text and image, including enhancement and replacement. Then on this basis, we propose three attribute-manipulated similarity learning methods, HAEN_Avg, HAEN_Rec, and HAEN_Cmb. With comprehensive validation on two real-world fashion datasets, we demonstrate that our methods can effectively leverage semantic knowledge to improve image retrieval performance, including attribute-guided and attribute-manipulated retrieval tasks.
Yongquan Wan, Cairong Yan, Guobing Zou, Bofeng Zhang
Intell. Data Anal.2
2023 Enhancing Multi-Behavior Recommendations Through Capturing Dynamic Preferences
abstract
Multi-behavior recommendation has gained significant attention in recent years for its ability to outperform single-behavior models. Current research related to multi-behavior models leaves room for improvement in the following two areas. First, the noise carried by individual behaviors and the additional noise generated during behavior processing is often overlooked, and these can ultimately degrade recommendation performance. Second, the specific time period of behavioral interactions and the frequency of interactions within that time period are also not taken into account. To address the above limitations, we propose a multi-behavior recommendation model integrating dynamic preferences (MB-DP) that captures dynamic interests while smoothing and denoising multi-behavior information. MB-DP extracts low and high-order semantics from various behaviors and unifies the measurements to generate interaction predictions. Additionally, it analyzes the interaction time and frequency of each behavior using gated recurrent units to capture the dynamic preferences of users and improve the prediction values. Extensive experimental results on two real-world datasets show that MB-DP significantly improves recommendation performance compared to the state-of-the-art baselines.
Cairong Yan, Xiaopeng Guan, Haixia Han, Zhaohui Zhang 0001, Yanting Zhang 0001
Int. J. Softw. Eng. Knowl. Eng.1
2023 MIN: multi-dimensional interest network for click-through rate prediction
Cairong Yan, Xiaoke Li, Yanting Zhang 0001, Zijian Wang 0010, Yongquan Wan
Knowl. Inf. Syst.1
2023 Dual attention composition network for fashion image retrieval with attribute manipulation
Yongquan Wan, Guobing Zou, Cairong Yan, Bofeng Zhang
Neural Comput. Appl.3
2022 Attribute-Guided Fashion Image Retrieval by Iterative Similarity Learning
abstract
Image retrieval methods in the fashion field mainly take advantage of query images that reflect user needs, without considering additional keywords that users can provide to specify the attributes in their interests. To achieve the fine-grained fashion retrieval, we propose an iterative similarity learning network (ISLN) for attribute-guided image retrieval, which takes a query image and a specified attribute as input, and outputs other images with the same or similar attribute values. The core of the network is the iterative similarity learning module, which leverages the aggressive learning ability of the deep neural network (DNN) to focus on the area of interest and extract a more accurate feature embedding during the learning process of image and text semantic mapping. Extensive experiments on FashionAI and DARN (+8.33% and +10.73% in mAP) datasets show that ISLN performs better than the state-of-the-art methods in fine-grained similarity retrieval tasks.
Cairong Yan, Yanting Zhang 0001, Yongquan Wan, Dandan Zhu 0001
ICME1
2022 On-Road Pedestrian Tracking Across Multiple Moving Cameras
abstract
With the rapid development of autonomous driving, tracking on-road pedestrians raises more attention in the public. Currently, most researches focus on single camera based tracking or tracking across multiple static cameras. Tracking across multiple moving cameras has not been well studied yet. In this paper, we propose a workflow for tracking pedestrians across multiple moving cameras, leveraging the state-of-the-art single camera based tracking method of FairMOT. We consider different factors such as appearance features, motion information, and camera spatial distribution to improve the tracking performance. The experimental results carried on a multi-target multi-moving camera tracking dataset show the feasibility of the proposed scheme in solving the tracking issue in a complex environmental setting.
Yanting Zhang 0001, Shuanghong Wang, Qingxiang Wang, Qiubo Huang, Cairong Yan
ICME5
2022 Learning Image Representation via Attribute-Aware Attention Networks for Fashion Classification
Yongquan Wan, Cairong Yan, Bofeng Zhang, Guobing Zou
MMM (1)2
2022 JointCTR: a joint CTR prediction framework combining feature interaction and sequential behavior learning
Cairong Yan, Xiaoke Li, Yanting Zhang 0001
Appl. Intell.1
2022 Detection-by-tracking of traffic signs in videos
Yanting Zhang 0001, Zijian Wang 0010, Ruoning Song, Cairong Yan, Yonggang Qi
Appl. Intell.4
2022 Recurrent spiking neural network with dynamic presynaptic currents based on backpropagation
abstract
In recent years, spiking neural networks (SNNs), which originated from the theoretical basis of neuroscience, have attracted neuromorphic computing and brain-like computing due to their advantages, such as neural dynamics and coding mechanism, which are similar to biological neurons. SNNs have become one of the mainstream frameworks in the field of brain-like computing. However, most of the Leaky Integrate-and-Fire (LIF) neuron models currently used by SNNs based on direct training of backpropagation (BP) do not consider the changes in the recurrent connections and the dynamic strength of neuron connections over time. This study presented the LIF neuron model with recurrent connections and a method for dynamically changing the presynaptic currents. Recurrent LIF neurons have an additional cyclic connection compared with classic LIF neurons. Their postsynaptic current stimulates a change in membrane potential at the next time point. Their dynamics were more similar to the activities of biological neurons. We also proposed an efficient and flexible BP training method for recurrent LIF neurons. On the basis of the above methods, we proposed the recurrent SNN with dynamic presynaptic currents based on backpropagation (RDS-BP). We test the proposed RDS-BP on three image data sets (MNIST, Fashion-MNIST and CIFAR-10) and two text data sets (IMDB and TREC). The results showed that the performance of RDS-BP not only exceeded the naive SNN models based on BP but also exceeded the SNN methods proposed in previous studies in recent years, which had excellent performance in previous experiments. Our work provides a new LIF neuron model with a recurrent connection and dynamic presynaptic current and a BP training arrangement for the proposed neuron, which could merit developments with neuromorphic and brain-like computing.
Zijian Wang 0010, Yanting Zhang 0001, Haibo Shi, Lei Cao 0002, Cairong Yan
Int. J. Intell. Syst.5
2022 Dynamic clustering based contextual combinatorial multi-armed bandit for online recommendation
abstract
Recommender systems still face a trade-off between exploring new items to maximize user satisfaction and exploiting those already interacted with to match user interests. This problem is widely recognized as the exploration/exploitation (EE) dilemma, and the multi-armed bandit (MAB) algorithm has proven to be an effective solution. As the scale of users and items in real-world application scenarios increases, their purchase interactions become sparser. Then three issues need to be investigated when building MAB-based recommender systems. First, large-scale users and sparse interactions increase the difficulty of user preference mining. Second, traditional bandits model items as arms and cannot deal with ever-growing items effectively. Third, widely used Bernoulli-based reward mechanisms only feedback 0 or 1, ignoring rich implicit feedback such as behaviors like click and add-to-cart. To address these problems, we propose an algorithm named Dynamic Clustering based Contextual Combinatorial Multi-Armed Bandits (DC3MAB), which consists of three configurable key components. Specifically, a dynamic user clustering strategy enables different users in the same cluster to cooperate in estimating the expected rewards of arms. A dynamic item partitioning approach based on collaborative filtering significantly reduces the scale of arms and produces a recommendation list instead of one item to provide diversity. In addition, a multi-class reward mechanism based on fine-grained implicit feedback helps better capture user preferences. Extensive empirical experiments on three real-world datasets demonstrate the superiority of our proposed DC3MAB over state-of-the-art bandits (On average, +75.8% in F1 and +54.3% in cumulative reward). The source code is available at https://github.com/HaixHan/DC3MAB.
Cairong Yan, Haixia Han, Yanting Zhang 0001, Dandan Zhu 0001, Yongquan Wan
Knowl. Based Syst.1
2021 Learning Fashion Similarity Based on Hierarchical Attribute Embedding
abstract
Embedding items directly into a common feature space, and then measuring the similarity by calculating the feature distance in this space, has become the main method for similarity learning in current fashion retrieval tasks. The method is simple and efficient, but it ignores the correlation among fashion attributes and the impact of these correlations on the feature space, thereby reducing the accuracy of retrieval. Since the number of fashion attributes is large and the semantic granularity is also different, how to capture the relationship between fashion attributes and perform refined embedding to accurately represent fashion items is a challenge. In this paper, by constructing an attribute tree, we propose a hierarchical attribute embedding method for representing fashion items to enhance the relationship between attributes and use masking technology to disentangle different attributes. Based on these modules, we propose a hierarchical attribute-aware embedding network (HAEN) which takes images and attributes as input, learns multiple attribute-specific embedding spaces, and measures fine-grained similarity in the corresponding spaces. The extensive experimental result on two fashion-related public datasets FashionAI and DARN shows the superiority (+5.11% and +3.09% in MAP, respectively) of our proposed HAEN compared with state-of-the-art methods.
Cairong Yan, Anan Ding, Yanting Zhang 0001, Zijian Wang 0010
DSAA1
2021 Two-Phase Multi-armed Bandit for Online Recommendation
abstract
Personalized online recommendations strive to adapt their services to individual users by making use of both item and user information. Despite recent progress, the issue of balancing exploitation-exploration (EE) [1] remains challenging. In this paper, we model the personalized online recommendation of e-commence as a two-phase multi-armed bandit problem. This is the first time that “big arm” and “small arm” are introduced into multi-armed bandit (MAB), and a two-stage strategy is adopted to provide target users with the most suitable recommendation list. In the first phase, MAB is used to obtain an item subset that users may be interested in from a large number of items. We use item categories as arms instead of individual items in existing related models to control the arm scale and reduce computational complexity. In the second phase, we directly use the items generated in the first phase as arms of MAB and obtain rewards through fine-grained implicit feedback from users. Empirical studies on three real-world datasets show that our proposed method TPBandit performs better than state-of-the-art bandit-based recommendation methods in several evaluation metrics such as Precision, Recall, and Hit Ratio. Moreover, the two-phase method improves the recommendation performance by nearly 50% compared to the one-phase method in the best case.
Cairong Yan, Haixia Han, Zijian Wang 0010, Yanting Zhang 0001
DSAA1
2021 A Multi-Task Learning Approach for Recommendation based on Knowledge Graph
abstract
Sparsity and cold start problem are two classic problems of collaborative filtering. To alleviate these issues, researchers usually add side information to the recommendation models to boost the performance. In this paper, we propose a multi-task learning approach for recommendation based on knowledge graph (KGeRec), which takes recommendation as the main task and the knowledge graph as an auxiliary task to provide side information for recommendation. To fully capture the correlation information between these two tasks, a feature interaction layer (FlU) based on cross networks is designed to share features between them. Besides, a side information embedding layer (SIE) is also designed in the recommendation task to exploit more feature information. We apply KGeRec to three public datasets about movie, book, and music. Experimental results show that the proposed KGeRec outperforms the state-of-the-art approaches (+2.2% in AUC, +2.6% in Accuracy, +2.5% and in F1-score, compared to the maximum value in Type I models; +1.3% in AUC, +0.8% in Accuracy, and +2% in F1-score, compared to the maximum value in Type II models) and it performs well in sparse datasets. We also validate the effectiveness of knowledge graphs in improving recommendation performance.
Cairong Yan, Yanting Zhang 0001, Zijian Wang 0010, Pengwei Wang 0001
IJCNN1
2021 Modeling Long- and Short-Term User Behaviors for Sequential Recommendation with Deep Neural Networks
abstract
In e-commerce platforms, a user's next behavior will be affected by his long-term constant interests and short-term temporal needs. Such information is usually hidden in the users' historical online behavior data, so how to capture long-term and short-term patterns becomes the key to design better recommendation models or algorithms. Current mainstream methods such as Markov chain, convolutional neural network, and recurrent neural network cannot well express the mixed dynamic characteristics. In this paper, we propose an attention-based deep neural network (ADNNet) to solve the problem. In ADNNet, a convolutional neural network is used to extract the short-term patterns in the behavior sequences, and a gated recurrent unit is used to mine the long-term patterns in the behavior sequences. The attention mechanism is adopted to help the network automatically learn the best fusion coefficient of these two patterns. Our experimental result on four real public datasets (+0.69% in Hit Ratio and +3.49% in MRR) shows the superiority of our proposed ADNNet compared with other state-of-the-art methods.
Cairong Yan, Yanting Zhang 0001, Zijian Wang 0010, Pengwei Wang 0001
IJCNN1
2021 Modeling low- and high-order feature interactions with FM and self-attention network
Cairong Yan, Yongquan Wan, Pengwei Wang 0001
Appl. Intell.1
2021 Similarity-based sales forecasting using improved ConvLSTM and prophet
abstract
Sales forecasting is an important part of e-commerce and is critical to smart business decisions. The traditional forecasting methods mainly focus on building a forecasting model, training the model through historical data, and then using it to forecast future sales. Such methods are feasible and effective for the products with rich historical data while they are not performing as well for the newly listed products with little or no historical data. In this paper, with the idea of collaborative filtering, a similarity-based sales forecasting (S-SF) method is proposed. The implementation framework of S-SF includes three modules in order. The similarity module is responsible for generating top-k similar products of a given new product. We calculate the similarity based on two data types: time series data of sales and text data such as product attributes. In the learning module, we propose an attention-based ConvLSTM model which we called AttConvLSTM, and optimize its loss function with the convex function information entropy. Then AttConvLSTM is integrated with Facebook Prophet model to forecast top-k similar products sales based on their historical data. The prediction results of all top-k similar products will be fused in the forecasting module through operations of alignment and scaling to forecast the target products sales. The experimental results show that the proposed S-SF method can simultaneously adapt to the sales forecasting of mature products and new products, which shows excellent diversity, and the forecasting idea based on similar products improves the accuracy of sales forecasting.
Yongquan Wan, Cairong Yan, Bofeng Zhang
Intell. Data Anal.3
2021 Attribute interaction aware matrix factorization method for recommendation
abstract
Matrix factorization (MF) models are effective and easy to expand and are widely used in industry, such as rating prediction and item recommendation. The basic MF model is relatively simple. In practical applications, side information such as attributes or implicit feedback is often combined to improve accuracy by modifying the model and optimizing the algorithm. In this paper, we propose an attribute interaction-aware matrix factorization (AIMF) method for recommendation tasks. We partition the original rating matrix into different sub-matrices according to the attribute interactions, train each sub-matrix independently, and merge all the latent vectors to generate the final score. Since the generated sub-matrices vary in size, an adaptive regularization coefficient optimization strategy and an adaptive latent vector dimension optimization strategy are proposed for sub-matrix training, and a variety of latent vector merging methods are put forward. The method AIMF has two advantages. When the original rating matrix is particularly large, the training time complexity of the MF-based model becomes higher and the update cost of the model is also higher. In AIMF, because each sub-matrix is usually much smaller than the original rating matrix, the training time complexity is greatly reduced after using parallel computing technology. Secondly, in AIMF, it is not necessary to modify the matrix factorization model to incorporate attributes and their interactive information into the model to improve the performance. The experimental results on the two classic public datasets MovieLens 1M and MovieLens 100k show that AIMF can not only effectively improve the accuracy of recommendation, but also make full use of parallel computing technology to improve training efficiency without modifying the matrix factorization model.
Yongquan Wan, Lihua Zhu, Cairong Yan, Bofeng Zhang
Intell. Data Anal.3
2021 Modeling implicit feedback based on bandit learning for recommendation
Cairong Yan, Junli Xian, Yongquan Wan, Pengwei Wang 0001
Neurocomputing1
2020 IT-Block: Inverted Triangle Block embedded U-Net for Medical Image Segmentation
abstract
Convolutional neural network (CNN) such as U-Net has demonstrated excellent performance for medical image segmentation. However, there are some limitations of its scalability. Specifically, the memory would increase significantly when embedding other functional modules into U-Net. Moreover, the kernel size used in U-Net is unitary, which makes it difficult to obtain the multi-level information and extract the target completely. In this paper, we only use 12 convolutional layers of U-Net as a backbone and design a novel architecture named Inverted Triangle (IT) Block embedded into it to address these problems. The IT-Block consists of Dense Connection, Residual Connection, and Inception, aiming to help the network obtain multi-level features and reuse them comprehensively. Furthermore, we optimize the dice loss to alleviate the butterfly effect, making the training process more stable during the backpropagation. The experimental results state that our framework is superior to U-Net in running time and accuracy.
Cairong Yan
IJCNN3
2020 MIRD-Net for Medical Image Segmentation
Cairong Yan
PAKDD (2)3
2019 Merging visual features and temporal dynamics in sequential recommendation
Cairong Yan
Neurocomputing1
2018 A rapid detection algorithm of corrupted data in cloud storage
Zhifeng Sun, Cairong Yan, Yanglan Gan
J. Parallel Distributed Comput.3
2017 An Intelligent Field-Aware Factorization Machine Model
Cairong Yan
DASFAA (1)1
2016 Dynamic epigenetic mode analysis using spatial temporal clustering
abstract
BACKGROUND: Differentiation of human embryonic stem cells requires precise control of gene expression that depends on specific spatial and temporal epigenetic regulation. Recently available temporal epigenomic data derived from cellular differentiation processes provides an unprecedented opportunity for characterizing fundamental properties of epigenomic dynamics and revealing regulatory roles of epigenetic modifications. RESULTS: This paper presents a spatial temporal clustering approach, named STCluster, which exploits the temporal variation information of epigenomes to characterize dynamic epigenetic mode during cellular differentiation. This approach identifies significant spatial temporal patterns of epigenetic modifications along human embryonic stem cell differentiation and cluster regulatory sequences by their spatial temporal epigenetic patterns. CONCLUSIONS: The results show that this approach is effective in capturing epigenetic modification patterns associated with specific cell types. In addition, STCluster allows straightforward identification of coherent epigenetic modes in multiple cell types, indicating the ability in the establishment of the most conserved epigenetic signatures during cellular differentiation process.
Yanglan Gan, Han Tao, Guobing Zou, Cairong Yan, Jihong Guan
BMC Bioinform.4
2015 Eliminating the Redundancy in MapReduce-Based Entity Resolution
abstract
Entity resolution is the basic operation of data quality management, and the key step to find the value of data. The parallel data processing framework based on MapReduce can deal with the challenge brought by big data. However, there exist two important issues, avoiding redundant pairs led by the multi-pass blocking method and optimizing candidate pairs based on the transitive relations of similarity. In this paper, we propose a multi-signature based parallel entity resolution method, called multi-sig-er, which supports unstructured data and structured data. Two redundancy elimination strategies are adopted to prune the candidate pairs and reduce the number of similarity computation without affecting the resolution accuracy. Experimental results on real-world datasets show that our method tends to handle large datasets and it is more suitable for complex similarity computation than simple object matching.
Cairong Yan, Yalong Song
CCGRID1
2014 Hmfs: Efficient Support of Small Files Processing over HDFS
Cairong Yan, Yanglan Gan
ICA3PP (2)1
2012 IncMR: Incremental Data Processing Based on MapReduce
abstract
MapReduce programming model is widely used for large scale and one-time data-intensive distributed computing, but lacks flexibility and efficiency of processing small incremental data. IncMR framework is proposed in this paper for incrementally processing new data of a large data set, which takes state as implicit input and combines it with new data. Map tasks are created according to new splits instead of entire splits while reduce tasks fetch their inputs including the state and the intermediate results of new map tasks from designate nodes or local nodes. Data locality is considered as one of the main optimization means for job scheduling. It is implemented based on Hadoop, compatible with the original MapReduce interfaces and transparent to users. Experiments show that non-iterative algorithms running in MapReduce framework can be migrated to IncMR directly to get efficient incremental and continuous processing without any modification. IncMR is competitive and in all studied cases runs faster than that processing the entire data set.
Cairong Yan, Xin Yang 0006, Min Li 0012, Xiaolin Li 0001
IEEE CLOUD1
2012 Affinity-aware Virtual Cluster Optimization for MapReduce Applications
abstract
Infrastructure-as-a-Service clouds are becoming ubiquitous for provisioning virtual machines on demand. Cloud service providers expect to use least resources to deliver best services. As users frequently request virtual machines to build virtual clusters and run MapReduce-like jobs for big data processing, cloud service providers intend to place virtual machines closely to minimize network latency and subsequently reduce data movement cost. In this paper we focus on the virtual machine placement issue for provisioning virtual clusters with minimum network latency in clouds. We define distance as the latency between virtual machines and use it to measure the affinity of virtual clusters. Such metric of distance indicates the considerations of virtual machine placement and topology of physical nodes in clouds. Then we formulate our problem as the classical shortest distance problem and solve it by modeling to integer programming problem. A greedy virtual machine placement algorithm is designed to get a compact virtual cluster. Furthermore, an improved heuristic algorithm is also presented for achieving a global resource optimization. The simulation results verify our algorithms and the experiment results validate the improvement achieved by our approaches.
Cairong Yan, Ming Zhu 0004, Xin Yang 0006, Min Li 0012, Youqun Shi, Xiaolin Li 0001
CLUSTER1