Rong Gao 0001

dblp:79/518-1 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0001-7935-7173ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SDLK-Net: Enhanced squeezed directional large kernel multi-scale multi-modal fusion network for salient object detection
Lingyu Yan, Rong Gao 0001, Zengmao Wang, Zhiwei Ye, Xinyun Wu
Appl. Intell.3
2026 TSD-Rec: Metric semantic noise enhanced diffusion based contrastive learning with topology prior for multi-behavior recommendation
Rong Gao 0001, Yabo Guo, Yonghong Yu, Zhiwei Ye, Li Zhang 0013, Lingyu Yan
Expert Syst. Appl.1
2026 Wasserstein distance-based graph contrastive learning for recommendation
Yonghong Yu, Yujie Liao, Li Zhang 0013, Rong Gao 0001
Expert Syst. Appl.5
2025 Cluster search optimisation of deep neural networks for audio emotion classification
abstract
Automated patient monitoring solutions greatly benefit from audio emotion classification, although the considerable variance in individual expression and interpretation of emotions poses a challenge. Current approaches often employ standard Audio Spectrogram Transformer (AST) and deep learning models such as Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN)-based networks. However, their performance can be enhanced by integrating neural architecture search techniques using swarm optimisation algorithms. In this research, we explore AST with hyperparameter optimisation for speech emotion recognition. Three deep learning architectures with optimisable τ b -block structures and variable filter numbers, i.e. 1DCNN, bidirectional LSTM (BiLSTM) and CNN-BiLSTM, are also proposed, enabling the optimisation of network depth and width. A novel Cluster Search Optimisation (CSO) algorithm is introduced. It incorporates Cluster Centroid Search, a Cluster Distance Improvement metric and reinforcement learning to dispatch different search actions based on clustering convergence and Q -learning strategies, respectively. A novel Noise Tempered K-means (NTKM) clustering model is also proposed with the integration of Gaussian-based noise insertion and cluster compactness-separation measurement, to further fine-tune the cluster centriods obtained using OPTICS clustering. CSO is used for hyperparameter and architecture search for AST and aforementioned deep networks. Attention mechanisms are also integrated with CSO-optimised networks to further enhance feature learning. We evaluate the resulting models against those devised by other optimisation algorithms across the EMO-DB, SAVEE, and TESS datasets. The empirical results demonstrate that CSO-optimised AST and CNN-BiLSTM with attention mechanisms outperform other architectures and yield favourable comparison results against those from existing state-of-the-art audio emotion classification methods. • Evolving transformer and deep networks are devised for audio emotion recognition. • A Cluster Search Optimisation algorithm is proposed to adapt hyperparameters. • It incorporates Noise Tempered K-means clustering and Cluster Distance Improvement. • The Q-learning algorithm is used to optimise search behaviours. • Our study indicates CSO-optimised deep networks’ effectiveness across datasets.
Sam Slade, Li Zhang 0013, Houshyar Asadi, Chee Peng Lim, Yonghong Yu, Dezong Zhao, Arjun Panesar, Philip Fei Wu, Rong Gao 0001
Knowl. Based Syst.9
2025 Self-supervised extracted contrast network for facial expression recognition
Lingyu Yan, Jinquan Yang, Jinyao Xia, Rong Gao 0001, Li Zhang 0013, Yuan Yan Tang
Multim. Tools Appl.4
2025 CFI-Former: Efficient lane detection by multi-granularity perceptual query attention transformer
Rong Gao 0001, Siqi Hu, Lingyu Yan, Lefei Zhang, Jia Wu 0001
Neural Networks1
2025 Contrastive Translation With Dynamical Temperature for Sequential Recommendation
abstract
Contrastive learning is a promising solution to the problem of data sparsity in the field of recommendation system since it is able to extract self-supervised signals from raw data. The traditional contrastive learning-based sequential recommendation algorithms generate augmentations of original item sequences by utilizing crop, mask and reorder operations. However, those augmentation schemes destroy the underlying semantics of item sequences, resulting in difficulty in accurately defining positive and negative samples. To address this issue, we propose a contrastive translation based sequential recommendation algorithm, namely, CT4Rec. Specifically, CT4Rec generates augmented views of item sequences by injecting noises into embeddings of users and items, which is able to guarantee that the underlying semantics of augmented views are consistent with those of original item sequence. Hence, CT4Rec is able to effectively learn the invariances among the augmented views. In addition, the personalized translation operations are utilized to model the third-order relationships among entities. Moreover, it is difficult for contrastive learning-based recommendation algorithms with static temperature to simultaneously capture the differences among individual users/items and among the clusters of users/items. Hence, we utilize a dynamic temperature strategy to enhance CT4Rec, which endows CT4Rec with the capabilities of group-wise discrimination and instance discrimination. Our validation on five benchmark datasets shows that CT4Rec outperforms SOTA sequential recommendation methods. Our code is released athttps://github.com/zar123123/CT4Rec.
Aoran Zhang 0001, Yonghong Yu, Li Zhang 0013, Rong Gao 0001, Hongzhi Yin
IEEE Trans. Syst. Man Cybern. Syst.4
2024 Hyperbolic Adversarial Learning for Personalized Item Recommendation
Aoran Zhang 0001, Yonghong Yu, Gongyou Xu, Rong Gao 0001, Li Zhang 0013, Hongzhi Yin
DASFAA (3)4
2024 Contrastive graph learning long and short-term interests for POI recommendation
abstract
Modeling users’ short-term dynamic and long-term static interests to enhance Point-of-Interests (POI) recommendation performance has shown lots of advantages. Since users’ check-in records can be viewed as a graph network, methods based on Graph Neural Networks (GNNs) have recently shown promising applicability for POI recommendation. However, existing GNN-based works have the following shortcomings: (1) ignoring the impact of complex higher-order relationships between user-POI dynamics over time; and (2) ignoring the difference in POI importance that cannot effectively capture the imbalances of geographical influence among POIs. To address these challenges, we propose a novel Self-supervised Long-and Short-term model (SLS-REC) for POI recommendation. Specifically, we first design a spatio-temporal Hawkes attention hypergraph neural network to capture the spatial dependence and temporal evolution in users’ short-term dynamic interests. Then we introduce a dynamic propagation mechanism of GNNs to learn the geographic influences underlying geographic imbalances among POIs. In addition, the contrastive learning framework over a fine-grained node dropout strategy is applied to maximize the mutual information of long and short-term interest representations. Finally, we adaptively unify the recommendation and self-supervised task with an attention-based mechanism to optimize the proposed SLS-REC model for POI recommendation. Experiments on real-world datasets show that the proposed model significantly outperforms state-of-the-art methods.
Jia-Run Fu, Rong Gao 0001, Yonghong Yu, Jia Wu 0001, Jing Li 0055, Donghua Liu, Zhiwei Ye
Expert Syst. Appl.2
2024 Video Deepfake classification using particle swarm optimization-based evolving ensemble models
abstract
The recent breakthrough of deep learning based generative models has led to the escalated generation of photo-realistic synthetic videos with significant visual quality. Automated reliable detection of such forged videos requires the extraction of fine-grained discriminative spatial-temporal cues. To tackle such challenges, we propose weighted and evolving ensemble models comprising 3D Convolutional Neural Networks (CNNs) and CNN-Recurrent Neural Networks (RNNs) with Particle Swarm Optimization (PSO) based network topology and hyper-parameter optimization for video authenticity classification. A new PSO algorithm is proposed, which embeds Muller's method and fixed-point iteration based leader enhancement, reinforcement learning-based optimal search action selection, a petal spiral simulated search mechanism, and cross-breed elite signal generation based on adaptive geometric surfaces. The PSO variant optimizes the RNN topologies in CNN-RNN, as well as key learning configurations of 3D CNNs, with the attempt to extract effective discriminative spatial-temporal cues. Both weighted and evolving ensemble strategies are used for ensemble formulation with aforementioned optimized networks as base classifiers. In particular, the proposed PSO algorithm is used to identify optimal subsets of optimized base networks for dynamic ensemble generation to balance between ensemble complexity and performance. Evaluated using several well-known synthetic video datasets, our approach outperforms existing studies and various ensemble models devised by other search methods with statistical significance for video authenticity classification. The proposed PSO model also illustrates statistical superiority over a number of search methods for solving optimization problems pertaining to a variety of artificial landscapes with diverse geometrical layouts.
Li Zhang 0013, Dezong Zhao, Chee Peng Lim, Houshyar Asadi, Haoqian Huang, Yonghong Yu, Rong Gao 0001
Knowl. Based Syst.7
2024 Low-light image enhancement base on brightness attention mechanism generative adversarial networks
Jia-Run Fu, Lingyu Yan, Yulin Peng, Kunpeng Zheng, Rong Gao 0001
Multim. Tools Appl.5
2024 Multi-level contrastive graph learning for academic abnormality prediction
Yong Ouyang, Yuanlin Wang, Rong Gao 0001, Yawen Zeng, Jinhang Liu, Zhiwei Ye
Neural Comput. Appl.3
2024 Hyperbolic Translation-Based Sequential Recommendation
abstract
The goal of sequential recommendation algorithms is to predict personalized sequential behaviors of users (i.e., next-item recommendation). Learning representations of entities (i.e., users and items) from sparse interaction behaviors and capturing the relationships between entities are the main challenges for sequential recommendation. However, most sequential recommendation algorithms model relationships among entities in Euclidean space, where it is difficult to capture hierarchical relationships among entities. Moreover, most of them utilize independent components to model the user preferences and the sequential behaviors, ignoring the correlation between them. To simultaneously capture the hierarchical structure relationships and model the user preferences and the sequential behaviors in a unified framework, we propose a general hyperbolic translation-based sequential recommendation framework, namely HTSR. Specifically, we first measure the distance between entities in hyperbolic space. Then, we utilize personalized hyperbolic translation operations to model the third-order relationships among a user, his/her latest visited item, and the next item to consume. In addition, we instantiate two hyperbolic translation-based sequential recommendation models, namely Poincaré translation-based sequential recommendation (PoTSR) and Lorentzian translation-based sequential recommendation (LoTSR). PoTSR and LoTSR utilize the Poincaré distance and Lorentzian distance to measure similarities between entities, respectively. Moreover, we utilize the tangent space optimization method to determine optimal model parameters. Experimental results on five real-world datasets show that our proposed hyperbolic translation-based sequential recommendation methods outperform the state-of-the-art sequential recommendation algorithms.
Yonghong Yu, Aoran Zhang 0001, Li Zhang 0013, Rong Gao 0001, Hongzhi Yin
IEEE Trans. Comput. Soc. Syst.4
2024 Hybrid graph transformer networks for multivariate time series anomaly detection
Rong Gao 0001, Lingyu Yan, Donghua Liu, Yonghong Yu, Zhiwei Ye
J. Supercomput.1
2024 Neural Inference Search for Multiloss Segmentation Models
abstract
Semantic segmentation is vital for many emerging surveillance applications, but current models cannot be relied upon to meet the required tolerance, particularly in complex tasks that involve multiple classes and varied environments. To improve performance, we propose a novel algorithm, neural inference search (NIS), for hyperparameter optimization pertaining to established deep learning segmentation models in conjunction with a new multiloss function. It incorporates three novel search behaviors, i.e., Maximized Standard Deviation Velocity Prediction, Local Best Velocity Prediction, and n -dimensional Whirlpool Search. The first two behaviors are exploratory, leveraging long short-term memory (LSTM)-convolutional neural network (CNN)-based velocity predictions, while the third employs n -dimensional matrix rotation for local exploitation. A scheduling mechanism is also introduced in NIS to manage the contributions of these three novel search behaviors in stages. NIS optimizes learning and multiloss parameters simultaneously. Compared with state-of-the-art segmentation methods and those optimized with other well-known search algorithms, NIS-optimized models show significant improvements across multiple performance metrics on five segmentation datasets. NIS also reliably yields better solutions as compared with a variety of search methods for solving numerical benchmark functions.
Sam Slade, Li Zhang 0013, Haoqian Huang, Houshyar Asadi, Chee Peng Lim, Yonghong Yu, Dezong Zhao, Hanhe Lin, Rong Gao 0001
IEEE Trans. Neural Networks Learn. Syst.9
2023 Self-supervised Dual Hypergraph learning with Intent Disentanglement for session-based recommendation
abstract
Existing works on session-based recommendation have shown the advantage in enhancing the prediction ability of recommendation with various deep learning techniques. However, the following challenges need to be addressed: (1) the hierarchy of item transition patterns is overlooked; (2) existing works fail to distinguish various factors of item transition within a single session for disentangling user intents. To cope with the above challenges, we propose a novel session based recommendation model called S elf-supervised D ual H ypergraph learning with I ntent D isentanglement model ( SDHID ). Specifically, we first propose a disentangled capsule hypergraph convolutional channel for ne-grained intent learning to capture the intra-session pattern. Accordingly, we introduce the hypergraph and capsule networks in disentangling to learn the item embedding for different factors, and then the representation of the intra-session pattern is obtained by aggregating item embedding with attention weights. Moreover, we build a novel dual–primal hypergraph convolutional channel by mapping the hypergraph to a dual–primal graph for learning the item transition pattern of inter-session. In addition, the above two channels are combined into a self-supervised contrastive learning framework by maximizing mutual information between the learned session representations. We unify the recommendation and the self-supervised tasks under a primary and auxiliary learning framework. The combined optimization of two tasks leads to a hierarchical joint learning item transition for intra- and inter-session. Extensive experiments on real datasets show that the proposed model outperforms several state-of-the-art models.
Rong Gao 0001, Yuhe Tao, Yonghong Yu, Jia Wu 0001, Xiongkai Shao, Jing Li 0055, Zhiwei Ye
Knowl. Based Syst.1
2023 Lightweight object detection model fused with feature pyramid
Zaoning Wang, Rong Gao 0001, Lingyu Yan
Multim. Tools Appl.4
2022 Human Action Recognition Using Hybrid Deep Evolving Neural Networks
abstract
Human action recognition can be applied in a multitude of fully diversified domains such as active large-scale surveillance, threat detection, personal safety in hazardous environments, human assistance, health monitoring, and intelligent robotics. Owing to its high demands in real-world applications, it has drawn significant attention. In this research, we propose hybrid deep neural networks, i.e. Convolutional Long Short-Term Memory (ConvLSTM) Networks, Long-term Recurrent Convolutional Networks (LRCN), for tackling video action classification. In particular, for the LRCN model, different CNN encoder architectures such as VGG16, ResNet50, DenseNet121 and MobileNet, as well as several Long Short-Term Memory (LSTM) variant decoder architectures, such as LSTM, bidirectional LSTM (BiLSTM) and Gated Recurrent Unit (GRU), are used for spatial-temporal feature extraction to test model performance. We adopt diverse experimental settings including using different numbers of frames per video and learning configurations to optimize performance. The empirical results indicate the superiority of MobileNet in combination with a BiLSTM network over other hybrid network settings for the action classification using the UCF50 dataset. Owing to the lightweight MobileNet encoder, this LRCN model also achieves a better trade-off between performance and training and inference computational costs, while outperforming existing state-of-the-art methods.
Pavan Dasari, Li Zhang 0013, Yonghong Yu, Haoqian Huang, Rong Gao 0001
IJCNN5
2022 Gated Dual Hypergraph Convolutional Networks for Recommendation with Self-supervised Learning
abstract
Recommender systems have become a crucial intelligent tool, which provides users with personalized services. Graph learning-based recommendation methods treat user-item interactions and the item transitions as pairwise relations but ignore the complex, higher-order interaction information between nodes. Moreover, since most users often interact with few or even no items, graph learning-based recommendation methods suffer from the data sparsity problem, as well as the unbalanced distribution of edges and nodes. To tackle these issues, we propose a Dual Hypergraph-based Self-supervised Learning recommendation model, named DHSL-GM. Specifically, we derive two dual hypergraphs from the user-item bipartite graph, which models the complex high-order user-item interactions by using hypergraph convolution with spectral hypergraph convolution operator. Meanwhile, we design a gated network-based message passing mechanism to dynamically guide message propagation, addressing the problem of the unbalanced distribution of edges and nodes. In addition, to alleviate the data sparsity problem, we design another dual hypergraph convolutional network based on a node discard strategy, which innovatively integrates self-supervised learning into the training of the hypergraph convolutional network. Experimental results on several real datasets demonstrate the superiority and effectiveness of the proposed model.
Rong Gao 0001, Jiakang Liu, Yonghong Yu, Donghua Liu, Xiongkai Shao, Zhiwei Ye
IJCNN1
2022 Social-path embedding-based transformer for graduation development prediction
Guangze Yang, Yong Ouyang, Zhiwei Ye, Rong Gao 0001, Yawen Zeng
Appl. Intell.4
2022 Hybrid neural networks based facial expression recognition for smart city
Lingyu Yan, Menghan Sheng, Rong Gao 0001
Multim. Tools Appl.4
2021 A hybrid neural network approach to combine textual information and rating information for item recommendation
Donghua Liu, Jing Li 0055, Bo Du 0001, Rong Gao 0001, Yujia Wu
Knowl. Inf. Syst.5
2020 Graph Neural Networks Boosted Personalized Tag Recommendation Algorithm
abstract
Personalized tag recommender systems recommend a set of tags for items based on users' historical behaviors, and play an important role in the collaborative tagging systems. However, traditional personalized tag recommendation methods cannot guarantee that the collaborative signal hidden in the interactions among entities is effectively encoded in the process of learning the representations of entities, resulting in insufficient expressive capacity for characterizing the preferences or attributes of entities. In this paper, we proposed a graph neural networks boosted personalized tag recommendation model, which integrates the graph neural networks into the pairwise interaction tensor factorization model. Specifically, we consider two types of interaction graph (i.e. the user-tag interaction graph and the item-tag interaction graph) that is derived from the tag assignments. For each interaction graph, we exploit the graph neural networks to capture the collaborative signal that is encoded in the interaction graph and integrate the collaborative signal into the learning of representations of entities by transmitting and assembling the representations of entity neighbors along the interaction graphs. In this way, we explicitly capture the collaborative signal, resulting in rich and meaningful representations of entities. Experimental results on real world datasets show that our proposed graph neural networks boosted personalized tag recommendation model outperforms the traditional tag recommendation models.
Yonghong Yu, Fengyixin Jiang, Li Zhang 0013, Rong Gao 0001, Haiyan Gao
IJCNN5
2020 Elective future: The influence factor mining of students' graduation development based on hierarchical attention neural network model with graph
Yong Ouyang, Yawen Zeng, Rong Gao 0001, Yonghong Yu
Appl. Intell.3
2019 DRCGR: Deep Reinforcement Learning Framework Incorporating CNN and GAN-Based for Interactive Recommendation
abstract
Recently, the application of deep reinforcement learning into the field of session-based interactive recommendation has attracted great attention from researchers. However, despite that some interactive recommendation models based on deep reinforcement learning have been proposed, they still suffer to the following limitations: (1) these works ignore the skip behaviors of sequential patterns in users' clicking behavior; (2) these works fail to incorporate positive feedback and negative feedback into the proposed deep reinforcement recommender system when the positive feedback is sparse. Therefore, to solve the problems mentioned above, a novel Deep Q-Network based recommendation framework incorporating CNN and GAN-based models is proposed to acquire robust performance, named DRCGR. Specifically, in DRCGR, a CNN model is used to capture the sequential features for positive feedback. Then, an adversarial training is adopted to learn optimal negative feedback representations Then, positive/negative representations are fed into DQN simultaneously, which are conducive to generating better action-value function The experimental results based on real-world e-commerce data demonstrate our framework's superiority over some state-of-the-art recommendation models.
Rong Gao 0001, Haifeng Xia, Jing Li 0055, Donghua Liu, Gang Chun
ICDM1
2019 DAML: Dual Attention Mutual Learning between Ratings and Reviews for Item Recommendation
abstract
Despite the great success of many matrix factorization based collaborative filtering approaches, there is still much space for improvement in recommender system field. One main obstacle is the cold-start and data sparseness problem, requiring better solutions. Recent studies have attempted to integrate review information into rating prediction. However, there are two main problems: (1) most of existing works utilize a static and independent method to extract the latent feature representation of user and item reviews ignoring the correlation between the latent features, which may fail to capture the preference of users comprehensively. (2) there is no effective framework that unifies ratings and reviews. Therefore, we propose a novel d ual a ttention m utual l earning between ratings and reviews for item recommendation, named DAML. Specifically, we utilize local and mutual attention of the convolutional neural network to jointly learn the features of reviews to enhance the interpretability of the proposed DAML model. Then the rating features and review features are integrated into a unified neural network model, and the higher-order nonlinear interaction of features are realized by the neural factorization machines to complete the final rating prediction. Experiments on the five real-world datasets show that DAML achieves significantly better rating prediction accuracy compared to the state-of-the-art methods. Furthermore, the attention mechanism can highlight the relevant information in reviews to increase the interpretability of rating prediction.
Donghua Liu, Jing Li 0055, Bo Du 0001, Rong Gao 0001
KDD5
2018 Geographical Proximity Boosted Recommendation Algorithms for Real Estate
Yonghong Yu, Can Wang 0004, Li Zhang 0013, Rong Gao 0001, Hua Wang 0002
WISE (2)4
2018 STSCR: Exploring spatial-temporal sequential influence and social information for location recommendation
Rong Gao 0001, Jing Li 0055, Xuefei Li 0001, Chengfang Song, Donghua Liu
Neurocomputing1
2018 A personalized point-of-interest recommendation model via fusion of geo-social information
Rong Gao 0001, Jing Li 0055, Xuefei Li 0001, Chengfang Song
Neurocomputing1