Shengrong Gong

dblp:27/10764 · also Sheng-Rong Gong · DBLP profile ↗
← Back
32ranked-venue papers
1as first author
16since 2021 · last 2025
0000-0003-0266-2422ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Computer networks · 3Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Enhancing model learning in reinforcement learning through Q-function-guided trajectory alignment
Shengrong Gong, Yali Si
Appl. Intell.3
2025 Subgraph retrieval and link scoring model for multi-hop question answering in knowledge graphs
Changshun Zhou, Wenhao Ying, Shengrong Gong
Appl. Intell.4
2025 Coupled flows as guidance for model-based policy optimization
Shengrong Gong, Yuya Sun, Lifan Zhou
Eng. Appl. Artif. Intell.1
2025 Trajectory self-correction and uncertainty estimation for enhanced model-based policy optimization
Kaijian Xia, Yizhang Jiang, Yangtao Xue, Shengrong Gong
Expert Syst. Appl.6
2024 CDA-MBPO: Corrected Data Aggregation for Model-Based Policy Optimization
abstract
Model-based reinforcement learning has shown promise in sample efficiency but suffers from errors accumulated during multi-step model sampling. To tackle this issue, we propose corrected data aggregation for model-based policy optimization. This approach involves aligning simulated trajectories with their real counterparts from random starting states and with varying sampling lengths to create paired real-simulated samples. The R-Q discriminator is incorporated to assess the quality of the simulated samples by computing the R-Q difference, modeled as a Gaussian distribution within each paired sample. We update the Q network and the dynamics model using all real samples and the simulated samples whose R-Q difference fall below a predefined threshold. The experimental results demonstrate that our method outperforms state-of-the-art model-based methods in sample efficiency and asymptotic performance across challenging tasks. Our code is available at https://github.com/duxin0618/CDA-MBPO.
Wenhao Ying, Shengrong Gong
ICASSP5
2024 Undersampling method based on minority class density for imbalanced data
Zhongqiang Sun, Wenhao Ying, Shengrong Gong
Expert Syst. Appl.4
2023 DARN: Crowd Counting Network Guided by Double Attention Refinement
Shuhan Chang, Lifan Zhou, Xuanyu Zhou, Shengrong Gong
PRCV (10)5
2023 TEFNet: Target-Aware Enhanced Fusion Network for RGB-T Tracking
Panfeng Chen, Shengrong Gong, Wenhao Ying
PRCV (10)2
2022 Cross-Modality Visible-Infrared Person Re-Identification with Multi-scale Attention and Part Aggregation
Shengrong Gong
ICONIP (6)2
2022 Spatial and Temporal Guidance for Semi-supervised Video Object Segmentation
Shengrong Gong, Lifan Zhou
ICONIP (3)2
2022 LASDNet: A Lightweight Anchor-Free Ship Detection Network for SAR Images
abstract
Deep convolutional neural networks (DCNN)-based methods have been applied widely to ship detection in SAR images. However, most DCNN-based ship target detectors that focus on the detection performance ignore the computation complexity. We propose a lightweight anchor-free ship detection network (LASDNet) for SAR images to tackle this problem. First, a lightweight backbone utilizing a double fusion with squeeze-and-excitation-bottleneck block under the CSPNet design (CSP-DFSEB) and three pooling blocks (i.e., EVE, FCT, and ME blocks) are constructed, which achieves a balance between accuracy and efficiency. Second, a transformer-based aggregation layer conducts feature fusion. Finally, an improved one-stage anchor-free detector FCOS is presented. The analyses of the High-Resolution SAR Images Dataset for Ship Detection and Instance Segmentation (HRSID) dataset show that the proposed detector has the second least number of parameters (1.15 MB), the lowest computation complexity (1.01 GFLOPs), and the highest average precision (59.25) compared with other state-of-the-art methods.
Lifan Zhou, Hanwen Yu, Yong Wang 0011, Shaojie Xu, Shengrong Gong, Mengdao Xing
IGARSS5
2022 CANet: An Unsupervised Deep Convolutional Neural Network for Efficient Cluster-Analysis-Based Multibaseline InSAR Phase Unwrapping
abstract
Multibaseline (MB) phase unwrapping (PU) is a vital processing procedure for MB synthetic aperture radar interferometry (InSAR) signal processing and can improve the traditional InSAR by changing the ill-posed problem to the well-posed problem. The existing research has shown that the MB PU problem can be successfully converted into an unsupervised cluster analysis problem. Using the high feature descriptiveness of the deep learning technique, an unsupervised deep convolutional neural network, referred to as CANet, is proposed to cluster all the pixels into different groups according to the input’s recognizable pattern of the ambiguity number of the MB interferometric phase. Subsequently, we extend our previous two-stage programming-based MB processing approach (TSPA) to processing MB PU on a sparse irregular network, which is established from the clustering result of CANet. Both theoretical analysis and experimental results show that the proposed method is an effective MB PU method, and its execution time is drastically lower than those of many classical MB PU methods.
Lifan Zhou, Hanwen Yu, Shengrong Gong, Mengdao Xing
IEEE Trans. Geosci. Remote. Sens.4
2021 SAR Image Colorization Using Multidomain Cycle-Consistency Generative Adversarial Network
abstract
Synthetic aperture radar (SAR) images are widely used for aerial and spatial image applications. However, Most of SAR images are usually grayscale images with no color information. Hence, the study of SAR image colorization is meaningful. At present, deep learning has become the mainstream method of SAR coloring, and with the most advanced pix2pix method, it achieves satisfactory results. However, such an approach is limited to the corresponding paired data, which may be difficult to get. We then notice that the cycle-consistency loss can remove this constraint to some extent. In this letter, we present a novel method to colorize the SAR image using a multidomain cycle-consistency generative adversarial network (MC-GAN). The proposed method improves the performance of coloring SAR images from two aspects: first, we propose a mask vector for images of every particular terrain combined with cycle-consistency loss, which does not need the paired SAR-optical images to train the model. Second, we define the multidomain classification loss, which can together get the correct output image with the color we hope it to be. We examined the proposed method on the newly SEN1-2 data set compared with the pix2pix and CycleGAN methods, which demonstrates the effectiveness of our proposed method.
Guang Ji, Lifan Zhou, Shengrong Gong
IEEE Geosci. Remote. Sens. Lett.6
2021 Learning Transferable Driven and Drone Assisted Sustainable and Robust Regional Disease Surveillance for Smart Healthcare
abstract
Smart healthcare has been applied in many fields such as disease surveillance and telemedicine, etc. However, there are some challenges for device deployment, data collection and guarantee of stainability in regional disease surveillance. First, it is difficult to deploy sensors and adjust the sensor network in unknown region for dynamic disease surveillance. Second, the limited life-cycle of sensor network may cause the loss of surveillance data. Thus, it is important to provide a sustainable and robust regional disease surveillance system. Given a set of Disease surveillance Area (DsA)s and Point of disease Surveillance (PoS)s, some sensors are deployed to monitor these PoSs, and a drone collect data from the sensors as well as charge the sensors to extend their life-cycles. The drone replenish its energy by relying on the bus network. We first formulate the drone assisted regional disease surveillance problem under the constraints of life-cycle of sensors and energy of drone, and propose an approximation algorithm to find a feasible cycle of drone to minimize the traveling time cost of drone. To satisfy the diversity requirements and dynamic scalability of regional disease surveillance, we deploy one robot in each DsA instead of sensors. We further formulate the learning transferable driven regional disease surveillance problem, and propose a joint schedule algorithm of drone and robots. The results of both theoretical analysis and extensive simulations show that the proposed algorithms can reduce the total time cost by 39.71 and 48.74 percent, average waiting time by 42.00 and 50.14 percent, and increase the average accessing ratio of PoSs by 15.53 and 22.30 percent, through the assistance of bus network and learning transferable features.
Yong Jin 0003, Zhenjiang Qian, Shengrong Gong, Weiyong Yang
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 Person Reidentification Based on Pose-Invariant Feature and B-KNN Reranking
abstract
Person reidentification, as one of the most important areas in the video surveillance field, is a crucial task in computer vision. It has attracted more and more attention in both academic research and industry due to its extremely high application value. However, the recognition accuracy of person reidentification is subjected to factors, such as illumination, pose, occlusion, and viewpoint. To alleviate the effect of such factors, the multiscale Retinex with color restoration (MSRSC) algorithm is adopted to preprocess the original images so that the color information can be restored and the illumination condition can be improved. To obtain pose-invariant features (PIFs), the convolution pose machine that can generate the body joint points of pedestrians is applied to divide the body into seven parts, the pose transform network is then used to align the body parts, and finally, the PIFs can be obtained from a designed pseudo-Siamese network by using the original and aligned images as the training samples. To further improve recognition accuracy, a reranking method based on bidirectional k-nearest neighbors (KNN) is presented to optimize the ranking list. Experimentally, the proposed method is implemented on three data sets: viewpoint invariant pedestrian recognition (VIPeR), CUHK03, and Market1501. The results demonstrate our method outperforms the other methods, as a result of both the representation of a more discriminative feature descriptor and the introduction of a reranking method.
Zongming Bao, Shengrong Gong, Kaijian Xia
IEEE Trans. Comput. Soc. Syst.3
2021 Behavior Prediction for Unmanned Driving Based on Dual Fusions of Feature and Decision
abstract
Behavioral decision systems may suffer from poor performance due to the failure in capturing the vibrations of environmental information. To better capture such vibrations and then make more accurate predictions, a parallel deep neural network based on dual fusions including feature and decision is proposed, called DFFD-Net. DFFD-NET is composed of two parts, the feature fusion network and the driving data network. The feature fusion model adopts two different operations, deconvolution and linear weighting, to fuse local features and global features, respectively. Deconvolution is applied between the convolutional layers, while linear weighting is operated among the outputs of SPP and LSTM. To further improve the accuracy of the prediction, the decisions generated from both networks are further weighed to get the final decision. Experimentally, DFFD-NET is implemented in the benchmarks BDDV and TORCS, and the results show that the final performance is benefited from both feature fusion and decision fusion. From the comparison, DFFD-NET can get state-of-the-art results on both perplexity and precision by only using the images captured from the front-facing camera as well as a few sensing data.
Shengrong Gong, Kaijian Xia, Yuchen Fu, Qiming Fu 0001, Hongsheng Yin 0001
IEEE Trans. Intell. Transp. Syst.3
2020 ACCLVOS: Atrous Convolution with Spatial-Temporal ConvLSTM for Video Object Segmentation
abstract
Semi-supervised video object segmentation aims at segmenting the target of interest throughout a video sequence when only the annotated mask of the first frame is given. A feasible method for segmentation is to capture the spatial-temporal coherence between frames. However, it may suffer from mask drift when the spatial-temporal coherence is unreliable. To relieve this problem, we propose an encoder-decoder-recurrent model for semi-supervised video object segmentation. The model adopts a U-shape architecture that combines atrous convolution and ConvLSTM to establish the coherence in both the spatial and temporal domains. Furthermore, the weight ratio for each block is also reconstructed to make the model more suitable for the VOS task. We evaluate our method on two benchmarks, DAVIS-2017 and Youtube-VOS, where state-of-the-art segmentation accuracy with a real-time inference speed of 21.3 frames per second on a Tesla P100 is obtained.
Muzhou Xu, Chunping Liu, Shengrong Gong
ICPR4
2020 Modeling-Learning-Based Actor-Critic Algorithm with Gaussian Process Approximator
Jack Tan, Husheng Dong, Xuemei Chen 0002, Shengrong Gong, Zhenjiang Qian
J. Grid Comput.5
2020 Spatio-Temporal Deep Residual Network with Hierarchical Attentions for Video Event Recognition
abstract
Event recognition in surveillance video has gained extensive attention from the computer vision community. This process still faces enormous challenges due to the tiny inter-class variations that are caused by various facets, such as severe occlusion, cluttered backgrounds, and so forth. To address these issues, we propose a spatio-temporal deep residual network with hierarchical attentions (STDRN-HA) for video event recognition. In the first attention layer, the ResNet fully connected feature guides the Faster R-CNN feature to generate object-based attention (O-attention) for target objects. In the second attention layer, the O-attention further guides the ResNet convolutional feature to yield the holistic attention (H-attention) in order to perceive more details of the occluded objects and the global background. In the third attention layer, the attention maps use the deep features to obtain the attention-enhanced features. Then, the attention-enhanced features are input into a deep residual recurrent network, which is used to mine more event clues from videos. Furthermore, an optimized loss function named softmax-RC is designed, which embeds the residual block regularization and center loss to solve the vanishing gradient in a deep network and enlarge the distance between inter-classes. We also build a temporal branch to exploit the long- and short-term motion information. The final results are obtained by fusing the outputs of the spatial and temporal streams. Experiments on the four realistic video datasets, CCV, VIRAT 1.0, VIRAT 2.0, and HMDB51, demonstrate that the proposed method has good performance and achieves state-of-the-art results.
Chunping Liu, Yi Ji 0001, Shengrong Gong, Haibao Xu
ACM Trans. Multim. Comput. Commun. Appl.4
2019 DXNet: An Encoder-Decoder Architecture with XSPP for Semantic Image Segmentation in Street Scenes
Yexin Shang, Shengrong Gong, Lifan Zhou, Wenhao Ying
ICONIP (5)3
2019 Region Selection Model with Saliency Constraint for Fine-Grained Recognition
Shaoxiong Zhou, Shengrong Gong, Wenhao Ying
ICONIP (1)2
2019 Trajectory-Pooled Spatial-Temporal Architecture of Deep Convolutional Neural Networks for Video Event Detection
abstract
Nowadays content-based video event detection faces great challenges due to complex scenes and blurred actions in surveillance videos. To alleviate these challenges, we propose a novel spatial-temporal architecture of deep convolutional neural networks for this task. By taking advantage of spatial-temporal information, we fine-tune two-stream networks, and then, fuse spatial and temporal features at convolution layers using a 2D pooling fusion method to enforce the consistence of spatial-temporal information. Based on the two-stream networks and spatial-temporal layer, a triple-channel model is obtained. Furthermore, we implement trajectory-constrained pooling to deep features and hand-crafted features to combine their merits. A fusion method on triple-channel yields the final detection result. The experiments on two benchmark surveillance video data sets including VIRAT 1.0 and VIRAT 2.0, which involve a suit of challenging events, such as person loading an object to a vehicle or person opening a vehicle trunk, manifest that the proposed method can achieve superior performance compared with the state-of-the-art methods on these event benchmarks.
Rui Ge 0005, Yi Ji 0001, Shengrong Gong, Chunping Liu
IEEE Trans. Circuits Syst. Video Technol.4
2018 Person re-identification by enhanced local maximal occurrence representation and generalized similarity metric learning
Husheng Dong, Chunping Liu, Yi Ji 0001, Shengrong Gong
Neurocomputing6
2018 Person re-identification by kernel null space marginal Fisher analysis
Husheng Dong, Chunping Liu, Yi Ji 0001, Shengrong Gong
Pattern Recognit. Lett.6
2018 Learning Multiple Kernel Metrics for Iterative Person Re-Identification
abstract
In person re-identification most metric learning methods learn from training data only once, and then they are deployed for testing. Although impressive performance has been achieved, the discriminative information from successfully identified test samples are ignored. In this work, we present a novel re-identification framework termed Iterative Multiple Kernel Metric Learning (IMKML). Specifically, there are two main modules in IMKML. In the first module, multiple metrics are learned via a new derived Kernel Marginal Nullspace Learning (KMNL) algorithm. Taking advantage of learning a discriminative nullspace from neighborhood manifold, KMNL can well tackle the Small Sample Size (SSS) problem in re-identification distance metric learning. The second module is to construct a pseudo training set by performing re-identification on the testing set. The pseudo training set, which consists of the test image pairs that are highly probable correct matches, is then inserted into the labeled training set to retrain the metrics. By iteratively alternating between the two modules, many more samples will be involved for training and significant performance gains can be achieved. Experiments on four challenging datasets, including VIPeR, PRID450S, CUHK01, and Market-1501, show that the proposed method performs favorably against the state-of-the-art approaches, especially on the lower ranks.
Husheng Dong, Chunping Liu, Yi Ji 0001, Shengrong Gong
ACM Trans. Multim. Comput. Commun. Appl.5
2017 Large margin relative distance learning for person re-identification
abstract
Distance metric learning has achieved great success in person re‐identification. Most existing methods that learn metrics from pairwise constraints suffer the problem of imbalanced data. In this study, the authors present a large margin relative distance learning (LMRDL) method which learns the metric from triplet constraints, so that the problem of imbalanced sample pairs can be bypassed. Different from existing triplet‐based methods, LMRDL employs an improved triplet loss that enforces penalisation on the triplets with minimal inter‐class distance, and this leads to a more stringent constraint to guide the learning. To suppress the large variations of pedestrian's appearance in different camera views, the authors propose to learn the metric over the intra‐class subspace. The proposed method is formulated as a logistic metric learning problem with positive semi‐definite constraint, and the authors derive an efficient optimisation scheme to solve it based on the accelerated proximal gradient approach. Experimental results show that the proposed method achieves state‐of‐the‐art performance on three challenging datasets (VIPeR, PRID450S, and GRID).
Husheng Dong, Shengrong Gong, Chunping Liu, Yi Ji 0001
IET Comput. Vis.2
2016 ESRS: An Efficient and Secure Relay Selection Algorithm for Mobile Social Networks
Xiaoshuang Xing, Xiuzhen Cheng, Shengrong Gong, Feng Zhao 0002, Hongbin Qiu
WASA4
2015 Fusion of spatially constrained attributes with kernelized ranking for person re-identification
abstract
The task of matching persons across non-overlapping camera views, known as person re-identification, is rather challenging due to strong visual similarity and large appearance changes caused by illumination, pose and occlusion. Most approaches rely on low-level features that are both discriminative and invariant. In this work, we propose a novel method to address this problem by fusing mid-level semantic attributes with kernelized ranking. First, a kernelized ranking model is learned, and it gives the initial ranking scores. Next, an adaptive similarity model based on spatially constrained attributes is used to refine the ranking list. Fusion of the two models leads to much better performance than each individual alone. Experiments demonstrate complements of the two models and the results achieve new state-of-the-art performance on two benchmark datasets.
Husheng Dong, Chunping Liu, Yi Ji 0001, Shengrong Gong
AVSS5
2015 Robust Dynamic Background Model with Adaptive Region Based on T2FS and GMM
abstract
For many tracking and surveillance applications, Gaussian mixture model (GMM) provides an effective mean to segment the foreground from background. Though, because of insufficient and noisy data in complex dynamic scenes, the estimated parameters of the GMM, which are based on the assumption that the pixel process meets multi-modal Gaussian distribution, may not accurately reflect the underlying distribution of the observations. And the existing block-based GMM (BGMM) method may be able to segment only rough foreground objects with time-consuming calculations. To solve these difficulties, this paper proposes to use type-2 fuzzy sets (T2FSs) to handle GMM’s uncertain parameters (T2GMM). Furthermore, this paper also introduces a novel representation of contextual spatial information including the color, edge and texture features for each block which is faster and almost lossless (T2BGMM). Experimental results demonstrate the efficiency of the proposed methods.
Yun Guo, Yi Ji 0001, Jutao Zhang, Shengrong Gong, Chunping Liu
KSEM4
2015 Learning topic of dynamic scene using belief propagation and weighted visual words approach
Chunping Liu, Shengrong Gong, Yi Ji 0001, Quan Liu 0004
Soft Comput.3
2013 Image categorization using a semantic hierarchy model with sparse set of salient regions
Chunping Liu, Shengrong Gong
Frontiers Comput. Sci.3
2013 Online belief propagation algorithm for probabilistic latent semantic analysis
Shengrong Gong, Chunping Liu
Frontiers Comput. Sci.2