Shengqin Jiang

dblp:190/4806 · DBLP profile ↗
← Back
24ranked-venue papers
17as first author
17since 2021 · last 2026
0000-0003-0026-8210ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Teacher Agent: A Knowledge Distillation-Free Framework for Rehearsal-Based Video Incremental Learning
Shengqin Jiang, Yaoyu Fang, Haokui Zhang, Qingshan Liu 0001, Yuankai Qi, Yang Yang 0002, Peng Wang 0023
Int. J. Comput. Vis.1
2025 Distillation-boosted heterogeneous architecture search for aphid counting
Shengqin Jiang, Qian Jie, Fengna Cheng, Kelu Yao
Expert Syst. Appl.1
2025 DuPt: Rehearsal-based continual learning with dual prompts
Shengqin Jiang, Daolong Zhang, Fengna Cheng, Xiaobo Lu, Qingshan Liu 0001
Neural Networks1
2025 Remote Sensing Object Counting With Online Knowledge Learning
abstract
Efficient models for remote sensing object counting are urgently required for applications in scenarios with limited computing resources, such as drones or embedded systems. A straightforward yet powerful technique to achieve this is knowledge distillation (KD), which steers the learning of student networks by leveraging the experience of already-trained teacher networks. However, it faces a pair of challenges. First, due to its two-stage training nature, a longer training period is essential, especially as the training samples increase. Second, despite the proficiency of teacher networks in transmitting assimilated knowledge, they tend to overlook the latent insights gained during their learning process. To address these challenges, we introduce an online distillation learning method for remote sensing object counting. It builds an end-to-end training framework that seamlessly integrates two distinct networks into a unified one. It comprises a shared shallow module, a teacher branch, and a student branch. The shared module serving as the foundation for both branches is dedicated to learning some primitive information. The teacher branch utilizes prior knowledge to reduce the difficulty of learning and guides the student branch in online learning. In parallel, the student branch achieves parameter reduction and rapid inference capabilities by means of channel reduction. This design empowers the student branch not only to receive privileged insights from the teacher branch but also to tap into the latent reservoir of knowledge held by the teacher branch during the learning process. Moreover, we propose a relation-in-relation distillation (RiRD) method that allows the student branch to effectively comprehend the evolution of the relationship of intralayer teacher features among different interlayer features. Extensive experiments on two challenging datasets demonstrate the effectiveness of our method, which achieves comparable performance to state-of-the-art (SOTA) methods despite using far fewer parameters.
Shengqin Jiang, Yuan Gao 0053, Fengna Cheng, Renlong Hang, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Cell-Wise Self-Optimization: Making Pretrained Model Better in Remote Sensing Counting
abstract
Recently, remote sensing counting has drawn a lot of attention due to its wide application requirements. However, most existing approaches tend to focus on optimizing the network backend or designing new loss functions, and overlook a foundational component, i.e., the feature extractor, which is usually based on a pre-trained model. This oversight limits the potential for further performance improvement. In this paper, we propose a cell-wise self-optimization method to enhance the feature extractor. By leveraging the powerful representation capabilities of a pre-trained model, our method further refines them for counting tasks, notably improving network performance on limited remote sensing data. Specifically, we design a lightweight cell-wise architecture optimization based on a network architecture search algorithm. It sequentially builds lightweight cells in parallel with the blocks in a pre-trained model, leveraging a newly proposed random path selection strategy for training the optimization framework. Additionally, we propose extracting high-frequency information from the blocks in a pre-trained model as self-guidance optimization to facilitate the learning of the searched cells. Experimental results on several representative datasets demonstrate that our proposed method significantly improves counting performance while substantially reducing parameters. For instance, compared with our baseline, it improves performance by 22.2% in MAE and 17.0% in MSE on the Building dataset, while reducing parameters by 61.2%.
Shengqin Jiang, Qian Jie, Fengna Cheng, Haokui Zhang, Yu Liu 0029, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Input-Regulated Remote Sensing Counting With Region Understanding
abstract
Remote sensing counting aims to automatically estimate the number of objects of interest from high-resolution aerial or satellite imagery, providing critical decision-making support in areas such as urban planning, traffic monitoring, and disaster response. While most existing methods leverage pre-trained models to enhance feature generalization, their performance is often hindered by the severe scarcity of annotated remote sensing data. This limits their generalizability in complex scenarios. To address these challenges, we propose a novel remote sensing counting network that effectively captures informative signals from relatively limited annotated data. Specifically, we first introduce a graph-driven input regulator that constructs a graph structure by modeling relationships among input features, effectively capturing intrinsic contextual dependencies. This structure allows the regulator to assign adaptive pixel-level weights to network inputs, prioritizing relevant signals while mitigating the risk of overfitting to a fixed data distribution. Second, we design a dynamic region-aware module that leverages fuzzy logic to adaptively identify and enhance highly discriminative local regions. In this way, it improves the robustness of the feature representations. Extensive experiments demonstrate the effectiveness of the proposed method compared with several state-of-the-art methods.
Shengqin Jiang, Haojian Long, Fengna Cheng, Yuankai Qi, Xiaobo Lu, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Spatial-Temporal Interleaved Network for Efficient Action Recognition
abstract
The decomposition of 3D convolution will considerably reduce the computing complexity of 3D convolutional neural networks, yet simple stacking restricts the performance of neural networks. To this end, we propose a spatial-temporal interleaved network for efficient action recognition. By deeply analyzing this task, it revisits the structure of 3D neural networks in action recognition from the following perspectives. To enhance the learning of robust spatial-temporal features, we initially propose an interleaved feature interaction module to comprehensively explore cross-layer features and capture the most discriminative information among them. With regards to being lightweight, a boosted parallel pseudo-3D module is introduced with the goal of circumventing a substantial number of computations from the lower to middle levels while enhancing temporal and spatial features in parallel at high levels. Furthermore, we exploit a spatial-temporal differential attention mechanism to suppress redundant features in different dimensions while reaping the benefits of nearly negligible parameters. Lastly, extensive experiments on four action recognition benchmarks are given to show the advantages and efficiency of our proposed method. Specifically, our method attains a 15.2% improvement in Top-1 accuracy compared to our baseline, a stack of full 3D convolutional layers, on the Something-Something V1 dataset while utilizing only 18.2% of the parameters.
Shengqin Jiang, Haokui Zhang, Yuankai Qi, Qingshan Liu 0001
IEEE Trans. Ind. Informatics1
2025 UST-SU: a U-shaped video prediction network based on partial autoregression
Zhaojun Cui, Shengqin Jiang
J. Supercomput.5
2025 Self-Reflection Neural Network for Class-Incremental Object Counting
abstract
In crowded scenarios, achieving the counting task of dynamically evolving categories is extremely challenging. In addition to grappling with challenges such as scale variations, severe occlusion and complex backgrounds, it is imperative to mitigate the issue of catastrophic forgetting. Previous approaches have heavily relied on leveraging historical data for knowledge distillation to tackle these difficulties. However, this strategy encounters two prominent obstacles: 1) Employing the teacher network from the previous stage for distillation incurs additional computational overhead during the training stage. 2) Although knowledge distillation can facilitate effective knowledge transfer, some inaccurate predictions from the teacher network may affect the knowledge acquisition in the current stage. To overcome these issues, we introduce a novel solution: a self-reflection neural network for class-incremental object counting. First, we construct a global-aware incremental regression branch that uses stacked transformer layers as backends to capture global information, while the final regression layers dynamically expand as categories increase. Furthermore, we introduce an uncertain estimation branch that selectively isolates certain feature maps to avoid some neurons updated with excessive gradient information, thereby enhancing the network plasticity while preserving stability. The output of this branch functions as a regularization signal, steering the learning process of the incremental regression branch. To foster a more robust retention of past knowledge, we propose a self-reflection loss. It employs the rectified outputs of global-aware incremental regression branch to encourage the network to reflect upon and refine its grasp of historical knowledge, effectively averting the pitfalls of inaccurate information. Our extensive experiments validate the effectiveness of our proposed method, achieving state-of-the-art results.
Shengqin Jiang, Linfei Li, Fengna Cheng, Yuankai Qi, Qingshan Liu 0001
IEEE Trans. Multim.1
2024 Tripartite-structure transformer for hyperspectral image classification
abstract
Abstract Hyperspectral images contain rich spatial and spectral information, which provides a strong basis for distinguishing different land‐cover objects. Therefore, hyperspectral image (HSI) classification has been a hot research topic. With the advent of deep learning, convolutional neural networks (CNNs) have become a popular method for hyperspectral image classification. However, convolutional neural network (CNN) has strong local feature extraction ability but cannot deal with long‐distance dependence well. Vision Transformer (ViT) is a recent development that can address this limitation, but it is not effective in extracting local features and has low computational efficiency. To overcome these drawbacks, we propose a hybrid classification network that combines the strengths of both CNN and ViT, names Spatial‐Spectral Former(SSF). The shallow layer employs 3D convolution to extract local features and reduce data dimensions. The deep layer employs a spectral‐spatial transformer module for global feature extraction and information enhancement in spectral and spatial dimensions. Our proposed model achieves promising results on widely used public HSI datasets compared to other deep learning methods, including CNN, ViT, and hybrid models.
Liuwei Wan, Meili Zhou, Shengqin Jiang, Zongwen Bai, Haokui Zhang
Comput. Intell.3
2024 RadarNet: A parallel spatiotemporal encoder network for radar extrapolation
Wei Tian 0002, Lei Yi, Xianghua Niu, Rong Fang, Shengqin Jiang
Neurocomputing8
2024 Exposure difference network for low-light image enhancement
Shengqin Jiang, Yongyue Mei
Pattern Recognit.1
2024 A Unified Object Counting Network With Object Occupation Prior
abstract
The counting task, which plays a fundamental role in numerous applications (e.g., crowd counting, traffic statistics), aims to predict the number of objects with various densities. Existing object counting tasks are designed for a single object class. However, it is inevitable to encounter newly coming data with new classes in our real world. We name this scenario as evolving object counting. In this paper, we build the first evolving object counting dataset and propose a unified object counting network as the first attempt to address this task. The proposed network consists of two key components: a class-agnostic mask module and a class-incremental module. The class-agnostic mask module learns generic object occupation prior by predicting a class-agnostic binary mask (e.g., 1 denotes there exists an object at the considering position in an image and 0 otherwise). The class-incremental module is used to handle new classes and provides discriminative class guidance for density map prediction. The combined outputs of the class-agnostic mask module and image feature extractor are used to predict the final density map. When new classes arrive, we first add new neural nodes to the last regression and classification layers of the class-incremental module. Then, instead of retraining the model from scratch, we utilize knowledge distillation to help the model retain and consolidate what it has previously learned. We also employ a support sample bank to store a small number of typical training samples for each class, which are used to prevent the model from forgetting key information from old data. With this design, our model can efficiently and effectively adapt to new classes while maintaining good performance on already-seen data without large-scale retraining. Extensive experiments on the collected dataset demonstrate favorable performance. The dataset and code will be available at:https://github.com/Tanyjiang/EOCO.
Shengqin Jiang, Fengna Cheng, Yuankai Qi, Qingshan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Compare and Focus: Multi-Scale View Aggregation for Crowd Counting
abstract
Recently, some state-of-the-art (SOTA) methods have designed dedicated context extractors to capture the global information that serves as a key clue for describing crowd density. A promising alternative is the transformer-based model which inherently captures long-range context dependencies. Recent related studies have made impressive progress, yet the following issues remain: (1) The size of the heads in the image is large near and small far away. The existing models fail to cope well with these variations. (2) There is an imbalance in the distribution of samples across different densities in the dataset, which leads to poor network performance on density distributions with a small number of samples. To address these issues, we propose to aggregate multi-scale views through Compare and Focus strategies. In terms of the first strategy, we mine differential hints from multi-scale view features to capture heads of varying sizes. This can effectively reduce the influence of redundant information while perceiving the subtleties of various view inputs, making it simpler to establish discriminative representations. As for the second strategy, we introduce a new activation function to formulate the Region of Interest (ROI) extraction module that enables the network to focus on relevant regions effectively. It can alleviate the extreme distribution imbalance of samples with different densities. Finally, several experiments show that our method achieves SOTA performance on four challenging datasets.
Shengqin Jiang, Jialu Cai, Haokui Zhang, Yu Liu 0029, Qingshan Liu 0001
IEEE Trans. Intell. Transp. Syst.1
2023 Uncertainty meets fixed-time control in neural networks
Shengqin Jiang, Yu Liu 0029, Shuiming Cai, Xiaobo Lu
Neurocomputing2
2021 Light fixed-time control for cluster synchronization of complex networks
Shengqin Jiang, Yuankai Qi, Shuiming Cai, Xiaobo Lu
Neurocomputing1
2021 D3D: Dual 3-D Convolutional Network for Real-Time Action Recognition
abstract
Three-dimensional convolutional neural networks (3D CNNs) have been explored to learn spatio-temporal information for video-based human action recognition. Expensive computational cost and memory demand resulted from standard 3D CNNs, however, hinder their application in practical scenarios. In this article, we address the aforementioned limitations by proposing a novel dual 3-D convolutional network (D3DNet) with two complementary lightweight branches. A coarse branch maintains large temporal receptive field by a fast temporal downsampling strategy and simulates the expensive 3-D convolutions using a combination of more efficient spatial convolutions and temporal convolutions. Meanwhile, a fine branch progressively downsamples the video in the temporal domain and adopts 3-D convolutional units with reduced channel capacities to capture multiresolution spatio-temporal information. Instead of learning these two branches independently, a shallow spatiotemporal downsampling module is shared for these two branches for efficient low-level feature learning. Besides, lateral connections are learned to effectively fuse the information from the two branches at multiple stages. The proposed network makes good balance between inference speed and action recognition performance. Based on RGB information only, it achieves competing performance on five popular video-based action recognition datasets, with inference speed of 3200 FPS on a single NVIDIA GTX 2080Ti card.
Shengqin Jiang, Yuankai Qi, Haokui Zhang, Zongwen Bai, Xiaobo Lu, Peng Wang 0023
IEEE Trans. Ind. Informatics1
2020 SCRM: self-correlated representation model for visual tracking
Shengqin Jiang, Xiaobo Lu, Fengna Cheng
Soft Comput.1
2020 Mask-Aware Networks for Crowd Counting
abstract
Crowd counting problem aims to count the number of objects within an image or a frame in the videos and is usually solved by estimating the density map generated from the object location annotations. The values in the density map, by nature, take two possible states: zero indicating no object around, a non-zero value indicating the existence of objects and the value denoting the local object density. In contrast to traditional methods which do not differentiate the density prediction of these two states, we propose to use a dedicated network branch to predict the object/non-object mask and then combine its prediction with the input image to produce the density map. Our rationale is that the mask prediction could be better modeled as a binary segmentation problem and the difficulty of estimating the density could be reduced if the mask is known. A key to the proposed scheme is the strategy of incorporating the mask prediction into the density map estimator. To this end, we study five possible solutions, and via analysis and experimental validation we identify the most effective one. Through extensive experiments on three public datasets, we demonstrate the superior performance of the proposed approach over the baselines and show that our network could achieve the state-of-the-art performance.
Shengqin Jiang, Xiaobo Lu, Yinjie Lei, Lingqiao Liu
IEEE Trans. Circuits Syst. Video Technol.1
2018 Bidirectionally aligned sparse representation for single image super-resolution
Shengqin Jiang, Xiaobo Lu
Multim. Tools Appl.3
2018 WeSamBE: A Weight-Sample-Based Method for Background Subtraction
abstract
Background subtraction techniques are often treated as fundamental and significant ways to analyze and understand video content. In this paper, we propose a weight-sample-based method for foreground detection. This method allows us to use a few samples with variable weights to achieve effective change detection. To rapidly adapt to changing scenarios, a minimum-weight update policy is first proposed to replace the most inefficient sample instead of the oldest sample or a random sample. In addition, a reward-and-penalty weighting strategy is put forward to reinforce active samples and punish others. In this way, the weights of relatively effective samples are increased and the false updating of effective samples with smaller weights is reduced. Moreover, some other strategies, such as spatial-diffusion policy and random time subsampling, are also incorporated to ensure the flexibility of the proposed method. Finally, in our experiments, an adaptive feedback technique is incorporated into our algorithm to adapt to more challenging videos, and the final results indicate that our method is superior to the state-of-the-art approaches on the challenging CDnet data set.
Shengqin Jiang, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.1
2017 Adaptive finite-time control for overlapping cluster synchronization in coupled complex networks
Shengqin Jiang, Xiaobo Lu, Shuiming Cai
Neurocomputing1
2017 Adaptive outer synchronization between two complex delayed dynamical networks via aperiodically intermittent pinning control
Xuqiang Lei, Shuiming Cai, Shengqin Jiang, Zengrong Liu
Neurocomputing3
2017 Multiscale self-similarity and sparse representation based single image super-resolution
Shengqin Jiang, Xiaobo Lu
Neurocomputing3