VLDB 2026 Research / reviewers in the wild / expert
Sheng Zhong 0006
dblp:53/4506-6
· DBLP profile ↗
14ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0001-8466-3730ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task LearningabstractGenerative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often undermine their effectiveness in downstream visual tasks. This paper introduces the Iterative Self-Training with Class-Aware Text-to-Image Synthesis (IST-CATS) framework, which addresses these challenges by integrating a class-aware text-to-image synthesis (CATS) component with an iterative self-training (IST) strategy. CATS innovatively introduces a class-aware chain approach to generate detailed descriptions. These descriptions act as prompts for a diffusion model, enabling the creation of a diverse of images accompanied by distinguishable objects against the background. The generated images can be easily pseudo-labeled by an unsupervised instance segmentation method, and then noisy pseudo labels can be effectively purified by a novel feature similarity-based filtering mechanism. The generated images underpin our IST, which progressively enhances vision models and refines pseudo labels through self-training and our proposed label filtering strategy (LabFilt). LabFilt meticulously improves the quality of pseudo labels by employing class-adaptive techniques at both the pixel and object levels, ensuring refined pseudo-label accuracy. IST-CATS demonstrates superior performance in object detection and semantic segmentation compared to traditional synthetic and semi/weakly-supervised methods, effectively addressing data collection and annotation challenges. Xiang Zhang 0018, Wanqing Zhao, Pengyang Li, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
AAAI | 6 |
| 2025 | Semantic image segmentation via dynamic curriculum learning
Xiang Zhang 0018, Wanqing Zhao, Chenji Wang, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
Appl. Intell. | 5 |
| 2024 | Enabling Near-Zero Cost Object Detection in Remote Sensing Imagery via Progressive Self-TrainingabstractDeep learning-based object detection models rely heavily on large-scale and precise annotations for training. However, manually annotating bounding-box annotations for such data is both time-consuming and costly, especially when dealing with high-resolution satellite imagery containing densely packed small-sized objects. To alleviate the burden of manual annotation, we propose a simple yet effective approach, called progressive self-training object detection (PSTDet), to enable accurate object detection in remote sensing imagery without relying on manual annotations. Our PSTDet framework consists of two main components: initial pseudo label generation (IPLG) and progressive self-training with relabeling (PST-R). In IPLG, we leverage unsupervised image clustering, unsupervised instance detection, and geometric constraints to automatically generate high-quality bounding-box annotations for the initial training dataset. This innovative approach significantly reduces the time and expense associated with data annotation, laying a solid foundation for the subsequent progressive self-training stage. The annotations produced by IPLG serve as the training data for PST-R, which enhances the detector and pseudo labels through progressive self-training and our proposed noisy pseudo label filtering strategy (NPLFilter). Our NPLFilter purifies the quality of pseudo labels by integrating geometric constraints, prior knowledge, and category-adaptive thresholds. Experimental results demonstrate that our method achieves significant performance improvement on challenging NWPU VHR-10.v2 and DIOR datasets. Notably, our method far outperforms state-of-the-art weakly supervised methods and compares favorably with fully supervised methods. Xiang Zhang 0018, Xiangteng Jiang, Qiyao Hu, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Hardware-Based Satellite Network Broadcast Storm Suppression Method
Keran Zhang, Hangzai Luo, Sheng Zhong 0006 |
Mob. Networks Appl. | 4 |
| 2023 | Learning to recover lost details from the dark
Maomei Liu, Sheng Zhong 0006, Hangzai Luo, Jinye Peng 0001 |
Pattern Recognit. Lett. | 3 |
| 2023 | Toward Blind-Adaptive Remote Sensing Image RestorationabstractWhile deep convolutional neural networks (CNNs) have substantially boosted the performance of low-level vision tasks, they remain largely under-explored in CNN-based remote sensing image restoration. This paper studies the JPEG-LS compressed remote sensing image restoration that faces the following problems. It requires a trade-off in preserving local context information and expanding spatial receptive fields. It needs blind restoration while achieving flexible performance. To this end, we propose a blind-adaptive restoration network, called TBANet, that integrates three modules into an end-to-end network to remedy these problems separately. Specifically, we build a scale-invariant wise-skip ResNet as the baseline to extract more context information. We present a receptive field expansion module by using scale-wise convolution for removing banding artifacts. We design a blind-adaptive controller to provide a deterministic result meanwhile meeting the needs of the user’s preference. In experiments, we compare the restoration accuracy among our model and many different variants of restoration methods on our collected remote sensing image dataset. The proposed network achieves superior performance against state-of-the-art methods in terms of both quantitative metrics and visual quality. Code and models are available at: https://github.com/lmmhh/TBANet. Maomei Liu, Lijia Fan, Sheng Zhong 0006, Hangzai Luo, Jinye Peng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Movable Object Detection in Remote Sensing Images via Dynamic Automatic LearningabstractThe performance of deep networks for object detection in remote sensing images (RSIs) largely depends on the availability of large-scale training images whose labels are given at the bounding-box level through a labor-intensive manual labeling process. To alleviate the huge burden of providing bounding-box annotations manually for movable objects, we propose a new approach, called dynamic automatic learning (DAL), to progressively learn object detectors. Specifically, a novel initial annotation generation (IAG) strategy is first designed to produce bounding-box annotations for movable objects in multi-temporal remote sensing images. During this process, image-level labels need to be manually labeled for the generated candidates. Next, a detection network learns the detection knowledge from multi-temporal remote sensing images with bounding-box annotations and then transfers the knowledge to generate pseudo boxes for the unlabeled data. Finally, with these pseudo boxes, the object detector can be optimized for generating accurate pseudo boxes iteratively. Furthermore, we introduce a pseudo box filtering (PBF) strategy to purify the quality of pseudo boxes to obtain accurate supervision. Our experiments on the challenging NWPU VHR-10.v2 and DIOR datasets have demonstrated that our DAL approach can achieve competitive results compared to state-of-the-art methods. Xiang Zhang 0018, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Boundary-Aware Bilateral Fusion Network for Cloud DetectionabstractCloud detection is one of the key technologies in the field of remote sensing. Although extensive deep learning-based cloud detection methods achieve good performance, their detection results in confusing areas such as cloud boundaries and thin clouds are often not satisfactory due to the potential inter-class similarity and intra-class inconsistency of objects. To this end, we propose a Boundary-Aware Bilateral Fusion network (BABFNet), which effectively enhances cloud detection in confusing areas by introducing a boundary prediction branch as an auxiliary. To avoid the loss of details, the boundary prediction branch is designed to run at full resolution with a shallow architecture, while some Semantic Enhancement Modules (SEMs) are used to supplement high-level semantic information by introducing multi-level encoder features of the cloud detection branch. This feature sharing in turn drives the cloud detection branch to focus more on cloud boundaries during training. At the end of the network, a Bilateral Fusion Module (BFM) is added for information complementarity between features from these two branches. The features from the cloud detection branch provide multi-scale features to the boundary prediction branch for more accurate boundary prediction, while the features from the boundary prediction branch further serve as prior knowledge to help the cloud detection branch aggregate contextual information. To verify the effectiveness of the proposed method, we select four different networks as cloud detection branches and conduct comparative experiments on two public datasets, GF-1 WFV and MODIS. The experimental results show that the proposed method significantly enhances cloud detection in confusing areas. Xiang Zhang 0018, Nailiang Kuang, Hangzai Luo, Sheng Zhong 0006, Jianping Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Automatic learning for object detection
Xiang Zhang 0018, Hangzai Luo, Wanqing Zhao, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 5 |
| 2022 | Detail-Aware Multiscale Context Fusion Network for Cloud DetectionabstractIn recent years, a large number of convolutional neural networks-based cloud detection algorithms have been proposed for remote sensing image preprocessing and most of them have an encoder-decoder structure. However, downsampling and upsampling operations, as the basic components of these methods, inevitably lead to the loss of detailed information in high-level features, which affects cloud detection performance. At the same time, the physical characteristics of the cloud, such as the variable size and irregular structure, also put forward requirements for the multi-scale feature representation ability of the network. To this end, we propose a novel cloud detection network named DMNet, which contains a Dense Feature Enhancement Module (DFEM) and a Multi-scale Context Fusion Spatial Attention Module (MCFSAM). DFEM aims to achieve information complementarity by exploiting the different properties of the features at different levels of the encoder, so as to strengthen the detailed information of high-level features and make the low-level features have more semantics. MCFSAM introduces a Multi-scale Context Fusion Block (MCFB) in spatial attention, which enables the network to densely capture contextual information at different scales and further emphasize useful features in the spatial dimension. Extensive experiments on GF-1 wide field-of-view satellite imagery (GF-1 WFV) dataset and Moderate-Resolution Imaging Spectroradiometer (MODIS) dataset demonstrate that our method outperforms other state-of-the-art cloud detection algorithms. Xiang Zhang 0018, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Hierarchical bilinear convolutional neural network for image classificationabstractAbstract Image classification is one of the mainstream tasks of computer vision. However, the most existing methods use labels of the same granularity level for training. This leads to ignoring the hierarchy that may help to differentiate different visual objects better. Embedding hierarchical information into the convolutional neural networks (CNNs) can effectively regulate the semantic space and thus reduce the ambiguity of prediction. To this end, a multi‐task learning framework, named as Hierarchical Bilinear Convolutional Neural Network (HB‐CNN), is developed by seamlessly integrating CNNs with multi‐task learning over the hierarchical visual concept structures. Specifically, the labels with a tree structure are used as the supervision to hierarchically train multiple branch networks. In this way, the model can not only learn additional information (e.g. context information) as the coarse‐level category features, but also focus the learned fine‐level category features on the object properties. To smoothly pass hierarchical conceptual information and encourage feature reuse, a connectivity pattern is proposed to connect features at different levels. Furthermore, a bilinear module is embedded to generalise various orderless texture feature descriptors so that our model can capture more discriminative features. The proposed method is extensively evaluated on the CIFAR‐10, CIFAR‐100, and ‘Orchid’ Plant image sets. The experimental results show the effectiveness and superiority of our method. Xiang Zhang 0018, Hangzai Luo, Sheng Zhong 0006, Ziyu Guan, Long Chen 0007, Jinye Peng 0001, Jianping Fan 0001 |
IET Comput. Vis. | 4 |
| 2021 | Learning noise-decoupled affine models for extreme low-light image enhancement
Maomei Liu, Sheng Zhong 0006, Hangzai Luo, Jinye Peng 0001 |
Neurocomputing | 3 |
| 2021 | Attention adjacency matrix based graph convolutional networks for skeleton-based action recognition
Qiguang Miao, Ruyi Liu 0001, Wentian Xin, Sheng Zhong 0006, Xuesong Gao |
Neurocomputing | 6 |
| 2014 | A multimodal investigation of in vivo muscle behavior: System design and data analysisabstractThe study is aimed to investigate in vivo behaviors of the rectus femoris muscle during isometric contraction by integrating simultaneously recorded electromyography (EMG), mechanomyography (MMG), and ultrasonography (US). We developed an experimental platform for simultaneous acquisition of EMG, MMG, US, as well as the torque, during isometric muscle contraction. Features from multimodal signals and images were then automatically extracted and calibrated to present time-varying characteristics of muscle behaviors. We further applied local polynomial regression (LPR) to reveal nonlinear and transient relationships between multimodal muscle features and torque. The results suggested that the proposed multimodal signal acquisition and integration are capable of providing novel and complete information about in vivo muscle contraction. The proposed experimental platform is a potentially useful tool for muscle assessment in various clinical and practical applications. Xin Chen 0025, Sheng Zhong 0006, Yangyang Niu, Siping Chen, Tianfu Wang 0001, S. C. Chan 0001, Zhiguo Zhang 0001 |
ISCAS | 2 |