Junpeng Zhang 0002

dblp:55/10048-2 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0001-8068-6767ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 SA-MixNet: Structure-Aware Mixup and Invariance Learning for Scribble-Supervised Road Extraction in Remote Sensing Images
abstract
Mainstreamed weakly supervised road extractors rely on highly confident pseudo-labels propagated from scribbles, and their performance often degrades gradually as the image scenes tend to vary. We argue that such degradation is due to the poor model’s invariance to scenes with different complexities, whereas existing solutions to this problem are commonly based on crafted priors that cannot be derived from scribbles. To eliminate the reliance on such priors, we propose a novel structure-aware mixup and invariance learning framework (SA-MixNet) for weakly supervised road extraction that improves the model invariance in a data-driven manner. Specifically, we design a structure-aware mixup (SA-Mix) scheme to paste road regions from one image onto another to create an image scene with increased complexity while preserving the road’s structural integrity. Then, an invariance regularization is imposed on the predictions of constructed and origin images to minimize their conflicts, which thus forces the model to behave consistently in various scenes. Moreover, a discriminator-based regularization is designed to enhance connectivity while preserving the structure of roads. Combining these designs, our framework demonstrates superior performance on the DeepGlobe, Wuhan, and Massachusetts datasets, outperforming the state-of-the-art techniques by 1.47%, 2.12%, and 4.09%, respectively, in IoU metrics, and showing its potential as a plug-and-play solution. Our source code is available athttps://github.com/xdu-jjgs.
Jie Feng 0003, Junpeng Zhang 0002, Weisheng Dong, Dingwen Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2025 Oriented Vehicle Joint Detection and Tracking in Satellite Video via Identifier-Free Point Supervision
abstract
Oriented vehicle detection and tracking play a crucial role in various real-world applications. Yet, existing advanced models heavily rely on abundant and accurate oriented bounding box and tracking identifier annotations, which are extremely labor-intensive for satellite videos. In this paper, we endeavor to employ identifier-free point annotation to achieve the competitive performance while minimizing annotation costs. Specifically, each instance across video frames are labeled by single points, without providing its instance identifier. Building upon this setting, we introduce an oriented vehicle joint detection and tracking framework for satellite video, focusing on enhancing model performance by carefully-designed sample acquisition and robust learning processes. Firstly, we leverage temporal and visual information to generate sequence-aligned pseudo-labels and visually-aligned synthetic objects, which complement each other during training by providing both exact appearance and annotation information. Secondly, a novel spatio-temporal consistency metric is developed to assess sample quality, which is then incorporated into a curriculum learning schedule. This strategy facilitates a gradual learning progression from high-quality data to low-quality or noisy examples, thereyby minimizing interference from potentially misleading samples. Finally, an end-to-end oriented object joint detection and tracking network is constructed to enable effective oriented vehicle dynamic analysis. Extensive ablation and experimental results on two satellite video datasets demonstrate the superiority of our proposed method.
Yuping Liang, Jinjian Wu, Junpeng Zhang 0002, Yuxuan Chang, Jie Feng 0003, Guangming Shi
IEEE Trans. Geosci. Remote. Sens.3
2025 S4DL: Shift-Sensitive Spatial-Spectral Disentangling Learning for Hyperspectral Image Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) techniques, extensively studied in hyperspectral image (HSI) classification, aim to use labeled source domain data and unlabeled target domain data to learn domain invariant features for cross-scene classification. Compared to natural images, numerous spectral bands of HSIs provide abundant semantic information, but they also increase the domain shift significantly. In most existing methods, both explicit alignment and implicit alignment simply align feature distribution, ignoring domain information in the spectrum. We noted that when the spectral channel between source and target domains is distinguished obviously, the transfer performance of these methods tends to deteriorate. Additionally, their performance fluctuates greatly owing to the varying domain shifts across various datasets. To address these problems, a novel shift-sensitive spatial-spectral disentangling learning (S4DL) approach is proposed. In S4DL, gradient-guided spatial-spectral decomposition (GSSD) is designed to separate domain-specific and domain-invariant representations by generating tailored masks under the guidance of the gradient from domain classification. A shift-sensitive adaptive monitor is defined to adjust the intensity of disentangling according to the magnitude of domain shift. Furthermore, a reversible neural network is constructed to retain domain information that lies not only in semantic but also the shallow-level detailed information. Extensive experimental results on several cross-scene HSI datasets consistently verified that S4DL is better than the state-of-the-art UDA methods. Our source code will be available athttps://github.com/xdu-jjgs/IEEE_TNNLS_S4DL.
Jie Feng 0003, Junpeng Zhang 0002, Ronghua Shang, Weisheng Dong, Guangming Shi, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.3
2024 A Dual-Branch Network for End-to-End Point-Supervised Object Detection on Remote Sensing Images
abstract
Learning object detectors for remote sensing images commonly requires for a huge number of annotated boundary boxes, which are not available without enormous manual efforts in annotating. Alternatively, points can indicate the existence of the objects of interests with reduced labeling cost. Existing Point-supervised object detection (PSOD) methods predominantly employ a two-stage training strategy, which involves propagating point annotations to pseudo boxes at the first stage then training an object detector with these pseudo boxes in a fully supervised manner. However, such paradigm substantially impedes the end-to-end flow of training gradients. In this work, we propose a novel dual-branch network (DBNet) for end-to-end weakly supervised object detection on remote sensing images. Firstly, a pseudo box generation network is attached to the object detector as a sibling branch, which produces semantic response maps for the objects of interest then extracts pseudo boxes by examining their spatial connectivity. Then, instead of training this pseudo box generation network separately, we jointly adjust the pseudo box generation network and the detection network through a multi-task loss. Experimental results on the DOTA-v1.0 dataset demonstrate the effectiveness of our proposed method, achieving an average precision (mAP50) of 32.3%.
Jie Feng 0003, Junpeng Zhang 0002, Ronghua Shang, Xiangrong Zhang, Licheng Jiao
IGARSS3
2024 CFDRM: Coarse-to-Fine Dynamic Refinement Model for Weakly Supervised Moving Vehicle Detection in Satellite Videos
abstract
Deep learning methods have gradually developed into the mainstream methods of moving vehicle detection in satellite videos. However, these methods require labor-intensive and time-consuming box-level annotations to predict accurate locations and sizes, which is challenging for large-scale satellite video datasets with hundreds of vehicles. To address this problem, a novel coarse-to-fine dynamic refinement framework (CFDRM) is proposed for moving vehicle detection in satellite videos only under the supervision of point-level annotations. CFDRM generates initial proposal boxes and performs spatio-guided matching with point annotations to obtain coarse box-level pseudo annotations. The initial priority of these coarse annotations is calculated by leveraging locally-consistent prior tailored to satellite videos. Then, a dynamic refinement detector is constructed to transfer coarse annotations to fine annotations with prior and predictive collaborative curriculum refinement. During the curriculum learning process, the coarse annotations are sequentially learned with a certain priority, where the priority is inferred by considering the prior knowledge from the locally-consistent prior and the knowledge itself from the predicted detector. Ultimately, a novel ambiguity-aware loss is designed to optimize the dynamic refinement detector from coarse annotations to fine annotations in an adaptively-weighted fashion. Extensive experiments have been conducted on the Jilin-1 and SkySat satellite video datasets demonstrate the superiority of CFDRM.
Jie Feng 0003, Quanpeng Jiang, Junpeng Zhang 0002, Yuping Liang, Ronghua Shang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2024 Weakly Supervised Solar Panel Mapping via Uncertainty Adjusted Label Transition in Aerial Images
abstract
This paper proposes a novel uncertainty-adjusted label transition (UALT) method for weakly supervised solar panel mapping (WS-SPM) in aerial Images. In weakly supervised learning (WSL), the noisy nature of pseudo labels (PLs) often leads to poor model performance. To address this problem, we formulate the task as a label-noise learning problem and build a statistically consistent mapping model by estimating the instance-dependent transition matrix (IDTM). We propose to estimate the IDTM with a parameterized label transition network describing the relationship between the latent clean labels and noisy PLs. A trace regularizer is employed to impose constraints on the form of IDTM for its stability. To further reduce the estimation difficulty of IDTM, we incorporate uncertainty estimation to first improve the accuracy of noisy dataset distillation and then mitigate the negative impacts of falsely distilled examples with an uncertainty-adjusted re-weighting strategy. Extensive experiments and ablation studies on two challenging aerial data sets support the validity of the proposed UALT.
Jue Zhang 0001, Xiuping Jia, Jun Zhou 0001, Junpeng Zhang 0002, Jiankun Hu
IEEE Trans. Image Process.4
2023 MidNet: An Anchor-and-Angle-Free Detector for Oriented Ship Detection in Aerial Images
abstract
Ship detection in aerial images remains an active yet challenging task due to its arbitrary object orientation and various aspect ratios from the bird’s-eye perspective. Most existing oriented objection detection methods rely on angular prediction or predefined anchor boxes, making these methods highly sensitive to unstable angular regression and excessive hyper-parameter setting. To address these issues, we replace the angular-based object encoding with an anchor-and-angle-free paradigm, and propose a novel detector deploying a center and four midpoints for encoding each oriented object, namely MidNet. Moreover, MidNet designs a novel symmetrical deformable convolution for enhanceing the features of midpoints, then the center and midpoints for an identical ship are adaptively matched by predicting corresponding centripetal shift and matching radius. Finally, a concise analytical geometry algorithm is proposed to calculate the ship orientation and refine the keypoints step-wisely for building precise oriented bounding boxes. On two public ship detection datasets, HRSC2016 and FGSD2021, MidNet outperforms the state-of-the-art detectors by achieving APs of 90.52% and 86.50%.
Yuping Liang, Jie Feng 0003, Xiangrong Zhang, Junpeng Zhang 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2023 SDANet: Semantic-Embedded Density Adaptive Network for Moving Vehicle Detection in Satellite Videos
abstract
In satellite videos, moving vehicles are extremely small-sized and densely clustered in vast scenes. Anchor-free detectors offer great potential by predicting the keypoints and boundaries of objects directly. However, for dense small-sized vehicles, most anchor-free detectors miss the dense objects without considering the density distribution. Furthermore, weak appearance features and massive interference in the satellite videos limit the application of anchor-free detectors. To address these problems, a novel semantic-embedded density adaptive network (SDANet) is proposed. In SDANet, the cluster-proposals, including a variable number of objects, and centers are generated parallelly through pixel-wise prediction. Then, a novel density matching algorithm is designed to obtain each object via partitioning the cluster-proposals and matching the corresponding centers hierarchically and recursively. Meanwhile, the isolated cluster-proposals and centers are suppressed. In SDANet, the road is segmented in vast scenes and its semantic features are embedded into the network by weakly supervised learning, which guides the detector to emphasize the regions of interest. By this way, SDANet reduces the false detection caused by massive interference. To alleviate the lack of appearance information on small-sized vehicles, a customized bi-directional conv-RNN module extracts the temporal information from consecutive input frames by aligning the disturbed background. The experimental results on Jilin-1 and SkySat satellite videos demonstrate the effectiveness of SDANet, especially for dense objects.
Jie Feng 0003, Yuping Liang, Xiangrong Zhang, Junpeng Zhang 0002, Licheng Jiao
IEEE Trans. Image Process.4
2022 Moving Vehicle Detection for Remote Sensing Video Surveillance With Nonstationary Satellite Platform
abstract
With satellite platforms gazing at a target territory, the captured satellite videos exhibit local misalignment and local intensity variation on some stationary objects that can be mistakenly extracted as moving objects and increase false alarm rates. Typical approaches for mitigating the effect of moving cameras in moving object detection (MOD) follow domain transformation technique, where the misalignment between consecutive frames is restricted to the image planar. However, such technique cannot properly handle satellite videos, as the local misalignment on them is caused by the varying projections from the 3D objects on the Earth's surface to 2D image planar. In order to suppress the effect of moving satellite platform in MOD, we propose a Moving-Confidence-Assisted Matrix Decomposition (MCMD) model, where foreground regularization is designed to promote real moving objects and ignore system movements with the assistance of a moving-confidence score estimated from dense optical flows. For solving the convex optimization problem in MCMD, both batch processing and online solutions are developed in this study, by adopting the alternating direction method and the stochastic optimization strategy, respectively. Experimental results on the videos captured by SkySat and Jilin-1 show that MCMD outperforms the state-of-the-art techniques with improved precision by suppressing effect of nonstationary satellite platforms.
Junpeng Zhang 0002, Xiuping Jia, Jiankun Hu, Kun Tan 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Learning Via Watching: A Weakly Supervised Moving Object Detector for Satellite Videos
abstract
Moving Object Detection (MOD) from satellite videos plays one of the most fundamental roles in satellite video surveillance. Training a supervised moving object detector typically requires boundary box annotations for object instances, however, this annotation process is time-consuming for satellite videos. In this paper, we propose a weakly supervised method for sidestepping this process, where the supervised information for detecting moving objects on a frame is instead provided by the unsupervised method based on motion knowledge across a video. We adopt the Extended Low-rank and Structured Sparse Decomposition (E-LSD) approach for generating pixel-wise pseudo labels for moving objects. Then the extracted pseudo labels are used for training a Deep Convolutional Neural Network (DCNN) following the Encoder-Decoder architecture with lateral connections for segmenting moving objects from a new frame. We demonstrate the effectiveness of the proposed method on a satellite video dataset, and, compared with five state-of-the-art MOD methods tested, it achieves both improved detection accuracy and promising frame rate.
Junpeng Zhang 0002, Jue Zhang 0001, Xiuping Jia
IGARSS1
2020 Error Bounded Foreground and Background Modeling for Moving Object Detection in Satellite Videos
abstract
Detecting moving objects from ground-based videos is commonly achieved by using background subtraction (BS) techniques. Low-rank matrix decomposition inspires a set of state-of-the-art approaches for this task. It is integrated with structured sparsity regularization to achieve BS in the developed method of low-rank and structured sparse decomposition (LSD). However, when this method is applied to satellite videos where spatial resolution is poor and targets' contrast to the background is low, its performance is limited as the data no longer fit adequately either the foreground structure or the background model. In this article, we handle these unexplained data explicitly and address the moving target detection from space as one of the pioneering studies. We propose a new technique by extending the decomposition formulation with bounded errors, named Extended LSD (E-LSD). This formulation integrates low-rank background, structured sparse foreground, as well as their residuals in a matrix decomposition problem. Solving this optimization problem is challenging. We provide an effective solution by introducing an alternative treatment and adopting the direct extension of alternating direction method of multipliers (ADMM). The proposed E-LSD was validated on two satellite videos, and the experimental results demonstrate the improvement in background modeling with boosted moving object detection precision over state-of-the-art methods.
Junpeng Zhang 0002, Xiuping Jia, Jiankun Hu
IEEE Trans. Geosci. Remote. Sens.1
2020 Online Structured Sparsity-Based Moving-Object Detection From Satellite Videos
abstract
Inspired by the recent developments in computer vision, low-rank and structured sparse matrix decomposition can be potentially be used for extract moving objects in satellite videos. This set of approaches seeks for rank minimization on the background that typically requires batch-based optimization over a sequence of frames, which causes delays in processing and limits their applications. To remedy this delay, we propose an online low-rank and structured sparse decomposition (O-LSD). O-LSD reformulates the batch-based low-rank matrix decomposition with the structured sparse penalty to its equivalent framewise separable counterpart, which then defines a stochastic optimization problem for online subspace basis estimation. In order to promote online processing, O-LSD conducts the foreground and background separations and the subspace basis update alternatingly for every frame in a video. We also show the convergence of O-LSD theoretically. Experimental results on two satellite videos demonstrate the performance of O-LSD in terms of accuracy, and the time consumption is comparable with the batch-based approaches with significantly reduced delay in processing.
Junpeng Zhang 0002, Xiuping Jia, Jiankun Hu, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2019 A Drone-Based Sensing System to Support Satellite Image Analysis for Rice Farm Mapping
abstract
With supervised machine learning algorithms, meaningful information can be extracted from satellite images to support rice farm mapping. The success of these algorithms depends largely on the availability and quality of ground-truth reference data. However, collecting such data is often laborious and time-consuming. The fast development of drone technique has opened up an efficient way for ground truthing. In this study, we construct a drone-based sensing system for the purpose of efficiently collecting ground-truth data. The drone carries dual cameras that provide multispectrsal images with high spatial resolution in the same sensed site. The system is dedicated to collecting training data for rice farm mapping in Australia. To demonstrate the ability of the constructed system, real-flight experiments were conducted. Drone images of rice crops were acquired in the Riverina region of Australia, during the 2018-2019 summer season. Three-dimensional models were constructed from multiple images captured by the drone, where the structural information of rice crops was extracted. Results show that the configuration of dual cameras and the construction of three-dimensional models are particularly advantageous for ground truthing, providing valuable information for reliably identifying rice crops.
Yiqing Guo, Xiuping Jia, David Paull, Junpeng Zhang 0002, Adnan Farooq
IGARSS4
2019 Improved Low Rank plus Structured Sparsity and Unstructured Sparsity Decomposition for Moving Object Detection in Satellite Videos
abstract
Detecting moving objects from satellite videos decomposes video frames into low rank background with additive sparse foreground, and moving objects are commonly rec-ognized by pixel-wise sparsity in the foreground. In prac-tice, moving objects are sets of spatially related pixels and tend to present structured sparsity rather than pixel-wise sparsity. Based on this spatial prior, Structured Sparsity Inducing Norm models moving objects as sparse sets of neighboring pixels, however, the isolated sparsity, which is unsuited to the spatial prior, is left to corrupt the low rank background. In order to address the corruption of unstructured sparsity to background, we proposed a new decomposition formulation named Improved Low Rank plus Structured Sparsity and Unstructured Sparsity Decomposition (ILRSUSD). An inexact Alternating Direction Method is then proposed to solve the improved decomposition formulation efficiently. Experimental result on a satellite video dataset demonstrates our improvement in background modeling with boosted moving object detection precision against state-of-art approaches.
Junpeng Zhang 0002, Xiuping Jia
IGARSS1
2018 An Effective Zoom-In Approach for Detecting DIM and Small Target Proposals in Satellite Imagery
abstract
Satellite high definition videos provide an opportunity to monitor moving targets over a large territory. However, the low spatial resolution and low contract of these videos make target detecting and tracking a challenging task. In this paper, we propose a zoom-in approach for detecting dim and small target proposals from each single frame of the videos to help with moving target tracking. Initialized by a coarse scale segmentation approach, dim and small targets are embedded in each superpixel due to limited size and weak signals. Similar superpixels are then merged using a graph-based approach based on the measurement of the overlap between their histograms. The background statistics become stronger and target pixels are more obvious in the merged superpixels, so that the target pixels can be extracted. Finally, the corresponding boundary box is generated for each spatially connected target pixels selected inside each superpixel. They form the dim and small target proposals. Experimental results show that our zoom-in scheme can generate less proposals with higher recall rate compared with state-of-the-art proposal extraction algorithms.
Junpeng Zhang 0002, Xiuping Jia, Jiankun Hu
IGARSS1