Wensheng Cheng

dblp:242/3532 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0001-7822-5216ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Blood Flow Speed Estimation with Optical Coherence Tomography Angiography Images
abstract
Estimating blood flow speed is essential in many medical and physiological applications, yet it is extremely challenging due to complex vascular structure and flow dynamics, particularly for cerebral cortex regions. Existing techniques, such as Optical Doppler Tomography (ODT), generally require complex hardware control and signal processing, and still suffer from inherent system-level artifacts. To address these challenges, we propose a new learning-based approach named OCTA-Flow, which directly estimates vascular blood flow speed from Optical Coherence Tomography Angiography (OCTA) images that are commonly used for vascular structure analysis. OCTA-Flow employs several novel components to achieve this goal. First, using an encoder-decoder architecture, OCTA-Flow leverages ODT data as pseudo labels during training, thus bypassing the difficulty of collecting ground truth data. Second, to capture the relationship between vessels of varying scales and their flow speed, we design an Adaptive Window Fusion module that employs multiscale window attention. Third, to mitigate ODT artifacts, we incorporate a Conditional Random Field Decoder that promotes smoothness and consistency in the estimated blood flow. Together, these innovations enable OCTA-Flow to effectively produce accurate flow estimation, suppress the artifacts in ODT, and enhance practicality, benefiting from the established techniques of OCTA data acquisition. The code and data are available at https://github.com/Spritea/OCTA-Flow.
Wensheng Cheng, Jiaxiang Ren 0002, Hyomin Jeong, Congwu Du, Yingtian Pan, Haibin Ling
CVPR1
2025 Sparse Reconstruction of Optical Doppler Tomography with Alternative State Space Model and Attention
Jiaxiang Ren 0002, Wensheng Cheng, Yanzuo Liu, Congwu Du, Yingtian Pan, Haibin Ling
MICCAI (16)3
2024 E2TRACK: Overcoming Occlusion in Aerial Tracking with Trajectory Estimation
abstract
This paper introduces E2Track, a novel pipeline designed to overcome challenges in aerial tracking, particularly focusing on target occlusion scenarios. While template matching-based aerial tracking methods have shown progress in handling appearance changes, they often struggle in scenes with significant occlusion. We identify a fundamental issue leading to the limitations of existing methods in occlusion scenarios: the lack of temporal trajectory modeling. To address this, we propose to leverage the track information as a key solution to the occlusion problem. E2Track comprises two essential modules: the trajectory prediction module and the trajectory decoding module. The trajectory prediction module utilizes the entire video’s temporal contextual features to predict target trajectories. On the other hand, the trajectory decoding module performs feature queries on the predicted trajectories to determine the target’s position in each frame. This dual-module architecture enhances the model’s ability to handle occlusion scenarios by incorporating temporal context and leveraging trajectory information. To validate the efficacy of our proposed approach, we conducted experiments comparing E2Track with state-of-the-art trackers on two publicly available datasets. The results demonstrate the superiority of E2Track over competing methods, emphasizing its effectiveness in addressing the challenges posed by occluded scenes.
Wensheng Cheng
IGARSS4
2024 Decoupling Representation for Nighttime Aerial Tracking
abstract
Nighttime aerial tracking is an indispensable step towards around-the-clock real-world applications. However, RGBbased tracking algorithms face significant challenges at night due to their vulnerability to illumination. Observing that different feature channels have varying sensitivity to illumination, we propose to decouple the representation for illuminationsensitive and illumination-insensitive embeddings. We devise a Nighttime aerial tracking scheme via Decouple Representations, termed NiDR, where the Illumination-Invariant Embedding (IIE) module and the Illumination-Sensitive Embedding (ISE) module are designed to decouple representations. We achieve this semantic decoupling by utilizing a pair of normlight and low-light images and regulating the reconstruction and consistency relations between features. Experiments on UAVDark135 exhibit the remarkable performance of NiDR under challenging nighttime scenarios, surpassing the secondbest competitor by a large margin of 3.1% on precision.
Xu Lei 0002, Yan Zhang 0115, Chang Xu 0027, Wen Yang 0001, Wensheng Cheng
IGARSS5
2024 Self-supervised 3D Skeleton Completion for Vascular Structures
Jiaxiang Ren 0002, Wensheng Cheng, Zhilin Zou, Kicheon Park, Yingtian Pan, Haibin Ling
MICCAI (11)3
2024 NiDR: Nighttime Aerial Tracking via Decoupled Representations
abstract
Vanilla aerial trackers exhibit sensitivity to low-light conditions (e.g., nighttime aerial tracking scenario). To mitigate this, existing methods incorporate the light enhancement method as a preprocessing for aerial tracking. Despite the advancements, these approaches are restricted to the disparity in task objectives between the enhancer and tracker. Motivated by the observation that feature channels exhibit varying sensitivity to illumination, we propose to decouple the feature representation into two distinct parts: 1) illumination-invariant feature embedding and 2) illumination-sensitive feature embedding. The former, realized by the illumination invariant embedding (IIE) module, enhances features that remain invariant to illumination changes. Meanwhile, the latter, facilitated by the illumination sensitive embedding (ISE) module, aims to mitigate the negative impact of illumination-sensitive features on tracking performance. Building upon this decoupling strategy, we introduce NiDR, a simple yet effective nighttime aerial tracker. The proposed NiDR exhibits strong performance on three nighttime aerial tracking benchmarks (i.e., UAVDark135, NAT2021, and DarkTrack2021). Notably, it outperforms previous competitors by large margins, e.g., 3.1 points on the UAVDark135 and 2.0 points on the Darktrack2021 in terms of precision for nighttime scenarios.
Xu Lei 0002, Yan Zhang 0115, Chang Xu 0027, Wensheng Cheng, Wen Yang 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Toward High Quality Multi-Object Tracking and Segmentation Without Mask Supervision
abstract
Recently studies have shown the potential of weakly supervised multi-object tracking and segmentation, but the drawbacks of coarse pseudo mask label and limited utilization of temporal information remain to be unresolved. To address these issues, we present a framework that directly uses box label to supervise the segmentation network without resorting to pseudo mask label. In addition, we propose to fully exploit the temporal information from two perspectives. Firstly, we integrate optical flow-based pairwise consistency to ensure mask consistency across frames, thereby improving mask quality for segmentation. Secondly, we propose a temporally adjacent pair-based sampling strategy to adapt instance embedding learning for data association in tracking. We combine these techniques into an end-to-end deep model, named BoxMOTS, which requires only box annotation without mask supervision. Extensive experiments demonstrate that our model surpasses current state-of-the-art by a large margin, and produces promising results on KITTI MOTS and BDD100K MOTS. The source code is available at https://github.com/Spritea/BoxMOTS.
Wensheng Cheng, Zhenyu Wu 0002, Haibin Ling, Gang Hua 0001
IEEE Trans. Image Process.1
2023 A3Track: Achieving Precise Target Tracking in Aerial Images With Receptive Field Alignment
abstract
Tracking arbitrary objects in aerial images presents formidable challenges to existing trackers. Among these challenges, the large scale variation and arbitrary geometry shape of visual targets are pronounced, resulting in two-fold mismatch issues between the feature receptive field and the tracking target. For one, there is a mismatch between the prior receptive field center and arbitrary-shaped targets. For another, the single receptive field mismatches the significantly scale-varied targets in the aerial imagery. To handle these challenges, we propose to Achieve precise Aerial tracking with receptive field Alignment, dubbed A3Track. The proposed A3Track is comprised of two modules: a Receptive Field Alignment (RFA) module and a Pyramid Receptive Field (PRF) module. First of all, we transform and update the receptive field center progressively, which drives the feature sampling location onto the targets’ main body, thus gradually yielding precise feature representation for arbitrary-shaped targets. We term this progressively updating process as the Receptive Field Alignment. Moreover, the PRF module constructs a set of pyramid features for the target, providing a multi-scale receptive field to handle the large scale variation of tracking objects. On four benchmarks, the new tracker A3Track achieves leading performance compared with existing methods and shows consistent improvements over baselines. The project is available at: https://chnleixu.github.io/A3Track-web/.
Xu Lei 0002, Chang Xu 0027, Wensheng Cheng, Wen Yang 0001, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.3
2021 GMOT-40: A Benchmark for Generic Multiple Object Tracking
abstract
Multiple Object Tracking (MOT) has witnessed remarkable advances in recent years. However, existing studies dominantly request prior knowledge of the tracking target (eg, pedestrians), and hence may not generalize well to unseen categories. In contrast, Generic Multiple Object Tracking (GMOT), which requires little prior information about the target, is largely under-explored. In this paper, we make contributions to boost the study of GMOT in three aspects. First, we construct the first publicly available dense GMOT dataset, dubbed GMOT-40, which contains 40 carefully annotated sequences evenly distributed among 10 object categories. In addition, two tracking protocols are adopted to evaluate different characteristics of tracking algorithms. Second, by noting the lack of devoted tracking algorithms, we have designed a series of baseline GMOT algorithms. Third, we perform a thorough evaluations on GMOT-40, involving popular MOT algorithms (with necessary modifications) and the proposed baselines. The GMOT-40 benchmark is publicly available at https://github.com/Spritea/GMOT40.
Hexin Bai, Wensheng Cheng, Peng Chu, Juehuan Liu, Kai Zhang 0001, Haibin Ling
CVPR2
2020 Look at the Big Picture: Building Area Extraction with Global Density Map
abstract
The automatic extraction of building areas from high-resolution satellite imagery has become an important and challenging research issue. Many recent studies have explored different deep learning-based semantic segmentation methods for better accuracy. However, the deep network usually takes sliding window cropped satellite images as inputs, which loses the global information and causes a high false positive rate. In this paper, we propose a density map guided attention mechanism for building area extraction to make the network look at the big picture. We exploit an FCN-based building density prediction network to generate a density heatmap from large satellite images. The density factors in heatmap control the classifier's threshold of building area extraction network that optimize the FP and recall rates. Furthermore, we propose a test-time overlap augmentation mechanism to improve the segmentation results. Our method outperforms state-of-the-art approaches and increases mIoU by about 3.08% to 93.31%, and decreases FP rate to 0.91%.
Haowen Guo, Wensheng Cheng, Wen Yang 0001, Gui-Song Xia
IGARSS2
2019 Context Aggregation Network for Semantic Labeling in Aerial Images
abstract
Multi-scale object recognition and accurate object localization are two major problems for semantic segmentation in high resolution aerial images. To handle these problems, we design a Context Fuse Module to aggregate multi-scale features and propose an Attention Mix Module to combine different level features for higher localization accuracy. We further employ a Residual Convolutional Module to refine features in all levels. Based on these modules, we construct a new end-to-end network for semantic labeling in aerial images. Experiments demonstrate that our network outperforms other state-of-the-art models on the large-scale ISPRS Vaihingen 2D Semantic Labeling Challenge dataset. The model implementation code is made publicly available1.
Wensheng Cheng, Wen Yang 0001, Youqi Pan, Haowen Guo
ICIP1