Gang Wang 0031

dblp:71/4292-31 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-1916-6110ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systems
abstract
Unmanned Aerial Vehicles (UAVs) are widely adopted across various fields, yet they raise significant privacy and safety concerns, demanding robust monitoring solutions. Existing anti-UAV methods primarily focus on position tracking but fail to capture UAV behavior and intent. To address this, we introduce a novel task—UAV Tracking and Intent Understanding (UTIU)—which aims to track UAVs while inferring and describing their motion states and intent for a more comprehensive monitoring approach. To tackle the task, we propose JTD-UAV, the first joint tracking, and intent description framework based on large language models. Our dual-branch architecture integrates UAV tracking with Visual Question Answering (VQA), allowing simultaneous localization and behavior description. To benchmark this task, we introduce the TDUAV dataset, the largest dataset for joint UAV tracking and intent understanding, featuring 1,328 challenging video sequences, over 163K annotated thermal frames, and 3K VQA pairs. Our benchmark demonstrates the effectiveness of JTD-UAV.
Jian Zhao 0006, Zhaoxin Fan, Xin Zhang 0093, Yudian Zhang, Lei Jin 0003, Gang Wang 0031, Mengxi Jia, Xuelong Li 0001
CVPR9
2025 Tracking Tiny Drones Against Clutter: Large-Scale Infrared Benchmark with Motion-Centric Adaptive Algorithm
Zongli Jiang, Jinli Zhang, Yixin Wei, Liang Li 0006, Yizheng Wang, Gang Wang 0031
ICCV7
2025 A Chaotic Dynamics Framework Inspired by Dorsal Stream for Event Signal Processing
abstract
Event cameras are bio-inspired vision sensors that encode visual information with high dynamic range, high temporal resolution, and low latency. Current state-of-the-art event stream processing methods rely on end-to-end deep learning techniques. However, these models are heavily dependent on data structures, limiting their stability and generalization capabilities across tasks, thereby hindering their deployment in real-world scenarios. To address this issue, we propose a chaotic dynamics event signal processing framework inspired by the dorsal visual pathway of the brain. Specifically, we utilize Continuous-coupled Neural Network (CCNN) to encode the event stream. CCNN encodes polarity-invariant event sequences as periodic signals and polarity-changing event sequences as chaotic signals. We then use continuous wavelet transforms to analyze the dynamical states of CCNN neurons and establish the high-order mappings of the event stream. The effectiveness of our method is validated through integration with conventional classification networks, achieving state-of-the-art classification accuracy on the N-Caltech101 and N-CARS datasets, with results of 84.3% and 99.9%, respectively. Our method improves the accuracy of event camera-based object classification while significantly enhancing the generalization and stability of event representation.
Jing Lian 0001, Zhaofei Yu, Jizhao Liu, Jisheng Dang, Gang Wang 0031
ICML6
2025 Optical Flow Estimation for Tiny Objects: New Problem, Specialized Benchmark, and Bioinspired Scheme
abstract
Optical flow is pivotal in video-based tasks, yet existing methods mostly focus on medium-/large-size objects, while underperforming when characterizing the motion of tiny objects. To bridge this gap, we introduce the On-off Time-delay with Hassenstein-Reichardt correlator (OTHR), a computationally efficient scheme inspired by the primate visual cortex's direction selectivity mechanism. OTHR kernels, applied across multiple frames, discern bright/dark luminance changes along a specific direction over a time delay, effectively estimating motion of tiny objects amidst noise and static backgrounds. Notably, OTHR integrates seamlessly with leading deep learning flow estimation models such as RAFT and FlowFormer. We also propose refined evaluation metrics for tiny objects and contribute a new dataset featuring such objects to aid algorithm development. Our experiments confirm OTHR's superiority over competing methods, particularly in enhancing state-of-the-art models' performance on tiny object motion estimation at minimal cost. Specifically, for objects less than 100 pixels, OTHR reduces RAFT and FlowFormer's errors by 22.03% and 83.50%, respectively. The codes will be accessible at https://github.com/JaneEliot/OTHR.
Xueyao Ji, Gang Wang 0031, Yizheng Wang
IJCAI2
2025 Adaptive memory fusion for multi-frame optical flow estimation
Shangsheng Li, Gang Wang 0031, Yizheng Wang
Neurocomputing4
2024 Unified Single-Stage Transformer Network for Efficient RGB-T Tracking
Jianqiang Xia, Dian-xi Shi, Linna Song, Songchang Jin, Chenran Zhao, Yu Cheng 0009, Lei Jin 0003, Jianan Li 0001, Gang Wang 0031, Junliang Xing, Jian Zhao 0006
IJCAI12
2024 Anti-UAV410: A Thermal Infrared Benchmark and Customized Scheme for Tracking Drones in the Wild
abstract
The perception of drones, also known as Unmanned Aerial Vehicles (UAVs), particularly in infrared videos, is crucial for effective anti-UAV tasks. However, existing datasets for UAV tracking have limitations in terms of target size and attribute distribution characteristics, which do not fully represent complex realistic scenes. To address this issue, we introduce a generalized infrared UAV tracking benchmark called Anti-UAV410. The benchmark comprises a total of 410 videos with over 438 K manually annotated bounding boxes. To tackle the challenges of UAV tracking in complex environments, we propose a novel method called Siamese drone tracker (SiamDT). SiamDT incorporates a dual-semantic feature extraction mechanism that explicitly models targets in dynamic background clutter, enabling effective tracking of small UAVs. The SiamDT method consists of three key steps: Dual-Semantic RPN Proposals (DS-RPN), Versatile R-CNN (VR-CNN), and Background Distractors Suppression. These steps are responsible for generating candidate proposals, refining prediction scores based on dual-semantic features, and enhancing the discriminative capacity of the trackers against dynamic background clutter, respectively. Extensive experiments conducted on the Anti-UAV410 dataset and three other large-scale benchmarks demonstrate the superior performance of the proposed SiamDT method compared to recent state-of-the-art trackers.
Bo Huang 0012, Jianan Li 0001, Gang Wang 0031, Jian Zhao 0006, Tingfa Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Region-adaptive and context-complementary cross modulation for RGB-T semantic segmentation
Fengguang Peng, Gang Wang 0031, Tianrui Hui, Si Liu 0001
Pattern Recognit.4
2024 Memory-Guided Collaborative Attention for Nighttime Thermal Infrared Image Colorization of Traffic Scenes
abstract
Robust imaging under challenging conditions, such as starlit nights, has broadened the adoption of thermal infrared (TIR) cameras for nighttime driving scenes. Given that TIR images are monochromatic, which makes them difficult to interpret by humans and limits the applicability of RGB-based algorithms, it is reasonable to perform colorization of nighttime TIR (NTIR) images by converting them into corresponding daytime color images (NTIR2DC). Despite the impressive results achieved by previous NTIR2DC methods, how to improve the colorization performance of small-sample categories without semantic annotation is under-explored. To address this issue, we propose a novel learning framework called Memory-guided cOllaboRative atteNtion Generative Adversarial Network (MornGAN), which is inspired by the analogical reasoning mechanisms of humans. Specifically, we first propose an online semantic distillation module to mine and refine the semantic cues of NTIR images. Then, a memory-guided sample selection strategy and adaptive collaborative attention loss are devised to enhance the semantic preservation of small-sample categories. Further, a new conditional gradient repair loss is introduced for reducing edge distortion during translation. Extensive experiments on the NTIR2DC task show that the proposed MornGAN significantly outperforms other image-to-image translation methods in terms of semantic preservation and edge consistency, which helps improve the object detection accuracy remarkably.
Fuya Luo, Yijun Cao, Kaifu Yang, Gang Wang 0031, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.4
2024 Linker: Learning Long Short-term Associations for Robust Visual Tracking
abstract
iamese and Transformer trackers have demon strated exceptional performance in visual object tracking. These methods utilize initial and potentially online templates to locate the target in subsequent frames. Despite their success, these trackers are vulnerable to changes in the target's appearance due to slow template updates and interference from similar objects, resulting from the absence of scene information. To address these issues, we introduce a reference region within our tracker. The reference region is updated rapidly, providing short-term scene information. By associating the initial template, reference region, and current search region, we enhance the tracker's ability to adapt to changes in target appearance and discriminate between the target and other objects. Additionally, we propose a novel Reference-Enhance (RE) module, which aggregates contextually relevant information from the reference region to enhance the template feature. Extensive experiments show our method achieves state-of-the-art performance on six popular visual object tracking benchmarks while running at over 40 FPS.
Zizheng Xun, Shangzhe Di, Yulu Gao, Zongheng Tang, Gang Wang 0031, Si Liu 0001, Bo Li 0006
IEEE Trans. Multim.5
2023 ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual Tracking
abstract
Recently, the transformer has enabled the speed-oriented trackers to approach state-of-the-art (SOTA) performance with high-speed thanks to the smaller input size or the lighter feature extraction backbone, though they still substantially lag behind their corresponding performance-oriented versions. In this paper, we demonstrate that it is possible to narrow or even close this gap while achieving high tracking speed based on the smaller input size. To this end, we non-uniformly resize the cropped image to have a smaller input size while the resolution of the area where the target is more likely to appear is higher and vice versa. This enables us to solve the dilemma of attending to a larger visual field while retaining more raw information for the target despite a smaller input size. Our formulation for the non-uniform resizing can be efficiently solved through quadratic programming (QP) and naturally integrated into most of the crop-based local trackers. Comprehensive experiments on five challenging datasets based on two kinds of transformer trackers, \ie, OSTrack and TransT, demonstrate consistent improvements over them. In particular, applying our method to the speed-oriented version of OSTrack even outperforms its performance-oriented counterpart by 0.6\% AUC on TNL2K, while running 50\% faster and saving over 55\% MACs. Codes and models are available at https://github.com/Kou-99/ZoomTrack.
Yutong Kou, Bing Li 0001, Gang Wang 0031, Weiming Hu 0004, Yizheng Wang, Liang Li 0006
NeurIPS4
2022 Recent advances on image edge detection: A comprehensive review
Junfeng Jing, Shenjuan Liu, Gang Wang 0031, Changming Sun
Neurocomputing3
2022 Thermal Infrared Image Colorization for Nighttime Driving Scenes With Top-Down Guided Attention
abstract
Benefitting from insensitivity to light and high penetration of foggy environments, infrared cameras are widely used for sensing in nighttime traffic scenes. However, the low contrast and lack of chromaticity of thermal infrared (TIR) images hinder the human interpretation and portability of high-level computer vision algorithms. Colorization to translate a nighttime TIR image into a daytime color (NTIR2DC) image may be a promising way to facilitate nighttime scene perception. Despite recent impressive advances in image translation, semantic encoding entanglement and geometric distortion in the NTIR2DC task remain under-addressed. Hence, we propose a toP-down attEntion And gRadient aLignment based generative adversarial network, referred to as PearlGAN. A top-down guided attention module and an elaborate attentional loss are first designed to reduce the semantic encoding ambiguity during translation. Then, a structured gradient alignment loss is introduced to encourage edge consistency between the translated and input images. In addition, pixel-level annotation is carried out on a subset of FLIR and KAIST datasets to evaluate the semantic preservation performance of multiple translation methods. Furthermore, a new metric is devised to evaluate the geometric consistency in the translation process. Extensive experiments demonstrate the superiority of the proposed PearlGAN over other image translation methods for the NTIR2DC task. The source code and labeled segmentation masks will be available athttps://github.com/FuyaLuo/PearlGAN/.
Fuya Luo, Yunhan Li, Gang Wang 0031, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.5
2021 Automated detection and counting of Artemia using U-shaped fully convolutional networks and deep convolutional networks
Gang Wang 0031, Gilbert Van Stappen, Bernard De Baets
Expert Syst. Appl.1
2020 Boundary detection using unbiased sparseness-constrained colour-opponent response and superpixel contrast
abstract
Boundaries play a crucial role in various image‐based tasks, but many existing non‐learning‐based boundary detection methods underperform in recognising authentic boundaries from a complex background. In this study, the authors address this problem using the sparseness‐constrained colour‐opponent response and the superpixel contrast. First, building on the biologically inspired colour‐opponency mechanism, the authors elaborate a method to compute the unbiased sparseness‐constrained colour‐opponent response. In this procedure, locations showing colour variations are enhanced, while the textural locations are preliminarily suppressed by the cue of local sparseness measure. Second, with the help of superpixel segmentation, the authors present an effective approach to obtain the superpixel contrast map. This approach helps to exploit the object shape information in suppressing textures. Consequently, the authors propose a non‐learning‐based method to detect boundaries in images, combining the unbiased sparseness‐constrained colour‐opponent response and the overall superpixel contrast map. Experiment results on widely adopted datasets manifest that the authors method outperforms most of the competing methods. In particular, compared with the state‐of‐the‐art surround‐modulation method, the proposed method obtains a comparable performance while consuming much less runtime.
Gang Wang 0031, Yongguang Chen, Suochang Yang, Fuqiang Feng, Bernard De Baets
IET Image Process.1
2020 High-ISO Long-Exposure Image Denoising Based on Quantitative Blob Characterization
abstract
Blob detection and image denoising are fundamental, sometimes related tasks in computer vision. In this paper, we present a computational method to quantitatively measure blob characteristics using normalized unilateral second-order Gaussian kernels. This method suppresses non-blob structures while yielding a quantitative measurement of the position, prominence and scale of blobs, which can facilitate the tasks of blob reconstruction and blob reduction. Subsequently, we propose a denoising scheme to address high-ISO long-exposure noise, which sometimes spatially shows a blob appearance, employing a blob reduction procedure as a cheap preprocessing for conventional denoising methods. We apply the proposed denoising methods to real-world noisy images as well as standard images that are corrupted by real noise. The experimental results demonstrate the superiority of the proposed methods over state-of-the-art denoising methods.
Gang Wang 0031, Carlos Lopez-Molina, Bernard De Baets
IEEE Trans. Image Process.1
2019 Noise-robust line detection using normalized and adaptive second-order anisotropic Gaussian kernels
Gang Wang 0031, Carlos Lopez-Molina, Guillermo Vidal-Diez de Ulzurrun, Bernard De Baets
Signal Process.1
2017 Blob Reconstruction Using Unilateral Second Order Gaussian Kernels with Application to High-ISO Long-Exposure Image Denoising
abstract
Blob detection and image denoising are fundamental, and sometimes related, tasks in computer vision. In this paper, we propose a blob reconstruction method using scale-invariant normalized unilateral second order Gaussian kernels. Unlike other blob detection methods, our method suppresses non-blob structures while also identifying blob parameters, i.e., position, prominence and scale, thereby facilitating blob reconstruction. We present an algorithm for high-ISO long-exposure noise removal that results from the combination of our blob reconstruction method and state-of-the-art denoising methods, i.e., the non-local means algorithm (NLM) and the color version of block-matching and 3-D filtering (CBM3D). Experiments on standard images corrupted by real high-ISO long-exposure noise and real-world noisy images demonstrate that our schemes incorporating the blob reduction procedure outperform both the original NLM and CBM3D.
Gang Wang 0031, Carlos Lopez-Molina, Bernard De Baets
ICCV1