Gang Wu 0010

dblp:99/6515-10 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
20since 2021 · last 2026
0009-0007-5003-3117ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GLoMOT: Efficient Online GNN-based Low-Frame-Rate Multi-Object Tracker
abstract
Low-frame-rate (LFR) Multi-Object Tracking (MOT) is crucial for efficient tracking on edge devices, as it significantly reduces computational and storage demands. However, existing trackers struggle in LFR settings due to large temporal gaps, extreme appearance changes, and motion non-linearity. While Graph Neural Network (GNN)-based trackers are effective at associating objects across these gaps, most operate offline, which prevents their use for online tracking. To address these limitations, we propose GLoMOT, a novel online GNN-based Low-Frame-Rate Multi-Object Tracker designed for robust performance in LFR videos. To bridge the large temporal gaps, we introduce a Dynamic Node Buffer Pool. This acts as a long-term memory, caching the states of absent objects to enable their robust re-association. To tackle extreme motion uncertainty, we propose an adaptive context-aware module that dynamically adjusts the weights of positional and appearance features, generating more robust features for predicting node connections. Furthermore, we propose a pseudo-depth feature calculation method. This provides the GNN with critical geometric context, which helps resolve spatial ambiguity arising from occlusions. Extensive experiments on several public MOT benchmarks, including DanceTrack, MOT17, and VisDrone, demonstrate GLoMOT's effectiveness and superiority, particularly in challenging Low-Frame-Rate conditions.
Yaxuan Hu 0001, Jie Hua 0005, Gang Wu 0010, Yuhong Yang 0001, Atsushi Suzuki 0002, Zhongyuan Wang 0001
AAAI3
2026 Learning from History: Task-agnostic Model Contrastive Learning for Image Restoration
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Wangmeng Zuo
Int. J. Comput. Vis.1
2026 Beyond Degradation Redundancy: Contrastive Prompt Learning for All-in-One Image Restoration
abstract
All-in-one image restoration, addressing diverse degradation types with a unified model, presents significant challenges in designing task-aware prompts that effectively guide restoration across multiple degradation scenarios. While adaptive prompt learning enables end-to-end optimization, it often yields overlapping or redundant task representations. Conversely, explicit prompts derived from pretrained classifiers enhance discriminability but may discard critical visual information for reconstruction. To address these limitations, we introduce Contrastive Prompt Learning (CPL), a novel framework that fundamentally enhances prompt-task alignment through two complementary innovations: a Sparse Prompt Module (SPM) that efficiently captures degradation-specific features while minimizing redundancy, and a Contrastive Prompt Regularization (CPR) that explicitly strengthens task boundaries by incorporating negative prompt samples across different degradation types. Unlike previous approaches that focus primarily on degradation classification, CPL optimizes the critical interaction between prompts and the restoration model itself. Extensive experiments across comprehensive benchmarks demonstrate that CPL consistently enhances state-of-the-art all-in-one restoration models, achieving significant improvements in both standard multi-task scenarios and challenging composite degradation settings. Our framework establishes new state-of-the-art performance while maintaining parameter efficiency, offering a principled solution for unified image restoration. The code is available at https://github.com/Aitical/CPLIR.
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 DSwinIR: Rethinking Window-Based Attention for Image Restoration
abstract
Image restoration has witnessed significant advancements with the development of deep learning models. Transformer-based models, particularly those using window-based self-attention, have become a dominant force. However, their performance is constrained by the rigid, non-overlapping window partitioning scheme, which leads to insufficient feature interaction across windows and limited receptive fields. This highlights the need for more adaptive and flexible attention mechanisms. In this paper, we propose the Deformable Sliding Window Transformer for Image Restoration (DSwinIR), a new attention mechanism: the Deformable Sliding Window (DSwin) Attention. This mechanism introduces a token-centric and content-aware paradigm that moves beyond the grid and fixed window partition. It comprises two complementary components. First, it replaces the rigid partitioning with a token-centric sliding window paradigm, making it effective at eliminating boundary artifacts. Second, it incorporates a content-aware deformable sampling strategy, which allows the attention mechanism to learn data-dependent offsets and actively shape its receptive field to focus on the most informative image regions. Extensive experiments show that DSwinIR achieves strong results, including state-of-the-art performance on several evaluated benchmarks. For instance, in all-in-one image restoration, our DSwinIR surpasses the most recent backbone GridFormer by 0.53 dB on the three-task benchmark and 0.87 dB on the five-task benchmark.
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 MSTDNet: Multi-scale traffic object detection network with smooth information perception
Jie Hua 0005, Zhongyuan Wang 0001, Hua Zou 0002, Gang Wu 0010, Jiayi Ma 0001
Pattern Recognit.5
2025 Debiased All-in-one Image Restoration with Task Uncertainty Regularization
abstract
All-in-one image restoration is a fundamental low-level vision task with significant real-world applications. The primary challenge lies in addressing diverse degradations within a single model. While current methods primarily exploit task prior information to guide the restoration models, they typically employ uniform multi-task learning, overlooking the heterogeneity in model optimization across different degradation tasks. To eliminate the bias, we propose a task-aware optimization strategy, that introduces adaptive task-specific regularization for multi-task image restoration learning. Specifically, our method dynamically weights and balances losses for different restoration tasks during training, encouraging the implementation of the most reasonable optimization route. In this way, we can achieve more robust and effective model training. Notably, our approach can serve as a plug-and-play strategy to enhance existing models without requiring modifications during inference. Extensive experiments in diverse all-in-one restoration settings demonstrate the superiority and generalization of our approach. For example, AirNet retrained with TUR achieves average improvements of 1.16 dB on three distinct tasks and 1.81 dB on five distinct all-in-one tasks. These results underscore TUR's effectiveness in advancing the SOTAs in all-in-one image restoration, paving the way for more robust and versatile image restoration.
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005
AAAI1
2025 DH-YOLO: Double-Head YOLO for Detecting Refrigerant Leakage Smoke in Air Conditioners
abstract
In the process of air conditioner dismantling and recycling at environmental protection companies, there is a risk of refrigerant leak smoke emissions. The release of fluorinated gases poses threats to both the environment and human health, making the detection of refrigerant leak smoke during dismantling crucial for environmental protection. However, existing smoke detection algorithms struggle to achieve high detection accuracy due to the low density and subtle characteristics of refrigerant leak smoke. In this paper, we observe the characteristics of leak smoke, which frequently appears near workers but rarely overlaps with them, thus adding a worker_fog co-location label to the previously established refrigerant leak detection dataset, ACRL-10K. This label assists the model in precisely localizing refrigerant leak smoke. To leverage the spatial correlation between workers and leak smoke, we propose a double-head yolo model, DHYOLO. The first detection head performs initial predictions on worker_fog regions, and the second head refines these predictions within the initially identified areas, enabling focused detection of refrigerant leak smoke. Additionally, we propose a multi-branch aggregation module, VoVRepGSCSP, which uses a multi-branch structure with convolutional kernels of various sizes to effectively capture image details, therefore enhancing YOLOv11’s sensitivity to refrigerant leak smoke. Experimental results on the supplemented ACRL-10K dataset demonstrate that DH-YOLO achieves an AP50 of up to 81.5% in leak detection tasks. Compared with the original YOLOv11, DH-YOLO improves AP50 across all sized models by 1.2% to 3.0%. In industrial applications, DH-YOLO shows promise for refrigerant leak warning in air conditioner dismantling processes.
Jie Hua 0005, Jinbi Liang, Gang Wu 0010
IJCNN5
2025 A Robust 3D CNN with Pyramidal Attention for Spatiotemporal Gait Recognition
abstract
Gait recognition has become an increasingly important biometric technique for identifying individuals from a distance without requiring their active cooperation. Since gait involves a sequence of motion patterns, effectively capturing temporal dynamics is essential for accurate recognition. Traditional methods that extract temporal features independently and fuse them at a later stage often fail to model the continuity and interdependence of motion across frames. To overcome this limitation, we propose a novel three-dimensional convolutional architecture named Robust Spatiotemporal 3D Convolutional Neural Network (RST3D), which jointly captures spatial and temporal correlations throughout gait sequences. The proposed architecture incorporates a comprehensive 3D convolutional block that operates along the temporal, height, and width dimensions, enabling the network to learn more expressive and coherent spatiotemporal representations. In addition, we introduce a Temporal Pyramidal Attention (TPA) block to enhance the network’s ability to model temporal dependencies by capturing discriminative motion patterns across multiple temporal scales. We evaluate our method on four large-scale gait recognition datasets: CASIA-B, OUMVLP, GREW, and Gait3D. Experimental results show that our approach consistently achieves superior performance compared to existing 3D CNN-based methods, particularly under challenging conditions such as view variation, clothing changes, and occlusion.
Jianyu Chen 0008, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Zengmin Xu, Gang Wu 0010, Zhongyuan Wang 0001
MMAsia6
2025 Multi-Modal Gait Recognition via Collaborative Feature Learning from Silhouettes and Skeletons
Jianyu Chen 0008, Zhongyuan Wang 0001, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Gang Wu 0010
PRCV (15)6
2025 A Survey on All-in-One Image Restoration: Taxonomy, Evaluation and Future Trends
abstract
Image restoration (IR) seeks to recover high-quality images from degraded observations caused by a wide range of factors, including noise, blur, compression, and adverse weather. While traditional IR methods have made notable progress by targeting individual degradation types, their specialization often comes at the cost of generalization, leaving them ill-equipped to handle the multifaceted distortions encountered in real-world applications. In response to this challenge, the all-in-one image restoration (AiOIR) paradigm has recently emerged, offering a unified framework that adeptly addresses multiple degradation types. These innovative models enhance the convenience and versatility by adaptively learning degradation-specific features while simultaneously leveraging shared knowledge across diverse corruptions. In this survey, we provide the first in-depth and systematic overview of AiOIR, delivering a structured taxonomy that categorizes existing methods by architectural designs, learning paradigms, and their core innovations. We systematically categorize current approaches and assess the challenges these models encounter, outlining research directions to propel this rapidly evolving field. To facilitate the evaluation of existing methods, we also consolidate widely-used datasets, evaluation protocols, and implementation practices, and compare and summarize the most advanced open-source models. As the first comprehensive review dedicated to AiOIR, this paper aims to map the conceptual landscape, synthesize prevailing techniques, and ignite further exploration toward more intelligent, unified, and adaptable visual restoration systems.
Junjun Jiang, Zengyuan Zuo, Gang Wu 0010, Kui Jiang, Xianming Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Learning Dynamic Prompts for All-in-One Image Restoration
abstract
All-in-one image restoration, which seeks to handle multiple types of degradation within a unified model, has become a prominent research topic in computer vision. While existing deep learning models have achieved remarkable success in specific restoration tasks, extending these models to heterogenous degradations presents significant challenges. Current all-in-one methods predominantly concentrate on extracting degradation priors, often employing learned and fixed task prompts to guide the restoration process. However, these static prompts are inclined to generate an average distribution characteristics of degradations, unable to accurately depict the unique attribute of the given input, consequently providing suboptimal restoration results. To tackle these challenges, we propose a novel dynamic prompt approach called Degradation Prototype Assignment and Prompt Distribution Learning (DPPD). Our approach decouples the degradation prior extraction into two novel components: Degradation Prototype Assignment (DPA) and Prompt Distribution Learning (PDL). DPA anchors the degradation representations to predefined prototypes, providing discriminative and scalable representations. In addition, PDL models prompts as distributions rather than fixed parameters, facilitating dynamic and adaptive prompt sampling. Extensive experiments demonstrate that our DPPD framework can achieve significant performance improvement on different image restoration tasks. Codes are available at our project page https://github.com/Aitical/DPPD.
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005, Liqiang Nie
IEEE Trans. Image Process.1
2025 DiffusionMOT: A Diffusion-Based Multiple Object Tracker
abstract
Recently, researchers have introduced diffusion models into multiple object tracking (MOT) tasks. However, existing diffusion-based MOT methods, such as DiffusionTrack, have significant limitations, including frequent ID switching, reduced performance when tracking nonlinear motion objects, and long inference time. To this end, we propose a more effective diffusion-based multiple object tracker named DiffusionMOT. In particular, we propose a mixed intersection over union (IoU) and Re-Identification (ReID) method for trajectory matching, which effectively reduces incorrect matches. Meanwhile, we propose a secondary calibration method for trajectory boxes, improving the accuracy of the generated detection boxes. Moreover, we introduce the parallel sampling technique from the field of image generation into object tracking and propose a parallel sampling module to enhance the model's inference speed while maintaining tracking accuracy. Furthermore, we design a pair-based two-stage matching (PTM) pipeline to more effectively utilize potential detection information. Extensive experiments on several public MOT benchmarks, including DanceTrack, SportsMOT, MOT20, and MOT17, demonstrate that our approach achieves state-of-the-art (SOTA) performance. The code and models are available at https://github.com/sad123-yx/DiffusionMOT.
Yaxuan Hu 0001, Jie Hua 0005, Zhen Han 0002, Hua Zou 0002, Gang Wu 0010, Zhongyuan Wang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Learning from History: Task-agnostic Model Contrastive Learning for Image Restoration
abstract
Contrastive learning has emerged as a prevailing paradigm for high-level vision tasks, which, by introducing properly negative samples, has also been exploited for low-level vision tasks to achieve a compact optimization space to account for their ill-posed nature. However, existing methods rely on manually predefined and task-oriented negatives, which often exhibit pronounced task-specific biases. To address this challenge, our paper introduces an innovative method termed 'learning from history', which dynamically generates negative samples from the target model itself. Our approach, named Model Contrastive Learning for Image Restoration (MCLIR), rejuvenates latency models as negative models, making it compatible with diverse image restoration tasks. We propose the Self-Prior guided Negative loss (SPN) to enable it. This approach significantly enhances existing models when retrained with the proposed model contrastive paradigm. The results show significant improvements in image restoration across various tasks and architectures. For example, models retrained with SPN outperform the original FFANet and DehazeFormer by 3.41 and 0.57 dB on the RESIDE indoor dataset for image dehazing. Similarly, they achieve notable improvements of 0.47 dB on SPA-Data over IDT for image deraining and 0.12 dB on Manga109 for a 4x scale super-resolution over lightweight SwinIR, respectively. Code and retrained models are available at https://github.com/Aitical/MCLIR.
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005
AAAI1
2024 Improving Domain Generalization in Self-supervised Monocular Depth Estimation via Stabilized Adversarial Training
Yuanqi Yao, Gang Wu 0010, Kui Jiang, Siao Liu, Jian Kuai, Xianming Liu 0005, Junjun Jiang
ECCV (24)2
2024 Zero-Mean Regularized Spectral Contrastive Learning: Implicitly Mitigating Wrong Connections in Positive-Pair Graphs
abstract
Contrastive learning has emerged as a popular paradigm of self-supervised learning that learns representations by encouraging representations of positive pairs to be similar while representations of negative pairs to be far apart. The spectral contrastive loss, in synergy with the notion of positive-pair graphs, offers valuable theoretical insights into the empirical successes of contrastive learning. In this paper, we propose incorporating an additive factor into the term of spectral contrastive loss involving negative pairs. This simple modification can be equivalently viewed as introducing a regularization term that enforces the mean of representations to be zero, which thus is referred to as *zero-mean regularization*. It intuitively relaxes the orthogonality of representations between negative pairs and implicitly alleviates the adverse effect of wrong connections in the positive-pair graph, leading to better performance and robustness. To clarify this, we thoroughly investigate the role of zero-mean regularized spectral contrastive loss in both unsupervised and supervised scenarios with respect to theoretical analysis and quantitative evaluation. These results highlight the potential of zero-mean regularized spectral contrastive learning to be a promising approach in various tasks.
Xianming Liu 0005, Feilong Zhang 0002, Gang Wu 0010, Deming Zhai, Junjun Jiang, Xiangyang Ji
ICLR4
2024 Exploiting Self-Supervised Constraints in image Super-Resolution
abstract
Recent advances in self-supervised learning, predominantly studied in high-level visual tasks, have been explored in low-level image processing. This paper introduces a novel self-supervised constraint for single image super-resolution, termed SSC-SR. SSC-SR uniquely addresses the divergence in image complexity by employing a dual asymmetric paradigm and a target model updated via exponential moving average to enhance stability. The proposed SSC-SR framework works as a plug-and-play paradigm and can be easily applied to existing SR models. Empirical evaluations reveal that our SSC-SR framework delivers substantial enhancements on a variety of benchmark datasets, achieving an average increase of 0.1 dB over EDSR and 0.06 dB over SwinIR. In addition, extensive ablation studies corroborate the effectiveness of each component in our SSC-SR framework. Codes are available at https://github.com/Aitical/SSCSR.
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005
ICME1
2024 Harmony in Diversity: Improving All-in-One Image Restoration via Multi-Task Collaboration
abstract
Deep learning-based all-in-one image restoration methods have garnered significant attention in recent years due to capable of addressing multiple degradation tasks. These methods focus on extracting task-oriented information to guide the unified model and have achieved promising results through elaborate architecture design. They commonly adopt a simple mix training paradigm, and the proper optimization strategy for all-in-one tasks has been scarcely investigated. This oversight neglects the intricate relationships and potential conflicts among various restoration tasks, consequently leading to inconsistent optimization rhythms. In this paper, we extend and redefine the conventional all-in-one image restoration task as a multi-task learning problem and propose a straightforward yet effective active-reweighting strategy, dubbed Art, to harmonize the optimization of multiple degradation tasks. Art is a plug-and-play optimization strategy designed to mitigate hidden conflicts among multi-task optimization processes. Through extensive experiments on a diverse range of all-in-one image restoration settings, Art has been demonstrated to substantially enhance the performance of existing methods. When incorporated into the AirNet and TransWeather models, it achieves average improvements of 1.16 dB and 1.21 dB on PSNR, respectively. We hope this work will provide a principled framework for collaborating multiple tasks in all-in-one image restoration and pave the way for more efficient and effective restoration models, ultimately advancing the state-of-the-art in this critical research domain. Code and pre-trained models are available at our project page https://github.com/Aitical/Art.
Gang Wu 0010, Junjun Jiang, Kui Jiang, Xianming Liu 0005
ACM Multimedia1
2024 Transforming Image Super-Resolution: A ConvFormer-Based Efficient Approach
abstract
Recent progress in single-image super-resolution (SISR) has achieved remarkable performance, yet the computational costs of these methods remain a challenge for deployment on resource-constrained devices. In particular, transformer-based methods, which leverage self-attention mechanisms, have led to significant breakthroughs but also introduce substantial computational costs. To tackle this issue, we introduce the Convolutional Transformer layer (ConvFormer) and propose a ConvFormer-based Super-Resolution network (CFSR), offering an effective and efficient solution for lightweight image super-resolution. The proposed method inherits the advantages of both convolution-based and transformer-based approaches. Specifically, CFSR utilizes large kernel convolutions as a feature mixer to replace the self-attention module, efficiently modeling long-range dependencies and extensive receptive fields with minimal computational overhead. Furthermore, we propose an edge-preserving feed-forward network (EFN) designed to achieve local feature aggregation while effectively preserving high-frequency information. Extensive experiments demonstrate that CFSR strikes an optimal balance between computational cost and performance compared to existing lightweight SR methods. When benchmarked against state-of-the-art methods such as ShuffleMixer, the proposed CFSR achieves a gain of 0.39 dB on the Urban100 dataset for the x2 super-resolution task while requiring 26% and 31% fewer parameters and FLOPs, respectively. The code and pre-trained models are available at https://github.com/Aitical/CFSR.
Gang Wu 0010, Junjun Jiang, Junpeng Jiang, Xianming Liu 0005
IEEE Trans. Image Process.1
2024 A Practical Contrastive Learning Framework for Single-Image Super-Resolution
abstract
Contrastive learning has achieved remarkable success on various high-level tasks, but there are fewer contrastive learning-based methods proposed for low-level tasks. It is challenging to adopt vanilla contrastive learning technologies proposed for high-level visual tasks to low-level image restoration problems straightly. Because the acquired high-level global visual representations are insufficient for low-level tasks requiring rich texture and context information. In this article, we investigate the contrastive learning-based single-image super-resolution (SISR) from two perspectives: positive and negative sample construction and feature embedding. The existing methods take naive sample construction approaches (e.g., considering the low-quality input as a negative sample and the ground truth as a positive sample) and adopt a prior model (e.g., pretrained very deep convolutional networks proposed by visual geometry group (VGG) model) to obtain the feature embedding. To this end, we propose a practical contrastive learning framework for SISR (PCL-SR). We involve the generation of many informative positive and hard negative samples in frequency space. Instead of utilizing an additional pretrained network, we design a simple but effective embedding network inherited from the discriminator network, which is more task-friendly. Compared with the existing benchmark methods, we retrain them by our proposed PCL-SR framework and achieve superior performance. Extensive experiments have been conducted to show the effectiveness and technical contributions of our proposed PCL-SR thorough ablation studies. The code and resulting models will be released via https://github.com/Aitical/PCL-SISR.
Gang Wu 0010, Junjun Jiang, Xianming Liu 0005
IEEE Trans. Neural Networks Learn. Syst.1
2023 No One Idles: Efficient Heterogeneous Federated Learning with Parallel Edge and Server Computation
abstract
Federated learning suffers from a latency bottleneck induced by network stragglers, which hampers the training efficiency significantly. In addition, due to the heterogeneous data distribution and security requirements, simple and fast averaging aggregation is not feasible anymore. Instead, complicated aggregation operations, such as knowledge distillation, are required. The time cost for complicated aggregation becomes a new bottleneck that limits the computational efficiency of FL. In this work, we claim that the root cause of training latency actually lies in the aggregation-then-broadcasting workflow of the server. By swapping the computational order of aggregation and broadcasting, we propose a novel and efficient parallel federated learning (PFL) framework that unlocks the edge nodes during global computation and the central server during local computation. This fully asynchronous and parallel pipeline enables handling complex aggregation and network stragglers, allowing flexible device participation as well as achieving scalability in computation. We theoretically prove that synchronous and asynchronous PFL can achieve a similar convergence rate as vanilla FL. Extensive experiments empirically show that our framework brings up to $5.56\times$ speedup compared with traditional FL. Code is available at: https://github.com/Hypervoyager/PFL.
Feilong Zhang 0002, Xianming Liu 0005, Gang Wu 0010, Junjun Jiang, Xiangyang Ji
ICML4