Hong Wang 0021

dblp:83/5522-21 · DBLP profile ↗
← Back
35ranked-venue papers
14as first author
34since 2021 · last 2026
0000-0002-6520-7681ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 9 first-author · 15 since 2021
YearPublicationVenuePosition
2026 A lightweight physics-driven dual exposure correction network for endoscopic images
Zhijian Wu, Hong Wang 0021, Yefeng Zheng 0001
Pattern Recognit.2
2026 Leveraging Text-Modulated Semantic Guidance for Low-Light Endoscopic Image Enhancement
abstract
Low light conditions in endoscopic imaging would lead to poor visibility, reduced contrast, and increased noise, which may hinder accurate diagnosis and surgical guidance. Against this low-light endoscopic image enhancement (LLEIE) task, inspired by the remarkable performance of pretrained CLIP in downstream vision tasks, in this paper, we carefully investigate the pretrained priors of CLIP and embed them into a text-modulated semantic-aware discriminator (TMSD). Through the adversarial learning mechanism, the discriminator can be easily integrated into different low-light enhancement baselines for helping them accomplish better visual restoration effects without incurring any extra inference cost. Specifically, to make the foundation model CLIP suitable for the LLEIE task, we initially propose a prompt learning procedure to obtain the text embedding and image semantics corresponding to the normal-light endoscopic imaging scenario. Building upon the acquired text prior and image semantic priors, we devise a text modulator to synergize these two priors, yielding a richer semantic representation. Leveraging the convolutional modulation and cross-attention mechanisms, we blend this semantic guidance information into the discriminator, thereby fostering the fine-grained distribution learning of normal-light endoscopic images in visual semantics and guiding different enhancement baselines achieving higher visual quality. Based on five public benchmark datasets, including three synthetic datasets, one real clinical dataset, and one clinical downstream segmentation dataset, we comprehensively evaluate the effectiveness of our proposed TMSD. Extensive experiments substantiate that the integration of the proposed TMSD enables seven representative baselines to obtain better perceptual quality, especially in the cross-domain clinical generalization scenario. Besides, the downstream segmentation accuracy can be evidently improved, showing the favorable application potential of the proposed TMSD. Moreover, to comprehensively evaluate the generality of our TMSD framework, we successfully apply it to a new and classic metal artifact reduction task. It is worth mentioning that our TMSD does not incur any extra computational cost during inference.
Hong Wang 0021, Zhijian Wu, Haodu Fang, Dong Wei 0004, Jinghan Sun, Yefeng Zheng 0001, Jianhua Ma 0001
IEEE Trans. Medical Imaging1
2026 A Refreshed Similarity-Based Upsampler for Direct High-Ratio Feature Upsampling
abstract
Feature upsampling is a fundamental and indispensable ingredient of almost all current network structures for dense prediction tasks. Very recently, a popular similarity-based feature upsampling pipeline has been proposed, which utilizes a high-resolution (HR) feature as guidance to help upsample the low-resolution (LR) deep feature based on their local similarity. Albeit achieving promising performance, this pipeline has specific limitations in methodological designs: 1) HR query and LR key features are not well aligned in a controllable manner; 2) the similarity between query-key features is computed based on the fixed inner product form, lacking flexibility; and 3) neighbor selection is coarsely operated on LR features, resulting in mosaic artifacts. These shortcomings make the existing methods along this pipeline primarily applicable to hierarchical network architectures with iterative features as guidance, and they are not readily extended to a broader range of structures, especially for a direct high-ratio upsampling. Against these issues, we thoroughly refresh this pipeline and meticulously optimize every methodological design. Specifically, we first propose an explicitly controllable query-key feature alignment from both semantic-aware and detail-aware perspectives and then construct a parameterized paired central difference convolution block for flexibly calculating the similarity between the well-aligned query-key features. Besides, we develop a fine-grained neighbor selection strategy on HR features, which is simple yet effective for alleviating mosaic artifacts. Based on these careful designs, we systematically construct a refreshed similarity-based feature upsampling framework named ReSFU. Based on 13 types of network backbones, comprehensive experiments substantiate that only in a simple and direct high-ratio upsampling manner, our ReSFU consistently achieves satisfactory performance on six tasks, including semantic segmentation, medical image segmentation, instance segmentation, panoptic segmentation, object detection, and monocular depth estimation, showing superior generality and ease of deployment beyond the existing upsamplers. Codes are available at https://github.com/zmhhmz/ReSFU.
Hong Wang 0021, Yefeng Zheng 0001, Deyu Meng
IEEE Trans. Neural Networks Learn. Syst.2
2025 A Simple and Better Baseline for Visual Grounding
abstract
Visual grounding aims to predict the locations of target objects specified by textual descriptions. For this task with linguistic and visual modalities, there is a latest research line that focuses on only selecting the linguistic-relevant visual regions for object localization to reduce the computational overhead. Albeit achieving impressive performance, it is iteratively performed on different image scales, and at every iteration, linguistic features and visual features need to be stored in a cache, incurring extra overhead. To facilitate the implementation, in this paper, we propose a feature selection-based simple yet effective baseline for visual grounding, called FSVG. Specifically, we directly encapsulate the linguistic and visual modalities into an overall network architecture without complicated iterative procedures, and utilize the language in parallel as guidance to facilitate the interaction between linguistic modal and visual modal for extracting effective visual features. Furthermore, to reduce the computational cost, during the visual feature learning, we introduce a similarity-based feature selection mechanism to only exploit language-related visual features for faster prediction. Extensive experiments conducted on several benchmark datasets comprehensively substantiate that the proposed FSVG achieves a better balance between accuracy and efficiency beyond the current state-of-the-art methods. Code is available at https://github.com/jcwang0602/FSVG.
Dingjiang Huang, Hong Wang 0021, Yefeng Zheng 0001
ICME4
2025 PD-INR: Prior-Driven Implicit Neural Representations for TOF-PET Reconstruction
Yuxuan Long, Hong Wang 0021, Xiaodong Kuang, Hailiang Huang 0001, Fan Rao, Huafeng Liu 0003, Yefeng Zheng 0001, Wentao Zhu 0002
MICCAI (3)3
2025 A Prior-Driven Lightweight Network for Endoscopic Exposure Correction
Zhijian Wu, Hong Wang 0021, Dingjiang Huang, Yefeng Zheng 0001
MICCAI (11)2
2025 A Multimodal Contrastive Learning for Detecting Aortic Dissection on 3D Non-contrast CT with Anatomy Simplification
Duoer Zhang, Yuxuan Qiu, Zhan Feng, Hong Wang 0021, Yefeng Zheng 0001, Wentao Zhu 0002
MICCAI (7)6
2025 SAMVSR: Leveraging Semantic Priors to Zone-Focused Mamba for Video Snow Removal
abstract
The outdoor vision systems are frequently degraded by snow particles, which obscure scene content and impair the performance of downstream vision tasks. While previous methods rely on physical priors, their performance often deteriorates under real-world conditions. Recently, semantic priors have proven effective in guiding image restoration, especially with the advent of the Segment Anything Model (SAM), which provides robust segmentation masks under adverse weather. However, leveraging SAM in video restoration remains underexplored due to the temporal inconsistency of inter-frame segmentation. In this work, we carefully construct the first framework to incorporate SAM-derived semantic priors into video snow removal, called SAMVSR. Specifically, to address temporal SAM label misalignment, we introduce an Entropy-wise Zone Propagation technique, which selects a reliable reference mask and semantically aligns instances across different frames via an entropy-guided label matching mechanism. Based on the aligned SAM semantic priors, we propose a Zone-Focused Mamba module, a novel Mamba-based architecture that restricts its scanning scope to semantically coherent zones, effectively mitigating irrelevant interactions and enhancing temporal-spatial consistency. Extensive experiments on both synthetic and real-world benchmarks finely validate the superiority of our proposed SAMVSR over existing state-of-the-art video desnowing techniques.
Jiaxuan Jiang 0001, Hong Wang 0021, Yefeng Zheng 0001
ACM Multimedia5
2025 Encoding sampling pattern for robust and generalized MRI reconstruction
Hong Wang 0021, Qi Xie 0002, Yefeng Zheng 0001, Deyu Meng
Pattern Recognit.2
2025 Unlocking the Potential of Weakly Labeled Data: A Co-Evolutionary Learning Framework for Abnormality Detection and Report Generation
abstract
Anatomical abnormality detection and report generation of chest X-ray (CXR) are two essential tasks in clinical practice. The former aims at localizing and characterizing cardiopulmonary radiological findings in CXRs, while the latter summarizes the findings in a detailed report for further diagnosis and treatment. Existing methods often focused on either task separately, ignoring their correlation. This work proposes a co-evolutionary abnormality detection and report generation (CoE-DG) framework. The framework utilizes both fully labeled (with bounding box annotations and clinical reports) and weakly labeled (with reports only) data to achieve mutual promotion between the abnormality detection and report generation tasks. Specifically, we introduce a bi-directional information interaction strategy with generator-guided information propagation (GIP) and detector-guided information propagation (DIP). For semi-supervised abnormality detection, GIP takes the informative feature extracted by the generator as an auxiliary input to the detector and uses the generator's prediction to refine the detector's pseudo labels. We further propose an intra-image-modal self-adaptive non-maximum suppression module (SA-NMS). This module dynamically rectifies pseudo detection labels generated by the teacher detection model with high-confidence predictions by the student. Inversely, for report generation, DIP takes the abnormalities' categories and locations predicted by the detector as input and guidance for the generator to improve the generated reports. Finally, a co-evolutionary training strategy is implemented to iteratively conduct GIP and DIP and consistently improve both tasks' performance. Experimental results on two public CXR datasets demonstrate CoE-DG's superior performance to several up-to-date object detection, report generation, and unified models. Our code is available at https://github.com/jinghanSunn/CoE-DG.
Jinghan Sun, Dong Wei 0004, Zhe Xu 0012, Donghuan Lu, Hong Wang 0021, Sotirios A. Tsaftaris, Steven McDonagh 0001, Yefeng Zheng 0001, Liansheng Wang 0002
IEEE Trans. Medical Imaging6
2025 Adaptive Weighting Based Metal Artifact Reduction in CT Images
abstract
Against the metal artifact reduction (MAR) task in computed tomography (CT) imaging, most of the existing deep-learning-based approaches generally select a single Hounsfield unit (HU) window followed by a normalization operation to preprocess CT images. However, in practical clinical scenarios, different body tissues and organs are often inspected under varying window settings for good contrast. The methods trained on a fixed single window would lead to insufficient removal of metal artifacts when being transferred to deal with other windows. To alleviate this problem, few works have proposed to reconstruct the CT images under multiple-window configurations. Albeit achieving good reconstruction performance for different windows, they adopt to directly supervise each window learning in an equal weighting way based on the training set. To improve the learning flexibility and model generalizability, in this paper, we propose an adaptive weighting algorithm, called AdaW, for the multiple-window metal artifact reduction, which can be applied to different deep MAR network backbones. Specifically, we first formulate the multiple window learning task as a bi-level optimization problem. Then we derive an adaptive weighting optimization algorithm where the learning process for MAR under each window is automatically weighted via a learning-to-learn paradigm based on the training set and validation set. This rationality is finely substantiated through theoretical analysis. Based on different network backbones, experimental comparisons executed on five datasets with different body sites comprehensively validate the effectiveness of AdaW in helping improve the generalization performance as well as its good applicability. We will release the code at https://github.com/hongwang01/AdaW.
Hong Wang 0021, Dong Wei 0004, Xian Wu 0001, Jianhua Ma 0001, Yefeng Zheng 0001
IEEE Trans. Medical Imaging1
2025 TRG-Net: An Interpretable and Controllable Rain Generator
abstract
Exploring and modeling the rain generation mechanism is critical for augmenting paired data to ease the training of rainy image processing models. Most of the conventional methods handle this task in an artificial physical rendering manner, through elaborately designing fundamental elements constituting rains. These kinds of methods, however, are over-dependent on human subjectivity, which limits their adaptability to real rains. In contrast, recent deep learning (DL) methods have achieved great success by training a neural network-based generator from pre-collected rainy image data. However, current methods usually design the generator in a "closed box" manner, increasing the learning difficulty and data requirements. To address these issues, this study proposes a novel DL-based rain generator, which fully takes the physical generation mechanism underlying rains into consideration and well encodes the learning of the fundamental rain factors (i.e., shape, orientation, length, width, and sparsity) explicitly into the deep network. Its significance lies in that the generator not only elaborately designs essential elements of the rain to simulate expected rains, like conventional artificial strategies, but also finely adapts to complicated and diverse practical rainy images, like DL methods. By rationally adopting the filter parameterization technique, the proposed rain generator is finely controllable with respect to rain factors and able to learn the distribution of these factors purely from data without the need for rain factor labels. Our unpaired generation experiments demonstrate that the rain generated by the proposed rain generator is not only of higher quality but also more effective for deraining and downstream tasks compared to current state-of-the-art rain generation methods. Besides, the paired data augmentation experiments, including both in-distribution and out-of-distribution (OOD), further validate the diversity of samples generated by our model for in-distribution deraining and OOD generalization tasks.
Zhiqiang Pang, Hong Wang 0021, Qi Xie 0002, Deyu Meng, Zongben Xu
IEEE Trans. Neural Networks Learn. Syst.2
2025 RSF-Conv: Rotation-and-Scale Equivariant Fourier Parameterized Convolution for Retinal Vessel Segmentation
abstract
Retinal vessel segmentation is of great clinical significance for the diagnosis of many eye-related diseases, but it is still a formidable challenge due to the intricate vascular morphology. With the skillful characterization of the translation symmetry existing in retinal vessels, convolutional neural networks (CNNs) have achieved great success in retinal vessel segmentation. However, the rotation-and-scale symmetry, as a more widespread image prior in retinal vessels, fails to be characterized by CNNs. Therefore, we propose a rotation-and-scale equivariant Fourier parameterized convolution (RSF-Conv) specifically for retinal vessel segmentation and provide the corresponding equivariance analysis. As a general module, RSF-Conv can be integrated into existing networks in a plug-and-play manner while significantly reducing the number of parameters. For instance, we replace the traditional convolution filters in U-Net, Iter-Net, DE-DCGCN-EE, and FR-UNet, with RSF-Convs, and faithfully conduct comprehensive experiments. RSF-Conv-enhanced methods not only have slight advantages under in-domain evaluation but also, more importantly, outperform all comparison methods by a significant margin under out-of-domain evaluation. It indicates that the remarkable generalization of RSF-Conv holds greater practical clinical significance for the prevalent cross-device and cross-hospital challenges in clinical practice. To comprehensively demonstrate the effectiveness of RSF-Conv, we also apply RSF-Conv + U-Net and RSF-Conv + Iter-Net to retinal artery/vein classification and achieve promising performance as well, indicating its clinical application potential. The code is available at https://github.com/szhc0gk/RSF-Conv.
Zihong Sun, Hong Wang 0021, Qi Xie 0002, Yefeng Zheng 0001, Deyu Meng
IEEE Trans. Neural Networks Learn. Syst.2
2024 Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto Optimization
abstract
Catastrophic forgetting remains a core challenge in continual learning (CL), where the models struggle to retain previous knowledge when learning new tasks. While existing replay-based CL methods have been proposed to tackle this challenge by utilizing a memory buffer to store data from previous tasks, they generally overlook the interdependence between previously learned tasks and fail to encapsulate the optimally integrated knowledge in previous tasks, leading to sub-optimal performance of the previous tasks. Against this issue, we first reformulate replay-based CL methods as a unified hierarchical gradient aggregation framework. We then incorporate the Pareto optimization to capture the interrelationship among previously learned tasks and design a Pareto-Optimized CL algorithm (POCL), which effectively enhances the overall performance of past tasks while ensuring the performance of the current task. Comprehensive empirical results demonstrate that the proposed POCL outperforms current state-of-the-art CL methods across multiple datasets and different settings.
Hong Wang 0021, Peilin Zhao, Yefeng Zheng 0001, Ying Wei 0001, Long-Kai Huang
ICML2
2024 Cross-Modal Vertical Federated Learning for MRI Reconstruction
abstract
Federated learning enables multiple hospitals to cooperatively learn a shared model without privacy disclosure. Existing methods often take a common assumption that the data from different hospitals have the same modalities. However, such a setting is difficult to fully satisfy in practical applications, since the imaging guidelines may be different between hospitals, which makes the number of individuals with the same set of modalities limited. To this end, we formulate this practical-yet-challenging cross-modal vertical federated learning task, in which data from multiple hospitals have different modalities with a small amount of multi-modality data collected from the same individuals. To tackle such a situation, we develop a novel framework, namely Federated Consistent Regularization constrained Feature Disentanglement (Fed-CRFD), for boosting MRI reconstruction by effectively exploring the overlapping samples (i.e., same patients with different modalities at different hospitals) and solving the domain shift problem caused by different modalities. Particularly, our Fed-CRFD involves an intra-client feature disentangle scheme to decouple data into modality-invariant and modality-specific features, where the modality-invariant features are leveraged to mitigate the domain shift problem. In addition, a cross-client latent representation consistency constraint is proposed specifically for the overlapping samples to further align the modality-invariant features extracted from different modalities. Hence, our method can fully exploit the multi-source data from hospitals while alleviating the domain shift problem. Extensive experiments on two typical MRI datasets demonstrate that our network clearly outperforms state-of-the-art MRI reconstruction methods.
Yunlu Yan, Hong Wang 0021, Yawen Huang, Nanjun He, Lei Zhu 0003, Yong Xu 0001, Yuexiang Li, Yefeng Zheng 0001
IEEE J. Biomed. Health Informatics2
2024 OSCNet: Orientation-Shared Convolutional Network for CT Metal Artifact Learning
abstract
X-ray computed tomography (CT) has been broadly adopted in clinical applications for disease diagnosis and image-guided interventions. However, metals within patients always cause unfavorable artifacts in the recovered CT images. Albeit attaining promising reconstruction results for this metal artifact reduction (MAR) task, most of the existing deep-learning-based approaches have some limitations. The critical issue is that most of these methods have not fully exploited the important prior knowledge underlying this specific MAR task. Therefore, in this paper, we carefully investigate the inherent characteristics of metal artifacts which present rotationally symmetrical streaking patterns. Then we specifically propose an orientation-shared convolution representation mechanism to adapt such physical prior structures and utilize Fourier-series-expansion-based filter parametrization for modelling artifacts, which can finely separate metal artifacts from body tissues. By adopting the classical proximal gradient algorithm to solve the model and then utilizing the deep unfolding technique, we easily build the corresponding orientation-shared convolutional network, termed as OSCNet. Furthermore, considering that different sizes and types of metals would lead to different artifact patterns (e.g., intensity of the artifacts), to better improve the flexibility of artifact learning and fully exploit the reconstructed results at iterative stages for information propagation, we design a simple-yet-effective sub-network for the dynamic convolution representation of artifacts. By easily integrating the sub-network into the proposed OSCNet framework, we further construct a more flexible network structure, called OSCNet+, which improves the generalization performance. Through extensive experiments conducted on synthetic and clinical datasets, we comprehensively substantiate the effectiveness of our proposed methods. Code will be released at https://github.com/hongwang01/OSCNet.
Hong Wang 0021, Qi Xie 0002, Dong Zeng, Jianhua Ma 0001, Deyu Meng, Yefeng Zheng 0001
IEEE Trans. Medical Imaging1
2024 Relational Experience Replay: Continual Learning by Adaptively Tuning Task-Wise Relationship
abstract
Continual learning is a promising machine learning paradigm to learn new tasks while retaining previously learned knowledge over streaming training data. Till now,rehearsal-basedmethods, keeping a small part of data from old tasks as a memory buffer, have shown good performance in mitigating catastrophic forgetting for previously learned knowledge. However, most of these methods typically treat each new task equally, which may not adequately consider the relationship or similarity between old and new tasks. Furthermore, these methods commonly neglect sample importance in the continual training process and result in sub-optimal performance on certain tasks. To address this challenging problem, we propose Relational Experience Replay (RER), a bi-level learning framework, to adaptively tune task-wise relationships and sample importance within each task to achieve a better ‘stability’ and ‘plasticity’ trade-off. As such, the proposed method is capable of accumulating new knowledge while consolidating previously learned old knowledge during continual learning. Extensive experiments conducted on three benchmark image datasets (CIFAR-10, CIFAR-100, and Tiny ImageNet) and two text datasets (20News and DBpedia) show that the proposed method can consistently improve the performance of all baselines and surpass current state-of-the-art methods.
Quanziang Wang, Renzhen Wang, Yuexiang Li, Dong Wei 0004, Hong Wang 0021, Kai Ma 0002, Yefeng Zheng 0001, Deyu Meng
IEEE Trans. Multim.5
2024 Low-Light Image Enhancement by Retinex-Based Algorithm Unrolling and Adjustment
abstract
Low-light image enhancement (LIE) has attracted tremendous research interests in recent years. Retinex theory-based deep learning methods, following a decomposition-adjustment pipeline, have achieved promising performance due to their physical interpretability. However, existing Retinex-based deep learning methods are still suboptimal, failing to leverage useful insights from traditional approaches. Meanwhile, the adjustment step is either oversimplified or overcomplicated, resulting in unsatisfactory performance in practice. To address these issues, we propose a novel deep-learning framework for LIE. The framework consists of a decomposition network (DecNet) inspired by algorithm unrolling and adjustment networks considering both global and local brightness. The algorithm unrolling allows the integration of both implicit priors learned from data and explicit priors inherited from traditional methods, facilitating better decomposition. Meanwhile, considering global and local brightness guides the design of effective yet lightweight adjustment networks. Moreover, we introduce a self-supervised fine-tuning strategy that achieves promising performance without manual hyperparameter tuning. Extensive experiments on benchmark LIE datasets demonstrate the superiority of our approach over existing state-of-the-art methods both quantitatively and qualitatively. Code is available at https://github.com/Xinyil256/RAUNA2023.
Qi Xie 0002, Qian Zhao 0002, Hong Wang 0021, Deyu Meng
IEEE Trans. Neural Networks Learn. Syst.4
2024 RCDNet: An Interpretable Rain Convolutional Dictionary Network for Single Image Deraining
abstract
As common weather, rain streaks adversely degrade the image quality and tend to negatively affect the performance of outdoor computer vision systems. Hence, removing rains from an image has become an important issue in the field. To handle such an ill-posed single image deraining task, in this article, we specifically build a novel deep architecture, called rain convolutional dictionary network (RCDNet), which embeds the intrinsic priors of rain streaks and has clear interpretability. In specific, we first establish a rain convolutional dictionary (RCD) model for representing rain streaks and utilize the proximal gradient descent technique to design an iterative algorithm only containing simple operators for solving the model. By unfolding it, we then build the RCDNet in which every network module has clear physical meanings and corresponds to each operation involved in the algorithm. This good interpretability greatly facilitates an easy visualization and analysis of what happens inside the network and why it works well in the inference process. Moreover, taking into account the domain gap issue in real scenarios, we further design a novel dynamic RCDNet, where the rain kernels can be dynamically inferred corresponding to input rainy images and then help shrink the space for rain layer estimation with few rain maps, so as to ensure a fine generalization performance in the inconsistent scenarios of rain types between training and testing data. By end-to-end training such an interpretable network, all involved rain kernels and proximal operators can be automatically extracted, faithfully characterizing the features of both rain and clean background layers and, thus, naturally leading to better deraining performance. Comprehensive experiments implemented on a series of representative synthetic and real datasets substantiate the superiority of our method, especially on its well generality to diverse testing scenarios and good interpretability for all its modules, compared with state-of-the-art single image derainers both visually and quantitatively. Code is available at https://github.com/hongwang01/DRCDNet.
Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Yuexiang Li, Yong Liang 0001, Yefeng Zheng 0001, Deyu Meng
IEEE Trans. Neural Networks Learn. Syst.1
2023 ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image Segmentation
abstract
Vision Transformers have recently shown impressive performances on medical image segmentation. Despite their strong capability of modeling long-range dependencies, the current methods still give rise to two main concerns in a class-level perspective: (1) intra-class problem: the existing methods lacked in extracting class-specific correspondences of different pixels, which may lead to poor object coverage and/or boundary prediction; (2) inter-class problem: the existing methods failed to model explicit category-dependencies among various objects, which may result in inaccurate localization. In light of these two issues, we propose a novel transformer, called ClassFormer, powered by two appealing transformers, i.e., intra-class dynamic transformer and inter-class interactive transformer, to address the challenge of fully exploration on compactness and discrepancy. Technically, the intra-class dynamic transformer is first designed to decouple representations of different categories with an adaptive selection mechanism for compact learning, which optimally highlights the informative features to reflect the salient keys/values from multiple scales. We further introduce the inter-class interactive transformer to capture the category dependency among different objects, and model class tokens as the representative class centers to guide a global semantic reasoning. As a consequence, the feature consistency is ensured with the expense of intra-class penalization, while inter-class constraint strengthens the feature discriminability between different categories. Extensive empirical evidence shows that ClassFormer can be easily plugged into any architecture, and yields improvements over the state-of-the-art methods in three public benchmarks.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hong Wang 0021, Yuexiang Li, Yawen Huang, Yefeng Zheng 0001
AAAI6
2023 SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic Segmentation
abstract
Semi-supervised learning improves data efficiency of deep models by leveraging unlabeled samples to alleviate the reliance on a large set of labeled samples. These successes concentrate on the pixel-wise consistency by using convolutional neural networks (CNNs) but fail to address both global learning capability and class-level features for unlabeled data. Recent works raise a new trend that Transformer achieves superior performance on the entire feature map in various tasks. In this paper, we unify the current dominant Mean-Teacher approaches by reconciling intra-model and inter-model properties for semi-supervised segmentation to produce a novel algorithm, SemiCVT, that absorbs the quintessence of CNNs and Transformer in a comprehensive way. Specifically, we first design a parallel CNN-Transformer architecture (CVT) with introducing an intra-model local-global interaction schema (LGI) in Fourier domain for full integration. The inter-model class-wise consistency is further presented to complement the class-level statistics of CNNs and Transformer in a cross-teaching manner. Extensive empirical evidence shows that SemiCVT yields consistent improvements over the state-of-the-art methods in two public benchmarks.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Hong Wang 0021, Yawen Huang, Yefeng Zheng 0001
CVPR7
2023 Interactive Segmentation as Gaussian Process Classification
abstract
Click-based interactive segmentation (IS) aims to extract the target objects under user interaction. For this task, most of the current deep learning (DL)-based methods mainly follow the general pipelines of semantic segmentation. Albeit achieving promising performance, they do not fully and explicitly utilize and propagate the click information, inevitably leading to unsatisfactory segmentation results, even at clicked points. Against this issue, in this paper, we propose to formulate the IS task as a Gaussian process (GP)-based pixel-wise binary classification model on each image. To solve this model, we utilize amortized variational inference to approximate the intractable GP posterior in a data-driven manner and then decouple the approximated GP posterior into double space forms for efficient sampling with linear complexity. Then, we correspondingly construct a GP classification framework, named GPCIS, which is integrated with the deep kernel learning mechanism for more flexibility. The main specificities of the proposed GPCIS lie in: 1) Under the explicit guidance of the derived GP posterior, the information contained in clicks can be finely propagated to the entire image and then boost the segmentation; 2) The accuracy of predictions at clicks has good theoretical support. These merits of GPCIS as well as its good generality and high efficiency are substantiated by comprehensive experiments on several benchmarks, as compared with representative methods both quantitatively and qualitatively. Codes will be released at https://github.com/zmhhlnz/GPCIS_CVPR2023.
Hong Wang 0021, Qian Zhao 0002, Yuexiang Li, Yawen Huang, Deyu Meng, Yefeng Zheng 0001
CVPR2
2023 MEPNet: A Model-Driven Equivariant Proximal Network for Joint Sparse-View Reconstruction and Metal Artifact Reduction in CT Images
Hong Wang 0021, Dong Wei 0004, Yuexiang Li, Yefeng Zheng 0001
MICCAI (10)1
2023 A deep weakly semi-supervised framework for endoscopic lesion segmentation
Hong Wang 0021, Haoqin Ji, Yuexiang Li, Nanjun He, Dong Wei 0004, Yawen Huang, Xinrong Chen, Yefeng Zheng 0001, Hongmeng Yu
Medical Image Anal.2
2023 InDuDoNet+: A deep unfolding dual domain network for metal artifact reduction in CT images
Hong Wang 0021, Yuexiang Li, Haimiao Zhang, Deyu Meng, Yefeng Zheng 0001
Medical Image Anal.1
2022 KXNet: A Model-Driven Deep Neural Network for Blind Super-Resolution
Jiahong Fu, Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu
ECCV (19)2
2022 Adaptive Convolutional Dictionary Network for CT Metal Artifact Reduction
abstract
Inspired by the great success of deep neural networks, learning-based methods have gained promising performances for metal artifact reduction (MAR) in computed tomography (CT) images. However, most of the existing approaches put less emphasis on modelling and embedding the intrinsic prior knowledge underlying this specific MAR task into their network designs. Against this issue, we propose an adaptive convolutional dictionary network (ACDNet), which leverages both model-based and learning-based methods. Specifically, we explore the prior structures of metal artifacts, e.g., non-local repetitive streaking patterns, and encode them as an explicit weighted convolutional dictionary model. Then, a simple-yet-effective algorithm is carefully designed to solve the model. By unfolding every iterative substep of the proposed algorithm into a network module, we explicitly embed the prior structure into a deep network , i.e., a clear interpretability for the MAR task. Furthermore, our ACDNet can automatically learn the prior for artifact-free CT images via training data and adaptively adjust the representation kernels for each input CT image based on its content. Hence, our method inherits the clear interpretability of model-based methods and maintains the powerful representation ability of learning-based methods. Comprehensive experiments executed on synthetic and clinical datasets show the superiority of our ACDNet in terms of effectiveness and model generalization. Code and supplementary material are available at https://github.com/hongwang01/ACDNet.
Hong Wang 0021, Yuexiang Li, Deyu Meng, Yefeng Zheng 0001
IJCAI1
2022 Orientation-Shared Convolution Representation for CT Metal Artifact Learning
Hong Wang 0021, Qi Xie 0002, Yuexiang Li, Yawen Huang, Deyu Meng, Yefeng Zheng 0001
MICCAI (6)1
2022 Survey on rain removal from videos or a single image
Hong Wang 0021, Minghan Li 0001, Qian Zhao 0002, Deyu Meng
Sci. China Inf. Sci.1
2022 DICDNet: Deep Interpretable Convolutional Dictionary Network for Metal Artifact Reduction in CT Images
abstract
Computed tomography (CT) images are often impaired by unfavorable artifacts caused by metallic implants within patients, which would adversely affect the subsequent clinical diagnosis and treatment. Although the existing deep-learning-based approaches have achieved promising success on metal artifact reduction (MAR) for CT images, most of them treated the task as a general image restoration problem and utilized off-the-shelf network modules for image quality enhancement. Hence, such frameworks always suffer from lack of sufficient model interpretability for the specific task. Besides, the existing MAR techniques largely neglect the intrinsic prior knowledge underlying metal-corrupted CT images which is beneficial for the MAR performance improvement. In this paper, we specifically propose a deep interpretable convolutional dictionary network (DICDNet) for the MAR task. Particularly, we first explore that the metal artifacts always present non-local streaking and star-shape patterns in CT images. Based on such observations, a convolutional dictionary model is deployed to encode the metal artifacts. To solve the model, we propose a novel optimization algorithm based on the proximal gradient technique. With only simple operators, the iterative steps of the proposed algorithm can be easily unfolded into corresponding network modules with specific physical meanings. Comprehensive experiments on synthesized and clinical datasets substantiate the effectiveness of the proposed DICDNet as well as its superior interpretability, compared to current state-of-the-art MAR methods. Code is available at https://github.com/hongwang01/DICDNet.
Hong Wang 0021, Yuexiang Li, Nanjun He, Kai Ma 0002, Deyu Meng, Yefeng Zheng 0001
IEEE Trans. Medical Imaging1
2021 From Rain Generation to Rain Removal
abstract
For the single image rain removal (SIRR) task, the performance of deep learning (DL)-based methods is mainly affected by the designed deraining models and training datasets. Most of current state-of-the-art focus on constructing powerful deep models to obtain better deraining results. In this paper, to further improve the deraining performance, we novelly attempt to handle the SIRR task from the perspective of training datasets by exploring a more efficient way to synthesize rainy images. Specifically, we build a full Bayesian generative model for rainy image where the rain layer is parameterized as a generator with the input as some latent variables representing the physical structural rain factors, e.g., direction, scale, and thickness. To solve this model, we employ the variational inference framework to approximate the expected statistical distribution of rainy image in a data-driven manner. With the learned generator, we can automatically and sufficiently generate diverse and non-repetitive training pairs so as to efficiently enrich and augment the existing benchmark datasets. User study qualitatively and quantitatively evaluates the realism of generated rainy images. Comprehensive experiments substantiate that the proposed model can faithfully extract the complex rain distribution that not only helps significantly improve the deraining performance of current deep single image derainers, but also largely loosens the requirement of large training sample pre-collection for the SIRR task. Code is available in https://github.com/hongwang01/VRGNet.
Hong Wang 0021, Zongsheng Yue, Qi Xie 0002, Qian Zhao 0002, Yefeng Zheng 0001, Deyu Meng
CVPR1
2021 InDuDoNet: An Interpretable Dual Domain Network for CT Metal Artifact Reduction
Hong Wang 0021, Yuexiang Li, Haimiao Zhang, Jiawei Chen 0009, Kai Ma 0002, Deyu Meng, Yefeng Zheng 0001
MICCAI (6)1
2021 Selective generative adversarial network for raindrop removal from a single image
Ming-Wen Shao, Hong Wang 0021, Deyu Meng
Neurocomputing3
2021 Structural residual learning for single image rain removal
Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Yong Liang 0001, Deyu Meng
Knowl. Based Syst.1
2020 A Model-Driven Deep Neural Network for Single Image Rain Removal
abstract
Deep learning (DL) methods have achieved state-of-the-art performance in the task of single image rain removal. Most of current DL architectures, however, are still lack of sufficient interpretability and not fully integrated with physical structures inside general rain streaks. To this issue, in this paper, we propose a model-driven deep neural network for the task, with fully interpretable network structures. Specifically, based on the convolutional dictionary learning mechanism for representing rain, we propose a novel single image deraining model and utilize the proximal gradient descent technique to design an iterative algorithm only containing simple operators for solving the model. Such a simple implementation scheme facilitates us to unfold it into a new deep network architecture, called rain convolutional dictionary network (RCDNet), with almost every network module one-to-one corresponding to each operation involved in the algorithm. By end-to-end training the proposed RCDNet, all the rain kernels and proximal operators can be automatically extracted, faithfully characterizing the features of both rain and clean background layers, and thus naturally lead to its better deraining performance, especially in real scenarios. Comprehensive experiments substantiate the superiority of the proposed network, especially its well generality to diverse testing scenarios and good interpretability for all its modules, as compared with state-of-the-arts both visually and quantitatively.
Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Deyu Meng
CVPR1