VLDB 2026 Research / reviewers in the wild / expert
Tao Wang 0052
dblp:12/5838-52
· DBLP profile ↗
36ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0002-0202-0174ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 23 since 2021Artificial intelligence and machine learning · 19 · 7 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable and Generalizable Correspondence Pruning via Geometry-Consistent Pre-TrainingabstractTwo-view correspondence pruning aims to identify reliable correspondences for camera pose estimation, serving as a fundamental step in many 3D vision tasks. Existing methods rely on geometric consistency to seek true correspondences (inliers) from numerous false correspondences (outliers). In this learning paradigm, outliers severely affect the representation learning of inliers, resulting in models that are neither robust nor generalizable. To address this issue, we propose a geometry-consistent pre-training paradigm that sculpts scalable and generalizable representations free from outlier interference. The paradigm features two appealing properties. 1) Implementation of geometry-consistent pre-training. We introduce masked inlier reconstruction as a pretext task and develop a simple yet effective pre-training framework based on a masked autoencoder. Specifically, due to the irregular and unordered nature of correspondences, which lack explicit positional information, we adopt a dual-branch structure that separately reconstructs the keypoints of two images. This enables indirect reconstruction of 4D correspondences, where keypoints from the paired image provide positional prompts. 2) Unified correspondence encoder. We propose a simple dual-stream encoder with built-in consensus interaction, providing a unified, extensible architecture that enhances representation learning. Extensive experiments demonstrate that our method, GeneralPruner, consistently outperforms state-of-the-art approaches in terms of robustness and generalization across various downstream tasks. Specifically, our method achieves 10.76%, 11.84%, and 8.65% performance gains in camera pose estimation, visual localization, and 3D registration, respectively. To the best of our knowledge, we are the first work to introduce a pre-training framework tailored for correspondence pruning, offering a more universal and scalable solution. Tangfei Liao, Xiaoqin Zhang 0002, Tao Wang 0052, Min Li 0052, Guobao Xiao, Mang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | CorrMoE: Mixture of Experts with De-Stylization Learning for Cross-Scene and Cross-Domain Correspondence PruningabstractEstablishing reliable correspondences between image pairs is a fundamental task in computer vision, underpinning applications such as 3D reconstruction and visual localization. Although recent methods have made progress in pruning outliers from dense correspondence sets, they often hypothesize consistent visual domains and overlook the challenges posed by diverse scene structures. In this paper, we propose CorrMoE, a novel correspondence pruning framework that enhances robustness under cross-domain and cross-scene variations. To address domain shift, we introduce a De-stylization Dual Branch, performing style mixing on both implicit and explicit graph features to mitigate the adverse influence of domain-specific representations. For scene diversity, we design a Bi-Fusion Mixture of Experts module that adaptively integrates multi-perspective features through linear-complexity attention and dynamic expert routing. Extensive experiments on benchmark datasets demonstrate that CorrMoE achieves superior accuracy and generalization compared to state-of-the-art methods. The code and pre-trained models are available at https://github.com/peiwenxia/CorrMoE. Peiwen Xia, Tangfei Liao, Danhuai Zhao, Jianjun Ke, Kaihao Zhang, Tong Lu 0002, Tao Wang 0052 |
ECAI | 8 |
| 2025 | LLFA: Fusing Global Illumination and Local Priors for Low-Light Face Image Enhancement with AdaptorabstractLow-light image enhancement problem has been widely studied. However, most existing methods do not perform well on low-light face images due to no specific facial characteristic considerations. We first create large-scale low-light face datasets with synthesized and real-world images to address the absence of suitable datasets. Our experiments show that existing LLIE and face restoration methods are limited in enhancing low-light face images. To overcome these challenges, we propose a novel framework, the Low-Light Face Adaptor (LLFA), featuring an auxiliary encoder and an adaptor module. The encoder captures global illumination information, while the adaptor module adaptively fuses this information with high-quality priors. We also introduce a joint learning strategy that optimizes the model by simultaneously learning face priors and the enhancement process. Comprehensive experiments demonstrate that LLFA significantly outperforms state-of-the-art methods. Ziqian Shao, Tao Wang 0052, Kaihao Zhang, Danhuai Zhao, Tong Lu 0002 |
ICASSP | 2 |
| 2025 | Segmentation-Guided Sparse Transformer for Under-Display Camera Image RestorationabstractUnder-display Camera is an emerging technology for full-screen display with a camera under the display. However, the current implementation of UDC causes serious image degradation. Incident light required for camera imaging undergoes attenuation and diffraction when passing through the display. Current UDC image restoration methods predominantly utilize convolutional networks, whereas transformer-based methods with superior performance are lacking. This paper proposes a Segmentation-Guided Sparse Transformer method (SGSFormer) for restoring images from UDC degraded images. Specifically, we utilize sparse self-attention to filter out redundant information and noise, directing the model’s attention to focus on the features more relevant to the degraded regions in need of reconstruction. Moreover, we integrate an instance segmentation map as prior information to guide sparse self-attention in filtering and focusing on the correct regions. Extensive experiments exhibit the superior performance of our model over the state-of-the-art methods. Jingyun Xue, Tao Wang 0052, Pengwen Dai, Kaihao Zhang |
ICASSP | 2 |
| 2025 | MaterialMVP: Illumination-Invariant Material Generation via Multi-View PBR Diffusion
Zebin He, Mingxin Yang, Tao Wang 0052, Kaihao Zhang, Guanying Chen, Jie Jiang 0015, Chunchao Guo, Wenhan Luo |
ICCV | 5 |
| 2025 | MOERL: When Mixture-Of-Experts Meet Reinforcement Learning for Adverse Weather Image Restoration
Tao Wang 0052, Peiwen Xia, Peng-Tao Jiang, Zhe Kong, Kaihao Zhang, Tong Lu 0002, Wenhan Luo |
ICCV | 1 |
| 2025 | GRIG: Data-Efficient Generative Residual Image InpaintingabstractImage inpainting is the task of filling in missing or masked regions of an image with semantically meaningful content. Recent methods have shown significant improvement in dealing with large missing regions. However, these methods usually require large training datasets to achieve satisfactory results, and there has been limited research into training such models on a small number of samples. To address this, we present a novel data-efficient generative residual image inpainting method that produces high-quality inpainting results. The core idea is to use an iterative residual reasoning method that incorporates convolutional neural networks (CNNs) for feature extraction and transformers for global reasoning within generative adversarial networks, along with image-level and patch-level discriminators. We also propose a novel forged-patch adversarial training strategy to create faithful textures and detailed appearances. Extensive evaluation shows that our method outperforms previous methods on the data-efficient image inpainting task, both quantitatively and qualitatively. Wanglong Lu, Xianta Jiang, Xiaogang Jin 0001, Minglun Gong, Kaijie Shi 0002, Tao Wang 0052, Hanli Zhao |
Comput. Vis. Media | 7 |
| 2025 | Visual style prompt learning using diffusion models for blind face restoration
Wanglong Lu, Tao Wang 0052, Kaihao Zhang, Xianta Jiang, Hanli Zhao |
Pattern Recognit. | 3 |
| 2025 | LLDiffusion: Learning degradation representations in diffusion models for low-light image enhancement
Tao Wang 0052, Kaihao Zhang, Yong Zhang 0034, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005 |
Pattern Recognit. | 1 |
| 2024 | CorrAdaptor: Adaptive Local Context Learning for Correspondence PruningabstractIn the fields of computer vision and robotics, accurate pixel-level correspondences are essential for enabling advanced tasks such as structure-from-motion and simultaneous localization and mapping. Recent correspondence pruning methods usually focus on learning local consistency through k-nearest neighbors, which makes it difficult to capture robust context for each correspondence. We propose CorrAdaptor, a novel architecture that introduces a dual-branch structure capable of adaptively adjusting local contexts through both explicit and implicit local graph learning. Specifically, the explicit branch uses KNN-based graphs tailored for initial neighborhood identification, while the implicit branch leverages a learnable matrix to softly assign neighbors and adaptively expand the local context scope, significantly enhancing the model’s robustness and adaptability to complex image variations. Moreover, we design a motion injection module to integrate motion consistency into the network to suppress the impact of outliers and refine local context learning, resulting in substantial performance improvements. The experimental results on extensive correspondence-based tasks indicate that our CorrAdaptor achieves state-of-the-art performance both qualitatively and quantitatively. Yuping He, Tangfei Liao, Xiaoqiu Xu, Tao Wang 0052, Tong Lu 0002 |
ECAI | 7 |
| 2024 | OMG: Occlusion-Friendly Personalized Multi-concept Generation in Diffusion Models
Zhe Kong, Yong Zhang 0034, Tianyu Yang 0003, Tao Wang 0052, Kaihao Zhang, Bizhu Wu, Guanying Chen, Wei Liu 0005, Wenhan Luo |
ECCV (31) | 4 |
| 2024 | Blind Face Video Restoration with Temporal Consistent Generative Prior and Degradation-Aware PromptabstractWithin the domain of blind face restoration (BFR), approaches lacking facial priors frequently result in excessively smoothed visual outputs. Exiting BFR methods predominantly utilize generative facial priors to achieve realistic and authentic details. However, these methods, primarily designed for images, encounter challenges in maintaining temporal consistency when applied to face video restoration. To tackle this issue, we introduce StableBFVR, an innovative Blind Face Video Restoration method based on Stable Diffusion that incorporates temporal information into the generative prior. This is achieved through the introduction of temporal layers in the diffusion process. These temporal layers consider both long-term and short-term information aggregation. Moreover, to improve generalizability, BFR methods employ complex, large-scale degradation during training, but it often sacrifices accuracy. Addressing this, StableBFVR features a novel mixed-degradation-aware prompt module, capable of encoding specific degradation information to dynamically steer the restoration process. Comprehensive experiments demonstrate that our proposed StableBFVR outperforms state-of-the-art methods. Jingfan Tan, Hyunhee Park, Tao Wang 0052, Kaihao Zhang, Pengwen Dai, Zikun Liu 0001, Wenhan Luo |
ACM Multimedia | 4 |
| 2024 | GridFormer: Residual Dense Transformer with Grid Structure for Image Restoration in Adverse Weather Conditions
Tao Wang 0052, Kaihao Zhang, Ziqian Shao, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005, Hongdong Li |
Int. J. Comput. Vis. | 1 |
| 2024 | Restoring vision in hazy weather with hierarchical contrastive learning
Tao Wang 0052, Guangpin Tao, Wanglong Lu, Kaihao Zhang, Wenhan Luo, Xiaoqin Zhang 0002, Tong Lu 0002 |
Pattern Recognit. | 1 |
| 2024 | Toward Real-World Blind Face Restoration With Generative Diffusion PriorabstractBlind face restoration is an important task in computer vision and has gained significant attention due to its wide-range applications. Previous works mainly exploit facial priors to restore face images and have demonstrated high-quality results. However, generating faithful facial details remains a challenging problem due to the limited prior knowledge obtained from finite data. In this work, we delve into the potential of leveraging the pretrained Stable Diffusion for blind face restoration. We propose BFRffusion which is thoughtfully designed to effectively extract features from low-quality face images and could restore realistic and faithful facial details with the generative prior of the pretrained Stable Diffusion. In addition, we build a privacy-preserving face dataset called PFHQ with balanced attributes like race, gender, and age. This dataset can serve as a viable alternative for training blind face restoration networks, effectively addressing privacy and bias concerns usually associated with the real face datasets. Through an extensive series of experiments, we demonstrate that our BFRffusion achieves state-of-the-art performance on both synthetic and real-world public testing datasets for blind face restoration and our PFHQ dataset is an available resource for training blind face restoration networks. The codes, pretrained models, and dataset are released at https://github.com/chenxx89/BFRffusion. Jingfan Tan, Tao Wang 0052, Kaihao Zhang, Wenhan Luo, Xiaochun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Dual Teacher Knowledge Distillation With Domain Alignment for Face Anti-SpoofingabstractFace recognition systems have raised concerns due to their vulnerability to different presentation attacks, and system security has become an increasingly critical concern. Although many face anti-spoofing (FAS) methods perform well in intra-dataset scenarios, their generalization remains a challenge. To address this issue, some methods adopt domain adversarial training (DAT) to extract domain-invariant features. Differently, in this paper, we propose a domain adversarial attack (DAA) method by adding perturbations to the input images, which makes them indistinguishable across domains and enables domain alignment. Moreover, since models trained on limited data and types of attacks cannot generalize well to unknown attacks, we propose a dual perceptual and generative knowledge distillation framework for face anti-spoofing that utilizes pre-trained face-related models containing rich face priors. Specifically, we adopt two different face-related models as teachers to transfer knowledge to the target student model. The pre-trained teacher models are not from the task of face anti-spoofing but from perceptual and generative tasks, respectively, which implicitly augment the data. By combining both DAA and dual-teacher knowledge distillation, we develop a dual teacher knowledge distillation with domain alignment framework (DTDA) for face anti-spoofing. The advantage of our proposed method has been verified through extensive ablation studies and comparison with state-of-the-art methods on public datasets across multiple protocols. Zhe Kong, Wentian Zhang, Tao Wang 0052, Kaihao Zhang, Yuexiang Li, Xiaoying Tang 0001, Wenhan Luo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Blind Face Restoration for Under-Display Camera via Dictionary Guided TransformerabstractBy hiding the front-facing camera below the display panel, Under-Display Camera (UDC) provides users with a full-screen experience. However, due to the characteristics of the display, images taken by UDC suffer from significant quality degradation. Methods have been proposed to tackle UDC image restoration and advances have been achieved. There are still no specialized methods and datasets for restoring UDC face images, which may be the most common problem in the UDC scene. To this end, considering color filtering, brightness attenuation, and diffraction in the imaging process of UDC, we propose a two-stage network UDC Degradation Model Network named UDC-DMNet to synthesize UDC images by modeling the processes of UDC imaging. Then we use UDC-DMNet and high-quality face images from FFHQ and CelebA-Test to create UDC face training datasets FFHQ-P/T and testing datasets CelebA-Test-P/T for UDC face restoration. We propose a novel dictionary-guided transformer network named DGFormer. Introducing the facial component dictionary and the characteristics of the UDC image in the restoration makes DGFormer capable of addressing blind face restoration in UDC scenarios. Experiments show that our DGFormer and UDC-DMNet achieve state-of-the-art performance. Jingfan Tan, Tao Wang 0052, Kaihao Zhang, Wenhan Luo, Xiaochun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | MC-Blur: A Comprehensive Benchmark for Image DeblurringabstractBlur artifacts can seriously degrade the visual quality of images, and numerous deblurring methods have been proposed for specific scenarios. However, in most real-world images, blur is caused by different factors, e.g., motion, and defocus. In this paper, we address how other deblurring methods perform in the case of multiple types of blur. For in-depth performance evaluation, we construct a new large-scale multi-cause image deblurring dataset (MC-Blur), including real-world and synthesized blurry images with different blur factors. The images in the proposed MC-Blur dataset are collected using other techniques: averaging sharp images captured by a 1000-fps high-speed camera, convolving Ultra-High-Definition (UHD) sharp images with large-size kernels, adding defocus to images, and real-world blurry images captured by various camera models. Based on the MC-Blur dataset, we conduct extensive benchmarking studies to compare SOTA methods in different scenarios, analyze their efficiency, and investigate the buildataset’s capacity. These benchmarking results provide a comprehensive overview of the advantages and limitations of current deblurring methods, revealing our dataset’s advances. The dataset is available to the public athttps://github.com/HDCVLab/MC-Blur-Dataset. Kaihao Zhang, Tao Wang 0052, Wenhan Luo, Wenqi Ren, Björn Stenger, Wei Liu 0005, Hongdong Li, Ming-Hsuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Multi-Prior Driven Network for RGB-D Salient Object DetectionabstractMost existing RGB-D salient object detection (SOD) methods rely on high-quality depth images. However, their performance is limited when processing low-quality depth maps. This paper exploits more complementary image priors to guide the model to learn on variable depth maps, and a novel multi-prior driven network called MPDNet is proposed for RGB-D SOD. MPDNet utilizes four processing pipelines to process RGB images and other priors, which include an RGB image processing pipeline, a depth map processing pipeline, a fine-grained and gradient prior processing pipeline, and an edge learning pipeline. Specifically, fine-grained and gradient priors are input to the same processing pipeline. For the depth maps, fine-grained and gradient priors, a prior channel attention module utilizes the channel attention mechanism to filter noises and highlights the salient cues. The RGB image processing pipeline uses a multi-feature progressive enhancement module to fuse and enhance features from depth maps. And a multi-feature prediction decoder decodes initial salient masks. In the edge learning pipeline, edge prior serves as an edge label and is captured by an edge capture module. Finally, the clear salient masks are obtained by fusing the salient information from the four pipelines. The experimental results on six benchmarks indicate that the proposed method outperforms thirteen state-of-the-art methods in six evaluation metrics. Xiaoqin Zhang 0002, Yuewang Xu, Tao Wang 0052, Tangfei Liao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Ultra-High-Definition Low-Light Image Enhancement: A Benchmark and Transformer-Based MethodabstractAs the quality of optical sensors improves, there is a need for processing large-scale images. In particular, the ability of devices to capture ultra-high definition (UHD) images and video places new demands on the image processing pipeline. In this paper, we consider the task of low-light image enhancement (LLIE) and introduce a large-scale database consisting of images at 4K and 8K resolution. We conduct systematic benchmarking studies and provide a comparison of current LLIE algorithms. As a second contribution, we introduce LLFormer, a transformer-based low-light enhancement method. The core components of LLFormer are the axis-based multi-head self-attention and cross-layer attention fusion block, which significantly reduces the linear complexity. Extensive experiments on the new dataset and existing public datasets show that LLFormer outperforms state-of-the-art methods. We also show that employing existing LLIE methods trained on our benchmark as a pre-processing step significantly improves the performance of downstream tasks, e.g., face detection in low-light conditions. The source code and pre-trained models are available at https://github.com/TaoWangzj/LLFormer. Tao Wang 0052, Kaihao Zhang, Tianrun Shen, Wenhan Luo, Björn Stenger, Tong Lu 0002 |
AAAI | 1 |
| 2023 | Graph Propagation Transformer for Graph Representation LearningabstractThis paper presents a novel transformer architecture for graph representation learning. The core insight of our method is to fully consider the information propagation among nodes and edges in a graph when building the attention module in the transformer blocks. Specifically, we propose a new attention mechanism called Graph Propagation Attention (GPA). It explicitly passes the information among nodes and edges in three ways, i.e. node-to-node, node-to-edge, and edge-to-node, which is essential for learning graph-structured data. On this basis, we design an effective transformer architecture named Graph Propagation Transformer (GPTrans) to further help learn graph data. We verify the performance of GPTrans in a wide range of graph learning experiments on several benchmark datasets. These results show that our method outperforms many state-of-the-art transformer-based graph models with better performance. The code will be released at https://github.com/czczup/GPTrans. Zhe Chen 0017, Tao Wang 0052, Tianrun Shen, Tong Lu 0002, Qiuying Peng |
IJCAI | 3 |
| 2023 | Punctuation-level Attack: Single-shot and Single Punctuation Can Fool Text ModelsabstractThe adversarial attacks have attracted increasing attention in various fields including natural language processing. The current textual attacking models primarily focus on fooling models by adding character-/word-/sentence-level perturbations, ignoring their influence on human perception. In this paper, for the first time in the community, we propose a novel mode of textual attack, punctuation-level attack. With various types of perturbations, including insertion, displacement, deletion, and replacement, the punctuation-level attack achieves promising fooling rates against SOTA models on typical textual tasks and maintains minimal influence on human perception and understanding of the text by mere perturbation of single-shot single punctuation. Furthermore, we propose a search method named Text Position Punctuation Embedding and Paraphrase (TPPEP) to accelerate the pursuit of optimal position to deploy the attack, without exhaustive search, and we present a mathematical interpretation of TPPEP. Thanks to the integrated Text Position Punctuation Embedding (TPPE), the punctuation attack can be applied at a constant cost of time. Experimental results on public datasets and SOTA models demonstrate the effectiveness of the punctuation attack and the proposed TPPE. We additionally apply the single punctuation attack to summarization, semantic-similarity-scoring, and text-to-image tasks, and achieve encouraging results. Chongyang Du, Tao Wang 0052, Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wei Liu 0005, Xiaochun Cao |
NeurIPS | 3 |
| 2022 | Uncertainty-Based Network for Few-Shot Image ClassificationabstractThe transductive inference is an effective technique in the few-shot learning task, where query sets update prototypes to improve themselves. However, these methods optimize the model by considering only the classification scores of the query instances as confidence while ignoring the uncertainty of these classification scores. In this paper, we propose a novel method called Uncertainty-Based Network, which models the uncertainty of classification results with the help of mutual information. Specifically, we first data augment and classify the query instance and calculate the mutual information of these classification scores. Then, mutual information is used as uncertainty to assign weights to classification scores, and the iterative update strategy based on classification scores and uncertainties assigns the optimal weights to query instances in prototype optimization. Extensive results on four benchmarks show that Uncertainty-Based Network achieves comparable performance in classification accuracy compared to state-of-the-art methods. Minglei Yuan, Chunhao Cai, Yin-Dong Zheng, Tao Wang 0052, Tong Lu 0002, Wenbin Li 0006 |
ICME | 5 |
| 2022 | Multi-Attention Convolutional Neural Network for Video DeblurringabstractVideo deblurring, which aims at restoring the sharp video from blurry video, is drawing increasing attention in the field of computer vision. In this paper, a method called Multi-Attention Convolutional Neural Network (MACNN) consisting of the temporal-spatial attention module, the frame channel attention module, and the feature extraction-reconstruction module is proposed. First, we use the temporal-spatial attention module and the frame channel attention module to capture features with temporal and spatial information existing across neighboring frames. Then, these captured features are fused and reconstructed to restore the sharp frame. Last but not least, we train MACNN together with a content loss and a perceptual loss in an end-to-end manner to recover realistic video details. Both quantitative and qualitative evaluation results on standard benchmarks demonstrate the proposed MACNN is superior to the state-of-the-art methods in terms of accuracy, efficiency, and visual effect. Xiaoqin Zhang 0002, Tao Wang 0052, Runhua Jiang, Li Zhao 0005, Yuewang Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Single Image Haze Removal Based on a Simple Additive Model With Haze Smoothness PriorabstractSingle image haze removal, which is to recover the clear version of a hazy image, is a challenging task in computer vision. In this paper, an additive haze model is proposed to approximate the hazy image formation process. In contrast with the traditional optical model, it regards the haze as an additive layer to a clean image. The model thus avoids estimating the medium transmission rate and the global atmospherical light. In addition, based on a critical observation that haze changes gradually and smoothly across the image, a haze smoothness prior is proposed to constrain this model. This prior assumes that the haze layer is much smoother than the clear image. Benefiting from this prior, we can directly separate the clean image from a single hazy image. Experimental results and comparisons with synthetic images and real-world images demonstrate that the proposed method outperforms state-of-the-art single image haze removal algorithms. Xiaoqin Zhang 0002, Tao Wang 0052, Guiying Tang, Li Zhao 0005, Yuewang Xu, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Hierarchical Feature Fusion With Mixed Convolution Attention for Single Image DehazingabstractSingle image dehazing, which aims at restoring a haze-free image from its correspondingly unconstrained hazy scene, is a fundamental yet challenging task and has gained immense popularity recently. However, the images recovered by some existing haze-removal methods often contain haze, artifacts, and color distortions, which severely degrade the visual quality and have negative impacts on subsequent computer vision tasks. To this end, we propose a network combining multi-scale hierarchical feature fusion and mixed convolution attention to progressively and adaptively enhance the dehazing performance. The haze levels and image structure information are accurately estimated by fusing multi-scale hierarchical features, thus the model restores images with less remaining haze. The proposed mixed convolution attention mechanism is capable of reducing feature redundancy, learning compact and effective internal representations and highlighting task-relevant features, thus, it can further help the model estimate images with sharper textural details and more vivid colors. Furthermore, a deep semantic loss is also proposed to highlight essential semantic information in deep features. The experimental results show that the proposed method outperforms state-of-the-art haze removal algorithms. Xiaoqin Zhang 0002, Tao Wang 0052, Runhua Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Video Deblurring via Spatiotemporal Pyramid Network and Adversarial Gradient Prior
Tao Wang 0052, Xiaoqin Zhang 0002, Runhua Jiang, Li Zhao 0005, Huiling Chen 0001, Wenhan Luo |
Comput. Vis. Image Underst. | 1 |
| 2021 | Self-filtering image dehazing with self-supporting module
Pengcheng Huang 0002, Li Zhao 0005, Runhua Jiang, Tao Wang 0052, Xiaoqin Zhang 0002 |
Neurocomputing | 4 |
| 2021 | Haze concentration adaptive network for image dehazing
Tao Wang 0052, Li Zhao 0005, Pengcheng Huang 0002, Xiaoqin Zhang 0002, Jiawei Xu 0004 |
Neurocomputing | 1 |
| 2021 | Attention-based interpolation network for video deblurring
Xiaoqin Zhang 0002, Runhua Jiang, Tao Wang 0052, Pengcheng Huang 0002, Li Zhao 0005 |
Neurocomputing | 3 |
| 2021 | Robust feature learning for adversarial defense via hierarchical feature alignment
Xiaoqin Zhang 0002, Tao Wang 0052, Runhua Jiang, Jiawei Xu 0004, Li Zhao 0005 |
Inf. Sci. | 3 |
| 2021 | Recursive Neural Network for Video DeblurringabstractVideo deblurring is still a challenging low-level vision task since spatio-temporal characteristics across both the spatial and temporal domains are difficult to model. In this article, to model the temporal information, we develop a non-local block which estimates inter-frame similarity and inter-frame difference. Specially, for modeling the spatial characteristics and restoring sharp frame details, we propose a recursive block that iteratively refines feature maps generated at the last iteration. In addition, a novel temporal loss function is introduced to ensure the temporal consistency of generated frames. Experimental results on public datasets demonstrate that our method achieves state-of-the-art performance both quantitatively and qualitatively. Xiaoqin Zhang 0002, Runhua Jiang, Tao Wang 0052 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Multi-Level Fusion and Attention-Guided CNN for Image DehazingabstractIn this paper, we tackle the problem of single image dehazing with a convolutional neural network. Within this network, we develop a multi-level fusion module to utilize both low-level and high-level features. The low-level features help to recover finer details, and the high-level features discover abstract semantics. They are complementary in the restoring of clear images. Moreover, a Residual Mixed-convolution Attention Module (RMAM) with an attention block is proposed to guide the network to focus on important features in the learning process. In this RMAM, group convolution, depth-wise convolution, and point-wise convolution are mixed, and thus it is much faster than its counterparts. With these two modules, we thus have an end-to-end network without explicitly estimating the atmospheric light intensity and the transmission map in the classical atmosphere scattering model. Both qualitative and quantitative experimental studies are carried out on public datasets including RESIDE, DCPDN-TestA, and the real-world dataset. The extensive results demonstrate both the effectiveness and efficiency of the proposed solution to single image dehazing. Xiaoqin Zhang 0002, Tao Wang 0052, Wenhan Luo, Pengcheng Huang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Single Image Dehazing via Dual-Path Recurrent NetworkabstractAn image can be decomposed into two parts: the basic content and details, which usually correspond to the low-frequency and high-frequency information of the image. For a hazy image, these two parts are often affected by haze in different levels, e.g., high-frequency parts are often affected more serious than low-frequency parts. In this paper, we approach the single image dehazing problem as two restoration problems of recovering basic content and image details, and propose a Dual-Path Recurrent Network (DPRN) to simultaneously tackle these two problems. Specifically, the core structure of DPRN is a dual-path block, which uses two parallel branches to learn the characteristics of the basic content and details of hazy images. Each branch consists of several Convolutional LSTM blocks and convolution layers. Moreover, a parallel interaction function is incorporated into the dual-path block, thus enables each branch to dynamically fuse the intermediate features of both the basic content and image details. In this way, both branches can benefit from each other, and recover the basic content and image details alternately, therefore alleviating the color distortion problem in the dehazing process. Experimental results show that the proposed DPRN outperforms state-of-the-art image dehazing methods in terms of both quantitative accuracy and qualitative visual effect. Xiaoqin Zhang 0002, Runhua Jiang, Tao Wang 0052, Wenhan Luo |
IEEE Trans. Image Process. | 3 |
| 2020 | Structured Dictionary Learning with Block Diagonal Regularization for Image ClassificationabstractSparse representation and dictionary learning have been successfully applied to encode dense data and facilitate image classification. Though existing dictionary learning methods achieve better performance than their counterparts, the class discriminative ability of learned dictionary is still limited. This paper proposes a novel supervised dictionary learning method based on the prior of the block diagonal phenomenon, i.e., each sample should be well reconstructed by the samples in the same class while poorly reconstructed by the samples in other class. Specifically, a block diagonal regularizer is imposed on the affinity matrix to enforce the sparse representation matrix to have an approximately block diagonal structure, which makes the learned dictionary more discriminative and suitable for classification tasks. Furthermore, we present an effective optimization strategy by combining the alternating minimization with the alternating direction method of multipliers (ADMM) for the proposed framework. Experimental results on six real-world datasets show that the proposed method is more effective than state-of-the-art dictionary learning methods. Manman Xu, Runhua Jiang, Tao Wang 0052, Di Wang 0008, Xiaoju Lu |
IEEE BigData | 3 |
| 2020 | Pyramid Channel-based Feature Attention Network for image dehazing
Xiaoqin Zhang 0002, Tao Wang 0052, Guiying Tang, Li Zhao 0005 |
Comput. Vis. Image Underst. | 2 |