VLDB 2026 Research / reviewers in the wild / expert
Li Zhao 0005
dblp:97/4708-5
· DBLP profile ↗
38ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0001-5787-2705ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 2 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 5 since 2021Databases, data management, data science and information retrieval · 9 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IGIANet: Illumination Guided Implicit Alignment Network for Infrared-Visible UAV DetectionabstractVisible-Infrared (RGB-IR) Unmanned Aerial Vehicle (UAV) object detection integrates complementary cues from visible and infrared sensors, offering broad application potential. However, due to sensor parallax, it still faces the challenge of weak spatial misalignment, which significantly limits its performance in UAV-based object detection. Existing methods emphasize strict alignment, overlooking spectral heterogeneity under varying illumination. To address these issues, we propose the Illumination Guided Implicit Alignment Network (IGIANet) to mitigate modality heterogeneity without explicit alignment. Specifically, we integrate three novel modules. First, we propose an illumination-guided frequency modulation module that adaptively allocates fusion weights to visible and infrared features based on global illumination estimation, effectively alleviating modality imbalance under varying lighting conditions. Second, we introduce a frequency-guided cross-modality differential enhancement module, which computes differential cues across frequency domains to enhance complementary information and highlight weakly aligned and low-contrast regions. Finally, we introduce an implicit alignment-driven dynamic fusion module that actively estimates offsets and generates dynamic, position-adaptive fusion kernels to align and fuse modalities. Extensive experiments demonstrate that IGIANet outperforms state-of-the-art models on various benchmarks, achieving 80.9% mAP on DroneVehicle, 57.1% mAP on VEDAI, and 49.4% mAP on FLIR. Xiangqi Chen, Dawei Zhang 0002, Li Zhao 0005, Chengzhuan Yang, Jungang Lou, Zhonglong Zheng, Sang-Woon Jeon, Hua Wang 0002 |
AAAI | 3 |
| 2026 | Template-Free Tracking Guidance for transformer trackers
Xuan Wang 0032, Li Zhao 0005, Dawei Zhang 0002, Chengzhuan Yang, Jungang Lou, Yunliang Jiang, Jinli Cao, Zhonglong Zheng |
Knowl. Based Syst. | 2 |
| 2026 | A cross-domain feature fusion network for nighttime drone-view object detection
Xiangqi Chen, Chengzhuan Yang, Jiashuaizi Mo, Li Zhao 0005, Zhonglong Zheng |
Pattern Recognit. | 5 |
| 2026 | APDiff: An Adaptive Physics-Guided Diffusion Framework for efficient unpaired image dehazing
Li Zhao 0005, Hanqi Wang, Chenxiang Fan, Haigen Hu, Wenqi Ren, Zhonglong Zheng |
Pattern Recognit. | 1 |
| 2026 | Corrections to "Exploring Fuzzy Priors From Multimapping GAN for Robust Image Dehazing"
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, Li Zhao 0005, En Fan, Feng Huang 0007 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2025 | SAM-TTT: Segment Anything Model via Reverse Parameter Configuration and Test-Time Training for Camouflaged Object DetectionabstractThis paper introduces a new Segment Anything Model (SAM) that leverages reverse parameter configuration and test-time training to enhance its performance on Camouflaged Object Detection (COD), named SAM-TTT. While most existing SAM-based COD models primarily focus on enhancing SAM by extracting favorable features and amplifying its advantageous parameters, a crucial gap is identified: insufficient attention to adverse parameters that impair SAM's semantic understanding in downstream tasks. To tackle this issue, the Reverse SAM Parameter Configuration Module is proposed to effectively mitigate the influence of adverse parameters in a train-free manner by configuring SAM's parameters. Building on this foundation, the T-Visioner Module is unveiled to strengthen advantageous parameters by integrating Test-Time Training layers, originally developed for language tasks, into vision tasks. Test-Time Training layers represent a new class of sequence modeling layers characterized by linear complexity and an expressive hidden state. By integrating two modules, SAM-TTT simultaneously suppresses adverse parameters while reinforcing advantageous ones, significantly improving SAM's semantic understanding in COD task. Our experimental results on various COD benchmarks demonstrate that the proposed approach achieves state-of-the-art performance, setting a new benchmark in the field. The code will be available at https://github.com/guobaoxiao/SAM-TTT. Zhenni Yu, Li Zhao 0005, Guobao Xiao, Xiaoqin Zhang 0002 |
ACM Multimedia | 2 |
| 2025 | Physics-Guided Diffusion Model for Unpaired Real-World Dehazing
Hanqi Wang, Chenxiang Fan, Haigen Hu, Li Zhao 0005, Xiaoqin Zhang 0002 |
PRCV (9) | 4 |
| 2025 | COMPrompter: reconceptualized segment anything model with multiprompt network for camouflaged object detection
Xiaoqin Zhang 0002, Zhenni Yu, Li Zhao 0005, Deng-Ping Fan, Guobao Xiao |
Sci. China Inf. Sci. | 3 |
| 2025 | DiFusionSeg: Diffusion-driven semantic segmentation with multi-modal image fusion for enhanced perception
Defeng He, Li Zhao 0005, Yayu Zheng, Xiaoqin Zhang 0002 |
Knowl. Based Syst. | 3 |
| 2025 | Global-local feature-mixed network with template update for visual tracking
Li Zhao 0005, Chenxiang Fan, Min Li 0052, Zhonglong Zheng, Xiaoqin Zhang 0002 |
Pattern Recognit. Lett. | 1 |
| 2025 | Exploring Fuzzy Priors From Multimapping GAN for Robust Image DehazingabstractSingle image dehazing has been extensively studied. While convolutional neural networks (CNNs) have driven notable progress in single image dehazing, their performance remains fundamentally constrained by the limited local receptive fields of convolutional operations, which impede the capture of global structural dependencies. In contrast, generative adversarial networks (GANs) have demonstrated exceptional capabilities in image synthesis, offering global insights into structure, texture, and color. The fuzzy prior, a probabilistic knowledge acquired through adversarial training in GANs, plays a pivotal role in robust dehazing. Motivated by this, we propose the fuzzy prior guided dehazing network (FPGDN). Our framework begins with a novel module that distills the fuzzy prior by translating an edge map into a color image, simultaneously capturing global structural, local textural, and color information. Subsequently, a dehazing network is constructed, leveraging this fuzzy prior. While the fuzzy prior captures rich color and texture features, the generated images may exhibit color shifts relative to the original scene. To remedy this, a CNN network is employed to capture local nuances and refine the dehazing outcome. Extensive experiments substantiate that the proposed FPGDN achieves superior dehazing performance on a variety of real and synthetic hazy images. Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, Li Zhao 0005, En Fan, Feng Huang 0007 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2025 | Mask-Guided Frequency Feature Fusion for Visible-Infrared Remote Sensing Object DetectionabstractVisible-infrared remote sensing object detection aims to achieve all-weather object detection by leveraging the complementary information from paired visible and infrared (RGB-IR) images. However, modality differences and weak alignment often limit its performance. Existing methods largely neglect the frequency discrepancies between modalities and require strict alignment, increasing complexity. To address these challenges, this study proposes a novel mask-guided frequency feature fusion (MGFF) method for RGB-IR object detection in remote sensing. Specifically, we develop a feature frequency decomposition and enhancement module using wavelet transform to reduce modality differences between RGB and IR images by restructuring and enhancing their frequency components. Additionally, we introduce a mask-guided feature reconstruction module and a feature-guided consistency loss, ensuring that even under weak alignment, the focus remains on integrating the target features from different modalities. Meanwhile, this loss is used to guide the reconstruction of features from different modalities. Finally, We design a multi-directional perception cross-modality fusion module to achieve deep fusion of multimodal information, which enhances object perception from different directions across modalities. Extensive evaluations on the widely recognized RGB-IR remote sensing benchmarks, including DroneVehicle and VEDAI, as well as the RGB-IR pedestrian dataset KAIST, substantiate the effectiveness of the proposed MGFF method. The results consistently demonstrate that the MGFF achieves a superior performance in terms of detection accuracy and robustness compared to existing state-of-the-art approaches. Xiangqi Chen, Li Zhao 0005, Chengzhuan Yang, Dawei Zhang 0002, Xiao Wang 0014, Xiaowei He 0003, Hua Wang 0002, Zhonglong Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | VSFormer: Visual-Spatial Fusion Transformer for Correspondence PruningabstractCorrespondence pruning aims to find correct matches (inliers) from an initial set of putative correspondences, which is a fundamental task for many applications. The process of finding is challenging, given the varying inlier ratios between scenes/image pairs due to significant visual differences. However, the performance of the existing methods is usually limited by the problem of lacking visual cues (e.g., texture, illumination, structure) of scenes. In this paper, we propose a Visual-Spatial Fusion Transformer (VSFormer) to identify inliers and recover camera poses accurately. Firstly, we obtain highly abstract visual cues of a scene with the cross attention between local features of two-view images. Then, we model these visual cues and correspondences by a joint visual-spatial fusion module, simultaneously embedding visual cues into correspondences for pruning. Additionally, to mine the consistency of correspondences, we also design a novel module that combines the KNN-based graph and the transformer, effectively capturing both local and global contexts. Extensive experiments have demonstrated that the proposed VSFormer outperforms state-of-the-art methods on outdoor and indoor benchmarks. Our code is provided at the following repository: https://github.com/sugar-fly/VSFormer. Tangfei Liao, Xiaoqin Zhang 0002, Li Zhao 0005, Tao Wang 0047, Guobao Xiao |
AAAI | 3 |
| 2024 | Exploring Deeper! Segment Anything Model with Depth Perception for Camouflaged Object DetectionabstractThis paper introduces a new Segment Anything Model with Depth Perception (DSAM) for Camouflaged Object Detection (COD). DSAM exploits the zero-shot capability of SAM to realize precise segmentation in the RGB-D domain. It consists of the Prompt-Deeper Module and the Finer Module. The Prompt-Deeper Module utilizes knowledge distillation and the Bias Correction Module to achieve the interaction between RGB features and depth features, especially using depth features to correct erroneous parts in RGB features. Then, the interacted features are combined with the box prompt in SAM to create a prompt with depth perception. The Finer Module explores the possibility of accurately segmenting highly camouflaged targets from a depth perspective. It uncovers depth cues in areas missed by SAM through mask reversion, self-filtering, and self-attention operations, compensating for its defects in the COD domain. DSAM represents the first step towards the SAM-based RGB-D COD model. It maximizes the utilization of depth features while synergizing with RGB features to achieve multimodal complementarity, thereby overcoming the segmentation limitations of SAM and improving its accuracy in COD. Experimental results on COD benchmarks demonstrate that DSAM achieves excellent segmentation performance and reaches the state-of-the-art (SOTA) on COD benchmarks with less consumption of training resources. The code will be available at https://github.com/guobaoxiao/DSAM. Zhenni Yu, Xiaoqin Zhang 0002, Li Zhao 0005, Yi Bin, Guobao Xiao |
ACM Multimedia | 3 |
| 2024 | Photo realistic synthetic dataset and multi-scale attention dehazing network
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, LinLin Shen, Li Zhao 0005, Jun Zhang 0011 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | A novel non-pretrained deep supervision network for polyp segmentation
Zhenni Yu, Li Zhao 0005, Tangfei Liao, Xiaoqin Zhang 0002, Geng Chen 0001, Guobao Xiao |
Pattern Recognit. | 2 |
| 2024 | Tensor Recovery With Weighted Tensor Average RankabstractIn this article, a curious phenomenon in the tensor recovery algorithm is considered: can the same recovered results be obtained when the observation tensors in the algorithm are transposed in different ways? If not, it is reasonable to imagine that some information within the data will be lost for the case of observation tensors under certain transpose operators. To solve this problem, a new tensor rank called weighted tensor average rank (WTAR) is proposed to learn the relationship between different resulting tensors by performing a series of transpose operators on an observation tensor. WTAR is applied to three-order tensor robust principal component analysis (TRPCA) to investigate its effectiveness. Meanwhile, to balance the effectiveness and solvability of the resulting model, a generalized model that involves the convex surrogate and a series of nonconvex surrogates are studied, and the corresponding worst case error bounds of the recovered tensor is given. Besides, a generalized tensor singular value thresholding (GTSVT) method and a generalized optimization algorithm based on GTSVT are proposed to solve the generalized model effectively. The experimental results indicate that the proposed method is effective. Xiaoqin Zhang 0002, Li Zhao 0005, Zhengyuan Zhou, Zhouchen Lin |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Random Reconstructed Unpaired Image-to-Image TranslationabstractThe goal of unpaired image-to-image translation is to learn a mapping from a source domain to a target domain without using any labeled examples of paired images. This problem can be solved by learning the conditional distribution of source images in the target domain. A major limitation of existing unpaired image-to-image translation algorithms is that they generate untruthful images which are overcolored and lack details, while the translation of realistic images must be rich in details. To address this limitation, in this article, we propose a random reconstructed unpaired image-to-image translation (RRUIT) framework by generative adversarial network, which uses random reconstruction to preserve the high-level features in the source and adopts an adversarial strategy to learn the distribution in the target. We update the proposed objective function with two loss functions. The auxiliary loss guides the generator to create a coarse image, while the coarse-to-fine block next to the generator block produces an image that obeys the distribution of the target domain. The coarse-to-fine block contains two submodules based on the densely connected atrous spatial pyramid pooling, which enriches the details of generated images. We conduct extensive experiments on photorealistic stylization and artistic stylization. The experimental results confirm the superiority of the proposed RRUIT. Xiaoqin Zhang 0002, Chenxiang Fan, Zhiheng Xiao, Li Zhao 0005, Huiling Chen 0001, Xiaojun Chang |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Multi-scale Residual Interaction for RGB-D Salient Object Detection
Mingjun Hu, Xiaoqin Zhang 0002, Li Zhao 0005 |
ACCV (3) | 3 |
| 2022 | Dual Priors Network for RGB-D Salient Object DetectionabstractAlthough the detection accuracy of RGB-D salient object detection by deep learning has met most of the task requirements, it is still a challenge to predict complete salient objects under the guidance of poor depth maps. In this paper, edge prior and depth prior are utilized to guide the network to detect salient objects, and a novel dual priors network called DPNet is proposed for RGB-D SOD. DPNet utilizes edge priors to compensate for the inadequate guidance of low-quality depth maps. Specially, the input RGB and depth maps are encoded by ResNet-50 backbone. To fuse these information effectively, multi-modal feature fusion module downsamples high-resolution features to enhance low-resolution features with its rich semantic information. Then the initial salient masks are decoded by a coarse-grained mask decoder. In addition, edge prior serves as label and is captured by an edge aware module. Finally, the fine-grained salient masks are obtained by fusing the initial salient masks and the salient edges. The experimental results on six benchmarks indicate that the proposed method outperforms ten state-of-the-art methods in six evaluation metrics. Yuewang Xu, Li Zhao 0005, Shaoli Cao, Shijie Feng |
IEEE Big Data | 2 |
| 2022 | Multi-Attention Convolutional Neural Network for Video DeblurringabstractVideo deblurring, which aims at restoring the sharp video from blurry video, is drawing increasing attention in the field of computer vision. In this paper, a method called Multi-Attention Convolutional Neural Network (MACNN) consisting of the temporal-spatial attention module, the frame channel attention module, and the feature extraction-reconstruction module is proposed. First, we use the temporal-spatial attention module and the frame channel attention module to capture features with temporal and spatial information existing across neighboring frames. Then, these captured features are fused and reconstructed to restore the sharp frame. Last but not least, we train MACNN together with a content loss and a perceptual loss in an end-to-end manner to recover realistic video details. Both quantitative and qualitative evaluation results on standard benchmarks demonstrate the proposed MACNN is superior to the state-of-the-art methods in terms of accuracy, efficiency, and visual effect. Xiaoqin Zhang 0002, Tao Wang 0052, Runhua Jiang, Li Zhao 0005, Yuewang Xu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Single Image Haze Removal Based on a Simple Additive Model With Haze Smoothness PriorabstractSingle image haze removal, which is to recover the clear version of a hazy image, is a challenging task in computer vision. In this paper, an additive haze model is proposed to approximate the hazy image formation process. In contrast with the traditional optical model, it regards the haze as an additive layer to a clean image. The model thus avoids estimating the medium transmission rate and the global atmospherical light. In addition, based on a critical observation that haze changes gradually and smoothly across the image, a haze smoothness prior is proposed to constrain this model. This prior assumes that the haze layer is much smoother than the clear image. Benefiting from this prior, we can directly separate the clean image from a single hazy image. Experimental results and comparisons with synthetic images and real-world images demonstrate that the proposed method outperforms state-of-the-art single image haze removal algorithms. Xiaoqin Zhang 0002, Tao Wang 0052, Guiying Tang, Li Zhao 0005, Yuewang Xu, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Two-Stage Image Dehazing with Depth Information and Cross-Scale Non-Local AttentionabstractImage dehazing using learning-based methods has achieved state-of-the-art performance in recent years. However, most of previous models for low-level vision tasks are always based on single-stage design. There is an issue with majority image dehazing approaches: a complex balance problem between spatial details and high-level contextualized information while recovering images. To address this issue, we propose a two-stage network, which consists of encoder-decoder subnetwork and original-resolution one. Specifically, the encoder-decoder subnetwork first learns the contextualized features, and then the features are combined with the original-resolution subnet that maintains enriched high-resolution features. For information exchange between two stages, we introduce a Cross-Scale Non-Local (CS-NL) attention module that can search more high-frequency details from low-resolution images in encoder-decoder stage and transfer directly them to the next stage. Moreover we embed the spatial feature transform (SFT) module into original-resolution subnet, which is incorporated with depth information to better achieve the purpose of image dehazing. The two-stage network, named as TSDCN-Net, which demonstrates its effectiveness according extensive experiments. The TSDCN-Net surpasses previous state-of-the-art single image dehazing methods by a large margin both quantitatively and qualitatively. Li Zhao 0005 |
IEEE BigData | 2 |
| 2021 | Video Deblurring via Spatiotemporal Pyramid Network and Adversarial Gradient Prior
Tao Wang 0052, Xiaoqin Zhang 0002, Runhua Jiang, Li Zhao 0005, Huiling Chen 0001, Wenhan Luo |
Comput. Vis. Image Underst. | 4 |
| 2021 | Self-filtering image dehazing with self-supporting module
Pengcheng Huang 0002, Li Zhao 0005, Runhua Jiang, Tao Wang 0052, Xiaoqin Zhang 0002 |
Neurocomputing | 2 |
| 2021 | Haze concentration adaptive network for image dehazing
Tao Wang 0052, Li Zhao 0005, Pengcheng Huang 0002, Xiaoqin Zhang 0002, Jiawei Xu 0004 |
Neurocomputing | 2 |
| 2021 | Attention-based interpolation network for video deblurring
Xiaoqin Zhang 0002, Runhua Jiang, Tao Wang 0052, Pengcheng Huang 0002, Li Zhao 0005 |
Neurocomputing | 5 |
| 2021 | Robust feature learning for adversarial defense via hierarchical feature alignment
Xiaoqin Zhang 0002, Tao Wang 0052, Runhua Jiang, Jiawei Xu 0004, Li Zhao 0005 |
Inf. Sci. | 6 |
| 2020 | Self-calibrated Attention Residual Network for Image Super-ResolutionabstractDeep Convolutional Neural Networks (DCNNs) have achieved remarkable performance in single image super-resolution (SISR). However, most SR methods restore high resolution (HR) images from single-scale region in the low resolution (LR) input, which limits the ability of method to infer multi-scales of details for high resolution (HR) output. In this paper a novel basic building block called self-Calibrated residual block (SARB) is proposed to solve this problem. SARB consists of carefully designed multi-scale paths, which can capture rich structure information from different scale. In addition, self-Calibrated residual block is introduced to adaptively learn informatively context to make network generate more discriminative representations. These blocks are composed of self-calibrated attention residual network (SARN) for image super-resolution. Experiments results on five benchmark datasets demonstrate that the proposed SARN achieves comparable results compared with the previous most of the state-of-the-art methods. Anqi Rong, Li Zhao 0005, Pengcheng Huang 0002, Jiawei Xu 0004 |
IEEE BigData | 2 |
| 2020 | A Nonlocal Denoising Framework Based on Tensor Robust Principal Component Analysis with ℓp normabstractThis paper have given a nonlocal denoising framework based on tensor robust principal component analysis with ℓpnorm for color image and video (NDFCIV), which have following three features: (1) it is capable of processing zero-mean Gaussian noise, impulse noise and any other noise that is created by mixing the two for color image and video at same time. (2) Meanwhile, nonlocal denoising strategy is adopted to promote the effectiveness of the denoising framework. (3) Moreover, we present a non-convex constraint method which can get more exact low rank tensor recovery result and enhance the denoising effect of the framework further. The experimental results demonstrate the effectiveness of the proposed denoising framework. Mengqing Sun, Li Zhao 0005, Jiawei Xu 0004 |
IEEE BigData | 2 |
| 2020 | Feature Fusion Based on Sparse Block for Image Super-resolutionabstractRecently, deep neural networks have been widely used in the task of single image super-resolution. However, existing deep neural networks always take huge parameters to map low-resolution images to high-resolution ones. In addition, most of them only consider high-level features to reconstruct high-resolution images. These two methodologies not only cause the difficulty of practical applications but also the inefficiency of restoring image details. Therefore, in this work, the authors propose a novel sparse block to learn high-level features. Based on this block, a fusion method is proposed to fuse features from multiple levels. By incorporating these two approaches, a lightweight neural network, i.e. Sparse Block Fusion Network (SBFN), is proposed for end-to-end training. Through extensive experiments, it is demonstrated the proposed methods can achieve comparable performance with few parameters. By making comprehensive comparisons, effectiveness of SBFN is also verified in multiple benchmark datasets. Shengping Wang, Li Zhao 0005, Runhua Jiang, Pengcheng Huang 0002, Jiawei Xu 0004 |
IEEE BigData | 2 |
| 2020 | Multi-level Feature Fusion Network for Single Image Super-ResolutionabstractRecently, deep convolution neural networks have achieved remarkable performance in the task of single image super-resolution (SISR). However, effectiveness of existing networks highly relies on their receptive field, which always increases with the depth of the network. In this work, we propose a novel module, named as residual group, to effectively learn feature maps by using dynamic receptive field. This residual group firstly uses a selective kernel convolution layer to dynamically learn multi-scale information from its input features. Then, several residual blocks are employed to further refine the learned feature. In addition, we also propose a selective feature fusion module to fuse appearance information in multi-level features. Within this module, the low-level features and high-level features are selectively fused to complement the high-level ones. Finally, by combining these two methods, we introduce a multi-level feature fusion network (MLFFN) for single image super-resolution (SISR). Through comprehensive experiments, we demonstrate that the proposed MLFFN achieves state-of-the-art performance both quantitatively and qualitatively. Xinxia Zhang, Xiaoqin Zhang 0002, Li Zhao 0005, Runhua Jiang, Pengcheng Huang 0002, Jiawei Xu 0004 |
IEEE BigData | 3 |
| 2020 | Pyramid Channel-based Feature Attention Network for image dehazing
Xiaoqin Zhang 0002, Tao Wang 0052, Guiying Tang, Li Zhao 0005 |
Comput. Vis. Image Underst. | 5 |
| 2020 | Exemplar-Based Denoising: A Unified Low-Rank Recovery FrameworkabstractExemplar-based image denoising algorithms have shown great potential for image restoration with a multitude of existing models. In this paper, we interpret nonlocal similar patch-based denoising as a problem of low-rank recovery. This offers a physically plausible model and unifies several existing techniques in a single low-rank recovery framework. The framework can handle complex noise models, such as zero-mean Gaussian noise, impulse noise, and any other noise that can be approximated by mixing these two kinds of noise. Moreover, we introduce a new nonconvex surrogate for the $l_{0}$ -norm and find the optimal solution of the optimization problems when the new norm is applied to low-rank recovery. The experimental results with different kinds of noise confirm the effectiveness of the proposed low-rank recovery framework and the new norm. Xiaoqin Zhang 0002, Di Wang 0008, Li Zhao 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Self-Taught Semisupervised Dictionary Learning With Nonnegative ConstraintabstractThis paper investigates classification by dictionary learning. A novel unified framework termed self-taught semisupervised dictionary learning with nonnegative constraint is proposed for simultaneously optimizing the components of a dictionary and a graph Laplacian. Specifically, an atom graph Laplacian regularization is built by using sparse coefficients to effectively capture the underlying manifold structure. It is more robust to noisy samples and outliers because atoms are more concise and representative than training samples. A nonnegative constraint imposed on the sparse coefficients guarantees that each sample is in the middle of its related atoms. In this way, the dependency between samples and atoms is made explicit. Furthermore, a self-taught mechanism is introduced to effectively feed back the manifold structure induced by atom graph Laplacian regularization and the supervised information hidden in unlabeled samples in order to learn a better dictionary. An efficient algorithm, combining a block coordinate descent method with the alternating direction method of multipliers, is derived to optimize the unified framework. Experimental results on several benchmark datasets show the effectiveness of the proposed model. Xiaoqin Zhang 0002, Di Wang 0008, Li Zhao 0005, Nannan Gu, Stephen J. Maybank |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Single Image Dehazing via Lightweight Multi-scale NetworksabstractSingle image haze removal is a challenging ill-posed problem in computer vision. Instead of leveraging the traditional model or handcrafted image priors, an end-to-end multi-scale convolutional neural network is proposed for single image haze removal task by directly mapping the hazy image to its corresponding haze-free image. To better retain the coarse and fine information, a multi-scale block is elaborated and embedded into the proposed architecture. This block can extract the feature at varying scales with a model size that is as small as possible. The global skip connection is adopted to promote the model performance. Extensive experiment results demonstrate that the proposed network outperforms the state-of-the-art single image haze removal algorithms on both synthetical and real-world images. In addition, the size of the model in this paper dominates among the high performance methods based on convolutional neural networks. Guiying Tang, Li Zhao 0005, Runhua Jiang, Xiaoqin Zhang 0002 |
IEEE BigData | 2 |
| 2019 | Single-Image Dehazing Using Color Attenuation Prior Based on Haze-LinesabstractIn this paper, we propose a new single-image dehazing method for synthetic and real-world hazy images. Based on the color attenuation prior, this proposed dehazing method improves it in two aspects. First, we estimate the atmospheric light with the haze-lines prior, which is based on the observation that pixel values of a hazy image can be modeled as lines in the RGB color space that intersects at the air-light. Second, the dynamic scattering coefficient, which is an exponential function of image depth, is proposed to replace the constant scattering coefficient. Experimental results demonstrate that the dehazed image of proposed algorithm is clearer and more natural than that of the color attenuation prior. The proposed algorithm can effectively improve the effect of dehazing. Qianru Wang, Li Zhao 0005, Guiying Tang, Hanli Zhao, Xiaoqin Zhang 0002 |
IEEE BigData | 2 |
| 2019 | Semantic Segmentation of Remote Sensing Images Using Multiscale Decoding NetworkabstractIn this letter, we propose a practical convolutional neural network architecture for semantic pixelwise segmentation of remote sensing images, named Multiscale Decoding Network. The proposed method is built on the success of fully convolutional networks (FCNs) and the transfer of pretrained networks. The decoding network of our architecture utilizes the combination of three paths, namely, unpooling path, transposed convolution path, and dilated convolution path, in the form of an inception module. The whole network is trained in the end-to-end manner and the parameters of the three paths are learned automatically. Since the proposed method transfers the feature of pretrained networks and has three simplified decoding paths with fewer parameters, it requires less training data and training time. Compared with the classical networks FCN, SegNet, and U-net, our network shows better performance on remote sensing images segmentation. Xiaoqin Zhang 0002, Zhiheng Xiao, Mingyu Fan, Li Zhao 0005 |
IEEE Geosci. Remote. Sens. Lett. | 5 |