EDBT 2026 Demo / reviewers in the wild / expert
Qian Zhao 0002
dblp:82/4299-2
· DBLP profile ↗
70ranked-venue papers
10as first author
30since 2021 · last 2026
0000-0001-9956-0064ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 8 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Computer networks · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Singular Value Fine-Tuning for Few-Shot Class-Incremental LearningabstractClass-Incremental Learning (CIL) aims to prevent catastrophic forgetting of previously learned classes while sequentially incorporating new ones. The more challenging Few-shot CIL (FSCIL) setting further complicates this by providing only a limited number of samples for each new class, increasing the risk of overfitting in addition to standard CIL challenges. While catastrophic forgetting has been extensively studied, overfitting in FSCIL, especially with large foundation models, has received less attention. To fill this gap, we propose the Singular Value Fine-tuning for FSCIL (SVFCL) and compared it with existing approaches for adapting foundation models to FSCIL, which primarily build on Parameter Efficient Fine-Tuning (PEFT) methods like prompt tuning and Low-Rank Adaptation (LoRA). Specifically, SVFCL applies singular value decomposition to the foundation model weights, keeping the singular vectors fixed while fine-tuning the singular values for each task, and then merging them. This simple yet effective approach not only alleviates the forgetting problem but also mitigates overfitting more effectively while significantly reducing trainable parameters. Extensive experiments on four benchmark datasets, along with visualizations and ablation studies, validate the effectiveness of SVFCL. The code will be made available. Zhiwu Wang, Renzhen Wang, Haokun Lin, Quanziang Wang, Qian Zhao 0002, Deyu Meng |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Enhancing Underwater Light Field Images via Global Geometry-Aware Diffusion ProcessabstractThis work studies the challenging problem of acquiring high-quality underwater images via 4-D light field (LF) imaging. To this end, we propose GeoDiff-LF, a novel diffusion-based framework built upon SD-Turbo to enhance underwater 4-D LF imaging by leveraging its spatial-angular structure. GeoDiff-LF consists of three key adaptations: 1) a modified U-Net architecture with convolutional and attention adapters to model geometric cues, 2) a geometry-guided loss function using tensor decomposition and progressive weighting to regularize global structure, and 3) an optimized sampling strategy with noise prediction to improve efficiency. By integrating diffusion priors and LF geometry, GeoDiff-LF effectively mitigates color distortion in underwater scenes. Extensive experiments demonstrate that our framework outperforms existing methods across both visual fidelity and quantitative performance, advancing the state-of-the-art in enhancing underwater imaging. The code will be publicly available at https://github.com/linlos1234/GeoDiff-LF. Yuji Lin, Qian Zhao 0002, Zongsheng Yue, Junhui Hou, Deyu Meng |
IEEE Trans. Image Process. | 2 |
| 2025 | Graph Domain Adaptation With Dual-Branch Encoder and Two-Level Alignment for Whole Slide Image-Based Survival Prediction
Yuntao Shou, Xiangyong Cao, Peiqiang Yan, Qiaohui, Qian Zhao 0002, Deyu Meng |
ICCV | 5 |
| 2024 | Blind Image Deconvolution by Generative-Based Kernel Prior and Initializer via Latent Encoding
Zongsheng Yue, Hui Wang 0103, Qian Zhao 0002, Deyu Meng |
ECCV (46) | 4 |
| 2024 | Deep Variational Network Toward Blind Image RestorationabstractBlind image restoration (IR) is a common yet challenging problem in computer vision. Classical model-based methods and recent deep learning (DL)-based methods represent two different methodologies for this problem, each with their own merits and drawbacks. In this paper, we propose a novel blind image restoration method, aiming to integrate both the advantages of them. Specifically, we construct a general Bayesian generative model for the blind IR, which explicitly depicts the degradation process. In this proposed model, a pixel-wise non-i.i.d. Gaussian distribution is employed to fit the image noise. It is with more flexibility than the simple i.i.d. Gaussian or Laplacian distributions as adopted in most of conventional methods, so as to handle more complicated noise types contained in the image degradation. To solve the model, we design a variational inference algorithm where all the expected posteriori distributions are parameterized as deep neural networks to increase their model capability. Notably, such an inference algorithm induces a unified framework to jointly deal with the tasks of degradation estimation and image restoration. Further, the degradation information estimated in the former task is utilized to guide the latter IR process. Experiments on two typical blind IR tasks, namely image denoising and super-resolution, demonstrate that the proposed method achieves superior performance over current state-of-the-arts. Zongsheng Yue, Hongwei Yong, Qian Zhao 0002, Lei Zhang 0006, Deyu Meng, Kwan-Yee Kenneth Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | HQ-IRN: Quantizing High-Frequency Features for Image RescalingabstractSaving and transmitting high-resolution (HR) images are often demanded in real life, especially in social media applications. The recently developed image rescaling techniques provide a storage and transmission economic way to deal with this problem, by jointly learning the downscaling and upscaling mappings for the images with the aid of the invertible neural network (INN). In the original pipeline, the high-frequency information is generally discarded in the downscaled low-resolution (LR) image, while is randomly sampled when doing upscaling, so that the storage and transmission cost can be minimal. However, the quality of the reconstructed HR image is limited due to the ignorance of the high-frequency information, and thus there are researchers trying to improve the upscaling performance by paying a bit more storage to partially save the high-frequency features. In this work, following this research line, we propose a new strategy to improve the image rescaling performance by more efficiently utilizing additional storage. Specifically, instead of saving the partial high-frequency features, we propose to quantize those features with a learned codebook and save the corresponding index matrix. Such a vector quantization strategy can recover as much as possible high-frequency features, and thus leads to a better image rescaling performance. Besides, the additional storage cost is the same or can be even less compared with existing methods. Experiments on a series of benchmark datasets demonstrate the effectiveness of the proposed method against current state-of-the-art ones. Zibo Song, Qian Zhao 0002, Deyu Meng |
IEEE Signal Process. Lett. | 2 |
| 2024 | Learnable Representative Coefficient Image Denoiser for Hyperspectral ImageabstractFully characterizing the spatial-spectral priors of hyperspectral images (HSI) is crucial for HSI denoising tasks. Recently, HSI denoising models based on representative coefficient images (RCIs) under the spectral low-rank decomposition framework have garnered significant attention due to their clever utilization of spatial-spectral information in HSI at a low cost. However, current methods either employ handcrafted classical denoisers or off-the-shelf deep denoisers to denoise RCIs, failing to fully capture the structural information of RCIs. In this paper, we propose a specific optimization framework for learning an RCI denoiser under the low-rank decomposition framework for the first time. Since low-rank decomposition can characterize the global low-rank property of HSI, our RCI denoiser only needs to learn the spatial prior of RCIs. Consequently, our optimization framework is inclined to learn a more powerful RCI denoiser. However, learning an RCI denoiser is not an easy task, primarily due to the lack of paired clean-noisy RCI data. To address this issue, we employ parametric techniques to represent the to-be-restored HSI as a function of RCI denoiser network parameters. In this way, the parameters of the RCI denoiser can thus be updated using noisy-clean HSI pairs. Furthermore, we adopt residual learning and Gaussian whitening techniques to enhance the RCI denoiser’s denoising ability for HSIs with various noise levels and different rank settings. Extensive experiments demonstrate that our method can achieve significant improvements in both denoising effectiveness and speed compared to state-of-the-art methods. The code of our algorithm is released at https://github.com/andrew-pengjj/RCILD.git. Jiangjun Peng, Hailin Wang 0001, Xiangyong Cao, Qian Zhao 0002, Jing Yao 0002, Hong-Ying Zhang 0001, Deyu Meng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Low-Light Image Enhancement by Retinex-Based Algorithm Unrolling and AdjustmentabstractLow-light image enhancement (LIE) has attracted tremendous research interests in recent years. Retinex theory-based deep learning methods, following a decomposition-adjustment pipeline, have achieved promising performance due to their physical interpretability. However, existing Retinex-based deep learning methods are still suboptimal, failing to leverage useful insights from traditional approaches. Meanwhile, the adjustment step is either oversimplified or overcomplicated, resulting in unsatisfactory performance in practice. To address these issues, we propose a novel deep-learning framework for LIE. The framework consists of a decomposition network (DecNet) inspired by algorithm unrolling and adjustment networks considering both global and local brightness. The algorithm unrolling allows the integration of both implicit priors learned from data and explicit priors inherited from traditional methods, facilitating better decomposition. Meanwhile, considering global and local brightness guides the design of effective yet lightweight adjustment networks. Moreover, we introduce a self-supervised fine-tuning strategy that achieves promising performance without manual hyperparameter tuning. Extensive experiments on benchmark LIE datasets demonstrate the superiority of our approach over existing state-of-the-art methods both quantitatively and qualitatively. Code is available at https://github.com/Xinyil256/RAUNA2023. Qi Xie 0002, Qian Zhao 0002, Hong Wang 0021, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | RCDNet: An Interpretable Rain Convolutional Dictionary Network for Single Image DerainingabstractAs common weather, rain streaks adversely degrade the image quality and tend to negatively affect the performance of outdoor computer vision systems. Hence, removing rains from an image has become an important issue in the field. To handle such an ill-posed single image deraining task, in this article, we specifically build a novel deep architecture, called rain convolutional dictionary network (RCDNet), which embeds the intrinsic priors of rain streaks and has clear interpretability. In specific, we first establish a rain convolutional dictionary (RCD) model for representing rain streaks and utilize the proximal gradient descent technique to design an iterative algorithm only containing simple operators for solving the model. By unfolding it, we then build the RCDNet in which every network module has clear physical meanings and corresponds to each operation involved in the algorithm. This good interpretability greatly facilitates an easy visualization and analysis of what happens inside the network and why it works well in the inference process. Moreover, taking into account the domain gap issue in real scenarios, we further design a novel dynamic RCDNet, where the rain kernels can be dynamically inferred corresponding to input rainy images and then help shrink the space for rain layer estimation with few rain maps, so as to ensure a fine generalization performance in the inconsistent scenarios of rain types between training and testing data. By end-to-end training such an interpretable network, all involved rain kernels and proximal operators can be automatically extracted, faithfully characterizing the features of both rain and clean background layers and, thus, naturally leading to better deraining performance. Comprehensive experiments implemented on a series of representative synthetic and real datasets substantiate the superiority of our method, especially on its well generality to diverse testing scenarios and good interpretability for all its modules, compared with state-of-the-art single image derainers both visually and quantitatively. Code is available at https://github.com/hongwang01/DRCDNet. Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Yuexiang Li, Yong Liang 0001, Yefeng Zheng 0001, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Interactive Segmentation as Gaussian Process ClassificationabstractClick-based interactive segmentation (IS) aims to extract the target objects under user interaction. For this task, most of the current deep learning (DL)-based methods mainly follow the general pipelines of semantic segmentation. Albeit achieving promising performance, they do not fully and explicitly utilize and propagate the click information, inevitably leading to unsatisfactory segmentation results, even at clicked points. Against this issue, in this paper, we propose to formulate the IS task as a Gaussian process (GP)-based pixel-wise binary classification model on each image. To solve this model, we utilize amortized variational inference to approximate the intractable GP posterior in a data-driven manner and then decouple the approximated GP posterior into double space forms for efficient sampling with linear complexity. Then, we correspondingly construct a GP classification framework, named GPCIS, which is integrated with the deep kernel learning mechanism for more flexibility. The main specificities of the proposed GPCIS lie in: 1) Under the explicit guidance of the derived GP posterior, the information contained in clicks can be finely propagated to the entire image and then boost the segmentation; 2) The accuracy of predictions at clicks has good theoretical support. These merits of GPCIS as well as its good generality and high efficiency are substantiated by comprehensive experiments on several benchmarks, as compared with representative methods both quantitatively and qualitatively. Codes will be released at https://github.com/zmhhlnz/GPCIS_CVPR2023. Hong Wang 0021, Qian Zhao 0002, Yuexiang Li, Yawen Huang, Deyu Meng, Yefeng Zheng 0001 |
CVPR | 3 |
| 2023 | Robust channel estimation based on the maximum entropy principle
Zhengyang Hu 0001, Jiang Xue 0001, Feng Li 0057, Qian Zhao 0002, Deyu Meng, Zongben Xu |
Sci. China Inf. Sci. | 4 |
| 2023 | Stein variational gradient descent with learned direction
Qian Zhao 0002, Hui Wang 0103, Xuehu Zhu, Deyu Meng |
Inf. Sci. | 1 |
| 2023 | MLR-SNet: Transferable LR Schedules for Heterogeneous TasksabstractThe learning rate (LR) is one of the most important hyperparameters in stochastic gradient descent (SGD) algorithm for training deep neural networks (DNN). However, current hand-designed LR schedules need to manually pre-specify a fixed form, which limits their ability to adapt to practical non-convex optimization problems due to the significant diversification of training dynamics. Meanwhile, it always needs to search proper LR schedules from scratch for new tasks, which, however, are often largely different with task variations, like data modalities, network architectures, or training data capacities. To address this learning-rate-schedule setting issue, we propose to parameterize LR schedules with an explicit mapping formulation, called MLR-SNet. The learnable parameterized structure brings more flexibility for MLR-SNet to learn a proper LR schedule to comply with the training dynamics of DNN. Image and text classification benchmark experiments substantiate the capability of our method for achieving proper LR schedules. Moreover, the explicit parameterized structure makes the meta-learned LR schedules capable of being transferable and plug-and-play, which can be easily generalized to new heterogeneous tasks. We transfer our meta-learned MLR-SNet to query tasks like different training epochs, network architectures, data modalities, dataset sizes from the training ones, and achieve comparable or even better performance compared with hand-designed LR schedules specifically designed for the query tasks. The robustness of MLR-SNet is also substantiated when the training data are biased with corrupted noise. We further prove the convergence of the SGD algorithm equipped with LR schedule produced by our MLR-SNet, with the convergence rate comparable to the best-known ones of the algorithm for solving the problem. The source code of our method is released at https://github.com/xjtushujun/MLR-SNet. Yanwen Zhu, Qian Zhao 0002, Deyu Meng, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Fourier Series Expansion Based Filter Parametrization for Equivariant ConvolutionsabstractIt has been shown that equivariant convolution is very helpful for many types of computer vision tasks. Recently, the 2D filter parametrization technique has played an important role for designing equivariant convolutions, and has achieved success in making use of rotation symmetry of images. However, the current filter parametrization strategy still has its evident drawbacks, where the most critical one lies in the accuracy problem of filter representation. To address this issue, in this paper we explore an ameliorated Fourier series expansion for 2D filters, and propose a new filter parametrization method based on it. The proposed filter parametrization method not only finely represents 2D filters with zero error when the filter is not rotated (similar as the classical Fourier series expansion), but also substantially alleviates the aliasing-effect-caused quality degradation when the filter is rotated (which usually arises in classical Fourier series expansion method). Accordingly, we construct a new equivariant convolution method based on the proposed filter parametrization method, named F-Conv. We prove that the equivariance of the proposed F-Conv is exact in the continuous domain, which becomes approximate only after discretization. Moreover, we provide theoretical error analysis for the case when the equivariance is approximate, showing that the approximation error is related to the mesh size and filter size. Extensive experiments show the superiority of the proposed method. Particularly, we adopt rotation equivariant convolution methods to a typical low-level image processing task, image super-resolution. It can be substantiated that the proposed F-Conv based method evidently outperforms classical convolution based methods. Compared with pervious filter parametrization based methods, the F-Conv performs more accurately on this low-level image processing task, reflecting its intrinsic capability of faithfully preserving rotation symmetries in local image features. Qi Xie 0002, Qian Zhao 0002, Zongben Xu, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | A Tensor-Based Online RPCA Model for Compressive Background SubtractionabstractBackground subtraction of videos has been a fundamental research topic in computer vision in the past decades. To alleviate the computation burden and enhance the efficiency, background subtraction from online compressive measurements has recently attracted much attention. However, current methods still have limitations. First, they are all based on matrix modeling, which breaks the spatial structure within video frames. Second, they generally ignore the complex disturbance within the background, which reduces the efficiency of the low-rank assumption. To alleviate this issue, we propose a tensor-based online compressive video reconstruction and background subtraction method, abbreviated as NIOTenRPCA, by explicitly modeling the background disturbance in different frames as nonidentical but correlated noise. By virtue of such sophisticated modeling, the proposed method can well adapt to complex video scenes and, thus, perform more robustly. Extensive experiments on a series of real-world video datasets have demonstrated the effectiveness of the proposed method compared with the existing state of the arts. The code of our method is released on the website: https://github.com/crystalzina/NIOTenRPCA. Zina Li, Yao Wang 0003, Qian Zhao 0002, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | A Probabilistic Formulation for Meta-Weight-NetabstractIn the last decade, deep neural networks (DNNs) have become dominant tools for various of supervised learning tasks, especially classification. However, it is demonstrated that they can easily overfit to training set biases, such as label noise and class imbalance. Example reweighting algorithms are simple and effective solutions against this issue, but most of them require manually specifying the weighting functions as well as additional hyperparameters. Recently, a meta-learning-based method Meta-Weight-Net (MW-Net) has been proposed to automatically learn the weighting function parameterized by an MLP via additional unbiased metadata, which significantly improves the robustness of prior arts. The method, however, is proposed in a deterministic manner, and short of intrinsic statistical support. In this work, we propose a probabilistic formulation for MW-Net, probabilistic MW-Net (PMW-Net) in short, which treats the weighting function in a probabilistic way, and can include the original MW-Net as a special case. By this probabilistic formulation, additional randomness is introduced while the flexibility of the weighting function can be further controlled during learning. Our experimental results on both synthetic and real datasets show that the proposed method improves the performance of the original MW-Net. Besides, the proposed PMW-Net can also be further extended to fully Bayesian models, to improve their robustness. Qian Zhao 0002, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Blind Image Super-resolution with Elaborate Degradation Modeling on Noise and KernelabstractWhile researches on model-based blind single image super-resolution (SISR) have achieved tremendous successes recently, most of them do not consider the image degradation sufficiently. Firstly, they always assume image noise obeys an independent and identically distributed (i.i.d.) Gaussian or Laplacian distribution, which largely underestimates the complexity of real noise. Secondly, previous commonly-used kernel priors (e.g., normalization, sparsity) are not effective enough to guarantee a rational kernel solution, and thus degenerates the performance of subsequent SISR task. To address the above issues, this paper proposes a model-based blind SISR method under the probabilistic framework, which elaborately models image degradation from the perspectives of noise and blur kernel. Specifically, instead of the traditional i.i.d. noise assumption, a patch-based non-i.i.d. noise model is proposed to tackle the complicated real noise, expecting to increase the degrees of freedom of the model for noise representation. As for the blur kernel, we novelly construct a concise yet effective kernel generator, and plug it into the proposed blind SISR method as an explicit kernel prior (EKP). To solve the proposed model, a theoretically grounded Monte Carlo EM algorithm is specifically designed. Comprehensive experiments demonstrate the superiority of our method over current state-of-the-arts on synthetic and real datasets. The source code is available at https://github.com/zsyOAOA/BSRDM. Zongsheng Yue, Qian Zhao 0002, Jianwen Xie, Lei Zhang 0006, Deyu Meng, Kwan-Yee Kenneth Wong |
CVPR | 2 |
| 2022 | KXNet: A Model-Driven Deep Neural Network for Blind Super-Resolution
Jiahong Fu, Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu |
ECCV (19) | 4 |
| 2022 | Survey on rain removal from videos or a single image
Hong Wang 0021, Minghan Li 0001, Qian Zhao 0002, Deyu Meng |
Sci. China Inf. Sci. | 4 |
| 2022 | A deep variational Bayesian framework for blind image deblurring
Qian Zhao 0002, Hui Wang 0103, Zongsheng Yue, Deyu Meng |
Knowl. Based Syst. | 1 |
| 2022 | MHF-Net: An Interpretable Deep Network for Multispectral and Hyperspectral Image FusionabstractMultispectral and hyperspectral image fusion (MS/HS fusion) aims to fuse a high-resolution multispectral (HrMS) and a low-resolution hyperspectral (LrHS) images to generate a high-resolution hyperspectral (HrHS) image, which has become one of the most commonly addressed problems for hyperspectral image processing. In this paper, we specifically designed a network architecture for the MS/HS fusion task, called MHF-net, which not only contains clear interpretability, but also reasonably embeds the well studied linear mapping that links the HrHS image to HrMS and LrHS images. In particular, we first construct an MS/HS fusion model which merges the generalization models of low-resolution images and the low-rankness prior knowledge of HrHS image into a concise formulation, and then we build the proposed network by unfolding the proximal gradient algorithm for solving the proposed model. As a result of the careful design for the model and algorithm, all the fundamental modules in MHF-net have clear physical meanings and are thus easily interpretable. This not only greatly facilitates an easy intuitive observation and analysis on what happens inside the network, but also leads to its good generalization capability. Based on the architecture of MHF-net, we further design two deep learning regimes for two general cases in practice: consistent MHF-net and blind MHF-net. The former is suitable in the case that spectral and spatial responses of training and testing data are consistent, just as considered in most of the pervious general supervised MS/HS fusion researches. The latter ensures a good generalization in mismatch cases of spectral and spatial responses in training and testing data, and even across different sensors, which is generally considered to be a challenging issue for general supervised MS/HS fusion methods. Experimental results on simulated and real data substantiate the superiority of our method both visually and quantitatively as compared with state-of-the-art methods along this line of research. Qi Xie 0002, Qian Zhao 0002, Zongben Xu, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Robust Online CSI Estimation in a Complex EnvironmentabstractChannel state information (CSI) estimation is one of the key techniques for improving the performance of wireless communication systems. Meanwhile, the fifth generation wireless communication systems require higher accuracy and lower latency for CSI estimation. In this paper, the methods of noise modeling and online learning are combined to improve the accuracy and reduce the latency. The complex noise environment (considering noise and interference together) is modeled as a specific mixture of Gaussian (MoG) distribution because of its widely approximation capability to any continuous distribution. The MoG CSI estimation (MoG-CE) model and expectation maximization (EM) algorithm are introduced as one of the baseline methods. Further, the parameters of the model can be updated in real time based on the prior knowledge of historical information. Therefore, the online MoG CSI estimation (O-MoG-CE) model and online MoG dynamic CSI estimation (O-MoG-D-CE) model are proposed for time-invariant and time-varying CSI estimations, respectively. The above models can not only self-adapt to various complex communication scenarios robustly but also achieve online and dynamic CSI estimation to improve the accuracy and reduce the latency significantly. In addition, the proposed models can be formulated as standard maximum a posteriori estimations and efficient online expectation maximization (OEM) algorithms are applied for the estimations in a pure machine learning fashion. Comparing with baseline methods, the simulation results demonstrate the superiority of the proposed methods in terms of the accuracy, latency and computation consumption. Jiang Xue 0001, Deyu Meng, Qian Zhao 0002, Zongben Xu |
IEEE Trans. Wirel. Commun. | 5 |
| 2021 | Learning to Purify Noisy Labels via Meta Soft Label CorrectorabstractRecent deep neural networks (DNNs) can easily overfit to biased training data with noisy labels. Label correction strategy is commonly used to alleviate this issue by identifying suspected noisy labels and then correcting them. Current approaches to correcting corrupted labels usually need manually pre-defined label correction rules, which makes it hard to apply in practice due to the large variations of such manual strategies with respect to different problems. To address this issue, we propose a meta-learning model, aiming at attaining an automatic scheme which can estimate soft labels through meta-gradient descent step under the guidance of a small amount of noise-free meta data. By viewing the label correction procedure as a meta-process and using a meta-learner to automatically correct labels, our method can adaptively obtain rectified soft labels gradually in iteration according to current training problems. Besides, our method is model-agnostic and can be combined with any other existing classification models with ease to make it available to noisy label cases. Comprehensive experiments substantiate the superiority of our method in both synthetic and real-world problems with noisy labels compared with current state-of-the-art label correction strategies. Qi Xie 0002, Qian Zhao 0002, Deyu Meng |
AAAI | 4 |
| 2021 | Learning an Explicit Weighting Scheme for Adapting Complex HSI NoiseabstractAn efficient approach for handling hyperspectral image (HSI) denoising issue is to impose weights on different HSI pixels to suppress negative influence brought by noisy elements. Such weighting scheme, however, largely depends on the prior understanding or subjective distribution assumption on HSI noises, making them easily biased to complicated real noises, and hardly generalizable to diverse practical scenarios. Against this issue, this paper proposes a new scheme aiming to capture general weighting principle in a data-driven manner. Specifically, such weighting principle is delivered by an explicit function, called hyper-weight-net (HWnet), mapping from an input noisy image to its properly imposed weights. A Bayesian framework as well as a variational inference algorithm for inferring HWnet parameters is elaborately designed, expecting to extract the latent weighting rule for general diverse and complicated noisy HSIs. Comprehensive experiments substantiate that the learned HWnet can be not only finely generalized to different noise types from those used in training, but also effectively transferred to other weighted models. Besides, as a sounder guidance, HWnet can help to more faithfully and robustly achieve deep hyperspectral prior(DHP). The extracted weights by HWnet are verified to be able to effectively capture complex noise knowledge underlying input HSI, revealing its working insight in experiments. Xiangyu Rui, Xiangyong Cao, Qi Xie 0002, Zongsheng Yue, Qian Zhao 0002, Deyu Meng |
CVPR | 5 |
| 2021 | From Rain Generation to Rain RemovalabstractFor the single image rain removal (SIRR) task, the performance of deep learning (DL)-based methods is mainly affected by the designed deraining models and training datasets. Most of current state-of-the-art focus on constructing powerful deep models to obtain better deraining results. In this paper, to further improve the deraining performance, we novelly attempt to handle the SIRR task from the perspective of training datasets by exploring a more efficient way to synthesize rainy images. Specifically, we build a full Bayesian generative model for rainy image where the rain layer is parameterized as a generator with the input as some latent variables representing the physical structural rain factors, e.g., direction, scale, and thickness. To solve this model, we employ the variational inference framework to approximate the expected statistical distribution of rainy image in a data-driven manner. With the learned generator, we can automatically and sufficiently generate diverse and non-repetitive training pairs so as to efficiently enrich and augment the existing benchmark datasets. User study qualitatively and quantitatively evaluates the realism of generated rainy images. Comprehensive experiments substantiate that the proposed model can faithfully extract the complex rain distribution that not only helps significantly improve the deraining performance of current deep single image derainers, but also largely loosens the requirement of large training sample pre-collection for the SIRR task. Code is available in https://github.com/hongwang01/VRGNet. Hong Wang 0021, Zongsheng Yue, Qi Xie 0002, Qian Zhao 0002, Yefeng Zheng 0001, Deyu Meng |
CVPR | 4 |
| 2021 | Semi-Supervised Video Deraining With Dynamical Rain GeneratorabstractWhile deep learning (DL)-based video deraining methods have achieved significant successes in recent years, they still have two major drawbacks. Firstly, most of them are insufficient to model the characteristics of rain layers contained in rainy videos. In fact, the rain layers exhibit strong visual properties (e.g., direction, scale, and thickness) in spatial dimension and causal properties (e.g., velocity and acceleration) in temporal dimension, and thus can be modeled by the spatial-temporal process in statistics. Secondly, current DL-based methods rely heavily on the labeled training data, whose rain layers are synthetic, thus leading to a deviation from real data. Such a gap between synthetic and real data sets results in poor performance when applying them to real scenarios. To address these issues, this paper proposes a new semi-supervised video deraining method, in which a dynamical rain generator is employed to fit the rain layer for the sake of better depicting its intrinsic characteristics. Specifically, the dynamical generator consists of one emission model and one transition model to simultaneously encode the spatial appearance and temporal dynamics of rain streaks, respectively, both of which are parameterized by deep neural networks (DNNs). Furthermore, different prior formats are designed for the labeled synthetic and unlabeled real data so as to fully exploit their underlying common knowledge. Last but not least, we design a Monte Carlo-based EM algorithm to learn the model. Extensive experiments are conducted to verify the superiority of the proposed semi-supervised deraining model. Zongsheng Yue, Jianwen Xie, Qian Zhao 0002, Deyu Meng |
CVPR | 3 |
| 2021 | Structural residual learning for single image rain removal
Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Yong Liang 0001, Deyu Meng |
Knowl. Based Syst. | 4 |
| 2021 | Robust online rain removal for surveillance videos with dynamic rains
Lixuan Yi, Qian Zhao 0002, Wei Wei 0006, Zongben Xu |
Knowl. Based Syst. | 2 |
| 2021 | SPLBoost: An Improved Robust Boosting Algorithm Based on Self-Paced LearningabstractIt is known that boosting can be interpreted as an optimization technique to minimize an underlying loss function. Specifically, the underlying loss being minimized by the traditional AdaBoost is the exponential loss, which proves to be very sensitive to random noise/outliers. Therefore, several boosting algorithms, e.g., LogitBoost and SavageBoost, have been proposed to improve the robustness of AdaBoost by replacing the exponential loss with some designed robust loss functions. In this article, we present a new way to robustify AdaBoost, that is, incorporating the robust learning idea of self-paced learning (SPL) into the boosting framework. Specifically, we design a new robust boosting algorithm based on the SPL regime, that is, SPLBoost, which can be easily implemented by slightly modifying off-the-shelf boosting packages. Extensive experiments and a theoretical characterization are also carried out to illustrate the merits of the proposed SPLBoost. Kaidong Wang, Yao Wang 0003, Qian Zhao 0002, Deyu Meng, Xiuwu Liao, Zongben Xu |
IEEE Trans. Cybern. | 3 |
| 2021 | Online Rain/Snow Removal From Surveillance VideosabstractVideo rain/snow removal from surveillance videos is an important task in the computer vision community since rain/snow existed in videos can severely degenerate the performance of many surveillance system. Various methods have been investigated extensively, but most only consider consistent rain/snow under stable background scenes. Rain/snow captured from practical surveillance camera, however, is always highly dynamic in time, and those videos also include occasionally transformed background scenes and background motions caused by waving leaves or water surfaces. To this issue, this paper proposes a novel rain/snow removal approach, which fully considers dynamic statistics of both rain/snow and background scenes taken from a video sequence. Specifically, the rain/snow is encoded as an online multi-scale convolutional sparse coding (OMS-CSC) model, which not only finely delivers the sparse scattering and multi-scale shapes of real rain/snow, but also well distinguish the components of background motion from rain/snow layer. The real-time ameliorated parameters in the model well encodes their temporally dynamic configurations. Furthermore, a transformation operator imposed on the background scenes is further embedded into the proposed model, which finely conveys the background transformations, such as rotations, scalings and distortions, inevitably existed in a real video sequence. The approach so constructed can naturally better adapt to the dynamic rain/snow as well as background changes, and also suitable to deal with the streaming video attributed its online learning mode. The proposed model is formulated in a concise maximum a posterior (MAP) framework and is readily solved by the alternating direction method of multipliers (ADMM). Compared with the state-of-the-art online and offline video rain/snow removal methods, the proposed method achieves best performance on synthetic and real videos datasets both visually and quantitatively. Specifically, our method can be implemented in relatively high efficiency, showing its potential to real-time video rain/snow removal. The code page is at: https://github.com/MinghanLi/OTMSCSC_matlab_2020. Minghan Li 0001, Xiangyong Cao, Qian Zhao 0002, Lei Zhang 0006, Deyu Meng |
IEEE Trans. Image Process. | 3 |
| 2020 | A Model-Driven Deep Neural Network for Single Image Rain RemovalabstractDeep learning (DL) methods have achieved state-of-the-art performance in the task of single image rain removal. Most of current DL architectures, however, are still lack of sufficient interpretability and not fully integrated with physical structures inside general rain streaks. To this issue, in this paper, we propose a model-driven deep neural network for the task, with fully interpretable network structures. Specifically, based on the convolutional dictionary learning mechanism for representing rain, we propose a novel single image deraining model and utilize the proximal gradient descent technique to design an iterative algorithm only containing simple operators for solving the model. Such a simple implementation scheme facilitates us to unfold it into a new deep network architecture, called rain convolutional dictionary network (RCDNet), with almost every network module one-to-one corresponding to each operation involved in the algorithm. By end-to-end training the proposed RCDNet, all the rain kernels and proximal operators can be automatically extracted, faithfully characterizing the features of both rain and clean background layers, and thus naturally lead to its better deraining performance, especially in real scenarios. Comprehensive experiments substantiate the superiority of the proposed network, especially its well generality to diverse testing scenarios and good interpretability for all its modules, as compared with state-of-the-arts both visually and quantitatively. Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Deyu Meng |
CVPR | 3 |
| 2020 | Dual Adversarial Network: Toward Real-World Noise Removal and Noise Generation
Zongsheng Yue, Qian Zhao 0002, Lei Zhang 0006, Deyu Meng |
ECCV (10) | 2 |
| 2020 | MEP-Based Channel Estimation under Complex Communication EnvironmentabstractIn this paper, we study the channel state information (CSI) estimation by utilizing maximum entropy principle (MEP) and noise modeling method. The new model can not only represent the characters of the complex communication environment, but can also adjust itself according to the environment by using machine learning. In addition, a new iteration algorithm is presented to derive numerical results. Adaptive parameters learning and features choosing capability make the proposed method outperform the existing methods. The accuracy of estimation is verified by the Monte Carlo simulations. Zhengyang Hu 0001, Jiang Xue 0001, Deyu Meng, Qian Zhao 0002, Zongben Xu |
ICC | 4 |
| 2020 | Color and direction-invariant nonlocal self-similarity prior and its application to color image denoising
Qi Xie 0002, Qian Zhao 0002, Zongben Xu, Deyu Meng |
Sci. China Inf. Sci. | 2 |
| 2020 | Discovering influential factors in variational autoencoders
Shiqi Liu 0001, Qian Zhao 0002, Xiangyong Cao, Huibin Li 0001, Deyu Meng, Hongying Meng, Sheng Liu 0033 |
Pattern Recognit. | 3 |
| 2020 | Enhanced 3DTV Regularization and Its Applications on HSI Denoising and Compressed SensingabstractThe total variation (TV) is a powerful regularization term encoding the local smoothness prior structure underlying images. By combining the TV regularization term with low rank prior, the 3D total variation (3DTV) regularizer has achieved advanced performance in general hyperspectral image (HSI) processing tasks. Intrinsically, 3DTV assumes i.i.d. sparsity structures on all bands of the gradient maps calculated along the spectrum and space of an HSI. This, however, largely deviates from the real-world cases, where the gradient maps generally have different while correlated gradient map structures across all bands. To alleviate this issue, we propose an enhanced 3DTV (E-3DTV) regularization term beyond the conventional. Instead of imposing sparsity on gradient maps themselves, the new term calculates sparsity on the subspace bases on gradient maps along all bands of an HSI, which naturally encodes the correlation and difference among all these bands, and thus more faithfully reflects the insightful configurations of an HSI. The E-3DTV term can easily replace the conventional 3DTV term and be embedded into an HSI processing model to ameliorate its performance. We made such attempts on two typical related tasks: HSI denoising and compressed sensing. The superiority of our proposed method is substantiated by extensive experiments on synthetic and real HSI data, visually and quantitatively on both tasks, as compared with current state-of-the-arts. The code of our algorithm is released athttps://github.com/andrew-pengjj/Enhanced-3DTV.git. Jiangjun Peng, Qi Xie 0002, Qian Zhao 0002, Yao Wang 0003, Yee Leung, Deyu Meng |
IEEE Trans. Image Process. | 3 |
| 2020 | Full-Spectrum-Knowledge-Aware Tensor Model for Energy-Resolved CT Iterative ReconstructionabstractEnergy-resolved computed tomography (ErCT) with a photon counting detector concurrently produces multiple CT images corresponding to different photon energy ranges. It has the potential to generate energy-dependent images with improved contrast-to-noise ratio and sufficient material-specific information. Since the number of detected photons in one energy bin in ErCT is smaller than that in conventional energy-integrating CT (EiCT), ErCT images are inherently more noisy than EiCT images, which leads to increased noise and bias in the subsequent material estimation. In this work, we first deeply analyze the intrinsic tensor properties of two-dimensional (2D) ErCT images acquired in different energy bins and then present a F ull- S pectrum-knowledge-aware Tensor analysis and processing (FSTensor) method for ErCT reconstruction to suppress noise-induced artifacts to obtain high-quality ErCT images and high-accuracy material images. The presented method is based on three considerations: (1) 2D ErCT images obtained in different energy bins can be treated as a 3-order tensor with three modes, i.e., width, height and energy bin, and a rich global correlation exists among the three modes, which can be characterized by tensor decomposition. (2) There is a locally piecewise smooth property in the 3-order ErCT images, and it can be captured by a tensor total variation regularization. (3) The images from the full spectrum are much better than the ErCT images with respect to noise variance and structural details and serve as external information to improve the reconstruction performance. We then develop an alternating direction method of multipliers algorithm to numerically solve the presented FSTensor method. We further utilize a genetic algorithm to tackle the parameter selection in ErCT reconstruction, instead of manually determining parameters. Simulation, preclinical and synthesized clinical ErCT results demonstrate that the presented FSTensor method leads to significant improvements over the filtered back-projection, robust principal component analysis, tensor-based dictionary learning and low-rank tensor decomposition with spatial-temporal total variation methods. Dong Zeng, Yongshuai Ge, Sui Li, Qi Xie 0002, Hao Zhang 0026, Zhaoying Bian, Qian Zhao 0002, Yuanqing Li 0001, Zongben Xu, Deyu Meng, Jianhua Ma 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Robust Multiview Subspace Learning With Nonindependently and Nonidentically Distributed Complex NoiseabstractMultiview Subspace Learning (MSL), which aims at obtaining a low-dimensional latent subspace from multiview data, has been widely used in practical applications. Most recent MSL approaches, however, only assume a simple independent identically distributed (i.i.d.) Gaussian or Laplacian noise for all views of data, which largely underestimates the noise complexity in practical multiview data. Actually, in real cases, noises among different views generally have three specific characteristics. First, in each view, the data noise always has a complex configuration beyond a simple Gaussian or Laplacian distribution. Second, the noise distributions of different views of data are generally nonidentical and with evident distinctiveness. Third, noises among all views are nonindependent but obviously correlated. Based on such understandings, we elaborately construct a new MSL model by more faithfully and comprehensively considering all these noise characteristics. First, the noise in each view is modeled as a Dirichlet process (DP) Gaussian mixture model (DPGMM), which can fit a wider range of complex noise types than conventional Gaussian or Laplacian. Second, the DPGMM parameters in each view are different from one another, which encodes the "nonidentical" noise property. Third, the DPGMMs on all views share the same high-level priors by using the technique of hierarchical DP, which encodes the "nonindependent" noise property. All the aforementioned ideas are incorporated into an integrated graphics model which can be appropriately solved by the variational Bayes algorithm. The superiority of the proposed method is verified by experiments on 3-D reconstruction simulations, multiview face modeling, and background subtraction, as compared with the current state-of-the-art MSL methods. Zongsheng Yue, Hongwei Yong, Deyu Meng, Qian Zhao 0002, Yee Leung, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Multilinear Multitask Learning by Rank-Product RegularizationabstractMultilinear multitask learning (MLMTL) considers an MTL problem in which tasks are arranged by multiple indices. By exploiting the higher order correlations among the tasks, MLMTL is expected to improve the performance of traditional MTL, which only considers the first-order correlation across all tasks, e.g., low-rank structure of the coefficient matrix. The key to MLMTL is designing a rational regularization term to represent the latent correlation structure underlying the coefficient tensor instead of matrix. In this paper, we propose a new MLMTL model by employing the rank-product regularization term in the objective, which on one hand can automatically rectify the weights along all its tensor modes and on the other hand have an explicit physical meaning. By using this regularization, the intrinsic high-order correlations among tasks can be more precisely described, and thus, the overall performance of all tasks can be improved. To solve the resulted optimization model, we design an efficient algorithm by applying the alternating direction method of multipliers (ADMM). We also analyze the convergence and show that the proposed algorithm, with certain restriction, is asymptotically regular. Experiments on both synthetic and real data sets substantiate the superiority of the proposed method beyond the existing MLMTL methods in terms of accuracy and efficiency. Qian Zhao 0002, Xiangyu Rui, Zhi Han, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Semi-Supervised Transfer Learning for Image Rain RemovalabstractSingle image rain removal is a typical inverse problem in computer vision. The deep learning technique has been verified to be effective for this task and achieved state-of-the-art performance. However, previous deep learning methods need to pre-collect a large set of image pairs with/without synthesized rain for training, which tends to make the neural network be biased toward learning the specific patterns of the synthesized rain, while be less able to generalize to real test samples whose rain types differ from those in the training data. To this issue, this paper firstly proposes a semi-supervised learning paradigm toward this task. Different from traditional deep learning methods which only use supervised image pairs with/without synthesized rains, we further put real rainy images, without need of their clean ones, into the network training process. This is realized by elaborately formulating the residual between an input rainy image and its expected network output (clear image without rain) as a concise mixture of Gaussians distribution. The network is therefore trained to transfer to adapting the real rain pattern domain instead of only the synthesis rain domain, and thus both the short-of-training-sample and bias-to-supervised-sample issues can be evidently alleviated. Experiments on synthetic and real data verify the superiority of our model compared to the state-of-the-arts. Wei Wei 0006, Deyu Meng, Qian Zhao 0002, Zongben Xu |
CVPR | 3 |
| 2019 | Multispectral and Hyperspectral Image Fusion by MS/HS Fusion NetabstractHyperspectral imaging can help better understand the characteristics of different materials, compared with traditional image systems. However, only high-resolution multispectral (HrMS) and low-resolution hyperspectral (LrHS) images can generally be captured at video rate in practice. In this paper, we propose a model-based deep learning approach for merging an HrMS and LrHS images to generate a high-resolution hyperspectral (HrHS) image. In specific, we construct a novel MS/HS fusion model which takes the observation models of low-resolution images and the low-rankness knowledge along the spectral mode of HrHS image into consideration. Then we design an iterative algorithm to solve the model by exploiting the proximal gradient method. And then, by unfolding the designed algorithm, we construct a deep network, called MS/HS Fusion Net, with learning the proximal operators and model parameters by convolutional neural networks. Experimental results on simulated and real data substantiate the superiority of our method both visually and quantitatively as compared with state-of-the-art methods along this line of research. Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Wangmeng Zuo, Zongben Xu |
CVPR | 3 |
| 2019 | Robust CSI Estimation Under Complex Communication EnvironmentabstractChannel estimation is the critical and fundamental problem in wireless communication techniques, however, the complexity environment, including interference and noise, post a fundamental limit on the accuracy of channel estimation on practical applications. Most existing channel estimation techniques are based on the simple assumption of Gaussian white noise, which makes the performance poorly within real communication environment. To address this problem, we propose a new channel estimation method by assuming the environment as Mixture of Gaussian (MoG) distributions and penalized MoG (PMoG) model by combining the penalized likelihood method with MoG distributions. This model is proposed by the first time in the research of wireless communication, and the superiority of this method lies on its approximation capability to wide range of scenarios of complex communication environments adaptively and analyzing the environment by learning the proper number of statistical components. Moreover, we design an Expectation Maximization (EM) algorithm to estimate the parameters of the PMoG model. The advantage of our method is demonstrated by simulation experiments. Haipei Zhang, Jiang Xue 0001, Deyu Meng, Qian Zhao 0002, Zongben Xu |
ICC | 4 |
| 2019 | Meta-Weight-Net: Learning an Explicit Mapping For Sample WeightingabstractCurrent deep neural networks(DNNs) can easily overfit to biased training data with corrupted labels or class imbalance. Sample re-weighting strategy is commonly used to alleviate this issue by designing a weighting function mapping from training loss to sample weight, and then iterating between weight recalculating and classifier updating. Current approaches, however, need manually pre-specify the weighting function as well as its additional hyper-parameters. It makes them fairly hard to be generally applied in practice due to the significant variation of proper weighting schemes relying on the investigated problem and training data. To address this issue, we propose a method capable of adaptively learning an explicit weighting function directly from data. The weighting function is an MLP with one hidden layer, constituting a universal approximator to almost any continuous functions, making the method able to fit a wide range of weighting function forms including those assumed in conventional research. Guided by a small amount of unbiased meta-data, the parameters of the weighting function can be finely updated simultaneously with the learning process of the classifiers. Synthetic and real experiments substantiate the capability of our method for achieving proper weighting functions in class imbalance and noisy label cases, fully complying with the common settings in traditional methods, and more complicated scenarios beyond conventional cases. This naturally leads to its better accuracy than other state-of-the-art methods. Qi Xie 0002, Lixuan Yi, Qian Zhao 0002, Sanping Zhou, Zongben Xu, Deyu Meng |
NeurIPS | 4 |
| 2019 | Variational Denoising Network: Toward Blind Noise Modeling and RemovalabstractBlind image denoising is an important yet very challenging problem in computer vision due to the complicated acquisition process of real images. In this work we propose a new variational inference method, which integrates both noise estimation and image denoising into a unique Bayesian framework, for blind image denoising. Specifically, an approximate posterior, parameterized by deep neural networks, is presented by taking the intrinsic clean image and noise variances as latent variables conditioned on the input noisy image. This posterior provides explicit parametric forms for all its involved hyper-parameters, and thus can be easily implemented for blind image denoising with automatic noise estimation for the test noisy image. On one hand, as other data-driven deep learning methods, our method, namely variational denoising network (VDN), can perform denoising efficiently due to its explicit form of posterior expression. On the other hand, VDN inherits the advantages of traditional model-driven approaches, especially the good generalization capability of generative models. VDN has good interpretability and can be flexibly utilized to estimate and remove complicated non-i.i.d. noise collected in real scenarios. Comprehensive experiments are performed to substantiate the superiority of our method in blind image denoising. Zongsheng Yue, Hongwei Yong, Qian Zhao 0002, Deyu Meng, Lei Zhang 0006 |
NeurIPS | 3 |
| 2019 | Nonconvex-Sparsity and Nonlocal-Smoothness-Based Blind Hyperspectral UnmixingabstractBlind hyperspectral unmixing (HU), as a crucial technique for hyperspectral data exploitation, aims to decompose mixed pixels into a collection of constituent materials weighted by the corresponding fractional abundances. In recent years, nonnegative matrix factorization (NMF) based methods have become more and more popular for this task and achieved promising performance. Among these methods, two types of properties upon the abundances, namely the sparseness and the structural smoothness, have been explored and shown to be important for blind HU. However, all of previous methods ignores another important insightful property possessed by a natural hyperspectral images (HSI), non-local smoothness, which means that similar patches in a larger region of an HSI are sharing the similar smoothness structure. Based on previous attempts on other tasks, such a prior structure reflects intrinsic configurations underlying a HSI, and is thus expected to largely improve the performance of the investigated HU problem. In this paper, we firstly consider such prior in HSI by encoding it as the nonlocal total variation (NLTV) regularizer. Furthermore, by fully exploring the intrinsic structure of HSI, we generalize NLTV to non-local HSI TV (NLHTV) to make the model more suitable for the bind HU task. By incorporating these two regularizers, together with a non-convex log-sum form regularizer characterizing the sparseness of abundance maps, to the NMF model, we propose novel blind HU models named NLTV/NLHTV and log-sum regularized NMF (NLTV-LSRNMF/NLHTV-LSRNMF), respectively. To solve the proposed models, an efficient algorithm is designed based on alternative optimization strategy (AOS) and alternating direction method of multipliers (ADMM). Extensive experiments conducted on both simulated and real hyperspectral data sets substantiate the superiority of the proposed approach over other competing ones for blind HU task. Jing Yao 0002, Deyu Meng, Qian Zhao 0002, Wenfei Cao, Zongben Xu |
IEEE Trans. Image Process. | 3 |
| 2018 | Video Rain Streak Removal by Multiscale Convolutional Sparse CodingabstractVideos captured by outdoor surveillance equipments sometimes contain unexpected rain streaks, which brings difficulty in subsequent video processing tasks. Rain streak removal from a video is thus an important topic in recent computer vision research. In this paper, we raise two intrinsic characteristics specifically possessed by rain streaks. Firstly, the rain streaks in a video contain repetitive local patterns sparsely scattered over different positions of the video. Secondly, the rain streaks are with multiscale configurations due to their occurrence on positions with different distances to the cameras. Based on such understanding, we specifically formulate both characteristics into a multiscale convolutional sparse coding (MS-CSC) model for the video rain streak removal task. Specifically, we use multiple convolutional filters convolved on the sparse feature maps to deliver the former characteristic, and further use multiscale filters to represent different scales of rain streaks. Such a new encoding manner makes the proposed method capable of properly extracting rain streaks from videos, thus getting fine video deraining effects. Experiments implemented on synthetic and real videos verify the superiority of the proposed method, as compared with the state-of-the-art ones along this research line, both visually and quantitatively. Minghan Li 0001, Qi Xie 0002, Qian Zhao 0002, Wei Wei 0006, Shuhang Gu, Deyu Meng |
CVPR | 3 |
| 2018 | Robust subspace clustering via penalized mixture of Gaussians
Jing Yao 0002, Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Zongben Xu |
Neurocomputing | 3 |
| 2018 | Kronecker-Basis-Representation Based Tensor Sparsity and Its Applications to Tensor RecoveryabstractAs a promising way for analyzing data, sparse modeling has achieved great success throughout science and engineering. It is well known that the sparsity/low-rank of a vector/matrix can be rationally measured by nonzero-entries-number ( norm)/nonzero- singular-values-number (rank), respectively. However, data from real applications are often generated by the interaction of multiple factors, which obviously cannot be sufficiently represented by a vector/matrix, while a high order tensor is expected to provide more faithful representation to deliver the intrinsic structure underlying such data ensembles. Unlike the vector/matrix case, constructing a rational high order sparsity measure for tensor is a relatively harder task. To this aim, in this paper we propose a measure for tensor sparsity, called Kronecker-basis-representation based tensor sparsity measure (KBR briefly), which encodes both sparsity insights delivered by Tucker and CANDECOMP/PARAFAC (CP) low-rank decompositions for a general tensor. Then we study the KBR regularization minimization (KBRM) problem, and design an effective ADMM algorithm for solving it, where each involved parameter can be updated with closed-form equations. Such an efficient solver makes it possible to extend KBR to various tasks like tensor completion and tensor robust principal component analysis. A series of experiments, including multispectral image (MSI) denoising, MSI completion and background subtraction, substantiate the superiority of the proposed methods beyond state-of-the-arts. Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Infrared small-dim target detection based on Markov random field guided noise modeling
Chenqiang Gao, Yongxing Xiao, Qian Zhao 0002, Deyu Meng |
Pattern Recognit. | 4 |
| 2018 | Denoising Hyperspectral Image With Non-i.i.d. Noise StructureabstractHyperspectral image (HSI) denoising has been attracting much research attention in remote sensing area due to its importance in improving the HSI qualities. The existing HSI denoising methods mainly focus on specific spectral and spatial prior knowledge in HSIs, and share a common underlying assumption that the embedded noise in HSI is independent and identically distributed (i.i.d.). In real scenarios, however, the noise existed in a natural HSI is always with much more complicated non-i.i.d. statistical structures and the under-estimation to this noise complexity often tends to evidently degenerate the robustness of current methods. To alleviate this issue, this paper attempts the first effort to model the HSI noise using a non-i.i.d. mixture of Gaussians (NMoGs) noise assumption, which finely accords with the noise characteristics possessed by a natural HSI and thus is capable of adapting various practical noise shapes. Then we integrate such noise modeling strategy into the low-rank matrix factorization (LRMF) model and propose an NMoG-LRMF model in the Bayesian framework. A variational Bayes algorithm is then designed to infer the posterior of the proposed model. As substantiated by our experiments implemented on synthetic and real noisy HSIs, the proposed method performs more robust beyond the state-of-the-arts. Yang Chen 0057, Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Zongben Xu |
IEEE Trans. Cybern. | 3 |
| 2018 | A Generalized Model for Robust Tensor Factorization With Noise Modeling by Mixture of GaussiansabstractThe low-rank tensor factorization (LRTF) technique has received increasing attention in many computer vision applications. Compared with the traditional matrix factorization technique, it can better preserve the intrinsic structure information and thus has a better low-dimensional subspace recovery performance. Basically, the desired low-rank tensor is recovered by minimizing the least square loss between the input data and its factorized representation. Since the least square loss is most optimal when the noise follows a Gaussian distribution, -norm-based methods are designed to deal with outliers. Unfortunately, they may lose their effectiveness when dealing with real data, which are often contaminated by complex noise. In this paper, we consider integrating the noise modeling technique into a generalized weighted LRTF (GWLRTF) procedure. This procedure treats the original issue as an LRTF problem and models the noise using a mixture of Gaussians (MoG), a procedure called MoG GWLRTF. To extend the applicability of the model, two typical tensor factorization operations, i.e., CANDECOMP/PARAFAC factorization and Tucker factorization, are incorporated into the LRTF procedure. Its parameters are updated under the expectation-maximization framework. Extensive experiments indicate the respective advantages of these two versions of MoG GWLRTF in various applications and also demonstrate their effectiveness compared with other competing methods. Xiai Chen, Zhi Han, Yao Wang 0003, Qian Zhao 0002, Deyu Meng, Lin Lin 0007, Yandong Tang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Should We Encode Rain Streaks in Video as Deterministic or Stochastic?abstractVideos taken in the wild sometimes contain unexpected rain streaks, which brings difficulty in subsequent video processing tasks. Rain streak removal in a video (RSRV) is thus an important issue and has been attracting much attention in computer vision. Different from previous RSRV methods formulating rain streaks as a deterministic message, this work first encodes the rains in a stochastic manner, i.e., a patch-based mixture of Gaussians. Such modification makes the proposed model capable of finely adapting a wider range of rain variations instead of certain types of rain configurations as traditional. By integrating with the spatiotemporal smoothness configuration of moving objects and low-rank structure of background scene, we propose a concise model for RSRV, containing one likelihood term imposed on the rain streak layer and two prior terms on the moving object and background scene layers of the video. Experiments implemented on videos with synthetic and real rains verify the superiority of the proposed method, as compared with the state-of-the-art methods, both visually and quantitatively in various performance metrics. Wei Wei 0006, Lixuan Yi, Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu |
ICCV | 4 |
| 2017 | Integration of 3-dimensional discrete wavelet transform and Markov random field for hyperspectral image classification
Xiangyong Cao, Lin Xu 0001, Deyu Meng, Qian Zhao 0002, Zongben Xu |
Neurocomputing | 4 |
| 2017 | A theoretical understanding of self-paced learning
Deyu Meng, Qian Zhao 0002, Lu Jiang 0004 |
Inf. Sci. | 2 |
| 2017 | Compressive Sensing of Hyperspectral Images via Joint Tensor Tucker Decomposition and Weighted Total Variation RegularizationabstractIn this letter, we consider the problem of compressive sensing of hyperspectral images (HSIs). We propose a novel tensor-based approach by modeling the global spatial-spectral correlation and local smoothness properties hidden in HSIs. Specifically, we use the tensor Tucker decomposition to describe the global spatial-spectral correlation among all HSI bands, and a weighted 3-D total variation to characterize the local smooth structure in both spatial and spectral modes. We then design an efficient algorithm to solve the resulting optimization problem by using the alternating direction method of multipliers. Experimental results on several HSI data sets demonstrate improved reconstruction performance of the proposed approach, as compared with other competing approaches. Yao Wang 0003, Lin Lin 0007, Qian Zhao 0002, Tianwei Yue, Deyu Meng, Yee Leung |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Robust Low-Dose CT Sinogram Preprocessing via Exploiting Noise-Generating MechanismabstractComputed tomography (CT) image recovery from low-mAs acquisitions without adequate treatment is always severely degraded due to a number of physical factors. In this paper, we formulate the low-dose CT sinogram preprocessing as a standard maximum a posteriori (MAP) estimation, which takes full consideration of the statistical properties of the two intrinsic noise sources in low-dose CT, i.e., the X-ray photon statistics and the electronic noise background. In addition, instead of using a general image prior as found in the traditional sinogram recovery models, we design a new prior formulation to more rationally encode the piecewise-linear configurations underlying a sinogram than previously used ones, like the TV prior term. As compared with the previous methods, especially the MAP-based ones, both the likelihood/loss and prior/regularization terms in the proposed model are ameliorated in a more accurate manner and better comply with the statistical essence of the generation mechanism of a practical sinogram. We further construct an efficient alternating direction method of multipliers algorithm to solve the proposed MAP framework. Experiments on simulated and real low-dose CT data demonstrate the superiority of the proposed method according to both visual inspection and comprehensive quantitative performance evaluation. Qi Xie 0002, Dong Zeng, Qian Zhao 0002, Deyu Meng, Zongben Xu, Zhengrong Liang, Jianhua Ma 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2016 | Robust Tensor Factorization with Unknown NoiseabstractBecause of the limitations of matrix factorization, such as losing spatial structure information, the concept of tensor factorization has been applied for the recovery of a low dimensional subspace from high dimensional visual data. Generally, the recovery is achieved by minimizing the loss function between the observed data and the factorization representation. Under different assumptions of the noise distribution, the loss functions are in various forms, like L1 and L2 norms. However, real data are often corrupted by noise with an unknown distribution. Then any specific form of loss function for one specific kind of noise often fails to tackle such real data with unknown noise. In this paper, we propose a tensor factorization algorithm to model the noise as a Mixture of Gaussians (MoG). As MoG has the ability of universally approximating any hybrids of continuous distributions, our algorithm can effectively recover the low dimensional subspace from various forms of noisy observations. The parameters of MoG are estimated under the EM framework and through a new developed algorithm of weighted low-rank tensor factorization (WLRTF). The effectiveness of our algorithm are substantiated by extensive experiments on both of synthetic data and real image data. Xiai Chen, Zhi Han, Yao Wang 0003, Qian Zhao 0002, Deyu Meng, Yandong Tang |
CVPR | 4 |
| 2016 | Multispectral Images Denoising by Intrinsic Tensor Sparsity RegularizationabstractMultispectral images (MSI) can help deliver more faithful representation for real scenes than the traditional image system, and enhance the performance of many computer vision tasks. In real cases, however, an MSI is always corrupted by various noises. In this paper, we propose a new tensor-based denoising approach by fully considering two intrinsic characteristics underlying an MSI, i.e., the global correlation along spectrum (GCS) and nonlocal self-similarity across space (NSS). In specific, we construct a new tensor sparsity measure, called intrinsic tensor sparsity (ITS) measure, which encodes both sparsity insights delivered by the most typical Tucker and CANDECOMP/ PARAFAC (CP) low-rank decomposition for a general tensor. Then we build a new MSI denoising model by applying the proposed ITS measure on tensors formed by non-local similar patches within the MSI. The intrinsic GCS and NSS knowledge can then be efficiently explored under the regularization of this tensor sparsity measure to finely rectify the recovery of a MSI from its corruption. A series of experiments on simulated and real MSI denoising problems show that our method outperforms all state-of-the-arts under comprehensive quantitative performance measures. Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu, Shuhang Gu, Wangmeng Zuo, Lei Zhang 0006 |
CVPR | 2 |
| 2016 | Robust Low-Rank Matrix Factorization Under General Mixture Noise DistributionsabstractMany computer vision problems can be posed as learning a low-dimensional subspace from high-dimensional data. The low rank matrix factorization (LRMF) represents a commonly utilized subspace learning strategy. Most of the current LRMF techniques are constructed on the optimization problems using L1-norm and L2-norm losses, which mainly deal with the Laplace and Gaussian noises, respectively. To make LRMF capable of adapting more complex noise, this paper proposes a new LRMF model by assuming noise as mixture of exponential power (MoEP) distributions and then proposes a penalized MoEP (PMoEP) model by combining the penalized likelihood method with MoEP distributions. Such setting facilitates the learned LRMF model capable of automatically fitting the real noise through MoEP distributions. Each component in this mixture distribution is adapted from a series of preliminary superor sub-Gaussian candidates. Moreover, by facilitating the local continuity of noise components, we embed Markov random field into the PMoEP model and then propose the PMoEP-MRF model. A generalized expectation maximization (GEM) algorithm and a variational GEM algorithm are designed to infer all parameters involved in the proposed PMoEP and the PMoEPMRF model, respectively. The superiority of our methods is demonstrated by extensive experiments on synthetic data, face modeling, hyperspectral image denoising, and background subtraction. Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Yang Chen 0057, Zongben Xu |
IEEE Trans. Image Process. | 2 |
| 2015 | Self-Paced Curriculum LearningabstractCurriculum learning (CL) or self-paced learning (SPL) represents a recently proposed learning regime inspired by the learning process of humans and animals that gradually proceeds from easy to more complex samples in training. The two methods share a similar conceptual learning paradigm, but differ in specific learning schemes. In CL, the curriculum is predetermined by prior knowledge, and remain fixed thereafter. Therefore, this type of method heavily relies on the quality of prior knowledge while ignoring feedback about the learner. In SPL, the curriculum is dynamically determined to adjust to the learning pace of the leaner. However, SPL is unable to deal with prior knowledge, rendering it prone to overfitting. In this paper, we discover the missing link between CL and SPL, and propose a unified framework named self-paced curriculum leaning (SPCL). SPCL is formulated as a concise optimization problem that takes into account both prior knowledge known before training and the learning progress during training. In comparison to human education, SPCL is analogous to "instructor-student-collaborative" learning mode, as opposed to "instructor-driven" in CL or "student-driven" in SPL. Empirically, we show that the advantage of SPCL on two tasks. Lu Jiang 0004, Deyu Meng, Qian Zhao 0002, Shiguang Shan, Alex Hauptmann 0001 |
AAAI | 3 |
| 2015 | Self-Paced Learning for Matrix FactorizationabstractMatrix factorization (MF) has been attracting much attention due to its wide applications. However, since MF models are generally non-convex, most of the existing methods are easily stuck into bad local minima, especially in the presence of outliers and missing data. To alleviate this deficiency, in this study we present a new MF learning methodology by gradually including matrix elements into MF training from easy to complex. This corresponds to a recently proposed learning fashion called self-paced learning (SPL), which has been demonstrated to be beneficial in avoiding bad local minima. We also generalize the conventional binary (hard) weighting scheme for SPL to a more effective real-valued (soft) weighting manner. The effectiveness of the proposed self-paced MF method is substantiated by a series of experiments on synthetic, structure from motion and background subtraction data. Qian Zhao 0002, Deyu Meng, Lu Jiang 0004, Qi Xie 0002, Zongben Xu, Alex Hauptmann 0001 |
AAAI | 1 |
| 2015 | Low-Rank Matrix Factorization under General Mixture Noise DistributionsabstractMany computer vision problems can be posed as learning a low-dimensional subspace from high dimensional data. The low rank matrix factorization (LRMF) represents a commonly utilized subspace learning strategy. Most of the current LRMF techniques are constructed on the optimization problem using L_1 norm and L_2 norm, which mainly deal with Laplacian and Gaussian noise, respectively. To make LRMF capable of adapting more complex noise, this paper proposes a new LRMF model by assuming noise as Mixture of Exponential Power (MoEP) distributions and proposes a penalized MoEP model by combining the penalized likelihood method with MoEP distributions. Such setting facilitates the learned LRMF model capable of automatically fitting the real noise through MoEP distributions. Each component in this mixture is adapted from a series of preliminary super-or sub-Gaussian candidates. An Expectation Maximization (EM) algorithm is also designed to infer the parameters involved in the proposed PMoEP model. The advantage of our method is demonstrated by extensive experiments on synthetic data, face modeling and hyperspectral image restoration. Xiangyong Cao, Yang Chen 0057, Qian Zhao 0002, Deyu Meng, Yao Wang 0003, Zongben Xu |
ICCV | 3 |
| 2015 | A Self-Paced Multiple-Instance Learning Framework for Co-Saliency DetectionabstractAs an interesting and emerging topic, co-saliency detection aims at simultaneously extracting common salient objects in a group of images. Traditional co-saliency detection approaches rely heavily on human knowledge for designing hand-crafted metrics to explore the intrinsic patterns underlying co-salient objects. Such strategies, however, always suffer from poor generalization capability to flexibly adapt various scenarios in real applications, especially due to their lack of insightful understanding of the biological mechanisms of human visual co-attention. To alleviate this problem, we propose a novel framework for this task, by naturally reformulating it as a multiple-instance learning (MIL) problem and further integrating it into a self-paced learning (SPL) regime. The proposed framework on one hand is capable of fitting insightful metric measurements and discovering common patterns under co-salient regions in a self-learning way by MIL, and on the other hand tends to promise the learning reliability and stability by simulating the human learning process through SPL. Experiments on benchmark datasets have demonstrated the effectiveness of the proposed framework as compared with the state-of-the-arts. Dingwen Zhang, Deyu Meng, Chao Li 0028, Lu Jiang 0004, Qian Zhao 0002, Junwei Han 0001 |
ICCV | 5 |
| 2015 | A Novel Sparsity Measure for Tensor RecoveryabstractIn this paper, we propose a new sparsity regularizer for measuring the low-rank structure underneath a tensor. The proposed sparsity measure has a natural physical meaning which is intrinsically the size of the fundamental Kronecker basis to express the tensor. By embedding the sparsity measure into the tensor completion and tensor robust PCA frameworks, we formulate new models to enhance their capability in tensor recovery. Through introducing relaxation forms of the proposed sparsity measure, we also adopt the alternating direction method of multipliers (ADMM) for solving the proposed models. Experiments implemented on synthetic and multispectral image data sets substantiate the effectiveness of the proposed methods. Qian Zhao 0002, Deyu Meng, Xu Kong, Qi Xie 0002, Wenfei Cao, Yao Wang 0003, Zongben Xu |
ICCV | 1 |
| 2015 | A block coordinate descent approach for sparse principal component analysis
Qian Zhao 0002, Deyu Meng, Zongben Xu, Chenqiang Gao |
Neurocomputing | 1 |
| 2015 | L1-Norm Low-Rank Matrix Factorization by Variational Bayesian MethodabstractThe L1 -norm low-rank matrix factorization (LRMF) has been attracting much attention due to its wide applications to computer vision and pattern recognition. In this paper, we construct a new hierarchical Bayesian generative model for the L1 -norm LRMF problem and design a mean-field variational method to automatically infer all the parameters involved in the model by closed-form equations. The variational Bayesian inference in the proposed method can be understood as solving a weighted LRMF problem with different weights on matrix elements based on their significance and with L2 -regularization penalties on parameters. Throughout the inference process of our method, the weights imposed on the matrix elements can be adaptively fitted so that the adverse influence of noises and outliers embedded in data can be largely suppressed, and the parameters can be appropriately regularized so that the generalization capability of the problem can be statistically guaranteed. The robustness and the efficiency of the proposed method are substantiated by a series of synthetic and real data experiments, as compared with the state-of-the-art L1 -norm LRMF methods. Especially, attributed to the intrinsic generalization capability of the Bayesian methodology, our method can always predict better on the unobserved ground truth data than existing methods. Qian Zhao 0002, Deyu Meng, Zongben Xu, Wangmeng Zuo, Yan Yan 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Robust Principal Component Analysis with Complex NoiseabstractThe research on robust principal component analysis (RPCA) has been attracting much attention recently. The original RPCA model assumes sparse noise, and use the L_1-norm to characterize the error term. In practice, however, the noise is much more complex and it is not appropriate to simply use a certain L_p-norm for noise modeling. We propose a generative RPCA model under the Bayesian framework by modeling data noise as a mixture of Gaussians (MoG). The MoG is a universal approximator to continuous distributions and thus our model is able to fit a wide range of noises such as Laplacian, Gaussian, sparse noises and any combinations of them. A variational Bayes algorithm is presented to infer the posterior of the proposed model. All involved parameters can be recursively updated in closed form. The advantage of our method is demonstrated by extensive experiments on synthetic data, face modeling and background subtraction. Qian Zhao 0002, Deyu Meng, Zongben Xu, Wangmeng Zuo, Lei Zhang 0006 |
ICML | 1 |
| 2014 | Robust sparse principal component analysis
Qian Zhao 0002, Deyu Meng, Zongben Xu |
Sci. China Inf. Sci. | 1 |
| 2013 | Learning dictionary from signals under global sparsity constraint
Deyu Meng, Qian Zhao 0002, Yee Leung, Zongben Xu |
Neurocomputing | 2 |
| 2012 | Improve robustness of sparse PCA by L1-norm maximization
Deyu Meng, Qian Zhao 0002, Zongben Xu |
Pattern Recognit. | 2 |