VLDB 2026 Research / reviewers in the wild / expert
Qi Xie 0002
dblp:45/2601-2
· DBLP profile ↗
44ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0003-0864-5623ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 6 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QMSANet: A quaternion multi-scale attention network for robust color image denoising
Qi Xie 0002, Yu Guo 0008, Boying Wu, Deyu Meng, Jean-Michel Morel, Qiyu Jin, Michael Kwok-Po Ng |
Neural Networks | 2 |
| 2026 | Multivariate neural directional total variation
Zelin Zeng, Guancheng Zhou, Yi-Si Luo, Xi-Le Zhao, Qi Xie 0002, Deyu Meng |
Pattern Recognit. | 5 |
| 2025 | A Regularization-Guided Equivariant Approach for Image RestorationabstractEquivariant and invariant deep learning models have been developed to exploit intrinsic symmetries in data, demonstrating significant effectiveness in certain scenarios. However, these methods often suffer from limited representation accuracy and rely on strict symmetry assumptions that may not hold in practice. These limitations pose a significant drawback for image restoration tasks, which demands high accuracy and precise symmetry representation. To address these challenges, we propose a rotation-equivariant regularization strategy that adaptively enforces the appropriate symmetry constraints on the data while preserving the network’s representational accuracy. Specifically, we introduce EQ-Reg, a regularizer designed to enhance rotation equivariance, which innovatively extends the insights of data-augmentation-based and equivariant-based methodologies. This is achieved through self-supervised learning and the spatial rotation and cyclic channel shift of feature maps deduce in the equivariant framework. Our approach firstly enables a non-strictly equivariant network suitable for image restoration, providing a simple and adaptive mechanism for adjusting equivariance based on task. Extensive experiments across three low-level tasks demonstrate the superior accuracy and generalization capability of our method, outperforming state-of-the-art approaches. Yulu Bai, Jiahong Fu, Qi Xie 0002, Deyu Meng |
CVPR | 3 |
| 2025 | Rotation-Equivariant Self-Supervised Method in Image DenoisingabstractSelf-supervised image denoising methods have garnered significant research attention in recent years, for this kind of method reduces the requirement of large training datasets. Compared to supervised methods, self-supervised methods rely more on the prior embedded in deep networks themselves. As a result, most of the self-supervised methods are designed with Convolution Neural Networks (CNNs) architectures, which well capture one of the most important image prior, translation equivariant prior. Inspired by the great success achieved by the introduction of translational equivariance, in this paper, we explore the way to further incorporate another important image prior. Specifically, we first apply high-accuracy rotation equivariant convolution to self-supervised image denoising. Through rigorous theoretical analysis, we have proved that simply replacing all the convolution layers with rotation equivariant convolution layers would modify the network into its rotation equivariant version. To the best of our knowledge, this is the first time that rotation equivariant image prior is introduced to self-supervised image denoising at the network architecture level with a comprehensive theoretical analysis of equivariance errors, which offers a new perspective to the field of self-supervised image denoising. Moreover, to further improve the performance, we design a new mask mechanism to fusion the output of rotation equivariant network and vanilla CNN-based network, and construct an adaptive rotation equivariant framework. Through extensive experiments on three typical methods, we have demonstrated the effectiveness of the proposed method. The code is available at: https://github.com/liuhanze623/AdaReNet. Hanze Liu, Jiahong Fu, Qi Xie 0002, Deyu Meng |
CVPR | 3 |
| 2025 | Online Functional Tensor Decomposition via Continual Learning for Streaming Data CompletionabstractOnline tensor decompositions are powerful and proven techniques that address the challenges in processing high-velocity streaming tensor data, such as traffic flow and weather system. The main aim of this work is to propose a novel online functional tensor decomposition (OFTD) framework, which represents a spatial-temporal continuous function using the CP tensor decomposition parameterized by coordinate-based implicit neural representations (INRs). The INRs allow for natural characterization of continually expanded streaming data by simply adding new coordinates into the network. Particularly, our method transforms the classical online tensor decomposition algorithm into a more dynamic continual learning paradigm of updating the INR weights to fit the new data without forgetting the previous tensor knowledge. To this end, we introduce a long-tail memory replay method that adapts to the local continuity property of INR. Extensive experiments for streaming tensor completion using traffic, weather, user-item, and video data verify the effectiveness of the OFTD approach for streaming data analysis. This endeavor serves as a pivotal inspiration for future research to connect classical online tensor tools with continual learning paradigms to better explore knowledge underlying streaming tensor data. Yanyi Li, Yi-Si Luo, Qi Xie 0002, Deyu Meng |
NeurIPS | 4 |
| 2025 | Polyline Path Masked Attention for Vision TransformerabstractGlobal dependency modeling and spatial position modeling are two core issues of the foundational architecture design in current deep learning frameworks. Recently, Vision Transformers (ViTs) have achieved remarkable success in computer vision, leveraging the powerful global dependency modeling capability of the self-attention mechanism. Furthermore, Mamba2 has demonstrated its significant potential in natural language processing tasks by explicitly modeling the spatial adjacency prior through the structured mask. In this paper, we propose Polyline Path Masked Attention (PPMA) that integrates the self-attention mechanism of ViTs with an enhanced structured mask of Mamba2, harnessing the complementary strengths of both architectures. Specifically, we first ameliorate the traditional structured mask of Mamba2 by introducing a 2D polyline path scanning strategy and derive its corresponding structured mask, polyline path mask, which better preserves the adjacency relationships among image tokens. Notably, we conduct a thorough theoretical analysis on the structural characteristics of the proposed polyline path mask and design an efficient algorithm for the computation of the polyline path mask. Next, we embed the polyline path mask into the self-attention mechanism of ViTs, enabling explicit modeling of spatial adjacency prior. Extensive experiments on standard benchmarks, including image classification, object detection, and segmentation, demonstrate that our model outperforms previous state-of-the-art approaches based on both state-space models and Transformers. For example, our proposed PPMA-T/S/B models achieve 48.7%/51.1%/52.3% mIoU on the ADE20K semantic segmentation task, surpassing RMT-T/S/B by 0.7%/1.3%/0.3%, respectively. Code is available at https://github.com/zhongchenzhao/PPMA. Zhongchen Zhao, Chaodong Xiao, Qi Xie 0002, Lei Zhang 0006, Deyu Meng |
NeurIPS | 4 |
| 2025 | MDFP-Net: A Model-Driven Deep Neural Network for Fourier PtychographyabstractFourier ptychography (FP) is a new computational imaging technique with the advantage of being able to provide super-resolution imaging. FP has a very complex degradation process. Merging with Fourier transforms and pupil aperture scanning causes difficulty in reconstructing high-resolution images by the commonly used deep neural network methods, e.g., based on convolutional neural networks (CNNs). In this paper, we propose a new optimization algorithm for FP, which is carefully designed so that it only constrains concise operations. Then, we unfold the proposed algorithm to design a new neural network, MDFP-Net, specifically for the FP task. MDFP-Net is consistent with a few stages, which well corresponds to the iterations of the proposed optimization algorithm for FP. This not only makes MDFP-Net more intuitively interpretable, but also makes MDFP-Net much more suitable for FP tasks than commonly used CNNs. Moreover, we have built a long-distance reflection FP measurement system and tested our neural network in real experiments. Simulation and real experimental results show that the proposed network can provide better reconstruction results than either traditional algorithms or other deep learning methods. Code is available at https://github.com/BP113/MDFPNET. Baopeng Li, Qi Xie 0002, Caiwen Ma, Zhibin Pan, Mingyang Yang, Xuewu Fan, Deyu Meng |
Comput. Vis. Media | 2 |
| 2025 | DS-Net: A model driven network framework for lesion segmentation on fundus image
Feiyu Tan, Qi Xie 0002, Jiahong Fu, Renzhen Wang, Deyu Meng |
Knowl. Based Syst. | 3 |
| 2025 | Rotation Equivariant Arbitrary-Scale Image Super-ResolutionabstractThe arbitrary-scale image super-resolution (ASISR), a recent popular topic in computer vision, aims to achieve arbitrary-scale high-resolution recoveries from a low-resolution input image. This task is realized by representing the image as a continuous implicit function through two fundamental modules, a deep-network-based encoder and an implicit neural representation (INR) module. Despite achieving notable progress, a crucial challenge of such a highly ill-posed setting is that many common geometric patterns, such as repetitive textures, edges, or shapes, are seriously warped and deformed in the low-resolution images, naturally leading to unexpected artifacts appearing in their high-resolution recoveries. Embedding rotation equivariance into the ASISR network is thus necessary, as it has been widely demonstrated that this enhancement enables the recovery to faithfully maintain the original orientations and structural integrity of geometric patterns underlying the input image. Motivated by this, we make efforts to construct a rotation equivariant ASISR method in this study. Specifically, we elaborately redesign the basic architectures of INR and encoder modules, incorporating intrinsic rotation equivariance capabilities beyond those of conventional ASISR networks. Through such amelioration, the ASISR network can, for the first time, be implemented with end-to-end rotational equivariance maintained from input to output. We also provide a solid theoretical analysis to evaluate its intrinsic equivariance error, demonstrating its inherent nature of embedding such an equivariance structure. The superiority of the proposed method is substantiated by experiments conducted on both simulated and real datasets. We also validate that the proposed framework can be readily integrated into current ASISR methods in a plug & play manner to further enhance their performance. Qi Xie 0002, Jiahong Fu, Zongben Xu, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Encoding sampling pattern for robust and generalized MRI reconstruction
Hong Wang 0021, Qi Xie 0002, Yefeng Zheng 0001, Deyu Meng |
Pattern Recognit. | 3 |
| 2025 | TRG-Net: An Interpretable and Controllable Rain GeneratorabstractExploring and modeling the rain generation mechanism is critical for augmenting paired data to ease the training of rainy image processing models. Most of the conventional methods handle this task in an artificial physical rendering manner, through elaborately designing fundamental elements constituting rains. These kinds of methods, however, are over-dependent on human subjectivity, which limits their adaptability to real rains. In contrast, recent deep learning (DL) methods have achieved great success by training a neural network-based generator from pre-collected rainy image data. However, current methods usually design the generator in a "closed box" manner, increasing the learning difficulty and data requirements. To address these issues, this study proposes a novel DL-based rain generator, which fully takes the physical generation mechanism underlying rains into consideration and well encodes the learning of the fundamental rain factors (i.e., shape, orientation, length, width, and sparsity) explicitly into the deep network. Its significance lies in that the generator not only elaborately designs essential elements of the rain to simulate expected rains, like conventional artificial strategies, but also finely adapts to complicated and diverse practical rainy images, like DL methods. By rationally adopting the filter parameterization technique, the proposed rain generator is finely controllable with respect to rain factors and able to learn the distribution of these factors purely from data without the need for rain factor labels. Our unpaired generation experiments demonstrate that the rain generated by the proposed rain generator is not only of higher quality but also more effective for deraining and downstream tasks compared to current state-of-the-art rain generation methods. Besides, the paired data augmentation experiments, including both in-distribution and out-of-distribution (OOD), further validate the diversity of samples generated by our model for in-distribution deraining and OOD generalization tasks. Zhiqiang Pang, Hong Wang 0021, Qi Xie 0002, Deyu Meng, Zongben Xu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | RSF-Conv: Rotation-and-Scale Equivariant Fourier Parameterized Convolution for Retinal Vessel SegmentationabstractRetinal vessel segmentation is of great clinical significance for the diagnosis of many eye-related diseases, but it is still a formidable challenge due to the intricate vascular morphology. With the skillful characterization of the translation symmetry existing in retinal vessels, convolutional neural networks (CNNs) have achieved great success in retinal vessel segmentation. However, the rotation-and-scale symmetry, as a more widespread image prior in retinal vessels, fails to be characterized by CNNs. Therefore, we propose a rotation-and-scale equivariant Fourier parameterized convolution (RSF-Conv) specifically for retinal vessel segmentation and provide the corresponding equivariance analysis. As a general module, RSF-Conv can be integrated into existing networks in a plug-and-play manner while significantly reducing the number of parameters. For instance, we replace the traditional convolution filters in U-Net, Iter-Net, DE-DCGCN-EE, and FR-UNet, with RSF-Convs, and faithfully conduct comprehensive experiments. RSF-Conv-enhanced methods not only have slight advantages under in-domain evaluation but also, more importantly, outperform all comparison methods by a significant margin under out-of-domain evaluation. It indicates that the remarkable generalization of RSF-Conv holds greater practical clinical significance for the prevalent cross-device and cross-hospital challenges in clinical practice. To comprehensively demonstrate the effectiveness of RSF-Conv, we also apply RSF-Conv + U-Net and RSF-Conv + Iter-Net to retinal artery/vein classification and achieve promising performance as well, indicating its clinical application potential. The code is available at https://github.com/szhc0gk/RSF-Conv. Zihong Sun, Hong Wang 0021, Qi Xie 0002, Yefeng Zheng 0001, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Rotation Equivariant Proximal Operator for Deep Unfolding Methods in Image RestorationabstractThe deep unfolding approach has attracted significant attention in computer vision tasks, which well connects conventional image processing modeling manners with more recent deep learning techniques. Specifically, by establishing a direct correspondence between algorithm operators at each implementation step and network modules within each layer, one can rationally construct an almost "white box" network architecture with high interpretability. In this architecture, only the predefined component of the proximal operator, known as a proximal network, needs manual configuration, enabling the network to automatically extract intrinsic image priors in a data-driven manner. In current deep unfolding methods, such a proximal network is generally designed as a CNN architecture, whose necessity has been proven by a recent theory. That is, CNN structure substantially delivers the translational symmetry image prior, which is the most universally possessed structural prior across various types of images. However, standard CNN-based proximal networks have essential limitations in capturing the rotation symmetry prior, another universal structural prior underlying general images. This leaves a large room for further performance improvement in deep unfolding approaches. To address this issue, this study makes efforts to suggest a high-accuracy rotation equivariant proximal network that effectively embeds rotation symmetry priors into the deep unfolding framework. Especially, we deduce, for the first time, the theoretical equivariant error for such a designed proximal network with arbitrary layers under arbitrary rotation degrees. This analysis should be the most refined theoretical conclusion for such error evaluation to date and is also indispensable for supporting the rationale behind such networks with intrinsic interpretability requirements. Through experimental validation on different vision tasks, including blind image super-resolution, medical image reconstruction, and image de-raining, the proposed method is validated to be capable of directly replacing the proximal network in current deep unfolding architecture and readily enhancing their state-of-the-art performance. This indicates its potential usability in general vision tasks. Jiahong Fu, Qi Xie 0002, Deyu Meng, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | OSCNet: Orientation-Shared Convolutional Network for CT Metal Artifact LearningabstractX-ray computed tomography (CT) has been broadly adopted in clinical applications for disease diagnosis and image-guided interventions. However, metals within patients always cause unfavorable artifacts in the recovered CT images. Albeit attaining promising reconstruction results for this metal artifact reduction (MAR) task, most of the existing deep-learning-based approaches have some limitations. The critical issue is that most of these methods have not fully exploited the important prior knowledge underlying this specific MAR task. Therefore, in this paper, we carefully investigate the inherent characteristics of metal artifacts which present rotationally symmetrical streaking patterns. Then we specifically propose an orientation-shared convolution representation mechanism to adapt such physical prior structures and utilize Fourier-series-expansion-based filter parametrization for modelling artifacts, which can finely separate metal artifacts from body tissues. By adopting the classical proximal gradient algorithm to solve the model and then utilizing the deep unfolding technique, we easily build the corresponding orientation-shared convolutional network, termed as OSCNet. Furthermore, considering that different sizes and types of metals would lead to different artifact patterns (e.g., intensity of the artifacts), to better improve the flexibility of artifact learning and fully exploit the reconstructed results at iterative stages for information propagation, we design a simple-yet-effective sub-network for the dynamic convolution representation of artifacts. By easily integrating the sub-network into the proposed OSCNet framework, we further construct a more flexible network structure, called OSCNet+, which improves the generalization performance. Through extensive experiments conducted on synthetic and clinical datasets, we comprehensively substantiate the effectiveness of our proposed methods. Code will be released at https://github.com/hongwang01/OSCNet. Hong Wang 0021, Qi Xie 0002, Dong Zeng, Jianhua Ma 0001, Deyu Meng, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Low-Light Image Enhancement by Retinex-Based Algorithm Unrolling and AdjustmentabstractLow-light image enhancement (LIE) has attracted tremendous research interests in recent years. Retinex theory-based deep learning methods, following a decomposition-adjustment pipeline, have achieved promising performance due to their physical interpretability. However, existing Retinex-based deep learning methods are still suboptimal, failing to leverage useful insights from traditional approaches. Meanwhile, the adjustment step is either oversimplified or overcomplicated, resulting in unsatisfactory performance in practice. To address these issues, we propose a novel deep-learning framework for LIE. The framework consists of a decomposition network (DecNet) inspired by algorithm unrolling and adjustment networks considering both global and local brightness. The algorithm unrolling allows the integration of both implicit priors learned from data and explicit priors inherited from traditional methods, facilitating better decomposition. Meanwhile, considering global and local brightness guides the design of effective yet lightweight adjustment networks. Moreover, we introduce a self-supervised fine-tuning strategy that achieves promising performance without manual hyperparameter tuning. Extensive experiments on benchmark LIE datasets demonstrate the superiority of our approach over existing state-of-the-art methods both quantitatively and qualitatively. Code is available at https://github.com/Xinyil256/RAUNA2023. Qi Xie 0002, Qian Zhao 0002, Hong Wang 0021, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | RCDNet: An Interpretable Rain Convolutional Dictionary Network for Single Image DerainingabstractAs common weather, rain streaks adversely degrade the image quality and tend to negatively affect the performance of outdoor computer vision systems. Hence, removing rains from an image has become an important issue in the field. To handle such an ill-posed single image deraining task, in this article, we specifically build a novel deep architecture, called rain convolutional dictionary network (RCDNet), which embeds the intrinsic priors of rain streaks and has clear interpretability. In specific, we first establish a rain convolutional dictionary (RCD) model for representing rain streaks and utilize the proximal gradient descent technique to design an iterative algorithm only containing simple operators for solving the model. By unfolding it, we then build the RCDNet in which every network module has clear physical meanings and corresponds to each operation involved in the algorithm. This good interpretability greatly facilitates an easy visualization and analysis of what happens inside the network and why it works well in the inference process. Moreover, taking into account the domain gap issue in real scenarios, we further design a novel dynamic RCDNet, where the rain kernels can be dynamically inferred corresponding to input rainy images and then help shrink the space for rain layer estimation with few rain maps, so as to ensure a fine generalization performance in the inconsistent scenarios of rain types between training and testing data. By end-to-end training such an interpretable network, all involved rain kernels and proximal operators can be automatically extracted, faithfully characterizing the features of both rain and clean background layers and, thus, naturally leading to better deraining performance. Comprehensive experiments implemented on a series of representative synthetic and real datasets substantiate the superiority of our method, especially on its well generality to diverse testing scenarios and good interpretability for all its modules, compared with state-of-the-art single image derainers both visually and quantitatively. Code is available at https://github.com/hongwang01/DRCDNet. Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Yuexiang Li, Yong Liang 0001, Yefeng Zheng 0001, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | A Learnable Optimization and Regularization Approach to Massive MIMO CSI FeedbackabstractChannel state information (CSI) plays a critical role in achieving the potential benefits of massive multiple input multiple output (MIMO) systems. In frequency division duplex (FDD) massive MIMO systems, the base station (BS) relies on sustained and accurate CSI feedback from users. However, due to the large number of antennas and users being served in massive MIMO systems, feedback overhead can become a bottleneck. In this paper, we propose a model-driven deep learning method for CSI feedback, called learnable optimization and regularization algorithm (LORA). Instead of using$l_{1}$-norm as the regularization term, LORA introduces a learnable regularization module that adapts to characteristics of CSI automatically. The conventional Iterative Shrinkage-Thresholding Algorithm (ISTA) is unfolded into a neural network, which can learn both the optimization process and the regularization term by end-to-end training. We show that LORA improves the CSI feedback accuracy and speed. Besides, a novel learnable quantization method and the corresponding training scheme are proposed, and it is shown that LORA can operate successfully at different bit rates, providing flexibility in terms of the CSI feedback overhead. Various realistic scenarios are considered to demonstrate the effectiveness and robustness of LORA through numerical simulations. Zhengyang Hu 0001, Guanzhang Liu, Qi Xie 0002, Jiang Xue 0001, Deyu Meng, Deniz Gündüz |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Memory-Augmented Deep Unfolding Network for Guided Image Super-resolution
Man Zhou 0003, Jinshan Pan, Wenqi Ren, Qi Xie 0002, Xiangyong Cao |
Int. J. Comput. Vis. | 5 |
| 2023 | Fourier Series Expansion Based Filter Parametrization for Equivariant ConvolutionsabstractIt has been shown that equivariant convolution is very helpful for many types of computer vision tasks. Recently, the 2D filter parametrization technique has played an important role for designing equivariant convolutions, and has achieved success in making use of rotation symmetry of images. However, the current filter parametrization strategy still has its evident drawbacks, where the most critical one lies in the accuracy problem of filter representation. To address this issue, in this paper we explore an ameliorated Fourier series expansion for 2D filters, and propose a new filter parametrization method based on it. The proposed filter parametrization method not only finely represents 2D filters with zero error when the filter is not rotated (similar as the classical Fourier series expansion), but also substantially alleviates the aliasing-effect-caused quality degradation when the filter is rotated (which usually arises in classical Fourier series expansion method). Accordingly, we construct a new equivariant convolution method based on the proposed filter parametrization method, named F-Conv. We prove that the equivariance of the proposed F-Conv is exact in the continuous domain, which becomes approximate only after discretization. Moreover, we provide theoretical error analysis for the case when the equivariance is approximate, showing that the approximation error is related to the mesh size and filter size. Extensive experiments show the superiority of the proposed method. Particularly, we adopt rotation equivariant convolution methods to a typical low-level image processing task, image super-resolution. It can be substantiated that the proposed F-Conv based method evidently outperforms classical convolution based methods. Compared with pervious filter parametrization based methods, the F-Conv performs more accurately on this low-level image processing task, reflecting its intrinsic capability of faithfully preserving rotation symmetries in local image features. Qi Xie 0002, Qian Zhao 0002, Zongben Xu, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | KXNet: A Model-Driven Deep Neural Network for Blind Super-Resolution
Jiahong Fu, Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu |
ECCV (19) | 3 |
| 2022 | Orientation-Shared Convolution Representation for CT Metal Artifact Learning
Hong Wang 0021, Qi Xie 0002, Yuexiang Li, Yawen Huang, Deyu Meng, Yefeng Zheng 0001 |
MICCAI (6) | 2 |
| 2022 | MHF-Net: An Interpretable Deep Network for Multispectral and Hyperspectral Image FusionabstractMultispectral and hyperspectral image fusion (MS/HS fusion) aims to fuse a high-resolution multispectral (HrMS) and a low-resolution hyperspectral (LrHS) images to generate a high-resolution hyperspectral (HrHS) image, which has become one of the most commonly addressed problems for hyperspectral image processing. In this paper, we specifically designed a network architecture for the MS/HS fusion task, called MHF-net, which not only contains clear interpretability, but also reasonably embeds the well studied linear mapping that links the HrHS image to HrMS and LrHS images. In particular, we first construct an MS/HS fusion model which merges the generalization models of low-resolution images and the low-rankness prior knowledge of HrHS image into a concise formulation, and then we build the proposed network by unfolding the proximal gradient algorithm for solving the proposed model. As a result of the careful design for the model and algorithm, all the fundamental modules in MHF-net have clear physical meanings and are thus easily interpretable. This not only greatly facilitates an easy intuitive observation and analysis on what happens inside the network, but also leads to its good generalization capability. Based on the architecture of MHF-net, we further design two deep learning regimes for two general cases in practice: consistent MHF-net and blind MHF-net. The former is suitable in the case that spectral and spatial responses of training and testing data are consistent, just as considered in most of the pervious general supervised MS/HS fusion researches. The latter ensures a good generalization in mismatch cases of spectral and spatial responses in training and testing data, and even across different sensors, which is generally considered to be a challenging issue for general supervised MS/HS fusion methods. Experimental results on simulated and real data substantiate the superiority of our method both visually and quantitatively as compared with state-of-the-art methods along this line of research. Qi Xie 0002, Qian Zhao 0002, Zongben Xu, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Learning to Purify Noisy Labels via Meta Soft Label CorrectorabstractRecent deep neural networks (DNNs) can easily overfit to biased training data with noisy labels. Label correction strategy is commonly used to alleviate this issue by identifying suspected noisy labels and then correcting them. Current approaches to correcting corrupted labels usually need manually pre-defined label correction rules, which makes it hard to apply in practice due to the large variations of such manual strategies with respect to different problems. To address this issue, we propose a meta-learning model, aiming at attaining an automatic scheme which can estimate soft labels through meta-gradient descent step under the guidance of a small amount of noise-free meta data. By viewing the label correction procedure as a meta-process and using a meta-learner to automatically correct labels, our method can adaptively obtain rectified soft labels gradually in iteration according to current training problems. Besides, our method is model-agnostic and can be combined with any other existing classification models with ease to make it available to noisy label cases. Comprehensive experiments substantiate the superiority of our method in both synthetic and real-world problems with noisy labels compared with current state-of-the-art label correction strategies. Qi Xie 0002, Qian Zhao 0002, Deyu Meng |
AAAI | 3 |
| 2021 | Learning an Explicit Weighting Scheme for Adapting Complex HSI NoiseabstractAn efficient approach for handling hyperspectral image (HSI) denoising issue is to impose weights on different HSI pixels to suppress negative influence brought by noisy elements. Such weighting scheme, however, largely depends on the prior understanding or subjective distribution assumption on HSI noises, making them easily biased to complicated real noises, and hardly generalizable to diverse practical scenarios. Against this issue, this paper proposes a new scheme aiming to capture general weighting principle in a data-driven manner. Specifically, such weighting principle is delivered by an explicit function, called hyper-weight-net (HWnet), mapping from an input noisy image to its properly imposed weights. A Bayesian framework as well as a variational inference algorithm for inferring HWnet parameters is elaborately designed, expecting to extract the latent weighting rule for general diverse and complicated noisy HSIs. Comprehensive experiments substantiate that the learned HWnet can be not only finely generalized to different noise types from those used in training, but also effectively transferred to other weighted models. Besides, as a sounder guidance, HWnet can help to more faithfully and robustly achieve deep hyperspectral prior(DHP). The extracted weights by HWnet are verified to be able to effectively capture complex noise knowledge underlying input HSI, revealing its working insight in experiments. Xiangyu Rui, Xiangyong Cao, Qi Xie 0002, Zongsheng Yue, Qian Zhao 0002, Deyu Meng |
CVPR | 3 |
| 2021 | From Rain Generation to Rain RemovalabstractFor the single image rain removal (SIRR) task, the performance of deep learning (DL)-based methods is mainly affected by the designed deraining models and training datasets. Most of current state-of-the-art focus on constructing powerful deep models to obtain better deraining results. In this paper, to further improve the deraining performance, we novelly attempt to handle the SIRR task from the perspective of training datasets by exploring a more efficient way to synthesize rainy images. Specifically, we build a full Bayesian generative model for rainy image where the rain layer is parameterized as a generator with the input as some latent variables representing the physical structural rain factors, e.g., direction, scale, and thickness. To solve this model, we employ the variational inference framework to approximate the expected statistical distribution of rainy image in a data-driven manner. With the learned generator, we can automatically and sufficiently generate diverse and non-repetitive training pairs so as to efficiently enrich and augment the existing benchmark datasets. User study qualitatively and quantitatively evaluates the realism of generated rainy images. Comprehensive experiments substantiate that the proposed model can faithfully extract the complex rain distribution that not only helps significantly improve the deraining performance of current deep single image derainers, but also largely loosens the requirement of large training sample pre-collection for the SIRR task. Code is available in https://github.com/hongwang01/VRGNet. Hong Wang 0021, Zongsheng Yue, Qi Xie 0002, Qian Zhao 0002, Yefeng Zheng 0001, Deyu Meng |
CVPR | 3 |
| 2021 | Structural residual learning for single image rain removal
Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Yong Liang 0001, Deyu Meng |
Knowl. Based Syst. | 3 |
| 2020 | A Model-Driven Deep Neural Network for Single Image Rain RemovalabstractDeep learning (DL) methods have achieved state-of-the-art performance in the task of single image rain removal. Most of current DL architectures, however, are still lack of sufficient interpretability and not fully integrated with physical structures inside general rain streaks. To this issue, in this paper, we propose a model-driven deep neural network for the task, with fully interpretable network structures. Specifically, based on the convolutional dictionary learning mechanism for representing rain, we propose a novel single image deraining model and utilize the proximal gradient descent technique to design an iterative algorithm only containing simple operators for solving the model. Such a simple implementation scheme facilitates us to unfold it into a new deep network architecture, called rain convolutional dictionary network (RCDNet), with almost every network module one-to-one corresponding to each operation involved in the algorithm. By end-to-end training the proposed RCDNet, all the rain kernels and proximal operators can be automatically extracted, faithfully characterizing the features of both rain and clean background layers, and thus naturally lead to its better deraining performance, especially in real scenarios. Comprehensive experiments substantiate the superiority of the proposed network, especially its well generality to diverse testing scenarios and good interpretability for all its modules, as compared with state-of-the-arts both visually and quantitatively. Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Deyu Meng |
CVPR | 2 |
| 2020 | Color and direction-invariant nonlocal self-similarity prior and its application to color image denoising
Qi Xie 0002, Qian Zhao 0002, Zongben Xu, Deyu Meng |
Sci. China Inf. Sci. | 1 |
| 2020 | Enhanced 3DTV Regularization and Its Applications on HSI Denoising and Compressed SensingabstractThe total variation (TV) is a powerful regularization term encoding the local smoothness prior structure underlying images. By combining the TV regularization term with low rank prior, the 3D total variation (3DTV) regularizer has achieved advanced performance in general hyperspectral image (HSI) processing tasks. Intrinsically, 3DTV assumes i.i.d. sparsity structures on all bands of the gradient maps calculated along the spectrum and space of an HSI. This, however, largely deviates from the real-world cases, where the gradient maps generally have different while correlated gradient map structures across all bands. To alleviate this issue, we propose an enhanced 3DTV (E-3DTV) regularization term beyond the conventional. Instead of imposing sparsity on gradient maps themselves, the new term calculates sparsity on the subspace bases on gradient maps along all bands of an HSI, which naturally encodes the correlation and difference among all these bands, and thus more faithfully reflects the insightful configurations of an HSI. The E-3DTV term can easily replace the conventional 3DTV term and be embedded into an HSI processing model to ameliorate its performance. We made such attempts on two typical related tasks: HSI denoising and compressed sensing. The superiority of our proposed method is substantiated by extensive experiments on synthetic and real HSI data, visually and quantitatively on both tasks, as compared with current state-of-the-arts. The code of our algorithm is released athttps://github.com/andrew-pengjj/Enhanced-3DTV.git. Jiangjun Peng, Qi Xie 0002, Qian Zhao 0002, Yao Wang 0003, Yee Leung, Deyu Meng |
IEEE Trans. Image Process. | 2 |
| 2020 | Full-Spectrum-Knowledge-Aware Tensor Model for Energy-Resolved CT Iterative ReconstructionabstractEnergy-resolved computed tomography (ErCT) with a photon counting detector concurrently produces multiple CT images corresponding to different photon energy ranges. It has the potential to generate energy-dependent images with improved contrast-to-noise ratio and sufficient material-specific information. Since the number of detected photons in one energy bin in ErCT is smaller than that in conventional energy-integrating CT (EiCT), ErCT images are inherently more noisy than EiCT images, which leads to increased noise and bias in the subsequent material estimation. In this work, we first deeply analyze the intrinsic tensor properties of two-dimensional (2D) ErCT images acquired in different energy bins and then present a F ull- S pectrum-knowledge-aware Tensor analysis and processing (FSTensor) method for ErCT reconstruction to suppress noise-induced artifacts to obtain high-quality ErCT images and high-accuracy material images. The presented method is based on three considerations: (1) 2D ErCT images obtained in different energy bins can be treated as a 3-order tensor with three modes, i.e., width, height and energy bin, and a rich global correlation exists among the three modes, which can be characterized by tensor decomposition. (2) There is a locally piecewise smooth property in the 3-order ErCT images, and it can be captured by a tensor total variation regularization. (3) The images from the full spectrum are much better than the ErCT images with respect to noise variance and structural details and serve as external information to improve the reconstruction performance. We then develop an alternating direction method of multipliers algorithm to numerically solve the presented FSTensor method. We further utilize a genetic algorithm to tackle the parameter selection in ErCT reconstruction, instead of manually determining parameters. Simulation, preclinical and synthesized clinical ErCT results demonstrate that the presented FSTensor method leads to significant improvements over the filtered back-projection, robust principal component analysis, tensor-based dictionary learning and low-rank tensor decomposition with spatial-temporal total variation methods. Dong Zeng, Yongshuai Ge, Sui Li, Qi Xie 0002, Hao Zhang 0026, Zhaoying Bian, Qian Zhao 0002, Yuanqing Li 0001, Zongben Xu, Deyu Meng, Jianhua Ma 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2019 | Multispectral and Hyperspectral Image Fusion by MS/HS Fusion NetabstractHyperspectral imaging can help better understand the characteristics of different materials, compared with traditional image systems. However, only high-resolution multispectral (HrMS) and low-resolution hyperspectral (LrHS) images can generally be captured at video rate in practice. In this paper, we propose a model-based deep learning approach for merging an HrMS and LrHS images to generate a high-resolution hyperspectral (HrHS) image. In specific, we construct a novel MS/HS fusion model which takes the observation models of low-resolution images and the low-rankness knowledge along the spectral mode of HrHS image into consideration. Then we design an iterative algorithm to solve the model by exploiting the proximal gradient method. And then, by unfolding the designed algorithm, we construct a deep network, called MS/HS Fusion Net, with learning the proximal operators and model parameters by convolutional neural networks. Experimental results on simulated and real data substantiate the superiority of our method both visually and quantitatively as compared with state-of-the-art methods along this line of research. Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Wangmeng Zuo, Zongben Xu |
CVPR | 1 |
| 2019 | Meta-Weight-Net: Learning an Explicit Mapping For Sample WeightingabstractCurrent deep neural networks(DNNs) can easily overfit to biased training data with corrupted labels or class imbalance. Sample re-weighting strategy is commonly used to alleviate this issue by designing a weighting function mapping from training loss to sample weight, and then iterating between weight recalculating and classifier updating. Current approaches, however, need manually pre-specify the weighting function as well as its additional hyper-parameters. It makes them fairly hard to be generally applied in practice due to the significant variation of proper weighting schemes relying on the investigated problem and training data. To address this issue, we propose a method capable of adaptively learning an explicit weighting function directly from data. The weighting function is an MLP with one hidden layer, constituting a universal approximator to almost any continuous functions, making the method able to fit a wide range of weighting function forms including those assumed in conventional research. Guided by a small amount of unbiased meta-data, the parameters of the weighting function can be finely updated simultaneously with the learning process of the classifiers. Synthetic and real experiments substantiate the capability of our method for achieving proper weighting functions in class imbalance and noisy label cases, fully complying with the common settings in traditional methods, and more complicated scenarios beyond conventional cases. This naturally leads to its better accuracy than other state-of-the-art methods. Qi Xie 0002, Lixuan Yi, Qian Zhao 0002, Sanping Zhou, Zongben Xu, Deyu Meng |
NeurIPS | 2 |
| 2019 | An Efficient Iterative Cerebral Perfusion CT Reconstruction via Low-Rank Tensor Decomposition With Spatial-Temporal Total Variation RegularizationabstractCerebrovascular diseases, i.e., acute stroke, are a common cause of serious long-term disability. Cerebral perfusion computed tomography (CPCT) can provide rapid, high-resolution, quantitative hemodynamic maps to assess and stratify perfusion in patients with acute stroke symptoms. However, CPCT imaging typically involves a substantial radiation dose due to its repeated scanning protocol. Therefore, in this paper, we present a low-dose CPCT image reconstruction method to yield high-quality CPCT images and high-precision hemodynamic maps by utilizing the great similarity information among the repeated scanned CPCT images. Specifically, a newly developed low-rank tensor decomposition with spatial-temporal total variation (LRTD-STTV) regularization is incorporated into the reconstruction model. In the LRTD-STTV regularization, the tensor Tucker decomposition is used to describe global spatial-temporal correlations hidden in the sequential CPCT images, and it is superior to the matricization model (i.e., low-rank model) that fails to fully investigate the prior knowledge of the intrinsic structures of the CPCT images after vectorizing the CPCT images. Moreover, the spatial-temporal TV regularization is used to characterize the local piecewise smooth structure in the spatial domain and the pixels' similarity with the adjacent frames in the temporal domain, because the intensity at each pixel in CPCT images is similar to its neighbors. Therefore, the presented LRTD-STTV model can efficiently deliver faithful underlying information of the CPCT images and preserve the spatial structures. An efficient alternating direction method of multipliers algorithm is also developed to solve the presented LRTD-STTV model. Extensive experimental results on numerical phantom and patient data are clearly demonstrated that the presented model can significantly improve the quality of CPCT images and provide accurate diagnostic features in hemodynamic maps for low-dose cases compared with the existing popular algorithms. Sui Li, Dong Zeng, Jiangjun Peng, Zhaoying Bian, Hao Zhang 0026, Qi Xie 0002, Yuting Liao, Shanli Zhang, Jing Huang 0018, Deyu Meng, Zongben Xu, Jianhua Ma 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Video Rain Streak Removal by Multiscale Convolutional Sparse CodingabstractVideos captured by outdoor surveillance equipments sometimes contain unexpected rain streaks, which brings difficulty in subsequent video processing tasks. Rain streak removal from a video is thus an important topic in recent computer vision research. In this paper, we raise two intrinsic characteristics specifically possessed by rain streaks. Firstly, the rain streaks in a video contain repetitive local patterns sparsely scattered over different positions of the video. Secondly, the rain streaks are with multiscale configurations due to their occurrence on positions with different distances to the cameras. Based on such understanding, we specifically formulate both characteristics into a multiscale convolutional sparse coding (MS-CSC) model for the video rain streak removal task. Specifically, we use multiple convolutional filters convolved on the sparse feature maps to deliver the former characteristic, and further use multiscale filters to represent different scales of rain streaks. Such a new encoding manner makes the proposed method capable of properly extracting rain streaks from videos, thus getting fine video deraining effects. Experiments implemented on synthetic and real videos verify the superiority of the proposed method, as compared with the state-of-the-art ones along this research line, both visually and quantitatively. Minghan Li 0001, Qi Xie 0002, Qian Zhao 0002, Wei Wei 0006, Shuhang Gu, Deyu Meng |
CVPR | 2 |
| 2018 | Kronecker-Basis-Representation Based Tensor Sparsity and Its Applications to Tensor RecoveryabstractAs a promising way for analyzing data, sparse modeling has achieved great success throughout science and engineering. It is well known that the sparsity/low-rank of a vector/matrix can be rationally measured by nonzero-entries-number ( norm)/nonzero- singular-values-number (rank), respectively. However, data from real applications are often generated by the interaction of multiple factors, which obviously cannot be sufficiently represented by a vector/matrix, while a high order tensor is expected to provide more faithful representation to deliver the intrinsic structure underlying such data ensembles. Unlike the vector/matrix case, constructing a rational high order sparsity measure for tensor is a relatively harder task. To this aim, in this paper we propose a measure for tensor sparsity, called Kronecker-basis-representation based tensor sparsity measure (KBR briefly), which encodes both sparsity insights delivered by Tucker and CANDECOMP/PARAFAC (CP) low-rank decompositions for a general tensor. Then we study the KBR regularization minimization (KBRM) problem, and design an effective ADMM algorithm for solving it, where each involved parameter can be updated with closed-form equations. Such an efficient solver makes it possible to extend KBR to various tasks like tensor completion and tensor robust principal component analysis. A series of experiments, including multispectral image (MSI) denoising, MSI completion and background subtraction, substantiate the superiority of the proposed methods beyond state-of-the-arts. Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Should We Encode Rain Streaks in Video as Deterministic or Stochastic?abstractVideos taken in the wild sometimes contain unexpected rain streaks, which brings difficulty in subsequent video processing tasks. Rain streak removal in a video (RSRV) is thus an important issue and has been attracting much attention in computer vision. Different from previous RSRV methods formulating rain streaks as a deterministic message, this work first encodes the rains in a stochastic manner, i.e., a patch-based mixture of Gaussians. Such modification makes the proposed model capable of finely adapting a wider range of rain variations instead of certain types of rain configurations as traditional. By integrating with the spatiotemporal smoothness configuration of moving objects and low-rank structure of background scene, we propose a concise model for RSRV, containing one likelihood term imposed on the rain streak layer and two prior terms on the moving object and background scene layers of the video. Experiments implemented on videos with synthetic and real rains verify the superiority of the proposed method, as compared with the state-of-the-art methods, both visually and quantitatively in various performance metrics. Wei Wei 0006, Lixuan Yi, Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu |
ICCV | 3 |
| 2017 | Self-Paced Co-trainingabstractCo-training is a well-known semi-supervised learning approach which trains classifiers on two different views and exchanges labels of unlabeled instances in an iterative way. During co-training process, labels of unlabeled instances in the training pool are very likely to be false especially in the initial training rounds, while the standard co-training algorithm utilizes a “draw without replacement” manner and does not remove these false labeled instances from training. This issue not only tends to degenerate its performance but also hampers its fundamental theory. Besides, there is no optimization model to explain what objective a cotraining process optimizes. To these issues, in this study we design a new co-training algorithm named self-paced cotraining (SPaCo) with a “draw with replacement” learning mode. The rationality of SPaCo can be proved under theoretical assumptions utilized in traditional co-training research, and furthermore, the algorithm exactly complies with the alternative optimization process for an optimization model of self-paced curriculum learning, which can be finely explained in robust learning manner. Experimental results substantiate the superiority of the proposed method as compared with current state-of-the-art co-training methods. Fan Ma, Deyu Meng, Qi Xie 0002, Zina Li, Xuanyi Dong |
ICML | 3 |
| 2017 | Weighted Nuclear Norm Minimization and Its Applications to Low Level Vision
Shuhang Gu, Qi Xie 0002, Deyu Meng, Wangmeng Zuo, Xiangchu Feng, Lei Zhang 0006 |
Int. J. Comput. Vis. | 2 |
| 2017 | Robust Low-Dose CT Sinogram Preprocessing via Exploiting Noise-Generating MechanismabstractComputed tomography (CT) image recovery from low-mAs acquisitions without adequate treatment is always severely degraded due to a number of physical factors. In this paper, we formulate the low-dose CT sinogram preprocessing as a standard maximum a posteriori (MAP) estimation, which takes full consideration of the statistical properties of the two intrinsic noise sources in low-dose CT, i.e., the X-ray photon statistics and the electronic noise background. In addition, instead of using a general image prior as found in the traditional sinogram recovery models, we design a new prior formulation to more rationally encode the piecewise-linear configurations underlying a sinogram than previously used ones, like the TV prior term. As compared with the previous methods, especially the MAP-based ones, both the likelihood/loss and prior/regularization terms in the proposed model are ameliorated in a more accurate manner and better comply with the statistical essence of the generation mechanism of a practical sinogram. We further construct an efficient alternating direction method of multipliers algorithm to solve the proposed MAP framework. Experiments on simulated and real low-dose CT data demonstrate the superiority of the proposed method according to both visual inspection and comprehensive quantitative performance evaluation. Qi Xie 0002, Dong Zeng, Qian Zhao 0002, Deyu Meng, Zongben Xu, Zhengrong Liang, Jianhua Ma 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2017 | Low-Dose Dynamic Cerebral Perfusion Computed Tomography Reconstruction via Kronecker-Basis-Representation Tensor Sparsity RegularizationabstractDynamic cerebral perfusion computed tomography (DCPCT) has the ability to evaluate the hemodynamic information throughout the brain. However, due to multiple 3-D image volume acquisitions protocol, DCPCT scanning imposes high radiation dose on the patients with growing concerns. To address this issue, in this paper, based on the robust principal component analysis (RPCA, or equivalently the low-rank and sparsity decomposition) model and the DCPCT imaging procedure, we propose a new DCPCT image reconstruction algorithm to improve low-dose DCPCT and perfusion maps quality via using a powerful measure, called Kronecker-basis-representation tensor sparsity regularization, for measuring low-rankness extent of a tensor. For simplicity, the first proposed model is termed tensor-based RPCA (T-RPCA). Specifically, the T-RPCA model views the DCPCT sequential images as a mixture of low-rank, sparse, and noise components to describe the maximum temporal coherence of spatial structure among phases in a tensor framework intrinsically. Moreover, the low-rank component corresponds to the "background" part with spatial-temporal correlations, e.g., static anatomical contribution, which is stationary over time about structure, and the sparse component represents the time-varying component with spatial-temporal continuity, e.g., dynamic perfusion enhanced information, which is approximately sparse over time. Furthermore, an improved nonlocal patch-based T-RPCA (NL-T-RPCA) model which describes the 3-D block groups of the "background" in a tensor is also proposed. The NL-T-RPCA model utilizes the intrinsic characteristics underlying the DCPCT images, i.e., nonlocal self-similarity and global correlation. Two efficient algorithms using alternating direction method of multipliers are developed to solve the proposed T-RPCA and NL-T-RPCA models, respectively. Extensive experiments with a digital brain perfusion phantom, preclinical monkey data, and clinical patient data clearly demonstrate that the two proposed models can achieve more gains than the existing popular algorithms in terms of both quantitative and visual quality evaluations from low-dose acquisitions, especially as low as 20 mAs. Dong Zeng, Qi Xie 0002, Wenfei Cao, Jiahui Lin, Hao Zhang 0026, Shanli Zhang, Jing Huang 0018, Zhaoying Bian, Deyu Meng, Zongben Xu, Zhengrong Liang, Wufan Chen, Jianhua Ma 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2016 | Multispectral Images Denoising by Intrinsic Tensor Sparsity RegularizationabstractMultispectral images (MSI) can help deliver more faithful representation for real scenes than the traditional image system, and enhance the performance of many computer vision tasks. In real cases, however, an MSI is always corrupted by various noises. In this paper, we propose a new tensor-based denoising approach by fully considering two intrinsic characteristics underlying an MSI, i.e., the global correlation along spectrum (GCS) and nonlocal self-similarity across space (NSS). In specific, we construct a new tensor sparsity measure, called intrinsic tensor sparsity (ITS) measure, which encodes both sparsity insights delivered by the most typical Tucker and CANDECOMP/ PARAFAC (CP) low-rank decomposition for a general tensor. Then we build a new MSI denoising model by applying the proposed ITS measure on tensors formed by non-local similar patches within the MSI. The intrinsic GCS and NSS knowledge can then be efficiently explored under the regularization of this tensor sparsity measure to finely rectify the recovery of a MSI from its corruption. A series of experiments on simulated and real MSI denoising problems show that our method outperforms all state-of-the-arts under comprehensive quantitative performance measures. Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu, Shuhang Gu, Wangmeng Zuo, Lei Zhang 0006 |
CVPR | 1 |
| 2015 | Self-Paced Learning for Matrix FactorizationabstractMatrix factorization (MF) has been attracting much attention due to its wide applications. However, since MF models are generally non-convex, most of the existing methods are easily stuck into bad local minima, especially in the presence of outliers and missing data. To alleviate this deficiency, in this study we present a new MF learning methodology by gradually including matrix elements into MF training from easy to complex. This corresponds to a recently proposed learning fashion called self-paced learning (SPL), which has been demonstrated to be beneficial in avoiding bad local minima. We also generalize the conventional binary (hard) weighting scheme for SPL to a more effective real-valued (soft) weighting manner. The effectiveness of the proposed self-paced MF method is substantiated by a series of experiments on synthetic, structure from motion and background subtraction data. Qian Zhao 0002, Deyu Meng, Lu Jiang 0004, Qi Xie 0002, Zongben Xu, Alex Hauptmann 0001 |
AAAI | 4 |
| 2015 | Convolutional Sparse Coding for Image Super-ResolutionabstractMost of the previous sparse coding (SC) based super resolution (SR) methods partition the image into overlapped patches, and process each patch separately. These methods, however, ignore the consistency of pixels in overlapped patches, which is a strong constraint for image reconstruction. In this paper, we propose a convolutional sparse coding (CSC) based SR (CSC-SR) method to address the consistency issue. Our CSC-SR involves three groups of parameters to be learned: (i) a set of filters to decompose the low resolution (LR) image into LR sparse feature maps, (ii) a mapping function to predict the high resolution (HR) feature maps from the LR ones, and (iii) a set of filters to reconstruct the HR images from the predicted HR feature maps via simple convolution operations. By working directly on the whole image, the proposed CSC-SR algorithm does not need to divide the image into overlapped patches, and can exploit the image global correlation to produce more robust reconstruction of image local structures. Experimental results clearly validate the advantages of CSC over patch based SC in SR application. Compared with state-of-the-art SR methods, the proposed CSC-SR method achieves highly competitive PSNR results, while demonstrating better edge and texture preservation performance. Shuhang Gu, Wangmeng Zuo, Qi Xie 0002, Deyu Meng, Xiangchu Feng, Lei Zhang 0006 |
ICCV | 3 |
| 2015 | A Novel Sparsity Measure for Tensor RecoveryabstractIn this paper, we propose a new sparsity regularizer for measuring the low-rank structure underneath a tensor. The proposed sparsity measure has a natural physical meaning which is intrinsically the size of the fundamental Kronecker basis to express the tensor. By embedding the sparsity measure into the tensor completion and tensor robust PCA frameworks, we formulate new models to enhance their capability in tensor recovery. Through introducing relaxation forms of the proposed sparsity measure, we also adopt the alternating direction method of multipliers (ADMM) for solving the proposed models. Experiments implemented on synthetic and multispectral image data sets substantiate the effectiveness of the proposed methods. Qian Zhao 0002, Deyu Meng, Xu Kong, Qi Xie 0002, Wenfei Cao, Yao Wang 0003, Zongben Xu |
ICCV | 4 |