Chenglong Bao

dblp:117/4815 · DBLP profile ↗
← Back
31ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0002-1201-1212ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 3 since 2021
YearPublicationVenuePosition
2025 A Tight Convergence Analysis of Inexact Stochastic Proximal Point Algorithm for Stochastic Composite Optimization Problems
abstract
The \textbf{i}nexact \textbf{s}tochastic \textbf{p}roximal \textbf{p}oint \textbf{a}lgorithm (isPPA) is popular for solving stochastic composite optimization problems with many applications in machine learning. While the convergence theory of the (inexact) PPA has been well established, the known convergence guarantees of isPPA require restrictive assumptions. In this paper, we establish the stability and almost sure convergence of isPPA under mild assumptions, where smoothness and (restrictive) strong convexity of the objective function are not required. Imposing a local Lipschitz condition on component functions and a quadratic growth condition on the objective function, we establish last-iterate iteration complexity bounds of isPPA regarding the distance to the solution set and the Karush–Kuhn–Tucker (KKT) residual. Moreover, we show that the established iteration complexity bounds are tight up to a constant by explicitly analyzing the bounds for the regularized Fr\'echet mean problem. We further validate the established convergence guarantees of isPPA by numerical experiments.
Shulan Zhu, Chenglong Bao, Defeng Sun, Yancheng Yuan
ICLR2
2025 A Regularized Newton Method for Nonconvex Optimization with Global and Local Complexity Guarantees
abstract
Finding an $\epsilon$-stationary point of a nonconvex function with a Lipschitz continuous Hessian is a central problem in optimization. Regularized Newton methods are a classical tool and have been studied extensively, yet they still face a trade‑off between global and local convergence. Whether a parameter-free algorithm of this type can simultaneously achieve optimal global complexity and quadratic local convergence remains an open question. To bridge this long-standing gap, we propose a new class of regularizers constructed from the current and previous gradients, and leverage the conjugate gradient approach with a negative curvature monitor to solve the regularized Newton equation. The proposed algorithm is adaptive, requiring no prior knowledge of the Hessian Lipschitz constant, and achieves a global complexity of $O(\epsilon^{-\frac{3}{2}})$ in terms of the second-order oracle calls, and $\tilde O(\epsilon^{-\frac{7}{4}})$ for Hessian-vector products, respectively. When the iterates converge to a point where the Hessian is positive definite, the method exhibits quadratic local convergence. Preliminary numerical results, including training the physics-informed neural networks, illustrate the competitiveness of our algorithm.
Bingrui Li, Chenglong Bao, Jun Zhu 0001
NeurIPS4
2025 Convection-Diffusion Equation: A Theoretically Certified Framework for Neural Networks
abstract
Differential equations have demonstrated intrinsic connections to network structures, linking discrete network layers through continuous equations. Most existing approaches focus on the interaction between ordinary differential equations (ODEs) and feature transformations, primarily working on input signals. In this paper, we study the partial differential equation (PDE) model of neural networks, viewing the neural network as a functional operating on a base model provided by the last layer of the classifier. Inspired by scale-space theory, we theoretically prove that this mapping can be formulated by a convection-diffusion equation, under interpretable and intuitive assumptions from both neural network and PDE perspectives. This theoretically certified framework covers various existing network structures and training techniques, offering a mathematical foundation and new insights into neural networks. Moreover, based on the convection-diffusion equation model, we design a new network structure that incorporates a diffusion mechanism into the network architecture from a PDE perspective. Extensive experiments on benchmark datasets and real-world applications confirm the effectiveness of the proposed model.
Tangjun Wang, Chenglong Bao, Zuoqiang Shi
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 SeNM-VAE: Semi-Supervised Noise Modeling with Hierarchical Variational Autoencoder
abstract
The data bottleneck has emerged as a fundamental challenge in learning based image restoration methods. Researchers have attempted to generate synthesized training data using paired or unpaired samples to address this challenge. This study proposes SeNM-VAE, a semi-supervised noise modeling method that leverages both paired and un-paired datasets to generate realistic degraded data. Our approach is based on modeling the conditional distribution of degraded and clean images with a specially designed graphical model. Under the variational inference framework, we develop an objective function for handling both paired and unpaired data. We employ our method to generate paired training samples for real-world image denoising and super-resolution tasks. Our approach excels in the quality of synthetic degraded images compared to other unpaired and paired noise modeling methods. Furthermore, our approach demonstrates remarkable performance in downstream image restoration tasks, even with limited paired data. With more paired data, our method achieves the best performance on the SIDD dataset.
Dihan Zheng, Yihang Zou, Chenglong Bao
CVPR4
2024 Diffusion Mechanism in Residual Neural Network: Theory and Applications
abstract
Diffusion, a fundamental internal mechanism emerging in many physical processes, describes the interaction among different objects. In many learning tasks with limited training samples, the diffusion connects the labeled and unlabeled data points and is a critical component for achieving high classification accuracy. Many existing deep learning approaches directly impose the fusion loss when training neural networks. In this work, inspired by the convection-diffusion ordinary differential equations (ODEs), we propose a novel diffusion residual network (Diff-ResNet), internally introduces diffusion into the architectures of neural networks. Under the structured data assumption, it is proved that the proposed diffusion block can increase the distance-diameter ratio that improves the separability of inter-class points and reduces the distance among local intra-class points. Moreover, this property can be easily adopted by the residual networks for constructing the separable hyperplanes. Extensive experiments of synthetic binary classification, semi-supervised graph node classification and few-shot image classification in various datasets validate the effectiveness of the proposed method.
Tangjun Wang, Zehao Dou, Chenglong Bao, Zuoqiang Shi
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Averaging Orientations with Molecular Symmetry in Cryo-EM
abstract
Abstract. Cryogenic electron microscopy (cryo-EM) is an invaluable technique for determining high-resolution three-dimensional structures of biological macromolecules using transmission particle images. The inherent symmetry in these macromolecules is advantageous, as it allows each image to represent multiple perspectives. However, data processing that incorporates symmetry can inadvertently average out asymmetric features. Therefore, a key preliminary step is to visualize two-dimensional asymmetric features in the particle images, which requires estimating orientation statistics under molecular symmetry constraints. Motivated by this challenge, we introduce a novel method for estimating the mean and variance of orientations with molecular symmetry. Utilizing tools from nonunique games, we show that our proposed nonconvex formulation can be simplified as a semidefinite programming problem. Moreover, we propose a novel rounding procedure to determine the representative values. Experimental results demonstrate that the proposed approach can find the global minima and the appropriate representatives with a high degree of probability. We release the code of our method as an open-source Python package named pySymStat. Finally, we apply pySymStat to visualize an asymmetric feature in an icosahedral virus, a feat that proved unachievable using the conventional two-dimensional classification method in RELION.
Chenglong Bao, Mingxu Hu
SIAM J. Imaging Sci.2
2024 PhaseNet: A Deep Learning Based Phase Reconstruction Method for Ground-Based Astronomy
abstract
Abstract. Ground-based astronomy utilizes modern telescopes to obtain information on the universe by analyzing recorded signals. Due to atmospheric turbulence, the reconstruction process requires solving a deconvolution problem with an unknown point spread function (PSF). The crucial step in PSF estimation is to obtain a high-resolution phase from low-resolution phase gradients, which is a challenging problem. In this paper, when multiple frames of low-resolution phase gradients are available, we introduce PhaseNet, a deep learning approach based on the Taylor frozen flow hypothesis. Our approach incorporates a data-driven residual regularization term, of which the gradient is parameterized by a network, into the Laplacian regularization based model. To solve the model, we unroll the Nesterov accelerated gradient algorithm so that the network can be efficiently and effectively trained. Finally, we evaluate the performance of PhaseNet under various atmospheric conditions and demonstrate its superiority over TV and Laplacian regularization based methods.
Dihan Zheng, Roland R. Wagner, Ronny Ramlau, Chenglong Bao, Raymond Chan 0001
SIAM J. Imaging Sci.5
2023 Learn From Unpaired Data for Image Restoration: A Variational Bayes Approach
abstract
Collecting paired training data is difficult in practice, but the unpaired samples broadly exist. Current approaches aim at generating synthesized training data from unpaired samples by exploring the relationship between the corrupted and clean data. This work proposes LUD-VAE, a deep generative method to learn the joint probability density function from data sampled from marginal distributions. Our approach is based on a carefully designed probabilistic graphical model in which the clean and corrupted data domains are conditionally independent. Using variational inference, we maximize the evidence lower bound (ELBO) to estimate the joint probability density function. Furthermore, we show that the ELBO is computable without paired samples under the inference invariant assumption. This property provides the mathematical rationale of our approach in the unpaired setting. Finally, we apply our method to real-world image denoising, super-resolution, and low-light image enhancement tasks and train the models using the synthetic data generated by the LUD-VAE. Experimental results validate the advantages of our method over other approaches.
Dihan Zheng, Kaisheng Ma, Chenglong Bao
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 A Class of Short-term Recurrence Anderson Mixing Methods and Their Applications
Fuchao Wei, Chenglong Bao, Yang Liu 0005
ICLR2
2022 A Variant of Anderson Mixing with Minimal Memory Size
abstract
Anderson mixing (AM) is a useful method that can accelerate fixed-point iterations by exploring the information from historical iterations. Despite its numerical success in various applications, the memory requirement in AM remains a bottleneck when solving large-scale optimization problems in a resource-limited machine. To address this problem, we propose a novel variant of AM method, called Min-AM, by storing only one vector pair, that is the minimal memory size requirement in AM. Our method forms a symmetric approximation to the inverse Hessian matrix and is proved to be equivalent to the full-memory Type-I AM for solving strongly convex quadratic optimization. Moreover, for general nonlinear optimization problems, we establish the convergence properties of Min-AM under reasonable assumptions and show that the mixing parameters can be adaptively chosen by estimating the eigenvalues of the Hessian. Finally, we extend Min-AM to solve stochastic programming problems. Experimental results on logistic regression and network training problems validate the effectiveness of the proposed Min-AM.
Fuchao Wei, Chenglong Bao
NeurIPS2
2022 Self-Distillation: Towards Efficient and Compact Neural Networks
abstract
Remarkable achievements have been obtained by deep neural networks in the last several years. However, the breakthrough in neural networks accuracy is always accompanied by explosive growth of computation and parameters, which leads to a severe limitation of model deployment. In this paper, we propose a novel knowledge distillation technique named self-distillation to address this problem. Self-distillation attaches several attention modules and shallow classifiers at different depths of neural networks and distills knowledge from the deepest classifier to the shallower classifiers. Different from the conventional knowledge distillation methods where the knowledge of the teacher model is transferred to another student model, self-distillation can be considered as knowledge transfer in the same model - from the deeper layers to the shallow layers. Moreover, the additional classifiers in self-distillation allow the neural network to work in a dynamic manner, which leads to a much higher acceleration. Experiments demonstrate that self-distillation has consistent and significant effectiveness on various neural networks and datasets. On average, 3.49 and 2.32 percent accuracy boost are observed on CIFAR100 and ImageNet. Besides, experiments show that self-distillation can be combined with other model compression methods, including knowledge distillation, pruning and lightweight model design.
Linfeng Zhang 0001, Chenglong Bao, Kaisheng Ma
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 An Unsupervised Deep Learning Approach for Real-World Image Denoising
Dihan Zheng, Sia Huat Tan, Zuoqiang Shi, Kaisheng Ma, Chenglong Bao
ICLR6
2021 Wavelet J-Net: A Frequency Perspective on Convolutional Neural Networks
abstract
It is well acknowledged in image processing domain that the information can be decomposed into different frequency parts and each part has its own merits. However, existing neural networks always ignore the distinctions and straightforwardly feed all the information into neural networks together, treating them equally. In this paper, we propose a novel neural networks framework named J-Net that decomposes images into different frequency bands and then processes them sequentially. Concretely, the images have been decomposed by wavelet transformation and then the wavelet coefficients are fed into neural networks gradually in different depth according to their decomposition levels. An attention module is utilized to facilitate the fusion of neural network features and injected information, yielding significant performance gain. Furthermore, we show how does the information with different frequency impact the accuracy of neural networks. Experiments show that 5.91%, 5.32% and 2.00% accuracy improvements on Caltech 101, Caltech256 and ImageNet, respectively.
Linfeng Zhang 0001, Xiaoman Zhang, Chenglong Bao, Kaisheng Ma
IJCNN3
2021 AFEC: Active Forgetting of Negative Transfer in Continual Learning
abstract
Continual learning aims to learn a sequence of tasks from dynamic data distributions. Without accessing to the old training samples, knowledge transfer from the old tasks to each new task is difficult to determine, which might be either positive or negative. If the old knowledge interferes with the learning of a new task, i.e., the forward knowledge transfer is negative, then precisely remembering the old tasks will further aggravate the interference, thus decreasing the performance of continual learning. By contrast, biological neural networks can actively forget the old knowledge that conflicts with the learning of a new experience, through regulating the learning-triggered synaptic expansion and synaptic convergence. Inspired by the biological active forgetting, we propose to actively forget the old knowledge that limits the learning of new tasks to benefit continual learning. Under the framework of Bayesian continual learning, we develop a novel approach named Active Forgetting with synaptic Expansion-Convergence (AFEC). Our method dynamically expands parameters to learn each new task and then selectively combines them, which is formally consistent with the underlying mechanism of biological active forgetting. We extensively evaluate AFEC on a variety of continual learning benchmarks, including CIFAR-10 regression tasks, visual classification tasks and Atari reinforcement tasks, where AFEC effectively improves the learning of new tasks and achieves the state-of-the-art performance in a plug-and-play way.
Mingtian Zhang, Zhongfan Jia, Qian Li 0040, Chenglong Bao, Kaisheng Ma, Jun Zhu 0001
NeurIPS5
2021 Stochastic Anderson Mixing for Nonconvex Stochastic Optimization
abstract
Anderson mixing (AM) is an acceleration method for fixed-point iterations. Despite its success and wide usage in scientific computing, the convergence theory of AM remains unclear, and its applications to machine learning problems are not well explored. In this paper, by introducing damped projection and adaptive regularization to the classical AM, we propose a Stochastic Anderson Mixing (SAM) scheme to solve nonconvex stochastic optimization problems. Under mild assumptions, we establish the convergence theory of SAM, including the almost sure convergence to stationary points and the worst-case iteration complexity. Moreover, the complexity bound can be improved when randomly choosing an iterate as the output. To further accelerate the convergence, we incorporate a variance reduction technique into the proposed SAM. We also propose a preconditioned mixing strategy for SAM which can empirically achieve faster convergence or better generalization ability. Finally, we apply the SAM method to train various neural networks including the vanilla CNN, ResNets, WideResNet, ResNeXt, DenseNet and LSTM. Experimental results on image classification and language model demonstrate the advantages of our method.
Fuchao Wei, Chenglong Bao, Yang Liu 0005
NeurIPS2
2020 Light-weight Calibrator: A Separable Component for Unsupervised Domain Adaptation
abstract
Existing domain adaptation methods aim at learning features that can be generalized among domains. These methods commonly require to update source classifier to adapt to the target domain and do not properly handle the trade-off between the source domain and the target domain. In this work, instead of training a classifier to adapt to the target domain, we use a separable component called data calibrator to help the fixed source classifier recover discrimination power in the target domain, while preserving the source domain's performance. When the difference between two domains is small, the source classifier's representation is sufficient to perform well in the target domain and outperforms GAN-based methods in digits. Otherwise, the proposed method can leverage synthetic images generated by GANs to boost performance and achieve state-of-the-art performance in digits datasets and driving scene semantic segmentation. Our method also empirically suggests the potential connection between domain adaptation and adversarial attacks.
Shaokai Ye, Kailu Wu, Mu Zhou, Sia Huat Tan, Kaidi Xu, Jiebo Song, Chenglong Bao, Kaisheng Ma
CVPR8
2020 Auxiliary Training: Towards Accurate and Robust Models
abstract
Training process is crucial for the deployment of the network in applications which have two strict requirements on both accuracy and robustness. However, most existing approaches are in a dilemma, i.e. model accuracy and robustness form an embarrassing tradeoff - the improvement of one leads to the drop of the other. The challenge remains as for we try to improve the accuracy and robustness simultaneously. In this paper, we propose a novel training method via introducing the auxiliary classifiers for training on corrupted samples, while the clean samples are normally trained with the primary classifier. In the training stage, a novel distillation method named input-aware self distillation is proposed to facilitate the primary classifier to learn the robust information from auxiliary classifiers. Along with it, a new normalization method - selective batch normalization is proposed to prevent the model from the negative influence of corrupted images. At the end of training period, a L2-norm penalty is applied to the weights of primary and auxiliary classifiers such that their weights are asymptotically identical. In the stage of inference, only the primary classifier is used and thus no extra computation and storage are needed. Extensive experiments on CIFAR10, CIFAR100 and ImageNet show that noticeable improvements on both accuracy and robustness can be observed by the proposed auxiliary training. On average, auxiliary training achieves 2.21% accuracy and 21.64% robustness (measured by corruption error) improvements over traditional training methods on CIFAR100. Codes has been released on github.
Linfeng Zhang 0001, Muzhou Yu, Zuoqiang Shi, Chenglong Bao, Kaisheng Ma
CVPR5
2020 Interpolation between Residual and Non-Residual Networks
abstract
Although ordinary differential equations (ODEs) provide insights for designing network architectures, its relationship with the non-residual convolutional neural networks (CNNs) is still unclear. In this paper, we present a novel ODE model by adding a damping term. It can be shown that the proposed model can recover both a ResNet and a CNN by adjusting an interpolation coefficient. Therefore, the damped ODE model provides a unified framework for the interpretation of residual and non-residual networks. The Lyapunov analysis reveals better stability of the proposed model, and thus yields robustness improvement of the learned networks. Experiments on a number of image classification benchmarks show that the proposed model substantially improves the accuracy of ResNet and ResNeXt over the perturbed inputs from both stochastic noise and adversarial attack methods. Moreover, the loss landscape analysis demonstrates the improved robustness of our method along the attack direction.
Zonghan Yang, Yang Liu 0005, Chenglong Bao, Zuoqiang Shi
ICML3
2020 Task-Oriented Feature Distillation
abstract
Feature distillation, a primary method in knowledge distillation, always leads to significant accuracy improvements. Most existing methods distill features in the teacher network through a manually designed transformation. In this paper, we propose a novel distillation method named task-oriented feature distillation (TOFD) where the transformation is convolutional layers that are trained in a data-driven manner by task loss. As a result, the task-oriented information in the features can be captured and distilled to students. Moreover, an orthogonal loss is applied to the feature resizing layer in TOFD to improve the performance of knowledge distillation. Experiments show that TOFD outperforms other distillation methods by a large margin on both image classification and 3D classification tasks. Codes have been released in Github.
Linfeng Zhang 0001, Yukang Shi, Zuoqiang Shi, Kaisheng Ma, Chenglong Bao
NeurIPS5
2019 Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation
abstract
Convolutional neural networks have been widely deployed in various application scenarios. In order to extend the applications' boundaries to some accuracy-crucial domains, researchers have been investigating approaches to boost accuracy through either deeper or wider network structures, which brings with them the exponential increment of the computational and storage cost, delaying the responding time. In this paper, we propose a general training framework named self distillation, which notably enhances the performance (accuracy) of convolutional neural networks through shrinking the size of the network rather than aggrandizing it. Different from traditional knowledge distillation - a knowledge transformation methodology among networks, which forces student neural networks to approximate the softmax layer outputs of pre-trained teacher neural networks, the proposed self distillation framework distills knowledge within network itself. The networks are firstly divided into several sections. Then the knowledge in the deeper portion of the networks is squeezed into the shallow ones. Experiments further prove the generalization of the proposed self distillation framework: enhancement of accuracy at average level is 2.65%, varying from 0.61% in ResNeXt as minimum to 4.07% in VGG19 as maximum. In addition, it can also provide flexibility of depth-wise scalable inference on resource-limited edge devices. Our codes have been released on github.
Linfeng Zhang 0001, Jiebo Song, Anni Gao, Chenglong Bao, Kaisheng Ma
ICCV5
2019 SCAN: A Scalable Neural Networks Framework Towards Compact and Efficient Models
abstract
Remarkable achievements have been attained by deep neural networks in various applications. However, the increasing depth and width of such models also lead to explosive growth in both storage and computation, which has restricted the deployment of deep neural networks on resource-limited edge devices. To address this problem, we propose the so-called SCAN framework for networks training and inference, which is orthogonal and complementary to existing acceleration and compression methods. The proposed SCAN firstly divides neural networks into multiple sections according to their depth and constructs shallow classifiers upon the intermediate features of different sections. Moreover, attention modules and knowledge distillation are utilized to enhance the accuracy of shallow classifiers. Based on this architecture, we further propose a threshold controlled scalable inference mechanism to approach human-like sample-specific inference. Experimental results show that SCAN can be easily equipped on various neural networks without any adjustment on hyper-parameters or neural networks architectures, yielding significant performance gain on CIFAR100 and ImageNet. Codes will be released on github soon.
Linfeng Zhang 0001, Zhanhong Tan, Jiebo Song, Chenglong Bao, Kaisheng Ma
NeurIPS5
2019 Barzilai-Borwein-based adaptive learning rate for deep learning
Jinxiu Liang, Yong Xu 0007, Chenglong Bao, Yuhui Quan, Hui Ji 0002
Pattern Recognit. Lett.3
2019 Whole Brain Susceptibility Mapping Using Harmonic Incompatibility Removal
abstract
Quantitative susceptibility mapping (QSM) uses the phase data in magnetic resonance signals to visualize a three-dimensional susceptibility distribution by solving the magnetic field to susceptibility inverse problem. Due to the presence of zeros of the integration kernel in the frequency domain, QSM is an ill-posed inverse problem. Although numerous regularization-based models have been proposed to overcome this problem, incompatibility in the field data, which leads to deterioration of the recovery, has not received enough attention. In this paper, we show that the data acquisition process of QSM inherently generates a harmonic incompatibility in the measured local field. Based on this discovery, we propose a novel regularization-based susceptibility reconstruction model with an additional sparsity-based regularization term on the harmonic incompatibility. Numerical experiments show that the proposed method achieves better performance than existing approaches.
Chenglong Bao, Jae Kyu Choi, Bin Dong 0001
SIAM J. Imaging Sci.1
2018 Coherence Retrieval Using Trace Regularization
abstract
The mutual intensity and its equivalent phase-space representations quantify an optical field's state of coherence and are important tools in the study of light propagation and dynamics, but they can only be estimated indirectly from measurements through a process called coherence retrieval, otherwise known as phase-space tomography. As practical considerations often rule out the availability of a complete set of measurements, coherence retrieval is usually a challenging high-dimensional ill-posed inverse problem. In this paper, we propose a trace-regularized optimization model for coherence retrieval and a provably convergent adaptive accelerated proximal gradient algorithm for solving the resulting problem. Applying our model and algorithm to both simulated and experimental data, we demonstrate an improvement in reconstruction quality over previous models as well as an increase in convergence speed compared to existing first-order methods.
Chenglong Bao, George Barbastathis, Hui Ji 0002, Zuowei Shen, Zhengyun Zhang
SIAM J. Imaging Sci.1
2018 PET-MRI Joint Reconstruction by Joint Sparsity Based Tight Frame Regularization
abstract
Recent technical advances lead to the coupling of PET and MRI scanners, enabling one to acquire functional and anatomical data simultaneously. In this paper, we propose a tight frame based PET-MRI joint reconstruction model via the joint sparsity of tight frame coefficients. In addition, a nonconvex balanced approach is adopted to take the different regularities of PET and MRI images into account. To solve the nonconvex and nonsmooth model, a proximal alternating minimization algorithm is proposed, and the global convergence is present based on the Kurdyka--Łojasiewicz property. Finally, the numerical experiments show that our proposed models achieve better performance over the existing PET-MRI joint reconstruction models.
Jae Kyu Choi, Chenglong Bao, Xiaoqun Zhang
SIAM J. Imaging Sci.2
2016 Equiangular Kernel Dictionary Learning with Applications to Dynamic Texture Analysis
abstract
Most existing dictionary learning algorithms consider a linear sparse model, which often cannot effectively characterize the nonlinear properties present in many types of visual data, e.g. dynamic texture (DT). Such nonlinear properties can be exploited by the so-called kernel sparse coding. This paper proposed an equiangular kernel dictionary learning method with optimal mutual coherence to exploit the nonlinear sparsity of high-dimensional visual data. Two main issues are addressed in the proposed method: (1) coding stability for redundant dictionary of infinite-dimensional space, and (2) computational efficiency for computing kernel matrix of training samples of high-dimensional data. The proposed kernel sparse coding method is applied to dynamic texture analysis with both local DT pattern extraction and global DT pattern characterization. The experimental results showed its performance gain over existing methods.
Yuhui Quan, Chenglong Bao, Hui Ji 0002
CVPR2
2016 Dictionary Learning for Sparse Coding: Algorithms and Convergence Analysis
abstract
In recent years, sparse coding has been widely used in many applications ranging from image processing to pattern recognition. Most existing sparse coding based applications require solving a class of challenging non-smooth and non-convex optimization problems. Despite the fact that many numerical methods have been developed for solving these problems, it remains an open problem to find a numerical method which is not only empirically fast, but also has mathematically guaranteed strong convergence. In this paper, we propose an alternating iteration scheme for solving such problems. A rigorous convergence analysis shows that the proposed method satisfies the global convergence property: the whole sequence of iterates is convergent and converges to a critical point. Besides the theoretical soundness, the practical benefit of the proposed method is validated in applications including image restoration and recognition. Experiments show that the proposed method achieves similar results with less computation when compared to widely used methods such as K-SVD.
Chenglong Bao, Hui Ji 0002, Yuhui Quan, Zuowei Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 L0 Norm Based Dictionary Learning by Proximal Methods with Global Convergence
abstract
Sparse coding and dictionary learning have seen their applications in many vision tasks, which usually is formulated as a non-convex optimization problem. Many iterative methods have been proposed to tackle such an optimization problem. However, it remains an open problem to have a method that is not only practically fast but also is globally convergent. In this paper, we proposed a fast proximal method for solving ℓ0norm based dictionary learning problems, and we proved that the whole sequence generated by the proposed method converges to a stationary point with sub-linear convergence rate. The benefit of having a fast and convergent dictionary learning method is demonstrated in the applications of image recovery and face recognition.
Chenglong Bao, Hui Ji 0002, Yuhui Quan, Zuowei Shen
CVPR1
2014 A Convergent Incoherent Dictionary Learning Algorithm for Sparse Coding
Chenglong Bao, Yuhui Quan, Hui Ji 0002
ECCV (6)1
2013 Fast Sparsity-Based Orthogonal Dictionary Learning for Image Restoration
abstract
In recent years, how to learn a dictionary from input images for sparse modelling has been one very active topic in image processing and recognition. Most existing dictionary learning methods consider an over-complete dictionary, e.g. the K-SVD method. Often they require solving some minimization problem that is very challenging in terms of computational feasibility and efficiency. However, if the correlations among dictionary atoms are not well constrained, the redundancy of the dictionary does not necessarily improve the performance of sparse coding. This paper proposed a fast orthogonal dictionary learning method for sparse image representation. With comparable performance on several image restoration tasks, the proposed method is much more computationally efficient than the over-complete dictionary based learning methods.
Chenglong Bao, Jian-Feng Cai 0001, Hui Ji 0002
ICCV1
2012 Real time robust L1 tracker using accelerated proximal gradient approach
abstract
Recently sparse representation has been applied to visual tracker by modeling the target appearance using a sparse approximation over a template set, which leads to the so-called L1 trackers as it needs to solve an ℓ1norm related minimization problem for many times. While these L1 trackers showed impressive tracking accuracies, they are very computationally demanding and the speed bottleneck is the solver to ℓ1norm minimizations. This paper aims at developing an L1 tracker that not only runs in real time but also enjoys better robustness than other L1 trackers. In our proposed L1 tracker, a new ℓ1norm related minimization model is proposed to improve the tracking accuracy by adding an ℓ1norm regularization on the coefficients associated with the trivial templates. Moreover, based on the accelerated proximal gradient approach, a very fast numerical solver is developed to solve the resulting ℓ1norm related minimization problem with guaranteed quadratic convergence. The great running time efficiency and tracking accuracy of the proposed tracker is validated with a comprehensive evaluation involving eight challenging sequences and five alternative state-of-the-art trackers.
Chenglong Bao, Yi Wu 0001, Haibin Ling, Hui Ji 0002
CVPR1