VLDB 2026 Research / reviewers in the wild / expert
Zhihui Zhu
dblp:71/8081
· DBLP profile ↗
71ranked-venue papers
15as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 9 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Theory of computation · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uniform design-embedded predictions of (tetra-)peptide physicochemical propertiesabstractMOTIVATION: Short peptides hold significant promise in drug discovery and materials science due to their biocompatibility, multifunctionality, ease of synthesis, etc. However, accurately predicting their physicochemical properties, a prerequisite for application development, remains a grand challenge due to the sheet quantity of peptides. RESULTS: This study presents an innovative approach integrating uniform design (UD) on the sampling over the whole space with artificial intelligence (AI) on the sampled data to enhance prediction of key physicochemical properties, including aggregation propensity (AP), hydrophilicity (logP), and isoelectric point (pI), within the complete sequence space of tetrapeptides (160 000 sequences). Using UD, we generate 31 distinct peptide datasets, with a consistent amino acid occupation fraction of 5% at each position, thereby creating unbiased training data without any amino acid preferences for training AI models. This work provides comprehensive datasets on the physicochemical properties of all tetrapeptides, develops robust AI-based predictive models, and quantitatively elucidates the relationships between key physicochemical attributes and self-assembly behaviors of short peptides by Shapley Additive Explanations (SHAP) analysis. By integrating the strategic experimental design (i.e. UD), AI modeling, and peptide domain knowledge, our approach facilitates the discovery and optimization of functional peptides, offering new opportunities for peptide-based therapeutic applications. AVAILABILITY AND IMPLEMENTATION: The complete datasets, source code, and pretrained models are made available at the Github repository (https://github.com/JiaqiBenWang/UD-AI-Peptide) and Zenodo (https://doi.org/10.5281/zenodo.17984124). Zhihui Zhu, Huapeng Liu, Haojin Zhou |
Bioinform. | 1 |
| 2026 | Evidence-verifiable intelligent interaction system and defect knowledge graph for operation-and-maintenance diagnosis of high-speed railway infrastructure
Zhihui Zhu, Weiqi Zheng, Jiajia Yan |
Expert Syst. Appl. | 1 |
| 2026 | On the Convergence of Gradient Descent on Learning Transformers With Residual ConnectionsabstractTransformer models have emerged as fundamental tools across various scientific and engineering disciplines, owing to their outstanding performance in diverse applications. Despite this empirical success, the theoretical foundations of Transformers remain relatively underdeveloped, particularly in understanding their training dynamics. Existing research predominantly examines isolated components–such as self-attention mechanisms and feedforward networks–without thoroughly investigating the interdependencies between these components, especially when residual connections are present. In this paper, we aim to bridge this gap by analyzing the convergence behavior of a structurally complete yet single-layer Transformer, comprising self-attention, a feedforward network, and residual connections. We demonstrate that, under appropriate initialization, gradient descent exhibits a linear convergence rate, where the convergence speed is determined by the minimum and maximum singular values of the output matrix from the attention layer. Moreover, our convergence analysis establishes a theoretical characterization of residual connections by showing that they alleviate the ill-conditioning of the attention output matrix, which arises from the low-rank structure induced by the softmax operation, thereby improving optimization stability. Empirical results corroborate our theoretical insights, illustrating the beneficial role of residual connections in promoting convergence stability. Jinxin Zhou, Jiachen Jiang, Zhihui Zhu |
IEEE Signal Process. Lett. | 4 |
| 2025 | The Distributional Reward Critic Framework for Reinforcement Learning Under Perturbed RewardsabstractThe reward signal plays a central role in defining the desired behaviors of agents in reinforcement learning (RL). Rewards collected from realistic environments could be perturbed, corrupted, or noisy due to an adversary, sensor error, or because they come from subjective human feedback. Thus, it is important to construct agents that can learn under such rewards. Existing methodologies for this problem make strong assumptions, including that the perturbation is known in advance, clean rewards are accessible, or that the perturbation preserves the optimal policy. We study a new, more general, class of unknown perturbations, and introduce a distributional reward critic framework for estimating reward distributions and perturbations during training. Our proposed methods are compatible with any RL algorithm. Despite their increased generality, we show that they achieve comparable or better rewards than existing methods in a variety of environments, including those with clean rewards. Under the challenging and generalized perturbations we study, we win/tie the highest return in 44/48 tested settings (compared to 11/48 for the best baseline). Our results broaden and deepen our ability to perform RL in reward-perturbed environments. Zhihui Zhu, Andrew Perrault |
AAAI | 2 |
| 2025 | Tracing Representation Progression: Analyzing and Enhancing Layer-Wise SimilarityabstractAnalyzing the similarity of internal representations within and across different models has been an important technique for understanding the behavior of deep neural networks. Most existing methods for analyzing the similarity between representations of high dimensions, such as those based on Centered Kernel Alignment (CKA), rely on statistical properties of the representations for a set of data points. In this paper, we focus on transformer models and study the similarity of representations between the hidden layers of individual transformers. In this context, we show that a simple sample-wise cosine similarity metric is capable of capturing the similarity and aligns with the complicated CKA. Our experimental results on common transformers reveal that representations across layers are positively correlated, with similarity increasing when layers get closer. We provide a theoretical justification for this phenomenon under the geodesic curve assumption for the learned transformer, a property that may approximately hold for residual networks. We then show that an increase in representation similarity implies an increase in predicted probability when directly applying the last-layer classifier to any hidden layer representation. This offers a justification for {\it saturation events}, where the model's top prediction remains unchanged across subsequent layers, indicating that the shallow layer has already learned the necessary knowledge. We then propose an aligned training method to improve the effectiveness of shallow layer by enhancing the similarity between internal representations, with trained models that enjoy the following properties: (1) more early saturation events, (2) layer-wise accuracies monotonically increase and reveal the minimal depth needed for the given task, (3) when served as multi-exit models, they achieve on-par performance with standard multi-exit architectures which consist of additional classifiers designed for early exiting in shallow layers. To our knowledge, our work is the first to show that one common classifier is sufficient for multi-exit models. We conduct experiments on both vision and NLP tasks to demonstrate the performance of the proposed aligned training. Jiachen Jiang, Jinxin Zhou, Zhihui Zhu |
ICLR | 3 |
| 2025 | Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language ModelsabstractAchieving better alignment between vision embeddings and Large Language Models (LLMs) is crucial for enhancing the abilities of Multimodal LLMs (MLLMs), particularly for recent models that rely on powerful pretrained vision encoders and LLMs. A common approach to connect the pretrained vision encoder and LLM is through a projector applied after the vision encoder. However, the projector is often trained to enable the LLM to generate captions, and hence the mechanism by which LLMs understand each vision token remains unclear. In this work, we first investigate the role of the projector in compressing vision embeddings and aligning them with word embeddings. We show that the projector significantly compresses visual information, removing redundant details while preserving essential elements necessary for the LLM to understand visual content. We then examine patch-level alignment---the alignment between each vision patch and its corresponding semantic words---and propose a $\textit{multi-semantic alignment hypothesis}$. Our analysis indicates that the projector trained by caption loss improves patch-level alignment but only to a limited extent, resulting in weak and coarse alignment. To address this issue, we propose $\textit{patch-aligned training}$ to efficiently enhance patch-level alignment. Our experiments show that patch-aligned training (1) achieves stronger compression capability and improved patch-level alignment, enabling the MLLM to generate higher-quality captions, (2) improves the MLLM's performance by 16% on referring expression grounding tasks, 4% on question-answering tasks, and 3% on modern instruction-following benchmarks when using the same supervised fine-tuning (SFT) setting. The proposed method can be easily extended to other multimodal models. Jiachen Jiang, Jinxin Zhou, Xia Ning, Zhihui Zhu |
NeurIPS | 5 |
| 2025 | Understanding Representation Dynamics of Diffusion Models via Low-Dimensional ModelingabstractDiffusion models, though originally designed for generative tasks, have demonstrated impressive self-supervised representation learning capabilities. A particularly intriguing phenomenon in these models is the emergence of unimodal representation dynamics, where the quality of learned features peaks at an intermediate noise level. In this work, we conduct a comprehensive theoretical and empirical investigation of this phenomenon. Leveraging the inherent low-dimensionality structure of image data, we theoretically demonstrate that the unimodal dynamic emerges when the diffusion model successfully captures the underlying data distribution. The unimodality arises from an interplay between denoising strength and class confidence across noise scales. Empirically, we further show that, in classification tasks, the presence of unimodal dynamics reliably reflects the diffusion model’s generalization: it emerges when the model generate novel images and gradually transitions to a monotonically decreasing curve as the model begins to memorize the training data. Zhihui Zhu, Peng Wang 0098, Qing Qu 0001 |
NeurIPS | 5 |
| 2025 | Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable DataabstractAmong many mysteries behind the success of deep networks lies the exceptional discriminative power of their learned representations as manifested by the intriguing Neural Collapse (NC) phenomenon, where simple feature structures emerge at the last layer of a trained neural network. Prior works on the theoretical understandings of NC have focused on analyzing the optimization landscape of matrix-factorization-like problems by considering the last-layer features as unconstrained free optimization variables and showing that their global minima exhibit NC. In this paper, we show that gradient flow on a two-layer ReLU network for classifying orthogonally separable data provably exhibits NC, thereby advancing prior results in two ways: First, we relax the assumption of unconstrained features, showing the effect of data structure and nonlinear activations on NC characterizations. Second, we reveal the role of the implicit bias of the training dynamics in facilitating the emergence of NC. Hancheng Min, Zhihui Zhu, René Vidal |
NeurIPS | 2 |
| 2025 | Solution and application of two-dimensional seismic wavefield evolution based on physics-informed neural networks
Zhihui Zhu, Zong Wang, Weiqi Zheng |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Simulation and application of two-dimensional seismic wavefields with random sources using Fourier Neural Operator
Zhihui Zhu, Zong Wang, Jiexin Zhong, Weiqi Zheng |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Understanding Deep Representation Learning via Layerwise Feature Compression and DiscriminationabstractOver the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data. However, it remains an open question how deep networks perform hierarchical feature learning across layers. In this work, we attempt to unveil this mystery by investigating the structures of intermediate features. Motivated by our empirical findings that linear layers mimic the roles of deep layers in nonlinear networks for feature learning, we explore how deep linear networks transform input data into output by investigating the output (i.e., features) of each layer after training in the context of multi-class classification problems. Toward this goal, we first define metrics to measure within-class compression and between-class discrimination of intermediate features, respectively. Through theoretical analysis of these two metrics, we show that the evolution of features follows a simple and quantitative pattern from shallow to deep layers when the input data is nearly orthogonal and the network weights are minimum-norm, balanced, and approximately low-rank: each layer of the linear network progressively compresses within-class features at a geometric rate and discriminates between-class features at a linear rate with respect to the number of layers that data have passed through. To the best of our knowledge, this is the first quantitative characterization of feature evolution in hierarchical representations of deep linear networks. Moreover, our extensive experiments not only validate our theoretical results but also reveal a similar pattern in deep nonlinear networks, which aligns well with recent empirical studies. Finally, we demonstrate the practical value of our results in transfer learning. Peng Wang 0098, Can Yaras, Zhihui Zhu, Laura Balzano, Qing Qu 0001 |
J. Mach. Learn. Res. | 4 |
| 2025 | Computational and Statistical Guarantees for Tensor-on-Tensor Regression With Tensor Train DecompositionabstractRecently, a tensor-on-tensor (ToT) regression model has been proposed to generalize tensor recovery, encompassing scenarios like scalar-on-tensor regression and tensor-on-vector regression. However, the exponential growth in tensor complexity poses challenges for storage and computation in ToT regression. To overcome this hurdle, tensor decompositions have been introduced, with the tensor train (TT)-based ToT model proving efficient in practice due to reduced memory requirements, enhanced computational efficiency, and decreased sampling complexity. Despite these practical benefits, a disparity exists between theoretical analysis and real-world performance. In this paper, we delve into the theoretical and algorithmic aspects of the TT-based ToT regression model. Assuming the regression operator satisfies the restricted isometry property (RIP), we conduct an error analysis for the solution to a constrained least-squares optimization problem. This analysis includes upper error bound and minimax lower bound, revealing that such error bounds polynomially depend on the order $N+M$N+M. To efficiently find solutions meeting such error bounds, we propose two optimization algorithms: the iterative hard thresholding (IHT) algorithm (employing gradient descent with TT-singular value decomposition (TT-SVD)) and the factorization approach using the Riemannian gradient descent (RGD) algorithm. When RIP is satisfied, spectral initialization facilitates proper initialization, and we establish the linear convergence rate of both IHT and RGD. Notably, compared to the IHT, which optimizes the entire tensor in each iteration while maintaining the TT structure through TT-SVD and poses a challenge for storage memory in practice, the RGD optimizes factors in the so-called left-orthogonal TT format, enforcing orthonormality among most of the factors, over the Stiefel manifold, thereby reducing the storage complexity of the IHT. However, this reduction in storage memory comes at a cost: the recovery of RGD is worse than that of IHT, while the error bounds of both algorithms depend on $N+M$N+M polynomially. Experimental validation substantiates the validity of our theoretical findings. Zhihui Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | DREAM: Diffusion Rectification and Estimation-Adaptive ModelsabstractWe present DREAM, a novel training framework representing Diffusion Rectification and Estimation-Adaptive Models, requiring minimal code changes (just three lines) yet significantly enhancing the alignment of training with sampling in diffusion models. DREAM features two components: diffusion rectification, which adjusts training to reflect the sampling process, and estimation adaptation, which balances perception against distortion. When applied to image super-resolution (SR), DREAM adeptly navigates the tradeoff between minimizing distortion and preserving high image quality. Experiments demonstrate DREAM's superiority over standard diffusion-based SR methods, showing a 2 to 3× faster training convergence and a 10 to 20× reduction in sampling steps to achieve comparable results. We hope DREAM will inspire a rethinking of diffusion model training paradigms. Our source code is available at link. Jinxin Zhou, Tianyu Ding, Jiachen Jiang, Ilya Zharkov, Zhihui Zhu, Luming Liang |
CVPR | 6 |
| 2024 | A Global Geometric Analysis of Maximal Coding Rate ReductionabstractThe maximal coding rate reduction (MCR$^2$) objective for learning structured and compact deep representations is drawing increasing attention, especially after its recent usage in the derivation of fully explainable and highly effective deep network architectures. However, it lacks a complete theoretical justification: only the properties of its global optima are known, and its global landscape has not been studied. In this work, we give a complete characterization of the properties of all its local and global optima as well as other types of critical points. Specifically, we show that each (local or global) maximizer of the MCR$^2$ problem corresponds to a low-dimensional, discriminative, and diverse representation, and furthermore, each critical point of the objective is either a local maximizer or a strict saddle point. Such a favorable landscape makes MCR$^2$ a natural choice of objective for learning diverse and discriminative representations via first-order optimization. To further verify our theoretical findings, we illustrate these properties with extensive experiments on both synthetic and real data sets. Peng Wang 0098, Huikang Liu, Druv Pai, Yaodong Yu, Zhihui Zhu, Qing Qu 0001, Yi Ma 0001 |
ICML | 5 |
| 2024 | Generalized Neural Collapse for a Large Number of ClassesabstractNeural collapse provides an elegant mathematical characterization of learned last layer representations (a.k.a. features) and classifier weights in deep classification models. Such results not only provide insights but also motivate new techniques for improving practical deep models. However, most of the existing empirical and theoretical studies in neural collapse focus on the case that the number of classes is small relative to the dimension of the feature space. This paper extends neural collapse to cases where the number of classes are much larger than the dimension of feature space, which broadly occur for language models, retrieval systems, and face recognition applications. We show that the features and classifier exhibit a generalized neural collapse phenomenon, where the minimum one-vs-rest margins is maximized. We provide empirical study to verify the occurrence of generalized neural collapse in practical deep neural networks. Moreover, we provide theoretical study to show that the generalized neural collapse provably occurs under unconstrained feature model with spherical constraint, under certain technical conditions on feature dimension and number of classes. Jiachen Jiang, Jinxin Zhou, Peng Wang 0098, Qing Qu 0001, Dustin G. Mixon, Chong You, Zhihui Zhu |
ICML | 7 |
| 2024 | Quantum Exploration-based Reinforcement Learning for Efficient Robot Path Planning in Sparse-Reward EnvironmentabstractWith the latest developments in sensors, battery, and Artificial Intelligence (AI) technologies, robots can perform missions in unstructured environments (e.g., disaster sites). However, their adaptation speed is still slow due to their limited onboard computing capability and dynamic disruption from environments, leading to inaccurate planning and delayed emergency reactions. Even though reinforcement learning increases robots’ exploration speed by fusing effective guidance, the numerous interactions make the learning time-consuming and even risky (e.g., collisions). To fundamentally improve robot adaptation speed, this work seeks help from quantum power. A novel Quantum Exploration based Dreamer model (QED) was developed to facilitate reinforcement learning explorations. QED based on stochastic quantum walker quickly explores environments, evaluates action quality, and obtains a global optimal exploration strategy; then, these high-quality exploration samples will be used to facilitate reinforcement learning speed. A theoretical benefit of QED is facilitating reinforcement learning speed without changing the underlying learning architecture, which makes the proposed QED applicable to general robot learning scenarios. To validate QED effectiveness, a robot path-planning task in an obstacle-dense environment was designed. The number of needed training episodes validated the effectiveness. The results show that QED achieves around ten times faster policy learning compared with Monte Carlo tree-based reinforcement learning methods and vanilla reinforcement learning methods. Yibei Guo, Zhihui Zhu, Mei Si 0001, Daniel Blankenberg |
RO-MAN | 3 |
| 2024 | Guaranteed Nonconvex Factorization Approach for Tensor Train RecoveryabstractTensor train (TT) decomposition represents an order-$N$ tensor using $O(N)$ order-$3$ tensors (i.e., factors of small dimension), achieved through products among these factors. Due to its compact representation, TT decomposition has been widely used in the fields of signal processing, machine learning, and quantum physics. It offers benefits such as reduced memory requirements, enhanced computational efficiency, and decreased sampling complexity. Nevertheless, existing optimization algorithms with guaranteed performance concentrate exclusively on using the TT format for reducing the optimization space in recovery problems, while still operating on the entire tensor in each iteration. There is a lack of comprehensive theoretical analysis for optimization involving the factors directly, despite the proven efficacy of such factorization methods in practice. In this paper, we provide the first convergence guarantee for the factorization approach in a TT-based recovery problem. Specifically, to avoid the scaling ambiguity and to facilitate theoretical analysis, we optimize over the so-called left-orthogonal TT format which enforces orthonormality among most of the factors. To ensure the orthonormal structure, we utilize the Riemannian gradient descent (RGD) for optimizing those factors over the Stiefel manifold. We first delve into the TT factorization/decomposition problem and establish the local linear convergence of RGD. Notably, the rate of convergence only experiences a linear decline as the tensor order increases. We then study the sensing problem that aims to recover a TT format tensor from linear measurements. Assuming the sensing operator satisfies the restricted isometry property (RIP), we show that with a proper initialization, which could be obtained through spectral initialization, RGD also converges to the ground-truth tensor at a linear rate. Furthermore, we expand our analysis to encompass scenarios involving Gaussian noise in the measurements. We prove that RGD can reliably recover the ground truth at a linear rate, with the recovery error exhibiting only polynomial growth in relation to the tensor order $N$. We conduct various experiments to validate our theoretical findings. Michael B. Wakin, Zhihui Zhu |
J. Mach. Learn. Res. | 3 |
| 2024 | Convergence Analysis for Learning Orthonormal Deep Linear Neural NetworksabstractEnforcing orthonormal or isometric property for the weight matrices has been shown to enhance the training of deep neural networks by mitigating gradient exploding/vanishing and increasing the robustness of the learned networks. However, despite its practical performance, the theoretical analysis of orthonormality in neural networks is still lacking; for example, how orthonormality affects the convergence of the training process. In this letter, we aim to bridge this gap by providing convergence analysis for training orthonormal deep linear neural networks. Specifically, we show that Riemannian gradient descent with an appropriate initialization converges at a linear rate for training orthonormal deep linear neural networks with a class of loss functions. Unlike existing works that enforce orthonormal weight matrices for all the layers, our approach excludes this requirement for one layer, which is crucial to establish the convergence guarantee. Our results shed light on how increasing the number of hidden layers can impact the convergence speed. Experimental results validate our theoretical analysis. Xuwei Tan, Zhihui Zhu |
IEEE Signal Process. Lett. | 3 |
| 2024 | Quantum State Tomography for Matrix Product Density OperatorsabstractThe reconstruction of quantum states from experimental measurements, often achieved using quantum state tomography (QST), is crucial for the verification and benchmarking of quantum devices. However, performing QST for a generic unstructured quantum state requires an enormous number of state copies that growsexponentiallywith the number of individual quanta in the system, even for the most optimal measurement settings. Fortunately, many physical quantum states, such as states generated by noisy, intermediate-scale quantum computers, are usually structured. In one dimension, such states are expected to be well approximated by matrix product operators (MPOs) with a matrix/bond dimension independent of the number of qubits, therefore enabling efficient state representation. Nevertheless, it is still unclear whether efficient QST can be performed for these states in general. In other words, there exist no rigorous bounds on the number of state copies required for reconstructing MPO states that scales polynomially with the number of qubits. In this paper, we attempt to bridge this gap and establish theoretical guarantees for the stable recovery of MPOs using tools from compressive sensing and the theory of empirical processes. We begin by studying two types of random measurement settings: Gaussian measurements and Haar random projective measurements. We show that the information contained in an MPO with a constant bond dimension can be preserved using a number of random measurements that depends onlylinearlyon the number of qubits, assuming no statistical error of the measurements. We then study MPO-based QST with Haar random projective measurements that can in principle be implemented on quantum computers. We prove that only apolynomialnumber of state copies in the number of qubits is required to guarantee bounded recovery error of an MPO state. Remarkably, such recovery can be achieved by measuring the state in each random basis only once, despite the large statistical error associated with the outcome of each measurement. Our work may be generalized to accommodate random local or t-design measurements that are more practical to implement on current quantum computers. It may also facilitate the discovery of efficient QST methods for other structured quantum states. Casey Jameson, Zhexuan Gong, Michael B. Wakin, Zhihui Zhu |
IEEE Trans. Inf. Theory | 5 |
| 2023 | OTOv2: Automatic, Generic, User-Friendly
Luming Liang, Tianyu Ding, Zhihui Zhu, Ilya Zharkov |
ICLR | 4 |
| 2023 | Seismic Data Reconstruction and Denoising by Enhanced Hankel Low-Rank Matrix EstimationabstractSeismic data reconstruction and denoising play a fundamental role in most seismic data processing algorithms which are often designed for regularly sampled and reliable data. By using the fact that the (block) Hankel matrix formulated from clean seismic data is low-rank if it corresponds to a few linear events, the low-rank based approach has been successfully used for seismic data reconstruction and denoising. In this paper, we simultaneously exploit the Hankel and low-rank structures within the (block) Hankel matrices of the clear and complete seismic data and formulate the problem of seismic data reconstruction and denoising as a Hankel low-rank reconstruction problem. The additional Hankel structure can further improve the performance for reconstruction and denoising. We then propose an iterative algorithm for solving the corresponding Hankel low-rank reconstruction problem. The algorithm is based on the alternating direction method of multipliers (ADMM) and has closed-form solutions for each update. Experiments on both synthetical and real seismic data demonstrate the superior performance of the proposed algorithm compared with conventional low-rank based methods for simultaneous seismic data reconstruction and denoising. Chong Wang 0020, Zhiyuan Gu, Zhihui Zhu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Provable Splitting Approach for Symmetric Nonnegative Matrix FactorizationabstractThe symmetric Nonnegative Matrix Factorization (NMF), a special but important class of the general NMF, has found numerous applications in data analysis such as various clustering tasks. Unfortunately, designing fast algorithms for the symmetric NMF is not as easy as for its nonsymmetric counterpart, since the latter admits the splitting property that allows state-of-the-art alternating-type algorithms. To overcome this issue, we first split the decision variable and transform the symmetric NMF to a penalized nonsymmetric one, paving the way for designing efficient alternating-type algorithms. We then show that solving the penalized nonsymmetric reformulation returns a solution to the original symmetric NMF. Moreover, we design a family of alternating-type algorithms and show that they all admit strong convergence guarantee: the generated sequence of iterates is convergent and converges at least sublinearly to a critical point of the original symmetric NMF. Finally, we conduct experiments on both synthetic data and real image clustering to support our theoretical results and demonstrate the performance of the alternating-type algorithms. Xiao Li 0009, Zhihui Zhu, Qiuwei Li, Kai Liu 0018 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Learning Approach For Fast Approximate Matrix FactorizationsabstractEfficiently computing an (approximate) orthonormal basis and low-rank approximation for the input data X plays a crucial role in data analysis. One of the most efficient algorithms for such tasks is the randomized algorithm, which proceeds by computing a projection XA with a random sketching matrix A of much smaller size, and then computing the orthonormal basis as well as low-rank factorizations of the tall matrix XA. While a random matrix A is the de facto choice, in this work, we improve upon its performance by utilizing a learning approach to find an adaptive sketching matrix A from a set of training data. We derive a closed-form formulation for the gradient of the training problem, enabling us to use efficient gradient-based algorithms. We also extend this approach for learning structured sketching matrix, such as the sparse sketching matrix that performs as selecting a few number of representative columns from the input data. Our experiments on both synthetical and real data show that both learned dense and sparse sketching matrices outperform the random ones in finding the approximate orthonormal basis and low-rank approximations. Zhihui Zhu |
ICASSP | 3 |
| 2022 | Robust Training under Label Noise by Over-parameterizationabstractRecently, over-parameterized deep networks, with increasingly more network parameters than training samples, have dominated the performances of modern machine learning. However, when the training data is corrupted, it has been well-known that over-parameterized networks tend to overfit and do not generalize. In this work, we propose a principled approach for robust training of over-parameterized deep networks in classification tasks where a proportion of training labels are corrupted. The main idea is yet very simple: label noise is sparse and incoherent with the network learned from clean data, so we model the noise and learn to separate it from the data. Specifically, we model the label noise via another sparse over-parameterization term, and exploit implicit algorithmic regularizations to recover and separate the underlying corruptions. Remarkably, when trained using such a simple method in practice, we demonstrate state-of-the-art test accuracy against label noise on a variety of real datasets. Furthermore, our experimental results are corroborated by theory on simplified linear models, showing that exact separation between sparse noise and low-rank data can be achieved under incoherent conditions. The work opens many interesting directions for improving over-parameterized models by using sparse over-parameterization and implicit regularization. Code is available at https://github.com/shengliu66/SOP. Zhihui Zhu, Qing Qu 0001, Chong You |
ICML | 2 |
| 2022 | On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained FeaturesabstractWhen training deep neural networks for classification tasks, an intriguing empirical phenomenon has been widely observed in the last-layer classifiers and features, where (i) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and (ii) cross-example within-class variability of last-layer activations collapses to zero. This phenomenon is called Neural Collapse (NC), which seems to take place regardless of the choice of loss functions. In this work, we justify NC under the mean squared error (MSE) loss, where recent empirical evidence shows that it performs comparably or even better than the de-facto cross-entropy loss. Under a simplified unconstrained feature model, we provide the first global landscape analysis for vanilla nonconvex MSE loss and show that the (only!) global minimizers are neural collapse solutions, while all other critical points are strict saddles whose Hessian exhibit negative curvature directions. Furthermore, we justify the usage of rescaled MSE loss by probing the optimization landscape around the NC solutions, showing that the landscape can be improved by tuning the rescaling hyperparameters. Finally, our theoretical findings are experimentally verified on practical network architectures. Jinxin Zhou, Xiao Li 0026, Tianyu Ding, Chong You, Qing Qu 0001, Zhihui Zhu |
ICML | 6 |
| 2022 | Revisiting Sparse Convolutional Model for Visual RecognitionabstractDespite strong empirical performance for image classification, deep neural networks are often regarded as ``black boxes'' and they are difficult to interpret. On the other hand, sparse convolutional models, which assume that a signal can be expressed by a linear combination of a few elements from a convolutional dictionary, are powerful tools for analyzing natural images with good theoretical interpretability and biological plausibility. However, such principled models have not demonstrated competitive performance when compared with empirically designed deep networks. This paper revisits the sparse convolutional modeling for image classification and bridges the gap between good empirical performance (of deep learning) and good interpretability (of sparse convolutional models). Our method uses differentiable optimization layers that are defined from convolutional sparse coding as drop-in replacements of standard convolutional layers in conventional deep neural networks. We show that such models have equally strong empirical performance on CIFAR-10, CIFAR-100 and ImageNet datasets when compared to conventional neural networks. By leveraging stable recovery property of sparse modeling, we further show that such models can be much more robust to input corruptions as well as adversarial perturbations in testing through a simple proper trade-off between sparse regularization and data reconstruction terms. Xili Dai, Pengyuan Zhai, Shengbang Tong, Xingjian Gao, Shao-Lun Huang, Zhihui Zhu, Chong You, Yi Ma 0001 |
NeurIPS | 7 |
| 2022 | Error Analysis of Tensor-Train Cross ApproximationabstractTensor train decomposition is widely used in machine learning and quantum physics due to its concise representation of high-dimensional tensors, overcoming the curse of dimensionality. Cross approximation---originally developed for representing a matrix from a set of selected rows and columns---is an efficient method for constructing a tensor train decomposition of a tensor from few of its entries. While tensor train cross approximation has achieved remarkable performance in practical applications, its theoretical analysis, in particular regarding the error of the approximation, is so far lacking. To our knowledge, existing results only provide element-wise approximation accuracy guarantees, which lead to a very loose bound when extended to the entire tensor. In this paper, we bridge this gap by providing accuracy guarantees in terms of the entire tensor for both exact and noisy measurements. Our results illustrate how the choice of selected subtensors affects the quality of the cross approximation and that the approximation error caused by model error and/or measurement error may not grow exponentially with the order of the tensor. These results are verified by numerical experiments, and may have important implications for the usefulness of cross approximations for high-order tensors, such as those encountered in the description of quantum many-body states. Alexander Lidiak, Zhexuan Gong, Gongguo Tang, Michael B. Wakin, Zhihui Zhu |
NeurIPS | 6 |
| 2022 | Neural Collapse with Normalized Features: A Geometric Analysis over the Riemannian ManifoldabstractWhen training overparameterized deep networks for classification tasks, it has been widely observed that the learned features exhibit a so-called "neural collapse'" phenomenon. More specifically, for the output features of the penultimate layer, for each class the within-class features converge to their means, and the means of different classes exhibit a certain tight frame structure, which is also aligned with the last layer's classifier. As feature normalization in the last layer becomes a common practice in modern representation learning, in this work we theoretically justify the neural collapse phenomenon under normalized features. Based on an unconstrained feature model, we simplify the empirical loss function in a multi-class classification task into a nonconvex optimization problem over the Riemannian manifold by constraining all features and classifiers over the sphere. In this context, we analyze the nonconvex landscape of the Riemannian optimization problem over the product of spheres, showing a benign global landscape in the sense that the only global minimizers are the neural collapse solutions while all other critical points are strict saddle points with negative curvature. Experimental results on practical deep networks corroborate our theory and demonstrate that better representations can be learned faster via feature normalization. Code for our experiments can be found at https://github.com/cjyaras/normalized-neural-collapse. Can Yaras, Peng Wang 0098, Zhihui Zhu, Laura Balzano, Qing Qu 0001 |
NeurIPS | 3 |
| 2022 | Are All Losses Created Equal: A Neural Collapse PerspectiveabstractWhile cross entropy (CE) is the most commonly used loss function to train deep neural networks for classification tasks, many alternative losses have been developed to obtain better empirical performance. Among them, which one is the best to use is still a mystery, because there seem to be multiple factors affecting the answer, such as properties of the dataset, the choice of network architecture, and so on. This paper studies the choice of loss function by examining the last-layer features of deep networks, drawing inspiration from a recent line work showing that the global optimal solution of CE and mean-square-error (MSE) losses exhibits a Neural Collapse phenomenon. That is, for sufficiently large networks trained until convergence, (i) all features of the same class collapse to the corresponding class mean and (ii) the means associated with different classes are in a configuration where their pairwise distances are all equal and maximized. We extend such results and show through global solution and landscape analyses that a broad family of loss functions including commonly used label smoothing (LS) and focal loss (FL) exhibits Neural Collapse. Hence, all relevant losses (i.e., CE, LS, FL, MSE) produce equivalent features on training data. In particular, based on the unconstrained feature model assumption, we provide either the global landscape analysis for LS loss or the local landscape analysis for FL loss and show that the (only!) global minimizers are neural collapse solutions, while all other critical points are strict saddles whose Hessian exhibit negative curvature directions either in the global scope for LS loss or in the local scope for FL loss near the optimal solution. The experiments further show that Neural Collapse features obtained from all relevant losses (i.e., CE, LS, FL, MSE) lead to largely identical performance on test data as well, provided that the network is sufficiently large and trained until convergence. Jinxin Zhou, Chong You, Xiao Li 0026, Kangning Liu, Qing Qu 0001, Zhihui Zhu |
NeurIPS | 7 |
| 2022 | Recovery and Generalization in Over-Realized Dictionary LearningabstractIn over two decades of research, the field of dictionary learning has gathered a large collection of successful applications, and theoretical guarantees for model recovery are known only whenever optimization is carried out in the same model class as that of the underlying dictionary. This work characterizes the surprising phenomenon that dictionary recovery can be facilitated by searching over the space of larger over-realized models. This observation is general and independent of the specific dictionary learning algorithm used. We thoroughly demonstrate this observation in practice and provide an analysis of this phenomenon by tying recovery measures to generalization bounds. In particular, we show that model recovery can be upper-bounded by the empirical risk, a model-dependent quantity and the generalization gap, reflecting our empirical findings. We further show that an efficient and provably correct distillation approach can be employed to recover the correct atoms from the over-realized model. As a result, our meta-algorithm provides dictionary estimates with consistently better recovery of the ground-truth model. Jeremias Sulam, Chong You, Zhihui Zhu |
J. Mach. Learn. Res. | 3 |
| 2021 | Dual Principal Component Pursuit for Learning a Union of Hyperplanes: Theory and AlgorithmsabstractState-of-the-art subspace clustering methods are based on convex formulations whose theoretical guarantees require the subspaces to be low-dimensional. Dual Principal Component Pursuit (DPCP) is a non-convex method that is specifically designed for learning high-dimensional subspaces, such as hyperplanes. However, existing analyses of DPCP in the multi-hyperplane case lack a precise characterization of the distribution of the data and involve quantities that are difficult to interpret. Moreover, the provable algorithm based on recursive linear programming is not efficient. In this paper, we introduce a new notion of geometric dominance, which explicitly captures the distribution of the data, and derive both geometric and probabilistic conditions under which a global solution to DPCP is a normal vector to a geometrically dominant hyperplane. We then prove that the DPCP problem for a union of hyperplanes satisfies a Riemannian regularity condition, and use this result to show that a scalable Riemannian subgradient method exhibits (local) linear convergence to the normal vector of the geometrically dominant hyperplane. Finally, we show that integrating DPCP into popular subspace clustering schemes, such as K-ensembles, leads to superior or competitive performance over the state-of-the-art in clustering hyperplanes. Tianyu Ding, Zhihui Zhu, Manolis C. Tsakiris, René Vidal, Daniel P. Robinson |
AISTATS | 2 |
| 2021 | CDFI: Compression-Driven Network Design for Frame InterpolationabstractDNN-based frame interpolation—that generates the intermediate frames given two consecutive frames—typically relies on heavy model architectures with a huge number of features, preventing them from being deployed on systems with limited resources, e.g., mobile devices. We propose a compression-driven network design for frame interpolation (CDFI), that leverages model pruning through sparsity-inducing optimization to significantly reduce the model size while achieving superior performance. Concretely, we first compress the recently proposed AdaCoF model and show that a 10× compressed AdaCoF performs similarly as its original counterpart; then we further improve this compressed model by introducing a multi-resolution warping module, which boosts visual consistencies with multi-level details. As a consequence, we achieve a significant performance gain with only a quarter in size compared with the original AdaCoF. Moreover, our model performs favorably against other state-of-the-arts in a broad range of datasets. Finally, the proposed compression-driven framework is generic and can be easily transferred to other DNN-based frame interpolation algorithm. Our source code is available at https://github.com/tding1/CDFI. Tianyu Ding, Luming Liang, Zhihui Zhu, Ilya Zharkov |
CVPR | 3 |
| 2021 | Dual Principal Component Pursuit for Robust Subspace Learning: Theory and Algorithms for a Holistic ApproachabstractThe Dual Principal Component Pursuit (DPCP) method has been proposed to robustly recover a subspace of high-relative dimension from corrupted data. Existing analyses and algorithms of DPCP, however, mainly focus on finding a normal to a single hyperplane that contains the inliers. Although these algorithms can be extended to a subspace of higher co-dimension through a recursive approach that sequentially finds a new basis element of the space orthogonal to the subspace, this procedure is computationally expensive and lacks convergence guarantees. In this paper, we consider a DPCP approach for simultaneously computing the entire basis of the orthogonal complement subspace (we call this a holistic approach) by solving a non-convex non-smooth optimization problem over the Grassmannian. We provide geometric and statistical analyses for the global optimality and prove that it can tolerate as many outliers as the square of the number of inliers, under both noiseless and noisy settings. We then present a Riemannian regularity condition for the problem, which is then used to prove that a Riemannian subgradient method converges linearly to a neighborhood of the orthogonal subspace with error proportional to the noise level. Tianyu Ding, Zhihui Zhu, René Vidal, Daniel P. Robinson |
ICML | 2 |
| 2021 | Only Train Once: A One-Shot Neural Network Training And Pruning FrameworkabstractStructured pruning is a commonly used technique in deploying deep neural networks (DNNs) onto resource-constrained devices. However, the existing pruning methods are usually heuristic, task-specified, and require an extra fine-tuning procedure. To overcome these limitations, we propose a framework that compresses DNNs into slimmer architectures with competitive performances and significant FLOPs reductions by Only-Train-Once (OTO). OTO contains two key steps: (i) we partition the parameters of DNNs into zero-invariant groups, enabling us to prune zero groups without affecting the output; and (ii) to promote zero groups, we then formulate a structured-sparsity optimization problem, and propose a novel optimization algorithm, Half-Space Stochastic Projected Gradient (HSPG), to solve it, which outperforms the standard proximal methods on group sparsity exploration, and maintains comparable convergence. To demonstrate the effectiveness of OTO, we train and compress full models simultaneously from scratch without fine-tuning for inference speedup and parameter reduction, and achieve state-of-the-art results on VGG16 for CIFAR10, ResNet50 for CIFAR10 and Bert for SQuAD and competitive result on ResNet50 for ImageNet. The source code is available at https://github.com/tianyic/onlytrainonce. Bo Ji 0003, Tianyu Ding, Biyi Fang, Guanyi Wang, Zhihui Zhu, Luming Liang, Yixin Shi, Xiao Tu |
NeurIPS | 6 |
| 2021 | Rank Overspecified Robust Matrix Recovery: Subgradient Method and Exact RecoveryabstractWe study the robust recovery of a low-rank matrix from sparsely and grossly corrupted Gaussian measurements, with no prior knowledge on the intrinsic rank. We consider the robust matrix factorization approach. We employ a robust $\ell_1$ loss function and deal with the challenge of the unknown rank by using an overspecified factored representation of the matrix variable. We then solve the associated nonconvex nonsmooth problem using a subgradient method with diminishing stepsizes. We show that under a regularity condition on the sensing matrices and corruption, which we call restricted direction preserving property (RDPP), even with rank overspecified, the subgradient method converges to the exact low-rank solution at a sublinear rate. Moreover, our result is more general in the sense that it automatically speeds up to a linear rate once the factor rank matches the unknown rank. On the other hand, we show that the RDPP condition holds under generic settings, such as Gaussian measurements under independent or adversarial sparse corruptions, where the result could be of independent interest. Both the exact recovery and the convergence rate of the proposed subgradient method are numerically verified in the overspecified regime. Moreover, our experiment further shows that our particular design of diminishing stepsize effectively prevents overfitting for robust recovery under overparameterized models, such as robust matrix sensing and learning robust deep image prior. This regularization effect is worth further investigation. Lijun Ding, Yudong Chen 0001, Qing Qu 0001, Zhihui Zhu |
NeurIPS | 5 |
| 2021 | Convolutional Normalization: Improving Deep Convolutional Network Robustness and TrainingabstractNormalization techniques have become a basic component in modern convolutional neural networks (ConvNets). In particular, many recent works demonstrate that promoting the orthogonality of the weights helps train deep models and improve robustness. For ConvNets, most existing methods are based on penalizing or normalizing weight matrices derived from concatenating or flattening the convolutional kernels. These methods often destroy or ignore the benign convolutional structure of the kernels; therefore, they are often expensive or impractical for deep ConvNets. In contrast, we introduce a simple and efficient ``Convolutional Normalization'' (ConvNorm) method that can fully exploit the convolutional structure in the Fourier domain and serve as a simple plug-and-play module to be conveniently incorporated into any ConvNets. Our method is inspired by recent work on preconditioning methods for convolutional sparse coding and can effectively promote each layer's channel-wise isometry. Furthermore, we show that our ConvNorm can reduce the layerwise spectral norm of the weight matrices and hence improve the Lipschitzness of the network, leading to easier training and improved robustness for deep ConvNets. Applied to classification under noise corruptions and generative adversarial network (GAN), we show that the ConvNorm improves the robustness of common ConvNets such as ResNet and the performance of GAN. We verify our findings via numerical experiments on CIFAR and ImageNet. Our implementation is available online at \url{https://github.com/shengliu66/ConvNorm}. Xiao Li 0026, Yuexiang Zhai, Chong You, Zhihui Zhu, Carlos Fernandez-Granda, Qing Qu 0001 |
NeurIPS | 5 |
| 2021 | A Geometric Analysis of Neural Collapse with Unconstrained FeaturesabstractWe provide the first global optimization landscape analysis of Neural Collapse -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of neural networks during the terminal phase of training. As recently reported by Papyan et al., this phenomenon implies that (i) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and (ii) cross-example within-class variability of last-layer activations collapses to zero. We study the problem based on a simplified unconstrained feature model, which isolates the topmost layers from the classifier of the neural network. In this context, we show that the classical cross-entropy loss with weight decay has a benign global landscape, in the sense that the only global minimizers are the Simplex ETFs while all other critical points are strict saddles whose Hessian exhibit negative curvature directions. Our analysis of the simplified model not only explains what kind of features are learned in the last layer, but also shows why they can be efficiently optimized, matching the empirical observations in practical deep network architectures. These findings provide important practical implications. As an example, our experiments demonstrate that one may set the feature dimension equal to the number of classes and fix the last-layer classifier to be a Simplex ETF for network training, which reduces memory cost by over 20% on ResNet18 without sacrificing the generalization performance. The source code is available at https://github.com/tding1/Neural-Collapse. Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li 0026, Chong You, Jeremias Sulam, Qing Qu 0001 |
NeurIPS | 1 |
| 2021 | Error-Tolerant Deep Learning for Remote Sensing Image Scene ClassificationabstractDue to its various application potentials, the remote sensing image scene classification (RSSC) has attracted a broad range of interests. While the deep convolutional neural network (CNN) has recently achieved tremendous success in RSSC, its superior performances highly depend on a large number of accurately labeled samples which require lots of time and manpower to generate for a large-scale remote sensing image scene dataset. In contrast, it is not only relatively easy to collect coarse and noisy labels but also inevitable to introduce label noise when collecting large-scale annotated data in the remote sensing scenario. Therefore, it is of great practical importance to robustly learn a superior CNN-based classification model from the remote sensing image scene dataset containing non-negligible or even significant error labels. To this end, this article proposes a new RSSC-oriented error-tolerant deep learning (RSSC-ETDL) approach to mitigate the adverse effect of incorrect labels of the remote sensing image scene dataset. In our proposed RSSC-ETDL method, learning multiview CNNs and correcting error labels are alternatively conducted in an iterative manner. It is noted that to make the alternative scheme work effectively, we propose a novel adaptive multifeature collaborative representation classifier (AMF-CRC) that benefits from adaptively combining multiple features of CNNs to correct the labels of uncertain samples. To quantitatively evaluate the performance of error-tolerant methods in the remote sensing domain, we construct remote sensing image scene datasets with: 1) simulated noisy labels by corrupting the open datasets with varying error rates and 2) real noisy labels by deploying the greedy annotation strategies that are practically used to accelerate the process of annotating remote sensing image scene datasets. Extensive experiments on these datasets demonstrate that our proposed RSSC-ETDL approach outperforms the state-of-the-art approaches. Yansheng Li 0001, Yongjun Zhang 0002, Zhihui Zhu |
IEEE Trans. Cybern. | 3 |
| 2021 | Learning Deep Cross-Modal Embedding Networks for Zero-Shot Remote Sensing Image Scene ClassificationabstractDue to its wide applications, remote sensing (RS) image scene classification has attracted increasing research interest. When each category has a sufficient number of labeled samples, RS image scene classification can be well addressed by deep learning. However, in the RS big data era, it is extremely difficult or even impossible to annotate RS scene samples for all the categories in one time as the RS scene classification often needs to be extended along with the emergence of new applications that inevitably involve a new class of RS images. Hence, the RS big data era fairly requires a zero-shot RS scene classification (ZSRSSC) paradigm in which the classification model learned from training RS scene categories obeys the inference ability to recognize the RS image scenes from unseen categories, in common with the humans’ evolutionary perception ability. Unfortunately, zero-shot classification is largely unexploited in the RS field. This article proposes a novel ZSRSSC method based on locality-preservation deep cross-modal embedding networks (LPDCMENs). The proposed LPDCMENs, which can fully assimilate the pairwise intramodal and intermodal supervision in an end-to-end manner, aim to alleviate the problem of class structure inconsistency between two hybrid spaces (i.e., the visual image space and the semantic space). To pursue a stable and generalization ability, which is highly desired for ZSRSSC, a set of explainable constraints is specially designed to optimize LPDCMENs. To fully verify the effectiveness of the proposed LPDCMENs, we collect a new large-scale RS scene data set, including the instance-level visual images and class-level semantic representations (RSSDIVCS), where the general and domain knowledge is exploited to construct the class-level semantic representations. Extensive experiments show that the proposed ZSRSSC method based on LPDCMENs can obviously outperform the state-of-the-art methods, and the domain knowledge further improves the performance of ZSRSSC compared with the general knowledge. The collected RSSDIVCS will be made publicly available along with this article. Yansheng Li 0001, Zhihui Zhu, Jin-Gang Yu, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | The Global Optimization Geometry of Low-Rank Matrix OptimizationabstractThis paper considers general rank-constrained optimization problems that minimize a general objective function${f}( {X})$over the set of rectangular${n}\times {m}$matrices that have rank at most r. To tackle the rank constraint and also to reduce the computational burden, we factorize$ {X}$into$ {U} {V} ^{\mathrm {T}}$where$ {U}$and$ {V}$are${n}\times {r}$and${m}\times {r}$matrices, respectively, and then optimize over the small matrices$ {U}$and$ {V}$. We characterize the global optimization geometry of the nonconvex factored problem and show that the corresponding objective function satisfies the robust strict saddle property as long as the original objective function f satisfies restricted strong convexity and smoothness properties, ensuring global convergence of many local search algorithms (such as noisy gradient descent) in polynomial time for solving the factored problem. We also provide a comprehensive analysis for the optimization geometry of a matrix factorization problem where we aim to find${n}\times {r}$and${m}\times {r}$matrices$ {U}$and$ {V}$such that$ {U} {V} ^{\mathrm {T}}$approximates a given matrix$ {X}^\star $. Aside from the robust strict saddle property, we show that the objective function of the matrix factorization problem has no spurious local minima and obeys the strict saddle property not only for the exact-parameterization case where$\mathrm {rank}( {X}^\star) = {r}$, but also for the over-parameterization case where$\mathrm {rank}( {X}^\star) < {r}$and the under-parameterization case where$\mathrm {rank}( {X}^\star) > {r}$. These geometric properties imply that a number of iterative optimization algorithms (such as gradient descent) converge to a global solution with random initialization. Zhihui Zhu, Qiuwei Li, Gongguo Tang, Michael B. Wakin |
IEEE Trans. Inf. Theory | 1 |
| 2021 | Factor-Bounded Nonnegative Matrix FactorizationabstractNonnegative Matrix Factorization (NMF) is broadly used to determine class membership in a variety of clustering applications. From movie recommendations and image clustering to visual feature extractions, NMF has applications to solve a large number of knowledge discovery and data mining problems. Traditional optimization methods, such as the Multiplicative Updating Algorithm (MUA), solves the NMF problem by utilizing an auxiliary function to ensure that the objective monotonically decreases. Although the objective in MUA converges, there exists no proof to show that the learned matrix factors converge as well. Without this rigorous analysis, the clustering performance and stability of the NMF algorithms cannot be guaranteed. To address this knowledge gap, in this article, we study the factor-bounded NMF problem and provide a solution algorithm with proven convergence by rigorous mathematical analysis, which ensures that both the objective and matrix factors converge. In addition, we show the relationship between MUA and our solution followed by an analysis of the convergence of MUA. Experiments on both toy data and real-world datasets validate the correctness of our proposed method and its utility as an effective clustering algorithm. Kai Liu 0018, Zhihui Zhu, Lodewijk Brand, Hua Wang 0007 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2020 | Viewpoint-Aware Loss with Angular Regularization for Person Re-IdentificationabstractAlthough great progress in supervised person re-identification (Re-ID) has been made recently, due to the viewpoint variation of a person, Re-ID remains a massive visual challenge. Most existing viewpoint-based person Re-ID methods project images from each viewpoint into separated and unrelated sub-feature spaces. They only model the identity-level distribution inside an individual viewpoint but ignore the underlying relationship between different viewpoints. To address this problem, we propose a novel approach, called Viewpoint-Aware Loss with Angular Regularization (VA-reID). Instead of one subspace for each viewpoint, our method projects the feature from different viewpoints into a unified hypersphere and effectively models the feature distribution on both the identity-level and the viewpoint-level. In addition, rather than modeling different viewpoints as hard labels used for conventional viewpoint classification, we introduce viewpoint-aware adaptive label smoothing regularization (VALSR) that assigns the adaptive soft label to feature representation. VALSR can effectively solve the ambiguity of the viewpoint cluster label assignment. Extensive experiments on the Market1501 and DukeMTMC-reID datasets demonstrated that our method outperforms the state-of-the-art supervised Re-ID methods. Zhihui Zhu, Xinyang Jiang, Feng Zheng 0001, Feiyue Huang, Xing Sun 0001, Wei-Shi Zheng 0001 |
AAAI | 1 |
| 2020 | Robust Homography Estimation via Dual Principal Component PursuitabstractWe revisit robust estimation of homographies over point correspondences between two or three views, a fundamental problem in geometric vision. The analysis serves as a platform to support a rigorous investigation of Dual Principal Component Pursuit (DPCP) as a valid and powerful alternative to RANSAC for robust model fitting in multiple-view geometry. Homography fitting is cast as a robust nullspace estimation problem over either homographic or epipolar/trifocal embeddings. We prove that the nullspace of epipolar or trifocal embeddings in the homographic scenario, of dimension 3 and 6 for two and three views respectively, is defined by unique, computable homographies. Experiments show that DPCP performs on par with USAC with local optimization, while requiring an order of magnitude less computing time, and it also outperforms a recent deep learning implementation for homography estimation. Tianjiao Ding, Yunchen Yang, Zhihui Zhu, Daniel P. Robinson, René Vidal, Laurent Kneip, Manolis C. Tsakiris |
CVPR | 3 |
| 2020 | Geometric Analysis of Nonconvex Optimization Landscapes for Overcomplete Learning
Qing Qu 0001, Yuexiang Zhai, Xiao Li 0009, Zhihui Zhu |
ICLR | 5 |
| 2020 | Robust Recovery via Implicit Bias of Discrepant Learning Rates for Double Over-parameterizationabstractRecent advances have shown that implicit bias of gradient descent on over-parameterized models enables the recovery of low-rank matrices from linear measurements, even with no prior knowledge on the intrinsic rank. In contrast, for {\em robust} low-rank matrix recovery from {\em grossly corrupted} measurements, over-parameterization leads to overfitting without prior knowledge on both the intrinsic rank and sparsity of corruption. This paper shows that with a {\em double over-parameterization} for both the low-rank matrix and sparse corruption, gradient descent with {\em discrepant learning rates} provably recovers the underlying matrix even without prior knowledge on neither rank of the matrix nor sparsity of the corruption. We further extend our approach for the robust recovery of natural images by over-parameterizing images with deep convolutional networks. Experiments show that our method handles different test images and varying corruption levels with a single learning pipeline where the network width and termination conditions do not need to be adjusted on a case-by-case basis. Underlying the success is again the implicit bias with discrepant learning rates on different over-parameterized parameters, which may bear on broader applications. Chong You, Zhihui Zhu, Qing Qu 0001, Yi Ma 0001 |
NeurIPS | 2 |
| 2020 | Orthant Based Proximal Stochastic Gradient Method for ℓ 1-Regularized Optimization
Tianyu Ding, Bo Ji 0003, Guanyi Wang, Yixin Shi, Xiao Tu, Zhihui Zhu |
ECML/PKDD (3) | 9 |
| 2020 | Exact Recovery of Multichannel Sparse Blind Deconvolution via Gradient DescentabstractWe study the multichannel sparse blind deconvolution (MCS-BD) problem, whose task is to simultaneously recover a kernel $a$ and multiple sparse inputs $\{x_i\}_{i=1}^p$ from their circulant convolution $y_i = a \;\circledast \;x_i $ ($i=1,\dots,p$). We formulate the task as a nonconvex optimization problem over the sphere. Under mild statistical assumptions of the data, we prove that the vanilla Riemannian gradient descent (RGD) method, with random initializations, provably recovers both the kernel $a$ and the signals $\{x_i\}_{i=1}^p$ up to a signed shift ambiguity. In comparison with state-of-the-art results, our work shows significant improvements in terms of sample complexity and computational efficiency. Our theoretical results are corroborated by numerical experiments, which demonstrate the superior performance of the proposed approach over the previous methods on both synthetic and real datasets. Qing Qu 0001, Xiao Li 0009, Zhihui Zhu |
SIAM J. Imaging Sci. | 3 |
| 2020 | The Global Geometry of Centralized and Distributed Low-rank Matrix Recovery Without RegularizationabstractLow-rank matrix recovery is a fundamental problem in signal processing and machine learning. A recent very popular approach to recovering a low-rank matrix X is to factorize it as a product of two smaller matrices, i.e., X = UVT, and then optimize over U, V instead of X. Despite the resulting non-convexity, recent results have shown that many factorized objective functions actually have benign global geometry-with no spurious local minima and satisfying the so-called strict saddle property-ensuring convergence to a global minimum for many local-search algorithms. Such results hold whenever the original objective function is restricted strongly convex and smooth. However, most of these results actually consider a modified cost function that includes a balancing regularizer. While useful for deriving theory, this balancing regularizer does not appear to be necessary in practice. In this work, we close this theory-practice gap by proving that the unaltered factorized non-convex problem, without the balancing regularizer, also has similar benign global geometry. Moreover, we also extend our theoretical results to the field of distributed optimization. Shuang Li 0003, Qiuwei Li, Zhihui Zhu, Gongguo Tang, Michael B. Wakin |
IEEE Signal Process. Lett. | 3 |
| 2020 | Design of Compressed Sensing System With Probability-Based Prior InformationabstractThis paper deals with the design of a sensing matrix along with a sparse recovery algorithm by utilizing the probability-based prior information for compressed sensing systems. With the knowledge of the probability for each atom of the dictionary being used, a diagonal weighted matrix is obtained and then the sensing matrix is designed by minimizing a weighted function such that the Gram of the equivalent dictionary is as close to the Gram of dictionary as possible. An analytical solution for the corresponding sensing matrix is derived that requires low computational complexity. We also exploit this prior information through the sparse recovery stage and propose a probability-driven orthogonal matching pursuit algorithm that improves the accuracy of the recovery. Simulations for synthetic data and application scenarios of video streaming are carried out to compare the performance of the proposed methods with some existing algorithms. The results reveal that the proposed compressed sensing (CS) approach outperforms existing CS systems. Qianru Jiang, Sheng Li 0005, Zhihui Zhu, Huang Bai, Xiongxiong He, Rodrigo C. de Lamare |
IEEE Trans. Multim. | 3 |
| 2019 | The Geometry of Equality-constrained Global Consensus ProblemsabstractA variety of unconstrained nonconvex optimization problems have been shown to have benign geometric landscapes that satisfy the strict saddle property and have no spurious local minima. We present a general result relating the geometry of an unconstrained centralized problem to its equality-constrained distributed extension. It follows that many global consensus problems inherit the benign geometry of their original centralized counterpart. Taking advantage of this fact, we demonstrate the favorable performance of the Gradient ADMM algorithm on a distributed low-rank matrix approximation problem. Qiuwei Li, Zhihui Zhu, Gongguo Tang, Michael B. Wakin |
ICASSP | 2 |
| 2019 | Noisy Dual Principal Component PursuitabstractDual Principal Component Pursuit (DPCP) is a recently proposed non-convex optimization based method for learning subspaces of high relative dimension from noiseless datasets contaminated by as many outliers as the square of the number of inliers. Experimentally, DPCP has proved to be robust to noise and outperform the popular RANSAC on 3D vision tasks such as road plane detection and relative poses estimation from three views. This paper extends the global optimality and convergence theory of DPCP to the case of data corrupted by noise, and further demonstrates its robustness using synthetic and real data. Tianyu Ding, Zhihui Zhu, Tianjiao Ding, Yunchen Yang, Daniel P. Robinson, Manolis C. Tsakiris, René Vidal |
ICML | 2 |
| 2019 | Alternating Minimizations Converge to Second-Order Optimal SolutionsabstractThis work studies the second-order convergence for both standard alternating minimization and proximal alternating minimization. We show that under mild assumptions on the (nonconvex) objective function, both algorithms avoid strict saddles almost surely from random initialization. Together with known first-order convergence results, this implies both algorithms converge to a second-order stationary point. This solves an open problem for the second-order convergence of alternating minimization algorithms that have been widely used in practice to solve large-scale nonconvex problems due to their simple implementation, fast convergence, and superb empirical performance. Qiuwei Li, Zhihui Zhu, Gongguo Tang |
ICML | 2 |
| 2019 | Learning Deep Networks under Noisy Labels for Remote Sensing Image Scene ClassificationabstractThe deep convolutional neural network (DCNN) has been successfully applied in remote sensing (RS) image scene classification, and the superior performances of DCNNs highly depend on not only the large number of samples, but also the high accuracy of labels. In practice, collecting definitely accurate labels for a large-scale RS image scene dataset needs large amounts of manual intervention, but collecting roughly noisy labels for a dataset would be significantly simplified in the RS scenario. Therefore, how to robustly learn the superior DCNN-based classification model from one RS image scene dataset containing some error labels is a practical problem of great importance. To this end, this paper proposes a simple but effective RS-oriented error-tolerant deep learning (RS-ETDL) approach to mitigate the adverse effect of incorrect labels of the corrupted RS image scene dataset. In our proposed RS-ETDL method, learning multi-view DCNN models and correcting error labels are jointly conducted in an iteratively alternative manner. Extensive experiments on the noisy RS image scene dataset demonstrate that our proposed method outperforms the state-of-the-art approaches with a large margin. Yansheng Li 0001, Yongjun Zhang 0002, Zhihui Zhu |
IGARSS | 3 |
| 2019 | A Nonconvex Approach for Exact and Efficient Multichannel Sparse Blind DeconvolutionabstractWe study the multi-channel sparse blind deconvolution (MCS-BD) problem, whose task is to simultaneously recover a kernel $\mathbf a$ and multiple sparse inputs $\{\mathbf x_i\}_{i=1}^p$ from their circulant convolution $\mathbf y_i = \mb a \circledast \mb x_i $ ($i=1,\cdots,p$). We formulate the task as a nonconvex optimization problem over the sphere. Under mild statistical assumptions of the data, we prove that the vanilla Riemannian gradient descent (RGD) method, with random initializations, provably recovers both the kernel $\mathbf a$ and the signals $\{\mathbf x_i\}_{i=1}^p$ up to a signed shift ambiguity. In comparison with state-of-the-art results, our work shows significant improvements in terms of sample complexity and computational efficiency. Our theoretical results are corroborated by numerical experiments, which demonstrate superior performance of the proposed approach over the previous methods on both synthetic and real datasets. Qing Qu 0001, Xiao Li 0009, Zhihui Zhu |
NeurIPS | 3 |
| 2019 | A Linearly Convergent Method for Non-Smooth Non-Convex Optimization on the Grassmannian with Applications to Robust Subspace and Dictionary LearningabstractMinimizing a non-smooth function over the Grassmannian appears in many applications in machine learning. In this paper we show that if the objective satisfies a certain Riemannian regularity condition with respect to some point in the Grassmannian, then a Riemannian subgradient method with appropriate initialization and geometrically diminishing step size converges at a linear rate to that point. We show that for both the robust subspace learning method Dual Principal Component Pursuit (DPCP) and the Orthogonal Dictionary Learning (ODL) problem, the Riemannian regularity condition is satisfied with respect to appropriate points of interest, namely the subspace orthogonal to the sought subspace for DPCP and the orthonormal dictionary atoms for ODL. Consequently, we obtain in a unified framework significant improvements for the convergence theory of both methods. Zhihui Zhu, Tianyu Ding, Daniel P. Robinson, Manolis C. Tsakiris, René Vidal |
NeurIPS | 1 |
| 2019 | Distributed Low-rank Matrix Factorization With Exact ConsensusabstractLow-rank matrix factorization is a problem of broad importance, owing to the ubiquity of low-rank models in machine learning contexts. In spite of its non- convexity, this problem has a well-behaved geometric landscape, permitting local search algorithms such as gradient descent to converge to global minimizers. In this paper, we study low-rank matrix factorization in the distributed setting, where local variables at each node encode parts of the overall matrix factors, and consensus is encouraged among certain such variables. We identify conditions under which this new problem also has a well-behaved geometric landscape, and we propose an extension of distributed gradient descent (DGD) to solve this problem. The favorable landscape allows us to prove convergence to global optimality with exact consensus, a stronger result than what is provided by off-the-shelf DGD theory. Zhihui Zhu, Qiuwei Li, Xinshuo Yang, Gongguo Tang, Michael B. Wakin |
NeurIPS | 1 |
| 2019 | Optimized structured sparse sensing matrices for compressive sensing
Tao Hong 0006, Xiao Li 0009, Zhihui Zhu, Qiuwei Li |
Signal Process. | 3 |
| 2019 | Speckle Suppression Based on Weighted Nuclear Norm Minimization and Grey TheoryabstractCoherent imaging systems are greatly affected by speckle noise, which makes visual analysis and features extraction a difficult task. In this paper, we propose a speckle suppression algorithm based on weighted nuclear norm minimization (WNNM) and Grey theory. First, we use logarithmic transformation to the noisy images such that the speckle noise is transformed into additive noise. Second, by matching the local blocks based on Grey theory, we will get approximate low-rank matrices grouped by the similar blocks of the reference patches. We then estimate the noise variance of the noisy images with the wavelet transform. Finally, we use WNNM method to denoise the image. The results show that our algorithm not only effectively improves the visual effect of the denoised image and preserves the local structure of the image better but also improves the objective indexes values of the denoised image. Shuaiqi Liu 0001, Jie Zhao 0008, Zhihui Zhu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2019 | Hankel Low-Rank Approximation for Seismic Noise AttenuationabstractThe low-rankness property of the Hankel matrix formulated from the clean seismic data corresponding to a few number of linear events has been successively leveraged in many low-rank (LR) approximation methods for seismic data denoising. The common scheme in these rank-reduction methods is to compute the best LR approximation of the formulated Hankel matrix and then obtain the denoised data from the LR matrix. However, without utilizing the Hankel structure when computing the LR approximation, if we rearrange the denoised data into a Hankel matrix, it is in general not exactly LR as expected. In this paper, we propose a Hankel LR (HLR) approximation method to simultaneously exploit both the Hankel structure and the LR property underlying the clean seismic data. The formulated HLR approximation problem is solved by an alternating-minimization-based algorithm. We provide rigorously convergence analysis of the proposed algorithm. The superior performance of the proposed HLR approximation method is demonstrated on both synthetic and field seismic data. Chong Wang 0020, Zhihui Zhu, Hanming Gu, Xinming Wu, Shuaiqi Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Dropping Symmetry for Fast Symmetric Nonnegative Matrix FactorizationabstractSymmetric nonnegative matrix factorization (NMF)---a special but important class of the general NMF---is demonstrated to be useful for data analysis and in particular for various clustering tasks. Unfortunately, designing fast algorithms for Symmetric NMF is not as easy as for the nonsymmetric counterpart, the latter admitting the splitting property that allows efficient alternating-type algorithms. To overcome this issue, we transfer the symmetric NMF to a nonsymmetric one, then we can adopt the idea from the state-of-the-art algorithms for nonsymmetric NMF to design fast algorithms solving symmetric NMF. We rigorously establish that solving nonsymmetric reformulation returns a solution for symmetric NMF and then apply fast alternating based algorithms for the corresponding reformulated problem. Furthermore, we show these fast algorithms admit strong convergence guarantee in the sense that the generated sequence is convergent at least at a sublinear rate and it converges globally to a critical point of the symmetric NMF. We conduct experiments on both synthetic data and image clustering to support our result. Zhihui Zhu, Xiao Li 0009, Kai Liu 0018, Qiuwei Li |
NeurIPS | 1 |
| 2018 | Dual Principal Component Pursuit: Improved Analysis and Efficient AlgorithmsabstractRecent methods for learning a linear subspace from data corrupted by outliers are based on convex L1 and nuclear norm optimization and require the dimension of the subspace and the number of outliers to be sufficiently small [27]. In sharp contrast, the recently proposed Dual Principal Component Pursuit (DPCP) method [22] can provably handle subspaces of high dimension by solving a non-convex L1 optimization problem on the sphere. However, its geometric analysis is based on quantities that are difficult to interpret and are not amenable to statistical analysis. In this paper we provide a refined geometric analysis and a new statistical analysis that show that DPCP can tolerate as many outliers as the square of the number of inliers, thus improving upon other provably correct robust PCA methods. We also propose a scalable Projected Sub-Gradient Descent method (DPCP-PSGD) for solving the DPCP problem and show it admits linear convergence even though the underlying optimization problem is non-convex and non-smooth. Experiments on road plane detection from 3D point cloud data demonstrate that DPCP-PSGD can be more efficient than the traditional RANSAC algorithm, which is one of the most popular methods for such computer vision applications. Zhihui Zhu, Daniel P. Robinson, Daniel Q. Naiman, René Vidal, Manolis C. Tsakiris |
NeurIPS | 1 |
| 2018 | On Collaborative Compressive Sensing Systems: The Framework, Design, and AlgorithmabstractBased on the maximum likelihood estimation principle, we derive a collaborative estimation framework that fuses several different estimators and yields a better estimate. Applying it to compressive sensing (CS), we propose a collaborative CS (CCS) scheme consisting of a bank of $K$ CS systems that share the same sensing matrix but have different sparsifying dictionaries. This CCS system is expected to yield better performance than each individual CS system, while requiring the same time as that needed for each individual CS system when a parallel computing strategy is used. We then provide an approach to designing optimal CCS systems by utilizing a measure that involves both the sensing matrix and dictionaries and hence allows us to simultaneously optimize the sensing matrix and all the $K$ dictionaries. An alternating minimization-based algorithm is derived for solving the corresponding optimal design problem. With a rigorous convergence analysis, we show that the proposed algorithm is convergent. Experiments are carried out to confirm the theoretical results and show that the proposed CCS system yields significant improvements over the existing CS systems in terms of the signal recovery accuracy. Zhihui Zhu, Gang Li 0010, Jiajun Ding, Qiuwei Li, Xiongxiong He |
SIAM J. Imaging Sci. | 1 |
| 2018 | An efficient method for robust projection matrix design
Tao Hong 0006, Zhihui Zhu |
Signal Process. | 2 |
| 2018 | Online learning sensing matrix and sparsifying dictionary simultaneously for compressive sensing
Tao Hong 0006, Zhihui Zhu |
Signal Process. | 2 |
| 2018 | The Eigenvalue Distribution of Discrete Periodic Time-Frequency Limiting OperatorsabstractBandlimiting and timelimiting operators play a fundamental role in analyzing bandlimited signals that are approximately timelimited (or vice versa). In this letter, we consider a time-frequency (in the discrete Fourier transform (DFT) domain) limiting operator whose eigenvectors are known as the periodic discrete prolate spheroidal sequences. We establish new nonasymptotic results on the eigenvalue distribution of this operator. As a byproduct, we also characterize the eigenvalue distribution of a set of submatrices of the DFT matrix, which is of independent interest. Zhihui Zhu, Santhosh Karnik, Mark A. Davenport, Justin K. Romberg, Michael B. Wakin |
IEEE Signal Process. Lett. | 1 |
| 2017 | Jazz: A companion to music for frequency estimation with missing dataabstractFrequency estimation is a classical problem in signal processing, with applications ranging from sensor array processing to wireless communications and structural health monitoring. Modern algorithms based on atomic norm minimization can cope with missing data but incur a high computational cost. To recover missing data from an ensemble of frequency-sparse signals, we propose a computationally efficient low-rank tensor completion algorithm that exploits the fact that each signal in the ensemble can be associated with a Toeplitz matrix. We name our algorithm JAZZ in the spirit of the classical MUSIC algorithm for frequency estimation and in tribute to the random, improvisational nature of jazz music. Qiuwei Li, Shuang Li 0003, Hassan Mansour, Michael B. Wakin, Dehui Yang, Zhihui Zhu |
ICASSP | 6 |
| 2017 | A new framework for designing incoherent sparsifying dictionariesabstractThis paper deals with designing incoherent sparsifying dictionaries. A new framework is proposed, in which the sparse representation error and mutual coherence are embedded. An alternating minimization method is developed for solving the optimal dictionary problem. One of the significant features of the proposed approach is that the dictionary is directly updated with each atom being normalized. A gradient-based algorithm is derived for this purpose. Experiments are carried out and the results show that the proposed approach outperforms some prevailing ones in terms of minimizing sparse representation error and mutual coherence. Gang Li 0010, Zhihui Zhu, Huang Bai, Aihua Yu |
ICASSP | 2 |
| 2017 | Fast orthogonal approximations of sampled sinusoids and bandlimited signalsabstractIn this paper, we provide a dictionary for representing the discrete vector one obtains when collecting a finite set of uniform samples from a baseband analog signal. Like the discrete prolate spheroidal sequences (DPSS's), the proposed orthogonal basis compactly captures most of the energy in oversampled bandlimited signals. The complexity of computing the representation of a signal using the proposed dictionary is comparable to the FFT, which is much less than that involving the DPSS basis. We also give non-asymptotic results to guarantee that the proposed basis not only provides a very high degree of approximation accuracy in an MSE sense for bandlimited sample vectors, but also that it can provide high-quality approximations of all sampled sinusoids within the band of interest. Zhihui Zhu, Santhosh Karnik, Michael B. Wakin, Mark A. Davenport, Justin K. Romberg |
ICASSP | 1 |
| 2017 | SAR Image Denoising via Sparse Representation in Shearlet Domain Based on Continuous Cycle SpinningabstractHow to suppress speckle noise effectively has become one of the key problems in remote sensing image processing. This problem also restricts the development of key technology severely, especially in military applications and so on. To overcome the shortcoming that the optimal solution of image denoising based on sparse representation does not have one-to-one mapping of the original signal space, in this paper, we propose a novel synthetic aperture radar (SAR) image denoising via sparse representation in Shearlet domain based on continuous cycle spinning. First, the Shearlet transform is applied to the noised SAR image. Second, a new optimal denoising model is constructed using the sparse representation model based on the cycle spinning theory. Finally, the alternate iteration algorithm is used to solve the optimal denoising model to obtain the denoised image. The experimental results show that the proposed method not only effectively suppresses the speckle noise and improves the peak signal-to-noise ratio of denoising SAR image, but also obviously improves the visual effect of the SAR image, especially by enhancing the texture of the SAR image. Shuaiqi Liu 0001, Peifei Li, Jie Zhao 0008, Zhihui Zhu, Xuehu Wang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | On the Asymptotic Equivalence of Circulant and Toeplitz MatricesabstractAny sequence of uniformly bounded N × N Hermitian Toeplitz matrices (HN) is asymptotically equivalent to a certain sequence of N×N circulant matrices (CN) derived from the Toeplitz matrices in the sense that ∥HN- CN∥F= o(√N) as N → ∞. This implies that certain collective behaviors of the eigenvalues of each Toeplitz matrix are reflected in those of the corresponding circulant matrix and supports the utilization of the computationally efficient fast Fourier transform (instead of the Karhunen-Loève transform) in applications like coding and filtering. In this paper, we study the asymptotic performance of the individual eigenvalue estimates. We show that the asymptotic equivalence of the circulant and Toeplitz matrices implies the individual asymptotic convergence of the eigenvalues for certain types of Toeplitz matrices. We also show that these estimates asymptotically approximate the largest and smallest eigenvalues for more general classes of Toeplitz matrices. Zhihui Zhu, Michael B. Wakin |
IEEE Trans. Inf. Theory | 1 |
| 2016 | An efficient algorithm for designing projection matrix in compressive sensing based on alternating optimization
Tao Hong 0006, Huang Bai, Sheng Li 0005, Zhihui Zhu |
Signal Process. | 4 |