Fanghui Liu 0001

dblp:119/1038 · DBLP profile ↗
← Back
55ranked-venue papers
26as first author
29since 2021 · last 2025
0000-0003-4133-7921ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 21 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 2 since 2021
YearPublicationVenuePosition
2025 How Gradient descent balances features: A dynamical analysis for two-layer neural networks
abstract
This paper investigates the fundamental regression task of learning $k$ neurons (\emph{a.k.a.} teachers) from Gaussian input, using two-layer ReLU neural networks with width $m$ (\emph{a.k.a.} students) and $m, k= \mathcal{O}(1)$, trained via gradient descent under proper initialization and a small step-size. Our analysis follows a three-phase structure: \emph{alignment} after weak recovery, \emph{tangential growth}, and \emph{local convergence}, providing deeper insights into the learning dynamics of gradient descent (GD). We prove the global convergence at the rate of $\mathcal{O}(T^{-3})$ for the zero loss of excess risk. Additionally, our results show that GD automatically groups and balances student neurons, revealing an implicit bias toward achieving the minimum ``balanced'' $\ell_2$-norm in the solution. Our work extends beyond previous studies in exact-parameterization setting ($m = k = 1$, (Yehudai and Ohad, 2020)) and single-neuron setting ($m \geq k = 1$, (Xu and Du, 2023)). The key technical challenge lies in handling the interactions between multiple teachers and students during training, which we address by refining the alignment analysis in Phase 1 and introducing a new dynamic system analysis for tangential components in Phase 2. Our results pave the way for further research on optimizing neural network training dynamics and understanding implicit biases in more complex architectures.
Fanghui Liu 0001, Volkan Cevher
ICLR2
2025 LoRA-One: One-Step Full Gradient Could Suffice for Fine-Tuning Large Language Models, Provably and Efficiently
abstract
This paper explores how theory can guide and enhance practical algorithms, using Low-Rank Adaptation (LoRA) (Hu et al., 2022) in large language models as a case study. We rigorously prove that, under gradient descent, LoRA adapters align with specific singular subspaces of the one-step full fine-tuning gradient. This result suggests that, by properly initializing the adapters using the one-step full gradient, subspace alignment can be achieved immediately—applicable to both linear and nonlinear models. Building on our theory, we propose a theory-driven algorithm, LoRA-One, where the linear convergence (as well as generalization) is built and incorporating preconditioners theoretically helps mitigate the effects of ill-conditioning. Besides, our theory reveals connections between LoRA-One and other gradient-alignment-based methods, helping to clarify misconceptions in the design of such algorithms. LoRA-One achieves significant empirical improvements over LoRA and its variants across benchmarks in natural language understanding, mathematical reasoning, and code generation. Code is available at: https://github.com/YuanheZ/LoRA-One.
Yuanhe Zhang, Fanghui Liu 0001, Yudong Chen 0001
ICML2
2025 The φ Curve: The Shape of Generalization through the Lens of Norm-based Capacity Control
Yudong Chen 0001, Lorenzo Rosasco, Fanghui Liu 0001
NeurIPS4
2024 The Role of Over-Parameterization in Machine Learning - the Good, the Bad, the Ugly
abstract
The conventional wisdom of simple models in machine learning misses the bigger picture, especially over-parameterized neural networks (NNs), where the number of parameters are much larger than the number of training data. Our goal is to explore the mystery behind over-parameterized models from a theoretical side. In this talk, I will discuss the role of over-parameterization in neural networks, to theoretically understand why they can perform well. First, I will discuss the role of over-parameterization in neural networks from the perspective of models, to theoretically understand why they can genralize well. Second, the effects of over-parameterization in robustness, privacy are discussed. Third, I will talk about the over-parameterization from kernel methods to neural networks in a function space theory view. Besides, from classical statistical learning to sequential decision making, I will talk about the benefits of over-parameterization on how deep reinforcement learning works well for function approximation. Potential future directions on theory of over-parameterization ML will also be discussed.
Fanghui Liu 0001
AAAI1
2024 Learning Scalable Model Soup on a Single GPU: An Efficient Subspace Training Strategy
Weisen Jiang, Fanghui Liu 0001, Xiaolin Huang, James T. Kwok
ECCV (65)3
2024 Efficient local linearity regularization to overcome catastrophic overfitting
abstract
Catastrophic overfitting (CO) in single-step adversarial training (AT) results in abrupt drops in the adversarial test accuracy (even down to $0$%). For models trained with multi-step AT, it has been observed that the loss function behaves locally linearly with respect to the input, this is however lost in single-step AT. To address CO in single-step AT, several methods have been proposed to enforce local linearity of the loss via regularization. However, these regularization terms considerably slow down training due to *Double Backpropagation*. Instead, in this work, we introduce a regularization term, called ELLE, to mitigate CO *effectively* and *efficiently* in classical AT evaluations, as well as some more difficult regimes, e.g., large adversarial perturbations and long training schedules. Our regularization term can be theoretically linked to curvature of the loss function and is computationally cheaper than previous methods by avoiding *Double Backpropagation*. Our thorough experimental validation demonstrates that our work does not suffer from CO, even in challenging settings where previous works suffer from it. We also notice that adapting our regularization parameter during training (ELLE-A) greatly improves the performance, specially in large $\epsilon$ setups. Our implementation is available in https://github.com/LIONS-EPFL/ELLE.
Elías Abad-Rocamora, Fanghui Liu 0001, Grigorios Chrysos 0002, Pablo M. Olmos, Volkan Cevher
ICLR2
2024 Generalization of Scaled Deep ResNets in the Mean-Field Regime
abstract
Despite the widespread empirical success of ResNet, the generalization properties of deep ResNet are rarely explored beyond the lazy training regime. In this work, we investigate scaled ResNet in the limit of infinitely deep and wide neural networks, of which the gradient flow is described by a partial differential equation in the large-neural network limit, i.e., the mean-field regime. To derive the generalization bounds under this setting, our analysis necessitates a shift from the conventional time-invariant Gram matrix employed in the lazy training regime to a time-variant, distribution-dependent version. To this end, we provide a global lower bound on the minimum eigenvalue of the Gram matrix under the mean-field regime. Besides, for the traceability of the dynamic of Kullback-Leibler (KL) divergence, we establish the linear convergence of the empirical error and estimate the upper bound of the KL divergence over parameters distribution. Finally, we build the uniform convergence for generalization bound via Rademacher complexity. Our results offer new insights into the generalization ability of deep ResNet beyond the lazy training regime and contribute to advancing the understanding of the fundamental properties of deep neural networks.
Yihang Chen 0003, Fanghui Liu 0001, Grigorios Chrysos 0002, Volkan Cevher
ICLR2
2024 Robust NAS under adversarial training: benchmark, theory, and beyond
abstract
Recent developments in neural architecture search (NAS) emphasize the significance of considering robust architectures against malicious data. However, there is a notable absence of benchmark evaluations and theoretical guarantees for searching these robust architectures, especially when adversarial training is considered. In this work, we aim to address these two challenges, making twofold contributions. First, we release a comprehensive data set that encompasses both clean accuracy and robust accuracy for a vast array of adversarially trained networks from the NAS-Bench-201 search space on image datasets. Then, leveraging the neural tangent kernel (NTK) tool from deep learning theory, we establish a generalization theory for searching architecture in terms of clean accuracy and robust accuracy under multi-objective adversarial training. We firmly believe that our benchmark and theoretical insights will significantly benefit the NAS community through reliable reproducibility, efficient assessment, and theoretical foundation, particularly in the pursuit of robust architectures.
Yongtao Wu, Fanghui Liu 0001, Carl-Johann Simon-Gabriel, Grigorios Chrysos 0002, Volkan Cevher
ICLR2
2024 Revisiting Character-level Adversarial Attacks for Language Models
abstract
Adversarial attacks in Natural Language Processing apply perturbations in the character or token levels. Token-level attacks, gaining prominence for their use of gradient-based methods, are susceptible to altering sentence semantics, leading to invalid adversarial examples. While character-level attacks easily maintain semantics, they have received less attention as they cannot easily adopt popular gradient-based methods, and are thought to be easy to defend. Challenging these beliefs, we introduce Charmer, an efficient query-based adversarial attack capable of achieving high attack success rate (ASR) while generating highly similar adversarial examples. Our method successfully targets both small (BERT) and large (Llama 2) models. Specifically, on BERT with SST-2, Charmer improves the ASR in $4.84$% points and the USE similarity in $8$% points with respect to the previous art. Our implementation is available in https://github.com/LIONS-EPFL/Charmer.
Elías Abad-Rocamora, Yongtao Wu, Fanghui Liu 0001, Grigorios Chrysos 0002, Volkan Cevher
ICML3
2024 High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
abstract
This paper studies kernel ridge regression in high dimensions under covariate shifts and analyzes the role of importance re-weighting. We first derive the asymptotic expansion of high dimensional kernels under covariate shifts. By a bias-variance decomposition, we theoretically demonstrate that the re-weighting strategy allows for decreasing the variance. For bias, we analyze the regularization of the arbitrary or well-chosen scale, showing that the bias can behave very differently under different regularization scales. In our analysis, the bias and variance can be characterized by the spectral decay of a data-dependent regularized kernel: the original kernel matrix associated with an additional re-weighting matrix, and thus the re-weighting strategy can be regarded as a data-dependent regularization for better understanding. Besides, our analysis provides asymptotic expansion of kernel functions/vectors under covariate shift, which has its own interest.
Yihang Chen 0003, Fanghui Liu 0001, Taiji Suzuki, Volkan Cevher
ICML2
2024 Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
abstract
Recent studies show that a reproducing kernel Hilbert space (RKHS) is not a suitable space to model functions by neural networks as the curse of dimensionality (CoD) cannot be evaded when trying to approximate even a single ReLU neuron. In this paper, we study a suitable function space for over-parameterized two-layer neural networks with bounded norms (e.g., the path norm, the Barron norm) in the perspective of sample complexity and generalization properties. First, we show that the path norm (as well as the Barron norm) is able to obtain width-independence sample complexity bounds, which allows for uniform convergence guarantees. Based on this result, we derive the improved result of metric entropy for $\epsilon$-covering up to $O(\epsilon^{-\frac{2d}{d+2}})$ ($d$ is the input dimension and the depending constant is at most polynomial order of $d$) via the convex hull technique, which demonstrates the separation with kernel methods with $\Omega(\epsilon^{-d})$ to learn the target function in a Barron space. Second, this metric entropy result allows for building a sharper generalization bound under a general moment hypothesis setting, achieving the rate at $O(n^{-\frac{d+2}{2d+2}})$. Our analysis is novel in that it offers a sharper and refined estimation for metric entropy (with a clear dependence relationship on the dimension $d$) and unbounded sampling in the estimation of the sample error and the output error.
Fanghui Liu 0001, Leello Tadesse Dadi, Volkan Cevher
J. Mach. Learn. Res.1
2024 Random fourier features for asymmetric kernels
Mingzhen He, Fanghui Liu 0001, Xiaolin Huang
Mach. Learn.3
2023 What can online reinforcement learning with function approximation benefit from general coverage conditions?
abstract
In online reinforcement learning (RL), instead of employing standard structural assumptions on Markov decision processes (MDPs), using a certain coverage condition (original from offline RL) is enough to ensure sample-efficient guarantees (Xie et al. 2023). In this work, we focus on this new direction by digging more possible and general coverage conditions, and study the potential and the utility of them in efficient online RL. We identify more concepts, including the $L^p$ variant of concentrability, the density ratio realizability, and trade-off on the partial/rest coverage condition, that can be also beneficial to sample-efficient online RL, achieving improved regret bound. Furthermore, if exploratory offline data are used, under our coverage conditions, both statistically and computationally efficient guarantees can be achieved for online RL. Besides, even though the MDP structure is given, e.g., linear MDP, we elucidate that, good coverage conditions are still beneficial to obtain faster regret bound beyond $\widetilde{\mathcal{O}}(\sqrt{T})$ and even a logarithmic order regret. These results provide a good justification for the usage of general coverage conditions in efficient online RL.
Fanghui Liu 0001, Luca Viano, Volkan Cevher
ICML1
2023 Benign Overfitting in Deep Neural Networks under Lazy Training
abstract
This paper focuses on over-parameterized deep neural networks (DNNs) with ReLU activation functions and proves that when the data distribution is well-separated, DNNs can achieve Bayes-optimal test error for classification while obtaining (nearly) zero-training error under the lazy training regime. For this purpose, we unify three interrelated concepts of overparameterization, benign overfitting, and the Lipschitz constant of DNNs. Our results indicate that interpolating with smoother functions leads to better generalization. Furthermore, we investigate the special case where interpolating smooth ground-truth functions is performed by DNNs under the Neural Tangent Kernel (NTK) regime for generalization. Our result demonstrates that the generalization error converges to a constant order that only depends on label noise and initialization noise, which theoretically verifies benign overfitting. Our analysis provides a tight lower bound on the normalized margin under non-smooth activation functions, as well as the minimum eigenvalue of NTK under high-dimensional settings, which has its own interest in learning theory.
Fanghui Liu 0001, Grigorios Chrysos 0002, Francesco Locatello, Volkan Cevher
ICML2
2023 Initialization Matters: Privacy-Utility Analysis of Overparameterized Neural Networks
abstract
We analytically investigate how over-parameterization of models in randomized machine learning algorithms impacts the information leakage about their training data. Specifically, we prove a privacy bound for the KL divergence between model distributions on worst-case neighboring datasets, and explore its dependence on the initialization, width, and depth of fully connected neural networks. We find that this KL privacy bound is largely determined by the expected squared gradient norm relative to model parameters during training. Notably, for the special setting of linearized network, our analysis indicates that the squared gradient norm (and therefore the escalation of privacy loss) is tied directly to the per-layer variance of the initialization distribution. By using this analysis, we demonstrate that privacy bound improves with increasing depth under certain initializations (LeCun and Xavier), while degrades with increasing depth under other initializations (He and NTK). Our work reveals a complex interplay between privacy and depth that depends on the chosen initialization distribution. We further prove excess empirical risk bounds under a fixed KL privacy budget, and show that the interplay between privacy utility trade-off and depth is similarly affected by the initialization.
Jiayuan Ye 0001, Fanghui Liu 0001, Reza Shokri, Volkan Cevher
NeurIPS3
2023 On the Convergence of Encoder-only Shallow Transformers
abstract
In this paper, we aim to build the global convergence theory of encoder-only shallow Transformers under a realistic setting from the perspective of architectures, initialization, and scaling under a finite width regime. The difficulty lies in how to tackle the softmax in self-attention mechanism, the core ingredient of Transformer. In particular, we diagnose the scaling scheme, carefully tackle the input/output of softmax, and prove that quadratic overparameterization is sufficient for global convergence of our shallow Transformers under commonly-used He/LeCun initialization in practice. Besides, neural tangent kernel (NTK) based analysis is also given, which facilitates a comprehensive comparison. Our theory demonstrates the separation on the importance of different scaling schemes and initialization. We believe our results can pave the way for a better understanding of modern Transformers, particularly on training dynamics.
Yongtao Wu, Fanghui Liu 0001, Grigorios Chrysos 0002, Volkan Cevher
NeurIPS2
2023 End-to-end kernel learning via generative random Fourier features
Kun Fang 0004, Fanghui Liu 0001, Xiaolin Huang, Jie Yang 0002
Pattern Recognit.2
2022 On the Double Descent of Random Features Models Trained with SGD
abstract
We study generalization properties of random features (RF) regression in high dimensions optimized by stochastic gradient descent (SGD) in under-/over-parameterized regime. In this work, we derive precise non-asymptotic error bounds of RF regression under both constant and polynomial-decay step-size SGD setting, and observe the double descent phenomenon both theoretically and empirically. Our analysis shows how to cope with multiple randomness sources of initialization, label noise, and data sampling (as well as stochastic gradients) with no closed-form solution, and also goes beyond the commonly-used Gaussian/spherical data assumption. Our theoretical results demonstrate that, with SGD training, RF regression still generalizes well for interpolation learning, and is able to characterize the double descent behavior by the unimodality of variance and monotonic decrease of bias. Besides, we also prove that the constant step-size SGD setting incurs no loss in convergence rate when compared to the exact minimum-norm interpolator, as a theoretical justification of using SGD in practice.
Fanghui Liu 0001, Johan A. K. Suykens, Volkan Cevher
NeurIPS1
2022 Understanding Deep Neural Function Approximation in Reinforcement Learning via $\epsilon$-Greedy Exploration
abstract
This paper provides a theoretical study of deep neural function approximation in reinforcement learning (RL) with the $\epsilon$-greedy exploration under the online setting. This problem setting is motivated by the successful deep Q-networks (DQN) framework that falls in this regime. In this work, we provide an initial attempt on theoretical understanding deep RL from the perspective of function class and neural networks architectures (e.g., width and depth) beyond the ``linear'' regime. To be specific, we focus on the value based algorithm with the $\epsilon$-greedy exploration via deep (and two-layer) neural networks endowed by Besov (and Barron) function spaces, respectively, which aims at approximating an $\alpha$-smooth Q-function in a $d$-dimensional feature space. We prove that, with $T$ episodes, scaling the width $m = \widetilde{\mathcal{O}}(T^{\frac{d}{2\alpha + d}})$ and the depth $L=\mathcal{O}(\log T)$ of the neural network for deep RL is sufficient for learning with sublinear regret in Besov spaces. Moreover, for a two layer neural network endowed by the Barron space, scaling the width $\Omega(\sqrt{T})$ is sufficient. To achieve this, the key issue in our analysis is how to estimate the temporal difference error under deep neural function approximation as the $\epsilon$-greedy exploration is not enough to ensure "optimism". Our analysis reformulates the temporal difference error in an $L^2(\mathrm{d}\mu)$-integrable space over a certain averaged measure $\mu$, and transforms it to a generalization problem under the non-iid setting. This might have its own interest in RL theory for better understanding $\epsilon$-greedy exploration in deep RL.
Fanghui Liu 0001, Luca Viano, Volkan Cevher
NeurIPS1
2022 Sound and Complete Verification of Polynomial Networks
abstract
Polynomial Networks (PNs) have demonstrated promising performance on face and image recognition recently. However, robustness of PNs is unclear and thus obtaining certificates becomes imperative for enabling their adoption in real-world applications. Existing verification algorithms on ReLU neural networks (NNs) based on classical branch and bound (BaB) techniques cannot be trivially applied to PN verification. In this work, we devise a new bounding method, equipped with BaB for global convergence guarantees, called Verification of Polynomial Networks or VPN for short. One key insight is that we obtain much tighter bounds than the interval bound propagation (IBP) and DeepT-Fast [Bonaert et al., 2021] baselines. This enables sound and complete PN verification with empirical validation on MNIST, CIFAR10 and STL10 datasets. We believe our method has its own interest to NN verification. The source code is publicly available at https://github.com/megaelius/PNVerification.
Elías Abad-Rocamora, Mehmet Fatih Sahin, Fanghui Liu 0001, Grigorios Chrysos 0002, Volkan Cevher
NeurIPS3
2022 Extrapolation and Spectral Bias of Neural Nets with Hadamard Product: a Polynomial Net Study
abstract
Neural tangent kernel (NTK) is a powerful tool to analyze training dynamics of neural networks and their generalization bounds. The study on NTK has been devoted to typical neural network architectures, but it is incomplete for neural networks with Hadamard products (NNs-Hp), e.g., StyleGAN and polynomial neural networks (PNNs). In this work, we derive the finite-width NTK formulation for a special class of NNs-Hp, i.e., polynomial neural networks. We prove their equivalence to the kernel regression predictor with the associated NTK, which expands the application scope of NTK. Based on our results, we elucidate the separation of PNNs over standard neural networks with respect to extrapolation and spectral bias. Our two key insights are that when compared to standard neural networks, PNNs can fit more complicated functions in the extrapolation regime and admit a slower eigenvalue decay of the respective NTK, leading to a faster learning towards high-frequency functions. Besides, our theoretical results can be extended to other types of NNs-Hp, which expand the scope of our work. Our empirical results validate the separations in broader classes of NNs-Hp, which provide a good justification for a deeper understanding of neural architectures.
Yongtao Wu, Fanghui Liu 0001, Grigorios Chrysos 0002, Volkan Cevher
NeurIPS3
2022 Generalization Properties of NAS under Activation and Skip Connection Search
abstract
Neural Architecture Search (NAS) has fostered the automatic discovery of state-of-the-art neural architectures. Despite the progress achieved with NAS, so far there is little attention to theoretical guarantees on NAS. In this work, we study the generalization properties of NAS under a unifying framework enabling (deep) layer skip connection search and activation function search. To this end, we derive the lower (and upper) bounds of the minimum eigenvalue of the Neural Tangent Kernel (NTK) under the (in)finite-width regime using a certain search space including mixed activation functions, fully connected, and residual neural networks. We use the minimum eigenvalue to establish generalization error bounds of NAS in the stochastic gradient descent training. Importantly, we theoretically and experimentally show how the derived results can guide NAS to select the top-performing architectures, even in the case without training, leading to a train-free algorithm based on our theory. Accordingly, our numerical validation shed light on the design of computationally efficient methods for NAS. Our analysis is non-trivial due to the coupling of various architectures and activation functions under the unifying framework and has its own interest in providing the lower bound of the minimum eigenvalue of NTK in deep learning theory.
Fanghui Liu 0001, Grigorios Chrysos 0002, Volkan Cevher
NeurIPS2
2022 Robustness in deep learning: The good (width), the bad (depth), and the ugly (initialization)
abstract
We study the average robustness notion in deep neural networks in (selected) wide and narrow, deep and shallow, as well as lazy and non-lazy training settings. We prove that in the under-parameterized setting, width has a negative effect while it improves robustness in the over-parameterized setting. The effect of depth closely depends on the initialization and the training mode. In particular, when initialized with LeCun initialization, depth helps robustness with the lazy training regime. In contrast, when initialized with Neural Tangent Kernel (NTK) and He-initialization, depth hurts the robustness. Moreover, under the non-lazy training regime, we demonstrate how the width of a two-layer ReLU network benefits robustness. Our theoretical developments improve the results by [Huang et al. NeurIPS21; Wu et al. NeurIPS21] and are consistent with [Bubeck and Sellke NeurIPS21; Bubeck et al. COLT21].
Fanghui Liu 0001, Grigorios Chrysos 0002, Volkan Cevher
NeurIPS2
2022 Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
abstract
The class of random features is one of the most popular techniques to speed up kernel methods in large-scale problems. Related works have been recognized by the NeurIPS Test-of-Time award in 2017 and the ICML Best Paper Finalist in 2019. The body of work on random features has grown rapidly, and hence it is desirable to have a comprehensive overview on this topic explaining the connections among various algorithms and theoretical results. In this survey, we systematically review the work on random features from the past ten years. First, the motivations, characteristics and contributions of representative random features based algorithms are summarized according to their sampling schemes, learning procedures, variance reduction properties and how they exploit training data. Second, we review theoretical results that center around the following key question: how many random features are needed to ensure a high approximation quality or no loss in the empirical/expected risks of the learned estimator. Third, we provide a comprehensive evaluation of popular random features based algorithms on several large-scale benchmark datasets and discuss their approximation quality and prediction performance for classification. Last, we discuss the relationship between random features and modern over-parameterized deep neural networks (DNNs), including the use of high dimensional random features in the analysis of DNNs as well as the gaps between current theoretical and empirical results. This survey may serve as a gentle introduction to this topic, and as a users' guide for practitioners interested in applying the representative algorithms and understanding theoretical results under various technical assumptions. We hope that this survey will facilitate discussion on the open problems in this topic, and more importantly, shed light on future research directions. Due to the page limit, we suggest the readers refer to the full version of this survey https://arxiv.org/abs/2004.11154.
Fanghui Liu 0001, Xiaolin Huang, Yudong Chen 0001, Johan A. K. Suykens
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Towards a Unified Quadrature Framework for Large-Scale Kernel Machines
abstract
In this paper, we develop a quadrature framework for large-scale kernel machines via a numerical integration representation. Considering that the integration domain and measure of typical kernels, e.g., Gaussian kernels, arc-cosine kernels, are fully symmetric, we leverage a numerical integration technique, deterministic fully symmetric interpolatory rules, to efficiently compute quadrature nodes and associated weights for kernel approximation. Thanks to the full symmetric property, the applied interpolatory rules are able to reduce the number of needed nodes while retaining a high approximation accuracy. Further, we randomize the above deterministic rules by the classical Monte-Carlo sampling and control variates techniques with two merits: 1) The proposed stochastic rules make the dimension of the feature mapping flexibly varying, such that we can control the discrepancy between the original and approximate kernels by tuning the dimnension. 2) Our stochastic rules have nice statistical properties of unbiasedness and variance reduction. In addition, we elucidate the relationship between our deterministic/stochastic interpolatory rules and current typical quadrature based rules for kernel approximation, thereby unifying these methods under our framework. Experimental results on several benchmark datasets show that our methods compare favorably with other representative kernel approximation based methods.
Fanghui Liu 0001, Xiaolin Huang, Yudong Chen 0001, Johan A. K. Suykens
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Fast Learning in Reproducing Kernel Krein Spaces via Signed Measures
abstract
In this paper, we attempt to solve a long-lasting open question for non-positive definite (non-PD) kernels in machine learning community: can a given non-PD kernel be decomposed into the difference of two PD kernels (termed as positive decomposition)? We cast this question as a distribution view by introducing the signed measure, which transforms positive decomposition to measure decomposition: a series of non-PD kernels can be associated with the linear combination of specific finite Borel measures. In this manner, our distribution-based framework provides a sufficient and necessary condition to answer this open question. Specifically, this solution is also computationally implementable in practice to scale non-PD kernels in large sample cases, which allows us to devise the first random features algorithm to obtain an unbiased estimator. Experimental results on several benchmark datasets verify the effectiveness of our algorithm over the existing methods.
Fanghui Liu 0001, Xiaolin Huang, Yingyi Chen, Johan A. K. Suykens
AISTATS1
2021 Kernel regression in high dimensions: Refined analysis beyond double descent
abstract
In this paper, we provide a precise characterization of generalization properties of high dimensional kernel ridge regression across the under- and over-parameterized regimes, depending on whether the number of training data n exceeds the feature dimension d. By establishing a bias-variance decomposition of the expected excess risk, we show that, while the bias is (almost) independent of d and monotonically decreases with n, the variance depends on n,d and can be unimodal or monotonically decreasing under different regularization schemes. Our refined analysis goes beyond the double descent theory by showing that, depending on the data eigen-profile and the level of regularization, the kernel regression risk curve can be a double-descent-like, bell-shaped, or monotonic function of n. Experiments on synthetic and real data are conducted to support our theoretical findings.
Fanghui Liu 0001, Zhenyu Liao 0001, Johan A. K. Suykens
AISTATS1
2021 Generalization Properties of hyper-RKHS and its Applications
abstract
This paper generalizes regularized regression problems in a hyper-reproducing kernel Hilbert space (hyper-RKHS), illustrates its utility for kernel learning and out-of-sample extensions, and proves asymptotic convergence results for the introduced regression models in an approximation theory view. Algorithmically, we consider two regularized regression models with bivariate forms in this space, including kernel ridge regression (KRR) and support vector regression (SVR) endowed with hyper-RKHS, and further combine divide-and-conquer with Nyström approximation for scalability in large sample cases. This framework is general: the underlying kernel is learned from a broad class, and can be positive definite or not, which adapts to various requirements in kernel learning. Theoretically, we study the convergence behavior of regularized regression algorithms in hyper-RKHS and derive the learning rates, which goes beyond the classical analysis on RKHS due to the non-trivial independence of pairwise samples and the characterisation of hyper-RKHS. Experimentally, results on several benchmarks suggest that the employed framework is able to learn a general kernel function form an arbitrary similarity matrix, and thus achieves a satisfactory performance on classification tasks.
Fanghui Liu 0001, Lei Shi 0010, Xiaolin Huang, Jie Yang 0002, Johan A. K. Suykens
J. Mach. Learn. Res.1
2021 Analysis of regularized least-squares in reproducing kernel Kreĭn spaces
Fanghui Liu 0001, Lei Shi 0010, Xiaolin Huang, Jie Yang 0002, Johan A. K. Suykens
Mach. Learn.1
2020 Random Fourier Features via Fast Surrogate Leverage Weighted Sampling
abstract
In this paper, we propose a fast surrogate leverage weighted sampling strategy to generate refined random Fourier features for kernel approximation. Compared to the current state-of-the-art method that uses the leverage weighted scheme (Li et al. 2019), our new strategy is simpler and more effective. It uses kernel alignment to guide the sampling process and it can avoid the matrix inversion operator when we compute the leverage function. Given n observations and s random features, our strategy can reduce the time complexity for sampling from O(ns2+s3) to O(ns2), while achieving comparable (or even slightly better) prediction performance when applied to kernel ridge regression (KRR). In addition, we provide theoretical guarantees on the generalization performance of our approach, and in particular characterize the number of random features required to achieve statistical guarantees in KRR. Experiments on several benchmark datasets demonstrate that our algorithm achieves comparable prediction performance and takes less time cost when compared to (Li et al. 2019).
Fanghui Liu 0001, Xiaolin Huang, Yudong Chen 0001, Jie Yang 0002, Johan A. K. Suykens
AAAI1
2020 Learning Data-adaptive Non-parametric Kernels
abstract
In this paper, we propose a data-adaptive non-parametric kernel learning framework in margin based kernel methods. In model formulation, given an initial kernel matrix, a data-adaptive matrix with two constraints is imposed in an entry-wise scheme. Learning this data-adaptive matrix in a formulation-free strategy enlarges the margin between classes and thus improves the model flexibility. The introduced two constraints are imposed either exactly (on small data sets) or approximately (on large data sets) in our model, which provides a controllable trade-off between model flexibility and complexity with theoretical demonstration. In algorithm optimization, the objective function of our learning framework is proven to be gradient-Lipschitz continuous. Thereby, kernel and classifier/regressor learning can be efficiently optimized in a unified framework via Nesterov's acceleration. For the scalability issue, we study a decomposition-based approach to our model in the large sample case. The effectiveness of this approximation is illustrated by both empirical studies and theoretical guarantees. Experimental results on various classification and regression benchmark data sets demonstrate that our non-parametric kernel learning framework achieves good performance when compared with other representative kernel learning based algorithms.
Fanghui Liu 0001, Xiaolin Huang, Chen Gong 0002, Jie Yang 0002, Li Li 0013
J. Mach. Learn. Res.1
2020 A Double-Variational Bayesian Framework in Random Fourier Features for Indefinite Kernels
abstract
Random Fourier features (RFFs) have been successfully employed to kernel approximation in large-scale situations. The rationale behind RFF relies on Bochner's theorem, but the condition is too strict and excludes many widely used kernels, e.g., dot-product kernels (violates the shift-invariant condition) and indefinite kernels [violates the positive definite (PD) condition]. In this article, we present a unified RFF framework for indefinite kernel approximation in the reproducing kernel Kreĭn spaces (RKKSs). Besides, our model is also suited to approximate a dot-product kernel on the unit sphere, as it can be transformed into a shift-invariant but indefinite kernel. By the Kolmogorov decomposition scheme, an indefinite kernel in RKKS can be decomposed into the difference of two unknown PD kernels. The spectral distribution of each underlying PD kernel can be formulated as a nonparametric Bayesian Gaussian mixtures model. Based on this, we propose a double-infinite Gaussian mixture model in RFF by placing the Dirichlet process prior. It takes full advantage of high flexibility on the number of components and has the capability of approximating indefinite kernels on a wide scale. In model inference, we develop a non-conjugate variational algorithm with a sub-sampling scheme for the posterior inference. It allows for the non-conjugate case in our model and is quite efficient due to the sub-sampling strategy. Experimental results on several large classification data sets demonstrate the effectiveness of our nonparametric Bayesian model for indefinite kernel approximation when compared to other representative random feature-based methods.
Fanghui Liu 0001, Xiaolin Huang, Lei Shi 0010, Jie Yang 0002, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.1
2019 Indefinite Kernel Logistic Regression With Concave-Inexact-Convex Procedure
abstract
In kernel methods, the kernels are often required to be positive definitethat restricts the use of many indefinite kernels. To consider those nonpositive definite kernels, in this paper, we aim to build an indefinite kernel learning framework for kernel logistic regression (KLR). The proposed indefinite KLR (IKLR) model is analyzed in the reproducing kernel Kreĭn spaces and then becomes nonconvex. Using the positive decomposition of a nonpositive definite kernel, the derived IKLR model can be decomposed into the difference of two convex functions. Accordingly, a concave-convex procedure (CCCP) is introduced to solve the nonconvex optimization problem. Since the CCCP has to solve a subproblem in each iteration, we propose a concave-inexact-convex procedure (CCICP) algorithm with an inexact solving scheme to accelerate the solving process. Besides, we propose a stochastic variant of CCICP to efficiently obtain a proximal solution, which achieves the similar purpose with the inexact solving scheme in CCICP. The convergence analyses of the above-mentioned two variants of CCCP are conducted. By doing so, our method works effectively not only in a deterministic setting but also in a stochastic setting. Experimental results on several benchmarks suggest that the proposed IKLR model performs favorably against the standard (positive definite) KLR and other competitive indefinite learning-based algorithms.
Fanghui Liu 0001, Xiaolin Huang, Chen Gong 0002, Jie Yang 0002, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.1
2018 Nonlinear Pairwise Layer and Its Training for Kernel Learning
abstract
Kernel learning is a fundamental technique that has been intensively studied in the past decades. For the complicated practical tasks, the traditional "shallow" kernels (e.g., Gaussian kernel and sigmoid kernel) are not flexible enough to produce satisfactory performance. To address this shortcoming, this paper introduces a nonlinear layer in kernel learning to enhance the model flexibility. This layer is pairwise, which fully considers the coupling information among examples. So our model contains a fixed single mapping layer (i.e. a Gaussian kernel) as well as a nonlinear pairwise layer, thereby achieving better flexibility than the existing kernel structures. Moreover, the proposed structure can be seamlessly embedded to Support Vector Machines (SVM), of which the training process can be formulated as a joint optimization problem including nonlinear function learning and standard SVM optimization. We theoretically prove that the objective function is gradient-Lipschitz continuous, which further guides us how to accelerate the optimization process in a deep kernel architecture. Experimentally, we find that the proposed structure outperforms other state-ofthe-art kernel-based algorithms on various benchmark datasets, and thus the effectiveness of the incorporated pairwise layer with its training approach is demonstrated.
Fanghui Liu 0001, Xiaolin Huang, Chen Gong 0002, Jie Yang 0002, Li Li 0013
AAAI1
2018 Deep-PUMR: Deep Positive and Unlabeled Learning with Manifold Regularization
Fanghui Liu 0001, Enmei Tu, Longbing Cao, Jie Yang 0002
ICONIP (1)2
2018 Online discriminative dictionary learning for robust object tracking
Tao Zhou 0002, Fanghui Liu 0001, Harish Bhaskar, Jie Yang 0002, Huanlong Zhang, Ping Cai
Neurocomputing2
2018 Densely Connected Discriminative Correlation Filters for Visual Tracking
abstract
Discriminative Correlation Filters (DCFs)-based approaches have recently achieved competitive performance in visual tracking. However, such conventional DCF-based trackers often lack the discriminative ability due to the shallow architecture. As a result, they can hardly tackle drastic appearance variations and easily drift when the target suffers heavy occlusions. To address this issue, a novel densely connected DCFs framework is proposed for visual tracking. We incorporate multiple nested DCFs into the deep learning architecture, and then train the compact network with the data-specific target. Specifically, feature maps and interim response maps are shared and reused throughout the whole network. By doing so, the implicit information carried out by each DCF is fully exploited to enhance the model representation ability during the tracking process. Moreover, a multiscale estimation scheme is developed to account for scale variations. Experimental results on the benchmarks demonstrate that the proposed approach achieves outstanding performance compared to the existing state-of-the-art trackers.
Cheng Peng 0004, Fanghui Liu 0001, Jie Yang 0002, Nikola K. Kasabov
IEEE Signal Process. Lett.2
2018 Robust Visual Tracking via Dirac-Weighted Cascading Correlation Filters
abstract
Correlation filter-based trackers (CFTs) have recently raised considerable attention in visual tracking and achieved competitive performance. Nevertheless, conventional structures of such CFTs build a shallow architecture with a single correlation filter, which cannot comprehensively depict the target appearance. Hence, these trackers lack strong discriminative ability and easily drift when the target suffers drastic appearance variations. To address the limitations, we propose Dirac-weighted cascading correlation filters (DWCCF) for visual tracking. It incorporates cascading characteristics of multiple filters to construct the target appearance model, and dynamically learns Dirac weights for each filter, which is accordingly robust to appearance variations. Besides, we design a boundary penalization strategy to adaptively reduce the boundary effects, which efficiently improves the detection precision for tracking. Qualitative and quantitative evaluations on OTB-2013 and OTB-2015 datasets demonstrate that the proposed DWCCF significantly outperforms other state-of-the-art methods.
Cheng Peng 0004, Fanghui Liu 0001, Jie Yang 0002, Nikola K. Kasabov
IEEE Signal Process. Lett.2
2018 Inverse Nonnegative Local Coordinate Factorization for Visual Tracking
abstract
Recently, nonnegative matrix factorization (NMF) with part-based representation has been widely used for appearance modeling in visual tracking. Unfortunately, not all the targets can be successfully decomposed as “parts” unless some rigorous conditions are satisfied. To avoid this problem, this paper introduces NMF's variants into the visual tracking framework in the view of data clustering for appearance modeling. First, an initial target appearance model based on NMF is proposed to describe the target's appearance with the incorporated local coordinate factorization constraint, orthogonality of the bases, and L1,1norm regularized sparse residual error constraint. Second, an inverse NMF model is proposed in which each learned base vector is regarded as a clustering center in a low-dimensional subspace. Potential target samples (from the foreground) will be clustered around base vectors, while the candidate samples (from the background) are very likely to spread irregularly over the entire clustering space. Such differences can be fully exploited by the inverse NMF model to produce more discriminative encoding vectors than the conventional NMF method. Furthermore, incremental updating model is introduced into the tracking framework for online updating the initial appearance model. Experiments on object tracking benchmark suggest that our tracker is able to achieve promising performance when compared with some state-of-the-art methods in deformation, occlusion, and other challenging situations.
Fanghui Liu 0001, Tao Zhou 0002, Chen Gong 0002, Keren Fu, Li Bai 0001, Jie Yang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2018 Robust Visual Tracking via Online Discriminative and Low-Rank Dictionary Learning
abstract
In this paper, we propose a novel and robust tracking framework based on online discriminative and low-rank dictionary learning. The primary aim of this paper is to obtain compact and low-rank dictionaries that can provide good discriminative representations of both target and background. We accomplish this by exploiting the recovery ability of low-rank matrices. That is if we assume that the data from the same class are linearly correlated, then the corresponding basis vectors learned from the training set of each class shall render the dictionary to become approximately low-rank. The proposed dictionary learning technique incorporates a reconstruction error that improves the reliability of classification. Also, a multiconstraint objective function is designed to enable active learning of a discriminative and robust dictionary. Further, an optimal solution is obtained by iteratively computing the dictionary, coefficients, and by simultaneously learning the classifier parameters. Finally, a simple yet effective likelihood function is implemented to estimate the optimal state of the target during tracking. Moreover, to make the dictionary adaptive to the variations of the target and background during tracking, an online update criterion is employed while learning the new dictionary. Experimental results on a publicly available benchmark dataset have demonstrated that the proposed tracking algorithm performs better than other state-of-the-art trackers.
Tao Zhou 0002, Fanghui Liu 0001, Harish Bhaskar, Jie Yang 0002
IEEE Trans. Cybern.2
2018 Robust Visual Tracking Revisited: From Correlation Filter to Template Matching
abstract
In this paper, we propose a novel matching based tracker by investigating the relationship between template matching and the recent popular correlation filter based trackers (CFTs). Compared to the correlation operation in CFTs, a sophisticated similarity metric termed mutual buddies similarity is proposed to exploit the relationship of multiple reciprocal nearest neighbors for target matching. By doing so, our tracker obtains powerful discriminative ability on distinguishing target and background as demonstrated by both empirical and theoretical analyses. Besides, instead of utilizing single template with the improper updating scheme in CFTs, we design a novel online template updating strategy named memory, which aims to select a certain amount of representative and reliable tracking results in history to construct the current stable and expressive template set. This scheme is beneficial for the proposed tracker to comprehensively understand the target appearance variations, recall some stable results. Both qualitative and quantitative evaluations on two benchmarks suggest that the proposed tracking method performs favorably against some recently developed CFTs and other competitive trackers.
Fanghui Liu 0001, Chen Gong 0002, Xiaolin Huang, Tao Zhou 0002, Jie Yang 0002, Dacheng Tao
IEEE Trans. Image Process.1
2017 Visual tracking via structural patch-based dictionary pair learning
abstract
In this paper, a novel visual tracking framework based on Structural Patch-based Dictionary Pair Learning (SPDPL), is proposed. The proposed representation model encapsulates partial and spatial structural variations of the target through a novel dictionary learning scheme. The proposed method facilitates learning a robust and discriminative dictionary by considering all patches from the same part of the target region as one class, thus transforming the tracking problem into a multi-class classification and reconstruction task. Finally, a simple yet effective observation model is designed to obtain the most optimal candidate during tracking. Systems experiments of the proposed tracking algorithm on the Object Tracker Benchmark (OTB) dataset have demonstrated improvements against several other state-of-the-art trackers.
Tao Zhou 0002, Fanghui Liu 0001, Harish Bhaskar, Jie Yang 0002, Ping Cai
ICIP2
2017 Robust Kernel Approximation for Classification
Fanghui Liu 0001, Xiaolin Huang, Cheng Peng 0004, Jie Yang 0002, Nikola K. Kasabov
ICONIP (1)1
2017 Correlation Filters with Adaptive Memories and Fusion for Visual Tracking
Cheng Peng 0004, Fanghui Liu 0001, Jie Yang 0002, Nikola K. Kasabov
ICONIP (3)2
2017 Indefinite Kernel Logistic Regression
abstract
Traditionally, kernel learning methods require positive definitiveness on the kernel, which is too strict and excludes many sophisticated similarities, that are indefinite. To utilize those indefinite kernels, indefinite learning methods are of great interests. This paper aims at the extension of the logistic regression from positive definite kernels to indefinite ones. The proposed model, named indefinite kernel logistic regression (IKLR), keeps consistency to the regular KLR in formulation but it essentially becomes non-convex. Thanks to the positive decomposition of an indefinite kernel, IKLR can be transformed into a difference of two convex models, which follows the use of concave-convex procedure. Moreover, aiming at large-scale problems in practice, a concave-inexact-convex procedure (CCICP) algorithm with an inexact solving scheme is proposed with convergence guarantees. Experimental results on multi-modal datasets demonstrate the superiority of the proposed IKLR model over kernel logistic regression with positive definite kernels and other state-of-the-art indefinite learning based methods.
Fanghui Liu 0001, Xiaolin Huang, Jie Yang 0002
ACM Multimedia1
2017 Online learning and joint optimization of combined spatial-temporal models for robust visual tracking
Tao Zhou 0002, Harish Bhaskar, Fanghui Liu 0001, Jie Yang 0002, Ping Cai
Neurocomputing3
2017 Kernelized temporal locality learning for real-time visual tracking
Fanghui Liu 0001, Tao Zhou 0002, Keren Fu, Jie Yang 0002
Pattern Recognit. Lett.1
2017 Graph Regularized and Locality-Constrained Coding for Robust Visual Tracking
abstract
Visual tracking is complicated due to factors, such as occlusion, background clutter, abrupt target motion, and illumination variations, among others. In recent years, subspace representation and sparse coding techniques have demonstrated significant improvements in tracking. However, performance gain in tracking has been at the expense of losing locality and similarity attributes among the instances to be encoded. In this paper, a graph regularized and locality-constrained coding (GRLC) technique that encapsulates local manifold structure of the data in order to preserve locality and similarity information among instances is proposed. The GRLC methodology incorporates a similarity-preserving term within the objective function of the locality-constrained linear coding model, thereby overcoming some of the inherent instability issues common to such coding methods. In the proposed GRLC scheme, a graph Laplacian regularizer is chosen as a smoothing operator to learn both the representation dictionary and the coefficients by preserving the local structure of the data. This graph Laplacian smoothing operator ensures that the representations vary smoothly along the geodesics of the data manifold. Thus, by deriving the objective function of the GRLC method, a discriminative dictionary of instances can be iteratively obtained and the corresponding coefficients for each candidate can be computed using this learned dictionary. Finally, an effective observation likelihood function based on reconstruction error and a simple dictionary update scheme for visual target tracking are also proposed. Experimental results on the CVPR2013 visual tracker benchmark have demonstrated a favorable performance of the proposed technique both in terms of accuracy and robustness.
Tao Zhou 0002, Harish Bhaskar, Fanghui Liu 0001, Jie Yang 0002
IEEE Trans. Circuits Syst. Video Technol.3
2017 Visual Tracking via Nonnegative Multiple Coding
abstract
It has been extensively observed that an accurate appearance model is critical to achieving satisfactory performance for robust object tracking. Most existing top-ranked methods rely on linear representation over a single dictionary, which brings about improper understanding on the target appearance. To address this problem, in this paper, we propose a novel appearance model named as “nonnegative multiple coding” (NMC) to accurately represent a target. First, a series of local dictionaries are created with different predefined numbers of nearest neighbors, and then the contributions of these dictionaries are automatically learned. As a result, this ensemble of dictionaries can comprehensively exploit the appearance information carried by all the constituted dictionaries. Second, the existing methods explicitly impose the nonnegative constraint to coefficient vectors, but in the proposed model, we directly deploy an efficient 12 norm regularization to achieve the similar nonnegative purpose with theoretical guarantees. Moreover, an efficient occlusion detection scheme is designed to alleviate tracking drifts, which investigates whether negative templates are selected to represent the severely occluded target. Experimental results on two benchmarks demonstrate that our NMC tracker are able to achieve superior performance to state-of-the-art methods.
Fanghui Liu 0001, Chen Gong 0002, Tao Zhou 0002, Keren Fu, Xiangjian He, Jie Yang 0002
IEEE Trans. Multim.1
2016 Robust visual tracking via inverse nonnegative matrix factorization
abstract
The establishment of robust target appearance model over time is an overriding concern in visual tracking. In this paper, we propose an inverse nonnegative matrix factorization (NMF) method for robust appearance modeling. Rather than using a linear combination of nonnegative basis vectors for each target image patch in conventional NMF, the proposed method is a reverse thought to conventional NMF tracker. It utilizes both the foreground and background information, and imposes a local coordinate constraint, where the basis matrix is sparse matrix from the linear combination of candidates with corresponding nonnegative coefficient vectors. Inverse NMF is used as a feature encoder, where the resulting coefficient vectors are fed into a SVM classifier for separating the target from the background. The proposed method is tested on several videos and compared with seven state-of-the-art methods. Our results have provided further support to the effectiveness and robustness of the proposed method.
Fanghui Liu 0001, Tao Zhou 0002, Keren Fu, Irene Y. H. Gu, Jie Yang 0002
ICASSP1
2016 Correlation filter tracking via bootstrap learning
abstract
In recent years, correlation filter based trackers outperform better than other trackers. Nevertheless, they only employ one feature and a single kernel, so they are usually not robust in complex scenes. In this paper, we derive a multi-feature and multi-kernel correlation filter based tracker which fully takes advantage of the invariance-discriminative power spectrums of various features and kernels to further improve the performance. A novel bootstrap learning method is utilized to obtain a strong classifier by fusing these weak kernel correlation filters (KCFs). Moreover, a new target scale estimation strategy is incorporated into our framework. The efficient and effective scale estimation method is based on target dictionary representation. The proposed method is tested on several videos and compared with seven state-of-the-art methods. Experimental results have provided further support to the effectiveness and robustness of the proposed method.
Kunqi Gu, Tao Zhou 0002, Fanghui Liu 0001, Jie Yang 0002, Yu Qiao 0003
ICIP3
2016 Incremental Robust Nonnegative Matrix Factorization for Object Tracking
Fanghui Liu 0001, Mingna Liu, Tao Zhou 0002, Yu Qiao 0003, Jie Yang 0002
ICONIP (2)1
2016 Geometric affine transformation estimation via correlation filter for visual tracking
Fanghui Liu 0001, Tao Zhou 0002, Jie Yang 0002
Neurocomputing1
2016 Robust visual tracking via constrained correlation filter coding
Fanghui Liu 0001, Tao Zhou 0002, Keren Fu, Jie Yang 0002
Pattern Recognit. Lett.1
2016 Co-saliency detection via inter and intra saliency propagation
Chenjie Ge, Keren Fu, Fanghui Liu 0001, Li Bai 0001, Jie Yang 0002
Signal Process. Image Commun.3