Bamdev Mishra

dblp:133/8291 · DBLP profile ↗
← Back
32ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0001-7430-2843ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Federated Learning on Riemannian Manifolds with Differential Privacy
Zhenwei Huang, Wen Huang 0001, Pratik Jawanpuria, Bamdev Mishra
Mach. Learn.4
2025 A Riemannian Approach to Ground Metric Learning for Optimal Transport
abstract
Optimal transport (OT) theory has attracted much attention in machine learning and signal processing applications. OT defines a notion of distance between probability distributions of source and target data points. A crucial factor that influences OT-based distances is the ground metric of the embedding space in which the source and target data points lie. In this work, we propose to learn a suitable latent ground metric parameterized by a symmetric positive definite matrix. We use the rich Riemannian geometry of symmetric positive definite matrices to jointly learn the OT distance along with the ground metric. Empirical results illustrate the efficacy of the learned metric in OT-based domain adaptation.
Pratik Jawanpuria, Dai Shi, Bamdev Mishra, Junbin Gao
ICASSP3
2024 Riemannian coordinate descent algorithms on matrix manifolds
abstract
Many machine learning applications are naturally formulated as optimization problems on Riemannian manifolds. The main idea behind Riemannian optimization is to maintain the feasibility of the variables while moving along a descent direction on the manifold. This results in updating all the variables at every iteration. In this work, we provide a general framework for developing computationally efficient coordinate descent (CD) algorithms on matrix manifolds that allows updating only a few variables at every iteration while adhering to the manifold constraint. In particular, we propose CD algorithms for various manifolds such as Stiefel, Grassmann, (generalized) hyperbolic, symplectic, and symmetric positive (semi)definite. While the cost per iteration of the proposed CD algorithms is low, we further develop a more efficient variant via a first-order approximation of the objective function. We analyze their convergence and complexity, and empirically illustrate their efficacy in several applications.
Andi Han, Pratik Jawanpuria, Bamdev Mishra
ICML3
2024 Submodular framework for structured-sparse optimal transport
abstract
Unbalanced optimal transport (UOT) has recently gained much attention due to its flexible framework for handling un-normalized measures and its robustness properties. In this work, we explore learning (structured) sparse transport plans in the UOT setting, i.e., transport plans have an upper bound on the number of non-sparse entries in each column (structured sparse pattern) or in the whole plan (general sparse pattern). We propose novel sparsity-constrained UOT formulations building on the recently explored maximum mean discrepancy based UOT. We show that the proposed optimization problem is equivalent to the maximization of a weakly submodular function over a uniform matroid or a partition matroid. We develop efficient gradient-based discrete greedy algorithms and provide the corresponding theoretical guarantees. Empirically, we observe that our proposed greedy algorithms select a diverse support set and we illustrate the efficacy of the proposed approach in various applications.
Piyushi Manupriya, Pratik Jawanpuria, Karthik S. Gurumoorthy, Saketha Nath Jagarlapudi, Bamdev Mishra
ICML5
2024 A Gauss-Newton Approach for Min-Max Optimization in Generative Adversarial Networks
abstract
A novel first-order method is proposed for training generative adversarial networks (GANs). It modifies the Gauss-Newton method to approximate the min-max Hessian and uses the Sherman-Morrison inversion formula to calculate the inverse. The method corresponds to a fixed-point method that ensures necessary contraction. To evaluate its effectiveness, numerical experiments are conducted on various datasets commonly used in image generation tasks, such as MNIST, Fashion MNIST, CIFAR10, FFHQ, and LSUN. Our method is capable of generating high-fidelity images with greater diversity across multiple datasets. It also achieves the highest inception score for CIFAR10 among all compared methods, including state-of-the-art second-order methods. Additionally, its execution time is comparable to that of first-order min-max methods.
Neel Mishra, Pawan Kumar 0001, Pratik Jawanpuria, Bamdev Mishra
IJCNN4
2024 SLTrain: a sparse plus low rank approach for parameter and memory efficient pretraining
abstract
Large language models (LLMs) have shown impressive capabilities across various tasks. However, training LLMs from scratch requires significant computational power and extensive memory capacity. Recent studies have explored low-rank structures on weights for efficient fine-tuning in terms of parameters and memory, either through low-rank adaptation or factorization. While effective for fine-tuning, low-rank structures are generally less suitable for pretraining because they restrict parameters to a low-dimensional subspace. In this work, we propose to parameterize the weights as a sum of low-rank and sparse matrices for pretraining, which we call SLTrain. The low-rank component is learned via matrix factorization, while for the sparse component, we employ a simple strategy of uniformly selecting the sparsity support at random and learning only the non-zero entries with the fixed support. While being simple, the random fixed-support sparse learning strategy significantly enhances pretraining when combined with low-rank learning. Our results show that SLTrain adds minimal extra parameters and memory costs compared to pretraining with low-rank parameterization, yet achieves substantially better performance, which is comparable to full-rank training. Remarkably, when combined with quantization and per-layer updates, SLTrain can reduce memory requirements by up to 73% when pretraining the LLaMA 7B model.
Andi Han, Wei Huang 0034, Mingyi Hong 0001, Akiko Takeda, Pratik Jawanpuria, Bamdev Mishra
NeurIPS7
2024 A Framework for Bilevel Optimization on Riemannian Manifolds
abstract
Bilevel optimization has gained prominence in various applications. In this study, we introduce a framework for solving bilevel optimization problems, where the variables in both the lower and upper levels are constrained on Riemannian manifolds. We present several hypergradient estimation strategies on manifolds and analyze their estimation errors. Furthermore, we provide comprehensive convergence and complexity analyses for the proposed hypergradient descent algorithm on manifolds. We also extend our framework to encompass stochastic bilevel optimization and incorporate the use of general retraction. The efficacy of the proposed framework is demonstrated through several applications.
Andi Han, Bamdev Mishra, Pratik Jawanpuria, Akiko Takeda
NeurIPS2
2024 Differentially private Riemannian optimization
abstract
Abstract In this paper, we study the differentially private empirical risk minimization problem where the parameter is constrained to a Riemannian manifold. We introduce a framework for performing differentially private Riemannian optimization by adding noise to the Riemannian gradient on the tangent space. The noise follows a Gaussian distribution intrinsically defined with respect to the Riemannian metric on the tangent space. We adapt the Gaussian mechanism from the Euclidean space to the tangent space compatible to such generalized Gaussian distribution. This approach presents a novel analysis as compared to directly adding noise on the manifold. We further prove privacy guarantees of the proposed differentially private Riemannian (stochastic) gradient descent using an extension of the moments accountant technique. Overall, we provide utility guarantees under geodesic (strongly) convex, general nonconvex objectives as well as under the Riemannian Polyak-Łojasiewicz condition. Empirical results illustrate the versatility and efficacy of the proposed framework in several applications.
Andi Han, Bamdev Mishra, Pratik Jawanpuria, Junbin Gao
Mach. Learn.2
2024 Riemannian block SPD coupling manifold and its application to optimal transport
abstract
Abstract In this work, we study the optimal transport (OT) problem between symmetric positive definite (SPD) matrix-valued measures. We formulate the above as a generalized optimal transport problem where the cost, the marginals, and the coupling are represented as block matrices and each component block is a SPD matrix. The summation of row blocks and column blocks in the coupling matrix are constrained by the given block-SPD marginals. We endow the set of such block-coupling matrices with a novel Riemannian manifold structure. This allows to exploit the versatile Riemannian optimization framework to solve generic SPD matrix-valued OT problems. We illustrate the usefulness of the proposed approach in several applications.
Andi Han, Bamdev Mishra, Pratik Jawanpuria, Junbin Gao
Mach. Learn.2
2023 Riemannian Accelerated Gradient Methods via Extrapolation
abstract
In this paper, we propose a convergence acceleration scheme for general Riemannian optimization problems by extrapolating iterates on manifolds. We show that when the iterates are generated from the Riemannian gradient descent method, the scheme achieves the optimal convergence rate asymptotically and is computationally more favorable than the recently proposed Riemannian Nesterov accelerated gradient methods. A salient feature of our analysis is the convergence guarantees with respect to the use of general retraction and vector transport. Empirically, we verify the practical benefits of the proposed acceleration strategy, including robustness to the choice of different averaging schemes on manifolds.
Andi Han, Bamdev Mishra, Pratik Jawanpuria, Junbin Gao
AISTATS2
2023 Light-weight Deep Extreme Multilabel Classification
abstract
Extreme multi-label (XML) classification refers to the task of supervised multi-label learning that involves a large number of labels. Hence, scalability of the classifier with increasing label dimension is an important consideration. In this paper, we develop a method called LightDXML which modifies the recently developed deep learning based XML framework by using label embeddings instead of feature embedding for negative sampling and iterating cyclically through three major phases: (1) proxy training of label embeddings (2) shortlisting of labels for negative sampling and (3) final classifier training using the negative samples. Consequently, LightDXML also removes the requirement of a re-ranker module, thereby, leading to further savings on time and memory requirements. The proposed method achieves the best of both worlds: while the training time, model size and prediction times are on par or better compared to the tree-based methods, it attains much better prediction accuracy that is on par with the deep learning based methods. Moreover, the proposed approach achieves the best tail-label prediction accuracy over most state-of-the-art XML methods on some of the large datasets.
Istasis Mishra, Arpan Dasgupta, Pratik Jawanpuria, Bamdev Mishra, Pawan Kumar 0001
IJCNN4
2022 ProtoBandit: Efficient Prototype Selection via Multi-Armed Bandits
Arghya Roy Chaudhuri, Pratik Jawanpuria, Bamdev Mishra
ACML3
2021 On Riemannian Optimization over Positive Definite Matrices with the Bures-Wasserstein Geometry
abstract
In this paper, we comparatively analyze the Bures-Wasserstein (BW) geometry with the popular Affine-Invariant (AI) geometry for Riemannian optimization on the symmetric positive definite (SPD) matrix manifold. Our study begins with an observation that the BW metric has a linear dependence on SPD matrices in contrast to the quadratic dependence of the AI metric. We build on this to show that the BW metric is a more suitable and robust choice for several Riemannian optimization problems over ill-conditioned SPD matrices. We show that the BW geometry has a non-negative curvature, which further improves convergence rates of algorithms over the non-positively curved AI geometry. Finally, we verify that several popular cost functions, which are known to be geodesic convex under the AI geometry, are also geodesic convex under the BW geometry. Extensive experiments on various applications support our findings.
Andi Han, Bamdev Mishra, Pratik Jawanpuria, Junbin Gao
NeurIPS2
2021 SPOT: A Framework for Selection of Prototypes Using Optimal Transport
Karthik S. Gurumoorthy, Pratik Jawanpuria, Bamdev Mishra
ECML/PKDD (4)3
2020 Geometry-aware domain adaptation for unsupervised alignment of word embeddings
abstract
We propose a novel manifold based geometric approach for learning unsupervised alignment of word embeddings between the source and the target languages.Our approach formulates the alignment learning problem as a domain adaptation problem over the manifold of doubly stochastic matrices.This viewpoint arises from the aim to align the second order information of the two language spaces.The rich geometry of the doubly stochastic manifold allows to employ efficient Riemannian conjugate gradient algorithm for the proposed formulation.Empirically, the proposed approach outperforms state-of-the-art optimal transport based approach on the bilingual lexicon induction task across several language pairs.The performance improvement is more significant for distant language pairs.
Pratik Jawanpuria, Mayank Meghwanshi, Bamdev Mishra
ACL3
2020 A Simple Approach to Learning Unsupervised Multilingual Embeddings
abstract
Recent progress on unsupervised cross-lingual embeddings in the bilingual setting has given the impetus to learning a shared embedding space for several languages.A popular framework to solve the latter problem is to solve the following two sub-problems jointly: 1) learning unsupervised word alignment between several language pairs, and 2) learning how to map the monolingual embeddings of every language to shared multilingual space.In contrast, we propose a simple approach by decoupling the above two sub-problems and solving them separately, one after another, using existing techniques.We show that this proposed approach obtains surprisingly good performance in tasks such as bilingual lexicon induction, cross-lingual word similarity, multilingual document classification, and multilingual dependency parsing.When distant languages are involved, the proposed approach shows robust behavior and outperforms existing unsupervised multilingual word embedding approaches.
Pratik Jawanpuria, Mayank Meghwanshi, Bamdev Mishra
EMNLP (1)3
2019 Riemannian adaptive stochastic gradient algorithms on matrix manifolds
abstract
Adaptive stochastic gradient algorithms in the Euclidean space have attracted much attention lately. Such explorations on Riemannian manifolds, on the other hand, are relatively new, limited, and challenging. This is because of the intrinsic non-linear structure of the underlying manifold and the absence of a canonical coordinate system. In machine learning applications, however, most manifolds of interest are represented as matrices with notions of row and column subspaces. In addition, the implicit manifold-related constraints may also lie on such subspaces. For example, the Grassmann manifold is the set of column subspaces. To this end, such a rich structure should not be lost by transforming matrices to just a stack of vectors while developing optimization algorithms on manifolds. We propose novel stochastic gradient algorithms for problems on Riemannian matrix manifolds by adapting the row and column subspaces of gradients. Our algorithms are provably convergent and they achieve the convergence rate of order $O(log(T)/sqrt(T))$, where $T$ is the number of iterations. Our experiments illustrate that the proposed algorithms outperform existing Riemannian adaptive stochastic algorithms.
Hiroyuki Kasai, Pratik Jawanpuria, Bamdev Mishra
ICML3
2019 Detection of Review Abuse via Semi-Supervised Binary Multi-Target Tensor Decomposition
abstract
Product reviews and ratings on e-commerce websites provide customers with detailed insights about various aspects of the product such as quality, usefulness, etc. Since they influence customers' buying decisions, product reviews have become a fertile ground for abuse by sellers (colluding with reviewers) to promote their own products or to tarnish the reputation of competitor's products. In this paper, our focus is on detecting such abusive entities (both sellers and reviewers) by applying tensor decomposition on the product reviews data. While tensor decomposition is mostly unsupervised, we formulate our problem as a semi-supervised binary multi-target tensor decomposition, to take advantage of currently known abusive entities. We empirically show that our multi-target semi-supervised model achieves higher precision and recall in detecting abusive entities as compared to unsupervised techniques. Finally, we show that our proposed stochastic partial natural gradient inference for our model empirically achieves faster convergence than stochastic gradient and Online-EM with sufficient statistics.
Anil R. Yelundur, Vineet Chaoji, Bamdev Mishra
KDD3
2019 A Riemannian gossip approach to subspace learning on Grassmann manifold
Bamdev Mishra, Hiroyuki Kasai, Pratik Jawanpuria, Atul Saroop
Mach. Learn.1
2019 Learning Multilingual Word Embeddings in Latent Metric Space: A Geometric Approach
abstract
Abstract We propose a novel geometric approach for learning bilingual mappings given monolingual embeddings and a bilingual dictionary. Our approach decouples the source-to-target language transformation into (a) language-specific rotations on the original embeddings to align them in a common, latent space, and (b) a language-independent similarity metric in this common space to better model the similarity between the embeddings. Overall, we pose the bilingual mapping problem as a classification problem on smooth Riemannian manifolds. Empirically, our approach outperforms previous approaches on the bilingual lexicon induction and cross-lingual word similarity tasks. We next generalize our framework to represent multiple languages in a common latent space. Language-specific rotations for all the languages and a common similarity metric in the latent space are learned jointly from bilingual dictionaries for multiple language pairs. We illustrate the effectiveness of joint learning for multiple languages in an indirect word translation setting.
Pratik Jawanpuria, Arjun Balgovind, Anoop Kunchukuttan, Bamdev Mishra
Trans. Assoc. Comput. Linguistics4
2018 Riemannian stochastic quasi-Newton algorithm with variance reduction and its convergence analysis
abstract
Stochastic variance reduction algorithms have recently become popular for minimizing the average of a large, but finite number of loss functions. The present paper proposes a Riemannian stochastic quasi-Newton algorithm with variance reduction (R-SQN-VR). The key challenges of averaging, adding, and subtracting multiple gradients are addressed with notions of retraction and vector transport. We present convergence analyses of R-SQN-VR on both non-convex and retraction-convex functions under retraction and vector transport operators. The proposed algorithm is evaluated on the Karcher mean computation on the symmetric positive-definite manifold and the low-rank matrix completion on the Grassmann manifold. In all cases, the proposed algorithm outperforms the state-of-the-art Riemannian batch and stochastic gradient algorithms.
Hiroyuki Kasai, Bamdev Mishra
AISTATS3
2018 Inductive Framework for Multi-Aspect Streaming Tensor Completion with Side Information
abstract
Low rank tensor completion is a well studied problem and has applications in various fields. However, in many real world applications the data is dynamic, i.e., new data arrives at different time intervals. As a result, the tensors used to represent the data grow in size. Besides the tensors, in many real world scenarios, side information is also available in the form of matrices which also grow in size with time. The problem of predicting missing values in the dynamically growing tensor is called dynamic tensor completion. Most of the previous work in dynamic tensor completion make an assumption that the tensor grows only in one mode. To the best of our Knowledge, there is no previous work which incorporates side information with dynamic tensor completion. We bridge this gap in this paper by proposing a dynamic tensor completion framework called Side Information infused Incremental Tensor Analysis (SIITA), which incorporates side information and works for general incremental tensors. We also show how non-negative constraints can be incorporated with SIITA, which is essential for mining interpretable latent clusters. We carry out extensive experiments on multiple real world datasets to demonstrate the effectiveness of SIITA in various different settings.
Madhav Nimishakavi, Bamdev Mishra, Manish Gupta 0001, Partha P. Talukdar
CIKM2
2018 A Unified Framework for Structured Low-rank Matrix Learning
abstract
We consider the problem of learning a low-rank matrix, constrained to lie in a linear subspace, and introduce a novel factorization for modeling such matrices. A salient feature of the proposed factorization scheme is it decouples the low-rank and the structural constraints onto separate factors. We formulate the optimization problem on the Riemannian spectrahedron manifold, where the Riemannian framework allows to develop computationally efficient conjugate gradient and trust-region algorithms. Experiments on problems such as standard/robust/non-negative matrix completion, Hankel matrix learning and multi-task learning demonstrate the efficacy of our approach.
Pratik Jawanpuria, Bamdev Mishra
ICML2
2018 Riemannian Stochastic Recursive Gradient Algorithm with Retraction and Vector Transport and Its Convergence Analysis
Hiroyuki Kasai, Bamdev Mishra
ICML3
2018 Inexact trust-region algorithms on Riemannian manifolds
abstract
We consider an inexact variant of the popular Riemannian trust-region algorithm for structured big-data minimization problems. The proposed algorithm approximates the gradient and the Hessian in addition to the solution of a trust-region sub-problem. Addressing large-scale finite-sum problems, we specifically propose sub-sampled algorithms with a fixed bound on sub-sampled Hessian and gradient sizes, where the gradient and Hessian are computed by a random sampling technique. Numerical evaluations demonstrate that the proposed algorithms outperform state-of-the-art Riemannian deterministic and stochastic gradient algorithms across different applications.
Hiroyuki Kasai, Bamdev Mishra
NeurIPS2
2018 A Dual Framework for Low-rank Tensor Completion
abstract
One of the popular approaches for low-rank tensor completion is to use the latent trace norm regularization. However, most existing works in this direction learn a sparse combination of tensors. In this work, we fill this gap by proposing a variant of the latent trace norm that helps in learning a non-sparse combination of tensors. We develop a dual framework for solving the low-rank tensor completion problem. We first show a novel characterization of the dual solution space with an interesting factorization of the optimal solution. Overall, the optimal solution is shown to lie on a Cartesian product of Riemannian manifolds. Furthermore, we exploit the versatile Riemannian optimization framework for proposing computationally efficient trust region algorithm. The experiments illustrate the efficacy of the proposed algorithm on several real-world datasets across applications.
Madhav Nimishakavi, Pratik Jawanpuria, Bamdev Mishra
NeurIPS3
2018 A Unified Framework for Domain Adaptation Using Metric Learning on Manifolds
Sridhar Mahadevan, Bamdev Mishra, Shalini Ghosh
ECML/PKDD (2)2
2017 A Sparse and Low-Rank Optimization Framework for Network Topology Control in Dense Fog-RAN
abstract
In this paper, we propose a sparse and low-rank optimization approach for network topology control in the partially connected fog radio access network (Fog-RAN). In this model, the sparsity of the modeling matrix represents the number of non-connected interference links, while the rank of the matrix models the achievable symmetric degrees-of-freedom (DoF) allocations. This model helps find the network topologies with the maximum number of allowed connected interference links. However, the sparse and low-rank optimization problem turns out to be highly intractable due to the non-convex sparsity objective and non- convex fixed-rank rank constraint. To address the coupled challenges in the objective and constraint, we propose a smoothed Riemannian optimization framework by exploiting the quotient manifold geometry of fixed-rank matrices, followed by a smoothed sparsity inducing surrogate. The proposed Rie- mannian algorithm has much lower computational cost compared with state-of-art matrix factorization parameterized methods. Simulation results further demonstrate the appealing sparsity and low-rankness tradeoff in the proposed model, thereby guiding the network deployment in dense Fog-RAN.
Yuanming Shi, Bamdev Mishra, Xuan Liu 0005, Wei Chen 0002
VTC Spring2
2017 Topological Interference Management With User Admission Control via Riemannian Optimization
abstract
Topological interference management (TIM) provides a promising way to manage interference only based on the network connectivity information. Previous works on the TIM problem mainly focus on using the index coding approach and graph theory to establish conditions of network topologies to achieve the feasibility of topological interference management. In this paper, we propose a novel user admission control approach via sparse and low-rank optimization to maximize the number of admitted users for achieving the feasibility of topological interference management. However, the resulting sparse and low-rank optimization problem is non-convex and highly intractable, for which the conventional convex relaxation approaches are inapplicable, e.g., a simple ℓ1-norm relaxation approach yields the objective unbounded and non-convex. To assist efficient algorithms design for the formulated rank-constrained (i.e., degrees-of-freedom (DoFs) allocation) ℓ0-norm maximization (i.e., user capacity maximization) problem, we propose a novel non-convex but smoothed ℓ1-regularized minimization approach to induce sparsity pattern with bounded objective values. We further develop a Riemannian trust-region algorithm to solve the resulting rank-constrained smooth non-convex optimization problem via exploiting the quotient manifold of fixed-rank matrices. Simulation results demonstrate the effectiveness and optimality of the proposed Riemannian algorithm to maximize the number of admitted users for topological interference management.
Yuanming Shi, Bamdev Mishra, Wei Chen 0002
IEEE Trans. Wirel. Commun.2
2016 Low-rank tensor completion: a Riemannian manifold preconditioning approach
abstract
We propose a novel Riemannian manifold preconditioning approach for the tensor completion problem with rank constraint. A novel Riemannian metric or inner product is proposed that exploits the least-squares structure of the cost function and takes into account the structured symmetry that exists in Tucker decomposition. The specific metric allows to use the versatile framework of Riemannian optimization on quotient manifolds to develop preconditioned nonlinear conjugate gradient and stochastic gradient descent algorithms in batch and online setups, respectively. Concrete matrix representations of various optimization-related ingredients are listed. Numerical comparisons suggest that our proposed algorithms robustly outperform state-of-the-art algorithms across different synthetic and real-world datasets.
Hiroyuki Kasai, Bamdev Mishra
ICML2
2016 Heterogeneous Tensor Decomposition for Clustering via Manifold Optimization
abstract
Tensor clustering is an important tool that exploits intrinsically rich structures in real-world multiarray or Tensor datasets. Often in dealing with those datasets, standard practice is to use subspace clustering that is based on vectorizing multiarray data. However, vectorization of tensorial data does not exploit complete structure information. In this paper, we propose a subspace clustering algorithm without adopting any vectorization process. Our approach is based on a novel heterogeneous Tucker decomposition model taking into account cluster membership information. We propose a new clustering algorithm that alternates between different modes of the proposed heterogeneous tensor model. All but the last mode have closed-form updates. Updating the last mode reduces to optimizing over the multinomial manifold for which we investigate second order Riemannian geometry and propose a trust-region algorithm. Numerical experiments show that our proposed algorithm compete effectively with state-of-the-art clustering algorithms that are based on tensor factorization.
Junbin Gao, Xia Hong 0001, Bamdev Mishra
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 Manopt, a matlab toolbox for optimization on manifolds
Nicolas Boumal, Bamdev Mishra, Pierre-Antoine Absil, Rodolphe Sepulchre
J. Mach. Learn. Res.2