VLDB 2026 Research / reviewers in the wild / expert
Zoltán Szabó 0001
dblp:73/2909-1
· DBLP profile ↗
30ranked-venue papers
13as first author
4since 2021 · last 2025
0000-0001-6183-7603ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 12 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
15 papers |
Kernel, tree and ensemble methods · 41% Probabilistic and Bayesian machine learning · 19% Learning theory · 17% | |
| Theoretical computer science
4 papers |
Information theory · 69% Computational complexity · 31% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods
kernel methods |
1.4 | 5 | 2024 | The Minimax Rate of HSIC Estimation for Translation-Invariant Kernels · NeurIPS 2024 Characteristic and Universal Tensor Product Kernels · J. Mach. Learn. Res. 2017 Gradient-free Hamiltonian Monte Carlo with Efficient Kernel Exponential Families · NIPS 2015 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
hilbert-schmidt independence criterion |
0.8 | 1 | 2024 | The Minimax Rate of HSIC Estimation for Translation-Invariant Kernels · NeurIPS 2024 |
Machine learning › Learning theory › statistical estimation › minimax estimation
minimax rates |
0.8 | 1 | 2024 | The Minimax Rate of HSIC Estimation for Translation-Invariant Kernels · NeurIPS 2024 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space |
0.6 | 2 | 2022 | Handling Hard Affine SDP Shape Constraints in RKHSs · J. Mach. Learn. Res. 2022 Learning Theory for Distribution Regression · J. Mach. Learn. Res. 2016 |
Machine learning › Trustworthy machine learning
shape constraints |
0.6 | 1 | 2022 | Handling Hard Affine SDP Shape Constraints in RKHSs · J. Mach. Learn. Res. 2022 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel machines |
0.4 | 1 | 2020 | Hard Shape-Constrained Kernel Machines · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
quantile regression |
0.4 | 1 | 2020 | Hard Shape-Constrained Kernel Machines · NeurIPS 2020 |
Machine learning › Efficient and distributed learning › model compression › sparsity
structured sparsity |
0.3 | 2 | 2014 | Spatio-temporal Event Classification Using Time-Series Kernel Based Structured Sparsity · ECCV (4) 2014 Online group-structured dictionary learning · CVPR 2011 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
goodness-of-fit testing |
0.3 | 1 | 2017 | A Linear-Time Kernel Goodness-of-Fit Test · NIPS 2017 |
Machine learning › Learning theory › hypothesis testing
independence testing |
0.3 | 1 | 2017 | An Adaptive Test of Independence with Analytic Kernel Embeddings · ICML 2017 |
Machine learning › Kernel, tree and ensemble methods
kernel embedding |
0.3 | 1 | 2017 | An Adaptive Test of Independence with Analytic Kernel Embeddings · ICML 2017 |
Machine learning › Learning theory › hypothesis testing › independence testing
kernel independence test |
0.3 | 1 | 2017 | An Adaptive Test of Independence with Analytic Kernel Embeddings · ICML 2017 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
distribution regression |
0.2 | 1 | 2016 | Learning Theory for Distribution Regression · J. Mach. Learn. Res. 2016 |
Machine learning › Learning theory
statistical learning theory |
0.2 | 1 | 2016 | Learning Theory for Distribution Regression · J. Mach. Learn. Res. 2016 |
Computational complexity › property testing
distribution testing |
0.2 | 1 | 2016 | Interpretable Distribution Features with Maximum Testing Power · NIPS 2016 |
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction |
0.2 | 1 | 2015 | Bayesian Manifold Learning: The Locally Linear Latent Variable Model (LL-LVM) · NIPS 2015 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
hamiltonian monte carlo |
0.2 | 1 | 2015 | Gradient-free Hamiltonian Monte Carlo with Efficient Kernel Exponential Families · NIPS 2015 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel approximation |
0.2 | 1 | 2015 | Optimal Rates for Random Fourier Features · NIPS 2015 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.2 | 1 | 2015 | Bayesian Manifold Learning: The Locally Linear Latent Variable Model (LL-LVM) · NIPS 2015 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › nonlinear manifold learning
locally linear embedding |
0.2 | 1 | 2015 | Bayesian Manifold Learning: The Locally Linear Latent Variable Model (LL-LVM) · NIPS 2015 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning |
0.2 | 1 | 2015 | Bayesian Manifold Learning: The Locally Linear Latent Variable Model (LL-LVM) · NIPS 2015 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.2 | 1 | 2015 | Gradient-free Hamiltonian Monte Carlo with Efficient Kernel Exponential Families · NIPS 2015 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random fourier features |
0.2 | 1 | 2015 | Optimal Rates for Random Fourier Features · NIPS 2015 |
Information theory › estimation theory
entropy estimation |
0.2 | 1 | 2014 | Information theoretical estimators toolbox · J. Mach. Learn. Res. 2014 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning |
0.1 | 1 | 2011 | Online group-structured dictionary learning · CVPR 2011 |
Machine learning › Learning theory
online learning |
0.1 | 1 | 2011 | Online group-structured dictionary learning · CVPR 2011 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
characteristic kernels |
0.1 | 1 | 2017 | Characteristic and Universal Tensor Product Kernels · J. Mach. Learn. Res. 2017 |
Machine learning › Representation and self-supervised learning
blind source separation |
0.1 | 1 | 2007 | Undercomplete Blind Subspace Deconvolution · J. Mach. Learn. Res. 2007 |
Information theory › signal processing
independent component analysis |
0.1 | 1 | 2007 | Undercomplete Blind Subspace Deconvolution · J. Mach. Learn. Res. 2007 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.1 | 1 | 2015 | Bayesian Manifold Learning: The Locally Linear Latent Variable Model (LL-LVM) · NIPS 2015 |
Methods — techniques the papers use, named apart from their topics
kernel methods · 1.1convex optimization · 1.0v-statistic · 0.8u-statistic · 0.8nyström approximation · 0.8second-order cone tightening · 0.6bahadur efficiency analysis · 0.6second-order cone programming · 0.4stein's method · 0.3hilbert-schmidt independence criterion · 0.3analytic kernel embeddings · 0.3semimetric optimization · 0.2maximum mean discrepancy · 0.2mutual information estimation · 0.2entropy estimation · 0.2subspace deconvolution · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Nyström Kernel Stein DiscrepancyabstractKernel methods underpin many of the most successful approaches in data science and statistics, and they allow representing probability measures as elements of a reproducing kernel Hilbert space without loss of information. Recently, the kernel Stein discrepancy (KSD), which combines Stein’s method with the flexibility of kernel techniques, gained considerable attention. Through the Stein operator, KSD allows the construction of powerful goodness-of-fit tests where it is sufficient to know the target distribution up to a multiplicative constant. However, the typical U- and V-statistic-based KSD estimators suffer from a quadratic runtime complexity, which hinders their application in large-scale settings. In this work, we propose a Nystr{ö}m-based KSD acceleration—with runtime $\mathcal{O} \left(mn+m^3\right)$ for $n$ samples and $m\ll n$ Nystr{ö}m points—, show its $\sqrt{n}$-consistency with a classical sub-Gaussian assumption, and demonstrate its applicability for goodness-of-fit testing on a suite of benchmarks. We also show the $\sqrt n$-consistency of the quadratic-time KSD estimator. Florian Kalinke, Zoltán Szabó 0001, Bharath K. Sriperumbudur |
AISTATS | 2 |
| 2024 | The Minimax Rate of HSIC Estimation for Translation-Invariant KernelsabstractKernel techniques are among the most influential approaches in data science and statistics. Under mild conditions, the reproducing kernel Hilbert space associated to a kernel is capable of encoding the independence of $M\ge2$ random variables. Probably the most widespread independence measure relying on kernels is the so-called Hilbert-Schmidt independence criterion (HSIC; also referred to as distance covariance in the statistics literature). Despite various existing HSIC estimators designed since its introduction close to two decades ago, the fundamental question of the rate at which HSIC can be estimated is still open. In this work, we prove that the minimax optimal rate of HSIC estimation on $\mathbb{R}^d$ for Borel measures containing the Gaussians with continuous bounded translation-invariant characteristic kernels is $\mathcal{O}\left(n^{-1/2}\right)$. Specifically, our result implies the optimality in the minimax sense of many of the most-frequently used estimators (including the U-statistic, the V-statistic, and the Nyström-based one) on $\mathbb{R}^d$. Florian Kalinke, Zoltán Szabó 0001 |
NeurIPS | 2 |
| 2023 | Nyström M-Hilbert-Schmidt independence criterionabstractKernel techniques are among the most popular and powerful approaches of data science. Among the key features that make kernels ubiquitous are (i) the number of domains they have been designed for, (ii) the Hilbert structure of the function class associated to kernels facilitating their statistical analysis, and (iii) their ability to represent probability distributions without loss of information. These properties give rise to the immense success of Hilbert-Schmidt independence criterion (HSIC) which is able to capture joint independence of random variables under mild conditions, and permits closed-form estimators with quadratic computational complexity (w.r.t. the sample size). In order to alleviate the quadratic computational bottleneck in large-scale applications, multiple HSIC approximations have been proposed, however these estimators are restricted to $M=2$ random variables, do not extend naturally to the $M\ge 2$ case, and lack theoretical guarantees. In this work, we propose an alternative Nyström-based HSIC estimator which handles the $M\ge 2$ case, prove its consistency, and demonstrate its applicability in multiple contexts, including synthetic examples, dependency testing of media annotations, and causal discovery. Florian Kalinke, Zoltán Szabó 0001 |
UAI | 2 |
| 2022 | Handling Hard Affine SDP Shape Constraints in RKHSsabstractShape constraints, such as non-negativity, monotonicity, convexity or supermodularity, play a key role in various applications of machine learning and statistics. However, incorporating this side information into predictive models in a hard way (for example at all points of an interval) for rich function classes is a notoriously challenging problem. We propose a unified and modular convex optimization framework, relying on second-order cone (SOC) tightening, to encode hard affine SDP constraints on function derivatives, for models belonging to vector-valued reproducing kernel Hilbert spaces (vRKHSs). The modular nature of the proposed approach allows to simultaneously handle multiple shape constraints, and to tighten an infinite number of constraints into finitely many. We prove the convergence of the proposed scheme and that of its adaptive variant, leveraging geometric properties of vRKHSs. Due to the covering-based construction of the tightening, the method is particularly well-suited to tasks with small to moderate input dimensions. The efficiency of the approach is illustrated in the context of shape optimization, safety-critical control, robotics and econometrics. Pierre-Cyril Aubin-Frankowski, Zoltán Szabó 0001 |
J. Mach. Learn. Res. | 2 |
| 2020 | Hard Shape-Constrained Kernel MachinesabstractShape constraints (such as non-negativity, monotonicity, convexity) play a central role in a large number of applications, as they usually improve performance for small sample size and help interpretability. However enforcing these shape requirements in a hard fashion is an extremely challenging problem. Classically, this task is tackled (i) in a soft way (without out-of-sample guarantees), (ii) by specialized transformation of the variables on a case-by-case basis, or (iii) by using highly restricted function classes, such as polynomials or polynomial splines. In this paper, we prove that hard affine shape constraints on function derivatives can be encoded in kernel machines which represent one of the most flexible and powerful tools in machine learning and statistics. Particularly, we present a tightened second-order cone constrained reformulation, that can be readily implemented in convex solvers. We prove performance guarantees on the solution, and demonstrate the efficiency of the approach in joint quantile regression with applications to economics and to the analysis of aircraft trajectories, among others. Pierre-Cyril Aubin-Frankowski, Zoltán Szabó 0001 |
NeurIPS | 2 |
| 2019 | On Kernel Derivative Approximation with Random Fourier FeaturesabstractRandom Fourier features (RFF) represent one of the most popular and wide-spread techniques in machine learning to scale up kernel algorithms. Despite the numerous successful applications of RFFs, unfortunately, quite little is understood theoretically on their optimality and limitations of their performance. Only recently, precise statistical-computational trade-offs have been established for RFFs in the approximation of kernel values, kernel ridge regression, kernel PCA and SVM classification. Our goal is to spark the investigation of optimality of RFF-based approximations in tasks involving not only function values but derivatives, which naturally lead to optimization problems with kernel derivatives. Particularly, in this paper, we focus on the approximation quality of RFFs for kernel derivatives and prove that the existing finite-sample guarantees can be improved exponentially in terms of the domain where they hold, using recent tools from unbounded empirical process theory. Our result implies that the same approximation guarantee is attainable for kernel derivatives using RFF as achieved for kernel values. Zoltán Szabó 0001, Bharath K. Sriperumbudur |
AISTATS | 1 |
| 2017 | An Adaptive Test of Independence with Analytic Kernel EmbeddingsabstractA new computationally efficient dependence measure, and an adaptive statistical test of independence, are proposed. The dependence measure is the difference between analytic embeddings of the joint distribution and the product of the marginals, evaluated at a finite set of locations (features). These features are chosen so as to maximize a lower bound on the test power, resulting in a test that is data-efficient, and that runs in linear time (with respect to the sample size n). The optimized features can be interpreted as evidence to reject the null hypothesis, indicating regions in the joint domain where the joint distribution and the product of the marginals differ most. Consistency of the independence test is established, for an appropriate choice of features. In real-world benchmarks, independence tests using the optimized features perform comparably to the state-of-the-art quadratic-time HSIC test, and outperform competing O(n) and O(n log n) tests. Wittawat Jitkrittum, Zoltán Szabó 0001, Arthur Gretton |
ICML | 2 |
| 2017 | A Linear-Time Kernel Goodness-of-Fit TestabstractWe propose a novel adaptive test of goodness-of-fit, with computational cost linear in the number of samples. We learn the test features that best indicate the differences between observed samples and a reference model, by minimizing the false negative rate. These features are constructed via Stein's method, meaning that it is not necessary to compute the normalising constant of the model. We analyse the asymptotic Bahadur efficiency of the new test, and prove that under a mean-shift alternative, our test always has greater relative efficiency than a previous linear-time kernel test, regardless of the choice of parameters for that test. In experiments, the performance of our method exceeds that of the earlier linear-time test, and matches or exceeds the power of a quadratic-time kernel test. In high dimensions and where model structure may be exploited, our goodness of fit test performs far better than a quadratic-time two-sample test based on the Maximum Mean Discrepancy, with samples drawn from the model. Wittawat Jitkrittum, Zoltán Szabó 0001, Kenji Fukumizu, Arthur Gretton |
NIPS | 3 |
| 2017 | Characteristic and Universal Tensor Product Kernels
Zoltán Szabó 0001, Bharath K. Sriperumbudur |
J. Mach. Learn. Res. | 1 |
| 2016 | Interpretable Distribution Features with Maximum Testing PowerabstractTwo semimetrics on probability distributions are proposed, given as the sum of differences of expectations of analytic functions evaluated at spatial or frequency locations (i.e, features). The features are chosen so as to maximize the distinguishability of the distributions, by optimizing a lower bound on test power for a statistical test using these features. The result is a parsimonious and interpretable indication of how and where two distributions differ locally. An empirical estimate of the test power criterion converges with increasing sample size, ensuring the quality of the returned features. In real-world benchmarks on high-dimensional text and image data, linear-time tests using the proposed semimetrics achieve comparable performance to the state-of-the-art quadratic-time maximum mean discrepancy test, while returning human-interpretable features that explain the test results. Wittawat Jitkrittum, Zoltán Szabó 0001, Kacper Chwialkowski, Arthur Gretton |
NIPS | 2 |
| 2016 | Learning Theory for Distribution RegressionabstractWe focus on the distribution regression problem: regressing to vector-valued outputs from probability measures. Many important machine learning and statistical tasks fit into this framework, including multi-instance learning and point estimation problems without analytical solution (such as hyperparameter or entropy estimation). Despite the large number of available heuristics in the literature, the inherent two-stage sampled nature of the problem makes the theoretical analysis quite challenging, since in practice only samples from sampled distributions are observable, and the estimates have to rely on similarities computed between sets of points. To the best of our knowledge, the only existing technique with consistency guarantees for distribution regression requires kernel density estimation as an intermediate step (which often performs poorly in practice), and the domain of the distributions to be compact Euclidean. In this paper, we study a simple, analytically computable, ridge regression-based alternative to distribution regression, where we embed the distributions to a reproducing kernel Hilbert space, and learn the regressor from the embeddings to the outputs. Our main contribution is to prove that this scheme is consistent in the two-stage sampled setup under mild conditions (on separable topological domains enriched with kernels): we present an exact computational-statistical efficiency trade-off analysis showing that our estimator is able to match the one-stage sampled minimax optimal rate (Caponnetto and De Vito, 2007; Steinwart et al., 2009). This result answers a $17 $-year-old open question, establishing the consistency of the classical set kernel (Haussler, 1999; Gärtner et al., 2002) in regression. We also cover consistency for more recent kernels on distributions, including those due to Christmann and Steinwart (2010). Zoltán Szabó 0001, Bharath K. Sriperumbudur, Barnabás Póczos, Arthur Gretton |
J. Mach. Learn. Res. | 1 |
| 2015 | Two-stage sampled learning theory on distributionsabstractWe focus on the distribution regression problem: regressing to a real-valued response from a probability distribution. Although there exist a large number of similarity measures between distributions, very little is known about their generalization performance in specific learning tasks. Learning problems formulated on distributions have an inherent two-stage sampled difficulty: in practice only samples from sampled distributions are observable, and one has to build an estimate on similarities computed between sets of points. To the best of our knowledge, the only existing method with consistency guarantees for distribution regression requires kernel density estimation as an intermediate step (which suffers from slow convergence issues in high dimensions), and the domain of the distributions to be compact Euclidean. In this paper, we provide theoretical guarantees for a remarkably simple algorithmic alternative to solve the distribution regression problem: embed the distributions to a reproducing kernel Hilbert space, and learn a ridge regressor from the embeddings to the outputs. Our main contribution is to prove the consistency of this technique in the two-stage sampled setting under mild conditions (on separable, topological domains endowed with kernels). As a special case, we answer a 15-year-old open question: we establish the consistency of the classical set kernel [Haussler, 1999; Gaertner et. al, 2002] in regression, and cover more recent kernels on distributions, including those due to [Christmann and Steinwart, 2010]. Zoltán Szabó 0001, Arthur Gretton, Barnabás Póczos, Bharath K. Sriperumbudur |
AISTATS | 1 |
| 2015 | Bayesian Manifold Learning: The Locally Linear Latent Variable Model (LL-LVM)abstractWe introduce the Locally Linear Latent Variable Model (LL-LVM), a probabilistic model for non-linear manifold discovery that describes a joint distribution over observations, their manifold coordinates and locally linear maps conditioned on a set of neighbourhood relationships. The model allows straightforward variational optimisation of the posterior distribution on coordinates and locally linear maps from the latent space to the observation space given the data. Thus, the LL-LVM encapsulates the local-geometry preserving intuitions that underlie non-probabilistic methods such as locally linear embedding (LLE). Its probabilistic semantics make it easy to evaluate the quality of hypothesised neighbourhood relationships, select the intrinsic dimensionality of the manifold, construct out-of-sample extensions and to combine the manifold model with additional probabilistic models that capture the structure of coordinates within the manifold. Mijung Park, Wittawat Jitkrittum, Ahmad Qamar, Zoltán Szabó 0001, Lars Buesing, Maneesh Sahani |
NIPS | 4 |
| 2015 | Optimal Rates for Random Fourier FeaturesabstractKernel methods represent one of the most powerful tools in machine learning to tackle problems expressed in terms of function values and derivatives due to their capability to represent and model complex relations. While these methods show good versatility, they are computationally intensive and have poor scalability to large data as they require operations on Gram matrices. In order to mitigate this serious computational limitation, recently randomized constructions have been proposed in the literature, which allow the application of fast linear algorithms. Random Fourier features (RFF) are among the most popular and widely applied constructions: they provide an easily computable, low-dimensional feature representation for shift-invariant kernels. Despite the popularity of RFFs, very little is understood theoretically about their approximation quality. In this paper, we provide a detailed finite-sample theoretical analysis about the approximation quality of RFFs by (i) establishing optimal (in terms of the RFF dimension, and growing set size) performance guarantees in uniform norm, and (ii) presenting guarantees in L^r (1 ≤ r < ∞) norms. We also propose an RFF approximation to derivatives of a kernel with a theoretical study on its approximation quality. Bharath K. Sriperumbudur, Zoltán Szabó 0001 |
NIPS | 2 |
| 2015 | Gradient-free Hamiltonian Monte Carlo with Efficient Kernel Exponential FamiliesabstractWe propose Kernel Hamiltonian Monte Carlo (KMC), a gradient-free adaptive MCMC algorithm based on Hamiltonian Monte Carlo (HMC). On target densities where classical HMC is not an option due to intractable gradients, KMC adaptively learns the target's gradient structure by fitting an exponential family model in a Reproducing Kernel Hilbert Space. Computational costs are reduced by two novel efficient approximations to this gradient. While being asymptotically exact, KMC mimics HMC in terms of sampling efficiency, and offers substantial mixing improvements over state-of-the-art gradient free samplers. We support our claims with experimental studies on both toy and real-world applications, including Approximate Bayesian Computation and exact-approximate MCMC. Heiko Strathmann, Dino Sejdinovic, Samuel Livingstone, Zoltán Szabó 0001, Arthur Gretton |
NIPS | 4 |
| 2015 | Kernel-Based Just-In-Time Learning for Passing Expectation Propagation Messages
Wittawat Jitkrittum, Arthur Gretton, Nicolas Heess, S. M. Ali Eslami, Balaji Lakshminarayanan, Dino Sejdinovic, Zoltán Szabó 0001 |
UAI | 7 |
| 2014 | Spatio-temporal Event Classification Using Time-Series Kernel Based Structured Sparsity
László A. Jeni, András Lörincz, Zoltán Szabó 0001, Jeffrey F. Cohn, Takeo Kanade |
ECCV (4) | 3 |
| 2014 | Information theoretical estimators toolbox
Zoltán Szabó 0001 |
J. Mach. Learn. Res. | 1 |
| 2013 | Explaining Unintelligible Words by Means of their Context
Balázs Pintér, Gyula Vörös, Zoltán Szabó 0001, András Lörincz |
ICPRAM | 3 |
| 2012 | 3D shape estimation in video sequences provides high precision evaluation of facial expressions
László A. Jeni, András Lörincz, Tamás Nagy, Zsolt Palotai, Judit Sebok, Zoltán Szabó 0001, Dániel Takács |
Image Vis. Comput. | 6 |
| 2012 | Separation theorem for independent subspace analysis and its consequences
Zoltán Szabó 0001, Barnabás Póczos, András Lörincz |
Pattern Recognit. | 1 |
| 2011 | Online group-structured dictionary learningabstractWe develop a dictionary learning method which is (i) online, (ii) enables overlapping group structures with (iii) non-convex sparsity-inducing regularization and (iv) handles the partially observable case. Structured sparsity and the related group norms have recently gained widespread attention in group-sparsity regularized problems in the case when the dictionary is assumed to be known and fixed. However, when the dictionary also needs to be learned, the problem is much more difficult. Only a few methods have been proposed to solve this problem, and they can handle two of these four desirable properties at most. To the best of our knowledge, our proposed method is the first one that possesses all of these properties. We investigate several interesting special cases of our framework, such as the online, structured, sparse non-negative matrix factorization, and demonstrate the efficiency of our algorithm with several numerical experiments. Zoltán Szabó 0001, Barnabás Póczos, András Lörincz |
CVPR | 1 |
| 2010 | Autoregressive independent process analysis with missing observations
Zoltán Szabó 0001 |
ESANN | 1 |
| 2010 | Auto-regressive independent process analysis without combinatorial efforts
Zoltán Szabó 0001, Barnabás Póczos, András Lörincz |
Pattern Anal. Appl. | 1 |
| 2009 | Controlled Complete ARMA Independent Process AnalysisabstractIn this paper we address the controlled complete AutoRegressive Moving Average Independent Process Analysis (ARMAX-IPA; X-exogenous input or control) problem, which is a generalization of the Blind SubSpace Deconvolution (BSSD) task. Compared to our previous work that dealt with the undercomplete situation, (i) here we extend the theory to complete systems, (ii) allow an autoregressive part to be present, (iii) and include exogenous control. We investigate the case when the observed signal is a linear mixture of independent multidimensional ARMA processes that can be controlled. Our objective is to estimate the ARMA processes, their driving noises as well as the mixing. We aim efficient estimation by choosing suitable control values. For the optimal choice of the control we adapt the D-optimality principle, also known as the ‘InfoMax method’. We solve the problem by reducing it to a fully observable D-optimal ARX task and Independent Subspace Analysis (ISA) that we can solve. Numerical examples illustrate the efficiency of the proposed method. Zoltán Szabó 0001, András Lörincz |
IJCNN | 1 |
| 2007 | Undercomplete Blind Subspace Deconvolution Via Linear Prediction
Zoltán Szabó 0001, Barnabás Póczos, András Lörincz |
ECML | 1 |
| 2007 | Post Nonlinear Independent Subspace Analysis
Zoltán Szabó 0001, Barnabás Póczos, Gábor Szirtes, András Lörincz |
ICANN (1) | 1 |
| 2007 | Neurally plausible, non-combinatorial iterative independent process analysis
András Lörincz, Zoltán Szabó 0001 |
Neurocomputing | 2 |
| 2007 | Undercomplete Blind Subspace Deconvolution
Zoltán Szabó 0001, Barnabás Póczos, András Lörincz |
J. Mach. Learn. Res. | 1 |
| 2004 | Hidden Markov model finds behavioral patterns of users working with a headmouse driven writing toolabstractWe studied user behaviors when the cursor is directed by a head in a simple control task. We used an intelligent writing tool called Dasher. Hidden Markov models (HMMs) were applied to separate behavioral patterns. We found that similar interpretations can be given to the hidden states upon learning. It is argued that the recognition of such general application specific behavioral patterns should be of help for adaptive human-computer interfaces. György Hévízi, Mihály Biczó, Barnabás Póczos, Zoltán Szabó 0001, Bálint Takács, András Lörincz |
IJCNN | 4 |