Aapo Hyvärinen

dblp:56/3623 · DBLP profile ↗
← Back
131ranked-venue papers
46as first author
13since 2021 · last 2025
0000-0002-5806-4432ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 124 · 41 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Density Ratio Estimation with Conditional Probability Paths
abstract
Density ratio estimation in high dimensions can be reframed as integrating a certain quantity, the time score, over probability paths which interpolate between the two densities. In practice, the time score has to be estimated based on samples from the two densities. However, existing methods for this problem remain computationally expensive and can yield inaccurate estimates. Inspired by recent advances in generative modeling, we introduce a novel framework for time score estimation, based on a conditioning variable. Choosing the conditioning variable judiciously enables a closed-form objective function. We demonstrate that, compared to previous approaches, our approach results in faster learning of the time score and competitive or better estimation accuracies of the density ratio on challenging tasks. Furthermore, we establish theoretical guarantees on the error of the estimated density ratio.
Hanlin Yu, Arto Klami, Aapo Hyvärinen, Anna Korba, Omar Chehab
ICML3
2024 Identifiable Feature Learning for Spatial Data with Nonlinear ICA
abstract
Recently, nonlinear ICA has surfaced as a popular alternative to the many heuristic models used in deep representation learning and disentanglement. An advantage of nonlinear ICA is that a sophisticated identifiability theory has been developed; in particular, it has been proven that the original components can be recovered under sufficiently strong latent dependencies. Despite this general theory, practical nonlinear ICA algorithms have so far been mainly limited to data with one-dimensional latent dependencies, especially time-series data. In this paper, we introduce a new nonlinear ICA framework that employs $t$-process (TP) latent components which apply naturally to data with higher-dimensional dependency structures, such as spatial and spatio-temporal data. In particular, we develop a new learning and inference algorithm that extends variational inference methods to handle the combination of a deep neural network mixing function with the TP prior, and employs the method of inducing points for computational efficacy. On the theoretical side, we show that such TP independent components are identifiable under very general conditions. Further, Gaussian Process (GP) nonlinear ICA is established as a limit of the TP Nonlinear ICA model, and we prove that the identifiability of the latent components at this GP limit is more restricted. Namely, those components are identifiable if and only if they have distinctly different covariance kernels. Our algorithm and identifiability theorems are explored on simulated spatial data and real world spatio-temporal data.
Hermanni Hälvä, Jonathan So, Richard E. Turner, Aapo Hyvärinen
AISTATS4
2024 Causal Representation Learning Made Identifiable by Grouping of Observational Variables
abstract
A topic of great current interest is Causal Representation Learning (CRL), whose goal is to learn a causal model for hidden features in a data-driven manner. Unfortunately, CRL is severely ill-posed since it is a combination of the two notoriously ill-posed problems of representation learning and causal discovery. Yet, finding practical identifiability conditions that guarantee a unique solution is crucial for its practical applicability. Most approaches so far have been based on assumptions on the latent causal mechanisms, such as temporal causality, or existence of supervision or interventions; these can be too restrictive in actual applications. Here, we show identifiability based on novel, weak constraints, which requires no temporal structure, intervention, nor weak supervision. The approach is based on assuming the observational mixing exhibits a suitable grouping of the observational variables. We also propose a novel self-supervised estimation framework consistent with the model, prove its statistical consistency, and experimentally show its superior CRL performances compared to the state-of-the-art baselines. We further demonstrate its robustness against latent confounders and causal cycles.
Hiroshi Morioka, Aapo Hyvärinen
ICML2
2023 Connectivity-contrastive learning: Combining causal discovery and representation learning for multimodal data
abstract
Causal discovery methods typically extract causal relations between multiple nodes (variables) based on univariate observations of each node. However, one frequently encounters situations where each node is multivariate, i.e. has multiple observational modalities. Furthermore, the observed modalities may be generated through an unknown mixing process, so that some original latent variables are entangled inside the nodes. In such a multimodal case, the existing frameworks cannot be applied. To analyze such data, we propose a new causal representation learning framework called connectivity-contrastive learning (CCL). CCL disentangles the observational mixing and extracts a set of mutually independent latent components, each having a separate causal structure between the nodes. The actual learning proceeds by a novel self-supervised learning method in which the pretext task is to predict the label of a pair of nodes from the observations of the node pairs. We present theorems which show that CCL can indeed identify both the latent components and the multimodal causal structure under weak technical assumptions, up to some indeterminacy. Finally, we experimentally show its superior causal discovery performance compared to state-of-the-art baselines, in particular demonstrating robustness against latent confounders.
Hiroshi Morioka, Aapo Hyvärinen
AISTATS2
2023 Provable benefits of annealing for estimating normalizing constants: Importance Sampling, Noise-Contrastive Estimation, and beyond
abstract
Recent research has developed several Monte Carlo methods for estimating the normalization constant (partition function) based on the idea of annealing. This means sampling successively from a path of distributions which interpolate between a tractable "proposal" distribution and the unnormalized "target" distribution. Prominent estimators in this family include annealed importance sampling and annealed noise-contrastive estimation (NCE). Such methods hinge on a number of design choices: which estimator to use, which path of distributions to use and whether to use a path at all; so far, there is no definitive theory on which choices are efficient. Here, we evaluate each design choice by the asymptotic estimation error it produces. First, we show that using NCE is more efficient than the importance sampling estimator, but in the limit of infinitesimal path steps, the difference vanishes. Second, we find that using the geometric path brings down the estimation error from an exponential to a polynomial function of the parameter distance between the target and proposal distributions. Third, we find that the arithmetic path, while rarely used, can offer optimality properties over the universally-used geometric path. In fact, in a particular limit, the optimal path is arithmetic. Based on this theory, we finally propose a two-step estimator to approximate the optimal path in an efficient way.
Omar Chehab, Aapo Hyvärinen, Andrej Risteski
NeurIPS2
2022 The optimal noise in noise-contrastive learning is not what you think
abstract
Learning a parametric model of a data distribution is a well-known statistical problem that has seen renewed interest as it is brought to scale in deep learning. Framing the problem as a self-supervised task, where data samples are discriminated from noise samples, is at the core of state-of-the-art methods, beginning with Noise-Contrastive Estimation (NCE). Yet, such contrastive learning requires a good noise distribution, which is hard to specify; domain-specific heuristics are therefore widely used. While a comprehensive theory is missing, it is widely assumed that the optimal noise should in practice be made equal to the data, both in distribution and proportion. This setting underlies Generative Adversarial Networks (GANs) in particular. Here, we empirically and theoretically challenge this assumption on the optimal noise. We show that deviating from this assumption can actually lead to better statistical estimators, in terms of asymptotic variance. In particular, the optimal noise distribution is different from the data’s and even from a different family.
Omar Chehab, Alexandre Gramfort, Aapo Hyvärinen
UAI3
2022 Binary independent component analysis: a non-stationarity-based approach
abstract
We consider independent component analysis of binary data. While fundamental in practice, this case has been much less developed than ICA for continuous data. We start by assuming a linear mixing model in a continuous-valued latent space, followed by a binary observation model. Importantly, we assume that the sources are non-stationary; this is necessary since any non-Gaussianity would essentially be destroyed by the binarization. Interestingly, the model allows for closed-form likelihood by employing the cumulative distribution function of the multivariate Gaussian distribution. In stark contrast to the continuous-valued case, we prove non-identifiability of the model with few observed variables; our empirical results imply identifiability when the number of observed variables is higher. We present a practical method for binary ICA that uses only pairwise marginals, which are faster to compute than the full multivariate likelihood. Experiments give insight into the requirements for the number of observed variables, segments, and latent sources that allow the model to be estimated.
Antti Hyttinen, Vitória Barin Pacela, Aapo Hyvärinen
UAI3
2021 Causal Autoregressive Flows
abstract
Two apparently unrelated fields — normalizing flows and causality — have recently received considerable attention in the machine learning community. In this work, we highlight an intrinsic correspondence between a simple family of autoregressive normalizing flows and identifiable causal models. We exploit the fact that autoregressive flow architectures define an ordering over variables, analogous to a causal ordering, to show that they are well-suited to performing a range of causal inference tasks, ranging from causal discovery to making interventional and counterfactual predictions. First, we show that causal models derived from both affine and additive autoregressive flows with fixed orderings over variables are identifiable, i.e. the true direction of causal influence can be recovered. This provides a generalization of the additive noise model well-known in causal discovery. Second, we derive a bivariate measure of causal direction based on likelihood ratios, leveraging the fact that flow models can estimate normalized log-densities of data. Third, we demonstrate that flows naturally allow for direct evaluation of both interventional and counterfactual queries, the latter case being possible due to the invertible nature of flows. Finally, throughout a series of experiments on synthetic and real data, the proposed method is shown to outperform current approaches for causal discovery as well as making accurate interventional and counterfactual predictions.
Ilyes Khemakhem, Ricardo Pio Monti, Robert Leech, Aapo Hyvärinen
AISTATS4
2021 Independent Innovation Analysis for Nonlinear Vector Autoregressive Process
abstract
The nonlinear vector autoregressive (NVAR) model provides an appealing framework to analyze multivariate time series obtained from a nonlinear dynamical system. However, the innovation (or error), which plays a key role by driving the dynamics, is almost always assumed to be additive. Additivity greatly limits the generality of the model, hindering analysis of general NVAR processes which have nonlinear interactions between the innovations. Here, we propose a new general framework called independent innovation analysis (IIA), which estimates the innovations from completely general NVAR. We assume mutual independence of the innovations as well as their modulation by an auxiliary variable (which is often taken as the time index and simply interpreted as nonstationarity). We show that IIA guarantees the identifiability of the innovations with arbitrary nonlinearities, up to a permutation and component-wise invertible nonlinearities. We also propose three estimation frameworks depending on the type of the auxiliary variable. We thus provide the first rigorous identifiability result for general NVAR, as well as very general tools for learning such models.
Hiroshi Morioka, Hermanni Hälvä, Aapo Hyvärinen
AISTATS3
2021 Disentangling Identifiable Features from Noisy Data with Structured Nonlinear ICA
abstract
We introduce a new general identifiable framework for principled disentanglement referred to as Structured Nonlinear Independent Component Analysis (SNICA). Our contribution is to extend the identifiability theory of deep generative models for a very broad class of structured models. While previous works have shown identifiability for specific classes of time-series models, our theorems extend this to more general temporal structures as well as to models with more complex structures such as spatial dependencies. In particular, we establish the major result that identifiability for this framework holds even in the presence of noise of unknown distribution. Finally, as an example of our framework's flexibility, we introduce the first nonlinear ICA model for time-series that combines the following very useful properties: it accounts for both nonstationarity and autocorrelation in a fully unsupervised setting; performs dimensionality reduction; models hidden states; and enables principled estimation and inference by variational maximum-likelihood.
Hermanni Hälvä, Sylvain Le Corff, Luc Lehéricy, Jonathan So, Yongjie Zhu, Elisabeth Gassiat, Aapo Hyvärinen
NeurIPS7
2021 Shared Independent Component Analysis for Multi-Subject Neuroimaging
abstract
We consider shared response modeling, a multi-view learning problem where one wants to identify common components from multiple datasets or views. We introduce Shared Independent Component Analysis (ShICA) that models eachview as a linear transform of shared independent components contaminated by additive Gaussian noise. We show that this model is identifiable if the components are either non-Gaussian or have enough diversity in noise variances. We then show that in some cases multi-set canonical correlation analysis can recover the correct unmixing matrices, but that even a small amount of sampling noise makes Multiset CCA fail. To solve this problem, we propose to use joint diagonalization after Multiset CCA, leading to a new approach called ShICA-J. We show via simulations that ShICA-J leads to improved results while being very fast to fit. While ShICA-J is based on second-order statistics, we further propose to leverage non-Gaussianity of the components using a maximum-likelihood method, ShICA-ML, that is both more accurate and more costly. Further, ShICA comes with a principled method for shared components estimation. Finally, we provide empirical evidence on fMRI and MEG datasets that ShICA yields more accurate estimation of the componentsthan alternatives.
Hugo Richard, Pierre Ablin, Bertrand Thirion, Alexandre Gramfort, Aapo Hyvärinen
NeurIPS5
2021 Information criteria for non-normalized models
abstract
Many statistical models are given in the form of non-normalized densities with an intractable normalization constant. Since maximum likelihood estimation is computationally intensive for these models, several estimation methods have been developed which do not require explicit computation of the normalization constant, such as noise contrastive estimation (NCE) and score matching. However, model selection methods for general nonnormalized models have not been proposed so far. In this study, we develop information criteria for non-normalized models estimated by NCE or score matching. They are approximately unbiased estimators of discrepancy measures for non-normalized models. Simulation results and applications to real data demonstrate that the proposed criteria enable selection of the appropriate non-normalized model in a data-driven manner.
Takeru Matsuda, Masatoshi Uehara, Aapo Hyvärinen
J. Mach. Learn. Res.3
2021 Direction Matters: On Influence-Preserving Graph Summarization and Max-Cut Principle for Directed Graphs
abstract
Summarizing large-scale directed graphs into small-scale representations is a useful but less-studied problem setting. Conventional clustering approaches, based on Min-Cut-style criteria, compress both the vertices and edges of the graph into the communities, which lead to a loss of directed edge information. On the other hand, compressing the vertices while preserving the directed-edge information provides a way to learn the small-scale representation of a directed graph. The reconstruction error, which measures the edge information preserved by the summarized graph, can be used to learn such representation. Compared to the original graphs, the summarized graphs are easier to analyze and are capable of extracting group-level features, useful for efficient interventions of population behavior. In this letter, we present a model, based on minimizing reconstruction error with nonnegative constraints, which relates to a Max-Cut criterion that simultaneously identifies the compressed nodes and the directed compressed relations between these nodes. A multiplicative update algorithm with column-wise normalization is proposed. We further provide theoretical results on the identifiability of the model and the convergence of the proposed algorithms. Experiments are conducted to demonstrate the accuracy and robustness of the proposed method.
Gang Niu 0001, Aapo Hyvärinen, Masashi Sugiyama
Neural Comput.3
2020 Variational Autoencoders and Nonlinear ICA: A Unifying Framework
abstract
The framework of variational autoencoders allows us to efficiently learn deep latent-variable models, such that the model’s marginal distribution over observed variables fits the data. Often, we’re interested in going a step further, and want to approximate the true joint distribution over observed and latent variables, including the true prior and posterior distributions over latent variables. This is known to be generally impossible due to unidentifiability of the model. We address this issue by showing that for a broad family of deep latent-variable models, identification of the true joint distribution over observed and latent variables is actually possible up to very simple transformations, thus achieving a principled and powerful form of disentanglement. Our result requires a factorized prior distribution over the latent variables that is conditioned on an additionally observed variable, such as a class label or almost any other observation. We build on recent developments in nonlinear ICA, which we extend to the case with noisy, undercomplete or discrete observations, integrated in a maximum likelihood framework. The result also trivially contains identifiable flow-based generative models as a special case.
Ilyes Khemakhem, Diederik P. Kingma, Ricardo Pio Monti, Aapo Hyvärinen
AISTATS4
2020 Relative gradient optimization of the Jacobian term in unsupervised deep learning
abstract
Learning expressive probabilistic models correctly describing the data is a ubiquitous problem in machine learning. A popular approach for solving it is mapping the observations into a representation space with a simple joint distribution, which can typically be written as a product of its marginals — thus drawing a connection with the field of nonlinear independent component analysis. Deep density models have been widely used for this task, but their maximum likelihood based training requires estimating the log-determinant of the Jacobian and is computationally expensive, thus imposing a trade-off between computation and expressive power. In this work, we propose a new approach for exact training of such neural networks. Based on relative gradients, we exploit the matrix structure of neural network parameters to compute updates efficiently even in high-dimensional spaces; the computational cost of the training is quadratic in the input size, in contrast with the cubic scaling of naive approaches. This allows fast training with objective functions involving the log-determinant of the Jacobian, without imposing constraints on its structure, in stark contrast to autoregressive normalizing flows.
Luigi Gresele, Giancarlo Fissore, Adrián Javaloy, Bernhard Schölkopf, Aapo Hyvärinen
NeurIPS5
2020 ICE-BeeM: Identifiable Conditional Energy-Based Deep Models Based on Nonlinear ICA
abstract
We consider the identifiability theory of probabilistic models and establish sufficient conditions under which the representations learnt by a very broad family of conditional energy-based models are unique in function space, up to a simple transformation. In our model family, the energy function is the dot-product between two feature extractors, one for the dependent variable, and one for the conditioning variable. We show that under mild conditions, the features are unique up to scaling and permutation. Our results extend recent developments in nonlinear ICA, and in fact, they lead to an important generalization of ICA models. In particular, we show that our model can be used for the estimation of the components in the framework of Independently Modulated Component Analysis (IMCA), a new generalization of nonlinear ICA that relaxes the independence assumption. A thorough empirical study show that representations learnt by our model from real-world image datasets are identifiable, and improve performance in transfer learning and semi-supervised learning tasks.
Ilyes Khemakhem, Ricardo Pio Monti, Diederik P. Kingma, Aapo Hyvärinen
NeurIPS4
2020 Modeling Shared responses in Neuroimaging Studies through MultiView ICA
abstract
Group studies involving large cohorts of subjects are important to draw general conclusions about brain functional organization. However, the aggregation of data coming from multiple subjects is challenging, since it requires accounting for large variability in anatomy, functional topography and stimulus response across individuals. Data modeling is especially hard for ecologically relevant conditions such as movie watching, where the experimental setup does not imply well-defined cognitive operations. We propose a novel MultiView Independent Component Analysis (ICA) model for group studies, where data from each subject are modeled as a linear combination of shared independent sources plus noise. Contrary to most group-ICA procedures, the likelihood of the model is available in closed form. We develop an alternate quasi-Newton method for maximizing the likelihood, which is robust and converges quickly. We demonstrate the usefulness of our approach first on fMRI data, where our model demonstrates improved sensitivity in identifying common sources among subjects. Moreover, the sources recovered by our model exhibit lower between-sessions variability than other methods. On magnetoencephalography (MEG) data, our method yields more accurate source localization on phantom data. Applied on 200 subjects from the Cam-CAN dataset, it reveals a clear sequence of evoked activity in sensor and source space.
Hugo Richard, Luigi Gresele, Aapo Hyvärinen, Bertrand Thirion, Alexandre Gramfort, Pierre Ablin
NeurIPS3
2020 Hidden Markov Nonlinear ICA: Unsupervised Learning from Nonstationary Time Series
abstract
Recent advances in nonlinear Independent Component Analysis (ICA) provide a principled framework for unsupervised feature learning and disentanglement. The central idea in such works is that the latent components are assumed to be independent conditional on some observed auxiliary variables, such as the time-segment index. This requires manual segmentation of data into non-stationary segments which is computationally expensive, inaccurate and often impossible. These models are thus not fully unsupervised. We remedy these limitations by combining nonlinear ICA with a Hidden Markov Model, resulting in a model where a latent state acts in place of the observed segment-index. We prove identifiability of the proposed model for a general mixing nonlinearity, such as a neural network. We also show how maximum likelihood estimation of the model can be done using the expectation-maximization algorithm. Thus, we achieve a new nonlinear ICA framework which is unsupervised, more efficient, as well as able to model underlying temporal dynamics.
Hermanni Hälvä, Aapo Hyvärinen
UAI2
2020 Robust contrastive learning and nonlinear ICA in the presence of outliers
abstract
Nonlinear independent component analysis (ICA) is a general framework for unsupervised representation learning, and aimed at recovering the latent variables in data. Recent practical methods perform nonlinear ICA by solving classification problems based on logistic regression. However, it is well-known that logistic regression is vulnerable to outliers, and thus the performance can be strongly weakened by outliers. In this paper, we first theoretically analyze nonlinear ICA models in the presence of outliers. Our analysis implies that estimation in nonlinear ICA can be seriously hampered when outliers exist on the tails of the (noncontaminated) target density, which happens in a typical case of contamination by outliers. We develop two robust nonlinear ICA methods based on the $\gamma$-divergence, which is a robust alternative to the KL-divergence in logistic regression. The proposed methods are theoretically shown to have desired robustness properties in the context of nonlinear ICA. We also experimentally demonstrate that the proposed methods are very robust and outperform existing methods in the presence of outliers. Finally, the proposed method is applied to ICA-based causal discovery and shown to find a plausible causal relationship on fMRI data.
Hiroaki Sasaki, Takashi Takenouchi, Ricardo Pio Monti, Aapo Hyvärinen
UAI4
2019 Nonlinear ICA Using Auxiliary Variables and Generalized Contrastive Learning
abstract
Nonlinear ICA is a fundamental problem for unsupervised representation learning, emphasizing the capacity to recover the underlying latent variables generating the data (i.e., identifiability). Recently, the very first identifiability proofs for nonlinear ICA have been proposed, leveraging the temporal structure of the independent components. Here, we propose a general framework for nonlinear ICA, which, as a special case, can make use of temporal structure. It is based on augmenting the data by an auxiliary variable, such as the time index, the history of the time series, or any other available information. We propose to learn nonlinear ICA by discriminating between true augmented data, or data in which the auxiliary variable has been randomized. This enables the framework to be implemented algorithmically through logistic regression, possibly in a neural network. We provide a comprehensive proof of the identifiability of the model as well as the consistency of our estimation method. The approach not only provides a general theoretical framework combining and generalizing previously proposed nonlinear ICA models and algorithms, but also brings practical advantages.
Aapo Hyvärinen, Hiroaki Sasaki, Richard E. Turner
AISTATS1
2019 Estimation of Non-Normalized Mixture Models
abstract
We develop a general method for estimating a finite mixture of non-normalized models. A non-normalized model is defined to be a parametric distribution with an intractable normalization constant. Existing methods for estimating non-normalized models without computing the normalization constant are not applicable to mixture models because they contain more than one intractable normalization constant. The proposed method is derived by extending noise contrastive estimation (NCE), which estimates non-normalized models by discriminating between the observed data and some artificially generated noise. In particular, the proposed method provides a probabilistically principled clustering method that is able to utilize a deep representation. Applications to clustering of natural images and neuroimaging data give promising results.
Takeru Matsuda, Aapo Hyvärinen
AISTATS2
2019 Causal Discovery with General Non-Linear Relationships using Non-Linear ICA
Ricardo Pio Monti, Kun Zhang 0001, Aapo Hyvärinen
UAI3
2019 Neural Empirical Bayes
abstract
We unify kernel density estimation and empirical Bayes and address a set of problems in unsupervised machine learning with a geometric interpretation of those methods, rooted in the concentration of measure phenomenon. Kernel density is viewed symbolically as $X\rightharpoonup Y$ where the random variable $X$ is smoothed to $Y= X+N(0,\sigma^2 I_d)$, and empirical Bayes is the machinery to denoise in a least-squares sense, which we express as $X \leftharpoondown Y$. A learning objective is derived by combining these two, symbolically captured by $X \rightleftharpoons Y$. Crucially, instead of using the original nonparametric estimators, we parametrize the energy function with a neural network denoted by $\phi$; at optimality, $\nabla \phi \approx -\nabla \log f$ where $f$ is the density of $Y$. The optimization problem is abstracted as interactions of high-dimensional spheres which emerge due to the concentration of isotropic Gaussians. We introduce two algorithmic frameworks based on this machinery: (i) a “walk-jump” sampling scheme that combines Langevin MCMC (walks) and empirical Bayes (jumps), and (ii) a probabilistic framework for associative memory, called NEBULA, defined a la Hopfield by the gradient flow of the learned energy to a set of attractors. We finish the paper by reporting the emergence of very rich “creative memories” as attractors of NEBULA for highly-overlapping spheres.
Saeed Saremi, Aapo Hyvärinen
J. Mach. Learn. Res.2
2018 A unified probabilistic model for learning latent factors and their connectivities from high-dimensional data
Ricardo Pio Monti, Aapo Hyvärinen
UAI2
2017 Nonlinear ICA of Temporally Dependent Stationary Sources
abstract
We develop a nonlinear generalization of independent component analysis (ICA) or blind source separation, based on temporal dependencies (e.g. autocorrelations). We introduce a nonlinear generative model where the independent sources are assumed to be temporally dependent, non-Gaussian, and stationary, and we observe arbitrarily nonlinear mixtures of them. We develop a method for estimating the model (i.e. separating the sources) based on logistic regression in a neural network which learns to discriminate between a short temporal window of the data vs. a temporal window of temporally permuted data. We prove that the method estimates the sources for general smooth mixing nonlinearities, assuming the sources have sufficiently strong temporal dependencies, and these dependencies are in a certain way different from dependencies found in Gaussian processes. For Gaussian (and similar) sources, the method estimates the nonlinear part of the mixing. We thus provide the first rigorous and general proof of identifiability of nonlinear ICA for temporally dependent sources, together with a practical method for its estimation.
Aapo Hyvärinen, Hiroshi Morioka
AISTATS1
2017 SPLICE: Fully Tractable Hierarchical Extension of ICA with Pooling
abstract
We present a novel probabilistic framework for a hierarchical extension of independent component analysis (ICA), with a particular motivation in neuroscientific data analysis and modeling. The framework incorporates a general subspace pooling with linear ICA-like layers stacked recursively. Unlike related previous models, our generative model is fully tractable: both the likelihood and the posterior estimates of latent variables can readily be computed with analytically simple formulae. The model is particularly simple in the case of complex-valued data since the pooling can be reduced to taking the modulus of complex numbers. Experiments on electroencephalography (EEG) and natural images demonstrate the validity of the method.
Aapo Hyvärinen, Motoaki Kawanabe
ICML2
2017 Mode-Seeking Clustering and Density Ridge Estimation via Direct Estimation of Density-Derivative-Ratios
Hiroaki Sasaki, Takafumi Kanamori, Aapo Hyvärinen, Gang Niu 0001, Masashi Sugiyama
J. Mach. Learn. Res.3
2017 Density Estimation in Infinite Dimensional Exponential Families
abstract
In this paper, we consider an infinite dimensional exponential family $\mathcal{P}$ of probability densities, which are parametrized by functions in a reproducing kernel Hilbert space $\mathcal{H}$, and show it to be quite rich in the sense that a broad class of densities on $\mathbb{R}^d$ can be approximated arbitrarily well in Kullback-Leibler (KL) divergence by elements in $\mathcal{P}$. Motivated by this approximation property, the paper addresses the question of estimating an unknown density $p_0$ through an element in $\mathcal{P}$. Standard techniques like maximum likelihood estimation (MLE) or pseudo MLE (based on the method of sieves), which are based on minimizing the KL divergence between $p_0$ and $\mathcal{P}$, do not yield practically useful estimators because of their inability to efficiently handle the log-partition function. We propose an estimator $\hat{p}_n$ based on minimizing the Fisher divergence, $J(p_0\Vert p)$ between $p_0$ and $p\in \mathcal{P}$, which involves solving a simple finite-dimensional linear system. When $p_0\in\mathcal{P}$, we show that the proposed estimator is consistent, and provide a convergence rate of $n^{-\min\left\{\frac{2}{3},\frac{2\beta+1}{2\beta+2}\right\}}$ in Fisher divergence under the smoothness assumption that $\log p_0\in\mathcal{R}(C^\beta)$ for some $\beta\ge 0$, where $C$ is a certain Hilbert-Schmidt operator on $\mathcal{H}$ and $\mathcal{R}(C^\beta)$ denotes the image of $C^\beta$. We also investigate the misspecified case of $p_0\notin\mathcal{P}$ and show that $J(p_0\Vert\hat{p}_n)\rightarrow \inf_{p\in\mathcal{P}}J(p_0\Vert p)$ as $n\rightarrow \infty$, and provide a rate for this convergence under a similar smoothness condition as above. Through numerical simulations we demonstrate that the proposed estimator outperforms the non- parametric kernel density estimator, and that the advantage of the proposed estimator grows as $d$ increases.
Bharath K. Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Aapo Hyvärinen, Revant Kumar
J. Mach. Learn. Res.4
2017 Simultaneous Estimation of Nongaussian Components and Their Correlation Structure
Hiroaki Sasaki, Michael Gutmann, Hayaru Shouno, Aapo Hyvärinen
Neural Comput.4
2017 A mixture of sparse coding models explaining properties of face neurons related to holistic and parts-based processing
abstract
Experimental studies have revealed evidence of both parts-based and holistic representations of objects and faces in the primate visual system. However, it is still a mystery how such seemingly contradictory types of processing can coexist within a single system. Here, we propose a novel theory called mixture of sparse coding models, inspired by the formation of category-specific subregions in the inferotemporal (IT) cortex. We developed a hierarchical network that constructed a mixture of two sparse coding submodels on top of a simple Gabor analysis. The submodels were each trained with face or non-face object images, which resulted in separate representations of facial parts and object parts. Importantly, evoked neural activities were modeled by Bayesian inference, which had a top-down explaining-away effect that enabled recognition of an individual part to depend strongly on the category of the whole input. We show that this explaining-away effect was indeed crucial for the units in the face submodel to exhibit significant selectivity to face images over object images in a similar way to actual face-selective neurons in the macaque IT cortex. Furthermore, the model explained, qualitatively and quantitatively, several tuning properties to facial features found in the middle patch of face processing in IT as documented by Freiwald, Tsao, and Livingstone (2009). These included, in particular, tuning to only a small number of facial features that were often related to geometrically large parts like face outline and hair, preference and anti-preference of extreme facial features (e.g., very large/small inter-eye distance), and reduction of the gain of feature tuning for partial face stimuli compared to whole face stimuli. Thus, we hypothesize that the coding principle of facial features in the middle patch of face processing in the macaque IT cortex may be closely related to mixture of sparse coding models.
Haruo Hosoya, Aapo Hyvärinen
PLoS Comput. Biol.2
2016 Unsupervised Feature Extraction by Time-Contrastive Learning and Nonlinear ICA
abstract
Nonlinear independent component analysis (ICA) provides an appealing framework for unsupervised feature learning, but the models proposed so far are not identifiable. Here, we first propose a new intuitive principle of unsupervised deep learning from time series which uses the nonstationary structure of the data. Our learning principle, time-contrastive learning (TCL), finds a representation which allows optimal discrimination of time segments (windows). Surprisingly, we show how TCL can be related to a nonlinear ICA model, when ICA is redefined to include temporal nonstationarities. In particular, we show that TCL combined with linear ICA estimates the nonlinear ICA model up to point-wise transformations of the sources, and this solution is unique --- thus providing the first identifiability result for nonlinear ICA which is rigorous, constructive, as well as very general.
Aapo Hyvärinen, Hiroshi Morioka
NIPS1
2016 Sparse and low-rank matrix regularization for learning time-varying Markov networks
Aapo Hyvärinen, Shin Ishii
Mach. Learn.2
2016 Learning Visual Spatial Pooling by Strong PCA Dimension Reduction
abstract
In visual modeling, invariance properties of visual cells are often explained by a pooling mechanism, in which outputs of neurons with similar selectivities to some stimulus parameters are integrated so as to gain some extent of invariance to other parameters. For example, the classical energy model of phase-invariant V1 complex cells pools model simple cells preferring similar orientation but different phases. Prior studies, such as independent subspace analysis, have shown that phase-invariance properties of V1 complex cells can be learned from spatial statistics of natural inputs. However, those previous approaches assumed a squaring nonlinearity on the neural outputs to capture energy correlation; such nonlinearity is arguably unnatural from a neurobiological viewpoint but hard to change due to its tight integration into their formalisms. Moreover, they used somewhat complicated objective functions requiring expensive computations for optimization. In this study, we show that visual spatial pooling can be learned in a much simpler way using strong dimension reduction based on principal component analysis. This approach learns to ignore a large part of detailed spatial structure of the input and thereby estimates a linear pooling matrix. Using this framework, we demonstrate that pooling of model V1 simple cells learned in this way, even with nonlinearities other than squaring, can reproduce standard tuning properties of V1 complex cells. For further understanding, we analyze several variants of the pooling model and argue that a reasonable pooling can generally be obtained from any kind of linear transformation that retains several of the first principal components and suppresses the remaining ones. In particular, we show how the classic Wiener filtering theory leads to one such variant.
Haruo Hosoya, Aapo Hyvärinen
Neural Comput.2
2016 Orthogonal Connectivity Factorization: Interpretable Decomposition of Variability in Correlation Matrices
abstract
In many multivariate time series, the correlation structure is nonstationary, that is, it changes over time. The correlation structure may also change as a function of other cofactors, for example, the identity of the subject in biomedical data. A fundamental approach for the analysis of such data is to estimate the correlation structure (connectivities) separately in short time windows or for different subjects and use existing machine learning methods, such as principal component analysis (PCA), to summarize or visualize the changes in connectivity. However, the visualization of such a straightforward PCA is problematic because the ensuing connectivity patterns are much more complex objects than, say, spatial patterns. Here, we develop a new framework for analyzing variability in connectivities using the PCA approach as the starting point. First, we show how to analyze and visualize the principal components of connectivity matrices by a tailor-made rank-two matrix approximation in which we use the outer product of two orthogonal vectors. This leads to a new kind of transformation of eigenvectors that is particularly suited for this purpose and often enables interpretation of the principal component as connectivity between two groups of variables. Second, we show how to incorporate the orthogonality and the rank-two constraint in the estimation of PCA itself to improve the results. We further provide an interpretation of these methods in terms of estimation of a probabilistic generative model related to blind separation of dependent sources. Experiments on brain imaging data give very promising results.
Aapo Hyvärinen, Vesa Kiviniemi, Motoaki Kawanabe
Neural Comput.1
2015 Independent component analysis with an inverse problem motivated penalty term
abstract
We describe a model where an independent component problem and a related linear inverse problem are modelled simultaneously, and construct an algorithm which in some circumstances produces demixing matrices of better quality than the basic ICA algorithms. The effect is achieved by adding a penalty term, motivated by the inverse problem, to the ICA objective function. Our method is related to the idea, which has received some attention in the brain imaging context, that solutions of independent component problems can be used as a basis for inverse methods.
Jouni Puuronen, Aapo Hyvärinen
IJCNN2
2015 Unifying Blind Separation and Clustering for Resting-State EEG/MEG Functional Connectivity Analysis
abstract
Unsupervised analysis of the dynamics (nonstationarity) of functional brain connectivity during rest has recently received a lot of attention in the neuroimaging and neuroengineering communities. Most studies have used functional magnetic resonance imaging, but electroencephalography (EEG) and magnetoencephalography (MEG) also hold great promise for analyzing nonstationary functional connectivity with high temporal resolution. Previous EEG/MEG analyses divided the problem into two consecutive stages: the separation of neural sources and then the connectivity analysis of the separated sources. Such nonoptimal division into two stages may bias the result because of the different prior assumptions made about the data in the two stages. We propose a unified method for separating EEG/MEG sources and learning their functional connectivity (coactivation) patterns. We combine blind source separation (BSS) with unsupervised clustering of the activity levels of the sources in a single probabilistic model. A BSS is performed on the Hilbert transforms of band-limited EEG/MEG signals, and coactivation patterns are learned by a mixture model of source envelopes. Simulation studies show that the unified approach often outperforms conventional two-stage methods, indicating further the benefit of using Hilbert transforms to deal with oscillatory sources. Experiments on resting-state EEG data, acquired in conjunction with a cued motor imagery or nonimagery task, also show that the states (clusters) obtained by the proposed method often correlate better with physiologically meaningful quantities than those obtained by a two-stage method.
Takeshi Ogawa, Aapo Hyvärinen
Neural Comput.3
2014 Estimating Dependency Structures for non-Gaussian Components with Linear and Energy Correlations
abstract
The statistical dependencies which independent component analysis (ICA) cannot remove often provide rich information beyond the ICA components. It would be very useful to estimate the dependency structure from data. However, most models have concentrated on higher-order correlations such as energy correlations, neglecting linear correlations. Linear correlations might be a strong and informative form of a dependency for some real data sets, but they are usually completely removed by ICA and related methods, and not analyzed at all. In this paper, we propose a probabilistic model of non-Gaussian components which are allowed to have both linear and energy correlations. The dependency structure of the components is explicitly parametrized by a parameter matrix, which defines an undirected graphical model over the latent components. Furthermore, the estimation of the parameter matrix is shown to be particularly simple because using score matching, the objective function is a quadratic form. Using artificial data, we demonstrate that the proposed method is able to estimate non-Gaussian components and their dependency structures, as it is designed to do. When applied to natural images and outputs of simulated complex cells in the primary visual cortex, novel dependencies between the estimated features are discovered.
Hiroaki Sasaki, Michael Gutmann, Hayaru Shouno, Aapo Hyvärinen
AISTATS4
2014 Clustering via Mode Seeking by Direct Estimation of the Gradient of a Log-Density
Hiroaki Sasaki, Aapo Hyvärinen, Masashi Sugiyama
ECML/PKDD (3)2
2014 ParceLiNGAM: A Causal Ordering Method Robust Against Latent Confounders
abstract
We consider learning a causal ordering of variables in a linear nongaussian acyclic model called LiNGAM. Several methods have been shown to consistently estimate a causal ordering assuming that all the model assumptions are correct. But the estimation results could be distorted if some assumptions are violated. In this letter, we propose a new algorithm for learning causal orders that is robust against one typical violation of the model assumptions: latent confounders. The key idea is to detect latent confounders by testing independence between estimated external influences and find subsets (parcels) that include variables unaffected by latent confounders. We demonstrate the effectiveness of our method using artificial data and simulated brain imaging data.
Tatsuya Tashiro, Shohei Shimizu, Aapo Hyvärinen, Takashi Washio
Neural Comput.3
2014 A Bayesian inverse solution using independent component analysis
Jouni Puuronen, Aapo Hyvärinen
Neural Networks2
2013 Pairwise likelihood ratios for estimation of non-Gaussian structural equation models
Aapo Hyvärinen, Stephen M. Smith 0001
J. Mach. Learn. Res.1
2013 Correlated topographic analysis: estimating an ordering of correlated components
Hiroaki Sasaki, Michael Gutmann, Hayaru Shouno, Aapo Hyvärinen
Mach. Learn.4
2012 Estimation of Causal Orders in a Linear Non-Gaussian Acyclic Model: A Method Robust against Latent Confounders
Tatsuya Tashiro, Shohei Shimizu, Aapo Hyvärinen, Takashi Washio
ICANN (1)3
2012 Learning a selectivity-invariance-selectivity feature extraction architecture for images
Michael Gutmann, Aapo Hyvärinen
ICPR2
2012 Noise-Contrastive Estimation of Unnormalized Statistical Models, with Applications to Natural Image Statistics
Michael Gutmann, Aapo Hyvärinen
J. Mach. Learn. Res.2
2011 Extracting Coactivated Features from Multiple Data Sets
Michael Gutmann, Aapo Hyvärinen
ICANN (1)2
2011 Complex-Valued Independent Component Analysis of Natural Images
Valero Laparra, Michael Gutmann, Jesús Malo, Aapo Hyvärinen
ICANN (2)4
2011 Hermite Polynomials and Measures of Non-gaussianity
Jouni Puuronen, Aapo Hyvärinen
ICANN (2)2
2011 Structural equations and divisive normalization for energy-dependent component analysis
abstract
Components estimated by independent component analysis and related methods are typically not independent in real data. A very common form of nonlinear dependency between the components is correlations in their variances or ener- gies. Here, we propose a principled probabilistic model to model the energy- correlations between the latent variables. Our two-stage model includes a linear mixing of latent signals into the observed ones like in ICA. The main new fea- ture is a model of the energy-correlations based on the structural equation model (SEM), in particular, a Linear Non-Gaussian SEM. The SEM is closely related to divisive normalization which effectively reduces energy correlation. Our new two- stage model enables estimation of both the linear mixing and the interactions re- lated to energy-correlations, without resorting to approximations of the likelihood function or other non-principled approaches. We demonstrate the applicability of our method with synthetic dataset, natural images and brain signals.
Aapo Hyvärinen
NIPS2
2011 DirectLiNGAM: A Direct Method for Learning a Linear Non-Gaussian Structural Equation Model
Shohei Shimizu, Takanori Inazumi, Yasuhiro Sogawa, Aapo Hyvärinen, Yoshinobu Kawahara, Takashi Washio, Patrik O. Hoyer, Kenneth Bollen
J. Mach. Learn. Res.4
2011 Estimating exogenous variables in data with more variables than observations
Yasuhiro Sogawa, Shohei Shimizu, Teppei Shimamura, Aapo Hyvärinen, Takashi Washio, Seiya Imoto
Neural Networks4
2010 Discovery of Exogenous Variables in Data with More Variables Than Observations
Yasuhiro Sogawa, Shohei Shimizu, Aapo Hyvärinen, Takashi Washio, Teppei Shimamura, Seiya Imoto
ICANN (1)3
2010 Sparse and Low-Rank Estimation of Time-Varying Markov Networks with Alternating Direction Method of Multipliers
Aapo Hyvärinen, Shin Ishii
ICONIP (1)2
2010 A Family of Computationally E cient and Simple Estimators for Unnormalized Statistical Models
Miika Pihlaja, Michael Gutmann, Aapo Hyvärinen
UAI3
2010 Source Separation and Higher-Order Causal Analysis of MEG and EEG
Kun Zhang 0001, Aapo Hyvärinen
UAI2
2010 Estimation of a Structural Vector Autoregression Model Using Non-Gaussianity
Aapo Hyvärinen, Kun Zhang 0001, Shohei Shimizu, Patrik O. Hoyer
J. Mach. Learn. Res.1
2010 A Two-Layer Model of Natural Stimuli Estimated with Score Matching
abstract
We consider a hierarchical two-layer model of natural signals in which both layers are learned from the data. Estimation is accomplished by score matching, a recently proposed estimation principle for energy-based models. If the first-layer outputs are squared and the second-layer weights are constrained to be nonnegative, the model learns responses similar to complex cells in primary visual cortex from natural images. The second layer pools a small number of features with similar orientation and frequency, but differing in spatial phase. For speech data, we obtain analogous results. The model unifies previous extensions to independent component analysis such as subspace and topographic models and provides new evidence that localized, oriented, phase-invariant features reflect the statistical properties of natural image patches.
Urs Köster, Aapo Hyvärinen
Neural Comput.2
2010 WordICA - emergence of linguistic representations for words by independent component analysis
abstract
Abstract We explore the use of independent component analysis (ICA) for the automatic extraction of linguistic roles or features of words. The extraction is based on the unsupervised analysis of text corpora. We contrast ICA with singular value decomposition (SVD), widely used in statistical text analysis, in general, and specifically in latent semantic analysis (LSA). However, the representations found using the SVD analysis cannot easily be interpreted by humans. In contrast, ICA applied on word context data gives distinct features which reflect linguistic categories. In this paper, we provide justification for our approach called WordICA, present the WordICA method in detail, compare the obtained results with traditional linguistic categories and with the results achieved using an SVD-based method, and discuss the use of the method in practical natural language engineering solutions such as machine translation systems. As the WordICA method is based on unsupervised learning and thus provides a general means for efficient knowledge acquisition, we foresee that the approach has a clear potential for practical applications.
Timo Honkela, Aapo Hyvärinen, Jaakko J. Väyrynen
Nat. Lang. Eng.2
2009 Learning reconstruction and prediction of natural stimuli by a population of spiking neurons
Michael Gutmann, Aapo Hyvärinen
ESANN2
2009 Learning Features by Contrasting Natural Images with Noise
Michael Gutmann, Aapo Hyvärinen
ICANN (2)2
2009 Modelling Image Complexity by Independent Component Analysis, with Application to Content-Based Image Retrieval
Jukka Perkiö, Aapo Hyvärinen
ICANN (2)2
2009 Causality Discovery with Additive Disturbances: An Information-Theoretical Perspective
Kun Zhang 0001, Aapo Hyvärinen
ECML/PKDD (2)2
2009 A direct method for estimating a causal ordering in a linear non-Gaussian acyclic model
Shohei Shimizu, Aapo Hyvärinen, Yoshinobu Kawahara
UAI2
2009 On the Identifiability of the Post-Nonlinear Causal Model
Kun Zhang 0001, Aapo Hyvärinen
UAI2
2009 Estimation of linear non-Gaussian acyclic models for latent factors
Shohei Shimizu, Patrik O. Hoyer, Aapo Hyvärinen
Neurocomputing3
2008 Causal modelling combining instantaneous and lagged effects: an identifiable model based on non-Gaussianity
abstract
Causal analysis of continuous-valued variables typically uses either autoregressive models or linear Gaussian Bayesian networks with instantaneous effects. Estimation of Gaussian Bayesian networks poses serious identifiability problems, which is why it was recently proposed to use non-Gaussian models. Here, we show how to combine the non-Gaussian instantaneous model with autoregressive models. We show that such a non-Gaussian model is identifiable without prior knowledge of network structure, and we propose an estimation method shown to be consistent. This approach also points out how neglecting instantaneous effects can lead to completely wrong estimates of the autoregressive coefficients.
Aapo Hyvärinen, Shohei Shimizu, Patrik O. Hoyer
ICML1
2008 Learning encoding and decoding filters for data representation with a spiking neuron
abstract
Data representation methods related to ICA and sparse coding have successfully been used to model neural representation. However, they are highly abstract methods, and the neural encoding does not correspond to a detailed neuron model. This limits their power to provide deeper insight into the sensory systems on a cellular level. We propose here data representation where the encoding happens with a spiking neuron. The data representation problem is formulated as an optimization problem: Encode the input so that it can be decoded from the spike train, and optionally, so that energy consumption is minimized. The optimization leads to a learning rule for the encoder and decoder which features synergistic interaction: The decoder provides feedback affecting the plasticity of the encoder while the encoder provides optimal learning data for the decoder.
Michael Gutmann, Aapo Hyvärinen, Kazuyuki Aihara
IJCNN2
2008 On the learning of nonlinear visual features from natural images by optimizing response energies
abstract
The operation of V1 simple cells in primates has been traditionally modelled with linear models resembling Gabor filters, whereas the functionality of subsequent visual cortical areas is less well understood. Here we explore the learning of mechanisms for further nonlinear processing by assuming a functional form of a product of two linear filter responses, and estimating a basis for the given visual data by optimizing for robust alternative of variance of the nonlinear model outputs. By a simple transformation of the learned model, we demonstrate that on natural images, both minimization and maximization in our setting lead to oriented, band-pass and localized linear filters whose responses are then nonlinearly combined. In minimization, the method learns to multiply the responses of two Gabor-like filters, whereas in maximization it learns to subtract the response magnitudes of two Gabor-like filters. Empirically, these learned nonlinear filters appear to function as conjunction detectors and as opponent orientation filters, respectively. We provide a preliminary explanation for our results in terms of filter energy correlations and fourth power optimization.
Jussi T. Lindgren, Aapo Hyvärinen
IJCNN2
2008 Unsupervised learning of dependencies between local luminance and contrast in natural images
abstract
Separate processing of local luminance and contrast in biological visual systems has been argued to be due to the independence of these two properties in natural image data. In this paper we examine spatial, retinotopic channels formed by these two quantities and use Independent Component Analysis to study the possible dependencies between the channels. As a result, oriented, localized bandpass filter pairs are learned, where one filter processes the luminance channel and the other the contrast channel. We study the relationship of the learned filters and their pairings, and show that these are due to dependencies existing between local luminance and contrast. Subsequently, our results suggest that the separate processing of local luminance and contrast can not be attributed to their independence in natural images.
Jussi T. Lindgren, Jarmo Hurri, Aapo Hyvärinen
IJCNN3
2008 Causal discovery of linear acyclic models with arbitrary distributions
Patrik O. Hoyer, Aapo Hyvärinen, Richard Scheines, Peter Spirtes, Joseph D. Ramsey, Gustavo Lacerda, Shohei Shimizu
UAI2
2008 Optimal Approximation of Signal Priors
abstract
In signal restoration by Bayesian inference, one typically uses a parametric model of the prior distribution of the signal. Here, we consider how the parameters of a prior model should be estimated from observations of uncorrupted signals. A lot of recent work has implicitly assumed that maximum likelihood estimation is the optimal estimation method. Our results imply that this is not the case. We first obtain an objective function that approximates the error occurred in signal restoration due to an imperfect prior model. Next, we show that in an important special case (small gaussian noise), the error is the same as the score-matching objective function, which was previously proposed as an alternative for likelihood based on purely computational considerations. Our analysis thus shows that score matching combines computational simplicity with statistical optimality in signal restoration, providing a viable alternative to maximum likelihood methods. We also show how the method leads to a new intuitive and geometric interpretation of structure inherent in probability distributions.
Aapo Hyvärinen
Neural Comput.1
2007 A Two-Layer ICA-Like Model Estimated by Score Matching
Urs Köster, Aapo Hyvärinen
ICANN (2)2
2007 Discovery of Linear Non-Gaussian Acyclic Models in the Presence of Latent Classes
Shohei Shimizu, Aapo Hyvärinen
ICONIP (1)2
2007 Equivalence of Some Common Linear Feature Extraction Techniques for Appearance-Based Object Recognition Tasks
abstract
Recently, a number of empirical studies have compared the performance of PCA and ICA as feature extraction methods in appearance-based object recognition systems, with mixed and seemingly contradictory results. In this paper, we briefly describe the connection between the two methods and argue that whitened PCA may yield identical results to ICA in some cases. Furthermore, we describe the specific situations in which ICA might significantly improve on PCA.
Maria Asunción Vicente, Patrik O. Hoyer, Aapo Hyvärinen
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Connections Between Score Matching, Contrastive Divergence, and Pseudolikelihood for Continuous-Valued Variables
abstract
Score matching (SM) and contrastive divergence (CD) are two recently proposed methods for estimation of nonnormalized statistical methods without computation of the normalization constant (partition function). Although they are based on very different approaches, we show in this letter that they are equivalent in a special case: in the limit of infinitesimal noise in a specific Monte Carlo method. Further, we show how these methods can be interpreted as approximations of pseudolikelihood.
Aapo Hyvärinen
IEEE Trans. Neural Networks1
2006 FastISA: A fast fixed-point algorithm for independent subspace analysis
Aapo Hyvärinen, Urs Köster
ESANN1
2006 A Quasi-stochastic Gradient Algorithm for Variance-Dependent Component Analysis
Aapo Hyvärinen, Shohei Shimizu
ICANN (2)1
2006 Learning to Segment Any Random Vector
abstract
We propose a method that takes observations of a random vector as input, and learns to segment each observation into two disjoint parts. We show how to use the internal coherence of segments to learn to segment almost any random variable. Coherence is formalized using the principle of autoprediction, i.e. two elements are similar if the observed values are similar to the predictions given by the elements for each other. To obtain a principled model and method, we formulate a generative model and show how it can be estimated in the limit of zero noise. The ensuing method is an abstract, adaptive (learning) generalization of well-known methods for image segmentation. It enables segmentation of random vectors in cases where intuitive prior information necessary for conventional segmentation methods is not available.
Aapo Hyvärinen, Jukka Perkiö
IJCNN1
2006 Emergence of conjunctive visual features by quadratic independent component analysis
abstract
In previous studies, quadratic modelling of natural images has resulted in cell models that react strongly to edges and bars. Here we apply quadratic Independent Component Analysis to natural image patches, and show that up to a small approximation error, the estimated components are computing conjunctions of two linear features. These conjunctive features appear to represent not only edges and bars, but also inherently two-dimensional stimuli, such as corners. In addition, we show that for many of the components, the underlying linear features have essentially V1 simple cell receptive field characteristics. Our results indicate that the development of the V2 cells preferring angles and corners may be partly explainable by the principle of unsupervised sparse coding of natural images.
Jussi T. Lindgren, Aapo Hyvärinen
NIPS2
2006 A Linear Non-Gaussian Acyclic Model for Causal Discovery
abstract
In recent years, several methods have been proposed for the discovery of causal structure from non-experimental data. Such methods make various assumptions on the data generating process to facilitate its identification from purely observational data. Continuing this line of research, we show how to discover the complete causal structure of continuous-valued data, under the assumptions that (a) the data generating process is linear, (b) there are no unobserved confounders, and (c) disturbance variables have non-Gaussian distributions of non-zero variances. The solution relies on the use of the statistical method known as independent component analysis, and does not require any pre-specified time-ordering of the variables. We provide a complete Matlab package for performing this LiNGAM analysis (short for Linear Non-Gaussian Acyclic Model), and demonstrate the effectiveness of the method using artificially generated data and real-world data.
Shohei Shimizu, Patrik O. Hoyer, Aapo Hyvärinen, Antti J. Kerminen
J. Mach. Learn. Res.3
2006 Consistency of Pseudolikelihood Estimation of Fully Visible Boltzmann Machines
abstract
A Boltzmann machine is a classic model of neural computation, and a number of methods have been proposed for its estimation. Most methods are plagued by either very slow convergence or asymptotic bias in the resulting estimates. Here we consider estimation in the basic case of fully visible Boltzmann machines. We show that the old principle of pseudolikelihood estimation provides an estimator that is computationally very simple yet statistically consistent.
Aapo Hyvärinen
Neural Comput.1
2005 Discovery of Non-gaussian Linear Causal Models using ICA
Shohei Shimizu, Aapo Hyvärinen, Yutaka Kano, Patrik O. Hoyer
UAI2
2005 Estimation of Non-Normalized Statistical Models by Score Matching
abstract
One often wants to estimate statistical models where the probability density function is known only up to a multiplicative normalization constant. Typically, one then has to resort to Markov Chain Monte Carlo methods, or approximations of the normalization constant. Here, we propose that such models can be estimated by minimizing the expected squared distance between the gradient of the log-density given by the model and the gradient of the log-density of the observed data. While the estimation of the gradient of log-density function is, in principle, a very difficult non-parametric problem, we prove a surprising result that gives a simple formula for this objective function. The density function of the observed data does not appear in this formula, which simplifies to a sample average of a sum of some derivatives of the log-density given by the model. The validity of the method is demonstrated on multivariate Gaussian and independent component analysis models, and by estimating an overcomplete filter set for natural image data.
Aapo Hyvärinen
J. Mach. Learn. Res.1
2005 A unifying model for blind separation of independent sources
Aapo Hyvärinen
Signal Process.1
2004 Linguistic feature extraction using independent component analysis
abstract
Our aim is to find syntactic and semantic relationships of words based on the analysis of corpora. We propose the application of independent component analysis, which seems to have clear advantages over two classic methods: latent semantic analysis and self-organizing maps. Latent semantic analysis is a simple method for automatic generation of concepts that are useful, e.g., in encoding documents for information retrieval purposes. However, these concepts cannot easily be interpreted by humans. Self-organizing maps can be used to generate an explicit diagram which characterizes the relationships between words. The resulting map reflects syntactic categories in the overall organization and semantic categories in the local level. The self-organizing map does not, however, provide any explicit distinct categories for the words. Independent component analysis applied on word context data gives distinct features which reflect syntactic and semantic categories. Thus, independent component analysis gives features or categories that are both explicit and can easily be interpreted by humans. This result can be obtained without any human supervision or tagged corpora that would have some predetermined morphological, syntactic or semantic information.
Timo Honkela, Aapo Hyvärinen
IJCNN2
2004 Spatiotemporal receptive fields maximizing temporal coherence in natural image sequences
Jarmo Hurri, Jaakko J. Väyrynen, Aapo Hyvärinen
Neurocomputing3
2004 A unifying framework for natural image statistics: spatiotemporal activity bubbles
Aapo Hyvärinen, Jarmo Hurri, Jaakko J. Väyrynen
Neurocomputing1
2004 Blind separation of sources that have spatiotemporal variance dependencies
Aapo Hyvärinen, Jarmo Hurri
Signal Process.1
2003 A two-layer temporal generative model of natural video exhibits complex-cell-like pooling of simple cell outputs
Jarmo Hurri, Aapo Hyvärinen
Neurocomputing2
2003 Connection between multilayer perceptrons and regression using independent component analysis
Aapo Hyvärinen, Ella Bingham
Neurocomputing1
2003 Simple-Cell-Like Receptive Fields Maximize Temporal Coherence in Natural Video
abstract
Recently, statistical models of natural images have shown the emergence of several properties of the visual cortex. Most models have considered the nongaussian properties of static image patches, leading to sparse coding or independent component analysis. Here we consider the basic time dependencies of image sequences instead of their nongaussianity. We show that simple-cell-type receptive fields emerge when temporal response strength correlation is maximized for natural image sequences. Thus, temporal response strength correlation, which is a nonlinear measure of temporal coherence, provides an alternative to sparseness in modeling simple-cell receptive field properties. Our results also suggest an interpretation of simple cells in terms of invariant coding principles, which have previously been used to explain complex-cell receptive fields.
Jarmo Hurri, Aapo Hyvärinen
Neural Comput.2
2002 Receptive Fields Similar to Simple Cells Maximize Temporal Coherence in Natural Video
Jarmo Hurri, Aapo Hyvärinen
ICANN2
2002 Interpreting Neural Response Variability as Monte Carlo Sampling of the Posterior
abstract
The responses of cortical sensory neurons are notoriously variable, with the number of spikes evoked by identical stimuli varying significantly from trial to trial. This variability is most often interpreted as ‘noise’, purely detrimental to the sensory system. In this paper, we propose an al- ternative view in which the variability is related to the uncertainty, about world parameters, which is inherent in the sensory stimulus. Specifi- cally, the responses of a population of neurons are interpreted as stochas- tic samples from the posterior distribution in a latent variable model. In addition to giving theoretical arguments supporting such a representa- tional scheme, we provide simulations suggesting how some aspects of response variability might be understood in this framework.
Patrik O. Hoyer, Aapo Hyvärinen
NIPS2
2002 Temporal Coherence, Natural Image Sequences, and the Visual Cortex
abstract
We show that two important properties of the primary visual cortex emerge when the principle of temporal coherence is applied to natural image sequences. The properties are simple-cell-like receptive fields and complex-cell-like pooling of simple cell outputs, which emerge when we apply two different approaches to temporal coherence. In the first approach we extract receptive fields whose outputs are as temporally co- herent as possible. This approach yields simple-cell-like receptive fields (oriented, localized, multiscale). Thus, temporal coherence is an alterna- tive to sparse coding in modeling the emergence of simple cell receptive fields. The second approach is based on a two-layer statistical generative model of natural image sequences. In addition to modeling the temporal coherence of individual simple cells, this model includes inter-cell tem- poral dependencies. Estimation of this model from natural data yields both simple-cell-like receptive fields, and complex-cell-like pooling of simple cell outputs. In this completely unsupervised learning, both lay- ers of the generative model are estimated simultaneously from scratch. This is a significant improvement on earlier statistical models of early vision, where only one layer has been learned, and others have been fixed a priori.
Jarmo Hurri, Aapo Hyvärinen
NIPS2
2002 Blind signal separation and independent component analysis
Shun-ichi Amari, Aapo Hyvärinen, Soo-Young Lee, Te-Won Lee, V. David Sánchez A.
Neurocomputing2
2002 Sparse coding of natural contours
Patrik O. Hoyer, Aapo Hyvärinen
Neurocomputing2
2002 An alternative approach to infomax and independent component analysis
Aapo Hyvärinen
Neurocomputing1
2002 Imposing sparsity on the mixing matrix in independent component analysis
Aapo Hyvärinen, Karthikesh Raju
Neurocomputing1
2002 Realizations of quantum computing using optical manipulations of atoms
Aapo Hyvärinen
Nat. Comput.1
2001 Topographic independent component analysis as a model of V1 organization and receptive fields
Aapo Hyvärinen, Patrik O. Hoyer
Neurocomputing1
2001 Complexity Pursuit: Separating Interesting Components from Time Series
abstract
A generalization of projection pursuit for time series, that is, signals with time structure, is introduced. The goal is to find projections of time series that have interesting structure, defined using criteria related to Kolmogoroff complexity or coding length. Interesting signals are those that can be coded with a short code length. We derive a simple approximation of coding length that takes into account both the nongaussianity and the autocorrelations of the time series. Also, we derive a simple algorithm for its approximative optimization. The resulting method is closely related to blind separation of nongaussian, time-dependent source signals.
Aapo Hyvärinen
Neural Comput.1
2001 Topographic Independent Component Analysis
abstract
In ordinary independent component analysis, the components are assumed to be completely independent, and they do not necessarily have any meaningful order relationships. In practice, however, the estimated "independent" components are often not at all independent. We propose that this residual dependence structure could be used to define a topographic order for the components. In particular, a distance between two components could be defined using their higher-order correlations, and this distance could be used to create a topographic representation. Thus, we obtain a linear decomposition into approximately independent components, where the dependence of two components is approximated by the proximity of the components in the topographic representation.
Aapo Hyvärinen, Patrik O. Hoyer, Mika Inki
Neural Comput.1
2001 Blind source separation by nonstationarity of variance: a cumulant-based approach
abstract
Blind separation of source signals usually relies either on the nonGaussianity of the signals or on their linear autocorrelations. A third approach was introduced by Matsuoka et al. (1995), who showed that source separation can be performed by using the nonstationarity of the signals, in particular the nonstationarity of their variances. In this paper, we show how to interpret the nonstationarity due to a smoothly changing variance in terms of higher order cross-cumulants. This is based on the time-correlation of the squares (energies) of the signals and leads to a simple optimization criterion. Using this criterion, we construct a fixed-point algorithm that is computationally very efficient.
Aapo Hyvärinen
IEEE Trans. Neural Networks1
2000 ICA of Complex Valued Signals: A Fast and Robust Deflationary Algorithm
abstract
Separation of complex valued signals is a frequently arising problem in signal processing. In this article it is assumed that the original, complex valued source signals are mutually statistically independent, and the problem is solved by the independent component analysis (ICA) model. ICA is a statistical method for transforming an observed multidimensional random vector into components that are mutually as independent as possible. In this article, a fast fixed-point type algorithm that is capable of separating complex valued, linearly mixed source signals is presented and its computational efficiency is shown by simulations. We also present a theorem on the local consistency of the estimator given by the algorithm.
Ella Bingham, Aapo Hyvärinen
IJCNN (3)2
2000 Feature Extraction from Color and Stereo Images Using ICA
abstract
Previous work has shown that independent component analysis (ICA) applied to natural image data yields features resembling Gabor functions and simple-cell receptive fields. This article considers the effects of including chromatic and stereo information. The inclusion of colour leads to features divided into separate red/green, blue/yellow, and bright/dark channels. Stereo image data, on the other hand, leads to binocular receptive fields which are tuned to various disparities. The similarities between these results and observed properties of simple cells in primary visual cortex are further evidence for the hypothesis that visual cortical neurons perform some type of redundancy reduction, which was one of the original motivations for ICA in the first place. In addition, ICA provides a principled method for feature extraction from colour and stereo images, such features could be used in image processing operations such as denoising, compression, and pattern recognition.
Patrik O. Hoyer, Aapo Hyvärinen
IJCNN (3)2
2000 Topographic ICA as a Model of V1 Receptive Fields
abstract
Independent component analysis (ICA), which is equivalent to linear sparse coding, has been recently used as a model of natural image statistics and V1 receptive fields. Olshausen and Field applied the principle of maximizing the sparseness of the coefficients of a linear representation to extract features from natural images. This leads to the emergence of oriented linear filters that have simultaneous localization in space and in frequency, thus resembling Gabor functions and V1 simple cell receptive fields. In this paper, we extend this model to explain emergence of V1 topography. This is done by ordering the basis vectors so that vectors with strong higher-order correlations are near to each other. This is a new principle of topographic organization, and may be more relevant to natural image statistics than the more conventional topographic ordering based on Euclidean distances. For example, this topographic ordering leads to simultaneous emergence of complex cell properties: each neighbourhood acts like a complex cell.
Aapo Hyvärinen, Patrik O. Hoyer, Mika Inki
IJCNN (4)1
2000 A Fast Fixed-Point Algorithm for Independent Component Analysis of Complex Valued Signals
abstract
Separation of complex valued signals is a frequently arising problem in signal processing. For example, separation of convolutively mixed source signals involves computations on complex valued signals. In this article, it is assumed that the original, complex valued source signals are mutually statistically independent, and the problem is solved by the independent component analysis (ICA) model. ICA is a statistical method for transforming an observed multidimensional random vector into components that are mutually as independent as possible. In this article, a fast fixed-point type algorithm that is capable of separating complex valued, linearly mixed source signals is presented and its computational efficiency is shown by simulations. Also, the local consistency of the estimator given by the algorithm is proved.
Ella Bingham, Aapo Hyvärinen
Int. J. Neural Syst.2
2000 Emergence of Phase- and Shift-Invariant Features by Decomposition of Natural Images into Independent Feature Subspaces
abstract
Olshausen and Field (1996) applied the principle of independence maximization by sparse coding to extract features from natural images. This leads to the emergence of oriented linear filters that have simultaneous localization in space and in frequency, thus resembling Gabor functions and simple cell receptive fields. In this article, we show that the same principle of independence maximization can explain the emergence of phase- and shift-invariant features, similar to those found in complex cells. This new kind of emergence is obtained by maximizing the independence between norms of projections on linear subspaces (instead of the independence of simple linear filter outputs). The norms of the projections on such "independent feature subspaces" then indicate the values of invariant features.
Aapo Hyvärinen, Patrik O. Hoyer
Neural Comput.1
2000 Independent component analysis: algorithms and applications
Aapo Hyvärinen, Erkki Oja
Neural Networks1
1999 Estimating signal-adapted wavelets using sparseness criteria
abstract
Multiresolution transforms have been shown to be effective for a variety of digital signal processing tasks. Recently, the task of adapting these usually fixed transforms to the statistics of the data has attracted much attention. So far, however, the methods proposed have been based exclusively on the second-order statistics of the signal. We show how to take into account higher order statistics to estimate a multiresolution transform from white data. The method is tested on speech data from the TIMIT database and is shown to give filters well adapted to the structure of the data.
Patrik O. Hoyer, Aapo Hyvärinen
IJCNN2
1999 A fast algorithm for estimating overcomplete ICA bases for image windows
abstract
We introduce a very fast method for estimating over-complete bases of independent components from image data. This is based on the concept of quasi-orthogonality, which means that in a very high-dimensional space, there can be a large, over-complete set of vectors that are almost orthogonal to each other. Thus we may estimate an over-complete basis by using one-unit ICA algorithms and forcing only partial decorrelation between the different independent components. The method can be implemented using a modification of the FastICA algorithm, which leads to a computationally highly efficient method.
Aapo Hyvärinen, Razvan Cristescu, Erkki Oja
IJCNN1
1999 Independent subspace analysis shows emergence of phase and shift invariant features from natural images
abstract
Olshausen and Field (1996, 1997) applied the principle of independence maximization by sparse coding to extract features from natural images. This leads to the emergence of oriented linear filters that have simultaneous localization in space and in frequency, thus resembling Gabor functions and simple cell receptive fields. In this paper, we show that the same principle of independence maximization can explain the emergence of phase and shift invariant features, similar to those found in complex cells. This new kind of emergence is obtained by maximizing the independence between norms of projections on linear subspaces (instead of the independence of simple linear filter outputs). The norms of the projections on such 'independent feature subspaces' then indicate the values of invariant features.
Aapo Hyvärinen, Patrik O. Hoyer
IJCNN1
1999 Emergence of Topography and Complex Cell Properties from Natural Images using Extensions of ICA
Aapo Hyvärinen, Patrik O. Hoyer
NIPS1
1999 Sparse Code Shrinkage: Denoising of Nongaussian Data by Maximum Likelihood Estimation
abstract
Sparse coding is a method for finding a representation of data in which each of the components of the representation is only rarely significantly active. Such a representation is closely related to redundancy reduction and independent component analysis, and has some neurophysiological plausibility. In this article, we show how sparse coding can be used for denoising. Using maximum likelihood estimation of nongaussian variables corrupted by gaussian noise, we show how to apply a soft-thresholding (shrinkage) operator on the components of sparse coding so as to reduce noise. Our method is closely related to the method of wavelet shrinkage, but it has the important benefit over wavelet methods that the representation is determined solely by the statistical properties of the data. The wavelet representation, on the other hand, relies heavily on certain mathematical properties (like self-similarity) that may be only weakly related to the properties of natural data.
Aapo Hyvärinen
Neural Comput.1
1999 Nonlinear independent component analysis: Existence and uniqueness results
Aapo Hyvärinen, Petteri Pajunen
Neural Networks1
1999 The Fixed-Point Algorithm and Maximum Likelihood Estimation for Independent Component Analysis
Aapo Hyvärinen
Neural Process. Lett.1
1999 Image Feature Extraction and Denoising by Sparse Coding
Erkki Oja, Aapo Hyvärinen, Patrik O. Hoyer
Pattern Anal. Appl.2
1999 Gaussian moments for noisy independent component analysis
abstract
A novel approach for the problem of estimating the data model of independent component analysis (or blind source separation) in the presence of Gaussian noise is introduced. We define the Gaussian moments of a random variable as the expectations of the Gaussian function (and some related functions) with different scale parameters, and show how the Gaussian moments of a random variable can be estimated from noisy observations. This enables us to use Gaussian moments as one-unit contrast functions that have no asymptotic bias even in the presence of noise, and that are robust against outliers. To implement the maximization of the contrast functions based on Gaussian moments, a modification of the fixed-point (FastICA) algorithm is introduced.
Aapo Hyvärinen
IEEE Signal Process. Lett.1
1999 Fast and robust fixed-point algorithms for independent component analysis
abstract
Independent component analysis (ICA) is a statistical method for transforming an observed multidimensional random vector into components that are statistically as independent from each other as possible. In this paper, we use a combination of two different approaches for linear ICA: Comon's information-theoretic approach and the projection pursuit approach. Using maximum entropy approximations of differential entropy, we introduce a family of new contrast (objective) functions for ICA. These contrast functions enable both the estimation of the whole decomposition by minimizing mutual information, and estimation of individual independent components as projection pursuit directions. The statistical properties of the estimators based on such contrast functions are analyzed under the assumption of the linear mixture model, and it is shown how to choose contrast functions that are robust and/or of minimum variance. Finally, we introduce simple fixed-point algorithms for practical optimization of the contrast functions. These algorithms optimize the contrast functions very fast and reliably.
Aapo Hyvärinen
IEEE Trans. Neural Networks1
1998 Image feature extraction by sparse coding and independent component analysis
abstract
Sparse coding is a method for finding a representation of data in which each of the components of the representation is only rarely significantly active. Such a representation is closely related to the techniques of independent component analysis and blind source separation. In this paper, we investigate the application of sparse coding for image feature extraction. We show how sparse coding can be used to extract wavelet-like features from natural image data. As an application of such a feature extraction scheme, we show how to apply a soft-thresholding operator on the components of sparse coding in order to reduce Gaussian noise. Methods based on sparse coding have the important benefit over wavelet methods that the features are determined solely by the statistical properties of the data, while the wavelet transformation relies heavily on certain abstract mathematical properties that may be only weakly related to the properties of the natural data.
Aapo Hyvärinen, Erkki Oja, Patrik O. Hoyer, Jarmo Hurri
ICPR1
1998 Sparse Code Shrinkage: Denoising by Nonlinear Maximum Likelihood Estimation
Aapo Hyvärinen, Patrik O. Hoyer, Erkki Oja
NIPS1
1998 Independent component analysis in the presence of Gaussian noise by maximizing joint likelihood
Aapo Hyvärinen
Neurocomputing1
1998 Independent component analysis by general nonlinear Hebbian-like learning rules
Aapo Hyvärinen, Erkki Oja
Signal Process.1
1997 From Neural Principal Components to Neural Independent Components
Erkki Oja, Juha Karhunen, Aapo Hyvärinen
ICANN3
1997 A family of fixed-point algorithms for independent component analysis
abstract
Independent component analysis (ICA) is a statistical signal processing technique whose main applications are blind source separation, blind deconvolution, and feature extraction. Estimation of ICA is usually performed by optimizing a 'contrast' function based on higher-order cumulants. It is shown how almost any error function can be used to construct a contrast function to perform the ICA estimation. In particular, this means that one can use contrast functions that are robust against outliers. As a practical method for finding the relevant extrema of such contrast functions, a fixed-point iteration scheme is then introduced. The resulting algorithms are quite simple and converge fast and reliably. These algorithms also enable estimation of the independent components one-by-one, using a simple deflation scheme.
Aapo Hyvärinen
ICASSP1
1997 Applications of neural blind separation to signal and image processing
abstract
In blind source separation one tries to separate statistically independent unknown source signals from their linear mixtures without knowing the mixing coefficients. Such techniques are currently studied actively both in statistical signal processing and unsupervised neural learning. We apply neural blind separation techniques developed in our laboratory to the extraction of features from natural images and to the separation of medical EEG signals. The new analysis method yields features that describe the underlying data better than for example classical principal component analysis. We discuss difficulties related with real-world applications of blind signal processing, too.
Juha Karhunen, Aapo Hyvärinen, Ricardo Vigário, Jarmo Hurri, Erkki Oja
ICASSP2
1997 New Approximations of Differential Entropy for Independent Component Analysis and Projection Pursuit
Aapo Hyvärinen
NIPS1
1997 A Fast Fixed-Point Algorithm for Independent Component Analysis
abstract
We introduce a novel fast algorithm for independent component analysis, which can be used for blind source separation and feature extraction. We show how a neural network learning rule can be transformed into a fixedpoint iteration, which provides an algorithm that is very simple, does not depend on any user-defined parameters, and is fast to converge to the most accurate solution allowed by the data. The algorithm finds, one at a time, all nongaussian independent components, regardless of their probability distributions. The computations can be performed in either batch mode or a semiadaptive manner. The convergence of the algorithm is rigorously proved, and the convergence speed is shown to be cubic. Some comparisons to gradient-based algorithms are made, showing that the new algorithm is usually 10 to 100 times faster, sometimes giving the solution in just a few iterations.
Aapo Hyvärinen, Erkki Oja
Neural Comput.1
1996 Purely Logical Neural Principal Component and Independent Component Learning
Aapo Hyvärinen
ICANN1
1996 One-unit Learning Rules for Independent Component Analysis
Aapo Hyvärinen, Erkki Oja
NIPS1
1996 Simple Neuron Models for Independent Component Analysis
Aapo Hyvärinen, Erkki Oja
Int. J. Neural Syst.1