Alexander Ilin

dblp:85/5835 · DBLP profile ↗
← Back
46ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0001-6419-3006ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 10 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Tuning Qwen2.5-VL to Improve Its Web Interaction Skills
abstract
Recent advances in vision–language models (VLMs) have sparked growing interest in using them to automate web tasks, yet their feasibility as independent agents that reason and act purely from visual input remains underexplored. We investigate this setting using Qwen2.5-VL-32B, one of the strongest open-source VLMs available, and focus on improving its reliability in web-based control. Through initial experimentation, we observe three key challenges: (i)~inaccurate localization of target elements, the cursor, and their relative positions, (ii)~sensitivity to instruction phrasing, and (iii)~an overoptimistic bias toward its own actions, often assuming they succeed rather than analyzing their actual outcomes. To address these issues, we fine-tune Qwen2.5-VL-32B for a basic web interaction task: moving the mouse and clicking on a page element described in natural language. Our training pipeline consists of two stages: (1)~teaching the model to determine whether the cursor already hovers over the target element or whether movement is required, and (2)~training it to execute a single command (a mouse move or a mouse click) at a time, verifying the resulting state of the environment before planning the next action. Evaluated on a custom benchmark of single-click web tasks, our approach increases success rates from 86% to 94% under the most challenging setting.
Alexandra Yakovleva, Henrik Pärssinen, Harri Valpola, Juho Kannala, Alexander Ilin
WWW5
2024 Generating Demonstrations for In-Context Compositional Generalization in Grounded Language Learning
abstract
In-Context-learning and few-shot prompting are viable methods compositional output generation.However, these methods can be very sensitive to the choice of support examples used.Retrieving good supports from the training data for a given test query is already a difficult problem, but in some cases solving this may not even be enough.We consider the setting of grounded language learning problems where finding relevant supports in the same or similar states as the query may be difficult.We design an agent which instead generates possible supports inputs and targets current state of the world, then uses them in-context-learning to solve the test query.We show substantially improved performance on a previously unsolved compositional generalization test without a loss of performance in other areas.The approach is general and can even scale to instructions expressed in natural language.0.61 h 0.57 ± .50 (60) 0.61 h 0.86 ± .34(266) 0.64 h 0.59 ± .49(743) 0.65 h 0.87 ± .33 (2655) 0.68 h 0.59 ± .49(2907) 0.69 h 0.88 ± .32 (8144) 0.71 h 0.63 ± .48(4412) 0.73 h 0.86 ± .35(7480) 0.74 h 0.80 ± .40 (6025) 0.77 h 0.78 ± .41(7283) 0.78 h 0.81 ± .39 (10824) 0.81 h 0.76 ± .43 (5358)
Sam Spilsbury, Pekka Marttinen, Alexander Ilin
EMNLP3
2023 Reader: Model-based language-instructed reinforcement learning
abstract
We explore how we can build accurate world models, which are partially specified by language, and how we can plan with them in the face of novelty and uncertainty. We propose the first model-based reinforcement learning approach to tackle the environment Read To Fight Monsters (Zhong et al., 2019), a grounded policy learning problem. In RTFM an agent has to reason over a set of rules and a goal, both described in a language manual, and the observations, while taking into account the uncertainty arising from the stochasticity of the environment, in order to generalize successfully its policy to test episodes. We demonstrate the superior performance and sample efficiency of our model-based approach to the existing model-free SOTA agents in eight variants of RTFM. Furthermore, we show how the agent’s plans can be inspected, which represents progress towards more interpretable agents.
Nicola Dainese, Pekka Marttinen, Alexander Ilin
EMNLP3
2023 Hierarchical Imitation Learning with Vector Quantized Models
abstract
The ability to plan actions on multiple levels of abstraction enables intelligent agents to solve complex tasks effectively. However, learning the models for both low and high-level planning from demonstrations has proven challenging, especially with higher-dimensional inputs. To address this issue, we propose to use reinforcement learning to identify subgoals in expert trajectories by associating the magnitude of the rewards with the predictability of low-level actions given the state and the chosen subgoal. We build a vector-quantized generative model for the identified subgoals to perform subgoal-level planning. In experiments, the algorithm excels at solving complex, long-horizon decision-making problems outperforming state-of-the-art. Because of its ability to plan, our algorithm can find better trajectories than the ones in the training set.
Kalle Kujanpää, Joni Pajarinen, Alexander Ilin
ICML3
2023 Improved Training of Physics-Informed Neural Networks with Model Ensembles
abstract
Learning the solution of partial differential equations (PDEs) with a neural network is an attractive alternative to traditional solvers due to its elegance, greater flexibility and the ease of incorporating observed data. However, training such physics-informed neural networks (PINNs) is notoriously difficult in practice since PINNs often converge to wrong solutions. In this paper, we address this problem by training an ensemble of PINNs. Our approach is motivated by the observation that individual PINN models find similar solutions in the vicinity of points with targets (e.g., observed data or initial conditions) while their solutions may substantially differ farther away from such points. Therefore, we propose to use the ensemble agreement as the criterion for gradual expansion of the solution interval, that is including new points for computing the loss derived from differential equations. Due to the flexibility of the domain expansion, our algorithm can easily incorporate measurements in arbitrary locations. In contrast to the existing PINN algorithms with time-adaptive strategies, the proposed algorithm does not need a predefined schedule of interval expansion and it treats time and space equally. We experimentally show that the proposed algorithm can stabilize PINN training and yield performance competitive to the recent variants of PINNs trained with time adaptation.
Katsiaryna Haitsiukevich, Alexander Ilin
IJCNN2
2023 Hybrid Search for Efficient Planning with Completeness Guarantees
abstract
Solving complex planning problems has been a long-standing challenge in computer science. Learning-based subgoal search methods have shown promise in tackling these problems, but they often suffer from a lack of completeness guarantees, meaning that they may fail to find a solution even if one exists. In this paper, we propose an efficient approach to augment a subgoal search method to achieve completeness in discrete action spaces. Specifically, we augment the high-level search with low-level actions to execute a multi-level (hybrid) search, which we call complete subgoal search. This solution achieves the best of both worlds: the practical efficiency of high-level search and the completeness of low-level search. We apply the proposed search method to a recently proposed subgoal search algorithm and evaluate the algorithm trained on offline data on complex planning problems. We demonstrate that our complete subgoal search not only guarantees completeness but can even improve performance in terms of search expansions for instances that the high-level could solve without low-level augmentations. Our approach makes it possible to apply subgoal-level planning for systems where completeness is a critical requirement.
Kalle Kujanpää, Joni Pajarinen, Alexander Ilin
NeurIPS3
2022 Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning
abstract
Offline reinforcement learning, by learning from a fixed dataset, makes it possible to learn agent behaviors without interacting with the environment.However, depending on the quality of the offline dataset, such pre-trained agents may have limited performance and would further need to be fine-tuned online by interacting with the environment.During online fine-tuning, the performance of the pre-trained agent may collapse quickly due to the sudden distribution shift from offline to online data.We propose to adaptively weigh the behavior cloning loss during online fine-tuning based on the agent's performance and training stability.Moreover, we use a randomized ensemble of Q functions to further increase the sample efficiency of online fine-tuning by performing a large number of learning updates.Experiments show that the proposed method yields state-of-the-art offline-to-online reinforcement learning performance on the popular D4RL benchmark.
Yi Zhao 0014, Rinu Boney, Alexander Ilin, Juho Kannala, Joni Pajarinen
ESANN3
2022 Learning Trajectories of Hamiltonian Systems with Neural Networks
Katsiaryna Haitsiukevich, Alexander Ilin
ICANN (1)2
2022 RNA secondary structure prediction with convolutional neural networks
abstract
BACKGROUND: Predicting the secondary, i.e. base-pairing structure of a folded RNA strand is an important problem in synthetic and computational biology. First-principle algorithmic approaches to this task are challenging because existing models of the folding process are inaccurate, and even if a perfect model existed, finding an optimal solution would be in general NP-complete. RESULTS: In this paper, we propose a simple, yet effective data-driven approach. We represent RNA sequences in the form of three-dimensional tensors in which we encode possible relations between all pairs of bases in a given sequence. We then use a convolutional neural network to predict a two-dimensional map which represents the correct pairings between the bases. Our model achieves significant accuracy improvements over existing methods on two standard datasets, RNAStrAlign and ArchiveII, for 10 RNA families, where our experiments show excellent performance of the model across a wide range of sequence lengths. Since our matrix representation and post-processing approaches do not require the structures to be pseudoknot-free, we get similar good performance also for pseudoknotted structures. CONCLUSION: We show how to use an artificial neural network design to predict the structure for a given RNA sequence with high accuracy only by learning from samples whose native structures have been experimentally characterized, independent of any energy model.
Mehdi Saman Booy, Alexander Ilin, Pekka Orponen
BMC Bioinform.2
2022 A Relational Model for One-Shot Classification of Images and Pen Strokes
abstract
We show that a deep learning model with built-in relational inductive bias can bring benefits to sample-efficient learning, without relying on extensive data augmentation. Our study shows that excellent results can be achieved with a model in which the relational inductive bias is applied to images, while building an efficient one-shot classifier on top of raw strokes is more challenging. The proposed one-shot classification model performs relational matching of a pair of inputs in the form of local and pairwise attention. Our approach solves with almost perfect accuracy the one-shot image classification Omniglot challenge when combined with a Hungarian matching algorithm and attains competitive results on the same task on characters represented as rotation-augmented strokes.
Arturs Polis, Alexander Ilin
Neurocomputing2
2022 Learning to Play Imperfect-Information Games by Imitating an Oracle Planner
abstract
We consider learning to play multiplayer imperfect-information games with simultaneous moves and large state-action spaces. Previous attempts to tackle such challenging games have largely focused on model-free learning methods, often requiring hundreds of years of experience to produce competitive agents. Our approach is based on model-based planning. We tackle the problem of partial observability by first building an (oracle) planner that has access to the full state of the environment and then distilling the knowledge of the oracle to a (follower) agent which is trained to play the imperfect-information game by imitating the oracle’s choices. We experimentally show that planning with naive Monte Carlo tree search does not perform very well in large combinatorial action spaces. We, therefore, propose planning with a fixed-depth tree search and decoupled TS for action selection. We show that the planner is able to discover efficient playing strategies in the games ofClash RoyaleandPommermanand the follower policy successfully learns to implement them by training on a few hundred battles.
Rinu Boney, Alexander Ilin, Juho Kannala, Jarno Seppänen
IEEE Trans. Games2
2021 A Relational Model for One-Shot Classification
abstract
We show that a deep learning model with built-in relational inductive bias can bring benefits to sample-efficient learning, without relying on extensive data augmentation.The proposed one-shot classification model performs relational matching of a pair of inputs in the form of local and pairwise attention.Our approach solves perfectly the one-shot image classification Omniglot challenge.Our model exceeds human level accuracy, as well as the previous state of the art, with no data augmentation.
Arturs Polis, Alexander Ilin
ESANN2
2021 Learning to Assist Agents by Observing Them
Antti Keurulainen, Isak Westerlund, Samuel Kaski, Alexander Ilin
ICANN (4)4
2021 Behaviour-Conditioned Policies for Cooperative Reinforcement Learning Tasks
Antti Keurulainen, Isak Westerlund, Ariel Kwiatkowski, Samuel Kaski, Alexander Ilin
ICANN (4)5
2021 Path-Link Graph Neural Network for IP Network Performance Prediction
Yangzhe Kong, Dmitry Petrov, Vilho Räisänen, Alexander Ilin
IM4
2021 A Grid-Structured Model of Tubular Reactors
abstract
We propose a grid-like computational model of tubular reactors. The architecture is inspired by the computations performed by solvers of partial differential equations which describe the dynamics of the chemical process inside a tubular reactor. The proposed model may be entirely based on the known form of the partial differential equations or it may contain generic machine learning components such as multi-layer perceptrons. We show that the proposed model can be trained using limited amounts of data to describe the state of a fixed-bed catalytic reactor. The trained model can reconstruct unmeasured states such as the catalyst activity using the measurements of inlet concentrations and temperatures along the reactor.
Katsiaryna Haitsiukevich, Samuli Bergman, Cesar de Araujo Filho, Francesco Corona, Alexander Ilin
INDIN5
2020 Conditional Spoken Digit Generation with StyleGAN
abstract
This paper adapts a StyleGAN model for speech generation with minimal or no conditioning on text. StyleGAN is a multi-scale convolutional GAN capable of hierarchically capturing data structure and latent variation on multiple spatial (or temporal) levels. The model has previously achieved impressive results on facial image generation, and it is appealing to audio applications due to similar multi-level structures present in the data. In this paper, we train a StyleGAN to generate mel-frequency spectrograms on the Speech Commands dataset, which contains spoken digits uttered by multiple speakers in varying acoustic conditions. In a conditional setting our model is conditioned on the digit identity, while learning the remaining data variation remains an unsupervised task. We compare our model to the current unsupervised state-of-the-art speech synthesis GAN architecture, the WaveGAN, and show that the proposed model outperforms according to numerical measures and subjective evaluation by listening tests.
Kasperi Palkama, Lauri Juvela, Alexander Ilin
INTERSPEECH3
2019 Active one-shot learning with Prototypical Networks
Rinu Boney, Alexander Ilin
ESANN2
2019 Regularizing Trajectory Optimization with Denoising Autoencoders
abstract
Trajectory optimization using a learned model of the environment is one of the core elements of model-based reinforcement learning. This procedure often suffers from exploiting inaccuracies of the learned model. We propose to regularize trajectory optimization by means of a denoising autoencoder that is trained on the same trajectories as the model of the environment. We show that the proposed regularization leads to improved planning with both gradient-based and gradient-free optimizers. We also demonstrate that using regularized trajectory optimization leads to rapid initial learning in a set of popular motor control tasks, which suggests that the proposed approach can be a useful tool for improving sample efficiency.
Rinu Boney, Norman Di Palo, Mathias Berglund, Alexander Ilin, Juho Kannala, Antti Rasmus, Harri Valpola
NeurIPS4
2017 Recurrent Ladder Networks
abstract
We propose a recurrent extension of the Ladder networks whose structure is motivated by the inference required in hierarchical latent variable models. We demonstrate that the recurrent Ladder is able to handle a wide variety of complex learning tasks that benefit from iterative inference and temporal modeling. The architecture shows close-to-optimal results on temporal modeling of video data, competitive results on music modeling, and improved perceptual grouping based on higher order abstractions, such as stochastic textures and motion cues. We present results for fully supervised, semi-supervised, and unsupervised tasks. The results suggest that the proposed architecture and principles are powerful tools for learning a hierarchy of abstractions, learning iterative inference and handling temporal information.
Isabeau Prémont-Schwarz, Alexander Ilin, Tele Hao, Antti Rasmus, Rinu Boney, Harri Valpola
NIPS2
2014 Linear State-Space Model with Time-Varying Dynamics
Jaakko Luttinen, Tapani Raiko, Alexander Ilin
ECML/PKDD (2)3
2013 A Two-Stage Pretraining Algorithm for Deep Boltzmann Machines
Kyunghyun Cho, Tapani Raiko, Alexander Ilin, Juha Karhunen
ICANN3
2013 Gaussian-Bernoulli restricted Boltzmann machines and automatic feature extraction for noise robust missing data mask estimation
abstract
A missing data mask estimation method based on Gaussian-Bernoulli restricted Boltzmann machine (GRBM) trained on cross-correlation representation of the audio signal is presented in the study. The automatically learned features by the GRBM are utilized in dividing the time-frequency units of the spectrographic mask into noise and speech dominant. The system is evaluated against two baseline mask estimation methods in a reverberant multisource environment speech recognition task. The proposed system is shown to provide a performance improvement in the speech recognition accuracy over the previous multifeature approaches.
Sami Keronen, Kyunghyun Cho, Tapani Raiko, Alexander Ilin, Kalle J. Palomäki
ICASSP4
2013 Gaussian-Bernoulli deep Boltzmann machine
abstract
In this paper, we study a model that we call Gaussian-Bernoulli deep Boltzmann machine (GDBM) and discuss potential improvements in training the model. GDBM is designed to be applicable to continuous data and it is constructed from Gaussian-Bernoulli restricted Boltzmann machine (GRBM) by adding multiple layers of binary hidden neurons. The studied improvements of the learning algorithm for GDBM include parallel tempering, enhanced gradient, adaptive learning rate and layer-wise pretraining. We empirically show that they help avoid some of the common difficulties found in training deep Boltzmann machines such as divergence of learning, the difficulty in choosing right learning rate scheduling, and the existence of meaningless higher layers.
Kyunghyun Cho, Tapani Raiko, Alexander Ilin
IJCNN3
2013 Enhanced Gradient for Training Restricted Boltzmann Machines
abstract
Restricted Boltzmann machines (RBMs) are often used as building blocks in greedy learning of deep networks. However, training this simple model can be laborious. Traditional learning algorithms often converge only with the right choice of metaparameters that specify, for example, learning rate scheduling and the scale of the initial weights. They are also sensitive to specific data representation. An equivalent RBM can be obtained by flipping some bits and changing the weights and biases accordingly, but traditional learning rules are not invariant to such transformations. Without careful tuning of these training settings, traditional algorithms can easily get stuck or even diverge. In this letter, we present an enhanced gradient that is derived to be invariant to bit-flipping transformations. We experimentally show that the enhanced gradient yields more stable training of RBMs both when used with a fixed learning rate and an adaptive one.
Kyunghyun Cho, Tapani Raiko, Alexander Ilin
Neural Comput.3
2012 Tikhonov-Type Regularization for Restricted Boltzmann Machines
Kyunghyun Cho, Alexander Ilin, Tapani Raiko
ICANN (1)2
2012 Gated Boltzmann Machine in Texture Modeling
Tele Hao, Tapani Raiko, Alexander Ilin, Juha Karhunen
ICANN (2)3
2012 Bayesian Robust PCA of Incomplete Data
Jaakko Luttinen, Alexander Ilin, Juha Karhunen
Neural Process. Lett.2
2011 Improved Learning of Gaussian-Bernoulli Restricted Boltzmann Machines
Kyunghyun Cho, Alexander Ilin, Tapani Raiko
ICANN (1)2
2011 Enhanced Gradient and Adaptive Learning Rate for Training Restricted Boltzmann Machines
Kyunghyun Cho, Tapani Raiko, Alexander Ilin
ICML3
2010 Parallel tempering is efficient for learning restricted Boltzmann machines
abstract
A new interest towards restricted Boltzmann machines (RBMs) has risen due to their usefulness in greedy learning of deep neural networks. While contrastive divergence learning has been considered an efficient way to learn an RBM, it has a drawback due to a biased approximation in the learning gradient. We propose to use an advanced Monte Carlo method called parallel tempering instead, and show experimentally that it works efficiently.
Kyunghyun Cho, Tapani Raiko, Alexander Ilin
IJCNN3
2010 Transformations in variational Bayesian factor analysis to speed up learning
Jaakko Luttinen, Alexander Ilin
Neurocomputing2
2010 Practical Approaches to Principal Component Analysis in the Presence of Missing Values
Alexander Ilin, Tapani Raiko
J. Mach. Learn. Res.1
2009 Transformations for variational factor analysis to speed up learning
Jaakko Luttinen, Alexander Ilin, Tapani Raiko
ESANN2
2009 Bayesian PCA for reconstruction of historical sea surface temperatures
abstract
In this work, reconstructions of historical global sea surface temperatures (SST) are performed using Bayesian principal component analysis (PCA). Two PCA models are examined: a model with isotropic noise and a model which takes into account data uncertainty due to sampling errors. Inference is done by variational Bayesian learning. The methods are compared with a more traditional technique, reduced space optimal interpolation (RSOI), that is currently used in producing standard historical SST analyses. New methods were applied to the MOHSST5, an observational data set for 1856-1991 period from the United Kingdom Meteorological Office, that was used in a previously published application of the RSOI. Data uncertainty specification was also identical to the one used in that RSOI application, hence the performances of all reconstructions are directly comparable. Reconstruction results for 1982-1991 period are tested via their comparison with the NOAA monthly 1deg OI (version 2) that blends in situ observations with the much better sampled satellite data. New reconstructions slightly outperform the published RSOI reconstruction in this test and suggest that further improvements are possible.
Alexander Ilin, Alexey Kaplan
IJCNN1
2009 Variational Gaussian-process factor analysis for modeling spatio-temporal data
abstract
We present a probabilistic latent factor model which can be used for studying spatio-temporal datasets. The spatial and temporal structure is modeled by using Gaussian process priors both for the loading matrix and the factors. The posterior distributions are approximated using the variational Bayesian framework. High computational cost of Gaussian process modeling is reduced by using sparse approximations. The model is used to compute the reconstructions of the global sea surface temperatures from a historical dataset. The results suggest that the proposed model can outperform the state-of-the-art reconstruction systems.
Jaakko Luttinen, Alexander Ilin
NIPS2
2007 Principal Component Analysis for Large Scale Problems with Lots of Missing Values
Tapani Raiko, Alexander Ilin, Juha Karhunen
ECML2
2007 Principal Component Analysis for Sparse High-Dimensional Data
Tapani Raiko, Alexander Ilin, Juha Karhunen
ICONIP (1)2
2006 Independent dynamics subspace analysis
Alexander Ilin
ESANN1
2006 Extraction of Components with Structured Variance
abstract
We present a method for exploratory data analysis of large spatiotemporal data sets such as global longtime climate measurements, extending our previous work on semiblind source separation of climate data. The method seeks fast changing components whose variances exhibit slow behavior with specific temporal structure. The algorithm is developed in the framework of denoising source separation. It finds sources iteratively and alternates between estimating the variance structure of extracted sources and using the structure to find new source estimates. The performance of the algorithm is first demonstrated on a simple example of a semiblind source separation problem with artificially generated signals. Then, the proposed technique is applied to the global surface temperature measurements coming from the NCEP/NCAR re-analysis project. Fast changing temperature components whose variances have prominent annual and decadal structures are extracted. The extracted annual components reflect higher temperature variability over the continents during winters. The components with slower changing variances might correspond to some interesting weather phenomena characterized by slowly changing temperature variability in specific regions.
Alexander Ilin, Harri Valpola, Erkki Oja
IJCNN1
2006 Exploratory analysis of climate data using source separation methods
Alexander Ilin, Harri Valpola, Erkki Oja
Neural Networks1
2005 Semiblind source separation of climate data detects El Nino as the component with the highest interannual variability
abstract
Denoising source separation (DSS), a recently developed source separation framework, was applied to extracting components exhibiting slow, interannual temporal behavior from climate data. Three datasets with daily measurements were used: surface temperature, sea level pressure and precipitation around the globe. For all datasets, the first extracted component captured the well-known El Nino-Southern Oscillation phenomenon and the second component was close to the derivative of the first one. Several other components with slow dynamics were extracted and together the components appear to capture essential features of the slow-dynamics state of the climate system. The first two components were identified reliably but the following components may have remained mixed, nonlinear DSS could identify the physically most meaningful rotation among them but only linear DSS was within the scope of this paper. This paper offers a simple demonstration of exploratory data analysis of climate data by DSS and suggests future lines of research.
Alexander Ilin, Harri Valpola, Erkki Oja
IJCNN1
2005 Frequency-Based Separation of Climate Signals
Alexander Ilin, Harri Valpola
PKDD1
2005 On the Effect of the Form of the Posterior Approximation in Variational Learning of ICA Models
Alexander Ilin, Harri Valpola
Neural Process. Lett.1
2004 Bayesian versus constrained structure approaches for source separation in post-nonlinear mixtures
abstract
The work presents experimental comparison of two approaches introduced for solving the nonlinear blind source separation (BSS) problem: the Bayesian methods developed at Helsinki University of Technology (HUT), and the BSS methods introduced for post-nonlinear (PNL mixtures at Institut National Polytechnique de Grenoble (INPG). The comparison is performed on artificial test problems containing PNL mixtures. Both the standard case when the number of sources is equal to the number of observations and the case of overdetermined mixtures are considered. A new interesting result of the experiments is that globally invertible PNL mixtures, but with non-invertible component-wise nonlinearities, can be identified and sources can be separated, which shows the relevance of exploiting more observations than sources.
Alexander Ilin, Sophie Achard, Christian Jutten
IJCNN1
2004 Nonlinear dynamical factor analysis for state change detection
abstract
Changes in a dynamical process are often detected by monitoring selected indicators directly obtained from the process observations, such as the mean values or variances. Standard change detection algorithms such as the Shewhart control charts or the cumulative sum (CUSUM) algorithm are often based on such first- and second-order statistics. Much better results can be obtained if the dynamical process is properly modeled, for example by a nonlinear state-space model, and then the accuracy of the model is monitored over time. The success of the latter approach depends largely on the quality of the model. In practical applications like industrial processes, the state variables, dynamics, and observation mapping are rarely known accurately. Learning from data must be used; however, methods for the simultaneous estimation of the state and the unknown nonlinear mappings are very limited. We use a novel method of learning a nonlinear state-space model, the nonlinear dynamical factor analysis (NDFA) algorithm. It takes a set of multivariate observations over time and fits blindly a generative dynamical latent variable model, resembling nonlinear independent component analysis. We compare the performance of the model in process change detection to various traditional methods. It is shown that NDFA outperforms the classical methods by a wide margin in a variety of cases where the underlying process dynamics changes.
Alexander Ilin, Harri Valpola, Erkki Oja
IEEE Trans. Neural Networks1