VLDB 2026 Research / reviewers in the wild / expert
Peter Tiño
dblp:t/PeterTino · also Peter Tino
· DBLP profile ↗
153ranked-venue papers
34as first author
28since 2021 · last 2026
0000-0003-2330-128XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 124 · 30 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Cooperation to Hierarchy: A Study of Dynamics of Hierarchy Emergence in a Multi-Agent System
Shanshan Mao, Peter Tiño |
EvoApplications (1) | 2 |
| 2025 | Universality of Real Minimal Complexity ReservoirabstractReservoir Computing (RC) models, a subclass of recurrent neural networks, are distinguished by their fixed, non-trainable input layer and dynamically coupled reservoir, with only the static readout layer being trained. This design circumvents the issues associated with backpropagating error signals through time, thereby enhancing both stability and training efficiency. RC models have been successfully applied across a broad range of application domains. Crucially, they have been demonstrated to be universal approximators of time-invariant dynamic filters with fading memory, under various settings of approximation norms and input driving sources. Simple Cycle Reservoirs (SCR) represent a specialized class of RC models with a highly constrained reservoir architecture, characterized by uniform ring connectivity and binary input-to-reservoir weights with an aperiodic sign pattern. For linear reservoirs, given the reservoir size, the reservoir construction has only one degree of freedom -- the reservoir cycle weight. Such architectures are particularly amenable to hardware implementations without significant performance degradation in many practical tasks. In this study we endow these observations with solid theoretical foundations by proving that SCRs operating in real domain are universal approximators of time-invariant dynamic filters with fading memory. Our results supplement recent research showing that SCRs in the complex domain can approximate, to arbitrary precision, any unrestricted linear reservoir with a non-linear readout. We furthermore introduce a novel method to drastically reduce the number of SCR units, making such highly constrained architectures natural candidates for low-complexity hardware implementations. Our findings are supported by empirical studies on real-world time series datasets. Robert Simon Fong, Peter Tiño |
AAAI | 3 |
| 2025 | Whole-genome phenotype prediction with machine learning: open problems in bacterial genomicsabstractMOTIVATION: How can we identify causal genetic mechanisms governing bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype yield high accuracy scores. However, attempts to extract meaningful interpretations from the predictive models are found to be corrupted by falsely identified 'causal' features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those approaches to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature. RESULTS: We identify major sources of non-injectivity in the formulation of the genotype-to-phenotype mapping function-linkage-disequilibrium, limited sampling, information loss in representations, unmeasured confounders and observational noise-and analyse their implications for machine learning applications. Using a collection of 4,140 Staphylococcus aureus isolates, we illustrate challenges surrounding the defined open problems. AVAILABILITY AND IMPLEMENTATION: Raw sequencing data are available from the European Nucleotide Archive (ENA) under project accessions ERP001012, PRJEB3174, PRJEB2655, PRJEB2756, and PRJEB2944. Assemblies and annotations were generated with the Sanger bacterial pipeline (https://github.com/sanger-pathogens/vr-codebase) and unitigs extracted using DBGWAS (https://gitlab.com/leoisl/dbgwas). Tamsin James, Ben Williamson, Peter Tiño, Nicole Wheeler |
Bioinform. | 3 |
| 2025 | Interpretable modelling and visualization of biomedical dataabstractApplications of interpretable machine learning (ML) techniques on medical datasets facilitate early and fast diagnoses, along with getting deeper insight into the data. Furthermore, the transparency of these models increase trust among application domain experts. Medical datasets face common issues such as heterogeneous measurements, imbalanced classes with limited sample size, and missing data, which hinder the straightforward application of ML techniques. In this paper we present a family of prototype-based (PB) interpretable models which are capable of handling these issues. Moreover we propose a strategy of harnessing the power of ensembles while maintaining the intrinsic interpretability of the PB models, by averaging over the model parameter manifolds. All the models were evaluated on a synthetic (publicly available dataset) in addition to detailed analyses of two real-world medical datasets (one publicly available). The models and strategies we introduce address the challenges of real-world medical data, while remaining computationally inexpensive and transparent. Moreover, they exhibit similar or superior in performance compared to alternative techniques. Sreejita Ghosh, Elizabeth Sarah Baranowski, Michael Biehl, Wiebke Arlt, Peter Tiño, Kerstin Bunte |
Neurocomputing | 5 |
| 2025 | Enzyme action optimizer: a novel bio-inspired optimization algorithmabstractThis paper presents the enzyme action optimization (EAO) algorithm, a novel bio-inspired optimization algorithm designed to simulate the adaptive enzyme mechanism in biological systems. EAO employs a novel strategy that dynamically balances between exploration and exploitation to efficiently navigate and optimize complex, multi-dimensional search spaces. EAO has been tested over diverse benchmark datasets, including the 23 classical benchmark functions, IEEE CEC2017, CEC2022 benchmark functions, where it has been compared with 14 recent and highly cited optimizers. The results show the superior performance of EAO over the compared optimizers in terms of finding the optimal solution, convergence speed, robustness, and overall performance. Furthermore, EAO was applied to solve five engineering design problems and demonstrated excellent performance results. The source code of EAO is publicly available for both MATLAB at: https://www.mathworks.com/matlabcentral/fileexchange/170296-enzyme-action-optimizer-a-novel-bio-inspired-optimization and PYTHON at: https://github.com/AliRodan/Enzyme-Action-Optimizer . Ali Rodan, Abdel Karim Al Tamimi, Loai M. Alnemer, Seyedali Mirjalili, Peter Tiño |
J. Supercomput. | 5 |
| 2024 | An Interpretable Alternative to Neural Representation Learning for Rating Prediction - Transparent Latent Class Modeling of User ReviewsabstractNowadays, neural network (NN) and deep learning (DL) techniques are widely adopted in many applications, including recommender systems. Given the sparse and stochastic nature of collaborative filtering (CF) data, recent works have critically analyzed the effective improvement of neural-based approaches compared to simpler and often transparent algorithms for recommendation. Previous results showed that NN and DL models can be outperformed by traditional algorithms in many tasks. Moreover, given the largely black-box nature of neural-based methods, interpretable results are not naturally obtained. Following on this debate, we first present a transparent probabilistic model that topologically organizes user and product latent classes based on the review information. In contrast to popular neural techniques for representation learning, we readily obtain a statistical, visualization-friendly tool that can be easily inspected to understand user and product characteristics from a textual-based perspective. Then, given the limitations of common embedding techniques, we investigate the possibility of using the estimated interpretable quantities as model input for a rating prediction task. To contribute to the recent debates, we evaluate our results in terms of both capacity for interpretability and predictive performances in comparison with popular text-based neural approaches. The results demonstrate that the proposed latent class representations can yield competitive predictive performances, compared to popular, but difficult-to-interpret approaches. Giuseppe Serra 0002, Peter Tiño, Zhao Xu 0001, Xin Yao 0001 |
IJCNN | 2 |
| 2024 | Predictive Modeling in the Reservoir Kernel Motif SpaceabstractThis work proposes a time series prediction method based on the kernel view of linear reservoirs. In particular, the time series motifs of the reservoir kernel are used as representational basis on which general readouts are constructed. We provide a geometric interpretation of our approach shedding light on how our approach is related to the core reservoir models and in what way the two approaches differ. Empirical experiments then compare predictive performances of our suggested model with those of recent state-of-art transformer based models, as well as the established recurrent network model - LSTM. The experiments are performed on both univariate and multivariate time series and with a variety of prediction horizons. Rather surprisingly we show that even when linear readout is employed, our method has the capacity to outperform transformer models on univariate time series and attain competitive results on multivariate benchmark datasets. We conclude that simple models with easily controllable capacity but capturing enough memory and subsequence structure can outperform potentially over-complicated deep learning models. This does not mean that reservoir motif based models are preferable to other more complex alternatives - rather, when introducing a new complex time series model one should employ as a sanity check simple, but potentially powerful alternatives/baselines such as reservoir models or the models introduced here. Peter Tiño, Robert Simon Fong, Roberto Fabio Leonarduzzi |
IJCNN | 1 |
| 2024 | Simple Cycle Reservoirs are UniversalabstractReservoir computation models form a subclass of recurrent neural networks with fixed non-trainable input and dynamic coupling weights. Only the static readout from the state space (reservoir) is trainable, thus avoiding the known problems with propagation of gradient information backwards through time. Reservoir models have been successfully applied in a variety of tasks and were shown to be universal approximators of time-invariant fading memory dynamic filters under various settings. Simple cycle reservoirs (SCR) have been suggested as severely restricted reservoir architecture, with equal weight ring connectivity of the reservoir units and input-to-reservoir weights of binary nature with the same absolute value. Such architectures are well suited for hardware implementations without performance degradation in many practical tasks. In this contribution, we rigorously study the expressive power of SCR in the complex domain and show that they are capable of universal approximation of any unrestricted linear reservoir system (with continuous readout) and hence any time-invariant fading memory filter over uniformly bounded input streams. Robert Simon Fong, Peter Tiño |
J. Mach. Learn. Res. | 3 |
| 2023 | An Approach for Dynamic Behavioural Prediction and Fault Injection in Cyber-Physical SystemsabstractModern technology integrates Cyber-Physical Systems (CPS), merging computational and physical processes. Ensuring CPS dependability is vital in averting adverse effects on critical applications due to unforeseen behaviour. To fortify CPS resilience, a novel technique for dynamic behavioural prediction and fault injection is introduced. It predicts dynamic CPS behaviour through system modelling under diverse operational scenarios, employing a fault model with diverse fault classes. Unlike the single model tenet, this approach engages multiple expert models to simulate both faultless and faulty behaviours. By adopting this approach, we can inject specialised faults and scale the analysis of the faults together or separately. Injecting faults assesses system reactions and reveals vulnerabilities. Tested on a water tank system, the approach proves effective in behaviour prediction and proactive fault handling, enhancing CPS design for robust, secure, and fault-tolerant systems. Hayatullahi Bolaji Adeyemo, Rami Bahsoon, Peter Tiño |
BDCAT | 3 |
| 2023 | Generalized Learning Vector Quantization With Log-Euclidean Metric Learning on Symmetric Positive-Definite ManifoldabstractIn many classification scenarios, the data to be analyzed can be naturally represented as points living on the curved Riemannian manifold of symmetric positive-definite (SPD) matrices. Due to its non-Euclidean geometry, usual Euclidean learning algorithms may deliver poor performance on such data. We propose a principled reformulation of the successful Euclidean generalized learning vector quantization (GLVQ) methodology to deal with such data, accounting for the nonlinear Riemannian geometry of the manifold through log-Euclidean metric (LEM). We first generalize GLVQ to the manifold of SPD matrices by exploiting the LEM-induced geodesic distance (GLVQ-LEM). We then extend GLVQ-LEM with metric learning. In particular, we study both 1) a more straightforward implementation of the metric learning idea by adapting metric in the space of vectorized log-transformed SPD matrices and 2) the full formulation of metric learning without matrix vectorization, thus preserving the second-order tensor structure. To obtain the distance metric in the full LEM learning (LEML) approaches, two algorithms are proposed. One method is to restrict the distance metric to be full rank, treating the distance metric tensor as an SPD matrix, and readily use the LEM framework (GLVQ-LEML-LEM). The other method is to cast no such restriction, treating the distance metric tensor as a fixed rank positive semidefinite matrix living on a quotient manifold with total space equipped with flat geometry (GLVQ-LEML-FM). Experiments on multiple datasets of different natures demonstrate the good performance of the proposed methods. Fengzhen Tang, Peter Tiño |
IEEE Trans. Cybern. | 2 |
| 2023 | LAAT: Locally Aligned Ant Technique for Discovering Multiple Faint Low Dimensional Structures of Varying DensityabstractDimensionality reduction and clustering are often used as preliminary steps for many complex machine learning tasks. The presence of noise and outliers can deteriorate the performance of such preprocessing and therefore impair the subsequent analysis tremendously. In manifold learning, several studies indicate solutions for removing background noise or noise close to the structure when the density is substantially higher than that exhibited by the noise. However, in many applications, including astronomical datasets, the density varies alongside manifolds that are buried in a noisy background. We propose a novel method to extract manifolds in the presence of noise based on the idea of Ant colony optimization. In contrast to the existing random walk solutions, our technique captures points that are locally aligned with major directions of the manifold. Moreover, we empirically show that the biologically inspired formulation of ant pheromone reinforces this behavior enabling it to recover multiple manifolds embedded in extremely noisy data clouds. The algorithm performance in comparison to state-of-the-art approaches for noise reduction in manifold detection and clustering is demonstrated, on several synthetic and real datasets, including an N-body simulation of a cosmological volume. Abolfazl Taghribi, Kerstin Bunte, Rory Smith, Michele Mastropietro, Reynier Peletier, Peter Tiño |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2023 | Hierarchical Reduced-Space Drift Detection Framework for Multivariate Supervised Data StreamsabstractIn a streaming environment, the characteristics of the data themselves and their relationship with the labels are likely to experience changes as time goes on. Most drift detection methods for supervised data streams are performance-based, that is, they detect changes only after the classication accuracy deteriorates. This may not be sufcient in many application areas where the reason behind a drift is also important. Another category of drift detectors are data distribution-based detectors. Although they can detect some drifts within the input space, changes affecting only the labelling mechanism cannot be identied. Furthermore, little work is available on drift detection for high-dimensional supervised data streams. In this paper we propose an advanced Hierarchical Reduced-space Drift Detection Framework for Supervised Data Streams (HRDS) which captures drifts regardless of their effects on classication performance. This framework suggests monitoring both marginal and class-conditional distributions within a lower-dimensional space specically relevant to the assigned classication task. Experimental comparisons have demonstrated that the proposed HRDS not only achieves high-quality performance on high-dimensional data streams, but also outperforms its competitors in terms of detection recall, precision and F-measure across a wide range of different concept drift types including subtle drifts. Peter Tiño, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Surrogate-based Digital Twin for Predictive Fault Modelling and Testing of Cyber Physical SystemsabstractCyber Physical Systems (CPS) pose a pressing need to ensure they are sufficiently reliable and continue to be dependable. It is, therefore, essential to test these systems to uncover any potential anomalies, which if not detected can lead to failure and/or cause loss or injury. Adequate or complete coverage of behaviours can be difficult to accomplish in CPS. We advocate a less expensive and easy-to-evaluate representation of the system via surrogate modelling. In this paper, we present a novel predictive fault modelling framework leveraging surrogate-based Digital Twin for probing for likely faults that can support software analysts and testers of CPS in their testing plans. The approach abstracts the CPS and uses a variant of Recurrent Neural Network known as Long Short-Term Memory (LSTM) surrogate model for forecasting. The forecasting can help in predicting multiple behaviours of the system components and the likely faults of systems under test; observations will consequently feed into the testing plans. Both direct and iterative (i.e. one-time and multiple-time varying steps) forecasting are supported as part of the framework. We evaluate our surrogate-based Digital Twins predictive modelling approach on two CPSs namely: water distribution system and air pollution detection system. The results show that our approach performed decently in predicting multiple time steps. Hayatullahi Bolaji Adeyemo, Rami Bahsoon, Peter Tiño |
BDCAT | 3 |
| 2022 | Duplication Scheduling with Bottom-Up Top-Down Recursive Neural Network
Vahab Samandi, Peter Tiño, Rami Bahsoon |
IDEAL | 2 |
| 2022 | Spatio-Temporal Activity Recognition for Evolutionary Search Behavior PredictionabstractTraditional methods for solving problems within computer science rely mostly upon the application of handcrafted algorithms. As however manual engineering of them can be considered to be a tedious process, it is interesting to consider how far internal mechanisms can be directly learned in an end-to-end manner instead. This is especially tempting to consider for metaheuristic and evolutionary optimization routines which inherently rely upon creating abundant amounts of data during run-time. To implement such an approach for these types of algorithms, it effectively requires a pipeline to first acquire deran-domized algorithm components in a domain-dependent manner and secondly a mapping to select them based upon characteristic features which unveil the black box character of an optimization problem. While in principle, within our prior work we proposed methods for extracting spatial features from metadata, these unfortunately fail to acknowledge the time-dependent nature of it. Thus, fail in scenarios when the inputs generated from initial iterations are not expressive enough. For this reason we specifically develop within this work architectures for spatio-temporal data processing. Particularly, we find that our proposed GCN-GRU and LSTM architectures, which take inspiration from CNN-LSTMs originally proposed for activity recognition in multimedia data-streams, demonstrate high efficiency and most consistent performance on time series of variable length. Further, we can also demonstrate that the class activation map (CAM) for interpretable learning with time series data helps to understand and reflects problem-dependent properties of the search behavior of an optimization algorithm. Stephen Friess, Peter Tiño, Stefan Menzel, Zhao Xu 0001, Bernhard Sendhoff, Xin Yao 0001 |
IJCNN | 2 |
| 2022 | Probabilistic modelling of general noisy multi-manifold data setsabstractThe intrinsic nature of noisy and complex data sets is often concealed in low-dimensional structures embedded in a higher dimensional space. Number of methodologies have been developed to extract and represent such structures in the form of manifolds (i.e. geometric structures that locally resemble continuously deformable intervals of Rj1). Usually a-priori knowledge of the manifold's intrinsic dimensionality is required. Additionally, their performance can often be hampered by the presence of a significant high-dimensional noise aligned along the low-dimensional core manifold. In real-world applications, the data can contain several low-dimensional structures of different dimensionalities. We propose a framework for dimensionality estimation and reconstruction of multiple noisy manifolds embedded in a noisy environment. To the best of our knowledge, this work represents the first attempt at detection and modelling of a set of coexisting general noisy manifolds by uniting two aspects of multi-manifold learning: the recovery and approximation of core noiseless manifolds and the construction of their probabilistic models. The easy-to-understand hyper-parameters can be manipulated to obtain an emerging picture of the multi-manifold structure of the data. We demonstrate the workings of the framework on two synthetic data sets, presenting challenging features for state-of-the-art techniques in Multi-Manifold learning. The first data set consists of multiple sampled noisy manifolds of different intrinsic dimensionalities, such as Möbius strip, toroid and spiral arm. The second one is a topologically complex set of three interlocked toroids. Given the absence of such unified methodologies in the literature, the comparison with existing techniques is organized along the two separate aspects of our approach mentioned above, namely manifold approximation and probabilistic modelling. The framework is then applied to a complex data set containing simulated gas volume particles from a particle simulation of a dwarf galaxy interacting with its host galaxy cluster. Detailed analysis of the recovered 1D and 2D manifolds can help us to understand the nature of Star Formation in such complex systems. Marco Canducci, Peter Tiño, Michele Mastropietro |
Artif. Intell. | 2 |
| 2022 | ASAP - A sub-sampling approach for preserving topological structures modeled with geodesic topographic mappingabstractTopological data analysis tools enjoy increasing popularity in a wide range of applications, such as Computer graphics, Image analysis, Machine learning, and Astronomy for extracting information. However, due to computational complexity, processing large numbers of samples of higher dimensionality quickly becomes infeasible. This contribution is twofold: We present an efficient novel sub-sampling strategy inspired by Coulomb’s law to decrease the number of data points in d-dimensional point clouds while preserving its homology. The method is not only capable of reducing the memory and computation time needed for the construction of different types of simplicial complexes but also preserves the size of the voids in d-dimensions, which is crucial e.g. for astronomical applications. Furthermore, we propose a technique to construct a probabilistic description of the border of significant cycles and cavities inside the point cloud. We demonstrate and empirically compare the strategy in several synthetic scenarios and an astronomical particle simulation of a dwarf galaxy for the detection of superbubbles (supernova signatures). Abolfazl Taghribi, Marco Canducci, Michele Mastropietro, Sven De Rijcke, Kerstin Bunte, Peter Tiño |
Neurocomputing | 6 |
| 2022 | Manifold Alignment Aware Ants: A Markovian Process for Manifold ExtractionabstractThe presence of manifolds is a common assumption in many applications, including astronomy and computer vision. For instance, in astronomy, low-dimensional stellar structures, such as streams, shells, and globular clusters, can be found in the neighborhood of big galaxies such as the Milky Way. Since these structures are often buried in very large data sets, an algorithm, which can not only recover the manifold but also remove the background noise (or outliers), is highly desirable. While other works try to recover manifolds either by pushing all points toward manifolds or by downsampling from dense regions, aiming to solve one of the problems, they generally fail to suppress the noise on manifolds and remove background noise simultaneously. Inspired by the collective behavior of biological ants in food-seeking process, we propose a new algorithm that employs several random walkers equipped with a local alignment measure to detect and denoise manifolds. During the walking process, the agents release pheromone on data points, which reinforces future movements. Over time the pheromone concentrates on the manifolds, while it fades in the background noise due to an evaporation procedure. We use the Markov chain (MC) framework to provide a theoretical analysis of the convergence of the algorithm and its performance. Moreover, an empirical analysis, based on synthetic and real-world data sets, is provided to demonstrate its applicability in different areas, such as improving the performance of t-distributed stochastic neighbor embedding (t-SNE) and spectral clustering using the underlying MC formulas, recovering astronomical low-dimensional structures, and improving the performance of the fast Parzen window density estimator. Mohammad Mohammadi 0004, Peter Tiño, Kerstin Bunte |
Neural Comput. | 2 |
| 2022 | Input-to-State Representation in Linear Reservoirs DynamicsabstractReservoir computing is a popular approach to design recurrent neural networks, due to its training simplicity and approximation performance. The recurrent part of these networks is not trained (e.g., via gradient descent), making them appealing for analytical studies by a large community of researchers with backgrounds spanning from dynamical systems to neuroscience. However, even in the simple linear case, the working principle of these networks is not fully understood and their design is usually driven by heuristics. A novel analysis of the dynamics of such networks is proposed, which allows the investigator to express the state evolution using the controllability matrix. Such a matrix encodes salient characteristics of the network dynamics; in particular, its rank represents an input-independent measure of the memory capacity of the network. Using the proposed approach, it is possible to compare different reservoir architectures and explain why a cyclic topology achieves favorable results as verified by practitioners. Pietro Verzelli, Cesare Alippi, Lorenzo Livi, Peter Tiño |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Attaining Meta-self-awareness through Assessment of Quality-of-KnowledgeabstractSelf-awareness is a crucial capability of autonomous service-based systems that enables them to self-adapt. There are different types of self-awareness whereby certain types of knowledge are captured at various levels. We argue that effective management of the trade-offs of dependability requirements can be achieved through “seamless” switching between different levels of awareness. However, the assessment of the quality of knowledge to enable dynamic switching between self-awareness levels has not been tackled yet. We propose a general architecture that exploits symbiotic simulation in order to tackle the complexity of assessing the quality of knowledge and attaining the meta-self-awareness property, wherein the system can reflect on its different levels of awareness. We conduct a thorough real-world study in the context of volunteer services. We conclude that a system made meta-self-aware using our approach achieves optimal performance by activating the most suitable awareness level. This comes at the cost of a modest computational overhead. Abdessalam Elhabbash, Rami Bahsoon, Peter Tiño, Peter R. Lewis 0001, Yehia El-khatib |
ICWS | 3 |
| 2021 | Tracking the Temporal-Evolution of Supernova Bubbles in Numerical Simulations
Marco Canducci, Abolfazl Taghribi, Michele Mastropietro, Sven De Rijcke, Reynier Peletier, Kerstin Bunte, Peter Tiño |
IDEAL | 7 |
| 2021 | SOMiMS - Topographic Mapping in the Model Space
Eder Zavala, Krasimira Tsaneva-Atanasova, Thomas Upton, Georgina Russell, Peter Tiño |
IDEAL | 7 |
| 2021 | Artificial Neural Networks as Feature Extractors in Continuous Evolutionary OptimizationabstractRecent years have seen the advancement of data-driven paradigms in population-based and evolutionary optimization. This reflects on one hand the mere abundance of available data, but on the other hand also progresses in the refinement of previously available machine learning methods. Surprisingly, deep pattern recognition methods emerging from the studies of neural networks have only been sparingly applied. This comes unexpected, as the complex data generated by evolutionary search algorithms can be considered tedious and intractable for manual analysis with mere practical intuitions. In this work, we therefore explore opportunities to employ deep networks to directly learn problem characteristics of continuous optimization problems. Particularly, with data obtained during initial runs of an optimization algorithm. We find that a graph neural network, trained upon a graph representation of continuous search spaces, shows in comparison to more traditional approaches higher validation accuracy and retrieves characteristics within the latent space which are better at distinguishing different continuous optimization problems. We hope that our study can pave the way towards new approaches which allow us to learn problem-dependent algorithm components and recall these from predictions of inputs generated during the run-time of an optimization algorithm. Stephen Friess, Peter Tiño, Zhao Xu 0001, Stefan Menzel, Bernhard Sendhoff, Xin Yao 0001 |
IJCNN | 2 |
| 2021 | Interpreting Node Embedding with Text-labeled GraphsabstractGraph neural networks have recently received increasing attention. These methods often map nodes into latent spaces and learn vector representations of the nodes for a variety of downstream tasks. To gain trust and to promote collaboration between AIs and humans, it would be better if those representations were interpretable for humans. However, most explainable AIs focus on a supervised learning setting and aim to answer the following question: “Why does the model predict y for an input x?”. For an unsupervised learning setting as node embedding, interpretation can be more complicated since the embedding vectors are usually not understandable for humans. On the other hand, nodes and edges in a graph are often associated with texts in many real-world applications. A question naturally arises: could we integrate the human-understandable textural data into graph learning to facilitate interpretable node embedding? In this paper we present interpretable graph neural networks (iGNN), a model to learn textual explanations for node representations modeling the extra information contained in the associated textual data. To validate the performance of the proposed method, we investigate the learned interpretability of the embedding vectors and use functional interpretability to measure it. Experimental results on multiple text-labeled graphs show the effectiveness of the iGNN model on learning textual explanations of node embedding while performing well in downstream tasks. Giuseppe Serra 0002, Zhao Xu 0001, Mathias Niepert, Carolin Lawrence, Peter Tiño, Xin Yao 0001 |
IJCNN | 5 |
| 2021 | Label-Assisted Memory Autoencoder for Unsupervised Out-of-Distribution Detection
Chao Pan 0005, Liyan Song, Ke Pei, Peter Tiño, Xin Yao 0001 |
ECML/PKDD (3) | 7 |
| 2021 | Designing Robust Models for Behaviour Prediction Using Sparse Data from Mobile Sensing: A Case Study of Office Workers' Availability for Well-being InterventionsabstractUnderstanding in which circumstances office workers take rest breaks is important for delivering effective mobile notifications and make inferences about their daily lifestyle, e.g., whether they are active and/or have a sedentary life. Previous studies designed for office workers show the effectiveness of rest breaks for preventing work-related conditions. In this article, we propose a hybrid personalised model involving a kernel density estimation model and a generalised linear mixed model to model office workers’ available moments for rest breaks during working hours. We adopt the experience-based sampling method through which we collected office workers’ responses regarding their availability through a mobile application with contextual information extracted by means of the mobile phone sensors. The experiment lasted 10 workdays and involved 19 office workers with a total of 528 responses. Our results show that time, location, ringer mode, and activity are effective features for predicting office workers’ availability. Our method can address sparse sample issues for building individual predictive behavioural models based on limited and unbalanced data. In particular, the proposed method can be considered as a potential solution to the “cold-start problem,” i.e., the negative impact of the lack of individual data when a new application is installed. Seyma Kucukozer Cavdar, Tugba Taskaya-Temizel, Abhinav Mehrotra, Mirco Musolesi, Peter Tiño |
ACM Trans. Comput. Heal. | 5 |
| 2021 | Probabilistic learning vector quantization on manifold of symmetric positive definite matricesabstractIn this paper, we develop a new classification method for manifold-valued data in the framework of probabilistic learning vector quantization. In many classification scenarios, the data can be naturally represented by symmetric positive definite matrices, which are inherently points that live on a curved Riemannian manifold. Due to the non-Euclidean geometry of Riemannian manifolds, traditional Euclidean machine learning algorithms yield poor results on such data. In this paper, we generalize the probabilistic learning vector quantization algorithm for data points living on the manifold of symmetric positive definite matrices equipped with Riemannian natural metric (affine-invariant metric). By exploiting the induced Riemannian distance, we derive the probabilistic learning Riemannian space quantization algorithm, obtaining the learning rule through Riemannian gradient descent. Empirical investigations on synthetic data, image data , and motor imagery electroencephalogram (EEG) data demonstrate the superior performance of the proposed method. Fengzhen Tang, Haifeng Feng, Peter Tiño, Bailu Si, Daxiong Ji |
Neural Networks | 3 |
| 2021 | Generalized Learning Riemannian Space Quantization: A Case Study on Riemannian Manifold of SPD MatricesabstractLearning vector quantization (LVQ) is a simple and efficient classification method, enjoying great popularity. However, in many classification scenarios, such as electroencephalogram (EEG) classification, the input features are represented by symmetric positive-definite (SPD) matrices that live in a curved manifold rather than vectors that live in the flat Euclidean space. In this article, we propose a new classification method for data points that live in the curved Riemannian manifolds in the framework of LVQ. The proposed method alters generalized LVQ (GLVQ) with the Euclidean distance to the one operating under the appropriate Riemannian metric. We instantiate the proposed method for the Riemannian manifold of SPD matrices equipped with the Riemannian natural metric. Empirical investigations on synthetic data and real-world motor imagery EEG data demonstrate that the performance of the proposed generalized learning Riemannian space quantization can significantly outperform the Euclidean GLVQ, generalized relevance LVQ (GRLVQ), and generalized matrix LVQ (GMLVQ). The proposed method also shows competitive performance to the state-of-the-art methods on the EEG classification of motor imagery tasks. Fengzhen Tang, Mengling Fan, Peter Tiño |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Representing Experience in Continuous Evolutionary optimisation through Problem-tailored Search OperatorsabstractEvolutionary algorithms are a class of population-based meta-heuristic methods partially inspired by natural evolution. Specifically, they rely on stochastic variation and selection processes to sequentially find optimal solutions of a function of interest. We attempt in this work to extract preferences in these stochastic evolutionary operators in form of empirical and improved distributions as basis for model-based mutation operators. The latter can be considered as representing problem-tailored search operators which exist independently from the optimisation run and thus can be transferred to similar problem instances. This offline approach is different to existing model-based optimisation techniques, e.g. EDA's, CMA-ES and Bayesian approaches, where adaption happens rather in an online manner without the influence of prior experience. Our approach can be rather considered to follow the recent line of research on knowledge transfer in optimisation, which until now heavily relies upon the transfer of candidate solutions across different optimisation tasks. We investigate in this paper the interplay between algorithm and optimisation task, its influence on the retrieved distributions and explore whether or not these can lead to performance improvements on a selected range of problems, as well as when transferring them across problems. At last, we make a comparison of built distributions in the hope of relating similarity in statistical distances between distributions to possible performance gains. Stephen Friess, Peter Tiño, Stefan Menzel, Bernhard Sendhoff, Xin Yao 0001 |
CEC | 2 |
| 2020 | ASAP - A Sub-sampling Approach for Preserving Topological Structures
Abolfazl Taghribi, Kerstin Bunte, Michele Mastropietro, Sven De Rijcke, Peter Tiño |
ESANN | 5 |
| 2020 | Visualisation and knowledge discovery from interpretable modelsabstractIncreasing number of sectors which affect human lives, are using Machine Learning (ML) tools. Hence the need for understanding their working mechanism and evaluating their fairness in decision-making, are becoming paramount, ushering in the era of Explainable AI (XAI). In this contribution we introduced a few intrinsically interpretable models which are also capable of dealing with missing values, in addition to extracting knowledge from the dataset and about the problem. These models are also capable of visualisation of the classifier and decision boundaries: they are the angle based variants of Learning Vector Quantization. We have demonstrated the algorithms on a synthetic dataset and a real-world one (heart disease dataset from the UCI repository). The newly developed classifiers helped in investigating the complexities of the UCI dataset as a multiclass problem. The performance of the developed classifiers were comparable to those reported in literature for this dataset, with additional value of interpretability, when the dataset was treated as a binary class problem. Sreejita Ghosh, Peter Tiño, Kerstin Bunte |
IJCNN | 2 |
| 2020 | Improving Sampling in Evolution Strategies Through Mixture-Based Distributions Built from Past Problem InstancesabstractThe notion of learning from different problem instances, although an old and known one, has in recent years regained popularity within the optimization community. Notable endeavors have been drawing inspiration from machine learning methods as a means for algorithm selection and solution transfer. However, surprisingly approaches which are centered around internal sampling models have not been revisited. Even though notable algorithms have been established in the last decades. In this work, we progress along this direction by investigating a method that allows us to learn an evolutionary search strategy reflecting rough characteristics of a fitness landscape. This latter model of a search strategy is represented through a flexible mixture-based distribution, which can subsequently be transferred and adapted for similar problems of interest. We validate this approach in two series of experiments in which we first demonstrate the efficacy of the recovered distributions and subsequently investigate the transfer with a systematic from the literature to generate benchmarking scenarios. Stephen Friess, Peter Tiño, Stefan Menzel, Bernhard Sendhoff, Xin Yao 0001 |
PPSN (1) | 2 |
| 2020 | Feature relevance determination for ordinal regression in the context of feature redundancies and privileged information
Lukas Pfannschmidt, Jonathan Jakob, Fabian Hinder, Michael Biehl, Peter Tiño, Barbara Hammer |
Neurocomputing | 5 |
| 2020 | Dynamical Systems as Temporal Feature SpacesabstractParametrised state space models in the form of recurrent networks are often used in machine learning to learn from data streams exhibiting temporal dependencies. To break the black box nature of such models it is important to understand the dynamical features of the input-driving time series that are formed in the state space. We propose a framework for rigorous analysis of such state representations in vanishing memory state space models such as echo state networks (ESN). In particular, we consider the state space a temporal feature space and the readout mapping from the state space a kernel machine operating in that feature space. We show that: (1) The usual ESN strategy of randomly generating input-to-state, as well as state coupling leads to shallow memory time series representations, corresponding to cross-correlation operator with fast exponentially decaying coefficients; (2) Imposing symmetry on dynamic coupling yields a constrained dynamic kernel matching the input time series with straightforward exponentially decaying motifs or exponentially decaying motifs of the highest frequency; (3) Simple ring (cycle) high-dimensional reservoir topology specified only through two free parameters can implement deep memory dynamic kernels with a rich variety of matching motifs. We quantify richness of feature representations imposed by dynamic kernels and demonstrate that for dynamic kernel associated with cycle reservoir topology, the kernel richness undergoes a phase transition close to the edge of stability. Peter Tiño |
J. Mach. Learn. Res. | 1 |
| 2020 | Sparsification of core set models in non-metric supervised learning
Frank-Michael Schleif, Christoph Raab, Peter Tiño |
Pattern Recognit. Lett. | 3 |
| 2019 | Exploiting Synthetically Generated Data with Semi-Supervised Learning for Small and Imbalanced DatasetsabstractData augmentation is rapidly gaining attention in machine learning. Synthetic data can be generated by simple transformations or through the data distribution. In the latter case, the main challenge is to estimate the label associated to new synthetic patterns. This paper studies the effect of generating synthetic data by convex combination of patterns and the use of these as unsupervised information in a semi-supervised learning framework with support vector machines, avoiding thus the need to label synthetic examples. We perform experiments on a total of 53 binary classification datasets. Our results show that this type of data over-sampling supports the well-known cluster assumption in semi-supervised learning, showing outstanding results for small high-dimensional datasets and imbalanced learning problems. María Pérez-Ortiz 0001, Peter Tiño, Rafal Mantiuk, César Hervás-Martínez |
AAAI | 2 |
| 2019 | Feature relevance bounds for ordinal regression
Lukas Pfannschmidt, Jonathan Jakob, Michael Biehl, Peter Tiño, Barbara Hammer |
ESANN | 4 |
| 2019 | Coevolutionary systems and PageRank
Siang Yew Chong, Peter Tiño, Jun He 0004 |
Artif. Intell. | 2 |
| 2019 | A New Framework for Analysis of Coevolutionary Systems - Directed Graph Representation and Random WalksabstractStudying coevolutionary systems in the context of simplified models (i.e., games with pairwise interactions between coevolving solutions modeled as self plays) remains an open challenge since the rich underlying structures associated with pairwise-comparison-based fitness measures are often not taken fully into account. Although cyclic dynamics have been demonstrated in several contexts (such as intransitivity in coevolutionary problems), there is no complete characterization of cycle structures and their effects on coevolutionary search. We develop a new framework to address this issue. At the core of our approach is the directed graph (digraph) representation of coevolutionary problems that fully captures structures in the relations between candidate solutions. Coevolutionary processes are modeled as a specific type of Markov chains-random walks on digraphs. Using this framework, we show that coevolutionary problems admit a qualitative characterization: a coevolutionary problem is either solvable (there is a subset of solutions that dominates the remaining candidate solutions) or not. This has an implication on coevolutionary search. We further develop our framework that provides the means to construct quantitative tools for analysis of coevolutionary processes and demonstrate their applications through case studies. We show that coevolution of solvable problems corresponds to an absorbing Markov chain for which we can compute the expected hitting time of the absorbing class. Otherwise, coevolution will cycle indefinitely and the quantity of interest will be the limiting invariant distribution of the Markov chain. We also provide an index for characterizing complexity in coevolutionary problems and show how they can be generated in a controlled manner. Siang Yew Chong, Peter Tiño, Jun He 0004, Xin Yao 0001 |
Evol. Comput. | 2 |
| 2019 | Self-awareness in Software Engineering: A Systematic Literature ReviewabstractBackground : Self-awareness has been recently receiving attention in computing systems for enriching autonomous software systems operating in dynamic environments. Objective : We aim to investigate the adoption of computational self-awareness concepts in autonomic software systems and motivate future research directions on self-awareness and related problems. Method : We conducted a systemic literature review to compile the studies related to the adoption of self-awareness in software engineering and explore how self-awareness is engineered and incorporated in software systems. From 865 studies, 74 studies have been selected as primary studies. We have analysed the studies from multiple perspectives, such as motivation, inspiration, and engineering approaches, among others. Results : Results have shown that self-awareness has been used to enable self-adaptation in systems that exhibit uncertain and dynamic behaviour. Though there have been recent attempts to define and engineer self-awareness in software engineering, there is no consensus on the definition of self-awareness. Also, the distinction between self-aware and self-adaptive systems has not been systematically treated. Conclusions : Our survey reveals that self-awareness for software systems is still a formative field and that there is growing attention to incorporate self-awareness for better reasoning about the adaptation decision in autonomic systems. Many pending issues and open problems outline possible research directions. Abdessalam Elhabbash, Maria Salama, Rami Bahsoon, Peter Tiño |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2018 | Machine learning and data analysis in astroinformatics
Michael Biehl, Kerstin Bunte, Giuseppe Longo, Peter Tiño |
ESANN | 4 |
| 2018 | Randomized Recurrent Neural Networks
Claudio Gallicchio, Alessio Micheli, Peter Tiño |
ESANN | 3 |
| 2018 | A mixture of experts model for predicting persistent weather patternsabstractWeather and atmospheric patterns are often persistent. The simplest weather forecasting method is the so-called persistence model, which assumes that the future state of a system will be similar (or equal) to the present state. Machine learning (ML) models are widely used in different weather forecasting applications, but they need to be compared to the persistence model to analyse whether they provide a competitive solution to the problem at hand. In this paper, we devise a new model for predicting low-visibility in airports using the concepts of mixture of experts. Visibility level is coded as two different ordered categorical variables: cloud height and runway visual height. The underlying system in this application is stagnant approximately in 90% of the cases, and standard ML models fail to improve on the performance of the persistence model. Because of this, instead of trying to simply beat the persistence model using ML, we use this persistence as a baseline and learn an ordinal neural network model that refines its results by focusing on learning weather fluctuations. The results show that the proposal outperforms persistence and other ordinal autoregressive models, especially for longer time horizon predictions and for the runway visual height variable. María Pérez-Ortiz 0001, Pedro Antonio Gutiérrez, Peter Tiño, Carlos Casanova-Mateo, Sancho Salcedo-Sanz |
IJCNN | 3 |
| 2018 | Supervised low rank indefinite kernel approximation using minimum enclosing balls
Frank-Michael Schleif, Andrej Gisbrecht, Peter Tiño |
Neurocomputing | 3 |
| 2018 | Asymptotic Fisher memory of randomized linear symmetric Echo State Networks
Peter Tiño |
Neurocomputing | 1 |
| 2017 | Comparison of strategies to learn from imbalanced classes for computer aided diagnosis of inborn steroidogenic disorders
Sreejita Ghosh, Elizabeth Sarah Baranowski, Rick van Veen, Gert-Jan de Vries, Michael Biehl, Wiebke Arlt, Peter Tiño, Kerstin Bunte |
ESANN | 7 |
| 2017 | Fisher memory of linear Wigner echo state networks
Peter Tiño |
ESANN | 1 |
| 2017 | Self-Awareness for Dynamic Knowledge Management in Self-Adaptive Volunteer ServicesabstractEngineering volunteer services calls for novel self-adaptive approaches for dynamically managing the process of selecting volunteer services. As these services tend to be published and withdrawn without restrictions, uncertainties, dynamisms and 'dilution of control' related to the decisions of selection and composition are complex problems. These services tend to exhibit periodic performance patterns, which are often repeated over a certain time period. Consequently, the awareness of such periodic patterns enables the prediction of the services performance leading to better adaptation. In this paper, we contribute to a self-adaptive approach, namely time-awareness, which combines self-aware principles with dynamic histograms to dynamically manage the periodic trends of services performance and their evolution trends. Such knowledge can inform the adaptation decisions, leading to increase in the precision of selecting and composing services. We evaluate the approach using a volunteer storage composition scenario. The evaluation results show the advantages of dynamic knowledge management in self-adaptive volunteer computing in selecting dependable services and satisfying higher number of requests. Abdessalam Elhabbash, Rami Bahsoon, Peter Tiño |
ICWS | 3 |
| 2017 | Linear dynamical based models for sequential domainsabstractThe aim of the paper is to explore how models based on a linear dynamic can be used in order to perform a prediction task in sequential domains. In the literature, it has already been shown that Linear Dynamical Systems (LDSs) can be quite useful when dealing with sequence learning tasks. Our aim is to study whether it is possible to use LDSs as building blocks for constructing more complex and powerful models. Specifically, we propose a model dubbed Linear System Network, that exploits several LDSs in order to compute a nonlinear projection of the input. Moreover, we explore whether is it possible to apply a co-learning technique in order to improve the performance of LDSs for the considered prediction task. Luca Pasa, Alessandro Sperduti, Peter Tiño |
IJCNN | 3 |
| 2017 | Classification of sparsely and irregularly sampled time series: A learning in model space approachabstractClassification of sparsely and irregularly sampled time series data is a challenging machine learning task. To tackle this problem, we present a learning in model space framework in which time-continuous dynamical system models are first inferred from individual time series and then the inferred models are used to represent these time series for the classification task. In contrast to the existing approaches using model point estimates to represent individual time series, we further employ posterior distributions over models, thus taking into account in a principled manner the uncertainty around the inferred model due to observation noise and data sparsity. Finally, we present a distributional classifier for classifying the posterior distributions. We evaluate the framework on a biological pathway model. In particular, we investigate the classification performance in the cases where model uncertainties in the training and test phases do not match. Peter Tiño, Krasimira Tsaneva-Atanasova |
IJCNN | 2 |
| 2017 | Probabilistic matching: Causal inference under measurement errorsabstractThe abundance of data produced daily from large variety of sources has boosted the need of novel approaches on causal inference analysis from observational data. Observational data often contain noisy or missing entries. Moreover, causal inference studies may require unobserved high-level information which needs to be inferred from other observed attributes. In such cases, inaccuracies of the applied inference methods will result in noisy outputs. In this study, we propose a novel approach for causal inference when one or more key variables are noisy. Our method utilizes the knowledge about the uncertainty of the real values of key variables in order to reduce the bias induced by noisy measurements. We evaluate our approach in comparison with existing methods both on simulated and real scenarios and we demonstrate that our method reduces the bias and avoids false causal inference conclusions in most cases. Fani Tsapeli, Peter Tiño, Mirco Musolesi |
IJCNN | 2 |
| 2017 | Ordinal regression based on learning vector quantization
Fengzhen Tang, Peter Tiño |
Neural Networks | 2 |
| 2017 | Indefinite Core Vector Machine
Frank-Michael Schleif, Peter Tiño |
Pattern Recognit. | 2 |
| 2016 | Learning in indefinite proximity spaces - recent trends
Frank-Michael Schleif, Peter Tiño, Yingyu Liang |
ESANN | 2 |
| 2016 | Probabilistic Modelling for Delay Estimation in Gravitationally Lensed Photon Streams
Sultanah Al Otaibi, Peter Tiño, Somak Raychaudhury |
IDEAL | 2 |
| 2016 | Model-coupled autoencoder for time series visualisation
Nikolaos Gianniotis, Sven Dennis Kügler, Peter Tiño, Kai Lars Polsterer |
Neurocomputing | 3 |
| 2016 | Oversampling the Minority Class in the Feature SpaceabstractThe imbalanced nature of some real-world data is one of the current challenges for machine learning researchers. One common approach oversamples the minority class through convex combination of its patterns. We explore the general idea of synthetic oversampling in the feature space induced by a kernel function (as opposed to input space). If the kernel function matches the underlying problem, the classes will be linearly separable and synthetically generated patterns will lie on the minority class region. Since the feature space is not directly accessible, we use the empirical feature space (EFS) (a Euclidean space isomorphic to the feature space) for oversampling purposes. The proposed method is framed in the context of support vector machines, where the imbalanced data sets can pose a serious hindrance. The idea is investigated in three scenarios: 1) oversampling in the full and reduced-rank EFSs; 2) a kernel learning technique maximizing the data class separation to study the influence of the feature space structure (implicitly defined by the kernel function); and 3) a unified framework for preferential oversampling that spans some of the previous approaches in the literature. We support our investigation with extensive experiments over 50 imbalanced data sets. María Pérez-Ortiz 0001, Pedro Antonio Gutiérrez, Peter Tiño, César Hervás-Martínez |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Autoencoding time series for visualisation
Nikolaos Gianniotis, Sven Dennis Kügler, Peter Tiño, Kai Lars Polsterer, Ranjeev Misra |
ESANN | 3 |
| 2015 | Probabilistic Classification Vector Machine at large scale
Frank-Michael Schleif, Andrej Gisbrecht, Peter Tiño |
ESANN | 3 |
| 2015 | Automated Detection of Galaxy Groups Through Probabilistic Hough Transform
Rafee T. Ibrahem, Peter Tiño, Richard J. Pearson, Trevor J. Ponman, Arif Babul |
ICONIP (3) | 2 |
| 2015 | Self-Adaptive Volunteered Services Composition through Stimulus- and Time-AwarenessabstractVolunteered Service Composition (VSC) refers to the process of composing volunteered services and resources. These services are typically published to a pool of voluntary resources. Selection and composition decisions tend to encounter numerous uncertainties: service consumers and applications have little control of these services and tend to be uncertain about their level of support for the desired functionalities and non-functionalities. In this paper, we contribute to a self-awareness framework that implements two levels of awareness, Stimulus-awareness and Time-awareness. The former responds to basic changes in the environment while the latter takes into consideration the historical performance of the services. We have used volunteer service computing as an example to demonstrate the benefits that self-awareness can introduce to self-adaptation. We have compared the Stimulus- and Time-awareness approaches with a recent Ranking approach from the literature. The results show that the Time-awareness level has the advantage of satisfying higher number of requests with lower time cost. Abdessalam Elhabbash, Rami Bahsoon, Peter Tiño, Peter R. Lewis 0001 |
ICWS | 3 |
| 2015 | Model Metric Co-Learning for Time Series Classification
Huanhuan Chen 0001, Fengzhen Tang, Peter Tiño, Anthony G. Cohn 0001, Xin Yao 0001 |
IJCAI | 3 |
| 2015 | Incremental probabilistic classification vector machine with linear costsabstractThe probabilistic classification vector machine is a very effective and generic probabilistic and sparse classifier. A recently published incremental version improved the runtime complexity to quadratic costs. We derive the Nyström approximation for asymmetric matrices to obtain linear runtime and memory complexity for the incremental probabilistic classification vector machine while keeping similar prediction performance. Frank-Michael Schleif, Peter Tiño |
IJCNN | 3 |
| 2015 | Indefinite Proximity Learning: A ReviewabstractEfficient learning of a data analysis task strongly depends on the data representation. Most methods rely on (symmetric) similarity or dissimilarity representations by means of metric inner products or distances, providing easy access to powerful mathematical formalisms like kernel or branch-and-bound approaches. Similarities and dissimilarities are, however, often naturally obtained by nonmetric proximity measures that cannot easily be handled by classical learning algorithms. Major efforts have been undertaken to provide approaches that can either directly be used for such data or to make standard methods available for these types of data. We provide a comprehensive survey for the field of learning with nonmetric proximities. First, we introduce the formalism used in nonmetric spaces and motivate specific treatments for nonmetric proximity data. Second, we provide a systematization of the various approaches. For each category of approaches, we provide a comparative discussion of the individual algorithms and address complexity issues and generalization properties. In a summarizing section, we provide a larger experimental study for the majority of the algorithms on standard data sets. We also address the problem of large-scale proximity learning, which is often overlooked in this context and of major importance to make the method relevant in practice. The algorithms we discuss are in general applicable for proximity-based clustering, one-class classification, classification, regression, and embedding approaches. In the experimental part, we focus on classification tasks. Frank-Michael Schleif, Peter Tiño |
Neural Comput. | 2 |
| 2015 | The Benefits of Modeling Slack Variables in SVMsabstractIn this letter, we explore the idea of modeling slack variables in support vector machine (SVM) approaches. The study is motivated by SVM+, which models the slacks through a smooth correcting function that is determined by additional (privileged) information about the training examples not available in the test phase. We take a closer look at the meaning and consequences of smooth modeling of slacks, as opposed to determining them in an unconstrained manner through the SVM optimization program. To better understand this difference we only allow the determination and modeling of slack values on the same information--that is, using the same training input in the original input space. We also explore whether it is possible to improve classification performance by combining (in a convex combination) the original SVM slacks with the modeled ones. We show experimentally that this approach not only leads to improved generalization performance but also yields more compact, lower-complexity models. Finally, we extend this idea to the context of ordinal regression, where a natural order among the classes exists. The experimental results confirm principal findings from the binary case. Fengzhen Tang, Peter Tiño, Pedro Antonio Gutiérrez, Huanhuan Chen 0001 |
Neural Comput. | 2 |
| 2014 | Recent trends in learning of structured and non-standard data
Frank-Michael Schleif, Peter Tiño, Thomas Villmann |
ESANN | 2 |
| 2014 | Support Vector Ordinal Regression using Privileged Information
Fengzhen Tang, Peter Tiño, Pedro Antonio Gutiérrez, Huanhuan Chen 0001 |
ESANN | 2 |
| 2014 | Learning the deterministically constructed Echo State NetworksabstractEcho State Networks (ESNs) have shown great promise in the applications of non-linear time series processing because of their powerful computational ability and efficient training strategy. However, the nature of randomization in the structure of the reservoir causes it be poorly understood and leaves room for further improvements for specific problems. A deterministically constructed reservoir model, Cycle Reservoir with Jumps (CRJ), shows superior generalization performance to standard ESN. However, the weights that govern the structure of the reservoir (reservoir weights) in CRJ model are obtained through exhaustive grid search which is very computational intensive. In this paper, we propose to learn the reservoir weights together with the linear readout weights using a hybrid optimization strategy. The reservoir weights are trained through nonlinear optimization techniques while the linear readout weights are obtained through linear algorithms. The experimental results demonstrate that the proposed strategy of training the CRJ network tremendously improves the computational efficiency without jeopardizing the generalization performance, sometimes even with better generalization performance. Fengzhen Tang, Peter Tiño, Huanhuan Chen 0001 |
IJCNN | 2 |
| 2014 | Combining learning in model space fault diagnosis with data validation/reconstruction: Application to the Barcelona water network
Joseba Quevedo, Huanhuan Chen 0001, Miquel Àngel Cugueró, Peter Tiño, Vicenç Puig, Ramon Sarrate, Xin Yao 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2014 | Ordinal regression neural networks based on concentric hyperspheres
Pedro Antonio Gutiérrez, Peter Tiño, César Hervás-Martínez |
Neural Networks | 2 |
| 2014 | Dimensionality reduction and topographic mapping of binary tensors
Jakub Mazgut, Peter Tiño, Mikael Bodén, Hong Yan 0001 |
Pattern Anal. Appl. | 2 |
| 2014 | Learning in the Model Space for Cognitive Fault DiagnosisabstractThe emergence of large sensor networks has facilitated the collection of large amounts of real-time data to monitor and control complex engineering systems. However, in many cases the collected data may be incomplete or inconsistent, while the underlying environment may be time-varying or unformulated. In this paper, we develop an innovative cognitive fault diagnosis framework that tackles the above challenges. This framework investigates fault diagnosis in the model space instead of the signal space. Learning in the model space is implemented by fitting a series of models using a series of signal segments selected with a sliding window. By investigating the learning techniques in the fitted model space, faulty models can be discriminated from healthy models using a one-class learning algorithm. The framework enables us to construct a fault library when unknown faults occur, which can be regarded as cognitive fault isolation. This paper also theoretically investigates how to measure the pairwise distance between two models in the model space and incorporates the model distance into the learning algorithm in the model space. The results on three benchmark applications and one simulated model for the Barcelona water distribution network confirm the effectiveness of the proposed framework. Huanhuan Chen 0001, Peter Tiño, Ali Rodan, Xin Yao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Efficient Probabilistic Classification Vector Machine With Incremental Basis Function SelectionabstractProbabilistic classification vector machine (PCVM) is a sparse learning approach aiming to address the stability problems of relevance vector machine for classification problems. Because PCVM is based on the expectation maximization algorithm, it suffers from sensitivity to initialization, convergence to local minima, and the limitation of Bayesian estimation making only point estimates. Another disadvantage is that PCVM was not efficient for large data sets. To address these problems, this paper proposes an efficient PCVM (EPCVM) by sequentially adding or deleting basis functions according to the marginal likelihood maximization for efficient training. Because of the truncated prior used in EPCVM, two approximation techniques, i.e., Laplace approximation and expectation propagation (EP), have been used to implement EPCVM to obtain full Bayesian solutions. We have verified Laplace approximation and EP with a hybrid Monte Carlo approach. The generalization performance and computational effectiveness of EPCVM are extensively evaluated. Theoretical discussions using Rademacher complexity reveal the relationship between the sparsity and the generalization bound of EPCVM. Huanhuan Chen 0001, Peter Tiño, Xin Yao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | A principled approach to mining from noisy logs using Heuristics MinerabstractNoise is a challenge for process mining algorithms, but there is no standard definition of noise nor accepted way to quantify it. This means it is not possible to mine with confidence from event logs which may not record the underlying process correctly. We discuss one way of thinking about noise in process mining. We consider mining from a `noisy log' as learning a probability distribution over traces, representing the true process, from a log which is a sample from multiple distributions: the `true' process model and one or more `noise' models. We apply this using a probabilistic analysis of the Heuristics Miner algorithm, and demonstrate on a simple example. We show that for a given model it is possible to predict how much data is needed to mine the underlying model without the noise, and identify differences in the the robustness of Heuristics Miner to different types of noise. Philip Weber 0001, Behzad Bordbar, Peter Tiño |
CIDM | 3 |
| 2013 | Ordinal-based metric learning for learning using privileged informationabstractLearning Using privileged Information (LUPI), originally proposed in [1], is an advanced learning paradigm that aims to improve the supervised learning in the presence of additional (privileged) information, available during training, but not in the test phase. We present a novel metric learning methodology that is specially designed for incorporating privileged information in ordinal classification tasks, where there is a natural order on the set of classes. This is done by changing the global metric in the input space, based on distance relations revealed by the privileged information. The proposed model is formulated in the context of ordinal prototype based classification with metric adaptation. Unlike the existing nominal version of LUPI in prototype models [8], [9], in ordinal classifications the proposed LUPI model takes explicitly into account the class order information during the input space metric learning. Experiments demonstrate that incorporating privileged information via the proposed ordinal-based metric learning can improve the ordinal classification performance. Shereen Fouad, Peter Tiño |
IJCNN | 2 |
| 2013 | Concept drift detection for online class imbalance learningabstractConcept drift detection methods are crucial components of many online learning approaches. Accurate drift detections allow prompt reaction to drifts and help to maintain high performance of online models over time. Although many methods have been proposed, no attention has been given to data streams with imbalanced class distributions, which commonly exist in real-world applications, such as fault diagnosis of control systems and intrusion detection in computer networks. This paper studies the concept drift problem for online class imbalance learning. We look into the impact of concept drift on single-class performance of online models based on three types of classifiers, under seven different scenarios with the presence of class imbalance. The analysis reveals that detecting drift in imbalanced data streams is a more difficult task than in balanced ones. Minority-class recall suffers from a significant drop after the drift involving the minority class. Overall accuracy is not suitable for drift detection. Based on the findings, we propose a new detection method DDM-OCI derived from the existing method DDM. DDM-OCI monitors minority-class recall online to capture the drift. The results show a quick response of the online model working with DDM-OCI to the new concept. Shuo Wang 0005, Leandro L. Minku, Davide Ghezzi, Daniele Caltabiano, Peter Tiño, Xin Yao 0001 |
IJCNN | 5 |
| 2013 | Model-based kernel for efficient time series analysisabstractWe present novel, efficient, model based kernels for time series data rooted in the reservoir computation framework. The kernels are implemented by fitting reservoir models sharing the same fixed deterministically constructed state transition part to individual time series. The proposed kernels can naturally handle time series of different length without the need to specify a parametric model class for the time series. Compared with most time series kernels, our kernels are computationally efficient. We show how the model distances used in the kernel can be calculated analytically or efficiently estimated. The experimental results on synthetic and benchmark time series classification tasks confirm the efficiency of the proposed kernel in terms of both generalization accuracy and computational speed. This paper also investigates on-line reservoir kernel construction for extremely long time series. Huanhuan Chen 0001, Fengzhen Tang, Peter Tiño, Xin Yao 0001 |
KDD | 3 |
| 2013 | A Spatial Mixture Approach to Inferring Sub-ROI Spatio-temporal Patterns from Rapid Event-Related fMRI Data
Stephen D. Mayhew, Zoe Kourtzi, Peter Tiño |
MICCAI (2) | 4 |
| 2013 | Novel approaches in machine learning and computational intelligence
Alessio Micheli, Frank-Michael Schleif, Peter Tiño |
Neurocomputing | 3 |
| 2013 | Time-dependent series variance learning with recurrent mixture density networks
Nikolay I. Nikolaev, Peter Tiño, Evgueni N. Smirnov |
Neurocomputing | 2 |
| 2013 | Short term memory in input-driven linear dynamical systems
Peter Tiño, Ali Rodan |
Neurocomputing | 1 |
| 2013 | Exploitation of Pairwise Class Distances for Ordinal ClassificationabstractOrdinal classification refers to classification problems in which the classes have a natural order imposed on them because of the nature of the concept studied. Some ordinal classification approaches perform a projection from the input space to one-dimensional (latent) space that is partitioned into a sequence of intervals (one for each class). Class identity of a novel input pattern is then decided based on the interval its projection falls into. This projection is trained only indirectly as part of the overall model fitting. As with any other latent model fitting, direct construction hints one may have about the desired form of the latent model can prove very useful for obtaining high-quality models. The key idea of this letter is to construct such a projection model directly, using insights about the class distribution obtained from pairwise distance calculations. The proposed approach is extensively evaluated with 8 nominal and ordinal classifiers methods, 10 real-world ordinal classification data sets, and 4 different performance measures. The new methodology obtained the best results in average ranking when considering three of the performance metrics, although significant differences are found for only some of the methods. Also, after observing other methods of internal behavior in the latent space, we conclude that the internal projections do not fully reflect the intraclass behavior of the patterns. Our method is intrinsically simple, intuitive, and easily understandable, yet highly competitive with state-of-the-art approaches to ordinal classification. Javier Sánchez-Monedero, Pedro Antonio Gutiérrez, Peter Tiño, César Hervás-Martínez |
Neural Comput. | 3 |
| 2013 | Scaling Up Estimation of Distribution Algorithms for Continuous OptimizationabstractSince estimation of distribution algorithms (EDAs) were proposed, many attempts have been made to improve EDAs' performance in the context of global optimization. So far, the studies or applications of multivariate probabilistic model-based EDAs in continuous domain are still mostly restricted to low-dimensional problems. Traditional EDAs have difficulties in solving higher dimensional problems because of the curse of dimensionality and rapidly increasing computational costs. However, scaling up continuous EDAs for large-scale optimization is still necessary, which is supported by the distinctive feature of EDAs: because a probabilistic model is explicitly estimated, from the learned model one can discover useful properties of the problem. Besides obtaining a good solution, understanding of the problem structure can be of great benefit, especially for black box optimization. We propose a novel EDA framework with model complexity control (EDA-MCC) to scale up continuous EDAs. By employing weakly dependent variable identification and subspace modeling, EDA-MCC shows significantly better performance than traditional EDAs on high-dimensional problems. Moreover, the computational cost and the requirement of large population sizes can be reduced in EDA-MCC. In addition to being able to find a good solution, EDA-MCC can also provide useful problem structure characterizations. EDA-MCC is the first successful instance of multivariate model-based EDAs that can be effectively applied to a general class of up to 500-D problems. It also outperforms some newly developed algorithms designed specifically for large-scale optimization. In order to understand the strengths and weaknesses of EDA-MCC, we have carried out extensive computational studies. Our results have revealed when EDA-MCC is likely to outperform others and on what kind of benchmark functions. Weishan Dong, Tianshi Chen 0002, Peter Tiño, Xin Yao 0001 |
IEEE Trans. Evol. Comput. | 3 |
| 2013 | Complex Coevolutionary Dynamics - Structural Stability and Finite Population EffectsabstractUnlike evolutionary dynamics, coevolutionary dynamics can exhibit a wide variety of complex regimes. This has been confirmed by numerical studies, e.g., in the context of evolutionary game theory (EGT) and population dynamics of simple two-strategy games with various types of replication and selection mechanisms. Using the framework of shadowing lemma, we study to what degree can such infinite population dynamics: 1) be reliably simulated on finite precision computers; and 2) be trusted to represent coevolutionary dynamics of possibly very large, but finite, populations. In a simple EGT setting of two-player symmetric games with two pure strategies and a polymorphic equilibrium, we prove that for (μ,λ), truncation, sequential tournament, best-of-group tournament, and linear ranking selections, the coevolutionary dynamics do not possess the shadowing property. In other words, infinite population simulations cannot be guaranteed to represent real trajectories or to be representative of coevolutionary dynamics of potentially very large, but finite, populations. Peter Tiño, Siang Yew Chong, Xin Yao 0001 |
IEEE Trans. Evol. Comput. | 1 |
| 2013 | Incorporating Privileged Information Through Metric LearningabstractIn some pattern analysis problems, there exists expert knowledge, in addition to the original data involved in the classification process. The vast majority of existing approaches simply ignore such auxiliary (privileged) knowledge. Recently a new paradigm-learning using privileged information-was introduced in the framework of SVM+. This approach is formulated for binary classification and, as typical for many kernel-based methods, can scale unfavorably with the number of training examples. While speeding up training methods and extensions of SVM+ to multiclass problems are possible, in this paper we present a more direct novel methodology for incorporating valuable privileged knowledge in the model construction phase, primarily formulated in the framework of generalized matrix learning vector quantization. This is done by changing the global metric in the input space, based on distance relations revealed by the privileged information. Hence, unlike in SVM+, any convenient classifier can be used after such metric modification, bringing more flexibility to the problem of incorporating privileged information during the training. Experiments demonstrate that the manipulation of an input space metric based on privileged data improves classification accuracy. Moreover, our methods can achieve competitive performance against the SVM+ formulations. Shereen Fouad, Peter Tiño, Somak Raychaudhury, Petra Schneider |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | A Framework for the Analysis of Process Mining AlgorithmsabstractThere are many process mining algorithms and representations, making it difficult to choose which algorithm to use or compare results. Process mining is essentially a machine learning task, but little work has been done on systematically analyzing algorithms to understand their fundamental properties, such as how much data are needed for confidence in mining. We propose a framework for analyzing process mining algorithms. Processes are viewed as distributions over traces of activities and mining algorithms as learning these distributions. We use probabilistic automata as a unifying representation to which other representation languages can be converted. We present an analysis of the Alpha algorithm under this framework and experimental results, which show that from the substructures in a model and behavior of the algorithm, the amount of data needed for mining can be predicted. This allows efficient use of data and quantification of the confidence which can be placed in the results. Philip Weber 0001, Behzad Bordbar, Peter Tiño |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2012 | Theory of Input Driven Dynamical Systems
Manjunath Gandhi, Peter Tiño, Herbert Jaeger |
ESANN | 2 |
| 2012 | Short Term Memory Quantifications in Input-Driven Linear Dynamical Systems
Peter Tiño, Ali Rodan |
ESANN | 1 |
| 2012 | Process Mining in Non-Stationary Environments
Philip Weber 0001, Peter Tiño, Behzad Bordbar |
ESANN | 2 |
| 2012 | Learning Using Privileged Information in Prototype Based Models
Shereen Fouad, Peter Tiño, Somak Raychaudhury, Petra Schneider |
ICANN (2) | 2 |
| 2012 | Prototype Based Modelling for Ordinal Classification
Shereen Fouad, Peter Tiño |
IDEAL | 2 |
| 2012 | Still Alive: Extending Keep-Alive Intervals in P2P Overlay Networks
Richard Price, Peter Tiño, Georgios Theodoropoulos 0001 |
Mob. Networks Appl. | 2 |
| 2012 | Adaptive Metric Learning Vector Quantization for Ordinal ClassificationabstractMany pattern analysis problems require classification of examples into naturally ordered classes. In such cases, nominal classification schemes will ignore the class order relationships, which can have a detrimental effect on classification accuracy. This article introduces two novel ordinal learning vector quantization (LVQ) schemes, with metric learning, specifically designed for classifying data items into ordered classes. In ordinal LVQ, unlike in nominal LVQ, the class order information is used during training in selecting the class prototypes to be adapted, as well as in determining the exact manner in which the prototypes get updated. Prototype-based models in general are more amenable to interpretations and can often be constructed at a smaller computational cost than alternative nonlinear classification models. Experiments demonstrate that the proposed ordinal LVQ formulations compare favorably with their nominal counterparts. Moreover, our methods achieve competitive performance against existing benchmark ordinal regression models. Shereen Fouad, Peter Tiño |
Neural Comput. | 2 |
| 2012 | Simple Deterministically Constructed Cycle Reservoirs with Regular JumpsabstractA new class of state-space models, reservoir models, with a fixed state transition structure (the "reservoir") and an adaptable readout from the state space, has recently emerged as a way for time series processing and modeling. Echo state network (ESN) is one of the simplest, yet powerful, reservoir models. ESN models are generally constructed in a randomized manner. In our previous study (Rodan & Tiňo, 2011), we showed that a very simple, cyclic, deterministically generated reservoir can yield performance competitive with standard ESN. In this contribution, we extend our previous study in three aspects. First, we introduce a novel simple deterministic reservoir model, cycle reservoir with jumps (CRJ), with highly constrained weight values, that has superior performance to standard ESN on a variety of temporal tasks of different origin and characteristics. Second, we elaborate on the possible link between reservoir characterizations, such as eigenvalue distribution of the reservoir matrix or pseudo-Lyapunov exponent of the input-driven reservoir dynamics, and the model performance. It has been suggested that a uniform coverage of the unit disk by such eigenvalues can lead to superior model performance. We show that despite highly constrained eigenvalue distribution, CRJ consistently outperforms ESN (which has much more uniform eigenvalue coverage of the unit disk). Also, unlike in the case of ESN, pseudo-Lyapunov exponents of the selected optimal CRJ models are consistently negative. Third, we present a new framework for determining the short-term memory capacity of linear reservoir models to a high degree of precision. Using the framework, we study the effect of shortcut connections in the CRJ reservoir topology on its memory capacity. Ali Rodan, Peter Tiño |
Neural Comput. | 2 |
| 2012 | Improving Generalization Performance in Co-Evolutionary LearningabstractRecently, the generalization framework in co-evolutionary learning has been theoretically formulated and demonstrated in the context of game-playing. Generalization performance of a strategy (solution) is estimated using a collection of random test strategies (test cases) by taking the average game outcomes, with confidence bounds provided by Chebyshev's theorem. Chebyshev's bounds have the advantage that they hold for any distribution of game outcomes. However, such a distribution-free framework leads to unnecessarily loose confidence bounds. In this paper, we have taken advantage of the near-Gaussian nature of average game outcomes and provided tighter bounds based on parametric testing. This enables us to use small samples of test strategies to guide and improve the co-evolutionary search. We demonstrate our approach in a series of empirical studies involving the iterated prisoner's dilemma (IPD) and the more complex Othello game in a competitive co-evolutionary learning setting. The new approach is shown to improve on the classical co-evolutionary learning in that we obtain increasingly higher generalization performance using relatively small samples of test strategies. This is achieved without large performance fluctuations typical of the classical approach. The new approach also leads to faster co-evolutionary search where we can strictly control the condition (sample sizes) under which the speedup is achieved (not at the cost of weakening precision in the estimates). Siang Yew Chong, Peter Tiño, Day Chyi Ku, Xin Yao 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2011 | Negatively Correlated Echo State Networks
Ali Rodan, Peter Tiño |
ESANN | 2 |
| 2011 | Time-Dependent Series Variance Estimation via Recurrent Neural Networks
Nikolay I. Nikolaev, Peter Tiño, Evgueni N. Smirnov |
ICANN (1) | 2 |
| 2011 | A Principled Approach to the Analysis of Process Mining Algorithms
Philip Weber 0001, Behzad Bordbar, Peter Tiño |
IDEAL | 3 |
| 2011 | One-Shot Learning of Poisson Distributions in Serial Analysis of Gene Expression
Peter Tiño |
ISNN (2) | 1 |
| 2011 | Searching for Coexpressed Genes in Three-Color cDNA Microarray Data Using a Probabilistic Model-Based Hough TransformabstractThe effects of a drug on the genomic scale can be assessed in a three-color cDNA microarray with the three color intensities represented through the so-called hexaMplot. In our recent study, we have shown that the Hough Transform (HT) applied to the hexaMplot can be used to detect groups of coexpressed genes in the normal-disease-drug samples. However, the standard HT is not well suited for the purpose because 1) the assayed genes need first to be hard-partitioned into equally and differentially expressed genes, with HT ignoring possible information in the former group; 2) the hexaMplot coordinates are negatively correlated and there is no direct way of expressing this in the standard HT and 3) it is not clear how to quantify the association of coexpressed genes with the line along which they cluster. We address these deficiencies by formulating a dedicated probabilistic model-based HT. The approach is demonstrated by assessing effects of the drug Rg1 on homocysteine-treated human umbilical vein endothetial cells. Compared with our previous study, we robustly detect stronger natural groupings of coexpressed genes. Moreover, the gene groups show coherent biological functions with high significance, as detected by the Gene Ontology analysis. Peter Tiño, Hongya Zhao, Hong Yan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | Minimum Complexity Echo State NetworkabstractReservoir computing (RC) refers to a new class of state-space models with a fixed state transition structure (the reservoir) and an adaptable readout form the state space. The reservoir is supposed to be sufficiently complex so as to capture a large number of features of the input stream that can be exploited by the reservoir-to-output readout mapping. The field of RC has been growing rapidly with many successful applications. However, RC has been criticized for not being principled enough. Reservoir construction is largely driven by a series of randomized model-building stages, with both researchers and practitioners having to rely on a series of trials and errors. To initialize a systematic study of the field, we concentrate on one of the most popular classes of RC methods, namely echo state network, and ask: What is the minimal complexity of reservoir construction for obtaining competitive models and what is the memory capacity (MC) of such simplified reservoirs? On a number of widely used time series benchmarks of different origin and characteristics, as well as by conducting a theoretical analysis we show that a simple deterministically constructed cycle reservoir is comparable to the standard echo state network methodology. The (short-term) MC of linear cyclic reservoirs can be made arbitrarily close to the proved optimal value. Ali Rodan, Peter Tiño |
IEEE Trans. Neural Networks | 2 |
| 2010 | On Reliability Of Simulations Of Complex Co-Evolutionary Processes
Peter Tiño, Siang Yew Chong, Xin Yao 0001 |
ECMS | 1 |
| 2010 | Multilinear Decomposition and Topographic Mapping of Binary Tensors
Jakub Mazgut, Peter Tiño, Mikael Bodén, Hong Yan 0001 |
ICANN (1) | 2 |
| 2010 | Simple Deterministically Constructed Recurrent Neural Networks
Ali Rodan, Peter Tiño |
IDEAL | 2 |
| 2010 | Uncovering delayed patterns in noisy and irregularly sampled time series: An astronomy application
Juan Carlos Cuevas-Tello, Peter Tiño, Somak Raychaudhury, Xin Yao 0001, Markus Harva |
Pattern Recognit. | 2 |
| 2009 | Still alive: Extending keep-alive intervals in P2P overlay networksabstractNodes within existing P2P networks typically exchange periodic keep-alive messages in order to maintain network connections between neighbours. This paper investigates a number of algorithms which allow each individual connections to extend the interval between successive keep-alive messages based u Richard Price, Peter Tiño |
CollaborateCom | 2 |
| 2009 | Topographic Mapping of Astronomical Light Curves via a Physically Inspired Probabilistic Model
Nikolaos Gianniotis, Peter Tiño, Steve Spreckley, Somak Raychaudhury |
ICANN (1) | 2 |
| 2009 | Fast parzen window density estimatorabstractParzen Windows (PW) is a popular nonparametric density estimation technique. In general the smoothing kernel is placed on all available data points, which makes the algorithm computationally expensive when large datasets are considered. Several approaches have been proposed in the past to reduce the computational cost of PW either by subsampling the dataset, or by imposing a sparsity in the density model. Typically the latter requires a rather involved and complex learning process. In this paper, we propose a new simple and efficient kernel-based method for non-parametric probability density function (pdf) estimation on large datasets. We cover the entire data space by a set of fixed radii hyper-balls with densities represented by full covariance Gaussians. The accuracy and efficiency of the new estimator is verified on both synthetic dataset and large datasets of astronomical simulations of the galaxy disruption process. Experiments demonstrate that the estimation accuracy of the new estimator is comparable to that of the previous approaches but with a significant speed-up. We also show that the pdf learnt by the new estimator could used to automatically find the most matching set in large scale astronomical simulations. Peter Tiño, Mark A. Fardal, Somak Raychaudhury, Arif Babul |
IJCNN | 2 |
| 2009 | Basic properties and information theory of Audic-Claverie statistic for analyzing cDNA arraysabstractBACKGROUND: The Audic-Claverie method 1 has been and still continues to be a popular approach for detection of differentially expressed genes in the SAGE framework. The method is based on the assumption that under the null hypothesis tag counts of the same gene in two libraries come from the same but unknown Poisson distribution. The problem is that each SAGE library represents only a single measurement. We ask: Given that the tag count samples from SAGE libraries are extremely limited, how useful actually is the Audic-Claverie methodology? We rigorously analyze the A-C statistic that forms a backbone of the methodology and represents our knowledge of the underlying tag generating process based on one observation. RESULTS: We show that the A-C statistic and the underlying Poisson distribution of the tag counts share the same mode structure. Moreover, the K-L divergence from the true unknown Poisson distribution to the A-C statistic is minimized when the A-C statistic is conditioned on the mode of the Poisson distribution. Most importantly, the expectation of this K-L divergence never exceeds 1/2 bit. CONCLUSION: A rigorous underpinning of the Audic-Claverie methodology has been missing. Our results constitute a rigorous argument supporting the use of Audic-Claverie method even though the SAGE libraries represent very sparse samples. Peter Tiño |
BMC Bioinform. | 1 |
| 2009 | Relationship Between Generalization and Diversity in Coevolutionary LearningabstractGames have long played an important role in the development and understanding of coevolutionary learning systems. In particular, the search process in coevolutionary learning is guided by strategic interactions between solutions in the population, which can be naturally framed as game playing. We study two important issues in coevolutionary learning - generalization performance and diversity - using games. The first one is concerned with the coevolutionary learning of strategies with high generalization performance, that is, strategies that can outperform against a large number of test strategies (opponents) that may not have been seen during coevolution. The second one is concerned with diversity levels in the population that may lead to the search of strategies with poor generalization performance. It is not known if there is a relationship between generalization and diversity in coevolutionary learning. This paper investigates whether there is such a relationship in coevolutionary learning through a detailed empirical study. We systematically investigate the impact of various diversity maintenance approaches on the generalization performance of coevolutionary learning quantitatively using case studies. The problem of the iterated prisoner's dilemma (IPD) game is considered. Unlike past studies, we can measure both the generalization performance and the diversity level of the population of evolved strategies. Results from our case studies show that the introduction and maintenance of diversity do not necessarily lead to the coevolutionary learning of strategies with high generalization performance. However, if individual strategies can be combined (e.g., using a gating mechanism), there is the potential of exploiting diversity in coevolutionary learning to improve generalization performance. Specifically, when the introduction and maintenance of diversity lead to a speciated population during coevolution, where each specialist strategy is capable of outperforming different opponents, the population as a whole can have a significantly higher generalization performance compared to individual strategies. Siang Yew Chong, Peter Tiño, Xin Yao 0001 |
IEEE Trans. Comput. Intell. AI Games | 2 |
| 2009 | Predictive Ensemble Pruning by Expectation PropagationabstractAn ensemble is a group of learners that work together as a committee to solve a problem. The existing ensemble learning algorithms often generate unnecessarily large ensembles, which consume extra computational resource and may degrade the generalization performance. Ensemble pruning algorithms aim to find a good subset of ensemble members to constitute a small ensemble, which saves the computational resource and performs as well as, or better than, the unpruned ensemble. This paper introduces a probabilistic ensemble pruning algorithm by choosing a set of ldquosparserdquo combination weights, most of which are zeros, to prune the ensemble. In order to obtain the set of sparse combination weights and satisfy the nonnegative constraint of the combination weights, a left-truncated, nonnegative, Gaussian prior is adopted over every combination weight. Expectation propagation (EP) algorithm is employed to approximate the posterior estimation of the weight vector. The leave-one-out (LOO) error can be obtained as a by-product in the training of EP without extra computation and is a good indication for the generalization error. Therefore, the LOO error is used together with the Bayesian evidence for model selection in this algorithm. An empirical study on several regression and classification benchmark data sets shows that our algorithm utilizes far less component learners but performs as well as, or better than, the unpruned ensemble. Our results are very competitive compared with other ensemble pruning algorithms. Huanhuan Chen 0001, Peter Tiño, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2009 | Probabilistic Classification Vector MachinesabstractIn this paper, a sparse learning algorithm, probabilistic classification vector machines (PCVMs), is proposed. We analyze relevance vector machines (RVMs) for classification problems and observe that adopting the same prior for different classes may lead to unstable solutions. In order to tackle this problem, a signed and truncated Gaussian prior is adopted over every weight in PCVMs, where the sign of prior is determined by the class label, i.e., +1 or -1. The truncated Gaussian prior not only restricts the sign of weights but also leads to a sparse estimation of weight vectors, and thus controls the complexity of the model. In PCVMs, the kernel parameters can be optimized simultaneously within the training algorithm. The performance of PCVMs is extensively evaluated on four synthetic data sets and 13 benchmark data sets using three performance metrics, error rate (ERR), area under the curve of receiver operating characteristic (AUC), and root mean squared error (RMSE). We compare PCVMs with soft-margin support vector machines (SVM(Soft)), hard-margin support vector machines (SVM(Hard)), SVM with the kernel parameters optimized by PCVMs (SVM(PCVM)), relevance vector machines (RVMs), and some other baseline classifiers. Through five replications of twofold cross-validation F test, i.e., 5 x 2 cross-validation F test, over single data sets and Friedman test with the corresponding post-hoc test to compare these algorithms over multiple data sets, we notice that PCVMs outperform other algorithms, including SVM(Soft), SVM(Hard), RVM, and SVM(PCVM), on most of the data sets under the three metrics, especially under AUC. Our results also reveal that the performance of SVM(PCVM) is slightly better than SVM(Soft), implying that the parameter optimization algorithm in PCVMs is better than cross validation in terms of performance and computational complexity. In this paper, we also discuss the superiority of PCVMs' formulation using maximum a posteriori (MAP) analysis and margin analysis, which explain the empirical success of PCVMs. Huanhuan Chen 0001, Peter Tiño, Xin Yao 0001 |
IEEE Trans. Neural Networks | 2 |
| 2008 | Predictive Modeling with Echo State Networks
Michal Cernanský, Peter Tiño |
ICANN (1) | 2 |
| 2008 | Multiple Manifolds Learning Framework Based on Hierarchical Mixture Density Model
Peter Tiño, Mark A. Fardal |
ECML/PKDD (2) | 2 |
| 2008 | Measuring Generalization Performance in Coevolutionary LearningabstractCoevolutionary learning involves a training process where training samples are instances of solutions that interact strategically to guide the evolutionary (learning) process. One main research issue is with the generalization performance, i.e., the search for solutions (e.g., input–output mappings) that best predict the required output for any new input that has not been seen during the evolutionary process. However, there is currently no such framework for determining the generalization performance in coevolutionary learning even though the notion of generalization is well-understood in machine learning. In this paper, we introduce a theoretical framework to address this research issue. We present the framework in terms of game-playing although our results are more general. Here, a strategy's generalization performance is its average performance against all test strategies. Given that the true value may not be determined by solving analytically a closed-form formula and is computationally prohibitive, we propose an estimation procedure that computes the average performance against a small sample of random test strategies instead. We perform a mathematical analysis to provide a statistical claim on the accuracy of our estimation procedure, which can be further improved by performing a second estimation on the variance of the random variable. For game-playing, it is well-known that one is more interested in the generalization performance against a biased and diverse sample of “good” test strategies. We introduce a simple approach to obtain such a test sample through the multiple partial enumerative search of the strategy space that does not require human expertise and is generally applicable to a wide range of domains. We illustrate the generalization framework on the coevolutionary learning of the iterated prisoner's dilemma (IPD) games. We investigate two definitions of generalization performance for the IPD game based on different perf Siang Yew Chong, Peter Tiño, Xin Yao 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2008 | Visualization of Tree-Structured Data Through Generative Topographic MappingabstractIn this paper, we present a probabilistic generative approach for constructing topographic maps of tree-structured data. Our model defines a low-dimensional manifold of local noise models, namely, (hidden) Markov tree models, induced by a smooth mapping from low-dimensional latent space. We contrast our approach with that of topographic map formation using recursive neural-based techniques, namely, the self-organizing map for structured data (SOMSD) (Hagenbuchner , 2003). The probabilistic nature of our model brings a number of benefits: 1) naturally defined cost function that drives the model optimization; 2) principled model comparison and testing for overfitting; 3) a potential for transparent interpretation of the map by inspecting the underlying local noise models; 4) natural accommodation of alternative local noise models implicitly expressing different notions of structured data similarity. Furthermore, in contrast with the recursive neural-based approaches, the smooth nature of the mapping from the latent space to the local model space allows for calculation of magnification factors--a useful tool for the detection of data clusters. We demonstrate our approach on three data sets: a toy data set, an artificially generated data set, and on a data set of images represented as quadtrees. Nikolaos Gianniotis, Peter Tiño |
IEEE Trans. Neural Networks | 2 |
| 2007 | Visualisation of tree-structured data through generative probabilistic modelling
Nikolaos Gianniotis, Peter Tiño |
ESANN | 2 |
| 2007 | Comparison of Echo State Networks with Simple Recurrent Networks and Variable-Length Markov Models on Symbolic Sequences
Michal Cernanský, Peter Tiño |
ICANN (1) | 2 |
| 2007 | Bifurcations of Renormalization Dynamics in Self-organizing Neural Networks
Peter Tiño |
ICONIP (1) | 1 |
| 2007 | Metric Properties of Structured Data Visualizations through Generative Probabilistic Modeling
Peter Tiño, Nikolaos Gianniotis |
IJCAI | 1 |
| 2007 | Equilibria of Iterative Softmax and Critical Temperatures for Intermittent Search in Self-Organizing Neural NetworksabstractKwok and Smith (2005) recently proposed a new kind of optimization dynamics using self-organizing neural networks (SONN) driven by softmax weight renormalization. Such dynamics is capable of powerful intermittent search for high-quality solutions in difficult assignment optimization problems. However, the search is sensitive to temperature setting in the softmax renormalization step. It has been hypothesized that the optimal temperature setting corresponds to the symmetry-breaking bifurcation of equilibria of the renormalization step, when viewed as an autonomous dynamical system called iterative softmax (ISM). We rigorously analyze equilibria of ISM by determining their number, position, and stability types. It is shown that most fixed points exist in the neighborhood of the maximum entropy equilibrium w= (N(-1), N(-1), ..., N(-1)), where N is the ISM dimensionality. We calculate the exact rate of decrease in the number of ISM equilibria as one moves away from w. Bounds on temperatures guaranteeing different stability types of ISM equilibria are also derived. Moreover, we offer analytical approximations to the critical symmetry-breaking bifurcation temperatures that are in good agreement with those found by numerical investigations. So far, the critical temperatures have been determined only by trial-and-error numerical simulations. On a set of N-queens problems for a wide range of problem sizes N, the analytically determined critical temperatures predict the optimal working temperatures for SONN intermittent search very well. It is also shown that no intermittent search can exist in SONN for temperatures greater than one-half. Peter Tiño |
Neural Comput. | 1 |
| 2006 | A Kernel-Based Approach to Estimating Phase Shifts Between Irregularly Sampled Time Series: An Application to Gravitational Lenses
Juan Carlos Cuevas-Tello, Peter Tiño, Somak Raychaudhury |
ECML | 2 |
| 2006 | Critical Temperatures for Intermittent Search in Self-Organizing Neural Networks
Peter Tiño |
PPSN | 1 |
| 2006 | Dynamics and Topographic Organization of Recursive Self-Organizing MapsabstractRecently there has been an outburst of interest in extending topographic maps of vectorial data to more general data structures, such as sequences or trees. However, there is no general consensus as to how best to process sequences using topographic maps, and this topic remains an active focus of neurocomputational research. The representational capabilities and internal representations of the models are not well understood. Here, we rigorously analyze a generalization of the self-organizing map (SOM) for processing sequential data, recursive SOM(RecSOM) (Voegtlin, 2002), as a nonautonomous dynamical system consisting of a set of fixed input maps. We argue that contractive fixed-input maps are likely to produce Markovian organizations of receptive fields on the RecSOM map. We derive bounds on parameter beta (weighting the importance of importing past information when processing sequences) under which contractiveness of the fixed-input maps is guaranteed. Some generalizations of SOM contain a dynamic module responsible for processing temporal contexts as an integral part of the model. We show that Markovian topographic maps of sequential data can be produced using a simple fixed (nonadaptable) dynamic module externally feeding a standard topographic model designed to process static vectorial data of fixed dimensionality (e.g., SOM). However, by allowing trainable feedback connections, one can obtain Markovian maps with superior memory depth and topography preservation. We elaborate on the importance of non-Markovian organizations in topographic maps of sequential data. Peter Tiño, Igor Farkas, Jort van Mourik |
Neural Comput. | 1 |
| 2006 | Learning Beyond Finite Memory in Recurrent Networks of Spiking NeuronsabstractWe investigate possibilities of inducing temporal structures without fading memory in recurrent networks of spiking neurons strictly operating in the pulse-coding regime. We extend the existing gradient-based algorithm for training feedforward spiking neuron networks, SpikeProp (Bohte, Kok, & La Poutré, 2002), to recurrent network topologies, so that temporal dependencies in the input stream are taken into account. It is shown that temporal structures with unbounded input memory specified by simple Moore machines (MM) can be induced by recurrent spiking neuron networks (RSNN). The networks are able to discover pulse-coded representations of abstract information processing states coding potentially unbounded histories of processed inputs. We show that it is often possible to extract from trained RSNN the target MM by grouping together similar spike trains appearing in the recurrent layer. Even when the target MM was not perfectly induced in a RSNN, the extraction procedure was able to reveal weaknesses of the induced mechanism and the extent to which the target machine had been learned. Peter Tiño, Ashley J. S. Mills |
Neural Comput. | 1 |
| 2005 | Recursive Self-organizing Map as a Contractive Iterative Function System
Peter Tiño, Igor Farkas, Jort van Mourik |
IDEAL | 1 |
| 2005 | Sequential relevance vector machine learning from time seriesabstractThis paper presents an approach to sequential training of the relevance vector machine suitable for Bayesian learning from time series. The key idea is to perform simultaneous incremental optimization of both the weight parameters and their prior hyperparameters using data arriving successively one at a time. Algorithms for efficient sequential regularized dynamic learning rate training of the weights and gradient-descent training of their corresponding individual priors are derived. It is shown that this fast sequential RVM can outperform similar Bayesian kernel methods, like: batch RVM, fast RVM, variational RVM, and Gaussian processes on multistep ahead forecasting of time series. Nikolay I. Nikolaev, Peter Tiño |
IJCNN | 2 |
| 2005 | Managing Diversity in Regression EnsemblesabstractEnsembles are a widely used and effective technique in machine learning---their success is commonly attributed to the degree of disagreement, or 'diversity', within the ensemble. For ensembles where the individual estimators output crisp class labels, this 'diversity' is not well understood and remains an open research issue. For ensembles of regression estimators, the diversity can be exactly formulated in terms of the covariance between individual estimator outputs, and the optimum level is expressed in terms of a bias-variance-covariance trade-off. Despite this, most approaches to learning ensembles use heuristics to encourage the right degree of diversity. In this work we show how to explicitly control diversity through the error function. The first contribution of this paper is to show that by taking the combination mechanism for the ensemble into account we can derive an error function for each individual that balances ensemble diversity with individual accuracy. We show the relationship between this error function and an existing algorithm called negative correlation learning, which uses a heuristic penalty term added to the mean squared error function. It is demonstrated that these methods control the bias-variance-covariance trade-off systematically, and can be utilised with any estimator capable of minimising a quadratic error function, for example MLPs, or RBF networks. As a second contribution, we derive a strict upper bound on the coefficient of the penalty term, which holds for any estimator that can be cast in a generalised linear regression framework, with mild assumptions on the basis functions. Finally we present the results of an empirical study, showing significant improvements over simple ensemble learning, and finding that this technique is competitive with a variety of methods, including boosting, bagging, mixtures of experts, and Gaussian processes, on a number of tasks. Gavin Brown 0001, Jeremy L. Wyatt, Peter Tiño |
J. Mach. Learn. Res. | 3 |
| 2005 | Semisupervised Learning of Hierarchical Latent Trait Models for Data VisualizationabstractRecently, we have developed the hierarchical generative topographic mapping (HGTM), an interactive method for visualization of large high-dimensional real-valued data sets. We propose a more general visualization system by extending HGTM in three ways, which allows the user to visualize a wider range of data sets and better support the model development process. 1) We integrate HGTM with noise models from the exponential family of distributions. The basic building block is the latent trait model (LTM). This enables us to visualize data of inherently discrete nature, e.g., collections of documents, in a hierarchical manner. 2) We give the user a choice of initializing the child plots of the current plot in either interactive, or automatic mode. In the interactive mode, the user selects "regions of interest", whereas in the automatic mode, an unsupervised minimum message length (MML)-inspired construction of a mixture of LTMs is employed. The unsupervised construction is particularly useful when high-level plots are covered with dense clusters of highly overlapping data projections, making it difficult to use the interactive mode. Such a situation often arises when visualizing large data sets. 3) We derive general formulas for magnification factors in latent trait models. Magnification factors are a useful tool to improve our understanding of the visualization plots, since they can highlight the boundaries between data clusters. We illustrate our approach on a toy example and evaluate it on three more complex real data sets. Ian T. Nabney, Yi Sun 0001, Peter Tiño, Ata Kabán |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2004 | A generative probabilistic approach to visualizing sets of symbolic sequencesabstractThere is a notable interest in extending probabilistic generative modeling principles to accommodate for more complex structured data types. In this paper we develop a generative probabilistic model for visualizing sets of discrete symbolic sequences. The model, a constrained mixture of discrete hidden Markov models, is a generalization of density-based visualization methods previously developed for static data sets. We illustrate our approach on sequences representing web-log data and chorals by J.S. Bach. Peter Tiño, Ata Kabán, Yi Sun 0001 |
KDD | 1 |
| 2004 | Evaluation of Adaptive Nature Inspired Task Allocation Against Alternate Decentralised Multiagent Strategies
Richard Price, Peter Tiño |
PPSN | 2 |
| 2004 | Making sense of sparse rating data in collaborative filtering via topographic organization of user preference patterns
Gabriela Grmanová, Peter Tiño |
Neural Networks | 2 |
| 2004 | Markovian architectural bias of recurrent neural networksabstractIn this paper, we elaborate upon the claim that clustering in the recurrent layer of recurrent neural networks (RNNs) reflects meaningful information processing states even prior to training [1], [2]. By concentrating on activation clusters in RNNs, while not throwing away the continuous state space network dynamics, we extract predictive models that we call neural prediction machines (NPMs). When RNNs with sigmoid activation functions are initialized with small weights (a common technique in the RNN community), the clusters of recurrent activations emerging prior to training are indeed meaningful and correspond to Markov prediction contexts. In this case, the extracted NPMs correspond to a class of Markov models, called variable memory length Markov models (VLMMs). In order to appreciate how much information has really been induced during the training, the RNN performance should always be compared with that of VLMMs and NPMs extracted before training as the "null" base models. Our arguments are supported by experiments on a chaotic symbolic sequence and a context-free language with a deep recursive structure. Index Terms-Complex symbolic sequences, information latching problem, iterative function systems, Markov models, recurrent neural networks (RNNs). Peter Tiño, Michal Cernanský, Lubica Benusková |
IEEE Trans. Neural Networks | 1 |
| 2003 | Recurrent Neural Networks with Small Weights Implement Definite Memory MachinesabstractRecent experimental studies indicate that recurrent neural networks initialized with “small” weights are inherently biased toward definite memory machines (Tiňno, Čerňanský, & Beňušková, 2002a, 2002b). This article establishes a theoretical counterpart: transition function of recurrent network with small weights and squashing activation function is a contraction. We prove that recurrent networks with contractive transition function can be approximated arbitrarily well on input sequences of unbounded length by a definite memory machine. Conversely, every definite memory machine can be simulated by a recurrent network with contractive transition function. Hence, initialization with small weights induces an architectural bias into learning with recurrent neural networks. This bias might have benefits from the point of view of statistical learning theory: it emphasizes one possible region of the weight space where generalization ability can be formally proved. It is well known that standard recurrent neural networks are not distribution independent learnable in the probably approximately correct (PAC) sense if arbitrary precision and inputs are considered. We prove that recurrent networks with contractive transition function with a fixed contraction parameter fulfill the so-called distribution independent uniform convergence of empirical distances property and hence, unlike general recurrent networks, are distribution independent PAC learnable. Barbara Hammer, Peter Tiño |
Neural Comput. | 2 |
| 2003 | Architectural Bias in Recurrent Neural Networks: Fractal AnalysisabstractWe have recently shown that when initialized with “small” weights, recurrent neural networks (RNNs) with standard sigmoid-type activation functions are inherently biased toward Markov models; even prior to any training, RNN dynamics can be readily used to extract finite memory machines (Hammer & Tiňo, 2002; Tiňo, Čerňanský, &Beňušková, 2002a, 2002b). Following Christiansen and Chater (1999), we refer to this phenomenon as the architectural bias of RNNs. In this article, we extend our work on the architectural bias in RNNs by performing a rigorous fractal analysis of recurrent activation patterns. We assume the network is driven by sequences obtained by traversing an underlying finite-state transition diagram&a scenario that has been frequently considered in the past, for example, when studying RNN-based learning and implementation of regular grammars and finite-state transducers. We obtain lower and upper bounds on various types of fractal dimensions, such as box counting and Hausdorff dimensions. It turns out that not only can the recurrent activations inside RNNs with small initial weights be explored to build Markovian predictive models, but also the activations form fractal clusters, the dimension of which can be bounded by the scaled entropy of the underlying driving source. The scaling factors are fixed and are given by the RNN parameters. Peter Tiño, Barbara Hammer |
Neural Comput. | 1 |
| 2002 | Architectural Bias in Recurrent Neural Networks - Fractal Analysis
Peter Tiño, Barbara Hammer |
ICANN | 1 |
| 2002 | A General Framework for a Principled Hierarchical Visualization of Multivariate Data
Ata Kabán, Peter Tiño, Mark A. Girolami |
IDEAL | 2 |
| 2002 | Hierarchical GTM: Constructing Localized Nonlinear Projection Manifolds in a Principled WayabstractIt has been argued that a single two-dimensional visualization plot may not be sufficient to capture all of the interesting aspects of complex data sets and, therefore, a hierarchical visualization system is desirable. In this paper, we extend an existing locally linear hierarchical visualization system PhiVis in several directions: 1) We allow for nonlinear projection manifolds. The basic building block is the Generative Topographic Mapping (GTM). 2) We introduce a general formulation of hierarchical probabilistic models consisting of local probabilistic models organized in a hierarchical tree. General training equations are derived, regardless of the position of the model in the tree. 3) Using tools from differential geometry, we derive expressions for local directional curvatures of the projection manifold. Like PhiVis, our system is statistically principled and is built interactively in a top-down fashion using the EM algorithm. It enables the user to interactively highlight those data in the ancestor visualization plots which are captured by a child model. We also incorporate into our system a hierarchical, locally selective representation of magnification factors and directional curvatures of the projection manifolds. Such information is important for further refinement of the hierarchical visualization plot, as well as for controlling the amount of regularization imposed on the local models. We demonstrate the principle of the approach on a toy data set and apply our system to two more complex 12- and 18-dimensional data sets. Peter Tiño, Ian T. Nabney |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Using Directional Curvatures to Visualize Folding Patterns of the GTM Projection Manifolds
Peter Tiño, Ian T. Nabney, Yi Sun 0001 |
ICANN | 1 |
| 2001 | Predicting the Future of Discrete Sequences from Fractal Representations of the Past
Peter Tiño, Georg Dorffner |
Mach. Learn. | 1 |
| 2001 | Attractive Periodic Sets in Discrete-Time Recurrent Networks (with Emphasis on Fixed-Point Stability and Bifurcations in Two-Neuron Networks)abstractWe perform a detailed fixed-point analysis of two-unit recurrent neural networks with sigmoid-shaped transfer functions. Using geometrical arguments in the space of transfer function derivatives, we partition the network state-space into distinct regions corresponding to stability types of the fixed points. Unlike in the previous studies, we do not assume any special form of connectivity pattern between the neurons, and all free parameters are allowed to vary. We also prove that when both neurons have excitatory self-connections and the mutual interaction pattern is the same (i.e., the neurons mutually inhibit or excite themselves), new attractive fixed points are created through the saddle-node bifurcation. Finally, for an N-neuron recurrent network, we give lower bounds on the rate of convergence of attractive periodic points toward the saturation values of neuron activations, as the absolute values of connection weights grow. Peter Tiño, Bill G. Horne, C. Lee Giles |
Neural Comput. | 1 |
| 2001 | Volatility Trading ia Temporal Pattern Recognition in Quantised Financial Time Series
Peter Tiño, Christian Schittenkopf, Georg Dorffner |
Pattern Anal. Appl. | 1 |
| 2001 | Financial volatility trading using recurrent neural networksabstractWe simulate daily trading of straddles on financial indexes. The straddles are traded based on predictions of daily volatility differences in the indexes. The main predictive models studied are recurrent neural nets (RNN). Such applications have often been studied in isolation. However, due to the special character of daily financial time-series, it is difficult to make full use of RNN representational power. Recurrent networks either tend to overestimate noisy data, or behave like finite-memory sources with shallow memory; they hardly beat classical fixed-order Markov models. To overcome data nonstationarity, we use a special technique that combines sophisticated models fitted on a larger data set, with a fixed set of simple-minded symbolic predictors using only recent inputs. Finally, we compare our predictors with the GARCH family of econometric models designed to capture time-dependent volatility structure in financial returns. GARCH models have been used to trade volatility. Experimental results show that while GARCH models cannot generate any significantly positive profit, by careful use of recurrent networks or Markov models, the market makers can generate a statistically significant excess profit, but then there is no reason to prefer RNN over much more simple and straightforward Markov models. We argue that any report containing RNN results on financial tasks should be accompanied by results achieved by simple finite-memory sources combined with simple techniques to fight nonstationarity in the data. Peter Tiño, Christian Schittenkopf, Georg Dorffner |
IEEE Trans. Neural Networks | 1 |
| 2000 | The profitability of trading volatility using real-valued and symbolic modelsabstractThere are two notions of volatility in literature: historical volatility and implied volatility. We concentrate on the latter by analyzing the profitability of a pure volatility trading strategy which is delta-neutral and independent of an option pricing model, for the German stock index DAX. Several very different methods ranging from linear and nonlinear, real-valued models to symbolic models of volatility changes are applied to predict the change in volatility to the next trading day and to gain profits by buying or selling straddles accordingly. The trading performance is evaluated for one historical and one implied volatility measure. The results are carefully evaluated concerning transaction costs, stationarity issues, and statistical significance. The main contribution of the paper is that, for the first time, the trading performance of models based on different modelling paradigms is compared. Christian Schittenkopf, Peter Tiño, Georg Dorffner |
CIFEr | 2 |
| 2000 | Building Predictive Models on Complex Symbolic Sequences with a Second-Order Recurrent BCM Network with Lateral InhibitionabstractWhen trained on symbolic sequences to perform the next-symbol prediction, recurrent neural networks (RNNs) tend to organize their state space so that "close" recurrent activation vectors correspond to histories of symbols yielding similar next-symbol distributions. In this paper we investigate an unsupervised alternative to the state space organization. In particular, we use a recurrent version of the Bienenstock-Cooper-Munro (BCM) network with lateral inhibition to map histories of symbols into activations of the recurrent layer. Recurrent BCM networks perform a kind of time-conditional projection pursuit. We compare the finite-context models built on top of BCM recurrent activations with those constructed on top of RNN recurrent activation vectors. As a test bed we use two complex symbolic sequences with rather deep memory structures. It is shown that the BCM-based model has a comparable or better performance than its RNN-based counterpart. This can be explained by the familiar information latching problem in recurrent networks when longer time spans are to be latched. Peter Tiño, Michal Stancík, Lubica Benusková |
IJCNN (2) | 1 |
| 1999 | Graded Grammaticality in Prediction Fractal Machines
Shan Parfitt, Peter Tiño, Georg Dorffner |
NIPS | 2 |
| 1999 | Building Predictive Models from Fractal Representations of Symbolic Sequences
Peter Tiño, Georg Dorffner |
NIPS | 1 |
| 1999 | Extracting finite-state representations from recurrent neural networks trained on chaotic symbolic sequencesabstractWhile much work has been done in neural-based modeling of real-valued chaotic time series, little effort has been devoted to address similar problems in the symbolic domain. We investigate the knowledge induction process associated with training recurrent neural networks (RNN's) on single long chaotic symbolic sequences. Even though training RNN's to predict the next symbol leaves the standard performance measures such as the mean square error on the network output virtually unchanged, the networks nevertheless do extract a lot of knowledge. We monitor the knowledge extraction process by considering the networks stochastic sources and letting them generate sequences which are then confronted with the training sequence via information theoretic entropy and cross-entropy measures. We also study the possibility of reformulating the knowledge gained by RNN's in a compact and easy-to-analyze form of finite-state stochastic machines. The experiments are performed on two sequences with different "complexities" measured by the size and state transition structure of the induced Crutchfield's epsilon-machines. We find that, with respect to the original RNN's, the extracted machines can achieve comparable or even better entropy and cross-entropy performance. Moreover, RNN's reflect the training sequence complexity in their dynamical state representations that can in turn be reformulated using finite-state means. Our findings are confirmed by a much more detailed analysis of model generated sequences through the statistical mechanical metaphor of entropy spectra. We also introduce a visual representation of allowed block structure in the studied sequences that, besides having nice theoretical properties, allows on the topological level for an illustrative insight into both RNN training and finite-state stochastic machine extraction processes. Peter Tiño, Miroslav Koteles |
IEEE Trans. Neural Networks | 1 |
| 1999 | Spatial representation of symbolic sequences through iterative function systemsabstractJeffrey proposed (1990) a graphic representation of DNA sequences using Barnsley's iterative function systems. In spite of further developments in this direction, the proposed graphic representation of DNA sequences has been lacking a rigorous connection between its spatial scaling characteristics and the statistical characteristics of the DNA sequences themselves. We 1) generalize Jeffrey's graphic representation to accommodate (possibly infinite) sequences over an arbitrary finite number of symbols; 2) establish a direct correspondence between the statistical characterization of symbolic sequences via Renyi entropy spectra (1959) and the multifractal characteristics (Renyi generalized dimensions) of the sequences' spatial representations; 3) show that for general symbolic dynamical systems, the multifractal f/sub H/-spectra in the sequence space coincide with the f/sub H/-spectra on spatial sequence representations. Peter Tiño |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 1997 | Extracting stochastic machines from recurrent neural networks trained on complex symbolic sequencesabstractWe train a recurrent neural network on a single, long, complex symbolic sequence with positive entropy. The training process is monitored through information theory based performance measures. We show that although the sequence is unpredictable, the network is able to code the sequence's topological and statistical structure in recurrent neuron activation scenarios. Such scenarios can be compactly represented through stochastic machines extracted from the trained network. Generative models, i.e. trained recurrent networks and extracted stochastic machines, are compared using entropy spectra of generated sequences. In addition, entropy spectra computed directly from the machines capture generalization abilities of extracted machines and are related to a machines' long term behavior. Peter Tiño, V. Vojtek |
KES (2) | 1 |
| 1996 | Learning long-term dependencies in NARX recurrent neural networksabstractIt has previously been shown that gradient-descent learning algorithms for recurrent neural networks can perform poorly on tasks that involve long-term dependencies, i.e. those problems for which the desired output depends on inputs presented at times far in the past. We show that the long-term dependencies problem is lessened for a class of architectures called nonlinear autoregressive models with exogenous (NARX) recurrent neural networks, which have powerful representational capabilities. We have previously reported that gradient descent learning can be more effective in NARX networks than in recurrent neural network architectures that have "hidden states" on problems including grammatical inference and nonlinear system identification. Typically, the network converges much faster and generalizes better than other networks. The results in this paper are consistent with this phenomenon. We present some experimental results which show that NARX networks can often retain information for two to three times as long as conventional recurrent neural networks. We show that although NARX networks do not circumvent the problem of long-term dependencies, they can greatly improve performance on long-term dependency problems. We also describe in detail some of the assumptions regarding what it means to latch information robustly and suggest possible ways to loosen these assumptions. Tsungnan Lin, Bill G. Horne, Peter Tiño, C. Lee Giles |
IEEE Trans. Neural Networks | 3 |
| 1995 | Learning long-term dependencies is not as difficult with NARX networks
Tsungnan Lin, Bill G. Horne, Peter Tiño, C. Lee Giles |
NIPS | 3 |
| 1995 | Learning and Extracting Initial Mealy Automata with a Modular Neural Network ModelabstractA hybrid recurrent neural network is shown to learn small initial mealy machines (that can be thought of as translation machines translating input strings to corresponding output strings, as opposed to recognition automata that classify strings as either grammatical or nongrammatical) from positive training samples. A well-trained neural net is then presented once again with the training set and a Kohonen self-organizing map with the “star” topology of neurons is used to quantize recurrent network state space into distinct regions representing corresponding states of a mealy machine being learned. This enables us to extract the learned mealy machine from the trained recurrent network. One neural network (Kohonen self-organizing map) is used to extract meaningful information from another network (recurrent neural network). Peter Tiño, Jozef Sajda |
Neural Comput. | 1 |