EDBT 2026 Demo / reviewers in the wild / expert
Sotirios Chatzis
dblp:25/6133 · also Sotirios P. Chatzis
· DBLP profile ↗
70ranked-venue papers
46as first author
8since 2021 · last 2023
0000-0002-4956-4013ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 40 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorSecurity and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | DISCOVER: Making Vision Networks Interpretable via Competition and DissectionabstractModern deep networks are highly complex and their inferential outcome very hard to interpret. This is a serious obstacle to their transparent deployment in safety-critical or bias-aware applications. This work contributes to *post-hoc* interpretability, and specifically Network Dissection. Our goal is to present a framework that makes it easier to *discover* the individual functionality of each neuron in a network trained on a vision task; discovery is performed in terms of textual description generation. To achieve this objective, we leverage: (i) recent advances in multimodal vision-text models and (ii) network layers founded upon the novel concept of stochastic local competition between linear units. In this setting, only a *small subset* of layer neurons are activated *for a given input*, leading to extremely high activation sparsity (as low as only $\approx 4\%$). Crucially, our proposed method infers (sparse) neuron activation patterns that enables the neurons to activate/specialize to inputs with specific characteristics, diversifying their individual functionality. This capacity of our method supercharges the potential of dissection processes: human understandable descriptions are generated only for the very few active neurons, thus facilitating the direct investigation of the network's decision process. As we experimentally show, our approach: (i) yields Vision Networks that retain or improve classification performance, and (ii) realizes a principled framework for text-based description and examination of the generated neuronal representations. Konstantinos P. Panousis, Sotirios Chatzis |
NeurIPS | 2 |
| 2023 | Continuous authentication with feature-level fusion of touch gestures and keystroke dynamics to solve security and usability issues
Ioannis Stylios, Sotirios Chatzis, Olga Thanou, Spyros Kokolakis |
Comput. Secur. | 2 |
| 2022 | Competing Mutual Information Constraints with Stochastic Competition-Based Activations for Learning Diversified RepresentationsabstractThis work aims to address the long-established problem of learning diversified representations. To this end, we combine information-theoretic arguments with stochastic competition-based activations, namely Stochastic Local Winner-Takes-All (LWTA) units. In this context, we ditch the conventional deep architectures commonly used in Representation Learning, that rely on non-linear activations; instead, we replace them with sets of locally and stochastically competing linear units. In this setting, each network layer yields sparse outputs, determined by the outcome of the competition between units that are organized into blocks of competitors. We adopt stochastic arguments for the competition mechanism, which perform posterior sampling to determine the winner of each block. We further endow the considered networks with the ability to infer the sub-part of the network that is essential for modeling the data at hand; we impose appropriate stick-breaking priors to this end. To further enrich the information of the emerging representations, we resort to information-theoretic principles, namely the Information Competing Process (ICP). Then, all the components are tied together under the stochastic Variational Bayes framework for inference. We perform a thorough experimental investigation for our approach using benchmark datasets on image classification. As we experimentally show, the resulting networks yield significant discriminative representation learning abilities. In addition, the introduced paradigm allows for a principled investigation mechanism of the emerging intermediate network representations. Konstantinos P. Panousis, Anastasios Antoniadis, Sotirios Chatzis |
AAAI | 3 |
| 2022 | Stochastic Deep Networks with Linear Competing Units for Model-Agnostic Meta-LearningabstractThis work addresses meta-learning (ML) by considering deep networks with stochastic local winner-takes-all (LWTA) activations. This type of network units results in sparse representations from each model layer, as the units are organized into blocks where only one unit generates a non-zero output. The main operating principle of the introduced units rely on stochastic principles, as the network performs posterior sampling over competing units to select the winner. Therefore, the proposed networks are explicitly designed to extract input data representations of sparse stochastic nature, as opposed to the currently standard deterministic representation paradigm. Our approach produces state-of-the-art predictive accuracy on few-shot image classification and regression experiments, as well as reduced predictive error on an active learning setting; these improvements come with an immensely reduced computational cost. Code is available at: https://github.com/Kkalais/StochLWTA-ML Konstantinos Kalais, Sotirios Chatzis |
ICML | 2 |
| 2022 | BioGames: a new paradigm and a behavioral biometrics collection tool for research purposesabstractPurpose The purpose of this paper is to present a new paradigm, named BioGames, for the extraction of behavioral biometrics (BB) conveniently and entertainingly. To apply the BioGames paradigm, the authors developed a BB collection tool for mobile devices named BioGames App. The BioGames App collects keystroke dynamics, touch gestures, and motion modalities and is available on GitHub. Interested researchers and practitioners may use it to create their datasets for research purposes. Design/methodology/approach One major challenge for BB and continuous authentication (CA) research is the lack of actual BB datasets for research purposes. The compilation and refinement of an appropriate set of BB data constitute a challenge and an open problem. The issue is aggravated by the fact that most users are reluctant to participate in long demanding procedures entailed in the collection of research biometric data. As a result, they do not complete the data collection procedure, or they do not complete it correctly. Therefore, the authors propose a new paradigm and introduce a BB collection tool, which they call BioGames, for the extraction of biometric features in a convenient way. The BioGames paradigm proposes a methodology where users play games without participating in an experimental painstaking process. The BioGames App collects keystroke dynamics, touch gestures, and motion modalities. Findings The authors proposed a new paradigm for the collection of BB on mobile devices and created the BioGames application. The BioGames App is an Android application that collects BB data on mobile devices and sends them to a database. The database design allows multiple users to store their sensor data at any time. Thus, there is no concern about data separation and synchronization. BioGames App is General Data Protection Regulation (GDPR) compliant as it collects and processes only anonymous data. Originality/value The BioGames App is a publicly available tool that combines the keystroke dynamics, touch gestures, and motion modalities. In addition, it uses a methodology where users play games without participating in an experimental painstaking process. Ioannis Stylios, Spyros Kokolakis, Andreas Skalkos, Sotirios Chatzis |
Inf. Comput. Secur. | 4 |
| 2022 | Key factors driving the adoption of behavioral biometrics and continuous authentication technology: an empirical researchabstractPurpose For the success of future investments in the implementation of continuous authentication systems, we should explore the key factors that influence technology adoption. The authors investigate the effect of various factors of behavioral intention through the new incorporation of a modified technology acceptance model (TAM) and diffusion of innovation theory (DOI). Also, the authors have created a new theoretical framework with constructs such as security and privacy risks (SPR), biometrics privacy concerns (BPC) and perceived risk of using the technology (PROU). In this paper, the authors conducted a structural equation modeling empirical research. This research is designed in such a way to respond to the trade-off between users’ concern for the protection of their biometrics privacy and their protection from risks. Design/methodology/approach The authors provide an extensive conceptual framework for both existing models (TAM and DOI) and the new constructs that the authors have added to the model. In addition, this research explores external factors, such as trust in technology (TT) and innovativeness (Innov). In addition, the authors have introduced significant constructs, to overcome the limitations of the TAM and to adapt it to the needs of the present research. The new theoretical framework the authors introduce in the present research concerns the constructs SPR, BPC and PROU. Findings The authors found that the main facilitators of behavioral intention to adopt the technology (BI) are TT, followed by compatibility (COMP), perceived usefulness (PU) and Innov. This research also shows that individuals are less interested in the ease of use of the technology and are willing to sacrifice it to achieve greater security. COMP and Innov also play a significant role. Individuals who believe that the usage of the behavioral biometrics continuous authentication (BBCA) technology would fit into their lifestyle and would like to experiment with new technologies have a positive intention to adopt the BBCA technology. The new constructs the authors have added are SPR, BPC and PROU. The authors’ results support the hypotheses that SPR is a facilitator to PU and PU acts as a facilitator to BI. Consequently, the hypothesis that individuals do not feel adequately protected by classical methods will consider the usefulness of the BBCA as a technology for their extra protection against risks is confirmed by the model. Also, with the constructs BPC and PROU, the authors examined if individuals’ concerns regarding their biometrics privacy act as inhibitors in the BI. The authors concluded that individuals consider that the benefits of using BBCA technology are much more important than the risks for their biometrics privacy since the hypothesis that the major inhibitor of BI is PROU is not supported by the model. Originality/value To the best of the authors’ knowledge, this research is among the first in the field that examines the factors that influence the individuals’ decision to adopt BBCA technology. Ioannis Stylios, Spyros Kokolakis, Olga Thanou, Sotirios Chatzis |
Inf. Comput. Secur. | 4 |
| 2021 | Local Competition and Stochasticity for Adversarial Robustness in Deep LearningabstractThis work addresses adversarial robustness in deep learning by considering deep networks with stochastic local winner-takes-all (LWTA) activations. This type of network units result in sparse representations from each model layer, as the units are organized in blocks where only one unit generates a non-zero output. The main operating principle of the introduced units lies on stochastic arguments, as the network performs posterior sampling over competing units to select the winner. We combine these LWTA arguments with tools from the field of Bayesian non-parametrics, specifically the stick-breaking construction of the Indian Buffet Process, to allow for inferring the sub-part of each layer that is essential for modeling the data at hand. Then, inference is performed by means of stochastic variational Bayes. We perform a thorough experimental evaluation of our model using benchmark datasets. As we show, our method achieves high robustness to adversarial perturbations, with state-of-the-art performance in powerful adversarial attack schemes. Konstantinos P. Panousis, Sotirios Chatzis, Antonios Alexos, Sergios Theodoridis |
AISTATS | 2 |
| 2021 | Stochastic Transformer Networks with Linear Competing Units: Application to end-to-end SL TranslationabstractAutomating sign language translation (SLT) is a challenging real-world application. Despite its societal importance, though, research progress in the field remains rather poor. Crucially, existing methods that yield viable performance necessitate the availability of laborious to obtain gloss sequence groundtruth. In this paper, we attenuate this need, by introducing an end-to-end SLT model that does not entail explicit use of glosses; the model only needs text groundtruth. This is in stark contrast to existing end-to-end models that use gloss sequence groundtruth, either in the form of a modality that is recognized at an intermediate model stage, or in the form of a parallel output process, jointly trained with the SLT model. Our approach constitutes a Transformer network with a novel type of layers that combines: (i) local winner-takes-all (LWTA) layers with stochastic winner sampling, instead of conventional ReLU layers, (ii) stochastic weights with posterior distributions estimated via variational inference, and (iii) a weight compression technique at inference time that exploits estimated posterior variance to perform massive, almost lossless compression. We demonstrate that our approach can reach the currently best reported BLEU-4 score on the PHOENIX 2014T benchmark, but without making use of glosses for model training, and with a memory footprint reduced by more than 70%. Andreas Voskou, Konstantinos P. Panousis, Dimitrios I. Kosmopoulos, Dimitris N. Metaxas, Sotirios Chatzis |
ICCV | 5 |
| 2020 | A Self-Attentive Emotion Recognition NetworkabstractAttention networks constitute the state-of-the-art paradigm for capturing long temporal dynamics. This paper examines the efficacy of this paradigm in the challenging task of emotion recognition in dyadic conversations. In this work, we introduce a novel attention mechanism capable of inferring the immensity of the effect of each past utterance on the current speaker emotional state. The proposed self-attention network captures the correlation patterns among consecutive encoder network states, thus enabling the robust and effective modeling of temporal dynamics over arbitrary long temporal horizons. We exhibit the effectiveness of our approach considering the challenging IEMOCAP benchmark. We show that, our devised methodology outperforms state-of-the-art alternatives and commonly used approaches, giving rise to promising new research directions in the context of Online Social Network (OSN) analysis tasks. Harris Partaourides, Kostantinos Papadamou, Nicolas Kourtellis, Ilias Leontiadis, Sotirios Chatzis |
ICASSP | 5 |
| 2020 | Gated Mixture Variational Autoencoders for Value Added Tax audit case selection
Christos Kleanthous, Sotirios Chatzis |
Knowl. Based Syst. | 2 |
| 2019 | Nonparametric Bayesian Deep Networks with Local CompetitionabstractThe aim of this work is to enable inference of deep networks that retain high accuracy for the least possible model complexity, with the latter deduced from the data during inference. To this end, we revisit deep networks that comprise competing linear units, as opposed to nonlinear units that do not entail any form of (local) competition. In this context, our main technical innovation consists in an inferential setup that leverages solid arguments from Bayesian nonparametrics. We infer both the needed set of connections or locally competing sets of units, as well as the required floating-point precision for storing the network parameters. Specifically, we introduce auxiliary discrete latent variables representing which initial network components are actually needed for modeling the data at hand, and perform Bayesian inference over them by imposing appropriate stick-breaking priors. As we experimentally show using benchmark datasets, our approach yields networks with less computational footprint than the state-of-the-art, and with no compromises in predictive accuracy. Konstantinos P. Panousis, Sotirios Chatzis, Sergios Theodoridis |
ICML | 2 |
| 2019 | t-Exponential Memory Networks for Question-Answering MachinesabstractRecent advances in deep learning have brought to the fore models that can make multiple computational steps in the service of completing a task; these are capable of describing long-term dependencies in sequential data. Novel recurrent attention models over possibly large external memory modules constitute the core mechanisms that enable these capabilities. Our work addresses learning subtler and more complex underlying temporal dynamics in language modeling tasks that deal with sparse sequential data. To this end, we improve upon these recent advances by adopting concepts from the field of Bayesian statistics, namely, variational inference. Our proposed approach consists in treating the network parameters as latent variables with a prior distribution imposed over them. Our statistical assumptions go beyond the standard practice of postulating Gaussian priors. Indeed, to allow for handling outliers, which are prevalent in long observed sequences of multivariate data, multivariate t -exponential distributions are imposed. On this basis, we proceed to infer corresponding posteriors; these can be used for inference and prediction at test time, in a way that accounts for the uncertainty in the available sparse training data. Specifically, to allow for our approach to best exploit the merits of the t -exponential family, our method considers a new t -divergence measure, which generalizes the concept of the Kullback-Leibler divergence. We perform an extensive experimental evaluation of our approach, using challenging language modeling benchmarks, and illustrate its superiority over existing state-of-the-art techniques. Kyriakos Tolias, Sotirios Chatzis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Indian Buffet Process Deep Generative Models for Semi-Supervised ClassificationabstractDeep generative models (DGMs) have brought about a major breakthrough, as well as renewed interest, in generative latent variable models. However, DGMs do not allow for performing data-driven inference of the number of latent features needed to represent the observed data. Traditional linear formulations address this issue by resorting to tools from the field of nonparametric statistics. Indeed, linear latent variable models imposed an Indian Buffet Process (IBP) prior have been extensively studied by the machine learning community; inference for such models can been performed either via exact sampling or via approximate variational techniques. Based on this inspiration, in this paper we examine whether similar ideas from the field of Bayesian nonparametrics can be utilized in the context of modern DGMs in order to address the latent variable dimensionality inference problem. To this end, we propose a novel DGM formulation, based on the imposition of an IBP prior. We devise an efficient Black-Box Variational inference algorithm for our model, and exhibit its efficacy in a number of semi-supervised classification experiments. In all cases, we use popular benchmark datasets, and compare to state-of-the-art DGMs. Sotirios Chatzis |
ICASSP | 1 |
| 2018 | A Recurrent Latent Variable Model for Supervised Modeling of High-Dimensional Sequential DataabstractIn this work, we attempt to ameliorate the impact of data sparsity in the context of supervised modeling applications dealing with high-dimensional sequential data. Specifically, we seek to devise a machine learning mechanism capable of extracting subtle and complex underlying temporal dynamics in the observed sequential data, so as to inform the predictive algorithm. To this end, we improve upon systems that utilize deep learning techniques with recurrently connected units; we do so by adopting concepts from the field of Bayesian statistics, namely variational inference. Our proposed approach consists in treating the network recurrent units as stochastic latent variables with a prior distribution imposed over them. On this basis, we proceed to infer corresponding posteriors; these can be used for prediction generation, in a way that accounts for the uncertainty in the available sparse training data. To allow for our approach to easily scale to large real-world datasets, we perform inference under an approximate amortized variational inference (AVI) setup, whereby the learned posteriors are parameterized via (conventional) neural networks. We perform an extensive experimental evaluation of our approach using challenging benchmark datasets, and illustrate its superiority over existing state-of-the-art techniques. Panayiotis Christodoulou, Sotirios Chatzis, Andreas S. Andreou |
INISTA | 2 |
| 2018 | Forecasting stock market crisis events using deep and statistical machine learning techniques
Sotirios Chatzis, Vasilis Siakoulis, Anastasios Petropoulos, Evangelos Stavroulakis, Nikos E. Vlachogiannakis |
Expert Syst. Appl. | 1 |
| 2018 | Deep learning with t-exponential Bayesian kitchen sinks
Harris Partaourides, Sotirios Chatzis |
Expert Syst. Appl. | 2 |
| 2018 | Latent subspace modeling of sequential data under the maximum entropy discrimination framework
Sotirios Chatzis |
Neurocomputing | 1 |
| 2017 | Recurrent latent variable conditional heteroscedasticityabstractGeneralized autoregressive conditional heteroscedasticity (GARCH) models have long been considered as one of the most successful families of approaches for volatility modeling in financial return signals. However, this family of methods employ quite rigid assumptions regarding the evolution of the variance. In this paper, we address these issues by introducing a recurrent latent variable model, capable of capturing highly flexible functional relationships for the variances. We derive a fast, scalable, and robust to overfitting Bayesian inference algorithm, by relying on amortized variational inference. This avoids the need to compute per-data point variational parameters, but can instead compute a set of global variational parameters valid for inference at both training and test time. We evaluate the efficacy of our approach in a number of benchmarks, and compare its performance to state-of-the-art methodologies. Sotirios Chatzis |
ICASSP | 1 |
| 2017 | Deep Bayesian Matrix Factorization
Sotirios Chatzis |
PAKDD (2) | 1 |
| 2017 | Deep Network Regularization via Bayesian Inference of Synaptic Connectivity
Harris Partaourides, Sotirios Chatzis |
PAKDD (1) | 2 |
| 2017 | A stacked generalization system for automated FOREX portfolio trading
Anastasios Petropoulos, Sotirios Chatzis, Vasilis Siakoulis, Nikos E. Vlachogiannakis |
Expert Syst. Appl. | 2 |
| 2017 | Asymmetric deep generative models
Harris Partaourides, Sotirios Chatzis |
Neurocomputing | 2 |
| 2017 | A hidden Markov model with dependence jumps for predictive modeling of multidimensional time-series
Anastasios Petropoulos, Sotirios Chatzis, Stelios Xanthopoulos |
Inf. Sci. | 2 |
| 2016 | A novel corporate credit rating system based on Student's-t hidden Markov models
Anastasios Petropoulos, Sotirios Chatzis, Stylianos Z. Xanthopoulos |
Expert Syst. Appl. | 2 |
| 2016 | Maximum entropy discrimination factor analyzers
Sotirios Chatzis |
Neurocomputing | 1 |
| 2016 | Software defect prediction using doubly stochastic Poisson processes driven by stochastic belief networks
Andreas S. Andreou, Sotirios Chatzis |
J. Syst. Softw. | 2 |
| 2015 | Inducing Space Dirichlet Process Mixture Large-Margin Entity RelationshipInference in Knowledge BasesabstractIn this paper, we focus on the problem of extending a given knowledge base by accurately predicting additional true facts based on the facts included in it. This is an essential problem of knowledge representation systems, since knowledge bases typically suffer from incompleteness and lack of ability to reason over their discrete entities and relationships. To achieve our goals, in our work we introduce an inducing space nonparametric Bayesian large-margin inference model, capable of reasoning over relationships between pairs of entities. Previous works addressing the entity relationship inference problem model each entity based on atomic entity vector representations. In contrast, our method exploits word feature vectors to directly obtain high-dimensional nonlinear inducing space representations for entity pairs. This way, we allow for extracting salient latent characteristics and interaction dynamics within entity pairs that can be useful for inferring their relationships. On this basis, our model performs the relations inference task by postulating a set of binary Dirichlet process mixture large-margin classifiers, presented with the derived inducing space representations of the considered entity pairs. Bayesian inference for this inducing space model is performed under the mean-field inference paradigm. This is made possible by leveraging a recently proposed latent variable formulation of regularized large-margin classifiers that facilitates mean-field parameter estimation. We exhibit the superiority of our approach over the state-of-the-art by considering the problem of predicting additional true relations between entities given subsets of the WordNet and FreeBase knowledge bases. Sotirios Chatzis |
CIKM | 1 |
| 2015 | A Nonparametric Bayesian Approach toward Stacked Convolutional Independent Component AnalysisabstractUnsupervised feature learning algorithms based on convolutional formulations of independent components analysis (ICA) have been demonstrated to yield state-of-the-art results in several action recognition benchmarks. However, existing approaches do not allow for the number of latent components (features) to be automatically inferred from the data in an unsupervised manner. This is a significant disadvantage of the state-of-the-art, as it results in considerable burden imposed on researchers and practitioners, who must resort to tedious cross-validation procedures to obtain the optimal number of latent features. To resolve these issues, in this paper we introduce a convolutional nonparametric Bayesian sparse ICA architecture for overcomplete feature learning from high-dimensional data. Our method utilizes an Indian buffet process prior to facilitate inference of the appropriate number of latent features under a hybrid variational inference algorithm, scalable to massive datasets. As we show, our model can be naturally used to obtain deep unsupervised hierarchical feature extractors, by greedily stacking successive model layers, similar to existing approaches. In addition, inference for this model is completely heuristics-free, thus, it obviates the need of tedious parameter tuning, which is a major challenge most deep learning approaches are faced with. We evaluate our method on several action recognition benchmarks, and exhibit its advantages over the state-of-the-art. Sotirios Chatzis, Dimitrios I. Kosmopoulos |
ICCV | 1 |
| 2015 | Sparse Bayesian Recurrent Neural Networks
Sotirios Chatzis |
ECML/PKDD (2) | 1 |
| 2015 | Maximum Entropy Discrimination Poisson Regression for Software Reliability ModelingabstractReliably predicting software defects is one of the most significant tasks in software engineering. Two of the major components of modern software reliability modeling approaches are: 1) extraction of salient features for software system representation, based on appropriately designed software metrics and 2) development of intricate regression models for count data, to allow effective software reliability data modeling and prediction. Surprisingly, research in the latter frontier of count data regression modeling has been rather limited. More specifically, a lack of simple and efficient algorithms for posterior computation has made the Bayesian approaches appear unattractive, and thus underdeveloped in the context of software reliability modeling. In this paper, we try to address these issues by introducing a novel Bayesian regression model for count data, based on the concept of max-margin data modeling, effected in the context of a fully Bayesian model treatment with simple and efficient posterior distribution updates. Our novel approach yields a more discriminative learning technique, making more effective use of our training data during model inference. In addition, it allows of better handling uncertainty in the modeled data, which can be a significant problem when the training data are limited. We derive elegant inference algorithms for our model under the mean-field paradigm and exhibit its effectiveness using the publicly available benchmark data sets. Sotirios Chatzis, Andreas S. Andreou |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | A Latent Manifold Markovian Dynamics Gaussian ProcessabstractIn this paper, we propose a Gaussian process (GP) model for analysis of nonlinear time series. Formulation of our model is based on the consideration that the observed data are functions of latent variables, with the associated mapping between observations and latent representations modeled through GP priors. In addition, to capture the temporal dynamics in the modeled data, we assume that subsequent latent representations depend on each other on the basis of a hidden Markov prior imposed over them. Derivation of our model is performed by marginalizing out the model parameters in closed form using GP priors for observation mappings, and appropriate stick-breaking priors for the latent variable (Markovian) dynamics. This way, we eventually obtain a nonparametric Bayesian model for dynamical systems that accounts for uncertainty in the modeled data. We provide efficient inference algorithms for our model on the basis of a truncated variational Bayesian approximation. We demonstrate the efficacy of our approach considering a number of applications dealing with real-world data, and compare it with the related state-of-the-art approaches. Sotirios Chatzis, Dimitrios I. Kosmopoulos |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Dynamic Bayesian Probabilistic Matrix FactorizationabstractCollaborative filtering algorithms generally rely on the assumption that user preference patterns remain stationary. However, real-world relational data are seldom stationary. User preference patterns may change over time, giving rise to the requirement of designing collaborative filtering systems capable of detecting and adapting to preference pattern shifts. Motivated by this observation, in this paper we propose a dynamic Bayesian probabilistic matrix factorization model, designed for modeling time-varying distributions. Formulation of our model is based on imposition of a dynamic hierarchical Dirichlet process (dHDP) prior over the space of probabilistic matrix factorization models to capture the time-evolving statistical properties of modeled sequential relational datasets. We develop a simple Markov Chain Monte Carlo sampler to perform inference. We present experimental results to demonstrate the superiority of our temporal model. Sotirios Chatzis |
AAAI | 1 |
| 2014 | Echo-State Conditional Restricted Boltzmann MachinesabstractRestricted Boltzmann machines (RBMs) are a powerful generative modeling technique, based on a complex graphical model of hidden (latent) variables. Conditional RBMs (CRBMs) are an extension of RBMs tailored to modeling temporal data. A drawback of CRBMs is their consideration of linear temporal dependencies, which limits their capability to capture complex temporal structure. They also require many variables to model long temporal dependencies, a fact that might provoke overfitting proneness. To resolve these issues, in this paper we propose the echo-state CRBM (ES-CRBM): our model uses an echo-state network reservoir in the context of CRBMs to efficiently capture long and complex temporal dynamics, with much fewer trainable parameters compared to conventional CRBMs. In addition, we introduce an (implicit) mixture of ES-CRBM experts (im-ES-CRBM) to enhance even further the capabilities of our ES-CRBM model. The introduced im-ES-CRBM allows for better modeling temporal observations which might comprise a number of latent or observable subpatterns that alternate in a dynamic fashion. It also allows for performing sequence segmentation using our framework. We apply our methods to sequential data modeling and classification experiments using public datasets. As we show, our approach outperforms both existing RBM-based approaches as well as related state-of-the-art methods, such as conditional random fields. Sotirios Chatzis |
AAAI | 1 |
| 2014 | A Non-stationary Infinite Partially-Observable Markov Decision Process
Sotirios Chatzis, Dimitrios I. Kosmopoulos |
ICANN | 1 |
| 2014 | Gaussian Process-Mixture Conditional HeteroscedasticityabstractGeneralized autoregressive conditional heteroscedasticity (GARCH) models have long been considered as one of the most successful families of approaches for volatility modeling in financial return series. In this paper, we propose an alternative approach based on methodologies widely used in the field of statistical machine learning. Specifically, we propose a novel nonparametric Bayesian mixture of Gaussian process regression models, each component of which models the noise variance process that contaminates the observed data as a separate latent Gaussian process driven by the observed data. This way, we essentially obtain a Gaussian process-mixture conditional heteroscedasticity (GPMCH) model for volatility modeling in financial return series. We impose a nonparametric prior with power-law nature over the distribution of the model mixture components, namely the Pitman-Yor process prior, to allow for better capturing modeled data distributions with heavy tails and skewness. Finally, we provide a copula-based approach for obtaining a predictive posterior for the covariances over the asset returns modeled by means of a postulated GPMCH model. We evaluate the efficacy of our approach in a number of benchmark scenarios, and compare its performance to state-of-the-art methodologies. Emmanouil A. Platanios, Sotirios Chatzis |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Nonparametric bayesian multitask collaborative filteringabstractThe dramatic rates new digital content becomes available has brought collaborative filtering systems to the epicenter of computer science research in the last decade. One of the greatest challenges collaborative filtering systems are confronted with is the data sparsity problem: users typically rate only very few items; thus, availability of historical data is not adequate to effectively perform prediction. To alleviate these issues, in this paper we propose a novel multitask collaborative filtering approach. Our approach is based on a coupled latent factor model of the users rating functions, which allows for coming up with an agile information sharing mechanism that extracts much richer task-correlation information compared to existing approaches. Formulation of our method is based on concepts from the field of Bayesian nonparametrics, specifically Indian Buffet Process priors, which allow for data-driven determination of the optimal number of underlying latent features (item characteristics and user traits) assumed in the context of the model. We experiment on several real-world datasets, demonstrating both the efficacy of our method, and its superiority over existing approaches. Sotirios Chatzis |
CIKM | 1 |
| 2013 | Infinite Markov-Switching Maximum Entropy Discrimination MachinesabstractIn this paper, we present a method that combines the merits of Bayesian nonparametrics, specifically stick-breaking priors, and large-margin kernel machines in the context of sequential data classification. The proposed model postulates a set of (theoretically) infinite interdependent large-margin classifiers as model components, that robustly capture local nonlinearity of complex data. The postulated large-margin classifiers are connected in the context of a Markov-switching construction that allows for capturing complex temporal dynamics in the modeled datasets. Appropriate stick-breaking priors are imposed over the component switching mechanism of our model to allow for data-driven determination of the optimal number of component large-margin classifiers, under a standard nonparametric Bayesian inference scheme. Efficient model training is performed under the maximum entropy discrimination (MED) framework, which integrates the large-margin principle with Bayesian posterior inference. We evaluate our method using several real-world datasets, and compare it to state-of-the-art alternatives. Sotirios Chatzis |
ICML (3) | 1 |
| 2013 | Margin-maximizing classification of sequential data with infinitely-long temporal dependencies
Sotirios Chatzis |
Expert Syst. Appl. | 1 |
| 2013 | Corrigendum to "Maximum-margin classification of sequential data with infinitely-long temporal dependencies" [Expert Systems with Applications 40 (11) (2013) 4519-4527]
Sotirios Chatzis |
Expert Syst. Appl. | 1 |
| 2013 | A latent variable Gaussian process model with Pitman-Yor process priors for multiclass classification
Sotirios Chatzis |
Neurocomputing | 1 |
| 2013 | The Infinite-Order Conditional Random Field Model for Sequential Data ModelingabstractSequential data labeling is a fundamental task in machine learning applications, with speech and natural language processing, activity recognition in video sequences, and biomedical data analysis being characteristic examples, to name just a few. The conditional random field (CRF), a log-linear model representing the conditional distribution of the observation labels, is one of the most successful approaches for sequential data labeling and classification, and has lately received significant attention in machine learning as it achieves superb prediction performance in a variety of scenarios. Nevertheless, existing CRF formulations can capture only one- or few-timestep interactions and neglect higher order dependences, which are potentially useful in many real-life sequential data modeling applications. To resolve these issues, in this paper we introduce a novel CRF formulation, based on the postulation of an energy function which entails infinitely long time-dependences between the modeled data. Building blocks of our novel approach are: 1) the sequence memoizer (SM), a recently proposed nonparametric Bayesian approach for modeling label sequences with infinitely long time dependences, and 2) a mean-field-like approximation of the model marginal likelihood, which allows for the derivation of computationally efficient inference algorithms for our model. The efficacy of the so-obtained infinite-order CRF (CRF(∞)) model is experimentally demonstrated. Sotirios Chatzis, Yiannis Demiris |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | A Markov random field-regulated Pitman-Yor process prior for spatially constrained data clustering
Sotirios Chatzis |
Pattern Recognit. | 1 |
| 2013 | A conditional random field-based model for joint sequence segmentation and classification
Sotirios Chatzis, Dimitrios I. Kosmopoulos, Paul Doliotis |
Pattern Recognit. | 1 |
| 2012 | A sparse nonparametric hierarchical Bayesian approach towards inductive transfer for preference modeling
Sotirios Chatzis, Yiannis Demiris |
Expert Syst. Appl. | 1 |
| 2012 | The echo state conditional random field model for sequential data modeling
Sotirios Chatzis, Yiannis Demiris |
Expert Syst. Appl. | 1 |
| 2012 | A spatially-constrained normalized Gamma process prior
Sotirios Chatzis, Dimitrios Korkinof, Yiannis Demiris |
Expert Syst. Appl. | 1 |
| 2012 | The copula echo state network
Sotirios Chatzis, Yiannis Demiris |
Pattern Recognit. | 1 |
| 2012 | A reservoir-driven non-stationary hidden Markov model
Sotirios Chatzis, Yiannis Demiris |
Pattern Recognit. | 1 |
| 2012 | A possibilistic clustering approach toward generative mixture models
Sotirios Chatzis, Gavriil Tsechpenakis |
Pattern Recognit. | 1 |
| 2012 | Visual Workflow Recognition Using a Variational Bayesian Treatment of Multistream Fused Hidden Markov ModelsabstractIn this paper, we provide a variational Bayesian (VB) treatment of multistream fused hidden Markov models (MFHMMs), and apply it in the context of active learning-based visual workflow recognition (WR). Contrary to training methods yielding point estimates, such as maximum likelihood or maximum a posteriori training, the VB approach provides an estimate of the posterior distribution over the MFHMM parameters. As a result, our approach provides an elegant solution toward the amelioration of the overfitting issues of point estimate-based methods. Additionally, it provides a measure of confidence in the accuracy of the learned model, thus allowing for the easy and cost-effective utilization of active learning in the context of MFHMMs. Two alternative active learning algorithms are considered in this paper: query by committee, which selects unlabeled data that minimize the classification variance, and a maximum information gain method that aims to maximize the alteration in model variance by proper data labeling. We demonstrate the efficacy of the proposed treatment of MFHMMs by examining two challenging WR scenarios, and show that the application of active learning, which is facilitated by our VB approach, allows for a significant reduction of the MFHMM training costs. Sotirios Chatzis, Dimitrios I. Kosmopoulos |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Nonparametric Mixtures of Gaussian Processes With Power-Law BehaviorabstractGaussian processes (GPs) constitute one of the most important Bayesian machine learning approaches, based on a particularly effective method for placing a prior distribution over the space of regression functions. Several researchers have considered postulating mixtures of GPs as a means of dealing with nonstationary covariance functions, discontinuities, multimodality, and overlapping output signals. In existing works, mixtures of GPs are based on the introduction of a gating function defined over the space of model input variables. This way, each postulated mixture component GP is effectively restricted in a limited subset of the input space. In this paper, we follow a different approach. We consider a fully generative nonparametric Bayesian model with power-law behavior, generating GPs over the whole input space of the learned task. We provide an efficient algorithm for model inference, based on the variational Bayesian framework, and prove its efficacy using benchmark and real-world datasets. Sotirios Chatzis, Yiannis Demiris |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | A Quantum-Statistical Approach Toward Robot Learning by DemonstrationabstractStatistical machine learning approaches have been at the epicenter of the ongoing research work in the field of robot learning by demonstration over the past few years. One of the most successful methodologies used for this purpose is a Gaussian mixture regression (GMR). In this paper, we propose an extension of GMR-based learning by demonstration models to incorporate concepts from the field of quantum mechanics. Indeed, conventional GMR models are formulated under the notion that all the observed data points can be assigned to a distinct number of model states (mixture components). In this paper, we reformulate GMR models, introducing some quantum states constructed by superposing conventional GMR states by means of linear combinations. The so-obtained quantum statistics-inspired mixture regression algorithm is subsequently applied to obtain a novel robot learning by demonstration methodology, offering a significantly increased quality of regenerated trajectories for computational costs comparable with currently state-of-the-art trajectory-based robot learning by demonstration approaches. We experimentally demonstrate the efficacy of the proposed approach. Sotirios Chatzis, Dimitrios Korkinof, Yiannis Demiris |
IEEE Trans. Robotics | 1 |
| 2011 | The One-Hidden Layer Non-parametric Bayesian Kernel MachineabstractIn this paper, we present a nonparametric Bayesian approach towards one-hidden-layer feed forward neural networks. Our approach is based on a random selection of the weights of the synapses between the input and the hidden layer neurons, and a Bayesian marginalization over the weights of the connections between the hidden layer neurons and the output neurons, giving rise to a kernel-based nonparametric Bayesian inference procedure for feed forward neural networks. Compared to existing approaches, our method presents a number of advantages, with the most significant being: (i) it offers a significant improvement in terms of the obtained generalization capabilities, (ii) being a nonparametric Bayesian learning approach, it entails inference instead of fitting to data, thus resolving the over fitting issues of non-Bayesian approaches, and (iii) it yields a full predictive posterior distribution, thus naturally providing a measure of uncertainty on the generated predictions (expressed by means of the variance of the predictive distribution), without the need of applying computationally intensive methods, e.g., bootstrap. We exhibit the merits of our approach by investigating its application to two difficult multimedia content classification applications: semantic characterization of audio scenes based on content, and yearly song classification, as well as a set of benchmark classification and regression tasks. Sotirios Chatzis, Dimitrios Korkinof, Yiannis Demiris |
ICTAI | 1 |
| 2011 | Deformable probability maps: Probabilistic shape and appearance-based object segmentation
Gavriil Tsechpenakis, Sotirios Chatzis |
Comput. Vis. Image Underst. | 2 |
| 2011 | A fuzzy c-means-type algorithm for clustering of data with mixed numeric and categorical attributes employing a probabilistic dissimilarity functional
Sotirios Chatzis |
Expert Syst. Appl. | 1 |
| 2011 | Numerical optimization using synergetic swarms of foraging bacterial populations
Sotirios Chatzis, Spyridon Koukas |
Expert Syst. Appl. | 1 |
| 2011 | A variational Bayesian methodology for hidden Markov models utilizing Student's-t mixtures
Sotirios Chatzis, Dimitrios I. Kosmopoulos |
Pattern Recognit. | 1 |
| 2011 | Echo State Gaussian ProcessabstractEcho state networks (ESNs) constitute a novel approach to recurrent neural network (RNN) training, with an RNN (the reservoir) being generated randomly, and only a readout being trained using a simple computationally efficient algorithm. ESNs have greatly facilitated the practical application of RNNs, outperforming classical approaches on a number of benchmark tasks. In this paper, we introduce a novel Bayesian approach toward ESNs, the echo state Gaussian process (ESGP). The ESGP combines the merits of ESNs and Gaussian processes to provide a more robust alternative to conventional reservoir computing networks while also offering a measure of confidence on the generated predictions (in the form of a predictive distribution). We exhibit the merits of our approach in a number of applications, considering both benchmark datasets and real-world applications, where we show that our method offers a significant enhancement in the dynamical data modeling capabilities of ESNs. Additionally, we also show that our method is orders of magnitude more computationally efficient compared to existing Gaussian process-based methods for dynamical data modeling, without compromises in the obtained predictive performance. Sotirios Chatzis, Yiannis Demiris |
IEEE Trans. Neural Networks | 1 |
| 2010 | A method for training finite mixture models under a fuzzy clustering principle
Sotirios Chatzis |
Fuzzy Sets Syst. | 1 |
| 2010 | Hidden Markov Models with Nonelliptically Contoured State DensitiesabstractHidden Markov models (HMMs) are a popular approach for modeling sequential data comprising continuous attributes. In such applications, the observation emission densities of the HMM hidden states are typically modeled by means of elliptically contoured distributions, usually multivariate Gaussian or Student's-t densities. However, elliptically contoured distributions cannot sufficiently model heavy-tailed or skewed populations which are typical in many fields, such as the financial and the communication signal processing domain.Employing finite mixtures of such elliptically contoured distributions to model the HMM state densities is a common approach for the amelioration of these issues.Nevertheless, the nature of the modeled data often requires postulation of a large number of mixture components for each HMM state, which might have a negative effect on both model efficiency and the training data set's size required to avoid overfitting. To resolve these issues, in this paper, we advocate for the utilization ofa nonelliptically contoured distribution, the multivariate normal inverse Gaussian (MNIG) distribution, for modeling the observation densities of HMMs. As we experimentally demonstrate, our selection allows for more effective modeling of skewed and heavy-tailed populations in a simple and computationally efficient manner. Sotirios Chatzis |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | The infinite hidden Markov random field modelabstractHidden Markov random field (HMRF) models are widely used for image segmentation, as they appear naturally in problems where a spatially constrained clustering scheme is asked for. A major limitation of HMRF models concerns the automatic selection of the proper number of their states, i.e., the number of region clusters derived by the image segmentation procedure. Existing methods, including likelihood- or entropy-based criteria, and reversible Markov chain Monte Carlo methods, usually tend to yield noisy model size estimates while imposing heavy computational requirements. Recently, Dirichlet process (DP, infinite) mixture models have emerged in the cornerstone of nonparametric Bayesian statistics as promising candidates for clustering applications where the number of clusters is unknown a priori; infinite mixture models based on the original DP or spatially constrained variants of it have been applied in unsupervised image segmentation applications showing promising results. Under this motivation, to resolve the aforementioned issues of HMRF models, in this paper, we introduce a nonparametric Bayesian formulation for the HMRF model, the infinite HMRF model, formulated on the basis of a joint Dirichlet process mixture (DPM) and Markov random field (MRF) construction. We derive an efficient variational Bayesian inference algorithm for the proposed model, and we experimentally demonstrate its advantages over competing methodologies. Sotirios Chatzis, Gavriil Tsechpenakis |
IEEE Trans. Neural Networks | 1 |
| 2009 | The infinite Hidden Markov random field modelabstractDirichlet process (DP) mixture models have recently emerged in the cornerstone of nonparametric Bayesian statistics as promising candidates for clustering applications where the number of clusters is unknown a priori. Hidden Markov random field (HMRF) models are parametric statistical models widely used for image segmentation, as they appear naturally in problems where a spatially-constrained clustering scheme is asked for. A major limitation of HMRF models concerns the automatic selection of the proper number of their states, i.e. the number of segments derived by the image segmentation procedure. Typically, for this purpose, various likelihood based criteria are employed. Nevertheless, such methods often fail to yield satisfactory results, exhibiting significant overfitting proneness. Recently, higher order conditional random field models using potentials defined on superpixels have been considered as alternatives tackling these issues. Still, these models are in general computationally inefficient, a fact that limits their widespread adoption in practical applications. To resolve these issues, in this paper we introduce a novel, nonparametric Bayesian formulation for the HMRF model, the infinite HMRF model. We describe an efficient variational Bayesian inference algorithm for the proposed model, and we apply it to a series of image segmentation problems, demonstrating its advantages over existing methodologies. Sotirios Chatzis, Gavriil Tsechpenakis |
ICCV | 1 |
| 2009 | Robust Sequential Data Modeling Using an Outlier Tolerant Hidden Markov ModelabstractHidden Markov (chain) models using finite Gaussian mixture models as their hidden state distributions have been successfully applied in sequential data modeling and classification applications. Nevertheless, Gaussian mixture models are well known to be highly intolerant to the presence of untypical data within the fitting data sets used for their estimation. Finite Student's t-mixture models have recently emerged as a heavier-tailed, robust alternative to Gaussian mixture models, overcoming these hurdles. To exploit these merits of Student's t-mixture models in the context of a sequential data modeling setting, we introduce, in this paper, a novel hidden Markov model where the hidden state distributions are considered to be finite mixtures of multivariate Student's t-densities. We derive an algorithm for the model parameters estimation under a maximum likelihood framework, assuming full, diagonal, and factor-analyzed covariance matrices. The advantages of the proposed model over conventional approaches are experimentally demonstrated through a series of sequential data modeling applications. Sotirios Chatzis, Dimitrios I. Kosmopoulos, Theodora A. Varvarigou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Vision-based production of personalized video
Dimitrios I. Kosmopoulos, Anastasios Doulamis, Alexandros Makris, Nikolaos D. Doulamis, Sotirios Chatzis, Stuart E. Middleton |
Signal Process. Image Commun. | 5 |
| 2009 | Factor Analysis Latent Subspace Modeling and Robust Fuzzy Clustering Using t -DistributionsabstractFactor analysis is a latent subspace model commonly used for local dimensionality reduction tasks. Fuzzyc-means (FCM) type fuzzy clustering approaches are closely related to Gaussian mixture models (GMMs), and expectation-maximization (EM) like algorithms have been employed in fuzzy clustering with regularized objective functions. Student'st-mixture models (SMMs) have been proposed recently as an alternative to GMMs, resolving their outlier vulnerability problems. In this paper, we propose a novel FCM-type fuzzy clustering scheme providing two significant benefits when compared with the existing approaches. First, it provides a well-established observation space dimensionality reduction framework for fuzzy clustering algorithms based on factor analysis, allowing concurrent performance of fuzzy clustering and, within each cluster, local dimensionality reduction. Second, it exploits the outlier tolerance advantages of SMMs to provide a novel, soundly founded, nonheuristic, robust fuzzy clustering framework by introducing the effective means to incorporate the explicit assumption about student'st-distributed data into the fuzzy clustering procedure. This way, the proposed model yields a significant performance increase for the fuzzy clustering algorithm, as we experimentally demonstrate. Sotirios Chatzis, Theodora A. Varvarigou |
IEEE Trans. Fuzzy Syst. | 1 |
| 2008 | A robust approach towards sequential data modeling and its application in automatic gesture recognitionabstractHidden Markov models using finite Gaussian mixture models as their hidden state distributions have been applied in modeling of time series that result from various noisy signals. Nevertheless, Gaussian mixture models are well-known to be highly intolerant to the presence of outliers within the fitting sets used for their estimation. Finite Student's-i mixture models have recently emerged as a heavier-tailed, robust alternative to Gaussian mixture models, overcoming these hurdles. To exploit those merits of Student's-i mixture models, we introduce in this paper a novel hidden Markov chain model where the hidden state distributions are considered to be finite mixtures of multivariate Student's-i densities and we derive an algorithm for the model parameters estimation under a maximum likelihood framework. We apply this novel approach in automatic gesture recognition and we show that our model provides a substantial improvement in data representation performance and computational efficiency over the standard Gaussian model. Sotirios Chatzis, Dimitrios I. Kosmopoulos, Theodora A. Varvarigou |
ICASSP | 1 |
| 2008 | Managing service level agreement contracts in OGSA-based Grids
Antonis Litke, Kleopatra Konstanteli, Vassiliki Andronikou, Sotirios Chatzis, Theodora A. Varvarigou |
Future Gener. Comput. Syst. | 4 |
| 2008 | Robust fuzzy clustering using mixtures of Student's-t distributions
Sotirios Chatzis, Theodora A. Varvarigou |
Pattern Recognit. Lett. | 1 |
| 2008 | A Fuzzy Clustering Approach Toward Hidden Markov Random Field Models for Enhanced Spatially Constrained Image SegmentationabstractHidden Markov random field (HMRF) models have been widely used for image segmentation, as they appear naturally in problems where a spatially constrained clustering scheme, taking into account the mutual influences of neighboring sites, is asked for. Fuzzyc-means (FCM) clustering has also been successfully applied in several image segmentation applications. In this paper, we combine the benefits of these two approaches, by proposing a novel treatment of HMRF models, formulated on the basis of a fuzzy clustering principle. We approach the HMRF model treatment problem as an FCM-type clustering problem, effected by introducing the explicit assumptions of the HMRF model into the fuzzy clustering procedure. Our approach utilizes a fuzzy objective function regularized by Kullback--Leibler divergence information, and is facilitated by application of a mean-field-like approximation of the MRF prior. We experimentally demonstrate the superiority of the proposed approach over competing methodologies, considering a series of synthetic and real-world image segmentation applications. Sotirios Chatzis, Theodora A. Varvarigou |
IEEE Trans. Fuzzy Syst. | 1 |
| 2006 | Video Representation and Retrieval Using Spatio-temporal Descriptors and Region Relations
Sotirios Chatzis, Anastasios Doulamis, Dimitrios I. Kosmopoulos, Theodora A. Varvarigou |
ICANN (2) | 1 |