Konstantinos Makantasis

dblp:119/1632 · DBLP profile ↗
← Back
27ranked-venue papers
15as first author
12since 2021 · last 2026
0000-0002-0889-2766ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1Computer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Litter detection from aerial imagery: a review of UAV-based approaches and deep learning techniques
abstract
Abstract Litter pollution remains a pressing environmental issue, motivating the search for automated monitoring methods that can scale effectively. The problem is heightened by the difficulty of detecting small and diverse litter objects across wide areas, prompting interest in Unmanned Aerial Vehicles (UAVs) and deep learning as viable solutions. We conducted a systematic literature review by initially using Google Scholar, and then iteratively expanding our search through bibliographic references to identify relevant studies and datasets. In this review, we: synthesize the applicability of nine publicly available litter datasets; compile and analyze computer vision integrations with Litter Management Systems, with a particular emphasis towards UAV-based solutions; and review relevant literature addressing this issue; among others. Our analysis includes four UAV-based datasets (BDW, UAVVaste, HAIDA, SODA) and five non-UAV datasets (TrashNet, TACO, MJU-Waste, PlastOPol, ZeroWaste), examining dataset characteristics, preprocessing techniques, model architectures, and evaluation metrics across studies. Our synthesis of the literature highlights the varied approaches different studies undertake, reflecting the complexity of the task and the absence of standardised protocols. We conclude by discussing priorities for future work, notably the need for more publicly available in-the-wild UAV-acquired datasets and the potential of newer model architectures to address current limitations in automated litter detection.
Matthias Bartolo, Gabriel Hili, Dylan Seychell, Matthew Montebello, Carl J. Debono, Saviour Formosa, Konstantinos Makantasis
Multim. Tools Appl.7
2025 Dynamic Ensembles Towards Out-of-Distribution Generalization of Affect Models
Sean Vella Caruana, Athanasios Papathanasiou, Konstantinos Makantasis
ICANN (1)3
2025 Privileged Contrastive Pretraining for Multimodal Affect Modelling
abstract
Affective Computing (AC) has made significant progress with the advent of deep learning, yet a persistent challenge remains: the reliable transfer of affective models from controlled laboratory settings (in-vitro) to uncontrolled real-world environments (in-vivo). To address this challenge we introduce the Privileged Contrastive Pretraining (PriCon) framework according to which models are first pretrained via supervised contrastive learning (SCL) and then act as teacher models within a Learning Using Privileged Information (LUPI) framework. PriCon both leverages privileged information during training and enhances the robustness of derived affect models via SCL. Experiments conducted on two benchmark affective corpora, RECOLA and AGAIN, demonstrate that models trained using PriCon consistently outperform LUPI and end to end models. Remarkably, in many cases, PriCon models achieve performance comparable to models trained with access to all modalities during both training and testing. The findings underscore the potential of PriCon as a paradigm towards further bridging the gap between in-vitro and in-vivo affective modelling, offering a scalable and practical solution for real-world applications.
Kosmas Pinitas, Konstantinos Makantasis, Georgios N. Yannakakis
ICMI2
2024 Varying the Context to Advance Affect Modelling: A Study on Game Engagement Prediction
abstract
Affective computing faces a pressing challenge: the limited ability of affect models to generalise amidst varying contextual factors within the same task. While well recognised, this challenge persists due to the absence of suitable large-scale corpora with rich and diverse contextual information within a domain. To address this challenge, this paper introduces a GameVibe, a novel corpus explicitly tailored to confront the lack of contextual diversity. The affect corpus is sourced from 30 First Person Shooter (FPS) games, showcasing diverse game modes and designs within the same domain. The corpus comprises 2 hours of annotated gameplay videos with engagement levels annotated by a total of 20 participants in a time-continuous manner. Our preliminary analysis on this corpus sheds light on the complexity of generalising affect predictions across contextual variations in similar affective computing tasks. These initial findings serve as a catalyst for further research, inspiring deeper inquiries into this critical, yet understudied, aspect of affect modelling.
Kosmas Pinitas, Nemanja Rasajski, Matthew Barthet, Maria Kaselimi, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis
ACII5
2024 From the Lab to the Wild: Affect Modeling Via Privileged Information
abstract
How can we reliably transfer affect models trained in controlled laboratory conditions (in-vitro) to uncontrolled real-world settings (in-vivo)? The information gap between in-vitro and in-vivo applications defines a core challenge of affective computing. This gap is caused by limitations related to affect sensing including intrusiveness, hardware malfunctions and availability of sensors. As a response to these limitations, we introduce the concept of privileged information for operating affect models in real-world scenarios (in the wild). Privileged information enables affect models to be trained across multiple modalities available in a lab, and ignore, without significant performance drops, those modalities that are not available when they operate in the wild. Our approach is tested in two multimodal affect databases one of which is designed for testing models of affect in the wild. By training our affect models using all modalities and then using solely raw footage frames for testing the models, we reach the performance of models that fuse all available modalities for both training and testing. The results are robust across both classification and regression affect modeling tasks which are dominant paradigms in affective computing. Our findings make a decisive step towards realizing affect interaction in the wild.
Konstantinos Makantasis, Kosmas Pinitas, Antonios Liapis, Georgios N. Yannakakis
IEEE Trans. Affect. Comput.1
2023 The Pixels and Sounds of Emotion: General-Purpose Representations of Arousal in Games
abstract
What if emotion could be captured in a general and subject-agnostic fashion? Is it possible, for instance, to design general-purpose representations that detect affect solely from the pixels and audio of a human-computer interaction video? In this article we address the above questions by evaluating the capacity of deep learned representations to predict affect by relying only on audiovisual information of videos. We assume that the pixels and audio of an interactive session embed the necessary information required to detect affect. We test our hypothesis in the domain of digital games and evaluate the degree to which deep classifiers and deep preference learning algorithms can learn to predict the arousal of players based only on the video footage of their gameplay. Our results from four dissimilar games suggest that general-purpose representations can be built across games as the arousal models obtain average accuracies as high as 85 percent using the challenging leave-one-video-out cross-validation scheme. The dissimilar audiovisual characteristics of the tested games showcase the strengths and limitations of the proposed method.
Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis
IEEE Trans. Affect. Comput.1
2022 Learning Task-Independent Game State Representations from Unlabeled Images
abstract
Self-supervised learning (SSL) techniques have been widely used to learn compact and informative representations from high-dimensional complex data. In many computer vision tasks, such as image classification, such methods achieve state-of-the-art results that surpass supervised learning approaches. In this paper, we investigate whether SSL methods can be leveraged for the task of learning accurate state representations of games, and if so, to what extent. For this purpose, we collect game footage frames and corresponding sequences of games’ internal state from three different 3D games: VizDoom, the CARLA racing simulator and the Google Research Football Environment. We train an image encoder with three widely used SSL algorithms using solely the raw frames, and then attempt to recover the internal state variables from the learned representations. Our results across all three games showcase significantly higher correlation between SSL representations and the game’s internal state compared to pre-trained baseline models such as ImageNet. Such findings suggest that SSL-based visual encoders can yield general—not tailored to a specific task—yet informative game representations solely from game pixel information. Such representations can, in turn, form the basis for boosting the performance of downstream learning tasks in games, including gameplaying, content generation and player modeling.
Chintan Trivedi, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis
CoG2
2022 Game State Learning via Game Scene Augmentation
abstract
Having access to accurate game state information is of utmost importance for any artificial intelligence task including game-playing, testing, player modeling, and procedural content generation. Self-Supervised Learning (SSL) techniques have shown to be capable of inferring accurate game state information from the high-dimensional pixel input of game footage into compressed latent representations. Contrastive Learning is a popular SSL paradigm where the visual understanding of the game’s images comes from contrasting dissimilar and similar game states defined by simple image augmentation methods. In this study, we introduce a new game scene augmentation technique—named GameCLR—that takes advantage of the game-engine to define and synthesize specific, highly-controlled renderings of different game states, thereby, boosting contrastive learning performance. We test our GameCLR technique on images of the CARLA driving simulator environment and compare it against the popular SimCLR baseline SSL method. Our results suggest that GameCLR can infer the game’s state information from game footage more accurately compared to the baseline. Our proposed approach allows us to conduct game artificial intelligence research by directly utilizing screen pixels as input.
Chintan Trivedi, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis
FDG2
2022 RankNEAT: outperforming stochastic gradient search in preference learning tasks
abstract
Stochastic gradient descent (SGD) is a premium optimization method for training neural networks, especially for learning objectively defined labels such as image objects and events. When a neural network is instead faced with subjectively defined labels---such as human demonstrations or annotations---SGD may struggle to explore the deceptive and noisy loss landscapes caused by the inherent bias and subjectivity of humans. While neural networks are often trained via preference learning algorithms in an effort to eliminate such data noise, the de facto training methods rely on gradient descent. Motivated by the lack of empirical studies on the impact of evolutionary search to the training of preference learners, we introduce the RankNEAT algorithm which learns to rank through neuroevolution of augmenting topologies. We test the hypothesis that RankNEAT outperforms traditional gradient-based preference learning within the affective computing domain, in particular predicting annotated player arousal from the game footage of three dissimilar games. RankNEAT yields superior performances compared to the gradient-based preference learner (RankNet) in the majority of experiments since its architecture optimization capacity acts as an eficient feature selection mechanism, thereby, eliminating overfitting. Results suggest that RankNEAT is a viable and highly eficient evolutionary alternative to preference learning.
Kosmas Pinitas, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis
GECCO2
2022 Automatic Inspection of Cultural Monuments Using Deep and Tensor-Based Learning on Hyperspectral Imagery
abstract
In Cultural Heritage, hyperspectral images are commonly used since they provide extended information regarding the optical properties of materials. Thus, the processing of such high-dimensional data becomes challenging from the perspective of machine learning techniques to be applied. In this paper, we propose a Rank-R tensor-based learning model to identify and classify material defects on Cultural Heritage monuments. In contrast to conventional deep learning approaches, the proposed high order tensor-based learning demonstrates greater accuracy and robustness against over-fitting. Experimental results on real-world data from UNESCO protected areas indicate the superiority of the proposed scheme compared to conventional deep learning models.
Ioannis N. Tzortzis, Ioannis Rallis, Konstantinos Makantasis, Anastasios Doulamis, Nikolaos D. Doulamis, Athanasios Voulodimos
ICIP3
2022 Supervised Contrastive Learning for Affect Modelling
abstract
Affect modeling is viewed, traditionally, as the process of mapping measurable affect manifestations from multiple modalities of user input to affect labels. That mapping is usually inferred through end-to-end (manifestation-to-affect) machine learning processes. What if, instead, one trains general, subject-invariant representations that consider affect information and then uses such representations to model affect? In this paper we assume that affect labels form an integral part, and not just the training signal, of an affect representation and we explore how the recent paradigm of contrastive learning can be employed to discover general high-level affect-infused representations for the purpose of modeling affect. We introduce three different supervised contrastive learning approaches for training representations that consider affect information. In this initial study we test the proposed methods for arousal prediction in the RECOLA dataset based on user information from multiple modalities. Results demonstrate the representation capacity of contrastive learning and its efficiency in boosting the accuracy of affect models. Beyond their evidenced higher performance compared to end-to-end arousal classification, the resulting representations are general-purpose and subject-agnostic, as training is guided though general affect information available in any multimodal corpus.
Kosmas Pinitas, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis
ICMI2
2021 Privileged Information for Modeling Affect In The Wild
abstract
A key challenge of affective computing research is discovering ways to reliably transfer affect models that are built in the laboratory to real world settings, namely in the wild. The existing gap between in vitro and in vivo affect applications is mainly caused by limitations related to affect sensing including intrusiveness, hardware malfunctions, availability of sensors, but also privacy and security. As a response to these limitations in this paper we are inspired by recent advances in machine learning and introduce the concept of privileged information for operating affect models in the wild. The presence of privileged information enables affect models to be trained across multiple modalities available in a lab setting and ignore modalities that are not available in the wild with no significant drop in their modeling performance. The proposed privileged information framework is tested in a game arousal corpus that contains physiological signals in the form of heart rate and electrodermal activity, game telemetry, and pixels of footage from two dissimilar games that are annotated with arousal traces. By training our arousal models using all modalities (in vitro) and using solely pixels for testing the models (in vivo), we reach levels of accuracy obtained from models that fuse all modalities both for training and testing. The findings of this paper make a decisive step towards realizing affect interaction in the wild.
Konstantinos Makantasis, Dávid Melhárt, Antonios Liapis, Georgios N. Yannakakis
ACII1
2020 Space-Time Domain Tensor Neural Networks: An Application on Human Pose Classification
abstract
Recent advances in sensing technologies require the design and development of pattern recognition models capable of processing spatiotemporal data efficiently. In this study, we propose a spatially and temporally aware tensor-based neural network for human pose classification using three-dimensional skeleton data. Our model employs three novel components. First, an input layer capable of constructing highly discriminative spatiotemporal features. Second, a tensor fusion operation that produces compact yet rich representations of the data, and third, a tensor-based neural network that processes data representations in their original tensor form. Our model is end-to-end trainable and characterized by a small number of trainable parameters making it suitable for problems where the annotated data is limited. Experimental evaluation of the proposed model indicates that it can achieve state-of-the-art performance.
Konstantinos Makantasis, Athanasios Voulodimos, Anastasios Doulamis, Nikolaos Bakalos, Nikolaos D. Doulamis
ICPR1
2019 From Pixels to Affect: A Study on Games and Player Experience
abstract
Is it possible to predict the affect of a user just by observing her behavioral interaction through a video? How can we, for instance, predict a user's arousal in games by merely looking at the screen during play? In this paper we address these questions by employing three dissimilar deep convolutional neural network architectures in our attempt to learn the underlying mapping between video streams of gameplay and the player's arousal. We test the algorithms in an annotated dataset of 50 gameplay videos of a survival shooter game and evaluate the deep learned models' capacity to classify high vs low arousal levels. Our key findings with the demanding leave-one-video-out validation method reveal accuracies of over 78 % on average and 98% at best. While this study focuses on games and player experience as a test domain, the findings and methodology are directly relevant to any affective computing area, introducing a general and user-agnostic approach for modeling affect.
Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis
ACII1
2019 Fusing Level and Ruleset Features for Multimodal Learning of Gameplay Outcomes
abstract
Which features of a game influence the dynamics of players interacting with it? Can a level's architecture change the balance between two competing players, or is it mainly determined by the character classes and roles that players choose before the game starts? This paper assesses how quantifiable gameplay outcomes such as score, duration and features of the heatmap can be predicted from different facets of the initial game state, specifically the architecture of the level and the character classes of the players. Experiments in this paper explore how different representations of a level and class parameters in a shooter game affect a deep learning model which attempts to predict gameplay outcomes in a large corpus of simulated matches. Findings in this paper indicate that a few features of the ruleset (i.e. character class parameters) are the main drivers for the model's accuracy in all tested gameplay outcomes, but the levels (especially when processed) can augment the model.
Antonios Liapis, Daniel Karavolos, Konstantinos Makantasis, Konstantinos Sfikas, Georgios N. Yannakakis
CoG3
2019 Common Mode Patterns for Supervised Tensor Subspace Learning
abstract
In this work we propose a method for reducing the dimensionality of tensor objects in a binary classification framework. The proposed Common Mode Patterns method takes into consideration the labels' information, and ensures that tensor objects that belong to different classes do not share common features after the reduction of their dimensionality. We experimentally validate the proposed supervised subspace learning technique and compared it against Multilinear Principal Component Analysis using a publicly available hyper-spectral imaging dataset. Experimental results indicate that the proposed CMP method can efficiently reduce the dimensionality of tensor objects, while, at the same time, increasing the inter-class separability.
Konstantinos Makantasis, Anastasios Doulamis, Nikolaos D. Doulamis, Athanasios Voulodimos
ICASSP1
2019 Hyperspectral Image Classification with Tensor-Based Rank-R Learning Models
abstract
In this paper, we present a general tensor-based nonlinear classifier, the Rank-R Feedforward Neural Network (FNN). In the proposed model, which is an extension of the Rank-1 FNN classifier, the network weights are constrained to satisfy a rank-R Canonical Polyadic Decomposition. By allowing a rank-R, instead of a rank-1, Canonical Polyadic Decomposition of the weights, the learning capacity of the model can be increased, which contributes to avoiding underfitting problems. The effectiveness of the proposed model is scrutinized on a hyperspectral image classification experimental setting, since hyperspectral data can naturally be represented as tensor objects. Performance evaluation results indicate that the proposed model outperforms other state-of-the-art models, including deep learning ones, especially in cases where the number of available training samples is small.
Konstantinos Makantasis, Athanasios Voulodimos, Anastasios Doulamis, Nikolaos D. Doulamis, Ioannis Georgoulas
ICIP1
2018 Tensor-Based Nonlinear Classifier for High-Order Data Analysis
abstract
In this paper we propose a tensor-based nonlinear model for high-order data classification. The advantages of the proposed scheme are that (i) it significantly reduces the number of weight parameters, and hence of required training samples, and (ii) it retains the spatial structure of the input samples. The proposed model, called Rank-1 FNN, is based on a modification of a feedforward neural network (FNN), such that its weights satisfy the rank-1 canonical decomposition. We also introduce a new learning algorithm to train the model, and we evaluate the Rank-1 FNN on third-order hyperspectral data. Experimental results and comparisons indicate that the proposed model outperforms state of the art classification methods, including deep learning based ones, especially in cases with small numbers of available training samples.
Konstantinos Makantasis, Anastasios Doulamis, Nikolaos D. Doulamis, Antonis Nikitakis, Athanasios Voulodimos
ICASSP1
2018 Data-Driven Background Subtraction Algorithm for In-Camera Acceleration in Thermal Imagery
abstract
Detection of moving objects in videos is a crucial step toward successful surveillance and monitoring applications. A key component for such tasks is called background subtraction and tries to extract regions of interest from the image background for further processing or action. For this reason, its accuracy and real-time performance are of great significance. Although effective background subtraction methods have been proposed, only a few of them take into consideration the special characteristics of thermal imagery. In this paper, we propose a background subtraction scheme, which models the thermal responses of each pixel as a mixture of Gaussians with unknown number of components. Following a Bayesian approach, our method automatically estimates the mixture structure, while simultaneously it avoids over-/underfitting. The pixel density estimate is followed by an efficient and highly accurate updating mechanism, which permits our system to be automatically adapted to dynamically changing operation conditions. We propose a reference implementation of our method in reconfigurable hardware achieving both adequate performance and low-power consumption. Adopting a high-level synthesis design and demanding floating point arithmetic operations are mapped in reconfigurable hardware, demonstrating fast prototyping and on-field customization at the same time.
Konstantinos Makantasis, Antonis Nikitakis, Anastasios Doulamis, Nikolaos D. Doulamis, Ioannis Papaefstathiou
IEEE Trans. Circuits Syst. Video Technol.1
2018 Tensor-Based Classification Models for Hyperspectral Data Analysis
abstract
In this paper, we present tensor-based linear and nonlinear models for hyperspectral data classification and analysis. By exploiting the principles of tensor algebra, we introduce new classification architectures, the weight parameters of which satisfy the rank-1 canonical decomposition property. Then, we propose learning algorithms to train both linear and nonlinear classifiers. The advantages of the proposed classification approach are that: 1) it significantly reduces the number of weight parameters required to train the model (and thus the respective number of training samples); 2) it provides a physical interpretation of model coefficients on the classification output; and 3) it retains the spatial and spectral coherency of the input samples. The linear tensor-based model exploits the principles of logistic regression, assuming the rank-1 canonical decomposition property among its weights. For the nonlinear classifier, we propose a modification of a feedforward neural network (FNN), called rank-1 FNN, since its weights satisfy again the rank-1 canonical decomposition property. An appropriate learning algorithm is also proposed to train the network. Experimental results and comparisons with state-of-the-art classification methods, either linear (e.g., linear support vector machine) or nonlinear (e.g., deep learning), indicate the outperformance of the proposed scheme, especially in the cases where a small number of training samples is available.
Konstantinos Makantasis, Anastasios Doulamis, Nikolaos D. Doulamis, Antonis Nikitakis
IEEE Trans. Geosci. Remote. Sens.1
2016 A novel background subtraction scheme for in-camera acceleration in thermal imagery
Antonis Nikitakis, Ioannis Papaefstathiou, Konstantinos Makantasis, Anastasios Doulamis
DATE3
2016 Deep learning based human behavior recognition in industrial workflows
abstract
We consider the fully automated behavior understanding through visual cues in industrial environments. In contrast to most existing work, which relies on domain knowledge to construct complex handcrafted features from inputs, we exploit a Convolutional Neural Network (CNN), which is a type of deep model and can act directly on the raw inputs, to automate the process of feature construction. Although such models are limited to handle still 2D inputs, in this paper we appropriately transform video input to incorporate temporal information into each frame. This way our model hierarchically constructs features from both spatial and temporal dimensions. We apply our model in real-world environment, on data taken from Nissan factory, and it achieves superior performance without relying on handcrafted features.
Konstantinos Makantasis, Anastasios Doulamis, Nikolaos D. Doulamis, Konstantinos Psychas
ICIP1
2016 In the wild image retrieval and clustering for 3D cultural heritage landmarks reconstruction
Konstantinos Makantasis, Anastasios Doulamis, Nikolaos D. Doulamis, Marinos Ioannides
Multim. Tools Appl.1
2016 3D measures exploitation for a monocular semi-supervised fall detection system
Konstantinos Makantasis, Eftychios Protopapadakis, Anastasios Doulamis, Nikolaos D. Doulamis, Nikolaos F. Matsatsinis
Multim. Tools Appl.1
2016 Semi-supervised vision-based maritime surveillance system using fused visual attention maps
Konstantinos Makantasis, Eftychios Protopapadakis, Anastasios Doulamis, Nikolaos F. Matsatsinis
Multim. Tools Appl.1
2015 Deep supervised learning for hyperspectral data classification through convolutional neural networks
abstract
Spectral observations along the spectrum in many narrow spectral bands through hyperspectral imaging provides valuable information towards material and object recognition, which can be consider as a classification task. Most of the existing studies and research efforts are following the conventional pattern recognition paradigm, which is based on the construction of complex handcrafted features. However, it is rarely known which features are important for the problem at hand. In contrast to these approaches, we propose a deep learning based classification method that hierarchically constructs high-level features in an automated way. Our method exploits a Convolutional Neural Network to encode pixels' spectral and spatial information and a Multi-Layer Perceptron to conduct the classification task. Experimental results and quantitative validation on widely used datasets showcasing the potential of the developed approach for accurate hyperspectral data classification.
Konstantinos Makantasis, Konstantinos Karantzalos, Anastasios Doulamis, Nikolaos D. Doulamis
IGARSS1
2013 Safe urban growth: An integrated ICT solution for unstandardized and distributed information handling
abstract
In 2007, our planet became predominantly urban as for the first time, more than half of the world's population was living in cities. This tendency for urbanization is characterized by urban sprawl phenomena and excessive environmental pollution degrading citizens' living standards. Although intelligent and safe urban planning and management is a challenging task, latest ICT advances can offer promising solutions to the aforementioned problem by providing efficient and effective ways for handling large amounts of unstandardized and distributed information. In this paper we present an integrated, fully automatic ICT solution for acquiring, processing and representing heterogeneous, unstandardized and distributed information about underground gas pipeline networks, in order to provide a safe way for urban growth.
Yiannis Agadakos, Konstantinos Makantasis, Panagiotis Partsinevelos, Anastasios Doulamis
WOWMOM2