Sotiris B. Kotsiantis

dblp:167/4988 · also Sotirios Kotsiantis, Sotiris Kotsiantis · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-2247-3082ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Reducing labeled data requirements in text classification with active learning and BERT-based transformers
abstract
Abstract It is a fact that natural language processing (NLP) has become an integral part of daily life, with research outcomes being integrated into various everyday implementations. A significant portion of this success can reasonably be attributed to the architecture of transformers. In this context, text classification problems constitute a large part of ongoing research. Simultaneously, there is a growing demand for high-quality labeled textual data. The latter is becoming increasingly urgent with the rising complexity and size of models. Based on this, the present work investigates the integration of active learning strategies into text classification problems using transformer-based models from the BERT family. Through an extensive experimental framework involving 10 datasets and 7 different BERT-based classifiers, we demonstrate that the incorporation of active learning in the context of text classification can significantly reduce the need for labeled data during the fine-tuning procedures. Specifically, our experimental results illustrate that without sacrificing model effectiveness–as measured by various evaluation metrics–we can achieve at least a 50% reduction in the dataset size in 70% of cases. Additionally, we show that the size of the dataset plays a crucial role in maintaining high performance levels.
Katerina Karanikola, Charalampos M. Liapis, Sotiris B. Kotsiantis
Neural Comput. Appl.3
2025 Combining normalizing flows with decision trees for interpretable unsupervised outlier detection
Vasilis Papastefanopoulos, Pantelis Linardatos, Sotiris B. Kotsiantis
Eng. Appl. Artif. Intell.3
2025 Enhancing sentiment analysis with distributional emotion embeddings
abstract
Sentiment classification tasks, such as emotion detection and sentiment analysis, are essential in modern natural language processing (NLP). Moreover, vector representation frameworks modeling semantic content underlie each state-of-the-art NLP algorithmic scheme. In sentiment classification, traditional methods often rely on such embedding vectors for semantic representation, yet they typically overlook the dynamic and sequential nature of emotions within textual data. In this work, we present a novel methodology that leverages the distributional patterns of emotions. An embedding framework that captures the inherent serial structure of emotional occurrences in text is introduced, modeling the interdependencies between emotion states as they unfold within a document. Our approach treats each sentence as an observation in a multivariate series of emotions, transforming the emotional flow of a text into a sequence of emotion strings. By applying distributional logic, emotion-based embeddings that represent both emotional and semantic information are derived. Through a comprehensive experimental framework, we demonstrate the effectiveness of these embeddings across various sentiment-related tasks, including emotion detection, irony identification, and hate speech classification, evaluated on multiple datasets. The results show that our distributional emotion embeddings significantly enhance the performance of sentiment classification models, offering improved generalization across diverse domains such as financial news and climate change discourse. Hence, this work highlights the potential of using distributional emotion embeddings to advance sentiment analysis, offering a more nuanced understanding of emotional language and its structured, context-dependent manifestations. • Introduction of distributional emotion embeddings, distinct from prior models. • Models emotions as evolving sequences for richer, context-aware sentiment analysis. • Compatible with any classification algorithm, from traditional to advanced. • Cross-domain generalization across various fields and datasets. • The proposed embeddings enhance any method used, demonstrating versatility.
Charalampos M. Liapis, Katerina Karanikola, Sotiris B. Kotsiantis
Neurocomputing3
2024 Data-efficient software defect prediction: A comparative analysis of active learning-enhanced models and voting ensembles
Charalampos M. Liapis, Katerina Karanikola, Sotiris B. Kotsiantis
Inf. Sci.3
2023 A multi-view-CNN framework for deep representation learning in image classification
Emmanuel Pintelas, Ioannis E. Livieris, Sotiris B. Kotsiantis, Panayiotis E. Pintelas
Comput. Vis. Image Underst.3
2023 A generic sparse regression imputation method for time series and tabular data
Athanasios Salamanis, George A. Gravvanis, Sotiris B. Kotsiantis, Konstantinos M. Giannoutakis
Knowl. Based Syst.3
2023 Classification of acoustical signals by combining active learning strategies with semi-supervised learning schemes
Stamatis Karlos, Christos K. Aridas, Vasileios G. Kanas, Sotiris B. Kotsiantis
Neural Comput. Appl.4
2023 A multivariate ensemble learning method for medium-term energy forecasting
abstract
Abstract In the contemporary context, both production and consumption of energy, being concepts intertwined through a condition of synchronicity, are pivotal for the orderly functioning of society, with their management being a building block in maintaining regularity. Hence, the pursuit to develop reliable computational tools for modeling such serial and time-dependent phenomena becomes similarly crucial. This paper investigates the use of ensemble learners for medium-term forecasting of the Greek energy system load using additional information from injected energy production from various sources. Through an extensive experimental process, over 435 regression schemes and 64 different modifications of the feature inputs were tested over five different prediction time frames, creating comparative rankings regarding two case studies: one related to methods and the other to feature setups. Evaluations according to six widely used metrics indicate an aggregate but clear dominance of a specific efficient and low-cost ensemble layout. In particular, an ensemble method that incorporates the orthogonal matching pursuit together with the Huber regressor according to an averaged combinatorial scheme is proposed. Moreover, it is shown that the use of multivariate setups improves the derived predictions.
Charalampos M. Liapis, Katerina Karanikola, Sotiris B. Kotsiantis
Neural Comput. Appl.3
2021 An adaptive cluster-based sparse autoregressive model for large-scale multi-step traffic forecasting
Athanasios Salamanis, Anastasia-Dimitra Lipitakis, George A. Gravvanis, Sotiris B. Kotsiantis, Dimosthenis Anagnostopoulos
Expert Syst. Appl.4
2021 A novel explainable image classification framework: case study on skin cancer and plant disease prediction
Emmanuel Pintelas, Meletis Liaskos, Ioannis E. Livieris, Sotiris B. Kotsiantis, Panayiotis E. Pintelas
Neural Comput. Appl.4
2021 Editorial: Applications of Fuzzy Systems in Data Science and Big Data
abstract
The papers in this special section focus on applications of fuzzy systems in data science and Big Data. In the era of Big Data, intelligent as well as fuzzy tools have become very important to understanding our changing Internet-of-Things-driven data-centric world. Extracting “intelligence” from massive amounts of data has allowed us to support decision making processes in many fields, ranging from common fields like medicine and engineering to more lucrative industries such as vehicular technology and environmental stressors. In line with the shift in many industries to analyzing big data, has added in the fuzzy perspective to allow new ways of reasoning. Examples include interpretability of computing schemes, which have become more and more complex. Yet, successful applications of biomathematical modeling in a fuzzy environment have shown a good alternative to mere common black-box tools. In the call for papers of this special issue, we had devoted our interest in research pertaining to the current state-of-the-art research in the application of fuzzy systems regarding data science and big data analytics. We actively solicited recent results mainly concerned with recent advances and challenges in the theory and applications of fuzzy systems in the fields of data sciences and big data environments.
Gautam Srivastava 0001, Jerry Chun-Wei Lin, Dragan Pamucar, Sotiris B. Kotsiantis
IEEE Trans. Fuzzy Syst.4
2020 An active learning ensemble method for regression tasks
abstract
Active learning is a typical approach for learning from both labeled and unlabeled examples aiming to build efficient and accurate predictive models at minimum expense under an expert’s guidance. Since there is a lack of labeled data in many scientific fields whilst, at the same time, the labeling cost of unlabeled data is typically high in terms of time and expenditure, active learning has grown rapidly over recent years with great success. This is reflected in various studies providing insights and analyzing several active learning methods, especially in the case of classification tasks, whereas, there is only a limited number of studies concerning the implementation of active learning methods for regression ones. Within this context, the present paper sets out to put forward a pool-based active learning regression algorithm employing the query by committee strategy to evaluate the informativeness of unlabeled examples. The experimental results on a plethora of benchmark datasets demonstrate the efficiency of the proposed method, since it prevails over the baseline active learning approach applying the random sampling strategy, as well as familiar supervised methods.
Nikos Fazakis, Georgios Kostopoulos, Stamatis Karlos, Sotiris B. Kotsiantis, Kyriakos N. Sgarbas
Intell. Data Anal.4
2019 A Deep Dense Neural Network for Bankruptcy Prediction
Stamatios-Aggelos N. Alexandropoulos, Christos K. Aridas, Sotiris B. Kotsiantis, Michael N. Vrahatis
EANN3
2019 Active learning Rotation Forest for multiclass classification
abstract
Abstract Although the achievements of the computer science field have facilitated the tasks of collecting, storing, and accessing vast amounts of data efficiently, its annotation still remains a non–easily resolvable problem since no automated mechanism can perform reliably enough. This fact is even more profound when the objective is the high quality of generalization ability. Active learning constitutes a scheme that is exploited for tackling such problems, controlling the demanded human effort under a trade‐off assumption concerning the achievement of higher accuracy rates. In this work, the well‐known Rotation Forest algorithm is integrated with the active learning theory for constructing a robust and accurate classifier. Thus, apart from exploiting the collected labeled data, unlabeled data mining takes place through appropriate queries, whereas human expert decisions over the most questionable of them enrich the initial ones. Comprehensive comparisons of the proposed algorithm against four distinct learners inside the same learning scheme were executed. Moreover, the baseline strategy of random sampling and the corresponding supervised scenarios were included. During the evaluation stage, 13 publicly available multiclass datasets were assessed, and the obtained results verified our assumptions, regarding also the significant supremacy of the proposed algorithm against the majority of its rivals.
Vangjel Kazllarof, Stamatis Karlos, Sotiris B. Kotsiantis
Comput. Intell.3
2019 A multi-scheme semi-supervised regression approach
Nikos Fazakis, Stamatis Karlos, Sotiris B. Kotsiantis, Kyriakos N. Sgarbas
Pattern Recognit. Lett.3
2017 Random Resampling in the One-Versus-All Strategy for Handling Multi-class Problems
Christos K. Aridas, Stamatios-Aggelos N. Alexandropoulos, Sotiris B. Kotsiantis, Michael N. Vrahatis
EANN3
2017 Using Active Learning Methods for Predicting Fraudulent Financial Statements
Stamatis Karlos, Georgios Kostopoulos, Sotiris B. Kotsiantis, Vassilis Tampakas
EANN3
2017 Predicting Student Performance in Distance Higher Education Using Active Learning
Georgios Kostopoulos, Anastasia-Dimitra Lipitakis, Sotiris B. Kotsiantis, George A. Gravvanis
EANN3
2015 Self-Train LogitBoost for Semi-supervised Learning
Stamatis Karlos, Nikos Fazakis, Sotiris B. Kotsiantis, Kyriakos N. Sgarbas
EANN3
2015 Predicting Student Performance in Distance Higher Education Using Semi-supervised Techniques
Georgios Kostopoulos, Sotiris B. Kotsiantis, Panayiotis E. Pintelas
MEDI2
2012 Integrating Global and Local Voting of Classifiers
abstract
Many data mining problems involve an investigation of the relationships between features in heterogeneous data sets, where different learning algorithms can be more appropriate for different regions. The author proposes herein a technique of integrating global and local voting of classifiers. A comparison with other well-known combining methods on standard benchmark data sets was performed, and the accuracy of the proposed method was greater.
Sotiris B. Kotsiantis
Cybern. Syst.1
2011 Cascade Generalization with Reweighting Data for Handling Imbalanced Problems
abstract
Many data sets exhibit skewed class distributions in which most cases are allocated to a class and far fewer cases to a smaller one. A classifier induced from an imbalanced data set has usually a low error rate for the majority class and an unacceptable error rate for the minority class. This paper provides a review on various methodologies that have tried to handle this problem. Afterwards, it presents an experimental study of these methodologies with a proposed cascade generalization ensemble that is applied in reweighted data and it concludes that such a framework can be a more effective solution to the problem. Our method improves the identification of a difficult small class, while keeping the classification ability of the other class in an acceptable level.
Sotiris B. Kotsiantis
Comput. J.1
2010 A combinational incremental ensemble of classifiers as a technique for predicting students' performance in distance education
Sotiris B. Kotsiantis, Kiriakos Patriarcheas, Michalis Nik Xenos
Knowl. Based Syst.1
2007 Towards an ontology-based system for intelligent prediction of firms with fraudulent financial statements
abstract
Existing models for facilitating predictions of firms with fraudulent financial statements (FFS) are not semantics-based. The objective of this work was to design an intelligent web portal to serve as service provider for predicting which Greek manufacturing firms issue FFS. The portal serves as a meeting point among manufacturing firms, which are listed on the Athens Stock Exchange (ASE). It has been conceived to help users (e.g., chartered accountants) to detect FFS. The information contained in the portal is dynamically updated and it is related with financial statements. The portal helps users to find Greek manufacturing firms that provide FFS. For this purpose, the knowledge of the formal financial statements domain has been represented by means of ontology, which has been used to guide the design of the application and to supply the system with semantic capabilities. Additionally, the ontological component allows for defining an ontology-guided search engine. The reasoning engine of the proposed system executes logic rules related with well-established financial variables, which characterize a financial statement as fraudulent or not.
Dimitris Kanellopoulos, Sotiris B. Kotsiantis, Vassilis Tampakas
ETFA2
2006 Financial Application of Neural Networks: Two Case Studies in Greece
Sotiris B. Kotsiantis, Euaggelos Koumanakos, Dimitris Tzelepis, Vassilis Tampakas
ICANN (2)1
2005 Predicting Students' Marks in Hellenic Open University
abstract
The ability to provide assistance for a student at the appropriate level is invaluable in the learning process. Not only does it aids the student's learning process but also prevents problems, such as student frustration and floundering. Students' key demographic characteristics and their marks in a small number of written assignments can constitute the training set for a regression method in order to predict the student's performance. The scope of this work compares some of the state of the art regression algorithms in the application domain of predicting students' marks. A number of experiments have been conducted with six algorithms, which were trained using datasets provided by the Hellenic Open University. Finally, a prototype version of software support tool for tutors has been constructed implementing the M5rules algorithm, which proved to be the most appropriate among the tested algorithms.
Sotiris B. Kotsiantis, Panayiotis E. Pintelas
ICALT1
2005 Local Bagging of Decision Stumps
Sotiris B. Kotsiantis, George E. Tsekouras, Panayiotis E. Pintelas
IEA/AIE1
2005 A Fuzzy Logic-Based Approach for Detecting Shifting Patterns in Cross-Cultural Data
George E. Tsekouras, Dimitris Papageorgiou, Sotiris B. Kotsiantis, Christos Kalloniatis, Panayiotis E. Pintelas
IEA/AIE3
2003 Preventing Student Dropout in Distance Learning Using Machine Learning Techniques
Sotiris B. Kotsiantis, Christos Pierrakeas, Panayiotis E. Pintelas
KES1