EDBT 2026 Demo / reviewers in the wild / expert
Vasileios Iosifidis
dblp:166/6355
· DBLP profile ↗
18ranked-venue papers
10as first author
4since 2021 · last 2023
0000-0002-3005-4507ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 8 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | AdaCC: cumulative cost-sensitive boosting for imbalanced classificationabstractAbstract Class imbalance poses a major challenge for machine learning as most supervised learning models might exhibit bias towards the majority class and under-perform in the minority class. Cost-sensitive learning tackles this problem by treating the classes differently, formulated typically via a user-defined fixed misclassification cost matrix provided as input to the learner. Such parameter tuning is a challenging task that requires domain knowledge and moreover, wrong adjustments might lead to overall predictive performance deterioration. In this work, we propose a novel cost-sensitive boosting approach for imbalanced data that dynamically adjusts the misclassification costs over the boosting rounds in response to model’s performance instead of using a fixed misclassification cost matrix. Our method, called AdaCC, is parameter-free as it relies on the cumulative behavior of the boosting model in order to adjust the misclassification costs for the next boosting round and comes with theoretical guarantees regarding the training error. Experiments on 27 real-world datasets from different domains with high class imbalance demonstrate the superiority of our method over 12 state-of-the-art cost-sensitive boosting approaches exhibiting consistent improvements in different measures, for instance, in the range of [0.3–28.56%] for AUC, [3.4–21.4%] for balanced accuracy, [4.8–45%] for gmean and [7.4–85.5%] for recall. Vasileios Iosifidis, Symeon Papadopoulos, Bodo Rosenhahn, Eirini Ntoutsi |
Knowl. Inf. Syst. | 1 |
| 2022 | Multi-fairness Under Class-Imbalance
Arjun Roy 0001, Vasileios Iosifidis, Eirini Ntoutsi |
DS | 2 |
| 2022 | Parity-based cumulative fairness-aware boosting
Vasileios Iosifidis, Arjun Roy 0001, Eirini Ntoutsi |
Knowl. Inf. Syst. | 1 |
| 2021 | LSTM Based Sentiment Analysis for Cryptocurrency Prediction
Xin Huang 0005, Wenbin Zhang 0002, Xuejiao Tang, Jayachander Surbiryala, Vasileios Iosifidis, Zhen Liu 0017, Ji Zhang 0001 |
DASFAA (3) | 6 |
| 2020 | Using Machine Learning to Automate Mammogram Images AnalysisabstractBreast cancer is the second leading cause of cancer-related death after lung cancer in women. Early detection of breast cancer in X-ray mammography is believed to have effectively reduced the mortality rate since 1989. However, a relatively high false positive rate and a low specificity in mammography technology still exist. In this work, a computer-aided automatic mammogram analysis system is proposed to process the mammogram images and automatically discriminate them as either normal or cancerous, consisting of three consecutive image processing, feature selection, and image classification stages. In designing the system, the discrete wavelet transforms (Daubechies 2, Daubechies 4, and Biorthogonal 6.8) and the Fourier cosine transform were first used to parse the mammogram images and extract statistical features. Then, an entropy-based feature selection method was implemented to reduce the number of features. Finally, different pattern recognition methods (including the Back-propagation Network, the Linear Discriminant Analysis, and the Naive Bayes Classifier) and a voting classification scheme were employed. The performance of each classification strategy was evaluated for sensitivity, specificity, and accuracy and for general performance using the Receiver Operating Curve. Our method is validated on the dataset from the Eastern Health in Newfoundland and Labrador of Canada. The experimental results demonstrated that the proposed automatic mammogram analysis system could effectively improve the classification performances. Xuejiao Tang, Liuhua Zhang, Wenbin Zhang 0002, Xin Huang 0005, Vasileios Iosifidis, Zhen Liu 0017, Enza Messina, Ji Zhang 0001 |
BIBM | 5 |
| 2020 | A Data-driven Human Responsibility Management SystemabstractAn ideal safe workplace is described as a place where staffs fulfill responsibilities in a well-organized order, potential hazardous events are being monitored in real-time, as well as the number of accidents and relevant damages are minimized. However, occupational-related death and injury are still increasing and have been highly attended in the last decades due to the lack of comprehensive safety management. A smart safety management system is therefore urgently needed, in which the staffs are instructed to fulfill responsibilities as well as automating risk evaluations and alerting staffs and departments when needed. In this paper, a smart system for safety management in the workplace based on responsibility big data analysis and the internet of things (IoT) are proposed. The real world implementation and assessment demonstrate that the proposed systems have superior accountability performance and improve the responsibility fulfillment through real-time supervision and self-reminder. Xuejiao Tang, Jiong Qiu, Wenbin Zhang 0002, Vasileios Iosifidis, Zhen Liu 0017, Ji Zhang 0001 |
IEEE BigData | 5 |
| 2020 | FairNN - Conjoint Learning of Fair Representations for Fair Decisions
Tongxin Hu, Vasileios Iosifidis, Wentong Liao, Michael Ying Yang, Eirini Ntoutsi, Bodo Rosenhahn |
DS | 2 |
| 2020 | FABBOO - Online Fairness-Aware Learning Under Class Imbalance
Vasileios Iosifidis, Eirini Ntoutsi |
DS | 1 |
| 2020 | Sentiment analysis on big sparse data streams with limited labels
Vasileios Iosifidis, Eirini Ntoutsi |
Knowl. Inf. Syst. | 1 |
| 2019 | FAE: A Fairness-Aware Ensemble FrameworkabstractAutomated decision making based on big data and machine learning (ML) algorithms can result in discriminatory decisions against certain protected groups defined upon personal data like gender, race, sexual orientation etc. Such algorithms designed to discover patterns in big data might not only pick up any encoded societal biases in the training data, but even worse, they might reinforce such biases resulting in more severe discrimination. The majority of thus far proposed fairness-aware machine learning approaches focus solely on the pre-, in- or post-processing steps of the machine learning process, that is, input data, learning algorithms or derived models, respectively. However, the fairness problem cannot be isolated to a single step of the ML process. Rather, discrimination is often a result of complex interactions between big data and algorithms, and therefore, a more holistic approach is required. The proposed FAE (Fairness-Aware Ensemble) framework combines fairness-related interventions at both pre-and post-processing steps of the data analysis process. In the pre-processing step, we tackle the problems of under-representation of the protected group (group imbalance) and of class-imbalance by generating balanced training samples. In the post-processing step, we tackle the problem of class overlapping by shifting the decision boundary in the direction of fairness. Vasileios Iosifidis, Besnik Fetahu, Eirini Ntoutsi |
IEEE BigData | 1 |
| 2019 | AdaFair: Cumulative Fairness Adaptive BoostingabstractThe widespread use of ML-based decision making in domains with high societal impact such as recidivism, job hiring and loan credit has raised a lot of concerns regarding potential discrimination. In particular, in certain cases it has been observed that ML algorithms can provide different decisions based on sensitive attributes such as gender or race and therefore can lead to discrimination. Although, several fairness-aware ML approaches have been proposed, their focus has been largely on preserving the overall classification accuracy while improving fairness in predictions for both protected and non-protected groups (defined based on the sensitive attribute(s)). The overall accuracy however is not a good indicator of performance in case of class imbalance, as it is biased towards the majority class. As we will see in our experiments, many of the fairness-related datasets suffer from class imbalance and therefore, tackling fairness requires also tackling the imbalance problem. To this end, we propose AdaFair, a fairness-aware classifier based on AdaBoost that further updates the weights of the instances in each boosting round taking into account a cumulative notion of fairness based upon all current ensemble members, while explicitly tackling class-imbalance by optimizing the number of ensemble members for balanced classification error. Our experiments show that our approach can achieve parity in true positive and true negative rates for both protected and non-protected groups, while it significantly outperforms existing fairness-aware methods up to 25% in terms of balanced error. Vasileios Iosifidis, Eirini Ntoutsi |
CIKM | 1 |
| 2019 | Fairness-Enhancing Interventions in Stream Classification
Vasileios Iosifidis, Thi Ngoc Han Tran, Eirini Ntoutsi |
DEXA (1) | 1 |
| 2018 | TweetsKB: A Public and Large-Scale RDF Corpus of Annotated Tweets
Pavlos Fafalios, Vasileios Iosifidis, Eirini Ntoutsi, Stefan Dietze |
ESWC | 2 |
| 2017 | Multi-aspect Entity-Centric Analysis of Big Social Media Archives
Pavlos Fafalios, Vasileios Iosifidis, Kostas Stefanidis, Eirini Ntoutsi |
TPDL | 2 |
| 2017 | Sentiment Classification over Opinionated Data Streams Through Informed Model Adaptation
Vasileios Iosifidis, Annina Oelschlager, Eirini Ntoutsi |
TPDL | 1 |
| 2017 | Large Scale Sentiment Learning with Limited LabelsabstractSentiment analysis is an important task in order to gain insights over the huge amounts of opinions that are generated in the social media on a daily basis. Although there is a lot of work on sentiment analysis, there are no many datasets available which one can use for developing new methods and for evaluation. To the best of our knowledge, the largest dataset for sentiment analysis is TSentiment [8], a 1.6 millions machine-annotated tweets dataset covering a period of about 3 months in 2009. This dataset however is too short and therefore insufficient to study heterogeneous, fast evolving streams. Therefore, we annotated the Twitter dataset of 2015 (228 million tweets without retweets and 275 million with retweets) and we make it publicly available for research. For the annotation we leverage the power of unlabeled data, together with labeled data using semi-supervised learning and in particular, Self-Learning and Co-Training. Our main contribution is the provision of the TSentiment15 dataset together with insights from the analysis, which includes a batch and a stream-processing of the data. In the former, all labeled and unlabeled data are available to the algorithms from the beginning, whereas in the later, they are revealed gradually based on their arrival time in the stream. Vasileios Iosifidis, Eirini Ntoutsi |
KDD | 1 |
| 2016 | Compressing Inverted Files using Modified LZW
Vasileios Iosifidis, Christos Makris 0001 |
WEBIST (1) | 1 |
| 2015 | Partial Order Preserving Encryption Search Trees
Kyriakos Ispoglou, Christos Makris 0001, Yannis C. Stamatiou, Elias C. Stavropoulos, Athanasios K. Tsakalidis, Vasileios Iosifidis |
DEXA (2) | 6 |