Christos Koutlis

dblp:164/5225 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-3682-408XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization
abstract
With the rapid advancement of sophisticated synthetic audio-visual content, e.g., for subtle malicious manipulations, ensuring the integrity of digital media has become paramount. This work presents a novel approach to temporal localization of deepfakes by leveraging Audio-Visual Speech Representation Reconstruction (AuViRe). Specifically, our approach reconstructs speech representations from one modality (e.g., lip movements) based on the other (e.g., audio waveform). Cross-modal reconstruction is significantly more challenging in manipulated video segments, leading to amplified discrepancies, thereby providing robust discriminative cues for precise temporal forgery localization. AuViRe outperforms the state of the art by +8.9 [email protected] on LAV-DF, +9.6 [email protected] on AV-Deepfake1M, and +5.1 AUC on an in-the-wild experiment. Code at https://github.com/mever-team/auvire.
Christos Koutlis, Symeon Papadopoulos
WACV1
2025 MAVias: Mitigate any Visual Bias
abstract
Mitigating biases in computer vision models is an essential step towards the trustworthiness of artificial intelligence models. Existing bias mitigation methods focus on a small set of predefined biases, limiting their applicability in visual datasets where multiple, possibly unknown biases exist. To address this limitation, we introduce MAVias, an open-set bias mitigation approach leveraging foundation models to discover spurious associations between visual attributes and target classes. MAVias first captures a wide variety of visual features in natural language via a foundation image tagging model, and then leverages a large language model to select those visual features defining the target class, resulting in a set of language-coded potential visual biases. We then translate this set of potential biases into vision-language embeddings and introduce an in-processing bias mitigation approach to prevent the model from encoding information related to them. Our experiments on diverse datasets, including CelebA, Waterbirds, ImageNet, and UrbanCars, show that MAVias effectively detects and mitigates a wide range of biases in visual recognition tasks outperforming current state-of-the-art.
Ioannis Sarridis, Christos Koutlis, Symeon Papadopoulos, Christos Diou
ICCV2
2025 Similarity Over Factuality: Are we Making Progress on Multimodal Out-of-Context Misinformation Detection?
abstract
Out-of-context (OOC) misinformation poses a significant challenge in multimodal fact-checking, where images are paired with texts that misrepresent their original context to support false narratives. Recent research in evidence-based OOC detection has seen a trend towards increasingly complex architectures, incorporating Transformers, foundation models, and large language models. In this study, we introduce a simple yet robust baseline, which assesses MUltimodal SimilaritiEs (MUSE), specifically the similarity between image-text pairs and external image and text evidence. Our results demonstrate that MUSE, when used with conventional classifiers like Decision Tree, Random Forest, and Multilayer Perceptron, can compete with and even surpass the state-of-the-art on the NewsCLIPpings and VERITE datasets at less than 1 % of their computational complexity. Furthermore, integrating MUSE in our proposed “Attentive Intermediate Transformer Representations” (AITR) significantly improved performance, by 3.3% and 7.5% on NewsCLIP-pings and VERITE, respectively. Nevertheless, the success of MUSE, relying on surface-level patterns and short-cuts, without examining factuality and logical inconsistencies, raises critical questions about how we define the task, construct datasets, collect external evidence and over-all, how we assess progress in the field. We release our code at: https://github.com/stevejpapad/outcontext-misinfo-progress.
Stefanos I. Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Panagiotis Petrantonakis
WACV2
2025 InDistill: Information flow-preserving knowledge distillation for model compression
abstract
In this paper, we introduce InDistill, a method that serves as a warmup stage for enhancing Knowledge Distillation (KD) effectiveness. InDistill focuses on transferring criti-cal information flow paths from a heavyweight teacher to a lightweight student. This is achieved via a training scheme based on curriculum learning that considers the distillation difficulty of each layer and the critical learning peri-ods when the information flow paths are established. This procedure can lead to a student model that is better pre-pared to learn from the teacher. To ensure the applicability of InDstill across a wide range of teacher-student pairs, we also incorporate a pruning operation when there is a discrepancy in the width of the teacher and student layers. This pruning operation reduces the width of the teacher's intermediate layers to match those of the student, allowing direct distillation without the need for an encoding stage. The proposed method is extensively evaluated using various pairs of teacher-student architectures on CIFAR-10, CIFAR-100, and ImageNet datasets demonstrating that preserving the information flow paths consistently increases the per-formance of the baseline KD approaches on both classi-fication and retrieval settings. The code is available at https://github.com/gsarridis/InDistill.
Ioannis Sarridis, Christos Koutlis, Giorgos Kordopatis-Zilos, Ioannis Kompatsiaris, Symeon Papadopoulos
WACV2
2025 FLAC: Fairness-Aware Representation Learning by Suppressing Attribute-Class Associations
abstract
Bias in computer vision systems can perpetuate or even amplify discrimination against certain populations. Considering that bias is often introduced by biased visual datasets, many recent research efforts focus on training fair models using such data. However, most of them heavily rely on the availability of protected attribute labels in the dataset, which limits their applicability, while label-unaware approaches, i.e., approaches operating without such labels, exhibit considerably lower performance. To overcome these limitations, this work introduces FLAC, a methodology that minimizes mutual information between the features extracted by the model and a protected attribute, without the use of attribute labels. To do that, FLAC proposes a sampling strategy that highlights underrepresented samples in the dataset, and casts the problem of learning fair representations as a probability matching problem that leverages representations extracted by a bias-capturing classifier. It is theoretically shown that FLAC can indeed lead to fair representations, that are independent of the protected attributes. FLAC surpasses the current state-of-the-art on Biased-MNIST, CelebA, and UTKFace, by 29.1%, 18.1%, and 21.9%, respectively. Additionally, FLAC exhibits 2.2% increased accuracy on ImageNet-A and up to 4.2% increased accuracy on Corrupted-Cifar10. Finally, in most experiments, FLAC even outperforms the bias label-aware state-of-the-art methods.
Ioannis Sarridis, Christos Koutlis, Symeon Papadopoulos, Christos Diou
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 RED-DOT: Multimodal Fact-Checking viaRelevant Evidence Detection
abstract
Online misinformation is often multimodal in nature, i.e., caused by misleading associations between texts and accompanying images. To support the fact-checking process, researchers have been recently developing automatic multimodal methods that gather and analyze external information, evidence, related to the image–text pairs under examination. However, prior works incorrectly assumed that all external information collected from the Web is relevant. In this study, we introduce a “relevant evidence detection” (RED) module to discern whether each piece of evidence is relevant, to support or refute the claim. Specifically, we develop the “relevant evidence detection directed transformer” (RED-DOT) and explore multiple architectural variants (e.g., single or dual-stage) and mechanisms (e.g., “guided attention”). Extensive ablation and comparative experiments demonstrate that RED-DOT outperforms the state-of-the-art (SotA), achieving up to 33.7% accuracy improvement on the VERITE benchmark. Furthermore, our evidence reranking and element-wise modality fusion led to RED-DOT surpassing the SotA on NewsCLIPpings+ by up to 3% without the need for numerous evidence or multiple backbone encoders. We release our code at:https://github.com/stevejpapad/relevant-evidencedetection.
Stefanos I. Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Panagiotis Petrantonakis
IEEE Trans. Comput. Soc. Syst.2
2024 Leveraging Representations from Intermediate Encoder-Blocks for Synthetic Image Detection
Christos Koutlis, Symeon Papadopoulos
ECCV (72)1
2024 SDFD: Building a Versatile Synthetic Face Image Dataset with Diverse Attributes
abstract
AI systems rely on extensive training on large datasets to address various tasks. However, image-based systems, particularly those used for demographic attribute prediction, face significant challenges. Many current face image datasets primarily focus on demographic factors such as age, gender, and skin tone, overlooking other crucial facial attributes like hairstyle and accessories. This narrow focus limits the diversity of the data and consequently the robustness of AI systems trained on them. This work aims to address this limitation by proposing a methodology for generating synthetic face image datasets that capture a broader spectrum of facial diversity. Specifically, our approach integrates a systematic prompt formulation strategy, encompassing not only demographics and biometrics but also non-permanent traits like make-up, hairstyle, and accessories. These prompts guide a state-of-the-art text-to-image model in generating a comprehensive dataset of high-quality realistic images and can be used as an evaluation set in face analysis systems. Compared to existing datasets, our proposed dataset proves equally or more challenging in image classification tasks while being much smaller in size.
Georgia Baltsou, Ioannis Sarridis, Christos Koutlis, Symeon Papadopoulos
FG3
2024 A CLIP-based Siamese Approach for Meme Classification
abstract
Memes are an increasingly prevalent element of online discourse in social networks, especially among young audiences. They carry ideas and messages that range from humorous to hateful, and are widely consumed. Their potentially high impact requires adequate means of control to moderate their use in large scale. In this work, we propose SimCLIP a deep learning-based architecture for cross-modal understanding of memes, leveraging a pre-trained CLIP encoder to produce context-aware embeddings and a Siamese fusion technique to capture the interactions between text and image. We perform an extensive experimentation on seven meme classification tasks across six datasets. We establish a new state of the art in Memotion7k with a 7.25% relative F1-score improvement, and achieve super-human performance on Harm-P with 13.73% F1-Score improvement. Our approach demonstrates the potential for compact meme classification models, enabling accurate and efficient meme monitoring. We share our code at https://github.com/jahuerta92/meme-classification-simclip.
Javier Huertas-Tato, Christos Koutlis, Symeon Papadopoulos, David Camacho, Ioannis Kompatsiaris
IJCNN2
2024 FairBranch: Mitigating Bias Transfer in Fair Multi-task Learning
abstract
The generalisation capacity of Multi-Task Learning (MTL) suffers when unrelated tasks negatively impact each other by updating shared parameters with conflicting gradients. This is known as negative transfer and leads to a drop in MTL accuracy compared to single-task learning (STL). Lately, there has been a growing focus on the fairness of MTL models, requiring the optimization of both accuracy and fairness for individual tasks. Analogously to negative transfer for accuracy, task-specific fairness considerations might adversely affect the fairness of other tasks when there is a conflict of fairness loss gradients between the jointly learned tasks - we refer to this as bias transfer. To address both negative- and bias-transfer in MTL, we propose a novel method called FairBranch, which branches the MTL model by assessing the similarity of learned parameters, thereby grouping related tasks to alleviate negative transfer. Moreover, it incorporates fairness loss gradient conflict correction between adjoining task-group branches to address bias transfer within these task groups. Our experiments on tabular and visual MTL problems show that FairBranch outperforms state-of-the-art MTLs on both fairness and accuracy. Our code is available on github.com/arjunroyihrpa/FairBranch
Arjun Roy 0001, Christos Koutlis, Symeon Papadopoulos, Eirini Ntoutsi
IJCNN2
2024 FaceX: Understanding Face Attribute Classifiers through Summary Model Explanations
abstract
EXplainable Artificial Intelligence (XAI) approaches are widely applied for identifying fairness issues in Artificial Intelligence (AI) systems. However, in the context of facial analysis, existing XAI approaches, such as pixel attribution methods, offer explanations for individual images, posing challenges in assessing the overall behavior of a model, which would require labor-intensive manual inspection of a very large number of instances and leaving to the human the task of drawing a general impression of the model behavior from the individual outputs. Addressing this limitation, we introduce FaceX, the first method that provides a comprehensive understanding of face attribute classifiers through summary model explanations. Specifically, FaceX leverages the presence of distinct regions across all facial images to compute a region-level aggregation of model activations, allowing for the visualization of the model's region attribution across 19 predefined regions of interest in facial images, such as hair, ears, or skin. Beyond spatial explanations, FaceX enhances interpretability by visualizing specific image patches with the highest impact on the model's decisions for each facial region within a test benchmark. Through extensive evaluation in various experimental setups, including scenarios with or without intentional biases and mitigation efforts on four benchmarks, namely CelebA, FairFace, CelebAMask-HQ, and Racial Faces in the Wild, FaceX demonstrates high effectiveness in identifying the models' biases.
Ioannis Sarridis, Christos Koutlis, Symeon Papadopoulos, Christos Diou
ICMR2
2024 Zero-Shot Content-Based Crossmodal Recommendation System
Federico D'Asaro, Sara De Luca, Lorenzo Bongiovanni, Giuseppe Rizzo 0002, Symeon Papadopoulos, Emmanouil Schinas, Christos Koutlis
Expert Syst. Appl.7
2023 MemeFier: Dual-stage Modality Fusion for Image Meme Classification
abstract
Hate speech is a societal problem that has significantly grown through the Internet. New forms of digital content such as image memes have given rise to spread of hate using multimodal means, being far more difficult to analyse and detect compared to the unimodal case. Accurate automatic processing, analysis and understanding of this kind of content will facilitate the endeavor of hindering hate speech proliferation through the digital world. To this end, we propose MemeFier, a deep learning-based architecture for fine-grained classification of Internet image memes, utilizing a dual-stage modality fusion module. The first fusion stage produces feature vectors containing modality alignment information that captures non-trivial connections between the text and image of a meme. The second fusion stage leverages the power of a Transformer encoder to learn inter-modality correlations at the token level and yield an informative representation. Additionally, we consider external knowledge as an additional input, and background image caption supervision as a regularizing component. Extensive experiments on three widely adopted benchmarks, i.e., Facebook Hateful Memes, Memotion7k and MultiOFF, indicate that our approach competes and in some cases surpasses state-of-the-art. Our code is available on GitHub1.
Christos Koutlis, Emmanouil Schinas, Symeon Papadopoulos
ICMR1
2023 VICTOR: Visual incompatibility detection with transformers and fashion-specific contrastive pre-training
Stefanos I. Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Ioannis Kompatsiaris
J. Vis. Commun. Image Represent.2
2019 Identification of Hidden Sources by Estimating Instantaneous Causality in High-Dimensional Biomedical Time Series
abstract
The study of connectivity patterns of a system's variables, such as multi-channel electroencephalograms (EEG), is of utmost importance towards a better understanding of its internal evolutionary mechanisms. Here, the problem of estimating the connectivity network from multivariate time series in the presence of prominent unobserved variables is addressed. The causality measure of partial mutual information from mixed embedding (PMIME), designed to estimate direct lag-causal effects in the presence of many observed variables, is adapted to estimate also zero-lag effects, the so-called instantaneous causality. We term the proposed advanced method, PMIME0. The estimation of instantaneous causality by PMIME0 is a signature of the presence of hidden source in the observed system, as demonstrated analytically in a toy model. It is further demonstrated that the PMIME0 identifies the true instantaneous with great accuracy in a variety of high-dimensional dynamical systems. The method is applied to EEG data with epileptiform discharges (EDs), and the results imply a strong impact of unobserved confounders during the EDs. This finding comes as a possible explanation for the increased levels of causality during epileptic seizures estimated by some measures affected by the presence of a common source.
Christos Koutlis, Vasilios K. Kimiskidis, Dimitris Kugiumtzis
Int. J. Neural Syst.1
2017 Dynamics of Epileptiform Discharges Induced by Transcranial Magnetic Stimulation in Genetic Generalized Epilepsy
abstract
OBJECTIVE: In patients with Genetic Generalized Epilepsy (GGE), transcranial magnetic stimulation (TMS) can induce epileptiform discharges (EDs) of varying duration. We hypothesized that (a) the ED duration is determined by the dynamic states of critical network nodes (brain areas) at the early post-TMS period, and (b) brain connectivity changes before, during and after the ED, as well as within the ED. METHODS: EEG recordings from two GGE patients were analyzed. For hypothesis (a), the characteristics of the brain dynamics at the early ED stage are measured with univariate and multivariate EEG measures and the dependence of the ED duration on these measures is evaluated. For hypothesis (b), effective connectivity measures are combined with network indices so as to quantify the brain network characteristics and identify changes in brain connectivity. RESULTS: A number of measures combined with specific channels computed on the first EEG segment post-TMS correlate with the ED duration. In addition, brain connectivity is altered from pre-ED to ED and post-ED and statistically significant changes were also detected across stages within the ED. CONCLUSION: ED duration is not purely stochastic, but depends on the dynamics of the post-TMS brain state. The brain network dynamics is significantly altered in the course of EDs.
Dimitris Kugiumtzis, Christos Koutlis, Alkiviadis Tsimpiris, Vasilios K. Kimiskidis
Int. J. Neural Syst.2
2015 Transcranial Magnetic Stimulation Combined with EEG Reveals Covert States of Elevated Excitability in the Human Epileptic Brain
abstract
BACKGROUND: Transcranial magnetic stimulation combined with electroencephalogram (TMS-EEG) can be used to explore the dynamical state of neuronal networks. In patients with epilepsy, TMS can induce epileptiform discharges (EDs) with a stochastic occurrence despite constant stimulation parameters. This observation raises the possibility that the pre-stimulation period contains multiple covert states of brain excitability some of which are associated with the generation of EDs. OBJECTIVE: To investigate whether the interictal period contains "high excitability" states that upon brain stimulation produce EDs and can be differentiated from "low excitability" states producing normal appearing TMS-EEG responses. METHODS: In a cohort of 25 patients with Genetic Generalized Epilepsies (GGE) we identified two subjects characterized by the intermittent development of TMS-induced EDs. The high-excitability in the pre-stimulation period was assessed using multiple measures of univariate time series analysis. Measures providing optimal discrimination were identified by feature selection techniques. The "high excitability" states emerged in multiple loci (indicating diffuse cortical hyperexcitability) and were clearly differentiated on the basis of 14 measures from "low excitability" states (accuracy = 0.7). CONCLUSION: In GGE, the interictal period contains multiple, quasi-stable covert states of excitability a class of which is associated with the generation of TMS-induced EDs. The relevance of these findings to theoretical models of ictogenesis is discussed.
Vasilios K. Kimiskidis, Christos Koutlis, Alkiviadis Tsimpiris, Reetta Kälviäinen, Philippe Ryvlin, Dimitris Kugiumtzis
Int. J. Neural Syst.2