Aristeidis Sotiras

dblp:28/7425 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-0795-8820ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 8 since 2021
YearPublicationVenuePosition
2026 Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model Adaptation
abstract
Vision-language models (VLMs) such as CLIP demonstrate strong performance but struggle when adapted to downstream tasks. Prompt learning has emerged as an efficient and effective strategy to adapt VLMs while preserving their pre-trained knowledge. However, existing methods still lead to overfitting and degrade zero-shot generalization. To address this challenge, we propose an optimal transport (OT)-guided prompt learning framework that mitigates forgetting by preserving the structural consistency of feature distributions between pre-trained and fine-tuned models. Unlike conventional point-wise constraints, OT naturally captures cross-instance relationships and expands the feasible parameter space for prompt tuning, allowing a better tradeoff between adaptation and generalization. Our approach enforces joint constraints on both vision and text representations, ensuring a holistic feature alignment. Extensive experiments on benchmark datasets demonstrate that our simple yet effective method outperforms existing prompt learning strategies in base-to-novel generalization, cross-dataset evaluation, and domain generalization, without requiring additional augmentation or ensemble techniques.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Haiyu Wu, Aristeidis Sotiras, Yalin Wang 0001, Abolfazl Razi
WACV7
2026 QCResUNet: Joint subject-level and voxel-level segmentation quality prediction
abstract
Deep learning has made significant strides in automated brain tumor segmentation from magnetic resonance imaging (MRI) scans in recent years. However, the reliability of these tools is hampered by the presence of poor-quality segmentation outliers, particularly in out-of-distribution samples, making their implementation in clinical practice difficult. Therefore, there is a need for quality control (QC) to screen the quality of the segmentation results. Although numerous automatic QC methods have been developed for segmentation quality screening, most were designed for cardiac MRI segmentation, which involves a single modality and a single tissue type. Furthermore, most prior works only provided subject-level predictions of segmentation quality and did not identify erroneous parts segmentation that may require refinement. To address these limitations, we proposed a novel multi-task deep learning architecture, termed QCResUNet, which produces subject-level segmentation-quality measures as well as voxel-level segmentation error maps for each available tissue class. To validate the effectiveness of the proposed method, we conducted experiments on assessing its performance on evaluating the quality of two distinct segmentation tasks. First, we aimed to assess the quality of brain tumor segmentation results. For this task, we performed experiments on one internal (Brain Tumor Segmentation (BraTS) Challenge 2021, n=1,251) and two external datasets (BraTS Challenge 2023 in Sub-Saharan Africa Patient Population (BraTS-SSA), n=40; Washington University School of Medicine (WUSM), n=175). Specifically, we first performed a three-fold cross-validation on the internal dataset using segmentations generated by different methods at various quality levels, followed by an evaluation on the external datasets. Second, we aimed to evaluate the segmentation quality of cardiac Magnetic Resonance Imaging (MRI) data from the Automated Cardiac Diagnosis Challenge (ACDC, n=100). The proposed method achieved high performance in predicting subject-level segmentation-quality metrics and accurately identifying segmentation errors on a voxel basis. This has the potential to be used to guide human-in-the-loop feedback to improve segmentations in clinical settings.
Peijie Qiu, Satrajit Chakrabarty, Soumyendu Sekhar Ghosh, Aristeidis Sotiras
Medical Image Anal.5
2025 Sequence Complementor: Complementing Transformers for Time Series Forecasting with Learnable Sequences
abstract
Since its introduction, the transformer has shifted the development trajectory away from traditional models (e.g., RNN, MLP) in time series forecasting, which is attributed to its ability to capture global dependencies within temporal tokens. Follow-up studies have largely involved altering the tokenization and self-attention modules to better adapt Transformers for addressing special challenges like non-stationarity, channel-wise dependency, and variable correlation in time series. However, we found that the expressive capability of sequence representation is a key factor influencing Transformer performance in time forecasting after investigating several representative methods, where there is an almost linear relationship between sequence representation entropy and mean square error, with more diverse representations performing better. In this paper, we propose a novel attention mechanism with Sequence Complementors and prove feasible from an information theory perspective, where these learnable sequences are able to provide complementary information beyond current input to feed attention. We further enhance the Sequence Complementors via a diversification loss that is theoretically covered. The empirical evaluation of both long-term and short-term forecasting has confirmed its superiority over the recent state-of-the-art methods.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Aristeidis Sotiras, Yalin Wang 0001, Abolfazl Razi
AAAI6
2025 Multimodal Variational Autoencoder: A Barycentric View
abstract
Multiple signal modalities, such as vision and sounds, are naturally present in real-world phenomena. Recently, there has been growing interest in learning generative models, in particular variational autoencoder (VAE), for multimodal representation learning especially in the case of missing modalities. The primary goal of these models is to learn a modality-invariant and modality-specific representation that characterizes information across multiple modalities. Previous attempts at multimodal VAEs approach this mainly through the lens of experts, aggregating unimodal inference distributions with a product of experts (PoE), a mixture of experts (MoE), or a combination of both. In this paper, we provide an alternative generic and theoretical formulation of multimodal VAE through the lens of barycenter. We first show that PoE and MoE are specific instances of barycenters, derived by minimizing the asymmetric weighted KL divergence to unimodal inference distributions. Our novel formulation extends these two barycenters to a more flexible choice by considering different types of divergences. In particular, we explore the Wasserstein barycenter defined by the 2-Wasserstein distance, which better preserves the geometry of unimodal distributions by capturing both modality-specific and modality-invariant representations compared to KL divergence. Empirical studies on three multimodal benchmarks demonstrated the effectiveness of the proposed method.
Peijie Qiu, Sayantan Kumar, Xiwen Chen, Abolfazl Razi, Yalin Wang 0001, Aristeidis Sotiras
AAAI9
2025 Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image Analysis
abstract
While multiple instance learning (MIL) has shown to be a promising approach for histopathological whole slide image (WSI) analysis, its reliance on permutation invariance significantly limits its capacity to effectively uncover semantic correlations between instances within WSIs. Based on our empirical and theoretical investigations, we argue that approaches that are not permutation-invariant but better capture spatial correlations between instances can offer more effective solutions. In light of these findings, we propose a novel alternative to existing MIL for WSI analysis by learning to restore the order of instances from their randomly shuffled arrangement. We term this task as cracking an instance jigsaw puzzle problem, where semantic correlations between instances are uncovered. To tackle the instance jigsaw puzzles, we propose a novel Siamese network solution, which is theoretically justified by optimal transport theory. We validate the proposed method on WSI classification and survival prediction tasks, where the proposed method outperforms the recent state-of-the-art MIL competitors. The code is available at https://github.com/xiwenc1/MIL-JigsawPuzzles.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Xuanzhao Dong, Yalin Wang 0001, Abolfazl Razi, Aristeidis Sotiras
ICCV11
2025 FIC-TSC: Learning Time Series Classification with Fisher Information Constraint
abstract
Analyzing time series data is crucial to a wide spectrum of applications, including economics, online marketplaces, and human healthcare. In particular, time series classification plays an indispensable role in segmenting different phases in stock markets, predicting customer behavior, and classifying worker actions and engagement levels. These aspects contribute significantly to the advancement of automated decision-making and system optimization in real-world applications. However, there is a large consensus that time series data often suffers from domain shifts between training and test sets, which dramatically degrades the classification performance. Despite the success of (reversible) instance normalization in handling the domain shifts for time series regression tasks, its performance in classification is unsatisfactory. In this paper, we propose $\textit{FIC-TSC}$, a training framework for time series classification that leverages Fisher information as the constraint. We theoretically and empirically show this is an efficient and effective solution to guide the model converges toward flatter minima, which enhances its generalizability to distribution shifts. We rigorously evaluate our method on 30 UEA multivariate and 85 UCR univariate datasets. Our empirical results demonstrate the superiority of the proposed method over 14 recent state-of-the-art methods.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Yalin Wang 0001, Aristeidis Sotiras, Abolfazl Razi
ICML8
2025 How Effective Can Dropout Be in Multiple Instance Learning ?
abstract
Multiple Instance Learning (MIL) is a popular weakly-supervised method for various applications, with a particular interest in histological whole slide image (WSI) classification. Due to the gigapixel resolution of WSI, applications of MIL in WSI typically necessitate a two-stage training scheme: first, extract features from the pre-trained backbone and then perform MIL aggregation. However, it is well-known that this suboptimal training scheme suffers from "noisy" feature embeddings from the backbone and inherent weak supervision, hindering MIL from learning rich and generalizable features. However, the most commonly used technique (i.e., dropout) for mitigating this issue has yet to be explored in MIL. In this paper, we empirically explore how effective the dropout can be in MIL. Interestingly, we observe that dropping the top-k most important instances within a bag leads to better performance and generalization even under noise attack. Based on this key observation, we propose a novel MIL-specific dropout method, termed MIL-Dropout, which systematically determines which instances to drop. Experiments on five MIL benchmark datasets and two WSI datasets demonstrate that MIL-Dropout boosts the performance of current MIL methods with a negligible computational cost. The code is available at https://github.com/ChongQingNoSubway/MILDropout.
Peijie Qiu, Xiwen Chen, Zhangsihao Yang, Aristeidis Sotiras, Abolfazl Razi, Yalin Wang 0001
ICML5
2025 Active Source-Free Cross-Domain and Cross-Modality Adaptation for Volumetric Medical Image Segmentation by Image Sensitivity and Organ Heterogeneity Sampling
Peijie Qiu, Daniel S. Marcus, Aristeidis Sotiras
MICCAI (6)5
2025 SC-VAE: Sparse coding-based variational autoencoder with learned ISTA
abstract
Learning rich data representations from unlabeled data is a key challenge towards applying deep learning algorithms in downstream tasks. Several variants of variational autoencoders (VAEs) have been proposed to learn compact data representations by encoding high-dimensional data in a lower dimensional space. Two main classes of VAEs methods may be distinguished depending on the characteristics of the meta-priors that are enforced in the representation learning step. The first class of methods derives a continuous encoding by assuming a static prior distribution in the latent space. The second class of methods learns instead a discrete latent representation using vector quantization (VQ) along with a codebook. However, both classes of methods suffer from certain challenges, which may lead to suboptimal image reconstruction results. The first class suffers from posterior collapse, whereas the second class suffers from codebook collapse. To address these challenges, we introduce a new VAE variant, termed sparse coding-based VAE with learned ISTA (SC-VAE), which integrates sparse coding within variational autoencoder framework. The proposed method learns sparse data representations that consist of a linear combination of a small number of predetermined orthogonal atoms. The sparse coding problem is solved using a learnable version of the iterative shrinkage thresholding algorithm (ISTA). Experiments on two image datasets demonstrate that our model achieves improved image reconstruction results compared to state-of-the-art methods. Moreover, we demonstrate that the use of learned sparse code vectors allows us to perform downstream tasks like image generation and unsupervised image segmentation through clustering image patches. The code is available at https://github.com/sotiraslab/SC-VAE . • SC-VAE integrates sparse coding into the VAE framework, avoiding codebook collapse. • SC-VAE outperforms existing methods in image reconstruction and generalization tasks. • SC-VAE enables effective disentanglement and interpolation. • SC-VAE’s sparse code vectors enable clustering, enhancing unsupervised segmentation.
Peijie Qiu, Sung Min Ha, Abdalla Bani, Aristeidis Sotiras
Pattern Recognit.6
2024 DGR-MIL: Exploring Diverse Global Representation in Multiple Instance Learning for Whole Slide Image Classification
Xiwen Chen, Peijie Qiu, Aristeidis Sotiras, Abolfazl Razi, Yalin Wang 0001
ECCV (38)4
2024 TimeMIL: Advancing Multivariate Time Series Classification via a Time-aware Multiple Instance Learning
abstract
Deep neural networks, including transformers and convolutional neural networks (CNNs), have significantly improved multivariate time series classification (MTSC). However, these methods often rely on supervised learning, which does not fully account for the sparsity and locality of patterns in time series data (e.g., quantification of diseases-related anomalous points in ECG and abnormal detection in signal). To address this challenge, we formally discuss and reformulate MTSC as a weakly supervised problem, introducing a novel multiple-instance learning (MIL) framework for better localization of patterns of interest and modeling time dependencies within time series. Our novel approach, TimeMIL, formulates the temporal correlation and ordering within a time-aware MIL pooling, leveraging a tokenized transformer with a specialized learnable wavelet positional token. The proposed method surpassed 26 recent state-of-the-art MTSC methods, underscoring the effectiveness of the weakly supervised TimeMIL in MTSC. The code is available https://github.com/xiwenc1/TimeMIL.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Aristeidis Sotiras, Yalin Wang 0001, Abolfazl Razi
ICML6
2024 SelfReg-UNet: Self-Regularized UNet for Medical Image Segmentation
Xiwen Chen, Peijie Qiu, Mohammad Farazi, Aristeidis Sotiras, Abolfazl Razi, Yalin Wang 0001
MICCAI (8)5
2023 QCResUNet: Joint Subject-Level and Voxel-Level Prediction of Segmentation Quality
Peijie Qiu, Satrajit Chakrabarty, Soumyendu Sekhar Ghosh, Aristeidis Sotiras
MICCAI (4)5
2022 Multi-scale semi-supervised clustering of brain images: Deriving disease subtypes
Junhao Wen 0002, Erdem Varol, Aristeidis Sotiras, Zhijian Yang, Ganesh B. Chand, Güray Erus, Haochang Shou, Ahmed Abdulkadir, Gyujoon Hwang, Dominic B. Dwyer, Alessandro Pigoni, Paola Dazzan, René S. Kahn, Hugo G. Schnack, Marcus V. Zanetti, Eva M. Meisenzahl, Geraldo Filho Bussato, Benedicto Crespo-Facorro, Rafael Romero-Garcia, Christos Pantelis, Stephen J. Wood, Chuanjun Zhuo, Russell T. Shinohara, Yong Fan 0001, Ruben C. Gur, Raquel E. Gur, Theodore D. Satterthwaite, Nikolaos Koutsouleris, Daniel H. Wolf, Christos Davatzikos
Medical Image Anal.3
2020 MAGIC: Multi-scale Heterogeneity Analysis and Clustering for Brain Diseases
Junhao Wen 0002, Erdem Varol, Ganesh B. Chand, Aristeidis Sotiras, Christos Davatzikos
MICCAI (7)4
2018 Generative Discriminative Models for Multivariate Inference and Statistical Mapping in Medical Imaging
Erdem Varol, Aristeidis Sotiras, Christos Davatzikos
MICCAI (3)2
2017 A Discrete MRF Framework for Integrated Multi-Atlas Registration and Segmentation
Stavros Alchatzidis, Aristeidis Sotiras, Evangelia I. Zacharaki, Nikos Paragios
Int. J. Comput. Vis.2
2016 Structured Outlier Detection in Neuroimaging Studies with Minimal Convex Polytopes
Erdem Varol, Aristeidis Sotiras, Christos Davatzikos
MICCAI (1)2
2016 Abnormality Detection via Iterative Deformable Registration and Basis-Pursuit Decomposition
abstract
We present a generic method for automatic detection of abnormal regions in medical images as deviations from a normative data base. The algorithm decomposes an image, or more broadly a function defined on the image grid, into the superposition of a normal part and a residual term. A statistical model is constructed with regional sparse learning to represent normative anatomical variations among a reference population (e.g., healthy controls), in conjunction with a Markov random field regularization that ensures mutual consistency of the regional learning among partially overlapping image blocks. The decomposition is performed in a principled way so that the normal part fits well with the learned normative model, while the residual term absorbs pathological patterns, which may then be detected through a statistical significance test. The decomposition is applied to multiple image features from an individual scan, detecting abnormalities using both intensity and shape information. We form an iterative scheme that interleaves abnormality detection with deformable registration, gradually improving robustness of the spatial normalization and precision of the detection. The algorithm is evaluated with simulated images and clinical data of brain lesions, and is shown to achieve robust deformable registration and localize pathological regions simultaneously. The algorithm is also applied on images from Alzheimer's disease patients to demonstrate the generality of the method.
Güray Erus, Aristeidis Sotiras, Russell T. Shinohara, Christos Davatzikos
IEEE Trans. Medical Imaging3
2015 Graph-Based Motion-Driven Segmentation of the Carotid Atherosclerotique Plaque in 2D Ultrasound Sequences
Aimilia Gastounioti, Aristeidis Sotiras, Konstantina S. Nikita, Nikos Paragios
MICCAI (3)2
2015 Disentangling Disease Heterogeneity with Max-Margin Multiple Hyperplane Classifier
Erdem Varol, Aristeidis Sotiras, Christos Davatzikos
MICCAI (1)2
2014 Discrete Multi Atlas Segmentation using Agreement Constraints
Stavros Alchatzidis, Aristeidis Sotiras, Nikos Paragios
BMVC2
2013 Deformable Medical Image Registration: A Survey
abstract
Deformable image registration is a fundamental task in medical image processing. Among its most important applications, one may cite: 1) multi-modality fusion, where information acquired by different imaging devices or protocols is fused to facilitate diagnosis and treatment planning; 2) longitudinal studies, where temporal structural or anatomical changes are investigated; and 3) population modeling and statistical atlases used to study normal anatomical variability. In this paper, we attempt to give an overview of deformable registration methods, putting emphasis on the most recent advances in the domain. Additional emphasis has been given to techniques applied to medical images. In order to study image registration methods in depth, their main components are identified and studied independently. The most recent techniques are presented in a systematic fashion. The contribution of this paper is to provide an extensive account of registration techniques in a systematic manner.
Aristeidis Sotiras, Christos Davatzikos, Nikos Paragios
IEEE Trans. Medical Imaging1
2011 Efficient parallel message computation for MAP inference
abstract
First order Markov Random Fields (MRFs) have become a predominant tool in Computer Vision over the past decade. Such a success was mostly due to the development of efficient optimization algorithms both in terms of speed as well as in terms of optimality properties. Message passing algorithms are among the most popular methods due to their good performance for a wide range of pairwise potential functions (PPFs). Their main bottleneck is computational complexity. In this paper, we revisit message computation as a distance transformation using a more formal setting than [8] to generalize it to arbitrary PPFs. The method is based on [20] yielding accurate results for a specific class of PPFs and in most other cases a close approximation. The proposed algorithm is parallel and thus enables us to fully take advantage of the computational power of parallel processing architectures. The proposed scheme coupled with an efficient belief propagation algorithm [8] and implemented on a massively parallel coprocessor provides results as accurate as state of the art inference methods, though is in general one order of magnitude faster in terms of speed.
Stavros Alchatzidis, Aristeidis Sotiras, Nikos Paragios
ICCV2
2011 DRAMMS: Deformable registration via attribute matching and mutual-saliency weighting
Yangming Ou, Aristeidis Sotiras, Nikos Paragios, Christos Davatzikos
Medical Image Anal.2
2010 Simultaneous Geometric - Iconic Registration
Aristeidis Sotiras, Yangming Ou, Ben Glocker, Christos Davatzikos, Nikos Paragios
MICCAI (2)1
2009 Graphical Models and Deformable Diffeomorphic Population Registration Using Global and Local Metrics
Aristeidis Sotiras, Nikos Komodakis, Ben Glocker, Jean-François Deux, Nikos Paragios
MICCAI (1)1