VLDB 2026 Research / reviewers in the wild / expert
Sotirios A. Tsaftaris
dblp:14/613
· DBLP profile ↗
74ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0002-8795-9294ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 17 since 2021Artificial intelligence and machine learning · 17 · 11 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Count2Density: Crowd density estimation without location-level annotationsabstractCrowd density estimation is a well-known computer vision task aimed at estimating the density distribution of people in an image. The primary challenge in this domain is the reliance on fine-grained location-level annotations — i.e., points placed on top of each individual — to train deep networks. Collecting such detailed annotations is both tedious, time-consuming, and poses a significant barrier to scalability for real-world applications. To alleviate this burden, we present Count2Density : a novel pipeline designed to predict meaningful density maps containing quantitative spatial information using only count-level annotations (i.e., the total number of people) during training. To achieve this, Count2Density generates pseudo-density maps leveraging past predictions stored in a Historical Map Bank, thereby reducing confirmation bias. This bank is initialised using an unsupervised saliency estimator to provide an initial spatial prior and is iteratively updated with an Exponential Moving Average of predicted density maps. These pseudo-density maps are obtained by sampling locations from estimated crowd areas using a hypergeometric distribution, with the number of samplings determined by the count-level annotations. To further enhance the spatial awareness of the model and promote robust feature learning, we add a self-supervised contrastive spatial regulariser to encourage similar feature representations within crowded regions while maximising dissimilarity with background regions. Experimental results demonstrate that our approach significantly outperforms cross-domain adaptation methods and achieves better results than recent state-of-the-art approaches in semi-supervised settings across several datasets. Additional analyses validate the effectiveness of each individual component of our pipeline, including the self-supervised contrastive regulariser, confirming the ability of Count2Density to effectively retrieve spatial information from count-level annotations and enabling accurate subregion counting. Mattia Litrico, Michael P. Pound, Sotirios A. Tsaftaris, Sebastiano Battiato, Mario Valerio Giuffrida |
Pattern Recognit. | 4 |
| 2026 | Position Paper: Artificial Intelligence in Medical Image Analysis: Advances, Clinical Translation, and Emerging FrontiersabstractOver the past five years, artificial intelligence (AI) has introduced new models and methods for addressing the challenges associated with the broader adoption of AI models and systems in medicine. This paper reviews recent advances in AI for medical image and video analysis, outlines emerging paradigms, highlights pathways for successful clinical translation, and provides recommendations for future work. Hybrid Convolutional Neural Network (CNN) Transformer architectures now deliver state-of-the-art results in segmentation, classification, reconstruction, synthesis, and registration. Foundation and generative AI models enable the use of transfer learning to smaller datasets with limited ground truth. Federated learning supports privacy-preserving collaboration across institutions. Explainable and trustworthy AI approaches have become essential to foster clinician trust, ensure regulatory compliance, and facilitate ethical deployment. Together, these developments pave the way for integrating AI into radiology, pathology, and wider healthcare workflows. Andreas Panayides, Hao Chen 0011, Nenad Filipovic, Tijana Geroski, Junlin Hou, Karim Lekadir, Kostas Marias, George K. Matsopoulos, Giorgos Papanastasiou, Pinaki Sarder, Georgia D. Tourassi, Sotirios A. Tsaftaris, Huazhu Fu, Efthyvoulos C. Kyriacou, Christos P. Loizou, Michalis E. Zervakis, Joel H. Saltz, Farah Shamout, Ken C. L. Wong, Jianhua Yao 0001, Amir A. Amini, Dimitrios I. Fotiadis, Constantinos S. Pattichis, Marios S. Pattichis |
IEEE J. Biomed. Health Informatics | 12 |
| 2025 | Towards Surgical Task Automation: Actor-Critic Models Meet Self-Supervised Imitation Learning*abstractSurgical robot task automation has recently attracted great attention due to its potential to benefit both surgeons and patients. Reinforcement learning (RL) based approaches have demonstrated promising ability to perform automated surgical manipulations on various tasks. To address the exploration challenge, expert demonstrations can be utilized to enhance the learning efficiency via imitation learning (IL) approaches. However, the successes of such methods normally rely on both states and action labels. Unfortunately, action labels can be hard to capture or their manual annotation is prohibitively expensive owing to the requirement for expert knowledge. Emulating expert behaviour using noisy or inaccurate labels poses significant risks, including unintended surgical errors that may result in patient discomfort or, in more severe cases, tissue damage. It therefore remains an appealing and open problem to leverage expert data composed of pure states into RL. In this work, we present an actor-critic RL framework, termed AC-SSIL, to overcome this challenge of improving learning process with state-only demonstrations collected by an unknown expert policy. It adopts a self-supervised IL method, dubbed SSIL, to effectively incorporate expert states into RL paradigms by retrieving from demonstrations the nearest neighbours of the query state and utilizing the bootstrapping of actor networks. It applies similarity-based regularization and improves its prediction capacity jointly with the actor network. We showcase through experiments on an open-source surgical simulation platform that our method delivers remarkable improvements over the RL baseline and exhibits comparable performance against action based IL methods, which implies the efficacy and potential of our method for expert demonstration-guided learning scenarios. Code will be made publicly available at https://github.com/Jingshuai-cqu/AC-SSIL. Jingshuai Liu, Alain Andres, Yonghang Jiang, Yuning Du, Wenmiao Shu, Can Pu, Sotirios A. Tsaftaris |
IROS | 8 |
| 2025 | GMT: Guided Mask Transformer for Leaf Instance SegmentationabstractLeaf instance segmentation is a challenging multi-instance segmentation task, aiming to separate and delin-eate each leaf in an image of a plant. Accurate segmentation of each leaf is crucial for plant-related applications such as the fine-grained monitoring of plant growth and crop yield estimation. This task is challenging because of the high similarity (in shape and colour), great size variation, and heavy occlusions among leaf instances. Furthermore, the typically small size of annotated leaf datasets makes it more difficult to learn the distinctive features needed for precise segmentation. We hypothesise that the key to overcoming the these challenges lies in the specific spatial patterns of leaf distri-bution. In this paper, we propose the Guided Mask Trans-former (GMT), which leverages and integrates leaf spa-tial distribution priors into a Transformer-based segmentor. These spatial priors are embedded in a set of guide functions that map leaves at different positions into a more sepa-rable embedding space. Our GMT consistently outperforms the state-of-the-art on three public plant datasets. Our code is available at https://github.com/vios-s/gmt-leaf-ins-seg. Sotirios A. Tsaftaris, Mario Valerio Giuffrida |
WACV | 2 |
| 2025 | MemControl: Mitigating Memorization in Diffusion Models via Automated Parameter SelectionabstractDiffusion models excel in generating images that closely resemble their training data but are also susceptible to data memorization, raising privacy, ethical, and legal concerns, particularly in sensitive domains such as medical imaging. We hypothesize that this memorization stems from the overparameterization of deep models and propose that regularizing model capacity during fine-tuning can mitigate this issue. Firstly, we empirically show that regulating the model capacity via Parameter-efficient fine-tuning (PEFT) mitigates memorization to some extent, however, it further requires the identification of the exact parameter subsets to be fine-tuned for high-quality generation. To identify these subsets, we introduce a bilevel optimization framework, MemControl, that automates parameter selection using memorization and generation quality metrics as rewards during fine-tuning. The parameter subsets discovered through MemControl achieve a superior tradeoff between generation quality and memorization. For the task of medical image generation, our approach outperforms existing state-of-the-art memorization mitigation strategies by fine-tuning as few as 0.019% of model parameters. Moreover, we demonstrate that the discovered parameter subsets are transferable to non-medical domains. Our framework is scalable to large datasets, agnostic to reward functions, and can be integrated with existing approaches for further memorization mitigation. To the best of our knowledge, this is the first study to empirically evaluate memorization in medical images and propose a targeted yet universal mitigation strategy. The code is available at this https URL. Raman Dutt, Ondrej Bohdal, Pedro Sanchez, Sotirios A. Tsaftaris, Timothy M. Hospedales |
WACV | 4 |
| 2025 | Zero-Shot Medical Phrase Grounding With Off-the-Shelf Diffusion ModelsabstractLocalizing the exact pathological regions in a given medical scan is an important imaging problem that traditionally requires a large amount of bounding box ground truth annotations to be accurately solved. However, there exist alternative, potentially weaker, forms of supervision, such as accompanying free-text reports, which are readily available. The task of performing localization with textual guidance is commonly referred to as phrase grounding. In this work, we use a publicly available Foundation Model, namely the Latent Diffusion Model, to perform this challenging task. This choice is supported by the fact that the Latent Diffusion Model, despite being generative in nature, contains cross-attention mechanisms that implicitly align visual and textual features, thus leading to intermediate representations that are suitable for the task at hand. In addition, we aim to perform this task in a zero-shot manner, i.e., without any training on the target task, meaning that the model's weights remain frozen. To this end, we devise strategies to select features and also refine them via post-processing without extra learnable parameters. We compare our proposed method with state-of-the-art approaches which explicitly enforce image-text alignment in a joint embedding space via contrastive learning. Results on a popular chest X-ray benchmark indicate that our method is competitive with SOTA on different types of pathology, and even outperforms them on average in terms of two metrics (mean IoU and AUC-ROC). Konstantinos Vilouras, Pedro Sanchez, Alison O'Neil, Sotirios A. Tsaftaris |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Unlocking the Potential of Weakly Labeled Data: A Co-Evolutionary Learning Framework for Abnormality Detection and Report GenerationabstractAnatomical abnormality detection and report generation of chest X-ray (CXR) are two essential tasks in clinical practice. The former aims at localizing and characterizing cardiopulmonary radiological findings in CXRs, while the latter summarizes the findings in a detailed report for further diagnosis and treatment. Existing methods often focused on either task separately, ignoring their correlation. This work proposes a co-evolutionary abnormality detection and report generation (CoE-DG) framework. The framework utilizes both fully labeled (with bounding box annotations and clinical reports) and weakly labeled (with reports only) data to achieve mutual promotion between the abnormality detection and report generation tasks. Specifically, we introduce a bi-directional information interaction strategy with generator-guided information propagation (GIP) and detector-guided information propagation (DIP). For semi-supervised abnormality detection, GIP takes the informative feature extracted by the generator as an auxiliary input to the detector and uses the generator's prediction to refine the detector's pseudo labels. We further propose an intra-image-modal self-adaptive non-maximum suppression module (SA-NMS). This module dynamically rectifies pseudo detection labels generated by the teacher detection model with high-confidence predictions by the student. Inversely, for report generation, DIP takes the abnormalities' categories and locations predicted by the detector as input and guidance for the generator to improve the generated reports. Finally, a co-evolutionary training strategy is implemented to iteratively conduct GIP and DIP and consistently improve both tasks' performance. Experimental results on two public CXR datasets demonstrate CoE-DG's superior performance to several up-to-date object detection, report generation, and unified models. Our code is available at https://github.com/jinghanSunn/CoE-DG. Jinghan Sun, Dong Wei 0004, Zhe Xu 0012, Donghuan Lu, Hong Wang 0021, Sotirios A. Tsaftaris, Steven McDonagh 0001, Yefeng Zheng 0001, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 7 |
| 2024 | FairTune: Optimizing Parameter Efficient Fine Tuning for Fairness in Medical Image AnalysisabstractTraining models with robust group fairness properties is crucial in ethically sensitive application areas such as medical diagnosis. Despite the growing body of work aiming to minimise demographic bias in AI, this problem remains challenging. A key reason for this challenge is the fairness generalisation gap: High-capacity deep learning models can fit all training data nearly perfectly, and thus also exhibit perfect fairness during training. In this case, bias emerges only during testing when generalisation performance differs across sub-groups. This motivates us to take a bi-level optimisation perspective on fair learning: Optimising the learning strategy based on validation fairness. Specifically, we consider the highly effective workflow of adapting pre-trained models to downstream medical imaging tasks using parameter-efficient fine-tuning (PEFT) techniques. There is a trade-off between updating more parameters, enabling a better fit to the task of interest vs. fewer parameters, potentially reducing the generalisation gap. To manage this tradeoff, we propose FairTune, a framework to optimise the choice of PEFT parameters with respect to fairness. We demonstrate empirically that FairTune leads to improved fairness on a range of medical imaging datasets. The code is available at https://github.com/Raman1121/FairTune. Raman Dutt, Ondrej Bohdal, Sotirios A. Tsaftaris, Timothy M. Hospedales |
ICLR | 3 |
| 2024 | The MRI Scanner as a Diagnostic: Image-Less Active Sampling
Yuning Du, Rohan Dharmakumar, Sotirios A. Tsaftaris |
MICCAI (3) | 3 |
| 2024 | Benchmarking Counterfactual Image GenerationabstractGenerative AI has revolutionised visual content editing, empowering users to effortlessly modify images and videos. However, not all edits are equal. To perform realistic edits in domains such as natural image or medical imaging, modifications must respect causal relationships inherent to the data generation process. Such image editing falls into the counterfactual image generation regime. Evaluating counterfactual image generation is substantially complex: not only it lacks observable ground truths, but also requires adherence to causal constraints. Although several counterfactual image generation methods and evaluation metrics exist a comprehensive comparison within a unified setting is lacking. We present a comparison framework to thoroughly benchmark counterfactual image generation methods. We evaluate the performance of three conditional image generation model families developed within the Structural Causal Model (SCM) framework. We incorporate several metrics that assess diverse aspects of counterfactuals, such as composition, effectiveness, minimality of interventions, and image realism. We integrate all models that have been used for the task at hand and expand them to novel datasets and causal graphs, demonstrating the superiority of Hierarchical VAEs across most datasets and metrics. Our framework is implemented in a user-friendly Python package that can be extended to incorporate additional SCMs, causal methods, generative models, and datasets for the community to build on. Code: https://github.com/gulnazaki/counterfactual-benchmark. Thomas Melistas, Nikos Spyrou, Nefeli Gkouti, Pedro Sanchez, Athanasios Vlontzos, Yannis Panagakis, Giorgos Papanastasiou, Sotirios A. Tsaftaris |
NeurIPS | 8 |
| 2024 | Long-short diffeomorphism memory network for weakly-supervised ultrasound landmark trackingabstractUltrasound is a promising medical imaging modality benefiting from low-cost and real-time acquisition. Accurate tracking of an anatomical landmark has been of high interest for various clinical workflows such as minimally invasive surgery and ultrasound-guided radiation therapy. However, tracking an anatomical landmark accurately in ultrasound video is very challenging, due to landmark deformation, visual ambiguity and partial observation. In this paper, we propose a long-short diffeomorphism memory network (LSDM), which is a multi-task framework with an auxiliary learnable deformation prior to supporting accurate landmark tracking. Specifically, we design a novel diffeomorphic representation, which contains both long and short temporal information stored in separate memory banks for delineating motion margins and reducing cumulative errors. We further propose an expectation maximization memory alignment (EMMA) algorithm to iteratively optimize both the long and short deformation memory, updating the memory queue for mitigating local anatomical ambiguity. The proposed multi-task system can be trained in a weakly-supervised manner, which only requires few landmark annotations for tracking and zero annotation for deformation learning. We conduct extensive experiments on both public and private ultrasound landmark tracking datasets. Experimental results show that LSDM can achieve better or competitive landmark tracking performance with a strong generalization capability across different scanner types and different ultrasound modalities, compared with other state-of-the-art methods. Xuejun Ni, Sotirios A. Tsaftaris, Huiyu Zhou 0001 |
Medical Image Anal. | 5 |
| 2024 | Compositionally Equivariant Representation LearningabstractDeep learning models often need sufficient supervision (i.e., labelled data) in order to be trained effectively. By contrast, humans can swiftly learn to identify important anatomy in medical images like MRI and CT scans, with minimal guidance. This recognition capability easily generalises to new images from different medical facilities and to new tasks in different settings. This rapid and generalisable learning ability is largely due to the compositional structure of image patterns in the human brain, which are not well represented in current medical models. In this paper, we study the utilisation of compositionality in learning more interpretable and generalisable representations for medical image segmentation. Overall, we propose that the underlying generative factors that are used to generate the medical images satisfy compositional equivariance property, where each factor is compositional (e.g., corresponds to human anatomy) and also equivariant to the task. Hence, a good representation that approximates well the ground truth factor has to be compositionally equivariant. By modelling the compositional representations with learnable von-Mises-Fisher (vMF) kernels, we explore how different design and learning biases can be used to enforce the representations to be more compositionally equivariant under un-, weakly-, and semi-supervised settings. Extensive results show that our methods achieve the best performance over several strong baselines on the task of semi-supervised domain-generalised medical image segmentation. Code will be made publicly available upon acceptance at https://github.com/vios-s. Xiao Liu 0037, Pedro Sanchez, Spyridon Thermos, Alison O'Neil, Sotirios A. Tsaftaris |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Diffusion Models for Causal Discovery via Topological Ordering
Pedro Sanchez, Xiao Liu 0037, Alison O'Neil, Sotirios A. Tsaftaris |
ICLR | 4 |
| 2023 | SLX: Similarity Learning for X-Ray Screening and Robust Automated Disassembled Object DetectionabstractBaggage screening is important in security-critical applications in airports for detecting threats, including firearms and parts of them. Existing approaches underperform to recognise prohibited objects that are disassembled, especially when learning from limited data and from images produced by different scanners with multi-view orientations. To address such limitations, in this paper, we develop the Similarity Learning X-ray screening (SLX) model for accurate and robust firearm component detection in cluttered scenes. We evaluate SLX on the X-ray Image Library (XIL) dataset that the UK Government has provided us with, for this research. SLX is based on a contrastive similarity learning approach combined with Out-of-Distribution (OoD) detection/ anomaly detection using a deep discriminative model, ResNet-152, for detecting and classifying forbidden items. The evaluation of SLX on the XIL dataset shows that it is effective, beneficial for detecting firearms and their parts, and outperforms other baseline models, on average, by approximately 12 points in accuracy. Nikolaos Dionelis, Richard Jackson, Sotirios A. Tsaftaris, Mehrdad Yaghoobi |
IJCNN | 3 |
| 2023 | MyoPS: A benchmark of myocardial pathology segmentation combining three-sequence cardiac magnetic resonance images
Lei Li 0020, Fuping Wu, Xinzhe Luo, Carlos Martín-Isla, Shuwei Zhai, Zhen Zhang 0057, Markus J. Ankenbrand, Haochuan Jiang, Linhong Wang, Tewodros Weldebirhan Arega, Elif Altunok, Jun Ma 0016, Xiaoping Yang 0001, Élodie Puybareau, Ilkay Öksüz, Stéphanie Bricq, Weisheng Li 0001, Kumaradevan Punithakumar, Sotirios A. Tsaftaris, Laura Maria Schreiber, Guocai Liu, Yong Xia 0001, Guotai Wang, Sergio Escalera, Xiahai Zhuang |
Medical Image Anal. | 25 |
| 2023 | The role of noise in denoising models for anomaly detection in medical images
Antanas Kascenas, Pedro Sanchez, Patrick Schrempf, William Clackett, Shadia Mikhael, Jeremy Voisey, Keith A. Goatman, Alexander J. Weir, Nicolas Pugeault, Sotirios A. Tsaftaris, Alison O'Neil |
Medical Image Anal. | 11 |
| 2023 | Survey: Leakage and Privacy at Inference TimeabstractLeakage of data from publicly available Machine Learning (ML) models is an area of growing significance since commercial and government applications of ML can draw on multiple sources of data, potentially including users' and clients' sensitive data. We provide a comprehensive survey of contemporary advances on several fronts, covering involuntary data leakage which is natural to ML models, potential malicious leakage which is caused by privacy attacks, and currently available defence mechanisms. We focus on inference-time leakage, as the most likely scenario for publicly available models. We first discuss what leakage is in the context of different data, tasks, and model architectures. We then propose a taxonomy across involuntary and malicious leakage, followed by description of currently available defences, assessment metrics, and applications. We conclude with outstanding challenges and open questions, outlining some promising directions for future research. Marija Jegorova, Chaitanya Kaul, Charlie Mayor, Alison O'Neil, Alexander J. Weir, Roderick Murray-Smith, Sotirios A. Tsaftaris |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | CTR: Contrastive Training Recognition Classifier for Few-Shot Open-World RecognitionabstractAI-enabled systems in security, autonomous systems, safety, and healthcare do not only need to effectively detect Out-of-Distribution (OoD) samples, but also to recognize Objects of Concern (OoC), e.g. multiple thorax diseases, efficiently with few-shots. Detecting OoD samples is crucial, because reporting an out-of-domain input as abnormal is better than falsely classifying it. Data samples, during inference, are not confined to a finite labelled set, and thus closed-set approaches are limiting, as they misclassify OoD inputs, and they may assign them high prediction confidence. Furthermore, although anomaly detection is possible, recognizing new OoC fast using only few-shot samples remains challenging. There is a lack of methodologies for joint anomaly detection and few-shot OoC classification. Our contribution is the development of a framework for joint few-shot OoC detection and classification and anomaly detection in the unknown previously-unseen, in the wild, environment, which is known as Open-World Recognition (OWR). We propose a novel methodology, the data distribution boundary Contrastive Training Recognition (CTR) classifier for few-shot OWR. CTR takes advantage of labels and classes to learn the normal (and few-shot abnormal) data better, to more accurately detect OoD. The proposed model: (i) reduces failures to detect anomalies in health- and safety-critical applications for avoiding unfavourable consequences, (ii) decreases false alarms, and (iii) improves performance overall. Our framework differs from existing approaches because: (a) anomaly and OoC detection are combined, which has several benefits, including improved OoD performance, (b) the performance, accuracy, and robustness of OoD and few-shot OoC detection are improved by strengthening the estimation of the normal class distribution at the boundary of its support, self-generating samples and setting them as abnormal, and (c) the knowledge base of models is also augmented by learning class-incrementally, alleviating forgetting. CTR outperforms baselines in several settings, including on the SVHN, CIFAR-FS, and BSCD-FS ChestX and ISIC datasets. Nikolaos Dionelis, Sotirios A. Tsaftaris, Mehrdad Yaghoobi |
ICPR | 2 |
| 2022 | vMFNet: Compositionality Meets Domain-Generalised Segmentation
Xiao Liu 0037, Spyridon Thermos, Pedro Sanchez, Alison O'Neil, Sotirios A. Tsaftaris |
MICCAI (8) | 5 |
| 2022 | Why Patient Data Cannot Be Easily Forgotten?
Ruolin Su, Xiao Liu 0037, Sotirios A. Tsaftaris |
MICCAI (8) | 3 |
| 2022 | FROB: Few-Shot ROBust Model for Joint Classification and Out-of-Distribution Detection
Nikolaos Dionelis, Sotirios A. Tsaftaris, Mehrdad Yaghoobi |
ECML/PKDD (3) | 2 |
| 2022 | Learning disentangled representations in the imaging domainabstractDisentangled representation learning has been proposed as an approach to learning general representations even in the absence of, or with limited, supervision. A good general representation can be fine-tuned for new target tasks using modest amounts of data, or used directly in unseen domains achieving remarkable performance in the corresponding task. This alleviation of the data and annotation requirements offers tantalising prospects for applications in computer vision and healthcare. In this tutorial paper, we motivate the need for disentangled representations, revisit key concepts, and describe practical building blocks and criteria for learning such representations. We survey applications in medical imaging emphasising choices made in exemplar key works, and then discuss links to computer vision applications. We conclude by presenting limitations, challenges, and opportunities. Xiao Liu 0037, Pedro Sanchez, Spyridon Thermos, Alison O'Neil, Sotirios A. Tsaftaris |
Medical Image Anal. | 5 |
| 2021 | Measuring the Biases and Effectiveness of Content-Style Disentanglement
Xiao Liu 0037, Spyridon Thermos, Gabriele Valvano, Agisilaos Chartsias, Alison O'Neil, Sotirios A. Tsaftaris |
BMVC | 6 |
| 2021 | Few-Shot Adaptive Detection of Objects of Concern Using Generative Models with Negative RetrainingabstractDetecting objects which we are interested in, Objects of Concern (OoC), is nowadays attracting attention. In aviation and transport, it is important to robustly detect OoC in security images. OoC are rare, differ from typical samples, and may be unknown during training. Most OoC detection methods need to be trained on large datasets to achieve good performance and have limited real-world generalization ability. A large variety of samples is needed, and it is expensive to collect and label large datasets due to the rarity of OoC. To address such limitations, we propose the negative REtraining with Few-shots Generative Adversarial Network (REFGAN) for detecting OoC. REFGAN aims at automatically identifying OoC by learning from Objects of No Concern (OoNC) and OoC. Our methodology comprises learning a prior using OoNC, and few-shot adaptation using the Few-Shot OoC (FSOoC). We propose a methodology to robustly perform few-shot adaptive detection of OoC using GANs and learned distribution boundaries. The evaluation of REFGAN on the Baggage SIXray dataset shows that when FSOoC are used, our model outperforms the prior improving OoC detection, and outperforms recent benchmarks by approximately 6.3% in mean values. REFGAN using few-shots of 80 samples shows a robust comparable performance compared to REFGAN using all the samples for retraining and model adaptation. REFGAN can detect unknown OoC and its evaluation on SIXray and CIFAR-10 shows robustness against the number of few-shot samples of OoC. REFGAN on CIFAR-10 outperforms benchmarks by approximately 16% using few-shots of 80 and of 10 samples. Nikolaos Dionelis, Mehrdad Yaghoobi, Sotirios A. Tsaftaris |
ICTAI | 3 |
| 2021 | Semi-supervised Meta-learning with Disentanglement for Domain-Generalised Medical Image Segmentation
Xiao Liu 0037, Spyridon Thermos, Alison O'Neil, Sotirios A. Tsaftaris |
MICCAI (2) | 4 |
| 2021 | Controllable Cardiac Synthesis via Disentangled Anatomy Arithmetic
Spyridon Thermos, Xiao Liu 0037, Alison O'Neil, Sotirios A. Tsaftaris |
MICCAI (3) | 4 |
| 2021 | Learning to synthesise the ageing brain without longitudinal data
Agisilaos Chartsias, Chengjia Wang, Sotirios A. Tsaftaris |
Medical Image Anal. | 4 |
| 2021 | Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation: The M&Ms ChallengeabstractThe emergence of deep learning has considerably advanced the state-of-the-art in cardiac magnetic resonance (CMR) segmentation. Many techniques have been proposed over the last few years, bringing the accuracy of automated segmentation close to human performance. However, these models have been all too often trained and validated using cardiac imaging samples from single clinical centres or homogeneous imaging protocols. This has prevented the development and validation of models that are generalizable across different clinical centres, imaging conditions or scanner vendors. To promote further research and scientific benchmarking in the field of generalizable deep learning for cardiac segmentation, this paper presents the results of the Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation (M&Ms) Challenge, which was recently organized as part of the MICCAI 2020 Conference. A total of 14 teams submitted different solutions to the problem, combining various baseline models, data augmentation strategies, and domain adaptation techniques. The obtained results indicate the importance of intensity-driven data augmentation, as well as the need for further research to improve generalizability towards unseen scanner vendors or new imaging protocols. Furthermore, we present a new resource of 375 heterogeneous CMR datasets acquired by using four different scanner vendors in six hospitals and three different countries (Spain, Canada and Germany), which we provide as open-access for the community to enable future research in the field. Víctor M. Campello, Polyxeni Gkontra, Cristian Izquierdo, Carlos Martín-Isla, Alireza Sojoudi, Peter M. Full, Klaus H. Maier-Hein, Yao Zhang 0010, Zhiqiang He 0002, Jun Ma 0016, Mario Parreño, Alberto Albiol, Fanwei Kong, Shawn C. Shadden, Jorge Corral Acero, Vaanathi Sundaresan, Mina Saber, Mustafa A. Alattar, Hongwei Li 0004, Bjoern Menze, Firas Khader, Christoph Haarburger, Cian M. Scannell, Mitko Veta, Adam Carscadden, Kumaradevan Punithakumar, Xiao Liu 0037, Sotirios A. Tsaftaris, Xiaoqiong Huang, Xin Yang 0009, Lei Li 0020, Xiahai Zhuang, David Viladés, Martín Luís Descalzo, Andrea Guala 0002, Lucia La Mura, Matthias G. W. Friedrich, Ria Garg, Julie Lebel, Filipe Henriques, Mahir Karakas, Ersin Çavus, Steffen E. Petersen, Sergio Escalera, Santi Seguí, Jose Rodriguez-Palomares, Karim Lekadir |
IEEE Trans. Medical Imaging | 28 |
| 2021 | Disentangle, Align and Fuse for Multimodal and Semi-Supervised Image SegmentationabstractMagnetic resonance (MR) protocols rely on several sequences to assess pathology and organ status properly. Despite advances in image analysis, we tend to treat each sequence, here termed modality, in isolation. Taking advantage of the common information shared between modalities (an organ's anatomy) is beneficial for multi-modality processing and learning. However, we must overcome inherent anatomical misregistrations and disparities in signal intensity across the modalities to obtain this benefit. We present a method that offers improved segmentation accuracy of the modality of interest (over a single input model), by learning to leverage information present in other modalities, even if few (semi-supervised) or no (unsupervised) annotations are available for this specific modality. Core to our method is learning a disentangled decomposition into anatomical and imaging factors. Shared anatomical factors from the different inputs are jointly processed and fused to extract more accurate segmentation masks. Image misregistrations are corrected with a Spatial Transformer Network, which non-linearly aligns the anatomical factors. The imaging factor captures signal intensity characteristics across different modality data and is used for image reconstruction, enabling semi-supervised learning. Temporal and slice pairing between inputs are learned dynamically. We demonstrate applications in Late Gadolinium Enhanced (LGE) and Blood Oxygenation Level Dependent (BOLD) cardiac segmentation, as well as in T2 abdominal segmentation. Code is available at https://github.com/vios-s/multimodal_segmentation. Agisilaos Chartsias, Giorgos Papanastasiou, Chengjia Wang, Scott Semple, David E. Newby, Rohan Dharmakumar, Sotirios A. Tsaftaris |
IEEE Trans. Medical Imaging | 7 |
| 2021 | Learning to Segment From Scribbles Using Multi-Scale Adversarial Attention GatesabstractLarge, fine-grained image segmentation datasets, annotated at pixel-level, are difficult to obtain, particularly in medical imaging, where annotations also require expert knowledge. Weakly-supervised learning can train models by relying on weaker forms of annotation, such as scribbles. Here, we learn to segment using scribble annotations in an adversarial game. With unpaired segmentation masks, we train a multi-scale GAN to generate realistic segmentation masks at multiple resolutions, while we use scribbles to learn their correct position in the image. Central to the model's success is a novel attention gating mechanism, which we condition with adversarial signals to act as a shape prior, resulting in better object localization at multiple scales. Subject to adversarial conditioning, the segmentor learns attention maps that are semantic, suppress the noisy activations outside the objects, and reduce the vanishing gradient problem in the deeper layers of the segmentor. We evaluated our model on several medical (ACDC, LVSC, CHAOS) and non-medical (PPSS) datasets, and we report performance levels matching those achieved by models trained with fully annotated segmentation masks. We also demonstrate extensions in a variety of settings: semi-supervised learning; combining multiple scribble sources (a crowdsourcing scenario) and multi-task learning (combining scribble and mask supervision). We release expert-made scribble annotations for the ACDC dataset, and the code used for the experiments, at https://vios-s.github.io/multiscale-adversarial-attention-gates. Gabriele Valvano, Andrea Leo, Sotirios A. Tsaftaris |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Boundary Of Distribution Support Generator (BDSG): Sample Generation On The BoundaryabstractGenerative models, such as Generative Adversarial Networks (GANs), have been used for unsupervised anomaly detection. While performance keeps improving, several limitations exist particularly attributed to difficulties at capturing multimodal supports and to the ability to approximate the underlying distribution closer to the tails, i.e. the boundary of the distribution's support. This paper proposes an approach that attempts to alleviate such shortcomings. We propose an invertible-residual-network-based model, the Boundary of Distribution Support Generator (BDSG). GANs generally do not guarantee the existence of a probability distribution and here, we use the recently developed Invertible Residual Network (IResNet) and Residual Flow (ResFlow), for density estimation. These models have not yet been used for anomaly detection. We leverage IResNet and ResFlow for Out-of-Distribution (OoD) sample detection and for sample generation on the boundary using a compound loss function that forces the samples to lie on the boundary. The BDSG addresses non-convex support, disjoint components, and multimodal distributions. Results on synthetic data and data from multimodal distributions, such as MNIST and CIFAR-10, demonstrate competitive performance compared to methods from the literature. Nikolaos Dionelis, Mehrdad Yaghoobi, Sotirios A. Tsaftaris |
ICIP | 3 |
| 2020 | INSIDE: Steering Spatial Attention with Non-imaging Information in CNNs
Grzegorz Jacenkow, Alison O'Neil, Brian Mohr, Sotirios A. Tsaftaris |
MICCAI (4) | 4 |
| 2020 | Have You Forgotten? A Method to Assess if Machine Learning Models Have Forgotten Data
Xiao Liu 0037, Sotirios A. Tsaftaris |
MICCAI (1) | 2 |
| 2020 | Pseudo-healthy synthesis with pathology disentanglement and adversarial learning
Agisilaos Chartsias, Sotirios A. Tsaftaris |
Medical Image Anal. | 3 |
| 2020 | Unsupervised Rotation Factorization in Restricted Boltzmann MachinesabstractFinding suitable image representations for the task at hand is critical in computer vision. Different approaches extending the original Restricted Boltzmann Machine (RBM) model have recently been proposed to offer rotation-invariant feature learning. In this paper, we present an extended novel RBM that learns rotation invariant features by explicitly factorizing for rotation nuisance in 2D image inputs within an unsupervised framework. While the goal is to learn invariant features, our model infers an orientation per input image during training, using information related to the reconstruction error. The training process is regularised by a Kullback-Leibler divergence, offering stability and consistency. We used the γ -score, a measure that calculates the amount of invariance, to mathematically and experimentally demonstrate that our approach indeed learns rotation invariant features. We show that our method outperforms the current state-of-the-art RBM approaches for rotation invariant feature learning on three different benchmark datasets, by measuring the performance with the test accuracy of an SVM classifier. Our implementation is available at https://bitbucket.org/tuttoweb/rotinvrbm. Mario Valerio Giuffrida, Sotirios A. Tsaftaris |
IEEE Trans. Image Process. | 2 |
| 2020 | AI in Medical Imaging Informatics: Current Challenges and Future DirectionsabstractThis paper reviews state-of-the-art research solutions across the spectrum of medical imaging informatics, discusses clinical translation, and provides future directions for advancing clinical practice. More specifically, it summarizes advances in medical imaging acquisition technologies for different modalities, highlighting the necessity for efficient medical data management strategies in the context of AI in big healthcare data analytics. It then provides a synopsis of contemporary and emerging algorithmic methods for disease classification and organ/ tissue segmentation, focusing on AI and deep learning architectures that have already become the de facto approach. The clinical benefits of in-silico modelling advances linked with evolving 3D reconstruction and visualization applications are further documented. Concluding, integrative analytics approaches driven by associate research branches highlighted in this study promise to revolutionize imaging informatics as known today across the healthcare continuum for both radiology and digital pathology applications. The latter, is projected to enable informed, more accurate diagnosis, timely prognosis, and effective treatment planning, underpinning precision medicine. Andreas Panayides, Amir A. Amini, Nenad Filipovic, Ashish Sharma 0001, Sotirios A. Tsaftaris, Alistair A. Young, David J. Foran, Nhan Do, Spyretta Golemati, Tahsin M. Kurç, Kun Huang 0001, Konstantina S. Nikita, Benjamin Veasey, Michalis E. Zervakis, Joel H. Saltz, Constantinos S. Pattichis |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | Consistent Brain Ageing Synthesis
Agisilaos Chartsias, Sotirios A. Tsaftaris |
MICCAI (4) | 3 |
| 2019 | Disentangled representation learning in cardiac image analysis
Agisilaos Chartsias, Thomas Joyce, Giorgos Papanastasiou, Scott Semple, Michelle C. Williams, David E. Newby, Rohan Dharmakumar, Sotirios A. Tsaftaris |
Medical Image Anal. | 8 |
| 2018 | Root Gap Correction with a Deep Inpainting Model
Hao Chen 0102, Mario Valerio Giuffrida, Sotirios A. Tsaftaris, Peter Doerner |
BMVC | 3 |
| 2018 | Factorised Spatial Representation Learning: Application in Semi-supervised Myocardial Segmentation
Agisilaos Chartsias, Thomas Joyce, Giorgos Papanastasiou, Scott Semple, Michelle C. Williams, David E. Newby, Rohan Dharmakumar, Sotirios A. Tsaftaris |
MICCAI (2) | 8 |
| 2018 | Statistical Shape Modeling of the Left Ventricle: Myocardial Infarct Classification ChallengeabstractStatistical shape modeling is a powerful tool for visualizing and quantifying geometric and functional patterns of the heart. After myocardial infarction (MI), the left ventricle typically remodels in response to physiological challenges. Several methods have been proposed in the literature to describe statistical shape changes. Which method best characterizes left ventricular remodeling after MI is an open research question. A better descriptor of remodeling is expected to provide a more accurate evaluation of disease status in MI patients. We therefore designed a challenge to test shape characterization in MI given a set of three-dimensional left ventricular surface points. The training set comprised 100 MI patients, and 100 asymptomatic volunteers (AV). The challenge was initiated in 2015 at the Statistical Atlases and Computational Models of the Heart workshop, in conjunction with the MICCAI conference. The training set with labels was provided to participants, who were asked to submit the likelihood of MI from a different (validation) set of 200 cases (100 AV and 100 MI). Sensitivity, specificity, accuracy and area under the receiver operating characteristic curve were used as the outcome measures. The goals of this challenge were to (1) establish a common dataset for evaluating statistical shape modeling algorithms in MI, and (2) test whether statistical shape modeling provides additional information characterizing MI patients over standard clinical measures. Eleven groups with a wide variety of classification and feature extraction approaches participated in this challenge. All methods achieved excellent classification results with accuracy ranges from 0.83 to 0.98. The areas under the receiver operating characteristic curves were all above 0.90. Four methods showed significantly higher performance than standard clinical measures. The dataset and software for evaluation are available from the Cardiac Atlas Project website1. Avan Suinesiaputra, Pierre Ablin, Xènia Albà, Martino Alessandrini, Jack Allen, Wenjia Bai, Serkan Çimen, Peter Claes, Brett R. Cowan, Jan D'hooge, Nicolas Duchateau, Jan Ehrhardt, Alejandro F. Frangi, Ali Gooya, Vicente Grau, Karim Lekadir, Allen Lu, Anirban Mukhopadhyay 0003, Ilkay Öksüz, Nripesh Parajuli, Xavier Pennec, Marco Pereañez, Catarina Pinto, Paolo Piras, Marc-Michel Rohé, Daniel Rueckert, Dennis Säring, Maxime Sermesant, Kaleem Siddiqi, Mahdi Tabassian, Luciano Teresi, Sotirios A. Tsaftaris, Matthias Wilms, Alistair A. Young, Pau Medrano-Gracia |
IEEE J. Biomed. Health Informatics | 32 |
| 2018 | Multimodal MR Synthesis via Modality-Invariant Latent RepresentationabstractWe propose a multi-input multi-output fully convolutional neural network model for MRI synthesis. The model is robust to missing data, as it benefits from, but does not require, additional input modalities. The model is trained end-to-end, and learns to embed all input modalities into a shared modality-invariant latent space. These latent representations are then combined into a single fused representation, which is transformed into the target output modality with a learnt decoder. We avoid the need for curriculum learning by exploiting the fact that the various input modalities are highly correlated. We also show that by incorporating information from segmentation masks the model can both decrease its error and generate data with synthetic lesions. We evaluate our model on the ISLES and BRATS data sets and demonstrate statistically significant improvements over state-of-the-art methods for single input tasks. This improvement increases further when multiple input modalities are used, demonstrating the benefits of learning a common latent space, again resulting in a statistically significant improvement over the current best method. Finally, we demonstrate our approach on non skull-stripped brain images, producing a statistically significant improvement over the previous best method. Code is made publicly available at https://github.com/agis85/multimodal_brain_synthesis. Agisilaos Chartsias, Thomas Joyce, Mario Valerio Giuffrida, Sotirios A. Tsaftaris |
IEEE Trans. Medical Imaging | 4 |
| 2018 | Simulation and Synthesis in Medical ImagingabstractThis editorial introduces the Special Issue on Simulation and Synthesis in Medical Imaging. In this editorial, we define so-far ambiguous terms of simulation and synthesis in medical imaging. We also briefly discuss the synergistic importance of mechanistic (hypothesis-driven) and phenomenological (data-driven) models of medical image generation. Finally, we introduce the twelve papers published in this issue covering both mechanistic (5) and phenomenological (7) medical image generation. This rich selection of papers covers applications in cardiology, retinopathy, histopathology, neurosciences, and oncology. It also covers all mainstream diagnostic medical imaging modalities. We conclude the editorial with a personal view on the field and highlight some existing challenges and future research opportunities. Alejandro F. Frangi, Sotirios A. Tsaftaris, Jerry L. Prince |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Robust Multi-modal MR Image Synthesis
Thomas Joyce, Agisilaos Chartsias, Sotirios A. Tsaftaris |
MICCAI (3) | 3 |
| 2017 | Unsupervised Myocardial Segmentation for Cardiac BOLDabstractA fully automated 2-D+time myocardial segmentation framework is proposed for cardiac magnetic resonance (CMR) blood-oxygen-level-dependent (BOLD) data sets. Ischemia detection with CINE BOLD CMR relies on spatio-temporal patterns in myocardial intensity, but these patterns also trouble supervised segmentation methods, the de facto standard for myocardial segmentation in cine MRI. Segmentation errors severely undermine the accurate extraction of these patterns. In this paper, we build a joint motion and appearance method that relies on dictionary learning to find a suitable subspace. Our method is based on variational pre-processing and spatial regularization using Markov random fields, to further improve performance. The superiority of the proposed segmentation technique is demonstrated on a data set containing cardiac phase-resolved BOLD MR and standard CINE MR image sequences acquired in baseline and ischemic condition across ten canine subjects. Our unsupervised approach outperforms even supervised state-of-the-art segmentation techniques by at least 10% when using Dice to measure accuracy on BOLD data and performs at par for standard CINE MR. Furthermore, a novel segmental analysis method attuned for BOLD time series is utilized to demonstrate the effectiveness of the proposed method in preserving key BOLD patterns. Ilkay Öksüz, Anirban Mukhopadhyay 0003, Rohan Dharmakumar, Sotirios A. Tsaftaris |
IEEE Trans. Medical Imaging | 4 |
| 2016 | Rotation-Invariant Restricted Boltzmann Machine Using Shared Gradient Filters
Mario Valerio Giuffrida, Sotirios A. Tsaftaris |
ICANN (2) | 2 |
| 2016 | Special issue on computer vision and image analysis in plant phenotyping
Hanno Scharr, Hannah M. Dee, Andrew P. French, Sotirios A. Tsaftaris |
Mach. Vis. Appl. | 4 |
| 2016 | Leaf segmentation in plant phenotyping: a collation study
Hanno Scharr, Massimo Minervini, Andrew P. French, Christian Klukas, David M. Kramer 0001, Xiaoming Liu 0002, Imanol Luengo, Jean-Michel Pape, Gerrit Polder, Danijela Vukadinovic, Xi Yin 0001, Sotirios A. Tsaftaris |
Mach. Vis. Appl. | 12 |
| 2016 | Finely-grained annotated datasets for image-based plant phenotyping
Massimo Minervini, Andreas Fischbach, Hanno Scharr, Sotirios A. Tsaftaris |
Pattern Recognit. Lett. | 4 |
| 2016 | Dictionary-Driven Ischemia Detection From Cardiac Phase-Resolved Myocardial BOLD MRI at RestabstractCardiac Phase-resolved Blood-Oxygen-Level Dependent (CP-BOLD) MRI provides a unique opportunity to image an ongoing ischemia at rest. However, it requires post-processing to evaluate the extent of ischemia. To address this, here we propose an unsupervised ischemia detection (UID) method which relies on the inherent spatio-temporal correlation between oxygenation and wall motion to formalize a joint learning and detection problem based on dictionary decomposition. Considering input data of a single subject, it treats ischemia as an anomaly and iteratively learns dictionaries to represent only normal observations (corresponding to myocardial territories remote to ischemia). Anomaly detection is based on a modified version of One-class Support Vector Machines (OCSVM) to regulate directly the margins by incorporating the dictionary-based representation errors. A measure of ischemic extent (IE) is estimated, reflecting the relative portion of the myocardium affected by ischemia. For visualization purposes an ischemia likelihood map is created by estimating posterior probabilities from the OCSVM outputs, thus obtaining how likely the classification is correct. UID is evaluated on synthetic data and in a 2D CP-BOLD data set from a canine experimental model emulating acute coronary syndromes. Comparing early ischemic territories identified with UID against infarct territories (after several hours of ischemia), we find that IE, as measured by UID, is highly correlated (Pearson's r=0.84) with respect to infarct size. When advances in automated registration and segmentation of CP-BOLD images and full coverage 3D acquisitions become available, we hope that this method can enable pixel-level assessment of ischemia with this truly non-invasive imaging technique. Marco Bevilacqua, Rohan Dharmakumar, Sotirios A. Tsaftaris |
IEEE Trans. Medical Imaging | 3 |
| 2016 | Supporting Autonomic Management of Clouds: Service Clustering With Random ForestabstractA promising solution for the management of services in clouds, as fostered by autonomic computing, is to resort to self-management. However, the obfuscation of underlying details of services in cloud computing, also due to privacy requirements, affects the effectiveness of autonomic managers. Data-driven approaches, in particular those relying on service clustering based on machine learning techniques, can assist the autonomic management and support decisions concerning, e.g., the scheduling and deployment of services. Unfortunately, applying such approaches is further complicated by the coexistence of different types of data within the information provided by the monitoring of cloud systems: both continuous (e.g., CPU load) and categorical (e.g., VM instance type) data are available. Current approaches deal with this problem in a heuristic fashion. In this paper, instead, we propose an approach that uses all types of data, and learns in a data-driven fashion the similarities and patterns among the services. More specifically, we design an unsupervised formulation of random forest to calculate service similarities and provide them as input to a clustering algorithm. For the sake of efficiency and to meet the dynamism requirement of autonomic clouds, our methodology consists of two steps: 1) off-line clustering and 2) on-line prediction. Using datasets from real-world clouds, we demonstrate the superiority of our solution with respect to others and validate the accuracy of the on-line prediction. Moreover, to show applicability of our approach, we devise a service scheduler that uses similarity among services, and evaluate its performance in a cloud test-bed using realistic data. Rafael Brundo Uriarte, Francesco Tiezzi 0001, Sotirios A. Tsaftaris |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2015 | Service Clustering for Autonomic Clouds Using Random ForestabstractManaging and optimising cloud services is one of the main challenges faced by industry and academia. A possible solution is resorting to self-management, as fostered by autonomic computing. However, the abstraction layer provided by cloud computing obfuscates several details of the provided services, which, in turn, hinders the effectiveness of autonomic managers. Data-driven approaches, particularly those relying on service clustering based on machine learning techniques, can assist the autonomic management and support decisions concerning, for example, the scheduling and deployment of services. One aspect that complicates this approach is that the information provided by the monitoring contains both continuous (e.g. CPU load) and categorical (e.g. VM instance type) data. Current approaches treat this problem in a heuristic fashion. This paper, instead, proposes an approach, which uses all kinds of data and learns in a data-driven fashion the similarities and resource usage patterns among the services. In particular, we use an unsupervised formulation of the Random Forest algorithm to calculate similarities and provide them as input to a clustering algorithm. For the sake of efficiency and meeting the dynamism requirement of autonomic clouds, our methodology consists of two steps: (i) off-line clustering and (ii) on-line prediction. Using datasets from real-world clouds, we demonstrate the superiority of our solution with respect to others and validate the accuracy of the on-line prediction. Moreover, to show the applicability of our approach, we devise a service scheduler that uses the notion of similarity among services and evaluate it in a cloud test-bed. Rafael Brundo Uriarte, Sotirios A. Tsaftaris, Francesco Tiezzi 0001 |
CCGRID | 2 |
| 2015 | Unsupervised Myocardial Segmentation for Cardiac MRI
Anirban Mukhopadhyay 0003, Ilkay Öksüz, Marco Bevilacqua, Rohan Dharmakumar, Sotirios A. Tsaftaris |
MICCAI (3) | 5 |
| 2015 | Dictionary Learning Based Image Descriptor for Myocardial Registration of CP-BOLD MR
Ilkay Öksüz, Anirban Mukhopadhyay 0003, Marco Bevilacqua, Rohan Dharmakumar, Sotirios A. Tsaftaris |
MICCAI (2) | 5 |
| 2015 | Classification-aware distortion metric for HEVC intra codingabstractIncreasingly many vision applications necessitate the transmission of acquired images and video to a remote location for automated processing. When the image data are consumed by analysis algorithms and possibly never seen by a human, tailoring compression to the application is beneficial from a bit rate perspective. We inject prior knowledge of the application in the encoder to make rate-distortion decisions based on an estimate of the accuracy that will be achieved when analyzing reconstructed image data. Focusing on classification (e.g., used for image segmentation), we propose a new application-aware distortion metric based on a geometric interpretation of classification error. We devise an implementation for the High Efficiency Video Coding standard, and derive optimal model parameters for the A-domain rate control algorithm by curve fitting procedures. We evaluate our approach on time-lapse sequences from plant phenotyping experiments and cell fluorescence microscopy encoded in intra-only mode, observing a reduction in segmentation error across bit rates. Massimo Minervini, Sotirios A. Tsaftaris |
VCIP | 2 |
| 2014 | Unsupervised and supervised approaches to color space transformation for image codingabstractThe linear transformation of input (typically RGB) data into a color space is important in image compression. Most schemes adopt fixed transforms to decorrelate the color channels. Energy compaction transforms such as the Karhunen-Loève (KLT) do entail a complexity increase. Here, we propose a new data-dependent transform (aKLT), that achieves compression performance comparable to the KLT, at a fraction of the computational complexity. More important, we also consider an application-aware setting, in which a classifier analyzes reconstructed images at the receiver's end. In this context, KLT-based approaches may not be optimal and transforms that maximize post-compression classifier performance are more suited. Relaxing energy compactness constraints, we propose for the first time a transform which can be found offline optimizing the Fisher discrimination criterion in a supervised fashion. In lieu of channel decorrelation, we obtain spatial decorrelation using the same color transform as a rudimentary classifier to detect objects of interest in the input image without adding any computational cost. We achieve higher savings encoding these regions at a higher quality, when combined with region-of-interest capable encoders, such as JPEG 2000. Massimo Minervini, Cristian Rusu, Sotirios A. Tsaftaris |
ICIP | 3 |
| 2014 | Structured Dictionaries for Ischemia Estimation in Cardiac BOLD MRI at Rest
Cristian Rusu, Sotirios A. Tsaftaris |
MICCAI (2) | 2 |
| 2014 | Synthetic Generation of Myocardial Blood-Oxygen-Level-Dependent MRI Time Series Via Structural Sparse Decomposition ModelingabstractThis paper aims to identify approaches that generate appropriate synthetic data (computer generated) for cardiac phase-resolved blood-oxygen-level-dependent (CP-BOLD) MRI. CP-BOLD MRI is a new contrast agent- and stress-free approach for examining changes in myocardial oxygenation in response to coronary artery disease. However, since signal intensity changes are subtle, rapid visualization is not possible with the naked eye. Quantifying and visualizing the extent of disease relies on myocardial segmentation and registration to isolate the myocardium and establish temporal correspondences and ischemia detection algorithms to identify temporal differences in BOLD signal intensity patterns. If transmurality of the defect is of interest pixel-level analysis is necessary and thus a higher precision in registration is required. Such precision is currently not available affecting the design and performance of the ischemia detection algorithms. In this work, to enable algorithmic developments of ischemia detection irrespective to registration accuracy, we propose an approach that generates synthetic pixel-level myocardial time series. We do this by 1) modeling the temporal changes in BOLD signal intensity based on sparse multi-component dictionary learning, whereby segmentally derived myocardial time series are extracted from canine experimental data to learn the model; and 2) demonstrating the resemblance between real and synthetic time series for validation purposes. We envision that the proposed approach has the capacity to accelerate development of tools for ischemia detection while markedly reducing experimental costs so that cardiac BOLD MRI can be rapidly translated into the clinical arena for the noninvasive assessment of ischemic heart disease. Cristian Rusu, Rita Morisi, Davide Boschetto, Rohan Dharmakumar, Sotirios A. Tsaftaris |
IEEE Trans. Medical Imaging | 5 |
| 2013 | Application-Aware Approach to Compression and Transmission of H.264 Encoded Video for Automated and Centralized Transportation SurveillanceabstractIn this paper, we present a transportation video coding and wireless transmission system specifically tailored to automated vehicle tracking applications. By taking into account the video characteristics and the lossy nature of the wireless channels, we propose video preprocessing and error control approaches to enhance tracking performance while conserving bandwidth resources and computational power at the transmitter. Compared with current state-of-the-art H.264-based implementations, our system is shown to yield over 80% bitrate savings for comparable tracking accuracy. Zhaofu Chen, Sotirios A. Tsaftaris, Eren Soyak, Aggelos K. Katsaggelos |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2011 | Tracking-optimized quantization for H.264 compression in transportation video surveillance applicationsabstractWe propose a tracking-aware system that removes video components of low tracking interest and optimizes the quantization during compression of frequency coefficients, particularly those that most influence trackers, significantly reducing bitrate while maintaining comparable tracking accuracy. We utilize tracking accuracy as our compression criterion in lieu of mean squared error metrics. The process of optimizing quantization tables suitable for automated tracking can be executed online or offline. The online implementation initializes the encoding procedure for a specific scene, but introduces delay. On the other hand, the offline procedure produces globally optimum quantization tables where the optimization occurs for a collection of video sequences. Our proposed system is designed with low processing power and memory requirements in mind, and as such can be deployed on remote nodes. Using H.264/AVC video coding and a commonly used state-of-the-art tracker we show that while maintaining comparable tracking accuracy our system allows for over 50% bitrate savings on top of existing savings from previous work. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2011 | Channel protection for H.264 compression in transportation video surveillance applicationsabstractThe compression of video and subsequent partial loss of the compressed bitstream can dramatically reduce the accuracy of automated tracking algorithms. This is problematic for centralized applications such as transportation surveillance systems, where remotely captured and compressed video is transmitted over lossy wireless links to a central location for tracking. We propose a low-complexity method for protecting compressed video against channel loss such that the tracking accuracy of decoded and concealed video is maximized. Our algorithm leverages a previous method of video processing that removes components of low tracking interest before compression to minimize bitrate, and uses some of the bitrate savings to introduce redundancy into the transmitted bitstream to reduce the probability of information loss. We show using a common tracker and loss concealment algorithm that our system allows for up to 100% increased tracking accuracy at a given bitrate, or 90% bitrate savings for comparable tracking quality. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2011 | Anomalous video event detection using spatiotemporal context
Junsong Yuan 0001, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
Comput. Vis. Image Underst. | 3 |
| 2011 | Low-Complexity Tracking-Aware H.264 Video Compression for Transportation SurveillanceabstractIn centralized transportation surveillance systems, video is captured and compressed at low processing power remote nodes and transmitted to a central location for processing. Such compression can reduce the accuracy of centrally run automated object tracking algorithms. In typical systems, the majority of communications bandwidth is spent on encoding temporal pixel variations such as acquisition noise or local changes to lighting. We propose a tracking-aware, H.264-compliant compression algorithm that removes temporal components of low tracking interest and optimizes the quantization of frequency coefficients, particularly those that most influence trackers, significantly reducing bitrate while maintaining comparable tracking accuracy. We utilize tracking accuracy as our compression criterion in lieu of mean squared error metrics. Our proposed system is designed with low processing power and memory requirements in mind, and as such can be deployed on remote nodes. Using H.264/AVC video coding and a commonly used state-of-the-art tracker we show that our algorithm allows for over 90% bitrate savings while maintaining comparable tracking accuracy. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Content-aware H.264 encoding for traffic video tracking applicationsabstractThe compression of video can reduce the accuracy of tracking algorithms, which is problematic for centralized applications that rely on remotely captured and compressed video for input. We show the effects of high compression on the features commonly used in real-time video object tracking. We propose a computationally efficient Region of Interest (ROI) extraction method, which is used during standard-compliant H.264 encoding to concentrate bitrate on regions in video most likely to contain objects of tracking interest (vehicles). This algorithm is shown to significantly increase tracking accuracy, which is measured by employing a commonly used automatic tracker. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 2010 | Video anomaly detection in spatiotemporal contextabstractCompared to other approaches that analyze object trajectories, we propose to detect anomalous video events at three levels considering spatiotemporal context of video objects, i.e., point anomaly, sequential anomaly, and co-occurrence anomaly. A hierarchical data mining approach is proposed to achieve this task. At each level, the frequency based analysis is performed to automatically discover regular rules of normal events. The events deviating from these rules are detected as anomalies. Experiments on real traffic video prove that the detected video anomalies are hazardous or illegal according to the traffic rule. Junsong Yuan 0001, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2010 | Quantization optimized H.264 encoding for traffic video tracking applicationsabstractThe compression of video can reduce the accuracy of post-compression tracking algorithms. This is problematic for centralized applications such as traffic surveillance systems, where remotely captured and compressed video is transmitted to a central location for tracking. We propose a low complexity optimization framework that automatically identifies video features critical to tracking and concentrates bitrate on these features via quantization tables. Using the H.264 video coding standard and two commonly used state-of-the-art trackers we show that our algorithm allows for over 60% bitrate savings while maintaining comparable tracking accuracy. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2008 | Local feature extraction for video copy detection in a databaseabstractIn this paper a new content-based copy identification method for video sequences is presented. It is robust to a number of image transformations and particulary robust to compression artifacts. A scale and rotation invariant local image descriptor for corner points in detected key frames is proposed based on a generalized Radon transform. In addition, a distance similarity metric is used that fuses intensity and geometry information to compare key frames extracted using a scene detection algorithm. Furthermore, to achieve low querying computational complexity a DP approach is employed. Experimental results demonstrate the effectiveness of our approach. Ehsan Maani, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2008 | A dynamic programming solution to tracking and elastically matching left ventricular walls in cardiac cine MRIabstractIn this paper an algorithm to detect and elastically match the contours of the epicardial walls of the left ventricle (LV) in cardiac phase-resolved 2-D magnetic resonance (MR) images is presented. For both tasks, dynamic programming (DP) is used. A mask conforming to the six segment model of the LV is fitted on a reference image and propagated utilizing the elastic matching information. At its present form the algorithm requires minimal parameter corrections among different sets of cine MRI images. Future extensions include comparisons with contours hand labeled by imaging experts. Sotirios A. Tsaftaris, Valentin Andermatt, Andre Schlegel, Aggelos K. Katsaggelos, Debiao Li, Rohan Dharmakumar |
ICIP | 1 |
| 2008 | Automated line flattening of Atomic Force Microscopy imagesabstractIn this paper, an automated algorithm to flatten lines from Atomic Force Microscopy (AFM) images is presented. Due to the mechanics of the AFM, there is a curvature distortion (bowing effect) present in the acquired images. At present, flattening such images requires human intervention to manually segment object data from the background, which is time consuming and highly inaccurate. The proposed method classifies the data into objects and background, and fits convex lines in an iterative fashion. Results on real images from DNA wrapped carbon nanotubes (DNA-CNTs) and synthetic experiments are presented, demonstrating the effectiveness of the proposed algorithm in increasing the resolution of the surface topography. Sotirios A. Tsaftaris, Jana Zujovic, Aggelos K. Katsaggelos |
ICIP | 1 |
| 2007 | DNA Microarray Image Intensity Extraction using EigenspotsabstractDNA microarrays are commonly used in the rapid analysis of gene expression in organisms. Image analysis is used to measure the average intensity of circular image areas (spots), which correspond to the level of expression of the genes. A crucial aspect of image analysis is the estimation of the background noise. Currently, background subtraction algorithms are used to estimate the local background noise and subtract it from the signal. In this paper we use principal component analysis (PCA) to de-correlate the signal from the noise, by projecting each spot on the space of eigenvectors, which we term eigenspots. PCA is well suited for such application due to the structural nature of the images. To compare the proposed method with other background estimation methods we use the industry standard signal-to-noise metric xdev. Sotirios A. Tsaftaris, Ramandeep Ahuja, Derek J. Shiell, Aggelos K. Katsaggelos |
ICIP (6) | 1 |
| 2006 | DNA Hybridization as a Similarity Criterion for Querying Digital Signals Stored in DNA DatabasesabstractWe demonstrate via simulation that hybridization of DNA molecules can be used as a similarity criterion for retrieving digital signals encoded and stored in a synthesized DNA database. After introducing some necessary DNA terminology, we briefly explain how digital signals are transformed to DNA sequences. Since retrieval is achieved through hybridization of query and data carrying DNA molecules, we present a mathematical model to estimate hybridization efficiency (also known as selectivity annealing). We show that selectivity annealing is inversely proportional to the mean squared error (MSE) of the encoded signal values. In addition, we show that the concentration of the molecules plays the same role as the decision threshold employed in digital signal matching algorithms. Finally, similar to the digital domain, we define a DNA signal-to-noise ratio (SNR) measure to assess the performance of the DNA-based retrieval scheme. Simulations are presented to validate our arguments Sotirios A. Tsaftaris, Vassily Hatzimanikatis, Aggelos K. Katsaggelos |
ICASSP (2) | 1 |
| 2006 | Joint source-channel coding for wireless object-based video communications utilizing data hidingabstractIn recent years, joint source-channel coding for multimedia communications has gained increased popularity. However, very limited work has been conducted to address the problem of joint source-channel coding for object-based video. In this paper, we propose a data hiding scheme that improves the error resilience of object-based video by adaptively embedding the shape and motion information into the texture data. Within a rate-distortion theoretical framework, the source coding, channel coding, data embedding, and decoder error concealment are jointly optimized based on knowledge of the transmission channel conditions. Our goal is to achieve the best video quality as expressed by the minimum total expected distortion. The optimization problem is solved using Lagrangian relaxation and dynamic programming. The performance of the proposed scheme is tested using simulations of a Rayleigh-fading wireless channel, and the algorithm is implemented based on the MPEG-4 verification model. Experimental results indicate that the proposed hybrid source-channel coding scheme significantly outperforms methods without data hiding or unequal error protection. Haohong Wang, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 2004 | DNA-based matching of digital signalsabstractAdleman with his pioneering work set the stage for the new field of bio-computing research (Science, vol.266, p.1021-1024, 1994). His main idea was to use actual chemistry to solve problems that are either unsolvable by conventional computers, or require an enormous amount of computation. The main focus of our research is to consider the application of molecular computing to the domain of digital signal processing (DSP). In this paper, we consider matching problems that arise in signal processing applications and are amenable to a DNA-based solution. Digital data are encoded in DNA sequences using a sophisticated codeword set that satisfies the noise tolerance constraint (NTC) that we introduce. NTC, one of the main contributions of our work, takes into account the presence of noise in digital signals by exploiting the annealing between non-perfect complementary sequences. We propose an algorithm to map binary values into DNA codewords by satisfying a number of constraints, including the NTC. Using that algorithm, we retrieved 128 codewords that enables us to use a DNA based approach to digital signal matching. Sotirios A. Tsaftaris, Aggelos K. Katsaggelos, Thrasyvoulos N. Pappas, Eleftherios T. Papoutsakis |
ICASSP (5) | 1 |
| 2002 | Compressed-domain video watermarking of MPEG streamsabstractA new technique for watermarking of MPEG compressed video streams is proposed. The watermarking scheme operates directly in the domain of MPEG program streams. Perceptual models are used during the embedding process in order to preserve the video quality. The watermark is embedded in the compressed domain and is detected without the use of the original video sequence. Experimental evaluation demonstrates that the proposed scheme is able to withstand a variety of attacks. The resulting watermarking system is fast and reliable, and is suitable for copyright protection and real-time content authentication applications. Dimitrios Simitopoulos, Sotirios A. Tsaftaris, Nikolaos V. Boulgouris, Michael G. Strintzis |
ICME (1) | 2 |