Fani Deligianni

dblp:26/1217 · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0003-1306-5017ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 7 first-authorArtificial intelligence and machine learning · 9 · 8 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Learning Semi-Supervised Medical Image Segmentation from Spatial Registration
abstract
Semi-supervised medical image segmentation has shown promise in training models with limited labeled data and abundant unlabeled data. However, state-of-the-art methods ignore a potentially valuable source of unsupervised semantic information-spatial registration transforms between image volumes. To address this, we propose CCT-R, a contrastive cross-teaching framework incorporating registration information. To leverage the semantic information available in registrations between volume pairs, CCT-R incorporates two proposed modules: Registration Supervision Loss (RSL) and Registration-Enhanced Positive Sampling (REPS). The RSL leverages segmentation knowledge derived from transforms between labeled and unlabeled volume pairs, providing an additional source of pseudo-labels. REPS enhances contrastive learning by identifying anatomically-corresponding positives across volumes using registration transforms. Experimental results on two challenging medical segmentation benchmarks demonstrate the effectiveness and superiority of CCT-R across various semi-supervised settings, with as few as one labeled case. Our code is available at https://github.com/kathyliu579/ContrastiveCross-teachingWithRegistration.
Qianying Liu, Paul Henderson, Xiao Gu 0003, Hang Dai, Fani Deligianni
WACV5
2025 Differentially Private Integrated Decision Gradients (IDG-DP) for Radar-Based Human Activity Recognition
abstract
Human motion analysis offers significant potential for healthcare monitoring and early detection of diseases. The advent of radar-based sensing systems has captured the spotlight for they are able to operate without physical contact and they can integrate with pre-existing Wi-Fi networks. They are also seen as less privacy-invasive compared to camera-based systems. However, recent research has shown high accuracy in recognizing subjects or gender from radar gait patterns, raising privacy concerns. This study addresses these issues by investigating privacy vulnerabilities in radar-based Human Activity Recognition (HAR) systems and proposing a novel method for privacy preservation using Differential Privacy (DP) driven by attributions derived with Integrated Decision Gradient (IDG) algorithm. We investigate Black-box Membership Inference Attack (MIA) Models in HAR settings across various levels of attacker-accessible information. We extensively evaluated the effectiveness of the proposed IDG-DP method by designing a CNN-based HAR model and rigorously assessing its resilience against MIAs. Experimental results demonstrate the potential of IDG-DP in mitigating privacy attacks while maintaining utility across all settings, particularly excelling against label-only and shadow model blackbox MIA attacks. This work represents a crucial step towards balancing the need for effective radar-based HAR with robust privacy protection in healthcare environments.
Idris Zakariyya, Linda Tran, Kaushik Bhargav Sivangi, Paul Henderson, Fani Deligianni
WACV5
2024 Fusion of Spatial and Riemannian Features to Enhance Detection of Gait Adaptation Mental States During Rhythmic Auditory Stimulation
abstract
Music has a powerful effect in entraining brain networks that influence both affective states and motor control. The use of Rhythmic Auditory Stimulation (RAS) has shown promising results in regularising and stabilising gait control in patients with neurological problems while alleviating associated depressive symptoms. Brain-computer interfaces (BCIs) can play a pivotal role in shaping music stimulus during these interventions. However, this requires robust detection of mental states during gait adaptation. In this work we investigate the use of Regularised Common Spatial Patterns (RCSP) and Riemannian geometry to detect gait states based on Electroencephalogram (EEG) signals. RCSP are particularly effective on small and noisy datasets while reducing overfitting. Riemannian geometry has proven powerful in analysing the covariance structure of EEG signals that reflect functional brain connectivity. Using a publicly available dataset, we extensively evaluate our methods using two dataset splits. We demonstrated statistically significant results in the dataset split 'individual subjects' with the combination of Regularised Common Spatial Patterns (RCSP) and Riemannian geometry.
Nicole Lai-Tan, Marios G. Philiastides, Fani Deligianni
ACII3
2024 Knowledge Distillation with Global Filters for Efficient Human Pose Estimation
Kaushik Bhargav Sivangi, Fani Deligianni
BMVC2
2024 CL-FML: Cluster-Based & Label-Aware Federated Meta-Learning for On-Demand Classification Tasks
abstract
Distributed analytics involving classification tasks demand robust model training. Real-time arbitrary classification tasks on distributed clients pose challenges due to constraints in data sharing. Federated (Meta)-Learning (FML) has been introduced for global distributed (meta)-model training, which generalizes well over distributed data and classification tasks. Current FML approaches assume fixed labels over unskewed class proportions and data distributions along with uniform task distributions. However, global meta-models can only be used for tasks that do not require addressing arbitrary out-of- distribution label issues. In real-world cases, class imbalance and label shifting are common issues in clients' data. On-demand tasks arriving at clients involve unseen labels. Therefore, 'one (meta)-model-fits-all‘ is not the best option. To address these challenges, we introduce multiple cluster-based meta-models, each one tailored to specific label distribution. Our framework, coined Cluster-based & Label-aware Federated Meta-Learning (CL-FML), involves distributed client clustering based on label shifting and cluster-based FML identifying the most suitable clients to engage per task. CL-FML leverages lightweight data augmentation to deal with arbitrary class-imbalanced tasks. Our comprehensive experiments and comparative assessment against baselines showcase that CL-FML efficiently achieves high accuracy by fast convergence, significantly reducing training rounds and communication load.
Tahani Aladwani, Christos Anagnostopoulos 0001, Shameem A. Puthiya Parambath, Fani Deligianni
DSAA4
2024 Predictive Modelling of Cognitive Workload in VR: An Eye-Tracking Approach
abstract
Cognitive training can boost and sharpen the brain’s abilities to remember, focus, and switch between different tasks. One of the key elements of cognitive training is cognitive load. It allows a manipulation of the intensity of the intervention to suit the participant’s ability level and keep the session enjoyable, i.e. neither too frustrating/hard nor too boring/easy). However, measuring cognitive workload in an objective way is still under-researched and difficult. Here, we have developed a novel sustained attention Virtual Reality (VR) task, using Unity, that aims to predict load in a controlled manner. We demonstrate promising results in that machine learning algorithms can identify perceived as well as objective difficulty of the game accurately, using a combination of eye-tracking and physiological data obtained directly within the VR environment.
Dominik Szczepaniak, Monika Harvey, Fani Deligianni
ETRA3
2024 Autofusion: Fusing Multi-Modalities with Interactions
abstract
In the context of escalating data flows through diverse channels, multimodal machine learning holds the potential to simultaneously process varied data formats from multiple sources, offering robust solutions for uncertainty in various applications. The study highlights the often under-explored correlation information among data modalities and emphasizes the importance of disentangling enriched interactions for more informed decision-making. This paper navigates the burgeoning field of multimodal artificial intelligence (AI) by proposing Autofusion, a pioneering framework addressing representation learning and fusion challenges. Our proposed approach integrates autoencoder structures to address overfitting issues in unimodal machine learning, simultaneously tackling information balancing challenges. The framework’s application in Alzheimer’s disease detection, using DementiaBank’s Pitt corpus, demonstrates promising results, outperforming unimodal methods and showcasing a substantial advantage over traditional fusion techniques. This research significantly contributes by introducing Autofusion as a comprehensive multimodal machine learning solution, demonstrating its efficacy through DementiaBank’s Pitt corpus to detect Alzheimer’s disease, and shedding light on the influential role of cross-modality interaction for enhanced performance in complex applications.
Thuy-Trinh Nguyen, Fani Deligianni, Hoang D. Nguyen
KES2
2024 The Price of Labelling: A Two-Phase Federated Self-learning Approach
Tahani Aladwani, Shameem A. Puthiya Parambath, Christos Anagnostopoulos 0001, Fani Deligianni
ECML/PKDD (4)4
2024 Threat Perception Captured by Emotion, Motor and Empathetic System Responses: A Systematic Review
abstract
The fight or flight phenomena is of evolutionary origin and responsible for the type of defensive behaviours enacted, when in the face of threat. This review attempts to draw the link between fear and aggression as motivational levers for fight or flight defensive behaviours. Furthermore, this review investigates whether human biological motion is modulated by the affective behaviours associated with the fight or flight phenomenon. It examines how threat informed emotion and motor systems have the potential to result in the modulation of empathetic appraisal. This is of interest to this systematic review, as empathetic modulation is crucial to prosocial drive, which has the potential to alleviate the perceived threat. Hence, this review investigates the role of affective computing in capturing the potential outcome of threat perception and associated empathetic responses. To gain a comprehensive understanding of the affective responses and biological motion evoked from threat scenarios, affective computing methods used to capture these neurophysiological and behavioural responses are discussed. A systematic review using Google Scholar and Web of Science was conducted as of 2023, and findings were supplemented by bibliographies of key articles. A total of 26 studies were analysed from initial web searches to explore the topics of empathy, threat perception, fight or flight, fear, aggression, and human motion. Relationships between affective behaviours (fear, aggression) and corresponding motor defensive behaviours (fight or flight) were examined within threat scenarios, and whether existing affective computing methods are effective in capturing these responses, identifying the varying consensus in the literature, challenges, and limitations of existing research.
Elizabeth M. Jacobs, Fani Deligianni, Frank E. Pollick
IEEE Trans. Affect. Comput.2
2023 Multi-Scale Cross Contrastive Learning for Semi-Supervised Medical Image Segmentation
Qianying Liu, Xiao Gu 0003, Paul Henderson, Fani Deligianni
BMVC4
2023 Optimizing Vision Transformers for Medical Image Segmentation
abstract
For medical image semantic segmentation (MISS), Vision Transformers have emerged as strong alternatives to convolutional neural networks thanks to their inherent ability to capture long-range correlations. However, existing research uses off-the-shelf vision Transformer blocks based on linear projections and feature processing which lack spatial and local context to refine organ boundaries. Furthermore, Transformers do not generalize well on small medical imaging datasets and rely on large-scale pre-training due to limited inductive biases. To address these problems, we demonstrate the design of a compact and accurate Transformer network for MISS, CS-Unet, which introduces convolutions in a multi-stage design for hierarchically enhancing spatial and local modeling ability of Transformers. This is mainly achieved by our well-designed Convolutional Swin Transformer (CST) block which merges convolutions with Multi-Head Self-Attention and Feed-Forward Networks for providing inherent localized spatial context and inductive biases. Experiments demonstrate CS-Unet without pre-training out- performs other counterparts by large margins on multi-organ and cardiac datasets with fewer parameters and achieves state-of-the-art performance. Our code is available at Github1.
Qianying Liu, Chaitanya Kaul, Jun Wang 0121, Christos Anagnostopoulos 0001, Roderick Murray-Smith, Fani Deligianni
ICASSP6
2023 Multimodal Machine Learning for Mental Disorder Detection: A Scoping Review
abstract
Recent advancements in machine learning and multimedia technologies have paved new ways for automatic medical diagnosis. In mental health, multimodal inputs such as visual and audible sensing data are promising to investigate the underlying mechanisms of many conditions, such as depression and bipolar disorders. With the increasing burden on healthcare systems, timely diagnosis of mental diseases using multiple modalities might benefit millions of people worldwide. This scoping review provides an exploratory overview of recent multimodal machine learning approaches for mental disorder screening. We also discuss a generalised end-to-end multimodal machine learning pipeline for future research and development of multimodal disease detection.
Thuy-Trinh Nguyen, Viet Hoang-Quoc Pham, Duc-Trong Le, Xuan-Son Vu, Fani Deligianni, Hoang D. Nguyen
KES5
2022 Eye-Tracking for Performance Evaluation and Workload Estimation in Space Telerobotic Training
abstract
Monitoring the mental workload of operators is of paramount importance in space telerobotic training and other teleoperation tasks. Instead of the estimation of task-specific workload, this article aims at investigating the impact of two significant confounding factors (time-pressure and latency) on space teleoperation and explored the use of eye-tracking technology for factor-induced mental workload estimation and performance evaluation. Ten subjects teleoperated a Canadarm2 robot to complete a complex on-orbit assembly task in our photo-realistic training simulator while wearing a head-mounted eye-tracker. To understand how time-pressure and latency influence eye-tracking features works, we first performed the statistical analysis on various features with respect to a single factor and across multiple groups. Next, eye-tracking features extracted from segment data and trial data is used to identify the mental workload induced by confounding factors, which can be used for developing personalized training programs and guaranteeing safe teleoperation. Furthermore, to improve the recognition performance using segment data, we propose the activity ratio and time ratio to characterize the informative segments. Finally, the relationship between simulator-defined performance measures and eye-tracking features is examined. Results show that fixation duration, saccade frequency and duration, pupil diameter, and index of pupillary activity are significant features that can be used in both factor-induced mental workload estimation and task performance evaluation.
Yao Guo 0002, Daniel R. Freer, Fani Deligianni, Guang-Zhong Yang
IEEE Trans. Hum. Mach. Syst.3
2021 Cross-Subject and Cross-Modal Transfer for Generalized Abnormal Gait Pattern Recognition
abstract
For abnormal gait recognition, pattern-specific features indicating abnormalities are interleaved with the subject-specific differences representing biometric traits. Deep representations are, therefore, prone to overfitting, and the models derived cannot generalize well to new subjects. Furthermore, there is limited availability of abnormal gait data obtained from precise Motion Capture (Mocap) systems because of regulatory issues and slow adaptation of new technologies in health care. On the other hand, data captured from markerless vision sensors or wearable sensors can be obtained in home environments, but noises from such devices may prevent the effective extraction of relevant features. To address these challenges, we propose a cascade of deep architectures that can encode cross-modal and cross-subject transfer for abnormal gait recognition. Cross-modal transfer maps noisy data obtained from RGBD and wearable sensors to accurate 4-D representations of the lower limb and joints obtained from the Mocap system. Subsequently, cross-subject transfer allows disentangling subject-specific from abnormal pattern-specific gait features based on a multiencoder autoencoder architecture. To validate the proposed methodology, we obtained multimodal gait data based on a multicamera motion capture system along with synchronized recordings of electromyography (EMG) data and 4-D skeleton data extracted from a single RGBD camera. Classification accuracy was improved significantly in both Mocap and noisy modalities.
Xiao Gu 0003, Yao Guo 0002, Fani Deligianni, Benny P. L. Lo, Guang-Zhong Yang
IEEE Trans. Neural Networks Learn. Syst.3
2020 Improving ECG Classification Interpretability using Saliency Maps
abstract
Cardiovascular disease is a large worldwide healthcare issue; symptoms often present suddenly with minimal warning. The electrocardiogram (ECG) is a fast, simple and reliable method of evaluating the health of the heart, by measuring electrical activity recorded through electrodes placed on the skin. ECGs often need to be analyzed by a cardiologist, taking time which could be spent on improving patient care and outcomes. Because of this, automatic ECG classification systems using machine learning have been proposed, which can learn complex interactions between ECG features and use this to detect abnormalities. However, algorithms built for this purpose often fail to generalize well to unseen data, reporting initially impressive results which drop dramatically when applied to new environments. Additionally, machine learning algorithms suffer a `black-box' issue, in which it is difficult to determine how a decision has been made. This is vital for applications in healthcare, as clinicians need to be able to verify the process of evaluation in order to trust the algorithm. This paper proposes a method for visualizing model decisions across each class in the MIT-BIH arrhythmia dataset, using adapted saliency maps averaged across complete classes to determine what patterns are being learned. We do this by building two algorithms based on state-of-the-art models. This paper highlights how these maps can be used to find problems in the model which could be affecting generalizability and model performance. Comparing saliency maps across complete classes gives an overall impression of confounding variables or other biases in the model, unlike what would be highlighted when comparing saliency maps on an ECG-by-ECG basis.
Yola Jones, Fani Deligianni, Jeff Dalton 0001
BIBE2
2020 Coupled Real-Synthetic Domain Adaptation for Real-World Deep Depth Enhancement
abstract
Advances in depth sensing technologies have allowed simultaneous acquisition of both color and depth data under different environments. However, most depth sensors have lower resolution than that of the associated color channels and such a mismatch can affect applications that require accurate depth recovery. Existing depth enhancement methods use simplistic noise models and cannot generalize well under real-world conditions. In this paper, a coupled real-synthetic domain adaptation method is proposed, which enables domain transfer between high-quality depth simulators and real depth camera information for super-resolution depth recovery. The method first enables the realistic degradation from synthetic images, and then enhances degraded depth data to high quality with a color-guided sub-network. The key advantage of the work is that it generalizes well to real-world datasets without further training or fine-tuning. Detailed quantitative and qualitative results are presented, and it is demonstrated that the proposed method achieves improved performance compared to previous methods fine-tuned on the specific datasets.
Xiao Gu 0003, Yao Guo 0002, Fani Deligianni, Guang-Zhong Yang
IEEE Trans. Image Process.3
2019 Comparison of Brain Networks Based on Predictive Models of Connectivity
abstract
In this study we adopt predictive modelling to identify simultaneously commonalities and differences in multimodal brain networks acquired within subjects. Typically, predictive modelling of functional connectomes from structural connectomes explores commonalities across multimodal imaging data. However, direct application of multivariate approaches such as sparse Canonical Correlation Analysis (sCCA) applies on the vectorised elements of functional connectivity across subjects and it does not guarantee that the predicted models of functional connectivity are Symmetric Positive Matrices (SPD). We suggest an elegant solution based on the transportation of the connectivity matrices on a Riemannian manifold, which notably improves the prediction performance of the model. Randomised lasso is used to alleviate the dependency of the sCCA on the lasso parameters and control the false positive rate. Subsequently, the binomial distribution is exploited to set a threshold statistic that reflects whether a connection is selected or rejected by chance. Finally, we estimate the sCCA loadings based on a de-noising approach that improves the estimation of the coefficients. We validate our approach based on resting-state fMRI and diffusion weighted MRI data. Quantitative validation of the prediction performance shows superior performance, whereas qualitative results of the identification process are promising.
Fani Deligianni, Jonathan D. Clayden, Guang-Zhong Yang
BIBE1
2019 Adaptive Riemannian BCI for Enhanced Motor Imagery Training Protocols
abstract
Traditional methods of training a Brain-Computer Interface (BCI) on motor imagery (MI) data generally involve multiple intensive sessions. The initial sessions produce simple prompts to users, while later sessions additionally provide realtime feedback to users, allowing for human adaptation to take place. However, this protocol only permits the BCI to update between sessions, with little real-time evaluation of how the classifier has improved. To solve this problem, we propose an adaptive BCI training framework which will update the classifier in real time to provide more accurate feedback to the user on 4-class motor imagery data. This framework will require only one session to fully train a BCI to a given subject. Three variations of an adaptive Riemannian BCI were implemented and compared on data from both our own recorded datasets and the commonly used BCI Competition IV Dataset 2a. Results indicate that the fastest and least computationally expensive adaptive BCI was able to correctly classify motor imagery data at a rate 5.8% higher than when using a standard protocol with limited data. In addition it was confirmed that the adaptive BCI automatically improved its performance as more data became available.
Daniel R. Freer, Fani Deligianni, Guang-Zhong Yang
BSN2
2019 From Emotions to Mood Disorders: A Survey on Gait Analysis Methodology
abstract
Mood disorders affect more than 300 million people worldwide and can cause devastating consequences. Elderly people and patients with neurological conditions are particularly susceptible to depression. Gait and body movements can be affected by mood disorders, and thus they can be used as a surrogate sign, as well as an objective index for pervasive monitoring of emotion and mood disorders in daily life. Here we review evidence that demonstrates the relationship between gait, emotions and mood disorders, highlighting the potential of a multimodal approach that couples gait data with physiological signals and home-based monitoring for early detection and management of mood disorders. This could enhance self-awareness, enable the development of objective biomarkers that identify high risk subjects and promote subject-specific treatment.
Fani Deligianni, Yao Guo 0002, Guang-Zhong Yang
IEEE J. Biomed. Health Informatics1
2018 Markerless gait analysis based on a single RGB camera
abstract
Gait analysis is an important tool for monitoring and preventing injuries as well as to quantify functional decline in neurological diseases and elderly people. In most cases, it is more meaningful to monitor patients in natural living environments with low-end equipment such as cameras and wearable sensors. However, inertial sensors cannot provide enough details on angular dynamics. This paper presents a method that uses a single RGB camera to track the 2D joint coordinates with state-of-the-art vision algorithms. Reconstruction of the 3D trajectories uses sparse representation of an active shape model. Subsequently, we extract gait features and validate our results in comparison with a state-of-the-art commercial multi-camera tracking system. Our results are comparable to those from the current literature based on depth cameras and optical markers to extract gait characteristics.
Xiao Gu 0003, Fani Deligianni, Benny P. L. Lo, Wei Chen 0015, Guang-Zhong Yang
BSN2
2017 Deep Learning for Health Informatics
abstract
With a massive influx of multimodality data, the role of data analytics in health informatics has grown rapidly in the last decade. This has also prompted increasing interests in the generation of analytical, data driven models based on machine learning in health informatics. Deep learning, a technique with its foundation in artificial neural networks, is emerging in recent years as a powerful tool for machine learning, promising to reshape the future of artificial intelligence. Rapid improvements in computational power, fast data storage, and parallelization have also contributed to the rapid uptake of the technology in addition to its predictive power and ability to generate automatically optimized high-level features and semantic interpretation from the input data. This article presents a comprehensive up-to-date review of research employing deep learning in health informatics, providing a critical analysis of the relative merit, and potential pitfalls of the technique as well as its future outlook. The paper mainly focuses on key applications of deep learning in the fields of translational bioinformatics, medical imaging, pervasive sensing, medical informatics, and public health.
Daniele Ravì, Charence Wong, Fani Deligianni, Melissa Berthelot, Javier Andreu-Perez, Benny P. L. Lo, Guang-Zhong Yang
IEEE J. Biomed. Health Informatics3
2013 A Framework for Inter-Subject Prediction of Functional Connectivity From Structural Networks
abstract
Functional connections between brain regions are supported by structural connectivity. Both functional and structural connectivity are estimated from in vivo magnetic resonance imaging and offer complementary information on brain organization and function. However, imaging only provides noisy measures, and we lack a good neuroscientific understanding of the links between structure and function. Therefore, inter-subject joint modeling of structural and functional connectivity, the key to multimodal biomarkers, is an open challenge. We present a probabilistic framework to learn across subjects a mapping from structural to functional brain connectivity. Expanding on our previous work [1], our approach is based on a predictive framework with multiple sparse linear regression. We rely on the randomized LASSO to identify relevant anatomo-functional links with some confidence interval. In addition, we describe resting-state functional magnetic resonance imaging in the setting of Gaussian graphical models, on the one hand imposing conditional independences from structural connectivity and on the other hand parameterizing the problem in terms of multivariate autoregressive models. We introduce an intrinsic measure of prediction error for functional connectivity that is independent of the parameterization chosen and provides the means for robust model selection. We demonstrate our methodology with regions within the default mode and the salience network as well as, atlas-based cortical parcellation.
Fani Deligianni, Gaël Varoquaux, Bertrand Thirion, David J. Sharp, Christian Ledig, Robert Leech, Daniel Rueckert
IEEE Trans. Medical Imaging1
2006 Non-rigid 2D-3D Registration with Catheter Tip EM Tracking for Patient Specific Bronchoscope Simulation
Fani Deligianni, Adrian James Chung, Guang-Zhong Yang
MICCAI (1)1
2006 Patient-specific bronchoscopy visualization through BRDF estimation and disocclusion correction
abstract
This paper presents an image-based method for virtual bronchoscope with photo-realistic rendering. The technique is based on recovering bidirectional reflectance distribution function (BRDF) parameters in an environment where the choice of viewing positions, directions, and illumination conditions are restricted. Video images of bronchoscopy examinations are combined with patient-specific three-dimensional (3-D) computed tomography data through two-dimensional (2-D)/3-D registration and shading model parameters are then recovered by exploiting the restricted lighting configurations imposed by the bronchoscope. With the proposed technique, the recovered BRDF is used to predict the expected shading intensity, allowing a texture map independent of lighting conditions to be extracted from each video frame. To correct for disocclusion artefacts, statistical texture synthesis was used to recreate the missing areas. New views not present in the original bronchoscopy video are rendered by evaluating the BRDF with different viewing and illumination parameters. This allows free navigation of the acquired 3-D model with enhanced photo-realism. To assess the practical value of the proposed technique, a detailed visual scoring that involves both real and rendered bronchoscope images is conducted.
Adrian James Chung, Fani Deligianni, Pallav L. Shah, Athol Wells, Guang-Zhong Yang
IEEE Trans. Medical Imaging2
2006 Nonrigid 2-D/3-D Registration for Patient Specific Bronchoscopy Simulation With Statistical Shape Modeling: Phantom Validation
abstract
This paper presents a nonrigid registration two-dimensional/three-dimensional (2-D/3-D) framework and its phantom validation for subject-specific bronchoscope simulation. The method exploits the recent development of five degrees-of-freedom miniaturized catheter tip electromagnetic trackers such that the position and orientation of the bronchoscope can be accurately determined. This allows the effective recovery of unknown camera rotation and airway deformation, which is modelled by an active shape model (ASM). ASM captures the intrinsic variability of the tracheo-bronchial tree during breathing and it is specific to the class of motion it represents. The method reduces the number of parameters that control the deformation, and thus greatly simplifies the optimisation procedure. Subsequently, pq-based registration is performed to recover both the camera pose and parameters of the ASM. Detailed assessment of the algorithm is performed on a deformable airway phantom, with the ground truth data being provided by an additional six degrees-of-freedom electromagnetic (EM) tracker to monitor the level of simulated respiratory motion.
Fani Deligianni, Adrian James Chung, Guang-Zhong Yang
IEEE Trans. Medical Imaging1
2005 Predictive Camera Tracking for Bronchoscope Simulation with CONDensation
Fani Deligianni, Adrian James Chung, Guang Zhong
MICCAI1
2005 Gaze-Contingent Soft Tissue Deformation Tracking for Minimally Invasive Robotic Surgery
George P. Mylonas, Danail Stoyanov, Fani Deligianni, Ara Darzi, Guang-Zhong Yang
MICCAI3
2005 Soft-Tissue Motion Tracking and Structure Estimation for Robotic Assisted MIS Procedures
Danail Stoyanov, George P. Mylonas, Fani Deligianni, Ara Darzi, Guang-Zhong Yang
MICCAI (2)3
2005 VIS-a-VE: Visual Augmentation for Virtual Environments in Surgical Training
abstract
Photo-realistic rendering combined with vision techniques is an important trend in developing next generation surgical simulation devices. Training with simulator is generally low in cost and more efficient than traditional methods that involve supervised learning on actual patients. Incorporating genuine patient data in the simulation can significantly improve the efficacy of training and skills assessment. In this paper, a photo-realistic simulation architecture is described that utilises patient-specific models for training in minimally invasive surgery. The datasets are constructed by combining computer tomographic images with bronchoscopy video of the same patient so that the three dimensional structures and visual appearance are accurately matched. Using simulators enriched by a library of datasets with sufficient patient variability, trainees can experience a wide range of realistic scenarios, including rare pathologies, with correct visual information. In this paper, the matching of CT and video data is accomplished by using a newly developed 2D/3D registration method that exploits a shape from shading similarity measure. Additionally, a method has been devised to allow shading parameter estimation by modelling the bidirectional reflectance distribution function (BRDF) of the visible surfaces. The derived BRDF is then used to predict the expected shading intensity such that a texture map independent of lighting conditions can be extracted. Thus new views can be generated that were not captured in the original bronchoscopy video, thus allowing free navigation of the acquired 3D model with enhanced photo-realism.
Adrian James Chung, Fani Deligianni, Pallav L. Shah, Athol Wells, Guang-Zhong Yang
EuroVis2
2005 Extraction of visual features with eye tracking for saliency driven 2D/3D registration
Adrian James Chung, Fani Deligianni, Xiaopeng Hu 0001, Guang-Zhong Yang
Image Vis. Comput.2
2004 Visual feature extraction via eye tracking for saliency driven 2D/3D registration
abstract
This paper presents a new technique for extracting visual saliency from experimental eye tracking data. An eye-tracking system is employed to determine which features that a group of human observers considered to be salient when viewing a set of video images. With this information, a biologically inspired saliency map is derived by transforming each observed video image into a feature space representation. By using a feature normalisation process based on the relative abundance of visual features within the background image and those dwelled on eye tracking scan paths, features related to visual attention are determined. These features are then back projected to the image domain to determine spatial areas of interest for unseen video images. The strengths and weaknesses of the method are demonstrated with feature correspondence for 2D to 3D image registration of endoscopy videos with computed tomography data. The biologically derived saliency map is employed to provide an image similarity measure that forms the heart of the 2D/3D registration method. It is shown that by only processing selective regions of interest as determined by the saliency map, rendering overhead can be greatly reduced. Significant improvements in pose estimation efficiency can be achieved without apparent reduction in registration accuracy when compared to that of using a non-saliency based similarity measure.
Adrian James Chung, Fani Deligianni, Xiaopeng Hu 0001, Guang-Zhong Yang
ETRA2
2004 Enhancement of Visual Realism with BRDF for Patient Specific Bronchoscopy Simulation
Adrian James Chung, Fani Deligianni, Pallav L. Shah, Athol Wells, Guang-Zhong Yang
MICCAI (2)2
2003 pq-Space Based 2D/3D Registration for Endoscope Tracking
Fani Deligianni, Adrian James Chung, Guang-Zhong Yang
MICCAI (1)1