Ruwan B. Tennakoon

dblp:127/9356 · also Ruwan Bandara Tennakoon · DBLP profile ↗
← Back
30ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-8909-5728ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Joint Modeling of Corruption-Driven and Information-Limited Uncertainty for Robust 3D Gaussian Splatting
abstract
Real-time 3D Gaussian Splatting (3DGS) has emerged as an efficient, high-fidelity alternative to neural radiance fields for novel view synthesis, enabling second-scale training and rendering via GPU rasterization. However, when input image collections contain transient disturbances (e.g., dynamic objects, exposure variations, motion blur) or suffer from sparse view coverage at scene boundaries, 3DGS performance degrades significantly due to reconstruction artifacts such as ghosting, floating points, and blurred surfaces. In this work, we present a unified framework that jointly addresses two types of artifacts: (1) corruption-driven artifacts, caused by transient or occluded content; and (2) information-limited artifacts, caused by insufficient multi-view observations. Our method leverages the training gradient signal, as well as the shape and spatial distribution of Gaussians, to adaptively suppress unreliable splats through a soft-masking strategy, without relying on any pretrained segmentation or feature networks. Extensive experiments on two real-world datasets with dynamic scenes and sparse camera trajectories demonstrate that our approach outperforms state-of-the-art robust 3DGS and uncertainty-pruning techniques in artifact suppression and reconstruction fidelity, while preserving real-time performance.
Zeji Hui, Amirali Khodadadian Gostar, Weiqin Chuah, Alireza Bab-Hadiashar, Ruwan B. Tennakoon
WACV5
2026 Logit-Adjusted Test-Time Adaptation under Partial Class Imbalance
abstract
Test-Time Adaptation (TTA) enables deep neural networks to handle distribution shifts without requiring labels at inference. However, existing methods commonly assume complete class overlap between source and target domains, which rarely holds in practice. We study the challenging setting of Partial Class Imbalance, where the target domain contains only a subset of source classes. We show that entropy minimization–based TTA methods degrade over long test sequences because batch normalization updates bias feature representations toward visible classes, resulting in skewed predictions. To address this, we propose Logit-Adjusted Entropy Minimization, a simple yet effective strategy that integrates target class priors into the adaptation objective. Our method is model-agnostic and can be seamlessly applied to a wide range of TTA algorithms. Extensive experiments on CIFAR-100-C, ImageNetC under diverse corruptions and severity levels, and the large-scale DomainNet-126 dataset demonstrate that our method consistently improves adaptation stability and accuracy for both CNNs and Vision Transformers. Compared to strong baselines, our approach reduces overfitting to visible classes and mitigates performance degradation in long-sequence adaptation. Code is available at https://github.com/thilinawee/latta
Thilina Weerasinghe, Ruwan B. Tennakoon, Weiqin Chuah, Alireza Bab-Hadiashar
WACV2
2026 On Micro-CT inspection of SPR joints in lightweight structures using vision-based machine learning approach
abstract
The paper presents an innovative non-destructive inspection (NDI) method, via employing a custom vision-based machine learning model, for evaluating the quality of Self-Piercing Rivet (SPR) joints, a commonly used joining technique in light vehicle production. Non-destructive inspection (NDI) is crucial for structural integrity analyses, including fracture and fatigue assessment. Utilizing μ-CT (Micro-Computed Tomography) scans of SPR joints and the Machine Vision AI (artificial intelligence) model, the present approach enables the detection and quantification of key quality parameters. The AI model is trained on images extracted from a comprehensive dataset of μ-CT scans of modified SPR joints. Samples used for training are modified carefully to represent the combined defects commonly found in SPR joints. The trained vision-based AI model is deployed to identify and quantify defects, accounting for cumulative errors from material inconsistencies, manufacturing imperfections, and dimensional tolerances. These parameters comprise cumulative defects that may affect joint integrity.The proposed method provides valuable cumulative measured inputs to numerical models of whole riveted structure, non-destructively, which are essential for evaluating the performance of complex structures incorporating SPRs. By significantly enhancing the prediction of the effects of manufacturing defects on lightweight structures, including fatigue life and joint durability, this approach enables more accurate and reliable assessments of individual joints; unlike current techniques, which depend on assumed, averaged, or estimated defect parameters without direct inspection. Ultimately, the AI-based NDI tool enables the enhancement of error detection precision, contributing to safer and more reliable structural integrity evaluations for a variety of manufacturing applications.
Weiqin Chuah, Ruwan B. Tennakoon, Mark Easton, Reza Hoseinnezhad, Alireza Bab-Hadiashar
Eng. Appl. Artif. Intell.2
2026 Uncertainty-Aware Information Pursuit for Interpretable and Reliable Medical Image Analysis
abstract
To be adopted in safety-critical domains like medical image analysis, AI systems must provide human-interpretable decisions. Variational Information Pursuit (VIP) offers an interpretable-by-design framework by sequentially querying input images for human-understandable concepts, using their presence or absence to make predictions. However, existing V-IP methods overlook sample-specific uncertainty in concept predictions, which can arise from ambiguous features or model limitations, leading to suboptimal query selection and reduced robustness. In this paper, we propose an interpretable and uncertainty-aware framework for medical imaging that addresses these limitations by accounting for upstream uncertainties in concept-based, interpretable-by-design models. Specifically, we introduce two uncertainty-aware models, EUAV-IP and IUA-VIP, that integrate uncertainty estimates into the V-IP querying process to prioritize more reliable concepts per sample. EUAV-IP skips uncertain concepts via masking, while IUAV-IP incorporates uncertainty into query selection implicitly for more informed and clinically aligned decisions. Our approach allows models to make reliable decisions based on a subset of concepts tailored to each individual sample, without human intervention, while maintaining overall interpretability. We evaluate our methods on five medical imaging datasets across four modalities: dermoscopy, X-ray, ultrasound, and blood cell imaging. The proposed IUAV-IP model achieves state-of-the-art accuracy among interpretable-by-design approaches on four of the five datasets, and generates more concise explanations by selecting fewer yet more informative concepts. These advances enable more reliable and clinically meaningful outcomes, enhancing model trustworthiness and supporting safer AI deployment in healthcare. Our code and models are available at: https://github.com/Nahiduzzaman09/ UAV-IP.
Md. Nahiduzzaman, Steven Korevaar, ZongYuan Ge, Feng Xia 0001, Alireza Bab-Hadiashar, Ruwan B. Tennakoon
IEEE Trans. Medical Imaging6
2025 Direct Estimation of Attenuation Information from Sinograms for Positron Emission Tomography Reconstruction
abstract
Positron Emission Tomography (PET) is a powerful imaging modality for assessing biochemical processes within the body. However, accurate image reconstruction is challenged by photon attenuation, particularly in dense structures such as bones, leading to quantification errors and reduced diagnostic confidence. Computed Tomography (CT) based attenuation correction is the standard approach but introduces additional radiation exposure, longer imaging times, and patient inconvenience, as well as potential registration errors, motion artifacts, and energy scaling inaccuracies. In this study, we propose a 3D U-Net based deep learning framework that directly estimates attenuation information from PET sinograms, eliminating the need for additional imaging modalities. Our approach integrates PET physics and employs custom skip connections to enhance cross-domain learning. We evaluate our model on a simulated brain dataset derived from real patient templates, achieving a Dice coefficient of 0.650 and an accuracy of 0.486 for bone structures. The clinical applicability of our method is further assessed by reconstructing PET images with the generated attenuation maps, yielding an MSE of 0.007 and an SSIM of 0.956, demonstrating strong structural consistency with CT-based attenuation correction. These results highlight the feasibility of performing PET image attenuation correction using PET sinograms alone, offering a promising alternative that reduces imaging time, radiation exposure, and patient burden while enabling faster and more efficient PET reconstruction.
Prabath Hetti Mudiyanselage, Ruwan B. Tennakoon, John Thangarajah, Robert Ware, Jason H. Callahan
IJCAI2
2025 Joint Structural-Functional Brain Graph Transformer
abstract
Multimodal brain graph transformers have become one of the foundational architectures of graph foundation models for brain science, relying on multimodal brain network fusion. However, most current multimodal brain network fusion methods primarily focus on modality-specific information fusion. The interplays within structural-functional brain networks are often ignored. Therefore, they fail to acquire essential coupling information, which is crucial for obtaining robust joint brain network representations. This oversight inevitably limits the effectiveness and generalization of these representations in various downstream tasks. To this end, we propose a novel joint structural-functional brain graph transformer model (namely sfBGT). Technically, we design a cross-network assortativity quantification mechanism to enable structural-functional brain network coupling, thus capturing the interplays of brain structure and function. We then employ a multimodal graph transformer to effectively learn joint representations of structural-functional brain networks along with their coupling relation representations. Experimental results on three real-world datasets demonstrate the superiority of sfBGT over state-of-the-art baselines.
Ciyuan Peng, Huafei Huang 0001, Tianqi Guo, Chengxuan Meng, Wenhong Zhao, Ruwan B. Tennakoon, Feng Xia 0001
ACM Trans. Intell. Syst. Technol.7
2025 IT-RUDA: Information Theory-Assisted Robust Unsupervised Domain Adaptation
abstract
Domain adaptation is a well-studied field in machine learning. Distribution shift between train (source) and test (target) datasets is a common problem encountered in machine learning applications. One approach to resolve this issue is to use the Unsupervised Domain Adaptation (UDA) technique that carries out knowledge transfer from a label-rich source domain to an unlabeled target domain. Outliers that exist in either source or target datasets can introduce additional challenges when using UDA in practice. In this article, \(\alpha\) -divergence is used as a measure to minimize the discrepancy between the source and target distributions while inheriting robustness, adjustable with a single parameter \(\alpha\) , as the prominent feature of this measure. Here, it is shown that the other well-known divergence-based UDA techniques can be derived as special cases of the proposed method. Furthermore, a theoretical upper bound is derived for the loss in the target domain in terms of the source loss and the \(\alpha\) -divergence between the joint distributions in the two domains. The robustness of the proposed method is validated through testing on several benchmarked datasets in open-set and partial UDA setups where extra classes existing in target and source datasets are considered as outliers. The code is publicly available at https://github.com/rashidis/IT-RUDA .
Shima Rashidi, Ruwan B. Tennakoon, Aref Miri Rekavandi, Papangkorn Jessadatavornwong, Amanda Freis, Garret Huff, Mark Easton, Adrian Mouritz, Reza Hoseinnezhad, Alireza Bab-Hadiashar
ACM Trans. Intell. Syst. Technol.2
2024 Single Domain Generalization via Normalised Cross-correlation Based Convolutions
abstract
Deep learning techniques often perform poorly in the presence of domain shift, where the test data follows a different distribution than the training data. The most practically desirable approach to address this issue is Single Domain Generalization (S-DG), which aims to train robust models using data from a single source. Prior work on S-DG has primarily focused on using data augmentation techniques to generate diverse training data. In this paper, we explore an alternative approach by investigating the robustness of linear operators, such as convolution and dense layers commonly used in deep learning. We propose a novel operator called "XCNorm" that computes the normalized cross-correlation between weights and an input feature patch. This approach is invariant to both affine shifts and changes in energy within a local feature patch and eliminates the need for commonly used non-linear activation functions. We show that deep neural networks composed of this operator are robust to common semantic distribution shifts. Furthermore, our empirical results on single-domain generalization benchmarks demonstrate that our proposed technique performs comparably to the stateof-the-art methods.1
Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
WACV2
2024 ALFREDO: Active Learning with FeatuRe disEntangelement and DOmain adaptation for medical image classification
Dwarikanath Mahapatra, Ruwan B. Tennakoon, Yasmeen M. George, Sudipta Roy 0002, Behzad Bozorgtabar, ZongYuan Ge, Mauricio Reyes 0001
Medical Image Anal.2
2023 Stress visualization in geometrically complex structures using Thermoelastic Stress Analysis and Augmented Reality
abstract
We present a framework for the visualization of mechanical stress using augmented reality (AR) using Thermoelastic Stress Analysis (TSA). The 2D stress images generated by TSA are converted to a 3D stress map using computer vision technology and then superimposed on the real object using AR. Our framework enables in-situ visualization of stress in geometrically complex structural components, which can assist in the design, manufacture, test, and through-life sustainment of failure-critical engineering assets. We also discuss the challenges of such a TSA-AR combination and present a case study that demonstrates the performance and significance of our system.
Ayman Mukhaimar, Ruwan B. Tennakoon, Nik Rajic, Fabio Zambetta, Pier Marzocca, Reza Hoseinnezhad, Chris Brooks, Stephen Van Der Velden, Kheang Khauv
VRST2
2023 Generalized framework for image and video object segmentation using affinity learning and message passing GNNS
abstract
Despite significant amount of work reported in the computer vision literature, segmenting images or videos based on multiple cues such as objectness, texture and motion, is still a challenge. This is particularly true when the number of objects to be segmented is not known or there are objects that are not classified in the training data (unknown objects). A possible remedy to this problem is to utiize graph-based clustering techniques such as Correlation Clustering. It is known that using long range affinities (Lifted multicut), makes correlation clustering more accurate than using only adjacent affinities (Multicut). However, the former is computationally expensive and hard to use. In this paper, we introduce a new framework to perform image/motion segmentation using an affinity learning module and a Message Passing Graph Neural Network (MPGNN). The affinity learning module uses a permutation invariant affinity representation to overcome the multi-object problem. The paper shows, both theoretically and empirically, that the proposed MPGNN aggregates higher order information and thereby converts the Lifted Multicut Problem (LMP) to a Multicut Problem (MP), which is easier and faster to solve. Importantly, the proposed method can be generalized to deal with different clustering problems with the same MPGNN architecture. For instance, our method produces competitive results for single image segmentation (on BSDS dataset) as well as unsupervised video object segmentation (on DAVIS17 dataset), by only changing the feature extraction part. In addition, using an ablation study on the proposed MPGNN architecture, we show that the way we update the parameterized affinities directly contributes to the accuracy of the results.
Sundaram Muthu, Ruwan B. Tennakoon, Tharindu Rathnayake, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
Comput. Vis. Image Underst.2
2023 An Information-Theoretic Method to Automatic Shortcut Avoidance and Domain Generalization for Dense Prediction Tasks
abstract
Deep convolutional neural networks for dense prediction tasks are commonly optimized using synthetic data, as generating pixel-wise annotations for real-world data is laborious. However, the synthetically trained models do not generalize well to real-world environments. This poor "synthetic to real" (S2R) generalization we address through the lens of shortcut learning. We demonstrate that the learning of feature representations in deep convolutional networks is heavily influenced by synthetic data artifacts (shortcut attributes). To mitigate this issue, we propose an Information-Theoretic Shortcut Avoidance (ITSA) approach to automatically restrict shortcut-related information from being encoded into the feature representations. Specifically, our proposed method minimizes the sensitivity of latent features to input variations: to regularize the learning of robust and shortcut-invariant features in synthetically trained models. To avoid the prohibitive computational cost of direct input sensitivity optimization, we propose a practical yet feasible algorithm to achieve robustness. Our results show that the proposed method can effectively improve S2R generalization in multiple distinct dense prediction tasks, such as stereo matching, optical flow, and semantic segmentation. Importantly, the proposed method enhances the robustness of the synthetically trained networks and outperforms their fine-tuned counterparts (on real data) for challenging out-of-domain applications.
Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 ITSA: An Information-Theoretic Approach to Automatic Shortcut Avoidance and Domain Generalization in Stereo Matching Networks
abstract
State-of-the-art stereo matching networks trained only on synthetic data often fail to generalize to more challenging real data domains. In this paper, we attempt to unfold an important factor that hinders the networks from generalizing across domains: through the lens of shortcut learning. We demonstrate that the learning of feature representations in stereo matching networks is heavily influenced by synthetic data artefacts (shortcut attributes). To mitigate this issue, we propose an Information-Theoretic Shortcut Avoidance (ITSA) approach to automatically restrict shortcut-related information from being encoded into the feature representations. As a result, our proposed method learns robust and shortcut-invariant features by minimizing the sensitivity of latent features to input variations. To avoid the prohibitive computational cost of direct input sensitivity optimization, we propose an effective yet feasible algorithm to achieve robustness. We show that using this method, state-of-the-art stereo matching networks that are trained purely on synthetic data can effectively generalize to challenging and previously unseen real data scenarios. Importantly, the proposed method enhances the robustness of the synthetic trained networks to the point that they outperform their fine-tuned counterparts (on real data) for challenging out-of-domain stereo datasets.
Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, Alireza Bab-Hadiashar, David Suter
CVPR2
2022 Maximum Consensus by Weighted Influences of Monotone Boolean Functions
abstract
Maximisation of Consensus (MaxCon) is one of the most widely used robust criteria in computer vision. Tennakoon et al. (CVPR2021), made a connection between MaxCon and estimation of influences of a Monotone Boolean function. In such, there are two distributions involved: the distribution defining the influence measure; and the distribution used for sampling to estimate the influence measure. This paper studies the concept of weighted influences for solving MaxCon. In particular, we study the Bernoulli measures. Theoretically, we prove the weighted influences, under this measure, of points belonging to larger structures are smaller than those of points belonging to smaller structures in general. We also consider another “natural” family of weighting strategies: sampling with uniform measure concentrated on a particular (Hamming) level of the cube. One can choose to have matching distributions: the same for defining the measure as for implementing the sampling. This has the advantage that the sampler is an unbiased estimator of the measure. Based on weighted sampling, we modify the algorithm of Tennakoon et al., and test on both synthetic and real datasets. We show some modest gains of Bernoulli sampling, and we illuminate some of the interactions between structure in data and weighted measures and weighted sampling.
Erchuan Zhang, David Suter, Ruwan B. Tennakoon, Tat-Jun Chin, Alireza Bab-Hadiashar, Giang Truong, Syed Zulqarnain Gilani
CVPR3
2022 Deep Learning-Based Incorporation of Planar Constraints for Robust Stereo Depth Estimation in Autonomous Vehicle Applications
abstract
In autonomous vehicles, depth information for the environment surrounding the vehicle is commonly extracted using time-of-flight (ToF) sensors such as LiDARs and RADARs. Those sensors have some limitations that may potentially degrade the quality and utility of the depth information to a substantial extent. An alternative solution is depth estimation from stereo pairs. However, stereo matching and depth estimation often fails at ill-posed regions including areas with repetitive patterns or textureless surfaces which are commonly found on planar surfaces. This paper focuses on designing an efficient framework for stereo depth estimation, using deep learning technique, that is robust against the mentioned ill-posed regions. With the observation that disparities of all pixels belonging to planar areas (scene plane) viewed by two rectified stereo images can be described using affine transformations, our proposed method predicts pixel-wise affine transformation parameters based on the depth information encoded in the aggregated cost volume. We also introduce a propagation term which enforces all pixels belonging to the same scene plane to be transformed using the same parameters. Disparity can then be computed by multiplying the predicted affine parameters with the corresponding pixel locations. The proposed method was evaluated on several benchmark datasets. We are able to obtain competitive results and at the same time reducing the processing time of common convolution neural network (CNN) in stereo matching by 50%. Analysis of the findings shows that our method can produce reliable results at the ill-posed regions which are challenging to the current state-of-the-arts methods.
Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, Alireza Bab-Hadiashar
IEEE Trans. Intell. Transp. Syst.2
2022 Semantic Guided Long Range Stereo Depth Estimation for Safer Autonomous Vehicle Applications
abstract
Autonomous vehicles in intelligent transportation systems must be able to perform reliable and safe navigation. This necessitates accurate object detection, which is commonly achieved by high-precision depth perception. Existing stereo vision-based depth estimation systems generally involve computation of pixel correspondences and estimation of disparities between rectified image pairs. The estimated disparity values will be converted into depth values in downstream applications. As most applications often work in the depth domain, the accuracy of depth estimation is often more compelling than disparity estimation. However, at large distances (> 50m), the accuracy of disparity estimation does not directly translate to the accuracy of depth estimation. In the context of learning-based stereo systems, this is mainly due to biases imposed by the choices of the disparity-based loss function and the training data. Consequently, the learning algorithms often produce unreliable depth estimates of under-represented foreground objects, particularly at large distances. To resolve this issue, we first analyze the effect of those biases and then propose a pair of depth-based loss functions for foreground objects and background separately. These loss functions can be tuned and can balance the inherent bias of the stereo learning algorithms. The efficacy of our solution is demonstrated by an extensive set of experiments, which are benchmarked against state of the art. We show on the KITTI 2015 benchmark that our proposed solution yields substantial improvements in disparity and depth estimation, particularly for objects located at distances beyond 50 meters, outperforming the previous state of the art by 10%.
Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
IEEE Trans. Intell. Transp. Syst.2
2021 Consensus Maximisation Using Influences of Monotone Boolean Functions
abstract
Consensus maximisation (MaxCon), which is widely used for robust fitting in computer vision, aims to find the largest subset of data that fits the model within some tolerance level. In this paper, we outline the connection between MaxCon problem and the abstract problem of finding the maximum upper zero of a Monotone Boolean Function (MBF) defined over the Boolean Cube. Then, we link the concept of influences (in a MBF) to the concept of outlier (in MaxCon) and show that influences of points belonging to the largest structure in data would generally be smaller under certain conditions. Based on this observation, we present an iterative algorithm to perform consensus maximisation. Results for both synthetic and real visual data experiments show that the MBF based algorithm is capable of generating a near optimal solution relatively quickly. This is particularly important where there are large number of outliers (gross or pseudo) in the observed data.
Ruwan B. Tennakoon, David Suter, Erchuan Zhang, Tat-Jun Chin, Alireza Bab-Hadiashar
CVPR1
2020 Cooperative sensor fusion in centralized sensor networks using Cauchy-Schwarz divergence
Amirali Khodadadian Gostar, Tharindu Rathnayake, Ruwan B. Tennakoon, Alireza Bab-Hadiashar, Giorgio Battistelli, Luigi Chisci, Reza Hoseinnezhad
Signal Process.3
2020 Motion Segmentation of RGB-D Sequences: Combining Semantic and Motion Information Using Statistical Inference
abstract
This paper presents an innovative method for motion segmentation in RGB-D dynamic videos with multiple moving objects. The focus is on finding static, small or slow moving objects (often overlooked by other methods) that their inclusion can improve the motion segmentation results. In our approach, semantic object based segmentation and motion cues are combined to estimate the number of moving objects, their motion parameters and perform segmentation. Selective object-based sampling and correspondence matching are used to estimate object specific motion parameters. The main issue with such an approach is the over segmentation of moving parts due to the fact that different objects can have the same motion (e.g. background objects). To resolve this issue, we propose to identify objects with similar motions by characterizing each motion by a distribution of a simple metric and using a statistical inference theory to assess their similarities. To demonstrate the significance of the proposed statistical inference, we present an ablation study, with and without static objects inclusion, on SLAM accuracy using the TUM-RGBD dataset. To test the effectiveness of the proposed method for finding small or slow moving objects, we applied the method to RGB-D MultiBody and SBM-RGBD motion segmentation datasets. The results showed that we can improve the accuracy of motion segmentation for small objects while remaining competitive on overall measures.
Sundaram Muthu, Ruwan B. Tennakoon, Tharindu Rathnayake, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
IEEE Trans. Image Process.2
2020 Classification of Volumetric Images Using Multi-Instance Learning and Extreme Value Theorem
abstract
Volumetric imaging is an essential diagnostic tool for medical practitioners. The use of popular techniques such as convolutional neural networks (CNN) for analysis of volumetric images is constrained by the availability of detailed (with local annotations) training data and GPU memory. In this paper, the volumetric image classification problem is posed as a multi-instance classification problem and a novel method is proposed to adaptively select positive instances from positive bags during the training phase. This method uses the extreme value theory to model the feature distribution of the images without a pathology and use it to identify positive instances of an imaged pathology. The experimental results, on three separate image classification tasks (i.e. classify retinal OCT images according to the presence or absence of fluid build-ups, emphysema detection in pulmonary 3D-CT images and detection of cancerous regions in 2D histopathology images) show that the proposed method produces classifiers that have similar performance to fully supervised methods and achieves the state of the art performance in all examined test cases.
Ruwan B. Tennakoon, Gerda Bortsova, Silas Nyboe Ørting, Amirali Khodadadian Gostar, Mathilde M. W. Wille, Zaigham Saghir, Reza Hoseinnezhad, Marleen de Bruijne, Alireza Bab-Hadiashar
IEEE Trans. Medical Imaging1
2019 Comparative Analysis of 3D Shape Recognition in the Presence of Data Inaccuracies
abstract
Classification of 3D shapes into physically meaningful categories is one of the most important tasks in understanding the immediate environment. Methods that leverage the recent advancements in deep learning have shown to outperform the traditional approaches. However, performances of those methods have only been analyzed with relatively clean data. Three-dimensional measurement sets (point clouds) produced by 3D scanners are rarely that accurate and often contain noise, outliers or missing points. This paper presents an extensive analysis of the robustness of the state-of-the-art neural network algorithms to realistic data inaccuracies. Our experiments show that the existence of these inaccuracies can significantly affect the performance of the deep learning-based algorithms.
Ayman Mukhaimar, Ruwan B. Tennakoon, Chow Yin Lai, Reza Hoseinnezhad, Alireza Bab-Hadiashar
ICIP2
2019 RETOUCH: The Retinal OCT Fluid Detection and Segmentation Benchmark and Challenge
abstract
Retinal swelling due to the accumulation of fluid is associated with the most vision-threatening retinal diseases. Optical coherence tomography (OCT) is the current standard of care in assessing the presence and quantity of retinal fluid and image-guided treatment management. Deep learning methods have made their impact across medical imaging, and many retinal OCT analysis methods have been proposed. However, it is currently not clear how successful they are in interpreting the retinal fluid on OCT, which is due to the lack of standardized benchmarks. To address this, we organized a challenge RETOUCH in conjunction with MICCAI 2017, with eight teams participating. The challenge consisted of two tasks: fluid detection and fluid segmentation. It featured for the first time: all three retinal fluid types, with annotated images provided by two clinical centers, which were acquired with the three most common OCT device vendors from patients with two different retinal diseases. The analysis revealed that in the detection task, the performance on the automated fluid detection was within the inter-grader variability. However, in the segmentation task, fusing the automated methods produced segmentations that were superior to all individual methods, indicating the need for further improvements in the segmentation performance.
Hrvoje Bogunovic, Freerk G. Venhuizen, Sophie Riedl 0001, Stefanos Apostolopoulos, Alireza Bab-Hadiashar, Ulas Bagci, Mirza Faisal Beg, Loza Bekalo, Qiang Chen 0004, Carlos Ciller, Karthik Gopinath, Amirali Khodadadian Gostar, Kiwan Jeon, Zexuan Ji, Sung Ho Kang, Dara Koozekanani, Donghuan Lu, Dustin Morley, Keshab K. Parhi, Hyoung Suk Park, Abdolreza Rashno, Marinko Sarunic, Saad Shaikh, Jayanthi Sivaswamy, Ruwan B. Tennakoon, Shivin Yadav, Sandro De Zanet, Sebastian M. Waldstein, Bianca S. Gerendas, Caroline C. W. Klaver, Clara I. Sánchez, Ursula Schmidt-Erfurth
IEEE Trans. Medical Imaging25
2018 Deep Multi-instance Volumetric Image Classification with Extreme Value Distributions
Ruwan B. Tennakoon, Amirali Khodadadian Gostar, Reza Hoseinnezhad, Marleen de Bruijne, Alireza Bab-Hadiashar
ACCV (3)1
2018 Non-Bayesian Track-Before-Detect Using Cauchy-Schwarz Divergence-Based Information Fusion
abstract
In this paper we present a novel non-Bayesian filtering method for tracking multiple objects with a particular application in time-lapse cell microscopic video sequence. In our method the heat-map of the frame sequence is extracted and represented as a pseudo-probability hypothesis density of the image. The pseudo-probability hypothesis density is used as measurements and fused with a prior Poisson random finite set density. We employed Cauchy-Schwarz divergence for information fusion. The presented algorithm was tested on a publicly available cell microscopic video sequence.
Amirali Khodadadian Gostar, Tharindu Rathnayake, Ruwan B. Tennakoon, Alireza Bab-Hadiashar, Reza Hoseinnezhad
FUSION3
2018 Visual Inspection of Storm-Water Pipe Systems using Deep Convolutional Neural Networks
abstract
Condition monitoring of storm-water pipe systems are carried-out regularly using semi-automated processors. Semi-automated inspection is time consuming, expensive and produces varying and relatively unreliable results due to operators fatigue and novicity. This paper propose an innovative method to automate the storm-water pipe inspection and condition assessment process which employs a computer vision algorithm based on deep-neural network architecture to classify the defect types automatically. With the proposed method, the operator only needs to guide the robot through each pipe and no longer needs to be an expert. The results obtained on a CCTV video dataset of storm-water pipes shows that the deep neural network architectures trained with data augmentation and transfer learning is capable of achieving high accuracies in identifying the defect types.
Ruwan B. Tennakoon, Reza Hoseinnezhad, Huu Tran, Alireza Bab-Hadiashar
ICINCO (1)1
2018 Robust visual data segmentation: Sampling from distribution of model parameters
abstract
This paper approaches the problem of geometric multi-model fitting as a data segmentation problem. The proposed solution is based on a sequence of sampling hyperedges from a hypergraph, model selection and hypergraph clustering steps. We developed a sampling method that significantly facilitates solving the segmentation problem using a new form of the Markov-Chain-Monte-Carlo (MCMC) method to effectively sample from hyperedge distribution. To sample from this distribution effectively, our proposed Markov Chain includes new ways of long and short jumps to perform exploration and exploitation of all structures. To enhance the quality of samples, a greedy algorithm is used to exploit nearby structure based on the minimization of the Least k th Order Statistics cost function . Unlike common sampling methods, ours does not require any specific prior knowledge about the distribution of models. The output set of samples leads to a clustering solution by which the final model parameters for each segment are obtained. The method competes favorably with the state-of-the-art both in terms of computation power and segmentation accuracy.
Alireza Sadri, Ruwan B. Tennakoon, Reza Hoseinnezhad, Alireza Bab-Hadiashar
Comput. Vis. Image Underst.2
2018 Effective Sampling: Fast Segmentation Using Robust Geometric Model Fitting
abstract
Identifying the underlying models in a set of data points that is contaminated by noise and outliers leads to a highly complex multi-model fitting problem. This problem can be posed as a clustering problem by the projection of higher-order affinities between data points into a graph, which can be clustered using spectral clustering. Calculating all possible higher-order affinities is computationally expensive. Hence, in most cases, only a subset is used. In this paper, we propose an effective sampling method for obtaining a highly accurate approximation of the full graph, which is required to solve multi-structural model fitting problems in computer vision. The proposed method is based on the observation that the usefulness of a graph for segmentation improves as the distribution of the hypotheses that are used to build the graph approaches the distribution of the actual parameters for the given data. In this paper, we approximate this actual parameter distribution by using a th-order statistics-based cost function, and the samples are generated using a greedy algorithm that is coupled with a data sub-sampling strategy. The experimental analysis shows that the proposed method is both accurate and computationally efficient compared with the state-of-the-art robust multi-model fitting techniques. The implementation of the method is publicly available from https://github.com/RuwanT/model-fitting-cbs.
Ruwan B. Tennakoon, Alireza Sadri, Reza Hoseinnezhad, Alireza Bab-Hadiashar
IEEE Trans. Image Process.1
2016 Robust Model Fitting Using Higher Than Minimal Subset Sampling
abstract
Identifying the underlying model in a set of data contaminated by noise and outliers is a fundamental task in computer vision. The cost function associated with such tasks is often highly complex, hence in most cases only an approximate solution is obtained by evaluating the cost function on discrete locations in the parameter (hypothesis) space. To be successful at least one hypothesis has to be in the vicinity of the solution. Due to noise hypotheses generated by minimal subsets can be far from the underlying model, even when the samples are from the said structure. In this paper we investigate the feasibility of using higher than minimal subset sampling for hypothesis generation. Our empirical studies showed that increasing the sample size beyond minimal size ( p ), in particular up to p+2, will significantly increase the probability of generating a hypothesis closer to the true model when subsets are selected from inliers. On the other hand, the probability of selecting an all inlier sample rapidly decreases with the sample size, making direct extension of existing methods unfeasible. Hence, we propose a new computationally tractable method for robust model fitting that uses higher than minimal subsets. Here, one starts from an arbitrary hypothesis (which does not need to be in the vicinity of the solution) and moves until either a structure in data is found or the process is re-initialized. The method also has the ability to identify when the algorithm has reached a hypothesis with adequate accuracy and stops appropriately, thereby saving computational time. The experimental analysis carried out using synthetic and real data shows that the proposed method is both accurate and efficient compared to the state-of-the-art robust model fitting techniques.
Ruwan B. Tennakoon, Alireza Bab-Hadiashar, Zhenwei Cao, Reza Hoseinnezhad, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Nonrigid Registration of Volumetric Images Using Ranked Order Statistics
abstract
Nonrigid image registration techniques using intensity based similarity measures are widely used in medical imaging applications. Due to high computational complexities of these techniques, particularly for volumetric images, finding appropriate registration methods to both reduce the computation burden and increase the registration accuracy has become an intensive area of research. In this paper, we propose a fast and accurate nonrigid registration method for intra-modality volumetric images. Our approach exploits the information provided by an order statistics based segmentation method, to find the important regions for registration and use an appropriate sampling scheme to target those areas and reduce the registration computation time. A unique advantage of the proposed method is its ability to identify the point of diminishing returns and stop the registration process. Our experiments on registration of end-inhale to end-exhale lung CT scan pairs, with expert annotated landmarks, show that the new method is both faster and more accurate than the state of the art sampling based techniques, particularly for registration of images with large deformations.
Ruwan B. Tennakoon, Alireza Bab-Hadiashar, Zhenwei Cao, Marleen de Bruijne
IEEE Trans. Medical Imaging1
2013 Quantification of Smoothing Requirement for 3D Optic Flow Calculation of Volumetric Images
abstract
Complexities of dynamic volumetric imaging challenge the available computer vision techniques on a number of different fronts. This paper examines the relationship between the estimation accuracy and required amount of smoothness for a general solution from a robust statistics perspective. We show that a (surprisingly) small amount of local smoothing is required to satisfy both the necessary and sufficient conditions for accurate optic flow estimation. This notion is called "just enough" smoothing, and its proper implementation has a profound effect on the preservation of local information in processing 3D dynamic scans. To demonstrate the effect of "just enough" smoothing, a robust 3D optic flow method with quantized local smoothing is presented, and the effect of local smoothing on the accuracy of motion estimation in dynamic lung CT images is examined using both synthetic and real image sequences with ground truth.
Alireza Bab-Hadiashar, Ruwan B. Tennakoon, Marleen de Bruijne
IEEE Trans. Image Process.2