Arcot Sowmya

dblp:83/3319 · DBLP profile ↗
← Back
105ranked-venue papers
2as first author
32since 2021 · last 2026
0000-0001-9236-5063ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 34 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 14 since 2021Systems, architecture and hardware · 8Software engineering, systems software and programming languages · 7 · 1 first-authorDatabases, data management, data science and information retrieval · 7Security and privacy · 2Human-computer interaction and ubiquitous computing · 2Computer networks · 1
YearPublicationVenuePosition
2026 Object detection and tracking in 360 degree omnidirectional images: A review
abstract
360 omnidirectional imagery enables comprehensive scene understanding, making it highly valuable for applications such as autonomous surveillance, robotic navigation, and immersive virtual environments. Unlike conventional imagery, 360 equirectangular videos introduce unique challenges, including viewpoint variations, geometric distortions, illumination shifts, scale changes, wraparound effects, and high computational demands. This review provides the first integrated survey to jointly analyse object detection, tracking, projection strategies, and datasets in omnidirectional imagery, highlighting challenges, limitations, and future research directions to enhance accuracy, efficiency, and robustness. This review synthesises current methodologies, examining key algorithms and frameworks for panoramic detection and tracking, while also discussing projection techniques that transform omnidirectional inputs into tractable representations. Relevant datasets for training and benchmarking are reviewed to contextualise practical evaluation and reproducibility. Diverse approaches, from deep learning architectures to geometric transformations, are evaluated with attention to their strengths, limitations, and applicability to real-world scenarios. Finally, the review outlines open challenges and emerging research opportunities, providing a roadmap for advancing panoramic object detection and tracking, serving as a comprehensive reference for researchers in the field.
Huma Hafeez, Sankaran Iyer, Arcot Sowmya, Jo Plested, Matthew A. Garratt
Neurocomputing3
2026 Medical hierarchical image classification via dual-geometry image-text learning
abstract
Hierarchical image classification is a fundamental challenge in medical image analysis, as tree-structured taxonomies inherently reflect biological and clinical relationships, spanning the general categorisation of disease entities and fine-grained cellular distinctions. Existing approaches primarily rely on multi-task learning and fine-grained detection, often requiring intricate model design and complex training strategies. In this paper, we aim to exploit the negative curvature property of hyperbolic space, which allows efficient representation of hierarchical structures. We propose a dual-geometry image-text framework, termed H 2 CL. Specifically, we introduce a lightweight classifier head on top of image backbones to extract both Euclidean and hyperbolic features, which are then combined to simultaneously preserve taxonomic consistency from an etiological perspective and enhance instance discrimination from a morphological perspective. Furthermore, a text branch is incorporated to integrate label semantics, where an entailment loss is employed to jointly model image–text alignment and inter-sample relationships. Extensive experiments on cervical cell, skin lesion, and gallbladder disease datasets demonstrate that our framework consistently outperforms advanced methods. Compared to the standard Swin Transformer, H 2 CL achieves an average accuracy improvement of 7% across all three datasets at the fine-grained level, with similarly consistent gains observed when integrated with other backbone models. The source code is publicly available at https://github.com/MCPathology/H2CL .
Lei Fan 0007, Arcot Sowmya, Erik Meijering, ZongYuan Ge, Yang Song 0001
Medical Image Anal.2
2026 CytoAL: Toward Label-Efficient Cytology Diagnosis via Cellularity-Guided Active Learning
abstract
Examining thyroid fine needle aspiration (FNA) can grade cancer risks, derive prognostic information, and guide follow-up care or surgery decision-making. However, thyroid cytology's diagnostic cues are more dispersed compared with pathology images in other disciplines, making standard annotation strategy for AI diagnosis labor-intensive. Inspired by how cytologists diagnose under the microscope, we propose an innovative cellularity-based active learning framework, namely Cyto-AL, to correlate cellularity with diagnostic categories for the active learning query. We also improve the Whole Slide Image (WSI) category of The Bethesda System for Reporting Thyroid Cytology (TBSRTC) prediction by proposing severe-stage pinpointed Multiple Instance Learning (MIL). Additionally, we introduce a lightweight score model to optimize the query in human in the loop (HITL) annotation strategy. Given scarce public thyroid cytology datasets, we release our collected and labeled images as benchmarks. The benchmark comprises 138 WSIs (27,496 valid image patches) collected from 2021-2023 across six classes, annotated by three pathologists using TBSRTC. At patch-level verification, Cyto-AL achieves a 2.2% average classification accuracy improvement over state-of-the-art methods with an equally labeled dataset, and its lightweight ranking-aware model reduces training time by around 65%. Moreover, the WSI-level MIL approach improves average accuracy by 10.7% and Macro-F1 score by 3.5%, outperforming standard sampling methods such as Monte Carlo sampling. The source code and dataset are available at https://github.com/Junchao-Zhu/Cyto-AL.
Junchao Zhu, Yiqing Shen 0003, Rui Fei Du, Arcot Sowmya, Caifeng Wan, Jing Ke
IEEE Trans. Image Process.4
2025 UniMRG: Refining Medical Semantic Understanding Across Modalities via LLM-Orchestrated Synergistic Evolution
Hongyan Xu 0002, Arcot Sowmya, Ian Katz, Dadong Wang
MICCAI (5)2
2025 Benchmarking ensemble machine learning algorithms for multi-class, multi-omics data integration in clinical outcome prediction
abstract
The complementary information found in different modalities of patient data can aid in more accurate modelling of a patient's disease state and a better understanding of the underlying biological processes of a disease. However, the analysis of multi-modal, multi-omics data presents many challenges. In this work, we compare the performance of a variety of ensemble machine learning (ML) algorithms that are capable of late integration of multi-class data from different modalities. The ensemble methods and their variations tested were (i) a voting ensemble, with hard and soft vote, (ii) a meta learner, and (iii) a multi-modal AdaBoost model using hard vote, soft vote, and meta learner to integrate the modalities on each boosting round, the PB-MVBoost model and a novel application of a mixture of expert's model. These were compared to simple concatenation. We examine these methods using data from an in-house study on hepatocellular carcinoma, plus validation datasets on studies from breast cancer and irritable bowel disease. We develop models that achieve an area under the receiver operating curve of up to 0.85 and find that two boosted methods, PB-MVBoost and AdaBoost with soft vote were the best performing models. We also examine the stability of features selected and the size of the clinical signature. Our work shows that integrating complementary omics and data modalities with effective ensemble ML models enhances accuracy in multi-class clinical outcome predictions and produces more stable predictive features than individual modalities or simple concatenation. We provide recommendations for the integration of multi-modal multi-class data.
Annette Spooner, Mohammad Karimi Moridani, Barbra Toplis, Jason Behary, Azadeh Safarchi, Salim Maher, Fatemeh Vafaee, Amany Zekry, Arcot Sowmya
Briefings Bioinform.9
2025 Identifying risk factors for Alzheimer's disease from multivariate longitudinal clinical data using temporal pattern mining
abstract
BACKGROUND: Patient data contain a wealth of information that could aid in understanding the onset and progression of disease. However, the task of modelling clinical data, which consist of multiple heterogeneous time series of different lengths, measured at different time intervals, is a complex one. A growing body of research has applied temporal pattern mining to this problem to identify common patterns in clinical attributes over time. However, the vast majority of these algorithms use techniques that are not ideally suited to clinical data. We present an efficient and scalable framework designed specifically for temporal pattern mining of real-world clinical data. Our framework combines temporal abstraction, an extended version of the efficient pattern-growth algorithm, TPMiner, the concepts of relative risk and the odds ratio to identify interesting and high-risk patterns and multiprocessing to improve computational efficiency. A complete set of cut-off values for discretisation and interpretation of the data is provided and is applicable to studies on ageing populations in general. We name this framework Clinical Temporal Pattern Mining or C-TPM. RESULTS: The framework is applied to data from two real-world studies of Alzheimer's disease (AD). The patterns discovered were predictive of AD in survival analysis models with a Concordance index of up to 0.87 and contain clinically relevant variables. A visualisation module provides a clear picture of the discovered patterns for ease of interpretability. CONCLUSIONS: The framework provides an effective and scalable method of modelling multivariate, longitudinal clinical data and can identify patterns in uncommon diseases and those that progress slowly over time. It is generalisable to clinical data from other medical domains as well as non-clinical data.
Annette Spooner, Gelareh Mohammadi, Perminder S. Sachdev, Henry Brodaty, Arcot Sowmya
BMC Bioinform.5
2025 Efficient transformer with compressed-attention for stereo image super-resolution
abstract
While self-attention mechanisms in transformers exhibit superior performance in image super-resolution tasks, improving efficiency remains a challenge. To enhance the efficiency of self-attention mechanisms for stereo image super-resolution, we propose an efficient transformer with compressed attention for stereo image super-resolution (ETCASSR). Specifically, we propose a simple yet effective compressed-attention mechanism that organizes channels from partial to full for attention operations. Using this mechanism, we develop a compressed window-based self-attention block and a compressed transposed self-attention block, enabling efficient intra-view feature extraction. To further enrich feature representation, we introduce a spatial local feature branch and a channel global feature branch to complement these two blocks. Furthermore, a compressed cross-attention block for cross-view feature extraction is designed by extending the compressed-attention mechanism. Combining these blocks, ETCASSR achieves state-of-the-art performance on stereo image super-resolution while maintaining low computational complexity and fast running speed. Additionally, we introduce ETCASR for single-image super-resolution by omitting the cross-view components from ETCASSR, also achieving superior performance with high efficiency. The proposed transformers offer significant potential applications in other vision tasks. Source code will be released at https://github.com/jianwensong/ETCASSR .
Jianwen Song, Arcot Sowmya, Changming Sun
Knowl. Based Syst.2
2025 Efficient frequency feature aggregation transformer for image super-resolution
abstract
Although vision transformers have shown remarkable performance in image super-resolution tasks, the key component, i.e., the self-attention mechanism, suffers from insufficient high-frequency information extraction capability and high computational costs, hindering further advancement. To address these limitations, we propose an efficient frequency feature aggregation transformer for single image super-resolution (EFATSR). Specifically, a frequency self-attention aggregation block is proposed to enhance the extraction of high-frequency information. This block incorporates a frequency spatial feature aggregation branch to supplement high-frequency feature extraction for a self-attention branch, enabling the model to capture high-frequency information more effectively. Additionally, a frequency channel-spatial aggregation block is proposed to extract channel and spatial features in the frequency domain, enhancing the efficiency of deep feature extraction. Extensive experiments on single image super-resolution demonstrate that EFATSR achieves state-of-the-art performance while maintaining low computational complexity. Furthermore, we extend EFATSR for stereo image super-resolution by incorporating a multi-head parallax-attention block, forming EFATSSR, which also shows remarkable performance and high efficiency. Source code is avaliable at https://github.com/jianwensong/EFATSR .
Jianwen Song, Arcot Sowmya, Changming Sun
Pattern Recognit.2
2024 SCD-NAS: Towards Zero-Cost Training in Melanoma Diagnosis
abstract
Diagnosing melanoma remains challenging despite advances in Convolutional Neural Networks (CNNs) for skin cancer detection. Their application in clinical settings is often limited by differences between natural and clinical images. To address this, we introduce the Skin Cancer Detection Neural Architecture Search (SCD-NAS) framework. In our method, Large Language Model (LLM) is leveraged as a proxy, which helps SCD-NAS achieve cost-free training. Additionally, to maximize the benefits of various architectural design spaces, we introduce a Search Space Expansion (SSE) methodology. This effectively combines the merits of diverse architectural configurations, thereby enhancing model performance. We conducted experiments on the ISIC 2020, MedMNISTv2, CIFAR-10 and CIFAR-100 datasets. Our SCD-NAS-derived ResNet50 model achieved an Area Under the Curve (AUC) of 91.23% on the ISIC 2020 dataset, improving the baseline by 5.93%. It also exceeded the CIFAR-10 benchmark by 2.45% in accuracy.
Hongyan Xu 0002, Xiu Su, Arcot Sowmya, Ian Katz, Dadong Wang
ICME3
2024 AMFP-net: Adaptive multi-scale feature pyramid network for diagnosis of pneumoconiosis from chest X-ray images
abstract
Early detection of pneumoconiosis by routine health screening of workers in the mining industry is critical for preventing the progression of this incurable disease. Automated pneumoconiosis classification in chest X-ray images is challenging due to the low contrast of opacities, inter-class similarity, intra-class variation and the existence of artifacts. Compared to traditional methods, convolutional neural networks have shown significant improvement in pneumoconiosis classification tasks, however, accurate classification remains challenging due to mainly the inability to focus on semantically meaningful lesion opacities. Most existing networks focus on high level abstract information and ignore low level detailed object information. Different from natural images where an object occupies large space, the classification of pneumoconiosis depends on the density of small opacities inside the lung. To address this issue, we propose a novel two-stage adaptive multi-scale feature pyramid network called AMFP-Net for the diagnosis of pneumoconiosis from chest X-rays. The proposed model consists of 1) an adaptive multi-scale context block to extract rich contextual and discriminative information and 2) a weighted feature fusion module to effectively combine low level detailed and high level global semantic information. This two-stage network first segments the lungs to focus more on relevant regions by excluding irrelevant parts of the image, and then utilises the segmented lungs to classify pneumoconiosis into different categories. Extensive experiments on public and private datasets demonstrate that the proposed approach can outperform state-of-the-art methods for both segmentation and classification.
Md. Shariful Alam, Dadong Wang, Arcot Sowmya
Artif. Intell. Medicine3
2024 Efficient masked feature and group attention network for stereo image super-resolution
abstract
Current stereo image super-resolution methods do not fully exploit cross-view and intra-view information, resulting in limited performance. While vision transformers have shown great potential in super-resolution, their application in stereo image super-resolution is hindered by high computational demands and insufficient channel interaction. This paper introduces an efficient masked feature and group attention network for stereo image super-resolution (EMGSSR) designed to integrate the strengths of transformers into stereo super-resolution while addressing their inherent limitations. Specifically, an efficient masked feature block is proposed to extract local features from critical areas within images, guided by sparse masks. A group-weighted cross-attention module consisting of group-weighted cross-view feature interactions along epipolar lines is proposed to fully extract cross-view information from stereo images. Additionally, a group-weighted self-attention module consisting of group-weighted self-attention feature extractions with different local windows is proposed to effectively extract intra-view information from stereo images. Experimental results demonstrate that the proposed EMGSSR outperforms state-of-the-art methods at relatively low computational costs. The proposed EMGSSR offers a robust solution that effectively extracts cross-view and intra-view information for stereo image super-resolution, bringing a promising direction for future research in high-fidelity stereo image super-resolution. Source codes will be released at https://github.com/jianwensong/EMGSSR .
Jianwen Song, Arcot Sowmya, Jien Kato, Changming Sun
Image Vis. Comput.2
2024 BioFusionNet: Deep Learning-Based Survival Risk Stratification in ER+ Breast Cancer Through Multifeature and Multimodal Data Fusion
abstract
Breast cancer is a significant health concern affecting millions of women worldwide. Accurate survival risk stratification plays a crucial role in guiding personalised treatment decisions and improving patient outcomes. Here we present BioFusionNet, a deep learning framework that fuses image-derived features with genetic and clinical data to obtain a holistic profile and achieve survival risk stratification of ER+ breast cancer patients. We employ multiple self-supervised feature extractors (DINO and MoCoV3) pretrained on histopathological patches to capture detailed image features. These features are then fused by a variational autoencoder and fed to a self-attention network generating patient-level features. A co-dual-cross-attention mechanism combines the histopathological features with genetic data, enabling the model to capture the interplay between them. Additionally, clinical data is incorporated using a feed-forward network, further enhancing predictive performance and achieving comprehensive multimodal feature integration. Furthermore, we introduce a weighted Cox loss function, specifically designed to handle imbalanced survival data, which is a common challenge. Our model achieves a mean concordance index of 0.77 and a time-dependent area under the curve of 0.84, outperforming state-of-the-art methods. It predicts risk (high versus low) with prognostic significance for overall survival in univariate analysis (HR=2.99, 95% CI: 1.88-4.78, p 0.005), and maintains independent significance in multivariate analysis incorporating standard clinicopathological variables (HR=2.91, 95% CI: 1.80-4.68, p 0.005).
Raktim Kumar Mondol, Ewan K. A. Millar, Arcot Sowmya, Erik Meijering
IEEE J. Biomed. Health Informatics3
2024 Efficient Hybrid Feature Interaction Network for Stereo Image Super-Resolution
abstract
It is very challenging to fully use cross-view information for stereo image super-resolution. Previous methods using pixel-based parallax-attention mechanisms do not consider neighborhood pixels. Also, they typically use convolutions for basic feature extraction, which may not be as effective as modern self-attention mechanisms in transformers. To address these limitations, we propose an efficient hybrid feature interaction network for stereo image super-resolution. Specifically, we propose a shifted cross-view interaction block that integrates neighborhood pixels and imposes constraints on the disparity range during cross-view interactions. In addition, we propose a hybrid feature interaction block consisting of local and global interaction branches for extracting intra-view features efficiently. In this block, we propose a design that incorporates lightweight attention connections and a partial downsampling operation to enhance spatial and channel feature interaction with high efficiency. Additionally, a dilated efficient channel attention mechanism is proposed to obtain cross-channel interactions within features. Experimental results evaluated on various metrics (PSNR, SSIM, and LPIPS) demonstrate that the proposed method achieves state-of-the-art stereo image super-resolution performance at relatively low computational cost. Moreover, the super-resolution images obtained by the proposed method achieve the smallest stereo matching errors compared to other methods. Source code will be publicly available athttps://github.com/jianwensong/EHFSSR.
Jianwen Song, Arcot Sowmya, Changming Sun
IEEE Trans. Multim.2
2023 Detection of Basal Cell Carcinoma in Whole Slide Images
Hongyan Xu 0002, Dadong Wang, Arcot Sowmya, Ian Katz
MICCAI (6)3
2023 Imbalanced classification for protein subcellular localization with multilabel oversampling
abstract
MOTIVATION: Subcellular localization of human proteins is essential to comprehend their functions and roles in physiological processes, which in turn helps in diagnostic and prognostic studies of pathological conditions and impacts clinical decision-making. Since proteins reside at multiple locations at the same time and few subcellular locations host far more proteins than other locations, the computational task for their subcellular localization is to train a multilabel classifier while handling data imbalance. In imbalanced data, minority classes are underrepresented, thus leading to a heavy bias towards the majority classes and the degradation of predictive capability for the minority classes. Furthermore, data imbalance in multilabel settings is an even more complex problem due to the coexistence of majority and minority classes. RESULTS: Our studies reveal that based on the extent of concurrence of majority and minority classes, oversampling of minority samples through appropriate data augmentation techniques holds promising scope for boosting the classification performance for the minority classes. We measured the magnitude of data imbalance per class and the concurrence of majority and minority classes in the dataset. Based on the obtained values, we identified minority and medium classes, and a new oversampling method is proposed that includes non-linear mixup, geometric and colour transformations for data augmentation and a sampling approach to prepare minibatches. Performance evaluation on the Human Protein Atlas Kaggle challenge dataset shows that the proposed method is capable of achieving better predictions for minority classes than existing methods. AVAILABILITY AND IMPLEMENTATION: Data used in this study are available at https://www.kaggle.com/competitions/human-protein-atlas-image-classification/data. Source code is available at https://github.com/priyarana/Protein-subcellular-localisation-method. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Priyanka Rana, Arcot Sowmya, Erik Meijering, Yang Song 0001
Bioinform.2
2023 Ensemble feature selection with data-driven thresholding for Alzheimer's disease biomarker discovery
abstract
BACKGROUND: Feature selection is often used to identify the important features in a dataset but can produce unstable results when applied to high-dimensional data. The stability of feature selection can be improved with the use of feature selection ensembles, which aggregate the results of multiple base feature selectors. However, a threshold must be applied to the final aggregated feature set to separate the relevant features from the redundant ones. A fixed threshold, which is typically used, offers no guarantee that the final set of selected features contains only relevant features. This work examines a selection of data-driven thresholds to automatically identify the relevant features in an ensemble feature selector and evaluates their predictive accuracy and stability. Ensemble feature selection with data-driven thresholding is applied to two real-world studies of Alzheimer's disease. Alzheimer's disease is a progressive neurodegenerative disease with no known cure, that begins at least 2-3 decades before overt symptoms appear, presenting an opportunity for researchers to identify early biomarkers that might identify patients at risk of developing Alzheimer's disease. RESULTS: The ensemble feature selectors, combined with data-driven thresholds, produced more stable results, on the whole, than the equivalent individual feature selectors, showing an improvement in stability of up to 34%. The most successful data-driven thresholds were the robust rank aggregation threshold and the threshold algorithm threshold from the field of information retrieval. The features identified by applying these methods to datasets from Alzheimer's disease studies reflect current findings in the AD literature. CONCLUSIONS: Data-driven thresholds applied to ensemble feature selectors provide more stable, and therefore more reproducible, selections of features than individual feature selectors, without loss of performance. The use of a data-driven threshold eliminates the need to choose a fixed threshold a-priori and can select a more meaningful set of features. A reliable and compact set of features can produce more interpretable models by identifying the factors that are important in understanding a disease.
Annette Spooner, Gelareh Mohammadi, Perminder S. Sachdev, Henry Brodaty, Arcot Sowmya
BMC Bioinform.5
2023 Super-Resolution Phase Retrieval Network for Single-Pattern Structured Light 3D Imaging
abstract
Structured light 3D imaging is often used for obtaining accurate 3D information via phase retrieval. Single-pattern structured light 3D imaging is much faster than multi-pattern versions. Current phase retrieval methods for single-pattern structured light 3D imaging are however not accurate enough. Besides, the projector resolution in a structured light 3D imaging system is expensive to improve due to hardware costs. To address the issues of low accuracy and low resolution of single-pattern structured light 3D imaging, this work proposes a super-resolution phase retrieval network (SRPRNet). Specifically, a phase-shifting module is proposed to extract multi-scale features with different phase shifts, and a refinement and super-resolution module is proposed to obtain refined and super-resolution phase components. After phase demodulation and unwrapping, high-resolution absolute phase is obtained. A sine shifting loss and a cosine shifting loss are also introduced to form the regularization term of the loss function. As far as can be ascertained, the proposed SRPRNet is the first network for super-resolution phase retrieval by using a single pattern, and it can also be used for standard-resolution phase retrieval. Experimental results on three datasets show that SRPRNet achieves state-of-the-art performance on $1\times $ , $2\times $ , and $4\times $ super-resolution phase retrieval tasks.
Jianwen Song, Kai Liu 0012, Arcot Sowmya, Changming Sun
IEEE Trans. Image Process.3
2023 A Multi-Scale Context Aware Attention Model for Medical Image Segmentation
abstract
Medical image segmentation is critical for efficient diagnosis of diseases and treatment planning. In recent years, convolutional neural networks (CNN)-based methods, particularly U-Net and its variants, have achieved remarkable results on medical image segmentation tasks. However, they do not always work consistently on images with complex structures and large variations in regions of interest (ROI). This could be due to the fixed geometric structure of the receptive fields used for feature extraction and repetitive down-sampling operations that lead to information loss. To overcome these problems, the standard U-Net architecture is modified in this work by replacing the convolution block with a dilated convolution block to extract multi-scale context features with varying sizes of receptive fields, and adding a dilated inception block between the encoder and decoder paths to alleviate the problem of information recession and the semantic gap between features. Furthermore, the input of each dilated convolution block is added to the output through a squeeze and excitation unit, which alleviates the vanishing gradient problem and improves overall feature representation by re-weighting the channel-wise feature responses. The original inception block is modified by reducing the size of the spatial filter and introducing dilated convolution to obtain a larger receptive field. The proposed network was validated on three challenging medical image segmentation tasks with varying size ROIs: lung segmentation on chest X-ray (CXR) images, skin lesion segmentation on dermoscopy images and nucleus segmentation on microscopy cell images. Improved performance compared to state-of-the-art techniques demonstrates the effectiveness and generalisability of the proposed Dilated Convolution and Inception blocks-based U-Net (DCI-UNet).
Md. Shariful Alam, Dadong Wang, Qiyu Liao, Arcot Sowmya
IEEE J. Biomed. Health Informatics4
2023 A Federated Learning System for Histopathology Image Analysis With an Orchestral Stain-Normalization GAN
abstract
Currently, data-driven based machine learning is considered one of the best choices in clinical pathology analysis, and its success is subject to the sufficiency of digitized slides, particularly those with deep annotations. Although centralized training on a large data set may be more reliable and more generalized, the slides to the examination are more often than not collected from many distributed medical institutes. This brings its own challenges, and the most important is the assurance of privacy and security of incoming data samples. In the discipline of histopathology image, the universal stain-variation issue adds to the difficulty of an automatic system as different clinical institutions provide distinct stain styles. To address these two important challenges in AI-based histopathology diagnoses, this work proposes a novel conditional Generative Adversarial Network (GAN) with one orchestration generator and multiple distributed discriminators, to cope with multiple-client based stain-style normalization. Implemented within a Federated Learning (FL) paradigm, this framework well preserves data privacy and security. Additionally, the training consistency and stability of the distributed system are further enhanced by a novel temporal self-distillation regularization scheme. Empirically, on large cohorts of histopathology datasets as a benchmark, the proposed model matches the performance of conventional centralized learning very closely. It also outperforms state-of-the-art stain-style transfer methods on the downstream Federated Learning image classification task, with an accuracy increase of over 20.0% in comparison to the baseline classification model.
Yiqing Shen 0003, Arcot Sowmya, Yulin Luo, Xiaoyao Liang, Dinggang Shen, Jing Ke
IEEE Trans. Medical Imaging2
2023 Cancer Survival Prediction From Whole Slide Images With Self-Supervised Learning and Slide Consistency
abstract
Histopathological Whole Slide Images (WSIs) at giga-pixel resolution are the gold standard for cancer analysis and prognosis. Due to the scarcity of pixel- or patch-level annotations of WSIs, many existing methods attempt to predict survival outcomes based on a three-stage strategy that includes patch selection, patch-level feature extraction and aggregation. However, the patch features are usually extracted by using truncated models (e.g. ResNet) pretrained on ImageNet without fine-tuning on WSI tasks, and the aggregation stage does not consider the many-to-one relationship between multiple WSIs and the patient. In this paper, we propose a novel survival prediction framework that consists of patch sampling, feature extraction and patient-level survival prediction. Specifically, we employ two kinds of self-supervised learning methods, i.e. colorization and cross-channel, as pretext tasks to train convnet-based models that are tailored for extracting features from WSIs. Then, at the patient-level survival prediction we explicitly aggregate features from multiple WSIs, using consistency and contrastive losses to normalize slide-level features at the patient level. We conduct extensive experiments on three large-scale datasets: TCGA-GBM, TCGA-LUSC and NLST. Experimental results demonstrate the effectiveness of our proposed framework, as it achieves state-of-the-art performance in comparison with previous studies, with concordance index of 0.670, 0.679 and 0.711 on TCGA-GBM, TCGA-LUSC and NLST, respectively.
Lei Fan 0007, Arcot Sowmya, Erik Meijering, Yang Song 0001
IEEE Trans. Medical Imaging2
2022 Data Agnostic Filter Gating For Efficient Deep Networks
abstract
Filter pruning is essential for deploying a well-trained CNN model on edge computation devices with a target computation budget (e.g., FLOPs). Current filter pruning methods mainly focus on leveraging feature maps to analyze the importance of filters, and prune those with less impact on the value of the CNN’s loss function, thereby ignoring the variance of input batches to differences in sparse structure over the filters. In this paper, we propose a data-agnostic filter pruning method that uses an auxiliary network named Dagger module to induce pruning with the pre-trained weights as input. Besides, to help prune filters with a preset FLOPs constraint, we utilize an explicit FLOPs-aware regularisation mechanism to directly promote pruning filters toward the target FLOPs. Experimental results on CIFAR-10 and ImageNet datasets show that the proposed filter pruning method surpasses the state-of-the-art.
Hongyan Xu 0002, Xiu Su, Shan You, Tao Huang 0020, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002, Dadong Wang, Arcot Sowmya
ICASSP10
2022 Vertebral Compression Fracture detection using Multiple Instance Learning and Majority Voting
abstract
Vertebral compression fractures (VCF) often miss detection in radiology scans, risking more severe secondary fractures in the future leading to permanent disability and death. Automated solutions are therefore desirable, however a frequent bottleneck in medical image analysis is the availability of radiologist’s time for annotations. To alleviate this problem, this work presents the first attempt at VCF detection using Multiple Instance Learning (MIL), a weakly supervised learning approach that can cope with a small annotated data set. The method involves localisation of the thoracic and lumbar spine regions by generating 6 bounding boxes from which 2D patches are extracted. These patches are then used as instances in a bag within an MIL setting to train a deep learning architecture using an algorithm employing an embedded space paradigm with a shared convolutional neural network (CNN) layer. Majority voting is then performed on the results of the 6 bounding boxes to achieve accuracy / F1 score of 81.05% / 80.74% for thoracic and 85.45 % / 85.61% for lumbar spine respectively.
Sankaran Iyer, Alan Blair 0001, Laughlin Dawes, Daniel Aaron Moses, Arcot Sowmya
ICPR6
2022 Multi-scale alignment and Spatial ROI Module for COVID-19 Diagnosis
abstract
Coronavirus Disease 2019 (COVID-19) has spread globally and become a health crisis faced by humanity since first reported. Radiology imaging technologies such as computer tomography (CT) and chest X-ray imaging (CXR) are effective tools for diagnosing COVID-19. However, in CT and CXR images, the infected area occupies only a small part of the image. Some common deep learning methods that integrate large-scale receptive fields may cause the loss of image detail, resulting in the omission of the region of interest (ROI) in COVID-19 images and are therefore not suitable for further processing. To this end, we propose a deep spatial pyramid pooling (D-SPP) module to integrate contextual information over different resolutions, aiming to extract information under different scales of COVID-19 images effectively. Besides, we propose a COVID-19 infection detection (CID) module to draw attention to the lesion area and remove interference from irrelevant information. Extensive experiments on four CT and CXR datasets have shown that our method produces higher accuracy of detecting COVID-19 lesions in CT and CXR images. It can be used as a computer-aided diagnosis tool to help doctors effectively diagnose and screen for COVID-19.
Hongyan Xu 0002, Dadong Wang, Arcot Sowmya
IJCNN3
2022 Fast FF-to-FFPE Whole Slide Image Translation via Laplacian Pyramid and Contrastive Learning
Lei Fan 0007, Arcot Sowmya, Erik Meijering, Yang Song 0001
MICCAI (2)2
2022 Training radiomics-based CNNs for clinical outcome prediction: Challenges, strategies and findings
Shuchao Pang, Matthew Field, Jason Dowling, Shalini K. Vinod, Lois Holloway, Arcot Sowmya
Artif. Intell. Medicine6
2022 Blood-based transcriptomic signature panel identification for cancer diagnosis: benchmarking of feature extraction methods
abstract
Liquid biopsy has shown promise for cancer diagnosis due to its minimally invasive nature and the potential for novel biomarker discovery. However, the low concentration of relevant blood-based biosources and the heterogeneity of samples (i.e. the variability of relative abundance of molecules identified), pose major challenges to biomarker discovery. Moreover, the number of molecular measurements or features (e.g. transcript read counts) per sample could be in the order of several thousand, whereas the number of samples is often substantially lower, leading to the curse of dimensionality. These challenges, among others, elucidate the importance of a robust biomarker panel identification or feature extraction step wherein relevant molecular measurements are identified prior to classification for cancer detection. In this work, we performed a benchmarking study on 12 feature extraction methods using transcriptomic profiles derived from different blood-based biosources. The methods were assessed both in terms of their predictive performance and the robustness of the biomarker panels in diagnosing cancer or stratifying cancer subtypes. While performing the comparison, the feature extraction methods are categorized into feature subset selection methods and transformation methods. A transformation feature extraction method, namely partial least square discriminant analysis, was found to perform consistently superior in terms of classification performance. As part of the benchmarking study, a generic pipeline has been created and made available as an R package to ensure reproducibility of the results and allow for easy extension of this study to other datasets (https://github.com/VafaeeLab/bloodbased-pancancer-diagnosis).
Abhishek Vijayan, Shadma Fatima, Arcot Sowmya, Fatemeh Vafaee
Briefings Bioinform.3
2022 Complex shearlets and rotary phase congruence tensor for corner detection
abstract
Corner detection algorithms based on multi-scale analysis attract more attention due to their promising performance. However, they only consider amplitude information, neglect phase information and partially utilize multi-scale decomposition coefficients to detect corners. This limits their detection accuracy, repeatability and localization ability. This paper describes a new multi-scale analysis based corner detector. To overcome the problems of bilateral margin responses, edge extension and lack of phase information in traditional shearlets, a novel complex shearlet transform is proposed to better localize distributed discontinuities and especially to extract phase information from geometrical features. Moreover, a new rotary phase congruence tensor is proposed to utilize all amplitude and phase information for corner detection. Its tolerances to noise and ability for corner localization are improved further by screening and normalizing the amplitude information. Experimental results demonstrate that the localization ability and detection accuracy of the proposed method are superior to current detectors, and its repeatability is generally higher than current detectors and recent machine learning based interest point detectors.
Changming Sun, Arcot Sowmya
Pattern Recognit.3
2021 PhonicsGAN: Synthesizing Graphical Videos from Phonics Songs
Nuha Aldausari, Arcot Sowmya, Nadine Marcus, Gelareh Mohammadi
ICANN (2)2
2021 Bidirectional Convolutional-LSTM based Network for lung segmentation of chest X-ray images
abstract
Deep Neural Networks (DNN)-based methods, particularly UNet, are considered as state-of-the-art for many medical imaging tasks. However, despite remarkable progress on segmenting the normal lung, performance of the UNet is unsatisfactory on challenging chest X-ray (CXR) images. This could be due to mainly two limiting factors: (1) skip connections that merge feature maps of similar size from encoding and decoding paths, and (2) loss of spatial information due to repetitive down-sampling operations. To overcome these problems, in this study, we propose a DNN-based new architecture that replaces the skip connections with a bidirectional convolutional-LSTM (BC-LSTM) module that allows exchange of more information between encoder and decoder paths and also capture spatiotemporal information. For further improvement, we add a multiple kernel pooling (MKP) block at the lowest level of UNet to encode more spatial information by different sized pooling operations. To evaluate the performance of our method, we use CXR images with different pulmonary diseases such as tuberculosis, pneumoconiosis, and Covid-19 from four public datasets as well as a private dataset and compare its performance with a standard UNet model. Results suggest that the proposed framework outperforms the UNet for all five datasets on lung segmentation, in terms of two evaluation metrics, namely Dice Coefficient (DC) and Jaccard Index (JI).
Md. Sharitul Alam, Dadong Wang, Arcot Sowmya
ICTAI3
2021 Learning Visual Features by Colorization for Slide-Consistent Survival Prediction from Whole Slide Images
Lei Fan 0007, Arcot Sowmya, Erik Meijering, Yang Song 0001
MICCAI (8)2
2021 Context-Enhanced Representation Learning for Single Image Deraining
Guoqing Wang 0001, Changming Sun, Arcot Sowmya
Int. J. Comput. Vis.3
2021 Attentive Feature Refinement Network for Single Rainy Image Restoration
abstract
Despite the fact that great progress has been made on single image deraining tasks, it is still challenging for existing models to produce satisfactory results directly, and it often requires a single or multiple refinement stages to gradually improve the quality. However, in this paper, we demonstrate that existing image-level refinement with a stage-independent learning design is problematic with the side effect of over/under-deraining. To resolve this issue, we for the first time propose the mechanism of learning to carry out refinement on the unsatisfactory features, and propose a novel attentive feature refinement (AFR) module. Specifically, AFR is designed as a two-branched network for simultaneous rain-distribution-aware attention map learning and attention guided hierarchy-preserving feature refinement. Guided by task-specific attention, coarse features are progressively refined to better model the diversified rainy effects. By using a separable convolution as the basic component, our AFR module introduces little computation overhead and can be readily integrated into most rainy-to-clean image translation networks for achieving better deraining results. By incorporating a series of AFR modules into a general encoder-decoder network, AFR-Net is constructed for deraining and it achieves new state-of-the-art results on both synthetic and real images. Furthermore, by using AFR-Net as a teacher model, we explore the use of knowledge distillation to successfully learn a student model that is also able to achieve state-of-the-art results but with a much faster inference speed (i.e., it only takes 0.08 second to process a 512×512 rainy image). Code and pre-trained models are available at 〈 https://github.com/RobinCSIRO/AFR-Net 〉 .
Guoqing Wang 0001, Changming Sun, Arcot Sowmya
IEEE Trans. Image Process.3
2020 Self-Supervised Remote Sensing Image Retrieval
abstract
Current remote sensing platforms generate a vast amount of imagery but the best current methods to index and retrieve that data require expensive and difficult to procure labels. In this paper, we aim to address this problem by presenting a performant content based image retrieval (CBIR) system that is capable of indexing and retrieval using only unlabelled data. We investigate the use of self-supervised learning, a method for end-to-end learning of visual features from large datasets. In particular, we investigate the performance of four state-of-the-art self-supervised learning methods: variational autoencoders, bidirectional GANs, colourisation networks and DeepCluster, and evaluate the quality of the representations learned on remote sensing CBIR problems. Experiments on two very high resolution datasets show that the best of these methods, DeepCluster, is able to achieve near parity with supervised transfer learning despite not using any label information.
Kane Walter, Matthew J. Gibson, Arcot Sowmya
IGARSS3
2020 Corner detection based on shearlet transform and multi-directional structure tensor
Changming Sun, Arcot Sowmya
Pattern Recognit.4
2020 Multi-Weighted Co-Occurrence Descriptor Encoding for Vein Recognition
abstract
Despite being highly secure, vein recognition suffers from the high inter-class similarity and intra-class variation resulting from the uncontrolled image capture, making the design of discriminative and robust representation very important. The recent success of convolutional neural network (CNN) for various image understanding tasks makes it a promising method for feature extraction. However, limited variability in small-scale datasets leads to systems derived from the direct training or fine-tuning not transferable and unreliable for practical biometric applications. This motivates the design of a multi-weighted co-occurrence descriptor encoding (MWCDE) model for vein recognition. Instead of directly conducting a feed-forward operation with a pre-trained CNN for obtaining the semantic features from the fully connected layers, co-occurrence features among convolutional filters are modeled first in MWCDE by a simple convolution between an indicator filter in a higher layer with a to-be-reweighted filter in a lower layer, and a redundancy-driven indicator filter selection algorithm is designed for filtering out some ambiguous representations. Second, another hard feature weighting strategy with a binary masking scheme is proposed for discarding noisy background and feature redundancy. The selected high-order descriptors are then embedded and aggregated into the compact feature vectors with a saliency driven spatial weighted Fisher vector algorithm, followed by the introduction of a generalized support vector machine for recognition. Extensive experiments with three benchmark vein datasets demonstrate that the proposed framework can achieve state-of-the-art results, and an additional experiment with the PolyU multispectral palmprint database illustrates its generalization ability. Code is available at (https://github.com/RobinCSIRO/MWCDE-for-Vein-Recognition).
Guoqing Wang 0001, Changming Sun, Arcot Sowmya
IEEE Trans. Inf. Forensics Secur.3
2020 Learning a Compact Vein Discrimination Model With GANerated Samples
abstract
Despite the great success achieved by convolutional neural networks (CNNs) in various image understanding tasks, it is still difficult for CNNs to be applied to vein recognition tasks due to the problems of insufficient training datasets, intra-class variations, and inter-class similarities. Besides, due to the essential requirement on the storage of millions of parameters for CNN, it is challenging to use a CNN for designing a vein-based embedded person identification system. In this paper, these two problems are addressed by learning a discriminative and compact vein recognition model. For the first problem, a hierarchical generative adversarial network (HGAN) consisting of a constrained CNN and a CycleGAN is proposed for data augmentation. Two similarity losses are defined for estimating the self-similarity and inter-class dissimilarity, and a CycleGAN model is properly trained with these two losses for better task-specific training sample generation. After obtaining a baseline vein recognition model fine-tuned on the augmented datasets, the existence of parameter redundancy in the over-parameterized network motivates the proposal of model compression by way of filter pruning and low rank approximation, thus making the compressed model more suitable for deployment on embedded systems. Through the vein recognition experiments with two different datasets and an additional palmprint recognition experiment, the proposed algorithms are shown to yield a highly compact model while keeping the accuracy acceptable for application.
Guoqing Wang 0001, Changming Sun, Arcot Sowmya
IEEE Trans. Inf. Forensics Secur.3
2020 Cascaded Attention Guidance Network for Single Rainy Image Restoration
abstract
Restoring a rainy image with raindrops or rainstreaks of varying scales, directions, and densities is an extremely challenging task. Recent approaches attempt to leverage the rain distribution (e.g., location) as prior to generate satisfactory results. However, concatenation of a single distribution map with the rainy image or with intermediate feature maps is too simplistic to fully exploit the advantages of such priors. To further explore this valuable information, an advanced cascaded attention guidance network, dubbed as CAG-Net, is formulated and designed as a three-stage model. In the first stage, a multitask learning network is constructed for producing the attention map and coarse de-raining results simultaneously. Subsequently, the coarse results and the rain distribution map are concatenated and fed to the second stage for results refinement. In this stage, the attention map generation network from the first stage is used to formulate a novel semantic consistency loss for better detail recovery. In the third stage, a novel pyramidal "whereand- how" learning mechanism is formulated. At each pyramid level, a two-branch network is designed to take the features from previous stages as inputs to generate better attention-guidance features and de-raining features, which are then combined via a gating scheme to produce the final de-raining results. Moreover, the uncertainty maps are also generated in this stage for more accurate pixel-wise loss calculation. Extensive experiments are carried out for removing raindrops or rainstreaks from both synthetic and real rainy images, and CAG-Net is demonstrated to produce significantly better results than state-of-the-art models. Code will be publicly available after paper acceptance.
Guoqing Wang 0001, Changming Sun, Arcot Sowmya
IEEE Trans. Image Process.3
2019 ERL-Net: Entangled Representation Learning for Single Image De-Raining
abstract
Despite the significant progress achieved in image de-raining by training an encoder-decoder network within the image-to-image translation formulation, blurry results with missing details indicate the deficiency of the existing models. By interpreting the de-raining encoder-decoder network as a conditional generator, within which the decoder acts as a generator conditioned on the embedding learned by the encoder, the unsatisfactory output can be attributed to the low-quality embedding learned by the encoder. In this paper, we hypothesize that there exists an inherent mapping between the low-quality embedding to a latent optimal one, with which the generator (decoder) can produce much better results. To improve the de-raining results significantly over existing models, we propose to learn this mapping by formulating a residual learning branch, that is capable of adaptively adding residuals to the original low-quality embedding in a representation entanglement manner. Using an embedding learned this way, the decoder is able to generate much more satisfactory de-raining results with better detail recovery and rain artefacts removal, providing new state-of-the-art results on four benchmark datasets with considerable improvement (i.e., on the challenging Rain100H data, an improvement of 4.19dB on PSNR and 5% on SSIM is obtained). The entanglement can be easily adopted into any encoder-decoder based image restoration networks. Besides, we propose a series of evaluation metrics to investigate the specific contribution of the proposed entangled representation learning mechanism. Codes are available at 〈https://github.com/RobinCSIRO/ERL-Net-for-Single-Image-Deraining〉.
Guoqing Wang 0001, Changming Sun, Arcot Sowmya
ICCV3
2019 Exploring Latent Structure Similarity for Bayesian Nonparameteric Model with Mixture of NHPP Sequence
Yongzhe Chang, Zhidong Li, Ling Luo 0002, Simon Luo, Arcot Sowmya, Yang Wang 0002, Fang Chen 0001
ICONIP (2)5
2019 Recovering DTW Distance Between Noise Superposed NHPP
Yongzhe Chang, Zhidong Li, Bang Zhang, Ling Luo 0002, Arcot Sowmya, Yang Wang 0002, Fang Chen 0001
PAKDD (2)5
2019 Hawkes Process with Stochastic Triggering Kernel
Feng Zhou 0011, Yixuan Zhang 0006, Zhidong Li, Xuhui Fan 0001, Yang Wang 0002, Arcot Sowmya, Fang Chen 0001
PAKDD (1)6
2018 A Refined MISD Algorithm Based on Gaussian Process Regression
Feng Zhou 0011, Zhidong Li, Xuhui Fan 0001, Yang Wang 0002, Arcot Sowmya, Fang Chen 0001
PAKDD (2)5
2017 A performance acceleration algorithm of spectral unmixing via subset selection
Jing Ke, Yi Guo 0001, Arcot Sowmya, Tomasz Bednarz
ESANN3
2017 Prediction of Protein X-ray Crystallisation Trial Image Time-courses
B. M. Thamali Lekamge, Arcot Sowmya, Janet Newman
ICPRAM2
2017 Deep Learning Approach for Classification of Mild Cognitive Impairment Subtypes
Upul Senanayake, Arcot Sowmya, Laughlin Dawes, Nicole A. Kochan, Wei Wen 0001, Perminder S. Sachdev
ICPRAM2
2016 Classification of Mild Cognitive Impairment Subtypes using Neuropsychological Data
abstract
While the research on Alzheimer’s disease (AD) is progressing, timely intervention before an individual becomes demented is often emphasized. Mild Cognitive Impairment (MCI), which is thought of as a prodromal syndrome to AD, may be useful in this context as potential interventions can be applied to individuals at increased risk of developing dementia. The current study attempts to address this problem using a selection of machine learning algorithms to discriminate between cognitively normal individuals and MCI individuals among a cohort of community dwelling individuals aged 70-90 years based on neuropsychological test performance. The overall best algorithm in our experiments was AdaBoost with decision trees while random forests was consistently stable. Ten-fold cross validation was used with ten repetitions to reduce variability and assess generalizing capabilities of the trained models. The results presented are consistently of the same calibre or better than the limited number of similar studies reported in the literature.
Upul Senanayake, Arcot Sowmya, Laughlin Dawes, Nicole A. Kochan, Wei Wen 0001, Perminder S. Sachdev
ICPRAM2
2016 Optimized GPU implementation for dynamic programming in image data processing
abstract
It is a trend now that computing power through parallelism is provided by multi-core systems or heterogeneous architectures for High Performance Computing (HPC) and scientific computing. Although many algorithms have been proposed and implemented using sequential computing, alternative parallel solutions provide more suitable and high performance solutions to the same problems. In this paper, three parallelization strategies are proposed and implemented for a dynamic programming based cloud smoothing application, using both shared memory and non-shared memory approaches. The experiments are performed on NVIDIA GeForce GT750m and Tesla K20m, two GPU accelerators of Kepler architecture. Detailed performance analysis is presented on partition granularity at block and thread levels, memory access efficiency and computational complexity. The evaluations described show high approximation of results with high efficiency in the parallel implementations, and these strategies can be adopted in similar data analysis and processing applications.
Jing Ke, Tomasz Bednarz, Arcot Sowmya
IPCCC3
2016 Soft Hough Forest-ERTs: Generalized Hough Transform based object detection from soft-labelled training data
Yang Wang 0002, Arcot Sowmya, Fang Chen 0001
Pattern Recognit.3
2015 An Underwater Color Image Quality Evaluation Metric
abstract
Quality evaluation of underwater images is a key goal of underwater video image retrieval and intelligent processing. To date, no metric has been proposed for underwater color image quality evaluation (UCIQE). The special absorption and scattering characteristics of the water medium do not allow direct application of natural color image quality metrics especially to different underwater environments. In this paper, subjective testing for underwater image quality has been organized. The statistical distribution of the underwater image pixels in the CIELab color space related to subjective evaluation indicates the sharpness and colorful factors correlate well with subjective image quality perception. Based on these, a new UCIQE metric, which is a linear combination of chroma, saturation, and contrast, is proposed to quantify the non-uniform color cast, blurring, and low-contrast that characterize underwater engineering and monitoring images. Experiments are conducted to illustrate the performance of the proposed UCIQE metric and its capability to measure the underwater image enhancement results. They show that the proposed metric has comparable performance to the leading natural color image quality metrics and the underwater grayscale image quality metrics available in the literature, and can predict with higher accuracy the relative amount of degradation with similar image content in underwater environments. Importantly, UCIQE is a simple and fast solution for real-time underwater video processing. The effectiveness of the presented measure is also demonstrated by subjective evaluation. The results show better correlation between the UCIQE and the subjective mean opinion score.
Arcot Sowmya
IEEE Trans. Image Process.2
2014 Human Action Description Based on Temporal Pyramid Histograms
abstract
In this paper, we present an approach to action description based on temporal pyramid histograms. Bag of features is a widely used action recognition framework based on local features, for example spatio-temporal feature points. Although it outperforms other approaches on several public datasets, sequencing information is ignored. Instead of only calculating the occurrence of code words, we also encode their temporal layout in this work. The proposed temporal pyramid histograms descriptor is a set of histogram atoms generated from the original video clip and its subsequences. To classify actions based on the temporal pyramid histograms descriptor, we design a kernel function to enhance the description ability of the descriptor. In the kernel function, weights are assigned to the histogram atoms according to the corresponding sequence lengths. We test the descriptor using nearest neighbour for classification. Experimental results show that, in comparison to the state-of-the-art, our description approach improves action recognition accuracy.
Arcot Sowmya
ICPRAM2
2014 New Image Quality Evaluation Metric for Underwater Video
abstract
This work presents a new vectorial underwater image quality metric, dubbed the CQ, that integrates the power spectrum. Unlike existing objective underwater image quality metrics, the proposed metric consists of a discriminator C based on the slope of the log-contrast power spectrum that is able to distinguish between marine habitats when a large number of images of different environments are to be processed, and a patch-based metric Q to predict the objective quality of underwater images. Experimental results illustrate that the proposed CQ metric is able to recognize underwater images with similar sharpness and correlates better with enhancement results compared to other methods, and also meet real-time requirements.
Arcot Sowmya
IEEE Signal Process. Lett.2
2013 Efficient graph cuts based extraction of vertebral column and ribs in lung MDCT images
abstract
A fully automatic novel algorithm based on graph cuts is presented for accurate and fast segmentation and isolation of human vertebral column and ribs in multi detector computed tomography (MDCT) images. The segmentation is followed by a two-step isolation method to remove mis-segmented parts such as the sternum, clavicle and scapula. The proposed algorithm was tested on 18 patient datasets, with 5 slices from each dataset compared to the reference delineation provided by a radiologist. The experiments were performed on both 2-D (with 4 and 8 neighbours) and 3-D (with 6 and 26 neighbours) graphs with wide range of parameter values. Based on our evaluation, the 2-D, 4 neighbours graph shows high performance (Dice similarity coefficient ≈ 92.5%) with low running time (57.86 s for a 346 slice dataset) and is recommended for accurate and fast segmentation of the vertebral column and ribs.
Banafsheh Pazokifard, Arcot Sowmya
ICIP2
2013 Robust human appearance matching across multi-cameras
abstract
In this paper, we present a novel solution to the problem of human appearance matching across multiple cameras. Humans are represented by a set of feature points sampled from upper bodies. The problem of appearance matching across multiple cameras is formulated as finding corresponding points in two upper bodies from different views based on dissimilarity of region signatures as well as geometric constraints between feature points. For dissimilarity of region signatures, we first use k-means clustering to describe the region around the feature point, then estimate the dissimilarity between different regions under integer optimization framework. For geometric constraints, we get the spatial information of feature points based on a scale and rotation invariant constraint method. Lastly, agglomerative clustering algorithm is used to find the correct cluster of candidate pairs. Our method is robust to both outliers and deformation, and the experimental results show promising matching results on multiple cameras.
Beihua Zhang, Xiongcai Cai, Arcot Sowmya
ICIP3
2013 A weakly supervised approach for object detection based on Soft-Label Boosting
abstract
Object detection is an important and challenging problem in the field of computer vision. Classical object detection approaches such as background subtraction and saliency detection do not require manual collection of training samples, but can be easily affected by noise factors, such as luminance changes and cluttered background. On the other hand, supervised learning based approaches such as Boosting and SVM usually have robust performance, but require substantial human effort to collect and label training samples. This study aims to combine the comparative advantages of both kinds of approaches, and its contributions are two-fold: (i) a weakly supervised approach for object detection, which does not require manual collection and labelling of training samples; (ii) an extension of Boosting algorithm denoted as Soft-Label Boosting, which is able to employ training samples with soft (probabilistic) labels instead of hard (binary) labels. Experimental results show that the proposed weakly supervised approach outperforms the state-of-the-art, and even achieves comparable performance to supervised approaches.
Yang Wang 0002, Fang Chen 0001, Arcot Sowmya
WACV4
2013 Analyzing an embedded sensor with timed automata in uppaal
abstract
An infrared sensor is modeled and analyzed in Uppaal. The sensor typifies the sort of component that engineers regularly integrate into larger systems by writing interface hardware and software. In all, three main models are developed. In the first model, the timing diagram of the sensor is interpreted and modeled as a timed safety automaton. This model serves as a specification for the complete system. A second model that emphasizes the separate roles of driver and sensor is then developed. It is validated against the timing diagram model using an existing construction that permits the verification of timed trace inclusion, for certain models, by reachability analysis (i.e., model checking). A transmission correctness property is also stated by means of an auxiliary automaton and shown to be satisfied by the model. A third model is created from an assembly language driver program, using a direct translation from the instruction set of a processor with simple timing behavior. This model is validated against the driver component of the second timing diagram model using the timed trace inclusion validation technique. The approach and its limitations offer insight into the nature and challenges of programming in real time.
Timothy Bourke, Arcot Sowmya
ACM Trans. Embed. Comput. Syst.2
2012 Sparse Dictionary Reconstruction for Textile Defect Detection
abstract
Inspired by the image de-noising techniques using learned dictionaries and sparse representation, we present a fabric defect detection scheme via sparse dictionary reconstruction. Fabric defects can be regarded as local anomalies against the relatively homogeneous texture background. Following from the flexibility of sparse representation, normal fabric samples can be efficiently represented using a linear combination of a few elements of a learned dictionary. When modeling new samples with a learned dictionary, tuned to the input data containing normal fabric structural features, abnormal or defective samples are likely to have larger dissimilarity than normal samples. We evaluate the proposed methods using ten different fabric types. Experimental results show that our method has many advantages in defect detection, especially in adapting variation of fabric textures.
Dimitri Semenovich, Arcot Sowmya
ICMLA (1)3
2012 Predicting onsets of genocide with sparse additive models
Dimitri Semenovich, Arcot Sowmya, Benjamin E. Goldsmith
ICPR2
2012 Perceptual Evaluation of Automatic 2.5D Cartoon Modelling
Fengqi An, Xiongcai Cai, Arcot Sowmya
PKAW3
2012 Robot Localisation Using Natural Landmarks
Peter Anderson 0001, Yongki Yusmanthia, Bernhard Hengst, Arcot Sowmya
RoboCup4
2011 Multiscale sparse representation of high-resolution computed tomography (HRCT) lung images for diffuse lung disease classification
abstract
A multiscale sparse representation scheme based on wavelet and contourlet transforms is employed to describe four patterns of diffuse lung disease patterns: normal, emphysema, ground glass opacity (GGO) and honey-combing based on HRCT lung images. First, using sparse representation, four discriminative dictionaries are trained for the four patterns respectively. After that, in the classification phase, a patch or ROI is assigned to the pattern with minimum resconstruction error. The method is tested on a collection of 89 slices from 38 patients, each slice of size 512 × 512, 16 bits/pixel in DICOM format. The dataset contains 73,000 ROIs of those slices marked by experienced radiologists. We employ this technique with 2-scale wavelet and [2 3] contourlet transform for diffuse lung disease classification. The technique presented here has the overall sensitivity of 91.05% and specificity 97.01%.
Kiet T. Vo, Arcot Sowmya
ICIP2
2011 Feature fusion for vehicle detection and tracking with low-angle cameras
abstract
In this paper, we address the problem of vehicle detection and tracking with low-angle cameras by combining windshield detection and feature points clustering, effectively fusing several primitive image features such as color, edge and interest point. By exploring various heterogenous features and multiple vehicle models, we achieve at least two improvements over the existing methods: higher detection accuracy and the ability to distinguish different vehicle types. Our experiments on real-world traffic video sequences demonstrate the benefits of feature fusion and the improved performance.
Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Zhidong Li, Bang Zhang, Jie Xu 0008
WACV3
2011 Distributed, multi-sensor tracking of multiple participants within immersive environments using a 2-cohort camera setup
Anuraag Sridhar, Arcot Sowmya
Mach. Vis. Appl.2
2010 Geometry Aware Local Kernels for Object Recognition
Dimitri Semenovich, Arcot Sowmya
ACCV (1)2
2010 Spatial-Temporal Affinity Propagation for Feature Clustering with Application to Traffic Video Analysis
Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Jie Xu 0008, Zhidong Li, Bang Zhang
ACCV (2)3
2010 Affinity Propagation Feature Clustering with Application to Vehicle Detection and Tracking in Road Traffic Surveillance
abstract
In this paper, we investigate the applicability of the newly proposed data clustering method, affinity propagation, in feature points clustering and the task of vehicle detection and tracking in road traffic surveillance. We propose a model-based temporal association scheme and novel preprocessing and postprocessing operations which together with affinity propagation make a quite successful method for the given task. Our experiments demonstrate the effectiveness and efficiency of our method and its superiority over the state-of-the-art algorithm.
Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Bang Zhang, Jie Xu 0008, Zhidong Li
AVSS3
2010 Vehicle detection and tracking with low-angle cameras
abstract
Vision-based vehicle detection is a critical task for traffic monitoring in modern Intelligent Traffic Systems (ITS). Due to the low-angle nature of most traffic surveillance cameras installed in the real world, vehicle detection in such case has to deal with one fundamental challenge — occlusion, which renders most traditional vehicle detection methods ineffective. In this paper, instead of detecting the vehicle as a whole, we propose a vehicle detection algorithm based on windshield model matching. By detecting windshield directly, the algorithm achieves robustness to occlusion. Together with camera calibration and vehicle tracking, the system is able to provide reliable traffic state estimation. Experiments on real traffic videos demonstrate the better performance of our system compared to the state-of-the-art algorithm.
Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Zhidong Li
ICIP3
2010 Scale-Space Representation of Lung HRCT Images for Diffuse Lung Disease Classification
Kiet T. Vo, Arcot Sowmya
ICISP2
2010 On-line, Incremental Learning for Real-Time Vision Based Movement Recognition
abstract
In this paper we tackle the problem of recognising movement classes in real-time surveillance video. We use a popular public dataset, the CAVIAR dataset, which contains ground truth labeling of people and their activities within a shopping centre environment. The task of movement classification is often performed using simple heuristic rules, and performance can suffer when an increased number of rules are added for the task. We provide a formal knowledge maintenance technique, known as Ripple Down Rules, to provide an elegant method of representing and updating the rules. Ripple Down Rules are an on-line, incremental learning strategy, and are highly suitable for this task due to their ability to incorporate new knowledge while maintaining past knowledge.
Anuraag Sridhar, Arcot Sowmya, Paul Compton
ICMLA2
2010 Tensor Power Method for Efficient MAP Inference in Higher-order MRFs
abstract
We present a new efficient algorithm for maximizing energy functions with higher order potentials suitable for MAP inference in discrete MRFs. Initially we relax integer constraints on the problem and obtain potential label assignments using higher-order (tensor) power method. Then we utilise an ascent procedure similar to the classic ICM algorithm to converge to a solution meeting the original integer constraints.
Dimitri Semenovich, Arcot Sowmya
ICPR2
2010 Incremental System Engineering Using Process Networks
Avishkar Misra, Arcot Sowmya, Paul Compton
PKAW2
2010 ACS: Automatic Converter Synthesis for SoC Bus Protocols
Karin Avnit, Arcot Sowmya, Jorgen Peddersen
TACAS2
2009 Directional Multi-scale Modeling of High-Resolution Computed Tomography (HRCT) Lung Images for Diffuse Lung Disease Classification
Kiet T. Vo, Arcot Sowmya
CAIP2
2009 A formal approach to design space exploration of protocol converters
abstract
In the field of chip design, hardware module reuse is a standard solution to the increasing complexity of chip architecture and the pressures to reduce time to market. In the absence of a single module interface standard, integration of pre-designed modules often requires the use of protocol converters. For an arbitrary pair of incompatible protocols it is likely that there exist more than one possible converter. However, existing approaches to automatic synthesis of protocol converters either produce a single suggested converter or provide a general nondeterministic solution, out of which a designer is required to extract a deterministic converter. In this work we present a novel approach for design space exploration of FSM based protocol converters. We present algorithms for extraction of minimal converters for a given pair of incompatible protocols. We demonstrate the process through a simple example, and report on results of experiments with converters for commercial protocols AMBA ASB, APB and the open core protocol (OCP). The experiments show a reduction in the number of states in the converter of as much as 62% (with an average reduction of 42%) and a reduction in the number of transitions of as much as 85% (with an average reduction of 61%), demonstrating the benefits of design space exploration.
Karin Avnit, Arcot Sowmya
DATE2
2009 A spectral method for context based disambiguation of image annotations
abstract
In this work we employ contextual information to improve the quality of image labellings provided by an existing automatic image annotation algorithm in a weakly supervised setting, where each training image is labelled but it is not known which part of the image its labels are referring to. We recast the problem into that of constructing a graph which encodes pairwise consistency of candidate annotations and observe that mutually consistent labels will form a compact cluster in this graph. We recover the clusters using a spectral theory based technique. The results are demonstrated on the Corel5k dataset. With improvements in the range of 25%-55% the performance in some cases approaches the state of the art despite using a very simple base algorithm.
Dimitri Semenovich, Arcot Sowmya
ICIP2
2009 Diffuse lung disease classification in HRCT lung images using generalized Gaussian density modeling of wavelets coefficients
abstract
The generalized Gaussian density model for wavelet subbands has been applied widely in texture image retrieval. In this paper, we employ wavelet-based texture extraction that is based on accurate modeling of the distribution of wavelet coefficients using generalized Gaussian density to classify four diffuse lung disease patterns: normal, emphysema, ground glass opacity and honey-combing. The evaluated classifiers are K-nearest neighbor (K-NN) and support vector machine (SVM). A collection of 124 slices from 45 patients has been investigated, each slice of size 512×512, 12bit/pixel in DICOM format. The dataset contains 6000 ROIs of those slices marked by experienced radiologists. We employ this technique at different wavelet transform scales and compare results to other wavelet-based classification techniques for diffuse lung disease classification. The technique presented here has the best overall accuracy of 92.25% for the multi-class case with 3-level wavelet transform and SVM classifier.
Kiet T. Vo, Arcot Sowmya
ICIP2
2009 Feature clustering for vehicle detection and tracking in road traffic surveillance
abstract
In this paper, we formulate the feature clustering problem for vehicle detection and tracking as a general MAP problem and solve it using MCMC. The proposed approach exhibits two advantages over existing methods: general Bayesian model can handle arbitrary objective functions and MCMC guarantees global optimal solution. Our algorithm is validated on real-world traffic video sequences, and is shown to outperform the state-of-the-art approach.
Jun Yang 0033, Yang Wang 0002, Getian Ye, Arcot Sowmya, Bang Zhang, Jie Xu 0008
ICIP4
2009 SparseSPOT: using a priori 3-D tracking for real-time multi-person voxel reconstruction
abstract
Voxel reconstruction has received increasing interest in recent times, driven by the need for efficient reconstructions of real world scenes from video images. The voxel model has proven useful for activity recognition and motion capture technologies. However most current voxel reconstruction algorithms operate on a fairly small 3-D real world volume and only allow for a single person to be reconstructed. In this paper we present SparseSPOT, an extension of the SPOT voxel reconstruction algorithm, that enables real-time reconstruction of multiple humans within a large environment. We compare SparseSPOT to SPOT and show (by extensive experimental evaluation) that the former achieves superior real time performance.
Anuraag Sridhar, Arcot Sowmya
VRST2
2009 Multi-level classification of emphysema in HRCT lung images
Mithun Nagendra Prasad, Arcot Sowmya, Peter Wilson 0001
Pattern Anal. Appl.2
2009 Provably correct on-chip communication: A formal approach to automatic protocol converter synthesis
abstract
Hardware module reuse is a standard solution to the problems of increasing complexity of chip architectures and pressure to reduce time to market. In the absence of a single module interface standard, predesigned modules for “plug-and-play” usually require a converter between incompatible interface protocols. Current approaches to automatic synthesis of protocol converters mostly lack formal foundations and either employ abstractions far removed from the HDL implementation level or grossly simplify the structure of the protocols considered. This work presents a state-machine-based formalism for modeling bus-based communication protocols and a notion of protocol compatibility and of correct conversion between incompatible protocols. This formalism is used to derive algorithms for checking protocol compatibility and for provably correct, automatic converter synthesis. Experiments with automatic converter synthesis between different configurations of widely used commercial bus protocols, such as AMBA AHB, ASB APB, and the Open Core Protocol (OCP) are discussed. The work here is unique in its combination of a completely formal approach and the use of a low abstraction level that enables precise modeling of protocol characteristics that is also close to HDL.
Karin Avnit, Vijay D'Silva, Arcot Sowmya, Sri Parameswaran
ACM Trans. Design Autom. Electr. Syst.3
2008 A Formal Approach To The Protocol Converter Problem
abstract
In the absence of a single module interface standard, integration of pre-designed modules in System-on-Chip design often requires the use of protocol converters. Existing approaches to automatic synthesis of protocol converters mostly lack formal foundations and either employ abstractions that ignore crucial low level behaviors, or grossly simplify the structure of the protocols considered. We present a state-machine based formal model for bus based communication protocols, and precisely define protocol compatibility, and correct protocol conversion. Our model is expressive enough to capture features of commercial protocols such as bursts, pipelined transfers, wait state insertion, and data persistence, in cycle accurate detail. We show that the most general, correct converter for a pair of protocols, can be described as the greatest fixed point of a function for updating buffer states. This characterization yields a natural algorithm for automatic synthesis of a provably correct converter by iterative computation of the fixed point. We report our experience with automatic converter synthesis between widely used commercial bus protocols, such as AMBA AHB, ASB, APB, and OCP, considering features which are beyond the scope of current techniques.
Karin Avnit, Vijay D'Silva, Arcot Sowmya, Sri Parameswaran
DATE3
2008 Situated Cognition in the Semantic Web Era
Paul Compton, Byeong Ho Kang 0001, Rodrigo Martínez-Béjar, Mamatha Rudrapatna, Arcot Sowmya
EKAW5
2008 Automatically transforming and relating Uppaal models of embedded systems
abstract
Relations between models are important for effective automatic validation, for comparing implementations with specifications, and for increased understanding of embedded systems designs. Timed automata may be used to model a system at multiple levels of abstraction, and timed trace inclusion is one way to relate the models.
Timothy Bourke, Arcot Sowmya
EMSOFT2
2008 Multi-level Classification of Emphysema in HRCT Lung Images Using Delegated Classifiers
Mithun Nagendra Prasad, Arcot Sowmya
MICCAI (1)2
2008 Designing Relevant Features for Continuous Data Sets Using ICA
abstract
Isolating relevant information and reducing the dimensionality of the original data set are key areas of interest in pattern recognition and machine learning. In this paper, a novel approach to reducing dimensionality of the feature space by employing independent component analysis (ICA) is introduced. While ICA is primarily a feature extraction technique, it is used here as a feature selection/construction technique in a generic way. The new technique, called feature selection based on independent component analysis (FS_ICA), efficiently builds a reduced set of features without loss in accuracy and also has a fast incremental version. When used as a first step in supervised learning, FS_ICA outperforms comparable methods in efficiency without loss of classification accuracy. For large data sets as in medical image segmentation of high-resolution computer tomography images, FS_ICA reduces dimensionality of the data set substantially and results in efficient and accurate classification.
Mithun Nagendra Prasad, Arcot Sowmya, Inge Koch
Int. J. Comput. Intell. Appl.2
2008 Segmentation of Lung Patterns in High-Resolution Computed Tomography Images of the Lung
abstract
An automated method is presented for segmentation of two-dimensional HRCT images of the lung into regions of four lung patterns: normal, emphysema, honeycombing, and ground-glass opacity (GGO). Segmentation was implemented in two stages. At the first stage, pixel-wise classification of the lung area was performed using local textural features extracted by the wavelet transform. At the second stage, classification results were refined by application of knowledge-based rules. Performance of the method was compared on two sets of HRCT images: one included HRCT images with characteristic examples of lung patterns and the other consisted of unselected HRCT images that represented a model of routine operations at a general radiology practice. On the first set of images, sensitivity of the method ranged from 0.92 to 0.99, and specificity ranged from 0.96 to 0.99. On the second set of images, sensitivity and specificity were, respectively, 0.49 and 0.95 for emphysema, 0.87 and 0.55 for normal, 0.34 and 0.99 for honeycombing, and 0.57 and 0.94 for GGO. The two-stage approach allowed for simple and effective application of high-level knowledge about appearance of lung patterns on HRCT images and did not require feature and region of interest size selection for the first stage of pixel-wise lung pattern classification.
Alena Shamsheyeva, Arcot Sowmya, Peter Wilson 0001
Int. J. Comput. Intell. Appl.2
2007 Level Learning Set: A Novel Classifier Based on Active Contour Models
Xiongcai Cai, Arcot Sowmya
ECML2
2007 An Adaptable Formal Model for Web Services Protocols
abstract
Agents require standard and reliable protocols to interact with service providers in order to provide high quality customer service over the Web. Many useful Web services protocols are coming on the market, but are often ambiguously specified by protocol designers and not fully verified. This can lead to interoperability problems among implementations of the same protocol as well as high software maintenance costs. We have recently proposed a formal hierarchical automata-based framework that aims to address these issues. In this paper, we extend our framework to overcome the two identified limitations, re-usability and adaptability, and describe how they are useful for conformance checking and for our observer-based technique for property specification. We also apply the extended framework on the WS-business activity's atomicoutcome protocol suite and discuss our experience using the model.
Pemadeep Ramsokul, Arcot Sowmya
ICIW2
2007 A Test Bed for Web Services Protocols
abstract
Transactions across composed Web services (WS) are usually non-trivial and require the use of some pre-agreed or standard protocols. Proper specification and implementation of these protocols are critical for the correct execution and termination of transactions; incomplete or ambiguous specifications can give rise to interoperability problems. We have recently proposed a framework for specifying and verifying WS protocols. In this paper, we propose a test bed based on this framework for finding bugs in implementations of WS protocols especially when the number of entities that can participate is not always the same, which is typical of transaction protocols. We also illustrate the utility of our test bed using the implementation of an actual WS protocol.
Pemadeep Ramsokul, Arcot Sowmya
ICIW2
2006 Learning Parameter Tuning for Object Extraction
Xiongcai Cai, Arcot Sowmya, John C. Trinder
ACCV (1)2
2006 A timing model for synchronous language implementations in simulink
abstract
We describe a simple scheme for mapping synchronous language models, in the form of Boolean Mealy Machines, into timed automata. The mapping captures certain idealized implementation details that are ignored, or assumed away, by the synchronous paradigm. In this regard, the scheme may be compared with other approaches such as the AASAP semantics. However, our model addresses input latching and reaction triggering differently. Additionally, the focus is not on model-checking but rather on creating a semantic model for simulating synchronous controllers within Simulink.The model considers both sample-driven and event-driven execution paradigms, and clarifies their similarities and differences. It provides a means of analyzing the timing behavior of small-scale embedded controllers.The integration of the timed automata models into Simulink is described and related work is discussed.
Timothy Bourke, Arcot Sowmya
EMSOFT2
2006 Automatic Detection of Tram Tracks on HRCT Images
abstract
On high resolution computed tomography (HRCT) images, dilated airways appear as two parallel lines that resemble tram tracks, when they lie in the plane of scan. Tram tracks, when visible, are characteristic of bronchiectasis, a disease caused by the irreversible dilatation of the bronchial tree. Detection of such patterns provides valuable diagnostic information. In this work, semi-supervised learning together with image analysis techniques have been used to detect tram tracks on HRCT images. The approach was tested on 1091 HRCT images belonging to 54 patients, and the results visually validated by radiologists. Sensitivity and specificity of 80% and 91% respectively were achieved.
Mamatha Rudrapatna, Prinith Amaratunga, Mithun Nagendra Prasad, Arcot Sowmya, Peter Wilson 0001
ICIP4
2006 Active Contour with Neural Networks-Based Information Fusion Kernel
Xiongcai Cai, Arcot Sowmya
ICONIP (2)2
2006 A Sniffer Based Approach to WS Protocols Conformance Checking
abstract
To reduce interoperability problems arising from ambiguous or incomplete Web services protocol specifications, we have recently introduced a formal framework, which allows modelling and automatic verification of such protocols. However, interoperability problems can still occur due to incorrect implementations. In this paper, we introduce a sniffer based approach to check the conformance of a protocol's implementation to its specification; messages of the actual implementations are captured, processed and checked against the specification's formal model. We also briefly illustrate the application of our framework using a version of the WS-AtomicTransaction protocol
Pemadeep Ramsokul, Arcot Sowmya
ISPDC2
2006 ASEHA: A Framework for Modelling and Verification ofWeb Services Protocols
abstract
Agents require standard and reliable protocols to interact with different service providers in order to provide high quality service to customers over the Web. Many useful protocols are coming into the market, but are often ambiguously specified by protocol designers and not fully verified. These can lead to interoperability problems among implementations of the same protocol and high software maintenance costs. In this paper, we propose a hierarchical automata-based framework to model the necessary features of protocols to verify their correctness. Our experience shows that the graphical models help uncover subtle scenarios and reduce, if not eliminate, ambiguities. We illustrate our formalism with a version of WS-atomic transaction protocol
Pemadeep Ramsokul, Arcot Sowmya
SEFM2
2005 Formal Models in Industry Standard Tools: an Argos Block within Simulink
abstract
Simulink is widely used within the industry for simulation and model-driven development, and reactive behaviors are often modeled using an add-on called Stateflow. Argos is one of the synchronous languages that have been proposed for the specification, validation and implementation of reactive systems. It is a rigorously defined graphical notation which, though not as powerful as Stateflow, is much less complicated. This paper describes the implementation of an Argos block for Simulink.
Timothy Bourke, Arcot Sowmya
Int. J. Softw. Eng. Knowl. Eng.2
2004 Synchronous Protocol Automata: A Framework for Modelling and Verification of SoC Communication Architectures
abstract
Plug-n-Play style Intellectual Property (IP) reuse in System on Chip (SoC) design is facilitated by the use of an on-chip bus architecture. We present a synchronous, Finite State Machine based framework for modelling communication aspects of such architectures. This formalism has been developed via interaction with designers and the industry and is intuitive and lightweight. We have developed cycle accurate methods to formally specify protocol compatibility and component composition and show how our model can be used for compatibility verification, interface synthesis and model checking with automated specification. We demonstrate the utility of our framework by modelling the AMBA bus architecture including details such as pipelined operation, burst and split transfers, the AHB-APB bridge and arbitration features.
Vijay D'Silva, S. Ramesh 0001, Arcot Sowmya
DATE3
2003 Support Vector Machines for Road Extraction from Remotely Sensed Images
Neil Yager, Arcot Sowmya
CAIP2
2002 k-time Forced Simulation: A Formal Verification Technique for IP Reuse
abstract
Automatic IP (Intellectual Property) matching is a key to reuse of IP cores. This paper presents an IP matching algorithm that can check whether a given programmable IP block can be adapted to match a given specification. When such adaptation is possible, the algorithm also generates a device driver to adapt the IP block. Though simulation, refinement and bisimulation based algorithms exist, they cannot be used to check the adaptability of an IP block, which is the essence of reuse. The IP matching algorithm is based on a formal verification technique called k-time forced simulation proposed in this paper k-time forced simulation may be used for identifying whether a given IP block (a device D) can be adapted to match a specification (a function F), given that D has a clock that is k-times faster than F. We demonstrate the applicability of the algorithm by reusing several IP blocks.
Partha S. Roop, Arcot Sowmya, S. Ramesh 0001
ICCD2
2001 A formal approach to component based development of synchronous programs
abstract
Synchronous languages may be used for specification and design of embedded systems. Assuming the availability of a library of synchronous programs, we propose a technique to enable reuse of these programs, via an algorithm for automatic matching of a design function to a program from the library. The algorithm, when successful, generates an interface which automatically adapts the program. The algorithm is based on a new simulation relation called synchronous forced simulation, which is shown to be necessary and sufficient for matching a given pair of function and program.
Partha S. Roop, Arcot Sowmya, S. Ramesh 0001
ASP-DAC2
2001 Forced simulation: A technique for automating component reuse in embedded systems
abstract
Component reuse techniques have been a recent focus of research because they are seen as the next-generation techniques to handle increasing system complexities. However, there are several unresolved issues to be addressed and prominent among them is the issue of component matching . As the number of reusable components in a component database grows, the task of manually matching a component to the user requirements becomes infeasible. Automating this matching can help in rapid system prototyping, improving quality and reducing cost. In addition, if the matching algorithm is sound, this approach can also reduce precious validation effort.In this article, we propose an algorithm for automatic matching of a design function to a device from a component database. The distinguishing feature of the algorithm is that when successful, it generates an interface that can automatically adapt the device to behave as the function. The algorithm is based on a new simulation relation called forced simulation that is shown to be a necessary and sufficient condition for component matching to be possible for a given pair of function and device. We demonstrate the application of the algorithm by reusing on some programmable components of the Intel family.
Partha S. Roop, Arcot Sowmya, S. Ramesh 0001
ACM Trans. Design Autom. Electr. Syst.2
1999 Learning Discriminatory and Descriptive Rules by an Inductive Logic Programming System
Maziar Palhang, Arcot Sowmya
ICML2
1998 Hidden time model for specification and verification of embedded systems
abstract
Embedded systems are application specific digital systems that are usually designed using a microprocessor, along with a set of programmable hardware and software components. Since these systems are real time in nature, specification of temporal constraints is a key issue. We have recently proposed the CFSMcharts language for component based specification of these systems. However this proposal had no features to specify quantitative temporal constraints that are crucial to embedded system specification. We propose a new model of time, called hidden time, for specification of temporal constraints in CFSMcharts and contrast it with existing schemes. The proposed scheme is hierarchical and hides away the quantitative temporal constraints from the top level specification. This leads to a simpler style for the specification of these constraints and simpler semantics for the top level specification. Another major contribution of the proposed scheme is that properties to be verified can be expressed in propositional temporal logic, whereas all the existing schemes have to use first order temporal logic. We also propose a new temporal logic called Hidden Propositional Temporal Logic (HPTL) as a requirement specification language. HPTL is based on the hidden time model and also supports module name qualifiers, which have applicability in a component based framework. Finally, we propose a scheme for automated verification.
Partha S. Roop, Arcot Sowmya
ECRTS2
1998 A real-time variable sampling technique: DIEM
abstract
We describe a sampling technique particularly suitable for active vision, dimensionally-independent exponential mapping (DIEM), in which each dimension of the original data is sampled in an exponentially increasing or decreasing series of steps, with bilateral symmetry about the data mid-point. Multidimensional data sampling is achieved by combining single dimension sampling coordinates. DIEM is simple, fast, flexible and very useful for active vision, but may also have applications in other domains possibly of higher dimensionality. Its most unusual feature, invertibility, is also one of its most useful features. The many advantages of DIEM are described. We also describe the functional characteristics of DIEM, provide formulae for DIEM specification and verification, and refer to how DIEM may best be exploited, giving our own work in visual robotics as an example.
Mark W. Peters, Arcot Sowmya
ICPR2
1998 Extending Statecharts with Temporal Logic
abstract
The task of designing large real-time reactive systems, which interact continuously with their environment and exhibit concurrency properties, is a challenging one. The authors explore the utility of a combination of behavior and function specification languages in specifying such systems and verifying their properties. An existing specification language, statecharts, is used to specify the behavior of real-time reactive systems, while a new logic-based language called FNLOG (based on first-order predicate calculus and temporal logic) is designed to express the system functions over real time. Two types of system properties, intrinsic and structural, are proposed. It is shown that both types of system properties are expressible in FNLOG and may be verified by logical deduction, and also hold for the corresponding behavior specification.
Arcot Sowmya, S. Ramesh 0001
IEEE Trans. Software Eng.1
1996 Automatic Model Building from Images for Multimedia Systems
Arcot Sowmya, Maziar Palhang
MMM1