VLDB 2026 Research / reviewers in the wild / expert
Fayyaz ul Amir Afsar Minhas
dblp:45/10683 · also Fayyaz A. Afsar, Fayyaz A. Minhas, Fayyaz Minhas
· DBLP profile ↗
26ranked-venue papers
2as first author
18since 2021 · last 2025
0000-0001-9129-1189ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HistoKernel: Whole slide image level Maximum Mean Discrepancy kernels for pan-cancer predictive modellingabstractIn computational pathology, labels are typically available only at the whole slide image (WSI) or patient level, necessitating weakly supervised learning methods that aggregate patch-level features or predictions to produce WSI-level scores for clinically significant tasks such as cancer subtype classification or survival analysis. However, existing approaches lack a theoretically grounded framework to capture the holistic distributional differences between the patch sets within WSIs, limiting their ability to accurately and comprehensively model the underlying pathology. To address this limitation, we introduce HistoKernel, a novel WSI-level Maximum Mean Discrepancy (MMD) kernel designed to quantify distributional similarity between WSIs using their local feature representation. HistoKernel enables a wide range of applications, including classification, regression, retrieval, clustering, survival analysis, multimodal data integration, and visualization of large WSI datasets. Additionally, HistoKernel offers a novel perturbation-based method for patch-level explainability. Our analysis over large pan-cancer datasets shows that HistoKernel achieves performance that typically matches or exceeds existing state-of-the-art methods across diverse tasks, including WSI retrieval (n = 9324), drug sensitivity regression (n = 551), point mutation classification (n = 3419), and survival analysis (n = 2291). By pioneering the use of kernel-based methods for a diverse range of WSI-level predictive tasks, HistoKernel opens new avenues for computational pathology research especially in terms of rapid prototyping on large and complex computational pathology datasets. Code and interactive visualization are available at: https://histokernel.dcs.warwick.ac.uk/. Piotr Keller, Muhammad Dawood, Brinder Singh Chohan, Fayyaz ul Amir Afsar Minhas |
Medical Image Anal. | 4 |
| 2024 | Physical Structure Representation and Environmental Data Fusion for Cyclone Intensity PredictionabstractAccurate tropical cyclone (TC) intensity prediction plays a crucial role in providing early warnings to coastal regions, helping prevent economic losses and ensuring the safety of individuals. Recently, deep learning (DL) for remote sensing images has been applied to TC intensity prediction and has achieved great success. However, most existing DL-based methods are unable to effectively leverage the changes in cyclone structure, which are important for TC intensity prediction. Moreover, these methods neglect the influence of physical environmental factors on TC intensity. In this paper, we propose a novel framework with Structural and Environmental Data Fusion Network (SEF-Net) to predict the maximum sustained wind (MSW) speed values near cyclone centers. Specifically, we extract spectral-spatial-temporal correlations by learning the changes in cyclone eye regions and cloud bands at different time steps. Subsequently, we perform multimodal fusion of image features and environmental factors to enhance the prediction accuracy of the network. Experimental results show that the proposed framework outperforms several state-of-the-art methods for TC intensity prediction on various cyclone datasets. Cong Wang 0003, Fayyaz ul Amir Afsar Minhas |
IGARSS | 3 |
| 2024 | A Two-Stage Neural Network Model for Automatic Detection of Skiffs from Satellite ImageryabstractThe detection of skiffs from satellite imagery has significant implications for maritime surveillance, search and rescue operations, and monitoring of illegal activities like piracy and smuggling. In contrast to larger vessels, skiffs can be very difficult to detect in images as they can easily blend with other objects or their surroundings. This paper presents a two-stage neural network method for the automatic detection of skiffs in satellite imagery. It also presents a new large benchmarking dataset of 9,420 annotated satellite images (1,525 images with multiple annotated skiffs and 7,895 images without skiffs) for the community. The proposed pipeline consists of two stages: the first stage is a filtering model to infer if there is a skiff in an input image or not, with the second stage being used to localize the skiffs in the image predicted positive by the filtering model. The first stage is implemented using a semantic segmentation model adapted as a binary classifier, achieving an AUC-ROC of 0.9964. The second stage is implemented using a Faster R-CNN object detector, achieving an average precision (AP) of 0.685 at intersection over union (IoU) of 0.5. This two-tiered approach not only significantly improves the detection rate of skiffs in satellite imagery, achieving an average precision of 0.736 at IoU of 0.5, but also presents a practical tool for the advancement of automated satellite imagery analysis. The implementation of the method as well as a demonstration and the associated dataset are available at the URLs: https://github.com/TasteDaBombDev/satcen-shipdetection. Tudor Ducaru, Stefan Stoica, Vlad Holban, Fayyaz ul Amir Afsar Minhas |
IJCNN | 4 |
| 2024 | Ouroboros: cross-linking protein expression perturbations and cancer histology imaging with generative-predictive modelingabstractSUMMARY: Imagine if we could simultaneously predict spatial protein expression in tissues from their routine Hematoxylin and Eosin (H&E) stained images, and create tissue images given protein expression profiles thus enabling virtual simulations of how protein expression alterations impact histology in complex diseases like cancer. Such an approach could lead to more informed diagnostic and therapeutic decisions for precision medicine at lower costs and shorter turnaround times, more detailed insights into underlying disease pathology as well as improvement in predictive and generative performance. In this study, we investigate the intricate correlation between protein expressions obtained from Hyperion mass cytometry and histopathological microstructures in conventional H&E stained glioblastoma (GBM) samples, unveiling morphological patterns and cellular-level spatial alterations associated with protein expression changes. To model these complex relationships, we propose a novel generative-predictive framework called Ouroboros for producing H&E images from protein expressions and simultaneously predicting protein expressions from H&E images. Our comprehensive sample-independent validation over 9920 tissue spots from 4 GBM samples encompassing visual image analysis, quantitative analysis, subspace alignment and perturbation experiments shows that the proposed generative-predictive approach offers significant improvements in predicting protein expression from images in comparison to baseline methods as well as accurate generation of virtual GBM sample images. This proof of concept study can contribute to advancing our understanding of histological responses to protein expression perturbations and lays the foundations for further developments in this area. AVAILABILITY AND IMPLEMENTATION: Implementation and associated data for the proposed approach are available at the URL: https://github.com/Srijay/Ouroboros. Srijay Deshpande, Sokratia Georgaka, Michael Haley, Robert Sellers, James Minshull, Jayakrupakar Nallala, Martin Fergie, Nicholas Stone, Nasir M. Rajpoot, Syed Murtuza Baker, Mudassar Iqbal, Kevin Couper, Federico Roncaroli, Fayyaz ul Amir Afsar Minhas |
Bioinform. | 14 |
| 2024 | SynCLay: Interactive synthesis of histology images from bespoke cellular layoutsabstractAutomated synthesis of histology images has several potential applications in computational pathology. However, no existing method can generate realistic tissue images with a bespoke cellular layout or user-defined histology parameters. In this work, we propose a novel framework called SynCLay (Synthesis from Cellular Layouts) that can construct realistic and high-quality histology images from user-defined cellular layouts along with annotated cellular boundaries. Tissue image generation based on bespoke cellular layouts through the proposed framework allows users to generate different histological patterns from arbitrary topological arrangement of different types of cells (e.g., neutrophils, lymphocytes, epithelial cells and others). SynCLay generated synthetic images can be helpful in studying the role of different types of cells present in the tumor microenvironment. Additionally, they can assist in balancing the distribution of cellular counts in tissue images for designing accurate cellular composition predictors by minimizing the effects of data imbalance. We train SynCLay in an adversarial manner and integrate a nuclear segmentation and classification model in its training to refine nuclear structures and generate nuclear masks in conjunction with synthetic images. During inference, we combine the model with another parametric model for generating colon images and associated cellular counts as annotations given the grade of differentiation and cellularities (cell densities) of different cells. We assess the generated images quantitatively using the Frechet Inception Distance and report on feedback from trained pathologists who assigned realism scores to a set of images generated by the framework. The average realism score across all pathologists for synthetic images was as high as that for the real images. Moreover, with the assistance from pathologists, we showcase the ability of the generated images to accurately differentiate between benign and malignant tumors, thus reinforcing their reliability. We demonstrate that the proposed framework can be used to add new cells to a tissue images and alter cellular positions. We also show that augmenting limited real data with the synthetic data generated by our framework can significantly boost prediction performance of the cellular composition prediction task. The implementation of the proposed SynCLay framework is available at https://github.com/Srijay/SynCLay-Framework. Srijay Deshpande, Muhammad Dawood, Fayyaz ul Amir Afsar Minhas, Nasir M. Rajpoot |
Medical Image Anal. | 3 |
| 2024 | CoNIC Challenge: Pushing the frontiers of nuclear detection, segmentation, classification and countingabstractNuclear detection, segmentation and morphometric profiling are essential in helping us further understand the relationship between histology and patient outcome. To drive innovation in this area, we setup a community-wide challenge using the largest available dataset of its kind to assess nuclear segmentation and cellular composition. Our challenge, named CoNIC, stimulated the development of reproducible algorithms for cellular recognition with real-time result inspection on public leaderboards. We conducted an extensive post-challenge analysis based on the top-performing models using 1,658 whole-slide images of colon tissue. With around 700 million detected nuclei per model, associated features were used for dysplasia grading and survival analysis, where we demonstrated that the challenge's improvement over the previous state-of-the-art led to significant boosts in downstream performance. Our findings also suggest that eosinophils and neutrophils play an important role in the tumour microevironment. We release challenge models and WSI-level results to foster the development of further methods for biomarker discovery. Simon Graham, Quoc Dang Vu, Mostafa Jahanifar, Martin Weigert 0001, Jun Zhang 0018, Sen Yang 0006, Jinxi Xiang, Josef Lorenz Rumberger, Elias Baumann, Peter Hirsch 0001, Chenyang Hong, Angelica I. Avilés-Rivero, Ayushi Jain, Heeyoung Ahn, Yiyu Hong, Hussam Azzuni, Min Xu 0009, Mohammad Yaqub, Marie-Claire Blache, Benoît Piégu, Bertrand Vernay, Tim Scherr, Moritz Böhland, Katharina Löffler, Weiqin Ying, Chixin Wang, David R. J. Snead, Shan E Ahmed Raza, Fayyaz ul Amir Afsar Minhas, Nasir M. Rajpoot |
Medical Image Anal. | 34 |
| 2024 | Mitosis detection, fast and slow: Robust and efficient detection of mitotic figuresabstractCounting of mitotic figures is a fundamental step in grading and prognostication of several cancers. However, manual mitosis counting is tedious and time-consuming. In addition, variation in the appearance of mitotic figures causes a high degree of discordance among pathologists. With advances in deep learning models, several automatic mitosis detection algorithms have been proposed but they are sensitive to domain shift often seen in histology images. We propose a robust and efficient two-stage mitosis detection framework, which comprises mitosis candidate segmentation (Detecting Fast) and candidate refinement (Detecting Slow) stages. The proposed candidate segmentation model, termed EUNet, is fast and accurate due to its architectural design. EUNet can precisely segment candidates at a lower resolution to considerably speed up candidate detection. Candidates are then refined using a deeper classifier network, EfficientNet-B7, in the second stage. We make sure both stages are robust against domain shift by incorporating domain generalization methods. We demonstrate state-of-the-art performance and generalizability of the proposed model on the three largest publicly available mitosis datasets, winning the two mitosis domain generalization challenge contests (MIDOG21 and MIDOG22). Finally, we showcase the utility of the proposed algorithm by processing the TCGA breast cancer cohort (1,124 whole-slide images) to generate and release a repository of more than 620K potential mitotic figures (not exhaustively validated). Mostafa Jahanifar, Adam J. Shephard, Neda Zamani Tajeddin, Simon Graham, Shan E Ahmed Raza, Fayyaz ul Amir Afsar Minhas, Nasir M. Rajpoot |
Medical Image Anal. | 6 |
| 2024 | Physics-Informed Learning for Tropical Cyclone Intensity PredictionabstractAccurate prediction of tropical cyclone (TC) intensity is important to disaster prevention. However, many deep learning (DL) models analyzing remote sensing images for TC intensity prediction lack explainability or interpretability due to insufficient consideration of physical knowledge. Therefore, we propose physics-informed cyclone intensity prediction networks (Pici-Nets) to forecast TC intensity measured by the maximum sustained wind (MSW) speed. There are three types of physical information closely related to MSW being used, the structural knowledge of cyclones reflected by multitemporal multispectral images (MSIs), the motional prior representing the dynamics of TC, and the environmental factors affecting the evolvement of TC provided by the public track data. As the multimodal information are fused by Pici-Nets, the MSW speed prediction process which imitates the manual forecast by considering multiple factors to reduce the prediction errors can be regarded as interpretable. Specifically, the first type of information guides Pici-Nets to sharpen the low-spatial-resolution MSIs, enhance temporal structural features, and produce explainable feature maps. Since the second or the third type of data are sometimes missing, we derive Pici-Net+ and Pici-Net++ from the vanilla Pici-Net to cope with the situations. In Pici-Net++, we even incorporate a physics simulation model (PSM) to simulate the missing data. This way, the models become applicable to incomplete datasets and more interpretable in MSW speed prediction. Experiments on different TC datasets validate the efficacy of Pici-Nets, as the results show that Pici-Nets outperform many classic and state-of-the-art (SOTA) models in TC intensity prediction. Cong Wang 0003, Fayyaz ul Amir Afsar Minhas |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Malignant Mesothelioma subtyping via sampling driven multiple instance prediction on tissue image and cell morphology dataabstractMalignant Mesothelioma is a difficult to diagnose and highly lethal cancer usually associated with asbestos exposure. It can be broadly classified into three subtypes: Epithelioid, Sarcomatoid, and a hybrid Biphasic subtype in which significant components of both of the previous subtypes are present. Early diagnosis and identification of the subtype informs treatment and can help improve patient outcome. However, the subtyping of malignant mesothelioma, and specifically the recognition of transitional features from routine histology slides has a high level of inter-observer variability. In this work, we propose an end-to-end multiple instance learning (MIL) approach for malignant mesothelioma subtyping. This uses an adaptive instance-based sampling scheme for training deep convolutional neural networks on bags of image patches that allows learning on a wider range of relevant instances compared to max or top-N based MIL approaches. We also investigate augmenting the instance representation to include aggregate cellular morphology features from cell segmentation. The proposed MIL approach enables identification of malignant mesothelial subtypes of specific tissue regions. From this a continuous characterisation of a sample according to predominance of sarcomatoid vs epithelioid regions is possible, thus avoiding the arbitrary and highly subjective categorisation by currently used subtypes. Instance scoring also enables studying tumor heterogeneity and identifying patterns associated with different subtypes. We have evaluated the proposed method on a dataset of 234 tissue micro-array cores with an AUROC of 0.89±0.05 for this task. The dataset and developed methodology is available for the community at: https://github.com/measty/PINS. Mark Eastwood, Silviu Tudor Marc, Xiaohong W. Gao, Heba Sailem, Judith Offman, Emmanouil Karteris, Angeles Montero Fernandez, Danny Jonigk, William Cookson, Miriam Moffatt, Sanjay Popat, Fayyaz ul Amir Afsar Minhas, Jan L. Robertus |
Artif. Intell. Medicine | 12 |
| 2023 | One model is all you need: Multi-task learning enables simultaneous histology image segmentation and classificationabstractThe recent surge in performance for image analysis of digitised pathology slides can largely be attributed to the advances in deep learning. Deep models can be used to initially localise various structures in the tissue and hence facilitate the extraction of interpretable features for biomarker discovery. However, these models are typically trained for a single task and therefore scale poorly as we wish to adapt the model for an increasing number of different tasks. Also, supervised deep learning models are very data hungry and therefore rely on large amounts of training data to perform well. In this paper, we present a multi-task learning approach for segmentation and classification of nuclei, glands, lumina and different tissue regions that leverages data from multiple independent data sources. While ensuring that our tasks are aligned by the same tissue type and resolution, we enable meaningful simultaneous prediction with a single network. As a result of feature sharing, we also show that the learned representation can be used to improve the performance of additional tasks via transfer learning, including nuclear classification and signet ring cell detection. As part of this work, we train our developed Cerberus model on a huge amount of data, consisting of over 600 thousand objects for segmentation and 440 thousand patches for classification. We use our approach to process 599 colorectal whole-slide images from TCGA, where we localise 377 million, 900 thousand and 2.1 million nuclei, glands and lumina respectively. We make this resource available to remove a major barrier in the development of explainable models for computational pathology. Simon Graham, Quoc Dang Vu, Mostafa Jahanifar, Shan E Ahmed Raza, Fayyaz ul Amir Afsar Minhas, David R. J. Snead, Nasir M. Rajpoot |
Medical Image Anal. | 5 |
| 2022 | Malignant Mesothelioma Subtyping of Tissue Images via Sampling Driven Multiple Instance Prediction
Mark Eastwood, Silviu Tudor Marc, Xiaohong W. Gao, Heba Sailem, Judith Offman, Emmanouil Karteris, Angeles Montero Fernandez, Danny Jonigk, William Cookson, Miriam Moffatt, Sanjay Popat, Fayyaz ul Amir Afsar Minhas, Jan L. Robertus |
AIME | 12 |
| 2022 | REET: robustness evaluation and enhancement toolbox for computational pathologyabstractMOTIVATION: Digitization of pathology laboratories through digital slide scanners and advances in deep learning approaches for objective histological assessment have resulted in rapid progress in the field of computational pathology (CPath) with wide-ranging applications in medical and pharmaceutical research as well as clinical workflows. However, the estimation of robustness of CPath models to variations in input images is an open problem with a significant impact on the downstream practical applicability, deployment and acceptability of these approaches. Furthermore, development of domain-specific strategies for enhancement of robustness of such models is of prime importance as well. RESULTS: In this work, we propose the first domain-specific Robustness Evaluation and Enhancement Toolbox (REET) for computational pathology applications. It provides a suite of algorithmic strategies for enabling robustness assessment of predictive models with respect to specialized image transformations such as staining, compression, focusing, blurring, changes in spatial resolution, brightness variations, geometric changes as well as pixel-level adversarial perturbations. Furthermore, REET also enables efficient and robust training of deep learning pipelines in computational pathology. Python implementation of REET is available at https://github.com/alexjfoote/reetoolbox. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alex Foote, Amina Asif, Nasir M. Rajpoot, Fayyaz ul Amir Afsar Minhas |
Bioinform. | 4 |
| 2022 | Insights into performance evaluation of compound-protein interaction prediction methodsabstractMOTIVATION: Machine-learning-based prediction of compound-protein interactions (CPIs) is important for drug design, screening and repurposing. Despite numerous recent publication with increasing methodological sophistication claiming consistent improvements in predictive accuracy, we have observed a number of fundamental issues in experiment design that produce overoptimistic estimates of model performance. RESULTS: We systematically analyze the impact of several factors affecting generalization performance of CPI predictors that are overlooked in existing work: (i) similarity between training and test examples in cross-validation; (ii) synthesizing negative examples in absence of experimentally verified negative examples and (iii) alignment of evaluation protocol and performance metrics with real-world use of CPI predictors in screening large compound libraries. Using both state-of-the-art approaches by other researchers as well as a simple kernel-based baseline, we have found that effective assessment of generalization performance of CPI predictors requires careful control over similarity between training and test examples. We show that, under stringent performance assessment protocols, a simple kernel-based approach can exceed the predictive performance of existing state-of-the-art methods. We also show that random pairing for generating synthetic negative examples for training and performance evaluation results in models with better generalization in comparison to more sophisticated strategies used in existing studies. Our analyses indicate that using proposed experiment design strategies can offer significant improvements for CPI prediction leading to effective target compound screening for drug repurposing and discovery of putative chemical ligands of SARS-CoV-2-Spike and Human-ACE2 proteins. AVAILABILITY AND IMPLEMENTATION: Code and supplementary material available at https://github.com/adibayaseen/HKRCPI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Adiba Yaseen, Imran Amin, Naeem Akhter, Asa Ben-Hur, Fayyaz ul Amir Afsar Minhas |
Bioinform. | 5 |
| 2022 | SAFRON: Stitching Across the Frontier Network for Generating Colorectal Cancer Histology Images
Srijay Deshpande, Fayyaz ul Amir Afsar Minhas, Simon Graham, Nasir M. Rajpoot |
Medical Image Anal. | 2 |
| 2022 | SlideGraph+: Whole slide image level graphs to predict HER2 status in breast cancerabstractHuman epidermal growth factor receptor 2 (HER2) is an important prognostic and predictive factor which is overexpressed in 15–20% of breast cancer (BCa). The determination of its status is a key clinical decision making step for selection of treatment regimen and prognostication. HER2 status is evaluated using transcriptomics or immunohistochemistry (IHC) through in-situ hybridisation (ISH) which incurs additional costs and tissue burden and is prone to analytical variabilities in terms of manual observational biases in scoring. In this study, we propose a novel graph neural network (GNN) based model (SlideGraph+) to predict HER2 status directly from whole-slide images of routine Haematoxylin and Eosin (H&E) stained slides. The network was trained and tested on slides from The Cancer Genome Atlas (TCGA) in addition to two independent test datasets. We demonstrate that the proposed model outperforms the state-of-the-art methods with area under the ROC curve (AUC) values > 0.75 on TCGA and 0.80 on independent test sets. Our experiments show that the proposed approach can be utilised for case triaging as well as pre-ordering diagnostic tests in a diagnostic setting. It can also be used for other weakly supervised prediction problems in computational pathology. The SlideGraph+ code repository is available at https://github.com/wenqi006/SlideGraph along with an IPython notebook showing an end-to-end use case at https://github.com/TissueImageAnalytics/tiatoolbox/blob/develop/examples/full-pipelines/slide-graph.ipynb. Wenqi Lu 0001, Michael Toss, Muhammad Dawood, Emad Rakha, Nasir M. Rajpoot, Fayyaz ul Amir Afsar Minhas |
Medical Image Anal. | 6 |
| 2022 | AMP0: Species-Specific Prediction of Anti-microbial Peptides Using Zero and Few Shot LearningabstractEvolution of drug-resistant microbial species is one of the major challenges to global health. Development of new antimicrobial treatments such as antimicrobial peptides needs to be accelerated to combat this threat. However, the discovery of novel antimicrobial peptides is hampered by low-throughput biochemical assays. Computational techniques can be used for rapid screening of promising antimicrobial peptide candidates prior to testing in the wet lab. The vast majority of existing antimicrobial peptide predictors arenon-targetedin nature, i.e., they can predict whether a given peptide sequence is antimicrobial, but they are unable to predict whether the sequence can target a particular microbial species. In this work, we have used zero and few shot machine learning to develop a targeted antimicrobial peptide activity predictor called AMP0. The proposed predictor takes the sequence of a peptide and any N/C-termini modifications together with the genomic sequence of a microbial species to generate targeted predictions. Cross-validation results show that the proposed scheme is particularly effective for targeted antimicrobial prediction in comparison to existing approaches and can be used for screening potential antimicrobial peptides in a targeted manner with only a small number of training examples for novel species. AMP0webserver is available athttp://ampzero.pythonanywhere.com. Sadaf Gull, Fayyaz ul Amir Afsar Minhas |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | A Generalized Meta-loss Function for Distillation Based Learning Using Privileged Information for Classification and Regression
Amina Asif, Muhammad Dawood, Fayyaz ul Amir Afsar Minhas |
ICANN (3) | 3 |
| 2021 | MILAMP: Multiple Instance Prediction of Amyloid ProteinsabstractAmyloid proteins are implicated in several diseases such as Parkinson's, Alzheimer's, prion diseases, etc. In order to characterize the amyloidogenicity of a given protein, it is important to locate the amyloid forming hotspot regions within the protein as well as to analyze the effects of mutations on these proteins. The biochemical and biological assays used for this purpose can be facilitated by computational means. This paper presents a machine learning method that can predict hotspot amyloidogenic regions within proteins and characterize changes in their amyloidogenicity due to point mutations. The proposed method called MILAMP (Multiple Instance Learning of AMyloid Proteins) achieves high accuracy for identification of amyloid proteins, hotspot localization, and prediction of mutation effects on amyloidogenicity by integrating heterogenous data sources and exploiting common predictive patterns across these tasks through multiple instance learning. The paper presents comprehensive benchmarking experiments to test the predictive performance of MILAMP in comparison to previously published state of the art techniques for amyloid prediction. The python code for the implementation and webserver for MILAMP is available at the URL: http://faculty.pieas.edu.pk/fayyaz/software.html#MILAMP. Farzeen Munir, Sadaf Gull, Amina Asif, Fayyaz ul Amir Afsar Minhas |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | Generalized Neural Framework for Learning with RejectionabstractLearning with Rejection (LWR) allows development of machine learning systems with the ability to discard low confidence decisions generated by a prediction model. That is, just like human experts, LWR allows machine models to abstain from generating a prediction when reliability of the prediction is expected to be low. Several frameworks for learning with rejection have been proposed in the literature. However, most of them work for classification problems only and regression with rejection has not been studied in much detail. In this work, we present a neural framework for LWR based on a generalized meta-loss function that involves simultaneous training of two neural network models: a predictor model for generating predictions and a rejecter model for deciding whether the prediction should be accepted or rejected. The proposed framework can be used for classification as well as regression and other related machine learning tasks. We have demonstrated the applicability and effectiveness of the method on synthetically generated data as well as benchmark datasets from UCI machine learning repository for both classification and regression problems. Despite being simpler in implementation, the proposed scheme for learning with rejection has shown to perform at par or better than previously proposed methods. Furthermore, we have applied the method to the problem of hurricane intensity prediction from satellite imagery. Significant improvement in performance as compared to conventional supervised methods shows the effectiveness of the proposed scheme in real-world regression problems. Python code files for the experiments can be found at: https://github.com/amina01/LWR. Amina Asif, Fayyaz ul Amir Afsar Minhas |
IJCNN | 2 |
| 2020 | PHURIE: hurricane intensity estimation from infrared satellite imagery using machine learning
Amina Asif, Muhammad Dawood, Bismillah Jan, Javaid Khurshid, Mark DeMaria, Fayyaz ul Amir Afsar Minhas |
Neural Comput. Appl. | 6 |
| 2020 | Deep-PHURIE: deep learning based hurricane intensity estimation from infrared satellite imagery
Muhammad Dawood, Amina Asif, Fayyaz ul Amir Afsar Minhas |
Neural Comput. Appl. | 3 |
| 2019 | An embarrassingly simple approach to neural multiple instance classification
Amina Asif, Fayyaz ul Amir Afsar Minhas |
Pattern Recognit. Lett. | 2 |
| 2018 | Learning protein binding affinity using privileged informationabstractBACKGROUND: Determining protein-protein interactions and their binding affinity are important in understanding cellular biological processes, discovery and design of novel therapeutics, protein engineering, and mutagenesis studies. Due to the time and effort required in wet lab experiments, computational prediction of binding affinity from sequence or structure is an important area of research. Structure-based methods, though more accurate than sequence-based techniques, are limited in their applicability due to limited availability of protein structure data. RESULTS: In this study, we propose a novel machine learning method for predicting binding affinity that uses protein 3D structure as privileged information at training time while expecting only protein sequence information during testing. Using the method, which is based on the framework of learning using privileged information (LUPI), we have achieved improved performance over corresponding sequence-based binding affinity prediction methods that do not have access to privileged information during training. Our experiments show that with the proposed framework which uses structure only during training, it is possible to achieve classification performance comparable to that which is obtained using structure-based features. Evaluation on an independent test set shows improved performance over the PPA-Pred2 method as well. CONCLUSIONS: The proposed method outperforms several baseline learners and a state-of-the-art binding affinity predictor not only in cross-validation, but also on an additional validation dataset, demonstrating the utility of the LUPI framework for problems that would benefit from classification using structure-based features. The implementation of LUPI developed for this work is expected to be useful in other areas of bioinformatics as well. Wajid Arshad Abbasi, Amina Asif, Asa Ben-Hur, Fayyaz ul Amir Afsar Minhas |
BMC Bioinform. | 4 |
| 2017 | Amino acid composition predicts prion activityabstractMany prion-forming proteins contain glutamine/asparagine (Q/N) rich domains, and there are conflicting opinions as to the role of primary sequence in their conversion to the prion form: is this phenomenon driven primarily by amino acid composition, or, as a recent computational analysis suggested, dependent on the presence of short sequence elements with high amyloid-forming potential. The argument for the importance of short sequence elements hinged on the relatively-high accuracy obtained using a method that utilizes a collection of length-six sequence elements with known amyloid-forming potential. We weigh in on this question and demonstrate that when those sequence elements are permuted, even higher accuracy is obtained; we also propose a novel multiple-instance machine learning method that uses sequence composition alone, and achieves better accuracy than all existing prion prediction approaches. While we expect there to be elements of primary sequence that affect the process, our experiments suggest that sequence composition alone is sufficient for predicting protein sequences that are likely to form prions. A web-server for the proposed method is available at http://faculty.pieas.edu.pk/fayyaz/prank.html, and the code for reproducing our experiments is available at http://doi.org/10.5281/zenodo.167136. Fayyaz ul Amir Afsar Minhas, Eric D. Ross, Asa Ben-Hur |
PLoS Comput. Biol. | 1 |
| 2012 | Multiple instance learning of Calmodulin binding sitesabstractMOTIVATION: Calmodulin (CaM) is a ubiquitously conserved protein that acts as a calcium sensor, and interacts with a large number of proteins. Detection of CaM binding proteins and their interaction sites experimentally requires a significant effort, so accurate methods for their prediction are important. RESULTS: We present a novel algorithm (MI-1 SVM) for binding site prediction and evaluate its performance on a set of CaM-binding proteins extracted from the Calmodulin Target Database. Our approach directly models the problem of binding site prediction as a large-margin classification problem, and is able to take into account uncertainty in binding site location. We show that the proposed algorithm performs better than the standard SVM formulation, and illustrate its ability to recover known CaM binding motifs. A highly accurate cascaded classification approach using the proposed binding site prediction method to predict CaM binding proteins in Arabidopsis thaliana is also presented. AVAILABILITY: Matlab code for training MI-1 SVM and the cascaded classification approach is available on request. CONTACT: [email protected] or [email protected]. Fayyaz ul Amir Afsar Minhas, Asa Ben-Hur |
Bioinform. | 1 |
| 2009 | Arif Index for Predicting the Classification Accuracy of Features and Its Application in Heart Beat Classification Problem
Muhammad Arif 0006, Fayyaz ul Amir Afsar Minhas, M. Usman Akram, Adnan Fida |
PAKDD | 2 |