EDBT 2026 Demo / reviewers in the wild / expert
Haris Mansoor
dblp:243/0054
· DBLP profile ↗
3ranked-venue papers in the field
1as first author
3since 2021 · last 2024
0000-0002-7611-4165ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 2 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Molecular sequence classification using efficient kernel based embedding
Sarwan Ali, Tamkanat E. Ali, Taslim Murad, Haris Mansoor, Murray Patterson |
Inf. Sci. | 4 |
| 2023 | Circular Arc Length-Based Kernel Matrix For Protein Sequence ClassificationabstractBiological sequence analysis is crucial in understanding the sequence structure, function, and evolutionary relationships. In traditional methods, using Euclidean distance metrics is common in measuring the similarity between sequence embeddings. However, they fail to capture sequence space’s inherent curvature and spherical nature. Therefore, we explore the application of spherical geometry and distance metric, circular arc length (CAL) based distance for comparing biological sequence embeddings. Spherical geometry is a non-Euclidean geometry that models the surface of a sphere, accounting for its curvature. By leveraging CAL, we can more accurately measure the pairwise distances between bio-sequence embeddings. In this study, we propose the utilization of CAL for comparing bio-sequence embeddings. We develop a function to compute the CAL distance between two sequence embeddings, enabling researchers to accurately measure the similarity between sequences while considering the underlying spherical geometry. Incorporating spherical geometry enables a more comprehensive understanding of the relationships and similarities between biological sequences, improving various downstream tasks such as classification, clustering, and evolutionary analysis. Our experimental evaluation demonstrates the advantages of using a spherical distance metric over Euclidean metrics for bio-sequence analysis. Our proposed CAL-based approach outperforms the Euclidean geometry-based baselines by depicting a huge performance improvement for the protein subcellular location classification task. In our experiments, accuracy is improved by 52.5% and 60. 3% compared to the PWM2Vec and Autoencoder methods, respectively, corresponding to the DT classifier. Taslim Murad, Sarwan Ali, Prakash Chourasia, Haris Mansoor, Murray Patterson |
IEEE Big Data | 4 |
| 2022 | Impact Of Missing Data Imputation On The Fairness And Accuracy Of Graph Node ClassifiersabstractAnalysis of the fairness of machine learning (ML) algorithms has attracted many researchers’ interest. Several studies have shown that ML methods produce a bias toward different groups, which limits the applicability of ML models in many applications, such as crime rate prediction. The data used for ML may have missing values, which, if not appropriately handled, are known to further harmfully affect fairness. To address this issue, many imputation methods have been proposed to deal with missing data. However, research on the effect of missing data imputation on fairness is still rather limited. In this paper, we analyze the impact of imputation on fairness in the context of graph data (node attributes) using different embedding and neural network methods. Extensive experiments on six datasets demonstrate several issues of fairness in graph node classification when dealing with missing data and various imputation techniques. We find that the choice of the imputation method affects both fairness and accuracy. Our results provide valuable insights into fairness ML over graph data and how to handle missingness in graphs efficiently. Haris Mansoor, Sarwan Ali, Shafiq Alam, Muhammad Asad Khan, Umair ul Hassan |
IEEE Big Data | 1 |