VLDB 2026 Research / reviewers in the wild / expert
Haris Mansoor
dblp:243/0054
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-7611-4165ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hist2Vec: A histogram and kernel-based embedding method for molecular sequence analysis
Sarwan Ali, Tamkanat E. Ali, Haris Mansoor, Prakash Chourasia, Murray Patterson |
Expert Syst. Appl. | 3 |
| 2024 | Gaussian Beltrami-Klein Model for Protein Sequence Classification: A Hyperbolic Approach
Sarwan Ali, Haris Mansoor, Prakash Chourasia, Murray Patterson |
ISBRA (1) | 2 |
| 2024 | Molecular sequence classification using efficient kernel based embedding
Sarwan Ali, Tamkanat E. Ali, Taslim Murad, Haris Mansoor, Murray Patterson |
Inf. Sci. | 4 |
| 2024 | Elliptic geometry-based kernel matrix for improved biological sequence classification
Sarwan Ali, Madiha Shabbir, Haris Mansoor, Prakash Chourasia, Murray Patterson |
Knowl. Based Syst. | 3 |
| 2023 | Circular Arc Length-Based Kernel Matrix For Protein Sequence ClassificationabstractBiological sequence analysis is crucial in understanding the sequence structure, function, and evolutionary relationships. In traditional methods, using Euclidean distance metrics is common in measuring the similarity between sequence embeddings. However, they fail to capture sequence space’s inherent curvature and spherical nature. Therefore, we explore the application of spherical geometry and distance metric, circular arc length (CAL) based distance for comparing biological sequence embeddings. Spherical geometry is a non-Euclidean geometry that models the surface of a sphere, accounting for its curvature. By leveraging CAL, we can more accurately measure the pairwise distances between bio-sequence embeddings. In this study, we propose the utilization of CAL for comparing bio-sequence embeddings. We develop a function to compute the CAL distance between two sequence embeddings, enabling researchers to accurately measure the similarity between sequences while considering the underlying spherical geometry. Incorporating spherical geometry enables a more comprehensive understanding of the relationships and similarities between biological sequences, improving various downstream tasks such as classification, clustering, and evolutionary analysis. Our experimental evaluation demonstrates the advantages of using a spherical distance metric over Euclidean metrics for bio-sequence analysis. Our proposed CAL-based approach outperforms the Euclidean geometry-based baselines by depicting a huge performance improvement for the protein subcellular location classification task. In our experiments, accuracy is improved by 52.5% and 60. 3% compared to the PWM2Vec and Autoencoder methods, respectively, corresponding to the DT classifier. Taslim Murad, Sarwan Ali, Prakash Chourasia, Haris Mansoor, Murray Patterson |
IEEE Big Data | 4 |
| 2023 | Hist2Vec: Kernel-Based Embeddings for Biological Sequence Classification
Sarwan Ali, Haris Mansoor, Prakash Chourasia, Murray Patterson |
ISBRA | 2 |
| 2023 | Short-Term Load Forecasting Using AMI DataabstractAccurate short-term load forecasting is essential for the efficient operation of the power sector. Forecasting load at a fine granularity such as hourly loads of individual households is challenging due to higher volatility and inherent stochasticity. At the aggregate levels, such as monthly load at a grid, the uncertainties and fluctuations are averaged out; hence predicting load is more straightforward. This paper proposes a method called Forecasting using Matrix Factorization (fmf) for short-term load forecasting (stlf). fmf only utilizes historical data from consumers’ smart meters to forecast future loads (does not use any non-calendar attributes, consumers’ demographics or activity patterns information, etc.) and can be applied to any locality. A prominent feature of fmf is that it works at any level of user-specified granularity, both in the temporal (from a single hour to days) and spatial dimensions (a single household to groups of consumers). We empirically evaluate fmf on three benchmark datasets and demonstrate that it significantly outperforms the state-of-the-art methods in terms of load forecasting. The computational complexity of fmf is also substantially less than known methods for stlf such as long short-term memory neural networks, random forest, support vector machines, and regression trees. Haris Mansoor, Sarwan Ali, Naveed Arshad, Muhammad Asad Khan, Safiullah Faizullah |
IEEE Internet Things J. | 1 |
| 2022 | Impact Of Missing Data Imputation On The Fairness And Accuracy Of Graph Node ClassifiersabstractAnalysis of the fairness of machine learning (ML) algorithms has attracted many researchers’ interest. Several studies have shown that ML methods produce a bias toward different groups, which limits the applicability of ML models in many applications, such as crime rate prediction. The data used for ML may have missing values, which, if not appropriately handled, are known to further harmfully affect fairness. To address this issue, many imputation methods have been proposed to deal with missing data. However, research on the effect of missing data imputation on fairness is still rather limited. In this paper, we analyze the impact of imputation on fairness in the context of graph data (node attributes) using different embedding and neural network methods. Extensive experiments on six datasets demonstrate several issues of fairness in graph node classification when dealing with missing data and various imputation techniques. We find that the choice of the imputation method affects both fairness and accuracy. Our results provide valuable insights into fairness ML over graph data and how to handle missingness in graphs efficiently. Haris Mansoor, Sarwan Ali, Shafiq Alam, Muhammad Asad Khan, Umair ul Hassan |
IEEE Big Data | 1 |