VLDB 2026 Research / reviewers in the wild / expert
M. Saqib Nawaz
dblp:192/8371 · also Muhammad Saqib Nawaz
· DBLP profile ↗
19ranked-venue papers
13as first author
15since 2021 · last 2026
0000-0001-9856-2885ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 8 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HieRMVir: Interpretable Viral Classification via Hierarchical Deep LearningabstractAccurate identification of pathogens-especially those with pandemic potential-remains a significant challenge, particularly when traditional sequence alignment methods fail. While recent genome sequence identification methods have shown promise, most do not account for the hierarchical structure of biological taxonomy or the varying informativeness of genomic features across classification levels. To address these limitations, we propose HieRMVir (Hierarchical Random forest and Mutual information-based Viral genome classifier), a novel hierarchical deep learning framework that integrates random forest (RF)-based feature weighting with mutual information (MI)-guided attention regularization for interpretable and accurate viral sequence classification. HieRMVir performs classification across three levels and leverages feature importance scores from RF to scale input features, while MI scores are used to guide the attention mechanism towards statistically informative k-mer patterns through regularized loss. Experimental results on over one million genome sequences demonstrate that HieRMVir achieves an average accuracy of 95.8% (95% CI: 95.3-96.4%), outperforming existing methods on multiple metrics. Evaluation using hierarchical performance metrics and the analysis of learned attention weights further highlight the biological relevance and interpretability of HieRMVir. M. Saqib Nawaz, Philippe Fournier-Viger, Shoaib Nawaz, Youxi Wu, Wei Song 0004 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | A Multipurpose Protein Compressor Based on MDL and Genetic AlgorithmabstractThe rapid expansion of protein sequence databases has created challenges for efficient storage, transmission, and analysis. Unlike genomic sequences with only four nucleotide bases, proteins are composed of twenty amino acids, making compression more complex. Existing specialized protein compressors, such as AC, AC2 and CPM-FCM, have achieved promising performance but still face limitations, including high computational cost, low adaptability and limited biological interpretability. This paper introduces GMP (Genetic algorithm-based MDL Protein compressor), a novel protein compression framework that leverages the Minimum Description Length (MDL) principle with a genetic algorithm to discover optimal patterns of amino acid subsequences (kAA-mers). Experimental results demonstrate that GMP attains compression performance comparable to state-of-the-art methods while additionally supporting tasks such as classification and clustering-capabilities absent from traditional protein compressors. This makes GMP not only an efficient compression framework but also a biologically interpretable tool for protein sequence analysis. GMP is available at github.com/MuhammadzohaibNawaz/GMP M. Zohaib Nawaz, M. Saqib Nawaz, Philippe Fournier-Viger, Xinzheng Niu, Mengqiu Li |
BIBM | 2 |
| 2025 | GRIMP: A Genetic Algorithm for Compression-Based Descriptive Pattern MiningabstractABSTRACT Traditional frequent pattern mining algorithms often report an overwhelming number of patterns in large datasets, many of which are redundant. To address this issue, Minimum Description Length (MDL)‐based methods have been employed, which use data compression to capture a smaller yet significant set of patterns. However, finding a good set of patterns according to MDL involves a very large search space, and current MDL‐based techniques often suffer from long runtimes and find suboptimal solutions. To discover better sets of patterns in less time, this paper introduces GRIMP (a Genetic algoRIthm for coMpression‐based descriptive Pattern mining), a novel framework that combines a genetic algorithm with MDL‐based pattern selection. Multiple genetic algorithm variants are explored within the GRIMP framework, and their effectiveness is compared using a large number of datasets. Experimental results demonstrate that GRIMP consistently outperforms previous methods by achieving higher compression ratios, generating more representative itemsets, and requiring less time. Additionally, the extracted patterns improve downstream classification tasks, highlighting the ability of GRIMP to find more representative patterns within the data. Muhammad Zohaib Nawaz, M. Saqib Nawaz, Philippe Fournier-Viger, Nazha Selmaoui-Folcher |
Expert Syst. J. Knowl. Eng. | 2 |
| 2025 | In silico framework for genome analysis
M. Saqib Nawaz, Muhammad Zohaib Nawaz, Yongshun Gong, Philippe Fournier-Viger, Abdoulaye Baniré Diallo |
Future Gener. Comput. Syst. | 1 |
| 2024 | SeqClin: Pattern-Based Analysis and Classification of Clinical DatasetsabstractAccurate analysis and classification of clinical datasets are crucial for understanding disease patterns, identifying risk factors and devising targeted interventions that ultimately contribute towards effective healthcare systems and improved patient outcomes. However, existing analysis and classification methods often fall short of effectively capturing complex sequential relationships within patient data and have limited interpretability. To overcome these challenges, we introduce SeqClin, a novel approach that utilizes frequent pattern mining to obtain valuable sequential information from clinical datasets. SeqClin first transforms clinical datasets into an appropriate format. Then, it employs sequential pattern mining algorithms to find frequent sequential patterns as well as rules of patient features in the datasets. These identified feature patterns and their respective values are then used for classification/detection. The performance of SeqClin is evaluated on four clinical datasets, where six classification models and evaluation metrics are employed for a comprehensive assessment. The obtained results show that the proposed approach surpassed previous approaches, with the extracted patterns and rules providing valuable insights into the key patient features and their values in clinical datasets. M. Saqib Nawaz, Philippe Fournier-Viger, Jimmy Ming-Tai Wu |
BIBM | 1 |
| 2024 | An MDL-Based Genetic Algorithm for Genome Sequence CompressionabstractThe exponential growth of genomic data has posed significant challenges for lossless compression of genome sequences. While recent reference-free genome compressors have shown promising results, they often fail to fully leverage the inherent sequential structure of genome sequences, require substantial computational resources and lack (or have limited) interpretability. This paper presents a novel genome compression method that employs the Minimum Description Length (MDL) principle, which is based on the idea that the best model for a given dataset is the one that provides the shortest description of that dataset. The proposed compressor, called GMG (Genetic algorithm for MDL-based Genome compression), integrates a genetic algorithm to identify optimal k-mers (patterns) in a model to best compress the genome data. Experimental results across various datasets demonstrate that GMG outperforms state-of-the-art genome compressors in terms of bits-per-base compression and computational efficiency. Furthermore, it is demonstrated that the optimal patterns identified by GMG for compression can also be utilized for genome classification, offering a multifunctional advantage over previous compressors. GMG is freely available at github.com/MuhammadzohaibNawaz/GMG Muhammad Zohaib Nawaz, M. Saqib Nawaz, Philippe Fournier-Viger, Vincent S. Tseng |
BIBM | 2 |
| 2023 | Using alignment-free and pattern mining methods for SARS-CoV-2 genome analysis
M. Saqib Nawaz, Philippe Fournier-Viger, Memoona Aslam, Wenjin Li, Yu-Lin He, Xinzheng Niu |
Appl. Intell. | 1 |
| 2023 | MDVA-GAN: multi-domain visual attribution generative adversarial networks
M. Saqib Nawaz, Feras N. Al-Obeidat, Abdallah Tubaishat, Tehseen Zia, Fahad Maqbool, Álvaro Rocha 0001 |
Neural Comput. Appl. | 1 |
| 2022 | S-PDB: Analysis and Classification of SARS-CoV-2 Spike Protein StructuresabstractThis paper proposes a novel and efficient method, called S-PDB, for the analysis and classification of Spike (S) protein structures of SARS-CoV-2 and other viruses/organisms in the Protein Data Bank (PDB). The method first finds and identifies protein structures in PDB that are similar to a protein structure of interest (SARS-CoV-2 S) via a protein structure comparison tool. The amino acid (AA) sequences of identified protein structures, downloaded from PDB, and their aligned amino acids (AAA) and secondary structure elements (ASSE), that are stored in three separate datasets, are then used for the reliable detection/classification of SARS-CoV-2 S protein structures. Three classifiers are used and their performance is compared by using six evaluation metrics. Obtained results show that two classifiers for text data (Multinomial Naive Bayes and Stochastic Gradient Descent) performed better and achieved high accuracy on the dataset that contains AAA of protein structures compared to the datasets for AA and ASSE, respectively. M. Saqib Nawaz, Philippe Fournier-Viger, Yu-Lin He |
BIBM | 1 |
| 2022 | Metaheuristic Algorithms for Proof Searching in HOL4abstractUser guided proof development in interactive theorem proving is a manual and time consuming activity.For automating proof searching and optimization in a higher-order logic proof assistant, we provide two metaheuristic algorithms that are based on Fitness Dependent Optimizer (FDO) and Bat Algorithm (BA).In both metaheuristic algorithms, random proof sequences are first created from a population of frequently occurring proof steps that are discovered using pattern mining techniques.Created proof sequences are then evolved till their fitness matches the fitness of the original (or target) proof sequences.Experiments are performed to investigate the performance of the proposed algorithms on different HOL4 theories.Moreover, the proposed FDO and BA-based proof searching approaches are compared with Simulated Annealing (SA) and Genetic Algorithm (GA)based methods.Results show that BA performs best, followed by FDO and SA for proof finding and optimization in HOL4. M. Saqib Nawaz, Muhammad Zohaib Nawaz, Osman Hasan, Philippe Fournier-Viger |
SEKE | 1 |
| 2022 | MalSPM: Metamorphic malware behavior analysis and classification using sequential pattern mining
M. Saqib Nawaz, Philippe Fournier-Viger, Muhammad Zohaib Nawaz, Guoting Chen, Youxi Wu |
Comput. Secur. | 1 |
| 2021 | Investigating Crossover Operators in Genetic Algorithms for High-Utility Itemset Mining
M. Saqib Nawaz, Philippe Fournier-Viger, Wei Song 0004, Jerry Chun-Wei Lin, Bernd Noack |
ACIIDS | 1 |
| 2021 | COVID-19 Genome Analysis Using Alignment-Free Methods
M. Saqib Nawaz, Philippe Fournier-Viger, Xinzheng Niu, Youxi Wu, Jerry Chun-Wei Lin |
IEA/AIE (1) | 1 |
| 2021 | Using artificial intelligence techniques for COVID-19 genome analysis
M. Saqib Nawaz, Philippe Fournier-Viger, Abbas Shojaee, Hamido Fujita |
Appl. Intell. | 1 |
| 2021 | Proof searching and prediction in HOL4 with evolutionary/heuristic and deep learning techniques
M. Saqib Nawaz, Muhammad Zohaib Nawaz, Osman Hasan, Philippe Fournier-Viger, Meng Sun 0002 |
Appl. Intell. | 1 |
| 2020 | Bibliometric Analysis of Social Media as a Platform for Knowledge ManagementabstractThe purpose of this study is to conduct a bibliometric analysis to examine the most influential journals, institutions, and countries in social media (SM) publications related to knowledge management (KM). Moreover, various research themes in SM KM publications are also explored. VOSviewer was employed to process 234 SM KM publications retrieved from Web of Science (WoS) in the time period 2009-2019. Different methodologies were used according to the nature of bibliometric analysis and explained in each section. Journal of Knowledge Management was the most influential journal in SM KM publications. USA and England ranked first and second respectively, while the Tampere University of Technology was the most productive institute in SM KM research. Four emerged themes indicated an explicit contribution of SM users in KM through big data, knowledge sharing, innovation, Enterprise 2.0, and social capital. This is the first bibliometric study that explores the overall contribution of SM publications in the KM field. Saleha Noor, Yi Guo 0009, Syed Hamad Hassan Shah, M. Saqib Nawaz, Atif Saleem Butt |
Int. J. Knowl. Manag. | 4 |
| 2020 | Research Synthesis and Thematic Analysis of Twitter Through Bibliometric AnalysisabstractIn literature, there is a shortage of comprehensive documents that can provide proper details about Twitter in research community. This study conducted a first descriptive bibliometric analysis to examine the most influential journals, institutions, and countries on Twitter. Similarly, bibliometric mapping analysis is carried out to explore different research themes in Twitter publications. VOSviewer was employed to process the 11,006 Twitter publications retrieved from the Web of Science (WoS) from 2009 to 2018. Obtained results suggest that USA and China received the highest number of publications on Twitter research, while the University of Illinois was the most productive institute. Furthermore, the five major themes have emerged in Twitter publications, and its remarkable role has been found in event detection, sentiment analysis, education, health, politics, and crisis as well as risk management. The authors believe that this study will open new doors for researchers to use online Twitter social networking communities in beauty salons, consulting companies, banks, and airlines. Saleha Noor, Yi Guo 0009, Syed Hamad Hassan Shah, M. Saqib Nawaz, Atif Saleem Butt |
Int. J. Semantic Web Inf. Syst. | 4 |
| 2018 | Reo2PVS: Formal Specification and Verification of Component ConnectorsabstractCompositional coordination models such as Reo provide powerful support for the development of large-scale distributed systems by allowing construction of complex connectors that coordinate behavior among different components.The reliability of such distributed systems highly depends on the correctness of connectors.In this paper, we use the proof assistant PVS for formal modeling, analysis and verification of component connectors.We first present the modeling of primitive channels and the composition operators that are used to combine channels for building complex connectors.Furthermore, we show how to model and analyze connector's behavior in PVS and prove some interesting connector properties.The model reflects the original topological structure of connectors simply and clearly.With the provided approach, different kinds of connector properties can be naturally formalized and proved in PVS. M. Saqib Nawaz, Meng Sun 0002 |
SEKE | 1 |
| 2017 | Finding Healthcare Issues with Search Engine Queries and Social Network DataabstractSearch engines and social networks are two entirely different data sources that can provide valuable information about Influenza. While search engine hosts can deliver popular queries (or terms) used for searching the Influenza related information, the social networks contain useful links of information sources that people have found valuable. The authors hypothesize that such data sources can provide vital first-hand information. In this article, they have proposed a methodology for detecting the information sources from social networks, particularly Twitter. The data filtering and source finding tasks are posed as classification tasks. Search engine queries are used for extracting related dataset. Results have shown that propose approach can be beneficial for extracting useful information regarding side effects, medications and to track geographical location of epidemics affected area. Muhammad Ikram Ullah Lali, Raza Ul-Mustafa, Kashif Saleem, M. Saqib Nawaz, Tehseen Zia, Basit Shahzad |
Int. J. Semantic Web Inf. Syst. | 4 |