VLDB 2026 Research / reviewers in the wild / expert
Muhammad Arif 0012
dblp:67/3312-12
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-3950-6618ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An artificial intelligence-based approach for identifying the proteins regulating liquid-liquid phase separationabstractLiquid-liquid phase separation (LLPS) is a biomolecular process that underpins the formation of membrane-less organelles within living cells. This phenomenon, along with the resulting condensate bodies, is increasingly recognized for its critical roles in various biological processes, such as ribonucleic acid (RNA) metabolism, chromatin rearrangement, and signal transduction. Notably, regulator proteins play a central role in the process of LLPS. They are essential for the formation, stabilization, and maintenance of the dynamic properties of LLPS, ensuring an appropriate phase separation response to cellular signals. Targeting these regulator proteins is the key to manipulating LLPS for applications in biotechnology, materials science, and medicine, including biomaterials, drug delivery, diagnostics, and synthetic biology. Given their importance, this study focused on an artificial intelligence-based approach to identify regulator proteins in LLPS. We constructed a dataset of 913 positive and 6584 negative protein sequences, and divided it into eight balanced training datasets and a test dataset. Semantic information from protein sequences was extracted using the ESM2_t36 pretrained protein language model, followed by training a multilayer perceptron classifier. The model achieved 0.78 accuracy on the test dataset, outperforming traditional sequence-based methods, one-hot encoding, and other pretrained embedding methods. SHapley Additive exPlanations (SHAP)-based interpretation revealed key biophysical patterns enriched in regulator proteins, including higher levels of charged and disordered residues. Our results show that deep contextual protein representations combined with neural network-based classifiers can accurately identify LLPS regulator proteins. This tool offers new opportunities for understanding condensate biology and designing synthetic phase-separating systems. All data and code are available at: https://github.com/bioplusAI/LLPS_regulators_pred. Zahoor Ahmed, Kiran Shahzadi, Yu-Qing Jiang, Yan-Ting Jin, Muhammad Arif 0012 |
Briefings Bioinform. | 6 |
| 2025 | PCBert-Kla: an efficient prediction method for lysine lactylation sites based on ProtBert and fusion of physicochemical featuresabstractProtein post-translational modifications (PTMs) play a critical role in regulating protein functionality and structural diversity. Among them, lysine lactylation (Kla), a newly identified PTM, is involved in energy metabolism, cellular reprogramming, and the progression of various diseases. In this study, we propose PCBert-Kla, a feature-fusion deep learning model based on ProtBert. This model leverages ProtBert to extract deep features from protein sequences, effectively capturing global and local contextual information. It integrated various physicochemical properties, including molecular weight, isoelectric point, amino acid composition, secondary structure content, hydrophobicity, and net charge. An attention mechanism in the fully connected layers enabled the model to select features automatically. PCBert-Kla exhibited exceptional accuracy and reliability in Kla site identification and demonstrated excellent generalization capability to outperform the existing models. In addition, we further enhanced the interpretability of the PCBert-Kla model by incorporating average attention maps. This model provided powerful tools for studying the functions of Kla and elucidating the mechanisms of related diseases, which can advance biomedical research and drug development. We also developed a free web service, available at http://pcbert-kla.lin-group.cn/, to provide users with easy access and usage. Hong-Qi Zhang, Yi-Xuan Qi, Huma Fida, Hao-Jiang Zhang, Muhammad Arif 0012, Pei-Yu Zhao, Tanvir Alam, Ye-Chen Qi, Xiao-Long Yu, Ke-Jun Deng |
Briefings Bioinform. | 5 |
| 2024 | DPI_CDF: druggable protein identifier using cascade deep forestabstractBACKGROUND: Drug targets in living beings perform pivotal roles in the discovery of potential drugs. Conventional wet-lab characterization of drug targets is although accurate but generally expensive, slow, and resource intensive. Therefore, computational methods are highly desirable as an alternative to expedite the large-scale identification of druggable proteins (DPs); however, the existing in silico predictor's performance is still not satisfactory. METHODS: In this study, we developed a novel deep learning-based model DPI_CDF for predicting DPs based on protein sequence only. DPI_CDF utilizes evolutionary-based (i.e., histograms of oriented gradients for position-specific scoring matrix), physiochemical-based (i.e., component protein sequence representation), and compositional-based (i.e., normalized qualitative characteristic) properties of protein sequence to generate features. Then a hierarchical deep forest model fuses these three encoding schemes to build the proposed model DPI_CDF. RESULTS: The empirical outcomes on 10-fold cross-validation demonstrate that the proposed model achieved 99.13 % accuracy and 0.982 of Matthew's-correlation-coefficient (MCC) on the training dataset. The generalization power of the trained model is further examined on an independent dataset and achieved 95.01% of maximum accuracy and 0.900 MCC. When compared to current state-of-the-art methods, DPI_CDF improves in terms of accuracy by 4.27% and 4.31% on training and testing datasets, respectively. We believe, DPI_CDF will support the research community to identify druggable proteins and escalate the drug discovery process. AVAILABILITY: The benchmark datasets and source codes are available in GitHub: http://github.com/Muhammad-Arif-NUST/DPI_CDF . Muhammad Arif 0012, Ge Fang, Ali Ghulam, Saleh Musleh, Tanvir Alam |
BMC Bioinform. | 1 |
| 2023 | VPatho: a deep learning-based two-stage approach for accurate prediction of gain-of-function and loss-of-function variantsabstractDetermining the pathogenicity and functional impact (i.e. gain-of-function; GOF or loss-of-function; LOF) of a variant is vital for unraveling the genetic level mechanisms of human diseases. To provide a 'one-stop' framework for the accurate identification of pathogenicity and functional impact of variants, we developed a two-stage deep-learning-based computational solution, termed VPatho, which was trained using a total of 9619 pathogenic GOF/LOF and 138 026 neutral variants curated from various databases. A total number of 138 variant-level, 262 protein-level and 103 genome-level features were extracted for constructing the models of VPatho. The development of VPatho consists of two stages: (i) a random under-sampling multi-scale residual neural network (ResNet) with a newly defined weighted-loss function (RUS-Wg-MSResNet) was proposed to predict variants' pathogenicity on the gnomAD_NV + GOF/LOF dataset; and (ii) an XGBOD model was constructed to predict the functional impact of the given variants. Benchmarking experiments demonstrated that RUS-Wg-MSResNet achieved the highest prediction performance with the weights calculated based on the ratios of neutral versus pathogenic variants. Independent tests showed that both RUS-Wg-MSResNet and XGBOD achieved outstanding performance. Moreover, assessed using variants from the CAGI6 competition, RUS-Wg-MSResNet achieved superior performance compared to state-of-the-art predictors. The fine-trained XGBOD models were further used to blind test the whole LOF data downloaded from gnomAD and accordingly, we identified 31 nonLOF variants that were previously labeled as LOF/uncertain variants. As an implementation of the developed approach, a webserver of VPatho is made publicly available at http://csbio.njust.edu.cn/bioinf/vpatho/ to facilitate community-wide efforts for profiling and prioritizing the query variants with respect to their pathogenicity and functional impact. Fang Ge, Chen Li 0021, Muhammad Arif 0012, Fuyi Li, Maha A. Thafar, Zihao Yan, Apilak Worachartcheewan, Jiangning Song, Dongjun Yu |
Briefings Bioinform. | 4 |
| 2023 | Lung-EffNet: Lung cancer classification using EfficientNet from CT-scan imagesabstractLung cancer (LC) remains a leading cause of death worldwide. Early diagnosis is critical to protect innocent human lives. Computed tomography (CT) scans are one of the primary imaging modalities for lung cancer diagnosis. However, manual CT scan analysis is time-consuming and prone to errors/not accurate. Considering these shortcomings, computational methods especially machine learning and deep learning algorithms are leveraged as an alternative to accelerate the accurate detection of CT scans as cancerous, and non-cancerous. In the present article, we proposed a novel transfer learning-based predictor called, Lung-EffNet for lung cancer classification. Lung-EffNet is built based on the architecture of EfficientNet and further modified by adding top layers in the classification head of the model. Lung-EffNet is evaluated by utilizing five variants of EfficientNet i.e., B0–B4. The experiments are conducted on the benchmark dataset “IQ-OTH/NCCD” for lung cancer patients grouped as benign, malignant, or normal based on the presence or absence of lung cancer. The class imbalance issue was handled through multiple data augmentation methods to overcome the biases. The developed model Lung-EffNet attained 99.10% of accuracy and a score of 0.97 to 0.99 of ROC on the test set. We compared the efficacy of the proposed fine-tuned pre-trained EfficientNet with other pre-trained CNN architectures. The predicted outcomes demonstrate that EfficientNetB1 based Lung-EffNet outperforms other CNNs in terms of both accuracy and efficiency. Moreover, it is faster and requires fewer parameters to train than other CNN based models, making it a good choice for large-scale deployment in clinical settings and a promising tool for automated lung cancer diagnosis from CT scan images. Rehan Raza, Fatima Zulfiqar, Muhammad Owais Khan, Muhammad Arif 0012, Atif Alvi, Muhammad Aksam Iftikhar, Tanvir Alam |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | Prediction of disease-associated nsSNPs by integrating multi-scale ResNet models with deep feature fusionabstractMore than 6000 human diseases have been recorded to be caused by non-synonymous single nucleotide polymorphisms (nsSNPs). Rapid and accurate prediction of pathogenic nsSNPs can improve our understanding of the principle and design of new drugs, which remains an unresolved challenge. In the present work, a new computational approach, termed MSRes-MutP, is proposed based on ResNet blocks with multi-scale kernel size to predict disease-associated nsSNPs. By feeding the serial concatenation of the extracted four types of features, the performance of MSRes-MutP does not obviously improve. To address this, a second model FFMSRes-MutP is developed, which utilizes deep feature fusion strategy and multi-scale 2D-ResNet and 1D-ResNet blocks to extract relevant two-dimensional features and physicochemical properties. FFMSRes-MutP with the concatenated features achieves a better performance than that with individual features. The performance of FFMSRes-MutP is benchmarked on five different datasets. It achieves the Matthew's correlation coefficient (MCC) of 0.593 and 0.618 on the PredictSNP and MMP datasets, which are 0.101 and 0.210 higher than that of the existing best method PredictSNP1. When tested on the HumDiv and HumVar datasets, it achieves MCC of 0.9605 and 0.9507, and area under curve (AUC) of 0.9796 and 0.9748, which are 0.1747 and 0.2669, 0.0853 and 0.1335, respectively, higher than the existing best methods PolyPhen-2 and FATHMM (weighted). In addition, on blind test using a third-party dataset, FFMSRes-MutP performs as the second-best predictor (with MCC and AUC of 0.5215 and 0.7633, respectively), when compared with the other four predictors. Extensive benchmarking experiments demonstrate that FFMSRes-MutP achieves effective feature fusion and can be explored as a useful approach for predicting disease-associated nsSNPs. The webserver is freely available at http://csbio.njust.edu.cn/bioinf/ffmsresmutp/ for academic use. Fang Ge, Ying Zhang 0053, Jian Xu 0009, Muhammad Arif 0012, Jiangning Song, Dongjun Yu |
Briefings Bioinform. | 4 |
| 2022 | DeepCPPred: A Deep Learning Framework for the Discrimination of Cell-Penetrating Peptides and Their Uptake EfficienciesabstractCell-penetrating peptides (CPPs) are special peptides capable of carrying a variety of bioactive molecules, such as genetic materials, short interfering RNAs and nanoparticles, into cells. Recently, research on CPP has gained substantial interest from researchers, and the biological mechanisms of CPPS have been assessed in the context of safe drug delivery agents and therapeutic applications. Correct identification and synthesis of CPPs using traditional biochemical methods is an extremely slow, expensive and laborious task particularly due to the large volume of unannotated peptide sequences accumulating in the World Bank repository. Hence, a powerful bioinformatics predictor that rapidly identifies CPPs with a high recognition rate is urgently needed. To date, numerous computational methods have been developed for CPP prediction. However, the available machine-learning (ML) tools are unable to distinguish both the CPPs and their uptake efficiencies. This study aimed to develop a two-layer deep learning framework named DeepCPPred to identify both CPPs in the first phase and peptide uptake efficiency in the second phase. The DeepCPPred predictor first uses four types of descriptors that cover evolutionary, energy estimation, reduced sequence and amino-acid contact information. Then, the extracted features are optimized through the elastic net algorithm and fed into a cascade deep forest algorithm to build the final CPP model. The proposed method achieved 99.45 percent overall accuracy with the CPP924 benchmark dataset in the first layer and 95.43 percent accuracy in the second layer with the CPPSite3 dataset using a 5-fold cross-validation test. Thus, our proposed bioinformatics tool surpassed all the existing state-of-the-art sequence-based CPP approaches. Muhammad Arif 0012, Muhammad Kabir, Abid Khan, Fang Ge, Adel Khelifi, Dongjun Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |