EDBT 2026 Demo / reviewers in the wild / expert
Hilal Tayara
dblp:192/6591
· DBLP profile ↗
17ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0001-5678-3479ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DDINet: A multi-task neural network for accurate drug-drug interaction prediction and effect analysis
Sabir Ali, Waleed Alam, Kil To Chong 0001, Hilal Tayara |
Knowl. Based Syst. | 4 |
| 2026 | Unveiling Viral Escape Mechanisms With Machine Learning: A Transformative Approach to Mutation Analysis for SARS-CoV-2 and BeyondabstractPersistent viruses like Influenza, HIV, and Coronavirus exemplify the challenge of viral escape, significantly hindering the development of long-lasting vaccines and effective treatments. This study leverages a Long Short-Term Memory (LSTM) based deep learning architecture to analyze an extensive dataset of over 3.1 million unique viral spike protein sequences, with SARS-CoV-2 serving as the primary example. Our model, Escape Elite Network (EEN) outperforms existing methods in detecting escape mutations across diverse datasets. In computational (Validation) and wet lab (Baum and Greaney) datasets, EEN achieved AUC scores of 0.949, 0.868, and 0.762, respectively, each with p-values less than $1\times 10^{-5}$, demonstrating high statistical significance. Specifically, EEN significantly outperformed the competitor models NPVE and SEN across all datasets. For the Greaney dataset, EEN achieved a 19.2% improvement over NPVE and a 5.7% improvement over SEN. Similarly, for the Baum dataset, EEN showed a 1.8% improvement over NPVE and a 12.8% improvement over SEN. This tool proactively predicts viral escape mutations, guiding the development of more effective vaccines and therapeutics. It identifies high-risk mutations before experimental validation, offering advantages in combating SARS-CoV-2 and potentially other viruses like HIV, Influenza, and African Swine Fever, with broad applications in virology and epidemiology. Prem Singh Bist, Kil To Chong 0001, Hilal Tayara |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | CpGFuse: a holistic approach for accurate identification of methylation states of DNA CpG sitesabstractAnomalous DNA methylation has wide-ranging implications, spanning from neurological disorders to cancer and cardiovascular complications. Current methods for single-cell DNA methylation analysis face limitations in coverage, leading to information loss and hampering our understanding of disease associations. The primary goal of this study is the imputation of CpG site methylation states in a given cell by leveraging the CpG states of other cells of the same type. To address this, we introduce CpGFuse, a novel methodology that combines information from diverse genomic features. Leveraging two benchmark datasets, we employed a careful preprocessing approach and conducted a comprehensive ablation study to assess the individual and collective contributions of DNA sequence, intercellular, and intracellular features. Our proposed model, CpGFuse, employs a convolutional neural network with an attention mechanism, surpassing existing models across HCCs and HepG2 datasets. The results highlight the effectiveness of our approach in enhancing accuracy and providing a robust tool for CpG site prediction in genomics. CpGFuse's success underscores the importance of integrating multiple genomic features for accurate identification of methylation states of CpG site. Sehi Park, Kil To Chong 0001, Hilal Tayara |
Briefings Bioinform. | 3 |
| 2025 | Ensemble insights: Unlocking the recombination losses in perovskite solar cells using stacked classifier
Basir Akbar, Kil To Chong 0001, Hilal Tayara |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | iAnOxPep: A Machine Learning Model for the Identification of Anti-Oxidative Peptides Using Ensemble LearningabstractDue to their safety, high activity, and plentiful sources, antioxidant peptides, particularly those produced from food, are thought to be prospective competitors to synthetic antioxidants in the fight against free radical-mediated illnesses. The lengthy and laborious trial-and-error method for identifying antioxidative peptides (AOP) has raised interest in creating computational-based methods. There exist two state-of-the-art AOP predictors; however, the restriction on peptide sequence length makes them inviable. By overcoming the aforementioned problem, a novel predictor might be useful in the context of AOP prediction. The method has been trained, tested, and evaluated on two datasets: a balanced one and an unbalanced one. We used seven different descriptors and five machine-learning (ML) classifiers to construct 35 baseline models. Five ML classifiers were further trained to create five meta-models using the combined output of 35 baseline models. Finally, these five meta-models were aggregated together through ensemble learning to create a robust predictive model named iAnOxPep. On both datasets, our proposed model demonstrated good prediction performance when compared to baseline models and meta-models, demonstrating the superiority of our approach in the identification of AOPs. For the purpose of screening and identifying possible AOPs, we anticipate that the iAnOxPep method will be an invaluable tool. Mir Tanveerul Hassan, Hilal Tayara, Kil To Chong 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | Enhanced prediction of hemolytic activity in antimicrobial peptides using deep learning-based sequence analysisabstractAntimicrobial peptides (AMPs) are a promising class of antimicrobial drugs due to their broad-spectrum activity against microorganisms. However, their clinical application is limited by their potential to cause hemolysis, the destruction of red blood cells. To address this issue, we propose a deep learning model based on convolutional neural networks (CNNs) for predicting the hemolytic activity of AMPs. Peptide sequences are represented using one-hot encoding, and the CNN architecture consists of multiple convolutional and fully connected layers. The model was trained on six different datasets: HemoPI-1, HemoPI-2, HemoPI-3, RNN-Hem, Hlppredfuse, and AMP-Combined, achieving Matthew's correlation coefficients of 0.9274, 0.5614, 0.6051, 0.6142, 0.8799, and 0.7484, respectively. Our model outperforms previously reported methods and can facilitate the development of novel AMPs with reduced hemolytic activity, which is crucial for their therapeutic use in treating bacterial infections. Ibrahim Abdelbaky, Mohamed Elhakeem, Hilal Tayara, E. M. Badr, Mustafa Abdul Salam |
BMC Bioinform. | 3 |
| 2024 | Generative AI in the Advancement of Viral Therapeutics for Predicting and Targeting Immune-Evasive SARS-CoV-2 MutationsabstractThe emergence of immune-evasive mutations in the SARS-CoV-2 spike protein is consistently challenging existing vaccines and therapies, making precise prediction of their escape potential a critical imperative. Artificial Intelligence(AI) holds great promise for deciphering the intricate language of protein. Here, we employed a Generative Adversarial Network to decipher the hidden escape pathways within the spike protein by generating spikes that closely resemble natural ones. Through comprehensive analysis, we demonstrated that generated sequences capture natural escape characteristics. Moreover, incorporating these sequences into an AI-based escape prediction model significantly enhanced its performance, achieving a 7% increase in detecting natural escape mutations on the experimentally validated Greaney dataset. Similar improvements were observed on other datasets, demonstrating the model's generalizability. Precisely predicting immune-evasive spikes not only enables the design of strategically targeted therapies but also has the potential to expedite future viral therapeutics. This breakthrough carries profound implications for shaping a more resilient future against viral threats. Prem Singh Bist, Hilal Tayara, Kil To Chong 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Sars-escape network for escape prediction of SARS-COV-2abstractMOTIVATION: Viruses have coevolved with their hosts for over millions of years and learned to escape the host's immune system. Although not all genetic changes in viruses are deleterious, some significant mutations lead to the escape of neutralizing antibodies and weaken the immune system, which increases infectivity and transmissibility, thereby impeding the development of antiviral drugs or vaccines. Accurate and reliable identification of viral escape mutational sequences could be a good indicator for therapeutic design. We developed a computational model that recognizes significant mutational sequences based on escape feature identification using natural language processing along with prior knowledge of experimentally validated escape mutants. RESULTS: Our machine learning-based computational approach can recognize the significant spike protein sequences of severe acute respiratory syndrome coronavirus 2 using sequence data alone. This modelling approach can be applied to other viruses, such as influenza, monkeypox and HIV using knowledge of escape mutants and relevant protein sequence datasets. AVAILABILITY: Complete source code and pre-trained models for escape prediction of severe acute respiratory syndrome coronavirus 2 protein sequences are available on Github at https://github.com/PremSinghBist/Sars-CoV-2-Escape-Model.git. The dataset is deposited to Zenodo at: doi: 10.5281/zenodo.7142638. The Python scripts are easy to run and customize as needed. CONTACT: [email protected]. Prem Singh Bist, Hilal Tayara, Kil To Chong 0001 |
Briefings Bioinform. | 2 |
| 2023 | ORI-Explorer: a unified cell-specific tool for origin of replication sites prediction by feature fusionabstractMOTIVATION: The origins of replication sites (ORIs) are precise regions inside the DNA sequence where the replication process begins. These locations are critical for preserving the genome's integrity during cell division and guaranteeing the faithful transfer of genetic data from generation to generation. The advent of experimental techniques has aided in the discovery of ORIs in many species. Experimentation, on the other hand, is often more time-consuming and pricey than computational approaches, and it necessitates specific equipment and knowledge. Recently, ORI sites have been predicted using computational techniques like motif-based searches and artificial intelligence algorithms based on sequence characteristics and chromatin states. RESULTS: In this article, we developed ORI-Explorer, a unique artificial intelligence-based technique that combines multiple feature engineering techniques to train CatBoost Classifier for recognizing ORIs from four distinct eukaryotic species. ORI-Explorer was created by utilizing a unique combination of three traditional feature-encoding techniques and a feature set obtained from a deep-learning neural network model. The ORI-Explorer has significantly outperformed current predictors on the testing dataset. Furthermore, by employing the sophisticated SHapley Additive exPlanation method, we give crucial insights that aid in comprehending model success, highlighting the most relevant features vital for forecasting cell-specific ORIs. ORI-Explorer is also intended to aid community-wide attempts in discovering potential ORIs and developing innovative verifiable biological hypotheses. AVAILABILITY AND IMPLEMENTATION: The used datasets along with the source code are made available through https://github.com/Z-Abbas/ORI-Explorer and https://zenodo.org/record/8358679. Zeeshan Abbas, Mobeen Ur Rehman, Hilal Tayara, Kil To Chong 0001 |
Bioinform. | 3 |
| 2023 | iCpG-Pos: an accurate computational approach for identification of CpG sites using positional features on single-cell whole genome sequence dataabstractMOTIVATION: The investigation of DNA methylation can shed light on the processes underlying human well-being and help determine overall human health. However, insufficient coverage makes it challenging to implement single-stranded DNA methylation sequencing technologies, highlighting the need for an efficient prediction model. Models are required to create an understanding of the underlying biological systems and to project single-cell (methylated) data accurately. RESULTS: In this study, we developed positional features for predicting CpG sites. Positional characteristics of the sequence are derived using data from CpG regions and the separation between nearby CpG sites. Multiple optimized classifiers and different ensemble learning approaches are evaluated. The OPTUNA framework is used to optimize the algorithms. The CatBoost algorithm followed by the stacking algorithm outperformed existing DNA methylation identifiers. AVAILABILITY AND IMPLEMENTATION: The data and methodologies used in this study are openly accessible to the research community. Researchers can access the positional features and algorithms used for predicting CpG site methylation patterns. To achieve superior performance, we employed the CatBoost algorithm followed by the stacking algorithm, which outperformed existing DNA methylation identifiers. The proposed iCpG-Pos approach utilizes only positional features, resulting in a substantial reduction in computational complexity compared to other known approaches for detecting CpG site methylation patterns. In conclusion, our study introduces a novel approach, iCpG-Pos, for predicting CpG site methylation patterns. By focusing on positional features, our model offers both accuracy and efficiency, making it a promising tool for advancing DNA methylation research and its applications in human health and well-being. Sehi Park, Mobeen Ur Rehman, Farman Ullah 0001, Hilal Tayara, Kil To Chong 0001 |
Bioinform. | 4 |
| 2023 | DL-m6A: Identification of N6-Methyladenosine Sites in Mammals Using Deep Learning Based on Different Encoding SchemesabstractN6-methyladenosine (m6A) is a common post-transcriptional alteration that plays a critical function in a variety of biological processes. Although experimental approaches for identifying m6A sites have been developed and deployed, they are currently expensive for transcriptome-wide m6A identification. Some computational strategies for identifying m6A sites have been presented as an effective complement to the experimental procedure. However, their performance still requires improvement. In this study, we have proposed a novel tool called DL-m6A for the identification of m6A sites in mammals using deep learning based on different encoding schemes. The proposed tool uses three encoding schemes which give the required contextual feature representation to the input RNA sequence. Later these contextual feature vectors individually go through several neural network layers for shallow feature extraction after which they are concatenated to a single feature vector. The concatenated feature map is then used by several other layers to extract the deep features so that the insight features of the sequence can be used for the prediction of m6A sites. The proposed tool is firstly evaluated on the tissue-specific dataset and later on a full transcript dataset. To ensure the generalizability of the tool we assessed the proposed model by training it on a full transcript dataset and test on the tissue-specific dataset. The achieved results by the proposed model have outperformed the existing tools. The results demonstrate that the proposed tool can be of great use for the biology experts and therefore a freely accessible web-server is created which can be accessed at: http://nsclbio.jbnu.ac.kr/tools/DL-m6A/. Mobeen Ur Rehman, Hilal Tayara, Kil To Chong 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | DeepCap-Kcr: accurate identification and investigation of protein lysine crotonylation sites based on capsule networkabstractLysine crotonylation (Kcr) is a posttranslational modification widely detected in histone and nonhistone proteins. It plays a vital role in human disease progression and various cellular processes, including cell cycle, cell organization, chromatin remodeling and a key mechanism to increase proteomic diversity. Thus, accurate information on such sites is beneficial for both drug development and basic research. Existing computational methods can be improved to more effectively identify Kcr sites in proteins. In this study, we proposed a deep learning model, DeepCap-Kcr, a capsule network (CapsNet) based on a convolutional neural network (CNN) and long short-term memory (LSTM) for robust prediction of Kcr sites on histone and nonhistone proteins (mammals). The proposed model outperformed the existing CNN architecture Deep-Kcr and other well-established tools in most cases and provided promising outcomes for practical use; in particular, the proposed model characterized the internal hierarchical representation as well as the important features from multiple levels of abstraction automatically learned from a small number of samples. The trained model was well generalized in other species (papaya). Moreover, we showed the features and properties generated by the internal capsule layer that can explore the internal data distribution related to biological significance (as a motif detector). The source code and data are freely available at https://github.com/Jhabindra-bioinfo/DeepCap-Kcr. Jhabindra Khanal, Hilal Tayara, Quan Zou 0001, Kil To Chong 0001 |
Briefings Bioinform. | 2 |
| 2022 | i6mA-Caps: a CapsuleNet-based framework for identifying DNA N6-methyladenine sitesabstractMOTIVATION: DNA N6-methyladenine (6mA) has been demonstrated to have an essential function in epigenetic modification in eukaryotic species in recent research. 6mA has been linked to various biological processes. It's critical to create a new algorithm that can rapidly and reliably detect 6mA sites in genomes to investigate their biological roles. The identification of 6mA marks in the genome is the first and most important step in understanding the underlying molecular processes, as well as their regulatory functions. RESULTS: In this article, we proposed a novel computational tool called i6mA-Caps which CapsuleNet based a framework for identifying the DNA N6-methyladenine sites. The proposed framework uses a single encoding scheme for numerical representation of the DNA sequence. The numerical data is then used by the set of convolution layers to extract low-level features. These features are then used by the capsule network to extract intermediate-level and later high-level features to classify the 6mA sites. The proposed network is evaluated on three datasets belonging to three genomes which are Rosaceae, Rice and Arabidopsis thaliana. Proposed method has attained an accuracy of 96.71%, 94% and 86.83% for independent Rosaceae dataset, Rice dataset and A.thaliana dataset respectively. The proposed framework has exhibited improved results when compared with the existing top-of-the-line methods. AVAILABILITY AND IMPLEMENTATION: A user-friendly web-server is made available for the biological experts which can be accessed at: http://nsclbio.jbnu.ac.kr/tools/i6mA-Caps/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mobeen Ur Rehman, Hilal Tayara, Quan Zou 0001, Kil To Chong 0001 |
Bioinform. | 2 |
| 2022 | ZayyuNet - A Unified Deep Learning Model for the Identification of Epigenetic Modifications Using Raw Genomic SequencesabstractEpigenetic modifications have a vital role in gene expression and are linked to cellular processes such as differentiation, development, and tumorigenesis. Thus, the availability of reliable and accurate methods for identifying and defining these changes facilitates greater insights into the regulatory mechanisms that rely on epigenetic modifications. The current experimental methods provide a genome-wide identification of epigenetic modifications; however, they are expensive and time-consuming. To date, several machine learning methods have been proposed for identifying modifications such as DNA N6-Methyladenine (6mA), RNA N6-Methyladenosine (m6A), DNA N4-methylcytosine (4mC), and RNA pseudouridine ( Ψ). However, these methods are task-specific computational tools and require different encoding representations of DNA/RNA sequences. In this study, we propose a unified deep learning model, called ZayyuNet, for the identification of various epigenetic modifications. The proposed model is based on an architecture called, SpinalNet, inspired by the human somatosensory system that can efficiently receive large inputs and achieve better performance. The proposed model has been evaluated on various epigenetic modifications such as 6mA, m6A, 4mC, and Ψ and the results achieved outperform current state-of-the-art models. A user-friendly web server has been built and made freely available at http://nsclbio.jbnu.ac.kr/tools/ZayyuNet/. Zeeshan Abbas, Hilal Tayara, Kil To Chong 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Identification of Functional piRNAs Using a Convolutional Neural NetworkabstractPiwi-interacting RNAs (piRNAs) are a distinct sub-class of small non-coding RNAs that are mainly responsible for germline stem cell maintenance, gene stability, and maintaining genome integrity by repression of transposable elements. piRNAs are also expressed aberrantly and associated with various kinds of cancers. To identify piRNAs and their role in guiding target mRNA deadenylation, the currently available computational methods require urgent improvements in performance. To facilitate this, we propose a robust predictor based on a lightweight and simplified deep learning architecture using a convolutional neural network (CNN) to extract significant features from raw RNA sequences without the need for more customized features. The proposed model's performance is comprehensively evaluated using k-fold cross-validation on a benchmark dataset. The proposed model significantly outperforms existing computational methods in the prediction of piRNAs and their role in target mRNA deadenylation. In addition, a user-friendly and publicly-accessible web server is available at http://nsclbio.jbnu.ac.kr/tools/2S-piRCNN/. Syed Danish Ali, Waleed Alam, Hilal Tayara, Kil To Chong 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Recent omics-based computational methods for COVID-19 drug discovery and repurposingabstractThe coronavirus disease 2019 (COVID-19) pandemic, caused by the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), is the main reason for the increasing number of deaths worldwide. Although strict quarantine measures were followed in many countries, the disease situation is still intractable. Thus, it is needed to utilize all possible means to confront this pandemic. Therefore, researchers are in a race against the time to produce potential treatments to cure or reduce the increasing infections of COVID-19. Computational methods are widely proving rapid successes in biological related problems, including diagnosis and treatment of diseases. Many efforts in recent months utilized Artificial Intelligence (AI) techniques in the context of fighting the spread of COVID-19. Providing periodic reviews and discussions of recent efforts saves the time of researchers and helps to link their endeavors for a faster and efficient confrontation of the pandemic. In this review, we discuss the recent promising studies that used Omics-based data and utilized AI algorithms and other computational tools to achieve this goal. We review the established datasets and the developed methods that were basically directed to new or repurposed drugs, vaccinations and diagnosis. The tools and methods varied depending on the level of details in the available information such as structures, sequences or metabolic data. Hilal Tayara, Ibrahim Abdelbaky, Kil To Chong 0001 |
Briefings Bioinform. | 1 |
| 2021 | Improved Predicting of The Sequence Specificities of RNA Binding Proteins by Deep LearningabstractRNA-binding proteins (RBPs) have a significant role in various regulatory tasks. However, the mechanism by which RBPs identify the subsequence target RNAs is still not clear. In recent years, several machine and deep learning-based computational models have been proposed for understanding the binding preferences of RBPs. These methods required integrating multiple features with raw RNA sequences such as secondary structure and their performances can be further improved. In this paper, we propose an efficient and simple convolution neural network, RBPCNN, that relies on the combination of the raw RNA sequence and evolutionary information. We show that conservation scores (evolutionary information) for the RNA sequences can significantly improve the overall performance of the proposed predictor. In addition, the automatic extraction of the binding sequence motifs can enhance our understanding of the binding specificities of RBPs. The experimental results show that RBPCNN outperforms significantly the current state-of-the-art methods. More specifically, the average area under the receiver operator curve was improved by 2.67 percent and the mean average precision was improved by 8.03 percent. The datasets and results can be downloaded from https://home.jbnu.ac.kr/NSCL/RBPCNN.htm. Hilal Tayara, Kil To Chong 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |