VLDB 2026 Research / reviewers in the wild / expert
Kamal Al-Nasr
dblp:99/7351
· DBLP profile ↗
15ranked-venue papers
8as first author
5since 2021 · last 2024
0000-0001-8459-9070ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 8 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Deep Learning Approach to Identify Protein's Secondary Structure Elements
Kamal Al-Nasr, Richard Mu, Mohammed Alamri |
ISBRA (1) | 2 |
| 2022 | LPTD: a novel linear programming-based topology determination method for cryo-EM mapsabstractSUMMARY: Topology determination is one of the most important intermediate steps toward building the atomic structure of proteins from their medium-resolution cryo-electron microscopy (cryo-EM) map. The main goal in the topology determination is to identify correct matches (i.e. assignment and direction) between secondary structure elements (SSEs) (α-helices and β-sheets) detected in a protein sequence and cryo-EM density map. Despite many recent advances in molecular biology technologies, the problem remains a challenging issue. To overcome the problem, this article proposes a linear programming-based topology determination (LPTD) method to solve the secondary structure topology problem in three-dimensional geometrical space. Through modeling of the protein's sequence with the aid of extracting highly reliable features and a distance-based scoring function, the secondary structure matching problem is transformed into a complete weighted bipartite graph matching problem. Subsequently, an algorithm based on linear programming is developed as a decision-making strategy to extract the true topology (native topology) between all possible topologies. The proposed automatic framework is verified using 12 experimental and 15 simulated α-β proteins. Results demonstrate that LPTD is highly efficient and extremely fast in such a way that for 77% of cases in the dataset, the native topology has been detected in the first rank topology in <2 s. Besides, this method is able to successfully handle large complex proteins with as many as 65 SSEs. Such a large number of SSEs have never been solved with current tools/methods. AVAILABILITY AND IMPLEMENTATION: The LPTD package (source code and data) is publicly available at https://github.com/B-Behkamal/LPTD. Moreover, two test samples as well as the instruction of utilizing the graphical user interface have been provided in the shared readme file. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bahareh Behkamal, Mahmoud Naghibzadeh, Andrea Pagnani, Mohammad Reza Saberi, Kamal Al-Nasr |
Bioinform. | 5 |
| 2021 | Assignment of Protein Secondary Structure Elements from Cα Backbone Trace: An Ensemble of Machine Learning ApproachesabstractSecondary structure elements in protein molecules refer to local sub-conformational regions stabilized by hydrogen bonding. Assigning Secondary Structure Elements is crucial in protein structure determination and function analysis. This work represents a recast of a previously developed classifier using ensemble of machine learning models. In this paper, we introduce new geometrical features to improve the accuracy, reduce training data set and process, and we develop and apply a post-processing step. The classifier is trained with 150K amino acids. We tested our classifier on a set of 20 protein structures and compared with previously developed classifier. The information from Protein Data Bank was used as a reference. The comparison shows that new method can produce assignments that are more aligned with PDB at 95.31% accuracy after applying a simple postprocessing step compared to 92.75% for the previous classifier. Kamal Al-Nasr, Ali Sekmen |
BIBM | 1 |
| 2021 | Deep Learning for Assignment of Protein Secondary Structure Elements from Cɑ CoordinatesabstractThis paper presents a Deep Neural network (DNN) system that uses a large set of geometric and categorical features for classification of secondary structure elements (SSEs) in the protein’s trace that consists of $C\alpha$ atoms on the backbone. A systematical approach is implemented for classification of protein SSE problem. This approach consists of two network architecture search (NAS) algorithms for selecting (1) network architecture and layer connectivity, and (2) regularization parameters. Each algorithm uses a different search space and they are used in succession to develop a DNN. The DNN system generates over 93% classification rate on average for multiple test sets without any post processing for amino acid configurations. Kamal Al-Nasr, Ali Sekmen, Bahadir Bilgin, Ahmet Bugra Koku |
BIBM | 1 |
| 2021 | Subspace Modeling for Classification of Protein Secondary Structure Elements from Cα TraceabstractThis paper presents a novel subspace segmentation algorithm that models protein $C\alpha$ traces of secondary structure elements (SSEs) as a union of subspaces. For each $C\alpha$, a set of general geometric features are considered. The algorithm first identifies the most relevant features for each SSE using a new matrix rank estimation technique and combinatorics. This is followed by grouping $C\alpha$ traces in a sliding-window so that each group represents a data point in a high-dimensional ambient space. Then, a lower dimensional subspace is matched for each SSE. When a group of unknown $C\alpha$ traces is presented, the algorithm determines a neighborhood around each $C\alpha$ and then uses two approaches to classify the $C\alpha$. In the first approach, the $C\alpha$ is represented as a data point in the ambient space and its distance to each subspace is calculated. In the second approach, a local subspace is matched to the $C\alpha$, and the separation of this local subspace from each SSE subspace is computed using geodesic distance on the Grassmannian manifold of the subspaces. The minimum point-to-subspace distance and minimum separation of subspaces are used to classify the $C\alpha$. This geometric and mathematical approach has been applied a large protein dataset and generated 85% classification rate without the need to train a large machine learning system. Ali Sekmen, Kamal Al-Nasr |
BIBM | 2 |
| 2020 | Machine Learning Approach to Assign Protein Secondary Structure Elements from Cα TraceabstractSecondary structure elements in protein molecules refer to local sub-conformational regions stabilized by hydrogen bonding. Secondary structure elements can be divided into helical, sheet, or loop. Secondary structure elements bolster the folding and topology of the protein. They are important for modern structural bioinformatics such as protein modeling and functional analysis. Therefore, assigning the types of secondary structures in proteins is crucial. Many methods have been developed to address the problem. Methods can be categorized into two approaches. One approach uses the information about hydrogen bonding and energy while the other approach uses protein trace geometry. If the information of some atoms is missing, the second approach is more feasible. In this paper, we develop a machine learning method that belongs to the second approach to assign secondary structure elements. We develop a 3-state machine learning classifier. The classifier uses protein's Ca information only. The classifier ensembles four (4) machine learning models: Random Forest, Support Vector Machine, Multilayer Perceptron, and eXtreme Gradient Boosting. The classifier is trained with 600K amino acids. We tested our classifier at two different data sets. One data set contains 150K amino acids. The accuracy of our system was 94.6%. In addition, the classifier was tested on a set of 20 protein structures and compared with PCASSO from the same category. The information from Protein Data Bank was used as a reference. The comparison shows that our method can produce assignments that are more aligned with PDB at 93% accuracy while PCASSO achieved 84% accuracy. Mohammad Al Sallal, Wei Chen 0003, Kamal Al-Nasr |
BIBM | 3 |
| 2019 | Supervised Regression Study for Electron Microscopy DataabstractThis study presents a supervised regression model to estimate the growth of Electron Microscopy experimental data for a decade ahead. The study employs the autoregression process model using the best curve-fitting that optimizes the level of confidence. Further, the proposed model retains the smallest normalized estimation error. The developed model was competently utilized to estimate the size of Electron Microscopy (EM) data expected to be released within years 2019-2028. One EM dataset was used to model and predict the annual growth of released 3DEM. Another EM dataset was used to model and predict the annual number of 3-Dimensional EM achieving resolution 10 Aoor better. Indeed, both models used EM data collected in the past 18 years, 2002-2018. The experimental results showed that the best curve-fitting orders to predict both datasets were AR(5) at 96.8% and AR(6) at 85% for the released 3DEM and 3DEM resolutions datasets, respectively. Therefore, the estimation findings disclose an exponential growing performance in the upcoming evolution for both, the released 3DEM and 3DEM resolutions datasets. However, the evolution rate of the released 3DEM confirms a faster exponential growth. Qasem Abu Al-Haija, Kamal Al-Nasr |
BIBM | 2 |
| 2017 | An Effective Computational Method Incorporating Multiple Secondary Structure Predictions in Topology Determination for Cryo-EM ImagesabstractA key idea in de novo modeling of a medium-resolution density image obtained from cryo-electron microscopy is to compute the optimal mapping between the secondary structure traces observed in the density image and those predicted on the protein sequence. When secondary structures are not determined precisely, either from the image or from the amino acid sequence of the protein, the computational problem becomes more complex. We present an efficient method that addresses the secondary structure placement problem in presence of multiple secondary structure predictions and computes the optimal mapping. We tested the method using 12 simulated images from α-proteins and two Cryo-EM images of α-β proteins. We observed that the rank of the true topologies is consistently improved by using multiple secondary structure predictions instead of a single prediction. The results show that the algorithm is robust and works well even when errors/misses in the predicted secondary structures are present in the image or the sequence. The results also show that the algorithm is efficient and is able to handle proteins with as many as 33 helices. Abhishek Biswas, Desh Ranjan, Mohammad Zubair, Stephanie Zeil, Kamal Al-Nasr, Jing He 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2016 | An efficient method for validating protein models using electron microscopy dataabstractCryo-Electron Microscopy is a powerful biophysical technique that is capable of generating 3-dimensional volume images for macromolecular assemblies and machines. De novo protein modeling uses these images to model the biological molecules. In de novo modeling, many candidate structures are generated at intermediate step. The candidates are evaluated conventionally by time-consuming approaches. We introduce an initial version of a geometrical screening method that uses the skeleton of the cryo-EM images to evaluate the candidate structures. A test of ten (10) proteins shows that our method was able to successfully detect good candidates in an efficient way. Kamal Al-Nasr, Bashar Aboona, Abdulrahman Alanazi |
BIBM | 1 |
| 2015 | Deriving Protein Backbone Using Traces Extracted from Density Maps at Medium Resolutions
Kamal Al-Nasr, Jing He 0002 |
ISBRA | 1 |
| 2014 | Solving the Secondary Structure MatchingProblem in Cryo-EM De Novo ModelingUsing a Constrained $K$-Shortest Path Graph AlgorithmabstractElectron cryomicroscopy is becoming a major experimental technique in solving the structures of large molecular assemblies. More and more three-dimensional images have been obtained at the medium resolutions between 5 and 10 Å. At this resolution range, major α-helices can be detected as cylindrical sticks and β-sheets can be detected as plain-like regions. A critical question in de novo modeling from cryo-EM images is to determine the match between the detected secondary structures from the image and those on the protein sequence. We formulate this matching problem into a constrained graph problem and present an O(Δ(2)N(2)2(N)) algorithm to this NP-Hard problem. The algorithm incorporates the dynamic programming approach into a constrained K-shortest path algorithm. Our method, DP-TOSS, has been tested using α-proteins with maximum 33 helices and α-β proteins up to five helices and 12 β-strands. The correct match was ranked within the top 35 for 19 of the 20 α-proteins and all nine α-β proteins tested. The results demonstrate that DP-TOSS improves accuracy, time and memory space in deriving the topologies of the secondary structure elements for proteins with a large number of secondary structures and a complex skeleton. Kamal Al-Nasr, Desh Ranjan, Mohammad Zubair, Lin Chen 0007, Jing He 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2013 | A Graph Approach to Bridge the Gaps in Volumetric Electron Cryo-microscopy Skeletons
Kamal Al-Nasr, Mugizi Robert Rwebangira, Legand L. Burge III |
ISBRA | 1 |
| 2013 | Intensity-Based Skeletonization of CryoEM Gray-Scale Images Using a True Segmentation-Free AlgorithmabstractCryo-electron microscopy is an experimental technique that is able to produce 3D gray-scale images of protein molecules. In contrast to other experimental techniques, cryo-electron microscopy is capable of visualizing large molecular complexes such as viruses and ribosomes. At medium resolution, the positions of the atoms are not visible and the process cannot proceed. The medium-resolution images produced by cryo-electron microscopy are used to derive the atomic structure of the proteins in de novo modeling. The skeletons of the 3D gray-scale images are used to interpret important information that is helpful in de novo modeling. Unfortunately, not all features of the image can be captured using a single segmentation. In this paper, we present a segmentation-free approach to extract the gray-scale curve-like skeletons. The approach relies on a novel representation of the 3D image, where the image is modeled as a graph and a set of volume trees. A test containing 36 synthesized maps and one authentic map shows that our approach can improve the performance of the two tested tools used in de novo modeling. The improvements were 62 and 13 percent for Gorgon and DP-TOSS, respectively. Kamal Al-Nasr, Mugizi Robert Rwebangira, Legand L. Burge III, Jing He 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | A Constraint Dynamic Graph Approach to Identify the Secondary Structure Topology from cryoEM Density Data in Presence of ErrorsabstractThe determination of the secondary structure topology is a critical step in deriving the atomic structure from the protein density map obtained from electron cryo-microscopy technique. This step often relies on the matching of two sources of information. One source comes from the secondary structures detected from the protein density map at the medium resolution, such as 5-10 A. The other source comes from the predicted secondary structures from the amino acid sequence. Due to the uncertainty in either source of information, a pool of possible secondary structure positions has to be sampled in order to include the true answer. A naive way to find the native topology is to exhaustively map the pool of possible secondary structures detected in the density map with the pool of the secondary structures predicted from the sequence and search for the topology with the lowest cost. This paper studies the question that is how to reduce the computation of the mapping when the uncertainty of the secondary structure predictions is considered. We present a method that combines the concept of dynamic graph with our previous work of using constrained shortest path to identify the topology of the secondary structures. We show a reduction of about 34.55% time as comparison to the naive way of handling the inaccuracies. To our knowledge, this is the Is computationally effective exact algorithm to identify the optimal topology of the secondary structures when the inaccuracy of the predicted data is considered. Abhishek Biswas, Dong Si, Kamal Al-Nasr, Desh Ranjan, Mohammad Zubair, Jing He 0002 |
BIBM | 3 |
| 2010 | Structure prediction for the helical skeletons detected from the low resolution protein density mapabstractBACKGROUND: The current advances in electron cryo-microscopy technique have made it possible to obtain protein density maps at about 6-10 A resolution. Although it is hard to derive the protein chain directly from such a low resolution map, the location of the secondary structures such as helices and strands can be computationally detected. It has been demonstrated that such low-resolution map can be used during the protein structure prediction process to enhance the structure prediction. RESULTS: We have developed an approach to predict the 3-dimensional structure for the helical skeletons that can be detected from the low resolution protein density map. This approach does not require the construction of the entire chain and distinguishes the structures based on the conformation of the helices. A test with 35 low resolution density maps shows that the highest ranked structure with the correct topology can be found within the top 1% of the list ranked by the effective energy formed by the helices. CONCLUSION: The results in this paper suggest that it is possible to eliminate the great majority of the bad conformations of the helices even without the construction of the entire chain of the protein. For many proteins, the effective contact energy formed by the secondary structures alone can distinguish a small set of likely structures from the pool. Kamal Al-Nasr, Weitao Sun, Jing He 0002 |
BMC Bioinform. | 1 |