Mohammad R. K. Mofrad

dblp:18/3704 · also Mohammad R. Kaazempur Mofrad · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0001-7004-4859ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 6 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Role of pore dilation in molecular transport through the nuclear pore complex: Insights from polymer scaling theory
abstract
The nuclear pore complex (NPC), a channel within the nuclear envelope filled with intrinsically disordered proteins, regulates the transport of macromolecules between the nucleus and the cytoplasm. Recent studies have highlighted the NPC's ability to adjust its diameter in response to the membrane tension, underscoring the importance of exploring how variations in pore size influence molecular transport through the NPC. In this study, we investigated the relationship between pore size and transport rate and proposed a mathematical model describing this connection. We began by theoretically analyzing how the pore size scales with the characteristic dimensions of the mesh-like structure within the pore. By introducing key assumptions about how the meshwork structure influences molecular diffusion, we derived a mathematical expression for the transport rate based on the size of the pore and the transported molecules. To validate our model, we conducted Brownian dynamics simulations using a coarse-grained representation of the NPC. These simulations, performed across a range of pore sizes, demonstrated strong agreement with our model's predictions, confirming its accuracy and applicability. Our model is specifically tailored for small-to-medium-sized molecules, approximately 5 nanometers in size, making it relevant to a wide range of transcription factors and signaling molecules. It also extends to molecules with weak and transient interactions with FG-Nups, such as importin-β. By presenting this model formula, our study offers a quantitative framework for analyzing the effects of pore dilation on nucleocytoplasmic transport.
Atsushi Matsuda, Mohammad R. K. Mofrad
PLoS Comput. Biol.2
2024 Fine-tuning protein embeddings for functional similarity evaluation
abstract
MOTIVATION: Proteins with unknown function are frequently compared to better characterized relatives, either using sequence similarity, or recently through similarity in a learned embedding space. Through comparison, protein sequence embeddings allow for interpretable and accurate annotation of proteins, as well as for downstream tasks such as clustering for unsupervised discovery of protein families. However, it is unclear whether embeddings can be deliberately designed to improve their use in these downstream tasks. RESULTS: We find that for functional annotation of proteins, as represented by Gene Ontology (GO) terms, direct fine-tuning of language models on a simple classification loss has an immediate positive impact on protein embedding quality. Fine-tuned embeddings show stronger performance as representations for K-nearest neighbor classifiers, reaching stronger performance for GO annotation than even directly comparable fine-tuned classifiers, while maintaining interpretability through protein similarity comparisons. They also maintain their quality in related tasks, such as rediscovering protein families with clustering. AVAILABILITY AND IMPLEMENTATION: github.com/mofradlab/go_metric.
Andrew Dickson, Mohammad R. K. Mofrad
Bioinform.2
2024 HGTDR: Advancing drug repurposing with heterogeneous graph transformers
abstract
MOTIVATION: Drug repurposing is a viable solution for reducing the time and cost associated with drug development. However, thus far, the proposed drug repurposing approaches still need to meet expectations. Therefore, it is crucial to offer a systematic approach for drug repurposing to achieve cost savings and enhance human lives. In recent years, using biological network-based methods for drug repurposing has generated promising results. Nevertheless, these methods have limitations. Primarily, the scope of these methods is generally limited concerning the size and variety of data they can effectively handle. Another issue arises from the treatment of heterogeneous data, which needs to be addressed or converted into homogeneous data, leading to a loss of information. A significant drawback is that most of these approaches lack end-to-end functionality, necessitating manual implementation and expert knowledge in certain stages. RESULTS: We propose a new solution, Heterogeneous Graph Transformer for Drug Repurposing (HGTDR), to address the challenges associated with drug repurposing. HGTDR is a three-step approach for knowledge graph-based drug repurposing: (1) constructing a heterogeneous knowledge graph, (2) utilizing a heterogeneous graph transformer network, and (3) computing relationship scores using a fully connected network. By leveraging HGTDR, users gain the ability to manipulate input graphs, extract information from diverse entities, and obtain their desired output. In the evaluation step, we demonstrate that HGTDR performs comparably to previous methods. Furthermore, we review medical studies to validate our method's top 10 drug repurposing suggestions, which have exhibited promising results. We also demonstrated HGTDR's capability to predict other types of relations through numerical and experimental validation, such as drug-protein and disease-protein inter-relations. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/bcb-sut/HGTDR and http://git.dml.ir/BCB/HGTDR.
Ali Gharizadeh, Karim Abbasi, Amin Ghareyazi, Mohammad R. K. Mofrad, Hamid R. Rabiee 0001
Bioinform.4
2023 GO Bench: shared hub for universal benchmarking of machine learning-based protein functional annotations
abstract
MOTIVATION: Gene annotation is the problem of mapping proteins to their functions represented as Gene Ontology (GO) terms, typically inferred based on the primary sequences. Gene annotation is a multi-label multi-class classification problem, which has generated growing interest for its uses in the characterization of millions of proteins with unknown functions. However, there is no standard GO dataset used for benchmarking the newly developed new machine learning models within the bioinformatics community. Thus, the significance of improvements for these models remains unclear. RESULTS: The Gene Benchmarking database is the first effort to provide an easy-to-use and configurable hub for the learning and evaluation of gene annotation models. It provides easy access to pre-specified datasets and takes the non-trivial steps of preprocessing and filtering all data according to custom presets using a web interface. The GO bench web application can also be used to evaluate and display any trained model on leaderboards for annotation tasks. AVAILABILITY AND IMPLEMENTATION: The GO Benchmarking dataset is freely available at www.gobench.org. Code is hosted at github.com/mofradlab, with repositories for website code, core utilities and examples of usage (Supplementary Section S.7). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andrew Dickson, Ehsaneddin Asgari, Alice C. McHardy, Mohammad R. K. Mofrad
Bioinform.4
2022 TripletProt: Deep Representation Learning of Proteins Based On Siamese Networks
abstract
Pretrained representations have recently gained attention in various machine learning applications. Nonetheless, the high computational costs associated with training these models have motivated alternative approaches for representation learning. Herein we introduce TripletProt, a new approach for protein representation learning based on the Siamese neural networks. Representation learning of biological entities which capture essential features can alleviate many of the challenges associated with supervised learning in bioinformatics. The most important distinction of our proposed method is relying on the protein-protein interaction (PPI) network. The computational cost of the generated representations for any potential application is significantly lower than comparable methods since the length of the representations is significantly smaller than that in other approaches. TripletProt offers great potentials for the protein informatics tasks and can be widely applied to similar tasks. We evaluate TripletProt comprehensively in protein functional annotation tasks including sub-cellular localization (14 categories) and gene ontology prediction (more than 2000 classes), which are both challenging multi-class, multi-label classification machine learning problems. We compare the performance of TripletProt with the state-of-the-art approaches including a recurrent language model-based approach (i.e., UniRep), as well as a protein-protein interaction (PPI) network and sequence-based method (i.e., DeepGO). Our TripletProt showed an overall improvement of F1 score in the above mentioned comprehensive functional annotation tasks, solely relying on the PPI network. Availability: The source code and datasets are available at https://github.com/EsmaeilNourani/TripletProt.
Esmaeil Nourani, Ehsaneddin Asgari, Alice C. McHardy, Mohammad R. K. Mofrad
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 EpitopeVec: linear epitope prediction using deep protein sequence embeddings
abstract
MOTIVATION: B-cell epitopes (BCEs) play a pivotal role in the development of peptide vaccines, immuno-diagnostic reagents and antibody production, and thus in infectious disease prevention and diagnostics in general. Experimental methods used to determine BCEs are costly and time-consuming. Therefore, it is essential to develop computational methods for the rapid identification of BCEs. Although several computational methods have been developed for this task, generalizability is still a major concern, where cross-testing of the classifiers trained and tested on different datasets has revealed accuracies of 51-53%. RESULTS: We describe a new method called EpitopeVec, which uses a combination of residue properties, modified antigenicity scales, and protein language model-based representations (protein vectors) as features of peptides for linear BCE predictions. Extensive benchmarking of EpitopeVec and other state-of-the-art methods for linear BCE prediction on several large and small datasets, as well as cross-testing, demonstrated an improvement in the performance of EpitopeVec over other methods in terms of accuracy and area under the curve. As the predictive performance depended on the species origin of the respective antigens (viral, bacterial and eukaryotic), we also trained our method on a large viral dataset to create a dedicated linear viral BCE predictor with improved cross-testing performance. AVAILABILITY AND IMPLEMENTATION: The software is available at https://github.com/hzi-bifo/epitope-prediction. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Akash Bahai, Ehsaneddin Asgari, Mohammad R. K. Mofrad, Andreas Kloetgen, Alice C. McHardy
Bioinform.3
2020 UniSent: Universal Adaptable Sentiment Lexica for 1000+ Languages
abstract
In this paper, we introduce UniSent universal sentiment lexica for 1000+ languages. Sentiment lexica are vital for sentiment analysis in absence of document-level annotations, a very common scenario for low-resource languages. To the best of our knowledge, UniSent is the largest sentiment resource to date in terms of the number of covered languages, including many low resource ones. In this work, we use a massively parallel Bible corpus to project sentiment information from English to other languages for sentiment analysis on Twitter data. We introduce a method called DomDrift to mitigate the huge domain mismatch between Bible and Twitter by a confidence weighting scheme that uses domain-specific embeddings to compare the nearest neighbors for a candidate sentiment word in the source (Bible) and target (Twitter) domain. We evaluate the quality of UniSent in a subset of languages for which manually created ground truth was available, Macedonian, Czech, German, Spanish, and French. We show that the quality of UniSent is comparable to manually created sentiment resources when it is used as the sentiment seed for the task of word sentiment prediction on top of embedding representations. In addition, we show that emoticon sentiments could be reliably predicted in the Twitter domain using only UniSent and monolingual embeddings in German, Spanish, French, and Italian. With the publication of this paper, we release the UniSent sentiment lexica at http://language-lab.info/unisent.
Ehsaneddin Asgari, Fabienne Braune, Benjamin Roth 0001, Christoph Ringlstetter, Mohammad R. K. Mofrad
LREC5
2019 MicroPheno: predicting environments and host phenotypes from 16S rRNA gene sequencing using a k-mer based representation of shallow sub-samples
abstract
Bioinformatics, bty296, https://doi.org/10.1093/bioinformatics/bty296 The author wishes to inform readers that the affiliations for the authors were incorrect in the original paper. They appear correctly above. The paper has been corrected online.
Ehsaneddin Asgari, Kiavash Garakani, Alice C. McHardy, Mohammad R. K. Mofrad
Bioinform.4
2019 DiTaxa: nucleotide-pair encoding of 16S rRNA for host phenotype and biomarker detection
abstract
SUMMARY: Identifying distinctive taxa for micro-biome-related diseases is considered key to the establishment of diagnosis and therapy options in precision medicine and imposes high demands on the accuracy of micro-biome analysis techniques. We propose an alignment- and reference- free subsequence based 16S rRNA data analysis, as a new paradigm for micro-biome phenotype and biomarker detection. Our method, called DiTaxa, substitutes standard operational taxonomic unit (OTU)-clustering by segmenting 16S rRNA reads into the most frequent variable-length subsequences. We compared the performance of DiTaxa to the state-of-the-art methods in phenotype and biomarker detection, using human-associated 16S rRNA samples for periodontal disease, rheumatoid arthritis and inflammatory bowel diseases, as well as a synthetic benchmark dataset. DiTaxa performed competitively to the k-mer based state-of-the-art approach in phenotype prediction while outperforming the OTU-based state-of-the-art approach in finding biomarkers in both resolution and coverage evaluated over known links from literature and synthetic benchmark datasets. AVAILABILITY AND IMPLEMENTATION: DiTaxa is available under the Apache 2 license at http://llp.berkeley.edu/ditaxa. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ehsaneddin Asgari, Philipp C. Münch, Till R. Lesker, Alice C. McHardy, Mohammad R. K. Mofrad
Bioinform.5
2018 MicroPheno: predicting environments and host phenotypes from 16S rRNA gene sequencing using a k-mer based representation of shallow sub-samples
abstract
Motivation: Microbial communities play important roles in the function and maintenance of various biosystems, ranging from the human body to the environment. A major challenge in microbiome research is the classification of microbial communities of different environments or host phenotypes. The most common and cost-effective approach for such studies to date is 16S rRNA gene sequencing. Recent falls in sequencing costs have increased the demand for simple, efficient and accurate methods for rapid detection or diagnosis with proved applications in medicine, agriculture and forensic science. We describe a reference- and alignment-free approach for predicting environments and host phenotypes from 16S rRNA gene sequencing based on k-mer representations that benefits from a bootstrapping framework for investigating the sufficiency of shallow sub-samples. Deep learning methods as well as classical approaches were explored for predicting environments and host phenotypes. Results: A k-mer distribution of shallow sub-samples outperformed Operational Taxonomic Unit (OTU) features in the tasks of body-site identification and Crohn's disease prediction. Aside from being more accurate, using k-mer features in shallow sub-samples allows (i) skipping computationally costly sequence alignments required in OTU-picking and (ii) provided a proof of concept for the sufficiency of shallow and short-length 16S rRNA sequencing for phenotype prediction. In addition, k-mer features predicted representative 16S rRNA gene sequences of 18 ecological environments, and 5 organismal environments with high macro-F1 scores of 0.88 and 0.87. For large datasets, deep learning outperformed classical methods such as Random Forest and Support Vector Machine. Availability and implementation: The software and datasets are available at https://llp.berkeley.edu/micropheno. Supplementary information: Supplementary data are available at Bioinformatics online.
Ehsaneddin Asgari, Kiavash Garakani, Alice C. McHardy, Mohammad R. K. Mofrad
Bioinform.4
2013 The Interaction of Vinculin with Actin
abstract
Vinculin can interact with F-actin both in recruitment of actin filaments to the growing focal adhesions and also in capping of actin filaments to regulate actin dynamics. Using molecular dynamics, both interactions are simulated using different vinculin conformations. Vinculin is simulated either with only its vinculin tail domain (Vt), with all residues in its closed conformation, with all residues in an open I conformation, and with all residues in an open II conformation. The open I conformation results from movement of domain 1 away from Vt; the open II conformation results from complete dissociation of Vt from the vinculin head domains. Simulation of vinculin binding along the actin filament showed that Vt alone can bind along the actin filaments, that vinculin in its closed conformation cannot bind along the actin filaments, and that vinculin in its open I conformation can bind along the actin filaments. The simulations confirm that movement of domain 1 away from Vt in formation of vinculin 1 is sufficient for allowing Vt to bind along the actin filament. Simulation of Vt capping actin filaments probe six possible bound structures and suggest that vinculin would cap actin filaments by interacting with both S1 and S3 of the barbed-end, using the surface of Vt normally occluded by D4 and nearby vinculin head domain residues. Simulation of D4 separation from Vt after D1 separation formed the open II conformation. Binding of open II vinculin to the barbed-end suggests this conformation allows for vinculin capping. Three binding sites on F-actin are suggested as regions that could link to vinculin. Vinculin is suggested to function as a variable switch at the focal adhesions. The conformation of vinculin and the precise F-actin binding conformation is dependent on the level of mechanical load on the focal adhesion.
Javad Golji, Mohammad R. K. Mofrad
PLoS Comput. Biol.2
2013 Localized Lipid Packing of Transmembrane Domains Impedes Integrin Clustering
abstract
Integrin clustering plays a pivotal role in a host of cell functions. Hetero-dimeric integrin adhesion receptors regulate cell migration, survival, and differentiation by communicating signals bidirectionally across the plasma membrane. Thus far, crystallographic structures of integrin components are solved only separately, and for some integrin types. Also, the sequence of interactions that leads to signal transduction remains ambiguous. Particularly, it remains controversial whether the homo-dimerization of integrin transmembrane domains occurs following the integrin activation (i.e. when integrin ectodomain is stretched out) or if it regulates integrin clustering. This study employs molecular dynamics modeling approaches to address these questions in molecular details and sheds light on the crucial effect of the plasma membrane. Conducting a normal mode analysis of the intact αllbβ3 integrin, it is demonstrated that the ectodomain and transmembrane-cytoplasmic domains are connected via a membrane-proximal hinge region, thus merely transmembrane-cytoplasmic domains are modeled. By measuring the free energy change and force required to form integrin homo-oligomers, this study suggests that the β-subunit homo-oligomerization potentially regulates integrin clustering, as opposed to α-subunit, which appears to be a poor regulator for the clustering process. If α-subunits are to regulate the clustering they should overcome a high-energy barrier formed by a stable lipid pack around them. Finally, an outside-in activation-clustering scenario is speculated, explaining how further loading the already-active integrin affects its homo-oligomerization so that focal adhesions grow in size.
Mehrdad Mehrbod, Mohammad R. K. Mofrad
PLoS Comput. Biol.2
2011 Brownian Dynamics Simulation of Nucleocytoplasmic Transport: A Coarse-Grained Model for the Functional State of the Nuclear Pore Complex
abstract
The nuclear pore complex (NPC) regulates molecular traffic across the nuclear envelope (NE). Selective transport happens on the order of milliseconds and the length scale of tens of nanometers; however, the transport mechanism remains elusive. Central to the transport process is the hydrophobic interactions between karyopherins (kaps) and Phe-Gly (FG) repeat domains. Taking into account the polymeric nature of FG-repeats grafted on the elastic structure of the NPC, and the kap-FG hydrophobic affinity, we have established a coarse-grained model of the NPC structure that mimics nucleocytoplasmic transport. To establish a foundation for future works, the methodology and biophysical rationale behind the model is explained in details. The model predicts that the first-passage time of a 15 nm cargo-complex is about 2.6±0.13 ms with an inverse Gaussian distribution for statistically adequate number of independent Brownian dynamics simulations. Moreover, the cargo-complex is primarily attached to the channel wall where it interacts with the FG-layer as it passes through the central channel. The kap-FG hydrophobic interaction is highly dynamic and fast, which ensures an efficient translocation through the NPC. Further, almost all eight hydrophobic binding spots on kap-β are occupied simultaneously during transport. Finally, as opposed to intact NPCs, cytoplasmic filaments-deficient NPCs show a high degree of permeability to inert cargos, implying the defining role of cytoplasmic filaments in the selectivity barrier.
Ruhollah Moussavi-Baygi, Yousef Jamali, Reza Karimi, Mohammad R. K. Mofrad
PLoS Comput. Biol.4
2009 Molecular Mechanics of the α-Actinin Rod Domain: Bending, Torsional, and Extensional Behavior
abstract
alpha-Actinin is an actin crosslinking molecule that can serve as a scaffold and maintain dynamic actin filament networks. As a crosslinker in the stressed cytoskeleton, alpha-actinin can retain conformation, function, and strength. alpha-Actinin has an actin binding domain and a calmodulin homology domain separated by a long rod domain. Using molecular dynamics and normal mode analysis, we suggest that the alpha-actinin rod domain has flexible terminal regions which can twist and extend under mechanical stress, yet has a highly rigid interior region stabilized by aromatic packing within each spectrin repeat, by electrostatic interactions between the spectrin repeats, and by strong salt bridges between its two anti-parallel monomers. By exploring the natural vibrations of the alpha-actinin rod domain and by conducting bending molecular dynamics simulations we also predict that bending of the rod domain is possible with minimal force. We introduce computational methods for analyzing the torsional strain of molecules using rotating constraints. Molecular dynamics extension of the alpha-actinin rod is also performed, demonstrating transduction of the unfolding forces across salt bridges to the associated monomer of the alpha-actinin rod domain.
Javad Golji, Robert Collins, Mohammad R. K. Mofrad
PLoS Comput. Biol.3