EDBT 2026 Demo / reviewers in the wild / expert
Ewy A. Mathé
dblp:197/8103
· DBLP profile ↗
13ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0003-4491-8107ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Aligning Orphanet Classification to Identify Disease Characteristics among Rare Disease ClustersabstractUnderstanding the underlying etiologies of rare diseases may facilitate research across multiple conditions, enabling basket trail design and drug repurposing. In this study, we aligned clusters of rare diseases with Orphanet classifications to represent their shared etiologies and establish a foundation for further investigation on underly biological mechanism discovery. By utilizing the linearized Orphanet categories, we connected 35 clusters of rare diseases into 18 classifications. Significant associations were found between the categories "Rare Developmental Defects During Embryogenesis" and "Rare Inborn Errors of Metabolism" and the clusters in this study, suggesting that many rare diseases originating in the prenatal period or related to metabolism may present a substantial opportunity for success in future investigation. Sungrim Moon, Jessica Maine, Ewy A. Mathé, Qian Zhu 0003 |
BIBM | 3 |
| 2023 | RaMP-DB 2.0: a renovated knowledgebase for deriving biological and chemical insight from metabolites, proteins, and genesabstractMOTIVATION: Functional interpretation of high-throughput metabolomic and transcriptomic results is a crucial step in generating insight from experimental data. However, pathway and functional information for genes and metabolites are distributed among many siloed resources, limiting the scope of analyses that rely on a single knowledge source. RESULTS: RaMP-DB 2.0 is a web interface, relational database, API and R package designed for straightforward and comprehensive functional interpretation of metabolomic and multi-omic data. RaMP-DB 2.0 has been upgraded with an expanded breadth and depth of functional and chemical annotations (ClassyFire, LIPID MAPS, SMILES, InChIs, etc.), with new data types related to metabolites and lipids incorporated. To streamline entity resolution across multiple source databases, we have implemented a new semi-automated process, thereby lessening the burden of harmonization and supporting more frequent updates. The associated RaMP-DB 2.0 R package now supports queries on pathways, common reactions (e.g. metabolite-enzyme relationship), chemical functional ontologies, chemical classes and chemical structures, as well as enrichment analyses on pathways (multi-omic) and chemical classes. Lastly, the RaMP-DB web interface has been completely redesigned using the Angular framework. AVAILABILITY AND IMPLEMENTATION: The code used to build all components of RaMP-DB 2.0 are freely available on GitHub at https://github.com/ncats/ramp-db, https://github.com/ncats/RaMP-Client/ and https://github.com/ncats/RaMP-Backend. The RaMP-DB web application can be accessed at https://rampdb.nih.gov/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. John C. Braisted, Andrew Patt, Cole Tindall, Timothy Sheils, Jorge Neyra, Kyle Spencer, Tara Eicher, Ewy A. Mathé |
Bioinform. | 8 |
| 2023 | Clustering rare diseases within an ontology-enriched knowledge graphabstractOBJECTIVE: Identifying sets of rare diseases with shared aspects of etiology and pathophysiology may enable drug repurposing. Toward that aim, we utilized an integrative knowledge graph to construct clusters of rare diseases. MATERIALS AND METHODS: Data on 3242 rare diseases were extracted from the National Center for Advancing Translational Science Genetic and Rare Diseases Information center internal data resources. The rare disease data enriched with additional biomedical data, including gene and phenotype ontologies, biological pathway data, and small molecule-target activity data, to create a knowledge graph (KG). Node embeddings were trained and clustered. We validated the disease clusters through semantic similarity and feature enrichment analysis. RESULTS: Thirty-seven disease clusters were created with a mean size of 87 diseases. We validate the clusters quantitatively via semantic similarity based on the Orphanet Rare Disease Ontology. In addition, the clusters were analyzed for enrichment of associated genes, revealing that the enriched genes within clusters are highly related. DISCUSSION: We demonstrate that node embeddings are an effective method for clustering diseases within a heterogenous KG. Semantically similar diseases and relevant enriched genes have been uncovered within the clusters. Connections between disease clusters and drugs are enumerated for follow-up efforts. CONCLUSION: We lay out a method for clustering rare diseases using graph node embeddings. We develop an easy-to-maintain pipeline that can be updated when new data on rare diseases emerges. The embeddings themselves can be paired with other representation learning methods for other data types, such as drugs, to address other predictive modeling problems. Jaleal Sanjak, Jessica Binder, Arjun Singh Yadaw, Qian Zhu 0003, Ewy A. Mathé |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Scientific Evidence Based Knowledge Graph in Rare DiseasesabstractRare diseases are naturally associated with low prevalence rate, which raises a big challenge due to less data available for supporting preclinical and clinical studies. Therefore, it is critical to fully utilize the accumulated scientific publications in rare diseases over years, in order to access full spectrum of scientific research and enable relevant scientific evidence extraction and generation. In this study, we obtained rare disease related PubMed articles, extracted multiple types of biomedical information, and semantically presented the data in a knowledge graph, which is hosted in Neo4j based on a predefined data model to support further rare disease research. Qian Zhu 0003, Ruizheng Liu, Gunjan Vatas, Andrew Clough, Yanji Xu, Dac-Trung Nguyen, Ewy A. Mathé, Eric Sid |
BIBM | 7 |
| 2021 | Network analyses in microbiome based on high-throughput multi-omics dataabstractTogether with various hosts and environments, ubiquitous microbes interact closely with each other forming an intertwined system or community. Of interest, shifts of the relationships between microbes and their hosts or environments are associated with critical diseases and ecological changes. While advances in high-throughput Omics technologies offer a great opportunity for understanding the structures and functions of microbiome, it is still challenging to analyse and interpret the omics data. Specifically, the heterogeneity and diversity of microbial communities, compounded with the large size of the datasets, impose a tremendous challenge to mechanistically elucidate the complex communities. Fortunately, network analyses provide an efficient way to tackle this problem, and several network approaches have been proposed to improve this understanding recently. Here, we systemically illustrate these network theories that have been used in biological and biomedical research. Then, we review existing network modelling methods of microbial studies at multiple layers from metagenomics to metabolomics and further to multi-omics. Lastly, we discuss the limitations of present studies and provide a perspective for further directions in support of the understanding of microbial communities. Zhaoqian Liu, Anjun Ma, Ewy A. Mathé, Marlena Merling, Qin Ma 0003, Bingqiang Liu |
Briefings Bioinform. | 3 |
| 2021 | Self-organizing maps with variable neighborhoods facilitate learning of chromatin accessibility signal shapes associated with regulatory elementsabstractBACKGROUND: Assigning chromatin states genome-wide (e.g. promoters, enhancers, etc.) is commonly performed to improve functional interpretation of these states. However, computational methods to assign chromatin state suffer from the following drawbacks: they typically require data from multiple assays, which may not be practically feasible to obtain, and they depend on peak calling algorithms, which require careful parameterization and often exclude the majority of the genome. To address these drawbacks, we propose a novel learning technique built upon the Self-Organizing Map (SOM), Self-Organizing Map with Variable Neighborhoods (SOM-VN), to learn a set of representative shapes from a single, genome-wide, chromatin accessibility dataset to associate with a chromatin state assignment in which a particular RE is prevalent. These shapes can then be used to assign chromatin state using our workflow. RESULTS: We validate the performance of the SOM-VN workflow on 14 different samples of varying quality, namely one assay each of A549 and GM12878 cell lines and two each of H1 and HeLa cell lines, primary B-cells, and brain, heart, and stomach tissue. We show that SOM-VN learns shapes that are (1) non-random, (2) associated with known chromatin states, (3) generalizable across sets of chromosomes, and (4) associated with magnitude and multimodality. We compare the accuracy of SOM-VN chromatin states against the Clustering Aggregation Tool (CAGT), an unsupervised method that learns chromatin accessibility signal shapes but does not associate these shapes with REs, and we show that overall precision and recall is increased when learning shapes using SOM-VN as compared to CAGT. We further compare enhancer state assignments from SOM-VN in signals above a set threshold to enhancer state assignments from Predicting Enhancers from ATAC-seq Data (PEAS), a deep learning method that assigns enhancer chromatin states to peaks. We show that the precision-recall area under the curve for the assignment of enhancer states is comparable to PEAS. CONCLUSIONS: Our work shows that the SOM-VN workflow can learn relationships between REs and chromatin accessibility signal shape, which is an important step toward the goal of assigning and comparing enhancer state across multiple experiments and phenotypic states. Tara Eicher, Jany Chan, Han Luu, Raghu Machiraju, Ewy A. Mathé |
BMC Bioinform. | 5 |
| 2020 | Correction to: The International Conference on Intelligent Biology and Medicine (ICIBM) 2019: bioinformatics methods and applications for human diseasesabstractAfter publication of this supplement article [1], it is requested the grant ID in the Funding section should be corrected from NSF grant IIS-7811367 to NSF grant IIS-1902617. Therefore, the correct 'Funding' section in this article should read: We thank the National Science Foundation (NSF grant IIS-1902617) for the financial support of ICIBM 2019. This article has not received sponsorship for publication. Zhongming Zhao, Yulin Dai, Ewy A. Mathé, Kai Wang 0063 |
BMC Bioinform. | 4 |
| 2019 | Challenges in proteogenomics: a comparison of analysis methods with the case study of the DREAM proteogenomics sub-challengeabstractBACKGROUND: Proteomic measurements, which closely reflect phenotypes, provide insights into gene expression regulations and mechanisms underlying altered phenotypes. Further, integration of data on proteome and transcriptome levels can validate gene signatures associated with a phenotype. However, proteomic data is not as abundant as genomic data, and it is thus beneficial to use genomic features to predict protein abundances when matching proteomic samples or measurements within samples are lacking. RESULTS: We evaluate and compare four data-driven models for prediction of proteomic data from mRNA measured in breast and ovarian cancers using the 2017 DREAM Proteogenomics Challenge data. Our results show that Bayesian network, random forests, LASSO, and fuzzy logic approaches can predict protein abundance levels with median ground truth-predicted correlation values between 0.2 and 0.5. However, the most accurately predicted proteins differ considerably between approaches. CONCLUSIONS: In addition to benchmarking aforementioned machine learning approaches for predicting protein levels from transcript levels, we discuss challenges and potential solutions in state-of-the-art proteogenomic analyses. Tara Eicher, Andrew Patt, Esko Kautto, Raghu Machiraju, Ewy A. Mathé |
BMC Bioinform. | 5 |
| 2019 | The International Conference on Intelligent Biology and Medicine (ICIBM) 2019: bioinformatics methods and applications for human diseasesabstractBetween June 9-11, 2019, the International Conference on Intelligent Biology and Medicine (ICIBM 2019) was held in Columbus, Ohio, USA. The conference included 12 scientific sessions, five tutorials or workshops, one poster session, four keynote talks and four eminent scholar talks that covered a wide range of topics in bioinformatics, medical informatics, systems biology and intelligent computing. Here, we describe 13 high quality research articles selected for publishing in BMC Bioinformatics. Zhongming Zhao, Yulin Dai, Ewy A. Mathé, Kai Wang 0063 |
BMC Bioinform. | 4 |
| 2018 | IntLIM: integration using linear models of metabolomics and gene expression dataabstractBACKGROUND: Integration of transcriptomic and metabolomic data improves functional interpretation of disease-related metabolomic phenotypes, and facilitates discovery of putative metabolite biomarkers and gene targets. For this reason, these data are increasingly collected in large (> 100 participants) cohorts, thereby driving a need for the development of user-friendly and open-source methods/tools for their integration. Of note, clinical/translational studies typically provide snapshot (e.g. one time point) gene and metabolite profiles and, oftentimes, most metabolites measured are not identified. Thus, in these types of studies, pathway/network approaches that take into account the complexity of transcript-metabolite relationships may neither be applicable nor readily uncover novel relationships. With this in mind, we propose a simple linear modeling approach to capture disease-(or other phenotype) specific gene-metabolite associations, with the assumption that co-regulation patterns reflect functionally related genes and metabolites. RESULTS: The proposed linear model, metabolite ~ gene + phenotype + gene:phenotype, specifically evaluates whether gene-metabolite relationships differ by phenotype, by testing whether the relationship in one phenotype is significantly different from the relationship in another phenotype (via a statistical interaction gene:phenotype p-value). Statistical interaction p-values for all possible gene-metabolite pairs are computed and significant pairs are then clustered by the directionality of associations (e.g. strong positive association in one phenotype, strong negative association in another phenotype). We implemented our approach as an R package, IntLIM, which includes a user-friendly R Shiny web interface, thereby making the integrative analyses accessible to non-computational experts. We applied IntLIM to two previously published datasets, collected in the NCI-60 cancer cell lines and in human breast tumor and non-tumor tissue, for which transcriptomic and metabolomic data are available. We demonstrate that IntLIM captures relevant tumor-specific gene-metabolite associations involved in known cancer-related pathways, including glutamine metabolism. Using IntLIM, we also uncover biologically relevant novel relationships that could be further tested experimentally. CONCLUSIONS: IntLIM provides a user-friendly, reproducible framework to integrate transcriptomic and metabolomic data and help interpret metabolomic data and uncover novel gene-metabolite relationships. The IntLIM R package is publicly available in GitHub ( https://github.com/mathelab/IntLIM ) and includes a user-friendly web application, vignettes, sample data and data/code to reproduce results. Jalal K. Siddiqui, Elizabeth Baskin, Carmen Z. Cantemir-Stone, Bofei Zhang, Russell Bonneville, Joseph P. McElroy, Kevin R. Coombes, Ewy A. Mathé |
BMC Bioinform. | 9 |
| 2017 | ALTRE: workflow for defining ALTered Regulatory Elements using chromatin accessibility dataabstractSummary: Regulatory elements regulate gene transcription, and their location and accessibility is cell-type specific, particularly for enhancers. Mapping and comparing chromatin accessibility between different cell types may identify mechanisms involved in cellular development and disease progression. To streamline and simplify differential analysis of regulatory elements genome-wide using chromatin accessibility data, such as DNase-seq, ATAC-seq, we developed ALTRE ( ALT ered R egulatory E lements), an R package and associated R Shiny web app. ALTRE makes such analysis accessible to a wide range of users-from novice to practiced computational biologists. Availability and Implementation: https://github.com/Mathelab/ALTRE. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Elizabeth Baskin, Rick Farouni, Ewy A. Mathé |
Bioinform. | 3 |
| 2017 | ALTRE: workflow for defining ALTered regulatory elements using chromatin accessibility dataabstractBioinformatics (2017) 33(5), 740–742 The authors of the above article wish to inform readers that Elizabeth Baskin and Rick Farouni were both equal contributors to the article. Elizabeth Baskin, Rick Farouni, Ewy A. Mathé |
Bioinform. | 3 |
| 2008 | Structure Based Functional Analysis of Bacteriophage f1 Gene V ProteinabstractA computational mutagenesis methodology utilizing a four-body, knowledge-based, statistical contact potential is applied toward globally quantifying relative structural changes (residual scores) in bacteriophage f1 gene V protein (GVP) due to single amino acid residue substitutions. We show that these residual scores correlate well with experimentally measured relative changes in protein function caused by the mutations. For each mutant, the approach also yields local measures of environmental perturbation occurring at every residue position (residual profile) in the protein. Implementation of the random forest algorithm, utilizing experimental GVP mutants whose feature vector components include environmental changes at the mutated position and at six nearest neighbors, correctly classifies mutants based on function with up to 72% accuracy while achieving 0.77 area under the receiver operating characteristic curve and a 0.42 correlation coefficient. An optimally trained random forest model is subsequently used to infer function for all remaining unexplored GVP mutants. Majid Masso, Ewy A. Mathé, Nida Parvez, Kahkeshan Hijazi, Iosif I. Vaisman |
BIBM | 2 |