VLDB 2026 Research / reviewers in the wild / expert
Gil Alterovitz
dblp:37/3236
· DBLP profile ↗
39ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Decoding Aging: Multi-Omics Insights into Oxidative Stress, Mitochondrial Dysfunction, and Cellular Senescence in FibroblastsabstractThis research investigates the complex biochemical mechanisms underlying aging by analyzing primary human fibroblasts using a longitudinal multi-omics dataset. This dataset includes cytology, DNA methylation and epigenetic clocks, bioenergetics, and cytokine profiling. Key findings indicate that mitochondrial efficiency declines with age, while glycolysis becomes more prevalent to compensate for energy demands. Epigenetic clocks, such as Hannum and PhenoAge, showed strong correlations with biological age (ρ > 0.650, p < 1e-6), validating the experimental setup and confirming that the cultured fibroblasts were aging appropriately. Fibroblasts with SURF1 mutations exhibited accelerated aging, marked by bioenergetic deficits, increased cell volume, and reduced proliferative capacity, underscoring the pivotal role of mitochondrial dysfunction in cellular senescence. Novel insights were gained from analyzing cytokines like IL-18 and PCSK9, some of which were linked to age-related diseases such as Alzheimer’s and cardiovascular disorders. Experimental treatments revealed distinct effects on cellular aging. Dexamethasone reduced inflammation but also increased DNA methylation, induced metabolic inefficiencies, and shortened cellular lifespan. By uncovering connections between mitochondrial dysfunction, epigenetic biomarkers, and immune dysregulation, this research identifies potential therapeutic targets for age-related diseases. Rajarshi Mandal, Gil Alterovitz |
CIBCB | 3 |
| 2025 | A unified protein embedding model with local and global structural sensitivityabstractAbstract Introduction Many research tasks depend on identifying structural homologs of proteins, including but not limited to evolutionary analysis, peptidomimetics, and functional an- notation. However, identifying structural homologs at scale requires efficient comparison algorithms, which currently are still limited. While sequence alignment algorithms like BLAST and MMSeqs2 have become extremely optimized, they are ineffective for structural comparisons, since many structural homologs differ vastly in sequence. Meanwhile, structural alignment algorithms like TM-Align [1] and DALI. [2] are computationally expensive due to high algorithmic complexity. Such algorithms often perform submatrix align- ment on Cα distance matrices to directly superimpose proteins, which has a quadratic time complexity or worse. In contrast, protein language models [3], or PLMs, can predict structural similarity in O(1) time. Utilizing the biophysi- cal prior that sequences can reconstruct structures, structural PLMs can generate sequence-based but structurally-aware em- beddings for proteins, which can then be compared via cosine similarity. Although efficient, existing structural PLMs do not recognize localized structural changes. One notable example is. The network outputs two types of embeddings: it directly produces per-residue embeddings, and then pools them into a global embedding for the protein. Predicted TM-scores are calculated as the cosine similarity between the global em- beddings of each protein. lDDT score prediction only occurs after filtering out low TM-score protein pairs; then, using the sequence alignment generated by TM-align, we predict the lDDT scores for each pair of aligned residues as the cosine similarity of the aligned embeddings. Training occurred over 5 epochs of 300, 000 protein pairs. Results We evaluated our model against two professionally curated datasets, one from the TM-Vec paper, and one from VIPUR. For the 886 proteins from the TM-Vec dataset, Table 1 evaluates our TM-score and lDDT-score prediction, as well as TM-Vec’s TM-score prediction. The TM-score error histogram is shown in Fig. 2, and the TM-score error box plot is shown in Fig. 3. TM-Vec [4], a recently-developed PLM trained to predict TM-scores (template-modeling scores), a global similarity metric for proteins. While highly accurate for TM-scores, TM-Vec is locally unaware. We aim to create a PLM that is both globally and locally structure-aware. Methodology Our PLM is a transformer-based Siamese neural network [5], consisting of two identical neural networks and a comparison head (see Fig. 1). Our PLM addresses both the inefficiency and local insensitivity of prior methods. We continue to generate sequence-based embeddings, resulting in efficient comparison times. Furthermore, our model utilizes a loss function that combines TM-score and a custom variation of lDDT scores (Local Distance Different Test scores) [6], which are per-residue structural similarity scores, ensuring the model captures both global and local structural features. Contrastive loss was based upon the function below. θt represents the model parameters at time t, which are updated based upon the loss of the previous parameters f (θt − 1). The degree to which each type of loss (TM, lDDT) contributes to the overall loss can be adjusted via the α, β hyperparameters. (In our model, α = 0.7, β = 0.3.) For the 364 proteins in the VIPUR dataset, Table 2 evaluates our TM-score and lDDT-score prediction, as well as TM-Vec’s TM-score prediction. The lDDT error histogram is shown in Fig. 4. Conclusion Our testing results confirmed the plausibility of our framework, as our model performed similarly in TM-score prediction when compared to highly accurate models like TM- Vec while also producing accurate lDDT scores. This dual capability makes the model a potential tool in downstream research tasks, especially mutation analysis, where it could aid in the identification of deleterious mutations as well as the recognition of affected subdomains. References 1. Y. Zhang, ‘Tm-align: a protein structure alignment algorithm based on the tm-score,’ Nucleic Acids Research, vol. 33, no. 7, p. 2302–2309, Apr. 2005. [Online]. Available: http://dx.doi.org/10.1093/nar/gki524 2. L. Holm and C. Sander, ‘Protein structure comparison by alignment of distance matrices,’ Journal of Molecular Biology, vol. 233, no. 1, p. 123–138, Sep. 1993. [Online]. Available: http://dx.doi.org/10.1006/jmbi.1993.1489 3. L. Wang, X. Li, H. Zhang, J. Wang, D. Jiang, Z. Xue, and Y. Wang, ‘A comprehensive review of protein language models,’ 2025. [Online]. Available: https://arxiv.org/abs/2502.06881 4. T. Hamamsy, J. T. Morton, R. Blackwell, D. Berenberg, N. Carriero, V. Gligorijevic, C. E. M. Strauss, J. K. Leman, K. Cho, and R. Bonneau, ‘Protein remote homology detection and structural alignment using deep learning,’ Nature Biotechnology, vol. 42, no. 6, p. 975–985, Sep. 2023. [Online]. Available: http://dx.doi.org/10.1038/s41587-023-01917-2 5. Y. Li, C. L. P. Chen, and T. Zhang, ‘A survey on siamese network: Methodologies, applications, and opportunities,’ IEEE Transactions on Artificial Intelligence, vol. 3, no. 6, p. 994–1014, Dec. 2022. [Online]. Available: http://dx.doi.org/10.1109/TAI.2022.3207112 6. V. Mariani, M. Biasini, A. Barbato, and T. Schwede, ‘lddt: a local superposition-free score for comparing protein structures and models using distance difference tests,’ Bioinformatics, vol. 29, no. 21, p. 2722–2728, Aug. 2013. [Online]. Available: http://dx.doi.org/10.1093/bioinformatics/btt473. Jerry Xu, Shaojun Pei, Gil Alterovitz |
Briefings Bioinform. | 3 |
| 2023 | Introducing HL7 FHIR Genomics Operations: a developer-friendly approach to genomics-EHR integrationabstractOBJECTIVE: Enabling clinicians to formulate individualized clinical management strategies from the sea of molecular data remains a fundamentally important but daunting task. Here, we describe efforts towards a new paradigm in genomics-electronic health record (HER) integration, using a standardized suite of FHIR Genomics Operations that encapsulates the complexity of molecular data so that precision medicine solution developers can focus on building applications. MATERIALS AND METHODS: FHIR Genomics Operations essentially "wrap" a genomics data repository, presenting a uniform interface to applications. More importantly, operations encapsulate the complexity of data within a repository and normalize redundant data representations-particularly relevant in genomics, where a tremendous amount of raw data exists in often-complex non-FHIR formats. RESULTS: Fifteen FHIR Genomics Operations have been developed, designed to support a wide range of clinical scenarios, such as variant discovery; clinical trial matching; hereditary condition and pharmacogenomic screening; and variant reanalysis. Operations are being matured through the HL7 balloting process, connectathons, pilots, and the HL7 FHIR Accelerator program. DISCUSSION: Next-generation sequencing can identify thousands to millions of variants, whose clinical significance can change over time as our knowledge evolves. To manage such a large volume of dynamic and complex data, new models of genomics-EHR integration are needed. Qualitative observations to date suggest that freeing application developers from the need to understand the nuances of genomic data, and instead base applications on standardized APIs can not only accelerate integration but also dramatically expand the applications of Omic data in driving precision care at scale for all. Robert H. Dolin, Bret S. E. Heale, Gil Alterovitz, Rohan Gupta, Justin Aronson, Aziz A. Boxwala, Shaileshbhai R. Gothi, David Haines, Arthur Hermann, Tonya Hongsermeier, Ammar Husami, Frank Naeymi-Rad, Barbara Rapchak, Chandan Ravishankar, James Shalaby, May Terry, Powell Zhang, Srikar Chamala |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | A research agenda to support the development and implementation of genomics-based clinical informatics tools and resourcesabstractOBJECTIVE: The Genomic Medicine Working Group of the National Advisory Council for Human Genome Research virtually hosted its 13th genomic medicine meeting titled "Developing a Clinical Genomic Informatics Research Agenda". The meeting's goal was to articulate a research strategy to develop Genomics-based Clinical Informatics Tools and Resources (GCIT) to improve the detection, treatment, and reporting of genetic disorders in clinical settings. MATERIALS AND METHODS: Experts from government agencies, the private sector, and academia in genomic medicine and clinical informatics were invited to address the meeting's goals. Invitees were also asked to complete a survey to assess important considerations needed to develop a genomic-based clinical informatics research strategy. RESULTS: Outcomes from the meeting included identifying short-term research needs, such as designing and implementing standards-based interfaces between laboratory information systems and electronic health records, as well as long-term projects, such as identifying and addressing barriers related to the establishment and implementation of genomic data exchange systems that, in turn, the research community could help address. DISCUSSION: Discussions centered on identifying gaps and barriers that impede the use of GCIT in genomic medicine. Emergent themes from the meeting included developing an implementation science framework, defining a value proposition for all stakeholders, fostering engagement with patients and partners to develop applications under patient control, promoting the use of relevant clinical workflows in research, and lowering related barriers to regulatory processes. Another key theme was recognizing pervasive biases in data and information systems, algorithms, access, value, and knowledge repositories and identifying ways to resolve them. Ken Wiley, Laura Findley, Madison Goldrich, Teji Rakhra-Burris, Ana Stevens, Pamela Williams, Carol J. Bult, Rex L. Chisholm, Patricia Deverka, Geoffrey S. Ginsburg, Eric D. Green, Gail P. Jarvik, George A. Mensah, Erin Ramos, Mary Relling, Dan M. Roden, Robb Rowley, Gil Alterovitz, Samuel J. Aronson, Lisa Bastarache, James J. Cimino, Erin L. Crowgey, Guilherme Del Fiol, Robert R. Freimuth, Mark A. Hoffman, Janina M. Jeff, Kevin B. Johnson, Kensaku Kawamoto, Subha Madhavan, Eneida A. Mendonça, Lucila Ohno-Machado, Siddharth Pratap, Casey Overby Taylor, Marylyn D. Ritchie, Nephi Walton, Chunhua Weng, Teresa Zayas-Cabán, Teri A. Manolio, Marc S. Williams |
J. Am. Medical Informatics Assoc. | 18 |
| 2021 | netAE: semi-supervised dimensionality reduction of single-cell RNA sequencing to facilitate cell labelingabstractMOTIVATION: Single-cell RNA sequencing allows us to study cell heterogeneity at an unprecedented cell-level resolution and identify known and new cell populations. Current cell labeling pipeline uses unsupervised clustering and assigns labels to clusters by manual inspection. However, this pipeline does not utilize available gold-standard labels because there are usually too few of them to be useful to most computational methods. This article aims to facilitate cell labeling with a semi-supervised method in an alternative pipeline, in which a few gold-standard labels are first identified and then extended to the rest of the cells computationally. RESULTS: We built a semi-supervised dimensionality reduction method, a network-enhanced autoencoder (netAE). Tested on three public datasets, netAE outperforms various dimensionality reduction baselines and achieves satisfactory classification accuracy even when the labeled set is very small, without disrupting the similarity structure of the original space. AVAILABILITY AND IMPLEMENTATION: The code of netAE is available on GitHub: https://github.com/LeoZDong/netAE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhengyang Dong, Gil Alterovitz |
Bioinform. | 2 |
| 2021 | vcf2fhir: a utility to convert VCF files into HL7 FHIR format for genomics-EHR integrationabstractBACKGROUND: VCF formatted files are the lingua franca of next-generation sequencing, whereas HL7 FHIR is emerging as a standard language for electronic health record interoperability. A growing number of FHIR-based clinical genomics applications are emerging. Here, we describe an open source utility for converting variants from VCF format into HL7 FHIR format. RESULTS: vcf2fhir converts VCF variants into a FHIR Genomics Diagnostic Report. Conversion translates each VCF row into a corresponding FHIR-formatted variant in the generated report. In scope are simple variants (SNVs, MNVs, Indels), along with zygosity and phase relationships, for autosomes, sex chromosomes, and mitochondrial DNA. Input parameters include VCF file and genome build ('GRCh37' or 'GRCh38'); and optionally a conversion region that indicates the region(s) to convert, a studied region that lists genomic regions studied by the lab, and a non-callable region that lists studied regions deemed uncallable by the lab. Conversion can be limited to a subset of VCF by supplying genomic coordinates of the conversion region(s). If studied and non-callable regions are also supplied, the output FHIR report will include 'region-studied' observations that detail which portions of the conversion region were studied, and of those studied regions, which portions were deemed uncallable. We illustrate the vcf2fhir utility via two case studies. The first, 'SMART Cancer Navigator', is a web application that offers clinical decision support by linking patient EHR information to cancerous gene variants. The second, 'Precision Genomics Integration Platform', intersects a patient's FHIR-formatted clinical and genomic data with knowledge bases in order to provide on-demand delivery of contextually relevant genomic findings and recommendations to the EHR. CONCLUSIONS: Experience to date shows that the vcf2fhir utility can be effectively woven into clinically useful genomic-EHR integration pipelines. Additional testing will be a critical step towards the clinical validation of this utility, enabling it to be integrated in a variety of real world data flow scenarios. For now, we propose the use of this utility primarily to accelerate FHIR Genomics understanding and to facilitate experimentation with further integration of genomics data into the EHR. Robert H. Dolin, Shaileshbhai R. Gothi, Aziz A. Boxwala, Bret S. E. Heale, Ammar Husami, Himanshu Khangar, Shubham Londhe, Frank Naeymi-Rad, Soujanya Rao, Barbara Rapchak, James Shalaby, Varun Suraj, Srikar Chamala, Gil Alterovitz |
BMC Bioinform. | 16 |
| 2021 | Interoperable genetic lab test reports: mapping key data elements to HL7 FHIR specifications and professional reporting guidelinesabstractOBJECTIVE: In many cases, genetic testing labs provide their test reports as portable document format files or scanned images, which limits the availability of the contained information to advanced informatics solutions, such as automated clinical decision support systems. One of the promising standards that aims to address this limitation is Health Level Seven International (HL7) Fast Healthcare Interoperability Resources Clinical Genomics Implementation Guide-Release 1 (FHIR CG IG STU1). This study aims to identify various data content of some genetic lab test reports and map them to FHIR CG IG specification to assess its coverage and to provide some suggestions for standard development and implementation. MATERIALS AND METHODS: We analyzed sample reports of 4 genetic tests and relevant professional reporting guidelines to identify their key data elements (KDEs) that were then mapped to FHIR CG IG. RESULTS: We identified 36 common KDEs among the analyzed genetic test reports, in addition to other unique KDEs for each genetic test. Relevant suggestions were made to guide the standard implementation and development. DISCUSSION AND CONCLUSION: The FHIR CG IG covers the majority of the identified KDEs. However, we suggested some FHIR extensions that might better represent some KDEs. These extensions may be relevant to FHIR implementations or future FHIR updates.The FHIR CG IG is an excellent step toward the interoperability of genetic lab test reports. However, it is a work-in-progress that needs informative and continuous input from the clinical genetics' community, specifically professional organizations, systems implementers, and genetic knowledgebase providers. Aly Khalifa, Clinton C. Mason, Jennifer H. Garvin, Marc S. Williams, Guilherme Del Fiol, Brian R. Jackson, Steven B. Bleyl, Gil Alterovitz, Stanley M. Huff |
J. Am. Medical Informatics Assoc. | 8 |
| 2021 | An explainable machine learning platform for pyrazinamide resistance prediction and genetic feature identification of Mycobacterium tuberculosisabstractOBJECTIVE: Tuberculosis is the leading cause of death from a single infectious agent. The emergence of antimicrobial resistant Mycobacterium tuberculosis strains makes the problem more severe. Pyrazinamide (PZA) is an important component for short-course treatment regimens and first- and second-line treatment regimens. This research aims for fast diagnosis of M. tuberculosis resistance to PZA and identification of genetic features causing resistance. MATERIALS AND METHODS: We use clinically collected genomic data of M. tuberculosis that are resistant or susceptible to PZA. A machine learning platform is built to diagnose PZA resistance using the whole genome sequence data, and to identify resistance genes and mutations. The platform consists of a deep convolutional neural network (DCNN) model for resistance diagnosis and a support vector machine (SVM) model as a surrogate to identify resistance genes and mutations. RESULTS: The DCNN model achieves a PZA resistance diagnosis accuracy of 93%. Each prediction takes less than a second. The SVM has revealed 2 novel genes, embB and gyrA, besides the well-known pncA gene, and 9 mutations that harbor PZA resistance. DISCUSSION: The DCNN and SVM machine learning platform, if used together with the real-time genome sequencing machines, could allow for rapid PZA diagnosis, allowing for critical time to ensure good patient outcomes, and preventing outbreaks of deadly infections. Furthermore, identifying pertinent resistance genes and mutations will help researchers better understand the biological mechanisms behind resistance. CONCLUSIONS: Machine learning can be used to achieve high-accuracy resistance prediction, and identify genes and mutations causing the resistance. Ling Teng, Gil Alterovitz |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | Personal Cognitive Health Library
Ning An 0001, Huitong Ding, Gil Alterovitz |
AMIA | 6 |
| 2020 | Recommendations for patient similarity classes: results of the AMIA 2019 workshop on defining patient similarityabstractDefining patient-to-patient similarity is essential for the development of precision medicine in clinical care and research. Conceptually, the identification of similar patient cohorts appears straightforward; however, universally accepted definitions remain elusive. Simultaneously, an explosion of vendors and published algorithms have emerged and all provide varied levels of functionality in identifying patient similarity categories. To provide clarity and a common framework for patient similarity, a workshop at the American Medical Informatics Association 2019 Annual Meeting was convened. This workshop included invited discussants from academics, the biotechnology industry, the FDA, and private practice oncology groups. Drawing from a broad range of backgrounds, workshop participants were able to coalesce around 4 major patient similarity classes: (1) feature, (2) outcome, (3) exposure, and (4) mixed-class. This perspective expands into these 4 subtypes more critically and offers the medical informatics community a means of communicating their work on this important topic. Nathan D. Seligson, Jeremy L. Warner, William S. Dalton, Robert S. Miller, Debra Patt, Kenneth L. Kehl, Matvey Palchuk, Gil Alterovitz, Laura K. Wiley, Ming Huang 0006, Feichen Shen, Yanshan Wang, Khoa A. Nguyen, Anthony F. Wong, Funda Meric-Bernstam, Elmer V. Bernstam, James L. Chen |
J. Am. Medical Informatics Assoc. | 9 |
| 2020 | An efficient causal structure learning algorithm for linear arbitrarily distributed continuous data
Jing Yang 0008, Ning An 0001, Yu Chen 0018, Gil Alterovitz |
J. Supercomput. | 5 |
| 2019 | Sparse online collaborative filtering with dynamic regularization
Kangkang Li 0001, Xiuze Zhou, Fan Lin, Wenhua Zeng, Beizhan Wang, Gil Alterovitz |
Inf. Sci. | 6 |
| 2018 | Subtype dependent biomarker identification and tumor classification from gene expression profiles
Aiguo Wang 0002, Ning An 0001, Guilin Chen, Li Liu 0001, Gil Alterovitz |
Knowl. Based Syst. | 5 |
| 2017 | Nonlinear dimensionality reduction methods for synthetic biology biobricks' visualizationabstractBACKGROUND: Visualizing data by dimensionality reduction is an important strategy in Bioinformatics, which could help to discover hidden data properties and detect data quality issues, e.g. data noise, inappropriately labeled data, etc. As crowdsourcing-based synthetic biology databases face similar data quality issues, we propose to visualize biobricks to tackle them. However, existing dimensionality reduction methods could not be directly applied on biobricks datasets. Hereby, we use normalized edit distance to enhance dimensionality reduction methods, including Isomap and Laplacian Eigenmaps. RESULTS: By extracting biobricks from synthetic biology database Registry of Standard Biological Parts, six combinations of various types of biobricks are tested. The visualization graphs illustrate discriminated biobricks and inappropriately labeled biobricks. Clustering algorithm K-means is adopted to quantify the reduction results. The average clustering accuracy for Isomap and Laplacian Eigenmaps are 0.857 and 0.844, respectively. Besides, Laplacian Eigenmaps is 5 times faster than Isomap, and its visualization graph is more concentrated to discriminate biobricks. CONCLUSIONS: By combining normalized edit distance with Isomap and Laplacian Eigenmaps, synthetic biology biobircks are successfully visualized in two dimensional space. Various types of biobricks could be discriminated and inappropriately labeled biobricks could be determined, which could help to assess crowdsourcing-based synthetic biology databases' quality, and make biobricks selection. Jiaoyun Yang, Huitong Ding, Ning An 0001, Gil Alterovitz |
BMC Bioinform. | 5 |
| 2016 | Gene expression prediction using low-rank matrix completionabstractBACKGROUND: An exponential growth of high-throughput biological information and data has occurred in the past decade, supported by technologies, such as microarrays and RNA-Seq. Most data generated using such methods are used to encode large amounts of rich information, and determine diagnostic and prognostic biomarkers. Although data storage costs have reduced, process of capturing data using aforementioned technologies is still expensive. Moreover, the time required for the assay, from sample preparation to raw value measurement is excessive (in the order of days). There is an opportunity to reduce both the cost and time for generating such expression datasets. RESULTS: We propose a framework in which complete gene expression values can be reliably predicted in-silico from partial measurements. This is achieved by modelling expression data as a low-rank matrix and then applying recently discovered techniques of matrix completion by using nonlinear convex optimisation. We evaluated prediction of gene expression data based on 133 studies, sourced from a combined total of 10,921 samples. It is shown that such datasets can be constructed with a low relative error even at high missing value rates (>50 %), and that such predicted datasets can be reliably used as surrogates for further analysis. CONCLUSION: This method has potentially far-reaching applications including how bio-medical data is sourced and generated, and transcriptomic prediction by optimisation. We show that gene expression data can be computationally constructed, thereby potentially reducing the costs of gene expression profiling. In conclusion, this method shows great promise of opening new avenues in research on low-rank matrix completion in biological sciences. Arnav Kapur, Kshitij Marwah, Gil Alterovitz |
BMC Bioinform. | 3 |
| 2016 | SMART precision cancer medicine: a FHIR-based app to provide genomic information at the point of careabstractBACKGROUND: Precision cancer medicine (PCM) will require ready access to genomic data within the clinical workflow and tools to assist clinical interpretation and enable decisions. Since most electronic health record (EHR) systems do not yet provide such functionality, we developed an EHR-agnostic, clinico-genomic mobile app to demonstrate several features that will be needed for point-of-care conversations. METHODS: Our prototype, called Substitutable Medical Applications and Reusable Technology (SMART)® PCM, visualizes genomic information in real time, comparing a patient's diagnosis-specific somatic gene mutations detected by PCR-based hotspot testing to a population-level set of comparable data. The initial prototype works for patient specimens with 0 or 1 detected mutation. Genomics extensions were created for the Health Level Seven® Fast Healthcare Interoperability Resources (FHIR)® standard; otherwise, the prototype is a normal SMART on FHIR app. RESULTS: The PCM prototype can rapidly present a visualization that compares a patient's somatic genomic alterations against a distribution built from more than 3000 patients, along with context-specific links to external knowledge bases. Initial evaluation by oncologists provided important feedback about the prototype's strengths and weaknesses. We added several requested enhancements and successfully demonstrated the app at the inaugural American Society of Clinical Oncology Interoperability Demonstration; we have also begun to expand visualization capabilities to include cancer specimens with multiple mutations. DISCUSSION: PCM is open-source software for clinicians to present the individual patient within the population-level spectrum of cancer somatic mutations. The app can be implemented on any SMART on FHIR-enabled EHRs, and future versions of PCM should be able to evolve in parallel with external knowledge bases. Jeremy L. Warner, Matthew J. Rioth, Kenneth D. Mandl, Joshua C. Mandel, David A. Kreda, Isaac S. Kohane, Daniel Carbone, Ross Oreto, Lucy Wang, Shilin Zhu, Heming Yao, Gil Alterovitz |
J. Am. Medical Informatics Assoc. | 12 |
| 2016 | Classification of hospital acquired complications using temporal clinical information from a large electronic health record
Jeremy L. Warner, Peijin Zhang, Jenny Liu, Gil Alterovitz |
J. Biomed. Informatics | 4 |
| 2016 | A Partial Correlation Statistic Structure Learning Algorithm Under Linear Structural Equation ModelsabstractA new algorithm, the Partial Correlation Statistic (PCS) algorithm, is presented for structure learning under linear Structural Equation Models. The PCS algorithm can deal with continuous data following linear arbitrary distribution rather than only a Gaussian distribution. This paper makes two specific contributions. First, for linear arbitrarily distributed datasets, which are generated by the linear structural equation models, if the sample size is sufficiently large, partial correlation coefficient statistic is proved to follow a Student's t-distribution. Second, the PCS algorithm combines hypothesis testing of partial correlation statistic and local learning to select potential neighbors of the target node. This significantly reduces the search space and achieves good time performance. The PCS algorithm does not need to choose optimal threshold of partial correlation by large amount of experiments. Especially, the PCS algorithm redefines the relevance from statistic theory and measure the relevance of the variables based on$p$-value. The effectiveness of the algorithm is compared with current state of the art methods on seven networks. A simulation shows that the PCS algorithm outperforms existing algorithms in terms of both accuracy and time performance on average. Jing Yang 0008, Ning An 0001, Gil Alterovitz |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Predicting hypertension without measurement: A non-invasive, questionnaire-based approach
Aiguo Wang 0002, Ning An 0001, Guilin Chen, Lian Li 0001, Gil Alterovitz |
Expert Syst. Appl. | 5 |
| 2015 | SMART on FHIR Genomics: facilitating standardized clinico-genomic appsabstractBACKGROUND: Supporting clinical decision support for personalized medicine will require linking genome and phenome variants to a patient's electronic health record (EHR), at times on a vast scale. Clinico-genomic data standards will be needed to unify how genomic variant data are accessed from different sequencing systems. METHODS: A specification for the basis of a clinic-genomic standard, building upon the current Health Level Seven International Fast Healthcare Interoperability Resources (FHIR®) standard, was developed. An FHIR application protocol interface (API) layer was attached to proprietary sequencing platforms and EHRs in order to expose gene variant data for presentation to the end-user. Three representative apps based on the SMART platform were built to test end-to-end feasibility, including integration of genomic and clinical data. RESULTS: Successful design, deployment, and use of the API was demonstrated and adopted by HL7 Clinical Genomics Workgroup. Feasibility was shown through development of three apps by various types of users with background levels and locations. CONCLUSION: This prototyping work suggests that an entirely data (and web) standards-based approach could prove both effective and efficient for advancing personalized medicine. Gil Alterovitz, Jeremy L. Warner, Peijin Zhang, Yishen Chen, Mollie Ullman-Cullere, David A. Kreda, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 1 |
| 2015 | Seeing the forest through the trees: uncovering phenomic complexity through interactive network visualizationabstractOur aim was to uncover unrecognized phenomic relationships using force-based network visualization methods, based on observed electronic medical record data. A primary phenotype was defined from actual patient profiles in the Multiparameter Intelligent Monitoring in Intensive Care II database. Network visualizations depicting primary relationships were compared to those incorporating secondary adjacencies. Interactivity was enabled through a phenotype visualization software concept: the Phenomics Advisor. Subendocardial infarction with cardiac arrest was demonstrated as a sample phenotype; there were 332 primarily adjacent diagnoses, with 5423 relationships. Primary network visualization suggested a treatment-related complication phenotype and several rare diagnoses; re-clustering by secondary relationships revealed an emergent cluster of smokers with the metabolic syndrome. Network visualization reveals phenotypic patterns that may have remained occult in pairwise correlation analysis. Visualization of complex data, potentially offered as point-of-care tools on mobile devices, may allow clinicians and researchers to quickly generate hypotheses and gain deeper understanding of patient subpopulations. Jeremy L. Warner, Joshua C. Denny, David A. Kreda, Gil Alterovitz |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | Accelerating wrapper-based feature selection with K-nearest-neighbor
Aiguo Wang 0002, Ning An 0001, Guilin Chen, Lian Li 0001, Gil Alterovitz |
Knowl. Based Syst. | 5 |
| 2014 | Accelerating incremental wrapper based gene selection with K-Nearest-NeighborabstractWrapper based gene selection methods tend to obtain better classification accuracy than filter methods, while it is much more time consuming. Accelerating this process without degrading the high accuracy is of great value for researchers to better analyze gene expression profiles. In this paper, we explore to reduce the time complexity of wrapper based gene selection method with K-Nearest-Neighbor (KNN) classifier embedded. Instead of taking KNN as a black box, we incrementally construct and maintain a classifier distance matrix to speed up the gene selection process. Experiments on six publicly available microarrays were first conducted to show the effectiveness of incremental wrapper based gene selection method with KNN. Then, to demonstrate the performance gain in time cost reduction, we analyzed the time complexity and experimentally evaluated it. Both theoretical analysis and experimental results prove that the proposed approach greatly accelerates the gene selection process without degrading the classification accuracy. Aiguo Wang 0002, Ning An 0001, Guilin Chen, Lian Li 0001, Gil Alterovitz |
BIBM | 5 |
| 2014 | Incremental wrapper based gene selection with Markov blanketabstractGene selection plays a crucial role in the analysis of microarray data with high dimensionality and small sample size. Incremental wrapper based feature subset selection (FSS) methods, among various feature selection approaches, tend to obtain high quality feature subset and better classification accuracy than filter methods, while it is much more time consuming since the interdependence and redundancy between features is evaluated in a wrapper way. In this paper, we explore to introduce Markov Blanket (MB) into incremental wrapper based FSS process. Rather than evaluate the quality of all the features ranked by a filter method, our proposal eliminates features that are redundant to the newly selected one via MB during the wrapper evaluation process to reduce the number of wrappers, enabling us to select the relevant features and eliminate redundant ones efficiently. To verify the effectiveness and efficiency of the proposed approach, experimental comparisons on six publicly available microarray data are conducted with two typical classifiers with different metrics, Naïve Bayes and 1-Nearest-Neighbor. Experimental results demonstrate that our approach greatly speeds up the feature selection process, obtains more compact feature subset and achieves better classification accuracy compared to that without MB for both two-category and multi-category problems. Aiguo Wang 0002, Ning An 0001, Guilin Chen, Jing Yang 0008, Lian Li 0001, Gil Alterovitz |
BIBM | 6 |
| 2013 | Constructing a Novel Cancer Ontology
Michael C. Gao, Jeremy L. Warner, Peter C. Yang, Gil Alterovitz |
AMIA | 4 |
| 2013 | Sharing of Genomic Information: Perspectives from Stakeholders
Jeremy L. Warner, Gil Alterovitz, Joshua C. Denny, Robert Fassett, Kevin S. Hughes |
AMIA | 2 |
| 2013 | Phenometric analysis of electronic health records: a new approach to visualization of high dimensional biomedical information
Jeremy L. Warner, Quan Ding, David A. Kreda, Zi'ou Zheng, Joshua C. Denny, Gil Alterovitz |
AMIA | 7 |
| 2013 | Causal discovery based on healthcare informationabstractCorrectly discovering causal relations from healthcare information can help people to understand disease mechanisms and discover disease causes. In some cases, the healthcare data do not follow a multivariate Gaussian distribution. We design a new causal structure learning algorithm. The algorithm can effectively combines ideas from local learning with simultaneous equations models techniques. In the first phase of the algorithm we select potential neighbors for each variable based on simultaneous equations models, and then perform a constrained hill-climbing search to orient the edges. Using the algorithm without prior knowledge, we analyze causal relations in the real data from the National Health and Nutrition Examination Survey. Jing Yang 0008, Ning An 0001, Gil Alterovitz, Lian Li 0001, Aiguo Wang 0002 |
BIBM | 3 |
| 2013 | Brief communication: External phenome analysis enables a rational federated query strategy to detect changing rates of treatment-related complications associated with multiple myelomaabstractElectronic health records (EHRs) are increasingly useful for health services research. For relatively uncommon conditions, such as multiple myeloma (MM) and its treatment-related complications, a combination of multiple EHR sources is essential for such research. The Shared Health Research Information Network (SHRINE) enables queries for aggregate results across participating institutions. Development of a rational search strategy in SHRINE may be augmented through analysis of pre-existing databases. We developed a SHRINE query for likely non-infectious treatment-related complications of MM, based upon an analysis of the Multiparameter Intelligent Monitoring in Intensive Care (MIMIC II) database. Using this query strategy, we found that the rate of likely treatment-related complications significantly increased from 2001 to 2007, by an average of 6% a year (p=0.01), across the participating SHRINE institutions. This finding is in keeping with increasingly aggressive strategies in the treatment of MM. This proof of concept demonstrates that a staged approach to federated queries, using external EHR data, can yield potentially clinically meaningful results. Jeremy L. Warner, Gil Alterovitz, Kelly Bodio, Robin M. Joyce |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Phenome-Based Analysis as a Means for Discovering Context-Dependent Clinical Reference Ranges
Jeremy L. Warner, Gil Alterovitz |
AMIA | 2 |
| 2012 | Reverse engineering biomolecular systems using -omic data: challenges, progress and opportunitiesabstractRecent advances in high-throughput biotechnologies have led to the rapid growing research interest in reverse engineering of biomolecular systems (REBMS). 'Data-driven' approaches, i.e. data mining, can be used to extract patterns from large volumes of biochemical data at molecular-level resolution while 'design-driven' approaches, i.e. systems modeling, can be used to simulate emergent system properties. Consequently, both data- and design-driven approaches applied to -omic data may lead to novel insights in reverse engineering biological systems that could not be expected before using low-throughput platforms. However, there exist several challenges in this fast growing field of reverse engineering biomolecular systems: (i) to integrate heterogeneous biochemical data for data mining, (ii) to combine top-down and bottom-up approaches for systems modeling and (iii) to validate system models experimentally. In addition to reviewing progress made by the community and opportunities encountered in addressing these challenges, we explore the emerging field of synthetic biology, which is an exciting approach to validate and analyze theoretical system models directly through experimental synthesis, i.e. analysis-by-synthesis. The ultimate goal is to address the present and future challenges in reverse engineering biomolecular systems (REBMS) using integrated workflow of data mining, systems modeling and synthetic biology. Chang F. Quo, Chanchala Kaddi, John H. Phan, Amin Zollanvari, Mingqing Xu, May D. Wang, Gil Alterovitz |
Briefings Bioinform. | 7 |
| 2010 | The challenges of informatics in synthetic biology: from biomolecular networks to artificial organismsabstractThe field of synthetic biology holds an inspiring vision for the future; it integrates computational analysis, biological data and the systems engineering paradigm in the design of new biological machines and systems. These biological machines are built from basic biomolecular components analogous to electrical devices, and the information flow among these components requires the augmentation of biological insight with the power of a formal approach to information management. Here we review the informatics challenges in synthetic biology along three dimensions: in silico, in vitro and in vivo. First, we describe state of the art of the in silico support of synthetic biology, from the specific data exchange formats, to the most popular software platforms and algorithms. Next, we cast in vitro synthetic biology in terms of information flow, and discuss genetic fidelity in DNA manipulation, development strategies of biological parts and the regulation of biomolecular networks. Finally, we explore how the engineering chassis can manipulate biological circuitries in vivo to give rise to future artificial organisms. Gil Alterovitz, Taro Muso, Marco Ramoni |
Briefings Bioinform. | 1 |
| 2010 | Mapping transcription mechanisms from multimodal genomic dataabstractBACKGROUND: Identification of expression quantitative trait loci (eQTLs) is an emerging area in genomic study. The task requires an integrated analysis of genome-wide single nucleotide polymorphism (SNP) data and gene expression data, raising a new computational challenge due to the tremendous size of data. RESULTS: We develop a method to identify eQTLs. The method represents eQTLs as information flux between genetic variants and transcripts. We use information theory to simultaneously interrogate SNP and gene expression data, resulting in a Transcriptional Information Map (TIM) which captures the network of transcriptional information that links genetic variations, gene expression and regulatory mechanisms. These maps are able to identify both cis- and trans- regulating eQTLs. The application on a dataset of leukemia patients identifies eQTLs in the regions of the GART, PCP4, DSCAM, and RIPK4 genes that regulate ADAMTS1, a known leukemia correlate. CONCLUSIONS: The information theory approach presented in this paper is able to infer the dependence networks between SNPs and transcripts, which in turn can identify cis- and trans-eQTLs. The application of our method to the leukemia study explains how genetic variants and gene expression are linked to leukemia. Hsun-Hsien Chang, Michael J. McGeachie, Gil Alterovitz, Marco Ramoni |
BMC Bioinform. | 3 |
| 2010 | Introduction to the special issue on information theory in molecular biology and neuroscienceabstractInformation theory--a field at the intersection of applied mathematics and electrical engineering--was primarily developed for the purpose of addressing problems arising in data storage and data transmission over (noisy) communication media. Consequently, information theory provides the formal basis for much of today’s storage and communication infrastructure. Olgica Milenkovic, Gil Alterovitz, Gerard Battail, Todd P. Coleman, Joachim Hagenauer, Sean P. Meyn, Nathan D. Price 0001, Marco Ramoni, Ilya Shmulevich, Wojciech Szpankowski |
IEEE Trans. Inf. Theory | 2 |
| 2008 | An information theoretic framework for genomic data analysisabstractThe breadth of biological data collected in the last decade has far outstripped the methods available to process it. To effectively investigate and explore this abundance of data, novel automated collection and analysis approaches must be devised. We have developed a new open software framework, the Open Genomic Analysis Platform (OGAP), to aid in the analysis of genomic data. It is capable of analyzing a variety of data source, and focuses on using information theory to characterize data. The frameworks has is capable of import a variety of genome tied data, and provides custom analysis and visualization of results. We then demonstrate the use of this framework analyzing the Prochlorococcus Marinus organism. We show a strong correlation between the information content of sequence data and up regulation of gene expression during lytic infection. Aaron McKenna, Gil Alterovitz |
BIBE | 2 |
| 2008 | Automated programming for bioinformatics algorithm deploymentabstractUNLABELLED: Many bioinformatics solutions suffer from the lack of usable interface/platform from which results can be analyzed and visualized. Overcoming this hurdle would allow for more widespread dissemination of bioinformatics algorithms within the biological and medical communities. The algorithms should be accessible without extensive technical support or programming knowledge. Here, we propose a dynamic wizard platform that provides users with a Graphical User Interface (GUI) for most Java bioinformatics library toolkits. The application interface is generated in real-time based on the original source code. This platform lets developers focus on designing algorithms and biologists/physicians on testing hypotheses and analyzing results. AVAILABILITY: The open source code can be downloaded from: http://bcl.med.harvard.edu/proteomics/proj/APBA/. Gil Alterovitz, Adnaan Jiwaji, Marco Ramoni |
Bioinform. | 1 |
| 2007 | Linking Protein Mass with Function via Organismal Massome NetworksabstractWith the human genome sequenced, attention has been shifting to proteins and their function. Several technologies including mass spectrometry and gel electrophoresis have traditionally been used to study proteins. These technologies rely on proteins' masses to characterize and/or identify them. Once identified, the discovered proteins' are often analyzed for their functional relevance. In this paper, we present an alternative approach to studying protein function directly, based on protein mass. By analyzing proteins with similar mass ranges (bins), we discovered that certain biological functions are associated with specific mass bins more frequently than would be expected by chance. Biological functional classes found in this manner could be seen across the three examined organisms: human, worm and fly. We found no sequence-based homology across significant mass-function proteins in the three different organisms. This investigation leads to useful property that the mass-based biological functional classes are preserved across the organisms. Thus, this paper describes a potential constraint for evolution of similar function based on mass constraints. Finally, this work yields a roadmap for experimental design for functional exploration of proteomes in mass-based technologies such as gel electrophoresis and mass spectrometry. Gil Alterovitz, Eugenia Lyashenko, Michael Xiang, Marco Ramoni |
BIBE | 1 |
| 2007 | A Systematic Approach to Quantifying Evolutionary Functional Trends Across the Universal Tree of LifeabstractThe accumulation of genomic and proteomic data of many organisms presents an opportunity to analyze entire phylogenetic trees in a systematic, quantified manner. The universal tree of life, constructed by genomic data, provides an evolutionary context for proteomic data of individual organisms. Using proteomic information mapped to biological functions, we survey the tree of life for evolutionary trends, where trends are low p-value linear regressions of relative protein diversity against an evolutionary timeline. These trends provide quantified perspectives across evolution that would be otherwise difficult to discern from the large amount of available bioinformatic data. We analyzed 22,762 functions across 167 proteomes for approximately 3.8 million information content calculations. Of particular interest is the rise of protein-protein interactions along the path of mammalian evolution. Of general interest are the numerous uncovered trends that open the door to future queries of focused scope. Gil Alterovitz, Taro Muso, Paresh Malalur, Marco Ramoni |
BIBE | 1 |
| 2006 | Discovering Biological Guilds through Topological Abstraction
Gil Alterovitz, Marco Ramoni |
AMIA | 1 |