Alexander V. Alekseyenko

dblp:71/2787 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
8since 2021 · last 2024
0000-0002-5748-2085ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 9 first-author · 8 since 2021
YearPublicationVenuePosition
2024 Enabling the clinical application of artificial intelligence in genomics: a perspective of the AMIA Genomics and Translational Bioinformatics Workgroup
abstract
OBJECTIVE: Given the importance AI in genomics and its potential impact on human health, the American Medical Informatics Association-Genomics and Translational Biomedical Informatics (GenTBI) Workgroup developed this assessment of factors that can further enable the clinical application of AI in this space. PROCESS: A list of relevant factors was developed through GenTBI workgroup discussions in multiple in-person and online meetings, along with review of pertinent publications. This list was then summarized and reviewed to achieve consensus among the group members. CONCLUSIONS: Substantial informatics research and development are needed to fully realize the clinical potential of such technologies. The development of larger datasets is crucial to emulating the success AI is achieving in other domains. It is important that AI methods do not exacerbate existing socio-economic, racial, and ethnic disparities. Genomic data standards are critical to effectively scale such technologies across institutions. With so much uncertainty, complexity and novelty in genomics and medicine, and with an evolving regulatory environment, the current focus should be on using these technologies in an interface with clinicians that emphasizes the value each brings to clinical decision-making.
Nephi Walton, Radhakrishnan Nagarajan, Chen Wang 0001, Murat Sincan, Robert R. Freimuth, David B. Everman, Derek C. Walton, Scott McGrath, Dominick J. Lemas, Panayiotis V. Benos, Alexander V. Alekseyenko, Qianqian Song 0002, Ece D. Gamsiz Uzun, Casey Overby Taylor, Alper Uzun, Thomas N. Person, Nadav Rappoport, Zhongming Zhao, Marc S. Williams
J. Am. Medical Informatics Assoc.11
2023 Not all phenotypes are created equal: covariates of success in e-phenotype specification
abstract
BACKGROUND: Electronic (e)-phenotype specification by noninformaticist investigators remains a challenge. Although validation of each patient returned by e-phenotype could ensure accuracy of cohort representation, this approach is not practical. Understanding the factors leading to successful e-phenotype specification may reveal generalizable strategies leading to better results. MATERIALS AND METHODS: Noninformaticist experts (n = 21) were recruited to produce expert-mediated e-phenotypes using i2b2 assisted by a honest data-broker and a project coordinator. Patient- and visit-sets were reidentified and a random sample of 20 charts matching each e-phenotype was returned to experts for chart-validation. Attributes of the queries and expert characteristics were captured and related to chart-validation rates using generalized linear regression models. RESULTS: E-phenotype validation rates varied according to experts' domains and query characteristics (mean = 61%, range 20-100%). Clinical domains that performed better included infectious, rheumatic, neonatal, and cancers, whereas other domains performed worse (psychiatric, GI, skin, and pulmonary). Match-rate was negatively impacted when specification of temporal constraints was required. In general, the increase in e-phenotype specificity contributed positively to match-rate. DISCUSSIONS AND CONCLUSIONS: Clinical experts and informaticists experience a variety of challenges when building e-phenotypes, including the inability to differentiate clinical events from patient characteristics or appropriately configure temporal constraints; a lack of access to available and quality data; and difficulty in specifying routes of medication administration. Biomedical query mediation by informaticists and honest data-brokers in designing e-phenotypes cannot be overstated. Although tools such as i2b2 may be widely available to noninformaticists, successful utilization depends not on users' confidence, but rather on creating highly specific e-phenotypes.
Bashir Hamidi, Patrick A. Flume, Kit N. Simpson, Alexander V. Alekseyenko
J. Am. Medical Informatics Assoc.4
2022 Building Determinants of Health Ontology of Mappable Elements (DHOME)
Lauren Cuppy, Tami Crawford, Nikunja Swain, Alexander V. Alekseyenko
AMIA4
2021 Colonizing Microbiome as a Determinant of COVID-19 Outcome: A Pilot Study
Alexander V. Alekseyenko, Bashir Hamidi, Stéphane M. Meystre
AMIA1
2021 Natural Language Processing and COVID-19 Predictive Analytics to Enable and Optimize SARS-CoV-2 Pooled Testing
Stéphane M. Meystre, Paul M. Heider, Jihad S. Obeid, Alexander V. Alekseyenko, James E. Madory
AMIA4
2021 A phylogenetic approach for weighting genetic sequences
abstract
BACKGROUND: Many important applications in bioinformatics, including sequence alignment and protein family profiling, employ sequence weighting schemes to mitigate the effects of non-independence of homologous sequences and under- or over-representation of certain taxa in a dataset. These schemes aim to assign high weights to sequences that are 'novel' compared to the others in the same dataset, and low weights to sequences that are over-represented. RESULTS: We formalise this principle by rigorously defining the evolutionary 'novelty' of a sequence within an alignment. This results in new sequence weights that we call 'phylogenetic novelty scores'. These scores have various desirable properties, and we showcase their use by considering, as an example application, the inference of character frequencies at an alignment column-important, for example, in protein family profiling. We give computationally efficient algorithms for calculating our scores and, using simulations, show that they are versatile and can improve the accuracy of character frequency estimation compared to existing sequence weighting schemes. CONCLUSIONS: Our phylogenetic novelty scores can be useful when an evolutionarily meaningful system for adjusting for uneven taxon sampling is desired. They have numerous possible applications, including estimation of evolutionary conservation scores and sequence logos, identification of targets in conservation biology, and improving and measuring sequence alignment accuracy.
Nicola De Maio, Alexander V. Alekseyenko, William J. Coleman-Smith, Fabio Pardi, Marc A. Suchard, Asif U. Tamuri, Jakub Truszkowski, Nick Goldman
BMC Bioinform.2
2021 Each patient is a research biorepository: informatics-enabled research on surplus clinical specimens via the living BioBank
abstract
The ability to analyze human specimens is the pillar of modern-day translational research. To enhance the research availability of relevant clinical specimens, we developed the Living BioBank (LBB) solution, which allows for just-in-time capture and delivery of phenotyped surplus laboratory medicine specimens. The LBB is a system-of-systems integrating research feasibility databases in i2b2, a real-time clinical data warehouse, and an informatics system for institutional research services management (SPARC). LBB delivers deidentified clinical data and laboratory specimens. We further present an extension to our solution, the Living µBiome Bank, that allows the user to request and receive phenotyped specimen microbiome data. We discuss the details of the implementation of the LBB system and the necessary regulatory oversight for this solution. The conducted institutional focus group of translational investigators indicates an overall positive sentiment towards potential scientific results generated with the use of LBB. Reference implementation of LBB is available at https://LivingBioBank.musc.edu.
Alexander V. Alekseyenko, Bashir Hamidi, Trevor D. Faith, Keith A. Crandall, Jennifer G. Powers, Christopher Metts, James E. Madory, Steven L. Carroll, Jihad S. Obeid, Leslie Lenert
J. Am. Medical Informatics Assoc.1
2021 Natural language processing enabling COVID-19 predictive analytics to support data-driven patient advising and pooled testing
abstract
OBJECTIVE: The COVID-19 (coronavirus disease 2019) pandemic response at the Medical University of South Carolina included virtual care visits for patients with suspected severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection. The telehealth system used for these visits only exports a text note to integrate with the electronic health record, but structured and coded information about COVID-19 (eg, exposure, risk factors, symptoms) was needed to support clinical care and early research as well as predictive analytics for data-driven patient advising and pooled testing. MATERIALS AND METHODS: To capture COVID-19 information from multiple sources, a new data mart and a new natural language processing (NLP) application prototype were developed. The NLP application combined reused components with dictionaries and rules crafted by domain experts. It was deployed as a Web service for hourly processing of new data from patients assessed or treated for COVID-19. The extracted information was then used to develop algorithms predicting SARS-CoV-2 diagnostic test results based on symptoms and exposure information. RESULTS: The dedicated data mart and NLP application were developed and deployed in a mere 10-day sprint in March 2020. The NLP application was evaluated with good accuracy (85.8% recall and 81.5% precision). The SARS-CoV-2 testing predictive analytics algorithms were configured to provide patients with data-driven COVID-19 testing advices with a sensitivity of 81% to 92% and to enable pooled testing with a negative predictive value of 90% to 91%, reducing the required tests to about 63%. CONCLUSIONS: SARS-CoV-2 testing predictive analytics and NLP successfully enabled data-driven patient advising and pooled testing.
Stéphane M. Meystre, Paul M. Heider, Jihad S. Obeid, James E. Madory, Alexander V. Alekseyenko
J. Am. Medical Informatics Assoc.7
2020 Informatics Architecture for Phenotyped Microbiome Data Generation from Just-in-Time Clinical Microbiology Specimens from CTSA Network Producers
Alexander V. Alekseyenko
AMIA1
2020 Informatics Challenges of COVID-19 Crisis: A Comprehensive Response from An Academic Health System
Alexander V. Alekseyenko, Leslie Lenert
AMIA1
2020 Actionable Opportunities for Improving Opioid Prescribing through Use of Informatics
Neel A. Shimpi, Karmen S. Williams, Alexander V. Alekseyenko
AMIA3
2019 Multivariate Pathway Enrichment Analysis for Interpreting Expression Profiles
Ali Shojaee Bakhtiari, Kristin Wallace, Alexander V. Alekseyenko
AMIA3
2019 Simulation Study of Just-in-Time Specimen Recruitment from University-Wide e-Phenotypes of Interest
Bashir Hamidi, Leslie Lenert, Alexander V. Alekseyenko
AMIA3
2019 Actionable opportunities for improving knowledge at intersectionof dental and medical data
Neel A. Shimpi, Karmen S. Williams, Alexander V. Alekseyenko
AMIA3
2018 The Hidden Microbiome Pipeline: Providing Access to Clinical Microbiome Specimens, Sequences, and Informatics Resources
Bashir Hamidi, Leslie Lenert, Jihad S. Obeid, Alexander V. Alekseyenko
AMIA4
2017 The Impact of Human Microbiome on Precision Medicine
Alexander V. Alekseyenko, Lita M. Proctor, Paul Carlson, Scott A. Jackson, Connie Chen
AMIA1
2017 Local Causal Networks Discover Predictive Cytokine Biomarkers of Scleroderma
Ali Shojaee Bakhtiari, Galina S. Bogatkevich, Alexander V. Alekseyenko
AMIA3
2017 A general framework for association analysis of microbial communities on a taxonomic tree
abstract
Motivation: : Association analysis of microbiome composition with disease-related outcomes provides invaluable knowledge towards understanding the roles of microbes in the underlying disease mechanisms. Proper analysis of sparse compositional microbiome data is challenging. Existing methods rely on strong assumptions on the data structure and fail to pinpoint the associated microbial communities. Results: : We develop a general framework to: (i) perform robust association tests for the microbial community that exhibits arbitrary inter-taxa dependencies; (ii) localize lineages on the taxonomic tree that are associated with covariates (e.g. disease status); and (iii) assess the overall association of the whole microbial community with the covariates. Unlike existing methods for microbiome association analysis, our framework does not make any distributional assumptions on the microbiome data; it allows for the adjustment of confounding variables and accommodates excessive zero observations; and it incorporates taxonomic information. We perform extensive simulation studies under a wide-range of scenarios to evaluate the new methods and demonstrate substantial power gain over existing methods. The advantages of the proposed framework are further demonstrated with real datasets from two microbiome studies. The relevant R package miLineage is publicly available. Availability and Implementation: : miLineage package, manual and tutorial are available at https://medschool.vanderbilt.edu/tang-lab/software/miLineage . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Zheng-Zheng Tang, Guanhua Chen 0002, Alexander V. Alekseyenko, Hongzhe Li
Bioinform.3
2016 Heteroscedastic Omics Data Analysis: Distance-Based Multivariate Approach
Alexander V. Alekseyenko
AMIA1
2016 The Human Microbiome: Informatics Challenges and Opportunities
Alexander V. Alekseyenko, Michael J. Becich, Todd Z. DeSantis, Jack A. Gilbert, Georg K. Gerber
AMIA1
2016 Comprehensive Evaluation of Univariate and Multivariate Distributions for Machine Learning With Microbiome Data
Ali Shojaee Bakhtiari, Alexander V. Alekseyenko
AMIA2
2016 Multivariate Welch t-test on distances
abstract
MOTIVATION: Permutational non-Euclidean analysis of variance, PERMANOVA, is routinely used in exploratory analysis of multivariate datasets to draw conclusions about the significance of patterns visualized through dimension reduction. This method recognizes that pairwise distance matrix between observations is sufficient to compute within and between group sums of squares necessary to form the (pseudo) F statistic. Moreover, not only Euclidean, but arbitrary distances can be used. This method, however, suffers from loss of power and type I error inflation in the presence of heteroscedasticity and sample size imbalances. RESULTS: We develop a solution in the form of a distance-based Welch t-test, [Formula: see text], for two sample potentially unbalanced and heteroscedastic data. We demonstrate empirically the desirable type I error and power characteristics of the new test. We compare the performance of PERMANOVA and [Formula: see text] in reanalysis of two existing microbiome datasets, where the methodology has originated. AVAILABILITY AND IMPLEMENTATION: The source code for methods and analysis of this article is available at https://github.com/alekseyenko/Tw2 Further guidance on application of these methods can be obtained from the author. CONTACT: [email protected].
Alexander V. Alekseyenko
Bioinform.1
2016 PERMANOVA-S: association test for microbial community composition that accommodates confounders and multiple distances
abstract
MOTIVATION: Recent advances in sequencing technology have made it possible to obtain high-throughput data on the composition of microbial communities and to study the effects of dysbiosis on the human host. Analysis of pairwise intersample distances quantifies the association between the microbiome diversity and covariates of interest (e.g. environmental factors, clinical outcomes, treatment groups). In the design of these analyses, multiple choices for distance metrics are available. Most distance-based methods, however, use a single distance and are underpowered if the distance is poorly chosen. In addition, distance-based tests cannot flexibly handle confounding variables, which can result in excessive false-positive findings. RESULTS: We derive presence-weighted UniFrac to complement the existing UniFrac distances for more powerful detection of the variation in species richness. We develop PERMANOVA-S, a new distance-based method that tests the association of microbiome composition with any covariates of interest. PERMANOVA-S improves the commonly-used Permutation Multivariate Analysis of Variance (PERMANOVA) test by allowing flexible confounder adjustments and ensembling multiple distances. We conducted extensive simulation studies to evaluate the performance of different distances under various patterns of association. Our simulation studies demonstrate that the power of the test relies on how well the selected distance captures the nature of the association. The PERMANOVA-S unified test combines multiple distances and achieves good power regardless of the patterns of the underlying association. We demonstrate the usefulness of our approach by reanalyzing several real microbiome datasets. AVAILABILITY AND IMPLEMENTATION: miProfile software is freely available at https://medschool.vanderbilt.edu/tang-lab/software/miProfile CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zheng-Zheng Tang, Guanhua Chen 0002, Alexander V. Alekseyenko
Bioinform.3
2007 Nested Containment List (NCList): a new algorithm for accelerating interval query of genome alignment and interval databases
abstract
MOTIVATION: The exponential growth of sequence databases poses a major challenge to bioinformatics tools for querying alignment and annotation databases. There is a pressing need for methods for finding overlapping sequence intervals that are highly scalable to database size, query interval size, result size and construction/updating of the interval database. RESULTS: We have developed a new interval database representation, the Nested Containment List (NCList), whose query time is O(n + log N), where N is the database size and n is the size of the result set. In all cases tested, this query algorithm is 5-500-fold faster than other indexing methods tested in this study, such as MySQL multi-column indexing, MySQL binning and R-Tree indexing. We provide performance comparisons both in simulated datasets and real-world genome alignment databases, across a wide range of database sizes and query interval widths. We also present an in-place NCList construction algorithm that yields database construction times that are approximately 100-fold faster than other methods available. The NCList data structure appears to provide a useful foundation for highly scalable interval database applications. AVAILABILITY: NCList data structure is part of Pygr, a bioinformatics graph database library, available at http://sourceforge.net/projects/pygr
Alexander V. Alekseyenko, Christopher J. Lee
Bioinform.1