Pavel Skums

dblp:95/3105 · also P. V. Skums · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0003-4007-5624ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 8 first-author · 12 since 2021
YearPublicationVenuePosition
2026 Phylogenetic inference under the balanced minimum evolution criterion via semidefinite programming
abstract
MOTIVATION: In this study, we investigate the application of Semidefinite Programming (SDP) to phylogenetics. SDP is a powerful optimization framework that seeks to optimize a linear objective function over the cone of positive semidefinite matrices. As a convex optimization problem, SDP generalizes linear programming and provides relaxations for many combinatorial optimization problems. However, despite its many applications, SDP remains largely unused in computational biology. RESULTS: We show how SDP relaxations can be designed and used for phylogenetic inference. We consider the Balanced Minimum Evolution (BME) problem, a widely used model in distance-based phylogenetics, and introduce an algorithm that combines an SDP relaxation with a rounding scheme that iteratively converts relaxed solutions into valid tree topologies. Experiments on simulated and empirical datasets show that the method enables accurate phylogenetic reconstruction. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://github.com/compbel/SDPTree (DOI 10.5281/zenodo.20838318).
Pavel Skums
Bioinform.1
2025 Editorial: Special Section on Computational Advances in Bio and Medical Sciences
Ion I. Mandoiu, Marmar Moussa, Sanguthevar Rajasekaran, Pavel Skums, Sharma V. Thankachan, Alex Zelikovsky
IEEE Trans. Comput. Biol. Bioinform.4
2024 Community Structure and Temporal Dynamics of Viral Epistatic Networks Allow for Early Detection of Emerging Variants with Altered Phenotypes
Fatemeh Mohebbi, Alex Zelikovsky, Serghei Mangul, Gerardo Chowell, Pavel Skums
RECOMB5
2023 Graph-Based Motif Discovery in Mimotope Profiles of Serum Antibody Repertoire
Hossein Saghaian, Pavel Skums, Yurij Ionov, Alex Zelikovsky
ISBRA2
2023 Guest Editors' Introduction to the Special Section on Bioinformatics Research and Applications
Zhipeng Cai 0001, Min Li 0007, Pavel Skums, Yanjie Wei
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 SOPHIE: Viral Outbreak Investigation and Transmission History Reconstruction in a Joint Phylogenetic and Network Theory Framework
Pavel Skums, Fatemeh Mohebbi, Vyacheslav Tsyvina, Pelin Icer, Sumathi Ramachandran, Yuri Khudyakov
RECOMB1
2022 Primary case inference in viral outbreaks through analysis of intra-host variant population
abstract
BACKGROUND: Investigation of outbreaks to identify the primary case is crucial for the interruption and prevention of transmission of infectious diseases. These individuals may have a higher risk of participating in near future transmission events when compared to the other patients in the outbreak, so directing more transmission prevention resources towards these individuals is a priority. Although the genetic characterization of intra-host viral populations can aid the identification of transmission clusters, it is not trivial to determine the directionality of transmissions during outbreaks, owing to complexity of viral evolution. Here, we present a new computational framework, PYCIVO: primary case inference in viral outbreaks. This framework expands upon our earlier work in development of QUENTIN, which builds a probabilistic disease transmission tree based on simulation of evolution of intra-host hepatitis C virus (HCV) variants between cases involved in direct transmission during an outbreak. PYCIVO improves upon QUENTIN by also adding a custom heterogeneity index and identifying the scenario when the primary case may have not been sampled. RESULTS: These approaches were validated using a set of 105 sequence samples from 11 distinct HCV transmission clusters identified during outbreak investigations, in which the primary case was epidemiologically verified. Both models can detect the correct primary case in 9 out of 11 transmission clusters (81.8%). However, while QUENTIN issues erroneous predictions on the remaining 2 transmission clusters, PYCIVO issues a null output for these clusters, giving it an effective prediction accuracy of 100%. To further evaluate accuracy of the inference, we created 10 modified transmission clusters in which the primary case had been removed. In this scenario, PYCIVO was able to correctly identify that there was no primary case in 8/10 (80%) of these modified clusters. This model was validated with HCV; however, this approach may be applicable to other microbial pathogens. CONCLUSIONS: PYCIVO improves upon QUENTIN by also implementing a custom heterogeneity index which empowers PYCIVO to make the important 'No primary case' prediction. One or more samples, possibly including the primary case, may have not been sampled, and this designation is meant to account for these scenarios.
Walker Gussler, David S. Campo, Zoya Dimitrova, Pavel Skums, Yuri Khudyakov
BMC Bioinform.4
2022 Guest Editors' Introduction to the Special Section on Bioinformatics Research and Applications
abstract
The papers in this special section were presented at the 15th International Symposium on Bioinformatics Research and Applications (ISBRA 2019), which was held at Technical University of Catalonia, Barcelona, Spain on June 3-6, 2019.
Zhipeng Cai 0001, Min Li 0007, Pavel Skums
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Guest Editors' Introduction to the Special Section on Bioinformatics Research and Applications
abstract
The papers in this special section were presented at the 16th International Symposium on Bioinformatics Research and Applications (ISBRA 2020), which was held virtually, on December 1-4, 2020. The ISBRA symposium provides a forum for the exchange of ideas and results among researchers, developers, and practitioners working on all aspects of Bioinformatics and computational biology and their applications.
Zhipeng Cai 0001, Giri Narasimhan, Pavel Skums
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Computational Approaches to Detect Illicit Drug Ads and Find Vendor Communities Within Social Media Platforms
abstract
The opioid abuse epidemic represents a major public health threat to global populations. The role social media may play in facilitating illicit drug trade is largely unknown due to limited research. However, it is known that social media use among adults in the US is widespread, there is vast capability for online promotion of illegal drugs with delayed or limited deterrence of such messaging, and further, general commercial sale applications provide safeguards for transactions; however, they do not discriminate between legal and illegal sale transactions. These characteristics of the social media environment present challenges to surveillance which is needed for advancing knowledge of online drug markets and the role they play in the drug abuse and overdose deaths. In this paper, we present a computational framework developed to automatically detect illicit drug ads and communities of vendors. The SVM- and CNN- based methods for detecting illicit drug ads, and a matrix factorization based method for discovering overlapping communities have been extensively validated on the large dataset collected from Google+, Flickr and Tumblr. Pilot test results demonstrate that our computational methods can effectively identify illicit drug ads and detect vendor-community with accuracy. These methods hold promise to advance scientific knowledge surrounding the role social media may play in perpetuating the drug abuse epidemic.
Fengpan Zhao, Pavel Skums, Alex Zelikovsky, Eric L. Sevigny, Monica Haavisto Swahn, Sheryl M. Strasser, Yan Huang 0032, Yubao Wu
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 A Novel Network Representation of SARS-CoV-2 Sequencing Data
Sergey Knyazev, Daniel Novikov, Mark Grinshpon, Harman Singh, Ram Ayyala, Varuni Sarwal, Roya Hosseini 0002, Pelin Icer Baykal, Pavel Skums, Ellsworth Campbell, Serghei Mangul, Alex Zelikovsky
ISBRA9
2021 Epidemiological data analysis of viral quasispecies in the next-generation sequencing era
abstract
The unprecedented coverage offered by next-generation sequencing (NGS) technology has facilitated the assessment of the population complexity of intra-host RNA viral populations at an unprecedented level of detail. Consequently, analysis of NGS datasets could be used to extract and infer crucial epidemiological and biomedical information on the levels of both infected individuals and susceptible populations, thus enabling the development of more effective prevention strategies and antiviral therapeutics. Such information includes drug resistance, infection stage, transmission clusters and structures of transmission networks. However, NGS data require sophisticated analysis dealing with millions of error-prone short reads per patient. Prior to the NGS era, epidemiological and phylogenetic analyses were geared toward Sanger sequencing technology; now, they must be redesigned to handle the large-scale NGS datasets and properly model the evolution of heterogeneous rapidly mutating viral populations. Additionally, dedicated epidemiological surveillance systems require big data analytics to handle millions of reads obtained from thousands of patients for rapid outbreak investigation and management. We survey bioinformatics tools analyzing NGS data for (i) characterization of intra-host viral population complexity including single nucleotide variant and haplotype calling; (ii) downstream epidemiological analysis and inference of drug-resistant mutations, age of infection and linkage between patients; and (iii) data collection and analytics in surveillance systems for fast response and control of outbreaks.
Sergey Knyazev, Lauren Hughes, Pavel Skums, Alex Zelikovsky
Briefings Bioinform.3
2020 Inference of mutability landscapes of tumors from single cell sequencing data
abstract
One of the hallmarks of cancer is the extremely high mutability and genetic instability of tumor cells. Inherent heterogeneity of intra-tumor populations manifests itself in high variability of clone instability rates. Analogously to fitness landscapes, the instability rates of clonal populations form their mutability landscapes. Here, we present MULAN (MUtability LANdscape inference), a maximum-likelihood computational framework for inference of mutation rates of individual cancer subclones using single-cell sequencing data. It utilizes the partial information about the orders of mutation events provided by cancer mutation trees and extends it by inferring full evolutionary history and mutability landscape of a tumor. Evaluation of mutation rates on the level of subclones rather than individual genes allows to capture the effects of genomic interactions and epistasis. We estimate the accuracy of our approach and demonstrate that it can be used to study the evolution of genetic instability and infer tumor evolutionary history from experimental data. MULAN is available at https://github.com/compbel/MULAN.
Viachaslau Tsyvina, Alex Zelikovsky, Sagi Snir, Pavel Skums
PLoS Comput. Biol.4
2019 Detecting Illicit Drug Ads in Google+ Using Machine Learning
Fengpan Zhao, Pavel Skums, Alex Zelikovsky, Eric L. Sevigny, Monica Haavisto Swahn, Sheryl M. Strasser, Yubao Wu
ISBRA2
2019 Inference of clonal selection in cancer populations using single-cell sequencing data
abstract
SUMMARY: Intra-tumor heterogeneity is one of the major factors influencing cancer progression and treatment outcome. However, evolutionary dynamics of cancer clone populations remain poorly understood. Quantification of clonal selection and inference of fitness landscapes of tumors is a key step to understanding evolutionary mechanisms driving cancer. These problems could be addressed using single-cell sequencing (scSeq), which provides an unprecedented insight into intra-tumor heterogeneity allowing to study and quantify selective advantages of individual clones. Here, we present Single Cell Inference of FItness Landscape (SCIFIL), a computational tool for inference of fitness landscapes of heterogeneous cancer clone populations from scSeq data. SCIFIL allows to estimate maximum likelihood fitnesses of clone variants, measure their selective advantages and order of appearance by fitting an evolutionary model into the tumor phylogeny. We demonstrate the accuracy our approach, and show how it could be applied to experimental tumor data to study clonal selection and infer evolutionary history. SCIFIL can be used to provide new insight into the evolutionary dynamics of cancer. AVAILABILITY AND IMPLEMENTATION: Its source code is available at https://github.com/compbel/SCIFIL.
Pavel Skums, Viachaslau Tsyvina, Alex Zelikovsky
Bioinform.1
2019 Guest Editors' Introduction to the Special Section on Bioinformatics Research and Applications
abstract
The papers in this special section were presented at the 12th International Symposium on Bioinformatics Research and Application (ISBRA), which was held at Belarusian State University in Minsk, Belarus on June 5-8, 2016.
Ion I. Mandoiu, Pavel Skums, Alex Zelikovsky
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 Predicting Opioid Epidemic by Using Twitter Data
Yubao Wu, Pavel Skums, Alex Zelikovsky, David S. Campo, Xueting Liao
ISBRA2
2018 QUENTIN: reconstruction of disease transmissions from viral quasispecies genomic data
abstract
Motivation: Genomic analysis has become one of the major tools for disease outbreak investigations. However, existing computational frameworks for inference of transmission history from viral genomic data often do not consider intra-host diversity of pathogens and heavily rely on additional epidemiological data, such as sampling times and exposure intervals. This impedes genomic analysis of outbreaks of highly mutable viruses associated with chronic infections, such as human immunodeficiency virus and hepatitis C virus, whose transmissions are often carried out through minor intra-host variants, while the additional epidemiological information often is either unavailable or has a limited use. Results: The proposed framework QUasispecies Evolution, Network-based Transmission INference (QUENTIN) addresses the above challenges by evolutionary analysis of intra-host viral populations sampled by deep sequencing and Bayesian inference using general properties of social networks relevant to infection dissemination. This method allows inference of transmission direction even without the supporting case-specific epidemiological information, identify transmission clusters and reconstruct transmission history. QUENTIN was validated on experimental and simulated data, and applied to investigate HCV transmission within a community of hosts with high-risk behavior. It is available at https://github.com/skumsp/QUENTIN. Contact: [email protected] or [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Pavel Skums, Alex Zelikovsky, Walker Gussler, Zoya Dimitrova, Sergey Knyazev, Igor Mandric, Sumathi Ramachandran, David S. Campo, Deeptanshu Jha, Leonid A. Bunimovich, Elizabeth Costenbader, Connie Sexton, Siobhán O'Connor 0002, Guo-liang Xia, Yuri Khudyakov
Bioinform.1
2018 Fast estimation of genetic relatedness between members of heterogeneous populations of closely related genomic variants
abstract
BACKGROUND: Many biological analysis tasks require extraction of families of genetically similar sequences from large datasets produced by Next-generation Sequencing (NGS). Such tasks include detection of viral transmissions by analysis of all genetically close pairs of sequences from viral datasets sampled from infected individuals or studying of evolution of viruses or immune repertoires by analysis of network of intra-host viral variants or antibody clonotypes formed by genetically close sequences. The most obvious naïeve algorithms to extract such sequence families are impractical in light of the massive size of modern NGS datasets. RESULTS: In this paper, we present fast and scalable k-mer-based framework to perform such sequence similarity queries efficiently, which specifically targets data produced by deep sequencing of heterogeneous populations such as viruses. It shows better filtering quality and time performance when comparing to other tools. The tool is freely available for download at https://github.com/vyacheslav-tsivina/signature-sj CONCLUSION: The proposed tool allows for efficient detection of genetic relatedness between genomic samples produced by deep sequencing of heterogeneous populations. It should be especially useful for analysis of relatedness of genomes of viruses with unevenly distributed variable genomic regions, such as HIV and HCV. For the future we envision, that besides applications in molecular epidemiology the tool can also be adapted to immunosequencing and metagenomics data.
Viachaslau Tsyvina, David S. Campo, Seth Sims, Alex Zelikovsky, Yuri Khudyakov, Pavel Skums
BMC Bioinform.6
2017 Agent-Based in Silico Evolution of HCV Quasispecies
Alexander Artyomenko, Pelin B. Icer, Pavel Skums, Sumathi Ramachandran, Yuri Khudyakov, Alex Zelikovsky
ISBRA3
2017 Modeling the Spread of HIV and HCV Infections Based on Identification and Characterization of High-Risk Communities Using Social Media
Deeptanshu Jha, Pavel Skums, Alex Zelikovsky, Yuri Khudyakov
ISBRA2
2015 Computational framework for next-generation sequencing of heterogeneous viral populations using combinatorial pooling
abstract
MOTIVATION: Next-generation sequencing (NGS) allows for analyzing a large number of viral sequences from infected patients, providing an opportunity to implement large-scale molecular surveillance of viral diseases. However, despite improvements in technology, traditional protocols for NGS of large numbers of samples are still highly cost and labor intensive. One of the possible cost-effective alternatives is combinatorial pooling. Although a number of pooling strategies for consensus sequencing of DNA samples and detection of SNPs have been proposed, these strategies cannot be applied to sequencing of highly heterogeneous viral populations. RESULTS: We developed a cost-effective and reliable protocol for sequencing of viral samples, that combines NGS using barcoding and combinatorial pooling and a computational framework including algorithms for optimal virus-specific pools design and deconvolution of individual samples from sequenced pools. Evaluation of the framework on experimental and simulated data for hepatitis C virus showed that it substantially reduces the sequencing costs and allows deconvolution of viral populations with a high accuracy. AVAILABILITY AND IMPLEMENTATION: The source code and experimental data sets are available at http://alan.cs.gsu.edu/NGS/?q=content/pooling.
Pavel Skums, Alexander Artyomenko, Olga Glebova, Sumathi Ramachandran, Ion I. Mandoiu, David S. Campo, Zoya Dimitrova, Alex Zelikovsky, Yuri Khudyakov
Bioinform.1
2013 Alignment of DNA Mass-Spectral Profiles Using Network Flows
Pavel Skums, Olga Glebova, Alex Zelikovsky, Zoya Dimitrova, David S. Campo, Lilia Ganova-Raeva, Yuri Khudyakov
ISBRA1
2013 Reconstruction of viral population structure from next-generation sequencing data using multicommodity flows
abstract
BACKGROUND: Highly mutable RNA viruses exist in infected hosts as heterogeneous populations of genetically close variants known as quasispecies. Next-generation sequencing (NGS) allows for analysing a large number of viral sequences from infected patients, presenting a novel opportunity for studying the structure of a viral population and understanding virus evolution, drug resistance and immune escape. Accurate reconstruction of genetic composition of intra-host viral populations involves assembling the NGS short reads into whole-genome sequences and estimating frequencies of individual viral variants. Although a few approaches were developed for this task, accurate reconstruction of quasispecies populations remains greatly unresolved. RESULTS: Two new methods, AmpMCF and ShotMCF, for reconstruction of the whole-genome intra-host viral variants and estimation of their frequencies were developed, based on Multicommodity Flows (MCFs). AmpMCF was designed for NGS reads obtained from individual PCR amplicons and ShotMCF for NGS shotgun reads. While AmpMCF, based on covering formulation, identifies a minimal set of quasispecies explaining all observed reads, ShotMCS, based on packing formulation, engages the maximal number of reads to generate the most probable set of quasispecies. Both methods were evaluated on simulated data in comparison to Maximum Bandwidth and ViSpA, previously developed state-of-the-art algorithms for estimating quasispecies spectra from the NGS amplicon and shotgun reads, respectively. Both algorithms were accurate in estimation of quasispecies frequencies, especially from large datasets. CONCLUSIONS: The problem of viral population reconstruction from amplicon or shotgun NGS reads was solved using the MCF formulation. The two methods, ShotMCF and AmpMCF, developed here afford accurate reconstruction of the structure of intra-host viral population from NGS reads. The implementations of the algorithms are available at http://alan.cs.gsu.edu/vira.html (AmpMCF) and http://alan.cs.gsu.edu/NGS/?q=content/shotmcf (ShotMCF).
Pavel Skums, Nicholas Mancuso, Alexander Artyomenko, Bassam Tork, Ion I. Mandoiu, Yuri Khudyakov, Alex Zelikovsky
BMC Bioinform.1
2012 Efficient error correction for next-generation sequencing of viral amplicons
abstract
BACKGROUND: Next-generation sequencing allows the analysis of an unprecedented number of viral sequence variants from infected patients, presenting a novel opportunity for understanding virus evolution, drug resistance and immune escape. However, sequencing in bulk is error prone. Thus, the generated data require error identification and correction. Most error-correction methods to date are not optimized for amplicon analysis and assume that the error rate is randomly distributed. Recent quality assessment of amplicon sequences obtained using 454-sequencing showed that the error rate is strongly linked to the presence and size of homopolymers, position in the sequence and length of the amplicon. All these parameters are strongly sequence specific and should be incorporated into the calibration of error-correction algorithms designed for amplicon sequencing. RESULTS: In this paper, we present two new efficient error correction algorithms optimized for viral amplicons: (i) k-mer-based error correction (KEC) and (ii) empirical frequency threshold (ET). Both were compared to a previously published clustering algorithm (SHORAH), in order to evaluate their relative performance on 24 experimental datasets obtained by 454-sequencing of amplicons with known sequences. All three algorithms show similar accuracy in finding true haplotypes. However, KEC and ET were significantly more efficient than SHORAH in removing false haplotypes and estimating the frequency of true ones. CONCLUSIONS: Both algorithms, KEC and ET, are highly suitable for rapid recovery of error-free haplotypes obtained by 454-sequencing of amplicons from heterogeneous viruses.The implementations of the algorithms and data sets used for their testing are available at: http://alan.cs.gsu.edu/NGS/?q=content/pyrosequencing-error-correction-algorithm.
Pavel Skums, Zoya Dimitrova, David S. Campo, Gilberto Vaughan, Livia Rossi, Joseph C. Forbi, Jonny Yokosawa, Alex Zelikovsky, Yuri Khudyakov
BMC Bioinform.1