EDBT 2026 Demo / reviewers in the wild / expert
Serghei Mangul
dblp:99/8436
· DBLP profile ↗
14ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0003-4770-3443ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Community Structure and Temporal Dynamics of Viral Epistatic Networks Allow for Early Detection of Emerging Variants with Altered Phenotypes
Fatemeh Mohebbi, Alex Zelikovsky, Serghei Mangul, Gerardo Chowell, Pavel Skums |
RECOMB | 3 |
| 2024 | VISTA: an integrated framework for structural variant discoveryabstractStructural variation (SV) refers to insertions, deletions, inversions, and duplications in human genomes. SVs are present in approximately 1.5% of the human genome. Still, this small subset of genetic variation has been implicated in the pathogenesis of psoriasis, Crohn's disease and other autoimmune disorders, autism spectrum and other neurodevelopmental disorders, and schizophrenia. Since identifying structural variants is an important problem in genetics, several specialized computational techniques have been developed to detect structural variants directly from sequencing data. With advances in whole-genome sequencing (WGS) technologies, a plethora of SV detection methods have been developed. However, dissecting SVs from WGS data remains a challenge, with the majority of SV detection methods prone to a high false-positive rate, and no existing method able to precisely detect a full range of SVs present in a sample. Previous studies have shown that none of the existing SV callers can maintain high accuracy across various SV lengths and genomic coverages. Here, we report an integrated structural variant calling framework, Variant Identification and Structural Variant Analysis (VISTA), that leverages the results of individual callers using a novel and robust filtering and merging algorithm. In contrast to existing consensus-based tools which ignore the length and coverage, VISTA overcomes this limitation by executing various combinations of top-performing callers based on variant length and genomic coverage to generate SV events with high accuracy. We evaluated the performance of VISTA on comprehensive gold-standard datasets across varying organisms and coverage. We benchmarked VISTA using the Genome-in-a-Bottle gold standard SV set, haplotype-resolved de novo assemblies from the Human Pangenome Reference Consortium, along with an in-house polymerase chain reaction (PCR)-validated mouse gold standard set. VISTA maintained the highest F1 score among top consensus-based tools measured using a comprehensive gold standard across both mouse and human genomes. VISTA also has an optimized mode, where the calls can be optimized for precision or recall. VISTA-optimized can attain 100% precision and the highest sensitivity among other variant callers. In conclusion, VISTA represents a significant advancement in structural variant calling, offering a robust and accurate framework that outperforms existing consensus-based tools and sets a new standard for SV detection in genomic research. Varuni Sarwal, Seungmo Lee, Jianzhi Yang, Sriram Sankararaman, Mark Chaisson, Eleazar Eskin, Serghei Mangul |
Briefings Bioinform. | 7 |
| 2023 | Response to 'comment on rigorous benchmarking of T cell receptor repertoire profiling methods for cancer RNA sequencing' by Davydov A.N.; Bolotin D.A.; Poslavsky S. V. and Chudakov D.MabstractHuang et al. reply: Davydov et al. discuss potential pitfalls in the publication ‘Rigorous benchmarking of T cell receptor repertoire profiling methods for cancer RNA sequencing’. Below, we clarified and demonstrated that conclusions cannot be drawn on the basis of their response. Davydov et al. highlighted that our comparison of tools like CATT, TRUST4, and ImRep—which do not consider Phred nucleotide quality in their analyses—against MiXCR was not fully standardized. They suggested that to ensure a fair comparison, the quality filter in MiXCR should also be disabled. We acknowledge the valuable suggestion from the Davydov et al. group regarding the implementation of Phred nucleotide quality filters on the reads before executing MIXCR. However, our primary objective in this publication is to faithfully replicate the typical workflow employed by scientific researchers when utilizing computational tools. Consequently, it would be inequitable to omit the default command when applying various bioinformatics tools to the identical set of samples. Consequently, we opted to employ uniform default criteria across all tools, executing each tool with its default settings. Furthermore, Davydov et al. pinpointed a specific issue related to clonotype abundance calculation in version 1.0.2 of TRUST4. We certainly appreciate the concerns raised by the Davydov et al. group regarding a potential bug in version 1.0.2 of TRUST4 that might affect clonotype abundance calculations. It’s crucial to maintain the highest standards of data integrity, and we take such observations seriously. However, it’s important to note that the version of TRUST4 used in our study was the version available at the time we were preparing and conducting our analyses for publication. Therefore, our methodology was in accordance with best practices available at that time. Davydov et al. pointed out that our evaluation of immune repertoire extraction tools focused solely on quantitative metrics, such as the number of reads containing CDR3 sequences and the number of reported clonotypes. And in doing so, we overlooked the crucial issue of false positive clonotypes, which can originate from regions unrelated to immune receptor genes and have a significant impact on downstream analyses and interpretations. We appreciate Davydov et al. group’s point regarding the potential impact of reads originating from genome regions unrelated to immune receptor genes (Variable (V), Diversity (D), Joining (J), or Constant (C)). The group led by Davydov et al. pointed out that one of the top clonotypes identified by TRUST from the PRJNA812076 samples—specifically, the CDR3 with the amino acid sequence CANTGELFF—is a false positive. However, after careful consideration, we respectfully disagree with the suggestion that our evaluation methodology is incomplete without assessing these aspects. Our focus on quantitative measures, for example the fractions of repertoire captured by RNA-Seq based on clonotypes confirmed by T cell receptor sequencing (TCR-Seq), was a deliberate choice. It’s important to note, that this particular clonotype (CANTGELFF) does not appear in our gold standard TCR-Seq sample. As such, we did not classify it as a true positive in our analysis, thereby eliminating the concern of reporting this false positive clonotype in our results. However, we did acknowledge the limitations of using TCR-Seq as a gold standard in our original publication. We noted instances where TCR clonotypes were identified through RNA-Seq-based techniques but were not detected using TCR-Seq. These discrepancies could either be false positives arising from RNA-Seq methods or false negatives from TCR-Seq—meaning actual TCR clonotypes missed by the TCR-Seq approach. Unfortunately, our benchmarking strategy lacks the capability to differentiate between these two scenarios. While we understand the concern about the possibility of false positives affecting the downstream analysis, we believe our current approach already addresses these issues adequately, and adding more layers of evaluation might dilute the focus of our research. We hope this clarifies our methodology but still appreciate the specific input on the concerns on the false positives. Davydov et al. pointed out that we excluded singletons (clones supported by only one read) from our analysis, noting that the implications of this filtering approach on the overall quality of the results remain ambiguous, particularly in light of the significant presence of false positives. While Davydov et al. advocate for retaining singletons in the analysis, we have a different stance on the matter. Given that TCR-Seq cannot be considered an infallible gold standard, we find it essential to implement specific criteria to bring it closer to serving as an accurate ground truth. Consequently, we identify singletons as likely false positives and opt to exclude them. The procedure of removing singletons serves as a method to enhance the reliability of the results derived from the gold standard. We used default settings for all TCR profiling tools to better reflect how these tools are commonly used in real-world scenarios. While we acknowledged a potential bug in TRUST4 version 1.0.2, we want to emphasize that we diligently adhered to the best practices available during the study. Our deliberate emphasis was on quantitative measures, specifically the fractions of repertoire captured by RNA-Seq, with TCR-Seq serving as our gold standard. To enhance the reliability of our results, we intentionally excluded singletons, which differs from Davydov et al.’s approach. This exclusion contributes to making TCR-Seq a more accurate gold standard for benchmarking purposes. Yu-Ning Huang, Mohammad Vahed, Kerui Peng, Houda Alachkar, Serghei Mangul |
Briefings Bioinform. | 5 |
| 2023 | Rigorous benchmarking of T-cell receptor repertoire profiling methods for cancer RNA sequencingabstractThe ability to identify and track T-cell receptor (TCR) sequences from patient samples is becoming central to the field of cancer research and immunotherapy. Tracking genetically engineered T cells expressing TCRs that target specific tumor antigens is important to determine the persistence of these cells and quantify tumor responses. The available high-throughput method to profile TCR repertoires is generally referred to as TCR sequencing (TCR-Seq). However, the available TCR-Seq data are limited compared with RNA sequencing (RNA-Seq). In this paper, we have benchmarked the ability of RNA-Seq-based methods to profile TCR repertoires by examining 19 bulk RNA-Seq samples across 4 cancer cohorts including both T-cell-rich and T-cell-poor tissue types. We have performed a comprehensive evaluation of the existing RNA-Seq-based repertoire profiling methods using targeted TCR-Seq as the gold standard. We also highlighted scenarios under which the RNA-Seq approach is suitable and can provide comparable accuracy to the TCR-Seq approach. Our results show that RNA-Seq-based methods are able to effectively capture the clonotypes and estimate the diversity of TCR repertoires, as well as provide relative frequencies of clonotypes in T-cell-rich tissues and low-diversity repertoires. However, RNA-Seq-based TCR profiling methods have limited power in T-cell-poor tissues, especially in highly diverse repertoires of T-cell-poor tissues. The results of our benchmarking provide an additional appealing argument to incorporate RNA-Seq into the immune repertoire screening of cancer patients as it offers broader knowledge into the transcriptomic changes that exceed the limited information provided by TCR-Seq. Kerui Peng, Theodore S. Nowicki, Katie Campbell, Mohammad Vahed, Dandan Peng, Yiting Meng, Anish Nagareddy, Yu-Ning Huang, Aaron Karlsberg, Zachary Miller, Jaqueline Joice Brito, Brian B. Nadel, Victoria M. Pak, Malak S. Abedalthagafi, Amanda M. Burkhardt, Houda Alachkar, Antoni Ribas, Serghei Mangul |
Briefings Bioinform. | 18 |
| 2022 | A comprehensive benchmarking of WGS-based deletion structural variant callersabstractAdvances in whole-genome sequencing (WGS) promise to enable the accurate and comprehensive structural variant (SV) discovery. Dissecting SVs from WGS data presents a substantial number of challenges and a plethora of SV detection methods have been developed. Currently, evidence that investigators can use to select appropriate SV detection tools is lacking. In this article, we have evaluated the performance of SV detection tools on mouse and human WGS data using a comprehensive polymerase chain reaction-confirmed gold standard set of SVs and the genome-in-a-bottle variant set, respectively. In contrast to the previous benchmarking studies, our gold standard dataset included a complete set of SVs allowing us to report both precision and sensitivity rates of the SV detection methods. Our study investigates the ability of the methods to detect deletions, thus providing an optimistic estimate of SV detection performance as the SV detection methods that fail to detect deletions are likely to miss more complex SVs. We found that SV detection tools varied widely in their performance, with several methods providing a good balance between sensitivity and precision. Additionally, we have determined the SV callers best suited for low- and ultralow-pass sequencing data as well as for different deletion length categories. Varuni Sarwal, Sebastian Niehus, Ram Ayyala, Aditya Sarkar, Sei Chang, Angela Lu, Neha Rajkumar, Nicholas Darci-Maher, Russell Littman, Karishma Chhugani, Arda Söylev, Zoia Comarova, Emily E. Wesel, Jacqueline Castellanos, Rahul Chikka, Margaret G. Distler, Eleazar Eskin, Jonathan Flint, Serghei Mangul |
Briefings Bioinform. | 20 |
| 2021 | A Novel Network Representation of SARS-CoV-2 Sequencing Data
Sergey Knyazev, Daniel Novikov, Mark Grinshpon, Harman Singh, Ram Ayyala, Varuni Sarwal, Roya Hosseini 0002, Pelin Icer Baykal, Pavel Skums, Ellsworth Campbell, Serghei Mangul, Alex Zelikovsky |
ISBRA | 11 |
| 2021 | Systematic evaluation of transcriptomics-based deconvolution methods and references using thousands of clinical samplesabstractEstimating cell type composition of blood and tissue samples is a biological challenge relevant in both laboratory studies and clinical care. In recent years, a number of computational tools have been developed to estimate cell type abundance using gene expression data. Although these tools use a variety of approaches, they all leverage expression profiles from purified cell types to evaluate the cell type composition within samples. In this study, we compare 12 cell type quantification tools and evaluate their performance while using each of 10 separate reference profiles. Specifically, we have run each tool on over 4000 samples with known cell type proportions, spanning both immune and stromal cell types. A total of 12 of these represent in vitro synthetic mixtures and 300 represent in silico synthetic mixtures prepared using single-cell data. A final 3728 clinical samples have been collected from the Framingham cohort, for which cell populations have been quantified using electrical impedance cell counting. When tools are applied to the Framingham dataset, the tool Estimating the Proportions of Immune and Cancer cells (EPIC) produces the highest correlation, whereas Gene Expression Deconvolution Interactive Tool (GEDIT) produces the lowest error. The best tool for other datasets is varied, but CIBERSORT and GEDIT most consistently produce accurate results. We find that optimal reference depends on the tool used, and report suggested references to be used with each tool. Most tools return results within minutes, but on large datasets runtimes for CIBERSORT can exceed hours or even days. We conclude that deconvolution methods are capable of returning high-quality results, but that proper reference selection is critical. Brian B. Nadel, Meritxell Oliva, Benjamin L. Shou, Keith Mitchell, Feiyang Ma, Dennis J. Montoya, Alice Mouton, Sarah Kim-Hellmuth, Barbara E. Stranger, Matteo Pellegrini, Serghei Mangul |
Briefings Bioinform. | 11 |
| 2016 | HapIso: An Accurate Method for the Haplotype-Specific Isoforms Reconstruction from Long Single-Molecule Reads
Serghei Mangul, Harry (Taegyun) Yang, Farhad Hormozdiari, Elizabeth Tseng, Alex Zelikovsky, Eleazar Eskin |
ISBRA | 1 |
| 2016 | Long Single-Molecule Reads Can Resolve the Complexity of the Influenza Virus Composed of Rare, Closely Related Mutant Variants
Alexander Artyomenko, Nicholas C. Wu, Serghei Mangul, Eleazar Eskin, Ren Sun, Alex Zelikovsky |
RECOMB | 3 |
| 2014 | Accurate viral population assembly from ultra-deep sequencing dataabstractMOTIVATION: Next-generation sequencing technologies sequence viruses with ultra-deep coverage, thus promising to revolutionize our understanding of the underlying diversity of viral populations. While the sequencing coverage is high enough that even rare viral variants are sequenced, the presence of sequencing errors makes it difficult to distinguish between rare variants and sequencing errors. RESULTS: In this article, we present a method to overcome the limitations of sequencing technologies and assemble a diverse viral population that allows for the detection of previously undiscovered rare variants. The proposed method consists of a high-fidelity sequencing protocol and an accurate viral population assembly method, referred to as Viral Genome Assembler (VGA). The proposed protocol is able to eliminate sequencing errors by using individual barcodes attached to the sequencing fragments. Highly accurate data in combination with deep coverage allow VGA to assemble rare variants. VGA uses an expectation-maximization algorithm to estimate abundances of the assembled viral variants in the population. RESULTS on both synthetic and real datasets show that our method is able to accurately assemble an HIV viral population and detect rare variants previously undetectable due to sequencing errors. VGA outperforms state-of-the-art methods for genome-wide viral assembly. Furthermore, our method is the first viral assembly method that scales to millions of sequencing reads. AVAILABILITY: Our tool VGA is freely available at http://genetics.cs.ucla.edu/vga/ Serghei Mangul, Nicholas C. Wu, Nicholas Mancuso, Alex Zelikovsky, Ren Sun, Eleazar Eskin |
Bioinform. | 1 |
| 2012 | TRIP: a method for novel transcript reconstruction from paired-end RNA-seq readsabstractRecent advances in DNA sequencing have made it possible to sequence the whole transcriptome by massively parallel sequencing, commonly referred as RNA-Seq. RNA-Seq is quickly becoming the technology of choice for transcriptome research and analyses. RNA-Seq allows to reduce the sequencing cost and significantly increase data throughput, but it is computationally challenging to use such RNA-Seq data for reconstructing of full length transcripts and accurately estimate their abundances across all cell types. A number of recent works have addressed the problem of transcriptome reconstruction from RNA-Seq reads. These methods fall into three categories: genome-guided, genome-independent and annotation-guided. In this work, we propose a novel statistical genome-guided method called “ T ranscriptome R econstruction using I nteger P rograming” (TRIP) that incorporates fragment length distribution into novel transcript reconstruction from paired-end RNA-Seq reads. To reconstruct novel transcripts, we create a splice graph based on inferred exon boundaries and RNA-Seq reads. A splice graph is a directed acyclic graph (DAG), whose vertices represent exons and edges represent splicing events. We enumerate all maximal paths in the splice graph using a depth-first-search (DFS) algorithm. These paths correspond to putative transcripts and are the input for the TRIP algorithm. To solve the transcriptome reconstruction problem we must select a set of putative transcripts with the highest support from the RNA-Seq reads. We formulate this problem as an integer program. The objective to select the smallest set of putative transcripts that yields a good statistical fit between the fragment length distribution empirically determined during library preparation and fragment lengths implied by mapping read pairs to selected transcripts. Preliminary experimental results on synthetic datasets generated with various sequencing parameters and distribution assumptions show that TRIP has increased transcriptome reconstruction accuracy compared to previous methods that ignore fragment length distribution information. Serghei Mangul, Adrian Caciula, Dumitru Brinza, Ion I. Mandoiu, Alex Zelikovsky |
BMC Bioinform. | 1 |
| 2011 | Maximum Likelihood Estimation of Incomplete Genomic Spectrum from HTS Data
Serghei Mangul, Irina Astrovskaya, Marius Nicolae, Bassam Tork, Ion I. Mandoiu, Alex Zelikovsky |
WABI | 1 |
| 2011 | Inferring viral quasispecies spectra from 454 pyrosequencing readsabstractBACKGROUND: RNA viruses infecting a host usually exist as a set of closely related sequences, referred to as quasispecies. The genomic diversity of viral quasispecies is a subject of great interest, particularly for chronic infections, since it can lead to resistance to existing therapies. High-throughput sequencing is a promising approach to characterizing viral diversity, but unfortunately standard assembly software was originally designed for single genome assembly and cannot be used to simultaneously assemble and estimate the abundance of multiple closely related quasispecies sequences. RESULTS: In this paper, we introduce a new Viral Spectrum Assembler (ViSpA) method for quasispecies spectrum reconstruction and compare it with the state-of-the-art ShoRAH tool on both simulated and real 454 pyrosequencing shotgun reads from HCV and HIV quasispecies. Experimental results show that ViSpA outperforms ShoRAH on simulated error-free reads, correctly assembling 10 out of 10 quasispecies and 29 sequences out of 40 quasispecies. While ShoRAH has a significant advantage over ViSpA on reads simulated with sequencing errors due to its advanced error correction algorithm, ViSpA is better at assembling the simulated reads after they have been corrected by ShoRAH. ViSpA also outperforms ShoRAH on real 454 reads. Indeed, 7 most frequent sequences reconstructed by ViSpA from a real HCV dataset are viable (do not contain internal stop codons), and the most frequent sequence was within 1% of the actual open reading frame obtained by cloning and Sanger sequencing. In contrast, only one of the sequences reconstructed by ShoRAH is viable. On a real HIV dataset, ShoRAH correctly inferred only 2 quasispecies sequences with at most 4 mismatches whereas ViSpA correctly reconstructed 5 quasispecies with at most 2 mismatches, and 2 out of 5 sequences were inferred without any mismatches. ViSpA source code is available at http://alla.cs.gsu.edu/~software/VISPA/vispa.html. CONCLUSIONS: ViSpA enables accurate viral quasispecies spectrum reconstruction from 454 pyrosequencing reads. We are currently exploring extensions applicable to the analysis of high-throughput sequencing data from bacterial metagenomic samples and ecological samples of eukaryote populations. Irina Astrovskaya, Bassam Tork, Serghei Mangul, Kelly Westbrooks, Ion I. Mandoiu, Peter Balfe, Alex Zelikovsky |
BMC Bioinform. | 3 |
| 2010 | Estimation of Alternative Splicing isoform Frequencies from RNA-Seq Data
Marius Nicolae, Serghei Mangul, Ion I. Mandoiu, Alex Zelikovsky |
WABI | 2 |