VLDB 2026 Research / reviewers in the wild / expert
Alexej Abyzov
dblp:43/6231
· DBLP profile ↗
14ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0001-5405-6729ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Genome-wide analysis and visualization of copy number with CNVpytor in igv.jsabstractSUMMARY: Copy number variation (CNV) and alteration (CNA) analysis is a crucial component in many genomic studies and its applications span from basic research to clinic diagnostics and personalized medicine. CNVpytor is a tool featuring a read depth-based caller and combined read depth and B-allele frequency (BAF) based 2D caller to find CNVs and CNAs. The tool stores processed intermediate data and CNV/CNA calls in a compact HDF5 file-pytor file. Here, we describe a new track in igv.js that utilizes pytor and whole genome variant files as input for on-the-fly read depth and BAF visualization, CNV/CNA calling and analysis. Embedding into HTML pages and Jupiter Notebooks enables convenient remote data access and visualization simplifying interpretation and analysis of omics data. AVAILABILITY AND IMPLEMENTATION: The CNVpytor track is integrated with igv.js and available at https://github.com/igvteam/igv.js. The documentation is available at https://github.com/igvteam/igv.js/wiki/cnvpytor. Usage can be tested in the IGV-Web app at https://igv.org/app and also on https://github.com/abyzovlab/CNVpytor. Arijit Panda, Milovan Suvakov, Helga Thorvaldsdóttir, Jill P. Mesirov, James T. Robinson, Alexej Abyzov |
Bioinform. | 6 |
| 2022 | OpBerg: Discovering Causal Sentences Using Optimal Alignments
Justin Wood, Nicholas J. Matiasz, Alcino J. Silva, William Hsu, Alexej Abyzov, Wei Wang 0010 |
DaWaK | 5 |
| 2022 | All2: A tool for selecting mosaic mutations from comprehensive multi-cell comparisonsabstractAccurate discovery of somatic mutations in a cell is a challenge that partially lays in immaturity of dedicated analytical approaches. Approaches comparing a cell's genome to a control bulk sample miss common mutations, while approaches to find such mutations from bulk suffer from low sensitivity. We developed a tool, All2, which enables accurate filtering of mutations in a cell without the need for data from bulk(s). It is based on pair-wise comparisons of all cells to each other where every call for base pair substitution and indel is classified as either a germline variant, mosaic mutation, or false positive. As All2 allows for considering dropped-out regions, it is applicable to whole genome and exome analysis of cloned and amplified cells. By applying the approach to a variety of available data, we showed that its application reduces false positives, enables sensitive discovery of high frequency mutations, and is indispensable for conducting high resolution cell lineage tracing. Vivekananda Sarangi, Yeongjun Jang, Milovan Suvakov, Taejeong Bae, Liana Fasching, Shobana Sekar, Livia Tomasini, Jessica Mariani, Flora Vaccarino, Alexej Abyzov |
PLoS Comput. Biol. | 10 |
| 2022 | Correction: All2: A tool for selecting mosaic mutations from comprehensive multi-cell comparisonsabstract[This corrects the article DOI: 10.1371/journal.pcbi.1009487.]. Vivekananda Sarangi, Yeongjun Jang, Milovan Suvakov, Taejeong Bae, Liana Fasching, Shobana Sekar, Livia Tomasini, Jessica Mariani, Flora Vaccarino, Alexej Abyzov |
PLoS Comput. Biol. | 10 |
| 2021 | LongAGE: defining breakpoints of genomic structural variants through optimal and memory efficient alignments of long readsabstractSUMMARY: Defining the precise location of structural variations (SVs) at single-nucleotide breakpoint resolution is a challenging problem due to large gaps in alignment. Previously, Alignment with Gap Excision (AGE) enabled us to define breakpoints of SVs at single-nucleotide resolution; however, AGE requires a vast amount of memory when aligning a pair of long sequences. To address this, we developed a memory-efficient implementation-LongAGE-based on the classical Hirschberg algorithm. We demonstrate an application of LongAGE for resolving breakpoints of SVs embedded into segmental duplications on Pacific Biosciences (PacBio) reads that can be longer than 10 kb. Furthermore, we observed different breakpoints for a deletion and a duplication in the same locus, providing direct evidence that such multi-allelic copy number variants (mCNVs) arise from two or more independent ancestral mutations. AVAILABILITY AND IMPLEMENTATION: LongAGE is implemented in C++ and available on Github at https://github.com/Coaxecva/LongAGE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Quang Tran 0002, Alexej Abyzov |
Bioinform. | 2 |
| 2020 | SCELLECTOR: ranking amplification bias in single cells using shallow sequencingabstractBACKGROUND: The study of mosaic mutation is important since it has been linked to cancer and various disorders. Single cell sequencing has become a powerful tool to study the genome of individual cells for the detection of mosaic mutations. The amount of DNA in a single cell needs to be amplified before sequencing and multiple displacement amplification (MDA) is widely used owing to its low error rate and long fragment length of amplified DNA. However, the phi29 polymerase used in MDA is sensitive to template fragmentation and presence of sites with DNA damage that can lead to biases such as allelic imbalance, uneven coverage and over representation of C to T mutations. It is therefore important to select cells with uniform amplification to decrease false positives and increase sensitivity for mosaic mutation detection. RESULTS: We propose a method, Scellector (single cell selector), which uses haplotype information to detect amplification quality in shallow coverage sequencing data. We tested Scellector on single human neuronal cells, obtained in vitro and amplified by MDA. Qualities were estimated from shallow sequencing with coverage as low as 0.3× per cell and then confirmed using 30× deep coverage sequencing. The high concordance between shallow and high coverage data validated the method. CONCLUSION: Scellector can potentially be used to rank amplifications obtained from single cell platforms relying on a MDA-like amplification step, such as Chromium Single Cell profiling solution. Vivekananda Sarangi, Alexandre Jourdon, Taejeong Bae, Arijit Panda, Flora Vaccarino, Alexej Abyzov |
BMC Bioinform. | 6 |
| 2017 | Landscape and variation of novel retroduplications in 26 human populationsabstractRetroduplications come from reverse transcription of mRNAs and their insertion back into the genome. Here, we performed comprehensive discovery and analysis of retroduplications in a large cohort of 2,535 individuals from 26 human populations, as part of 1000 Genomes Phase 3. We developed an integrated approach to discover novel retroduplications combining high-coverage exome and low-coverage whole-genome sequencing data, utilizing information from both exon-exon junctions and discordant paired-end reads. We found 503 parent genes having novel retroduplications absent from the reference genome. Based solely on retroduplication variation, we built phylogenetic trees of human populations; these represent superpopulation structure well and indicate that variable retroduplications are effective population markers. We further identified 43 retroduplication parent genes differentiating superpopulations. This group contains several interesting insertion events, including a SLMO2 retroduplication and insertion into CAV3, which has a potential disease association. We also found retroduplications to be associated with a variety of genomic features: (1) Insertion sites were correlated with regular nucleosome positioning. (2) They, predictably, tend to avoid conserved functional regions, such as exons, but, somewhat surprisingly, also avoid introns. (3) Retroduplications tend to be co-inserted with young L1 elements, indicating recent retrotranspositional activity, and (4) they have a weak tendency to originate from highly expressed parent genes. Our investigation provides insight into the functional impact and association with genomic elements of retroduplications. We anticipate our approach and analytical methodology to have application in a more clinical context, where exome sequencing data is abundant and the discovery of retroduplications can potentially improve the accuracy of SNP calling. Yan Zhang 0032, Shantao Li, Alexej Abyzov, Mark Gerstein |
PLoS Comput. Biol. | 3 |
| 2015 | MetaSV: an accurate and integrative structural-variant caller for next generation sequencingabstractUNLABELLED: Structural variations (SVs) are large genomic rearrangements that vary significantly in size, making them challenging to detect with the relatively short reads from next-generation sequencing (NGS). Different SV detection methods have been developed; however, each is limited to specific kinds of SVs with varying accuracy and resolution. Previous works have attempted to combine different methods, but they still suffer from poor accuracy particularly for insertions. We propose MetaSV, an integrated SV caller which leverages multiple orthogonal SV signals for high accuracy and resolution. MetaSV proceeds by merging SVs from multiple tools for all types of SVs. It also analyzes soft-clipped reads from alignment to detect insertions accurately since existing tools underestimate insertion SVs. Local assembly in combination with dynamic programming is used to improve breakpoint resolution. Paired-end and coverage information is used to predict SV genotypes. Using simulation and experimental data, we demonstrate the effectiveness of MetaSV across various SV types and sizes. AVAILABILITY AND IMPLEMENTATION: Code in Python is at http://bioinform.github.io/metasv/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marghoob Mohiyuddin, John C. Mu, Narges Bani Asadi, Mark Gerstein, Alexej Abyzov, Wing H. Wong, Hugo Y. K. Lam |
Bioinform. | 6 |
| 2015 | VarSim: a high-fidelity simulation and validation framework for high-throughput genome sequencing with cancer applicationsabstractSUMMARY: VarSim is a framework for assessing alignment and variant calling accuracy in high-throughput genome sequencing through simulation or real data. In contrast to simulating a random mutation spectrum, it synthesizes diploid genomes with germline and somatic mutations based on a realistic model. This model leverages information such as previously reported mutations to make the synthetic genomes biologically relevant. VarSim simulates and validates a wide range of variants, including single nucleotide variants, small indels and large structural variants. It is an automated, comprehensive compute framework supporting parallel computation and multiple read simulators. Furthermore, we developed a novel map data structure to validate read alignments, a strategy to compare variants binned in size ranges and a lightweight, interactive, graphical report to visualize validation results with detailed statistics. Thus far, it is the most comprehensive validation tool for secondary analysis in next generation sequencing. AVAILABILITY AND IMPLEMENTATION: Code in Java and Python along with instructions to download the reads and variants is at http://bioinform.github.io/varsim. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. John C. Mu, Marghoob Mohiyuddin, Narges Bani Asadi, Mark Gerstein, Alexej Abyzov, Wing H. Wong, Hugo Y. K. Lam |
Bioinform. | 6 |
| 2011 | AGE: defining breakpoints of genomic structural variants at single-nucleotide resolution, through optimal alignments with gap excisionabstractMOTIVATION: Defining the precise location of structural variations (SVs) at single-nucleotide breakpoint resolution is an important problem, as it is a prerequisite for classifying SVs, evaluating their functional impact and reconstructing personal genome sequences. Given approximate breakpoint locations and a bridging assembly or split read, the problem essentially reduces to finding a correct sequence alignment. Classical algorithms for alignment and their generalizations guarantee finding the optimal (in terms of scoring) global or local alignment of two sequences. However, they cannot generally be applied to finding the biologically correct alignment of genomic sequences containing SVs because of the need to simultaneously span the SV (e.g. make a large gap) and perform precise local alignments at the flanking ends. RESULTS: Here, we formulate the computations involved in this problem and describe a dynamic-programming algorithm for its solution. Specifically, our algorithm, called AGE for Alignment with Gap Excision, finds the optimal solution by simultaneously aligning the 5' and 3' ends of two given sequences and introducing a 'large-gap jump' between the local end alignments to maximize the total alignment score. We also describe extensions allowing the application of AGE to tandem duplications, inversions and complex events involving two large gaps. We develop a memory-efficient implementation of AGE (allowing application to long contigs) and make it available as a downloadable software package. Finally, we applied AGE for breakpoint determination and standardization in the 1000 Genomes Project by aligning locally assembled contigs to the human genome. AVAILABILITY AND IMPLEMENTATION: AGE is freely available at http://sv.gersteinlab.org/age. Alexej Abyzov, Mark Gerstein |
Bioinform. | 1 |
| 2010 | Analysis of Combinatorial Regulation: Scaling of Partnerships between Regulators with the Number of Governed TargetsabstractThrough combinatorial regulation, regulators partner with each other to control common targets and this allows a small number of regulators to govern many targets. One interesting question is that given this combinatorial regulation, how does the number of regulators scale with the number of targets? Here, we address this question by building and analyzing co-regulation (co-transcription and co-phosphorylation) networks that describe partnerships between regulators controlling common genes. We carry out analyses across five diverse species: Escherichia coli to human. These reveal many properties of partnership networks, such as the absence of a classical power-law degree distribution despite the existence of nodes with many partners. We also find that the number of co-regulatory partnerships follows an exponential saturation curve in relation to the number of targets. (For E. coli and Bacillus subtilis, only the beginning linear part of this curve is evident due to arrangement of genes into operons.) To gain intuition into the saturation process, we relate the biological regulation to more commonplace social contexts where a small number of individuals can form an intricate web of connections on the internet. Indeed, we find that the size of partnership networks saturates even as the complexity of their output increases. We also present a variety of models to account for the saturation phenomenon. In particular, we develop a simple analytical model to show how new partnerships are acquired with an increasing number of target genes; with certain assumptions, it reproduces the observed saturation. Then, we build a more general simulation of network growth and find agreement with a wide range of real networks. Finally, we perform various down-sampling calculations on the observed data to illustrate the robustness of our conclusions. Nitin Bhardwaj, Matthew B. Carson, Alexej Abyzov, Koon-Kiu Yan, Hui Lu 0004, Mark Gerstein |
PLoS Comput. Biol. | 3 |
| 2008 | An AP Endonuclease 1-DNA Polymerase β Complex: Theoretical Prediction of Interacting SurfacesabstractAbasic (AP) sites in DNA arise through both endogenous and exogenous mechanisms. Since AP sites can prevent replication and transcription, the cell contains systems for their identification and repair. AP endonuclease (APEX1) cleaves the phosphodiester backbone 5' to the AP site. The cleavage, a key step in the base excision repair pathway, is followed by nucleotide insertion and removal of the downstream deoxyribose moiety, performed most often by DNA polymerase beta (pol-beta). While yeast two-hybrid studies and electrophoretic mobility shift assays provide evidence for interaction of APEX1 and pol-beta, the specifics remain obscure. We describe a theoretical study designed to predict detailed interacting surfaces between APEX1 and pol-beta based on published co-crystal structures of each enzyme bound to DNA. Several potentially interacting complexes were identified by sliding the protein molecules along DNA: two with pol-beta located downstream of APEX1 (3' to the damaged site) and three with pol-beta located upstream of APEX1 (5' to the damaged site). Molecular dynamics (MD) simulations, ensuring geometrical complementarity of interfaces, enabled us to predict interacting residues and calculate binding energies, which in two cases were sufficient (approximately -10.0 kcal/mol) to form a stable complex and in one case a weakly interacting complex. Analysis of interface behavior during MD simulation and visual inspection of interfaces allowed us to conclude that complexes with pol-beta at the 3'-side of APEX1 are those most likely to occur in vivo. Additional multiple sequence analyses of APEX1 and pol-beta in related organisms identified a set of correlated mutations of specific residues at the predicted interfaces. Based on these results, we propose that pol-beta in the open or closed conformation interacts and makes a stable interface with APEX1 bound to a cleaved abasic site on the 3' side. The method described here can be used for analysis in any DNA-metabolizing pathway where weak interactions are the principal mode of cross-talk among participants and co-crystal structures of the individual components are available. Alexej Abyzov, Alper Uzun, Phyllis R. Strauss, Valentin A. Ilyin |
PLoS Comput. Biol. | 1 |
| 2005 | Friend, an integrated analytical front-end application for bioinformaticsabstractUNLABELLED: Friend is a bioinformatics application designed for simultaneous analysis and visualization of multiple structures and sequences of proteins and/or DNA/RNA. The application provides basic functionalities, such as structure visualization, with different rendering and coloring, sequence alignment and simple phylogeny analysis, along with a number of extended features to perform more complex analyses of sequence structure relationships, including structural alignment of proteins, investigation of specific interaction motifs, studies of protein-protein and protein-DNA interactions and protein super-families. It is also useful for functional annotation of proteins, protein modeling and protein folding studies. Friend provides three levels of usage: (1) an extensive GUI for a scientist with no programming experience, (2) a command line interface for scripting for a scientist with some programming experience and (3) the ability to extend Friend with user written libraries for an experienced programmer. The application is linked and communicates with local and remote sequence and structure databases. AVAILABILITY: http://mozart.bio.neu.edu/friend. Alexej Abyzov, Mounir Errami, Chesley M. Leslin, Valentin A. Ilyin |
Bioinform. | 1 |
| 2004 | Structural exon database, SEDB, mapping exon boundaries on multiple protein structuresabstractUNLABELLED: Comparative analysis of exon/intron organization of genes and their resulting protein structures is important for understanding evolutionary relationships between species, rules of protein organization and protein functionality. We present Structural Exon Database (SEDB), with a Web interface, an application that allows users to retrieve the exon/intron organization of genes and map the location of the exon boundaries and the intron phase onto a multiple structural alignment. SEDB is linked with Friend, an integrated analytical multiple sequence/structure viewer, which allows simultaneous visualization of exon boundaries on structure and sequence alignments. With SEDB researchers can study the correlations of gene structure with the properties of the encoded three-dimensional protein structures across eukaryotic organisms. AVAILABILITY: SEDB is publicly available at http://glinka.bio.neu.edu/SEDB/SEDB.html SUPPLEMENTARY INFORMATION: On the SEDB Web site. Chesley M. Leslin, Alexej Abyzov, Valentin A. Ilyin |
Bioinform. | 2 |