Aleksandra A. Galitsyna

dblp:234/3820 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0001-8969-5694ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Bioframe: operations on genomic intervals in Pandas dataframes
abstract
MOTIVATION: Genomic intervals are one of the most prevalent data structures in computational genome biology, and used to represent features ranging from genes, to DNA binding sites, to disease variants. Operations on genomic intervals provide a language for asking questions about relationships between features. While there are excellent interval arithmetic tools for the command line, they are not smoothly integrated into Python, one of the most popular general-purpose computational and visualization environments. RESULTS: Bioframe is a library to enable flexible and performant operations on genomic interval dataframes in Python. Bioframe extends the Python data science stack to use cases for computational genome biology by building directly on top of two of the most commonly-used Python libraries, NumPy and Pandas. The bioframe API enables flexible name and column orders, and decouples operations from data formats to avoid unnecessary conversions, a common scourge for bioinformaticians. Bioframe achieves these goals while maintaining high performance and a rich set of features. AVAILABILITY AND IMPLEMENTATION: Bioframe is open-source under MIT license, cross-platform, and can be installed from the Python Package Index. The source code is maintained by Open2C on GitHub at https://github.com/open2c/bioframe.
Nezar Abdennur, Geoffrey Fudenberg, Ilya M. Flyamer, Aleksandra A. Galitsyna, Anton Goloborodko, Maxim Imakaev, Sergey Venev
Bioinform.4
2024 Cooltools: Enabling high-resolution Hi-C analysis in Python
abstract
Chromosome conformation capture (3C) technologies reveal the incredible complexity of genome organization. Maps of increasing size, depth, and resolution are now used to probe genome architecture across cell states, types, and organisms. Larger datasets add challenges at each step of computational analysis, from storage and memory constraints to researchers' time; however, analysis tools that meet these increased resource demands have not kept pace. Furthermore, existing tools offer limited support for customizing analysis for specific use cases or new biology. Here we introduce cooltools (https://github.com/open2c/cooltools), a suite of computational tools that enables flexible, scalable, and reproducible analysis of high-resolution contact frequency data. Cooltools leverages the widely-adopted cooler format which handles storage and access for high-resolution datasets. Cooltools provides a paired command line interface (CLI) and Python application programming interface (API), which respectively facilitate workflows on high-performance computing clusters and in interactive analysis environments. In short, cooltools enables the effective use of the latest and largest genome folding datasets.
Nezar Abdennur, Sameer Abraham, Geoffrey Fudenberg, Ilya M. Flyamer, Aleksandra A. Galitsyna, Anton Goloborodko, Maxim Imakaev, Betul A. Oksuz, Sergey Venev
PLoS Comput. Biol.5
2024 Pairtools: From sequencing data to chromosome contacts
abstract
The field of 3D genome organization produces large amounts of sequencing data from Hi-C and a rapidly-expanding set of other chromosome conformation protocols (3C+). Massive and heterogeneous 3C+ data require high-performance and flexible processing of sequenced reads into contact pairs. To meet these challenges, we present pairtools-a flexible suite of tools for contact extraction from sequencing data. Pairtools provides modular command-line interface (CLI) tools that can be flexibly chained into data processing pipelines. The core operations provided by pairtools are parsing of.sam alignments into Hi-C pairs, sorting and removal of PCR duplicates. In addition, pairtools provides auxiliary tools for building feature-rich 3C+ pipelines, including contact pair manipulation, filtration, and quality control. Benchmarking pairtools against popular 3C+ data pipelines shows advantages of pairtools for high-performance and flexible 3C+ analysis. Finally, pairtools provides protocol-specific tools for restriction-based protocols, haplotype-resolved contacts, and single-cell Hi-C. The combination of CLI tools and tight integration with Python data analysis libraries makes pairtools a versatile foundation for a broad range of 3C+ pipelines.
Nezar Abdennur, Geoffrey Fudenberg, Ilya M. Flyamer, Aleksandra A. Galitsyna, Anton Goloborodko, Maxim Imakaev, Sergey Venev
PLoS Comput. Biol.4
2023 HiConfidence: a novel approach uncovering the biological signal in Hi-C data affected by technical biases
abstract
The chromatin interaction assays, particularly Hi-C, enable detailed studies of genome architecture in multiple organisms and model systems, resulting in a deeper understanding of gene expression regulation mechanisms mediated by epigenetics. However, the analysis and interpretation of Hi-C data remain challenging due to technical biases, limiting direct comparisons of datasets obtained in different experiments and laboratories. As a result, removing biases from Hi-C-generated chromatin contact matrices is a critical data analysis step. Our novel approach, HiConfidence, eliminates biases from the Hi-C data by weighing chromatin contacts according to their consistency between replicates so that low-quality replicates do not substantially influence the result. The algorithm is effective for the analysis of global changes in chromatin structures such as compartments and topologically associating domains. We apply the HiConfidence approach to several Hi-C datasets with significant technical biases, that could not be analyzed effectively using existing methods, and obtain meaningful biological conclusions. In particular, HiConfidence aids in the study of how changes in histone acetylation pattern affect chromatin organization in Drosophila melanogaster S2 cells. The method is freely available at GitHub: https://github.com/victorykobets/HiConfidence.
Victoria A Kobets, Sergey V. Ulianov, Aleksandra A. Galitsyna, Semen A. Doronin, Elena A. Mikhaleva, Mikhail S. Gelfand, Yuri Y. Shevelyov, Sergey V. Razin, Ekaterina Khrameeva
Briefings Bioinform.3
2021 Single-cell Hi-C data analysis: safety in numbers
abstract
Over the past decade, genome-wide assays for chromatin interactions in single cells have enabled the study of individual nuclei at unprecedented resolution and throughput. Current chromosome conformation capture techniques survey contacts for up to tens of thousands of individual cells, improving our understanding of genome function in 3D. However, these methods recover a small fraction of all contacts in single cells, requiring specialised processing of sparse interactome data. In this review, we highlight recent advances in methods for the interpretation of single-cell genomic contacts. After discussing the strengths and limitations of these methods, we outline frontiers for future development in this rapidly moving field.
Aleksandra A. Galitsyna, Mikhail S. Gelfand
Briefings Bioinform.1
2021 Perspectives for the reconstruction of 3D chromatin conformation using single cell Hi-C data
abstract
Construction of chromosomes 3D models based on single cell Hi-C data constitute an important challenge. We present a reconstruction approach, DPDchrom, that incorporates basic knowledge whether the reconstructed conformation should be coil-like or globular and spring relaxation at contact sites. In contrast to previously published protocols, DPDchrom can naturally form globular conformation due to the presence of explicit solvent. Benchmarking of this and several other methods on artificial polymer models reveals similar reconstruction accuracy at high contact density and DPDchrom advantage at low contact density. To compare 3D structures insensitively to spatial orientation and scale, we propose the Modified Jaccard Index. We analyzed two sources of the contact dropout: contact radius change and random contact sampling. We found that the reconstruction accuracy exponentially depends on the number of contacts per genomic bin allowing to estimate the reconstruction accuracy in advance. We applied DPDchrom to model chromosome configurations based on single-cell Hi-C data of mouse oocytes and found that these configurations differ significantly from a random one, that is consistent with other studies.
Pavel Kos, Aleksandra A. Galitsyna, Sergey V. Ulianov, Mikhail S. Gelfand, Sergey V. Razin, Alexander Chertovich
PLoS Comput. Biol.2
2020 HiChew: a Tool for TAD Clustering in Embryogenesis
Nikolai S. Bykov, Olga M. Sigalova, Mikhail S. Gelfand, Aleksandra A. Galitsyna
ISBRA4
2018 Large-scale analysis of RNA-DNA interactions
Aleksandra A. Galitsyna, Anastasia Zharikova, Mariya D. Logacheva, Sergey V. Razin, Andrey Mironov, Alexey Gavrilov
BIBM1
2018 Reconstruction of the chromatin 3D conformation from single cell Hi-C data
Pavel Kos, Aleksandra A. Galitsyna, Sergey V. Ulianov, Mikhail S. Gelfand, Sergey V. Razin, Alexander Chertovich
BIBM2
2018 The chromatin structure of Dictyostelium discoideum
Olga Tsoy, Aleksandra A. Galitsyna, Ekaterina Khrameeva, Sergey V. Ulianov, Mikhail S. Gelfand, Sergey V. Razin
BIBM2
2018 Nuclear lamina maintains global spatial organization of chromatin in Drosophila cultured cells
Sergey V. Ulianov, Semen A. Doronin, Ekaterina Khrameeva, Pavel Kos, Sergey S. Starikov, Aleksandra A. Galitsyna, Artem Luzhin, Mikhail S. Gelfand, Alexander Chertovich, Sergey V. Razin, Yuri Y. Shevelyov
BIBM6
2018 Single-cell Hi-C demonstrates that TADs are stable units of Drosophila genome folding that persist in individual cells
Vlada S. Zakharova, Aleksandra A. Galitsyna, Kirill E. Polovnikov, Ekaterina Khrameeva, Mariya D. Logacheva, Elena A. Mikhaleva, Egor S. Vassetzky, Aleksey A. Gavrilov, Yuri Y. Shevelev, Sergey K. Nechaev, Sergey V. Ulianov, Sergey V. Razin
BIBM2