Brona Brejová

dblp:83/4890 · also Bronislava Brejová · DBLP profile ↗
← Back
39ranked-venue papers
15as first author
9since 2021 · last 2026
0000-0002-9483-1766ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 5 first-author · 8 since 2021Theory of computation · 10 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author
YearPublicationVenuePosition
2026 Efficient Algorithms for Pangenome Personalization
abstract
A pangenome graph is a representation of the genomes of multiple individuals of the same species. Using a pangenome graph reference instead of a single linear reference genome can increase accuracy of read mapping and downstream tasks, e.g., variant calling, but can also lead to increasing computational demands and false positives. In 2024, Sirén et al. proposed to select only parts of the pangenome mostly likely to match a studied individual, introducing the so-called personalized pangenome reference. Their algorithm is based on greedily selecting sections of paths representing individual haplotypes comprising the pangenome. In this article, we formulate the problem of pangenome personalization purely in terms of pangenome vertices and edges, as finding two paths using vertices supported by sequencing data. We provide several algorithms for solving the problem, ranging from a simple linear-time greedy algorithm with approximation ratio analysis, through dynamic programming and application of minimum-cost flow. Our implementation misses only a small percentage of vertices belonging to the studied individual and improves the sensitivity of read mapping compared to the linear reference.
Denys Andrukhovskyi, Martin Madzin, Luca Denti, Tomás Vinar, Brona Brejová
WABI5
2024 Efficient Analysis of Annotation Colocalization Accounting for Genomic Contexts
Askar Gafurov, Tomás Vinar, Paul Medvedev, Brona Brejová
RECOMB4
2023 PlasBin-flow: a flow-based MILP algorithm for plasmid contigs binning
abstract
MOTIVATION: The analysis of bacterial isolates to detect plasmids is important due to their role in the propagation of antimicrobial resistance. In short-read sequence assemblies, both plasmids and bacterial chromosomes are typically split into several contigs of various lengths, making identification of plasmids a challenging problem. In plasmid contig binning, the goal is to distinguish short-read assembly contigs based on their origin into plasmid and chromosomal contigs and subsequently sort plasmid contigs into bins, each bin corresponding to a single plasmid. Previous works on this problem consist of de novo approaches and reference-based approaches. De novo methods rely on contig features such as length, circularity, read coverage, or GC content. Reference-based approaches compare contigs to databases of known plasmids or plasmid markers from finished bacterial genomes. RESULTS: Recent developments suggest that leveraging information contained in the assembly graph improves the accuracy of plasmid binning. We present PlasBin-flow, a hybrid method that defines contig bins as subgraphs of the assembly graph. PlasBin-flow identifies such plasmid subgraphs through a mixed integer linear programming model that relies on the concept of network flow to account for sequencing coverage, while also accounting for the presence of plasmid genes and the GC content that often distinguishes plasmids from chromosomes. We demonstrate the performance of PlasBin-flow on a real dataset of bacterial samples. AVAILABILITY AND IMPLEMENTATION: https://github.com/cchauve/PlasBin-flow.
Aniket C. Mane, Mahsa Faizrahnemoon, Tomás Vinar, Brona Brejová, Cédric Chauve
Bioinform.4
2023 WarpSTR: determining tandem repeat lengths using raw nanopore signals
abstract
MOTIVATION: Short tandem repeats (STRs) are regions of a genome containing many consecutive copies of the same short motif, possibly with small variations. Analysis of STRs has many clinical uses but is limited by technology mainly due to STRs surpassing the used read length. Nanopore sequencing, as one of long-read sequencing technologies, produces very long reads, thus offering more possibilities to study and analyze STRs. Basecalling of nanopore reads is however particularly unreliable in repeating regions, and therefore direct analysis from raw nanopore data is required. RESULTS: Here, we present WarpSTR, a novel method for characterizing both simple and complex tandem repeats directly from raw nanopore signals using a finite-state automaton and a search algorithm analogous to dynamic time warping. By applying this approach to determine the lengths of 241 STRs, we demonstrate that our approach decreases the mean absolute error of the STR length estimate compared to basecalling and STRique. AVAILABILITY AND IMPLEMENTATION: WarpSTR is freely available at https://github.com/fmfi-compbio/warpstr.
Jozef Sitarcík, Tomás Vinar, Brona Brejová, Werner Krampl, Jaroslav Budis, Ján Radvánszky, Mária Lucká
Bioinform.3
2022 Markov chains improve the significance computation of overlapping genome annotations
abstract
MOTIVATION: Genome annotations are a common way to represent genomic features such as genes, regulatory elements or epigenetic modifications. The amount of overlap between two annotations is often used to ascertain if there is an underlying biological connection between them. In order to distinguish between true biological association and overlap by pure chance, a robust measure of significance is required. One common way to do this is to determine if the number of intervals in the reference annotation that intersect the query annotation is statistically significant. However, currently employed statistical frameworks are often either inefficient or inaccurate when computing P-values on the scale of the whole human genome. RESULTS: We show that finding the P-values under the typically used 'gold' null hypothesis is NP-hard. This motivates us to reformulate the null hypothesis using Markov chains. To be able to measure the fidelity of our Markovian null hypothesis, we develop a fast direct sampling algorithm to estimate the P-value under the gold null hypothesis. We then present an open-source software tool MCDP that computes the P-values under the Markovian null hypothesis in O(m2+n) time and O(m) memory, where m and n are the numbers of intervals in the reference and query annotations, respectively. Notably, MCDP runtime and memory usage are independent from the genome length, allowing it to outperform previous approaches in runtime and memory usage by orders of magnitude on human genome annotations, while maintaining the same level of accuracy. AVAILABILITY AND IMPLEMENTATION: The software is available at https://github.com/fmfi-compbio/mc-overlaps. All data for reproducibility are available at https://github.com/fmfi-compbio/mc-overlaps-reproducibility. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Askar Gafurov, Brona Brejová, Paul Medvedev
Bioinform.2
2022 VirPool: model-based estimation of SARS-CoV-2 variant proportions in wastewater samples
abstract
BACKGROUND: The genomes of SARS-CoV-2 are classified into variants, some of which are monitored as variants of concern (e.g. the Delta variant B.1.617.2 or Omicron variant B.1.1.529). Proportions of these variants circulating in a human population are typically estimated by large-scale sequencing of individual patient samples. Sequencing a mixture of SARS-CoV-2 RNA molecules from wastewater provides a cost-effective alternative, but requires methods for estimating variant proportions in a mixed sample. RESULTS: We propose a new method based on a probabilistic model of sequencing reads, capturing sequence diversity present within individual variants, as well as sequencing errors. The algorithm is implemented in an open source Python program called VirPool. We evaluate the accuracy of VirPool on several simulated and real sequencing data sets from both Illumina and nanopore sequencing platforms, including wastewater samples from Austria and France monitoring the onset of the Alpha variant. CONCLUSIONS: VirPool is a versatile tool for wastewater and other mixed-sample analysis that can handle both short- and long-read sequencing data. Our approach does not require pre-selection of characteristic mutations for variant profiles, it is able to use the entire length of reads instead of just the most informative positions, and can also capture haplotype dependencies within a single read.
Askar Gafurov, Andrej Baláz, Fabian Amman, Kristína Borsová, Viktória Cabanová, Boris Klempa, Andreas Bergthaler, Tomás Vinar, Brona Brejová
BMC Bioinform.9
2022 Dynamic Pooling Improves Nanopore Base Calling Accuracy
abstract
In nanopore sequencing, electrical signal is measured as DNA molecules pass through the sequencing pores. Translating these signals into DNA bases (base calling) is a highly non-trivial task, and its quality has a large impact on the sequencing accuracy. The most successful nanopore base callers to date use convolutional neural networks (CNN) to accomplish the task. Convolutional layers in CNNs are typically composed of filters with constant window size, performing best in analysis of signals with uniform speed. However, the speed of nanopore sequencing varies greatly both within reads and between sequencing runs. Here, we present dynamic pooling, a novel neural network component, which addresses this problem by adaptively adjusting the pooling ratio. To demonstrate the usefulness of dynamic pooling, we developed two base callers: Heron and Osprey. Heron improves the accuracy beyond the experimental high-accuracy base caller Bonito developed by Oxford Nanopore. Osprey is a fast base caller that can compete in accuracy with Guppy high-accuracy mode, but does not require GPU acceleration and achieves a near real-time speed on common desktop CPUs. Availability: https://github.com/fmfi-compbio/osprey, https://github.com/fmfi-compbio/heron.
Vladimír Boza, Peter Peresíni, Brona Brejová, Tomás Vinar
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 Probabilistic Models of k-mer Frequencies (Extended Abstract)
Askar Gafurov, Tomás Vinar, Brona Brejová
CiE3
2021 Nanopore base calling on the edge
abstract
MOTIVATION: MinION is a portable nanopore sequencing device that can be easily operated in the field with features including monitoring of run progress and selective sequencing. To fully exploit these features, real-time base calling is required. Up to date, this has only been achieved at the cost of high computing requirements that pose limitations in terms of hardware availability in common laptops and energy consumption. RESULTS: We developed a new base caller DeepNano-coral for nanopore sequencing, which is optimized to run on the Coral Edge Tensor Processing Unit, a small USB-attached hardware accelerator. To achieve this goal, we have designed new versions of two key components used in convolutional neural networks for speech recognition and base calling. In our components, we propose a new way of factorization of a full convolution into smaller operations, which decreases memory access operations, memory access being a bottleneck on this device. DeepNano-coral achieves real-time base calling during sequencing with the accuracy slightly better than the fast mode of the Guppy base caller and is extremely energy efficient, using only 10 W of power. AVAILABILITY AND IMPLEMENTATION: https://github.com/fmfi-compbio/coral-basecaller. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Peter Peresíni, Vladimír Boza, Brona Brejová, Tomás Vinar
Bioinform.3
2020 DeepNano-blitz: a fast base caller for MinION nanopore sequencers
abstract
MOTIVATION: Oxford Nanopore MinION is a portable DNA sequencer that is marketed as a device that can be deployed anywhere. Current base callers, however, require a powerful GPU to analyze data produced by MinION in real time, which hampers field applications. RESULTS: We have developed a fast base caller DeepNano-blitz that can analyze stream from up to two MinION runs in real time using a common laptop CPU (i7-7700HQ), with no GPU requirements. The base caller settings allow trading accuracy for speed and the results can be used for real time run monitoring (i.e. sample composition, barcode balance, species identification, etc.) or prefiltering of results for more detailed analysis (i.e. filtering out human DNA from human-pathogen runs). AVAILABILITY AND IMPLEMENTATION: DeepNano-blitz has been developed and tested on Linux and Intel processors and is available under MIT license at https://github.com/fmfi-compbio/deepnano-blitz. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vladimír Boza, Peter Peresíni, Brona Brejová, Tomás Vinar
Bioinform.3
2019 Dante: genotyping of known complex and expanded short tandem repeats
abstract
MOTIVATION: Short tandem repeats (STRs) are stretches of repetitive DNA in which short sequences, typically made of 2-6 nucleotides, are repeated several times. Since STRs have many important biological roles and also belong to the most polymorphic parts of the human genome, they became utilized in several molecular-genetic applications. Precise genotyping of STR alleles, therefore, was of high relevance during the last decades. Despite this, massively parallel sequencing (MPS) still lacks the analysis methods to fully utilize the information value of STRs in genome scale assays. RESULTS: We propose an alignment-free algorithm, called Dante, for genotyping and characterization of STR alleles at user-specified known loci based on sequence reads originating from STR loci of interest. The method accounts for natural deviations from the expected sequence, such as variation in the repeat count, sequencing errors, ambiguous bases and complex loci containing several different motifs. In addition, we implemented a correction for copy number defects caused by the polymerase induced stutter effect as well as a prediction of STR expansions that, according to the conventional view, cannot be fully captured by inherently short MPS reads. We tested Dante on simulated datasets and on datasets obtained by targeted sequencing of protein coding parts of thousands of selected clinically relevant genes. In both these datasets, Dante outperformed HipSTR and GATK genotyping tools. Furthermore, Dante was able to predict allele expansions in all tested clinical cases. AVAILABILITY AND IMPLEMENTATION: Dante is open source software, freely available for download at https://github.com/jbudis/dante. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jaroslav Budis, Marcel Kucharík, Frantisek Duris, Juraj Gazdarica, Michaela Zrubcová, Andrej Ficek, Tomás Szemes, Brona Brejová, Ján Radvánszky
Bioinform.8
2016 Isometric Gene Tree Reconciliation Revisited
Brona Brejová, Askar Gafurov, Dana Pardubská, Michal Sabo, Tomás Vinar
WABI1
2016 RNA motif search with data-driven element ordering
abstract
BACKGROUND: In this paper, we study the problem of RNA motif search in long genomic sequences. This approach uses a combination of sequence and structure constraints to uncover new distant homologs of known functional RNAs. The problem is NP-hard and is traditionally solved by backtracking algorithms. RESULTS: We have designed a new algorithm for RNA motif search and implemented a new motif search tool RNArobo. The tool enhances the RNAbob descriptor language, allowing insertions in helices, which enables better characterization of ribozymes and aptamers. A typical RNA motif consists of multiple elements and the running time of the algorithm is highly dependent on their ordering. By approaching the element ordering problem in a principled way, we demonstrate more than 100-fold speedup of the search for complex motifs compared to previously published tools. CONCLUSIONS: We have developed a new method for RNA motif search that allows for a significant speedup of the search of complex motifs that include pseudoknots. Such speed improvements are crucial at a time when the rate of DNA sequencing outpaces growth in computing. RNArobo is available at http://compbio.fmph.uniba.sk/rnarobo .
Ladislav Rampásek, Randi M. Jimenez, Andrej Lupták, Tomás Vinar, Brona Brejová
BMC Bioinform.5
2015 Fishing in Read Collections: Memory Efficient Indexing for Sequence Assembly
Vladimír Boza, Jakub Jursa, Brona Brejová, Tomás Vinar
SPIRE3
2015 How Big is that Genome? Estimating Genome Size and Coverage from k-mer Abundance Spectra
Michal Hozza, Tomás Vinar, Brona Brejová
SPIRE3
2015 Sequence annotation with HMMs: New problems and their complexity
Michal Nánási, Tomás Vinar, Brona Brejová
Inf. Process. Lett.3
2014 GAML: Genome Assembly by Maximum Likelihood
Vladimír Boza, Brona Brejová, Tomás Vinar
WABI2
2013 Probabilistic Approaches to Alignment with Tandem Repeats
Michal Nánási, Tomás Vinar, Brona Brejová
WABI3
2013 Efficient routing in carrier-based mobile networks
Brona Brejová, Stefan Dobrev, Rastislav Kralovic, Tomás Vinar
Theor. Comput. Sci.1
2011 Routing in Carrier-Based Mobile Networks
Brona Brejová, Stefan Dobrev, Rastislav Kralovic, Tomás Vinar
SIROCCO1
2011 Fast Computation of a String Duplication History under No-Breakpoint-Reuse - (Extended Abstract)
Brona Brejová, Gad M. Landau, Tomás Vinar
SPIRE1
2011 Automated Segmentation of DNA Sequences with Complex Evolutionary Histories
Brona Brejová, Michal Burger, Tomás Vinar
WABI1
2011 A Practical Algorithm for Ancestral Rearrangement Reconstruction
Jakub Kovác, Brona Brejová, Tomás Vinar
WABI2
2010 The Highest Expected Reward Decoding for HMMs with Application to Recombination Detection
Michal Nánási, Tomás Vinar, Brona Brejová
CPM3
2009 Predicting Gene Structures from Multiple RT-PCR Tests
Jakub Kovác, Tomás Vinar, Brona Brejová
WABI3
2007 On-Line Viterbi Algorithm for Analysis of Long Biological Sequences
Rastislav Srámek, Brona Brejová, Tomás Vinar
WABI2
2007 The most probable annotation problem in HMMs and its application to bioinformatics
Brona Brejová, Dan Brown 0001, Tomás Vinar
J. Comput. Syst. Sci.1
2006 New Bounds for Motif Finding in Strong Instances
Brona Brejová, Dan Brown 0001, Ian M. Harrower, Tomás Vinar
CPM1
2005 Sharper Upper and Lower Bounds for an Approximation Scheme for Consensus-Pattern
Brona Brejová, Dan Brown 0001, Ian M. Harrower, Alejandro López-Ortiz, Tomás Vinar
CPM1
2005 Vector seeds: An extension to spaced seeds
Brona Brejová, Dan Brown 0001, Tomás Vinar
J. Comput. Syst. Sci.1
2004 The Most Probable Labeling Problem in HMMs and Its Application to Bioinformatics
Brona Brejová, Dan Brown 0001, Tomás Vinar
WABI1
2004 Finding hidden independent sets in interval graphs
Therese Biedl, Brona Brejová, Erik D. Demaine, Angèle M. Foley, Alejandro López-Ortiz, Tomás Vinar
Theor. Comput. Sci.2
2003 Finding Hidden Independent Sets in Interval Graphs
Therese Biedl, Brona Brejová, Erik D. Demaine, Angèle M. Foley, Alejandro López-Ortiz, Tomás Vinar
COCOON2
2003 Optimal Spaced Seeds for Hidden Markov Models, with Application to Homologous Coding Regions
Brona Brejová, Dan Brown 0001, Tomás Vinar
CPM1
2003 Vector Seeds: An Extension to Spaced Seeds Allows Substantial Improvements in Sensitivity and Specifity
Brona Brejová, Dan Brown 0001, Tomás Vinar
WABI1
2003 Optimal DNA Signal Recognition Models with a Fixed Amount of Intrasignal Dependency
Brona Brejová, Dan Brown 0001, Tomás Vinar
WABI1
2002 A Better Method for Length Distribution Modeling in HMMs and Its Application to Gene Finding
Brona Brejová, Tomás Vinar
CPM1
2001 Analyzing variants of Shellsort
Brona Brejová
Inf. Process. Lett.1
2000 Simplifying Flow Networks
Therese Biedl, Brona Brejová, Tomás Vinar
MFCS2