VLDB 2026 Research / reviewers in the wild / expert
Matej Lexa
dblp:30/5910
· DBLP profile ↗
17ranked-venue papers
6as first author
3since 2021 · last 2024
0000-0002-4213-5259ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 3 since 2021Systems, architecture and hardware · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › epigenomics
3d genome organization |
0.6 | 1 | 2022 | HiC-TE: a computational pipeline for Hi-C data analysis to study the role of repeat family interactions in the genome 3D organization · Bioinform. 2022 |
Bioinformatics and computational biology › epigenomics
hi-c data analysis |
0.6 | 1 | 2022 | HiC-TE: a computational pipeline for Hi-C data analysis to study the role of repeat family interactions in the genome 3D organization · Bioinform. 2022 |
Bioinformatics and computational biology
sequence analysis |
0.5 | 6 | 2016 | Triplex: an R/Bioconductor package for identification and visualization of potential intramolecular triplex patterns in DNA sequences · Bioinform. 2013 A dynamic programming algorithm for identification of triplex-forming sequences · Bioinform. 2011 Hammock: a hidden Markov model-based peptide clustering algorithm to identify protein-interaction consensus motifs in large datasets · Bioinform. 2016 |
Bioinformatics and computational biology › structural bioinformatics › molecular structure prediction
DNA structure prediction |
0.5 | 2 | 2017 | pqsfinder: an exhaustive and imperfection-tolerant search tool for potential quadruplex-forming sequences in R · Bioinform. 2017 Triplex: an R/Bioconductor package for identification and visualization of potential intramolecular triplex patterns in DNA sequences · Bioinform. 2013 |
Bioinformatics and computational biology › molecular evolution › evolutionary bioinformatics › evolutionary genomics
genome evolution |
0.4 | 1 | 2020 | TE-greedy-nester: structure-based detection of LTR retrotransposons and their nesting · Bioinform. 2020 |
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics › RNA structure prediction
g-quadruplex prediction |
0.4 | 1 | 2020 | pqsfinder web: G-quadruplex prediction using optimized pqsfinder algorithm · Bioinform. 2020 |
Bioinformatics and computational biology › structural bioinformatics › molecular structure prediction
nucleic acid structure prediction |
0.4 | 1 | 2020 | pqsfinder web: G-quadruplex prediction using optimized pqsfinder algorithm · Bioinform. 2020 |
Bioinformatics and computational biology › genomics › transposable elements
transposable element detection |
0.4 | 1 | 2020 | TE-greedy-nester: structure-based detection of LTR retrotransposons and their nesting · Bioinform. 2020 |
Bioinformatics and computational biology › protein analysis
protein-protein interaction |
0.2 | 1 | 2016 | Hammock: a hidden Markov model-based peptide clustering algorithm to identify protein-interaction consensus motifs in large datasets · Bioinform. 2016 |
Bioinformatics and computational biology
multiple sequence alignment |
0.1 | 1 | 2016 | Hammock: a hidden Markov model-based peptide clustering algorithm to identify protein-interaction consensus motifs in large datasets · Bioinform. 2016 |
Bioinformatics and computational biology › genomics › repetitive DNA analysis
de novo repeat detection |
0.1 | 1 | 2005 | RAP: a new computer program for de novo identification of repeated sequences in whole genomes · Bioinform. 2005 |
Bioinformatics and computational biology › sequence analysis
repeat detection |
0.1 | 1 | 2005 | RAP: a new computer program for de novo identification of repeated sequences in whole genomes · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
nextflow pipeline · 0.6heatmap visualization · 0.6structural analysis · 0.4sequence similarity · 0.4recursive algorithm · 0.4greedy recursive algorithm · 0.4scoring model parametrization · 0.3peptide clustering · 0.2hidden markov model · 0.2triplex DNA search algorithm · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A perspective on genetic and polygenic risk scores - advances and limitations and overview of associated toolsabstractPolygenetic Risk Scores are used to evaluate an individual's vulnerability to developing specific diseases or conditions based on their genetic composition, by taking into account numerous genetic variations. This article provides an overview of the concept of Polygenic Risk Scores (PRS). We elucidate the historical advancements of PRS, their advantages and shortcomings in comparison with other predictive methods, and discuss their conceptual limitations in light of the complexity of biological systems. Furthermore, we provide a survey of published tools for computing PRS and associated resources. The various tools and software packages are categorized based on their technical utility for users or prospective developers. Understanding the array of available tools and their limitations is crucial for accurately assessing and predicting disease risks, facilitating early interventions, and guiding personalized healthcare decisions. Additionally, we also identify potential new avenues for future bioinformatic analyzes and advancements related to PRS. Jana Schwarzerova, Martin Hurta, Vojtech Barton, Matej Lexa, Dirk Walther 0001, Valentine Provazník, Wolfram Weckwerth |
Briefings Bioinform. | 4 |
| 2022 | HiC-TE: a computational pipeline for Hi-C data analysis to study the role of repeat family interactions in the genome 3D organizationabstractMOTIVATION: The role of repetitive DNA in the 3D organization of the interphase nucleus is a subject of intensive study. In studies of 3D nucleus organization, mutual contacts of various loci can be identified by Hi-C sequencing. Typical analyses use binning of read pairs by location to reduce noise. We use binning by repeat families instead to make similar conclusions about repeat regions. RESULTS: To achieve this, we combined Hi-C data, reference genome data and tools for repeat analysis into a Nextflow pipeline identifying and quantifying the contacts of specific repeat families. As an output, our pipeline produces heatmaps showing contact frequency and circular diagrams visualizing repeat contact localization. Using our pipeline with tomato data, we revealed the preferential homotypic interactions of ribosomal DNA, centromeric satellites and some LTR retrotransposon families and, as expected, little contact between organellar and nuclear DNA elements. While the pipeline can be applied to any eukaryotic genome, results in plants provide better coverage, since the built-in TE-greedy-nester software only detects tandems and LTR retrotransposons. Other repeats can be fed via GFF3 files. This pipeline represents a novel and reproducible way to analyze the role of repetitive elements in the 3D organization of genomes. AVAILABILITY AND IMPLEMENTATION: https://gitlab.fi.muni.cz/lexa/hic-te/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matej Lexa, Monika Cechova, Son Hoang Nguyen, Pavel Jedlicka, Viktor Tokan, Zdenek Kubat, Roman Hobza, Eduard Kejnovský |
Bioinform. | 1 |
| 2021 | SARS-CoV-2 hot-spot mutations are significantly enriched within inverted repeats and CpG island lociabstractSARS-CoV-2 is an intensively investigated virus from the order Nidovirales (Coronaviridae family) that causes COVID-19 disease in humans. Through enormous scientific effort, thousands of viral strains have been sequenced to date, thereby creating a strong background for deep bioinformatics studies of the SARS-CoV-2 genome. In this study, we inspected high-frequency mutations of SARS-CoV-2 and carried out systematic analyses of their overlay with inverted repeat (IR) loci and CpG islands. The main conclusion of our study is that SARS-CoV-2 hot-spot mutations are significantly enriched within both IRs and CpG island loci. This points to their role in genomic instability and may predict further mutational drive of the SARS-CoV-2 genome. Moreover, CpG islands are strongly enriched upstream from viral ORFs and thus could play important roles in transcription and the viral life cycle. We hypothesize that hypermethylation of these loci will decrease the transcription of viral ORFs and could therefore limit the progression of the disease. Pratik Goswami, Martin Bartas, Matej Lexa, Natália Bohálová, Adriana Volná, Jirí Cerven, Veronika Cervenová, Petr Pecinka, Vladimír Spunda, Miroslav Fojta, Václav Brázda |
Briefings Bioinform. | 3 |
| 2020 | pqsfinder web: G-quadruplex prediction using optimized pqsfinder algorithmabstractMOTIVATION: G-quadruplex is a DNA or RNA form in which four guanine-rich regions are held together by base pairing between guanine nucleotides in coordination with potassium ions. G-quadruplexes are increasingly seen as a biologically important component of genomes. Their detection in vivo is problematic; however, sequencing and spectrometric techniques exist for their in vitro detection. We previously devised the pqsfinder algorithm for PQS identification, implemented it in C++ and published as an R/Bioconductor package. We looked for ways to optimize pqsfinder for faster and user-friendly sequence analysis. RESULTS: We identified two weak points where pqsfinder could be optimized. We modified the internals of the recursive algorithm to avoid matching and scoring many sub-optimal PQS conformations that are later discarded. To accommodate the needs of a broader range of users, we created a website for submission of sequence analysis jobs that does not require knowledge of R to use pqsfinder. AVAILABILITY AND IMPLEMENTATION: https://pqsfinder.fi.muni.cz, https://bioconductor.org/packages/pqsfinder. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dominika Labudová, Jirí Hon, Matej Lexa |
Bioinform. | 3 |
| 2020 | TE-greedy-nester: structure-based detection of LTR retrotransposons and their nestingabstractMOTIVATION: Transposable elements (TEs) in eukaryotes often get inserted into one another, forming sequences that become a complex mixture of full-length elements and their fragments. The reconstruction of full-length elements and the order in which they have been inserted is important for genome and transposon evolution studies. However, the accumulation of mutations and genome rearrangements over evolutionary time makes this process error-prone and decreases the efficiency of software aiming to recover all nested full-length TEs. RESULTS: We created software that uses a greedy recursive algorithm to mine increasingly fragmented copies of full-length LTR retrotransposons in assembled genomes and other sequence data. The software called TE-greedy-nester considers not only sequence similarity but also the structure of elements. This new tool was tested on a set of natural and synthetic sequences and its accuracy was compared to similar software. We found TE-greedy-nester to be superior in a number of parameters, namely computation time and full-length TE recovery in highly nested regions. AVAILABILITY AND IMPLEMENTATION: http://gitlab.fi.muni.cz/lexa/nested. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matej Lexa, Pavel Jedlicka, Ivan Vanat, Michal Cervenanský, Eduard Kejnovský |
Bioinform. | 1 |
| 2018 | TE-nester: a recursive software tool for structure-based discovery of nested transposable elements
Matej Lexa, Radovan Lapar, Pavel Jedlicka, Ivan Vanat, Michal Cervenanský, Eduard Kejnovský |
BIBM | 1 |
| 2017 | pqsfinder: an exhaustive and imperfection-tolerant search tool for potential quadruplex-forming sequences in RabstractMOTIVATION: G-quadruplexes (G4s) are one of the non-B DNA structures easily observed in vitro and assumed to form in vivo. The latest experiments with G4-specific antibodies and G4-unwinding helicase mutants confirm this conjecture. These four-stranded structures have also been shown to influence a range of molecular processes in cells. As G4s are intensively studied, it is often desirable to screen DNA sequences and pinpoint the precise locations where they might form. RESULTS: We describe and have tested a newly developed Bioconductor package for identifying potential quadruplex-forming sequences (PQS). The package is easy-to-use, flexible and customizable. It allows for sequence searches that accommodate possible divergences from the optimal G4 base composition. A novel aspect of our research was the creation and training (parametrization) of an advanced scoring model which resulted in increased precision compared to similar tools. We demonstrate that the algorithm behind the searches has a 96% accuracy on 392 currently known and experimentally observed G4 structures. We also carried out searches against the recent G4-seq data to verify how well we can identify the structures detected by that technology. The correlation with pqsfinder predictions was 0.622, higher than the correlation 0.491 obtained with the second best G4Hunter. AVAILABILITY AND IMPLEMENTATION: http://bioconductor.org/packages/pqsfinder/ This paper is based on pqsfinder-1.4.1. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jirí Hon, Tomás Martínek, Jaroslav Zendulka, Matej Lexa |
Bioinform. | 4 |
| 2016 | Hammock: a hidden Markov model-based peptide clustering algorithm to identify protein-interaction consensus motifs in large datasetsabstractMOTIVATION: Proteins often recognize their interaction partners on the basis of short linear motifs located in disordered regions on proteins' surface. Experimental techniques that study such motifs use short peptides to mimic the structural properties of interacting proteins. Continued development of these methods allows for large-scale screening, resulting in vast amounts of peptide sequences, potentially containing information on multiple protein-protein interactions. Processing of such datasets is a complex but essential task for large-scale studies investigating protein-protein interactions. RESULTS: The software tool presented in this article is able to rapidly identify multiple clusters of sequences carrying shared specificity motifs in massive datasets from various sources and generate multiple sequence alignments of identified clusters. The method was applied on a previously published smaller dataset containing distinct classes of ligands for SH3 domains, as well as on a new, an order of magnitude larger dataset containing epitopes for several monoclonal antibodies. The software successfully identified clusters of sequences mimicking epitopes of antibody targets, as well as secondary clusters revealing that the antibodies accept some deviations from original epitope sequences. Another test indicates that processing of even much larger datasets is computationally feasible. AVAILABILITY AND IMPLEMENTATION: Hammock is published under GNU GPL v. 3 license and is freely available as a standalone program (from http://www.recamo.cz/en/software/hammock-cluster-peptides/) or as a tool for the Galaxy toolbox (from https://toolshed.g2.bx.psu.edu/view/hammock/hammock). The source code can be downloaded from https://github.com/hammock-dev/hammock/releases. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Adam Krejci, Ted R. Hupp, Matej Lexa, Borivoj Vojtesek, Petr Müller |
Bioinform. | 3 |
| 2013 | Triplex: an R/Bioconductor package for identification and visualization of potential intramolecular triplex patterns in DNA sequencesabstractMOTIVATION: Upgrade and integration of triplex software into the R/Bioconductor framework. RESULTS: We combined a previously published implementation of a triplex DNA search algorithm with visualization to create a versatile R/Bioconductor package 'triplex'. The new package provides functions that can be used to search Bioconductor genomes and other DNA sequence data for occurrence of nucleotide patterns capable of forming intramolecular triplexes (H-DNA). Functions producing 2D and 3D diagrams of the identified triplexes allow instant visualization of the search results. Leveraging the power of Biostrings and GRanges classes, the results get fully integrated into the existing Bioconductor framework, allowing their passage to other Genome visualization and annotation packages, such as GenomeGraphs, rtracklayer or Gviz. AVAILABILITY: R package 'triplex' is available from Bioconductor (bioconductor.org). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jirí Hon, Tomás Martínek, Kamil Rajdl, Matej Lexa |
Bioinform. | 4 |
| 2011 | Architecture model for approximate tandem repeat detectionabstractAlgorithms for biological sequence analysis, such as approximate string matching or algorithms for identification of sequence patterns supporting specific structural elements, present good opportunities for hardware acceleration. Implementation of these algorithms often results in architectures based on multidimensional arrays of computing elements. Mapping effectively these computational structures on FPGAs remains one of the challenging problems. This paper focuses on a specific hardware architecture for detection of approximate tandem repeats in sequences. We show how to create a parametrized architecture model combined with an automatic technique that can determine appropriate circuit dimensions with respect to input task parameters and the target platform properties. Tomás Martínek, Matej Lexa |
ASAP | 2 |
| 2011 | A dynamic programming algorithm for identification of triplex-forming sequencesabstractMOTIVATION: Current methods for identification of potential triplex-forming sequences in genomes and similar sequence sets rely primarily on detecting homopurine and homopyrimidine tracts. Procedures capable of detecting sequences supporting imperfect, but structurally feasible intramolecular triplex structures are needed for better sequence analysis. RESULTS: We modified an algorithm for detection of approximate palindromes, so as to account for the special nature of triplex DNA structures. From available literature, we conclude that approximate triplexes tolerate two classes of errors. One, analogical to mismatches in duplex DNA, involves nucleotides in triplets that do not readily form Hoogsteen bonds. The other class involves geometrically incompatible neighboring triplets hindering proper alignment of strands for optimal hydrogen bonding and stacking. We tested the statistical properties of the algorithm, as well as its correctness when confronted with known triplex sequences. The proposed algorithm satisfactorily detects sequences with intramolecular triplex-forming potential. Its complexity is directly comparable to palindrome searching. AVAILABILITY: Our implementation of the algorithm is available at http://www.fi.muni.cz/lexa/triplex as source code and a web-based search tool. The source code compiles into a library providing searching capability to other programs, as well as into a stand-alone command-line application based on this library. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matej Lexa, Tomás Martínek, Ivana Burgetová, Daniel Kopecek, Marie Brázdová |
Bioinform. | 1 |
| 2010 | Hardware Acceleration of Approximate Tandem Repeat DetectionabstractUnderstanding the structure and function of DNA sequences represents an important area of research in modern biology. Unfortunately, analysis of such data is often complicated by the presence of mutations introduced by evolutionary processes. At the lowest scale, these usually occur in biological sequences as character substitutions, insertions or deletions (indel). They increase the time-complexity of algorithms for sequence analysis by introducing an element of uncertainty, complicating their practical usage. One class of such algorithms has been designed to search for tandem repeats with possible errors - approximate tandem repeats. This paper investigates the possibilities for hardware acceleration of approximate tandem repeat searching and describes a parametrized architecture suitable for chips with FPGA technology. The proposed architecture is able to detect tandems with both types of errors (mismatches and indels) and does not limit the length of detected tandem. A prototype of the circuit was implemented in VHDL language and synthesized for Virtex5 technology. Application on test sequences shows that the circuit is able to speed up tandem searching in orders of thousands in comparison with the best-known software method relying on suffix arrays. Tomás Martínek, Matej Lexa |
FCCM | 2 |
| 2009 | Architecture model for approximate palindrome detectionabstractUnderstanding the structure and function of DNA sequences represents an important area of research in modern biology. One of the interesting structures occurring in DNA is a palindrome. Biologists believe that palindromes play an important role in regulation of gene activity and other cell processes because they are often observed near promoters, introns and specific untranslated regions. Unfortunately, the time complexity of algorithms for palindrome detection increases when mutations in the form of character insertions, deletions or substitutions are taken into consideration. In recent years, several works have been aimed at acceleration of such algorithms using dedicated circuits capable of potentially large-scale searching. However, widespread use of such circuits is often complicated by varying user task details or the need to use a specific target platform. The objective of this work is therefore to create a model of hardware architecture for approximate palindrome detection and develop a technique for automatic mapping of this model to the target platform without intervention of an experienced designer. The proposed model and the mapping technique are implemented and evaluated on a family of chips with Virtex5 technology. Tomás Martínek, Jan Vozenilek, Matej Lexa |
DDECS | 3 |
| 2008 | Hardware acceleration of approximate palindromes searchingabstractUnderstanding the structure and function of DNA sequences represents an important area of research in modern biology. Unfortunately, analysis of such data is often complicated by the presence of mutations introduced by evolutionary processes. They increase the time-complexity of algorithms for sequence analysis by introducing an element of uncertainty, complicating their practical usage. One class of such algorithms has been designed to search for palindromes with possible errors-approximate palindromes. The best state-of-the-art methods implemented in software show time-complexity between linear and quadratic, depending on required input parameters. This paper investigates the possibilities for hardware acceleration of approximate palindrome searching and describes a parametrized architecture suitable for chips with FPGA technology. A prototype of the proposed architecture was implemented in VHDL language and synthesized for Virtex technology. Application on test sequences shows that the circuit is able to speed up palindrome searching by up to 8000× in comparison with the best-known software method relying on suffix arrays. Tomás Martínek, Matej Lexa |
FPT | 2 |
| 2005 | RAP: a new computer program for de novo identification of repeated sequences in whole genomesabstractMOTIVATION: DNA repeats are a common feature of most genomic sequences. Their de novo identification is still difficult despite being a crucial step in genomic analysis and oligonucleotides design. Several efficient algorithms based on word counting are available, but too short words decrease specificity while long words decrease sensitivity, particularly in degenerated repeats. RESULTS: The Repeat Analysis Program (RAP) is based on a new word-counting algorithm optimized for high resolution repeat identification using gapped words. Many different overlapping gapped words can be counted at the same genomic position, thus producing a better signal than the single ungapped word. This results in better specificity both in terms of low-frequency detection, being able to identify sequences repeated only once, and highly divergent detection, producing a generally high score in most intron sequences. AVAILABILITY: The program is freely available for non-profit organizations, upon request to the authors. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: The program has been tested on the Caenorhabditis elegans genome using word lengths of 12, 14 and 16 bases. The full analysis has been implemented in the UCSC Genome Browser and is accessible at http://genome.cribi.unipd.it. Davide Campagna, Chiara Romualdi, Nicola Vitulo, Micky Del Favero, Matej Lexa, Nicola Cannata, Giorgio Valle |
Bioinform. | 5 |
| 2003 | PRIMEX: rapid identification of oligonucleotide matches in whole genomesabstractSUMMARY: PRIMEX (PRImer Match EXtractor) can detect oligonucleotide sequences in whole genomes, allowing for mismatches. Using a word lookup table and server functionality, PRIMEX accepts queries from client software and returns matches rapidly. We find it faster and more sensitive than currently available tools. AVAILABILITY: Running applications and source code have been made available at http://bioinformatics.cribi.unipd.it/primex Matej Lexa, Giorgio Valle |
Bioinform. | 1 |
| 2001 | Virtual PCRabstractAbstract Summary: We present an algorithm that uses public sequence data to predict PCR products. The algorithm is implemented as a CGI script. Output is compared to real-world PCR. Availability: Perl code and instructions for installation are freely available over the internet at http://www.sci.muni.cz/LMFR/vpcr.html Contact: [email protected] * To whom correspondence should be addressed. Matej Lexa, J. Horak, Bretislav Brzobohaty |
Bioinform. | 1 |