Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Marc Rehmsmeier

dblp:35/3795 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
3since 2021 · last 2022
0000-0002-5021-7721ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
sequence analysis
0.112009
mkESA: enhanced suffix array construction tool · Bioinform. 2009
Algorithms and data structures › sequence algorithms › string algorithms
string indexing
0.112009
mkESA: enhanced suffix array construction tool · Bioinform. 2009
Algorithms and data structures › sequence algorithms › string algorithms › string indexing
suffix array construction
0.112009
mkESA: enhanced suffix array construction tool · Bioinform. 2009
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics › RNA structure prediction
RNA secondary structure prediction
0.112006
RNAshapes: an integrated RNA analysis package based on abstract shapes · Bioinform. 2006
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search
0.122000
Sequence Database Search Using Jumping Alignments · ISMB 2000
WWW access to the SYSTERS protein sequence cluster set · Bioinform. 1999
Bioinformatics and computational biology
sequence alignment
0.012000
Sequence Database Search Using Jumping Alignments · ISMB 2000
Bioinformatics and computational biology
protein sequence analysis
0.011999
WWW access to the SYSTERS protein sequence cluster set · Bioinform. 1999
Bioinformatics and computational biology › protein sequence analysis
protein sequence clustering
0.011999
WWW access to the SYSTERS protein sequence cluster set · Bioinform. 1999
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics › RNA structure prediction
RNA folding
0.012006
RNAshapes: an integrated RNA analysis package based on abstract shapes · Bioinform. 2006

Methods — techniques the papers use, named apart from their topics

precision-recall curve · 0.3ROC curve · 0.3parallelization · 0.2deep-shallow suffix array construction · 0.2dynamic programming · 0.1jumping alignment · 0.0sequence similarity search · 0.0multiple sequence alignment · 0.0
YearPublicationVenuePosition
2022 MGcount: a total RNA-seq quantification tool to address multi-mapping and multi-overlapping alignments ambiguity in non-coding transcripts
abstract
BACKGROUND: Total-RNA sequencing (total-RNA-seq) allows the simultaneous study of both the coding and the non-coding transcriptome. Yet, computational pipelines have traditionally focused on particular biotypes, making assumptions that are not fullfilled by total-RNA-seq datasets. Transcripts from distinct RNA biotypes vary in length, biogenesis, and function, can overlap in a genomic region, and may be present in the genome with a high copy number. Consequently, reads from total-RNA-seq libraries may cause ambiguous genomic alignments, demanding for flexible quantification approaches. RESULTS: Here we present Multi-Graph count (MGcount), a total-RNA-seq quantification tool combining two strategies for handling ambiguous alignments. First, MGcount assigns reads hierarchically to small-RNA and long-RNA features to account for length disparity when transcripts overlap in the same genomic position. Next, MGcount aggregates RNA products with similar sequences where reads systematically multi-map using a graph-based approach. MGcount outputs a transcriptomic count matrix compatible with RNA-sequencing downstream analysis pipelines, with both bulk and single-cell resolution, and the graphs that model repeated transcript structures for different biotypes. The software can be used as a python module or as a single-file executable program. CONCLUSIONS: MGcount is a flexible total-RNA-seq quantification tool that successfully integrates reads that align to multiple genomic locations or that overlap with multiple gene features. Its approach is suitable for the simultaneous estimation of protein-coding, long non-coding and small non-coding transcript concentration, in both precursor and processed forms. Both source code and compiled software are available at https://github.com/hitaandrea/MGcount .
Andrea Hita, Gilles Brocart, Ana Fernandez, Marc Rehmsmeier, Anna Alemany, Sol Schvartzman
BMC Bioinform.4
2022 Correction: MGcount: a total RNA-seq quantification tool to address multi-mapping and multi-overlapping alignments ambiguity in non-coding transcripts
abstract
identified an error in Fig. 2 and caption 2c.The correct figure is given below, and the caption has been updated from ''Reads ri (i = 1, 10)'' to '' Reads ri (i = 1
Andrea Hita, Gilles Brocart, Ana Fernandez, Marc Rehmsmeier, Anna Alemany, Sol Schvartzman
BMC Bioinform.4
2021 MOCCA: a flexible suite for modelling DNA sequence motif occurrence combinatorics
abstract
BACKGROUND: Cis-regulatory elements (CREs) are DNA sequence segments that regulate gene expression. Among CREs are promoters, enhancers, Boundary Elements (BEs) and Polycomb Response Elements (PREs), all of which are enriched in specific sequence motifs that form particular occurrence landscapes. We have recently introduced a hierarchical machine learning approach (SVM-MOCCA) in which Support Vector Machines (SVMs) are applied on the level of individual motif occurrences, modelling local sequence composition, and then combined for the prediction of whole regulatory elements. We used SVM-MOCCA to predict PREs in Drosophila and found that it was superior to other methods. However, we did not publish a polished implementation of SVM-MOCCA, which can be useful for other researchers, and we only tested SVM-MOCCA with IUPAC motifs and PREs. RESULTS: We here present an expanded suite for modelling CRE sequences in terms of motif occurrence combinatorics-Motif Occurrence Combinatorics Classification Algorithms (MOCCA). MOCCA contains efficient implementations of several modelling methods, including SVM-MOCCA, and a new method, RF-MOCCA, a Random Forest-derivative of SVM-MOCCA. We used SVM-MOCCA and RF-MOCCA to model Drosophila PREs and BEs in cross-validation experiments, making this the first study to model PREs with Random Forests and the first study that applies the hierarchical MOCCA approach to the prediction of BEs. Both models significantly improve generalization to PREs and boundary elements beyond that of previous methods-including 4-spectrum and motif occurrence frequency Support Vector Machines and Random Forests-, with RF-MOCCA yielding the best results. CONCLUSION: MOCCA is a flexible and powerful suite of tools for the motif-based modelling of CRE sequences in terms of motif composition. MOCCA can be applied to any new CRE modelling problems where motifs have been identified. MOCCA supports IUPAC and Position Weight Matrix (PWM) motifs. For ease of use, MOCCA implements generation of negative training data, and additionally a mode that requires only that the user specifies positives, motifs and a genome. MOCCA is licensed under the MIT license and is available on Github at https://github.com/bjornbredesen/MOCCA .
Bjørn André Bredesen, Marc Rehmsmeier
BMC Bioinform.2
2017 Precrec: fast and accurate precision-recall and ROC curve calculations in R
abstract
The precision-recall plot is more informative than the ROC plot when evaluating classifiers on imbalanced datasets, but fast and accurate curve calculation tools for precision-recall plots are currently not available. We have developed Precrec, an R library that aims to overcome this limitation of the plot. Our tool provides fast and accurate precision-recall calculations together with multiple functionalities that work efficiently under different conditions. AVAILABILITY AND IMPLEMENTATION: Precrec is licensed under GPL-3 and freely available from CRAN (https://cran.r-project.org/package=precrec). It is implemented in R with C ++. CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Takaya Saito, Marc Rehmsmeier
Bioinform.2
2009 mkESA: enhanced suffix array construction tool
abstract
Abstract Summary: We introduce the tool mkESA, an open source program for constructing enhanced suffix arrays (ESAs), striving for low memory consumption, yet high practical speed. mkESA is a user-friendly program written in portable C99, based on a parallelized version of the Deep-Shallow suffix array construction algorithm, which is known for its high speed and small memory usage. The tool handles large FASTA files with multiple sequences, and computes suffix arrays and various additional tables, such as the LCP table (longest common prefix) or the inverse suffix array, from given sequence data. Availability: The source code of mkESA is freely available under the terms of the GNU General Public License (GPL) version 2 at http://bibiserv.techfak.uni-bielefeld.de/mkesa/. Contact: [email protected]
Robert Homann, David Fleer, Robert Giegerich, Marc Rehmsmeier
Bioinform.4
2006 RNAshapes: an integrated RNA analysis package based on abstract shapes
abstract
Abstract Summary: We introduce RNAshapes, a new software package that integrates three RNA analysis tools based on the abstract shapes approach: the analysis of shape representatives, the calculation of shape probabilities and the consensus shapes approach. This new package is completely reimplemented in C and outruns the original implementations significantly in runtime and memory requirements. Additionally, we added a number of useful features like suboptimal folding with correct dangling energies, structure graph output, shape matching and a sliding window approach. Availability: RNAshapes is freely available at as C source code, and as compiled binaries for the most common computer architectures. For Microsoft Windows, we also offer a graphical user interface with convenient access to the complete functionality of the package. Contact: [email protected]
Peter Steffen, Björn Voß, Marc Rehmsmeier, Jens Reeder, Robert Giegerich
Bioinform.3
2002 Phase4: Automatic Evaluation of Database Search Methods
abstract
It has become standard to evaluate newly devised database search methods in terms of sensitivity and selectivity and to compare them with existing methods. This involves the construction of a suitable evaluation scenario, the execution of the methods, the assessment of their performances, and the presentation of the results. Each of these four phases and their smooth connection usually imposes formidable work. To relieve the evaluator of this burden, a system has been designed with which evaluations can be effected rapidly. It is implemented in the programming language Python whose object-oriented features are used to offer a great flexibility in changing the evaluation design. A graphical user interface is provided which offers the usual amenities such as radio- and checkbuttons or file browsing facilities.
Marc Rehmsmeier
Briefings Bioinform.1
2000 Sequence Database Search Using Jumping Alignments
Rainer Spang, Marc Rehmsmeier, Jens Stoye
ISMB2
1999 WWW access to the SYSTERS protein sequence cluster set
abstract
SUMMARY: We present a Web server where the SYSTERS cluster set of the non-redundant protein database consisting of sequences from SWISS-PROT and PIR is being made available for querying and browsing. The cluster set can be searched with a new sequence using the SSMAL search tool. Additionally, a multiple alignment is generated for each cluster and annotated with domain information from the Pfam protein family database. AVAILABILITY: The server address is http://www.dkfz-heidelberg.de/tbi/services/cluster/ systersform
Antje Krause, Pierre Nicodème, Erich Bornberg-Bauer, Marc Rehmsmeier, Martin Vingron
Bioinform.4