Marc Kirchner

dblp:59/7099 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 44% Machine translation · 44% Probabilistic and Bayesian machine learning · 13%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
data annotation
0.812024
On Efficient and Statistical Quality Estimation for Data Annotation · ACL (1) 2024
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
0.812024
On Efficient and Statistical Quality Estimation for Data Annotation · ACL (1) 2024
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
0.332011
libfbi: a C++ implementation for fast box intersection and application to sparse mass spectrometry data · Bioinform. 2011
Non-linear classification for on-the-fly fractional mass filtering and targeted precursor fragmentation in mass spectrometry experiments · Bioinform. 2010
Computational protein profile similarity screening for quantitative mass spectrometry experiments · Bioinform. 2010
Bioinformatics and computational biology
proteomics
0.332011
SIMA: Simultaneous Multiple Alignment of LC/MS Peak Lists · Bioinform. 2011
Deuteration distribution estimation with improved sequence coverage for HX/MS experiments · Bioinform. 2010
Non-linear classification for on-the-fly fractional mass filtering and targeted precursor fragmentation in mass spectrometry experiments · Bioinform. 2010
Machine learning and data management
model evaluation
0.212022
Neo: Generalizing Confusion Matrix Visualization to Hierarchical and Multi-Output Labels · CHI 2022
Bioinformatics and computational biology › metabolomics
LC-MS data analysis
0.112011
SIMA: Simultaneous Multiple Alignment of LC/MS Peak Lists · Bioinform. 2011
Bioinformatics and computational biology › metabolomics
retention time alignment
0.112011
SIMA: Simultaneous Multiple Alignment of LC/MS Peak Lists · Bioinform. 2011
Bioinformatics and computational biology › proteomics
quantitative proteomics
0.112010
Computational protein profile similarity screening for quantitative mass spectrometry experiments · Bioinform. 2010

Methods — techniques the papers use, named apart from their topics

probability distribution algebra · 1.1formative study · 1.1confidence intervals · 0.8acceptance sampling · 0.8space partitioning data structure · 0.1multidimensional kernel function · 0.1maximum likelihood estimation · 0.1supervised classification · 0.1statistical test for equality · 0.1random forest classification · 0.1l1-regularized feature extraction · 0.1discrete mapping · 0.1correlation-based distance measure · 0.1
YearPublicationVenuePosition
2024 On Efficient and Statistical Quality Estimation for Data Annotation
abstract
Annotated datasets are an essential ingredient to train, evaluate, compare and productionalize supervised machine learning models.It is therefore imperative that annotations are of high quality.For their creation, good quality management and thereby reliable quality estimates are needed.Then, if quality is insufficient during the annotation process, rectifying measures can be taken to improve it.Quality estimation is often performed by having experts manually label instances as correct or incorrect.But checking all annotated instances tends to be expensive.Therefore, in practice, usually only subsets are inspected; sizes are chosen mostly without justification or regard to statistical power and more often than not, are relatively small.Basing estimates on small sample sizes, however, can lead to imprecise values for the error rate.Using unnecessarily large sample sizes costs money that could be better spent, for instance on more annotations.Therefore, we first describe in detail how to use confidence intervals for finding the minimal sample size needed to estimate the annotation error rate.Then, we propose applying acceptance sampling as an alternative to error rate estimation We show that acceptance sampling can reduce the required sample sizes up to 50% while providing the same statistical guarantees.
Jan-Christoph Klie, Juan Haladjian, Marc Kirchner
ACL (1)3
2022 Neo: Generalizing Confusion Matrix Visualization to Hierarchical and Multi-Output Labels
abstract
The confusion matrix, a ubiquitous visualization for helping people evaluate machine learning models, is a tabular layout that compares predicted class labels against actual class labels over all data instances. We conduct formative research with machine learning practitioners at Apple and find that conventional confusion matrices do not support more complex data-structures found in modern-day applications, such as hierarchical and multi-output labels. To express such variations of confusion matrices, we design an algebra that models confusion matrices as probability distributions. Based on this algebra, we develop Neo, a visual analytics system that enables practitioners to flexibly author and interact with hierarchical and multi-output confusion matrices, visualize derived metrics, renormalize confusions, and share matrix specifications. Finally, we demonstrate Neo’s utility with three model evaluation scenarios that help people better understand model performance and reveal hidden confusions.
Jochen Görtler, Fred Hohman, Dominik Moritz, Kanit Wongsuphasawat, Donghao Ren, Marc Kirchner, Kayur Patel
CHI7
2011 libfbi: a C++ implementation for fast box intersection and application to sparse mass spectrometry data
abstract
MOTIVATION: Algorithms for sparse data require fast search and subset selection capabilities for the determination of point neighborhoods. A natural data representation for such cases are space partitioning data structures. However, the associated range queries assume noise-free observations and cannot take into account observation-specific uncertainty estimates that are present in e.g. modern mass spectrometry data. In order to accommodate the inhomogeneous noise characteristics of sparse real-world datasets, point queries need to be reformulated in terms of box intersection queries, where box sizes correspond to uncertainty regions for each observation. RESULTS: This contribution introduces libfbi, a standard C++, header-only template implementation for fast box intersection in an arbitrary number of dimensions, with arbitrary data types in each dimension. The implementation is applied to a data aggregation task on state-of-the-art liquid chromatography/mass spectrometry data, where it shows excellent run time properties. AVAILABILITY: The library is available under an MIT license and can be downloaded from http://software.steenlab.org/libfbi. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marc Kirchner, Buote Xu, Hanno Steen, Judith A. J. Steen
Bioinform.1
2011 SIMA: Simultaneous Multiple Alignment of LC/MS Peak Lists
abstract
MOTIVATION: Alignment of multiple liquid chromatography/mass spectrometry (LC/MS) experiments is a necessity today, which arises from the need for biological and technical repeats. Due to limits in sampling frequency and poor reproducibility of retention times, current LC systems suffer from missing observations and non-linear distortions of the retention times across runs. Existing approaches for peak correspondence estimation focus almost exclusively on solving the pairwise alignment problem, yielding straightforward but suboptimal results for multiple alignment problems. RESULTS: We propose SIMA, a novel automated procedure for alignment of peak lists from multiple LC/MS runs. SIMA combines hierarchical pairwise correspondence estimation with simultaneous alignment and global retention time correction. It employs a tailored multidimensional kernel function and a procedure based on maximum likelihood estimation to find the retention time distortion function that best fits the observed data. SIMA does not require a dedicated reference spectrum, is robust with regard to outliers, needs only two intuitive parameters and naturally incorporates incomplete correspondence information. In a comparison with seven alternative methods on four different datasets, we show that SIMA yields competitive and superior performance on real-world data. AVAILABILITY: A C++ implementation of the SIMA algorithm is available from http://hci.iwr.uni-heidelberg.de/MIP/Software.
Björn Voß, Michael Hanselmann, Bernhard Y. Renard, Martin S. Lindner, Ullrich Köthe, Marc Kirchner, Fred A. Hamprecht
Bioinform.6
2010 Computational protein profile similarity screening for quantitative mass spectrometry experiments
abstract
MOTIVATION: The qualitative and quantitative characterization of protein abundance profiles over a series of time points or a set of environmental conditions is becoming increasingly important. Using isobaric mass tagging experiments, mass spectrometry-based quantitative proteomics deliver accurate peptide abundance profiles for relative quantitation. Associated data analysis workflows need to provide tailored statistical treatment that (i) takes the correlation structure of the normalized peptide abundance profiles into account and (ii) allows inference of protein-level similarity. We introduce a suitable distance measure for relative abundance profiles, derive a statistical test for equality and propose a protein-level representation of peptide-level measurements. This yields a workflow that delivers a similarity ranking of protein abundance profiles with respect to a defined reference. All procedures have in common that they operate based on the true correlation structure that underlies the measurements. This optimizes power and delivers more intuitive and efficient results than existing methods that do not take these circumstances into account. RESULTS: We use protein profile similarity screening to identify candidate proteins whose abundances are post-transcriptionally controlled by the Anaphase Promoting Complex/Cyclosome (APC/C), a specific E3 ubiquitin ligase that is a master regulator of the cell cycle. Results are compared with an established protein correlation profiling method. The proposed procedure yields a 50.9-fold enrichment of co-regulated protein candidates and a 2.5-fold improvement over the previous method. AVAILABILITY: A MATLAB toolbox is available from http://hci.iwr.uni-heidelberg.de/mip/proteomics.
Marc Kirchner, Bernhard Y. Renard, Ullrich Köthe, Darryl J. Pappin, Fred A. Hamprecht, Hanno Steen, Judith A. J. Steen
Bioinform.1
2010 Non-linear classification for on-the-fly fractional mass filtering and targeted precursor fragmentation in mass spectrometry experiments
abstract
MOTIVATION: Mass spectrometry (MS) has become the method of choice for protein/peptide sequence and modification analysis. The technology employs a two-step approach: ionized peptide precursor masses are detected, selected for fragmentation, and the fragment mass spectra are collected for computational analysis. Current precursor selection schemes are based on data- or information-dependent acquisition (DDA/IDA), where fragmentation mass candidates are selected by intensity and are subsequently included in a dynamic exclusion list to avoid constant refragmentation of highly abundant species. DDA/IDA methods do not exploit valuable information that is contained in the fractional mass of high-accuracy precursor mass measurements delivered by current instrumentation. RESULTS: We extend previous contributions that suggest that fractional mass information allows targeted fragmentation of analytes of interest. We introduce a non-linear Random Forest classification and a discrete mapping approach, which can be trained to discriminate among arbitrary fractional mass patterns for an arbitrary number of classes of analytes. These methods can be used to increase fragmentation efficiency for specific subsets of analytes or to select suitable fragmentation technologies on-the-fly. We show that theoretical generalization error estimates transfer into practical application, and that their quality depends on the accuracy of prior distribution estimate of the analyte classes. The methods are applied to two real-world proteomics datasets. AVAILABILITY: All software used in this study is available from http://software.steenlab.org/fmf CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marc Kirchner, Wiebke Timm, Peying Fong, Philine Wangemann, Hanno Steen
Bioinform.1
2010 Deuteration distribution estimation with improved sequence coverage for HX/MS experiments
abstract
MOTIVATION: Time-resolved hydrogen exchange (HX) followed by mass spectrometry (MS) is a key technology for studying protein structure, dynamics and interactions. HX experiments deliver a time-dependent distribution of deuteration levels of peptide sequences of the protein of interest. The robust and complete estimation of this distribution for as many peptide fragments as possible is instrumental to understanding dynamic protein-level HX behavior. Currently, this data interpretation step still is a bottleneck in the overall HX/MS workflow. RESULTS: We propose HeXicon, a novel algorithmic workflow for automatic deuteration distribution estimation at increased sequence coverage. Based on an L(1)-regularized feature extraction routine, HeXicon extracts the full deuteration distribution, which allows insight into possible bimodal exchange behavior of proteins, rather than just an average deuteration for each time point. Further, it is capable of addressing ill-posed estimation problems, yielding sparse and physically reasonable results. HeXicon makes use of existing peptide sequence information, which is augmented by an inferred list of peptide candidates derived from a known protein sequence. In conjunction with a supervised classification procedure that balances sensitivity and specificity, HeXicon can deliver results with increased sequence coverage. AVAILABILITY: The entire HeXicon workflow has been implemented in C++ and includes a graphical user interface. It is available at http://hci.iwr.uni-heidelberg.de/software.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xinghua Lou, Marc Kirchner, Bernhard Y. Renard, Ullrich Köthe, Sebastian Boppel, Christian Graf 0002, Chung-Tien Lee, Judith A. J. Steen, Hanno Steen, Matthias P. Mayer, Fred A. Hamprecht
Bioinform.2
2008 NITPICK: peak identification for mass spectrometry data
abstract
BACKGROUND: The reliable extraction of features from mass spectra is a fundamental step in the automated analysis of proteomic mass spectrometry (MS) experiments. RESULTS: This contribution proposes a sparse template regression approach to peak picking called NITPICK. NITPICK is a Non-greedy, Iterative Template-based peak PICKer that deconvolves complex overlapping isotope distributions in multicomponent mass spectra. NITPICK is based on fractional averaging, a novel extension to Senko's well-known averaging model, and on a modified version of sparse, non-negative least angle regression, for which a suitable, statistically motivated early stopping criterion has been derived. The strength of NITPICK is the deconvolution of overlapping mixture mass spectra. CONCLUSION: Extensive comparative evaluation has been carried out and results are provided for simulated and real-world data sets. NITPICK outperforms pepex, to date the only alternate, publicly available, non-greedy feature extraction routine. NITPICK is available as software package for the R programming language and can be downloaded from (http://hci.iwr.uni-heidelberg.de/mip/proteomics/).
Bernhard Y. Renard, Marc Kirchner, Hanno Steen, Judith A. J. Steen, Fred A. Hamprecht
BMC Bioinform.2