Bernd Fischer 0003

dblp:27/3809-3 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
0since 2021 · last 2015
0000-0001-9437-2099ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 95% Medical and health informatics · 5%
Databases, data mining, and information retrieval
2 papers
Data mining · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 61% Segmentation and scene understanding · 39%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
cancer genomics
0.212015
SomaticSignatures: inferring mutational signatures from single-nucleotide variants · Bioinform. 2015
Bioinformatics and computational biology
genomics
0.212014
h5vc: scalable nucleotide tallies with HDF5 · Bioinform. 2014
Bioinformatics and computational biology › genomics › computational genomics
SNP detection
0.212014
h5vc: scalable nucleotide tallies with HDF5 · Bioinform. 2014
Bioinformatics and computational biology › genomics
variant calling
0.212014
h5vc: scalable nucleotide tallies with HDF5 · Bioinform. 2014
Bioinformatics and computational biology
high-content screening
0.212013
CellH5: a format for data exchange in high-content screening · Bioinform. 2013
Bioinformatics and computational biology
proteomics
0.122007
PepSplice: cache-efficient search algorithms for comprehensive identification of tandem mass spectra · Bioinform. 2007
A Hidden Markov Model for de Novo Peptide Sequencing · NIPS 2004
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
generalized linear model
0.112008
The Group-Lasso for generalized linear models: uniqueness of solutions and efficient algorithms · ICML 2008
Bioinformatics and computational biology › computational neuroscience
computational neuroanatomy
0.112008
Probabilistic image registration and anomaly detection by nonlinear warping · CVPR 2008
Bioinformatics and computational biology › bioimage informatics
electron microscopy image analysis
0.112008
Probabilistic image registration and anomaly detection by nonlinear warping · CVPR 2008
Medical and health informatics › medical imaging › medical image analysis
image registration
0.112008
Probabilistic image registration and anomaly detection by nonlinear warping · CVPR 2008
Data mining
clustering
0.122003
Bagging for Path-Based Clustering · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Clustering with the Connectivity Kernel · NIPS 2003
Mathematical optimization › regularization
group lasso
0.112008
The Group-Lasso for generalized linear models: uniqueness of solutions and efficient algorithms · ICML 2008
Mathematical optimization › statistical estimation › regression
regularized regression
0.112008
The Group-Lasso for generalized linear models: uniqueness of solutions and efficient algorithms · ICML 2008
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number estimation
0.112014
h5vc: scalable nucleotide tallies with HDF5 · Bioinform. 2014
Bioinformatics and computational biology › proteomics › peptide sequencing
de novo peptide sequencing
0.012004
A Hidden Markov Model for de Novo Peptide Sequencing · NIPS 2004
Computer vision › Segmentation and scene understanding
perceptual grouping
0.012003
Path-Based Clustering for Grouping of Smooth Curves and Texture Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Data mining › clustering
ensemble clustering
0.012003
Bagging for Path-Based Clustering · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Data mining › clustering
kernel clustering
0.012003
Clustering with the Connectivity Kernel · NIPS 2003
Image and video processing › geometric correction
geometric distortion correction
0.012008
Probabilistic image registration and anomaly detection by nonlinear warping · CVPR 2008
Image and video processing
image restoration
0.012008
Probabilistic image registration and anomaly detection by nonlinear warping · CVPR 2008
Bioinformatics and computational biology › proteomics › post-translational modification
post-translational modification identification
0.012007
PepSplice: cache-efficient search algorithms for comprehensive identification of tandem mass spectra · Bioinform. 2007
Computer vision › Segmentation and scene understanding
image segmentation
0.012003
Bagging for Path-Based Clustering · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Image and video processing › image segmentation
texture segmentation
0.012003
Path-Based Clustering for Grouping of Smooth Curves and Texture Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2003

Methods — techniques the papers use, named apart from their topics

pattern decomposition · 0.2non-negative matrix factorization · 0.2statistical computing · 0.2HDF5 · 0.2HDF5 format · 0.2solution path algorithm · 0.2polynomial kernel expansion · 0.2outlier detection · 0.2expectation-maximization · 0.2active group identification · 0.2resampling · 0.1connectedness measure · 0.1bagging · 0.1agglomerative optimization · 0.1agglomerative algorithm · 0.1hypergeometric scoring · 0.1cache-efficient search algorithm · 0.1spectral clustering · 0.0
YearPublicationVenuePosition
2015 SomaticSignatures: inferring mutational signatures from single-nucleotide variants
abstract
UNLABELLED: Mutational signatures are patterns in the occurrence of somatic single-nucleotide variants that can reflect underlying mutational processes. The SomaticSignatures package provides flexible, interoperable and easy-to-use tools that identify such signatures in cancer sequencing data. It facilitates large-scale, cross-dataset estimation of mutational signatures, implements existing methods for pattern decomposition, supports extension through user-defined approaches and integrates with existing Bioconductor workflows. AVAILABILITY AND IMPLEMENTATION: The R package SomaticSignatures is available as part of the Bioconductor project. Its documentation provides additional details on the methods and demonstrates applications to biological datasets. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Julian Gehring, Bernd Fischer 0003, Michael F. Lawrence, Wolfgang Huber
Bioinform.2
2014 h5vc: scalable nucleotide tallies with HDF5
abstract
SUMMARY: As applications of genome sequencing, including exomes and whole genomes, are expanding, there is a need for analysis tools that are scalable to large sets of samples and/or ultra-deep coverage. Many current tool chains are based on the widely used file formats BAM and VCF or VCF-derivatives. However, for some desirable analyses, data management with these formats creates substantial implementation overhead, and much time is spent parsing files and collating data. We observe that a tally data structure, i.e. the table of counts of nucleotides × samples × strands × genomic positions, provides a reasonable intermediate level of abstraction for many genomics analyses, including single nucleotide variant (SNV) and InDel calling, copy-number estimation and mutation spectrum analysis. Here we present h5vc, a data structure and associated software for managing tallies. The software contains functionality for creating tallies from BAM files, flexible and scalable data visualization, data quality assessment, computing statistics relevant to variant calling and other applications. Through the simplicity of its API, we envision making low-level analysis of large sets of genome sequencing data accessible to a wider range of researchers. AVAILABILITY AND IMPLEMENTATION: The package H5VC for the statistical environment R is available through the Bioconductor project. The HDF5 system is used as the core of our implementation. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Paul Theodor Pyl, Julian Gehring, Bernd Fischer 0003, Wolfgang Huber
Bioinform.3
2013 CellH5: a format for data exchange in high-content screening
abstract
UNLABELLED: High-throughput microscopy data require a diversity of analytical approaches. However, the construction of workflows that use algorithms from different software packages is difficult owing to a lack of interoperability. To overcome this limitation, we present CellH5, an HDF5 data format for cell-based assays in high-throughput microscopy, which stores high-dimensional image data along with inter-object relations in graphs. CellH5Browser, an interactive gallery image browser, demonstrates the versatility and performance of the file format on live imaging data of dividing human cells. CellH5 provides new opportunities for integrated data analysis by multiple software platforms. AVAILABILITY: Source code is freely available at www.github.com/cellh5 under the GPL license and at www.bioconductor.org/packages/release/bioc/html/rhdf5.html under the Artistic-2.0 license. Demo datasets and the CellH5Browser are available at www.cellh5.org. A Fiji importer for cellh5 will be released soon.
Christoph Sommer 0002, Michael Held, Bernd Fischer 0003, Wolfgang Huber, Daniel Gerlich
Bioinform.3
2011 Extracting quantitative genetic interaction phenotypes from matrix combinatorial RNAi
abstract
BACKGROUND: Systematic measurement of genetic interactions by combinatorial RNAi (co-RNAi) is a powerful tool for mapping functional modules and discovering components. It also provides insights into the role of epistasis on the way from genotype to phenotype. The interpretation of co-RNAi data requires computational and statistical analysis in order to detect interactions reliably and sensitively. RESULTS: We present a comprehensive approach to the analysis of univariate phenotype measurements, such as cell growth. The method is based on a quantitative model and is demonstrated on two example Drosophila cell culture data sets. We discuss adjustments for technical variability, data quality assessment, model parameter fitting and fit diagnostics, choice of scale, and assessment of statistical significance. CONCLUSIONS: As a result, we obtain quantitative genetic interactions and interaction networks reflecting known biological relationships between target genes. The reliable extraction of presence, absence, and strength of interactions provides insights into molecular mechanisms.
Elin Axelsson, Thomas Sandmann, Thomas Horn, Michael Boutros, Wolfgang Huber, Bernd Fischer 0003
BMC Bioinform.6
2009 Optimal Transitions for Targeted Protein Quantification: Best Conditioned Submatrix Selection
Rastislav Srámek, Bernd Fischer 0003, Elias Vicari, Peter Widmayer
COCOON2
2009 Adaptive bandwidth selection for biomarker discovery in mass spectrometry
Bernd Fischer 0003, Volker Roth 0001, Joachim M. Buhmann
Artif. Intell. Medicine1
2008 Probabilistic image registration and anomaly detection by nonlinear warping
abstract
Automatic, defect tolerant registration of transmission electron microscopy (TEM) images poses an important and challenging problem for biomedical image analysis, e.g. in computational neuroanatomy. In this paper we demonstrate a fully automatic stitching and distortion correction method for TEM images and propose a probabilistic approach for image registration. The technique identifies image defects due to sample preparation and image acquisition by outlier detection. A polynomial kernel expansion is used to estimate a non-linear image transformation based on intensities and spatial features. Corresponding points in the images are not determined beforehand, but they are estimated via an EM-algorithm during the registration process which is preferable in the case of (noisy) TEM images. Our registration model is successfully applied to two large image stacks of serial section TEM images acquired from brain tissue samples in a computational neuroanatomy project and shows significant improvement over existing image registration methods on these large datasets.
Verena Kaynig, Bernd Fischer 0003, Joachim M. Buhmann
CVPR2
2008 The Group-Lasso for generalized linear models: uniqueness of solutions and efficient algorithms
abstract
The Group-Lasso method for finding important explanatory factors suffers from the potential non-uniqueness of solutions and also from high computational costs. We formulate conditions for the uniqueness of Group-Lasso solutions which lead to an easily implementable test procedure that allows us to identify all potentially active groups. These results are used to derive an efficient algorithm that can deal with input dimensions in the millions and can approximate the solution path efficiently. The derived methods are applied to large-scale learning problems where they exhibit excellent performance and where the testing procedure helps to avoid misinterpretations of the solutions.
Volker Roth 0001, Bernd Fischer 0003
ICML2
2007 PepSplice: cache-efficient search algorithms for comprehensive identification of tandem mass spectra
abstract
Abstract Motivation: Tandem mass spectrometry allows for high-throughput identification of complex protein samples. Searching tandem mass spectra against sequence databases is the main analysis method nowadays. Since many peptide variations are possible, including them in the search space seems only logical. However, the search space usually grows exponentially with the number of independent variations and may therefore overwhelm computational resources. Results: We provide fast, cache-efficient search algorithms to screen large peptide search spaces including non-tryptic peptides, whole genomes, dozens of posttranslational modifications, unannotated point mutations and even unannotated splice sites. All these search spaces can be screened simultaneously. By optimizing the cache usage, we achieve a calculation speed that closely approaches the limits of the hardware. At the same time, we control the size of the overall search space by limiting the combinations of variations that can co-occur on the same peptide. Using a hypergeometric scoring scheme, we applied these algorithms to a dataset of 1 420 632 spectra. We were able to identify a considerable number of peptide variations within a modest amount of computing time on standard desktop computers. Availability: PepSplice is available as a C++ application for Linux, Windows and OSX at www.ti.inf.ethz.ch/pw/software/pepsplice/. It is open source under the revised BSD license. Contact: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Franz F. Roos, Riko Jacob, Jonas Grossmann, Bernd Fischer 0003, Joachim M. Buhmann, Wilhelm Gruissem, Sacha Baginsky, Peter Widmayer
Bioinform.4
2007 Time-series alignment by non-negative multiple generalized canonical correlation analysis
abstract
BACKGROUND: Quantitative analysis of differential protein expressions requires to align temporal elution measurements from liquid chromatography coupled to mass spectrometry (LC/MS). We propose multiple Canonical Correlation Analysis (mCCA) as a method to align the non-linearly distorted time scales of repeated LC/MS experiments in a robust way. RESULTS: Multiple canonical correlation analysis is able to map several time series to a consensus time scale. The alignment function is learned in a supervised fashion. We compare our approach with previously published methods for aligning mass spectrometry data on a large proteomics dataset. The proposed method significantly increases the number of proteins that are identified as being differentially expressed in different biological samples. CONCLUSION: Jointly aligning multiple liquid chromatography/mass spectrometry samples by mCCA substantially increases the detection rate of potential bio-markers which significantly improves the interpretability of LC/MS data.
Bernd Fischer 0003, Volker Roth 0001, Joachim M. Buhmann
BMC Bioinform.1
2007 Improved functional prediction of proteins by learning kernel combinations in multilabel settings
abstract
BACKGROUND: We develop a probabilistic model for combining kernel matrices to predict the function of proteins. It extends previous approaches in that it can handle multiple labels which naturally appear in the context of protein function. RESULTS: Explicit modeling of multilabels significantly improves the capability of learning protein function from multiple kernels. The performance and the interpretability of the inference model are further improved by simultaneously predicting the subcellular localization of proteins and by combining pairwise classifiers to consistent class membership estimates. CONCLUSION: For the purpose of functional prediction of proteins, multilabels provide valuable information that should be included adequately in the training process of classifiers. Learning of functional categories gains from co-prediction of subcellular localization. Pairwise separation rules allow very detailed insights into the relevance of different measurements like sequence, structure, interaction data, or expression data. A preliminary version of the software can be downloaded from http://www.inf.ethz.ch/personal/vroth/KernelHMM/.
Volker Roth 0001, Bernd Fischer 0003
BMC Bioinform.2
2004 A Hidden Markov Model for de Novo Peptide Sequencing
abstract
De novo Sequencing of peptides is a challenging task in proteome re- search. While there exist reliable DNA-sequencing methods, the high- throughput de novo sequencing of proteins by mass spectrometry is still an open problem. Current approaches suffer from a lack in precision to detect mass peaks in the spectrograms. In this paper we present a novel method for de novo peptide sequencing based on a hidden Markov model. Experiments effectively demonstrate that this new method signif- icantly outperforms standard approaches in matching quality.
Bernd Fischer 0003, Volker Roth 0001, Joachim M. Buhmann, Jonas Grossmann, Sacha Baginsky, Wilhelm Gruissem, Franz F. Roos, Peter Widmayer
NIPS1
2003 Clustering with the Connectivity Kernel
abstract
Clustering aims at extracting hidden structure in dataset. While the prob- lem of finding compact clusters has been widely studied in the litera- ture, extracting arbitrarily formed elongated structures is considered a much harder problem. In this paper we present a novel clustering algo- rithm which tackles the problem by a two step procedure: first the data are transformed in such a way that elongated structures become compact ones. In a second step, these new objects are clustered by optimizing a compactness-based criterion. The advantages of the method over related approaches are threefold: (i) robustness properties of compactness-based criteria naturally transfer to the problem of extracting elongated struc- tures, leading to a model which is highly robust against outlier objects; (ii) the transformed distances induce a Mercer kernel which allows us to formulate a polynomial approximation scheme to the generally NP- hard clustering problem; (iii) the new method does not contain free kernel parameters in contrast to methods like spectral clustering or mean-shift clustering.
Bernd Fischer 0003, Volker Roth 0001, Joachim M. Buhmann
NIPS1
2003 Path-Based Clustering for Grouping of Smooth Curves and Texture Segmentation
abstract
Perceptual grouping organizes image parts in clusters based on psychophysically plausible similarity measures. We propose a novel grouping method in this paper, which stresses connectedness of image elements via mediating elements rather than favoring high mutual similarity. This grouping principle yields superior clustering results when objects are distributed on low-dimensional extended manifolds in a feature space, and not as local point clouds. In addition to extracting connected structures, objects are singled out as outliers when they are too far away from any cluster structure. The objective function for this perceptual organization principle is optimized by a fast agglomerative algorithm. We report on perceptual organization experiments where small edge elements are grouped to smooth curves. The generality of the method is emphasized by results from grouping textured images with texture gradients in an unsupervised fashion.
Bernd Fischer 0003, Joachim M. Buhmann
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Bagging for Path-Based Clustering
abstract
A resampling scheme for clustering with similarity to bootstrap aggregation (bagging) is presented. Bagging is used to improve the quality of path-based clustering, a data clustering method that can extract elongated structures from data in a noise robust way. The results of an agglomerative optimization method are influenced by small fluctuations of the input data. To increase the reliability of clustering solutions, a stochastic resampling method is developed to infer consensus clusters. A related reliability measure allows us to estimate the number of clusters, based on the stability of an optimized cluster solution under resampling. The quality of path-based clustering with resampling is evaluated on a large image data set of human segmentations.
Bernd Fischer 0003, Joachim M. Buhmann
IEEE Trans. Pattern Anal. Mach. Intell.1