Michael G. Schimek

dblp:47/5796 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0002-1712-6668ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
array CGH analysis
0.112009
MSMAD: a computationally efficient method for the analysis of noisy array CGH data · Bioinform. 2009
Bioinformatics and computational biology › genomics
breakpoint detection
0.112009
MSMAD: a computationally efficient method for the analysis of noisy array CGH data · Bioinform. 2009
Bioinformatics and computational biology › cancer genomics
copy number analysis
0.112009
MSMAD: a computationally efficient method for the analysis of noisy array CGH data · Bioinform. 2009
Bioinformatics and computational biology › genomics
genome analysis
0.112009
MSMAD: a computationally efficient method for the analysis of noisy array CGH data · Bioinform. 2009

Methods — techniques the papers use, named apart from their topics

median smoothing · 0.1median absolute deviation · 0.1
YearPublicationVenuePosition
2024 Effective signal reconstruction from multiple ranked lists via convex optimization
abstract
Abstract The ranking of objects is widely used to rate their relative quality or relevance across multiple assessments. Beyond classical rank aggregation, it is of interest to estimate the usually unobservable latent signals that inform a consensus ranking. Under the only assumption of independent assessments, which can be incomplete, we introduce indirect inference via convex optimization in combination with computationally efficient Poisson Bootstrap. Two different objective functions are suggested, one linear and the other quadratic. The mathematical formulation of the signal estimation problem is based on pairwise comparisons of all objects with respect to their rank positions. Sets of constraints represent the order relations. The transitivity property of rank scales allows us to reduce substantially the number of constraints associated with the full set of object comparisons. The key idea is to globally reduce the errors induced by the rankers until optimal latent signals can be obtained. Its main advantage is low computational costs, even when handling $$n < < p$$ n < < p data problems. Exploratory tools can be developed based on the bootstrap signal estimates and standard errors. Simulation evidence, a comparison with the state-of-the-art rank centrality method, and two applications, one in higher education evaluation and the other in molecular cancer research, are presented.
Michael G. Schimek, Luca Vitale, Bastian Pfeifer, Michele La Rocca 0001
Data Min. Knowl. Discov.1
2024 Correction to: Effective signal reconstruction from multiple ranked lists via convex optimization
Michael G. Schimek, Luca Vitale, Bastian Pfeifer, Michele La Rocca 0001
Data Min. Knowl. Discov.1
2023 Parea: Multi-view ensemble clustering for cancer subtype discovery
abstract
Multi-view clustering methods are essential for the stratification of patients into sub-groups of similar molecular characteristics. In recent years, a wide range of methods have been developed for this purpose. However, due to the high diversity of cancer-related data, a single method may not perform sufficiently well in all cases. We present Parea, a multi-view hierarchical ensemble clustering approach for disease subtype discovery. We demonstrate its performance on several machine learning benchmark datasets. We apply and validate our methodology on real-world multi-view patient data, comprising seven types of cancer. Parea outperforms the current state-of-the-art on six out of seven analysed cancer types. We have integrated the Parea method into our Python package Pyrea (https://github.com/mdbloice/Pyrea), which enables the effortless and flexible design of ensemble workflows while incorporating a wide range of fusion and clustering algorithms.
Bastian Pfeifer, Marcus D. Bloice, Michael G. Schimek
J. Biomed. Informatics3
2021 Integrative hierarchical ensemble clustering for improved disease subtype discovery
abstract
Multi-omics clustering methods are used for the stratification of patients into sub-groups of similar molecular characteristics. In recent years, a wide range of methods has been developed for this purpose. However, due to the high diversity of cancer-related data, a single method may not perform sufficiently well in all cases. Here, we propose a comprehensive framework for multi-omics hierarchical ensemble clustering. We provide a flexible environment that allows to build hierarchical clustering ensembles suitable for the available data and research goals. Survival analyses for data from The Cancer Genome Atlas (TCGA) indicate that our proposed ensembles provide more robust, and thus more reliable results than the state-of-the-art. We have implemented our architecture within the R-package HC-fused, which is freely available on Github.
Bastian Pfeifer, Andrei Voicu-Spineanu, Michael G. Schimek, Nikolaos Alachiotis 0001
BIBM3
2021 A hierarchical clustering and data fusion approach for disease subtype discovery
abstract
Recent advances in multi-omics clustering methods enable a more fine-tuned separation of cancer patients into clinical relevant clusters. These advancements have the potential to provide a deeper understanding of cancer progression and may facilitate the treatment of cancer patients. Here, we present a simple hierarchical clustering and data fusion approach, named HC-fused, for the detection of disease subtypes. Unlike other methods, the proposed approach naturally reports on the individual contribution of each single-omic to the data fusion process. We perform multi-view simulations with disjoint and disjunct cluster elements across the views to highlight fundamentally different data integration behavior of various state-of-the-art methods. HC-fused combines the strengths of some recently published methods and shows superior performance on real world cancer data from the TCGA (The Cancer Genome Atlas) database. An R implementation of our method is available on GitHub (pievos101/HC-fused).
Bastian Pfeifer, Michael G. Schimek
J. Biomed. Informatics2
2009 MSMAD: a computationally efficient method for the analysis of noisy array CGH data
abstract
MOTIVATION: Genome analysis has become one of the most important tools for understanding the complex process of cancerogenesis. With increasing resolution of CGH arrays, the demand for computationally efficient algorithms arises, which are effective in the detection of aberrations even in very noisy data. RESULTS: We developed a rather simple, non-parametric technique of high computational efficiency for CGH array analysis that adopts a median absolute deviation concept for breakpoint detection, comprising median smoothing for pre-processing. The resulting algorithm has the potential to outperform any single smoothing approach as well as several recently proposed segmentation techniques. We show its performance through the application of simulated and real datasets in comparison to three other methods for array CGH analysis. IMPLEMENTATION: Our approach is implemented in the R-language and environment for statistical computing (version 2.6.1 for Windows, R-project, 2007). The code is available at: http://www.iba.muni.cz/~budinska/msmad.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Eva Budinska, Eva Gelnarova, Michael G. Schimek
Bioinform.3