Hans-Georg Müller

dblp:67/657 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Functional data analysis for multivariate distributions through Wasserstein slicing
abstract
The modeling of samples of distributions is a major challenge since distributions do not form a vector space. While various approaches exist for univariate distributions, including transformations to a Hilbert space, far less is known about the multivariate case. We utilize a transformation approach to map multivariate distributions to a Hilbert space via a Wasserstein slicing method that is invertible. This approach combines functional data analysis tools, such as functional principal component analysis and modes of variation, with the facility to map back to interpretable distributions. We also provide convergence guarantees for the Hilbert space representations under a broad class of such transforms. The method is illustrated using joint systolic and diastolic blood pressure data.
Hans-Georg Müller
NeurIPS2
2025 Fréchet Geodesic Boosting
abstract
Gradient boosting has become a cornerstone of machine learning, enabling base learners such as decision trees to achieve exceptional predictive performance. While existing algorithms primarily handle scalar or Euclidean outputs, increasingly prevalent complex-structured data, such as distributions, networks, and manifold-valued outputs, present challenges for traditional methods. Such non-Euclidean data lack algebraic structures such as addition, subtraction, or scalar multiplication required by standard gradient boosting frameworks. To address these challenges, we introduce Fréchet geodesic boosting (FGBoost), a novel approach tailored for outputs residing in geodesic metric spaces. FGBoost leverages geodesics as proxies for residuals and constructs ensembles in a way that respects the intrinsic geometry of the output space. Through theoretical analysis, extensive simulations, and real-world applications, we demonstrate the strong performance and adaptability of FGBoost, showcasing its potential for modeling complex data.
Yidong Zhou 0001, Su I Iao, Hans-Georg Müller
NeurIPS3
2025 Conditional Wasserstein Barycenters and Interpolation/Extrapolation of Distributions
abstract
Increasingly complex data analysis tasks motivate the study of the dependency of distributions of multivariate continuous random variables on scalar or vector predictors. Statistical regression models for distributional responses are a key technique for the emerging field of distributional data analysis, but so far have primarily been investigated for the case of one-dimensional response distributions. We investigate here the challenging case of conditional Fréchet means for multivariate response distributions under the Wasserstein metric, which has not been studied previously but is relevant for statistical data analysis in various fields, including climatology and health, as we demonstrate with real data applications. A second innovation is that we harness the notion of conditional barycenters and geodesics in the Wasserstein space to interpolate as well as extrapolate multivariate distributions under suitable regularity conditions, where even the simpler case of extrapolating one-dimensional distributions has not been studied before. We cover both global parametric-like and local smoothing-like models to implement conditional Wasserstein barycenters and establish asymptotic convergence properties for the corresponding estimates under suitable regularity assumptions. The utility of distributional inter- and extrapolation is explored in both simulations and examples. Conditional Wasserstein barycenters and distribution extrapolation are specifically illustrated with applications in epidemiology and climatology.
Jianing Fan, Hans-Georg Müller
IEEE Trans. Inf. Theory2
2023 Go with the Flow: Personalized Task Sequencing Improves Online Language Learning
Nathalie Rzepka, Katharina Simbeck, Hans-Georg Müller, Niels Pinkwart
AIED3
2022 Keep It Up: In-session Dropout Prediction to Support Blended Classroom Scenarios
Nathalie Rzepka, Katharina Simbeck, Hans-Georg Müller, Niels Pinkwart
CSEDU (2)3
2022 An Online Controlled Experiment Design to Support the Transformation of Digital Learning towards Adaptive Learning Platforms
Nathalie Rzepka, Katharina Simbeck, Hans-Georg Müller, Niels Pinkwart
CSEDU (2)3
2022 Fairness of In-session Dropout Prediction
Nathalie Rzepka, Katharina Simbeck, Hans-Georg Müller, Niels Pinkwart
CSEDU (2)3
2022 Network Regression with Graph Laplacians
abstract
Network data are increasingly available in various research fields, motivating statistical analysis for populations of networks, where a network as a whole is viewed as a data point. The study of how a network changes as a function of covariates is often of paramount interest. However, due to the non-Euclidean nature of networks, basic statistical tools available for scalar and vector data are no longer applicable. This motivates an extension of the notion of regression to the case where responses are network data. Here we propose to adopt conditional Fréchet means implemented as M-estimators that depend on weights derived from both global and local least squares regression, extending the Fréchet regression framework to networks that are quantified by their graph Laplacians. The challenge is to characterize the space of graph Laplacians to justify the application of Fréchet regression. This characterization then leads to asymptotic rates of convergence for the corresponding M-estimators by applying empirical process methods. We demonstrate the usefulness and good practical performance of the proposed framework with simulations and with network data arising from resting-state fMRI in neuroimaging, as well as New York taxi records.
Yidong Zhou 0001, Hans-Georg Müller
J. Mach. Learn. Res.2
2022 Cox Point Process Regression
abstract
Point processes in time have a wide range of applications that include the claims arrival process in insurance or the analysis of queues in operations research. Due to advances in technology, such samples of point processes are increasingly encountered. A key object of interest is the local intensity function. It has a straightforward interpretation that allows to understand and explore point process data. We consider functional approaches for point processes, where one has a sample of repeated realizations of the point process. This situation is inherently connected with Cox processes, where the intensity functions of the replications are modeled as random functions. Here we study a situation where one records covariates for each replication of the process, such as the daily temperature for bike rentals. For modeling point processes as responses with vector covariates as predictors we propose a novel regression approach for the intensity function that is intrinsically nonparametric. While the intensity function of a point process that is only observed once on a fixed domain cannot be identified, we show how covariates and repeated observations of the process can be utilized to make consistent estimation possible, and we also derive asymptotic rates of convergence without invoking parametric assumptions.
Álvaro Gajardo, Hans-Georg Müller
IEEE Trans. Inf. Theory2
2021 What you apply is not what you learn! Examining students' strategies in German capitalization tasks
Nathalie Rzepka, Hans-Georg Müller, Katharina Simbeck
EDM2
2010 Functional embedding for the classification of gene expression profiles
abstract
Abstract Motivation: Low sample size n high-dimensional large p data with n≪p are commonly encountered in genomics and statistical genetics. Ill-conditioning of the variance-covariance matrix for such data renders the traditional multivariate data analytical approaches unattractive. On the other side, functional data analysis (FDA) approaches are designed for infinite-dimensional data and therefore may have potential for the analysis of large p data. We herein propose a functional embedding (FEM) technique, which exploits the interface between multivariate and functional data, aiming at borrowing strength across the sample through FDA techniques in order to resolve the difficulties caused by the high dimension p. Results: Using pairwise dissimilarities among predictor variables, one obtains a univariate configuration of these covariates. This is interpreted as variable ordination that defines the domain of a suitable function space, thus leading to the FEM of the high-dimensional data. The embedding may then be followed by functional logistic regression for the classification of high-dimensional multivariate data as an example for downstream analysis. The resulting functional classification is evaluated on several published gene expression array datasets and a mass spectrometric data, and is shown to compare favorably with various methods that have been employed previously for the classification of these high-dimensional gene expression profiles. Availability: The implementation of FEM and Classification via Functional Embedding (CFEM) as described in this article was done with the PACE package written in Matlab. The latest version of PACE is publicly accessible at http://anson.ucdavis.edu/∼mueller/data/programs.html. An example MATLAB script for FEM is available at http://www.lehigh.edu/∼psw205/psw205.html Contact: [email protected]; [email protected]
Ping-Shi Wu, Hans-Georg Müller
Bioinform.2
2008 Inferring gene expression dynamics via functional regression analysis
abstract
BACKGROUND: Temporal gene expression profiles characterize the time-dynamics of expression of specific genes and are increasingly collected in current gene expression experiments. In the analysis of experiments where gene expression is obtained over the life cycle, it is of interest to relate temporal patterns of gene expression associated with different developmental stages to each other to study patterns of long-term developmental gene regulation. We use tools from functional data analysis to study dynamic changes by relating temporal gene expression profiles of different developmental stages to each other. RESULTS: We demonstrate that functional regression methodology can pinpoint relationships that exist between temporary gene expression profiles for different life cycle phases and incorporates dimension reduction as needed for these high-dimensional data. By applying these tools, gene expression profiles for pupa and adult phases are found to be strongly related to the profiles of the same genes obtained during the embryo phase. Moreover, one can distinguish between gene groups that exhibit relationships with positive and others with negative associations between later life and embryonal expression profiles. Specifically, we find a positive relationship in expression for muscle development related genes, and a negative relationship for strictly maternal genes for Drosophila, using temporal gene expression profiles. CONCLUSION: Our findings point to specific reactivation patterns of gene expression during the Drosophila life cycle which differ in characteristic ways between various gene groups. Functional regression emerges as a useful tool for relating gene expression patterns from different developmental stages, and avoids the problems with large numbers of parameters and multiple testing that affect alternative approaches.
Hans-Georg Müller, Jeng-Min Chiou, Xiaoyan Leng
BMC Bioinform.1
2006 Classification using functional data analysis for temporal gene expression data
abstract
MOTIVATION: Temporal gene expression profiles provide an important characterization of gene function, as biological systems are predominantly developmental and dynamic. We propose a method of classifying collections of temporal gene expression curves in which individual expression profiles are modeled as independent realizations of a stochastic process. The method uses a recently developed functional logistic regression tool based on functional principal components, aimed at classifying gene expression curves into known gene groups. The number of eigenfunctions in the classifier can be chosen by leave-one-out cross-validation with the aim of minimizing the classification error. RESULTS: We demonstrate that this methodology provides low-error-rate classification for both yeast cell-cycle gene expression profiles and Dictyostelium cell-type specific gene expression patterns. It also works well in simulations. We compare our functional principal components approach with a B-spline implementation of functional discriminant analysis for the yeast cell-cycle data and simulations. This indicates comparative advantages of our approach which uses fewer eigenfunctions/base functions. The proposed methodology is promising for the analysis of temporal gene expression data and beyond. AVAILABILITY: MATLAB programs are available upon request.
Xiaoyan Leng, Hans-Georg Müller
Bioinform.2
2003 Modes and clustering for time-warped gene expression profile data
abstract
MOTIVATION: The study of the dynamics of regulatory processes has led to increased interest for the analysis of temporal gene expression level data. To address the dynamics of regulation, expression data are collected repeatedly over time. It is difficult to statistically represent the resulting high-dimensional data. When regulatory processes determine gene expression, time-warping is likely to be present, i.e. the sample of gene expression trajectories reflects variation not only in terms of the expression amplitudes, but also in terms of the temporal structure of gene expression. RESULTS: A non-parametric time-synchronized iterative mean updating technique is proposed to find an overall representation that corresponds to a mode of a sample of expression profiles, viewed as a random sample in function space. The proposed algorithm explores the application of previous work of Hall and Heckman to genome-wide expression data and provides an extension that includes random time-warping with the aim to synchronize timescales across genes. The proposed algorithm is universally applicable for the construction of modes for functional data with time-warping. We demonstrate the construction of mode functions for a sample of Drosophila gene expression data. The algorithm can be applied to define clusters among the observed trajectories of gene expression, without any kind of prior non-time-warped clustering, as illustrated in the numerical example.
Hans-Georg Müller
Bioinform.2