Nicolai Meinshausen

dblp:21/2269 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Probabilistic and Bayesian machine learning · 33% Learning theory · 21% Trustworthy machine learning · 15%
Theoretical computer science
4 papers
Algorithms and data structures · 76% Algorithmic game theory and mechanism design · 24%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 88% Computational science and engineering · 12%

Topics — the 21 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
conditional density estimation
0.822023
Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression · J. Mach. Learn. Res. 2022
Confidence and Uncertainty Assessment for Distributional Random Forests · J. Mach. Learn. Res. 2023
Machine learning › Learning theory › statistical estimation › asymptotic estimation theory
asymptotic distribution
0.712023
Confidence and Uncertainty Assessment for Distributional Random Forests · J. Mach. Learn. Res. 2023
Natural language and speech › Information extraction and text analysis
bootstrapping
0.712023
Confidence and Uncertainty Assessment for Distributional Random Forests · J. Mach. Learn. Res. 2023
Machine learning › Learning theory › statistical estimation › confidence set construction
confidence intervals
0.712023
Confidence and Uncertainty Assessment for Distributional Random Forests · J. Mach. Learn. Res. 2023
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest
0.712023
Confidence and Uncertainty Assessment for Distributional Random Forests · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning
statistical inference
0.712023
Confidence and Uncertainty Assessment for Distributional Random Forests · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
distribution regression
0.612022
Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression · J. Mach. Learn. Res. 2022
Machine learning › Trustworthy machine learning › interpretability
sensitivity analysis
0.612022
Scalable Sensitivity and Uncertainty Analyses for Causal-Effect Estimates of Continuous-Valued Interventions · NeurIPS 2022
Machine learning › Trustworthy machine learning
uncertainty estimation
0.612022
Scalable Sensitivity and Uncertainty Analyses for Causal-Effect Estimates of Continuous-Valued Interventions · NeurIPS 2022
Computational social science and digital humanities
causal inference
0.612022
Scalable Sensitivity and Uncertainty Analyses for Causal-Effect Estimates of Continuous-Valued Interventions · NeurIPS 2022
Machine learning › Learning theory
high-dimensional regression
0.312018
The xyz algorithm for fast interaction search in high-dimensional data · J. Mach. Learn. Res. 2018
Algorithms and data structures
randomized algorithms
0.312018
The xyz algorithm for fast interaction search in high-dimensional data · J. Mach. Learn. Res. 2018
Machine learning › Optimization for machine learning
large-scale regression
0.312017
On $b$-bit Min-wise Hashing for Large-scale Regression and Classification with Sparse Data · J. Mach. Learn. Res. 2017
Algorithms and data structures › data structure design › search structures › hashing
minwise hashing
0.312017
On $b$-bit Min-wise Hashing for Large-scale Regression and Classification with Sparse Data · J. Mach. Learn. Res. 2017
Machine learning › Optimization for machine learning › adaptive optimization
adaptive stochastic optimization
0.212016
Scalable Adaptive Stochastic Optimization Using Random Projections · NIPS 2016
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.212016
Scalable Adaptive Stochastic Optimization Using Random Projections · NIPS 2016
Machine learning › Optimization for machine learning
stochastic gradient descent
0.212016
Scalable Adaptive Stochastic Optimization Using Random Projections · NIPS 2016
Algorithmic game theory and mechanism design › social choice
robust aggregation
0.212016
Magging: Maximin Aggregation for Inhomogeneous Large-Scale Data · Proc. IEEE 2016
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.212015
BACKSHIFT: Learning causal cyclic graphs from unknown shift interventions · NIPS 2015
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
cyclic causal models
0.212015
BACKSHIFT: Learning causal cyclic graphs from unknown shift interventions · NIPS 2015
Machine learning › Probabilistic and Bayesian machine learning › causal inference
latent confounders
0.212022
Scalable Sensitivity and Uncertainty Analyses for Causal-Effect Estimates of Continuous-Valued Interventions · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

random forest · 1.2uncertainty-aware deep model · 1.1neural network · 1.1closest pair problem · 0.7bootstrap · 0.7maximum mean discrepancy · 0.6kernel weighting · 0.6subsampling · 0.5robust statistics · 0.5mean aggregation · 0.5randomized algorithms · 0.3randomized algorithm · 0.3random projection · 0.2adagrad · 0.2
YearPublicationVenuePosition
2023 Confidence and Uncertainty Assessment for Distributional Random Forests
abstract
The Distributional Random Forest (DRF) is a recently introduced Random Forest algorithm to estimate multivariate conditional distributions. Due to its general estimation procedure, it can be employed to estimate a wide range of targets such as conditional average treatment effects, conditional quantiles, and conditional correlations. However, only results about the consistency and convergence rate of the DRF prediction are available so far. We characterize the asymptotic distribution of DRF and develop a bootstrap approximation of it. This allows us to derive inferential tools for quantifying standard errors and the construction of confidence regions that have asymptotic coverage guarantees. In simulation studies, we empirically validate the developed theory for inference of low-dimensional targets and for testing distributional differences between two populations
Jeffrey Näf, Corinne Emmenegger, Peter Bühlmann, Nicolai Meinshausen
J. Mach. Learn. Res.4
2022 Scalable Sensitivity and Uncertainty Analyses for Causal-Effect Estimates of Continuous-Valued Interventions
abstract
Estimating the effects of continuous-valued interventions from observational data is a critically important task for climate science, healthcare, and economics. Recent work focuses on designing neural network architectures and regularization functions to allow for scalable estimation of average and individual-level dose-response curves from high-dimensional, large-sample data. Such methodologies assume ignorability (observation of all confounding variables) and positivity (observation of all treatment levels for every covariate value describing a set of units), assumptions problematic in the continuous treatment regime. Scalable sensitivity and uncertainty analyses to understand the ignorance induced in causal estimates when these assumptions are relaxed are less studied. Here, we develop a continuous treatment-effect marginal sensitivity model (CMSM) and derive bounds that agree with the observed data and a researcher-defined level of hidden confounding. We introduce a scalable algorithm and uncertainty-aware deep models to derive and estimate these bounds for high-dimensional, large-sample observational data. We work in concert with climate scientists interested in the climatological impacts of human emissions on cloud properties using satellite observations from the past 15 years. This problem is known to be complicated by many unobserved confounders.
Andrew Jesson, Alyson Douglas, Peter Manshausen, Maëlys Solal, Nicolai Meinshausen, Philip Stier, Yarin Gal, Uri Shalit
NeurIPS5
2022 Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression
abstract
Random Forest is a successful and widely used regression and classification algorithm. Part of its appeal and reason for its versatility is its (implicit) construction of a kernel-type weighting function on training data, which can also be used for targets other than the original mean estimation. We propose a novel forest construction for multivariate responses based on their joint conditional distribution, independent of the estimation target and the data model. It uses a new splitting criterion based on the MMD distributional metric, which is suitable for detecting heterogeneity in multivariate distributions. The induced weights define an estimate of the full conditional distribution, which in turn can be used for arbitrary and potentially complicated targets of interest. The method is very versatile and convenient to use, as we illustrate on a wide range of examples. The code is available as Python and R packages drf.
Domagoj Cevid, Loris Michel, Jeffrey Näf, Peter Bühlmann, Nicolai Meinshausen
J. Mach. Learn. Res.5
2021 Conditional variance penalties and domain shift robustness
abstract
Abstract When training a deep neural network for image classification, one can broadly distinguish between two types of latent features of images that will drive the classification. We can divide latent features into (i) ‘core’ or ‘conditionally invariant’ features $$C$$ C whose distribution $$C\vert Y$$ C|Y , conditional on the classY, does not change substantially across domains and (ii) ‘style’ features $$S$$ S whose distribution $$S\vert Y$$ S|Y can change substantially across domains. Examples for style features include position, rotation, image quality or brightness but also more complex ones like hair color, image quality or posture for images of persons. Our goal is to minimize a loss that is robust under changes in the distribution of these style features. In contrast to previous work, we assume that the domain itself is not observed and hence a latent variable. We do assume that we can sometimes observe a typically discrete identifier or “ $$\mathrm {ID}$$ ID variable”. In some applications we know, for example, that two images show the same person, and $$\mathrm {ID}$$ ID then refers to the identity of the person. The proposed method requires only a small fraction of images to have $$\mathrm {ID}$$ ID information. We group observations if they share the same class and identifier $$(Y,\mathrm {ID})=(y,\mathrm {id})$$ (Y,ID)=(y,id) and penalize the conditional variance of the prediction or the loss if we condition on $$(Y,\mathrm {ID})$$ (Y,ID) . Using a causal framework, this conditional variance regularization (CoRe) is shown to protect asymptotically against shifts in the distribution of the style variables in a partially linear structural equation model. Empirically, we show that the CoRepenalty improves predictive accuracy substantially in settings where domain changes occur in terms of image quality, brightness and color while we also look at more complex changes such as changes in movement and posture.
Christina Heinze-Deml, Nicolai Meinshausen
Mach. Learn.2
2018 The xyz algorithm for fast interaction search in high-dimensional data
abstract
When performing regression on a data set with $p$ variables, it is often of interest to go beyond using main linear effects and include interactions as products between individual variables. For small-scale problems, these interactions can be computed explicitly but this leads to a computational complexity of at least $\mathcal{O}(p^2)$ if done naively. This cost can be prohibitive if $p$ is very large. We introduce a new randomised algorithm that is able to discover interactions with high probability and under mild conditions has a runtime that is subquadratic in $p$. We show that strong interactions can be discovered in almost linear time, whilst finding weaker interactions requires $\mathcal{O}(p^\alpha)$ operations for $1<\alpha<2$ depending on their strength. The underlying idea is to transform interaction search into a closest pair problem which can be solved efficiently in subquadratic time. The algorithm is called $xyz$ and is implemented in the language R. We demonstrate its efficiency for application to genome-wide association studies, where more than $10^{11}$ interactions can be screened in under $280$ seconds with a single-core $1.2$ GHz CPU.
Gian-Andrea Thanei, Nicolai Meinshausen, Rajen Dinesh Shah
J. Mach. Learn. Res.2
2017 On $b$-bit Min-wise Hashing for Large-scale Regression and Classification with Sparse Data
Rajen Dinesh Shah, Nicolai Meinshausen
J. Mach. Learn. Res.2
2016 DUAL-LOCO: Distributing Statistical Estimation Using Random Projections
abstract
We present DUAL-LOCO, a communication-efficient algorithm for distributed statistical estimation. DUAL-LOCO assumes that the data is distributed across workers according to the features rather than the samples. It requires only a single round of communication where low-dimensional random projections are used to approximate the dependencies between features available to different workers. We show that DUAL-LOCO has bounded approximation error which only depends weakly on the number of workers. We compare DUAL-LOCO against a state-of-the-art distributed optimization method on a variety of real world datasets and show that it obtains better speedups while retaining good accuracy. In particular, DUAL-LOCO allows for fast cross validation as only part of the algorithm depends on the regularization parameter.
Christina Heinze-Deml, Brian McWilliams, Nicolai Meinshausen
AISTATS3
2016 Scalable Adaptive Stochastic Optimization Using Random Projections
abstract
Adaptive stochastic gradient methods such as AdaGrad have gained popularity in particular for training deep neural networks. The most commonly used and studied variant maintains a diagonal matrix approximation to second order information by accumulating past gradients which are used to tune the step size adaptively. In certain situations the full-matrix variant of AdaGrad is expected to attain better performance, however in high dimensions it is computationally impractical. We present Ada-LR and RadaGrad two computationally efficient approximations to full-matrix AdaGrad based on randomized dimensionality reduction. They are able to capture dependencies between features and achieve similar performance to full-matrix AdaGrad but at a much smaller computational cost. We show that the regret of Ada-LR is close to the regret of full-matrix AdaGrad which can have an up-to exponentially smaller dependence on the dimension than the diagonal variant. Empirically, we show that Ada-LR and RadaGrad perform similarly to full-matrix AdaGrad. On the task of training convolutional neural networks as well as recurrent neural networks, RadaGrad achieves faster convergence than diagonal AdaGrad.
Gabriel Krummenacher, Brian McWilliams, Yannic Kilcher, Joachim M. Buhmann, Nicolai Meinshausen
NIPS5
2016 Magging: Maximin Aggregation for Inhomogeneous Large-Scale Data
abstract
Large-scale data analysis poses both statistical and computational problems which need to be addressed simultaneously. A solution is often straightforward if the data are homogeneous: one can use classical ideas of subsampling and mean aggregation to get a computationally efficient solution with acceptable statistical accuracy, where the aggregation step simply averages the results obtained on distinct subsets of the data. However, if the data exhibit inhomogeneities (and typically they do), the same approach will be inadequate, as it will be unduly influenced by effects that are not persistent across all the data due to, for example, outliers or time-varying effects. We show that a tweak to the aggregation step can produce an estimator of effects which are common to all data, and hence interesting for interpretation and often leading to better prediction than pooled effects.
Peter Bühlmann, Nicolai Meinshausen
Proc. IEEE2
2015 BACKSHIFT: Learning causal cyclic graphs from unknown shift interventions
abstract
We propose a simple method to learn linear causal cyclic models in the presence of latent variables. The method relies on equilibrium data of the model recorded under a specific kind of interventions (``shift interventions''). The location and strength of these interventions do not have to be known and can be estimated from the data. Our method, called BACKSHIFT, only uses second moments of the data and performs simple joint matrix diagonalization, applied to differences between covariance matrices. We give a sufficient and necessary condition for identifiability of the system, which is fulfilled almost surely under some quite general assumptions if and only if there are at least three distinct experimental settings, one of which can be pure observational data. We demonstrate the performance on some simulated data and applications in flow cytometry and financial time series.
Dominik Rothenhäusler, Christina Heinze-Deml, Jonas Peters, Nicolai Meinshausen
NIPS4
2014 Random intersection trees
Rajen Dinesh Shah, Nicolai Meinshausen
J. Mach. Learn. Res.2
2006 Quantile Regression Forests
abstract
Random forests were introduced as a machine learning tool in Breiman (2001) and have since proven to be very popular and powerful for high-dimensional regression and classification. For regression, random forests give an accurate approximation of the conditional mean of a response variable. It is shown here that random forests provide information about the full conditional distribution of the response variable, not only about the conditional mean. Conditional quantiles can be inferred with quantile regression forests, a generalisation of random forests. Quantile regression forests give a non-parametric and accurate way of estimating conditional quantiles for high-dimensional predictor variables. The algorithm is shown to be consistent. Numerical examples suggest that the algorithm is competitive in terms of predictive power.
Nicolai Meinshausen
J. Mach. Learn. Res.1