VLDB 2026 Research / reviewers in the wild / expert
Matthew J. Holland
dblp:148/9989
· DBLP profile ↗
19ranked-venue papers
19as first author
11since 2021 · last 2024
0000-0002-6704-1769ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 17 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Survey of Learning Criteria Going beyond the Usual Risk (Abstract Reprint)abstractVirtually all machine learning tasks are characterized using some form of loss function, and "good performance" is typically stated in terms of a sufficiently small average loss, taken over the random draw of test data. While optimizing for performance on average is intuitive, convenient to analyze in theory, and easy to implement in practice, such a choice brings about trade-offs. In this work, we survey and introduce a wide variety of non-traditional criteria used to design and evaluate machine learning algorithms, place the classical paradigm within the proper historical context, and propose a view of learning problems which emphasizes the question of "what makes for a desirable loss distribution?" in place of tacit use of the expected loss. Matthew J. Holland, Kazuki Tanabe |
AAAI | 1 |
| 2024 | Robust variance-regularized risk minimization with concomitant scaling
Matthew J. Holland |
AISTATS | 1 |
| 2024 | Criterion Collapse and Loss Distribution ControlabstractIn this work, we consider the notion of "criterion collapse," in which optimization of one metric implies optimality in another, with a particular focus on conditions for collapse into error probability minimizers under a wide variety of learning criteria, ranging from DRO and OCE risks (CVaR, tilted ERM) to non-monotonic criteria underlying recent ascent-descent algorithms explored in the literature (Flooding, SoftAD). We show how collapse in the context of losses with a Bernoulli distribution goes far beyond existing results for CVaR and DRO, then expand our scope to include surrogate losses, showing conditions where monotonic criteria such as tilted ERM cannot avoid collapse, whereas non-monotonic alternatives can. Matthew J. Holland |
ICML | 1 |
| 2024 | Soft ascent-descent as a stable and flexible alternative to floodingabstractAs a heuristic for improving test accuracy in classification, the "flooding" method proposed by Ishida et al. (2020) sets a threshold for the average surrogate loss at training time; above the threshold, gradient descent is run as usual, but below the threshold, a switch to gradient *ascent* is made. While setting the threshold is non-trivial and is usually done with validation data, this simple technique has proved remarkably effective in terms of accuracy. On the other hand, what if we are also interested in other metrics such as model complexity or average surrogate loss at test time? As an attempt to achieve better overall performance with less fine-tuning, we propose a softened, pointwise mechanism called SoftAD (soft ascent-descent) that downweights points on the borderline, limits the effects of outliers, and retains the ascent-descent effect of flooding, with no additional computational overhead. We contrast formal stationarity guarantees with those for flooding, and empirically demonstrate how SoftAD can realize classification accuracy competitive with flooding (and the more expensive alternative SAM) while enjoying a much smaller loss generalization gap and model norm. Matthew J. Holland, Kosuke Nakatani |
NeurIPS | 1 |
| 2023 | Flexible risk design using bi-directional dispersionabstractMany novel notions of “risk” (e.g., CVaR, tilted risk, DRO risk) have been proposed and studied, but these risks are all at least as sensitive as the mean to loss tails on the upside, and tend to ignore deviations on the downside. We study a complementary new risk class that penalizes loss deviations in a bi-directional manner, while having more flexibility in terms of tail sensitivity than is offered by mean-variance. This class lets us derive high-probability learning guarantees without explicit gradient clipping, and empirical tests using both simulated and real data illustrate a high degree of control over key properties of the test loss distribution of gradient-based learners. Matthew J. Holland |
AISTATS | 1 |
| 2023 | A Survey of Learning Criteria Going Beyond the Usual RiskabstractVirtually all machine learning tasks are characterized using some form of loss function, and “good performance” is typically stated in terms of a sufficiently small average loss, taken over the random draw of test data. While optimizing for performance on average is intuitive, convenient to analyze in theory, and easy to implement in practice, such a choice brings about trade-offs. In this work, we survey and introduce a wide variety of non-traditional criteria used to design and evaluate machine learning algorithms, place the classical paradigm within the proper historical context, and propose a view of learning problems which emphasizes the question of “what makes for a desirable loss distribution?” in place of tacit use of the expected loss. Matthew J. Holland, Kazuki Tanabe |
J. Artif. Intell. Res. | 1 |
| 2022 | Anytime Guarantees under Heavy-Tailed Data
Matthew J. Holland |
AAAI | 1 |
| 2022 | Spectral risk-based learning using unbounded lossesabstractIn this work, we consider the setting of learning problems under a wide class of spectral risk (or "L-risk") functions, where a Lipschitz-continuous spectral density is used to flexibly assign weight to extreme loss values. We obtain excess risk guarantees for a derivative-free learning procedure under unbounded heavy-tailed loss distributions, and propose a computationally efficient implementation which empirically outperforms traditional risk minimizers in terms of balancing spectral risk and misclassification error. Matthew J. Holland, El Mehdi Haress |
AISTATS | 1 |
| 2022 | Learning with risks based on M-locationabstractAbstract In this work, we study a new class of risks defined in terms of the location and deviation of the loss distribution, generalizing far beyond classical mean-variance risk functions. The class is easily implemented as a wrapper around any smooth loss, it admits finite-sample stationarity guarantees for stochastic gradient methods, it is straightforward to interpret and adjust, with close links to M-estimators of the loss location, and has a salient effect on the test loss distribution, giving us control over symmetry and deviations that are not possible under naive ERM. Matthew J. Holland |
Mach. Learn. | 1 |
| 2021 | Scaling-Up Robust Gradient Descent TechniquesabstractWe study a scalable alternative to robust gradient descent (RGD) techniques that can be used when losses and/or gradients can be heavy-tailed, though this will be unknown to the learner. The core technique is simple: instead of trying to robustly aggregate gradients at each step, which is costly and leads to sub-optimal dimension dependence in risk bounds, we choose a candidate which does not diverge too far from the majority of cheap stochastic sub-processes run over partitioned data. This lets us retain the formal strength of RGD methods at a fraction of the cost. Matthew J. Holland |
AAAI | 1 |
| 2021 | Learning with risk-averse feedback under potentially heavy tailsabstractWe study learning algorithms that seek to minimize the conditional value-at-risk (CVaR), when all the learner knows is that the losses (and gradients) incurred may be heavy-tailed. We begin by studying a general-purpose estimator of CVaR for potentially heavy-tailed random variables, which is easy to implement in practice, and requires nothing more than finite variance and a distribution function that does not change too fast or slow around just the quantile of interest. With this estimator in hand, we then derive a new learning algorithm which robustly chooses among candidates produced by stochastic gradient-driven sub-processes, obtain excess CVaR bounds, and finally complement the theory with a regression application. Matthew J. Holland, El Mehdi Haress |
AISTATS | 1 |
| 2019 | Robust descent using smoothed multiplicative noiseabstractIn this work, we propose a novel robust gradient descent procedure which makes use of a smoothed multiplicative noise applied directly to observations before constructing a sum of soft-truncated gradient coordinates. We show that the procedure has competitive theoretical guarantees, with the major advantage of a simple implementation that does not require an iterative sub-routine for robustification. Empirical tests reinforce the theory, showing more efficient generalization over a much wider class of data distributions. Matthew J. Holland |
AISTATS | 1 |
| 2019 | Classification using margin pursuitabstractIn this work, we study a new approach to optimizing the margin distribution realized by binary classifiers, in which the learner searches the hypothesis space in such a way that a pre-set margin level ends up being a distribution-robust estimator of the margin location. This procedure is easily implemented using gradient descent, and admits finite-sample bounds on the excess risk under unbounded inputs, yielding competitive rates under mild assumptions. Empirical tests on real-world benchmark data reinforce the basic principles highlighted by the theory. Matthew J. Holland |
AISTATS | 1 |
| 2019 | Better generalization with less data using robust gradient descentabstractFor learning tasks where the data (or losses) may be heavy-tailed, algorithms based on empirical risk minimization may require a substantial number of observations in order to perform well off-sample. In pursuit of stronger performance under weaker assumptions, we propose a technique which uses a cheap and robust iterative estimate of the risk gradient, which can be easily fed into any steepest descent procedure. Finite-sample risk bounds are provided under weak moment assumptions on the loss gradient. The algorithm is simple to implement, and empirical tests using simulations and real-world data illustrate that more efficient and reliable learning is possible without prior knowledge of the loss tails. Matthew J. Holland, Kazushi Ikeda |
ICML | 1 |
| 2019 | PAC-Bayes under potentially heavy tailsabstractWe derive PAC-Bayesian learning guarantees for heavy-tailed losses, and obtain a novel optimal Gibbs posterior which enjoys finite-sample excess risk bounds at logarithmic confidence. Our core technique itself makes use of PAC-Bayesian inequalities in order to derive a robust risk estimator, which by design is easy to compute. In particular, only assuming that the first three moments of the loss distribution are bounded, the learning algorithm derived from this estimator achieves nearly sub-Gaussian statistical error, up to the quality of the prior. Matthew J. Holland |
NeurIPS | 1 |
| 2019 | Efficient learning with robust gradient descent
Matthew J. Holland, Kazushi Ikeda |
Mach. Learn. | 1 |
| 2017 | Robust regression using biased objectives
Matthew J. Holland, Kazushi Ikeda |
Mach. Learn. | 1 |
| 2015 | Location robust estimation of predictive Weibull parameters in short-term wind speed forecastingabstractFrom turbine control systems at wind farms to extreme weather early-warning systems, short-term probabilistic wind speed forecasts are seeing widespread use in industrial applications. Successful modern forecast methods, often Weibull-based, have been shown to be extremely sensitive to even minor changes in location. We contend that this lack of robustness stems not from model selection, but rather the parameter estimation methods used, and propose a new proper scoring rule to be dynamically minimized. Tested on a weather array spanning the islands of Japan, we verify both superior short-term forecasting performance and model fit of the proposed method over all standard references, and empirically confirm the desired location robustness. Matthew J. Holland, Kazushi Ikeda |
ICASSP | 1 |
| 2014 | Forecasting in wind energy applications with site-adaptive Weibull estimationabstractFrom optimal supply decisions to anticipatory control systems, wind-based energy applications rely heavily upon accurate, local, short-term forecasts of future wind speed. Recent studies have shown continuous ranked probability score (CRPS) minimizing models with Gaussian assumptions to be effective for well-researched sites where those assumptions are appropriate. We consider the more general case where Gaussianity is not assumed and access to historical data may be constrained. Deriving a CRPS expression for a minimum Extreme Value distribution, we use it to propose a site-adaptive Weibull-based CRPS-minimizing model, which is tested and shown to perform better than both deterministic and probabilistic reference models on a ground-based array of weather observation sites in northern Japan. Matthew J. Holland, Kazushi Ikeda |
ICASSP | 1 |