Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Christian J. Walder

dblp:w/CJWalder · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Probabilistic and Bayesian machine learning · 69% Learning theory · 14% Trustworthy machine learning · 11%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Theoretical computer science
3 papers
Mathematical optimization · 64% Algorithms and data structures · 36%

Topics — the 28 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
1.232020
Quantile Propagation for Wasserstein-Approximate Gaussian Processes · NeurIPS 2020
All your loss are belong to Bayes · NeurIPS 2020
Fast Bayesian Intensity Estimation for the Permanental Process · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process › temporal point process
hawkes process
0.822020
Variational Inference for Sparse Gaussian Process Modulated Hawkes Process · AAAI 2020
Efficient Non-parametric Bayesian Hawkes Processes · IJCAI 2019
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
0.822020
Variational Inference for Sparse Gaussian Process Modulated Hawkes Process · AAAI 2020
Efficient Non-parametric Bayesian Hawkes Processes · IJCAI 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.722020
Quantile Propagation for Wasserstein-Approximate Gaussian Processes · NeurIPS 2020
Fast Bayesian Intensity Estimation for the Permanental Process · ICML 2017
Machine learning › Learning theory › loss function
proper scoring rules
0.712023
LegendreTron: Uprising Proper Multiclass Loss Learning · ICML 2023
Machine learning › Deep learning architectures and training
sequence modeling
0.722018
Neural Dynamic Programming for Musical Self Similarity · ICML 2018
Self-Bounded Prediction Suffix Tree via Approximate String Matching · ICML 2018
Recommender systems › diversified recommendation
determinantal point process
0.612022
Determinantal Point Process Likelihoods for Sequential Recommendation · SIGIR 2022
Recommender systems
loss function design
0.612022
Determinantal Point Process Likelihoods for Sequential Recommendation · SIGIR 2022
Recommender systems
sequential recommendation
0.612022
Determinantal Point Process Likelihoods for Sequential Recommendation · SIGIR 2022
Machine learning › Probabilistic and Bayesian machine learning
bayesian decision theory
0.412020
All your loss are belong to Bayes · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.412020
Variational Inference for Sparse Gaussian Process Modulated Hawkes Process · AAAI 2020
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
expectation propagation
0.412020
Quantile Propagation for Wasserstein-Approximate Gaussian Processes · NeurIPS 2020
Machine learning › Learning theory › loss function
proper loss
0.412020
All your loss are belong to Bayes · NeurIPS 2020
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412019
Monge blunts Bayes: Hardness Results for Adversarial Training · ICML 2019
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.412019
Monge blunts Bayes: Hardness Results for Adversarial Training · ICML 2019
Machine learning › Trustworthy machine learning
robustness
0.412019
Monge blunts Bayes: Hardness Results for Adversarial Training · ICML 2019
Machine learning › Learning theory
statistical learning theory
0.412019
Monge blunts Bayes: Hardness Results for Adversarial Training · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
prediction suffix trees
0.312018
Self-Bounded Prediction Suffix Tree via Approximate String Matching · ICML 2018
Audio and music processing › music technology › computer music
symbolic music modeling
0.312018
Neural Dynamic Programming for Musical Self Similarity · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.312017
Fast Bayesian Intensity Estimation for the Permanental Process · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
cox process
0.312017
Fast Bayesian Intensity Estimation for the Permanental Process · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
laplace approximation
0.312017
Fast Bayesian Intensity Estimation for the Permanental Process · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
permanental process
0.312017
Fast Bayesian Intensity Estimation for the Permanental Process · ICML 2017
Mathematical optimization
variational inference
0.112020
Variational Inference for Sparse Gaussian Process Modulated Hawkes Process · AAAI 2020
Mathematical optimization
continuous optimization
0.112019
Monge blunts Bayes: Hardness Results for Adversarial Training · ICML 2019
Mathematical optimization
optimal transport
0.112019
Monge blunts Bayes: Hardness Results for Adversarial Training · ICML 2019
Algorithms and data structures › sequence algorithms › string algorithms
edit distance
0.112018
Neural Dynamic Programming for Musical Self Similarity · ICML 2018
Algorithms and data structures › sequence algorithms
string algorithms
0.112018
Neural Dynamic Programming for Musical Self Similarity · ICML 2018

Methods — techniques the papers use, named apart from their topics

expectation-maximization · 1.2variational inference · 0.9sparse gaussian process · 0.9neural network · 0.7monotonicity of gradients · 0.7legendre transformation · 0.7dynamic programming · 0.7convex functions · 0.7likelihood optimization · 0.6determinantal point process · 0.6wasserstein distance · 0.4source function · 0.4quantile propagation · 0.4information geometry · 0.4optimal transport · 0.4lipschitz classifiers · 0.4integral probability metrics · 0.4
YearPublicationVenuePosition
2023 LegendreTron: Uprising Proper Multiclass Loss Learning
abstract
Loss functions serve as the foundation of supervised learning and are often chosen prior to model development. To avoid potentially ad hoc choices of losses, statistical decision theory describes a desirable property for losses known as *properness*, which asserts that Bayes' rule is optimal. Recent works have sought to *learn losses* and models jointly. Existing methods do this by fitting an inverse canonical link function which monotonically maps $\mathbb{R}$ to $[0,1]$ to estimate probabilities for binary problems. In this paper, we extend monotonicity to maps between $\mathbb{R}^{C-1}$ and the projected probability simplex $\tilde{\Delta}^{C-1}$ by using monotonicity of gradients of convex functions. We present LegendreTron as a novel and practical method that jointly learns *proper canonical losses* and probabilities for multiclass problems. Tested on a benchmark of domains with up to 1,000 classes, our experimental results show that our method consistently outperforms the natural multiclass baseline under a $t$-test at 99% significance on all datasets with greater than $10$ classes.
Kevin H. Lam, Christian J. Walder, Spiridon I. Penev, Richard Nock
ICML2
2022 Determinantal Point Process Likelihoods for Sequential Recommendation
abstract
Sequential recommendation is a popular task in academic research and close to real-world application scenarios, where the goal is to predict the next action(s) of the user based on his/her previous sequence of actions. In the training process of recommender systems, the loss function plays an essential role in guiding the optimization of recommendation models to generate accurate suggestions for users. However, most existing sequential recommendation tech- niques focus on designing algorithms or neural network architectures, and few efforts have been made to tailor loss functions that fit naturally into the practical application scenario of sequential recommender systems.
Yuli Liu, Christian J. Walder, Lexing Xie
SIGIR2
2020 Variational Inference for Sparse Gaussian Process Modulated Hawkes Process
abstract
The Hawkes process (HP) has been widely applied to modeling self-exciting events including neuron spikes, earthquakes and tweets. To avoid designing parametric triggering kernel and to be able to quantify the prediction confidence, the non-parametric Bayesian HP has been proposed. However, the inference of such models suffers from unscalability or slow convergence. In this paper, we aim to solve both problems. Specifically, first, we propose a new non-parametric Bayesian HP in which the triggering kernel is modeled as a squared sparse Gaussian process. Then, we propose a novel variational inference schema for model optimization. We employ the branching structure of the HP so that maximization of evidence lower bound (ELBO) is tractable by the expectation-maximization algorithm. We propose a tighter ELBO which improves the fitting performance. Further, we accelerate the novel variational inference schema to linear time complexity by leveraging the stationarity of the triggering kernel. Different from prior acceleration methods, ours enjoys higher efficiency. Finally, we exploit synthetic data and two large social media datasets to evaluate our method. We show that our approach outperforms state-of-the-art non-parametric frequentist and Bayesian methods. We validate the efficiency of our accelerated variational inference schema and practical utility of our tighter ELBO for model selection. We observe that the tighter ELBO exceeds the common one in model selection.
Christian J. Walder, Marian-Andrei Rizoiu
AAAI2
2020 All your loss are belong to Bayes
abstract
Loss functions are a cornerstone of machine learning and the starting point of most algorithms. Statistics and Bayesian decision theory have contributed, via properness, to elicit over the past decades a wide set of admissible losses in supervised learning, to which most popular choices belong (logistic, square, Matsushita, etc.). Rather than making a potentially biased ad hoc choice of the loss, there has recently been a boost in efforts to fit the loss to the domain at hand while training the model itself. The key approaches fit a canonical link, a function which monotonically relates the closed unit interval to R and can provide a proper loss via integration. In this paper, we rely on a broader view of proper composite losses and a recent construct from information geometry, source functions, whose fitting alleviates constraints faced by canonical links. We introduce a trick on squared Gaussian Processes to obtain a random process whose paths are compliant source functions with many desirable properties in the context of link estimation. Experimental results demonstrate substantial improvements over the state of the art.
Christian J. Walder, Richard Nock
NeurIPS1
2020 Quantile Propagation for Wasserstein-Approximate Gaussian Processes
abstract
Approximate inference techniques are the cornerstone of probabilistic methods based on Gaussian process priors. Despite this, most work approximately optimizes standard divergence measures such as the Kullback-Leibler (KL) divergence, which lack the basic desiderata for the task at hand, while chiefly offering merely technical convenience. We develop a new approximate inference method for Gaussian process models which overcomes the technical challenges arising from abandoning these convenient divergences. Our method---dubbed Quantile Propagation (QP)---is similar to expectation propagation (EP) but minimizes the $L_2$ Wasserstein distance (WD) instead of the KL divergence. The WD exhibits all the required properties of a distance metric, while respecting the geometry of the underlying sample space. We show that QP matches quantile functions rather than moments as in EP and has the same mean update but a smaller variance update than EP, thereby alleviating EP's tendency to over-estimate posterior variances. Crucially, despite the significant complexity of dealing with the WD, QP has the same favorable locality property as EP, and thereby admits an efficient algorithm. Experiments on classification and Poisson regression show that QP outperforms both EP and variational Bayes.
Christian J. Walder, Edwin V. Bonilla, Marian-Andrei Rizoiu, Lexing Xie
NeurIPS2
2019 Monge blunts Bayes: Hardness Results for Adversarial Training
abstract
The last few years have seen a staggering number of empirical studies of the robustness of neural networks in a model of adversarial perturbations of their inputs. Most rely on an adversary which carries out local modifications within prescribed balls. None however has so far questioned the broader picture: how to frame a resource-bounded adversary so that it can be severely detrimental to learning, a non-trivial problem which entails at a minimum the choice of loss and classifiers. We suggest a formal answer for losses that satisfy the minimal statistical requirement of being proper. We pin down a simple sufficient property for any given class of adversaries to be detrimental to learning, involving a central measure of “harmfulness” which generalizes the well-known class of integral probability metrics. A key feature of our result is that it holds for all proper losses, and for a popular subset of these, the optimisation of this central measure appears to be independent of the loss. When classifiers are Lipschitz – a now popular approach in adversarial training –, this optimisation resorts to optimal transport to make a low-budget compression of class marginals. Toy experiments reveal a finding recently separately observed: training against a sufficiently budgeted adversary of this kind improves generalization.
Zac Cranko, Aditya Krishna Menon, Richard Nock, Cheng Soon Ong, Christian J. Walder
ICML6
2019 Efficient Non-parametric Bayesian Hawkes Processes
abstract
In this paper, we develop an efficient non-parametric Bayesian estimation of the kernel function of Hawkes processes. The non-parametric Bayesian approach is important because it provides flexible Hawkes kernels and quantifies their uncertainty. Our method is based on the cluster representation of Hawkes processes. Utilizing the stationarity of the Hawkes process, we efficiently sample random branching structures and thus, we split the Hawkes process into clusters of Poisson processes. We derive two algorithms --- a block Gibbs sampler and a maximum a posteriori estimator based on expectation maximization --- and we show that our methods have a linear time complexity, both theoretically and empirically. On synthetic data, we show our methods to be able to infer flexible Hawkes triggering kernels. On two large-scale Twitter diffusion datasets, we show that our methods outperform the current state-of-the-art in goodness-of-fit and that the time complexity is linear in the size of the dataset. We also observe that on diffusions related to online videos, the learned kernels reflect the perceived longevity for different content types such as music or pets videos.
Christian J. Walder, Marian-Andrei Rizoiu, Lexing Xie
IJCAI2
2018 Self-Bounded Prediction Suffix Tree via Approximate String Matching
abstract
Prediction suffix trees (PST) provide an effective tool for sequence modelling and prediction. Current prediction techniques for PSTs rely on exact matching between the suffix of the current sequence and the previously observed sequence. We present a provably correct algorithm for learning a PST with approximate suffix matching by relaxing the exact matching condition. We then present a self-bounded enhancement of our algorithm where the depth of suffix tree grows automatically in response to the model performance on a training sequence. Through experiments on synthetic datasets as well as three real-world datasets, we show that the approximate matching PST results in better predictive performance than the other variants of PST.
Dongwoo Kim 0002, Christian J. Walder
ICML2
2018 Neural Dynamic Programming for Musical Self Similarity
abstract
We present a neural sequence model designed specifically for symbolic music. The model is based on a learned edit distance mechanism which generalises a classic recursion from computer science, leading to a neural dynamic program. Repeated motifs are detected by learning the transformations between them. We represent the arising computational dependencies using a novel data structure, the edit tree; this perspective suggests natural approximations which afford the scaling up of our otherwise cubic time algorithm. We demonstrate our model on real and synthetic data; in all cases it out-performs a strong stacked long short-term memory benchmark.
Christian J. Walder, Dongwoo Kim 0002
ICML1
2018 Neural Causality Detection for Multi-dimensional Point Processes
Christian J. Walder, Tom Gedeon
ICONIP (4)2
2017 Computer Assisted Composition with Recurrent Neural Networks
abstract
Sequence modeling with neural networks has lead to powerful models of symbolic music data. We address the problem of exploiting these models to reach creative musical goals, by combining with human input. To this end we generalise previous work, which sampled Markovian sequence models under the constraint that the sequence belong to the language of a given finite state machine provided by the human. We consider more expressive non-Markov models, thereby requiring approximate sampling which we provide in the form of an efficient sequential Monte Carlo method. In addition we provide and compare with a beam search strategy for conditional probability maximisation. Our algorithms are capable of convincingly re-harmonising famous musical works. To demonstrate this we provide visualisations, quantitative experiments, a human listening test and audio examples. We find both the sampling and optimisation procedures to be effective, yet complementary in character. For the case of highly permissive constraint sets, we find that sampling is to be preferred due to the overly regular nature of the optimisation based results. The generality of our algorithms permits countless other creative applications.
Christian J. Walder, Dongwoo Kim 0002
ACML1
2017 Fast Bayesian Intensity Estimation for the Permanental Process
abstract
The Cox process is a stochastic process which generalises the Poisson process by letting the underlying intensity function itself be a stochastic process. In this paper we present a fast Bayesian inference scheme for the permanental process, a Cox process under which the square root of the intensity is a Gaussian process. In particular we exploit connections with reproducing kernel Hilbert spaces, to derive efficient approximate Bayesian inference algorithms based on the Laplace approximation to the predictive distribution and marginal likelihood. We obtain a simple algorithm which we apply to toy and real-world problems, obtaining orders of magnitude speed improvements over previous work.
Christian J. Walder, Adrian N. Bishop
ICML1