Ulrich Paquet

dblp:24/3808 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Probabilistic and Bayesian machine learning · 34% Graph learning · 12% Trustworthy machine learning · 11%
Databases, data mining, and information retrieval
6 papers
Recommender systems · 82% Information retrieval · 16% Machine learning and data management · 2%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 100%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.612022
Role of Human-AI Interaction in Selective Prediction · AAAI 2022
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.422016
Sequential Neural Models with Stochastic Layers · NIPS 2016
Perturbative corrections for approximate inference in Gaussian latent variable models · J. Mach. Learn. Res. 2013
Recommender systems
collaborative filtering
0.422016
Beyond Collaborative Filtering: The List Recommendation Problem · WWW 2016
One-class collaborative filtering with random graphs · WWW 2013
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.332013
Perturbative corrections for approximate inference in Gaussian latent variable models · J. Mach. Learn. Res. 2013
Perturbation Corrections in Approximate Inference: Mixture Modelling Applications · J. Mach. Learn. Res. 2009
Improving on Expectation Propagation · NIPS 2008
Machine learning › Graph learning
graph neural network
0.312018
Recurrent Relational Networks · NeurIPS 2018
Machine learning › Graph learning › graph neural network
relational network
0.312018
Recurrent Relational Networks · NeurIPS 2018
Knowledge, reasoning and agents › Knowledge representation and reasoning
relational reasoning
0.312018
Recurrent Relational Networks · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
bayesian active learning
0.312017
Knowing What to Ask: A Bayesian Active Learning Approach to the Surveying Problem · AAAI 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
determinantal point process
0.312017
Low-Rank Factorization of Determinantal Point Processes · AAAI 2017
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.312017
A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning · NIPS 2017
Machine learning › Generative modeling
variational autoencoder
0.312017
A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning · NIPS 2017
Recommender systems › music recommendation
playlist generation
0.312017
Groove Radio: A Bayesian Hierarchical Model for Personalized Playlist Generation · WSDM 2017
Recommender systems › e-commerce recommendation
product recommendation
0.312017
Low-Rank Factorization of Determinantal Point Processes · AAAI 2017
Machine learning › Reinforcement learning
policy learning
0.212016
Collective Noise Contrastive Estimation for Policy Transfer Learning · AAAI 2016
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy transfer
0.212016
Collective Noise Contrastive Estimation for Policy Transfer Learning · AAAI 2016
Machine learning › Deep learning architectures and training
state space model
0.212016
Sequential Neural Models with Stochastic Layers · NIPS 2016
Machine learning › Deep learning architectures and training › recurrent neural network
stochastic recurrent neural network
0.212016
Sequential Neural Models with Stochastic Layers · NIPS 2016
Recommender systems › collaborative filtering › ranking-based collaborative filtering
listwise collaborative filtering
0.212016
Beyond Collaborative Filtering: The List Recommendation Problem · WWW 2016
Information retrieval › similarity search › nearest neighbor search
maximum inner product search
0.212016
Indexable Probabilistic Matrix Factorization for Maximum Inner Product Search · AAAI 2016
Human-AI interaction › human-in-the-loop
human-in-the-loop evaluation
0.212022
Role of Human-AI Interaction in Selective Prediction · AAAI 2022
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent gaussian model
0.212013
Perturbative corrections for approximate inference in Gaussian latent variable models · J. Mach. Learn. Res. 2013
Recommender systems › collaborative filtering
implicit feedback
0.212013
One-class collaborative filtering with random graphs · WWW 2013
Recommender systems › collaborative filtering
one-class collaborative filtering
0.212013
One-class collaborative filtering with random graphs · WWW 2013
Information retrieval › user interaction
personalization
0.112012
Transparent user models for personalization · KDD 2012
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.112009
Convexity and Bayesian constrained local models · CVPR 2009
Computer vision › Face, body and person analysis › face alignment
constrained local models
0.112009
Convexity and Bayesian constrained local models · CVPR 2009
Computer vision › Face, body and person analysis
face alignment
0.112009
Convexity and Bayesian constrained local models · CVPR 2009
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.112009
Perturbation Corrections in Approximate Inference: Mixture Modelling Applications · J. Mach. Learn. Res. 2009
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling
hierarchical bayesian model
0.112017
Groove Radio: A Bayesian Hierarchical Model for Personalized Playlist Generation · WSDM 2017
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning › state representation learning
latent dynamics model
0.112017
A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning · NIPS 2017

Methods — techniques the papers use, named apart from their topics

selective prediction · 1.1low-rank factorization · 0.6kernel learning · 0.6human-in-the-loop experiments · 0.6human-in-the-loop experiment · 0.6recurrent relational network · 0.3graph representation · 0.3kalman filter · 0.3end-to-end training · 0.3bayesian hierarchical modeling · 0.3bayesian dimensionality reduction · 0.3augmented linear regression · 0.3probabilistic matrix factorization · 0.2inverse propensity scoring · 0.2collaborative filtering · 0.2stochastic gradient descent · 0.2random graph · 0.2mean-field variational inference · 0.2
YearPublicationVenuePosition
2022 Role of Human-AI Interaction in Selective Prediction
abstract
Recent work has shown the potential benefit of selective prediction systems that can learn to defer to a human when the predictions of the AI are unreliable, particularly to improve the reliability of AI systems in high-stakes applications like healthcare or conservation. However, most prior work assumes that human behavior remains unchanged when they solve a prediction task as part of a human-AI team as opposed to by themselves. We show that this is not the case by performing experiments to quantify human-AI interaction in the context of selective prediction. In particular, we study the impact of communicating different types of information to humans about the AI system's decision to defer. Using real-world conservation data and a selective prediction system that improves expected accuracy over that of the human or AI system working individually, we show that this messaging has a significant impact on the accuracy of human judgements. Our results study two components of the messaging strategy: 1) Whether humans are informed about the prediction of the AI system and 2) Whether they are informed about the decision of the selective prediction system to defer. By manipulating these messaging components, we show that it is possible to significantly boost human performance by informing the human of the decision to defer, but not revealing the prediction of the AI. We therefore show that it is vital to consider how the decision to defer is communicated to a human when designing selective prediction systems, and that the composite accuracy of a human-AI team must be carefully evaluated using a human-in-the-loop framework.
Elizabeth Bondi-Kelly, Raphael Koster, Hannah Sheahan, Martin J. Chadwick, Yoram Bachrach, A. Taylan Cemgil, Ulrich Paquet, Krishnamurthy Dvijotham
AAAI7
2018 Recurrent Relational Networks
abstract
This paper is concerned with learning to solve tasks that require a chain of interde- pendent steps of relational inference, like answering complex questions about the relationships between objects, or solving puzzles where the smaller elements of a solution mutually constrain each other. We introduce the recurrent relational net- work, a general purpose module that operates on a graph representation of objects. As a generalization of Santoro et al. [2017]’s relational network, it can augment any neural network model with the capacity to do many-step relational reasoning. We achieve state of the art results on the bAbI textual question-answering dataset with the recurrent relational network, consistently solving 20/20 tasks. As bAbI is not particularly challenging from a relational reasoning point of view, we introduce Pretty-CLEVR, a new diagnostic dataset for relational reasoning. In the Pretty- CLEVR set-up, we can vary the question to control for the number of relational reasoning steps that are required to obtain the answer. Using Pretty-CLEVR, we probe the limitations of multi-layer perceptrons, relational and recurrent relational networks. Finally, we show how recurrent relational networks can learn to solve Sudoku puzzles from supervised training data, a challenging task requiring upwards of 64 steps of relational reasoning. We achieve state-of-the-art results amongst comparable methods by solving 96.6% of the hardest Sudoku puzzles.
Rasmus Berg Palm, Ulrich Paquet, Ole Winther
NeurIPS2
2017 Low-Rank Factorization of Determinantal Point Processes
abstract
Determinantal point processes (DPPs) have garnered attention as an elegant probabilistic model of set diversity. They are useful for a number of subset selection tasks, including product recommendation. DPPs are parametrized by a positive semi-definite kernel matrix. In this work we present a new method for learning the DPP kernel from observed data using a low-rank factorization of this kernel. We show that this low-rank factorization enables a learning algorithm that is nearly an order of magnitude faster than previous approaches, while also providing for a method for computing product recommendation predictions that is far faster (up to 20x faster or more for large item catalogs) than previous techniques that involve a full-rank DPP kernel. Furthermore, we show that our method provides equivalent or sometimes better test log-likelihood than prior full-rank DPP approaches.
Mike Gartrell, Ulrich Paquet, Noam Koenigstein
AAAI2
2017 Knowing What to Ask: A Bayesian Active Learning Approach to the Surveying Problem
abstract
We examine the surveying problem, where we attempt to predict how a target user is likely to respond to questions by iteratively querying that user, collaboratively based on the responses of a sample set of users. We focus on an active learning approach, where the next question we select to ask the user depends on their responses to the previous questions. We propose a method for solving the problem based on a Bayesian dimensionality reduction technique. We empirically evaluate our method, contrasting it to benchmark approaches based on augmented linear regression, and show that it achieves much better predictive performance, and is much more robust when there is missing data.
Yoad Lewenberg, Yoram Bachrach, Ulrich Paquet, Jeffrey S. Rosenschein
AAAI3
2017 A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning
abstract
This paper takes a step towards temporal reasoning in a dynamically changing video, not in the pixel space that constitutes its frames, but in a latent space that describes the non-linear dynamics of the objects in its world. We introduce the Kalman variational auto-encoder, a framework for unsupervised learning of sequential data that disentangles two latent representations: an object's representation, coming from a recognition model, and a latent state describing its dynamics. As a result, the evolution of the world can be imagined and missing data imputed, both without the need to generate high dimensional frames at each time step. The model is trained end-to-end on videos of a variety of simulated physical systems, and outperforms competing methods in generative and missing data imputation tasks.
Marco Fraccaro, Simon Kamronn, Ulrich Paquet, Ole Winther
NIPS3
2017 Groove Radio: A Bayesian Hierarchical Model for Personalized Playlist Generation
abstract
This paper describes an algorithm designed for Microsoft's Groove music service, which serves millions of users world wide. We consider the problem of automatically generating personalized music playlists based on queries containing a ``seed'' artist and the listener's user ID. Playlist generation may be informed by a number of information sources including: user specific listening patterns, domain knowledge encoded in a taxonomy, acoustic features of audio tracks, and overall popularity of tracks and artists. The importance assigned to each of these information sources may vary depending on the specific combination of user and seed~artist.
Shay Ben-Elazar, Gal Lavee, Noam Koenigstein, Oren Barkan, Hilik Berezin, Ulrich Paquet, Tal Zaccai
WSDM6
2016 Indexable Probabilistic Matrix Factorization for Maximum Inner Product Search
Marco Fraccaro, Ulrich Paquet, Ole Winther
AAAI2
2016 Collective Noise Contrastive Estimation for Policy Transfer Learning
abstract
We address the problem of learning behaviour policies to optimise online metrics from heterogeneous usage data. While online metrics, e.g., click-through rate, can be optimised effectively using exploration data, such data is costly to collect in practice, as it temporarily degrades the user experience. Leveraging related data sources to improve online performance would be extremely valuable, but is not possible using current approaches. We formulate this task as a policy transfer learning problem, and propose a first solution, called collective noise contrastive estimation (collective NCE). NCE is an efficient solution to approximating the gradient of a log-softmax objective. Our approach jointly optimises embeddings of heterogeneous data to transfer knowledge from the source domain to the target domain. We demonstrate the effectiveness of our approach by learning an effective policy for an online radio station jointly from user-generated playlists, and usage data collected in an exploration bucket.
Weinan Zhang 0001, Ulrich Paquet, Katja Hofmann
AAAI2
2016 Sequential Neural Models with Stochastic Layers
abstract
How can we efficiently propagate uncertainty in a latent state representation with recurrent neural networks? This paper introduces stochastic recurrent neural networks which glue a deterministic recurrent neural network and a state space model together to form a stochastic and sequential neural generative model. The clear separation of deterministic and stochastic layers allows a structured variational inference network to track the factorization of the model’s posterior distribution. By retaining both the nonlinear recursive structure of a recurrent neural network and averaging over the uncertainty in a latent path, like a state space model, we improve the state of the art results on the Blizzard and TIMIT speech modeling data sets by a large margin, while achieving comparable performances to competing methods on polyphonic music modeling.
Marco Fraccaro, Søren Kaae Sønderby, Ulrich Paquet, Ole Winther
NIPS3
2016 Bayesian Low-Rank Determinantal Point Processes
abstract
Determinantal point processes (DPPs) are an emerging model for encoding probabilities over subsets, such as shopping baskets, selected from a ground set, such as an item catalog. They have recently proved to be appealing models for a number of machine learning tasks, including product recommendation. DPPs are parametrized by a positive semi-definite kernel matrix. Prior work has shown that using a low-rank factorization of this kernel provides scalability improvements that open the door to training on large-scale datasets and computing online recommendations, both of which are infeasible with standard DPP models that use a full-rank kernel. A low-rank DPP model can be trained using an optimization-based method, such as stochastic gradient ascent, to find a point estimate of the kernel parameters, which can be performed efficiently on large-scale datasets. However, this approach requires careful tuning of regularization parameters to prevent overfitting and provide good predictive performance, which can be computationally expensive. In this paper we present a Bayesian method for learning a low-rank factorization of this kernel, which provides automatic control of regularization. We show that our Bayesian low-rank DPP model can be trained efficiently using stochastic gradient Hamiltonian Monte Carlo (SGHMC). Our Bayesian model generally provides better predictive performance on several real-world product recommendation datasets than optimization-based low-rank DPP models trained using stochastic gradient ascent, and better performance than several state-of-the art recommendation methods in many cases.
Mike Gartrell, Ulrich Paquet, Noam Koenigstein
RecSys2
2016 Beyond Collaborative Filtering: The List Recommendation Problem
abstract
Most Collaborative Filtering (CF) algorithms are optimized using a dataset of isolated user-item tuples. However, in commercial applications recommended items are usually served as an ordered list of several items and not as isolated items. In this setting, inter-item interactions have an effect on the list's Click-Through Rate (CTR) that is unaccounted for using traditional CF approaches. Most CF approaches also ignore additional important factors like click propensity variation, item fatigue, etc. In this work, we introduce the list recommendation problem. We present useful insights gleaned from user behavior and consumption patterns from a large scale real world recommender system. We then propose a novel two-layered framework that builds upon existing CF algorithms to optimize a list's click probability. Our approach accounts for inter-item interactions as well as additional information such as item fatigue, trendiness patterns, contextual information etc. Finally, we evaluate our approach using a novel adaptation of Inverse Propensity Scoring (IPS) which facilitates off-policy estimation of our method's CTR and showcases its effectiveness in real-world settings.
Oren Sar Shalom, Noam Koenigstein, Ulrich Paquet, Hastagiri P. Vanchinathan
WWW3
2014 Speeding up the Xbox recommender system using a euclidean transformation for inner-product spaces
abstract
A prominent approach in collaborative filtering based recommender systems is using dimensionality reduction (matrix factorization) techniques to map users and items into low-dimensional vectors. In such systems, a higher inner product between a user vector and an item vector indicates that the item better suits the user's preference. Traditionally, retrieving the most suitable items is done by scoring and sorting all items. Real world online recommender systems must adhere to strict response-time constraints, so when the number of items is large, scoring all items is intractable.
Yoram Bachrach, Yehuda Finkelstein, Ran Gilad-Bachrach, Liran Katzir 0001, Noam Koenigstein, Nir Nice, Ulrich Paquet
RecSys7
2013 Xbox movies recommendations: variational bayes matrix factorization with embedded feature selection
abstract
We present a matrix factorization model inspired by challenges we encountered while working on the Xbox movies recommendation system. The item catalog in a recommender system is typically equipped with meta-data features in the form of labels. However, only part of these features are informative or useful with regard to collaborative filtering. By incorporating a novel sparsity prior on feature parameters, the model automatically discerns and utilizes informative features while simultaneously pruning non-informative features.
Noam Koenigstein, Ulrich Paquet
RecSys2
2013 One-class collaborative filtering with random graphs
abstract
The bane of one-class collaborative filtering is interpreting and modelling the latent signal from the missing class. In this paper we present a novel Bayesian generative model for implicit collaborative filtering. It forms a core component of the Xbox Live architecture, and unlike previous approaches, delineates the odds of a user disliking an item from simply being unaware of it. The latent signal is treated as an unobserved random graph connecting users with items they might have encountered. We demonstrate how large-scale distributed learning can be achieved through a combination of stochastic gradient descent and mean field variational inference over random graph samples. A fine-grained comparison is done against a state of the art baseline on real world data.
Ulrich Paquet, Noam Koenigstein
WWW1
2013 Perturbative corrections for approximate inference in Gaussian latent variable models
Manfred Opper, Ulrich Paquet, Ole Winther
J. Mach. Learn. Res.2
2012 Transparent user models for personalization
abstract
Personalization is a ubiquitous phenomenon in our daily online experience. While such technology is critical for helping us combat the overload of information we face, in many cases, we may not even realize that our results are being tailored to our personal tastes and preferences. Worse yet, when such a system makes a mistake, we have little recourse to correct it.
Khalid El-Arini, Ulrich Paquet, Ralf Herbrich, Jurgen Van Gael, Blaise Agüera y Arcas
KDD2
2012 The Xbox recommender system
abstract
A recent addition to Microsoft's Xbox Live Marketplace is a recommender system which allows users to explore both movies and games in a personalized context. The system largely relies on implicit feedback, and runs on a large scale, serving tens of millions of daily users. We describe the system design, and review the core recommendation algorithm.
Noam Koenigstein, Nir Nice, Ulrich Paquet, Nir Schleyen
RecSys3
2012 Collaborative learning of preference rankings
abstract
We propose a model for learning user preference rankings for the purpose of making product recommendations. The model allows us to learn from pairwise preference statements or from (incomplete) rankings over more than two items. We present two algorithms for performing inference in this model, both with excellent scaling in the number of users and items. The superior predictive performance of the new method is demonstrated on the well-known sushi preference data set. In addition, we show how the model can be used effectively in an active learning setting where we select only a small number of informative items for learning.
Tim Salimans, Ulrich Paquet, Thore Graepel
RecSys2
2009 Convexity and Bayesian constrained local models
abstract
The accurate localization of facial features plays a fundamental role in any face recognition pipeline. Constrained local models (CLM) provide an effective approach to localization by coupling ensembles of local patch detectors for non-rigid object alignment. A recent improvement has been made by using generic convex quadratic fitting (CQF), which elegantly addresses the CLM warp update by enforcing convexity of the patch response surfaces. In this paper, CQF is generalized to a Bayesian inference problem, in which it appears as a particular maximum likelihood solution. The Bayesian viewpoint holds many advantages: for example, the task of feature localization can explicitly build on previous face detection stages, and multiple sets of patch responses can be seamlessly incorporated. A second contribution of the paper is an analytic solution to finding convex approximations to patch response surfaces, which removes CQF's reliance on a numeric optimizer. Improvements in feature localization performance are illustrated on the Labeled Faces in the Wild and BioID data sets.
Ulrich Paquet
CVPR1
2009 Perturbation Corrections in Approximate Inference: Mixture Modelling Applications
Ulrich Paquet, Ole Winther, Manfred Opper
J. Mach. Learn. Res.1
2008 Improving on Expectation Propagation
abstract
We develop as series of corrections to Expectation Propagation (EP), which is one of the most popular methods for approximate probabilistic inference. These corrections can lead to improvements of the inference approximation or serve as a sanity check, indicating when EP yields unrealiable results.
Manfred Opper, Ulrich Paquet, Ole Winther
NIPS2
2007 Particle Swarms for Linearly Constrained Optimisation
Ulrich Paquet, Andries P. Engelbrecht
Fundam. Informaticae1
2005 On the Explicit Use of Example Weights in the Construction of Classifiers
Andrew Naish, Sean B. Holden, Ulrich Paquet
ICANN (2)3
2005 Bayesian Hierarchical Ordinal Regression
Ulrich Paquet, Sean B. Holden, Andrew Naish
ICANN (2)1
2003 A new particle swarm optimiser for linearly constrained optimisation
abstract
A new PSO algorithm, the linear PSO (LPSO), is developed to optimise functions constrained by linear constraints of the form Ax = b. A crucial property of the LPSO is that the possible movement of particles through vector spaces is guaranteed by the velocity and position update equations. This property makes the LPSO ideal in optimising linearly constrained problems. The LPSO is extended to the converging linear PSO, which is guaranteed to always find at least a local minimum.
Ulrich Paquet, Andries P. Engelbrecht
IEEE Congress on Evolutionary Computation1
2003 Training support vector machines with particle swarms
abstract
Training a support vector machine requires solving a constrained quadratic programming problem. Linear particle swarm optimization is intuitive and simple to implement, and is presented as an alternative to current numeric SVM training methods. Performance of the new algorithm is demonstrated on the MNIST character recognition dataset.
Ulrich Paquet, Andries P. Engelbrecht
IJCNN1