VLDB 2026 Research / reviewers in the wild / expert
John Guiver
dblp:52/2457
· DBLP profile ↗
12ranked-venue papers
3as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 2 first-authorArtificial intelligence and machine learning · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Data mining · 38% Information retrieval · 37% Recommender systems · 10% | |
| Artificial intelligence
5 papers |
Probabilistic and Bayesian machine learning · 100% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › crowdsourcing
crowdsourced annotation |
0.2 | 1 | 2015 | Language Understanding in the Wild: Combining Crowdsourcing and Machine Learning · WWW 2015 |
Machine learning › Probabilistic and Bayesian machine learning
probabilistic programming |
0.2 | 1 | 2014 | Tabular: a schema-driven probabilistic programming language · POPL 2014 |
Data mining
crowdsourcing |
0.2 | 1 | 2014 | Community-based bayesian aggregation models for crowdsourcing · WWW 2014 |
Data mining › crowdsourcing
label aggregation |
0.2 | 1 | 2014 | Community-based bayesian aggregation models for crowdsourcing · WWW 2014 |
Information retrieval
evaluation |
0.2 | 2 | 2009 | A few good topics: Experiments in topic set reduction for retrieval evaluation · ACM Trans. Inf. Syst. 2009 SoftRank: optimizing non-smooth rank metrics · WSDM 2008 |
Information retrieval › ranking
learning to rank |
0.2 | 2 | 2008 | SoftRank: optimizing non-smooth rank metrics · WSDM 2008 Learning to rank with SoftRank and Gaussian processes · SIGIR 2008 |
Information retrieval › retrieval evaluation
ranking evaluation |
0.2 | 2 | 2008 | SoftRank: optimizing non-smooth rank metrics · WSDM 2008 Learning to rank with SoftRank and Gaussian processes · SIGIR 2008 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian graphical model |
0.1 | 1 | 2012 | How To Grade a Test Without Knowing the Answers - A Bayesian Graphical Model for Adaptive Crowdsourcing and Aptitude Testing · ICML 2012 |
Data mining › text mining › topic modeling
latent dirichlet allocation |
0.1 | 1 | 2010 | Statistical models of music-listening sessions in social media · WWW 2010 |
Recommender systems
music recommendation |
0.1 | 1 | 2010 | Statistical models of music-listening sessions in social media · WWW 2010 |
Data mining › text mining
topic modeling |
0.1 | 1 | 2010 | Statistical models of music-listening sessions in social media · WWW 2010 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.1 | 1 | 2009 | Bayesian inference for Plackett-Luce ranking models · ICML 2009 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
expectation propagation |
0.1 | 1 | 2009 | Bayesian inference for Plackett-Luce ranking models · ICML 2009 |
Recommender systems › choice modeling
plackett-luce model |
0.1 | 1 | 2009 | Bayesian inference for Plackett-Luce ranking models · ICML 2009 |
Information retrieval › ranking
ranking model |
0.1 | 1 | 2009 | Bayesian inference for Plackett-Luce ranking models · ICML 2009 |
Information retrieval › ranking › learning to rank
NDCG optimization |
0.1 | 1 | 2008 | Learning to rank with SoftRank and Gaussian processes · SIGIR 2008 |
Information retrieval
retrieval models |
0.1 | 1 | 2008 | SoftRank: optimizing non-smooth rank metrics · WSDM 2008 |
Web and social media mining › social media analysis
social media text analysis |
0.1 | 1 | 2015 | Language Understanding in the Wild: Combining Crowdsourcing and Machine Learning · WWW 2015 |
Data integration and cleaning
imputation |
0.1 | 1 | 2014 | Tabular: a schema-driven probabilistic programming language · POPL 2014 |
Information retrieval › evaluation
effectiveness metrics |
0.0 | 1 | 2009 | A few good topics: Experiments in topic set reduction for retrieval evaluation · ACM Trans. Inf. Syst. 2009 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.0 | 1 | 2008 | Learning to rank with SoftRank and Gaussian processes · SIGIR 2008 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
sparse gaussian process |
0.0 | 1 | 2008 | Learning to rank with SoftRank and Gaussian processes · SIGIR 2008 |
Methods — techniques the papers use, named apart from their topics
bayesian inference · 0.7graphical model · 0.4probabilistic model expressions · 0.4formal semantics · 0.4confusion matrix · 0.4crowdsourcing · 0.2bayesian modeling · 0.2maximum likelihood · 0.2type systems · 0.2type system · 0.2latent dirichlet allocation · 0.1Power EP · 0.1sparse gaussian process · 0.1softrank · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Time-Sensitive Bayesian Information Aggregation for Crowdsourcing SystemsabstractMany aspects of the design of efficient crowdsourcing processes, such as defining workers bonuses, fair prices and time limits of the tasks, involve knowledge of the likely duration of the task at hand. In this work we introduce a new timesensitive Bayesian aggregation method that simultaneously estimates a tasks duration and obtains reliable aggregations of crowdsourced judgments. Our method, called BCCTime, uses latent variables to represent the uncertainty about the workers completion time, the tasks duration and the workers accuracy. To relate the quality of a judgment to the time a worker spends on a task, our model assumes that each task is completed within a latent time window within which all workers with a propensity to genuinely attempt the labelling task (i.e., no spammers) are expected to submit their judgments. In contrast, workers with a lower propensity to valid labelling, such as spammers, bots or lazy labellers, are assumed to perform tasks considerably faster or slower than the time required by normal workers. Specifically, we use efficient message-passing Bayesian inference to learn approximate posterior probabilities of (i) the confusion matrix of each worker, (ii) the propensity to valid labelling of each worker, (iii) the unbiased duration of each task and (iv) the true label of each task. Using two real- world public datasets for entity linking tasks, we show that BCCTime produces up to 11% more accurate classifications and up to 100% more informative estimates of a tasks duration compared to stateoftheart methods. Matteo Venanzi, John Guiver, Pushmeet Kohli, Nicholas R. Jennings |
J. Artif. Intell. Res. | 2 |
| 2015 | Language Understanding in the Wild: Combining Crowdsourcing and Machine LearningabstractSocial media has led to the democratisation of opinion sharing. A wealth of information about public opinions, current events, and authors' insights into specific topics can be gained by understanding the text written by users. However, there is a wide variation in the language used by different authors in different contexts on the web. This diversity in language makes interpretation an extremely challenging task. Crowdsourcing presents an opportunity to interpret the sentiment, or topic, of free-text. However, the subjectivity and bias of human interpreters raise challenges in inferring the semantics expressed by the text. To overcome this problem, we present a novel Bayesian approach to language understanding that relies on aggregated crowdsourced judgements. Our model encodes the relationships between labels and text features in documents, such as tweets, web articles, and blog posts, accounting for the varying reliability of human labellers. It allows inference of annotations that scales to arbitrarily large pools of documents. Our evaluation using two challenging crowdsourcing datasets shows that by efficiently exploiting language models learnt from aggregated crowdsourced labels, we can provide up to 25% improved classifications when only a small portion, less than 4% of documents has been labelled. Compared to the six state-of-the-art methods, we reduce by up to 67% the number of crowd responses required to achieve comparable accuracy. Our method was a joint winner of the CrowdFlower - CrowdScale 2013 Shared Task challenge at the conference on Human Computation and Crowdsourcing (HCOMP 2013). Edwin Simpson, Matteo Venanzi, Steven Reece, Pushmeet Kohli, John Guiver, Stephen J. Roberts, Nicholas R. Jennings |
WWW | 5 |
| 2014 | Students, Teachers, Exams and MOOCs: Predicting and Optimizing Attainment in Web-Based Education Using a Probabilistic Graphical Model
Bar Shalem, Yoram Bachrach, John Guiver, Christopher M. Bishop |
ECML/PKDD (3) | 3 |
| 2014 | Tabular: a schema-driven probabilistic programming languageabstractWe propose a new kind of probabilistic programming language for machine learning. We write programs simply by annotating existing relational schemas with probabilistic model expressions. We describe a detailed design of our language, Tabular, complete with formal semantics and type system. A rich series of examples illustrates the expressiveness of Tabular. We report an implementation, and show evidence of the succinctness of our notation relative to current best practice. Finally, we describe and verify a transformation of Tabular schemas so as to predict missing values in a concrete database. The ability to query for missing values provides a uniform interface to a wide variety of tasks, including classification, clustering, recommendation, and ranking. Andrew D. Gordon 0001, Thore Graepel, Nicolas Rolland, Claudio V. Russo, Johannes Borgström, John Guiver |
POPL | 6 |
| 2014 | Community-based bayesian aggregation models for crowdsourcingabstractThis paper addresses the problem of extracting accurate labels from crowdsourced datasets, a key challenge in crowdsourcing. Prior work has focused on modeling the reliability of individual workers, for instance, by way of confusion matrices, and using these latent traits to estimate the true labels more accurately. However, this strategy becomes ineffective when there are too few labels per worker to reliably estimate their quality. To mitigate this issue, we propose a novel community-based Bayesian label aggregation model, CommunityBCC, which assumes that crowd workers conform to a few different types, where each type represents a group of workers with similar confusion matrices. We assume that each worker belongs to a certain community, where the worker's confusion matrix is similar to (a perturbation of) the community's confusion matrix. Our model can then learn a set of key latent features: (i) the confusion matrix of each community, (ii) the community membership of each user, and (iii) the aggregated label of each item. We compare the performance of our model against established aggregation methods on a number of large-scale, real-world crowdsourcing datasets. Our experimental results show that our CommunityBCC model consistently outperforms state-of-the-art label aggregation methods, requiring, on average, 50% less data to pass the 90% accuracy mark. Matteo Venanzi, John Guiver, Gabriella Kazai, Pushmeet Kohli, Milad Shokouhi |
WWW | 2 |
| 2012 | How To Grade a Test Without Knowing the Answers - A Bayesian Graphical Model for Adaptive Crowdsourcing and Aptitude Testing
Yoram Bachrach, Thore Graepel, Tom Minka, John Guiver |
ICML | 4 |
| 2012 | De-Layering Social Networks by Shared Tastes of Friendships
Laura Dietz, Ben Gamari, John Guiver, Edward Lloyd Snelson, Ralf Herbrich |
ICWSM | 3 |
| 2010 | Statistical models of music-listening sessions in social mediaabstractUser experience in social media involves rich interactions with the media content and other participants in the community. In order to support such communities, it is important to understand the factors that drive the users' engagement. In this paper we show how to define statistical models of different complexity to describe patterns of song listening in an online music community. First, we adapt the LDA model to capture music taste from listening activities across users and identify both the groups of songs associated with the specific taste and the groups of listeners who share the same taste. Second, we define a graphical model that takes into account listening sessions and captures the listening moods of users in the community. Our session model leads to groups of songs and groups of listeners with similar behavior across listening sessions and enables faster inference when compared to the LDA model. Our experiments with the data from an online media site demonstrate that the session model is better in terms of the perplexity compared to two other models: the LDA-based taste model that does not incorporate cross-session information and a baseline model that does not use latent groupings of songs. Elena Zheleva, John Guiver, Eduarda Mendes Rodrigues, Natasa Milic-Frayling |
WWW | 2 |
| 2009 | Bayesian inference for Plackett-Luce ranking modelsabstractThis paper gives an efficient Bayesian method for inferring the parameters of a Plackett-Luce ranking model. Such models are parameterised distributions over rankings of a finite set of objects, and have typically been studied and applied within the psychometric, sociometric and econometric literature. The inference scheme is an application of Power EP (expectation propagation). The scheme is robust and can be readily applied to large scale data sets. The inference algorithm extends to variations of the basic Plackett-Luce model, including partial rankings. We show a number of advantages of the EP approach over the traditional maximum likelihood method. We apply the method to aggregate rankings of NASCAR racing drivers over the 2002 season, and also to rankings of movie genres. John Guiver, Edward Lloyd Snelson |
ICML | 1 |
| 2009 | A few good topics: Experiments in topic set reduction for retrieval evaluationabstractWe consider the issue of evaluating information retrieval systems on the basis of a limited number of topics. In contrast to statistically-based work on sample sizes, we hypothesize that some topics or topic sets are better than others at predicting true system effectiveness, and that with the right choice of topics, accurate predictions can be obtained from small topics sets. Using a variety of effectiveness metrics and measures of goodness of prediction, a study of a set of TREC and NTCIR results confirms this hypothesis, and provides evidence that the value of a topic set for this purpose does generalize. John Guiver, Stefano Mizzaro, Stephen E. Robertson |
ACM Trans. Inf. Syst. | 1 |
| 2008 | Learning to rank with SoftRank and Gaussian processesabstractIn this paper we address the issue of learning to rank for document retrieval using Thurstonian models based on sparse Gaussian processes. Thurstonian models represent each document for a given query as a probability distribution in a score space; these distributions over scores naturally give rise to distributions over document rankings. However, in general we do not have observed rankings with which to train the model; instead, each document in the training set is judged to have a particular relevance level: for example "Bad", "Fair", "Good", or "Excellent". The performance of the model is then evaluated using information retrieval (IR) metrics such as Normalised Discounted Cumulative Gain (NDCG). Recently Taylor et al. presented a method called SoftRank which allows the direct gradient optimisation of a smoothed version of NDCG using a Thurstonian model. In this approach, document scores are represented by the outputs of a neural network, and score distributions are created artificially by adding random noise to the scores. The SoftRank mechanism is a general one; it can be applied to different IR metrics, and make use of different underlying models. In this paper we extend the SoftRank framework to make use of the score uncertainties which are naturally provided by a Gaussian process (GP), which is a probabilistic non-linear regression model. We further develop the model by using sparse Gaussian process techniques, which give improved performance and efficiency, and show competitive results against baseline methods when tested on the publicly available LETOR OHSUMED data set. We also explore how the available uncertainty information can be used in prediction and how it affects model performance. John Guiver, Edward Lloyd Snelson |
SIGIR | 1 |
| 2008 | SoftRank: optimizing non-smooth rank metricsabstractWe address the problem of learning large complex ranking functions. Most IR applications use evaluation metrics that depend only upon the ranks of documents. However, most ranking functions generate document scores, which are sorted to produce a ranking. Hence IR metrics are innately non-smooth with respect to the scores, due to the sort. Unfortunately, many machine learning algorithms require the gradient of a training objective in order to perform the optimization of the model parameters,and because IR metrics are non-smooth,we need to find a smooth proxy objective that can be used for training. We present a new family of training objectives that are derived from the rank distributions of documents, induced by smoothed scores. We call this approach SoftRank. We focus on a smoothed approximation to Normalized Discounted Cumulative Gain (NDCG), called SoftNDCG and we compare it with three other training objectives in the recent literature. We present two main results. First, SoftRank yields a very good way of optimizing NDCG. Second, we show that it is possible to achieve state of the art test set NDCG results by optimizing a soft NDCG objective on the training set with a different discount function Michael J. Taylor 0001, John Guiver, Stephen E. Robertson, Tom Minka |
WSDM | 2 |