VLDB 2026 Research / reviewers in the wild / expert
Peter B. Golbus
dblp:95/10439
· DBLP profile ↗
8ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 4 first-authorArtificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Information retrieval · 85% Recommender systems · 11% Data mining · 4% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
evaluation |
0.9 | 5 | 2014 | Contextual and dimensional relevance judgments for reusable SERP-level evaluation · WWW 2014 On the information difference between standard retrieval models · SIGIR 2014 A mutual information-based framework for the analysis of information retrieval systems · SIGIR 2013 |
Information retrieval › evaluation
relevance judgment |
0.5 | 3 | 2014 | Contextual and dimensional relevance judgments for reusable SERP-level evaluation · WWW 2014 A document rating system for preference judgements · SIGIR 2013 IR system evaluation using nugget-based test collections · WSDM 2012 |
Information retrieval
retrieval models |
0.5 | 2 | 2020 | Deep Ranking for Style-Aware Room Recommendations (Student Abstract) · AAAI 2020 On the information difference between standard retrieval models · SIGIR 2014 |
Recommender systems
content-based recommendation |
0.4 | 1 | 2020 | Deep Ranking for Style-Aware Room Recommendations (Student Abstract) · AAAI 2020 |
Information retrieval › retrieval models › neural retrieval
neural ranking model |
0.4 | 1 | 2020 | Deep Ranking for Style-Aware Room Recommendations (Student Abstract) · AAAI 2020 |
Information retrieval
ranking |
0.4 | 1 | 2020 | Deep Ranking for Style-Aware Room Recommendations (Student Abstract) · AAAI 2020 |
Information retrieval › evaluation
effectiveness metrics |
0.2 | 1 | 2014 | Contextual and dimensional relevance judgments for reusable SERP-level evaluation · WWW 2014 |
Data mining › dimensionality reduction › feature selection
mutual information |
0.2 | 1 | 2013 | A mutual information-based framework for the analysis of information retrieval systems · SIGIR 2013 |
Information retrieval › evaluation › relevance judgment
preference judgments |
0.2 | 1 | 2013 | A document rating system for preference judgements · SIGIR 2013 |
Information retrieval › evaluation › evaluation methodology
nugget-based evaluation |
0.1 | 1 | 2012 | IR system evaluation using nugget-based test collections · WSDM 2012 |
Information retrieval › evaluation
test collection |
0.1 | 1 | 2012 | IR system evaluation using nugget-based test collections · WSDM 2012 |
Methods — techniques the papers use, named apart from their topics
deep learning · 0.4comparison learning · 0.4probabilistic framework · 0.4user preference correlation · 0.2information theory · 0.2mutual information · 0.2elo rating system · 0.2matching function · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Deep Ranking for Style-Aware Room Recommendations (Student Abstract)abstractWe present a deep learning based room image retrieval framework that is based on style understanding. Given a dataset of room images labeled by interior design experts, we map the noisy style labels to comparison labels. Our framework learns the style spectrum of each image from the generated comparisons and makes significantly more accurate recommendations compared to discrete classification baselines. Ilkay Yildiz, Esra Ataer Cansizoglu, Hantian Liu, Peter B. Golbus, Ozan Tezcan, Jaewoo Choi 0003 |
AAAI | 4 |
| 2014 | On the information difference between standard retrieval modelsabstractRecent work introduced a probabilistic framework that measures search engine performance information-theoretically. This allows for novel meta-evaluation measures such as Information Difference, which measures the magnitude of the difference between search engines in their ranking of documents. for which we have relevance information. Using Information Difference we can compare the behavior of search engines-which documents the search engine prefers, as well as search engine performance-how likely the search engine is to satisfy a hypothetical user. In this work, we a) extend this probabilistic framework to precision-oriented contexts, b) show that Information Difference can be used to detect similar search engines at shallow ranks, and c) demonstrate the utility of the Information Difference methodology by showing that well-tuned search engines employing different retrieval models are more similar than a well-tuned and a poorly tuned implementation of the same retrieval model. Peter B. Golbus, Javed A. Aslam |
SIGIR | 1 |
| 2014 | Contextual and dimensional relevance judgments for reusable SERP-level evaluationabstractDocument-level relevance judgments are a major component in the calculation of effectiveness metrics. Collecting high-quality judgments is therefore a critical step in information retrieval evaluation. However, the nature of and the assumptions underlying relevance judgment collection have not received much attention. In particular, relevance judgments are typically collected for each document in isolation, although users read each document in the context of other documents. In this work, we aim to investigate the nature of relevance judgment collection. We collect relevance labels in both isolated and conditional setting, and ask for judgments in various dimensions of relevance as well as overall relevance. Then we compare the relevance metrics based on various types of judgments with other metrics of quality such as user preference. Our analyses illuminate how these settings for judgment collection affect the quality and the characteristics of the judgments. We also find that the metrics based on conditional judgments show higher correlation with user preference than isolated judgments. Peter B. Golbus, Imed Zitouni, Jin Young Kim 0005, Ahmed Awadallah 0001, Fernando Diaz 0001 |
WWW | 1 |
| 2013 | A document rating system for preference judgementsabstractHigh quality relevance judgments are essential for the evaluation of information retrieval systems. Traditional methods of collecting relevance judgments are based on collecting binary or graded nominal judgments, but such judgments are limited by factors such as inter-assessor disagreement and the arbitrariness of grades. Previous research has shown that it is easier for assessors to make pairwise preference judgments. However, unless the preferences collected are largely transitive, it is not clear how to combine them in order to obtain document relevance scores. Another difficulty is that the number of pairs that need to be assessed is quadratic in the number of documents. In this work, we consider the problem of inferring document relevance scores from pairwise preference judgments by analogy to tournaments using the Elo rating system. We show how to combine a linear number of pairwise preference judgments from multiple assessors to compute relevance scores for every document. Maryam Bashir, Jesse Anderton, Jie Wu 0024, Peter B. Golbus, Virgil Pavlu, Javed A. Aslam |
SIGIR | 4 |
| 2013 | A mutual information-based framework for the analysis of information retrieval systemsabstractWe consider the problem of information retrieval evaluation and the methods and metrics used for such evaluations. We propose a probabilistic framework for evaluation which we use to develop new information-theoretic evaluation metrics. We demonstrate that these new metrics are powerful and generalizable, enabling evaluations heretofore not possible. Peter B. Golbus, Javed A. Aslam |
SIGIR | 1 |
| 2013 | Increasing evaluation sensitivity to diversity
Peter B. Golbus, Javed A. Aslam, Charles L. A. Clarke |
Inf. Retr. | 1 |
| 2012 | IR system evaluation using nugget-based test collectionsabstractThe development of information retrieval systems such as search engines relies on good test collections, including assessments of retrieved content. The widely employed Cranfield paradigm dictates that the information relevant to a topic be encoded at the level of documents, therefore requiring effectively complete document relevance assessments. As this is no longer practical for modern corpora, numerous problems arise, including scalability, reusability, and applicability. We propose a new method for relevance assessment based on relevant information, not relevant documents. Once the relevant 'nuggets' are collected, our matching method can assess any document for relevance with high accuracy, and so any retrieved list of documents can be assessed for performance. In this paper we analyze the performance of the matching function by looking at specific cases and by comparing with other methods. We then show how these inferred relevance assessments can be used to perform IR system evaluation, and we discuss in particular reusability and scalability. Our main contribution is a methodology for producing test collections that are highly accurate, more complete, scalable, reusable, and can be generated with similar amounts of effort as existing methods, with great potential for future applications. Virgil Pavlu, Shahzad Rajput, Peter B. Golbus, Javed A. Aslam |
WSDM | 3 |
| 2011 | A nugget-based test collection construction paradigmabstractThe problem of building test collections is central to the development of information retrieval systems such as search engines. Starting with a few relevant "nuggets" of information manually extracted from existing TREC corpora, we implement and test a methodology that finds and correctly assesses the vast majority of relevant documents found by TREC assessors - as well as up to four times more additional relevant documents. Our methodology produces highly accurate test collections that hold the promise of addressing the issues of scalability, reusability, and applicability. Shahzad Rajput, Virgil Pavlu, Peter B. Golbus, Javed A. Aslam |
CIKM | 3 |