EDBT 2026 Demo / reviewers in the wild / expert
Chengxi Luo
dblp:297/0799
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2022
0009-0000-3733-1717ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
dueling bandit |
0.6 | 1 | 2022 | Human Preferences as Dueling Bandits · SIGIR 2022 |
Information retrieval › evaluation › relevance judgment
preference judgments |
0.6 | 1 | 2022 | Human Preferences as Dueling Bandits · SIGIR 2022 |
Information retrieval
retrieval evaluation |
0.6 | 1 | 2022 | Human Preferences as Dueling Bandits · SIGIR 2022 |
Information retrieval › evaluation › test collection
test collection construction |
0.6 | 1 | 2022 | Human Preferences as Dueling Bandits · SIGIR 2022 |
Information retrieval
evaluation |
0.5 | 1 | 2021 | Evaluation Measures Based on Preference Graphs · SIGIR 2021 |
Information retrieval › evaluation › user-oriented evaluation
preference-based evaluation |
0.5 | 1 | 2021 | Evaluation Measures Based on Preference Graphs · SIGIR 2021 |
Information retrieval
ranking |
0.5 | 1 | 2021 | Evaluation Measures Based on Preference Graphs · SIGIR 2021 |
Information retrieval › similarity measure
rank similarity measure |
0.5 | 1 | 2021 | Evaluation Measures Based on Preference Graphs · SIGIR 2021 |
Information retrieval › retrieval models › neural retrieval
neural ranking model |
0.2 | 1 | 2022 | Human Preferences as Dueling Bandits · SIGIR 2022 |
Information retrieval
retrieval models |
0.2 | 1 | 2022 | Human Preferences as Dueling Bandits · SIGIR 2022 |
Information retrieval › evaluation › effectiveness metrics
discounted cumulative gain |
0.1 | 1 | 2021 | Evaluation Measures Based on Preference Graphs · SIGIR 2021 |
Methods — techniques the papers use, named apart from their topics
interleaving · 0.6dueling bandits · 0.6rank-biased overlap · 0.5preference graphs · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Human Preferences as Dueling BanditsabstractThe dramatic improvements in core information retrieval tasks engendered by neural rankers create a need for novel evaluation methods. If every ranker returns highly relevant items in the top ranks, it becomes difficult to recognize meaningful differences between them and to build reusable test collections. Several recent papers explore pairwise preference judgments as an alternative to traditional graded relevance assessments. Rather than viewing items one at a time, assessors view items side-by-side and indicate the one that provides the better response to a query, allowing fine-grained distinctions. If we employ preference judgments to identify the probably best items for each query, we can measure rankers by their ability to place these items as high as possible. We frame the problem of finding best items as a dueling bandits problem. While many papers explore dueling bandits for online ranker evaluation via interleaving, they have not been considered as a framework for offline evaluation via human preference judgments. We review the literature for possible solutions. For human preference judgments, any usable algorithm must tolerate ties, since two items may appear nearly equal to assessors, and it must minimize the number of judgments required for any specific pair, since each such comparison requires an independent assessor. Since the theoretical guarantees provided by most algorithms depend on assumptions that are not satisfied by human preference judgments, we simulate selected algorithms on representative test cases to provide insight into their practical utility. Based on these simulations, one algorithm stands out for its potential. Our simulations suggest modifications to further improve its performance. Using the modified algorithm, we collect over 10,000 preference judgments for pools derived from submissions to the TREC 2021 Deep Learning Track, confirming its suitability. We test the idea of best-item evaluation and suggest ideas for further theoretical and practical progress. Xinyi Yan, Chengxi Luo, Charles L. A. Clarke, Nick Craswell, Ellen M. Voorhees, Pablo Castells |
SIGIR | 2 |
| 2021 | Evaluation Measures Based on Preference GraphsabstractThe offline evaluation of search requires us to define a standard against which we measure the quality of results returned by a ranker. Frequently this standard is defined in absolute terms through relevance grades, but it can also be defined in relative terms through preferences. These preferences might be created through explicit preference judgments, derived from relevance grades, or inferred from clicks and other signals. Preferences from multiple sources might even be combined. In contrast to absolute grades, preferences avoid complex definitions of relevance, indicating only that a ranker should favor one result over another. Despite the simplicity and flexibility of preferences, widespread adoption has been limited by the lack of established evaluation measures. Recent work in this direction has taken two approaches: 1) measures based on weighted counts of agreements and disagreements between a set of preferences and an actual ranking generated by a ranker; and 2) measures that translate preferences into gain values for use with traditional measures, such as nDCG. Both approaches require methods for specifying weights or gains that have little or no theoretical foundation, and the values of these measures have no clear and meaningful interpretation. To address these problems, we propose an evaluation measure that computes the similarity between a directed multigraph of preferences and an actual ranking generated by a ranker. The measure computes an ordering for the vertices of the preference graph that maximizes its similarity to the actual ranking under a rank similarity measure. This maximum similarity becomes the value of the measure. Preference graphs are often acyclic, or nearly so, and to compute the measure we extend an approximate greedy algorithm that is known to produce good results for nearly acyclic graphs. For the rank similarity measure we employ Rank Biased Overlap (RBO) which was explicitly created to match the requirements of search and related applications. We validate the new measure over several collections of preferences explored in recent work. Charles L. A. Clarke, Chengxi Luo, Mark D. Smucker |
SIGIR | 2 |