VLDB 2026 Research / reviewers in the wild / expert
Teerapong Leelanupab
dblp:39/7567
· DBLP profile ↗
9ranked-venue papers in the field
2as first author
5since 2021 · last 2026
0000-0002-8117-0612ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 9 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Where Relevance Emerges: A Layer-Wise Study of Internal Attention for Zero-Shot Re-RankingabstractZero-shot document re-ranking with Large Language Models (LLMs) has evolved from Pointwise methods to Listwise and Setwise approaches that optimize computational efficiency. Despite their success, these methods predominantly rely on generative scoring or output logits, which face bottlenecks in inference latency and result consistency. In-Context Re-ranking (ICR) has recently been proposed as an O(1) alternative method. ICR extracts internal attention signals directly, avoiding the overhead of text generation. However, existing ICR methods simply aggregate signals across all layers; layer-wise contributions and their consistency across architectures have been left unexplored. Furthermore, no unified study has compared internal attention with traditional generative and likelihood-based mechanisms across diverse ranking frameworks under consistent conditions. Shengyao Zhuang, Zheng Yao 0004, Guido Zuccon, Teerapong Leelanupab |
SIGIR | 5 |
| 2026 | When LLM Judges Inflate Scores: Exploring Overrating in Relevance AssessmentabstractHuman relevance assessment is time-consuming and cognitively intensive, limiting the scalability of Information Retrieval evaluation. This has led to growing interest in using large language models (LLMs) as proxies for human judges. However, it remains an open question whether LLM-based relevance judgments are reliable, stable, and rigorous enough to match humans for relevance assessment. In this work, we conduct a study of overrating behavior in LLM-based relevance judgments across model backbones, evaluation paradigms (pointwise and pairwise), and passage modification strategies. We show that models consistently assign inflated relevance scores-often with high confidence-to passages that do not genuinely satisfy the underlying information need, revealing a system-wide bias rather than random fluctuations in judgment. Furthermore, controlled experiments show that LLM-based relevance judgments can be highly sensitive to passage length and surface-level lexical cues. These results raise concerns about the usage of LLMs as drop-in replacements for human relevance assessors, and highlight the urgent need for careful diagnostic evaluation frameworks when applying LLMs for relevance assessments. Our code and results are publicly available. https://github.com/chutingyu/Exploring-Overrating Chuting Yu, Hang Li 0009, Guido Zuccon, Joel Mackenzie, Teerapong Leelanupab |
SIGIR | 5 |
| 2025 | DenseReviewer: A Screening Prioritisation Tool for Systematic Review Based on Dense Retrieval
Xinyu Mao 0001, Teerapong Leelanupab, Harrisen Scells, Guido Zuccon |
ECIR (5) | 2 |
| 2025 | AiReview: An Open Platform for Accelerating Systematic Reviews with LLMsabstractSystematic reviews are fundamental to evidence-based medicine.Creating one is time-consuming and labour-intensive, mainly due to the need to screen, or assess, many studies for inclusion in the review.Existing tools help streamline this process, mostly using traditional machine learning.Large language models (LLMs) offer new opportunities to speed up screening, yet no tool currently enables users to directly apply LLMs or ensures systematic and transparent use of these methods.This paper presents (i) a flexible framework for using LLMs in systematic review tasks, especially title and abstract screening, and (ii) a web-based interface for LLMassisted screening.Together, they form AiReview-a novel platform that connects cutting-edge LLM-assisted screening methods with real-world systematic review practice.The live tool is available at https://aireview.ielab.io.We also release the code publicly at https://github.com/ielab/ai-review. Xinyu Mao 0001, Teerapong Leelanupab, Martin Potthast, Harrisen Scells, Guido Zuccon |
SIGIR | 2 |
| 2024 | Embark on DenseQuest: A System for Selecting the Best Dense Retriever for a Custom CollectionabstractIn this demo we present a web-based application for selecting an effective pre-trained dense retriever to use on a private collection. Our system, DenseQuest, provides unsupervised selection and ranking capabilities to predict the best dense retriever among a pool of available dense retrievers, tailored to an uploaded target collection. DenseQuest implements a number of existing approaches, including a recent, highly effective method powered by Large Language Models (LLMs), which requires neither queries nor relevance judgments. The system is designed to be intuitive and easy to use for those information retrieval engineers and researchers who need to identify a general-purpose dense retrieval model to encode or search a new private target collection. Our demonstration illustrates conceptual architecture and the different use case scenarios of the system implemented on the cloud, enabling universal access and use. DenseQuest is available at https://densequest.ielab.io. Ekaterina Khramtsova, Teerapong Leelanupab, Shengyao Zhuang, Mahsa Baktash, Guido Zuccon |
SIGIR | 2 |
| 2013 | Is Intent-Aware Expected Reciprocal Rank Sufficient to Evaluate Diversity?
Teerapong Leelanupab, Guido Zuccon, Joemon M. Jose |
ECIR | 1 |
| 2013 | Crowdsourcing interactions: using crowdsourcing for evaluating interactive information retrieval systems
Guido Zuccon, Teerapong Leelanupab, Stewart Whiting, Emine Yilmaz, Joemon M. Jose, Leif Azzopardi |
Inf. Retr. | 2 |
| 2012 | A comprehensive analysis of parameter settings for novelty-biased cumulative gainabstractIn the TREC Web Diversity track, novelty-biased cumulative gain (α-NDCG) is one of the official measures to assess retrieval performance of IR systems. The measure is characterised by a parameter, α, the effect of which has not been thoroughly investigated. We find that common settings of α, i.e. α=0.5, may prevent the measure from behaving as desired when evaluating result diversification. This is because it excessively penalises systems that cover many intents while it rewards those that redundantly cover only few intents. This issue is crucial since it highly influences systems at top ranks. We revisit our previously proposed threshold, suggesting α be set on a query-basis. The intuitiveness of the measure is then studied by examining actual rankings from TREC 09-10 Web track submissions. By varying α according to our query-based threshold, the discriminative power of α-NDCG is not harmed and in fact, our approach improves α-NDCG's robustness. Experimental results show that the threshold for α can turn the measure to be more intuitive than using its common settings. Teerapong Leelanupab, Guido Zuccon, Joemon M. Jose |
CIKM | 1 |
| 2012 | CrowdTiles: presenting crowd-based information for event-driven information needsabstractTime plays a central role in many web search information needs relating to recent events. For recency queries where fresh information is most desirable, there is likely to be a great deal of highly-relevant information created very recently by crowds of people across the world, particularly on platforms such as Wikipedia and Twitter. With so many users, mainstream events are often very quickly reflected in these sources. The English Wikipedia encyclopedia consists of a vast collection of user-edited articles covering a range of topics. During events, users collaboratively create and edit existing articles in near real-time. Simultaneously, users on Twitter disseminate and discuss event details, with a small number of users becoming influential for the topic. Stewart Whiting, Ke Zhou 0003, Joemon M. Jose, Omar Alonso, Teerapong Leelanupab |
CIKM | 5 |