VLDB 2026 Research / reviewers in the wild / expert
Utshab Kumar Ghosh
dblp:335/8323
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2026
0000-0003-3096-6909ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › retrieval models › neural retrieval
late interaction retrieval |
1.0 | 1 | 2026 | Reproduction Beyond Benchmarks: ConstBERT and ColBERT-v2 Across Backends and Query Distributions · SIGIR 2026 |
Information retrieval › retrieval models › neural retrieval
multi-vector retrieval |
1.0 | 1 | 2026 | Reproduction Beyond Benchmarks: ConstBERT and ColBERT-v2 Across Backends and Query Distributions · SIGIR 2026 |
Information retrieval › evaluation › evaluation methodology
reproducibility |
1.0 | 1 | 2026 | Reproduction Beyond Benchmarks: ConstBERT and ColBERT-v2 Across Backends and Query Distributions · SIGIR 2026 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 1.0ablation study · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reproduction Beyond Benchmarks: ConstBERT and ColBERT-v2 Across Backends and Query DistributionsabstractReproducibility must validate architectural robustness, not just numerical accuracy. We evaluate ColBERT-v2 and ConstBERT across five dimensions, finding that while ConstBERT reproduces within 0.05% MRR@10 on MS-MARCO, both models show a drop of 86–97% on long, narrative queries (TREC ToT 2025). Ablations prove this failure is architectural: performance plateaus at 20 words because the MaxSim operator's uniform token weighting cannot distinguish signal from filler noise. Furthermore, undocumented backend parameters create an 8-point gap due to ConstBERT's sparse centroid coverage, and fine-tuning with 3× more data actually degrades performance by up to 29%. We conclude that architectural constraints in multi-vector retrieval cannot be overcome by adaptation alone. Code: https://github.com/utshabkg/multi-vector-reproducibility. Utshab Kumar Ghosh, Ashish David, Shubham Chatterjee |
SIGIR | 1 |