VLDB 2026 Research / reviewers in the wild / expert
Samarth Bhargav 0001
dblp:228/6970
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2023
0000-0001-5204-8514ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Market-Aware Models for Efficient Cross-Market Recommendation
Samarth Bhargav 0001, Mohammad Aliannejadi, Evangelos Kanoulas |
ECIR (1) | 1 |
| 2023 | When the Music Stops: Tip-of-the-Tongue Retrieval for MusicabstractWe present a study of Tip-of-the-tongue (ToT) retrieval for music, where a searcher is trying to find an existing music entity, but is unable to succeed as they cannot accurately recall important identifying information. ToT information needs are characterized by complexity, verbosity, uncertainty, and possible false memories. We make four contributions. (1) We collect a dataset - TOTMUSIC--of 2,278 information needs and ground truth answers. (2) We introduce a schema for these information needs and show that they often involve multiple modalities encompassing several Music IR sub-tasks such as lyric search, audio-based search, audio fingerprinting, and text search. (3) We underscore the difficulty of this task by benchmarking a standard text retrieval approach on this dataset. (4) We investigate the efficacy of query reformulations generated by a Large Language Model (LLM), and show that they are not as effective as simply employing the entire information need as a query--leaving several open questions for future research. Samarth Bhargav 0001, Anne Schuth, Claudia Hauff |
SIGIR | 1 |
| 2022 | Reproducibility as a Mechanism for Teaching Fairness, Accountability, Confidentiality, and Transparency in Artificial IntelligenceabstractIn this work, we explain the setup for a technical, graduate-level course on Fairness, Accountability, Confidentiality, and Transparency in Artificial Intelligence (FACT-AI) at the University of Amsterdam, which teaches FACT-AI concepts through the lens of reproducibility. The focal point of the course is a group project based on reproducing existing FACT-AI algorithms from top AI conferences and writing a corresponding report. In the first iteration of the course, we created an open source repository with the code implementations from the group projects. In the second iteration, we encouraged students to submit their group projects to the Machine Learning Reproducibility Challenge, resulting in 9 reports from our course being accepted for publication in the ReScience journal. We reflect on our experience teaching the course over two years, where one year coincided with a global pandemic, and propose guidelines for teaching FACT-AI through reproducibility in graduate-level AI study programs. We hope this can be a useful resource for instructors who want to set up similar courses in the future. Ana Lucic, Maurits J. R. Bleeker, Sami Jullien, Samarth Bhargav 0001, Maarten de Rijke |
AAAI | 4 |
| 2022 | 'It's on the tip of my tongue': A new Dataset for Known-Item RetrievalabstractThe tip of the tongue known-item retrieval (TOT-KIR) task involves the 'one-off' retrieval of an item for which a user cannot recall a precise identifier. The emergence of several online communities where users pose known-item queries to other users indicates the inability of existing search systems to answer such queries. Research in this domain is hampered by the lack of large, open or realistic datasets. Prior datasets relied on either annotation by crowd workers, which can be expensive and time-consuming, or generating synthetic queries, which can be unrealistic. Additionally, small datasets make the application of modern (neural) retrieval methods unviable, since they require a large number of data-points. In this paper, we collect the largest dataset yet with 15K query-item pairs in two domains, namely, Movies and Books, from an online community using heuristics, rendering expensive annotation unnecessary while ensuring that queries are realistic. We show that our data collection method is accurate by conducting a data study. We further demonstrate that methods like BM25 fall short of answering such queries, corroborating prior research. The size of the dataset makes neural methods feasible, which we show outperforms lexical baselines, indicating that neural/dense retrieval is superior for the TOT-KIR task. Samarth Bhargav 0001, Georgios Sidiropoulos, Evangelos Kanoulas |
WSDM | 1 |
| 2021 | Robustness Evaluation of Entity Disambiguation Using Prior Probes: the Case of Entity OvershadowingabstractEntity disambiguation (ED) is the last step of entity linking (EL), when candidate entities are reranked according to the context they appear in.All datasets for training and evaluating models for EL consist of convenience samples, such as news articles and tweets, that propagate the prior probability bias of the entity distribution towards more frequently occurring entities.It was previously shown that performance of EL systems on such datasets is overestimated, since it is possible to obtain higher accuracy scores by merely learning the prior.To provide a more adequate evaluation benchmark, we introduce the ShadowLink dataset, which includes 16K short text snippets annotated with entity mentions.We evaluate and report the performance of several popular EL systems on the ShadowLink benchmark.The results show a considerable difference in accuracy between common and uncommon ambiguous entities that require disambiguation, for all of the EL systems under evaluation, demonstrating the effects of prior probability bias and entity overshadowing. Vera Provatorova, Samarth Bhargav 0001, Svitlana Vakulenko, Evangelos Kanoulas |
EMNLP (1) | 2 |
| 2019 | Sinkhorn AutoEncoders
Giorgio Patrini, Rianne van den Berg, Patrick Forré, Marcello Carioni, Samarth Bhargav 0001, Max Welling, Tim Genewein, Frank Nielsen |
UAI | 5 |