EDBT 2026 Demo / reviewers in the wild / expert
Pranav Kasela
dblp:331/3588
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-0972-2424ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reproducible experiments on visual exploration framework of geospatial vector big data
Zebang Liu, Anran Yang, Mengyu Ma, Jiali Zhou, Ning Jing, Jichong Yin, Pranav Kasela, Raúl Martín-Santamaría |
Inf. Syst. | 8 |
| 2026 | Toward Exploring Mixed-Initiative Conversation Generation Based on Community Question AnsweringabstractConversational search addresses users’ information needs through multi-turn and context-aware interactions. Given that user queries are often ambiguous, the use of clarifying questions can effectively reduce uncertainty and enable a mixed-initiative conversational system. However, current datasets for clarifying questions remain limited in the following three aspects: (1) underrepresented multi-turn conversational data, (2) limited diversity, and (3) heavily reliance on crowdsourcing, thereby suffering from limitations such as high annotation cost. To address these issues, we propose a large language model (LLM)-based three-stage framework that relies on an existing community question answering dataset. It encompasses: (1) extracting essential information from the initial user query with the relevant contextual information, (2) generating clarifying questions paired with corresponding answers, and (3) refining conversations to ensure coherence and a natural conversational flow. We assess our multi-stage method against a baseline that directly prompts LLMs to generate conversations in a single-step process, evaluating on an answer retrieval task using recall, precision, normalized discounted cumulative gain and mean average precision. Results show that our three-stage generation approach consistently outperforms the baseline particularly in recall, while also achieving competitive results across other metrics. Human and automatic evaluations further indicate the high quality of generated conversations and fine-tuning on them improves retrieval performance, highlighting the pipeline’s potential. Lili Lu, Pranav Kasela, Federico Ravenda, Chuan Meng, Gabriella Pasi, Fabio Crestani |
ACM Trans. Inf. Syst. | 2 |
| 2025 | Leveraging Cognitive Complexity of Texts for Contextualization in Dense RetrievalabstractDense Retrieval Models (DRMs) estimate the semantic similarity between queries and documents based on their embeddings.Prior studies highlight the importance of embedding contextualization in enhancing retrieval performance.To this aim, existing approaches primarily leverage token-level information derived from query/document interactions.In this paper, we introduce a novel DRM, namely DenseC3, which leverages query/document interactions based on the full embedding representations generated by a Transformer-based model.To enhance similarity estimation, DenseC3 integrates external linguistic information about the Cognitive Complexity of texts, enriching the contextualization of embeddings.We empirically evaluate our approach across seven benchmarks and three different IR tasks to assess the impact of Cognitive Complexity-aware query and document embeddings for contextualization in dense retrieval.Results show that our approach consistently outperforms standard fine-tuning techniques on lightweight biencoders (e.g., BERT-based) and traditional late-interaction models (i.e., ColBERT) across all benchmarks.On larger retrieval-optimized bi-encoders like Contriever, our model achieves comparable or higher performance on four of the considered evaluation benchmarks.Our findings suggest that Cognitive Complexityaware embeddings enhance query and document representations, improving Effrosyni Sokli, Georgios Peikos, Pranav Kasela, Gabriella Pasi |
EMNLP | 3 |
| 2025 | Investigating Task Arithmetic for Zero-Shot Information RetrievalabstractLarge Language Models (LLMs) have shown impressive zero-shot performance across a variety of Natural Language Processing tasks, including document re-ranking. However, their effectiveness degrades on unseen tasks and domains, largely due to shifts in vocabulary and word distributions. In this paper, we investigate Task Arithmetic, a technique that combines the weights of LLMs pre-trained on different tasks or domains via simple mathematical operations, such as addition or subtraction, to adapt retrieval models without requiring additional fine-tuning. Our method is able to synthesize diverse tasks and domain knowledge into a single model, enabling effective zero-shot adaptation in different retrieval contexts. Extensive experiments on publicly available scientific, biomedical, and multilingual datasets show that our method improves state-of-the-art re-ranking performance by up to 18% in NDCG@10 and 15% in P@10. In addition to these empirical gains, our analysis provides insights into the strengths and limitations of Task Arithmetic as a practical strategy for zero-shot learning and model adaptation. We make our code publicly available at https://github.com/DetectiveMB/Task-Arithmetic-for-ZS-IR. Marco Braga 0001, Pranav Kasela, Alessandro Raganato, Gabriella Pasi |
SIGIR | 2 |
| 2025 | PARK: Personalized academic retrieval with knowledge-graphsabstractAcademic Search is a search task aimed to manage and retrieve scientific documents like journal articles and conference papers. Personalization in this context meets individual researchers’ needs by leveraging, through user profiles, the user related information (e.g. documents authored by a researcher), to improve search effectiveness and to reduce the information overload. While citation graphs are a valuable means to support the outcome of recommender systems, their use in personalized academic search (with, e.g. nodes as papers and edges as citations) is still under-explored. Existing personalized models for academic search often struggle to fully capture users’ academic interests. To address this, we propose a two-step approach: first, training a neural language model for retrieval, then converting the academic graph into a knowledge graph and embedding it into a shared semantic space with the language model using translational embedding techniques. This allows user models to capture both explicit relationships and hidden structures in citation graphs and paper content. We evaluate our approach in four academic search domains, outperforming traditional graph-based and personalized models in three out of four, with up to a 10% improvement in MAP@100 over the second-best model. This highlights the potential of knowledge graph-based user models to enhance retrieval effectiveness. Pranav Kasela, Gabriella Pasi, Raffaele Perego 0001 |
Inf. Syst. | 1 |
| 2024 | DESIRE-ME: Domain-Enhanced Supervised Information Retrieval Using Mixture-of-Experts
Pranav Kasela, Gabriella Pasi, Raffaele Perego 0001, Nicola Tonellotto |
ECIR (2) | 1 |
| 2022 | A Multi-Domain Benchmark for Personalized Search EvaluationabstractPersonalization in Information Retrieval has been a hot topic in both academia and industry for the past two decades. However, there is still a lack of high-quality standard benchmark datasets for conducting offline comparative evaluations in this context. To mitigate this problem, in the past few years, approaches to derive synthetic datasets suited for evaluating Personalized Search models have been proposed. In this paper, we put forward a novel evaluation benchmark for Personalized Search with more than 18 million documents and 1.9 million queries across four domains. We present a detailed description of the benchmark construction procedure, highlighting its characteristics and challenges. We provide baseline performance including pre-trained neural models, opening room for the evaluation of personalized approaches, as well as domain adaptation and transfer learning scenarios. We make both datasets and models available for future research. Elias Bassani, Pranav Kasela, Alessandro Raganato, Gabriella Pasi |
CIKM | 2 |