EDBT 2026 Demo / reviewers in the wild / expert
Prasenjit Majumder
dblp:28/5982
· DBLP profile ↗
21ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-0840-9313ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SqCLIRIL: Spoken query cross-lingual information retrieval in Indian languagesabstractThis paper presents SqCLIRIL (Spoken Query Cross-lingual Information Retrieval in Indian Languages), a comprehensive benchmark designed to evaluate spoken query-based cross-lingual retrieval across five Indian languages: Hindi, Gujarati, Bengali, Kannada, and English. The task encompasses monolingual and cross-lingual retrieval settings across five language pairs, incorporating spoken queries (male and female voices) and document collections in text. The primary objective is to assess retrieval effectiveness in a low-resource, multilingual setting that reflects real-world language diversity and access constraints. We investigate four retrieval architectures: (i) sparse lexical matching using BM25, (ii) dense semantic retrieval via bi-encoder models, (iii) hybrid ranking through Reciprocal Rank Fusion (RRF) of sparse and dense scores, and (iv) a Large Language Model (LLM)-based pointwise fusion strategy (LPF) that integrates generative semantic alignment. Experimental evaluations on the human-translated TREC DL’19 and DL’20 query sets in the above five languages show that while dense retrieval improves substantially over traditional sparse models, fusion-based approaches—particularly LPF—consistently yield superior nDCG scores across most query-document language pairs. The results underscore the utility of generative AI in enhancing retrieval performance in multilingual, low-resource, speech-centric IR scenarios. This benchmark contributes to developing scalable, speech-first, and language-agnostic retrieval systems, with implications for inclusive information access in linguistically fragmented regions. Bhargav Dave, Prasenjit Majumder |
Pattern Recognit. Lett. | 2 |
| 2026 | Editorial: Special Section Forum for Information Retrieval Evaluation (FIRE) 2024
Thomas Mandl 0001, Prasenjit Majumder |
Pattern Recognit. Lett. | 2 |
| 2024 | Third Workshop on Augmented Intelligence in Technology-Assisted Review Systems (ALTARS)
Giorgio Maria Di Nunzio, Evangelos Kanoulas, Prasenjit Majumder |
ECIR (5) | 3 |
| 2023 | 2nd Workshop on Augmented Intelligence in Technology-Assisted Review Systems (ALTARS)
Giorgio Maria Di Nunzio, Evangelos Kanoulas, Prasenjit Majumder |
ECIR (3) | 3 |
| 2023 | Text representation for direction prediction of share market
Surupendu Gangopadhyay, Prasenjit Majumder |
Expert Syst. Appl. | 2 |
| 2023 | Detecting offensive speech in conversational code-mixed dialogue on social media: A contextual dataset and benchmark experiments
Hiren Madhu, Shrey Satapara, Sandip Modha, Thomas Mandl 0001, Prasenjit Majumder |
Expert Syst. Appl. | 5 |
| 2022 | Augmented Intelligence in Technology-Assisted Review Systems (ALTARS 2022): Evaluation Metrics and Protocols for eDiscovery and Systematic Review Systems
Giorgio Maria Di Nunzio, Evangelos Kanoulas, Prasenjit Majumder |
ECIR (2) | 3 |
| 2022 | An empirical evaluation of text representation schemes to filter the social media streamabstractModeling text in a numerical representation is a prime task for any Natural Language Processing downstream task such as text classification. This paper attempts to study the effectiveness of text representation schemes on the text classification task, such as aggressive text detection, a special case of Hate speech from social media. Aggression levels are categorized into three predefined classes, namely: ‘Non-aggressive’ (NAG), ‘Overtly Aggressive’ (OAG), and ‘Covertly Aggressive’ (CAG). Various text representation schemes based on BoW techniques, word embedding, contextual word embedding, sentence embedding on traditional classifiers, and deep neural models are compared on a text classification problem. The weighted F1 score is used as a primary evaluation metric. The results show that text representation using Googles’ universal sentence encoder (USE) performs better than word embedding and BoW techniques on traditional classifiers, such as SVM, while pre-trained word embedding models perform better on classifiers based on the deep neural models on the English dataset. Recent pre-trained transfer learning models like Elmo, ULMFi, and BERT are fine-tuned for the aggression classification task. However, results are not at par with the pre-trained word embedding model. Overall, word embedding using pre-trained fastText vectors produces the best weighted F1-score than Word2Vec and Glove. On the Hindi dataset, BoW techniques perform better than word embeddings on traditional classifiers such as SVM. In contrast, pre-trained word embedding models perform better on classifiers based on the deep neural nets. Statistical significance tests are employed to ensure the significance of the classification results. Deep neural models are more robust against the bias induced by the training dataset. They perform substantially better than traditional classifiers, such as SVM, logistic regression, and Naive Bayes classifiers on the Twitter test dataset. Sandip Modha, Prasenjit Majumder, Thomas Mandl 0001 |
J. Exp. Theor. Artif. Intell. | 2 |
| 2020 | Detecting and visualizing hate speech in social media: A cyber Watchdog for surveillance
Sandip Modha, Prasenjit Majumder, Thomas Mandl 0001, Chintak Mandalia |
Expert Syst. Appl. | 2 |
| 2020 | Query specific graph-based query reformulation using UMLS for clinical information access
Jainisha Sankhavara, Rishi Dave, Bhargav Dave, Prasenjit Majumder |
J. Biomed. Informatics | 4 |
| 2020 | Translating Morphologically Rich Indian Languages under Zero-Resource ConditionsabstractThis work presents an in-depth analysis of machine translations of morphologically-rich Indo-Aryan and Dravidian languages under zero-resource conditions. It focuses on Zero-Shot Systems for these languages and leverages transfer-learning by exploiting target-side monolingual corpora and parallel translations from other languages. These systems are compared with direct translations using the BLEU and TER metrics. Further, Zero-Shot Systems are used as pre-trained models for fine-tuning with real human-generated data taken in different proportions that range from 100 sentences to the entire training set. Performances of the Indo-Aryan and Dravidian languages are compared with a focus on their morphological complexity. The systems with a Dravidian source language performed much better and reached very near to the level of direct translations. This is observed likely due to morphological richness and complexity in the language, which in turn provided more room for transfer-learning in this case. A comparative analysis based on language families has been done. These systems were fine-tuned further, which in turn outperformed direct translations with just 500 parallel sentences for a Dravidian source language. However, systems with an Indo-Aryan source language showed similar performance after getting fine-tuned with 10,000 sentences. Ashwani Tanwar, Prasenjit Majumder |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2018 | Content Based Weighted Consensus Summarization
Parth Mehta 0001, Prasenjit Majumder |
ECIR | 2 |
| 2018 | Effective aggregation of various summarization techniques
Parth Mehta 0001, Prasenjit Majumder |
Inf. Process. Manag. | 2 |
| 2016 | Improving Information Retrieval Performance on OCRed Text in the Absence of Clean Text Ground Truth
Kripabandhu Ghosh, Anirban Chakraborty 0002, Swapan K. Parui, Prasenjit Majumder |
Inf. Process. Manag. | 4 |
| 2015 | Learning combination weights in data fusion using Genetic Algorithms
Kripabandhu Ghosh, Swapan K. Parui, Prasenjit Majumder |
Inf. Process. Manag. | 3 |
| 2015 | Approaches to Temporal Expression Recognition in HindiabstractTemporal annotation of plain text is considered a useful component of modern information retrieval tasks. In this work, different approaches for identification and classification of temporal expressions in Hindi are developed and analyzed. First, a rule-based approach is developed, which takes plain text as input and based on a set of hand-crafted rules, produces a tagged output with identified temporal expressions. This approach performs with a strict F1-measure of 0.83. In another approach, a CRF-based classifier is trained with human tagged data and is then tested on a test dataset. The trained classifier identifies the time expressions from plain text and further classifies them to various classes. This approach performs with a strict F1-measure of 0.78. Next, the CRF is replaced by an SVM-based classifier and the same experiment is performed with the same features. This approach is shown to be comparable to the CRF and performs with a strict F1-measure of 0.77. Using the rule base information as an additional feature enhances the performances to 0.86 and 0.84 for the CRF and SVM respectively. With three different comparable systems performing the extraction task, merging them to take advantage of their positives is the next step. As the first merge experiment, rule-based tagged data is fed to the CRF and SVM classifiers as additional training data. Evaluation results report an increase in F1-measure of the CRF from 0.78 to 0.8. Second, a voting-based approach is implemented, which chooses the best class for each token from the outputs of the three approaches. This approach results in the best performance for this task with a strict F1-measure of 0.88. In this process a reusable gold standard dataset for temporal tagging in Hindi is also developed. Named the ILTIMEX2012 corpus, it consists of 300 manually tagged Hindi news documents. Nitin Ramrakhiyani, Prasenjit Majumder |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2013 | Optimum Parameter Selection for K.L.D. Based Authorship Attribution in Gujarati
Parth Mehta 0001, Prasenjit Majumder |
IJCNLP | 2 |
| 2010 | Introduction to the Special Issue on Indian Language Information Retrieval Part Iabstract[not available] Donna K. Harman, Noriko Kando, Prasenjit Majumder, Mandar Mitra, Carol Peters |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2010 | The FIRE 2008 Evaluation ExerciseabstractThe aim of the Forum for Information Retrieval Evaluation (FIRE) is to create an evaluation framework in the spirit of TREC (Text REtrieval Conference), CLEF (Cross-Language Evaluation Forum), and NTCIR (NII Test Collection for IR Systems), for Indian language Information Retrieval. The first evaluation exercise conducted by FIRE was completed in 2008. This article describes the test collections used at FIRE 2008, summarizes the approaches adopted by various participants, discusses the limitations of the datasets, and outlines the tasks planned for the next iteration of FIRE. Prasenjit Majumder, Mandar Mitra, Dipasree Pal, Ayan Bandyopadhyay, Samaresh Maiti, Sukomal Pal, Deboshree Modak, Sucharita Sanyal |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2008 | Text collections for FIREabstractThe aim of the Forum for Information Retrieval Evaluation (FIRE) is to create a Cranfield-like evaluation framework in the spirit of TREC, CLEF and NTCIR, for Indian Language Information Retrieval. For the first year, six Indian languages have been selected: Bengali, Hindi, Marathi, Punjabi, Tamil, and Telugu. This poster describes the tasks as well as the document and topic collections that are to be used at the FIRE workshop. Prasenjit Majumder, Mandar Mitra, Dipasree Pal, Ayan Bandyopadhyay, Samaresh Maiti, Sukanya Mitra, Aparajita Sen, Sukomal Pal |
SIGIR | 1 |
| 2007 | YASS: Yet another suffix stripperabstractStemmers attempt to reduce a word to its stem or root form and are used widely in information retrieval tasks to increase the recall rate. Most popular stemmers encode a large number of language-specific rules built over a length of time. Such stemmers with comprehensive rules are available only for a few languages. In the absence of extensive linguistic resources for certain languages, statistical language processing tools have been successfully used to improve the performance of IR systems. In this article, we describe a clustering-based approach to discover equivalence classes of root words and their morphological variants. A set of string distance measures are defined, and the lexicon for a given text collection is clustered using the distance measures to identify these equivalence classes. The proposed approach is compared with Porter's and Lovin's stemmers on the AP and WSJ subcollections of the Tipster dataset using 200 queries. Its performance is comparable to that of Porter's and Lovin's stemmers, both in terms of average precision and the total number of relevant documents retrieved. The proposed stemming algorithm also provides consistent improvements in retrieval performance for French and Bengali, which are currently resource-poor. Prasenjit Majumder, Mandar Mitra, Swapan K. Parui, Gobinda Kole, Pabitra Mitra, Kalyankumar Datta |
ACM Trans. Inf. Syst. | 1 |