VLDB 2026 Research / reviewers in the wild / expert
Sudip Kumar Naskar
dblp:14/987
· DBLP profile ↗
37ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-1588-4665ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and RankersabstractLeveraging both labeled (input-output associations) and unlabeled data (wider contextual grounding) may provide complementary benefits in retrieval augmented generation (RAG). However, effectively combining evidence from these heterogeneous sources is challenging as the respective similarity scores are not inter-comparable. Additionally, aggregating beliefs from the outputs of multiple rankers can improve the effectiveness of RAG. Our proposed method first aggregates the top-documents from a number of IR models using a standard rank fusion technique for each source (labeled and unlabeled). Next, we standardize the retrieval score distributions within each source by applying z-score transformation before merging the top-retrieved documents from the two sources. We evaluate our approach on the fact verification task, demonstrating that it consistently improves over the best-performing individual ranker or source and also shows better out-of-domain generalization. Payel Santra, Madhusudan Ghosh, Debasis Ganguly, Partha Basuchowdhuri, Sudip Kumar Naskar |
CIKM | 5 |
| 2024 | IndicFinNLP: Financial Natural Language Processing for Indian LanguagesabstractApplications of Natural Language Processing (NLP) in the finance domain have been very popular of late. For financial NLP, (FinNLP) while various datasets exist for widely spoken languages like English and Chinese, datasets are scarce for low resource languages,particularly for Indian languages. In this paper, we address this challenges by presenting IndicFinNLP – a collection of 9 datasets consisting of three tasks relating to FinNLP for three Indian languages. These tasks are Exaggerated Numeral Detection, Sustainability Classification, and ESG Theme Determination of financial texts in Hindi, Bengali, and Telugu. Moreover, we release the datasets under CC BY-NC-SA 4.0 license for the benefit of the research community. Sohom Ghosh, Arnab Maji, Aswartha Narayana, Sudip Kumar Naskar |
LREC/COLING | 4 |
| 2024 | "The Absence of Evidence is Not the Evidence of Absence": Fact Verification via Information Retrieval-Based In-Context Learning
Payel Santra, Madhusudan Ghosh, Debasis Ganguly, Partha Basuchowdhuri, Sudip Kumar Naskar |
DaWaK | 5 |
| 2023 | Extracting Methodology Components from AI Research Papers: A Data-driven Factored Sequence Labeling ApproachabstractExtraction of methodology component names from scientific articles is a challenging task due to the diversified contexts around the occurrences of these entities, and the different levels of granularity and containment relationships exhibited by these entities. We hypothesize that standard sequence labeling approaches may not adequately model the dependence of methodology name mentions with their contexts, due to the problems of their large, fast evolving, and domain-specific vocabulary. As a solution, we propose a factored approach, where the mention-context dependencies are represented in a more fine-grained manner, thus allowing the model parameters to better adjust to the different characteristic patterns inherent within the data. In particular, we experiment with two variants of this factored approach - one that uses the per-entity category information derived from an ontology, and the other that makes use of the topology of the sentence embedding space to infer a category for each entity constituting that sentence. We demonstrate that both these factored variants of SciBERT outperform their non-factored counterpart, a state-of-the-art model for scientific concept extraction. Madhusudan Ghosh, Debasis Ganguly, Partha Basuchowdhuri, Sudip Kumar Naskar |
CIKM | 4 |
| 2020 | The Transference Architecture for Automatic Post-EditingabstractIn automatic post-editing (APE) it makes sense to condition post-editing (pe) decisions on both the source (src) and the machine translated text (mt) as input.This has led to multi-encoder based neural APE approaches.A research challenge now is the search for architectures that best support the capture, preparation and provision of src and mt information and its integration with pe decisions.In this paper we present an efficient multi-encoder based APE model, called transference.Unlike previous approaches, it (i) uses a transformer encoder block for src, (ii) followed by a decoder block, but without masking for self-attention on mt, which effectively acts as second encoder combining src → mt, and (iii) feeds this representation into a final decoder block generating pe.Our model outperforms the best performing systems by 1 BLEU point on the WMT 2016, 2017, and 2018 English-German APE shared tasks (PBSMT and NMT).Furthermore, the results of our model on the WMT 2019 APE task using NMT data shows performance at the level of the state-of-the-art.The inference time of our model is similar to the vanilla transformer-based NMT system although our model deals with two separate encoders.We further investigate the importance of our newly introduced second encoder and find that decreasing the number of layers hurts performance, while reducing the number of layers of the decoder does not matter much. Santanu Pal, Hongfei Xu, Nico Herbig 0001, Sudip Kumar Naskar, Antonio Krüger, Josef van Genabith |
COLING | 4 |
| 2020 | Online Bangla handwritten word recognition using HMM and language model
Shibaprasad Sen, Ankan Bhattacharyya, Mridul Mitra, Kaushik Roy 0004, Sudip Kumar Naskar, Ram Sarkar |
Neural Comput. Appl. | 5 |
| 2019 | Improving CAT Tools in the Translation Workflow: New Approaches and Evaluation
Mihaela Vela, Santanu Pal, Marcos Zampieri, Sudip Kumar Naskar, Josef van Genabith |
MTSummit (2) | 4 |
| 2019 | Word Difficulty Prediction Using Convolutional Neural NetworksabstractMost text-simplification systems require an indicator of the complexity of the words. The prevalent approaches to word difficulty prediction are based on manual feature engineering. Using deep learning based models are largely left unexplored due to their comparatively poor performance. In this paper we explore the use of one of such in predicting the difficulty of words. We treat the problem as a binary classification problem. We train traditional machine learning models and evaluate their performance on the task. Removing dependency on frequency of previously acquired words for measuring difficulty was one of our primary aims. Then we analyze a convolutional neural network based prediction model which operates at the character level and evaluate its efficiency compared to others. Arpan Basu, Avishek Garain, Sudip Kumar Naskar |
TENCON | 3 |
| 2018 | Recognizing Textual Entailment Using Weighted Dependency Relations
Tanik Saikh, Sudip Kumar Naskar, Asif Ekbal |
CICLing (2) | 2 |
| 2018 | Says Who? Deep Learning Models for Joint Speech Recognition, Segmentation and DiarizationabstractThe field of speech recognition has seen tremendous advances in the recent past owing to the development of powerful deep learning architectures. However, the closely related fields of speech segmentation and di-arization are still primarily dominated by sophisticated variants of hierarchical clustering algorithms. We propose a powerful adaptation of the state-of-the-art Speech Recognition models for these tasks and demonstrate the effectiveness of our techniques on standard datasets. Our architectures are a combination of Bidirectional Long Short Term Memory (LSTM) Networks, Convolutional Networks, and Fully Connected Networks, trained by Gradient Descent to minimize the Cross Entropy and the Connectionist Temporal Classification (CTC) losses. We adapt the Libri Speech corpus for the task of segmentation and diarization. We obtained comparable results with respect to state-of-the-art in both tasks. Amitrajit Sarkar, Surajit Dasgupta, Sudip Kumar Naskar, Sivaji Bandyopadhyay |
ICASSP | 3 |
| 2017 | Textual Entailment Using Machine Translation Evaluation Metrics
Tanik Saikh, Sudip Kumar Naskar, Asif Ekbal, Sivaji Bandyopadhyay |
CICLing (1) | 2 |
| 2017 | Feature Selection and Class-Weight Tuning Using Genetic Algorithm for Bio-molecular Event Extraction
Amit Majumder, Asif Ekbal, Sudip Kumar Naskar |
NLDB | 3 |
| 2017 | Towards Generating Object-Oriented Programs Automatically from Natural Language Texts for Solving Mathematical Word Problems
Sourav Mandal 0001, Sudip Kumar Naskar |
NLDB | 2 |
| 2016 | Forest to String Based Statistical Machine Translation with Hybrid Word Alignments
Santanu Pal, Sudip Kumar Naskar, Josef van Genabith |
CICLing (2) | 2 |
| 2016 | Multi-Engine and Multi-Alignment Based Automatic Post-Editing and its Impact on Translation ProductivityabstractIn this paper we combine two strands of machine translation (MT) research: automatic post-editing (APE) and multi-engine (system combination) MT. APE systems learn a target-language-side second stage MT system from the data produced by human corrected output of a first stage MT system, to improve the output of the first stage MT in what is essentially a sequential MT system combination architecture. At the same time, there is a rich research literature on parallel MT system combination where the same input is fed to multiple engines and the best output is selected or smaller sections of the outputs are combined to obtain improved translation output. In the paper we show that parallel system combination in the APE stage of a sequential MT-APE combination yields substantial translation improvements both measured in terms of automatic evaluation metrics as well as in terms of productivity improvements measured in a post-editing experiment. We also show that system combination on the level of APE alignments yields further improvements. Overall our APE system yields statistically significant improvement of 5.9% relative BLEU over a strong baseline (English–Italian Google MT) and 21.76% productivity increase in a human post-editing experiment with professional translators. Santanu Pal, Sudip Kumar Naskar, Josef van Genabith |
COLING | 2 |
| 2016 | Statistical Natural Language Generation from Tabular Non-textual Data
Joy Mahapatra, Sudip Kumar Naskar, Sivaji Bandyopadhyay |
INLG | 2 |
| 2016 | CATaLog Online: Porting a Post-editing Tool to the Web
Santanu Pal, Marcos Zampieri, Sudip Kumar Naskar, Tapas Nayak, Mihaela Vela, Josef van Genabith |
LREC | 3 |
| 2015 | Textual Entailment Using Different Similarity Metrics
Tanik Saikh, Sudip Kumar Naskar, Chandan Giri, Sivaji Bandyopadhyay |
CICLing (1) | 2 |
| 2014 | Role of Paraphrases in PB-SMT
Santanu Pal, Pintu Lohar, Sudip Kumar Naskar |
CICLing (2) | 3 |
| 2014 | Word Alignment-Based Reordering of Source Chunks in PB-SMT
Santanu Pal, Sudip Kumar Naskar, Sivaji Bandyopadhyay |
LREC | 2 |
| 2013 | A Diagnostic Evaluation Approach for English to Hindi MT Using Linguistic Checkpoints and Error Rates
Renu Balyan, Sudip Kumar Naskar, Antonio Toral, Niladri Chatterjee |
CICLing (2) | 2 |
| 2013 | Meta-Evaluation of a Diagnostic Quality Metric for Machine Translation
Sudip Kumar Naskar, Antonio Toral, Federico Gaspari, Declan Groves |
MTSummit | 1 |
| 2013 | MWE Alignment in Phrase Based Statistical Machine Translation
Santanu Pal, Sudip Kumar Naskar, Sivaji Bandyopadhyay |
MTSummit | 2 |
| 2013 | A Web Application for the Diagnostic Evaluation of Machine Translation over Specific Linguistic Phenomena
Antonio Toral, Sudip Kumar Naskar, Joris Vreeke, Federico Gaspari, Declan Groves |
HLT-NAACL | 2 |
| 2012 | Translation Quality-Based Supplementary Data Selection by Incremental Update of Translation Models
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith |
COLING | 2 |
| 2012 | Domain Adaptation in SMT of User-Generated Forum Content Guided by OOV Word Reduction: Normalization and/or Supplementary Data
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith |
EAMT | 2 |
| 2011 | Combining Semantic and Syntactic Generalization in Example-Based Machine Translation
Sarah Ebling, Andy Way, Martin Volk 0001, Sudip Kumar Naskar |
EAMT | 4 |
| 2011 | A Comparative Evaluation of Research vs. Online MT Systems
Antonio Toral, Federico Gaspari, Sudip Kumar Naskar, Andy Way |
EAMT | 3 |
| 2011 | Domain Adaptation in Statistical Machine Translation of User-Forum Data using Component Level Mixture Modelling
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith |
MTSummit | 2 |
| 2011 | A Framework for Diagnostic Evaluation of MT Based on Linguistic Checkpoints
Sudip Kumar Naskar, Antonio Toral, Federico Gaspari, Andy Way |
MTSummit | 1 |
| 2011 | Integrating source-language context into phrase-based statistical machine translation
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way |
Mach. Transl. | 2 |
| 2010 | Mitigating Problems in Analogy-based EBMT with SMT and vice versa: A Case Study with Named Entity Transliteration
Sandipan Dandapat, Sara Morrissey, Sudip Kumar Naskar, Harold L. Somers |
PACLIC | 3 |
| 2009 | Using Supertags as Source Language Context in SMT
Rejwanul Haque, Sudip Kumar Naskar, Yanjun Ma, Andy Way |
EAMT | 2 |
| 2009 | Dependency Relations as Source Context in Phrase-Based SMT
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way |
PACLIC | 2 |
| 2009 | Experiments on Domain Adaptation for English--Hindi SMT
Rejwanul Haque, Sudip Kumar Naskar, Josef van Genabith, Andy Way |
PACLIC | 2 |
| 2008 | Bengali, Hindi and Telugu to English Ad-hoc Bilingual Task
Sivaji Bandyopadhyay, Tapabrata Mondal, Sudip Kumar Naskar, Asif Ekbal, Rejwanul Haque, Srinivasa Rao Godhavarthy |
IJCNLP | 3 |
| 2006 | A Modified Joint Source-Channel Model for Transliteration
Asif Ekbal, Sudip Kumar Naskar, Sivaji Bandyopadhyay |
ACL | 2 |