EDBT 2026 Demo / reviewers in the wild / expert
Debarshi Kumar Sanyal
dblp:60/7052
· DBLP profile ↗
15ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0001-8723-5002ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 9 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DTECT: Dynamic Topic Explorer & Context TrackerabstractTo address the challenge of interpreting evolving themes in temporal text, we present DTECT (Dynamic Topic Explorer & Context Tracker), an interactive, end-to-end system for uncovering thematic dynamics. The system integrates a complete pipeline that supports data preprocessing, multiple model architectures, and dedicated metrics to analyze temporal topic quality. To enhance interpretability, DTECT features LLM-driven automatic topic labeling, trend analysis, interactive visualizations with document summarization, and a natural language chat interface. This cohesive platform empowers users to intuitively explore how topics change over time. Suman Adhya, Debarshi Kumar Sanyal |
AAAI | 2 |
| 2025 | S2WTM: Spherical Sliced-Wasserstein Autoencoder for Topic ModelingabstractModeling latent representations in a hyperspherical space has proven effective for capturing directional similarities in high-dimensional text data, benefiting topic modeling.Variational autoencoder-based neural topic models (VAE-NTMs) commonly adopt the von Mises-Fisher prior to encode hyperspherical structure.However, VAE-NTMs often suffer from posterior collapse, where the KL divergence term in the objective function highly diminishes, leading to ineffective latent representations.To mitigate this issue while modeling hyperspherical structure in the latent space, we propose the Spherical Sliced Wasserstein Autoencoder for Topic Modeling (S2WTM).S2WTM employs a prior distribution supported on the unit hypersphere and leverages the Spherical Sliced-Wasserstein distance to align the aggregated posterior distribution with the prior.Experimental results demonstrate that S2WTM outperforms state-of-the-art topic models, generating more coherent and diverse topics while improving performance on downstream tasks. github.com/AdhyaSuman/S2WTM Suman Adhya, Debarshi Kumar Sanyal |
ACL (1) | 2 |
| 2025 | NLP-QA: A Large-scale Benchmark for Informative Question Answering over Natural Language Processing Documents
Avishek Lahiri, Debarshi Kumar Sanyal, Imon Mukherjee |
CIKM | 2 |
| 2025 | TaxoAlign: Scholarly Taxonomy Generation Using Language ModelsabstractTaxonomies play a crucial role in helping researchers structure and navigate knowledge in a hierarchical manner.They also form an important part in the creation of comprehensive literature surveys.The existing approaches to automatic survey generation do not compare the structure of the generated surveys with those written by human experts.To address this gap, we present our own method for automated taxonomy creation that can bridge the gap between human-generated and automaticallycreated taxonomies.For this purpose, we create the CS-TAXOBENCH benchmark which consists of 460 taxonomies that have been extracted from human-written survey papers.We also include an additional test set of 80 taxonomies curated from conference survey papers.We propose TAXOALIGN, a threephase topic-based instruction-guided method for scholarly taxonomy generation.Additionally, we propose a stringent automated evaluation framework that measures the structural alignment and semantic coherence of automatically generated taxonomies in comparison to those created by human experts.We evaluate our method and various baselines on CS-TAXOBENCH, using both automated evaluation metrics and human evaluation studies.The results show that TAXOALIGN consistently surpasses the baselines on nearly all metrics.The code and data can be found at https: //github.com/AvishekLahiri/TaxoAlign. Avishek Lahiri, Debarshi Kumar Sanyal |
EMNLP | 3 |
| 2025 | Fine-tuned encoder models with data augmentation beat ChatGPT in agricultural named entity recognition and relation extraction
Sayan De, Debarshi Kumar Sanyal, Imon Mukherjee |
Expert Syst. Appl. | 2 |
| 2024 | GINopic: Topic Modeling with Graph Isomorphism NetworkabstractSuman Adhya, Debarshi Kumar Sanyal. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Suman Adhya, Debarshi Kumar Sanyal |
NAACL-HLT | 2 |
| 2023 | Do Neural Topic Models Really Need Dropout? Analysis of the Effect of Dropout in Topic ModelingabstractDropout is a widely used regularization trick to resolve the overfitting issue in large feedforward neural networks trained on a small dataset, which performs poorly on the held-out test subset.Although the effectiveness of this regularization trick has been extensively studied for convolutional neural networks, there is a lack of analysis of it for unsupervised models and in particular, VAE-based neural topic models.In this paper, we have analyzed the consequences of dropout in the encoder as well as in the decoder of the VAE architecture in three widely used neural topic models, namely, contextualized topic model (CTM), ProdLDA, and embedded topic model (ETM) using four publicly available datasets.We characterize the dropout effect on these models in terms of the quality and predictive performance of the generated topics. Suman Adhya, Avishek Lahiri, Debarshi Kumar Sanyal |
EACL | 3 |
| 2023 | Improving Neural Topic Models with Wasserstein Knowledge Distillation
Suman Adhya, Debarshi Kumar Sanyal |
ECIR (2) | 2 |
| 2023 | Label informed hierarchical transformers for sequential sentence classification in scientific abstractsabstractAbstract Segmenting scientific abstracts into discourse categories like background, objective, method, result, and conclusion is useful in many downstream tasks like search, recommendation and summarization. This task of classifying each sentence in the abstract into one of a given set of discourse categories is called sequential sentence classification. Existing machine learning‐based approaches to this problem consider the content of only the abstract to obtain the neural representation of each sentence, which is then labelled with a discourse category. But this ignores the semantic information offered by the discourse labels themselves. In this paper, we propose LIHT, Label Informed Hierarchical Transformers – a method for sequential sentence classification that explicitly and hierarchically exploits the semantic information in the labels to learn label‐aware neural sentence representations. The hierarchical model helps to capture not only the fine‐grained interactions between the discourse labels and the words in the abstract at the sentence level but also the potential dependencies that may exist in the label sequence. Thus, LIHT generates label‐aware contextual sentence representations that are then labelled with a conditional random field. We evaluate LIHT on three publicly available datasets, namely, PUBMED‐RCT, NICTA‐PIBOSO and CSAbstract. The incremental gain in F1‐score in all the three cases over the respective state‐of‐the‐art approaches is around . Though the gains are modest, LIHT establishes a new performance benchmark for this task and is a novel technique of independent interest. We also perform an ablation study to identify the contribution of each component of LIHT in the observed performance, and a case study to visualize the roles of the different components of our model. T. Y. S. S. Santosh, Sai Saketh Aluru, Anoop Vallabhajosyula, Debarshi Kumar Sanyal, Partha Pratim Das 0001 |
Expert Syst. J. Knowl. Eng. | 4 |
| 2022 | Locating Code Omission Error due to Incorrect Polymorphic Method CallabstractDynamic program slicing methods are widely used for debugging because many statements can be ignored in the process of localizing a bug. A dynamic program slice for a variable contains only those statements that influenced this variable. One limitation of dynamic slicing-based techniques is that they cannot capture execution omission errors, which may cause the execution of certain critical statements in a program to be omitted and thus result in failures. In this paper, we propose a solution to locate execution omission errors by using dynamic slices for a variable in the presence of methods that are virtual in Java programs. Note that any method which is not a static, private or final method can be considered a virtual method in a Java program. Due to the assignment of objects of wrong derived classes to base class reference, some different versions of polymorphic methods can be called which may cause failure. We designed a slicing method called polymorphic relevant slice which can be used to force the execution of the omitted code by initializing objects of all possible classes derived from the same base class with the available data, assigning them to base class reference one-by-one and switch execution for all alternative virtual functions to check if any of them meets specification. We have used a system dependence graph to identify all the derived class objects which could have been assigned to the base class reference. Sudakshina Dutta, Debarshi Kumar Sanyal |
ICST | 2 |
| 2021 | HiCoVA: Hierarchical Conditional Variational Autoencoder for Keyphrase GenerationabstractThe task of keyphrase generation, unlike extraction, aims to generate the phrases which succinctly capture the key information of the source text, that are even absent in the document (i.e., do not match any contiguous sub-sequence of source text). Despite the significant progress achieved by sequence-to-sequence (seq2seq) models in modelling such high entropy task, they are limited by their deterministic modelling capability which limits the generation of a diverse set of keyphrases. To address the above limitation, in this paper, we propose to incorporate Conditional Variational Autoencoder (CoVA) into seq2seq models for its ability to represent a set of keyphrases as a probabilistic distribution which improves the diversity of the generated keyphrases. We model the probabilistic distribution using a hierarchical latent structure where a global latent variable tries to model the diversity among the keyphrases and local latent variables control the generation of each keyphrase to make them coherent. Experimental results on four benchmark datasets of research papers demonstrate the effectiveness of our proposed approach in achieving a large improvement in diversity along with modest gains in quality with respect to previous models. T. Y. S. S. Santosh, Nikhil Reddy Varimalla, Anoop Vallabhajosyula, Debarshi Kumar Sanyal, Partha Pratim Das 0001 |
CIKM | 4 |
| 2021 | Gazetteer-Guided Keyphrase Generation from Research Papers
T. Y. S. S. Santosh, Debarshi Kumar Sanyal, Plaban Kumar Bhowmick, Partha Pratim Das 0001 |
PAKDD (1) | 2 |
| 2020 | SaSAKE: Syntax and Semantics Aware Keyphrase Extraction from Research PapersabstractKeyphrases in a research paper succinctly capture the primary content of the paper and also assist in indexing the paper at a concept level. Given the huge rate at which scientific papers are published today, it is important to have effective ways of automatically extracting keyphrases from a research paper. In this paper, we present a novel method, Syntax and Semantics Aware Keyphrase Extraction (SaSAKE), to extract keyphrases from research papers. It uses a transformer architecture, stacking up sentence encoders to incorporate sequential information, and graph encoders to incorporate syntactic and semantic dependency graph information. Incorporation of these dependency graphs helps to alleviate long-range dependency problems and identify the boundaries of multi-word keyphrases effectively. Experimental results on three benchmark datasets show that our proposed method SaSAKE achieves state-of-the-art performance in keyphrase extraction from scientific papers. T. Y. S. S. Santosh, Debarshi Kumar Sanyal, Plaban Kumar Bhowmick, Partha Pratim Das 0001 |
COLING | 2 |
| 2020 | DAKE: Document-Level Attention for Keyphrase Extraction
T. Y. S. S. Santosh, Debarshi Kumar Sanyal, Plaban Kumar Bhowmick, Partha Pratim Das 0001 |
ECIR (2) | 2 |
| 2017 | Dimensionality reduction of EEG signal using Fuzzy Discernibility MatrixabstractHigh dimensionality of feature space is a problem in supervised machine learning. Redundant or superfluous features either slow down the training process or dilute the quality of classification. Many methods are available in literature for dimensionality reduction. Earlier studies explored a discernibility matrix (DM) based reduct calculation for dimensionality reduction. Discernibility matrix works only on discrete values. But most real-world datasets are continuous in nature. Use of traditional discernibility matrix approach inevitably incurs information loss due to discretization. In this paper, we propose a fuzzified adaptation of discernibility matrix with four variants of dissimilarity measure to deal with continuous data. The proposed algorithm has been applied on EEG dataset-III from BCI competition-II. The reduced dataset is then classified using Support Vector Machine (SVM). The performance of the proposed Fuzzy Discernibility Matrix (FDM) variants are compared with original discernibility matrix based method and Principal Component Analysis (PCA). In our empirical study, the proposed method outperforms the other two methods, thus suggesting that it is competitive with them. Rajdeep Chatterjee, Tathagata Bandyopadhyay, Debarshi Kumar Sanyal, Dibyajyoti Guha |
HSI | 3 |