EDBT 2026 Demo / reviewers in the wild / expert
Abhik Jana
dblp:200/8353
· DBLP profile ↗
14ranked-venue papers
8as first author
7since 2021 · last 2026
0000-0002-4485-0002ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 26% Information extraction and text analysis · 22% Planning, search and constraint satisfaction · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Spatial and temporal data management · 100% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › fairness › fair data pre-processing
dataset debiasing |
0.9 | 1 | 2025 | Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection · EMNLP 2025 |
Machine learning › Trustworthy machine learning › dataset bias
modality bias |
0.9 | 1 | 2025 | Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection · EMNLP 2025 |
Computer vision › Vision and language › multimodal dialogue
multimodal intent recognition |
0.9 | 1 | 2025 | Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection · EMNLP 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
travel planning |
0.9 | 1 | 2025 | TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis › document understanding › legal text analysis
legal text understanding |
0.6 | 1 | 2022 | LexGLUE: A Benchmark Dataset for Legal Language Understanding in English · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis › lexical semantics
compositionality prediction |
0.4 | 1 | 2019 | On the Compositionality Prediction of Noun Phrases using Poincaré Embeddings · ACL (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge engineering › knowledge integration
domain knowledge integration |
0.4 | 1 | 2019 | Incorporating Domain Knowledge into Medical NLI using Knowledge Graphs · EMNLP/IJCNLP (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph |
0.4 | 1 | 2019 | Incorporating Domain Knowledge into Medical NLI using Knowledge Graphs · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis › lexical semantics
multiword expression |
0.4 | 1 | 2019 | On the Compositionality Prediction of Noun Phrases using Poincaré Embeddings · ACL (1) 2019 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.4 | 1 | 2019 | Incorporating Domain Knowledge into Medical NLI using Knowledge Graphs · EMNLP/IJCNLP (1) 2019 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.4 | 1 | 2019 | On the Compositionality Prediction of Noun Phrases using Poincaré Embeddings · ACL (1) 2019 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2025 | Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
benchmark construction · 1.7multimodal fusion · 0.9large language model · 0.9benchmark dataset · 0.6knowledge graph embedding · 0.4hypernymy information · 0.4distributional semantic models · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Early Fusion with Contrastive Learning: A Lightweight Alternative for Multi-modal Classification
Felix Wernlein, Abhik Jana, Sandipan Sikdar |
LREC | 2 |
| 2025 | TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel PlanningabstractSoumyabrata Chaudhuri, Pranav Purkar, Ritwik Raghav, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Soumyabrata Chaudhuri, Pranav Purkar, Ritwik Raghav, Shubhojit Mallick, Abhik Jana, Shreya Ghosh 0002 |
ACL (1) | 6 |
| 2025 | Text Takes Over: A Study of Modality Bias in Multimodal Intent DetectionabstractThe rise of multimodal data, integrating text, audio, and visuals, has created new opportunities for studying multimodal tasks such as intent detection.This work investigates the effectiveness of Large Language Models (LLMs) and non-LLMs, including text-only and multimodal models, in the multimodal intent detection task.Our study reveals that Mistral-7B, a text-only LLM, outperforms most competitive multimodal models by approximately 9% on MIntRec-1 and 4% on MIntRec2.0dataset.This performance advantage comes from a strong textual bias in these datasets, where over 90% of the samples require textual input, either alone or in combination with other modalities, for correct classification.We confirm the modality bias of these datasets via human evaluation, too.Next, we propose a framework to debias the datasets, and upon debiasing, more than 70% of the samples in MIntRec-1 and more than 50% in MIntRec2.0get removed, resulting in significant performance degradation across all models, with smaller multimodal fusion models being the most affected with an accuracy drop of over 50 -60%.Further, we analyze the context-specific relevance of different modalities through empirical analysis.Our findings highlight the challenges posed by modality bias in multimodal intent datasets and emphasize the need for unbiased datasets to evaluate multimodal models effectively.We release both the code and the dataset used for this work.1 Ankan Mullick, Saransh Sharma, Abhik Jana, Pawan Goyal 0002 |
EMNLP | 3 |
| 2024 | On Zero-Shot Counterspeech Generation by LLMsabstractWith the emergence of numerous Large Language Models (LLM), the usage of such models in various Natural Language Processing (NLP) applications is increasing extensively. Counterspeech generation is one such key task where efforts are made to develop generative models by fine-tuning LLMs with hatespeech - counterspeech pairs, but none of these attempts explores the intrinsic properties of large language models in zero-shot settings. In this work, we present a comprehensive analysis of the performances of four LLMs namely GPT-2, DialoGPT, ChatGPT and FlanT5 in zero-shot settings for counterspeech generation, which is the first of its kind. For GPT-2 and DialoGPT, we further investigate the deviation in performance with respect to the sizes (small, medium, large) of the models. On the other hand, we propose three different prompting strategies for generating different types of counterspeech and analyse the impact of such strategies on the performance of the models. Our analysis shows that there is an improvement in generation quality for two datasets (17%), however the toxicity increase (25%) with increase in model size. Considering type of model, GPT-2 and FlanT5 models are significantly better in terms of counterspeech quality but also have high toxicity as compared to DialoGPT. ChatGPT are much better at generating counter speech than other models across all metrics. In terms of prompting, we find that our proposed strategies help in improving counter speech generation across all the models. Punyajoy Saha, Aalok Agrawal, Abhik Jana, Chris Biemann, Animesh Mukherjee 0001 |
LREC/COLING | 3 |
| 2022 | LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishabstractIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, Nikolaos Aletras. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II, Ion Androutsopoulos, Daniel Martin Katz, Nikolaos Aletras |
ACL (1) | 2 |
| 2022 | Network embeddings from distributional thesauri for improving static word representations
Abhik Jana, Siddhant Haldar, Pawan Goyal 0002 |
Expert Syst. Appl. | 1 |
| 2022 | Hypernymy Detection for Low-resource Languages: A Study for Hindi, Bengali, and AmharicabstractNumerous attempts for hypernymy relation (e.g., dog “is-a” animal) detection have been made for resourceful languages like English, whereas efforts made for low-resource languages are scarce primarily due to lack of gold-standard datasets and suitable distributional models. Therefore, we introduce four gold-standard datasets for hypernymy detection for each of the two languages, namely, Hindi and Bengali, and two gold-standard datasets for Amharic. Another major contribution of this work is to prepare distributional thesaurus (DT) embeddings for all three languages using three different network embedding methods (DeepWalk, role2vec, and M-NMF) for the first time on these languages and to show their utility for hypernymy detection. Posing this problem as a binary classification task, we experiment with supervised classifiers like Support Vector Machine, Random Forest, and so on, and we show that these classifiers fed with DT embeddings can obtain promising results while evaluated against proposed gold-standard datasets, specifically in an experimental setup that counteracts lexical memorization. We further incorporate DT embeddings and pre-trained fastText embeddings together using two different hybrid approaches, both of which produce an excellent performance. Additionally, we validate our methodology on gold-standard English datasets as well, where we reach a comparable performance to state-of-the-art models for hypernymy detection. Abhik Jana, Gopalakrishnan Venkatesh, Seid Muhie Yimam, Chris Biemann |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2020 | Using Distributional Thesaurus Embedding for Co-hyponymy DetectionabstractDiscriminating lexical relations among distributionally similar words has always been a challenge for natural language processing (NLP) community. In this paper, we investigate whether the network embedding of distributional thesaurus can be effectively utilized to detect co-hyponymy relations. By extensive experiments over three benchmark datasets, we show that the vector representation obtained by applying node2vec on distributional thesaurus outperforms the state-of-the-art models for binary classification of co-hyponymy vs. hypernymy, as well as co-hyponymy vs. meronymy, by huge margins. Abhik Jana, Nikhil Reddy Varimalla, Pawan Goyal 0002 |
LREC | 1 |
| 2020 | Network measures: A new paradigm towards reliable novel word sense detection
Abhik Jana, Animesh Mukherjee 0001, Pawan Goyal 0002 |
Inf. Process. Manag. | 1 |
| 2019 | On the Compositionality Prediction of Noun Phrases using Poincaré EmbeddingsabstractThe compositionality degree of multiword expressions indicates to what extent the meaning of a phrase can be derived from the meaning of its constituents and their grammatical relations. Prediction of (non)-compositionality is a task that has been frequently addressed with distributional semantic models. We introduce a novel technique to blend hierarchical information with distributional information for predicting compositionality. In particular, we use hypernymy information of the multiword and its constituents encoded in the form of the recently introduced Poincaré embeddings in addition to the distributional information to detect compositionality for noun phrases. Using a weighted average of the distributional similarity and a Poincaré similarity function, we obtain consistent and substantial, statistically significant improvement across three gold standard datasets over state-of-the-art models based on distributional information only. Unlike traditional approaches that solely use an unsupervised setting, we have also framed the problem as a supervised task, obtaining comparable improvements. Further, we publicly release our Poincaré embeddings, which are trained on the output of handcrafted lexical-syntactic patterns on a large corpus. Abhik Jana, Dmitry Puzyrev, Alexander Panchenko, Pawan Goyal 0002, Chris Biemann, Animesh Mukherjee 0001 |
ACL (1) | 1 |
| 2019 | Incorporating Domain Knowledge into Medical NLI using Knowledge GraphsabstractSoumya Sharma, Bishal Santra, Abhik Jana, Santosh Tokala, Niloy Ganguly, Pawan Goyal. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Soumya Sharma, Bishal Santra, Abhik Jana, Santosh Tokala, Niloy Ganguly, Pawan Goyal 0002 |
EMNLP/IJCNLP (1) | 3 |
| 2018 | WikiRef: Wikilinks as a route to recommending appropriate references for scientific Wikipedia pagesabstractThe exponential increase in the usage of Wikipedia as a key source of scientific knowledge among the researchers is making it absolutely necessary to metamorphose this knowledge repository into an integral and self-contained source of information for direct utilization. Unfortunately, the references which support the content of each Wikipedia entity page, are far from complete. Why are the reference section ill-formed for most Wikipedia pages? Is this section edited as frequently as the other sections of a page? Can there be appropriate surrogates that can automatically enhance the reference section? In this paper, we propose a novel two step approach – WikiRef – that (i) leverages the wikilinks present in a scientific Wikipedia target page and, thereby, (ii) recommends highly relevant references to be included in that target page appropriately and automatically borrowed from the reference section of the wikilinks. In the first step, we build a classifier to ascertain whether a wikilink is a potential source of reference or not. In the following step, we recommend references to the target page from the reference section of the wikilinks that are classified as potential sources of references in the first step. We perform an extensive evaluation of our approach on datasets from two different domains – Computer Science and Physics. For Computer Science we achieve a notably good performance with a precision@1 of 0.44 for reference recommendation as opposed to 0.38 obtained from the most competitive baseline. For the Physics dataset, we obtain a similar performance boost of 10% with respect to the most competitive baseline. Abhik Jana, Pranjal Kanojiya, Pawan Goyal 0002, Animesh Mukherjee 0001 |
COLING | 1 |
| 2018 | Network Features Based Co-hyponymy Detection
Abhik Jana, Pawan Goyal 0002 |
LREC | 1 |
| 2018 | Can Network Embedding of Distributional Thesaurus Be Combined with Word Vectors for Better Representation?abstractDistributed representations of words learned from text have proved to be successful in various natural language processing tasks in recent times. While some methods represent words as vectors computed from text using predictive model (Word2vec) or dense count based model (GloVe), others attempt to represent these in a distributional thesaurus network structure where the neighborhood of a word is a set of words having adequate context overlap. Being motivated by recent surge of research in network embedding techniques (DeepWalk, LINE, node2vec etc.), we turn a distributional thesaurus network into dense word vectors and investigate the usefulness of distributional thesaurus embedding in improving overall word representation. This is the first attempt where we show that combining the proposed word representation obtained by distributional thesaurus embedding with the state-of-the-art word representations helps in improving the performance by a significant margin when evaluated against NLP tasks like word similarity and relatedness, synonym detection, analogy detection. Additionally, we show that even without using any handcrafted lexical resources we can come up with representations having comparable performance in the word similarity and relatedness tasks compared to the representations where a lexical resource has been used. Abhik Jana, Pawan Goyal 0002 |
NAACL-HLT | 1 |