VLDB 2026 Research / reviewers in the wild / expert
Jivnesh Sandhan
dblp:213/9499
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-4515-4295ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 58% Representation and self-supervised learning · 37% Language models and text generation · 6% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › named entity processing
entity discovery and linking |
0.9 | 1 | 2025 | Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking · EMNLP 2025 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing |
0.8 | 1 | 2024 | CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
self-supervised contrastive learning |
0.8 | 1 | 2024 | CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.8 | 1 | 2024 | CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024 |
Computational social science and digital humanities › cultural analysis
literary analysis |
0.3 | 1 | 2025 | Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking · EMNLP 2025 |
Natural language and speech › Language models and text generation
multilingual language models |
0.2 | 1 | 2024 | CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
entity linking models · 1.7coreference resolution · 1.7position encoding removal · 0.8data augmentation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mahānāma: A Unique Testbed for Literary Entity Discovery and LinkingabstractHigh lexical variation, ambiguous references, and long-range dependencies make entity resolution in literary texts particularly challenging.We present Mahānāma, the first large-scale dataset for end-to-end Entity Discovery and Linking (EDL) in Sanskrit, a morphologically rich and under-resourced language.Derived from the Mahābhārata, the world's longest epic, the dataset comprises over 109K named entity mentions mapped to 5.5K unique entities, and is aligned with an English knowledge base to support cross-lingual linking.The complex narrative structure of Mahānāma, coupled with extensive name variation and ambiguity, poses significant challenges to resolution systems.Our evaluation reveals that current coreference and entity linking models struggle when evaluated on the global context of the test set.These results highlight the limitations of current approaches in resolving entities within such complex discourse.Mahānāma thus provides a unique benchmark for advancing entity resolution, especially in literary domains. 1 Sujoy Sarkar, Gourav Sarkar, Manoj Balaji Jagadeeshan, Jivnesh Sandhan, Amrith Krishna, Pawan Goyal 0002 |
EMNLP | 4 |
| 2025 | Tagsim: Topic-Informed Attention Guided Similarity Metric for Image Caption ComparisonabstractExisting image caption evaluation metrics, such as BLEU, ROUGE, and CIDEr primarily rely on high-level similarities like n-gram matching. Here, we propose TAGSim, a novel metric that automatically incorporates topics or concepts as the caption’s topic contains the main summary of the image. TAGSim creates topic-weighted latent representations, thereby acquiring semantics by integrating both con-text and topic information. It leverages a regression model to create an attention-based novel similarity metric, while internally building caption representations. TAGSim moves be-yond lexical overlap by focusing on meaningful semantic relationships, better aligning captions with the core image topic. It handles paraphrasing and diverse expressions, ensuring a more nuanced and reliable evaluation across languages and styles. It outperforms the strong similarity baseline by an average 0.72 points (SCI evaluation index) across 3 datasets. It is language agnostic, empirically established that it is a pseudometric, and correlates well with standard caption evaluation metrics. Our code and datasets will be publicly available. Vipul Chanchlani, Vishal Himmatsinghka, Ayush Himmatsinghka, Jivnesh Sandhan, Tushar Sandhan |
ICIP | 4 |
| 2024 | CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource LanguagesabstractNeural dependency parsing has achieved remarkable performance for Low-Resource Morphologically-Rich languages.It has also been well-studied that Morphologically-Rich languages exhibit relatively free-word-order.This prompts a fundamental investigation: Is there a way to enhance dependency parsing performance, making the model robust to word order variations utilizing the relatively freeword-order nature of Morphologically-Rich languages?In this work, we examine the robustness of graph-based parsing architectures on 7 relatively free-word-order languages.We focus on scrutinizing essential modifications such as data augmentation and the removal of position encoding required to adapt these architectures accordingly.To this end, we propose a contrastive self-supervised learning method to make the model robust to word order variations.Furthermore, our proposed modification demonstrates a substantial average gain of 3.03/2.95points in 7 relatively free-word-order languages, as measured by the UAS/LAS Score metric when compared to the best performing baseline. Pretam Ray, Jivnesh Sandhan, Amrith Krishna, Pawan Goyal 0002 |
EMNLP | 2 |
| 2023 | Systematic Investigation of Strategies Tailored for Low-Resource Settings for Low-Resource Dependency ParsingabstractIn this work, we focus on low-resource dependency parsing for multiple languages.Several strategies are tailored to enhance performance in low-resource scenarios.While these are well-known to the community, it is not trivial to select the best-performing combination of these strategies for a low-resource language that we are interested in, and not much attention has been given to measuring the efficacy of these strategies.We experiment with 5 lowresource strategies for our ensembled approach on 7 Universal Dependency (UD) low-resource languages.Our exhaustive experimentation on these languages supports the effective improvements for languages not covered in pretrained models.We show a successful application of the ensembled system on a truly low-resource language Sanskrit. 1 Jivnesh Sandhan, Laxmidhar Behera, Pawan Goyal 0002 |
EACL | 1 |
| 2022 | A Novel Multi-Task Learning Approach for Context-Sensitive Compound Type Identification in SanskritabstractThe phenomenon of compounding is ubiquitous in Sanskrit. It serves for achieving brevity in expressing thoughts, while simultaneously enriching the lexical and structural formation of the language. In this work, we focus on the Sanskrit Compound Type Identification (SaCTI) task, where we consider the problem of identifying semantic relations between the components of a compound word. Earlier approaches solely rely on the lexical information obtained from the components and ignore the most crucial contextual and syntactic information useful for SaCTI. However, the SaCTI task is challenging primarily due to the implicitly encoded context-sensitive semantic relation between the compound components. Thus, we propose a novel multi-task learning architecture which incorporates the contextual information and enriches the complementary syntactic information using morphological tagging and dependency parsing as two auxiliary tasks. Experiments on the benchmark datasets for SaCTI show 6.1 points (Accuracy) and 7.7 points (F1-score) absolute gain compared to the state-of-the-art system. Further, our multi-lingual experiments demonstrate the efficacy of the proposed architecture in English and Marathi languages. Jivnesh Sandhan, Hrishikesh Terdalkar, Tushar Sandhan, Suvendu Samanta, Laxmidhar Behera, Pawan Goyal 0002 |
COLING | 1 |