VLDB 2026 Research / reviewers in the wild / expert
Sujoy Sarkar
dblp:138/0044
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › named entity processing
entity discovery and linking |
0.9 | 1 | 2025 | Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking · EMNLP 2025 |
Computational social science and digital humanities › cultural analysis
literary analysis |
0.3 | 1 | 2025 | Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
entity linking models · 1.7coreference resolution · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mahānāma: A Unique Testbed for Literary Entity Discovery and LinkingabstractHigh lexical variation, ambiguous references, and long-range dependencies make entity resolution in literary texts particularly challenging.We present Mahānāma, the first large-scale dataset for end-to-end Entity Discovery and Linking (EDL) in Sanskrit, a morphologically rich and under-resourced language.Derived from the Mahābhārata, the world's longest epic, the dataset comprises over 109K named entity mentions mapped to 5.5K unique entities, and is aligned with an English knowledge base to support cross-lingual linking.The complex narrative structure of Mahānāma, coupled with extensive name variation and ambiguity, poses significant challenges to resolution systems.Our evaluation reveals that current coreference and entity linking models struggle when evaluated on the global context of the test set.These results highlight the limitations of current approaches in resolving entities within such complex discourse.Mahānāma thus provides a unique benchmark for advancing entity resolution, especially in literary domains. 1 Sujoy Sarkar, Gourav Sarkar, Manoj Balaji Jagadeeshan, Jivnesh Sandhan, Amrith Krishna, Pawan Goyal 0002 |
EMNLP | 1 |