Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Amrith Krishna

dblp:160/4306 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Information extraction and text analysis · 65% Representation and self-supervised learning · 23% Language models and text generation · 12%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
1.222024
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024
Keep it Surprisingly Simple: A Simple First Order Graph Based Parsing Model for Joint Morphosyntactic Parsing in Sanskrit · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis › named entity processing
entity discovery and linking
0.912025
Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking · EMNLP 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024
Machine learning › Representation and self-supervised learning › contrastive learning
self-supervised contrastive learning
0.812024
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.812024
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024
Natural language and speech › Language models and text generation › text generation › surface realization
linearization
0.412019
Poetry to Prose Conversion in Sanskrit as a Linearisation Task: A Case for Low-Resource Languages · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › syntactic parsing
word ordering
0.412019
Poetry to Prose Conversion in Sanskrit as a Linearisation Task: A Case for Low-Resource Languages · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological tagging
0.312018
Free as in Free Word Order: An Energy Based Model for Word Segmentation and Morphological Tagging in Sanskrit · EMNLP 2018
Natural language and speech › Information extraction and text analysis
word segmentation
0.312018
Free as in Free Word Order: An Energy Based Model for Word Segmentation and Morphological Tagging in Sanskrit · EMNLP 2018
Computational social science and digital humanities › cultural analysis
literary analysis
0.312025
Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking · EMNLP 2025
Natural language and speech › Language models and text generation
multilingual language models
0.212024
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

entity linking models · 1.7coreference resolution · 1.7graph pruning · 0.9exact search · 0.9energy-based model · 0.8position encoding removal · 0.8data augmentation · 0.8energy-based models · 0.4token embedding · 0.4seq2seq · 0.4pre-training · 0.4
YearPublicationVenuePosition
2026 Linguistically informed automatic speech recognition in Sanskrit
Rishabh Kumar, Devaraj Adiga, Rishav Ranjan, Amrith Krishna, Ganesh Ramakrishnan, Pawan Goyal 0002, Preethi Jyothi
Comput. Speech Lang.4
2025 Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking
abstract
High lexical variation, ambiguous references, and long-range dependencies make entity resolution in literary texts particularly challenging.We present Mahānāma, the first large-scale dataset for end-to-end Entity Discovery and Linking (EDL) in Sanskrit, a morphologically rich and under-resourced language.Derived from the Mahābhārata, the world's longest epic, the dataset comprises over 109K named entity mentions mapped to 5.5K unique entities, and is aligned with an English knowledge base to support cross-lingual linking.The complex narrative structure of Mahānāma, coupled with extensive name variation and ambiguity, poses significant challenges to resolution systems.Our evaluation reveals that current coreference and entity linking models struggle when evaluated on the global context of the test set.These results highlight the limitations of current approaches in resolving entities within such complex discourse.Mahānāma thus provides a unique benchmark for advancing entity resolution, especially in literary domains. 1
Sujoy Sarkar, Gourav Sarkar, Manoj Balaji Jagadeeshan, Jivnesh Sandhan, Amrith Krishna, Pawan Goyal 0002
EMNLP5
2024 Samayik: A Benchmark and Dataset for English-Sanskrit Translation
abstract
We release Saamayik, a dataset of around 53,000 parallel English-Sanskrit sentences, written in contemporary prose. Sanskrit is a classical language still in sustenance and has a rich documented heritage. However, due to the limited availability of digitized content, it still remains a low-resource language. Existing Sanskrit corpora, whether monolingual or bilingual, have predominantly focused on poetry and offer limited coverage of contemporary written materials. Saamayik is curated from a diverse range of domains, including language instruction material, textual teaching pedagogy, and online tutorials, among others. It stands out as a unique resource that specifically caters to the contemporary usage of Sanskrit, with a primary emphasis on prose writing. Translation models trained on our dataset demonstrate statistically significant improvements when translating out-of-domain contemporary corpora, outperforming models trained on older classical-era poetry datasets. Finally, we also release benchmark models by adapting four multilingual pre-trained models, three of them have not been previously exposed to Sanskrit for translating between English and Sanskrit while one of them is multi-lingual pre-trained translation model including English and Sanskrit. The dataset and source code can be found at https://github.com/ayushbits/saamayik.
Ayush Maheshwari, Ashim Gupta, Amrith Krishna, Atul Kumar Singh, Ganesh Ramakrishnan, Anil Kumar Gourishetty, Jitin Singla
LREC/COLING3
2024 CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages
abstract
Neural dependency parsing has achieved remarkable performance for Low-Resource Morphologically-Rich languages.It has also been well-studied that Morphologically-Rich languages exhibit relatively free-word-order.This prompts a fundamental investigation: Is there a way to enhance dependency parsing performance, making the model robust to word order variations utilizing the relatively freeword-order nature of Morphologically-Rich languages?In this work, we examine the robustness of graph-based parsing architectures on 7 relatively free-word-order languages.We focus on scrutinizing essential modifications such as data augmentation and the removal of position encoding required to adapt these architectures accordingly.To this end, we propose a contrastive self-supervised learning method to make the model robust to word order variations.Furthermore, our proposed modification demonstrates a substantial average gain of 3.03/2.95points in 7 relatively free-word-order languages, as measured by the UAS/LAS Score metric when compared to the best performing baseline.
Pretam Ray, Jivnesh Sandhan, Amrith Krishna, Pawan Goyal 0002
EMNLP3
2022 Does Meta-learning Help mBERT for Few-shot Question Generation in a Cross-lingual Transfer Setting for Indic Languages?
abstract
Few-shot Question Generation (QG) is an important and challenging problem in the Natural Language Generation (NLG) domain. Multilingual BERT (mBERT) has been successfully used in various Natural Language Understanding (NLU) applications. However, the question of how to utilize mBERT for few-shot QG, possibly with cross-lingual transfer, remains. In this paper, we try to explore how mBERT performs in few-shot QG (cross-lingual transfer) and also whether applying meta-learning on mBERT further improves the results. In our setting, we consider mBERT as the base model and fine-tune it using a seq-to-seq language modeling framework in a cross-lingual setting. Further, we apply the model agnostic meta-learning approach to our base model. We evaluate our model for two low-resource Indian languages, Bengali and Telugu, using the TyDi QA dataset. The proposed approach consistently improves the performance of the base model in few-shot settings and even works better than some heavily parameterized models. Human evaluation also confirms the effectiveness of our approach.
Aniruddha Roy, Rupak Kumar Thakur, Isha Sharma, Ashim Gupta, Amrith Krishna, Sudeshna Sarkar, Pawan Goyal 0002
COLING5
2022 Linguistically Informed Post-processing for ASR Error correction in Sanskrit
Rishabh Kumar, Devaraj Adiga, Rishav Ranjan, Amrith Krishna, Ganesh Ramakrishnan, Pawan Goyal 0002, Preethi Jyothi
INTERSPEECH4
2022 ProoFVer: Natural Logic Theorem Proving for Fact Verification
abstract
Abstract Fact verification systems typically rely on neural network classifiers for veracity prediction, which lack explainability. This paper proposes ProoFVer, which uses a seq2seq model to generate natural logic-based inferences as proofs. These proofs consist of lexical mutations between spans in the claim and the evidence retrieved, each marked with a natural logic operator. Claim veracity is determined solely based on the sequence of these operators. Hence, these proofs are faithful explanations, and this makes ProoFVer faithful by construction. Currently, ProoFVer has the highest label accuracy and the second best score in the FEVER leaderboard. Furthermore, it improves by 13.21% points over the next best model on a dataset with counterfactual instances, demonstrating its robustness. As explanations, the proofs show better overlap with human rationales than attention-based highlights and the proofs help humans predict model decisions correctly more often than using the evidence directly.1
Amrith Krishna, Sebastian Riedel 0001, Andreas Vlachos 0001
Trans. Assoc. Comput. Linguistics1
2020 Keep it Surprisingly Simple: A Simple First Order Graph Based Parsing Model for Joint Morphosyntactic Parsing in Sanskrit
abstract
Morphologically rich languages seem to benefit from joint processing of morphology and syntax, as compared to pipeline architectures. We propose a graph-based model for joint morphological parsing and dependency parsing in Sanskrit. Here, we extend the Energy based model framework (Krishna et al., 2020), proposed for several structured prediction tasks in Sanskrit, in 2 simple yet significant ways. First, the framework’s default input graph generation method is modified to generate a multigraph, which enables the use of an exact search inference. Second, we prune the input search space using a linguistically motivated approach, rooted in the traditional grammatical analysis of Sanskrit. Our experiments show that the morphological parsing from our joint model outperforms standalone morphological parsers. We report state of the art results in morphological parsing, and in dependency parsing, both in standalone (with gold morphological tags) and joint morphosyntactic parsing setting.
Amrith Krishna, Ashim Gupta, Deepak Garasangi, Pavankumar Satuluri, Pawan Goyal 0002
EMNLP (1)1
2020 SHR++: An Interface for Morpho-syntactic Annotation of Sanskrit Corpora
abstract
We propose a web-based annotation framework, SHR++, for morpho-syntactic annotation of corpora in Sanskrit. SHR++ is designed to generate annotations for the word-segmentation, morphological parsing and dependency analysis tasks in Sanskrit. It incorporates analyses and predictions from various tools designed for processing texts in Sanskrit, and utilise them to ease the cognitive load of the human annotators. Specifically, SHR++ uses Sanskrit Heritage Reader, a lexicon driven shallow parser for enumerating all the phonetically and lexically valid word splits along with their morphological analyses for a given string. This would help the annotators in choosing the solutions, rather than performing the segmentations by themselves. Further, predictions from a word segmentation tool are added as suggestions that can aid the human annotators in their decision making. Our evaluation shows that enabling this segmentation suggestion component reduces the annotation time by 20.15 %. SHR++ can be accessed online at http://vidhyut97.pythonanywhere.com/ and the codebase, for the independent deployment of the system elsewhere, is hosted at https://github.com/iamdsc/smart-sanskrit-annotator.
Amrith Krishna, Shiv Vidhyut, Dilpreet Chawla, Sruti Sambhavi, Pawan Goyal 0002
LREC1
2020 A Graph-Based Framework for Structured Prediction Tasks in Sanskrit
abstract
We propose a framework using energy-based models for multiple structured prediction tasks in Sanskrit. Ours is an arc-factored model, similar to the graph-based parsing approaches, and we consider the tasks of word segmentation, morphological parsing, dependency parsing, syntactic linearization, and prosodification, a “prosody-level” task we introduce in this work. Ours is a search-based structured prediction framework, which expects a graph as input, where relevant linguistic information is encoded in the nodes, and the edges are then used to indicate the association between these nodes. Typically, the state-of-the-art models for morphosyntactic tasks in morphologically rich languages still rely on hand-crafted features for their performance. But here, we automate the learning of the feature function. The feature function so learned, along with the search space we construct, encode relevant linguistic information for the tasks we consider. This enables us to substantially reduce the training data requirements to as low as 10%, as compared to the data requirements for the neural state-of-the-art models. Our experiments in Czech and Sanskrit show the language-agnostic nature of the framework, where we train highly competitive models for both the languages. Moreover, our framework enables us to incorporate language-specific constraints to prune the search space and to filter the candidates during inference. We obtain significant improvements in morphosyntactic tasks for Sanskrit by incorporating language-specific constraints into the model. In all the tasks we discuss for Sanskrit, we either achieve state-of-the-art results or ours is the only data-driven solution for those tasks.
Amrith Krishna, Bishal Santra, Ashim Gupta, Pavankumar Satuluri, Pawan Goyal 0002
Comput. Linguistics1
2019 Poetry to Prose Conversion in Sanskrit as a Linearisation Task: A Case for Low-Resource Languages
abstract
The word ordering in a Sanskrit verse is often not aligned with its corresponding prose order.Conversion of the verse to its corresponding prose helps in better comprehension of the construction.Owing to the resource constraints, we formulate this task as a word ordering (linearisation) task.In doing so, we completely ignore the word arrangement at the verse side.kāvya guru, the approach we propose, essentially consists of a pipeline of two pretraining steps followed by a seq2seq model.The first pretraining step learns task specific token embeddings from pretrained embeddings.In the next step, we generate multiple hypotheses for possible word arrangements of the input (Wang et al., 2018).We then use them as inputs to a neural seq2seq model for the final prediction.We empirically show that the hypotheses generated by our pretraining step result in predictions that consistently outperform predictions based on the original order in the verse.Overall, kāvya guru outperforms current state of the art models in linearisation for the poetry to prose conversion task in Sanskrit.
Amrith Krishna, Vishnu Dutt Sharma, Bishal Santra, Aishik Chakraborty, Pavankumar Satuluri, Pawan Goyal 0002
ACL (1)1
2018 Upcycle Your OCR: Reusing OCRs for Post-OCR Text Correction in Romanised Sanskrit
abstract
We propose a post-OCR text correction approach for digitising texts in Romanised Sanskrit.Owing to the lack of resources our approach uses OCR models trained for other languages written in Roman.Currently, there exists no dataset available for Romanised Sanskrit OCR.So, we bootstrap a dataset of 430 images, scanned in two different settings and their corresponding ground truth.For training, we synthetically generate training images for both the settings.We find that the use of copying mechanism (Gu et al., 2016) yields a percentage increase of 7.69 in Character Recognition Rate (CRR) than the current state of the art model in solving monotone sequence-tosequence tasks (Schnober et al., 2016).We find that our system is robust in combating OCR-prone errors, as it obtains a CRR of 87.01%from an OCR output with CRR of 35.76% for one of the dataset settings.A human judgement survey performed on the models shows that our proposed model results in predictions which are faster to comprehend and faster to improve for a human than the other systems 1 .
Amrith Krishna, Bodhisattwa Prasad Majumder, Rajesh Shreedhar Bhat, Pawan Goyal 0002
CoNLL1
2018 Free as in Free Word Order: An Energy Based Model for Word Segmentation and Morphological Tagging in Sanskrit
abstract
Amrith Krishna, Bishal Santra, Sasi Prasanth Bandaru, Gaurav Sahu, Vishnu Dutt Sharma, Pavankumar Satuluri, Pawan Goyal. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Amrith Krishna, Bishal Santra, Sasi Prasanth Bandaru, Gaurav Sahu, Vishnu Dutt Sharma, Pavankumar Satuluri, Pawan Goyal 0002
EMNLP1
2018 Building a Word Segmenter for Sanskrit Overnight
Vikas Reddy, Amrith Krishna, Vishnu Dutt Sharma, Prateek Gupta, Vineeth M. R, Pawan Goyal 0002
LREC2
2016 Word Segmentation in Sanskrit Using Path Constrained Random Walks
abstract
In Sanskrit, the phonemes at the word boundaries undergo changes to form new phonemes through a process called as sandhi. A fused sentence can be segmented into multiple possible segmentations. We propose a word segmentation approach that predicts the most semantically valid segmentation for a given sentence. We treat the problem as a query expansion problem and use the path-constrained random walks framework to predict the correct segments.
Amrith Krishna, Bishal Santra, Pavankumar Satuluri, Sasi Prasanth Bandaru, Bhumi Faldu, Yajuvendra Singh, Pawan Goyal 0002
COLING1
2016 FeRoSA: A Faceted Recommendation System for Scientific Articles
Tanmoy Chakraborty 0002, Amrith Krishna, Mayank Singh 0001, Niloy Ganguly, Pawan Goyal 0002, Animesh Mukherjee 0001
PAKDD (2)2