Prachuryya Kaushik

dblp:410/8559 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2026
0009-0007-9299-4426ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Information extraction and text analysis · 78% Machine translation · 17% Language models and text generation · 5%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › named entity recognition
fine-grained entity recognition
1.922026
SampurNER: Fine-Grained Named Entity Recognition Dataset for 22 Indian Languages · AAAI 2026
TAFSIL: Taxonomy Adaptable Fine-grained Entity Recognition through Distant Supervision for Indian Languages · SIGIR 2025
Natural language and speech › Information extraction and text analysis
named entity recognition
1.922026
SampurNER: Fine-Grained Named Entity Recognition Dataset for 22 Indian Languages · AAAI 2026
TAFSIL: Taxonomy Adaptable Fine-grained Entity Recognition through Distant Supervision for Indian Languages · SIGIR 2025
Natural language and speech › Machine translation
low-resource machine translation
1.012026
SampurNER: Fine-Grained Named Entity Recognition Dataset for 22 Indian Languages · AAAI 2026
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision
0.912025
TAFSIL: Taxonomy Adaptable Fine-grained Entity Recognition through Distant Supervision for Indian Languages · SIGIR 2025
Natural language and speech › Language models and text generation
multilingual language models
0.312026
SampurNER: Fine-Grained Named Entity Recognition Dataset for 22 Indian Languages · AAAI 2026
Knowledge graphs
knowledge graph construction
0.312025
TAFSIL: Taxonomy Adaptable Fine-grained Entity Recognition through Distant Supervision for Indian Languages · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

distant supervision · 2.7fuzzy matching · 1.7entity-anchored machine translation · 1.0
YearPublicationVenuePosition
2026 SampurNER: Fine-Grained Named Entity Recognition Dataset for 22 Indian Languages
abstract
We introduce SampurNER, a fine-grained named entity recognition (FgNER) dataset encompassing all 22 scheduled Indian languages spoken by more than two billion people across various countries. While manual annotation for FgNER resources is often labor-intensive and expensive, distant supervision methods have been employed as a viable solution. However, such datasets are often noisy, with entity mentions tagged with multiple types, requiring computationally intensive noise-aware models for effective FgNER. Moreover, resources for both coarse-grained and fine-grained named entity recognition tasks in Indian languages remain scarce. To address this, we propose an entity-anchored machine translation (EaMaTa) framework that leverages the largest manually annotated English FgNER dataset, FewNERD, to create a large-scale FgNER dataset in 22 languages. On average, the dataset comprises over 153k sentences, 354k entities, and 3.3M tokens in each language. The languages covered are: Assamese (as), Bengali (bn), Bodo (brx), Dogri (doi), Gujarati (gu), Hindi (hi), Kannada (kn), Kashmiri (ks), Konkani (gom), Maithili (mai), Malayalam (ml), Manipuri (mni), Marathi (mr), Nepali (ne), Odia (or), Punjabi (pa), Sanskrit (sa), Santali (sat), Sindhi (sd), Tamil (ta), Telugu (te), and Urdu (ur). Various rigorous analyses and human evaluations confirm the high quality of the dataset and demonstrate the effectiveness of the entity-anchored machine translation framework with up to 9% increase in F1-score against the current state-of-the-art. Additionally, we extend our analysis to zero-shot, multilingual, and cross-lingual settings, investigating the influence of language family and script similarity on cross-lingual FgNER performance.
Prachuryya Kaushik, Ashish Anand
AAAI1
2026 FiNERVINER: Fine-grained Named Entity Recognition for Vulnerable Languages of India's North Eastern Region
Prachuryya Kaushik, Ashish Anand
LREC1
2026 APTFiNER: Annotation Preserving Translation for Fine-grained Named Entity Recognition
Prachuryya Kaushik, Adittya Gupta, Ajanta Maurya, Gautam Sharma, V. Vijaya Saradhi, Ashish Anand
LREC1
2025 TAFSIL: Taxonomy Adaptable Fine-grained Entity Recognition through Distant Supervision for Indian Languages
abstract
Several studies have used distant supervision to create resources for fine-grained entity recognition (FgER) to mitigate the challenges of manual annotation.However, most of these methods are primarily developed for English and cannot be efficiently adapted to many other languages, including Indian languages.Moreover, the emergence of new and unseen entity types deteriorates the performance of the supervised models trained on FgER datasets with different predefined sets of entity types.This work introduces TAFSIL, a taxonomy-adaptable FgER framework to create FgER datasets in six Indian languages.The chosen languages are spoken by more than a billion speakers across various countries.TAFSIL utilizes the high interlink between the knowledge base WikiData and linked corpora Wikipedia through multi-stage heuristics and improves annotation through fuzzy match and quality sentence selection.TAFSIL enables us to create datasets of a total size of around three million samples for six languages Hindi (Hi), Marathi (Mr), Sanskrit (Sa), Tamil (Ta), Telugu (Te), and Urdu (Ur) belonging to two language families Indo-European and Dravidian.We evaluate the robustness of TAFSIL by creating various datasets in four taxonomies FIGER, OntoNotes, HAnDS, and MultiCoNER2.Our extensive experiments suggest the sound quality of the datasets as there is a relative improvement of 83% in average F1 score over zero-shot performance across various FgER state-of-the-art models.The resource is publicly available at https://huggingface.co/datasets/prachuryyaI- ITG/TAFSIL.
Prachuryya Kaushik, Shivansh Mishra, Ashish Anand
SIGIR1