Shubhanker Banerjee

dblp:276/7121 · DBLP profile ↗
← Back
4ranked-venue papers in the field
3as first author
4since 2021 · last 2025
0000-0002-3969-5183ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 3 (2 first)Information Retrieval & Web Search · 1 (1 first)
YearPublicationVenuePosition
2025 Benchmarking Hindi Term Extraction in Education: A Dataset and Analysis
abstract
This paper introduces the HTEC HindiTerm Extraction Dataset 2.0, a resourcedesigned to support terminology extractionand classification tasks within the education domain. HTEC 2.0 has been developed with the objective of providing a high-quality benchmark dataset for the evaluation of term recognition and classification methodologies in Hindi educationaldiscourse. The dataset consists of 97 documents sourced from Hindi Wikipedia, covering a diverse range of topics relevant tothe education sector. Within these documents, 1,702 terms have been manuallyannotated where each term is defined as asingle-word or multi-word expression thatconveys a domain-specific meaning. Theannotated terms in HTEC 2.0 are systematically categorized into seven distinct classes.Furthermore, this paper outlines the development of annotation guidelines, detailingthe criteria used to determine term boundaries and category assignments. By offeringa structured dataset with clearly definedterm classifications, HTEC 2.0 serves as avaluable resource for researchers workingon terminology extraction, domain-specificnamed entity recognition, and text classification in Hindi.
Shubhanker Banerjee, Bharathi Raja Chakravarthi, John P. McCrae
LDK1
2025 Cuaċ: Fast and Small Universal Representations of Corpora
abstract
The increasing size and diversity of corpora in natural language processing requires highly efficient processing frameworks. Building on the universal corpus format, Teanga, we present Cuaċ, a format for the compact representation of corpora. We describe this methodology based on short-string compression and indexing techniques and show that the files created with this methodology are similar to compressed human-readable serializations and can be further compressed using lossless compression. We also show that this introduces no computational penalty on the time to process files. This methodology aims to speed up natural language processing pipelines and is the basis for a fast database system for corpora.
John P. McCrae, Bernardo Stearns, Alamgir Munir Qazi, Shubhanker Banerjee, Atul Kr. Ojha
LDK4
2024 Large Language Models for Few-Shot Automatic Term Extraction
Shubhanker Banerjee, Bharathi Raja Chakravarthi, John P. McCrae
NLDB (1)1
2023 MG2P: An Empirical Study Of Multilingual Training for Manx G2P
Shubhanker Banerjee, Bharathi Raja Chakravarthi, John P. McCrae
LDK1