Bal Krishna Bal

dblp:56/8153 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-0345-8917ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
cross-language information retrieval
1.012026
Towards Building a Standard Benchmark for Low-Resource Nepali Information Retrieval · SIGIR 2026
Information retrieval
evaluation
1.012026
Towards Building a Standard Benchmark for Low-Resource Nepali Information Retrieval · SIGIR 2026
Information retrieval › evaluation
test collection
1.012026
Towards Building a Standard Benchmark for Low-Resource Nepali Information Retrieval · SIGIR 2026
Information retrieval
retrieval models
0.312026
Towards Building a Standard Benchmark for Low-Resource Nepali Information Retrieval · SIGIR 2026
YearPublicationVenuePosition
2026 English-Nepali-Tamang: A Trilingual Parallel Corpus and Benchmark for Low-Resource Machine Translation
abstract
This article describes the research project aimed at developing a Trilingual Machine translation System for English, Nepali, and Tamang language pairs. This project is expected to address knowledge and communication gaps caused by language barriers and mitigate disparities in the availability of information and knowledge sources in Tamang and Nepali.
Praveen Acharya, Rupak Raj Ghimire, Prakash Poudyal, Balaram Prasain, Bal Krishna Bal
EAMT (2)5
2026 Parallel Corpus Development Toolkit (PCDT): A Web-Based Platform for Multilingual Parallel Data Creation
abstract
This paper presents PCDT, a web-based platform for collecting sentence-aligned parallel corpora through a community-driven approach to support machine translation for under-resourced languages. The tool decentralizes the translation task to the target community and subsequently reviewed by language experts.
Praveen Acharya, Rupak Raj Ghimire, Bipesh Subedi, Prakash Poudyal, Balaram Prasain, Bal Krishna Bal
EAMT (2)6
2026 NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
Rupak Raj Ghimire, Bipesh Subedi, Balaram Prasain, Prakash Poudyal, Praveen Acharya, Nischal Karki, Rupak Tiwari, Rishikesh Kumar Sharma, Jenny Poudel, Bal Krishna Bal
LREC10
2026 Nepal Script Text Recognition from Ancient Artifacts: Challenges and Opportunities
Swornim Nakarmi, Sarin Sthapit, Sahil Ratna Tuladhar, Arya Shakya, Bal Krishna Bal, Rajani Chulyadyo
LREC5
2026 Nepali Lemmatization with Multilingual Transformers: Intrinsic and Extrinsic Evaluation in a Low-Resource Setting
Sunil Regmi, Sundeep Dawadi, Bal Krishna Bal
LREC3
2026 Towards Building a Standard Benchmark for Low-Resource Nepali Information Retrieval
abstract
Despite being spoken by over 26 million people, Nepali lacks a standardized Information Retrieval (IR) benchmark, preventing systematic evaluation of retrieval models in this low-resource setting. While progress has been made in several Nepali NLP tasks, no publicly available test collection with queries and relevance judgments exists for ad-hoc retrieval task. We propose the creation of a first standardized Nepali IR benchmark supporting monolingual (Nepali queries \rightarrow Nepali documents), cross-lingual (English queries \rightarrow Nepali documents), and code-mixed (Nepali-English queries \rightarrow Nepali documents) retrieval. We argue that Nepali provides a compelling testbed for studying morphology-aware retrieval, code-mixed query processing, and multilingual embedding transfer.
Praveen Acharya, Bal Krishna Bal
SIGIR2
2024 Exploring the Potential of Large Language Models (LLMs) for Low-resource Languages: A Study on Named-Entity Recognition (NER) and Part-Of-Speech (POS) Tagging for Nepali Language
abstract
Large Language Models (LLMs) have made significant advancements in Natural Language Processing (NLP) by excelling in various NLP tasks. This study specifically focuses on evaluating the performance of LLMs for Named Entity Recognition (NER) and Part-of-Speech (POS) tagging for a low-resource language, Nepali. The aim is to study the effectiveness of these models for languages with limited resources by conducting experiments involving various parameters and fine-tuning and evaluating two datasets namely, ILPRL and EBIQUITY. In this work, we have experimented with eight LLMs for Nepali NER and POS tagging. While some prior works utilized larger datasets than ours, our contribution lies in presenting a comprehensive analysis of multiple LLMs in a unified setting. The findings indicate that NepBERTa, trained solely in the Nepali language, demonstrated the highest performance with F1-scores of 0.76 and 0.90 in ILPRL dataset. Similarly, it achieved 0.79 and 0.97 in EBIQUITY dataset for NER and POS respectively. This study not only highlights the potential of LLMs in performing classification tasks for low-resource languages but also compares their performance with that of alternative approaches deployed for the tasks.
Bipesh Subedi, Sunil Regmi, Bal Krishna Bal, Praveen Acharya
LREC/COLING3
2023 Large Vocabulary Continous Speech Recognition for Nepali Language using CNN and Transformer
Shishir Paudel, Bal Krishna Bal, Dhiraj Shrestha
LDK2
2020 Aspect Based Abusive Sentiment Detection in Nepali Social Media Texts
abstract
With the increase in internet access and the ease of writing comments in the Nepali language, fine-grained sentiment analysis of social media comments is becoming more and more pertinent. There are a number of benchmarked datasets for high-resource languages (English, French, and German) in specific domains like restaurants, hotels or electronic goods but not in low-resource languages like Nepali. In this paper, we present our work to create a dataset for the targeted aspect-based sentiment analysis in the social media domain, set up a dataset benchmark and evaluate using various machine learning models. The dataset comprises of code-mixed and code-switched comments extracted from Nepali YouTube videos. We present convincing baselines using a multilingual BERT model for the Aspect Term Extraction task and BiLSTM model for the Sentiment Classification Task achieving 57.978% and 81.60% F1 score respectively.
Oyesh Mann Singh, Sandesh Timilsina, Bal Krishna Bal, Anupam Joshi
ASONAM3
2014 Issues in Encoding the Writing of Nepal's Languages
Patrick A. V. Hall, Bal Krishna Bal, Sagun Dhakhwa, Bhim Narayan Regmi
CICLing (1)2
2010 Towards Building Annotated Resources for Analyzing Opinions and Argumentation in News Editorials
Bal Krishna Bal, Patrick Saint-Dizier
LREC1