Chandresh Maurya

dblp:303/4435 · also Chandresh Kumar Maurya · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-3519-600XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CLAOCS-TX: Cross-Lingual Triplet Extraction with Aspect-Opinion-Aware Code-Switched Prompting and LLM-Guided Contrastive Distillation
abstract
Cross-lingual learning enables the transfer of structured sentiment knowledge from high-resource languages to unlabeled or low-resource languages, but prior work has largely focused on coarse-grained sentiment classification or aspect extraction.In contrast, zero-shot cross-lingual aspectopinion-sentiment triplet extraction (ASTE), which extracts sentiment triplets of the form (aspect term, opinion term, sentiment polarity), remains underexplored.We propose a unified framework that leverages large language models (LLMs) as both structured pseudo-label generators and semantic teachers for ASTE.Our approach employs stepwise structured prompting over aspect-and opinion-aware code-switched variants to generate reliable pseudo triplets, followed by a multi-variant consistency filter to retain high-confidence supervision.We further introduce a triplet-aware contrastive distillation objective that aligns student triplet representations with LLM-encoded semantic embeddings.During inference, only the student ASTE model is used, without requiring LLM access.Experiments on four non-Indic and four low-resource Indic target languages show consistent improvements over strong cross-lingual and LLM-based baselines.The proposed method yields an absolute micro-F1 improvement of 5.3 points on non-Indic languages and 3.8 points on low-resource Indic languages compared to the best competing approach.Ablation results further validate the complementary roles of aspect-and opinion-aware code-switched prompting and triplet-aware contrastive distillation, with larger relative gains observed in low-resource Indic settings.
Lipika Dewangan, Chandresh Maurya
ACL (1)2
2026 MaitH 1.0: A Parallel Corpus and Baseline for Low-Resource Maithili-Hindi Translation
Kamanksha Prasad Dubey, Chandresh Maurya, Kumar Padmanabh
LREC2
2026 Continual End-to-End Speech-to-Text translation using augmented bi-sampler
Balaram Sarkar, Pranav Karande, Ankit Malviya, Chandresh Maurya
Comput. Speech Lang.4
2026 CMF_Hit: Enhancing code-Mixed aspect-Based sentiment analysis via language-Aware gradient-Based tokenization and feature fusion
Lipika Dewangan, Chandresh Maurya
Expert Syst. Appl.2
2025 Benchmark Creation for Aspect-Based Sentiment Analysis in Low-Resource Odia Language and Evaluation through Fine-Tuning of Multilingual Models
abstract
The rapid growth of online product reviews spurs significant interest in Aspect-Based Sentiment Analysis (ABSA), which involves identifying aspect terms and their associated sentiment polarity. While ABSA is widely studied in resource-rich languages like English, Chinese, and Spanish, it remains underexplored in low-resource languages such as Odia. To address this gap, we create a reliable resource for aspect-based sentiment analysis in Odia. The dataset is annotated for two specific tasks: Aspect Term Extraction (ATE) and Aspect Polarity Classification (APC), spanning seven domains and aligned with the SemEval-2014 benchmark. Furthermore, we employ an ensemble data augmentation approach combining back-translation with a fine-tuned T5 paraphrase generation model to enhance the dataset and apply a semantic similarity filter using a Universal Sentence Encoder (USE) to remove low-quality data and ensure a balanced distribution of sample difficulty in the newly augmented dataset. Finally, we validate our dataset by fine-tuning multilingual pre-trained models, XLM-R and IndicBERT, on ATE and APC tasks. Additionally, we use three classical baseline models to evaluate the quality of the proposed dataset for these tasks. We hope the Odia dataset will spur more work for the ABSA task.
Lipika Dewangan, Zoyah Afsheen Sayeed, Chandresh Maurya
COLING3
2025 Improving Bird Classification with Primary Color Additives
Ezhini Rasendiran R, Chandresh Maurya
INTERSPEECH2
2025 HiProIBM: unsupervised continual learning through hierarchical prototypical cross-level discrimination along with information bottleneck subnetwork masking
Ankit Malviya, Chandresh Maurya
Appl. Intell.2
2025 End-to-End Speech-to-Text Translation: A Survey
Nivedita Sethiya, Chandresh Maurya
Comput. Speech Lang.2
2025 Direct speech-to-speech neural machine translation: A survey
Mahendra Gupta, Maitreyee Dutta, Chandresh Maurya
Speech Commun.3
2025 Indic-ST: A Large-Scale Multilingual Corpus for Low-Resource Speech-to-Text Translation
abstract
We introduce Indic-ST, a novel dataset for speech-to-text translation (ST) task from English to Indic languages to bridge the performance gap. ST involves converting spoken input in one language into written text in another, playing a key role in real-world applications like subtitling, lecture transcription, and multilingual communication systems. Despite several efforts like Meta’s seamless m4t, OpenAI’s Whisper, or Google USM model, the performance of ST models on low-resource languages lags to that of English (or high-resource languages like European languages). Indic-ST is compiled from four distinct domains: conversational audio, religious texts, education, and news, which combined results in the Indic-ST dataset. To the best of our knowledge, this is the largest low-resource ST data covering approximately 6,800 hours of English speech in the real human voice and text in 15 Indic languages with diverse scripts totaling approximately 900 GB in size. To assess the usefulness of the dataset, we present the baseline performance of individual language pairs using state-of-the-art ST models. We also present a unified multilingual English-to-Indic-ST model. The code and dataset are available at https://github.com/Nivedita5/Indic-ST .
Nivedita Sethiya, Saanvi Nair, Puneet Walia, Chandresh Maurya
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2024 Indic-TEDST: Datasets and Baselines for Low-Resource Speech to Text Translation
abstract
Speech-to-text (ST) task is the translation of speech in a language to text in a different language. It has use cases in subtitling, dubbing, etc. Traditionally, ST task has been solved by cascading automatic speech recognition (ASR) and machine translation (MT) models which leads to error propagation, high latency, and training time. To minimize such issues, end-to-end models have been proposed recently. However, we find that only a few works have reported results of ST models on a limited number of low-resource languages. To take a step further in this direction, we release datasets and baselines for low-resource ST tasks. Concretely, our dataset has 9 language pairs and benchmarking has been done against SOTA ST models. The low performance of SOTA ST models on Indic-TEDST data indicates the necessity of the development of ST models specifically designed for low-resource languages.
Nivedita Sethiya, Saanvi Nair, Chandresh Maurya
LREC/COLING3
2021 FiLMing Multimodal Sarcasm Detection with Attention
Sundesh Gupta, Aditya Shah, Miten Shah, Laribok Syiemlieh, Chandresh Maurya
ICONIP (5)5
2021 Distributed Sparse Class-Imbalance Learning and Its Applications
abstract
In the present work, the study on class imbalance problems in adistributedsetting exploiting sparsity structure in the data has been carried out. We formulate the class-imbalance learning problem as a cost-sensitive learning problem with$L_1$regularization. The cost-sensitive loss function is a cost-weighted smooth hinge loss. The resultant optimization problem is minimized within theDistributed Alternating Direction Method of Multiplier(DADMM) framework. We partition the data matrix across samples. This operation splits the original problem into a distributed$L_2$regularized smooth loss minimization and a$L_1$regularized squared loss minimization.$L_2$regularized subproblem is solved via Limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) and random coordinate descent method in parallel at multiple processing nodes usingMPIwhereas$L_1$regularized problem is just a simple soft-thresholding operation. We show, empirically, that the distributed solution approximates the centralized solution on many benchmark data sets. The centralized solution is obtained via Cost-Sensitive Stochastic Coordinate Descent (CSSCD). Empirical results on small and large-scale benchmark datasets show some promising avenues to further investigate the real-world applications of the proposed algorithms such as anomaly detection, class-imbalance learning, etc. To the best of our knowledge, ours is the first work to study class-imbalance in adistributedenvironment on large-scalesparsedata.
Chandresh Maurya, Durga Toshniwal, Gopalan Vijendran Venkoparao
IEEE Trans. Big Data1
2018 Prediction of Invoice Payment Status in Account Payable Business Process
Tarun Tater, Sampath Dechu, Senthil Mani, Chandresh Maurya
ICSOC4
2018 Large-Scale Distributed Sparse Class-Imbalance Learning
Chandresh Maurya, Durga Toshniwal
Inf. Sci.1
2016 Online sparse class imbalance learning on big data
Chandresh Maurya, Durga Toshniwal, Gopalan Vijendran Venkoparao
Neurocomputing1