VLDB 2026 Research / reviewers in the wild / expert
Chandresh Maurya
dblp:303/4435 · also Chandresh Kumar Maurya
· DBLP profile ↗
16ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-3519-600XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLAOCS-TX: Cross-Lingual Triplet Extraction with Aspect-Opinion-Aware Code-Switched Prompting and LLM-Guided Contrastive DistillationabstractCross-lingual learning enables the transfer of structured sentiment knowledge from high-resource languages to unlabeled or low-resource languages, but prior work has largely focused on coarse-grained sentiment classification or aspect extraction.In contrast, zero-shot cross-lingual aspectopinion-sentiment triplet extraction (ASTE), which extracts sentiment triplets of the form (aspect term, opinion term, sentiment polarity), remains underexplored.We propose a unified framework that leverages large language models (LLMs) as both structured pseudo-label generators and semantic teachers for ASTE.Our approach employs stepwise structured prompting over aspect-and opinion-aware code-switched variants to generate reliable pseudo triplets, followed by a multi-variant consistency filter to retain high-confidence supervision.We further introduce a triplet-aware contrastive distillation objective that aligns student triplet representations with LLM-encoded semantic embeddings.During inference, only the student ASTE model is used, without requiring LLM access.Experiments on four non-Indic and four low-resource Indic target languages show consistent improvements over strong cross-lingual and LLM-based baselines.The proposed method yields an absolute micro-F1 improvement of 5.3 points on non-Indic languages and 3.8 points on low-resource Indic languages compared to the best competing approach.Ablation results further validate the complementary roles of aspect-and opinion-aware code-switched prompting and triplet-aware contrastive distillation, with larger relative gains observed in low-resource Indic settings. Lipika Dewangan, Chandresh Maurya |
ACL (1) | 2 |
| 2026 | MaitH 1.0: A Parallel Corpus and Baseline for Low-Resource Maithili-Hindi Translation
Kamanksha Prasad Dubey, Chandresh Maurya, Kumar Padmanabh |
LREC | 2 |
| 2026 | Continual End-to-End Speech-to-Text translation using augmented bi-sampler
Balaram Sarkar, Pranav Karande, Ankit Malviya, Chandresh Maurya |
Comput. Speech Lang. | 4 |
| 2026 | CMF_Hit: Enhancing code-Mixed aspect-Based sentiment analysis via language-Aware gradient-Based tokenization and feature fusion
Lipika Dewangan, Chandresh Maurya |
Expert Syst. Appl. | 2 |
| 2025 | Benchmark Creation for Aspect-Based Sentiment Analysis in Low-Resource Odia Language and Evaluation through Fine-Tuning of Multilingual ModelsabstractThe rapid growth of online product reviews spurs significant interest in Aspect-Based Sentiment Analysis (ABSA), which involves identifying aspect terms and their associated sentiment polarity. While ABSA is widely studied in resource-rich languages like English, Chinese, and Spanish, it remains underexplored in low-resource languages such as Odia. To address this gap, we create a reliable resource for aspect-based sentiment analysis in Odia. The dataset is annotated for two specific tasks: Aspect Term Extraction (ATE) and Aspect Polarity Classification (APC), spanning seven domains and aligned with the SemEval-2014 benchmark. Furthermore, we employ an ensemble data augmentation approach combining back-translation with a fine-tuned T5 paraphrase generation model to enhance the dataset and apply a semantic similarity filter using a Universal Sentence Encoder (USE) to remove low-quality data and ensure a balanced distribution of sample difficulty in the newly augmented dataset. Finally, we validate our dataset by fine-tuning multilingual pre-trained models, XLM-R and IndicBERT, on ATE and APC tasks. Additionally, we use three classical baseline models to evaluate the quality of the proposed dataset for these tasks. We hope the Odia dataset will spur more work for the ABSA task. Lipika Dewangan, Zoyah Afsheen Sayeed, Chandresh Maurya |
COLING | 3 |
| 2025 | Improving Bird Classification with Primary Color Additives
Ezhini Rasendiran R, Chandresh Maurya |
INTERSPEECH | 2 |
| 2025 | HiProIBM: unsupervised continual learning through hierarchical prototypical cross-level discrimination along with information bottleneck subnetwork masking
Ankit Malviya, Chandresh Maurya |
Appl. Intell. | 2 |
| 2025 | End-to-End Speech-to-Text Translation: A Survey
Nivedita Sethiya, Chandresh Maurya |
Comput. Speech Lang. | 2 |
| 2025 | Direct speech-to-speech neural machine translation: A survey
Mahendra Gupta, Maitreyee Dutta, Chandresh Maurya |
Speech Commun. | 3 |
| 2025 | Indic-ST: A Large-Scale Multilingual Corpus for Low-Resource Speech-to-Text TranslationabstractWe introduce Indic-ST, a novel dataset for speech-to-text translation (ST) task from English to Indic languages to bridge the performance gap. ST involves converting spoken input in one language into written text in another, playing a key role in real-world applications like subtitling, lecture transcription, and multilingual communication systems. Despite several efforts like Meta’s seamless m4t, OpenAI’s Whisper, or Google USM model, the performance of ST models on low-resource languages lags to that of English (or high-resource languages like European languages). Indic-ST is compiled from four distinct domains: conversational audio, religious texts, education, and news, which combined results in the Indic-ST dataset. To the best of our knowledge, this is the largest low-resource ST data covering approximately 6,800 hours of English speech in the real human voice and text in 15 Indic languages with diverse scripts totaling approximately 900 GB in size. To assess the usefulness of the dataset, we present the baseline performance of individual language pairs using state-of-the-art ST models. We also present a unified multilingual English-to-Indic-ST model. The code and dataset are available at https://github.com/Nivedita5/Indic-ST . Nivedita Sethiya, Saanvi Nair, Puneet Walia, Chandresh Maurya |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2024 | Indic-TEDST: Datasets and Baselines for Low-Resource Speech to Text TranslationabstractSpeech-to-text (ST) task is the translation of speech in a language to text in a different language. It has use cases in subtitling, dubbing, etc. Traditionally, ST task has been solved by cascading automatic speech recognition (ASR) and machine translation (MT) models which leads to error propagation, high latency, and training time. To minimize such issues, end-to-end models have been proposed recently. However, we find that only a few works have reported results of ST models on a limited number of low-resource languages. To take a step further in this direction, we release datasets and baselines for low-resource ST tasks. Concretely, our dataset has 9 language pairs and benchmarking has been done against SOTA ST models. The low performance of SOTA ST models on Indic-TEDST data indicates the necessity of the development of ST models specifically designed for low-resource languages. Nivedita Sethiya, Saanvi Nair, Chandresh Maurya |
LREC/COLING | 3 |
| 2021 | FiLMing Multimodal Sarcasm Detection with Attention
Sundesh Gupta, Aditya Shah, Miten Shah, Laribok Syiemlieh, Chandresh Maurya |
ICONIP (5) | 5 |
| 2021 | Distributed Sparse Class-Imbalance Learning and Its ApplicationsabstractIn the present work, the study on class imbalance problems in adistributedsetting exploiting sparsity structure in the data has been carried out. We formulate the class-imbalance learning problem as a cost-sensitive learning problem with$L_1$regularization. The cost-sensitive loss function is a cost-weighted smooth hinge loss. The resultant optimization problem is minimized within theDistributed Alternating Direction Method of Multiplier(DADMM) framework. We partition the data matrix across samples. This operation splits the original problem into a distributed$L_2$regularized smooth loss minimization and a$L_1$regularized squared loss minimization.$L_2$regularized subproblem is solved via Limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) and random coordinate descent method in parallel at multiple processing nodes usingMPIwhereas$L_1$regularized problem is just a simple soft-thresholding operation. We show, empirically, that the distributed solution approximates the centralized solution on many benchmark data sets. The centralized solution is obtained via Cost-Sensitive Stochastic Coordinate Descent (CSSCD). Empirical results on small and large-scale benchmark datasets show some promising avenues to further investigate the real-world applications of the proposed algorithms such as anomaly detection, class-imbalance learning, etc. To the best of our knowledge, ours is the first work to study class-imbalance in adistributedenvironment on large-scalesparsedata. Chandresh Maurya, Durga Toshniwal, Gopalan Vijendran Venkoparao |
IEEE Trans. Big Data | 1 |
| 2018 | Prediction of Invoice Payment Status in Account Payable Business Process
Tarun Tater, Sampath Dechu, Senthil Mani, Chandresh Maurya |
ICSOC | 4 |
| 2018 | Large-Scale Distributed Sparse Class-Imbalance Learning
Chandresh Maurya, Durga Toshniwal |
Inf. Sci. | 1 |
| 2016 | Online sparse class imbalance learning on big data
Chandresh Maurya, Durga Toshniwal, Gopalan Vijendran Venkoparao |
Neurocomputing | 1 |