Rishabh Maheshwary

dblp:282/0038 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Augmenting LLM Reasoning with Dynamic Notes Writing for Complex MultiHop QA
Rishabh Maheshwary, Masoud Hashemi, Khyati Mahajan, Shiva Krishna Reddy Malay, Sai Rajeswar, Sathwik Tejaswi Madhusudhan, Spandana Gella, Vikas Yadav
LREC1
2025 M-RewardBench: Evaluating Reward Models in Multilingual Settings
abstract
Srishti Gureja, Lester James Validad Miranda, Shayekh Bin Islam, Rishabh Maheshwary, Drishti Sharma, Gusti Triandi Winata, Nathan Lambert, Sebastian Ruder, Sara Hooker, Marzieh Fadaee. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Srishti Gureja, Lester James V. Miranda, Shayekh Bin Islam, Rishabh Maheshwary, Drishti Sharma, Gusti Winata, Nathan Lambert 0001, Sebastian Ruder, Sara Hooker, Marzieh Fadaee
ACL (1)4
2025 INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
abstract
The performance differential of large language models (LLM) between languages hinders their effective deployment in many regions, inhibiting the potential economic and societal value of generative AI tools in many communities. However, the development of functional LLMs in many languages (i.e., multilingual LLMs) is bottlenecked by the lack of high-quality evaluation resources in languages other than English. Moreover, current practices in multilingual benchmark construction often translate English resources, ignoring the regional and cultural knowledge of the environments in which multilingual systems would be used. In this work, we construct an evaluation suite of 197,243 QA pairs from local exam sources to measure the capabilities of multilingual LLMs in a variety of regional contexts. Our novel resource, INCLUDE, is a comprehensive knowledge- and reasoning-centric benchmark across 44 written languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
Angelika Romanou, Negar Foroutan Eghlidi, Anna Sotnikova, Zeming Chen 0001, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Mohamed A. Haggag, Imanol Schlag, Marzieh Fadaee, Sara Hooker, Antoine Bosselut, Snegha A, Alfonso Amayuelas, Azril Hafizi Amirudin, Viraat Aryabumi, Danylo Boiko, Jenny Chim, Gal Cohen, Aditya Kumar Dalmia, Abraham Diress, Sharad Duwal, Daniil Dzenhaliou, Daniel Fernando Erazo Florez, Fabian Farestam, Joseph Marvin Imperial, Shayekh Bin Islam, Perttu Isotalo, Maral Jabbarishiviari, Börje Karlsson 0001, Eldar Khalilov, Christopher Klamm, Fajri Koto, Dominik Krzeminski, Gabriel Adriano de Melo, Syrielle Montariol, Yiyang Nan, Joel Niklaus, Jekaterina Novikova, Johan S. Obando-Ceron, Debjit Paul, Esther Ploeger, Jebish Purbey, Swati Rajwal, Selvan Sunitha Ravi, Sara Rydell, Roshan Santhosh, Drishti Sharma, Marjana Prifti Skenduli, Arshia Soltani Moakhar, Bardia Soltani Moakhar, Ran Tamir, Ayush K. Tarun, Azmine Toushik Wasi, Thenuka Ovin Weerasinghe, Serhan Yilmaz, Mike Zhang
ICLR7
2025 M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
abstract
Rishabh Maheshwary, Vikas Yadav, Hoang H Nguyen, Khyati Mahajan, Sathwik Tejaswi Madhusudhan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Rishabh Maheshwary, Vikas Yadav, Hoang Nguyen 0006, Khyati Mahajan, Sathwik Tejaswi Madhusudhan
NAACL (Long Papers)1
2025 Prompting with Phonemes: Enhancing LLMs' Multilinguality for Non-Latin Script Languages
abstract
Hoang H Nguyen, Khyati Mahajan, Vikas Yadav, Julian Salazar, Philip S. Yu, Masoud Hashemi, Rishabh Maheshwary. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hoang Nguyen 0006, Khyati Mahajan, Vikas Yadav, Julian Salazar, Philip S. Yu, Masoud Hashemi, Rishabh Maheshwary
NAACL (Long Papers)7
2023 Improving Selective Visual Question Answering by Learning from Your Peers
abstract
Despite advances in Visual Question Answering (VQA), the ability of models to assess their own correctness remains under-explored. Recent work has shown that VQA models, out-of-the-box, can have difficulties abstaining from answering when they are wrong. The option to abstain, also called Selective Prediction, is highly relevant when deploying systems to users who must trust the system's output (e.g., VQA assistants for users with visual impairments). For such scenarios, abstention can be especially important as users may provide out-of-distribution (OOD) or adversarial inputs that make incorrect answers more likely. In this work, we explore Selective VQA in both in-distribution (ID) and OOD scenarios, where models are presented with mixtures of ID and OOD data. The goal is to maximize the number of questions answered while minimizing the risk of error on those questions. We propose a simple yet effective Learning from Your Peers (LYP) approach for training multimodal selection functions for making abstention decisions. Our approach uses predictions from models trained on distinct subsets of the training data as targets for optimizing a Selective VQA model. It does not require additional manual labels or held-out data and provides a signal for identifying examples that are easy/difficult to generalize to. In our extensive evaluations, we show this benefits a number of models across different architectures and scales. Overall, for ID, we reach 32.92% in the selective prediction metric coverage at 1 % risk of error$(\mathcal{C} {@} 1\%)$which doubles the previous best coverage of 15.79% on this task. For mixed ID/OOD, using models' softmax confidences for abstention decisions performs very poorly, answering$\mathcal{C}$@1%.
Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam, Stefan Scherer, Xinlei Chen, Matthieu Cord, Marcus Rohrbach
CVPR3
2022 Practice Makes a Solver Perfect: Data Augmentation for Math Word Problem Solvers
abstract
Existing Math Word Problem (MWP) solvers have achieved high accuracy on benchmark datasets.However, prior works have shown that such solvers do not generalize well and rely on superficial cues to achieve high performance.In this paper, we first conduct experiments to showcase that this behaviour is mainly associated with the limited size and diversity present in existing MWP datasets.Next, we propose several data augmentation techniques broadly categorized into Substitution and Paraphrasing based methods.By deploying these methods we increase the size of existing datasets by five folds.Extensive experiments on two benchmark datasets across three state-of-the-art MWP solvers shows that proposed methods increase the generalization and robustness of existing solvers.On average, proposed methods significantly increase the state-of-the-art results by over five percentage points on benchmark datasets.Further, the solvers trained on the augmented dataset performs comparatively better on the challenge test set.We also show the effectiveness of proposed techniques through ablation studies and verify the quality of augmented samples through human evaluation. Original ProblemProblem: Nancy grew 8 potatoes.Sandy grew 5 potatoes.How many potatoes did they grow in total ?True Equation: X = 8+5Paraphrasing Method Problem: How many potatoes did they grow in all given that nancy grew 8 potatoes and sandy grew 5 potatoes.Equation Label: X = 8+5 Substitution Method Problem: Dwight grew 8 potatoes.Juliette grew 5 potatoes.How many potatoes did they grow together ?Equation Label: X = 8+5
Rishabh Maheshwary, Vikram Pudi
NAACL-HLT2
2021 Generating Natural Language Attacks in a Hard Label Black Box Setting
abstract
We study an important and challenging task of attacking natural language processing models in a hard label black box setting. We propose a decision-based attack strategy that crafts high quality adversarial examples on text classification and entailment tasks. Our proposed attack strategy leverages population-based optimization algorithm to craft plausible and semantically similar adversarial examples by observing only the top label predicted by the target model. At each iteration, the optimization procedure allow word replacements that maximizes the overall semantic similarity between the original and the adversarial text. Further, our approach does not rely on using substitute models or any kind of training data. We demonstrate the efficacy of our proposed approach through extensive experimentation and ablation studies on five state-of-the-art target models across seven benchmark datasets. In comparison to attacks proposed in prior literature, we are able to achieve a higher success rate with lower word perturbation percentage that too in a highly restricted setting.
Rishabh Maheshwary, Saket Maheshwary, Vikram Pudi
AAAI1
2021 A Context Aware Approach for Generating Natural Language Attacks
abstract
We study an important task of attacking natural language processing models in a black box setting. We propose an attack strategy that crafts semantically similar adversarial examples on text classification and entailment tasks. Our proposed attack finds candidate words by considering the information of both the original word and its surrounding context. It jointly leverages masked language modelling and next sentence prediction for context understanding. In comparison to attacks proposed in prior literature, we are able to generate high quality adversarial examples that do significantly better both in terms of success rate and word perturbation percentage.
Rishabh Maheshwary, Saket Maheshwary, Vikram Pudi
AAAI1
2021 A Strong Baseline for Query Efficient Attacks in a Black Box Setting
abstract
Existing black box search methods have achieved high success rate in generating adversarial attacks against NLP models.However, such search methods are inefficient as they do not consider the amount of queries required to generate adversarial attacks.Also, prior attacks do not maintain a consistent search space while comparing different search methods.In this paper, we propose a query efficient attack strategy to generate plausible adversarial examples on text classification and entailment tasks.Our attack jointly leverages attention mechanism and locality sensitive hashing (LSH) to reduce the query count.We demonstrate the efficacy of our approach by comparing our attack with four baselines across three different search spaces.Further, we benchmark our results across the same search space used in prior attacks.In comparison to attacks proposed, on an average, we are able to reduce the query count by 75% across all datasets and target models.We also demonstrate that our attack achieves a higher success rate when compared to prior attacks in a limited query setting.
Rishabh Maheshwary, Saket Maheshwary, Vikram Pudi
EMNLP (1)1