EDBT 2026 Demo / reviewers in the wild / expert
Mihir Parmar
dblp:253/6105
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 14 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoAct: Co-Active LLM Preference Learning with Human-AI SynergyabstractLearning from preference-based feedback has become an effective approach for aligning LLMs across diverse tasks.However, highquality human-annotated preference data remains expensive and scarce.Existing methods address this challenge through either selfrewarding, which scales by using purely AIgenerated labels but risks unreliability, or active learning, which ensures quality through oracle annotation but cannot fully leverage unlabeled data.In this paper, we present COACT, a novel framework that synergistically combines selfrewarding and active learning through strategic human-AI collaboration.COACT leverages self-consistency to identify both reliable selflabeled data and samples that are requiring oracle verification.Additionally, oracle feedback guides the model to generate new instructions within its solvable capability.Evaluated on three reasoning benchmarks across two model families, COACT achieves average improvements of +13.25% on GSM8K, +8.19% on MATH, and +13.16% on WebInstruct, consistently outperforming all baselines.1 Ruiyao Xu, Mihir Parmar, Tiankai Yang 0001, Zhengyu Hu, Yue Zhao 0016, Kaize Ding |
ACL (1) | 2 |
| 2025 | PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem SolvingabstractMihir Parmar, Palash Goyal, Xin Liu, Yiwen Song, Mingyang Ling, Chitta Baral, Hamid Palangi, Tomas Pfister. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Mihir Parmar, Palash Goyal, Yiwen Song, Chitta Baral, Hamid Palangi, Tomas Pfister |
EMNLP | 1 |
| 2025 | PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem SolvingabstractMihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen, Long Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Mihir Parmar, Palash Goyal, Yanfei Chen, Long T. Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang 0002, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi |
EMNLP | 1 |
| 2025 | ThinkTuning: Instilling Cognitive Reflections without DistillationabstractRecent advances in test-time scaling have led to the emergence of thinking LLMs that exhibit self-reflective behaviors and multi-step reasoning.While RL drives this self-improvement paradigm, a recent study (Gandhi et al., 2025) shows that RL alone does not truly instill these new reasoning abilities -it merely draws out behaviors already present in the base models.This raises a question: How can we train models that don't exhibit such thinking behavior to develop it in the first place?To this end, we propose THINKTUNING, a GRPO-based interactive training approach where we augment the rollouts of a student model with the guidance from a teacher model.A simple idea from classroom practice inspires our method: a teacher poses a problem, lets the student try an answer, then gives corrective feedback-enough to point the mind in the right direction and then show the solution.Each piece of feedback reshapes the student's thoughts, leading them to arrive at the correct solution.Similarly, we find that this type of implicit supervision through feedback from a teacher model of the same size improves the reasoning capabilities of the student model.In particular, on average, our method shows a 3.85% improvement over zero-shot baselines across benchmarks, and on MATH-500, AIME and GPQA-Diamond it shows 2.08%, 2.23% and 3.99% improvements over the vanilla-GRPO baseline 1 . Aswin RRV, Jacob Dineen, Divij Handa, Md Nayem Uddin, Mihir Parmar, Chitta Baral, Ben Zhou |
EMNLP | 5 |
| 2024 | LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsabstractMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura, Man Luo, Santosh Mashetty, Arindam Mitra, Chitta Baral. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura, Man Luo 0003, Santosh Mashetty, Arindam Mitra, Chitta Baral |
ACL (1) | 1 |
| 2024 | Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMsabstractExtractive summarization plays a pivotal role in natural language processing due to its widerange applications in summarizing diverse content efficiently, while also being faithful to the original content.Despite significant advancement achieved in extractive summarization by Large Language Models (LLMs), these summaries frequently exhibit incoherence.An important aspect of the coherent summary is its readability for intended users.Although there have been many datasets and benchmarks proposed for creating coherent extractive summaries, none of them currently incorporate user intent to improve coherence in extractive summarization.Motivated by this, we propose a systematically created human-annotated dataset consisting of coherent summaries for five publicly available datasets and natural language user feedback, offering valuable insights into how to improve coherence in extractive summaries.We utilize this dataset for aligning LLMs through supervised fine-tuning with natural language human feedback to enhance the coherence of their generated summaries.Preliminary experiments with Falcon-40B and Llama-2-13B show significant performance improvements (∼ 10% Rouge-L) in terms of producing coherent summaries.We further utilize human feedback to benchmark results over instruction-tuned models such as FLAN-T5 which resulted in several interesting findings 1 . Mihir Parmar, Hanieh Deilamsalehy, Franck Dernoncourt, Seunghyun Yoon 0002, Ryan Rossi, Trung Bui |
EMNLP | 1 |
| 2024 | Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language ModelsabstractNisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja, Mutsumi Nakamura, Neeraj Varshney, Chitta Baral. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Nisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja, Mutsumi Nakamura, Neeraj Varshney, Chitta Baral |
EMNLP | 3 |
| 2024 | Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?abstractNemika Tyagi, Mihir Parmar, Mohith Kulkarni, Aswin Rrv, Nisarg Patel, Mutsumi Nakamura, Arindam Mitra, Chitta Baral. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Nemika Tyagi, Mihir Parmar, Mohith Kulkarni, Aswin RRV, Nisarg Patel, Mutsumi Nakamura, Arindam Mitra, Chitta Baral |
EMNLP | 2 |
| 2023 | Don't Blame the Annotator: Bias Already Starts in the Annotation InstructionsabstractIn recent years, progress in NLU has been driven by benchmarks.These benchmarks are typically collected by crowdsourcing, where annotators write examples based on annotation instructions crafted by dataset creators.In this work, we hypothesize that annotators pick up on patterns in the crowdsourcing instructions, which bias them to write many similar examples that are then over-represented in the collected data.We study this form of bias, termed instruction bias, in 14 recent NLU benchmarks, showing that instruction examples often exhibit concrete patterns, which are propagated by crowdworkers to the collected data.This extends previous work (Geva et al., 2019) and raises a new concern of whether we are modeling the dataset creator's instructions, rather than the task.Through a series of experiments, we show that, indeed, instruction bias can lead to overestimation of model performance, and that models struggle to generalize beyond biases originating in the crowdsourcing instructions.We further analyze the influence of instruction bias in terms of pattern frequency and model size, and derive concrete recommendations for creating future NLU benchmarks. 1 Mihir Parmar, Swaroop Mishra, Mor Geva, Chitta Baral |
EACL | 1 |
| 2022 | Less is More: Summary of Long Instructions is Better for Program SynthesisabstractDespite the success of large pre-trained language models (LMs) such as Codex, they show below-par performance on the larger and more complicated programming related questions.We show that LMs benefit from the summarized version of complicated questions.Our findings show that superfluous information often present in problem description such as human characters, background stories, and names (which are included to help humans in understanding a task) does not help models in understanding a task.To this extent, we create a metadataset from the frequently used APPS dataset and the newly created CodeContests dataset for the program synthesis task.Our meta-dataset consists of human and synthesized summaries of the long and complicated programming questions.Experimental results on Codex show that our proposed approach outperforms baseline by 8.13% on the APPS dataset and 11.88% on the CodeContests dataset on average in terms of strict accuracy.Our analysis shows that summaries significantly improve performance for introductory (9.86%) and interview (11.48%) programming questions.However, it shows improvement by a small margin (∼ 2%) for competitive programming questions, implying scope for future research in this direction.1 * Equal Contribution 1 Code and data is available at https://github.com/kurbster/ Prompt-Summarization 2 Detailed related work is presented in Appendix A However, LMs such as Codex show below-par performance on the long and complicated programming questions.We observe that the natural language description of the program becomes long and complicated when there is superfluous information (see section 2.1.1).The goal of adding this information to the description is to make it more understandable to humans.However, we find that this information confuses the model in understanding a task 3 .We propose that removing the excess information and providing the model with the exact specifications of the problem can improve the performance of the LMs.To remove excess information 4 , we summarize the descriptions of the program in such a way that it does not lose important specifications.We use the APPS dataset (Hendrycks et al., 2021) and Code-Contests dataset (Li et al., 2022) which are a collection of coding problems from different online sources and create a meta-dataset consisting of human and synthesized summaries.We perform all experiments using the GPT-based Codex model (Chen et al., 2021) on the proposed meta-dataset and show that the summarized version of complicated questions improves strict accuracy by 8.13% on the APPS dataset and 11.85% on CodeContests.From our analysis, we can see significant improvement for introductory (9.86%) and interview (11.48%) related programming questions.However, it shows improvement by a small margin (∼ 2%) for competitive programming questions.Considering that automatic evaluation of a program does not reward for partial correctness, we perform qualitative evaluation on our meta-dataset and find that original questions often confuse models in understanding the underlying problem, as models latch on to some spurious words in the text (e.g. the word 'list' in question makes the model 3 See example in Appendix C 4 Instructions for creating summaries given in Appendix N 4532 design a list even though the underlying problem is on graphs).We further analyze model performance on different types of summaries (i.e., basic, expert, and synthetic) and provide instruction-design principles that can help future research on prompting in program synthesis.2 Method 2.1 Dataset We use the APPS (Hendrycks et al., 2021) and CodeContests (Li et al., 2022) datasets to create summaries.We crowd-sourced the creation of human summaries.The result was 373 human summaries for APPS and 80 summaries for CodeContests along with and 8663 synthetic summaries using both datasets.Table 1 shows the statistics of the generated summaries. Kirby Kuznia, Swaroop Mishra, Mihir Parmar, Chitta Baral |
EMNLP | 3 |
| 2022 | Is a Question Decomposition Unit All We Need?abstractLarge Language Models (LMs) have achieved state-of-the-art performance on many Natural Language Processing (NLP) benchmarks.With the growing number of new benchmarks, we build bigger and more complex LMs.However, building new LMs may not be an ideal option owing to the cost, time and environmental impact associated with it.We explore an alternative route: can we modify data by expressing it in terms of the model's strengths, so that a question becomes easier for models to answer?We investigate if humans can decompose a hard question into a set of simpler questions that are relatively easier for models to solve.We analyze a range of datasets involving various forms of reasoning and find that it is indeed possible to significantly improve model performance (24% for GPT3 and 29% for RoBERTa-SQuAD along with a symbolic calculator) via decomposition.Our approach provides a viable option to involve people in NLP research in a meaningful way.Our findings indicate that Human-in-the-loop Question Decomposition (HQD) can potentially provide an alternate path to building large LMs 1 . Pruthvi Patel, Swaroop Mishra, Mihir Parmar, Chitta Baral |
EMNLP | 3 |
| 2022 | Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP TasksabstractYizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, Xudong Shen. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma 0001, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit |
EMNLP | 22 |
| 2021 | Fundamental Challenges in Deep Learning for Stiff Contact DynamicsabstractFrictional contact has been extensively studied as the core underlying behavior of legged locomotion and manipulation, and its nearly-discontinuous nature makes planning and control difficult even when an accurate model of the robot is available. Here, we present empirical evidence that learning an accurate model in the first place can be confounded by contact, as modern deep learning approaches are not designed to capture this non-smoothness. We isolate the effects of contact’s non-smoothness by varying the mechanical stiffness of a compliant contact simulator. Even for a simple system, we find that stiffness alone dramatically degrades training processes, generalization, and data-efficiency. Our results raise serious questions about simulated testing environments which do not accurately reflect the stiffness of rigid robotic hardware. Significant additional investigation will be necessary to fully understand and mitigate these effects, and we suggest several avenues for future study. Mihir Parmar, Mathew Halm, Michael Posa |
IROS | 1 |
| 2021 | Residual Neural Network precisely quantifies dysarthria severity-level based on short-duration speech segments
Siddhant Gupta, Ankur T. Patil, Mirali Purohit, Mihir Parmar, Maitreya Patel, Hemant A. Patil, Rodrigo Capobianco Guido |
Neural Networks | 4 |