Ashutosh Modi

dblp:139/0873 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
22since 2021 · last 2025
0000-0002-0962-8350ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 4 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Calibration Across Layers: Understanding Calibration Evolution in LLMs
abstract
Large Language Models (LLMs) have demonstrated inherent calibration capabilities, where predicted probabilities align well with correctness, despite prior findings that deep neural networks are often overconfident.Recent studies have linked this behavior to specific components in the final layer, such as entropy neurons and the unembedding matrix's null space.In this work, we provide a complementary perspective by investigating how calibration evolves throughout the network's depth.Analyzing multiple open-weight models on the MMLU benchmark, we uncover a distinct confidence correction phase in the upper/later layers, where model confidence is actively recalibrated after decision certainty has been reached.Furthermore, we identify a low-dimensional calibration direction in the residual stream whose perturbation significantly improves calibration metrics (ECE and MCE) without harming accuracy.Our findings suggest that calibration is a distributed phenomenon, shaped throughout the network's forward pass, not just in its final projection, providing new insights into how confidence-regulating mechanisms operate within LLMs.
Abhinav Joshi, Areeb Ahmad, Ashutosh Modi
EMNLP3
2025 PoseStitch-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
abstract
Sign language translation remains a challenging task due to the scarcity of large-scale, sentence-aligned datasets.Prior arts have focused on various feature extraction and architectural changes to support neural machine translation for sign languages.We propose POSESTITCH-SLT, a novel pre-training scheme that is inspired by linguistic-templatesbased sentence generation technique.With translation comparison on two sign language datasets, How2Sign and iSign, we show that a simple transformer-based encoder-decoder architecture outperforms the prior art when considering template-generated sentence pairs in training.We achieve BLEU-4 score improvements from 1.97 to 4.56 on How2Sign and from 0.55 to 3.43 on iSign, surpassing prior state-ofthe-art methods for pose-based gloss-free translation.The results demonstrate the effectiveness of template-driven synthetic supervision in low-resource sign language settings.
Abhinav Joshi, Sanjeet Singh, Ashutosh Modi
EMNLP4
2025 IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
abstract
Identifying/retrieving relevant statutes and prior cases/precedents for a given legal situation are common tasks exercised by law practitioners.Researchers to date have addressed the two tasks independently, thus developing completely different datasets and models for each task; however, both retrieval tasks are inherently related, e.g., similar cases tend to cite similar statutes (due to similar factual situation).In this paper, we address this gap.We propose IL-PCSR (Indian Legal corpus for Prior Case and Statute Retrieval), which is a unique corpus that provides a common testbed for developing models for both the tasks (Statute Retrieval and Precedent Retrieval) that can exploit the dependence between the two.We experiment extensively with several baseline models on the tasks, including lexical models, semantic models and ensemble based on GNNs.Further, to exploit the dependence between the two tasks, we develop an LLM-based re-ranking approach that gives the best performance.
Shounak Paul, Dhananjay Ghumare, Pawan Goyal 0002, Saptarshi Ghosh 0001, Ashutosh Modi
EMNLP5
2025 Towards Quantifying Commonsense Reasoning with Mechanistic Insights
abstract
Abhinav Joshi, Areeb Ahmad, Divyaksh Shukla, Ashutosh Modi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Abhinav Joshi, Areeb Ahmad, Divyaksh Shukla, Ashutosh Modi
NAACL (Long Papers)4
2025 Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
abstract
Transformer-based language models exhibit complex and distributed behavior, yet their internal computations remain poorly understood. Existing mechanistic interpretability methods typically treat attention heads and multilayer perceptron layers (MLPs) (the building blocks of a transformer architecture) as indivisible units, overlooking possibilities of functional substructure learned within them. In this work, we introduce a more fine-grained perspective that decomposes these components into orthogonal singular directions, revealing superposed and independent computations within a single head or MLP. We validate our perspective on widely used standard tasks like Indirect Object Identification (IOI), Gender Pronoun (GP), and Greater Than (GT), showing that previously identified canonical functional heads, such as the “name mover,” encode multiple overlapping subfunctions aligned with distinct singular directions. Nodes in a computational graph, that are previously identified as circuit elements show strong activation along specific low-rank directions, suggesting that meaningful computations reside in compact subspaces. While some directions remain challenging to interpret fully, our results highlight that transformer computations are more distributed, structured, and compositional than previously assumed. This perspective opens new avenues for fine-grained mechanistic interpretability and a deeper understanding of model internals.
Areeb Ahmad, Abhinav Joshi, Ashutosh Modi
NeurIPS3
2025 Geometry of Decision Making in Language Models
abstract
Large Language Models (LLMs) show strong generalization across diverse tasks, yet the internal decision-making processes behind their predictions remain opaque. In this work, we study the geometry of hidden representations in LLMs through the lens of intrinsic dimension (ID), focusing specifically on decision-making dynamics in a multiple-choice question answering (MCQA) setting. We perform a large-scale study, with 28 open-weight transformer models and estimate ID across layers using multiple estimators, while also quantifying per-layer performance on MCQA tasks. Our findings reveal a consistent ID pattern across models: early layers operate on low-dimensional manifolds, middle layers expand this space, and later layers compress it again, converging to decision-relevant representations. Together, these results suggest LLMs implicitly learn to project linguistic inputs onto structured, low-dimensional manifolds aligned with task-specific decisions, providing new geometric insights into how generalization and reasoning emerge in language models.
Abhinav Joshi, Divyanshu Bhatt, Ashutosh Modi
NeurIPS3
2024 IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
abstract
Abhinav Joshi, Shounak Paul, Akshat Sharma, Pawan Goyal, Saptarshi Ghosh, Ashutosh Modi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Abhinav Joshi, Shounak Paul, Akshat Sharma, Pawan Goyal 0002, Saptarshi Ghosh 0001, Ashutosh Modi
ACL (1)6
2024 Towards Measuring and Modeling "Culture" in LLMs: A Survey
abstract
Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Shivdutt Singh, Alham Fikri Aji, Jacki O’Neill, Ashutosh Modi, Monojit Choudhury. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Alham Fikri Aji, Jacki O'Neill, Ashutosh Modi, Monojit Choudhury
EMNLP7
2024 BookSQL: A Large Scale Text-to-SQL Dataset for Accounting Domain
abstract
Rahul Kumar, Amar Raja Dibbu, Shrutendra Harsola, Vignesh Subrahmaniam, Ashutosh Modi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Amar Raja Dibbu, Shrutendra Harsola, Vignesh Subrahmaniam, Ashutosh Modi
NAACL-HLT5
2024 COLD: Causal reasOning in cLosed Daily activities
abstract
Large Language Models (LLMs) have shown state-of-the-art performance in a variety of tasks, including arithmetic and reasoning; however, to gauge the intellectual capabilities of LLMs, causal reasoning has become a reliable proxy for validating a general understanding of the mechanics and intricacies of the world similar to humans. Previous works in natural language processing (NLP) have either focused on open-ended causal reasoning via causal commonsense reasoning (CCR) or framed a symbolic representation-based question answering for theoretically backed-up analysis via a causal inference engine. The former adds an advantage of real-world grounding but lacks theoretically backed-up analysis/validation, whereas the latter is far from real-world grounding. In this work, we bridge this gap by proposing the COLD (Causal reasOning in cLosed Daily activities) framework, which is built upon human understanding of daily real-world activities to reason about the causal nature of events. We show that the proposed framework facilitates the creation of enormous causal queries (∼ 9 million) and comes close to the mini-turing test, simulating causal reasoning to evaluate the understanding of a daily real-world task. We evaluate multiple LLMs on the created causal queries and find that causal reasoning is challenging even for activities trivial to humans. We further explore (the causal reasoning abilities of LLMs) using the backdoor criterion to determine the causal strength between events.
Abhinav Joshi, Areeb Ahmad, Ashutosh Modi
NeurIPS3
2024 Text-Based Fine-Grained Emotion Prediction
abstract
Text-based emotion prediction is an important task in the field of affective computing. Most prior work has been restricted to predicting emotions corresponding to a few high-level emotion classes. This paper explores and experiments with various techniques for fine-grained (27 classes) emotion prediction†which appeared at ACII 2021. In particular, (1) we present a method to incorporate multiple annotations from different raters, (2) we analyze the model's performance on fused emotion classes and with sub-sampled training data, (3) we present a method to leverage the correlations among the emotion categories, and (4) we propose a new framework for text-based fine-grained emotion prediction through emotion definition modeling. The emotion definition-based model outperforms the existing state-of-the-art for fine-grained emotion dataset GoEmotions. The approach involves a multi-task learning framework that models definitions of emotions as an auxiliary task while being trained on the primary task of emotion prediction. We model definitions using masked language modeling and class definition prediction tasks. We show that this trained model can be used for transfer learning on other benchmark datasets in emotion prediction with varying emotion label sets, domains, and sizes. The proposed models outperform the baselines on transfer learning experiments demonstrating the model's generalization capability.
Gargi Singh, Dhanajit Brahma, Piyush Rai, Ashutosh Modi
IEEE Trans. Affect. Comput.4
2023 U-CREAT: Unsupervised Case Retrieval using Events extrAcTion
abstract
The task of Prior Case Retrieval (PCR) in the legal domain is about automatically citing relevant (based on facts and precedence) prior legal cases in a given query case.To further promote research in PCR, in this paper, we propose a new large benchmark (in English) for the PCR task: IL-PCR (Indian Legal Prior Case Retrieval) corpus.Given the complex nature of case relevance and the long size of legal documents, BM25 remains a strong baseline for ranking the cited prior documents.In this work, we explore the role of events in legal case retrieval and propose an unsupervised retrieval method-based pipeline U-CREAT (Unsupervised Case Retrieval using Events Extraction).We find that the proposed unsupervised retrieval method significantly increases performance compared to BM25 and makes retrieval faster by a considerable margin, making it applicable to real-time case retrieval systems.Our proposed system is generic, we show that it generalizes across two different legal systems (Indian and Canadian), and it shows state-ofthe-art performance on the benchmarks for both the legal systems (IL-PCR and COLIEE corpora).
Abhinav Joshi, Akshat Sharma, Sai Kiran Tanikella, Ashutosh Modi
ACL (1)4
2023 EtiCor: Corpus for Analyzing LLMs for Etiquettes
abstract
Etiquettes are an essential ingredient of dayto-day interactions among people.Moreover, etiquettes are region-specific, and etiquettes in one region might contradict those in other regions.In this paper, we propose EtiCor, an Etiquettes Corpus, having texts about social norms from five different regions across the globe.The corpus provides a test bed for evaluating LLMs for knowledge and understanding of region-specific etiquettes.Additionally, we propose the task of Etiquette Sensitivity.We experiment with state-of-the-art LLMs (Delphi, Falcon40B, and GPT-3.5).Initial results indicate that LLMs, mostly fail to understand etiquettes from regions from non-Western world. * Equal Contributions 1In this paper, we use the term etiquette and social norm inter-changeably Wikipedia DataPercentage
Ashutosh Dwivedi, Pradhyumna Lavania, Ashutosh Modi
EMNLP3
2023 ScriptWorld: Text Based Environment for Learning Procedural Knowledge
abstract
Text-based games provide a framework for developing natural language understanding and commonsense knowledge about the world in reinforcement learning based agents. Existing text-based environments often rely on fictional situations and characters to create a gaming framework and are far from real-world scenarios. In this paper, we introduce ScriptWorld: a text-based environment for teaching agents about real-world daily chores and hence imparting commonsense knowledge. To the best of our knowledge, it is the first interactive text-based gaming framework that consists of daily real-world human activities designed using scripts dataset. We provide gaming environments for 10 daily activities and perform a detailed analysis of the proposed environment. We develop RL-based baseline models/agents to play the games in ScriptWorld. To understand the role of language models in such environments, we leverage features obtained from pre-trained language models in the RL agents. Our experiments show that prior knowledge obtained from a pre-trained language model helps to solve real-world text-based gaming environments.
Abhinav Joshi, Areeb Ahmad, Umang Pandey, Ashutosh Modi
IJCAI4
2023 ASR for Low Resource and Multilingual Noisy Code-Mixed Speech
Tushar Verma, Atul Shree, Ashutosh Modi
INTERSPEECH3
2022 CISLR: Corpus for Indian Sign Language Recognition
abstract
Indian Sign Language, though used by a diverse community, still lacks well-annotated resources for developing systems that would enable sign language processing.In recent years researchers have actively worked for sign languages like American Sign Languages, however, Indian Sign language is still far from datadriven tasks like machine translation.To address this gap, in this paper, we introduce a new dataset CISLR (Corpus for Indian Sign Language Recognition) for word-level recognition in Indian Sign Language using videos.The corpus has a large vocabulary of around 4700 words covering different topics and domains.Further, we propose a baseline model for word recognition from sign language videos.To handle the low resource problem in the Indian Sign Language, the proposed model consists of a prototype-based one-shot learner that leverages resource-rich American Sign Language to learn generalized features for improving predictions in Indian Sign Language.Our experiments show that gesture features learned in another sign language can help perform one-shot predictions in CISLR.
Abhinav Joshi, Ashwani Bhat, Pradeep S, Priya Gole, Shashwat Gupta, Shreyansh Agarwal, Ashutosh Modi
EMNLP7
2022 Generalized Product-of-Experts for Learning Multimodal Representations in Noisy Environments
abstract
A real-world application or setting involves interaction between different modalities (e.g., video, speech, text). In order to process the multimodal information automatically and use it for an end application, Multimodal Representation Learning (MRL) has emerged as an active area of research in recent times. MRL involves learning reliable and robust representations of information from heterogeneous sources and fusing them. However, in practice, the data acquired from different sources are typically noisy. In some extreme cases, a noise of large magnitude can completely alter the semantics of the data leading to inconsistencies in the parallel multimodal data. In this paper, we propose a novel method for multimodal representation learning in a noisy environment via the generalized product of experts technique. In the proposed method, we train a separate network for each modality to assess the credibility of information coming from that modality, and subsequently, the contribution from each modality is dynamically varied while estimating the joint distribution. We evaluate our method on two challenging benchmarks from two diverse domains: multimodal 3D hand-pose estimation and multimodal surgical video segmentation. We attain state-of-the-art performance on both benchmarks. Our extensive quantitative and qualitative evaluations show the advantages of our method compared to previous approaches.
Abhinav Joshi, Jinang Shah, Binod Bhattarai, Ashutosh Modi, Danail Stoyanov
ICMI5
2022 Corpus for Automatic Structuring of Legal Documents
abstract
In populous countries, pending legal cases have been growing exponentially. There is a need for developing techniques for processing and organizing legal documents. In this paper, we introduce a new corpus for structuring legal documents. In particular, we introduce a corpus of legal judgment documents in English that are segmented into topical and coherent parts. Each of these parts is annotated with a label coming from a list of pre-defined Rhetorical Roles. We develop baseline models for automatically predicting rhetorical roles in a legal document based on the annotated corpus. Further, we show the application of rhetorical roles to improve performance on the tasks of summarization and legal judgment prediction. We release the corpus and baseline model code along with the paper.
Prathamesh Kalamkar, Aman Tiwari, Astha Agarwal, Saurabh Karn, Smita Gupta, Vivek Raghavan, Ashutosh Modi
LREC7
2022 COGMEN: COntextualized GNN based Multimodal Emotion recognitioN
abstract
Abhinav Joshi, Ashwani Bhat, Ayush Jain, Atin Singh, Ashutosh Modi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Abhinav Joshi, Ashwani Bhat, Atin Vikram Singh, Ashutosh Modi
NAACL-HLT5
2021 Fine-Grained Emotion Prediction by Modeling Emotion Definitions
abstract
In this paper, we propose a new framework for fine-grained emotion prediction in the text through emotion definition modeling. Our approach involves a multi-task learning framework that models definitions of emotions as an auxiliary task while being trained on the primary task of emotion prediction. We model definitions using masked language modeling and class definition prediction tasks. Our models outperform existing state-of-the-art for fine-grained emotion dataset GoEmotions. We further show that this trained model can be used for transfer learning on other benchmark datasets in emotion prediction with varying emotion label sets, domains, and sizes. The proposed models outperform the baselines on transfer learning experiments demonstrating the generalization capability of the models.
Gargi Singh, Dhanajit Brahma, Piyush Rai, Ashutosh Modi
ACII4
2021 ILDC for CJPE: Indian Legal Documents Corpus for Court Judgment Prediction and Explanation
abstract
Vijit Malik, Rishabh Sanjay, Shubham Kumar Nigam, Kripabandhu Ghosh, Shouvik Kumar Guha, Arnab Bhattacharya, Ashutosh Modi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Vijit Malik, Rishabh Sanjay, Shubham Kumar Nigam, Kripabandhu Ghosh, Shouvik Kumar Guha, Arnab Bhattacharya 0001, Ashutosh Modi
ACL/IJCNLP (1)7
2021 Adv-OLM: Generating Textual Adversaries via OLM
abstract
Deep learning models are susceptible to adversarial examples that have imperceptible perturbations in the original input, resulting in adversarial attacks against these models.Analysis of these attacks on the state of the art transformers in NLP can help improve the robustness of these models against such adversarial inputs.In this paper, we present Adv-OLM, a black-box attack method that adapts the idea of Occlusion and Language Models (OLM) to the current state of the art attack methods.OLM is used to rank words of a sentence, which are later substituted using word replacement strategies.We experimentally show that our approach outperforms other attack methods for several text classification tasks.
Vijit Malik, Ashwani Bhat, Ashutosh Modi
EACL3
2020 Adapting a Language Model for Controlled Affective Text Generation
abstract
Human use language not just to convey information but also to express their inner feelings and mental states. In this work, we adapt the state-of-the-art language generation models to generate affective (emotional) text. We posit a model capable of generating affect-driven and topic focused sentences without losing grammatical correctness as the affect intensity increases. We propose to incorporate emotion as prior for the probabilistic state-of-the-art text generation model such as GPT-2. The model gives a user the flexibility to control the category and intensity of emotion as well as the topic of the generated text. Previous attempts at modelling fine-grained emotions fall out on grammatical correctness at extreme intensities, but our model is resilient to this and delivers robust results at all intensities. We conduct automated evaluations and human studies to test the performance of our model, and provide a detailed comparison of the results with other models. In all evaluations, our model outperforms existing affective text generation models.
Tushar Goswamy, Ishika Singh, Ahsan Barkati, Ashutosh Modi
COLING4
2018 MCScript: A Novel Dataset for Assessing Machine Comprehension Using Script Knowledge
Simon Ostermann 0002, Ashutosh Modi, Michael Roth 0001, Stefan Thater, Manfred Pinkal
LREC2
2018 Multi-layer Annotation of the Rigveda
Oliver Hellwig, Heinrich Hettrich, Ashutosh Modi, Manfred Pinkal
LREC3
2017 Modelling Semantic Expectation: Using Script Knowledge for Referent Prediction
abstract
Recent research in psycholinguistics has provided increasing evidence that humans predict upcoming content. Prediction also affects perception and might be a key to robustness in human language processing. In this paper, we investigate the factors that affect human prediction by building a computational model that can predict upcoming discourse referents based on linguistic knowledge alone vs. linguistic knowledge jointly with common-sense knowledge in the form of scripts. We find that script knowledge significantly improves model estimates of human predictions. In a second study, we test the highly controversial hypothesis that predictability influences referring expression type but do not find evidence for such an effect.
Ashutosh Modi, Ivan Titov 0001, Vera Demberg, Asad B. Sayeed, Manfred Pinkal
Trans. Assoc. Comput. Linguistics1
2016 Event Embeddings for Semantic Script Modeling
abstract
Semantic scripts is a conceptual representation which defines how events are organized into higher level activities.Practically all the previous approaches to inducing script knowledge from text relied on count-based techniques (e.g., generative models) and have not attempted to compositionally model events.In this work, we introduce a neural network model which relies on distributed compositional representations of events.The model captures statistical dependencies between events in a scenario, overcomes some of the shortcomings of previous approaches (e.g., by more effectively dealing with data sparsity) and outperforms count-based counterparts on the narrative cloze task.
Ashutosh Modi
CoNLL1
2016 InScript: Narrative texts annotated with script information
Ashutosh Modi, Tatiana Anikina, Simon Ostermann 0002, Manfred Pinkal
LREC1
2014 Inducing Neural Models of Script Knowledge
abstract
Induction of common sense knowledge about prototypical sequence of events has recently received much attention (e.g., Chambers and Jurafsky (2008); Regneri et al. (2010)). Instead of inducing this knowledge in the form of graphs, as in much of the previous work, in our method, distributed representations of event real-izations are computed based on distributed representations of predicates and their ar-guments, and then these representations are used to predict prototypical event or-derings. The parameters of the composi-tional process for computing the event rep-resentations and the ranking component of the model are jointly estimated. We show that this approach results in a sub-stantial boost in performance on the event ordering task with respect to the previous approaches, both on natural and crowd-sourced texts. 1
Ashutosh Modi, Ivan Titov 0001
CoNLL1