Abhinav Joshi

dblp:308/0603 · DBLP profile ↗
← Back
13ranked-venue papers
12as first author
13since 2021 · last 2025
0000-0001-6756-1126ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 10 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Calibration Across Layers: Understanding Calibration Evolution in LLMs
abstract
Large Language Models (LLMs) have demonstrated inherent calibration capabilities, where predicted probabilities align well with correctness, despite prior findings that deep neural networks are often overconfident.Recent studies have linked this behavior to specific components in the final layer, such as entropy neurons and the unembedding matrix's null space.In this work, we provide a complementary perspective by investigating how calibration evolves throughout the network's depth.Analyzing multiple open-weight models on the MMLU benchmark, we uncover a distinct confidence correction phase in the upper/later layers, where model confidence is actively recalibrated after decision certainty has been reached.Furthermore, we identify a low-dimensional calibration direction in the residual stream whose perturbation significantly improves calibration metrics (ECE and MCE) without harming accuracy.Our findings suggest that calibration is a distributed phenomenon, shaped throughout the network's forward pass, not just in its final projection, providing new insights into how confidence-regulating mechanisms operate within LLMs.
Abhinav Joshi, Areeb Ahmad, Ashutosh Modi
EMNLP1
2025 PoseStitch-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
abstract
Sign language translation remains a challenging task due to the scarcity of large-scale, sentence-aligned datasets.Prior arts have focused on various feature extraction and architectural changes to support neural machine translation for sign languages.We propose POSESTITCH-SLT, a novel pre-training scheme that is inspired by linguistic-templatesbased sentence generation technique.With translation comparison on two sign language datasets, How2Sign and iSign, we show that a simple transformer-based encoder-decoder architecture outperforms the prior art when considering template-generated sentence pairs in training.We achieve BLEU-4 score improvements from 1.97 to 4.56 on How2Sign and from 0.55 to 3.43 on iSign, surpassing prior state-ofthe-art methods for pose-based gloss-free translation.The results demonstrate the effectiveness of template-driven synthetic supervision in low-resource sign language settings.
Abhinav Joshi, Sanjeet Singh, Ashutosh Modi
EMNLP1
2025 Towards Quantifying Commonsense Reasoning with Mechanistic Insights
abstract
Abhinav Joshi, Areeb Ahmad, Divyaksh Shukla, Ashutosh Modi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Abhinav Joshi, Areeb Ahmad, Divyaksh Shukla, Ashutosh Modi
NAACL (Long Papers)1
2025 Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
abstract
Transformer-based language models exhibit complex and distributed behavior, yet their internal computations remain poorly understood. Existing mechanistic interpretability methods typically treat attention heads and multilayer perceptron layers (MLPs) (the building blocks of a transformer architecture) as indivisible units, overlooking possibilities of functional substructure learned within them. In this work, we introduce a more fine-grained perspective that decomposes these components into orthogonal singular directions, revealing superposed and independent computations within a single head or MLP. We validate our perspective on widely used standard tasks like Indirect Object Identification (IOI), Gender Pronoun (GP), and Greater Than (GT), showing that previously identified canonical functional heads, such as the “name mover,” encode multiple overlapping subfunctions aligned with distinct singular directions. Nodes in a computational graph, that are previously identified as circuit elements show strong activation along specific low-rank directions, suggesting that meaningful computations reside in compact subspaces. While some directions remain challenging to interpret fully, our results highlight that transformer computations are more distributed, structured, and compositional than previously assumed. This perspective opens new avenues for fine-grained mechanistic interpretability and a deeper understanding of model internals.
Areeb Ahmad, Abhinav Joshi, Ashutosh Modi
NeurIPS2
2025 Geometry of Decision Making in Language Models
abstract
Large Language Models (LLMs) show strong generalization across diverse tasks, yet the internal decision-making processes behind their predictions remain opaque. In this work, we study the geometry of hidden representations in LLMs through the lens of intrinsic dimension (ID), focusing specifically on decision-making dynamics in a multiple-choice question answering (MCQA) setting. We perform a large-scale study, with 28 open-weight transformer models and estimate ID across layers using multiple estimators, while also quantifying per-layer performance on MCQA tasks. Our findings reveal a consistent ID pattern across models: early layers operate on low-dimensional manifolds, middle layers expand this space, and later layers compress it again, converging to decision-relevant representations. Together, these results suggest LLMs implicitly learn to project linguistic inputs onto structured, low-dimensional manifolds aligned with task-specific decisions, providing new geometric insights into how generalization and reasoning emerge in language models.
Abhinav Joshi, Divyanshu Bhatt, Ashutosh Modi
NeurIPS1
2024 IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
abstract
Abhinav Joshi, Shounak Paul, Akshat Sharma, Pawan Goyal, Saptarshi Ghosh, Ashutosh Modi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Abhinav Joshi, Shounak Paul, Akshat Sharma, Pawan Goyal 0002, Saptarshi Ghosh 0001, Ashutosh Modi
ACL (1)1
2024 COLD: Causal reasOning in cLosed Daily activities
abstract
Large Language Models (LLMs) have shown state-of-the-art performance in a variety of tasks, including arithmetic and reasoning; however, to gauge the intellectual capabilities of LLMs, causal reasoning has become a reliable proxy for validating a general understanding of the mechanics and intricacies of the world similar to humans. Previous works in natural language processing (NLP) have either focused on open-ended causal reasoning via causal commonsense reasoning (CCR) or framed a symbolic representation-based question answering for theoretically backed-up analysis via a causal inference engine. The former adds an advantage of real-world grounding but lacks theoretically backed-up analysis/validation, whereas the latter is far from real-world grounding. In this work, we bridge this gap by proposing the COLD (Causal reasOning in cLosed Daily activities) framework, which is built upon human understanding of daily real-world activities to reason about the causal nature of events. We show that the proposed framework facilitates the creation of enormous causal queries (∼ 9 million) and comes close to the mini-turing test, simulating causal reasoning to evaluate the understanding of a daily real-world task. We evaluate multiple LLMs on the created causal queries and find that causal reasoning is challenging even for activities trivial to humans. We further explore (the causal reasoning abilities of LLMs) using the backdoor criterion to determine the causal strength between events.
Abhinav Joshi, Areeb Ahmad, Ashutosh Modi
NeurIPS1
2023 U-CREAT: Unsupervised Case Retrieval using Events extrAcTion
abstract
The task of Prior Case Retrieval (PCR) in the legal domain is about automatically citing relevant (based on facts and precedence) prior legal cases in a given query case.To further promote research in PCR, in this paper, we propose a new large benchmark (in English) for the PCR task: IL-PCR (Indian Legal Prior Case Retrieval) corpus.Given the complex nature of case relevance and the long size of legal documents, BM25 remains a strong baseline for ranking the cited prior documents.In this work, we explore the role of events in legal case retrieval and propose an unsupervised retrieval method-based pipeline U-CREAT (Unsupervised Case Retrieval using Events Extraction).We find that the proposed unsupervised retrieval method significantly increases performance compared to BM25 and makes retrieval faster by a considerable margin, making it applicable to real-time case retrieval systems.Our proposed system is generic, we show that it generalizes across two different legal systems (Indian and Canadian), and it shows state-ofthe-art performance on the benchmarks for both the legal systems (IL-PCR and COLIEE corpora).
Abhinav Joshi, Akshat Sharma, Sai Kiran Tanikella, Ashutosh Modi
ACL (1)1
2023 ScriptWorld: Text Based Environment for Learning Procedural Knowledge
abstract
Text-based games provide a framework for developing natural language understanding and commonsense knowledge about the world in reinforcement learning based agents. Existing text-based environments often rely on fictional situations and characters to create a gaming framework and are far from real-world scenarios. In this paper, we introduce ScriptWorld: a text-based environment for teaching agents about real-world daily chores and hence imparting commonsense knowledge. To the best of our knowledge, it is the first interactive text-based gaming framework that consists of daily real-world human activities designed using scripts dataset. We provide gaming environments for 10 daily activities and perform a detailed analysis of the proposed environment. We develop RL-based baseline models/agents to play the games in ScriptWorld. To understand the role of language models in such environments, we leverage features obtained from pre-trained language models in the RL agents. Our experiments show that prior knowledge obtained from a pre-trained language model helps to solve real-world text-based gaming environments.
Abhinav Joshi, Areeb Ahmad, Umang Pandey, Ashutosh Modi
IJCAI1
2022 CISLR: Corpus for Indian Sign Language Recognition
abstract
Indian Sign Language, though used by a diverse community, still lacks well-annotated resources for developing systems that would enable sign language processing.In recent years researchers have actively worked for sign languages like American Sign Languages, however, Indian Sign language is still far from datadriven tasks like machine translation.To address this gap, in this paper, we introduce a new dataset CISLR (Corpus for Indian Sign Language Recognition) for word-level recognition in Indian Sign Language using videos.The corpus has a large vocabulary of around 4700 words covering different topics and domains.Further, we propose a baseline model for word recognition from sign language videos.To handle the low resource problem in the Indian Sign Language, the proposed model consists of a prototype-based one-shot learner that leverages resource-rich American Sign Language to learn generalized features for improving predictions in Indian Sign Language.Our experiments show that gesture features learned in another sign language can help perform one-shot predictions in CISLR.
Abhinav Joshi, Ashwani Bhat, Pradeep S, Priya Gole, Shashwat Gupta, Shreyansh Agarwal, Ashutosh Modi
EMNLP1
2022 Multimodal Representation Learning For Real-World Applications
abstract
Multimodal representation learning has shown tremendous improvements in recent years. An extensive set of works for fusing multiple modalities have shown promising results on the public benchmarks. However, most famous works target unrealistic settings or toy datasets, and a considerable gap exists between the real-world implications of the existing methods. In this work, we aim to bridge the gap between the well-defined benchmark settings and the real-world use cases. We aim to explore architectures inspired by existing promising approaches that have the potential to be implemented in real-world instances. Moreover, we also try to move the research forward by addressing questions that can be solved using multimodal approaches and have a considerable impact on the community. With this work, we attempt to leverage the multimodal representation learning methods, which directly apply to real-world settings.
Abhinav Joshi
ICMI1
2022 Generalized Product-of-Experts for Learning Multimodal Representations in Noisy Environments
abstract
A real-world application or setting involves interaction between different modalities (e.g., video, speech, text). In order to process the multimodal information automatically and use it for an end application, Multimodal Representation Learning (MRL) has emerged as an active area of research in recent times. MRL involves learning reliable and robust representations of information from heterogeneous sources and fusing them. However, in practice, the data acquired from different sources are typically noisy. In some extreme cases, a noise of large magnitude can completely alter the semantics of the data leading to inconsistencies in the parallel multimodal data. In this paper, we propose a novel method for multimodal representation learning in a noisy environment via the generalized product of experts technique. In the proposed method, we train a separate network for each modality to assess the credibility of information coming from that modality, and subsequently, the contribution from each modality is dynamically varied while estimating the joint distribution. We evaluate our method on two challenging benchmarks from two diverse domains: multimodal 3D hand-pose estimation and multimodal surgical video segmentation. We attain state-of-the-art performance on both benchmarks. Our extensive quantitative and qualitative evaluations show the advantages of our method compared to previous approaches.
Abhinav Joshi, Jinang Shah, Binod Bhattarai, Ashutosh Modi, Danail Stoyanov
ICMI1
2022 COGMEN: COntextualized GNN based Multimodal Emotion recognitioN
abstract
Abhinav Joshi, Ashwani Bhat, Ayush Jain, Atin Singh, Ashutosh Modi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Abhinav Joshi, Ashwani Bhat, Atin Vikram Singh, Ashutosh Modi
NAACL-HLT1