VLDB 2026 Research / reviewers in the wild / expert
Somak Aditya
dblp:165/0785
· DBLP profile ↗
24ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-0113-2545ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language ModelsabstractPresent day LLMs face the challenge of managing affordance-based safety risks—situations where outputs inadvertently facilitate harmful actions due to overlooked logical implications. Traditional safety solutions, such as scalar outcome-based reward models, parameter tuning, or heuristic decoding strategies, lack the granularity and proactive nature needed to reliably detect and intervene during subtle yet crucial reasoning steps. Addressing this fundamental gap, we introduce AURA, an innovative, multi-layered framework centered around Process Reward Models (PRMs), providing comprehensive, step level evaluations across logical coherence and safety-awareness. Our framework seamlessly combines introspective self-critique, fine-grained PRM assessments, and adaptive safety-aware decoding to dynamically and proactively guide models toward safer reasoning trajectories. Empirical evidence clearly demonstrates that this approach significantly surpasses existing methods, significantly improving the logical integrity and affordance-sensitive safety of model outputs. This research represents a pivotal step toward safer, more responsible, and contextually aware AI, setting a new benchmark for alignment-sensitive applications. Sayantan Adak, Pratyush Chatterjee, Somnath Banerjee 0002, Rima Hazra, Somak Aditya, Animesh Mukherjee 0001 |
AAAI | 5 |
| 2026 | PRAGWORLD: A Benchmark Evaluating LLMs' Local World Model Under Minimal Linguistic Alterations and Conversational DynamicsabstractReal-world conversations are rich with pragmatic elements, such as entity mentions, references, and implicatures. Understanding such nuances is a requirement for successful natural communication, and often requires building a local _world model_ which encodes such elements and captures the dynamics of their evolving states. However, it is not well-understood whether language models (LMs) construct or maintain a robust implicit representation of conversations. In this work, we evaluate the ability of LMs to encode and update their internal world model in dyadic conversations and test their _malleability_ under linguistic alterations. To facilitate this, we apply seven minimal linguistic alterations to conversations sourced from popular conversational QA datasets and construct a benchmark with two variants (i.e., Manual and Synthetic) comprising yes-no questions. We evaluate nine open and one closed source LMs and observe that they struggle to maintain robust accuracy. Our analysis unveils that LMs struggle to memorize crucial details, such as tracking entities under linguistic alterations to conversations. We then propose a dual-perspective interpretability framework which identifies transformer layers that are _useful_ or _harmful_ and highlights linguistic alterations most influenced by harmful layers, typically due to encoding spurious signals or relying on shortcuts. Inspired by these insights, we propose two layer-regularization based fine-tuning strategies that suppress the effect of the harmful layers. Sachin Vashistha, Aryan Bibhuti, Atharva Naik, Martin Tutek, Somak Aditya |
AAAI | 5 |
| 2026 | AFGNN: API Misuse Detection using Graph Neural Networks and ClusteringabstractApplication Programming Interfaces (APIs) are crucial to software development, enabling integration of existing systems with new applications by reusing tried and tested code, saving development time and increasing software safety. In particular, the Java standard library APIs, along with numerous third-party APIs, are extensively utilized in the development of enterprise application software. However, their misuse remains a significant source of bugs and vulnerabilities. Furthermore, due to the limited examples in the official API documentation, developers often rely on online portals and generative AI models to learn unfamiliar APIs, but using such examples may introduce unintentional errors in the software. In this paper, we present AFGNN, a novel Graph Neural Network (GNN)-based framework for efficiently detecting API misuses in Java code. AFGNN uses a novel API Flow Graph (AFG) representation that captures the API execution sequence, data, and control flow information present in the code to model the API usage patterns. AFGNN uses self-supervised pre-training with AFG representation to effectively compute the embeddings for unknown API usage examples and cluster them to identify different usage patterns. Experiments on popular API usage datasets show that AFGNN significantly outperforms state-of-the-art small language models and API misuse detectors. Ponnampalam Pirapuraj, Tamal Mondal, Sharanya Gupta, Akash Lal, Somak Aditya, Jyothi Vedurada |
MSR | 5 |
| 2026 | AD2 : Analysis and Detection of Adversarial Threats in Visual Perception for End-to-End Autonomous Driving SystemsabstractEnd-to-end autonomous driving systems have achieved significant progress, yet their adversarial robustness remains largely underexplored. In this work, we conduct a closed-loop evaluation of state-of-the-art autonomous driving agents under black-box adversarial threat models in CARLA. Specifically, we consider three representative attack vectors on the visual perception pipeline: (i) a physics-based blur attack induced by acoustic waves, (ii) an electromagnetic interference attack that distorts captured images, and (iii) a digital attack that adds ghost objects as carefully crafted bounded perturbations on images. Our experiments on two advanced agents, Transfuser and Interfuser, reveal severe vulnerabilities to such attacks, with driving scores dropping by up to 99% in the worst case, raising valid safety concerns. To help mitigate such threats, we further propose a lightweight Attack Detection model for Autonomous Driving systems (AD2) based on attention mechanisms that capture spatial-temporal consistency. Comprehensive experiments across multi-camera inputs on CARLA show that our detector achieves superior detection capability and computational efficiency compared to existing approaches. Ishan Sahu, Somnath Hazra, Somak Aditya, Soumyajit Dey |
WACV | 3 |
| 2025 | EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture VideosabstractAs digital platforms redefine educational paradigms, ensuring interactivity remains vital for effective learning.This paper explores using Multimodal Large Language Models (MLLMs) to automatically respond to student questions from online lectures -a novel question answering task of real world significance.We introduce the EduVidQA Dataset with 5252 question-answer pairs (both synthetic and realworld) from 296 computer science videos covering diverse topics and difficulty levels.To understand the needs of the dataset and task evaluation, we empirically study the qualitative preferences of students, which we provide as an important contribution to this line of work.Our benchmarking experiments consist of 6 stateof-the-art MLLMs, through which we study the effectiveness of our synthetic data for finetuning, as well as showing the challenging nature of the task.We evaluate the models using both text-based and qualitative metrics, thus showing a nuanced perspective of the models' performance, which is paramount to future work.This work not only sets a benchmark for this important problem, but also opens exciting avenues for future research in the field of Natural Language Processing for Education. Sourjyadip Ray, Somak Aditya, Pawan Goyal 0002 |
EMNLP | 3 |
| 2025 | SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text GenerationabstractSaurabh Kumar Pandey, Sachin Vashistha, Debrup Das, Somak Aditya, Monojit Choudhury. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Saurabh Kumar Pandey, Sachin Vashistha, Debrup Das, Somak Aditya, Monojit Choudhury |
NAACL (Long Papers) | 4 |
| 2024 | Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting JailbreaksabstractRecent explorations with commercial Large Language Models (LLMs) have shown that non-expert users can jailbreak LLMs by simply manipulating their prompts; resulting in degenerate output behavior, privacy and security breaches, offensive outputs, and violations of content regulator policies. Limited studies have been conducted to formalize and analyze these attacks and their mitigations. We bridge this gap by proposing a formalism and a taxonomy of known (and possible) jailbreaks. We survey existing jailbreak methods and their effectiveness on open-source and commercial LLMs (such as GPT-based models, OPT, BLOOM, and FLAN-T5-XXL). We further discuss the challenges of jailbreak detection in terms of their effectiveness against known attacks. For further analysis, we release a dataset of model outputs across 3700 jailbreak prompts over 4 tasks. Abhinav Rao, Atharva Naik, Sachin Vashistha, Somak Aditya, Monojit Choudhury |
LREC/COLING | 4 |
| 2024 | Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMsabstractReasoning is a fundamental component of language understanding.Recent prompting techniques, such as chain of thought, have consistently improved LLMs' performance on various reasoning tasks.Nevertheless, there is still little understanding of what triggers reasoning abilities in LLMs in the inference stage.In this paper, we investigate the effect of the input representation on the reasoning abilities of LLMs.We hypothesize that representing natural language tasks as code can enhance specific reasoning abilities such as entity tracking or logical reasoning.To study this, we propose code prompting, a methodology we operationalize as a chain of prompts that transforms a natural language problem into code and directly prompts the LLM using the generated code without resorting to external code execution.We find that code prompting exhibits a high-performance boost for multiple LLMs (up to 22.52 percentage points on GPT 3.5, 7.75 on Mixtral, and 16.78 on Mistral) across multiple conditional reasoning datasets.We then conduct comprehensive experiments to understand how the code representation triggers reasoning abilities and which capabilities are elicited in the underlying models.Our analysis on GPT 3.5 reveals that the code formatting of the input problem is essential for performance improvement.Furthermore, the code representation improves sample efficiency of in-context learning and facilitates state tracking of entities.1 Haritz Puerto, Martin Tutek, Somak Aditya, Xiaodan Zhu 0001, Iryna Gurevych |
EMNLP | 3 |
| 2024 | ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital EnvironmentsabstractThe global shortage of healthcare workers has demanded the development of smart healthcare assistants, which can help monitor and alert healthcare workers when necessary.We examine the healthcare knowledge of existing Large Vision Language Models (LVLMs) via the Visual Question Answering (VQA) task in hospital settings through expert annotated open-ended questions.We introduce the Emergency Room Visual Question Answering (ERVQA) dataset, consisting of triplets covering diverse emergency room scenarios, a seminal benchmark for LVLMs.By developing a detailed error taxonomy and analyzing answer trends, we reveal the nuanced nature of the task.We benchmark state-of-the-art open-source and closed LVLMs using traditional and adapted VQA metrics: Entailment Score and CLIPScore Confidence.Analyzing errors across models, we infer trends based on properties like decoder type, model size, and in-context examples.Our findings suggest the ERVQA dataset presents a highly complex task, highlighting the need for specialized, domain-specific solutions. Sourjyadip Ray, Kushal Gupta, Soumi Kundu, Payal Arvind Kasat, Somak Aditya, Pawan Goyal 0002 |
EMNLP | 5 |
| 2024 | MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical ReasoningabstractDebrup Das, Debopriyo Banerjee, Somak Aditya, Ashish Kulkarni. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Debrup Das, Debopriyo Banerjee, Somak Aditya, Ashish Kulkarni |
NAACL-HLT | 3 |
| 2023 | Prover: Generating Intermediate Steps for NLI with Commonsense Knowledge Retrieval and Next-Step PredictionabstractDeepanway Ghosal, Somak Aditya, Monojit Choudhury. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Deepanway Ghosal, Somak Aditya, Monojit Choudhury |
IJCNLP (1) | 2 |
| 2023 | SYNC: A Structurally Guided Hard Negative Curricula for Generalizable Neural Code SearchabstractAtharva Naik, Soumitra Das, Jyothi Vedurada, Somak Aditya. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Atharva Naik, Soumitra Das, Jyothi Vedurada, Somak Aditya |
IJCNLP (1) | 4 |
| 2022 | LITMUS Predictor: An AI Assistant for Building Reliable, High-Performing and Fair Multilingual NLP SystemsabstractPre-trained multilingual language models are gaining popularity due to their cross-lingual zero-shot transfer ability, but these models do not perform equally well in all languages. Evaluating task-specific performance of a model in a large number of languages is often a challenge due to lack of labeled data, as is targeting improvements in low performing languages through few-shot learning. We present a tool - LITMUS Predictor - that can make reliable performance projections for a fine-tuned task-specific model in a set of languages without test and training data, and help strategize data labeling efforts to optimize performance and fairness objectives. Anirudh Srinivasan, Gauri Kholkar, Rahul Kejriwal, Tanuja Ganu, Sandipan Dandapat, Sunayana Sitaram, Balakrishnan Santhanam, Somak Aditya, Kalika Bali, Monojit Choudhury |
AAAI | 8 |
| 2021 | First Workshop on Knowledge Injection in Neural Networks (KINN)abstractDeep learning (DL) has made rapid progress in the last decade, with neural network-based language and vision models achieving state-of-the-art performance in various tasks. Yet purely data-driven neural network models exhibit several issues impacting real-world deployment of such models adversely. These include reliance on large quantities of training data, poor robustness, lack of generalization, poor explainability, and glaring gaps in implicit and commonsense knowledge. The availability of rich structured (or semi-structured) knowledge sources has spurred the research community into exploring Knowledge Injection in Neural Networks (KINN) as a means of mitigating the above-mentioned challenges. This has led to the development of hybrid AI systems that combine the purely data-driven learning of the neural network models with an infusion of knowledge from external sources. Such KINN systems include the development of retrieval augmented neural models, neuro symbolic systems and a plethora of combinations of NNs and knowledge graphs and structured knowledge bases. Vasudev Lal, Somak Aditya, Yezhou Yang, Pasquale Minervini, Sandya Mannarswamy |
CIKM | 2 |
| 2020 | Exploratory Navigation and Selective Reading
Natwar Modani, Paridhi Maheshwari, Harsh Deshpande, Saurab Sirpurkar, Diviya, Somak Aditya |
AAAI | 6 |
| 2020 | TaxiNLI: Taking a Ride up the NLU HillabstractPre-trained Transformer-based neural architectures have consistently achieved state-of-theart performance in the Natural Language Inference (NLI) task.Since NLI examples encompass a variety of linguistic, logical, and reasoning phenomena, it remains unclear as to which specific concepts are learnt by the trained systems and where they can achieve strong generalization.To investigate this question, we propose a taxonomic hierarchy of categories that are relevant for the NLI task.We introduce TAXINLI, a new dataset, that has 10k examples from the MNLI dataset (Williams et al., 2018) with these taxonomic labels.Through various experiments on TAXINLI, we observe that whereas for certain taxonomic categories SOTA neural models have achieved near perfect accuracies-a large jump over the previous models-some categories still remain difficult.Our work adds to the growing body of literature that shows the gaps in the current NLI systems and datasets through a systematic presentation and analysis of reasoning categories. Pratik Joshi, Somak Aditya, Aalok Sathe, Monojit Choudhury |
CoNLL | 2 |
| 2019 | Integrating Knowledge and Reasoning in Image UnderstandingabstractDeep learning based data-driven approaches have been successfully applied in various image understanding applications ranging from object recognition, semantic segmentation to visual question answering. However, the lack of knowledge integration as well as higher-level reasoning capabilities with the methods still pose a hindrance. In this work, we present a brief survey of a few representative reasoning mechanisms, knowledge integration methods and their corresponding image understanding applications developed by various groups of researchers, approaching the problem from a variety of angles. Furthermore, we discuss upon key efforts on integrating external knowledge with neural networks. Taking cues from these efforts, we conclude by discussing potential pathways to improve reasoning capabilities. Somak Aditya, Yezhou Yang, Chitta Baral |
IJCAI | 1 |
| 2019 | Spatial Knowledge Distillation to Aid Visual ReasoningabstractFor tasks involving language and vision, the current state-of-the-art methods tend not to leverage any additional information that might be present to gather relevant (commonsense) knowledge. A representative task is Visual Question Answering where large diagnostic datasets have been proposed to test a system's capability of answering questions about images. The training data is often accompanied by annotations of individual object properties and spatial locations. In this work, we take a step towards integrating this additional privileged information in the form of spatial knowledge to aid in visual reasoning. We propose a framework that combines recent advances in knowledge distillation (teacher-student framework), relational reasoning and probabilistic logical languages to incorporate such knowledge in existing neural networks for the task of Visual Question Answering. Specifically, for a question posed against an image, we use a probabilistic logical language to encode the spatial knowledge and the spatial understanding about the question in the form of a mask that is directly provided to the teacher network. The student network learns from the ground-truth information as well as the teachers prediction via distillation. We also demonstrate the impact of predicting such a mask inside the teachers network using attention. Empirically, we show that both the methods improve the test accuracy over a state-of-the-art approach on a publicly available dataset. Somak Aditya, Rudra Saha, Yezhou Yang, Chitta Baral |
WACV | 1 |
| 2018 | Explicit Reasoning over End-to-End Neural Architectures for Visual Question AnsweringabstractMany vision and language tasks require commonsense reasoning beyond data-driven image and natural language processing. Here we adopt Visual Question Answering (VQA) as an example task, where a system is expected to answer a question in natural language about an image. Current state-of-the-art systems attempted to solve the task using deep neural architectures and achieved promising performance. However, the resulting systems are generally opaque and they struggle in understanding questions for which extra knowledge is required. In this paper, we present an explicit reasoning layer on top of a set of penultimate neural network based systems. The reasoning layer enables reasoning and answering questions where additional knowledge is required, and at the same time provides an interpretable interface to the end users. Specifically, the reasoning layer adopts a Probabilistic Soft Logic (PSL) based engine to reason over a basket of inputs: visual relations, the semantic parse of the question, and background ontological knowledge from word2vec and ConceptNet. Experimental analysis of the answers and the key evidential predicates generated on the VQA dataset validate our approach. Somak Aditya, Yezhou Yang, Chitta Baral |
AAAI | 1 |
| 2018 | Combining Knowledge and Reasoning through Probabilistic Soft Logic for Image Puzzle Solving
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos |
UAI | 1 |
| 2018 | Image Understanding using vision and reasoning through Scene Description Graph
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos, Cornelia Fermüller |
Comput. Vis. Image Underst. | 1 |
| 2017 | Explainable Image Understanding Using Vision and ReasoningabstractImage Understanding is fundamental to intelligent agents.Researchers have explored Caption Generation and VisualQuestion Answering as independent aspects of Image Understanding (Johnson et al. 2015; Xiong, Merity, and Socher2016). Common to most of the successful approaches, are the learning of end-to-end signal mapping (image-to-caption, image and question to answer). The accuracy is impressive. It is also important to explain a decision to end-user(justify the results, and rectify based on feedback). Very recently, there has been some focus (Hendricks et al. 2016;Liu et al. ) on explaining some aspects of the learning systems. In my research, I look towards building explainableImage Understanding systems that can be used to generate captions and answer questions. Humans learn both from examples (learning) and by reading (knowledge). Inspired by such an intuition, researchers have constructed Knowledge-Bases that encode (probabilistic) commonsense and background knowledge. In this work, we look towards efficiently using this probabilistic knowledge on top of machine learning capabilities, to rectify noise in visual detections and generate captions or answers to posed questions. Somak Aditya |
AAAI | 1 |
| 2015 | Towards Addressing the Winograd Schema Challenge - Building and Using a Semantic Parser and a Knowledge Hunting Module
Arpit Sharma 0001, Nguyen Ha Vo, Somak Aditya, Chitta Baral |
IJCAI | 3 |
| 2015 | Recognizing Social Constructs from Textual ConversationabstractSomak Aditya, Chitta Baral, Nguyen Ha Vo, Joohyung Lee, Jieping Ye, Zaw Naung, Barry Lumpkin, Jenny Hastings, Richard Scherl, Dawn M. Sweet, Daniela Inclezan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Somak Aditya, Chitta Baral, Nguyen Ha Vo, Jieping Ye, Zaw Naung, Barry Lumpkin, Jenny Hastings, Richard B. Scherl, Dawn M. Sweet, Daniela Inclezan |
HLT-NAACL | 1 |