Abhilekh Borah

dblp:392/9232 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 33% Trustworthy machine learning · 30% Representation and self-supervised learning · 17%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.012026
DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization · AAAI 2026
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
1.012026
DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization · AAAI 2026
Machine learning › Trustworthy machine learning
fairness
1.012026
DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization · AAAI 2026
Machine learning › Trustworthy machine learning › adversarial machine learning › adversarial defense
jailbreak defense
1.012026
Always Refuse: Steering LLMs Against Jailbreaks with Contrastive Activations (Student Abstract) · AAAI 2026
Machine learning › Trustworthy machine learning
robustness
1.012026
Always Refuse: Steering LLMs Against Jailbreaks with Contrastive Activations (Student Abstract) · AAAI 2026
Computer vision › Vision and language › cross-modal alignment › image-text alignment
text-to-image alignment
1.012026
DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization · AAAI 2026
Natural language and speech › Language models and text generation
alignment
0.912025
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations · EMNLP 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture · EMNLP 2025
Machine learning › Representation and self-supervised learning › representation learning › representation geometry
latent space geometry
0.912025
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations · EMNLP 2025
Machine learning › Representation and self-supervised learning
representation analysis
0.912025
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations · EMNLP 2025
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering
0.312026
Always Refuse: Steering LLMs Against Jailbreaks with Contrastive Activations (Student Abstract) · AAAI 2026
Natural language and speech › Language models and text generation › natural language understanding
multimodal language understanding
0.312025
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

kernel methods · 1.0heavy-tailed self-regularization · 1.0divergence measure · 1.0direct preference optimization · 1.0contrastive activation steering · 1.0layer-wise pooled representations · 0.9clustering · 0.9benchmarking · 0.9
YearPublicationVenuePosition
2026 Always Refuse: Steering LLMs Against Jailbreaks with Contrastive Activations (Student Abstract)
abstract
Refusals must be resilient, not brittle.” Yet guarding refusals against adversarial phrasing and shifting user contexts remains difficult: large language models (LLMs) still yield to jailbreak prompts that evade safety filters and surface harmful content. We propose Refusal Activation Steering (RAS), a training-free, inference-time method that uses contrastive activations to shift LLM responses, biasing generation trajectories toward refusals without altering model weights. The approach is modular and domain-targetable, avoiding collateral refusals on benign queries while strengthening activation- space boundaries for unsafe content. On adversarial evaluations with an 8B instruction-tuned model, we find that steering improves refusal rate by ∼ 52% and reduces attack success rate by ∼ 40%, establishing a lightweight and interpretable safety layer for robust refusal consistency. To foster further research in this domain, we have made our implementation publicly available.
Abhilekh Borah, Chebrolu Niranjan, Kokil Jaidka
AAAI1
2026 DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization
abstract
Alignment is crucial for text-to-image (T2I) models to ensure that the generated images faithfully capture user intent while maintaining safety and fairness. Direct Preference Optimization (DPO) has emerged as a key alignment technique for large language models (LLMs), and its influence is now extending to T2I systems. This paper introduces DPO-Kernels for T2I models, a novel extension of DPO that enhances alignment across three key dimensions: (i) Hybrid Loss, which integrates embedding-based objectives with the traditional probability-based loss to improve optimization; (ii) Kernelized Representations, leveraging Radial Basis Function (RBF), Polynomial, and Wavelet kernels to enable richer feature transformations, ensuring better separation between safe and unsafe inputs; and (iii) Divergence Selection, expanding beyond DPO’s default Kullback–Leibler (KL) regularizer by incorporating alternative divergence measures such as Wasserstein and Rényi divergences to enhance stability and robustness in alignment training. We introduce DETONATE, the first large-scale benchmark of its kind, comprising approximately 100K curated image pairs, categorized as chosen and rejected. This benchmark encapsulates three critical axes of social bias and discrimination: Race, Gender, and Disability. The prompts are sourced from the hate speech datasets, while the images are generated using state-of-the-art T2I models, including Stable Diffusion 3.5 Large (SD-3.5), Stable Diffusion XL (SD-XL), and Midjourney. Furthermore, to evaluate alignment beyond surface metrics, we introduce the Alignment Quality Index (AQI) for T2I systems: a novel geometric measure that quantifies latent space separability of safe/unsafe image activations, revealing hidden model vulnerabilities. While alignment techniques often risk overfitting, we empirically demonstrate that DPO-Kernels preserve strong generalization bounds using the theory of Heavy-Tailed Self-Regularization (HT-SR).
Renjith Prasad Kaippilly Mana, Abhilekh Borah, Hasnat Md Abdullah, Chathurangi Shyalika, Ritvik Garimella, Rajarshi Roy 0007, Harshul Raj Surana, Nasrin Imanpour, Suranjana Trivedy, Amit P. Sheth, Amitava Das 0001
AAAI2
2026 ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
abstract
The rapid proliferation of Large Language Models (LLMs) has significantly contributed to the development of equitable AI systems capable of factual question-answering (QA). However, no known study tests the LLMs' robustness when presented with obfuscated versions of questions. To systematically evaluate these limitations, we propose a novel technique, ObfusQAte, and leveraging the same, introduce ObfusQA, a comprehensive, first-of-its-kind framework with multi-tiered obfuscation levels designed to examine LLM capabilities across three distinct dimensions: (i) Named-Entity Indirection, (ii) Distractor Indirection, and (iii) Contextual Overload. By capturing these fine-grained distinctions in language, ObfusQA provides a comprehensive benchmark for evaluating LLM robustness and adaptability. Our study observes that LLMs exhibit a tendency to fail or generate hallucinated responses when confronted with these increasingly nuanced variations. To foster research in this direction, we make ObfusQAte publicly available.
Shubhra Ghosh, Abhilekh Borah, Aditya Kumar Guru, Kripabandhu Ghosh
LREC2
2025 Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
abstract
Abhilekh Borah, Chhavi Sharma, Danush Khanna, Utkarsh Bhatt, Gurpreet Singh, Hasnat Md Abdullah, Raghav Kaushik Ravi, Vinija Jain, Jyoti Patel, Shubham Singh, Vasu Sharma, Arpita Vats, Rahul Raja, Aman Chadha, Amitava Das. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Abhilekh Borah, Chhavi Sharma, Danush Khanna, Utkarsh Bhatt, Hasnat Md Abdullah, Raghav Kaushik Ravi, Vinija Jain, Jyoti Patel, Vasu Sharma, Arpita Vats, Rahul Raja, Aman Chadha, Amitava Das 0001
EMNLP1
2025 DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
abstract
Arijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Nishant Mishra, Sriparna Saha. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Arijit Maji, Raghvendra Kumar 0003, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Sriparna Saha 0001
EMNLP6