Adithya Bhaskar

dblp:334/7656 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 53% Trustworthy machine learning · 40% Information extraction and text analysis · 7%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.522024
Finding Transformer Circuits With Edge Pruning · NeurIPS 2024
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation
alignment
0.912025
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization · ICLR 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.912025
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization · ICLR 2025
Machine learning › Trustworthy machine learning › language model interpretability
attention head analysis
0.812024
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models · ACL (1) 2024
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit discovery
0.812024
Finding Transformer Circuits With Edge Pruning · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.812024
Finding Transformer Circuits With Edge Pruning · NeurIPS 2024
Natural language and speech › Language models and text generation › pre-trained language model
pretrained language model analysis
0.812024
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation › natural language understanding
ambiguity handling
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Language models and text generation › decoding
constrained decoding
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Language models and text generation
decoding
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Language models and text generation
in-context learning
0.212024
Finding Transformer Circuits With Edge Pruning · NeurIPS 2024
Natural language and speech › Language models and text generation › linguistic generalization
syntactic generalization
0.212024
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models · ACL (1) 2024
Data models and query languages › SQL
SQL query generation
0.212023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

plan-based template generation · 1.3constrained infilling · 1.3beam search · 1.3centered hidden embedding similarity · 0.9subnetwork analysis · 0.8optimization · 0.8gradient-based pruning · 0.8edge pruning · 0.8attention head analysis · 0.8
YearPublicationVenuePosition
2025 Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
abstract
Direct Preference Optimization (DPO) and its variants are increasingly used for aligning language models with human preferences. Although these methods are designed to teach a model to generate preferred responses more frequently relative to dispreferred responses, prior work has observed that the likelihood of preferred responses often decreases during training. The current work sheds light on the causes and implications of this counterintuitive phenomenon, which we term *likelihood displacement*. We demonstrate that likelihood displacement can be *catastrophic*, shifting probability mass from preferred responses to responses with an opposite meaning. As a simple example, training a model to prefer $\texttt{No}$ over $\texttt{Never}$ can sharply increase the probability of $\texttt{Yes}$. Moreover, when aligning the model to refuse unsafe prompts, we show that such displacement can *unintentionally lead to unalignment*, by shifting probability mass from preferred refusal responses to harmful responses (e.g., reducing the refusal rate of Llama-3-8B-Instruct from 74.4% to 33.4%). We theoretically characterize that likelihood displacement is driven by preferences that induce similar embeddings, as measured by a *centered hidden embedding similarity (CHES)* score. Empirically, the CHES score enables identifying which training samples contribute most to likelihood displacement in a given dataset. Filtering out these samples effectively mitigated unintentional unalignment in our experiments. More broadly, our results highlight the importance of curating data with sufficiently distinct preferences, for which we believe the CHES score may prove valuable.
Noam Razin, Sadhika Malladi, Adithya Bhaskar, Danqi Chen 0001, Sanjeev Arora, Boris Hanin
ICLR3
2024 The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
abstract
Prior work has found that pretrained language models (LMs) fine-tuned with different random seeds can achieve similar in-domain performance but generalize differently on tests of syntactic generalization.In this work, we show that, even within a single model, we can find multiple subnetworks that perform similarly indomain, but generalize vastly differently.To better understand these phenomena, we investigate if they can be understood in terms of "competing subnetworks": the model initially represents a variety of distinct algorithms, corresponding to different subnetworks, and generalization occurs when it ultimately converges to one.This explanation has been used to account for generalization in simple algorithmic tasks ("grokking").Instead of finding competing subnetworks, we find that all subnetworkswhether they generalize or not-share a set of attention heads, which we refer to as the heuristic core.Further analysis suggests that these attention heads emerge early in training and compute shallow, non-generalizing features.The model learns to generalize by incorporating additional attention heads, which depend on the outputs of the "heuristic" heads to compute higher-level features.Overall, our results offer a more detailed picture of the mechanisms for syntactic generalization in pretrained LMs. 1
Adithya Bhaskar, Dan Friedman, Danqi Chen 0001
ACL (1)1
2024 Finding Transformer Circuits With Edge Pruning
abstract
The path to interpreting a language model often proceeds via analysis of circuits---sparse computational subgraphs of the model that capture specific aspects of its behavior. Recent work has automated the task of discovering circuits. Yet, these methods have practical limitations, as they either rely on inefficient search algorithms or inaccurate approximations. In this paper, we frame circuit discovery as an optimization problem and propose _Edge Pruning_ as an effective and scalable solution. Edge Pruning leverages gradient-based pruning techniques, but instead of removing neurons or components, prunes the _edges_ between components. Our method finds circuits in GPT-2 that use less than half the number of edges than circuits found by previous methods while being equally faithful to the full model predictions on standard circuit-finding tasks. Edge Pruning is efficient on tasks involving up to 100,000 examples, outperforming previous methods in speed and producing substantially better circuits. It also perfectly recovers the ground-truth circuits in two models compiled with Tracr. Thanks to its efficiency, we scale Edge Pruning to CodeLlama-13B, a model over 100x the size of GPT-2. We use this setting for a case study, where we compare the mechanisms behind instruction prompting and in-context learning. We find two circuits with more than 99.96% sparsity that match the performance of the full model. Further analysis reveals that the mechanisms in the two settings overlap substantially. This shows that Edge Pruning is a practical and scalable tool for interpretability, which can shed light on behaviors that only emerge in large models.
Adithya Bhaskar, Alexander Wettig, Dan Friedman, Danqi Chen 0001
NeurIPS1
2024 Performance bounds for LASSO under multiplicative LogNormal noise: Applications to pooled RT-PCR testing
Richeek Das, Aaron Jerry Ninan, Adithya Bhaskar, Ajit Rajwade 0001
Signal Process.3
2023 Benchmarking and Improving Text-to-SQL Generation under Ambiguity
abstract
Research in Text-to-SQL conversion has been largely benchmarked against datasets where each text query corresponds to one correct SQL.However, natural language queries over reallife databases frequently involve significant ambiguity about the intended SQL due to overlapping schema names and multiple confusing relationship paths.To bridge this gap, we develop a novel benchmark called AmbiQT with over 3000 examples where each text is interpretable as two plausible SQLs due to lexical and/or structural ambiguity.When faced with ambiguity, an ideal top-k decoder should generate all valid interpretations for possible disambiguation by the user (Elgohary et al., 2021;Zhong et al., 2022).We evaluate several Text-to-SQL systems and decoding algorithms, including those employing state-of-the-art LLMs, and find them to be far from this ideal.The primary reason is that the prevalent beam search algorithm and its variants, treat SQL queries as a string and produce unhelpful token-level diversity in the top-k.We propose LogicalBeam, a new decoding algorithm that navigates the SQL logic space using a blend of plan-based template generation and constrained infilling.Counterfactually generated plans diversify templates while in-filling with a beam-search, that branches solely on schema names, provides value diversity.Log-icalBeam is up to 2.5× more effective than state-of-the-art models at generating all candidate SQLs in the top-k ranked outputs.It also enhances the top-5 Exact and Execution Match Accuracies on SPIDER and Kaggle DBQA 1 .
Adithya Bhaskar, Tushar Tomar, Ashutosh Sathe, Sunita Sarawagi
EMNLP1