VLDB 2026 Research / reviewers in the wild / expert
Aryaman Arora
dblp:263/6933
· DBLP profile ↗
10ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-4977-8206ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 45% Language models and text generation · 29% Image recognition and object detection · 6% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
3.6 | 5 | 2025 | Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability · J. Mach. Learn. Res. 2025 Improved Representation Steering for Language Models · NeurIPS 2025 AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders · ICML 2025 |
Natural language and speech › Language models and text generation › model steering
representation steering |
1.7 | 2 | 2025 | Improved Representation Steering for Language Models · NeurIPS 2025 AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
causal abstraction |
0.9 | 1 | 2025 | Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability · J. Mach. Learn. Res. 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.9 | 1 | 2025 | Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability · J. Mach. Learn. Res. 2025 |
Computer vision › Image recognition and object detection › visual concept learning
concept detection |
0.9 | 1 | 2025 | AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability › representation engineering
concept steering |
0.9 | 1 | 2025 | Improved Representation Steering for Language Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation › model steering
language model steering |
0.9 | 1 | 2025 | AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
0.9 | 1 | 2025 | Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability · J. Mach. Learn. Res. 2025 |
Natural language and speech › Language models and text generation
model steering |
0.9 | 1 | 2025 | Improved Representation Steering for Language Models · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › interpretability
causal explanation |
0.8 | 1 | 2024 | CausalGym: Benchmarking causal interpretability methods on linguistic tasks · ACL (1) 2024 |
Natural language and speech › Language models and text generation › large language model
large language model adaptation |
0.8 | 1 | 2024 | ReFT: Representation Finetuning for Language Models · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.8 | 1 | 2024 | ReFT: Representation Finetuning for Language Models · NeurIPS 2024 |
Natural language and speech › Speech recognition and synthesis › pronunciation modeling
grapheme-to-phoneme conversion |
0.4 | 1 | 2020 | Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and Punjabi · ACL 2020 |
Natural language and speech › Language models and text generation
prompting |
0.3 | 1 | 2025 | Improved Representation Steering for Language Models · NeurIPS 2025 |
Natural language and speech › Information extraction and text analysis
linguistic probing |
0.2 | 1 | 2024 | CausalGym: Benchmarking causal interpretability methods on linguistic tasks · ACL (1) 2024 |
Computational social science and digital humanities
language diversity |
0.2 | 1 | 2022 | Computational Historical Linguistics and Language Diversity in South Asia · ACL (1) 2022 |
Natural language and speech › Speech recognition and synthesis
pronunciation lexicon |
0.1 | 1 | 2020 | Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and Punjabi · ACL 2020 |
Methods — techniques the papers use, named apart from their topics
sparse autoencoder · 1.7representation finetuning · 1.6preference optimization · 0.9linear probe · 0.9difference-in-means · 0.9concept suppression · 0.9causal mediation analysis · 0.9bidirectional steering · 0.9activation patching · 0.9distributed alignment search · 0.8comparative linguistics · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersabstractFine-grained steering of language model outputs is essential for safety and reliability. Prompting and finetuning are widely used to achieve these goals, but interpretability researchers have proposed a variety of representation-based techniques as well, including sparse autoencoders (SAEs), linear artificial tomography, supervised steering vectors, linear probes, and representation finetuning. At present, there is no benchmark for making direct comparisons between these proposals. Therefore, we introduce AxBench, a large-scale benchmark for steering and concept detection, and report experiments on Gemma-2-2B and 9B. For steering, we find that prompting outperforms all existing methods, followed by finetuning. For concept detection, representation-based methods such as difference-in-means, perform the best. On both evaluations, SAEs are not competitive. We introduce a novel weakly-supervised representational method (Rank-1 Representation Finetuning; ReFT-r1), which is competitive on both tasks while providing the interpretability advantages that prompting lacks. Along with AxBench, we train and publicly release SAE-scale feature dictionaries for ReFT-r1 and DiffMean. Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang 0078, Jing Huang 0014, Daniel Jurafsky, Christopher D. Manning, Christopher Potts |
ICML | 2 |
| 2025 | Improved Representation Steering for Language ModelsabstractSteering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representations to adjust behavior. Recent work has shown that adjusting weights or representations is often less effective than steering by prompting, for instance when wanting to introduce or suppress a particular concept. We demonstrate how to improve representation steering via our new Reference-free Preference Steering (RePS), a bidirectional preference-optimization objective that jointly does concept steering and suppression. We train three parameterizations of RePS and evaluate them on AxBench, a large-scale model steering benchmark. On Gemma models with sizes ranging from 2B to 27B, RePS outperforms all existing steering methods trained with a language modeling objective and substantially narrows the gap with prompting -- while promoting interpretability and minimizing parameter count. In suppression, RePS matches the language-modeling objective on Gemma-2 and outperforms it on the larger Gemma-3 variants while remaining resilient to prompt-based jailbreaking attacks that defeat prompting. Overall, our results suggest that RePS provides an interpretable and robust alternative to prompting for both steering and suppression. Zhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning, Christopher Potts |
NeurIPS | 3 |
| 2025 | Causal Abstraction: A Theoretical Foundation for Mechanistic InterpretabilityabstractCausal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level details of black box AI models. Our contributions are (1) generalizing the theory of causal abstraction from mechanism replacement (i.e., hard and soft interventions) to arbitrary mechanism transformation (i.e., functionals from old mechanisms to new mechanisms), (2) providing a flexible, yet precise formalization for the core concepts of polysemantic neurons, the linear representation hypothesis, modular features, and graded faithfulness, and (3) unifying a variety of mechanistic interpretability methods in the common language of causal abstraction, namely, activation and path patching, causal mediation analysis, causal scrubbing, causal tracing, circuit analysis, concept erasure, sparse autoencoders, differential binary masking, distributed alignment search, and steering. Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang 0014, Aryaman Arora, Zhengxuan Wu, Noah D. Goodman, Christopher Potts, Thomas Icard |
J. Mach. Learn. Res. | 7 |
| 2024 | CausalGym: Benchmarking causal interpretability methods on linguistic tasksabstractLanguage models (LMs) have proven to be powerful tools for psycholinguistic research, but most prior work has focused on purely behavioural measures (e.g., surprisal comparisons).At the same time, research in model interpretability has begun to illuminate the abstract causal mechanisms shaping LM behavior.To help bring these strands of research closer together, we introduce CausalGym.We adapt and expand the Syntax-Gym suite of tasks to benchmark the ability of interpretability methods to causally affect model behaviour.To illustrate how CausalGym can be used, we study the pythia models (14M-6.9B)and assess the causal efficacy of a wide range of interpretability methods, including linear probing and distributed alignment search (DAS).We find that DAS outperforms the other methods, and so we use it to study the learning trajectory of two difficult linguistic phenomena in pythia-1b: negative polarity item licensing and filler-gap dependencies.Our analysis shows that the mechanism implementing both of these tasks is learned in discrete stages, not gradually.https://github.com/aryamanarora/ causalgym Aryaman Arora, Daniel Jurafsky, Christopher Potts |
ACL (1) | 1 |
| 2024 | ReFT: Representation Finetuning for Language ModelsabstractParameter-efficient finetuning (PEFT) methods seek to adapt large neural models via updates to a small number of *weights*. However, much prior interpretability work has shown that *representations* encode rich semantic information, suggesting that editing representations might be a more powerful alternative. We pursue this hypothesis by developing a family of **Representation Finetuning (ReFT)** methods. ReFT methods operate on a frozen base model and learn task-specific interventions on hidden representations. We define a strong instance of the ReFT family, Low-rank Linear Subspace ReFT (LoReFT), and we identify an ablation of this method that trades some performance for increased efficiency. Both are drop-in replacements for existing PEFTs and learn interventions that are 15x--65x more parameter-efficient than LoRA. We showcase LoReFT on eight commonsense reasoning tasks, four arithmetic reasoning tasks, instruction-tuning, and GLUE. In all these evaluations, our ReFTs deliver the best balance of efficiency and performance, and almost always outperform state-of-the-art PEFTs. Upon publication, we will publicly release our generic ReFT training library. Zhengxuan Wu, Aryaman Arora, Zheng Wang 0078, Atticus Geiger, Daniel Jurafsky, Christopher D. Manning, Christopher Potts |
NeurIPS | 2 |
| 2022 | Computational Historical Linguistics and Language Diversity in South AsiaabstractSouth Asia is home to a plethora of languages, many of which severely lack access to new language technologies.This linguistic diversity also results in a research environment conducive to the study of comparative, contact, and historical linguistics-fields which necessitate the gathering of extensive data from many languages.We claim that data scatteredness (rather than scarcity) is the primary obstacle in the development of South Asian language technology, and suggest that the study of language history is uniquely aligned with surmounting this obstacle.We review recent developments in and at the intersection of South Asian NLP and historical-comparative linguistics, describing our and others' current efforts in this area.We also offer new strategies towards breaking the data barrier. Aryaman Arora, Adam Farris, Samopriya Basu, Suresh Kolichala |
ACL (1) | 1 |
| 2022 | Universal Dependencies for PunjabiabstractWe introduce the first Universal Dependencies treebank for Punjabi (written in the Gurmukhi script) and discuss corpus design and linguistic phenomena encountered in annotation. The treebank covers a variety of genres and has been annotated for POS tags, dependency relations, and graph-based Enhanced Dependencies. We aim to expand the diversity of coverage of Indo-Aryan languages in UD. Aryaman Arora |
LREC | 1 |
| 2022 | MASALA: Modelling and Analysing the Semantics of Adpositions in Linguistic Annotation of HindiabstractWe present a completed, publicly available corpus of annotated semantic relations of adpositions and case markers in Hindi. We used the multilingual SNACS annotation scheme, which has been applied to a variety of typologically diverse languages. Building on past work examining linguistic problems in SNACS annotation, we use language models to attempt automatic labelling of SNACS supersenses in Hindi and achieve results competitive with past work on English. We look towards upstream applications in semantic role labelling and extension to related languages such as Gujarati. Aryaman Arora, Nitin Venkateswaran, Nathan Schneider 0001 |
LREC | 1 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 75 |
| 2020 | Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and PunjabiabstractHindi grapheme-to-phoneme (G2P) conversion is mostly trivial, with one exception: whether a schwa represented in the orthography is pronounced or unpronounced (deleted).Previous work has attempted to predict schwa deletion in a rule-based fashion using prosodic or phonetic analysis.We present the first statistical schwa deletion classifier for Hindi, which relies solely on the orthography as the input and outperforms previous approaches.We trained our model on a newly-compiled pronunciation lexicon extracted from various online dictionaries.Our best Hindi model achieves state of the art performance, and also achieves good performance on a closely related language, Punjabi, without modification. Aryaman Arora, Luke Gessler, Nathan Schneider 0001 |
ACL | 1 |