Shikhar Murty

dblp:202/2040 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations
abstract
Large Language Models (LLMs) exhibit a robust mastery of syntax when processing and generating text.While this suggests internalized understanding of hierarchical syntax and dependency relations, the precise mechanism by which they represent syntactic structure is an open area within interpretability research.Probing provides one way to identify syntactic mechanisms linearly encoded in activations; however, no comprehensive study has yet established whether a model's probing accuracy reliably predicts its downstream syntactic performance.Adopting a "mechanisms vs. outcomes" framework, we evaluate 32 open-weight transformer models and find that syntactic features extracted via probing fail to predict outcomes of targeted syntax evaluations across English linguistic phenomena.Our results highlight a substantial disconnect between latent syntactic representations found via probing and observable syntactic behaviors in downstream tasks.
Ananth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar Murty
EMNLP4
2025 MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
abstract
Models that rely on subword tokenization have significant drawbacks, such as sensitivity to character-level noise like spelling errors and inconsistent compression rates across different languages and scripts. While character- or byte-level models like ByT5 attempt to address these concerns, they have not gained widespread adoption—processing raw byte streams without tokenization results in significantly longer sequence lengths, making training and inference inefficient. This work introduces MrT5 (MergeT5), a more efficient variant of ByT5 that integrates a token deletion mechanism in its encoder to dynamically shorten the input sequence length. After processing through a fixed number of encoder layers, a learned delete gate determines which tokens are to be removed and which are to be retained for subsequent layers. MrT5 effectively "merges" critical information from deleted tokens into a more compact sequence, leveraging contextual information from the remaining tokens. In continued pre-training experiments, we find that MrT5 can achieve significant gains in inference runtime with minimal effect on performance, as measured by bits-per-byte. Additionally, with multilingual training, MrT5 adapts to the orthographic characteristics of each language, learning language-specific compression rates. Furthermore, MrT5 shows comparable accuracy to ByT5 on downstream evaluations such as XNLI, TyDi QA, and character-level tasks while reducing sequence lengths by up to 75%. Our approach presents a solution to the practical limitations of existing byte-level models.
Julie Kallini, Shikhar Murty, Christopher D. Manning, Christopher Potts, Róbert Csordás
ICLR2
2025 Sneaking Syntax into Transformer Language Models with Tree Regularization
abstract
Ananjan Nandi, Christopher D Manning, Shikhar Murty. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ananjan Nandi, Christopher D. Manning, Shikhar Murty
NAACL (Long Papers)3
2024 BAGEL: Bootstrapping Agents by Guiding Exploration with Language
abstract
Following natural language instructions by executing actions in digital environments (e.g. web-browsers and REST APIs) is a challenging task for language model (LM) agents. Unfortunately, LM agents often fail to generalize to new environments without human demonstrations. This work presents BAGEL, a method for bootstrapping LM agents without human supervision. BAGEL converts a seed set of randomly explored trajectories to synthetic demonstrations via round-trips between two noisy LM components: an LM labeler which converts a trajectory into a synthetic instruction, and a zero-shot LM agent which maps the synthetic instruction into a refined trajectory. By performing these round-trips iteratively, BAGEL quickly converts the initial distribution of trajectories towards those that are well-described by natural language. We adapt the base LM agent at test time with in-context learning by retrieving relevant BAGEL demonstrations based on the instruction, and find improvements of over 2-13% absolute on ToolQA and MiniWob++, with up to 13x reduction in execution failures.
Shikhar Murty, Christopher D. Manning, Peter Shaw 0004, Mandar Joshi, Kenton Lee
ICML1
2023 Pushdown Layers: Encoding Recursive Structure in Transformer Language Models
abstract
Recursion is a prominent feature of human language, and fundamentally challenging for self-attention due to the lack of an explicit recursive-state tracking mechanism.Consequently, Transformer language models poorly capture long-tail recursive structure and exhibit sample-inefficient syntactic generalization.This work introduces Pushdown Layers, a new self-attention layer that models recursive state via a stack tape that tracks estimated depths of every token in an incremental parse of the observed prefix.Transformer LMs with Pushdown Layers are syntactic language models that autoregressively and synchronously update this stack tape as they predict new tokens, in turn using the stack tape to softly modulate attention over tokens-for instance, learning to "skip" over closed constituents.When trained on a corpus of strings annotated with silver constituency parses, Transformers equipped with Pushdown Layers achieve dramatically better and 3-5x more sample-efficient syntactic generalization, while maintaining similar perplexities.Pushdown Layers are a drop-in replacement for standard self-attention.We illustrate this by finetuning GPT2-medium with Pushdown Layers on an automatically parsed WikiText-103, leading to improvements on several GLUE text classification tasks.
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. Manning
EMNLP1
2023 Characterizing intrinsic compositionality in transformers with Tree Projections
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. Manning
ICLR1
2022 Fixing Model Bugs with Natural Language Patches
abstract
Current approaches for fixing systematic problems in NLP models (e.g., regex patches, finetuning on more data) are either brittle, or labor-intensive and liable to shortcuts.In contrast, humans often provide corrections to each other through natural language.Taking inspiration from this, we explore natural language patches-declarative statements that allow developers to provide corrective feedback at the right level of abstraction, either overriding the model ("if a review gives 2 stars, the sentiment is negative") or providing additional information the model may lack ("if something is described as the bomb, then it is good").We model the task of determining if a patch applies separately from the task of integrating patch information, and show that with a small amount of synthetic data, we can teach models to effectively use real patches on real data-1 to 7 patches improve accuracy by ~1-4 accuracy points on different slices of a sentiment analysis dataset, and F1 by 7 points on a relation extraction dataset.Finally, we show that finetuning on as many as 100 labeled examples may be needed to match the performance of a small set of language patches.
Shikhar Murty, Christopher D. Manning, Scott M. Lundberg, Marco Túlio Ribeiro
EMNLP1
2022 On Measuring the Intrinsic Few-Shot Hardness of Datasets
abstract
While advances in pre-training have led to dramatic improvements in few-shot learning of NLP tasks, there is limited understanding of what drives successful few-shot adaptation in datasets.In particular, given a new dataset and a pre-trained model, what properties of the dataset make it few-shot learnable and are these properties independent of the specific adaptation techniques used?We consider an extensive set of recent few-shot learning methods, and show that their performance across a large number of datasets is highly correlated, showing that few-shot hardness may be intrinsic to datasets, for a given pre-trained model.To estimate intrinsic few-shot hardness, we then propose a simple and lightweight metric called Spread that captures the intuition that fewshot learning is made possible by exploiting feature-space invariances between training and test samples.Our metric better accounts for few-shot hardness compared to existing notions of hardness, and is ~8-100x faster to compute. ⋆ Equal ContributionMethod D1 D2 LMBFF 45.3 -0.4 NullPrompts 43.0 -5.7 BitFit 46.3 -3.5 AdaPET 44.9 -0.3 P-Tuning 46.3 0.3 Few-shot (Avg) 45.2 -2 Full Fine-tuning 45.3 35
Shikhar Murty, Christopher D. Manning
EMNLP2
2021 DReCa: A General Task Augmentation Strategy for Few-Shot Natural Language Inference
abstract
Shikhar Murty, Tatsunori B. Hashimoto, Christopher Manning. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Shikhar Murty, Tatsunori B. Hashimoto, Christopher D. Manning
NAACL-HLT1
2020 ExpBERT: Representation Engineering with Natural Language Explanations
abstract
Suppose we want to specify the inductive bias that married couples typically go on honeymoons for the task of extracting pairs of spouses from text.In this paper, we allow model developers to specify these types of inductive biases as natural language explanations.We use BERT fine-tuned on MultiNLI to "interpret" these explanations with respect to the input sentence, producing explanationguided representations of the input.Across three relation extraction tasks, our method, ExpBERT, matches a BERT baseline but with 3-20× less labeled data and improves on the baseline by 3-10 F1 points with the same amount of labeled data.
Shikhar Murty, Pang Wei Koh, Percy Liang
ACL1
2019 Systematic Generalization: What Is Required and Can It Be Learned?
Dzmitry Bahdanau, Shikhar Murty, Michael Noukhovitch, Thien Huu Nguyen, Harm de Vries, Aaron C. Courville
ICLR (Poster)2
2018 Probabilistic Embedding of Knowledge Graphs with Box Lattice Measures
abstract
Embedding methods which enforce a partial order or lattice structure over the concept space, such as Order Embeddings (OE) (Vendrov et al., 2016), are a natural way to model transitive relational data (e.g.entailment graphs).However, OE learns a deterministic knowledge base, limiting expressiveness of queries and the ability to use uncertainty for both prediction and learning (e.g.learning from expectations).Probabilistic extensions of OE (Lai and Hockenmaier, 2017) have provided the ability to somewhat calibrate these denotational probabilities while retaining the consistency and inductive bias of ordered models, but lack the ability to model the negative correlations found in real-world knowledge.In this work we show that a broad class of models that assign probability measures to OE can never capture negative correlation, which motivates our construction of a novel box lattice and accompanying probability measure to capture anticorrelation and even disjoint concepts, while still providing the benefits of probabilistic modeling, such as the ability to perform rich joint and conditional queries over arbitrary sets of concepts, and both learning from and predicting calibrated uncertainty.We show improvements over previous approaches in modeling the Flickr and WordNet entailment graphs, and investigate the power of the model. * Equal contribution.
Luke Vilnis, Xiang Li 0069, Shikhar Murty, Andrew McCallum
ACL (1)3
2018 Hierarchical Losses and New Resources for Fine-grained Entity Typing and Linking
abstract
Extraction from raw text to a knowledge base of entities and fine-grained types is often cast as prediction into a flat set of entity and type labels, neglecting the rich hierarchies over types and entities contained in curated ontologies.Previous attempts to incorporate hierarchical structure have yielded little benefit and are restricted to shallow ontologies.This paper presents new methods using real and complex bilinear mappings for integrating hierarchical information, yielding substantial improvement over flat predictions in entity linking and fine-grained entity typing, and achieving new state-of-the-art results for end-to-end models on the benchmark FIGER dataset.We also present two new human-annotated datasets containing wide and deep hierarchies which we will release to the community to encourage further research in this direction: MedMentions, a collection of PubMed abstracts in which 246k mentions have been mapped to the massive UMLS ontology; and Type-Net, which aligns Freebase types with the WordNet hierarchy to obtain nearly 2k entity types.In experiments on all three datasets we show substantial gains from hierarchy-aware training.
Shikhar Murty, Patrick Verga, Luke Vilnis, Irena Radovanovic, Andrew McCallum
ACL (1)1
2018 Embedded-State Latent Conditional Random Fields for Sequence Labeling
abstract
Complex textual information extraction tasks are often posed as sequence labeling or shallow parsing, where fields are extracted using local labels made consistent through probabilistic inference in a graphical model with constrained transitions.Recently, it has become common to locally parametrize these models using rich features extracted by recurrent neural networks (such as LSTM), while enforcing consistent outputs through a simple linear-chain model, representing Markovian dependencies between successive labels.However, the simple graphical model structure belies the often complex non-local constraints between output labels.For example, many fields, such as a first name, can only occur a fixed number of times, or in the presence of other fields.While RNNs have provided increasingly powerful context-aware local features for sequence tagging, they have yet to be integrated with a global graphical model of similar expressivity in the output distribution.Our model goes beyond the linear chain CRF to incorporate multiple hidden states per output label, but parametrizes their transitions parsimoniously with low-rank logpotential scoring matrices, effectively learning an embedding space for hidden states.This augmented latent space of inference variables complements the rich feature representation of the RNN, and allows exact global inference obeying complex, learned non-local output constraints.We experiment with several datasets and show that the model outperforms baseline CRF+RNN models when global output constraints are necessary at inference-time, and explore the interpretable latent structure.
Dung Thai, Sree Harsha Ramesh, Shikhar Murty, Luke Vilnis, Andrew McCallum
CoNLL3
2018 Mitigating the Effect of Out-of-Vocabulary Entity Pairs in Matrix Factorization for KB Inference
abstract
This paper analyzes the varied performance of Matrix Factorization (MF) on the related tasks of relation extraction and knowledge-base completion, which have been unified recently into a single framework of knowledge-base inference (KBI) [Toutanova et al., 2015]. We first propose a new evaluation protocol that makes comparisons between MF and Tensor Factorization (TF) models fair. We find that this results in a steep drop in MF performance. Our analysis attributes this to the high out-of-vocabulary (OOV) rate of entity pairs in test folds of commonly-used datasets. To alleviate this issue, we propose three extensions to MF. Our best model is a TF-augmented MF model. This hybrid model is robust and obtains strong results across various KBI datasets.
Prachi Jain 0001, Shikhar Murty, Mausam, Soumen Chakrabarti
IJCAI2