Harsh Jhamtani

dblp:146/6263 · DBLP profile ↗
← Back
24ranked-venue papers
11as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 8 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Steering Large Language Models between Code Execution and Textual Reasoning
abstract
While a lot of recent research focuses on enhancing the textual reasoning capabilities of Large Language Models (LLMs) by optimizing the multi-agent framework or reasoning chains, several benchmark tasks can be solved with 100\% success through direct coding, which is more scalable and avoids the computational overhead associated with textual iterating and searching. Textual reasoning has inherent limitations in solving tasks with challenges in math, logics, optimization, and searching, which is unlikely to be solved by simply scaling up the model and data size. The recently released OpenAI GPT Code Interpreter and multi-agent frameworks such as AutoGen have demonstrated remarkable proficiency of integrating code generation and execution to solve complex tasks using LLMs. However, based on our experiments on 7 existing popular methods for steering code/text generation in both single- and multi-turn settings with 14 tasks and 6 types of LLMs (including the new O1-preview), currently there is no optimal method to correctly steer LLMs to write code when needed. We discover some interesting patterns on when models use code vs. textual reasoning with the evolution to task complexity and model sizes, which even result in an astonishingly inverse scaling behavior. We also discover that results from LLM written code are not always better than using textual reasoning, even if the task could be solved through code. To mitigate the above issues, we propose three methods to better steer LLM code/text generation and achieve a notable improvement. The costs of token lengths and runtime are thoroughly discussed for all the methods. We believe the problem of steering LLM code/text generation is critical for future research and has much space for further improvement. Project Page, Datasets, and Codes are available at https://yongchao98.github.io/CodeSteer/.
Yongchao Chen, Harsh Jhamtani, Srinagesh Sharma, Chuchu Fan
ICLR2
2024 Language-to-Code Translation with a Single Labeled Example
abstract
Kaj Bostrom, Harsh Jhamtani, Hao Fang, Sam Thomson, Richard Shin, Patrick Xia, Benjamin Van Durme, Jason Eisner, Jacob Andreas. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Kaj Bostrom, Harsh Jhamtani, Hao Fang 0002, Sam Thomson, Richard Shin, Patrick Xia 0002, Benjamin Van Durme, Jason Eisner, Jacob Andreas
EMNLP2
2024 Learning to Retrieve Iteratively for In-Context Learning
abstract
We introduce iterative retrieval, a novel framework that empowers retrievers to make iterative decisions through policy optimization.Finding an optimal portfolio of retrieved items is a combinatorial optimization problem, generally considered NP-hard.This approach provides a learned approximation to such a solution, meeting specific task requirements under a given family of large language models (LLMs).We propose a training procedure based on reinforcement learning, incorporating feedback from LLMs.We instantiate an iterative retriever for composing in-context learning (ICL) exemplars and apply it to various semantic parsing tasks that demand synthesized programs as outputs.By adding only 4M additional parameters for state encoding, we convert an offthe-shelf dense retriever into a stateful iterative retriever, outperforming previous methods in selecting ICL exemplars on semantic parsing datasets such as SMCALFLOW, TREEDST, and MTOP.Additionally, the trained iterative retriever generalizes across different inference LLMs beyond the one used during training.
Yunmo Chen, Tongfei Chen, Harsh Jhamtani, Patrick Xia 0002, Richard Shin, Jason Eisner, Benjamin Van Durme
EMNLP3
2024 Ontologically Faithful Generation of Non-Player Character Dialogues
abstract
We introduce a language generation dataset grounded in a popular video game.KNUDGE (KNowledge Constrained User-NPC Dialogue GEneration) requires models to produce trees of dialogue between video game characters that accurately reflect quest and entity specifications stated in natural language.KNUDGE is constructed from side quest dialogues drawn directly from game data of Obsidian Entertainment's The Outer Worlds, leading to real-world complexities in generation: (1) utterances must remain faithful to the game lore, including character personas and backstories; (2) a dialogue must accurately reveal new quest details to the human player; and (3) dialogues are large trees as opposed to linear chains of utterances.We report results for a set of neural generation models using supervised and in-context learning techniques; we find competent performance but room for future work addressing the challenges of creating realistic, game-quality dialogues.
Nathaniel Weir, Ryan Thomas, Randolph D'Amore, Kellie Hill, Benjamin Van Durme, Harsh Jhamtani
EMNLP6
2024 Natural Language Decomposition and Interpretation of Complex Utterances
Harsh Jhamtani, Hao Fang 0002, Patrick Xia 0002, Eran Levy, Jacob Andreas, Benjamin Van Durme
IJCAI1
2022 Achieving Conversational Goals with Unsupervised Post-hoc Knowledge Injection
abstract
Bodhisattwa Prasad Majumder, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Julian McAuley. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Bodhisattwa Prasad Majumder, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Julian J. McAuley
ACL (1)2
2022 PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification Data for Learning Enhanced Generation
abstract
A personification is a figure of speech that endows inanimate entities with properties and actions typically seen as requiring animacy. In this paper, we explore the task of personification generation. To this end, we propose PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification data for Learning Enhanced generation. We curate a corpus of personifications called PersonifCorp, together with automatically generated de-personified literalizations of these personifications. We demonstrate the usefulness of this parallel corpus by training a seq2seq model to personify a given literal input. Both automatic and human evaluations show that fine-tuning with PersonifCorp leads to significant gains in personification-related qualities such as animacy and interestingness. A detailed qualitative analysis also highlights key strengths and imperfections of PINEAPPLE over baselines, demonstrating a strong ability to generate diverse and creative personifications that enhance the overall appeal of a sentence.
Sedrick Keh, Varun Gangal, Steven Y. Feng, Harsh Jhamtani, Malihe Alikhani, Eduard H. Hovy
COLING5
2021 Truth-Conditional Captions for Time Series Data
abstract
In this paper, we explore the task of automatically generating natural language descriptions of salient patterns in a time series, such as stock prices of a company over a week.A model for this task should be able to extract high-level patterns such as presence of a peak or a dip.While typical contemporary neural models with attention mechanisms can generate fluent output descriptions for this task, they often generate factually incorrect descriptions.We propose a computational model with a truth-conditional architecture which first runs small learned programs on the input time series, then identifies the programs/patterns which hold true for the given input, and finally conditions on only the chosen valid program (rather than the input time series) to generate the output text description.A program in our model is constructed from modules, which are small neural networks that are designed to capture numerical patterns and temporal information.The modules are shared across multiple programs, enabling compositionality as well as efficient learning of module parameters.The modules, as well as the composition of the modules, are unobserved in data, and we learn them in an end-to-end fashion with the only training signal coming from the accompanying natural language text descriptions.We find that the proposed model is able to generate high-precision captions even though we consider a small and simple space of module types.
Harsh Jhamtani, Taylor Berg-Kirkpatrick
EMNLP (1)1
2021 Investigating Robustness of Dialog Models to Popular Figurative Language Constructs
abstract
Humans often employ figurative language use in communication, including during interactions with dialog systems.Thus, it is important for real-world dialog systems to be able to handle popular figurative language constructs like metaphor and simile.In this work, we analyze the performance of existing dialog models in situations where the input dialog context exhibits use of figurative language.We observe large gaps in handling of figurative language when evaluating the models on two open domain dialog datasets.When faced with dialog contexts consisting of figurative language, some models show very large drops in performance compared to contexts without figurative language.We encourage future research in dialog modeling to separately analyze and report results on figurative language in order to better test model capabilities relevant to real-world use.Finally, we propose lightweight solutions to help existing models become more robust to figurative language by simply using an external resource to translate figurative language to literal (non-figurative) forms while preserving the meaning to the best extent possible.
Harsh Jhamtani, Varun Gangal, Eduard H. Hovy, Taylor Berg-Kirkpatrick
EMNLP (1)1
2021 Formulating Neural Sentence Ordering as the Asymmetric Traveling Salesman Problem
abstract
Input s1: Our son really likes his new bike.s2: It's almost Christmas
Vishal Keswani, Harsh Jhamtani
INLG2
2020 Domain Adaptation via Context Prediction for Engineering Diagram Search
Harsh Jhamtani, Taylor Berg-Kirkpatrick
ECIR (2)1
2020 Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-Answering
abstract
Despite the rapid progress in multihop question-answering (QA), models still have trouble explaining why an answer is correct, with limited explanation training data available to learn from.To address this, we introduce three explanation datasets in which explanations formed from corpus facts are annotated.Our first dataset, eQASC, contains over 98K explanation annotations for the multihop question answering dataset QASC, and is the first that annotates multiple candidate explanations for each answer.The second dataset eQASC-perturbed is constructed by crowd-sourcing perturbations (while preserving their validity) of a subset of explanations in QASC, to test consistency and generalization of explanation prediction models.The third dataset eOBQA is constructed by adding explanation annotations to the OBQA dataset to test generalization of models trained on eQASC.We show that this data can be used to significantly improve explanation quality (+14% absolute F1 over a strong retrieval baseline) using a BERT-based classifier, but still behind the upper bound, offering a new challenge for future research.We also explore a delexicalized chain representation in which repeated noun phrases are replaced by variables, thus turning them into generalized reasoning chains (for example: "X is a Y" AND "Y has Z" IMPLIES "X has Z").We find that generalized chains maintain performance while also being more robust to certain perturbations. 1
Harsh Jhamtani, Peter Clark
EMNLP (1)1
2020 Like hiking? You probably enjoy nature: Persona-grounded Dialog with Commonsense Expansions
abstract
Existing persona-grounded dialog models often fail to capture simple implications of given persona descriptions, something which humans are able to do seamlessly.For example, state-of-the-art models cannot infer that interest in hiking might imply love for nature or longing for a break.In this paper, we propose to expand available persona sentences using existing commonsense knowledge bases and paraphrasing resources to imbue dialog models with access to an expanded and richer set of persona descriptions.Additionally, we introduce fine-grained grounding on personas by encouraging the model to make a discrete choice among persona sentences while synthesizing a dialog response.Since such a choice is not observed in the data, we model it using a discrete latent random variable and use variational learning to sample from hundreds of persona expansions.Our model outperforms competitive baselines on the PERSONA-CHAT dataset in terms of dialog quality and diversity while achieving persona-consistent and controllable dialog generation.
Bodhisattwa Prasad Majumder, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Julian J. McAuley
EMNLP (1)2
2019 Learning Rhyming Constraints using Structured Adversaries
abstract
Harsh Jhamtani, Sanket Vaibhav Mehta, Jaime Carbonell, Taylor Berg-Kirkpatrick. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Harsh Jhamtani, Sanket Vaibhav Mehta, Jaime G. Carbonell, Taylor Berg-Kirkpatrick
EMNLP/IJCNLP (1)1
2018 SPINE: SParse Interpretable Neural Embeddings
abstract
Prediction without justification has limited utility. Much of the success of neural models can be attributed to their ability to learn rich, dense and expressive representations. While these representations capture the underlying complexity and latent trends in the data, they are far from being interpretable. We propose a novel variant of denoising k-sparse autoencoders that generates highly efficient and interpretable distributed word representations (word embeddings), beginning with existing word representations from state-of-the-art methods like GloVe and word2vec. Through large scale human evaluation, we report that our resulting word embedddings are much more interpretable than the original GloVe and word2vec embeddings. Moreover, our embeddings outperform existing popular word embeddings on a diverse suite of benchmark downstream tasks.
Anant Subramanian, Danish Pruthi, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Eduard H. Hovy
AAAI3
2018 Learning to Generate Move-by-Move Commentary for Chess Games from Large-Scale Social Forum Data
abstract
Harsh Jhamtani, Varun Gangal, Eduard Hovy, Graham Neubig, Taylor Berg-Kirkpatrick. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Harsh Jhamtani, Varun Gangal, Eduard H. Hovy, Graham Neubig, Taylor Berg-Kirkpatrick
ACL (1)1
2018 Learning to Describe Differences Between Pairs of Similar Images
abstract
In this paper, we introduce the task of automatically generating text to describe the differences between two similar images.We collect a new dataset by crowd-sourcing difference descriptions for pairs of image frames extracted from video-surveillance footage.Annotators were asked to succinctly describe all the differences in a short paragraph.As a result, our novel dataset provides an opportunity to explore models that align language and vision, and capture visual salience.The dataset may also be a useful benchmark for coherent multi-sentence generation.We perform a firstpass visual analysis that exposes clusters of differing pixels as a proxy for object-level differences.We propose a model that captures visual salience by using a latent variable to align clusters of differing pixels with output sentences.We find that, for both single-sentence generation and as well as multi-sentence generation, the proposed model outperforms the models that use attention alone.
Harsh Jhamtani, Taylor Berg-Kirkpatrick
EMNLP1
2017 Generating Appealing Brand Names
Gaurush Hiranandani, Pranav Maneriker, Harsh Jhamtani
CICLing (2)3
2017 Leveraging Site Search Logs to Identify Missing Content on Enterprise Webpages
Harsh Jhamtani, Rishiraj Saha Roy, Niyati Chhaya, Eric Nyberg
ECIR1
2017 Charmanteau: Character Embedding Models For Portmanteau Creation
abstract
Portmanteaus are a word formation phenomenon where two words are combined to form a new word.We propose character-level neural sequence-tosequence (S2S) methods for the task of portmanteau generation that are end-toend-trainable, language independent, and do not explicitly use additional phonetic information.We propose a noisy-channelstyle model, which allows for the incorporation of unsupervised word lists, improving performance over a standard sourceto-target model.This model is made possible by an exhaustive candidate generation strategy specifically enabled by the features of the portmanteau task.Experiments find our approach superior to a state-of-the-art FST-based baseline with respect to ground truth accuracy and human evaluation.
Varun Gangal, Harsh Jhamtani, Graham Neubig, Eduard H. Hovy, Eric Nyberg
EMNLP2
2016 A Supervised Approach for Text Illustration
abstract
In this paper we propose a novel method to illustrate text articles with pictures from a tagged repository. Certain types of documents, like news articles, are often accompanied by a few pictures only. Prior works leverage topics or key phrases from the text to suggest relevant pictures. We propose a supervised model based on features like readability, picturability, sentiment polarity, and presence of important phrases, to identify and rank key sentences. The proposed method then suggests some relevant pictures based on the top ranked sentences thus identified.
Harsh Jhamtani, Shubham Varma, Midhun Gundapuneni, Siddhartha Kumar Dutta
ACM Multimedia1
2016 Generating Multiple Diverse Summaries
Natwar Modani, Balaji Vasan Srinivasan, Harsh Jhamtani
WISE (1)3
2014 Identifying Purchase Intent from Social Posts
Vineet Gupta 0002, Devesh Varshney, Harsh Jhamtani, Deepam Kedia, Shweta Karwa
ICWSM3
2014 Word-level Language Identification in Bi-lingual Code-switched Texts
Harsh Jhamtani, Suleep Kumar Bhogi, Vaskar Raychoudhury
PACLIC1