Jordan Henkel

dblp:217/1922 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0003-3862-249XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Developer Prompts in Practice: An Empirical Study of Bias, Security, and Optimization
abstract
Background: Modern software increasingly relies on Developer Prompts (Dev Prompts)-snippets of natural language embedded directly in source code-to leverage the capabilities of Large Language Models (LLMs) for tasks like classification, summarization, and content generation. Yet, despite the rapid adoption of LLMs and Dev Prompts, it remains unclear to what extent these prompts unintentionally encode biases, invite injection attacks, or underperform due to sub-optimal phrasing. Aims: To address this gap, we present a large-scale empirical analysis of Dev Prompts found in real open-source software projects to assess the prevalence of bias, security vulnerabilities, and performance issues. Then, we propose and validate approaches to mitigate these issues, and demonstrate the practical feasibility of addressing them. Method: We systematically sampled 2,320 Dev Prompts from a set of 40,573 found in open-source software projects, to identify the prevalence of the aforementioned issues. We also implemented a lightweight tool that automatically rewrites flawed prompts. Results: We find evidence of easy-to-fix issues across multiple dimensions: 3.46% of prompts contain language likely to lead to biased model responses, while over 10.75% are vulnerable to straightforward injection attacks, and we posit that many more are amenable to performance improvement through minor adjustments. Our prototype successfully mitigated bias in 68.29% of cases, prevented injection vulnerabilities in 41.81%, and improved performance in 37.1% of tested prompts. Conclusions: Our findings highlight an urgent need for future research and dedicated tool-support to help software developers write safer, fairer, and more effective prompts. To facilitate ongoing work in this emerging area, we share our data and analysis infrastructure publicly. We encourage the community to further explore the implications of Dev Prompts in modern software.
Dhia Elhaq Rzig, Dhruba Jyothi Paul, Kaiser Pister, Jordan Henkel, Foyzul Hassan
ESEM4
2024 NL2SQL is a solved problem... Not!
Avrilia Floratou, Fotis Psallidas, Fuheng Zhao, Shaleen Deep, Gunther Hagleither, Wangda Tan, Joyce Cahoon, Rana Alotaibi, Jordan Henkel, Abhik Singla, Alex Van Grootel, Brandon Chow, Katherine Lin, Marcos Campos, K. Venkatesh Emani, Vivek Pandit, Victor Shnayder, Wenjing Wang 0005, Carlo Curino
CIDR9
2024 ReAcTable: Enhancing ReAct for Table Question Answering
abstract
Table Question Answering (TQA) presents a substantial challenge at the intersection of natural language processing and data analytics. This task involves answering natural language (NL) questions on top of tabular data, demanding proficiency in logical reasoning, understanding of data semantics, and fundamental analytical capabilities. Due to its significance, a substantial volume of research has been dedicated to exploring a wide range of strategies aimed at tackling this challenge including approaches that leverage Large Language Models (LLMs) through in-context learning or Chain-of-Thought (CoT) prompting as well as approaches that train and fine-tune custom models. Nonetheless, a conspicuous gap exists in the research landscape, where there is limited exploration of how innovative foundational research, which integrates incremental reasoning with external tools in the context of LLMs, as exemplified by the ReAct paradigm, could potentially bring advantages to the TQA task. In this paper, we aim to fill this gap, by introducing ReAcTable ( ReAct for Table Question Answering tasks), a framework inspired by the ReAct paradigm that is carefully enhanced to address the challenges uniquely appearing in TQA tasks such as interpreting complex data semantics, dealing with errors generated by inconsistent data and generating intricate data transformations. ReAcTable relies on external tools such as SQL and Python code executors, to progressively enhance the data by generating intermediate data representations, ultimately transforming it into a more accessible format for answering the user's questions with greater ease. Through extensive empirical evaluations using three popular TQA benchmarks, we demonstrate that ReAcTable achieves remarkable performance even when compared to fine-tuned approaches. In particular, it outperforms the best prior result on the WikiTQ benchmark by 2.1%, achieving an accuracy of 68.0% without requiring training a new model or fine-tuning.
Jordan Henkel, Avrilia Floratou, Joyce Cahoon, Shaleen Deep, Jignesh M. Patel
Proc. VLDB Endow.2
2022 Semantic Robustness of Models of Source Code
abstract
Deep neural networks are vulnerable to adversarial examples-small input perturbations that result in incorrect predictions. We study this problem for models of source code, where we want the neural network to be robust to source-code modifications that preserve code functionality. To facilitate training robust models, we define a powerful and generic adversary that can employ sequences of parametric, semantics-preserving program transformations. We then explore how, with such an adversary, one can train models that are robust to adversarial program transformations. We conduct a thorough evaluation of our approach and find several surprising facts: we find robust training to beat dataset augmentation in every evaluation we performed; we find that a state-of-the-art architecture (code2seq) for models of code is harder to make robust than a simpler baseline; additionally, we find code2seq to have surprising weaknesses not present in our simpler baseline model; finally, we find that robust models perform better against unseen data from different sources (as one might hope)-however, we also find that robust models are not clearly better in the cross-language transfer task. To the best of our knowledge, we are the first to study the interplay between robustness of models of code and the domain-adaptation and cross-language- transfer tasks.
Jordan Henkel, Goutham Ramakrishnan, Zi Wang 0016, Aws Albarghouthi, Somesh Jha, Thomas W. Reps
SANER1
2021 Shipwright: A Human-in-the-Loop System for Dockerfile Repair
Jordan Henkel, Denini Silva, Leopoldo Teixeira, Marcelo d'Amorim, Thomas W. Reps
ICSE1
2020 Learning from, understanding, and supporting DevOps artifacts for docker
abstract
With the growing use of DevOps tools and frameworks, there is an increased need for tools and techniques that support more than code. The current state-of-the-art in static developer assistance for tools like Docker is limited to shallow syntactic validation. We identify three core challenges in the realm of learning from, understanding, and supporting developers writing DevOps artifacts: (i) nested languages in DevOps artifacts, (ii) rule mining, and (iii) the lack of semantic rule-based analysis. To address these challenges we introduce a toolset, binnacle, that enabled us to ingest 900,000 GitHub repositories.
Jordan Henkel, Christian Bird, Shuvendu K. Lahiri, Thomas W. Reps
ICSE1
2020 A Dataset of Dockerfiles
abstract
Dockerfiles are one of the most prevalent kinds of DevOps artifacts used in industry. Despite their prevalence, there is a lack of sophisticated semantics-aware static analysis of Dockerfiles. In this paper, we introduce a dataset of approximately 178,000 unique Dockerfiles collected from GitHub. To enhance the usability of this data, we describe five representations we have devised for working with, mining from, and analyzing these Dockerfiles. Each Dockerfile representation builds upon the previous ones, and the final representation, created by three levels of nested parsing and abstraction, makes tasks such as mining and static checking tractable. The Dockerfiles, in each of the five representations, along with metadata and the tools used to shepard the data from one representation to the next are all available at: https://doi.org/10.5281/zenodo.3628771.
Jordan Henkel, Christian Bird, Shuvendu K. Lahiri, Thomas W. Reps
MSR1
2018 Code vectors: understanding programs through embedded abstracted symbolic traces
abstract
With the rise of machine learning, there is a great deal of interest in treating programs as data to be fed to learning algorithms. However, programs do not start off in a form that is immediately amenable to most off-the-shelf learning techniques. Instead, it is necessary to transform the program to a suitable representation before a learning technique can be applied.
Jordan Henkel, Shuvendu K. Lahiri, Ben Liblit, Thomas W. Reps
ESEC/SIGSOFT FSE1