VLDB 2026 Research / reviewers in the wild / expert
Sheena Panthaplackel
dblp:255/5631
· DBLP profile ↗
7ranked-venue papers
5as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
6 papers |
Software maintenance and evolution · 32% Program synthesis and code generation · 31% Empirical software engineering · 10% | |
| Artificial intelligence
4 papers |
Language models and text generation · 86% Information extraction and text analysis · 14% |
Topics — the 14 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software maintenance and evolution › software documentation
code comment consistency |
1.1 | 3 | 2021 | Deep Just-In-Time Inconsistency Detection Between Comments and Source Code · AAAI 2021 Learning to Update Natural Language Comments Based on Code Changes · ACL 2020 Associating Natural Language Comment and Source Code Entities · AAAI 2020 |
Empirical software engineering › AI for software engineering
evaluation of language models for code |
0.8 | 1 | 2024 | Unsupervised Evaluation of Code LLMs with Round-Trip Correctness · ICML 2024 |
Program verification › equivalence checking
semantic equivalence checking |
0.8 | 1 | 2024 | Unsupervised Evaluation of Code LLMs with Round-Trip Correctness · ICML 2024 |
Software testing
test oracle |
0.8 | 1 | 2024 | Unsupervised Evaluation of Code LLMs with Round-Trip Correctness · ICML 2024 |
Natural language and speech › Language models and text generation › text summarization › abstractive summarization
headline generation |
0.6 | 1 | 2022 | Updated Headline Generation: Creating Updated Summaries for Evolving News Stories · ACL (1) 2022 |
Natural language and speech › Language models and text generation
text summarization |
0.6 | 1 | 2022 | Updated Headline Generation: Creating Updated Summaries for Evolving News Stories · ACL (1) 2022 |
Natural language and speech › Language models and text generation › text summarization › temporal summarization
update summarization |
0.6 | 1 | 2022 | Updated Headline Generation: Creating Updated Summaries for Evolving News Stories · ACL (1) 2022 |
Program synthesis and code generation
code generation with language models |
0.6 | 1 | 2022 | CoditT5: Pretraining for Source Code and Natural Language Editing · ASE 2022 |
Program synthesis and code generation › code generation with language models
edit-based code generation |
0.6 | 1 | 2022 | CoditT5: Pretraining for Source Code and Natural Language Editing · ASE 2022 |
Software maintenance and evolution › software documentation
code comment maintenance |
0.4 | 1 | 2020 | Learning to Update Natural Language Comments Based on Code Changes · ACL 2020 |
Compilers and program optimization
code generation |
0.4 | 1 | 2020 | Learning to Update Natural Language Comments Based on Code Changes · ACL 2020 |
Software maintenance and evolution › software documentation
comment updating |
0.4 | 1 | 2020 | Learning to Update Natural Language Comments Based on Code Changes · ACL 2020 |
Software maintenance and evolution › software maintenance
bug fixing |
0.2 | 1 | 2022 | CoditT5: Pretraining for Source Code and Natural Language Editing · ASE 2022 |
Software maintenance and evolution
code review |
0.2 | 1 | 2022 | CoditT5: Pretraining for Source Code and Natural Language Editing · ASE 2022 |
Methods — techniques the papers use, named apart from their topics
sequence-to-sequence model · 1.0neural sequence models · 1.0deep learning · 1.0copy mechanism · 1.0beam search · 1.0binary classification · 0.9large language model · 0.8reranking · 0.6pre-trained language model · 0.6edit modeling · 0.6contextual summarization · 0.6conditional generation · 0.6sequence labeling · 0.4feature engineering · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Unsupervised Evaluation of Code LLMs with Round-Trip CorrectnessabstractTo evaluate code large language models (LLMs), research has relied on a few small manually curated benchmarks, such as HumanEval and MBPP, which represent a narrow part of the real-world software domains. In this work, we introduce round-trip correctness (RTC) as an alternative evaluation method. RTC allows Code LLM evaluation on a broader spectrum of real-world software domains without the need for costly human curation. RTC rests on the idea that we can ask a model to make a prediction (e.g., describe some code using natural language), feed that prediction back (e.g., synthesize code from the predicted description), and check if this round-trip leads to code that is semantically equivalent to the original input. We show how to employ RTC to evaluate code synthesis and editing. We find that RTC strongly correlates with model performance on existing narrow-domain code synthesis benchmarks while allowing us to expand to a much broader set of domains and tasks which was not previously possible without costly human annotations. Miltiadis Allamanis, Sheena Panthaplackel |
ICML | 2 |
| 2022 | Updated Headline Generation: Creating Updated Summaries for Evolving News StoriesabstractWe propose the task of updated headline generation, in which a system generates a headline for an updated article, considering both the previous article and headline.The system must identify the novel information in the article update, and modify the existing headline accordingly.We create data for this task using the NewsEdits corpus (Spangher and May, 2021) by automatically identifying contiguous article versions that are likely to require a substantive headline update.We find that models conditioned on the prior headline and body revisions produce headlines judged by humans to be as factual as gold headlines while making fewer unnecessary edits compared to a standard headline generation model.Our experiments establish benchmarks for this new contextual summarization task. Sheena Panthaplackel, Adrian Benton, Mark Dredze |
ACL (1) | 1 |
| 2022 | CoditT5: Pretraining for Source Code and Natural Language EditingabstractPretrained language models have been shown to be effective in many software-related generation tasks; however, they are not well-suited for editing tasks as they are not designed to reason about edits. To address this, we propose a novel pretraining objective which explicitly models edits and use it to build CoditT5, a large language model for software-related editing tasks that is pretrained on large amounts of source code and natural language comments. We fine-tune it on various downstream editing tasks, including comment updating, bug fixing, and automated code review. By outperforming standard generation-based models, we demonstrate the generalizability of our approach and its suitability for editing tasks. We also show how a standard generation model and our edit-based model can complement one another through simple reranking strategies, with which we achieve state-of-the-art performance for the three downstream editing tasks. Jiyang Zhang 0003, Sheena Panthaplackel, Pengyu Nie 0001, Junyi Jessy Li, Milos Gligoric 0001 |
ASE | 2 |
| 2021 | Copy That! Editing Sequences by Copying SpansabstractNeural sequence-to-sequence models are finding increasing use in editing of documents, for example in correcting a text document or repairing source code. In this paper, we argue that common seq2seq models (with a facility to copy single tokens) are not a natural fit for such tasks, as they have to explicitly copy each unchanged token. We present an extension of seq2seq models capable of copying entire spans of the input to the output in one step, greatly reducing the number of decisions required during inference. This extension means that there are now many ways of generating the same output, which we handle by deriving a new objective for training and a variation of beam search for inference that explicitly handles this problem. In our experiments on a range of editing tasks of natural language and source code, we show that our new model consistently outperforms simpler baselines. Sheena Panthaplackel, Miltiadis Allamanis, Marc Brockschmidt |
AAAI | 1 |
| 2021 | Deep Just-In-Time Inconsistency Detection Between Comments and Source CodeabstractNatural language comments convey key aspects of source code such as implementation, usage, and pre- and post-conditions. Failure to update comments accordingly when the corresponding code is modified introduces inconsistencies, which is known to lead to confusion and software bugs. In this paper, we aim to detect whether a comment becomes inconsistent as a result of changes to the corresponding body of code, in order to catch potential inconsistencies just-in-time, i.e., before they are committed to a code base. To achieve this, we develop a deep-learning approach that learns to correlate a comment with code changes. By evaluating on a large corpus of comment/code pairs spanning various comment types, we show that our model outperforms multiple baselines by significant margins. For extrinsic evaluation, we show the usefulness of our approach by combining it with a comment update model to build a more comprehensive automatic comment maintenance system which can both detect and resolve inconsistent comments based on code changes. Sheena Panthaplackel, Junyi Jessy Li, Milos Gligoric 0001, Raymond J. Mooney |
AAAI | 1 |
| 2020 | Associating Natural Language Comment and Source Code EntitiesabstractComments are an integral part of software development; they are natural language descriptions associated with source code elements. Understanding explicit associations can be useful in improving code comprehensibility and maintaining the consistency between code and comments. As an initial step towards this larger goal, we address the task of associating entities in Javadoc comments with elements in Java source code. We propose an approach for automatically extracting supervised data using revision histories of open source projects and present a manually annotated evaluation dataset for this task. We develop a binary classifier and a sequence labeling model by crafting a rich feature set which encompasses various aspects of code, comments, and the relationships between them. Experiments show that our systems outperform several baselines learning from the proposed supervision. Sheena Panthaplackel, Milos Gligoric 0001, Raymond J. Mooney, Junyi Jessy Li |
AAAI | 1 |
| 2020 | Learning to Update Natural Language Comments Based on Code ChangesabstractWe formulate the novel task of automatically updating an existing natural language comment based on changes in the body of code it accompanies.We propose an approach that learns to correlate changes across two distinct language representations, to generate a sequence of edits that are applied to the existing comment to reflect the source code modifications.We train and evaluate our model using a dataset that we collected from commit histories of open-source software projects, with each example consisting of a concurrent update to a method and its corresponding comment.We compare our approach against multiple baselines using both automatic metrics and human evaluation.Results reflect the challenge of this task and that our model outperforms baselines with respect to making edits. Sheena Panthaplackel, Pengyu Nie 0001, Milos Gligoric 0001, Junyi Jessy Li, Raymond J. Mooney |
ACL | 1 |