VLDB 2026 Research / reviewers in the wild / expert
Dawn Drain
dblp:274/2078
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2022
0000-0002-6606-4141ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Generating Accurate Assert Statements for Unit Test Cases using Pretrained TransformersabstractUnit testing represents the foundational basis of the software testing pyramid, beneath integration and end-to-end testing. Automated software testing researchers have proposed a variety of techniques to assist developers in this time-consuming task. Michele Tufano, Dawn Drain, Alexey Svyatkovskiy, Neel Sundaresan |
AST | 2 |
| 2022 | Exploring and evaluating personalized models for code generationabstractLarge Transformer models achieved the state-of-the-art status for Natural Language Understanding tasks and are increasingly becoming the baseline model architecture for modeling source code. Transformers are usually pre-trained on large unsupervised corpora, learning token representations and transformations relevant to modeling generally available text, and are then fine-tuned on a particular downstream task of interest. While fine-tuning is a tried-and-true method for adapting a model to a new domain -- for example, question-answering on a given topic -- generalization remains an on-going challenge. In this paper, we explore and evaluate transformer model fine-tuning for personalization. In the context of generating unit tests for Java methods, we evaluate learning to personalize to a specific software project using several personalization techniques. We consider three key approaches: (i) custom fine-tuning, which allows all the model parameters to be tuned; (ii) lightweight fine-tuning, which freezes most of the model's parameters, allowing tuning of the token embeddings and softmax layer only or the final layer alone; (iii) prefix tuning, which keeps model parameters frozen, but optimizes a small project-specific prefix vector. Each of these techniques offers a trade-off in total compute cost and predictive performance, which we evaluate by code and task-specific metrics, training time, and total computational operations. We compare these fine-tuning strategies for code generation and discuss the potential generalization and cost benefits of each in various deployment scenarios. Andrei Zlotchevski, Dawn Drain, Alexey Svyatkovskiy, Colin B. Clement, Neel Sundaresan, Michele Tufano |
ESEC/SIGSOFT FSE | 2 |
| 2021 | Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax HierarchyabstractColin Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano, Dawn Drain, Nan Duan, Neel Sundaresan, Alexey Svyatkovskiy. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Colin B. Clement, Michele Tufano, Dawn Drain, Nan Duan 0001, Neel Sundaresan, Alexey Svyatkovskiy |
EMNLP (1) | 5 |
| 2021 | GraphCodeBERT: Pre-training Code Representations with Data Flow
Daya Guo, Shuo Ren 0002, Zhangyin Feng, Duyu Tang, Shujie Liu 0001, Nan Duan 0001, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin 0001, Daxin Jiang, Ming Zhou 0001 |
ICLR | 14 |
| 2020 | PyMT5: multi-mode translation of natural language and Python code with transformersabstractSimultaneously modeling source code and natural language has many exciting applications in automated software development and understanding.Pursuant to achieving such technology, we introduce PYMT5, the PYTHON method text-to-text transfer transformer, which is trained to translate between all pairs of PYTHON method feature combinations: a single model that can both predict whole methods from natural language documentation strings (docstrings) and summarize code into docstrings of any common style.We present an analysis and modeling effort of a large-scale parallel corpus of 26 million PYTHON methods and 7.7 million method-docstring pairs, demonstrating that for docstring and method generation, PYMT5 outperforms similarlysized auto-regressive language models (GPT2) which were English pre-trained or randomly initialized.On the CODE-SEARCHNET test set, our best model predicts 92.1% syntactically correct method bodies, achieved a BLEU score of 8.59 for method generation and 16.3 for docstring * Corresponding author † Work done during a Microsoft internship generation (summarization), and achieved a ROUGE-L F-score of 24.8 for method generation and 36.7 for docstring generation. Colin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, Neel Sundaresan |
EMNLP (1) | 2 |