Colin B. Clement

dblp:241/5448 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2023
0000-0002-3727-7308ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Program Translation via Code Distillation
abstract
Software version migration and program translation are an important and costly part of the lifecycle of large codebases.Traditional machine translation relies on parallel corpora for supervised translation, which is not feasible for program translation due to a dearth of aligned data.Recent unsupervised neural machine translation techniques have overcome data limitations by included techniques such as back translation and low level compiler intermediate representations (IR).These methods face significant challenges due to the noise in code snippet alignment and the diversity of IRs respectively.In this paper we propose a novel model called Code Distillation (CoDist) whereby we capture the semantic and structural equivalence of code in a language agnostic intermediate representation.Distilled code serves as a translation pivot for any programming language, leading by construction to parallel corpora which scale to all available source code by simply applying the distillation compiler.We demonstrate that our approach achieves state-of-the-art performance on CodeXGLUE and TransCoder GeeksForGeeks translation benchmarks, with an average absolute increase of 12.7% on the TransCoder GeeksforGeeks translation benchmark compare to TransCoder-ST.
Yufan Huang, Mengnan Qi, Yongqiang Yao, Maoquan Wang, Bin Gu 0001, Colin B. Clement, Neel Sundaresan
EMNLP6
2023 SUT: Active Defects Probing for Transcompiler Models
abstract
Program translation, i.e. transcompilation has been attracting increasing attention from researchers due to its enormous application value.However, we observe that current program translating models still make elementary syntax errors, particularly when the source language uses syntax elements not present in the target language, which is exactly what developers are concerned about while may not be well exposed by frequently used metrics such as BLEU, CodeBLEU and Computation Accuracy.In this paper, we focus on evaluating the model's ability to address these basic syntax errors and developed an novel active defects probing suite, the Syntactic Unit Tests (SUT) and highly interpretable evaluation harness including Syntax Unit Test Accuracy (SUT Acc) metric and Syntax Element Test Score (SETS), to help diagnose and promote progress in this area.Our Syntactic Unit Test fills the gap in the community for a fine-grained evaluation dataset for program translation.Experimental analysis shows that our evaluation harness is more accurate, reliable, and in line with human judgments compared to previous metrics.
Mengnan Qi, Yufan Huang, Maoquan Wang, Yongqiang Yao, Bin Gu 0001, Colin B. Clement, Neel Sundaresan
EMNLP7
2022 Learning to Reduce False Positives in Analytic Bug Detectors
abstract
Due to increasingly complex software design and rapid iterative development, code defects and security vulnerabilities are prevalent in modern software. In response, programmers rely on static analysis tools to regularly scan their codebases and find potential bugs. In order to maximize coverage, however, these tools generally tend to report a significant number of false positives, requiring developers to manually verify each warning. To address this problem, we propose a Transformer-based learning approach to identify false positive bug warnings. We demonstrate that our models can improve the precision of static analysis by 17.5%. In addition, we validated the generalizability of this approach across two major bug types: null dereference and resource leak.
Anant Kharkar, Roshanak Zilouchian Moghaddam, Matthew Jin, Colin B. Clement, Neel Sundaresan
ICSE6
2022 Generating Examples from CLI Usage: Can Transformers Help?
abstract
Continuous evolution in modern software often causes documentation, tutorials, and examples to be out of sync with changing interfaces and frameworks. Relying on outdated documentation and examples can lead programs to fail or be less efficient or even less secure. In response, programmers need to regularly turn to other resources on the web, such as StackOverflow for examples to guide them in writing software. We recognize that this inconvenient, error-prone, and expensive process can be improved by using machine learning applied to software usage data. In this paper, we present a practical system, which uses machine learning on large-scale telemetry data and documentation corpora, generating appropriate and complex examples that can be used to improve documentation. We discuss both feature-based and transformer-based machine learning approaches and demonstrate that our system achieves 100% coverage for the used functionalities in the product, providing up-to-date examples upon every release and reduces the numbers of PRs submitted by software owners writing and editing documentation by >68%. We also share valuable lessons learnt during the 3 years that our production quality system has been deployed for Azure Cloud Command Line Interface (Azure CLI)
Roshanak Zilouchian Moghaddam, Spandan Garg, Colin B. Clement, Yevhen Mohylevskyy, Neel Sundaresan
KDD3
2022 DeepDev-PERF: a deep learning-based approach for improving software performance
abstract
Improving software performance is an important yet challenging part of the software development cycle. Today, the majority of performance inefficiencies are identified and patched by performance experts. Recent advancements in deep learning approaches and the wide-spread availability of open-source data creates a great opportunity to automate the identification and patching of performance problems. In this paper, we present DeepDev-PERF, a transformer-based approach to suggest performance improvements for C# applications. We pretrain DeepDev-PERF on English and Source code corpora, followed by finetuning for the task of generating performance improvement patches for C# applications. Our evaluation shows that our model can generate the same performance improvement suggestion as the developer fix in ‍53
Spandan Garg, Roshanak Zilouchian Moghaddam, Colin B. Clement, Neel Sundaresan
ESEC/SIGSOFT FSE3
2022 Exploring and evaluating personalized models for code generation
abstract
Large Transformer models achieved the state-of-the-art status for Natural Language Understanding tasks and are increasingly becoming the baseline model architecture for modeling source code. Transformers are usually pre-trained on large unsupervised corpora, learning token representations and transformations relevant to modeling generally available text, and are then fine-tuned on a particular downstream task of interest. While fine-tuning is a tried-and-true method for adapting a model to a new domain -- for example, question-answering on a given topic -- generalization remains an on-going challenge. In this paper, we explore and evaluate transformer model fine-tuning for personalization. In the context of generating unit tests for Java methods, we evaluate learning to personalize to a specific software project using several personalization techniques. We consider three key approaches: (i) custom fine-tuning, which allows all the model parameters to be tuned; (ii) lightweight fine-tuning, which freezes most of the model's parameters, allowing tuning of the token embeddings and softmax layer only or the final layer alone; (iii) prefix tuning, which keeps model parameters frozen, but optimizes a small project-specific prefix vector. Each of these techniques offers a trade-off in total compute cost and predictive performance, which we evaluate by code and task-specific metrics, training time, and total computational operations. We compare these fine-tuning strategies for code generation and discuss the potential generalization and cost benefits of each in various deployment scenarios.
Andrei Zlotchevski, Dawn Drain, Alexey Svyatkovskiy, Colin B. Clement, Neel Sundaresan, Michele Tufano
ESEC/SIGSOFT FSE4
2021 Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy
abstract
Colin Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano, Dawn Drain, Nan Duan, Neel Sundaresan, Alexey Svyatkovskiy. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Colin B. Clement, Michele Tufano, Dawn Drain, Nan Duan 0001, Neel Sundaresan, Alexey Svyatkovskiy
EMNLP (1)1
2021 GraphCodeBERT: Pre-training Code Representations with Data Flow
Daya Guo, Shuo Ren 0002, Zhangyin Feng, Duyu Tang, Shujie Liu 0001, Nan Duan 0001, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin 0001, Daxin Jiang, Ming Zhou 0001
ICLR13
2020 PyMT5: multi-mode translation of natural language and Python code with transformers
abstract
Simultaneously modeling source code and natural language has many exciting applications in automated software development and understanding.Pursuant to achieving such technology, we introduce PYMT5, the PYTHON method text-to-text transfer transformer, which is trained to translate between all pairs of PYTHON method feature combinations: a single model that can both predict whole methods from natural language documentation strings (docstrings) and summarize code into docstrings of any common style.We present an analysis and modeling effort of a large-scale parallel corpus of 26 million PYTHON methods and 7.7 million method-docstring pairs, demonstrating that for docstring and method generation, PYMT5 outperforms similarlysized auto-regressive language models (GPT2) which were English pre-trained or randomly initialized.On the CODE-SEARCHNET test set, our best model predicts 92.1% syntactically correct method bodies, achieved a BLEU score of 8.59 for method generation and 16.3 for docstring * Corresponding author † Work done during a Microsoft internship generation (summarization), and achieved a ROUGE-L F-score of 24.8 for method generation and 36.7 for docstring generation.
Colin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, Neel Sundaresan
EMNLP (1)1