EDBT 2026 Demo / reviewers in the wild / expert
Saikat Chakraborty 0001
dblp:137/5220-1
· DBLP profile ↗
19ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0002-6889-7171ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 13 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Neural Synthesis for SMT-Assisted Proof-Oriented ProgrammingabstractProof-oriented programs mix computational content with proofs of program correctness. However, the human effort involved in programming and proving is still substantial, despite the use of Satisfiability Modulo Theories (SMT) solvers to automate proofs in languages such as F*. Seeking to spur research on using AI to automate the construction of proof-oriented programs, we curate a dataset of 600K lines of open-source F* programs and proofs, including software used in production systems ranging from Windows and Linux, to Python and Firefox. Our dataset includes around 32K top-level F* definitions, each representing a type-directed program and proof synthesis problem-producing a definition given a formal specification expressed as an F* type. We provide a program-fragment checker that queries F* to check the correctness of candidate solutions. We believe this is the largest corpus of SMT-assisted program proofs coupled with a reproducible program-fragment checker. Grounded in this dataset, we investigate the use of AI to synthesize programs and their proofs in F*, with promising results. Our main finding in that the performance of fine-tuned smaller language models (such as Phi-2 or StarCoder) compare favorably with large language models (such as GPT-4), at a much lower computational cost. We also identify various type-based retrieval augmentation techniques and find that they boost performance significantly. With detailed error analysis and case studies, we identify potential strengths and weaknesses of models and techniques and suggest directions for future improvements. Saikat Chakraborty 0001, Gabriel Ebner, Siddharth Bhat, Sarah Fakhoury, Sakina Fatima, Shuvendu K. Lahiri, Nikhil Swamy |
ICSE | 1 |
| 2024 | Leveraging LLMs for Program Verification
Adharsh Kamath, J. Nausheen Mohammed, Aditya Senthilnathan, Saikat Chakraborty 0001, Pantazis Deligiannis, Shuvendu K. Lahiri, Akash Lal, Aseem Rastogi, Subhajit Roy 0001, Rahul Sharma 0001 |
FMCAD | 4 |
| 2024 | Towards Causal Deep Learning for Vulnerability DetectionabstractDeep learning vulnerability detection has shown promising results in recent years. However, an important challenge that still blocks it from being very useful in practice is that the model is not robust under perturbation and it cannot generalize well over the out-of-distribution (OOD) data, e.g., applying a trained model to unseen projects in real world. We hypothesize that this is because the model learned non-robust features, e.g., variable names, that have spurious correlations with labels. When the perturbed and OOD datasets no longer have the same spurious features, the model prediction fails. To address the challenge, in this paper, we introduced causality into deep learning vulnerability detection. Our approach CausalVul consists of two phases. First, we designed novel perturbations to discover spurious features that the model may use to make predictions. Second, we applied the causal learning algorithms, specifically, do-calculus, on top of existing deep learning models to systematically remove the use of spurious features and thus promote causal based prediction. Our results show that CausalVul consistently improved the model accuracy, robustness and OOD performance for all the state-of-the-art models and datasets we experimented. To the best of our knowledge, this is the first work that introduces do calculus based causal learning to software engineering models and shows it's indeed useful for improving the model accuracy, robustness and generalization. Our replication package is located at https://figshare.com/s/0ffda320dcb96c249ef2. Ira Ceka, Chengzhi Mao, Saikat Chakraborty 0001, Baishakhi Ray, Wei Le |
ICSE | 4 |
| 2024 | Reinforest: Reinforcing Semantic Code Similarity for Cross-Lingual Code Search ModelsabstractThis paper introduces a novel code-to-code search technique that enhances the performance of Large Language Models (LLMs) by including both static and dynamic features as well as utilizing both similar and dissimilar examples during training. We present the first-ever code search method that encodes dynamic runtime information during training without the need to execute either the corpus under search or the search query at inference time and the first code search technique that trains on both positive and negative reference samples. To validate the efficacy of our approach, we perform a set of studies demonstrating the capability of enhanced LLMs to perform cross-language code-to-code search. Our evaluation demonstrates that the effectiveness of our approach is consistent across various model architectures and programming languages. We outperform the state-of-the-art cross-language search tool by up to 44.7%. Moreover, our ablation studies reveal that even a single positive and negative reference sample in the training process results in substantial performance improvements demonstrating both similar and dissimilar references are important parts of code search. Importantly, we show that enhanced well-crafted, fine-tuned models consistently outperform enhanced larger modern LLMs without fine tuning, even when enhancing the largest available LLMs highlighting the importance for open-sourced models. To ensure the reproducibility and extensibility of our research, we present an open-sourced implementation of our tool and training procedures called Reinforest. Anthony Saieva, Saikat Chakraborty 0001, Gail E. Kaiser |
SCAM | 2 |
| 2024 | LLM-Based Test-Driven Interactive Code Generation: User Study and Empirical EvaluationabstractLarge language models (LLMs) have shown great potential in automating significant aspects of coding by producing natural code from informal natural language (NL) intent. However, given NL is informal, it does not lend easily to checking that the generated code correctly satisfies the user intent. In this paper, we propose a novel interactive workflowTiCoderfor guided intent clarification (i.e., partial formalization) through tests to support the generation of more accurate code suggestions. Through a mixed methods user study with 15 programmers, we present an empirical evaluation of the effectiveness of the workflow to improve code generation accuracy. We find that participants using the proposed workflow are significantly more likely to correctly evaluate AI generated code, and report significantly less task-induced cognitive load. Furthermore, we test the potential of the workflow at scale with four different state-of-the-art LLMs on two python datasets, using an idealized proxy for a user feedback. We observe an average absolute improvement of 45.97% in the pass@1 code generation accuracy for both datasets and across all LLMs within 5 user interactions, in addition to the automatic generation of accompanying unit tests. Sarah Fakhoury, Aaditya Naik, Georgios Sakkas, Saikat Chakraborty 0001, Shuvendu K. Lahiri |
IEEE Trans. Software Eng. | 4 |
| 2024 | Automated Code Editing With Search-Generate-ModifyabstractCode editing is essential in evolving software development. In literature, several automated code editing tools are proposed, which leverage Information Retrieval-based techniques and Machine Learning-based code generation and code editing models. Each technique comes with its own promises and perils, and for this reason, they are often used together to complement their strengths and compensate for their weaknesses. This paper proposes a hybrid approach to better synthesize code edits by leveraging the power of code search, generation, and modification.Our key observation is that a patch that is obtained by search & retrieval, even if incorrect, can provide helpful guidance to a code generation model. However, a retrieval-guided patch produced by a code generation model can still be a few tokens off from the intended patch. Such generated patches can be slightly modified to create the intended patches. We developed a novel tool to solve this challenge: SARGAM, which is designed to follow a real developer’s code editing behavior. Given an original code version, the developer maysearchfor the related patches,generateor write the code, and thenmodifythe generated code to adapt it to the right context. Our evaluation of SARGAM on edit generation shows superior performance w.r.t. the current state-of-the-art techniques. SARGAM also shows its effectiveness on automated program repair tasks. Changshu Liu, Pelin Çetin, Yogesh Patodia, Baishakhi Ray, Saikat Chakraborty 0001, Yangruibo Ding |
IEEE Trans. Software Eng. | 5 |
| 2023 | Summarize and Generate to Back-translate: Unsupervised Translation of Programming LanguagesabstractBack-translation is widely known for its effectiveness in neural machine translation when there is little to no parallel data.In this approach, a source-to-target model is coupled with a target-to-source model trained in parallel.The target-to-source model generates noisy sources, while the source-to-target model is trained to reconstruct the targets and vice versa.Recent developments of multilingual pre-trained sequence-to-sequence models for programming languages have been very effective for a broad spectrum of downstream software engineering tasks.Hence, training them to build programming language translation systems via back-translation is compelling.However, these models cannot be further trained via back-translation since they learn to output sequences in the same language as the inputs during pre-training.As an alternative, we propose performing back-translation via code summarization and generation.In code summarization, a model learns to generate natural language (NL) summaries given code snippets.In code generation, the model learns to do the opposite.Therefore, target-to-source generation in back-translation can be viewed as a target-to-NL-to-source generation.We show that our proposed approach performs competitively with state-of-the-art methods.We have made the code publicly available.1 Wasi Uddin Ahmad, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001 |
EACL | 2 |
| 2023 | CONCORD: Clone-Aware Contrastive Learning for Source CodeabstractDeep Learning (DL) models to analyze source code have shown immense promise during the past few years. More recently, self-supervised pre-training has gained traction for learning generic code representations valuable for many downstream SE tasks, such as clone and bug detection. Yangruibo Ding, Saikat Chakraborty 0001, Luca Buratti, Saurabh Pujar, Alessandro Morari, Gail E. Kaiser, Baishakhi Ray |
ISSTA | 2 |
| 2023 | Grace: Language Models Meet Code EditsabstractDevelopers spend a significant amount of time in editing code for a variety of reasons such as bug fixing or adding new features. Designing effective methods to predict code edits has been an active yet challenging area of research due to the diversity of code edits and the difficulty of capturing the developer intent. In this work, we address these challenges by endowing pre-trained large language models (LLMs) with the knowledge of relevant prior associated edits, which we call the Grace (Generation conditioned on Associated Code Edits) method. The generative capability of the LLMs helps address the diversity in code changes and conditioning code generation on prior edits helps capture the latent developer intent. We evaluate two well-known LLMs, codex and CodeT5, in zero-shot and fine-tuning settings respectively. In our experiments with two datasets, Grace boosts the performance of the LLMs significantly, enabling them to generate 29% and 54% more correctly edited code in top-1 suggestions relative to the current state-of-the-art symbolic and neural approaches, respectively. Priyanshu Gupta, Avishree Khare, Yasharth Bajpai, Saikat Chakraborty 0001, Sumit Gulwani, Aditya Kanade 0001, Arjun Radhakrishna, Gustavo Soares, Ashish Tiwari 0001 |
ESEC/SIGSOFT FSE | 4 |
| 2022 | Towards Learning (Dis)-Similarity of Source Code from Program ContrastsabstractYangruibo Ding, Luca Buratti, Saurabh Pujar, Alessandro Morari, Baishakhi Ray, Saikat Chakraborty. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yangruibo Ding, Luca Buratti, Saurabh Pujar, Alessandro Morari, Baishakhi Ray, Saikat Chakraborty 0001 |
ACL (1) | 6 |
| 2022 | NatGen: generative pre-training by "naturalizing" source codeabstractPre-trained Generative Language models (e.g., PLBART, CodeT5, SPT-Code) for source code yielded strong results on several tasks in the past few years, including code generation and translation. These models have adopted varying pre-training objectives to learn statistics of code construction from very large-scale corpora in a self-supervised fashion; the success of pre-trained models largely hinges on these pre-training objectives. This paper proposes a new pre-training objective, “Naturalizing” of source code, exploiting code’s bimodal, dual-channel (formal & natural channels) nature. Unlike natural language, code’s bimodal, dual-channel nature allows us to generate semantically equivalent code at scale. We introduce six classes of semantic preserving transformations to introduce unnatural forms of code, and then force our model to produce more natural original programs written by developers. Learning to generate equivalent, but more natural code, at scale, over large corpora of open-source code, without explicit manual supervision, helps the model learn to both ingest & generate code. We fine-tune our model in three generative Software Engineering tasks: code generation, code translation, and code refinement with limited human-curated labeled data and achieve state-of-the-art performance rivaling CodeT5. We show that our pre-trained model is especially competitive at zero-shot and few-shot learning, and better at learning code properties (e.g., syntax, data flow) Saikat Chakraborty 0001, Toufique Ahmed, Yangruibo Ding, Premkumar T. Devanbu, Baishakhi Ray |
ESEC/SIGSOFT FSE | 1 |
| 2022 | CODIT: Code Editing With Tree-Based Neural ModelsabstractThe way developers edit day-to-day code tends to be repetitive, often using existing code elements. Many researchers have tried to automate repetitive code changes by learning from specific change templates which are applied to limited scope. The advancement of deep neural networks and the availability of vast open-source evolutionary data opens up the possibility of automatically learning those templates from the wild. However, deep neural network based modeling for code changes and code in general introduces some specific problems that needs specific attention from research community. For instance, compared to natural language, source code vocabulary can be significantly larger. Further, good changes in code do not break its syntactic structure. Thus, deploying state-of-the-art neural network models without adapting the methods to the source code domain yields sub-optimal results. To this end, we propose a novel tree-based neural network system to model source code changes and learn code change patterns from the wild. Specifically, we propose a tree-based neural machine translation model to learn the probability distribution of changes in code. We realize our model with a change suggestion engine,Codit, and train the model with more than 24k real-world changes and evaluate it on 5k patches. Our evaluation shows the effectiveness ofCoditin learning and suggesting patches.Coditcan also learn specific bug fix pattern from bug fixing patches and can fix 25 bugs out of 80 bugs in Defects4J. Saikat Chakraborty 0001, Yangruibo Ding, Miltiadis Allamanis, Baishakhi Ray |
IEEE Trans. Software Eng. | 1 |
| 2022 | Deep Learning Based Vulnerability Detection: Are We There Yet?abstractAutomated detection of software vulnerabilities is a fundamental problem in software security. Existing program analysis techniques either suffer from high false positives or false negatives. Recent progress in Deep Learning (DL) has resulted in a surge of interest in applying DL for automated vulnerability detection. Several recent studies have demonstrated promising results achieving an accuracy of up to 95 percent at detecting vulnerabilities. In this paper, we ask,“how well do the state-of-the-art DL-based techniques perform in a real-world vulnerability prediction scenario?”To our surprise, we find that their performance drops by more than 50 percent. A systematic investigation of what causes such precipitous performance drop reveals that existing DL-based vulnerability prediction approaches suffer from challenges with the training data (e.g., data duplication, unrealistic distribution of vulnerable classes, etc.) and with the model choices (e.g., simple token-based models). As a result, these approaches often do not learn features related to the actual cause of the vulnerabilities. Instead, they learn unrelated artifacts from the dataset (e.g., specific variable/function names, etc.). Leveraging these empirical findings, we demonstrate how a more principled approach to data collection and model design, based on realistic settings of vulnerability prediction, can lead to better solutions. The resulting tools perform significantly better than the studied baseline—up to 33.57 percent boost in precision and 128.38 percent boost in recall compared to the best performing model in the literature. Overall, this paper elucidates existing DL-based vulnerability prediction systems’ potential issues and draws a roadmap for future DL-based vulnerability prediction research. Saikat Chakraborty 0001, Rahul Krishna, Yangruibo Ding, Baishakhi Ray |
IEEE Trans. Software Eng. | 1 |
| 2021 | On Multi-Modal Learning of Editing Source CodeabstractIn recent years, Neural Machine Translator (NMT) has shown promise in automatically editing source code. Typical NMT based code editor only considers the code that needs to be changed as input and suggests developers with a ranked list of patched code to choose from - where the correct one may not always be at the top of the list. While NMT based code editing systems generate a broad spectrum of plausible patches, the correct one depends on the developers’ requirement and often on the context where the patch is applied. Thus, if developers provide some hints, using natural language, or providing patch context, NMT models can benefit from them.As a proof of concept, in this research, we leverage three modalities of information: edit location, edit code context, commit messages (as a proxy of developers’ hint in natural language) to automatically generate edits with NMT models. To that end, we build Modit, a multi-modal NMT based code editing engine. With in-depth investigation and analysis, we show that developers’ hint as an input modality can narrow the search space for patches and outperform state-of-the-art models to generate correctly patched code in top-1 position. Saikat Chakraborty 0001, Baishakhi Ray |
ASE | 1 |
| 2021 | Unified Pre-training for Program Understanding and GenerationabstractWasi Ahmad, Saikat Chakraborty, Baishakhi Ray, Kai-Wei Chang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Wasi Uddin Ahmad, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001 |
NAACL-HLT | 2 |
| 2020 | A Transformer-based Approach for Source Code SummarizationabstractGenerating a readable summary that describes the functionality of a program is known as source code summarization.In this task, learning code representation by modeling the pairwise relationship between code tokens to capture their long-range dependencies is crucial.To learn code representation for summarization, we explore the Transformer model that uses a self-attention mechanism and has shown to be effective in capturing long-range dependencies.In this work, we show that despite the approach is simple, it outperforms the state-of-the-art techniques by a significant margin.We perform extensive analysis and ablation studies that reveal several important findings, e.g., the absolute encoding of source code tokens' position hinders, while relative encoding significantly improves the summarization performance.We have made our code publicly available 1 to facilitate future research. Wasi Uddin Ahmad, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001 |
ACL | 2 |
| 2019 | Toward Optimal Selection of Information Retrieval Models for Software Engineering TasksabstractInformation Retrieval (IR) plays a pivotal role in diverse Software Engineering (SE) tasks, e.g., bug localization and triaging, bug report routing, code retrieval, requirements analysis, etc. SE tasks operate on diverse types of documents including code, text, stack-traces, and structured, semi-structured and unstructured meta-data that often contain specialized vocabularies. As the performance of any IR-based tool critically depends on the underlying document types, and given the diversity of SE corpora, it is essential to understand which models work best for which types of SE documents and tasks. We empirically investigate the interaction between IR models and document types for two representative SE tasks (bug localization and relevant project search), carefully chosen as they require a diverse set of SE artifacts (mixtures of code and text), and confirm that the models' performance varies significantly with mix of document types. Leveraging this insight, we propose a generalized framework, SRCH, to automatically select the most favorable IR model(s) for a given SE task. We evaluate SRCH w.r.t. these two tasks and confirm its effectiveness. Our preliminary user study shows that SRCH's intelligent adaption of the IR model(s) to the task at hand not only improves precision and recall for SE tasks but may also improve users' satisfaction. Md. Masudur Rahman 0001, Saikat Chakraborty 0001, Gail E. Kaiser, Baishakhi Ray |
SCAM | 2 |
| 2018 | Building Language Models for Text with Named EntitiesabstractText in many domains involves a significant amount of named entities.Predicting the entity names is often challenging for a language model as they appear less frequent on the training corpus.In this paper, we propose a novel and effective approach to building a discriminative language model which can learn the entity names by leveraging their entity type information.We also introduce two benchmark datasets based on recipes and Java programming codes, on which we evaluate the proposed model.Experimental results show that our model achieves 52.2% better perplexity in recipe generation and 22.06% on code generation than the stateof-the-art language models. Md. Rizwan Parvez, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001 |
ACL (1) | 2 |
| 2017 | A Heuristic Initialized Stochastic Memetic Algorithm for MDPVRP With Interdependent Depot OperationsabstractThe vehicle routing problem (VRP) is a widely studied combinatorial optimization problem. We introduce a variant of the multidepot and periodic VRP (MDPVRP) and propose a heuristic initialized stochastic memetic algorithm to solve it. The main challenge in designing such an algorithm for a large combinatorial optimization problem is to avoid premature convergence by maintaining a balance between exploration and exploitation of the search space. We employ intelligent initialization and stochastic learning to address this challenge. The intelligent initialization technique constructs a population by a mix of random and heuristic generated solutions. The stochastic learning enhances the solutions' quality selectively using simulated annealing with a set of random and heuristic operators. The hybridization of randomness and greediness in the initialization and learning process helps to maintain the balance between exploration and exploitation. Our proposed algorithm has been tested extensively on the existing benchmark problems and outperformed the baseline algorithms by a large margin. We further compared our results with that of the state-of-the-art algorithms working under MDPVRP formulation and found a significant improvement over their results. Abdus Salam Azad, Saikat Chakraborty 0001 |
IEEE Trans. Cybern. | 3 |