Veselin Raychev

dblp:28/8412 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 16 · 8 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Security and privacy · 2Theory of computation · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
20 papers
Program synthesis and code generation · 32% Program analysis · 26% Concurrent programming · 11%
Network and information security
5 papers
Systems and software security · 97% Malware analysis · 3%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
code generation with language models
1.422025
BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025
TFix: Learning to Fix Coding Errors with a Text-to-Text Transformer · ICML 2021
Systems and software security › secure software development
secure code generation
0.912025
BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025
Systems and software security › secure software development
vulnerability prevention
0.912025
BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025
Software testing
test generation
0.912025
BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025
Concurrent programming
concurrency bugs
0.842015
Stateless model checking of event-driven applications · OOPSLA 2015
Scalable race detection for Android applications · OOPSLA 2015
Commutativity race detection · PLDI 2014
Program analysis
dynamic analysis
0.842015
Stateless model checking of event-driven applications · OOPSLA 2015
Scalable race detection for Android applications · OOPSLA 2015
Commutativity race detection · PLDI 2014
Program analysis
static analysis
0.832019
Unsupervised learning of API aliasing specifications · PLDI 2019
Learning a Static Analyzer from Data · CAV (1) 2017
Scalable taint specification inference with big code · PLDI 2019
Empirical software engineering
mining software repositories
0.722021
Learning to find naming issues with big code and small supervision · PLDI 2021
Predicting Program Properties from "Big Code" · POPL 2015
Program synthesis and code generation
code completion
0.732016
Learning programs from noisy data · POPL 2016
Probabilistic model for code with decision trees · OOPSLA 2016
Code completion with statistical language models · PLDI 2014
Concurrent programming › concurrency bug detection
data race detection
0.632015
Scalable race detection for Android applications · OOPSLA 2015
Commutativity race detection · PLDI 2014
Effective race detection for event-driven programs · OOPSLA 2013
Debugging and program repair
automated program repair
0.512021
TFix: Learning to Fix Coding Errors with a Text-to-Text Transformer · ICML 2021
Program synthesis and code generation › code completion
statistical code completion
0.422016
Learning programs from noisy data · POPL 2016
Code completion with statistical language models · PLDI 2014
Systems and software security › information flow tracking
taint analysis
0.412019
Scalable taint specification inference with big code · PLDI 2019
Program analysis
specification mining
0.412019
Unsupervised learning of API aliasing specifications · PLDI 2019
Systems and software security › vulnerability analysis
API misuse detection
0.312018
Inferring crypto API rules from code changes · PLDI 2018
Systems and software security › software vulnerability
cryptographic API misuse
0.312018
Inferring crypto API rules from code changes · PLDI 2018
Program analysis
binary analysis
0.312018
Debin: Predicting Debug Information in Stripped Binaries · CCS 2018
Software maintenance and evolution
code change analysis
0.312018
Inferring crypto API rules from code changes · PLDI 2018
Program analysis › binary analysis
stripped binary analysis
0.312018
Debin: Predicting Debug Information in Stripped Binaries · CCS 2018
Program synthesis and code generation › neural program synthesis
LLM-based program synthesis
0.312017
Program Synthesis for Character Level Language Modeling · ICLR (Poster) 2017
Systems and software security
program analysis
0.212016
Statistical Deobfuscation of Android Applications · CCS 2016
Program synthesis and code generation
programming by example
0.212016
Learning programs from noisy data · POPL 2016
Query processing and optimization
aggregation
0.212015
Parallelizing user-defined aggregations using symbolic execution · SOSP 2015
Query processing and optimization › parallel query processing
parallel aggregation
0.212015
Parallelizing user-defined aggregations using symbolic execution · SOSP 2015
Program analysis › dynamic analysis
happens-before analysis
0.212015
Scalable race detection for Android applications · OOPSLA 2015
Program verification › model checking
stateless model checking
0.212015
Stateless model checking of event-driven applications · OOPSLA 2015
Parallel and multicore computing
parallel programming models
0.212015
Parallelizing user-defined aggregations using symbolic execution · SOSP 2015
Software maintenance and evolution › refactoring
automated refactoring
0.212013
Refactoring with synthesis · OOPSLA 2013
Software maintenance and evolution
refactoring
0.212013
Refactoring with synthesis · OOPSLA 2013
Software maintenance and evolution
software libraries
0.112019
Unsupervised learning of API aliasing specifications · PLDI 2019

Methods — techniques the papers use, named apart from their topics

large language model · 1.7end-to-end exploit execution · 1.7machine learning · 1.2big code · 0.8program synthesis · 0.7structured prediction · 0.6probabilistic graphical models · 0.6unsupervised pattern mining · 0.5text-to-text transformer · 0.5supervised classification · 0.5pre-training · 0.5fine-tuning · 0.5symbolic execution · 0.4semi-supervised learning · 0.4language modeling · 0.3statistical learning · 0.2
YearPublicationVenuePosition
2025 BaxBench: Can LLMs Generate Correct and Secure Backends?
abstract
Automatic program generation has long been a fundamental challenge in computer science. Recent benchmarks have shown that large language models (LLMs) can effectively generate code at the function level, make code edits, and solve algorithmic coding tasks. However, to achieve full automation, LLMs should be able to generate production-quality, self-contained application modules. To evaluate the capabilities of LLMs in solving this challenge, we introduce BaxBench, a novel evaluation benchmark consisting of 392 tasks for the generation of backend applications. We focus on backends for three critical reasons: (i) they are practically relevant, building the core components of most modern web and cloud software, (ii) they are difficult to get right, requiring multiple functions and files to achieve the desired functionality, and (iii) they are security-critical, as they are exposed to untrusted third-parties, making secure solutions that prevent deployment-time attacks an imperative. BaxBench validates the functionality of the generated applications with comprehensive test cases, and assesses their security exposure by executing end-to-end exploits. Our experiments reveal key limitations of current LLMs in both functionality and security: (i) even the best model, OpenAI o1, achieves a mere 62% on code correctness; (ii) on average, we could successfully execute security exploits on around half of the correct programs generated by each LLM; and (iii) in less popular backend frameworks, models further struggle to generate correct and secure applications. Progress on BaxBench signifies important steps towards autonomous and secure software development with LLMs.
Mark Vero, Niels Mündler, Victor Chibotaru, Veselin Raychev, Maximilian Baader, Nikola Jovanovic 0001, Martin T. Vechev
ICML4
2021 TFix: Learning to Fix Coding Errors with a Text-to-Text Transformer
abstract
The problem of fixing errors in programs has attracted substantial interest over the years. The key challenge for building an effective code fixing tool is to capture a wide range of errors and meanwhile maintain high accuracy. In this paper, we address this challenge and present a new learning-based system, called TFix. TFix works directly on program text and phrases the problem of code fixing as a text-to-text task. In turn, this enables it to leverage a powerful Transformer based model pre-trained on natural language and fine-tuned to generate code fixes (via a large, high-quality dataset obtained from GitHub commits). TFix is not specific to a particular programming language or class of defects and, in fact, improved its precision by simultaneously fine-tuning on 52 different error types reported by a popular static analyzer. Our evaluation on a massive dataset of JavaScript programs shows that TFix is practically effective: it is able to synthesize code that fixes the error in 67 percent of cases and significantly outperforms existing learning-based approaches.
Berkay Berabi, Veselin Raychev, Martin T. Vechev
ICML3
2021 Learning to find naming issues with big code and small supervision
abstract
We introduce a new approach for finding and fixing naming issues in source code. The method is based on a careful combination of unsupervised and supervised procedures: (i) unsupervised mining of patterns from Big Code that express common naming idioms. Program fragments violating such idioms indicates likely naming issues, and (ii) supervised learning of a classifier on a small labeled dataset which filters potential false positives from the violations.
Cheng-Chun Lee, Veselin Raychev, Martin T. Vechev
PLDI3
2019 Scalable taint specification inference with big code
abstract
We present a new scalable, semi-supervised method for inferring taint analysis specifications by learning from a large dataset of programs. Taint specifications capture the role of library APIs (source, sink, sanitizer) and are a critical ingredient of any taint analyzer that aims to detect security violations based on information flow.
Victor Chibotaru, Benjamin Bichsel, Veselin Raychev, Martin T. Vechev
PLDI3
2019 Unsupervised learning of API aliasing specifications
abstract
Real world applications make heavy use of powerful libraries and frameworks, posing a significant challenge for static analysis as the library implementation may be very complex or unavailable. Thus, obtaining specifications that summarize the behaviors of the library is important as it enables static analyzers to precisely track the effects of APIs on the client program, without requiring the actual API implementation.
Jan Eberhardt, Samuel Steffen, Veselin Raychev, Martin T. Vechev
PLDI3
2018 Debin: Predicting Debug Information in Stripped Binaries
abstract
We present a novel approach for predicting debug information in stripped binaries. Using machine learning, we first train probabilistic models on thousands of non-stripped binaries and then use these models to predict properties of meaningful elements in unseen stripped binaries. Our focus is on recovering symbol names, types and locations, which are critical source-level information wiped off during compilation and stripping. Our learning approach is able to distinguish and extract key elements such as register-allocated and memory-allocated variables usually not evident in the stripped binary. To predict names and types of extracted elements, we use scalable structured prediction algorithms in probabilistic graphical models with an extensive set of features which capture key characteristics of binary code. Based on this approach, we implemented an automated tool, called Debin, which handles ELF binaries on three of the most popular architectures: x86, x64 and ARM. Given a stripped binary, Debin outputs a binary augmented with the predicted debug information. Our experimental results indicate that Debin is practically useful: for x64, it predicts symbol names and types with 68.8% precision and 68.3% recall. We also show that Debin is helpful for the task of inspecting real-world malware -- it revealed suspicious library usage and behaviors such as DNS resolver reader.
Pesho Ivanov, Petar Tsankov, Veselin Raychev, Martin T. Vechev
CCS4
2018 Inferring crypto API rules from code changes
abstract
Creating and maintaining an up-to-date set of security rules that match misuses of crypto APIs is challenging, as crypto APIs constantly evolve over time with new cryptographic primitives and settings, making existing ones obsolete.
Rumen Paletov, Petar Tsankov, Veselin Raychev, Martin T. Vechev
PLDI3
2017 Learning a Static Analyzer from Data
Pavol Bielik, Veselin Raychev, Martin T. Vechev
CAV (1)2
2017 Program Synthesis for Character Level Language Modeling
Pavol Bielik, Veselin Raychev, Martin T. Vechev
ICLR (Poster)2
2016 Statistical Deobfuscation of Android Applications
abstract
This work presents a new approach for deobfuscating Android APKs based on probabilistic learning of large code bases (termed "Big Code"). The key idea is to learn a probabilistic model over thousands of non-obfuscated Android applications and to use this probabilistic model to deobfuscate new, unseen Android APKs. The concrete focus of the paper is on reversing layout obfuscation, a popular transformation which renames key program elements such as classes, packages, and methods, thus making it difficult to understand what the program does. Concretely, the paper: (i) phrases the layout deobfuscation problem of Android APKs as structured prediction in a probabilistic graphical model, (ii) instantiates this model with a rich set of features and constraints that capture the Android setting, ensuring both semantic equivalence and high prediction accuracy, and (iii) shows how to leverage powerful inference and learning algorithms to achieve overall precision and scalability of the probabilistic predictions.
Benjamin Bichsel, Veselin Raychev, Petar Tsankov, Martin T. Vechev
CCS2
2016 PHOG: Probabilistic Model for Code
abstract
We introduce a new generative model for code called probabilistic higher order grammar (PHOG). PHOG generalizes probabilistic context free grammars (PCFGs) by allowing conditioning of a production rule beyond the parent non-terminal, thus capturing rich contexts relevant to programs. Even though PHOG is more powerful than a PCFG, it can be learned from data just as efficiently. We trained a PHOG model on a large JavaScript code corpus and show that it is more precise than existing models, while similarly fast. As a result, PHOG can immediately benefit existing programming tools based on probabilistic models of code.
Pavol Bielik, Veselin Raychev, Martin T. Vechev
ICML2
2016 Probabilistic model for code with decision trees
abstract
In this paper we introduce a new approach for learning precise and general probabilistic models of code based on decision tree learning. Our approach directly benefits an emerging class of statistical programming tools which leverage probabilistic models of code learned over large codebases (e.g., GitHub) to make predictions about new programs (e.g., code completion, repair, etc).
Veselin Raychev, Pavol Bielik, Martin T. Vechev
OOPSLA1
2016 Learning programs from noisy data
abstract
We present a new approach for learning programs from noisy datasets. Our approach is based on two new concepts: a regularized program generator which produces a candidate program based on a small sample of the entire dataset while avoiding overfitting, and a dataset sampler which carefully samples the dataset by leveraging the candidate program's score on that dataset. The two components are connected in a continuous feedback-directed loop. We show how to apply this approach to two settings: one where the dataset has a bound on the noise, and another without a noise bound. The second setting leads to a new way of performing approximate empirical risk minimization on hypotheses classes formed by a discrete search space. We then present two new kinds of program synthesizers which target the two noise settings. First, we introduce a novel regularized bitstream synthesizer that successfully generates programs even in the presence of incorrect examples. We show that the synthesizer can detect errors in the examples while combating overfitting -- a major problem in existing synthesis techniques. We also show how the approach can be used in a setting where the dataset grows dynamically via new examples (e.g., provided by a human). Second, we present a novel technique for constructing statistical code completion systems. These are systems trained on massive datasets of open source programs, also known as ``Big Code''. The key idea is to introduce a domain specific language (DSL) over trees and to learn functions in that DSL directly from the dataset. These learned functions then condition the predictions made by the system. This is a flexible and powerful technique which generalizes several existing works as we no longer need to decide a priori on what the prediction should be conditioned (another benefit is that the learned functions are a natural mechanism for explaining the prediction). As a result, our code completion system surpasses the prediction capabilities of existing, hard-wired systems.
Veselin Raychev, Pavol Bielik, Martin T. Vechev, Andreas Krause 0001
POPL1
2015 Scalable race detection for Android applications
abstract
We present a complete end-to-end dynamic analysis system for finding data races in mobile Android applications. The capabilities of our system significantly exceed the state of the art: our system can analyze real-world application interactions in minutes rather than hours, finds errors inherently beyond the reach of existing approaches, while still (critically) reporting very few false positives. Our system is based on three key concepts: (i) a thorough happens-before model of Android-specific concurrency, (ii) a scalable analysis algorithm for efficiently building and querying the happens-before graph, and (iii) an effective set of domain-specific filters that reduce the number of reported data races by several orders of magnitude. We evaluated the usability and performance of our system on 354 real-world Android applications (e.g., Facebook). Our system analyzes a minute of end-user interaction with the application in about 24 seconds, while current approaches take hours to complete. Inspecting the results for 8 large open-source applications revealed 15 harmful bugs of diverse kinds. Some of the bugs we reported were confirmed and fixed by developers.
Pavol Bielik, Veselin Raychev, Martin T. Vechev
OOPSLA2
2015 Stateless model checking of event-driven applications
abstract
Modern event-driven applications, such as, web pages and mobile apps, rely on asynchrony to ensure smooth end-user experience. Unfortunately, even though these applications are executed by a single event-loop thread, they can still exhibit nondeterministic behaviors depending on the execution order of interfering asynchronous events. As in classic shared-memory concurrency, this nondeterminism makes it challenging to discover errors that manifest only in specific schedules of events. In this work we propose the first stateless model checker for event-driven applications, called R4. Our algorithm systematically explores the nondeterminism in the application and concisely exposes its overall effect, which is useful for bug discovery. The algorithm builds on a combination of three key insights: (i) a dynamic partial order reduction (DPOR) technique for reducing the search space, tailored to the domain of event-driven applications, (ii) conflict-reversal bounding based on a hypothesis that most errors occur with a small number of event reorderings, and (iii) approximate replay of event sequences, which is critical for separating harmless from harmful nondeterminism. We instantiate R4 for the domain of client-side web applications and use it to analyze event interference in a number of real-world programs. The experimental results indicate that the precision and overall exploration capabilities of our system significantly exceed that of existing techniques.
Casper Svenning Jensen, Anders Møller, Veselin Raychev, Dimitar Dimitrov 0004, Martin T. Vechev
OOPSLA3
2015 Predicting Program Properties from "Big Code"
abstract
We present a new approach for predicting program properties from massive codebases (aka "Big Code"). Our approach first learns a probabilistic model from existing data and then uses this model to predict properties of new, unseen programs.
Veselin Raychev, Martin T. Vechev, Andreas Krause 0001
POPL1
2015 Parallelizing user-defined aggregations using symbolic execution
abstract
User-defined aggregations (UDAs) are integral to large-scale data-processing systems, such as MapReduce and Hadoop, because they let programmers express application-specific aggregation logic. System-supported associative aggregations, such as counting or finding the maximum, are data-parallel and thus these systems optimize their execution, leading in many cases to orders-of-magnitude performance improvements. These optimizations, however, are not possible on arbitrary UDAs.
Veselin Raychev, Madan Musuvathi, Todd Mytkowicz
SOSP1
2014 Commutativity race detection
abstract
This paper introduces the concept of a commutativity race. A commutativity race occurs in a given execution when two library method invocations can happen concurrently yet they do not commute. Commutativity races are an elegant concept enabling reasoning about concurrent interaction at the library interface.
Dimitar Dimitrov 0004, Veselin Raychev, Martin T. Vechev, Eric Koskinen
PLDI2
2014 Code completion with statistical language models
abstract
We address the problem of synthesizing code completions for programs using APIs. Given a program with holes, we synthesize completions for holes with the most likely sequences of method calls.
Veselin Raychev, Martin T. Vechev, Eran Yahav
PLDI1
2013 Refactoring with synthesis
abstract
Refactoring has become an integral part of modern software development, with wide support in popular integrated development environments (IDEs). Modern IDEs provide a fixed set of supported refactorings, listed in a refactoring menu. But with IDEs supporting more and more refactorings, it is becoming increasingly difficult for programmers to discover and memorize all their names and meanings. Also, since the set of refactorings is hard-coded, if a programmer wants to achieve a slightly different code transformation, she has to either apply a (possibly non-obvious) sequence of several built-in refactorings, or just perform the transformation by hand.
Veselin Raychev, Max Schäfer, Manu Sridharan, Martin T. Vechev
OOPSLA1
2013 Effective race detection for event-driven programs
abstract
Like shared-memory multi-threaded programs, event-driven programs such as client-side web applications are susceptible to data races that are hard to reproduce and debug. Race detection for such programs is hampered by their pervasive use of ad hoc synchronization, which can lead to a prohibitive number of false positives. Race detection also faces a scalability challenge, as a large number of short-running event handlers can quickly overwhelm standard vector-clock-based techniques.
Veselin Raychev, Martin T. Vechev, Manu Sridharan
OOPSLA1
2013 Automatic Synthesis of Deterministic Concurrency
Veselin Raychev, Martin T. Vechev, Eran Yahav
SAS1
2010 Fast Routing in Very Large Public Transportation Networks Using Transfer Patterns
Hannah Bast, Erik Carlsson, Arno Eigenwillig, Robert Geisberger, Chris Harrelson, Veselin Raychev, Fabien Viger
ESA (1)6