Niels Mündler

dblp:245/7560 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-3851-2557ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
4 papers
Program synthesis and code generation · 40% Software testing · 32% Programming languages and type systems · 12%
Network and information security
2 papers
Systems and software security · 80% Security and privacy of machine learning · 20%
Artificial intelligence
1 paper
Language models and text generation · 56% Trustworthy machine learning · 44%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
code generation with language models
2.032025
Type-Constrained Code Generation with Language Models · Proc. ACM Program. Lang. 2025
BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025
Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025
Software testing
test generation
1.622025
BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024
Security and privacy of machine learning
adversarial attack
0.912025
Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025
Systems and software security
insecure code generation
0.912025
Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025
Systems and software security › secure software development
secure code generation
0.912025
BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025
Systems and software security
vulnerability discovery
0.912025
Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025
Systems and software security › secure software development
vulnerability prevention
0.912025
BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025
Programming languages and type systems
type systems
0.912025
Type-Constrained Code Generation with Language Models · Proc. ACM Program. Lang. 2025
Machine learning › Trustworthy machine learning
hallucination
0.812024
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation · ICLR 2024
Software testing › test generation
automated test generation
0.812024
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024
Program synthesis and code generation
code agent
0.812024
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024
Compilers and program optimization
code generation
0.812024
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024
Program verification
functional correctness
0.312025
Type-Constrained Code Generation with Language Models · Proc. ACM Program. Lang. 2025
Program synthesis and code generation › code completion
LLM-based code completion
0.312025
Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025
Natural language and speech › Language models and text generation
prompting
0.212024
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation · ICLR 2024
Debugging and program repair
automated program repair
0.212024
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

large language model · 2.5query-based optimization · 1.7end-to-end exploit execution · 1.7black-box optimization · 1.7search over inhabitable types · 0.9prefix automata · 0.9constrained decoding · 0.9prompting-based detection · 0.8iterative refinement · 0.8code agents · 0.8black-box LM · 0.8
YearPublicationVenuePosition
2025 Black-Box Adversarial Attacks on LLM-Based Code Completion
abstract
Modern code completion engines, powered by large language models (LLMs), assist millions of developers with their strong capabilities to generate functionally correct code. Due to this popularity, it is crucial to investigate the security implications of relying on LLM-based code completion. In this work, we demonstrate that state-of-the-art black-box LLM-based code completion engines can be stealthily biased by adversaries to significantly increase their rate of insecure code generation. We present the first attack, named INSEC, that achieves this goal. INSEC works by injecting an attack string as a short comment in the completion input. The attack string is crafted through a query-based optimization procedure starting from a set of carefully designed initialization schemes. We demonstrate INSEC's broad applicability and effectiveness by evaluating it on various state-of-the-art open-source models and black-box commercial services (e.g., OpenAI API and GitHub Copilot). On a diverse set of security-critical test cases, covering 16 CWEs across 5 programming languages, INSEC increases the rate of generated insecure code by more than 50%, while maintaining the functional correctness of generated code. We consider INSEC practical - it requires low resources and costs less than 10 US dollars to develop on commodity hardware. Moreover, we showcase the attack's real-world deployability, by developing an IDE plug-in that stealthily injects INSEC into the GitHub Copilot extension.
Slobodan Jenko, Niels Mündler, Mark Vero, Martin T. Vechev
ICML2
2025 BaxBench: Can LLMs Generate Correct and Secure Backends?
abstract
Automatic program generation has long been a fundamental challenge in computer science. Recent benchmarks have shown that large language models (LLMs) can effectively generate code at the function level, make code edits, and solve algorithmic coding tasks. However, to achieve full automation, LLMs should be able to generate production-quality, self-contained application modules. To evaluate the capabilities of LLMs in solving this challenge, we introduce BaxBench, a novel evaluation benchmark consisting of 392 tasks for the generation of backend applications. We focus on backends for three critical reasons: (i) they are practically relevant, building the core components of most modern web and cloud software, (ii) they are difficult to get right, requiring multiple functions and files to achieve the desired functionality, and (iii) they are security-critical, as they are exposed to untrusted third-parties, making secure solutions that prevent deployment-time attacks an imperative. BaxBench validates the functionality of the generated applications with comprehensive test cases, and assesses their security exposure by executing end-to-end exploits. Our experiments reveal key limitations of current LLMs in both functionality and security: (i) even the best model, OpenAI o1, achieves a mere 62% on code correctness; (ii) on average, we could successfully execute security exploits on around half of the correct programs generated by each LLM; and (iii) in less popular backend frameworks, models further struggle to generate correct and secure applications. Progress on BaxBench signifies important steps towards autonomous and secure software development with LLMs.
Mark Vero, Niels Mündler, Victor Chibotaru, Veselin Raychev, Maximilian Baader, Nikola Jovanovic 0001, Martin T. Vechev
ICML2
2025 Type-Constrained Code Generation with Language Models
abstract
Large language models (LLMs) have achieved notable success in code generation. However, they still frequently produce uncompilable output because their next-token inference procedure does not model formal aspects of code. Although constrained decoding is a promising approach to alleviate this issue, it has only been applied to handle either domain-specific languages or syntactic features of general-purpose programming languages. However, LLMs frequently generate code with typing errors, which are beyond the domain of syntax and generally hard to adequately constrain. To address this challenge, we introduce a type-constrained decoding approach that leverages type systems to guide code generation. For this purpose, we develop novel prefix automata and a search over inhabitable types, forming a sound approach to enforce well-typedness on LLM-generated code. We formalize our approach on a foundational simply-typed language and extend it to TypeScript to demonstrate practicality. Our evaluation on the HumanEval and MBPP datasets shows that our approach reduces compilation errors by more than half and significantly increases functional correctness in code synthesis, translation, and repair tasks across LLMs of various sizes and model families, including state-of-the-art open-weight models with more than 30B parameters. The results demonstrate the generality and effectiveness of our approach in constraining LLM code generation with formal rules of type systems.
Niels Mündler, Hao Wang 0112, Koushik Sen, Dawn Song, Martin T. Vechev
Proc. ACM Program. Lang.1
2024 Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
abstract
Large language models (large LMs) are susceptible to producing text that contains hallucinated content. An important instance of this problem is self-contradiction, where the LM generates two contradictory sentences within the same context. In this work, we present a comprehensive investigation into self-contradiction for various instruction-tuned LMs, covering evaluation, detection, and mitigation. Our primary evaluation task is open-domain text generation, but we also demonstrate the applicability of our approach to shorter question answering. Our analysis reveals the prevalence of self-contradictions, e.g., in 17.7% of all sentences produced by ChatGPT. We then propose a novel prompting-based framework designed to effectively detect and mitigate self-contradictions. Our detector achieves high accuracy, e.g., around 80% F1 score when prompting ChatGPT. The mitigation algorithm iteratively refines the generated text to remove contradictory information while preserving text fluency and informativeness. Importantly, our entire framework is applicable to black-box LMs and does not require retrieval of external knowledge. Rather, our method complements retrieval-based methods, as a large portion of self-contradictions (e.g., 35.2% for ChatGPT) cannot be verified using online text. Our approach is practically effective and has been released as a push-button tool to benefit the public at https://chatprotect.ai/.
Niels Mündler, Slobodan Jenko, Martin T. Vechev
ICLR1
2024 SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
abstract
Rigorous software testing is crucial for developing and maintaining high-quality code, making automated test generation a promising avenue for both improving software quality and boosting the effectiveness of code generation methods. However, while code generation with Large Language Models (LLMs) is an extraordinarily active research area, test generation remains relatively unexplored. We address this gap and investigate the capability of LLM-based Code Agents to formalize user issues into test cases. To this end, we propose a novel benchmark based on popular GitHub repositories, containing real-world issues, ground-truth bug-fixes, and golden tests. We find that LLMs generally perform surprisingly well at generating relevant test cases, with Code Agents designed for code repair exceeding the performance of systems designed specifically for test generation. Further, as test generation is a similar but more structured task than code generation, it allows for a more fine-grained analysis using issue reproduction rate and coverage changes, providing a dual metric for analyzing systems designed for code repair. Finally, we find that generated tests are an effective filter for proposed code fixes, doubling the precision of SWE-Agent. We release all data and code at https://github.com/logic-star-ai/SWT-Bench.
Niels Mündler, Mark Niklas Müller, Martin T. Vechev
NeurIPS1
2022 A Verified Implementation of B+-Trees in Isabelle/HOL
Niels Mündler, Tobias Nipkow
ICTAC1