EDBT 2026 Demo / reviewers in the wild / expert
Niels Mündler
dblp:245/7560
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-3851-2557ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
4 papers |
Program synthesis and code generation · 40% Software testing · 32% Programming languages and type systems · 12% | |
| Network and information security
2 papers |
Systems and software security · 80% Security and privacy of machine learning · 20% | |
| Artificial intelligence
1 paper |
Language models and text generation · 56% Trustworthy machine learning · 44% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program synthesis and code generation
code generation with language models |
2.0 | 3 | 2025 | Type-Constrained Code Generation with Language Models · Proc. ACM Program. Lang. 2025 BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025 Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025 |
Software testing
test generation |
1.6 | 2 | 2025 | BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025 SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024 |
Security and privacy of machine learning
adversarial attack |
0.9 | 1 | 2025 | Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025 |
Systems and software security
insecure code generation |
0.9 | 1 | 2025 | Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025 |
Systems and software security › secure software development
secure code generation |
0.9 | 1 | 2025 | BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025 |
Systems and software security
vulnerability discovery |
0.9 | 1 | 2025 | Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025 |
Systems and software security › secure software development
vulnerability prevention |
0.9 | 1 | 2025 | BaxBench: Can LLMs Generate Correct and Secure Backends? · ICML 2025 |
Programming languages and type systems
type systems |
0.9 | 1 | 2025 | Type-Constrained Code Generation with Language Models · Proc. ACM Program. Lang. 2025 |
Machine learning › Trustworthy machine learning
hallucination |
0.8 | 1 | 2024 | Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation · ICLR 2024 |
Software testing › test generation
automated test generation |
0.8 | 1 | 2024 | SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024 |
Program synthesis and code generation
code agent |
0.8 | 1 | 2024 | SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024 |
Compilers and program optimization
code generation |
0.8 | 1 | 2024 | SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024 |
Program verification
functional correctness |
0.3 | 1 | 2025 | Type-Constrained Code Generation with Language Models · Proc. ACM Program. Lang. 2025 |
Program synthesis and code generation › code completion
LLM-based code completion |
0.3 | 1 | 2025 | Black-Box Adversarial Attacks on LLM-Based Code Completion · ICML 2025 |
Natural language and speech › Language models and text generation
prompting |
0.2 | 1 | 2024 | Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation · ICLR 2024 |
Debugging and program repair
automated program repair |
0.2 | 1 | 2024 | SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.5query-based optimization · 1.7end-to-end exploit execution · 1.7black-box optimization · 1.7search over inhabitable types · 0.9prefix automata · 0.9constrained decoding · 0.9prompting-based detection · 0.8iterative refinement · 0.8code agents · 0.8black-box LM · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Black-Box Adversarial Attacks on LLM-Based Code CompletionabstractModern code completion engines, powered by large language models (LLMs), assist millions of developers with their strong capabilities to generate functionally correct code. Due to this popularity, it is crucial to investigate the security implications of relying on LLM-based code completion. In this work, we demonstrate that state-of-the-art black-box LLM-based code completion engines can be stealthily biased by adversaries to significantly increase their rate of insecure code generation. We present the first attack, named INSEC, that achieves this goal. INSEC works by injecting an attack string as a short comment in the completion input. The attack string is crafted through a query-based optimization procedure starting from a set of carefully designed initialization schemes. We demonstrate INSEC's broad applicability and effectiveness by evaluating it on various state-of-the-art open-source models and black-box commercial services (e.g., OpenAI API and GitHub Copilot). On a diverse set of security-critical test cases, covering 16 CWEs across 5 programming languages, INSEC increases the rate of generated insecure code by more than 50%, while maintaining the functional correctness of generated code. We consider INSEC practical - it requires low resources and costs less than 10 US dollars to develop on commodity hardware. Moreover, we showcase the attack's real-world deployability, by developing an IDE plug-in that stealthily injects INSEC into the GitHub Copilot extension. Slobodan Jenko, Niels Mündler, Mark Vero, Martin T. Vechev |
ICML | 2 |
| 2025 | BaxBench: Can LLMs Generate Correct and Secure Backends?abstractAutomatic program generation has long been a fundamental challenge in computer science. Recent benchmarks have shown that large language models (LLMs) can effectively generate code at the function level, make code edits, and solve algorithmic coding tasks. However, to achieve full automation, LLMs should be able to generate production-quality, self-contained application modules. To evaluate the capabilities of LLMs in solving this challenge, we introduce BaxBench, a novel evaluation benchmark consisting of 392 tasks for the generation of backend applications. We focus on backends for three critical reasons: (i) they are practically relevant, building the core components of most modern web and cloud software, (ii) they are difficult to get right, requiring multiple functions and files to achieve the desired functionality, and (iii) they are security-critical, as they are exposed to untrusted third-parties, making secure solutions that prevent deployment-time attacks an imperative. BaxBench validates the functionality of the generated applications with comprehensive test cases, and assesses their security exposure by executing end-to-end exploits. Our experiments reveal key limitations of current LLMs in both functionality and security: (i) even the best model, OpenAI o1, achieves a mere 62% on code correctness; (ii) on average, we could successfully execute security exploits on around half of the correct programs generated by each LLM; and (iii) in less popular backend frameworks, models further struggle to generate correct and secure applications. Progress on BaxBench signifies important steps towards autonomous and secure software development with LLMs. Mark Vero, Niels Mündler, Victor Chibotaru, Veselin Raychev, Maximilian Baader, Nikola Jovanovic 0001, Martin T. Vechev |
ICML | 2 |
| 2025 | Type-Constrained Code Generation with Language ModelsabstractLarge language models (LLMs) have achieved notable success in code generation. However, they still frequently produce uncompilable output because their next-token inference procedure does not model formal aspects of code. Although constrained decoding is a promising approach to alleviate this issue, it has only been applied to handle either domain-specific languages or syntactic features of general-purpose programming languages. However, LLMs frequently generate code with typing errors, which are beyond the domain of syntax and generally hard to adequately constrain. To address this challenge, we introduce a type-constrained decoding approach that leverages type systems to guide code generation. For this purpose, we develop novel prefix automata and a search over inhabitable types, forming a sound approach to enforce well-typedness on LLM-generated code. We formalize our approach on a foundational simply-typed language and extend it to TypeScript to demonstrate practicality. Our evaluation on the HumanEval and MBPP datasets shows that our approach reduces compilation errors by more than half and significantly increases functional correctness in code synthesis, translation, and repair tasks across LLMs of various sizes and model families, including state-of-the-art open-weight models with more than 30B parameters. The results demonstrate the generality and effectiveness of our approach in constraining LLM code generation with formal rules of type systems. Niels Mündler, Hao Wang 0112, Koushik Sen, Dawn Song, Martin T. Vechev |
Proc. ACM Program. Lang. | 1 |
| 2024 | Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and MitigationabstractLarge language models (large LMs) are susceptible to producing text that contains hallucinated content. An important instance of this problem is self-contradiction, where the LM generates two contradictory sentences within the same context. In this work, we present a comprehensive investigation into self-contradiction for various instruction-tuned LMs, covering evaluation, detection, and mitigation. Our primary evaluation task is open-domain text generation, but we also demonstrate the applicability of our approach to shorter question answering. Our analysis reveals the prevalence of self-contradictions, e.g., in 17.7% of all sentences produced by ChatGPT. We then propose a novel prompting-based framework designed to effectively detect and mitigate self-contradictions. Our detector achieves high accuracy, e.g., around 80% F1 score when prompting ChatGPT. The mitigation algorithm iteratively refines the generated text to remove contradictory information while preserving text fluency and informativeness. Importantly, our entire framework is applicable to black-box LMs and does not require retrieval of external knowledge. Rather, our method complements retrieval-based methods, as a large portion of self-contradictions (e.g., 35.2% for ChatGPT) cannot be verified using online text. Our approach is practically effective and has been released as a push-button tool to benefit the public at https://chatprotect.ai/. Niels Mündler, Slobodan Jenko, Martin T. Vechev |
ICLR | 1 |
| 2024 | SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsabstractRigorous software testing is crucial for developing and maintaining high-quality code, making automated test generation a promising avenue for both improving software quality and boosting the effectiveness of code generation methods. However, while code generation with Large Language Models (LLMs) is an extraordinarily active research area, test generation remains relatively unexplored. We address this gap and investigate the capability of LLM-based Code Agents to formalize user issues into test cases. To this end, we propose a novel benchmark based on popular GitHub repositories, containing real-world issues, ground-truth bug-fixes, and golden tests. We find that LLMs generally perform surprisingly well at generating relevant test cases, with Code Agents designed for code repair exceeding the performance of systems designed specifically for test generation. Further, as test generation is a similar but more structured task than code generation, it allows for a more fine-grained analysis using issue reproduction rate and coverage changes, providing a dual metric for analyzing systems designed for code repair. Finally, we find that generated tests are an effective filter for proposed code fixes, doubling the precision of SWE-Agent. We release all data and code at https://github.com/logic-star-ai/SWT-Bench. Niels Mündler, Mark Niklas Müller, Martin T. Vechev |
NeurIPS | 1 |
| 2022 | A Verified Implementation of B+-Trees in Isabelle/HOL
Niels Mündler, Tobias Nipkow |
ICTAC | 1 |