VLDB 2026 Research / reviewers in the wild / expert
Jasper Dekoninck
dblp:361/7298
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0009-0621-170XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 62% Learning theory · 13% Efficient and distributed learning · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Emerging computing paradigms · 100% | |
| Theoretical computer science
2 papers |
Logic in computer science · 66% Mathematical optimization · 34% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model evaluation |
2.5 | 3 | 2025 | MathConstruct: Challenging LLM Reasoning with Constructive Proofs · ICML 2025 Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation · ICLR 2025 ConStat: Performance-Based Contamination Detection in Large Language Models · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › inference efficiency
inference cost optimization |
0.9 | 1 | 2025 | A Unified Approach to Routing and Cascading for LLMs · ICML 2025 |
Natural language and speech › Language models and text generation › evaluation of language models › reasoning evaluation
mathematical reasoning benchmark |
0.9 | 1 | 2025 | MathConstruct: Challenging LLM Reasoning with Constructive Proofs · ICML 2025 |
Machine learning › Learning theory
model selection |
0.9 | 1 | 2025 | A Unified Approach to Routing and Cascading for LLMs · ICML 2025 |
Logic in computer science › proof theory
constructive proof |
0.9 | 1 | 2025 | MathConstruct: Challenging LLM Reasoning with Constructive Proofs · ICML 2025 |
Machine learning › Trustworthy machine learning › Data-centric AI
data contamination detection |
0.8 | 1 | 2024 | ConStat: Performance-Based Contamination Detection in Large Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation
text generation |
0.8 | 1 | 2024 | Controlled Text Generation via Language Model Arithmetic · ICLR 2024 |
Emerging computing paradigms › quantum computer architecture
quantum circuit synthesis |
0.8 | 1 | 2024 | Synthetiq: Fast and Versatile Quantum Circuit Synthesis · Proc. ACM Program. Lang. 2024 |
Emerging computing paradigms
quantum computing |
0.8 | 1 | 2024 | Synthetiq: Fast and Versatile Quantum Circuit Synthesis · Proc. ACM Program. Lang. 2024 |
Emerging computing paradigms › quantum computer architecture › quantum circuit synthesis
quantum gate decomposition |
0.8 | 1 | 2024 | Synthetiq: Fast and Versatile Quantum Circuit Synthesis · Proc. ACM Program. Lang. 2024 |
Mathematical optimization
combinatorial optimization |
0.2 | 1 | 2024 | Synthetiq: Fast and Versatile Quantum Circuit Synthesis · Proc. ACM Program. Lang. 2024 |
Mathematical optimization › metaheuristic optimization
simulated annealing |
0.2 | 1 | 2024 | Synthetiq: Fast and Versatile Quantum Circuit Synthesis · Proc. ACM Program. Lang. 2024 |
Methods — techniques the papers use, named apart from their topics
problem variation generation · 1.7automated verification · 1.7simulated annealing · 1.5circuit simplification · 1.5maximum a posteriori estimation · 0.9statistical testing · 0.8speculative sampling · 0.8model arithmetic · 0.8benchmark comparison · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM EvaluationabstractRating-based human evaluation has become an essential tool to accurately evaluate the impressive performance of large language models (LLMs). However, current rating systems suffer from several important limitations: first, they fail to account for biases that significantly influence evaluation results, second, they require large and expensive preference datasets to obtain accurate ratings, and third, they do not facilitate meaningful comparisons of model ratings across different tasks. To address these issues, we introduce Polyrating, an expressive and flexible rating system based on maximum a posteriori estimation that enables a more nuanced and thorough analysis of model performance at lower costs. Polyrating can detect and quantify biases affecting human preferences, ensuring fairer model comparisons. Further, Polyrating can reduce the cost of human evaluations by up to $41$% for new models and up to $77$% for new tasks by leveraging existing benchmark scores. Lastly, Polyrating enables direct comparisons of ratings across different tasks, providing a comprehensive understanding of an LLMs' strengths, weaknesses, and relative performance across different applications. Jasper Dekoninck, Maximilian Baader, Martin T. Vechev |
ICLR | 1 |
| 2025 | MathConstruct: Challenging LLM Reasoning with Constructive ProofsabstractWhile Large Language Models (LLMs) demonstrate impressive performance in mathematics, existing math benchmarks come with significant limitations. Many focus on problems with fixed ground-truth answers, and are often saturated due to problem simplicity or the viability of guessing or memorization. Crucially, they capture only a narrow subset of relevant math problems. To address this research gap, we introduce MathConstruct, a new benchmark of 127 challenging problems sourced from various math competitions, which targets constructive proofs, a widely encountered problem type requiring the construction of mathematical objects with specific properties. These proofs are particularly suitable for LLM evaluation, as solution correctness can be easily verified. Our automated verifiers also enable MathConstruct to generate problem variations, used to evaluate robustness. State-of-the-art LLMs solve only 41% of MathConstruct problems, highlighting its complexity and importance for LLM evaluation. Mislav Balunovic, Jasper Dekoninck, Nikola Jovanovic 0001, Ivo Petrov, Martin T. Vechev |
ICML | 2 |
| 2025 | A Unified Approach to Routing and Cascading for LLMsabstractThe availability of a wide range of large language models (LLMs) embedded in various agentic systems has significantly increased the potential of model selection strategies to improve the cost-performance tradeoff. Existing strategies involve either routing, where a single model is chosen per query, or cascading, which sequentially runs increasingly larger models until a satisfactory answer is found. However, current approaches face three key limitations: they (1) lack formal proofs of optimality, (2) fail to identify the conditions under which these strategies are most effective to improve the cost-performance tradeoff, and (3) are unable to combine both paradigms for further improvements. To address these issues, we first derive a novel optimal strategy for cascading and prove the optimality of an existing routing strategy. Further, we propose *cascade routing*, a unified framework that integrates routing and cascading into a theoretically optimal strategy. Through our analysis, we identify good quality estimators as the critical factor for the success of model selection paradigms. Finally, in our experiments, we show that cascade routing consistently outperforms the individual approaches by a large margin and we analyze quality estimators to determine when routing and/or cascading are useful paradigms for model selection. Jasper Dekoninck, Maximilian Baader, Martin T. Vechev |
ICML | 1 |
| 2025 | MathArena: Evaluating LLMs on Uncontaminated Math CompetitionsabstractThe rapid advancement of reasoning capabilities in large language models (LLMs) has led to notable improvements on mathematical benchmarks. However, many of the most commonly used evaluation datasets (e.g., AIME 2024) are widely available online, making it difficult to disentangle genuine reasoning from potential memorization. Furthermore, these benchmarks do not evaluate proof-writing capabilities, which are crucial for many mathematical tasks.To address this, we introduce MathArena, a new benchmark based on the following key insight: recurring math competitions provide a stream of high-quality, challenging problems that can be used for real-time evaluation of LLMs. By evaluating models as soon as new problems are released, we effectively eliminate the risk of contamination.Using this framework, we find strong signs of contamination in AIME 2024. Nonetheless, evaluations on harder competitions, such as CMIMC 2025, demonstrate impressive reasoning capabilities in top-performing models.MathArena is also the first benchmark for proof-writing capabilities. On IMO 2025, top models achieve slightly less than 40\%, demonstrating both notable progress and significant room for improvement.So far, we have evaluated over $50$ models across seven competitions, totaling $162$ problems. As an evolving benchmark, MathArena will continue to track the progress of LLMs on newly released competitions, ensuring rigorous and up-to-date evaluation of mathematical reasoning. Mislav Balunovic, Jasper Dekoninck, Ivo Petrov, Nikola Jovanovic 0001, Martin T. Vechev |
NeurIPS | 2 |
| 2024 | Controlled Text Generation via Language Model ArithmeticabstractAs Large Language Models (LLMs) are deployed more widely, customization with respect to vocabulary, style, and character becomes more important. In this work, we introduce model arithmetic, a novel inference framework for composing and biasing LLMs without the need for model (re)training or highly specific datasets. In addition, the framework allows for more precise control of generated text than direct prompting and prior controlled text generation (CTG) techniques. Using model arithmetic, we can express prior CTG techniques as simple formulas and naturally extend them to new and more effective formulations. Further, we show that speculative sampling, a technique for efficient LLM sampling, extends to our setting. This enables highly efficient text generation with multiple composed models with only marginal overhead over a single model. Our empirical evaluation demonstrates that model arithmetic allows fine-grained control of generated text while outperforming state-of-the-art on the task of toxicity reduction. We release an open source easy-to-use implementation of our framework at https://github.com/eth-sri/language-model-arithmetic. Jasper Dekoninck, Marc Fischer 0002, Luca Beurer-Kellner, Martin T. Vechev |
ICLR | 1 |
| 2024 | ConStat: Performance-Based Contamination Detection in Large Language ModelsabstractPublic benchmarks play an essential role in the evaluation of large language models. However, data contamination can lead to inflated performance, rendering them unreliable for model comparison. It is therefore crucial to detect contamination and estimate its impact on measured performance. Unfortunately, existing detection methods can be easily evaded and fail to quantify contamination. To overcome these limitations, we propose a novel definition of *contamination as artificially inflated and non-generalizing benchmark performance* instead of the inclusion of benchmark samples in the training data. This perspective enables us to detect *any* model with inflated performance, i.e., performance that does not generalize to rephrased samples, synthetic samples from the same distribution, or different benchmarks for the same task. Based on this insight, we develop ConStat, a statistical method that reliably detects and quantifies contamination by comparing performance between a primary and reference benchmark relative to a set of reference models. We demonstrate the effectiveness of ConStat in an extensive evaluation of diverse model architectures, benchmarks, and contamination scenarios and find high levels of contamination in multiple popular models including Mistral, Llama, Yi, and the top-3 Open LLM Leaderboard models. Jasper Dekoninck, Mark Niklas Müller, Martin T. Vechev |
NeurIPS | 1 |
| 2024 | Synthetiq: Fast and Versatile Quantum Circuit SynthesisabstractTo implement quantum algorithms on quantum computers it is crucial to decompose their operators into the limited gate set supported by those computers. Unfortunately, existing works automating this essential task are generally slow and only applicable to narrow use cases.We present Synthetiq, a method to synthesize quantum circuits implementing a given specification over arbitrary finite gate sets, which is faster and more versatile than existing works. Synthetiq utilizes Simulated Annealing instantiated with a novel, domain-specific energy function that allows developers to leverage partial specifications for better efficiency. Synthetiq further couples this synthesis method with a custom simplification pass, to ensure efficiency of the found circuits. We experimentally demonstrate that Synthetiq can generate better implementations than were previously known for multiple relevant quantum operators including RCCCX, CCT, CCiSWAP, C√SWAP, and C√iSWAP. Our extensive evaluation also demonstrates Synthetiq frequently outperforms a wide variety of more specialized tools in their own domains, including (i) the well-studied task of synthesizing fully specified operators in the Clifford+T gate set, (ii) є-approximate synthesis of multi-qubit operators in the same gate set, and (iii) synthesis tasks with custom gate sets. On all those tasks, Synthetiq is typically one to two orders of magnitude faster than previous state-of-the-art and can tackle problems that were previously out of the reach of any synthesis tool. Anouk Paradis, Jasper Dekoninck, Benjamin Bichsel, Martin T. Vechev |
Proc. ACM Program. Lang. | 2 |