VLDB 2026 Research / reviewers in the wild / expert
David Delgado 0003
dblp:332/5900
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2026
0009-0006-5633-2874ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How much does an LLM know about my programming language?abstractLarge Language Models (LLMs) are increasingly used for code generation, yet their support for programming languages is uneven, particularly for low-resource languages and infrequently used language constructs. Existing evaluation methodologies primarily assess functional correctness through test execution, which fails to reveal blind spots in a model's knowledge of a target language. In this paper, we introduce a systematic framework to identify and quantify the syntactic knowledge exhibited by an LLM with respect to a given programming language. We define syntactic coverage as a set of measurable indicators of how comprehensively language constructs are represented in datasets and generated code. Our approach analyzes code snippets against grammatical rules and structural features to compute interpretable coverage metrics, detect underrepresented constructs, and expose gaps in language support. The proposed framework supports comparative analysis across datasets, languages, and models through an automated pipeline that produces actionable diagnostic reports. Our experimental results demonstrate that syntactic coverage analysis reveals substantial differences in language support that are not captured by functional correctness alone, enabling more informed model selection, benchmark design, and dataset improvement strategies. David Delgado 0003, Loli Burgueño, Robert Clarisó |
SLE | 1 |
| 2026 | A framework for assessing the capabilities of code generation of constraint domain-specific languages with large language modelsabstractLarge language models (LLMs) can be used to support software development tasks, e.g. , through code completion or code generation. However, their effectiveness drops significantly when considering less popular programming languages such as domain-specific languages (DSLs). In this paper, we propose a generic framework for evaluating the capabilities of LLMs generating DSL code from textual specifications. The generated code is assessed from the perspectives of well-formedness and correctness. This framework is applied to a particular type of DSL, constraint languages, focusing our experiments on OCL and Alloy and comparing their results to those achieved for Python, a popular general-purpose programming language. Experimental results show that, in general, LLMs have better performance for Python than for OCL and Alloy. LLMs with smaller context windows such as open-source LLMs may be unable to generate constraint-related code, as this requires managing both the constraint and the domain model where it is defined. Moreover, some improvements to the code generation process such as code repair (asking an LLM to fix incorrect code) or multiple attempts (generating several candidates for each coding task) can improve the quality of the generated code. Meanwhile, other decisions like the choice of a prompt template have less impact. All these dimensions can be systematically analyzed using our evaluation framework, making it possible to decide the most effective way to set up code generation for a particular type of task. David Delgado 0003, Loli Burgueño, Robert Clarisó |
J. Syst. Softw. | 1 |