VLDB 2026 Research / reviewers in the wild / expert
Shahin Honarvar
dblp:276/3442
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-7421-8462ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Correct-Consistency and Robustness in Code-Generating LLMsabstractEnsuring the reliability of large language models (LLMs) for code generation is crucial for their safe integration into software engineering practices, especially in safety-critical domains. Despite advances in LLM tuning, they frequently generate incorrect code, raising concerns about robustness and trustworthiness. This PhD research introduces Turbulence, a novel benchmark and evaluation framework designed to assess both the correct-consistency and robustness of LLMs through a structured coding question neighbourhood approach. By evaluating model performance across sets of semantically related but non-equivalent coding tasks, Turbulence identifies discontinuities in LLM generalisation, revealing patterns of success and failure that standard correctness evaluations often overlook. Applied to 22 instruction-tuned LLMs across Python coding question neighbourhoods, the benchmark highlights significant variability in correctness, including error patterns persisting even under deterministic settings. Future work will extend the question neighbourhood concept to Capture The Flag (CTF) challenges, enabling a deeper analysis of model reasoning capabilities in progressively complex tasks. This extension has attracted interest from the UK AI Safety Institute, which recognises the frame-work's potential for advancing rigorous evaluation methodologies in the context of safe and trusted AI for software engineering. Shahin Honarvar |
ICST | 1 |
| 2025 | Turbulence: Systematically and Automatically Testing Instruction-Tuned Large Language Models for CodeabstractWe present a method for systematically evaluating the correctness and robustness of instruction-tuned large language models (LLMs) for code generation via a new benchmark, Turbulence. Turbulence consists of a large set of natural language question templates, each of which is a programming problem, parameterised so that it can be asked in many different forms. Each question template has an associated test oracle that judges whether a code solution returned by an LLM is correct. Thus, from a single question template, it is possible to ask an LLM a neighbourhood of very similar programming questions, and assess the correctness of the result returned for each question. This allows gaps in an LLM's code generation abilities to be identified, including anomalies where the LLM correctly solves almost all questions in a neighbourhood but fails for particular parameter instantiations. We present experiments against five LLMs from OpenAI, Cohere and Meta, each at two temperature configurations. Our findings show that, across the board, Turbulence is able to reveal gaps in LLM reasoning ability. This goes beyond merely highlighting that LLMs sometimes produce wrong code (which is no surprise): by systematically identifying cases where LLMs are able to solve some problems in a neighbourhood but do not manage to generalise to solve the whole neighbourhood, our method is effective at highlighting robustness issues. We present data and examples that shed light on the kinds of mistakes that LLMs make when they return incorrect code results. Shahin Honarvar, Mark van der Wilk, Alastair F. Donaldson |
ICST | 1 |
| 2025 | The "Question Neighbourhood" Approach for Systematic Evaluation of Code-Generating LLMsabstractWe present the concept of aquestion neighbourhoodfor systematically evaluating instruction-tuned large language models (LLMs) for code generation via a new benchmark, Turbulence. Turbulence consists of a large set of natural languagequestion templates, each of which is a programming problem, parameterised so that it can be asked in many different forms. Each question template has an associatedtest oraclethat judges whether a code solution returned by an LLM is correct. Thus, from a single question template, it is possible to ask an LLM aneighbourhoodof very similar programming questions, and assess the correctness of the result returned for each question. This allows gaps in an LLM’s code generation abilities to be identified, includinganomalieswhere the LLM correctly solvesmanyquestions in a neighbourhood but fails for particular parameter instantiations. We present experiments against 22 state-of-the-art proprietary and open-source LLMs, each at two temperature configurations. Our evaluation is based on three complementary scores: accuracy score, correctness-potential score, and consistent-correctness score. Our findings show that, across the board, Turbulence is able to reveal cases where LLMs do not behave in a correct and consistent manner, highlighting gaps in their reasoning ability. This goes beyond merely highlighting that LLMs sometimes produce wrong code (which is no surprise): by systematically identifying cases where LLMs are able to solve some problems in a neighbourhood but do not manage to generalise to solve the whole neighbourhood, our method provides detailed insight into the behavioural characteristics of current code-generating LLMs. We present data and examples that shed light on the kinds of mistakes that LLMs make when they return incorrect code results. Shahin Honarvar, Marek Rei, Alastair F. Donaldson |
IEEE Trans. Software Eng. | 1 |