Vincenzo Scotti 0001

dblp:241/0860-1 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2027
0000-0002-8765-604XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Large language models in model-driven engineering: a systematic mapping study
abstract
Abstract The application of Large Language Models (LLMs) in Model-Driven Engineering (MDE) has emerged as a rapidly evolving research area. While existing systematic literature reviews have examined specific technical approaches, a comprehensive mapping of the broader research landscape (e.g., development trends) remains lacking. This study presents a systematic mapping study of LLM applications in MDE, analyzing 86 primary studies collected from five databases, covering publications from 2022 to early 2026. Guided by five research questions, we characterize the field across five dimensions: MDE task distribution and research contribution types, LLM technologies and interaction strategies, artifact representation and processing, validation practices, and publication landscape. Our findings reveal that current LLM4MDE research is heavily concentrated on Model Generation, while tasks such as Model Migration, DSL Engineering, and Metamodeling remain marginal. Most approaches rely on black-box OpenAI models accessed via remote APIs and adapted through prompt engineering, with fine-tuning and retrieval-augmented generation rarely employed. Inputs are predominantly natural-language artifacts, while outputs are model-oriented but usually expressed in lightweight textual formats rather than native MDE exchange formats. Validation is centered on quantitative experimentation, with 42% of studies reporting no baseline and cost efficiency reported in fewer than one quarter of studies. The field has grown rapidly, from one paper in 2022 to 42 in 2025, with research concentrated in Europe and Canada and limited industry involvement. Based on these findings, we identify gaps and opportunities across task coverage, technical configuration, and evaluation practice, offering a knowledge map to guide future work in this cross-disciplinary field.
Yuhong Fu, Haowei Cheng, Maximilian Hummel, Vincenzo Scotti 0001, Nathan Hagel, Georg Grossmann, Markus Stumptner, Regina Hebig, Daniel Strüber 0001, Anne Koziolek
Empir. Softw. Eng.6
2026 A Conceptual Model of Operationalizable Trustworthiness for Autonomous Systems
abstract
Autonomous systems are increasingly deployed across critical domains, yet research on trustworthy autonomous systems (TAS) appears to lack an agreed-upon, operationalizable taxonomy for trustworthiness characteristics. While trustworthy artificial intelligence (TAI) frameworks offer measurable characteristics with more established metrics, TAS literature tends to be fragmented with high-level models that can be difficult to apply in practice. This paper investigates the extent of alignment in TAS trustworthiness research and identifies potential sources of misalignment, including differences in system definitions, autonomy properties, stakeholder perspectives, and domain requirements. Based on this analysis, we synthesize a foundational two-pillar core of trustworthiness, the system’s ability to dependably do what it is intended to do and the extent to which stakeholders can assess and justify confidence in this behavior. We propose an initial operationalization of this core using characteristics adapted from established TAI frameworks for which metrics and evaluation methods exist, and outline how it can be applied to specific application contexts. An example shows how stakeholder analysis guides characteristic adaptation.
Bálint Máté, Diego Perez-Palacin, Vincenzo Scotti 0001, Raffaela Mirandola
COMPSAC3
2025 How Toxic Can You Get? Search-Based Toxicity Testing for Large Language Models
abstract
Language is a deep-rooted means of perpetration of stereotypes and discrimination. Large Language Models (LLMs), now a pervasive technology in our everyday lives, can cause extensive harm when prone to generating toxic responses. The standard way to address this issue is to align the LLM, which, however, dampens the issue without constituting a definitive solution. Therefore, testing LLM even after alignment efforts remains crucial for detecting any residual deviations with respect to ethical standards. We present EvoTox, an automated testing framework for LLMs’ inclination to toxicity, providing a way to quantitatively assess how much LLMs can be pushed towards toxic responses even in the presence of alignment. The framework adopts an iterative evolution strategy that exploits the interplay between two LLMs, the System Under Test (SUT) and the Prompt Generator steering SUT responses toward higher toxicity. The toxicity level is assessed by an automated oracle based on an existing toxicity classifier. We conduct a quantitative and qualitative empirical evaluation using five state-of-the-art LLMs as evaluation subjects having increasing complexity (7–671B parameters). Our quantitative evaluation assesses the cost-effectiveness of four alternative versions of EvoTox against existing baseline methods, based on random search, curated datasets of toxic prompts, and adversarial attacks. Our qualitative assessment engages human evaluators to rate the fluency of the generated prompts and the perceived toxicity of the responses collected during the testing sessions. Results indicate that the effectiveness, in terms of detected toxicity level, is significantly higher than the selected baseline methods (effect size up to 1.0 against random search and up to 0.99 against adversarial attacks). Furthermore, EvoTox yields a limited cost overhead (from 22% to 35% on average).This work includes examples of toxic degeneration by LLMs, which may be considered profane or offensive to some readers. Reader discretion is advised.
Simone Corbo, Luca Bancale, Valeria De Gennaro, Livia Lestingi, Vincenzo Scotti 0001, Matteo Camilli
IEEE Trans. Software Eng.5