VLDB 2026 Research / reviewers in the wild / expert
José Antonio Hernández López
dblp:274/0650
· DBLP profile ↗
19ranked-venue papers
13as first author
18since 2021 · last 2026
0000-0003-2439-2136ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 13 first-author · 18 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Grounding Generative AI in Software Engineering: Are We There Yet?
Mootez Saad, José Antonio Hernández López, Boqi Chen, Neil A. Ernst, Dániel Varró, Tushar Sharma 0001 |
SANER | 2 |
| 2026 | Syntactic multilingual probing of pre-trained language models of codeabstractPre-trained language models (PLMs) have demonstrated remarkable abilities in coding tasks, establishing themselves as a state-of-the-art technique in machine learning for code. However, due to their deep neural network-based structure, PLMs function as black-box systems, making it crucial to understand the types of information they actually learn. Recent studies indicate that PLMs possess cross-lingual capabilities, allowing them to generalize to unseen programming languages and outperform monolingual models when trained in a multilingual setting. Nonetheless, the reasons behind these cross-lingual abilities remain largely uncharted and remain open questions. In this paper, we explore this phenomenon through a syntactic perspective. Specifically, we build on our prior work, the AST-Probe, a probing methodology that evaluates whether a PLM encodes the complete grammatical structure of a programming language. This probe identifies a syntactic subspace within the PLM’s vector representations, which is then used to reconstruct ASTs. We extend this approach in two ways. First, we conducted experiments on eight programming languages and eight PLMs and found that: (1) this syntactic structure can be extracted in all cases, (2) CodeBERT and GraphCodeBERT excel at encoding ASTs, and (3) syntactic knowledge resides in the middle layers of all PLMs, with a distribution that is independent of the programming language. Secondly, we mathematically adapt the AST-Probe to a multilingual setting and apply it to CodeBERT. Our findings provide evidence that CodeBERT learns cross-lingual representations of programming languages syntax. José Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, Houari Sahraoui |
J. Syst. Softw. | 1 |
| 2025 | The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed Languagesabstract[Motivation] Automated bug detection in dynamically typed languages such as Python is essential for maintaining code quality. The lack of mandatory type annotations in such languages can lead to errors that are challenging to identify early with traditional static analysis tools. Recent progress in deep neural networks has led to increased use of neural bug detectors. In statically typed languages, a type checker is integrated into the compiler and thus taken into consideration when the neural bug detector is designed for these languages. [Problem] However, prior studies overlook this aspect during the training and testing of neural bug detectors for dynamically typed languages. When an optional type checker is used, assessing existing neural bug detectors on bugs easily detectable by type checkers may impact their performance estimation. Moreover, including these bugs in the training set of neural bug detectors can shift their detection focus toward the wrong type of bugs. [Contribution] We explore the impact of type checking on various neural bug detectors for variable misuse bugs, a common type targeted by neural bug detectors. Existing synthetic and real-world datasets are type-checked to evaluate the prevalence of type-related bugs. Then, we investigate how type-related bugs influence the training and testing of the neural bug detectors. [Findings] Our findings indicate that existing bug detection datasets contain a significant proportion of type-related bugs. Building on this insight, we discover integrating the neural bug detector with a type checker can be beneficial, especially when the code is annotated with types. Further investigation reveals neural bug detectors perform better on type-related bugs than other bugs. Moreover, removing type-related bugs from the training data helps improve neural bug detectors' ability to identify bugs beyond the scope of type checkers. Boqi Chen, José Antonio Hernández López, Gunter Mussbacher, Dániel Varró |
ICSE | 2 |
| 2025 | SHERPA: A Model-Driven Framework for Large Language Model ExecutionabstractRecently, large language models (LLMs) have achieved widespread application across various fields. Despite their impressive capabilities, LLMs suffer from a lack of structured reasoning ability, particularly for complex tasks requiring domain-specific best practices, which are often unavailable in the training data. Although multi-step prompting methods incorporating human best practices, such as chain-of-thought and tree-of-thought, have gained popularity, they lack a general mechanism to control LLM behavior. In this paper, we propose SHERPA, a model-driven framework to improve the LLM performance on complex tasks by explicitly incorporating domain-specific best practices into hierarchical state machines. By structuring the LLM execution processes using state machines, SHERPA enables more fine-grained control over their behavior via rules or decisions driven by machine learning-based approaches, including LLMs. We show that SHERPA is applicable to a wide variety of tasks-specifically, code generation, class name generation, and question answering-replicating previously proposed approaches while further improving the performance. We demonstrate the effectiveness of SHERPA for the aforementioned tasks using various LLMs. Our systematic evaluation compares different state machine configurations against baseline approaches without state machines. Results show that integrating well-designed state machines significantly improves the quality of LLM outputs, and is particularly beneficial for complex tasks with well-established human best practices but lacking data used for training LLMs. Boqi Chen, Kua Chene, José Antonio Hernández López, Gunter Mussbacher, Dániel Varró, Amir Feizpour |
MODELS | 3 |
| 2025 | Generative AI in Simulation-Based Test Environments for Large-Scale Cyber-Physical Systems: An Industrial Study
Masoud Sadrnezhaad, José Antonio Hernández López, Torvald Mårtensson, Dániel Varró |
PROFES | 2 |
| 2025 | ModelXGlue: a benchmarking framework for ML tools in MDEabstractAbstract The integration of machine learning (ML) into model-driven engineering (MDE) holds the potential to enhance the efficiency of modelers and elevate the quality of modeling tools. However, a consensus is yet to be reached on which MDE tasks can derive substantial benefits from ML and how progress in these tasks should be measured. This paper introduces ModelXGlue , a dedicated benchmarking framework to empower researchers when constructing benchmarks for evaluating the application of ML to address MDE tasks. A benchmark is built by referencing datasets and ML models provided by other researchers, and by selecting an evaluation strategy and a set of metrics. ModelXGlue is designed with automation in mind and each component operates in an isolated execution environment (via Docker containers or Python environments), which allows the execution of approaches implemented with diverse technologies like Java, Python, R, etc. We used ModelXGlue to build reference benchmarks for three distinct MDE tasks: model classification, clustering, and feature name recommendation. To build the benchmarks we integrated existing third-party approaches in ModelXGlue . This shows that ModelXGlue is able to accommodate heterogeneous ML models, MDE tasks and different technological requirements. Moreover, we have obtained, for the first time, comparable results for these tasks. Altogether, it emerges that ModelXGlue is a valuable tool for advancing the understanding and evaluation of ML tools within the context of MDE. José Antonio Hernández López, Jesús Sánchez Cuadrado, Riccardo Rubei, Davide Di Ruscio |
Softw. Syst. Model. | 1 |
| 2025 | Experimenting with modeling-specific word embeddings
José Antonio Hernández López, Carlos Durá, Jesús Sánchez Cuadrado |
Softw. Syst. Model. | 1 |
| 2025 | On Inter-Dataset Code Duplication and Data Leakage in Large Language ModelsabstractMotivation.Large language models (LLMs) have exhibited remarkable proficiency in diverse software engineering (SE) tasks, such as code summarization, code translation, and code search. Handling such tasks typically involves acquiring foundational coding knowledge on large, general-purpose datasets during a pre-training phase, and subsequently refining on smaller, task-specific datasets as part of a fine-tuning phase.Problem statement.Data leakagei.e.,using information of the test set to perform the model training, is a well-known issue in training of machine learning models. A manifestation of this issue is the intersection of the training and testing splits. Whileintra-datasetcode duplication examines this intersection within a given dataset and has been addressed in prior research,inter-dataset code duplication, which gauges the overlap between different datasets, remains largely unexplored. If this phenomenon exists, it could compromise the integrity ofLLMevaluations because of the inclusion of fine-tuning test samples that were already encountered during pre-training, resulting in inflated performance metrics.Contribution.This paper explores the phenomenon of inter-dataset code duplication and its impact on evaluatingLLMs across diverseSEtasks.Study design.We conduct an empirical study using theCodeSearchNetdataset (csn), a widely adopted pre-training dataset, and five fine-tuning datasets used for variousSEtasks. We first identify the intersection between the pre-training and fine-tuning datasets using a deduplication process. Next, we pre-train two versions ofLLMs using a subset ofcsn: one leakyLLM, which includes the identified intersection in its pre-training set, and one non-leakyLLMthat excludes these samples. Finally, we fine-tune both models and compare their performances using fine-tuning test samples that are part of the intersection.Results.Our findings reveal a potential threat to the evaluation ofLLMs across multipleSEtasks, stemming from the inter-dataset code duplication phenomenon. We also demonstrate that this threat is accentuated by the chosen fine-tuning technique. Furthermore, we provide evidence that open-source models such asCodeBERT,GraphCodeBERT, andUnixCodercould be affected by inter-dataset duplication. Based on our findings, we delve into prior research that may be susceptible to this threat. Additionally, we offer guidance toSEresearchers on strategies to prevent inter-dataset code duplication. José Antonio Hernández López, Boqi Chen, Mootez Saad, Tushar Sharma 0001, Dániel Varró |
IEEE Trans. Software Eng. | 1 |
| 2025 | Why Do Machine Learning Notebooks Crash? An Empirical Study on Public Python Jupyter NotebooksabstractJupyter notebooks have become central in data science, integrating code, text and output in a flexible environment. With the rise of machine learning (ML), notebooks are increasingly used for prototyping and data analysis. However, due to their dependence on complex ML libraries and the flexible notebook semantics that allow cells to be run in any order, notebooks are susceptible to software bugs that may lead to program crashes. This paper presents a comprehensive empirical study focusing on crashes in publicly available Python ML notebooks. We collect 64,031 notebooks containing 92,542 crashes from GitHub and Kaggle, and manually analyze a sample of 746 crashes across various aspects, including crash types and root causes. Our analysis identifies unique ML-specific crash types, such as tensor shape mismatches and dataset value errors that violate API constraints. Additionally, we highlight unique root causes tied to notebook semantics, including out-of-order execution and residual errors from previous cells, which have been largely overlooked in prior research. Furthermore, we identify the most error-prone ML libraries, and analyze crash distribution across ML pipeline stages. We find that over 40% of crashes stem from API misuse and notebook-specific issues. Crashes frequently occur when using ML libraries like TensorFlow/Keras and Torch. Additionally, over 70% of the crashes occur during data preparation, model training, and evaluation or prediction stages of the ML pipeline, while data visualization errors tend to be unique to ML notebooks. Willem Meijer, José Antonio Hernández López, Ulf Nilsson, Dániel Varró |
IEEE Trans. Software Eng. | 3 |
| 2024 | ModelMate: A recommender for textual modeling languages based on pre-trained language modelsabstractCurrent DSL environments lack smart editing facilities intended to enhance modeler productivity and cannot keep pace of current developments of integrated development environments based on AI. In this paper, we propose an approach to address this shortcoming through a recommender system specifically tailored for textual DSLs based on the fine-tuning of pre-trained language models. We identify three main tasks: identifier suggestion, line completion, and block completion, which we implement over the same fine-tuned model and we propose a workflow to apply these tasks to any textual DSL. We have evaluated our approach with different pre-trained models for three DSLs: Emfatic, Xtext and a DSL to specify domain entities, showing that the system performs well and provides accurate suggestions. We compare it against existing approaches in the feature name recommendation task showing that our system outperforms the alternatives. Moreover, we evaluate the inference time of our approach obtaining low latencies, which makes the system adequate for live assistance. Finally, we contribute a concrete recommender, named ModelMate, which implements the training, evaluation and inference steps of the workflow as well as providing integration into Eclipse-based textual editors. Carlos Durá, José Antonio Hernández López, Jesús Sánchez Cuadrado |
MODELS | 2 |
| 2024 | Text2VQL: Teaching a Model Query Language to Open-Source Language Models with ChatGPTabstractWhile large language models (LLMs) like ChatGPT has demonstrated impressive capabilities in addressing various software engineering tasks, their use in a model-driven engineering (MDE) context is still in an early stage. Since the technology is proprietary and accessible solely through an API, its use may be incompatible with the strict protection of intellectual properties in industrial models. While there are open-source LLM alternatives, they often lack the power of proprietary models and require extensive data fine-tuning to realize their full potential. Furthermore, open-source datasets tailored for MDE tasks are scarce, posing challenges for training such models effectively. José Antonio Hernández López, Máté Földiák, Dániel Varró |
MODELS | 1 |
| 2024 | ModelSet: A labelled dataset of software models for machine learningabstractCurated collections of models are essential for the success of Machine Learning (ML) and Data Analytics in Model-Driven Engineering (MDE). However, current datasets are either too small or not properly curated. In this paper, we present ModelSet, a dataset composed of 5,466 Ecore models and 5,120 UML models which have been manually labelled to support ML tasks. We describe the structure of the dataset and explain how to use the associated library to develop ML applications in Python. Finally, we present some applications which can be addressed using ModelSet. Tool Website: https://github.com/modelset José Antonio Hernández López, Javier Luis Cánovas Izquierdo, Jesús Sánchez Cuadrado |
Sci. Comput. Program. | 1 |
| 2023 | Generating Structurally Realistic Models With Deep Autoregressive NetworksabstractModel generators are important tools in model-based systems engineering to automate the creation of software models for tasks like testing and benchmarking. Previous works have established four properties that a generator should satisfy: consistency, diversity, scalability, and structural realism. Although several generators have been proposed, none of them is focused on realism. As a result, automatically generated models are typically simple and appear synthetic. This work proposes a new architecture for model generators which is specifically designed to be structurally realistic. Given a dataset consisting of several models deemed as real models, this type of generators is able to produce new models which are structurally similar to the models in the dataset, but are fundamentally novel models. Our implementation, namedModelMime(M2), is based on a deep autoregressive model which combines a Graph Neural Network with a Recurrent Neural Network. We decompose each model into a sequence of edit operations, and the neural network is trained in the task of predicting the next edit operation given a partial model. At inference time, the system produces new models by sampling edit operations and iteratively completing the model. We have evaluated M2 with respect to three state-of-the-art generators, showing that 1) our generator outperforms the others in terms of the structurally realistic property 2) the models generated by M2 are most of the time consistent, 3) the diversity of the generated models is at least the same as the real ones and, 4) the generation process is scalable once the generator is trained. José Antonio Hernández López, Jesús Sánchez Cuadrado |
IEEE Trans. Software Eng. | 1 |
| 2022 | AST-Probe: Recovering abstract syntax trees from hidden representations of pre-trained language modelsabstractThe objective of pre-trained language models is to learn contextual representations of textual data. Pre-trained language models have become mainstream in natural language processing and code modeling. Using probes, a technique to study the linguistic properties of hidden vector spaces, previous works have shown that these pre-trained language models encode simple linguistic properties in their hidden representations. However, none of the previous work assessed whether these models encode the whole grammatical structure of a programming language. In this paper, we prove the existence of a syntactic subspace, lying in the hidden representations of pre-trained language models, which contain the syntactic information of the programming language. We show that this subspace can be extracted from the models’ representations and define a novel probing method, the AST-Probe, that enables recovering the whole abstract syntax tree (AST) of an input code snippet. In our experimentations, we show that this syntactic subspace exists in five state-of-the-art pre-trained language models. In addition, we highlight that the middle layers of the models are the ones that encode most of the AST information. Finally, we estimate the optimal size of this syntactic subspace and show that its dimension is substantially lower than those of the models’ representation spaces. This suggests that pre-trained language models use a small part of their representation spaces to encode syntactic information of the programming languages. José Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, Houari Sahraoui |
ASE | 1 |
| 2022 | Machine learning methods for model classification: a comparative studyabstractIn the quest to reuse modeling artifacts, academics and industry have proposed several model repositories over the last decade. Different storage and indexing techniques have been conceived to facilitate searching capabilities to help users find reusable artifacts that might fit the situation at hand. In this respect, machine learning (ML) techniques have been proposed to categorize and group large sets of modeling artifacts automatically. This paper reports the results of a comparative study of different ML classification techniques employed to automatically label models stored in model repositories. We have built a framework to systematically compare different ML models (feed-forward neural networks, graph neural networks, k-nearest neighbors, support version machines, etc.) with varying model encodings (TF-IDF, word embeddings, graphs and paths). We apply this framework to two datasets of about 5,000 Ecore and 5,000 UML models. We show that specific ML models and encodings perform better than others depending on the characteristics of the available datasets (e.g., the presence of duplicates) and on the goals to be achieved. José Antonio Hernández López, Riccardo Rubei, Jesús Sánchez Cuadrado, Davide Di Ruscio |
MoDELS | 1 |
| 2022 | An efficient and scalable search engine for modelsabstractAbstract Search engines extract data from relevant sources and make them available to users via queries. A search engine typically crawls the web to gather data, analyses and indexes it and provides some query mechanism to obtain ranked results. There exist search engines for websites, images, code, etc., but the specific properties required to build a search engine for models have not been explored much. In the previous work, we presented MAR, a search engine for models which has been designed to support a query-by-example mechanism with fast response times and improved precision over simple text search engines. The goal of MAR is to assist developers in the task of finding relevant models. In this paper, we report new developments of MAR which are aimed at making it a useful and stable resource for the community. We present the crawling and analysis architecture with which we have processed about 600,000 models. The indexing process is now incremental and a new index for keyword-based search has been added. We have also added a web user interface intended to facilitate writing queries and exploring the results. Finally, we have evaluated the indexing times, the response time and search precision using different configurations. MAR has currently indexed over 500,000 valid models of different kinds, including Ecore meta-models, BPMN diagrams, UML models and Petri nets. MAR is available at http://mar-search.org . José Antonio Hernández López, Jesús Sánchez Cuadrado |
Softw. Syst. Model. | 1 |
| 2022 | ModelSet: a dataset for machine learning in model-driven engineeringabstractAbstract The application of machine learning (ML) algorithms to address problems related to model-driven engineering (MDE) is currently hindered by the lack of curated datasets of software models. There are several reasons for this, including the lack of large collections of good quality models, the difficulty to label models due to the required domain expertise, and the relative immaturity of the application of ML to MDE. In this work, we present ModelSet , a labelled dataset of software models intended to enable the application of ML to address software modelling problems. To create it we have devised a method designed to facilitate the exploration and labelling of model datasets by interactively grouping similar models using off-the-shelf technologies like a search engine. We have built an Eclipse plug-in to support the labelling process, which we have used to label 5,466 Ecore meta-models and 5,120 UML models with its category as the main label plus additional secondary labels of interest. We have evaluated the ability of our labelling method to create meaningful groups of models in order to speed up the process, improving the effectiveness of classical clustering methods. We showcase the usefulness of the dataset by applying it in a real scenario: enhancing the MAR search engine. We use ModelSet to train models able to infer useful metadata to navigate search results. The dataset and the tooling are available at https://figshare.com/s/5a6c02fa8ed20782935c and a live version at http://modelset.github.io . José Antonio Hernández López, Javier Luis Cánovas Izquierdo, Jesús Sánchez Cuadrado |
Softw. Syst. Model. | 1 |
| 2021 | Towards the Characterization of Realistic Model Generators using Graph Neural NetworksabstractThe automatic generation of software models is an important element in many software and systems engineering scenarios such as software tool certification, validation of cyber-physical systems, or benchmarking graph databases. Several model generators are nowadays available, but the topic of whether they generate realistic models has been little studied. The state-of-the-art approach to check the realistic property in software models is to rely on simple comparisons using graph metrics and statistics. This generates a bottleneck due to the compression of all the information contained in the model into a small set of metrics. Furthermore, there is a lack of interpretation in these approaches since there are no hints of why the generated models are not realistic. Therefore, in this paper, we tackle the problem of assessing how realistic a generator is by mapping it to a classification problem in which a Graph Neural Network (GnN) will be trained to distinguish between the two sets of models (real and synthetic ones). Then, to assess how realistic a generator is we perform the Classifier Two-Sample Test (C2ST). Our approach allows for interpretation of the results by inspecting the attention layer of the GNN. We use our approach to assess four state-of-the-art model generators applied to three different domains. The results show that none of the generators can be considered realistic. José Antonio Hernández López, Jesús Sánchez Cuadrado |
MoDELS | 1 |
| 2020 | MAR: a structure-based search engine for modelsabstractThe availability of shared software models provides opportunities for reusing, adapting and learning from them. Public models are typically stored in a variety of locations, including model repositories, regular source code repositories, web pages, etc. To profit from them developers need effective search mechanisms to locate the models relevant for their tasks. However, to date, there has been little success in creating a generic and efficient search engine specially tailored to the modelling domain. José Antonio Hernández López, Jesús Sánchez Cuadrado |
MoDELS | 1 |