VLDB 2026 Research / reviewers in the wild / expert
Darko Durisic
dblp:128/3213
· DBLP profile ↗
13ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0002-3901-873XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Train-Test Data Leakage in Automotive Image DatasetsabstractReliable evaluation of machine learning (ML)-enabled perception systems for intelligent vehicles critically depends on the integrity of training and test datasets. A major risk arises when near-duplicate or visually similar images appear across subsets, leading to inflated performance estimates. This study systematically quantifies train-test similarity in six widely used automotive datasets - KITTI, ZOD, BDD100k, ONCE, Cirrus, and SODA10M - in their default splits. We employ perceptual hashing (pHash) and deep feature embeddings to measure image-level redundancy. Results show 4,513 pairs of KITTI and 562 of ZOD images were almost identical, corresponding to 25% and 5% of their test images, respectively. The other examined datasets contain only at most 3 pairs (almost 0%) of almost identical images in their existing train-test splits. These findings underscore the need for similarity analysis during dataset preparation, particularly for video-based collections with strong spatio-temporal dependencies. By exposing dataset-specific risks of data leakage in popular datasets, this study contributes practical insights for both dataset curators and ML practitioners. These insights are valuable when using public benchmark datasets in safety-critical domains such as autonomous driving (AD). Md. Abu Ahammed Babu, Miroslaw Staron, András Bálint, Darko Durisic, Sushant Kumar Pandey |
IV | 4 |
| 2025 | D-LeDe: A Data Leakage Detection Method for Automotive Perception SystemsabstractData leakage is a very common problem that is often overlooked during splitting data into train and test sets before training any ML/DL model. The model performance gets artificially inflated with the presence of data leakage during the evaluation phase which often leads the model to erroneous prediction on real-time deployment. However, detecting the presence of such leakage is challenging, particularly in the object detection context of perception systems where the model needs to be supplied with image data for training. In this study, we conduct a computational experiment to develop a method for detecting data leakage. We then conducted an initial evaluation of the method as a first step on a public dataset, “Kitti”, which is a popular and widely accepted benchmark dataset in the automotive domain. The evaluation results show that our proposed D-LeDe method are able to successfully detect potential data leakage caused by image similarity. A further validation was also provided to justify the evaluation outcome by conducting pair-wise image similarity analysis using perceptual hash (pHash) distance. Md. Abu Ahammed Babu, Sushant Kumar Pandey, Darko Durisic, Ashok Chaitanya Koppisetty, Miroslaw Staron |
VEHITS | 3 |
| 2025 | Design pattern recognition: a study of large language modelsabstractAbstract Context As Software Engineering (SE) practices evolve due to extensive increases in software size and complexity, the importance of tools to analyze and understand source code grows significantly. Objective This study aims to evaluate the abilities of Large Language Models (LLMs) in identifying DPs in source code, which can facilitate the development of better Design Pattern Recognition (DPR) tools. We compare the effectiveness of different LLMs in capturing semantic information relevant to the DPR task. Methods We studied Gang of Four (GoF) DPs from the P-MARt repository of curated Java projects. State-of-the-art language models, including Code2Vec, CodeBERT, CodeGPT, CodeT5, and RoBERTa, are used to generate embeddings from source code. These embeddings are then used for DPR via a k-nearest neighbors prediction. Precision, recall, and F1-score metrics are computed to evaluate performance. Results RoBERTa is the top performer, followed by CodeGPT and CodeBERT, which showed mean F1 Scores of 0.91, 0.79, and 0.77, respectively. The results show that LLMs without explicit pre-training can effectively store semantics and syntactic information, which can be used in building better DPR tools. Conclusion The performance of LLMs in DPR is comparable to existing state-of-the-art methods but with less effort in identifying pattern-specific rules and pre-training. Factors influencing prediction performance in Java files/programs are analyzed. These findings can advance software engineering practices and show the importance and abilities of LLMs for effective DPR in source code. Sushant Kumar Pandey, Sivajeet Chand, Jennifer Horkoff, Miroslaw Staron, Miroslaw Ochodek, Darko Durisic |
Empir. Softw. Eng. | 6 |
| 2023 | TransDPR: Design Pattern Recognition Using Programming Language ModelsabstractCurrent Design Pattern Recognition (DPR) methods have limitations, such as the reliance on semantic information, limited recognition of novel or modified pattern versions, and other factors. We present an introductory DPR technique by using a Programming Language Model (PLM) called TransDPR, which utilizes a Facebook pre-trained model (TransCoder), which is a Cross-lingual programming Language Model (XLM) based on a transformer architecture. We leverage an n-dimensional vector representation of programs and apply logistic regression to learn design patterns (DPs). Our approach utilizes the GitHub repository to collect singleton and prototype DP programs written in$C$++ source code. Our results indicate that TransDPR achieves 90% accuracy and an F1-score of 0.88 on open-source projects. We evaluate the proposed model on two developed modules from Volvo Cars and invite the original developers to validate the prediction results. Sushant Kumar Pandey, Miroslaw Staron, Jennifer Horkoff, Miroslaw Ochodek, Nicholas Mucci, Darko Durisic |
ESEM | 6 |
| 2022 | Comparing Input Prioritization Techniques for Testing Deep Learning AlgorithmsabstractDeep learning (DL) systems are becoming an essential part of software systems, so it is necessary to test them thoroughly. This is a challenging task since the test sets can grow over time as the new data is being acquired, and it becomes time-consuming. Input prioritization is necessary to reduce the testing time since prioritized test inputs are more likely to reveal the erroneous behavior of a DL system earlier during test execution. Input prioritization approaches have been rudimentary analyzed against each other, this study compares different input prioritization techniques regarding their effectiveness and efficiency. This work considers surprise adequacy, autoencoder-based, and similarity-based input prioritization approaches in the example of testing a DL image classification algorithms applied on MNIST, Fashion-MNIST, CIFAR-10, and STL-10 datasets. To measure effectiveness and efficiency, we use a modified APFD (Average Percentage of Fault Detected), and set up & execution time, respectively. We observe that the surprise adequacy is the most effective (0.785 to 0.914 APFD). The autoencoder-based and similarity-based techniques are less effective, with the performance from 0.532 to 0.744 APFD and 0.579 to 0.709 APFD, respectively. In contrast, the similarity-based and surprise adequacy-based approaches are the most and least efficient, respectively. The findings in this work demonstrate the trade-off between the considered input prioritization techniques to understanding their practical applicability for testing DL algorithms. Vasilii Mosin, Miroslaw Staron, Darko Durisic, Francisco Gomes de Oliveira Neto, Sushant Kumar Pandey, Ashok Chaitanya Koppisetty |
SEAA | 3 |
| 2019 | Assessing the impact of meta-model evolution: a measure and its automotive application
Darko Durisic, Miroslaw Staron, Matthias Tichy, Jörgen Hansson |
Softw. Syst. Model. | 1 |
| 2017 | Co-Evolution of Meta-Modeling Syntax and Informal Semantics in Domain-Specific Modeling Environments - A Case Study of AUTOSARabstractOne domain-specific modeling environment is centered around a domain-specific meta-model which defines syntax (modeling elements, e.g., classes) for the domain models. However, in order for the system designers to be able to construct meaningful models, semantics of the domain-specific meta-model needs to be described as well. This semantics is often provided in a form of informal natural language specifications that contain a set of design requirements, each describing the intended use of one or more modeling elements. Intuitively, introduction of new concepts into the modeling environment is expected to require changes in both meta-modeling syntax and informal semantics in such a way that their co-evolution is highly correlated. In order to test this hypothesis, we analyzed the relation between added classes, attributes, and connectors, as meta-modeling syntax, and modified/added design requirements, as meta-modeling semantics, in a case study of the AUTOSAR meta-modeling environment. We found that new AUTOSAR concepts usually require both new modeling elements and new design requirements, but surprisingly adding more elements is not always followed by more requirements. This finding is also validated by the moderately strong correlation between the evolution of these two AUTOSAR meta-modeling artifacts (Spearman's rho 0,63 and Kendall's tau 0,49). For system designers, this means that both meta-modeling syntax and informal semantics is important to be considered in the analysis of domain-specific meta-model evolution, but it may not be enough for understanding the use of all modeling elements. For designers responsible for the maintenance of domain-specific meta-models, this means that more effort shall be put into describing the semantics of all introduced modeling elements. Darko Durisic, Corrado Motta, Miroslaw Staron, Matthias Tichy |
MoDELS | 1 |
| 2017 | Measuring the Evolution of Meta-models - A Case Study of Modelica and UML Meta-modelsabstractThe evolution of both general purpose and domain-specific meta-models and its impact on the existing models and modeling tools has been discussed extensively in the modeling research community. To assess the impact of domain-specific meta-model evolution on the modeling tools, a number of measures have been proposed by Durisic et al., NoC (Number of Changes) being the most prominent one. The proposed measures are evaluated on a case of AUTOSAR meta-model that specifies the language for designing automotive system architectures. In this paper, we assess the applicability of these measure and the underlying data-model for their calculation in a case study of Modelica and UML meta-models. Our preliminary results show that the proposed data-model and the measures can be applied to both analyzed meta-models as we were able to capture 68/77 changes on average per Modelica/UML release. However, only a subset of the data-model elements is applicable for analyzing the evolution of Modelica and also certain transformation of the data-model is required in case of UML. Despite these encouraging results, further studies are needed to assess the usefulness of the actual measures, e.g., NoC, in assessing the impact of Modelica/UML meta-model evolution on the modeling tools. Maxime Jimenez, Darko Durisic, Miroslaw Staron |
MODELSWARD | 2 |
| 2016 | Addressing the Need for Strict Meta-modeling in Practice - A Case Study of AUTOSARabstractMeta-modeling has been a topic of interest in the modeling community for many years, yielding substantialnumber of papers describing its theoretical concepts. Many of them are aiming to solve the problem of traditionalUML based domain-specific meta-modeling related to its non-compliance to the strict meta-modelingprinciple, such as the deep meta-modeling approach. In this paper, we show the practical use of meta-modelsin the automotive development process based on AUTOSAR and visualize places in the AUTOSAR metamodelwhich are broken according to the strict meta-modeling principle. We then explain how the AUTOSARmeta-modeling environment can be re-worked in order to comply to this principle by applying three individualapproaches, each one combined with the concept of Orthogonal Classification Architecture: UML extension,prototypical pattern and deep instantiation. Finally we discuss the applicability of these approaches in practiceand contrast the identified issues with the actual problems faced by the automotive meta-modeling practitioners.Our objective is to bridge the current gap between the theoretical and practical concerns in meta-modeling. Darko Durisic, Miroslaw Staron, Matthias Tichy, Jörgen Hansson |
MODELSWARD | 1 |
| 2016 | Should We Adopt a New Version of a Standard? - A Method and Its Evaluation on AUTOSAR
Corrado Motta, Darko Durisic, Miroslaw Staron |
PROFES | 2 |
| 2015 | ARCA - Automated Analysis of AUTOSAR Meta-model ChangesabstractThe software architecture of automotive software systems on the European market and wider is designed following the AUTOSAR standard. This requires continuous adoption of new AUTOSAR releases in the development projects in order to enable new innovative solutions in cars. Under these circumstances, the analysis of impact of the AUTOSAR meta-model changes on the modeling tools used in the development is crucial for avoiding delays and increased cost. However due to tens of new features combined with thousands of meta-model changes between consecutive releases of AUTOSAR, tool support is needed for such analysis. In this paper we present a systematic method and a tool - ARCA - for automated analysis of the AUTOSAR meta-model changes. The tool is able to identify relevant changes affecting modeling tools used by different roles in the development process and present the optimal set of new features to be adopted in the projects. The goal of the tool is to enable faster and cheaper software innovation cycles in cars. Darko Durisic, Miroslaw Staron, Matthias Tichy |
MiSE@ICSE | 1 |
| 2014 | Quantifying Long-Term Evolution of Industrial Meta-Models - A Case StudyabstractMeasurement in software engineering is an important activity for successful planning and management of projects under development. However knowing what to measure and how is crucial for the correct interpretation of the measurement results. In this paper, we assess the applicability of a number of software metrics for measuring a set of meta-model properties - size, length, complexity, coupling and cohesion. The goal is to identify which of these properties are mostly affected by the evolution of industrial meta-models and also which metrics should be used for their successful monitoring. In order to assess the applicability of the chosen set of metrics, we calculate them on a set of releases of the standardized meta-model used in the development of automotive software systems - the AUTOSAR meta-model - in a case study at Volvo Car Corporation. To identify the most applicable metrics, we used Principal Component Analysis (PCA). The results of these metrics shall be used by software designers in planning software development projects based on multiple AUTOSAR meta-model versions. We concluded that the evolution of the AUTOSAR meta-model is quite even with respect to all 5 properties and that the metrics based on fan-in complexity and package cohesion quantify the evolution most accurately. Darko Durisic, Miroslaw Staron, Matthias Tichy, Jörgen Hansson |
IWSM/Mensura | 1 |
| 2013 | Measuring the impact of changes to the complexity and coupling properties of automotive software systems
Darko Durisic, Martin Nilsson 0002, Miroslaw Staron, Jörgen Hansson |
J. Syst. Softw. | 1 |