Sanja Lazarova-Molnar

dblp:42/3496 · DBLP profile ↗
← Back
26ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-6052-0863ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Automated Formalization Of Expert Knowledge For Data-Knowledge-Fused Digital Twin Model Extraction Using Pre-Trained Language Models
abstract
Integrating expert knowledge into Digital Twin model extraction remains an open challenge. Although expert knowledge can enhance model robustness and contextual validity, it is typically expressed in unstructured natural language, leading to ambiguity, heterogeneity, and limited operationalizability. Consequently, systematic and automated formalization mechanisms are required to enable effective fusion of expert knowledge with data for extracting data-knowledge-fused Digital Twin models. To address this gap, we propose an automated approach for formalizing expert knowledge expressed in natural language. We evaluate two formalization approaches using twelve publicly available pre-trained language models under a set of explicitly defined constraints. Based on the empirical evaluation, encoder-decoder language models emerge as the most suitable candidates, as they demonstrate the strongest overall performance within the considered evaluation setting. We further illustrate the role of the proposed approach within data-knowledge-fused Digital Twin model extraction through a reliability case study. The proposed approach reduces effort required for integrating expert knowledge and supports automated model extraction.
Michelle Jungmann, Sanja Lazarova-Molnar
ECMS2
2026 RoadAceVision: Deep Multimodal Learning for Road Surface Classification
Lukas Michael Schwemer, Mustafa Demetgül, Sanja Lazarova-Molnar
IV3
2026 MDPNML: A Multidimensional Petri Net Markup Language Enabling Construction and Simulation of Comprehensive Digital Twin Models
abstract
A Digital Twin (DT) is a data-driven virtual representation of a physical system that updates in near-real-time and incorporates models of system physics and behaviors. Multidimensional Stochastic Petri Nets (MDSPNs), as an extension of traditional Stochastic Petri Nets (SPNs), provide an intuitive formalism for modeling and analyzing complex systems across multiple dimensions, enabling the development of comprehensive DTs. In MDSPNs, system objectives can be associated with different relevant dimensions, including time, energy, and waste. The Petri Net Markup Language (PNML), a standard XML-based interchange format, is widely used for sharing and executing Petri net models across tools. PNML, however, lacks multidimensional semantics. A PNML-compatible format for MDSPNs would enable portable model exchange across tools and allow conversion to time-oriented SPNs for generic PNML tools, supporting end-to-end traditional and comprehensive DT workflows. Such an extended PNML format also enhances model reproducibility and reduces the effort required to integrate dimension-specific behavior. In this paper, we introduce the Multidimensional Petri Net Markup Language (MDPNML) and identify necessary adaptations to the PNML format to represent MDSPNs. MDPNML supports multidimensional attributes, different dimensions in transitions, and simulation parameters for MDSPNs. Through an illustrative case study, we demonstrate MDPNML generation and execution for multidimensional simulation.
Atieh Khodadadi, Sanja Lazarova-Molnar
MODELSWARD2
2026 Frequency-Based Partitioning for Modular Validation of Stochastic Petri Net Digital Twin Models
abstract
As technological advancements enable the manufacturing industry’s adoption of cyber-physical production systems, the critical need for Digital Twins has also become more apparent. They allow for optimization, flexibility, and continuous improvement by replicating physical systems in a digital environment. However, they can only be utilized if their underlying models remain a valid representation of the corresponding real-world systems. As manufacturing systems and their Digital Twins grow in complexity, model validation and decomposition become increasingly challenging, particularly due to heterogeneous system dynamics. In complex manufacturing systems, different events occur at different rates, and a uniform validation approach of the corresponding Digital Twin models may fail to account for these dynamic behaviors. Depending on the rate of occurrence of different events, validation mechanisms must decide whether to preserve, recalibrate, or re-extract different parts of a model. To address this challenge, a modular validation framework was introduced, partitioning Digital Twin models into sub-models and dynamically adjusting validation policies. The effectiveness of such a framework, however, critically depends on how the model is partitioned and whether the partitions capture heterogeneous system behavior while preserving dependencies. In this paper, we present a data-driven, behavior-based approach for partitioning stochastic Petri net Digital Twin models that automatically identifies meaningful partition boundaries by analyzing reachable marking patterns and transition firing frequencies. The approach preserves inter-partition dependencies while enabling independent validation of sub-models. We demonstrate the method through a reliability-focused manufacturing case study, showing that frequency-based partitioning can capture heterogeneous dynamics while preserving system behavior.
Ashkan Zare, Sanja Lazarova-Molnar
SIGSIM-PADS2
2026 Navigating the Use of Generative AI in Modeling and Simulation: A Case-Study-Driven Methodology for Agentic Workflows in Synthetic Data Generation for Surrogate Modeling
Dusan Sturek, Sanja Lazarova-Molnar
SIMULTECH2
2025 Multidimensional Stochastic Petri Nets: A Novel Approach to Modeling and Simulation of Stochastic Discrete-Event Systems
abstract
Process Mining (PM) has been proven valuable for extracting process flows from data, also in the form of stochastic Petri net (SPN) models of systems. SPNs are widely recognized for their ability to model complex, stochastic systems and are extensively used in combination with PM. While SPNs provide an intuitive and straightforward way to model complex systems, representing changes across multiple dimensions, such as energy and waste, remains challenging in their standard frameworks. In this paper, we introduce an extension of stochastic Petri nets, termed Multidimensional SPNs (MDSPNs), by extending the SPN framework to capture dynamics along different dimensions. MDSPNs facilitate a comprehensive modeling of systems’ behaviors from multiple perspectives, which can correspond to the diverse objectives of systems. To facilitate design and simulation of MDSPNs, we designed and developed MDPySPN, a Python library, which we also introduce in this paper. MDPySPN enables the simulation of MDSPNs by supporting alterations of multiple values at system events. With MDPySPN, we aim to provide researchers, engineers, and simulation professionals with a practical and extensible toolkit to model, simulate, and analyze MDSPNs, thereby supporting multi-objective optimization of stochastic processes in systems. Through a case study, we demonstrate the capabilities of modeling and simulation of MDSPNs using MDPySPN.
Atieh Khodadadi, Sanja Lazarova-Molnar
SMC2
2025 Data Requirements for Tire-Road Monitoring: A Roadmap for Data Collection, Processing, and Decision Making
abstract
Today, noise pollution and road anomalies are one of the most significant problems for city authorities. Moreover, one of the biggest sources of this noise is car noise, especially tire-road noise. With the widespread use of electric vehicles, engine noise has been reduced to a minimum. Many engineering fields are carrying out studies to monitor this noise and road condition in tires, roads and car mechanics. The objectives are to improve passenger comfort, to ensure maneuverability and speed control of unmanned vehicles in accordance with road conditions, to reduce accidents and road-related vehicle damage, and to reduce noise pollution. The most important part of these studies is tire-road noise and road condition monitoring with AI. In this article, we examine the existing studies on tire-road, the open-source data problem in this field, the solution proposal, the detection of road-related problems, and the studies on noise estimation in this field. It also aims to provide a road map on existing databases, data processing, and smart decision-making in this field and to support future studies in this field.
Mustafa Demetgül, Jiawen Meng, Sanja Lazarova-Molnar, Alexey V. Vinel
VTC2025-Fall3
2025 Acoustic Traffic Attribute Classification for ITS: A Comparative Study of Machine Learning and CNN Approaches
abstract
Reliable information about traffic attributes, including vehicle type, speed range, and direction, is essential for traffic management and intelligent transportation systems (ITS). Although radar- and camera-based solutions can provide accurate data, they are often expensive, vulnerable to environmental conditions, and may raise privacy concerns.In this study, we investigate passive acoustic sensing as a cost-effective and privacy-preserving alternative. Using stereo vehicle pass-by noise encoded as Mel-Frequency Cepstral Coefficients (MFCCs), we simultaneously classify vehicle type, speed range (inferred from road-specific speed limits), and movement direction. Three modeling strategies are evaluated: (1) traditional machine learning on time-averaged MFCCs, (2) a hybrid model combining ResNet50-based feature extraction with LightGBM classification, and (3) end-to-end convolutional neural networks (CNNs), including a lightweight multi-task variant enhanced with attention mechanisms.Experiments are conducted on the public IDMT-Traffic dataset, comprising 17,506 stereo audio clips. Our multi-task CNN achieves the best overall performance with only 646K parameters, reaching approximately 99% accuracy across all three tasks. These results highlight the viability of stereo acoustic sensing as a lightweight, scalable, and privacy-preserving solution for real-time traffic perception and V2X infrastructure support.
Jiawen Meng, Mustafa Demetgül, Sanja Lazarova-Molnar, Frank Gauterin, Alexey V. Vinel
VTC2025-Fall3
2025 Towards integrating process mining with agent-based modeling and simulation: State of the art and outlook
abstract
Agent-based modeling and simulation (ABMS) is a valuable tool for assessing complex socio-technical systems and is becoming increasingly advanced through the integration with data-driven capabilities. Process mining is an emerging data-driven discipline that combines elements from data mining and process modeling to gain insights into process execution through tasks such as process discovery, conformance checking, and process enhancement using event data. This study explores the role of process mining and its impact on the ABMS paradigm, identifying the current state of the art, gaps in the literature, and future directions for integrating process mining with ABMS. A systematic literature review is conducted to examine how ABMS and process mining techniques are jointly employed to address challenges reported in the literature. From an initial pool of 189 publications, a final set of 20 papers was synthesized, their primary contributions were discussed, and open issues and challenges for future research were identified. Although the integrated field of process mining and ABMS shows an upward trend in publications, it remains modest and requires further efforts to achieve synergistic improvements in socio-technical systems. The findings offer initial guidance for promising research directions.
Rob H. Bemthuis, Sanja Lazarova-Molnar
Expert Syst. Appl.2
2024 Data-Driven Agent-Based Modeling and Simulation of Price Competition in the Danish Pharmaceutical Market
Ruhollah Jamali, Sanja Lazarova-Molnar
EUMAS2
2024 A Vision for Advancing Digital Twins Intelligence: Key Insights and Lessons from Decades of Research and Experience with Simulation
abstract
Digital Twins have revolutionized the domain of Modeling and Simulation by making use of the growing and cost-efficient possibilities to extract data from systems, as well as the increasing computational power. At the same time, Digital Twins have enabled tremendous advances in diverse cyber-physical systems by enabling better monitoring, predictive maintenance, design optimization, and informed decision-making. As their popularity evolved, the understanding of what a Digital Twins has become more and more dispersed and unclear. Here, we offer understanding of what a Digital Twin is based on our experience in research within its native domain of Modeling and Simulation, with a concrete focus on the key considerations that need to be made when developing Digital Twins or working with them. We, furthermore, emphasize the need to include all available knowledge for better-informed Digital Twins. To illustrate our ideas and vision, we use case studies from our research.
Sanja Lazarova-Molnar
SIMULTECH1
2023 Using Process Mining for Face Validity Assessment in Agent-Based Simulation Models: An Exploratory Case Study
Rob H. Bemthuis, Ruben Govers, Sanja Lazarova-Molnar
CoopIS3
2023 An Approach for Face Validity Assessment of Agent-Based Simulation Models Through Outlier Detection with Process Mining
Rob H. Bemthuis, Sanja Lazarova-Molnar
EDOC2
2023 Predictive Process Monitoring for Prediction of Remaining Cycle Time in Automated Manufacturing: A Case Study
abstract
Predicting remaining cycle times of products in manufacturing systems is critical to ensure on-time deliveries to customers, schedule resources and actions for expected order completions, and address excessive production stops proactively rather than retroactively. Recent advances in Predictive Process Monitoring (PPM), a sub-discipline of Process Mining, enable the use of machine learning to predict remaining cycle times based on event data. We apply PPM to the automated manufacturing domain and demonstrate the approach using a case study from a water meter manufacturer. For prediction of remaining cycle times, PPM relies on regression methods, such as Decision Trees, Random Forests, and Gradient Boosting Machines based on event data. We compare the prediction accuracy of these methods and show that PPM can deliver relevant insights for production lines without imposing extensive data requirements.
Jonas Friederich, Jonas Kristoffer Lindeløv, Sanja Lazarova-Molnar
ETFA3
2023 Towards Developing an Agent-Based Model of Price Competition in the European Pharmaceutical Parallel Trade Market
Ruhollah Jamali, Sanja Lazarova-Molnar
EUMAS2
2023 Equipment-centric Data-driven Reliability Assessment of Complex Manufacturing Systems
abstract
Complex manufacturing systems produce highly engineered products with long product cycle times and are characterized by complex production process behaviors. Ensuring the reliability of these systems is critical to meet customer demands, improve product quality and minimize production losses. The collection and storage of data by sensors and information systems respectively enable the automatic generation and analysis of reliability models of complex manufacturing systems, reducing the need for expert knowledge of the processes. In this article, we propose a novel approach to generate data-driven reliability models of complex manufacturing systems using stochastic Petri nets as the modeling formalism. Our method extracts models from event logs that capture relevant events related to material flow in a system, and state logs, that capture operational state changes in a system’s production resources using process mining. We, furthermore, simulate the derived data-driven reliability models using discrete-event simulation and validate the models to ensure their robustness. We demonstrate the successful application of our method using a case study from the wafer fabrication domain. The results of our case study indicate that data-driven reliability assessment of complex manufacturing systems is feasible and can provide rapid insights into such systems. In addition, the extracted models can be used to support decisions related to maintenance planning, parts procurement and system configuration.
Jonas Friederich, Wentong Cai 0001, Boon-Ping Gan, Sanja Lazarova-Molnar
SIGSIM-PADS4
2023 Data-driven extraction and analysis of repairable fault trees from time series data
abstract
Fault tree analysis is a probability-based technique for estimating the risk of an undesired top event, typically a system failure. Traditionally, building a fault tree requires involvement of knowledgeable experts from different fields, relevant for the system under study. Nowadays’ systems, however, integrate numerous Internet of Things (IoT) devices and are able to generate large amounts of data that can be utilized to extract fault trees that reflect the true fault-related behavior of the corresponding systems. This is especially relevant as systems typically change their behaviors during their lifetimes, rendering initial fault trees obsolete. For this reason, we are interested in extracting fault trees from data that is generated from systems during their lifetimes. We present DDFTAnb algorithm for learning fault trees of systems using time series data from observed faults, enhanced with Naïve Bayes classifiers for estimating the future fault-related behavior of the system for unobserved combinations of basic events, where the state of the top event is unknown. Our proposed algorithm extracts repairable fault trees from multinomial time series data, classifies the top event for the unseen combinations of basic events, and then uses proxel-based simulation to estimate the system’s reliability. We, furthermore, assess the sensitivity of our algorithm to different percentages of data availabilities. Results indicate DDFTAnb’s high performance for low levels of data availability, however, when there are sufficient or high amounts of data, there is no need for classifying the top event.
Parisa Niloofar, Sanja Lazarova-Molnar
Expert Syst. Appl.2
2022 Teaching Modeling, Simulation, and Performance Evaluation Course Online with Jupyter Notebook: Course Development and Lessons Learned
abstract
This innovative practice full paper presents a case of teaching modeling, simulation, and performance evaluation course online during the COVID-19 pandemic for graduate students. The course includes theoretical and practical sessions with varying complexity. Students should have a good background in math, statistics, and probability. Students must also have good experience in one of the computer programming languages to solve the homework and work on the term project. The challenge is how to teach these topics online and engage the students in the course as they learn in face-to-face classes. Delivering the course using PowerPoint slides and a whiteboard is not suitable for teaching the class online. Therefore, there must be an alternative way to deliver the course online. We noticed a growing interest in using Jupyter Notebook in teaching, which motivated us to apply it to the mentioned course with some innovations. Jupyter Notebook is an open-source web application that allows us to create and share documents that contain live code, equations, visualizations, and narrative text. We want to share our experience teaching this course online using Jupyter Notebook in this innovative practice. We will share the course development plan, delivery mode, lessons learned, and student feedback. I will also highlight maximizing the benefits of the Jupyter Notebook using add-ins and tools useful for teaching, such as converting the Jupyter Notebook to a slide show. Developing courses in Jupyter Notebook could be time-consuming and frustrating, especially if there are a lot of math equations, tables, and drawings. This effort pays off in terms of the quality of the instruction and learning, and it gives students a tool to help them practice and engage with the course material.
Farag M. Sallabi, Sanja Lazarova-Molnar
FIE2
2019 The Synergy of Simulation and Time Series Forecasting for Live Performance Testing of Smart Buildings
abstract
Differences in requirements for reliability in buildings imply the different needs for calculation of expected building behaviour. In this paper we examine four techniques for calculating expected behaviour of buildings. Two of them are simulation techniques, namely, a white box EnergyPlus model and a æ static tool as per the requirements of the Danish government. The other two are machine learning techniques, namely an ARIMA model, and an long short-term memory artificial recurrent neural network, used in deep learning. We compare and contrast these four techniques based on their accuracy of forecast, as well as execution time to forecast a new data point. Furthermore, we provide an algorithm for selection of forecasting technique based on terms such as availability, accuracy, and execution time requirements, to facilitate real time threshold generation in light of building performance testing.
Elena Markoska, Sanja Lazarova-Molnar
iiWAS2
2016 Reliability of cyber physical systems with focus on building management systems
abstract
Cyber-physical systems are slowly emerging to dominate our world. Cyber-physical systems (CPS) are systems that tightly integrates users, devices and software. Whereas many of these systems are obviously safety-critical systems, some of them become so under special circumstances. This is the case with our focus CPS, i.e. building management systems (BMS), which are not always safety critical per se, but under special circumstances they can become such. This certainly depends on the purpose of the building. We can easily imagine BMS of hospital buildings as safety-critical, but also BMS of buildings that store sensitive materials and equipment that could be of biological nature or encompassing sensitive technology that would need special temperature, humidity and light settings. For this reason, in this paper we would like to emphasize on the importance of reliability of CPS in general, with a special focus on BMS, as our area of interest. Furthermore, we also propose a classification of buildings with respect to the necessity of having their reliability evaluated.
Sanja Lazarova-Molnar, Hamid Reza Shaker, Nader Mohamed
IPCCC1
2016 Middleware to support cyber-physical systems
abstract
Middleware can provide novel and practical approaches for enhancing Cyber-Physical Systems (CPS) application development processes and operations. This paper discusses the roles and advantages of using middleware to build CPS. In addition, the paper studies the required features needed in such middleware for development and operation of CPS.
Nader Mohamed, Jameela Al-Jaroodi, Sanja Lazarova-Molnar, Imad Jawhar
IPCCC3
2016 Software Engineering Issues for Cyber-Physical Systems
abstract
Cyber-Physical Systems (CPS) provide many smart features for enhancing physical processes. These systems are designed with a set of distributed hardware, software, and network components that are embedded in physical systems and environments or attached to humans. Together they function seamlessly to offer specific functionalities or features that help enhance human lives, operations or environments. While different CPS components play important roles in a successful CPS development, the software plays the most important role among them. Acquiring and using high quality CPS components is the first step; however, designing and implementing the right software to integrate and use them effectively is essential. The software facilitates better interfaces, more control and adds smart services, high flexibility and many other added values and features to the CPS. However, software development for CPS is not a trivial task. This paper provides an overview discussion of software engineering issues related to the analysis, design, development, verification and validation, and quality assurance of CPS software. Some of these issues are related to the nature/type of CPS while others are related to the complexity of the software development processes used to develop such systems.
Jameela Al-Jaroodi, Nader Mohamed, Imad Jawhar, Sanja Lazarova-Molnar
SMARTCOMP4
2013 Performance Modeling of Data Dissemination in Vehicular Ad Hoc Networks
abstract
Vehicular Ad hoc Networks (VANETs) have become a cornerstone component of Intelligent Transportation Systems (ITS). VANET applications present a huge potential for improving road safety and travel comfort, hence the growing interest of both academia and industry. The main advantage of VANETs is its ad hoc nature which does not require fixed infrastructure or centralized administration. However, designing scalable information dissemination techniques for VANET applications remains a challenging task due to the inherent nature of such highly dynamic environments. Existing dissemination techniques often resort to simulation for performance evaluation and there are only few studies that offer mathematical modeling. In this paper we provide a comparative study of existing performance modeling approaches for data dissemination techniques designed for different VANET applications.
Moumena Chaqfeh, Abderrahmane Lakas, Sanja Lazarova-Molnar
DS-RT3
2011 A genetic algorithm to enhance transmembrane helices prediction
abstract
A transmembrane helix (TMH) topology prediction is becoming a central problem in bioinformatics because the structure of TM proteins is difficult to determine by experimental means. Therefore, methods which could predict the TMHs topologies computationally are highly desired. In this paper we introduce TMHindex, a method for detecting TMH segments solely by the amino acid sequence information. Each amino acid in a protein sequence is represented by a Compositional Index deduced from a combination of the difference in amino acid appearances in TMH and non-TMH segments in training protein sequences and the amino acid composition information. Furthermore, genetic algorithm was employed to find the optimal threshold value to separate TMH segments from non-TMH segments. The method successfully predicted 376 out of the 378 TMH segments in 70 testing protein sequences. The level of accuracy achieved using TMHindex in comparison to recent methods for predicting the topology of TM proteins is a strong argument in favor of our method.
Nazar Zaki, Salah Bouktif, Sanja Lazarova-Molnar
GECCO3
2010 Modeling Human Decision Behaviors for Accurate Prediction of Project Schedule Duration
Sanja Lazarova-Molnar, Rabeb Mizouni
EOMAS1
2009 Protein-protein interaction based on pairwise similarity
abstract
BACKGROUND: Protein-protein interaction (PPI) is essential to most biological processes. Abnormal interactions may have implications in a number of neurological syndromes. Given that the association and dissociation of protein molecules is crucial, computational tools capable of effectively identifying PPI are desirable. In this paper, we propose a simple yet effective method to detect PPI based on pairwise similarity and using only the primary structure of the protein. The PPI based on Pairwise Similarity (PPI-PS) method consists of a representation of each protein sequence by a vector of pairwise similarities against large subsequences of amino acids created by a shifting window which passes over concatenated protein training sequences. Each coordinate of this vector is typically the E-value of the Smith-Waterman score. These vectors are then used to compute the kernel matrix which will be exploited in conjunction with support vector machines. RESULTS: To assess the ability of the proposed method to recognize the difference between "interacted" and "non-interacted" proteins pairs, we applied it on different datasets from the available yeast saccharomyces cerevisiae protein interaction. The proposed method achieved reasonable improvement over the existing state-of-the-art methods for PPI prediction. CONCLUSION: Pairwise similarity score provides a relevant measure of similarity between protein sequences. This similarity incorporates biological knowledge about proteins and it is extremely powerful when combined with support vector machine to predict PPI.
Nazar Zaki, Sanja Lazarova-Molnar, Wassim El-Hajj, Piers Campbell
BMC Bioinform.2