VLDB 2026 Research / reviewers in the wild / expert
Maria Ulan
dblp:221/1630
· DBLP profile ↗
5ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-3906-7611ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAFEXPLAIN: a Complete Approach Towards Trustworthy AI-Based Safety-Critical SystemsabstractAI becomes increasingly important in safetycritical systems, especially in the case of autonomous systems, since navigation relies on AI for object detection and collision avoidance. However, safety-critical systems must adhere to functional safety standards that enforce software to be correct-by-construction, component decomposition to simplify design and validation, and the use of data only for testing purposes not to design the system itself. AI in general, and Deep Learning (DL) in particular have opposed characteristics since they have error rates (e.g., due to mispredictions), AI/DL modules can only be designed and validated monolithically, and they build on data for their design (i.e. for training purposes). Hence, DL solutions are at odds with the development process of safetycritical systems. A number of standards have recently emerged in different domains to reconcile the requirements of safety-critical systems with the characteristics of DL solutions, such as ISO 21448, ISO/IEC TR 5469, and ISO 8800, among others. However, there is a lack of realistic practice to design a DL-based safety-critical system in accordance with those regulations, and existing solutions only cover some aspects in isolation, and are often incompatible among them. SAFEXPLAIN is a 3-year Horizon Europe project addressing this challenge. SAFEXPLAIN, which finishes in September 2025, has already reached its main goals providing specific and complementary solutions to all those challenges so that AIbased safety-critical systems can be designed, implemented and validated adhering to the relevant functional safety standards in domains such as automotive, space and railway. In particular, SAFEXPLAIN provides the concepts, processes, tools and frameworks addressing the challenge end-to-end, from concept to solution. This is proven by the successful application of the SAFEXPLAIN approach in three case studies from the automotive, space and railway domains, whose results will see the light very soon. Jaume Abella 0001, Irune Agirre, Thanh Hai Bui, Frank Geujen, Gabriele Giordana, Carlo Donzella, Francisco J. Cazorla, Enrico Mezzetti, Axel Brando, Javier Fernández 0004, Irune Yarza, Joanes Plazaola, Maria Ulan, Rob Lavreysen, Lucas Tosi, Ilaria Bloise, Lorenzo Feruglio, Ilaria Cinelli, Stefano Lodico, William Guarienti, Giuseppe Nicosia, Valeria Dallara |
DSD | 14 |
| 2025 | Talk is Cheap, Energy is Not: Towards a Green, Context-Aware Metrics Framework for Automatic Speech RecognitionabstractAutomatic Speech Recognition (ASR) systems are increasingly deployed across diverse computing environments, from cloud servers to edge devices. While accuracy has traditionally been the primary evaluation metric, the inference efficiency of these systems, including energy consumption, memory usage, and hardware utilisation, significantly impacts their practical usability. This paper introduces a novel benchmarking framework that assesses ASR models during inference from both performance and sustainability perspectives. We introduce a multi-metric evaluation approach quantifying Word Error Rate (WER), Real-Time Factor (RTF), Energy Per Audio Second (EPAS), inference latency, GPU Memory Efficiency (GME), and Hardware Utilisation Rate (HUR). Our framework includes configurable weighting schemes tailored for various deployment scenarios: balanced general-purpose evaluation, resource-constrained environments, high-throughput batch inference, and real-time processing. To demonstrate the utility of the framework, we benchmark several state-of-the-art ASR architectures (Whisper, Wav2Vec2, HuBERT, WavLM, UniSpeech, and SpeechT5) in both FP16 and FP32 precision on NVIDIA Jetson AGX Orin hardware. The proposed methodology supports researchers and practitioners in making informed model selection decisions based on context-specific inference requirements. By illuminating performance–consumption trade-offs, the metrics framework can help to reduce computational costs and the carbon footprint of ASR systems, while maintaining acceptable accuracy. Maria Ulan, Erik Johannes Husom, Jeriek Van den Abeele |
ECML/PKDD (9) | 1 |
| 2021 | Weighted software metrics aggregation and its application to defect predictionabstractAbstract It is a well-known practice in software engineering to aggregate software metrics to assess software artifacts for various purposes, such as their maintainability or their proneness to contain bugs. For different purposes, different metrics might be relevant. However, weighting these software metrics according to their contribution to the respective purpose is a challenging task. Manual approaches based on experts do not scale with the number of metrics. Also, experts get confused if the metrics are not independent, which is rarely the case. Automated approaches based on supervised learning require reliable and generalizable training data, a ground truth, which is rarely available. We propose an automated approach to weighted metrics aggregation that is based on unsupervised learning. It sets metrics scores and their weights based on probability theory and aggregates them. To evaluate the effectiveness, we conducted two empirical studies on defect prediction, one on ca. 200 000 code changes, and another ca. 5 000 software classes. The results show that our approach can be used as an agnostic unsupervised predictor in the absence of a ground truth. Maria Ulan, Welf Löwe, Morgan Ericsson, Anna Wingkvist |
Empir. Softw. Eng. | 1 |
| 2021 | Copula-based software metrics aggregationabstractAbstract A quality model is a conceptual decomposition of an abstract notion of quality into relevant, possibly conflicting characteristics and further into measurable metrics. For quality assessment and decision making, metrics values are aggregated to characteristics and ultimately to quality scores. Aggregation has often been problematic as quality models do not provide the semantics of aggregation. This makes it hard to formally reason about metrics, characteristics, and quality. We argue that aggregation needs to be interpretable and mathematically well defined in order to assess, to compare, and to improve quality. To address this challenge, we propose a probabilistic approach to aggregation and define quality scores based on joint distributions of absolute metrics values. To evaluate the proposed approach and its implementation under realistic conditions, we conduct empirical studies on bug prediction of ca. 5000 software classes, maintainability of ca. 15000 open-source software systems, and on the information quality of ca. 100000 real-world technical documents. We found that our approach is feasible, accurate, and scalable in performance. Maria Ulan, Welf Löwe, Morgan Ericsson, Anna Wingkvist |
Softw. Qual. J. | 1 |
| 2018 | Quality Models Inside Out: Interactive Visualization of Software Metrics by Means of Joint ProbabilitiesabstractAssessing software quality, in general, is hard; each metric has a different interpretation, scale, range of values, or measurement method. Combining these metrics automatically is especially difficult, because they measure different aspects of software quality, and creating a single global final quality score limits the evaluation of the specific quality aspects and trade-offs that exist when looking at different metrics. We present a way to visualize multiple aspects of software quality. In general, software quality can be decomposed hierarchically into characteristics, which can be assessed by various direct and indirect metrics. These characteristics are then combined and aggregated to assess the quality of the software system as a whole. We introduce an approach for quality assessment based on joint distributions of metrics values. Visualizations of these distributions allow users to explore and compare the quality metrics of software systems and their artifacts, and to detect patterns, correlations, and anomalies. Furthermore, it is possible to identify common properties and flaws, as our visualization approach provides rich interactions for visual queries to the quality models' multivariate data. We evaluate our approach in two use cases based on: 30 real-world technical documentation projects with 20,000 XML documents, and an open source project written in Java with 1000 classes. Our results show that the proposed approach allows an analyst to detect possible causes of bad or good quality. Maria Ulan, Sebastian Hönel, Rafael Messias Martins, Morgan Ericsson, Welf Löwe, Anna Wingkvist, Andreas Kerren |
VISSOFT | 1 |