VLDB 2026 Research / reviewers in the wild / expert
Matias Duran
dblp:313/5990
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-5109-4102ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Applying Metamorphic Testing for Pose Estimation in the Context of Rugby Analysis: Lessons Learned and FindingsabstractAnalysis of rugby match and training footage is particularly useful for coaches and players to understand and improve their tackling technique, and potentially lower the rate of injuries. Machine learning models (in particular for pose estimation) promise to streamline rugby analysis. However models trained for “general purpose” computer vision tasks, such as pose estimation and object detection, frequently fail as a result of the challenging conditions and significant domain shift that rugby footage presents: high-impact, close-contact play causes problems such as occlusions, motion blur, and unconventional body orientations. It is therefore crucial to understand the specific conditions which cause these systems to fail so they can be prioritised during pre-processing and expensive manual data collection. In this paper we leverage Met-Pose, a metamorphic testing system to understand the specific conditions that cause pose estimation systems to fail. Metamorphic testing is particularly advantageous as this approach side-steps the need for costly, manually labelled data. Our ongoing project on applying pose estimation for rugby analysis employs MediaPipe, a popular, widely used pose estimation system, on rugby broadcast footage. We show how applying metamorphic testing to a sport analytics application can reveal situations that challenge the model without the need for any manual data labelling. For example, our results show that in this context, MediaPipe is particularly sensitive to motion blur and colour loss, but less so to lighting and resolution changes. Furthermore, we show how this process can be adapted to focus on particular aspects of an application by proposing a new metamorphic rule exploring the effect of including or excluding context on MediaPipe’s results. Our results show where MediaPipe struggles in complex, real-world sporting scenarios and also offer concrete insights for improving data augmentation, data collection and system design in sports analytics. Matias Duran, Will Connors, Thomas Laurent 0003, Ellen Rushe, Anthony Ventresque |
ECAI | 1 |
| 2025 | Metamorphic Testing for Pose Estimation SystemsabstractPose estimation systems are used in a variety of fields, from sports analytics to livestock care. Given their potential impact, it is paramount to systematically test their behaviour and potential for failure. This is a complex task due to the oracle problem and the high cost of manual labelling necessary to build ground truth keypoints. This problem is exacerbated by the fact that different applications require systems to focus on different subjects (e.g., human versus animal) or landmarks (e.g., only extremities versus whole body and face), which makes labelled test data rarely reusable. To combat these problems we propose MET-POSE, a metamorphic testing framework for pose estimation systems that bypasses the need for manual annotation while assessing the performance of these systems under different circumstances. MET-POSE thus allows users of pose estimation systems to assess the systems in conditions that more closely relate to their application without having to label an ad-hoc test dataset or rely only on available datasets, which may not be adapted to their application domain. While we define Met-pose in general terms, we also present a non-exhaustive list of metamorphic rules that represent common challenges in computer vision applications, as well as a specific way to evaluate these rules. We then experimentally show the effectiveness of Met-pose by applying it to Mediapipe Holistic, a state of the art human pose estimation system, with the FLIC and PHOENIX datasets. With these experiments, we outline numerous ways in which the outputs of Met-pose can uncover faults in pose estimation systems at a similar or higher rate than classic testing using hand labelled data, and show that users can tailor the rule set they use to the faults and level of accuracy relevant to their application. Matias Duran, Thomas Laurent 0003, Ellen Rushe, Anthony Ventresque |
ICST | 1 |
| 2023 | Adaptive Search-based Repair of Deep Neural NetworksabstractDeep Neural Networks (DNNs) are finding a place at the heart of more and more critical systems, and it is necessary to ensure they perform in as correct a way as possible. Search-based repair methods, that search for new values for target neuron weights in the network to better process fault-inducing inputs, have shown promising results. These methods rely on fault localisation to determine what weights the search should target. However, as the search progresses and the network evolves, the weights responsible for the faults in the system will change, and the search will lose in effectiveness. In this work, we propose an adaptive search method for DNN repair that adaptively updates the target weights during the search by performing fault localisation on the current state of the model. We propose and implement two methods to decide when to update the target weights, based on the progress of the search's fitness value or on the evolution of fault localisation results. We apply our technique to two image classification DNN architectures against a dataset of autonomous driving images, and compare it with a state-of-the art search-based DNN repair approach. Davide Li Calsi, Matias Duran, Thomas Laurent 0003, Xiao-Yi Zhang 0005, Paolo Arcaini, Fuyuki Ishikawa |
GECCO | 2 |
| 2023 | Distributed Repair of Deep Neural NetworksabstractDeep Neural Networks (DNNs) are applied in several safety-critical domains and their trustworthiness is of paramount importance. For example, DNNs used in autonomous driving as classifiers should not misclassify detected objects; however, since obtaining perfect accuracy is not possible, special attention should be given to the most critical cases, e.g., pedestrians. This has been confirmed by the consortium of our partners from the automotive domain that provided us with specific risk levels for different misclassifications. A recent approach to improve DNN performance is to localise DNN weights responsible for the misclassifications and then adjust (repair) them to improve the misclassifications. However, they under-perform when they need to consider multiple misclassifications, and they do not consider the risk levels of the different misclassifications. To tackle this, we propose DISTRREP, a distributed repair approach that first finds the best fixes for each critical misclassification, and then integrates them in a single repaired DNN model, by considering the risk levels. We assess DISTRREP over three DNN models and a dataset of autonomous driving images, by considering requirements specified by our industrial partners. Experiments show that DISTRREP is more effective than baseline approaches based on retraining, and other risk-unaware repair approaches. Davide Li Calsi, Matias Duran, Xiao-Yi Zhang 0005, Paolo Arcaini, Fuyuki Ishikawa |
ICST | 2 |
| 2021 | What to Blame? On the Granularity of Fault Localization for Deep Neural NetworksabstractValidating Deep Neural Networks (DNNs) used for classification is of paramount importance; an approach for this consists in (i) executing the DNN over the test dataset, (ii) collecting information about classifications, and (iii) applying fault localization (FL) techniques to identify the neurons responsible for the misclassifications. DNNs can have multiple misclassification types, and so neurons responsible for one type could be different from those responsible for another type. However, depending on the granularity of the analyzed dataset, FL may not reveal these differences: failure types more frequent in the dataset may mask less frequent ones. We here propose a way to perform FL for DNNs that avoids this masking effect by selecting test data in a granular way. We conduct an empirical study, using a spectrum-based FL approach for DNNs, to assess how FL results change by changing the granularity of the analyzed test data. Namely, we perform FL by using test data with two different granularities: following a state-of-the-art approach that considers all misclassifications for a given class together, and the proposed fine-grained approach. Results show that FL should be done for each misclassification, such that practitioners have a more detailed analysis of the DNN faults and can make a more informed decision on what to fix in the DNN. Matias Duran, Xiao-Yi Zhang 0005, Paolo Arcaini, Fuyuki Ishikawa |
ISSRE | 1 |