Andrea Visentin

dblp:185/1434 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0003-3702-4826ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 A Tool for Fairness Assessment and Red-Lining Detection in AI Systems
abstract
Fairness in AI is essential to data analysis, ensuring ethical decision-making. Various fairness metrics are used depending on the specific use case. In this work, we propose a fairness assessment tool that provides a comprehensive analysis, incorporating state-of-the-art fairness metrics and association analysis to detect and include the red-lining effect. Features that exhibit strong associations with sensitive attributes can contribute to indirect discrimination. For this reason, we identify hidden biases through association analysis, detecting proxy variables that may perpetuate discrimination and lead to a more discriminatory machine learning model. We introduce a generalized fairness analysis framework capable of addressing complex scenarios, supported by a query mechanism designed to capture extensive contextual information for a more representative evaluation. Notably, fairness concerns in the business sector often involve complex, context-dependent scenarios. Our tool enables users to formulate such problems effectively and retrieve important insights. In this work, we present the flow of the tool along with some indicative results on complex fairness problems.
Dimitrios Bikoulis, Panagiotis E. Kyziropoulos, Eduardo Vyhmeister, Gabriel G. Castañé, Andrea Visentin
SMARTCOMP5
2025 Predicting Daily Depression Scores Using Passive Sensing: A Behaviour-Aware Approach to Missing Value Handling
abstract
Depression is a serious mental health issue among college students, yet early detection and continuous monitoring remain challenging due to reliance on infrequent and subjective self-reports. Recent developments in mobile and wearable sensor technologies provide a promising alternative by using passive, real-time monitoring of behavioural patterns. However, such data are often missing, and traditional imputation methods typically assume data are missing (completely) at random. In this study, we introduced a sensor-aware preprocessing pipeline that adjusts specific missing values when the absence of data may reflect meaningful user behaviour rather than sensor failure. The pipeline is applied to longitudinal mobile sensing data and weekly Patient Health Questionnaire-4 (PHQ-4) depression scores from college students over three academic years. By interpolating PHQ-4 scores to generate daily labels, we train and evaluate regression models to predict depression scores. Our sensor-aware adjustment approach improved model performance compared to standard preprocessing methods. These findings show the value of including contextual knowledge in missing data handling and support the feasibility of passive, continuous mental health monitoring in student populations.
Sarah Corbett, Nicola Rossberg, Malik Muhammad Qirtas, Andrea Visentin
SMARTCOMP4
2025 DEMO: EaseTalk: An LLM-Driven Speech Practice Tool for Real-Life Scenarios
abstract
Stuttering can make everyday conversations challenging, especially in situations that involve unfamiliar people or social pressure. While traditional therapy provides structured support, many individuals struggle to apply those techniques in real-world scenarios. Digital speech tools exist, but most focus on repetitive drills and rarely provide realistic, contextdriven speaking practice. We present EaseTalk, an AI-driven mobile application that allows users to practice real-life conversational scenarios in a supportive, self-paced environment. The app uses speech recognition and prompt-engineered large language models (LLMs) to detect speech disfluencies such as repetitions, prolongations and blocks. EaseTalk aims to empower users through independent, scenario-based speech practice with potential applications extending beyond stuttering into areas like social anxiety, interview preparation and language learning.
Marco Faggiani, Malik Muhammad Qirtas, Pauline Frizelle, Fiona Ryan, Nicole Muller, Andrea Visentin
SMARTCOMP6
2025 Estimating Perceived Fatigue Using Machine Learning and Biomechanical Features from Wearable Sensors
abstract
Physical fatigue is a state of reduced physical ability caused by prolonged activity or repetitive tasks. It can affect performance in tasks that require effort, focus or precision and can lead to reduced strength and increased risk of injuries. Detecting physical fatigue is important for timely interventions and enhancing safety, efficiency and overall well-being in workplaces, sports and rehabilitation. In this study, we propose a robust and generalizable framework for fatigue detection using wearable sensor data, specifically using Inertial Measurement Unit and Electromyography sensors. A comprehensive set of biomechanical features was extracted from raw sensor data to capture both kinematic and neuromuscular aspects of fatigue progression. These features were evaluated across shoulder internal rotation and external rotation movements under different resistance levels. We trained and compared multiple regression models for fatigue estimation using subjective fatigue ratings based on the Borg Rating of Perceived Exertion scale and performed feature importance analysis to get model interpretability. The extracted feature set showed strong generalizability specifically for IR movements, as proved by leave one task out cross-validation, where models maintained robust performance across unseen movement-resistance task settings. This work highlights the potential of combining IMU and EMG data, along with biomechanical features extracted from these two sensor modalities for accurate and interpretable fatigue estimation. It opens the way for real-world applications in dynamic and diverse environments for effective fatigue estimation.
Malik Muhammad Qirtas, Merve Nur Yasar, Marco Sica, Salvatore Tedesco, Andrea Visentin
SMARTCOMP5
2025 Unsupervised Induction Motor Anomaly Detection Using a Deep Convolutional Autoencoder Based on Multi-Sensor Data Fusion
abstract
Induction motors are the primary way to convert electrical current into mechanical power. They are a fundamental component of industrial processes and equipment. Early fault detection and preventive maintenance are of great concern. In the last few years, many deep learning data-driven approaches have been used to detect faults in electric motors. This problem comes with two significant challenges: some faults are easier to detect using a specific sensor (e.g., vibration or current); in industrial applications, it is hard to obtain fault measurements. In most cases, only measurements of normal behaviour are available. This paper presents a multi-signal unsupervised anomaly detection system based on deep convolutional variational autoencoders (VAE). We use three sensors to sample from operating industrial motors: vibration, current, and magnetic flux. We divide the dataset into a training set, in which the network fits the nominal working condition of the motor. The system is then deployed in detection mode, analyzing the stream of data provided by the sensors. The experimental results show that the system accurately detects anomalies and has sufficient sensitivity to recognize changes in the motor load and behavior in practice.
Andrea Visentin, Marco Dalla, Benjamin Provan-Bessell, Barry O'Sullivan
SMARTCOMP1
2025 Data Monetization Through Tailored Demand Representation
abstract
In today's digital economy, data represents a critical strategic resource, necessitating innovative approaches to its monetization and value realization. This study evaluates various methodologies for data monetization through customized demand representations. By examining diverse consumer demand models, we capture the intricate behaviors of digital market consumers. Our contribution includes improving existing modeling techniques by incorporating nuanced dependencies of data value, such as price sensitivity, quality perception, and trustworthiness of services. Moreover, we address non-discrete service consumption scenarios relevant to contemporary offerings like Information-as-a-Service (IaaS) and Answers-as-a-Service (AaaS). This work contributes to the linkage between theoretical models and practical strategies specific to data markets, expanding current understanding and providing actionable insights for effective data monetization strategies. Additionally, the study highlights potential areas for future research, particularly regarding the integration of emerging technological advancements and evolving regulatory landscapes. These insights can guide firms in adapting their data monetization frameworks to maintain a competitive advantage in rapidly changing markets.
Eduardo Vyhmeister, Lorenzo Reyes-Bozo, Dimitrios Bikoulis, Panagiotis E. Kyziropoulos, Gabriel G. Castañé, Andrea Visentin
SMARTCOMP6
2024 Explainable Algorithm Selection for the Capacitated Lot Sizing Problem
Andrea Visentin, Aodh Ó Gallchóir, Jens Kärcher, Herbert Meyr
CPAIOR (2)1
2024 SAT Instances Generation Using Graph Variational Autoencoders
abstract
This paper presents a SAT instance generator using a Graph Variational Autoencoder (GVAE2SAT ) architecture that outperforms existing generative deep learning models in speed and requires minimal post-processing.Our computational analyses benchmark this model against current deep learning techniques, introducing advanced metrics for more accurate evaluation.This new model is unique in its ability to maintain partial satisfiability of SAT instances while significantly reducing computational time.Although no method perfectly addresses all challenges in generating SAT instances, our approach marks a significant step forward in the efficiency and effectiveness of SAT instance generation.
Daniel Crowley, Marco Dalla, Barry O'Sullivan, Andrea Visentin
ESANN4
2024 A Machine Learning Approach to Model Counting
abstract
Model counting (#SAT) is the problem of computing the number of satisfying assignments for a given Boolean formula. It has a significant theoretical and practical interest. Tackling it can be challenging since the number of potential solution grows exponentially with the number of variables. Due to the inherent complexity of the problem, approaches to approximate model counting have been developed as a practical alternative. These methods extract the number of solutions within user-specified tolerance and confidence levels and in a fraction of the time required by exact model counters. However, even these methods require extensive computations, restricting their applicability to relatively small instances. In this paper, we propose a new approximate machine learning model counter that overcome this limitation. Predicting the number of solutions can be seen as a regression problem. We deploy an array of machine learning techniques trained to infer the approximate number of solutions based on statistical features extracted from a SAT propositional formula. Extensive numerical experiments performed on synthetic crafted and benchmark datasets show that learning approaches can provide a good approximation of the number of solutions with a much lower computational time and resource cost than the state-of-the-art approximate and exact model counters, making it possible to approximate the model count of instances previously out of reach. We then investigated the structural factors that lead to a high model count using AI explainability approaches.
Marco Dalla, Andrea Visentin, Barry O'Sullivan
ICTAI2
2023 Performance and Energy Savings Trade-Off with Uncertainty-Aware Cloud Workload Forecasting
abstract
Cloud computing has seen widespread adoption because it increases the productivity and efficiency of industries and allows for effective scalability of their business [1]. Guaranteeing performance levels is at the core of cloud services and requires huge computational resources, especially with the latest advances in technologies such as Artificial Intelligence and the Internet of Things [2]. Typically, customers subscribe to agreements where cloud providers ensure specific levels of reliability, availability and responsiveness to systems and applications and describe penalties if the service levels are not met. At the same time, massive computational resources are a cost for providers and have a significant environmental impact, which will increase in the future. It is estimated that the energy consumption of data centres (which host cloud services) will grow from 292 TWh in 2016 to 353 TWh in 2030 [3], and greenhouse gas emissions will increase over 14% in 2040, compared to a 1-1.6% increase in the 2007–2016 [4].
Diego Carraro, Andrea Rossi 0010, Andrea Visentin, Steven D. Prestwich, Kenneth N. Brown
ICNP3
2023 Using Machine Learning Classifiers in SAT Branching [Extended Abstract]
abstract
The Boolean Satisfiability Problem (SAT) can be framed as a binary classification task. Recently, numerous machine and deep learning techniques have been successfully deployed to predict whether a CNF has a solution. However, these approaches do not provide a variables assignment when the instance is satisfiable and have not been used as part of SAT solvers. In this work, we investigate the possibility of using a machine-learning SAT/UNSAT classifier to assign a truth value to a variable. A heuristic solver can be created by iteratively assigning one variable to the value that leads to higher predicted satisfiability. We test our approach with and without probing features and compare it to a heuristic assignment based on the variable's purity. We consider as objective the maximisation of the number of literals fixed before making the CNF unsatisfiable. The preliminary results show that this iterative procedure can consistently fix variables without compromising the formula's satisfiability, finding a complete assignment in almost all test instances.
Ruth Helen Bergin, Marco Dalla, Andrea Visentin, Barry O'Sullivan, Gregory M. Provan
SOCS3
2023 SAT Feature Analysis for Machine Learning Classification Tasks
abstract
The extraction of meaningful features from CNF instances is crucial to applying machine learning to SAT solving, enabling algorithm selection and configuration for solver portfolios and satisfiability classification. While many approaches have been proposed for feature extraction, their relevance to these tasks is unclear. Their applicability and comparison of the information extracted and the computational effort needed are complicated by the lack of working or updated implementations, negatively affecting reproducibility. In this paper, we analyse the performance of five sets of features presented in the literature on SAT/UNSAT and problem category classification over a dataset of 3000 instances across ten problem classes distributed equally between SAT and UNSAT. To increase reproducibility and encourage research in this area, we released a Python library containing an updated and clear implementation of structural, graph-based, statistical and probing features presented in the literature for SAT CNF instances; and we define a clear pipeline to compare feature sets in a given learning task robustly. We analysed which of the computed features are relevant for the specific task and the tradeoff they provide between accuracy and computational effort. The results of the analysis provide insights into which features mostly affect an instance's satisfiability and which can be used to identify the problem's type. These insights can be used to develop more effective solver portfolios and satisfiability classification algorithms.
Marco Dalla, Benjamin Provan-Bessell, Andrea Visentin, Barry O'Sullivan
SOCS3
2022 Bayesian Uncertainty Modelling for Cloud Workload Prediction
abstract
Providers of cloud computing systems need to manage resources carefully to meet the desired Quality of Service and reduce waste due to overallocation. An accurate prediction of future demand is crucial to allocate resources to service requests without excessive delays. Current state-of-the-art methods such as Long Short-Term Memory-based models make only point forecasts of demand without considering the uncertainty in their predictions. Forecasting a distribution would provide a more comprehensive picture and inform resource scheduler decisions. We investigate Bayesian Neural Networks and deep learning models to predict workload distribution and evaluate them on the time series forecasting of CPU and memory workload of 8 clusters on the Google Cloud data centre. Experiments show that the proposed models provide accurate demand prediction and better estimations of resource usage bounds, reducing overprediction and total predicted resources, while avoiding underprediction. These approaches have good runtime performance making them applicable for practitioners.
Andrea Rossi 0010, Andrea Visentin, Steven D. Prestwich, Kenneth N. Brown
CLOUD2
2021 Automated SAT Problem Feature Extraction using Convolutional Autoencoders
abstract
The Boolean Satisfiability Problem (SAT) was the first known NP-complete problem and has a very broad literature focusing on it. It has been applied successfully to various real-world problems, such as scheduling, planning and cryptography. SAT problem feature extraction plays an essential role in this field. SAT solvers are complex, fine-tuned systems that exploit problem structure. The ability to represent/encode a large SAT problem using a compact set of features has broad practical use in instance classification, algorithm portfolios, and solver configuration. The performance of these techniques relies on the ability of feature extraction to convey helpful information. Researchers often craft these features "by hand" to capture particular structures of the problem. Instead, in this paper, we extract features using semi-supervised deep learning. We train a convolutional autoencoder (AE) to compress the SAT problem into a limited latent space and reconstruct it minimizing the reconstruction error. The latent space projection should preserve much of the structural features of the problem. We compare our approach to a set of features commonly used for algorithm selection. Firstly, we train classifiers on the projection to predict if the problems are satisfiable or not. If the compression conveys valuable information, a classifier should be able to take correct decisions. In the second experiment, we check if the classifiers can identify the original problem that was encoded as SAT. The empirical analysis shows that the autoencoder is able to represent problem features in a limited latent space efficiently, as well as convey more information than current feature extraction methods.
Marco Dalla, Andrea Visentin, Barry O'Sullivan
ICTAI2
2019 Predicting Judicial Decisions: A Statistically Rigorous Approach and a New Ensemble Classifier
abstract
Natural language processing and machine learning are gaining wide popularity in supporting judicial decision-making. Research in this area is particularly active. However, a methodological issue in the use of AI methods can lead to poor statistical soundness in the results. We consider and improve the work of Aletras et. al. [1] for predicting the outcome of cases at the European Court of Human Rights. We replicate their experiments using a more statistically reliable methodology and analyzed the results using state-of-the-art Bayesian techniques for classifier comparison. We also improved classification accuracy using an ensemble-based approach. These techniques will widely improve the statistical soundness of machine learning applications in law by providing robust baselines for comparison.
Andrea Visentin, Alessia Nardotto, Barry O'Sullivan
ICTAI1
2016 Robust Principal Component Analysis by Reverse Iterative Linear Programming
Andrea Visentin, Steven D. Prestwich, Armagan Tarim
ECML/PKDD (2)1