Alba González-Cebrián

dblp:335/9987 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-7519-4917ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SMARDY: the CORE of zero-trust FAIR marketplace for research data
abstract
Supporting discovery through good and open management of existing datasets is the core of progress in data-rich research environments. Open Data and FAIR (Findable, Accessible, Interoperable, and Reusable) principles drive the exploitation of current results to a better and more trustworthy scientific era, but the wide adoption in Science is hindered by concerns regarding the proper handling of sensitive or extra-valuable copyrighted datasets. We present Smardy, our proposal for extending Open Data Repositories for research datasets with components meant to protect data sovereignty and trust in transfer. Our implementation of a cross-platform is the core of a FAIR Dataset Marketplace, which allows the authors to trade their datasets with the unconditional security of a Zero-Trust environment, and helps them to protect their IP over data using undisputable, Blockchain-based proofs of their authorship. The depicted aspects include application-code design, functional schemes, fingerprinting, and encryption steps for properly handling datasets and generating authorship proofs.
Cosmin-Andrei Ionite, Ion-Dorinel Filip, Alba González-Cebrián, Ciprian Dobre
Connect. Sci.3
2024 Data Drift for Automatic FAIR-compliant Dataset Versioning in Large Repositories
abstract
Construed as a shift in the distribution or structure of data over time, data drift can adversely affect the performance of machine learning models and data-driven decisions. This study examines two data drift metrics, denoted as dE,PCAand dE,AE, that are derived from unsupervised ML models: the reconstruction error-based metrics of Principal Component Analysis (PCA) and Autoencoders (AE). To investigate the robustness of these metrics, we have systematically accessed time-series datasets from the European Data Portal. Our experiments have examined data versioning through three basic events: creation, update, and deletion. The results are summarised and aggregated for all datasets, and unsupervised analysis based on Robust PCA and AE has been performed to examine patterns within the impact of dataset characteristics on data drift detection and computational efficiency. Our results indicate that both metrics aligned closely in performance with new records, suggesting consistent drift detection under normal conditions with FAIR compliance. However, high-dimensional datasets posed challenges for both PCA and AE models. Update events revealed discrepancies between the two metrics, suggesting that non-linear shifts affected AE-based metrics more than PCA-based ones. Deletion events demonstrated the resilience of these metrics against data loss, but also revealed variability in the reliability of the PCA model; i.e., data drift metrics derived from PCA and AE can be effective but sensitive to certain dataset characteristics.
Alba González-Cebrián, Iulian Ciolacu, Michael Bradford, Ciprian Dobre, Horacio González-Vélez
e-Science1
2022 Smardy: Zero-Trust FAIR Marketplace for Research Data
abstract
Over the past five years, different organisations have increasingly called for science to become more open and reproducible. They have endorsed a set of data-management principles known as the FAIR (Findable, Accessible, Interoperable, Reusable) principles. As such, there is a growing trend towards the open availability of research data, as researchers continue to enhance reproducibility by enabling sharing and opening of their findings and datasets. However, there is not yet a standardised way to openly enable access to datasets while keeping control of their final use, potentially obtaining benefits from their utilisation. This paper introduces Smardy, an EU-funded project which is deploying a traceable FAIR-compliant open innovation marketplace for data. Its innovative method for data exchange consists of the use of blockchain for controlling access rights to data, with data models able to grant access according to policies completely kept under the control of the data owner/producer. We also describe how Smardy employs dimensionality reduction techniques to automatically generate FAIR–compliant metadata, statistical fingerprinting to identify derivated datasets, and watermarking to help data owners trace the distribution of multiple copies of a dataset.
Ion-Dorinel Filip, Cosmin Ionite, Alba González-Cebrián, Mihaela Balanescu, Ciprian Dobre, Adriana E. Chis, Dave Feenan, Adrian-Alexandru Buga, Ioan-Mihai Constantin, George Suciu, George V. Iordache, Horacio González-Vélez
IEEE Big Data3
2022 Automatic Versioning of Time Series Datasets: a FAIR Algorithmic Approach
abstract
As one of the fundamental concepts underpinning the FAIR (Findability, Accessibility, Interoperability, and Reusability) guiding principles, data provenance entails keeping track of each version for a given dataset from its original to its latest version. However, standard terms to determine and include versioning information in the metadata of a given dataset are still ambiguous and do not explicitly define how to assess the overlap of information between items along a versioning stream. In this work, we propose a novel approach for automatic versioning of time series datasets, based on the use of parameters from two dimensionality reduction approaches, namely Principal Component Analysis and Autoencoders. That is to say, we systematically detect and measure similarities (information distances) in datasets via dimensionality reduction, encode them as different versions, and then automatically generate provenance metadata via a FAIR versioning service using the W3C DCAT 3.0 nomenclature. We illustrate this approach with two time series datasets and demonstrate how the proposed parameters effectively assess the similarity between different data versions. Our results have shown that the proposed version similarity metrics are robust$(s^{(0,1)}=1)$to the alteration of up to 60% of cells, the removal of up to 60% of rows, and the log-scale transformation of variables. In contrast, row-wise transformations (e.g. converting absolute values to a percentage of a second variable) yield minimal similarity values$(s^{(0,1)} < 0.75)$. Our code and datasets are openly available to enable reproducibility.
Alba González-Cebrián, Luke A. McGuinness, Michael Bradford, Adriana E. Chis, Horacio González-Vélez
e-Science1