Nemania Borovits

dblp:275/3307 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0002-5661-4795ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Addressing Data Scarcity with Synthetic Data: A Secure and GDPR-Compliant Cloud-Based Platform
abstract
This study presents a cloud-based platform for synthetic data generation, validation, and evaluation, developed to address data scarcity in the telecommunications sector while ensuring compliance with the General Data Protection Regulation (GDPR). In collaboration with a Dutch telecommunications provider facing data scarcity due to low user-consent rates, we developed a platform that allows synthetic data vendors to securely generate synthetic data based on schema input without accessing sensitive information. Vendors uploaded containerized executables for synthetic data generation and the platform automated infrastructure provisioning, ensuring no access to personal data. A validation mechanism minimized the risk of re-identification by ensuring that the synthetic data did not inadvertently replicate real data points. We mutually agreed with the vendors on five evaluation metrics and the platform logged and calculated performance for each, allowing them to refine their algorithms. To validate the platform’s performance, we conducted an offline study with the TV viewership team, using each vendor’s synthetic data to generate viewership categories. The vendor with the best evaluation metrics also produced categories most similar to the real data, confirming the platform’s effectiveness. This study, involving two vendors and a telecommunications company, demonstrated the platform’s applicability in addressing business challenges while ensuring privacy compliance.
Nemania Borovits, Gianluigi Bardelloni, Hossein Hashemi 0005, Masoom Tulsiani, Damian A. Tamburri, Willem-Jan van den Heuvel
ACM Trans. Softw. Eng. Methodol.1
2023 Anonymization-as-a-Service: The Service Center Transcripts Industrial Case
Nemania Borovits, Gianluigi Bardelloni, Damian A. Tamburri, Willem-Jan van den Heuvel
ICSOC (2)1
2022 FindICI: Using machine learning to detect linguistic inconsistencies between code and natural language descriptions in infrastructure-as-code
abstract
Linguistic anti-patterns are recurring poor practices concerning inconsistencies in the naming, documentation, and implementation of an entity. They impede the readability, understandability, and maintainability of source code. This paper attempts to detect linguistic anti-patterns in Infrastructure-as-Code (IaC) scripts used to provision and manage computing environments. In particular, we consider inconsistencies between the logic/body of IaC code units and their short text names. To this end, we propose FindICI a novel automated approach that employs word embedding and classification algorithms. We build and use the abstract syntax tree of IaC code units to create code embeddings used by machine learning techniques to detect inconsistent IaC code units. We evaluated our approach with two experiments on Ansible tasks systematically extracted from open source repositories for various word embedding models and classification algorithms. Classical machine learning models and novel deep learning models with different word embedding methods showed comparable and satisfactory results in detecting inconsistent Ansible tasks related to the top-10 used Ansible modules.
Nemania Borovits, Indika Kumara, Dario Di Nucci, Parvathy Krishnan, Stefano Dalla Palma, Fabio Palomba, Damian A. Tamburri, Willem-Jan van den Heuvel
Empir. Softw. Eng.1