Flavio Bertini 0001

dblp:58/3571-1 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-6925-5712ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Large Language Models Evaluation for PubMed Extractive Summarisation
abstract
The increasingly large amount of available biomedical literature is making it difficult to gather and synthesise all the necessary information. Moreover, this domain-specific task demands a high level of reliability in the generated text and concepts. Pre-trained large language models have recently shown promising results. Given the specific requirements of biomedical text summarisation, our evaluation focuses on extractive models, prioritising the accuracy of the generated text. In this article, we evaluate the capabilities of 18 general-domain and biomedical pre-trained language models in various configurations on the biomedical extractive summarisation task using one single and two multi-document datasets consisting of 33,000, 5,000 and 470,000 PubMed articles, respectively. We performed the comparison using several well-known metrics, namely ROUGE-1, ROUGE-2, ROUGE-L, BERTScore, BLEU and METEOR. The main contribution of this work lies in providing a detailed performance analysis, highlighting the differences between general-domain and biomedical models, and identifying key factors that influence model performance in extractive summarisation tasks within the biomedical domain. Experimental results show that biomedical models tend to result in higher recall, while general-domain models produce higher precision, while general-domain models produce higher precision. This corresponds to more expressive summaries for biomedical models and shorter summaries for general-domain models.
Tian Cheng Xia, Flavio Bertini 0001, Danilo Montesi
ACM Trans. Comput. Heal.2
2023 Graph-based Tool for Exploring PubMed Knowledge Base
abstract
Studies have shown that data retrieval and visualization tools can help health professionals to improve their understanding and communication with patients, their relationship with stakeholders, and their decision-making process. However, not many efforts have been made in this direction. In this paper, we present a prototype system for the indexing, annotation, and visualization of the PubMed knowledge base to enable the search and retrieval of health-related evidence. The proposed tool builds and keeps updated an enriched graph based on PubMed articles associating them with concepts extracted from the Unified Medical Language System (UMLS) Metathesaurus. Moreover, it allows a full-text search and graph-based navigation and supports an overview of concepts and related publications. The proposed architecture enables scale-up thanks to its containerized nature and parallelization capabilities. The code is open-source under the Apache V2 license.
Simone Bottoni, Alberto Trombetta, Flavio Bertini 0001, Danilo Montesi, Francesca Bonin, Alessandra Pascale, Martin Gleize, Pierpaolo Tommasi
ICDE3
2023 BIOCHAIN: towards a platform for securely sharing microbiological data
abstract
There is a need to persuade public and private entities to share their currently unexposed bio-data banks by preserving ownership and secrecy. The reason is to make available results that can be obtained by massively exploiting the content of such data by modern machine learning approaches. Digital catalogues of data collections are being provided. However, they are not developed to protect private content that may be shared according to privileges assigned by the owners. Here, we present BIOCHAIN, a data-sharing module which will be the basis for a computational platform aimed at performing federated data analysis. The platform is intended to be used by a consortium of private and public institutions in the field of microbiology. BIOCHAIN makes use of blockchain technology to guarantee fairness among entities of the consortium by allowing them to securely share their data.
Vincenzo Bonnici, Vincenzo Arceri, Alessio Diana, Flavio Bertini 0001, Eleonora Iotti, Alessia Levante, Valentina Bernini, Erasmo Neviani, Alessandro Dal Palù
IDEAS4
2023 A Web-Based Application for Screening Alzheimer's Disease in the Preclinical Phase
abstract
As a result of an increasing elderly population, the number of people with age-related diseases is increasing worldwide. Alzheimer's disease is thus becoming an emergency health and social problem. Neuropsychological evaluation and biomarker identification represent the two main approaches to identifying subjects with Alzheimer's. In this paper, we propose a web application designed to be sensitive to the cognitive changes distinctive of the early Mild Cognitive Impairment, which is a condition in which someone experiences minor cognitive problems, and the preclinical phase of Alzheimer's disease. The application is conceived to be self-administered in a comfortable and non-stressful environment. It was designed to be quick to administer, automatic to score, and able to preserve privacy because of the highly sensitive data collected. The preliminary evaluation of the application was done by enrolling 518 subjects characterised by several risk factors and the presence of a family history, which underwent standard neuropsychological screening.
Flavio Bertini 0001, Daniela Beltrami, Pegah Barakati, Laura Calzà, Enrico Ghidoni, Danilo Montesi
ISCC1
2022 Grayscale Text Watermarking
abstract
In an increasingly connected world, where information is easily spread through multiple channels and platforms, digital watermarking has been broadly investigated in authorship attribution and intellectual property protection of digital content. However, text contents pose many challenges due to a low capacity to embed a watermark. In this paper, we propose a new structural watermarking method for small pieces of text that may allow to hide upwards of 15 bits of watermark per character manipulating the underlying font grayscale values. The proposed method ensures length preservation and is robust to the copy and paste activities. Moreover, the method is able to embed a password-based watermark returning visually indistinguishable watermarked text.
Simone Branchetti, Flavio Bertini 0001, Danilo Montesi
IDEAS2
2022 An automatic Alzheimer's disease classifier based on spontaneous spoken English
Flavio Bertini 0001, Davide Allevi, Gianluca Lutero, Laura Calzà, Danilo Montesi
Comput. Speech Lang.1
2022 Automatic Speech Classifier for Mild Cognitive Impairment and Early Dementia
abstract
The World Health Organization estimates that 50 million people are currently living with dementia worldwide and this figure will almost triple by 2050. Current pharmacological treatments are only symptomatic, and drugs or other therapies are ineffective in slowing down or curing the neurodegenerative process at the basis of dementia. Therefore, early detection of cognitive decline is of the utmost importance to respond significantly and deliver preventive interventions. Recently, the researchers showed that speech alterations might be one of the earliest signs of cognitive defect, observable well in advance before other cognitive deficits become manifest. In this article, we propose a full automated method able to classify the audio file of the subjects according to the progress level of the pathology. In particular, we trained a specific type of artificial neural network, called autoencoder, using the visual representation of the audio signal of the subjects, that is, the spectrogram. Moreover, we used a data augmentation approach to overcome the problem of the large amount of annotated data usually required during the training phase, which represents one of the most major obstacles in deep learning. We evaluated the proposed method using a dataset of 288 audio files from 96 subjects: 48 healthy controls and 48 cognitively impaired participants. The proposed method obtained good classification results compared to the state-of-the-art neuropsychological screening tests and, with an accuracy of 90.57%, outperformed the methods based on manual transcription and annotation of speech.
Flavio Bertini 0001, Davide Allevi, Gianluca Lutero, Danilo Montesi, Laura Calzà
ACM Trans. Comput. Heal.1
2020 Hierarchical embedding for DAG reachability queries
abstract
Current hierarchical embeddings are inaccurate in both reconstructing the original taxonomy and answering reachability queries over Direct Acyclic Graph. In this paper, we propose a new hierarchical embedding, the Euclidean Embedding (EE), that is correct by design due to its mathematical formulation and associated lemmas. Such embedding can be constructed during the visit of a taxonomy, thus making it faster to generate if compared to other learning-based embeddings. After proposing a novel set of metrics for determining the embedding accuracy with respect to the reachability queries, we compare our proposed embedding with state-of-the-art approaches using full trees from 3 to 1555 nodes and over a real-world Direct Acyclic Graph of 1170 nodes. The benchmark shows that EE outperforms our competitors in both accuracy and efficiency.
Giacomo Bergami, Flavio Bertini 0001, Danilo Montesi
IDEAS2
2019 On approximate nesting of multiple social network graphs: a preliminary study
abstract
A fundamental problem in Social Network Analysis is how to move from single-layer to multi-layer, which provide a holistic view. User profiles resolution has received considerable attention since it allows to match users on different online social networks (OSNs). However, to the best of our knowledge, no study has focused on nesting operation for merging OSNs graphs. This work is a first step in the direction of defining the data model and the algorithm to perform approximate nesting of multiple OSNs graphs, based on user features. We provide initial experimental evidence based on synthetic data.
Giacomo Bergami, Flavio Bertini 0001, Danilo Montesi
IDEAS2
2019 Fine-grain watermarking for intellectual property protection
abstract
The current online digital world, consisting of thousands of newspapers, blogs, social media, and cloud file sharing services, is providing easy and unlimited access to a large treasure of text contents. Making copies of these text contents is simple and virtually costless. As a result, producers and owners of text content are interested in the protection of their intellectual property (IP) rights. Digital watermarking has become crucially important in the protection of digital contents. Out of all, text watermarking poses many challenges, since text is characterized by a low capacity to embed a watermark and allows only a restricted number of alternative syntactic and semantic permutations. This becomes even harder when authors want to protect not just a whole book or article, but each single sentence or paragraph, a problem well known to copyright law. In this paper, we present a fine-grain text watermarking method that protects even small portions of the digital content. The core method is based on homoglyph characters substitution for latin symbols and whitespaces. It allows to produce a watermarked version of the original text, preserving the anonymity of the users according to the right to privacy. In particular, the embedding and extraction algorithms allow to continuously protect the watermark through the whole document in a fine-grain fashion. It ensures visual indistinguishability and length preservation, meaning that it does not cause overhead to the original document, and it is robust to the copy and past of small excerpts of the text. We use a real dataset of 1.8 million New York articles to evaluate our method. We evaluate and compare the robustness against common attacks, and we propose a new measure for partial copy and paste robustness. The results show the effectiveness of our approach providing an average length of 101 characters needed to embed the watermark and allowing to protect paragraph-long excerpt or smaller the 94.5% of the times.
Stefano Giovanni Rizzo, Flavio Bertini 0001, Danilo Montesi
EURASIP J. Inf. Secur.2
2018 A Cluster-based Approach of Smartphone Camera Fingerprint for User Profiles Resolution within Social Network
abstract
In the last decades, Social Networks (SNs) have deeply changed interactions and habits of the users that are also prone to create more than one profile on the same SN. On the flip side, fake profiles (i.e., impersonating profiles), have become a considerable problem in digital investigations. In this paper, we propose a method for user profiles resolution through a cluster-based approach of the smartphone fingerprints extracted from the images being posted on SNs. The proposed method is thus able to detect fake profiles. To evaluate our approach, we use a real dataset of 1,500 images from 10 different smartphone devices and Facebook and WhatsApp platforms. The results show that the average of sensitivity and specificity for user profiles resolution is about 98%.
Rahimeh Rouhi, Flavio Bertini 0001, Danilo Montesi
IDEAS2
2018 Predicting Frailty Condition in Elderly Using Multidimensional Socioclinical Databases
abstract
Smart cities face the challenge of combining sustainable national welfare with high living standards. In the last decades, life expectancy increased globally, leading to various age-related issues in almost all developed countries. Frailty affects elderly who are experiencing daily life limitations due to cognitive and functional impairments and represents a remarkable burden for national health systems. In this paper, we proposed two different predictive models for frailty by exploiting 12 socioclinical databases. Emergency hospitalization or all-cause mortality within a year were used as surrogates of frailty. The first model was able to assign a frailty risk score to each subject older than 65 years old, identifying five different classes for tailor made interventions. The second prediction model assigned a worsening risk score to each subject in the first nonfrail class, namely the probability to move in a higher frailty class within the year. We conducted a retrospective cohort study based on the whole elderly population of the Municipality of Bologna, Italy. We created a baseline cohort of 95 368 subjects for the frailty risk model and a baseline cohort of 58 789 subjects for the worsening risk model, respectively. To evaluate the predictive ability of our models through calibration and discrimination estimates, we used, respectively, a six-year and a four-year observation period. Good discriminatory power and calibration were obtained, demonstrating a good predictive ability of the models.
Flavio Bertini 0001, Giacomo Bergami, Danilo Montesi, Giacomo Veronese, Giulio Marchesini, Paolo Pandolfi
Proc. IEEE1
2017 Text Watermarking in Social Media
abstract
One of the most shared content in Social Media (SM) is text, making it vulnerable to copy and authorship misappropriation. Due to the low data noise, watermark embedding is very hard. This problem is exacerbated in the context of SM, where the amount of data in a single message can be extremely small, like in Twitter. Firstly, in this paper we investigate whether SM do applies watermarks on the texts. Then, we propose a text watermarking method able to work on all the SM platforms considered, while ensuring visual indistinguishability and length preservation of the original text and robustness to copy and paste. We conduct an extended evaluation on eighteen different SM platforms by using 6,000 posts from six public figures' profiles.
Stefano Giovanni Rizzo, Flavio Bertini 0001, Danilo Montesi, Carlo Stomeo
ASONAM2
2016 Content-preserving Text Watermarking through Unicode Homoglyph Substitution
abstract
Digital watermarking has become crucially important in authentication and copyright protection of the digital contents, since more and more data are daily generated and shared online through digital archives, blogs and social networks. Out of all, text watermarking is a more difficult task in comparison to other media watermarking. Text cannot be always converted into image, it accounts for a far smaller amount of data (eg. social network posts) and the changes in short texts would strongly affect the meaning or the overall visual form. In this paper we propose a text watermarking technique based on homoglyph characters substitution for latin symbols1. The proposed method is able to efficiently embed a password based watermark in short texts by strictly preserving the content. In particular, it uses alternative Unicode symbols to ensure visual indistinguishability and length preservation, namely content-preservation. To evaluate our method, we use a real dataset of 1.8 million New York articles. The results show the effectiveness of our approach providing an average length of 101 characters needed to embed a 64bit password based watermark.
Stefano Giovanni Rizzo, Flavio Bertini 0001, Danilo Montesi
IDEAS2
2015 Smartphone Verification and User Profiles Linking Across Social Networks by Camera Fingerprinting
Flavio Bertini 0001, Rajesh Sharma 0002, Andrea Ianni, Danilo Montesi
ICDF2C1
2015 Profile resolution across multilayer networks through smartphone camera fingerprint
abstract
In the last decade, various social platforms have been introduced on the web. Due to their specific orientation (friendship, professional connections, image sharing, etc.) users often join multiple networks. An important problem across these networks is the resolution of users profiles. That is, to identify if set of user profiles from different networks with different user ids or nicknames belong to the same user. The problem is more meaningful for resolving different profiles in digital forensic and criminal investigations. In this paper, we propose a method for profile resolution with the help of pictures being posted on different social platforms. We use the smartphone cameras which have become the source of instant image capturing and uploading process. In particular, we exploit the characteristic noise present in the images due to the manufacturing defects, to match user profiles across social platforms. To test our approach we select five different smartphones with two pairs of identical models, and three social platforms, namely Facebook, Google+ and WhatsApp. We evaluate our approach using real dataset of 1000 high-resolution pictures. The results indicate that even in the worst case our approach can provide profile matching upto 89.83%.
Flavio Bertini 0001, Rajesh Sharma 0002, Andrea Ianni, Danilo Montesi
IDEAS1