Manuel Francisco

dblp:276/9369 · also Manuel Francisco Aparicio · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0001-9748-2269ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Multi-Label Quantification
abstract
Quantification, variously called supervised prevalence estimation or learning to quantify , is the supervised learning task of generating predictors of the relative frequencies (a.k.a. prevalence values ) of the classes of interest in unlabelled data samples. While many quantification methods have been proposed in the past for binary problems and, to a lesser extent, single-label multiclass problems, the multi-label setting (i.e., the scenario in which the classes of interest are not mutually exclusive) remains by and large unexplored. A straightforward solution to the multi-label quantification problem could simply consist of recasting the problem as a set of independent binary quantification problems. Such a solution is simple but naïve, since the independence assumption upon which it rests is, in most cases, not satisfied. In these cases, knowing the relative frequency of one class could be of help in determining the prevalence of other related classes. We propose the first truly multi-label quantification methods, i.e., methods for inferring estimators of class prevalence values that strive to leverage the stochastic dependencies among the classes of interest in order to predict their relative frequencies more accurately. We show empirical evidence that natively multi-label solutions outperform the naïve approaches by a large margin. The code to reproduce all our experiments is available online.
Alejandro Moreo, Manuel Francisco, Fabrizio Sebastiani 0001
ACM Trans. Knowl. Discov. Data2
2023 A Methodology to Quickly Perform Opinion Mining and Build Supervised Datasets Using Social Networks Mechanics
abstract
Social Networking Sites (SNS) offer a full set of possibilities to perform opinion studies such as polling or market analysis. Normally, artificial intelligence techniques are applied, and they often require supervised datasets. The process of building these is complex, time-consuming and expensive. In this paper, we propose to assist the labelling task by taking advantage of social network mechanics. In order to do that, we introduce theco-retweetrelation to build a graph that allows us to propagate user labels to their similarity neighbourhood. Therefore, it is possible to iteratively build supervised datasets with significant less human effort and with higher accuracy than other weak-supervision techniques. We tested our proposal with 3 datasets labelled by an expert committee, and results shows that it outperforms other weak-supervision techniques. This methodology may be adapted to other social networks and topics, and it is relevant for applications like informed decision-making (e.g., content moderation), specially when interpretability is required.
Manuel Francisco, Juan Luis Castro
IEEE Trans. Knowl. Data Eng.1
2022 Similarity Fuzzy Semantic Network for Social Media Analysis
Juan Luis Castro, Manuel Francisco
IPMU (1)2
2020 A fuzzy model to enhance user profiles in microblogging sites using deep relations
Manuel Francisco, Juan Luis Castro
Fuzzy Sets Syst.1