VLDB 2026 Research / reviewers in the wild / expert
Francesco Archetti
dblp:17/814
· DBLP profile ↗
20ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0003-1131-3830ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorTheory of computation · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gaussian Process regression over discrete probability measures: on the non-stationarity relation between Euclidean and Wasserstein Squared Exponential KernelsabstractAbstract Gaussian Process regression is a kernel method successfully adopted in many real-life applications. Recently, there is a growing interest on extending this method to non-Euclidean input spaces, like the one considered in this paper, consisting of probability measures. Although a Positive Definite kernel can be defined by using a suitable distance—the Wasserstein distance— the common procedure for learning the Gaussian Process model can fail due to numerical issues, arising earlier and more frequently than in the case of an Euclidean input space and, as demonstrated, impossible to avoid by adding artificial noise ( nugget effect ) as usually done. This paper uncovers the main reason of these issues, that is a non-stationarity relation between the Wasserstein-based squared exponential kernel and its Euclidean counterpart. As a relevant result, we learn a Gaussian Process model by assuming the input space as Euclidean and then use an algebraic transformation, based on the uncovered relation, to transform it into a non-stationary and Wasserstein-based Gaussian Process model over probability measures. This algebraic transformation is simpler than log-exp maps used on data belonging to Riemannian manifolds and recently extended to consider the pseudo-Riemannian structure of an input space equipped with the Wasserstein distance. Antonio Candelieri, Andrea Ponti, Francesco Archetti |
J. Glob. Optim. | 3 |
| 2024 | Fair and green hyperparameter optimization via multi-objective and multiple information source Bayesian optimizationabstractAbstract It has been recently remarked that focusing only on accuracy in searching for optimal Machine Learning models amplifies biases contained in the data, leading to unfair predictions and decision supports. Recently, multi-objective hyperparameter optimization has been proposed to search for Machine Learning models which offer equally Pareto-efficient trade-offs between accuracy and fairness. Although these approaches proved to be more versatile than fairness-aware Machine Learning algorithms—which instead optimize accuracy constrained to some threshold on fairness—their carbon footprint could be dramatic, due to the large amount of energy required in the case of large datasets. We propose an approach named FanG-HPO: fair and green hyperparameter optimization (HPO), based on both multi-objective and multiple information source Bayesian optimization. FanG-HPO uses subsets of the large dataset to obtain cheap approximations (aka information sources) of both accuracy and fairness, and multi-objective Bayesian optimization to efficiently identify Pareto-efficient (accurate and fair) Machine Learning models. Experiments consider four benchmark (fairness) datasets and four Machine Learning algorithms, and provide an assessment of FanG-HPO against both fairness-aware Machine Learning approaches and two state-of-the-art Bayesian optimization tools addressing multi-objective and energy-aware optimization. Antonio Candelieri, Andrea Ponti, Francesco Archetti |
Mach. Learn. | 3 |
| 2022 | AutoTinyML for microcontrollers: Dealing with black-box deployability
Riccardo Perego, Antonio Candelieri, Francesco Archetti, Danilo Pau |
Expert Syst. Appl. | 3 |
| 2021 | Green machine learning via augmented Gaussian processes and multi-information source optimizationabstractAbstract Searching for accurate machine and deep learning models is a computationally expensive and awfully energivorous process. A strategy which has been recently gaining importance to drastically reduce computational time and energy consumed is to exploit the availability of different information sources, with different computational costs and different “fidelity,” typically smaller portions of a large dataset. The multi-source optimization strategy fits into the scheme of Gaussian Process-based Bayesian Optimization. An Augmented Gaussian Process method exploiting multiple information sources (namely, AGP-MISO) is proposed. The Augmented Gaussian Process is trained using only “reliable” information among available sources. A novel acquisition function is defined according to the Augmented Gaussian Process. Computational results are reported related to the optimization of the hyperparameters of a Support Vector Machine (SVM) classifier using two sources: a large dataset—the most expensive one—and a smaller portion of it. A comparison with a traditional Bayesian Optimization approach to optimize the hyperparameters of the SVM classifier on the large dataset only is reported. Antonio Candelieri, Riccardo Perego, Francesco Archetti |
Soft Comput. | 3 |
| 2020 | Tuning Deep Neural Network's Hyperparameters Constrained to Deployability on Tiny Systems
Riccardo Perego, Antonio Candelieri, Francesco Archetti, Danilo Pau |
ICANN (2) | 3 |
| 2020 | Modelling human active search in optimizing black-box functionsabstractAbstract Modelling human function learning has been the subject of intense research in cognitive sciences. The topic is relevant in black-box optimization where information about the objective and/or constraints is not available and must be learned through function evaluations. In this paper, we focus on the relation between the behaviour of humans searching for the maximum and the probabilistic model used in Bayesian optimization. As surrogate models of the unknown function, both Gaussian processes and random forest have been considered: the Bayesian learning paradigm is central in the development of active learning approaches balancing exploration/exploitation in uncertain conditions towards effective generalization in large decision spaces. In this paper, we analyse experimentally how Bayesian optimization compares to humans searching for the maximum of an unknown 2D function. A set of controlled experiments with 60 subjects, using both surrogate models, confirm that Bayesian optimization provides a general model to represent individual patterns of active learning in humans. Antonio Candelieri, Riccardo Perego, Ilaria Giordani, Andrea Ponti, Francesco Archetti |
Soft Comput. | 5 |
| 2019 | Global optimization in machine learning: the design of a predictive analytics application
Antonio Candelieri, Francesco Archetti |
Soft Comput. | 2 |
| 2018 | Automated Rehabilitation Exercises Assessment in Wearable Sensor Data StreamsabstractThis work stems from the Italian project H-CIM (Health-Care Intelligent Monitoring), aimed at developing a wearable sensor data streams based home-monitoring system to support self-rehabilitation of elderly outpatients. Different from the pervasive data stream applications, which are always accompanied by the evolution of unstable class concepts, this project requires stable standard and personalized rehabilitation exercises patterns be provided to assess outpatient's self-therapy progress at home. In this designed pipeline, the representation sequences of the personal standard rehabilitation exercises in wearable sensor streams is therefore first benchmarked, then an assessment system which integrates multistage data processing and analyzing is proposed to enable elders to manage their own rehabilitation progress properly. The system proved to be an effective tool for supporting compliance monitoring and personalized self-rehabilitation; it is currently under further development within the Italian project Home-IoT, with the aim to become a more general data stream analytics service, not only devoted to rehabilitation exercises assessment. Antonio Candelieri, Wenbin Zhang 0002, Enza Messina, Francesco Archetti |
IEEE BigData | 4 |
| 2018 | Bayesian optimization of pump operations in water distribution systemsabstractBayesian optimization has become a widely used tool in the optimization and machine learning communities. It is suitable to problems as simulation/optimization and/or with an objective function computationally expensive to evaluate. Bayesian optimization is based on a surrogate probabilistic model of the objective whose mean and variance are sequentially updated using the observations and an “acquisition” function based on the model, which sets the next observation at the most “promising” point. The most used surrogate model is the Gaussian Process which is the basis of well-known Kriging algorithms. In this paper, the authors consider the pump scheduling optimization problem in a Water Distribution Network with both ON/OFF and variable speed pumps. In a global optimization model, accounting for time patterns of demand and energy price allows significant cost savings. Nonlinearities, and binary decisions in the case of ON/OFF pumps, make pump scheduling optimization computationally challenging, even for small Water Distribution Networks. The well-known EPANET simulator is used to compute the energy cost associated to a pump schedule and to verify that hydraulic constraints are not violated and demand is met. Two Bayesian Optimization approaches are proposed in this paper, where the surrogate model is based on a Gaussian Process and a Random Forest, respectively. Both approaches are tested with different acquisition functions on a set of test functions, a benchmark Water Distribution Network from the literature and a large-scale real-life Water Distribution Network in Milan, Italy. Antonio Candelieri, Raffaele Perego 0002, Francesco Archetti |
J. Glob. Optim. | 3 |
| 2014 | A p-Median approach for predicting drug response in tumour cellsabstractBACKGROUND: The complexity of biological data related to the genetic origins of tumour cells, originates significant challenges to glean valuable knowledge that can be used to predict therapeutic responses. In order to discover a link between gene expression profiles and drug responses, a computational framework based on Consensus p-Median clustering is proposed. The main goal is to simultaneously predict (in silico) anticancer responses by extracting common patterns among tumour cell lines, selecting genes that could potentially explain the therapy outcome and finally learning a probabilistic model able to predict the therapeutic responses. RESULTS: The experimental investigation performed on the NCI60 dataset highlights three main findings: (1) Consensus p-Median is able to create groups of cell lines that are highly correlated both in terms of gene expression and drug response; (2) from a biological point of view, the proposed approach enables the selection of genes that are strongly involved in several cancer processes; (3) the final prediction of drug responses, built upon Consensus p-Median and the selected genes, represents a promising step for predicting potential useful drugs. CONCLUSION: The proposed learning framework represents a promising approach predicting drug response in tumour cells. Elisabetta Fersini, Enza Messina, Francesco Archetti |
BMC Bioinform. | 3 |
| 2012 | Emotional states in judicial courtrooms: An experimental investigation
Elisabetta Fersini, Enza Messina, Francesco Archetti |
Speech Commun. | 3 |
| 2010 | Semantics and Machine Learning: A New Generation of Court Management Systems
Elisabetta Fersini, Enza Messina, Francesco Archetti, Mauro Cislaghi |
IC3K | 3 |
| 2010 | Web Page Classification: A Probabilistic Model with Relational Uncertainty
Elisabetta Fersini, Enza Messina, Francesco Archetti |
IPMU | 3 |
| 2010 | A probabilistic relational approach for web document clustering
Elisabetta Fersini, Enza Messina, Francesco Archetti |
Inf. Process. Manag. | 3 |
| 2009 | An integrated communications framework for context aware continuous monitoring with body sensor networksabstractThis paper deals with a wireless pervasive communication system to support advanced healthcare applications. The proposed system is based on an ad hoc interaction of mobile body sensor networks with independent wireless sensor networks already deployed within the environments in order to allow a continuous and context aware health monitoring for patients along their daily life scenarios with an unprecedented precision and flexibility of sensing. After an accurate protocol characterization, simulation results are provided, underlining remarkable performance with respect to existing solutions, for different mobility models and node density values. Francesco Chiti, Romano Fantacci, Francesco Archetti, Enza Messina, Daniele Toscani |
IEEE J. Sel. Areas Commun. | 3 |
| 2008 | Enhancing web page classification through image-block importance analysis
Elisabetta Fersini, Enza Messina, Francesco Archetti |
Inf. Process. Manag. | 3 |
| 2006 | Foreground-to-Ghost Discrimination in Single-Difference Pre-processing
Francesco Archetti, Cristina E. Manfredotti, Enza Messina, Domenico G. Sorrenti |
ACIVS | 1 |
| 2006 | A Hierarchical Document Clustering Environment Based on the Induced Bisecting k-Means
Francesco Archetti, P. Campanelli, Elisabetta Fersini, Enza Messina |
FQAS | 1 |
| 2006 | Genetic programming for human oral bioavailability of drugsabstractAutomatically assessing the value of bioavailability from the chemical structure of a molecule is a very important issue in biomedicine and pharmacology. In this paper, we present an empirical study of some well known Machine Learning techniques, including various versions of Genetic Programming, which have been trained to this aim using a dataset of molecules with known bioavailability. Genetic Programming has proven the most promising technique among the ones that have been considered both from the point of view of the accurateness of the solutions proposed, of the generalization capabilities and of the correlation between predicted data and correct ones. Our work represents a first answer to the demand for quantitative bioavailability estimation methods proposed in literature, since the previous contributions focus on the classification of molecules into classes with similar bioavailability. Francesco Archetti, Stefano Lanzeni, Enza Messina, Leonardo Vanneschi |
GECCO | 1 |
| 1990 | Petri net-based emulation for a highly concurrent pick-and-place machineabstractThe emulation of pick-and-place machines to compute their throughput rate for different products is a key ingredient for capacity planning and scheduling of high-volume assembly lines where these machines are used. It is shown how an emulator can be built in the framework of Petri nets and how it can be instantiated for a particular setup and sequence of placements. Three advantages are argued for the proposed emulator. First, the emulation code can be generated automatically. Second, the model graphically represents the concurrency and synchronization aspects of the machine's operations. Third, it allows a formal representation of the machine's operations.> Anna Sciomachen, Stephen J. Grotzinger, Francesco Archetti |
IEEE Trans. Robotics Autom. | 3 |