Lorenzo Baldacci

dblp:94/1446 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0002-1638-4972ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
cost estimation
0.412019
A Cost Model for SPARK SQL · IEEE Trans. Knowl. Data Eng. 2019
Query processing and optimization
cost model
0.412019
A Cost Model for SPARK SQL · IEEE Trans. Knowl. Data Eng. 2019
Bioinformatics and computational biology › molecular informatics › cheminformatics
atom mapping
0.212016
Reaction Decoder Tool (RDT): extracting features from chemical reactions · Bioinform. 2016
Bioinformatics and computational biology
enzymatic reaction analysis
0.212016
Reaction Decoder Tool (RDT): extracting features from chemical reactions · Bioinform. 2016
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
metabolic pathway analysis
0.212016
Reaction Decoder Tool (RDT): extracting features from chemical reactions · Bioinform. 2016
Cloud and datacenter computing
cluster resource management and scheduling
0.112019
A Cost Model for SPARK SQL · IEEE Trans. Knowl. Data Eng. 2019

Methods — techniques the papers use, named apart from their topics

straggler handling · 0.8analytical modeling · 0.8dynamic programming · 0.2
YearPublicationVenuePosition
2019 A Cost Model for SPARK SQL
abstract
In this paper, we propose a novel cost model for Spark SQL. The cost model covers the class of Generalized Projection, Selection, Join (GPSJ) queries. The cost model keeps into account the network and IO costs as well as the most relevant CPU costs. The execution cost is computed starting from a physical plan produced by Spark. The set of operations adopted by Spark when executing a GPSJ query are analytically modeled based on the cluster and application parameters, together with a set of database statistics. Experimental results carried out on three benchmarks and on two clusters of different sizes and with different computation features show that our model can estimate the actual execution time with about the 20 percent of errors on the average. Such an accuracy is good enough to let the system choose the most effective plan even when the execution time differences are limited. The error can be reduced to 14 percent, if the analytic model is coupled with our straggler handling strategy.
Lorenzo Baldacci, Matteo Golfarelli
IEEE Trans. Knowl. Data Eng.1
2017 QETL: An approach to on-demand ETL from non-owned data sources
Lorenzo Baldacci, Matteo Golfarelli, Simone Graziani, Stefano Rizzi
Data Knowl. Eng.1
2016 Reaction Decoder Tool (RDT): extracting features from chemical reactions
abstract
UNLABELLED: Extracting chemical features like Atom-Atom Mapping (AAM), Bond Changes (BCs) and Reaction Centres from biochemical reactions helps us understand the chemical composition of enzymatic reactions. Reaction Decoder is a robust command line tool, which performs this task with high accuracy. It supports standard chemical input/output exchange formats i.e. RXN/SMILES, computes AAM, highlights BCs and creates images of the mapped reaction. This aids in the analysis of metabolic pathways and the ability to perform comparative studies of chemical reactions based on these features. AVAILABILITY AND IMPLEMENTATION: This software is implemented in Java, supported on Windows, Linux and Mac OSX, and freely available at https://github.com/asad/ReactionDecoder CONTACT: : [email protected] or [email protected].
Syed Asad Rahman, Gilliean Torrance, Lorenzo Baldacci, Sergio Martínez Cuesta, Franz Fenninger, Nimish Gopal, Saket Choudhary, John W. May, Gemma L. Holliday, Christoph Steinbeck, Janet M. Thornton
Bioinform.3
2016 Natural gas consumption forecasting for anomaly detection
Lorenzo Baldacci, Matteo Golfarelli, Davide Lombardi, Franco Sami
Expert Syst. Appl.1
2014 GOLAM: A Framework for Analyzing Genomic Data
abstract
The emerging medical models aim at leveraging on high-throughput genome sequencing technologies to better target drugs to patients' personal profiles so as to increase their effectiveness. However, the huge amount of data made available by these technologies calls for sophisticated and automated analysis techniques. In this direction we present GOLAM, a framework for OLAP analysis and mining of matches between genomic regions extracted from ENCODE, a worldwide-available collection of shared genomic data. The goal of GOLAM is to overcome the current limitations of genome analysis methods, that are normally based on browsing. This is done by partially automating and speeding-up the analysis process on the one hand, by making it more flexible and introducing a multi-resolution view of data on the other. The framework has been partially implemented so far; in this paper we focus on conveying its potential and on describing its functional architecture and the underlying data models.
Lorenzo Baldacci, Matteo Golfarelli, Simone Graziani, Stefano Rizzi
DOLAP1
2006 Clustering techniques for protein surfaces
Lorenzo Baldacci, Matteo Golfarelli, Alessandra Lumini, Stefano Rizzi
Pattern Recognit.1