Giovanni Matteo Fumarola

dblp:164/7907 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2Computer networks · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 85% Embedded and real-time systems · 9% Storage systems · 3%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%
Network and information security
1 paper
Authentication and access control · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.932019
Hydra: a federated resource manager for data-center scale analytics · NSDI 2019
Preemption-aware planning on big-data systems · PPoPP 2016
History-Based Harvesting of Spare Cycles and Storage in Large-Scale Datacenters · OSDI 2016
Query processing and optimization › query optimization › transformation-based optimization
query plan rewrite
0.812024
Membrane - Safe and Performant Data Access Controls in Apache Spark in the Presence of Imperative Code · Proc. VLDB Endow. 2024
Authentication and access control › access control
data access control
0.812024
Membrane - Safe and Performant Data Access Controls in Apache Spark in the Presence of Imperative Code · Proc. VLDB Endow. 2024
Cloud and datacenter computing › cluster resource management and scheduling
resource scheduling
0.622019
Hydra: a federated resource manager for data-center scale analytics · NSDI 2019
History-Based Harvesting of Spare Cycles and Storage in Large-Scale Datacenters · OSDI 2016
Embedded and real-time systems › real-time scheduling
reservation-based scheduling
0.212016
Preemption-aware planning on big-data systems · PPoPP 2016
Cloud and datacenter computing › serverless computing
serverless analytics
0.212024
Membrane - Safe and Performant Data Access Controls in Apache Spark in the Presence of Imperative Code · Proc. VLDB Endow. 2024
Cloud and datacenter computing
cluster resource management and scheduling
0.212015
Mercury: Hybrid Centralized and Distributed Scheduling in Large Shared Clusters · USENIX ATC 2015
Storage systems
distributed storage
0.112016
History-Based Harvesting of Spare Cycles and Storage in Large-Scale Datacenters · OSDI 2016
Distributed systems
distributed coordination
0.112015
Mercury: Hybrid Centralized and Distributed Scheduling in Large Shared Clusters · USENIX ATC 2015
Cloud and datacenter computing
resource allocation
0.112015
Mercury: Hybrid Centralized and Distributed Scheduling in Large Shared Clusters · USENIX ATC 2015

Methods — techniques the papers use, named apart from their topics

query plan rewriting · 2.3container isolation · 2.3planning algorithm · 0.2
YearPublicationVenuePosition
2024 Membrane - Safe and Performant Data Access Controls in Apache Spark in the Presence of Imperative Code
abstract
Data Governance is an increasingly critical feature of modern cloud database systems, enabling administrators to set granular access policies on their data. AWS customers want to define row or column filtering on their blob storage data and access it using popular tools such as Apache Spark. AWS EMR provides a managed and serverless solution that lets users run Spark jobs in the AWS cloud with imperative and declarative programming against their data, while securely enforcing the fine-grained access controls defined on those datasets. Spark runs its compiler and scheduler alongside the user application and embeds user-defined functions in query plans, giving a threat actor direct access to its memory space. This introduces attack vectors such as information disclosure or privilege escalation during policy enforcement, in addition to well-researched threats such as SQL side channel attacks. In this paper, we present Membrane: a novel approach to secure query plans with declarative and imperative code. The innovation comes from splitting the Spark driver in two in order to rewrite query plans with security boundaries while avoiding traditional tradeoffs when using container isolation techniques. The approach described herein enables applying fine grained data access controls to both SQL and map-reduce Spark jobs, with negligible performance and cost differences.
Andrei Paduroiu, Sungheun Wi, Roni Burd, Ruhollah A Farchtchi, Giovanni Matteo Fumarola
Proc. VLDB Endow.6
2019 Hydra: a federated resource manager for data-center scale analytics
Carlo Curino, Subru Krishnan, Konstantinos Karanasos, Sriram Rao, Giovanni Matteo Fumarola, Botong Huang, Kishore Chaliparambil, Arun Suresh, Young Chen, Solom Heddaya, Roni Burd, Sarvesh Sakalanaga, Chris Douglas, Bill Ramsey, Raghu Ramakrishnan 0001
NSDI5
2016 History-Based Harvesting of Spare Cycles and Storage in Large-Scale Datacenters
George Prekas, Giovanni Matteo Fumarola, Marcus Fontoura, Íñigo Goiri, Ricardo Bianchini
OSDI3
2016 Preemption-aware planning on big-data systems
abstract
Recent developments in Big Data frameworks are moving towards reservation based approaches as a mean to manage the increasingly complex mix of computations, whereas preemption techniques are employed to meet strict jobs deadlines. Within this work we propose and evaluate a new planning algorithm in the context of reservation based scheduling. Our approach is able to achieve high cluster utilization while minimizing the need for preemption that causes system overheads and planning mispredictions.
Marco Rabozzi, Matteo Mazzucchelli, Roberto Cordone, Giovanni Matteo Fumarola, Marco D. Santambrogio
PPoPP4
2015 Mercury: Hybrid Centralized and Distributed Scheduling in Large Shared Clusters
Konstantinos Karanasos, Sriram Rao, Carlo Curino, Chris Douglas, Kishore Chaliparambil, Giovanni Matteo Fumarola, Solom Heddaya, Raghu Ramakrishnan 0001, Sarvesh Sakalanaga
USENIX ATC6