Andrea Lottarini

dblp:142/3206 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Hardware accelerators and domain-specific architectures · 46% Interconnection networks and networks-on-chip · 23% Cloud and datacenter computing · 13%
Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 67% Database system architecture and tuning · 33%

Topics — the 8 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
database accelerator
0.522017
Network Synthesis for Database Processing Units · DAC 2017
Q100: the architecture and design of a database processing unit · ASPLOS 2014
Query processing and optimization
analytical query processing
0.412019
Master of none acceleration: a comparison of accelerator architectures for analytical query processing · ISCA 2019
Interconnection networks and networks-on-chip › network topology › network topology design
network-on-chip topology
0.312017
Network Synthesis for Database Processing Units · DAC 2017
Interconnection networks and networks-on-chip
network topology
0.312017
Network Synthesis for Database Processing Units · DAC 2017
Hardware accelerators and domain-specific architectures › database accelerator
query accelerator
0.312017
Network Synthesis for Database Processing Units · DAC 2017
Performance modeling and evaluation
benchmarking
0.112018
vbench: Benchmarking Video Transcoding in the Cloud · ASPLOS 2018
GPUs and heterogeneous computing
GPU computing
0.112018
vbench: Benchmarking Video Transcoding in the Cloud · ASPLOS 2018
Performance modeling and evaluation
workload characterization
0.112018
vbench: Benchmarking Video Transcoding in the Cloud · ASPLOS 2018

Methods — techniques the papers use, named apart from their topics

simulation · 0.4coarse-grained instruction · 0.4ASIC tile · 0.4microarchitectural profiling · 0.3topology exploration · 0.3
YearPublicationVenuePosition
2019 Master of none acceleration: a comparison of accelerator architectures for analytical query processing
abstract
Hardware accelerators are one promising solution to contend with the end of Dennard scaling and the slowdown of Moore's law. For mature workloads that are regular and have high compute per byte, hardening an application into one or more hardware modules is a standard approach. However, for some applications, we find that a programmable homogeneous architecture is preferable.
Andrea Lottarini, Joao Pedro Cerqueira, Thomas J. Repetti, Stephen A. Edwards, Kenneth A. Ross, Mingoo Seok, Martha A. Kim
ISCA1
2018 vbench: Benchmarking Video Transcoding in the Cloud
abstract
This paper presents vbench, a publicly available benchmark for cloud video services. We are the first study, to the best of our knowledge, to characterize the emerging video-as-a-service workload. Unlike prior video processing benchmarks, vbench's videos are algorithmically selected to represent a large commercial corpus of millions of videos. Reflecting the complex infrastructure that processes and hosts these videos, vbench includes carefully constructed metrics and baselines. The combination of validated corpus, baselines, and metrics reveal nuanced tradeoffs between speed, quality, and compression. We demonstrate the importance of video selection with a microarchitectural study of cache, branch, and SIMD behavior. vbench reveals trends from the commercial corpus that are not visible in other video corpuses. Our experiments with GPUs under vbench's scoring scenarios reveal that context is critical: GPUs are well suited for live-streaming, while for video-on-demand shift costs from compute to storage and network. Counterintuitively, they are not viable for popular videos, for which highly compressed, high quality copies are required. We instead find that popular videos are currently well-served by the current trajectory of software encoders.
Andrea Lottarini, Alex Ramírez, Joel Coburn, Martha A. Kim, Parthasarathy Ranganathan, Daniel Stodolsky, Mark Wachsler
ASPLOS1
2017 Network Synthesis for Database Processing Units
abstract
We explore on-chip network topologies for the Q100, an analytic query accelerator for relational databases. In such data-centric accelerators, interconnects play a critical role by moving large volumes of data. In this paper we show that various interconnect topologies can trade a factor of 2.5x in performance for 3.3x area. Moreover, standard topologies (e.g., ring or mesh) are not optimal.
Andrea Lottarini, Stephen A. Edwards, Kenneth A. Ross, Martha A. Kim
DAC1
2014 Q100: the architecture and design of a database processing unit
abstract
In this paper, we propose Database Processing Units, or DPUs, a class of domain-specific database processors that can efficiently handle database applications. As a proof of concept, we present the instruction set architecture, microarchitecture, and hardware implementation of one DPU, called Q100. The Q100 has a collection of heterogeneous ASIC tiles that process relational tables and columns quickly and energy-efficiently. The architecture uses coarse grained in- structions that manipulate streams of data, thereby maximizing pipeline and data parallelism, and minimizing the need to time multiplex the accelerator tiles and spill inter- mediate results to memory. This work explores a Q100 de- sign space of 150 configurations, selecting three for further analysis: a small, power-conscious implementation, a high- performance implementation, and a balanced design that maximizes performance per Watt. We then demonstrate that the power-conscious Q100 handles the TPC-H queries with three orders of magnitude less energy than a state of the art software DBMS, while the performance-oriented design out- performs the same DBMS by 70X.
Lisa Wu Wills, Andrea Lottarini, Timothy K. Paine, Martha A. Kim, Kenneth A. Ross
ASPLOS2