Melina Demertzi

dblp:70/4410 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Reconfigurable computing and FPGAs · 62% Hardware accelerators and domain-specific architectures · 38%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA accelerator
0.112009
Computation reuse in domain-specific optimization of signal recognition · FPGA 2009
Hardware accelerators and domain-specific architectures
computation reuse
0.012009
Computation reuse in domain-specific optimization of signal recognition · FPGA 2009
Hardware accelerators and domain-specific architectures
domain-specific optimization
0.012009
Computation reuse in domain-specific optimization of signal recognition · FPGA 2009

Methods — techniques the papers use, named apart from their topics

walsh wavelet packets · 0.2computation reuse scheduling · 0.2bestbasis algorithm · 0.2
YearPublicationVenuePosition
2017 Accelerating Java Streams With A Data AnalyticsHardware Accelerator
abstract
In this paper, we demonstrate how new technology from Oracle can be utilized to provide big data analytics acceleration in a streamlined fashion. Specifically, our approach leverages the acceleration capabilities of the Data Analytics Accelerator (DAX) unit provided by Oracle's T7/M7/S7 SPARC processors and the Java Stream API to seamlessly accelerate Java applications by up to 20X, while using drastically fewer resources.
Karthik Ganesan 0006, Ahmed Khawaja, Shrinivas Joshi, Michelle Szucs, Melina Demertzi, Yao-Min Chen
ICPE6
2011 Domain-Specific Optimization of Signal Recognition Targeting FPGAs
abstract
Domain-specific optimizations on matrix computations exploiting specific arithmetic and matrix representation formats have achieved significant performance/area gains in Field-Programmable Gate Array (FPGA) hardware designs. In this article, we explore the application of data-driven optimizations to reduce both storage and computation requirements to the problem of signal recognition from a known dictionary. By starting with a high-level mathematical representation of a signal recognition problem, we perform optimizations across the layers of the system, exploiting mathematical structure to improve implementation efficiency. Specifically, we use Walsh wavelet packets in conjunction with a BestBasis algorithm to distinguish between spoken digits. The resulting transform matrices are quite sparse, and exhibit a rich algebraic structure that contains significant overlap across rows. As a consequence, dot-product computations of the transform matrix and signal vectors exhibit significant computation reuse, or repeated identical computations. We present an algorithm for identifying this computation reuse and scheduling of the row computations. We exploit this reuse to derive FPGA hardware implementations that reduce the amount of computation for an individual matrix by as much as 6.35× and an average of 2× for a single dot-product unit. The implementation that exploits reuse achieves a 2× computation reduction compared to three concurrently-executing simpler accumulator units with the same aggregate design area and outperforms software implementations on high-end desktop personal computers.
Melina Demertzi, Pedro C. Diniz, Mary W. Hall, Anna Gilbert 0001
ACM Trans. Reconfigurable Technol. Syst.1
2009 Computation reuse in domain-specific optimization of signal recognition
abstract
Domain-specific optimizations that exploit specific arithmetic and representation formats have been shown to achieve significant performance/area gains in FPGA hardware designs. In this work, we describe an approach to domain-specific optimization that goes beyond this representation level. We perform a joint optimization from a high-level mathematical abstract representation and hardware implementation point of view. We focus on a signal recognition system that distinguishes between spoken digits. We construct transform matrices from Walsh wavelet packets in conjunction with a BestBasis algorithm. The resulting transform matrices exhibit a rich algebraic structure and contain significant overlap across rows, exhibiting significant computation reuse in the dot-product operation of the transform matrix applied to the signal vector. We have developed an algorithm for identifying the computation reuse and scheduling the row computations across various computation units to significantly reduce the overall amount of computation.
Melina Demertzi, Pedro C. Diniz, Mary W. Hall, Anna Gilbert 0001
FPGA1
2008 The potential of computation reuse in high-level optimization of a signal recognition system
abstract
This paper evaluates the potential of exploiting computation reuse in a signal recognition system that is jointly optimized from mathematical representation, algorithm design and final implementation. Walsh wavelet packets in conjunction with a BestBasis algorithm are used to derive transforms that discriminate between signals. The FPGA implementation of this computation exploits the structure of the resulting transform matrices in several ways to derive a highly optimized hardware representation of this signal recognition problem. Specifically, we observe in the transform matrices a significant amount of reuse of subrows, thus indicating redundant computation. Through analysis of this reuse, we discover the potential for a 3times reduction in the amount of computation of combining a transform matrix and signal. In this paper, we focus on how the implementation might exploit this reuse in a profitable way. By exploiting a subset of this computation reuse, the system can navigate the tradeoff space of reducing computation and the extra storage required.
Melina Demertzi, Pedro C. Diniz, Mary W. Hall, Anna Gilbert 0001
IPDPS1