Juan Sebastian Piedrahita Giraldo

dblp:176/6369 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
3since 2021 · last 2021
0000-0002-1691-2915ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2021 ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators
abstract
Building efficient embedded deep learning systems requires a tight co-design between DNN algorithms, hardware, and algorithm-to-hardware mapping, a.k.a. dataflow. However, owing to the large joint design space, finding an optimal solution through physical implementation becomes infeasible. To tackle this problem, several design space exploration (DSE) frameworks have emerged recently, yet they either suffer from long runtimes or a limited exploration space. This article introduces ZigZag, a rapid DSE framework for DNN accelerator architecture and mapping. ZigZag extends the common DSE with uneven mapping opportunities and smart mapping search strategies. Uneven mapping decouples operands (W/I/O), memory hierarchy, and mappings (temporal/spatial), opening up a whole new space for DSE, and thus better design points are found by ZigZag compared to other SotAs. For this, ZigZag uses an enhanced nested-for-loop format as a uniform representation to integrate algorithm, accelerator, and algorithm-to-accelerator mapping. ZigZag consists of three key components: 1) an analytical energy-performance-area Hardware Cost Estimator, 2) two Mapping Search Engines that support spatial and temporal even/uneven mapping search, and 3) an Architecture Generator that auto-explores the wide memory hierarchy design space. Benchmarking experiments against published works, in-house accelerator, and existing DSE frameworks, together with three case studies, show the reliability and capability of ZigZag. Up to 64 percent more energy-efficient solutions are found compared to other SotAs, due to ZigZag's uneven mapping capabilities.
Linyan Mei, Pouya Houshmand, Vikram Jain, Juan Sebastian Piedrahita Giraldo, Marian Verhelst
IEEE Trans. Computers4
2021 Hardware Acceleration for Embedded Keyword Spotting: Tutorial and Survey
abstract
In recent years, Keyword Spotting (KWS) has become a crucial human–machine interface for mobile devices, allowing users to interact more naturally with their gadgets by leveraging their own voice. Due to privacy, latency and energy requirements, the execution of KWS tasks on the embedded device itself instead of in the cloud, has attracted significant attention from the research community. However, the constraints associated with embedded systems, including limited energy, memory, and computational capacity, represent a real challenge for the embedded deployment of such interfaces. In this article, we explore and guide the reader through the design of KWS systems. To support this overview, we extensively survey the different approaches taken by the recent state-of-the-art (SotA) at the algorithmic, architectural, and circuit level to enable KWS tasks in edge, devices. A quantitative and qualitative comparison between relevant SotA hardware platforms is carried out, highlighting the current design trends, as well as pointing out future research directions in the development of this technology.
Juan Sebastian Piedrahita Giraldo, Marian Verhelst
ACM Trans. Embed. Comput. Syst.1
2021 Efficient Execution of Temporal Convolutional Networks for Embedded Keyword Spotting
abstract
Recently, the use of keyword spotting (KWS) has become prevalent in mobile devices. State-of-the-art deep learning algorithms such as temporal convolutional networks (TCNs) have been applied to this task achieving superior accuracy results. These models can, however, be mapped in multiple ways onto embedded devices, ranging from real-time streaming inference with or without computational sprinting to delayed batched inference. Although functionally equivalent, these deployment settings, however, strongly impacts average power consumption and latency of this real time task, hence requiring a thorough optimization. This work analyzes the challenges, benefits, and drawbacks of the different execution modes available for TCN-based KWS inference on dedicated hardware. With this objective, this research contributes to: 1) presenting a complete deep learning accelerator optimized for TCN inference; 2) evaluating the impact on performance and power of the different deployment options for TCN inference applied to KWS obtaining up to 8$\mu \text{W}$for real-time operation; and 3) optimizing real-time power consumption for KWS inference by exploiting the use of cascaded neural networks (NNs), achieving up to 35% additional power savings.
Juan Sebastian Piedrahita Giraldo, Vikram Jain, Marian Verhelst
IEEE Trans. Very Large Scale Integr. Syst.1
2019 Efficient Keyword Spotting through Hardware-Aware Conditional Execution of Deep Neural Networks
abstract
Keyword spotting is a task that requires ultra-low power due to its always-on operation. State-of-the-art approaches achieve this by drastically pruning model size, yet often at the expense of accuracy. This work tackles this fundamental conflict between operating efficiency and accuracy in three ways: 1.) Exploiting dynamic neural network cascades for keyword spotting using an end-to-end hardware-aware training; 2.) Deriving the optimal number of stages and stage dimensions in function of the input class distributions; 3.) Using the low-latency response of the first stage for speculative execution of the later stages, training the dynamic cascade through a hardware-aware cost function. Results show the framework can generate cascade models optimized in function of the class distribution (background noise, target keywords and other-speech), reducing computational cost by 87% for always-on operation while maintaining the baseline accuracy of the most complex model of the cascade. On top of this, the hardware-aware speculative execution provides an additional 2x energy savings over the non-speculative case.
Juan Sebastian Piedrahita Giraldo, Chris O'Connor, Marian Verhelst
AICCSA1
2015 Evaluation of energy savings on a VLIW processor through dynamic issue-width adaptation
abstract
The development of energy efficient hardware has been a trend in microprocessor design for the last two decades. VLIW processors are a representative example, since they have a simpler design and competitive performance, because their ILP exploitation is done statically by the compiler. In this paper, we study the energy savings that could be obtained by adapting such microarchitecture according to the current program phase. Our contribution is twofold. First, by executing a set of benchmarks on the ρ-vex configurable softcore VLIW processor, and by modifying the number of issues, we show the potentials of energy reduction. Then, with this information in hand, we developed an oracle experiment to dynamically vary the issue width of the processor according to the phase behavior, considering two different phase granularites. The potential energy savings using this policy could be as high as 81.5% when compared with the static version, executing the MiBench set.
Juan Sebastian Piedrahita Giraldo, Anderson Luiz Sartor, Luigi Carro, Stephan Wong, Antonio Carlos Schneider Beck
RSP1