Alberto Delmas Lascorz

dblp:202/1799 · also Alberto Delmas · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
2since 2021 · last 2024
0009-0006-2487-250XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Hardware accelerators and domain-specific architectures · 71% Memory systems · 21% Processor architecture and microarchitecture · 5%
Artificial intelligence
3 papers
Efficient and distributed learning · 100%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
1.532024
Atalanta: A Bit is Worth a "Thousand" Tensor Values · ASPLOS (2) 2024
Laconic deep learning inference acceleration · ISCA 2019
Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks · ASPLOS 2019
Memory systems
memory compression
1.332024
Atalanta: A Bit is Worth a "Thousand" Tensor Values · ASPLOS (2) 2024
ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019
Late Breaking Results: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick · DAC 2020
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.132020
Late Breaking Results: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick · DAC 2020
ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018
Machine learning › Efficient and distributed learning
inference acceleration
0.412019
Laconic deep learning inference acceleration · ISCA 2019
Hardware accelerators and domain-specific architectures › sparsity exploitation
bit-level sparsity exploitation
0.412019
Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks · ASPLOS 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN inference accelerator
0.312018
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN accelerator
precision-scalable accelerator
0.312018
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018
Image and video processing › image restoration
denoising
0.312017
IDEAL: image denoising accelerator · MICRO 2017
Image and video processing
image restoration
0.312017
IDEAL: image denoising accelerator · MICRO 2017
Processor architecture and microarchitecture
data-parallel architecture
0.312017
Bit-pragmatic deep neural network computing · MICRO 2017
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.312017
Bit-pragmatic deep neural network computing · MICRO 2017
Machine learning › Efficient and distributed learning › model compression
quantization
0.222019
ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018
Machine learning › Efficient and distributed learning
model compression
0.112018
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018
Energy-efficient computing
energy characterization
0.112017
IDEAL: image denoising accelerator · MICRO 2017
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.112017
Bit-pragmatic deep neural network computing · MICRO 2017

Methods — techniques the papers use, named apart from their topics

hardware-software co-design · 0.8fine-grain data width encoding · 0.8dynamic width selection · 0.8arithmetic coding · 0.8per-layer precision · 0.7bit-parallel multiply-accumulate · 0.7lossless compression · 0.4fixed-point quantization · 0.4static scheduling · 0.4sparse shuffling network · 0.4
YearPublicationVenuePosition
2024 Atalanta: A Bit is Worth a "Thousand" Tensor Values
abstract
Atalanta is a lossless, hardware/software co-designed compression technique for the tensors of fixed-point quantized deep neural networks. Atalanta increases effective memory capacity, reduces off-die traffic, and/or helps to achieve the desired performance/energy targets while using smaller off-die memories during inference. Atalanta is architected to deliver nearly identical coding efficiency compared to Arithmetic Coding while avoiding its complexity, overhead, and bandwidth limitations. Indicatively, the Atalanta decoder and encoder units each use less than 50B of internal storage. In hardware, Atalanta is implemented as an assist over any machine learning accelerator transparently compressing/decompressing tensors just before the off-die memory controller. This work shows the performance and energy efficiency of Atalanta when implemented in a 65nm technology node. Atalanta reduces data footprint of weights and activations to 60% and 48% respectively on average over a wide set of 8-bit quantized models and complements a wide range of quantization methods. Integrated with a Tensorcore-based accelerator, Atalanta boosts the speedup and energy efficiency to 1.44× and 1.37×, respectively. Atalanta is effective at compressing the stashed activations during training for fixed-point inference.
Alberto Delmas Lascorz, Mostafa Mahmoud, Ali Hadi Zadeh, Milos Nikolic 0002, Kareem Ibrahim, Christina Giannoula, Ameer Abdelhadi, Andreas Moshovos
ASPLOS (2)1
2024 BitPruning: Learning Bitlengths for Aggressive and Accurate Quantization
abstract
BitPruning is a training method for minimizing inference bitlengths at any granularity while maintaining accuracy. BitPruning extends the meaning of fixed-point bitlenghts into the continuous domain by interpolating between the nearest two integers, enabling gradient descent to learn bitlengths together with other parameters. A novel regularizer penalizes large bitlength representations and can be modified to minimize other quantifiable criteria, such as number of operations or memory footprint. BitPruning learns thrifty representations while maintaining accuracy: With ImageNet, it produces an average per layer bitlength of 3.76 and 4.36 bits on ResNet18 and MobileNet V2 respectively, remaining within 0.5% of the base TOP-1 accuracy. Simple modifications of the BitPruning regularizer can be used to further reduce compute workload by up to 24%, as well as memory footprint in activation or weight-heavy tasks by up to 14% and 8% respectively.
Milos Nikolic 0002, Ghouthi Boukli Hacene, Ciaran Bannon, Alberto Delmas Lascorz, Matthieu Courbariaux, Omar Mohamed Awad, Isak Edo Vivancos, Yoshua Bengio, Vincent Gripon, Andreas Moshovos
ISCAS4
2020 Late Breaking Results: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick
abstract
Data accesses between on- and off-chip memories account for a large fraction of overall energy consumption during inference with deep learning networks. We present Boveda, a lossless on-chip memory compression technique for neural networks operating on fixed-point values. Boveda reduces the datawidth used per block of values to be only as long as necessary: since most values are of small magnitude Boveda drastically reduces their footprint. Boveda can be used to increase the effective on-chip capacity, to reduce off-chip traffic, or to reduce the on-chip memory capacity needed to achieve a performance/energy target. Boveda reduces total model footprint to 53%.
Isak Edo Vivancos, Sayeh Sharify, Milos Nikolic 0002, Ciaran Bannon, Mostafa Mahmoud, Alberto Delmas Lascorz, Andreas Moshovos
DAC6
2019 Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks
abstract
Weight and activation sparsity can be leveraged in hardware to boost the performance and energy efficiency of Deep Neural Networks during inference. Fully capitalizing on sparsity requires re-scheduling and mapping the execution stream to deliver non-zero weight/activation pairs to multiplier units for maximal utilization and reuse. However, permitting arbitrary value re-scheduling in memory space and in time places a considerable burden on hardware to perform dynamic at-runtime routing and matching of values, and incurs significant energy inefficiencies. Bit-Tactical (TCL) is a neural network accelerator where the responsibility for exploiting weight sparsity is shared between a novel static scheduling middleware, and a co-designed hardware front-end with a lightweight sparse shuffling network comprising two (2- to 8-input) multiplexers per activation input. We empirically motivate two back-end designs chosen to target bit-sparsity in activations, rather than value-sparsity, with two benefits: a) we avoid handling the dynamically sparse whole-value activation stream, and b) we uncover more ineffectual work. TCL outperforms other state-of-the-art accelerators that target sparsity for weights and activations, the dynamic precision requirements of activations, or their bit-level sparsity for a variety of neural networks.
Alberto Delmas Lascorz, Patrick Judd, Dylan Malone Stuart, Zissis Poulos, Mostafa Mahmoud, Sayeh Sharify, Milos Nikolic 0002, Kevin Siu, Andreas Moshovos
ASPLOS1
2019 Laconic deep learning inference acceleration
abstract
We present a method for transparently identifying ineffectual computations during inference with Deep Learning models. Specifically, by decomposing multiplications down to the bit level, the amount of work needed by multiplications during inference can be potentially reduced by at least 40× across a wide selection of neural networks (8b and 16b). This method produces numerically identical results and does not affect overall accuracy. We present Laconic, a hardware accelerator that implements this approach to boost energy efficiency for inference with Deep Learning Networks. Laconic judiciously gives up some of the work reduction potential to yield a low-cost, simple, and energy efficient design that outperforms other state-of-the-art accelerators: an optimized DaDianNao-like design [13], Eyeriss [15], SCNN [71], Pragmatic [3], and BitFusion [83]. We study 16b, 8b, and 1b/2b fixed-point quantized models.
Sayeh Sharify, Alberto Delmas Lascorz, Mostafa Mahmoud, Milos Nikolic 0002, Kevin Siu, Dylan Malone Stuart, Zissis Poulos, Andreas Moshovos
ISCA2
2019 ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning
abstract
We show that selecting a data width for all values in Deep Neural Networks, quantized or not and even if that width is different per layer, amounts to worst-case design. Much shorter data widths can be used if we target the common case by adjusting the data type width at a much finer granularity. We propose ShapeShifter, where we group weights and activations and encode them using a width specific to each group and where typical group sizes vary from 16 to 256 values. The per group widths are selected statically for the weights and dynamically by hardware for the activations. We present two applications of ShapeShifter. In the first, that is applicable to any system, ShapeShifter reduces off- and on-chip storage and communication. This ShapeShifter-based memory compression is simple and low cost yet reduces off-chip traffic to 33% and 36% for 8-bit and 16-bit models respectively. This makes it possible to sustain higher performance for a given off-chip memory interface while also boosting energy efficiency. In the second application, we show how ShapeShifter can be implemented as a surgical extension over designs that exploit variable precision in time.
Alberto Delmas Lascorz, Sayeh Sharify, Isak Edo Vivancos, Dylan Malone Stuart, Omar Mohamed Awad, Patrick Judd, Mostafa Mahmoud, Milos Nikolic 0002, Kevin Siu, Zissis Poulos, Andreas Moshovos
MICRO1
2018 Loom: exploiting weight and activation precisions to accelerate convolutional neural networks
abstract
Loom (LM), a hardware inference accelerator for Convolutional Neural Networks (CNNs) is presented. In LM every bit of data precision that can be saved translates to proportional performance gains. For both weights and activations LM exploits profile-derived per layer precisions. However, at runtime LM further trims activation precisions at a much smaller than a layer granularity. On average, across several image classification CNNs and for a configuration that can perform the equivalent of 128 16b × 16b multiply-accumulate operations per cycle LM outperforms a state-of-the-art bit-parallel accelerator [3] by 3.19× without any loss in accuracy while being 2.59× more energy efficient. LM can trade-off accuracy for additional improvements in execution performance and energy efficiency and compares favorably to an accelerator that targeted only activation precisions.
Sayeh Sharify, Alberto Delmas Lascorz, Kevin Siu, Patrick Judd, Andreas Moshovos
DAC2
2017 Bit-pragmatic deep neural network computing
abstract
Deep Neural Networks expose a high degree of parallelism, making them amenable to highly data parallel architectures. However, data-parallel architectures often accept inefficiency in individual computations for the sake of overall efficiency. We show that on average, activation values of convolutional layers during inference in modern Deep Convolutional Neural Networks (CNNs) contain 92% zero bits. Processing these zero bits entails ineffectual computations that could be skipped. We propose Pragmatic (PRA), a massively data-parallel architecture that eliminates most of the ineffectual computations on-the-fly, improving performance and energy efficiency compared to state-of-the-art high-performance accelerators [5]. The idea behind PRA is deceptively simple: use serial-parallel shift-and-add multiplication while skipping the zero bits of the serial input. However, a straightforward implementation based on shift-and-add multiplication yields unacceptable area, power and memory access overheads compared to a conventional bit-parallel design. PRA incorporates a set of design decisions to yield a practical, area and energy efficient design.
Jorge Albericio, Alberto Delmas Lascorz, Patrick Judd, Sayeh Sharify, Gerard O'Leary, Roman Genov, Andreas Moshovos
MICRO2
2017 IDEAL: image denoising accelerator
abstract
Computational imaging pipelines (CIPs) convert the raw output of imaging sensors into the high-quality images that are used for further processing. This work studies how Block-Matching and 3D filtering (BM3D), a state-of-the-art denoising algorithm can be implemented to meet the demands of user-interactive (UI) applications. Denoising is the most computationally demanding stage of a CIP taking more than 95% of time on a highly-optimized software implementation [29]. We analyze the performance and energy consumption of optimized software implementations on three commodity platforms and find that their performance is inadequate.
Mostafa Mahmoud, Bojian Zheng, Alberto Delmas Lascorz, Felix Heide, Jonathan Assouline, Paul Boucher, Emmanuel Onzon, Andreas Moshovos
MICRO3
2014 An improved FPGA-based specific processor for Blokus Duo
abstract
This article presents a hardware design of a specific processor for Blokus Duo game. This design is an evolution of our previous work presented in the ICFPT'13 Design Competition. In order to improve its performance we have designed parallel hardware blocks to speed up the most time-consuming tasks, and included additional techniques to reduce the search space. As a consequence we can process a board six times faster than in our previous version and we prune the game-tree much more efficiently.
Javier Olivito, Alberto Delmas Lascorz, Javier Resano
FPT2