VLDB 2026 Research / reviewers in the wild / expert
Isak Edo Vivancos
dblp:248/8951 · also Isak Edo
· DBLP profile ↗
7ranked-venue papers
1as first author
2since 2021 · last 2024
0009-0004-1189-2259ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Hardware accelerators and domain-specific architectures · 57% Memory systems · 30% Performance modeling and evaluation · 8% | |
| Artificial intelligence
3 papers |
Efficient and distributed learning · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
2.2 | 5 | 2021 | FPRaker: A Processing Element For Accelerating Neural Network Training · MICRO 2021 GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference · MICRO 2020 TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training · MICRO 2020 |
Memory systems
memory compression |
0.6 | 3 | 2020 | ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019 GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference · MICRO 2020 Late Breaking Results: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick · DAC 2020 |
Machine learning › Efficient and distributed learning
low-precision training |
0.6 | 2 | 2021 | FPRaker: A Processing Element For Accelerating Neural Network Training · MICRO 2021 TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training · MICRO 2020 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2020 | TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training · MICRO 2020 |
Machine learning › Efficient and distributed learning › model compression › sparsity
sparsity exploitation |
0.4 | 1 | 2020 | TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training · MICRO 2020 |
Memory systems
memory controller |
0.4 | 1 | 2020 | Mocktails: Capturing the Memory Behaviour of Proprietary Mobile Architectures · ISCA 2020 |
Memory systems › memory controller
memory scheduling |
0.4 | 1 | 2020 | Mocktails: Capturing the Memory Behaviour of Proprietary Mobile Architectures · ISCA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
quantized inference |
0.4 | 1 | 2020 | GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference · MICRO 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN training accelerator
sparse DNN training accelerator |
0.4 | 1 | 2020 | TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training · MICRO 2020 |
Performance modeling and evaluation
workload characterization |
0.4 | 1 | 2020 | Mocktails: Capturing the Memory Behaviour of Proprietary Mobile Architectures · ISCA 2020 |
Storage systems › data compression
delta compression |
0.1 | 1 | 2021 | FPRaker: A Processing Element For Accelerating Neural Network Training · MICRO 2021 |
Memory systems
memory bandwidth |
0.1 | 1 | 2021 | FPRaker: A Processing Element For Accelerating Neural Network Training · MICRO 2021 |
Embedded and real-time systems › embedded hardware platform
heterogeneous system-on-chip |
0.1 | 1 | 2020 | Mocktails: Capturing the Memory Behaviour of Proprietary Mobile Architectures · ISCA 2020 |
Hardware accelerators and domain-specific architectures › model compression
weight compression |
0.1 | 1 | 2020 | GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference · MICRO 2020 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.1 | 1 | 2019 | ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019 |
Methods — techniques the papers use, named apart from their topics
quantization · 1.4pruning · 1.0floating-point multiply-accumulate · 1.0delta encoding · 1.0dynamic width selection · 0.8synthetic trace generation · 0.4simulation · 0.4lossless compression · 0.4fixed-point quantization · 0.43-bit weight quantization · 0.4fine-grain data width encoding · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | BitPruning: Learning Bitlengths for Aggressive and Accurate QuantizationabstractBitPruning is a training method for minimizing inference bitlengths at any granularity while maintaining accuracy. BitPruning extends the meaning of fixed-point bitlenghts into the continuous domain by interpolating between the nearest two integers, enabling gradient descent to learn bitlengths together with other parameters. A novel regularizer penalizes large bitlength representations and can be modified to minimize other quantifiable criteria, such as number of operations or memory footprint. BitPruning learns thrifty representations while maintaining accuracy: With ImageNet, it produces an average per layer bitlength of 3.76 and 4.36 bits on ResNet18 and MobileNet V2 respectively, remaining within 0.5% of the base TOP-1 accuracy. Simple modifications of the BitPruning regularizer can be used to further reduce compute workload by up to 24%, as well as memory footprint in activation or weight-heavy tasks by up to 14% and 8% respectively. Milos Nikolic 0002, Ghouthi Boukli Hacene, Ciaran Bannon, Alberto Delmas Lascorz, Matthieu Courbariaux, Omar Mohamed Awad, Isak Edo Vivancos, Yoshua Bengio, Vincent Gripon, Andreas Moshovos |
ISCAS | 7 |
| 2021 | FPRaker: A Processing Element For Accelerating Neural Network TrainingabstractWe present FPRaker, a processing element for composing training accelerators. FPRaker processes several floating-point multiply-accumulation operations concurrently and accumulates their result into a higher precision accumulator. FPRaker boosts performance and energy efficiency during training by taking advantage of the values that naturally appear during training. It processes the significand of the operands of each multiply-accumulate as a series of signed powers of two. The conversion to this form is done on-the-fly. This exposes ineffectual work that can be skipped: values when encoded have few terms and some of them can be discarded as they would fall outside the range of the accumulator given the limited precision of floating-point. FPRaker also takes advantage of spatial correlation in values across channels and uses delta-encoding off-chip to reduce memory footprint and bandwidth. We demonstrate that FPRaker can be used to compose an accelerator for training and that it can improve performance and energy efficiency compared to using optimized bit-parallel floating-point units under iso-compute area constraints. We also demonstrate that FPRaker delivers additional benefits when training incorporates pruning and quantization. Finally, we show that FPRaker naturally amplifies performance with training methods that use a different precision per layer. Omar Mohamed Awad, Mostafa Mahmoud, Isak Edo Vivancos, Ali Hadi Zadeh, Ciaran Bannon, Anand Jayarajan, Gennady Pekhimenko, Andreas Moshovos |
MICRO | 3 |
| 2020 | Late Breaking Results: Building an On-Chip Deep Learning Memory Hierarchy Brick by BrickabstractData accesses between on- and off-chip memories account for a large fraction of overall energy consumption during inference with deep learning networks. We present Boveda, a lossless on-chip memory compression technique for neural networks operating on fixed-point values. Boveda reduces the datawidth used per block of values to be only as long as necessary: since most values are of small magnitude Boveda drastically reduces their footprint. Boveda can be used to increase the effective on-chip capacity, to reduce off-chip traffic, or to reduce the on-chip memory capacity needed to achieve a performance/energy target. Boveda reduces total model footprint to 53%. Isak Edo Vivancos, Sayeh Sharify, Milos Nikolic 0002, Ciaran Bannon, Mostafa Mahmoud, Alberto Delmas Lascorz, Andreas Moshovos |
DAC | 1 |
| 2020 | Mocktails: Capturing the Memory Behaviour of Proprietary Mobile ArchitecturesabstractComputation demands on mobile and edge devices are increasing dramatically. Mobile devices, such as smart phones, incorporate a large number of dedicated accelerators and fixed-function hardware blocks to deliver the required performance and power efficiency. Due to the heterogeneous nature of these devices, they feature vastly larger design spaces than traditional systems featuring only a CPU. Currently, academia struggles to fully evaluate such heterogeneous systems on chip due to the limited access and availability of proprietary workloads. To address these challenges, we propose Mocktails: a methodology to synthetically recreate the varying spatio-temporal memory access behaviour of proprietary heterogeneous compute devices. We focus on capturing the interspersed address streams of the workload and the burstiness of the injection process for proprietary compute devices commonly found in mobile systems. We evaluate Mocktails in simulation with proprietary memory traces of IP blocks. Mocktails accurately recreates the dynamic behaviour of memory access scheduling for memory controller metrics including read row hits (at most 7.3% error) and write row hits (at most 2.8% error). Architects can use Mocktails in their simulations as a substitute for a proprietary compute device, making the tool a useful conduit between industry and academia. Mario Badr, Carlo Delconte, Isak Edo Vivancos, Radhika Jagtap, Matteo Andreozzi, Natalie D. Enright Jerger |
ISCA | 3 |
| 2020 | TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network TrainingabstractTensorDash is a hardware-based technique that enables data-parallel MAC units to take advantage of sparsity in their input operand streams. When used to compose a hardware accelerator for deep learning, TensorDash can speedup the training process while also increasing energy efficiency. TensorDash combines a low-cost sparse input operand interconnect with an area-efficient hardware scheduler. The scheduler can effectively extract sparsity in the activations, the weights, and the gradients. Over a wide set of state-of-the-art models covering various applications, TensorDash accelerates the training process by 1.95× while being 1.5× more energy efficient when incorporated on top of a Tensorcore-based accelerator at less than 5% area overhead. TensorDash is datatype agnostic and we demonstrate it with IEEE standard mixed-precision floating-point units and a popular optimized for machine learning floating-point format (BFloat16). Mostafa Mahmoud, Isak Edo Vivancos, Ali Hadi Zadeh, Omar Mohamed Awad, Gennady Pekhimenko, Jorge Albericio, Andreas Moshovos |
MICRO | 2 |
| 2020 | GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient InferenceabstractAttention-based models have demonstrated remarkable success in various natural language understanding tasks. However, efficient execution remains a challenge for these models which are memory-bound due to their massive number of parameters. We present GOBO, a model quantization technique that compresses the vast majority (typically 99.9%) of the 32-bit floating-point parameters of state-of-the-art BERT models and their variants to 3 bits while maintaining their accuracy. Unlike other quantization methods, GOBO does not require fine-tuning nor retraining to compensate for the quantization error. We present two practical hardware applications of GOBO. In the first GOBO reduces memory storage and traffic and as a result inference latency and energy consumption. This GOBO memory compression mechanism is plug-in compatible with many architectures; we demonstrate it with the TPU, Eyeriss, and an architecture using Tensor Cores-like units. Second, we present a co-designed hardware architecture that also reduces computation. Uniquely, the GOBO architecture maintains most of the weights in 3b even during computation, a property that: (i) makes the processing elements area efficient, allowing us to pack more compute power per unit area, (ii) replaces most multiply-accumulations with additions, and (iii) reduces the off-chip traffic by amplifying on-chip memory capacity. Ali Hadi Zadeh, Isak Edo Vivancos, Omar Mohamed Awad, Andreas Moshovos |
MICRO | 2 |
| 2019 | ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep LearningabstractWe show that selecting a data width for all values in Deep Neural Networks, quantized or not and even if that width is different per layer, amounts to worst-case design. Much shorter data widths can be used if we target the common case by adjusting the data type width at a much finer granularity. We propose ShapeShifter, where we group weights and activations and encode them using a width specific to each group and where typical group sizes vary from 16 to 256 values. The per group widths are selected statically for the weights and dynamically by hardware for the activations. We present two applications of ShapeShifter. In the first, that is applicable to any system, ShapeShifter reduces off- and on-chip storage and communication. This ShapeShifter-based memory compression is simple and low cost yet reduces off-chip traffic to 33% and 36% for 8-bit and 16-bit models respectively. This makes it possible to sustain higher performance for a given off-chip memory interface while also boosting energy efficiency. In the second application, we show how ShapeShifter can be implemented as a surgical extension over designs that exploit variable precision in time. Alberto Delmas Lascorz, Sayeh Sharify, Isak Edo Vivancos, Dylan Malone Stuart, Omar Mohamed Awad, Patrick Judd, Mostafa Mahmoud, Milos Nikolic 0002, Kevin Siu, Zissis Poulos, Andreas Moshovos |
MICRO | 3 |