VLDB 2026 Research / reviewers in the wild / expert
Alberto Mannari
dblp:298/8871
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Eliminating Redundancy: Ultra-compact Code Generation for Programmable Dataflow AcceleratorsabstractModern AI accelerators adopt dataflow architectures to achieve both high peak throughput (TOPS) and energy efficiency (TOPS/W). These designs feature wide datapaths and hierarchical scratchpad memories that supply dense compute arrays with high-bandwidth data access and extensive operand reuse. Complementing the compute–memory subsystem is a lightweight control path that orchestrates data movement, program loading, and register initialization. To reduce energy and area overheads, conventional processor features—such as instruction caches, execution stacks, and branch speculation—are deliberately omitted. While this streamlined design maximizes efficiency, it shifts a critical responsibility onto the compiler: transforming complex kernels into highly compact instruction streams that must fit entirely within the limited instruction buffers (IBUFFs) of the accelerator’s programmable units.In this paper, we introduce two novel compiler transformations—Loop Absorption (LA) and Loop Index Set Merging (LISM) for ultra compact code generation. Loop Absorption merges isomorphic sibling operations into a single loop body, while LISM unifies adjacent loops with similar bodies into a unified iteration space. Together, these complementary techniques eliminate redundant code patterns and produce compact hierarchical loop nests. We implement LA and LISM in the IBM Spyre compiler and evaluate them on diverse deep learning workloads including ResNet-50, Inception-v3, SSD, and BERT-Large. Across these models, our combined approach achieves a geometric mean compression of 1.48× over the baseline, enabling layers that previously exceeded IBUFF capacity to compile successfully. Prasanth Chatarasi, Alex Gatea, Bardia Mahjour, Alberto Mannari, Chris Bowler, Shubham Jain 0004, Masoud Ataei Jaliseh, Nicole Khoun, Vijayalakshmi Srinivasan, Swagath Venkataramani |
CGO | 5 |
| 2026 | Enabling Spill-Free Compilation via Affine-Based Live Range Reduction OptimizationabstractAI Accelerators employ dataflow architectures to achieve impressive peak compute performance (TOPS) and processing efficiencies (TOPS/W). Typically, dataflow architectures use wide data-paths to connect off-chip memory to dense compute arrays (via hierarchy of on-chip memories/vector register files) for efficient data movement with reuse, as well as compute. Such architectures often possess an independent lightweight control-path for loading programs and initializing registers, and lack traditional architectural features like instruction cache and execution stacks. This poses a unique challenge to compiler requiring program generation of complex compute kernels to fit within an instruction buffer and allocating a limited set of scalar registers without support to spill to memory.This paper contributes a significant step towards spill-free compilation and proposes a Live range reduction optimization based on Affine expression propagation analysis. Our solution performs a global, compiler-directed analysis to model variable values as affine expressions of in-scope variables, enabling safe symbolic re-materialization of values at their use-sites leveraging near- by variables without introducing new operations. This shortens variable lifetimes, while significantly reducing register pressure without incurring program binary and execution overhead. The static nature and regular memory access patterns of AI applications make them well-suited for the proposed optimization. We demonstrate the effectiveness of the technique in the context of IBM Spyre accelerator and its compiler. Our results over a range of AI workloads spanning transformer and CNN models demonstrate spill-free code generation, with most of the workloads requiring less than 50% of the available registers. Prasanth Chatarasi, Alex Gatea, Wei Wang 0333, Chris Bowler, Shubham Jain 0004, Masoud Ataei Jaliseh, Nicole Khoun, Alberto Mannari, Bardia Mahjour, Vijayalakshmi Srinivasan, Swagath Venkataramani |
CGO | 8 |
| 2025 | Live Demonstration: Improving efficiency of speech recognition with neuro-inspired units on AIU SpyreabstractThis demonstration implements efficient speech recognition through the use of biologically-inspired units implemented on the recently introduced Artificial Intelligence Unit (AIU Spyre). Speech recognition models approach human-level accuracy, but are significantly more power-hungry compared to the human brain. Efficient biological processes can be incorporated into deep learning models substantially reducing the computational cost and inference time. Simultaneously, novel chips, such as the AIU Spyre are designed from the ground up for energy-efficient execution of artificial neural networks. We demonstrate the scalability and efficiency of our neuro-inspired speech model through an implementation on the AIU Spyre. Yannick Schnider, Thomas Ortner, Stanislaw Wozniak, Alberto Mannari, Angeliki Pantazi |
ISCAS | 4 |
| 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and InferenceabstractThe growing prevalence and computational demands of Artificial Intelligence (AI) workloads has led to widespread use of hardware accelerators in their execution. Scaling the performance of AI accelerators across generations is pivotal to their success in commercial deployments. The intrinsic error-resilient nature of AI workloads present a unique opportunity for performance/energy improvement through precision scaling. Motivated by the recent algorithmic advances in precision scaling for inference and training, we designed RaPiD1, a 4-core AI accelerator chip supporting a spectrum of precisions, namely, 16 and 8-bit floating-point and 4 and 2-bit fixed-point. The 36mm2RaPiD chip fabricated in 7nm EUV technology delivers a peak 3.5 TFLOPS/W in HFP8 mode and 16.5 TOPS/W in INT4 mode at nominal voltage. Using a performance model calibrated to within 1% of the measurement results, we evaluated DNN inference using 4-bit fixed-point representation for a 4-core 1 RaPiD chip system and DNN training using 8-bit floating point representation for a 768 TFLOPs AI system comprising 4 32-core RaPiD chips. Our results show INT4 inference for batch size of 1 achieves 3 - 13.5 (average 7) TOPS/W and FP8 training for a mini-batch of 512 achieves a sustained 102 - 588 (average 203) TFLOPS across a wide range of applications. Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang 0333, Sanchari Sen, Ankur Agrawal, Monodeep Kar, Shubham Jain 0004, Alberto Mannari, Hoang Tran, Eri Ogawa, Kazuaki Ishizaki, Hiroshi Inoue, Marcel Schaal, Mauricio J. Serrano, Jungwook Choi, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Allison Allain, James Bonanno, Nianzheng Cao, Robert Casatuta, Matthew Cohen, Bruce M. Fleischer, Michael Guillorn, Howard Haynie, Jinwook Jung, Mingu Kang, Kyu-Hyoun Kim, Siyu Koswatta, Sae Kyu Lee, Martin Lutz, Silvia M. Müller, Jinwook Oh, Ashish Ranjan 0001, Zhibin Ren, Scot Rider, Kerstin Schelm, Michael Scheuermann, Joel Silberman, Vidhi Zalani, Xin Zhang 0025, Ching Zhou, Matthew M. Ziegler, Vinay Shah, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Leland Chang, Kailash Gopalakrishnan |
ISCA | 9 |