VLDB 2026 Research / reviewers in the wild / expert
Vikram Jain
dblp:257/5287
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0002-1267-1683ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning AccelerationabstractBit-serial computation facilitates bit-wise sequential data processing, offering numerous benefits, such as a reduced area footprint and dynamically-adaptive computational precision. It has emerged as a prominent approach, particularly in leveraging bit-level sparsity in Deep Neural Networks (DNNs). However, existing bit-serial accelerators exploit bit-level sparsity to reduce computations by skipping zero bits, but they suffer from inefficient memory accesses due to the irregular indices of the non-zero bits. As memory accesses typically are the dominant contributor to DNN accelerator performance, this paper introduces a novel computing approach called “bit-column-serial” and a compatible architecture design named “BitWave.” BitWave harnesses the advantages of the “bit-column-serial” approach, leveraging structured bit-level sparsity in combination with dynamic dataflow techniques. This achieves a reduction in computations and memory footprints through redundant computation skipping and weight compression. BitWave is able to mitigate the performance drop or the need for retraining that is typically associated with sparsity-enhancing techniques using a post-training optimization involving selected weight bit-flips. Empirical studies conducted on four deep-learning benchmarks demonstrate the achievements of BitWave: (1) Maximally realize 13.25x higher speedup, 7.71 x efficiency compared to state-of-the-art sparsity-aware accelerators. (2) Occupying 1.138 mm2area and consuming 17.56 mW power in 16nm FinFet process node. Man Shi, Vikram Jain, Antony Joseph, Maurice Meijer, Marian Verhelst |
HPCA | 2 |
| 2024 | Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the EdgeabstractHybrid vision transformers combine the elements of conventional neural networks (NN) and vision transformers (ViT) to enable lightweight and accurate detection. However, several challenges remain for their efficient deployment on resource-constrained edge devices. The hybrid models suffer from a widely diverse set of NN layer types and large intermediate data tensors, hampering efficient hardware acceleration. To enable their execution at the edge, this paper proposes innovations across the hardware-scheduling stack: a.) At the lowest level, a configurable PE array supports all hybrid ViT layer types; b.) temporal loop re-ordering within one layer, enabling hardware support for normalization and softmax layers, minimizing on-chip data transfers; c.) further scheduling optimization employs layer fusion across inverted bottleneck layers to drastically reduce off-chip memory transfers. The resulting accelerator is implemented in 28nm CMOS, achieving a peak energy efficiency of 1.39 TOPS/W at 25.6 GMACs/s. Joren Dumoulin, Pouya Houshmand, Vikram Jain, Marian Verhelst |
ISCAS | 3 |
| 2024 | Design Approach for Die-to-Die Interfaces to Enable Energy-Efficient Chiplet SystemsabstractHeterogeneous chiplet integration and advanced packaging have given a new lease to scaling of compute in a post-Moore era. A critical aspect of designing chiplet systems is die-to-die interfaces for aggregation of smaller disaggregated chiplets. In recent years, interconnects and packaging have made a huge leap forward, thereby, enabling high bandwidth and energy-efficient parallel die-to-die (D2D) interfaces. Instead of bespoke solutions, building a standardized D2D interface provides a mechanism for interoperability between heterogeneous chiplets and facilitates low power and energy-efficient design by prescribing implementation strategy. In this paper, we discuss some of the recent efforts in the standardization of die-to-die interfaces. Starting with an overview of Advanced Interface Bus (AIB) PHY, we emphasize its simplicity and high energy efficiency. Followed by a case study that demonstrates an effective chiplet integration employing AIB. A more recent open standard for chiplet integration is the Universal Chiplet Interconnect Express (UCIe). We discuss the key aspects of UCIe, including its electrical and packaging characteristics, as well as low-power features that lead to a 10x power reduction compared to typical off-package I/O. Finally, we discuss the UCIe-lite controller, an effort to democratize the chiplet infrastructure by providing a simplified open-source RTL generator of the D2D interface. The generator is highly parameterizable and lightweight, enabling chiplet systems for energy-efficient applications. Vikram Jain, Wei Tang 0010, Zuoguo Wu, Viansa Schmulbach, Sophia Shao, Zhengya Zhang, Borivoje Nikolic |
ISLPED | 1 |
| 2023 | PATRONoC: Parallel AXI Transport Reducing Overhead for Networks-on-Chip targeting Multi-Accelerator DNN Platforms at the EdgeabstractEmerging deep neural network (DNN) applications require high-performance multi-core hardware acceleration with large data bursts. Classical network-on-chips (NoCs) use serial packet-based protocols suffering from significant protocol translation overheads towards the endpoints. This paper proposes PATRONoC, an open-source fully AXI-compliant NoC fabric to better address the specific needs of multi-core DNN computing platforms. Evaluation of PATRONoC in a 2D-mesh topology shows 34 % higher area efficiency compared to a state-of-the-art classical NoC at 1 GHz. PATRONoC’s throughput outperforms a baseline NoC by 2-8× on uniform random traffic and provides a high aggregated throughput of up to 350 GiB/s on synthetic and DNN workload traffic. Vikram Jain, Matheus A. Cavalcante, Nazareno Bruschi, Michael Rogenmoser, Thomas Benz, Andreas Kurth, Davide Rossi 0001, Luca Benini, Marian Verhelst |
DAC | 1 |
| 2023 | PetaOps/W edge-AI $\mu$ Processors: Myth or reality?abstractWith the rise of deep learning (DL), our world braces for artificial intelligence (AI) in every edge device, creating an urgent need for edge-AI SoCs. This SoC hardware needs to support high throughput, reliable and secure AI processing at ultra-low power (ULP), with a very short time to market. With its strong legacy in edge solutions and open processing platforms, the EU is well-positioned to become a leader in this SoC market. However, this requires AI edge processing to become at least 100 times more energy-efficient, while offering sufficient flexibility and scalability to deal with AI as a fast-moving target. Since the design space of these complex SoCs is huge, advanced tooling is needed to make their design tractable. The CONVOLVE project (currently in Inital stage) addresses these roadblocks. It takes a holistic approach with innovations at all levels of the design hierarchy. Starting with an overview of SOTA DL processing support and our project methodology, this paper presents 8 important design choices largely impacting the energy efficiency and flexibility of DL hardware. Finding good solutions is key to making smart-edge computing a reality. Manil Dev Gomony, Floran de Putter, Anteneh Gebregiorgis, Gianna Paulin, Linyan Mei, Vikram Jain, Said Hamdioui, Victor Sanchez, Tobias Grosser, Marc Geilen, Marian Verhelst, Friedemann Zenke, Frank K. Gürkaynak, Barry de Bruin, Sander Stuijk, Simon Davidson, Sayandip De, Mounir Ghogho, Alexandra Jimborean, Sherif Eissa, Luca Benini, Dimitrios Soudris, Rajendra Bishnoi, Sam Ainsworth 0001, Federico Corradi, Ouassim Karrakchou, Tim Güneysu, Henk Corporaal |
DATE | 6 |
| 2021 | ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN AcceleratorsabstractBuilding efficient embedded deep learning systems requires a tight co-design between DNN algorithms, hardware, and algorithm-to-hardware mapping, a.k.a. dataflow. However, owing to the large joint design space, finding an optimal solution through physical implementation becomes infeasible. To tackle this problem, several design space exploration (DSE) frameworks have emerged recently, yet they either suffer from long runtimes or a limited exploration space. This article introduces ZigZag, a rapid DSE framework for DNN accelerator architecture and mapping. ZigZag extends the common DSE with uneven mapping opportunities and smart mapping search strategies. Uneven mapping decouples operands (W/I/O), memory hierarchy, and mappings (temporal/spatial), opening up a whole new space for DSE, and thus better design points are found by ZigZag compared to other SotAs. For this, ZigZag uses an enhanced nested-for-loop format as a uniform representation to integrate algorithm, accelerator, and algorithm-to-accelerator mapping. ZigZag consists of three key components: 1) an analytical energy-performance-area Hardware Cost Estimator, 2) two Mapping Search Engines that support spatial and temporal even/uneven mapping search, and 3) an Architecture Generator that auto-explores the wide memory hierarchy design space. Benchmarking experiments against published works, in-house accelerator, and existing DSE frameworks, together with three case studies, show the reliability and capability of ZigZag. Up to 64 percent more energy-efficient solutions are found compared to other SotAs, due to ZigZag's uneven mapping capabilities. Linyan Mei, Pouya Houshmand, Vikram Jain, Juan Sebastian Piedrahita Giraldo, Marian Verhelst |
IEEE Trans. Computers | 3 |
| 2021 | Variable-Rate VLSI Architecture for 400-Gb/s Hard-Decision Product DecoderabstractVariable-rate transceivers, which adapt to the conditions, will be central to energy-efficient communication. However, fiber-optic communication systems with high bit-rate requirements make design of flexible transceivers challenging, since additional circuits needed to orchestrate the flexibility will increase area and degrade speed. We propose a variable-rate VLSI architecture of a forward error correction (FEC) decoder based on hard-decision product codes. Variable shortening of component codes provides a mechanism by which code rate can be varied, the number of iterations offers a knob to control the coding gain, while a key-equation solver module that can swap between error-locator polynomial coefficients provides a means to change error correction capability. Our evaluations based on 28-nm netlists show that a variable-rate decoder implementation can offer a net coding gain (NCG) range of 9.96-10.38dB at a post-FEC bit-error rate of 10-15. The decoder achieves throughputs in excess of 400Gb/s, latencies below 53ns, and energy efficiencies of 1.14pJ/bit or less. While the area of the variable-rate decoder is 31% larger than a decoder with a fixed rate, the power dissipation is a mere 5% higher. The variable error correction capability feature increases the NCG range further, to above 10.5dB, but at a significant area cost. Vikram Jain, Christoffer Fougstedt, Per Larsson-Edefors |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Efficient Execution of Temporal Convolutional Networks for Embedded Keyword SpottingabstractRecently, the use of keyword spotting (KWS) has become prevalent in mobile devices. State-of-the-art deep learning algorithms such as temporal convolutional networks (TCNs) have been applied to this task achieving superior accuracy results. These models can, however, be mapped in multiple ways onto embedded devices, ranging from real-time streaming inference with or without computational sprinting to delayed batched inference. Although functionally equivalent, these deployment settings, however, strongly impacts average power consumption and latency of this real time task, hence requiring a thorough optimization. This work analyzes the challenges, benefits, and drawbacks of the different execution modes available for TCN-based KWS inference on dedicated hardware. With this objective, this research contributes to: 1) presenting a complete deep learning accelerator optimized for TCN inference; 2) evaluating the impact on performance and power of the different deployment options for TCN inference applied to KWS obtaining up to 8$\mu \text{W}$for real-time operation; and 3) optimizing real-time power consumption for KWS inference by exploiting the use of cascaded neural networks (NNs), achieving up to 35% additional power savings. Juan Sebastian Piedrahita Giraldo, Vikram Jain, Marian Verhelst |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |