Janak Sharda

dblp:244/2428 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-1438-2439ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
abstract
As Large Language Models (LLMs) continue to evolve, Mixture of Experts (MoE) architecture has emerged as a prevailing design for achieving state-of-the-art performance across a wide range of tasks.MoE models use sparse gating to activate only a handful of expert sub-networks per input, achieving billion-parameter capacity with inference costs akin to much smaller models.However, such models often pose challenges for hardware deployment due to the massive data volume introduced by the MoE layers.To address the challenges of serving MoE models, we propose Stratum, a system-hardware co-design approach that combines the novel memory technology Monolithic 3D-Stackable DRAM (Mono3D DRAM), near-memory processing (NMP), and GPU acceleration.The logic and Mono3D DRAM dies are connected through hybrid bonding, whereas the Mono3D DRAM stack and GPU are interconnected via silicon interposer.Mono3D DRAM offers higher internal bandwidth than HBM thanks to the dense vertical interconnect pitch enabled by its monolithic structure, which supports implementations of higher-performance near-memory processing.Furthermore, we tackle the latency differences introduced by aggressive vertical scaling of Mono3D DRAM along the 𝑧-dimension by constructing internal memory tiers and assigning data across layers based on * Equal contribution
Yue Pan 0009, Zihan Xia 0002, Po-Kai Hsu, Lanxiang Hu, Hyungyo Kim, Janak Sharda, Minxuan Zhou, Nam Sung Kim, Shimeng Yu, Tajana Rosing, Mingu Kang
MICRO6
2024 Thermally Constrained Codesign of Heterogeneous 3-D Integration of Compute-in-Memory, Digital ML Accelerator, and RISC-V Cores for Mixed ML and Non-ML Workloads
abstract
Heterogeneous 3-D (H3D) integration not only reduces the chip form factor and fabrication cost but also allows the merging of diverse compute paradigms that suit different applications. This is especially attractive when modern algorithms, such as the augmented reality/virtual reality (AR/VR) workloads, consist of mixed machine learning (ML) and non-ML workloads. To date, codesign that considers the thermal, latency, and power constraints of H3D hardware is largely unexplored. In this work, a thermally aware framework for H3D hardware design is developed to evaluate the thermal, latency, and power trade-offs for a heterogeneous system with compute-in-memory (CIM), digital ML cores, and RISC-V cores. The framework solves for runtime tunable operating points described as the optimal speedup factor, the number of activated RISC-V cores, the cooling coefficient, and the activity rate based on user-defined criteria, achieving up to 135 TOPS and 215 TOPS/W under$74~^{\circ }$C for the AR/VR workloads.
Yuan-Chun Luo, Anni Lu, Janak Sharda, Moritz Scherer 0001, Jorge Gomez 0001, Syed Shakib Sarwar, Ziyun Li 0001, Reid Frederick Pinkham, Barbara De Salvo, Shimeng Yu
IEEE Trans. Very Large Scale Integr. Syst.3
2023 Temporal Frame Filtering for Autonomous Driving Using 3D-Stacked Global Shutter CIS With IWO Buffer Memory and Near-Pixel Compute
abstract
With the advancement of deep learning to solve autonomous driving problems, the computation and memory requirements have been growing rapidly. Near-pixel compute-based CMOS image sensors (CIS) have been investigated as a potential candidate to perform the initial computations of workloads close to the pixel and reduce data movement. In this work, we design a near-pixel compute CIS capable of implementing a temporal frame filtering network, which rejects redundant image frames targeting autonomous driving applications. To improve performance and avoid image distortion, 3D-stacked global shutter CIS is proposed. This architecture integrates photodiodes with memory and compute units using Cu-Cu hybrid bonding. We propose to use back-end-of-line (BEOL) compatible Tungsten-doped Indium Oxide Transistors (IWO FETs) based embedded DRAM as buffer memory to achieve refresh-free storage and high bandwidth connections between various components. Near-pixel compute circuit is optimized by including sparsity-aware adder tree and using NOR gates as data buffers. The two-tier system comprises photodiodes on tier-1 in 40 nm node, and near-pixel compute and buffer memory on tier-2 in 22 nm node. We perform simulations in Cadence, obtaining an energy efficiency of 65 TOPS/W and a compute density of 1.04 TOPS/mm2 for$8\times8\text{b}$MAC, with a total latency of 1.15 ms/frame.
Janak Sharda, Wantong Li 0002, Qiucheng Wu, Shiyu Chang, Shimeng Yu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 A Crossbar Array of Analog-Digital-Hybrid Volatile Memory Synapse Cells for Energy-Efficient On-Chip Learning
abstract
Conventional-silicon-transistor-based Volatile Memory (VM) synapse has been proposed as an alternative to Non Volatile Memory (NVM) synapse in crossbar-array-based neuromorphic/ in-memory-computing systems. Here, through SPICE simulations, we have designed an analog-digital-hybrid Volatile Memory Synapse Cell (VMSC) for such a crossbar array of VM synapses. In our VMSC, the transistor synapse stores nearly analog values of weight. But the other transistors, which carry out the weight update for the transistor synapse, are designed following the principle of static CMOS logic (digital), making our design energy-efficient. Through system-level study, we report classification accuracy, speed, and energy consumption for on-chip learning on the VMSC-based crossbar designed here, using popular machine learning data sets. We show that despite a low value of capacitance of our MOSFET synapses (low area- footprint hence), the weights are retained in them long enough for our VMSC-based crossbar to exhibit comparable accuracy as a NVM-synapse-based crossbar.
Janak Sharda, Ritvik Sharma, Debanjan Bhowmik
ISCAS1