Vidushi Goyal

dblp:161/3106 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
4since 2021 · last 2024
0000-0002-5008-3049ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Duet: A Collaborative User Driven Recommendation System for Edge Devices
abstract
Recommendation systems are the backbone for numerous user applications on edge devices. However, the compute and memory-intensive nature of recommendation models renders them unsuitable for edge devices. Nevertheless, by decoupling the model fraction related to user history (e.g., past visited pages, liked posts) and user attributes (such as age, gender), we can offload partial recommendation models onto local edge devices. Hence, we present Duet, a novel collaborative edge-cloud recommendation system that intelligently decomposes the recommendation model into two smaller models - user and item models - that execute simultaneously on the edge device and cloud before coming together to deliver final recommendations. Further, we propose a lightweight Duet architecture to support user models on resource-constrained edge devices. Duet reduces the average latency by 6.4X and improves energy efficiency by 4.6X across five recommendation models.
Vidushi Goyal, Valeria Bertacco, Reetuparna Das
DAC1
2022 Hardware-friendly User-specific Machine Learning for Edge Devices
abstract
Machine learning (ML) on resource-constrained edge devices is expensive and often requires offloading computation to the cloud, which may compromise the privacy of user data. In contrast, the type of data processed at edge devices is user-specific and limited to a few inference classes. In this work, we explore building smaller, user-specific machine learning models, rather than utilizing a generic, compute-intensive machine learning model that caters to a diverse range of users. We first present a hardware-friendly, lightweight pruning technique to create user-specific models directly on mobile platforms, while simultaneously executing inferences. The proposed technique leverages compute sharing between pruning and inference, customizes the backward pass of training, and chooses a pruning granularity for efficient processing on edge. We then propose architectural support to prune user-specific models on a systolic edge ML inference accelerator. We demonstrate that user-specific models provide a speedup of 2.9× and 2.3× on the mobile CPUs for the ResNet-50 and Inception-V3 models.
Vidushi Goyal, Reetuparna Das, Valeria Bertacco
ACM Trans. Embed. Comput. Syst.1
2021 MyML: User-Driven Machine Learning
abstract
Machine learning (ML) on resource-constrained edge devices is expensive and often requires offloading computation to the cloud, which may compromise the privacy of user data. In contrast, the type of data processed at edge devices is user specific and limited to few inference classes. In this work, we explore the opportunity of building smaller, user-specific machine learning models, rather than utilizing a generic, compute-intensive machine learning model that caters to a diverse range of users. We first present a hardware-friendly, light-weight pruning technique to create user-specific models directly on mobile platforms, while simultaneously executing inferences. The proposed technique leverages compute sharing between pruning and inference, customizes the retraining backward-pass and chooses a pruning granularity for efficient processing on edge. We then propose architectural support to prune user-specific models on a systolic edge ML inference accelerator. We demonstrate that user-specific models provide a speedup of $2.3\times$ over the generic model on mobile CPUs.
Vidushi Goyal, Valeria Bertacco, Reetuparna Das
DAC1
2021 Compute-Capable Block RAMs for Efficient Deep Learning Acceleration on FPGAs
abstract
The density of FPGA on-chip memory has been continuously increasing with modern FPGAs having thousands of block RAMs (BRAMs) distributed across their reconfigurable fabric. These distributed BRAMs can provide a tremendous amount of on-chip bandwidth for efficient acceleration of data-intensive applications. In this work, we propose enhancing the ubiquitous FPGA BRAMs with in-memory compute-capabilities. As a result, BRAMs can act as normal storage units or their bitlines can be re-purposed as SIMD lanes executing bit-serial arithmetic operations. Our proposed architectural change results in 1.6× and 2.3× increase in the peak multiply-accumulate throughput of a large Stratix 10 FPGA, at a minimal cost of only 1.8% increase in the FPGA die size and no change to the BRAM's interface to the programmable routing. Then, we present RIMA, a reconfigurable in-memory accelerator architecture for deep learning (DL) inference. RIMA exploits the proposed compute-capable BRAMs and the FPGA's reconfigurability to achieve 1.25× and 3× higher performance compared to the state-of-the-art Brainwave DL soft processor for 8-bit integer and block floating-point precisions, respectively. In addition, RIMA implemented on a Stratix 10 FPGA enhanced with compute-capable BRAMs can achieve an order of magnitude higher performance compared to a same-generation GPU.
Xiaowei Wang 0005, Vidushi Goyal, Jiecao Yu, Valeria Bertacco, Andrew Boutros, Eriko Nurvitadhi, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das
FCCM2
2020 Seesaw: End-to-end Dynamic Sensing for IoT using Machine Learning
abstract
IoT edge devices’ small form factor places tight constraints on their battery life. These devices are often equipped with multiple sensors, a few of them responsible for most of the energy usage. Naïvely lowering the sensing rate of these power-hungry sensors reduces energy consumption, but also degrades application’s output quality. In this work, we observe that it is possible to leverage low-power sensors in a system to predict the impact of throttling a power-intensive sensor on application’s output accuracy. We thus propose Seesaw, an end-to-end ML-based solution that automatically identifies correlations between power-intensive and lightweight sensors without human expertise. Further, Seesaw deploys a low-overhead decision tree predictor to determine the optimal sensing rates for power-intensive sensors, thus avoiding significant quality degradation. We show that Seesaw improves battery life for (1) video recording on mountable video cameras and (2) route tracking on fitness trackers by 32% and 66%, respectively, without significant accuracy loss.
Vidushi Goyal, Valeria Bertacco, Reetuparna Das
DAC1
2020 Neksus: An Interconnect for Heterogeneous System-In-Package Architectures
abstract
In the embedded systems industry today, skyrocketing design and manufacturing costs of Systems-on-Chip (SoCs) are key limiting factors for growth. Emerging 2.5D-based System-In-Package (SiP) architectures show potential to lower these costs by enabling the reuse of hard core units and providing higher manufacturing yields due to small chiplet sizes.In this paper, we present Neksus, a novel architecture designed to lower SiP manufacturing costs, support modular "plug-and-play" chiplet integration, and leverage the unique properties of interposers. Key to Neksus is a new dedicated interconnect chiplet that addresses the limitations of SiP packaging technology by leveraging direct communication over a mini-chain IP-connection topology. In addition to satisfying SiP technology constraints, because our mini-chain design provides high-bandwidth IP-to-IP communication, it is particularly well-suited for bandwidth-intensive mobile applications. Our evaluation shows Neksus provides up to 28% performance improvement and 31% energy savings over recent SiP architecture.
Vidushi Goyal, Xiaowei Wang 0005, Valeria Bertacco, Reetuparna Das
IPDPS1