EDBT 2026 Demo / reviewers in the wild / expert
Kshitij Bhardwaj
dblp:53/11157
· DBLP profile ↗
18ranked-venue papers
10as first author
8since 2021 · last 2024
0000-0001-7076-9251ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 9 first-author · 5 since 2021Software engineering, systems software and programming languages · 7 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MulBERRY: Enabling Bit-Error Robustness for Energy-Efficient Multi-Agent Autonomous SystemsabstractThe adoption of autonomous swarms, consisting of a multitude of unmanned aerial vehicles (UAVs), operating in a collaborative manner, has become prevalent in mainstream application domains for both military and civilian purposes. These swarms are expected to collaboratively carry out navigation tasks and employ complex reinforcement learning (RL) models within the stringent onboard size, weight, and power constraints. While techniques such as reducing onboard operating voltage can improve the energy efficiency of both computation and flight missions, they can lead to on-chip bit failures that are detrimental to mission safety and performance. Zishen Wan, Nandhini Chandramoorthy, Karthik Swaminathan, Kshitij Bhardwaj, Vijay Janapa Reddi, Arijit Raychowdhury |
ASPLOS (2) | 5 |
| 2024 | Designing an Energy-Efficient Fully-Asynchronous Deep Learning Convolution EngineabstractIn the face of exponential growth in semiconductor energy usage, there is a significant push towards highly energy-efficient microelectronics design. While the traditional circuit designs typically employ clocks to synchronize the computing operations, these circuits incur significant performance and energy overheads due to their data-independent worst-case operation and complex clock tree networks. In this paper, we explore asynchronous or clockless techniques where clocks are replaced by request, acknowledge handshaking signals. To quantify the potential energy and performance gains of asynchronous logic, we design a highly energy -efficient asynchronous deep learning convolution engine, which uses 87 % of total DL accelerator energy. Our asynchronous design shows 5.06x lower energy and 5.09 x lower delay than the synchronous one. Mattia Vezzoli, Lukas Nel, Kshitij Bhardwaj, Rajit Manohar, Maya B. Gokhale |
DATE | 3 |
| 2024 | Nuclear Fusion Diamond Polishing DatasetabstractIn the Inertial Confinement Fusion (ICF) process, roughly a 2mm spherical shell made of high-density carbon is used as a target for laser beams, which compress and heat it to energy levels needed for high fusion yield in nuclear fusion. These shells are polished meticulously to meet the standards for a fusion shot. However, the polishing of these shells involves multiple stages, with each stage taking several hours. To make sure that the polishing process is advancing in the right direction, we are able to measure the shell surface roughness. This measurement, however, is very labor-intensive, time-consuming, and requires a human operator. To help improve the polishing process we have released the first dataset to the public that consists of raw vibration signals with the corresponding polishing surface roughness changes. We show that this dataset can be used with a variety of neural network based methods for prediction of the change of polishing surface roughness, hence eliminating the need for the time-consuming manual process. This is the first dataset of its kind to be released in public and its use will allow the operator to make any necessary changes to the ICF polishing process for optimal results. This dataset contains the raw vibration data of multiple polishing runs with their extracted statistical features and the corresponding surface roughness values. Additionally, to generalize the prediction models to different polishing conditions, we also apply domain adaptation techniques to improve prediction accuracy for conditions unseen by the trained model. The dataset is available in \url{https://junzeliu.github.io/Diamond-Polishing-Dataset/}. Antonios Alexos, Junze Liu, Shashank Galla, Sean Hayes, Kshitij Bhardwaj, Alexander Schwartz, Monika Biener, Pierre Baldi, Satish T. S. Bukkapatnam, Suhas Bhandarkar |
NeurIPS | 5 |
| 2023 | Real-Time Fully Unsupervised Domain Adaptation for Lane Detection in Autonomous DrivingabstractWhile deep neural networks are being utilized heavily for autonomous driving, they need to be adapted to new unseen environmental conditions for which they were not trained. We focus on a safety critical application of lane detection, and propose a lightweight, fully unsupervised, real-time adaptation approach that only adapts the batch-normalization parameters of the model. We demonstrate that our technique can perform inference, followed by on-device adaptation, under a tight constraint of 30 FPS on Nvidia Jetson Orin. It shows similar accuracy (avg. of 92.19%) as a state-of-the-art semi-supervised adaptation algorithm but which does not support real-time adaptation. Kshitij Bhardwaj, Zishen Wan, Arijit Raychowdhury, Ryan A. Goldhahn |
DATE | 1 |
| 2022 | Unsupervised Test-Time Adaptation of Deep Neural Networks at the Edge: A Case StudyabstractDeep learning is being increasingly used in mobile and edge autonomous systems. The prediction accuracy of deep neural networks (DNNs), however, can degrade after deployment due to encountering data samples whose distributions are differ-ent than the training samples. To continue to robustly predict, DNNs must be able to adapt themselves post-deployment. Such adaptation at the edge is challenging as new labeled data may not be available, and it has to be performed on a resource-constrained device. This paper performs a case study to evaluate the cost of test-time fully unsupervised adaptation strategies on a real-world edge platform: Nvidia Jetson Xavier NX. In particular, we adapt pretrained state-of-the-art robust DNNs (trained using data augmentation) to improve the accuracy on image classification data that contains various image corruptions. During this prediction-time on-device adaptation, the model parameters of a DNN are updated using a single backpropagation pass while optimizing entropy loss. The effects of following three simple model updates are compared in terms of accuracy, adaptation time and energy: updating only convolutional (Conv-Tune); only fully-connected (FC-Tune); and only batch-norm parameters (BN-Tune). Our study shows that BN-Tune and Conv-Tune are more effective than FC-Tune in terms of improving accuracy for corrupted images data (average of 6.6%, 4.97%, and 4.02%, respectively over no adaptation). However, FC-Tune leads to significantly faster and more energy efficient solution with a small loss in accuracy. Even when using FC-Tune, the extra overheads of on-device fine-tuning are significant to meet tight real-time deadlines (209ms). This study motivates the need for designing hardware-aware robust algorithms for efficient on-device adaptation at the autonomous edge. Kshitij Bhardwaj, James Diffenderfer, Bhavya Kailkhura, Maya B. Gokhale |
DATE | 1 |
| 2022 | Benchmarking Test-Time Unsupervised Deep Neural Network Adaptation on Edge DevicesabstractThe prediction accuracy of deep neural networks (DNNs) after deployment at the edge can suffer with time due to shifts in the distribution of the new data. To improve robustness of DNNs, they must be able to update themselves. However, DNN adaptation at the edge is challenging due to lack of resources. Recently, lightweight prediction-time unsupervised DNN adaptation techniques have been introduced that improve prediction accuracy of the models for noisy data by re-tuning the batch normalization parameters. This paper performs a comprehensive measurement study of such techniques to quantify their performance and energy on various edge devices as well as find bottlenecks and propose optimization opportunities. Kshitij Bhardwaj, James Diffenderfer, Bhavya Kailkhura, Maya B. Gokhale |
ISPASS | 1 |
| 2022 | Roofline Model for UAVs: A Bottleneck Analysis Tool for Onboard Compute Characterization of Autonomous Unmanned Aerial VehiclesabstractWe introduce an early-phase bottleneck analysis and characterization model called the F-1 for designing computing systems that target autonomous Unmanned Aerial Vehicles (UAVs). The model provides insights by exploiting the fundamental relationships between various components in the autonomous UAV, such as sensor, compute, and body dynamics. To guarantee safe operation while maximizing the performance (e.g., velocity) of the UAV, the compute, sensor, and other mechanical properties must be carefully selected or designed. The F-1 model provides visual insights that can aid a system architect in understanding the optimal compute design or selection for autonomous UAVs. The model is experimentally validated using real UAVs, and the error is between 5.1% to 9.5% compared to real-world flight tests. An interactive web-based tool for the F-1 model called Skyline is available for free of cost use at: https://bit.ly/skyline-tool Srivatsan Krishnan, Zishen Wan, Kshitij Bhardwaj, Ninad Jadhav, Aleksandra Faust, Vijay Janapa Reddi |
ISPASS | 3 |
| 2022 | Automatic Domain-Specific SoC Design for Autonomous Unmanned Aerial VehiclesabstractBuilding domain-specific accelerators is becoming increasingly paramount to meet the high-performance requirements under stringent power and real-time constraints. However, emerging application domains like autonomous vehicles are complex systems with constraints extending beyond the computing stack. Manually selecting and navigating the design space to design custom and efficient domain-specific SoCs (DSSoC) is tedious and expensive. Hence, there is a need for automated DSSoC design methodologies. In this paper, we use agile and autonomous UAVs as a case study to understand how to automate domain-specific SoCs design for autonomous vehicles. Architecting a UAV DSSoC requires consideration of parameters such as sensor rate, compute throughput, and other physical characteristics (e.g., payload weight, thrust-to-weight ratio) that affect overall performance. Iterating over several component choices results in a combinatorial explosion of the number of possible combinations: from tens of thousands to billions, depending on implementation details. To navigate the DSSoC design space efficiently, we introduce AutoPilot, a systematic methodology for automatically designing DSSoC for autonomous UAVs. AutoPilot uses machine learning to navigate the large DSSoC design space and automatically select a combination of autonomy algorithm and hardware accelerator while considering the cross-product effect across different UAV components. AutoPilot consistently outperforms general-purpose hardware selections like Xavier NX and Jetson TX2, as well as dedicated hardware accelerators built for autonomous UAVs. DSSoC designs generated by AutoPilot increase the number of missions on average by up to 2.25×, 1.62×, and 1.43× for nano, micro, and mini-UAVs, respectively, over baselines. Further, we discuss the potential application of AutoPilot methodology to other related autonomous vehicles. Srivatsan Krishnan, Zishen Wan, Kshitij Bhardwaj, Paul N. Whatmough, Aleksandra Faust, Sabrina M. Neuman, Gu-Yeon Wei, David Brooks 0001, Vijay Janapa Reddi |
MICRO | 3 |
| 2020 | A comprehensive methodology to determine optimal coherence interfaces for many-accelerator SoCsabstractModern systems-on-chip (SoCs) include not only general-purpose CPUs but also specialized hardware accelerators. Typically, there are three coherence model choices to integrate an accelerator with the memory hierarchy: no coherence, coherent with the last-level cache (LLC), and private cache based full coherence. However, there has been very limited research on finding which coherence models are optimal for the accelerators of a complex many-accelerator SoC. This paper focuses on determining a cost-aware coherence interface for an SoC and its target application: find the best coherence models for the accelerators that optimize their power and performance, considering both workload characteristics and system-level contention. A novel comprehensive methodology is proposed that uses Bayesian optimization to efficiently find the cost-aware coherence interfaces for SoCs that are modeled using the gem5-Aladdin architectural simulator. For a complete analysis, gem5-Aladdin is extended to support LLC coherence in addition to already-supported no coherence and full coherence. For a heterogeneous SoC targeting applications with varying amount of accelerator-level parallelism, the proposed framework rapidly finds cost-aware coherence interfaces that show significant performance and power benefits over the other commonly-used coherence interfaces. Kshitij Bhardwaj, Marton Havasi, Yuan Yao 0006, David Brooks 0001, José Miguel Hernández-Lobato, Gu-Yeon Wei |
ISLPED | 1 |
| 2020 | SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning WorkloadsabstractIn recent years, there has been tremendous advances in hardware acceleration of deep neural networks. However, most of the research has focused on optimizing accelerator microarchitecture for higher performance and energy efficiency on a per-layer basis. We find that for overall single-batch inference latency, the accelerator may only make up 25–40%, with the rest spent on data movement and in the deep learning software framework. Thus far, it has been very difficult to study end-to-end DNN performance during early stage design (before RTL is available), because there are no existing DNN frameworks that support end-to-end simulation with easy custom hardware accelerator integration. To address this gap in research infrastructure, we present SMAUG, the first DNN framework that is purpose-built for simulation of end-to-end deep learning applications. SMAUG offers researchers a wide range of capabilities for evaluating DNN workloads, from diverse network topologies to easy accelerator modeling and SoC integration. To demonstrate the power and value of SMAUG, we present case studies that show how we can optimize overall performance and energy efficiency for up to 1.8×–5× speedup over a baseline system, without changing any part of the accelerator microarchitecture, as well as show how SMAUG can tune an SoC for a camera-powered deep learning pipeline. Sam Likun Xi, Yuan Yao 0006, Kshitij Bhardwaj, Paul N. Whatmough, Gu-Yeon Wei, David Brooks 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2019 | Towards a Complete Methodology for Synthesizing Bundled-Data Asynchronous Circuits on FPGAsabstractAsynchronous circuits are gaining momentum as a promising low-power alternative to the conventional synchronous design approaches. In particular, single-rail bundled-data design style has seen significant interest both for designing GALS systems and in the emerging area of neuromorphic computing. However, there has been only limited research on implementing these asynchronous circuits on commercial FPGAs, which can be challenging due to the use of relative timing constraints in these designs for correct operation. This paper proposes a systematic CAD methodology to synthesize efficiently bundled-data asynchronous circuits on commercial FPGAs, achieving a two-fold goal for the target implementation: robustness and high performance. The methodology is targeted to the existing Xilinx Vivado tool set. As a case study, two asynchronous NoC switches are prototyped on Xilinx Virtex 7 in 28 nm: one supporting unicast, and the other also handling multicast. The former shows significant energy and idle power improvements, with some performance benefits, over a high-performance synchronous FPGA-based switch. The asynchronous multicast router also shows promising energy and performance results. Although a NoC case study is used, the proposed approach is general and can be used for other bundled-data asynchronous circuits. Kshitij Bhardwaj, Paolo Mantovani, Luca P. Carloni, Steven M. Nowick |
ISLPED | 1 |
| 2019 | A Continuous-Time Replication Strategy for Efficient Multicast in Asynchronous NoCsabstractMulticast communication (one-to-many) is common in parallel architectures and emerging areas, such as neuromorphic computing. However, there is very limited research in supporting multicast in asynchronous networks-on-chip (NoCs). This paper proposes a new parallel multicast asynchronous NoC with a 2-D mesh topology. To the best of our knowledge, this is the first general-purpose asynchronous NoC to support multicast in 2-D meshes. A critical feature of this NoC is the use of a new continuous-time replication strategy, where the flits of a multicast packet are routed through the distinct outputs of the router according to each output's own rate, in parallel, and in continuous time. This unique asynchronous continuous-time replication, not discretized to clock cycles, can handle subtle variations in network congestion and exploit “subcycle” differentials in operating speeds. A new continuous-time multiway read (CMR) buffer is proposed to enable this replication strategy. Only a single CMR buffer is used per input port, with multiple independent read pointers, which is accessed by different outputs. For diverse multicast benchmarks, the new parallel multicast network is achieved significant latency and throughput gains over a serial baseline. Interestingly, consistent latency improvements were observed for unicast, in spite of the extra instrumentation. Kshitij Bhardwaj, Steven M. Nowick |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Achieving Lightweight Multicast in Asynchronous NoCs Using a Continuous-Time Multi-Way Read BufferabstractMulticast communication (1-to-many) is common in parallel architectures and emerging areas such as neuromorphic computing. However, there is very limited research in supporting multicast in asynchronous NoCs. This paper proposes a new parallel multicast asynchronous NoC with 2D-mesh topology. To the best of our knowledge, this is the first general-purpose asynchronous NoC to support multicast in 2D meshes. A critical feature of this NoC is the use of a new continuous-time replication strategy, where the flits of a multicast packet are routed through the distinct outputs of the router according to each output's own rate, in parallel and in continuous time. This unique asynchronous continuous-time replication, not discretized to clock cycles, can handle subtle variations in network congestion and exploit "sub-cycle" differentials in operating speeds. A new continuous-time multi-way read (CMR) buffer is proposed to enable this replication strategy. Only a single CMR buffer is used per input port, with multiple independent read pointers, which are accessed by different outputs. For diverse multicast benchmarks, the new parallel multicast network achieved significant latency and throughput gains over a serial baseline. Moderate energy overhead was seen for one benchmark with a small multicast portion, but major reductions were achieved for higher amounts of multicast. Interestingly, consistent latency improvements were observed for unicast, in spite of the extra instrumentation. Experiments on isolated multicast packet transmissions also showed over an order-of-magnitude improvement in delivery time. Kshitij Bhardwaj, Weiwei Jiang 0002, Steven M. Nowick |
NOCS | 1 |
| 2016 | Achieving lightweight multicast in asynchronous networks-on-chip using local speculationabstractWe propose a lightweight parallel multicast targeting an asynchronous NoC with a variant Mesh-of-Trees topology. A novel strategy, local speculation, is introduced, where a subset of switches are speculative and always broadcast. These switches are surrounded by non-speculative switches, which throttle any redundant packets, restricting these packets to small regions. Speculative switches have simplified designs, thereby improving network performance. A hybrid network architecture is proposed to mix the speculative and non-speculative switches. For multicast benchmarks, significant performance improvements with small power savings are obtained by the new approach over a tree-based non-speculative approach. Interestingly, similar improvements are also shown for unicast. Finally, another benefit is to reduce the address field size in multicast packets. Kshitij Bhardwaj, Steven M. Nowick |
DAC | 1 |
| 2015 | A Lightweight Early Arbitration Method for Low-Latency Asynchronous 2D-Mesh NoC's
Weiwei Jiang 0002, Kshitij Bhardwaj, Geoffray Lacourba, Steven M. Nowick |
DAC | 2 |
| 2015 | Wearout Resilience in NoCs Through an Aging Aware Adaptive Routing AlgorithmabstractContinuous technology scaling has made aging mechanisms, such as negative bias temperature instability and electromigration primary concerns in network-on-chip (NoC) designs. In this paper, we extensively analyze the effects of these aging mechanisms on NoC routers and links. We observe a critical need of a robust aging-aware routing algorithm that not only reduces power-performance overheads caused due to aging degradation, but also minimizes the stress experienced by heavily utilized routers and links. To solve this problem, we propose an aging-aware adaptive routing algorithm and a router microarchitecture that routes the packets along the paths, which are both least congested and experience minimum aging degradation. After an extensive experimental analysis using real workloads, we observe 13% and 12.17% average overhead reduction in network latency and energy-delay product per flit, a 10.4% improvement in performance, and a 60% improvement in mean time to failure using our aging-aware routing algorithm. Dean Michael Ancajas, Kshitij Bhardwaj, Koushik Chakraborty, Sanghamitra Roy |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Towards graceful aging degradation in NoCs through an adaptive routing algorithmabstractContinuous technology scaling has made aging mechanisms such as Negative Bias Temperature Instability (NBTI) and electromigration primary concerns in Network-on-Chip (NoC) designs. In this paper, we model the effects of these aging mechanisms on NoC components such as routers and links using a novel reliability metric called Traffic Threshold per Epoch (TTpE). We observe a critical need of a robust aging-aware routing algorithm that not only reduces power-performance overheads caused due to aging degradation but also minimizes the stress experienced by heavily utilized routers and links. To solve this problem, we propose an aging-aware adaptive routing algorithm and a router microarchitecture that routes the packets along the paths which are both least congested and experience minimum aging stress. After an extensive experimental analysis using real workloads, we observe a 13%, 12.7% average overhead reduction in network latency and Energy-Delay-Product-Per-Flit (EDPPF) and a 10.4% improvement in performance using our aging-aware routing algorithm. Kshitij Bhardwaj, Koushik Chakraborty, Sanghamitra Roy |
DAC | 1 |
| 2012 | An MILP-based aging-aware routing algorithm for NoCsabstractNetwork-on-Chip (NoC) architectures have emerged as a better replacement of the traditional bus-based communication in the many-core era. However, continuous technology scaling has made aging mechanisms such as Negative Bias Temperature Instability (NBTI) and electromigration primary concerns in NoC design. In this paper1, we propose a novel system-level aging model to model the effects of asymmetric aging in NoCs. We observe a critical need of a holistic aging analysis, which when combined with power-performance optimization, poses a multi-objective design challenge. To solve this problem, we propose a Mixed Integer Linear Programming (MILP)-based aging-aware routing algorithm that optimizes the various design constraints using a multi-objective formulation. After an extensive experimental analysis using real workloads, we observe a 62.7%, 46% average overhead reduction in network latency and Energy-Delay-Product-Per-Flit (EDPPF) and a 41% improvement in Instructions Per Cycle (IPC) using our aging-aware routing algorithm. Kshitij Bhardwaj, Koushik Chakraborty, Sanghamitra Roy |
DATE | 1 |