EDBT 2026 Demo / reviewers in the wild / expert
Rajesh Kedia
dblp:198/1986
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-3711-0396ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Software-Based Approximate Multiplication on Multiplierless CPUs Using Custom InstructionabstractMany lightweight, low-power microcontrollers contain CPU cores without hardware multipliers. When deployed for applications involving multiplication, such devices emulate multiplication using computationally expensive software techniques. While existing works used approximate hardware multipliers to reduce the complexity of multiplication, there are limited works on approximate multiplication in software. Translating ideas from approximate hardware multipliers directly to a software code does not provide considerable improvements. In this work, we study common approaches for approximate hardware multipliers and then propose a custom instruction for leading-one detection (LOD) which significantly reduces the complexity of performing approximate multiplication in software.We implement LOD instruction in a RISC-V based core and use it to develop seven different software routines for approximate multiplication, and evaluate them on four different popular kernels. These routines are based on existing approximate hardware multipliers and three newly proposed techniques. These routines provide multiple points of trade-off between error rate and computation cycles, enabling configurable choices to the designer. When multiplying 1 million random pairs of numbers, one of the proposed techniques, RoBA-(UpDn)2, enables a very low error rate of 0.05% on average and 0.83% as maximum error; while consuming only 51% of the original CPU cycles. When deployed on two image processing applications, RoBA-(UpDn)2 can provide a signal to noise ratio (SNR) larger than 50 dB; while consuming about 70% of the original computation cycles. Shalu Prathmesh Rajiv, Rajesh Kedia |
DATE | 2 |
| 2025 | EXPRESS: A Framework for Execution Time Prediction of Concurrent CNNs on Xilinx DPU AcceleratorabstractDeep learning Processor Unit (DPU) is a highly configurable CNN accelerator that supports a variety of CNNs and can be implemented with multiple instances on the same FPGA. Many applications deploy concurrent execution of different CNNs and in such a setting, an execution time predictor can help “optimize” the DPU configurations to meet the performance requirements of different tasks. We characterize CNN execution on DPUs and reduce the variability in execution time due to interference from the operating system. Subsequently, we propose a machine learning-based framework (EXPRESS) to predict the execution time of any given CNN on a DPU configuration, considering CNN, DPU, and bus characteristics. We improvise EXPRESS to support heterogeneous CNNs in EXPRESS-2.0 by making features independent of the number of CNNs. Our entire experimentation is based on data from a real FPGA board for 16 standard CNNs. Our frameworks, EXPRESS and EXPRESS-2.0, significantly outperform state-of-the-art by achieving an average execution time prediction error of 2.2% and 0.7%, respectively. We illustrate the effectiveness of this low prediction error for design space exploration, which is very useful for embedded system application developers. Shikha Goel, Rajesh Kedia, Rijurekha Sen, M. Balakrishnan |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | Dynamic Thermal Management of 3D Memory through Rotating Low Power States and Partial Channel ClosureabstractModern high-performance and high-bandwidth three-dimensional (3D) memories are characterized by frequent heating. Prior art suggests turning off hot channels and migrating data to the background DDR memory, incurring significant performance and energy overheads. We propose three Dynamic Thermal Management (DTM) approaches for 3D memories, reducing these overheads. The first approach, Rotating-channel Low-power-state-based DTM (RL-DTM) , minimizes the energy overheads by avoiding data migration. RL-DTM places 3D memory channels into low power states instead of turning them off. Since data accesses are disallowed during low power state, RL-DTM balances each channel’s low-power-state duration. The second approach, Masked rotating-channel Low-power-state-based DTM (ML-DTM) , is a fine-grained policy that minimizes the energy-delay product (EDP) and improves the performance of RL-DTM by considering the channel access rate. The third strategy, Partial channel closure and ML-DTM , minimizes performance overheads of existing channel-level turn-off-based policies by closing a channel only partially and integrating ML-DTM, reducing the number of channels being turned off. We evaluate the proposed DTM policies using various mixes of SPEC benchmarks and multi-threaded workloads and observe them to significantly improve performance, energy, and EDP over state-of-the-art approaches for different 3D memory architectures. Lokesh Siddhu, Aritra Bagchi, Rajesh Kedia, Isaar Ahmad, Shailja Pandey, Preeti Ranjan Panda |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | CoreMemDTM: Integrated Processor Core and 3D Memory Dynamic Thermal Management for Improved PerformanceabstractThe growing performance of processors and 3D memories has resulted in higher power densities and temperatures. Dynamic thermal management (DTM) policies for processor cores and memory have received significant research attention, but existing solutions address processors and 3D memories independently, which causes overcompensation, and there is a need to coordinate the DTM of the two subsystems. Further, existing CPU DTM policies slow down heated cores significantly, increasing the overall execution time and performance overheads. We propose CoreMemDTM, a technique for integrating processor core and 3D memory DTM policies that attempts to minimize performance overheads. We suggest employing DTM depending on the thermal margin since safe temperature thresholds might differ for the two subsystems. We propose a stall-balanced core DVFS policy for core thermal management that enables distributed cooling, decreasing overheads. We evaluate CoreMemDTM using ten different SPEC CPU2017 workloads across various safe temperature thresholds and observe average execution time and energy improvements of 14% and 36% compared to state-of-the-art DTM policies. Lokesh Siddhu, Rajesh Kedia, Preeti Ranjan Panda |
DATE | 2 |
| 2022 | EXPRESS: CNN EXecution Time PREdiction for DPU DeSign Space ExplorationabstractDeep learning Processor Units (DPUs) from Xilinx are design-time configurable CNN accelerators for FPGAs. We propose EXPRESS, which predicts the execution time of any given CNN on a DPU. EXPRESS incorporates the effect of bus connections into prediction. As a DPU is invoked by a host CPU to process a CNN layer by layer, EXPRESS considers the CPU and the DPU execution time for predicting the end-to-end processing time. EXPRESS has an average prediction error of 2.2% and significantly outperforms state-of-the-art. Shikha Goel, Rajesh Kedia, Rijurekha Sen, M. Balakrishnan |
FPT | 2 |
| 2022 | CoMeT: An Integrated Interval Thermal Simulation Toolchain for 2D, 2.5D, and 3D Processor-Memory SystemsabstractProcessing cores and the accompanying main memory working in tandem enable modern processors. Dissipating heat produced from computation remains a significant problem for processors. Therefore, the thermal management of processors continues to be an active subject of research. Most thermal management research is performed using simulations, given the challenges in measuring temperatures in real processors. Fast yet accurate interval thermal simulation toolchains remain the research tool of choice to study thermal management in processors at the system level. However, the existing toolchains focus on the thermal management of cores in the processors, since they exhibit much higher power densities than memory. The memory bandwidth limitations associated with 2D processors lead to high-density 2.5D and 3D packaging technology: 2.5D packaging technology places cores and memory on the same package; 3D packaging technology takes it further by stacking layers of memory on the top of cores themselves. These new packagings significantly increase the power density of the processors, making them prone to overheating. Therefore, mitigating thermal issues in high-density processors (packaged with stacked memory) becomes even more pressing. However, given the lack of thermal modeling for memories in existing interval thermal simulation toolchains, they are unsuitable for studying thermal management for high-density processors. To address this issue, we present the first integrated Core and Memory interval Thermal (CoMeT) simulation toolchain.CoMeTcomprehensively supports thermal simulation of high- and low-density processors corresponding to four different core-memory (integration) configurations—off-chip DDR memory, off-chip 3D memory, 2.5D, and 3D.CoMeTsupports several novel features that facilitate overlying system research.CoMeTadds only an additional ~5% simulation-time overhead compared to an equivalent state-of-the-art core-only toolchain. The source code ofCoMeThas been made open for public use under theMITlicense. Lokesh Siddhu, Rajesh Kedia, Shailja Pandey, Martin Rapp, Anuj Pathania, Jörg Henkel, Preeti Ranjan Panda |
ACM Trans. Archit. Code Optim. | 2 |
| 2022 | FN-CACTI: Advanced CACTI for FinFET and NC-FinFET TechnologiesabstractCache memories are an indispensable component of many processor-based systems and contribute significantly to the overall area, power consumption, and delay. This leads to an important role played by modeling tools for estimating the area, power consumption, and access time of cache memories. However, existing modeling tools such as CACTI and its various extensions have been primarily designed using data from various projections. For the first time, we propose an entire flow for obtaining/calibrating the transistor characteristics from a commercial technology and use these characteristics within CACTI. We also improve the modeling approach to make them more fine-grained and follow recent manufacturing trends suitable for FinFET technology. Further, for the first time, we extend CACTI to support negative capacitance fin field effect transistor (NC-FinFET), an emerging technology depicting negative capacitance whose current and capacitive characteristics are very different compared to those of the FinFET. We use the proposed tool (FN-CACTI) to identify NC-FinFET-based caches to be significantly more energy-efficient than corresponding FinFET-based caches. We also study an application of FN-CACTI to determine optimal voltages corresponding to the lowest energy consumption for NC-FinFET and FinFET-based caches of various sizes. Divya Praneetha Ravipati, Rajesh Kedia, Victor M. van Santen, Jörg Henkel, Preeti Ranjan Panda, Hussam Amrouch |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | Leakage-Aware Dynamic Thermal Management of 3D Memoriesabstract3D memory systems offer several advantages in terms of area, bandwidth, and energy efficiency. However, thermal issues arising out of higher power densities have limited their widespread use. While prior works have looked at reducing dynamic power through reduced memory accesses, in these memories, both leakage and dynamic power consumption are comparable. Furthermore, as the temperature rises, the leakage power increases, creating a thermal-leakage loop. We study the impact of leakage power on 3D memory temperature and propose turning OFF specific memory channels to meet thermal constraints. Data is migrated to a 2D memory before closing a 3D channel. We introduce an analytical model to assess the 2D memory delay and use the model to guide data migration decisions. The above strategy is referred to asFastCooland provides an improvement of 22%, 19%, and 32% on average (up to 57%, 72%, and 82%) in performance, memory energy, and energy-delay product (EDP), respectively, on different workloads consisting of SPEC CPU2006 benchmarks. We further propose a thermal management strategy namedEnergy-Efficient FastCool (EEFC), which improves upon FastCool by selecting the channels to be closed by considering temperature, leakage, access rate, and position of various 3D memory channels at runtime. Our experiments demonstrate that EEFC leads to an additional improvement of up to 30%, 30%, and 51% in performance, memory energy, and EDP compared to FastCool. Finally, we analyze the effects of process variations on the efficiency of the proposed FC and EEFC strategies. Variation in the manufacturing process causes changes in the leakage power and temperature profile. Since EEFC considers both while selecting channels for closure, it is more resilient to process variations and achieves a lower application execution time and memory energy compared to FastCool. Lokesh Siddhu, Rajesh Kedia, Preeti Ranjan Panda |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2019 | GRanDE: Graphical Representation and Design Space Exploration of Embedded SystemsabstractTasks executing computer vision and machine learning algorithms are becoming popular on embedded platforms. A key characteristic of such tasks is the presence of modes providing different levels of application performance in terms of metrics like accuracy. The system designer has the flexibility to select an appropriate mode for executing such tasks. Secondly, the designer also has the traditional flexibility of choosing suitable components to build the execution platform. Thirdly, the system performance might vary with various external factors (known as context), and during the initial stages of system design, the designer might have the flexibility to support only a subset of the possible contexts. This three-fold flexibility in the hands of the designer has not been explored simultaneously in prior works and raises the complexity of designing embedded systems many-fold. In this paper, we address the design of such systems through a novel framework named GRanDE (Graphical Representation and Design Space Exploration). GRanDE consists of a comprehensive graphical representation to capture the three aspects of the design space discussed earlier. Further, we transform this representation into Constraint Logic Programming (CLP) constructs, which could be used to interactively explore and prune the design space. We demonstrate the applicability of the proposed framework on an embedded system named MAVI having ~1.3 million design points. The generated CLP program could prune up to 99.74% of the design space of MAVI. Rajesh Kedia, M. Balakrishnan, Kolin Paul |
DSD | 1 |