EDBT 2026 Demo / reviewers in the wild / expert
Venkata Sai Praneeth Karempudi
dblp:253/8601
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-0609-6894ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Up the Sustainability of Photonic Tensor Cores With Device-Circuit-Signaling Co-DesignabstractPhotonic tensor cores have grown in popularity over the past few years for accelerating tensor-based kernels found in abundance in deep learning workloads because they offer potentially massive spatial parallelism (across wavelengths and waveguides), subnanosecond-scale start-to-solution latency, and near-dissipation-free dynamic operation. However, several shortcomings severely limit the practically achievable parallelism, processing throughput, and energy efficiency in existing photonic tensor core architectures. For instance, the wavelength-selective analog operation of existing designs makes them highly prone to crosstalk noise and other optical signal penalties and losses. These penalties and losses interplay with an already tight optical power budget, leading to strong trade-offs for achievable spatial parallelism, operating data rate, and analog precision. This paper shows how co-designing low-dissipation, low-noise, and high-speed electro-photonic devices, crosstalk-minimal circuit organizations, and mixed unary/analog signaling methods can overcome these shortcomings to realize photonic tensor cores with scaled-up throughput and operational energy-efficiency benefits for accelerating a variety of deep learning workloads. Ishan G. Thakkar, Sairam Sri Vatsavai, Venkata Sai Praneeth Karempudi, Oluwaseun Adewunmi Alo |
ICCD | 3 |
| 2025 | HEANA: A Hybrid Time-Amplitude Analog Optical Accelerator with Flexible Dataflows for Energy-Efficient CNN InferenceabstractSeveral photonic microring resonator (MRR)-based analog accelerators have been proposed to accelerate the inference of integer-quantized Convolutional Neural Networks (CNNs) with remarkably higher throughput and energy efficiency compared to their electronic counterparts. However, the existing analog photonic accelerators suffer from three shortcomings: (1) severe hampering of wavelength parallelism due to various crosstalk effects, (2) inflexibility of supporting various dataflows with temporal accumulations, and (3) failure in fully leveraging the ability of photodetectors to perform in situ accumulations. These shortcomings collectively hamper the performance and energy efficiency of prior accelerators. To tackle these shortcomings, we present a novel H ybrid tim E - A mplitude a N alog optical A ccelerator, called HEANA. HEANA employs hybrid time-amplitude analog optical modulators (TAOMs) in a spectrally hitless arrangement, which significantly reduces optical signal losses and crosstalk effects, thereby increasing the wavelength parallelism in HEANA. HEANA employs our invented balanced photo-charge accumulators (BPCAs) that enable buffer-less, in situ, spatio-temporal accumulations to eliminate the need to use reduction networks in HEANA, relieving it from related latency and energy overheads. Moreover, TAOMs and BPCAs increase the flexibility of HEANA to efficiently support spatio-temporal accumulations for various dataflows. Our evaluation for the inference of four modern CNNs indicates that HEANA provides improvements of at least 25× and 32× in frames per second (FPS) and FPS/W (energy efficiency), respectively, for equal-area comparisons on gmean over two MRR-based analog CNN accelerators from prior work. Sairam Sri Vatsavai, Venkata Sai Praneeth Karempudi, Ishan G. Thakkar |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2024 | An Analysis of Various Design Pathways Towards Multi-Terabit Photonic On-Interposer InterconnectsabstractIn the wake of dwindling Moore’s Law, to address the rapidly increasing complexity and cost of fabricating large-scale, monolithic systems-on-chip (SoCs), the industry has adopted dis-aggregation as a solution, wherein a large monolithic SoC is partitioned into multiple smaller chiplets that are then assembled into a large system-in-package (SiP) using advanced packaging substrates such as silicon interposer. For such interposer-based SiPs, there is a push to realize on-interposer inter-chiplet communication bandwidth of multi-Tb/s and end-to-end communication latency of no more than 10 ns. This push comes as the natural progression from some recent prior works on SiP design, and is driven by the proliferating bandwidth demand of modern data-intensive workloads. To meet this bandwidth and latency goal, prior works have focused on a potential solution of using the silicon photonic interposer (SiPhI) for integrating and interconnecting a large number of chiplets into an SiP. Despite the early promise, the existing designs of on-SiPhI interconnects still have to evolve by leaps and bounds to meet the goal of multi-Tb/s bandwidth. However, the possible design pathways, upon which such an evolution can be achieved, have not been explored in any prior works yet. In this paper, we have identified several design pathways that can help evolve on-SiPhI interconnects to achieve multi-Tb/s aggregate bandwidth. We perform an extensive link-level and system-level analysis in which we explore these design pathways in isolation and in different combinations of each other. From our link-level analysis, we have observed that the design pathways that simultaneously enhance the spectral range and optical power budget available for wavelength multiplexing can render aggregate bandwidth of up to 4 Tb/s per on-SiPhI link. We also show that such high-bandwidth on-SiPhI links can substantially improve the performance and energy-efficiency of the state-of-the-art CPU and GPU chiplets based SiPs. Venkata Sai Praneeth Karempudi, Janibul Bashir, Ishan G. Thakkar |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2023 | High-Speed and Energy-Efficient Non-Binary Computing with Polymorphic Electro-Optic Circuits and ArchitecturesabstractIn this paper, we present microring resonator (MRR) based polymorphic E-O circuits and architectures that can be employed for high-speed and energy-efficient non-binary reconfigurable computing. Our polymorphic E-O circuits can be dynamically programmed to implement different logic and arithmetic functions at different times. They can provide compactness and polymorphism to consequently improve operand handling, reduce idle time, and increase amortization of area and static power overheads. When combined with flexible photodetectors with the innate ability to accumulate a high number of optical pulses in situ, our circuits can support energy-efficient processing of data in non-binary formats such as stochastic/unary and high-dimensional reservoir formats. Furthermore, our polymorphic E-O circuits enable configurable E-O computing accelerator architectures for processing binarized and integer quantized convolutional neural networks (CNNs). We compare our designed polymorphic E-O circuits and architectures to several circuits and architectures from prior works in terms of area, latency, and energy consumption. Ishan G. Thakkar, Sairam Sri Vatsavai, Venkata Sai Praneeth Karempudi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | SCONNA: A Stochastic Computing Based Optical Accelerator for Ultra-Fast, Energy-Efficient Inference of Integer-Quantized CNNsabstractConvolutional Neural Networks (CNNs) are used extensively for artificial intelligence applications due to their record-breaking accuracy. For efficient and swift hardware-based acceleration, CNNs are typically quantized to have integer input/weight parameters. The acceleration of a CNN inference task uses convolution operations that are typically transformed into vector-dot-product (VDP) operations. Several photonic microring resonators (MRRs) based hardware architectures have been proposed to accelerate integer-quantized CNNs with remarkably higher throughput and energy efficiency compared to their electronic counterparts. However, the existing photonic MRR-based analog accelerators exhibit a very strong trade-off between the achievable input/weight precision and VDP operation size, which severely restricts their achievable VDP operation size for the quantized input/weight precision of 4 bits and higher. The restricted VDP operation size ultimately suppresses computing throughput to severely diminish the achievable performance benefits. To address this shortcoming, we for the first time present a merger of stochastic computing and MRR-based CNN accelerators. To leverage the innate precision flexibility of stochastic computing, we invent an MRR-based optical stochastic multiplier (OSM). We employ multiple OSMs in a cascaded manner using dense wavelength division multiplexing, to forge a novel Stochastic Computing based Optical Neural Network Accelerator (SCONNA). SCONNA achieves significantly high throughput and energy efficiency for accelerating inferences of high-precision quantized CNNs. Our evaluation for the inference of four modern CNNs at 8-bit input/weight precision indicates that SCONNA provides improvements of up to 66.5×, 90× and 91× in frames-per-second (FPS), FPS/W and FPS/W/mm2respectively, on average over two photonic MRR-based analog CNN accelerators from prior work, with Top-1 accuracy drop of only up to 0.4% for large CNNs and up to 1.5% for small CNNs. We developed a transaction-level, event-driven python-based simulator for the evaluation of SCONNA and other accelerators (https://github.com/uky-UCAT/SC_ONN_SIM.git). Sairam Sri Vatsavai, Venkata Sai Praneeth Karempudi, Ishan G. Thakkar, Sayed Ahmad Salehi, Jeffrey Todd Hastings |
IPDPS | 2 |
| 2022 | Photonic Networks-on-Chip Employing Multilevel Signaling: A Cross-Layer Comparative StudyabstractPhotonic network-on-chip (PNoC) architectures employ photonic links with dense wavelength-division multiplexing (DWDM) to enable high throughput on-chip transfers. Unfortunately, increasing the DWDM degree (i.e., using a larger number of wavelengths) to achieve a higher aggregated data rate in photonic links and, hence, higher throughput in PNoCs, requires sophisticated and costly laser sources along with extra photonic hardware. This extra hardware can introduce undesired noise to the photonic link and increase the bit error rate (BER), power, and area consumption of PNoCs. To mitigate these issues, the use of 4-pulse amplitude modulation (4-PAM) signaling, instead of the conventional on-off keying (OOK) signaling, can halve the wavelength signals utilized in photonic links for achieving the target aggregate data rate while reducing the overhead of crosstalk noise, BER, and photonic hardware. There are various designs of 4-PAM modulators reported in the literature. For example, the signal superposition (SS)–, electrical digital-to-analog converter (EDAC)–, and optical digital-to-analog converter (ODAC)–based designs of 4-PAM modulators have been reported. However, it is yet to be explored how these SS-, EDAC-, and ODAC-based 4-PAM modulators can be utilized to design DWDM-based photonic links and PNoC architectures. In this article, we provide a systematic analysis of the SS, EDAC, and ODAC types of 4-PAM modulators from prior work with regards to their applicability and utilization overheads. We then present a heuristic-based search method to employ these 4-PAM modulators for designing DWDM-based SS, EDAC, and ODAC types of 4-PAM photonic links with two different design goals: (i) to attain the desired BER of 10 -9 at the expense of higher optical power and lower aggregate data rate and (ii) to attain maximum aggregate data rate with the desired BER of 10 -9 at the expense of longer packet transfer latency. We then employ our designed 4-PAM SS–, 4-PAM EDAC–, 4-PAM ODAC–, and conventional OOK modulator–based photonic links to constitute corresponding variants of the well-known CLOS and SWIFT PNoC architectures. We eventually compare our designed SS-, EDAC-, and ODAC-based variants of 4-PAM links and PNoCs with the conventional OOK links and PNoCs in terms of performance and energy efficiency in the presence of inter-channel crosstalk. From our link-level and PNoC-level evaluation, we have observed that the 4-PAM EDAC–based variants of photonic links and PNoCs exhibit better performance and energy efficiency compared with the OOK-, 4-PAM SS–, and 4-PAM ODAC–based links and PNoCs. Venkata Sai Praneeth Karempudi, Febin Sunny, Ishan G. Thakkar, Sai Vineel Reddy Chittamuru, Mahdi Nikdast, Sudeep Pasricha |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2021 | Design Exploration and Scalability Analysis of a CMOS-Integrated, Polymorphic, Nanophotonic Arithmetic-Logic UnitabstractOver the past two decades, the clock speed, and hence, the singlecore performance of microprocessors has already stagnated. Following this, the recent faltering of Moore's law due to the CMOS fabrication technology reaching its unavoidable physical limit has presaged daunting challenges for designing power-efficient and ultrafast microprocessors. To overcome these challenges, vigorous efforts have been made to develop new more-than-Moore technologies and architectures for computing. Among these, nanophotonic integrated circuits based computing architectures have shown revolutionary potential. Among recent demonstrations of nanophotonic circuits for computing, a polymorphic, nanophotonic ALU (PoN-ALU) carries a notable importance since it has shown very high flexibility, high speed, and low power consumption for computing. In this paper, we carry out a design space exploration of this PoN-ALU to derive new design guidelines that can help scale the speed and energy efficiency of PoNALU even further. Venkata Sai Praneeth Karempudi, Shreyan Datta, Ishan G. Thakkar |
SenSys | 1 |
| 2020 | Redesigning Photonic Interconnects with Silicon-on-Sapphire Device Platform for Ultra-Low-Energy On-Chip CommunicationabstractTraditional silicon-on-insulator (SOI) platform based on-chip photonic interconnects have limited energy-bandwidth scalability due to the optical non-linearity induced power constraints of the constituent photonic devices. In this paper, we propose to break this scalability barrier using a new silicon-on-sapphire (SOS) based photonic device platform. Our physical-layer characterization results show that SOS based photonic devices have negligible optical non-linearity effects in the mid-infrared region near 4m, which drastically alleviates their power constraints. Our link-level analysis shows that SOS based photonic devices can be used to realize photonic links with aggregated data rate of more than 1 Tb/s, which recently has been deemed unattainable for the SOI-based photonic on-chip links. We also show that such high-throughput SOS-based photonic links can significantly improve the energy-efficiency of on-chip photonic communication architectures. Our system-level analysis results position SOS-based photonic interconnects to pave the way for realizing ultra-low-energy (< 1 pJ/bit) on-chip data transfers. Venkata Sai Praneeth Karempudi, Sairam Sri Vatsavai, Ishan G. Thakkar |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | PROTEUS: Rule-Based Self-Adaptation in Photonic NoCs for Loss-Aware Co-Management of Laser Power and PerformanceabstractThe performance of on-chip communication in the state-of-the-art multi-core processors that use the traditional electronic NoCs has already become severely energy-constrained. To that end, emerging photonic NoCs (PNoC) are seen as a potential solution to improve the energy-efficiency (performance per watt) of on-chip communication. However, existing PNoC designs cannot realize their full potential due to their excessive laser power consumption. Prior works that attempt to improve laser power efficiency in PNoCs do not consider all key factors that affect the laser power requirement of PNoCs. Therefore, they cannot yield the desired balance between the reduction in laser power, achieved performance and energy-efficiency in PNoCs. In this paper, we present PROTEUS framework that employs rule-based self-adaptation in PNoCs. Our approach not only reduces the laser power consumption, but also minimizes the average packet latency by opportunistically increasing the communication data rate in PNoCs, and thus, yields the desired balance between the laser power reduction, performance, and energy-efficiency in PNoCs. Our evaluation with PARSEC benchmarks shows that our PROTEUS framework can achieve up to 24.5% less laser power consumption, up to 31% less average packet latency, and up to 20% less energy-per-bit, compared to another laser power management technique from prior work. Sairam Sri Vatsavai, Venkata Sai Praneeth Karempudi, Ishan G. Thakkar |
NOCS | 2 |