EDBT 2026 Demo / reviewers in the wild / expert
Sudhakar Pamarti
dblp:24/5690
· DBLP profile ↗
25ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-1457-7508ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frequency Channelized Filtering by Aliasing
Anmol Kachroo, Sudhakar Pamarti |
ISCAS | 2 |
| 2025 | A Comparative Analysis of Low Temperature and Room Temperature Circuit OperationabstractLow-temperature (LT) conditions can potentially lead to lower power consumption and enhanced performance in circuit operations by reducing the transistor leakage current, increasing carrier mobility, reducing wear-out, and reducing interconnect resistance. We develop PROCEED-LT, a pathfinding framework to co-optimize devices and circuits over a wide performance range. Our results demonstrate that circuit operations at LT (−196 °C) reduce power compared to room temperature (RT, 85 °C) by$15\times $to over$23.8\times $depending on performance level. Alternatively, LT improves performance by$2.4\times $(high-power, high-performance)$- 7.0\times $(low-power, low-performance) at the same power point. These gains are further improved in low-activity circuits and when using multivoltage configurations. Meanwhile, we highlight the need for improvement in$V_{\text {th}}$variation to leverage benefits at cryogenic temperatures. Ali H. Hassan, Rhesa Muhammad Ramadhan, Yingheng Li, Chih-Kong Ken Yang, Sudhakar Pamarti, Puneet Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | SCIMITAR: Stochastic Computing In-Memory In-Situ Tracking ARchitecture for Event-Based CamerasabstractEvent-based cameras offer low latency and high-dynamic range imaging data in a sparse format that is well-suited for high-speed object tracking. Processing this sparse data in the same way as traditional camera data requires a great deal of unnecessary computation, making it difficult to take advantage of the high-effective frame rate for real-time processing. In this work, we propose an accelerator for high-speed object tracking on event-based camera data. SCIMITAR combines digital in-memory stochastic computing, in-situ stochastic stream generation, and multiple optimizations for utilizing input sparsity. SCIMITAR provides unparalleled performance with latency and energy that scale with sparsity. We demonstrate SCIMITAR performance on an object tracking application using circuit-level simulations of custom-designed compute-in-memory (CIM) macros and digital circuits. We achieve a frame processing rate of 26k frames/s with 100 regions-of-interest per frame and equivalent or better than state-of-the-art tracking accuracy. The accelerator achieves a peak throughput of 71 TOP/S and energy efficiency of 733 to 1702 TOP/S/W demonstrated on a range of event-based vision datasets, which is$5\times $higher than other CIM solutions. Wojciech Romaszkan, Jiyue Yang, Alexander Graening, Vinod Kurian Jacob, Jishnu Sen, Sudhakar Pamarti, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Digital Residual Alias Cancellation for Filtering-by-Aliasing ReceiversabstractThe filtering-by-aliasing (FA) receivers have demonstrated sharp analog finite impulse response (FIR) filtering by combining periodically time-varying (PTV) circuit elements with uniform sampling. Seen at the sampled output, the impulse response of an FA receiver can be controlled by the PTV resistor’s resistance that varies over time, which realizes very sharp filtering with 50–70-dB stopband rejection. However, the finite stopband rejection and the inherent sampling in FA result in unwanted in-band residual blocker aliases, which cannot be suppressed further by downstream linear time-invariant (LTI) stages, unlike in a conventional receiver, and the overall blocker rejection may be insufficient in some applications. This paper describes a digital residual alias cancellation technique tailored for the FA receivers. By using a second channel to capture the blocker and equalizing the blocker in the FA and the second channels with digital baseband filters, measurement results, with the cancellation algorithm implemented in MATLAB, show that the residual blocker aliases can be suppressed by another ~15 dB, in addition to the FA analog filtering. Shi Bu, Vinod Kurian Jacob, Sudhakar Pamarti |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | FPGA Crystal Oscillator Circuit Emulation Based on Wave Digital FilterabstractThe design cycle of analog and mixed signal (AMS) components requires the designer to iteratively perform analog simulations, layout, fabrication, and hardware testing. Unlike digital designs, system verification is a difficult task in analog designs, primarily due to a lack of emulation. Thus, a method to emulate AMS components on digital hardware would be highly beneficial. In this work, a high-quality factor crystal oscillator circuit is implemented on a Xilinx Vertex 7 field-programmable gate array (FPGA) using the wave digital filter (WDF) based model with a nonlinear lookup table for modeling transistor characteristics. The number of required hardware resources was minimized while ensuring that the accuracy of the emulation shows an almost perfect match with the SPICE simulations. The WDF model was designed with a tree structure so that it only requires 32 clock cycles to compute a complete sample. The resulting emulation computes a sample at 18.75 MHz while running on an FPGA with a 600 MHz clock. Abdulaziz Alshaya, Sudhakar Pamarti, Christos Papavassiliou |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | REX-SC: Range-Extended Stochastic Computing Accumulation for Neural Network AccelerationabstractDeep learning has grown in capability and size in recent years, prompting research on alternative computing methods to cope with the increased compute cost. Stochastic computing (SC) promises higher compute efficiency with its compact compute units, but accuracy issues have prevented wide adoption, and accuracy-improving techniques have sacrificed runtime or training performance. In this work, we propose extended range SC—Range-Extended SC Accumulation to deal with the accuracy issues of SC. By modifying the functionality of OR-based SC accumulation, we increase SC computation accuracy without sacrificing the performance benefits. Our approach achieves a$2\times $reduction in stream length for the same accuracy compared to SC with OR-based accumulation and an up to$3.6\times $improvement in energy compared to SC with binary addition. With proper modeling, our approach improves training performance for SC-based neural networks and makes training SC models practical for large datasets like ImageNet. Tianmu Li, Wojciech Romaszkan, Sudhakar Pamarti, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | A Digital Alias Cancellation Technique for Filtering-by-Aliasing ReceiversabstractThere has been a growing trend of developing programmable transceivers, which are the key to realizing a true software-defined radio [1]. Some prior works, such as N-path filters and mixer-first receivers, utilized periodically time-varying (PTV) circuits and have shown some promises and achieved what conventional time-invariant circuits are incapable of. Among them, the filtering-by-aliasing (FA) receivers [2], [3] have demonstrated good potential by showcasing sharp bandpass filters with good linearity and programmability. Shi Bu, Vinod Kurian Jacob, Sudhakar Pamarti |
ISCAS | 3 |
| 2022 | A 14-bit 1-GS/s SiGe Bootstrap Sampler for High Resolution ADC with 250-MHz InputabstractAn 86.6-dB SFDR, 1-GS/s differential bootstrap sampler in a 0.18-um SiGe BiCMOS technology is presented. The performance is achieved using an amplitude-modulated bootstrap circuit. The results show 14-bit linearity over nearly 500-MHz bandwidth, while consuming less power compared to a conventional MOSFET switched-capacitor bootstrap circuit due to less parasitic capacitance and the use of high ftHBT. Jiazhang Song, Li-Yang Chen, Mau-Chung Frank Chang, Sudhakar Pamarti, Chih-Kong Ken Yang |
ISCAS | 4 |
| 2021 | Designing a 2048-Chiplet, 14336-Core Waferscale ProcessorabstractWaferscale processor systems can provide the large number of cores, and memory bandwidth required by today’s highly parallel workloads. One approach to building waferscale systems is to use a chiplet-based architecture where pre-tested chiplets are integrated on a passive silicon-interconnect wafer. This technology allows heterogeneous integration and can provide significant performance and cost benefits. However, designing such a system has several challenges such as power delivery, clock distribution, waferscale-network design, design for testability and fault-tolerance. In this work, we discuss these challenges and the solutions we employed to design a 2048-chiplet, 14,336-core waferscale processor system. Saptadeep Pal, Irina Alam, Nick Cebry, Haris Suhail, Shi Bu, Subramanian S. Iyer, Sudhakar Pamarti, Rakesh Kumar 0002, Puneet Gupta 0001 |
DAC | 8 |
| 2021 | GEO: Generation and Execution Optimized Stochastic Computing Accelerator for Neural NetworksabstractStochastic computing (SC) has seen a renaissance in recent years as a means for machine learning acceleration due to its compact arithmetic and approximation properties. Still, SC accuracy remains an issue, with prior works either not fully utilizing the computational density or suffering from significant accuracy losses. In this work, we propose GEO - Generation and Execution Optimized Stochastic Computing Accelerator for Neural Networks, which optimizes stream generation and execution components of SC, and bridges the accuracy gap between stochastic computing and fixed-point neural networks. It improves accuracy by coupling controlled stream sharing with training and balancing OR and binary accumulations. GEO further optimizes the SC execution through progressive shadow buffering and architectural optimizations. GEO can improve accuracy compared to state-of-the-art SC by 2.2-4.0% points while being up to 4.4X faster and 5.3X more energy efficient. GEO eliminates the accuracy gap between SC and fixed-point architectures while delivering up to 5.6X higher throughput and 2.6X lower energy. Tianmu Li, Wojciech Romaszkan, Sudhakar Pamarti, Puneet Gupta 0001 |
DATE | 3 |
| 2020 | ACOUSTIC: Accelerating Convolutional Neural Networks through Or-Unipolar Skipped Stochastic ComputingabstractAs privacy and latency requirements force a move towards edge Machine Learning inference, resource constrained devices are struggling to cope with large and computationally complex models. For Convolutional Neural Networks, those limitations can be overcome by taking advantage of enormous data reuse opportunities and amenability to reduced precision. To do that however, a level of compute density unattainable for conventional binary arithmetic is required. Stochastic Computing can deliver such density, but it has not lived up to its full potential because of multiple underlying precision issues. We present ACOUSTIC: Accelerating Convolutions through Or-Unipolar Skipped sTochastIc Computing, an accelerator framework that enables fully stochastic, high-density CNN inference. Leveraging split-unipolar representation, OR-based accumulation and novel computation-skipping approach, ACOUSTIC delivers server-class parallelism within a mobile area and power budget - a 12mm2accelerator can be as much as 38.7x more energy efficient and 72.5x faster than conventional fixed-point accelerators. It can also be up to 79.6x more energy efficient than state-of-the-art stochastic accelerators. At the lower-end ACOUSTIC achieves 8x-120X inference throughput improvement with similar energy and area when compared to recent mixed-signal/neuromorphic accelerators. Wojciech Romaszkan, Tianmu Li, Tristan Melton, Sudhakar Pamarti, Puneet Gupta 0001 |
DATE | 4 |
| 2016 | Wave digital filter based analog circuit emulation on FPGAabstractUnlike well accepted FPGA emulation for digital circuits, there is no winning emulation solution for analog and mixed-signal (AMS) circuits. This paper presents an analog circuit emulation based on wave digital filters (WDFs), which covers the entire flow of transforming an AMS circuit from SPICE netlist to hardware implementation in FPGA. More specifically, it presents the theoretical support of how to map linear and nonlinear circuit components to WDF. The detail implementation of each WDF component in FPGA is not elaborated due to the page limit. Experiments show that there is a virtually perfect match between FPGA emulation and HSPICE simulations on two small but representative analog circuits, indicating high accuracy of the proposed emulation, and the FPGA-based WDF emulation can process analog signal sampled at as high as 512KHz, which is adequate for a variety of biomedical sensing applications. Yen-Lung Chen, Chien-Nan Jimmy Liu, Jing-Yang Jou, Sudhakar Pamarti, Lei He 0001 |
ISCAS | 6 |
| 2015 | Toward Wave Digital Filter based Analog Circuit Emulation on FPGA (Abstract Only)abstractSoftware simulation of analog and mixed-signal circuits often takes a long computing time. Unlike digital circuits that can be validated by FPGA emulation, there is no winning emulation solution for analog circuits. As the first step to applying wave digital filter (WDF) to emulate post-layout analog circuits, we present how to map linear and nonlinear components in an original circuit to WDFs with exactly same behaviors. To validate, we implement the emulation circuit (i.e., WDFs) in FPGA. To be more specific, each emulation time step is executed as a finite state machine, while all the computing resource, e.g. floating point units (FPU), are shared as a resource pool and used only when it is necessary, which result in a very small resource consumption on FPGA. Virtually perfect match is obtained between the Verilog and SPICE simulations for a number of primitive analog circuits, indicating the high accuracy of the proposed emulation. In terms of runtime, the WDF implementation is about 3-4x faster than HSPICE on a small two-stage differential amplifier circuit. And better speedup can be anticipated when it scales to larger circuits because of the underlying binary tree structure of the WDF implementation. Yen-Lung Chen, Chien-Nan Jimmy Liu, Sudhakar Pamarti, Lei He 0001 |
FPGA | 5 |
| 2015 | Frequency-domain analysis of a mixer-first receiver using conversion matricesabstractThe analysis of a mixer-first receiver using conversion matrices is presented. Conversion matrices provide a systematic approach to analyze linear periodically time-varying (LPTV) circuits. Using conversion matrices of LPTV components in a frequency domain equivalent circuit allows analysis similar to a linear time invariant (LTI) circuit. For example, Ohm's law, Kirchhoff's voltage and current laws, impedance combination rules, etc., can all be used in such equivalent circuits. On applying this method to a mixer-first receiver, a common LPTV circuit, results already established in prior art are reproduced accurately. Further, effects of a few important non-idealities, such as clock overlaps, imperfect clock edges and parasitic components that were not considered previously, are also calculated and verified. Sameed Hameed, Mansour Rachid, Babak Daneshrad, Sudhakar Pamarti |
ISCAS | 4 |
| 2013 | Filtering of subtractive discrete dither in quantizers: Some new resultsabstractSubtractive dither in quantizers is examined as a means to mitigate quantizer non-linearity. The effects of filtering the dither signal to shape its spectral content outside the signal band while maintaining its benefits are studied in detail. Design strategies for finite impulse response (FIR) filters that accomplish spectral shaping as well as allay quantizer non-linearity are derived theoretically. Simulation results for low/medium resolution quantizers are presented to validate the derived conditions on the filter structures. Abhishek Ghosh, Sudhakar Pamarti |
ICASSP | 2 |
| 2013 | Adaptive signal conditioning algorithms to enable wideband signal digitizationabstractData converters are an essential component in any communication system link. Time-interleaving several data converters is an effective way to enhance the overall bandwidth of the data-converter. However, mismatches between individual channels in time-interleaved data-converters manifest as a lower available dynamic range. A novel adaptive signal-conditioning technique is proposed to correct for the channel-mismatch errors with minimal hardware requirements. Behavioral simulations corroborating the same are presented. Abhishek Ghosh, Sudhakar Pamarti |
ICC | 2 |
| 2013 | Mitigating timing errors in time-interleaved ADCs: A signal conditioning approachabstractNovel techniques based on signal-conditioning are presented to mitigate timing errors in time-interleaved ADCs. A theoretical bound on the achievable spurious signal content, on applying the techniques, is also derived. Behavioral simulations corroborating the same are presented. Abhishek Ghosh, Sudhakar Pamarti |
ISCAS | 2 |
| 2012 | Robustness of xampling-based RF receivers against analog mismatchesabstractThe analog imperfections in RF direct conversion receiver, of which I/Q imbalance is a major detriment, are examined. Existing literatures on I/Q imbalance compensation try to compensate for the imbalance by estimating the mismatches. In this work, xampling-based decoding algorithm is examined. This algorithm is shown to be very efficient in handling the impairments in the analog components. Simulation results are also presented to illustrate some of the benefits of the proposed approach. Chiranjib Choudhuri, Abhishek Ghosh, Urbashi Mitra, Sudhakar Pamarti |
ICASSP | 4 |
| 2012 | Worst-Case Estimation for Data-Dependent Timing Jitter and Amplitude Noise in High-Speed Differential LinkabstractDifferential signaling has been widely used in high-speed interconnects. Signal integrity issues, such as inter-symbol interference (ISI) and crosstalk between the differential pair, however, still cause significant timing jitter and amplitude noise and heavily limit the performance of the differential link. The pre-emphasis filter is commonly used to reduce ISI but may potentially change the crosstalk behavior. In this paper, we first propose formula-based jitter and noise models considering the combined effect of ISI, crosstalk, and pre-emphasis filter. With the same set of input patterns, experiment shows our models achieve within 5% difference compared with SPICE simulation. By utilizing these formula-based models, we then develop algorithms to directly find out the input patterns for worst-case jitter and worst-case amplitude noise through pseudo-Boolean optimization (PBO) and mathematical programming. In addition, a heuristic algorithm is proposed to further reduce runtime. Experiments show our algorithms obtain more reliable worst-case jitter and noise compared with pseudorandom bit sequences simulation and, meanwhile, reduce runtime by 25× when using a general PBO solver and by 150× when using our proposed heuristic algorithm. Wei Yao 0002, Yiyu Shi 0001, Lei He 0001, Sudhakar Pamarti |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | A progammable baseband anti-alias filter for a passive-mixer-based, SAW-less, multi-band, multi-mode WEDGE transmitterabstractA programmable baseband anti-alias filter (AAF) for a passive-mixer-based, 1.8V, SAW-less, multi-band, multi-mode WEDGE (WCDMA/HSUPA/EGPRS) cellular transmitter (TX) is described. This paper presents an AAF which results in ultra-low, -170dBc/Hz, receive-band noise, and enables the first single-mixer-based, multi-mode, SAW-less TX. By providing the noise and linearity performance only on-demand from system requirements, the presented AAF enables power savings of 14mW, or 34% of the total 41mW TX power. The AAF has been fabricated as part of a single-chip 0.13μm CMOS transceiver. Sandeep D'Souza, Mau-Chung Frank Chang, Sudhakar Pamarti, Bipul Agarwal, Hossein Zarei, Tirdad Sowlati, Roc Berenguer |
ISCAS | 3 |
| 2011 | A novel reconfigurable alias interference cancellation technique for A-to-D conversionabstractWe propose a novel technique for the controlled suppression of aliasing interferers during analog-to-digital conversion. The technique exploits the aliasing inherent in the sampling operation to perform interference cancellation. This is achieved by optimally spreading and weighting the aliasing bands in the frequency domain prior to the analog-to-digital converter (ADC). The resulting filtering response scales automatically with the ADC bandwidth and can be digitally reconfigured to target specific interference profiles to maximize the signal-to-interference-and-noise ratio (SINR). The proposed technique eliminates the need for complex reconfigurable filters in a varying spectral content scenario while maintaining minimal ADC bandwidth and resolution requirements. Mansour Rachid, Sudhakar Pamarti, Babak Daneshrad |
ISCAS | 2 |
| 2010 | A power amplifier with minimal efficiency degradation under back-offabstractA zero voltage switching technique that can maintain its peak efficiency over a 12dB dynamic range of output power is presented. The application of the technique is illustrated by the design and simulation based verification of an 800 MHz, 130nm CMOS PA that achieves close to a constant drain efficiency of about 40% over a 12 dB dynamic range of output power. System level simulations show that the PA achieves an ACPR of -54dBc at 400 KHz and -60dBc at 600 KHz frequency offsets respectively. Nitesh Singhal, Nitin Nidhi, Sudhakar Pamarti |
ISCAS | 3 |
| 2009 | Joint design-time and post-silicon optimization for digitally tuned analog circuitsabstractJoint design time and post-silicon optimization for analog circuits has been an open problem in literature because of the complex nature of analog circuit modeling and optimization. In this paper we formulate the co-optimization problem for digitally tuned analog circuits to optimize the parametric yield, subject to power and area constraints. A general optimization framework combing the branch-and-bound algorithm and gradient ascent method is proposed. We demonstrate our framework with two examples in high-speed serial link, the transmitter design and the phase-locked-loop (PLL) design. Simulation results show that compared with the design heuristic from analog designers' perspective, joint design-time and post-silicon optimization can improve the yield by up to 47% for transmitter design and up to 56% for PLL design under the same area and power constraints. To the best of the authors' knowledge, this is the first in-depth study on yield-driven analog circuit design technique that optimizes post-silicon tuning together with the design-time optimization. Wei Yao 0002, Yiyu Shi 0001, Lei He 0001, Sudhakar Pamarti |
ICCAD | 4 |
| 2007 | A Theoretical Analysis of Split Delta-Sigma ADCsabstractThe recently proposed split analog-to-digital converter (ADC) owes its advantages to the quantization noise from its constituent analog delta-sigma modulators being uniformly distributed, white, and uncorrelated with the input and with each other. A theoretical analysis of the statistics of the quantization noise is presented in the context of a generic split-ADC. Sufficient conditions are derived that ensure that the quantization noise has the aforementioned properties; the application of the conditions is illustrated in the case of two example split-ADCs Sudhakar Pamarti |
ISCAS | 1 |
| 2006 | Power-efficient pulse width modulation DC/DC converters with zero voltage switching controlabstractThis paper proposes a power-efficient PWM DC/DC converter design with a novel zero voltage switching (ZVS) control technique. The ZVS control is realized by an inner feedback loop which is implemented by simple digital circuitry between the input and output of the power transistors and achieves real-time zero voltage switching (ZVS) for various loading and device parameters with power efficiencies over 90.0%. In addition, an outer feedback loop is used to ensure that the output precisely tracks a reference voltage level. We have also built the relationship between the output voltage ripple and the speed of the voltage comparators which has shown to introduce new low-frequency signals to the loops and cause significant output voltage ripples. Experiment results show that the output ripple could be reduced by 4x by carefully handling the generation and propagation of these low frequency signals. Changbo Long, Sasank Reddy, Sudhakar Pamarti, Lei He 0001, Tanay Karnik |
ISLPED | 3 |