Siddharth Joshi 0001

dblp:63/6495-1 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-9201-9678ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Hadamard-Walsh Channelized Receivers: Theory, Implementation, and Applications
abstract
Ultra-wideband (UWB) communication and sensing, known for its high data rate and low latency, has emerged as a prominent technology in internet-of-things (IoT) and 5G communication applications. In this paper, we propose to utilize the Hadamard-Walsh Transformation (HWT) as an efficient and accurate channelization technique, in order to relax the per-channel receiver specifications. We present a comprehensive analysis of HWT and how it enables channelization, which is supported by simulation results. The study demonstrates HWT’s potential for applications such as RF analog-to-digital converter (ADC), correlator, and compressive sensing (CS).
Adyant Balaji, Siddharth Joshi 0001, Gert Cauwenberghs
ISCAS3
2025 CAMEL: Capacitive Analog In-Memory Equalization for RF Signal Processing
abstract
As next-generation wireline and wireless systems are scaled to meet increasing data demands, existing signal processing approaches face significant power and latency challenges. To address these demands, we present CAMEL (Capacitive Analog In-Memory Equalization), a mixed-signal, discrete-time, analog in-memory switched-capacitor finite impulse response (FIR) filter designed in Intel16. Using this filter as a core, we develop a 16-tap antenna-domain I/Q equalizer, with 8-bit accuracy, consuming 90 mW from a 1 V supply, while achieving a data rate of 2 Gbps at a bit error rate (BER) of 10−4in a realistic channel at 18 dB signal-to-noise ratio (SNR). Mismatch analysis and scaling studies indicate that this design can be extended to 12 bit and 48-tap configurations with linear increase in power, while delivering full digital reconfigurability, and datarates exceeding 5 Gbps with a power efficiency of 9.81 pJ/bit.
Md. Shahrul Islam, Kshama Lakshmi Ranganatha, Richard Dorrance, Siddharth Joshi 0001
ISCAS4
2025 SRAM-Based Ring Oscillators as Nonlinear Compute-in-Memory for Low-Power Communication
abstract
This paper presents two designs of digitally controlled ring oscillators (DCRO) using SRAM-inspired cells in a commercially available 22 nm technology node. The proposed design offers a compact, compilable, DCRO capable of rapid transition between digitally selectable frequencies. Two designs are conducted on a commercially available 22 nm technology node, a compact minimal footprint design occupying 25.58μm2area (Design-I) and a scaled up design occupying 50.5μm2(Design-II). Extensive simulations validate low transit time between frequency-states (≤15ns on post-extracted netlists), while offering a tuning range of 700 MHz and 1.736 GHz for Design-I and Design-II respectively. Design-I performs competitively with state-of-the-art ring oscillators, consuming 42.5 μW from a 0.8 V supply while achieving a figure of merit (FoM) of 144.79 dBc/Hz for a 1.05 GHz frequency, While Design-II, consuming − 83.8 μW from a 0.8 V supply demonstrates an of –174.01 dBc/Hz for a 1.92 GHz local oscillator signal. This outperforms existing freerunning oscillators, delivering results competitive with those achieved using a PLL.
Kshama Lakshmi Ranganatha, Md. Shahrul Islam, Sudipto Chakraborty, Siddharth Joshi 0001
ISCAS4
2024 Edge Inference with Fully Differentiable Quantized Mixed Precision Neural Networks
abstract
The large computing and memory cost of deep neural networks (DNNs) often precludes their use in resource-constrained devices. Quantizing the parameters and operations to lower bit-precision offers substantial memory and energy savings for neural network inference, facilitating the use of DNNs on edge computing platforms. Recent efforts at quantizing DNNs have employed a range of techniques en-compassing progressive quantization, step-size adaptation, and gradient scaling. This paper proposes a new quantization approach for mixed precision convolutional neural networks (CNNs) targeting edge-computing. Our method establishes a new Pareto frontier in model accuracy and memory footprint demonstrating a range of pre-trained quantized models, delivering best-in-class accuracy below 4.3 MB of weights and activations without modifying the model architecture. Our main contributions are: (i) a method for tensor-sliced learned precision with a hardware-aware cost function for heterogeneous differentiable quantization, (ii) targeted gradient modification for weights and activations to mitigate quantization errors, and (iii) a multi-phase learning schedule to address instability in learning arising from updates to the learned quantizer and model parameters. We demonstrate the effectiveness of our techniques on the ImageNet dataset across a range of models including EfficientNet-Lite0 (e.g., 4.14 MB of weights and activations at 67.66% accuracy) and MobileNetV2 (e.g., 3.51 MB weights and activations at 65.39% accuracy).
Clemens JS Schaefer, Siddharth Joshi 0001, Raúl Blázquez
WACV2
2023 Micro/Nano Circuits and Systems Design and Design Automation: Challenges and Opportunities
abstract
The field of design and design automation of micro-/nano-circuits and systems has played a pivotal role in advancing information technologies that are an inseparable part of all our lives. Without the fundamental principles and tools created in this field, modern-day electronic systems that form the foundations of today's information age would not be a reality. Though the field has achieved tremendous success in the past few decades, it is now facing some unprecedented challenges, stemming from foundational technologies all the way to new applications. Business-as-usual approaches are plateauing. New, fundamental research and innovation are needed to sustain the demanded growth. This paper aims to summarize the key challenges and future research directions in the field of micro/nano circuits and systems design and design automation.
Gert Cauwenberghs, Jason Cong, Xiaobo Sharon Hu, Siddharth Joshi 0001, Subhasish Mitra, Wolfgang Porod, H.-S. Philip Wong
Proc. IEEE4
2022 Ruby: Improving Hardware Efficiency for Tensor Algebra Accelerators Through Imperfect Factorization
abstract
Finding high-quality mappings of Deep Neural Network (DNN) models onto tensor accelerators is critical for efficiency. State-of-the-art mapping exploration tools use remainderless (i.e., perfect) factorization to allocate hardware resources, through tiling the tensors, based on factors of tensor dimensions. This limits the size of the search space, (i.e., mapspace), but can lead to low resource utilization. We introduce a new mapspace, Ruby, that adds remainders (i.e., imperfect factorization) to expand the mapspace with high-quality mappings for user-defined architectures. This expansion allows us to allocate resources more precisely by generating tile sizes that better conform to hardware resources. However, this mapspace expansion also incurs an increase in the number of unique mappings. Consequently, this paper studies the trade-off between Ruby’s mapspace expansion and mapping quality. We propose Ruby-S (Spatial) to only employ imperfect factorization towards improved parallelism. Ruby-S incurs a moderate mapspace expansion while reducing energy-delay product (EDP) up to 50% when implementing ResNet-50 on an Eyeriss-like architecture with an average improvement of 20%. For the most part, this improvement can be attributed to higher compute utilization. EDP on a Simba-like architecture improves up to 40% with an average of 10%. For DeepBench workloads Ruby-S yields improvements of up to 45% with an average improvement of 10% on an Eyeriss-like architecture. Ruby-S is robust to accelerator configurations and improves EDP by 20% on average, with a maximum improvement of 55% when implementing ResNet-50 on different accelerator configurations. Ruby-S mappings form a new Pareto frontier, improving the performance of previous configurations by an average of 30% and 20% for ResNet-50 and DeepBench workloads respectively.
Mark Horeni, Pooria Taheri, Po-An Tsai, Angshuman Parashar, Joel S. Emer, Siddharth Joshi 0001
ISPASS6
2021 LSTMs for Keyword Spotting with ReRAM-Based Compute-In-Memory Architectures
abstract
The increasingly central role of speech based human computer interaction necessitates on-device, low-latency, low- power, high-accuracy key word spotting (KWS). State-of-the- art accuracies on speech-related tasks have been achieved by long short-term memory (LSTM) neural network (NN) models. Such models are typically computationally intensive because of their heavy use of Matrix vector multiplication (MVM) operations. Compute-in-Memory (CIM) architectures, while well suited to MVM operations, have not seen widespread adoption for LSTMs. In this paper we adapt resistive random access memory based CIM architectures for KWS using LSTMs. We find that a hybrid system composed of CIM cores and digital cores achieves 90% test accuracy on the google speech data set at the cost of 25 uJ/decision. Our optimized architecture uses 5-bit inputs, and analog weights to produce 6-bit outputs. All digital computation are performed with 8-bit precision leading to a 3.7× improvement in computational efficiency compared to equivalent digital systems at that accuracy.
Clemens JS Schaefer, Mark Horeni, Pooria Taheri, Siddharth Joshi 0001
ISCAS4
2020 A Device Non-Ideality Resilient Approach for Mapping Neural Networks to Crossbar Arrays
abstract
We propose a technology-independent method, referred to as adjacent connection matrix (ACM), to efficiently map signed weight matrices to non-negative crossbar arrays. When compared to same-hardware-overhead mapping methods, using ACM leads to improvements of up to 20% in training accuracy for ResNet-20 with the CIFAR-10 dataset when training with 5-bit precision crossbar arrays or lower. When compared with strategies that use two elements to represent a weight, ACM achieves comparable training accuracies, while also offering area and read energy reductions of 2.3× and 7×, respectively. ACM also has a mild regularization effect that improves inference accuracy in crossbar arrays without any retraining or costly device/variation-aware training.
Arman Kazemi, Cristobal Alessandri, Alan C. Seabaugh, Xiaobo Sharon Hu, Michael T. Niemier, Siddharth Joshi 0001
DAC6
2020 A 4.2-pJ/Conv 10-b Asynchronous ADC with Hybrid Two-Tier Level-Crossing Event Coding
abstract
An asynchronous continuous-time level-crossing analog-to-digital converter (LC-ADC) for high-throughput, high-resolution applications is presented. The proposed 10-bit ADC architecture comprises two stages of level-crossing ADCs, the first stage resolving for 5 MSBs and the second folded residue stage for 5 LSBs. Gray encoding of the output bits ensure single-bit transitions between adjacent digital outputs. Compared to uniform-sampling synchronous ADCs, LC-ADCs generate fewer samples for sparse signals, useful in many applications for biomedical signal acquisition, event-driven computer vision, etc. Unlike conventional LC-ADCs with a few comparators tuned for lower power consumption to acquire sparse signals, this two-tier LC-ADC is optimized for high-resolution tracking of continuous signals, like Electrocardiogram (ECG). Designed and fabricated in 0.18-μm CMOS technology, chip area of the proposed ADC is 1310 × 125 μm2. Operating at 1.8 V supply, the ADC consumes 160-426 μW for 1 Hz to 200 kHz input frequencies at full scale amplitude and achieves an energy efficiency figure-of-merit of 4.16-pJ/conv.
Rajkumar Kubendran, Jongkil Park 0001, Ritvik Sharma, Chul Kim, Siddharth Joshi 0001, Gert Cauwenberghs, Sohmyung Ha
ISCAS5
2020 Memory Organization for Energy-Efficient Learning and Inference in Digital Neuromorphic Accelerators
abstract
The energy efficiency of neuromorphic hardware is greatly affected by the energy of storing, accessing, and updating synaptic parameters. Various methods of memory organisation targeting energy-efficient digital accelerators have been investigated in the past, however, they do not completely encapsulate the energy costs at a system level. To address this shortcoming and to account for various overheads, we synthesize the controller and memory for different encoding schemes and extract the energy costs from these synthesized blocks. Additionally, we introduce functional encoding for structured connectivity such as the connectivity in convolutional layers. Functional encoding offers a 58% reduction in the energy to implement a backward pass and weight update in such layers compared to existing index-based solutions. We show that for a 2 layer spiking neural network trained to retain a spatio-temporal pattern, bitmap (PB-BMP) based organization can encode the sparser networks more efficiently. This form of encoding delivers a 1.37× improvement in energy efficiency coming at the cost of a 4% degradation in network retention accuracy as measured by the van Rossum distance.
Clemens JS Schaefer, Patrick Faley, Emre Neftci, Siddharth Joshi 0001
ISCAS4
2020 Embedding error correction into crossbars for reliable matrix vector multiplication using emerging devices
abstract
Emerging memory devices are an attractive choice for implementing very energy-efficient in-situ matrix-vector multiplication (MVM) for use in intelligent edge platforms. Despite their great potential, device-level non-idealities have a large impact on the application-level accuracy of deep neural network (DNN) inference. We introduce a low-density parity-check code (LDPC) based approach to correct non-ideality induced errors encountered during in-situ MVM. We first encode the weights using error correcting codes (ECC), perform MVM on the encoded weights, and then decode the result after in-situ MVM. We show that partial encoding of weights can maintain DNN inference accuracy while minimizing the overhead of LDPC decoding. Within two iterations, our ECC method recovers 60% of the accuracy in MVM computations when 5% of underlying computations are error-prone. Compared to an alternative ECC method which uses arithmetic codes, using LDPC improves AlexNet classification accuracy by 0.8% at iso-energy. Similarly, at iso-energy, we demonstrate an improvement in CIFAR-10 classification accuracy of 54% with VGG-11 when compared to a strategy that uses 2× redundancy in weights. Further design space explorations demonstrate that we can leverage the resilience endowed by ECC to improve energy efficiency (by reducing operating voltage). A 3.3× energy efficiency improvement in DNN inference on CIFAR-10 dataset with VGG-11 is achieved at iso-accuracy.
Qiuwen Lou, Tianqi Gao, Patrick Faley, Michael T. Niemier, Xiaobo Sharon Hu, Siddharth Joshi 0001
ISLPED6
2017 Memristor for computing: Myth or reality?
abstract
CMOS technology and its sustainable scaling have been the enablers for the design and manufacturing of computer architectures that have been fuelling a wider range of applications. Today, however, both the technology and the computer architectures are suffering from serious challenges/ walls making them incapable to deliver the right computing power at pre-defined constraints. This motivates the need of exploring new architectures and new technologies; not only to maintain the economic benefit of scaling, but also to enable the solutions of emerging computer power and data storage hungry applications such as big-data and data-intensive applications. This paper discusses the emerging memristor device as complementary (or alternative) to CMOS device and shows how this device can enable new ways of computing that will at least solve the challenges of today's architectures for some applications. The paper shows not only the potential of memristor devices in enabling new memory technologies and new logic design styles, but also their potential in enabling memory intensive architectures as well as neuromorphic computing due to their unique properties such as the tight integration with CMOS and the ability to learn and adapt.
Said Hamdioui, Shahar Kvatinsky, Gert Cauwenberghs, Lei Xie 0005, Nimrod Wald, Siddharth Joshi 0001, Hesham Mostafa Elsayed, Henk Corporaal, Koen Bertels
DATE6
2017 Hierarchical Address Event Routing for Reconfigurable Large-Scale Neuromorphic Systems
abstract
We present a hierarchical address-event routing (HiAER) architecture for scalable communication of neural and synaptic spike events between neuromorphic processors, implemented with five Xilinx Spartan-6 field-programmable gate arrays and four custom analog neuromophic integrated circuits serving 262k neurons and 262M synapses. The architecture extends the single-bus address-event representation protocol to a hierarchy of multiple nested buses, routing events across increasing scales of spatial distance. The HiAER protocol provides individually programmable axonal delay in addition to strength for each synapse, lending itself toward biologically plausible neural network architectures, and scales across a range of hierarchies suitable for multichip and multiboard systems in reconfigurable large-scale neuromorphic systems. We show approximately linear scaling of net global synaptic event throughput with number of routing nodes in the network, at $3.6\times 10^{7}$ synaptic events per second per 16k-neuron node in the hierarchy.
Jongkil Park 0001, Theodore Yu, Siddharth Joshi 0001, Christoph Maier, Gert Cauwenberghs
IEEE Trans. Neural Networks Learn. Syst.3
2016 A 6μW/MHz charge buffer with 7fF input capacitance in 65nm CMOS for non-contact electropotential sensing
abstract
Capacitive non-contact electric field sensing is a prime modality for signal detection and communication in a variety of contexts ranging from bio-potential measurements and proximity sensing to body centered communication and human computer interaction. Although recent developments in the space of wearable sensors have greatly expanded the sensory capability, they simultaneously place more stringent requirements on the front-end. Thus, low-noise power efficient front-ends that can demonstrate low input parasitic capacitances can accelerate the adoption of sensing and communication technologies that exploit this form of capacitive coupling. To this end, we present results from a 193μm2, 6μW/MHz, unity-gain, charge buffer, fabricated in 65nm CMOS for use in electric field sensing.
Siddharth Joshi 0001, Chul Kim, Gert Cauwenberghs
ISCAS1
2012 Live demonstration: Hierarchical Address-Event Routing architecture for reconfigurable large scale neuromorphic systems
abstract
Recent advances in neuromorphic engineering for brain-like computing and neural prostheses are converging towards realization of electronic synaptic arrays approaching the integration density and energy efficiency of the human brain. A major impediment in this development is the real-time synaptic routing in a large-scale spiking neuron architecture. Here we present a hierarchical address-event routing (HiAER) communication architecture for routing neural events in a scaleable reconfigurable large-scale neuromorphic system. The neural events are routed in real-time through synaptic connections with configurable parameters governing connectivity, synaptic strength, and axonal delay. The HiAER architecture is implemented on a hardware platform with five Xilinx Spartan-6 FPGA cores.
Jongkil Park 0001, Theodore Yu, Christoph Maier, Siddharth Joshi 0001, Gert Cauwenberghs
ISCAS4