EDBT 2026 Demo / reviewers in the wild / expert
Gaurav Singh 0005
dblp:07/2202-5
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0001-8728-3936ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRIM: Thermal Auto-Compensation for Resistive In-Memory ComputingabstractIn-Memory Computing (IMC) has emerged as one of the most promising architectures to efficiently compute artificial intelligence tasks on hardware, particularly Deep Neural Networks (DNNs). IMC can make use of analog computation principles alongside emerging Non-Volatile Memory (eNVM) technologies, potentially offering several orders of magnitude increased energy efficiency compared to generic processing units. Yet, the use of analog circuitry, potentially integrated with emerging technologies post-processed on top of silicon wafers, increases the susceptibility of hardware to a large spectrum of variations, for instance manufacturing, noise or temperature sensitivity. Hence, this susceptibility can hamper the large-scale deployment of IMC circuits into the market. To tackle the reliability of analog resistive-based IMC circuits regarding temperature variations, this paper presents TRIM, a thermal on-chip auto-compensation method aimed at fully calibrating first-order temperature effects. TRIM is designed to maintain the computational accuracy of IMC cores in DNN applications over a wide temperature range, while being highly scalable and adaptable. In essence, the temperature compensation is realized through a Complementary-To-Absolute-Temperature (CTAT) voltage reference integrated inside a voltage regulator and applied at the zero reference node of a Multiplying Digital-to-Analog Converter (MDAC), eliminating the need for external circuits or look-up tables. The proposed methodology is demonstrated on a proof-of-concept 65 nm CMOS resistive IMC column. Measurement results showcase that the proof-of-concept auto-compensation system significantly enhances inference and Multiply-And-Accumulate (MAC) operation accuracy of any first-order resistive crossbar column, achieving inference accuracy recovery of 100% over a temperature range of -20 ∘C to 60 ∘C and a 91.3 in MAC operation accuracy, with an area overhead of 2% and power overhead of <0.02%. Dipesh C. Monga, Gaurav Singh 0005, Omar Numan, Kazybek Adam, Martin Andraud, Kari Halonen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Urine-Powered Batteryless Sensor Node With Printed Harvesters and Sensors for Smart DiapersabstractAs the global population ages, caregivers encounter significant challenges in monitoring wet diapers and assessing the volumes of voided fluids in adult incontinence products. Recent advancements in smart-diaper technology address these issues, but challenges persist with integrating electronics, frequent removal and reapplication, and managing batteries. This study presents a batteryless urine-powered IoT sensor node using printed energy harvesters and capacitive sensors integrated with an ultra-low-power integrated circuit for disposable smart diapers. Coplanar capacitive sensors and electrochemical energy harvesters are printed on the flexible substrate using environment-friendly materials. An ultra-low power frontend interface is developed, which is powered by the harvested energy from urine. Laboratory measurements validate the functionality of on-chip electronics and the performance of in-diaper sensors and energy harvesters. The system demonstrates effective energy harvesting by maintaining a stable regulated voltage of 1.1 V with urine volumes of 90 ml or more, powering the front-end electronics continuously for 6 hours. The proposed system successfully demonstrated the detection of multiple urination events in the diaper and quantified the voided volume as low as 30 ml. The results demonstrated in this work pave the way for cost-effective, disposable, and environmentally friendly solutions for smart diapers, enhancing both efficiency and comfort for caregivers and elderly individuals. Muhammad Tanweer, Dipesh C. Monga, Gaurav Singh 0005, Liam Gillan, Raimo Sepponen, I. Oguz Tanzer, Kari Halonen |
IEEE Internet Things J. | 3 |
| 2024 | On-chip Built-In Self-Calibration of Thermal Variations for Mixed-Signal In-Memory ComputingabstractIn-memory computing (IMC) accelerators have become a pivotal architecture for enhancing AI algorithm computations, particularly critical for embedding deep neural networks (DNNs) in edge devices. The efficiency of these systems is paramount, yet IMC cores are prone to fluctuations due to process, temperature, and voltage variations, which can detrimentally impact DNN accuracy. This research introduces an innovative Built-In Self-Calibration (BISC) methodology, specifically designed to compensate for temperature-induced variations in mixed-signal IMC cores. The methodology enables real-time, on-chip adjustment of DNN weights during computation within the IMC core without modifying the computation path. The proposed approach, implemented on a silicon prototype, not only maintained DNN computation accuracy under substantial temperature variations but also fully compensated for almost 90% of the offset caused by these variations, without introducing any non-idealities. Gaurav Singh 0005, Omar Numan, Dipesh C. Monga, Martin Andraud, Kari Halonen |
ETS | 1 |
| 2024 | On Hardware-efficient Inference in Probabilistic CircuitsabstractProbabilistic circuits (PCs) offer a promising avenue to perform embedded reasoning under uncertainty. They support efficient and exact computation of various probabilistic inference tasks by design. Hence, hardware-efficient computation of PCs is highly interesting for edge computing applications. As computations in PCs are based on arithmetic with probability values, they are typically performed in the log domain to avoid underflow. Unfortunately, performing the log operation on hardware is costly. Hence, prior work has focused on computations in the linear domain, resulting in high resolution and energy requirements. This work proposes the first dedicated approximate computing framework for PCs that allows for low-resolution logarithm computations. We leverage Addition As Int, resulting in linear PC computation with simple hardware elements. Further, we provide a theoretical approximation error analysis and present an error compensation mechanism. Empirically, our method obtains up to 357{\texttimes} and 649{\texttimes} energy reduction on custom hardware for evidence and MAP queries respectively with little or no computational error. Lingyun Yao, Martin Trapp 0001, Jelin Leslin, Gaurav Singh 0005, Peng Zhang 0028, Karthekeyan Periasamy, Martin Andraud |
UAI | 4 |
| 2024 | A 22-nm All-Digital Time-Domain Neural Network Accelerator for Precision In-Sensor ProcessingabstractDeep neural network (DNN) accelerators are increasingly integrated into sensing applications, such as wearables and sensor networks, to provide advanced in-sensor processing capabilities. Given wearables’ strict size and power requirements, minimizing the area and energy consumption of DNN accelerators is a critical concern. In that regard, computing DNN models in the time domain is a promising architecture, taking advantage of both technology scaling friendliness and efficiency. Yet, time-domain accelerators are typically not fully digital, limiting the full benefits of time-domain computation. In this work, we propose an all-digital time-domain accelerator with a small size and low energy consumption to target precision in-sensor processing like human activity recognition (HAR). The proposed accelerator features a simple and efficient architecture without dependencies on analog nonidealities such as leakage and charge errors. An eight-neuron layer (core computation layer) is implemented in 22-nm FD-SOI technology. The layer occupies$70 \times \,70\,\mu $m while supporting multibit inputs (8-bit) and weights (8-bit) with signed accumulation up to 18 bits. The power dissipation of the computation layer is 576$\mu $W at 0.72-V supply and 500-MHz clock frequency achieving an average area efficiency of 24.74 GOPS/mm2 (up to 544.22 GOPS/mm2), an average energy efficiency of 0.21 TOPS/W (up to 4.63 TOPS/W), and a normalized energy efficiency of 13.46 1b-TOPS/W (up to 296.30 1b-TOPS/W). Ahmed M. Mohey, Jelin Leslin, Gaurav Singh 0005, Marko Kosunen, Jussi Ryynänen, Martin Andraud |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |