EDBT 2026 Demo / reviewers in the wild / expert
Aarthy Mani
dblp:263/0870
· DBLP profile ↗
8ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0002-6159-6974ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | 1.63 pJ/SOP Neuromorphic Processor With Integrated Partial Sum Routers for In-Network ComputingabstractNeuromorphic computing is promising to achieve unprecedented energy efficiency by emulating the human brain’s mechanism. Conventional neuromorphic accelerators employ split-and-merge method to map spiking neural networks’ inputs to surpass the fan-in capabilities of a single neuron core. However, this approach gives rise to the risk of accuracy compromise and extra core usage for the merging process. Moreover, it requires excessive data movement and clock cycles to aggregate spikes generated by partial sums instead of total sums obtained from different cores with substantial power and energy overhead. This work presents a novel approach to addressing the challenges imposed by the split-and-merge method. We propose an energy-efficient, reconfigurable neuromorphic processor that leverages several key techniques to mitigate the above issues. First, we introduce a partial sum router circuitry that enables in-network computing (INC), eliminating the need for extra merge cores. Second, we adopt software-defined Networks-on-Chip (NoCs) by leveraging predefined, efficient routing, eliminating power-hungry routing computation. At last, we incorporate fine-grained power gating and clock gating techniques for further power reduction. Experimental results from our test chip demonstrate the lossless mapping of the algorithm and exceptional energy efficiency, achieving an energy consumption of 1.63 pJ/SOP at 0.48 V. This energy efficiency represents a 22.4% improvement compared to the state-of-the-art results. Our proposed neuromorphic processor provides an efficient and flexible solution for neural network processing, mitigating the limitations of the traditional split-and-merge approach while delivering superior energy efficiency. Dongrui Li, Ming Ming Wong, Yi Sheng Chong, Jun Zhou 0014, Mohit Upadhyay, Ananta Narayanan Balaji, Aarthy Mani, Weng-Fai Wong, Li-Shiuan Peh, Anh-Tuan Do, Bo Wang 0020 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2023 | A 129.83 TOPS/W Area Efficient Digital SOT/STT MRAM-Based Computing-In-Memory for Advanced Edge AI ChipsabstractThis paper proposes a spin-orbit torque (SOT) magnetoresistive random access memory (MRAM)-based digital computing in memory (CIM) structure for advanced CIM edge AI chips. To avoid frequent data reloading and reduce the area overhead caused by the large transistor in the write path, 11 transistors and 4 shared heavy metal (HM) SOT MRAM bitcell is proposed. It contains 4b weight in a single cell in a compact area able to hold a large capacity with reduced latency. Compared to the previous SRAM+NOR digital CIM design, the SOT/STT MRAM CIM designs occupy only 40% and 30% bitcell area respectively. Additionally, SOT MRAM has a 1.16% leakage current of SRAM at the TT corner at room temperature. The proposed design is verified using statistic simulations in 28nm technology. It achieves 129.83 TOPS/W at 1b/1b/8b precision. Lu Lu 0013, Aarthy Mani, Anh-Tuan Do |
ISCAS | 2 |
| 2023 | 1.7pJ/SOP Neuromorphic Processor with Integrated Partial Sum Routers for In-Network ComputingabstractConventional neuromorphic accelerators primarily leverage split-merge method to accommodate a neural network that is beyond a single core's size, leading to possible accuracy loss, extra core usage and significant power and energy overhead. This work presents an energy-efficient, reconfigurable neuro-morphic processor to address the problem by (i) a partial sum router circuitry that enables in-network computing to remove the need of extra merge cores; (ii) software-defined Networks-on-Chip that eliminates the power-hungry routing compute and (iii) fine-grained power gating and clock gating technique for power reduction. Our test chip achieves lossless mapping as the algorithm and an energy efficiency of 1.7pJ/SOP at 0.5V, 19% lower than state-of-the-art result. Bo Wang 0020, Ming Ming Wong, Dongrui Li, Yi Sheng Chong, Jun Zhou 0014, Weng-Fai Wong, Li-Shiuan Peh, Aarthy Mani, Mohit Upadhyay, Ananta Narayanan Balaji, Anh-Tuan Do |
ISCAS | 8 |
| 2023 | Stability Analysis of 6T SRAM at Deep Cryogenic Temperature for Quantum Computing ApplicationsabstractCMOS circuits operating at cryogenic temperature are gaining interest as one of the most promising approaches to efficiently scale up quantum processors in near- and medium future. However, there are major challenges such as (1) strict power dissipation limit at 4K plate due to the limited cooling power of the dilution fridge and (2) significant shifts in CMOS device behavior (i.e. variations, threshold voltage, charge carrier mobility and sub-threshold slope) which are not accurately captured in the standard BSIM models from the foundries. Although there have been extensive works experimentally characterizing and analyzing CMOS transistors at ∼4 K, there is a lack of digital and memory subsystem study. Since on-chip SRAM is one of the most power-consuming and the most vulnerable element in cryogenic SoC, this work analyzes the stability of low-voltage 6T-SRAM at deep cryogenic temperature (i.e 77K and 8K), in comparison with 300K operation. Our DC analysis showed that in general, Write static noise margins of the SRAM cell improves when temperature changes from 300K to 8K, even at low-voltage condition. Regarding the Read static noise margin, our simulation showed that the inverters exhibit pseudo-static hysteresis and interestingly this leads to an improvement of read static noise margin of the cell, similar to what observed in a Schmitt-Trigger SRAM. These results suggest that although CMOS transistors exhibit higher threshold voltage in cryogenic temperature, it is still possible to operate the SRAM at low-voltage for power saving in quantum computing applications. Seong-Beom Kim, Aarthy Mani, Leong Xu Heng Victor, Yuanjin Zheng, Anh-Tuan Do |
ISCAS | 2 |
| 2023 | 1V, 1.13μm pixel pitch Liquid Crystal Driver with Charge-Balancing Scheme for SLM ApplicationsabstractThis work proposes a compact 9T SRAM-based pixel design for low-voltage and high-speed modulator for spatially varying modulation of light (i.e. SLM). To reduce the supply voltage to the CMOS pixel backplane, the operating point is shifted towards the linear window by dynamically pulsing both the top & bottom electrodes in each pixel. A test chip in a standard 40nm CMOS technology was implemented, supporting upto 90 frames/second VGA display. Our testing results demonstrated that the proposed HCS switching scheme is efficient in achieving optical modulation up to 8-bit resolution, making it a viable candidate for state of the art SLMs applications with high frame rate and low power requirements. Aarthy Mani, Chong Yi Sheng, Rasna Maruthiyodan Veetil, Moitra Parikshit, Tobias Wilhelm W. Mass, Chong Ser Choong, Xuewu Xu, Ramon José Paniagua Domínguez, Arseniy I. Kuznetsov, P. Krishna, P. Keyi, Kevin Tshun Chuan Chai, Anh-Tuan Do |
ISCAS | 1 |
| 2023 | LAXOR: A Bit-Accurate BNN Accelerator with Latch-XOR Logic for Local ComputingabstractBinary Neural Network (BNN) accelerators are attractive solutions for Artificial Internet-of-Things (AIoT) applications thanks to the compact models and low computational cost while maintaining satisfactory classification performance. Various analog/mix-signal compute-in-memory macros have been proposed to boost the energy efficiency of binary convolution tasks. However, this approach incurs inaccurate computation due to its sensitivity to temperature, noise, and process variations. In this work, we present a full-digital BNN architecture that leverages a novel Latch-XOR logic array for local bitwise multiplication, suppressing massive data movement and achieving 4.2× lower energy per operation compared to the decoupled standard cell approach. An optimized population count circuitry is also proposed for data accumulation, which obtains 1.37× Energy-Delay-Area saving compared to Binary-Adder-Tree-based implementation. To enable seamless hardware-software co-optimization, we have developed an in-house simulator for design space exploration as well as flexible mapping with various network topologies and kernel sizes. Our experiment shows the Latch-XOR-based architecture in 28nm CMOS technology achieves an enhanced energy efficiency of 2315 TOPS/W, 3.4× higher compared to the state-of-the-art synthesized digital architecture. This manifests that the proposed accelerator is highly suited for AIoT applications. Dongrui Li, Tomomasa Yamasaki, Aarthy Mani, Anh-Tuan Do, Niangjun Chen, Bo Wang 0020 |
ISLPED | 3 |
| 2020 | Ultra-Low Leakage, High Fan-Out Neuro Connection Map with TCAM-Based LUT, Localized Priority Encoder and Decoder-Less SRAMabstractIn this paper we present an energy and area efficient Neuro Connection Map (NCM) which features high fan-out connection to improve the core utilization for large scale neuromorphic system. To meet the stringent power and area requirements, we propose to deploy multi-Vth memory cell design for performance-leakage optimization, localized priority encoder for area/power minimization and finally a decoder-less SRAM for area saving. Our measurements at 1 V supply show that the whole design consumes only 819nW of leakage power. Read and search power are at 5.9μW/MHz and 14.8μW/MHz, respectively. When used in a neuro core with average 100 spikes per neuron per second, the NCM consumes only 0.2pJ/Synaptic event at 1V, making it highly suitable for energy-efficient neuromorphic computing applications. Aarthy Mani, Fei Li 0015, Ming Ming Wong, Luo Tao, Vishnu Paramasivam, Anh-Tuan Do |
ISCAS | 1 |
| 2020 | Scalable Block-Based Spiking Neural Network Hardware with a Multiplierless Neuron ModelabstractThis paper proposes a scalable hardware architecture for block-based spiking neural networks utilizing a multiplierless spiking neuron model. These blocks were implemented as a neurocore mesh generated from an interconnect algorithm, allowing for seamless scalability of the network size while mitigating connectivity errors. The routing fabric asynchronous protocol allows for critical timing paths between blocks to be relaxed. The proposed neuron model consumed less logic compared to standard models with multipliers, reducing up to 16% of the neurocore logic cell area. The network was implemented alongside a computing subsystem as an FPGA-based system-on-chip, communicating via a fabric interconnect bridge. Experimental results validate the functionality of the proposed system, and achieved comparable classification accuracy to existing works. Vishnu P. Nambiar, Eng-Kiat Koh, Junran Pu, Aarthy Mani, Ming Ming Wong, Wang Ling Goh, Anh-Tuan Do |
ISCAS | 4 |