EDBT 2026 Demo / reviewers in the wild / expert
Ming Ming Wong
dblp:01/8059
· DBLP profile ↗
15ranked-venue papers
5as first author
7since 2021 · last 2024
0000-0002-6420-1202ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | 1.63 pJ/SOP Neuromorphic Processor With Integrated Partial Sum Routers for In-Network ComputingabstractNeuromorphic computing is promising to achieve unprecedented energy efficiency by emulating the human brain’s mechanism. Conventional neuromorphic accelerators employ split-and-merge method to map spiking neural networks’ inputs to surpass the fan-in capabilities of a single neuron core. However, this approach gives rise to the risk of accuracy compromise and extra core usage for the merging process. Moreover, it requires excessive data movement and clock cycles to aggregate spikes generated by partial sums instead of total sums obtained from different cores with substantial power and energy overhead. This work presents a novel approach to addressing the challenges imposed by the split-and-merge method. We propose an energy-efficient, reconfigurable neuromorphic processor that leverages several key techniques to mitigate the above issues. First, we introduce a partial sum router circuitry that enables in-network computing (INC), eliminating the need for extra merge cores. Second, we adopt software-defined Networks-on-Chip (NoCs) by leveraging predefined, efficient routing, eliminating power-hungry routing computation. At last, we incorporate fine-grained power gating and clock gating techniques for further power reduction. Experimental results from our test chip demonstrate the lossless mapping of the algorithm and exceptional energy efficiency, achieving an energy consumption of 1.63 pJ/SOP at 0.48 V. This energy efficiency represents a 22.4% improvement compared to the state-of-the-art results. Our proposed neuromorphic processor provides an efficient and flexible solution for neural network processing, mitigating the limitations of the traditional split-and-merge approach while delivering superior energy efficiency. Dongrui Li, Ming Ming Wong, Yi Sheng Chong, Jun Zhou 0014, Mohit Upadhyay, Ananta Narayanan Balaji, Aarthy Mani, Weng-Fai Wong, Li-Shiuan Peh, Anh-Tuan Do, Bo Wang 0020 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | 1.7pJ/SOP Neuromorphic Processor with Integrated Partial Sum Routers for In-Network ComputingabstractConventional neuromorphic accelerators primarily leverage split-merge method to accommodate a neural network that is beyond a single core's size, leading to possible accuracy loss, extra core usage and significant power and energy overhead. This work presents an energy-efficient, reconfigurable neuro-morphic processor to address the problem by (i) a partial sum router circuitry that enables in-network computing to remove the need of extra merge cores; (ii) software-defined Networks-on-Chip that eliminates the power-hungry routing compute and (iii) fine-grained power gating and clock gating technique for power reduction. Our test chip achieves lossless mapping as the algorithm and an energy efficiency of 1.7pJ/SOP at 0.5V, 19% lower than state-of-the-art result. Bo Wang 0020, Ming Ming Wong, Dongrui Li, Yi Sheng Chong, Jun Zhou 0014, Weng-Fai Wong, Li-Shiuan Peh, Aarthy Mani, Mohit Upadhyay, Ananta Narayanan Balaji, Anh-Tuan Do |
ISCAS | 2 |
| 2022 | A 1800μm2, 953Gbps/W AES Accelerator for IoT Applications in 40nm CMOSabstractA compact and energy-efficient AES accelerator for area and power-constrained IoT applications was fabricated in a 40nm CMOS process. By eliminating the need of intermediate data registers for MixColumns and ShiftRows in our proposed AES accelerator, we were able to reduce the total flip-flops to only 269 bits. Further, by reusing functional blocks and swapping the D flip-flops in data storage with scan flip-flops, our chip occupies only a tiny area of $1800 \mu \text{m}^{2}$ with an extremely low number of 657 gates. In addition, clock gating method and near-threshold voltage were used in our design. Thus, our accelerator consumes only $3.2 \mu \text{W}$ with an operation efficiency of 953 Gbps/W using a 0.48 V supply voltage. Compared with prior arts, our design has savings of 53% on area and 55% on the number of gates. When operated with a supply voltage of 0.48 V at 25°C, we can also achieve lower energy efficiency. Jingjing Lan, Vishnu P. Nambiar, Ming Ming Wong, Fei Li 0015, Yuan Gao 0011, Kevin Tshun Chuan Chai, Anh-Tuan Do |
ISCAS | 3 |
| 2022 | Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong |
Neurocomputing | 8 |
| 2022 | Corrigendum to "Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency" [Neurocomputing (2022) 128-140]
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong |
Neurocomputing | 8 |
| 2021 | A 25 TOPS/W High Power Efficiency Deterministic and Split Stochastic MAC (SC-MAC) DesignabstractThis work presented a stochastic computing (SC) split multiply-and-accumulate (MAC) unit that is operated using deterministic sequence and is able to achieve time latency and power reductions without accuracy degrading. In this improved deterministic SC design, the conventional Stochastic Number Generator (SNG) with large overhead is replaced with a lightweight decoder that effectively generates uncorrelated and segmented stochastic number (SN) without the need for random sources (PRNG). The proposed deterministic and split SC-MAC is implemented in ASIC 40nm technology for detailed hardware evaluation and its functionality is also verified in convolutional neural network (CNN) using MNIST data sets. The new SCMAC is found to be higher in power efficiency (GMACS/mW) and lower in energy consumption (pJ/MAC) as compared to the conventional SC-MAC as well as the prior arts. Ming Ming Wong, Anh-Tuan Do |
VLSI-SoC | 1 |
| 2021 | A 5.28-mm² 4.5-pJ/SOP Energy-Efficient Spiking Neural Network Hardware With Reconfigurable High Processing Speed Neuron Core and Congestion-Aware RouterabstractIn recent years, fast computation, low power, and small footprint are the key motivations for building SNN hardware. The unique features of SNN hardware have not been fully exploited, where the computation speed and energy efficiency of the SNN hardware can be improved according to the sparse spiking and non-uniform traffic of SNN. In this paper, we propose a 5.28-mm$^{2}~4096$-neuron 1M-synapse energy-efficient digital SNN hardware that can achieve ultra-low energy per synaptic operation of 4.5 pJ. The proposed neuron computing unit is implemented in pipeline architecture to achieve high synaptic processing speed. The proposed spike processing unit can significantly increase the processing speed of the neuron core by$1.9\times $and$9.4\times $when the spike injection rate is 50% and 10%, respectively. Besides, the increase in the processing speed of the neuron core leads to a reduction in energy consumption of up to 81.5%. An event-driven clock gating circuit that can reduce the power consumption of the proposed neuron block by more than 70% is proposed in this paper. This paper proposes a supervised STDP+ algorithm for SNN training, and the classification accuracy of the MNIST digits is 89.6% with 73.6% weight sparsity of the output layer. Junran Pu, Wang Ling Goh, Vishnu P. Nambiar, Ming Ming Wong, Anh-Tuan Do |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2020 | Ultra-Low Leakage, High Fan-Out Neuro Connection Map with TCAM-Based LUT, Localized Priority Encoder and Decoder-Less SRAMabstractIn this paper we present an energy and area efficient Neuro Connection Map (NCM) which features high fan-out connection to improve the core utilization for large scale neuromorphic system. To meet the stringent power and area requirements, we propose to deploy multi-Vth memory cell design for performance-leakage optimization, localized priority encoder for area/power minimization and finally a decoder-less SRAM for area saving. Our measurements at 1 V supply show that the whole design consumes only 819nW of leakage power. Read and search power are at 5.9μW/MHz and 14.8μW/MHz, respectively. When used in a neuro core with average 100 spikes per neuron per second, the NCM consumes only 0.2pJ/Synaptic event at 1V, making it highly suitable for energy-efficient neuromorphic computing applications. Aarthy Mani, Fei Li 0015, Ming Ming Wong, Luo Tao, Vishnu Paramasivam, Anh-Tuan Do |
ISCAS | 3 |
| 2020 | Scalable Block-Based Spiking Neural Network Hardware with a Multiplierless Neuron ModelabstractThis paper proposes a scalable hardware architecture for block-based spiking neural networks utilizing a multiplierless spiking neuron model. These blocks were implemented as a neurocore mesh generated from an interconnect algorithm, allowing for seamless scalability of the network size while mitigating connectivity errors. The routing fabric asynchronous protocol allows for critical timing paths between blocks to be relaxed. The proposed neuron model consumed less logic compared to standard models with multipliers, reducing up to 16% of the neurocore logic cell area. The network was implemented alongside a computing subsystem as an FPGA-based system-on-chip, communicating via a fabric interconnect bridge. Experimental results validate the functionality of the proposed system, and achieved comparable classification accuracy to existing works. Vishnu P. Nambiar, Eng-Kiat Koh, Junran Pu, Aarthy Mani, Ming Ming Wong, Wang Ling Goh, Anh-Tuan Do |
ISCAS | 5 |
| 2018 | A New High Throughput and Area Efficient SHA-3 ImplementationabstractHigh performance and area efficient Secure Hash Algorithm (SHA-3) hardware realization is investigated and proposed in this work. In addition to the new and simplified round constant (RC) generator, the presented SHA-3 hash implementations employed architectural optimization approaches based on the concepts of unrolling, pipelining and subpipelining. This has therefore produced a total of five implementations of SHA-3 which are denoted as Cases I-V in both FPGA and ASIC. Considering the trade-offs between the performance and hardware cost, the best architecture in term of the throughput and area efficiency is identified in Case V. The architecture has the highest throughput of 16.51 Gbps and area efficiency of 11.47 Mbps/slices for the FPGA implementation. While in ASIC, our best implementation (Case V) achieves the highest throughput of 48 Gbps. Ming Ming Wong, Jawad Haj-Yahya, Suman Sau, Anupam Chattopadhyay |
ISCAS | 1 |
| 2018 | Construction of a Low Multiplicative Complexity GF (24) Inversion Circuit for Compact AES S-BoxabstractIn this work, we construct a compact composite AES S-Box by deriving a new low multiplicative complexity GF (24) inversion circuit. A deterministic tree search algorithm is applied to search for constructions that are optimum in terms of multiplicative complexity. From the results, the circuit with the smallest gate count is selected for GF (24) inversion. To the best of our knowledge, the proposed AES S-Box requires the smallest gate count to date with the size of 112 gates and depth of 25 gates. Jia Jun Tay, Mou Ling Dennis Wong, Ming Ming Wong, Cishen Zhang, Ismat Hijazin |
TENCON | 3 |
| 2018 | Lightweight and High Performance SHA-256 using Architectural Folding and 4-2 Adder CompressorabstractThe modern era of Internet-of-Things (IoT) is naturally imposing a tight area/runtime constraint on the computing kernels. Security kernels, as part of the standardized protocols as well as custom defense techniques, are among the most common tasks executed on every digital device. Therefore, low area cost and high performance implementation of security kernels is an important goal of current system designers. In this paper, we revisit the state-of-the-art implementations of SHA-256, a standardized security primitive for authentication and propose novel optimizations. Our optimizations, based on architectural folding and 4-2 adder compressor, are geared toward both lightweight and high performance implementations. Detailed experiments of our optimized architecture on different FPGA fabrics clearly demonstrate their benefits. Our presented design point successfully attained the highest hardware efficiency (throughput/area) figures among the published literature so far. Ming Ming Wong, Vikramkumar Pudi, Anupam Chattopadhyay |
VLSI-SoC | 1 |
| 2018 | A tree search algorithm for low multiplicative complexity logic design
Jia Jun Tay, Mou Ling Dennis Wong, Ming Ming Wong, Cishen Zhang, Ismat Hijazin |
Future Gener. Comput. Syst. | 3 |
| 2012 | Compact Multiplicative Inverter for Hardware Elliptic Curve Cryptosystem
Ming Ming Wong, Mou Ling Dennis Wong, Ka Lok Man |
NPC | 1 |
| 2012 | Construction of Optimum Composite Field Architecture for Compact High-Throughput AES S-BoxesabstractIn this work, we derive three novel composite field arithmetic (CFA) Advanced Encryption Standard (AES) S-boxes of the field GF(((22)2)2). The best construction is selected after a sequence of algorithmic and architectural optimization processes. Furthermore, for each composite field constructions, there exists eight possible isomorphic mappings. Therefore, after the exploitation of a new common subexpression elimination algorithm, the isomorphic mapping that results in the minimal implementation area cost is chosen. High throughput hardware implementations of our proposed CFA AES S-boxes are reported towards the end of this paper. Through the exploitation of both algebraic normal form and seven stages fine-grained pipelining, our best case achieves a throughput 3.49 Gbps on a Cyclone II EP2C5T144C6 field-programmable gate array. Ming Ming Wong, Mou Ling Dennis Wong, Asoke K. Nandi, Ismat Hijazin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |