Zhengzhe Wei

dblp:352/9082 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2025
0009-0009-1339-0398ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2025 A 2.793 μW Near-Threshold Neuronal Population Dynamics Trajectory Filter for Reliable Simultaneous Localization and Mapping
abstract
This work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for the trajectory error correction task within a simultaneous localization and mapping workflow. A custom discretized procedural algorithm approximating a neuronal population dynamics-based inference operation is developed for mapping onto an ultra-lightweight digital macro featuring massively parallel in-situ processing techniques. Fabricated using a 40nm technology, the test chip features a$22\times 22$neuron array with 0.1358mm2 core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming sub-10-$\mu $W power.
Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0096, Tony Tae-Hyoung Kim, Yuanjin Zheng
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 A Graph-Based Accelerator of Retinex Model With Bit-Serial Computing for Image Enhancements
abstract
This work proposes the Poisson equation formulation of the Retinex model for image enhancements using a low-power graph hardware accelerator performing finite difference updates on a lattice graph processing element (PE) array. By encapsulating the underlying algorithm in a graph hardware structure, a highly localized dataflow that takes advantage of the physical placement of the PEs is enabled to minimize data movement and maximize data reuse. The on-chip dataflow that achieves data sharing, and reuse among neighboring PEs during massively parallel updates is generated in each PE driven by two external control signals. Using a custom accumulator design intended for bit-serial computing, this work enables precision on demand and extensive on-chip data reuse with minimal area overhead, accommodating a non-overlap image mapping scheme in which a$20\times 20$image tile can be processed without external memory access at a time. With increasing user-configurable update count, image noise and shadow can be progressively removed with the inevitable loss of image details. Fabricated using a 65nm technology, the test chip occupies 0.2955mm2 core area and consumes 2.191mW operating at 1V, 25.6MHz, and a reconfigurable 10- or 14-bit precision.
Zhengzhe Wei, Junjie Mu, Yuanjin Zheng, Tony Tae-Hyoung Kim, Bongjin Kim
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 A 2.793µW Near-Threshold Neuronal Population Dynamics Simulator for Reliable Simultaneous Localization and Mapping
abstract
This work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for a component within the back-end of simultaneous localization and mapping. A custom discretized procedural algorithm including injection, finite difference update, activation, and inhibition to approximate neuronal population dynamics is developed for digital implementation. Fabricated using a 40nm technology, the test chip features a scalable neuron 22 × 22 array with 0.1358mm2core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming 2.793µW of power.
Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0016, Tony Tae-Hyoung Kim, Yuanjin Zheng
ISCAS1
2023 A Graph-Based Accelerator of Retinex Model with Bit-Serial Computing for Image Processing
abstract
This work implements the Poisson equation formulation of the Retinex model for image enhancements using a graph hardware accelerator performing finite difference updates on a 2D lattice graph PE array. A single clock gating control signal manages the data flow, data sharing, and reuse pattern among neighboring PEs during massively parallel updates. With increasing user-configurable update count, image noise and shadow can be progressively removed with the inevitable loss of image details. Accommodating a non-overlap image mapping scheme in which a$20\times 20$image tile can be processed without external memory access at a time, the proposed accelerator consists of 18$\times 18$regular PEs surrounded by$4\times 20$boundary PEs with reconfigurable data flow and 4 boundary cache registers. Fabricated using a 65nm technology, the test chip occupies 0.2955mm2core area, and consumes 2.191mW operating at 1V, 25.6MHz, and a reconfigurable 10- or 14-bit precision.
Zhengzhe Wei, Junjie Mu, Zhongzhiguang Lu, Yuanjin Zheng, Tony Tae-Hyoung Kim, Bongjin Kim
ISCAS1