EDBT 2026 Demo / reviewers in the wild / expert
Boyi Dong
dblp:365/7913
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-5729-9654ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Photoacoustic Quality-Enhancement Neural Network Processor with the Scalable and End-to-End Architecture by Improving the Sparsity LevelabstractRecent advancements have marked significant progress in photoacoustic imaging as an effective method for acquiring deep bio-tissue visuals in modern medical clinical therapy and the efficacy of U-Net and its variants has been established for imaging quality enhancement in this field. Unlike common computer vision datasets such as ImageNet [1] and PASCAL VOC [2], biomedical images exhibit highly structured patterns, low spatial resolution, and single-channel modality, as shown in Fig. 1. Additionally, the U-Net parameters trained for medical super-resolution tasks demonstrate a high sparsity ratio, making them suitable for implementation on edge-computing platforms. Therefore, developing an energy-efficient photoacoustic imaging setup in this area is a natural progression. However, this development is constrained by the current neural network architectures, which are built around a U-Net backbone. The multi-stage feature extractor, skip connection integration across different blocks, and the encoder-decoder backbone design pose significant challenges to cutting-edge computational hardware platforms. In this study, a scalable, sparsity-supported neural network accelerator architecture for bio-tissue imaging quality enhancement is proposed to meet the stringent requirements of latency and energy efficiency, as depicted in Fig. 2. This architecture achieves desired performance improvements by exploring the sparsity possibilities in neural network during the training process and implementing an end-to-end pixel-first hardware design to minimize data movement and support sparsity computation. Compared with the state-of-the-art related works, this optimized architecture has achieved minimum on-chip storage overhead and the fastest frame for the application of photoacoustic imaging quality enhancement. The scalable architecture has also been implemented on a Xilinx XCZU9EG FPGA and attains a performance of PSNR@ 24 dB and a frame rate of 164 fps at a working frequency of 250 MHz. Zhengyuan Zhang 0002, Caijie Liang, Boyi Dong, Yange Wang, Zhongzhiguang Lu, Xiangjun Yin, Shenglong Zhuo, Yifan Wu 0009, Yingjie Cao, Tianyang Zhou, Jian Qian, Patrick Chiang 0001, Lei Qiu 0002, Yuanjin Zheng |
ISCAS | 5 |
| 2025 | Compact Sub-THz Frequency Conversion Module in 28-nm CMOS for D-Band Radar TransceiverabstractThis paper proposes a compact sub-THz frequency conversion module for D-band transceivers, fabricated using a 28-nm CMOS process. The module integrates an injection-locked frequency multiplier (ILFM) for Tx signal frequency up-conversion and an active Gilbert double-balanced mixer for Rx signal frequency down-conversion. The system was tested on a probe station. Utilizing a tunable coupling-coil technique and optimized inductance, the ILFM achieves a locking range of 105.2-125.5 GHz with an output power of -4.5 dBm. The Gilbert mixer demonstrates a conversion loss of -6.5 dB across the same range with an LO power of -4.2 dBm. The active region of ILFM and mixer chips occupy areas of 0.31 mm2and 0.51 mm2, respectively. Yange Wang, Guanghao Fan, Boyi Dong, Zhongzhiguang Lu, Cao Wan, Yuanjin Zheng |
ISCAS | 3 |
| 2025 | A 2.793 μW Near-Threshold Neuronal Population Dynamics Trajectory Filter for Reliable Simultaneous Localization and MappingabstractThis work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for the trajectory error correction task within a simultaneous localization and mapping workflow. A custom discretized procedural algorithm approximating a neuronal population dynamics-based inference operation is developed for mapping onto an ultra-lightweight digital macro featuring massively parallel in-situ processing techniques. Fabricated using a 40nm technology, the test chip features a$22\times 22$neuron array with 0.1358mm2 core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming sub-10-$\mu $W power. Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0096, Tony Tae-Hyoung Kim, Yuanjin Zheng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | A 2.793µW Near-Threshold Neuronal Population Dynamics Simulator for Reliable Simultaneous Localization and MappingabstractThis work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for a component within the back-end of simultaneous localization and mapping. A custom discretized procedural algorithm including injection, finite difference update, activation, and inhibition to approximate neuronal population dynamics is developed for digital implementation. Fabricated using a 40nm technology, the test chip features a scalable neuron 22 × 22 array with 0.1358mm2core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming 2.793µW of power. Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0016, Tony Tae-Hyoung Kim, Yuanjin Zheng |
ISCAS | 2 |
| 2024 | A Fully Probabilistic Model for Sigmoid Approximation and Its Hardware- Efficient ImplementationabstractThe sigmoid function is a representative activation function in shallow neural networks. Its hardware realization is challenging due to the complex exponential and reciprocal operations. Existing studies applied piecewise models to approximate sigmoid function and employed numerical methods or non-uniform input segmentations to mitigate fitting inaccuracies. However, the breakpoints introduce inevitable approximation precision loss. Besides, additional fitting processes greatly increase hardware complexity and power consumption. This paper presents a hardware-friendly sigmoidal approximation from the perspective of probability theory. We find that for a given input, the output of a sigmoid function can be approximated by the probability that the sum of this input and a Gaussian random variable is greater than or equal to zero. As the derived theorem does not involve piecewise expressions, the precision loss caused by the breakpoint issue is avoided. A low-complexity binary-search-based address localization method is proposed to optimize our theorem for hardware implementation. For the optimized scheme, an efficient implemented circuit is also presented. Our scheme’s approximation ability and hardware efficiency are validated through software modeling and FPGA- and ASIC-based experiments. Feedforward neural network-based classification applications demonstrate that building networks with the proposed sigmoid approximator has only a tiny recognition rate loss. Wenhao Lu, Minshan Lu, Xiangfen Zhang, Zhongzhiguang Lu, Boyi Dong |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |