EDBT 2026 Demo / reviewers in the wild / expert
Zhaoteng Meng
dblp:312/7886
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0000-4476-8669ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimizing Area and Power of MAC Arrays in DNN Accelerators via Overflow-Aware Partial Sum ManagementabstractOn-device fixed-point deep neural network (DNN) accelerators are widely used, but the multiply-accumulate (MAC) units that perform the atomic operations in DNNs have become inefficient due to the overestimation of guard bits. To address this issue, we propose an overflow-aware management mechanism. This mechanism promptly adds the partial sum that is about to overflow to the previous partial sum stored in the output buffer and then writes it back into the buffer, thereby reducing the local guard bit overhead in the processing elements (PEs). We implemented and evaluated a PE array based on this mechanism, which reduced area overhead for adders by 43% and for registers by 35%, along with a 14% decrease in energy consumption. Zhaoteng Meng, Kailin Lv, Long Xiao |
ISCAS | 1 |
| 2025 | MASL-AFU: A High Memory Access Efficiency 2-D Scalable LUT-Based Activation Function Unit for On-Device DNN TrainingabstractOn-device deep neural network (DNN) training faces constraints in storage capacity and energy supply. Existing works primarily focus on optimizing the training of convolutional and batch normalization (BN) layers to improve the compute-to-communication (CTC) ratio and reduce the energy cost of off-chip memory access (MA). However, the training of activation layers remains challenging due to the additional off-chip MA required for derivative calculations. This article proposes MASL-AFU, an architecture designed to accelerate the activation layer in on-device DNN training. MASL-AFU leverages nonuniform piecewise linear (NUPWL) functions to speed up the forward propagation (FP) in the activation layer. During the error propagation (EP) process, retrieving derivatives from a lookup table (LUT) eliminates the need for redundant retrieval of the input data used in FP. By storing LUT indices instead of the original activation inputs, MASL-AFU significantly reduces and accelerates MA. Compared to other activation function units, MASL-AFU offers up to a$5.8\times $increase in computational and off-chip MA efficiency. In addition, MASL-AFU incorporates two dimensions of scalability: data precision and the number of LUT entries. These scalable, hardware-friendly methods enhance MASL-AFU’s area efficiency by up to$3.24\times $and energy efficiency by up to$3.85\times $. Zhaoteng Meng, Jianing Zeng, Kailin Lv, Haoyue Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | BitHist: A Precision-Scalable Sparse-Awareness DNN Accelerator Based on Bit Slices Products Histogram
Zhaoteng Meng, Long Xiao, Xiaoyao Gao |
Euro-Par | 1 |
| 2022 | A Flexible Real-Time Stereo Vision Architecture for Multiple Data Streams with Runtime Configurable ParametersabstractIt is significant for a stereo vision real-time computing system to flexibly adapt to different parameters of stereo matching without re-customizing hardwares. In this paper, a configurable pipelined hardware architecture based on the sum of absolute differences (SAD) algorithm is proposed. We split the SAD calculation into two parts to accommodate pipelined computing. The architecture can be configured with different resolutions, window sizes, and disparity levels without stopping and restarting. In addition, it can be configured as a multiple-data-stream mode and we have developed a configuration gen-eration algorithm for the mode. The presented architecture is synthesized and implemented on a Xilinx ZCUI04 board. The evaluation results demonstrate that the real-time computing of 480P, 720P, and 1080P video streams can be process at 250MHz with the peak computing performance of 480P/784fps at the disparity level of 125. It uses 60% LUTs, 34% registers, and 39 % BRAM, producing flexible configurability and superior computing performance than the other similar work. Zhaoteng Meng |
FPL | 1 |
| 2021 | A Change-Aware Approach for Relative Motion SegmentationabstractAnalysis on changes of image features is an effective means to leverage multi-dimensional information in motion segmentation. Methods without considering temporal perspective are agnostic of motion state, and inherently unsuitable for motion detection. However, the effect of existing spatio-temporal approaches is lagging behind that of spatial methods by a margin. To make better use of temporal information, this paper tackles the task of moving object segmentation by constructing a Change-aware Siamese neural network(ChaSiam) to detect relative foreground and changes. Further, a reference frame update strategy is attached to our network for overcoming the weakness of spatio-temporal approaches in cases with camera ego-motion. Extensive experiments show that our proposed model outperforms previous state-of-the-art spatio-temporal methods on Change Detection dataset, and compared with spatial methods our model has similar performance with better generalization. Zhuojun Zou, Zhaoteng Meng |
ICME | 2 |