EDBT 2026 Demo / reviewers in the wild / expert
Meng Zhang 0010
dblp:04/6901-10
· DBLP profile ↗
25ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-2188-8195ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Synergistic Bayesian Optimization and Reinforcement Learning with Bidirectional Interaction for Efficient VLSI Constraint Tuning
Jiayi Tu 0001, Jindong Tu, Meng Zhang 0010, Tinghuan Chen |
ASP-DAC | 5 |
| 2026 | A Survey of the First TinyML@ICCAD Contest for Ventricular Arrhythmia Detection by Artificial Intelligence on Low-power MicroprocessorabstractArtificial intelligence has achieved remarkable success in various real-world applications. However, the challenge lies in its implementation on hardware platforms with constrained resources and low power while maintaining real-time capabilities. Edge artificial intelligence, in particular, stands as a pivotal field for the practical deployment of AI. The 41st IEEE/ACM International Conference on Computer-Aided Design introduced the inaugural TinyML Design Contest in 2022. The contest entailed a rigorous, multi-month research and development competition, focusing on the creation of real-time detection algorithms for life-threatening ventricular arrhythmia. These algorithms were required to be deployable on the low-power microprocessor NUCLEO-L432KC. Open to multi-person teams worldwide, the contest garnered 150 teams participation teams from 50+ organizations, with 41 teams successfully completing the challenge. Our SEUer team secured the second place. This article provides a detailed exposition of the contest, offering insights into its structure and objectives. Furthermore, it analyzes and discusses the methods developed by some of the entries as well as representative results. Finally, the article concludes with directions for future improvements. Meng Zhang 0010, Tinghuan Chen, Jun Yang 0006 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | TKD: An Efficient Deep Learning Compiler with Cross-Device Knowledge DistillationabstractGenerating high-performance tensor programs on resource-constrained devices is challenging for current Deep Learning (DL) compilers that use learning-based cost models to predict the performance of tensor programs. Due to the inability of cost models to leverage cross-device information, it is extremely time-consuming to collect data and train a new cost model. To address this problem, this paper proposes TKD, a novel DL compiler that can be efficiently adapted to devices that are resource-constrained. TKD reduces the time budget by over 11x through an adaptive tensor program filter that eliminates redundant and unimportant measurements of tensor programs. Furthermore, by refining the cost model architecture with a multi-head attention module and distilling transferable knowledge from source devices, TKD outperforms state-of-the-art methods in prediction accuracy, compilation time, and compilation quality. We conducted experiments on the edge GPU, NVIDIA Jetson TX2, and the results show that compared to TenSet and TLP, TKD reduces compilation time by 1.58x and 1.16x, while achieving 1.40x and 1.27x speedups of the tensor programs, respectively. Chaoyao Shen, Linfeng Jiang, Meng Zhang 0010 |
DATE | 5 |
| 2025 | Real-Time Semantic Segmentation for UAV Perspectives on Embedded Platforms
Chaoyao Shen, Yuning Ji, Linfeng Jiang, Meng Zhang 0010 |
ICIC (1) | 6 |
| 2025 | Algorithm-Hardware Co-design for Accelerating Depthwise Separable CNNsabstractDepthwise separable convolution (DSC) is a popular method for constructing lightweight neural networks. However, the pointwise convolution (PWC) has a much larger number of parameters than the depthwise convolution (DWC), causing the imbalanced parameter ratio of PWC to DWC. In this article, we propose an efficient and hardware-efficiency convolution (Shared Kernel sliding on channel Convolution, SKC) to replace the redundant PWC in DSC for a balanced parameter ratio, where SKC customizes the sharing kernel in the channel dimension to reduce the number of parameters, and the local connection in the channel dimension reduces the computation. Furthermore, the proposed SKC is suitable for Winograd acceleration, and the large kernel decomposition method is introduced to facilitate its use. We implement the first Winograd-based FPGA hardware accelerator for DSCNets. The shared 1D and 2D Winograd convolution computing engine is proposed to compute the proposed DSC consisting of DWC and SKC efficiently. An alternating loading and reusing storage approach is developed to efficiently load SKC input feature maps. Experimental results show our DSC-based accelerator can achieve 20× higher power efficiency at the cost of a small loss of accuracy by algorithm-hardware co-design compared with traditional accelerators. RenGang Li, Tinghuan Chen, Meng Zhang 0010, Henk Corporaal |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2024 | An analysis of TinyML@ICCAD for implementing AI on low-power microprocessor
Meng Zhang 0010, Tinghuan Chen, Jun Yang 0006 |
Sci. China Inf. Sci. | 3 |
| 2024 | Fast Constraints Tuning via Transfer Learning and Multiobjective OptimizationabstractAs the complexity of very-large-scale integration (VLSI) increases, empirically determining the design constraints necessary to achieve the optimal performance, power, and area (PPA) within the electronic design automation (EDA) workflow becomes more challenging. Design space exploration is capable of effectively and automatically identifying the design constraints required to attain the optimal PPA in VLSI designs. However, the absence of prior knowledge can lead to less efficient explorations. This paper proposes a novel fast constraint tuning framework via transfer learning and multi-objective Bayesian optimization (MOBO) to find the optimal design constraints. Firstly, we introduce transfer learning into multi-objective Bayesian optimization by Gaussian Copula and transform the PPA data into residual observations. We propose to transfer the prior information of the implemented technologies to the advanced technology to optimize the parameter design space under the advanced technology. Secondly, we propose Gaussian process regression with an auto-encoder-based deep kernel as a surrogate model in MOBO. The auto-encoder-based deep kernel can extract more input features to make the surrogate model more precise. We employ the batch uncertainty-aware search acquisition function to improve exploration efficiency. Using this surrogate model and this acquisition function in MOBO can reduce the amount that EDA tools need to run. The average EDA tools running times of the proposed model is 204, and the average ADRS is 0.0373. Compared to state-of-the-art approaches, experiments on a CPU design reveal that a higher-quality Pareto frontier can be provided with a shorter running time. Meng Zhang 0010, Yifan Niu, Zewei Chen, Yajun Ha, Tinghuan Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Uint-Packing: Multiply Your DNN Accelerator Performance via Unsigned Integer DSP PackingabstractDSP blocks are undoubtedly efficient solutions for implementing multiply-accumulate (MAC) operations on FPGA. Since DSP resources are scarce in FPGA, the advanced solution is to pack parallel multiplication operations into a single DSP. However, available methods are based on signed-type multiplication, leading to both loss of accuracy and increased area. To solve these issues simultaneously, we propose an unsigned integer DSP packing generalization model called uint-packing. Guided by this generalization model, we design the novel computational structure of the DNN accelerator. Our system design is state-of-the-art, with 2.8× throughput and 4× energy efficiency compared to the third-place DAC-SDC’22 design. Meng Zhang 0010, Xinye Cao |
DAC | 2 |
| 2022 | A fast parameter tuning framework via transfer learning and multi-objective bayesian optimizationabstractDesign space exploration (DSE) can automatically and effectively determine design parameters to achieve the optimal performance, power and area (PPA) in very large-scale integration (VLSI) design. The lack of prior knowledge causes low efficient exploration. In this paper, a fast parameter tuning framework via transfer learning and multi-objective Bayesian optimization is proposed to quickly find the optimal design parameters. Gaussian Copula is utilized to establish the correlation of the implemented technology. The prior knowledge is integrated into multi-objective Bayesian optimization through transforming the PPA data to residual observation. The uncertainty-aware search acquisition function is employed to explore design space efficiently. Experiments on a CPU design show that this framework can achieve a higher quality of Pareto frontier with less design flow running than state-of-the-art methodologies. Tinghuan Chen, Jiaxin Huang 0010, Meng Zhang 0010 |
DAC | 4 |
| 2022 | An Efficient FPGA Implementation for Real-Time and Low-Power UAV Object DetectionabstractIn this paper, an efficient real-time hardware accelerator based on Field Programmable Gate Array (FPGA) is proposed for unmanned aerial vehicle (UAV) object detection. We first analyze the iSmart3-SkyNet (a popular UAV object detection network), using the roofline model. Then, a series of optimization strategies are proposed for low power and real-time UAV object detection based on FPGA. Stackable shared PE and regulable loop count improve the computing roof and the utilization of computing resources. Channel augmentation is used to increase the memory bandwidth, and improve the computing efficiency for shallow layers. Regulable Loop Count reduces unnecessary computation, and pre-load workflow improves the overall parallelism of heterogeneous systems. The results show that our accelerator (SEUT) achieves 78.6 frames per second and 0. 068J per image with 0.73 Intersection over Union for object detection. Source code will be available at https://github.com/aicerl/accob. Meng Zhang 0010, Henk Corporaal |
ISCAS | 3 |
| 2022 | Efficient channel expansion and pyramid depthwise-pointwise-depthwise neural networks
Meng Zhang 0010, Ruixia Wu, Dongpeng Weng |
Appl. Intell. | 2 |
| 2022 | Efficient depthwise separable convolution accelerator for classification and UAV object detection
Meng Zhang 0010, Ruixia Wu, Xinye Cao, Wenzhao Liu |
Neurocomputing | 3 |
| 2022 | SCWC: Structured channel weight sharing to compress convolutional neural networks
Meng Zhang 0010, Jiuyang Wang, Dongpeng Weng, Henk Corporaal |
Inf. Sci. | 2 |
| 2022 | OGCNet: Overlapped group convolution for deep convolutional neural networks
Meng Zhang 0010, Qianru Zhang |
Knowl. Based Syst. | 2 |
| 2022 | An Efficient Sharing Grouped Convolution via Bayesian LearningabstractCompared with traditional convolutions, grouped convolutional neural networks are promising for both model performance and network parameters. However, existing models with the grouped convolution still have parameter redundancy. In this article, concerning the grouped convolution, we propose a sharing grouped convolution structure to reduce parameters. To efficiently eliminate parameter redundancy and improve model performance, we propose a Bayesian sharing framework to transfer the vanilla grouped convolution to be the sharing structure. Intragroup correlation and intergroup importance are introduced into the prior of the parameters. We handle the Maximum Type II likelihood estimation problem of the intragroup correlation and intergroup importance by a group LASSO-type algorithm. The prior mean of the sharing kernels is iteratively updated. Extensive experiments are conducted to demonstrate that on different grouped convolutional neural networks, the proposed sharing grouped convolution structure with the Bayesian sharing framework can reduce parameters and improve prediction accuracy. The proposed sharing framework can reduce parameters up to 64.17%. For ResNeXt-50 with the sharing grouped convolution on ImageNet dataset, network parameters can be reduced by 96.875% in all grouped convolutional layers, and accuracies are improved to 78.86% and 94.54% for top-1 and top-5, respectively. Tinghuan Chen, Qi Sun 0002, Meng Zhang 0010, Hao Geng, Qianru Zhang, Bei Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Efficient densely connected convolutional neural networks
Meng Zhang 0010, Jiaojie Li, Feng Lv, Guodong Tong |
Pattern Recognit. | 2 |
| 2021 | Micro-Doppler Signature-Based Detection, Classification, and Localization of Small UAV With Long Short-Term Memory Neural NetworkabstractAlong with the popularization of small unmanned aerial vehicles (UAVs), societal concerns related to security, privacy, and public safety have gained more attention, thus opening a new avenue for small UAV surveillance. However, the conventional radar technologies pose challenges for the surveillance of small UAVs due to the high cost, small radar cross section, low flying altitude, and slow flying speed. In this article, we propose a novel micro-Doppler signature-based surveillance method using machine learning techniques, for detection, classification, and localization of small UAVs. Via extensive experiments, we demonstrate the performance gain of our proposed method by applying long short-term memory neural network. Yingxiang Sun, Samith Abeywickrama, Lahiru Jayasinghe, Chau Yuen, Jiajia Chen 0002, Meng Zhang 0010 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | Deep Learning Based Defect Detection for Solder Joints on Industrial X-Ray Circuit Board ImagesabstractQuality control is of vital importance during electronics production. As the methods of producing electronic circuits improve, there is an increasing chance of solder defects during assembling the printed circuit board (PCB). Many technologies have been incorporated for inspecting failed soldering, such as X-ray imaging, optical imaging, and thermal imaging. With some advanced algorithms, the new technologies are expected to control the production quality based on the digital images. However, current algorithms sometimes are not accurate enough to meet the quality control. Specialists are needed to do a follow-up checking. For automated X-ray inspection, joint of interest on the X-ray image is located by region of interest (ROI) and inspected by some algorithms. Some incorrect ROIs deteriorate the inspection algorithm. The high dimension of X-ray images and the varying sizes of image dimensions also challenge the inspection algorithms. On the other hand, recent advances on deep learning shed light on image-based tasks and are competitive to human levels. In this paper, deep learning is incorporated in X-ray imaging based quality control during PCB quality inspection. Two artificial intelligence (AI) based models are proposed and compared for joint defect detection. The noised ROI problem and the varying sizes of imaging dimension problem are addressed. The efficacy of the proposed methods are verified through experimenting on a real-world 3D X-ray dataset. By incorporating the proposed methods, specialist inspection workload is largely saved. Qianru Zhang, Meng Zhang 0010, Chinthaka Gamanayake, Chau Yuen, Zehao Geng, Hirunima Jayasekaraand, Xuewen Zhang, Chia-wei Woo, Jenny Chen Ni Low, Xiang Liu 0001 |
INDIN | 2 |
| 2020 | Robust deep auto-encoding Gaussian process regression for unsupervised anomaly detection
Jinan Fan, Qianru Zhang, Jialei Zhu, Meng Zhang 0010, Hanxiang Cao |
Neurocomputing | 4 |
| 2019 | Recent advances in convolutional neural network acceleration
Qianru Zhang, Meng Zhang 0010, Tinghuan Chen, Zhifei Sun, Yuzhe Ma, Bei Yu 0001 |
Neurocomputing | 2 |
| 2018 | Electricity Theft Detection Using Generative ModelsabstractAdvanced metering infrastructure (AMI) plays an important role in smart grid. On one hand, AMI makes the smart grid more vulnerable to cyber attacks. On the other hand, large amount of available usage data helps detect energy thefts using machine learning methods. In this paper, we focus on energy theft that results in customer usage pattern change in utility database. To overcome the imbalance problem between normal and anomaly behavior data, we propose an anomaly detection framework called semi-supervised generative Gaussian mixture model, which can be controlled with detection indicator thresholds to adjust the intensity of detection. Human knowledge is successfully introduced into the model using detection indicators. We analyze it with various machine learning based methods including one-class SVM and autoencoder, and show that our framework has the most effective performance validated by simulation that is based on real-world energy consumption data. Qianru Zhang, Meng Zhang 0010, Tinghuan Chen, Jinan Fan |
ICTAI | 2 |
| 2017 | Delay-cost tradeoff for virtual machine migration in cloud data centers
Xiumin Wang 0005, Xiaoming Chen 0001, Chau Yuen, Weiwei Wu 0001, Meng Zhang 0010, Cheng Zhan |
J. Netw. Comput. Appl. | 5 |
| 2017 | Context Management Scheme Optimization of Coarse-Grained Reconfigurable Architecture for Multimedia ApplicationsabstractDue to the combination of flexibility and efficiency, coarse-grained reconfigurable architectures (CGRAs) are suitable for the implementation of computing-intensive applications. However, with the growing performance requirements, the scale of CGRA increases exponentially, which leads to configuration performance degradation and configuration power rise. Based on the analysis of configuration context features, we optimize the context management scheme of CGRA from the aspects of context cache structure and replacement strategy. The context cache is structured hierarchically to reduce the memory overhead without configuration performance degradation and a hybrid context replacement algorithm is proposed to further increase the configuration efficiency with a novel context frequency weight factor. Experimental results show that the proposed context management scheme improves the configuration performance of the base CGRA significantly by 13.6%-20.5% for H.264 decoding and 13.6%-20.5% for MPEG2 decoding with only 43% context cache cost. Compared with other works, the proposed context management scheme shows the advantages of 2.3-6× less normalized context cache size and 2.3-2.7× cache efficiency. Peng Cao 0002, Bo Liu 0019, Jinjiang Yang, Jun Yang 0006, Meng Zhang 0010, Longxing Shi |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2012 | Improved timing carrier synchronizer and error controller for wireless multimedia sensor networksabstractThe improved synchronization algorithm and error controller structure for OFDM system in wireless multimedia sensor networks are presented. The timing and carrier synchronization are acquired by coarse estimate and fine synchronization. The simulation results show that the improved synchronization algorithm performance of timing and carrier synchronization is better than traditional algorithm, merely poorer than Park algorithm. The OFDM synchronizer and error controller are verified on the platform of Altera FPGA. The results validate that each performance of both improved synchronization and error controlling method meets WMSN system requirements better. Timing error is only one sampling point and the frequency offset is 0.25% of sub-carrier frequency interval. Meng Zhang 0010, Yingdong Zhou, Youchao Dong, Jianhui Wu 0001 |
WOWMOM | 1 |
| 2012 | All-Digital Wide Range Precharge Logic 50% Duty Cycle CorrectorabstractA novel all-digital 50% duty cycle corrector (DCC) is pro- posed in this paper. The DCC features include a delay unit based on precharge logic gates with low delay time and a robust SR latch under process voltage and temperature variations for final edge combination over wide frequency and duty-cycle ranges. The rising edge of the output clock has a constant delay when comparing to the input clock, which makes it easy to cooperate with a delay locked loop. The circuit is fabricated in Chartered 0.18-μm CMOS process. The acceptable input clock frequency ranges from 400 MHz to 2 GHz. The correcting error is ±3.5% at 1 GHz or ±1% at 400 MHz. Junhui Gu, Jianhui Wu 0001, Danhong Gu, Meng Zhang 0010, Longxing Shi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |