EDBT 2026 Demo / reviewers in the wild / expert
Haikang Diao
dblp:289/3420
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-3347-5590ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SEGA-DCIM: Design Space Exploration-Guided Automatic Digital CIM Compiler with Multiple Precision SupportabstractDigital computing-in-memory (DCIM) has been a popular solution for addressing the memory wall problem in recent years. However, the DCIM design still heavily relies on manual efforts, and the optimization of DCIM is often based on human experience. These disadvantages limit the time to market while increasing the design difficulty of DCIMs. This work proposes a design space exploration-guided automatic DCIM compiler (SEGA-DCIM) with multiple precision support, including integer and floating-point data precision operations. SEGA-DCIM can automatically generate netlists and layouts of DCIM designs by leveraging a template-based method. With a multi-objective genetic algorithm (MOGA)-based design space explorer, SEGA-DCIM can easily select appropriate DCIM designs for a specific application considering the trade-offs among area, power, and delay. As demonstrated by the experimental results, SEGA-DCIM offers solutions with wide design space, including integer and floating-point precision designs, while maintaining competitive performance compared to state-of-the-art (SOTA) DCIMs. Haikang Diao, Haoyi Zhang, Haoyang Luo, Yibo Lin, Runsheng Wang, Yuan Wang 0001, Xiyuan Tang |
DATE | 1 |
| 2025 | Adder-DCIM: A Parallel Bit-Flexible Digital CIM Accelerator Joint Model Compression Framework for AdderNet InferenceabstractHeavily constrained edge-side tasks necessitate AI chips with low power consumption, low latency, and low cost. In recent years, digital compute-in-memory (CIM) has emerged as a promising solution to enhance energy efficiency and throughput density. However, digital CIM still faces various challenges: significant power and area overhead from multiplication, difficulty in exploiting fine-grained sparsity, and throughput degradation associated with bit-serial architecture. In this work, we propose Adder-DCIM: an efficient parallel bit-flexible DCIM accelerator joint model compression framework for AdderNet inference, in which the key contributions are: 1) a CIM-friendly model compression framework that includes operator decomposition, lossless fine-grained sparsity, and Kullback-Leibler-divergence(KLD)-based Cin-wise mixed-precision quantization; 2) a synchronous parallel DCIM architecture for throughput improvement with mix-precision quantization; 3) a bit-flexible minimal selector circuit for efficient mixed-precision computation. The experimental results demonstrate that under a 28-nm process, the proposed Adder-DCIM achieves a peak energy efficiency of 134 TOPS/W and a peak throughput density of 6.49 TOPS/mm2at INT8. When running ResNet20 on CIFAR10 and ResNet50 on ImageNet, the proposed Adder-DCIM achieves 255 TOPS/[email protected] and 243 TOPS/[email protected] with only a slight decrease in accuracy by 0.34% and 0.7%, respectively. Compared to multiply-based DCIM, Adder-DCIM improves energy efficiency × throughput density metrics by 20.7× for ResNet50 inference. Haikang Diao, Chuyue Tang, Bocheng Xu, Haoyang Luo, Meng Li 0004, Yuan Wang 0001, Xiyuan Tang |
ICCAD | 1 |
| 2022 | Redistribution of Weights and Activations for AdderNet QuantizationabstractAdder Neural Network (AdderNet) provides a new way for developing energy-efficient neural networks by replacing the expensive multiplications in convolution with cheaper additions (i.e., L1-norm). To achieve higher hardware efficiency, it is necessary to further study the low-bit quantization of AdderNet. Due to the limitation that the commutative law in multiplication does not hold in L1-norm, the well-established quantization methods on convolutional networks cannot be applied on AdderNets. Thus, the existing AdderNet quantization techniques propose to use only one shared scale to quantize both the weights and activations simultaneously. Admittedly, such an approach can keep the commutative law in the L1-norm quantization process, while the accuracy drop after low-bit quantization cannot be ignored. To this end, we first thoroughly analyze the difference on distributions of weights and activations in AdderNet and then propose a new quantization algorithm by redistributing the weights and the activations. Specifically, the pre-trained full-precision weights in different kernels are clustered into different groups, then the intra-group sharing and inter-group independent scales can be adopted. To further compensate the accuracy drop caused by the distribution difference, we then develop a lossless range clamp scheme for weights and a simple yet effective outliers clamp strategy for activations. Thus, the functionality of full-precision weights and the representation ability of full-precision activations can be fully preserved. The effectiveness of the proposed quantization method for AdderNet is well verified on several benchmarks, e.g., our 4-bit post-training quantized adder ResNet-18 achieves an 66.5% top-1 accuracy on the ImageNet with comparable energy efficiency, which is about 8.5% higher than that of the previous AdderNet quantization methods. Code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/AdderQuant. Kai Han 0002, Haikang Diao, Chuanjian Liu, Enhua Wu, Yunhe Wang 0001 |
NeurIPS | 3 |
| 2022 | Real-Time and Cost-Effective Smart Mat System Based on Frequency Channel Selection for Sleep Posture Recognition in IoMTabstractSleep posture, which affects the quality of sleep and could lead to medical conditions, such as pressure ulcers, is a key metric for sleep analysis in Internet of Medical Things (IoMT). In this article, a real-time and low-cost smart mat system for sleep posture recognition based on frequency channel selection is proposed. The system can recognize postures unobtrusively with a dense flexible sensor array. In addition, to enable real-time recognition with a relatively low-cost STM32 processor system, a lightweight algorithm that includes frequency channel selection, model pretraining, and real-time classification is proposed. Through a series of short-term and overnight experiments with 21 subjects, the feasibility and reliability of the proposed system were evaluated. Experimental results show that the accuracy of the short-term experiment is up to 95.43% and of the overnight experiment is up to 86.80% for four posture categories (supine, prone, right, and left) classification. The model size is just 56 kB which is much smaller than other methods. The runtime of the complete algorithm is about 6 ms with a low-power STM32 embedded system, which shows the system’s ability to provide real-time posture recognition. As an edge device, the proposed system could lead to the development of fast, convenient, and low-cost sleep posture recognition products for IoMT. Haikang Diao, Chen Chen 0039, Wei Yuan 0005, Amara Amara, Toshiyo Tamura, Benny P. L. Lo, Long Meng, Sio-Hang Pun, Yuan-Ting Zhang, Wei Chen 0015 |
IEEE Internet Things J. | 1 |
| 2021 | Unobtrusive Smart Mat System for Sleep Posture RecognitionabstractSleep posture, as a crucial index for sleep quality assessment and pressure ulcer prevention, has been widely studied for medical diagnoses and sleep disease treatment. In this paper, an unobtrusive smart mat system for sleep posture recognition is proposed, which is based on a dense flexible sensor array and printed electrodes and along with an algorithmic framework. With the dense flexible sensor array, the system offers a comfortable and high-resolution solution for long-term pressure sensing. Meanwhile, compared with other large-area and low-density mat systems, it reduces the area to minimize manufacturing cost and computational complexity, while also increases the density of the sensor to improve accuracy. To distinguish the sleep postures, the algorithmic framework that includes pre-processing, feature extraction, and posture classification is developed. Pilot studies in two scenarios including subject-dependent and subject- independent classification are performed with 7 persons for 4 different postures recognition. The experimental results show that the accuracy of the smart mat system can achieve over 78% using Support Vector Machines (SVMs) and k-Nearest Neighbor (kNN) for the subject-independent scenario. For the subject-dependent scenario, the accuracy can reach over 95%. It proves that the proposed method can recognize different sleep postures effectively. Haikang Diao, Chen Chen 0039, Wei Chen 0015, Wei Yuan 0005, Amara Amara |
ISCAS | 1 |