EDBT 2026 Demo / reviewers in the wild / expert
Yonghao Tan
dblp:333/5165
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0001-5372-5863ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Adaptive Depth Processing System Based on Foundation Transformers and Time-of-Flight FusionabstractGenerating high-quality depth maps with accurate values is a critical research topic, and various depth estimation methods, such as Time-of-Flight (ToF) and Monocular Depth Estimation (MDE), are advancing rapidly. However, each single-sensor approach is inherently limited by the characteristics of its respective modality. Meanwhile, fused approaches using multiple sensors or dedicated trained models are often plagued by system complexity and limited generalization. In this paper, we propose an adaptive depth processing system based on the foundation transformer and ToF fusion, aiming to harness the precision of ToF data and the high-quality depth priors provided by the foundation model in a monocular and zero-shot manner. A hardware–algorithm co-design partitions the computation between a dedicated ASIC for compute-intensive foundation transformer acceleration and an MPSoC FPGA for potential reconfigurable fusion, yielding 36.4 ms latency and 1.9 W power under a 27.6 GOPS/frame workload. Zero-shot evaluations on various ToF data sources including FLAT, TICaM, and in-house captures confirm strong cross-domain generalization, demonstrating the system’s adaptability, efficiency, and accuracy. An-Nan Xiong, Pingcheng Dong, Yonghao Tan, Yuzhong Jiao, Manto Yung, Ann Li, Luhong Liang, Mansun Chan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-DesignabstractDNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of highprecision partial sums (PSUMs) leads to excessive memory demands in architectures utilizing input/weight stationary dataflows. Traditional compression strategies have typically overlooked PSUM quantization, which may account for 69% of power consumption. This study introduces a novel Additive Partial Sum Quantization (APSQ) method, seamlessly integrating PSUM accumulation into the quantization framework. A grouping strategy that combines APSQ with PSUM quantization enhanced by a reconfigurable architecture is further proposed. The APSQ performs nearly lossless on NLP and CV tasks across BERT, Segformer, and EfficientViT models while compressing PSUMs to INT8. This leads to a notable reduction in energy costs by $\mathbf{2 8-8 7 \%}$. Extended experiments on LLaMA2-7B demonstrate the potential of APSQ for large language models. Code is available at https://github.com/Yonghao-Tan/APSQ. Yonghao Tan, Pingcheng Dong, Yongkun Wu, Yu Liu 0007, Shih-Yang Liu, Xijie Huang, Luhong Liang, Kwang-Ting Cheng |
DAC | 1 |
| 2025 | CoXplorer: Multi-Staged Co-Exploration Framework for AI Model Compression and Accelerator DesignabstractThe rapid evolution of artificial intelligence (AI) algorithms demands efficient computing chips, positioning algorithm-hardware co-design as a crucial optimization strategy. However, automating the co-design process remains challenging due to the lack of a unified exploration framework for both algorithmic and hardware domains, as existing tools - hardware design space exploration (DSE) and compression neural architecture search (Compression NAS) - operate independently, relying entirely on manual collaboration. This paper presents CoXplorer, a co-exploration framework that connects model-compression optimization space and architecture design space. We make three key contributions: (1) a multi-staged co-design space decomposition method that enables systematic exploration of compression-hardware design choices with reduced complexity, (2) an AC-Copilot toolchain enhanced with multi-grained performance modeling driven by hardware simulation-compilation hierarchical cooperation to fulfill various evaluation requirements of co-exploration, enabling balanced simulation accuracy-efficiency trade-offs, and (3) a co-exploration workflow with hierarchical and bottleneck-guided search for harmonizing optimization objectives of both model and hardware design spaces, resulting in improved search efficiency. We validate the CoXplorer on two edge chips, which achieve 53.7% throughput and 45.8% energy efficiency improvements for the CNN acceleration, and 7.5× speedup with 9.9× energy efficiency boost for the Transformer acceleration. A case study on large language model acceleration shows CoXplorer’s extensibility to emerging workloads, enhancing LLAMA2-7B inference throughput from 6.75 to 25.46 tokens/s via co-optimization with compression and near-memory computing architecture. Songchen Ma, Yonghao Tan, Pingcheng Dong, Di Pang, Yu Liu 0007, Luhong Liang, Kwang-Ting Cheng, Fengbin Tu |
ICCAD | 3 |
| 2025 | Exploiting the Memory-Compute-Coupling Feature for CIM Accelerator Design OptimizationabstractSRAM computing-in-memory (CIM) accelerators have evolved as a promising solution to the memory wall problem in neural network (NN) models. By integrating memory and compute resources in each macro, CIM accelerators offer massive in-situ computing parallelism and large memory capacity, enabling spatial mapping with layer fusion and potentially keeping layers stationary in CIM. However, CIM’s memory-compute coupling (MCC) feature poses challenges in designing CIM accelerators. From an architecture aspect, designers must balance CIM’s memory and compute resources by optimizing the macro’s memory-compute ratio (MCR) configuration across diverse scenarios. From a mapping aspect, conventional mappings, which allocate each macro exclusively to each layer, face two major problems: a layer-fusion dilemma (the accelerator suffers from excessive memory access due to layer replications or performance degradation due to load imbalance) and a layer-eviction issue (storing layers stationary in CIM is usually infeasible due to limited CIM capacity). To address these challenges, this paper introduces MCC-DSE, an MCC-aware Design Space Exploration framework for architecture-mapping co-optimization of CIM accelerators. We also propose a three-axis CIM division mapping, which interleaves multiple layers in each macro to concurrently optimize memory access and performance during layer fusion as well as reserves a part of CIM memory in each macro for layer pinning. Compared to baseline architecture and mapping, MCC-DSE shows a 1.4x 8.3x EDP reduction across various workloads and chip areas. Moreover, MCC-DSE provides insights into CIM accelerator optimization, such as selecting optimal MCR and configuring CIM dynamically for different scenarios. Yongkun Wu, Jia Chen 0032, Zhenhua Zhu 0002, Jingyu He, Pingcheng Dong, Yonghao Tan, Xin Zhao 0044, Liang Chang 0002, Yu Wang 0002, Fengbin Tu, Chi-Ying Tsui, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | Genetic Quantization-Aware Approximation for Non-Linear Operations in TransformersabstractNon-linear functions are prevalent in Transformers and their lightweight variants, incurring substantial and frequently underestimated hardware costs. Previous state-of-the-art works optimize these operations by piece-wise linear approximation and store the parameters in look-up tables (LUT), but most of them require unfriendly high-precision arithmetics such as FP/INT 32 and lack consideration of integer-only INT quantization. This paper proposed a genetic LUT-Approximation algorithm namely GQA-LUT that can automatically determine the parameters with quantization awareness. The results demonstrate that GQA-LUT achieves negligible degradation on the challenging semantic segmentation task for both vanilla and linear Transformer models. Besides, proposed GQA-LUT enables the employment of INT8-based LUT-Approximation that achieves an area savings of 81.3~81.7% and a power reduction of 79.3~80.2% compared to the high-precision FP/INT 32 alternatives. Code is available at https://github.com/PingchengDong/GQA-LUT. Pingcheng Dong, Yonghao Tan, Tianwei Ni, Yu Liu 0007, Luhong Liang, Shih-Yang Liu, Xijie Huang, Huaiyu Zhu 0004, Fengwei An, Kwang-Ting Cheng |
DAC | 2 |