Pingcheng Dong

dblp:309/0566 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0003-3975-286XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Adaptive Depth Processing System Based on Foundation Transformers and Time-of-Flight Fusion
abstract
Generating high-quality depth maps with accurate values is a critical research topic, and various depth estimation methods, such as Time-of-Flight (ToF) and Monocular Depth Estimation (MDE), are advancing rapidly. However, each single-sensor approach is inherently limited by the characteristics of its respective modality. Meanwhile, fused approaches using multiple sensors or dedicated trained models are often plagued by system complexity and limited generalization. In this paper, we propose an adaptive depth processing system based on the foundation transformer and ToF fusion, aiming to harness the precision of ToF data and the high-quality depth priors provided by the foundation model in a monocular and zero-shot manner. A hardware–algorithm co-design partitions the computation between a dedicated ASIC for compute-intensive foundation transformer acceleration and an MPSoC FPGA for potential reconfigurable fusion, yielding 36.4 ms latency and 1.9 W power under a 27.6 GOPS/frame workload. Zero-shot evaluations on various ToF data sources including FLAT, TICaM, and in-house captures confirm strong cross-domain generalization, demonstrating the system’s adaptability, efficiency, and accuracy.
An-Nan Xiong, Pingcheng Dong, Yonghao Tan, Yuzhong Jiao, Manto Yung, Ann Li, Luhong Liang, Mansun Chan
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
abstract
DNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of highprecision partial sums (PSUMs) leads to excessive memory demands in architectures utilizing input/weight stationary dataflows. Traditional compression strategies have typically overlooked PSUM quantization, which may account for 69% of power consumption. This study introduces a novel Additive Partial Sum Quantization (APSQ) method, seamlessly integrating PSUM accumulation into the quantization framework. A grouping strategy that combines APSQ with PSUM quantization enhanced by a reconfigurable architecture is further proposed. The APSQ performs nearly lossless on NLP and CV tasks across BERT, Segformer, and EfficientViT models while compressing PSUMs to INT8. This leads to a notable reduction in energy costs by $\mathbf{2 8-8 7 \%}$. Extended experiments on LLaMA2-7B demonstrate the potential of APSQ for large language models. Code is available at https://github.com/Yonghao-Tan/APSQ.
Yonghao Tan, Pingcheng Dong, Yongkun Wu, Yu Liu 0007, Shih-Yang Liu, Xijie Huang, Luhong Liang, Kwang-Ting Cheng
DAC2
2025 CoXplorer: Multi-Staged Co-Exploration Framework for AI Model Compression and Accelerator Design
abstract
The rapid evolution of artificial intelligence (AI) algorithms demands efficient computing chips, positioning algorithm-hardware co-design as a crucial optimization strategy. However, automating the co-design process remains challenging due to the lack of a unified exploration framework for both algorithmic and hardware domains, as existing tools - hardware design space exploration (DSE) and compression neural architecture search (Compression NAS) - operate independently, relying entirely on manual collaboration. This paper presents CoXplorer, a co-exploration framework that connects model-compression optimization space and architecture design space. We make three key contributions: (1) a multi-staged co-design space decomposition method that enables systematic exploration of compression-hardware design choices with reduced complexity, (2) an AC-Copilot toolchain enhanced with multi-grained performance modeling driven by hardware simulation-compilation hierarchical cooperation to fulfill various evaluation requirements of co-exploration, enabling balanced simulation accuracy-efficiency trade-offs, and (3) a co-exploration workflow with hierarchical and bottleneck-guided search for harmonizing optimization objectives of both model and hardware design spaces, resulting in improved search efficiency. We validate the CoXplorer on two edge chips, which achieve 53.7% throughput and 45.8% energy efficiency improvements for the CNN acceleration, and 7.5× speedup with 9.9× energy efficiency boost for the Transformer acceleration. A case study on large language model acceleration shows CoXplorer’s extensibility to emerging workloads, enhancing LLAMA2-7B inference throughput from 6.75 to 25.46 tokens/s via co-optimization with compression and near-memory computing architecture.
Songchen Ma, Yonghao Tan, Pingcheng Dong, Di Pang, Yu Liu 0007, Luhong Liang, Kwang-Ting Cheng, Fengbin Tu
ICCAD4
2025 Memory Efficient Transformer Adapter for Dense Predictions
abstract
While current Vision Transformer (ViT) adapter methods have shown promising accuracy, their inference speed is implicitly hindered by inefficient memory access operations, e.g., standard normalization and frequent reshaping. In this work, we propose META, a simple and fast ViT adapter that can improve the model's memory efficiency and decrease memory time consumption by reducing the inefficient memory access operations. Our method features a memory-efficient adapter block that enables the common sharing of layer normalization between the self-attention and feed-forward network layers, thereby reducing the model's reliance on normalization operations. Within the proposed block, the cross-shaped self-attention is employed to reduce the model's frequent reshaping operations. Moreover, we augment the adapter block with a lightweight convolutional branch that can enhance local inductive biases, particularly beneficial for the dense prediction tasks, e.g., object detection, instance segmentation, and semantic segmentation. The adapter block is finally formulated in a cascaded manner to compute diverse head features, thereby enriching the variety of feature representations. Empirically, extensive evaluations on multiple representative datasets validate that META substantially enhances the predicted quality, while achieving a new state-of-the-art accuracy-efficiency trade-off. Theoretically, we demonstrate that META exhibits superior generalization capability and stronger adaptability.
Pingcheng Dong, Kwang-Ting Cheng
ICLR3
2025 Cyclic Contrastive Knowledge Transfer for Open-Vocabulary Object Detection
abstract
In pursuit of detecting unstinted objects that extend beyond predefined categories, prior arts of open-vocabulary object detection (OVD) typically resort to pretrained vision-language models (VLMs) for base-to-novel category generalization. However, to mitigate the misalignment between upstream image-text pretraining and downstream region-level perception, additional supervisions are indispensable, e.g., image-text pairs or pseudo annotations generated via self-training strategies. In this work, we propose CCKT-Det trained without any extra supervision. The proposed framework constructs a cyclic and dynamic knowledge transfer from language queries and visual region features extracted from VLMs, which forces the detector to closely align with the visual-semantic space of VLMs. Specifically, 1) we prefilter and inject semantic priors to guide the learning of queries, and 2) introduce a regional contrastive loss to improve the awareness of queries on novel objects. CCKT-Det can consistently improve performance as the scale of VLMs increases, all while requiring the detector at a moderate level of computation overhead. Comprehensive experimental results demonstrate that our method achieves performance gain of +2.9% and +10.2% AP_{50} over previous state-of-the-arts on the challenging COCO benchmark, both without and with a stronger teacher model.
Pingcheng Dong, Long Chen 0016
ICLR3
2025 Exploiting the Memory-Compute-Coupling Feature for CIM Accelerator Design Optimization
abstract
SRAM computing-in-memory (CIM) accelerators have evolved as a promising solution to the memory wall problem in neural network (NN) models. By integrating memory and compute resources in each macro, CIM accelerators offer massive in-situ computing parallelism and large memory capacity, enabling spatial mapping with layer fusion and potentially keeping layers stationary in CIM. However, CIM’s memory-compute coupling (MCC) feature poses challenges in designing CIM accelerators. From an architecture aspect, designers must balance CIM’s memory and compute resources by optimizing the macro’s memory-compute ratio (MCR) configuration across diverse scenarios. From a mapping aspect, conventional mappings, which allocate each macro exclusively to each layer, face two major problems: a layer-fusion dilemma (the accelerator suffers from excessive memory access due to layer replications or performance degradation due to load imbalance) and a layer-eviction issue (storing layers stationary in CIM is usually infeasible due to limited CIM capacity). To address these challenges, this paper introduces MCC-DSE, an MCC-aware Design Space Exploration framework for architecture-mapping co-optimization of CIM accelerators. We also propose a three-axis CIM division mapping, which interleaves multiple layers in each macro to concurrently optimize memory access and performance during layer fusion as well as reserves a part of CIM memory in each macro for layer pinning. Compared to baseline architecture and mapping, MCC-DSE shows a 1.4x 8.3x EDP reduction across various workloads and chip areas. Moreover, MCC-DSE provides insights into CIM accelerator optimization, such as selecting optimal MCR and configuring CIM dynamically for different scenarios.
Yongkun Wu, Jia Chen 0032, Zhenhua Zhu 0002, Jingyu He, Pingcheng Dong, Yonghao Tan, Xin Zhao 0044, Liang Chang 0002, Yu Wang 0002, Fengbin Tu, Chi-Ying Tsui, Kwang-Ting Cheng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 Eliminating Semantic Ambiguity in Human Pose Estimation via Stable Feature Upsampling
abstract
Human pose estimation is a challenging research task in the computer vision community due to the semantic ambiguity problem caused by inevitable occlusions, varying body shapes, and complex articulations. Although deep learning-based methods have significantly improved the performance of this task, existing feature upsampling operations,e.g., bilinear interpolation and transposed convolution, within current convolutional neural networks and Transformer frameworks suffer from a multitude of limitations, including the inability to adapt to specific tasks and the loss of fine-grained semantic details. In this work, we propose a simple yet effective two-step stable feature upsampling (SIU) strategy that addresses these limitations by leveraging a learnable and efficient upsampling operation. Specifically, wefirstapply periodic shuffling to increase the resolution of the feature maps.Secondly, we utilize convolution layers to adjust the size of feature channels to match those of the input feature maps. The proposed SIU enables the entire network to adapt to the specific feature requirements of the human pose estimation task, making it more effective in preserving spatial information. Quantitatively, extensive experimental results on the challenging COCO-WholeBody dataset validate that our approach outperforms state-of-the-art methods accurately and efficiently, and possesses strong transferability, making it applicable to a wide range of baselines. Moreover, the qualitative results validate that SIU can effectively eliminate the semantic ambiguity problem in challenging pose scenarios, such as occlusions and overlapping1.
Rui Yan 0010, Xiangbo Shu, Pingcheng Dong, Long Chen 0016, Xiaoyu Du 0002
IEEE Trans. Circuits Syst. Video Technol.5
2024 Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
abstract
Non-linear functions are prevalent in Transformers and their lightweight variants, incurring substantial and frequently underestimated hardware costs. Previous state-of-the-art works optimize these operations by piece-wise linear approximation and store the parameters in look-up tables (LUT), but most of them require unfriendly high-precision arithmetics such as FP/INT 32 and lack consideration of integer-only INT quantization. This paper proposed a genetic LUT-Approximation algorithm namely GQA-LUT that can automatically determine the parameters with quantization awareness. The results demonstrate that GQA-LUT achieves negligible degradation on the challenging semantic segmentation task for both vanilla and linear Transformer models. Besides, proposed GQA-LUT enables the employment of INT8-based LUT-Approximation that achieves an area savings of 81.3~81.7% and a power reduction of 79.3~80.2% compared to the high-precision FP/INT 32 alternatives. Code is available at https://github.com/PingchengDong/GQA-LUT.
Pingcheng Dong, Yonghao Tan, Tianwei Ni, Yu Liu 0007, Luhong Liang, Shih-Yang Liu, Xijie Huang, Huaiyu Zhu 0004, Fengwei An, Kwang-Ting Cheng
DAC1
2024 Live Demonstration: A 1920×1080 129fps 4.3pJ/pixel Stereo-Matching Processor for Low-power Applications
abstract
This demonstration presents an advanced stereo vision system with high energy efficiency. An ov5640 binocular camera, operating at a maximum of 30 frames/second with FHD (1920×1080) resolution, is employed to capture image pairs. A Spartan-7 FPGA rectifies these images with a calibration map matrix and then channels the pixel stream to the stereo-matching processor in a 28nm CMOS process for depth estimation. The resulting depth map, crucial for tasks like obstacle detection and navigation, is displayed in real-time on the monitor for low-power stereo vision applications.
Zhuoyu Chen, Shengming Zhou, Pingcheng Dong, Fengwei An, Lei Chen 0001
ISCAS3
2024 Stereo Matching Accelerator With Re-Computation Scheme and Data-Reused Pipeline for Autonomous Vehicles
abstract
Binocular stereo vision is a depth estimation technique by imitating human eyes. It is widely used in various fields, such as self-driving cars, SLAM, and 3D reconstruction. However, designing a hardware architecture that can balance resource utilization, processing speed, and estimation accuracy remains a significant challenge. This paper proposes a compact and efficient hardware-based design that incorporates linear fitting-based cost fusion, disparity optimization with subpixel interpolation, and multi-directional occlusion filling techniques. Firstly, a gradient-enhanced pipelined matching costs architecture with a resource-saving scheme and re-computation paradigm is proposed to improve the accuracy of edge information. Then, we approximate the nonlinear exponential function by linear fitting to save the hardware resource. Moreover, we design the subpixel interpolation with an SRT radix-4 divider to refine the disparity, which significantly enhances the accuracy of the disparity map in real situations. Finally, we proposed a resource-reused architecture for synchronous hole filling and median filter in post-processing. The disparity map quality of the proposed architecture is evaluated on KITTI2015 datasets, which delivers leading accuracy compared to other state-of-the-art works. The architecture has been successfully implemented and demonstrated on the Stratix-V FPGA platform and achieved 54 frames per second operating at 112 MHz under a resolution of$1920\times 1080$.
Xiwei Fang, Yunhao Ma, Pingcheng Dong, Zhuoyu Chen, Lei Chen 0001, Fengwei An
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 LLM-FP4: 4-Bit Floating-Point Quantized Transformers
abstract
We propose LLM-FP4 for quantizing both weights and activations in large language models (LLMs) down to 4-bit floating-point values, in a post-training manner.Existing posttraining quantization (PTQ) solutions are primarily integer-based and struggle with bit widths below 8 bits.Compared to integer quantization, floating-point (FP) quantization is more flexible and can better handle long-tail or bell-shaped distributions, and it has emerged as a default choice in many hardware platforms.One characteristic of FP quantization is that its performance largely depends on the choice of exponent bits and clipping range.In this regard, we construct a strong FP-PTQ baseline by searching for the optimal quantization parameters.Furthermore, we observe a high interchannel variance and low intra-channel variance pattern in activation distributions, which adds activation quantization difficulty.We recognize this pattern to be consistent across a spectrum of transformer models designed for diverse tasks, such as LLMs, BERT, and Vision Transformer models.To tackle this, we propose per-channel activation quantization and show that these additional scaling factors can be reparameterized as exponential biases of weights, incurring a negligible cost.Our method, for the first time, can quantize both weights and activations in the LLaMA-13B to only 4-bit and achieves an average score of 63.1 on the common sense zero-shot reasoning tasks, which is only 5.8 lower than the full-precision model, significantly outperforming the previous stateof-the-art by 12.7 points.Code is available at: https://github.com/nbasyl/LLM-FP4.
Shih-Yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong, Kwang-Ting Cheng
EMNLP4
2022 A 4.29nJ/pixel Stereo Depth Coprocessor With Pixel Level Pipeline and Region Optimized Semi-Global Matching for IoT Application
abstract
The semi-global matching (SGM) algorithm in stereo vision is a well-known depth-estimation method since it can generate dense and robust disparity maps. However, the real-time processing and low power dissipation, the specifications of the Internet-of-Thing (IoT) applications, are challenging for their computational complexity. In this paper, we propose a hardware-oriented SGM algorithm with pixel-level pipeline and region-optimized cost aggregation for high-speed processing and low hardware-resource usage. Firstly, the matching costs in a region are integrated with an optimization strategy to significantly reduce memory usage and improve the processing speed of the cost aggregation. Then, a two-layer parallel two-stage pipeline (TPTP) architecture, which enables pixel-level processing, is designed to calculate two directions (0° and 135°) aggregation to further solve the crucial computational bottleneck of the SGM algorithm. Finally, the architecture is demonstrated on a low-cost XILINX Spartan-7 device and an advanced Stratix-V FPGA device for VGA ($640\times 480$) depth estimation. The experimental results show that the proposed architecture with compact hardware architecture also ensures accuracy. The pixel-level pipeline architecture enables a processing speed of 355 frames per second (fps) at 109MHz on the Spartan-7 FPGA device and 508 fps at 156MHz on the Stratix-V FPGA. Besides, the coprocessor respectively achieves an energy efficiency of 4.74 nJ/pixel with a power dissipation of 517mW and 4.29nJ/pixel with a power dissipation of 669mW on these two FPGAs.
Pingcheng Dong, Zhuoyu Chen, Zhuoao Li, Yuzhe Fu, Lei Chen 0001, Fengwei An
IEEE Trans. Circuits Syst. I Regul. Pap.1