Tiancheng Cao

dblp:248/8567 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-7259-5192ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A High-Accuracy Probabilistic-Based Sigmoid Approximator Incorporating Memory-Saving and Time-Efficient Strategies
abstract
The sigmoid function, as a widely used activation function in neural networks, has gained much attention for its approximation and associated usage in edge devices. A recent study applied the Gaussian cumulative function to approximate the sigmoid function. Although this probabilistic method simplifies hardware implementation through a low-complexity binary search, it requires intensive random access memory (RAM) storage, and the search process is time-consuming. Besides, it targets minimizing the maximum mapping error rather than ensuring accurate approximation across all inputs. To address these issues, this article proposes a hardware-friendly and high-accuracy probabilistic-based sigmoid approximator. We first present that given an input, the output of a sigmoid function is strictly equivalent to the probability of a logistic random variable less than or equal to this input. Then, an indirect random variable quantizing strategy is exhibited to reduce memory usage and concurrently minimize precision loss. The latency for the proposed scheme is also optimized. Afterward, a resource-efficient and low-latency sigmoid approximator is developed on digital circuits. Finally, we derive an upper bound on the absolute error between the approximator's output and the true value. Experiments verify the usefulness of our scheme and showcase superior performance in approximation accuracy and resource cost.
Wenhao Lu, Andrew Chi-Sing Leung, Tiancheng Cao, Yucen Shi, Yiping Ke, Zhenya Zang
IEEE Trans. Neural Networks Learn. Syst.4
2026 FPGA-Based Real-Time ECG Classification System Using Quantized Inception-ResNeXt Neural Network and CWT Approximation
abstract
This article presents a software–hardware codesigned field-programmable gate array (FPGA)-based real-time electrocardiogram (ECG) classification system that combines methodological and practical innovations to achieve state-of-the-art performance with an ultracompact model. On the software side, we introduce a hardware-adaptive, configurable quantization-aware training (QAT) framework that enables layerwise precision assignment and flexible quantization, ensuring that the trained model is highly accurate and hardware-friendly even at ultralow bit widths. On the hardware side, we propose a resource-efficient FPGA accelerator featuring a streaming architecture and a cosine-approximated continuous wavelet transform (CWT) module, optimized for low-power and real-time inference. Implemented in an FPGA, we demonstrate that a six-layer Inception-ResNeXt (IRN) network can achieve 99.5% inference accuracy on the MIT-BIH ECG dataset with 200-mW dynamic power and 0.0767-mJ/inference energy efficiency.
Tiancheng Cao, Wei Soon Ng, Wang Ling Goh, Yuan Gao 0011, Hen-Wei Huang
IEEE Trans. Very Large Scale Integr. Syst.1
2026 FPGA Implementation of PoolFormer Network Using Python-Driven High-Level Synthesis Framework for Edge-AIoT Speech Recognition
Tiancheng Cao, Wei Soon Ng, Wang Ling Goh, Yuan Gao 0011
IEEE Trans. Very Large Scale Integr. Syst.1
2025 Neuromorphic FeRAM-Based Co-Design for Imaging Enhancement in Handheld Photoacoustic Systems
abstract
This paper introduces a novel platform designed to enhance the imaging quality of handheld photoacoustic imaging (PAI) systems, addressing the limitations of current portable PAI devices. The platform integrates the MultiResU-Net imaging enhancement algorithm with a Ferroelectric random-access memory (FeRAM) crossbar array, enabling efficient in-memory computing that is highly suitable for deep neural networks involving extensive matrix multiplications. The hardware implementation is optimized for low-power operation on edge devices, and a specifically designed algorithmic strategy is introduced to accurately simulate hardware variations with a time complexity of O(mn). The feasibility and effectiveness of this approach are demonstrated through simulations using synthesized and in vivo data, showing a more than tenfold improvement in imaging resolution. The neural network inference is significantly accelerated, completing within microseconds, thereby fully supporting real-time imaging. The entire platform is compact, with dimensions of 25×25×20 cm3, making it a portable, high-resolution, real-time imaging solution for personalized healthcare.
Tiancheng Cao, Zhengyuan Zhang 0002, Shuailin Tao, Chen Liu 0009, Wang Ling Goh, Yuanjing Zheng, Yuan Gao 0011
ISCAS1
2025 A Gait Data Compression and Reconstruction Framework for Edge Device using Low-Dimensional Attention Model with Autoencoder
abstract
This paper presents a gait data compression and reconstruction framework based on a low-dimensional attention model with autoencoder. By reducing the size of the attention filter to match the maximum matrix rank, the dimensionality of the attention filter can be reduced to enhance the compression ratio. Extensive evaluations using MHEALTH dataset demonstrated that the proposed method can achieve compression ratio of 24 with low reconstruction error of Percent Root Mean Square Difference (PRD) of 0.0323, Correlation Coefficient (CC) of 0.9510, and Signal-to-Noise Ratio Loss (SNRL) of 1.51 dB. The proposed compression model is implemented in hardware using microcontroller. Fixed-point quantization and optimized Softmax layer representation are performed to reduce the hardware resources requirement.
Shuailin Tao, Wang Ling Goh, Tiancheng Cao, Yuan Gao 0011
ISCAS3
2025 Cybersecure End-to-End FPGA-Accelerated ECG Monitoring for Precision Diagnosis With Personalized CWT and Adversarial Defense
abstract
Advancements in wearable technology and edge computing have transformed cardiovascular monitoring, driving the demand for private, secure, and real-time diagnostic solutions. This paper presents an edge-wearable ECG monitoring system that integrates personalized continuous wavelet transform (CWT) preprocessing, a DeepFool-FGSM adversarial defense, and an optimized parallel PoolFormer architecture for resource-constrained FPGA deployment. The personalized CWT captures individual-specific ECG features and mitigates model-inversion privacy risks. The defense approach balances robustness and computational efficiency and reduces hardware complexity and energy via quantization-aware training (QAT). Evaluations on field programmable gate array (FPGA) confirm high diagnostic accuracy (98.93%), real-time inference (latency <1.7 ms), and improved robustness against adversarial perturbations, with 0.055 W FPGA-core power. Together, the system delivers confidentiality, integrity, and availability for cybersecure, personalized ECG monitoring at the edge.
Tiancheng Cao, Wei Soon Ng, Rong Tan, Dawei Wang 0011, Hen-Wei Huang
IEEE J. Biomed. Health Informatics1
2025 Edge PoolFormer: Modeling and Training of PoolFormer Network on RRAM Crossbar for Edge-AI Applications
abstract
PoolFormer is a subset of Transformer neural network with a key difference of replacing computationally demanding token mixer with pooling function. In this work, a memristor-based PoolFormer network modeling and training framework for edge-artificial intelligence (AI) applications is presented. The original PoolFormer structure is further optimized for hardware implementation on RRAM crossbar by replacing the normalization operation with scaling. In addition, the nonidealities of RRAM crossbar from device to array level as well as peripheral readout circuits are analyzed. By integrating these factors into one training framework, the overall neural network performance is evaluated holistically and the impact of nonidealities to the network performance can be effectively mitigated. Implemented in Python and PyTorch, a 16-block PoolFormer network is built with$64\times 64$four-level RRAM crossbar array model extracted from measurement results. The total number of the proposed Edge PoolFormer network parameters is 0.246 M, which is at least one order smaller than the conventional CNN implementation. This network achieved inference accuracy of 88.07% for CIFAR-10 image classification tasks with accuracy degradation of 1.5% compared to the ideal software model with FP32 precision weights.
Tiancheng Cao, Weihao Yu 0001, Yuan Gao 0011, Chen Liu 0009, Shuicheng Yan, Wang Ling Goh
IEEE Trans. Very Large Scale Integr. Syst.1
2023 RRAM-PoolFormer: A Resistive Memristor-based PoolFormer Modeling and Training Framework for Edge-AI Applications
abstract
PoolFormer is a type of neural network architecture that is abstracted from Transformer where the computationally heavy token mixer module is replaced with simple pooling function. This paper presents a memristor-based PoolFormer modeling and training framework for edge-AI applications. To fit for implementation on resistive crossbar array, original PoolFormer structure is further optimized by replacing normalization operation with hardware friendly scaling operation. In addition, the non-idealities of RRAM crossbar from device to array level as well as peripheral readout circuits are also included. By incorporating these elements under a single framework for network training, their impact to the network performance can be effectively mitigated. This framework is implemented in a combination of Python and PyTorch. A 16-block PoolFormer network is designed and optimized for CIFAR-10 image classification tasks using measured$\mathbf{64}\times \mathbf{64}$RRAM crossbar array results. The total network weights are only 0.26M, which is at least one order of magnitude smaller than that of the conventional DNN implementation. When compared to the ideal model with FP64 weight bit-length, 85.86% inference accuracy is reached with only 4-level weight resolutions and less than 4% accuracy loss.
Tiancheng Cao, Weihao Yu 0001, Yuan Gao 0011, Chen Liu 0009, Shuicheng Yan, Wang Ling Goh
ISCAS1
2019 A Method of Ontology Evolution and Concept Evaluation Based on Knowledge Discovery in the Heavy Haul Railway Risk System
Tiancheng Cao, Wenxin Mu, Aurélie Montarnal, Anne-Marie Barthe-Delanoë
PRO-VE1