VLDB 2026 Research / reviewers in the wild / expert
Tsung-Te Liu
dblp:84/8270
· DBLP profile ↗
18ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-5433-9830ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Lightweight RISC-V Multiply-Accumulate Extension with Data Preloading Scheme
Meng-Hsueh Lee, Tsung-Te Liu |
ISCAS | 3 |
| 2026 | A Lightweight Speaker Verification Processor based on Large-Scale Dataset
Jui-Yang Hsu, Yun-Hwa Tsou, Tsung-Te Liu |
ISCAS | 3 |
| 2025 | Device Noises Resilient Training and Inference Framework for Smart Sensing on Analog Computing In MemoryabstractNeural networks have demonstrated superior performance over rule-based and model-based approaches in processing noisy sensing data. However, their substantial computational and energy demands hinder deployment in battery-powered embedded systems. Computing-in-Memory (CIM) devices offer a promising alternative by significantly reducing energy consumption. Prior work [2] achieves this by leveraging a non-von Neumann architecture, which minimizes data movement between memory and compute units, thereby mitigating the memory wall bottleneck. Despite these advantages, analog CIM (ACIM) systems face several key challenges, including analog noise, limited numerical precision, and increased hardware complexity. While cloud-based neural networks are still dominant, emerging applications increasingly demand real-time, privacy-preserving inference on-device. For instance, facial authentication requires local execution on edge devices to ensure low-latency responsiveness and to protect user privacy. CIM architectures are particularly well-suited to these scenarios due to their tightly integrated memory-compute structure, offering low-latency and energy-efficient inference capabilities. Xin-You Liu, Chi-Sheng Shih 0001, Tsung-Te Liu, Pei-Kuei Tsung, Chih-Wei Chen, Chieh-Fang Teng |
CASES | 3 |
| 2024 | Low-Complexity Algorithmic Test Generation for Neuromorphic ChipsabstractNeuromorphic chips are promising hardware implementations for artificial intelligence (AI) applications owing to their low power consumption. However, neuromorphic chips are difficult to test since they have many potential configurations but lack design for testability (DfT). We propose an algorithmic test generation method for neuromorphic chips without DfT, including fault activation and fault propagation. Fault activation differentiates a neuron's good output and faulty output. Fault propagation sensitizes fault effects to differentiate outputs of faulty chips and good chips. On an L-layer Spiking Neural Network (SNN) model, we achieve 100% fault coverage using O(L) test configurations and test patterns under negligible or no weight variation. Our results show that test effectiveness is maintained even with 4-bit weight quantization. We incur no test escape and overkill even under 10% weight variation. Our total test length is over 73K times shorter than previous works. Hsu-Yu Huang, Chu-Yun Hsiao, Tsung-Te Liu, Chien-Mo James Li |
DAC | 3 |
| 2024 | Highly Reliable PUF Circuits Using Efficient Post-Processing Stabilization TechniqueabstractA Physically Unclonable Function (PUF) circuit is purposefully engineered to leverage the inherent variations in manufacturing processes to generate distinct identities. However, the efficacy of PUFs can be compromised due to the influence of external factors such as noise and environmental variations, which can lead to instability of identities. In this paper, a novel post-processing stabilization technique is introduced, utilizing a mismatch recombination algorithm and resulting in a 31× reduction of the Bit Error Rate (BER) to 0.19%. Through the application of this stabilization technique in conjunction with a modified self-compared Ring Oscillator (RO) architecture, the proposed PUF achieves a uniqueness of 49.96%. Moreover, the proposed PUF exhibits substantial resilience against temperature variations, maintaining a BER of under 2% across the temperature range from 0°C to 45°C. The proposed PUF was implemented and verified on a Xilinx Artix-7 FPGA. Yu-Hsiang Tseng, Shao-Hong Yang, Tsung-Te Liu |
ISCAS | 3 |
| 2024 | An Efficient FPGA-Based Dilated and Transposed Convolutional Neural Network AcceleratorabstractThis work presents a Field Programmable Gate Array (FPGA)-based deep neural network (DNN) accelerator that can maintain consistently high efficiency when executing various neural network architectures, including convolutional neural network (CNN), transposed and dilated convolution (TD-convolution) operations for modern computer vision (CV) tasks. To deal with the utilization degradation issue with a large processing unit (PE) array, a 3-D mapping strategy that adaptively tailors different layer configurations is proposed to optimize the parallelism dimensions of the PE, which significantly increases the hardware utilization to enhance the accelerator efficiency. Moreover, to minimize the implementation and performance overhead resulting from the TD-convolution operations, a unified processing flow is proposed to realize an integrated operation of traditional and TD-convolution. This allows the accelerator to bypass redundant zero operations, further boosting overall efficiency. The 4096-PE accelerator implementation on Intel Stratix 10 FPGA achieves a throughput performance of 2.597–2.870 TOPS with an efficiency of 0.63-0.70 GOPS/DSP across various DNN networks. This represents$1.72\times $and$1.73\times $improvement in throughput and efficiency, respectively, compared to the state-of-the-art designs. Tsung-Hsi Wu, Tsung-Te Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | LC4SV: A Denoising Framework Learning to Compensate for Unseen Speaker Verification ModelsabstractThe performance of speaker verification (SV) models may drop dramatically in noisy environments. A speech enhancement (SE) module can be used as a front-end strategy. However, existing SE methods may fail to bring performance improvements to downstream SV systems due to artifacts in the predicted signals of SE models. To compensate for artifacts, we propose a generic denoising framework named LC4SV, which can serve as a pre-processor for various unknown downstream SV models. In LC4SV, we employ a learning-based interpolation agent to automatically generate the appropriate coefficients between the enhanced signal and its noisy input to improve SV performance in noisy environments. Our experimental results demonstrate that LC4SV consistently improves the performance of various unseen SV systems. To the best of our knowledge, this work is the first attempt to develop a learning-based interpolation scheme aiming at improving SV performance in noisy environments. Chi-Chang Lee, Chu-Song Chen, Hsin-Min Wang, Tsung-Te Liu, Yu Tsao 0001 |
ASRU | 5 |
| 2023 | A Speculative Computation Approach for Energy-Efficient Deep Neural NetworkabstractDeep neural networks (DNNs) have been widely used for data processing and analysis nowadays. Many computational techniques have been proposed to improve the energy efficiency of executing DNNs, which is critical for emerging smart edge applications. This article presents a speculative computation approach to improving the energy efficiency of DNN computations. The proposed approach employs the techniques of input channel partitioning and threshold-based negative masking to predict and eliminate unnecessary computations. Moreover, a systematic procedure of threshold optimization is proposed to achieve the best tradeoff between the energy and accuracy performance. Finally, an energy-efficient DNN processor architecture was designed and implemented to support the proposed speculative computation approach. The experimental results show that the proposed DNN processor with speculative computation can enhance the energy efficiency of the processor by 22.8%, with only 0.96% accuracy degradation and 1% implementation overhead. Rui-Xuan Zheng, Ya-Cheng Ko, Tsung-Te Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | A Highly Stable Physically Unclonable Function Using Algorithm-Based Mismatch Hardening Technique in 28-nm CMOSabstractPhysically unclonable function (PUF) is an emerging security solution for Internet of Things (IoT) devices. However, PUF faces a critical design challenge: responses should remain the same regardless of environmental conditions. This article presents an algorithm-based mismatch hardening technique that provides an effective and efficient solution to stabilizing PUF responses, which could be applied to various PUF circuits. The proposed stabilization technique combines multiple mismatches to achieve high reliability with minimum loss of utility. Moreover, the proposed approach requires only the available digitized PUF responses, avoiding any auxiliary measurement circuit to minimize additional implementation and testing overhead. The PUF test chip implemented in 28-nm CMOS technology shows that the proposed stabilization technique achieves a highly stable performance, lowering the nominal bit error rate (BER) to 0.0016%. It also exhibits excellent reliability with a worst-case BER of 0.16% across 0.7 to 1.2 V and −50 to 125°C, proving to be a promising candidate for security primitive for IoT applications. Shao-Hong Yang, Tsung-Te Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Multi-Robot Formation Control using Collective Behavior Model and Reinforcement LearningabstractA multi-robot system has advantages in complex tasks, where formation control is one of the most critical and fundamental tasks. For small-sized, autonomous, and enduring robots, realizing high energy and area efficiency is extremely important. This paper presents a approach that combines swarm intelligence and reinforcement learning to realize accurate and reliable operations. An area-energy-efficient hardware architecture is proposed to perform formation control in a distributed robotic system. The proposed system demonstrates substantially lower cost and power consumption when compared with the state-of-the-art designs. Jung-Chun Liu, Tsung-Te Liu |
ISCAS | 2 |
| 2022 | A Robust Area-Efficient Physically Unclonable Function With High Machine Learning Attack Resilience in 28-nm CMOSabstractStrong physically unclonable function (PUF) offers a promising solution to low-cost hardware identification and authentication for Internet of Things (IoT) applications. The continuous advancement of machine learning (ML) technology makes the PUF resilience to ML attacks a major design priority. This paper presents a robust and area-efficient strong PUF design with high ML attack resilience. The proposed PUF architecture based on inverter amplifiers operating in the subthreshold region achieves both low energy consumption and high supply and temperature scalability. The proposed nonlinearity topology effectively enhances PUF resilience to various ML attacks with low implementation area and cost. The proposed strong PUF design was designed and implemented using a 28-nm CMOS process. The measurement results show that the proposed PUF design achieves a nearly ideal ML attack resilience of 49.96 % with a small area of 239,857 F2, and demonstrates a stable operation across a wide range of supply voltage from 0.5–1.4 V and temperature from −40–100 °C. This represents$3\times $improvement in area efficiency,$2.25\times $and$1.08\times $improvement in operating voltage and temperature range, respectively, compared to the state-of-the-art results. You-Cheng Lai, Chun-Yen Yao, Shao-Hong Yang, Ying-Wei Wu, Tsung-Te Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | Clock-Less DFT and BIST for Dual-Rail Asynchronous Circuits
Tsai-Chieh Chen, Chia-Cheng Pai, Yi-Zhan Hsieh, Hsiao-Yin Tseng, Chien-Mo James Li, Tsung-Te Liu, I-Wei Chiu |
J. Electron. Test. | 6 |
| 2019 | Variation-Resilient Design Techniques for Energy-Constrained SystemsabstractProcess, voltage, and temperature (PVT) variations substantially increase the variability of digital CMOS logics and reduce the operation robustness, especially for energy-constrained systems with aggressive voltage scaling. This paper reviews several variation-resilient design techniques for addressing PVT variations to improve the energy efficiency of digital CMOS VLSI circuits. The scope includes static and adaptive design techniques for design-time and run-time optimization, respectively. In addition, an emerging adaptive design strategy combining the fully integrated voltage regulator for system-level optimization is also introduced. Bing-Chen Wu, Tsung-Te Liu |
IOLTS | 2 |
| 2013 | TEASE: a systematic analysis framework for early evaluation of FinFET-based advanced technology nodesabstractThis paper proposes TEASE (Technology Exploration and Analysis for SoC-level Evaluation), a framework to systematically analyze and evaluate system design in finFET-based technology node. The proposed framework combines both lithography and electrical constraints of a particular technology node to optimize the standard cell library performance. Growing complexity of logic design at nodes below 20nm causes to adopt a design style that can embrace the simplicity required to enable manufacturing, along with a process technology that can be finely tuned to the desired performance constraints. Additionally, the introduction of finFET based devices poses a new challenge for the designers to come up with an efficient standard cell template. The proposed framework can be used to detect the technology constraints that act as the bottleneck for the enablement of design at these advanced nodes. Results presented in this paper show by optimizing these bottlenecks we can improve the performance of a standard cell library significantly. Furthermore, adapting to such an analysis framework at an early stage of technology development helps to take the design constraints into the decision loop for realization of technology research into real products. Arindam Mallik, Paul Zuber, Tsung-Te Liu, Bharani Chava, Bhavana Ballal, Pablo Royer Del Bario, Rogier Baert, Kris Croes, Julien Ryckaert, Mustafa Badaroglu, Abdelkarim Mercha, Diederik Verkest |
DAC | 3 |
| 2011 | A low-leakage parallel CRC generator for ultra-low power applicationsabstractUnlike static CMOS circuits, the standby energy in sense amplifier-based pass transistor logic (SAPTL) circuits can be decoupled from its performance, allowing separate optimization strategies for leakage and speed. In this paper, a 64-byte parallel cyclic-redundancy check (CRC) generator is designed and implemented using asynchronous 90nm SAPTL circuits, with a simulated minimum energy point that is 25% lower than the equivalent complementary static CMOS implementation. The low leakage operation results in (a) an 7.9X reduction in measured energy when VDDis reduced from 1V to 0.3V at α = 0.1 and (b) a 10% reduction in measured delay with a stack forward body bias of 0.4V, with no corresponding increase in energy. Louis P. Alarcón, Tsung-Te Liu, Jan M. Rabaey |
ISCAS | 2 |
| 2011 | Linearity analysis of CMOS passive mixerabstractThe analysis of distortion behavior in a CMOS passive mixer is presented. We use a simple device model with continuous equation to characterize the MOSFET switch in different operating regions. The nonlinear behaviors of passive mixers based on different switch topologies are investigated with power series analysis, which demonstrates that the CMOS switch based approach exhibits a better linearity performance due to the distortion cancellation of even order nonlinearity. The calculated conversion gain and linearity performances are compared with simulation results. Tsung-Te Liu, Jan M. Rabaey |
ISCAS | 1 |
| 2010 | Ultralow-Power Design in Near-Threshold RegionabstractOperation in the subthreshold region most often is synonymous to minimum-energy operation. Yet, the penalty in performance is huge. In this paper, we explore how design in the moderate inversion region helps to recover some of that lost performance, while staying quite close to the minimum-energy point. An energy-delay modeling framework that extends over the weak, moderate, and strong inversion regions is developed. The impact of activity and design parameters such as supply voltage and transistor sizing on the energy and performance in this operational region is derived. The quantitative benefits of operating in near-threshold region are established using some simple examples. The paper shows that a 20% increase in energy from the minimum-energy point gives back ten times in performance. Based on these observations, a pass-transistor based logic family that excels in this operational region is introduced. The logic family operates most of its logic in the above-threshold mode (using low-threshold transistors), yet containing leakage to only those in subthreshold. Operation below minimum-energy point of CMOS is demonstrated. In leakage-dominated ultralow-power designs, time-multiplexing will be shown to yield not only area, but also energy reduction due to lower leakage. Finally, the paper demonstrates the use of ultralow-power design techniques in chip synthesis. Dejan Markovic, Cheng C. Wang, Louis P. Alarcón, Tsung-Te Liu, Jan M. Rabaey |
Proc. IEEE | 4 |
| 2009 | Asynchronous Computing in Sense Amplifier-Based Pass Transistor LogicabstractThis paper presents the design and implementation of a low-energy asynchronous logic topology using sense amplifier-based pass transistor logic (SAPTL). The SAPTL structure can realize very low energy computation by using low-leakage pass transistor networks at low supply voltages. The introduction of asynchronous operation in SAPTL further improves energy-delay performance without a significant increase in hardware complexity. We show two different self-timed approaches: 1) the bundled data and 2) the dual-rail handshaking protocol. The proposed self-timed SAPTL architectures provide robust and efficient asynchronous computation using a glitch-free protocol to avoid possible dynamic timing hazards. Simulation and measurement results show that the self-timed SAPTL with dual-rail protocol exhibits energy-delay characteristics better than synchronous and bundled data self-timed approaches in 90-nm CMOS. Tsung-Te Liu, Louis P. Alarcón, Matthew D. Pierson, Jan M. Rabaey |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |