Yuan Liang 0004

dblp:18/835-4 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-0640-8544ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A Scalable External Memory Access and On-Chip Storage Architecture for Edge-AI Accelerators : - Multi-Path Rolling Data Refresh and Layer-Wise Bank Allocation -
abstract
For resource-constrained AI accelerators applied in edge computing, achieving high power efficiency in neural network (NN) model computation is crucial. However, current designs often overlook the efficiency of off-chip/on-chip data interaction, leading to high latency, which in turn results in suboptimal power efficiency during computation. Additionally, inefficient memory bank allocation further exacerbates latency by causing underutilization of storage resources, thereby contributing to higher overall latency and energy consumption. To address these challenges, this paper proposes a scalable multi-path rolling data refresh and layer-wise bank allocation architecture. The rolling data refresh mechanism enables efficient data interaction between off-chip and on-chip storage, reducing latency and minimizing the area overhead of on-chip memories. The layer-wise bank allocation optimizes on-chip memory utilization according to specific application requirements, improving memory efficiency. A case study on a 28nm AI accelerator demonstrates a 30.6% reduction in area, achieves a power efficiency of 7.36–10.28 TOPS/W, and reduces external memory access by 2.63% to 37.24% on VGG16 and ViT-Small.
Huizi Zhang, Qiufeng Li, Yuan Liang 0004, Zhenzhe Chen, Jinjun Xiong, Mingqiang Huang, Longyang Lin, Masanori Hashimoto
ISLPED4
2025 Genshin: A Generalized Framework with Software-Hardware Co-design and Pruned Fault Injection for Reliability Analysis
abstract
Reliability-demanding devices often require numerous fault injections (FIs) for reliability analysis in the product cycle. However, software-based FI typically demonstrates extremely low efficiency due to low simulation throughput, especially for large-scale designs, while hardware-based FI presents challenges related to complexity of setup and limited scalability. Additionally, FIs often occur in intervals where errors do not affect the system’s outcome, e.g., after final read before next write, necessitating efficient pruning of non-impactful FIs. To address this, a general-purpose FI-specialized framework, Genshin, is proposed for rapid reliability analysis. On the hardware side, we provide an FI-specialized design, which works with Design Under Test (DUT) chips on PCB boards and supports FI control based on the scan chain (SC). An integrated programmable logic allows for flexible and custom FI pattern definitions. Furthermore, an architecturally correct execution (ACE) analysis generates pruned fault tables for DUTs. In Genshin, the SC logic achieves 3,802-65,388 cycles/FI across SC lengths ranging from 2,795 to 61,393 in different DUTs, while the programmable logic enables custom error patterns such as layout-aware multi-bit upset (MBU). Furthermore, the pruned fault tables achieve fault reduction rates from 45.80% to 83.21%.
Hao-Yang Chi, Chien-Hsing Liang, Yu-Hong Chao, Huizi Zhang, Yuan Liang 0004, Wang Liao 0001, Jinjun Xiong, Jing-Jia Liou, Masanori Hashimoto, Longyang Lin
ITC6
2025 An 88.5 fsrms Integrated Jitter and -76.2 dBc Reference Spur mmW PLL Utilizing a Ripple Compensation Phase/Frequency Detector
abstract
Millimeter-wave (mmW) phase-locked loops (PLLs) typically favor a wide loop bandwidth for stronger suppression of the out-of-band phase noise from a voltage-controlled oscillator (VCO). Unfortunately, doing so lowers the degree of attenuation to the PLL reference spurs. This paper proposes a ripple compensation phase detector (RCPD) for extending PLL loop bandwidth and phase noise suppression without sacrificing reference spur performance. The RCPD inherently consists of a pair of PDs that generate respective ripple simultaneously, with each PD’s ripple current compensating the other, resulting in a glitch-free RCPD output. A calibrator is also introduced to reduce device mismatches. With the proposed techniques, the proposed mmW PLL was implemented using 22 nm bulk CMOS technology. The mmW PLL operates from 32.7 to 39.4 GHz, achieving an integrated jitter and reference spur of 88.5 fsrms (1 kHz to 100 MHz) and –76.2 dBc, respectively, with a figure-of-merit (FoM) of –247.5 dB.
Yuan Liang 0004, Zhongyuan Fang, Masanori Hashimoto
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 Millimetre-Wave and Terahertz Antennas and Directional Coupler Enabled by Wafer-Level Packaging Platform with Interposer
abstract
The recent development of wafer-level, low-cost packaging platforms based on the silicon interposer and direct wafer bonding has paved a new way toward high-performance millimeter-wave to terahertz 2.5/3D heterogeneous integration, enabling high-speed wireless communication and chip-to-chip interconnects. Using this emerging technology, several passive components are studied in this paper toward the terahertz applications. One 240 GHz distributed mushroom antenna is designed using interposers to form multiple resonances unit- cell, achieving 7.16 dBi gain with 76% radiation efficiency. A 300 GHz differential patch antenna is designed, delivering 5.6 dBi gain with 88% radiation efficiency. A 60 GHz directional coupler based on interleaving topology is proposed and simulated, showing a 3.3 dB coupling factor with more than 30 GHz bandwidth. These preliminary results reveal a promising solution by using the wafer-to-wafer packaging platform to build high-performance passive building blocks.
Yuan Liang 0004, Chirn Chye Boon, Qian Chen 0027, Yangtao Dong
ISCAS1
2020 A 3GS/s Highly Linear Energy Efficient Constant-Slope Based Voltage-to-Time Converter
abstract
This paper presents a high speed highly linear energy-efficient constant-slope based voltage-to-time converter (VTC). By combining sample-and-hold with constant charging process, we achieve precise sampled step and linear charging ramp concurrently without using the extra control clock. Simple calibration has been implemented to overcome conversion gain variation due to process-voltage-temperature (PVT) variation. The post-layout simulation results show that the SFDR/-THD of the proposed VTC reaches 58.7dB/56.3dB at 3GS/s near Nyquist (typical corner). The VTC achieves 144ps output range with 0.83mW power consumption at 3GHz. It occupies an active area of 41.3um × 34.4um implemented in 65nm CMOS.
Qian Chen 0027, Yuan Liang 0004, Bongjin Kim, Chirn Chye Boon
ISCAS2
2020 Multi-Channel FSK Inter/Intra-Chip Communication by Exploiting Field-Confined Slow-Wave Transmission Line
abstract
Using on-chip slow-wave transmission line (SW-TL) has paved a new way towards millimeter-wave (mm-wave) to terahertz (THz) low power and high speed inter-/intra-chip communications. This work presents an on-chip SW-TL featured by periodic comb-shape grooves with capability to strongly localize electric-field. A gradient groove structure is proposed to serve as the mode converter and performs the mode transformation between the quasi-TEM wave and the slow-wave with low return loss. Due to field confinement, when two SW-TL are only 2.4 μm apart, more than 19 dB crosstalk suppression is observed compared with two conventional TL with the same metal spacing. A dual-channel 160 GHz frequency-shift keying (FSK) transceiver is designed in 65 nm CMOS technology. The preliminary results show that by exploiting SW-TL as the silicon channel, the receiver can recover error-free 4 Gb/s dual-channel data, whereas the eye diagram of the transceiver using traditional transmission line (TL) is fully distorted. The transceiver consumes 36 mW DC power from a 1.2 V power supply.
Qian Chen 0027, Chirn Chye Boon, Xueyong Zhang, Chenyang Li 0008, Yuan Liang 0004, Zhe Liu 0038, Ting Guo 0001
ISCAS5
2020 A 6bit 1.2GS/s Symmetric Successive Approximation Energy-Efficient Time-to-Digital Converter in 40nm CMOS
abstract
This work presents a 6bit 1.2GS/s symmetric successive approximation (SSA) energy-efficient time-to-digital converter (TDC). The delay offset of the successive approximation (SA) TDC has been alleviated by employing the balanced architecture and optimizing the phase detector (PD). Size-optimized inverter chain is deployed as the delay unit with good linearity to increase the conversion rate and reduce the power consumption. In addition, dynamic logic is implemented to further improve the speed and energy efficiency. As a proof-of-concept design, the TDC is verified by post-layout simulation (Transient noise + Monte Carlo (MC)) in 40nm low power CMOS technology, achieving 0.98LSB /1.03LSB worst case differential nonlinearity (DNL)/integral nonlinearity (INL) and 0.014pJ/conversion-step figure-of-merit (FOM). The simulated single-shot precision (SSP) of the proposed TDC is 0.71 LSB.
Qian Chen 0027, Yuan Liang 0004, Chirn Chye Boon
ISCAS2
2019 Design and Analysis of $D$ -Band On-Chip Modulator and Signal Source Based on Split-Ring Resonator
abstract
In an effort toward high-speed and low-power I/O data link in the future exascale data server, this paper presents a signal source and a modulator in the D-band. The split-ring resonator (SRR) structures are used to boost both the signal power and the extinction ratio (ER). The modulator manifests itself as a compact SRR whose magnetic resonance frequency can be modulated by high-speed data. Such a magnetic metamaterial achieves a significant reduction of radiation loss with high ER by stacking two auxiliary SRR unit cells with interleaved placement. The high-Q tank for oscillation is realized by a stacked SRR decorated with slow-wave transmission line (T-line) for electric field confinement. A four-way power-combined fundamental 80-GHz coupled-oscillator network is magnetically synchronized by the slow-wave T-line, which is frequency doubled to 160 GHz. Fabricated in the 65-nm CMOS process, the measured results show that: 1) the modulator achieves 3-dB insertion loss at the onstate with 43-dB isolation at the off-state, leading to a 40-dB ER at 125 GHz within an area of only 40 μm×67 μm and 2) the signal source achieves 6.3% frequency tuning range (FTR) with 3.7-mW peak output power at 160 GHz within 0.053-mm2active area. It has a measured phase noise of -105 dBc/Hz at 10-MHz offset, 5.5% dc-to-RF power efficiency, 70.1-mW/mm2power density, FOM of -171 dBc/Hz, and FOMT of -172.7 dBc/Hz.
Yuan Liang 0004, Chirn Chye Boon, Chenyang Li 0008, Xiao-Lan Tang, Herman Jalli Ng, Dietmar Kissinger, Hao Yu 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2015 An energy efficient and low cross-talk CMOS sub-THz I/O with surface-wave modulator and interconnect
abstract
Free-space EM-wave based GHz interconnect has significant loss and crosstalk that cannot be deployed as low-power and dense I/Os for future network-on-chip (NoC) integration of many-core and memory. This paper proposes an energy-efficient and low-crosstalk sub-THz (0.1T-1T) I/O with use of surface-wave based modulator and interconnects in CMOS. By introducing sub-wavelength periodical corrugation structure onto transmission line, the surface-wave is established to propagate signal that is strongly localized on surface of top-layer metal wire, which results in low coupling into lossy substrate and neighboring metal wires. As such, significant power saving and cross-talk reduction can be observed with high communication bandwidth. In addition, a high on/off-ratio surface-wave modulator is also proposed to support on-chip THz communication. As designed in 65nm CMOS, the results have shown that the proposed surface-wave I/O interface achieves 25Gbps data rate and 0.016pJ/bit/mm energy efficiency at 140GHz carrier frequency over 20mm surface-wave channels. They can be placed with 2.4μm channel spacing and a -20dB crosstalk ratio. The surface-wave modulator also achieves significant reduction of radiation loss with 23dB extinction ratio.
Yuan Liang 0004, Hao Yu 0001, Junfeng Zhao 0003, Yuangang Wang
ISLPED1