VLDB 2026 Research / reviewers in the wild / expert
Patrick Chiang 0001
dblp:13/1055 · also Patrick Yin Chiang
· DBLP profile ↗
34ranked-venue papers
0as first author
12since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 4x106 Gb/s PAM4 Current-Reuse VCSEL Driver with a High-Bandwidth VGA in 130 nm BiCMOS
Wei Chen 0176, Minhao Li, Pisen Zhou, Patrick Chiang 0001, Ronghua Ni |
ISCAS | 7 |
| 2025 | The Photoacoustic Quality-Enhancement Neural Network Processor with the Scalable and End-to-End Architecture by Improving the Sparsity LevelabstractRecent advancements have marked significant progress in photoacoustic imaging as an effective method for acquiring deep bio-tissue visuals in modern medical clinical therapy and the efficacy of U-Net and its variants has been established for imaging quality enhancement in this field. Unlike common computer vision datasets such as ImageNet [1] and PASCAL VOC [2], biomedical images exhibit highly structured patterns, low spatial resolution, and single-channel modality, as shown in Fig. 1. Additionally, the U-Net parameters trained for medical super-resolution tasks demonstrate a high sparsity ratio, making them suitable for implementation on edge-computing platforms. Therefore, developing an energy-efficient photoacoustic imaging setup in this area is a natural progression. However, this development is constrained by the current neural network architectures, which are built around a U-Net backbone. The multi-stage feature extractor, skip connection integration across different blocks, and the encoder-decoder backbone design pose significant challenges to cutting-edge computational hardware platforms. In this study, a scalable, sparsity-supported neural network accelerator architecture for bio-tissue imaging quality enhancement is proposed to meet the stringent requirements of latency and energy efficiency, as depicted in Fig. 2. This architecture achieves desired performance improvements by exploring the sparsity possibilities in neural network during the training process and implementing an end-to-end pixel-first hardware design to minimize data movement and support sparsity computation. Compared with the state-of-the-art related works, this optimized architecture has achieved minimum on-chip storage overhead and the fastest frame for the application of photoacoustic imaging quality enhancement. The scalable architecture has also been implemented on a Xilinx XCZU9EG FPGA and attains a performance of PSNR@ 24 dB and a frame rate of 164 fps at a working frequency of 250 MHz. Zhengyuan Zhang 0002, Caijie Liang, Boyi Dong, Yange Wang, Zhongzhiguang Lu, Xiangjun Yin, Shenglong Zhuo, Yifan Wu 0009, Yingjie Cao, Tianyang Zhou, Jian Qian, Patrick Chiang 0001, Lei Qiu 0002, Yuanjin Zheng |
ISCAS | 15 |
| 2025 | A depth-aware geometric fusion based view synthesis method for sparse RGB-D input
Xiaobo Lu, Sifan Zhou, Patrick Chiang 0001 |
Expert Syst. Appl. | 4 |
| 2025 | An Adaptive Beam-Steering dToF LiDAR System Using Addressable Multi-Channel VCSEL Transmitter, 128 × 80 SPAD Sensor, and ML-Based Edge-Computing Object DetectionabstractIn this work, a solid-state direct time-of-flight (dToF) and adaptive beam-steering Light Detection and Ranging (LiDAR) system is proposed for machine learning (ML) based object detection. To leverage the capabilities of software and hardware, a co-optimization design from a neural network based algorithm to the architecture of transmitter, receiver and optical components is realized. Firstly, an object detection neural network is proposed for the depth-only input algorithm, which indicates the Region of Interest (ROI) in the illuminating field and gives hints of opened scan channels in the next two frames to decrease the total cost of the laser driver and sensor array. Next, the proposed network utilizes the Cross-Stage-Patrial (CSP) block to replace the residual structure in the backbone to achieve a lightweight performance and is implemented on the NVIDIA-Jetson to verify the system-level adaptive beam steering feature. To realize the smart working mode, a customized multi-channel and addressable TX is designed for adaptive and optical control to save power consumption and extend the ranging distance. At the same time, a 128×80 resolution RX which consists of Single-Photon Avalanche Diodes (SPADs) and column-wise Time-to-Digital Converter (TDC) is incorporated to capture the returned photons for combining sub-regions into an entire depth map. Next, to customize the specific scanning mechanism, for the optical setup, a cylindrical lens array is designed to reshape the laser beam, which matches the pattern of the transmitter to illuminate different targeted objects. Both the laser driver chip and the sensor chip with a 128×80 SPAD array are fabricated in the 180-nm Bipolar-CMOS-DMOS (BCD) process. Finally, the laser driver chip realizes the power of 5 W with an adjustable pulse width of 1.5 ns and the SPAD array integrates the depth accuracy of 5 cm at 15 m. Due to that the neural network realizes an accuracy up to 0.8, a low-power solid-state LiDAR prototype with adaptive beam steering is demonstrated. Yifan Wu 0009, Sifan Zhou, Lei Wang 0187, Jier Wang, Yuan Li 0074, Rui Bai 0001, Xuefeng Chen 0004, Yuanjin Zheng, Patrick Chiang 0001, Shenglong Zhuo, Lei Qiu 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 12 |
| 2024 | Sub-SA: Strengthen In-Context Learning via Submodular Selective AnnotationabstractIn-context learning (ICL) leverages in-context examples as prompts for the predictions of Large Language Models (LLMs). These prompts play a crucial role in achieving strong performance. However, the selection of suitable prompts from a large pool of labeled examples often entails significant annotation costs. To address this challenge, we propose Sub-SA (Submodular Selective Annotation), a submodule-based selective annotation method. The aim of Sub-SA is to reduce annotation costs while improving the quality of in-context examples and minimizing the time consumption of the selection process. In Sub-SA, we design a submodular function that facilitates effective subset selection for annotation and demonstrates the characteristics of monotonically and submodularity from the theoretical perspective. Specifically, we propose RPR (Reward and Penalty Regularization) to better balance the diversity and representativeness of the unlabeled dataset attributed to a reward term and a penalty term, respectively. Consequently, the selection for annotations can be effectively addressed with a simple yet effective greedy search algorithm based on the submodular function. Finally, we apply the similarity prompt retrieval to get the examples for ICL. Compared to existing selective annotation approaches, Sub-SA offers two main advantages. (1.) Sub-SA operates in an end-to-end, unsupervised manner, and significantly reduces the time consumption of the selection process (from hours-level to millisecond-level). (2.) Sub-SA enables a better balance between data diversity and representativeness and obtains state-of-the-art performance. Meanwhile, the theoretical support guarantees their reliability and scalability in practical scenarios. Extensive experiments conducted on diverse models and datasets demonstrate the superiority of Sub-SA over previous methods, achieving millisecond(ms)-level time selection and remarkable performance gains. The efficiency and effectiveness of Sub-SA make it highly suitable for real-world ICL scenarios. Our codes are available at https://github.com/JamesQian11/SubSA Jian Qian, Sifan Zhou, Ruizhi Hun, Patrick Chiang 0001 |
ECAI | 6 |
| 2023 | A 2GHz On-Chip-Oscilloscope with High Accuracy Pulse Width Detection for Auto-Peak-Power Controller & Peak-Current Detector in Voltage-Mode DToF DriverabstractThis paper presents a high-speed and high-resolution On-Chip Oscilloscope (OCO) for Auto-Peak-Power Controller (APPC) and Peak-Current Detector in voltage mode Direct Time-of-Flight(DToF) driver. The OCO supports both optical input and electrical input, with front-end circuit of TIA and PGA respectively. The front-end circuit bandwidth is up to 2GHz to support 500ps pulse width detection. The optical input is for detecting VCSEL peak power, and feedback to the integrated Boost Converter to adjust VCSEL supply voltage to find the target peak optical power. The electrical input is for detecting the VCSEL peak current to avoid driver over-current, by preventing APPC from adapting to excessive peak current due to abnormal VCSEL or PD. The front-end circuit is followed by an 11-bit asynchronous logic SAR-ADC with up to 40MHz sampling rate. The OSC time accuracy is 6.6ps and detected pulse width is up to 12.8ns. The chip was fabricated in BCD180nm process. According to the test results, the measured waveform of OCO is comparable with commercial oscilloscope. By turning on the OCO-based APPC feature, the variation of optical power is controlled within 3% in the temperature range of 25-85°C. Yuan Li 0074, Jiqing Xu, Yuxiang Tang 0001, Shenglong Zhuo, Xuefeng Chen 0004, Hengwei Yu, Huanli Jiang, Patrick Chiang 0001 |
ISCAS | 11 |
| 2023 | An Angle-Insensitive Time-Interleaved 4-Channels 138dB Dynamic Range Light Sensor with Flicker Detection for Smart Lighting ApplicationabstractThis paper presents an angle-insensitive light sensor by means of two pairs of Dual-Photodiodes (PDs) for X-axis and Y-axis direction respectively. A time-interleaved architecture is proposed to support 4-Channels (CHs) simultaneous flicker detection, sharing two Trans-Impedance-Amplifiers (TIAs) and one 16-bit third-order incremental Δ∑ADC. 7-Levels Auto-Gain-Control (AGC) is implemented to extend dynamic range and Dark-PD is introduced to track and cancel temperature-dependent offset. The chip was fabricated using 180nm CMOS technology. Test results show that the proposed light sensor achieves −30°∼30° flat angular response with +/−67° field-of-view (FOV). AGC extends input dynamic range to 138dB(0.012∼103klux), with linearity error of 0.169%. Up to 610Hz 4-Channels flicker can be detected in parallel at 0.053lux light input. Dark offset variation is controlled within +/−1.2LSB in the range of −30°C∼85°C. The chip draws 212uA from 1.8V supply and chip area is$2\text{mm}\times 0.65\text{mn}$. Yuan Li 0074, Aaron Wang, Chill Wang, Qianfan Ran, Shenglong Zhuo, Yifan Wu 0009, Hengwei Yu, Patrick Chiang 0001 |
ISCAS | 9 |
| 2023 | A 2KSPS 123dB Dynamic-Range SPAD-based Optical Sensor SoC with On-chip Auto-Gain-Control and FFT Processor for 100 μlux Light Illuminance and Flicker DetectionabstractSilicon photodetectors, such as photodiode (PD), are extensively used in optical sensing as their proper performance and low cost to achieve screen brightness adjustment, ambient light flicker detection, and other functions. With the rise of under-screen optical sensing, PD is no longer sufficient for low-light level light detection and real-time low light level flicker detection. While an emerging detector, single photon avalanche diode (SPAD), shows its advantage to ultralow light sensing. However, SPAD tends to saturate at high light levels, which limits its widespread application. To inspire ambient light sensor design and address these issues, this paper proposes a high sample per second (SPS), wide dynamic range (DR) and ultralow illuminance detection SPAD array system on chip (SoC). A$16\times 16$SPAD array architecture along with counter and charge-pump is designed in 180nm BCD process to sense 100μdux ambient light at 660nm. On-chip FFT processor of 2048 points with 32-bit accuracy is realized to process sampled optical waveform for flicker detection. Besides, an area-efficient integrated charge pump is implemented to regulate high-voltage (19~25V) to SPADs with on-chip digital auto-gain control (AGC), which improves the optical DR from 87db to 123db. This work extends the optical sensing range significantly and demonstrates the effectiveness of the proposed SoC with systematical measurements. Yifan Wu 0009, Jier Wang, Hengwei Yu, Xiangyu Fang, Mengxin Yu, Jiqing Xu, Lei Qiu 0002, Patrick Chiang 0001, Shenglong Zhuo |
ISCAS | 11 |
| 2023 | A Fully Integrated dToF System-on-Chip with High Precision Using Adaptive Optical Power Control and Shifted Histogram-Bin BinningabstractThis paper demonstrates a ranging sensor system with a configurable array of$16 \times 16$single photon avalanche diodes (SPADs), a 940nm vertical cavity surface-emitting laser (VCSEL), a co-design VCSEL driver with tunable widths from 400ps to 3630ps full-width at half-maximum (FWHM) optical pulses and peak power from 30mW to 170mW, and an embedded core to implement adaptive optical power control and distance extraction. An adaptive optical power control and a shifted bins binning of the histogram (SBbH) method to achieve high-precision distance measurement both at short-range and long-range. We achieved a minimum distance of 20mm with an absolute error better than 20% (4mm), still with a precision better than 10mm at 7m. All of results are obtained with 80% reflectivity of target at 100 frame rate. Hengwei Yu, Shenglong Zhuo, Yifan Wu 0009, Jiqing Xu, Jier Wang, Patrick Chiang 0001 |
ISCAS | 8 |
| 2022 | An Integrated 200MHz 4A Pulsed Laser Driver with DLL-Based Time Interpolator for Indirect Time-of-Flight ApplicationsabstractThis paper presents an integrated pulsed laser driver in 180nm BCD process for indirect time-of-flight applications. Pulses with 4A peak current at a modulation frequency up to 200MHz can be generated. Automatic optical power control (AOPC) is realized for laser eye safety protection. A DLL-based time interpolator is proposed to eliminate the mechanical moving components used during system calibration. The measured absolute accuracy of the time interpolator is +/−50ps with a resolution of 800ps and tuning range of 49. 6ns. This work effectively extends the measurement distance of the I-ToF system. The calibration efficiency is also improved with the proposed method. Shenglong Zhuo, Yifan Wu 0009, Lichun Xie, Yajie Qin, Rui Bai 0001, Patrick Chiang 0001 |
ISCAS | 12 |
| 2022 | A 56-Gb/s Reconfigurable Silicon-Photonics Transmitter Using High-Swing Distributed Driver and 2-Tap In-Segment Feed-Forward Equalizer in 65-nm CMOSabstractThis article presents a reconfigurable silicon- photonics transmitter (TX) for short-reach optical interconnects. The proposed hybrid-integrated TX combines a 65-nm CMOS driver with a 180-nm SOI-CMOS silicon-photonic Mach-Zehnder Modulator (MZM). The driver integrated with in- segment fractional-UI spaced feed-forward equalizer (FFE) is proposed to support the non-return-zero (NRZ) signaling, electrical- and optical-domain 4-level pulse-amplitude modulation (PAM-4) signaling. The driver employs a reconfigurable distributed topology to achieve high swing, wide bandwidth and flexible operation. The MZM is driven differentially in a push-pull configuration for high modulation efficiency. Measurement results show that the proposed TX operates up to 50-Gb/s NRZ data rate with 4-Vppd swing and 1.92-ps RMS jitter. In the optical PAM-4 mode, it reaches 56-Gb/s data rate and achieves >5-dB extinction ratio (ER) at the cost of 10.9-pJ/bit power efficiency. Yuguang Zhang, Qiwen Liao, Zhao Zhang 0004, Miaofeng Li, Jingbo Shi, Jian Liu 0021, Nanjian Wu, Yong Chen 0005, Patrick Chiang 0001, Ningmei Yu, Xi Xiao 0004, Nan Qi 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 12 |
| 2021 | A 4 × 10 Gb/s Adaptive Optical Receiver Utilizing Current-Reuse and Crosstalk-RemoveabstractThis article presents a$4 \times 10$Gb/s low-noise adaptive optical receiver, utilizing a current-reuse architecture and channel crosstalk-remove techniques in 65-nm CMOS process. The receiver integrates a transimpedance amplifier (TIA), a continuous-time linear equalizer (EQ), high-gain and high-bandwidth limiting amplifiers, and a 50-$\Omega $output driver into a single die. The TIA employs a common-source-based pseudo-differential topology with input series inductive peaking and$g_{m}$-enhancement to improve bandwidth and noise performance. An automatic-gain-control loop scheme is presented which solves the bandwidth variation issue caused by variation of the TIA input impedance, across a large dynamic range of a small input to a maximum overload current. Multiple crosstalk reduction techniques are adopted to improve the channel isolation performance. The 850-nm vertical cavity surface emitting laser (VCSEL)-based full-link measurement results show that the optical receiver achieves$20.4~\mu \text{A}_{\mathrm {pp}}$sensitivity bit-error-rate (BER$ < 1\text{e}$-12) with a 200-fF photodiode, competitive with other prior CMOS-based TIAs. The crosstalk over different channels is also measured, demonstrating only a sensitivity penalty of 1 dB. Xuefeng Chen 0004, Rui Bai 0001, Patrick Chiang 0001, Quan Pan 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | FotonNet: A Hardware-efficient Object Detection System using 3D-depth Segmentation and 2D-deep Neural Network Classifier
Sun Miao, Shi Shi, Patrick Chiang 0001 |
ICPRAM | 4 |
| 2020 | A 1.575GHz, 1.63mW CMOS Injection-Locked Ring Oscillator Powered by FBAR-Based PLL ReferenceabstractAn injection-locked ring oscillator (ILRO) used for multi-phase clock generation in multi-gigahertz serial links and time-interleaved ADCs, is co-designed with a MEMS-Film Bulk Acoustic Wave Resonator (FBAR)-based, 1.575GHz Integer-N PLL in CMOS. Low-jitter, multi-phase clock generation is obtained using a wide-tracking bandwidth, ring oscillator injection-locked with the FBAR-based PLL. Measured results show an output jitter of 1.2ps RMS at 1.575GHz and phase noise of -116dBc/Hz @1MHz offset. Excluding the power in the output buffers for testing, the ILRO and FBAR-based PLL dissipate 0.88mW and 0.75mW respectively, resulting in a total power of 1.63mW. Kangmin Hu, Julie R. Hu, Brian P. Otis, Patrick Chiang 0001 |
ISCAS | 4 |
| 2020 | A 112-Gb/s PAM-4 Linear Optical Receiver in 130-nm SiGe BiCMOSabstractIn this paper, we present a linear optical receiver for 112-Gb/s PAM-4 optical link. We propose a transimpedance front-end that optimizes thermal noise, power supply noise rejection, linearity and bandwidth altogether. The pseudo-differential structure is employed to achieve both low thermal noise and good power supply noise rejection. A transimpedance amplifier (TIA) gain control technique is proposed to improve linearity at both topology and transistor level while maintaining stability. An NIC-CTLE combo extends bandwidth with optimized frequency response. Designed in a 130nm SiGe BiCMOS process, the receiver realizes 37 GHz total bandwidth and input-referred noise of 19.8 pA/√Hz. The transimpedance gain can vary from 70 dB Ω to 50 dBΩ, which enables maximum input overload current of 1.8 mApp with <; 5% THD at differential output swing of 600 mVpp. The receiver consumes 77mA from 3.3V supply. Dan Li 0011, Shengwei Gao, Yongjun Shi, Xiaoyan Gui, Nan Qi 0002, Zhiyong Li 0014, Quan Pan 0002, Patrick Chiang 0001, Li Geng |
ISCAS | 8 |
| 2016 | A Robust Energy/Area-Efficient Forwarded-Clock Receiver With All-Digital Clock and Data Recovery in 28-nm CMOS for High-Density InterconnectsabstractThis paper presents a robust energy/area-efficient receiver fabricated in a 28-nm CMOS process. The receiver consists of eight data lanes plus one forwarded-clock lane supporting the hypertransport standard for high-density chip-to-chip links. The proposed all-digital clock and data recovery (ADCDR) circuit, which is well suited for today's CMOS process scaling, enables the receiver to achieve low power and area consumption. The ADCDR can enter into open loop after lock-in to save power and avoid clock dithering phenomenon. Moreover, to compensate the open loop, a phase tracking procedure is proposed to enable the ADCDR to track the phase drift due to the voltage and temperature variations. Furthermore, the all-digital delay-locked loop circuit integrated in the ADCDR can generate accurate multiphase clocks with the proposed calibrated locking algorithm in the presence of process variations. The precise multiphase clocks are essential for the half-rate sampling and Alexander-type phase detecting. Measurement results show that the receiver can operate at a data rate of 6.4 Gbits/s with a bit error rate-12, consuming 7.5-mW per lane (1.2 pJ/bit) under a 0.85 V power supply. With ADCDR's phase tracking, the receiver performs better in jitter tolerance and achieves a 500-kHz bandwidth, which is high enough to track the phase drift. The receiver core occupies an area of 0.02 mm2per lane. Hao Li 0047, Patrick Chiang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Rate-adaptive compressed-sensing and sparsity variance of biomedical signalsabstractBiomedical signals exhibit substantial variance in their sparsity, preventing conventional a-priori open-loop setting of the compressed sensing (CS) compression factor. In this work, we propose, analyze, and experimentally verify a rate-adaptive compressed-sensing system where the compression factor is modified automatically, based upon the sparsity of the input signal. Experimental results based on an embedded sensor platform exhibit a 16.2% improvement in power consumption for the proposed rate-adaptive CS versus traditional CS with a fixed compression factor. We also demonstrate the potential to improve this number to 24% through the use of an ultra low power processor in our embedded system. Vahid Behravan, Neil E. Glover, Rutger Farry, Patrick Chiang 0001, Mohammed Shoaib |
BSN | 4 |
| 2015 | Platform IO and system memory test using L3 cache based test (CBT) and parallel execution of CPGC Intel BIST engineabstractAs the memory industry pushes to increase memory density, device variation is creating more defects. Furthermore, new form factors (phone, tablet, mobile and client PC) and low cost board and platform limits physical access to JTAG or TAP. Taking advantage of the x86 architecture's high functional bandwidth to memory, the quickest way to access and test memory is through CPU core, by storing and running parallel CPGC/BIST test content via CPU L3 cache, one CPGC BIST engine per memory channel. We propose a cache based testing framework that speeds up test time 60× to 170× compared to JTAG or TAP based testing using the same test content. We will present the cache based test (CBT) architecture and infrastructure (MRC/NEM setup, CPGC/IBIST), test content, results, and a side by side comparison of test time to JTAG or TAP. Finally we will discuss and compare this approach to generalized cache based tests. Bruce Querbach, Tan Peter Yanyang, Lovelace Van, David Blankenbeckler, Rahul Khanna, Sudeep Puligundla, Patrick Chiang 0001 |
ITC | 7 |
| 2014 | A reusable BIST with software assisted repair technology for improved memory and IO debug, validation and test timeabstractAs silicon integration complexity increases with 3D stacking and Through-Silicon-Via (TSV), so does the occurrence of memory and IO defects and associated test and validation time. This ultimately leads to an overall cost increase. On a 14nm Intel SOC, a reusable BIST engine called Converged-Pattern-Generator-Checker (CPGC) are architected to detect memory and IO defects, and combined with the software assisted repair technology to automatically repair memory cell defects on 3D stacked Wide-IO DRAM. Additionally, we also present the CPGC gate count, power, simulation, and silicon results. The reusable CPGC IP is designed to connect to a standard IP interface, which enables a quick turn-key SOC development cycle. Silicon results show CPGC can speed up validation by 5x, improve test time from minutes down to seconds, and decrease debug time by 5x including root-cause of boot failures of the memory interface. CPGC is also used in memory training and initialization, which makes it a critical part of Intel SOC. Bruce Querbach, Rahul Khanna, David Blankenbeckler, Yulan Zhang, Ronald T. Anderson, David Gage Ellis, Zale T. Schoenborn, Sabyasachi Deyati, Patrick Chiang 0001 |
ITC | 9 |
| 2013 | SWIFT: A Low-Power Network-On-Chip Implementing the Token Flow Control Router Architecture With Swing-Reduced InterconnectsabstractA 64-bit, 8 × 8 mesh network-on-chip (NoC) is presented that uses both new architectural and circuit design techniques to improve on-chip network energy-efficiency, latency, and throughput. First, we propose token flow control, which enables bypassing of flit buffering in routers, thereby reducing buffer size and their power consumption. We also incorporate reduced-swing signaling in on-chip links and crossbars to minimize datapath interconnect energy. The 64-node NoC is experimentally validated with a 2 × 2 test chip in 90 nm, 1.2 V CMOS that incorporates traffic generators to emulate the traffic of the full network. Compared with a fully synthesized baseline 8 × 8 NoC architecture designed to meet the same peak throughput, the fabricated prototype reduces network latency by 20% under uniform random traffic, when both networks are run at their maximum operating frequencies. When operated at the same frequencies, the SWIFT NoC reduces network power by 38% and 25% at saturation and low loads, respectively. Jacob Postman, Tushar Krishna, Christopher Edmonds, Li-Shiuan Peh, Patrick Chiang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2012 | Regaining throughput using completion detection for error-resilient, near-threshold logicabstractOperating in the near-threshold regime can result in significant energy savings. Unfortunately, the increased timing variation prevents conventional error-detection techniques from properly functioning. This paper introduces two circuit-level timing error detection techniques that aim to increase throughput while operating in the near-threshold voltage regime: current-sensing completion detection and transition-aware completion detection. Each method allows any digital circuit to operate at speeds not limited by the worst-case critical path. Throughput improvements and energy savings are reported for implementations on a 16-bit adder. Joseph Crop, Robert Pawlowski, Patrick Chiang 0001 |
DAC | 3 |
| 2012 | Lane decoupling for improving the timing-error resiliency of wide-SIMD architecturesabstractA significant portion of the energy dissipated in modern integrated circuits is consumed by the overhead associated with timing guardbands that ensure reliable execution. Timing speculation, where the pipeline operates at an unsafe voltage with any rare errors detected and resolved by the architecture, has been demonstrated to significantly improve the energy-efficiency of scalar processor designs. Unfortunately, applying the same timing-speculative approach to wide-SIMD architectures, such as those used in highly-efficient GPUs, may not provide similar gains. In this work, we make two important contributions. The first is a set of models describing a parametrized general error probability function that is based on measurements of a fabricated chip and the expected efficiency benefits of timing speculation in a SIMD context. The second contribution is a decoupled SIMD pipeline that more effectively utilizes timing speculation and recovery, when compared with a standard SIMD design that uses only conventional timing speculations. The proposed lane decoupling enables each SIMD lane to tolerate timing errors independent of other adjacent lanes, resulting in higher throughput and improved scalability. We validate our modes and evaluate our design using a cycle-based GPU simulator, describe the conditions where efficiency improvements can be obtained, and explore the benefits of decoupling across a wide range of parameters. Our results show that timing speculation can achieve up to 10.3% improvement in efficiency. Evgeni Krimer, Patrick Chiang 0001, Mattan Erez |
ISCA | 2 |
| 2012 | A 12-bit 7 µW/channel 1 kHz/channel incremental ADC for biosensor interface circuitsabstractA two-channel micro-power incremental ADC, designed for biosensor interface circuits, is reported. It uses a noise-coupled multi-bit delta-sigma loop, integrated with a novel digital decimation filter operating in near-threshold. It was realized in the IBM 90 nm CMOS technology. The fabricated 90nm CMOS prototype device, for a 1 Vppdifferential input range, experimentally shows a 74dB SNDR up to 2 kHz (1 kHz/channel) signal bandwidth. The total measured power consumption of the modulator is 13.5 μW. Joseph Crop, Jeongseok Chae, Patrick Chiang 0001, Gabor C. Temes |
ISCAS | 4 |
| 2012 | A low-leakage dynamic register file with unclocked wordline and sub-segmentation for improved bitline scalabilityabstractA register file is presented that uses an unclocked wordline with dynamic bitline sub-segmentation and stack-forcing to reduce both active and standby leakage. The proposed technique improves scalability to a larger number of entries per local bitline segment. For an example 32nm-CMOS 4-write, 6-read port 32 bits x 168 entries register file, the proposed bitline segmentation technique improves local bitline delay by 33%, lowers standby leakage by 70%, and reduces the number of clocked wordlines by 83% when compared with a conventional design. Furthermore, wordline sharing is shown to reduce the number of wordlines by as much as 58%. Eric Donkoh, Patrick Chiang 0001 |
ISLPED | 2 |
| 2012 | Register file write data gating techniques and break-even analysis modelabstractRegister Files account for 30% of 32nm Intel WSM Core dynamic power of which 25% is due to write data distribution. We analyze Register File data gating strategies used to reduce write bitline dynamic power by as much as 96%. We explore the tradeoff of various data gating topologies (Global, Midway, Local), logic implementations (NAND, NOR, Tri-State), and techniques (Stack-Forcing, State-Forcing) to reduce both dynamic and leakage power. We then present a simple and accurate data gating break-even analysis model. The model comprehends "Data" and "Enable" switching activity, signal probability, logic implementation overhead, demonstrating an average error range of ±5%. Eric Donkoh, Teck Siong Ong, Yan Nee Too, Patrick Chiang 0001 |
ISLPED | 4 |
| 2012 | A Comparative Study of 20-Gb/s NRZ and Duobinary Signaling Using Statistical AnalysisabstractA statistical analysis technique for estimating bit-error rate (BER) and eye opening is presented for both non-return-to-zero (NRZ) and duobinary signaling schemes. This method enables fast and accurate BER distribution simulation of a serial link transceiver including channel and circuit imperfections, such as finite pulse rise/fall time, duty cycle variation and both receiver and transmitter forwarded-clock jitter. A comparison between 20-Gb/s NRZ and duobinary transmitters using this simulator shows that while duobinary transmission relaxes the requirements on the receiver equalizer due to the lower Nyquist frequency of the transmitted data, significant eye-opening and BER degradation can arise from clock non-idealities. The proposed statistical analysis is verified against traditional time-domain, transient eye-diagram simulations at 20-Gb/s, transmitted through measured s-parameter channel characteristics. Kangmin Hu, Larry Wu, Patrick Chiang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | An energy-efficient 64-QAM MIMO detector for emerging wireless standardsabstractA power/area aware design is mandatory for the MIMO (Multi-Input Multi-Output) detectors used in LTE and WiMAX standards. The 64-QAM modulation used in the MIMO detector requires more detection effort compared to the smaller constellation sizes widely implemented in the literature. In this work we propose a new architecture for the K-best detector, which unlike the popular multi-stage architecture used for K-best detectors, implements just one core. Also, we introduce a slight modification to the K-best algorithm that reduces the number of multiplications by 44%, and reduces the total power consumption by 27%, without any noticeable performance degradation. The overall architecture consumes only 24KGate, which is the smallest area compared to the other implementations in the literature. It also results in an at least 4-fold greater throughput-efficiency (Mbps/KiloGate) compared to the other detectors, while consuming a small power. The decoder implemented in a commercial 130nm process provides a data-rate of 107Mbps, and consumes 54.4mW. Nariman Moezzi Madani, Thorlindur Thorolfsson, Joseph Crop, Patrick Chiang 0001, William Rhett Davis |
DATE | 4 |
| 2011 | Network coding in multicore processorsabstractIn this paper, we present our initial results of improving communication performance in chip multicore processors by applying Network Coding (NC) at intermediate routers. Routers with NC capability enable two flit flows on the same communication links at the same time. Consequently, simulation results show that throughput and delay are improved with the proposed NC architecture. For saturated uni-cast communications in a 6-core network-on-a-chip, throughput can be improved up to 2.3 times with a negligible penalty of average latency at about 2.5 cycles. On the other hand, for saturated multi-cast communications in a 9-core network-on-a-chip, average latency can be reduced from over 1000 cycles to just about 37 cycles and throughput can be increased by 1.4 times. Thuan Duong-Ba, Thinh P. Nguyen, Patrick Chiang 0001 |
IPCCC | 3 |
| 2010 | SWIFT: A SWing-reduced interconnect for a Token-based Network-on-Chip in 90nm CMOSabstractWith the advent of chip multi-processors (CMPs), on-chip networks are critical for providing low-power communications that scale to high core counts. With this motivation, we present a 64-bit, 8×8 mesh Network-on-Chip in 90nm CMOS that: (a) bypasses flit buffering in routers using Token Flow Control, thereby reducing buffer power along the control path, and (b) uses low-voltage-swing crossbars and links to reduce interconnect energy in the data path. These approaches enable 38% power savings and 39% latency reduction, when compared with an equivalent baseline network. An experimental 2×2 core prototype, operating at 400 MHz, validates our design. Tushar Krishna, Jacob Postman, Christopher Edmonds, Li-Shiuan Peh, Patrick Chiang 0001 |
ICCD | 5 |
| 2009 | A 10Gb/s Wire-line Transceiver with Half Rate Period Calibration CDRabstractThis paper presents the design of a 10 Gb/s low power wire-line transceiver in 65 nm CMOS process with 1 V supply voltage. The transmitter occupies an area of 430 mum times 240 mum, consumes 50.56 mW power and has a 5-order programmable pre-emphasis equalizer. The receiver occupies an area of 300 mum times 500 mum. With the novel half rate period calibration clock data recovery (CDR) circuit, the receiver consumes only 52 mW power. The receiver combines a low power wideband programmable continuous time linear equalizer (CTLE) and a 3-order decision feedback equalizer (DFE). Zhuo Gao, Patrick Chiang 0001, Feng Zhang 0014 |
ISCAS | 3 |
| 2009 | Comparison of On-die Global Clock Distribution Methods for Parallel Serial LinksabstractThis paper presents a comparative study of clock distribution methods for serial links, including inverter chain, CML chain, transmission line, inductive load and capacitively driven wires in regards to delay, jitter and power consumption. Analysis, simulation and design insights are given for each method for 2.5 GHz clock propagation by on-die 5 mm wire in a 90 nm CMOS process. Simulations show the transmission line achieves least jitter and delay, while capacitively driven wire illustrates the best power-jitter and power-delay product. Kangmin Hu, Tao Jiang 0005, Patrick Chiang 0001 |
ISCAS | 3 |
| 2009 | Sense Amplifier Power and Delay Characterization for Operation under Low-Vdd and Low-voltage Clock SwingabstractTwo critical aspects of sense amplifiers (SA), power consumption and clock-to-data delay, are studied and presented for operation under low-supply voltage and driven by low-swing clock. Trade-offs and simulation results are given for a 4-stack StrongARM and a 3-stack double-tail SA, showing up to 50% power reduction in the SA itself and 25% in the clock generation circuit, with acceptable delay degradation. Tao Jiang 0005, Patrick Chiang 0001 |
ISCAS | 2 |
| 2009 | Measuring and Compensating for Process Mismatch-induced, Reference Spurs in Phase-locked Loops using a Sub-sampled DSPabstractA DSP method based on sub-sampling followed by M-point FFT of the sub-sampled signal is used to reduce the phase-locked loop's reference spur. To validate the system's effectiveness a digital calibration loop in SIMULINK is designed. The results show that the reference spur can be improved by 22 dBc with a 1% residual current mismatch and a 1 nA net value of leakage current. Zhuo Gao, D. Kesharwani, Patrick Chiang 0001, Weiwu Hu |
ISCAS | 3 |
| 2007 | Process Variation Compensation of a 2.4GHz LNA in 0.18um CMOS Using Digitally Switchable CapacitanceabstractA 0.18mum CMOS 2.4GHz LNA (low noise amplifier) with digitally switchable capacitance has been designed to investigate its ability to compensate for performance variation across worst case process conditions. The effects of transistor model and passive component variation are first simulated to quantify the range of performance uncertainty. The use of various methods of switchable capacitance is then investigated at multiple LNA nodes to observe the effect on S11, S21, and NF. After calibration, the new design allows for 300MHz of frequency tuning and improvements in S11by 7.2dB, S21by 2.4dB, and NF by 0.6dB. Yike Cui, Baoyong Chi, Minjie Liu, Yongming Li 0004, Zhihua Wang 0001, Patrick Chiang 0001 |
ISCAS | 7 |