EDBT 2026 Demo / reviewers in the wild / expert
Chuanjin Richard Shi
dblp:205/4088 · also C.-J. Richard Shi, Richard Shi
· DBLP profile ↗
119ranked-venue papers
15as first author
26since 2021 · last 2026
0000-0002-3157-3464ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 112 · 14 first-author · 23 since 2021Software engineering, systems software and programming languages · 8Artificial intelligence and machine learning · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI EnginesabstractCommercial FPGAs, such as AMD Versal devices, increasingly incorporate AI engines that exploit low-precision packed-SIMD fused multiply–accumulate (FMA) to achieve proportional throughput gains. However, trans-precision FMA (e.g., multiplying two FP16 numbers and adding their result to an FP32 accumulator), which preserves numerical stability by accumulating in higher precision, remains bottlenecked by the highest-precision, lowest-throughput operation. Dot-product accumulation (DPA) (e.g., performing a dot-product on two 4-element FP8 vectors and adding its result to an FP32 accumulator) can fully utilize the input/output bandwidth and computational resources. Existing flexible open-source FPUs, such as FPnew, do not support DPA and implement SIMD FMA on low-precision formats by replicating independent FMA lanes, which increases area, underutilizes shared arithmetic resources, and complicates the integration of DPA operations.This paper presents TransDot, a reconfigurable FPU that unifies multi-precision SIMD FMA and trans-precision DPA within a shared, reconfigurable datapath. TransDot extends the baseline design with 2-term FP16, 4-term FP8, and 8-term FP4 dot-product accumulation into FP32 using reconfigurable subcomponents. Evaluation shows that TransDot delivers 2× FP16, 4× FP8, and 8× FP4 throughput via DPA with FP32 accumulation, and 1.46× area efficiency in FP16 DPA and 2.92× area efficiency in FP8 DPA, at the cost of 37.3% larger area on average and an additional pipeline stage in dot-product mode compared to the FPnew baseline. These results demonstrate that TransDot’s area-efficient design enables scalable deployment in next-generation AMD Versal AI engines. Maohua Nie, Sin-Chen Lin, Chuanjin Richard Shi, Ang Li 0011 |
FCCM | 4 |
| 2026 | An Injection-Locked Eight Phase Clock Generator with Edge Replacement and Injection Error Calibration Achieving -253.7dB FOMJitter-N
Sirou Li, Weijia Zeng, Kaiyun Cao, Liangjian Lyu, Chuanjin Richard Shi, Hao Min |
ISCAS | 5 |
| 2026 | An All-MOS 1-nA Current Reference Insensitive to Process and Voltage Variations
Kexin Shan, Jiaqing Rui, Yuting Wan, Wenxian Gu, Xing Wu 0005, Chuanjin Richard Shi, Liangjian Lyu |
ISCAS | 6 |
| 2026 | An Optimized Pre-Emphasis Spike Detector for High-Density Neural Interfaces
Jingjie Tang, Yulun Peng, Hengchang Bi, Xing Wu 0005, Chuanjin Richard Shi, Liangjian Lyu |
ISCAS | 5 |
| 2026 | A 28nm 56.45TOPS/W Hardware-software Co-optimized Sparse Accelerator for Traditional DNNs and Transformer-based LLMs
Chuanjin Richard Shi |
ISCAS | 5 |
| 2026 | A 39.4-μW 915-MHz Third-Harmonic Mixing Receiver With On-Chip LO Achieving -86-dBm Sensitivity and Multichannel SelectionabstractThis paper presents a 915 MHz ultra-low-power (ULP) receiver based on a single-path third-harmonic mixing (SPTHM) architecture. Unlike conventional multi-path sub-harmonic receivers that require precise multi-phase local oscillators (LOs), this work simplifies the receiver to a single-mixer path architecture by jointly optimizing the LO harmonic order and duty cycle. The receiver is driven by a 10%-duty-cycle LO operating at one-third of the carrier frequency ($f_{\mathrm {c}}$). Compared to the typical 50%-duty-cycle LO in the SPTHM configuration, the proposed receiver improves the conversion gain by 12.1 dB and noise figure (NF) by 5.2 dB. The architecture also exhibits a front-end NF variation of less than 1 dB across the 3%-13% LO duty-cycle range, thereby relaxing constraints on pulse generation. To facilitate ULP channel selection, a comparison-skipped frequency-locked loop (FLL) is used, consuming just$5~\mu $W. A high-Q IF amplifier with an improved active inductor load is incorporated to enhance in-band interference rejection, achieving 31 dB signal-to-interference ratio (SIR) at 5 MHz offset. Fabricated in a 65 nm CMOS, the receiver achieves a sensitivity of −86 dBm at 250 kbps data rate with a$39.4~\mu $W power consumption, including an on-chip LO. It indicates a competitive figure-of-merit (FoM) of 184 dB within a compact active area of 0.16 mm2. Heyu Ren, Wenjun Gong, Sirou Li, Xing Wu 0005, Liangjian Lyu, Chuanjin Richard Shi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | ReaLLM: A Trace-Driven Framework for Rapid Simulation of Large-Scale LLM InferenceabstractAs Large Language Models (LLMs) continue to scale, optimizing their deployment requires efficient hardware and system co-design. However, current LLM performance evaluation frameworks fail to capture both chip-level execution details and system-wide behavior, making it difficult to assess realistic performance bottlenecks. In this work, we introduce ReaLLM, a trace-driven simulation framework designed to bridge the gap between detailed accelerator design and large-scale inference evaluation. Unlike prior simulators, ReaLLM integrates kernel profiling derived from detailed microarchitectural simulations with a new trace-driven end-to-end system simulator, enabling precise evaluation of parallelism strategies, batching techniques, and scheduling policies. To address the high computational cost of exhaustive simulations, ReaLLM constructs a precomputed kernel library based on hypothesized scenarios, interpolating results to efficiently explore a vast design space of LLM inference systems. Our validation against real hardware demonstrates the framework's accuracy, achieving an average end-to-end latency prediction error of only 9.1% when simulating inference tasks running on 4 NVIDIA H100 GPUs. We further use ReaLLM to evaluate popular LLMs' end-to-end performance across traces from different applications and identify key system bottlenecks, showing that modern GPU-based LLM inference is increasingly compute-bound rather than memory-bandwidth bound at large scale. Additionally, we significantly reduce simulation time with our precomputed kernel library by a factor of$6 \times$for full-simulations and$164 \times$for workload SLO exploration. ReaLLM is open-source and available at https://github.com/bespoke-silicon-group/reallm. Huwan Peng, Scott Davidson 0004, Chuanjin Richard Shi, Michael B. Taylor |
ASAP | 3 |
| 2025 | HCSAs: Hybrid Computing Systolic Arrays for Accelerating Mamba Models with Unified State Space Buffers and Energy-Efficient Dataflow
Maohua Nie, Chuanjin Richard Shi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | LEGO: A 12nm FinFET Analog Cell Library for Analog/Mixed-Signal Applications
Xindi Liu, Chien-Jian Tseng, Chuanjin Richard Shi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | A 40µW 915MHz Receiver with Sub-Passive Third-Harmonic Mixer Achieving -88dBm Sensitivity and Multi-Channel SelectionabstractThis paper presents a 915MHz multi-channel receiver for ultra-low-power (ULP) applications. A sub-passive third-harmonic mixer is proposed to improve the front-end conversion performance at the ULP budget. Driven by the local oscillator (LO) that operates at one-third of the RF carrier frequency with a duty cycle of about 10%, the proposed mixer improves the front-end’s noise figure (NF) to 15dB while reducing the LO power consumption by three times. The mixer also alleviates the requirement for LO phase accuracy. With the LO duty cycle ranging from 4% to 14%, the variation of the mixer performance is less than 10%, which allows a simple pulse generation approach without a precise duty cycle control circuit, thereby saving power consumption. A low-power frequency-locked loop (FLL) with a 196kHz tuning step facilitates channel selection. Additionally, a high-Q intermediate-frequency (IF) amplifier with an active inductor is used to suppress in-band interference and noise bandwidth. Implemented in a 65nm CMOS process and based on post-simulation results, the receiver achieves a sensitivity of −88dBm while consuming 40µW at a 250kb/s data rate. The receiver also performs an improved figure-of-merit (FoM) of 186dB with 26dB interference tolerance and a compact active area of 0.16mm2. Heyu Ren, Wenjun Gong, Sirou Li, Liangjian Lyu, Chuanjin Richard Shi |
ISCAS | 6 |
| 2025 | A 915MHz 97nW Low-Area Wake-Up Receiver with an Envelope-Tracking Mixer Achieving -73.2dBm SensitivityabstractThis paper presents a 915MHz bit-level duty-cycled (BLDC) wake-up receiver (WuRX) designed for ultra-low-power Internet-of-Things (IoT) systems. Utilizing an envelope-tracking (ET) mixer, the WuRX generates a phase-following (PF) local oscillator (LO) from the incoming radio-frequency (RF) signals. The recovered PF LO has a low duty cycle of about 10% and aligns the peak and valley voltages of the RF signal. As a result, the on-off keying data is accurately downconverted to the baseband with improved gain. Moreover, the high-speed sampling capability of the ET mixer allows the receiver’s front-end to operate at a low duty cycle of 0.024%, significantly lower than other mixer-based BLDC WuRXs. Dynamic baseband amplifiers with pre-charging and auto-zeroing abilities are employed to save area and reduce power consumption. Fabricated in a 65nm CMOS process, the proposed BLDC WuRX achieves a sensitivity of −73.2dBm at a 1kbps data rate, consuming only 97nW and occupying an ultra-low-area of 0.036mm2. Heyu Ren, Wenjun Gong, Liangjian Lyu, Chuanjin Richard Shi |
ISCAS | 5 |
| 2025 | A 240-mV 33.8-μV/°C 1.5-nW Voltage Detector for Energy HarvestingabstractThis paper presents an ultra-low-voltage (ULV) and ultra-low-power (ULP) voltage detector (VD) for energy harvesting systems (EHS). The VD is designed using a 2-transistor (2T) voltage detection circuit and a 2T bias circuit. By leveraging two types of transistors with small threshold voltage differences, the detection voltage is reduced to 240 mV, enhancing the feasibility and practicality of the design for ULV EHS. The detection voltage can be programmed by incorporating an appropriate voltage divider into the VD, while the proposed 2T bias technique reduces its variation across different process corners. The design is implemented in a 65 nm CMOS process and occupies a chip area of 260 µm2. Post-layout simulation results indicate that the VD achieves a detection voltage of 240 mV, with a standard deviation (σ) of 6.6 mV, drawing only 5 nA of current at a supply voltage of 0.3 V. The average temperature coefficient (TC) across five process corners is 33.8 µV/°C, simulated within a temperature range of -40 °C to 125 °C. Wenjun Gong, Liangjian Lyu, Chuanjin Richard Shi |
ISCAS | 4 |
| 2025 | A 0.473 μJ/class Seizure Detection Processor with LSVM Classifier and LPF-Based Feature ExtractionabstractThe closed-loop deep brain stimulation system demands high-performance seizure detection, especially in terms of ultra-low power consumption and patient specificity. In this paper, we propose a seizure detection approach featuring low-pass filters for feature extraction and a programmable linear support vector machine for classification. This approach effectively reduces power consumption while retaining the signal energy near the cutoff frequencies and preserving the correlation between adjacent frequency bands. To reduce the false alarm rate, a Hidden Markov Model is utilized for post-processing. The proposed processor also employed calculation bit-width optimization and time-division multiplexing to minimize power and area consumption, while maintaining minimal accuracy loss. Implemented in a 65-nm CMOS process, the processor occupies an active area of 0.14 mm2. It achieves an energy classification efficiency of 0.473 μJ/class with 0.7-V supply and 16.384-kHz system clock. The measurement results show a sensitivity of 95.92%, a specificity of 98.11%, and a false alarm rate of 1.78 times/h, as validated by the CHB-MIT dataset. Wenxian Gu, Xudong Hao, Hengchang Bi, Xing Wu 0005, Chuanjin Richard Shi, Liangjian Lyu |
ISCAS | 5 |
| 2025 | A Reference Double-Sampling PLL-Based Eight Phase Clock Generator Achieving 0.18mW/GHz/phase and -251.9dB FOMJitter-NabstractA reference double-sampling phase-locked-loop-based (RDSPLL-based) multi-phase clock generator (MPCG) for DDR PHY is presented. The reference double-sampling architecture is utilized to achieve low phase noise. A CDAC-embedded voltage offset calibration is proposed to reduce jitter and reference spur, and a CMP-ADC hybrid phase detector is adopted to accelerate the locking process. Fabricated in 65nm, the proposed MPCG achieves better than 1° phase accuracy with a 100MHz reference clock. The reference spur is reduced from -56dBc to -80dBc and the locking time is reduced from 10.5us to 1.6us. The measured RMS jitter is 674fs at 2.4GHz with 3.43mW, yielding the FOMJitter-Nof -251.9dB. Sirou Li, Weijia Zeng, Kaiyun Cao, Liangjian Lyu, Chuanjin Richard Shi |
ISCAS | 5 |
| 2025 | An Integer-N Reference-Double-Sampling PLL for Frequency-Multiplied Octa-Phase Clock Generation Achieving -251.9 dB FOMJitter-NabstractThis paper presents a reference double-sampling phase-locked loop (RDSPLL) that integrates frequency multiplication and octa-phase clock generation into a single system, significantly reducing power consumption. A differential ring oscillator (DRO) is employed to generate octa-phase clocks with high phase accuracy. The reference double-sampling technique extends the loop bandwidth, effectively suppressing phase noise from the ring oscillator and thereby reducing jitter. To achieve accurate and efficient phase error detection, we proposed a novel offset-compensated hybrid phase detector (OCH-PD), featuring an offset calibration and a comparator-ADC hybrid quantizer. The offset calibration utilizes the CDAC to dynamically compensate for the mismatch in double-sampling, improving jitter and spur performance. The hybrid quantizer supports dynamic mode switching based on different locking states: during the coarse frequency locking phase, it operates in the ADC mode to accelerate the locking process; once a stable lock is achieved, it switches to the comparator mode to enable low-power, high-speed quantization. Fabricated in a 65-nm CMOS process, the prototype achieves 674 fs RMS jitter at 2.4 GHz while consuming only 3.43 mW, resulting in a$\text {FOM}_{\text {Jitter-N}}$of -251.9 dB. With offset calibration, the reference spur at 100 MHz is suppressed from -56 dBc to -80 dBc, and the jitter is reduced from 1.42 ps to 674 fs. The locking time improves from$10.5~{\mu }$s to$1.6~{\mu }$s using the hybrid quantizer. The eight-phase accuracy remains better than 1° over the frequency range of 2-2.8 GHz. Sirou Li, Rongjin Xu, Weijia Zeng, Kaiyun Cao, Heyu Ren, Xing Wu 0005, Liangjian Lyu, Chuanjin Richard Shi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | A Workload-Balance-Aware Accelerator Enabling Dense-to-Arbitrary-Sparse Neural NetworksabstractDeep neural networks (DNNs) have proved their great potential over various perceptual and cognitive tasks with the cost of ever-growing storage capacity and computation complexity. Sparse representations in neural networks have emerged as a compelling method to achieve substantial reductions in computational overhead and energy consumption. However, the introduction of sparsity presents challenges such as irregular memory accesses and wasted computation cycles. Traditional methods have attempted to address these challenges with varied success, unfortunately, often involving complex hardware designs or sacrificing sparsity’s benefits. In this article, we propose a software-hardware codesign solution, comprising an offline workload balancing encoding algorithm toward arbitrary sparsity in DNNs and a dedicated Processing Element array–based accelerator with a lightweight switch network. Extensive experiments are conducted to demonstrate the conclusion that our proposal is feasible to support various network structures and a wide range of sparsity ratios. With the encoding algorithm, a 1.16×–2.61× acceleration is achieved compared to the baseline. The system-on-chip measures 7.9 mm 2 , achieving an energy efficiency of 0.7 TOPS/W (dense), 2.1 TOPS/W (at 75% sparsity), and 10.3 TOPS/W (at 99.9% sparsity) at 0.9 V and 1,066-MHz clock frequency. Maohua Nie, Longfei Gou, Junmin He, Yongchuan Dong, Qiaosha Zou, Chuanjin Richard Shi |
ACM Trans. Embed. Comput. Syst. | 11 |
| 2024 | SSGCNet: A Sparse Spectra Graph Convolutional Network for Epileptic EEG Signal ClassificationabstractIn this article, we propose a sparse spectra graph convolutional network (SSGCNet) for epileptic electroencephalogram (EEG) signal classification. The goal is to develop a lightweighted deep learning model while retaining a high level of classification accuracy. To do so, we propose a weighted neighborhood field graph (WNFG) to represent EEG signals. The WNFG reduces redundant edges between graph nodes and has lower graph generation time and memory usage than the baseline solution. The sequential graph convolutional network is further developed from a WNFG by combining sparse weight pruning and the alternating direction method of multipliers (ADMM). Compared with the state-of-the-art method, our method has the same classification accuracy on the Bonn public dataset and the spikes and slow waves (SSW) clinical real dataset when the connection rate is ten times smaller. Chuanjin Richard Shi |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | AutoMap: Automatic Mapping of Neural Networks to Deep Learning Accelerators for Edge DevicesabstractEmerging deep neural networks (DNNs) have been emerging in applications (object detection, automatic speech recognition, etc.) deployed on edge devices. To improve the energy efficiency of edge devices, domain-specific deep learning accelerators (DLAs) are designed with limited on-chip resources. The manifold DLA designs and evolving DNN topologies bring challenges for applications mapping and scheduling on hardware resources. In this article, we propose an automatic DNN mapping framework named AutoMap, given the hardware backend information. First, a computational graph representation called extended directed weighted graph (EDWG) is proposed, which realizes unified expression for both spatial and temporal network interlayer connections. Second, an associated partitioner is implemented for splitting an EDWG into subEDWGs, which incorporates the on-chip memory constraint and facilitates weight data reuse on chip. Finally, a dynamic memory allocation strategy is utilized to alleviate the feature storing burden introduced by the multivarious network sizes and connections. Compared to the baseline mapping methods, experimental results show that our proposed automatic mapping framework can help to speedup the execution of several DNNs on state-of-the-art DLAs, ranging from$1.27\times $to$3.45\times $. The utilization of the PE array can increase from 20% to 64%. Maohua Nie, Qiaosha Zou, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | Augmenting aspect-level sentiment classification with distance-related local context input
Yongchuan Dong, Qiaosha Zou, Chuanjin Richard Shi |
J. Supercomput. | 3 |
| 2023 | Accelerating Distributed GNN Training by CodesabstractEmerging graph neural network (GNN) has recently attracted much attention and has been used extensively in many real-world applications thanks to its powerful expression ability of unstructured data. The real-world graph datasets are very large-scale, which can contain up to billions of nodes and tens of billions of edges. It usually requires distributed system to train GNN on such huge datasets. As a result, the data communication overheads between machines become the bottleneck of GNN computation. Our profiling results show that getting attributes from remote machines during sampling phase in GNN occupies$> $75% of the time of the training process. To address this issue, in this article, we propose Coded Neighbor Sampling (CNS) framework, which introduces codes technique to reduce the communication overheads of GNN. In the proposed CNS framework, the codes technique is coupled with GNN sampling method to exploit the data excess among different machines caused by unstructured nature of graph data. An analytical performance model is built for the proposed CNS framework, whose results are corroborated by the simulation and validate the benefit of the proposed CNS framework over both conventional GNN training method and conventional codes technique. Performance metrics, such as communication overheads, runtime, and throughput, of the proposed CNS framework are evaluated on a distributed GNN training simulation system implemented on MPI4py platform. The results show that, on average, the proposed CNS framework can save communication overhead by 40.6%, 35.5%, and 16.5%, reduce the runtime by 12.1%, 17.0%, and 10.0%, and improve the throughput by 16.2%, 24.4%, and 11.2%, respectively, when training GNN models with Cora, PubMed, and Large Taobao. Tianchan Guan, Dimin Niu, Qiaosha Zou, Hongzhong Zheng, Chuanjin Richard Shi, Yuan Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | A 2.0-2.9 GHz ring-based injection-locked clock multiplier using a self-alignment frequency-tracking loop for reference spur reduction
Rongjin Xu, Dawei Ye, Chuanjin Richard Shi |
Integr. | 3 |
| 2022 | Analysis and Design of Digital Injection-Locked Clock Multipliers Using Bang-Bang Phase DetectorsabstractThe analysis of injection-locked clock multipliers (ILCMs) using bang-bang phase detectors (BBPDs) is challenging due to the nonlinear BBPD and the multi-rate injected oscillator. This paper presents an explicit analysis of digital ILCMs using BBPDs and proposed an intuitive approach to optimizing parameters with given noise sources. A time-domain analysis in the single-clock domain is presented to solve the closed-form expression of jitter in the ILCM. The proposed approach exhibits good consistency with simulations, for various design parameters and noise cases. With the predicted input-referred jitter, the equivalent BBPD gain is resolved to derive the frequency-domain noise transfer functions in concise forms. To achieve the desired performance with given specifications, recommended design procedures are summarized based on the proposed analysis and verified by simulations. Rongjin Xu, Dawei Ye, Chuanjin Richard Shi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | SAIL: A Deep-Learning-Based System for Automatic Gait Assessment From TUG VideosabstractGait disorders are common in the elderly people, seriously hinder patients’ mobility and sometimes indicate underlying severe neurological diseases. Timely and automatic diagnosis of gait disorders is greatly desired. Existing methods with wearable devices put burdens on patients. We establish a video-based algorithm namedSAILto perform contactless gait assessment automatically. The SAIL contains three parts, namely,skeleton detector,parameter extractor, andgait classifier. Using a pose estimation algorithm, the skeleton detector converts RGB videos to a human skeleton sequence. Then, the parameter extractor extracts gait parameters from skeletons with a signal detection technique. Finally, a trained Support vector machine is used as a gait classifier to detect abnormal gait. The SAIL achieves 86.2% sensitivity and 98.5% specificity for abnormal gait detection on ourSAIL-TUGdataset, outperforming general clinic doctors with 76.4% and 97.4%, respectively. Nine gait parameters and the binary gait classification result are included in the final gait report. We implement an automatic gait assessment system based on SAIL and deployed the user-interface software in more than 60 hospitals for practical applications. More than 30 000 gait reports have been automatically generated. Moreover, we establish a publicly available dataset namedSAIL-TUGincluding 404 annotated Timed “Up & Go” videos. Qiaosha Zou, Yanmin Tang, Chuanjin Richard Shi |
IEEE Trans. Hum. Mach. Syst. | 7 |
| 2021 | Systolic-Array Deep-Learning Acceleration Exploring Pattern-Indexed Coordinate-Assisted Sparsity for Real-Time On-Device Speech ProcessingabstractThis paper presents a hardware-software co-design for efficient sparse deep neural networks (DNNs) implementation in a regular systolic array for real-time on-device speech processing. The weight pruning format, exploring pattern-based coordinate-assisted (PICA) sparsity, expands the pattern-based pruning into both convolutional neural networks (CNNs) and recurrent neural networks (RNNs). It reduces the index storage overhead as well as avoids accuracy degradation. The proposed systolic accelerator leverages the intrinsic data reuse and locality to accommodate the PICA-based sparsity without using complex data distribution networks. It also supports DNNs with different topologies. By reducing the model size by 16x, PICA sparsification reduces 6.02x index storage overhead while still achieving 20.7% WER in TIMIT dataset. For the pruned WaveNet and LSTM, the accelerator achieves 0.62 and 2.69 TOPS/W energy efficiency, 1.7x to 10x higher than the state-of-the-art. Shiwei Liu 0002, Qiaosha Zou, Chuanjin Richard Shi |
ACM Great Lakes Symposium on VLSI | 6 |
| 2021 | A Fully-Synthesizable Fast-Response Digital LDO Using Automatic Offset Control and ReuseabstractThis paper proposes a fully synthesizable digital low-dropout (DLDO) regulator using an automatic offset control and reuse technique implemented in a 65-nm CMOS technology. To realize the fully synthesizable DLDO design, all components of core blocks are made with standard logic cells. The proposed offset control and reuse technique is adopted to cancel the offset voltage from the logic cells automatically, to provide the adaptive equivalent thresholds for the voltage comparison window, and to speed up the dropout voltage response. Besides, the modified comparator-triggered oscillator introduces the output of the comparator to the oscillation loop to choose the optimum clock edge in the delay line to feed back, which ensures both of the fast response and enough time margin for the comparator delay at the same time. The core area of the output-capacitance-less DLDO is 0.027 mm2. The simulated results show that a 100 mA step in load current produces a voltage drop of 140 mV with the response time of 1.2 ns. The steady-state error is less than 4 mV. The peak current efficiency is 99.9%. Chuanjin Richard Shi |
ISCAS | 3 |
| 2021 | Saliency-based YOLO for single target detection
Junying Hu, Chuanjin Richard Shi, Jiangshe Zhang 0001 |
Knowl. Inf. Syst. | 2 |
| 2020 | ReBoc: Accelerating Block-Circulant Neural Networks in ReRAMabstractDeep neural networks (DNNs) emerge as a key component in various applications. However, the ever-growing DNN size hinders efficient processing on hardware. To tackle this problem, on the algorithmic side, compressed DNN models are explored, of which block-circulant DNN models are memory efficient and hardware-friendly; on the hardware side, resistive random-access memory (ReRAM) based accelerators are promising for in-situ processing of DNNs. In this work, we design an accelerator named ReBoc for accelerating block-circulant DNNs in ReRAM to reap the benefits of light-weight models and efficient in-situ processing simultaneously. We propose a novel mapping scheme which utilizes Horizontal Weight Slicing and Intra-Crossbar Weight Duplication to map block-circulant DNN models onto ReRAM crossbars with significant improved crossbar utilization. Moreover, two specific techniques, namely Input Slice Reusing and Input Tile Sharing are introduced to take advantage of the circulant calculation feature in block- circulant DNNs to reduce data access and buffer size. In REBOC, a DNN model is executed within an intra-layer processing pipeline and achieves respectively 96× and 8.86× power efficiency improvement compared to the state-of-the-art FPGA and ASIC accelerators for block-circulant neural networks. Compared to ReRAM-based DNN accelerators, REBOC achieves averagely 4.1× speedup and 2.6× energy reduction. Yitu Wang, Fan Chen 0001, Linghao Song, Chuanjin Richard Shi, Hai Li 0001, Yiran Chen 0001 |
DATE | 4 |
| 2020 | A 400 MHz, 8-Bit, 1.75-ps Resolution Pipelined-Two-Step Time-to-Digital Converter with Dynamic Time AmplificationabstractThis work proposes a high-speed pipelined-two-step time-to-digital converter (TDC) with a dynamic time amplification (DTA) to improve the resolution at low power. The key element of this TDC is the DTA. It samples the residual time errors as voltages held in the MOM capacitors and discharges them to generate the amplified time difference. Thanks to the dynamic time-voltage-time conversion, the DTA realizes high linearity and power efficiency, and can be employed to build a pipeline TDC architecture with high sampling frequency because of its sample and hold operation. Moreover, the DTA maintains constant gain, so only a one-time forground calibration for gain mismatch is required in this TDC. Simulations show that the TDC designed in 65 nm CMOS achieves 8-bit, 1.75 ps of time resolution, and 1 LSB INL and 1.6 LSB DNL with one-time foreground calibration at 400 MHz sampling frequency while just consuming 726 μW power, which corresponds to 18.45 fJ/Conv. FoM. Yuting Tu, Rongjin Xu, Dawei Ye, Liangjian Lyu, Chuanjin Richard Shi |
ISCAS | 5 |
| 2020 | A 0.6V 1.07 μW/Channel neural interface IC using level-shifted feedback
Liangjian Lyu, Yu Wang 0046, Chixiao Chen, Chuanjin Richard Shi |
Integr. | 4 |
| 2020 | Analysis of Passive Charge Sharing-Based Segmented SAR ADCsabstractThis article presents the theoretical analysis of passive charge sharing-based segmented successive-approximation-register (SAR) analog-to-digital converter (ADC), where the precise reference source in a capacitive digital-to-analog converter (CDAC) is replaced by a capacitor that is$\beta $times larger than its bit capacitor and precharged to the reference level, known as a reference charge reservoir (RCR). A segmented SAR-ADC uses a coarse SAR-ADC to compute some most significant bits (MSBs). Four methods, namely aligned switching (AS) with bitwise RCRs, AS with a subsample-wise RCR, detect-and-skip aligned switching (DAS-AS) with bitwise RCRs, and DAS-AS with a subsample-wise RCR are introduced for setting fine MSBs. Closed-form analytic expressions of the reference error due to the finite reference capacitance are derived and validated by behavioral modeling and circuit simulation of an 11-bit 50 MS/s segmented SAR ADC in 65-nm CMOS technology. The error expressions can be used to select one of the four methods for setting the fine MSBs and to determine$\beta $for the required linearity or for implementing digital circuitry for precise error correction. Aili Wang 0002, Chuanjin Richard Shi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | A low-voltage low-power multi-channel neural interface IC using level-shifted feedback technologyabstractA low-voltage low-power 16-channel neural interface front-end IC for in-vivo neural recording applications is presented in this paper. A current reuse telescope amplifier is used to achieve better noise efficiency factor (NEF). Power efficiency factor (PEF) is further improved by reducing supply voltage with the proposed level-shifted feedback (LSFB) technique. The neural interface is fabricated in a 65 nm CMOS process. It operates under 0.6V supply voltage consuming 1.07 μW/channel. An input referred noise of 5.18 μV is measured, leading to a NEF of 2.94 and a PEF of 5.19 over 10 kHz bandwidth. Liangjian Lyu, Yu Wang 0046, Chixiao Chen, Chuanjin Richard Shi |
ASP-DAC | 4 |
| 2019 | Hierarchical Representation Learning for Bipartite GraphsabstractRecommender systems on E-Commerce platforms track users' online behaviors and recommend relevant items according to each user’s interests and needs. Bipartite graphs that capture both user/item feature and use-item interactions have been demonstrated to be highly effective for this purpose. Recently, graph neural network (GNN) has been successfully applied in representation of bipartite graphs in industrial recommender systems. Providing individualized recommendation on a dynamic platform with billions of users is extremely challenging. A key observation is that the users of an online E-Commerce platform can be naturally clustered into a set of communities. We propose to cluster the users into a set of communities and make recommendations based on the information of the users in the community collectively. More specifically, embeddings are assigned to the communities and the user embedding is decomposed into two parts, each of which captures the community-level generalizations and individualized preferences respectively. The community embedding can be considered as an enhancement to the GNN methods that are inherently flat and do not learn hierarchical representations of graphs. The performance of the proposed algorithm is demonstrated on a public dataset and a world-leading E-Commerce company dataset. Kunyang Jia, Chuanjin Richard Shi, Hongxia Yang |
IJCAI | 4 |
| 2019 | A 340nW/Channel Neural Recording Analog Front-End using Replica-Biasing LNAs to Tolerate 200mVpp Interfere from 350mV Power SupplyabstractThis paper presents an 8-channel power-efficient neural recording analog front-end (AFE) with high power-supply rejection ratio (PSRR) and wide dynamic range. The ultra-low power is achieved by using a low supply voltage current-reusing input stage in the low noise amplifier (LNA). In order to improve the PSRR in low supply voltage amplifiers, we propose a replica biasing circuit to generate the biasing current, which is insensitive to the supply noise. Furthermore, the dynamic range is enlarged by utilizing an averaged local field potential (A-LFP) feedback loop. The prototype is fabricated in a 65nm CMOS process. Each channel of the AFE occupies 0.04mm2and only consumes 340nW from 0.35V/0.7V dual supply. The AFE provides a maximum gain of 54dB with 6.7μV input-referred noise integrating from 0.5Hz to 6.5 kHz. The proposed 0.35V input stage can tolerate a supply interferer up to 200mVpp, while maintaining a PSRR of 74dB. Liangjian Lyu, Dawei Ye, Chuanjin Richard Shi |
ISCAS | 3 |
| 2019 | A 2.46GHz, -88dBm Sensitivity CMOS Passive Mixer-First Nonlinear Receiver with >50dB Tolerance to In-Band InterfererabstractThis paper presents a -88 dBm sensitivity, 150Kbp/s OOK mixer-first nonlinear receiver in 65nm CMOS operating at the 2.46GHz ISM band. Since the LNA in the 1stIF band can be saturated by the strong in-band interferer, the shifted limiter (SL) is used to improve the interference resilience. Hence, by using an input power detection block, the 1stgain stage in the 1stIF band can alternatively turn on the LNA or the SL to improve the dynamic range. The in-band SIR at +/-1, 3 and 5MHz are measured to be -43/-11, -53/-54 and -53/-56dB respectively, while just consumes 380 to 610μW. Dawei Ye, Rongjin Xu, Liangjian Lyu, Chuanjin Richard Shi |
ISCAS | 4 |
| 2019 | Analysis of Bitwise and Samplewise Switched Passive Charge Sharing SAR ADCsabstractThis paper presents the analysis of bitwise and samplewise switched passive charge sharing for successive approximation register (SAR) analog-to-digital conversion (ADC). Closed-form analytic expressions of ADC transfer functions are derived based on charge conservation and validated by behavioral and schematic simulations. This leads to two elegant results for SAR ADCs with bitwise switched reference charge reservoirs (BS-RCRs). First, a binary-weighted SAR ADC implemented with BS-RCRs is transformed into a subradix-2 ADC. Second, the reference error caused by finite reservoir capacitance appears in the form of bit weight error. This error can be corrected digitally or by selecting a sufficiently large bit reference capacitance to bit weight capacitance ratio$\beta $. However, the reference error with samplewise switched reference charge reservoir (SS-RCR) is input dependent. In addition, an equivalent-circuit model-based analysis method is introduced, which shows more circuit intuition why BS-RCRs have better linearity than SS-RCR. A case study of an 11-bit 100-MS/s SAR ADC in 65-nm CMOS is presented. Chuanjin Richard Shi, Aili Wang 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Exploring the programmability for deep learning processors: from architecture to tensorizationabstractThis paper presents an instruction and Fabric Programmable Neuron Array (iFPNA) architecture, its 28nm CMOS chip prototype, and a compiler for the acceleration of a variety of deep learning neural networks (DNNs) including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and fully connected (FC) networks on chip. The iFPNA architecture combines instruction-level programmability as in an Instruction Set Architecture (ISA) with logic-level reconfigurability as in a Field-Programmable Gate Array (FPGA) in a sliced structure for scalability. Four data flow models, namely weight stationary, input stationary, row stationary and tunnel stationary, are described as the abstraction of various DNN data and computational dependence. The iFPNA compiler partitions a large-size DNN to smaller networks, each being mapped to, optimized and code generated for, the underlying iFPNA processor using one or a mixture of the four data-flow models. Experimental results have shown that state-of-art large-size CNNs, RNNs, and FC networks can be mapped to the iFPNA processor achieving the near ASIC performance. Chixiao Chen, Huwan Peng, Xindi Liu, Chuanjin Richard Shi |
DAC | 5 |
| 2018 | Constrained Optimization Based Low-Rank Approximation of Deep Neural Networks
Chuanjin Richard Shi |
ECCV (10) | 2 |
| 2018 | Pin-Efficient 12-Bit 8-Wire 8-Level Permutation Coding for High-Speed Parallel Wireline TranceiversabstractA 12-bit 8-wire 8-level (12B8W8L) permutation coding scheme is designed for high speed parallel I/O interfaces. The proposed 12B8W8L permutation coding scheme improves pin efficiency to 150%, compared to 50% pin efficiency of the non-return-to-zero (NRZ) signaling, while keeping the encoding and decoding logic simple. It reduces the baud rate to 1/3 of NRZ signaling compensating for reduced SNR. The proposed coding scheme eliminates simultaneous switching noise (SSN) and reference voltage noise. A prototype transceiver is designed, simulated and verified in a 65 nm CMOS process, achieving 6 Gb/s/pin data rate and 0.99 pJ/bit energy efficiency over 4-inch FR4 channels. Ailing Piao, Aili Wang 0002, Chuanjin Richard Shi |
ISCAS | 3 |
| 2018 | A 10-bit 50-MS/s SAR ADC with 1 fJ/Conversion in 14 nm SOI FinFET CMOS
Aili Wang 0002, Chuanjin Richard Shi |
Integr. | 2 |
| 2016 | A sub-nW mV-range programmable threshold comparator for near-zero-energy sensingabstractThis paper describes a comparator capable of detecting mV-range input voltage signals reliably using sub-nW power consumption. The comparator uses a current-mirror-based positive feedback and hysteresis to generate the mV-range threshold. It has three new ideas: (1) to use the input signal to bias the current mirror, which only activates the current mirror when the input signal reaches a detectable threshold; (2) to achieve dynamically controllable detectable threshold by using control signals to adjust the mirror transistor sizes; and (3) to use negative feedback to compensate the process-voltage-temperature variation. Preliminary SPICE simulation results using state-of-the-art 9HP (BiCMOS) have shown that the proposed hysteresis-based comparator can achieve the following performance metrics: 1) Very low comparator threshold (in the range of mV); furthermore, these threshold voltages can be programmable at run-time from a few to tens of mV. 2) Steep equivalent subthreshold slope in the range of sub-mV/decade. 3) Ultra-low leakage power (in the sub-nW range for a 0.1mV input voltage) and switching energy (in the nW range). The proposed hysteresis-based comparator only requires one standard logic supply source (working over 1V to 0.5V). Furthermore, initial simulations have shown that this comparator threshold is reliable over the process and voltage (PV) variations, and linearly proportional to temperature (T) variation. Aili Wang 0002, Allen Waters, Chuanjin Richard Shi |
ISCAS | 3 |
| 2016 | Highly time-interleaved noise-shaped SAR ADC with reconfigurable orderabstractA novel technique for implementing noise-shaping in time-interleaved SAR ADCs is presented. The noise-shaping is implemented passively and is easily reconfigured from first to second order. In both cases it achieves better SNR improvement than a conventional loop filter for a range of low OSR, and without the need for a power-hungry integrator. The architecture provides an energy-efficient ADC solution for a wide range of sample rates and resolution. To the best of the authors' knowledge, it is the first example of a noise-shaped ADC with more than two interleaved channels. Allen Waters, Aili Wang 0002, Chuanjin Richard Shi |
ISCAS | 3 |
| 2012 | Low-power LDPC decoding based on iteration predictionabstractLow-density parity-check (LDPC) codes have very broad applications. Low-power LDPC decoder design is becoming increasingly important for wireless and other power-constraint systems. Compared to disabling or simplifying the decoder circuit, reducing the power supply voltage can bring more power reduction. Through predicting the number of iterations needed for convergence, this paper proposes to make full use of the time available for decoding and scale down the power supply voltage in the remaining decoding iterations. Novel iteration prediction schemes are developed. The proposed schemes require very small hardware overhead and do not lead to noticeable error-correcting performance loss. Compared to using the original supply voltage and powering off the decoder after convergence, the proposed schemes can bring 50% dynamic power reduction for an example LDPC code, and the power saving further increases with the signal-to-noise ratio. Xinmiao Zhang 0001, Fang Cai, Chuanjin Richard Shi |
ISCAS | 3 |
| 2010 | Mixed-signal system-on-chip verification using a recursively-verifying-modeling (RVM) methodologyabstractThe verification of mixed-signal SoC is emerging as the most significant challenge, and with its cost surpassing the chip design cost. This paper presents a new automated verification methodology namely RVM (recursively verifying and modeling) and a set of supporting electronic design automation tools. The RVM methodology is built on the existing design flow and environments but with three major innovations to cope with custom-designed transistor blocks: a tool for automatically generating and validating simulation-efficiently behavioral models from a circuit netlist, a tool for characterizing and verifying the electrical rule correctness of analog blocks, and a hierarchical environment that allows designers to control the modeling and verification complexity. With the RVM methodology, analog circuits are verified in a way similar to the well-established digital verification. A set of industry benchmark results have shown that the RVM methodology is cable of reducing the verification time by potentially 100× to 1000×. With the increasing complexity of full-chip mixed-signal system-on-chip design, the RVM methodology is emerging as the only scalable verification solution. Chuanjin Richard Shi |
ISCAS | 1 |
| 2008 | Symmetry-aware placement with transitive closure graphs for analog layout designabstractA new scheme is proposed to use transitive closure graph (TCG) to explore the full symmetry solution space in analog layout design. We define a set of TCG symmetric-feasible conditions and show that it is extremely useful in reducing the solution space. A method is presented for generating random symmetric-feasible TCGs in O(n) time preserving the TCG closure property. Experimental results have confirmed the effectiveness of the proposed symmetry-aware TCG placement algorithm. Chuanjin Richard Shi, Yingtao Jiang |
ASP-DAC | 2 |
| 2008 | A 6-11GHz multi-phase VCO design with active inductorsabstractA multiphase VCO using differential active inductors is designed and fabricated in an IBM 0.13um CMOS process. Using active inductors, the core VCO occupies 0.3×0.4 mm2, and the central frequency can vary from 6GHz to 11GHz, exhibiting a 58% tunable frequency ranges. Exploiting current reuse, the phase noise is measured to be less than −95dBc/Hz at 1MHz offset with the measured power consumption ranging from 12 to 31mW in 6–11GHz. A signal coupling technique with phase shifting is employed to generate eight phase signals. The achieved phase errors of this design are smaller than 3°. In addition, bias compensation is shown to reduce the frequency variation under 10% while the temperature varies from −55°C to 125°C. Yu-Te Liao, Chuanjin Richard Shi |
ISCAS | 2 |
| 2008 | A quantum-dot light-harvesting architecture using deterministic phase controlabstractEfficient solar-energy harvesting is fundamental to solar cell technology. Much research effort has been devoted to the construction of new light-harvesting structures, including the use of semiconductor quantum dots (QDs), to improve the widespread availability of solar cells. In this paper, a new light-harvesting architecture is considered, which utilizes quantum dots. The proposed architecture is composed of quantum phase-locked loops (QPLLs) to enhance the harvesting efficiency of QD solar cells by utilizing feedback control principles. The purpose of QPLL is to synchronize the phases of monochromatic light harvested by the antenna systems. This paper addresses a deterministic modeling and control formulation of the QPLL. The QPLL consists of a tracking controller and a proportional-integral (PI) controller. Simulation results for the controllers are presented and discussed. Cherry Wakayama, Wolf Kohn, Zelda B. Zabinsky, Chuanjin Richard Shi |
ISCAS | 4 |
| 2008 | Simulation of Closely Related Dynamic Nonlinear Systems With Application to Process-Voltage-Temperature Corner AnalysisabstractEven for a single circuit, it has become increasingly time consuming to simulate at the SPICE level. However, the situation is getting worse when it comes to simulating thousands of such circuits, where often one circuit is closely related to another. This problem arises in applications such as process-voltage-temperature corner circuit simulation or simulation-in-the-loop circuit optimization. The traditional approach to solving this problem is to repeatedly invoke SPICE simulation on each of those circuits. Such an approach does not exploit the similarity among those circuits and could lead to prohibitively high computational cost. This paper presents a new simulation approach capable of simulating hundreds and thousands of closely related systems with the computational cost comparable to or even less than that of a few simulations, yet with the same simulation accuracy and robustness. The proposed approach is based on the combination of the LU-factorization-based direct method (used to construct preconditioners) and Krylov-subspace-based iterative methods (used to solve circuit equations) to explore the common characteristics shared by a set of closely related systems. The key novelty is a systematic method that uses the fewest direct solving for underlying linearized systems and then solves the rest using Krylov-subspace-based iterative methods, with preconditioners computed from those LU factors. In addition, a method of automatically constructing preconditioners from device equations has been developed based on model compilation and demonstrated on MOS transistor Berkeley short-channel IGFET model (BSIM) models. Several circuit examples are included to show the effectiveness of the proposed approach. Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Parasitic-Aware Optimization and Retargeting of Analog Layouts: A Symbolic-Template ApproachabstractLayout parasitics can significantly affect the performance of analog integrated circuits (ICs). In this paper, a systematic method of optimizing an existing analog layout considering parasitics is presented for technology migration and performance retargeting. This method represents the locations of layout rectangle edges as variables and extracts circuit and layout integrity such as device symmetry, matching, and design rules as constraints. To ensure the desired circuit performance, bounds of layout parasitics are determined first. These bounds are used to constrain the layout geometries while retargeting existing high-quality layouts across technologies and specification sets. The problem is then solved by a graph-based algorithm combined with nonlinear optimization. The proposed method has been implemented in a parasitic-aware automatic layout optimization and retargeting tool (intellectual property reuse-based analog IC layout). Its efficiency and effectiveness are demonstrated by successfully retargeting operational amplifiers within 1 min of CPU time. Nuttorn Jangkrajarng, Sambuddha Bhattacharya, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | A Graph Reduction Approach to Symbolic Circuit AnalysisabstractA new graph reduction approach to symbolic circuit analysis is developed in this paper. A Binary Decision Diagram (BDD) mechanism is formulated, together with a specially designed graph reduction process and a recursive sign determination algorithm. A symbolic analog circuit simulator is developed using a combination of these techniques. The simulator is able to analyze large analog circuits in the frequency domain. Experimental results are reported. Guoyong Shi, Chuanjin Richard Shi |
ASP-DAC | 3 |
| 2007 | Maximizing the throughput-area efficiency of fully-parallel low-density parity-check decoding with C-slow retiming and asynchronous deep pipeliningabstractIn this paper, we apply C-slow retiming and asynchronous deep pipelining to maximize the throughput-area efficiency of fully parallel low-density-parity-check (LDPC) decoding. Pipelined decoders are implemented in a 0.18 mum FDSOI CMOS process. Experimental results show that our pipelining technique is an efficient approach to maximizing LDPC decoding throughput while minimizing the area consumption. First, pipelined decoders can achieve extraordinary high throughput which non-pipelined design cannot. Second, for the same throughput, pipelined decoders use less area than non-pipelined design. Our approach can improve the throughput of a published implementation by 4 times with only about 80% area overhead. Without using clocks, proposed asynchronous pipelined decoders are more scalable in design complexity and more robust to process-voltage-temperature variations than existing clock-based LDPC decoders. Ming Su, Chuanjin Richard Shi |
ICCD | 3 |
| 2007 | Implementing a 2-Gbs 1024-bit 1/2-rate low-density parity-check code decoder in three-dimensional integrated circuitsabstractA 1024-bit, 1/2-rate fully parallel low-density parity-check (LDPC) code decoder has been designed and implemented using a three-dimensional (3D) 0.18 mum fully depleted silicon-on-insulator (FDSOI) CMOS technology based on wafer bonding. The 3D-IC decoder was implemented with about 8M transistors, placed on three tiers, each with one active layer and three metal layers, using 6.9 mm by 7.0 mm of die area. It was simulated to have a 2 Gbps throughput, and consume only 260 mW. This first large-scale 3D application-specific integrated circuit (ASIC) with fine-grain (5mum) vertical interconnects is made possible by jointly developing a complete automated 3D design flow from a commercial 2-D design flow combined with the needed 3D-design tools. The 3D implementation is estimated to offer more than 10 xpower-delay-area product improvement over its corresponding 2D implementation. The work demonstrated the benefits of fine-grain 3D integration for interconnect-heavy very-large-scale digital ASIC implementation. Cherry Wakayama, Robin Panda, Nuttorn Jangkrajarng, Chuanjin Richard Shi |
ICCD | 6 |
| 2007 | VHDL-AMS based modeling and simulation of mixed-technology microsystems: a tutorial
Pavel V. Nikitin, Chuanjin Richard Shi |
Integr. | 2 |
| 2007 | CASCADE: A Standard Supercell Design Methodology With Congestion-Driven Placement for Three-Dimensional Interconnect-Heavy Very Large-Scale Integrated CircuitsabstractIn this paper, CASCADE, a standard supercell-based design methodology, its supporting automated design flow, and associated design tools, are presented for 3D implementations of a class of interconnect-heavy application-specific very large-scale integrated circuits. In CASCADE, a system is first partitioned and synthesized using standard 2D design tools to a set of supercells with the same height and varying widths. With this, the 3D design is reduced to 3D supercell placement and 3D-via assignment. A congestion-driven simulated-annealing method is used to find a 3D placement of supercells to minimize the total wire length, the longest wire length, and the number of 3D vias and routing density. To efficiently estimate the routing density of a 3D grid space within the optimization loop, a simple probabilistic congestion model with an incremental congestion computation has been developed. Once the supercell placement is fixed, the problem of assigning 3D vias to accomplish minimal 2D routing densities and uniform 3D-via distribution is solved by an efficient min-cost-max-flow method. The proposed methods have been implemented and tested on a set of ISPD98 circuit benchmarks. Experimental results have shown that the proposed congestion-driven 3D supercell placement and flow-based 3D-via-assignment tools have yielded satisfactory placement with small-area, low-congestion, short-wire-length, few, and uniformly distributed 3D vias. Furthermore, an excellent correlation between routing-density estimation by our model and the actual routing performed by a commercial router has been observed. We have applied the proposed 3D design methodology, tools, and flows to tape out an over 4-million-gate low-density parity-check decoder in a three-tier 0.18- fully depleted silicon-on-insulator 3D CMOS process manufactured by MIT Lincoln Laboratory. The postlayout simulation of this DRC-clean layout design showed an about ten times improvement on the power-delay-area product compared to a 2D implementation in the same process. Cherry Wakayama, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | A quasi-newton preconditioned Newton-Krylov method for robust and efficient time-domain simulation of integrated circuits with strong parasitic couplingsabstractIn this paper, the Newton-Krylov method is explored for robust and efficient time-domain VLSI circuit simulation. Different from the LU-factorization based direct method, the Newton-Krylov method uses a preconditioned Krylov-subspace iterative method for linear system solving. Our key contribution is to introduce an effective quasi-Newton preconditioning scheme for Krylov-subspace methods to reduce the number and cost of LU factorizations during time-domain circuit simulation. Experimental results on a collection of digital, analog and RF circuits have shown that the quasi-Newton preconditioned Krylov-subspace method is as robust and accurate as SPICE3. The proposed Newton-Krylov method is especially attractive for simulating circuits with a large amount of parasitic RLC elements for post-layout verification. Chuanjin Richard Shi |
ASP-DAC | 2 |
| 2006 | A high-throughput low-power fully parallel 1024-bit 1/2-rate low density parity check code decoder in 3-dimensional integrated circuitsabstractA 1024-bit, 1/2 -rate fully parallel low density parity check (LDPC) code decoder has been designed and implemented using a 3D 0.18/spl mu/m fully depleted silicon-on-insulator (FDSOI) CMOS technology based on wafer bonding. The taped-out 3D decoder with about 8M transistors was simulated to have a high throughput of 2Gb/s and a low power consumption of only 430mW using 6.4/spl mu/m by 6.3/spl mu/m of die area. The 3D implementation is estimated to offer more than 10/spl times/ power-delay-area product improvement over its corresponding 2D implementation. This first large-scale 3D ASIC with fine-grain (5/spl mu/m) vertical interconnects is made possible by jointly developing a complete automated 3D design flow from a commercial 2D design flow combined with the needed 3D-design point tools. Cherry Wakayama, Nuttorn Jangkrajarng, Chuanjin Richard Shi |
ASP-DAC | 5 |
| 2006 | Template-based parasitic-aware optimization and retargeting of analog and RF integrated circuit layoutsabstractParasitic effects are extremely significant for the performance of analog and RF integrated circuits. Although layout retargeting for technology migration or specification update is able to preserve designers' intent, the associated layout parasitics cannot be guaranteed to meet the performance requirements. In this paper, we present a novel algorithm that performs parasitic-aware automatic layout retargeting for analog/RF integrated circuits. Given parasitic resistance/capacitance bounds and matching constraints ensuring desired circuit performance, the algorithm creates a reduced-template-graph from original layouts and adds parasitic constraints. Using a two-dimensional hybrid scheme of graph-based optimization and nonlinear programming, the nonlinear problem is solved effectively and efficiently. The algorithm has successfully retargeted operational amplifiers and an RF low-noise amplifier within minutes of CPU time. Nuttorn Jangkrajarng, Sambuddha Bhattacharya, Nathan Kohagen, Chuanjin Richard Shi |
ICCAD | 5 |
| 2006 | Improved automatic differentiation method for efficient model compilerabstractThe growing complexity of semiconductor devices presents a serious challenge to model developers. Compact device model compilers have been developed over the years to address this challenge. However, those compilers cannot generate very computationally efficient device models. In this paper, we present several effective techniques to improve the automatic differentiation method of model compiler MCAST so that it can generate very efficient device models. Those techniques include function strength reduction, minimum division, minimum function calls, and hierarchical automatic differentiation. Experimental results in modeling industry level semiconductor devices such as BSIM3, BSIM4 and B3SOI, demonstrated that the presented techniques lead to very computationally efficient models that are comparable to manually tuned ones Chuanjin Richard Shi |
ISCAS | 2 |
| 2006 | FROSTY: A program for fast extraction of high-level structural representation from circuit description for industrial CMOS circuits
Lei Yang 0019, Chuanjin Richard Shi |
Integr. | 2 |
| 2006 | Multilevel symmetry-constraint generation for retargeting large analog layoutsabstractThe strong impact of layout intricacies on analog-circuit performance poses great challenges to analog layout automation. Recently, template-based methods have been shown to be effective in reuse-centric layout automation for CMOS analog blocks such as operational amplifiers. The layout-retargeting method first creates a template by extracting a set of constraints from an existing layout representation. From this template, new layouts are then generated corresponding to new technology processes and new device specifications. For large analog layouts, however, this method results in an unmanageable template due to a tremendous increase in the number of constraints, especially those emerging from layout symmetries. In this paper, we present a new method of multilevel symmetry-constraint generation by utilizing the inherent circuit structure and hierarchy information from the extracted netlist. The method has been implemented in a layout-retargeting system called Intellectual Property Reuse-based Analog IC Layout (IPRAIL) and demonstrated 18 times reduction in the number of symmetry constraints required for retargeting an analog-to-digital converter layout. This enables our retargeting engine to successfully handle the complexities associated with large analog layouts. While manual relayout is known to take weeks, our layout-retargeting tool generates the target layout in hours and achieves comparable electrical performance Sambuddha Bhattacharya, Nuttorn Jangkrajarng, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | SILCA: SPICE-accurate iterative linear-centric analysis for efficient time-domain Simulation of VLSI circuits with strong parasitic couplingsabstractA new circuit analysis method, named SPICE-accurate iterative linear-centric analysis (SILCA), is proposed for the efficient and accurate time-domain simulation of deep submicron very large scale integrated (VLSI) circuits with strong parasitic couplings. SILCA consists of two key linear-centric techniques applied to time-domain nonlinear circuit simulation. For numerical integration, explicit-formula substitution and iterative-formula transformation are presented to convert implicit variable time-step integration to fixed leading coefficient (FLC) variable time-step integration. This paper characterizes both convergence and stability properties of the resulting FLC integration formulae. For nonlinear iteration, a successive variable chord (SVC) method is used as an alternative to the Newton-Raphson method. Further, the low-rank update technique is implemented for fast LU factorization. With these techniques, the number and cost of required LU factorizations are reduced dramatically. Experimental results on nonlinear circuits coupled with substrate and power/ground networks have demonstrated that SILCA achieves more than an order of magnitude speedup over SPICE3 in terms of both the cost of LU factorization and the overall CPU time. SILCA is suitable for efficient SPICE-like time-domain simulation of parasitic-coupled VLSI circuits, where the number of linear parasitic elements dominates the number of nonlinear devices. Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | A Quasi-Newton Preconditioned Newton-Krylov Method for Robust and Efficient Time-Domain Simulation of Integrated Circuits With Strong Parasitic CouplingsabstractIn this paper, the Newton-Krylov method is explored for robust and efficient time-domain simulation of integrated circuits with large amount of parasitic elements. Different from LU-factorization-based direct methods used in SPICE-like circuit simulators, the Newton-Krylov method uses a preconditioned Krylov-subspace iterative method for solving linearized-circuit equations. A key contribution of this paper is to introduce an effective quasi-Newton preconditioning scheme for Krylov-subspace methods to reduce the number and cost of LU factorization during an entire time-domain circuit simulation. The proposed quasi-Newton preconditioning scheme consists of four key techniques: 1) a systematic method for adaptively controlling time step sizes; 2) automatically generated piecewise weakly nonlinear (PWNL) definition of nonlinear devices to construct quasi-Newton preconditioners; 3) low-rank update techniques for incrementally updating preconditioners; and 4) incomplete-LU preconditioning for efficiency. Experimental results on a collection of digital, analog, and RF circuits have shown that the quasi-Newton preconditioned Krylov-subspace method is as robust and accurate as the direct method used in SPICE. The proposed Newton-Krylov method is attractive for simulating circuits with massive parasitic RLC elements for postlayout verification. For a nonlinear circuit with power/ground networks with tens-of-thousand elements, the CPU time speedup over SPICE3 is over 20X, and it is expected to increase further with the circuit size Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | On symbolic model order reductionabstractSymbolic model order reduction (SMOR) is a macromodeling technique that generates reduced-order models while retaining the parameters in the original models. Such symbolic reduced-order models can be repeatedly simulated with a greater efficiency for varying model parameters. Although the model-order-reduction concept has been extensively developed in literature and widely applied in a variety of problems, model order reduction from a symbolic perspective has not been well studied. Several methods developed in this paper include symbol isolation, nominal projection, and first-order approximation. These methods can be applied to models having only a few parametric elements and to models having many symbolic elements. Of special practical interest are models that have slightly varying parameters such as process related variations, for which efficient reduction procedures can be developed. Each technique proposed in this paper has been tested by circuit examples. Experiments show that the proposed methods are efficient and effective for many circuit problems Guoyong Shi, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Efficient DC fault simulation of nonlinear analog circuits: one-step relaxation and adaptive simulation continuationabstractEfficient dc fault simulation of nonlinear analog circuits is addressed in this paper. Two techniques, one-step relaxation and adaptive simulation continuation, are proposed. By one-step relaxation, only one Newton-Raphson iteration is performed for each faulty circuit with the dc solution of the good circuit as the initial point, and the approximate solution is used for detecting the fault. The paper shows experimentally and justifies theoretically that approximate dc fault simulation by one-step relaxation can accomplish almost the same fault coverage as exact dc fault simulation. Exact dc fault simulation by adaptive simulation continuation is first to order faulty circuits based on the results of one-step relaxation, and then to use the solution of the previous faulty circuit as the initial point for the Newton-Raphson iteration of the next faulty circuit. Experiments on a set of 29 MCNC Circuit Simulation and Modeling Workshop benchmark circuits show that exact dc fault simulation by adaptive simulation continuation can achieve an average speedup of 4.4 and as high as 15 over traditional stand-alone fault simulation. Chuanjin Richard Shi, Michael W. Tian, Guoyong Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | An FPGA implementation of low-density parity-check code decoder with multi-rate capabilityabstractWith superior error correction capability, low-density parity-check (LDPC) has initiated wide scale interests in wireless telecommunication fields. In the past, various structures of single code rate LDPC decoders have been implemented for different applications. However, in order to cover a wide range of service requirements and diverse interference conditions in wireless applications, LDPC decoders that can operate in both high and low code rates are desired. In this paper, a new multi-rate LDPC decoder architecture is presented and implemented in a Xilinx FPGA device. Through selection pins, three operating modes with the irregular 1/2 rate, regular 5/8 rate and regular 7/8 rate are supported. The measurement results show LDPC decoder can achieve BER below 10-5 at SNR of 1.4dB in the most critical case with the irregular 1/2 mode. Lei Yang 0019, Manyuan Shen, Hui Liu 0011, Chuanjin Richard Shi |
ASP-DAC | 4 |
| 2005 | Template-driven parasitic-aware optimization of analog integrated circuit layoutsabstractLayout parasitics have great impact on analog circuit performance. This paper presents an algorithm for explicit parasitic control during layout retargeting of analog integrated circuits. In order to ensure desired circuit performance, bounds on layout parasitics' magnitudes are determined first. Then, graph techniques are coupled with mathematical programming to constrain layout geometry based on these parasitic bounds. The algorithm has been demonstrated to ensure desired circuit performance during technology migration and performance specification changes. Sambuddha Bhattacharya, Nuttorn Jangkrajarng, Chuanjin Richard Shi |
DAC | 3 |
| 2005 | An Efficiently Preconditioned GMRES Method for Fast Parasitic-Sensitive Deep-Submicron VLSI Circuit SimulationabstractWe propose an efficiently preconditioned generalized minimal residual (GMRES) method for fast SPICE-accurate transient simulation of parasitic-sensitive deep-submicron VLSI circuits. First, when time step-sizes vary within a predefined range, the preconditioned GMRES method is applied to solve circuit matrix equations rather than LU factorization. The preconditioner we use comes directly from the previously factorized L and U matrices. Second, to keep using the same preconditioner during nonlinear iteration, the successive variable chord method is applied as an alternative to the Newton-Raphson method. An improved piecewise weakly nonlinear definition of MOSFET is adopted and the low-rank update technique is implemented to refresh the preconditioner efficiently. With these techniques, the number of required LU factorizations during transient simulation is reduced dramatically. Experimental results on power/ground networks have demonstrated that the proposed method yields SPICE-like accuracy with an about 18/spl times/ overall CPU time speedup over SPICE3 for circuits with tens of thousands elements. Chuanjin Richard Shi |
DATE | 2 |
| 2005 | VLSI implementation of a low-error-floor and capacity-approaching low-density parity-check code decoder with multi-rate capacityabstractWith the superior error correction capability, low-density parity-check (LDPC) codes have initiated wide scale interests in wireless communication and storage fields. In the past, various structures of single code-rate LDPC decoders have been reported. However, to cover a wide range of service requirements and diverse interference conditions in wireless applications, LDPC decoders that can operate at both high and low code rates are desirable. In this paper, a 9k code length multi-rate LDPC decoder architecture is presented and implemented on a Xilinx FPGA device. Using pin selection, three operating modes, namely, the irregular 1/2 code, the regular 5/8 code and the regular 7/8 code, are supported. Furthermore, to suppress the error floor level, a characterization on the conditions for short cycles in a LDPC code matrix expanded from a small base matrix is presented, and a cycle elimination algorithm is developed to detect and break such short cycles. The effectiveness of the cycle elimination algorithm has been verified by both simulation and hardware measurements, which show that the error floor is suppressed to a much lower level without incurring any performance penalty. The implemented decoder is tested in an experimental LDPC-OFDM system and achieves the superior measured performance of block error rate below 10/sup -7/ at SNR 1.8 dB. Lei Yang 0019, Hui Liu 0011, Chuanjin Richard Shi |
GLOBECOM | 3 |
| 2005 | Noise aware behavioral modeling of the E-Delta fractional-N frequency synthesizerabstractThis paper presents the behavioral model of a Ε-Δ fractional-N frequency synthesizer in terms of different noise sources and non-ideal effects. To accurately predict the phase noise of the synthesizer, different jitter noise sources such as phase modulation (PM) noise in phase-frequency detector and divider, frequency modulation (FM) noise in VCO are properly depicted. The Ε-Δ modulator, with its divider value dithered and quantization noise dynamically injected to the PLL, is described in behavioral model, which allows the designer to study the quantization noise impaction to the PLL phase noise. All the models are implemented in VHDL-AMS and simulated using Mentor Graphics ADvance-MS (ADMS). Our behavioral modeling method enables a fast simulation of the PLL system and an accurate phase noise prediction. Lei Yang 0019, Cherry Wakayama, Chuanjin Richard Shi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Fast-yet-accurate PVT simulation by combined direct and iterative methodsabstractThe operations and performances of deep-submicron integrated circuits are affected significantly by the variations of process parameters, power supply voltages and operating temperatures. Circuit simulation for all the combinations of process-voltage-temperature (PVT) conditions, known as PVT simulation, is emerging as a must not only for analog and RF circuit designs, but also for the designs of digital library cells and critical paths. With the number of PVT conditions in the range of hundreds and even thousands, existing solution of invoking a simulator repeatedly for all these PVT conditions is becoming extremely time consuming. This work presents a new simulation approach capable of simulating hundreds and thousands PVT corners with the computational cost comparable to or even less than that of a few corner simulations, yet with the same simulation accuracy and robustness. The proposed approach is based on the combination of the LU-factorization based direct method and Krylov subspace based iterative methods to explore the common characteristics shared by a circuit under all PVT corners. The key novelty is a systematic method that uses as few LU based direct solving as possible for underlying linearized systems, and then solves the rest of linearized systems across the entire PVT linear system space using Krylov subspace based iterative methods with preconditioners computed from those LU factors. Chuanjin Richard Shi |
ICCAD | 2 |
| 2004 | Hierarchical extraction and verification of symmetry constraints for analog layout automation
Sambuddha Bhattacharya, Nuttorn Jangkrajarng, Roy Hartono, Chuanjin Richard Shi |
ASP-DAC | 4 |
| 2004 | Multiple specifications radio-frequency integrated circuit design with automatic template-driven layout retargeting
Nuttorn Jangkrajarng, Sambuddha Bhattacharya, Roy Hartono, Chuanjin Richard Shi |
ASP-DAC | 4 |
| 2004 | CrtSmile: a CAD tool for CMOS RF transistor substrate modeling incorporating layout effects
Ravikanth Suravarapu, Roy Hartono, Sambuddha Bhattacharya, Kartikeya Mayaram, Chuanjin Richard Shi |
ASP-DAC | 6 |
| 2004 | Parametric reduced order modeling for interconnect analysis
Guoyong Shi, Chuanjin Richard Shi |
ASP-DAC | 2 |
| 2004 | Correct-by-construction layout-centric retargeting of large analog designsabstractAggressive design cycles in the semiconductor industry demand a design-reuse principle for analog circuits. The strong impact of layout intricacies on analog circuit performance necessitates design reuse with special focus on layout aspects. This paper presents a computer-aided design tool and the methodology for a layout-centric reuse of large analog intellectual-property blocks. From an existing layout representation, an analog circuit is retargeted to different processes and performances; the corresponding correct-by-construction layouts are generated automatically and have performances comparable to manually crafted layouts. The tool and the methodology are validated on large analog intellectual-property blocks. While manual re-design and re-layout is known to take weeks to months, our reuse tool-suite achieves comparable performance in hours. Sambuddha Bhattacharya, Nuttorn Jangkrajarng, Roy Hartono, Chuanjin Richard Shi |
DAC | 4 |
| 2004 | Hierarchical Multi-Dimensional Table Lookup for Model Compiler Based Circuit SimulationabstractIn this paper, a systematic method for automatically generating hierarchical multi-dimensional table lookup models for compact device and behavioral models with any number of terminals is presented. The method is based on an Abstract Syntax Tree representation of analytic equations. Expensive part of the computations represented by abstract syntax trees are identified and replaced by two-dimensional table lookup models. An error-control based optimization algorithm is developed to generate table lookup models with the minimal amount of table data for a given accuracy requirement. The proposed method has been implemented in the model compiler MCAST and the circuit simulator SPICE3. Experimental results show that, compared to non-optimized compilation based simulation, the simulation using the proposed table lookup optimization method is about 40 times faster and achieves sufficiently accurate results with error less than 1-2%. Chuanjin Richard Shi |
DATE | 2 |
| 2004 | Efficient approximation of symbolic expressions for analog behavioral modeling and analysisabstractEfficient algorithms are presented to generate approximate expressions for transfer functions and characteristics of large linear-analog circuits. The algorithms are based on a compact determinant decision diagram (DDD) representation of exact transfer functions and characteristics. Several theoretical properties of DDDs are characterized, and three algorithms, namely, based on dynamic programming, based on consecutive k-shortest path (SP), and based on incremental k-SP, are presented in this paper. We show theoretically that all three algorithms have time complexity linearly proportional to |DDD|, the number of vertices of a DDD, and that the incremental k-SP-based algorithm is fastest and the most flexible one. Experimental results confirm that the proposed algorithms are the most efficient ones reported so far, and are capable of generating thousands of dominant terms for typical analog blocks in CPU seconds on a modern computer workstation. Sheldon X.-D. Tan, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Efficient DDD-based term generation algorithm for analog circuit behavioral modelingabstractAn efficient approach to generating symbolic product terms for behavioral modeling of large linear analog circuits is presented. The approach is based on a compact determinant decision diagram (DDD) representation of transfer functions and characteristics of analog circuits. The new algorithm is based on the concept that a dominant term in a DDD graph can be found by searching the shortest path in the graph. But instead of traversing a whole DDD graph each time, we show that a shortest path can be found by just updating a small number of the newly added vertices after the first shortest path is found. Experimental results indicate that the new symbolic term generation algorithm outperforms both pure shortest path based algorithm and dynamic programming based algorithm, which is the fastest symbolic term generation algorithm published so far. Sheldon X.-D. Tan, Chuanjin Richard Shi |
ASP-DAC | 2 |
| 2003 | Symbolic analysis of analog circuits with hard nonlinearityabstractA new methodology is presented to solve a strongly nonlinear circuit, characterized by Piece-Wise Linear (PWL) functions, symbolically and explicitly in terms of its circuit parameters and is amenable to computer implementation. The method is based on a modified nodal formulation of piecewise linear circuit equations as a mixed Linear Complementarity Problem (MLCP). The technique of determinant-decision diagrams is applied to implement the symbolic transformation of the MLCP to the standard LCP. Complementarity-decision diagrams are used to represent the resulting LCP. Examples are presented that demonstrate the accuracy and efficiency of the proposed method. Alicia Manthe, Chuanjin Richard Shi |
DAC | 3 |
| 2003 | Symbolic Analysis of Nonlinear Analog Circuits
Alicia Manthe, Chuanjin Richard Shi, Kartikeya Mayaram |
DATE | 3 |
| 2003 | SILCA: Fast-Yet-Accurate Time-Domain Simulation of VLSI Circuits with Strong Parasitic Coupling Effects
Chuanjin Richard Shi |
ICCAD | 2 |
| 2003 | FROSTY: A Fast Hierarchy Extractor for Industrial CMOS Circuits
Lei Yang 0019, Chuanjin Richard Shi |
ICCAD | 2 |
| 2003 | Parametric Equivalent Circuit Extraction for VLSI Structures
Pavel V. Nikitin, Winnie Yam, Chuanjin Richard Shi |
VLSI-SOC | 3 |
| 2003 | IPRAIL - intellectual property reuse-based analog IC layout automation
Nuttorn Jangkrajarng, Sambuddha Bhattacharya, Roy Hartono, Chuanjin Richard Shi |
Integr. | 4 |
| 2003 | Balanced multi-level multi-way partitioning of analog integrated circuits for hierarchical symbolic analysis
Sheldon X.-D. Tan, Chuanjin Richard Shi |
Integr. | 2 |
| 2003 | Efficient very large scale integration power/ground network sizing based on equivalent circuit modelingabstractWe present an efficient method of minimizing the area of power/ground (P/G) networks in integrated circuit layouts subject to reliability constraints. Instead of directly sizing the original P/G network extracted from a circuit layout, as done previously, the new method first constructs a reduced but electrically equivalent P/G network. Then the sequence of linear programming method is applied to optimize the reduced network. The solution of the original network is then backsolved from the optimized reduced network. The new method exploits the regularities in the P/G networks to reduce the complexities of P/G networks. Experimental results show that the sizes of reduced networks are typically significantly smaller than that of the original networks. The resulting algorithm is fast enough that P/G networks with more than one million branches can be sized in a few minutes on modern Sun workstations. Sheldon X.-D. Tan, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Reliability-constrained area optimization of VLSI power/ground networks via sequence of linear programmingsabstractThis paper presents a new method of sizing the widths of the power and ground routes in integrated circuits so that the chip area required by the routes is minimized subject to electromigration and IR voltage drop constraints. The basic idea is to transform the underlying constrained nonlinear programming problem into a sequence of linear programs. Theoretically, we show (that the sequence of linear programs always converges to the optimum solution of the relaxed convex optimization problem. Experimental results demonstrate that the proposed sequence-of-linear-program method Is orders of magnitude faster than the best-known method based on conjugate gradients with constantly better solution qualities. Sheldon X.-D. Tan, Chuanjin Richard Shi, Jyh-Chwen Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Fast Power/Ground Network Optimization Based on Equivalent Circuit ModelingabstractThis paper presents an efficient algorithm for optimizing the area of power or ground networks in integrated circuits subject to the reliability constraints. Instead of solving the original power/ground networks extracted from circuit layouts as previous methods did, the new method first builds the equivalent models for many series resistors in the original networks, then the sequence of linear programming method [9] is used to solve the simplified networks. The solutions of the original networks then are back solved from the optimized, simplified networks. The new algorithm simply exploits the regularities in the power/ground networks. Experimental results show that the complexities of simplified networks are typically significantly smaller than that of the original circuits, which renders the new algorithm extremely fast. For instance, power/ground networks with more than one million branches can be sized in a few minutes on modern SUN workstations. Sheldon X.-D. Tan, Chuanjin Richard Shi |
DAC | 2 |
| 2001 | Distributed Event-Driven Simulation of VHDL-SPICE Mixed-Signal CircuitsabstractPresents a new framework and its prototype implementation for the distributed simulation of a mixed-signal system where some parts are modeled by differential and algebraic equations (as in SPICE) and other parts are modeled by discrete events (as in VHDL or Verilog). The work is built on top of a general-purpose framework of parallel and distributed simulation that combines both conservative and optimistic synchronization methods to extract the maximum concurrency available in such mixed analog and digital systems. We demonstrate that the maximum speedup can be achieved with digital components using optimistic scheduling and with analog components using conservative scheduling. A technique to further increase the parallelism among simultaneous events without sacrificing the simulation accuracy is proposed, by using a coarser time grain for the internal step events of analog solvers. Experimental results demonstrate a 6.2/spl times/ speedup on eight processors for a circuit described in VHDL and SPICE. This paper details our implementation, including how to make the numerical integration used by analog electrical-level event-driven simulation and to synchronize it with digital behavioral and gate-level simulation, analog/digital conversions, and how to resolve the non-convergence arising from straightforward analog and digital simulator integration. This paper also shows that optimistic synchronization for distributed analog simulation is feasible and quite efficient for small-to-medium-size analog blocks. Dragos Lungeanu, Chuanjin Richard Shi |
ICCD | 2 |
| 2001 | Lower Bound Based DDD Minimization for Efficient Symbolic Circuit AnalysisabstractThe determinant decision diagram (DDD) is a variant of binary decision diagrams (BDDs) for representing symbolic matrix determinants and cofactors in symbolic circuit analysis. Inspired by the ideas of Rudell (1993) and Drechsler et al. (2001) on BDD minimization, we present a lower-bound based sifting algorithm for reordering the DDD vertices to minimize the DDD size. Our contributions are (1) an adaptation of Rudell's sifting technique for DDD minimization with new rules for determining vertex signs, and (2) tighter lower bounds developed specifically for DDDs. On a set of DDD examples from symbolic circuit analysis, experimental results have demonstrated that the proposed lower-bound based reordering algorithm can effectively reduce DDD sizes. It has also been demonstrated that sifting with lower bounds uses about 50% less computation compared to sifting without using lower bounds, and sifting with the new lower bounds reduces the computation further by up to 8% compared to sifting with Drechsler's lower bounds for BDDs. Alicia Manthe, Chuanjin Richard Shi |
ICCD | 2 |
| 2001 | Compact representation and efficient generation of s-expandedsymbolic network functions for computer-aided analog circuit designabstractA graph-based approach is presented for the generation of exact symbolic network functions in the form of rational polynomials of the complex frequency variable s for analog integrated circuits. The approach employs determinant decision diagrams (DDDs) to represent the determinant of a circuit matrix and its cofactors. A notion of multiroot DDDs is introduced, where each root represents a symbolic expression for an individual coefficient of the powers of s in the numerator and denominator of a network function, and multiple roots share their common subgraphs. A DDD-based algorithm is presented for generating s-expanded network functions. We prove theoretically and validate experimentally that the algorithm constructs in O(kl|DDD|) time an s-expanded DDD with no more than kl|DDD| vertices, where k is the degree of the denominator s polynomial, l is the maximum number of devices that connect to a circuit node, and |DDD| is the number of DDD vertices representing the circuit-matrix determinant. For a practical circuit, |DDD| is often many orders-of-magnitude less than the number of product terms. In contrast, previous approaches require the time and space complexities proportional to the number of product terms, which grows exponentially with the size of a circuit. Experimental results have demonstrated that the new approach can produce exact s-expanded-symbolic network functions for /spl mu/A741 operational amplifiers in several CPU seconds on an UltraSparc-I workstation. The expressive power of multiroot s-expanded DDDs is so remarkable that in one instance, over 10/sup 35/ symbolic product terms have been represented by a multiroot DDD with less than 17 K vertices. The compactness of DDDs is further demonstrated in the context of symbolic noise evaluation, where potentially many transfer functions, each being used for a noise source in the circuit, can be represented by a single DDD with the size comparable to that for a few transfer functions. This provides a powerful tool for solving many symbolic analysis problems such as deriving interpretable symbolic expressions, dominant pole/zero estimation, and analog testability analysis. We have also demonstrated that repetitive numerical evaluation with the derived s-expanded symbolic expressions for frequency-domain simulation and small-signal noise analysis can be much faster than SPICE-like simulators and the resulting expressions for a circuit block can be used as behavioral models for high-level simulation. Chuanjin Richard Shi, Sheldon X.-D. Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2000 | Analog-testability analysis by determinant-decision-diagrams based symbolic analysisabstractThe use of the column-rank of the sensitivity matrix as a testability measure for parametric faults in linear analog circuits was pioneered by Saeks in 1970s, and later re-discovered by several others.Its practical use has been, however, limited by how it can be calculated.Numerical algorithms suffer from inevitable round-off errors, while classical symbolic techniques can only handle very small circuits.In this paper , an innovative and efficient graph based symbolic analysis approach, called Determinant Decision Diagrams, is applied to testability measurement and selection for optimum test vectors.The new approach is promising in testability analysis of much larger analog circuits. Tao Pi, Chuanjin Richard Shi |
ASP-DAC | 2 |
| 2000 | Symbolic circuit-noise analysis and modeling with determinant decision diagramsabstractAbstract | In this paper, a new symbolic noise analysis and modeling technique is presented. The new method exploits the sharing of symbolic expressions in the noise models by using a recently introduced graph, called determinant decision diagrams (DDDs), for symbolic determinant representations. With efcient DDD-based graph manipulations, we are able to generate the exact noise models for analog blocks. Symbolic noise analysis and modeling on real analog circuit examples are presented and compared with SPICE noise simulation. 1. Sheldon X.-D. Tan, Chuanjin Richard Shi |
ASP-DAC | 2 |
| 2000 | Multi-terminal determinant decision diagrams: a new approach to semi-symbolic analysis of analog integrated circuitsabstractA graph representation called Multi-Terminal Determinant Decision Diagrams (MTDDD's) is proposed for the semi-symbolic transfer functions of analog integrated circuits. With multiple numeric terminals (instead of only terminal 1 and terminal 0 in Determinant Decision Diagrams -- DDD's), MTDDD's can describe naturally numeric coefficients that arise from semi-symbolic analysis, where some circuit parameters are considered symbols and the others are given as numeric values. Similar to DDD's, MTDDD's are capable of representing a huge number of symbolic/semi-symbolic product terms in a compact manner. An efficient DDD-based algorithm is described to construct MTDDD's. Experimental results have demonstrated that semi-symbolic transfer functions for practical analog circuits like µA741 can be generated in less than a minute on a Pentium-II 450MHz PC. Tao Pi, Chuanjin Richard Shi |
DAC | 2 |
| 2000 | Layout Compaction for Yield Optimization via Critical Area MinimizationabstractThis paper presents a new compaction algorithm to improve the yield of IC layout. The yield is improved by reducing the area where the faults are more likely to happen known as critical area. Instead of assuming that the critical area could probably be present everywhere in the layout, the algorithm first finds where this area can actually exist, and then attempts to minimize it. The algorithm takes benefit from a fast multi-layer critical area computation to extract the rectangles that compose it. Afterwards, the extracted rectangles are involved into the layer minimization process which is the second phase of the compaction procedure to minimize their area. A new formulation of the layer minimization problem is used in such a way that the critical area minimization adds neither extra variables nor extra constraints to the original compaction algorithm. The algorithm has been tested on actual layouts. Youcef Bourai, Chuanjin Richard Shi |
DATE | 2 |
| 2000 | Parallel and Distributed VHDL SimulationabstractThis paper presents a methodology for parallel and distributed simulation of VHDL using the PDES (parallel discrete-event simulation) paradigm. To achieve better features and performance, some PDES protocols assume that simultaneous events may be processed in arbitrary order. We describe a solution of how to apply these algorithms to have a correct simulation of the distributed VHDL cycle, including the delta cycle. The solution is based on tie-breaking the simultaneous events using Lamport's logical clocks to causally order them according to the VHDL simulation cycle, and defining the VHDL virtual time as a pair of simulation physical time and cycle/phase logical time. The paper also shows how to use this method with a PDES protocol that relaxes the simulation of simultaneous events to arbitrary order; allowing the LPs to self-adapt to optimistic or conservative mode, without the lookahead requirement. The lookahead is application-dependent and for some systems may be zero or unknown. The parallel simulation of VHDL designs ranging from 5531 to 14704 LPs using these methods obtained a promising, almost linear speedup. Dragos Lungeanu, Chuanjin Richard Shi |
DATE | 2 |
| 2000 | Canonical symbolic analysis of large analog circuits withdeterminant decision diagramsabstractSymbolic analysis has many applications in the design of analog circuits. Existing approaches rely on two forms of symbolic-expression representation: expanded sum-of-product form and arbitrarily nested form. Expanded form suffers the problem that the number of product terms grows exponentially with the size of a circuit. Nested form is neither canonical nor amenable to symbolic manipulation. In this paper, we present a new approach to exact and canonical symbolic analysis by exploiting the sparsity and sharing of product terms. It consists of representing the symbolic determinant of a circuit matrix by a graph-called a determinant decision diagram (DDD)-and performing symbolic analysis by graph manipulations. We show that DDD construction, as well as many symbolic analysis algorithms, takes time almost linear in the number of DDD vertices. We describe an efficient DDD-vertex ordering heuristic and prove that it is optimum for ladder-structured circuits. For practical analog circuits, the numbers of DDD vertices are several orders of magnitude less than the numbers of product terms. The algorithms have been implemented and compared respectively to symbolic analyzers ISAAC and Maple-V in generating the expanded sum-of-product expressions, and SCAPP in generating the nested sequences of expressions. Chuanjin Richard Shi, Sheldon X.-D. Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2000 | Hierarchical symbolic analysis of analog integrated circuits viadeterminant decision diagramsabstractA new method is proposed for hierarchical symbolic analysis of large analog integrated circuits. It consists of performing symbolic suppression of each subcircuit to its terminals in terms of subcircuit matrix determinants and cofactors, and applying Cramer's rule to symbolically solve the set of equations at the top level of the circuit hierarchy. An annotated, directed, and acyclic graph, called determinant decision diagram (DDD), is used to represent symbolic determinants of subcircuit matrices and cofactors used in subcircuit suppression, as well as symbolic determinants of the top-level circuit matrix and cofactors required in applying Cramer's rule. DDD enables us to systematically exploit the inherent sparsity of circuit matrices and the sharing of symbolic expressions. It is capable of representing a huge number of symbolic product terms in a canonical and highly compact manner. The proposed method is illustrated using a Cauer parameter low-pass filter. It has been implemented in a symbolic analyzer and compared to best-known hierarchical symbolic analyzer SCAPP and numerical simulator SPICE. Experimental results on several analog circuits including the /spl mu/A741 operational amplifier - a circuit with less structural regularities - are described. Sheldon X.-D. Tan, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Symmetry Detection for Automatic Analog-Layout RecyclingabstractLayout symmetry is used to minimize the impact of mismatch on the performance of analog circuits. In this paper, an efficient algorithm is presented to detect automatically the mask layout symmetry. It consists of identifying signal nets, isolating circuit devices and detecting their symmetry, and finally, synthesizing the layout symmetry. Combined with layout compaction with symmetry constraints, this technique provides a methodology for automatic analog-layout recycling. Youcef Bourai, Chuanjin Richard Shi |
ASP-DAC | 2 |
| 1999 | Balanced Multi-Level Multi-Way Partitioning of Large Analog Circuits for Hierarchical Symbolic AnalysisabstractSymbolic analysis of analog circuits is important in analog design automation. However, it is limited to the analysis of small analog circuits where exact symbolic expressions are required. In this paper, we present an efficient algorithm for partitioning large general analog circuits into smaller subcircuits so that symbolic analysis can be performed hierarchically. Experimental results have demonstrated that our method outperforms the best partitioning-based symbolic analyzer SCAPP. Sheldon X.-D. Tan, Chuanjin Richard Shi |
ASP-DAC | 2 |
| 1999 | Reliability-Constrained Area Optimization of VLSI Power/Ground Networks via Sequence of Linear ProgrammingsabstractThis paper presents a new method for determining the widths of the power and ground routes in integrated circuits so that the area required by the routes is minimized subject to the reliability constraints. The basic idea is to transform the resulting constrained nonlinear programming problem into a sequence of linear programs. Theoretically, we show that the sequence of linear programs always converges to the optimum solution of the relaxed convex problem. Experimental results demonstrate that the sequence-of-linear-programming method is orders of magnitude faster than the best-known method based on conjugate gradients, with constantly better optimization solutions. Sheldon X.-D. Tan, Chuanjin Richard Shi, Dragos Lungeanu, Jyh-Chwen Lee, Li-Pen Yuan |
DAC | 2 |
| 1999 | Interpretable Symbolic Small-Signal Characterization of Large Analog Circuits using Determinant Decision DiagramsabstractA new approach is proposed to generate interpretable symbolic expressions of small-signal characteristics for large analog circuits. The approach is based on a complete, exact, yet compact representation of symbolic expressions via determinant decision diagrams (DDDs). We show that two key tasks of generating interpretable symbolic expressions-term de-cancellation and term simplification-can be performed in linear time in terms of the number of DDD vertices. With the number of DDD vertices many-orders-of-magnitude less than the number of product terms, the proposed approach has been shown to be much more efficient than other start-of-the-art approaches. Sheldon X.-D. Tan, Chuanjin Richard Shi |
DATE | 2 |
| 1999 | Distributed simulation of VLSI systems via lookahead-free self-adaptive optimistic and conservative synchronizationabstractThe paper presents a novel protocol for parallel and distributed simulation of VLSI systems. It is novel in two aspects: first, it combines optimistic and conservative synchronization methods, allowing processes to self-adapt for maximal utilization of concurrency. Second, it does not require any application-dependent information like lookahead, which in many cases is unknown, zero, or difficult to automatically obtain from a design in a hardware description language. All these features make it very convenient and practical, extending the class of applications to at least all VHDL circuits, including delta cycle. The proposed protocol has been implemented and used for VHDL simulation. Experimental results on several large VHDL circuits (between 1411 and 14704 processes) have shown promising linear speedups. We also observed that the dynamic synchronization, in which processes automatically adapt to optimistic or conservative behavior, follows closely or finds a very good configuration. This protocol may have a strong impact for mixed-signal circuit simulation, where digital parts may be optimistic and heavy-state analog parts, conservative. Dragos Lungeanu, Chuanjin Richard Shi |
ICCAD | 2 |
| 1999 | A Characterization of Signed Hypergraphs and Its Applications to VLSI Via Minimization and Logic Synthesis
Chuanjin Richard Shi, Janusz A. Brzozowski |
Discret. Appl. Math. | 1 |
| 1999 | Simulation and sensitivity of linear analog circuits under parameter variations by Robust interval analysisabstractAn interval-mathematic approach is presented for frequency-domain simulation and sensitivity analysis of linear analog circuits under parameter variations. With uncertain parameters represented as intervals, bounding frequency-domain responses is formulated as the problem of solving systems of linear interval equations. The formulation is based on a variant of modified nodal analysis, and is particularly amenable to interval analysis. Some characterization of the solution sets of systems of linear interval equations are derived. With these characterizations, an elegant and efficient algorithm is proposed to solve systems of linear interval equations. While the widely used Monte Carlo approach requires many circuit simulations to achieve even moderate accuracy, the computational cost of the proposed approach is about twice that of one circuit simulation. The computed response bounds contain provably, or are usually very close to, the actual response bounds. Further, sensitivity under parameter variations can be computed from the response bounds at minor computational cost. The algorithms are implemented in SPICE3F5, using sparse-matrix techniques and tested on several practical analog circuits. Chuanjin Richard Shi, Michael W. Tian |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 1998 | Mixed-Signal Hardware Description Languages in the Era of System-on-Silicon: Challenges and Opportunities (Abstract of Embedded Tutorial)abstractSPICE-based simulation is recognized as a vital tool in shortening time-to-market, reducing product cost, and improving system reliabi lity. It continues to have a profound impact on today's electronic industry. Behavioral modeling and simulation with Analog and Mixed-Signal Hardware Description Languages (AHDLs) are becoming critical to address the challenge of designing today's increasingly complicated electronic systems. These systems may have millions of transistors, are radically diverse (e.g., micro-electro-me chanical) and governed by tightly coupled physical (e.g. thermal-electron ic and deep-submicron) effects. The trend towards behavioral modeling and simulation has been exemplified by the success of proprietary modeling languages and simulators and by the emergence of several language standards (VHDL-AMS and Verilog-AMS) and their application from full system verification to design reuse, virtual prototyping, testing, and even synthesis. In this tutorial, we will overview some of recent developments in AHDLs and their applications, with particular emphasis on challenges and opportunities. Chuanjin Richard Shi |
ASP-DAC | 1 |
| 1998 | Automatic Test Generation for Linear Analog Circuits under Parameter Variations
Chuanjin Richard Shi, Michael W. Tian |
ASP-DAC | 1 |
| 1998 | Efficient DC Fault Simulation of Nonlinear Analog CircuitsabstractThis paper describes a method to improve the efficiency of nonlinear DC fault simulation. The method uses the Newton-Raphson algorithm to simulate each faulty circuit. The key idea is to order the given list of faults in such a way that the solution of previous faulty circuit can serve as a good initial point for the simulation of the next faulty circuit. To build a good ordering, one step Newton-Raphson iteration is performed for all the faulty circuits once, and the results are used to quantify how faulty circuits and the good circuit are close in their behaviors. With one-step Newton-Raphson iteration implemented by Householder's formula, the proposed method has virtually no overhead. Experimental results on a set of 36 MCNC benchmark circuits show an average speedup of 4.4 and as high as 15 over traditional stand-alone fault simulation. Michael W. Tian, Chuanjin Richard Shi |
DATE | 2 |
| 1998 | Nonlinear Analog DC Fault Simulation by One-Step RelaxationabstractEfficient methods have been developed for fault simulation of linear analog circuits. However, DC fault simulation of nonlinear analog circuits-a more practically-relevant problem-remains largely unexplored. In this paper, we propose an one-step relaxation approach to nonlinear DC fault simulation. In this approach, only one Newton-Raphson iteration is performed for the faulty circuit with the DC solution of the good circuit as the initial point, and the results are used to approximate the actual results of exact fault simulation. With one-step relaxation implemented using Householder's formula, the proposed approach is numerically stable, and computationally efficient. It has a very simple circuit interpretation: the nonlinear circuit under test is modeled by a linearized circuit at its operating point, and faults are modeled as faults in the linearized circuit. Experiment results have demonstrated that the proposed approach achieves almost the same fault coverage as exact fault simulation for 29 MCNC Circuit Simulation Workshop benchmark circuits. Michael W. Tian, Chuanjin Richard Shi |
VTS | 2 |
| 1998 | Behavioral Level Noise Modeling and Jitter Simulation of Phase-Locked Loops with Faults Using VHDL-AMS
Nihal J. Godambe, Chuanjin Richard Shi |
J. Electron. Test. | 2 |
| 1998 | Cluster-cover a theoretical framework for a class of VLSI-CAD optimization problemsabstractThis article introduces a mathematical framework called cluster-cover. We show that this framework captures the combinatorial structure of a class of VLSI design optimization problems, including two-level logic minimization, constrained encoding, multilayer topological planar routing, application timing assignment for delay-fault testing, and minimization of monitoring logic for BIST enchancement. These apparently unrelated problems can all be cast into two metaproblems in our framework: finding a maximum cluster and finding a minimum cover. We describe paradigms for developing algorithms for these problems. First, a simple heuristic called greedy peeling is presented and characterized. We derive sufficient conditions that guarantee optimum solutions by greedy peeling. We generalize the performance analysis of a multilayer topological planar routing heuristic to greedy peeling for the general cluster-cover problems. We propose a performance bound of greedy set covering that can be computed efficiently for a given problem instance; this bound is much tighter than the previously known bounds. Second, prime covering—orignally developed for logic minimization—is generalized to finding exact solutions for cluster-cover problems. Previously, only the connection between logic minimizaton and constrained encoding was known. Chuanjin Richard Shi, Janusz A. Brzozowski |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 1997 | Block-level fault isolation using partition theory and logic minimization techniquesabstractMultichip modules are emerging as a key packaging technology for mixed-signal circuits and systems. In this paper, we consider how to localize a failure within a chip boundary as rapidly as possible in order to expedite the rework process and to minimize its overall impact on manufacturing throughput and cycle time. A key contribution of this paper is to provide a unified block-level fault isolation framework for analog and digital circuits, and to show that optimum fault isolation reduces to set covering. This allows us to apply directly powerful set covering techniques and solvers developed recently in logic minimization. In addition, we present a greedy peeling heuristic with performance bound computation. Some preliminary experimental results are included to demonstrate the feasibility and performance of the proposed approach. Chuanjin Richard Shi |
ASP-DAC | 1 |
| 1997 | Solving constrained via minimization by compact linear programmingabstractVia minimization is an important problem in integrated circuit layout and printed circuit board design. A linear (non-integral) programming approach to two-layer constrained via minimization (CVM) is presented. The approach finds optimum solutions for routings containing no more than three way splits, and guarantees provably good results for the general case. Most importantly, the size of linear programming formulation is polynomial in terms of the size of the CVM problem. The significance of the work lies in three aspects. First, since linear programming can be solved in polynomial time, the work thus provides, for the first time, a mathematical programming solution with computational efficiency comparable to known combinatorial CVM algorithms. Second, the compact linear programming approach is provably good and natural for general CVM, while previous restricted CVM algorithms are difficult to he extended to the general case. Third, the approach can handle additional constraints in a unified manner, and thus provides an efficient method for performance-driven layer assignment. The approach is based on some new graph-theoretic and polyhedron-combinatorial results presented on the structure of the CVM problem. Chuanjin Richard Shi |
ASP-DAC | 1 |
| 1997 | Rapid Frequency-Domain Analog Fault Simulation Under Parameter TolerancesabstractFault-driven analog and mixed-signal testing calls for rapid fault simulationtechniques. A problem that has not been addressed effectivelyby existing research is that circuit parameters have tolerance ranges.In this paper, we propose representing parameters under variations asintervals, and present an efficient algorithm - based on interval analysisand Householder's formula - to compute the worst-case responsebounds of good and faulty linear(ized) circuits under parameter variations.Our approach takes CPU time comparable to one nominal circuitsimulation, and always produces correct and conservative results. Thealgorithm has been implemented into SPICE3F5. Experimental resultsshow an acceptable accuracy. Michael W. Tian, Chuanjin Richard Shi |
DAC | 2 |
| 1997 | Symbolic analysis of large analog circuits with determinant decision diagramsabstractSymbolic analog-circuit analysis has many applications, and is especially useful for analog synthesis and testability analysis. We present a new approach to exact and canonical symbolic analysis by exploiting the sparsity and sharing of product terms. It consists of representing the symbolic determinant of a circuit matrix by a graph-called determinant decision diagram (DDD)-and performing symbolic analysis by graph manipulations. We showed that DDD construction and DDD-based symbolic analysis can be performed in time complexity proportional to the number of DDD vertices. We described a vertex ordering heuristic, and showed that the number of DDD vertices can be quite small-usually orders-of-magnitude less than the number of product terms. The algorithm has been implemented. An order-of-magnitude improvement in both CPU time and memory usage over existing symbolic analyzers ISAAC and Maple-V has been observed for large analog circuits. Chuanjin Richard Shi, Sheldon X.-D. Tan |
ICCAD | 1 |
| 1997 | Behavioral level noise modeling and jitter simulation of phase-locked loops with faults using VHDL-AMSabstractIt is important to predict noise at the early stages of a top down design. In this paper, we propose a methodology to model phase noise or jitter, a key specification for phase-locked loops, using a mixed-signal hardware description language, and to simulate the effects of catastrophic faults on the phase jitter at the behavioral level. In contrast to existing approaches which either require dedicated noise simulators or postpone noise and fault simulation to the transistor level, we have successfully demonstrated that noise in a voltage-controlled oscillator, power supply noise, and their effects on the overall phase jitter within a faulty phase locked loop can be modeled and simulated earlier on at the behavioral level. Our simulation results are consistent with experimentally verified, theoretical predictions. Nihal J. Godambe, Chuanjin Richard Shi |
VTS | 2 |
| 1996 | Exact Dichotomy-based Constrained EncodiabstractConstrained encoding has several applications in the synthesis of finite state machines (FSMs), e.g., it can be used to generate asynchronous FSM state assignment that guarantees a critical hazard-free implementation, or to generate synchronous FSM state assignment with minimum PLA implementation. This paper presents ZEDICHO, an original zero-suppressed binary decision diagram (ZBDD) based algorithm that solves exactly the dichotomy-based constrained encoding problem. Olivier Coudert, Chuanjin Richard Shi |
ICCD | 2 |
| 1995 | A framework for the analysis and design of algorithms for a class of VLSI-CAD optimization problemsabstractNo abstract available. Chuanjin Richard Shi, Janusz A. Brzozowski |
ASP-DAC | 1 |
| 1992 | A signed hypergraph model of constrained via minimizationabstractThe author proposes a use of the notion of hypergraphs to describe the general constrained via minimization (CVM) problem. He shows that the formulation of the general CVM by means of hypergraphs turns out to be surprisingly simple and general. In the case of two-layer routing, a signed hypergraph model is introduced. On the basis of this model, the author develops a fast (linear-time) heuristic and obtains promising results; he also presents two methods of modeling multiway splits by graphs, producing better results than all the previous methods.> Chuanjin Richard Shi |
Great Lakes Symposium on VLSI | 1 |
| 1991 | Group delay as an estimate of delay in logicabstractIt is an accepted practice in signal delay estimation to model MOS digital circuits as RC circuits. In most cases Elmore's definition is exactly equivalent to the group delay of the network at zero frequency. A computationally efficient noniterative method to calculate this delay for networks with any linear elements and arbitrary topology is presented. It is shown that in RC networks under certain conditions, the Elmore delay and the 50% unit step response delay are related by a constant which is largely independent of the element values and topology. An efficient method to obtain sensitivities of the delay with respect to any element in the network is presented.> Jirí Vlach, James A. Barby, Anthony Vannelli, T. Talkhan, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |