EDBT 2026 Demo / reviewers in the wild / expert
Hao Yu 0001
dblp:64/4832-1
· DBLP profile ↗
132ranked-venue papers
15as first author
39since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 121 · 15 first-author · 33 since 2021Software engineering, systems software and programming languages · 17 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IR Drop-Aware ECO: A Fast Approach to Minimize Layout and Timing DisturbanceabstractEnsuring power integrity in advanced IC design is increasingly challenging, as excessive IR drop can severely impact circuit performance and reliability, especially during the late-stage Engineering Change Order (ECO) process. In this work, we propose a novel IR drop-aware ECO framework that addresses IR drop violations through targeted cell displacement while minimizing timing and layout disruption. Our approach incorporates vertical IR drop mitigation and horizontal timing fix, and employs a rail severity scoring mechanism that combines current correlation and spatial proximity to evaluate IR drop severity. Experimental results on three post-routed benchmark designs demonstrate that our method achieves significant reductions in worst-case dynamic voltage drop for certain designs and mitigates local timing degradation. Additionally, the proposed severity score accurately reflects trends in IR drop risk, providing valuable guidance for ECO optimization. Jingchao Hu, Yibo Lin, Hao Yu 0001, Quan Chen 0007, Zhou Jin 0001, Cheng Zhuo |
ASP-DAC | 3 |
| 2026 | A 442.42 TOPS/W RRAM-based Digital Computing-in-Memory Accelerator for BF16×1-bit Vision Transformer
Jingyun Gu, Jiaqi Yang 0009, Jingyao Dong, Qilong Chen, Dingbang Liu, Wei Mao 0002, Hao Yu 0001 |
ISCAS | 10 |
| 2026 | Frieren: A Fault-Tolerant Reconfigurable Energy-Efficient Computing Architecture With Enhanced Reliability in Harsh EnvironmentsabstractIn harsh environments such as space, strong radiation effects often induce single-event effects that threaten the reliability of computing systems. Meanwhile, edge artificial intelligence (AI) processors deployed in these conditions must not only tolerate faults but also operate under stringent resource constraints, while still ensuring efficient task execution. Achieving high-performance and energy-efficient computation with adaptive reliability in such harsh conditions is therefore of great importance. This work presents Frieren, a fault-tolerant and reconfigurable computing architecture for reliable operation in harsh environments. A 22 nm system-on-chip (SoC) prototype is implemented to validate Frieren and evaluate its resilience to soft errors. Frieren operates in three primary modes: (1) a high-throughput computation engine mode, (2) a multi-core mode featuring adaptive dual-core lockstep (DCLS) for fault tolerance and programmable parallel computing, and (3) a JTAG-assisted scan-chain-based fault injection (FI) mode. The first two modes fully share processing elements and memory resources, ensuring zero data movement during mode transitions, while the third mode supports pre-deployment reliability evaluation by emulating transient faults. Both irradiation and hardware-level FI experiments are conducted to verify reliability, confirming the robustness of Frieren. Radiation tests of the SoC indicate that DCLS can correct up to about 83% of RISC-V errors, while customized parallel computing in multi-core mode achieves a 17.77× latency reduction. Moreover, the SoC delivers up to 17.18 TOPS/W in computation engine mode and 1.92 TOPS/W in multi-core mode, demonstrating an energy-efficient and resilient platform for AI deployment under harsh conditions. In real workloads, the SoC achieves peak energy efficiencies of 14.72 TOPS/W on SuperYOLO and 12.33 TOPS/W on DROID-SLAM. Qiufeng Li, Weirong Dong, Mingqiang Huang, Hao Yu 0001, Yiyu Shi 0001, Hiromitsu Awano, Takashi Sato 0001, Mehdi Saligane, Longyang Lin, Masanori Hashimoto |
IEEE Trans. Computers | 7 |
| 2025 | An MLA-LLM Hardware Acceleration with Tensor-Train Decomposition on Group Vector Systolic AcceleratorabstractLarge language models (LLMs) are both storageintensive and computation-intensive, posing significant challenges when deployed on resource-constrained hardware. As linear layers in LLMs are mainly resource consuming parts and KV cache affects the deployment of LLMs, this paper develops a posttraining process for converting multi-head attention (MHA) into multi-head latent attention (MLA) and a tensor-train decomposition (TTD) for LLMs with a further hardware implementation on FPGA. The post-training MLA process reduces the KV cache by 25 %. And TTD compression is applied to the linear layers in LLaMA2-7B models with compression ratios (CRs) for the whole network of$2.45 \times$. The compressed LLMs are further implemented on FPGA hardware within a highly efficient group vector systolic array (GVSA) architecture, which has DSP-shared parallel vector PEs for TTD inference, as well as optimized data communication in this accelerator. Experimental results show that the corresponding TTD based LLM accelerator implemented on FPGA achieves a peak speed of 64.20 tokens/s and$1.54 \times$reduction in first token delay for LLaMA2-7B models. Compared with NVIDIA A100 GPU, it achieves 46 % higher throughput. Sixiao Huang, Tintin Wang, Keyao Jiang, Kai Li 0024, Mingqiang Huang, Hao Yu 0001 |
ASAP | 8 |
| 2025 | A Layer-wised Mixed-Precision CIM Accelerator with Bit-level Sparsity-aware ADCs for NAS-Optimized CNNsabstractExploring multiple precisions as well as sparsities for a computingin-memory (CIM) based convolutional accelerators is challenging. To further improve energy efficiency with minimal accuracy loss, this paper develops a neural architecture search (NAS) method to identify precision for each layer of the CNN and further leverages bit-level sparsity. The results indicate that following this approach, ResNet-18 and VGG-16 not only maintain their accuracy but also implement layer-wised mixed-precision effectively. Furthermore, there is a substantial enhancement in the bit-level sparsity of weights within each layer, with an average bit-level sparsity exceeding 90% per bit, thus providing broader possibilities for hardware-level sparsity optimization. In terms of hardware design, a mixed-precision (2/4/8-bit) readout circuit as well as a bit-level sparsity-aware Analog-to-Digital Converter (ADC) are both proposed to reduce system power consumption. Based on bit-level sparsity mixed-precision CNNs benchmarks, post-layout simulation results in 28nm reveal that the proposed accelerator achieves up to 245.72 TOPS/W energy efficiency, which shows about 2.52 -- 6.57× improvement compared to the state-of-the-art SRAM-based CIM accelerators. Haoxiang Zhou, Zikun Wei, Dingbang Liu, Liuyang Zhang, Chenchen Ding, Jiaqi Yang 0009, Wei Mao 0002, Hao Yu 0001 |
ASP-DAC | 8 |
| 2025 | Fast Routing Algorithm for Mask Stitching Region of Ultra Large Wafer Scale IntegrationabstractInterposer-based packaging has gained tremendous popularity in integrating advanced logic and memory chiplets for artificial intelligence and high-performance computing systems. The size of the silicon interposer is the critical bottleneck in improving the performance of integrated systems by mounting more and more advanced chiplets, such as high bandwidth memory (HBM). Nowadays, ultra large wafer scale integration is a popular alternative to integrate large amounts of advanced chiplets on a big wafer scale silicon interposer. However, wafer scale silicon interposers cannot be manufactured by one mask due to the reticle limitation. Therefore, the mask stitching technique is used to manufacture ultra large systems by applying multiple masks for different sub-regions of an ultra large silicon interposer. To achieve the alignment of two adjacent sub-regions manufactured by different masks, the two sub-regions have an overlapped stitching region. Previous algorithms cannot handle the special design rules of mask stitching regions and are not efficient enough to generate high-quality routing solutions. In this work, we propose a fast routing algorithm for mask stitching regions to efficiently solve the special design rules. The time complexity of the proposed algorithm is O(n log n), where n is the number of nets. Compared with state-of-the-art work, our algorithm can achieve 100% routability with an effective reduction of wirelength. Furthermore, the proposed algorithm can achieve a speedup of thousands of times. Zhen Zhuang, Quan Chen 0007, Hao Yu 0001, Tsung-Yi Ho |
ASP-DAC | 3 |
| 2025 | HachiFI: A Lightweight SoC Architecture-Independent Fault-Injection Framework for SEU Impact EvaluationabstractSingle-Event Upsets (SEUs), triggered by energetic particles, manifest as unexpected bit-flips in memory cells or registers, potentially causing significant anomalies in electronic devices. Driven by the needs of safety-critical applications, it is crucial to evaluate the reliability of these electronic devices before they are deployed. However, traditional reliability analysis techniques, such as irradiation experiments, are costly, while fault injection (FI) simulations often fail to provide full coverage and have limited effectiveness and accuracy. To address these issues, we introduce HachiFI, a lightweight, architecture-independent framework that automates fault injection with 100% coverage via memory and scan-chain accesses and simulates the behavior of SEUs based on specific cross-sections. HachiFI supports configurable fault injection patterns for both system-level and module-level reliability analysis. Using HachiFI, we demonstrate a low hardware overhead (2=0.984) between FI and irradiation experiments, verified on a 22nm edge-AI chip. Wang Liao 0001, Hao Yu 0001, Longyang Lin, Masanori Hashimoto |
DATE | 4 |
| 2025 | LLM-Barber: Block-Aware Rebuilder for Sparsity Mask in One-Shot for Large Language ModelsabstractLarge language models (LLMs) have seen substantial growth, necessitating efficient model pruning techniques. Existing post-training pruning methods primarily measure weight importance in converged dense models, often overlooking changes in weight significance during the pruning process, leading to performance degradation. To address this issue, we present LLM-Barber (Block-Aware Rebuilder for Sparsity Mask in One-Shot), a novel one-shot pruning framework that rebuilds the sparsity mask of pruned models without any retraining or weight reconstruction. LLM-Barber incorporates block-aware error optimization across Self-Attention and MLP blocks, facilitating global performance optimization. We are the first to employ the product of weights and gradients as a pruning metric in the context of LLM post-training pruning. This enables accurate identification of weight importance in massive models and significantly reduces computational complexity compared to methods using second-order information. Our experiments show that LLM-Barber efficiently prunes models from LLaMA and OPT families (7B to 13B) on a single A100 GPU in just 30 minutes, achieving state-of-the-art results in both perplexity and zero-shot performance across various language benchmarks. Yupeng Su, Xiaoqun Liu, Tianlai Jin, Dongkuan Wu, Zhengfei Chen, Graziano Chesi, Ngai Wong 0001, Hao Yu 0001 |
ICCAD | 9 |
| 2025 | A 20.98TOPS/W Energy-Efficient Binary BERT Model on Group Vector Systolic CIM AcceleratorabstractTransformer-based large language models (LLMs) impose significant bandwidth and compute challenges when deployed on edge devices. SRAM-based compute-in-memory (CIM) accelerators offer a promising solution to reduce data movement but are still limited by model size. This work develops a ternary weight splitting (TWS) binarization to obtain Brain-Floating-Point-16×INT1 (BF16×1-b) and INT8×INT1 (8-b×1-b) based transformers that exhibit competitive accuracy while significantly reducing model size compared to full precision counterparts. Then, a fully digital SRAM-based CIM accelerator is designed incorporating a bit-parallel SRAM macro within a highly efficient group vector systolic architecture, which can store one column of BERT-Tiny model with stationary systolic data reuse. The design in a 28nm technology only requires 2KB SRAM with an area of 2mm2. It achieves a throughput of 6.55TOPS and consumes a total power of 312.5mW and 221mW at 400MHz, resulting in a state-of-the-art area efficiency of 3.3TOPS/mm2and normalized energy efficiency of 20.98TOPS/W and 34.35TOPS/W for BF16×1-b and 8-b×1-b respectively on BERT-Tiny model, demonstrating a 10.25× improvement in area efficiency and a 2.23× improvement in energy efficiency compared to other state-of-the-art counterparts. Additionally, our proposed configuration compresses the model size by 32% with only a 0.5% accuracy loss on SST-2. Dingbang Liu, Qilong Chen, Jingyun Gu, Jiaqi Yang 0009, Kai Li 0024, Wei Mao 0002, Ngai Wong 0001, Chang Wen Chen, Hao Yu 0001 |
ISLPED | 10 |
| 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and RegressionabstractEvent cameras draw inspiration from biological systems, boasting low latency and high dynamic range while consuming minimal power. The most current approach to processing Event Cloud often involves converting it into frame-based representations, which neglects the sparsity of events, loses fine-grained temporal information, and increases the computational burden. In contrast, Point Cloud is a popular representation for processing 3-dimensional data and serves as an alternative method to exploit local and global spatial features. Nevertheless, previous point-based methods show an unsatisfactory performance compared to the frame-based method in dealing with spatio-temporal event streams. In order to bridge the gap, we propose EventMamba, an efficient and effective framework based on Point Cloud representation by rethinking the distinction between Event Cloud and Point Cloud, emphasizing vital temporal information. The Event Cloud is subsequently fed into a hierarchical structure with staged modules to process both implicit and explicit temporal features. Specifically, we redesign the global extractor to enhance explicit temporal extraction among a long sequence of events with temporal aggregation and State Space Model (SSM) based Mamba. Our model consumes minimal computational resources in the experiments and still exhibits SOTA point-based performance on six different scales of action recognition datasets. It even outperformed all frame-based methods on both Camera Pose Relocalization (CPR) and eye-tracking regression tasks. Yue Zhou 0010, Jiadong Zhu, Xiaopeng Lin, Haotian Fu, Yulong Huang 0001, Yuetong Fang, Fei Ma 0006, Hao Yu 0001, Bojun Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2025 | EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language ModelsabstractThe rapid advancements in artificial intelligence (AI), particularly the Large Language Models (LLMs), have profoundly affected our daily work and communication forms. However, it is still a challenge to deploy LLMs on resource-constrained edge devices (such as robots), due to the intensive computation requirements, heavy memory access, diverse operator types and difficulties in compilation. In this work, we proposed EdgeLLM to address the above issues. Firstly, focusing on the computation, we designed mix-precision processing element array together with group systolic architecture, that can efficiently support both FP$16\ast $FP16 for the MHA block (Multi-Head Attention) and FP$16\ast $INT4 for the FFN layer (Feed-Forward Network). Meanwhile specific optimization on log-scale structured weight sparsity, has been used to further increase the efficiency. Secondly, to address the compilation and deployment issue, we analyzed the whole operators within LLM models and developed a universal data parallelism scheme, by which all of the input and output features maintain the same data shape, enabling to process different operators without any data rearrangement. Then we proposed an end-to-end compiler to map the whole LLM model on CPU-FPGA heterogeneous system (AMD Xilinx VCU128 FPGA). The accelerator achieves$1.91\times $higher throughput and$7.55\times $higher energy efficiency than the commercial GPU (NVIDIA A100-SXM4-80G). When compared with state-of-the-art FPGA accelerator of FlightLLM, it shows 10-24% better performance in terms of HBM bandwidth utilization, energy efficiency and LLM throughput. Mingqiang Huang, Kai Li 0024, Haoxiang Peng, Yupeng Su, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | A 28-nm 135.19 TOPS/W Bootstrapped-SRAM Compute-in-Memory Accelerator With Layer-Wise Precision and SparsityabstractArtificial intelligence (AI) edge devices demand high energy efficiency as well as inference accuracy. SRAM-based compute-in-memory (CIM) accelerators have great potential for power reduction but still need to exploit higher throughput and better linearity performance. To meet edge-AI computing demands by CIM works, it is crucial to optimize algorithms and parameters for specific circuit systems to achieve hardware acceleration. This work firstly employs neural network search (NAS) method to find out the layer-wise optimized precisions and sparsities for convolutional neural networks (CNNs). Then, a 144-Kb charge-domain signed mixed-precision (2/4/8-bit) CIM accelerator employing bootstrapped SRAM cells with 9-transistors and 1-capacitor (9T1C) structure is proposed that incorporates a bit-level sparsity-aware analog-to-digital converter (ADC). This work not only achieves highly linear parallel accumulation operations to meet AI computing demands but also implements a hardware and software co-optimization system tailored to specific data characteristics. The design is verified on NAS-optimized networks VGG-16 and ResNet-18 using Cifar-10 dataset, which could achieve an equivalent accuracy at 4-bit of 68.68% while maintaining a high energy efficiency at 2-bit of 135.19TOPS/W by measurements. Wei Mao 0002, Dingbang Liu, Haoxiang Zhou, Fuyi Li, Kai Li 0024, Qiuping Wu, Jiaqi Yang 0009, Liuyang Zhang, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2024 | How accurately can soft error impact be estimated in black-box/white-box cases? - a case study with an edge AI SoC -abstractArtificial intelligence (AI) edge devices often feature numerous storage units and sequential logic circuits, making them vulnerable to soft errors. For reliable and critical edge AI applications, assessing System-on-Chip (SoC) reliability in advance is essential. Here, there are two cases: a self-designed SoC (white-box), or a commercial off-the-shelf (COTS) chip (black-box). This study uses alpha particle irradiation results on our 22nm AI SoC as a golden reference to estimate soft error impacts, injecting faults across the entire chip in the white-box case and into the accessible memory and registers in the black-box case. The results demonstrate a high degree of consistency between the white-box case and golden reference, meaning that pre-silicon reliability assessment is feasible. As for the black-box case, the proportion of memory in the SoC remains unchanged and is still significantly larger than that of registers, and hence the simulation results between black-box and white-box are not substantially different. Qiufeng Li, Longyang Lin, Wang Liao 0001, Liuyao Dai, Hao Yu 0001, Masanori Hashimoto |
DAC | 6 |
| 2024 | APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language ModelsabstractLarge Language Models (LLMs) have greatly advanced the natural language processing paradigm. However, the high computational load and huge model sizes pose a grand challenge for deployment on edge devices. To this end, we propose APTQ (Attention-aware Post-Training Mixed-Precision Quantization) for LLMs, which considers not only the second-order information of each layer's weights, but also, for the first time, the nonlinear effect of attention outputs on the entire model. We leverage the Hessian trace as a sensitivity metric for mixed-precision quantization, ensuring an informed precision reduction that retains model performance. Experiments show APTQ surpasses previous quantization methods, achieving an average of 4 bit width a 5.22 perplexity nearly equivalent to full precision in the C4 dataset. In addition, APTQ attains state-of-the-art zero-shot accuracy of 68.24% and 70.48% at an average bitwidth of 3.8 in LLaMa-7B and LLaMa-13B, respectively, demonstrating its effectiveness to produce high-quality quantized LLMs. Hantao Huang, Yupeng Su, Ngai Wong 0001, Hao Yu 0001 |
DAC | 6 |
| 2024 | An Isotropic Shift-Pointwise Network for Crossbar-Efficient Neural Network DesignabstractResistive random-access memory (RRAM), with its programmable and nonvolatile conductance, permits compute-in-memory (CIM) at a much higher energy efficiency than the traditional von Neumann architecture, making it a promising candidate for edge AI. Nonetheless, the fixed-size crossbar tiles on RRAM are inherently unfit for conventional pyramid-shape convolutional neural networks (CNNs) that incur low crossbar utilization. To this end, we recognize the mixed-signal (digital-analog) nature in RRAM circuits and customize an isotropic shift-pointwise network that exploits digital shift operations for efficient spatial mixing and analog pointwise operations for channel mixing. To fast ablate various shift-pointwise topologies, a new recon-figurable energy-efficient shift module is designed and packaged into a seamless mixed-domain simulator. The optimized design achieves a near-100% crossbar utilization, providing a state-of-the-art INT8 accuracy of 94.88% (76.55%) on the CIFAR-10 (CIFAR-100) dataset with 1.6M parameters, which sets a new standard for RRAM-based AI accelerators. Muqun Niu, Hantao Huang, Graziano Chesi, Hao Yu 0001, Ngai Wong 0001 |
DATE | 7 |
| 2024 | FMTT: Fused Multi-Head Transformer with Tensor-Compression for 3D Point Clouds Detection on Edge DevicesabstractThe real-time detection of 3D objects represents a grand challenge on edge devices. Existing 3D point clouds models are over-parameterized with heavy computation load. This paper proposes a highly compact model for 3D point clouds detection using tensor-compression. Compared to conventional methods, we propose a fused multi-head transformer tensor-compression (FMTT) to achieve both compact size yet with high accuracy. The FMTT leverages different ranks to extract both high and low-level features and then fuses them together to improve the accuracy. Experiments on the KITTI dataset show that the proposed FMTT can achieve 6.04× smaller than the uncompressed model from 55.09MB to 9.12MB such that the compressed model can be implemented on edge devices. It also achieves 2.62% improved accuracy in easy mode and 0.28% improved accuracy in hard mode. Zikun Wei, Chenchen Ding, Hantao Huang, Hao Yu 0001 |
DATE | 7 |
| 2024 | LAMPS: A Layer-wised Mixed-Precision-and-Sparsity Accelerator for NAS-Optimized CNNs on FPGAabstractThe increasing model size and computation load of convolutional neural networks (CNN) pose a grand challenge to deploy CNN models on edge computing devices. To further improve performance without significant accuracy loss, this paper developed a neural architecture search (NAS) method to achieve a layer-wise mixed-precision-and-sparsity (LAMPS) CNN. However, this optimization cannot be fully utilized and directly mapped to existing AI accelerators due to the irregu- lar computation of sparse and multi-precision data. To tackle this challenge, this work proposed a LAMPS vector systolic accelerator and demonstrated state-of-the-art results. Experi- mental results show that the LAMPS accelerator on Xilinx ZCU102 achieves an average performance of 756.83 GOPS and 470.25 GOPS when accelerating the NAS-optimized VGG16 and Resnet18, respectively, leading to 1.3-6.0x speed-up over the state- of-the-art accelerators on FPGA. Shuxin Yang, Chenchen Ding, Mingqiang Huang, Kai Li 0024, Chenghao Li 0010, Zikun Wei, Sixiao Huang, Jingyao Dong, Liuyang Zhang, Hao Yu 0001 |
FCCM | 10 |
| 2023 | Agile Hardware and Software Co-Design for RISC-V-Based Multi-Precision Deep Learning MicroprocessorabstractRecent network architecture search (NAS) has been widely applied to simplify deep learning neural networks, which typically result in a multi-precision network. Many multi-precision accelerators have been developed as well to support computing multi-precision networks manually. A software-hardware interface is thereby needed to automatically map multi-precision networks onto multi-precision accelerators. In this paper, we have developed an agile hardware and software co-design for RISC-V-based multi-precision deep learning microprocessor. We have designed custom RISC-V instructions with a framework to automatically compile multi-precision CNN networks onto multi-precision CNN accelerators, demonstrated on FPGA. Experiments show that with NAS optimized multi-precision CNN models (LeNet, VGG16, ResNet, MobileNet), the RISC-V core with multi-precision accelerators can reach the highest throughput in 2,4,8-bit precisions respectively on a Xilinx ZCU102 FPGA. Zicheng He, Qiufeng Li, Hao Yu 0001 |
ASP-DAC | 5 |
| 2023 | RankSearch: An Automatic Rank Search Towards Optimal Tensor Compression for Video LSTM Networks on EdgeabstractVarious industrial and domestic applications call for optimized lightweight video LSTM network models on edge. The recent tensor-train method can transform space-time features into tensors, which can be further decomposed into low-rank network models for lightweight video analysis on edge. The rank selection of tensor is however manually performed with no optimization. This paper formulates a rank search algorithm to automatically decide tensor ranks with consideration of the trade-off between network accuracy and complexity. A fast rank search method, called RankSearch, is developed to find optimized low-rank video LSTM network models on edge. Results from experiments show that RankSearch achieves a$4.84 >$reduction in model complexity, and$1.96\times$speed-up in run time while delivering a 3.86% accuracy improvement compared with the manual-ranked models. Changhai Man, Chenchen Ding, Shaobo Luo, Rumin Zhang, Ngai Wong 0001, Hao Yu 0001 |
DATE | 11 |
| 2023 | Multi-bit-width CNN Accelerator with Systolic-in-Systolic Dataflow and Single DSP Multiple Multiplication SchemeabstractMulti-bit-width neural network enlightens a promising method for high performance yet energy efficient edge computing due to its balance between software algorithm accuracy and hardware efficiency. To date, FPGA has been one of the core hardware platforms for deploying various neural networks. However, it is still difficult to fully make use of the dedicated digital signal processing (DSP) blocks in FPGA for accelerating the multi-bit-width network. In this work, we develop state-of-the-art multi-bit-width convolutional neural network accelerator with novel systolic-in-systolic type of dataflow and single DSP multiple multiplication (SDMM) INT2/4/8 execution scheme. Multi-level optimizations have also been adopted to further improve the performance, including group-vector systolic array for maximizing the circuit efficiency as well as minimizing the systolic delay, and differential neural architecture search (NAS) method for the high accuracy multi-bit-width network generation. The proposed accelerator has been practically deployed on Xilinx ZCU102 with accelerating NAS optimized VGG16 and Resnet18 networks as case studies. Average performance on accelerating the convolutional layer in VGG16 and Resnet18 is 1289GOPs and 1155GOPs, respectively. Throughput for running the full multi-bit-width VGG16 network is 870.73 GOPS at 250MHz, which has exceeded all of previous CNN accelerators on the same platform. Mingqiang Huang, Yucen Liu, Sixiao Huang, Kai Li 0024, Qiuping Wu, Hao Yu 0001 |
FPGA | 6 |
| 2023 | Reliability Exploration of System-on-Chip With Multi-Bit-Width Accelerator for Multi-Precision Deep Neural NetworksabstractDeep neural networks (DNNs) in safety-critical applications demand high reliability even when running on edge-computing devices. Recent works on System-on-Chip (SoC) design with state-of-the-art (SOTA) hardware artificial intelligence (AI) accelerators and corresponding multi-bit-width (MBW) convolutional neural network (CNN) generation strategies show that MBW CNNs can effectively explore the trade-off between network accuracy and hardware efficiency. However, reliability has not been considered in such trade-off analysis, even though highly quantized CNNs may elevate the impact of bit flips in the hardware. Also, the reliability of the microcontroller and its interface operating with the AI accelerator are not studied. This work evaluates the reliability of DNN computation in an SoC that includes a processor, SOTA AI accelerator, and NN models highly optimized for computation efficiency using a neural architecture search (NAS) method. Focusing on neutron-induced soft error, which is the primary source of bit-flip errors in a terrestrial environment, we perform fault injection and neutron beam experiments. For these experiments, we prototype the SoC on a flash-based FPGA platform, in which the configuration memory is robust to neutron irradiation. Then, we analyze the experimental data and identify vulnerable components in the system. Furthermore, we evaluate how the SoC running different NAS-optimized MBW LeNet5 networks impact the performance, radiation sensitivity, failure rate of MBW accelerator, and crash rate of the system on the FPGAs. Our results show that instruction and data tightly coupled memory (I/DTCM) are the most vulnerable parts and the control status registers (CSRs) in our accelerator are the second most vulnerable component. Moreover, MBW networks have higher susceptibility to critical errors than single-precision networks, low-precision data are more likely to affect the classification results, and the high bits are more sensitive to faults. Mingqiang Huang, Changhai Man, Liuyao Dai, Hao Yu 0001, Masanori Hashimoto |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | An Integer-Only and Group-Vector Systolic Accelerator for Efficiently Mapping Vision Transformer on EdgeabstractTransformer-like network has shown remarkable high performance in both natural language processing and computer vision. However, the huge computational demands in non-linear floating-point arithmetic and the irregular memory access requirement in self-attention mechanism make it still a challenge to deploy Transformer on edge. To address the above issues, we propose integer-only quantization scheme for the simplification of non-linear operations (such as LayerNorm, Softmax and Gelu), meanwhile algorithm-hardware co-design strategy is applied to guarantee both the high accuracy and high efficiency. Besides, we construct general-purpose group vector systolic array to efficiently accelerate the matrix multiplication operations including both regular matrix-multiplication/convolution and the irregular multi-head self-attention mechanism. Unified data-package strategy and flexible on-/off-chip data storage management strategy are also proposed to further improve the performance. The design has been deployed on Xilinx ZCU102 FPGA platform, achieving an overall inference latency of 4.077ms and 11.15ms per image for ViT-tiny and ViT-s, respectively. The average throughput can reach as high as 762.7 GOPs, which shows significant improvement over the previous state-of-the-art FPGA Transformer accelerator. Mingqiang Huang, Junyi Luo, Chenchen Ding, Zikun Wei, Sixiao Huang, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | An Energy-Efficient Bit-Split-and-Combination Systolic Accelerator for NAS-Based Multi-Precision Convolution Neural NetworksabstractOptimized convolutional neural network (CNN) models and energy-efficient hardware design are of great importance in edge-computing applications. The neural architecture search (NAS) methods are employed for CNN model optimization with multi-precision networks. To satisfy the computation requirements, multi-precision convolution accelerators are highly desired. The existing high-precision-split (HPS) designs reduce the additional logics for reconfiguration while resulting in low throughput for low precisions. The low-precision-combination (LPC) designs improve the low-precision throughput with large hardware cost. In this work, a bit-split-and-combination (BSC) systolic accelerator is proposed to overcome the bottlenecks. Firstly, BSC-based multiply-accumulate (MAC) unit is designed to support multi-precision computation operations. Secondly, multi-precision systolic dataflow is developed with improved data-reuse and transmission efficiency. The proposed work is designed by Chisel and synthesized in 28-nm process. The BSC MAC unit achieves maximum 2.40× and 1.64× energy efficiency than HPS and LPC units, respectively. Compared with published accelerator designs Gemmini, Bit-fusion and Bit-serial, the proposed accelerator achieves up to 2.94 × area efficiency and 6.38 × energy-saving performance on the multi-precision VGG-16, ResNet-18 and LeNet-5 benchmarks. Liuyao Dai, Gengbin Huang, Junzhuo Zhou, Kai Li 0024, Wei Mao 0002, Hao Yu 0001 |
ASP-DAC | 8 |
| 2022 | A Precision-Scalable Energy-Efficient Bit-Split-and-Combination Vector Systolic Accelerator for NAS-Optimized DNNs on EdgeabstractOptimized model and energy-efficient hardware are both required for deep neural networks (DNNs) in edge-computing area. Neural architecture search (NAS) methods are employed for DNN model optimization with resulted multi-precision networks. Previous works have proposed low-precision-combination (LPC) and high-precision-split (HPS) methods for multi-precision networks, which are not energy-efficient for precision-scalable vector implementation. In this paper, a bit-split-and-combination (BSC) based vector systolic accelerator is developed for a precision-scalable energy-efficient convolution on edge. The maximum energy efficiency of the proposed BSC vector processing element (PE) is up to 1.95× higher in 2-bit, 4-bit and 8-bit operations when compared with LPC and HPS PEs. Further with NAS optimized multi-precision CNN networks, the averaged energy efficiency of the proposed vector systolic BSC PE array achieves up to 2.18× higher in 2-bit, 4-bit and 8-bit operations than that of LPC and HPS PE arrays. Kai Li 0024, Junzhuo Zhou, Junyi Luo, Zhengke Yang, Shuxin Yang, Wei Mao 0002, Mingqiang Huang, Hao Yu 0001 |
DATE | 9 |
| 2022 | A High Throughput Multi-bit-width 3D Systolic Accelerator for NAS Optimized Deep Neural Networks on FPGAabstractNeural architecture search (NAS) optimized multi-bit-width convolutional neural network (CNN) maintains the balance between network performance and efficiency, thus enlightening a promising method for accurate yet energy-efficient edge computing. In this work, we propose a high throughput three-dimensional (3D) systolic accelerator for NAS optimized CNNs, in which the input feature matrix, weight matrix and output feature matrix are delivering vertically, horizontally and perpendicularly through the systolic array respectively. With 3D systolic data flow, the processing time and logic resources consumption can be both reduced compared to the classical non-stationary systolic array. Besides, Booth-based multi-bit-width (INT2/4/8) multiply-add-accumulation (MAC) unit is developed within the 3D systolic accelerator. Deployed on FPGA platform Xilinx ZCU102, peek performance of the convolutional layer can reach as high as 2775 GOPS for INT2, 1650 GOPS for INT4, and 816 GOPS for INT8 respectively. The average performance on accelerating full NAS VGG16 network is 647 GOPS. Mingqiang Huang, Yucen Liu, Shuxin Yang, Kai Li 0024, Junyi Luo, Zhengke Yang, Qiufeng Li, Hao Yu 0001, Changhai Man |
FPGA | 9 |
| 2022 | FASSST: Fast Attention Based Single-Stage Segmentation Net for Real-Time Instance SegmentationabstractReal-time instance segmentation is crucial in various AI applications. This work designs a network named Fast Attention based Single-Stage Segmentation NeT (FASSST) that performs instance segmentation with video-grade speed. Using an instance attention module (IAM), FASSST quickly locates target instances and segments with region of interest (ROI) feature fusion (RFF) aggregating ROI features from pyramid mask layers. The module employs an efficient single-stage feature regression, straight from features to instance coordinates and class probabilities. Experiments on COCO and CityScapes datasets show that FASSST achieves state-of-the-art performance under competitive accuracy: real-time inference of 47.5FPS on a GTX1080Ti GPU and 5.3FPS on a Jetson Xavier NX board with only 71.6 GFLOPs. Peining Zhen, Tianshu Hou, Chiu Wa Ng, Haibao Chen, Hao Yu 0001, Ngai Wong 0001 |
WACV | 7 |
| 2022 | ANT-UNet: Accurate and Noise-Tolerant Segmentation for Pathology Image ProcessingabstractPathology image segmentation is an essential step in early detection and diagnosis for various diseases. Due to its complex nature, precise segmentation is not a trivial task. Recently, deep learning has been proved as an effective option for pathology image processing. However, its efficiency is highly restricted by inconsistent annotation quality. In this article, we propose an accurate and noise-tolerant segmentation approach to overcome the aforementioned issues. This approach consists of two main parts: a preprocessing module for data augmentation and a new neural network architecture, ANT-UNet. Experimental results demonstrate that, even on a noisy dataset, the proposed approach can achieve more accurate segmentation with 6% to 35% accuracy improvement versus other commonly used segmentation methods. In addition, the proposed architecture is hardware friendly, which can reduce the amount of parameters to one-tenth of the original and achieve 1.7× speed-up. Yufei Chen 0007, Tingtao Li, Qinming Zhang, Wei Mao 0002, Nan Guan, Hao Yu 0001, Cheng Zhuo |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2022 | A High Performance Multi-Bit-Width Booth Vector Systolic Accelerator for NAS Optimized Deep Learning Neural NetworksabstractMulti-bit-width convolutional neural network (CNN) maintains the balance between network accuracy and hardware efficiency, thus enlightening a promising method for accurate yet energy-efficient edge computing. In this work, we develop state-of-the-art multi-bit-width accelerator for NAS Optimized deep learning neural networks. To efficiently process the multi-bit-width network inferencing, multi-level optimizations have been proposed. Firstly, differential Neural Architecture Search (NAS) method is adopted for the high accuracy multi-bit-width network generation. Secondly, hybrid Booth based multi-bit-width multiply-add-accumulation (MAC) unit is developed for data processing. Thirdly, vector systolic array is proposed for effectively accelerating the matrix multiplications. With vector-style systolic dataflow, both the processing time and logic resources consumption can be reduced when compared with the classical systolic array. Finally, The proposed multi-bit-width CNN acceleration scheme has been practically deployed on FPGA platform of Xilinx ZCU102. Average performance on accelerating the full NAS optimized VGG16 network is 784.2 GOPS, and peek performance of the convolutional layer can reach as high as 871.26 GOPS for INT8, 1676.96 GOPS for INT4, and 2863.29 GOPS for INT2 respectively, which is among the best results in previous CNN accelerator benchmarks. Mingqiang Huang, Yucen Liu, Changhai Man, Kai Li 0024, Wei Mao 0002, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | A Fall Detection Network by 2D/3D Spatio-temporal Joint Models with Tensor Compression on EdgeabstractFalling is ranked highly among the threats in elderly healthcare, which promotes the development of automatic fall detection systems with extensive concern. With the fast development of the Internet of Things (IoT) and Artificial Intelligence (AI), camera vision-based solutions have drawn much attention for single-frame prediction and video understanding on fall detection in the elderly by using Convolutional Neural Network (CNN) and 3D-CNN, respectively. However, these methods hardly supervise the intermediate features with good accurate and efficient performance on edge devices, which makes the system difficult to be applied in practice. This work introduces a fast and lightweight video fall detection network based on a spatio-temporal joint-point model to overcome these hurdles. Instead of detecting fall motion by the traditional CNNs, we propose a Long Short-Term Memory (LSTM) model based on time-series joint-point features extracted from a pose extractor . We also introduce the increasingly mature RGB-D camera and propose 3D pose estimation network to further improve the accuracy of the system. We propose to apply tensor train decomposition on the model to reduce storage and computational consumption so the deployment on edge devices can to realized. Experiments are conducted to verify the proposed framework. For fall detection task, the proposed video fall detection framework achieves a high sensitivity of 98.46% on Multiple Cameras Fall, 100% on UR Fall, and 98.01% on NTU RGB-D 120. For pose estimation task, our 2D model attains 73.3 mAP in the COCO keypoint challenge, which outperforms the OpenPose by 8%. Our 3D model attains 78.6% mAP on NTU RGB-D dataset with 3.6× faster speed than OpenPose. Shuwei Li, Changhai Man, Wei Mao 0002, Shaobo Luo, Rumin Zhang, Hao Yu 0001 |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2022 | Correlated Multi-objective Multi-fidelity Optimization for HLS Directives DesignabstractHigh-level synthesis (HLS) tools have gained great attention in recent years because it emancipates engineers from the complicated and heavy hardware description language writing and facilitates the implementations of modern applications (e.g., deep learning models) on Field-programmable Gate Array (FPGA) , by using high-level languages and HLS directives. However, finding good HLS directives is challenging, due to the time-consuming design processes, the balances among different design objectives, and the diverse fidelities (accuracies of data) of the performance values between the consecutive FPGA design stages. To find good HLS directives, a novel automatic optimization algorithm is proposed to explore the Pareto designs of the multiple objectives while making full use of the data with different fidelities from different FPGA design stages. Firstly, a non-linear Gaussian process (GP) is proposed to model the relationships among the different FPGA design stages. Secondly, for the first time, the GP model is enhanced as correlated GP (CGP) by considering the correlations between the multiple design objectives, to find better Pareto designs. Furthermore, we extend our model to be a deep version deep CGP (DCGP) by using the deep neural network to improve the kernel functions in Gaussian process models, to improve the characterization capability of the models, and learn better feature representations. We test our design method on some public benchmarks (including general matrix multiplication and sparse matrix-vector multiplication) and deep learning-based object detection model iSmart2 on FPGA. Experimental results show that our methods outperform the baselines significantly and facilitate the deep learning designs on FPGA. Qi Sun 0002, Tinghuan Chen, Siting Liu 0002, Jianli Chen, Hao Yu 0001, Bei Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2022 | An Energy-Efficient Mixed-Bitwidth Systolic Accelerator for NAS-Optimized Deep Neural NetworksabstractOptimized deep neural network (DNN) models and energy-efficient hardware designs are of great importance in edge-computing applications. The neural architecture search (NAS) methods are employed for DNN model optimization with mixed-bitwidth networks. To satisfy the computation requirements, mixed-bitwidth convolution accelerators are highly desired for low-power and high-throughput performance. There exist several methods to support mixed-bitwidth multiply-accumulate (MAC) operations in DNN accelerator designs. The low-bitwidth-combination (LBC) method improves the low-bitwidth throughput with a large hardware cost. The high-bitwidth-split (HBS) method minimizes the additional logic gates for configuration. However, the throughput performance in the low-bitwidth mode is poor. In this work, a bit-split-and-combination (BSC) systolic accelerator is proposed. The BSC-based MAC unit is designed to support mixed-bitwidth operations with the best overall performance. Besides, interprocessing element (PE) systolic and intra-PE paralleled dataflow not only improves throughput performance in mixed-bitwidth modes, but also saves power performance for data transmission. The proposed work is designed and synthesized in a 28-nm process. The BSC MAC unit achieves a maximum $2.08\times $ and $1.75\times $ energy efficiency improvement than the HBS and LBC unit, respectively. Compared with the state-of-the-art accelerators, the proposed work also achieves excellent energy-efficient performance with 20.02, 23.55, and 30.17 TOPS/W on mixed-bitwidth VGG-16, ResNet-18, and LeNet-5 benchmarks at 0.6 V, respectively. Wei Mao 0002, Liuyao Dai, Kai Li 0024, Laimin Du, Shaobo Luo, Mingqiang Huang, Hao Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2022 | A Configurable Floating-Point Multiple-Precision Processing Element for HPC and AI Converged ComputingabstractThere is an emerging need to design configurable accelerators for the high-performance computing (HPC) and artificial intelligence (AI) applications in different precisions. Thus, the floating-point (FP) processing element (PE), which is the key basic unit of the accelerators, is necessary to meet multiple-precision requirements with energy-efficient operations. However, the existing structures by using high-precision-split (HPS) and low-precision-combination (LPC) methods result in low utilization rate of the multiplication array and long multiterm processing period, respectively. In this article, a configurable FP multiple-precision PE design is proposed with the LPC structure. Half precision, single precision, and double precision are supported. The 100% multiplier utilization rate of the multiplication array for all precisions is achieved with improved speed in the comparison and summation process. The proposed design is realized in a 28-nm process with 1.429-GHz clock frequency. Compared with the existing multiple-precision FP methods, the proposed structure achieves 63% and 88% area-saving performance for FP16 and FP32 operations, respectively. The$4\times $and$20\times $maximum throughput rates are obtained when compared with fixed FP32 and FP64 operations. Compared with the previous multiple-precision PEs, the proposed one achieves the best energy-efficiency performance with 975.13 GFLOPS/W. Wei Mao 0002, Kai Li 0024, Liuyao Dai, Xinang Xie, He Li 0008, Longyang Lin, Hao Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2021 | A Video-based Fall Detection Network by Spatio-temporal Joint-point Model on Edge DevicesabstractTripping or falling is among the top threats in elderly healthcare, and the development of automatic fall detection systems are of considerable importance. With the fast development of the Internet of Things (IoT), camera vision-based solutions have drawn much attention in recent years. The traditional fall video analysis on the cloud has significant communication overhead. This work introduces a fast and lightweight video fall detection network based on a spatio-temporal joint-point model to overcome these hurdles. Instead of detecting falling motion by the traditional Convolutional Neural Networks (CNNs), we propose a Long Short-Term Memory (LSTM) model based on time-series joint-point features, extracted from a pose extractor and then filtered from a geometric joint-point filter. Experiments are conducted to verify the proposed framework, which shows a high sensitivity of 98.46% on Multiple Cameras Fall Dataset and 100% on UR Fall Dataset. Furthermore, our model can achieve pose estimation tasks simultaneously, attaining 73.3 mAP in the COCO keypoint challenge dataset, which outperforms the OpenPose work by 8%. Shuwei Li, Changhai Man, Wei Mao 0002, Ngai Wong 0001, Hao Yu 0001 |
DATE | 7 |
| 2021 | A Reconfigurable Multiple-Precision Floating-Point Dot Product Unit for High-Performance ComputingabstractThere is an emerging need to optimize floating-point (FP) dot product units (DPU) for high-performance scientific computing as well as training deep learning models. Due to different precision requirements of applications, a reconfigurable multiple-precision DPU operation can largely reduce the cost of area and power. However, the existing methods could result in redundant bits for unit multipliers, but also leave idle hardware resources for the operations in different precisions. In this paper, a reconfigurable multiple-precision FP DPU design is proposed for high-performance computing (HPC) applications. The FP DPU can be reconfigured as follows. A bit-partitioning method is provided to minimize the redundant bits with a configurable mixed-precision multiplier for three-mode operations: 20 half-precision Dot Product (DP), 5 single-precision DP, and 1 double-precision DP operations. Any of the modes can be executed in two successive clock cycles without idle hardware resources. The proposed design is realized by using the UMC 55-nm process with simulation results. Compared with the existing multiple-precision FP methods, the proposed DPU achieves 88.9% and 35.8% area-saving performance for FP16 and FP32 operations, respectively. Moreover, when using benchmarked HPC applications where multiple precisions can be used, the proposed reconfigurable DPU can accelerate up to 4× and 20× maximum throughput rates when compared with fixed FP32 and FP64 operations, respectively. Wei Mao 0002, Kai Li 0024, Xinang Xie, Shirui Zhao, He Li 0008, Hao Yu 0001 |
DATE | 6 |
| 2021 | Correlated Multi-objective Multi-fidelity Optimization for HLS Directives DesignabstractHigh-level synthesis (HLS) tools have gained great attention in recent years because it emancipates engineers from the complicated and heavy hardware description language writing, by using high-level languages and HLS directives. However, previous works seem powerless, due to the time-consuming design processes, the contradictions among design objectives, and the accuracy difference between the three stages (fidelities). To find good HLS directives, in this paper, a novel correlated multi-objective non-linear optimization algorithm is proposed to explore the Pareto solutions while making full use of data from different fidelities. A non-linear Gaussian process is proposed to model relationships among the analysis reports from different fidelities for the same objective. For the first time, correlated multivariate Gaussian process models are introduced into this domain to characterize the complex relationships of multiple objectives in each design fidelity. A tree-based method is proposed to erase invalid solutions and obviously non-optimal solutions. Experimental results show that our non-linear and pioneering correlated models can approximate the Pareto-frontier of the directive design space in a shorter time with much better performance and good stability, compared with the state-of-the-art. Qi Sun 0002, Tinghuan Chen, Siting Liu 0002, Jin Miao, Jianli Chen, Hao Yu 0001, Bei Yu 0001 |
DATE | 6 |
| 2021 | S3-Net: A Fast and Lightweight Video Scene Understanding Network by Single-shot SegmentationabstractReal-time understanding in video is crucial in various AI applications such as autonomous driving. This work presents a fast single-shot segmentation strategy for video scene understanding. The proposed net, called S3-Net, quickly locates and segments target sub-scenes, meanwhile extracts structured time-series semantic features as inputs to an LSTM-based spatio-temporal model. Utilizing ten-sorization and quantization techniques, S3-Net is intended to be lightweight for edge computing. Experiments using CityScapes, UCF11, HMDB51 and MOMENTS datasets demonstrate that the proposed S3-Net achieves an accuracy improvement of 8.1% versus the 3D-CNN based approach on UCF11, a storage reduction of 6.9× and an inference speed of 22.8 FPS on CityScapes with a GTX1080Ti GPU. Haibao Chen, Ngai Wong 0001, Hao Yu 0001 |
WACV | 5 |
| 2021 | TEANS: A Target Enhancement and Attenuated Nonmaximum Suppression Object Detector for Remote Sensing ImagesabstractIn this letter, we propose an effective approach to learn a convolutional neural network (CNN) model with target enhancement and attenuated nonmaximum suppression (NMS) technique (TEANS) for object detection in optical remote sensing images. TEANS mainly consists of two steps. First, the target enhancement architecture, including target upsampling and reconvolution, is designed into a given deep ResNet-101 model for accurate object detection, especially for small ones. Second, the attenuated NMS technique is used for overcoming wrong eliminations of serried object proposals. For verifying the effectiveness of the TEANS method, evaluations are implemented on a publicly available 15-class optical remote sensing object detection data set. Experimental results show that TEANS can achieve 5.55%, 18.77%, 26.81%, 55.07%, 28.48%, 6.01%, and 5.51% improvements in mean Average Precision (mAP), respectively, compared with standard Faster R-CNN, R-FCN, YOLOv2, SSD, USB-BBR, YOLOv3, and MS-VANs frameworks. Haibao Chen, Guanghui He 0002, Bingyi Zhang, Hao Yu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2021 | Fast Video Facial Expression Recognition by a Deeply Tensor-Compressed LSTM Neural Network for Mobile DevicesabstractMobile devices usually suffer from limited computation and storage resources, which seriously hinders them from deep neural network applications. In this article, we introduce a deeply tensor-compressed long short-term memory (LSTM) neural network for fast video-based facial expression recognition on mobile devices. First, a spatio-temporal facial expression recognition LSTM model is built by extracting time-series feature maps from facial clips. The LSTM-based spatio-temporal model is further deeply compressed by means of quantization and tensorization for mobile device implementation. Based on datasets of Extended Cohn-Kanade (CK+), MMI, and Acted Facial Expression in Wild 7.0, experimental results show that the proposed method achieves 97.96%, 97.33%, and 55.60% classification accuracy and significantly compresses the size of network model up to 221× with reduced training time per epoch by 60%. Our work is further implemented on the RK3399Pro mobile device with a Neural Process Engine. The latency of the feature extractor and LSTM predictor can be reduced 30.20× and 6.62× , respectively, on board with the leveraged compression methods. Furthermore, the spatio-temporal model costs only 57.19 MB of DRAM and 5.67W of power when running on the board. Peining Zhen, Haibao Chen, Zhigang Ji, Hao Yu 0001 |
ACM Trans. Internet Things | 6 |
| 2021 | S3-Net: A Fast Scene Understanding Network by Single-Shot Segmentation for Autonomous DrivingabstractReal-time segmentation and understanding of driving scenes are crucial in autonomous driving. Traditional pixel-wise approaches extract scene information by segmenting all pixels in a frame, and hence are inefficient and slow. Proposal-wise approaches only learn from the proposed object candidates, but still require multiple steps on the expensive proposal methods. Instead, this work presents a fast single-shot segmentation strategy for video scene understanding. The proposed net, called S3-Net, quickly locates and segments target sub-scenes , and meanwhile extracts attention-aware time-series sub-scene features ( ats-features ) as inputs to an attention-aware spatio-temporal model (ASM) . Utilizing tensorization and quantization techniques, S3-Net is intended to be lightweight for edge computing. Experiments results on CityScapes, UCF11, HMDB51, and MOMENTS datasets demonstrate that the proposed S3-Net achieves an accuracy improvement of 8.1% versus the 3D-CNN based approach on UCF11, a storage reduction of 6.9× and an inference speed of 22.8 FPS on CityScapes with a GTX1080Ti GPU. Haibao Chen, Ngai Wong 0001, Hao Yu 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2020 | An Anomaly Comprehension Neural Network for Surveillance Videos on Terminal DevicesabstractAnomaly comprehension in surveillance videos is more challenging than detection. This work introduces the design of a lightweight and fast anomaly comprehension neural network. For comprehension, a spatio-temporal LSTM model is developed based on the structured, tensorized time-series features extracted from surveillance videos. Deep compression of network size is achieved by tensorization and quantization for the implementation on terminal devices. Experiments on large-scale video anomaly dataset UCF-Crime demonstrate that the proposed network can achieve an impressive inference speed of 266 FPS on a GTX-1080Ti GPU, which is 4.29 faster than ConvLSTM-based method; a 3.34% AUC improvement with 5.55% accuracy niche versus the 3D-CNN based approach; and at least 15k× parameter reduction and 228× storage compression over the RNN-based approaches. Moreover, the proposed framework has been realized on an ARM-core based IOT board with only 2.4W power consumption. Guangtai Huang, Peining Zhen, Haibao Chen, Ngai Wong 0001, Hao Yu 0001 |
DATE | 7 |
| 2020 | Energy-Efficient Machine Learning Accelerator for Binary Neural NetworksabstractBinary neural network (BNN) has shown great potential to be implemented with power efficiency and high throughput. Compared with its counterpart, the convolutional neural network (CNN), BNN is trained with binary constrained weights and activations, which are more suitable for edge devices with less computing and storage resource requirements. In this paper, we introduce the BNN characteristics, basic operations and the binarized-network optimization methods. Then we summarize several accelerator designs for BNN hardware implementation by using three mainstream structures, i.e., ReRAM-based crossbar, FPGA and ASIC. Based on the BNN characteristics and hardware custom designs, all these methods achieve massively parallelized computations and highly pipelined data flow to enhance its latency and throughput performance. In addition, the intermediate data with the binary format are stored and processed on chip by constructing the computing-in-memory (CIM) architecture to reduce the off-chip communication costs, including power and latency. Wei Mao 0002, Zhihua Xiao, Peng Xu 0035, Dingbang Liu, Shirui Zhao, Fengwei An, Hao Yu 0001 |
ACM Great Lakes Symposium on VLSI | 8 |
| 2020 | DEEPEYE: A Deeply Tensor-Compressed Neural Network for Video Comprehension on Terminal DevicesabstractVideo object detection and action recognition typically require deep neural networks (DNNs) with huge number of parameters. It is thereby challenging to develop a DNN video comprehension unit in resource-constrained terminal devices. In this article, we introduce a deeply tensor-compressed video comprehension neural network, called DEEPEYE, for inference on terminal devices. Instead of building a Long Short-Term Memory (LSTM) network directly from high-dimensional raw video data input, we construct an LSTM-based spatio-temporal model from structured, tensorized time-series features for object detection and action recognition. A deep compression is achieved by tensor decomposition and trained quantization of the time-series feature-based LSTM network. We have implemented DEEPEYE on an ARM-core-based IOT board with 31 FPS consuming only 2.4W power. Using the video datasets MOMENTS, UCF11 and HMDB51 as benchmarks, DEEPEYE achieves a 228.1× model compression with only 0.47% mAP reduction; as well as 15 k × parameter reduction with up to 8.01% accuracy improvement over other competing approaches. Guangya Li, Ngai Wong 0001, Haibao Chen, Hao Yu 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2019 | DEEPEYE: A Deeply Tensor-Compressed Neural Network Hardware Accelerator: Invited PaperabstractVideo detection and classification constantly involve high dimensional data that requires a deep neural network (DNN) with huge number of parameters. It is thereby quite challenging to develop a DNN video comprehension at terminal devices. In this paper, we introduce a deeply tensor compressed video comprehension neural network called DEEPEYE for inference at terminal devices. Instead of building a Long Short-Term Memory (LSTM) network directly from raw video data, we build a LSTM-based spatio-temporal model from tensorized time-series features for object detection and action recognition. Moreover, a deep compression is achieved by tensor decomposition and trained quantization of the time-series feature-based spatio-temporal model. We have implemented DEEPEYE on an ARM-core based IOT board with only 2.4W power consumption. Using the video datasets MOMENTS and UCF11 as benchmarks, DEEPEYE achieves a 228.1× model compression with only 0.47% mAP deduction; as well as 15k× parameter reduction yet 16.27% accuracy improvement. Guangya Li, Ngai Wong 0001, Haibao Chen, Hao Yu 0001 |
ICCAD | 5 |
| 2019 | A large-scale in-memory computing for deep neural network with trained quantization
Haibao Chen, Hao Yu 0001 |
Integr. | 4 |
| 2019 | Spoof Plasmon Interconnects - Communications Beyond RC LimitabstractThe inception of spoof surface plasmon polariton (SSPP) mode realized in planar, patterned conductors to manage light beyond diffraction limit at a chosen frequency garnered significant attention of late. We show that, an SSPP channel can be chosen to act in two distinct ways: first, as a regular RC limited electrical interconnect at low frequencies; and second, as an exotic, beyond RC limit communication channel near its resonant frequency by binding the electromagnetic field on its surface to the elimination of capacitance C. A dynamic transformation between these two modes can constitute an energy economic, tera-scale inter-chip hybrid communication network. We have investigated theoretical limits on the information transfer capability of SSPP interconnects. We show that, a geometry dependent tradeoff relation between cross-talk limited bandwidth density and information traveling length emerges in SSPP-based communication networks. According to our analysis, a bandwidth density of 1 Gbps/μm is attainable in SSPP communication network with ~10-mm information transfer distance, where each channel can carry ~300-Gb/s information with nominal crosstalk. Soumitra Roy Joy, Mikhail Erementchouk, Hao Yu 0001, Pinaki Mazumder |
IEEE Trans. Commun. | 3 |
| 2019 | LTNN: A Layerwise Tensorized Compression of Multilayer Neural NetworkabstractAn efficient deep learning requires a memory-efficient construction of a neural network. This paper introduces a layerwise tensorized formulation of a multilayer neural network, called LTNN, such that the weight matrix can be significantly compressed during training. By reshaping the multilayer neural network weight matrix into a high-dimensional tensor with a low-rank approximation, significant network compression can be achieved with maintained accuracy. An according layerwise training is developed by a modified alternating least-squares method with backward propagation for fine-tuning only. LTNN can provide the state-of-the-art results on various benchmarks with significant compression. For MNIST benchmark, LTNN shows 64 × compression rate without accuracy drop. For Imagenet12 benchmark, our proposed LTNN achieves 35.84 × compression of the neural network with around 2% accuracy drop. We have also shown 1.615 × faster on inference speed than the existing works due to the smaller tensor core ranks. Hantao Huang, Hao Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Design and Analysis of $D$ -Band On-Chip Modulator and Signal Source Based on Split-Ring ResonatorabstractIn an effort toward high-speed and low-power I/O data link in the future exascale data server, this paper presents a signal source and a modulator in the D-band. The split-ring resonator (SRR) structures are used to boost both the signal power and the extinction ratio (ER). The modulator manifests itself as a compact SRR whose magnetic resonance frequency can be modulated by high-speed data. Such a magnetic metamaterial achieves a significant reduction of radiation loss with high ER by stacking two auxiliary SRR unit cells with interleaved placement. The high-Q tank for oscillation is realized by a stacked SRR decorated with slow-wave transmission line (T-line) for electric field confinement. A four-way power-combined fundamental 80-GHz coupled-oscillator network is magnetically synchronized by the slow-wave T-line, which is frequency doubled to 160 GHz. Fabricated in the 65-nm CMOS process, the measured results show that: 1) the modulator achieves 3-dB insertion loss at the onstate with 43-dB isolation at the off-state, leading to a 40-dB ER at 125 GHz within an area of only 40 μm×67 μm and 2) the signal source achieves 6.3% frequency tuning range (FTR) with 3.7-mW peak output power at 160 GHz within 0.053-mm2active area. It has a measured phase noise of -105 dBc/Hz at 10-MHz offset, 5.5% dc-to-RF power efficiency, 70.1-mW/mm2power density, FOM of -171 dBc/Hz, and FOMT of -172.7 dBc/Hz. Yuan Liang 0004, Chirn Chye Boon, Chenyang Li 0008, Xiao-Lan Tang, Herman Jalli Ng, Dietmar Kissinger, Hao Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2018 | SqueezedText: A Real-Time Scene Text Recognition by Binary Convolutional Encoder-Decoder NetworkabstractA new approach for real-time scene text recognition is proposed in this paper. A novel binary convolutional encoder-decoder network (B-CEDNet) together with a bidirectional recurrent neural network (Bi-RNN). The B-CEDNet is engaged as a visual front-end to provide elaborated character detection, and a back-end Bi-RNN performs character-level sequential correction and classification based on learned contextual knowledge. The front-end B-CEDNet can process multiple regions containing characters using a one-off forward operation, and is trained under binary constraints with significant compression. Hence it leads to both remarkable inference run-time speedup as well as memory usage reduction. With the elaborated character detection, the back-end Bi-RNN merely processes a low dimension feature sequence with category and spatial information of extracted characters for sequence correction and classification. By training with over 1,000,000 synthetic scene text images, the B-CEDNet achieves a recall rate of 0.86, precision of 0.88 and F-score of 0.87 on ICDAR-03 and ICDAR-13. With the correction and classification by Bi-RNN, the proposed real-time scene text recognition achieves state-of-the-art accuracy while only consumes less than 1-ms inference run-time. The flow processing flow is realized on GPU with a small network size of 1.01 MB for B-CEDNet and 3.23 MB for Bi-RNN, which is much faster and smaller than the existing solutions. Zichuan Liu, Yixing Li, Fengbo Ren, Wang Ling Goh, Hao Yu 0001 |
AAAI | 5 |
| 2018 | A GPU-Outperforming FPGA Accelerator Architecture for Binary Convolutional Neural NetworksabstractFPGA-based hardware accelerators for convolutional neural networks (CNNs) have received attention due to their higher energy efficiency than GPUs. However, it is challenging for FPGA-based solutions to achieve a higher throughput than GPU counterparts. In this article, we demonstrate that FPGA acceleration can be a superior solution in terms of both throughput and energy efficiency when a CNN is trained with binary constraints on weights and activations. Specifically, we propose an optimized fully mapped FPGA accelerator architecture tailored for bitwise convolution and normalization that features massive spatial parallelism with deep pipelines stages. A key advantage of the FPGA accelerator is that its performance is insensitive to data batch size, while the performance of GPU acceleration varies largely depending on the batch size of the data. Experiment results show that the proposed accelerator architecture for binary CNNs running on a Virtex-7 FPGA is 8.3× faster and 75× more energy-efficient than a Titan X GPU for processing online individual requests in small batch sizes. For processing static data in large batch sizes, the proposed solution is on a par with a Titan X GPU in terms of throughput while delivering 9.5× higher energy efficiency. Yixing Li, Zichuan Liu, Kai Xu 0007, Hao Yu 0001, Fengbo Ren |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2018 | Distributed Machine Learning on Smart-Gateway Network toward Real-Time Smart-Grid Energy Management with Behavior CognitionabstractReal-time data analytics for smart-grid energy management is challenging with consideration of both occupant behavior profiles and energy profiles. This article proposes a distributed and networked machine-learning platform on smart-gateway-based smart-grid in residential buildings. It can analyze occupant behaviors, provide short-term load forecasting, and allocate renewable energy resources. First, occupant behavior profile is captured by real-time indoor positioning system with WiFi data analytics; and the energy profile is extracted by real-time meter system with electricity load data analytics. Then, the 24-hour occupant behavior profile and energy profile are fused with prediction using an online distributed machine-learning algorithm with real-time data update. Based on the forecasted occupant behavior profile and energy profile, solar energy source is allocated to reduce peak demand on the main electricity power-grid. The whole management flow can be operated on the distributed smart-gateway network with limited computational resources but with a supported general machine-learning engine. Experimental results on occupant behavior extraction show that the proposed algorithm can achieve 91.2% positioning accuracy within 3.64m. Moreover, 50× and 38× speed-up is obtained during data testing and training, respectively, when compared to traditional support vector machine (SVM) method. For short-term load forecasting, it is 14.83% more accurate when compared to SVM-based data analytics. Based on the predicted occupant behavior profile and energy profile, our proposed energy management system can achieve 19.66% more peak load reduction and 26.41% more cost saving as compared to the SVM-based method. Hantao Huang, Hang Xu 0001, Yuehua Cai, Suleman Khalid Rai, Hao Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2017 | A 7.663-TOPS 8.2-W Energy-efficient FPGA Accelerator for Binary Convolutional Neural Networks (Abstract Only)
Yixing Li, Zichuan Liu, Kai Xu 0007, Hao Yu 0001, Fengbo Ren |
FPGA | 4 |
| 2017 | An energy-efficient and high-throughput bitwise CNN on sneak-path-free digital ReRAM crossbarabstractConvolutional neural network (CNN) based machine learning requires a highly parallel as well as low power consumption (including leakage power) hardware accelerator. In this paper, we will present a digital ReRAM crossbar based CNN accelerator that can achieve significantly higher throughput and lower power consumption than state-of-arts. The CNN is trained with binary constraints on both weights and activations such that all operations become bitwise. With further use of 1-bit comparator, the bitwise CNN model can be naturally realized on a digital ReRAM-crossbar device. A novel sneak-path-free ReRAM-crossbar is further utilized for large-scale realization. Simulation experiments show that the bitwise CNN accelerator on the digital ReRAM crossbar achieves 98.3% and 91.4% accuracy on MNIST and CIFAR-10 benchmarks, respectively. Moreover, it has a peak throughput of 792GOPS at the power consumption of 6.3mW, which is 18.86 times higher throughput and 44.1 times lower power than CMOS CNN (non-binary) accelerators. Leibin Ni, Zichuan Liu, J. Joshua Yang, Hao Yu 0001, Kanwen Wang, Yuangang Wang |
ISLPED | 5 |
| 2017 | Distributed In-Memory Computing on Binary RRAM CrossbarabstractThe recently emerging resistive random-access memory (RRAM) can provide nonvolatile memory storage but also intrinsic computing for matrix-vector multiplication, which is ideal for the low-power and high-throughput data analytics accelerator performed in memory. However, the existing RRAM crossbar--based computing is mainly assumed as a multilevel analog computing, whose result is sensitive to process nonuniformity as well as additional overhead from AD-conversion and I/O. In this article, we explore the matrix-vector multiplication accelerator on a binary RRAM crossbar with adaptive 1-bit-comparator--based parallel conversion. Moreover, a distributed in-memory computing architecture is also developed with the according control protocol. Both memory array and logic accelerator are implemented on the binary RRAM crossbar, where the logic-memory pair can be distributed with the control bus protocol. Experimental results have shown that compared to the analog RRAM crossbar, the proposed binary RRAM crossbar can achieve significant area savings with better calculation accuracy. Moreover, significant speedup can be achieved for matrix-vector multiplication in neural network--based machine learning such that the overall training and testing time can be both reduced. In addition, large energy savings can be also achieved when compared to the traditional CMOS-based out-of-memory computing architecture. Leibin Ni, Hantao Huang, Zichuan Liu, Rajiv V. Joshi, Hao Yu 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2017 | A Multiagent Minority-Game-Based Demand-Response Management of Smart Buildings Toward Peak Load ReductionabstractThis paper presents a cyber-physical management of smart buildings based on smart-gateway network with distributed and real-time energy data collection and analytics. We consider a building with multiple rooms supplied with one main electricity grid and one additional solar energy grid. Based on smart-gateway network, energy signatures of rooms are first extracted with consideration of uncertainty and further classified as different types of agents. Then, a multiagent minority-game (MG)-based demand-response management is introduced to reduce peak demand on the main electricity grid and also to fairly allocate solar energy on the additional grid. Experiment results show that compared to the traditional static and centralized energy-management system (EMS), and the recent multiagent EMS using price-demand competition, the proposed uncertainty-aware MG-EMS can achieve up to 50× and 145× utilization rate improvements, respectively, regarding to the fairness of solar energy resource allocation. More importantly, the peak load from the main electricity grid is reduced by 38.50% in summer and 15.83% in winter based on benchmarked energy data of building. Lastly, an average 23% uncertainty can be reduced with an according 37% balanced energy allocation improved comparing to the MG-EMS without consideration of uncertainty. Hantao Huang, Yuehua Cai, Hang Xu 0001, Hao Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | An energy-efficient matrix multiplication accelerator by distributed in-memory computing on binary RRAM crossbarabstractEmerging resistive random-access memory (RRAM) can provide non-volatile memory storage but also intrinsic logic for matrix-vector multiplication, which is ideal for low-power and high-throughput data analytics accelerator performed in memory. However, the existing RRAM-based computing device is mainly assumed on a multi-level analog computing, whose result is sensitive to process non-uniformity as well as additional AD- conversion and I/O overhead. This paper explores the data analytics accelerator on binary RRAM-crossbar. Accordingly, one distributed in-memory computing architecture is proposed with design of according component and control protocol. Both memory array and logic accelerator can be implemented by RRAM-crossbar purely in binary, where logic-memory pairs can be distributed with protocol of control bus. Based on numerical results for fingerprint matching that is mapped on the proposed RRAM-crossbar, the proposed architecture has shown 2.86x faster speed, 154x better energy efficiency, and 100x smaller area when compared to the same design by CMOS-based ASIC. Leibin Ni, Yuhao Wang 0002, Hao Yu 0001, Chuliang Weng, Junfeng Zhao 0003 |
ASP-DAC | 3 |
| 2016 | Distributed-neuron-network based machine learning on smart-gateway network towards real-time indoor data analytics
Hantao Huang, Yuehua Cai, Hao Yu 0001 |
DATE | 3 |
| 2016 | Lab-on-CMOS: A multi-modal CMOS sensor platform towards personalized DNA sequencingabstractPrecision medicine requires scalable bioinstrument for a personalized DNA sequencing, which can be label-free, cost-efficient, and high-throughput. This paper mainly presents three kinds of CMOS-based label-free sensors, including: i) a high-sensitivity ion-sensitive field-effect transistor (ISFET) sensor with pH-to-time-to-voltage conversion (pH-TVC); ii) a dual-mode sensor with image and chemical modes for high accuracy; and iii) a THz metamaterial sensor with electrical resonance detection. The developed CMOS multi-modal sensor platform can show a scaled solution for future personalized DNA sequencing. Yu Jiang 0004, Xu Liu 0002, Xiwei Huang, Yang Shang, Mei Yan, Hao Yu 0001 |
ISCAS | 6 |
| 2016 | On-line machine learning accelerator on digital RRAM-crossbarabstractOn-line machine learning has become the need for future data analytics. This work will show an ℓ2norm based hardware solver for on-line machine learning that can significantly reduce training time when compared to the traditional gradient-based solution using backward propagation. We will show that the intensive matrix-vector multiplication in ℓ2norm solution can be mapped onto a distributed in-memory accelerator using the recent resistive switching random access memory (RRAM) device. A digitized matrix-vector multiplication accelerator will be developed based on the distributed RRAM-crossbar. Such a distributed RRAM-crossbar architecture can utilize the reformulated ℓ2norm solver with a scalable and energy-efficient solution for real-time training and testing in image recognition. Experiment results have shown that significant speedup can be achieved for matrix-vector multiplication in the ℓ2norm solver such hat the overall training and testing time can be reduced respectively. In addition, large energy saving can be also achieved when compared to the traditional CMOS-based out-of-memory computing architecture. Leibin Ni, Hantao Huang, Hao Yu 0001 |
ISCAS | 3 |
| 2016 | A Compressive-sensing based Testing Vehicle for 3D TSV Pre-bond and Post-bond Testing DataabstractOnline testing vehicle is required for 3D TSV pre-bond and post-bond testing due to high probability of TSV failures. It has become a challenge to deal with large sets of generated testing data with limited probing when transmitting the data out. In this paper, a lossless compressive-sensing based testing vehicle is developed for online testing of TSVs. By exploring sparsity of the testing data under constraint of failure bound of TSV, sparse-representation based encoding can be deployed by XOR and AND network on chip to deal with large volume of testing data. Experimental results (with benchmarks) have shown that 89.70% pre-bond data compression rate can be achieved under 0.5% probability of failures; and 88.18% post-bond data compression rate can be achieved with 5% probability of failures. Hantao Huang, Hao Yu 0001, Cheng Zhuo, Fengbo Ren |
ISPD | 2 |
| 2016 | A Q-Learning Based Self-Adaptive I/O Communication for 2.5D Integrated Many-Core Microprocessor and MemoryabstractA self-adaptive output-voltage swing adjustment is introduced in the design of energy-efficient I/O communication for 2.5D integrated many-core microprocessor and memory. Instead of transmitting signal with large voltage swing, a Q-learning based I/O management is deployed to adaptively adjust the I/O output-voltage swing under constraints of both communication power and bit error rate (BER). Simulation results show that the proposed adaptive 2.5D I/Os (in 65 nm CMOS) can achieve an average of 12.5 mW I/O power, 4 GHz bandwidth and 3.125 pJ/bit energy efficiency for one channel under 10-6BER. With the use of conventional Q-learning and further accelerated Q-learning, we can achieve 12.95 and 18.89 percent power reduction and 14 and 15.11 percent energy efficiency improvement when compared to the use of uniform output-voltage swing based I/O communication. Sai Manoj Pudukotai Dinakarrao, Hao Yu 0001, Hantao Huang, Dongjun Xu |
IEEE Trans. Computers | 2 |
| 2016 | A Zonotoped Macromodeling for Eye-Diagram Verification of High-Speed I/O Links With Jitter and Parameter VariationsabstractIt is challenging to efficiently evaluate the performance bound of high-precision analog circuits with input and parameter variations at nano-scale. With the use of zonotope to model uncertainty of input data pattern (or jitter) and multiple parameters, a reachability-based verification is developed in this paper to compute the worst-case eye-diagram. The proposed zonotope-based reachability analysis can consider both spatial and temporal variations in one-time simulation. Moreover, a nonlinear zonotoped macromodeling is further developed to reduce the computational complexity. Performance bound for I/O links considering the parameter variations are evaluated. In addition, the eye-diagrams are generated by the proposed zonotoped macromodel for performance evaluation considering both temporal and spatial variations. As shown by experiments, the zonotoped macromodel achieves up to 450× speedup compared to the Monte Carlo simulation of the original model within small error under specified macromodel order for high-speed I/O links eye-diagram verification. Leibin Ni, Sai Manoj Pudukotai Dinakarrao, Chenjie Gu, Hao Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | DW-AES: A Domain-Wall Nanowire-Based AES for High Throughput and Energy-Efficient Data Encryption in Non-Volatile MemoryabstractBig-data storage poses significant challenges to anonymization of sensitive information against data sniffing. Not only will the encryption bandwidth be limited by the I/O traffic, the transfer of data between the processor and the memory will also expose the input-output mapping of intermediate computations on I/O channels that are susceptible to semi-invasive and non-invasive attacks. Limited by the simplistic cell-level logic, existing logic-in-memory computing architectures are incapable of performing the complete encryption process within the memory at reasonable throughput and energy efficiency. In this paper, a block-level in-memory architecture for advanced encryption standard (AES) is proposed. The proposed technique, called DW-AES, maps all AES operations directly to the domain-wall nanowires. The entire encryption process can be completed within a homogeneous, high-density, and standby-power-free non-volatile spintronic-based memory array without exposing the intermediate results to external I/O interface. Domain-wall nanowire-based pipelining and multi-issue pipelining methods are also proposed to increase the throughput of the baseline DW-AES with an insignificant area overhead and negligible difference on leakage power and energy consumption. The experimental results show that DW-AES can reduce the leakage power and area by the orders of magnitude compared with existing CMOS ASIC accelerators. It has an energy efficiency of 22 pJ/b, which is 5× and 3× better than the CMOS ASIC and memristive CMOL-based implementations, respectively. Under the same area budget, the proposed DW-AES achieves 4.6× higher throughput than the latest CMOS ASIC AES with similar power consumption. The throughput improvement increases to 11× for pipelined DW-AES at the expense of doubling the power consumption. Yuhao Wang 0002, Leibin Ni, Chip-Hong Chang, Hao Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2015 | A 64×64 1200fps dual-mode CMOS ion-image sensor for accurate DNA sequencingabstractA dual-mode (chemical/optical) CMOS ion-image sensor is demonstrated towards accurate DNA sequencing by integrating the ion-sensitive field-effect transistor (ISFET) with 4-T CMOS image sensor (CIS) pixel fabricated in standard 0.18μm CIS process. With accurate determination of physical locations for microbeads through contact imaging, local pH for one DNA slice attached on the microbead can be obtained with accurate correlation to improve sequencing accuracy from system perspective. Moreover, towards high-throughput large-arrayed sequencing, pixel-to-pixel ISFET threshold voltage mismatch is reduced by correlated double sampling (CDS) readout that supports both image and pH modes. Measurement results show a readout sensitivity of 103.8mV/pH, a fixed-pattern-noise (FPN) reduction from 4% to 0.3%, and a readout speed of 1200 frames/second (fps). Xiwei Huang, Mei Yan, Hao Yu 0001 |
ASP-DAC | 4 |
| 2015 | An energy-efficient non-volatile in-memory accelerator for sparse-representation based face recognition
Yuhao Wang 0002, Hantao Huang, Leibin Ni, Hao Yu 0001, Mei Yan, Chuliang Weng, Junfeng Zhao 0003 |
DATE | 4 |
| 2015 | Indoor positioning by distributed machine-learning based data analytics on smart gateway networkabstractReal-time data analysis on sensor nodes is challenging due to limited computing resources. A changing environment where received signal strength (RSSI) varies with time makes it more complex to update position predictors for real-time indoor positioning. Based on the distributed collection and analytics of RSSI values in a gateway network, a time-efficient workload-based (WL) distributed support vector machine (WL-DSVM) algorithm is introduced in this paper to perform the indoor positioning. Experimental results show that with 5 distributed sensor nodes running in parallel, the proposed WL-DSVM can achieve a performance improvement in run time up to 3.2× with a stable positioning accuracy. Yuehua Cai, Suleman Khalid Rai, Hao Yu 0001 |
IPIN | 3 |
| 2015 | A 16-channel 24-V 1.8-mA power efficiency enhanced neural/muscular stimulator with exponentially decaying stimulation currentabstractThis paper presents a current-mode neural/muscular stimulator with an exponentially decaying stimulation current. The use of exponentially decaying current makes the voltage on the stimulating electrode constant during the stimulation, which eliminates the headroom and increases the power efficiency. A simple exponentially decaying current generator is proposed based on Taylor series approximation and implemented in a 16-channel prototype stimulator IC. The prototype IC is fabricated in a 0.18-μm CMOS process with high-voltage LDMOS option, occupying a core area of 1.65 mm × 1.65 mm. The stimulator is tested with different loads, which mimics the electrode impedances, and the measured results show that maximum stimulation power efficiency of 95.9% can be achieved at the output stage of the stimulator. Depending on the electrode impedance and stimulation current, the power efficiency can be improved by nearly 10% at the output stage, compared to traditional constant-current stimulator. Xu Liu 0002, Mei Yan, Shih-Cheng Yen, Hao Yu 0001, Minkyu Je, Yong Ping Xu |
ISCAS | 6 |
| 2015 | A body-biasing of readout circuit for STT-RAM with improved thermal reliabilityabstractAs the integration density rockets up for contemporary VLSI circuits, power consumption limits the scalability of technology advancement of CMOS. Spin transfer torque-magnetic random access memory (STT-MRAM), as one of the emerging non-CMOS technologies, has the promising prospect of low standby power, fast access speed and compatibility with the CMOS fabrication process. However, with the technology node scaling down, typical 1 Transistor-1 Magnetic Tunnel Junction (1T-1MTJ) STT-RAM cell suffers from severe reliability challenges, especially for read operation under temperature fluctuation. In this paper, we quantitatively analyze the temperature effect on read reliability of STT-RAM cell and propose a novel body-biasing feedback readout circuit design to improve the read sensing margin under different temperatures. The experiments based on 40nm CMOS technology and MTJ compact model validate the effectiveness of the proposed method. The improved sensing margin also permits a smaller sensing current for reading such that higher read energy efficiency can be achieved. Lun Yang, Yuanqing Cheng, Yuhao Wang 0002, Hao Yu 0001, Weisheng Zhao 0001, Aida Todri |
ISCAS | 4 |
| 2015 | An energy efficient and low cross-talk CMOS sub-THz I/O with surface-wave modulator and interconnectabstractFree-space EM-wave based GHz interconnect has significant loss and crosstalk that cannot be deployed as low-power and dense I/Os for future network-on-chip (NoC) integration of many-core and memory. This paper proposes an energy-efficient and low-crosstalk sub-THz (0.1T-1T) I/O with use of surface-wave based modulator and interconnects in CMOS. By introducing sub-wavelength periodical corrugation structure onto transmission line, the surface-wave is established to propagate signal that is strongly localized on surface of top-layer metal wire, which results in low coupling into lossy substrate and neighboring metal wires. As such, significant power saving and cross-talk reduction can be observed with high communication bandwidth. In addition, a high on/off-ratio surface-wave modulator is also proposed to support on-chip THz communication. As designed in 65nm CMOS, the results have shown that the proposed surface-wave I/O interface achieves 25Gbps data rate and 0.016pJ/bit/mm energy efficiency at 140GHz carrier frequency over 20mm surface-wave channels. They can be placed with 2.4μm channel spacing and a -20dB crosstalk ratio. The surface-wave modulator also achieves significant reduction of radiation loss with 23dB extinction ratio. Yuan Liang 0004, Hao Yu 0001, Junfeng Zhao 0003, Yuangang Wang |
ISLPED | 2 |
| 2015 | Optimizing Boolean embedding matrix for compressive sensing in RRAM crossbarabstractThe emerging resistive random-access-memory (RRAM) crossbar provides an intrinsic fabric for matrix-vector multiplication, which can be leveraged as power efficient linear embedding hardware for data analytics such as compressive sensing. As the matrix elements are represented by resistance of RRAM cells, it imposes constraints for the embedding matrix due to limited RRAM programming resolution. A random Boolean embedding can be efficiently mapped to the RRAM crossbar but suffers from poor performance. Learning-based embedding matrices can deliver optimized performance but are continuous-valued which prevents it from being mapped to RRAM crossbar structure directly. In this paper, we have proposed one algorithm that can find an optimal Boolean embedding matrix for a given learned real-valued embedding matrix, so that it can be effectively mapped to the RRAM crossbar structure while high performance is preserved. The numerical experiments demonstrate that the proposed optimized Boolean embedding can reduce the embedding distortion by 2.7x, and image recovery error by 2.5x compared to the random Boolean embedding, both mapped on RRAM crossbar. In addition, optimized Boolean embedding on RRAM crossbar exhibits 10x faster speed, 17x better energy efficiency, and three orders of magnitude smaller area with slight accuracy penalty, when compared to the optimized real-valued embedding on CMOS ASIC platform. Yuhao Wang 0002, Xin Li 0001, Hao Yu 0001, Leibin Ni, Chuliang Weng, Junfeng Zhao 0003 |
ISLPED | 3 |
| 2015 | A robust recognition error recovery for micro-flow cytometer by machine-learning enhanced single-frame super-resolution processingabstractWith the recent advancement in microfluidics based lab-on-a-chip technology, lensless imaging system integrating microfluidic channel with CMOS image sensor has become a promising solution for the system minimization of flow cytometer. The design challenge for such an imaging-based micro-flow cytometer under poor resolution is how to recover cell recognition error under various flow rates. A microfluidic lensless imaging system is developed in this paper using extreme-learning-machine enhanced single-frame super-resolution processing, which can effectively recover the recognition error when increasing flow rate for throughput. As shown in the experiments, with mixed flowing HepG2 and Huh7 cells as inputs, the developed scheme shows that 23% better recognition accuracy can be achieved compared to the one without error recovery. Meanwhile, it also achieves an average of 98.5% resource saving compared to the previous multi-frame super-resolution processing. Xiwei Huang, Mei Yan, Hao Yu 0001 |
Integr. | 4 |
| 2015 | 3D Many-Core Microprocessor Power Management by Space-Time Multiplexing Based Demand-Supply MatchingabstractA reconfigurable power switch network is proposed to perform a demand-supply matched power management between 3D-integrated microprocessor cores and power converters. The power switch network makes physical connections between cores and converters by 3D through-silicon-vias (TSVs). Space-time multiplexing is achieved by the configuration of power switch network and is realized by learning and classifying power-signature of workloads. As such, by classifying workloads based on magnitude and phase of power-signature, space-time multiplexing can be performed with the minimum number of converters allocated to cluster of cores. Furthermore, a demand-response based workload scheduling is performed to reduce peak-power and to balance workload. The proposed power management is verified by system models with physical design parameters and benched power traces of workloads. For a 64-core case, experiment results show 40.53 percent peak-power reduction and 2.50x balanced workload along with a 42.86 percent reduction in the required number of power converters compared to the work without using STM based power management. Sai Manoj Pudukotai Dinakarrao, Hao Yu 0001, Kanwen Wang |
IEEE Trans. Computers | 2 |
| 2015 | A GPU-Accelerated Parallel Shooting Algorithm for Analysis of Radio Frequency and Microwave Integrated CircuitsabstractThis paper presents a new parallel shooting-Newton method based on a graphic processing unit (GPU)-accelerated periodic Arnoldi shooting solver (GAPAS) for fast periodic steady-state analysis of radio frequency/millimeter-wave integrated circuits. The new algorithm first explores a periodic structure of the state matrix by using a periodic Arnoldi algorithm for computing the resulting structured Krylov subspace in the generalized minimal residual (GMRES) solver. The resulting periodic Arnoldi shooting method is very amenable for massive parallel computing, such as GPUs. Second, the periodic Arnoldi-based GMRES solver in the shooting-Newton method is parallelized on the recent NVIDIA Tesla GPU platforms. We further explore CUDA GPUs features, such as coalesced memory access and overlapping transfers with computation to boost the efficiency of the resulting parallel GAPAS method. Experimental results from several industrial examples show that when compared with the state-of-the-art implicit GMRES method under the same accuracy, the new parallel shooting-Newton method can lead up to $8\times$ speedup. Xuexin Liu, Hao Yu 0001, Sheldon X.-D. Tan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | A robustness optimization of SRAM dynamic stability by sensitivity-based reachability analysisabstractA robustness optimization of SRAM dynamic stability at nano-scale is developed in this paper by zonotope-based reachability analysis. A backward Euler method is developed to efficiently perform reachability analysis by zonotope to deal with multiple device parameters with tuning ranges. Moreover, a sensitivity calculation of zonotope is developed to optimize safety distance by simultaneously tuning multiple SRAM device parameters without multiple repeated computations. As such, sequential robustness optimizations can be performed such that the optimized SRAM designs can depart from unsafe region but converge into safe region. The proposed method is implemented inside a SPICE-like simulator. As shown by numerical experiments, the proposed method can achieve 600× speedup on average compared to the traditional verification method by Monte-Carlo under the similar accuracy. In addition, compared to the traditional small-signal based sensitivity optimization, the proposed method can converge faster with high accuracy. Sai Manoj Pudukotai Dinakarrao, Hao Yu 0001 |
ASP-DAC | 3 |
| 2014 | Energy efficient in-memory machine learning for data intensive image-processing by non-volatile domain-wall memoryabstractImage processing in conventional logic-memory I/O-integrated systems will incur significant communication congestion at memory I/Os for excessive big image data at exa-scale. This paper explores an in-memory machine learning on neural network architecture by utilizing the newly introduced domain-wall nanowire, called DW-NN. We show that all operations involved in machine learning on neural network can be mapped to a logic-in-memory architecture by non-volatile domain-wall nanowire. Domain-wall nanowire based logic is customized for in machine learning within image data storage. As such, both neural network training and processing can be performed locally within the memory. The experimental results show that system throughput in DW-NN is improved by 11.6x and the energy efficiency is improved by 92x when compared to conventional image processing system. Hao Yu 0001, Yuhao Wang 0002, Wei Fei, Chuliang Weng, Junfeng Zhao 0003, Zhulin Wei |
ASP-DAC | 1 |
| 2014 | Zonotope-based nonlinear model order reduction for fast performance bound analysis of analog circuits with multiple-interval-valued parameter variationsabstractIt is challenging to efficiently evaluate performance bound of high-precision analog circuits with multiple parameter variations at nano-scale. In this paper, a nonlinear model order reduction is proposed to deploy zonotope-based model for multiple-interval-valued parameter variations. As such, one can have a zonotope-based reachability analysis to generate a set of trajectories with performance bound defined. By further constructing local parameterized subspaces to approximate a number of zonotopes along the set of trajectories, one can perform nonlinear model order reduction to generate the performance bound under parameter variations. As shown by numerical experiments, the zonotope-based nonlinear macromodeling by order of 19 achieves up to 500× speedup when compared to Monte Carlo simulations of the original model; and up to 50% smaller error when compared to previous parameterized nonlinear macromodeling under the same order. Sai Manoj Pudukotai Dinakarrao, Hao Yu 0001 |
DATE | 3 |
| 2014 | Energy efficient in-memory AES encryption based on nonvolatile domain-wall nanowireabstractThe widely applied Advanced Encryption Standard (AES) encryption algorithm is critical in secure big-data storage. Data oriented applications have imposed high throughput and low power, i.e., energy efficiency (J/bit), requirements when applying AES encryption. This paper explores an in-memory AES encryption using the newly introduced domain-wall nanowire. We show that all AES operations can be fully mapped to a logic-in-memory architecture by non-volatile domain-wall nanowire, called DW-AES. The experimental results show that DW-AES can achieve the best energy efficiency of 24 pJ/bit, which is 9X and 6.5X times better than CMOS ASIC and memristive CMOL implementations, respectively. Under the same area budget, the proposed DW-AES exhibits 6.4X higher throughput and 29% power saving compared to a CMOS ASIC implementation; 1.7X higher throughput and 74% power reduction compared to a memristive CMOL implementation. Yuhao Wang 0002, Hao Yu 0001, Dennis Sylvester, Pingfan Kong |
DATE | 2 |
| 2014 | A thermal resilient integration of many-core microprocessors and main memory by 2.5D TSI I/OsabstractOne memory-logic-integration design platform is developed in this paper with thermal reliability analysis provided for 2.5D through-silicon-interposer (TSI) and 3D through-silicon-via (TSV) based integrations. Temperature-dependent delay and power models have been developed at microarchitecture level for 2.5D and 3D integrations of many-core microprocessors and main memory, respectively. Experiments are performed by general-purpose benchmarks from SPEC CPU2006 and also cloud-oriented benchmarks from Phoenix with the following observations. The memory-logic integration by 3D RC-interconnected TSV I/Os can result in thermal runaway failures due to strong electrical-thermal couplings. On the other hand, the one by 2.5D transmission-line-interconnected TSI I/Os has shown almost the same energy efficiency and better thermal resilience. Sih-Sian Wu, Kanwen Wang, Sai Manoj Pudukotai Dinakarrao, Tsung-Yi Ho, Mingbin Yu, Hao Yu 0001 |
DATE | 6 |
| 2014 | A zonotoped macromodeling for reachability verification of eye-diagram in high-speed I/O links with jitterabstractWith the use of zonotope to model uncertainty of input data pattern (or jitter), a reachability-based verification is developed in this paper to compute the worst-case eye-diagram. The proposed zonotope-based reachability analysis can consider both spatial and temporal variations in one-time simulation of high-speed I/O links. Moreover, nonlinear zonotoped macromodeling is developed to reduce the verification complexity. As shown by experiments, the zonotoped macromodel achieves up to 450× speedup compared to the Monte Carlo simulation of the original model within small error under specified macromodel order for highspeed I/O links verification. Sai Manoj Pudukotai Dinakarrao, Hao Yu 0001, Chenjie Gu, Cheng Zhuo |
ICCAD | 2 |
| 2014 | Reinforcement learning based self-adaptive voltage-swing adjustment of 2.5D I/Os for many-core microprocessor and memory communicationabstractA reinforcement learning based I/O management is developed for energy-efficient communication between many-core microprocessor and memory. Instead of transmitting data under a fixed large voltage-swing, an online reinforcement Q-learning algorithm is developed to perform a self-adaptive voltage-swing control of 2.5D through-silicon interposer (TSI) I/O circuits. Such a voltage-swing adjustment is formulated as a Markov decision process (MDP) problem solved by model-free reinforcement learning under constraints of both power budget and bit-error-rate (BER). Experimental results show that the adaptive 2.5D TSI I/Os designed in 65nm CMOS can achieve an average of 12.5mw I/O power, 4GHz bandwidth and 3.125pJ/bit energy efficiency for one channel under 10-6BER, which has 18.89% power saving and 15.11% improvement of energy efficiency on average. Hantao Huang, Sai Manoj Pudukotai Dinakarrao, Dongjun Xu, Hao Yu 0001, Zhigang Hao |
ICCAD | 4 |
| 2014 | An energy-efficient 2.5D through-silicon interposer I/O with self-adaptive adjustment of output-voltage swingabstractA self-adaptive output swing adjustment is introduced for the design of energy-efficient 2.5D through-silicon interposer (TSI) I/Os. Instead of transmitting signal with large voltage swing, Q-learning based self-adaptive adjustment is deployed to adjust I/O output-voltage swing under constraints of both power budget and bit error rate (BER). Experimental results show that the adaptive 2.5D TSI I/Os designed in 65nm CMOS can achieve an average of 13mW I/O power, 4GHz bandwidth and 3.25pJ/bit energy efficiency for one channel under 10-6 BER, which has ~21.42% reduction of power and ~14.47% energy efficiency improvement. Dongjun Xu, Sai Manoj Pudukotai Dinakarrao, Hantao Huang, Ningmei Yu, Hao Yu 0001 |
ISLPED | 5 |
| 2014 | Reachability-Based Robustness Verification and Optimization of SRAM Dynamic Stability Under Process VariationsabstractThe dynamic stability margin of SRAM is largely suppressed at nanoscale due to not only dynamic noise but also process variation. This paper introduces an analog verification for SRAM dynamic stability under threshold-voltage variations. A zonotope-based reachability analysis by the backward Euler method is deployed for SRAM dynamic stability in state space with consideration of SRAM nonlinear dynamics. It can simultaneously consider multiple SRAM variation sources without multiple repeated computations. What is more, sensitivity analysis is developed for zonotope to optimize SRAM designs departing from unsafe regions by simultaneously tuning multiple SRAM device parameters. In addition, compared to the SRAM optimization by single-parameter small-signal sensitivity, the proposed method can converge faster with higher accuracy. As shown by numerical experiments, the proposed optimization method can achieve 600× speedup on average when compared to the repeated Monte Carlo simulations under the similar accuracy. Hao Yu 0001, Sai Manoj Pudukotai Dinakarrao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Nonvolatile CBRAM-Crossbar-Based 3-D-Integrated Hybrid Memory for Data RetentionabstractThis paper explores the design of 3-D-integrated hybrid memory by conductive-bridge random-access-memory (CBRAM). Considering internal states, height, and radius of the conductive bridge of one CBRAM device, an accurate CBRAM device model is developed for CBRAM-crossbar-based nonvolatile memory design with efficient estimation of area, access time, and power. Based on this design platform, one 3-D-integrated hybrid memory is designed by stacking one tier of CBRAMcrossbar with tiers of static random access memory (SRAM) and dynamic random access memory (DRAM), where the tier of CBRAM-crossbar is deployed for data retention during power gating of SRAM/DRAM tiers. One corresponding block-level data retention is developed to only write back dirty data from SRAM/DRAM to CBRAM-crossbar. When compared with phase-change random-access-memory-based system-level data retention, our design achieves 11× faster data-migration speed and 10× less data-migration power. When compared with ferroelectric random-access-memory-based bit-level data retention, our design also achieves 17× smaller area and 56× smaller power under the same data-migration speed. Yuhao Wang 0002, Hao Yu 0001, Wei Zhang 0012 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Thermal simulator of 3D-IC with modeling of anisotropic TSV conductance and microchannel entrance effectsabstractThis paper presents a fast and accurate steady state thermal simulator for heatsink and microfluid-cooled 3D-ICs. This model considers the thermal effect of TSVs at fine-granularity by calculating the anisotropic equivalent thermal conductances of a solid grid cell if TSVs are inserted. Entrance effect of microchannels is also investigated for accurate modeling of microfluidic cooling. The proposed thermal simulator is verified against commercial multiphysics solver COMSOL and compared with Hotspot and 3D-ICE. Simulation results shows that for heatsink cooling, the proposed simulator is as accurate as Hotspot but runs much faster at moderate granularity. For microfluidic cooling, our proposed simulator is much more accurate than 3D-ICE in its estimation of steady state temperature and thermal distribution. Hanhua Qian, Hao Liang 0003, Chip-Hong Chang, Wei Zhang 0012, Hao Yu 0001 |
ASP-DAC | 5 |
| 2013 | Thermal-reliable 3D clock-tree synthesis considering nonlinear electrical-thermal-coupled TSV modelabstract3D physical design needs accurate device model of through-silicon vias (TSVs). In this paper, physics-based electrical-thermal model is introduced for both signal and dummy thermal TSVs with the consideration of nonlinear electrical-thermal dependence. Taking thermal-reliable 3D clock-tree synthesis as a case-study to verify the effectiveness of the proposed TSV model, one nonlinear programming-based clock-skew reduction problem is formulated to allocate thermal TSVs for clock-skew reduction under non-uniform temperature distribution. With a number of 3D clock-tree benchmarks, experiments show that under the nonlinear electrical-thermal TSV model, insertion of thermal TSVs can effectively reduce temperature-gradient introduced clock-skew by 58.4% on average, and has 11.6% higher clock-skew reduction than the result under linear electrical-thermal model. Yang Shang, Chun Zhang 0003, Hao Yu 0001, Chuan Seng Tan, Xin Zhao 0001, Sung Kyu Lim |
ASP-DAC | 3 |
| 2013 | Stable backward reachability correction for PLL verification with consideration of environmental noise induced jitterabstractIt is unknown to perform efficient PLL system-level verification with consideration of jitter induced by substrate or power-supply noise. With the consideration of nonlinear phase noise macromodel, this paper introduces a forward reachability analysis with stable backward correction for PLL system-level verification with jitter. By refining initial state of PLL through backward correction, one can perform an efficient PLL verification to automatically adjust the locking range with consideration of environmental noise induced jitter. Moreover, to overcome the unstable nature during backward correction, a stability calibration is introduced in this paper to limit error. To validate our method, the proposed approach is applied to verify a number of PLL designs including single-LC or coupled-LC oscillators described by system-level behavioral model with jitter. Experimental results show that our forward reachability analysis with backward correction can succeed in reaching the adjusted locking range by correcting initial states in presence of environmental noise induced jitter. Haipeng Fu, Hao Yu 0001, Guoyong Shi |
ASP-DAC | 3 |
| 2013 | Peak power reduction and workload balancing by space-time multiplexing based demand-supply matching for 3D thousand-core microprocessorabstractSpace-time multiplexing is utilized for demand-supply matching between many-core microprocessors and power converters. Adaptive clustering is developed to classify cores by similar power level in space and similar power behavior in time. In each power management cycle, minimum number of power converters are allocated for space-time multiplexed matching, which is physically enabled by 3D through-silicon-vias. Moreover, demand-response based task adjustment is applied to reduce peak power and to balance workload. The proposed power management system is verified by system models with physical design parameters and benched power traces, which show 38.10% peak power reduction and 2.60x balanced workload. Sai Manoj Pudukotai Dinakarrao, Kanwen Wang, Hao Yu 0001 |
DAC | 3 |
| 2013 | 3D reconfigurable power switch network for demand-supply matching between multi-output power converters and many-core microprocessorsabstractA 3D reconfigurable power switch network is introduced to optimally provide demand-supply matching between on-chip multi-output power converters and many-core microprocessors. For effective DVFS power management of many cores by area-efficient on-chip power converters, the reconfigurable power switch network supports space and time multiplexed access between power converters and cores. An integer linear programming is deployed to find one configuration of space-time multiplexing that can match between supply and demand with balanced utilization. The overall power management system is verified in SystemC-AMS based models. Experiment results show that the proposed design achieves 35.36% power saving on average when compared to the one without using the proposed power management. Kanwen Wang, Hao Yu 0001, Benfei Wang, Chun Zhang 0003 |
DATE | 2 |
| 2013 | Cyber-physical management for heterogeneously integrated 3D thousand-core on-chip microprocessorabstractThough 3D TSV/TSI technology provides the promising platform for heterogeneous system integration with design drivers ranged from thousand-core microprocessor to millimeter-cubic sensor, the fundamental challenge is lack of light to deal with significantly increased design complexity. From device level, new state of variables from different physical domains such as MEMS, microfluidic and NVM devices have to be identified and described together with conventional states from CMOS VLSI; and from system level, cyber management of states of voltage-level and temperature has to be maintained under a real-time demand response fashion. Moreover, a cyber-physical link is required to compress and virtualize device level state details during system level state control. This paper shows device-level 3D integration by example of MEMS and CMOS VLSI. In addition, a cyber-physical thermal management for 3D integrated many-core microprocessors is discussed. Sai Manoj Pudukotai Dinakarrao, Hao Yu 0001 |
ISCAS | 2 |
| 2013 | A 5-bit 1.25GS/s 4.7mW delay-based pipelined ADC in 65nm CMOSabstractThis paper presents a delay based pipeline (DBP) analog to digital converter (ADC) suitable for high speed and low power applications. Active sample and hold and residue amplifier used in conventional pipeline ADCs are replaced by an analog delay line. The analog delay line is implemented by time-interleaved sampling of the signal in each stage of the ADC. A novel multi-phase clock generator is introduced to generate ADC timing signals. A 5-bit, 1.25 GS/s DBP ADC is designed in 65nm CMOS process. Post-layout simulations confirm that the proposed ADC achieves a peak SNDR of 30.5dB while consuming 4.7mW from a single 1.2V power supply. Ali Mesgarani, Haipeng Fu, Mei Yan, A. Tekin, Hao Yu 0001, Suat U. Ay |
ISCAS | 5 |
| 2013 | An ultralow-power memory-based big-data computing platform by nonvolatile domain-wall nanowire devicesabstractAs one recently introduced non-volatile memory (NVM) device, domain-wall nanowire (or race-track) has shown potential for main memory storage but also computing capability. In this paper, the domain-wall nanowire is studied for a memory-based computing platform towards ultra-low-power big-data processing. One domain-wall nanowire based logic-in-memory architecture is proposed for big-data processing, where the domain-wall nanowire memory is deployed as main memory for data storage as well as XOR-logic for comparison and addition operations. The domain-wall nanowire based logic-in-memory circuits are evaluated by SPICE-level verifications. Further evaluated by applications of general-purpose SPEC2006 benchmark and also web-searching oriented Phoenix benchmark, the proposed computing platform can exhibit a significant power saving on both main memory and ALU under the similar performance when compared to CMOS based designs. Yuhao Wang 0002, Hao Yu 0001 |
ISLPED | 2 |
| 2013 | SRAM dynamic stability verification by reachability analysis with consideration of threshold voltage variationabstractDynamic stability margin of SRAM is largely suppressed at nano-scale due to not only dynamic noise but also process variation. A novel dynamic stability verification is developed in this paper based on analog reachability analysis for checking SRAM failure. In the presence of mismatch such as threshold voltage variation of all transistors, zonotope-based reachability analysis is deployed to efficiently verify SRAM failure at transistor level. The threshold voltage variation is considered by the modified input range of SRAM. As such, the suppressed stability margin and further failure region can be verified by performing a time-evolved reachability analysis with formed zonotope to distinguish safe and failure regions. One can perform efficient verification of the SRAM dynamic stability without repeated yet time-consuming Monte-Carlo simulations considering variations from all transistors. As demonstrated by numerical experiment results, the developed reachability analysis can accurately verify the SRAM dynamic stability under threshold voltage variations from all transistors. Speedup of more than 400x in runtime can be achieved over the Monte Carlo approach of 500 samples with the similar accuracy. Hao Yu 0001, Sai Manoj Pudukotai Dinakarrao, Guoyong Shi |
ISPD | 2 |
| 2013 | SPECO: Stochastic Perturbation based Clock tree Optimization considering temperature uncertainty
Sina Basir-Kazeruni, Hao Yu 0001, Fang Gong, Yu Hu 0002, Lei He 0001 |
Integr. | 2 |
| 2013 | An efficient channel clustering and flow rate allocation algorithm for non-uniform microfluidic cooling of 3D integrated circuits
Hanhua Qian, Chip-Hong Chang, Hao Yu 0001 |
Integr. | 3 |
| 2013 | Reliable 3-D Clock-Tree Synthesis Considering Nonlinear Capacitive TSV Model With Electrical-Thermal-Mechanical CouplingabstractA robust physical design of 3-D IC requires investigation on through-silicon via (TSV). The large temperatures and stress gradients can severely affect TSV delay with large variation. The traditional physical model treats TSV as a resistor with linear electrical-thermal dependence, which ignores the fundamental device physics. In this paper, a physics-based electrical-thermal–mechanical delay model is developed for signal TSVs in 3-D IC. With consideration of liner material and also stress, a nonlinear model is established between electrical delay with temperature and stress. Moreover, sensitivity analysis is performed to relate the reduction of temperature and stress gradients with respect to dummy TSVs insertion. Taking the design of 3-D clock tree as a case study, we have formulated a nonlinear optimization problem for clock-skew reduction. By allocating dummy TSVs to reduce the temperature and stress gradients, the clock skew introduced by signal TSVs and drivers can be minimized. A number of 3-D clock-tree benchmarks are utilized in experiments. We have observed that with the use of dummy TSV insertion, clock skew can be reduced by 61.3% on average when the accurate nonlinear electrical-thermal–mechanical delay model is applied. Sai Manoj Pudukotai Dinakarrao, Hao Yu 0001, Yang Shang, Chuan Seng Tan, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Stochastic Behavioral Modeling and Analysis for Analog/Mixed-Signal CircuitsabstractIt has become increasingly challenging to model the stochastic behavior of analog/mixed-signal (AMS) circuits under large-scale process variations. In this paper, a novel moment-matching-based method has been proposed to accurately extract the probabilistic behavioral distributions of AMS circuits. This method first utilizes Latin hypercube sampling coupling with a correlation control technique to generate a few samples (e.g., sample size is linear with number of variable parameters) and further analytically evaluate the high-order moments of the circuit behavior with high accuracy. In this way, the arbitrary probabilistic distributions of the circuit behavior can be extracted using moment-matching method. More importantly, the proposed method has been successfully applied to high-dimensional problems with linear complexity. The experiments demonstrate that the proposed method can provide up to 1666X speedup over crude Monte Carlo method for the same accuracy. Fang Gong, Sina Basir-Kazeruni, Lei He 0001, Hao Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | Fast simulation of hybrid CMOS and STT-MTJ circuits with identified internal state variablesabstractHybrid integration of CMOS and non-volatile memory (NVM) devices has become the technology foundation for emerging non-volatile memory based computing. The primary challenge to validate a hybrid system with both CMOS and non-volatile devices is to develop a SPICE-like simulator that can simulate the dynamic behavior of hybrid system accurately and efficiently. Since spin-transfer-toque magnetic-tunneling-junction (STT-MTJ) device is one of the most promising candidates of next generation NVM devices, it is under great interest in including this new device in the standard CMOS design flow. The previous approaches require complex equivalent circuits to represent the STT-MTJ device, and ignore dynamic effect without consideration of internal states. This paper proposes a new modified nodal analysis for STT-MTJ device with identified internal state variables. As demonstrated by a number of experiment examples on hybrid systems with both CMOS and STT-MTJ devices, our newly developed SPICE-like simulator can deal with the dynamic behavior of STT-MTJ device under arbitrary driving condition and reduce the CPU time by more than 20 times for memory circuits when compared to the previous equivalent circuit approaches. Yang Shang, Wei Fei, Hao Yu 0001 |
ASP-DAC | 3 |
| 2012 | A GPU-accelerated envelope-following method for switching power converter simulationabstractIn this paper, we propose a new envelope-following parallel transient analysis method for the general switching power converters. The new method first exploits the parallelisim in the envelope-following method and parallelize the Newton update solving part, which is the most computational expensive, in GPU platforms to boost the simulation performance. To further speed up the iterative GMRES solving for Newton update equation in the envelope-following method, we apply the matrix-free Krylov basis generation technique, which was previously used for RF simulation. Last, the new method also applies more robust Gear-2 integration to compute the sensitivity matrix instead of traditional integration methods. Experimental results from several integrated on-chip power converters show that the proposed GPU envelope-following algorithm leads to about 10× speedup compared to its CPU counterpart, and 100× faster than the traditional envelop-following methods while still keeps the similar accuracy. Xuexin Liu, Sheldon X.-D. Tan, Hai Wang 0002, Hao Yu 0001 |
DATE | 4 |
| 2012 | Fair energy resource allocation by minority game algorithm for smart buildingsabstractReal-time and decentralized energy resource allocation has become the main feature to develop for the next generation energy management system (EMS). In this paper, a minority game (MG)-based EMS (MG-EMS) is proposed for smart buildings with hybrid energy sources: main energy resource from electrical power-grid and renewable energy resource from solar photovoltaic (PV) cells. Compared to the traditional static and centralized EMS (SC-EMS), and the recent multi-agent-based EMS (MA-EMS) based on price-demand competition, our proposed MG-EMS can achieve up to 51× and 147× utilization rate improvements respectively regarding to the fairness of solar energy resource allocation. In addition, the proposed MG-EMS can also reduce peak energy demand for main power-grid by 30.6%. As such, one can significantly reduce the cost and improve the stability of micro-grid of smart buildings with a high utilization rate of solar energy. Chun Zhang 0003, Hantao Huang, Hao Yu 0001 |
DATE | 4 |
| 2012 | Distributed thermal-aware task scheduling for 3D Network-on-ChipabstractThe development of 3D integration technology significantly improves the bandwidth of network-on-chip (NoC) system. However, the 3D technology-enabled high integration density also brings severe concerns of temperature increase, which may impair system reliability and degrade the performance. Task scheduling has been regarded as one effective approach in eliminating thermal hotspot without introducing hardware overhead. However, centralized thermal-aware task scheduling algorithms for 3D-NoC have been limited for incurring high computational complexity as the system scale increase. In this paper, we propose a distributed agent-based thermal-aware task scheduling algorithm for 3D-NoC which shows high scheduling efficiency and high scalability. Experimental results have shown that when compared to the centralized algorithms, our algorithm can achieve up to 13 °C reduction in peak temperature of the system without sacrificing performance. Yingnan Cui, Wei Zhang 0012, Hao Yu 0001 |
ICCD | 3 |
| 2012 | Decentralized agent based re-clustering for task mapping of tera-scale network-on-chip systemabstractWith the rapid increasing demand for high-performance computing, such as cloud computing, Tera (flops) scale high-performance computing system composed of hundreds of on-chip processing cores has become the recent interest. Given a large-scale computing system such as network-on-chip (NoC) with hundreds of cores, bandwidth and power density are the fundamental limits dominated by on-chip communication. This has brought extreme challenge when mapping application tasks onto Tera-scale NoC system. Previous task mapping scheme is mainly centralized and static, and hence results in large communication volume, not scalable for runtime task mapping required by Tera-scale NoC system. In order to improve on-chip traffic and reduce power density for the need of Tera-scale NoC system, we have proposed a de-centralized re-clustering algorithm. The processing cores in the NoC system are organized into clusters with an efficient decentralized re-clustering scheme to adjust the cluster size for the task mapping. As such, the communication volume can be significantly reduced and result in decreased power. Experimental results have demonstrated that our proposed algorithm can achieve reduction of communication traffic (up to 66.7%). The energy consumption profile has also been efficiently improved to reduce the hotspots. Yingnan Cui, Wei Zhang 0012, Hao Yu 0001 |
ISCAS | 3 |
| 2012 | Design of low power 3D hybrid memory by non-volatile CBRAM-crossbar with block-level data-retentionabstractAs one of the newly introduced resistive random access memory (ReRAM) devices, this paper has shown an in-depth study of conductive-bridging random access memory (CBRAM) for non-volatile memory (NVM) computing. Firstly, a CBRAM-crossbar based memory is evaluated with accurate physical-level model and circuit-level characterization. It is then deployed as NVM component with a 3D hybrid integration of SRAM/DRAM, where one layer of CBRAM-crossbar is designed for data-retention under power gating to reduce leakage power from SRAM/DRAM at other layers. Moreover, a block-level data-retention scheme is designed to only write back dirty data from SRAM/DRAM to CBRAM-crossbar. When compared to the hybrid memory using phase-change random access memory (PCRAM) as data-retention, our CBRAM-based hybrid memory achieves 16x faster migration time and 4x less migration power for hibernating transition. When compared to the FeRAM-based bit-wise data-retention, our approach also achieves 17x smaller area and 8x smaller power under the same data migration speed. Yuhao Wang 0002, Chun Zhang 0003, Hao Yu 0001, Wei Zhang 0012 |
ISLPED | 3 |
| 2012 | Fast timing analysis of clock networks considering environmental uncertainty
Hai Wang 0002, Hao Yu 0001, Sheldon X.-D. Tan |
Integr. | 2 |
| 2012 | A Fast Non-Monte-Carlo Yield Analysis and Optimization by Stochastic Orthogonal PolynomialsabstractPerformance failure has become a significant threat to the reliability and robustness of analog circuits. In this article, we first develop an efficient non-Monte-Carlo (NMC) transient mismatch analysis, where transient response is represented by stochastic orthogonal polynomial (SOP) expansion under PVT variations and probabilistic distribution of transient response is solved. We further define performance yield and derive stochastic sensitivity for yield within the framework of SOP, and finally develop a gradient-based multiobjective optimization to improve yield while satisfying other performance constraints. Extensive experiments show that compared to Monte Carlo-based yield estimation, our NMC method achieves up to 700 X speedup and maintains 98% accuracy. Furthermore, multiobjective optimization not only improves yield by up to 95.3% with performance constraints, it also provides better efficiency than other existing methods. Fang Gong, Xuexin Liu, Hao Yu 0001, Sheldon X.-D. Tan, Junyan Ren, Lei He 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2012 | Design Exploration of Hybrid CMOS and Memristor Circuit by New Modified Nodal AnalysisabstractDesign of hybrid circuits and systems based on CMOS and nano-device requires rethinking of fundamental circuit analysis to aid design exploration. Conventional circuit analysis with modified nodal analysis (MNA) cannot consider new nano-devices such as memristor together with the traditional CMOS devices. This paper has introduced a new MNA method with magnetic flux (Φ) as new state variable. New SPICE-like circuit simulator is thereby developed for the design of hybrid CMOS and memristor circuits. A number of CMOS and memristor-based designs are explored, such as oscillator, chaotic circuit, programmable logic, analog-learning circuit, and crossbar memory, where their functionality, performance, reliability and power can be efficiently verified by the newly developed simulator. Specifically, one new 3-D-crossbar architecture with diode-added memristor is also proposed to improve integration density and to avoid sneak path during read-write operation. Wei Fei, Hao Yu 0001, Wei Zhang 0012, Kiat Seng Yeo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | A Parallel and Incremental Extraction of Variational Capacitance With Stochastic Geometric MomentsabstractThis paper presents a parallel and incremental solver for stochastic capacitance extraction. The random geometrical variation is described by stochastic geometrical moments, which lead to a densely augmented system equation. To efficiently extract the capacitance and solve the system equation, a parallel fast-multipole-method (FMM) is developed in the framework of stochastic geometrical moments. This can efficiently estimate the stochastic potential interaction and its matrix-vector product (MVP) with charge. Moreover, a generalized minimal residual (GMRES) method with incremental update is developed to calculate both the nominal value and the variance. Our overall extraction show is called piCAP. A number of experiments show that piCAP efficiently handles a large-scale on-chip capacitance extraction with variations. Specifically, a parallel MVP in piCAP is up 3 × to faster than a serial MVP, and an incremental GMRES in piCAP is up to 15× faster than non-incremental GMRES methods. Fang Gong, Hao Yu 0001, Lingli Wang, Lei He 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | A structured parallel periodic Arnoldi shooting algorithm for RF-PSS analysis based on GPU platformsabstractThe recent multi/many-core CPUs or GPUs have provided an ideal parallel computing platform to accelerate the time-consuming analysis of radio-frequency/millimeter-wave (RF/MM) integrated circuit (IC). This paper develops a structured shooting algorithm that can fully take advantage of parallelism in periodic steady state (PSS) analysis. Utilizing periodic structure of the state matrix of RF/MM-IC simulation, a cyclic-block-structured shooting-Newton method has been parallelized and mapped onto recent GPU platforms. We first present the formulation of the parallel cyclic-block-structured shooting-Newton algorithm, called periodic Arnoldi shooting method. Then we will present its parallel implementation details on GPU. Results from several industrial examples show that the structured parallel shooting-Newton method on Tesla's GPU can lead to speedups of more than 20× compared to the state-of-the-art implicit GMRES methods under the same accuracy on the CPU. Xuexin Liu, Hao Yu 0001, Jacob Relles, Sheldon X.-D. Tan |
ASP-DAC | 2 |
| 2011 | Fast non-monte-carlo transient noise analysis for high-precision analog/RF circuits by stochastic orthogonal polynomialsabstractStochastic device noise has become a significant challenge for high-precision analog/RF circuits, and it is particularly difficult to correctly include both white noise and flicker noise in the traditional transient verification with an efficient numerical solution. In this paper, a Non-Monte-Carlo transient noise analysis is developed. Both white noise and flicker noise are considered in Itô integral based stochastic differential algebraic equation (SDAE), which is solved by one-time calculation of variance using stochastic orthogonal polynomials (SoPs). Our work is the first in literature to provide the SoP-based SDAE solution with application for transient noise analysis. Experiments on a number of different analog circuits demonstrate that the proposed method is up to 488X faster than Monte Carlo method with similar accuracy, and achieves on average 6.8X speedup over the existing non-Monte-Carlo approaches. Fang Gong, Hao Yu 0001, Lei He 0001 |
DAC | 2 |
| 2011 | On the preconditioner of conjugate gradient method - A power grid simulation perspectiveabstractPreconditioned Conjugate Gradient (PCG) method has been demonstrated to be effective in solving large-scale linear systems for sparse and symmetric positive definite matrices. One critical problem in PCG is to design a good preconditioner, which can significantly reduce the runtime while keeping memory usage efficient. Universal preconditioners are simple and easy to construct, but their effectiveness is highly problem-dependent. On the other hand, domain-specific preconditioners that explore the underlying physical meaning of the matrices usually work better, but are difficult to design. In this paper, we study the problem in the context of power grid simulation, and develop a novel preconditioner based on the power grid structure through simple circuit simulations. Experimental results show 43% reduction in the number of iterations and 23% speedup over existing universal preconditioners. Chung-Han Chou, Nien-Yu Tsai, Hao Yu 0001, Che-Rung Lee, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
ICCAD | 3 |
| 2011 | Stochastic analog circuit behavior modeling by point estimation methodabstractStochastic device parameter variations have dramatically increased beyond the scale of 65nm and can significantly lead to large mismatch for analog circuits. To estimate unknown analog circuit behavior in performance space under the given stochastic variations in parameter space, many state-of-art approaches have been developed recently. However, either Gaussian distribution or response surface model (RSM) with analytical formulae has to be assumed when connecting performance space and parameter space. A novel point-estimation based approach has been proposed in this paper to capture arbitrary stochastic distributions for analog circuit behaviors in performance space. First, to evaluate high-order moments of circuit behavior in an accurate fashion, the point-estimation method has been applied with only a few number of simulations. Then, probability density function (PDF) of circuit behavior can be efficiently extracted by the obtained high-order moments. This method is further extended for multiple parameters under linear complexity. Extensive numerical experiments on a number of different circuits have demonstrated that the proposed point-estimation method can provide up to 181X runtime speedup with the same accuracy, when compared with Monte Carlo method. Moreover, it can further achieve up to 15X speedup over the RSM-based method such as APEX with the similar accuracy. Fang Gong, Hao Yu 0001, Lei He 0001 |
ISPD | 2 |
| 2010 | A fast analog mismatch analysis by an incremental and stochastic trajectory piecewise linear macromodelabstractTo cope with an increasing complexity when analyzing analog mismatch in sub-90nm designs, this paper presents a fast non-Monte-Carlo method to calculate mismatch in time domain. The local random mismatch is described by a noise source with an explicit dependence on geometric parameters, and is further expanded by stochastic orthogonal polynomials (SOPs). This forms a stochastic differential-algebra-equation (SDAE). To deal with large-scale problems, the SDAE is linearized at a number of snapshots along the nominal transient trajectory, and hence is naturally embedded into a trajectory-piecewise-linear (TPWL) macromodeling. The TPWL is improved with a novel incremental aggregation of subspaces identified at those snapshots. Experiments show that the proposed method, isTPWL, is hundreds of times faster than Monte-Carlo method with a similar accuracy. In addition, our macromodel further reduces runtime by up to 25X, and is faster to build and more accurate to simulate compared to existing approaches. Hao Yu 0001, Xuexin Liu, Hai Wang 0002, Sheldon X.-D. Tan |
ASP-DAC | 1 |
| 2010 | QuickYield: an efficient global-search based parametric yield estimation with performance constraintsabstractWith technology scaling down to 90nm and below, many yield-driven design and optimization methodologies have been proposed to cope with the prominent process variation and to increase the yield. A critical issue that affects the efficiency of those methods is to estimate the yield when given design parameters under variations. Existing methods either use Monte Carlo method in performance domain where thousands of simulations are required, or use local search in parameter domain where a number of simulations are required to characterize the point on the yield boundary defined by performance constraints. To improve efficiency, in this paper we propose QuickYield, a yield surface boundary determination by surface-point finding and global-search. Experiments on a number of different circuits show that for the same accuracy, QuickYield is up to 519X faster compared with the Monte Carlo approach, and up to 4.7X faster compared with YENSS, the fastest approach reported in literature. Fang Gong, Hao Yu 0001, Yiyu Shi 0001, Daesoo Kim, Junyan Ren, Lei He 0001 |
DAC | 2 |
| 2010 | A robust periodic arnoldi shooting algorithm for efficient analysis of large-scale RF/MM ICsabstractThe verification of large radio-frequency/millimeter-wave (RF/MM) integrated circuits (ICs) has regained attention for high-performance designs beyond 90nm and 60GHz. The traditional time-domain verification by standard Krylov-subspace based shooting method might not be able to deal with newly increased verification complexity. The numerical algorithms with small computational cost yet superior convergence are highly desired to extend designers' creativity to probe those extremely challenging designs of RF/MM ICs. This paper presents a new shooting algorithm for periodic RF/MM-IC systems. Utilizing a periodic structure of the state matrix, a periodic Arnoldi shooting algorithm is developed to exploit the structured Krylov-subspace. This leads to an improved efficiency and convergence. Results from several industrial examples show that the proposed periodic Arnoldi shooting method, called PAS, is 1000 times faster than the direct-LU and the explicit GMRES methods. Moreover, when compared to the existing industrial standard, a matrix-free GMRES with non-structured Krylov-subspace, the new PAS method reduces iteration number and runtime by 3 times with the same accuracy. Xuexin Liu, Hao Yu 0001, Sheldon X.-D. Tan |
DAC | 2 |
| 2010 | A new modified nodal analysis for nano-scale memristor circuit simulationabstractIt is unclear how to include the newly discovered memristor together with traditional electronic devices into a circuit simulator such as SPICE. To perform a fast prototyping of circuits composed of memristors at nano-scale, this paper introduces a new modified nodal analysis (MNA) to include this fourth circuit element into a first-order differential-algebra-equation (DAE). In the new MNA, magnetic flux is employed as the state variable for a flux-controlled memristor, called memductor. The new MNA is implemented in a circuit simulator to efficiently provide time-domain transient simulations for a number of nano-scale memristor-based circuits. Hao Yu 0001, Wei Fei |
ISCAS | 1 |
| 2010 | Fast Analysis of a Large-Scale Inductive Interconnect by Block-Structure-Preserved MacromodelingabstractAbstract—To efficiently analyze the large-scale interconnect dominant circuits with inductive couplings (mutual inductances), this paper introduces a new state matrix, called VNA, to stamp inverse-inductance elements by replacing inductive-branch current with flux. The state matrix under VNA is diagonal-dominant, sparse, and passive. To further explore the sparsity and hierarchy at the block level, a new matrix-stretching method is introduced to reorder coupled fluxes into a decoupled state matrix with a bordered block diagonal (BBD) structure. A corresponding block-structure-preserved model-order reduction, called BVOR, is developed to preserve the sparsity and hierarchy of the BBD matrix at the block level. This enables us to efficiently build and simulate the macromodel within a SPICE-like circuit simulator. Experiments show that our method achieves up to 7 faster modeling building time, up to 33 faster simulation time, and as much as 67 smaller waveform error compared to SAPOR [a second-order reduction based on nodal analysis (NA)] and PACT (a first-order 2 2 structured reduction based on modified NA). Index Terms—Circuit simulation, high-speed interconnect model, model-order reduction. I. Hao Yu 0001, Chunta Chu, Yiyu Shi 0001, David Smart, Lei He 0001, Sheldon X.-D. Tan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Fast analysis of nontree-clock network considering environmental uncertainty by parameterized and incremental macromodelingabstractIt is challenging to verify clock-skew for large-scale nontree clock network with environmental uncertainties such as supply voltage fluctuation and thermal temperature gradient. This paper presents a fast clock-skew analysis via parameterized incremental truncated-balanced-realization, called piTBR method. Environmental uncertainties are parametrically and structurally added into the state equation of clock network. A compact macromodel is obtained by the subspace projection constructed from the singular value decomposition (SVD) of circuit output waveforms. To reduce the computational cost, we propose an incremental SVD method that only needs to partially update the projection matrix by analyzing the perturbed output waveform owning to environmental uncertainties. Experiments on a number of clock networks show that compared with the macromodeling by the fast TBR method, our method reduces the computational cost in the order of 100× with a similar accuracy. In addition, compared with the macromodeling by the Krylov-subspace-based method, our method reduces the waveform error by 2× with a similar runtime. Hai Wang 0002, Hao Yu 0001, Sheldon X.-D. Tan |
ASP-DAC | 2 |
| 2009 | PiCAP: a parallel and incremental capacitance extraction considering stochastic process variationabstractIt is unknown how to include stochastic process variation into fast-multipole-method (FMM) for a full chip capacitance extraction. This paper presents a parallel FMM extraction using stochastic polynomial expanded geometrical moments. It utilizes multi-processors to evaluate in parallel for the stochastic potential interaction and its matrix-vector product (MVP) with charge. Moreover, a generalized minimal residual (GMRES) method with deflation is modified to incrementally consider the nominal value and the variance. The overall extraction flow is called piCAP. Experiments show that the parallel MVP in piCAP is up to 3X faster than the serial MVP, and the incremental GMRES in pi-CAP is up to 15X faster than non-incremental GMRES methods. Fang Gong, Hao Yu 0001, Lei He 0001 |
DAC | 2 |
| 2009 | Allocating power ground vias in 3D ICs for simultaneous power and thermal integrityabstractThe existing work on via allocation in 3D ICs ignores power/ground vias' ability to simultaneously reduce voltage bounce and remove heat. This article develops the first in-depth study on the allocation of power/ground vias in 3D ICs with simultaneous consideration of power and thermal integrity. By identifying principal ports and parameters, effective electrical and thermal macromodels are employed to provide dynamic power and thermal integrity as well as sensitivity with respect to via density. With the use of sensitivity, an efficient via allocation simultaneously driven by power and thermal integrity is developed. Experiments show that, compared to sequential power and thermal optimization using static integrity, sequential optimization using the dynamic integrity reduces nonsignal vias by up to 18%, and simultaneous optimization using dynamic integrity further reduces nonsignal vias by up to 45.5%. Hao Yu 0001, Joanna Ho, Lei He 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2008 | Thermal Via Allocation for 3-D ICs Considering Temporally and Spatially Variant Thermal PowerabstractThe existing 3-D thermal-via allocation methods are based on the steady-state thermal analysis and may lead to excessive number of thermal vias. This paper develops an accurate and efficient thermal-via allocation considering the temporally and spatially variant thermal-power. The transient temperature is calculated by macromodel with a one-time structured and parameterized model reduction, which also generates temperature sensitivity with respect to thermal-via density. The proposed thermal-via allocation minimizes the time-integral of temperature violation, and is solved by a sequential quadratic programming algorithm with use of sensitivities from the macromodel. Compared to the existing method using the steady-state thermal analysis, our method in experiments is 126$\times$faster to obtain temperature, and reduces the number of thermal vias by 2.04$\times$under the same temperature bound. Hao Yu 0001, Yiyu Shi 0001, Lei He 0001, Tanay Karnik |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | Off-chip Decoupling Capacitor Allocation for Chip Package Co-DesignabstractOff-chip decoupling capacitor (decap) allocation is a demanding task during package and chip codesign. Existing approaches can not handle large numbers of I/O counts and large numbers of legal decap positions. In this paper, we propose a fast decoupling capacitor allocation method. By applying a spectral clustering, a small amount of principal I/Os can be found. Accordingly, the large power supply network is partitioned into several blocks each with only one principal I/O. This enables a localized macromodeling for each block by a triangular-structured reduction. In addition, to systemically consider a large legal position map in a manageable fashion, the map of legal positions is decomposed into multiple rings, which are further parameterized in each block. The decaps are then allocated according to the sensitivity obtained from the parameterized macro-model for each block. Compared to the PRIMA-based macromodeling, experiments show that our method (TBS2) is 25X faster and has 3.04X smaller error. Moreover, our decap allocation reduces the optimization time by 97X, and reduces decap cost by up to 16% to meet the same power-integrty target. Hao Yu 0001, Chunta Chu, Lei He 0001 |
DAC | 1 |
| 2007 | Minimal skew clock embedding considering time variant temperature gradientabstractThe existing temperature-aware clock embedding assumes a time-invariant temperature gradient. However, it is not solved how to find the worst-case temperature gradient leading to the worst case skew. In this paper, we develop a PErturbation based Clock Optimization (PECO) considering the timevariant temperature gradient. For a given clock topology, we minimize the worst case skew without asking for the worst case temperature map. We decide the merging point level by level based on the sensitivity of the skew with respect to the change of merging point. Such sensitivity is calculated using a parameterized model, which is compressed by a singularvalue-decomposition (SVD) and K-means based clustering considering the temperature correlation. The experimental results show that our algorithm reduces worst-case skew by up to 5X compared to the existing zero skew based ZST/DME method with small (up to 1%) wirelength overhead. Hao Yu 0001, Yu Hu 0002, Lei He 0001 |
ISPD | 1 |
| 2007 | Circuit-simulated obstacle-aware Steiner routingabstractThis article develops circuit-simulated routing algorithms. We model the routing graph by an RC network with terminals as inputs, and show that the faster an output reaches its peak, the higher the possibility for the corresponding Hanan or escape node to become a Steiner point. This enables us to select Steiner points and then apply any minimum spanning tree algorithm to obtain obstacle-free or obstacle-aware Steiner routing. Compared with existing algorithms, our algorithms have significant gain on either wirelength or runtime for obstacle-free routing, and on both wirelength and runtime for obstacle-aware routing. Yiyu Shi 0001, Paul Mesa, Hao Yu 0001, Lei He 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2006 | Circuit simulation based obstacle-aware Steiner routingabstractSteiner routing is a fundamental yet NP-hard problem in VLSI design and other research fields. In this paper, we propose to model the routing graph by an RC network with routing terminals as input ports and Hanan nodes as output ports. We show that the faster an output reaches its peak, the higher the possibility for the correspondent Hanan node to be a Steiner point. Iteratively adding one or multiple selected Steiner points to build and improve Steiner trees leads to 1-cktSteiner and Blocked-cktSteiner (in short, B-cktSteiner) algorithms, respectively. When there are no routing obstacles, 1-cktSteiner obtains similar wirelength compared with the best existing algorithm FastSteiner. Both are less than 1% worse than the exact solution, but 1-cktSteiner is up to 11.3X faster than FastSteiner. Compared with the fastest existing heuristic FLUTE, B-cktSteiner has similar runtime but up to 1.9% shorter wirelength. Different from FastSteiner and FLUTE which are only applicable to non-obstacle cases, 1-cktSteiner and B-cktSteiner can be applied to routing with obstacles with minimal runtime increase. Compared with the best existing obstacle-avoiding algorithm An-OARSMan, 1-cktSteiner has similar runtime and reduces wirelength by 6.12%, and B-cktSteiner has an average speedup of 352X with a similar wirelength. Yiyu Shi 0001, Paul Mesa, Hao Yu 0001, Lei He 0001 |
DAC | 3 |
| 2006 | Fast analysis of structured power grid by triangularization based structure preserving model order reductionabstractIn this paper, a Triangularization Based Structure preserving (TBS) model order reduction is proposed to verify power integrity of on-chip structured power grid. The power grid is represented by interconnected basic blocks according to current density, and basic blocks are further clustered into compact blocks, each with a unique pole distribution. Then, the system is transformed into a triangular system, where compact blocks are in its diagonal andthe system poles are determined only by the diagonal blocks. Finally, projection matrices are constructed and applied for compact blocks separately. The resulting macromodel has more matched poles and is more accurate than the one using flat projection. It is also sparse and enables a two-level analysis for simulation time reduction. Compared to existing approaches, TBS in experiments achieves up to 133X and 109X speedup in macromodel buildingand simulation respectively, and reduces waveform error by 33X. Hao Yu 0001, Yiyu Shi 0001, Lei He 0001 |
DAC | 1 |
| 2006 | Simultaneous power and thermal integrity driven via stapling in 3D ICsabstractThe existing work on via-stapling in 3D integrated circuits optimizes power and thermal integrity separately and uses steadystate thermal analysis. This paper presents the first in-depth study on simultaneous power and thermal integrity driven viastapling in 3D design. The transient temperature and supply voltage violations are calculated by a structured and parameterized model reduction, which also generates parameterized temperature and voltage violation sensitivities with respect to the via pattern and density. Using parameterized sensitivities, an efficient yet effective greedy optimization is presented to optimize power and thermal integrity simultaneously. Experiments with two active device layers show that compared to sequential power and thermal optimization using steady-state thermal analysis, sequential optimization using transient thermal analysis reduces non-signal vias by on average 11.5%, and simultaneous optimization using transient thermal analysis reduces non-signal vias by on average 34%. The via reduction would be higher for the 3D design with more device layers. Hao Yu 0001, Joanna Ho, Lei He 0001 |
ICCAD | 1 |
| 2006 | A fast block structure preserving model order reduction for inverse inductance circuitsabstractMost existing RCL-1 circuit reductions stamp inverse inductance L-1 elements by a second-order nodal analysis (NA). The NA formulation uses nodal voltage variables and describes inductance by nodal susceptance. This leads to a singular matrix stamping in general. We introduce a new circuit stamping for RCL-1 circuits using branch vector potentials. The new circuit stamping results in a first-order circuit matrix that is semi-positive definite and non-singular. We call this as vectorpotential based nodal analysis (VNA). It enables an accurate and passive reduction. In addition, to preserve the structure of state matrices such as sparsity and hierarchy, we represent the flat VNA matrix in a bordered-block diagonal (BBD) form. This enables us to build and simulate the macromodel efficiently. In experiments performed on several test cases, our method achieves up to 15X faster modeling building time, up to 33X faster simulation time, and as much as 67X smaller waveform error compared to SAPOR, the best existing second order RCL-1 reduction method. Hao Yu 0001, Yiyu Shi 0001, Lei He 0001, David Smart |
ICCAD | 1 |
| 2006 | Thermal via allocation for 3D ICs considering temporally and spatially variant thermal powerabstractAll existing methods for thermal-via allocation are based on a steady-state thermal analysis and may lead to excessive number of thermal vias. This paper develops an accurate and efficient thermal-via allocation considering temporally and spatially variant thermal-power. The transient temperature is calculated using macromodel by a structured and parameterized model reduction, which generates temperature sensitivity with respect to thermal-via density. By defining a thermal-violation integral based on the transient temperature, a nonlinear optimization problem is formulated to allocate thermal-vias and minimize thermal violation integral. This optimization problem is transformed into a sequence of subproblems by Lagrangian relaxation, and each subproblem is solved by quadratic programming using sensitives from the macromodel. Experiments show that compared to the existing method using steady-state thermal analysis, our method is 126X faster to obtain the temperature profile, and reduces the number of thermal vias by 2.04X under the same temperature bound. Hao Yu 0001, Yiyu Shi 0001, Lei He 0001, Tanay Karnik |
ISLPED | 1 |
| 2006 | SAMSON: a generalized second-order arnoldi method for reducing multiple source linear network with susceptanceabstractPower integrity analysis of in-package and on-chip power supply needs to consider a large number of ports and handle magnetic coupling that is better represented by susceptance. The existing moment matching methods are not able to accurately model both large number of ports and susceptance. In this paper, we propose a generalized Second-order Arnoldi method for reducing Multiple Source Linear Network (SAMSON) with susceptance. We employ a right-hand-side excitation current vector to replace the port incident matrix such that an MIMO (Multiple-input-multiple-output) system is transformed into an equivalent superposed SIMO (Single-input-multiple-output) system to avoid accuracy loss in block moment matching, and develop a generalized second-order Arnoldi method based orthonormalization to accurately handle susceptance and non-impulse current sources. Compared with existing EKS and IEKS approaches able to consider non-impulse sources but not susceptance, SAMSON is slightly faster and is more accurate in high frequency range and at dc. With same model order, SAMSON reduces time domain waveform error by 33X compared to EKS/IEKS and by 47X compared with the best block moment matching method applicable to susceptance. Yiyu Shi 0001, Hao Yu 0001, Lei He 0001 |
ISPD | 2 |
| 2006 | Wideband passive multiport model order reduction and realization of RLCM circuitsabstractThis paper presents a novel compact passive modeling technique for high-performance RF passive and interconnect circuits modeled as high-order resistor-inductor-capacitor-mutual inductance circuits. The new method is based on a recently proposed general s-domain hierarchical modeling and analysis method and vector potential equivalent circuit model for self and mutual inductances. Theoretically, this paper shows that s-domain hierarchical reduction is equivalent to implicit moment matching at around s=0 and that the existing hierarchical reduction method by one-point expansion is numerically stable for general tree-structured circuits. It is also shown that hierarchical reduction preserves the reciprocity of passive circuit matrices. Practically, a hierarchical multipoint reduction scheme to obtain accurate-order reduced admittance matrices of general passive circuits is proposed. A novel explicit waveform-matching algorithm is proposed for searching dominant poles and residues from different expansion points based on the unique hierarchical reduction framework. To enforce passivity, state-space-based optimization is applied to the model order reduced admittance matrix. Then, a general multiport network realization method to realize the passivity-enforced reduced admittance based on the relaxed one-port network synthesis technique using Foster's canonical form is proposed. The resulting modeling algorithm can generate the multiport passive SPICE-compatible model for any linear passive network with easily controlled model accuracy and complexity. Experimental results on an RF spiral inductor and a number of high-speed transmission line circuits are presented. In comparison with other approaches, the proposed reduction is as accurate as passive reduced-order interconnect macromodeling algorithm in the high-frequency domain due to the enhanced multipoint expansion, but leads to smaller realized circuit models. In addition, under the same reduction ratio, realized models by the new method have less error compared with reduced circuits by time-constant-based reduction techniques in time domain. Zhenyu Qi 0002, Hao Yu 0001, Pu Liu, Sheldon X.-D. Tan, Lei He 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Wideband modeling of RF/Analog circuits via hierarchical multi-point model order reductionabstractThis paper proposes a novel wideband modeling technique for high-performance RF passives and linear(ized) analog circuits. The new method is based on a recently proposed sdomain hierarchical modeling and analysis method [27]. Theoretically, we show that the s-domain hierarchical reduction is equivalent to implicit moment matching around s = 0, and that the existing hierarchical reduction method by one-point expansion is numerically stable for general tree-structured circuits. Practically, we propose a hierarchical multi-point reduction scheme for high-fidelity, wideband modeling of general passive or active linear circuits. A novel explicit waveform matching algorithm is proposed for searching the dominant poles and residues from different expansion points based on the unique hierarchical reduction framework. Experimental results with large analog circuits, on-chip spiral inductors are presented to validate the proposed method. Zhenyu Qi 0002, Sheldon X.-D. Tan, Hao Yu 0001, Lei He 0001 |
ASP-DAC | 3 |
| 2005 | A wideband hierarchical circuit reduction for massively coupled interconnectsabstractWe develop a realizable circuit reduction to generate the interconnect macro-model for parasitic estimation in wideband applications. The inductance is represented by VPEC (vector potential equivalent circuit) model, which not only enables the passive sparsification but also gives correct low-frequency response, whereas the recent circuit reduction intrinsically has inaccurate value and low-frequency response due to nodal-susceptance formulation. Applying hierarchical circuit-reduction enhanced by multi-point expansions, we can obtain an accurate high-order impedance function to capture the high-frequency response. The impedance function is further enforced passivity by convex programming, and realized by a Foster's synthesis. Experiments show that our method is as accurate as PRIMA in high frequency range, but leads to a realized circuit model with up to 10X times less complexity and up to 8X smaller simulation time. In addition, under the same reduction ratio, its error margin is less than that for the time-constant based reduction in both time-domain and frequency-domain simulations. Hao Yu 0001, Lei He 0001, Zhenyu Qi 0002, Sheldon X.-D. Tan |
ASP-DAC | 1 |
| 2005 | A provably passive and cost-efficient model for inductive interconnectsabstractTo reduce the model complexity for inductive interconnects, the vector potential equivalent circuit (VPEC) model was introduced recently and a localized VPEC model was developed based on geometry integration. In this paper, the authors show that the localized VPEC model is not accurate for interconnects with nontrivial sizes. They derive an accurate VPEC model by inverting the inductance matrix under the partial element equivalent circuit (PEEC) model and prove that the effective resistance matrix under the resulting full VPEC model is passive and strictly diagonal dominant. This diagonal dominance enables truncating small-valued off-diagonal elements to obtain a sparsified VPEC model named truncated VPEC (tVPEC) model with guaranteed passivity. To avoid inverting the entire inductance matrix, the authors further present another sparsified VPEC model with preserved passivity, the windowed VPEC (wVPEC) model, based on inverting a number of inductance submatrices. Both full and sparsified VPEC models are SPICE compatible. Experiments show that the full VPEC model is as accurate as the full PEEC model but consumes less simulation time than the full PEEC model does. Moreover, the sparsified VPEC model is orders of magnitude (1000/spl times/) faster and produces a waveform with small errors (3%) compared to the full PEEC model, and wVPEC uses less (up to 90/spl times/) model building time yet is more accurate compared to the tVPEC model. Hao Yu 0001, Lei He 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2003 | Vector potential equivalent circuit based on PEEC inversionabstractThe geometry-integration based vector potential equivalent circuit (VPEC) was introduced to obtain a localized circuit model for inductive interconnects in [1]. In this paper, we show that the method in [1] is accurate only for the two body problem. We derive N-body VPEC models based on geometry integration and inversion of inductance matrix under the PEEC model, respectively. Both VPEC models are derived from first principles and are accurate compared to the full PEEC model. The resulting circuit matrix G can be analyzed directly by existing simulation tools such as SPICE, and the simulation time of VPEC model is 47X less than that for PEEC model for a bus structure with 256 wires. It is also passive and strictly diagonal dominant, which leads to efficient circuit sparsification methods such as numerical and geometry based sparsifications. Compared to the full PEEC model, the sparsified VPEC models are orders of magnitude faster and produce waveforms with very small error. Hao Yu 0001, Lei He 0001 |
DAC | 1 |