Yang Guo 0003

dblp:73/1810-3 · DBLP profile ↗
← Back
76ranked-venue papers
3as first author
35since 2021 · last 2026
0000-0001-9050-0866ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 61 · 1 first-author · 31 since 2021Software engineering, systems software and programming languages · 7 · 3 since 2021Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Row-wise Inter-Phase Pipelining for Hardware-Efficient GCN Acceleration
abstract
Graph Convolutional Networks (GCNs) are widely deployed for learning on graph-structured data. Yet, their canonical two-phase execution, consisting of feature transformation followed by neighborhood aggregation, suffers from severe inefficiencies on existing accelerators. The prevailing phase-decoupled model explicitly stores dense intermediate results to off-chip memory after transformation and reloads them for aggregation, incurring substantial inter-phase data movement. Additionally, irregular access patterns cause intra-phase redundancy through repeated fetches of sparse inputs. To overcome both forms of redundancy, we propose IPRS-GCN, a streaming accelerator built on Inter-Phase Row-Streaming (IPRS). IPRS fuses transformation and aggregation at row granularity into a single pipeline, forwarding each node’s transformed features directly to its neighbors via the Row Queue as soon as they are computed. This eliminates off-chip storage of intermediates and enforces single-pass streaming access to inputs while keeping weights resident on-chip. The architecture realizes this dataflow through lightweight mechanisms including conflict-aware non-zero packing and degree-adaptive buffering, achieving near-ideal compute utilization with minimal on-chip footprint. Evaluated across five real-world datasets, IPRS-GCN achieves a geometric mean speedup of 107.74 × over an NVIDIA V100 GPU, outperforms HyGCN and AWB-GCN by 7.74 × and 2.12 ×, respectively, and improves energy efficiency by 1.37 ×.
Junsheng Chang, Yang Guo 0003, Li Shen 0007
CF3
2026 Enabling Ultra-Reliable Memories: A Practical Framework for Zero Mis-correction SEC-DED-DAEC Codes for Safety-Critical Systems
abstract
In safety-critical systems such as autonomous driving and aerospace, memory reliability standards are evolving from "high-reliability" to "ultra-reliability," demanding the eradication of all foreseeable, deterministic failure modes. To address the prevalent challenge of Double Adjacent Errors (DAE) induced from radiation, the design of SEC-DED-DAEC codes faces a critical dilemma: efficient but flawed Hsiao-based codes that risk miscorrection, versus correct-by-construction but costly and inflexible OLS-based codes. This trade-off between efficiency and correctness presents a key barrier to designing ultra-reliable systems. To resolve this impasse, this paper introduces MCTS-CDB, a novel Computer-Aided Design (CAD) framework. By integrating a CDCL-inspired search with Monte Carlo Tree Search (MCTS) guidance, it systematically constructs codes that achieve a zero-miscorrection guarantee within the highly-efficient Hsiao architecture. Experimental results validate our approach, showing that compared to a wide range of existing schemes, our generated codes achieve the correctness while reducing average encoding and decoding delays by 24.15% and 13.66%. This work provides a practical solution for designing the ultra-reliable memory subsystems required by next-generation safety-critical applications.
Guixiang Chen, Sheng Liu 0001, Yang Guo 0003
DATE4
2026 3D Chiplet Partitioning and Floorplanning Interaction with Vertical Bonding Consideration
abstract
The emerging technologies of 3D chiplet integration offer a promising path to increase functional density and communication efficiency beyond the limitations of traditional 2D designs. Among various methods, hybrid bonding enables high-density vertical connections with reduced parasitics and latency. However, few existing approaches explicitly support this vertical interconnect scenario, and most treat partitioning and floorplanning as separate stages. To fully exploit the benefits of vertical integration, partitioning and floorplanning need to be considered jointly to balance bond demand and bond supply. In this paper, we propose a unified framework for 3D chiplet partitioning and floorplanning under fine-pitch bonding technologies. Experiments on industry-standard benchmarks demonstrate a 10%∼40% reduction in HPWL and over 60% decrease in inter-die vertical connection overflow, confirming the effectiveness and superiority of our approach over current methods.
Mengen Chen, Yao Wang 0002, Yang Guo 0003
DATE5
2026 CACWS: Congestion-Aware Coordinated Warp Scheduler for Partitioned GPGPU
Sheng Liu 0001, Yang Guo 0003, Jianfeng Cui, Zekun Jiang
IPDPS3
2026 WSGraph: A Framework for Tackling Redundant and Irregular Data Access in Streaming Graph Processing
abstract
The demand for real-time streaming graph analysis has grown significantly, as hundreds of thousands of updates come every second. Monotonic graph algorithms such as Shortest Path are widely used in real-time analytics, but there are two bottlenecks that limit their performance, specially on planar graphs. One is massive redundant data accesses due to irregular state propagations and the other is high memory latency caused by irregular data accesses. We observe that existing systems mainly focus on general-purpose graph algorithms. If the properties of specific graph algorithms are exploited, the analysis performance can be further improved. Moreover, these systems typically tackle these two bottlenecks separately through either software or hardware mechanisms, but not both. However, both bottlenecks need to be addressed simultaneously in real scenarios such as road navigation. This article proposes WSGraph, a software-hardware co-design framework for high-performance streaming graph processing. WSGraph tackles these two challenges by enforcing regularized processing orders and enabling precise data prefetching. Specifically, at the software level, WSGraph integrates a priority-based work scheduler with sliding-window bucket mapping scheme to regulate state propagations, thereby drastically reducing redundant data accesses. At the hardware level, WSGraph incorporates a lightweight in-core Proactive Data Engine (PDE). By exploiting intra-vertex access regularity, the PDE accurately prefetches relevant graph data to effectively hide the high latency of irregular memory accesses. Experimental results demonstrate that WSGraph achieves significant performance improvements over existing systems. Compared with the state-of-the-art software system KickStarter, WSGraph gains a 2.13× speedup primarily by reducing graph data accesses by an average of 78.6%.
Xuanyi Li, Chen Li 0015, Zhengyi Dai, Jianzhuang Lu, Yang Guo 0003
ACM Trans. Archit. Code Optim.6
2026 A Low-Overhead SEU Hardening Method for CML Latch in High-Speed Frequency Divider Circuits
abstract
The increasing susceptibility of aerospace integrated circuits (ICs) to single-event effects (SEEs) demands advanced radiation-hardening-by-design (RHBD) methodologies, especially for analog/mixed-signal circuits operating in extreme environments. While digital circuit hardening has seen extensive advancements, circuit-level radiation hardening techniques specifically targeting analog building blocks—such as high-speed current-mode logic (CML) latches used in frequency dividers—have received relatively less attention, largely due to the inherent trade-offs between radiation tolerance and high-speed performance. This work proposes a low-overhead RHBD strategy for CML latches in high-speed dividers, addressing single-event upsets (SEUs) through the integration of pMOS transistors at storage nodes to counteract charge collection induced by SEEs. The methodology is rigorously validated using calibrated double-exponential current source simulations, emulating SEEs. The hardened design demonstrates robust radiation tolerance across varied pMOS control signals and sizing parameters, with minimal sensitivity introduced at additional nodes. Postlayout simulations in sub-20 nm FinFET technology reveal a maximum operating frequency of 46 GHz, accompanied by ultralow implementation overheads: a 2% area penalty and 22% power increase. The strategy resolves critical trade-offs between radiation resilience and high-speed performance, offering a scalable solution for aerospace-grade circuits in radiation-intensive environments. This advancement underscores the viability of tailored RHBD approaches for analog/mixed-signal systems, enabling next-generation space electronics that harmonize reliability with cutting-edge functionality.
Yahao Fang, Yaqing Chi, Deng Luo, Hanhan Sun, Guofang Yu, Yang Guo 0003
IEEE Trans. Very Large Scale Integr. Syst.8
2025 SI-Aware Wire Timing Prediction at Pre-Routing Stage with Multi-Corner Consideration
abstract
Timing closure is a critical but effort-taking task in VLSI designs. Early design stages have relatively ample room for changes that can fix timing problems in a proactive manner. However, accurate timing prediction is very challenging at early stages due to the absence of information determined by later stages in the design flow. At the pre-routing stage, it is generally believed that the prediction of wire delay is more complicated than that of gate delay, since the former is highly dependent on the routing information and PVT conditions. Addressing that, in this work, the prediction model is studied and the importance of multiple features, including Signal-Integrity (SI) related ones, is explored, with the purpose of boosting the turnaround time of physical design and reducing the performance penalty caused by worst-case scenario assumptions. Experimental results show that the proposed timing predictor has achieved a correlation of over 0.98 with the sign-off timing results in SI-mode under multi-corner scenarios.
Renjun Zhao, Yao Wang 0002, Chang Liu 0019, Yang Guo 0003
ASP-DAC6
2025 LintLLM: An Open-Source Verilog Linting Framework Based on Large Language Models
Zhigang Fang 0002, Renzhi Chen, Yang Guo 0003, Huadong Dai, Lei Wang 0011
ACM Great Lakes Symposium on VLSI4
2025 LAMP: A Locality-Aware TLB Sharing Framework for Multi-Tenant GPUs via TLB Comprehensive Profiling
abstract
With the growing prevalence of cloud services, GPUs are increasingly shared across multiple tenants, making address translation a performance-critical path. Our study makes three key observations: 1) Shared L2 TLB becomes the bottleneck in multi-tenant systems. Intense competition for entries can lead to high miss rates and thrashing under heavy pressure. 2) Benchmarks vary significantly in TLB sensitivity due to differing access patterns and memory reuse potential. 3) The prevailing management strategies, including free-for-all sharing and static partitioning, fail to account for tenant access behavior. And the state-of-the-art token-based strategy relies solely on miss rate, an inadequate metric that fails to capture per-tenant access patterns. Based on these observations, we propose LAMP, a localityaware TLB sharing framework. LAMP comprehensively profiles the access pattern considering three key aspects: spatial locality, temporal locality, and access density. These aspects allow LAMP to assess each tenant's sensitivity to TLB capacity and their reuse potential. Based on these insights, LAMP optimizes the sharing by dynamically partitioning L2 TLB, granting more entries to reuse-efficient tenants while limiting wasteful occupancy. This reduces thrashing and cross-tenant interference. Experimental results show that LAMP improves system performance by$\mathbf{1 2. 3 \%}$on average across a range of multi-tenant workloads, with negligible hardware overhead.
Chen Li 0015, Xuanyi Li, Jianzhuang Lu, Yang Guo 0003
HPCC6
2025 SICA: A Multicore Neuromorphic Processor Featuring Sparse Integration and Communication-Aware Optimization
abstract
Neuromorphic computing has emerged as a promising paradigm due to its event-driven operation and energy efficiency, driving extensive research in neuromorphic processor development. When implementing spiking neural networks (SNNs) on such processors, two critical aspects must be addressed: neuron computation and spike communication. For neuron computation, previous work primarily relies on parallel accumulation via adder trees but fails to leverage the inherent sparsity in SNNs. For spike communication, conventional mesh topologies suffer from long-distance communication inefficiencies, while suboptimal mapping strategies further exacerbate latency issues. To address these challenges, we propose a low-overhead fast sparse detection mechanism that effectively exploits spike sparsity and optimizes the processor's workflow, thereby achieving efficient synaptic integration with minimal overhead. For spike communication, we employ an on-chip broadcast mechanism combined with a hybrid torus-mesh topology to significantly reduce communication latency, while systematically evaluating the impact of three distinct mapping strategies-random, sequential, and communication-aware mapping-on overall performance. Experimental results demonstrate significant improvements, with our solution delivering speedups of$1.26 \times$and$1.24 \times$compared to LSMCore on the N-MNIST and MNIST datasets, respectively. Furthermore, the communication-aware mapping strategy achieves a 24.87% reduction in communication latency, while the torus topology contributes an additional$\mathbf{1 4. 4 8} \boldsymbol{\%}$latency reduction.
Junbo Tie, Xun Xiao, Yuanfeng Luo, Yang Guo 0003, Lei Wang 0011
HPCC12
2025 A Multi-objective Mapping Optimization Strategy for Large-Scale NoC Design
Chen Li 0015, Jianzhuang Lu, Yang Guo 0003
ICA3PP (5)5
2025 RTLBench: A Multi-Dimensional Benchmark Suite for Evaluating LLM-Generated RTL Code
abstract
The rapid advancement of large language models (LLMs) has enabled automated Register Transfer Level (RTL) code generation, accelerating chip design workflows. However, existing benchmarks focus mainly on syntax and functionality, overlooking critical engineering aspects such as lint compliance, readability, and coding style. To address this gap, we propose RTLBench, a benchmark suite of 160 copyright-free RTL cases sourced from textbooks and open-source projects. RTLBench features a multi-dimensional evaluation framework covering syntax, functionality, lint compliance, readability, and style consistency. To assess subjective code quality metrics, it also incorporates an LLM-as-a-judge mechanism. We evaluated 24 state-of-the-art LLMs using RTLBench, finding that while several models perform well in syntax and functionality, most fall short on engineering quality. To address this, we propose Log2BetterRTL, a log-driven feedback system that transforms EDA tool diagnostics into iterative improvement prompts. It improves syntax correctness by up to 18.13 %, boosts functional correctness by 14.38 %, reduces lint violations by up to 229, and raises clarity scores by 0.51. These results demonstrate RTLBench's effectiveness in evaluating and enhancing LLMgenerated RTL, bridging the gap between generative AI and industrial-grade hardware design. The suite and scripts are available at: https://fangzhigang32.github.io/RTLBench.
Zhigang Fang 0002, Renzhi Chen, Yang Guo 0003, Huadong Dai, Lei Wang 0011
ICCD3
2025 AICAWS: Arithmetic Intensity Based Cache-Conscious Adaptive Warp Scheduler
abstract
General-Purpose Graphics Processing Units (GPGPUs) are crucial for parallel computing in artificial intelligence and big data with their performance heavily relying on efficient warp scheduling. Traditional schedulers, such as Round-Robin (RR) and Greedy-Then-Oldest (GTO), employ static strategies that struggle with adapting to diverse workloads, causing performance disparities across different applications. Prior work has focused on aspects like critical warps and memory access locality but has often overlooked the arithmetic intensity of workloads. Drawing inspiration from the Roofline model and recognizing that different workloads exhibit distinct computational intensities, we propose an Arithmetic Intensity based CacheConscious Adaptive Warp Scheduler (AICAWS). It operates by first analyzing the kernel's static arithmetic intensity through compiler, which serves as a baseline for the hardware. Subsequently, during warp execution, AICAWS dynamically monitors the warp's execution progress, analyzes its runtime arithmetic intensity, and adjusts warp scheduling strategies based on this. Furthermore, AICAWS considers cache locality during warp execution, enabling fine-grained classification of warps based on this. This synergistic mechanism enables AICAWS to effectively hide long-latency memory access operations. Evaluations on diverse benchmarks demonstrate that AICAWS achieves an average performance improvement of 26.3% compared to the baseline scheduler, with a peak improvement of 77.9%.
Sheng Liu 0001, Zekun Jiang, Jianfeng Cui, Yang Guo 0003
ICCD5
2025 Super Microscaling: Enhancing Precision and Hardware Efficiency in Deep Learning Quantization
abstract
The Microscaling (MX) data format, an state-of-the-art quantization technique for deep learning tensor operations, suffers from precision loss at low bit-widths and potential hardware overhead. To overcome these limitations, we propose the Super Microscaling (SMX) data format and its accompanying hardware engine. SMX ensures higher accuracy by optimizing the conversion from Floating Point (FP) to MX, particularly in low-bit width scenarios. We also present spatial-temporal reuse strategy and two-level dequantization hardware architecture, which boosts efficiency in terms of hardware area and energy consumption. Experimental results demonstrate that SMX outperforms MX, achieving average accuracy improvements of 55.86% for BERT, 40.89% for ResNet50, and 37.51% for VGG, while reducing hardware area and energy consumption by 22.74% and 25.46%, respectively.
Zhuang Cao, Xin Ju 0005, Mei Wen, Yang Guo 0003
ISCAS5
2025 A perspective on digital signal processor based leadership performance accelerator for AI and HPC
Yang Guo 0003, Sheng Ma
Frontiers Comput. Sci.1
2025 A quantized network processor for mixed-precision deep learning models based on enhanced Microscaling format
Mei Wen, Xin Ju 0005, Zhuang Cao, Yang Guo 0003
J. Syst. Archit.5
2025 EDA-Copilot: A RAG-Powered Intelligent Assistant for EDA Tools
abstract
With the rise of Large Language Models (LLMs), researchers have become increasingly interested in their applications in EDA flows, particularly in specific subdomains such as serving as knowledge assistants and generating RTL code. In this study, we present a Retrieval-Augmented Generation (RAG) framework tailored to EDA task processing, named EDA-Adaptive RAG. This framework addresses the implicit semantics of EDA data and facilitates efficient knowledge acquisition through classification and enhanced retrieval, significantly enhancing LLMs ability to acquire EDA knowledge. Furthermore, we aim to integrate RAG into the design process as an EDA assistant application. Using RTL code generation as a case study, we demonstrate that the performance of RTL code generation can be enhanced through highly relevant retrievals provided by our RAG. The experimental analysis involves EDA Q&A tasks and RTL code generation evaluation. It is shown that our method outperforms the latest works in terms of both answer stability and code quality.
Haoying Wu, Bei Yu 0001, Yang Guo 0003
ACM Trans. Design Autom. Electr. Syst.5
2024 CoDPoC IP: A Configurable Data Protection Circuit to Support Multiple Key Agreement Scheme
Yijing Peng, Zhenyu Wang 0014, Zhenbin Guo, Ding Deng, Shaoqing Li, Yang Guo 0003
ICICS (2)8
2024 Enhancing the PE Utilization for Multi-Precision Systolic Array via Optimizing Computation Latency
abstract
Systolic array (SA) architectures are widely recognized as the optimal choice for Convolutional Neural Networks (CNNs). However, existing SAs suffer from reduced computational efficiency when confronted with an inadequate workload scale. Furthermore, this issue becomes even more pronounced in the accelerators that support multiple precisions. In this paper, by analyzing the under-utilization of processing element (PE), we propose a SA accelerator that optimizes computation latency for multi-precision scenarios. Considering dynamic changes in data precision, we incorporate a switching strategy to further enhance computational efficiency. Experimental results demonstrate that our proposed design only incurs a 1.203% increase in area compared to the classic approach, while achieving an average performance improvement of 20% on small-scale CNN models.
Mei Wen, Xin Ju 0005, Junzhong Shen, Yang Guo 0003
ISCAS5
2024 SCD-PUF: Shuffled Chaotic-dual-PUF With High Machine Learning Attack Resilience
abstract
Physical unclonable functions (PUFs) provide a promising solution for enhancing security and device authentication. Strong PUFs can generate quantities of challenge-response pairs(CRPs) but are vulnerable to machine learning (ML) attacks. Weak PUFs must restrict direct access to the original response because they have limited CRPs. In this article, we present a Shuffled Chaotic-dual-PUF structure(SCD-PUF) to defeat against ML attacks. Its working procedure is divided into two main stages: In the first phase, the weak PUF is used to generate secret bits as parameters for the chaotic configuration and along with the secret bits generated by the chaotic process, serve as the obfuscation configuration for the second phase. The second stage involves the Knuth-Durstenfeld shuffle algorithm, concatenation and XOR operations to obfuscate the challenges and responses at the same time. To prove the effectiveness of our proposal, we implement an example of SCD-PUF using Static Random-Access Memory(SRAM) PUF and Arbiter PUF(APUF) on Xilinx ZedBoard FPGAs. Using Logistic Regression (LR), Support Vector Machine (SVM), and Artificial Neural Networks (ANN) as attacking methods, the learning accuracy is maintained at around 51% even when the training data increase to one million, which proves our proposal has enough resistance to ML attacks. Also, the area overhead of our proposal is appropriate and acceptable.
Yijing Peng, Ding Deng, Zhenyu Wang 0014, Yang Guo 0003
ITC-Asia4
2024 Improving the Ability of Thermal Radiation Based Hardware Trojan Detection
Ting Su 0009, Lusi Zhang, Simin Feng, Jialong Song, Yongkang Tang, Shaoqing Li, Yang Guo 0003, Hengzhu Liu
USENIX Security Symposium11
2024 A survey of compute nodes with 100 TFLOPS and beyond for supercomputers
Junsheng Chang, Kai Lu 0001, Yang Guo 0003, Yongwen Wang, Libo Huang 0002, Yao Wang 0002, Biwei Zhang
CCF Trans. High Perform. Comput.3
2023 Design of three-factor secure and efficient authentication and key-sharing protocol for IoT devices
Zhenyu Wang 0014, Ding Deng, Shen Hou, Yang Guo 0003, Shaoqing Li
Comput. Commun.4
2023 A General Layout Pattern Clustering Using Geometric Matching-based Clip Relocation and Lower-bound Aided Optimization
abstract
With the continuous shrinking of feature size, detection of lithography hotspots has been raised as one of the major concerns in Design-for-Manufacturability (DFM) of semiconductor processing. Hotspot detection, along with other DFM measures, trades off turnaround time for the yield of IC manufacturing, and thus a simplified but wide-ranging pattern definition is a key to the problem. Layout pattern clustering methods, which group geometrically similar layout clips into clusters, have been vastly proposed to identify layout patterns efficiently. To minimize the clustering number for subsequent DFM processing, in this article, we propose a geometric-matching-based clip relocation technique to increase the opportunity of pattern clustering. Particularly, we formulate the lower bound of the clustering number as a maximum-clique problem, and we have also proved that the clustering problem can be solved by the result of the maximum-clique very efficiently. Compared with the experimental results of the state-of-the-art approaches on ICCAD 2016 Contest benchmarks, the proposed method can achieve the optimal solutions for all benchmarks with very competitive runtime. To evaluate the scalability, the ICCAD 2016 Contest benchmarks are extended and evaluated. Moreover, experimental results on the extended benchmarks demonstrate that our method can reduce the cluster number by 16.59% on average, while the runtime is 74.11% faster on large-scale benchmarks compared with previous works.
Yao Wang 0002, Zhiyong Fu, Yang Guo 0003
ACM Trans. Design Autom. Electr. Syst.5
2023 A Soft-Error Mitigation Approach Using Pulse Quenching Enhancement at Detailed Placement for Combinational Circuits
abstract
As technology continuously shrinks, radiation-induced soft errors have become a great threat to the circuit reliability. Among all the causes, the Single-Event Transient (SET) effect is the dominating one for the radiation-induced soft errors. SET-induced soft errors can be mitigated by multiple methods. In terms of area and power overhead, blocking SET propagation is considered to be the most efficient way for soft error reduction. It is found that the SET pulse width can be shrunk by a pulse quenching effect, which can be utilized to mitigate soft errors without introducing any area and power overhead. In this article, we present an effective detailed placer to exploit the pulse quenching effect for soft error reduction in combinational circuits. In our method, the quenching effect enhancement is globally optimized while the cell displacement is minimized. The experimental results demonstrate that our method reduces the soft error vulnerability of the circuits by 29.53% versus 18.38% of the state-of-the-art solution. Meanwhile, our method has a minimal effect on the displacement and half-perimeter wire length (HPWL) compared with the previous solutions, which means a minimum timing influence to the original design.
Yao Wang 0002, Chang Liu 0019, Qiang Wu 0015, Juan Luo, Yang Guo 0003
ACM Trans. Design Autom. Electr. Syst.6
2022 Accurate timing prediction at placement stage with look-ahead RC network
abstract
Timing closure is a critical but effort-taking task in VLSI designs. In placement stage, a fast and accurate net delay estimator is highly desirable to guide the timing optimization prior to routing, and thus reduce the timing pessimism and shorten the design turn-around time. To handle the timing uncertainty at the placement stage, we propose a fast net delay timing predictor based on machine learning, which extract the fully timing features using a look-ahead RC network. Experimental results show that the proposed timing predictor has achieved average correlation over 0.99 with the post-routing sign-off timing results obtained in Synopsys PrimeTime.
Zhiyong Fu, Yao Wang 0002, Chang Liu 0019, Yang Guo 0003
DAC5
2022 Mentha: Enabling Sparse-Packing Computation on Systolic Arrays
abstract
Generalized Sparse Matrix-Matrix Multiplication (SpGEMM) is a critical kernel in domains like graph analytic and scientific computation. As a kind of classical special-purpose architecture, systolic arrays were first used for complex computing problems, e.g., matrix multiplication. However, classical systolic arrays are not efficient enough when handling sparse matrices due to the fact that the PEs containing zero-valued entries perform unnecessary operations that do not contribute to the result. Accordingly, in this paper, we propose Mentha, a framework that enables systolic arrays to accelerate sparse matrix computation by employing a sparse-packing algorithm suitable for various dataflow of systolic array. Firstly, Mentha supports both online and offline methods. By packing the rows or columns of the sparse matrix, the zero-valued items in the matrix are significantly reduced and the density of the matrix is improved. In addition, acceleration benefits can be obtained by the adaptation scheme even with limited resources. Moreover, we reconfigure PEs in systolic arrays at a low cost (1.28x in area, 1.21x in power) and find that our method outperforms TPU-like systolic arrays by 1.2~3.3x in terms of SpMM and 1.3~4.4x in terms of SpGEMM when dealing with moderately sparse matrices (sparsity < 0.9), while its performance is at least 9.7x better than cuSPARSE. Furthermore, experimental results show a FLOPs reduction of roughly 3.4x in the neural network.
Minjin Tang, Mei Wen, Yasong Cao, Junzhong Shen, Jianchao Yang, Jiawei Fei, Yang Guo 0003, Sheng Liu 0001
ICPP7
2022 MT-3000: a heterogeneous multi-zone processor for HPC
Kai Lu 0001, Yang Guo 0003, Chun Huang 0006, Sheng Liu 0001, Ruibo Wang, Jianbin Fang, Tao Tang 0001, Zhaoyun Chen, Biwei Liu, Zhong Liu 0003, Yuanwu Lei, Haiyan Sun
CCF Trans. High Perform. Comput.3
2022 Heterogeneous Systolic Array Architecture for Compact CNNs Hardware Accelerators
abstract
Compact convolutional neural networks have become a hot research topic. However, we find that the systolic array accelerators are extremely inefficient in dealing with compact models, especially when processing depthwise convolutional layers in the neural networks. To make systolic arrays more efficient for compact convolutional neural networks, we propose the heterogeneous systolic array (HeSA) architecture. It introduces heterogeneous processing elements that support multiple modes of dataflow, which can further exploit the reuse data chance of depthwise convolutional layers and without changing the scale or structure of the nave systolic array. By increasing the utilization rate of processing elements in the array, the HeSA improves the performance, throughput, and energy efficiency compared to the standard baseline. In addition, we design the flexible buffer structure for the communication between the computing array and external buffer. Through flexible routing, the HeSA can achieve a large-scale array design with maintaining high processing elements utilization rate and low communication costs. Based on our evaluation with typical workloads, the HeSA improves the utilization rate of the computing resource in depthwise convolutional layers by 4.5 - 11.2 and acquires 1.6 - 3.1 total performance speedup compared to the standard systolic array architecture. In the large-scale array design, the HeSA can reduce the data traffic by 40% while maintaining the same performance as the scaling-out method. By improving the on-chip data reuse chance and reducing data traffic, the HeSA saves over 20% in energy consumption. Meanwhile, the area of the HeSA is basically unchanged compared to the baseline due to its simple design.
Sheng Ma, Yang Guo 0003, Dongsheng Li 0001, Yuran Qiao
IEEE Trans. Parallel Distributed Syst.4
2021 HeSA: Heterogeneous Systolic Array Architecture for Compact CNNs Hardware Accelerators
abstract
Compact convolutional neural networks have become a hot research topic. However, we find that the hardware accelerator with systolic arrays processing compact models is extremely performance-inefficient, especially when processing depthwise convolutional layers in the networks. To make systolic arrays efficient for compact convolutional neural networks, we propose the heterogeneous systolic array (HeSA) architecture. It introduces heterogeneous processing elements that support multiple modes of dataflow, which can further exploit the reuse data chance of depthwise convolutional layers and without changing the architecture of the naïve systolic array. By increasing the utilization rate of processing elements in the array, HeSA improves the performance, throughput, and energy efficiency compared to the standard baseline. Based on our evaluation with typical workloads, HeSA improves the utilization rate of the computing resource in depthwise convolutional layers by 4.5×-5.5× and acquires 1.5-2.2× total performance speedup compared to the standard systolic array architecture. HeSA also improves the on-chip data reuse chance and saves over 20% of energy consumption. Meanwhile, the area of HeSA is basically unchanged compared to the baseline due to its simple design.
Sheng Ma, Yang Guo 0003
DATE4
2021 Improving Inter-kernel Data Reuse With CTA-Page Coordination in GPGPU
abstract
Although modern GPUs are equipped with expanding memory, accommodating the entire working set of large-scale workloads can still be a challenge. With the support of unified virtual memory and demand paging, programmers can transparently oversubscribe the main memory. However, this transparent management still comes at a severe performance cost, especially for applications with inter-kernel data sharing. While there have been many efforts to reduce additional data migrations caused by the memory oversubscription, few consider the reuse of shared data during the boundary of adjacent kernels. Due to limited memory capacity, we observe that adjacent kernel often demands shared pages that were evicted by the previous kernel, resulting in a significant number of costly data migrations. In this paper, we propose a CTA-Page collaborative framework, called CPC, that transparently reduces the impact of memory oversubscription using CTA dispatch switching and page replacement switching coordinately to reuse inter-kernel shared data. We evaluate CPC with a variety of GPGPU benchmark suites. Experimental results show that the system performance is improved by 65 % compared with the state-of-the-art technique for applications with inter-kernel data sharing.
Xuanyi Li, Chen Li 0015, Yang Guo 0003, Rachata Ausavarungnirun
ICCAD3
2021 Editorial for the special issue on reliability and power efficiency for HPC
Jifeng He 0001, Chenggang Wu 0002, Huawei Li 0001, Yang Guo 0003, Tao Li 0022
CCF Trans. High Perform. Comput.4
2021 A dynamically configurable LFSR-based PUF design against machine learning attacks
Shen Hou, Ding Deng, Zhenyu Wang 0014, Jiahe Shi, Shaoqing Li, Yang Guo 0003
CCF Trans. High Perform. Comput.6
2021 Advancing DSP into HPC, AI, and beyond: challenges, mechanisms, and future directions
Chen Li 0015, Chang Liu 0019, Sheng Liu 0001, Yuanwu Lei, Jian Zhang 0022, Yang Guo 0003
CCF Trans. High Perform. Comput.8
2021 Configurable Multi-directional Systolic Array Architecture for Convolutional Neural Networks
abstract
The systolic array architecture is one of the most popular choices for convolutional neural network hardware accelerators. The biggest advantage of the systolic array architecture is its simple and efficient design principle. Without complicated control and dataflow, hardware accelerators with the systolic array can calculate traditional convolution very efficiently. However, this advantage also brings new challenges to the systolic array. When computing special types of convolution, such as the small-scale convolution or depthwise convolution, the processing element (PE) utilization rate of the array decreases sharply. The main reason is that the simple architecture design limits the flexibility of the systolic array. In this article, we design a configurable multi-directional systolic array (CMSA) to address these issues. First, we added a data path to the systolic array. It allows users to split the systolic array through configuration to speed up the calculation of small-scale convolution. Second, we redesigned the PE unit so that the array has multiple data transmission modes and dataflow strategies. This allows users to switch the dataflow of the PE array to speed up the calculation of depthwise convolution. In addition, unlike other works, we only make a few changes and modifications to the existing systolic array architecture. It avoids additional hardware overheads and can be easily deployed in application scenarios that require small systolic arrays such as mobile terminals. Based on our evaluation, CMSA can increase the PE utilization rate by up to 1.6 times compared to the typical systolic array when running the last layers of ResNet-18. When running depthwise convolution in MobileNet, CMSA can increase the utilization rate by up to 14.8 times. At the same time, CMSA and the traditional systolic arrays are similar in area and energy consumption.
Sheng Ma, Xinhai Chen 0001, Yang Guo 0003
ACM Trans. Archit. Code Optim.5
2020 Hierarchical Clustering With Hard-Batch Triplet Loss for Person Re-Identification
abstract
For clustering-guided fully unsupervised person reidentification (re-ID) methods, the quality of pseudo labels generated by clustering directly decides the model performance. In order to improve the quality of pseudo labels in existing methods, we propose the HCT method which combines hierarchical clustering with hard-batch triplet loss. The key idea of HCT is to make full use of the similarity among samples in the target dataset through hierarchical clustering, reduce the influence of hard examples through hard-batch triplet loss, so as to generate high quality pseudo labels and improve model performance. Specifically, (1) we use hierarchical clustering to generate pseudo labels, (2) we use PK sampling in each iteration to generate a new dataset for training, (3) we conduct training with hard-batch triplet loss and evaluate model performance in each iteration. We evaluate our model on Market-1501 and DukeMTMC-reID. Results show that HCT achieves 56.4% mAP on Market-1501 and 50.7% mAP on DukeMTMC-reID which surpasses state-of-the-arts a lot in fully unsupervised re-ID and even better than most unsupervised domain adaptation (UDA) methods which use the labeled source dataset. Code will be released soon on https://github.com/zengkaiwei/HCT
Kaiwei Zeng, Munan Ning, Yang Guo 0003
CVPR4
2020 Maximum Clique Based Method for Optimal Solution of Pattern Classification
abstract
As the transistor feature size continuously shrinks, design for manufacturability (DFM) has become a crucial concern. Layout pattern classification, which groups geometrically similar layout clips into clusters, has been widely utilized in a variety of DFM applications, such as hotspot library generation, hierarchical data storage, and systematic yield optimization. In this paper, we have proposed a maximum clique based method to obtain the lower bound of the clustering count and have proven that the lower bound is the theoretical optimal solution. To solve the clustering problem, we formulate it as a Set-Covering Problem (SCP) and utilize the result of the maximum clique to help the SCP quickly converge. Compared with the experimental results of the state-of-the-art approaches on ICCAD 2016 Contest benchmarks, our proposed method can achieve optimal solutions for all benchmarks with an approximate minimum run-time.
Zhiyong Fu, Yao Wang 0002, Yang Guo 0003
ICCD5
2020 CMSA: Configurable Multi-directional Systolic Array for Convolutional Neural Networks
abstract
The systolic array is one of the most popular choices for convolutional neural network accelerators. However, when computing special convolution, such as small-scale convolution or depthwise convolution, the utilization rate of the array fluctuates or even declines sharply. To address these issues, we design a configurable multi-directional systolic array (CMSA). The array can switch data mapping or dataflow for special convolution by changing the data transmission direction and configuring the array. Meanwhile, it keeps the original systolic array architecture and computing mode. Our design makes the systolic array flexible. Based on our evaluation, CMSA can increase the units utilization rate by up to 1.6× compared to the typical systolic array when running last layers of ResNet. When running depthwise convolution in MobileNet, CMSA can increase the utilization rate by up to 14.8×.
Sheng Ma, Yang Guo 0003
ICCD4
2020 A Macro-Micro Weakly-Supervised Framework for AS-OCT Tissue Segmentation
Munan Ning, Cheng Bian, Donghuan Lu, Chenglang Yuan, Yang Guo 0003, Kai Ma 0002, Yefeng Zheng 0001
MICCAI (5)7
2020 FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and Caching
abstract
Main memory, composed of DRAM, is a performance bottleneck for many applications, due to the high DRAM access latency. In-DRAM caches work to mitigate this latency by augmenting regular-latency DRAM with small-but-fast regions of DRAM that serve as a cache for the data held in the regular-latency (i.e., slow) region of DRAM. While an effective in-DRAM cache can allow a large fraction of memory requests to be served from a fast DRAM region, the latency savings are often hindered by inefficient mechanisms for migrating (i.e., relocating) copies of data into and out of the fast regions. Existing in-DRAM caches have two sources of inefficiency: (1) their data relocation granularity is an entire multi-kilobyte row of DRAM, even though much of the row may never be accessed due to poor data locality; and (2) because the relocation latency increases with the physical distance between the slow and fast regions, multiple fast regions are physically interleaved among slow regions to reduce the relocation latency, resulting in increased hardware area and manufacturing complexityWe propose a new substrate, FIGARO, that uses existing shared global buffers among subarrays within a DRAM bank to provide support for in-DRAM data relocation across subar-rays at the granularity of a single cache block. FIGARO has a distance-independent latency within a DRAM bank, and avoids complex modifications to DRAM (such as the interleaving of fast and slow regions). Using FIGARO, we design a fine-grained in-DRAM cache called FIGCache. The key idea of FIGCache is to cache only small, frequently-accessed portions of different DRAM rows in a designated region of DRAM. By caching only the parts of each row that are expected to be accessed in the near future, we can pack more of the frequently-accessed data into FIGCache, and can benefit from additional row hits in DRAM (i.e., accesses to an already-open row, which have a lower latency than accesses to an unopened row). FIGCache provides benefits for systems with both heterogeneous DRAM banks (i.e., banks with fast regions and slow regions) and conventional homogeneous DRAM banks (i.e., banks with only slow regions)Our evaluations across a wide variety of applications show that FIGCache improves the average performance of a system using DDR4 DRAM by 16.3% and reduces average DRAM energy consumption by 7.8% for 8-core workloads, over a conventional system without in-DRAM caching. We show that FIGCache outperforms state-of-the-art in-DRAM caching techniques, and that its performance gains are robust across many system and mechanism parameters.
Lois Orosa 0001, Xiangjun Peng, Yang Guo 0003, Saugata Ghose, Minesh Patel, Jeremie S. Kim, Juan Gómez-Luna, Mohammad Sadrosadati, Nika Mansouri-Ghiasi, Onur Mutlu
MICRO4
2020 Energy clustering for unsupervised person re-identification
Kaiwei Zeng, Munan Ning, Yang Guo 0003
Image Vis. Comput.4
2020 Deviation based clustering for unsupervised person re-identification
Munan Ning, Kaiwei Zeng, Yang Guo 0003
Pattern Recognit. Lett.3
2020 Novel Design Strategy Toward A2 Trojan Detection Based on Built-In Acceleration Structure
abstract
With the separation of design and manufacture in semiconductor industry, self-designed circuits are exposed to hardware Trojan attacks when they are outsourced to an untrustworthy foundry. Trojans activated by digital logic have gained extensive attention. However, those with analog trigger component remain as a serious issue such as the A2 Trojan. Existing defense against A2 Trojans mainly relies on runtime detection mechanism which needs large monitor hardware overhead and complicated identification/handling software. To address these limitations, this article proposes a built-in structure to accelerate the activation of A2 Trojans, which consists of several composite-logic ring oscillators and a multiple-purpose controller. In addition, two post-fabrication detection schemes named time-division mode-switching (TDMS) and scan-based fault test compatible (SFTC) are proposed. TDMS detection scheme can discover A2 Trojans when running functional patterns by inserting oscillating operations every other cycle. SFTC detection scheme can detect A2 Trojans during scan-based fault test by introducing oscillation before each capture operation. Evaluations across a wide range of A2 Trojans and benchmarks show that our proposal is more power-efficient and area-saving compared with existing monitor structure.
Ding Deng, Yang Guo 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 A Dynamic and Proactive GPU Preemption Mechanism Using Checkpointing
abstract
The demand for multitasking GPUs increases whenever the GPU may be shared by multiple applications, either spatially or temporally. This requires that GPUs can be preempted and switch context to a new application while already executing one. Unlike CPUs, context switching in GPUs is prohibitively expensive due to the large context states to swap out. There have been a number of efforts on reducing the overhead of preemption, through reducing the context sizes or overlapping context switching with execution. All those techniques are reactive approaches, meaning that context switching occurs when the preemption request arrives. In this paper, we propose a dynamic and proactive mechanism to reduce the latency of preemption. We observe that kernel execution is almost always preceded by known commands in both CUDA and OpenCL implementations. Hence, a preemption can be anticipated before the actual request arrives. We study such lead time and develop a prediction scheme to perform an early state saving. When the actual preemption is invoked, an incremental update relative to the previous saved state is performed, much like the conventional checkpointing mechanism. Our design can also choose to drain or checkpointing dynamically and accurately according to the feature of kernels in the runtime. This design effectively reduces the stall time of the preempting kernel due to context switching by 58.6%. Moreover, through careful handling of the saved state, we can also reduce the overall size of saved state by an average of 23.3%, compared with a full context switching.
Chen Li 0015, Andrew Zigerelli, Jun Yang 0002, Youtao Zhang, Sheng Ma, Yang Guo 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2020 Lithography Hotspot Detection with FFT-based Feature Extraction and Imbalanced Learning Rate
abstract
With the increasing gap between transistor feature size and lithography manufacturing capability, the detection of lithography hotspots becomes a key stage of physical verification flow to enhance manufacturing yield. Although machine learning approaches are distinguished for their high detection efficiency, they still suffer from problems such as large-scale layout and class imbalance. In this article, we develop a hotspot detection model based on machine learning with high performance. In the proposed model, we first apply an Fast Fourier Transform--based feature extraction method that can compress large-scale layout to a multi-dimensional representation with much smaller size while preserving the discriminative layout pattern information to improve the detection efficiency. Second, addressing the class imbalance problem, we propose a new technique called imbalanced learning rate and embed it into the convolutional neural network model to further reduce false alarms without accuracy decay. Compared with the results of current state-of-the-art approaches on ICCAD 2012 Contest benchmarks, our proposed model can achieve better solutions in many evaluation metrics, including the official metrics.
Shizhe Zhou, Rui Li 0019, Yao Wang 0002, Yang Guo 0003
ACM Trans. Design Autom. Electr. Syst.6
2019 A Framework for Memory Oversubscription Management in Graphics Processing Units
abstract
Modern discrete GPUs support unified memory and demand paging. Automatic management of data movement between CPU memory and GPU memory dramatically reduces developer effort. However, when application working sets exceed physical memory capacity, the resulting data movement can cause great performance loss.
Chen Li 0015, Rachata Ausavarungnirun, Christopher J. Rossbach, Youtao Zhang, Onur Mutlu, Yang Guo 0003, Jun Yang 0002
ASPLOS6
2019 Improving the DRAM Access Efficiency for Matrix Multiplication on Multicore Accelerators
abstract
The parallelization of matrix multiplication on multicore accelerators divides a matrix into several partitions. The existing design deploys an independent DMA transfer for each core to access its own partition from DRAM. This design has poor memory access efficiency, since memory access streams of multiple concurrent DMA transfers interfere with each other. We propose Distributed-DMA (D-DMA), which invokes one transfer to serve all cores. D-DMA accesses data in a row-major manner to efficiently exploit inter-partition locality to improve the DRAM access efficiency. Compared with a baseline design, D-DMA improves the bandwidth by 84.8% and reduces DRAM energy consumption by 43.1% for micro-benchmarks. It achieves higher performance for the GEMM benchmark. With much lower hardware cost, D-DMA significantly outperforms an out-of-order memory controller.
Sheng Ma, Yang Guo 0003, Shenggang Chen, Libo Huang 0002, Zhiying Wang 0003
DATE2
2019 An Efficient Direct Memory Access (DMA) Controller for Scientific Computing Accelerators
abstract
We design an efficient DMA controller for scientific computing accelerators. It supports several flexible and powerful transfers, including reshape transfers, parameter linking mechanism, and transfer chaining meachnism. We also optimize the DMA controller for critical scientific computing kernels. It supports high bandwidth matrix transposition during data movement. It improves the memory access efficiency for matrix multiplication. Experimental results show that the data movement bandwidth achieved by the DMA controller is similar to the theoretical maximum one. It also performs very closely to an ideal design for real applications.
Sheng Ma, Libo Huang 0002, Yuanwu Lei, Yang Guo 0003, Zhiying Wang 0003
ISCAS4
2019 Coordinated DMA: Improving the DRAM Access Efficiency for Matrix Multiplication
abstract
High performance implementation of matrix multiplication is essential for scientific computing. The memory access procedure is quite possible to be the bottleneck of matrix multiplication. The widely used GotoBLAS GEMM implementation divides the integral matrix into several partitions to be assigned to different cores for parallelization. Traditionally, each core deploys a DMA transfer to access its own partition in the DRAM memory. However, deploying an independent DMA transfer for each core cannot efficiently exploit the inter-core locality. Also, multiple concurrent DMA transfers interfere with each other, further reducing the DRAM access efficiency. We observe that the same row of neighboring partitions is in the same DRAM page, which means that there is significant locality inherent in the address layout. We propose the coordinated DMA to efficiently exploit the locality. It invokes one transfer to serve all cores and moves data in a row-major manner to improve the DRAM access efficiency. Compared with a baseline design, the coordinated DMA improves the bandwidth by 84.8 percent and reduces DRAM energy consumption by 43.1 percent for micro-benchmarks. It achieves higher performance for the GEMM and Linpack benchmark. With much less hardware costs, the coordinated DMA significantly outperforms an out-of-order memory controller.
Sheng Ma, Zhong Liu 0003, Shenggang Chen, Libo Huang 0002, Yang Guo 0003, Zhiying Wang 0003, Meidi Zhang
IEEE Trans. Parallel Distributed Syst.5
2018 PEP: proactive checkpointing for efficient preemption on GPUs
abstract
The demand for multitasking GPUs increases whenever the GPU may be shared by multiple applications, either spatially or temporally. This requires that GPUs can be preempted and switch context to a new application while already executing one. Unlike CPUs, context switching in GPUs is prohibitively expensive due to the large context states to swap out. There have been a number of efforts on reducing the overhead of preemption, through reducing the context sizes or overlapping context switching with execution. All those techniques are reactive approaches, meaning that context switching occurs when the preemption request arrives.
Chen Li 0015, Andrew Zigerelli, Jun Yang 0002, Yang Guo 0003
DAC4
2018 Accelerating CNNs Using Optimized Scheduling Strategy
Sheng Ma, Wenwu Li, Yang Guo 0003
ICA3PP (3)4
2018 Adaptive VC Partitioning for NoCs in GPGPUs
abstract
The design of efficient Networks-on-Chip (NoCs) is essential for GPGPUs. The asymmetry of GPGPU traffic has a significant effect on the overall system performance. An existing VC partitioning design statically assigns more VCs to the heavier reply traffic. Yet, its static partitioning cannot adapt to the dynamic variation of NoC traffic. Thus, we propose an adaptive VC partitioning (A-VCP) mechanism , which dynamically chooses the optimal VC partitioning by sampling the traffic status. Compared with the static configuration, A-VCP averagely improves the system performance by 10.5%, and reduces the energy-delay product by 9.1%.
Sheng Ma, Hongyi Lu, Libo Huang 0002, Li Shen 0007, Yang Guo 0003, Zhiying Wang 0003, Wenliang Xue
ISCAS5
2018 Performance Analysis of Different Convolution Algorithms in GPU Environment
abstract
Convolutional neural networks (CNNs) have a wide range of applications in image and video recognition, recommender systems and natural language processing. But CNNs are computationally intensive, and its computational cost is hard to accept. In order to speed up the calculations, people focus on optimizing convolution that account for most of the proportion of CNNs' operation. So, many algorithms have been proposed to accelerate the operation of convolution layers. However, each algorithm has its advantages and disadvantages, and there is no one algorithm that can handle all situations. In this paper, we examine the performance of various algorithms in GPU environment. By building a customized CNN model, we have fully explored the impact of the neural structure on the performance of algorithms, including inference/training speed, memory consumption and power consumption. In addition to the algorithms, we also focus on how their implementations in GPU environment affect their performance. We trace the kernel functions of these implementations to further generalize the characteristics of these algorithms. Finally, we summarize the characteristics of each algorithm., and design a strategy to assigns the appropriate implementation for different convolutional layers in CNNs. With our strategy, we can make AlexNet run 1.2× to 2.8× faster than other strategies in GPU environment. This work has very important meaning for understanding these algorithms and may provide insights for further optimizations of the architecture of GPUs and accelerators.
Sheng Ma, Yang Guo 0003
NAS3
2018 Detailed placement for pulse quenching enhancement in anti-radiation combinational circuit design
Chang Liu 0019, Yang Guo 0003
Integr.4
2018 An SAT-Based Method to Multithreaded Program Verification for Mobile Crowdsourcing Networks
abstract
This paper focused on the safety verification of the multithreaded programs for mobile crowdsourcing networks. A novel algorithm was proposed to find a way to apply IC3, which is typically the fastest algorithm for SAT‐based finite state model checking, in a very clever manner to solve the safety problem of multithreaded programs. By computing a series of overapproximation reachability, the safety properties can be verified by the SAT‐based model checking algorithms. The results show that the new algorithm outperforms all the recently published works, especially on memory consumption (an advantage that comes from IC3).
Long Zhang 0004, Wanxia Qu, Yinjia Huo, Yang Guo 0003, Sikun Li
Wirel. Commun. Mob. Comput.4
2017 A Mixed-Size Monolithic 3D Placer with 2D Layout Inheritance
abstract
Monolithic 3D IC is a high integration density emerging technology in the age of both "More Moore" and "More-than-Moore". In this paper, we propose a novel method of generating mixed-size 3D placement based on transforming a 2D placement result. Experimental results indicate that, when compared with the input 2D placement, the 3D placer can reduce the wirelength by 57%, and provide a four-layer 3D chip footprint of about one quarter of the 2D counterpart. Moreover, our placer can preserve the layout information from the 2D placement input, which means that the 2D placement quality can be inherited in the 3D placement results. Compared with an analytical wirelength-driven placer, our placer achieves 34% benefit on 2D layout inheritance and 12% benefit on runtime with acceptable (4%) wirelength cost.
Yao Wang 0002, Yang Guo 0003, Sorin Cotofana
ACM Great Lakes Symposium on VLSI3
2017 Fairness-Oriented and Location-Aware NUCA for Many-Core SoC
abstract
Non-uniform cache architecture (NUCA) is often employed to organize the last level cache (LLC) by Networks-on-Chip (NoC). However, along with the scaling up for network size of Systems-on-Chip (SoC), two trends gradually begin to emerge. First, the network latency is becoming the major source of the cache access latency. Second, the communication distance and latency gap between different cores is increasing. Such gap can seriously cause the network latency imbalance problem, aggravate the degree of non-uniform for cache access latencies, and then worsen the system performance.
Zicong Wang, Chen Li 0015, Yang Guo 0003
NOCS4
2017 A Parallel Test Application Method towards Power Reduction
Ding Deng, Yang Guo 0003, Zhentao Li
J. Electron. Test.2
2016 DLL: A dynamic latency-aware load-balancing strategy in 2.5D NoC architecture
abstract
As the 3D stacking technology still faces several challenges, the 2.5D stacking technology gains better application prospects nowadays. With the silicon interposer, the 2.5D stacking can improve the bandwidth and capacity of the memory system. To satisfy the communication requirements of the integrated memory system, the free routing resources in the interposer should be explored to implement an additional network. Yet, the performance is strongly limited by the unbalanced loads between the CPU-layer network and the interposer-layer network. In this paper, to address this issue, we propose a dynamic latency-aware load-balancing (DLL) strategy. Our key innovations are detecting congestion of the network layer via the average latency of recent packets and making the network layer selection at each source node. We leverage the free routing resources in the interposer to implement a latency propagation ring. With the ring, the latency information tracked at destination nodes is propagated back to source nodes. We achieve load-balance by using these information. Experimental results show that compared with the baseline design, a destination-detection strategy and a buffer-aware strategy, our DLL strategy achieves 45%, 14.9% and 6.5% of average throughput improvements with minor overheads.
Chen Li 0015, Sheng Ma, Lu Wang 0019, Zicong Wang, Xia Zhao 0004, Yang Guo 0003
ICCD6
2016 Ripple 2.0: Improved Movement of Cells in Routability-Driven Placement
abstract
Routability is one of the most important problems in high-performance circuit designs. From the viewpoint of placement design, two major factors cause routing congestion: (i) interconnections between cells and (ii) connections on macro blockages. In this article, we present a routability-driven placer, Ripple 2.0, which emphasizes both kinds of routing congestion. Several techniques will be presented, including (i) cell inflation with routing path consideration, (ii) congested cluster optimization, (iii) routability-driven cell spreading, and (iv) simultaneous routing and placement for routability refinement. With the official evaluation protocol, Ripple 2.0 outperforms other published academic routability-driven placers. Compared with top results in the ICCAD 2012 contest, Ripple 2.0 achieves a better detailed routing solution obtained by a commercial router.
Yao Wang 0002, Yang Guo 0003, Evangeline F. Y. Young
ACM Trans. Design Autom. Electr. Syst.3
2013 Application specified soft error failure rate analysis using sequential equivalence checking techniques
abstract
Soft errors have become a critical challenge as a result of technology scaling. However, to evaluate the influence of soft errors in flip-flop (FF) on the failure of circuit is a hard verification problem. Here, we proposed a novel flip-flop soft error failure rate analysis methodology using sequential equivalence checking (SEC) and taking the application behaviors into consideration, which combines the advantage of formal techniques based approaches in completeness and the advantage of application behaviors in accuracy in differentiating vulnerability of FFs. As a result, all the FFs in a circuit are sorted by their failure rates and designers can use this information to perform optimal hardening of selected sequential components against soft errors. Experimental results on an implementation of a SpaceWire end node and the set of the largest ISCAS'89 benchmark sequential circuits demonstrate the efficiency of our approach. Case study on an instruction decoder of a practical 32 bits microprocessor shows the applicable of our methodology.
Tun Li 0002, Sikun Li, Yang Guo 0003
ASP-DAC4
2013 Translation validation of scheduling in high level synthesis
abstract
The growing design-productivity gap has made designers shift toward using high-level synthesis (HLS) techniques to generate register transfer level design from high-level languages. Unfortunately, this translation process is very complex and may introduce bugs into the generated design, which can create a mismatch between what a designer intends and what is actually implemented in the circuit. In this paper, we present an equivalence checking method to validate the result of HLS scheduling against the initial high-level program. Finite state machine with data path (FSMD) models were used to represent designs before and after scheduling. The proposed method uses a bisimulation relation approach to prove equivalence. The automatically established bisimulation relation guarantees that for each execution sequence in the design before scheduling, a related and equivalent execution sequence exists in the design after scheduling and vice versa. Our method provides a unified way to deal with various scheduling optimizations. We have implemented our validation technique and compared it with a state-of-the-art HLS scheduling verification method. The promising results show the effectiveness and efficiency of our method.
Tun Li 0002, Yang Guo 0003, Wanwei Liu, Mingsheng Tang
ACM Great Lakes Symposium on VLSI2
2012 State space reduction in modeling checking parameterized cache coherence protocol by two-dimensional abstraction
abstract
Scalability of cache coherence protocol is a key component in future shared-memory multi-core or multi-processor systems. The state space explosion is the first hurdle while applying model-checking to scalable protocols. In order to validate parameterized cache coherence protocols effectively, we present a new method of reducing the state space of parameterized systems, two-dimensional abstraction (TDA). Drawing inspiration from the design principle of parameterized systems, an abstract model of an unbounded system is constructed out of finite states. The mathematical principles underlying TDA is presented. Theoretical reasoning demonstrates that TDA is correct and sound. An example of parameterized cache coherence protocol based on MESI illustrates how to produce a much smaller abstract model by TDA. We also demonstrate the power of our method by applying it to various well-known classes of protocols. During the development of TH-1A supercomputer system, TDA was used to verify the coherence protocol in FT-1000 CPU and showed the potential advantages in reducing the verification complexity.
Yang Guo 0003, Wanxia Qu, Long Zhang 0004
J. Supercomput.1
2007 Coverage Driven Test Generation Framework for RTL Functional Verification
abstract
Functional verification is widely recognized as the bottleneck of the hardware design cycle. The coverage-driven verification approach makes coverage the core engine that drives the whole verification flow, which enables reaching high quality verification in a timely manner. In this paper, we present a coverage driven test generation methodology and a set of tools. We present a novel method for automatic generating simulation vectors from HDL descriptions based on path coverage and constraint solving. We present a novel approach to generate functional vectors based on assertions for RTL design verification. Our approach combines program-slicing based design extraction, word-level SAT and dynamic searching techniques. We also present a coverage analysis method based on VCD file, which only replaying the simulation of the control statements in the HDL description. Experimental results show the efficiency of our methodology.
Yang Guo 0003, Wanxia Qu, Tun Li 0002, Sikun Li
CAD/Graphics1
2007 A Novel Collaborative Verification Environment for SoC Co-Verification
abstract
We designs and implements a system-on-chip SW/HW co-verification environment SoC-Gen, which collaborates formal verification and simulation techniques for SoC co-verification. This paper first give an overview of SoC-Gen, and then focus on the simulation based verification environment: SoC-CBSHVE, which based on componential design and integration methodology The environment adopts automatic software, hardware and simulation wrappers. Simulator adopts asynchronous parallel algorithm and bus-based communication mechanism. Five data buses support verification components simulation communication. Standard message format and unify simulation interfaces easy SoC components design and verification. Experimental results show that SoC-CBSHVE enables easy debugging, rich portability, and high verification speed, at a low cost for system-on-chip system-level software and hardware co-verification.
Tun Li 0002, Sikun Li, Jinshan Yu, Yang Guo 0003
CSCWD4
2006 Scheduling of Transactions Based on Extended Scheduling Timed Petri Nets for SoC System-Level Test-Case Generation
Jinshan Yu, Tun Li 0002, Yang Guo 0003, QingPing Tan
EUC3
2006 TraceDo: An On-Chip Trace System for Real-Time Debug and Optimization in Multiprocessor SoC
Xiao Hu 0004, Pengyong Ma, Shuming Chen, Yang Guo 0003
ISPA4
2005 Automatic functional test program generation for microprocessor verification
abstract
A novel specification driven and constraints solving based method to automatically generate test programs from simple to complex ones for advanced microprocessors is presented in this paper. Our microprocessor architectural automatic test program generator (MA2TG) can produce not only random test programs but also a sequence of instructions for a specific constraint by specifying a user constraints file. The proposed methodology makes three important contributions. First, it simplifies the microprocessor architecture modeling and eases adoption of architecture modification via architecture description language (ADL) specification. Second, it generates test programs for specific constraints utilizing the power of state-to-art constraints solving techniques. Finally, the number of test program for microprocessor verification and the verification time are dramatically reduced. We applied this method on DLX processor to illustrate the usefulness of our approach.
Tun Li 0002, Yang Guo 0003, Sikun Li
ASP-DAC4
2005 Predicate Abstraction of RTL Verilog Descriptions Using Constraint Logic Programming
Tun Li 0002, Yang Guo 0003, Sikun Li, GongJie Liu
ATVA2
2005 Functional Vectors Generation for RT-Level Verilog Descriptions Based on Path Enumeration and Constraint Logic Programming
abstract
This paper presents a novel method for automatic functional vectors generation from RT-level HDL descriptions based on path coverage and constraint solving. Compared with existing method, the advantage of this method includes: 1) it avoids generating redundant constraints, which will accelerate the test generation process, 2) it solves the problem of how to propagate the internal values to the primary inputs with decision models, 3) it can handle various HDL description styles, and various styles of designs. Experimental results conduct on several practical designs show that our method can efficiently improve the functional vectors generation process. The prototype system has been applied to verify RTL description of a real 32-bits microprocessor core and complex bugs remained hidden in the RTL descriptions are detected.
Tun Li 0002, Yang Guo 0003, GongJie Liu, Sikun Li
DSD2
2005 MA2TG: A Functional Test Program Generator for Microprocessor Verification
abstract
A novel specification driven and constraints solving based method to automatically generate test programs from simple to complex ones for advanced microprocessors is presented in this paper. Our microprocessor architectural automatic test program generator (MA/sup 2/TG) can produce not only random test programs but also a sequence of instructions for a specific constraint by specifying a user constraints file. The proposed methodology makes three important contributions. First, it simplifies the microprocessor architecture modeling and eases adoption of architecture modification via architecture description language (ADL) specification. Second, it generates test programs for specific constraints utilizing the power of state-to-art constraints solving techniques. Finally, the number of test program for microprocessor verification and the verification time are dramatically reduced. We applied this method on DLX processor to illustrate the usefulness of our approach.
Tun Li 0002, Yang Guo 0003, GongJie Liu, Sikun Li
DSD3
2004 Parallel verilog simulation: architecture and circuit partition
Tun Li 0002, Yang Guo 0003, Sikun Li, Fujiang Ao, Gongjie Li
ASP-DAC2
2004 CLP Based Static Property Checking
Tun Li 0002, Yang Guo 0003, Sikun Li
ATVA2
2004 Assertion-based automated functional vectors generation using constraint logic programming
abstract
We present a novel approach to generate functional vectors based on assertions for RTL design verification. Our approach combines program-slicing based design extraction, word-level SAT and dynamic searching techniques. Through design extraction, vectors generation need only concern about the design parts related to the given assertion, thus large practical designs can be handled. Constraints Logic Programming (CLP) naturally models mixed bit-level and word-level constraints, and word-level SAT techniques solve the mixed constraints in a unified framework, which gain perfect performance. Initial states derived from dynamic simulation can dramatically accelerate the searching process of functional vectors generation. A prototype system has been built, and the experimental results on some public benchmarks and industrial circuits demonstrate the efficiency of our approach and its applicability to large practical designs.
Tun Li 0002, Yang Guo 0003, Sikun Li
ACM Great Lakes Symposium on VLSI2
2004 Automatic Circuit Extractor for HDL Description Using Program Slicing
Tun Li 0002, Yang Guo 0003, Sikun Li
J. Comput. Sci. Technol.2
2003 An Automatic Circuit Extractor for RTL Verification
abstract
For RTL verification, we have to separate the control and datapath parts contained in the whole design, and apply different verification techniques for different parts. This paper presents a new circuit extraction method using program slicing technique, and develops an elegant theoretical basis based on program slicing for circuit extraction from Verilog description. The technique can obtain a chaining slice for given signals of interest. Compared with related researches, the main advantages of our method include: it is fine grain; it has no HDL coding style limitation; it is precise and is capable of dealing with various Verilog constructions. The technique has been integrated with a commercial simulation environment and incorporated into a design process. The experimental results on practical designs show the significant benefits of the proposed approach.
Tun Li 0002, Yang Guo 0003, Sikun Li
Asian Test Symposium2