EDBT 2026 Demo / reviewers in the wild / expert
Guihai Yan
dblp:12/7075
· DBLP profile ↗
68ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-1254-3278ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 61 · 10 first-author · 23 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAPID: Accelerating Point Cloud Diffusion Models via Space-Aware Mix-Precision QuantizationabstractPoint cloud diffusion models, as an emerging 3D generation method, hold broad prospects in 3D modeling, AR/VR, and so on. However, their reliance on costly full-precision neural network computations during extended denoising process limits their practical application. To address this challenge, we propose RAPID, an accelerator co-designed with a space-aware quantization method. First, RAPID uses K-means to partition points into groups and computes scaling factors in each, mitigating accuracy issues caused by uneven distribution. Second, it employs a mixed-precision quantization scheme that uses low precision for internal point groups and high precision for detail-rich edge groups, ensuring generation quality while minimizing bit-width. Third, it reuses computation results for groups with little change between timesteps, reducing redundant calculations. Moreover, RAPID’s hardware features a mixed-precision PE array for efficient computations at various bit-widths, and a filter for dynamic bit-width allocation and result reuse. Evaluations show that, compared to the NVIDIA RTX A5000 GPU and state-of-the-art accelerators, RAPID achieves average speedups of 9.22×, 4.66×, 3.69×, and 3.01×, and energy savings of 61.74×, 4.30×, 3.94×, and 2.76×, with negligible accuracy loss. Qichu Sun, Linxi Lu, Haishuang Fan, Jingya Wu, Huawei Li 0001, Xiaowei Li 0001, Guihai Yan |
DATE | 8 |
| 2026 | DPU for Cybersecurity: Enabling Inline Defense and Self-Protection
Xiaowei Li 0001, Yunkun Liao, Guihai Yan |
J. Comput. Sci. Technol. | 3 |
| 2025 | APTO: Accelerating Serialization-Based Point Cloud Transformers with Position-Aware PruningabstractPoint cloud processing has broad applications in autonomous driving and robotics. Serialization-based point cloud transformers map unordered point clouds onto directed curves, use sparse convolution for down-sampling and apply attention in local windows to capture spatial relationships. Despite achieving great accuracy, these models face inference latency challenges: neighbor search in sparse convolution exhibits low parallelism; attention computation remains complex, especially with larger window sizes; softmax introduces data dependencies. This paper proposes APTO, an accelerator for serialization-based models. It uses voxels' z-curve indices to perform neighbor searches in parallel, employs a position-aware pruning strategy using neighboring voxel counts to eliminate useless attention computations, and adopts a fine-grained attention dataflow for parallel processes with minimal data dependencies. Besides, its hardware has dedicated computation cores for efficient processing. Evaluations show that APTO achieves average 10.22×, 3.53× and 2.70× speedups over RTX 4090 GPU, PointAcc, and SpOctA, with 153.59×, 8.57× and 7.25× energy savings. Qichu Sun, Haishuang Fan, Fangqiang Ding, Linxi Lu, Jingya Wu, Xiaowei Li 0001, Guihai Yan |
ASP-DAC | 8 |
| 2025 | SNO: Securing Network Function Offloading on FPGA-based SmartNICs in Untrusted CloudsabstractAs network bandwidth outpaces host CPU compute capability, Smart Network Interface Cards (SmartNICs) are increasingly deployed to offload network functions from the host CPU. FPGA-based SmartNICs excel due to their programmability at hardware speed, enabling high-performance and customized offloading. Securing offloaded network functions on FPGA-based SmartNICs is a critical challenge in the cloud, as the sensitive user cannot fully trust the cloud service provider (CSP). CPU Trusted Execution Environments (TEEs) protect software code, not FPGA hardware circuits. Existing FPGA TEEs fail to provide packet I/O protection, System-on-Chip (SoC) CPU utilization, and user-friendly memory access interfaces. To address this gap, we introduce SNO, the first TEE for FPGA-based SmartNICs with the secure boot, the SNO Manager for attestation and network function lifecycle orchestration, and the SNO Guard for I/O encryption and authentication. SNO increases SoC CPU utilization (+6.6% for 8-CPU SoC) by co-locating the SNO Manager with the CSP software while isolating the security-critical components of SNO Manager inside SoC CPU TEE, reduces performance overhead by integrating a fully-pipelined AES-GCM engine and overlapped execution, and offers a user-friendly (86.9% user code reduction) streaming interface. The experimental results show that SNO introduces a relative latency overhead of 7.7–143.2% (corresponding to absolute overheads up to 96 nanoseconds) across five network functions, significantly offset by microsecond-level latency savings from offloading. Yunkun Liao, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
ICCAD | 6 |
| 2025 | Flame: A Multiplier-Free LLM Accelerator with Dynamic Block Floating PointabstractRecently, large language models (LLMs) have achieved remarkable success across various machine learning tasks. However, deploying LLM inference still remains challenging due to the high computing overhead and substantial memory footprint. In this paper, we propose Flame, a software-hardware co-design accelerator that enables efficient LLM inference with dynamic block floating point (BFP). Flame employs a layer-wise adaptive precision search algorithm to optimize BFP mantissa bitwidth to balance model accuracy and inference speedup. To mitigate accuracy degradation introduced by BFP conversion, we propose a channel reorder approach that adjusts the value distribution within each tensor group. Finally, we leverage a novel hardware-friendly linear complexity multiplication algorithm to implement an efficient hardware accelerator featuring multiplierfree processing units. Evaluation results show that Flame achieves up to$6.32 \times$speedup and$4.04 \times$energy efficiency compared to GPU, while maintaining superior model accuracy. Ao Lyu, Haishuang Fan, Guihai Yan |
ICCD | 3 |
| 2025 | Hermes: Accelerating Packet Processing in DPU with Neural NetworkabstractThis paper presents Hermes, an approach to address two bottlenecks in Open vSwitch (OvS) implemented on Data Processing Units (DPUs). The first bottleneck stems from memory bandwidth contention in the hardware path, while the second bottleneck arises from increased upcalls to software during OpenFlow ruleset updates. Hermes leverages the Range-Query Recursive Model Index (RQRMI) to overcome these bottlenecks through two methods: 1) a three-level hardware path design that combines hash-based flow tables and RQRMI inference module, which reduces memory bandwidth consumption, and 2) a hardware-accelerated training module that enables rapid model retraining for synchronizing hardware path with OpenFlow rulesets. Our prototype shows Hermes increases throughput by up to$3.7 \times$over traditional OvS offloading schemes, while reducing upcalls by 71% and improving throughput by$1.6 \times$during ruleset updates. Xinyu Chen 0001, Hanyue Lin, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
ICCD | 7 |
| 2025 | FUS: FPGA-based Universal Sketch with homogeneous and heterogeneous memory architectures
Yunkun Liao, Jingya Wu, Wenyan Lu, Huawei Li 0001, Xiaowei Li 0001, Guihai Yan |
CCF Trans. High Perform. Comput. | 6 |
| 2025 | KPU: Kernel Processing Unit for in-Memory Analytical Query ProcessingabstractDomain-specific architecture has greatly improved performance and energy efficiency in in-memory databases, especially for accelerating single-functional computing logic in analytic query processing, such as sort, join and aggregation. However, as data volumes surge exponentially, these dedicated accelerators are struggling to satisfy the burgeoning demand for handling intricate and multifaceted workloads. A major challenge lies in establishing a flexible framework that engages these ‘coarse-grained’ units without incurring extra overheads from hardware integration, programming, compilation, runtime and operating systems.In this paper, the kernel processing unit (KPU) is proposed to optimize CPU-accelerator heterogeneous systems for in-memory databases. KPU provides a unified interface to consolidate all database query operators. In terms of KPU hardware architecture, kernel customization and data transmission are two critical bottlenecks. To address the challenges, multiple independently designed homogeneous table cores are integrated to support flexible high-performance SQL queries, and a customized efficient data management system (DMS) works collaboratively to maximize the utilization of on-chip memory bandwidth. Additionally, a database application-specific KPU instruction set architecture (KISA) dedicated to parallel analytical query processing is proposed to enable parallel KPU programming. To trade off between accelerator computing capacity and data transfer latency, KPU designs an offloading mechanism to map SQL queries between the CPU and accelerator adaptively based on a performance model and a function simulator. The experiments demonstrate that KPU surpasses the general-purpose CPU and GPU by an average of 24.5× and 8.75×, respectively. Jingya Wu, Wenyan Lu, Haishuang Fan, Hao Kong 0005, Xiaowei Li 0001, Guihai Yan |
IEEE Trans. Computers | 6 |
| 2025 | GRACE: An End-to-End Graph Processing Accelerator on FPGA With Graph Reordering EngineabstractGraphs play an important role in various applications. With the rapid expansion of vertices in real life, existing large-scale graph processing frameworks on CPUs and GPUs encounter challenges in optimizing cache usage due to irregular memory access patterns. To address this, graph reordering has been proposed to improve the locality of the graph, but introduces significant overhead without delivering substantial end-to-end performance improvement. While there have been many FPGA-based accelerators for graph processing, achieving high throughput often requires complex graph prepossessing on CPUs. Therefore, implementing an efficient end-to-end graph processing system remains challenging. This article introduces GRACE, an end-to-end FPGA-based graph processing accelerator with a graph reordering engine and a pull-based vertex-centric programming model (PL-VCPM) Engine. First, GRACE employs a customized high-degree vertex cache (HDC) to improve memory access efficiency. Second, GRACE offloads the graph preprocessing to FPGA. We customize an efficient graph reordering engine to complete preprocessing. Third, GRACE adopts a graph pruning strategy to remove the activation and computation redundancy in graph processing. Finally, GRACE introduces a graph conflict board (GCB) to resolve data conflicts and a multiport cache to enhance parallel efficiency. Experimental results demonstrate that GRACE achieves$7.1 \times $end-to-end performance speedup over CPU and$1.8 \times $over GPU, as well as$27.3 \times $and$8.7 \times $energy efficiency over CPU and GPU. Moreover, GRACE delivers up to$34.9 \times $performance speedup compared to the state-of-the-art FPGA accelerator. Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | Co-ViSu: Accelerating Video Super-Resolution With Codec Information ReuseabstractHigh-resolution (HR) videos have gained popularity with the widespread adoption of high-definition displays. Super-resolution (SR) techniques aim to recover HR frames from low-resolution (LR) frames. While deep neural network (DNN)-based SR methods have outperformed traditional techniques in quality, they face performance challenges. FPGA-based SR accelerators have been developed to optimize the performance and power efficiency. However, most of these accelerators process only uncompressed video frames and perform per-frame DNN inference, overlooking the temporal-spatial information inherent in compressed video bitstreams. We propose a novel compressed video SR workflow that includes a codec information reuse algorithm and a dedicated FPGA accelerator named Co-ViSu. Our approach leverages the observation that non-key frames can be reconstructed using codec information and HR key-frames, significantly reducing DNN computations. The Co-ViSu algorithm employs subpixel interpolation to enhance high-frequency details and an MV-aware method to improve SR reconstruction quality. The Co-ViSu hardware integrates decoder, SR, and encoder engines within a parallel pipeline architecture, utilizing codec information reuse to bypass non-key frame decoding, eliminate complex DNN computations, and accelerate encoding processes. Experimental results demonstrate that Co-ViSu achieves performance improvements ranging from$3.6\times $to$9.4\times $and a$4.2\times $gain in energy efficiency with minimal quality loss compared to traditional flow. Additionally, Co-ViSu offers a$2.1\times $increase in throughput compared to state-of-the-art solutions. Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | TianMen: a DPU-based storage network offloading structure for disaggregated datacentersabstractIn modern disaggregated datacenters, the storage network which interconnects the compute and memory pools becomes the performance bottleneck. The high-end RDMA devices cannot meet the complex requirements of storage networks, due to the limited RDMA semantics and throughput. Existing solutions essentially follow the monolithic design, so they suffer from underutilized resources and high scaling costs. In this paper, we design TianMen, which offloads the storage network by extending RDMA semantics and customizing communication hardware structure. Specifically, we use DPU as the infrastructure, leveraging the rich storage and compute resources. TianMen enables fully disaggregated storage system that bypasses the server-side CPU, and supports elastic resource pools. Experimental results show that, compared with state-of-the-art solutions: 1) Tian-Men achieves 1 RTT for GET/PUT operations and up to 6× access acceleration; 2) TianMen provides CPU bypass storage network management, including 3.2× speedup of metadata consistency management, per-request load balancing, and 10s microsecond-level fault recovery latency; 3) TianMen saturates the communication bandwidth when processing small payload, increasing the bandwidth utilization by 34.2%. And TianMen achieves 2.27× throughput compared to the commercial RNICs. Weiyue Zhao, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
SoCC | 5 |
| 2024 | Co-Via: A Video Frame Interpolation Accelerator Exploiting Codec Information ReuseabstractVideo Frame Interpolation (VFI) aims to generate intermediate frames between consecutive frames. Recent DNN-based VFI offers superior quality but suffers from performance issues. However, very few studies have focused on VFI hardware acceleration and existing work overlooks temporal information from compressed video bitstreams. In this paper, we propose a novel compressed VFI workflow and an accelerator, Co-Via. Co-Via exploits codec information reuse to reduce complex DNN computations and alleviate hardware pressure. FPGA-based Co-Via outperforms an RTX 4090 GPU 10.31X, offering a 43.08X energy efficiency boost. Its ASIC version achieves 2.4X higher throughput and 3.6X energy efficiency than the state-of-the-art solution. Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
DAC | 6 |
| 2024 | PHD: Parallel Huffman Decoder on FPGA for Extreme Performance and Energy EfficiencyabstractHuffman decoding is crucial in data compression, and the self-synchronization-based parallel decoding algorithm enables subsequence-level parallelism. This paper introduces PHD, the first accelerator designed for self-synchronization-based parallel Huffman decoding on a Field-Programmable Gate Array (FPGA). Designing PHD poses challenges, including managing fine-grained parallelism, addressing limited on-chip memory, and handling inter-codeword dependency. PHD incorporates bit-level, subsequence-level, and tile-level parallelism, utilizes hybrid memory to store the codebook efficiently, and introduces the ONCE MORE optimization to reduce decoding loop iterations. Experimental results demonstrate that PHD outperforms the state-of-the-art GPU-based baseline regarding latency (9.4X to 12.8X reduction) and energy consumption (12.4X to 18.2X reduction). Yunkun Liao, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
DAC | 5 |
| 2024 | Athena: Add More Intelligence to RMT-Based Network Data Plane with Low-Bit Quantization
Yunkun Liao, Hanyue Lin, Jingya Wu, Wenyan Lu, Huawei Li 0001, Xiaowei Li 0001, Guihai Yan |
Euro-Par (2) | 7 |
| 2024 | Efficient RNIC Cache Side-Channel Attack Detection Through DPU-Driven Architecture
Yunkun Liao, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
Euro-Par (2) | 5 |
| 2024 | AMST: Accelerating Large-Scale Graph Minimum Spanning Tree Computation on FPGAabstractThe minimum spanning tree (MST) plays an important role in variant fields, such as chip design and network analysis. With the rapid expansion of vertices in real-life graphs, the bottleneck problem of MST algorithms in large-scale graphs grows more prominent. While there have been many FPGA-based accelerators for large-scale graph algorithms such as Graph Random Walk, and various algorithms to accelerate MST on CPUs and GPUs, effectively implementing MST algorithms for large-scale graphs on FPGAs remains quite challenging. This is due to several reasons: The neighbor vertices in the graph require extensive random memory access and the memory access characteristics vary across different stages and iterations. There are a large number of useless computations due to the existence of internal edges within a component (intra-edge). Parallel MST algorithm suffers from significant communication overhead due to the minimum edge data update conflicts and memory read-write conflicts.This paper proposes AMST to accelerate large-scale graph MST computation on FPGA. First, AMST employs a customized hash-based high-degree vertex cache (HDC) to improve memory access efficiency. Second, AMST adopts a graph pruning strategy that skips intra-edge and sorts edges by weight to eliminate useless computation and memory access. Finally, AMST utilizes a sorting networking module and a multi-port HDC to improve parallel efficiency. The experimental results demonstrate that AMST achieves an average performance speedup of 17.52× over CPU and 1.89× over GPU, as well as 74.96× over CPU and 10.45× over GPU on energy efficiency. Haishuang Fan, Qichu Sun, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
IPDPS | 7 |
| 2024 | DPU-Direct: Unleashing Remote Accelerators via Enhanced RDMA for Disaggregated DatacentersabstractThis paper presents DPU-Direct, an accelerator disaggregation system that connects accelerator nodes (ANs) and CPU nodes (CNs) over a standard Remote Direct Memory Access (RDMA) network. DPU-Direct eliminates the latency introduced by the CPU-based network stack, and PCIe interconnects between network I/O and the accelerator. The DPU-Direct system architecture includes a DPU Wrapper hardware architecture, an RDMA-based Accelerator Access Pattern (RAAP), and a CN-side programming model. The DPU Wrapper connects accelerators directly with the RDMA engine, turning ANs into disaggregation-native devices. The RAAP provides the CN with low-latency and high throughput accelerator semantics based on standard RDMA operations. Our FPGA prototype demonstrates DPU-Direct’s efficacy with two proof-of-concept applications: AES encryption and key-value cache, which are computationally intensive and latency-sensitive. DPU-Direct yields a 400x speedup in AES encryption over the CPU baseline and matches the performance of the locally integrated AES accelerator. For key-value cache, DPU-Direct reduces the average end-to-end latency by 1.66x for GETs and 1.30x for SETs over the CPU-RDMA-Polling baseline, reducing latency jitter by over 10x for both operations. Yunkun Liao, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
IEEE Trans. Computers | 5 |
| 2024 | Satisfying Energy-Efficiency Constraints for Mobile SystemsabstractEnergy-efficiency is one of the most important design criteria for mobile systems, such as smartphones and tablets. But current mobile systems always over-provision resources to satisfy users. The root cause is that, we have no knowledge on how much of system performance/energy will exactly satisfy users. Psychophysics defines the quantified link between physical stimuli and human-perceived stimuli. So, we will leverage psychophysics to study the quantified correlation between computer architecture resources (i.e., physical stimuli) and user satisfaction (i.e., human-perceived stimuli). We then exploit such correlation to precisely apportion resources to operate tasks and accurately satisfy users. Benefiting from our precisely-defined user satisfaction criteria and well-designed algorithms, we can reduce energy consumption of computer architectures by up to 42.9% without harming user experience. To the best of our knowledge, we for the first time theoretically and accurately model such substantial correlation. Our work opens a new research domain for fundamentally improving mobiles’ energy-efficiency. Xueliang Li 0002, Shicong Hong, Junyang Chen 0001, Junkai Ji, Chengwen Luo 0001, Guihai Yan, Zhibin Yu 0001, Jianqiang Li 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | Co-ViSu: a Video Super-Resolution Accelerator Exploiting Codec Information ReuseabstractHigh-resolution (HR) videos have become popular due to the widespread adoption of high-definition displays. Super-resolution (SR) techniques aim to recover HR frames from low-resolution (LR) frames. Recently, deep neural network (DNN)-based SR methods have achieved superior quality compared to traditional methods. FPGA-based SR accelerators have been proposed to optimize performance and power efficiency. However, most accelerators tailored for video SR only accept uncompressed video frames and operate per-frame DNN inference, ignoring the temporal-spatial information in compressed video bitstreams. In contrast, we observe that non-key frames can be directly constructed using codec information and HR key-frames, saving a significant amount of DNN computing. In this paper, we propose a novel compressed video SR flow and a specific FPGA accelerator called Co-ViSu that integrates decoder, SR, and encoder engines. Co-ViSu exploits codec information reuse scheme to skip non-key frame decoding, avoid complex DNN computation and speed up encoding. Our experimental results show that Co-ViSu achieves 3.6x to 9.4x performance, 4.2x energy efficiency gain with only 0.17dB quality loss compared to the traditional flow, and 2.1x throughput than state-of-the-art. Haishuang Fan, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
FPL | 5 |
| 2023 | M2VT: A Multi-Output Encoder Accelerator for Multiple-Way Video TranscodingabstractVideo transcoding is a general but compute-intensive technology in video streaming services. Traditional single-encoder accelerators transcode multiple streams independently in the multi-output scenario. However, this mode neglects redundant computation and introduces high hardware complexity. To solve these issues, we propose a multi-encoder accelerator supporting reuse scheme. We introduce four fast algorithms based on parameter sharing to simplify encoding complexity. To further optimize the architecture, we also propose the standalone stream insertion (SSI) to increase the pipeline efficiency, and co-optimize memory access. Implementation results show that multi-encoder can reduce 68.03% computation complexity. Moreover, the area and power efficiency improve 3.05x and 2.62x. Haishuang Fan, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | KPU-SQL: Kernel Processing Unit for High-Performance SQL AccelerationabstractApplication-specific accelerator is a prominent way for analytic query processing. To achieve a substantial improvement over the state-of-the-art in performance while maintaining programmability, we propose a kernel processing unit (KPU) framework and apply it to SQL acceleration. Kernel customization and data transmission are two critical bottlenecks, we separately optimize them in the key core and shadow core with a self-designed data management system. A software stack named RACE with a performance model and function simulator is also introduced. The experiments demonstrate that KPU-SQL outperforms the CPU and GPU by 24.5x and 8.75x on average, respectively. Hao Kong 0005, Haishuang Fan, Jingya Wu, Liyun Cheng, Wenyan Lu, Guihai Yan, Xiaowei Li 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2023 | Optimize the TX Architecture of RDMA NIC for Performance Isolation in the Cloud EnvironmentabstractRemote Direct Memory Access (RDMA) is a promising technology for achieving low latency and high bandwidth access to remote memory. However, performance interference exists when multiple tenants share an RDMA Network Interface Card (RNIC) in the cloud environment. Although some initial studies have investigated the root cause and possible solutions to RDMA performance interference, there is no research to analyze and solve the performance interference from the RNIC architecture. Compared with the existing software approach, optimizing RNIC architecture can introduce less performance and CPU overhead. This paper addresses performance isolation by modeling, analyzing, and optimizing the transmit-side (TX) RNIC architecture. First, we introduce a baseline TX RNIC architecture to explain the existing performance interference. Then, we propose separate caching and slicing execution to avoid the bandwidth-sensitive tenants affecting latency-sensitive tenants. Later, we add isolated backpressure and adaptive Weighted Round-robin scheduling to ensure the bandwidth-sensitive tenants share the bandwidth equally. Our experiments show that these optimizations achieve near-optimal performance isolation. Yunkun Liao, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | BitColor: Accelerating Large-Scale Graph Coloring on FPGA with Parallel Bit-Wise EnginesabstractThe graph coloring algorithm plays a crucial role in many applications such as social network analysis. However, since the minimal graph coloring problem is NP-complete, which is increasingly computationally and memory-intensive as the number of vertices in the graph grows rapidly. Despite numerous FPGA-based works proposed to accelerate large-scale graph processing algorithms, such as Single Source Shortest Path, and various coloring algorithms, such as linear programming algorithms, efficiently implementing the greedy coloring algorithm for large-scale graphs on FPGA still remains highly challenging due to several reasons: ① The coloring algorithm requires color state traversal to determine the final color after traversing neighbor vertices. The time complexity of color traversal is equal to the neighbor vertices traversal, which is inefficient. ② Neighbor vertices traversal requires extensive random memory accesses on vertex color data. ③ Coloring different vertices in parallel is difficult due to potential color update conflicts between adjacent vertices. Haishuang Fan, Jingya Wu, Wenyan Lu, Xiaowei Li 0001, Guihai Yan |
ICPP | 6 |
| 2023 | DOE: database offloading engine for accelerating SQL processing
Hao Kong 0005, Wenyan Lu, Jingya Wu, Yu Zhang 0027, Guihai Yan, Xiaowei Li 0001 |
Distributed Parallel Databases | 6 |
| 2022 | Using Psychophysics to Guide Power Adaptation for Input Methods on Mobile ArchitecturesabstractThe predominant user activities on mobile architectures (e.g., smartphones) involve entering text in instant messaging apps, short message services, and social networking services. Recent research reveals that the normal use of input methods drains approximately half of the battery capacity due to their above-average power requirements and frequent use. In this paper, we first study the power characteristics of mobile input methods and find that they consistently over-provision resources to satisfy users. For example, the psychophysical evidence available indicates the response of spell-checking features within a time threshold makes users feel they have instant feedback. However, current systems perform it very quickly, which is imperceptible to users and costly in terms of energy use. Given this over- provisioning, the system can be slowed down to save energy while retaining the feeling of instant response. Inspired by this observation, we also exploit several other psychophysical facts to identify the exact criteria to satisfy users. As a result, we present a user experience-oriented technology, utexia, to optimize the energy use of mobile input methods. The evaluation shows that utexia conserves up to 42.9% in energy use while strictly ensuring a good user experience. Xueliang Li 0002, Shicong Hong, Junyang Chen 0001, Guihai Yan, Kaishun Wu |
HPCA | 4 |
| 2021 | ShuntFlowPlus: An Efficient and Scalable Dataflow Accelerator Architecture for Stream ApplicationsabstractStreaming processing is an important and growing class of applications for analyzing continuous streams in real time. In such applications, sliding-window aggregation (SWAG) is a widely used approa... Shijun Gong, Wenyan Lu, Guihai Yan, Xiaowei Li 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2020 | BZIP: A compact data memory system for UTXO-based blockchains
Shuhao Jiang, Shijun Gong, Junchao Yan, Guihai Yan, Yi Sun 0004, Xiaowei Li 0001 |
J. Syst. Archit. | 5 |
| 2019 | TNPU: an efficient accelerator architecture for training convolutional neural networksabstractTraining large scale convolutional neural networks (CNNs) is an extremely computation and memory intensive task that requires massive computational resources and training time. Recently, many accelerator solutions have been proposed to improve the performance and efficiency of CNNs. Existing approaches mainly focus on the inference phase of CNN, and can hardly address the new challenges posed in CNN training: the resource requirement diversity and bidirectional data dependency between convolutional layers (CVLs) and fully-connected layers (FCLs). To overcome this problem, this paper presents a new accelerator architecture for CNN training, called TNPU, which leverages the complementary effect of the resource requirements between CVLs and FCLs. Unlike prior approaches optimizing CVLs and FCLs in separate way, we take an alternative by smartly orchestrating the computation of CVLs and FCLs in single computing unit to work concurrently so that both computing and memory resources will maintain high utilization, thereby boosting the performance. We also proposed a simplified out-of-order scheduling mechanism to address the bidirectional data dependency issues in CNN training. The experiments show that TNPU achieves a speedup of 1.5x and 1.3x, with an average energy reduction of 35.7% and 24.1% over comparably provisioned state-of-the-art accelerators (DNPU and DaDianNao), respectively. Guihai Yan, Wenyan Lu, Shuhao Jiang, Shijun Gong, Jingya Wu, Junchao Yan, Xiaowei Li 0001 |
ASP-DAC | 2 |
| 2019 | ShuntFlow: An Efficient and Scalable Dataflow Accelerator Architecture for Streaming ApplicationsabstractStreaming processing is an important and growing class of applications for analyzing continuous streams of real time data. Sliding-window aggregations (SWAGs) dominate the computation time in such applications and dictate an unprecedented computation capacity which poses a great challenge to the computing architectures. General-purpose processors cannot efficiently handle SWAGs because of the specific computation patterns. This paper proposes an efficient accelerator architecture for ubiquitous SWAGs, called ShuntFlow. ShuntFlow is a typical type of Kernel Processing Unit (KPU) where "Kernel" represent two main categories of SWAG operations widely used in streaming processing. Meanwhile, we propose a shunt rule to enable ShuntFlow to efficiently handle SWAGs with arbitrary parameters. As a case study, we implemented ShuntFlow on an Altera Arria 10 AX115N FPGA board at 150 MHz and compared it to previous approaches. The experimental results show that ShuntFlow provides a tremendous throughput and latency advantage over CPU and GPU implementations on both reduce-like and index-like SWAGs. Shijun Gong, Wenyan Lu, Guihai Yan, Xiaowei Li 0001 |
DAC | 4 |
| 2019 | SqueezeFlow: A Sparse CNN Accelerator Exploiting Concise Convolution RulesabstractConvolutional Neural Networks (CNNs) have been widely used in machine learning tasks. While delivering state-of-the-art accuracy, CNNs are known as both compute- and memory-intensive. This paper presents the SqueezeFlow accelerator architecture that exploits sparsity of CNN models for increased efficiency. Unlike prior accelerators that trade complexity for flexibility, SqueezeFlow exploits concise convolution rules to benefit from the reduction of computation and memory accesses as well as the acceleration of existing dense architectures without intrusive PE modifications. Specifically, SqueezeFlow employs a PT-OS-sparse dataflow that removes the ineffective computations while maintaining the regularity of CNN computations. We present a full design down to the layout at 65 nm, with an area of 4.80 mm2and power of 536.09 mW. The experiments show that SqueezeFlow achieves a speedup of 2:9× on VGG16 compared to the dense architectures, with an area and power overhead of only 8.8 and 15.3 percent, respectively. On three representative sparse CNNs, SqueezeFlow improves the performance and energy efficiency by 1:8× and 1:5× over the state-of-the-art sparse accelerators. Shuhao Jiang, Shijun Gong, Jingya Wu, Junchao Yan, Guihai Yan, Xiaowei Li 0001 |
IEEE Trans. Computers | 6 |
| 2019 | Promoting the Harmony between Sparsity and Regularity: A Relaxed Synchronous Architecture for Convolutional Neural NetworksabstractThere are two approaches to improve the performance of Convolutional Neural Networks (CNNs): 1) accelerating computation and 2) reducing the amount of computation. The acceleration approaches take the advantage of CNN computing regularity which enables abundant fine-grained parallelisms in feature maps, neurons, and synapses. Alternatively, reducing computations leverages the intrinsic sparsity of CNN neurons and synapses. The sparsity represents as the computing “bubbles”, i.e., zero or tiny-valued neurons and synapses. These bubbles can be removed to reduce the volume of computations. Although distinctly different from each other in principle, we find that the two types of approaches are not orthogonal to each other. Even worse, they may conflict to each other when working together. The conditional branches introduced by some bubble-removing mechanisms in the original computations destroy the regularity of deeply nested loops, thereby impairing the intrinsic parallelisms. Therefore, enabling the synergy between the two types of approaches is critical to arrive at superior performance. This paper proposed a relaxed synchronous computing architecture, FlexFlow-Pro, to fulfill this purpose. Compared with the state-of-the-art accelerators, the FlexFlow-Pro gains more than 2.5× performance on average and 2× energy efficiency. Wenyan Lu, Guihai Yan, Shijun Gong, Shuhao Jiang, Jingya Wu, Xiaowei Li 0001 |
IEEE Trans. Computers | 2 |
| 2019 | ShuttleNoC: Power-Adaptable Communication Infrastructure for Many-Core ProcessorsabstractNetworks-on-chip (NoCs), as the communication infrastructure in many-core processors, has demonstrated remarkable power consumption along with the technology scaling. However, due to the temporal and spatial heterogeneity of the on-chip traffic, one critical problem is that the NoC power consumption cannot effectively adapt to the variation of its traffic intensity, also known as localized power adaptation, hence yielding a suboptimal power efficiency. Prior approaches either resort to the over-provisioned NoC design or coarse-grained bandwidth scaling to partially alleviate excessive power consumption brought by the traffic temporal or spatial heterogeneity. While in this paper, we propose a novel NoC architecture called Shuttle NoC (ShuttleNoC) to address this challenge. It leverages the link reconfiguration to enable flexible packet traversing between multiple subnetworks, and specialized punch lines to accelerate latency sensitive traffic. With the support of the dedicated power adaptation mechanisms, it is shown in the evaluation that the proposed ShuttleNoC architecture could effectively tackle the power and performance tradeoff and significantly boost the power efficiency compared with the state-of-the-art baselines. Yisong Chang, Guihai Yan, Ning Lin, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | SynergyFlow: An Elastic Accelerator Architecture Supporting Batch Processing of Large-Scale Deep Neural NetworksabstractNeural networks (NNs) have achieved great success in a broad range of applications. As NN-based methods are often both computation and memory intensive, accelerator solutions have been proved to be highly promising in terms of both performance and energy efficiency. Although prior solutions can deliver high computational throughput for convolutional layers, they could incur severe performance degradation when accommodating the entire network model, because there exist very diverse computing and memory bandwidth requirements between convolutional layers and fully connected layers and, furthermore, among different NN models. To overcome this problem, we proposed an elastic accelerator architecture, called SynergyFlow, which intrinsically supports layer-level and model-level parallelism for large-scale deep neural networks. SynergyFlow boosts the resource utilization by exploiting the complementary effect of resource demanding in different layers and different NN models. SynergyFlow can dynamically reconfigure itself according to the workload characteristics, maintaining a high performance and high resource utilization among various models. As a case study, we implement SynergyFlow on a P395-AB FPGA board. Under 100MHz working frequency, our implementation improves the performance by 33.8% on average (up to 67.2% on AlexNet) compared to comparable provisioned previous architectures. Guihai Yan, Wenyan Lu, Shijun Gong, Shuhao Jiang, Jingya Wu, Xiaowei Li 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2018 | RiskCap: Minimizing Effort of Error Regulation for Approximate ComputingabstractQuality management, which is responsible for controlling approximation quality to meet user requirement, plays a key role in the applicability of approximate computing. An effective and efficient quality management needs to be accurate to detect intolerable errors meanwhile light-weight in nature. However, it is difficult to design such a quality management satisfying both the two demands and existing work usually optimizes for one demand at the expense of the other. In this paper, we aim to achieve higher energy efficiency of quality management by optimizing detection accuracy and overhead simultaneously. We observe that the detection difficulty varies across inputs and there exists much redundant computation in detection process. Based on this observation, a cascaded quality management which can minimize the overhead and doesn't lower detection accuracy is proposed. The proposed solution pays more proper computation effort according to different detection difficulties of inputs so as to avoid unnecessary energy consumption. What's more, by exploring the design space sufficiently and effectively, we can assure the highest energy-efficiency of the proposed topology. The experiment results demonstrate that our approach can achieve much greater energy-efficiency than existing solutions. Shuhao Jiang, Xin He 0011, Guihai Yan, Xuan Zhang 0001, Xiaowei Li 0001 |
ATS | 4 |
| 2018 | CCR: A concise convolution rule for sparse neural network acceleratorsabstractConvolutional Neural networks (CNNs) have achieved great success in a broad range of applications. As CNN-based methods are often both computation and memory intensive, sparse CNNs have emerged as an effective solution to reduce the amount of computation and memory accesses while maintaining the high accuracy. However, dense CNN accelerators can hardly benefit from the reduction of computations and memory accesses due to the lack of support for irregular and sparse models. This paper proposed a concise convolution rule (CCR) to diminish the gap between sparse CNNs and dense CNN accelerators. CCR transforms a sparse convolution into multiple effective and ineffective ones. The ineffective convolutions in which either the neurons or synapses are all zeros do not contribute to the final results and the computations and memory accesses can be eliminated. The effective convolutions in which both the neurons and synapses are dense can be easily mapped to the existing dense CNN accelerators. Unlike prior approaches which trade complexity for flexibility, CCR advocates a novel approach to reaping the benefits from the reduction of computation and memory accesses as well as the acceleration of the existing dense architectures without intrusive PE modifications. As a case study, we implemented a sparse CNN accelerator, SparseK, following the rationale of CCR. The experiments show that SparseK achieved a speedup of 2.9× on VGG16 compared to a comparably provisioned dense architecture. Compared with state-of-the-art sparse accelerators, SparseK can improve the performance and energy efficiency by 1.8× and 1.5×, respectively. Guihai Yan, Wenyan Lu, Shuhao Jiang, Shijun Gong, Jingya Wu, Xiaowei Li 0001 |
DATE | 2 |
| 2018 | SmartShuttle: Optimizing off-chip memory accesses for deep learning acceleratorsabstractConvolutional Neural Network (CNN) accelerators are rapidly growing in popularity as a promising solution for deep learning based applications. Though optimizations on computation have been intensively studied, the energy efficiency of such accelerators remains limited by off-chip memory accesses since their energy cost is magnitudes higher than other operations. Minimizing off-chip memory access volume, therefore, is the key to further improving energy efficiency. However, we observed that sticking to minimizing the accesses of one data type as many prior work did cannot fit the varying shapes of convolutional layers in CNNs. Hence, there exists a dilemma of minimizing the accesses of which data type. To overcome the problem, this paper proposed an adaptive layer partitioning and scheduling scheme, called SmartShuttle, to minimize off-chip memory accesses for CNN accelerators. Smartshuttle can adaptively switch among different data reuse schemes and the corresponding tiling factor settings to dynamically match different convolutional layers. Moreover, SmartShuttle thoroughly investigates the impact of data reusability and sparsity on the memory access volume. The experimental results show that SmartShuttle processes the convolutional layers at 434.8 multiply and accumulations (MACs)/DRAM access for VGG16 (batch size = 3), and 526.3 MACs/DRAM access for AlexNet (batch size = 4), which outperforms the state-of-the-art approach (Eyeriss) by 52.2% and 52.6%, respectively. Guihai Yan, Wenyan Lu, Shuhao Jiang, Shijun Gong, Jingya Wu, Xiaowei Li 0001 |
DATE | 2 |
| 2018 | Tetris: re-architecting convolutional neural network computation for machine learning acceleratorsabstractInference efficiency is the predominant consideration in designing deep learning accelerators. Previous work mainly focuses on skipping zero values to deal with remarkable ineffectual computation, while zero bits in non-zero values, as another major source of ineffectual computation, is often ignored. The reason lies on the difficulty of extracting essential bits during operating multiply-and-accumulate (MAC) in the processing element. Based on the fact that zero bits occupy as high as 68.9% fraction in the overall weights of modern deep convolutional neural network models, this paper firstly proposes a weight kneading technique that could eliminate ineffectual computation caused by either zero value weights or zero bits in non-zero weights, simultaneously. Besides, a split-and-accumulate (SAC) computing pattern in replacement of conventional MAC, as well as the corresponding hardware accelerator design called Tetris are proposed to support weight kneading at the hardware level. Experimental results prove that Tetris could speed up inference up to 1.50x, and improve power efficiency up to 5.33x compared with the state-of-the-art baselines. Ning Lin, Guihai Yan, Xiaowei Li 0001 |
ICCAD | 4 |
| 2018 | AxTrain: Hardware-Oriented Neural Network Training for Approximate InferenceabstractThe intrinsic error tolerance of neural network (NN) makes approximate computing a promising technique to improve the energy efficiency of NN inference. Conventional approximate computing focuses on balancing the efficiency-accuracy trade-off for existing pre-trained networks, which can lead to suboptimal solutions. In this paper, we propose AxTrain, a hardware-oriented training framework to facilitate approximate computing for NN inference. Specifically, AxTrain leverages the synergy between two orthogonal methods---one actively searches for a network parameters distribution with high error tolerance, and the other passively learns resilient weights by numerically incorporating the noise distributions of the approximate hardware in the forward pass during the training phase. Experimental results from various datasets with near-threshold computing and approximation multiplication strategies demonstrate AxTrain's ability to obtain resilient neural network parameters and system energy efficiency improvement. Xin He 0011, Liu Ke 0001, Wenyan Lu, Guihai Yan, Xuan Zhang 0001 |
ISLPED | 4 |
| 2018 | Fault tolerance on-chip: a reliable computing paradigm using self-test, self-diagnosis, and self-repair (3S) approach
Xiaowei Li 0001, Guihai Yan, Jing Ye 0001, Ying Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2018 | CPicker: Leveraging Performance-Equivalent Configurations to Improve Data Center Energy Efficiency
Faqiang Sun, Guihai Yan, Xin He 0011, Huawei Li 0001, Yinhe Han 0001 |
J. Comput. Sci. Technol. | 2 |
| 2017 | ApproxEye: Enabling approximate computation reuse for microrobotic computer visionabstractAiming at real-life problems, microrobotic systems have gained more and more attention. However, limited achievable performance of microrobotic system prevents it from carrying out complex tasks. Current research work propose customize designs for different applications and incorporate dedicated accelerator for high energy efficiency. However, not only such techniques require significant manual effort and expertise for specified applications, but also the accelerator itself dictates unnegligible amount of chip resources. So in this paper we propose ApproxEye, a partial approximate computation reuse framework to accelerate microrobotic computer vision. Leveraging computation locality, ApproxEye reuses previous “similar” computations to reduce redundant computations. To squeeze every piece of computation reuse opportunity, ApproxEye proposes to 1) heuristically define optimal reuse granularity and 2) apply adaptive reuse requirements for different computations. Moreover, to reduce latency of computation reuse, ApproxEye tailors a parallel implemented search scheme for approximate computation reuse. Experimental results show ApproxEye could effectively exploit the potential of computation reuse and achieve 57.05% speedup on average. Xin He 0011, Guihai Yan, Faqiang Sun, Yinhe Han 0001, Xiaowei Li 0001 |
ASP-DAC | 2 |
| 2017 | FlexFlow: A Flexible Dataflow Accelerator Architecture for Convolutional Neural NetworksabstractConvolutional Neural Networks (CNN) are very computation-intensive. Recently, a lot of CNN accelerators based on the CNN intrinsic parallelism are proposed. However, we observed that there is a big mismatch between the parallel types supported by computing engine and the dominant parallel types of CNN workloads. This mismatch seriously degrades resource utilization of existing accelerators. In this paper, we propose a flexible dataflow architecture (FlexFlow) that can leverage the complementary effects among feature map, neuron, and synapse parallelism to mitigate the mismatch. We evaluated our design with six typical practical workloads, it acquires 2-10x performance speedup and 2.5-10x power efficiency improvement compared with three state-of-the-art accelerator architectures. Meanwhile, FlexFlow is highly scalable with growing computing engine scale. Wenyan Lu, Guihai Yan, Shijun Gong, Yinhe Han 0001, Xiaowei Li 0001 |
HPCA | 2 |
| 2016 | ACR: Enabling computation reuse for approximate computingabstractApproximate computing, which trades off computation quality (e.g, accuracy) and computation efforts, has becoming a promising technique to improve performance for many mission-non-critical and error-tolerant applications. The computations in such applications usually exhibit superior value locality, i.e, computations performed by a function or code region are very likely to reproduce “similar” results. Reusing the similar results can bypass redundant computations, as long as “exact” results are not mandatory. However, conventional computation reuse techniques are less effective in approximate computing paradigm. The input values of two computation instances have to be identical to reuse one for another, hence “exact” in nature.We propose ACR, an approximate computation reuse framework, to enable computation reuse for approximate computing. ACR relaxes the exact matching requirement in inputs to some extent regulated by “similarity” quantification, thereby shifting the exact computation reuse paradigm to its approximate counterpart. We furthermore propose an input significance-aware similarity quantification scheme through statistical approaches. Experimental result shows ACR could effectively exploit the potential of computation reuse for approximate computing and reduce 47.6% computations on average for a set of approximate applications. Xin He 0011, Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
ASP-DAC | 2 |
| 2016 | Wide Operational Range Processor Power Delivery Design for Both Super-Threshold Voltage and Near-Threshold Voltage Computing
Xin He 0011, Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
J. Comput. Sci. Technol. | 2 |
| 2016 | An Analytical Framework for Estimating Scale-Out and Scale-Up Power Efficiency of Heterogeneous ManycoresabstractHeterogeneous manycore architectures have shown to be highly promising to boost power efficiency through two independent ways: (1) enabling massive thread-level parallelism, called “scale-out” approach, and (2) enabling thread migration between heterogeneous cores, called “scale-up” approach. How to accurately model the profitability of power efficiency of the two ways, particularly in an analytical and computational-effective manner, is essential to reap the power efficiency of such architectures. We propose a comprehensive analytical model to predict the power efficiency from the two independent ways. Given power efficiency is measured by performance per watt, this model is composed of a performance and a power model. The performance model is built by two orthogonal functions a and β. Function a describes the scale-out speedup from multithreading; function β presents the scale-up speedup from core heterogeneity. Thus, the performance model can clearly capture the overall speedup of any multithreading and thread-to-core mapping strategies. The power model predicts the power of corresponding scale-out and scale-up configurations. It simultaneously captures the power variations caused by thread synchronization and thread migration between heterogeneous cores. We build both performance and power model in an analytical way and keep the computational complexity in mind. This merit leads to a suit of comprehensive and low-complexity models for runtime management. These models are validated on large-scale heterogeneous manycore architecture with full-system simulations. For performance prediction, the average error is below 12 percent, lower than that of the state-of-the-art methods. For power prediction, the average error is 7.74 percent. On top of the models, we introduce two heuristic scheduling algorithms, performance-oriented MAX-P and power efficiency-oriented MAX-E, to demonstrate the usage of these models. The results show that MAX-P outperforms the state-of-the-art methods by 18 percent in performance averagely; MAX-E outperforms the baseline by 70 percent in power efficiency on average. Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
IEEE Trans. Computers | 2 |
| 2016 | CoreRank: Redeeming "Sick Silicon" by Dynamically Quantifying Core-Level Healthy ConditionabstractIn field degradation of manycore processors poses a grand challenge to core management, largely because the degradation is hard to quantify. We propose a novel core-level degradation quantification scheme, CoreRank, to facilitate the management. We first develop a new degradation metric, called “healthy condition”, to capture the implication of performance degradation of a core with specific degraded components. Then, we propose a performance sampling scheme by using micro-operation streams, called snippet, to statistically quantify cores’ healthy condition. We find that similar snippets exhibit stable performance distribution, which makes them ideal micro-benchmarks to testify the core-level healthy conditions. We develop a hardware-implemented version of CoreRank based on bloom filter and hash table. Unlike the traditional “faulty” or “fault-free” judgement, CoreRank provides a key facility to make better use of those imperfect cores that suffered from various progressive aging mechanisms such as NBTI, HCI. Experimental results show that CoreRank successfully hides significant performance degradation of a defective manycore processor in which even more than half of the cores are salvaged from various defects. Guihai Yan, Faqiang Sun, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Computers | 1 |
| 2016 | EcoUp: Towards Economical Datacenter UpgradingabstractThe rapid growth of cloud services dictates increasingly powerful datacenters to maintain the high quality of service (QoS). It's a common practice in virtually all tiers of datacenters to continuously upgrade the datacenters, i.e. replacing outdated and failed servers with more advanced and efficient ones. However, how to upgrade a datacenter in the most cost-efficient strategy remains unclear, and however this problem goes increasingly challenging given the great diversity of applications. In practice, the datacenters' operators usually resort to expending the scale of servers. The preferred servers are either expensive but high-performance, or, by contrast, cheap but low-power. Whatever sever preferences, how to justify the cost-efficiency is still an open problem. We claim that a cost-efficient upgrading strategy should be fully aware of not only the capacity and cost of various servers, but also the resource demands of target applications. We model this strategy as a recommendation problem: recommending the “best” servers to a datacenter. We propose “EcoUp”, a model-based framework that faithfully rates the cost efficiency of server candidates, relying on which an optimal server portfolio can be derived. The performance prediction on candidate servers is realized by employing a sophisticated latent factor model (LFM). The cost mainly involves the server purchasing cost and energy bill. Given the application distribution, EcoUp can give an optimal server portfolio under a certain capital budget. We use Google trace, a big profiling dataset opened by Google, to validate the performance prediction. Experimental results show that the error rate is below 8 percent on average. Meanwhile, we build a comprehensive upgrading procedure on a local cluster to evaluate the potential of EcoUp. The results show that our approach significantly outperforms two conventional upgrading strategies by 12.3 and 33.6 percent in terms of system throughput, respectively. Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | ShuttleNoC: Boosting on-chip communication efficiency by enabling localized power adaptationabstractNetworks-on-Chip (NoC) gradually becomes a main contributor of chip-level power consumption. Due to the temporal and spatial heterogeneity of on-chip traffic, existing power management approaches cannot adapt the NoC power consumption to its traffic intensity, and hence lead to a suboptimal power efficiency. They either resort to over-provisioned NoC design that only suits for traffic spatial distribution, or coarse-grained power gating that only serves traffic temporal variation. In this paper, we propose a novel NoC architecture called Shuttle Networks-on-Chip (ShuttleNoC). By permitting packets shuttling between multiple subnetworks, localized power adaptation can be achieved. Experimental results show that ShuttleNoC could achieve optimal power efficiency with up to 23.5% power savings and 22.3% performance boost in comparison with traditional heterogeneity-agnostic NoC designs. Guihai Yan, Yinhe Han 0001, Ying Wang 0001, Xiaowei Li 0001 |
ASP-DAC | 2 |
| 2015 | RISO: Enforce Noninterfered Performance With Relaxed Network-on-Chip Isolation in Many-Core Cloud ProcessorsabstractWorkload consolidation is widely used in modern cloud processors to reduce total cost of ownership. Performance isolation has to be enforced between consolidated workloads to achieve controllable quality of service. Networks-on-chip (NoCs), as a major shared resource, often incur traffic interference and violate performance isolation criteria. Previous work resorts to strict isolation strategy that partitions NoC into independent regions to isolate core-to-core communication traffic. However, strict isolation either results in low consolidation density or degrades network performance, and more importantly, cannot be applied to memory access traffic. To address these weaknesses, we propose a novel performance isolation strategy in NoC, called relaxed isolation (RISO). It permits underutilized routers and links to be shared by multiple applications, and, at the same time, it keeps the aggregated traffic in check to enforce performance isolation. Experimental results show that RISO could effectively improve consolidation density and network performance in synergy. Binzhang Fu, Ying Wang 0001, Yinhe Han 0001, Guihai Yan, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2014 | Amphisbaena: Modeling two orthogonal ways to hunt on heterogeneous many-coresabstractHeterogeneous many-cores can deliver high performance or energy efficiency. There are two orthogonal ways to improve performance: 1) scale-out by exploiting thread-level parallelism, and 2) scale-up by enabling core heterogeneity. Predicting the performance of such architecture is increasingly challenging. We propose a comprehensive performance model Amphisbaena, or Φ, built from two orthogonal functions α and β. Function α describes the scale-out speedup and function β handles the scale-up speedup. The Φ model can clearly tell not only the overall speedup of a given multithreading and core mapping strategy, but also how to improve the multithreading and core mapping, hence should be a promising performance predictor for future heterogenous many-cores. The results show that Φ model's error rate is within 12%, which is lower than state-of-the-art methods. We demonstrate the application of Φ model by introducing a heuristic scheduling algorithm, which outperforms the baselines by 13% on average. Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
ASP-DAC | 2 |
| 2014 | On-Chip Delay Sensor for Environments with Large Temperature FluctuationsabstractThe precision of on-chip delay sensors is degraded by temperature fluctuations, which hinders these sensors from applying to on-line fault predicting and DVFS. We present a novel path delay measuring technique which is immune to large temperature fluctuation. The delay reference are generated by gate biasing temperature compensation devices in which the pull-up and pull-down network are tuned to set the measurement circuit working in temperature insensitive point, thereby eliminating the precision degradation due to temperature variations. Video image scaling IP is used as experimental circuit to validate the effectiveness of the proposed technique. Experimental results show that within temperature range of -55°C to 125°C, the measurement error is reduced from 19.56% to 0.5%, compared with the techniques without temperature resilience. Jibing Qiu, Guihai Yan, Xiaowei Li 0001 |
ATS | 2 |
| 2014 | SuperRange: Wide operational range power delivery design for both STV and NTV computingabstractThe load power range of modern processors is greatly enlarged because many advanced power management techniques like dynamic voltage frequency scaling, Turbo boosting, and Near Threshold Voltage technologies are incorporated. However, the power saving may be offset by power loss in power delivery; moreover, as the efficiency of power delivery varies greatly with different load conditions, conventional power delivery designs cannot maintain high efficiency over the entire voltage range. We propose SuperRange, a wide operational range power delivery scheme. SuperRange complements the power delivery capability of on-chip voltage regulator and off-chip voltage regulator. Experimental results show SuperRange has an average 70% power conversion efficiency over wide operational range which outperforms conventional power delivery schemes. And it also exhibits superior resilience to power-constrained systems. Xin He 0011, Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2014 | SmartCap: Using Machine Learning for Power Adaptation of Smartphone's Application ProcessorabstractPower efficiency is increasingly critical to battery-powered smartphones. Given that the using experience is most valued by the user, we propose that the power optimization should directly respect the user experience. We conduct a statistical sample survey and study the correlation among the user experience, system runtime activities, and computational performance of an application processor. We find that there exists a minimal frequency requirement, called “saturated frequency”. Above this frequency, the device consumes more power but provides little improvements in user experience. This study motivates an intelligent self-adaptive scheme, SmartCap, that automatically identifies the most power-efficient state of the application processor. Compared to prior Linux power adaptation schemes, SmartCap can help save power from 11% to 84%, depending on applications, with little decline in user experience. Xueliang Li 0004, Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | Orchestrator: Guarding Against Voltage Emergencies in Multithreaded ApplicationsabstractVoltage emergency (VE) has become a critical challenge with decreasing feature size and increasing power capacity. Destructive core interference is one main source of VE in multicore processors. We observed that the applications following single program and multiple data programming model tend to spark domain-wide destructive core interference because multiple threads exhibit similar power activity. We analyze and quantify this effect and propose one low-cost solution, Orchestrator, to avoid voltage droop synergy among cores. Orchestrator leverages the thread diversity to smooth voltage droops in multicore architectures based on thread scheduling. The thread migration impact on performance is also considered. Experimental results show that Orchestrator can significantly reduce VEs, thereby improving performance. Xing Hu 0001, Guihai Yan, Yu Hu 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | RISO: relaxed network-on-chip isolation for cloud processorsabstractCloud service providers use workload consolidation technique in many-core cloud processors to optimize system utilization and augment performance for ever extending scale-out workloads. Performance isolation usually has to be enforced for the consolidated workloads sharing the same many-core resources. Networks-on-chip (NoC) serves as a major shared resource, also needs to be isolated to avoid violating performance isolation. Prior work uses strict network isolation to fulfill performance isolation. However, strict network isolation either results in low consolidation density, or complex routing mechanisms which indicates prohibitive high hardware cost and large latency. In view of this limitation, we propose a novel NoC isolation strategy for many-core cloud processors, called relaxed isolation (RISO). It permits underutilized links to be shared by multiple applications, at the same time keeps the aggregated traffic in check to enforce performance isolation. The experimental results show that the consolidation density is improved more than 12% in comparison with previous strict isolation scheme, meanwhile reducing network latency by 38.4% on average. Guihai Yan, Yinhe Han 0001, Binzhang Fu, Xiaowei Li 0001 |
DAC | 2 |
| 2013 | Orchestrator: a low-cost solution to reduce voltage emergencies for multi-threaded applicationsabstractVoltage emergencies have become a major challenge to multi-core processors because core-to-core resonance may put all cores into danger which jeopardizes system reliability. We observed that the applications following SPMD (Single Program and Multiple Data) programming model tend to spark domain-wide voltage resonance because multiple threads sharing the same function body exhibit similar power activity. When threads are judiciously relocated among the cores, the voltage droops can be greatly reduced. We propose “Orchestrator”, a sensor-free non-intrusive scheme for multi-core architectures to smooth the voltage droops. Orchestrator focuses on the inter-core voltage interactions, and maximally leverages the thread diversity to avoid voltage droops synergy among cores. Experimental results show that Orchestrator can reduce up to 64% voltage emergencies on average, meanwhile improving performance. Xing Hu 0001, Guihai Yan, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2013 | SmartCap: user experience-oriented power adaptation for smartphone's application processorabstractPower efficiency is increasingly critical to battery-powered smartphones. Given the using experience is most valued by the user, we propose that the power optimization should directly respect the user experience. We conduct a statistical sample survey and study the correlation among the user experience, the system runtime activities, and the minimal required frequency of an application processor. This study motivates an intelligent self-adaptive scheme, SmartCap, which automatically identifies the most power-efficient state of the application processor according to system activities. Compared to prior Linux power adaptation schemes, SmartCap can help save power from 11% to 84%, depending on applications, with little decline in user experience. Xueliang Li 0004, Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2012 | AgileRegulator: A hybrid voltage regulator scheme redeeming dark silicon for power efficiency in a multicore architectureabstractThe widening gap between the fast-increasing transistor budget but slow-growing power delivery and system cooling capability calls for novel architectural solutions to boost energy efficiency. Leveraging the fact of surging “dark silicon” area, we propose a hybrid scheme to use both on-chip and off-chip voltage regulators, called “AgileRegulator”, for a multicore system to explore both coarse-grain and fine-grain power phases. We present two complementary algorithms: Sensitivity-Aware Application Scheduling (SAAS) and Responsiveness-Aware Application Scheduling (RAAS) to maximally achieve the energy saving potential of the hybrid regulator scheme. Experimental results show that the hybrid scheme achieves performance-energy efficiency close to per-core DVFS, without imposing much design cost. Meanwhile, the silicon overhead of this scheme is well contained into the “dark silicon”. Unlike other application specific schemes based on accelerators, the proposed scheme itself is a simple and universal solution for chip area and energy trade-offs. Guihai Yan, Yingmin Li, Yinhe Han 0001, Xiaowei Li 0001, Minyi Guo, Xiaoyao Liang |
HPCA | 1 |
| 2011 | Online timing variation tolerance for digital integrated circuitsabstractEnsuring safe timing increasingly becomes a paramount challenge with the technology scaling to nanoscale. This study aims to provide timing variation detection and tolerance solutions. We first propose a versatile online timing variation detection scheme which can handle multiple types of faults. With the capability of detection, we further propose two tolerance schemes to eliminate runtime margin in DVFS applications and improve lifetime reliability under progressive aging mechanisms, respectively. Lastly, given the more complicated PVT variations whose primary circuit implication is also timing variations, we propose TEA-TM, a novel architectural scheme to reduce timing emergencies. Collectively, we aims to build a comprehensive framework for timing variation tolerance and demonstrate several specific applications. Guihai Yan, Xiaowei Li 0001 |
ITC | 1 |
| 2011 | ReviveNet: A Self-Adaptive Architecture for Improving Lifetime Reliability via Localized Timing AdaptationabstractThe aggressive technology scaling poses serious challenges to lifetime reliability. A parament challenge comes from a variety of aging mechanisms that can cause gradual performance degradation of circuits. Prior work shows that such progressive degradation can be reliably detected by dedicated aging sensors, which provides a good foundation for proposing a new scheme to improve lifetime reliability. In this paper, we propose ReviveNet, a hardware-implemented aging-aware and self-adaptive architecture. Aging awareness is realized by deploying dedicated aging sensors, and self-adaptation is achieved by employing a group of synergistic agents. Each agent implements a localized timing adaptation mechanism to tolerate aging-induced delay on critical paths. On the evaluation, a reliability model based on widely used weibull distribution is presented. Experimental results show that, without compromising with any nominal architectural performance, ReviveNet can improve the Mean-Time-To-Failure by up to 48.7 percent, at the expense of 9.5 percent area overhead and small power increase. Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
IEEE Trans. Computers | 1 |
| 2011 | MicroFix: Using timing interpolation and delay sensors for power reductionabstractTraditional DVFS schemes are oblivious to fine-grained adaptability resulting from path-grained timing imbalance. With the awareness of such fine-grained adaptability, better power-performance efficiency can be obtained. We propose a new scheme, MicroFix, to exploit such fine-grained adaptability. We first show the potential resulted from the path-grained timing imbalance and then present a new technique, Timing Interpolation, to reap the fine-grained adaptability for power reduction. Moreover, to eliminate the conservative margins of traditional DVFS, unlike the previous approaches such as Razor that reactively handle the delay errors (induced by aggressively scaled voltage/frequcncy) by enabling error detection and recovery, we propose a proactive approach by error prediction, thereby obviate the high-cost recovery routines. MicroFix was evaluated based on ISCAS89 benchmarks and the floating-point unit adopted by OpenSPARC T1 processor. Compared to ideal traditional DVFS schemes, the experimental results show that for most of the evaluated circuits, MicroFix can help saving up to 20% power consumption without compromising with frequency, at the expense of less than 5% area overhead. Compared to nonideal DVFS schemes (with 10% voltage margin), the power reduction can even reach up to 38% on average. Guihai Yan, Yinhe Han 0001, Xiaoyao Liang, Xiaowei Li 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2011 | SVFD: A Versatile Online Fault Detection Scheme via Checking of Stability ViolationabstractIn ultra-deep submicrometer technology, soft errors and device aging are two of the paramount reliability concerns. Although many studies have been done to tackle the two challenges, most take them separately so far, thereby failing to reach better performance-cost tradeoffs. To support a more efficient design tradeoff, we propose a unified fault detection scheme—stability violation-based fault detection (SVFD), by which the soft errors (both single event upset and single event transient), aging delay, and delay faults can be uniformly dealt with. SVFD grounds on a new fault model, stability violation, derived from analysis of signal behavior. SVFD has been validated by conducting a set of intensive Hspice simulations targeting the next-generation 32-nm CMOS technology. An application of SVFD to a floating-point unit (FPU) is also evaluated. Experimental results show that SVFD has more versatile fault detection capability for fault detection than several schemes recently proposed at comparable overhead in terms of area, power, and performance. Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | Leveraging the core-level complementary effects of PVT variations to reduce timing emergencies in multi-core processorsabstractProcess, Voltage, and Temperature (PVT) variations can significantly degrade the performance benefits expected from next nanoscale technology. The primary circuit implication of the PVT variations is the resultant timing emergencies. In a multi-core processor running multiple programs, variations create spatial and temporal unbalance across the processing cores. Most prior schemes are dedicated to tolerating PVT variations individually for a single core, but ignore the opportunity of leveraging the complementary effects between variations and the intrinsic variation unbalance among individual cores. We find that the notorious delay impacts from different variations are not necessary aggregated. Cores with mild variations can share the violent workload from cores suffering large variations. If operated correctly, variations on different cores can help mitigating each other and result in a variation-mild environment. In this paper, we propose Timing Emergency Aware Thread Migration (TEA-TM), a delay sensor-based scheme to reduce system timing emergencies under PVT variations. Fourier transform and frequency domain analysis are conducted to provide the insights and the potential of the PVT co-optimization scheme. Experimental results show on average TEA-TM can help save up to 24% throughput loss, at the same time improve the system fairness by 85%. Guihai Yan, Xiaoyao Liang, Yinhe Han 0001, Xiaowei Li 0001 |
ISCA | 1 |
| 2010 | Performance-asymmetry-aware scheduling for Chip Multiprocessors with static core coupling
Jianbo Dong, Lei Zhang 0008, Yinhe Han 0001, Guihai Yan, Xiaowei Li 0001 |
J. Syst. Archit. | 4 |
| 2009 | M-IVC: Using Multiple Input Vectors to Minimize Aging-Induced DelayabstractNegative bias temperature instability (NBTI) has been a significant reliability concern in current digital circuit design due to its effect of increasing the path delay with time and in turn degrading the circuit performance. NBTI degradation has strong dependence on input pattern and duty cycles. Based on this observation, we propose to apply multiple input vectors to the combination circuit in a non-uniform way during standby mode. Multiple input vectors can enhance the capability to control the circuit nodes, achieve smaller duty cycles to reduce the stress time of gates and thus mitigate static NBTI. A constrained multi-object optimization model is formalized to find the optimal combination of duty cycles for timing-critical paths, which in turn minimizes the increase of path delay. An ATPG-like procedure is then presented to generate the corresponding input vectors. Experimental results demonstrate that the delay increase of timing-critical paths can be mitigated significantly under long time NBTI effect (10-year) by only applying a small number of vectors. Yinhe Han 0001, Lei Zhang 0008, Huawei Li 0001, Xiaowei Li 0001, Guihai Yan |
Asian Test Symposium | 6 |
| 2009 | A unified online Fault Detection scheme via checking of Stability ViolationabstractIn ultra-deep submicro technology, two of the paramount reliability concerns are soft errors and device aging. Although intensive studies have been done to face the two challenges, most take them separately so far, thereby failing to reach better performance-cost tradeoffs. To support a more efficient design tradeoff, we present a new fault model, stability violation, derived from analysis of signal behavior. Furthermore, we propose a unified fault detection scheme-stability violation based fault detection (SVFD), by which the soft errors (both single event upset and single event transient), aging delay, and delay faults can be uniformly handled. SVFD can greatly facilitate soft error-resistant and aging-aware designs. SVFD is validated by conducting a set of intensive Hspice simulations targeting 65 nm CMOS technology. Experimental results show that SVFD has more robust capability for fault detection than previous schemes at comparable overhead in terms of area, power, and performance. Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
DATE | 1 |
| 2009 | MicroFix: exploiting path-grained timing adaptability for improving power-performance efficiencyabstractTraditional DVFS schemes are oblivious to fine-grained adaptability resulting from path-grained timing imbalance. With the awareness of such fine-grained adaptability, better power-performance efficiency can be obtained. We propose a new approach, MicroFix, to exploit such fine-grained adaptability. We first reveal the potential of the path-grained timing imbalance and then present a novel implementation of MicroFix. Moreover, to eliminate the conservative margins of traditional DVFS, unlike the previous approaches that reactively handle the delay errors (induced by aggressively scaled voltage/frequcncy) by error detection and recovery strategies, we propose a proactive approach by error prediction. MicroFix was evaluated based on the floating-point unit adopted by OpenSPARC T1 processor. Compared against traditional DVFS schemes, the experimental results shows that MicroFix improves the EDP (Energy-Delay Product) up to 35% for high-performance circuits and PDP (Power-Delay Product) to 28% for low-power circuits, while at the expense of only 7% area overhead. Guihai Yan, Yinhe Han 0001, Xiaoyao Liang, Xiaowei Li 0001 |
ISLPED | 1 |
| 2009 | Variation-Aware Scheduling for Chip Multiprocessors with Thread Level RedundancyabstractThread-level redundancy in Chip Multiprocessors(TLR-CMP) is efficient for soft error tolerance. Process variation causes core-to-core (C2C) performance asymmetry across a chip, which should be taken into consideration for application scheduling. In this paper, two types of variations beyond C2C are introduced, i.e., inter-pair and intra-pair variation in TLR-CMP. Intra-pair performance asymmetry can affect the performance of applications differently. Based on the above observation, we firstly formalize the variation aware scheduling in TLR-CMP as a 0-1 programming problem,to maximize the system weighted throughput. An efficient scheduling algorithm, named IntraVarF&AppSen, is then proposed to tackle this problem, which can be proved to be optimal when the number of applications to be scheduled is equal to the number of core pairs. Simulation on a 64-core CMP shows 2.8%-4% improvement in weighted throughput when compared to prior VarF&AppIPC algorithm. Jianbo Dong, Lei Zhang 0008, Yinhe Han 0001, Guihai Yan, Xiaowei Li 0001 |
PRDC | 4 |