Yihan Hu 0003

dblp:268/5967-3 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2024
0009-0004-8469-1845ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2024 An Agile Deploying Approach for Large-Scale Workloads on CGRA-CPU Architecture
abstract
Adopting specialized accelerators such as Coarse-Grained Reconfigurable Architectures (CGRAs) alongside CPUs to enhance performance within specific domains is an astute choice. However, the integration of heterogeneous architectures introduces complex challenges for compiler design. Simultaneously, the ever-expanding scale of workloads imposes substantial burdens on deployment. To address above challenges, this paper introduces CGRV-OPT, a user-friendly multi-level compiler designed to deploy large-scale workloads to CGRA and RISC-V CPU architecture. Built upon the MLIR framework, CGRV-OPT serves as a pivotal bridge, facilitating the seamless conversion of high-level workload descriptions into low-level intermediate representations (IRs) for different architectures. A salient feature of our approach is the automation of a comprehensive suite of optimizations and transformations, which speed up each kernel computing within the intricate SoC. Additionally, we have seamlessly integrated an automated software-hardware partitioning mechanism, guided by our multi-level optimizations, resulting in a remarkable 2.14 × speed up over large-scale workloads. The CGRV-OPT framework significantly alleviates the challenges faced by software developers, including those with limited expertise in hardware architectures.
Jiahang Lou, Xuchen Gao, Yiqing Mao, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang
DATE5
2024 HETA: A Heterogeneous Temporal CGRA Modeling and Design Space Exploration via Bayesian Optimization
abstract
Due to its high energy efficiency and flexibility, coarse-grained reconfigurable architecture (CGRA) has gained increasing attention. Temporal CGRA is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform spatial and temporal computations. Although multiple temporal CGRAs have been proposed, an architecture with rich design parameters and heterogeneous modeling is still lacking. To this end, we propose a highly parameterized heterogeneous temporal CGRA, called HETA. However, the highly parameterized and heterogeneous design introduces a challenging design space for manual exploration. To address this challenge, we introduce a Bayesian-optimization (BO)-based design space exploration (DSE) of homogeneous and heterogeneous architectures. Different from other DSE processes that require defining the heterogeneous exploration strategy, our approach adopts a searching-pruning-based method without manual intervention. To improve the efficiency of DSE, we develop a fast statistic model for area evaluation, whose error is below 1%. In addition, a pipeline mapping (PiPMap) algorithm is developed to alleviate the restrictions caused by data synchronization and unleash the potential of the proposed architecture. Experimental results show that HETA can achieve 89%, 52%, and 47% improvement in throughput, area efficiency, and energy efficiency over the neighbor-to-neighbor (N2N)-based interconnect CGRA, respectively. Compared with the Switch-based interconnect CGRA, HETA’s area efficiency is increased by 61%. Furthermore, compared with the homogeneous architecture of HETA, the optimized heterogeneous architecture improves area efficiency and energy efficiency by 14.7% and 4.8%, respectively.
Yuan Dai, Jingyuan Li 0003, Qilong Zhu, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang
IEEE Trans. Very Large Scale Integr. Syst.5
2021 LETA: A lightweight exchangeable-track accelerator for efficientnet based on FPGA
abstract
Lightweight convolutional neural networks (CNNs) have become increasingly popular due to their lower computational complexity and fewer memory accesses with equivalent accuracy compared to previous CNN models. However, the newly proposed networks bring new challenges to efficient hardware design, such as, in EfficientNet, depthwise convolution, squeeze-and-excitation (SE) module, and swish/sigmoid functions. Although individual engine architecture could achieve a high computing efficiency for the standard convolution or the depth-wise convolution, it is still not efficient for EfficientNet because the workload imbalance between two types of convolutional engines causes inevitable idling. To overcome this problem, we present a lightweight reconfigurable computational kernel based on FPGA with an exchangeable-track datapath scheme. In addition, a low-accuracy-loss function replacement strategy is proposed for swish/sigmoid functions. Furthermore, the low-cost hardware architecture to implement the replaced functions is designed. The proposed accelerator (LETA) can implement EfficientNet on Xilinx XCVU37P with a 300 MHz system clock and a 600 MHz kernel clock. The linear growth of resource usage in the 4-kernel implementation in 1 super logic region (SLR) with the same clock frequencies justifies the scalability of LETA. The experimental results show that LETA can achieve 2× throughput/DSP compared to the latest FPGA-based accelerator with 1.6% (0.7%) top-1 (top-5) accuracy loss on EfficientNet-B3.
Jingbo Gao, Yihan Hu 0003, Xitian Fan, Wai-Shing Luk, Wei Cao 0002, Lingli Wang
FPT3