Zhihua Fan

dblp:85/937 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
28since 2021 · last 2026
0000-0002-5950-7370ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 5 first-author · 28 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 BitRed: Taming Non-Uniform Bit-Level Sparsity with a Programmable RISC-V ISA for DNN Acceleration
abstract
The non-uniform and dynamic nature of Bit-Level Sparsity (BLS) poses a critical load-imbalance challenge for parallel hardware accelerators. While the Bit-Interleaving paradigm, represented by state-of-the-art accelerators like Bitlet, shows promise, it is fundamentally constrained by a rigid datapath and severe inter-channel load imbalance. This paper introduces BitRed, an accelerator that embodies a new ''programmable adaptive bit-interleaving'' philosophy. Rather than a monolithic design, BitRed's core is an Adaptive-Sparse Processing Unit (ASPU) that deconstructs the acceleration process into a set of orthogonal RISC-V ISA extensions for pre-processing (cal.pre), adaptive distillation with dynamic load balancing (cal.adis), and PDP-optimal reduction (cal.red). By transforming a rigid hardware problem into a flexible scheduling problem, this ISA-based approach provides a fundamentally more adaptable and extensible solution. Empirical studies on a broad set of benchmarks highlight the following results (normalized to a SCNN baseline): (1) up to 9.4× speedup over the Bitlet, and 5.6× over the latest bit-serial SOTA BitWave; (2) up to 7.6× higher inference efficiency than Bitlet on representative models; (3) 5.072mm2 area and scalable power consumption from 550.43mW (float32) to 495.12mW (16b) and 457.90mW (8b) @28nm TSMC; and (4) high versatility across precisions, and up to 18.9×13.5× higher than NVIDIA A100 and Jetson Orin 32GB, demonstrating significant competitiveness against GPUs.
Yanhuan Liu, Kunming Zhang, Yuqun Liu, Siao Wen, Lexin Wang, Tianyu Liu 0007, Zhihua Fan, Xiaochun Ye, Dongrui Fan, Xuejun An
ASPLOS (2)9
2026 RISC-V ISA Extensions for Vectorized Unstructured Sparse SpMM in LLM Inference
abstract
Unstructured sparsity has emerged as a key enabler for pruning large language models (LLMs) while preserving accuracy. However, its highly irregular pattern makes it notoriously difficult to accelerate, creating severe bottlenecks in vectorization and memory access that prevent efficient deployment on edge hardware with tight power and area constraints. We present SCG, a vectorizable sparse matrix format designed to unlock high-performance unstructured sparse matrix–matrix multiplication (SpMM), the dominant kernel in LLM feed-forward networks and Q/K/V/O projections. To exploit SCG, we introduce custom RISC-V instructions and extend the BOOM processor with two lightweight pipelines for efficient parallel execution. This format–instruction–hardware co-design directly addresses the long-standing challenge of unstructured sparse acceleration in general-purpose processors. On real LLM workloads, our design achieves 11.9×, 12.7×, and 13.4× average speedups over baseline BOOM on LLaMA2-7B, OPT-1.3B, and TinyLLaMA-1.1B, respectively, with negligible hardware overhead. Compared to state-of-the-art sparse accelerators, it delivers up to 1.72× higher area efficiency.
Tengfei Xia, Zhihua Fan, Shantian Qin, Xiaochun Ye
DATE2
2026 A2RT: Efficient Ray Tracing Accelerator with Approximate-Accurate Computing and Quantization
abstract
Ray tracing (RT) has revolutionized photorealistic rendering by simulating light transport, but existing methods face a trade-off between computational efficiency and rendering accuracy. To address this, we present A2RT, a software-hardware co-designed RT accelerator employing the end to end optimization of "quantization → computation". On the software side, we introduce a customized data flow mechanism with type-specific quantization for bounding boxes, ray origins, and directions, and we organize BVH nodes into Group- and Sub-Nodes. At the hardware level, a heterogeneous RT engine allocates resources based on node criticality: accurate computing units handle Group-Nodes, while approximate units process Sub-Nodes. A custom INT-FLOAT approximate multiplier further accelerates the approximate units. Experimental results show that A2RT achieves 45.51% energy consumption and 2.29× speedup over RT Core, and consumes 81.79% of energy while delivering 1.57× performance improvement compared to state-of-the-art accelerators.
Zhihua Fan, Yudong Mu, Zhen Wang 0045, Xiaochun Ye, Xuejun An
DATE2
2026 HGNNMap: Heterogeneous Graph Neural Network-Based Mapping for Spatial Accelerators
Shengzhong Tang, Zhihua Fan, Tianyu Liu 0007, Xuejun An, Xiaochun Ye
Euro-Par (1)2
2026 MLX: Multi-Layer Execution for Structured LLM Workload Acceleration on Spatial Architectures
Zhihua Fan, Zirui Ma, Yuqun Liu, Tengfei Xia, Yanhuan Liu, Kunming Zhang, Xiaochun Ye, Dongrui Fan, Jian Weng 0002
ISCA3
2026 HALO: A heterogeneous accelerator for low-latency and energy-efficient edge LLM inference
Kunming Zhang, Zhihua Fan, Yanhuan Liu, Lexin Wang, Yuqun Liu, Xiaochun Ye
Future Gener. Comput. Syst.2
2026 A real-time edge SAR imaging acceleration architecture utilizing multi-level dataflow parallelism
Yinshen Wang, Zhengxuan Hu, Zhihua Fan, Xuejun An, Xiaochun Ye
J. Syst. Archit.4
2026 A RISC-V Extended Infrastructure for Edge FHE Through Software and Hardware Co-Design
abstract
Fully Homomorphic Encryption (FHE) is a foundational technique in privacy-preserving computation, enabling secure data processing without decryption. However, existing acceleration approaches for edge-side FHE suffer from several challenges, including limited speedup, complex hardware designs, and underutilization of CPU resources. To address these issues, we propose a lightweight yet effective hardware-software co-design acceleration scheme based on RISC-V architecture. First, we propose a custom RISC-V instruction set extension tailored for fully homomorphic encryption, enabling efficient and fine-grained acceleration of modular arithmetic operations. Second, we design an efficient FHE acceleration architecture, including customized circuit implementations and a pipelined execution unit, to support low-latency and energy-efficient computation. Third, we introduce a parallel software acceleration strategy based on butterfly computation patterns and multi-threading techniques, fully utilizing the RISC-V vector extension and multi-core resources. Experimental results demonstrate that our solution achieves an average speedup of 11.5× and an energy efficiency improvement of 7.05×. Compared to the current state-of-the-art design, our approach delivers an average performance gain of 2.27× with a significantly reduced design complexity.
Zhihua Fan, Xuejun An, Xiaochun Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 NFMap: Node Fusion Optimization for Efficient CGRA Mapping with Reinforcement Learning
Yudong Mu, Zhihua Fan, Xuejun An, Xiaochun Ye
APPT3
2025 Accelerating Authenticated Block Ciphers via RISC-V Custom Cryptography Instructions
abstract
As one of the standardized encryption algorithms, authenticated block ciphers based on Galois/Counter Mode (GCM) is a widely-used method to guarantee the accuracy and reliability in data transmission. The profiling work demonstrates that across the execution process of GCM mode, the authentication operation is the main performance bottleneck because it introduces operations in high-dimensional Galois field (GF), which could not be efficiently executed via existing ISA. To overcome this problem, we propose a custom ISA extension and cooperate it with RISC-V cryptography extension to accelerate the whole process of authenticated block ciphers. Besides, we design a specific crypto core including a fully-pipelined GF(2128) multiplier to support the extended instructions and integrate it into the multi-issue out-of-order core XT910 without introducing any clock frequency overhead. The proposed design significantly reduces the the number of instructions required in the main operations of authenticated block ciphers. We compare the performance of our designs to other existing acceleration scheme based on RISC-V ISA extension. Experimental result shows that our design outperforms other related work and achieves up to 17 × speedup with a lightweight hardware overhead.
Tianyu Liu 0007, Zhen Wang 0045, Zhihua Fan, Xiaochun Ye, Dongrui Fan
DATE6
2025 FDHA: Fusion-Driven Heterogeneous Accelerator for Efficient Diffusion Model Inference
Yudong Mu, Zhihua Fan, Xiaoxia Yao, Honglie Wang, Xuejun An, Xiaochun Ye
Euro-Par (2)2
2025 TSCNN: Compressing and Accelerating Sparse CNNs Using Sign-Reserved Toeplitz Filters
abstract
Exploiting the sparsity in convolutional neural networks (CNNs) is crucial to accelerate computing and reduce energy consumption. However, unstructured sparsity often introduces irregularity in convolutional operations, which complicates the control logic and undermines the benefits of sparsification. Structured sparsity alleviates these problems but sacrifices the adaptability to arbitrary sparse patterns. In this paper, we propose TSCNN, an algorithm-hardware co-design solution that aims to compress and accelerate sparse CNNs while balancing both adaptability to sparsity and computational efficiency. In terms of algorithm, TSCNN adopts pruned filters compressed with sign-reserved Toeplitz matrix format (Tfilters), which systematically enhances the regularity of data reuse and flexibly reduces network parameters by$44 \%-86 \%$while maintaining accuracy. In terms of hardware, TSCNN accelerator employs custom computing components to adapt to the structure of Tfilters and support the adaptive dataflow, further optimizing the computational efficiency. Experiments show that TSCNN outperforms a dense accelerator, SCNN and CSCNN, achieving$4.49 \times, 2.29 \times, 2.08 \times$and$1.29 \times$speedup and reducing energy consumption by$74.65 \%, 41.04 \%, 49.29 \%$and 43.66%, respectively.
Zhen Wang 0045, Tianyu Liu 0007, Zhihua Fan, Xiaochun Ye, Dongrui Fan
HPCC3
2025 StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
abstract
Multimodal Transformers are emerging artificial intelligence (AI) models designed to process a mixture of signals from diverse modalities. Digital computing-in-memory (CIM) architectures are considered promising for achieving high efficiency while maintaining high accuracy. However, current digital CIM-based accelerators exhibit inflexibility in microarchitecture, dataflow, and pipeline to effectively accelerate multimodal Transformer. In this paper, we propose StreamDCIM, a tile-based streaming digital CIM accelerator for multimodal Transformers. It overcomes the above challenges with three features: First, we present a tile-based reconfigurable CIM macro microarchitecture with normal and hybrid reconfigurable modes to improve intra-macro CIM utilization. Second, we implement a mixed-stationary cross-forwarding dataflow with tile-based execution decoupling to exploit tile-level computation parallelism. Third, we introduce a ping-pong-like fine-grained compute-rewriting pipeline to overlap high-latency on-chip CIM rewriting. Experimental results show that StreamDCIM outperforms non-streaming and layer-based streaming CIM-based solutions by geomean 2.63×and 1.28× on typical multimodal Transformer models.
Shantian Qin, Ziqing Qiang, Zhihua Fan, Xuejun An, Xiaochun Ye, Dongrui Fan
ISCAS3
2025 LightCacheRL: A Lightweight Reinforcement Learning Framework for Unified Cache Management
Kunming Zhang, Zhihua Fan, Yingchun Fu, Yanhuan Liu, Lexin Wang, Yuqun Liu
NPC (1)2
2025 Accelerating tensor multiplication by exploring hybrid product with hardware and software co-design
Zhihua Fan, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Xuejun An
J. Syst. Archit.2
2025 GenCNN: A Partition-Aware Multi-Objective Mapping Framework for CNN Accelerators Based on Genetic Algorithm
abstract
Convolutional Neural Networks (CNNs) require partitioning to efficiently run on CNN accelerators, which offer multiple parallel processing dimensions, such as Processing Element (PE) array topologies and Single Instruction Multiple Data (SIMD) execution. The choice of parallelization strategy directly impacts accelerator performance. However, the vast search space for CNN partitioning and parallelization makes manual optimization costly and complex, especially when addressing both aspects simultaneously. This highlights the need for an automated framework to efficiently map CNNs onto accelerators. Our key insight is that existing approaches suffer from inadequate accelerator performance modeling and a lack of multi-objective optimization strategies that jointly consider task partitioning and convolution parallelization. To address this, we propose GenCNN, a multi-objective genetic algorithm-based mapping framework for CNN accelerators. GenCNN first constructs a fine-grained performance model that captures both off-chip data access and on-chip data processing. It then applies the Non-dominated Sorting Genetic Algorithm II improved by Multi-Objective Bayesian Optimization to derive a Pareto-optimal partitioning and parallelization strategy that balances off-chip latency and PE utilization. Finally, GenCNN optimizes scheduling and routing to minimize data transfers. Experimental results show that GenCNN achieves up to 17.66× speedup in compilation and 6.47× in execution compared with state-of-the-art mapping frameworks.
Yudong Mu, Zhihua Fan, Xuejun An, Dongrui Fan, Xiaochun Ye
ACM Trans. Archit. Code Optim.2
2025 PANDA: Adaptive Prefetching and Decentralized Scheduling for Dataflow Architectures
abstract
Dataflow architectures are considered promising architecture, offering a commendable balance of performance, efficiency, and flexibility. Abundant prior works have been proposed to improve the performance of dataflow architectures. Nevertheless, these solutions can be further improved due to the lack of efficient data prefetching and flexible task scheduling. In this article, we propose a novel dataflow architecture with adaptive p refetching an d d ecentr a lized scheduling (PANDA). First, we present an application-adaptive data prefetching method and on-chip memory microarchitecture designed to overlap memory access latency. Second, we introduce a decentralized dataflow scheduling approach and processing element (PE) microarchitecture aimed at improving hardware utilization. Experimental results show that in a wide range of real-world applications, PANDA attains up to 2.53× performance improvement and 1.79× energy efficiency improvement over the state-of-the-art dataflow architectures.
Shantian Qin, Zhihua Fan, Zhen Wang 0045, Xuejun An, Xiaochun Ye, Dongrui Fan
ACM Trans. Archit. Code Optim.2
2025 A RISC-V Extended Infrastructure for CNNs Through Pipelined Computing and Data Dependence Optimization
abstract
With the rapid development of artificial intelligence (AI), convolutional neural networks (CNNs) have been widely applied in fields like computer vision and recommendation systems. This growth has intensified the demand for hardware acceleration of CNNs. Existing accelerators are either designed as co-processors or improve performance through extended instructions. While these methods can significantly improve performance, they often result in limited programming and execution flexibility. In this paper, we design custom RISC-V instructions specifically for CNNs to maximize data reuse and exploit parallelism. Then, to efficiently execute CNNs instructions, we extend a Pipelined Vector Computing Unit (PPVCU). Finally, we incorporate Pattern Detection Logic (PDL) to identify common data dependence patterns in CNNs, enabling the Data Dependence Computing Unit (DDCU) to process instructions within each pattern in parallel. Experimental results show that our approach achieves, on average, 9.54× performance improvement and 6.7× energy efficiency improvement compared to our baseline, 8.34× performance improvement and 3.1× energy efficiency improvement compared to state-of-the-art designs.
Teng Luo, Tengfei Xia, Zhihua Fan, Yudong Mu, Xuejun An, Xiaochun Ye, Dongrui Fan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 DFU-E: A Dataflow Architecture for Edge DSP and AI Applications
abstract
Edge computing aims to enable swift, real-time data processing, analysis, and storage close to the data source. However, edge computing platforms are often constrained by limited processing power and efficiency. This paper presents DFU-E, a dataflow-based accelerator specifically designed to meet the demands of edge digital signal processing (DSP) and artificial intelligence (AI) applications. Our design addresses real-world requirements with three main innovations. First, to accommodate the diverse algorithms utilized at the edge, we propose a multi-layer dataflow mechanism capable of exploiting task-level, instruction block-level, instruction-level, and data-level parallelism. Second, we develop an edge dataflow architecture that includes a customized processing element (PE) array, memory, and on-chip network microarchitecture optimized for the multi-layer dataflow mechanism. Third, we design an edge dataflow software stack that enables automatic optimizations through operator fusion, dataflow graph mapping, and task scheduling. We utilize representative real-world DSP and AI applications for evaluation. Comparing with Nvidia's state-of-the-art edge computing processor, DFU-E achieves up to 1.42× geometric mean performance improvement and 1.27× energy efficiency improvement.
Zhihua Fan, Tianyu Liu 0007, Zhen Wang 0045, Meng Wu 0006, Kunming Zhang, Yanhuan Liu, Ninghui Sun, Xiaochun Ye, Dongrui Fan
IEEE Trans. Parallel Distributed Syst.2
2024 LeakageFreeSpec: Applying the Wiping Approach to Defend Against Transient Execution Attacks
abstract
Transient Execution Attacks, such as Spectre and Meltdown, pose significant challenges to processor security. While various defense strategies have been proposed, they often suffer from issues like high-performance overhead and incomplete coverage. This paper introduces a novel hardware-based mitigation approach known as LeakageFreeSpec to address transient execution attacks.
Fahong Yu, Zhihua Fan
CF4
2024 OBSD: On-The-Fly Block-Wise Sparse Distillation Accelerating SpGEMMs in DNN Applications
abstract
The mainstream deep neural networks (DNNs) widely adopt pruning techniques to alleviate network overfitting and computational complexity. Simultaneously, the weight and activation data exhibit increasingly noticeable block-wise sparsity. Current DNN-specific accelerators mainly focus on element-wise while neglecting sparse features of block granularity. Meanwhile, one of the state-of-the-art work, HIRAC, combining matrix tiling with the fast packing algorithm, SorPack, to achieve effective acceleration of Sparse General Matrix Multiplications (SpGEMMs). However, the SorPack algorithm executes on the host CPU, and its runtime lies on the critical path of the overall execution. In this work, a specific SpGEMM accelerator named OBSD is proposed which achieves high performance and energy efficiency. An on-the-fly block-wise distillation approach is proposed for leveraging the block-wise sparsity in both weight and activation data, which is implemented during the data loading process, consuming minimal additional time. Moreover, a data-flow architecture is devised to improve efficiency of data exchange among data processing elements (DPs) with investigating the relationship for different partition sizes, block-wise sparsity, and block-wise distance. The evaluation results demonstrate that: (1) OBSD achieves average of 1.89× speedup as compared to HIRAC for representative matrices in DNN workloads. (2) An end-to-end evaluation on a DNN model shows a 2.41× speedup and energy efficiency improvement of 1.99× over the HIRAC. (3) OBSD possesses a power consumption 12.9W with 56.6mm2area @28nm TSMC.
Yanhuan Liu, Kunming Zhang, Zhihua Fan, Lexin Wang, Tianyu Liu 0007, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Xuejun An
HPCC4
2024 Improving Utilization of Dataflow Unit for Multi-Batch Processing
abstract
Dataflow architectures can achieve much better performance and higher efficiency than general-purpose core, approaching the performance of a specialized design while retaining programmability. However, advanced application scenarios place higher demands on the hardware in terms of cross-domain and multi-batch processing. In this article, we propose a unified scale-vector architecture that can work in multiple modes and adapt to diverse algorithms and requirements efficiently. First, a novel reconfigurable interconnection structure is proposed, which can organize execution units into different cluster typologies as a way to accommodate different data-level parallelism. Second, we decouple threads within each DFG node into consecutive pipeline stages and provide architectural support. By time-multiplexing during these stages, dataflow hardware can achieve much higher utilization and performance. In addition, the task-based program model can also exploit multi-level parallelism and deploy applications efficiently. Evaluated in a wide range of benchmarks, including digital signal processing algorithms, CNNs, and scientific computing algorithms, our design attains up to 11.95× energy efficiency (performance-per-watt) improvement over GPU (V100), and 2.01× energy efficiency improvement over state-of-the-art dataflow architectures.
Zhihua Fan, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Ninghui Sun, Xuejun An
ACM Trans. Archit. Code Optim.1
2023 Improving Utilization of Dataflow Architectures Through Software and Hardware Co-Design
Zhihua Fan, Shengzhong Tang, Xuejun An, Xiaochun Ye, Dongrui Fan
Euro-Par1
2023 DFGC: DFG-aware NoC Control based on Time Stamp Prediction for Dataflow Architecture
abstract
Coarse-grained reconfigurable architectures (CGRAs) have been regarded as promising accelerating paradigm for the ever-evolving algorithms in multi domains. Obtaining high energy-efficiency on CGRAs relies heavily on the combination of mapping, timing (issuing) and routing to decrease run-time idle. Statically configuration-driven designs are widely adopted, but the leak of hardware flexibility leads to a heavy reliance on burdensome scheduling of compiler to avoid over-serialization. This draw a trade-off between hardware/software co-design in CGRA scheduling.Unlike a static-schedule oriented approach, we propose DFGC (DFG-aware NoC Control), a dataflow-driven CGRA which takes the advantages of fully exploring data parallelism in different kernels by using dataflow dynamic firing with low overhead. The DFGC compiler is responsible for analyzing critical data paths and generating TimeStamp predictions instead of per cycle configuration. The rough predicted results enable router and PE to be sensitive to the entire dataflow graph, thereby accelerating the whole computation process. The DFG-aware NoC design realize a combined scheduling technique of hardware dynamic decision-making together with static prediction. DFGC represents a paradigm of CGRA that is worth exploring, achieving hardware/software co-design without relying on sophisticated designed compiler. Experiments show that DFGC achieves 1.32× energy efficiency improvement over a dataflow architecture and 1.8× energy efficiency improvement over a state-of-the-art static configured CGRA.
Tianyu Liu 0007, Zhihua Fan
ICCD3
2023 Alleviating Transfer Latency in DataFlow Accelerator for DSP Applications
abstract
Towards multiple domains, dataflow accelerators show superiority for their flexible programmability and high efficiency. This efficiency relies highly on data communication between processing elements (PEs), which is sensitive to PE location, array scale and workload size. Laying out instructions as a dataflow graph on the PE array creates more instruction-level parallelism. However, the farther distance between remote PEs and memory banks introduces extra transfer latency, bringing performance degradation to high real-time applications. This paper examines the workloads of digital signal processing across different data scales and classifies latency problems related to data transfers and kernel switching. Specifically, we propose a novel forwarding network on chip to alleviate transfer latency and improve multi-destination sharing in the dataflow execution. Moreover, we devise bandwidth reusing mechanism to speedup kernel switching. The experiment results show that our scalable design achieves up to 2.19× (1.45× on average) speedup while reducing switching overhead by 9.85×, with an area overhead of 10.82% over the conventional dataflow accelerator.
Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Junying Huang, Shengzhong Tang, Yanhuan Liu, Kunming Zhang, Xiaochun Ye, Dongrui Fan
ICCD3
2023 Accelerating Convolutional Neural Networks by Exploiting the Sparsity of Output Activation
abstract
Deep Convolutional Neural Networks (CNNs) are the most widely used family of machine learning methods that have had a transformative effect on a wide range of applications. Previous studies have made great breakthroughs in accelerating CNNs, but they only target on the input sparsity of activation and weight, thus do not eliminate the unnecessary computations due to the fact that more zeros in the output results are not directly caused by the zero-valued positions of the input data. In this paper, we take advantage of the output activation sparsity to reduce the execution time and energy consumption of CNNs. First, we propose an effective prediction method that leverages the output activation sparsity. Our method first predicts the output activation polarity of convolutional layers based on the singular value decomposition (SVD) approach. Then, it uses the predicted negative value to skip invalid computations. Second, an effective accelerator is designed to take advantage of sparsity to achieve CNN inference acceleration. Each PE is equipped with a prediction unit and a non-zero value detection unit to remove invalid computation blocks. And an instruction bypass technique is proposed which further exploits the sparsity of the weights. The efficient dataflow graph mapping approach and pipeline execution ensure high computational resource utilization. Experiments show that our approach achieves up to 1.63× speedup and 55.30% energy reduction compared with dense networks with a slight loss of accuracy. Compared with Eyeriss, our accelerator achieves on average 1.31 × performance improvement and 54% energy reduction. Our accelerator also achieves a similar performance to SnaPEA, but with a better energy efficiency.
Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Yanhuan Liu, Meng Wu 0006, Xinxin Wu, Xiaochun Ye, Dongrui Fan, Ninghui Sun, Xuejun An
IEEE Trans. Parallel Distributed Syst.1
2022 LRP: Predictive output activation based on SVD approach for CNN s acceleration
abstract
Convolutional Neural Networks (CNNs) achieve state-of-the-art performance in a wide range of applications. CNNs contain millions of parameters, and a large number of computations challenge hardware design. In this paper, we take advantage of the output activation sparsity of CNNs to reduce the execution time and energy consumption of the network. We propose Low Rank Prediction (LRP), an effective prediction method that leverages the output activation sparsity. LRP first predicts the output activation polarity of the convolutional layer based on the singular value decomposition (SVD) approach of the convolution kernel. And then it uses the predicted negative value to skip invalid computation in the original convolution. In addition, an effective accelerator, LRPPU, is proposed to take advantage of sparsity to achieve network inference acceleration. Experiments show that our LRPPU achieves 1.48 x speedup and 2.02 x energy reduction compared with dense networks with slight loss of accuracy. Also, it achieves on average 2.57 x speedup over Eyeriss and has similar performance and less accuracy loss compared with SnaPFA.
Xinxin Wu, Zhihua Fan, Tianyu Liu 0007, Xiaochun Ye, Dongrui Fan
DATE2
2022 A Routing-Aware Mapping Method for Dataflow Architectures
Zhihua Fan, Tianyu Liu 0007, Xuejun An, Xiaochun Ye, Dongrui Fan
NPC1
2011 Metadata Distribution and Consistency Techniques for Large-Scale Cluster File Systems
abstract
Most supercomputers nowadays are based on large clusters, which call for sophisticated, scalable, and decentralized metadata processing techniques. From the perspective of maximizing metadata throughput, an ideal metadata distribution policy should automatically balance the namespace locality and even distribution without manual intervention. None of existing metadata distribution schemes is designed to make such a balance. We propose a novel metadata distribution policy, Dynamic Dir-Grain (DDG), which seeks to balance the requirements of keeping namespace locality and even distribution of the load by dynamic partitioning of the namespace into size-adjustable hierarchical units. Extensive simulation and measurement results show that DDG policies with a proper granularity significantly outperform traditional techniques such as the Random policy and the Subtree policy by 40 percent to 62 times. In addition, from the perspective of file system reliability, metadata consistency is an equally important issue. However, it is complicated by dynamic metadata distribution. Metadata consistency of cross-metadata server operations cannot be solved by traditional metadata journaling on each server. While traditional two-phase commit (2PC) algorithm can be used, it is too costly for distributed file systems. We proposed a consistent metadata processing protocol, S2PC-MP, which combines the two-phase commit algorithm with metadata processing to reduce overheads. Our measurement results show that S2PC-MP not only ensures fast recovery, but also greatly reduces fail-free execution overheads.
Jin Xiong, Yiming Hu, Guojie Li, Rongfeng Tang, Zhihua Fan
IEEE Trans. Parallel Distributed Syst.5
2008 Least visible path analysis in raster terrain
abstract
Least visible path analysis is a basic function in terrain visibility analysis. However, current least visible path planning is constrained to least‐cost path computing on a cost surface obtained from visibility information of all the terrain points on digital elevation models. This kind of method ignores the visibility correlation caused by the overlapped part of the adjacent points' reverse viewsheds. With regard to such a correlation, this paper proposes a new method using the amalgamation of reverse viewsheds to find an optimal path with the minimal visibility dominance. Both methods are implemented using C++ programming, and the experimental results show that the new method produces more accurate least visible paths than the traditional one does.
Jinfang Zhang, Zhihua Fan
Int. J. Geogr. Inf. Sci.4
2007 Max/Min Path Visual Coverage Problems in Raster Terrain
abstract
Path visual coverage is the region which could be seen from the path in the terrain. This paper advances the concept of "average horizon " which represents the wide extent of the visual field of each path point to analyze and model the max and min path visual coverage problems, and resolves them utilizing simulated annealing algorithm based on the operation of viewshed amalgamation with pre-computed and stored viewshed information.
Jinfang Zhang, Zhihua Fan
CAD/Graphics4