Anping He

dblp:02/6123 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-1913-2057ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Hyper-Parallel Superscalar Asynchronous RISC-V Processor Based on Event-Driven Logic
abstract
Event-driven neuromorphic computing involves sparse and asynchronous signal activity, which leads to irregular computation patterns and fine-grained concurrency. As a result, processing architectures need to support both high parallelism and energy efficiency. Among existing architectural solutions, superscalar designs exhibit significant potential for addressing high parallelism demands. However, conventional superscalar processors, which rely on synchronous circuits, maintain high-frequency clocking at all times, leading to substantial power inefficiency in sparse computation scenarios. To address this issue, we propose an asynchronous superscalar architecture that replaces global clocking with fully local handshake-based control, implemented using a bundled-data asynchronous protocol. The design supports decoding of up to 64 scalar instructions per cycle and implements the RISC-V RV32IMC instruction set. A prototype was fabricated using a 110 nm complementary metal oxide semiconductor (CMOS) process and was evaluated through post-layout simulation. Operating at 1.2 V, the processor delivers a peak INT8 throughput of 669.4 GOPS, with a static power consumption of 421 mW.
Kangli Zhao, Anping He, Lixian Zhu, Qunxi Dong, Fuze Tian, Qingguo Zhou, Qinglin Zhao
IEEE Trans. Comput. Soc. Syst.2
2025 Asynchronous Architecture Design and Implementation of Physical Memory Protection for RISC-V
abstract
Trusted Execution Environment(TEE) delineates se- cure zones within the system, ensuring the security of the execution environment and data confidentiality. However, due to the relatively late start of the RISC-V instruction set ar- chitecture, there are currently few designs for RISC-V-based TEEs, which face challenges such as performance bottlenecks, insufficient support for privilege modes, and the lack of a unified model. Moreover, synchronous circuits’ reliance on extensive clock networks increases area and power usage. This paper introduces an asynchronous Physical Memory Protection(PMP) design for RISC-V, addressing these issues. This scheme provides key hardware support for implementing TEE for RISC-V. The architecture supports user, supervisor, and machine modes, offering 4 B to 16 GB granular protection and up to 64 PMP entries. The initial address and the number of PMP entries can be configured according to requirements, offering high flexibility. After testing, the overall and sub-module functions are correct and align with expectations. Finally, in the 110-nm process, the area of this architecture is 0.47 mm2and the power consumption is only 3.6829 mW.
Anping He, Yunpeng Xing, Jingye Zhong, Yinglong Li, Jun Ma 0037
ISCAS1
2024 Research on High-Efficiency Asynchronous Superscalar Processors: (PhD Forum Paper)
abstract
With the rapid popularization of mobile devices and smart hardware, copuled with the growing trend towards intelligence, the application of processor has been extended to nearly all fields of information technology. These fields now demand more stringent requirements in terms of processor performance and power. Traditional superscalar processors, which utilize a global clock control design, face limitations in enhancing pipeline depth and superscalar width as circuit sizes and transistor densities increase. This constraint impedes the potential performance enhancements of the processors. At the same time the clock circuit itself is not involved in data operations, it contributes to significant dynamic power consumption issues. Therefore, this paper proposes a novel asynchronous superscalar fine-grained processing method that utilizes local fine-grained, multilevel asynchronous micropipelines for overall control. This enhances both the superscalar width and pipeline depth. Based on this method, an asynchronous superscalar processor architecture is implemented, featuring an on-demand operational mechanism where start-stop transitions do not require reinitialization, enabling improved performance while reducing dynamic power.
Kangli Zhao, Anping He
ASAP2
2024 HSAMM: A Hybrid-Strassen Algorithm-Based Asynchronous Architecture for Sparse Matrix Multiplication
abstract
General sparse matrix-matrix multiplication (SpGEMM) is a fundamental computational method with wide-ranging applications in scientific simulations, machine learning, and image processing. However, when tackling large-scale SpGEMM, single-core processors fall short in managing the computation-intensive tasks, while multi-core and many-core processors encounter challenges such as complex scheduling, high communication overhead, and substantial energy consumption. Existing platforms for SpGEMM that base on synchronous circuit design require complex state machines to implement parsing mechanisms, and they suffer from power consumption issues associated with clock trees. Therefore, the exploration of asynchronous sparse matrix accelerators is essential to address these issues. This paper presents an asynchronous sparse matrix multiplication accelerator architecture designed for 128×128 sparse matrices, called HSAMM, which is based on the Hybrid-Strassen algorithm and employs the BCSR and BCSC formats for the storage of input and output matrices. HSAMM is validated through prototype implementation and performance evaluation on the FACE-VUP platform. Simulation results demonstrate significant performance enhancements in SpGEMM (sparsity ≤ 0.8%) for int32 data type, achieving average accelerations of 3.2×, 8.2× and 83.3× over the Intel MKL, the Eigen library and the matlab, respectively.
Lingzhuang Zhang, Rongqing Hu, Yilong Jiang, Jun Ma 0037, Yinglong Li, Anping He
HPCC7
2019 How to Accelerate FPGA Application in an Asynchronous Way?
abstract
FPGA with massive customizable parallel computation capacity, is potentially good for fast time-to-market applications. However, its complex placing and routing lead to a relatively large latency and low frequency. Besides, the clock problems might make a design hard and slow, especially the one with complex control or variant computations. All of those defects seem to be from the essence of the synchronous design methodology and there does not exist an easy way to solve by clocks. Although the FPGA vendors do not supply an asynchronous design routine, flow or tool, it is still possible to implement a clockless design with a concrete FPGA chip and then harness lots of benefits that synchronous one misses, which is shown in this paper. We adopt link-joint as the asynchronous communication mechanism that discards clock limitation, but equips high throughput due to the fast handshake among neighbor clicks. The simplest link-joint circuit is click that conforms to Bundled Bound Data (BBD) protocol for local communication. Multiple clicks can be constructed and trimmed to types of micro-pipeline structures, feasibly and flexibly. With above considerations, we propose an innovative asynchronous design method for Xilinx FPGA applications, as well as the asynchronous control framework by dedicated micro-pipeline structures. Furthermore, we introduce delay maching technologies as well as whole design flow and tool-chain. All of these supply an applicable way of accelerating an asynchronous design for a FPGA. The case-studies show that communication between neighbor clicks is less than 1.1ns and the asynchronous method accelerates FPGA latency extremely.
Anping He, Jinlin Zhang, Lvying Yu
FPGA1
2018 Click-Based Asynchronous Mesh Network with Bounded Bundled Data
abstract
We have implemented an asynchronous mesh network. This paper describes our innovative design using a Click controller. Compared to designs that use other asynchronous circuit families with C-elements and four-phase bundled data, our two-phase Click-based Bounded Bundled Data design is faster, but introduces phase skews when handling concurrent traffic at a single node. Instead of eliminating the phase skews, we use them as computation slots. Our network uses a novel asynchronous arbiter with a queue that can accept data from both the four cardinal directions as well as from a local source, five directions in all. We have implemented our network design in 1 × 1, 2 × 2 and 4 × 4 sizes, larger network could be implemented easier since the isomorphism and modularity of the routing nodes. Our experiments show that an initial data item passes through a node in 157ns v.s. 81ns for non-delay-branch and delay-branch designs separately. Following items take about 65% as long. But for a network, the average latency of a node keeps almost same for different paths. We believe that with the non-delay-branch designs, our asynchronous mesh network could offer 10.1M routes per second for a 1 × 1 network and 5.33M routes per second for 2 × 2 or 5.06M for 4 × 4 networks, and work at the rate of 17.3M, 10.1M and 11.7M with the enhanced delay-branch way. For both cases, its latency is approximately linear with scale.
Anping He, Guangbo Feng, Yong Hei, Hong Chen 0002
ICPP1
2017 A Hybrid Model Equipped with the Minimum Cycle Decomposition Concept for Short-Term Forecasting of Electrical Load Time Series
abstract
Electricity load forecasting is an essential, however complicated work. Due to the influence of a large number of uncertain factors, it shows complicated nonlinear combination features. Therefore, it is difficult to improve the prediction accuracy and the tremendous breadth of applicability especially for using a single method. In order to improve the performance including accuracy and applicability of electricity load forecasting, in this paper, a concept named minimum cycle decomposition (MCD) that the raw data are grouped according to the minimum cycle was proposed for the first time. In addition, a hybrid prediction model (HMM) based on one-order difference, ensemble empirical model decomposition (EEMD), mind evolutionary computation (MEC) and wavelet neural network (WNN) was also proposed in this study. The HMM model consists of two parts. Part one, pre-processing, known as one order difference to remove the trend of subsequence and EEMD to reduce the noise, was performed by HMM model on each subset. Part two, the WNN optimized by MEC (WNN $$+$$ MEC) was applied on resultant subseries. Finally, a number of different models were used as the comparative experiment to validate the effectiveness of the presented method, such as back propagation neural network (BP-1), BPNN combined MCD (BP-2), WNN combined MCD (WNNM), a HMM (DEEPLSSVM) based on one-order difference, EEMD, particle swarm optimization and least squares support vector machine and a hybrid model (DEESGRNN) based on one-order difference, EEMD, simulate anneal and generalized regression neural network. Certain evaluation measurements are taken into account to assess the performance. Experiments were carried out on QLD (Queensland) and NSW (New South Wales) electricity markets historical data, and the experimental results show that the MCD has the advantages of improving model accuracy and of generalization ability. In addition, the simulation results also suggested that the proposed hybrid model has better performance.
Zhaoshuang He, Anping He
Neural Process. Lett.4
2016 Modular Timing Constraints for Delay-Insensitive Systems
Hoon Park, Anping He, Marly Roncken, Ivan E. Sutherland
J. Comput. Sci. Technol.2
2015 A discrete invasive weed optimization algorithm for solving traveling salesman problem
Yongquan Zhou, Qifang Luo, Anping He
Neurocomputing4