Honglan Zhan

dblp:265/0821 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0002-2840-9639ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2024 Multi: Reduce Energy Overhead of Criticality-Aware Dynamic Instruction Scheduling for Energy Efficiency
abstract
Criticality-aware dynamic instruction scheduling (DIS) focuses on prioritizing the execution of critical instructions, thereby significantly improving performance. However, criticality-aware DIS is energy-intensive. Energy overhead mainly comes from the design of DIS strategy and the multi-port tables used for extracting critical instruction slices. In this study, Multi is proposed, which develops an energy-efficient DIS strategy and an adaptive multi-port table coordinated with the DIS strategy to reduce the energy overhead of criticality-aware DIS. Evaluation on the SPEC CPU2017 benchmark shows that Multi outperforms the state-of-the-art work in terms of core energy consumption with a geometric mean reduction of 5.9%. Moreover, post-layout simulation results demonstrate a 61.2% reduction in energy consumption of the proposed adaptive multi-port table. It also proves the applicability of Multi in ultra-wide processors with unified or distributed schedulers, contributing to enhanced energy efficiency.
Honglan Zhan, Xianhua Liu 0001, Xu Cheng 0001
ICCD1
2023 MBAPIS: Multi-Level Behavior Analysis Guided Program Interval Selection for Microarchitecture Studies
abstract
Understanding program behavior is crucial in computer architecture research, but the growing size of benchmarks makes analyzing and simulating entire programs increasingly challenging. In practice, researchers often select representative program intervals for analysis and testing. These intervals are different sections of continuous execution of a program. SimPoint is a well-known method for selecting representative intervals using hardware-independent information. However, when focusing on a specific microarchitecture study, it is desirable to select intervals that are more relevant to that study. For instance, intervals with more branch mispredictions are more appropriate for branch prediction studies. We refer to these intervals as “tailored intervals” for branch prediction studies. This paper presents a Multi-level Behavior Analysis guided Program Interval Selection (MBAPIS) for selecting tailored intervals. For a given microarchitecture study, the first level of MBAPIS uses hardware performance counters to prioritize selecting the intervals that exhibit clearer microarchitectural characteristics relevant to that study. The second level analyzes the processor performance bottlenecks to further select the intervals where the concerned microarchitecture design more strongly impacts performance. Finally, MBAPIS performs clustering analysis with the basic block information of each interval selected by the first two levels, and selects the representative intervals among them while preserving the diverse software behavior. Additionally, we present a general and extensible interval-replaying design to accurately re-execute selected intervals. The SPEC CPU2006 and CPU2017 benchmarks are used for evaluation. The results demonstrate that MBAPIS can select representative and tailored intervals for two typical microarchi-tecture studies and deliver accurate estimates of the concerned hardware events for all tailored intervals in each benchmark, with an average error rate of less than 1.5%. Moreover, the interval-replaying design effectively restores the hardware behavior of the intervals selected by MBAPIS, with an average relative error rate of annroximately 1.4%.
Hongwei Cui, Honglan Zhan, Shuhao Liang, Xianhua Liu 0001, Xu Cheng 0001
PACT3
2023 A Hardware-Software Cooperative Interval-Replaying for FPGA-based Architecture Evaluation
abstract
Open-source processors and FPGA provide more real and accurate results of the new microarchitecture design, but the long execution time for running large benchmarks on FPGA boards still hinders researchers. This paper proposes a hardware-software cooperative interval-replaying. It uses simula-tors to create checkpoints for arbitrary program intervals and provides an extensible and portable checkpoint loader to re-execute selected intervals. In addition, this paper extends RISC-VISA and proposes an event-based sampling design to find hot program intervals with more representative microarchitecture characteristics. By using checkpoints in hot regions, researchers can quickly verify the effectiveness of microarchitecture designs on FPGA and alleviate the speed bottleneck of FPGA. The correctness and effectiveness of the checkpoint scheme and the event-based sampling design are evaluated on FPGA. The experimental results show that the solution is effective.
Hongwei Cui, Shuhao Liang, Honglan Zhan, Xianhua Liu 0001, Xu Cheng 0001
DATE5
2023 High-Speed and Energy-Efficient Single-Port Content Addressable Memory to Achieve Dual-Port Operation
abstract
High-speed and energy-efficient multi-port content addressable memory (CAM) is very important to modern superscalar processors. In order to overcome the disadvantages of multi-port CAM and improve the performance of searching stage, a high-speed and energy-efficient single-port (SP) CAM is introduced to achieve dual-port (DP) operation. For different bit cell topologies - the traditional 9T CAM cell and 6T SRAM cell, two novel peripheral schemes - CShare and VClamp are proposed. The proposed schemes are verified using all possible corners, a wide range of temperature and detailed Monte-Carlo variation analysis. With 65-nm process and 1.2 V supply, the search delay of CShare and VClamp is 0.55 ns and 0.6 ns, respectively, a reduction of approximately 87% compared to the state-of-the-art works. In addition, compared with the recently proposed IOT BCAM, CShare and VClamp can provide 84.9% and 85.1% energy reduction in the TT corner, respectively. Experimental results in an 8 Kb CAM at 1.2 V supply and across different corners show that the energy efficiency is improved by 45.56% (CShare) and 45.64% (VClamp) on average in comparison with DP CAM.
Honglan Zhan, Hongwei Cui, Xianhua Liu 0001, Xu Cheng 0001
DATE1
2020 In-Memory Computing With Double Word Lines and Three Read Ports for Four Operands
abstract
The von Neumann architecture is approaching its limits in terms of scalability and power consumption. In-memory computation is a possible approach to mitigate this limitation. This brief proposes a configurable 8T static random access memory (SRAM) cell with double word lines and three read ports for in-memory computing. In addition to the normal SRAM function, XOR/XNOR and compound Boolean logic operations of three or four operands, such as AND-OR, AND-OR-INVERT, OR-AND, and OR-AND-INVERT, can be performed in one cycle by fully utilizing the three read ports to obtain 13.2-fJ/bit consumption at 0.6 V. The logic operation frequency is 793 MHz at 1.2 V. The proposed SRAM effectively resolves the bottleneck of the existing in-memory computation schemes that only support compound Boolean logic operations with more than two cycles. In addition, the proposed SRAM array scheme can be configured and used as a binary content-addressable memory or a ternary content-addressable memory for searching operations; it achieves 0.24 fJ/search/bit at 0.6 V in the worst case. At 1.2 V, the searching frequency is up to 813 MHz when searching 128 bits with 65-nm technology.
Zhi-Ting Lin, Honglan Zhan, Chunyu Peng, Wenjuan Lu, Xiulong Wu, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.2