Chenji Han

dblp:360/6071 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
0009-0007-1247-9644ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 ACRS: Adjacent Computation Resource Sharing among Partitioned GPU Sub-Cores
abstract
Modern GPUs typically segment Streaming Multiprocessors (SMs) into sub-cores (e.g. 4 sub-cores) to reduce power consumption and chip area. However, this partitioned design prevents potential task distributions across sub-cores, impairing overall execution efficiency. In this paper, we explore the performance benefit of sharing hardware resources among sub-cores and identify functional units (FUs) as critical components for compute-intensive applications. Moreover, our observations reveal that instructions residing in operand collectors can be obstructed by back-end FUs, but there is a high probability that unoccupied FUs are available in adjacent sub-cores during such blockages. In response, we introduce the adjacent computation resource sharing (ACRS) framework to efficiently utilize these unoccupied units among sub-cores. ACRS has two key modules: Shared FU Issue (SF_ISSUE) and Shared FU Write Back (SF_WriteBack). SF_ISSUE monitors the status of operand collectors and functional units, and offloads instructions from blocked sub-cores to unoccupied resources. Meanwhile, SF_WriteBack routes results back to the original sub-core.To minimize wiring overhead, each sub-core is assigned a fixed target core for sharing. We design a series of matching policies and finally filter out the most effective sequential method. Evaluation results show that ACRS improves performance by up to $46.4 \%$, with an average of $14.1 \%$ over the traditional partitioned architecture, while reducing energy consumption by $8.3 \%$. Besides, ACRS achieves an additional 12.3% performance improvement compared with the SOTA method.
Penghao Song, Chongxi Wang, Chenji Han
DAC3
2025 MeMo: Enhancing Representative Sampling via Mechanistic Micro-Model Signatures
abstract
Representative Sampling, exemplified by SimPoint is widely utilized in pre-silicon performance evaluation. It relies on the code signature to characterize programs and select representative simulation points to estimate the performance. Several code signatures have been proposed, including BBV, MIC, BBV-LDV, and PMC. Although these code signatures can achieve low average performance estimation errors for benchmark suites, they still yield high errors for specific programs within those suites, due to their inherent limitations of incomplete characteristic representations and indirect performance correlations. In this work, we propose the utilization of mechanistic micromodel signatures, MeMo, to enhance representative sampling. MeMo contains various signatures sourced from a series of micromodels under different configurations. These models encompass fetch, issue, and cache models, each concentrating on certain$\mu$Arch structure constraints while idealizing the others. MeMo enables the comprehensive depiction of different program characteristics and their performance responses while not confined to specific$\mu$Arch. We thoroughly evaluated MeMo against current code signatures on SPEC CPU2017. The experiment demonstrated that compared to the most commonly utilized BBV, MeMo can substantially reduce the average CPI estimation error from 3.96% to 1.63% and significantly suppress the maximum estimation error from 19.94% to 6.49%.
Chenji Han, Huai Xu, Guangyao Guo, Fuxin Zhang
ISPASS1
2025 SnsBooster: Enhancing Sampling-based μArch Evaluation Efficiency through Online Performance Sensitivity Analysis
abstract
Sampling-based methods, such as SimPoint, are widely used for efficient pre-silicon μ Arch evaluations, where the costs are the number of simulation points multiplied by the number of evaluated μ Arch designs. However, these costs keep growing with an increasing number of simulation points and expanding μ Arch design space. Although techniques have been developed to accelerate the μ Arch design space exploration, less attention has been given to further reducing the simulation budget of each μ Arch evaluation. Common strategies like reducing simulation coverage or sampling fewer simulation points typically compromise estimation accuracy. Therefore, further reducing the simulation budget without compromising estimation accuracy remains a critical research problem. In this work, we propose SnsBooster to enhance sampling-based μ Arch evaluation efficiency, based on two insights: (a) large portions of simulation points’ performance changes are typically insensitive to the evaluated μ Arch changes, and (b) simulation points’ performance sensitivities under specific μ Arch change correlate with their inherent characteristics. By online building a μ Arch-specific performance sensitivity classifier via progressive simulation and continuous validation, SnsBooster can identify and selectively evaluate only performance-sensitive points, thus reducing the simulation budget without compromising estimation accuracy. When applied across various μ Arch changes, SnsBooster achieves an average simulation budget reduction of 39.04% with an accuracy loss of only 0.14%, compared to simulating all the sampled points. Under the same accuracy loss, SnsBooster’s simulation budgets are only 64.73% and 65.60% of those required by methods of reducing simulation coverage or sampling fewer points. Besides, under identical simulation budgets, the average accuracy losses of these methods are 1.41% and 1.23%, which is substantially higher than that of SnsBooster.
Chenji Han, Zifei Zhang 0001, Xinyu Li 0010, Qi Guo 0001, Fuxin Zhang
ACM Trans. Archit. Code Optim.1
2025 Tiaozhuan: A General and Efficient Indirect Branch Optimization for Binary Translation
abstract
Binary translation enables transparent execution, analysis, and modification of the binary program, serving as a core technology that facilitates instruction set emulation, cross-platform compatibility of software, and program instrumentation. Handling indirect branch instructions is widely recognized as a significant performance bottleneck in binary translation. While the target of a direct branch can be determined during the translation phase, an indirect branch requires a runtime lookup from the guest program counter to the host program counter, significantly influencing the performance of translator. Although several methods have been proposed to accelerate this process, each guest indirect branch instruction still translates into approximately 10 host instructions, resulting in considerable overhead. This article introduces Tiaozhuan, which addresses this issue by employing two optimization schemes. First, full address mapping uses a larger address space to store address mappings from guest to host, effectively reducing the number of instructions required to lookup the target of an indirect branch. Second, exceptionassisted branch elimination further eliminates branch instructions that check target correctness of targets in the lookup process. These two approaches enable indirect branches target lookup to be completed within one to two instructions, noticeably decreasing the overhead of indirect branches. Compared to state-of-the-art mechanisms, the SPEC CPU2006 benchmark suite showed a reduction in the number of instructions by an average of 4.2%, with the highest observed performance improvement reaching 19.4% and an average increase of 3.9%.
Xinyu Li 0010, Guangyao Guo, Yanzhi Lan, Chenji Han, Gen Niu, Fuxin Zhang
ACM Trans. Archit. Code Optim.5
2025 Augur: Semantics-Aware Temporal Prefetching for Linked Data Structure
abstract
Linked data structures (LDS), such as lists and trees, are widely used in modern applications. Traversing LDS typically involves a significant amount of pointer chasing. Due to the serial nature of memory access in pointer chasing, the incurred long memory latency of traversing LDS has become a critical performance bottleneck. Furthermore, the poor spatial locality in LDS makes it difficult for spatial prefetchers to predict access addresses. Although temporal prefetchers can handle irregular memory access patterns, hindered by the challenges of collecting semantic information, current state-of-the-art temporal prefetchers suffer from significant metadata redundancy and frequent metadata conflicts. Consequently, there remain substantial opportunities to enhance the LDS prefetching. To solve this problem, we propose Augur, a semantics-aware temporal prefetcher to enhance LDS performance. Augur utilizes a novel pruning method to obtain semantic information and effectively extracts node address correlations from the perspective of nodes in LDS, thereby diminishing the metadata redundancy and conflicts. Additionally, Augur employs efficient metadata management strategies that guarantee a minimal storage overhead. Evaluated on LDS workloads, Augur achieves an average performance speedup of 17.8% and 11.7% over the baseline stride prefetcher and state-of-the-art spatial prefetcher Berti, respectively. Furthermore, Augur outperforms the state-of-the-art temporal prefetcher MISB, Triage, and Triangel, by 17.4%, 12.8%, and 6.3%, respectively, with a significantly lower storage overhead of only 1.26 KB.
Junliang Wu, Chenji Han, Xinyu Li 0010, Fuxin Zhang
ACM Trans. Archit. Code Optim.3
2025 ETBench: Characterizing Hybrid Vision Transformer Workloads Across Edge Devices
abstract
Lightweight Convolution and Vision Transformer hybrid models have increasingly dominated the frontiers of deep learning (DL) on edge devices; however, to the best of our knowledge, no prior work has provided comprehensive evaluation on hybrid models’ performance and analyzed their characteristics by diving deep into the edge ecosystem with diversified modern DL inference engines and heterogeneous hardware. This paper proposes a comprehensive open-source benchmark suite,ETBench, to allow power-efficiency, performance and accuracy assessment for state-of-the-art (SOTA) hybrid models across 11 most widely-used DL engines deployed on diverse edge devices. After building ETBench that satisfies 6 design requirements proposed in our work, we conduct extensive experiments on 14 devices including 19 CPUs, 11 GPUs and 5 NPUs, and obtain benchmark results from all deployment scenarios (combinations of models, quantization formats, software engines, and hardware platforms). Valuable observations and insightful implications are finally summarized. For example, within current DL engines, the INT8 quantization is significantly underperformed in terms of accuracy and speed against FP16 for hybrid models. Overall, ETBench serves as a collaborative platform that assists model architects in better evaluating their models and makes it possible for future co-optimizations of DL engines and hardware accelerators.
Yingkun Zhou, Zhengshuyuan Tian, Jinpeng Ye, Chenji Han, Fuxin Zhang
IEEE Trans. Computers6
2024 A dependence graph pattern mining method for processor performance analysis
Chenji Han, Fuxin Zhang
Perform. Evaluation2
2024 Tyche: An Efficient and General Prefetcher for Indirect Memory Accesses
abstract
Indirect memory accesses (IMAs, i.e., A [ f ( B [ i ])]) are typical memory access patterns in applications such as graph analysis, machine learning, and database. IMAs are composed of producer-consumer pairs, where the consumers’ memory addresses are derived from the producers’ memory data. Due to the built-in value-dependent feature, IMAs exhibit poor locality, making prefetching ineffective. Hindered by the challenges of recording the potentially complex graphs of instruction dependencies among IMA producers and consumers, current state-of-the-art hardware prefetchers either (a) exhibit inadequate IMA identification abilities or (b) rely on the run-ahead mechanism to prefetch IMAs intermittently and insufficiently. To solve this problem, we propose Tyche, 1 an efficient and general hardware prefetcher to enhance IMA performance. Tyche adopts a bilateral propagation mechanism to precisely excavate the instruction dependencies in simple chains with moderate length (rather than complex graphs). Based on the exact instruction dependencies, Tyche can accurately identify various IMA patterns, including nonlinear ones, and generate accurate prefetching requests continuously. Evaluated on broad benchmarks, Tyche achieves an average performance speedup of 16.2% over the state-of-the-art spatial prefetcher Berti. More importantly, Tyche outperforms the state-of-the-art IMA prefetchers IMP, Gretch, and Vector Runahead, by 15.9%, 12.8%, and 10.7%, respectively, with a lower storage overhead of only 0.57 KB.
Chenji Han, Xinyu Li 0010, Junliang Wu, Yifan Hao 0001, Zidong Du, Qi Guo 0001, Fuxin Zhang
ACM Trans. Archit. Code Optim.2
2023 SCFM: A Statistical Coarse-to-Fine Method to Select Cross-Microarchitecture Reliable Simulation Points
Chenji Han, Hongze Tan 0001, Xinyu Li 0010, Ruiyang Wu 0001, Fuxin Zhang
APPT1
2023 CCS: A Motif-Based Storage Format for Micro-execution Dependence Graph
Chenji Han
APPT2