VLDB 2026 Research / reviewers in the wild / expert
Xinyu Li 0010
dblp:88/2359-10
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-2640-8173ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SnsBooster: Enhancing Sampling-based μArch Evaluation Efficiency through Online Performance Sensitivity AnalysisabstractSampling-based methods, such as SimPoint, are widely used for efficient pre-silicon μ Arch evaluations, where the costs are the number of simulation points multiplied by the number of evaluated μ Arch designs. However, these costs keep growing with an increasing number of simulation points and expanding μ Arch design space. Although techniques have been developed to accelerate the μ Arch design space exploration, less attention has been given to further reducing the simulation budget of each μ Arch evaluation. Common strategies like reducing simulation coverage or sampling fewer simulation points typically compromise estimation accuracy. Therefore, further reducing the simulation budget without compromising estimation accuracy remains a critical research problem. In this work, we propose SnsBooster to enhance sampling-based μ Arch evaluation efficiency, based on two insights: (a) large portions of simulation points’ performance changes are typically insensitive to the evaluated μ Arch changes, and (b) simulation points’ performance sensitivities under specific μ Arch change correlate with their inherent characteristics. By online building a μ Arch-specific performance sensitivity classifier via progressive simulation and continuous validation, SnsBooster can identify and selectively evaluate only performance-sensitive points, thus reducing the simulation budget without compromising estimation accuracy. When applied across various μ Arch changes, SnsBooster achieves an average simulation budget reduction of 39.04% with an accuracy loss of only 0.14%, compared to simulating all the sampled points. Under the same accuracy loss, SnsBooster’s simulation budgets are only 64.73% and 65.60% of those required by methods of reducing simulation coverage or sampling fewer points. Besides, under identical simulation budgets, the average accuracy losses of these methods are 1.41% and 1.23%, which is substantially higher than that of SnsBooster. Chenji Han, Zifei Zhang 0001, Xinyu Li 0010, Qi Guo 0001, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | Tiaozhuan: A General and Efficient Indirect Branch Optimization for Binary TranslationabstractBinary translation enables transparent execution, analysis, and modification of the binary program, serving as a core technology that facilitates instruction set emulation, cross-platform compatibility of software, and program instrumentation. Handling indirect branch instructions is widely recognized as a significant performance bottleneck in binary translation. While the target of a direct branch can be determined during the translation phase, an indirect branch requires a runtime lookup from the guest program counter to the host program counter, significantly influencing the performance of translator. Although several methods have been proposed to accelerate this process, each guest indirect branch instruction still translates into approximately 10 host instructions, resulting in considerable overhead. This article introduces Tiaozhuan, which addresses this issue by employing two optimization schemes. First, full address mapping uses a larger address space to store address mappings from guest to host, effectively reducing the number of instructions required to lookup the target of an indirect branch. Second, exceptionassisted branch elimination further eliminates branch instructions that check target correctness of targets in the lookup process. These two approaches enable indirect branches target lookup to be completed within one to two instructions, noticeably decreasing the overhead of indirect branches. Compared to state-of-the-art mechanisms, the SPEC CPU2006 benchmark suite showed a reduction in the number of instructions by an average of 4.2%, with the highest observed performance improvement reaching 19.4% and an average increase of 3.9%. Xinyu Li 0010, Guangyao Guo, Yanzhi Lan, Chenji Han, Gen Niu, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | Augur: Semantics-Aware Temporal Prefetching for Linked Data StructureabstractLinked data structures (LDS), such as lists and trees, are widely used in modern applications. Traversing LDS typically involves a significant amount of pointer chasing. Due to the serial nature of memory access in pointer chasing, the incurred long memory latency of traversing LDS has become a critical performance bottleneck. Furthermore, the poor spatial locality in LDS makes it difficult for spatial prefetchers to predict access addresses. Although temporal prefetchers can handle irregular memory access patterns, hindered by the challenges of collecting semantic information, current state-of-the-art temporal prefetchers suffer from significant metadata redundancy and frequent metadata conflicts. Consequently, there remain substantial opportunities to enhance the LDS prefetching. To solve this problem, we propose Augur, a semantics-aware temporal prefetcher to enhance LDS performance. Augur utilizes a novel pruning method to obtain semantic information and effectively extracts node address correlations from the perspective of nodes in LDS, thereby diminishing the metadata redundancy and conflicts. Additionally, Augur employs efficient metadata management strategies that guarantee a minimal storage overhead. Evaluated on LDS workloads, Augur achieves an average performance speedup of 17.8% and 11.7% over the baseline stride prefetcher and state-of-the-art spatial prefetcher Berti, respectively. Furthermore, Augur outperforms the state-of-the-art temporal prefetcher MISB, Triage, and Triangel, by 17.4%, 12.8%, and 6.3%, respectively, with a significantly lower storage overhead of only 1.26 KB. Junliang Wu, Chenji Han, Xinyu Li 0010, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | AVM-BTB: Adaptive and Virtualized Multi-level Branch Target BufferabstractBranch Target Buffer (BTB) plays an important role in modern processors. It is used to identify branches in the instruction stream and predict branch targets. The accuracy of BTB is highly impacted by BTB capacity. However, expanding BTB capacity using traditional methods requires valuable on-chip SRAM. Both timing and area restriction make these approaches unsustainable. Moreover, these methods overlook the different demands of various applications, leading to increased power consumption and resource waste in some cases. To address this problem, we propose AVM-BTB. The key observations behind AVM-BTB come from three aspects: 1) BTB requirements vary over different applications and even over different running stages of the same application. 2) Micro-operation Cache (Uop Cache) and ICache exhibit inefficiency when confronted with instruction footprints that greatly exceed their capacity. 3) In specific scenarios of frontend overload, reducing cache capacity and increasing BTB size can effectively mitigate expensive branch prediction errors. Simultaneously, the implementation of Fetch Directed Instruction Prefetching (FDIP) can offset the limitations in cache capacity to some extent. These observations reveal the feasibility of dynamically borrowing cache capacity as temporary BTB and returning these BTB to cache when they are not needed, further resulting in an adaptive and virtualized multi-level BTB scheme. However, such a BTB structure is non-trivial. In this work, from the perspective of instructions, the cache hierarchy stores instruction data, while the BTB stores metadata used for branch prediction and instruction prefetch. Targeting high performance, AVM-BTB maintains a dynamic balance between data and metadata by monitoring the BTB error rate and effective accesses. Evaluation with 1253 traces shows that AVM-BTB is suitable for both frontend-bound and frontend-friendly scenarios, without consuming additional SRAM and with reasonable implementation efforts. Compared to baseline, AVM-BTB delivers an average performance boost of $18.22\%$ and a power consumption reduction of $2.77\%$. It also outperforms the five state-of-the-art solutions by $6.26 \%-18.26 \%$ on average in terms of IPC. Yunzhe Liu 0005, Xinyu Li 0010, Qi Guo 0001, Fuxin Zhang |
ISCA | 2 |
| 2024 | BTBench: A Benchmark for Comprehensive Binary Translation Performance EvaluationabstractBinary translation serves as a fundamental technol-ogy for instruction set emulation, system virtualization, runtime instrumentation, and numerous other applications. Many techniques have been proposed to enhance the efficiency of binary translation systems. However, imprecise performance evaluation leads to potential performance shortcomings of binary translators in real-world applications. Previous studies primarily employ CPU benchmarks, which may overlook performance issues spe-cific to binary translators and fail to guide for optimizing binary translation. To address this issue, we propose a new benchmark suite named BTBench(Binary Translation Benchmark), which provides a convenient, portable, and comprehensive solution. BT-Bench takes into account the inherent attributes of binary trans-lators, such as translation and code-cache lookup overhead. We carefully select benchmarks that offer comprehensive coverage and align with real-world application scenarios. To validate the effectiveness of BTBench, we conducted rigorous experimentation on four widely-used binary translators. The analysis of the results reveals that, compared to existing CPU benchmarks, BTBench is better suited for identifying potential performance shortcomings, providing invaluable insights for future optimization efforts. The BTBench benchmark suite is publicly available1• Xinyu Li 0010, Yanzhi Lan, Gen Niu, Fuxin Zhang |
ISPASS | 1 |
| 2024 | An Instruction Inflation Analyzing Framework for Dynamic Binary TranslatorsabstractDynamic binary translators (DBTs) are widely used to migrate applications between different instruction set architectures (ISAs). Despite extensive research to improve DBT performance, noticeable overhead remains, preventing near-native performance, especially when translating from complex instruction set computer (CISC) to reduced instruction set computer (RISC). For computational workloads, the main overhead stems from translated code quality. Experimental data show that state-of-the-art DBT products have dynamic code inflation of at least 1.46. This indicates that on average, more than 1.46 host instructions are needed to emulate one guest instruction. Worse, inflation closely correlates with translated code quality. However, the detailed sources of instruction inflation remain unclear. To understand the sources of inflation, we present Deflater , an instruction inflation analysis framework comprising a mathematical model, a collection of black-box unit tests called BenchMIAOes , and a trace-based simulator called InflatSim . The mathematical model calculates overall inflation based on the inflation of individual instructions and translation block optimizations. BenchMIAOes extract model parameters from DBTs without accessing DBT source code. InflatSim implements the model and uses the extracted parameters from BenchMIAOes to simulate a given DBT’s behavior. Deflater is a valuable tool to guide DBT analysis and improvement. Using Deflater, we simulated inflation for three state-of-the-art CISC-to-RISC DBTs: ExaGear, Rosetta2, and LATX, with inflation errors of 5.63%, 5.15%, and 3.44%, respectively for SPEC CPU 2017, gaining insights into these commercial DBTs. Deflater also efficiently models inflation for the open source DBT QEMU and suggests optimizations that can substantially reduce inflation. Implementing the suggested optimizations confirms Deflater’s effective guidance, with 4.65% inflation error, and gains 5.47x performance improvement. Benyi Xie, Chenghao Yan, Sicheng Tao, Xinyu Li 0010, Yanzhi Lan, Xiang Wu 0016, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | Tyche: An Efficient and General Prefetcher for Indirect Memory AccessesabstractIndirect memory accesses (IMAs, i.e., A [ f ( B [ i ])]) are typical memory access patterns in applications such as graph analysis, machine learning, and database. IMAs are composed of producer-consumer pairs, where the consumers’ memory addresses are derived from the producers’ memory data. Due to the built-in value-dependent feature, IMAs exhibit poor locality, making prefetching ineffective. Hindered by the challenges of recording the potentially complex graphs of instruction dependencies among IMA producers and consumers, current state-of-the-art hardware prefetchers either (a) exhibit inadequate IMA identification abilities or (b) rely on the run-ahead mechanism to prefetch IMAs intermittently and insufficiently. To solve this problem, we propose Tyche, 1 an efficient and general hardware prefetcher to enhance IMA performance. Tyche adopts a bilateral propagation mechanism to precisely excavate the instruction dependencies in simple chains with moderate length (rather than complex graphs). Based on the exact instruction dependencies, Tyche can accurately identify various IMA patterns, including nonlinear ones, and generate accurate prefetching requests continuously. Evaluated on broad benchmarks, Tyche achieves an average performance speedup of 16.2% over the state-of-the-art spatial prefetcher Berti. More importantly, Tyche outperforms the state-of-the-art IMA prefetchers IMP, Gretch, and Vector Runahead, by 15.9%, 12.8%, and 10.7%, respectively, with a lower storage overhead of only 0.57 KB. Chenji Han, Xinyu Li 0010, Junliang Wu, Yifan Hao 0001, Zidong Du, Qi Guo 0001, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 3 |
| 2023 | SCFM: A Statistical Coarse-to-Fine Method to Select Cross-Microarchitecture Reliable Simulation Points
Chenji Han, Hongze Tan 0001, Xinyu Li 0010, Ruiyang Wu 0001, Fuxin Zhang |
APPT | 4 |
| 2023 | On-Demand Triggered Memory Management Unit in Dynamic Binary Translator
Benyi Xie, Xinyu Li 0010, Chenghao Yan, Fuxin Zhang |
APPT | 2 |
| 2023 | LAST: An Efficient In-place Static Binary Translator for RISC Architectures
Yanzhi Lan, Gen Niu, Xinyu Li 0010, Liangpu Wang, Fuxin Zhang |
ICA3PP (2) | 4 |
| 2022 | Eliminate the overhead of interrupt checking in full-system dynamic binary translatorabstractDynamic binary translation (DBT) is a ubiquitous technique for program emulation, instrumentation and debugging. Full-system dynamic binary translators, which can run operating systems, are required to emulate interrupt delivery. Existing full-system dynamic binary translators use a simple scheme to do so, by attaching to each translated code block a prologue that checks for pending interrupts. However, this approach is inefficient, as interrupts are delivered infrequently, relatively to the execution of translated blocks, and therefore most of the interrupt checks are unnecessary and wasteful. Gen Niu, Fuxin Zhang, Xinyu Li 0010 |
SYSTOR | 3 |