EDBT 2026 Demo / reviewers in the wild / expert
Gen Niu
dblp:234/0158
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-2401-8021ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tiaozhuan: A General and Efficient Indirect Branch Optimization for Binary TranslationabstractBinary translation enables transparent execution, analysis, and modification of the binary program, serving as a core technology that facilitates instruction set emulation, cross-platform compatibility of software, and program instrumentation. Handling indirect branch instructions is widely recognized as a significant performance bottleneck in binary translation. While the target of a direct branch can be determined during the translation phase, an indirect branch requires a runtime lookup from the guest program counter to the host program counter, significantly influencing the performance of translator. Although several methods have been proposed to accelerate this process, each guest indirect branch instruction still translates into approximately 10 host instructions, resulting in considerable overhead. This article introduces Tiaozhuan, which addresses this issue by employing two optimization schemes. First, full address mapping uses a larger address space to store address mappings from guest to host, effectively reducing the number of instructions required to lookup the target of an indirect branch. Second, exceptionassisted branch elimination further eliminates branch instructions that check target correctness of targets in the lookup process. These two approaches enable indirect branches target lookup to be completed within one to two instructions, noticeably decreasing the overhead of indirect branches. Compared to state-of-the-art mechanisms, the SPEC CPU2006 benchmark suite showed a reduction in the number of instructions by an average of 4.2%, with the highest observed performance improvement reaching 19.4% and an average increase of 3.9%. Xinyu Li 0010, Guangyao Guo, Yanzhi Lan, Chenji Han, Gen Niu, Fuxin Zhang |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | BTBench: A Benchmark for Comprehensive Binary Translation Performance EvaluationabstractBinary translation serves as a fundamental technol-ogy for instruction set emulation, system virtualization, runtime instrumentation, and numerous other applications. Many techniques have been proposed to enhance the efficiency of binary translation systems. However, imprecise performance evaluation leads to potential performance shortcomings of binary translators in real-world applications. Previous studies primarily employ CPU benchmarks, which may overlook performance issues spe-cific to binary translators and fail to guide for optimizing binary translation. To address this issue, we propose a new benchmark suite named BTBench(Binary Translation Benchmark), which provides a convenient, portable, and comprehensive solution. BT-Bench takes into account the inherent attributes of binary trans-lators, such as translation and code-cache lookup overhead. We carefully select benchmarks that offer comprehensive coverage and align with real-world application scenarios. To validate the effectiveness of BTBench, we conducted rigorous experimentation on four widely-used binary translators. The analysis of the results reveals that, compared to existing CPU benchmarks, BTBench is better suited for identifying potential performance shortcomings, providing invaluable insights for future optimization efforts. The BTBench benchmark suite is publicly available1• Xinyu Li 0010, Yanzhi Lan, Gen Niu, Fuxin Zhang |
ISPASS | 3 |
| 2023 | LAST: An Efficient In-place Static Binary Translator for RISC Architectures
Yanzhi Lan, Gen Niu, Xinyu Li 0010, Liangpu Wang, Fuxin Zhang |
ICA3PP (2) | 3 |
| 2022 | Eliminate the overhead of interrupt checking in full-system dynamic binary translatorabstractDynamic binary translation (DBT) is a ubiquitous technique for program emulation, instrumentation and debugging. Full-system dynamic binary translators, which can run operating systems, are required to emulate interrupt delivery. Existing full-system dynamic binary translators use a simple scheme to do so, by attaching to each translated code block a prologue that checks for pending interrupts. However, this approach is inefficient, as interrupts are delivered infrequently, relatively to the execution of translated blocks, and therefore most of the interrupt checks are unnecessary and wasteful. Gen Niu, Fuxin Zhang, Xinyu Li 0010 |
SYSTOR | 1 |
| 2021 | BTMMU: an efficient and versatile cross-ISA memory virtualizationabstractFull system dynamic binary translation (DBT) has many important applications, but it is typically much slower than the native host. One major overhead in full system DBT comes from cross-ISA memory virtualization, where multi-level memory address translation is needed to map guest virtual address into host physical address. Like the SoftMMU used in the popular open-source emulator QEMU, software-based memory virtualization solutions are not efficient. Meanwhile, mature techniques for same-ISA virtualization such as shadow page table or second level address translation are not directly applicable due to cross-ISA difficulties. Some previous studies achieved significant speedup by utilizing existing hardware (TLB or virtualization hardware) of the host. However, since the hardware is not designed with cross-ISA in mind, those solutions had some limitations that were hard to overcome. Most of them only supported guests with smaller virtual address space than the host. Some supported only guests with the same page size. And some did not support privileged memory accesses. Kele Huang, Fuxin Zhang, Gen Niu, Junrong Wu |
VEE | 4 |
| 2018 | Improving Reliability of Deduplication-Based Storage Systems with Per-File ParityabstractThe reliability issue in deduplication-based storage systems has not received adequate attention. Existing approaches introduce data redundancy after files have been deduplicated, either by replication on critical data chunks, i.e., chunks with high reference count, or RAID schemes on unique data chunks, which means that these schemes are based on individual unique data chunks rather than individual files. This can leave individual files vulnerable to losses, particularly in the presence of transient and unrecoverable data chunk errors such as latent sector errors. To address this file reliability issue, this paper proposes a Per-File Parity (short for PFP) scheme to improve the reliability of deduplication-based storage systems. PFP computes the XOR parity within parity groups of data chunks of each file after the chunking process but before the data chunks are deduplicated. Therefore, PFP can provide parity redundancy protection for all files by intra-file recovery and a higher-level protection for data chunks with high reference counts by inter-file recovery. Our reliability analysis and extensive data-driven, failure-injection based experiments conducted on a prototype implementation of PFP show that PFP significantly outperforms the existing redundancy solutions, DTR and RCR, in system reliability, tolerating multiple data chunk failures and guaranteeing file availability upon multiple data chunk failures. Moreover, a performance evaluation shows that PFP only incurs an average of 5.7% performance degradation to the deduplication-based storage system. Suzhen Wu, Huagao Luan, Bo Mao 0003, Hong Jiang 0001, Gen Niu, Hui Rao, Jindong Zhou |
SRDS | 5 |