Bo Huang 0002

dblp:95/6229-2 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-5126-7192ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 since 2021Software engineering, systems software and programming languages · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Dr.avx: A Dynamic Compilation System for Seamlessly Executing Hardware-Unsupported Vectorization Instructions
abstract
Modern processors are breaking a fundamental rule: backward compatibility within their own ISA families. We term this Generational ISA Fragmentation (GIF), where newer processors cannot execute instructions supported by prior generations within the same ISA family. This phenomenon is exemplified by Intel’s removal of AVX-512 from Alder Lake processors after years of deployment, ARM’s inconsistent support for SVE across cores, and RISC-V’s incompatible vector specifications. GIF causes illegal instruction crashes when running applications optimized for earlier processors on newer hardware, threatening the foundation of software portability that has underpinned decades of computing evolution.We introduce Dr.avx, a dynamic compilation system that enables seamless execution of AVX-512 instructions on hardware that lacks native support. Dr.avx addresses the most instructive GIF instance, x86 AVX-512 fragmentation, by targeting the integer and floating-point operations that dominate real workloads. Our rewrite engine performs a fine-grained classification of AVX-512 opcode-operand patterns and employs three complementary strategies: Instr Mirroring, AVX Lowering, and Scalar Fallback. Experiments show that Dr.avx incurs a geometric mean overhead of 1.44× on SPEC CINT2017 relative to native AVX-512 execution, 17.3% better than Intel’s closed-source SDE. On production databases, Dr.avx sustains 75%–88% (MySQL) and 86%–99% (MongoDB) of native throughput, yielding 2.0–2.7× higher throughput than SDE. For LLM inference (llama.cpp), Dr.avx keeps 95%–99% of native tokens/s and delivers 2.5–4.8× speedup over SDE. Unlike Intel’s proprietary SDE, which provides no visibility into its implementation details, Dr.avx achieves functional correctness while providing an open, extensible, near-native performance implementation. Our work offers both a remedy for AVX-512 fragmentation and a blueprint for addressing similar compatibility challenges emerging across all major ISAs.
Mianzhi Wu, Haoyu Liao, Jianmei Guo, Bo Huang 0002
CGO6
2026 FlexInstru: A flexible instrumentation framework for tracing long-running native workloads
Wenlong Mu, Ning Li 0054, Zimo Ji, Jianmei Guo, Bo Huang 0002
J. Syst. Softw.5
2025 AOBO: A Fast-Switching Online Binary Optimizer on AArch64
abstract
As the complexity of real-world server applications continues to grow, performance optimizations for large-scale applications are becoming increasingly challenging. The success of online optimization offered by OCOLOS and Dynimize proves that binary rewriting based on edge profiling data can significantly accelerate these applications. However, no similar online binary optimizer is currently available on the AArch64 platform. In response to the growing adoption of the AArch64 platform, this article introduces AOBO, a fast-switching online binary optimizer specifically designed for AArch64. In addition to providing practical and efficient engineering support for AArch64-specific features, AOBO overcomes the challenge of lacking hardware counters for edge profiling on most commercially available AArch64 servers. In particular, AOBO embraces a novel edge weight estimation scheme to deliver more accurate edge estimation, which in turn allows AOBO’s binary rewriter to generate more efficient code. Furthermore, time spent on AOBO’s online code replacement stage is optimized to work at a subsecond level, thus enabling a fast switch from running the original binary to running the optimized one. We evaluate AOBO with CINT2017, GCC, MySQL and MongoDB, measuring the accuracy and coverage of the estimated edge weights, the performance improvements of the optimized binaries, and the online optimization cost. To make a fair comparison, we are using the performance data of the binaries generated by the default compilation scripts in the software packages as a baseline. Experimental data shows that AOBO can offer a more accurate edge weight estimation and generate binaries with superior performance. Furthermore, AOBO achieves online optimization with a very small overhead and significantly improves the performance of large-scale applications. Compared with the baselines, AOBO’s online optimization can achieve 24.7% and 31.11% performance improvement respectively for MySQL and MongoDB. Notably, application pause time is reduced from 1,599.8 milliseconds to 462.1 milliseconds for MySQL, and from 1,765.9 milliseconds to 507.1 milliseconds for MongoDB.
Wenlong Mu, Bo Huang 0002, Jianmei Guo
ACM Trans. Archit. Code Optim.3
2025 Retrospecting Available CPU Resources: SMT-Aware Scheduling to Prevent SLA Violations in Data Centers
abstract
The article focuses on an understudied yet fundamental problem: existing methods typically average the utilization of multiple hardware threads to evaluate the available CPU resources. However, the approach could underestimate the actual usage of the underlying physical core for Simultaneous Multi-Threading (SMT) processors, leading to an overestimation of remaining resources. The overestimation propagates from microarchitecture to operating systems and cloud schedulers, which may misguide scheduling decisions, exacerbate CPU overcommitment, and increase Service Level Agreement (SLA) violations. To address the potential overestimation problem, we propose an SMT-aware and purely data-driven approach namedRemaining CPU(RCPU) that reserves more CPU resources to restrict CPU overcommitment and prevent SLA violations. RCPU requires only a few modifications to the existing cloud infrastructures and can be scaled up to large data centers. Extensive evaluations in the data center proved that RCPU contributes to a reduction of SLA violations by 18% on average for 98% of all latency-sensitive applications. Under a benchmarking experiment, we prove that RCPU increases the accuracy by 69% in terms of Mean Absolute Error (MAE) compared to the state-of-the-art.
Haoyu Liao, Tong-Yu Liu, Jianmei Guo, Bo Huang 0002, Dingyu Yang, Jonathan Ding
IEEE Trans. Parallel Distributed Syst.4
2024 DeployFix: Dynamic Repair of Software Deployment Failures via Constraint Solving
abstract
Software deployment misconfiguration often happens and has been one of the major causes of deployment failures that give rise to service interruptions. However, there is currently no existing approach to automatically repairing deployment failures. We propose DeployFix, which automatically repairs software deployment failures via constraint solving in the dynamic-changing deployment environments. DeployFix first defines DeployIR as a unified intermediate representation to achieve the translation of heterogeneous specifications from different schedulers with different syntaxes. By reducing the root-cause analysis of deployment failures to the conflict resolution in propositional logic, DeployFix uses off-the-shelf constraint solvers to achieve automatic localization and diagnosis of conflicting constraints, which are the root causes of deployment failures. DeployFix finally resolves the conflicting constraints and generates repaired deployment configurations in terms of practical requirements. We evaluate DeployFix in both simulation and production environments with tens of thousands of nodes at Alibaba, on which tens of thousands of applications are running guided by hundreds of thousands of deployment constraints. Experimental results demonstrate that DeployFix outperforms the state of the art and it correctly repairs the deployment failures in minutes, even in a large production data center.
Haoyu Liao, Jianmei Guo, Bo Huang 0002, Yujie Han, Dingyu Yang, Kai Shi 0006, Jonathan Ding, Guoyao Xu, Liping Zhang 0013
ASE3
2024 EFACT: An External Function Auto-Completion Tool to strengthen static binary lifting
Haoyu Liao, Bo Huang 0002, Jianmei Guo
J. Syst. Softw.4
2024 Efficient Cross-platform Multiplexing of Hardware Performance Counters via Adaptive Grouping
abstract
Collecting sufficient microarchitecture performance data is essential for performance evaluation and workload characterization. There are many events to be monitored in a modern processor while only a few hardware performance monitoring counters (PMCs) can be used, so multiplexing is commonly adopted. However, inefficiency commonly exists in state-of-the-art profiling tools when grouping events for multiplexing PMCs. It has the risk of inaccurate measurement and misleading analysis. Commercial tools can leverage PMCs, but they are closed source and only support their specified platforms. To this end, we propose an approach for efficient cross-platform microarchitecture performance measurement via adaptive grouping, aiming to improve the metrics’ sampling ratios. The approach generates event groups based on the number of available PMCs detected on arbitrary machines while avoiding the scheduling pitfall of Linux perf_event subsystem. We evaluate our approach with SPEC CPU 2017 on four mainstream x86-64 and AArch64 processors and conduct comparative analyses of efficiency with two other state-of-the-art tools, LIKWID and ARM Top-down Tool. The experimental results indicate that our approach gains around 50% improvement in the average sampling ratio of metrics without compromising the correctness and reliability.
Tong-Yu Liu, Jianmei Guo, Bo Huang 0002
ACM Trans. Archit. Code Optim.3
2024 TCSA: Efficient Localization of Busy-Wait Synchronization Bugs for Latency-Critical Applications
abstract
Busy-wait synchronization is often used for latency-critical applications to ensure low latency. Unfortunately, its performance bugs due to thread contention may lead to request failures or even system crashes. Localizing the performance bugs of busy-wait synchronization is not trivial because we have to pinpoint the exact moment of occurrence from a relatively long measurement period and simultaneously identify candidate busy-wait threads from numerous concurrent threads. Existing methods often rely on hotspot-driven analysis of lock-related functions, but they still need extensive manual work to localize busy-wait threads. This paper proposes timing call stack analysis (TCSA), an efficient approach to localizing busy-wait synchronization bugs. The key idea is to time-serialize the function call stacks of applications and identify consecutive identical call stacks to catch busy-wait threads. TCSA can handle any application regardless of its programming language and identify various busy-wait patterns, including spinlocks, chaining spinlocks, futexes, and safepoint checks within the Java Virtual Machine. Compared to the state-of-the-art, TCSA can effectively diminish the quantity of examined records (e.g., threads and functions) by 1 to 3 orders of magnitude. TCSA has been deployed to a large cloud service provider, demonstrating its effectiveness, efficiency, and practicality in four real latency-critical applications.
Ning Li 0054, Jianmei Guo, Bo Huang 0002, Chengdong Li, Wenxin Huang
IEEE Trans. Parallel Distributed Syst.3
2023 A Hotspot-Driven Semi-automated Competitive Analysis Framework for Identifying Compiler Key Optimizations
abstract
High-performance compilers play an important role in improving the run-time performance of a program, and it is hard and time-consuming to identify the key optimizations implemented in a high-performance compiler with traditional program analysis. In this paper, we propose a hotspot-driven semi-automated competitive analysis framework for identifying key optimizations through comparing the hotspot codes generated by any two different compilers. Our framework is platform-agnostic and works well on both AArch64 and X64 platforms, which automates the stages of hotspot detection and dynamic binary instrumentation only for selected hotspots. With the instrumented instruction characterization information, the framework users can analyze the binary code within a much smaller scope to explore practical optimizations implemented in any of the compilers compared. To demonstrate the effectiveness and practicality, we conduct experiments on SPECspeed 2017 Integer benchmarks(CINT2017) and their binaries generated by open-source GCC compiler versus proprietary Huawei BiSheng and Intel ICC compilers on AArch64 and X64 platforms respectively. Empirical studies show that our methods can identify several significant optimizations that have been implemented by proprietary compilers and as well can be implemented in open-source compilers. To Hangzhou Hongjun Microelectronics Technology(Hjmicro), the identified key optimizations shed great light on optimizing their GCC-based product compiler, which delivers 20.83% improvement for SPECrate 2017 Integer on AArch64 platform.
Wenlong Mu, Bo Huang 0002, Jianmei Guo, Shiqiang Cui
CC3
2011 HiTune: Dataflow-Based Performance Analysis for Big Data Cloud
Jinquan Dai, Shengsheng Huang, Bo Huang 0002
USENIX ATC4
2009 SharK: A Web 2.0 Service Infrastructure for Knowledge Sharing
Bo Huang 0002, Junyong Ding, Jinquan Dai
CSEDU (1)1
2009 Control flow obfuscation with information flow tracking
abstract
Recent micro-architectural research has proposed various schemes to enhance processors with additional tags to track various properties of a program. Such a technique, which is usually referred to as information flow tracking, has been widely applied to secure software execution (e.g., taint tracking), protect software privacy and improve performance (e.g., control speculation).
Haibo Chen 0001, Liwei Yuan 0003, Xi Wu 0001, Binyu Zang, Bo Huang 0002, Pen-Chung Yew
MICRO5
2007 Pipelined Execution of Critical Sections Using Software-Controlled Caching in Network Processors
abstract
To keep up with the explosive Internet packet processing demands, modern network processors (NPs) employ a highly parallel, multi-threaded and multi-core architecture. In such a parallel paradigm, accesses to the shared variables in the external memory (and the associated memory latency) are contained in the critical sections, so that they can be executed atomically and sequentially by different threads in the network processor. In this paper, we present a novel program transformation that is used in the Intelreg Auto-partitioning C Compiler for IXP to exploit the inherent finer-grained parallelism of those critical sections, using the software-controlled caching mechanism available in the NPs. Consequently, those critical sections can be executed in a pipelined fashion by different threads, thereby effectively hiding the memory latency and improving the performance of network applications. Experimental results show that the proposed transformation provides impressive speedup (up-to 9.9times) and scalability (up-to 80 threads) of the performance for the real-world network application (a 10Gbps Ethernet Core/Metro Router)
Jinquan Dai, Bo Huang 0002
CGO3
2007 Metadata driven memory optimizations in dynamic binary translator
abstract
A dynamic binary translator offers solutions for translating and running source architecture binaries on target architecture at runtime. Regardless of its growing popularity, practical dynamic binary translators usually suffer from the limited optimizations performed when generating the translated code due to the lack of useful information available in the executable files and the requirement to conform to the binary-level compatibility. Trying to generate more efficient translated code, we propose in this paper a novel method of passing performance critical information to a dynamic binary translator through the metadata section generated during the static compilation phase. With the performance critical metadata, the dynamic binary translator is able to perform aggressive optimizations to generate higher quality code. We implemented a general and extensible framework in GCC 4.0 and IA-32® Execution Layer, and selected metadata related to memory optimizations as our target. The metadata enables IA-32 EL to perform memory optimizations such as registerization, memory ordering relaxation and address disambiguation of memory instructions. Experimental data shows an overall performance improvement of 15.03% for SPECfp2000 and 1.21% for SPECint2000. For some specific benchmarks, the performance improvement is even up to 37.09%.
Chaohao Xu, Bo Huang 0002
VEE5
2006 Optimizing Dynamic Binary Translation for SIMD Instructions
abstract
Dynamic binary translation technology allows a program written for one architecture to be executed on a second architecture without recompiling the source code. Effective dynamic binary translation of SIMD (single instruction multiple data) instructions has become more and more important as SIMD extensions have gained popularity among general-purpose CPUs within the last decade. Many SIMD extensions allow an SIMD register to hold data of different types at different times. Supporting multiple data types within the same register complicates the task of a dynamic translator, which may or may not be able to determine the type of the register at translation time. We propose the SIMD data type tracking algorithm to translate the SIMD instructions and three algorithms to further optimize the translation. Our results show that the three optimizing algorithms give overall 3.89% performance improvement for SPEC2K INT benchmarks and 6.61%for SPEC2KFP benchmarks.
Shu Xu 0004, Bo Huang 0002
CGO4
2005 Boosting the Performance of Multimedia Applications Using SIMD Instructions
Weihua Jiang, Bo Huang 0002, Binyu Zang, Chuanqi Zhu
CC3
2005 Automatically partitioning packet processing applications for pipelined architectures
abstract
Modern network processors employs parallel processing engines (PEs) to keep up with explosive internet packet processing demands. Most network processors further allow processing engines to be organized in a pipelined fashion to enable higher processing throughput and flexibility. In this paper, we present a novel program transformation technique to exploit parallel and pipelined computing power of modern network processors. Our proposed method automatically partitions a sequential packet processing application into coordinated pipelined parallel subtasks which can be naturally mapped to contemporary high-performance network processors. Our transformation technique ensures that packet processing tasks are balanced among pipeline stages and that data transmission between pipeline stages is minimized. We have implemented the proposed transformation method in an auto-partitioning C compiler product for Intel Network Processors. Experimental results show that our method provides impressive speed up for the commonly used NPF IPv4 forwarding and IP forwarding benchmarks. For a 9-stage pipeline, our auto-partitioning C compiler obtained more than 4X speedup for the IPv4 forwarding PPS and the IP forwarding PPS (for both the IPv4 traffic and IPv6 traffic).
Jinquan Dai, Bo Huang 0002, Luddy Harrison
PLDI2
2005 Automatic multithreading and multiprocessing of C programs for IXP
abstract
Effective compilation of packet processing applications onto the Intel IXP network processors requires, among other things, the automatic use of multiple threads on one or more processing elements, and the automatic introduction of synchronization as required to correctly enforce dependences between such threads. We describe the program transformation that is used in the Intel Auto-partitioning C Compiler for IXP to automatically multithread/multi-process a program for the IXP. This transformation consists of steps that introduce inter-thread signaling to enforce dependences, optimize the placement of such signaling, reduce the number of signals in use to the number available in hardware, and transform the initialization code for correct execution in the multithreaded version. Experimental results show that our method provides impressive speedup for six PPSes (Packet Processing Stages) in the widely used NPF IP forwarding benchmarks. For most packet processing stages, our algorithms can achieve almost linear performance improvement after automatic multi-threading transformation. The automatic multi-processing transformation help further boost the speedup of two PPSes.
Bo Huang 0002, Jinquan Dai, Luddy Harrison
PPoPP2
2001 A New Approach to Pointer Analysis for Assignments
Bo Huang 0002, Binyu Zang, Chuanqi Zhu
J. Comput. Sci. Technol.1