Haoyu Liao

dblp:377/6009 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0009-0004-1611-5360ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Dr.avx: A Dynamic Compilation System for Seamlessly Executing Hardware-Unsupported Vectorization Instructions
abstract
Modern processors are breaking a fundamental rule: backward compatibility within their own ISA families. We term this Generational ISA Fragmentation (GIF), where newer processors cannot execute instructions supported by prior generations within the same ISA family. This phenomenon is exemplified by Intel’s removal of AVX-512 from Alder Lake processors after years of deployment, ARM’s inconsistent support for SVE across cores, and RISC-V’s incompatible vector specifications. GIF causes illegal instruction crashes when running applications optimized for earlier processors on newer hardware, threatening the foundation of software portability that has underpinned decades of computing evolution.We introduce Dr.avx, a dynamic compilation system that enables seamless execution of AVX-512 instructions on hardware that lacks native support. Dr.avx addresses the most instructive GIF instance, x86 AVX-512 fragmentation, by targeting the integer and floating-point operations that dominate real workloads. Our rewrite engine performs a fine-grained classification of AVX-512 opcode-operand patterns and employs three complementary strategies: Instr Mirroring, AVX Lowering, and Scalar Fallback. Experiments show that Dr.avx incurs a geometric mean overhead of 1.44× on SPEC CINT2017 relative to native AVX-512 execution, 17.3% better than Intel’s closed-source SDE. On production databases, Dr.avx sustains 75%–88% (MySQL) and 86%–99% (MongoDB) of native throughput, yielding 2.0–2.7× higher throughput than SDE. For LLM inference (llama.cpp), Dr.avx keeps 95%–99% of native tokens/s and delivers 2.5–4.8× speedup over SDE. Unlike Intel’s proprietary SDE, which provides no visibility into its implementation details, Dr.avx achieves functional correctness while providing an open, extensible, near-native performance implementation. Our work offers both a remedy for AVX-512 fragmentation and a blueprint for addressing similar compatibility challenges emerging across all major ISAs.
Mianzhi Wu, Haoyu Liao, Jianmei Guo, Bo Huang 0002
CGO4
2026 Constructing quantum circuits for Grover key search in MK-3 featuring 16-bit S-boxes
abstract
Abstract Quantum computing poses a significant threat to modern cryptosystems, prompting nations to initiate the development of next-generation cryptographic algorithms where quantum security assessment constitutes a critical task. In this work, we present the first complete quantum attack circuit construction and security evaluation of quantum resistance for the authenticated encryption algorithm MK-3, which features 16-bit S-boxes. Firstly, we achieve breakthroughs in finite field arithmetic. The arrangement method of Toffoli gates in the subfield $$F_{2^4}$$ F 2 4 enables two types of optimized $$F_{2^8}$$ F 2 8 multiplication circuits to reduce DW-cost by at least 50%. For $$F_{2^8}$$ F 2 8 inversion, by proposing a new intermediate state uncomputing technique, the DW-cost of the inversion circuit is reduced by at least 40% compared with existing solutions. This can also be used to optimize the quantum circuits of block ciphers such as AES and SM4. Secondly, we propose three types of quantum circuits for MK-3’s 16-bit S-box. The basic circuit reduces multiplicative bit operations to 115, representing at least 20% reduction. We propose a more compact S-box structure, reducing the number of circuit qubits for two types of S-boxes from 52 to 40. Finally, we implement full-algorithm quantum attacks and conduct security evaluations for MK-3. Complete quantum circuits supporting 128-bit and 256-bit keys are constructed, enabling Grover key search attacks. Security assessments confirm the 128-bit version meets NIST Security Category 1 (with reasonable inference of achieving Category 2), while the 256-bit version satisfies the highest Category 5 standards, demonstrating resistance against large-scale quantum computing attacks. This work achieves the first quantum circuit implementation for 16-bit S-boxes and provides critical evaluation benchmarks for post-quantum cryptography standardization.
Haoyu Liao, Qingbin Luo, Yuanmeng Zheng, Lang Ding
Cybersecur.1
2025 Quantum Circuit Synthesis for AES with Low DW-Cost
Haoyu Liao, Qingbin Luo
ASIACRYPT (2)1
2025 PSCA: A FPGA-based Protein Structure Comparison Accelerator with Symmetric Simplified Matrix
Hui Su, Xingyun Qi, Qiang Wang 0006, Puguang Liu, Haoyu Liao
ICA3PP (6)6
2025 Retrospecting Available CPU Resources: SMT-Aware Scheduling to Prevent SLA Violations in Data Centers
abstract
The article focuses on an understudied yet fundamental problem: existing methods typically average the utilization of multiple hardware threads to evaluate the available CPU resources. However, the approach could underestimate the actual usage of the underlying physical core for Simultaneous Multi-Threading (SMT) processors, leading to an overestimation of remaining resources. The overestimation propagates from microarchitecture to operating systems and cloud schedulers, which may misguide scheduling decisions, exacerbate CPU overcommitment, and increase Service Level Agreement (SLA) violations. To address the potential overestimation problem, we propose an SMT-aware and purely data-driven approach namedRemaining CPU(RCPU) that reserves more CPU resources to restrict CPU overcommitment and prevent SLA violations. RCPU requires only a few modifications to the existing cloud infrastructures and can be scaled up to large data centers. Extensive evaluations in the data center proved that RCPU contributes to a reduction of SLA violations by 18% on average for 98% of all latency-sensitive applications. Under a benchmarking experiment, we prove that RCPU increases the accuracy by 69% in terms of Mean Absolute Error (MAE) compared to the state-of-the-art.
Haoyu Liao, Tong-Yu Liu, Jianmei Guo, Bo Huang 0002, Dingyu Yang, Jonathan Ding
IEEE Trans. Parallel Distributed Syst.1
2024 Optimization of TDM Using Single-ended Transmission for Multi-FPGA Platforms
abstract
In large-scale designs, the multi-FPGA system is a popular approach for hardware acceleration and pre-silicon verification due to its scalable capabilities. Researchers aim to enhance communication bandwidth and decrease latency between FPGAs, all while working within the constraints of limited physical pins available. To address this issue, we propose an optimization for time-division multiplexing(TDM). This optimization combines the advantages of traditional logic multiplexer circuits and high-speed serial circuits by utilizing the serial-to-parallel converter (ISERDES) and parallel-to-serial converter (OSERDES). The I/OSERDES approach typically employs the low-voltage differential signaling standard (LVDS) for high-speed data transmission, necessitating two physical pins. To save one physical pin, we advocate for a single-ended transmission solution without a forward clock. Moreover, we propose a new algorithm at the receiver to align the phase of the receiver’s clock. In comparison to the LVDS-based solution, our proposed interface achieves double the communication bandwidth of inter-FPGA chips without introducing additional system latency under the same TDM rate.
Haoyu Liao, Puguang Liu, Xingyun Qi
ISCAS1
2024 A High-performance Hardware Accelerator for Genome Alignment
abstract
Genome alignment is a vital process in genome sequencing and bio-informatics research. It entails aligning short DNA sequence fragments (usually tens to hundreds of base pairs) with a reference genome sequence, uncovering crucial biological information and variations. However, due to the high computational complexity of current sequence matching algorithms and the rapid growth of genetic data, there exists a computational bottleneck in sequence alignment workflows. Therefore, scholars have turned to using hardware to accelerate this computation, with a focus on high-performance computing. While existing acceleration solutions have mostly concentrated on the classical algorithm, showing significant improvements, there has been limited work on accelerating alignment algorithms in popular bio-informatics software tools (enhanced version). In this paper, we proposed a hardware accelerator for the alignment tasks in the popular sequence alignment tool minimap2 using the KSW2 algorithm. Our approach utilizes an anti-diagonal processing element(PE) array for parallel computation, implements the BAND technology in hardware, and enables support for longer input sequences without sacrificing accuracy. Additionally, we integrated the traceback stage of the KSW2 algorithm in hardware, reducing communication data overhead. We attained a 15.32x acceleration compared to software optimized with Streaming SIMD Extensions (SSE) instruction sets and hyper-threading technology. Additionally, our design shows a 3.42x speedup compared to other high-performance hardware.
Haoyu Liao, Hui Su, Xingyun Qi
ISPA1
2024 DeployFix: Dynamic Repair of Software Deployment Failures via Constraint Solving
abstract
Software deployment misconfiguration often happens and has been one of the major causes of deployment failures that give rise to service interruptions. However, there is currently no existing approach to automatically repairing deployment failures. We propose DeployFix, which automatically repairs software deployment failures via constraint solving in the dynamic-changing deployment environments. DeployFix first defines DeployIR as a unified intermediate representation to achieve the translation of heterogeneous specifications from different schedulers with different syntaxes. By reducing the root-cause analysis of deployment failures to the conflict resolution in propositional logic, DeployFix uses off-the-shelf constraint solvers to achieve automatic localization and diagnosis of conflicting constraints, which are the root causes of deployment failures. DeployFix finally resolves the conflicting constraints and generates repaired deployment configurations in terms of practical requirements. We evaluate DeployFix in both simulation and production environments with tens of thousands of nodes at Alibaba, on which tens of thousands of applications are running guided by hundreds of thousands of deployment constraints. Experimental results demonstrate that DeployFix outperforms the state of the art and it correctly repairs the deployment failures in minutes, even in a large production data center.
Haoyu Liao, Jianmei Guo, Bo Huang 0002, Yujie Han, Dingyu Yang, Kai Shi 0006, Jonathan Ding, Guoyao Xu, Liping Zhang 0013
ASE1
2024 EFACT: An External Function Auto-Completion Tool to strengthen static binary lifting
Haoyu Liao, Bo Huang 0002, Jianmei Guo
J. Syst. Softw.2