Haojun Xia

dblp:297/4387 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Software-Defined platform management for data center: security, low entropy, and efficiency
abstract
Abstract The trend of heterogeneous servers and the rise of Software-Defined Data Center (SDDC) have transformed data center management. Collaborative management of hardware and software is crucial for rapid deployment and migration. As the boundary between physical infrastructure and virtual infrastructure blurs, data center management faces challenges in fine-grained resource provisioning, energy efficiency optimization, and security assurance. To address these challenges, this paper proposes a novel Software-Defined Platform Management (SDPM) architecture based on out-of-band management. This architecture extends server platform management capabilities from physical infrastructure to virtual machines. By abstracting heterogeneous resources into execution points managed by a centralized control plane and consolidating standard industry interfaces, the architecture introduces capabilities for resource provisioning, energy consumption regulation, as well as access control and trusted computing support. A prototype implementation on a real server and experimental results demonstrate that the architecture can dynamically allocate resources based on predictions of virtual machine workloads, optimize energy consumption through workload-aware and temperature-driven fan control, and support secure communication channels to implement advanced access control policies. These results highlight SDPM’s potential in advancing resource provisioning, energy efficiency, and security in modern data centers.
Haojun Xia, Bibo Tu
Cybersecur.2
2025 Interference Monitoring for Colocated Workloads in Low-Entropy Computing Systems
abstract
Users' expectations of internet applications have extended beyond basic functionality to prioritize tail latency and minimal jitter, factors often impacted by competitive interference among colocated workloads in operating systems. In the context of low-entropy computing, it is essential to monitor thread-level interference among colocated processes in a real-time and software-defined manner. This paper introduces NoiseCatcher, an interference monitoring solution that provides nanosecond-level reporting with minimal overhead on mission-critical threads. By leveraging eBPF (Extended Berkeley Packet Filter) and in-kernel shared data structures, our approach efficiently captures essential system events without requiring costly context switches. Aggregated data is then transmitted to the Baseboard Management Controller (BMC) via the system bus for persistent storage, with access provided through open RESTful APIs for flexible querying. The solution is evaluated on a production server, with comparative analysis against the Linux kernel's OSNOISE tracer. Results indicate that our solution not only matches OSNOISE in functionality but also achieves nanosecond-level precision, minimized performance degradation.
Haojun Xia, Yanchang Feng, Bibo Tu
CSCWD2
2025 BLFair: enabling proportional I/O sharing for NVMe SSD in SPDK para-virtualization architecture
abstract
Abstract In data centers, the Storage Performance Development Kit (SPDK) para-virtualization architecture is an efficient solution for non-volatile memory express (NVMe) solid-state drive (SSD) virtualization but faces challenges in maintaining performance fairness and isolation due to storage resource competition among multi-tenants. However, the existing Quality of Service method in SPDK fails to ensure proportional I/O sharing among multi-tenants. Providing fairness and isolation while maintaining high storage utilization in SPDK remains a challenge. In this paper, we propose BLFair to address this problem. Specifically, BLFair implements proportional I/O sharing for multi-tenants in the SPDK. The design of BLFair can effectively reduce the high time complexity caused by the ordering and the overhead of maintaining the virtual clock. Moreover, BLFair allows for achieving a trade-off between proportional I/O sharing and maximizing storage utilization. BLFair also uses the lockless ring mechanism to achieve scalability for cross-core operation. We have implemented a prototype system of BLFair in SPDK. Finally, we conduct evaluations with different workloads in both local storage and NVMe over RDMA fabric environments. The results show that our method can achieve fairness and scalability. BLFair can achieve up to 7.09x 99.99th latency reduction compared to the system with no fairness. Evaluation results in realistic workloads also show that BLFair outperforms other methods.
Yanchang Feng, Haojun Xia, Bibo Tu
Comput. J.3
2025 Thermal Elasticity-Aware Host Resource Provision for Carbon Efficiency on Virtualized Servers
Haojun Xia, Yanchang Feng, Haohao Liu, Bibo Tu
IEEE Trans. Computers2
2024 Desktop Virtualization Optimization Methods Based on IDV Architecture
abstract
With the growing demand for users’ flexible use of office desktops and enterprises’ centralized management of information resources, desktop virtualization technologies, represented by remote desktops, have become a prominent approach in current desktop management. Among the mainstream desktop virtualization technologies, Intelligent Desktop Virtualization (IDV) architecture offers significant advantages regarding network dependency and resource utilization. However, the IDV scenario presents challenges on the server, such as the increasing number of centrally managed images, resource consumption due to frequent image pulling, and low efficiency in image synchronization. Moreover, the terminal faces an issue of not fully leveraging hardware resources. Considering the IDV-specific characteristics, we design and implement a set of optimization methods for desktop virtualization, which outperform traditional IDV solutions.
Haojun Xia, Chen Li 0066, Bibo Tu
CSCWD2
2024 An Efficient Caching Mechanism for End-host Network Functions
abstract
The end host serves as a natural enforcement point for various network functions (NFs), such as network address translators (NATs), firewalls, and load balancers. However, due to the limitations of the Linux networking stack, NFs struggle to achieve high performance when utilizing high-speed network interfaces. The eXpress Data Path (XDP) is a high-performance framework for packet processing within the Linux kernel. It operates as an optimized execution point before the networking stack. In contrast to kernel-bypass solutions like DPDK, XDP offers an appealing alternative by providing comparable performance with lower CPU usage.In this paper, we propose PFC, a novel approach that leverages XDP for Pre-Function table Caching. PFC acts as a packet header processor for incoming packets, it consists of two distinct facets: one involves traffic management within the end host, and the other focuses on processing requests from distributed applications. Experimental results show that PFC can significantly increase throughput and achieve the equivalent performance compared to DPDK. Furthermore, PFC can integrate seamlessly with existing systems without requiring any modifications to the applications.
Haojun Xia, Chen Li 0066, Bibo Tu
CSCWD2
2024 Continuous Authentication Technology Based on Device Driver Behavior
abstract
At present, the existing peripheral interface authentication is mostly based on static features, which cannot effectively resist the security threats in the new form. This paper proposes a continuous authentication technology of device driver behavior from a dynamic perspective, aiming to achieve continuous authentication of user identity through real-time monitoring and analysis of long-time interaction behavior between the device and the system. According to the process of peripheral access to the system, the device driver behavior is divided into the device driver behavior in the enumeration phase and the device driver behavior in the user usage phase. In the enumeration phase, this paper proposes an authentication technique based on device enumeration fingerprints; while in the user usage phase, continuous authentication of user identity is realized by continuous authentication technique in peripheral communication mode. We use three supervised learning algorithms for experiments, and their accuracy rate reaches more than 97.25%. Compared with the traditional identity authentication with static features, the dynamic continuous authentication through these two phases significantly enhances the security of identity authentication in the whole process of the peripheral.
Haojun Xia, Haohao Liu, Bibo Tu
CSCWD1
2024 EI-XIDS: An explainable intrusion detection system based on integration framework
abstract
The application of Deep Learning (DL) in Intrusion Detection Systems (IDS) has become a focal point of research due to its outstanding performance. However, the black-box nature of these systems has raised concerns within the research community. Addressing this challenge, this paper draws upon the concept of ensemble learning and introduces an Explainable Intrusion Detection System (X-IDS), EI-XIDS. This system integrates a variety of advanced Explainable Artificial Intelligence (XAI) methods and adaptively selects them according to different scenarios through reinforcement learning. Comparative experiments demonstrate that EI-XIDS outperforms the current state-of-the-art explanation methods, achieving label flip rates of 97% and 96% on the NSL-KDD and UNSW-NB15 datasets, respectively. These results underscore EI-XIDS’s superior interpretability accuracy, robustness, and sparsity, showcasing its significant potential in the field of network security.
Chen Li 0066, Kun Zhang 0016, Haojun Xia, Bibo Tu
CSCWD4
2024 CloudFusion: Multi-Source Intrusion Detection in Cloud Environments
abstract
Addressing the multifaceted security challenges inherent in cloud environments, our study delineates a robust, real-time threat detection framework. This methodology integrates three cardinal technologies: a memory access mechanism rooted in Virtual Machine Monitor (VMM) analytics, offering profound insights into the operational dynamics of virtual machines; a semantic reconstruction method, informed by software architectural tenets, adept at discerning intricate adversarial activities; and a log-oriented decoding and rule alignment mechanism tailored for sophisticated handling of cloud-based log data. In unison, these technologies forge a proficient, instantaneous threat detection paradigm. Applied and authenticated on a cloud platform, the proposed framework buttressed by judiciously crafted security protocols enables the contemporaneous surveillance of malicious incursions affecting virtual machines, host systems, and network data streams. Both functionality and efficiency assessments attest to the system’s adeptness in precise threat identification while ensuring minimal performance disruption for host and client systems.
Kun Zhang 0016, Haojun Xia, Bibo Tu, Chen Li 0066
CSCWD3
2024 MonoNN: Enabling a New Monolithic Optimization Space for Neural Network Inference Tasks on Modern GPU-Centric Architectures
Donglin Zhuang, Zhen Zheng, Haojun Xia, Xiafei Qiu, Wei Lin 0016, Shuaiwen Song
OSDI3
2024 A Self-Supervised Targeted Process Anomaly Detection Method Based on the Minimum Set of Observed Events
abstract
In scenarios involving a single targeted application service, it is essential to monitor the security of business-oriented processes. However, there is currently a lack of lightweight, real-time online anomaly detection methods for targeted processes that do not require labeled data. This paper presents a self-supervised anomaly detection model for targeted processes, based on a minimal observation event set and utilizing eBPF and deep learning techniques. The model first selects a minimal set of observed events for the targeted process, which includes critical system calls, process scheduling, resource usage, and I/O operations. Unlike traditional system call sequence features, this model focuses on the rate of change in the frequency of feature selection calls. The detection model employs the VAE-LSTM algorithm, where the VAE module constructs robust short windows and the LSTM module estimates long-term correlations within the sequence. Through self-supervised learning, the model learns the normal behavior of targeted processes, extracts robust behavioral features, and reduces feature dimensions. By performing online detection of these learned features against runtime processes, the model achieves second-level anomaly detection for targeted processes. Finally, by simulating and constructing eight different types of process attack scenarios, the experimental results demonstrate a detection accuracy exceeding 92%, with a performance overhead of less than 10%.
Haojun Xia, Limin Sun 0001, Zhanwei Song, Bibo Tu
TrustCom1
2024 Quant-LLM: Accelerating the Serving of Large Language Models via FP6-Centric Algorithm-System Co-Design on Modern GPUs
Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen 0004, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, Olatunji Ruwase, Yuxiong He, Shuaiwen Song
USENIX ATC1
2023 Flash-LLM: Enabling Low-Cost and Highly-Efficient Large Generative Model Inference With Unstructured Sparsity
abstract
With the fast growth of parameter size, it becomes increasingly challenging to deploy large generative models as they typically require large GPU memory consumption and massive computation. Unstructured model pruning has been a common approach to reduce both GPU memory footprint and the overall computation while retaining good model accuracy. However, the existing solutions do not provide an efficient support for handling unstructured sparsity on modern GPUs, especially on the highly-structured tensor core hardware. Therefore, we propose Flash-LLM for enabling low-cost and highly efficient large generative model inference with the sophisticated support of unstructured sparsity on high-performance but highly restrictive tensor cores. Based on our key observation that the main bottleneck of generative model inference is the several skinny matrix multiplications for which tensor cores would be significantly under-utilized due to low computational intensity, we propose a general Load-as-Sparse and Compute-as-Dense methodology for unstructured sparse matrix multiplication (SpMM). The basic insight is to address the significant memory bandwidth bottleneck while tolerating redundant computations that are not critical for end-to-end performance on tensor cores. Based on this, we design an effective software framework for tensor core based unstructured SpMM, leveraging on-chip resources for efficient sparse data extraction and computation/memory-access overlapping. Extensive evaluations demonstrate that (1) at SpMM kernel level, Flash-LLM significantly outperforms the state-of-the-art library, i.e., Sputnik and SparTA by an average of 2.9X and 1.5X, respectively.(2) At end-to-end framework level on OPT-30B/66B/175B models, for tokens per GPU-second , Flash-LLM achieves up to 3.8X and 3.6X improvement over DeepSpeed and FasterTransformer, respectively, with significantly lower inference cost.
Haojun Xia, Zhen Zheng, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li 0045, Wei Lin 0016, Shuaiwen Song
Proc. VLDB Endow.1
2023 Enabling Fast and Memory-Efficient Acceleration for Pattern Matching Workloads: The Lightweight Automata Processing Engine
abstract
Growing pattern matching applications are employing finite automata as their basic processing model. These applications match tens to thousands of patterns on a large amount of data, which brings a great challenge to conventional processors. Therefore hardware-based solutions have emerged frequently and achieved high throuphput automata processing. However, existing methods are generally difficult to achieve both processing speed and storage efficiency, and are often too heavy to be integrated into a small chip and have to rely on off-chip DRAMs or other high capacity memories even on some simple data sets, leading to the potential area and power consumption issues. In this paper, we focus on building a more lightweight automata processing engine, hoping to store the whole automata model into on-chip memory and run effectively and independently. We propose LAP, a lightweight automata processing engine. Powered with a novel automata model (A-DFA) and efficient packing algorithms, extremely high storage efficiency compared with traditional DFA is achieved in LAP. Meanwhile, we identify the key parallelization factors in the A-DFA model and then propose a specialized microarchitecture with novel instructions to further accelerate the state transition process. As a result, LAP can obtain more effective trade-off between processing speed and storage efficiency. Evaluation results show that LAP achieves extremely high storage efficiency on simple data sets, exceeding IBM's RegX by 8×, and achieves significant improvements in processing speed ranging from 1.32× to 1.91× compared with previous lightweight hardware implementations. Moreover, LAP has good scalability in hardware architecture. It is easy to build an acceleration system with higher throughput by increasing the number of cores. We prototype a 16-core system into Xilinx ZC702 FPGA and a 64-core system into Xilinx ZCU102 FPGA respectively. The prototype system on ZC702 on average achieves 3.5 GB/s throughput on simple data sets, and the prototype system on ZCU102 can obtain higher throughput and compute density values on part of large datasets in ANMLZoo compared with modern in-memory NFA-based solutions.
Lei Gong 0003, Chao Wang 0003, Haojun Xia, Xianglan Chen, Xi Li 0003, Xuehai Zhou
IEEE Trans. Computers3
2022 Evaluation and Optimization on Virtualization Performance Cost under Semantic Gap
abstract
Virtualization is a key enabling technology in modern data centers. While it provides numerous benefits, it also creates new problems. Virtualization requires the hypervisor to treat the virtual machine as a black box, limiting the ability of information exchange between the hypervisor and the virtual machine, bringing a problem known as the semantic gap. Currently, much research on the semantic gap mainly focuses on bridging the semantic gap. The evaluation of the semantic gap, on the other hand, is a neglected but crucial problem, and relevant research is currently lacking. Therefore, this paper proposes a corresponding virtualization performance cost model to better evaluate the semantic gap. Based on this cost model, we summarize solutions that can be used to alleviate the semantic gap. Furthermore, we propose a novel evaluation method for the CPU double scheduling semantic gap. Finally, we propose an effective virtio-balloon based dynamic memory tuning strategy to alleviate the memory semantic gap. The experiments show that for 400.perlbench, our strategy saves 551MB of memory on average during running and reclaims 990MB of memory after running, with a performance cost of only 1.1%. For 429.mcf, our strategy saves 815MB of memory on average during running and reclaims 2130MB of memory after running, although with the performance cost of 27.6%, it prevents performance cliff-like drop caused by memory shortage.
Haojun Xia, Kun Zhang 0016, Bibo Tu
CSCWD2
2021 LAP: A Lightweight Automata Processor for Pattern Matching Tasks
abstract
Growing applications are employing finite automata as their basic computational model. These applications match tens to thousands of patterns on a large amount of data, which brings great challenges to conventional processors. Hardware-based solutions have achieved high throughputs automata processing. However, they are too heavy to be integrated into small chips. Besides, they have to rely on DRAMs or other high capacity memories to store their underlying automata models. We focus on building a more lightweight automata processor, which can store the whole automata model into SRAMs with limited size and run independently. We propose LAP, a lightweight automata processor. Extremely high storage efficiency is achieved in LAP, leveraging a novel automata model (ADFA) and efficient packing algorithms. Besides, we exploit software-hardware co-design to achieve faster processing speed. We observe that ADFA's traversal algorithm is parallelizable. Thus, we propose novel hardware instructions to parallel the additional memory accesses in ADFA model and hide their access overhead. LAP is organized into a four-stage pipeline and prototyped into Xilinx Artix-7 FPGA at 263 MHz frequency. Evaluations show that LAP achieves extremely high storage efficiency, exceeding IBM's RegX and Micron's AP by 8×. Besides, LAP achieves significant improvements in processing speed ranging from 32% to 91% compared with previous lightweight implementations. As a result, a low-power CPU equipped with five LAP cores can achieve 9.5 Gbps processing throughput matching 400 patterns simultaneously.
Haojun Xia, Lei Gong 0003, Chao Wang 0003, Xianglan Chen, Xuehai Zhou
DATE1
2021 HyperKRP: A Kernel Runtime Security Architecture with A Tiny Hypervisor on Commodity Hardware
abstract
The large body of kernel code provides broad attack surfaces to exploitable bugs or misconfigurations. Current mitigations are difficult to be integrated together or have a non-trivial performance or code size impact. Thus, systematical protection for the kernel is of critical importance and is required. In this paper, we propose a kernel runtime security architecture, called HyperKRP, to provide systematical protection for kernel code, critical kernel data, and efficient kernel page tables. We have implemented a fully working prototype for a recent Linux kernel running on the Intel x86 processor. Our prototype is compromised of three protection engines based on a small size hypervisor. The evaluation shows that HyperKRP effectively ensures kernel runtime security with acceptable overhead.
Kunli Lin, Wenqing Liu, Kun Zhang 0016, Haojun Xia, Bibo Tu
GLOBECOM4
2021 η-LSTM: Co-Designing Highly-Efficient Large LSTM Training via Exploiting Memory-Saving and Architectural Design Opportunities
abstract
Recently, the recurrent neural network, or its most popular type—the Long Short Term Memory (LSTM) network— has achieved great success in a broad spectrum of real-world application domains, such as autonomous driving, natural language processing, sentiment analysis, and epidemiology. Due to the complex features of the real-world tasks, current LSTM models become increasingly bigger and more complicated for enhancing the learning ability and prediction accuracy. However, through our in-depth characterization on the state-of-the-art general-purpose deep-learning accelerators, we observe that the LSTM training execution grows inefficient in terms of storage, performance, and energy consumption, under an increasing model size. With further algorithmic and architectural analysis, we identify the root cause for large LSTM training inefficiency: massive intermediate variables. To enable a highly-efficient LSTM training solution for the ever-growing model size, we exploit some unique memory-saving and performance improvement opportunities from the LSTM training procedure, and leverage them to propose the first cross-stack training solution, η-LSTM, for large LSTM models. η-LSTM comprises both software-level and hardware-level innovations that effectively lower the memory footprint upper-bound and excessive data movements during large LSTM training, while also drastically improving training performance and energy efficiency. Experimental results on six real-world large LSTM training benchmarks demonstrate that η-LSTM reduces the required memory footprint by an average of 57.5% (up to 75.8%) and brings down the data movements for weight matrices, activation data, and intermediate variables by 40.9%, 32.9%, and 80.0%, respectively. Furthermore, it outperforms the state-of-the-art GPU implementation for LSTM training by an average of 3.99× (up to 5.73×) on performance and 2.75× (up to 4.25) on energy. We hope this work can shed some light on how to design high logic utilization for future NPUs.
Xingyao Zhang 0002, Haojun Xia, Donglin Zhuang, Xin Fu 0001, Michael B. Taylor, Shuaiwen Song
ISCA2
2021 Shift-BNN: Highly-Efficient Probabilistic Bayesian Neural Network Training via Memory-Friendly Pattern Retrieving
abstract
Bayesian Neural Networks (BNNs) that possess a property of uncertainty estimation have been increasingly adopted in a wide range of safety-critical AI applications which demand reliable and robust decision making, e.g., self-driving, rescue robots, medical image diagnosis. The training procedure of a probabilistic BNN model involves training an ensemble of sampled DNN models, which induces orders of magnitude larger volume of data movement than training a single DNN model. In this paper, we reveal that the root cause for BNN training inefficiency originates from the massive off-chip data transfer by Gaussian Random Variables (GRVs). To tackle this challenge, we propose a novel design that eliminates all the off-chip data transfer by GRVs through the reversed shifting of Linear Feedback Shift Registers (LFSRs) without incurring any training accuracy loss. To efficiently support our LFSR reversion strategy at the hardware level, we explore the design space of the current DNN accelerators and identify the optimal computation mapping scheme to best accommodate our strategy. By leveraging this finding, we design and prototype the first highly efficient BNN training accelerator, named Shift-BNN, that is low-cost and scalable. Extensive evaluation on five representative BNN models demonstrates that Shift-BNN achieves an average of 4.9 × (up to 10.8 ×) boost in energy efficiency and 1.6 × (up to 2.8 ×) speedup over the baseline DNN training accelerator.
Qiyu Wan, Haojun Xia, Xingyao Zhang 0002, Shuaiwen Song, Xin Fu 0001
MICRO2