Hang Huang

dblp:05/125 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 AEP: Achieving Hierarchical Fault Tolerance in DSM Through Atomic Execution Protection
abstract
Recent advances in CXL 3.0 have revitalized Distributed Shared Memory (DSM), enabling zero-copy sharing across machine nodes at near-NUMA latency and making shared-memory communication practical for distributed applications. However, DSM memory allocators remain vulnerable to partial failures: client crashes during critical, non-atomic memory operations can corrupt metadata and stall the entire system. Existing solutions either block all clients during recovery or impose heavy runtime overhead due to distributed fault-tolerance protocols. We propose Atomic Execution Protection (AEP), a kernel-assisted fault tolerance mechanism that guarantees atomic completion of user-space critical operations by deferring termination until protected regions finish. AEP tolerates intra-node partial failures without costly distributed coordination. To extend beyond a single node, we further design AEP-DSM, a hierarchical DSM allocator that combines AEP-protected node-level allocators with cross-node coordination. Our Linux AEP prototype integrated with a cross-node DSM consistency protocol (CXL-SHM) achieves near-native local performance and improves throughput by up to 13.3x over state-of-the-art DSM allocators.
Zixuan Wang 0030, Hang Huang, Jia Rao, Hui Lu 0001, Hao Fan 0006, Song Wu 0001, Hai Jin 0001
EuroSys3
2025 EPPDL: An efficient privacy-preserving distributed ledger for digital asset transfer in Web3.0
Lichuan Ma, Hang Huang, Youyang Qu
Future Gener. Comput. Syst.3
2024 vKernel: Enhancing Container Isolation via Private Code and Data
abstract
Container technology is increasingly adopted in cloud environments. However, the lack of isolation in the shared kernel becomes a significant barrier to the wide adoption of containers. The challenges lie in how to simultaneously attain high performance and isolation. On the one hand, kernel-level isolation mechanisms, such asseccomp,capabilities, andapparmor, achieve good performance without much overhead, but lack the support for per-container customization. On the other hand, user-level and VM-based isolation offer superior security guarantees and allow for customization since a container is assigned a dedicated kernel, however, at the cost of high overhead. We presentvKernel, a kernel isolation framework. It maintains a minimal set of code and data that are either sensitive or are prone to interference in a virtual kernel instance (vKI). vKernel relies on inline hooks to intercept and redirect requests sent to the host kernel to a vKI, where container-specific security rules, functions, and data are implemented. Through case studies, we demonstrate that under vKernel user-defined data isolation and kernel customization can be supported with a reasonable engineering effort. An evaluation of vKernel with micro-benchmarks, cloud services, real-world applications show that vKernel achieves good security guarantees, but with much less overhead.
Hang Huang, Jia Rao, Song Wu 0001, Hao Fan 0006, Chen Yu 0003, Hai Jin 0001, Kun Suo, Lisong Pan
IEEE Trans. Computers1
2023 PVM: Efficient Shadow Paging for Deploying Secure Containers in Cloud-native Environment
abstract
In cloud-native environments, containers are often deployed within lightweight virtual machines (VMs) to ensure strong security isolation and privacy protection. With the growing demand for customized cloud services, third-party vendors are turning to infrastructure-as-a-service (IaaS) cloud providers to build their own cloud-native platforms, necessitating the need to run a VM or a guest that hosts containers inside another VM instance leased from an IaaS cloud. State-of-the-art nested virtualization in the x86 architecture relies heavily on the host hypervisor to expose hardware virtualization support to the guest hypervisor, not only complicating cloud management but also raising concerns about an increased attack surface at the host hypervisor.
Hang Huang, Jiangshan Lai, Jia Rao, Hui Lu 0001, Wenlong Hou, Zhengyu He, Weidong Han 0003, Tao Ma 0006, Song Wu 0001
SOSP1
2023 Characterizing and optimizing Kernel resource isolation for containers
Kun Wang 0005, Song Wu 0001, Kun Suo, Hang Huang, Hai Jin 0001
Future Gener. Comput. Syst.5
2023 Adapt Burstable Containers to Variable CPU Resources
abstract
In the age of the cloud-native, container technology, referred as OS-level virtualization, is increasingly adopted to deploy cloud applications. Compared with virtual machines, containers are lightweight and flexible in resource management. An important quality-of-service (QoS) class in container management is burstable container, whose resource limits are higher than the actual requests allowing a container to expand whenever demands ramp up and additional resources become available. However, efficiently managing burstable containers is challenging, especially for CPU resources. On the one hand, burstable containers should maintain sufficient concurrency, in the form of threads, to utilize extendable CPU resources. On the other hand, the degree of concurrency necessary for utilizing peak CPU resources leads to suboptimal performance when a container's CPU allocation is constrained. In this paper, we recommend that the number of threads in burstable containers should always be set to the CPU limit to guarantee extensibility. However, modern operating systems (OSes) fall short of efficiently managing thread oversubscription. First, the OS CPU scheduler is inefficient for scheduling excessive threads and lacks container awareness. Second, the existing blocking synchronization supported by the OS kernel is inefficient in handling the sleep and wakeup of excessive threads. Finally, the non-blocking synchronization may waste CPUs performing busy waiting when more than one thread in the run queue. To this end, we present a user-level adaptive container scheduler and two OS mechanisms,virtual blockingandbusy-waiting detection, to avoid inefficiency in managing burstable containers without requiring program code changes. Experimental results show that our approaches can keep burstable containers efficient while allowing the applications in containers to take advantage of additional CPUs. The performance gain under high system load is up to 29.7×.
Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Duoqiang Wang, Kun Suo, Lisong Pan
IEEE Trans. Computers1
2021 Towards Exploiting CPU Elasticity via Efficient Thread Oversubscription
abstract
Elasticity is an essential feature of cloud computing, which allows users to dynamically add or remove resources in response to workload changes. However, building applications that truly exploit elasticity is non-trivial. Traditional applications need to be modified to efficiently utilize variable resources. This paper explores thread oversubscription, i.e., provisioning more threads than the available cores, to exploit CPU elasticity in the cloud. While maintaining sufficient concurrency allows applications to utilize additional CPUs when more are made available, it is widely believed that thread oversubscription introduces prohibitive overheads due to excessive context switches, loss of locality, and contention on shared resources.
Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Hong Jiang 0001, Hao Che, Xiaofeng Wu 0002
HPDC1
2021 SwitchFlow: preemptive multitasking for deep learning
abstract
Accelerators, such as GPU, are a scarce resource in deep learning (DL). Effectively and efficiently sharing GPU leads to improved hardware utilization as well as user experiences, who may need to wait for hours to access GPU before a long training job is done. Spatial and temporal multitasking on GPU have been studied in the literature, but popular deep learning frameworks, such as Tensor-Flow and PyTorch, lack the support of GPU sharing among multiple DL models, which are typically represented as computation graphs, heavily optimized by underlying DL libraries, and run on a complex pipeline spanning CPU and GPU. Our study shows that GPU kernels, spawned from computation graphs, can barely execute simultaneously on a single GPU and time slicing may lead to low GPU utilization.
Xiaofeng Wu 0002, Jia Rao, Wei Chen 0038, Hang Huang, Chris Ding, Heng Huang 0001
Middleware4
2019 Adaptive Resource Views for Containers
abstract
As OS-level virtualization advances, containers have become a viable alternative to virtual machines in deploying applications in the cloud. Unlike virtual machines, which allow guest OSes to run atop virtual hardware, containers have direct access to physical hardware and share one OS kernel. While the absence of virtual hardware abstractions eliminates most virtualization overhead, it presents unique challenges for containerized applications to efficiently utilize the underlying hardware. The lack of hardware abstraction exposes the total amount of resources that are shared among all containers to each individual container. Parallel runtimes (e.g., OpenMP) and managed programming languages (e.g., Java) that rely on OS-exported information for resource management could suffer from suboptimal performance. In this paper, we develop a per-container view of resources to export information on the actual resource allocation to containerized applications. The central design of the resource view is a per-container sys\_namespace that calculates the effective capacity of CPU and memory in the presence of resource sharing among containers. We further create a virtual sysfs to seamlessly interface user space applications with sys\_namespace. We use two case studies to demonstrate how to leverage the continuously updated resource view to enable elasticity in the HotSpot JVM and OpenMP. Experimental results show that an accurate view of resource allocation leads to more appropriate configurations and improved performance in a variety of containerized applications.
Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Kun Suo, Xiaofeng Wu 0002
HPDC1
2018 Dynamic vertical memory scalability for OpenJDK cloud applications
abstract
The cloud is an increasingly popular platform to deploy applications as it lets cloud users to provide resources to their applications as needed. Furthermore, cloud providers are now starting to offer a "pay-as-you-use" model in which users are only charged for the resources that are really used instead of paying for a statically sized instance. This new model allows cloud users to save money, and cloud providers to better utilize their hardware.
Rodrigo Bruno, Paulo Ferreira 0001, Ruslan Synytsky, Tetiana Fydorenchyk, Jia Rao, Hang Huang, Song Wu 0001
ISMM6
2016 A Performance Study of Containers in Cloud Environment
Bowen Ruan, Hang Huang, Song Wu 0001, Hai Jin 0001
APSCC2
2002 XIDEN: Crosstalk Target Identification Framework
abstract
An efficient crosstalk target identification framework called XIDEN has been developed that is used prior to the computationally expensive processes of crosstalk validation and test generation. XIDEN is mainly composed of a set of extractors and filters that together identify the prime crosstalk targets. These prime targets include all error producing targets, i.e. targets that can potentially create Boolean errors. A methodology has been developed to determine the sequence of extractors and filters to identify a small set of targets with low computational cost. The effects of process variation and extraction accuracy as well as complexity are considered in XIDEN. After performing a training process on sample circuits, XIDEN produces a set of effective extractor filter sequences to be used for production circuits.
Shahin Nazarian, Hang Huang, Suriyaprakash Natarajan, Sandeep Gupta 0001, Melvin A. Breuer
ITC2