Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhuobin Huang

dblp:190/4367 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 48% Memory systems · 15% GPUs and heterogeneous computing · 15%
Software engineering, system software, and programming languages
3 papers
Operating systems · 100%

Topics — the 10 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
serverless computing
1.622025
Towards Serialization/Deserialization-free State Transfer in Serverless Workflows · ACM Trans. Comput. Syst. 2025
Serialization/Deserialization-free State Transfer in Serverless Workflows · EuroSys 2024
Operating systems › fault tolerance
checkpoint and rollback
0.912025
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation · SOSP 2025
Operating systems
interprocess communication
0.912025
Towards Serialization/Deserialization-free State Transfer in Serverless Workflows · ACM Trans. Comput. Syst. 2025
Memory systems
cache management
0.912025
CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data Paths · SIGCOMM 2025
Distributed systems
fault tolerance
0.312025
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation · SOSP 2025
Distributed systems › distributed system architecture › distributed operating systems
process migration
0.312025
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation · SOSP 2025
Cloud and datacenter computing › serverless computing
serverless platforms
0.312025
Towards Serialization/Deserialization-free State Transfer in Serverless Workflows · ACM Trans. Comput. Syst. 2025
Hardware accelerators and domain-specific architectures › network accelerator
SmartNIC
0.312025
CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data Paths · SIGCOMM 2025
Interconnection networks and networks-on-chip
remote direct memory access
0.212024
Serialization/Deserialization-free State Transfer in Serverless Workflows · EuroSys 2024
Distributed systems › distributed communication
remote memory access
0.212024
Serialization/Deserialization-free State Transfer in Serverless Workflows · EuroSys 2024

Methods — techniques the papers use, named apart from their topics

RDMA · 3.3validated speculation · 1.7copy-on-write · 1.7language runtime integration · 1.5OS primitive co-design · 1.5proactive rate control · 0.9elastic buffering · 0.9
YearPublicationVenuePosition
2025 CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data Paths
abstract
Efficient Input/Output (I/O) data path between NICs and CPUs/DRAMs is critical for supporting datacenter applications with high-performance network transmission, especially as link speed scales to 100Gbps and beyond. Traditional I/O acceleration strategies, such as Data Direct I/O (DDIO) and Remote Direct Memory Access (RDMA), perform suboptimally due to the inefficient utilization of the Last-Level Cache (LLC). This paper presents CEIO, a novel cache-efficient network I/O architecture that employs proactive rate control and elastic buffering to achieve zero LLC misses in the I/O data path while ensuring the effectiveness of DDIO and RDMA under various network conditions. We have implemented CEIO on commodity SmartNICs and incorporated it into widely-used DPDK and RDMA libraries. Experiments with well-optimized RPC framework and distributed file system under realistic workloads demonstrate that CEIO achieves up to 2.9× higher throughput and 1.9× lower P99.9 latency over prior work.
Bowen Liu 0002, Qijing Li, Zhuobin Huang, Yijun Sun, Wenxue Li 0004, Junxue Zhang 0001, Ping Yin, Kai Chen 0005
SIGCOMM4
2025 PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
abstract
PhoenixOS (PhOS) is the first OS service that can concurrently checkpoint and restore (C/R) GPU processes—a fundamental capability for critical tasks such as fault tolerance, process migration, and fast startup. While concurrent C/R is well-established on CPUs, it poses unique challenges on GPUs due to their lack of essential features for efficiently tracing concurrent memory reads and writes, such as specific hardware capabilities (e.g., dirty bits) and OS-mediated data paths (e.g., copy-on-write).
Xingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao, Rong Chen 0001, Mingcong Han, Jinyu Gu 0001, Haibo Chen 0001
SOSP2
2025 Towards Serialization/Deserialization-free State Transfer in Serverless Workflows
abstract
Serialization and deserialization dominate the state transfer time of serverless workflows, leading to substantial performance penalties when executing various serverless workflow applications. We identify the key reason for serialization and deserialization as a lack of ability to efficiently access the (remote) memory of another function. To this end, we propose RMMap , an OS primitive for remote memory map, which allows a serverless function to directly access the memory of another function, even if it is located remotely. RMMap is the first to completely eliminate serialization and deserialization overhead when transferring states between any pairs of functions in (unmodified) serverless workflows. To make remote memory map efficient and feasible, we co-design it with modern networking (RDMA), OS, language runtime, and serverless platform. Evaluations using real-world serverless workloads show that integrating RMMap with Knative reduces the serverless workflow execution time on Knative by up to 2.6× and improves resource utilizations by 86.3%.
Xingda Wei, Fangming Lu, Zhuobin Huang, Rong Chen 0001, Mingyu Wu 0001, Haibo Chen 0001
ACM Trans. Comput. Syst.3
2024 Serialization/Deserialization-free State Transfer in Serverless Workflows
abstract
Serialization and deserialization play a dominant role in the state transfer time of serverless workflows, leading to substantial performance penalties during workflow execution. We identify the key reason as a lack of ability to efficiently access the (remote) memory of another function. We propose RMMap, an OS primitive for remote memory map. It allows a serverless function to directly access the memory of another function, even if it is located remotely. RMMap is the first to completely eliminates serialization and deserialization when transferring states between any pairs of functions in (unmodified) serverless workflows. To make remote memory map efficient and feasible, we co-design it with fast networking (RDMA), OS, language runtime, and serverless platform. Evaluations using real-world serverless workloads show that integrating RMMap with Knative reduces the serverless workflow execution time on Knative by up to 2.6 × and improves resource utilizations by 86.3%.
Fangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen 0001, Mingyu Wu 0001, Haibo Chen 0001
EuroSys3
2022 Neuro-RDM: An Explainable Neural Network Landscape of Reaction-Diffusion Model for Cognitive Task Recognition
Tingting Dan, Hongmin Cai, Zhuobin Huang, Paul J. Laurienti, Won Hwa Kim, Guorong Wu 0001
MICCAI (8)3
2022 Learning Brain Dynamics of Evolving Manifold Functional MRI Data Using Geometric-Attention Neural Network
abstract
Functional connectivities (FC) of brain network manifest remarkable geometric patterns, which is the gateway to understanding brain dynamics. In this work, we present a novel geometric-attention neural network to characterize the time-evolving brain state change from the functional neuroimages by tracking the trajectory of functional dynamics on high-dimension Riemannian manifold of symmetric positive definite (SPD) matrices. Specifically, we put the spotlight on learning the common state-specific manifold signatures that represent the underlying cognition. In this context, the driving force of our neural network is tied up with the learning of the evolution functionals on the Riemannian manifold of SPD matrix that underlies the known evolving brain states. To do so, we train a convolution neural network (CNN) on the Riemannian manifold of SPD matrices to seek for the putative low-dimension feature representations, followed by an end-to-end recurrent neural network (RNN) to yield the time-varying mapping function of SPD matrices which fits the evolutionary trajectories of the underlying states. Furthermore, we devise a geometric attention mechanism in CNN, allowing us to discover the latent geometric patterns in SPD matrices that are associated with the underlying states. Notably, our work has the potential to understand how brain function emerges behavior by investigating the geometrical patterns from functional brain networks, which is essentially a correlation matrix of neuronal activity signals. Our proposed manifold-based neural network achieves promising results in predicting brain state changes on both simulated data and task functional neuroimaging data from Human Connectome Project, which implies great applicability in neuroscience studies.
Tingting Dan, Zhuobin Huang, Hongmin Cai, Paul J. Laurienti, Guorong Wu 0001
IEEE Trans. Medical Imaging2
2021 Detecting Brain State Changes via Manifold Mean Shifting
abstract
The topology of human functional networks is assumed to oscillate during brain states changes. The functional neuroimage is employed to offer a non-invasive window to understand cognition and behaviors by characterizing the functional connections between spatially distinct brain regions. Consequently, identifying the transitions of functional connectivities is the critical step to understanding the mechanism of cognition that might be underlined with neurological disorders. However, little attention has been paid to studying the geometry of the entire functional brain network. To tackle this issue, this paper models the cognition changes on functional brain networks as a set of landmarks residing on a Riemannian manifold. Accordingly, we propose a Riemannian manifold mean shift method to detect cognition changes by identifying the representative function networks of the distribution of functional networks. The manifold mean shift (MMS) method is applied on both simulated data and real functional neuroimaging data, downloaded from Human Connectome Project (HCP). Experimental results demonstrated the MMS achieved highly accurate and consistent cognition change, by comparing three state-of-the-art methods.
Zhuobin Huang, Tingting Dan, Jiazhou Chen 0001, Hongmin Cai, Guorong Wu 0001
BIBM1
2021 NeuralMon: Graph Neural Network for Flow Measurement Allocation
abstract
Fine-grained and accurate network flow measurements are essential for various network management tasks. In recent years, the evolution of programmable networks enables flow measurement on the switch. However, limited hardware resources on programmable switches drive the shift of measurement from a single switch to network-wide coordinations. This paper aims to optimize the allocation strategy of flow measurement among switches under the objective of measurement coverage and accuracy in network-wide measurement scenarios. We design a Graph Neural Network model, NeuralMon, that can model and solve the above problem precisely. NeuralMon converts network topologies and network flows into a hypergraph and transforms the flow measurement task allocation problem into a node classification problem. NeuralMon is effective in learning the task allocation solution from the network topologies and flows directly. Even on untrained real-world network topologies, NeuralMon still provides excellent performance.
Yang Wang 0053, Xiong Wang 0001, Zhuobin Huang, Ci He, Shizhong Xu
GLOBECOM3
2021 Detecting Brain State Changes by Geometric Deep Learning of Functional Dynamics on Riemannian Manifold
Zhuobin Huang, Hongmin Cai, Tingting Dan, Paul J. Laurienti, Guorong Wu 0001
MICCAI (7)1
2021 Fusion of multi-source retinal fundus images via automatic registration for clinical diagnosis
Tingting Dan, Yu Hu 0004, Chu Han, Zhihao Fan, Zhuobin Huang, Bin Zhang 0050, Guihua Tao, Baoyi Liu, Honghua Yu, Hongmin Cai
Neurocomputing5