VLDB 2026 Research / reviewers in the wild / expert
Haichuan Hu
dblp:333/1845
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rearchitecting Buffered I/O in the Era of High-Bandwidth SSDs
Yekang Zhan, Tianze Wang, Zheng Peng 0017, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
FAST | 4 |
| 2026 | ComPass: Contrastive Learning for Automated Patch Correctness Assessment in Program Repair
Quanjun Zhang, Ye Shang, Haichuan Hu, Chunrong Fang, Zhenyu Chen 0001 |
Empir. Softw. Eng. | 3 |
| 2025 | Rethinking the Request-to-IO Transformation Process of File Systems for Full Utilization of High-Bandwidth SSDs
Yekang Zhan, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
FAST | 2 |
| 2025 | AIS: An Active Idleness I/O Scheduler to Reduce Buffer-Exhausted Degradation of Solid-State DrivesabstractModern solid-state drives (SSDs) continue to boost storage density and I/O bandwidth at the cost of flash-access I/O latency, especially for write, hence they prevalently deploy a build-in buffer to absorb incoming writes. However, when the buffer is used up, the applications suffer from a sudden and long performance decline, i.e., buffer-exhausted degradation (BED). To holistically understand BED and recovery, we design an automated testing toolset (SSDTest) to measure six commodity NVMe SSDs and find: (1) the occurrence of the BED strictly relies on the written-data amount, (2) BED dramatically increases I/O latency of SSDs, especially write and read-after-write, (3) BED can be conditionally reduced and recovered only after a period of idle time, and (4) a read without preceding writes is largely immune to BED, but prolongs the required idle time to recover the available buffer. Furthermore, we build a black-box SSD buffer-recovery model to quantitatively characterize the idleness-recovery behaviors and design an SSD BED predictor to make BED occurrence and buffer recovery predictable. Leveraging this model, we further design an Active Idleness I/O Scheduler (AIS) with small-sized auxiliary storage to actively regulate the I/O idle-intervals to maximize the internal buffer recovery of SSD. AIS adaptively steers incoming data to the auxiliary storage to (1) strategically keep SSD idle to reduce the occurrence of BED and (2) mitigate the tail latency of SSDs caused by read-after-writes during BED. We perform extensive evaluations under a variety of workloads. The results show that AIS improves average, 99th, 99.9th, and 99.99th-percentile latencies of SSDs by up to 29.3%, 37.3%, 78.7%, and 67.2% respectively, with up to 512MB auxiliary storage. Yekang Zhan, Xiangrui Yang 0001, Haichuan Hu, Qiang Cao 0001, Yifan Zhang 0012, Jie Yao 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | Working Smarter Not Harder: Hybrid Cooling for Deep Learning in Edge DatacentersabstractThe proliferation of deep-learning-based mobile and IoT applications has driven the increasing deployment of edge datacenters equipped with domain-specific accelerators. The unprecedented computing power offered by these accelerators puts a heavy burden on the cooling system, motivating more potent cooling techniques like cold water cooling. However, we observe that cold water cooling results in significant energy waste in edge datacenters due to the fluctuating resource utilization both spatially and temporally. To tackle this issue, we propose the concept of “working smarter” by slowing down accelerators deliberately whenever possible and enabling warm water cooling during these times to achieve cooling efficiency. Based on this concept, we develop Hyco—a hybrid water cooling system tailored for edge datacenters running deep learning workloads. First, Hyco features a zone-based cooling architecture enabling dynamic switching between cold water and warm water cooling. Then, based on a lightweight latency estimation method, Hyco incorporates a learning-based scheduling scheme to determine “which” accelerator workers and “when” to slow down through an adaptive and intelligent power-latency trade-off for deep learning models. The simulation with real-world traces shows that Hyco reduces the cooling energy consumption by up to 34.74× while satisfying latency constraints more than 99% of the time for deep-learning-based applications. Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu |
IEEE Trans. Sustain. Comput. | 3 |
| 2024 | RomeFS: A CXL-SSD Aware File System Exploiting Synergy of Memory-Block Dual PathsabstractCompute eXpress Link (CXL) based Solid-State Drives (CXL-SSDs), such as the Samsung CMM-H model, promise to offer CXL.mem memory and CXL.io block dual-mode interfaces. Nonetheless, whether and how cloud applications with diverse and varying access patterns benefit from such dual-mode CXL-SSD remains an open question for academia and industry. Yekang Zhan, Haichuan Hu, Xiangrui Yang 0001, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001 |
SoCC | 2 |
| 2024 | SchInFS: A File System Integrating Functions of the Block I/O Scheduler for ZNS SSDsabstractEmerging Zoned Namespace (ZNS) SSDs divide address space into sequentially written zones and transfer garbage collection (GC) to the host, thereby providing more stable performance, increased capacity, and extended device lifespan. However, the sequential write constraint poses some problems for file system design on ZNS devices, particularly leading to bottlenecks in multi-threaded performance for concurrent write requests. Through comprehensive experiments, we analyze the scalability issues of existing POSIX file systems on NVMe ZNS SSDs and identify the root causes: (1) current file systems generally fail to simultaneously utilize the throughput of multiple zones, and (2) their methods for concurrent writing within a single zone are inefficient. To fully exploit the concurrent performance of ZNS SSDs, we propose SchInFS, a novel multi-head logging ZNS SSD file system that integrates the functions of the block I/O scheduler. Firstly, SchInFS employs a multi-head logging design to leverage the throughput of multiple zones concurrently. Secondly, it provides an independent merge queue for each log, facilitating efficient cross-thread write blocks. Finally, SchInFS uses a Block I/O Submission Controller (BSC) to ensure the timely submission of requests and ordered writing within a single zone. Evaluation on real devices demonstrates the effectiveness of SchInFS, showcasing a substantial improvement in the concurrency performance of ZNS SSDs by up to 71.94% compared to current ZNS SSD file systems. Jintong Zhang, Haichuan Hu, Jianxi Chen, Yekang Zhan |
ICCD | 2 |
| 2024 | λGrapher: A Resource-Efficient Serverless System for GNN Serving through Graph SharingabstractGraph Neural Networks (GNNs) have been increasingly adopted for graph analysis in web applications such as social networks. Yet, efficient GNN serving remains a critical challenge due to high workload fluctuations and intricate GNN operations. Serverless computing, thanks to its flexibility and agility, offers on-demand serving of GNN inference requests. Alas, the request-centric serverless model is still too coarse-grained to avoid resource waste. Haichuan Hu, Fangming Liu, Qiangyu Pei, Yongjie Yuan, Zichen Xu 0001, Lin Wang 0015 |
WWW | 1 |
| 2024 | Exploring nonintrusive measurements of spatio-temporal portrait of microservicesabstractAbstract As cloud native technology advances, the scale and complexity of applications built on microservice architecture continue to expand, leading to increasingly intricate differences between software within the same application. Microservice applications, offering high flexibility, are deployed in data centers as black boxes from the users' perspective, leaving them with no insight into the orchestration of cloud service providers. Consequently, users face challenges in promptly recognizing performance imbalances within their deployed applications. Meanwhile, cloud service providers may cut costs by offering a mix of qualified and unqualified services, potentially deceiving users. To enhance the understanding of microservice application organization, we propose a non‐intrusive measurement framework, termed NMPI. NMPI facilitates rapid identification of microservice application defects, offering insights into cloud services and detecting fraudulent behavior in microservice‐based applications. We model microservice applications using a queue analysis‐based approach and filter the dominant frequency components of average response time signals by employing k‐means on the fast fourier transform (FFT). Our model constructs a library of performance portraits for various software, with these portraits resembling human fingerprints that carry and mark the software's internal information. Utilizing a two‐tier microservices‐based application incorporating a database as a case study allows us to demonstrate the effectiveness of NMPI. Our experimental results show that NMPI can produce differentiable profiles of data service performance portraits across a diverse and extensive range of workloads, enabling the identification of software types and the analysis of performance conditions. Zichen Xu 0001, Dan Wu 0010, Xiaoling Li 0002, Biyong Liu, Haichuan Hu, Shuang Tan, Yusong Tan, Chenren Xu, Christopher Stewart, Qihe Zhou |
Softw. Pract. Exp. | 6 |
| 2024 | Machine Translation Testing via Syntactic Tree PruningabstractMachine translation systems have been widely adopted in our daily life, making life easier and more convenient. Unfortunately, erroneous translations may result in severe consequences, such as financial losses. This requires to improve the accuracy and the reliability of machine translation systems. However, it is challenging to test machine translation systems because of the complexity and intractability of the underlying neural models. To tackle these challenges, we propose a novel metamorphic testing approach by syntactic tree pruning (STP) to validate machine translation systems. Our key insight is that a pruned sentence should have similar crucial semantics compared with the original sentence. Specifically, STP (1) proposes a core semantics-preserving pruning strategy by basic sentence structures and dependency relations on the level of syntactic tree representation, (2) generates source sentence pairs based on the metamorphic relation, and (3) reports suspicious issues whose translations break the consistency property by a bag-of-words model. We further evaluate STP on two state-of-the-art machine translation systems (i.e., Google Translate and Bing Microsoft Translator) with 1,200 source sentences as inputs. The results show that STP accurately finds 5,073 unique erroneous translations in Google Translate and 5,100 unique erroneous translations in Bing Microsoft Translator (400% more than state-of-the-art techniques), with 64.5% and 65.4% precision, respectively. The reported erroneous translations vary in types and more than 90% of them are not found by state-of-the-art techniques. There are 9,393 erroneous translations unique to STP, which is 711.9% more than state-of-the-art techniques. Moreover, STP is quite effective in detecting translation errors for the original sentences with a recall reaching 74.0%, improving state-of-the-art techniques by 55.1% on average. Quanjun Zhang, Juan Zhai, Chunrong Fang, Weisong Sun, Haichuan Hu |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2023 | AsyFunc: A High-Performance and Resource-Efficient Serverless Inference System via Asymmetric FunctionsabstractRecent advances in deep learning (DL) have spawned various intelligent cloud services with well-trained DL models. Nevertheless, it is nontrivial to maintain the desired end-to-end latency under bursty workloads, raising critical challenges on high-performance while resource-efficient inference services. To handle burstiness, some inference services have migrated to the serverless paradigm for its rapid elasticity. However, they neglect the impact of the time-consuming and resource-hungry model-loading process when scaling out function instances, leading to considerable resource inefficiency for maintaining high performance under burstiness. Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Fangming Liu |
SoCC | 3 |
| 2023 | Fast and Scalable Gate-Level Simulation in Massively Parallel SystemsabstractThe natural bijection between a proposed circuit design and its graph representation shall allow any graph optimization algorithm deploying into many-core systems efficiently. However, this process suffers from the exponentially growing overhead and heavy memory footprint with the signal propagation. To conquer the unique challenge, we systematically study the simulation with millions of gates, and identify that the processing complexity could grow exponentially from the signal inputs, the skewness of the computational graph stays. Thus, we present ZhouBi, a fast and scalable gate-level simulation framework to fully exploit the parallelism from many-core systems. ZhouBi contributes in threefolds, (I) a graph representation that colors gate-level netlists and identifies skew partitions based on the graph skewness; (II) A set of heuristic algorithms that picks opportunistic and conservative algorithms to accelerate the simulation; (III) A system facility that supports selective mapping between simulation and many-core, providing a tradeoff between the risk of concurrent simulation fail and performance gain. We have prototyped ZhouBi and evaluated it with practical baselines. ZhouBi can achieve a 27.6× performance gain, as compared to the state-of-the-practice Veriwell without compromising any correctness. Our framework supports large graphs enabling scale-out gate-level simulations for chip design. Haichuan Hu, Zichen Xu 0001, Yuhao Wang 0001, Fangming Liu |
ICCAD | 1 |
| 2023 | CCSMP: an efficient closed contiguous sequential pattern mining algorithm with a pattern relation graph
Haichuan Hu, Ruiqing Xia, Shichao Liu 0002 |
Appl. Intell. | 1 |
| 2022 | HBtree: A Heterogeneous B+tree with Multi-granularity for Hybrid NVM-SSD StorageabstractTraditional index structures build on homogeneous storage consisting of same-granularity blocks and maintain a map between logical offset and storage-blocks. Non-Volatile Memory (NVM) and Solid-State Drive have their own optimal I/O sizes. However, these homogeneous indexes cannot uniformly and efficiently manage hybrid NVM-SSD storage space with heterogenous blocks. This paper proposes a heterogeneous B+tree, referred as to HBtree, to index multi-granularity blocks for hybrid NVM-SSD storage. The indexing node of HBtree is same to B+tree with largest-granularity blocks. However, a part of leaves in the legacy B+tree are extended to index small granularity blocks. Therefore, HBtree has the low tree-height while indexing different granularity blocks simultaneously. We implement HBtree and evaluate it on hybrid NVM-SSD storage. Compared to legacy B+tree, HBtree improves insert and search performance by up to 4.54x and 1.96x respectively and improving space utilization by up to 3.49x. Yekang Zhan, Haichuan Hu, Qiang Cao 0001 |
NAS | 2 |