VLDB 2026 Research / reviewers in the wild / expert
Haitao Song 0001
dblp:45/6289-1
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0003-2113-2224ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 61% Hardware accelerators and domain-specific architectures · 30% Memory systems · 9% | |
| Software engineering, system software, and programming languages
1 paper |
Runtime systems and virtual machines · 100% | |
| Network and information security
1 paper |
Hardware security and side channels · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU cache |
1.0 | 1 | 2026 | Unified and Near-optimal Multi-GPU Cache for Embedding-based Deep Learning · ACM Trans. Comput. Syst. 2026 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.0 | 1 | 2026 | Unified and Near-optimal Multi-GPU Cache for Embedding-based Deep Learning · ACM Trans. Comput. Syst. 2026 |
GPUs and heterogeneous computing
multi-GPU computing |
1.0 | 1 | 2026 | Unified and Near-optimal Multi-GPU Cache for Embedding-based Deep Learning · ACM Trans. Comput. Syst. 2026 |
Runtime systems and virtual machines
garbage collection |
0.8 | 1 | 2024 | Toward an SGX-Friendly Java Runtime · IEEE Trans. Computers 2024 |
Runtime systems and virtual machines › virtual machine implementation
java virtual machine |
0.8 | 1 | 2024 | Toward an SGX-Friendly Java Runtime · IEEE Trans. Computers 2024 |
Memory systems
cache management |
0.3 | 1 | 2026 | Unified and Near-optimal Multi-GPU Cache for Embedding-based Deep Learning · ACM Trans. Comput. Syst. 2026 |
Hardware security and side channels › trusted execution environments
Intel SGX |
0.2 | 1 | 2024 | Toward an SGX-Friendly Java Runtime · IEEE Trans. Computers 2024 |
Hardware security and side channels
trusted execution environments |
0.2 | 1 | 2024 | Toward an SGX-Friendly Java Runtime · IEEE Trans. Computers 2024 |
Methods — techniques the papers use, named apart from their topics
address-conscious launching · 1.5SGX-aware heap layout · 1.5factored extraction · 1.0cache policy optimization · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unified and Near-optimal Multi-GPU Cache for Embedding-based Deep LearningabstractThis article presents UGache , a unified multi-GPU cache system designed for embedding-based deep learning (EmbDL). UGache is primarily motivated by the unique characteristics of EmbDL applications, namely read-only and skewed embedding accesses with affinity and predictability. UGache introduces a novel factored extraction mechanism that avoids bandwidth congestion to fully exploit high-speed cross-GPU interconnects (e.g., NVLink and NVSwitch). Based on a hotness metric, UGache also provides a near-optimal cache policy that balances local and remote access to minimize the extraction time for diverse GPU interconnect topologies. We have implemented UGache and integrated it into two representative frameworks, TensorFlow and PyTorch. Evaluation using two typical types of EmbDL applications, namely graph neural network (GNN) training and deep learning recommendation (DLR) inference, shows that UGache outperforms state-of-the-art replication and partition designs by an average of 1.93× and 1.63× (up to 5.25× and 3.45×), respectively. Furthermore, we demonstrate the applicability of UGache ’s principle beyond embedding-based deep learning, with an example of text-to-image generation on an inference cluster. Xiaoniu Song, Rong Chen 0001, Haitao Song 0001, Haibo Chen 0001 |
ACM Trans. Comput. Syst. | 3 |
| 2024 | On-demand and Parallel Checkpoint/Restore for GPU ApplicationsabstractLeveraging serverless computing for cloud-based machine learning services is on the rise, promising cost-efficiency and flexibility are crucial for ML applications relying on high-performance GPUs and substantial memory. However, despite modern serverless platforms handling diverse devices like GPUs seamlessly on a pay-as-you-go basis, a longstanding challenge remains: startup latency, a well-studied issue when serverless is CPU-centric. For example, initializing GPU apps with minor GPU models, like MobileNet, demands several seconds. For more intricate models such as GPT-2, startup latency can escalate to around 10 seconds, vastly overshadowing the short computation time for GPU-based inference. Prior solutions tailored for CPU serverless setups, like fork() and Checkpoint/Restore, cannot be directly and effectively applied due to differences between CPUs and GPUs. Yanning Yang, Dong Du 0003, Haitao Song 0001, Yubin Xia |
SoCC | 3 |
| 2024 | Toward an SGX-Friendly Java RuntimeabstractHardware enclaves assist in constructing a trusted execution environment (TEE) to store private code and data and thus become an appealing solution to enhance applications’ security. Nevertheless, state-of-the-art enclave implementations like Intel Software Guard Extensions (SGX) have severe performance issues and hinder the deployment of more complicated applications, especially those written in high-level languages like Java. To reduce the performance overhead, prior work has partitioned applications or rebuilt lightweight language runtimes, but they either require manual labor from developers or fail to provide full-fledged support for existing applications. This work instead providesSAJ, a runtime built upon a full-fledged Java virtual machine (JVM) and thus requires no modifications to applications.SAJfirst analyzes the performance of vanilla JVMs running in enclaves and finds that the memory management overhead and boot phase are culprits for performance slowdown. For memory management,SAJintroduces SGX-aware heap layout and garbage collector, which reduces both GC and application execution time. As for the boot phase,SAJintroduces an address-conscious launching mechanism to improve the boot performance. The evaluation under representative Java applications shows thatSAJcan reduce the overall GC pause time, application time, and boot time by 2.93$\boldsymbol{\times}$, 2.58$\boldsymbol{\times}$, and 2.73$\boldsymbol{\times}$on average, respectively. Mingyu Wu 0001, Zhe Li 0037, Haibo Chen 0001, Binyu Zang, Sanhong Li, Haitao Song 0001 |
IEEE Trans. Computers | 8 |
| 2023 | FSAD-Net: Feedback Spatial Attention Dehazing NetworkabstractRecent dehazing networks learn more discriminative high-level features by designing deeper networks or introducing complicated structures, while ignoring inherent feature correlations in intermediate layers. In this article, we establish a novel and effective end-to-end dehazing method, named feedback spatial attention dehazing network (FSAD-Net). FSAD-Net is based on the recurrent structure and consists of four modules: a shallow feature extraction block (SFEB), a feedback block (FB), multiple advanced residual blocks (ARBs), and a reconstruction block (RB). FB is designed to handle feedback connections, and it can improve the dehazing performance by exploiting the dependencies of deep features across stages. ARB implements a novel attention-based estimation on a residual block to adapt to pixels with different distributions. Finally, RB helps restore haze-free images. It can be seen from the experimental results that FSAD-Net almost outperforms the state-of-the-arts in terms of five quantitative metrics. Moreover, the qualitatively comparisons on real-world images also demonstrate the superiority of the proposed FSAD-Net. Considering the efficiency and effectiveness of FSAD-Net, it can be expected to serve as a suitable image dehazing baseline in the future. Yu Zhou 0066, Ping Li 0016, Haitao Song 0001, C. L. Philip Chen, Bin Sheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | NHBS-Net: A Feature Fusion Attention Network for Ultrasound Neonatal Hip Bone SegmentationabstractUltrasound is a widely used technology for diagnosing developmental dysplasia of the hip (DDH) because it does not use radiation. Due to its low cost and convenience, 2-D ultrasound is still the most common examination in DDH diagnosis. In clinical usage, the complexity of both ultrasound image standardization and measurement leads to a high error rate for sonographers. The automatic segmentation results of key structures in the hip joint can be used to develop a standard plane detection method that helps sonographers decrease the error rate. However, current automatic segmentation methods still face challenges in robustness and accuracy. Thus, we propose a neonatal hip bone segmentation network (NHBS-Net) for the first time for the segmentation of seven key structures. We design three improvements, an enhanced dual attention module, a two-class feature fusion module, and a coordinate convolution output head, to help segment different structures. Compared with current state-of-the-art networks, NHBS-Net gains outstanding performance accuracy and generalizability, as shown in the experiments. Additionally, image standardization is a common need in ultrasonography. The ability of segmentation-based standard plane detection is tested on a 50-image standard dataset. The experiments show that our method can help healthcare workers decrease their error rate from 6%-10% to 2%. In addition, the segmentation performance in another ultrasound dataset (fetal heart) demonstrates the ability of our network. Ruhan Liu, Mengyao Liu 0004, Bin Sheng 0001, Huating Li, Ping Li 0016, Haitao Song 0001, Ping Zhang 0016, Lixin Jiang, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |