Shihui Song

dblp:310/3843 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Designing Domain-Specific Compilers for Lossy Compression: A Case Study on Wafer-Scale Engine
Shihui Song, Robert Underwood, Sheng Di, Peng Jiang 0004, Franck Cappello
IPDPS1
2026 Near-Zero Cost KV Cache Compression for Large Language Model Inference
Boyuan Zhang 0002, Yafan Huang, Shihui Song, Jinda Jia, Chengming Zhang 0006, Zhi Zhang 0005
IPDPS4
2025 A Memory-Efficient and Computation-Balanced Lossy Compressor on Wafer-Scale Engine
abstract
Cerebras system has demonstrated immense potential across various scientific domains. However, modern scientific simulations frequently generate vast volumes of data in a short time, leading to bottlenecks in runtime performance and memory footprint. While an ultra-fast error-bounded lossy compressor can mitigate such limitations with high compression ratios and guaranteed data quality, deploying it into Cerebras dataflow architecture poses significant difficulties. Specifically, Cerebras faces memory challenges, such as the absence of shared memory and limited local memory, alongside computational challenges, including specialized parallelism and sensitivity to imbalanced workloads. In this work, we propose CERESZII, an error-bounded lossy compressor that computes within Cerebras system. CereSZ-II addresses these challenges with a carefully optimized four-stage compression workflow, consisting of Pre-quantization, Lightweight Prediction, Fixed-size Huffman Encoding, and Spatial-aware Offset Computation, ensuring both memory efficiency and computational balance. Evaluation of several real-world scientific datasets shows that CERESZ-II achieves over 800 GB/s throughput, delivering high compression ratios and reliable reconstructed data quality.
Shihui Song, Robert Underwood, Sheng Di, Yafan Huang, Peng Jiang 0004, Franck Cappello
IPDPS1
2025 What to Support When You're Compressing: The State of Practice Gaps and Opportunities for Scientific Data Compression
abstract
Over the last nearly 20 years, lossy compression has become an essential aspect of HPC applications’ data pipelines, allowing them to overcome limitations in storage capacity and bandwidth and, in some cases, increase computational throughput and capacity. However, with the adoption of lossy compression comes the requirement to assess and control the impact lossy compression has on scientific outcomes. In this work, we take a major step forward in describing the state of practice and by characterizing workloads. We examine applications’ needs and compressors’ capabilities across 9 different supercomputing application domains. We present 24 takeaways that provide best practices for applications, operational impacts for facilities achieving compressed data, and gaps in application needs not addressed by production compressors that point towards opportunities for future compression research.
Franck Cappello, Robert Underwood, Yuri Alexeev, Allison H. Baker, Ebru Bozdag, Martin Burtscher, Kyle Chard, Sheng Di, Kyle Gerard Felker, Paul Christopher O'Grady, Hanqi Guo 0001, Yafan Huang, Peng Jiang 0004, Sian Jin, Petter Johansson, Shaomeng Li, Xin Liang 0001, Erik Lindahl, Peter Lindstrom 0001, Zarija Lukic, Magnus Lundborg, Danylo Lykov, Masaru Nagaso, Kento Sato, Amarjit Singh, Seung Woo Son 0001, Shihui Song, William Tang 0002, Dingwen Tao, Jiannan Tian, Kazutomo Yoshii, Kai Zhao 0008
SC27
2024 CereSZ: Enabling and Scaling Error-bounded Lossy Compression on Cerebras CS-2
abstract
Today's scientific applications running on supercomputers produce large volumes of data, leading to critical data storage and communication challenges. To tackle the challenges, error-bounded lossy compression is commonly adopted since it can reduce data size drastically within a user-defined error threshold. Previous work has shown that compression techniques can significantly reduce the storage and I/O overhead while retaining good data quality. However, the existing compressors are mainly designed for CPU and GPU. As new AI chips are being incorporated into supercomputers and increasingly used for accelerating scientific computing, there is a growing demand for efficient data compression on the new architecture. In this paper, we propose an efficient lossy compressor, CereSZ, based on the Cerebras CS-2 system. The compression algorithm is mapped onto Cerebras using both data parallelism and pipeline parallelism. In order to achieve a balanced workload on each processing unit, we propose an algorithm to evenly distribute the pipeline stages. Our experiments with six scientific datasets demonstrate that CereSZ can achieve a throughput from 227.93 GB/s to 773.8 GB/s, 2.43x to 10.98x faster than existing GPU compressors.
Shihui Song, Yafan Huang, Peng Jiang 0004, Xiaodong Yu 0001, Weijian Zheng, Sheng Di, Qinglei Cao, Yunhe Feng, Franck Cappello
HPDC1
2024 Sequential seeding policy on social influence maximization: a Q-learning-driven discrete differential evolution optimization
Jianxin Tang, Shihui Song, Qian Du 0011, Jitao Qu
J. Supercomput.2
2024 Identifying top-k influential nodes in social networks: a discrete hybrid optimizer by integrating butterfly optimization algorithm with differential evolution
Jianxin Tang, Lihong Han, Shihui Song
J. Supercomput.4
2023 Steering the spread of influence adaptively in social networks via a discrete scheduled particle swarm optimization
Jianxin Tang, Shihui Song, Jimao Lan, Fuqing Zhao
Appl. Intell.2
2022 Rethinking graph data placement for graph neural network training on multiple GPUs
abstract
Graph partitioning is commonly used for dividing graph data for parallel processing. While they achieve good performance for the traditional graph processing algorithms, the existing graph partitioning methods are unsatisfactory for data-parallel GNN training on GPUs. In this work, we rethink the graph data placement problem for large-scale GNN training on multiple GPUs. We find that loading input features is a performance bottleneck for GNN training on large graphs that cannot be stored on GPU. To reduce the data loading overhead, we first propose a performance model of data movement among CPU and GPUs in GNN training. Then, based on the performance model, we provide an efficient algorithm to divide and distribute the graph data onto multiple GPUs so that the data loading time is minimized. For cases where data placement alone cannot achieve good performance, we propose a locality-aware neighbor sampling technique to further reduce the data movement overhead without losing accuracy. Our experiments with graphs of different sizes on different numbers of GPUs show that our techniques not only achieve smaller data loading time but also incur much less preprocessing overhead than the existing graph partitioning methods.
Shihui Song, Peng Jiang 0004
ICS1
2022 Exposing and Exploiting Fine-Grained Block Structures for Fast and Accurate Sparse Training
abstract
Sparse training is a popular technique to reduce the overhead of training large models. Although previous work has shown promising results for nonstructured sparse models, it is still unclear whether a sparse model with structural constraints can be trained from scratch to high accuracy. In this work, we study the dynamic sparse training for a class of sparse models with shuffled block structures. Compared to nonstructured models, such fine-grained structured models are more hardware-friendly and can effectively accelerate the training process. We propose an algorithm that keeps adapting the sparse model while maintaining the active parameters in shuffled blocks. We conduct experiments on a variety of networks and datasets and obtain positive results. In particular, on ImageNet, we achieve dense accuracy for ResNet50 and ResNet18 at 0.5 sparsity. On CIFAR10/100, we show that dense accuracy can be recovered at 0.6 sparsity for various models. At higher sparsity, our algorithm can still match the accuracy of nonstructured sparse training in most cases, while reducing the training time by up to 5x due to the fine-grained block structures in the models.
Peng Jiang 0004, Lihan Hu, Shihui Song
NeurIPS3
2022 Rethinking graph data placement for graph neural network training on multiple GPUs
abstract
The existing Graph Neural Network (GNN) systems adopt graph partitioning to divide the graph data for multi-GPU training. Although they support large graphs, we find that the existing techniques lead to large data loading overhead. In this work, we for the first time model the data movement overhead among CPU and GPUs in GNN training. Based on the performance model, we provide an efficient algorithm to divide and distribute the graph data onto multiple GPUs so that the data loading time is minimized. The experiments show that our technique achieves smaller data loading time compared with the existing graph partitioning methods.
Shihui Song, Peng Jiang 0004
PPoPP1
2022 MATT. A Multiple-instance Attention Mechanism for Long-tail Music Genre Classification
abstract
Imbalanced music genre classification is a crucial task in the Music Information Retrieval (MIR) field for identifying the long-tail, data-poor genre based on the related music audio segments, which is very prevalent in real-world scenarios. Most of the existing models are designed for class-balanced music datasets, resulting in poor performance in accuracy and generalization when identifying the music genres at the tail of the distribution. Inspired by the success of introducing Multi-instance Learning (MIL) in various classification tasks, we propose a novel mechanism named Multi-instance Attention (MATT)1to boost the performance for identifying tail classes. Specifically, we first construct the bag-level datasets by generating the album-artist pair bags. Second, we leverage neural networks to encode the music audio segments. Finally, under the guidance of a multi-instance attention mechanism, the neural network-based models could select the most informative genre to match the given music segment. Comprehensive experimental results on a large-scale music genre benchmark dataset with long-tail distribution demonstrate MATT significantly outperforms other state-of-the-art baselines.1Github: https://github.com/JohannesLiu/Music-Genre-Classification
Shihui Song, Menghua Zhang, Yafan Huang
SMC2
2021 Rumor Detection on Social Media with Out-In-Degree Graph Convolutional Networks
abstract
With the tremendous development in hardware computing and the widespread use of mobile terminal devices, there are increasingly more people who prefer to share their lives and opinions on social media. Though social media plat-forms allow everyone to express their opinions freely, they create convenience for rumor propagation in the meantime, which brings huge negative influence on the public and makes rumor detection extremely necessary. Currently, the most effective methods regard rumor propagation network as a graph and adopt graph convolutional networks (GCN) to detect rumor automatically. Such methods achieve promising performance in rumor detection, however, we argue that they have two critical defects: 1) they neglect the position contributions of rumor nodes in a graph, reducing the accuracy of rumor detection results; 2) they are inadequate in dealing with imbalanced data, which also indicates the inflexibility and the poor generalization ability of the model. To overcome these issues, we incorporate Katz centrality into spectral-domain graph convolution and propose a novel model named Out-In-Degree Graph Convolutional Networks (OID-GCN). Specifically, besides enhancing accuracy, Katz centrality can efficiently capture the position information of nodes, while the rest structure of OID-GCN shows a superb ability in dealing with imbalanced data. Comprehensive experimental results on two real-world datasets Twitter-15 and Twitter-16 demonstrate our OID-GCN outperforms existing methods.
Shihui Song, Yafan Huang, Hongwei Lu
SMC1