VLDB 2026 Research / reviewers in the wild / expert
Yuhong Song
dblp:254/2591
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Computational Performance Bounds Prediction in Quantum Computing With Unstable NoiseabstractQuantum computing has significantly advanced in recent years, boasting devices with hundreds of quantum bits (qubits), hinting at its potential quantum advantage over classical computing. Yet, noise in quantum devices poses significant barriers to realizing this supremacy. Understanding noise’s impact is crucial for reproducibility and application reuse; moreover, the next-generation quantum-centric supercomputing essentially requires efficient and accurate noise characterization to support system management (e.g., job scheduling), where ensuring correct functional performance (i.e., fidelity) of jobs on available quantum devices can even be higher-priority than traditional objectives. However, noise fluctuates over time, even on the same quantum device, which makes predicting the computational bounds for on-the-fly noise is vital. Noisy quantum simulation can offer insights but faces efficiency and scalability issues. In this work, we propose a data-driven workflow, namely QuBound, to predict computational performance bounds. It decomposes historical performance traces to isolate noise sources and devises a novel encoder to embed circuit and noise information processed by a Long Short-Term Memory (LSTM) network. For evaluation, we compare QuBound with a state-of-the-art learning-based predictor, which only generates a single performance value instead of a bound. Experimental results show that the result of the existing approach falls outside of performance bounds, while all predictions from our QuBound with the assistance of performance decomposition better fit the bounds. Moreover, QuBound can efficiently produce practical bounds for various circuits with over 106 speedup over simulation; in addition, the range from QuBound is over 10× narrower than the state-of-the-art analytical approach. Jinyang Li 0001, Samudra Dasgupta, Yuhong Song, Lei Yang 0018, Travis S. Humble, Weiwen Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | Minimizing Communications of Quantum Circuit Simulations on Distributed SystemsabstractEfficient full-state quantum circuit simulations are useful tools for the design of quantum algorithms. Multi-node distributed systems are commonly employed as such simulations require a large amount of computation power and memory space. In distributed systems, communication overhead can be the performance bottleneck. This paper presents a distributed simulation framework called QuanTrans. A quantum circuit is composed of many levels of quantum gates. The simulation is conducted level by level. For circuits with particular structures, it employs a hybrid simulation approach to replace intermediate multi-level communications with one level of final merge operation, whose communication volume is comparable to that of one level of simulation in previous work. A circuit without such structures is sliced to find applicable sub-circuits with a single or multiple consecutive level(s). One level of communication is required for each sub-circuit, so we further propose a polynomial-time optimal circuit slicing algorithm. It can transform any circuit such that the number of sliced sub-circuits is the minimum after transformation. Experimental results show that QuanTrans can effectively reduce communication time and simulation time. Longshan Xu, Edwin H.-M. Sha, Yuhong Song, Yunfan Chi, Qingfeng Zhuge |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Ensuring Data Freshness for In-Storage Computing with Cooperative Buffer ManagerabstractIn-storage computing (ISC) aims to mitigate the excessive data movement between the host memory and storage by offloading computation to storage devices for in-situ execution. However, ensuring data freshness remains a key challenge for practical ISC. For performance considerations, many data processing systems implement a buffer manager to cache part of the on-disk data in the host memory. While the host applications commit updates to the in-memory cached copies of the data, ISC operators offloaded to the device only have access to the on-disk persistent data. Thus, ISC may miss the most recent updates from the host and produce incorrect results after reading the stale and inconsistent data from the persistent storage. With this limitation, current ISC can only be used in read-only settings where the on-disk data are not subject to concurrent updates. To tackle this problem, we propose a cooperative buffer manager for ISC to transparently provide data freshness guarantees to host applications. Proposed methods allow the device to synchronize with the host buffer manager and decide whether to read the most recent copy of data from host memory or flash memory. We implement our method based on a real hardware platform and perform evaluation with a Bs--tree based key-value store. Experiments show that our method can provide transparent data freshness for host applications with reduced latency. Jin Xue, Yuhong Song, Zili Shao |
DATE | 2 |
| 2025 | MirrorFS: Cross-Layered File System for SSD-based In-Storage Computing
Jin Xue, Yuhong Song, Zili Shao |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | MuDP: multi-granularity data placement for uniform loops on SPM-DRAM architectures to minimize latency
Edwin H.-M. Sha, Yuhong Song, Yibo Guo, Longshan Xu, Qingfeng Zhuge |
Frontiers Comput. Sci. | 3 |
| 2024 | Mera: Memory Reduction and Acceleration for Quantum Circuit Simulation via Redundancy ExplorationabstractWith the development of quantum computing, quantum processor demonstrates the potential supremacy in specific applications, such as Grover's database search and popular quantum neural networks (QNNs). For better calibrating the quantum algorithms and machines, quantum circuit simulation on classical computers becomes crucial. However, as the number of quantum bits (qubits) increases, the memory requirement grows exponentially. In order to reduce memory usage and accelerate simulation, we propose a multi-level optimization, namely Mera, by exploring memory and computation redundancy. First, for a large number of sparse quantum gates, we propose two compressed structures for low-level full-state simulation. The corresponding gate operations are designed for practical implementations, which are relieved from the longtime compression and decompression. Second, for the dense Hadamard gate, which is definitely used to construct the superposition, we design a customized structure for significant memory saving as a regularity-oriented simulation. Meanwhile, an ondemand amplitude updating process is optimized for execution acceleration. Experiments show that our compressed structures increase the number of qubits from 17 to 35, and achieve up to$6.9 \times$acceleration for QNN. Yuhong Song, Edwin H.-M. Sha, Longshan Xu, Qingfeng Zhuge, Zili Shao |
ICCD | 1 |
| 2023 | Hardware-aware neural architecture search for stochastic computing-based neural networks on tiny devices
Yuhong Song, Edwin H.-M. Sha, Qingfeng Zhuge, Rui Xu 0013, Xiaowei Xu 0004, Bingzhe Li, Lei Yang 0018 |
J. Syst. Archit. | 1 |
| 2023 | Loop interchange and tiling for multi-dimensional loops to minimize write operations on NVMs
Rui Xu 0013, Edwin H.-M. Sha, Qingfeng Zhuge, Yuhong Song, Han Wang 0051 |
J. Syst. Archit. | 4 |
| 2023 | Spatiotemporal knowledge teacher-student reinforcement learning to detect liver tumors without contrast agents
Chenchu Xu, Yuhong Song, Dong Zhang 0009, Leonardo Kayat Bittencourt, Sree Harsha Tirumani, Shuo Li 0001 |
Medical Image Anal. | 2 |
| 2023 | Optimizing Data Placement for Hybrid SRAM+Racetrack Memory SPM in Embedded SystemsabstractNonvolatile memory (NVM) has the potential as the medium for scratchpad memory (SPM) in embedded devices. Racetrack memory (RM), in particular, is a developing memory technology that possesses high density and read latency comparable to SRAM. The RM’s access operations, however, are based on shift operations. Multiple shift operations will lead to long access latency and high energy. In this article, SRAM is borrowed to help the shifts reduction. Thus, a novel hybrid SRAM+RM SPM is presented to make use of SRAM’s random access and RM’s high density. But, there are some challenges to the proposed architecture: 1) the large capacity of SRAM is not available due to its low density and 2) due to the drawbacks of RM mentioned above, data that are randomly accessed are not expected to be stored on RM. Therefore, a data placement scheme and an instruction scheduling strategy are presented for the proposed architecture. First, an access instruction scheduling strategy is introduced to obtain a relatively sequential access sequence to help with the shifts and SRAM size reduction; second, to help with data placement, a metric for representing the data access cost is proposed; third, a data placement strategy based on the metric is proposed; and finally, a solution for decreasing SRAM size is suggested to maximize the capacity of SPM (or minimize the size of SPM). Experiments show that the suggested scheme can significantly improve the performance of the hybrid SPM while also reducing the shifts on RM with minimal SRAM. Rui Xu 0013, Edwin H.-M. Sha, Qingfeng Zhuge, Yuhong Song, Han Wang 0051, Liang Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | BSC: Block-based Stochastic Computing to Enable Accurate and Efficient TinyMLabstractAlong with the progress of AI democratization, machine learning (ML) has been successfully applied to edge applications, such as smart phones and automated driving. Nowadays, more applications require ML on tiny devices with extremely limited resources, like implantable cardioverter de-fibrillator (ICD), which is known as TinyML. Unlike ML on the edge, TinyML with a limited energy supply has higher demands on low-power execution. Stochastic computing (SC) using bitstreams for data representation is promising for TinyML since it can perform the fundamental ML operations using simple logical gates, instead of the complicated binary adder and multiplier. However, SC commonly suffers from low accuracy for ML tasks due to low data precision and inaccuracy of arithmetic units. Increasing the length of the bitstream in the existing works can mitigate the precision issue but incur higher latency. In this work, we propose a novel SC architecture, namely Block-based Stochastic Computing (BSC). BSC divides inputs into blocks, such that the latency can be reduced by exploiting high data parallelism. Moreover, optimized arithmetic units and output revision (OUR) scheme are proposed to improve accuracy. On top of it, a global optimization approach is devised to determine the number of blocks, which can make a better latency-power trade-off. Experimental results show that BSC can outperform the existing designs in achieving over 10% higher accuracy on ML tasks and over$6\times$power reduction. Yuhong Song, Edwin H.-M. Sha, Qingfeng Zhuge, Rui Xu 0013, Yongzhuo Zhang, Bingzhe Li, Lei Yang 0018 |
ASP-DAC | 1 |
| 2022 | Optimal Loop Tiling for Minimizing Write Operations on NVMs with Complete Memory Latency HidingabstractNon-volatile memory (NVM) is expected to be the second level memory (named remote memory) in two-level memory hierarchy in the future. However, NVM has the limited write endurance, thus it is vital to reduce the number of write operations on NVM. Meanwhile, in two-level memory hierarchy, prefetch is widely used for fetching certain data before it is actually required, to hide the remote memory access latency. In general, large-scale nested loop is the performance bottleneck in one program due to the write operations on NVM caused by the first level memory (named local memory) miss and data reuse. Loop tiling is the key technique for grouping iterations so as to reduce the communication with remote memory used in compiler. In this paper, we propose a new loop tiling approach for minimizing the write operations on NVMs and completely hiding the NVM access latency. Specifically, we introduce a series of theorems to help loop tiling. Then, a legal tile shape and an optimal tile size selection strategy is proposed according to data dependency and local memory capacity. Furthermore, we propose a pipeline scheduling policy to completely hide the remote memory latency. Extensive experiments show that the proposed techniques can reduce write operations on NVMs by 95.1% on average, and NVM latency can be completely hidden. Rui Xu 0013, Edwin H.-M. Sha, Qingfeng Zhuge, Yuhong Song, Jingzhi Lin |
ASP-DAC | 4 |
| 2021 | Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile DevicesabstractA pruning-based AutoML framework for run-time reconfigurability, namely RT3, is proposed in this work. This enables Transformer-based large Natural Language Processing (NLP) models to be efficiently executed on resource-constrained mobile devices and reconfigured (i.e., switching models for dynamic hardware conditions) at run-time. Such reconfigurability is the key to save energy for battery-powered mobile devices, which widely use dynamic voltage and frequency scaling (DVFS) technique for hardware reconfiguration to prolong battery life. In this work, we creatively explore a hybrid block-structured pruning (BP) and pattern pruning (PP) for Transformer-based models and first attempt to combine hardware and software reconfiguration to maximally save energy for battery-powered mobile devices. Specifically, RT3integrates two-level optimizations: First, it utilizes an efficient BP as the first-step compression for resource-constrained mobile devices; then, RT3heuristically generates a shrunken search space based on the first level optimization and searches multiple pattern sets with diverse sparsity for PP via reinforcement learning to support lightweight software reconfiguration, which corresponds to available frequency levels of DVFS (i.e., hardware reconfiguration). At run-time, RT3can switch the lightweight pattern sets within 45ms to guarantee the required real-time constraint at different frequency levels. Results further show that RT3can prolong battery life over $ 4\times$ improvement with less than 1% accuracy loss for Transformer and 1.5% score decrease for DistilBERT. Yuhong Song, Weiwen Jiang, Panjie Qi, Qingfeng Zhuge, Edwin H.-M. Sha, Sakyasingha Dasgupta, Yiyu Shi 0001, Caiwen Ding |
DAC | 1 |
| 2021 | Accommodating Transformer onto FPGA: Coupling the Balanced Model Compression and FPGA-Implementation OptimizationabstractRecently, Transformers gradually gain popularity and perform outstanding for many Natural Language Processing (NLP) tasks. However, Transformers suffer from heavy computation and memory footprint, making it difficult to deploy on embedded devices. The field-programmable gate array (FPGA) is widely used to accelerate deep learning algorithms for its advantages. However, the trained Transformer models are too large to accommodate to an FPGA fabric. To accommodate Transformer onto FPGA and achieve efficient execution, we propose an acceleration framework coupling the balanced model compression at the algorithm level and FPGA-implementation optimization at the hardware level. At algorithm level, we adopt a block-balanced pruning and propose an efficient sparse matrix storage format for this pruning technique, named Compressed Block Row (CBR). At the hardware level, we design an accelerator for sparse model. And we also abstract a performance analytic model to evaluate the performance of accelerator. Experiments show that our CBR format perform better than general formats and can significantly save storage space. And our accelerator can achieve $38\times$ and $1.93\times$ speedup compared to other works on CPU and GPU respectively. Panjie Qi, Yuhong Song, Hongwu Peng, Shaoyi Huang, Qingfeng Zhuge, Edwin H.-M. Sha |
ACM Great Lakes Symposium on VLSI | 2 |
| 2021 | Accelerating Framework of Transformer by Hardware Design and Model Compression Co-OptimizationabstractState-of-the-art Transformer-based models, with gigantic parameters, are difficult to be accommodated on resource constrained embedded devices. Moreover, with the development of technology, more and more embedded devices are available to run a Transformer model. For a Transformer model with different constraints (tight or loose), it can be deployed onto devices with different computing power. However, in previous work, designers did not choose the best device among multiple devices. Instead, they just used an existing device to deploy model, which was not necessarily the best fit and may lead to underutilization of resources. To address the deployment challenge of Transformer and the problem to select the best device, we propose an algorithm$\leftrightarrows$hardware closed-loop acceleration framework. Given a dataset, a model, latency constraint LC and accuracy constraint AC, our framework can provide a best device satisfying both constraints. In order to generate a compressed model with high sparsity ratio, we propose a novel pruning technique, hierarchical pruning (HP). We optimize the sparse matrix storage format for HP matrix to further reduce memory usage for FPGA implementation. We design a accelerator that takes advantage of HP to solve the problem of concurrent random access. Experiments on Transformer and TinyBert model show that our framework can find different devices for various LC and AC, covering from low-end devices to high-end devices. Our HP can achieve higher sparsity ratio and is more flexible than other sparsity pattern. Our framework can achieve 37 x, 1.9 x, 1.7x speedup compared to CPU, GPU and FPGA, respectively. Panjie Qi, Edwin H.-M. Sha, Qingfeng Zhuge, Hongwu Peng, Shaoyi Huang, Zhenglun Kong, Yuhong Song |
ICCAD | 7 |