VLDB 2026 Research / reviewers in the wild / expert
Shui Jiang
dblp:72/2238
· DBLP profile ↗
16ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 12 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TIMBER: A Fast Algorithm for Timing and Power Optimization using Multi-bit Flip-flopsabstractMulti-bit flip-flop (MBFF) banking and debanking is a widely adopted technique for optimizing power and total negative slack (TNS) during the post-placement stage of digital design. While banking flipflops can reduce both power and area, excessive banking may lead to increased TNS due to significant register displacement, as well as bin density violations (BDVs) caused by over-placing MBFFs in legalized regions. To address these challenges, the EDA community recently organized a CAD Contest seeking innovative solutions from both academia and industry. In response, we present TIMBER, a fast and effective optimization algorithm that balances competing objectives in MBFF placement. Unlike existing methods, TIMBER employs a bin-density-aware placement strategy that simultaneously minimizes BDVs and TNS, while also achieving gains in power and area efficiency. To further enhance the runtime performance, TIMBER incorporates a parallelization strategy. Experimental results on the official 2024 CAD Contest benchmarks demonstrate that TIMBER outperforms the first-place winner, delivering on average $13.08 \times$ better solution quality, zero BDVs, $5.06 \times$ faster single-threaded runtime, $3.56 \times$ lower memory usage and up to $72.49 \times$ speedup in multi-threaded execution. Aditya Das Sarma, Shui Jiang, Wan-Luan Lee, Tsung-Yi Ho, Tsung-Wei Huang |
ASP-DAC | 2 |
| 2026 | G-kway: Multilevel GPU-Accelerated k-way Graph Partitioner using Task Graph ParallelismabstractGraph partitioning is important for the design of many CAD algorithms. However, as the graph size continues to grow, graph partitioning becomes increasingly time-consuming. Recent research has introduced parallel graph partitioners using either multi-core CPUs or GPUs. However, the speedup of existing CPU graph partitioners is typically limited to a few cores, while the performance of GPU-based solutions is algorithmically limited by available GPU memory. To overcome these challenges, we propose G-kway, an efficient multilevel GPU-accelerated k -way graph partitioner. G-kway introduces an effective union find-based coarsening and a novel independent set-based refinement algorithm to significantly accelerate both the coarsening and uncoarsening stages. Furthermore, when kernel launch overhead becomes substantial in the refinement algorithm, G-kway employs CUDA Graph-based uncoarsening to reduce the overhead and improve performance. Experimental results have shown that G-kway outperforms both the state-of-the-art CPU-based and GPU-based parallel partitioners with an average speedup of 8.6× and 3.8×, respectively, while achieving comparable partitioning quality. Additionally, G-kway with CUDA Graph-based uncoarsening can further accelerate graph partitioning, achieving up to 1.93× speedup over the default G-kway. Wan-Luan Lee, Dian-Lun Lin, Shui Jiang, Cheng-Hsiang Chiu, Yibo Lin, Bei Yu 0001, Tsung-Yi Ho, Tsung-Wei Huang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | BQSim: GPU-accelerated Batch Quantum Circuit Simulation using Decision DiagramabstractQuantum circuit simulation (QCS) plays an important role in the designs and analysis of a quantum algorithm, as it assists researchers in understanding how quantum operations work without accessing expensive quantum computers. Despite many QCS methods, they are largely limited to simulating one input at a time. However, many simulation-driven quantum computing applications, such as testing and verification, require simulating multiple inputs to reason a quantum algorithm under different scenarios. We refer to this type of QCS as batch quantum circuit simulation (BQCS). In this paper, we present BQSim, a GPU-accelerated batch quantum circuit simulator. BQSim is inspired by the state-of-the-art decision diagram (DD) that can compactly represent quantum gate matrices, but overcomes its limitation of CPU-centric simulation. Specifically, BQSim uses DD to optimize a quantum circuit for reduced BQCS computation and converts DD to a GPU-efficient data structure. Additionally, BQSim employs a task graph-based execution strategy to minimize repetitive kernel call overhead and efficiently overlap kernel execution with data movement. Compared with three state-of-the-art quantum circuit simulators, cuQuantum, Qiskit Aer, and FlatDD, BQSim is 3.25×, 159.06×, and 311.42× faster on average. Shui Jiang, Yi-Hua Chung, Chih-Chun Chang, Tsung-Yi Ho, Tsung-Wei Huang |
ASPLOS (2) | 1 |
| 2025 | iG-kway: Incremental k-way Graph Partitioning on GPUabstractRecent advances in GPU-accelerated graph partitioning have achieved significant performance gains but remain limited to full graph partitioning, lacking support for incremental updates. This limitation is critical in CAD applications, where circuit graphs undergo iterative, incremental modifications during optimization. We present iG-kway, the first GPU-based incremental k-way graph partitioner. iG-kway features an incrementality-aware data structure and a refinement kernel that efficiently updates only affected vertices with minimal quality loss. Experiments show that iG-kway delivers up to $84 \times$ speedup over the state-of-the-art G-kway with comparable partitioning quality. Wan-Luan Lee, Shui Jiang, Dian-Lun Lin, Che Chang, Boyang Zhang 0007, Yi-Hua Chung, Ulf Schlichtmann, Tsung-Yi Ho, Tsung-Wei Huang |
DAC | 2 |
| 2025 | SimPart: A Simple Yet Effective Replication-Aided Partitioning Algorithm for Logic Simulation on GPU
Yi-Hua Chung, Shui Jiang, Wan-Luan Lee, Yanqing Zhang 0002, Haoxing Ren, Tsung-Yi Ho, Tsung-Wei Huang |
Euro-Par (3) | 2 |
| 2024 | G-kway: Multilevel GPU-Accelerated k-way Graph PartitionerabstractGraph partitioning is important for the design of many CAD algorithms. However, as the graph size continues to grow, graph partitioning becomes increasingly time-consuming. To overcome these challenges, we propose G-kway, an efficient multilevel GPU-accelerated k-way graph partitioner. G-kway introduces an effective union find-based coarsening and a novel independent set-based refinement algorithm to significantly accelerate both the coarsening and uncoarsening stages. Experimental results have shown that G-kway outperforms both the state-of-the-art CPU-based and GPU-based parallel partitioners with an average speedup of 8.6× and 3.8×, respectively, while achieving comparable partitioning quality. Wan-Luan Lee, Dian-Lun Lin, Tsung-Wei Huang, Shui Jiang, Tsung-Yi Ho, Yibo Lin, Bei Yu 0001 |
DAC | 4 |
| 2024 | FlatDD: A High-Performance Quantum Circuit Simulator using Decision Diagram and Flat ArrayabstractQuantum circuit simulator (QCS) is essential for designing quantum algorithms because it assists researchers in understanding how quantum operations work without access to expensive quantum computers. Traditional array-based QCSs suffer from exponential time and memory complexities. To address this problem, Decision Diagram (DD) was introduced to compress simulation data by exploring the circuit regularity. However, for irregular circuit structures, DD-based simulation incurs significant runtime and memory overhead. To overcome this challenge, we present FlatDD, a high-performance QCS that capitalizes on the strength of both DD- and array-based approaches. FlatDD parallelizes the simulation workload at multiple levels and leverages an efficient caching technique to reuse historical results. To further enhance the simulation performance for deep circuits, FlatDD introduces a gate-fusion algorithm to reduce the computational cost. Compared to state-of-the-art QCSs on commonly used quantum circuits, FlatDD achieves 34.81× speed-up and 1.93× memory reduction. Shui Jiang, Rongliang Fu, Lukas Burgholzer, Robert Wille, Tsung-Yi Ho, Tsung-Wei Huang |
ICPP | 1 |
| 2024 | Fed-MPS: Federated learning with local differential privacy using model parameter selection for resource-constrained CPS
Shui Jiang, Xiaoding Wang 0001, Youxiong Que, Hui Lin 0007 |
J. Syst. Archit. | 1 |
| 2024 | CapSpeaker: Injecting Commands to Voice Assistants Via CapacitorsabstractRecent studies have exposed that voice assistants can be manipulated by various voice commands without being noticed, however, existing attacks require a nearby speaker to play the attack commands. In this paper, we demonstrate that even without a speaker, we can use capacitors inside electronic devices to produce malicious voice commands, i.e., we convert capacitors into speakers and call itCapSpeaker. The underlying principle ofCapSpeakeris the inverse piezoelectric effect, i.e., varying the voltage across a capacitor to make it vibrate and thus emit acoustic noises. Forcing capacitors to emit target voice commands is challenging because (1) capacitors' response frequency is out of the range of audible voices. (2) We can not directly control the voltage across capacitors to manipulate their emit sounds. To overcome these challenges, we propose a PWM-based modulation scheme to embed the malicious audio onto a high-frequency carrier, e.g., above 20 kHz, and we create malware to induce the designed voltage across the capacitors such thatCapSpeakerplays the chosen malicious commands. Our evaluation of 7 commercial devices demonstrates thatCapSpeakeris feasible to inject voice commands, e.g., ”open the door”, at a distance of up to 10.5 cm. Xiaoyu Ji 0001, Juchuan Zhang, Yancheng Jiang, Shui Jiang, Wenyuan Xu 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | Scalable Scan-Chain-Based Extraction of Neural Network ModelsabstractScan chains have greatly improved hardware testability while introducing security breaches for confidential data. Scan-chain attacks have extended their scope from cryptoprocessors to AI edge devices. The recently proposed scan-chain-based neural network (NN) model extraction attack (lCCAD 2021) made it possible to achieve fine-grained extraction and is multiple orders of magnitude more efficient both in queries and accuracy than its coarse-grained mathematical counterparts. However, both query formulation complexity and constraint solver failures increase drastically with network depth/size. We demonstrate a more powerful adversary, who is capable of improving scalability while maintaining accuracy, by relaxing high-fidelity constraints to formulate an approximate-fidelity-based layer-constrained least-squares extraction using random queries. We conduct our extraction attack on neural network inference topologies of different depths and sizes, targeting the MNIST digit recognition task. The results show that our method outperforms the scan-chain attack proposed in ICCAD 2021 by an average increase in the extracted neural network's functional accuracy of ≈ 32% and 2–3 orders of reduction in queries. Furthermore, we demonstrated that our attack is highly effective even in the presence of countermeasures against adversarial samples. Shui Jiang, Seetal Potluri, Tsung-Yi Ho |
DATE | 1 |
| 2023 | SNICIT: Accelerating Sparse Neural Network Inference via Compression at Inference Time on GPUabstractSparse deep neural network (DNN) has become an important technique for reducing the inference cost of large DNNs. However, computing large sparse DNNs is very challenging because inference iterations can incur highly irregular patterns and unbalanced loads. To address this challenge, the recent HPEC Graph Challenge seeks novel high-performance inference methods for large sparse DNNs. Despite the rapid progress over the past four years, solutions have largely focused on static model compression or sparse multiplication kernels, while ignoring dynamic data compression at inference time which can achieve significant yet untapped performance benefits. Consequently, we propose SNICIT, a new GPU algorithm to accelerate large sparse DNN inference via compression at inference time. SNICIT leverages data clustering to transform intermediate results into a sparser representation that largely reduces computation over inference iterations. Evaluated on both HPEC Graph Challenge benchmarks and conventional DNNs (MNIST, CIFAR-10), SNICIT achieves 6 ∼ 444 × and 1.36 ∼ 1.95 × speed-ups over the previous champions, respectively. Shui Jiang, Tsung-Wei Huang, Bei Yu 0001, Tsung-Yi Ho |
ICPP | 1 |
| 2023 | Security Closure of IC Layouts Against Hardware TrojansabstractDue to cost benefits, supply chains of integrated circuits (ICs) are largely outsourced nowadays. However, passing ICs through various third-party providers gives rise to many threats, like piracy of IC intellectual property or insertion of hardware Trojans, i.e., malicious circuit modifications. Qijing Wang, Bangqi Fu, Shui Jiang, Xiaopeng Zhang 0009, Lilas Alrahis, Ozgur Sinanoglu, Johann Knechtel, Tsung-Yi Ho, Evangeline F. Y. Young |
ISPD | 4 |
| 2022 | OutletGuarder: Detecting DarkSide Ransomware by Power Factor Correction Signals in an Electrical OutletabstractRansomware is a kind of computer malware that has spread widely in recent years, such as DarkSide, which spread around the world recently. It’s reported that DarkSide extorted ${\$}$ 90 million in nine months. It extorts ransom from users by encrypting user files and other methods, causing huge economic losses to users, including commercial organizations and individuals. Existing ransomware detection methods include the hostbased methods and the network-based methods. However, these methods are either hard to deploy or have the possibility to be evaded. In this paper, we propose OutletGuarder, a non-intrusive detection method against DarkSide ransomware based on the signal generated by the Power Factor Correction module of the host computer’s power supply in electrical outlets, which carries the power consumption information of the host computer during the execution of DarkSide. By utilizing the power consumption variation among different programs, especially the power consumption caused by frequent encryption and I/O operations during the execution of DarkSide, OutletGuarder achieves a detection F1 Score of 97.50%. The impact of classification models and untrained programs, as well as the model transferability and robustness are evaluated. Shan Zou, Juchuan Zhang, Shui Jiang, Yushi Cheng, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
ICPADS | 3 |
| 2021 | CapSpeaker: Injecting Voices to Microphones via CapacitorsabstractVoice assistants can be manipulated by various malicious voice commands, yet existing attacks require a nearby speaker to play the attack commands. In this paper, we show that even when no speakers are available, we can play malicious commands by utilizing the capacitors inside electronic devices, i.e., we convert capacitors into speakers and call it CapSpeaker. Essentially, capacitors can emit acoustic noises due to the inverse piezoelectric effect, i.e., varying the voltage across a capacitor can make it vibrate and thus emit acoustic noises. Forcing capacitors to play malicious voice commands is challenging because (1) the frequency responses of capacitors as speakers have poor performance in the range of audible voices, and (2) we have no direct control over the voltage across capacitors to manipulate their emitting sounds. To overcome the challenges, we use a PWM-based modulation scheme to embed the malicious audio onto a high-frequency carrier, e.g., above 20 kHz, and we create malware that can induce the right voltage across the capacitors such that CapSpeaker plays the chosen malicious commands. We conducted extensive experiments with 2 LED lamps (a modified one and a commercial one) and 5 victim devices (iPhone 4s, iPad mini 5, Huawei Nova 5i, etc.). Evaluation results demonstrate that CapSpeaker is feasible at a distance up to 10.5 cm, triggering a smartphone to receive voice commands, e.g., "open the door''. Xiaoyu Ji 0001, Juchuan Zhang, Shui Jiang, Jishen Li, Wenyuan Xu 0001 |
CCS | 3 |
| 2006 | ShanghaiGrid: an Information Service GridabstractAbstract The goal of the ShanghaiGrid is to provide information services to the people. It aims to construct a metropolitan‐area information service infrastructure and establish an open standard for widespread upper‐layer applications from both communities and the government. The Information Service Grid Toolkit and a typical application called the Traffic Information Grid are discussed in detail. Copyright © 2005 John Wiley & Sons, Ltd. Minglu Li 0001, Min-You Wu, Ying Li 0013, Jian Cao 0001, Linpeng Huang, Qianni Deng, Xinhua Lin, Weiqin Tong, Yadong Gui, Aoying Zhou, Xinhong Wu, Shui Jiang |
Concurr. Comput. Pract. Exp. | 13 |
| 2005 | ShanghaiGrid: A Grid Prototype for Metropolis Information Services
Minglu Li 0001, Min-You Wu, Ying Li 0013, Linpeng Huang, Qianni Deng, Jian Cao 0001, Guangtao Xue, Chuliang Weng, Xinhua Lin, Xinda Lu, Weiqin Tong, Yadong Gui, Aoying Zhou, Xinhong Wu, Shui Jiang |
APWeb | 16 |