VLDB 2026 Research / reviewers in the wild / expert
Shouzhen Gu
dblp:136/4933
· DBLP profile ↗
26ranked-venue papers
6as first author
10since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 3 first-author · 8 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improving F2FS fsync() Latency Through Parallelizing Dnode and Data Page WritebackabstractF2FS improves performance and longevity through out-place updates, and is now wildly used in the real world. However, additional overhead is introduced to support such design, which leads to a significant performance bottleneck when performing fsync(). Through a series of experimental observations, the paper reveals the impact of dnode page writeback on throughput and fsync() latency. The serial flushing of data pages and dnode pages in the current fsync() design limits the potential for parallel write back and fails to fully utilize the parallelism of flash devices. We then deeply look into the current fsync() design and find the main difficulty of paralleling fsync() is the dependency between dnode pages and data pages. Based on these findings, we propose a new dual-thread design that significantly reduces the total latency of fsync() and improves throughput by pre-allocating data pages. Then, we give two optimizations to reduce overhead. A red-black tree is introduced to cache old block addresses for better node management performance. A linked list is introduced to avoid contention of the page cache for better page performance. We implemented our method in Linux Kernel, and the experimental results show that our dual-thread design can decrease fsync() latency and increase write throughput in different situations. Mengyang Ma, Yumiao Zhao, Yunpeng Song, Shouzhen Gu |
NAS | 5 |
| 2024 | Zoned-WB: WriteBooster Design with Zoned Storage for User Experience on SmartphonesabstractWriteBooster is widely adopted as a non-volatile write buffer to enhance user experience for smartphones. How-ever, the host suffers sub-optimal write performance when using WriteBooster due to the lack of utilization of rich semantic infor-mation. With the zoned storage being included in smartphones, WriteBooster design presents new opportunities. In this paper, we propose Zoned-WB, a WriteBooster management scheme based on zoned storage to utilize the rich semantic information on the host to improve user experience. Specifically, Zoned- Wbincludes two parts, zoned storage-based WB and foreground request-aware WB. First, the zoned storage-based WB is aimed at designing WriteBooster management scheme based on zoned storage. Second, foreground request-aware WB is designed to adaptively adjust the capacity quota for different types of requests in Writebooster based on rich semantic information on the host. We evaluate Zoned-WB on a zoned storage emulator with workloads collected from smartphones. Evaluation results show that Zoned- Wbcan effectively improve user experience. Dingcui Yu, Ziang Huang, Wentong Li 0002, Zonghuan Yan, Shouzhen Gu, Liang Shi 0001 |
NAS | 5 |
| 2023 | An Efficient Decoder for a Linear Distance Quantum LDPC CodeabstractRecent developments have shown the existence of quantum low-density parity check (qLDPC) codes with constant rate and linear distance. A natural question concerns the efficient decodability of these codes. In this paper, we present a linear time decoder for the recent quantum Tanner codes construction of asymptotically good qLDPC codes, which can correct all errors of weight up to a constant fraction of the blocklength. Our decoder is an iterative algorithm which searches for corrections within constant-sized regions. At each step, the corrections are found by reducing a locally defined and efficiently computable cost function which serves as a proxy for the weight of the remaining error. Shouzhen Gu, Christopher A. Pattison, Eugene Tang |
STOC | 1 |
| 2023 | A Differentially Private Federated Learning Model Against Poisoning Attacks in Edge ComputingabstractFederated learning is increasingly popular, as it allows us to circumvent challenges due to data islands, by training a global model using data from one or more data owners/sources. However, in edge computing, resource-constrained end devices are vulnerable to be compromised and abused to facilitate poisoning attacks. Privacy-preserving is another important property to consider when dealing with sensitive user data on end devices. Most existing approaches only consider either defending against poisoning attacks or supporting privacy, but not both properties simultaneously. In this paper, we propose a differentially private federated learning model against poisoning attacks, designed for edge computing deployment. First, we design a weight-based algorithm to perform anomaly detection on the parameters uploaded by end devices in edge nodes, which improves detection rate using only small-size validation datasets and minimizes the communication cost. Then, differential privacy technology is leveraged to protect the privacy of both data and model in an edge computing setting. We also evaluate and compare the detection performance in the presence of random and customized malicious end devices with the state-of-the-art, in terms of attack resiliency, communication and computation costs. Experimental results demonstrate that our scheme can achieve an optimal tradeoff between security, efficiency and accuracy. Jun Zhou 0018, Yisong Wang 0002, Shouzhen Gu, Zhenfu Cao, Xiaolei Dong, Kim-Kwang Raymond Choo |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | Work-in-Progress: Cooperative MLP-Mixer Networks Inference On Heterogeneous Edge Devices through Partition and FusionabstractAs a newly proposed DNN architecture, MLP-Mixer is attracting increasing attention due to its competitive results compared to CNNs and attention-base networks in various tasks. Although MLP-Mixer only contains MLP layers, it still suffers from high communication costs in edge computing scenarios, resulting in long inference time. To improve the inference performance of an MLP-Mixer model on correlated resource-constrained heterogeneous edge devices, this paper proposes a novel partition and fusion method specific for MLP-Mixer layers, which can significantly reduce the communication costs. Experimental results show that, when the number of devices increases from 2 to 6, our partition and fusion method can archive 1.01-1.27x and 1.54-3.12x speedup in scenarios with heterogeneous and homogeneous devices, respectively. Shouzhen Gu, Mingsong Chen 0001 |
CASES | 2 |
| 2022 | Fairness Scheduling for Tasks with Different Real-time Level on Heterogeneous SystemsabstractFor a real-time task-intensive systems, the fairness of task execution in dynamic scheduling is an important research area. However, many exist scheduling algorithms are unable to guarantee that tasks can be completed by the deadline and executed with a fair priority. In this paper, we proposed an efficient Multi-DAG real-time scheduling algorithm, HSDFW, which employs a fair priority calculation method to enable tasks with different real-time levels can be completed by the deadline, and a rejection policy to improve the performance of schedule. We proposed an INLP model and an evaluation simulator to verify the efficiency of HSDFW algorithm. The evaluation results show that our proposed algorithm has excellent performance in terms of average scheduling length and resource utilization. Shifan Shao, Shouzhen Gu, Edwin H.-M. Sha, Qingfeng Zhuge |
ICPADS | 2 |
| 2022 | Transient computing for energy harvesting systems: A survey
Min Jia 0002, Edwin H.-M. Sha, Qingfeng Zhuge, Shouzhen Gu |
J. Syst. Archit. | 4 |
| 2021 | Relaxed Placement: Minimizing Shift Operations for Racetrack Memory in Hybrid SPMabstractRacetrack memory (RM) has high access performance comparable to SRAM. It is a kind of non-volatile memory (NVM), which consists of data block clusters (DBCs) and access ports. However, data accessing on RM is based on shift operations, which will decrease the performance of RM. This paper proposes techniques by using SRAM to reduce the shifts and improve the accessing performance of RM. The key idea is to place randomly accessed data on SRAM ahead of time to relax the data placement on RM. First, a greedy scheduling strategy is proposed to reduce the requirement of SRAM. Second, to further reduce shifts, data with similar association degree are grouped and allocated to each DBC. Experimental results show that the proposed techniques reduce the shifts by 72.3% with only 256-byte SRAM compared to pure RM. Rui Xu 0013, Edwin H.-M. Sha, Qingfeng Zhuge, Liang Shi 0001, Shouzhen Gu, Yan Hou |
ACM Great Lakes Symposium on VLSI | 5 |
| 2021 | Performance optimization for parallel systems with shared DWM via retiming, loop scheduling, and data placement
Shouzhen Gu, Rui Xu 0013, Edwin H.-M. Sha, Qingfeng Zhuge |
J. Syst. Archit. | 2 |
| 2021 | Optimizing the data placement and scheduling on multi-port DWM in multi-core embedded system
Edwin H.-M. Sha, Shouzhen Gu, Qingfeng Zhuge |
J. Syst. Archit. | 3 |
| 2020 | An Empirical Study of Hybrid SSD with Optane and QLC FlashabstractEmerging non-volatile memory (NVM) technologies provide a new way to solve the I/O bottleneck problem. As one of the widely respected solutions, hybrid storage device performance in the real environment is worth studying. Previously, due to the delayed progress of NVM, most of the studies are proceeded on simulated devices. In this paper, an empirical study is presented on the state-of-the-art hybrid storage device - Intel Optane H10, which is designed with Optane Memory and Quad-Level Cell (QLC) NAND flash. Several interesting findings are concluded with the study, which should be well considered during the employment. Yina Lv, Changlong Li 0006, Shouzhen Gu, Liang Shi 0001 |
ICCD | 4 |
| 2020 | Optimizing Data Placement for Hybrid SPM with SRAM and Racetrack MemoryabstractIn this paper, a novel hybrid scratchpad memory (SPM) with SRAM and racetrack memory (RM) is proposed. The basic idea is to smartly place data on SPM by taking the advantages of these two memories. First, a metric is proposed to represent the access cost of data; Second, a data placement scheme is proposed based on the metric; Finally, to maximize the size of SPM, a scheme is further proposed to minimize the size of SRAM. Experimental results show that the proposed scheme reduces the shift operations of RM by 80.12% and reduces the cost of SPM by 80.72% with only 17.63% SRAM compared with a baseline SPM with pure RM. Rui Xu 0013, Edwin H.-M. Sha, Qingfeng Zhuge, Shouzhen Gu, Liang Shi 0001 |
ICCD | 4 |
| 2020 | Towards the design of efficient hash-based indexing scheme for growing databases on non-volatile memory
Zhulin Ma, Edwin H.-M. Sha, Qingfeng Zhuge, Weiwen Jiang, Runyu Zhang 0002, Shouzhen Gu |
Future Gener. Comput. Syst. | 6 |
| 2020 | Hardware/Software Co-Exploration of Neural ArchitecturesabstractWe propose a novel hardware and software co-exploration framework for efficient neural architecture search (NAS). Different from existing hardware-aware NAS which assumes a fixed hardware design and explores theNAS spaceonly, our framework simultaneously explores both the architecture search space and thehardware design spaceto identify the best neural architecture and hardware pairs that maximize both test accuracy and hardware efficiency. Such a practice greatly opens up the design freedom and pushes forward the Pareto frontier between hardware efficiency and test accuracy for better design tradeoffs. The framework iteratively performs a two-level (fast and slow) exploration. Without lengthy training, the fast exploration can effectively fine-tune hyperparameters and prune inferior architectures in terms of hardware specifications, which significantly accelerates the NAS process. Then, the slow exploration trains candidates on a validation set and updates a controller using the reinforcement learning to maximize the expected accuracy together with the hardware efficiency. In this article, we demonstrate that the co-exploration framework can effectively expand the search space to incorporate models with high accuracy, and we theoretically show that the proposed two-level optimization can efficiently prune inferior solutions to better explore the search space. The experimental results on ImageNet show that the co-exploration NAS can find solutions with the same accuracy, 35.24% higher throughput, 54.05% higher energy efficiency, compared with the hardware-aware NAS. Weiwen Jiang, Lei Yang 0018, Edwin H.-M. Sha, Qingfeng Zhuge, Shouzhen Gu, Sakyasingha Dasgupta, Yiyu Shi 0001, Jingtong Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | A Wear-Leveling-Aware Fine-Grained Allocator for Non-Volatile MemoryabstractEmerging non-volatile memories (NVMs) are promising main memory for their advanced characteristics. However, the low endurance of NVM cells makes them vulnerable to frequent fine-grained updates. This paper proposes a Wear-leveling Aware Fine-grained Allocator (WAFA) for NVM. WAFA divides pages into basic memory units to support fine-grained updates. WAFA allocates the basic memory units of a page in a rotational manner to distribute fine-grained updates evenly on memory cells. The fragmented basic memory units of each page caused by the memory allocation and deallocation operations are reorganized by reform operation. We implement WAFA in Linux kernel 4.4.4. Experimental results show that WAFA can reduce 81.1% and 40.1% of the total writes of pages over NVMalloc and nvm_alloc, the state-of-the-art wear-conscious allocator for NVM. Meanwhile, WAFA shows 48.6% and 42.3% performance improvement over NVMalloc and nvm_alloc, respectively. Xianzhang Chen, Qingfeng Zhuge, Edwin H.-M. Sha, Shouzhen Gu, Chaoshu Yang, Chun Jason Xue |
DAC | 5 |
| 2018 | Write-Aware Data Allocation on Heterogeneous Memory Architecture with Minimum CostabstractMore and more Non-Volatile Memories (NVM) have been widely applied to various embedded systems to build the heterogeneous memory architecture. However, the write-endurance of NVM remains a great challenge. Hence, we should take full consideration of the write-endurance of NVM when allocating data on heterogeneous memory architecture. There is an observation that, for most real workloads, about 10% of data account for 90% write operations. This brings us an opportunity to reduce the write wear of NVM through carefully allocating write-intensive data. In this paper, we explore the problem that how to find a balance between the system cost and write-endurance of NVM for data allocation on heterogeneous memory architecture. We propose a write-aware data allocation algorithm, WADA. WADA can not only greatly reduce the write wear of NVM, but also guarantee the near-optimal system cost. We also propose an integer linear programming (ILP) model to generate an optimal data allocation, which can obtain the minimum cost. The result of ILP can be used as a standard to evaluate the efficiency of other algorithms. Experiments show that WADA outperforms all the other algorithms on both system cost and write wear of NVM. Compared to previous algorithms, WADA can reduce up to 47.77% system cost and 60.89% write wear of NVM. Compared to ILP, WADA can achieve the near-optimal system cost within just 2% difference. Yanbo Zhou, Shouzhen Gu, Lixia Zheng, Edwin H.-M. Sha, Qingfeng Zhuge, Lin Wu 0002 |
RTCSA | 2 |
| 2016 | A Time, Energy, and Area Efficient Domain Wall Memory-Based SPM for Embedded SystemsabstractApplications that run in the embedded systems normally should be finished within a timing constraint in energy-efficient fashion. Due to these two requirements, the embedded systems often employ software-controlled scratch pad memory (SPM) instead of hardware-controlled cache as their on-chip memory. The data accesses in SPMs are controlled purely by the software, which provides better time-predictability and precise time-control. In this paper, we propose a time, energy, and area efficient domain wall memory (DWM)-based SPM for embedded systems. To efficiently manage this type of novel SPM, an integer nonlinear programming formulation and the instructions group schedule algorithm are proposed to generate memory access instruction scheduling and data placement. In addition, the longest move reduce algorithm is also proposed to configure different types of DWM memory cells to achieve minimal area size. Experimental results show that the proposed techniques can generate a configuration of DWM-based SPM with minimal area size while satisfying time constraint. Shouzhen Gu, Edwin H.-M. Sha, Qingfeng Zhuge, Yiran Chen 0001, Jingtong Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Area and performance co-optimization for domain wall memory in application-specific embedded systemsabstractDomain Wall Memory (DWM), a recently developed spin-based non-volatile memory technology, inherently offers unprecedented benefits in density by storing multiple bits in the domains of a ferromagnetic nanowire, which logically resembles a bit-serial tape. However, this structure also leads to a unique challenge that the bits must be sequentially accessed by performing \shift" operations, resulting in variable and potential higher access latencies. In this paper, we propose a hardware and software co-optimize approach to improve area efficiency and performance for DWM in application-specific embedded systems. For an application-specific embedded system, this technique can obtain a DWM which consists of both micro-cell DWM and macro-cell DWM with minimal area size. Meanwhile, instruction schedule and data allocation with minimal memory access overhead are generated. Experimental results show that the proposed method can minimize the DWM area size while satisfying a system performance constraint. Shouzhen Gu, Edwin H.-M. Sha, Qingfeng Zhuge, Yiran Chen 0001, Jingtong Hu |
DAC | 1 |
| 2015 | Optimizing Task and Data Assignment on Multi-Core Systems with Multi-Port SPMsabstractMulti-core processors have been adopted in modern embedded systems to meet the ever increasing performance requirements. Scratchpad memory (SPM), a software-controlled on-chip memory, has been used in embedded systems as an alternative to hardware-controlled cache due to its advantage in die area, power consumption, and timing predictability. SPMs in multi-core systems can be accessed by both local core and remote cores. In order to alleviate data contention on a SPM unit, multi-port SPMs are employed in multi-core systems. In such systems, proper task scheduling and data assignment can significantly improve the overall performance by exploring the parallelism of computation tasks and concurrent data accesses on SPMs. Since scheduling for multi-core systems is NP-Complete in general. In this paper, we propose an ILP formulation to optimally determine the task scheduling and data assignment on multi-core systems with multi-port SPMs. Since ILP takes exponential time to finish, we also propose a heuristic method, including the task assignment with remote access reduced (TARAR) algorithm and the minimum memory access cost (MMAC) algorithm, to obtain near optimal solutions within polynomial time. According to the experimental results, the ILP formulation can improve the system performance by 23.02 percent over the HAFF algorithm on average, while the heuristic algorithm can improve the system performance by 16.48 percent over HAFF on average. Shouzhen Gu, Qingfeng Zhuge, Juan Yi, Jingtong Hu, Edwin H.-M. Sha |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Minimum-cost data allocation with guaranteed probability on multiple types of memoryabstractAs the advance of memory technologies, multiple types of memory such as different kinds of non-volatile memory (NVM), SRAM, DRAM, etc. provide a flexible configuration considering performance, energy and cost. For improving the performance of systems with multiple types of memory, data allocation is one of the most important tasks. The previous studies on data allocation problem assume the worst (fixed) case of data-access frequencies. However, the data allocation produced by employing worst case usually leads to an inferior performance for most of time. In this paper, we model this problem by probabilities and design efficient algorithms that can give optimal-cost data allocation with a guaranteed probability. The proposed DAGP algorithm produces a set of feasible data allocation solutions which generates the minimum access time or cost guaranteed by a given probability. The experiments show that our technique can significantly reduce the access time or cost compared with the technique considering worst case scenario. For example, comparing with the optimal result generated by employing the worst cases, our technique can reduce memory access time by 10.35% on average when guaranteed probability is set to be 0.8. Moreover, for 80 percents of cases, memory access time is reduced by 23.98% on average. Shouzhen Gu, Qingfeng Zhuge, Jingtong Hu, Juan Yi, Edwin H.-M. Sha |
RTCSA | 1 |
| 2014 | A space allocation and reuse strategy for PCM-based embedded systems
Linbo Long, Duo Liu 0002, Jingtong Hu, Shouzhen Gu, Qingfeng Zhuge, Edwin H.-M. Sha |
J. Syst. Archit. | 4 |
| 2014 | Scheduling to Optimize Cache Utilization for Non-Volatile Main MemoriesabstractIn power and size sensitive embedded systems, non-volatile memories (NVMs) are replacing DRAM as the main memory since they have higher density, lower static power consumption, and lower costs. Unfortunately, these technologies are limited by their endurance and long write latencies. To minimize the main memory access time and extend the lifetime of the NVM, we optimally schedule tasks by an ILP formulation. We also present a heuristic, Concatenation Scheduling, to solve large problems in a reasonable amount of time. Our experimental results show that when compared with list scheduling, concatenation scheduling can reduce the total memory access time by an average of 9.99% and increase the lifetime of the NVM by 26.66%. When compared with list scheduling, ILP can reduce the total memory access time by an average of 12.39% and increase the lifetime of the NVM by 38.74%. Jingtong Hu, Qingfeng Zhuge, Chun Jason Xue, Wei-Che Tseng, Shouzhen Gu, Edwin H.-M. Sha |
IEEE Trans. Computers | 5 |
| 2014 | Minimizing System Cost with Efficient Task Assignment on Heterogeneous Multicore Processors Considering Time ConstraintabstractHigh-performance computing systems typically employ heterogeneous multicore design to improve both execution performance and efficiency. Task assignment is critical in exploiting the diversity of computation capability, energy consumption, as well as communication cost on heterogeneous multicore processors. In this paper, we explore the opportunity of task assignment on heterogeneous multicore processors to minimize execution and communication costs considering time constraint. The general heterogeneous task assignment problem is NP-Complete. However, we find that optimal task assignment can be achieved for widely used, tree-shaped task graphs using dynamic programming. We first propose a dynamic programming algorithm, the Optimal Tree Assign (OTA) algorithm, to generate optimal assignments for trees. Then, we develop the Integer Linear Programming model of the general task assignment problem for Directed Acyclic Graphs. A polynomial-time heuristic, the Extended Tree Assignment algorithm, is also proposed to produce near-optimal solutions for the general heterogeneous task assignment problem efficiently. The experimental results show that the proposed algorithms outperform both homogeneous task assignment method and greedy strategy for all the benchmarks. The OTA algorithm reduces the total system time by 42.5 percent and 23.5 percent on average compared with the homogeneous task assignment method and greedy algorithm, respectively. Qingfeng Zhuge, Shouzhen Gu, Jingtong Hu, Edwin H.-M. Sha |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Efficient task assignment and scheduling for MPSoC DSPS with VS-SPM considering concurrent accesses through data allocationabstractVirtually Shared Scratch-Pad Memory (VS-SPM) with multiple memory banks can be used as on-chip memory on multiprocessor systems-on-chips (MPSoCs) to close the speed gap between fast processors and slow memories. By exploring the parallelism of computation tasks on processors and concurrent data accesses on each SPM, the results of task assignment and data allocation can significantly affect the overall performance of a schedule. In this paper, we propose ILP formulations for solving the problem of task assignment and scheduling on MPSoCs with multi-bank VS-SPM.We also propose a polynomial-time algorithm, the Potential Remote Access Prediction (PRAP) algorithm, to generate near-optimal results efficiently. The experimental results demonstrate the effectiveness of our technique. Shouzhen Gu, Qingfeng Zhuge, Jingtong Hu, Juan Yi, Edwin H.-M. Sha |
ICASSP | 1 |
| 2013 | A space-based wear leveling for PCM-based embedded systemsabstractPhase change memory (PCM) has emerged as a promising candidate to replace DRAM in embedded systems. However, it can only sustain a limited number of write operations. To solve this issue, this paper proposes a novel and effective wear-leveling technique in software level to prolong the lifetime of PCM-based embedded systems. A polynomial-time algorithm, Multi-Space Wear Leveling Algorithm (MWL), is proposed to achieve effective wear-leveling. The experimental results show our technique can greatly extend the lifetime of PCM-based embedded systems compared with the previous work. Compared with the method without adopting wear-leveling, it introduces no more than 0.7% extra writes and 0.6% running overhead. Linbo Long, Duo Liu 0002, Jingtong Hu, Shouzhen Gu, Qingfeng Zhuge, Edwin H.-M. Sha |
RTCSA | 4 |
| 2013 | Optimizing task assignment for heterogeneous multiprocessor system with guaranteed reliability and timing constraintabstractEffective task assignment, which is essential for achieving high performance in a heterogeneous multiprocessor system, remains a challenging problem despite extensive studies. This paper addresses the task assignment problem with guaranteed reliability and timing constraint for heterogeneous multiprocessor system. Inherently, heterogeneous systems are more complex than homogeneous systems. The added complexity could increase the potential for system failures. In this paper, we describe a method to determine an assignment which satisfies the timing constraint and the reliability requirement. We develop an Integer Linear Programming (ILP) formulation to find the optimal solutions. For the general problem, the task assignment problem is NP-Complete. Therefore, we propose a polynomial-time heuristic algorithm, DAG Heu algorithm, to solve the general problem. Experimental results on benchmark task graphs of several well-known parallel applications show that the proposed algorithm and the ILP formulation significantly outperform existing algorithms. Juan Yi, Qingfeng Zhuge, Jingtong Hu, Shouzhen Gu, Mingwen Qin, Edwin H.-M. Sha |
RTCSA | 4 |