Xiaoyang Wang 0006

dblp:81/1832-6 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-6629-0176ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 NIO-Cache: Device-Affinitive Page Cache Placement Mechanism for NUMA Systems
Jiazheng Zhang, Xiaoyang Wang 0006, Jiwu Shu
CCGrid2
2026 A 28nm RAID-on-Chip SoC Using On-Chip SRAM and Xfer-Driven Transfers
abstract
Modern data centers and HPC platforms demand storage controllers that sustain high throughput and low latency under massive parallel I/O. However, existing RAID-on-Chip (ROC) designs commonly depend on large off-chip DRAM buffers, which become bottlenecks under write-amplified workloads due to internal data traffic and inefficient I/O scheduling.
Xiaoyang Wang 0006, Minjie Fan 0003, Jiwu Shu
ACM Great Lakes Symposium on VLSI3
2026 Revisiting the RAID Performance Bottleneck in the SSD Era: A Roofline Model Perspective
abstract
RAID is widely employed in modern storage systems due to its high bandwidth and reliability. While traditional RAID controllers utilize large DRAM caches to mitigate HDD latency, the advent of SSDs has shifted potential bottlenecks to the controller architecture itself. In this paper, we introduce a Roofline model adapted for RAID systems, defining I/O Processing Intensity (IOPI) as the ratio of I/O requests to total bytes transferred. We demonstrate that write performance is constrained by DRAM bandwidth under low IOPI and by computational overhead under high IOPI. To mitigate the DRAM bottleneck, we propose a high-bandwidth SRAM-based architecture, bringing parity-based RAID write performance closer to theoretical limits.
Xiaoyang Wang 0006, Congming Gao, Jiwu Shu
ISPASS2
2025 Write-Once, Prove-Once: A Reusable Framework for Secure Boot Verification in Rocq
abstract
Secure boot is a fundamental mechanism for establishing a hardware-rooted chain of trust in modern computing systems. While formal verification using interactive theorem provers like Rocq offers strong correctness guarantees, existing approaches are often tightly coupled to specific hardware configurations and cryptographic algorithms, making them brittle and difficult to reuse when system components change. In this work, we present a flexible and modular framework for formally verifying secure boot processes in Rocq. Our approach introduces a three-layer abstraction model-Feature, Axiom, and Theorem-that decouples high-level security properties from lowlevel implementation details. We abstract critical components such as storage media and cryptographic primitives (e.g., encryption, signatures, hash functions) as interfaces with welldefined operators and axiomatic correctness properties. All highlevel proofs are conducted based on these axioms, without depending on concrete implementations. When specific algorithms or hardware are chosen, the axioms become proof obligations to be verified locally. This separation ensures that changes in implementation-such as switching cryptographic primitives or storage types-do not invalidate existing proofs, enabling significant proof reuse. We demonstrate the applicability of our framework in modeling a realistic secure boot chain. Our work lays the foundation for scalable and maintainable formal verification of trust-sensitive system initialization processes.
Minjie Fan 0003, Kuai Yu, Gaosong Xu, Xiaoyang Wang 0006, Jiwu Shu
ICPADS5
2023 FD-CNN: A Frequency-Domain FPGA Acceleration Scheme for CNN-Based Image-Processing Applications
abstract
In the emerging edge-computing scenarios, FPGAs have been widely adopted to accelerate convolutional neural network (CNN)–based image-processing applications, such as image classification, object detection, and image segmentation, and so on. A standard image-processing pipeline first decodes the collected compressed images from Internet of Things (IoTs) to RGB data, then feeds them into CNN engines to compute the results. Previous works mainly focus on optimizing the CNN inference parts. However, we notice that on the popular ZYNQ FPGA platforms, image decoding can also become the bottleneck due to the poor performance of embedded ARM CPUs. Even with a hardware accelerator, the decoding operations still incur considerable latency. Moreover, conventional RGB-based CNNs have too few input channels at the first layer, which can hardly utilize the high parallelism of CNN engines and greatly slows down the network inference. To overcome these problems, in this article, we propose FD-CNN, a novel CNN accelerator leveraging the partial-decoding technique to accelerate CNNs directly in the frequency domain. Specifically, we omit the most time-consuming IDCT (Inverse Discrete Cosine Transform) operations of image decoding and directly feed the DCT coefficients (i.e., the frequency data) into CNNs. By this means, the image decoder can be greatly simplified. Moreover, compared to the RGB data, frequency data has a narrower input resolution but has 64× more channels. Such an input shape is more hardware friendly than RGB data and can substantially reduce the CNN inference time. We then systematically discuss the algorithm, architecture, and command set design of FD-CNN. To deal with the irregularity of different CNN applications, we propose an image-decoding-aware design-space exploration (DSE) workflow to optimize the pipeline. We further propose an early stopping strategy to tackle the time-consuming progressive JPEG decoding. Comprehensive experiments demonstrate that FD-CNN achieves, on average, 3.24×, 4.29× throughput improvement, 2.55×, 2.54× energy reduction and 2.38×, 2.58× lower latency on ZC-706 and ZCU-102 platforms, respectively, compared to the baseline image-processing pipelines.
Xiaoyang Wang 0006, Zhe Zhou 0002, Zhihang Yuan, Jingchen Zhu, Kangrui Sun, Guangyu Sun 0003
ACM Trans. Embed. Comput. Syst.1
2022 GNNear: Accelerating Full-Batch Training of Graph Neural Networks with near-Memory Processing
abstract
Recently, Graph Neural Networks (GNNs) have become state-of-the-art algorithms for analyzing non-euclidean graph data. However, to realize efficient GNN training is challenging, especially on large graphs. The reasons are many-folded: 1) GNN training incurs a substantial memory footprint. Full-batch training on large graphs even requires hundreds to thousands of gigabytes of memory. 2) GNN training involves both memory-intensive and computation-intensive operations, challenging current CPU/GPU platforms. 3) The irregularity of graphs can result in severe resource under-utilization and load-imbalance problems.
Zhe Zhou 0002, Cong Li 0008, Xuechao Wei, Xiaoyang Wang 0006, Guangyu Sun 0003
PACT4
2020 Hardware-assisted Service Live Migration in Resource-limited Edge Computing Systems
abstract
Service live migration means migrating the running services from one machine to another with negligible service downtime. It has been considered as a powerful mechanism to facilitate service management. However, conventional live migration methods always come with expensive cost of data transmission, and thus can hardly be applied to a real-world edge computing system directly due to the limited network bandwidth. To tackle this problem, some recent works present various techniques to reduce the data transmission.However, these techniques for data transmission reduction always introduce extra computational costs, which have a great impact on the quality of service (QoS), especially in edge systems containing lots of nodes with insufficient computational resources. To alleviate this issue, we propose an insight to offload data reduction computations to a specific hardware accelerator, thus reducing the burden of CPU cores. To this end, we present a novel hardware accelerator design to speed up the data transmission reduction computations to accelerate the service live migration. For evaluation, we implement a prototype on an FPGA platform. Compared to the normal CPU-based approaches, our specialized accelerator is 3.1× faster, 2.9× more-energy efficient, and can reduce 29%∼47% of total migrating time and 24%∼40% of service downtime in our cases. Furthermore, our architecture has great scalability and is easy-configurable to achieve a balance between cost and performance.
Zhe Zhou 0002, Xiaoyang Wang 0006, Zheng Liang 0003, Guangyu Sun 0003, Guojie Luo
DAC3
2020 Edge-Stream: a Stream Processing Approach for Distributed Applications on a Hierarchical Edge-computing System
abstract
With the rapid growth of IoT devices, the traditional cloud computing scheme is inefficient for many IoT based applications, mainly due to network data flood, long latency, and privacy issues. To this end, the edge computing scheme is proposed to mitigate these problems. However, in an edge computing system, the application development becomes more complicated as it involves increasing levels of edge nodes. Although some efforts have been introduced, existing edge computing frameworks still have some limitations in various application scenarios. To overcome these limitations, we propose a new programming model called Edge-Stream. It is a simple and programmer-friendly model, which can cover typical scenarios in edge-computing. Besides, we address several new issues, such as data sharing and area awareness, in this model. We also implement a prototype of edge-computing framework based on the Edge-Stream model. A comprehensive evaluation is provided based on the prototype. Experimental results demonstrate the effectiveness of the model.
Xiaoyang Wang 0006, Zhe Zhou 0002, Ping Han, Tong Meng, Guangyu Sun 0003, Jidong Zhai
SEC1
2020 Bigflow: A General Optimization Layer for Distributed Computing Frameworks
Yuncong Zhang, Xiaoyang Wang 0006, Guangyu Sun 0003, Gong-Lin Zheng, Shan-Hui Yin, Xian-Jin Ye, Zhan Song, Dong-Dong Miao
J. Comput. Sci. Technol.2
2019 RC-NVM: Dual-Addressing Non-Volatile Memory Architecture Supporting Both Row and Column Memory Accesses
abstract
Although emerging non-volatile memories (NVMs) have been comprehensively studied to design next-generation memory systems, the symmetry of the crossbar structure adopted by most NVMs has not been addressed. In this work, we argue that crossbar-based NVMs can enable dual-addressing memory architecture, i.e., RC-NVM, to support both row- and column-oriented memory accesses for workloads with different access patterns. Through circuit-level analysis, we first prove that such a dual-addressing architecture is only practical with crossbar-based NVMs rather than DRAM. Then, we introduce the RC-NVM architecture from bank, chip and module levels, and propose RC-NVM aware memory controller. We also address the challenges to implement the end-to-end RC-NVM system. Especially, we design a novel protocol to solve the cache synonym problem with very little overhead. Finally, we introduce the deployment of RC-NVM for in-memory databases (IMDBs) and evaluate its performance with IMDBs and well-optimized general matrix multiply (GEMM) workloads. Experimental results show that with only 10 percent area overhead 1) the memory access performance of IMDBs can be improved up to 14.5X, and 2) for GEMM, RC-NVM naturally supports SIMD operations and outperforms the best tiled layout by 19 percent.
Shuo Li 0007, Nong Xiao 0001, Peng Wang 0025, Guangyu Sun 0003, Xiaoyang Wang 0006, Yiran Chen 0001, Hai Li 0001, Jason Cong, Tao Zhang 0032
IEEE Trans. Computers5
2018 RC-NVM: Enabling Symmetric Row and Column Memory Accesses for In-memory Databases
abstract
Ever increasing DRAM capacity has fostered the development of in-memory databases (IMDB). The massive performance improvements provided by IMDBs have enabled transactions and analytics on the same database. In other words, the integration of OLTP (on-line transactional processing) and OLAP (on-line analytical processing) systems is becoming a general trend. However, conventional DRAM-based main memory is optimized for row-oriented accesses generated by OLTP workloads in row-based databases. OLAP queries scanning on specified columns cause so-called strided accesses and result in poor memory performance. Since memory access latency dominates in IMDB processing time, it can degrade overall performance significantly. To overcome this problem, we propose a dual-addressable memory architecture based on non-volatile memory, called RC-NVM, to support both row-oriented and column-oriented accesses. We first present circuit-level analysis to prove that such a dual-addressable architecture is only practical with RC-NVM rather than DRAM technology. Then, we rethink the addressing schemes, data layouts, cache synonym, and coherence issues of RC-NVM in architectural level to make it applicable for IMDBs. Finally, we propose a group caching technique that combines the IMDB knowledge with the memory architecture to further optimize the system. Experimental results show that the memory access performance can be improved up to 14.5X with only 15% area overhead.
Peng Wang 0025, Shuo Li 0007, Guangyu Sun 0003, Xiaoyang Wang 0006, Yiran Chen 0001, Hai Li 0001, Jason Cong, Nong Xiao 0001, Tao Zhang 0032
HPCA4