Ziyu Song

dblp:337/3000 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 57% Storage systems · 24% Electronic design automation · 19%
Artificial intelligence
1 paper
Trustworthy machine learning · 50% Image recognition and object detection · 50%
Computer networks
1 paper
Datacenter networks · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.912025
Vision Transformers Beat WideResNets on Small Scale Datasets Adversarial Robustness · AAAI 2025
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.912025
Vision Transformers Beat WideResNets on Small Scale Datasets Adversarial Robustness · AAAI 2025
Computer vision › Image recognition and object detection
image classification
0.912025
Vision Transformers Beat WideResNets on Small Scale Datasets Adversarial Robustness · AAAI 2025
Computer vision › Image recognition and object detection › image classification
robust image classification
0.912025
Vision Transformers Beat WideResNets on Small Scale Datasets Adversarial Robustness · AAAI 2025
Datacenter networks › RDMA
RDMA congestion control
0.912025
SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine · USENIX ATC 2025
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems
0.912025
CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access · ICDE 2025
Storage systems
data placement
0.912025
Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN Training · SC 2025
GPUs and heterogeneous computing
GPU storage access
0.912025
CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access · ICDE 2025
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU training
0.912025
Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN Training · SC 2025
Electronic design automation › design optimization
topology optimization
0.912025
Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN Training · SC 2025
Datacenter networks
RDMA
0.312025
SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine · USENIX ATC 2025
Storage systems › flash and SSD › solid-state drive
NVMe SSD
0.312025
CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access · ICDE 2025

Methods — techniques the papers use, named apart from their topics

vision transformer · 0.9software-programmable congestion control · 0.9max-flow · 0.9knapsack algorithms · 0.9diffusion model · 0.9data augmentation · 0.9batching storage access · 0.9asynchronous APIs · 0.9adversarial training · 0.9
YearPublicationVenuePosition
2025 Vision Transformers Beat WideResNets on Small Scale Datasets Adversarial Robustness
abstract
For an extensive period, Vision Transformers (ViTs) have been deemed unsuitable for attaining robust performance on small-scale datasets, with WideResNet models maintaining dominance in this domain. While WideResNet models have persistently set the state-of-the-art (SOTA) benchmarks for robust accuracy on datasets such as CIFAR-10 and CIFAR-100, this paper challenges the prevailing belief that only WideResNet can excel in this context. We pose the critical question of whether ViTs can surpass the robust accuracy of WideResNet models. Our results provide a resounding affirmative answer. By employing ViT, enhanced with data generated by a diffusion model for adversarial training, we demonstrate that ViTs can indeed outshine WideResNet in terms of robust accuracy. Specifically, under the Infty-norm threat model with epsilon = 8/255, our approach achieves robust accuracies of 74.97% on CIFAR-10 and 44.07% on CIFAR-100, representing improvements of +3.9% and +1.4%, respectively, over the previous SOTA models. Notably, our ViT-B/2 model, with 3 times fewer parameters, surpasses the previously best-performing WRN-70-16. Our achievement opens a new avenue, suggesting that future models employing ViTs or other novel efficient architectures could eventually replace the long-dominant WRN models.
Ziyu Song, Shujun Xie, Longxin Lin
AAAI2
2025 CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access
abstract
With the wide adoption of GPU and the explosion in data volumes, existing accelerator-centric systems require massive storage access. They adopt high-performance storage devices like NVMe SSDs to scale up single-node systems cost-effectively and leverage the CPU to manage these SSDs. However, they suffer from performance bottlenecks because of the high CPU OS kernel overhead and the CPU memory intermediated data transfer. To address this issue, GPU-initiated and GPU-managed SSD management is proposed to allow the GPU to fully manipulate SSDs: 1) direct data transfer from SSD to GPU memory (data plane) and 2) GPU-managed SSD control (control plane). This can potentially enable these GPU systems to fully leverage the SSD bandwidth. However, we still identify two severe issues. First, the GPU-management SSD control leads to low GPU Streaming Multiprocessor utilization. Second, it leads to the serial execution of SSD accesses with GPU computation, which slows down the overall computing task. To this end, we propose CAM, the first asynchronous GPU-initialized, CPU-managed SSD management for batching storage access. It 1) offloads the SSD control plane from GPU to CPU, thus maximizing GPU streaming multiprocessor utilization, and 2) adopts asynchronous user-friendly APIs that allow programmers to easily overlap GPU computation and SSD I/O operations while keeping a synchronous programming experience. As such, CAM enables us to achieve the best of two worlds: high performance and high programmability. The experimental results show that CAM can perform GNN model training, mergesort, and GEMM up to$\mathbf{1.84}\times, \mathbf{1.5}\times$, and$\mathbf{1.84}\times \mathbf{faster}$compared to the existing state-of-the-art GPU systems, while keeping high programmability.
Ziyu Song, Jie Zhang 0081, Jie Sun 0017, Mo Sun 0001, Zihan Yang 0004, Xuzheng Chen, Fei Wu 0001, Huajin Tang, Zeke Wang
ICDE1
2025 Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN Training
abstract
Graph Neural Networks (GNNs) are widely employed in applications like recommendation systems, social network analysis, and fraud detection, but training large-scale GNNs is challenging due to its memory limitations. Existing systems face a trade-off between throughput and monetary cost: Distributed systems require expensive memory scaling, while single-machine out-of-core systems are limited by GPU/PCIe throughput. To this end, we propose Moment, a physical communication topology and data placement co-optimizer to enable high-throughput and low-cost GNN training in a single multi-GPU machine. Moment addresses communication contention and GPU load imbalance issues by modeling the physical topology as capacity-constrained directed graphs and formulating communication scheduling as a max-flow problem. It also introduces a data-distribution-aware knapsack algorithm for optimized data placement. Experimental results show that Moment outperforms out-of-core systems by up to 6.51 × and distributed systems by up to 3.02 ×, with only 50% monetary cost.
Zuocheng Shi, Jie Sun 0017, Ziyu Song, Mo Sun 0001, Fei Wu 0001, Zeke Wang
SC3
2025 SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine
Hongjing Huang, Jie Zhang 0081, Xuzheng Chen, Ziyu Song, Jiajun Qin, Zeke Wang
USENIX ATC4
2025 Advancing explainability of adversarial trained Convolutional Neural Networks for robust engineering applications
Dehua Zhou, Ziyu Song, Zicong Chen, Xianting Huang, Congming Ji, Saru Kumari, Chien-Ming Chen 0001, Sachin Kumar 0002
Eng. Appl. Artif. Intell.2
2023 Ant colony based fish crowding degree optimization algorithm for magnetic resonance imaging segmentation in sports knee joint injury assessment
abstract
Abstract The aim was to analyse the application value of magnetic resonance imaging (MRI) in the injury evaluation of posterolateral corner (PLC) caused by sports, and the ant colony‐based crowding degree of fish swarm optimization algorithm for image segmentation optimization of MRI images. The crowding degree of fish swarm was introduced to construct the ant colony optimization algorithm (ACOA), and the maximum and minimum ant system (MMAS) and ant colony algorithm based on variation features (VF‐based ACA) were introduced to compare with ACOA. Besides, ACOA was applied to the MRI diagnosis of 98 patients with PLC injuries. Then the number of Degree I, Degree II, and Degree III in PLC patients were compared. The results showed that the iteration times and running time of the optimal solution were obtained by ACOA when the number of cities was 10, which were compared with those obtained by MMAS and VF‐based ACA with no great differences ( P > 0.05), and the iteration times and running time of ACOA were greater than those of MMAS and VF‐based ACA when the number of cities was 50 ( P < 0.05). There were 73 patients with PLC injuries involved 2 or more structures, 11 patients with PLC injuries only involved the lateral collateral ligament (LCL), and 7 patients with PLC injuries only involved the popliteal tendon (PLT). The injury rates of LCL, medial collateral ligament, anterior cruciate ligament (ACL), posterior cruciate ligament (PCL), and PLT were 87.02%, 51.33%, 72.45%, 41.75%, and 84.06%, respectively. The number of patients with PLC injuries combined with Grade I of LCL, PCL, ACL, PCL, biceps femoris tendon, and PLT was higher than that of patients with Grade II and III ( P < 0.05). In conclusion, ACOA was much better than MMAS and VF‐based ACA in dealing with complex work, which effectively improved the ergodicity of search and reduced the running time. There were more common in PLC injuries combined with LCL, PCL, ACL, PCL, and PLT, MRI images based on ACOA could clearly show the degree of ligament and tendon injuries.
Ziyu Song
Expert Syst. J. Knowl. Eng.1
2023 Research on bud counting of cut lily flowers based on machine vision
Ziyu Song, Yancheng Zhang
Multim. Tools Appl.2