Changyou Zhang

dblp:85/351 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-4025-0736ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorArtificial intelligence and machine learning · 2Theory of computation · 2Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AMGNet: Dual-Domain Scale-Specific Graph Learning for Multivariate Time Series Forecasting
Shuzi Niu, Huiyuan Li 0002, Changyou Zhang, Changmao Wu
ICIC (3)4
2025 A Pattern-Aware Finite Element Matrix Assembly Method on GPUs
abstract
The Finite Element Method (FEM) is a fundamental technique for solving large-scale and complex engineering problems. During the construction of the system equations, the efficiency of finite element matrix assembly plays a crucial role in the overall performance. However, existing approaches often overlook the sensitivity of assembly algorithm performance to mesh characteristics, making it difficult to achieve optimal performance across diverse problems. In this work, we propose a novel pattern-aware FEM matrix assembly method on GPUs. To this end, we thoroughly analyze the key factors affecting performance and extract a set of potentially influential mesh features and density representations. Based on this, we construct a Deep learning-based prediction model that fully captures the input mesh characteristics to predict the performance-optimal assembly strategy. Experimental results on mesh datasets with a wide range of feature variations demonstrate that our method achieves remarkable prediction accuracy and delivers up to$7.34 \times$speedup in execution time compared to state-of-the-art approaches. To the best of our knowledge, this is the first work that introduces auto-tuning for the FEM matrix assembly process.
Changyou Zhang, Zhuo Tian, Guangzhao Li, Chen Ju
CLUSTER3
2025 MG-αGCD: Accelerating Graph Community Detection on Multi-GPU Platforms
abstract
Graph community detection is widely applied in fields such as genetic engineering and social network analysis.As the scale of input graphs continues to grow and multi-GPU platforms become increasingly prevalent, utilizing the storage capacity and computational power of these platforms to scale graph community detection algorithms has become more feasible.However, existing multi-GPU graph community detection methods are constrained by the traditional CPUdominated communication model, and fail to simultaneously account for the irregular sparse memory access patterns and the latency disparities between local and remote communication.Consequently, they do not fully exploit modern high-speed GPU interconnect technologies.Furthermore, while current solutions propose various strategies to mitigate the decline in clustering quality during parallelization, these approaches are often inefficient or compromise clustering quality.Finally, to address the high memory overhead of the algorithms, existing Multi-GPU solutions extend the graph size limit at the cost of reduced performance.To address these challenges, we propose MG-𝛼GCD, a novel graph community detection algorithm designed for multi-GPU platforms.First, MG-𝛼GCD introduces a loading balancing and latency-aware computation-communication pipeline that effectively mitigates the overhead of highlatency remote communication.Second, MG-𝛼GCD incorporates a bidirectional probing heuristic to enhance execution efficiency while outperforming existing methods in clustering quality.Lastly, MG-𝛼GCD employs a two-phase graph coarsening algorithm consisting of a symbolic phase and a numeric phase, which significantly reduces GPU peak memory usage and minimizes data transfers between CPU and
Changyou Zhang
ICS2
2023 Unify the Usage of Lexicon in Chinese Named Entity Recognition
Wenjia Wu, Changyou Zhang, Shuzi Niu
DASFAA (3)2
2023 DeltaSPARSE: High-Performance Sparse General Matrix-Matrix Multiplication on Multi-GPU Systems
abstract
Sparse General Matrix-Matrix Multiplication (SpGEMM) serves as a fundamental operation in the domains of sparse linear algebra and graph data processing. The majority of existing research predominantly concentrates on optimizing SpGEMM in the context of single GPU scenarios. Nevertheless, the growing prevalence of multi-GPU systems offers opportunities to harness the computational capabilities of multiple GPUs, thereby enhancing the performance of sparse general matrix-matrix multiplication. The efficacy of multi-GPU SpGEMM is chiefly constrained by two factors: (1) the irregular sparse pattern of sparse matrices, and (2) the load imbalance among multiple GPUs. To address these challenges, this paper presents DeltaSPARSE, the first algorithm to achieve significant speed-up for large-scale SpGEMM on multiple GPUs, to the best of our knowledge. Our algorithm incorporates hybrid accumulators, which dynamically choose the most suitable accumulator algorithm for rows exhibiting varying levels of sparsity. Moreover, we suggest a hierarchical task scheduling approach to partition and allocate tasks across diverse levels of parallel hardware, such as GPU s, blocks, warps, and threads. Experimental outcomes utilizing the SuiteS parse matrix dataset reveal that DeltaSPARSE displays near-linear scalability in multi-GPU configurations. Furthermore, it attains substantial speed enhancements in comparison to the present state-of-the-art single GPU SpGEMM methods, including NSPARSE, spECK, bhSPARSE, and cuSPARSE, across matrices with various sparse characteristics.
Changyou Zhang
HiPC2
2023 Accelerating Sparse General Matrix-Matrix Multiplication for NVIDIA Volta GPU and Hygon DCU
abstract
Sparse general matrix-matrix multiplication (SpGEMM) is challenging especially on graphic accelerators. Existing solutions do not fully utilize the shared memory of the graphics accelerator. Our proposal could effectively utilize the graphics accelerator's on-chip shared memory and dynamically assign the device resources by grouping the rows based on a hybrid strategy for load balancing. Experiments show that our proposal achieves speedups of up to x7.43 in double precision compared to existing SpGEMM libraries. Our implementation is fully general and our optimization strategy adaptively processes the SpGEMM workload row-wise to substantially improve performance by decreasing the work complexity and utilizing the memory hierarchy more effectively.
Zhuo Tian, Changyou Zhang
HPDC3
2023 ShapeRef: A Representation Method of Industrial Abnormal Time-Series Waveform Based on Shape Reference
abstract
Time-series waveform data widely exist in various industrial fields, such as equipment monitoring and fault diagnosis. The current time series representation methods have limitations when dealing with industrial abnormal time-series waveforms, such as limited applicability, semantic ambiguity, and time distortion. This work proposes a novel shape reference-based representation method for industrial abnormal time-series waveform (ShapeRef), which takes the shape of the standard waveform as a reference to represent the anomaly deviation. Specifically, ShapeRef first establishes a time-series shape reference frame, then proposes the minimum shape difference-based mapping method to describe the mapping process of coordinates, and finally reduces multi-intersection points in the mapping process to achieve uniform mapping of the abnormal time-series waveform. Experimental results show that ShapeRef can effectively represent abnormal time-series waveforms and outperforms several baseline methods in the clustering task of a real industrial equipment waveform dataset. This work enhances the accuracy and reliability of industrial equipment monitoring and fault diagnosis, which could have significant practical implications.
Changyou Zhang, Wenjia Wu, Wen Bo
SMC2
2022 An Asynchronous Parallel Algorithm to Improve the Scalability of Finite Element Solvers
abstract
Large-scale finite element equations are usually solved by the Preconditioned Conjugate Gradient (PCG) iterative method, and the computational hotspots are sparse matrix-vector multiplication and inner product, which requires local and global communication. But, the latency of the high-performance cluster is too long for the PCG algorithm. This paper proposes a new asynchronous algorithm that could reduce the number of communication among cluster nodes to improve the scalability of finite element solvers. The performance could be improved up to 6.31 x.
Zhuo Tian, Changyou Zhang
CLUSTER2
2020 Boosting performance of virtualized desktop infrastructure with physical GPU and SPICE
Shupan Li, Chungang Shi, Liequan Che, Changyou Zhang, Yuanzhang Li 0001
Sci. China Inf. Sci.5
2019 Privacy-preserving governmental data publishing: A fog-computing-based differential privacy approach
Chunhui Piao, Yajuan Shi, Jiaqi Yan 0002, Changyou Zhang
Future Gener. Comput. Syst.4
2019 A packet-reordering covert channel over VoLTE voice and video traffics
Xiaosong Zhang 0002, Liehuang Zhu, Xianmin Wang, Changyou Zhang, Yu-an Tan 0001
J. Netw. Comput. Appl.4
2018 Distributed Parallel Simulation of Primary Sample Space Metropolis Light Transport
Changmao Wu, Changyou Zhang, Qiao Sun 0005
ICA3PP (1)2
2018 Bandwidth Reduced Parallel SpMV on the SW26010 Many-Core Platform
abstract
SpMV (Sparse Matrix-Vector multiplication), in its simplest form y = Ax, multiplies a sparse matrix with a dense vector and is a widely used computing primitive in the domain of HPC. On the newly SW26010 many-core platform, we propose a highly efficient CSR (Compressed Storage Row) based implementation of parallel SpMV, referred to as SWCSR-SpMV in the sequel. SpMV in the CSR format can be trivially parallelized but its performance is majorly impeded by memory access efficiency, and therefore to leverage high-throughput memory access mechanism while avoiding redundant bandwidth usage becomes the major goal of designing high performance SpMV on the target platform. The original problem is sequentially partitioned into row-slices, each of which can reside in the fast scratchpad memory, so that the loaded x'es can be reused; meanwhile, a dynamic look-ahead scheme is applied to avoid redundant memory access; we split the many-core mesh into smaller communication scope to facilitate the sharing of the common data across the working threads via the high speed on-mesh data bus. Beyond the above, to leverage massive parallelism balanced workload is ensured by both static and dynamic means. Performance evaluation is done on a benchmark of 36 frequently used sparse matrices in the fields of graph computing, data mining, computational fluid dynamics, etc.. While the performance upper-bound is defined by the ratio between the minimal data access volume required against the practically optimal bandwidth, ignoring the computing overhead, SWCSR-SpMV can achieve an efficiency of nearly 87%, maintaining over 75% for 1/3 of the testing matrices. SWCSR-SpMV is further applied in a PETSc based application, a 1.75x-2.6x speedup is sustained in a multi-process environment on the Sunway TaiHuLight supercomputer.
Qiao Sun 0005, Changyou Zhang, Changmao Wu, Leisheng Li
ICPP2
2018 A novel anti-detection criterion for covert storage channel threat estimation
Changyou Zhang, Yongji Wang 0002
Sci. China Inf. Sci.2
2018 An interleaved depth-first search method for the linear optimization problem with disjunctive constraints
Yinrun Lyu, Changyou Zhang, Dacheng Qu, Nasro Min-Allah, Yongji Wang 0002
J. Glob. Optim.3
2018 An extra-parity energy saving data layout for video surveillance
Xiao Yu 0005, Changyou Zhang, Yuanzhang Li 0001, Yu-an Tan 0001
Multim. Tools Appl.2
2018 A code protection scheme by process memory relocation for android devices
Xiaosong Zhang 0002, Yu-an Tan 0001, Changyou Zhang, Yuanzhang Li 0001, Jun Zheng 0007
Multim. Tools Appl.3
2018 An optimized data hiding scheme for Deflate codes
Yu-an Tan 0001, Changyou Zhang, Jun Zheng 0007
Soft Comput.4
2017 A round-optimal lattice-based blind signature scheme for cloud services
Yu-an Tan 0001, Xiaosong Zhang 0002, Liehuang Zhu, Changyou Zhang, Jun Zheng 0007
Future Gener. Comput. Syst.5
2017 Solving linear optimization over arithmetic constraint formula
Yinrun Lyu, JingZheng Wu, Changyou Zhang, Nasro Min-Allah, Jamal Alhiyafi, Yongji Wang 0002
J. Glob. Optim.5
2017 A car-face region-based image retrieval method with attention of SIFT features
Changyou Zhang
Multim. Tools Appl.1
2016 Proactive service selection based on acquaintance model and LS-SVM
Changyou Zhang
Neurocomputing3
2015 UniDegree: A GPU-Based Graph Representation for SSSP
Changyou Zhang, Zhiyou Liu
ICA3PP (2)1
2010 Auto-tuning Dense Matrix Multiplication for GPGPU with Cache
abstract
In this paper we discuss about our experiences in improving the performance of GEMM (both single and double precision) on Fermi architecture using CUDA, and how the new features of Fermi such as cache affect performance. It is found that the addition of cache in GPU on one hand helps the processers take advantage of data locality occurred in runtime but on the other hand renders the dependency of performance on algorithmic parameters less predictable. Auto tuning then becomes a useful technique to address this issue. Our auto-tuned SGEMM and DGEMM reach 563 GFlops and 253 GFlops respectively on Tesla C2050. The design and implementation entirely use CUDA and C and have not benefited from tuning at the level of binary code.
Xiang Cui, Changyou Zhang, Hong Mei 0001
ICPADS3