Yichong Zhang

dblp:256/0378 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Graph learning · 30% Probabilistic and Bayesian machine learning · 20% Efficient and distributed learning · 16%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 44% Distributed systems · 22% Processor architecture and microarchitecture · 22%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph clustering
1.122022
Detecting Latent Communities in Network Formation Models · J. Mach. Learn. Res. 2022
Determining the Number of Communities in Degree-corrected Stochastic Block Models · J. Mach. Learn. Res. 2021
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › relational model
statistical network models
1.122022
Detecting Latent Communities in Network Formation Models · J. Mach. Learn. Res. 2022
Determining the Number of Communities in Degree-corrected Stochastic Block Models · J. Mach. Learn. Res. 2021
Distributed systems › communication optimization
communication-computation overlap
1.012026
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering · EuroSys 2026
GPUs and heterogeneous computing › GPU communication
inter-GPU communication
1.012026
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering · EuroSys 2026
Processor architecture and microarchitecture
latency hiding
1.012026
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering · EuroSys 2026
GPUs and heterogeneous computing
multi-GPU computing
1.012026
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering · EuroSys 2026
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention
0.912025
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models · NeurIPS 2025
Machine learning › Generative modeling
visual generation
0.912025
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models · NeurIPS 2025
Machine learning › Graph learning
stochastic block model
0.512021
Determining the Number of Communities in Degree-corrected Stochastic Block Models · J. Mach. Learn. Res. 2021
Data mining
clustering
0.412020
Strong Consistency of Spectral Clustering for Stochastic Block Models · IEEE Trans. Inf. Theory 2020
Data mining › structured data mining › graph mining
community detection
0.412020
Strong Consistency of Spectral Clustering for Stochastic Block Models · IEEE Trans. Inf. Theory 2020
Data mining › clustering
spectral clustering
0.412020
Strong Consistency of Spectral Clustering for Stochastic Block Models · IEEE Trans. Inf. Theory 2020
Data mining › structured data mining › graph mining › community detection
stochastic block model
0.412020
Strong Consistency of Spectral Clustering for Stochastic Block Models · IEEE Trans. Inf. Theory 2020
Parallel and multicore computing › parallel scheduling
communication scheduling
0.312026
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering · EuroSys 2026
Parallel and multicore computing
parallel programming models
0.312026
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering · EuroSys 2026
Mathematical optimization › continuous optimization › convex optimization › norm optimization
nuclear norm minimization
0.212022
Detecting Latent Communities in Network Formation Models · J. Mach. Learn. Res. 2022
Information theory › hypothesis testing
likelihood ratio test
0.112021
Determining the Number of Communities in Degree-corrected Stochastic Block Models · J. Mach. Learn. Res. 2021

Methods — techniques the papers use, named apart from their topics

spectral clustering · 2.6nuclear norm regularization · 1.1iterative logistic regression · 1.1tile-wise overlapping · 1.0signaling · 1.0reordering · 1.0pseudo likelihood ratio · 1.0binary segmentation · 1.0token reordering · 0.9sparsification · 0.9quantization · 0.9k-means · 0.4graph laplacian · 0.4
YearPublicationVenuePosition
2026 Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering
abstract
Generative models have achieved remarkable success across various applications, driving the demand for multi-GPU computing. Inter-GPU communication becomes a bottleneck in multi-GPU computing systems, particularly on consumer-grade GPUs. By exploiting concurrent hardware execution, overlapping computation and communication latency becomes an effective technique for mitigating the communication overhead. We identify that an efficient and adaptable overlapping design should satisfy (1) tile-wise overlapping to maximize the overlapping opportunity, (2) interference-free computation to maintain the original computational performance, and (3) communication agnosticism to reduce the development burden against varying communication primitives. Nevertheless, current designs fail to simultaneously optimize for all of those features.
Ke Hong, Minxu Liu, Qiuli Mao, Zixiao Huang 0001, Lufang Chen, Yichong Zhang, Zhenhua Zhu 0002, Guohao Dai 0001, Yu Wang 0002
EuroSys9
2025 PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models
abstract
In visual generation, the quadratic complexity of attention mechanisms results in high memory and computational costs, especially for longer token sequences required in high-resolution image or multi-frame video generation. To address this, prior research has explored techniques such as sparsification and quantization. However, these techniques face significant challenges under low density and reduced bitwidths. Through systematic analysis, we identify that the core difficulty stems from the dispersed and irregular characteristics of visual attention patterns. Therefore, instead of introducing specialized sparsification and quantization design to accommodate such patterns, we propose an alternative strategy: "reorganizing" the attention pattern to alleviate the challenges. Inspired by the local aggregatin nature of visual feature extraction, we design a novel **P**attern-**A**ware token **R**e**O**rdering (**PARO**) technique, which unifies the diverse attention patterns into a hardware-friendly block-wise pattern. This unification substantially simplifies and enhances both sparsification and quantization. We evaluate the performance-efficiency trade-offs of various design choices and finalize a methodology tailored for the unified pattern. Our approach, **PAROAttention**, achieves video and image generation with lossless metrics, and nearly identical results from full-precision (FP) baselines, while operating at notably lower density (**20%-30%**) and bitwidth (**INT8/INT4**), achieving a **1.9 - 2.7x** end-to-end latency speedup.
Tianchen Zhao, Ke Hong, Xuefeng Xiao 0001, Huixia Li, Ruiqi Xie, Yichong Zhang, Yu Wang 0002
NeurIPS10
2022 Detecting Latent Communities in Network Formation Models
abstract
This paper proposes a logistic undirected network formation model which allows for assortative matching on observed individual characteristics and the presence of edge-wise fixed effects. We model the coefficients of observed characteristics to have a latent community structure and the edge-wise fixed effects to be of low rank. We propose a multi-step estimation procedure involving nuclear norm regularization, sample splitting, iterative logistic regression and spectral clustering to detect the latent communities. We show that the latent communities can be exactly recovered when the expected degree of the network is of order logn or higher, where n is the number of nodes in the network. The finite sample performance of the new estimation and inference methods is illustrated through both simulated and real datasets.
Shujie Ma, Liangjun Su, Yichong Zhang
J. Mach. Learn. Res.3
2021 Determining the Number of Communities in Degree-corrected Stochastic Block Models
abstract
We propose to estimate the number of communities in degree-corrected stochastic block models based on a pseudo likelihood ratio statistic. To this end, we introduce a method that combines spectral clustering with binary segmentation. This approach guarantees an upper bound for the pseudo likelihood ratio statistic when the model is over-fitted. We also derive its limiting distribution when the model is under-fitted. Based on these properties, we establish the consistency of our estimator for the true number of communities. Developing these theoretical properties require a mild condition on the average degrees - growing at a rate no slower than log(n), where n is the number of nodes. Our proposed method is further illustrated by simulation studies and analysis of real-world networks. The numerical results show that our approach has satisfactory performance when the network is semi-dense.
Shujie Ma, Liangjun Su, Yichong Zhang
J. Mach. Learn. Res.3
2020 Strong Consistency of Spectral Clustering for Stochastic Block Models
abstract
In this paper we prove the strong consistency of several methods based on the spectral clustering techniques that are widely used to study the community detection problem in stochastic block models (SBMs). We show that under some weak conditions on the minimal degree, the number of communities, and the eigenvalues of the probability block matrix, the K-means algorithm applied to the eigenvectors of the graph Laplacian associated with its first few largest eigenvalues can classify all individuals into the true community uniformly correctly almost surely. Extensions to both regularized spectral clustering and degree-corrected SBMs are also considered. We illustrate the performance of different methods on simulated networks.
Liangjun Su, Wuyi Wang, Yichong Zhang
IEEE Trans. Inf. Theory3