Yangzihao Wang

dblp:157/4935 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 56% GPUs and heterogeneous computing · 44%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel algorithms › graph algorithms
breadth-first search
0.522016
Gunrock: a high-performance graph processing library on the GPU · PPoPP 2016
Gunrock: a high-performance graph processing library on the GPU · PPoPP 2015
GPUs and heterogeneous computing
GPU graph processing
0.522016
Gunrock: a high-performance graph processing library on the GPU · PPoPP 2016
Gunrock: a high-performance graph processing library on the GPU · PPoPP 2015
Machine learning › Efficient and distributed learning
distributed training
0.412020
Distributed Equivalent Substitution Training for Large-Scale Recommender Systems · SIGIR 2020
Recommender systems
click-through rate prediction
0.412020
Distributed Equivalent Substitution Training for Large-Scale Recommender Systems · SIGIR 2020
Recommender systems
neural recommendation
0.412020
Distributed Equivalent Substitution Training for Large-Scale Recommender Systems · SIGIR 2020
Parallel and multicore computing › graph processing
parallel graph analytics
0.122016
Gunrock: a high-performance graph processing library on the GPU · PPoPP 2016
Gunrock: a high-performance graph processing library on the GPU · PPoPP 2015
Recommender systems
large-scale recommendation
0.112020
Distributed Equivalent Substitution Training for Large-Scale Recommender Systems · SIGIR 2020

Methods — techniques the papers use, named apart from their topics

sparse feature aggregation · 0.9distributed equivalent substitution · 0.9frontier-based abstraction · 0.5bulk-synchronous abstraction · 0.5
YearPublicationVenuePosition
2020 Distributed Equivalent Substitution Training for Large-Scale Recommender Systems
abstract
We present Distributed Equivalent Substitution (DES) training, a novel distributed training framework for large-scale recommender systems with dynamic sparse features. DES introduces fully synchronous training to large-scale recommendation system for the first time by reducing communication, thus making the training of commercial recommender systems converge faster and reach better CTR. DES requires much less communication by substituting the weights-rich operators with the computationally equivalent sub-operators and aggregating partial results instead of transmitting the huge sparse weights directly through the network. Due to the use of synchronous training on large-scale Deep Learning Recommendation Models (DLRMs), DES achieves higher AUC(Area Under ROC). We successfully apply DES training on multiple popular DLRMs of industrial scenarios. Experiments show that our implementation outperforms the state-of-the-art PS-based training framework, achieving up to 68.7% communication savings and higher throughput compared to other PS-based recommender systems.
Haidong Rong, Yangzihao Wang, Feihu Zhou, Junjie Zhai, Rui Lan, Yuekui Yang
SIGIR2
2017 Multi-GPU Graph Analytics
abstract
We present a single-node, multi-GPU programmable graph processing library that allows programmers to easily extend single-GPU graph algorithms to achieve scalable performance on large graphs with billions of edges. Directly using the single-GPU implementations, our design only requires programmers to specify a few algorithm-dependent concerns, hiding most multi-GPU related implementation details. We analyze the theoretical and practical limits to scalability in the context of varying graph primitives and datasets. We describe several optimizations, such as direction optimizing traversal, and a just-enough memory allocation scheme, for better performance and smaller memory consumption. Compared to previous work, we achieve best-of-class performance across operations and datasets, including excellent strong and weak scalability on most primitives as we increase the number of GPUs in the system
Yuechao Pan, Yangzihao Wang, Yuduo Wu, Carl Yang 0002, John D. Owens
IPDPS2
2016 Gunrock: a high-performance graph processing library on the GPU
abstract
For large-scale graph analytics on the GPU, the irregularity of data access/control flow and the complexity of programming GPUs have been two significant challenges for developing a programmable high-performance graph library. "Gunrock," our high-level bulk-synchronous graph-processing system targeting the GPU, takes a new approach to abstracting GPU graph analytics: rather than designing an abstraction around computation, Gunrock instead implements a novel data-centric abstraction centered on operations on a vertex or edge frontier. Gunrock achieves a balance between performance and expressiveness by coupling high-performance GPU computing primitives and optimization strategies with a high-level programming model that allows programmers to quickly develop new graph primitives with small code size and minimal GPU programming knowledge. We evaluate Gunrock on five graph primitives (BFS, BC, SSSP, CC, and PageRank) and show that Gunrock has on average at least an order of magnitude speedup over Boost and PowerGraph, comparable performance to the fastest GPU hardwired primitives, and better performance than any other GPU high-level graph library.
Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andrew Riffel, John D. Owens
PPoPP1
2015 Gunrock: a high-performance graph processing library on the GPU
abstract
For large-scale graph analytics on the GPU, the irregularity of data access and control flow and the complexity of programming GPUs have been two significant challenges for developing a programmable high-performance graph library. "Gunrock", our graph-processing system, uses a high-level bulk-synchronous abstraction with traversal and computation steps, designed specifically for the GPU. Gunrock couples high performance with a high-level programming model that allows programmers to quickly develop new graph primitives with less than 300 lines of code. We evaluate Gunrock on five graph primitives and show that Gunrock has at least an order of magnitude speedup over Boost and PowerGraph, comparable performance to the fastest GPU hardwired primitives, and better performance than any other GPU high-level graph library.
Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andrew Riffel, John D. Owens
PPoPP1