EDBT 2026 Demo / reviewers in the wild / expert
Yangzihao Wang
dblp:157/4935
· DBLP profile ↗
4ranked-venue papers
2as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 56% GPUs and heterogeneous computing · 44% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel algorithms › graph algorithms
breadth-first search |
0.5 | 2 | 2016 | Gunrock: a high-performance graph processing library on the GPU · PPoPP 2016 Gunrock: a high-performance graph processing library on the GPU · PPoPP 2015 |
GPUs and heterogeneous computing
GPU graph processing |
0.5 | 2 | 2016 | Gunrock: a high-performance graph processing library on the GPU · PPoPP 2016 Gunrock: a high-performance graph processing library on the GPU · PPoPP 2015 |
Machine learning › Efficient and distributed learning
distributed training |
0.4 | 1 | 2020 | Distributed Equivalent Substitution Training for Large-Scale Recommender Systems · SIGIR 2020 |
Recommender systems
click-through rate prediction |
0.4 | 1 | 2020 | Distributed Equivalent Substitution Training for Large-Scale Recommender Systems · SIGIR 2020 |
Recommender systems
neural recommendation |
0.4 | 1 | 2020 | Distributed Equivalent Substitution Training for Large-Scale Recommender Systems · SIGIR 2020 |
Parallel and multicore computing › graph processing
parallel graph analytics |
0.1 | 2 | 2016 | Gunrock: a high-performance graph processing library on the GPU · PPoPP 2016 Gunrock: a high-performance graph processing library on the GPU · PPoPP 2015 |
Recommender systems
large-scale recommendation |
0.1 | 1 | 2020 | Distributed Equivalent Substitution Training for Large-Scale Recommender Systems · SIGIR 2020 |
Methods — techniques the papers use, named apart from their topics
sparse feature aggregation · 0.9distributed equivalent substitution · 0.9frontier-based abstraction · 0.5bulk-synchronous abstraction · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Distributed Equivalent Substitution Training for Large-Scale Recommender SystemsabstractWe present Distributed Equivalent Substitution (DES) training, a novel distributed training framework for large-scale recommender systems with dynamic sparse features. DES introduces fully synchronous training to large-scale recommendation system for the first time by reducing communication, thus making the training of commercial recommender systems converge faster and reach better CTR. DES requires much less communication by substituting the weights-rich operators with the computationally equivalent sub-operators and aggregating partial results instead of transmitting the huge sparse weights directly through the network. Due to the use of synchronous training on large-scale Deep Learning Recommendation Models (DLRMs), DES achieves higher AUC(Area Under ROC). We successfully apply DES training on multiple popular DLRMs of industrial scenarios. Experiments show that our implementation outperforms the state-of-the-art PS-based training framework, achieving up to 68.7% communication savings and higher throughput compared to other PS-based recommender systems. Haidong Rong, Yangzihao Wang, Feihu Zhou, Junjie Zhai, Rui Lan, Yuekui Yang |
SIGIR | 2 |
| 2017 | Multi-GPU Graph AnalyticsabstractWe present a single-node, multi-GPU programmable graph processing library that allows programmers to easily extend single-GPU graph algorithms to achieve scalable performance on large graphs with billions of edges. Directly using the single-GPU implementations, our design only requires programmers to specify a few algorithm-dependent concerns, hiding most multi-GPU related implementation details. We analyze the theoretical and practical limits to scalability in the context of varying graph primitives and datasets. We describe several optimizations, such as direction optimizing traversal, and a just-enough memory allocation scheme, for better performance and smaller memory consumption. Compared to previous work, we achieve best-of-class performance across operations and datasets, including excellent strong and weak scalability on most primitives as we increase the number of GPUs in the system Yuechao Pan, Yangzihao Wang, Yuduo Wu, Carl Yang 0002, John D. Owens |
IPDPS | 2 |
| 2016 | Gunrock: a high-performance graph processing library on the GPUabstractFor large-scale graph analytics on the GPU, the irregularity of data access/control flow and the complexity of programming GPUs have been two significant challenges for developing a programmable high-performance graph library. "Gunrock," our high-level bulk-synchronous graph-processing system targeting the GPU, takes a new approach to abstracting GPU graph analytics: rather than designing an abstraction around computation, Gunrock instead implements a novel data-centric abstraction centered on operations on a vertex or edge frontier. Gunrock achieves a balance between performance and expressiveness by coupling high-performance GPU computing primitives and optimization strategies with a high-level programming model that allows programmers to quickly develop new graph primitives with small code size and minimal GPU programming knowledge. We evaluate Gunrock on five graph primitives (BFS, BC, SSSP, CC, and PageRank) and show that Gunrock has on average at least an order of magnitude speedup over Boost and PowerGraph, comparable performance to the fastest GPU hardwired primitives, and better performance than any other GPU high-level graph library. Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andrew Riffel, John D. Owens |
PPoPP | 1 |
| 2015 | Gunrock: a high-performance graph processing library on the GPUabstractFor large-scale graph analytics on the GPU, the irregularity of data access and control flow and the complexity of programming GPUs have been two significant challenges for developing a programmable high-performance graph library. "Gunrock", our graph-processing system, uses a high-level bulk-synchronous abstraction with traversal and computation steps, designed specifically for the GPU. Gunrock couples high performance with a high-level programming model that allows programmers to quickly develop new graph primitives with less than 300 lines of code. We evaluate Gunrock on five graph primitives and show that Gunrock has at least an order of magnitude speedup over Boost and PowerGraph, comparable performance to the fastest GPU hardwired primitives, and better performance than any other GPU high-level graph library. Yangzihao Wang, Andrew A. Davidson, Yuechao Pan, Yuduo Wu, Andrew Riffel, John D. Owens |
PPoPP | 1 |