EDBT 2026 Demo / reviewers in the wild / expert
Jinliang Wei
dblp:124/4315
· DBLP profile ↗
12ranked-venue papers
2as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Distributed systems · 47% Parallel and multicore computing · 31% Cloud and datacenter computing · 15% | |
| Artificial intelligence
6 papers |
Efficient and distributed learning · 72% Graph learning · 22% Language models and text generation · 4% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
1.5 | 4 | 2023 | Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023 Automating Dependence-Aware Parallelization of Machine Learning Training on Distributed Shared Memory · EuroSys 2019 LightLDA: Big Topic Models on Modest Computer Clusters · WWW 2015 |
Parallel and multicore computing
parallel programming models |
1.0 | 2 | 2023 | Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023 Automating Dependence-Aware Parallelization of Machine Learning Training on Distributed Shared Memory · EuroSys 2019 |
Machine learning › Efficient and distributed learning › distributed training
model parallelism |
0.7 | 1 | 2023 | Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023 |
Distributed systems › communication optimization
communication-computation overlap |
0.7 | 1 | 2023 | Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023 |
Machine learning › Graph learning › graph neural network training
distributed GNN training |
0.5 | 1 | 2021 | Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads · OSDI 2021 |
Machine learning › Graph learning
graph neural network training |
0.5 | 1 | 2021 | Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads · OSDI 2021 |
Cloud and datacenter computing
serverless computing |
0.5 | 1 | 2021 | Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads · OSDI 2021 |
Machine learning › Efficient and distributed learning › distributed training
data parallel training |
0.4 | 1 | 2019 | Automating Dependence-Aware Parallelization of Machine Learning Training on Distributed Shared Memory · EuroSys 2019 |
Parallel and multicore computing › parallel programming models
automatic parallelization |
0.4 | 1 | 2019 | Automating Dependence-Aware Parallelization of Machine Learning Training on Distributed Shared Memory · EuroSys 2019 |
Distributed systems › distributed system architecture
communication architecture |
0.3 | 1 | 2017 | Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters · USENIX ATC 2017 |
Distributed systems › distributed machine learning
distributed deep learning |
0.3 | 1 | 2017 | Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters · USENIX ATC 2017 |
Distributed systems › distributed machine learning
distributed training communication |
0.3 | 1 | 2017 | Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters · USENIX ATC 2017 |
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster |
0.3 | 1 | 2017 | Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters · USENIX ATC 2017 |
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
parameter server |
0.3 | 2 | 2015 | High-Performance Distributed ML at Scale through Parameter Server Consistency Models · AAAI 2015 Petuum: A New Platform for Distributed Machine Learning on Big Data · KDD 2015 |
Machine learning › Efficient and distributed learning › distributed training
model and data parallelism |
0.2 | 1 | 2015 | Petuum: A New Platform for Distributed Machine Learning on Big Data · KDD 2015 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2015 | LightLDA: Big Topic Models on Modest Computer Clusters · WWW 2015 |
Distributed systems
consistency models |
0.2 | 1 | 2015 | High-Performance Distributed ML at Scale through Parameter Server Consistency Models · AAAI 2015 |
Distributed systems
distributed coordination |
0.2 | 1 | 2015 | High-Performance Distributed ML at Scale through Parameter Server Consistency Models · AAAI 2015 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2023 | Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023 |
Cloud and datacenter computing
big data analytics |
0.2 | 1 | 2014 | Exploiting Bounded Staleness to Speed Up Big Data Analytics · USENIX ATC 2014 |
Distributed systems › consistency models
bounded staleness |
0.2 | 1 | 2014 | Exploiting Bounded Staleness to Speed Up Big Data Analytics · USENIX ATC 2014 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2015 | LightLDA: Big Topic Models on Modest Computer Clusters · WWW 2015 |
Methods — techniques the papers use, named apart from their topics
intra-layer model parallelism · 1.3computation decomposition · 1.3distributed shared memory · 1.1dependence analysis · 1.1serverless computing · 1.0distributed training · 1.0theoretical convergence analysis · 0.4eager communication · 0.4asynchronous parameter updates · 0.4bounded-latency network synchronization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning ModelsabstractLarge deep learning models have shown great potential with state-of-the-art results in many tasks. However, running these large models is quite challenging on an accelerator (GPU or TPU) because the on-device memory is too limited for the size of these models. Intra-layer model parallelism is an approach to address the issues by partitioning individual layers or operators across multiple devices in a distributed accelerator cluster. But, the data communications generated by intra-layer model parallelism can contribute to a significant proportion of the overall execution time and severely hurt the computational efficiency. Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hechtman, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Tongfei Guo, Yuanzhong Xu, Zongwei Zhou |
ASPLOS (1) | 2 |
| 2021 | Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads
John Thorpe, Yifan Qiao 0002, Jon Eyolfson, Shen Teng, Guanzhou Hu, Jinliang Wei, Keval Vora, Ravi Netravali, Miryung Kim, Guoqing Harry Xu |
OSDI | 7 |
| 2019 | Automating Dependence-Aware Parallelization of Machine Learning Training on Distributed Shared MemoryabstractMachine learning (ML) training is commonly parallelized using data parallelism. A fundamental limitation of data parallelism is that conflicting (concurrent) parameter accesses during ML training usually diminishes or even negates the benefits provided by additional parallel compute resources. Although it is possible to avoid conflicting parameter accesses by carefully scheduling the computation, existing systems rely on programmer manual parallelization and it remains a question when such parallelization is possible. Jinliang Wei, Garth A. Gibson, Phillip B. Gibbons, Eric P. Xing |
EuroSys | 1 |
| 2017 | Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters
Hao Zhang 0025, Shizhen Xu, Wei Dai 0003, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, Eric P. Xing |
USENIX ATC | 8 |
| 2016 | Addressing the straggler problem for iterative convergent parallel MLabstractFlexRR provides a scalable, efficient solution to the straggler problem for iterative machine learning (ML). The frequent (e.g., per iteration) barriers used in traditional BSP-based distributed ML implementations cause every transient slowdown of any worker thread to delay all others. FlexRR combines a more flexible synchronization model with dynamic peer-to-peer re-assignment of work among workers to address straggler threads. Experiments with real straggler behavior observed on Amazon EC2 and Microsoft Azure, as well as injected straggler behavior stress tests, confirm the significance of the problem and the effectiveness of FlexRR's solution. Using FlexRR, we consistently observe near-ideal run-times (relative to no performance jitter) across all real and injected straggler behaviors tested. Aaron Harlap, Henggang Cui, Wei Dai 0003, Jinliang Wei, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, Eric P. Xing |
SoCC | 4 |
| 2015 | High-Performance Distributed ML at Scale through Parameter Server Consistency ModelsabstractAs Machine Learning (ML) applications embrace greater data size and model complexity, practitioners turn to distributed clusters to satisfy the increased computational and memory demands. Effective use of clusters for ML programs requires considerable expertise in writing distributed code, but existing highly-abstracted frameworks like Hadoop that pose low barriers to distributed-programming have not, in practice, matched the performance seen in highly specialized and advanced ML implementations. The recent Parameter Server (PS) paradigm is a middle ground between these extremes, allowing easy conversion of single-machine parallel ML programs into distributed ones, while maintaining high throughput through relaxed ``consistency models" that allow asynchronous (and, hence, inconsistent) parameter reads. However, due to insufficient theoretical study, it is not clear which of these consistency models can really ensure correct ML algorithm output; at the same time, there remain many theoretically-motivated but undiscovered opportunities to maximize computational throughput. Inspired by this challenge, we study both the theoretical guarantees and empirical behavior of iterative-convergent ML algorithms in existing PS consistency models. We then use the gleaned insights to improve a consistency model using an "eager" PS communication mechanism, and implement it as a new PS system that enables ML programs to reach their solution more quickly. Wei Dai 0003, Abhimanu Kumar, Jinliang Wei, Qirong Ho, Garth A. Gibson, Eric P. Xing |
AAAI | 3 |
| 2015 | Managed communication and consistency for fast data-parallel iterative analyticsabstractAt the core of Machine Learning (ML) analytics is often an expert-suggested model, whose parameters are refined by iteratively processing a training dataset until convergence. The completion time (i.e. convergence time) and quality of the learned model not only depends on the rate at which the refinements are generated but also the quality of each refinement. While data-parallel ML applications often employ a loose consistency model when updating shared model parameters to maximize parallelism, the accumulated error may seriously impact the quality of refinements and thus delay completion time, a problem that usually gets worse with scale. Although more immediate propagation of updates reduces the accumulated error, this strategy is limited by physical network bandwidth. Additionally, the performance of the widely used stochastic gradient descent (SGD) algorithm is sensitive to step size. Simply increasing communication often fails to bring improvement without tuning step size accordingly and tedious hand tuning is usually needed to achieve optimal performance. Jinliang Wei, Wei Dai 0003, Aurick Qiao, Qirong Ho, Henggang Cui, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, Eric P. Xing |
SoCC | 1 |
| 2015 | Petuum: A New Platform for Distributed Machine Learning on Big DataabstractHow can one build a distributed framework that allows efficient deployment of a wide spectrum of modern advanced machine learning (ML) programs for industrial-scale problems using Big Models (100s of billions of parameters) on Big Data (terabytes or petabytes)- Contemporary parallelization strategies employ fine-grained operations and scheduling beyond the classic bulk-synchronous processing paradigm popularized by MapReduce, or even specialized operators relying on graphical representations of ML programs. The variety of approaches tends to pull systems and algorithms design in different directions, and it remains difficult to find a universal platform applicable to a wide range of different ML programs at scale. We propose a general-purpose framework that systematically addresses data- and model-parallel challenges in large-scale ML, by leveraging several fundamental properties underlying ML programs that make them different from conventional operation-centric programs: error tolerance, dynamic structure, and nonuniform convergence; all stem from the optimization-centric nature shared in ML programs' mathematical definitions, and the iterative-convergent behavior of their algorithmic solutions. These properties present unique opportunities for an integrative system design, built on bounded-latency network synchronization and dynamic load-balancing scheduling, which is efficient, programmable, and enjoys provable correctness guarantees. We demonstrate how such a design in light of ML-first principles leads to significant performance improvements versus well-known implementations of several ML programs, allowing them to run in much less time and at considerably larger model sizes, on modestly-sized computer clusters. Eric P. Xing, Qirong Ho, Wei Dai 0003, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, Yaoliang Yu |
KDD | 5 |
| 2015 | LightLDA: Big Topic Models on Modest Computer ClustersabstractWhen building large-scale machine learning (ML) programs, such as massive topic models or deep neural networks with up to trillions of parameters and training examples, one usually assumes that such massive tasks can only be attempted with industrial-sized clusters with thousands of nodes, which are out of reach for most practitioners and academic researchers. We consider this challenge in the context of topic modeling on web-scale corpora, and show that with a modest cluster of as few as 8 machines, we can train a topic model with 1 million topics and a 1-million-word vocabulary (for a total of 1 trillion parameters), on a document collection with 200 billion tokens --- a scale not yet reported even with thousands of machines. Our major contributions include: 1) a new, highly-efficient O(1) Metropolis-Hastings sampling algorithm, whose running cost is (surprisingly) agnostic of model size, and empirically converges nearly an order of magnitude more quickly than current state-of-the-art Gibbs samplers; 2) a model-scheduling scheme to handle the big model challenge, where each worker machine schedules the fetch/use of sub-models as needed, resulting in a frugal use of limited memory capacity and network bandwidth; 3) a differential data-structure for model storage, which uses separate data structures for high- and low-frequency words to allow extremely large models to fit in memory, while maintaining high inference speed. These contributions are built on top of the Petuum open-source distributed ML framework, and we provide experimental evidence showing how this development puts massive data and models within reach on a small cluster, while still enjoying proportional time cost reductions with increasing cluster size. Jinhui Yuan, Fei Gao 0018, Qirong Ho, Wei Dai 0003, Jinliang Wei, Xun Zheng, Eric P. Xing, Tie-Yan Liu, Wei-Ying Ma |
WWW | 5 |
| 2015 | Petuum: A New Platform for Distributed Machine Learning on Big DataabstractWhat is a systematic way to efficiently apply a wide spectrum of advanced ML programs to industrial scale problems, using Big Models (up to 100 s of billions of parameters) on Big Data (up to terabytes or petabytes)? Modern parallelization strategies employ fine-grained operations and scheduling beyond the classic bulk-synchronous processing paradigm popularized by MapReduce, or even specialized graph-based execution that relies on graph representations of ML programs. The variety of approaches tends to pull systems and algorithms design in different directions, and it remains difficult to find a universal platform applicable to a wide range of ML programs at scale. We propose a general-purpose framework, Petuum, that systematically addresses data- and model-parallel challenges in large-scale ML, by observing that many ML programs are fundamentally optimization-centric and admit error-tolerant, iterative-convergent algorithmic solutions. This presents unique opportunities for an integrative system design, such as bounded-error network synchronization and dynamic scheduling based on ML program structure. We demonstrate the efficacy of these system designs versus well-known implementations of modern ML algorithms, showing that Petuum allows ML programs to run in much less time and at considerably larger model sizes, even on modestly-sized compute clusters. Eric P. Xing, Qirong Ho, Wei Dai 0003, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, Yaoliang Yu |
IEEE Trans. Big Data | 5 |
| 2014 | Exploiting iterative-ness for parallel ML computationsabstractMany large-scale machine learning (ML) applications use iterative algorithms to converge on parameter values that make the chosen model fit the input data. Often, this approach results in the same sequence of accesses to parameters repeating each iteration. This paper shows that these repeating patterns can and should be exploited to improve the efficiency of the parallel and distributed ML applications that will be a mainstay in cloud computing environments. Focusing on the increasingly popular "parameter server" approach to sharing model parameters among worker threads, we describe and demonstrate how the repeating patterns can be exploited. Examples include replacing dynamic cache and server structures with static pre-serialized structures, informing prefetch and partitioning decisions, and determining which data should be cached at each thread to avoid both contention and slow accesses to memory banks attached to other sockets. Experiments show that such exploitation reduces per-iteration time by 33--98%, for three real ML workloads, and that these improvements are robust to variation in the patterns over time. Henggang Cui, Alexey Tumanov, Jinliang Wei, Lianghong Xu, Wei Dai 0003, Jesse Haber-Kucharsky, Qirong Ho, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, Eric P. Xing |
SoCC | 3 |
| 2014 | Exploiting Bounded Staleness to Speed Up Big Data Analytics
Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai 0003, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, Eric P. Xing |
USENIX ATC | 7 |