Joo Seong Jeong

dblp:163/1557 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4Computer networks · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 41% GPUs and heterogeneous computing · 36% Parallel and multicore computing · 16%
Artificial intelligence
4 papers
Efficient and distributed learning · 57% Deep learning architectures and training · 24% Optimization for machine learning · 19%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 59% Programming languages and type systems · 41%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
0.922022
Hippo: Sharing Computations in Hyper-Parameter Optimization · Proc. VLDB Endow. 2022
Apache REEF: Retainable Evaluator Execution Framework · ACM Trans. Comput. Syst. 2017
Machine learning › Efficient and distributed learning
computation reuse
0.612022
Hippo: Sharing Computations in Hyper-Parameter Optimization · Proc. VLDB Endow. 2022
Machine learning › Optimization for machine learning
hyperparameter optimization
0.612022
Hippo: Sharing Computations in Hyper-Parameter Optimization · Proc. VLDB Endow. 2022
Machine learning and data management
inference serving
0.612022
Orca: A Distributed Serving System for Transformer-Based Generative Models · OSDI 2022
Parallel and multicore computing › task scheduling › task graph scheduling
critical path scheduling
0.612022
Hippo: Sharing Computations in Hyper-Parameter Optimization · Proc. VLDB Endow. 2022
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous mobile processors
0.612022
Band: coordinated multi-DNN inference on heterogeneous mobile processors · MobiSys 2022
Cloud and datacenter computing
inference serving
0.612022
Orca: A Distributed Serving System for Transformer-Based Generative Models · OSDI 2022
GPUs and heterogeneous computing › deep learning on GPUs
multi-DNN inference
0.612022
Band: coordinated multi-DNN inference on heterogeneous mobile processors · MobiSys 2022
Machine learning › Efficient and distributed learning › distributed training
data parallel training
0.412019
Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks · EuroSys 2019
Machine learning › Efficient and distributed learning
distributed training
0.412019
Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks · EuroSys 2019
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
parameter server
0.412019
Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks · EuroSys 2019
Compilers and program optimization
deep learning compiler
0.412019
JANUS: Fast and Flexible Deep Learning via Symbolic Graph Execution of Imperative Programs · NSDI 2019
Machine learning › Deep learning architectures and training
recursive neural network
0.312018
Improving the expressiveness of deep learning frameworks with recursion · EuroSys 2018
Programming languages and type systems › control structures
recursion
0.312018
Improving the expressiveness of deep learning frameworks with recursion · EuroSys 2018
Embedded and real-time systems › on-device inference
mobile inference
0.212022
Band: coordinated multi-DNN inference on heterogeneous mobile processors · MobiSys 2022
GPUs and heterogeneous computing
GPU training
0.112019
Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks · EuroSys 2019
Distributed systems
fault tolerance
0.112017
Apache REEF: Retainable Evaluator Execution Framework · ACM Trans. Comput. Syst. 2017

Methods — techniques the papers use, named apart from their topics

transformer · 1.1stage tree · 1.1search plan · 1.1generative model · 1.1sparsity-aware communication · 0.8hybrid parameter server and allreduce · 0.8operator partitioning · 0.6coordinated scheduling · 0.6collaborative filtering · 0.3
YearPublicationVenuePosition
2022 Band: coordinated multi-DNN inference on heterogeneous mobile processors
abstract
The rapid development of deep learning algorithms, as well as innovative hardware advancements, encourages multi-DNN workloads such as augmented reality applications. However, existing mobile inference frameworks like TensorFlow Lite and MNN fail to efficiently utilize heterogeneous processors available on mobile platforms, because they focus on running a single DNN on a specific processor. As mobile processors are too resource-limited to deliver reasonable performance for such workloads by their own, it is challenging to serve multi-DNN workloads with existing frameworks.
Joo Seong Jeong, Jingyu Lee, Changmin Jeon, Changjin Jeong, Youngki Lee 0001, Byung-Gon Chun
MobiSys1
2022 Orca: A Distributed Serving System for Transformer-Based Generative Models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun
OSDI2
2022 Hippo: Sharing Computations in Hyper-Parameter Optimization
abstract
Hyper-parameter optimization is crucial for pushing the accuracy of a deep learning model to its limits. However, a hyper-parameter optimization job, referred to as a study, involves numerous trials of training a model using different training knobs, and therefore is very computation-heavy, typically taking hours and days to finish. We observe that trials issued from hyper-parameter optimization algorithms often share common hyper-parameter sequence prefixes. Based on this observation, we propose Hippo, a hyper-parameter optimization system that reuses computation across trials to reduce the overall amount of computation significantly. Instead of treating each trial independently as in existing hyper-parameter optimization systems, Hippo breaks down the hyper-parameter sequences into stages and merges common stages to form a tree of stages (a stage tree). Hippo maintains an internal data structure, search plan, to manage the current status and history of a study, and employs a critical path based scheduler to minimize the overall study completion time. Hippo applies to not only single studies but multi-study scenarios as well. Evaluations show that Hippo's stage-based execution strategy outperforms trial-based methods for several models and hyper-parameter optimization algorithms, reducing end-to-end training time by up to 2.76X (3.53x) and GPU-hours by up to 4.81X (6.77x), for single (multiple) studies.
Ahnjae Shin, Joo Seong Jeong, Do Yoon Kim, Soyoung Jung, Byung-Gon Chun
Proc. VLDB Endow.2
2019 Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks
abstract
The employment of high-performance servers and GPU accelerators for training deep neural network models have greatly accelerated recent advances in deep learning (DL). DL frameworks, such as TensorFlow, MXNet, and Caffe2, have emerged to assist DL researchers to train their models in a distributed manner. Although current DL frameworks scale well for image classification models, there remain opportunities for scalable distributed training on natural language processing (NLP) models. We found that current frameworks show relatively low scalability on training NLP models due to the lack of consideration to the difference in sparsity of model parameters. In this paper, we propose Parallax, a framework that optimizes data parallel training by utilizing the sparsity of model parameters. Parallax introduces a hybrid approach that combines Parameter Server and AllReduce architectures to optimize the amount of data transfer according to the sparsity. Experiments show that Parallax built atop Tensor-Flow achieves scalable training throughput on both dense and sparse models while requiring little effort from its users. Parallax achieves up to 2.8x, 6.02x speedup for NLP models than TensorFlow and Horovod with 48 GPUs, respectively. The training speed for the image classification models is equal to Horovod and 1.53x faster than TensorFlow.
Soojeong Kim, Gyeong-In Yu, Hojin Park, Sungwoo Cho, Eunji Jeong, Hyeonmin Ha, Sanha Lee, Joo Seong Jeong, Byung-Gon Chun
EuroSys8
2019 Automating System Configuration of Distributed Machine Learning
abstract
The performance of distributed machine learning systems is dependent on their system configuration. However, configuring the system for optimal performance is challenging and time consuming even for experts due to the diverse runtime factors such as workloads or the system environment. We present cost-based optimization to automatically find a good system configuration for parameter server (PS) machine learning (ML) frameworks. We design and implement Cruise that applies the optimization technique to tune distributed PS ML execution automatically. Evaluation results on three ML applications verify that Cruise automates the system configuration of the applications to achieve good performance with minor reconfiguration costs.
Woo-Yeon Lee, Markus Weimer, Byung-Gon Chun, Yunseong Lee, Joo Seong Jeong, Gyeong-In Yu, Hojin Park, Beomyeol Jeon, Won Wook Song, Gunhee Kim
ICDCS6
2019 JANUS: Fast and Flexible Deep Learning via Symbolic Graph Execution of Imperative Programs
Eunji Jeong, Sungwoo Cho, Gyeong-In Yu, Joo Seong Jeong, Dongjin Shin, Byung-Gon Chun
NSDI4
2018 Improving the expressiveness of deep learning frameworks with recursion
abstract
Recursive neural networks have widely been used by researchers to handle applications with recursively or hierarchically structured data. However, embedded control flow deep learning frameworks such as TensorFlow, Theano, Caffe2, and MXNet fail to efficiently represent and execute such neural networks, due to lack of support for recursion. In this paper, we add recursion to the programming model of existing frameworks by complementing their design with recursive execution of dataflow graphs as well as additional APIs for recursive definitions. Unlike iterative implementations, which can only understand the topological index of each node in recursive data structures, our recursive implementation is able to exploit the recursive relationships between nodes for efficient execution based on parallel computation. We present an implementation on TensorFlow and evaluation results with various recursive neural network models, showing that our recursive implementation not only conveys the recursive nature of recursive neural networks better than other implementations, but also uses given resources more effectively to reduce training and inference time.
Eunji Jeong, Joo Seong Jeong, Soojeong Kim, Gyeong-In Yu, Byung-Gon Chun
EuroSys2
2017 Apache REEF: Retainable Evaluator Execution Framework
abstract
Resource Managers like YARN and Mesos have emerged as a critical layer in the cloud computing system stack, but the developer abstractions for leasing cluster resources and instantiating application logic are very low level. This flexibility comes at a high cost in terms of developer effort, as each application must repeatedly tackle the same challenges (e.g., fault tolerance, task scheduling and coordination) and reimplement common mechanisms (e.g., caching, bulk-data transfers). This article presents REEF, a development framework that provides a control plane for scheduling and coordinating task-level (data-plane) work on cluster resources obtained from a Resource Manager. REEF provides mechanisms that facilitate resource reuse for data caching and state management abstractions that greatly ease the development of elastic data processing pipelines on cloud platforms that support a Resource Manager service. We illustrate the power of REEF by showing applications built atop: a distributed shell application, a machine-learning framework, a distributed in-memory caching system, and a port of the CORFU system. REEF is currently an Apache top-level project that has attracted contributors from several institutions and it is being used to develop several commercial offerings such as the Azure Stream Analytics service.
Byung-Gon Chun, Tyson Condie, Yingda Chen, Carlo Curino, Chris Douglas, Matteo Interlandi, Beomyeol Jeon, Joo Seong Jeong, Gyewon Lee, Yunseong Lee, Tony Majestro, Dahlia Malkhi, Sergiy Matusevych, Brandon Myers, Mariia Mykhailova, Shravan M. Narayanamurthy, Joseph Noor, Raghu Ramakrishnan 0001, Sriram Rao, Russell Sears, Beysim Sezgin, Taegeon Um, Julia Wang, Markus Weimer, Youngseok Yang
ACM Trans. Comput. Syst.10
2015 Elastic Memory: Bring Elasticity Back to In-Memory Big Data Analytics
Joo Seong Jeong, Woo-Yeon Lee, Yunseong Lee, Youngseok Yang, Byung-Gon Chun
HotOS1