Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jae-Won Chung

dblp:261/9463 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0001-5924-3427ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Energy-efficient computing · 71% Performance modeling and evaluation · 20% Cloud and datacenter computing · 5%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.912025
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization · NeurIPS 2025
Energy-efficient computing
energy measurement
0.912025
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization · NeurIPS 2025
Machine learning › Efficient and distributed learning
distributed training
0.812024
Reducing Energy Bloat in Large Model Training · SOSP 2024
Energy-efficient computing
GPU power consumption
0.712023
Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training · NSDI 2023
Machine learning and data management
machine learning systems
0.312025
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization · NeurIPS 2025
GPUs and heterogeneous computing
GPU computing
0.212023
Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training · NSDI 2023

Methods — techniques the papers use, named apart from their topics

automated optimization recommendation · 1.7energy profiling · 1.5
YearPublicationVenuePosition
2025 The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
abstract
As the adoption of Generative AI in real-world services grow explosively, energy has emerged as a critical bottleneck resource. However, energy remains a metric that is often overlooked, under-explored, or poorly understood in the context of building ML systems. We present the ML.ENERGY Benchmark, a benchmark suite and tool for measuring inference energy consumption under realistic service environments, and the corresponding ML.ENERGY Leaderboard, which have served as a valuable resource for those hoping to understand and optimize the energy consumption of their generative AI services. In this paper, we explain four key design principles for benchmarking ML energy we have acquired over time, and then describe how they are implemented in the ML.ENERGY Benchmark. We then highlight results from the early 2025 iteration of the benchmark, including energy measurements of 40 widely used model architectures across 6 different tasks, case studies of how ML design choices impact energy consumption, and how automated optimization recommendations can lead to significant (sometimes more than 40%) energy savings without changing what is being computed by the model. The ML.ENERGY Benchmark is open-source and can be easily extended to various customized models and application scenarios.
Jae-Won Chung, Jeff J. Ma, Oh Jun Kweon, Yuxuan Xia, Zhiyu Wu, Mosharaf Chowdhury
NeurIPS1
2024 Reducing Energy Bloat in Large Model Training
abstract
Training large AI models on numerous GPUs consumes a massive amount of energy, making power delivery one of the largest limiting factors in building and operating datacenters for AI workloads. However, we observe that not all energy consumed during training directly contributes to end-to-end throughput; a significant portion can be removed without slowing down training. We call this portion energy bloat.
Jae-Won Chung, Yile Gu, Insu Jang, Luoxi Meng, Nikhil Bansal 0001, Mosharaf Chowdhury
SOSP1
2023 Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training
Jae-Won Chung, Mosharaf Chowdhury
NSDI2
2020 ShadowTutor: Distributed Partial Distillation for Mobile Video DNN Inference
abstract
Following the recent success of deep neural networks (DNN) on video computer vision tasks, performing DNN inferences on videos that originate from mobile devices has gained practical significance. As such, previous approaches developed methods to offload DNN inference computations for images to cloud servers to manage the resource constraints of mobile devices. However, when it comes to video data, communicating information of every frame consumes excessive network bandwidth and renders the entire system susceptible to adverse network conditions such as congestion. Thus, in this work, we seek to exploit the temporal coherence between nearby frames of a video stream to mitigate network pressure. That is, we propose ShadowTutor, a distributed video DNN inference framework that reduces the number of network transmissions through intermittent knowledge distillation to a student model. Moreover, we update only a subset of the student’s parameters, which we call partial distillation, to reduce the data size of each network transmission. Specifically, the server runs a large and general teacher model, and the mobile device only runs an extremely small but specialized student model. On sparsely selected key frames, the server partially trains the student model by targeting the teacher’s response and sends the updated part to the mobile device. We investigate the effectiveness of ShadowTutor with HD video semantic segmentation. Evaluations show that network data transfer is reduced by 95% on average. Moreover, the throughput of the system is improved by over three times and shows robustness to changes in network bandwidth.
Jae-Won Chung, Jae-Yun Kim, Soo-Mook Moon
ICPP1