VLDB 2026 Research / reviewers in the wild / expert
Jongseok Park 0002
dblp:137/1527-2 · also Jong Seok Park 0002
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0003-3910-7182ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 31% Embedded and real-time systems · 31% GPUs and heterogeneous computing · 21% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 100% | |
| Computer networks
1 paper |
Edge and fog computing · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed inference |
0.8 | 1 | 2024 | CoActo: CoActive Neural Network Inference Offloading with Fine-grained and Concurrent Execution · MobiSys 2024 |
Edge and fog computing › edge inference
collaborative inference |
0.8 | 1 | 2024 | CoActo: CoActive Neural Network Inference Offloading with Fine-grained and Concurrent Execution · MobiSys 2024 |
Machine learning › Efficient and distributed learning
inference systems |
0.7 | 1 | 2023 | ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks · NeurIPS 2023 |
Parallel and multicore computing › task scheduling
dynamic scheduling |
0.7 | 1 | 2023 | ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks · NeurIPS 2023 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
convolution optimization |
0.6 | 1 | 2022 | mGEMM: low-latency convolution with minimal memory overhead optimized for mobile devices · MobiSys 2022 |
Embedded and real-time systems
low-latency inference |
0.6 | 1 | 2022 | mGEMM: low-latency convolution with minimal memory overhead optimized for mobile devices · MobiSys 2022 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.6 | 1 | 2022 | mGEMM: low-latency convolution with minimal memory overhead optimized for mobile devices · MobiSys 2022 |
GPUs and heterogeneous computing › embedded GPU
mobile GPU |
0.6 | 1 | 2022 | mGEMM: low-latency convolution with minimal memory overhead optimized for mobile devices · MobiSys 2022 |
Embedded and real-time systems › on-device inference
mobile inference |
0.6 | 1 | 2022 | mGEMM: low-latency convolution with minimal memory overhead optimized for mobile devices · MobiSys 2022 |
Compilers and program optimization › intermediate representation
dataflow graph |
0.2 | 1 | 2023 | ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks · NeurIPS 2023 |
Compilers and program optimization
deep learning compiler |
0.2 | 1 | 2023 | ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
tiling · 2.0dynamic scheduling · 2.0data flow graph · 1.3data flow graphs · 0.7convolution · 0.6GEMM · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CoActo: CoActive Neural Network Inference Offloading with Fine-grained and Concurrent ExecutionabstractCollaborative inference is the current state-of-the-art solution for mobile-server neural network inference offloading. However, we find that existing collaborative inference solutions only focus on partitioning the DNN computation, which is only a small part of achieving an efficient DNN offloading system. What ultimately determines the performance of DNN offloading is how the execution system utilizes the characteristics of the given DNN offloading task on the mobile, network, and server resources of the offloading environment. To this end, we design CoActo, a DNN execution system built from the ground up for mobile-server inference offloading. Our key design philosophy is Coactive Inference Offloading, which is a new, improved concept of DNN offloading that adds two properties, 1) fine-grained expression of DNNs and 2) concurrency of runtime resources, to existing collaborative inference. In CoActo, system components go beyond simple model splitting of existing approaches and operate more proactively to achieve the coactive execution of inference workloads. CoActo dynamically schedules concurrent interleaving of the mobile, server, and network operations to actively increase resource utilization, enabling lower end-to-end latency. We implement CoActo for various mobile devices and server environments and evaluate our system with distinct environment settings and DNN models. The experimental results show that our system achieves up to 2.1 times speed-up compared to the state-of-the-art collaborative inference solutions. Kyungmin Bin, Jongseok Park 0002, Chanjeong Park, Seyeon Kim 0001, Kyunghan Lee |
MobiSys | 2 |
| 2023 | ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural NetworksabstractModern Deep Neural Network (DNN) frameworks use tensor operators as the main building blocks of DNNs. However, we observe that operator-based construction of DNNs incurs significant drawbacks in parallelism in the form of synchronization barriers. Synchronization barriers of operators confine the scope of parallel computation to each operator and obscure the rich parallel computation opportunities that exist across operators. To this end, we present ASPEN, a novel parallel computation solution for DNNs that achieves fine-grained dynamic execution of DNNs, which (1) removes the operator barriers and expresses DNNs in dataflow graphs of fine-grained tiles to expose the parallel computation opportunities across operators, and (2) exploits these opportunities by dynamically locating and scheduling them in runtime. This novel approach of ASPEN enables opportunistic parallelism, a new class of parallelism for DNNs that is unavailable in the existing operator-based approaches. ASPEN also achieves high resource utilization and memory reuse by letting each resource asynchronously traverse depthwise in the DNN graph to its full computing potential. We provide challenges and solutions to our approach and show that our proof-of-concept implementation of ASPEN on CPU shows exceptional performance, outperforming state-of-the-art inference systems of TorchScript and TVM by up to 3.2$\times$ and 4.3$\times$, respectively. Jongseok Park 0002, Kyungmin Bin, Gibum Park, Sangtae Ha, Kyunghan Lee |
NeurIPS | 1 |
| 2022 | mGEMM: low-latency convolution with minimal memory overhead optimized for mobile devicesabstractThe convolution layer is the key building block in many neural network designs. Most high-performance implementations of the convolution operation rely on GEMM (General Matrix Multiplication) to achieve high computational throughput with a large workload size. However, in mobile environments, the user experience priority puts focus on low-latency inferences over a single or limited batch size. This signifies two major problems of current GEMM-based solutions: 1) GEMM-based solutions require mapping the convolution operation to GEMM, causing overheads in both computation and memory, 2) GEMM-based solutions lose large opportunities of data reuse while mapping, leading to under-utilization of the given hardware. Through an in-depth analysis of current GEMM-based solutions, we identify the root cause of these problems, and we propose mGEMM, a convolution solution that overcomes the aforementioned problems, without changes in accuracy. mGEMM expands the structure of GEMM in such a way that it can accommodate the convolution operation without any overhead, while the existing algorithms suffer from inefficiencies in converting the convolution operation to a static GEMM algorithm. Our extensive evaluations done over various neural networks and test devices show that mGEMM outperforms the existing solutions in the aspects of latency, memory overhead, and energy consumption. In running a real-world application, YoloV3-Tiny object detection, mGEMM achieves up to 1.29× and 1.58× speedup in total latency and convolution latency compared to the state-of-the-art, resulting in 15.5% reduction in energy consumption while using only near-minimum heap memory. Jongseok Park 0002, Kyungmin Bin, Kyunghan Lee |
MobiSys | 1 |