Intaik Park

dblp:27/6921 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-6958-5997ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 6 since 2021Systems, architecture and hardware · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems
abstract
In large-scale industrial recommendation systems, model checkpoints are instrumental in maintaining training goodput and numerical correctness during system failures and job preemptions. The increasing prevalence of multi-terabyte models has rendered frequent regular model checkpoints impractical, resulting in substantial lost progress when recovering from failures. As model sizes continue to grow, researchers and practitioners are compelled to investigate more efficient and scalable solutions. This paper presents DECK, a novel approach to delta model checkpointing designed for real-world industrial systems. Specifically, DECK focuses on extracting delta states with near-zero overhead, staging and streaming delta checkpoints without interrupting the training process, and merging delta checkpoints in an optimal and decoupled manner. Experimental results demonstrate that DECK achieves a 12-fold increase in checkpoint frequency while maintaining negligible impact on training throughput, thereby attaining state-of-the-art (SOTA) production performance.
Sibasish Acharya, Sihui Han, Yongxiong Ren, Yanli Zhao, Chucheng Wang, Pradeep Fernando, Siqi Yan, Yicong Du, Elzbieta Krepska, Intaik Park, Min Ni, Qunshu Zhang
Proc. VLDB Endow.13
2024 Toward 100TB Recommendation Models with Embedding Offloading
abstract
Training recommendation models become memory-bound with large embedding tables, and fast GPU memory is scarce. In this paper, we explore embedding caches and prefetch pipelines to effectively leverage large but slow host memory for embedding tables. We introduce Locality-Aware Sharding and iterative planning that automatically size caches optimally and produce effective sharding plans. Embedding Offloading, a system that combines all of these components and techniques, is implemented on top of Meta’s open-source libraries, FBGEMM GPU and TorchRec, and it is used to improve scalability and efficiency of industry-scale production models. Embedding Offloading achieved 37x model scale to 100TB model size with only 26% training speed regression.
Intaik Park, Ehsan K. Ardestani, Damian Reeves, Sarunya Pumma, Henry Tsang, Levy Zhao, Jian He 0004, Joshua Deng, Dennis Van Der Staay
RecSys1
2021 MTCNet: Multi-Task Complex Network for Concurrent Channel Estimation and Equalization
abstract
Convolutional neural networks (CNNs) have been widely adopted in various fields and the wireless communication domain is no exception; researchers leveraged CNNs for communication tasks such as channel estimation and equalization. They showed the potential of CNNs for these applications but were insufficient to be considered for real world implementation due to the limited applicability and high execution latency. In this paper, we propose a novel complex number-based neural network architecture, a multi-task loss function, and a training scheme to learn channel estimation and equalization tasks simultaneously in an end-to-end manner. The resulting single-stage CNN architecture, Multi-Task Complex Network (MTCNet), achieves better performance than any other conventional methods. In addition, to the best of our knowledge, MTCNet is the first CNN-based channel estimation and equalization that can operate with various subcarrier settings with inferencing time below 1 ms.
Dongha Bahn, Jae-Il Jung, Junik Jang, Changbae Yoon, Chanjong Park, Intaik Park
GLOBECOM6
2021 One for All: Traffic Prediction at Heterogeneous 5G Edge with Data-Efficient Transfer Learning
abstract
By placing the computing, storage and networking resources close to the end users, distributed edge computing greatly benefits the performance of 5G communication systems. However, as a tradeoff, resources on the edge are usually limited and imbalanced among the heterogeneous edge nodes. To overcome this drawback, this paper proposes a Transfer Learning based Prediction (TLP) framework that allows the edge nodes to share their resources and data in an efficient manner. In particular, the TLP framework focuses on the prediction of the future traffic load, which is a key reference for many automated network functions. To enhance the efficiency of data and bandwidth, TLP first learns a base model on a data-abundant edge node (the source), and then transfers this model (instead of data) to other data-limited nodes (the targets). To achieve a delicate balance between maintaining common features and learning target-specific features, we develop a new transfer learning technique named Similarity-based Elastic Weight Con-solidation (SEWC), and integrate it into TLP. Experiments on real-world data illustrate that, compared to the state-of-the-art methods, TLP-SEWC reduces the Mean Absolute Error (MAE) of traffic prediction by up to 57.9%.
Xi Chen 0009, Ju Wang 0003, Yi Tian Xu, Di Wu 0044, Xue Liu 0004, Gregory Dudek, Taeseop Lee, Intaik Park
GLOBECOM9
2021 AFB: Improving Communication Load Forecasting Accuracy with Adaptive Feature Boosting
abstract
Prediction of key system characteristics, such as the communication load, is required to overcome the delays in wireless communication systems. State-of-The-Art (SOTA) approaches mostly apply existing Neural Network (NN) structures, and extract latent features purely based on their sensitivity to the forecasting accuracy. This way of feature extraction may neglect some non-obvious yet informative dimensions in the model input, leading to inaccurate forecasting results. In this paper, we present an Adaptive Feature Boosting (AFB) approach, which integrates multiple AutoEncoders (AEs) to automatically extract robust and comprehensive latent features for communication load forecasting. The recurrent and residual connections among the AEs make sure that the extracted latent features are representative for all input dimensions. With more comprehensive information extracted from the history, the forecasting accuracy is thus improved. We evaluate AFB against existing approaches on a real-world dataset that contains Call Detail Records (CDRs) of the Milan city over a period of two months. The evaluation shows that our AFB-based approach achieves 35.2% more accurate load forecasting results than the SOTA deep approaches.
Chengming Hu, Xi Chen 0009, Ju Wang 0003, Jikun Kang, Yi Tian Xu, Xue Liu 0004, Di Wu 0044, Seowoo Jang, Intaik Park, Gregory Dudek
GLOBECOM10
2021 Deep Reinforcement Learning for cell on/off energy saving on Wireless Networks
abstract
Increased network traffic demands have led to ex-tremely dense network deployments. This translates to significant growth in energy consumption at the radio access networks, resulting in high network operation costs (OPEX). In this work, we apply deep reinforcement learning to reduce the energy consumption at the base station in dense wireless networks, by allowing cells that overlap in geographical areas to be put in standby mode according to the changing network conditions. We start by formulating the problem of the cell on/off energy saving in dense wireless networks as a Markov decision process. Then, a deep reinforcement learning (DRL) solution is proposed. This DRL solution takes into account different key performance indicators (KPIs) of both the network and user equipment and aims to reduce the energy consumed by the network without significantly impacting the overall KPIs. The performance of the proposed solution is evaluated using a practical network simulator.
Joan S. Pujol-Roigl, Shangbin Wu, Yue Wang 0008, Minsuk Choi, Intaik Park
GLOBECOM5
2021 Load Balancing for Communication Networks via Data-Efficient Deep Reinforcement Learning
abstract
Within a cellular network, load balancing between different cells is of critical importance to network performance and quality of service. Most existing load balancing algorithms are manually designed and tuned rule-based methods where near-optimality is almost impossible to achieve. These rule-based meth-ods are difficult to adapt quickly to traffic changes in real-world environments. Given the success of Reinforcement Learning (RL) algorithms in many application domains, there have been a number of efforts to tackle load balancing for communication systems using RL-based methods. To our knowledge, none of these efforts have addressed the need for data efficiency within the RL framework, which is one of the main obstacles in applying RL to wireless network load balancing. In this paper, we formulate the communication load balancing problem as a Markov Decision Process and propose a data-efficient transfer deep reinforcement learning algorithm to address it. Experimental results show that the proposed method can significantly improve the system performance over other baselines and is more robust to environmental changes.
Di Wu 0044, Jikun Kang, Yi Tian Xu, Jimmy Li 0001, Xi Chen 0009, Dmitriy Rivkin, Michael R. M. Jenkin, Taeseop Lee, Intaik Park, Xue Liu 0004, Gregory Dudek
GLOBECOM10
2021 Hierarchical Policy Learning for Hybrid Communication Load Balancing
abstract
Due to the uneven demographic distribution and people’s daily activities, communication systems usually experience highly imbalanced load across different cells. This imbalance leads to unsatisfied users in the congested cells and under-utilized resources in the less-loaded cells. To deal with this issue, existing work migrates the load from heavily loaded cells to lightly loaded cells, by either handing over active mode User Equipment (UEs) to other serving cells, or re-selecting the camping cells for idle mode UEs. In this paper, we further advance the research on Load Balancing (LB) with a hybrid control of both active and idle UEs. This task is challenging, due to the conflicts between Active-UE LB (AULB) and Idle-UE LB (IULB) policies. To overcome this challenge, we propose a Hierarchical Policy Learning (HPL) framework, which coordinates the actions between LB policies with a two-level learning structure. In this way, HPL produces AULB and IULB policies that are better aligned with each other. Extensive simulation results illustrate the efficiency and efficacy of the proposed HPL.
Jikun Kang, Xi Chen 0009, Di Wu 0044, Yi Tian Xu, Xue Liu 0004, Gregory Dudek, Taeseop Lee, Intaik Park
ICC8
2008 Launch-on-Shift-Capture Transition Tests
abstract
The two most popular transition tests are launch-on-shift (LOS) test and launch-on-capture (LOC) test. The LOS and LOC tests differ in their launch mechanisms, creating their own pros and cons. In this paper, new hybrids of LOS and LOC tests that launch transitions using both launch mechanisms are introduced. The new transition tests improved fault coverage without significant test length penalty. This paper presents the concepts and pattern generation methods of these new transition tests as well as experimental results that demonstrate the benefits of these tests.
Intaik Park, Edward J. McCluskey
ITC1
2008 Error Sequence Analysis
abstract
With increasing IC process variation and increased operating speed, it is more likely that even subtle defects will lead to the malfunctioning of a circuit. Various fault models, such as the transition fault model and the path-delay model, have been used to aid delay defect detection. However, these models are not efficient for small-delay defect coverage or for test pattern generation time. Error sequence analysis utilizes the order in which the errors occur during a frequency sweep of a transition test to identify small-delay defects that may escape the same test applied in the conventional way. Moreover, it can detect such defects even in the presence of inter-die process variations, such as lot-to-lot and wafer-to-wafer process variation. In addition, error sequence analysis is very effective in separating devices with delay defects from devices that have failed due to process variation.
Jaekwang Lee, Intaik Park, Edward J. McCluskey
VTS2
2008 Inconsistent Fail due to Limited Tester Timing Accuracy
abstract
Delay testing is a technique to determine if a chip will function correctly at a specified frequency. If a chip passes delay tests, it will presumably function at the specified frequency in the field. This paper presents experimental results that show how chips can pass very thorough delay tests and still fail in the field. It is shown that some chips sometimes pass and sometimes fail when the same delay test is applied multiple times under the same test conditions. These chips are called inconsistent fails. This paper shows how tester timing edge placement accuracy can cause inconsistent fails and suggests the minimum requirements for guardbands that avoid the inconsistent test results.
Intaik Park, Donghwi Lee, Erik Chmelar, Edward J. McCluskey
VTS1
2005 Effective TARO Pattern Generation
abstract
TARO test patterns are transition fault test patterns that sensitize each transition fault to all of the outputs that can be reached from the fault location. We were not able to identify any ATPG tool that can generate TARO test patterns directly. This paper describes a technique to use an existing transition fault ATPG tool to efficiently generate TARO test patterns. This technique was used to generate TARO patterns for the ELF35 test chip. When these patterns were applied to the ELF35 chips, all of the defective chips were discovered (no test escapes).
Intaik Park, Ahmad A. Al-Yamani, Edward J. McCluskey
VTS1