EDBT 2026 Demo / reviewers in the wild / expert
Daewoo Kim
dblp:199/8703
· DBLP profile ↗
16ranked-venue papers
8as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021Computer networks · 5 · 4 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MTM: Rethinking Memory Profiling and Migration for Multi-Tiered Large MemoryabstractMulti-terabyte large memory systems are often characterized by more than two memory tiers with different latency and bandwidth. Multi-tiered large memory systems call for rethinking of memory profiling and migration because of the unique problems unseen in the traditional memory systems with smaller capacity and fewer tiers. We develop MTM, an application-transparent Multi-Tiered Memory management framework, based on three principles: (1) connecting the control of profiling overhead with the profiling mechanism for high-quality profiling; (2) building a universal page migration policy on the complex multi-tiered memory for high performance; and (3) introducing huge page awareness. We evaluate MTM using common big-data applications with realistic working sets (hundreds of GB to 1 TB). MTM outperforms seven solutions by up to 42% (17% on average). Jie Ren 0015, Dong Xu 0024, Junhee Ryu, Kwangsik Shin, Daewoo Kim, Dong Li 0001 |
EuroSys | 5 |
| 2024 | Are Your Epochs Too Epic? Batch Free Can Be HarmfulabstractEpoch based memory reclamation (EBR) is one of the most popular techniques for reclaiming memory in lock-free and optimistic locking data structures, due to its ease of use and good performance in practice. However, EBR is known to be sensitive to thread delays, which can result in performance degradation. Moreover, the exact mechanism for this performance degradation is not well understood. Daewoo Kim, Trevor Brown 0001, Ajay Singh 0002 |
PPoPP | 1 |
| 2024 | Efficient Tensor Offloading for Large Deep-Learning Model Training based on Compute Express LinkabstractThe deep learning models (DL) are becoming bigger, easily beyond the memory capacity of a single accelerator. The recent progress in large DL training utilizes CPU memory as an extension of accelerator memory and offloads tensors to CPU memory to save accelerator memory. This solution transfers tensors between the two memories, creating a major performance bottleneck. We identify two problems during tensor transfers: (1) the coarse-grained tensor transfer creating difficulty in hiding transfer overhead, and (2) the redundant transfer that unnecessarily migrates value-unchanged bytes from CPU to accelerator. We introduce a cache coherence interconnect based on Compute Express Link (CXL) to build a cache coherence domain between CPU memory and accelerator memory. By slightly extending CXL to support an update cache-coherence protocol and avoiding unnecessary data transfers, we reduce training time by $33.7 \%$ (up to $55.4 \%$) without changing model convergence and accuracy, compared with the state-of-the-art work in DeepSpeed [62]. Dong Xu 0024, Kwangsik Shin, Daewoo Kim, Hyeran Jeon, Dong Li 0001 |
SC | 4 |
| 2022 | Symphony: Learning Realistic and Diverse Agents for Autonomous Driving SimulationabstractSimulation is a crucial tool for accelerating the development of autonomous vehicles. Making simulation realistic requires models of the human road users who interact with such cars. Such models can be obtained by applying learning from demonstration (LfD) to trajectories observed by cars already on the road. However, existing LfD methods are typically insufficient, yielding policies that frequently collide or drive off the road. To address this problem, we propose Symphony, which greatly improves realism by combining conventional policies with a parallel beam search. The beam search refines these policies on the fly by pruning branches that are unfavourably evaluated by a discriminator. However, it can also harm diversity, i.e., how well the agents cover the entire distribution of realistic behaviour, as pruning can encourage mode collapse. Symphony addresses this issue with a hierarchical approach, factoring agent behaviour into goal generation and goal conditioning. The use of such goals ensures that agent diversity neither disappears during adversarial training nor is pruned away by the beam search. Experiments on both proprietary and open Waymo datasets confirm that Symphony agents learn more realistic and diverse behaviour than several baselines. Maximilian Igl, Daewoo Kim, Alex Kuefler, Paul Mougin, Punit Shah, Kyriacos Shiarlis, Dragomir Anguelov, Mark Palatucci, Brandyn White, Shimon Whiteson |
ICRA | 2 |
| 2020 | Distributed Slot Scheduling for QoS Guarantee over TSCH-based IoT Networks via Adaptive ParameterizationabstractInternet of Things (IoT), which connects a large number of devices with wireless connectivity, has come into the spotlight. As the scope of IoT applications becomes wider, we observe a surge of missioncritical IoT services, e.g., industrial automation systems and medical IoT systems, requiring to satisfy stringent latency, reliability, and/or energy efficiency guarantees. For this purpose, a new MAC, called Time Slotted Channel Hopping (TSCH), has been standardized in IEEE 802.15.4e. However, it is challenging to design a distributed scheduling protocol that achieves the required QoS and energy efficiency at the same time due to complicated tradeoff (providing enough number of slots for QoS vs. minimizing scheduled slots for energy efficiency). In this paper, we propose a novel framework for providing QoS, called SSAP, which is designed to maximize network lifetime in a distributed fashion while satisfying given reliability and latency requirements. To this end, we decompose our goal into two crucial design components: (i) scheduling of slot and channel, and (ii) control of medium access period, each of which is performed by low-complexity and distributed mechanisms. To the best of our knowledge, this paper is the first work to comprehensively handle multiple QoSes for TSCH-based IoT networks. We implement SSAP in Contiki OS and perform extensive simulations and real experiments under various scenarios. Our evaluation results demonstrate that SSAP satisfies highly reliable communication and latency requirements while having the network lifetime that is 1.6 times longer compared to existing protocols for TSCH. Jinhwan Jung, Daewoo Kim, Joohyun Kang, Namjo Ahn, Yung Yi |
IPSN | 2 |
| 2020 | Bird-MAC: Energy-Efficient MAC for Quasi-Periodic IoT Applications by Avoiding Early Wake-upabstractWe propose a new MAC protocol for IoT applications, called Bird-MAC, which is highly energy efficient in the applications where IoT sensors report monitoring status in a quasi-periodic manner, as in structural health monitoring and static environmental monitoring. Two key design ideas of Bird-MAC are: (a) no need of early-wake-up of transmitters and (b) taking the right balance between synchronization and coordination costs. The idea (a) is possible by allowing a node (whether it is a transmitter or receiver) to wake up just with its given wake-up schedule, and letting a late bird (which wakes up later) notify its wake-up status to its corresponding early bird (which wakes up earlier), where the early bird just infrequently waits for the late bird's wake-up signal. The idea (b) is realized by designing Bird-MAC to be placed in a scheme between purely synchronous and asynchronous schemes. We provide a rigorous mathematical analysis that is used to choose the right protocol parameters of Bird-MAC. We demonstrate the performance of Bird-MAC through extensive simulations, and real experiments. The experiment on our testbed using a 26 node testbed at an underground parking lot of our office building to monitor its structural health shows that energy consumption is reduced by about up to 45 percent over existing sensor MAC protocols. We also confirm the applicability of Bird-MAC in a challenging and realistic scenario through the experiment on Yeongjong Grand Bridge in South Korea. Daewoo Kim, Jinhwan Jung, Yoonpyo Koo, Yung Yi |
IEEE Trans. Mob. Comput. | 1 |
| 2020 | Economics of Fog Computing: Interplay Among Infrastructure and Service Providers, Users, and Edge Resource OwnersabstractFog computing is a paradigm which brings computing, storage, and networking closer to end users and end devices for better service provisioning. One of the crucial factors in the success of fog computing is on how to incentivize the individual users' edge resources and provide them to end users such that fog computing is economically beneficial to all involved economic players. In this paper, we model and analyze a market of fog computing, from which we aim at drawing practical implications to uncover how the fog computing market should operate. To this end, we conduct an economic analysis of such user-oriented fog computing by modeling a market consisting of Infrastructure and Service Provider (ISP), end Service Users (SUs), and Edge Resource Owners (EROs) as a non-cooperative game. In this market, ISP, which provides a platform for fog computing, behaves as a mediator or a broker which leases EROs' edge resources and provides various services to SUs. In our model, a two-stage dynamic game is used where in each stage, there exists a dynamic game, one for between ISP and EROs and another for between ISP and SUs, to model the market more practically. Despite this complex game structure, we provide a closed-form equilibrium analysis which gives an insight on how much economic benefit is obtained by ISP, SUs, and EROs from user-oriented fog computing under what conditions, and we figure out the economic factors that have a significant impact on the success of fog computing. Daewoo Kim, Hyojung Lee, HyungSeok Song, Nakjung Choi, Yung Yi |
IEEE Trans. Mob. Comput. | 1 |
| 2019 | On-chip memory optimization for high-level synthesis of multi-dimensional data on FPGAabstractIt is very challenging to design an on-chip memory architecture for high-performance kernels with large amount of computation and data. The on-chip memory architecture must support efficient data access from both the computation part and the external memory part, which often have very different expectations about how data should be accessed and stored. Previous work provides only a limited set of optimizations. In this paper we show how to fundamentally restructure on-chip buffers, by decoupling logical array view from the physical buffer view, and providing general mapping schemes for the two. Our framework considers the entire data flow from the external memory to the computation part in order to minimize resource usage without creating performance bottleneck. Our experimental results demonstrate that our proposed technique can generate solutions that reduce memory usage significantly (2X over the conventional method), and successfully generate optimized on-chip buffer architectures without costly design iterations for highly optimized computation kernels. Daewoo Kim, Sugil Lee, Jongeun Lee |
ASP-DAC | 1 |
| 2019 | Learning to Schedule Communication in Multi-agent Reinforcement Learning
Daewoo Kim, David Hostallero, Wan Ju Kang, Kyunghwan Son, Yung Yi |
ICLR (Poster) | 1 |
| 2019 | QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement LearningabstractWe explore value-based solutions for multi-agent reinforcement learning (MARL) tasks in the centralized training with decentralized execution (CTDE) regime popularized recently. However, VDN and QMIX are representative examples that use the idea of factorization of the joint action-value function into individual ones for decentralized execution. VDN and QMIX address only a fraction of factorizable MARL tasks due to their structural constraint in factorization such as additivity and monotonicity. In this paper, we propose a new factorization method for MARL, QTRAN, which is free from such structural constraints and takes on a new approach to transforming the original joint action-value function into an easily factorizable one, with the same optimal actions. QTRAN guarantees more general factorization than VDN or QMIX, thus covering a much wider class of MARL tasks than does previous methods. Our experiments for the tasks of multi-domain Gaussian-squeeze and modified predator-prey demonstrate QTRAN’s superior performance with especially larger margins in games whose payoffs penalize non-cooperative behavior more aggressively. Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Hostallero, Yung Yi |
ICML | 2 |
| 2019 | Double MAC on a DSP: Boosting the Performance of Convolutional Neural Networks on FPGAsabstractDeep learning workloads, such as convolutional neural networks (CNNs) are important due to increasingly demanding high-performance hardware acceleration. One distinguishing feature of a deep learning workload is that it is inherently resilient to small numerical errors and thus works very well with low precision hardware. We propose a novel method called double multiply-and-accumulate (MAC) to theoretically double the computation rate of CNN accelerators by packing two MAC operations into one digital signal processing block of off-the-shelf field-programmable gate arrays (FPGAs). We overcame several technical challenges by exploiting the mode of operation in the CNN accelerator. We have validated our method through FPGA synthesis and Verilog simulation, and evaluated our method by applying it to the state-of-the-art CNN accelerator. The double MAC approach used can double the computation throughput of a CNN layer. On the network level (all convolution layers combined), the performance improvement varies depending on the CNN application and FPGA size, from 14% to more than 80% over a highly optimized state-of-the-art accelerator solution, without sacrificing the output quality significantly. Sugil Lee, Daewoo Kim, Dong Nguyen 0001, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | On the Economics of Fog Computing: Inter-Play among Infrastructure and Service Providers, Users, and Edge Resource OwnersabstractFog computing is a paradigm which brings computing, storage, and networking closer to end users and devices for better service provisioning. One of the crucial factors towards the success of fog computing is how to incentivize the individual users' edge resources, thereby opening the era of user- participated fog computing. In this paper, we provide an economic analysis of such user-oriented fog computing by modeling a market consisting of ISP (Infrastructure and Service Provider), SUs (end Service Users), and EROs (Edge Resource Owners) as a noncooperative game. In this market, ISP, which provides a platform of fog computing, behaves as a mediator or a broker to lease the edge resources from EROs and provide various services to SUs. In our game formulation, a two-stage dynamic game is used, where in each stage there exists another dynamic game, one for between ISP and EROs and another for between ISP and SUs, to model the market more practically. Despite this complex game structure, we provide a closed- form equilibrium analysis, which gives an insight of how much economic benefits are obtained by ISP, SUs, and EROs under what conditions. Daewoo Kim, Hyojung Lee, HyungSeok Song, Nakjung Choi, Yung Yi |
ICC | 1 |
| 2018 | Cost-Performance Comparison of Various Accelerator Implementation Platforms for Deep Convolutional Neural Network
Yechan Yu, HoJin Kim, Jinjoo Ha, Daewoo Kim, Kang Yi |
PDCAT | 4 |
| 2017 | Double MAC: Doubling the performance of convolutional neural networks on modern FPGAsabstractThis paper presents a novel method to double the computation rate of convolutional neural network (CNN) accelerators by packing two multiply-and-accumulate (MAC) operations into one DSP block of off-the-shelf FPGAs (called Double MAC). While a general SIMD MAC using a single DSP block seems impossible, our solution is tailored for the kind of MAC operations required for a convolution layer. Our preliminary evaluation shows that not only can our Double MAC approach increase the computation throughput of a CNN layer by twice with essentially the same resource utilization, the network level performance can also be improved by 14~84% over a highly optimized state-of-the-art accelerator solution depending on the CNN hyper-parameters. Dong Nguyen 0001, Daewoo Kim, Jongeun Lee |
DATE | 2 |
| 2017 | FPGA implementation of convolutional neural network based on stochastic computingabstractThere has been a body of research to use stochastic computing (SC) for the implementation of neural networks, in the hope that it will reduce the area cost and energy consumption. However, no working neural network system based on stochastic computing has been demonstrated to support the viability of SC-based deep neural networks in terms of both recognition accuracy and cost/energy efficiency. In this demonstration we present an SC-based deep nenural network system that is highly accurate and efficient. Our system takes an input image and processes it with a convolutional neural network implemented on an FPGA using stochastic computing to recognize the input image, with nearly the same accuracy as conventional binary implementations. Daewoo Kim, Mansureh S. Moghaddam, Hossein Moradian, Hyeon Uk Sim, Jongeun Lee, Kiyoung Choi |
FPT | 1 |
| 2017 | Revisiting Sensor MAC for Periodic Monitoring: Why Should Transmitters Be Early Birds?abstractWe propose a new sensor MAC protocol, called Bird-MAC, which is highly energy efficient in the applications where sensors periodically report monitoring status with a very low rate, as in structural health monitoring and static environmental monitoring. Two key design ideas of Bird-MAC are: (a) no need of early-wake-up of transmitters and (b) taking the right balance between synchronization and coordination costs. The idea (a) is possible by allowing a node (whether it is a transmitter or receiver) to wake up just with its given wake-up schedule, and letting a late bird (which wakes up later) notify its wake-up status to its corresponding early bird (which wakes up earlier), where the early bird just infrequently waits (i.e., nods) for the late bird's wake-up signal. The idea (b) is realized by designing Bird-MAC to be placed in a scheme between purely synchronous and asynchronous schemes. We provide rigorous mathematical analysis that is used to choose the right protocol parameters of Bird-MAC. We demonstrate the performance of Bird-MAC through extensive simulations, and real experiments using a 26 node testbed at an underground parking lot of our office building to monitor its structural health, where we confirm that energy consumption is reduced by about up to 45% over existing sensor MAC protocols. Daewoo Kim, Jinhwan Jung, Yoonpyo Koo, Yung Yi |
SECON | 1 |