VLDB 2026 Research / reviewers in the wild / expert
Daejin Park
dblp:119/5645
· DBLP profile ↗
22ranked-venue papers
2as first author
15since 2021 · last 2025
0000-0002-5560-873XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bit-Separable Transformer Accelerator Leveraging Output Activation Sparsity for Efficient DRAM Access
Seunghyun Park 0002, Daejin Park |
HCS | 2 |
| 2025 | Lightweight Empirical Reinforcement Learning Driven Adaptive Super-Twisting Control with Fused Linear-Nonlinear Sliding Surfaces for Embedded Vehicle ControlabstractThis paper proposes a lightweight empirical reinforcement learning-based sliding mode control framework with a hybrid sliding surface, designed to enhance robustness and smoothness of control input under various disturbances and driving conditions. While the conventional Super-Twisting Algorithm (STA) provides strong robustness, its fixed control parameters limit adaptability and generalization. To address this issue, adaptive control parameter tuning is introduced through machine learning techniques such as reinforcement learning. However, such approaches often suffer from high computational cost and instability during the early stages of learning, making them less practical for real-time embedded systems. To overcome these limitations, this study presents a lightweight empirical reinforcement learning method that adjusts control parameters in real time based on system errors and sliding surface dynamics. A hybrid sliding surface structure that dynamically blends linear and nonlinear components is also introduced to allow the control system to adapt its convergence behavior organically according to real-time conditions. This results in an optimal balance between control performance and input smoothness. Experimental evaluations demonstrate that the proposed framework enables flexible and balanced control behavior, making it highly suitable for real-time, robust, and energy-efficient vehicle control applications. Hyunjoong Lee, Daejin Park |
IECON | 2 |
| 2025 | ML-Based Fast and Precise Target Docking of Autonomous Mobile Robots for Intelligent Transportation Systems Using 2-D LiDAR
Sunghoon Hong, Hyukjun Kwon, Gyuhun Sim, Kwangyong Choi, Daejin Park |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | S3A-NPU: A High-Performance Hardware Accelerator for Spiking Self-Supervised Learning With Dynamic Adaptive Memory OptimizationabstractSpiking self-supervised learning (SSL) has become prevalent for low power consumption and low-latency properties, as well as the ability to learn from large quantities of unlabeled data. However, the computational intensity and resource requirements are significant challenges to apply to accelerators. In this article, we propose the scalable, spiking self-supervised learning, streamline optimization accelerator ($S^{3}$A)-neural processing unit (NPU), a highly optimized accelerator for spiking SSL models. This architecture minimizes memory access by leveraging input data provided by the user and optimizes computation through the maximization of data reuse. By dynamically optimizing memory based on model characteristics and implementing specialized operations for data preprocessing, which are critical in SSL, computational efficiency can be significantly improved. The parallel processing lanes account for the two encoders in the SSL architecture, combined with a pipelined structure that considers the temporal data accumulation of spiking neural networks (SNNs) to enhance computational efficiency. We evaluate the design on field-programmable gate array (FPGA), where a 16-bit quantized spiking residual network (ResNet) model trained on the Canadian Institute for Advanced Research (CIFAR) and MNIST dataset has top 94.08% accuracy.$S^{3}$A-NPU optimization significantly improved computational resource utilization, resulting in a 25% reduction in latency. Moreover, as the first spiking self-supervised accelerator, it demonstrated highly efficient computation compared to existing accelerators, utilizing only 29k look up tables (LUTs) and eight block random access memories (BRAMs). This makes it highly suitable for resource-constrained applications, particularly in the context of spiking SSL models on edge devices. We implemented it on a silicon chip using a 130-nm process design kit (PDK), and the design was less than$1~\text {cm}^{2}$. Heuijee Yun, Daejin Park |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Cloud Memory Enabled Code Generation via Online Computing for Seamless Edge AI OperationabstractThis paper introduces an innovative architecture designed to enhance the execution of Artificial Intelligence (AI) software on edge devices, which are often constrained by limited hardware resources. The core of proposal is to dynamically adapt AI models through server-mediated parameter updates and learning, thus allowing edge devices to efficiently process AI tasks in real-time and adapt to various operational conditions. By leveraging the computational power of cloud resources for the heavy lifting of AI model training, the computational burden on edge devices is alleviated, enabling them to focus on inference tasks with updated models. This approach significantly improves the operational efficiency and adaptability of edge computing in AI applications. Our architecture employs server-based emulation to monitor and dynamically update edge devices, ensuring their execution is optimized for current conditions. Experimental results demonstrate a substantial reduction in operational time up to 75% compared to traditional edge devices without accelerators and 49 % when compared to devices equipped with accelerators. Moreover, proposed model shows an ability to improve accuracy by 20 % in scenarios with biased inputs through continuous learning and parameter updating, highlighting its adaptability to changing environments. This research contributes to the field of edge computing by demonstrating a viable solution for deploying sophisticated AI models in resource-constrained environments. By offloading computationally intensive tasks to the cloud, proposed architecture ensures that edge devices can operate more efficiently and handle a broader range of AI applications. This study not only underscores the potential of integrating cloud and edge computing to overcome the limitations of edge devices but also opens new avenues for future research in intelligent edge computing systems. Myeongjin Kang, Daejin Park |
COMPSAC | 2 |
| 2024 | Implementation of Dynamic Round Robin Scheduling on Bare-Metal Shallow Multi-OS for Lightweighted MicrocontrollersabstractIn recent years, there has been a trend towards integrating functions using a small number of microcontrollers instead of employing multiple microcontrollers across various environments. This shift underscores the need for a hypervisor capable of efficiently utilizing resources while imposing minimal overhead. Addressing this demand, this paper introduces a hyper-visor employing dynamic round-robin scheduling, which flexibly adjusts time quantum allocation based on the urgency of each OS. Furthermore, a monitor mode is devised to oversee resource allocation among multiple OS. To enhance responsiveness while managing these OS, ultra-light context-switching is implemented within the monitor mode. The proposed system demonstrates a notable reduction in execution time, approximately 19% compared to traditional round-robin scheduling. Additionally, in terms of energy efficiency, the proposed system yields a 34% reduction in energy consumption compared to existing methods. Notably, the ultra-light context-switching mechanism consumes only about 5% of the processing cycle when compared to FreeRTOS. Daejin Park |
COMPSAC | 2 |
| 2024 | Optimizing Multithreaded Access to Global Variables in SPM through Compiler-Enhanced Dependency Analysis
Gihyeon Jeon, Daejin Park |
TENCON | 2 |
| 2024 | Differential Image-Based Scalable YOLOv7-Tiny Implementation for Clustered Embedded SystemsabstractConvolutional neural networks (CNNs) for powerful visual image analysis are gaining popularity in artificial intelligence. The main difference in CNNs compared to other artificial neural networks is that many convolutional layers are added, which improve the performance of visual image analysis by extracting the feature maps required for image classification. However, algorithm optimization is required to run applications that require low-latency in edge compute modules with limited processing resources. In this paper, we propose a novel algorithm optimization method for fast CNNs by using continuous differential images. The main idea is to reduce computation variably by using the differential value of the input in each convolutional layer. Also, the proposed method is compatible with all types of CNNs, and the performance is better when the pixel value difference of continuous images is low. We use the DarkNet framework to evaluate our algorithm using fast convolution and half convolution approaches on a clustered system. As a result, when the input frame rate is 10 fps, FLOPs are reduced by about 4.92 times compared to the original YOLOv7-tiny. By reducing the FLOPs of the convolutional layer, the inference speed increases to about 4.86 FPS, performing 1.57 times faster than the original YOLOv7-tiny. In the case of parallel processing that used two edge compute modules for using half convolution approach, FLOPs reduced more, and the response speed improved. In addition, faster Object detection implementation is possible by additionally expanding up to 7 compute modules in a scalable clustered embedded system as much as the user wants. Sunghoon Hong, Daejin Park |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Work-in-Progress: Searching Optimal Compiler Optimization Passes Sequence for Reducing Runtime Memory Profile using Ensemble Reinforcement LearningabstractThe order in which compiler optimization passes are applied has a significant impact on program performance. However, widely used compiler optimization options use handpicked sets of optimization passes, optimized for specific benchmarks. In this paper, we propose an ensemble reinforcement learning (RL) model that optimizes LLVM transform passes sequence to reduce the runtime memory profile, which is an important consideration in resource-constrained embedded systems. We developed an LLVM intermediate representation (IR) analysis pass to extract static program features. The extracted features are processed with PCA for dimension reduction. We also generated datasets using a random program generator, and clustered them according to the PCA results of their extracted features. The ensemble RL model was trained on each clustered dataset. Experiments showed that the proposed model reduced 37% more memory profile than the standard optimization option. Juneseo Chang, Daejin Park |
EMSOFT | 2 |
| 2023 | Work-in-Progress: Micro-Accelerator-in-the-Loop Framework for MCU Integrated Accelerator Peripheral Fast PrototypingabstractThe resource constraints of MCU-based platforms limits their ability to utilize high-performance accelerators such as GPUs or servers, mainly due to insufficient resources for ML applications. Currently, solutions utilizing accelerators connected as peripherals to the on-chip bus of microcontroller units (MCUs) are being proposed. We define this approach as a Micro-Accelerator (MA). Due to the necessity of connecting the MA to the MCU core and the on-chip bus within the chip, conducting a iterative full system evaluation of the embedded software that drives the MA poses significant challenges. To address this challenge, we propose a framework that enables rapid prototyping of custom-designed MA and facilitates profiling of its acceleration performance. Experimental results evaluating the performance of the MA for two tiny machine learning (TinyML) applications within the proposed framework demonstrate a cycle latency reduction of 84.32% and 61.32% compared to a general machine learning framework, respectively. Jisu Kwon, Daejin Park |
EMSOFT | 2 |
| 2023 | RTOS-Based Task-Driven Scheduling for Vehicle Independent Brushless Direct Current Motor ControlabstractHigh-level, self-driving vehicles and unmanned vehicles are completely electronic instead of having traditional mechanical steering methods. Among the electronic steering systems, a motor is assigned to each wheel to simplify the structure of the vehicle, and additional efficiency and functions can be obtained through unique movements. Brushless direct current (BLDC) motors, which operate by applying a voltage pattern controlled by pulse-width modulation(PWM), are advantageous for these systems. When a steering command is entered into the system that outputs a motor signal that changes the PWM pattern, the motor signal output is interrupted during the steering task, which negatively affects the motor's operation. In this situation, if the motor's specific section output signal is ensured, the unexpected effect on the motor's operation can be reduced. This paper proposes a real-time scheduling system led by a motor signal task that ensures phase cycles according to the operating characteristics of a BLDC motor. After generating a motor operation output of a specific phase cycle, the motor signal task checks the priority table of the tasks currently to be performed and continues performing the task or blocks itself. In this case, the length of the phase cycle varies according to the current motor operation state. The scheduler manages the execution status of tasks after the motor signal task is blocked, and updates the priority table by receiving external requests in the form of interrupts. In situations where the motor is accelerating, the proposed method provides a disturbance ratio, that is 20.14% less than that of the general method. In situations where the motor is driven at constant speed, the proposed method provides a disturbance ratio, that is 8.82% less than that of the general method. In both cases, the proposed method showed the highest amount of stability increase in the low-speed section. Dongkyu Jung, Daejin Park |
IECON | 2 |
| 2023 | Digital-Twin Consistency Checking Based on Observed Timed Events With Unobservable Transitions in Smart ManufacturingabstractSmart factories manage digital twins (DTs) to evaluate the performance of various what-if production scenarios. This article presents a DT consistency-checking approach to maintain DT in high fidelity by checking whether each sensed timed event from the physical manufacturing plant is under its corresponding DT-based estimations in runtime. The approach targets DTs developed using time colored Petri net (TCPN). To build the candidates of the next observable event with observable time margins, we considered the stochastic property of the plant, frequent external actuation caused by a new order, machine maintenance, etc., as well as intermediate unobservable state transitions reaching the sensible events. Based on the considerations, we propose an iterative method to build the virtual estimates for streaming physical events using efficiently evolved state-class graphs (SCGs). We also propose a TCPN partitioning method to accelerate the SCG-evolution and make DT maintenance easier by supporting the isolation of inconsistent subnets being diagnosed. We applied the approach to a USB flash-drive factory to prove the concept and evaluated the performance under various situations to show speedups of the SCG evolution, that is the crucial overhead of the estimation. Moon Gi Seok, Wen Jun Tan, Wentong Cai 0001, Daejin Park |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Work-in-Progress: Accuracy-Area Efficient Online Fault Detection for Robust Neural Network Software-Embedded MicrocontrollersabstractDetecting transient faults in safety-critical neural network (NN) applications operated on embedded systems has become a concern, but it is challenging to achieve high accuracy because of the open context problem and resource constraints. This study proposes an accuracy-area efficient, data-analysis-based online soft errors (SEs) and control flow errors (CFEs) detection, applicable to any NN application with low overhead. We insert code for runtime monitoring data assertion, and the data are distributed to shallow or deep detection models selectively. The shallow detection model detects CFEs by verifying runtime signatures with values obtained from simulations, and detects SEs of data having constant values according to program input. SEs of other data are verified by a deep detection model using a sliding window one-class support vector machine. Fault injection experiments on an image classification NN showed that our detector has significant detection accuracy in fault conditions. Juneseo Chang, Sejong Oh, Daejin Park |
EMSOFT | 3 |
| 2021 | Toward Data-Adaptable TinyML using Model Partial Replacement for Resource Frugal Edge DeviceabstractDemand to perform machine learning (ML) tasks in microcontroller unit (MCU)-based edge devices instead of the server, that have limited resources, is gradually increasing. TinyML framework makes possible that creating ML firmware in a language that can be ported to the MCU. This paper aims at a technique that flexibly responds to various inputs by partial replacement of the network model part among the ML firmware operating in the MCU. Before implementing the proposed technique, a preliminary experiment was performed. As the number of words trained on the network in the speech command dataset increases, the size of the model increases, but the evaluation accuracy decreases. The experimental results show the possibility of a technique that replaces small learning models corresponded to each domain, instead of using a huge model that trains all input data variations for different domains. Jisu Kwon, Daejin Park |
HPC Asia | 2 |
| 2021 | Metamorphic Edge Processor Simulation Framework Using Flexible Runtime Partial Replacement of Software-Embedded Verilog RTL ModelsabstractIterative register-transfer level (RTL) simulation is essential for the edge processor design, but the RTL simulation speed is significantly slower in a system where various RTL models are complicatedly integrated. In this paper, we propose a novel metamorphic edge processor simulation framework that partitions the software part and virtualizes it in the system emulator to eject from full RTL simulation. The system emulator, which is written in a high-level language, and the Verilog simulation have different abstraction levels, thus the Verilog procedural interface (VPI) module is plugged into the Verilog simulator to connect with the virtual layer interface. In the system emulator, a Verilog RTL simulation session corresponding to a specific parameter set can be dynamically loaded at runtime to provide metamorphism by flexible partial parameter-driven RTL model replacement. We applied the proposed framework to finite impulse response (FIR) filter, and it is successfully demonstrated and achieved simulation speedup for given parameters. Jisu Kwon, Sejong Oh, Daejin Park |
ISCAS | 3 |
| 2020 | Runtime Abstraction-Level Conversion of Discrete-Event Wafer-fabrication Models for Simulation AccelerationabstractSpeeding up the simulation of discrete-event wafer fab models is essential because optimizing the scheduling and dispatching policies under various circumstances requires repeated evaluation of the decision candidates during parameter-space exploration. In this paper, we present a runtime abstraction-level conversion approach for discrete-event wafer-fabrication (wafer-fab) models to gain simulation speedup. During the simulation, if a machine group of the wafer fab models reaches a steady state, then the proposed approach attempts to substitute this group model with a mean-delay model (MDM) as a high abstraction level model. The MDM abstracts the detailed operations of the group's sub-component models into an average delay based on the queueing modeling, which can guarantee acceptable accuracy under steady state. The proposed abstraction-level converter (ALC) observes the queueing parameters of low-level groups to identify the convergence of each group's work-in-progress (WIP) level through a statistical test. When a group's WIP level is converged, the output-to-input couplings between the models are revised to change a wafer-lot process flow from the low-level group to a mean-delay model. When the ALC detects a divergence caused by a re-entrant flow or a machine-down, the high-level model is switched back to its corresponding low-level group model. The ALC then generates dummy wafer-lot events to synchronize the busyness of high-level steady state. The proposed method was applied to case studies of wafer-fab systems and achieves simulation speedup from 6.1 to 11.8 times with corresponding 2.5 to 5.9% degradation inaccuracy. Moon Gi Seok, Chew Wye Chan, Wentong Cai 0001, Hessam S. Sarjoughian, Daejin Park |
SIGSIM-PADS | 5 |
| 2019 | A high-level modeling and simulation approach using test-driven cellular automata for fast performance analysis of RTL NoC designsabstractThe simulation speedup of designed RTL NoC regarding the packet transmission is essential to analyze the performance or to optimize NoC parameters for various combinations of intellectual-property (IP) blocks, which requires repeated computations for parameter-space exploration. In this paper, we propose a high-level modeling and simulation (M&S) approach using a revised cellular automata (CA) concept to speed up simulation of dynamic flit movements and queue occupancy within target RTL NoC. The CA abstracts the detailed RTL operations with the view of deciding a cell's state of actions (related to moving packet flits and changing the connection between CA cells) using its own high-level states and those of neighbors, and executing relevant operations to the decided action states. During the performing the operations including connection requests and acceptances, architecture-independent and user-developed routing and arbitration functions are utilized. The decision regarding the action states follows a rule set, which is generated by the proposed test environment. The proposed method was applied to an open-source Verilog NoC, which achieves simulation speedup by approximately 8 to 31 times for a given parameter set. Moon Gi Seok, Hessam S. Sarjoughian, Daejin Park |
ASP-DAC | 3 |
| 2019 | Homogeneity patch search method for voting-based efficient vehicle color classification using front-of-vehicle image
Yoosoo Jeong, Kil-Houm Park, Daejin Park |
Multim. Tools Appl. | 3 |
| 2017 | An HLA-Based Distributed Cosimulation Framework in Mixed-Signal System-on-Chip DesignabstractIn mixed-signal system-on-chip (SoC) design, distributed cosimulation is one of the practical approaches for unifying various abstracted hardware models using different description languages. Conventional ad hoc distributed cosimulation solutions do not have formal theoretical backgrounds of simulator integration into their solutions. In this brief, we propose a general cosimulation framework based on the high-level architecture (HLA) and newly defined programming language interface for interoperation (PLI-I) as a formal simulator interface. Based on the PLI-I and HLA, we propose formal integration and interoperation procedures. To reduce integration costs, the procedures have been developed into a common library and then merged with model-dependent signal-event converter to handle differently abstracted in/out signals. During the interoperation, to resolve the different time-advance mechanisms of the digital and analog simulators, the adapter executes an advanced HLA-based synchronization based on the presimulation concepts. The case study shows the reduced design effort in integrating and validating the heterogeneous models and simulators using the proposed framework in mixed-signal SoC design. Moon Gi Seok, Tag Gon Kim, Chang Beom Choi, Daejin Park |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | Framework for simulation of the Verilog/SPICE mixed model: Interoperation of Verilog and SPICE simulators using HLA/RTI for model reusabilityabstractDesigning a mixed-signal integrated hardware requires the mixed simulation for legacy digital blocks and analog circuits, which are usually represented by the Verilog description language for digital blocks and the SPICE circuit netlist of analog circuits. Without model translations or source-level modifications and to simulate mixed legacy Verilog models and SPICE circuit netlists that are usually developed based on the different SPICE languages, parameters and primitives, this paper proposes a simulation framework whose concept is connecting a legacy Verilog and proper SPICE simulator for the target SPICE model using a run-time infrastructure (RTI) based on high level architecture (HLA) and adapters that are pluggable libraries to enable the interoperation and integration of simulators through HLA. For the interoperation, to exchange analog/digital signals, the adapter converts analog/digital signals to events or events to analog/digital signals using user-defined, signal-event converters. To synchronize different time advance policies, the adapter performs time synchronization procedures based on the pre-simulation concept. For the integration of Verilog/SPICE simulators and the RTI, adapters are developed following each component interface, which are IEEE-std Verilog procedural interface, proposed SPICE procedural interface and IEEE-std HLA interface. The proposed framework was applied to the digitally controlled buck converter simulation. Moon Gi Seok, Daejin Park, Geun Rae Cho, Tag Gon Kim |
VLSI-SoC | 2 |
| 2014 | Built-In Binary Code Inversion Technique for On-Chip Flash Memory Sense Amplifier With Reduced Read Current ConsumptionabstractThe bit-line sense amplifier (S/A) for on-chip flash memory compares cell current with reference current to identify data that are programmed. The S/A for 0 (erased) cell data consumes a large sink current, which is greater than off-current for 1 (programmed) cell data. This brief proposes a built-in write/read path based on binary inversion methods to reduce the sensing current of S/A. An original binary code is programmed into flash memory with an inverted binary code based on the proposed bit inversion techniques. The de-inversion hardware, which is implemented with small logic gates to restore original binary data, only consumes logic current instead of analog sink current in the S/A. The proposed techniques are evaluated for the DSPStone benchmark and are applied to the modified S/A for ARM Cortex-M3-based microcontroller with 128-kB on-chip flash memory based on a 0.18-um EEPROM technology. The circuit-level simulation result for the DSPStone benchmark shows that a newly implemented chip with the S/A based on the proposed technique consumes approximately less than 22% of the operating power that conventional S/A uses. Daejin Park, Tag Gon Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | A Low-Power Fractional-Order Synchronizer for Syncless Time-Sequential Synchronization of 3-D TV Active Shutter GlassesabstractThe 3-D TV active shutter glasses (SGs) technology requires the communication of the sync timing for time-sequential frame synchronization, in contrast to the syncless film-type patterned retarder approach. To advance our previous work based on the fractional-order timer, this paper proposes a fractional-order synchronizer, including an adaptive sync reconstructor (ASR) based on FRT for syncless frame synchronization. The FRT enables accurate synchronization regardless of sync-clock speed. The hybrid cooperation of FRT and ASR reduces the required frequency of communication of the sync packets and turns off the emitter on the TV side for perfect syncless operation. The implemented one-chip solution uses less than roughly 11% of the operating current of major commercial SG. The syncless SG technique also minimizes synchronization failure and 3-D vision crosstalk from interruption in sync packets. Daejin Park, Chang-Min Kim, Sungho Kwak, Tag Gon Kim |
IEEE Trans. Circuits Syst. Video Technol. | 1 |