Seongsoo Hong

dblp:73/5284 · DBLP profile ↗
← Back
52ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 14 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Deep Software Stack Optimization for AI-Enabled Embedded Systems
abstract
On-device AI enables real-time, privacy-preserving inference but remains constrained by the limited resources of edge devices. To address these constraints, static optimization techniques are employed to reduce model complexity, yet they often fail to handle dynamic challenges such as bursty input streams. We propose pipelined DNN inference, where a model is partitioned into sub-models that execute concurrently across heterogeneous accelerators. We examine the supporting software stack, introduce a model slicer that ensures correct partitioning, and present a multithreaded inference architecture that overlaps pipeline stages to maintain stable throughput even under surges in input rate.
Namcheol Lee, Geonha Park, Seongsoo Hong
EMSOFT4
2025 DNNPipe: Dynamic programming-based optimal DNN partitioning for pipelined inference on IoT networks
abstract
ABSTRACT Pipeline parallelization is an effective technique that enables the efficient execution of deep neural network (DNN) inference on resource-constrained IoT devices. To enable pipeline parallelization across computing nodes with asymmetric performance profiles, interconnected via low-latency, high-bandwidth networks, we propose DNNPipe, a DNN partitioning algorithm that constructs a pipeline plan for a given DNN. The primary objective of DNNPipe is to maximize the throughput of DNN inference while minimizing the runtime overhead of DNN partitioning, which is repeatedly executed online in dynamically changing IoT environments. To achieve this, DNNPipe uses dynamic programming (DP) with pruning techniques that preserve optimality to explore the search space and find the optimal pipeline plan whose maximum stage time is no greater than that of any other possible pipeline plan. Specifically, it aggressively prunes suboptimal pipeline plans using two pruning techniques: upper-bound-based pruning and under-utilized-stage pruning . Our experimental results demonstrate that pipelined inference using an obtained optimal pipeline plan improves DNN throughput by up to 1.78 times compared to the highest performing single device and DNNPipe achieves up to 98.26% lower runtime overhead compared to PipeEdge, the fastest known optimal DNN partitioning algorithm.
Woobean Seo, Saehwa Kim, Seongsoo Hong
J. Syst. Archit.3
2025 Inference framework supporting parallel execution across heterogeneous accelerators
Philkyue Shin, Myungsun Kim, Seongsoo Hong
J. Syst. Archit.3
2024 Dynamic Mapping of Mixed-Criticality Applications onto a Mixed-Criticality Runtime System with Probabilistic Guarantees
abstract
The emergence of software-defined vehicles (SDVs) introduces significant challenges in dynamically deploying services with diverse criticality semantics. To address this issue, we present a framework for the dynamic mapping of mixed-criticality applications (MCAs) onto a mixed-criticality runtime system (MCR) with probabilistic guarantees. We model an SDV service, such as a Docker container, as an MCA and provide an MCR based on a finite-state machine. We present an approach that maps the criticality levels of an MCA to those of the MCR, tracks available resources in the MCR, converts the resource demands of an MCA, and performs admission control to ensure the MCR remains schedulable. This framework enables the reliable and prioritized execution of critical SDV functions while appropriately managing less critical tasks.
Namcheol Lee, Seongsoo Hong, Saehwa Kim
ICDCS2
2024 Partitioning Deep Neural Networks for Optimally Pipelined Inference on Heterogeneous IoT Devices with Low Latency Networks
abstract
Pipeline parallelization is an effective technique that enables the efficient execution of deep neural network (DNN) inference on resource-constrained IoT devices. To support pipeline parallelization on heterogeneous computing nodes with low-latency networks, we propose DNNPipe, a DNN partitioning algorithm that constructs a pipeline plan for a given DNN. The primary objective of DNNPipe is to maximize the throughput of DNN inference while minimizing the runtime overhead of DNN partitioning, which is repeatedly executed in dynamically changing IoT environments. To achieve this, DNNPipe uses dynamic programming for an exhaustive exploration to find the optimal pipeline plan whose maximum stage execution time is no greater than that of any other possible pipeline plan. Additionally, it aggressively prunes suboptimal pipeline plans using an upper bound on the minimum value among all possible pipeline plans' maximum stage execution times. Experimental results demonstrate that DNNPipe significantly reduces its execution time and iteration counts compared to PipeEdge, the fastest known optimal DNN partitioning algorithm.
Woobean Seo, Saehwa Kim, Seongsoo Hong
ICDCS3
2023 Memory-Aware DVFS Governing Policy for Improved Energy-Saving in the Linux Kernel
abstract
Energy-aware computing is one of the most critical issues in modern computing systems. Linux has introduced the schedutil governor since its 4.7 kernel release, which dynamically scales the processor frequency level to reduce energy consumption. Although it has been widely used in most Linux-based systems, the governor often makes inaccurate decisions in frequency selection, thus leading to unnecessary energy consumption. In this paper, we propose an enhanced governor as an alternative. We first rigorously analyze the schedutil's policy and then find out that it does not take into account memory stalls when characterizing CPU performance via CPI (cycles per instruction). This yields inherent inaccuracy, particularly in modern SoCs where multiple CPU cores and accelerators incur a huge amount of memory traffic over the system bus. To rectify this problem, we reformulate the CPU performance estimation function of the schedutil governor via the memory stall cycle ratio so that it can dynamically reflect the effects of changes in the processor's frequency and the system's memory contention. We show that our CPU performance estimation function is easily integrated into schedutil. We also show that the memory stall cycle ratio, the key element of the function, can be efficiently calculated at runtime with performance monitoring units (PMU) commonly available in most modern SoCs. We have implemented our governor and conducted extensive experiments to validate its effectiveness. Experimental results show that our governor saves more energy than schedutil by up to 28.91% without noticeable performance degradation.
Philkyue Shin, Dahun Kim, Seongsoo Hong
RTCSA3
2018 Fair-share scheduling in single-ISA asymmetric multicore architecture via scaled virtual runtime and load redistribution
Myungsun Kim, Soonhyun Noh, Jinhwa Hyeon, Seongsoo Hong
J. Parallel Distributed Comput.4
2017 Location detection for navigation using IMUs with a map through coarse-grained machine learning
abstract
Location detection or localization supporting navigation has assumed significant importance in the recent past. In particular, techniques that exploit cheap inertial measurement units (IMU), the gyroscope and the accelerometer, have garnered attention, especially in an embedded computing context. However, these sensors measurements are quite unreliable, and it is widely believed that these sensors by themselves are too noisy for localization with acceptable accuracy. Consequently, several lines of work embody other costly alternatives to lower the impact of accumulated errors associated with IMU based approaches, invariably leading to very high energy costs resulting in lowered battery life. In this paper, we show that IMUs are sufficient by themselves if we augment them with known structural or geographical information about the physical area being explored by the user. By using the map of the region being explored and the fact that humans typically walk in a structured manner, our approach sidesteps the challenges created by noise and concomitant accumulation of error. Specifically, we show that a simple coarse-grained machine learning approach mitigates the effect of the noisy perturbations in the information from our IMUs, provided we have accurate maps. Throughout, we rely on the principle of inexactness in an overarching manner and relax the need for absolute accuracy in return for significant lowering of resource (energy) costs. Notably, our approach is completely independent of any external guidance from sources including GPS, Bluetooth or WiFi support, and is this privacy preserving. Specifically, we show through experimental results that by relying on gyroscope and accelerometer data alone, we can correctly identify the path-segment where the user is walking/running on a known map, as well as the position within the path with an accuracy of 4.3 meters on the average using 0.44 Joules. This is a factor of 27X cheaper in energy lower than the “gold standard” that one could consider based on GPS support which, surprisingly, has an associated error of 8.7 meters on the average.
E. J. Jose Gonzalez, Chen Luo 0001, Anshumali Shrivastava, Krishna V. Palem, Yongshik Moon, Soonhyun Noh, Daedong Park, Seongsoo Hong
DATE8
2017 Reducing Energy Consumption of a Modem via Selective Packet Transmission Delaying
abstract
Most practical energy-saving mechanisms for a modem rely on a packet transmission delaying technique where the smartphone attempts to reduce the energy consumption of the modem by delaying transmissions of delay-tolerant packets and piggybacking them onto later packets. However, unconditional packet transmission delaying may lead to unanticipated energy loss at the modem since the exact radio resource control state of the modem is not considered. In order to address this problem, we propose a mechanism that selectively delays the packet transmissions only if such delays are expected to achieve energy savings. Our mechanism consists of three key components, which are deferrable packet identifier, pattern-based next packet transmission predictor and packet transmission time designator. In order to demonstrate the effectiveness, we have implemented all three components into Android 4.4 KitKat running on a Google Nexus 5 smartphone and then performed extensive experiments. Our experiments show that the proposed mechanism reduces the energy consumption of an LTE modem by up to 22.5 percent when compared to that of a legacy LTE modem.
Daedong Park, Philkyue Shin, Seongsoo Hong
MDM3
2017 Providing fair-share scheduling on multicore computing systems via progress balancing
Sungju Huh, Seongsoo Hong
J. Syst. Softw.2
2015 Fair-Share Scheduling for Performance-Asymmetric Multicore Architecture via Scaled Virtual Runtime
abstract
As users begin to demand applications with superior user experience and high service quality, asymmetric multicore processors are increasingly adopted in embedded systems due to their architectural benefits in improved performance and power savings. While fair-share scheduling is a crucial kernel service for such applications, it is still in an early stage when it comes to performance-asymmetric multicore architecture. In this paper, we propose a new fair-share scheduler by adopting the notion of scaled CPU time which reflects performance asymmetry between different types of cores. Our scheduler can work with kernel's dynamic resource control mechanisms since it makes use of a varying performance ratio between cores and thus captures dynamic performance asymmetry such as a core's changing operating frequency. We develop our approach on top of ARM's big. LITTLE architecture which runs Linaro's scheduling framework. Since Linaro's relies on the completely fair scheduler(CFS) of the Linux kernel and CFS is virtual runtime based, we revise the notion of virtual runtime using the scaled CPU time and incorporate it into the proposed approach. As a result, our approach achieves fair-share scheduling by simply balancing tasks' virtual runtimes. To demonstrate its effectiveness, we have implemented the proposed scheduler and performed a series of experiments on ARM's Versatile Express TC2 board. We ran the SPEC CPU2006 and PARSEC benchmarks for three minutes and measured tasks' virtual runtimes. We observed that the maximum virtual runtime difference was only 0.69 seconds in our approach while the original CFS yielded the maximum difference of 8.35 seconds.
Myungsun Kim, Soonhyun Noh, Sungju Huh, Seongsoo Hong
RTCSA4
2015 Cross-layer resource control and scheduling for improving interactivity in Android
abstract
Summary Android smartphones are often reported to suffer from sluggish user interactions due to poor interactivity. This is partly because Android and its task scheduler, the completely fair scheduler (CFS), may incur perceptibly long response time to user‐interactive tasks. Particularly, the Android framework cannot systemically favor user‐interactive tasks over other background tasks since it does not distinguish between them. Furthermore, user‐interactive tasks can suffer from high dispatch latency due to the non‐preemptive nature of CFS. To address these problems, this paper presents framework‐assisted task characterization and virtual time‐based CFS. The former is a cross‐layer resource control mechanism between the Android framework and the underlying Linux kernel. It identifies user‐interactive tasks at the framework‐level, by using the notion of a user‐interactive task chain. It then enables the kernel scheduler to selectively promote the priorities of worker tasks appearing in the task chain to reduce the preemption latency. The latter is a cross‐layer refinement of CFS in terms of interactivity. It allows a task to be preempted at every predefined period. It also adjusts the virtual runtimes of the identified user‐interactive tasks to ensure that they are always scheduled prior to the other tasks in the run‐queue when they wake up. As a result, the dispatch latency of a user‐interactive task is reduced to a small value. We have implemented our approach into Android 4.1.2 running with Linux kernel 3.0.31. Experimental results show that the response time of a user interaction is reduced by up to 77.35% while incurring only negligible overhead. Copyright © 2014 John Wiley & Sons, Ltd.
Sungju Huh, Jonghun Yoo, Seongsoo Hong
Softw. Pract. Exp.3
2014 Utilization-aware load balancing for the energy efficient operation of the big.LITTLE processor
abstract
ARM's big.LITTLE architecture introduces the opportunity to optimize power consumption by selecting the core type most suitable for a level of processing demand. To take advantage of this new axis of optimization, we introduce processor utilization into the Linux kernel's load balancing algorithm. Our method improves the Linux kernel's ability to schedule tasks in an energy efficient manner without making it directly aware of the available core types. Experimental results show an energy consumption improvement over the standard Linux scheduler up to 11.35% with almost no reduction in performance.
Myungsun Kim, Kibeom Kim, James R. Geraci, Seongsoo Hong
DATE4
2013 Petri Net-Based FTL Architecture for Parametric WCET Estimation via FTL Operation Sequence Derivation
abstract
A flash translation layer (FTL) provides file systems with transparent access to NAND flash memory. Although many applications running on it require real-time guarantees, it is difficult to provide tight worst case execution time (WCET) bounds with conventional static WCET analysis since an FTL exhibits a large variance in execution time depending on its runtime state. Parametric WCET analysis could be an effective alternative but it is also challenging to formulate a parametric WCET function for an FTL program because traditional FTL architecture does not properly model the runtime availability of flash resources in its code structure. To overcome such a limitation, we propose Petri net-based FTL architecture where a Petri net explicitly specifies dependencies between FTL operations and the runtime resource availability. It comes with an FTL operation sequencer that derives at runtime the shortest sequence of FTL operations for servicing an incoming FTL request under the current resource availability. The sequencer computes the WCET of the request by merely summing the WCETs of only those FTL operations in the sequence. Our experimental results show the effectiveness of our FTL architecture. It allowed for tight WCET estimation that yielded WCETs shorter by a factor of 54 than statically analyzed ones.
Jonghun Yoo, Jaesoo Lee, Seongsoo Hong
IEEE Trans. Computers3
2012 Providing Fair Share Scheduling on Multicore Cloud Servers via Virtual Runtime-based Task Migration Algorithm
abstract
While Linux is the most favored operating system for an open source-based cloud data center, it falls short of expectations when it comes to fair share multicore scheduling. The primary task scheduler of the mainline Linux kernel, CFS, cannot provide a desired level of fairness in a multicore system. CFS uses a weight-based load balancing mechanism to evenly distribute task weights among all cores. Contrary to expectations, this mechanism cannot guarantee fair share scheduling since balancing loads among cores has nothing to do with bounding differences in the virtual runtimes of tasks. To make matters worse, CFS allows a persistent load imbalance among cores. This paper presents a virtual runtime-based task migration algorithm which directly bounds the maximum virtual runtime difference among tasks. For a given pair of cores, our algorithm periodically partitions run able tasks into two groups depending on their virtual runtimes and assigns each group to a dedicated core. In doing so, it bounds the load difference between two cores by the largest weight in the task set and makes the core with larger virtual runtimes receive a larger load and thus run more slowly. It bounds the virtual runtime difference of any pair of tasks running on these cores by a constant. We have implemented the algorithm into the Linux kernel 2.6.38.8. Experimental results show that the maximal virtual runtime difference is 50.53 time units while incurring only 0.14% more run-time overhead than CFS.
Sungju Huh, Jonghun Yoo, Myungsun Kim, Seongsoo Hong
ICDCS4
2012 Improving Interactivity via VT-CFS and Framework-Assisted Task Characterization for Linux/Android Smartphones
abstract
Android smart phones are often reported to suffer from sluggish user interactions due to poor interactivity. This is because the Linux kernel may incur perceptibly long response time to user interactive tasks. Particularly, the completely fair scheduler (CFS) of Linux cannot systematically favor a user interactive task over background tasks since it fails to effectively distinguish between them. Even if a user interactive task is successfully identified, such a task can suffer from a high scheduling latency due to the non-preemptive nature of CFS. This paper presents a framework-assisted task characterization and virtual runtime-based CFS (VT-CFS) to address these problems. The former is a cooperative mechanism between the Android application framework and the kernel. It identifies a user interactive task at the framework level and then enables the task scheduler to selectively promote the priority of the identified task at the kernel level. VT-CFS is an extension of the CFS. It allows a task to be preempted at any preemption tick so that the scheduling latency of a user interactive task is bounded by the tick interval. We have implemented our approach into Android 2.2 running with Linux kernel 2.6.32. Experimental results show that the response time of a user interactive task is reduced by up to 31.4% while incurring only 0.9% more run-time overhead than the legacy system.
Sungju Huh, Jonghun Yoo, Seongsoo Hong
RTCSA3
2011 Providing Network Performance Isolation in VDE-Based Cloud Computing Systems
abstract
In a cloud computing system, virtual machines owned by different clients are co-hosted on a single physical machine. It is vital to isolate network performance between the clients for ensuring fair usage of the constrained and shared network resources of the physical machine. Unfortunately, the existing network performance isolation techniques are not effective for cloud computing systems because they are difficult to be adopted in a large scale and require non-trivial modification to the network stack of a guest OS. In this paper, we propose a performance isolation-enabled virtual distributed Ethernet (PIE-VDE) to overcome such difficulties. It is a network virtualization software module running on a host OS. It intends to (1) allocate fair share of outgoing link bandwidth to the co-hosted clients and (2) divide a client's share to the virtual machines owned by it in a fair way. Our approach supports full virtualization of a guest OS, ease in wide scale adoption, limited modification to the existing system, low run-time overhead and work-conserving servicing. Experimental results show the effectiveness of the proposed mechanism. Every client received at least 99.5% of its bandwidth share as specified by its weight.
Vijeta Rathore, Jonghun Yoo, Jaesoo Lee, Seongsoo Hong
HPCC4
2011 Preventing TCP performance interference on asymmetric links using ACKs-first variable-size queuing
Daedong Park, Seongsoo Hong, Jungkeun Park
Comput. Commun.3
2008 Scheduler-Assisted Prefetching: Efficient Demand Paging for Embedded Systems
abstract
Embedded systems tend to use demand paging in order to provide more memory to applications in a cost-effective manner. However, demand paging drastically degrades the performance when the page fault rate is high. Prefetching has been known as a common remedy for page fault overhead. Although many prefetching mechanisms have been proposed, they are either effective only for specific page access patterns or too straight-forward to decrease a page fault rate to an acceptable level. We propose a scheduler-assisted prefetching mechanism which does not have such fundamental defects. As a proof of concept, our mechanism was completely implemented in Linux. We have also conducted a series of experiments to show its effectiveness. The experimental results showed a significant improvement: the number of the major page faults and the scheduling latency decreased by 30% and 51%, respectively.
Stanislav A. Belogolov, Jungkeun Park, Seongsoo Hong
RTCSA4
2008 CREAM: A Generic Build-Time Component Framework for Distributed Embedded Systems
abstract
A component framework plays an important role in CBSD as it determines how software components are developed, packaged, assembled and deployed. A desirable component framework for developing diverse cross-domain embedded applications should meet such requirements as (1) lightweight on memory use, (2) integrated task execution model, (3) fast inter-component communication, (4) support for distributed processing, and (5) transparency from underlying communication middleware. Although current embedded system component frameworks address some of the above requirements, they fail to meet all of them taken together. We thus propose a new embedded system component framework called CREAM (component-based remote-communicating embedded application model). It achieves these goals by using build-time code generation, explicit control of task creation and execution in the component framework, static analysis of component composition to generate efficient component binding, and abstraction of the componentpsilas application logic from the communication middleware. We have implemented the CREAM component framework and conducted a series of experiments to compare its performance characteristics to a raw socket-based communication implementation and the lightweight-CCM implementation by MicoCCM.
Chetan Raj, Jungkeun Park, Seongsoo Hong
RTCSA4
2008 A Goal-oriented Mixed-granularity Component Selection Method for Huge Component Repositories
Xiaolin Xi, Seongsoo Hong
SEKE4
2008 Quasistatic shared libraries and XIP for memory footprint reduction in MMU-less embedded systems
abstract
Despite a rapid decrease in the price of solid state memory devices, system memory is still a very precious resource in embedded systems. The use of shared libraries and execution-in-place (XIP) is known to be effective in significantly reducing memory usage. Unfortunately, many resource-constrained embedded systems lack an MMU, making it extremely difficult to support these techniques. To address this problem, we propose a novel shared library technique called a quasi-static shared library and an XIP, both based on our enhanced position independent code technique. In our quasistatic shared libraries, global symbols are bound to pseudoaddresses at linking time and actual physical addresses are bound at loading time. Unlike conventional shared libraries, they do not require symbol tables that take up valuable memory space and, therefore, allow for expedited address translation at runtime. Our XIP technique is facilitated by our enhanced position independent code where a data section can be arbitrarily located. Both the shared library and XIP techniques are made possible by emulating an MMU's memory mapping feature with a data section base register (DSBR) and a data section base table (DSBT). We have implemented these proposed techniques in a commercial ADSL (Asymmetric Digital Subscriber Line) home network gateway equipped with an MMU-less ARM7TDMI processor core, 2MB flash memory, and 16MB RAM. We measured its memory usage and evaluated its performance overhead by conducting a series of experiments. These experiments clearly demonstrate the effectiveness of our techniques in reducing memory usage. The results are impressive: 35% reduction in flash memory usage when using only the shared library and 30% reduction in RAM usage when using the shared library and XIP together. These results were achieved with only a negligible performance penalty of less than 4%. Even though these techniques were applied to uClinux-based embedded systems, they can be used for any MMU-less real-time operating system.
Jaesoo Lee, Saehwa Kim, Seongsoo Hong
ACM Trans. Embed. Comput. Syst.4
2007 Preventing network performance interference with ACK-separation queuing mechanism in a home network gateway using an asymmetric link
abstract
In network-enabled consumer electronics development, much of the time and effort is spent analyzing and solving network performance problems. In this paper, we define an instance of such problems discovered while developing a commercial home network gateway. We then analyze its cause and propose a solution mechanism. Our home network gateway uses an asymmetric link (ADSL) and suffers from an undesirable phenomenon where downlink traffic interferes with upload speed. We call this phenomenon the network performance interference problem. While this problem can easily be confused with receive livelock caused by packet contention at the input queue, we find that this is not the case. By performing extensive experiments and analysis, we reveal that our problem is caused by packet contention at the output queue and certain intrinsic characteristics of TCP. We devise an ACK-separation queuing mechanism for this problem and implement it in the home network gateway. Our experiments show that it effectively solves the problem.
Seongsoo Hong
RTCSA2
2006 The robot software communications architecture (RSCA): embedded middleware for networked service robots
abstract
In this paper, we present a robot middleware technology named robot software communications architecture (RSCA) for its use in networked home service robots. The RSCA provides a standard operating environment for the robot applications together with a framework that expedites the development of such applications. The operating environment is comprised of a real-time operating system, a communication middleware, and a deployment middleware. Particularly, the deployment middleware supports the reconfiguration of component-based robot applications including installation, creation, start, stop, tear-down, and un-installation. In designing RSCA, we have adopted a middleware called SCA from the software defined radio domain and extend it since the original SCA lacks the real-time guarantees and appropriate event services. We have fully implemented RSCA and performed measurements to quantify its run-time performance. Our implementation clearly shows the viability of RSCA.
Seongsoo Hong, Jaesoo Lee, Hyeonsang Eom, Gwangil Jeon
IPDPS1
2006 Design Patterns for Releasing Applications in C++ Implementations of JTRS Software Communications Architecture
abstract
The software communications architecture (SCA), which has been adopted as an SDR (software defined radio) Forum standard, provides a framework that successfully exploits common design patterns of distributed, real-time, and object-oriented embedded systems software. We have fully implemented the SCA v2.2 in C++. During this implementation process, we have encountered the lack of a suitable design pattern for releasing the SCA applications. Unfortunately, design patterns for releasing objects have been neither extensively addressed nor well investigated as opposed to creational design patterns. This is largely due to the fact that such releasing design patterns are highly dependent on programming languages. In this paper, we investigate three viable design patterns for releasing the SCA applications in C++ and discuss their pros and cons. In addition, we select the most portable and thus most reusable pattern, which we name Vulture design pattern, among those alternatives and detail our specific implementation.
Michael Barth, Jonghun Yoo, Saehwa Kim, Seongsoo Hong
ISORC4
2006 Scenario-based multitasking for real-time object-oriented models
Saehwa Kim, Seongsoo Hong
Inf. Softw. Technol.3
2006 Rapid performance re-engineering of distributed embedded systems via latency analysis and k-level diagonal search
Jungkeun Park, Minsoo Ryu, Seongsoo Hong, Lucia Lo Bello
J. Parallel Distributed Comput.3
2006 Q-SCA: Incorporating QoS support into software communications architecture for SDR waveform processing
Jaesoo Lee, Saehwa Kim, Seongsoo Hong
Real Time Syst.4
2005 Extending Software Communications Architecture for QoS Support in SDR Signal Processing
abstract
The software communications architecture (SCA) is middleware for providing component interfaces and dynamic reconfigurability to software define radio (SDR) modems and it is widely adopted by the SDR forum. Although its main application domains include real-time signal processing, SCA lacks QoS support. We thus propose a QoS-enabled extension of SCA called Q-SCA. Our extended architecture incorporates a new waveform model, extended domain profiles that describe QoS requirements, and a modified application instantiation process that supports both admission control and resource allocation. We show the viability of Q-SCA through full implementation and experiments.
Jaesoo Lee, Seunghyun Han, Seongsoo Hong
RTCSA4
2005 Resource conscious development of middleware for control environments: a case of CORBA-based middleware for the CAN bus systems
Seongsoo Hong, Tae-Hyung Kim 0003
Inf. Softw. Technol.1
2004 Experimental Assessment of Scenario-Based Multithreading for Real-Time Object-Oriented Models: A Case Study with PBX Systems
Saehwa Kim, Michael Buettner, Mark Hermeling, Seongsoo Hong
EUC4
2004 State Machine Based Operating System Architecture for Wireless Sensor Networks
Tae-Hyung Kim 0003, Seongsoo Hong
PDCAT2
2003 Embedded Linux Outlook in the PostPC Industry
abstract
This paper presents an analysis of the future of Embedded Linux in the PostPC industry. This analysis is carried out by first examining the current forces at work in the PostPC market and how they effect Linux. Next, we look at the future trends in the PostPC market for the many types of devices we see in the market now and will see in the future and how well these trends can be addressed by Embedded Linux and other RTOSes in the marketplace. The third step in this analysis is to examine the qualities of Linux itself and what gives it an advantage in the marketplace as well as the challenges that accompany using Linux as an embedded platform. We address the challenges that Embedded Linux faces by proposing the need for the standardization of Embedded Linux. This means building a core standard for a kernel that is optimized for embedded systems, standard middleware profiles for powering different applications, and a standard for addressing intellectual property and software components.
Seongsoo Hong
ISORC1
2003 Deterministic and Statistical Deadline Guarantees for a Mixed Set of Periodic and Aperiodic Tasks
Minsoo Ryu, Seongsoo Hong
RTCSA2
2001 Rapid Re-Engineering of Embedded Real-Time Systems via Cost-Benefit Analysis with K-Level Diagonal Searc
abstract
This paper formulates a problem of embedded real-time system re-engineering and presents a systematic solution. Embedded real-time system re-engineering is defined as an understanding and alteration of a legacy system to guarantee newly imposed performance requirements. The performance requirements may include a real-time throughput and an input-to-output latency. The proposed approach is based on bottleneck analysis and nonlinear optimization. The inputs to the approach include a system design specified with a process network accompanied by task graphs and task schedules, and a new real-time throughput requirement specified as a system's period constraint. The output is a set of scaling factors that represent the ratios of performance upgrades for processing elements. The solution approach works in two steps. First, it identifies bottleneck processes by estimating process latencies and by analyzing resource sharing among processes. It then derives a set of linear constraints from the new throughput requirement for bottleneck processes. Second, it formulates an integer nonlinear optimization problem and solves it for scaling factors with an objective of minimizing the hardware upgrade cost. Resultant scaling factors are used for cost-effective upgrades of processing elements. To efficiently find feasible solutions, we propose the k-level diagonal search algorithm which runs in a polynomial time with respect to the number of processing elements. Simulation results also confirm this assertion.
Jungkeun Park, Minsoo Ryu, Seongsoo Hong, Lucia Lo Bello
RTSS3
2001 Timing Constraint Remapping to Achieve Time Equi-Continuity in Distributed Real-Time Systems
abstract
Discretely synchronized, distributed real-time systems may suffer from a time discontinuity problem in that local clocks observe the disappearance or reappearance of time intervals. This problem occurs since traditional discrete clock synchronization algorithms adjust local clocks instantaneously. Such time discontinuities may lead to runtime faults due to the loss or gain of critical time points such as task release times and deadlines. In this paper, we propose a dynamic constraint transformation technique we call a constraint transformation for equi-continuity (CTEC) to correctly enforce timing requirements in a distributed real-time system possessing periodically synchronized distributed local clocks. While continuous clock synchronization is generally suggested to avoid the time discontinuity problem, it incurs too much runtime overhead to be implemented in software. The proposed CTEC technique can solve the time discontinuity problem without modifying discrete clock synchronization algorithms. The CTEC, working as an added component of discrete clock synchronization, moves timing constraints out of discontinuous time intervals. In doing so, it makes use of a mapping derived from continuous clock synchronization in order to exploit its continuity property. We formally prove the correctness of CTEC by showing that the CTEC with discrete clock synchronization generates the same task schedule as continuous clock synchronization. In order to show the effectiveness of CTEC, we have implemented it on a distributed platform based on the CAN bus and performed extensive experiments. The experimental results indicate that time discontinuities present a consistency problem to real-world systems. They also show that CTEC is an effective solution to the problem while incurring little run-tine overhead.
Minsoo Ryu, Jungkeun Park, Seongsoo Hong
IEEE Trans. Computers3
2001 Bounding Cache-Related Preemption Delay for Real-Time Systems
abstract
Cache memory is used in almost all computer systems today to bridge the ever increasing speed gap between the processor and main memory. However, its use in multitasking computer systems introduces additional preemption delay due to the reloading of memory blocks that are replaced during preemption. This cache-related preemption delay poses a serious problem in realtime computing systems where predictability is of utmost importance. We propose an enhanced technique for analyzing and thus bounding the cache-related preemption delay in fixed-priority preemptive scheduling focusing on instruction caching. The proposed technique improves upon previous techniques in two important ways. First, the technique takes into account the relationship between a preempted task and the set of tasks that execute during the preemption when calculating the cache-related preemption delay. Second, the technique considers the phasing of tasks to eliminate many infeasible task interactions. These two features are expressed as constraints of a linear programming problem whose solution gives a guaranteed upper bound on the cache-related preemption delay. This paper also compares the proposed technique with previous techniques using randomly generated task sets. The results show that the improvement on the worst-case response time prediction by the proposed technique over previous techniques ranges between 5 percent and 18 percent depending on the cache refill time when the task set utilization is 0.6. The results also show that as the cache refill time increases, the improvement increases, which indicates that accurate prediction of cache-related preemption delay by the proposed technique becomes increasingly important if the current trend of widening speed gap between the processor and main memory continues.
Chang-Gun Lee, Kwangpo Lee, Joosun Hahn, Yang-Min Seo, Sang Lyul Min, Rhan Ha, Seongsoo Hong, Chang Yun Park, Minsuk Lee, Chong-Sang Kim
IEEE Trans. Software Eng.7
2000 Resource-Conscious Customization of CORBA for CAN-Based Distributed Embedded Systems
abstract
While recently emerging middleware technologies such as CORBA and DCOM address the complexity of distributed programming, they cannot be directly applied to distributed control system design due to their excessive resource demand and inadequate communication models. We propose a new CORBA design for CAN-based distributed embedded control systems. Our design goal is to minimize its resource need and make it support group communication without losing the IDL (interface definition language) level compliance to the OMG standards. To achieve this, we develop a transport protocol on the CAN and a group communication scheme based on the well-known publisher/subscriber model. The protocol effectively realizes subject-based addressing and supports anonymous publisher/subscriber communication. We also customize the method invocation and message passing protocol, referred to as the general inter-ORB protocol (GIOP), of CORBA so that CORBA method invocations are efficiently serviced on a low-bandwidth network such as the CAN. This customization includes packed data encoding and variable-length integer encoding for compact representation of IDL data types. The new CORBA design clearly demonstrates that it is feasible to use CORBA in developing distributed embedded systems on real-time networks possessing severe resource limitations.
Kimoon Kim, Gwangil Jeon, Seongsoo Hong, Sunil Kim, Tae-Hyung Kim 0003
ISORC3
1999 A Period Assignment Algorithm for Real-Time System Design
Minsoo Ryu, Seongsoo Hong
TACAS2
1999 Experimental Assessment of the Period Calibration Method: A Case Study
Namyun Kim, Minsoo Ryu, Seongsoo Hong, Heonshik Shin
Real Time Syst.3
1998 Network conscious design of distributed real-time systems
Jung Woo Park, Young Shin Kim, Seongsoo Hong, Manas Saksena, Sam H. Noh, Wook Hyun Kwon
J. Syst. Archit.3
1998 Analysis of Cache-Related Preemption Delay in Fixed-Priority Preemtive Scheduling
abstract
We propose a technique for analyzing cache-related preemption delays of tasks that cause unpredictable variation in task execution time in the context of fixed-priority preemptive scheduling. The proposed technique consists of two steps. The first step performs a per-task analysis to estimate cache-related preemption cost for each execution point in a given task. The second step computes the worst case response time of each task that includes the cache-related preemption delay using a response time equation and a linear programming technique. This step takes as its input the preemption cost information of tasks obtained in the first step. This paper also compares the proposed approach with previous approaches. The results show that the proposed approach gives a prediction of the worst case cache-related preemption delay that is up to 60 percent tighter than those obtained from the previous approaches.
Chang-Gun Lee, Joosun Hahn, Yang-Min Seo, Sang Lyul Min, Rhan Ha, Seongsoo Hong, Chang Yun Park, Minsuk Lee, Chong-Sang Kim
IEEE Trans. Computers6
1997 Enhanced analysis of cache-related preemption delay in fixed-priority preemptive scheduling
abstract
We propose an enhanced technique for analyzing, and thus bounding cache related preemption delay in fixed priority preemptive scheduling focusing on instruction caching. The proposed technique improves upon previous techniques in two important ways. First, the technique takes into account the relationship between a preempted task and the set of tasks that execute during the preemption when calculating the cache related preemption delay. Second, the technique considers phasing of tasks to eliminate many infeasible task interactions. These two features are expressed as constraints of a linear programming problem whose solution gives a guaranteed upper bound on the cache related preemption delay. The paper also compares the proposed technique with previous techniques. The results show that the proposed technique gives up to 60% tighter prediction of the worst case response time than the previous techniques.
Chang-Gun Lee, Joosun Hahn, Yang-Min Seo, Sang Lyul Min, Rhan Ha, Seongsoo Hong, Chang Yun Park, Minsuk Lee, Chong-Sang Kim
RTSS6
1997 Slicing Real-Time Programs for Enhanced Schedulability
abstract
In this article we present a compiler-based technique to help develop correct real-time systems. The domain we consider is that of multiprogrammed real-time applications, in which periodic tasks control physical systems via interacting with external sensors and actuators. While a system is up and running, these operations must be performed as specified—otherwise the system may fail. Correctness depends not only on each program individually, but also on the time-multiplexed behavior of all of the programs running together. Errors due to overloaded resources are exposed very late in a development process, and often at runtime. They are usually remedied by human-intensive activities such as instrumentation, measurement, code tuning and redesign. We describe a static alternative to this process, which relies on well-accepted technologies from optimizing compilers and fixed-priority scheduling. Specifically, when a set of tasks are found to be overloaded, a scheduling analyzer determines candidate tasks to be transformed via program slicing. The slicing engine decomposes each of the selected tasks into two fragments: one that is “time critical” and the other “unobservable.” The unobservable part is then spliced to the end of the time-critical code, with the external semantics being maintained. The benefit is that the scheduler may postpone the unobservable code beyond its original deadline, which can enhance overall schedulability. While the optimization is completely local, the improvement is realized globally, for the entire task set.
Richard Gerber 0001, Seongsoo Hong
ACM Trans. Program. Lang. Syst.2
1996 Resource Conscious Design of Distributed Real-Time Systems: An End-to-End Approach
abstract
We present a resource conscious approach to designing distributed real-time systems. This work extends our original solution (Gerber et al., 1995), which was limited to single processor systems. Starting from a given task graph, and a set of end-to-end constraints, we systematically generate task attributes (e.g., periods and deadlines) such that: the task set is schedulable; and the end-to-end constraints are satisfied. The methodology can be mostly automated, and provides useful feedback to a designer when it fails to find a solution. We expect that the techniques presented in this paper will help reduce the laborious process of designing a real-time system, by bringing resource contention and schedulability aspects early into the design process.
Manas Saksena, Seongsoo Hong
ICECCS2
1996 Visual assessment of a real-time system design: a case study on a CNC controller
abstract
We describe our experiments on a real-time system design, focusing on design alternatives such as scheduling jitter, sensor-to-output latency, intertask communication schemes and the system utilization. The prime objective of these experiments was to evaluate a real-time design produced using the period calibration method (Gerber et al., 1995) and thus identify the limitations of the method. We chose a computerized numerical control (CNC) machine as our target real-time system and built a realistic controller and a plant simulator. Our results were extracted from a controlled series of more than a hundred test controllers obtained by varying four test variables. This study unveils many interesting facts: average sensor-to-output latency is one of the most dominating factors in determining control quality; the effect of scheduling jitter appears only when the average sensor-to-output latency is sufficiently small; and loop processing periods are another dominating factor of performance. Based on these results, we propose a new communication scheme and a new objective function for the period calibration method.
Namyun Kim, Minsoo Ryu, Seongsoo Hong, Manas Saksena, Chong-Ho Choi, Heonshik Shin
RTSS3
1996 Analysis of cache-related preemption delay in fixed-priority preemptive scheduling
abstract
We propose a technique for analyzing cache-related preemption delays of tasks that cause unpredictable variation in task execution time in the context of fixed-priority preemptive scheduling. The proposed technique consists of two steps. The first step performs a per-task analysis to estimate cache-related preemption cost for each execution point in a given task from the number of useful cache blocks at the execution point. The second step computes the worst case response time of each task using a response time equation and a linear programming technique which takes as its input the preemption cost information of tasks obtained in the first step. Our experimental results show that the proposed technique gives a prediction of the worst case cache-related preemption delay that is up to 60% tighter than that obtained from previous approaches.
Chang-Gun Lee, Joosun Hahn, Sang Lyul Min, Rhan Ha, Seongsoo Hong, Chang Yun Park, Minsuk Lee, Chong-Sang Kim
RTSS5
1995 Compiling Real-Time Programs With Timing Constraint Refinement and Structural Code Motion
abstract
We present a programming language called TCEL (Time-Constrained Event Language), whose semantics are based on time-constrained relationships between observable events. Such a semantics infers only those timing constraints necessary to achieve real-time correctness, without overconstraining the system. Moreover, an optimizing compiler can exploit this looser semantics to help tune the code, so that its worst-case execution time is consistent with its real-time requirements. In this paper we describe such a transformation system, which works in two phases. First, the TCEL source code is translated into an intermediate representation. Then an instruction-scheduling algorithm rearranges selected unobservable operations and synthesizes tasks guaranteed to respect the original event-based constraints.>
Richard Gerber 0001, Seongsoo Hong
IEEE Trans. Software Eng.2
1995 Guaranteeing Real-Time Requirements With Resource-Based Calibration of Periodic Processes
abstract
The paper presents a comprehensive design methodology for guaranteeing end to end requirements of real time systems. Applications are structured as a set of process components connected by asynchronous channels, in which the end points are the system's external inputs and outputs. Timing constraints are then postulated between these inputs and outputs; they express properties such as end to end propagation delay, temporal input sampling correlation, and allowable separation times between updated output values. The automated design method works as follows: First new tasks are created to correlate related inputs, and an optimization algorithm, whose objective is to minimize CPU utilization, transforms the end to end requirements into a set of intermediate rate constraints on the tasks. If the algorithm fails, a restructuring tool attempts to eliminate bottlenecks by transforming the application, which is then resubmitted into the assignment algorithm. The final result is a schedulable set of fully periodic tasks, which collaboratively maintain the end to end constraints.>
Richard Gerber 0001, Seongsoo Hong, Manas Saksena
IEEE Trans. Software Eng.2
1994 Guaranteeing End-to-End Timing Constraints by Calibrating Intermediate Processes
abstract
This paper presents a comprehensive design methodology for guaranteeing end-to-end requirements of real-time systems.Applications are structured as a set of process components connected by asynchronous channels, in which the endpoints are the system's external inputs and outputs.Timing constraints are then postulated between these inputs and outputs; they express properties such as end-to-end propagation delay, temporal input-sampling correlation, and allowable separation times between updated output values.The automated design method works as follows: First the end-to-end requirements are transformed into a set of intermediate rate constraints on the tasks, and new tasks are created to correlate related inputs.The intermediate constraints are then solved by an optimization algorithm, whose objective is to minimize CPU utilization.If the algorithm fails, a restructuring tool attempts to eliminate bottlenecks by transforming the application, which is then re-submitted into the assignment algorithm.The nal result is a schedulable set of fully periodic tasks, which collaboratively maintain the end-to-end constraints.
Richard Gerber 0001, Seongsoo Hong, Manas Saksena
RTSS2
1993 Compiling Real-Time Programs into Schedulable Code
abstract
We present a programming language with first-class timing constructs, whose semantics is based on timeconstrained relationships between observable events. Since a system specification postulates timing relationships between events, realizing the specification in a program becomes a more straightforward process. Using these constraints, as well as those imposed by data and control flow properties, our objective is to transform the code so that its worst-case execution time is consistent with its real-time requirements. To accomplish this goal we first translate an event-based source program into intermediate code, in which the timing constraints are imposed on the code itself, and then use a compilation technique which synthesizes feasible code from the original source program. 1 Overview The construction of a hard real-time system can, in practice, be a painful process. Many factors conspire to make this the case, among which are inflexible scheduling paradigms and the lack of high-le...
Seongsoo Hong, Richard Gerber 0001
PLDI1
1993 Semantics-based compiler transformations for enhanced schedulability
abstract
We present TCEL (time-constrained event language), whose timing semantics is based solely on the constrained relationships between observable events. Using this semantics, the unobservable code can be automatically moved to convert an unschedulable task set into a schedulable one. We illustrate this by an application of program-slicing, which we use to automatically tune control-domain systems driven by rate monotonic scheduling.>
Richard Gerber 0001, Seongsoo Hong
RTSS2