EDBT 2026 Demo / reviewers in the wild / expert
Soonhoi Ha
dblp:h/SoonhoiHa
· DBLP profile ↗
115ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0001-7472-9142ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 101 · 7 first-author · 16 since 2021Software engineering, systems software and programming languages · 21 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TSIM4ICS: Trace-Driven SystemC-TLM Simulation Framework for I/O Die-Based Multi-Chiplet SystemsabstractThe growing demand for high-performance, energy-efficient, and heterogeneous computing has spurred research on chiplet architectures, which enable modular, scalable, and cost-effective system designs. While most prior studies focus on distributed Network-on-Package(NoP)-based chiplet architectures, this paper addresses multi-chiplet systems that employ an I/O die as a communication hub. The I/O die integrates and standardizes inter-chiplet and external interfaces, thereby enhancing modularity, scalability, multi-vendor integration, and performance-power optimization. We present TSIM4ICS, a trace-driven SystemC-TLM simulator for multi-chiplet systems that estimates end-to-end application performance. Each chiplet model generates communication traces to the I/O die, capturing die-to-die(D2D) links and DDR-interface latency/bandwidth to reveal the impact of remote accesses. Using a multi-NPU chiplet model running partitioned CNN workloads, our simulator allows exploration of the design space across the number of NPUs, chiplets, and workload partitioning, supporting HW/SW co-design. TSIM4ICS is publicly available online1to promote reproducible and practical evaluation of chiplet systems. Youngchul Yoon, Soonhoi Ha |
DATE | 2 |
| 2025 | TSim4CXL: Trace-Driven Simulation Framework for CXL-Based High-Performance Computing Systems
Jaewoo Son, Youngchul Yoon, Soonhoi Ha |
Euro-Par (1) | 3 |
| 2025 | Enabling Decoder-only Language Model Inference on a CNN AcceleratorabstractThe remarkable success of Transformer architectures in Natural Language Processing (NLP) has led to increased demand for embedded systems capable of efficiently handling NLP tasks along with traditional vision tasks based on Convolutional Neural Networks (CNNs) or Vision Transformers. Given that CNN accelerators are already widely adopted commercially, this paper investigates the feasibility of leveraging existing CNN accelerators to handle NLP workloads—particularly decoder-only language models—thus avoiding the high cost of developing dedicated NLP hardware. However, direct use of CNN accelerators for language model poses several challenges due to the distinct characteristics of this network. These include non-computational operations in attention layers, such as head merging/splitting and tensor transpositions, as well as the need for floating point precision in nonlinear operations such as Softmax and RMSNorm, which CNN accelerators typically do not support. Furthermore, inference throughput drops significantly during the generation phase due to memory-bound operations, in contrast to the compute-bound nature of convolution. To address these challenges, we propose a set of minimal hardware extensions to existing CNN accelerators to enable efficient decoder-only Transformer inference. We propose a method to execute multi-head attention layers without relying on dedicated reshaping hardware and augment the architecture with a lightweight SIMD coprocessor to handle non-linear operations. Additionally, we present techniques to improve power efficiency during the generation stage. Our experimental results demonstrate that even low-power CNN accelerators can achieve NLP inference throughput comparable to GPUs, opening a promising path toward versatile and cost-effective embedded AI hardware. Seongwoo Choi, Hyunsu Moh, Changjae Yi, Joon Choi, Soonhoi Ha |
ICCAD | 5 |
| 2025 | Pioneering Pathways: The Evolution of Embedded Software Design Methodologies (Keynote)abstractEmbedded software design is growing ever more challenging as increasing user demands drive up the complexity of underlying hardware. Designers must consider real-time performance and resource constraints alongside functional requirements. Since software bugs can result in critical failures, it is essential to ensure design correctness at compile time. In this keynote, I present the evolution of embedded software design methodologies developed to address these challenges. The journey begins with the Ptolemy framework, which specifies system behavior through a hierarchical combination of formal models of computation, including synchronous dataflow (SDF) and finite state machine (FSM) models. Building upon Ptolemy, the PeaCE methodology extends these principles to hardware/software codesign, focusing on efficient software code generation from an extended SDF model. PeaCE further refines the approach by constraining the combination of formal models within a two-level hierarchical task graph. While this methodology enables embedded software generation for specific hardware platforms, its low task granularity limits its broader applicability. The next generation methodology, HOPES is a parallel embedded SW design framework, introducing a novel “programming platform” concept called the universal execution model(UEM), that hides the underlying system architecture from the programmer. UEM employs specification models similar to those in PeaCE but increases task granularity, using dataflow tasks as the scheduling units of the operating system. Formal models offer several advantages for software design. First, critical design errors such as deadlock and buffer overflow can be detected through static analysis of these models. Second, resource requirements and real-time performance can be estimated at compile time. Finally, target code can be automatically synthesized from UEM, significantly reducing manual coding effort. By preserving UEM semantics, the synthesized code is correct by construction. A notable drawback of formal models is the learning curve required to understand their semantics. To simplify behavior specification for users, we propose an evolutionary design methodology in which the mission of multiple robots is specified at the service level. In the HiSARM framework, a service-level script is translated into a UEM specification from which SW code is automatically generated by the HOPES framework. While these design frameworks show great promise in academic settings, their adoption in real-world embedded software design remains a critical next step. Looking forward, we anticipate further evolution by integrating AI technology into each design stage. The pursuit of better and easier ways to design embedded software will continue to drive innovation, presenting exciting new research challenges for the future. Soonhoi Ha |
LCTES | 1 |
| 2025 | Worst case response time analysis for completely fair scheduling in Linux systemsabstractAbstract The popularity of Linux in embedded systems has grown because of its reliability, flexibility, and performance. To ensure these systems meet specific real-time requirements, such as deadlines and throughput, Linux provides support for real-time schedulers and the PREEMPT_RT patch. However, while these tools prioritize high-priority tasks, they can inadvertently compromise the performance of other tasks. Completely Fair Scheduling (CFS) has served as the default scheduling policy in Linux until recently. The CFS is based on the principle that all runnable tasks should share the processor fairly, which helps balance task performance with overall system responsiveness. Despite its benefits, there has been no established method to assess whether real-time requirements are met under CFS. This paper introduces a novel analysis method to estimate the worst-case response time (WCRT) of tasks under CFS, providing a new solution for running real-time tasks in embedded Linux systems. Due to the dynamic nature of CFS, traditional WCRT analysis techniques are not applicable directly. Our technique analyzes how tasks sharing the same processor affect each other, focusing on their vruntime. By examining the bounds of vruntime variation and calculating the maximum interference from other tasks, we effectively estimate the WCRT. We also introduce algorithms that assign the nice values to tasks based on our proposed WCRT analysis technique, ensuring that the real-time requirements are met. We validate the proposed approach through comparative experiments using both a self-developed CFS simulator and an actual Linux system. Our simulator allows for rapid simulations and efficient exploration of various execution scenarios. Through extensive experiments, we empirically validate that our proposed analysis method is efficient with an acceptable level of overestimation. These make it a valuable tool for system verification and design optimization in Linux-based real-time systems. Kyonghwan Yoon, Eunjin Jeong, Woosuk Kang, Jonghyun Choe, Soonhoi Ha |
Real Time Syst. | 5 |
| 2025 | Optimization of Task Allocation for Resource-Constrained Swarm RobotsabstractWhile task allocation of swarm robots has been extensively researched, resource constraints of robots are rarely considered. In this work, we propose two novel task allocation methods robust to robot failures while considering the resource constraint, limited communication range, and deadline constraint of tasks. The first method, STA (static task allocation) method, finds an optimal task allocation solution at compile-time in terms of the minimum expected finish time, using answer set programming. On the other hand, the DTA (dynamic task allocation) method determines the task candidates for each robot at compile-time considering the resource constraint. It lets each robot select a task autonomously at run-time iteratively by exchanging the task allocation information with its neighbor robots. We assess the efficacy of our methods across three distinct environments: a numerical simulation, a swarm robotics simulation, and real robots. Experimental results show that the proposed methods can effectively tolerate robot failures, and the DTA method is superior to the STA method as the probability of robot failure increases. However, the STA method also exhibits consistent performance and superiority when faced with limitations in inter-robot communication. Additionally, we validate the feasibility of our method in a real-world context by conducting experiments with actual robots.Note to Practitioners—The motivation of this work is to explore how to allocate tasks efficiently to swarm robots to ensure timely completion despite occasional robot failures. In search-and-rescue scenarios, such as in the aftermath of a disaster, the effective use of swarm robots is vital, and the time taken to search is crucial to rescuing individuals within a critical time. Various approaches have been proposed to tackle this problem, taking into account time constraints. However, few studies have considered the impact of hardware constraints on robots. To address this issue, this paper proposes two new strategies to find an optimal allocation: the Static Task Allocation (STA) method and the Dynamic Task Allocation (DTA) method. Our methods are evaluated both on real robots and in simulation environments, demonstrating their suitability for practical application. Woosuk Kang, Eunjin Jeong, Sungjun Shim, Soonhoi Ha |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Software Optimization and Design Methodology for Low Power Computer Vision SystemsabstractThis tutorial article addresses a low power computer vision system as an example of a growing application domain of neural networks, exploring various technologies developed to enhance accuracy within the resource and performance constraints imposed by the hardware platform. Focused on a given hardware platform and network model, software optimization techniques, including pruning, quantization, low-rank approximation, and parallelization, aim to satisfy resource and performance constraints while minimizing accuracy loss. Due to the interdependence of model compression approaches, their systematic application is crucial, as evidenced by winning solutions in the Lower Power Image Recognition Challenge (LPIRC) of 2017 and 2018. Recognizing the typical heterogeneity of processing elements in contemporary hardware platforms, the effective utilization through parallelizing neural networks emerges as increasingly vital for performance enhancement. The article advocates for a more impactful strategy—designing a network architecture tailored to a specific hardware platform. For detailed information on each technique, the article provides corresponding references. Soonhoi Ha, Eunjin Jeong |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2025 | A Framework for Multi-Robot Programming: From High-Level Specification to Retargetable DeploymentabstractIn addition to the various requirements that a multi-robot framework should meet, swarm robotics applications also demand robustness, flexibility, and scalability. While several frameworks have been developed for multi-robot operation, they mostly fall short of adequately supporting some of these essential requirements. In this work, we introduce a novel multi-robot programming framework called HiSARM (High-level Specification, Automatic code generation, and Retargetable deployment for Multi-robot systems), designed to assist both mission planners and robot software developers. HiSARM employs a high-level language to enable mission planners to specify collaborative tasks among multiple robots intuitively. From these scripts written in a high-level language, executable robot code is automatically generated. Mission planners can easily customize the executable robot code according to their specific needs within HiSARM. Additionally, HiSARM provides multiple verification environments by introducing a formal intermediate representation and enabling retargetable deployment of the binary file across various domains, including simulation and real robots. For robot software developers, HiSARM eases the development process for robot software developers by categorizing software components and auto-generating swarm-related functions. We tested HiSARM with multiple scenarios in simulation and real robot environments, demonstrating that it effectively supports all the necessary features for multi-robot applications, including swarm operations. Woosuk Kang, Eunjin Jeong, Kyonghwan Yoon, Soonhoi Ha |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2024 | Vision Transformer Inference on a CNN AcceleratorabstractFollowing the remarkable performance demonstrated by the Transformer architecture in the field of computer vision as well as natural language processing (NLP), there is a growing demand for embedded systems capable of executing Vision Transformer (ViT) applications as well as Convolutional Neural Network (CNN) applications efficiently. Since CNN accelerators are already widely used commercially, this paper explores the possibility of using existing CNN accelerators to support ViT rather than developing separate accelerators for each. CNN accelerators inherently have some limitations in efficiently handling operations in transformers: matrix multiplication (MM) operations with two non-constant matrices and nonlinear operations. To overcome these limitations, we first propose a novel technique to efficiently handle MM operations without special reshaping hardware in an adder-tree type CNN accelerator. And we propose an optimal scheduling method to minimize the idle time caused by offloading computation of nonlinear operations of the Transformer. Additionally, we investigate the possibility of executing layer normalization and GELU operations on the accelerator with minor extensions. The experimental results validate the effectiveness of the proposed methods. Changjae Yi, Hyunsu Moh, Soonhoi Ha |
ICCD | 3 |
| 2023 | How to Boost Deep Neural Networks for Computer VisionabstractAs the range of neural network applications has exploded, various model compression techniques have been developed to increase the accuracy of neural networks under the resource constraints given by the hardware platform and the performance constraints required by users. In this perspective paper, the current status and future prospects of individual techniques are briefly summarized. And it presents the importance of understanding the characteristics of the hardware platform and the systematic methodology of applying these techniques harmoniously. Soonhoi Ha |
DAC | 1 |
| 2023 | Fast and Accurate Virtual Prototyping of an NPU with Analytical Memory ModelingabstractAs the application area of convolutional neural networks (CNNs) is fast expanding, the demand for a customized hardware accelerator called a neural processing unit (NPU), is increasing to process them efficiently in terms of execution time and energy consumption. In the design of an NPU, building a fast and accurate virtual prototype enables us to develop a compiler concurrently with the hardware and to explore the micro-architectural design space. Since the memory access latency has a great effect on performance, it is necessary to model the memory access overhead accurately in the virtual prototype. In this work, we propose a novel analytical model for memory access latency, improving the performance estimation accuracy significantly compared with the previous state-of-the-art analytical model by considering the effect of memory access patterns of the NPU on the latency. The proposed high-level virtual prototype achieves an estimated execution time gap within a 6.8% difference from the RTL simulation result. To demonstrate the usefulness of a fast and accurate prototype, we propose a compiler optimization technique and a new DMA logic tailored for the NPU for further performance improvement. Choonghoon Park, Hyunsu Moh, Changjae Yi, Soonhoi Ha |
RSP | 5 |
| 2023 | Energy-Aware Scenario-Based Mapping of Deep Learning Applications Onto Heterogeneous Processors Under Real-Time ConstraintsabstractTo cope with the increasing demand for deep learning applications in embedded systems, emerging embedded devices tend to equip multiple heterogeneous processors, including GPU and deep learning hardware accelerator, called neural processing unit (NPU). It becomes popular to run multiple deep learning (DL) applications simultaneously to provide several functionalities. In this work, we assume that applications have real-time constraints that may vary at run time. While extensive studies have been conducted recently to find an efficient mapping of multiple DL applications on various hardware platforms, they do not consider the constraints imposed by the NPU and the associated software development kit (SDK) in a real embedded platform. In this paper, we propose a novel energy-aware mapping methodology of multiple DL applications onto a real embedded system that has multiple heterogeneous processors. The objective is to minimize energy consumption while satisfying the real-time constraints of all applications. In the proposed scheme, we first select Pareto-optimal mapping solutions for each application. Then mapping combination is explored, considering the scenario that indicates the dynamism of applications while satisfying the constraints. Also, we reduce energy consumption by tuning the frequency of processors. We could satisfy up to 40% higher deadline constraints and reduce the energy consumption by 22% ∼ 31% compared to the static mapping methods with real-life applications and different scenarios on a real platform. Jangryul Kim, Soonhoi Ha |
IEEE Trans. Computers | 2 |
| 2022 | Hardware-Software Codesign of a CNN AcceleratorabstractThe explosive growth of deep learning applications based on convolutional neural network (CNN) in embedded sys-tems is spurring the development of a hardware CNN accelerator, called a neural processing unit (NPU). In this work, we present how the hardware-software codesign methodology could be applied to the design of a novel adder-type NPU. After devising a baseline datapath that enables fully-pipelined execution of layers, we define a high-level behavior model based on which a high-level compiler and a virtual prototyping system are built concurrently. Since it is easy to change the microarchitecture of an NPU by modifying the simulation models of the hardware modules, we could explore the design space of NPU microarchitecture easily. In addition, we could evaluate the effect of hardware extensions to support various types of non-convolutional operations that recent CNN models use widely. After the final datapath is determined, we design the control structure and low-level compiler and implement the NPU prototype. Implementation results on an FPGA prototype show the viability of the proposed methodology and its outcome. Changjae Yi, Soonhoi Ha |
DSD | 3 |
| 2022 | Multi-Bank On-Chip Memory Management Techniques for CNN AcceleratorsabstractSince off-chip DRAM access affects both performance and power consumption significantly, convolutional neural network (CNN) accelerators commonly aim to maximize data reuse in on-chip memory. By organizing the on-chip memory to multiple banks, we may hide off-chip DRAM access delay by prefetching data to unused banks during computation. When and where to prefetch data and how to reuse the feature map data between layers define the multi-bank on-chip memory management (MOMM) problem. In this paper, we propose compiler techniques to solve the MOMM problem with two different objectives: one is to minimize the off-chip memory access volume, and the other is to minimize the processing delay caused by unhidden DRAM accesses. By running CNN benchmarks on a cycle-level NPU simulator, we demonstrate the trade-off relation between two objectives. Compared with the baseline approach that does not reuse the feature map between layers, we could reduce the DRAM access volume and the processing delay up to 55.0 and 79.4 percent, respectively. Moreover, we extend the proposed techniques to consider layer fusion that aims to reuse feature maps between layers. Experiment results confirm the superiority of the proposed hybrid fusion technique to the per-layer processing technique and the pure fusion technique. Duseok Kang, Soonhoi Ha |
IEEE Trans. Computers | 3 |
| 2022 | SNAS: Fast Hardware-Aware Neural Architecture Search MethodologyabstractRecently, automated neural architecture search (NAS) emerges as the default technique to find a state-of-the-art (SOTA) convolutional neural network (CNN) architecture with higher accuracy than manually designed architectures for image classification. In this article, we present a fast hardware-aware NAS methodology, called S3NAS, reflecting the latest research results. It consists of three steps: 1) supernet design; 2) Single-Path NAS for fast architecture exploration; and 3) scaling and post-processing. In the first step, we design a supernet, superset of candidate networks with two features: one is to allow stages to have a different number of blocks, and the other is to enable blocks to have parallel layers of different kernel sizes (MixConv). Next, we perform a differential search by extending the Single-Path NAS technique to support the MixConv layer and to add a latency-aware loss term to reduce the hyperparameter search overhead. Finally, we use compound scaling to scale up the network maximally within the latency constraint. In addition, we add squeeze-and-excitation (SE) blocks and h-swish activation functions if beneficial in the post-processing step. Experiments with the proposed methodology on four different hardware platforms demonstrate the effectiveness of the proposed methodology. It is capable of finding networks with better latency–accuracy tradeoff than SOTA networks, and the network search can be done within 4 h using TPUv3. Jaeseong Lee 0002, Jungsub Rhim, Duseok Kang, Soonhoi Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | TensorRT-Based Framework and Optimization Methodology for Deep Learning Inference on Jetson BoardsabstractAs deep learning inference applications are increasing in embedded devices, an embedded device tends to equip neural processing units (NPUs) in addition to a multi-core CPU and a GPU. NVIDIA Jetson AGX Xavier is an example. For fast and efficient development of deep learning applications, TensorRT is provided as the SDK for high-performance inference, including an optimizer and runtime that delivers low latency and high throughput for deep learning inference applications. Like most deep learning frameworks, TensorRT assumes that the inference is executed on a single processing element, GPU or NPU, not both. In this article, we present a TensorRT-based framework supporting various optimization parameters to accelerate a deep learning application targeted on an NVIDIA Jetson embedded platform with heterogeneous processors, including multi-threading, pipelining, buffer assignment, and network duplication. Since the design space of allocating layers to diverse processing elements and optimizing other parameters is huge, we devise a parameter optimization methodology that consists of a heuristic for balancing pipeline stages among heterogeneous processors and fine-tuning the process for optimizing parameters. With nine real-life benchmarks, we could achieve 101%~680% performance improvement and up to 55% energy reduction over the baseline inference using a GPU only. Eunjin Jeong, Jangryul Kim, Soonhoi Ha |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | Hierarchical Scheduling of an SDF/L Graph onto Multiple ProcessorsabstractAlthough dataflow models are known to thrive at exploiting task-level parallelism of an application, it is difficult to exploit the parallelism of data, represented well with loop structures, since these structures are not explicitly specified in existing dataflow models. SDF/L model overcomes this shortcoming by specifying the loop structures explicitly in a hierarchical fashion. We introduce a scheduling technique of an application represented by the SDF/L model onto heterogeneous processors. In the proposed method, we explore the mapping of tasks using an evolutionary meta-heuristic and schedule hierarchically in a bottom-up fashion, creating parallel loop schedules at lower levels first and then re-using them when constructing the schedule at a higher level. The efficiency of the proposed scheduling methodology is verified with benchmark examples and randomly generated SDF/L graphs. Mari-Liis Oldja, Jangryul Kim, Dowhan Jeong, Soonhoi Ha |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2021 | Fast Simulation of a Many-NPU Network-on-Chip for Microarchitectural Design Space ExplorationabstractA viable solution to cope with the ever-increasing computation complexity of deep learning applications is to integrate many neural processing units (NPUs) in a chip where a network-on-chip (NoC) is used as the communication fabric. Since the design space of an NoC is huge, the network topology is first selected based on the communication patterns of applications with a high-level performance estimation method. After the network topology is selected, the microarchitectural design space exploration is performed with a cycle-level NoC simulator. However, the existing NoC simulator is so slow that design space exploration of the microarchitecture is usually conducted manually in a narrow space. Since a synthetic trace is used, the simulation accuracy is also limited. To overcome these weak-nesses, we present a simulation technique that is fast and accurate enough for microarchitectural design space of an NoC. In the proposed technique, we use the real communication trace from the many-NPU simulation without NoC consideration. To this end, we define the trace format that defines the interface between a many-NPU simulator and the NoC simulator. To accelerate simulation speed, we propose a parallelization technique at the cluster level in the simulation of the hierarchical NoC. The key technique is to manage the timestamps of events at the cluster boundary to do without time synchronization error. And, we adjust the abstraction level of simulation models to reduce the number of modules in the SystemC NoC simulation. With the proposed technique, we could achieve up to 40 times speed-up for 32 NPU system, compared with the FlexNoC simulator. Jintaek Kang, Changjae Yi, Keonjoo Lee, Seungwook Lee, Soojung Ryu, Soonhoi Ha |
DSD | 6 |
| 2021 | Dataflow Model-based Software Synthesis Framework for Parallel and Distributed Embedded SystemsabstractExisting software development methodologies mostly assume that an application runs on a single device without concern about the non-functional requirements of an embedded system such as latency and resource consumption. Besides, embedded software is usually developed after the hardware platform is determined, since a non-negligible portion of the code depends on the hardware platform. In this article, we present a novel model-based software synthesis framework for parallel and distributed embedded systems. An application is specified as a set of tasks with the given rules for execution and communication. Having such rules enables us to perform static analysis to check some software errors at compile-time to reduce the verification difficulty. Platform-specific programs are synthesized automatically after the mapping of tasks onto processing elements is determined. The proposed framework is expandable to support new hardware platforms easily. The proposed communication code synthesis method is extensible and flexible to support various communication methods between devices. In addition, the fault-tolerant feature can be added by modifying the task graph automatically according to the selected fault-tolerance configurations by the user. The viability of the proposed software development methodology is evaluated with a real-life surveillance application that runs on six processing elements. Eunjin Jeong, Dowhan Jeong, Soonhoi Ha |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2020 | Tensor Virtualization Technique to Support Efficient Data Reorganization for CNN AcceleratorsabstractThere is a growing need for data reorganization in recent neural networks for various applications such as Generative Adversarial Networks(GANs) that use transposed convolution and U-Net that requires upsampling. We propose a novel technique, called tensor virtualization technique, to perform data reorganization efficiently with a minimal hardware addition for adder-tree based CNN accelerators. In the proposed technique, a data reorganization request is specified with a few parameters and data reorganization is performed in the virtual space without overhead in the physical memory. It allows existing adder-tree-based CNN accelerators to accelerate a wide range of neural networks that require data reorganization, including U-Net, DCGAN, and SRGAN. Soonhoi Ha |
DAC | 2 |
| 2020 | Software Development Framework for Cooperating Robots with High-level Mission SpecificationabstractIn recent years, there has been a growing interest in multiple robots performing a single task through different types of collaboration. There are two software challenges when deploying collaborative robots: how to specify a cooperative mission and how to program each robot to accomplish its mission. In this paper, we propose a novel software development framework to support distributed robot systems, swarm robots, and their hybrid. We extend the service-oriented and model-based (SeMo) framework [1] to improve the robustness, scalability, and flexibility of robot collaboration. To enable a casual user to specify various types of cooperative missions easily, the high-level mission scripting language is extended with new features such as team hierarchy, group service, one-to-many communication. The script program is refined to the robot codes through two intermediate steps, strategy description and task graph generation, in the proposed framework. The viability of the proposed framework is evidenced by two preliminary experiments using real robots and a robot simulator. Hyesun Hong, Woosuk Kang, Soonhoi Ha |
IROS | 3 |
| 2019 | Fast Performance Estimation and Design Space Exploration of Manycore-based Neural ProcessorsabstractIn the design of a neural processor, a cycle-accurate simulator is usually built to estimate the performance before hardware implementation. Since using the simulator to perform design space exploration (DSE) of hardware architecture is quite time consuming, we propose a novel method to use a high-level analytical model for fast DSE. In the model, non-deterministic execution delay is modeled with some parameters whose contribution to the performance is estimated statically by simulation. The viability of the proposed methodology is confirmed with two neural processors with different manycore architectures, achieving 2000 times speed-up within 3% accuracy error, compared with simulator-based DSE. Jintaek Kang, Dowhan Jung, Kwanghyun Chung, Soonhoi Ha |
DAC | 4 |
| 2019 | A Novel Convolutional Neural Network Accelerator That Enables Fully-Pipelined Execution of LayersabstractIn this paper, we propose a novel CNN accelerator, called MIDAP, aiming to maximize the utilization of MAC units by enabling fully-pipelined execution of layers. To this end, MIDAP adopts two-level pipelining, macro pipelining and micro pipelining, and large on-chip SRAMs. The macro pipeline consists of three modules for convolution, activation, and pooling layer and each module accesses separate memory units without access conflict. For micro-pipelining inside the convolution module, the datapath is designed to be free from dynamic resource contention. Also, the inter-layer feature map reuse is maximized by compile-time analysis. From the simulation results on 1GHz frequency, MIDAP shows the remarkable end-to-end performance of several well-known CNNs: for instance, 144 fps for Inception V3 model and 892 fps for Mobilenet V1 model with 1024 MACs. A high-level synthesis result reveals that our accelerator is able to achieve about 2.0 TOPs/W with a small area less than 2mm2with 8nm CMOS technology, thanks to its simple datapath. Jintaek Kang, Hyungdal Kwon, Hyunsik Park, Soonhoi Ha |
ICCD | 5 |
| 2019 | Optimization of Fault-Tolerant Mixed-Criticality Multi-Core Systems with Enhanced WCRT AnalysisabstractThis article proposes a novel optimization technique of fault-tolerant mixed-criticality multi-core systems with worst-case response time (WCRT) guarantees. Typically, in fault-tolerant multi-core systems, tasks can be replicated or re-executed in order to enhance the reliability. In addition, based on the policy of mixed-criticality scheduling, low-criticality tasks can be dropped at runtime. Such uncertainties caused by hardening and mixed-criticality scheduling make WCRT analysis very difficult. We show that previous analysis techniques are pessimistic as they consider avoidably extreme cases that can be safely ignored within the given reliability constraint. We improve the analysis in order to tighten the pessimism of WCRT estimates by considering the maximum number of faults to be tolerated. Further, we improve the mixed-criticality scheduling by allowing partial dropping of low-criticality tasks. On top of those, we explore the design space of hardening, task-to-core mapping, and quality-of-service of the multi-core mixed-criticality systems. The effectiveness of the proposed technique is verified by extensive experiments with synthetic and real-life benchmarks. Junchul Choi, Hoeseok Yang, Soonhoi Ha |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2018 | Architectures and algorithms for user customization of CNNsabstractIn this paper we present a convolutional neural network architecture that supports user customization through incremental transfer learning. The architecture consists of a large basic inference engine and a small augmenting engine. After training the basic inference engine and augmenting engine on a large general dataset, the basic inference engine is fixed. For user customization, only the augmenting engine is re-trained on-device using a small user specific dataset provided by the user. To accelerate the training of the augmenting engine we map this to a coarsegrained reconfigurable array processor. The complete network architecture is evaluated using the Caffe framework, and a C-code equivalent network is implemented and tested on a CGRA processor. Experiments with NIST'19 and our user-specific datasets show an increase in accuracy of the system from 76.3% to 93.2% after user customization. Mapping this code to a CGRA gives us a speed up of 45x and a 49-and 3-fold reduced energy consumption over an ARMv7 processor and a 3-way VLIW processor, respectively, showing the potential of CGRAs as DNN processors. Barend Harris, Mansureh S. Moghaddam, Duseok Kang, Inpyo Bae, Euiseok Kim, Hyemi Min, Hansu Cho, Sukjin Kim, Bernhard Egger 0002, Soonhoi Ha, Kiyoung Choi |
ASP-DAC | 10 |
| 2018 | NNsim: fast performance estimation based on sampled simulation of GPGPU kernels for neural networksabstractExistent GPU simulators are too slow to use for neural networks implemented in GPUs. For fast performance estimation, we propose a novel hybrid method of analytical performance modeling and sampled simulation of GPUs. By taking full advantage of repeated computation of neural networks, three sampling techniques are devised: Inter-Kernel sampling, Intra-Kernel sampling, and Streaming Multiprocessor sampling. The key technique is to estimate the average IPC through sampled simulation, considering the effect of the warp scheduler and memory access contention. Compared with GPGPU-Sim, the proposed technique reduces the simulation time by up to 450 times with less than 5.0% of accuracy loss. Jintaek Kang, Kwanghyun Chung, Youngmin Yi, Soonhoi Ha |
DAC | 4 |
| 2018 | End-to-end latency analysis of cause-effect chains in an engine management systemabstractAn engine management system consists of periodic or sporadic real-time tasks. A task is a set of runnables that may be fully preemptive or partially at runnable boundaries. A cause-effect chain is defined as a chain of runnables that are connected by the read/write dependency. We propose a novel analytical technique to estimate the end-to-end latency of a cause-effect chain by considering conservatively estimated schedule time bounds of associated runnables. The proposed approach is verified with an industrial-strength automotive benchmark. Junchul Choi, Soonhoi Ha |
DATE | 3 |
| 2018 | Joint optimization of speed, accuracy, and energy for embedded image recognition systemsabstractThis paper presents the image recognition system that won the first prize in the LPIRC (Low Power Image Recognition Challenge) in 2017. The goal of the challenge is to maximize the ratio between the accuracy and energy consumption within a time limit of 10 minutes for the processing of 20,000 images. Among three conflicting goals of accuracy, speed, and energy consumption, we considered the trade-off between accuracy and speed first to select Nvidia Jetson TX2 as the hardware platform and Tiny YOLO as the image recognition algorithm. Next, we applied a series of software optimization techniques to improve throughput, such as pipelining, multithreading, Tucker decomposition, and 16-bit quantization. Lastly, we explored the CPU and GPU frequencies to minimize the total energy consumption. As a result, we could achieve an accuracy of 0.24 mAP with energy consumption of 2.08Wh, which corresponds to the score of 0.11931, 2.7 times higher than the winner of LPIRC 2016. Duseok Kang, Jintaek Kang, Sungjoo Yoo, Soonhoi Ha |
DATE | 5 |
| 2018 | C-GOOD: C-code generation framework for optimized on-device deep learningabstractExecuting deep learning algorithms on mobile embedded devices is challenging because embedded devices usually have tight constraints on the computational power, memory size, and energy consumption while the resource requirements of deep learning algorithms achieving high accuracy continue to increase. Thus it is typical to use an energy-efficient accelerator such as mobile GPU, DSP array, and customized neural processor chip. Moreover, new deep learning algorithms that aim to balance accuracy, speed, and resource requirements are developed on a deep learning framework such as Caffe[16] and Tensorflow[1] that is assumed to run directly on the target hardware. However, embedded devices may not be able to run those frameworks directly due to hardware limitations or missing OS support. To overcome this difficulty, we develop a deep learning software framework that generates a C code that can be run on any devices. The framework is facilitated with various options for software optimization that can be performed according to the optimization methodology proposed in this paper. Another benefit is that it can generate various styles of C code, tailored for a specific compiler or the accelerator architecture. Experiments on three platforms, NVIDIA Jetson TX2[23], Odroid XU4[10], and SRP (Samsung Reconfigurable Processor)[32], demonstrate the potential of the proposed approach. Duseok Kang, Euiseok Kim, Inpyo Bae, Bernhard Egger 0002, Soonhoi Ha |
ICCAD | 5 |
| 2018 | A hybrid performance analysis technique for distributed real-time embedded systems
Junchul Choi, Hyunok Oh, Soonhoi Ha |
Real Time Syst. | 3 |
| 2018 | EditorialabstractThis large volume of special issue includes all regular papers presented at Embedded Systems Week (ESWEEK) 2018 that brings together three leading conferences (CASES, CODES+ISSS, and EMSOFT) in the embedded systems area. ESWEEK is a unique premier event that covers all aspects of embedded systems design and hardware/software architectures. ESWEEK presents a wide range of topics unveiling state-of-the-art techniques as can be found in this special issue. Following the journal-integrated publication model started last year, the three conferences conducted the journal-like two-stage peer-reviewed process before final decision. Acceptance rates have been about 25.5% for all conferences with a total number of 270 submissions to the journal track. Soonhoi Ha, Petru Eles |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | SeMo: Service-Oriented and Model-Based Software Framework for Cooperating RobotsabstractIn the near future, it will be common that a variety of robots are cooperating to perform a mission in various fields. A key technical challenge to realize this vision is software challenge on how to specify the mission at the user level and how to program each robot separately. In this paper, we propose a novel software development framework that separates mission specification and robot behavior programming. For mission specification, a novel scripting language is proposed with the expression capability of dynamic mode change and multitasking. For robot behavior programming, an extended dataflow model is used for task-level behavior specification that does not depend on the robot hardware platform. In the proposed framework, the actual robot software is automatically generated from the model. The viability of the proposed framework is demonstrated with two real-life experiments. Hyesun Hong, Hanwoong Jung, KangKyu Park, Soonhoi Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Incremental training of CNNs for user customization: work-in-progressabstractThis paper presents a convolutional neural network architecture that supports transfer learning for user customization. The architecture consists of a large basic inference engine and a small augmenting engine. Initially, both engines are trained using a large dataset. Only the augmenting engine is tuned to the user-specific dataset. To preserve the accuracy for the original dataset, the novel concept of quality factor is proposed. The final network is evaluated with the Caffe framework, and our own implementation on a coarse-grained reconfigurable array (CGRA) processor. Experiments with MNIST, NIST'19, and our user-specific datasets show the effectiveness of the proposed approach and the potential of CGRAs as DNN processors. Mansureh S. Moghaddam, Barend Harris, Duseok Kang, Inpyo Bae, Euiseok Kim, Hyemi Min, Hansu Cho, Sukjin Kim, Bernhard Egger 0002, Soonhoi Ha, Kiyoung Choi |
CASES | 10 |
| 2017 | A space- and energy-efficient code Compression/Decompression technique for coarse-grained reconfigurable architectures
Bernhard Egger 0002, Duseok Kang, Mansureh S. Moghaddam, Youngchul Cho, Yeonbok Lee, Sukjin Kim, Soonhoi Ha, Kiyoung Choi |
CGO | 8 |
| 2017 | Hierarchical Dataflow Modeling of Iterative ApplicationsabstractEven though dataflow models are good at exploiting task-level parallelism of an application, it is difficult to exploit the parallelism of loop structures since they are not explicitly specified in existent dataflow models. To overcome this drawback, we propose a novel extension to the SDF model, called SDF/L graph, specifying the loop structures explicitly in a hierarchical fashion. With a given SDF/L graph specification and the mapping and scheduling information, an application can be automatically parallelized on a multicore system. The enhanced expression capability by the proposed extension is verified with two applications, k-means clustering and deep neural network application. Hyesun Hong, Hyunok Oh, Soonhoi Ha |
DAC | 3 |
| 2017 | FIFA: A Kernel-Level Fault Injection Framework for ARM-Based Embedded Linux SystemabstractEmulating fault scenarios by injecting faults intentionally is commonly used to test and verify the robustness of a system. As the number of hardware devices integrated into an embedded system tends to increase consistently and the chance of hardware failure is expected to increase in an SoC, it becomes important to emulate fault scenarios caused by hardware-related errors. To this end, we present a kernel-level fault injection framework for ARM-based embedded Linux systems, called FIFA, aiming to investigate the effect of an individual hardware error in a real hardware platform rather than performing statistical analysis by random experiments. FIFA consists of two complementary fault injection techniques, one is based on the Kernel GNU Debugger and the other on hardware breakpoints. Compared with the previous work that emulates bit-flip errors only, FIFA supports other types of errors such as time delay and device failure. The viability of the proposed framework is proved by real-life experiments with an ODROID-XU4 system. Eunjin Jeong, Namgoo Lee, Jinhan Kim, Duseok Kang, Soonhoi Ha |
ICST | 5 |
| 2017 | Worst-Case Response Time Analysis of a Synchronous Dataflow Graph in a Multiprocessor System with Real-Time TasksabstractIn this article, we propose a novel technique that estimates a tight upper bound of the worst-case response time (WCRT) of a synchronous dataflow (SDF) graph when the SDF graph shares processors with other real-time tasks. When an SDF graph is executed at runtime under a self-timed or static assignment scheduling policy on a multi-processor system, static scheduling of the SDF graph does not guarantee the satisfaction of latency constraints since changes to the schedule may result in timing anomalies. To estimate the WCRT of an SDF graph with a given mapping and scheduling result, we first construct a task instance dependency graph that depicts the dependency between node executions in a static schedule. The proposed technique combines two techniques in a novel way: schedule time bound analysis and response time analysis. The former is used to consider the interference between task instances in the same SDF graph, and the latter is used to consider the interference from other real-time tasks. Through extensive experiments with synthetic examples and benchmarks, we verify the superior performance of the proposed technique compared to other existent techniques. Junchul Choi, Soonhoi Ha |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2017 | Multiprocessor Scheduling of a Multi-Mode Dataflow Graph Considering Mode Transition DelayabstractThe Synchronous Data Flow (SDF) model is widely used for specifying signal processing or streaming applications. Since modern embedded applications become more complex with dynamic behavior changes at runtime, several extensions of the SDF model have been proposed to specify the dynamic behavior changes while preserving static analyzability of the SDF model. They assume that an application has a finite number of behaviors (or modes), and each behavior (mode) is represented by an SDF graph. They are classified as multi-mode dataflow models in this article. While there exist several scheduling techniques for multi-mode dataflow models, no one allows task migration between modes. By observing that the resource requirement can be additionally reduced if task migration is allowed, we propose a multiprocessor scheduling technique of a multi-mode dataflow graph considering task migration between modes. Based on a genetic algorithm, the proposed technique schedules all SDF graphs in all modes simultaneously to minimize the resource requirement. To satisfy the throughput constraint, the proposed technique calculates the actual throughput requirement of each mode and the output buffer size for tolerating throughput jitter. We compare the proposed technique with a method that analyzes SDF graphs in each execution mode separately, a method that does not allow task migration, and a method that does not allow mode-overlapped schedule for synthetic examples and five real applications: H.264 decoder, lane detection, vocoder, MP3 decoder, and printer pipeline. Hanwoong Jung, Hyunok Oh, Soonhoi Ha |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2016 | Real-time co-scheduling of multiple dataflow graphs on multi-processor systemsabstractIt is challenging to schedule multiple dataflow applications concurrently on multi-processor embedded systems with processor sharing. As a viable solution, an approach has been proposed recently, in which the dataflow graphs are transformed into a set of independent realtime tasks. However, it may produce poor resource utilization and excessive buffer usage. Alternatively, we propose a novel two-phase scheduling technique. In the first phase, a set of static schedules is produced for each dataflow considering the resource sharing possibility; Then, we use a meta-heuristic to find the combination of per-graph schedules to minimize the resource requirement by processor sharing. We show that the proposed technique exhibits better resource and buffer efficiency. Shin-Haeng Kang, Duseok Kang, Hoeseok Yang, Soonhoi Ha |
DAC | 4 |
| 2016 | Conservative modeling of shared resource contention for dependent tasks in partitioned multi-core systems
Junchul Choi, Soonhoi Ha |
DATE | 3 |
| 2016 | TQSIM: A fast cycle-approximate processor simulator based on QEMU
Shin-Haeng Kang, Donghoon Yoo, Soonhoi Ha |
J. Syst. Archit. | 3 |
| 2016 | A Formal Approach to Power Optimization in CPSs With Delay-Workload Dependence AwarenessabstractThe design of cyber-physical systems (CPSs) faces various new challenges that are unheard of in the design of classical real-time systems. Power optimization is one of the major design goals that is witnessing such new challenges. The presence of interaction between the cyber and physical components of a CPS leads to dependence between the time delay of a computational task and the amount of workload in the next iteration. We demonstrate that it is essential to take this delay-workload dependence into consideration in order to achieve low power consumption. In this paper, we identify this new challenge, and present the first formal and comprehensive model to enable rigorous investigations on this topic. We propose a simple power management policy, and show that this policy achieves a best possible notion of optimality. In fact, we show that the optimal power consumption is attained in a “steady-state” operation and a simple policy of finding and entering this steady state suffices, which can be quite surprising considering the added complexity of this problem. Finally, we validated the efficiency of our policy with experiments. Hyung-Chan An, Hoeseok Yang, Soonhoi Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | Modeling and power optimization of cyber-physical systems with energy-workload tradeoffabstractIn this paper, we propose to take the relationship between delay and workload into account in the optimization of cyber-physical systems (CPSs). Since the components at the physical side continuously change their values or properties, a longer delay at the cyber part may result in a bigger workload for the next computation. We formulate this tradeoff and apply it to the power optimization of CPS. In doing so, we examine the schedulability of the given CPS with respect to the given parameters and initial workload. Then, we propose to keep the system operate in the stable state with minimum scaling factor and prove that it is better than any alternating sequences. We verify the validity of the proposed delay-workload model by measuring the execution delay of real-life examples. The effectiveness of the proposed power optimization policy is demonstrated with simulation results. Hoeseok Yang, Soonhoi Ha |
ISLPED | 2 |
| 2015 | Fast GPU-in-the-loop simulation technique at OpenGL ES API level for Android Graphics ApplicationsabstractBuilding a full system simulator for a CPU-GPU heterogeneous architecture recently draws keen attention of mobile device developers who want to run full software stacks without modification. A common practice is to integrate a GPU simulator with a CPU simulator, which runs very slow or is not applicable when the GPU simulator does not exist. To overcome these drawbacks, a HIL (Hardware-in-the Loop) simulation approach that integrates a real GPU with a CPU simulator has been proposed recently, in which HIL simulation is performed at the device driver level, supporting a specific GPU only. For design space exploration of CPU-GPU heterogeneous architecture, it is necessary to support various GPUs. To this end, we propose a GPU-HIL simulation technique that integrates a real GPU at the OpenGL ES API level, aiming to make a good compromise between speed and timing accuracy. Technical challenges and proposed solutions are presented in detail. Using three OpenGL ES Android benchmarks, preliminary experiments show some use cases of the proposed simulation framework for design space exploration and API-level dynamic behavior monitoring. Youngsub Ko, Youngmin Yi, Joongbaik Kim, Soonhoi Ha |
RSP | 4 |
| 2015 | Optimal Checkpoint Selection with Dual-Modular Redundancy HardeningabstractWith the continuous scaling of semiconductor technology, failure rate is increasing significantly so that reliability becomes an important issue in multiprocessor system-on-chip (MPSoC) design. We propose an optimal checkpoint selection with task duplication hardening to tolerate transient faults. A target application is specified in a task graph, and the schedule/checkpoint placements are determined at design time. The proposed optimal algorithm minimizes the checkpoint overhead with a latency constraint. Experimental results show that the proposed algorithm effectively reduces the minimum end-to-end latency to perform a fault-tolerant schedule. In addition, the proposed algorithm dramatically decreases the checkpointing overhead on uniprocessor and multiprocessor systems compared with a greedy approach and an equidistant algorithm. Shin-Haeng Kang, Hae-woo Park, Sungchan Kim, Hyunok Oh, Soonhoi Ha |
IEEE Trans. Computers | 5 |
| 2014 | Static Mapping of Mixed-Critical Applications for Fault-Tolerant MPSoCsabstractThis paper presents a static mapping optimization technique for fault-tolerant mixed-criticality MPSoCs. The uncertainties imposed by system hardening and mixed criticality algorithms, such as dynamic task dropping, make the worst-case response time analysis difficult for such systems. We tackle this challenge and propose a worst-case analysis framework that considers both reliability and mixed-criticality concerns. On top of that, we build up a design space exploration engine that optimizes fault-tolerant mixed-criticality MPSoCs and provides worst-case guarantees. We study the mapping optimization considering judicious task dropping, that may impose a certain service degradation. Extensive experiments with real-life and synthetic benchmarks confirm the effectiveness of the proposed technique. Shin-Haeng Kang, Hoeseok Yang, Sungchan Kim, Iuliana Bacivarov, Soonhoi Ha, Lothar Thiele |
DAC | 5 |
| 2014 | Hardware-in-the-loop Simulation for CPU/GPU Heterogeneous PlatformsabstractMulti-core CPU/GPU heterogeneous platforms became popular in embedded systems. A full system simulator is typically used to observe the internal system behavior by running complete software stacks without modification on simulation models of CPUs and other devices in the system. However, there are few known full system simulators for CPU/GPU heterogeneous platforms and existent GPU simulators are prohibitively slow for running application software. In this paper, we propose a hardware-in-the-loop simulation technique that integrates GPU hardware into a full system simulator. A novel interfacing mechanism between CPU simulator and the development board, where GPU hardware is integrated, is devised. In the experiments, we took Exynos 4412 as a case study, where gem5 simulator is used to simulate mainly a quad-core ARM CPU in the platform and an Exynos development board is used to run the Mali GPU hardware. We could successfully run Android apps on the proposed hardware-in-the-loop simulation framework with up to 1.5 M cycles per second performance. Youngsub Ko, Youngmin Yi, Myungsun Kim, Soonhoi Ha |
DAC | 5 |
| 2014 | Reliability-aware mapping optimization of multi-core systems with mixed-criticalityabstractThis paper presents a novel mapping optimization technique for mixed critical multi-core systems with different reliability requirements. For this scope, we derived a quantitative reliability metric and presented a scheduling analysis that certifies given mixed-criticality constraints. Our framework is capable of investigating re-execution, passive replication, and modular redundancy with optimized voter placement, while typical hardening approaches consider only one or two of these techniques. The proposed technique complies with existing safety standards and is power-efficient, as demonstrated by our experiments. Shin-Haeng Kang, Hoeseok Yang, Sungchan Kim, Iuliana Bacivarov, Soonhoi Ha, Lothar Thiele |
DATE | 5 |
| 2014 | Dynamic Behavior Specification and Dynamic Mapping for Real-Time Embedded Systems: HOPES ApproachabstractAs the number of processors in a chip increases and more functions are integrated, the system status will change dynamically due to various factors such as the workload variation, QoS requirement, and unexpected component failure. A typical method to deal with the dynamics of the system is to decide the mapping decision at runtime, based on the local information of the system status. It is very challenging to guarantee any real-time performance of a certain application in such a dynamically varying system. To solve this problem, we propose a hybrid specification of dataflow and FSM models to specify the dynamic behavior of a system distinguishing inter- and intra-application dynamism. At the top level, each application is specified by a dataflow task and the dynamic behavior is modeled as a control task that supervises the execution of applications. Inside a dataflow task, we specify the dynamic behavior using a similar way as FSM-based SADF in which an application is specified by a synchronous dataflow graph for each mode of operation. It enables us to perform compile-time scheduling of each graph to maximize the throughput varying the number of allocated processors, and store the scheduling information. When a change in system state is detected at runtime, the number of allocated processors to the active tasks is determined dynamically utilizing the stored scheduling information of those tasks in order to meet the real-time requirements. The proposed technique is implemented in the HOPES design environment. Through preliminary experiments with a simple smartphone example, we show the viability of the proposed methodology. Hanwoong Jung, Chanhee Lee 0002, Shin-Haeng Kang, Sungchan Kim, Hyunok Oh, Soonhoi Ha |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2013 | A novel analytical method for worst case response time estimation of distributed embedded systemsabstractIn this paper, we propose a novel analytical method, called scheduling time bound analysis, to find a tight upper bound of the worst-case response time in a distributed real-time embedded system, considering execution time variations of tasks, jitter of input arrivals, and scheduling anomaly behavior in a multi-tasking system all together. By analyzing the graph topology and worst-case scheduling scenarios, we measure the conservative scheduling time bound of each task. The proposed method supports an arbitrary mixture of preemptive and non-preemptive processing elements. Its speed is comparable to compositional approaches while it gives a much tighter bound. The advantages of the proposed approach compared with related work were verified by experimental results with randomly generated task graphs and a real-life automotive application. Hyunok Oh, Junchul Choi, Hyojin Ha, Soonhoi Ha |
DAC | 5 |
| 2012 | Relaxed synchronization technique for speeding-up the parallel simulation of multiprocessor systemsabstractFor design verification of an MPSoC, a virtual prototyping system has been widely used as a cheap and fast method without a hardware prototype. It usually consists of component simulators working together in a single simulation host. As the number of component simulators increases, the simulation performance degrades significantly due to occurrence of frequent inter-simulator communication. In this paper, to boost up the simulation speed further, we propose a novel technique, called relaxed synchronization, which uses a simulation cache at each component simulator for simulation purpose. Like an architectural cache that reduces the main memory access frequency, a simulation cache reduces the count of synchronous communication effectively between the corresponding component simulator and the simulation backplane. When a read or write request to a shared memory is made, a cache line, not a single element, is transferred to utilize the space and temporal locality for simulation. The proposed technique is based on an assumption that the application program uses a relaxed memory model. Through experiments with real-life applications, it is proved that the proposed approach improves the simulation performance by up to 330 %. Dukyoung Yun, Sungchan Kim, Soonhoi Ha |
ASP-DAC | 3 |
| 2012 | Executing synchronous dataflow graphs on a SPM-based multicore architectureabstractIn this paper we are concerned about executing synchronous dataflow (SDF) applications on a multicore architecture where a core has a limited size of scratchpad memory (SPM). Unlike traditional multi-processor scheduling of SDF graphs, we consider the SPM size limitation that incurs code and data overlay overhead. Since the scheduling problem is intractable, we propose an EA(evolutionary algorithm)-based technique. To hide memory latency, prefetching is aggressively performed in the proposed technique. The experimental results show that our approach reduces the overlay overhead significantly compared to a non-optimized approach and the previous approach. Junchul Choi, Hyunok Oh, Sungchan Kim, Soonhoi Ha |
DAC | 4 |
| 2012 | A cycle-level parallel simulation technique exploiting both space and time parallelismabstractAs the number of processors increases in an MPSoC, the simulation performance degrades significantly if all component simulators run sequentially. Recently a novel parallel simulation technique was proposed to exploit space-parallelism by distributing component simulators to multiple host cores. In this paper, we boost the performance further by exploiting time-parallelism in case that an application is specified as a task graph following the data-flow semantics, such as a KPN (Kahn Process Network) or a data flow graph. Time-parallel simulation enables parallel execution of tasks in different intervals in the timeline by resolving data dependencies between them with redundant host code execution. The proposed technique provides higher degree of parallelism beyond the number of processors in the target architecture. Experiments with real-life multimedia examples prove the effectiveness of the proposed approach. Dukyoung Yun, Youngmin Yi, Sungchan Kim, Soonhoi Ha |
RSP | 4 |
| 2012 | An ILP-based Worst-case Performance Analysis Technique for Distributed Real-time Embedded SystemsabstractFinding a tight upper bound of the worst-case response time in a distributed real-time embedded system is a very challenging problem since we have to consider execution time variations of tasks, jitter of input arrivals, scheduling anomaly behavior in a multi-tasking system, all together. In this paper, we translate the problem as an optimization problem and propose a novel solution based on ILP (Integer Linear Programming). In the proposed technique, we formulate a set of ILP formulas in a compositional way for modeling flexibility, but solve the problem holistically to achieve tighter upper bounds. To mitigate the time complexity of the ILP method, we perform static analysis based on a scheduling heuristic to reduce the number of variables and confine the variable ranges. Preliminary experiments with the benchmarks used in the related work and a real-life example show promising results that give tight bounds in an affordable solution time. Hyunok Oh, Hyojin Ha, Shin-Haeng Kang, Junchul Choi, Soonhoi Ha |
RTSS | 6 |
| 2012 | A parallel and distributed meta-heuristic framework based on partially ordered knowledge sharing
Minyoung Kim 0002, Mark-Oliver Stehr, Hyunok Oh, Soonhoi Ha |
J. Parallel Distributed Comput. | 5 |
| 2012 | Prolog to the Section on Hardware/Software Codesign
Soonhoi Ha |
Proc. IEEE | 1 |
| 2012 | A Parallel Simulation Technique for Multicore Embedded Systems and Its Performance AnalysisabstractA virtual prototyping system is constructed by replacing real processing components with component simulators running concurrently. The performance of such a distributed simulation decreases drastically as the number of component simulators increases. Thus, we propose a novel parallel simulation technique to boost up the simulation speed. In the proposed technique, a simulator wrapper performs time synchronization with the simulation backplane on behalf of the associated component simulator itself. Component simulators send null messages periodically to the backplane to enable parallel simulation without any causality problems. Since excessive communication may degrade the simulation performance, we also propose a novel performance analysis technique to determine an optimal period of null message transfer, considering both the characteristics of a target application and the configurations of the simulation host. Through intensive experiments, we show that the proposed parallel simulation achieves almost linear speedup to the number of processor cores if the frequency of null message transfer is optimally decided. The proposed analysis technique could predict the simulation performance with more than 90% accuracy in the worst case for various target applications and simulation environments we have used for experiments. Dukyoung Yun, Sungchan Kim, Soonhoi Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Minimizing buffer requirements for throughput constrained parallel execution of synchronous dataflow graphabstractThis paper concerns throughput-constrained parallel execution of synchronous data flow graphs. This paper assumes static mapping and dynamic scheduling of nodes, which has several benefits over static scheduling approaches. We determine the buffer size of all arcs to minimize the total buffer size while satisfying a throughput constraint. Dynamic scheduling is able to achieve the similar throughput performance as the static scheduling does by unfolding the given SDF graph. A key issue of dynamic scheduling is how to assign the priority to each node invocation, which is also discussed in this paper. Since the problem is NP-hard, we present a heuristic based on a genetic algorithm. The experimental results confirm the viability of the proposed technique. Tae-ho Shin, Hyunok Oh, Soonhoi Ha |
ASP-DAC | 3 |
| 2011 | Simulation environment configuration for parallel simulation of multicore embedded systemsabstractIncreasing complexity of multicore embedded systems makes careful construction of virtual prototyping system crucial to shorten design turnaround time due to the growing demand of simulation time. Parallel simulation aims to accelerate the simulation speed by running component simulators concurrently. But extra overhead of communication and synchronization between simulators may overshadow the benefits of parallel simulation. In this paper we propose a technique to configure the simulation environment optimally considering the application characteristics. Particularly, we focus on three design axes, simulation platform selection, mapping of component simulators to participating host processors and period of null message transfer for time synchronization. As a result, the proposed technique enables the efficient exploitation of parallelism by 1) well-balanced distribution of simulation workloads to host processors and 2) the minimized overhead for null message transfer, in turn, leading to the maximal simulation performance. The experimental results show that the proposed technique robustly found the optimal configurations for wide variance of application characteristics and simulation platform. Dukyoung Yun, Sungchan Kim, Soonhoi Ha |
DAC | 4 |
| 2011 | Fast Communication Architecture Exploration of Processor Pool-Based MPSoC via Static Performance AnalysisabstractMultiprocessor systems-on-chip (MPSoCs) are evolving toward processor pool-based architecture that employs a hierarchical on-chip network for inter-processor and intra-processor pool communication. This letter presents a systematic exploration method of the cascaded bus matrix-based on-chip network design for processor pool-based MPSoCs. It uses an evolutionary algorithm to find optimal architectures in terms of on-chip area while satisfying a given performance constraint. Since simulation is too time-consuming to evaluate the performance of complex on-chip networks during architecture exploration, we propose to prune the design space efficiently using two novel static analysis techniques: 1) bandwidth analysis considering task execution dependences, and 2) memory contention analysis for accurate performance estimation. Thanks to fast and accurate evaluation by the proposed analysis techniques, we achieved an order of magnitude speed improvement for the architecture exploration without performance loss, compared with a simulation-based approach. Youngpyo Joo, Sungchan Kim, Soonhoi Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Library Support in an Actor-Based Parallel Programming PlatformabstractActor model-based design is actively researched for parallel embedded SW design since the model exposes the potential parallelism explicitly in an architecture-neutral form. In most actor-oriented models, actors are self-contained and data channels are the only sharable object between actors, and they compose a system in a flat layer. In contrast, it is common to use shared library functions and construct vertically layered software for efficiency and modularity. To fill this gap between modeling and implementation, we propose a special actor, library task, with new types of ports: library master port and library slave port. It is a sharable and mappable object that defines a set of function interfaces inside. N:1 master-slave connection allows sharing a library task and the master-slave connection can specify vertically layered software and client-server applications naturally. To support the library task in our embedded software design environment, we develop an automatic mapping algorithm as well as an automatic code generator. The design environment with the library task is applied for two target platforms: IBM CELL Broad band Engine and an ARM-based multicore simulator. Preliminary experiments show that the special actor, or library task, extends the expression power of the previous actor model with efficiently generated codes. Hae-woo Park, Hanwoong Jung, Hyunok Oh, Soonhoi Ha |
IEEE Trans. Ind. Informatics | 4 |
| 2010 | An Application Framework for Loosely Coupled Networked Cyber-Physical SystemsabstractNetworked Cyber-Physical Systems (NCPSs) present many challenges since they require a tight combination with the physical world as well as a balance between autonomous operation and coordination among heterogeneous nodes. These fundamental challenges range from how NCPSs are architected, implemented, composed, and programmed to how they can be validated. In this paper, we describe a new paradigm for programming an NCPS that enables users to specify their needs and nodes to contribute capabilities and resources. This new paradigm is based on the partially ordered knowledge-sharing model that makes explicit the abstract structure of a computation in space and time. Based on this model, we propose an application framework that provides a uniform abstraction for a wide range of NCPS applications, especially those concerned with distributed sensing, optimization, and control. The proposed framework provides a generic service to represent, manipulate, and share knowledge across the network under minimal assumptions on connectivity. Our framework is tested on a new distributed version of an evolutionary optimization algorithm that runs on a computing cluster and is also used to solve a dynamic distributed optimization problem in a simulated NCPS that uses mobile robots as controllable data mules. Minyoung Kim 0002, Mark-Oliver Stehr, Soonhoi Ha |
EUC | 4 |
| 2010 | An MILP-Based Performance Analysis Technique for Non-Preemptive Multitasking MPSoCabstractFor real-time applications, it is necessary to estimate the worst-case performance early in the design process without actual hardware implementation. While the non-preemptive task scheduling is pertinent to multi-core platforms because of easy implementation and high performance, its scheduling anomaly behavior makes the worst-case performance estimation extremely difficult. In this paper, we propose an analysis technique based on mixed integer linear programming (MILP) to estimate the worst-case performance of each task in a non-preemptive multitask application on multi-processor system-on-chip architecture. MILP provides a systematic way to describe the complex interaction among task scheduling, communication architecture, and task execution, which affects the worst-case behavior dynamically. The proposed analysis technique overcomes several limitations that previous work usually has; it allows multiple tasks with different periods and models contention on the communication architecture. We show that the proposed analysis takes affordable computation time to make it of practical value even though it has exponential complexity in theory. The proposed technique estimates a safe bound on task latency statistically, which is demonstrated by extensive random simulations. Hoeseok Yang, Sungchan Kim, Soonhoi Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | Serialized parallel code generation framework for MPSoCabstractThe models of computations that express concurrency naturally are preferred for initial specification of MPSoC system, since popular programming languages such as C and C++ are designed for sequential execution. In our previous work, we proposed a design framework where two models are used for the initial specification of the system behavior; task model at the top level and dataflow model inside each task. After the partition and mapping process is performed with each architecture candidate, the target code is automatically generated for both Design-Space Exploration (DSE) and final implementation. In this article, we focus on parallel code generation for MPSoC, proposing two main techniques. The first is to express functional and data parallelism differently following the partition and mapping decision. In the proposed technique, the generated code consists of multiple tasks running concurrently, which achieves functional parallelism. On the other hand, we use OpenMP directives to express data parallelism inside a task. Second is to adopt the code serialization technique to execute a multitasking application without OS scheduler, aiming to generate the highly portable code on various platforms for an efficient DSE process. We extend the previous code serialization techniques to multiprocessor systems and utilize the formal properties of the dataflow model for efficient code generation. The experiments including H.263 codec example show the viability of the proposed technique and the efficiency of the generated code. Seongnam Kwon, Soonhoi Ha |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2009 | On-chip communication architecture exploration for processor-pool-based MPSoCabstractMPSoC is evolving towards processor-pool (PP)-based architectures, which employ hierarchical on-chip network for inter- and intra-PP communication. Since the design space of PP-based MPSoC is extremely wide, application-specific optimization of on-chip communication is a nontrivial task. This paper presents a systematic methodology for on-chip network design of PP-based MPSoC. The proposed approach allows independent configurations of PPs, which leads to efficient solutions than previous work. Since time-consuming simulation is inevitable to evaluate complicated on-chip network during exploration, we do early pruning of design space by a bandwidth analysis technique that considers task execution dependencies. Our approach yields the Pareto-optimal solutions between clock frequency and area requirements. The experiments show that the proposed technique finds more efficient architectures compared with the previous approaches. Youngpyo Joo, Sungchan Kim, Soonhoi Ha |
DATE | 3 |
| 2009 | Programming MPSoC platforms: Road works ahead!
Rainer Leupers, András Vajda, Marco Bekooij, Soonhoi Ha, Rainer Dömer, Achim Nohl |
DATE | 4 |
| 2009 | Pipelined data parallel task mapping/scheduling technique for MPSoCabstractIn this paper, we propose a multi-task mapping/scheduling technique for heterogeneous and scalable MPSoC. To utilize the large number of cores embedded in MPSoC, the proposed technique considers temporal and data parallelisms as well as task parallelism. We define a multi-task mapping/scheduling problem with all these parallelisms and propose a QEA (quantum-inspired evolutionary algorithm)-based heuristic. Compared with an ILP (Integer Linear Programming) approach, experiments with real-life examples show the feasibility and the efficiency of the proposed technique. Hoeseok Yang, Soonhoi Ha |
DATE | 2 |
| 2009 | Introduction
Felix Wolf 0001, Andy D. Pimentel, Luiz De Rose, Soonhoi Ha, Thilo Kielmann, Anna Sikora |
Euro-Par | 4 |
| 2008 | Architecture Exploration of NAND Flash-based Multimedia CardabstractIn this paper, we present an architecture exploration methodology for low-end embedded systems where the reduction of cost is a primary design concern. The architecture exploration of such systems needs to explore a wide design space spanned by detailed architecture parameters through cycle-accurate performance estimation. For fast exploration, the proposed methodology is based on an efficient evolutionary algorithm, called QEA, and trace-driven simulation to evaluate architecture candidates quickly. We applied the proposed methodology to NAND flash-based Multimedia Card as a case study considering the following design parameters: buffer size, flash memory configuration, clock, communication architecture, and memory allocation. The experimental results validate the proposed methodology by showing the optimal architecture configurations with varying performance constraints and design parameters. Sungchan Kim, Chanik Park, Soonhoi Ha |
DATE | 3 |
| 2008 | Overcoming performance bottlenecks in using OpenMP on SMP clusters
Woo-Chul Jeun, Yang-Suk Kee, Soonhoi Ha, Changdon Kee |
Parallel Comput. | 3 |
| 2008 | Introduction to embedded systems week 2006 special issueabstractintroduction Share on Introduction to embedded systems week 2006 special issue Editors: Soonhoi Ha Seoul National University Seoul National UniversityView Profile , Kiyoung Choi Seoul National University Seoul National UniversityView Profile , Taewhan Kim Seoul National University Seoul National UniversityView Profile , Krisztian Flautner ARM Ltd. U.K. ARM Ltd. U.K.View Profile , Sanglyul Min Seoul National University Seoul National UniversityView Profile , Wang Yi Uppsala University Uppsala UniversityView Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 7Issue 2Article No.: 8pp 1–3https://doi.org/10.1145/1331331.1331332Published:29 January 2008Publication History 0citation445DownloadsMetricsTotal Citations0Total Downloads445Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Soonhoi Ha, Kiyoung Choi, Krisztián Flautner, Sang Lyul Min, Wang Yi 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2008 | A retargetable parallel-programming framework for MPSoCabstractAs more processing elements are integrated in a single chip, embedded software design becomes more challenging: It becomes a parallel programming for nontrivial heterogeneous multiprocessors with diverse communication architectures, and design constraints such as hardware cost, power, and timeliness. In the current practice of parallel programming with MPI or OpenMP, the programmer should manually optimize the parallel code for each target architecture and for the design constraints. Thus, the design-space exploration of MPSoC (multiprocessor systems-on-chip) costs become prohibitively large as software development overhead increases drastically. To solve this problem, we develop a parallel-programming framework based on a novel programming model called common intermediate code (CIC). In a CIC, functional parallelism and data parallelism of application tasks are specified independently of the target architecture and design constraints. Then, the CIC translator translates the CIC into the final parallel code, considering the target architecture and design constraints to make the CIC retargetable. Experiments with preliminary examples, including the H.263 decoder, show that the proposed parallel-programming framework increases the design productivity of MPSoC software significantly. Seongnam Kwon, Woo-Chul Jeun, Soonhoi Ha, Yunheung Paek |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2007 | Model-based Programming Environment of Embedded Software for MPSoCabstractA noble model-based programming environment of embedded software for MPSoC is proposed. By defining a common intermediate code (CIC), it separates modeling of the software and implementation optimized for target architecture. It also allows us to use diverse models for initial specification. Another feature is to provide multi-phase debugging capabilities: at the modeling stage, at the code generation stage, and at the simulation stage. Preliminary experiments with a Divx player confirm the feasibility and validity of the proposed technique. Soonhoi Ha |
ASP-DAC | 1 |
| 2007 | Effective OpenMP Implementation and Translation For Multiprocessor System-On-Chip without Using OSabstractIt is attractive to use the OpenMP as a parallel programming model on a multiprocessor system-on-chip (MPSoC) because it is easy to write a parallel program in the OpenMP and there is no standard method for parallel programming on an MPSoC. In this paper, we propose an effective OpenMP implementation and translation for major OpenMP directives on an MPSoC with physically shared memories, hardware semaphores, and no operating system. Woo-Chul Jeun, Soonhoi Ha |
ASP-DAC | 2 |
| 2007 | Performance evaluation and optimization of dual-port SDRAM architecture for mobile embedded systemsabstractRecently dual-port SDRAM (DPSDRAM) architecture tailored for dual-processor based mobile embedded systems has been announced where a single memory chip plays the role of the local memories and the shared memory for both processors. In order to keep memory consistency from simultaneous accesses of both ports, every access to the shared memory should be protected by a synchronization mechanism, which can result in substantial access latency. We propose two optimization techniques by exploiting the communication patterns of target application: lock-priority scheme and static-copy scheme. Further, by dividing the shared bank into multiple blocks, we enable simultaneous accesses to different blocks and achieve considerable performance gain. Experiments on a virtual prototyping system show a promising result that we achieve about 20-50% performance gain compared to the base DPSDRAM architecture. Hoeseok Yang, Sungchan Kim, Hae-woo Park, Soonhoi Ha |
CASES | 5 |
| 2007 | CATS: cycle accurate transaction-driven simulation with multiple processor simulatorsabstractThis paper focuses on enhancing performance of cycle accurate simulation with multiple processor simulators. Simulation performance is determined by how often simulators exchange events with one another and how accurately simulators model their behavior. Previous techniques have limited their applicability or sacrificed accuracy for performance. In this paper, we notice that inaccuracy comes from events which arrive between event exchange boundaries. To solve the problem, we propose cycle accurate transaction-driven simulation which maintains event exchange boundaries at bus transactions but compensates for accuracy. The proposed technique is implemented in a publicly available CATS framework and our experiment with 64 processors achieves 1.2M processor cycles/s (200K instructions/s) which is faster than other cycle accurate frameworks by an order of magnitude Dohyung Kim 0007, Soonhoi Ha, Rajesh K. Gupta 0001 |
DATE | 2 |
| 2007 | A novel technique to use scratch-pad memory for stack management
Hae-woo Park, Soonhoi Ha |
DATE | 3 |
| 2007 | Fast and Accurate Cosimulation of MPSoC Using Trace-Driven Virtual SynchronizationabstractAs MPSoC has become an effective solution to ever-increasing design complexity of modern embedded systems, fast and accurate cosimulation of such systems is becoming a tough challenge. Cosimulation performance is in inverse proportion to the number of processor simulators in conventional cosimulation frameworks with lock-step synchronization schemes. To overcome this problem, we propose a novel time synchronization technique called trace-driven virtual synchronization. Having separate phases of event generation and event alignment in the cosimulation, time synchronization overhead is reduced to almost zero, boosting cosimulation speed while accuracy is almost preserved. In addition, this technique enables (1) a fast mixed level cosimulation where different abstraction level simulators are easily integrated communicating with traces and (2) a distributed parallel cosimulation where each simulator can run at its full speed without synchronizing with other simulator too frequently. We compared the performance and the accuracy with MaxSim, a well-known commercial System C simulation framework, and the proposed framework showed 11 times faster performance for H.263 decoder example, while the error was below 5%. Youngmin Yi, Dohyung Kim 0007, Soonhoi Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | PeaCE: A hardware-software codesign environment for multimedia embedded systemsabstractExistent hardware-software (HW-SW) codesign tools mainly focus on HW-SW cosimulation to build a virtual prototyping environment that enables software design and system verification without need of making a hardware prototype. Not only HW-SW cosimulation, but also HW-SW codesign methodology involves system specification, functional simulation, design-space exploration, and hardware-software cosynthesis. The PeaCE codesign environment is the first full-fledged HW-SW codesign environment that provides seamless codesign flow from functional simulation to system synthesis. Targeting for multimedia applications with real-time constraints, PeaCE specifies the system behavior with a heterogeneous composition of three models of computation and utilizes features of the formal models maximally during the whole design process. It is also a reconfigurable framework in the sense that third-party design tools can be integrated to build a customized tool chain. Experiments with industry-strength examples prove the viability of the proposed technique. Soonhoi Ha, Sungchan Kim, Choonseung Lee, Youngmin Yi, Seongnam Kwon, Youngpyo Joo |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2006 | Conversion of reference C code to dataflow model: H.264 encoder case studyabstractModel-based design is widely accepted in developing complex embedded system under intense time-to-market pressure. While it promises improved design productivity, the main bottleneck lies not in the design methodology but in constructing the initial algorithm representation in the specified model. It is particularly true if a complicated multimedia application is given in the form of a sequential reference C code. In this paper we propose a systematic procedure for converting a sequential C code to a dataflow specification that has been widely used in many design environments for DSP systems. The proposed technique is successfully applied to H.264 encoder algorithm as a case study. Hyeyoung Hwang, Taewook Oh, Hyunuk Jung, Soonhoi Ha |
ASP-DAC | 4 |
| 2006 | Memory optimal single appearance schedule with dynamic loop count for synchronous dataflow graphsabstractIn this paper, we propose a new single appearance schedule for synchronous dataflow programs to minimize data memory and code memory size simultaneously. While a single appearance schedule promises only one appearance of each node definition in the generated code, it requires significant amount of data memory overhead compared with a buffer optimal schedule allowing multiple appearance. The key idea of the proposed technique is to make a dynamic decision of loop count to make a schedule quasi-static. The proposed quasi-static schedule produces a single appearance schedule code with minimum data memory requirement. We prove that every buffer optimal schedule can be transformed to our single appearance schedule which requires optimal buffer size for arbitrary synchronous dataflow graphs. The only penalty for the proposed technique is slight performance overhead of computing loop counts dynamically. In order to minimize the overhead we propose optimization techniques. Experimental results show that the proposed algorithm reduces 20% total memory with less than 1% performance overhead compared with the previous single appearance schedule algorithms. Hyunok Oh, Nikil Dutt, Soonhoi Ha |
ASP-DAC | 3 |
| 2006 | Parallel co-simulation using virtual synchronization with redundant host executionabstractIn traditional parallel co-simulation approaches, the simulation speed is heavily limited by time synchronization overhead between simulators and idle time caused by data dependency. Recent work has shown that the time synchronization overhead can be reduced significantly by predicting the next synchronization points more effectively or by separating trace-driven architecture simulation from trace generation from component simulators. The latter is known as virtual synchronization technique. In this paper, we propose redundant host execution to minimize the simulation idle time caused by data dependency in simulation models. By combining virtual synchronization and redundant host execution techniques we could make parallel execution of multiple simulators a viable solution for fast but cycle-accurate co-simulation. Experiments show about 40% performance gain over a technique which uses virtual synchronization only Dohyung Kim 0007, Soonhoi Ha, Rajesh K. Gupta 0001 |
DATE | 2 |
| 2006 | Dynamic code overlay of SDF-modeled programs on low-end embedded systemsabstractIn this paper we propose a dynamic code overlay technique of synchronous data-flow (SDF)-modeled program for low-end embedded systems which lack MMU-support. With this technique, the system can utilize expensive SRAM memory more efficiently by using flash memory as code storage. SRAM is divided into several regions called overlay slots. A data-flow block or a cluster of data-flow blocks is loaded into the corresponding overlay slot on demand at run-time. Which blocks are clustered together and which overlay slots are allocated to the clusters are statically decided by the clustering and placement algorithm. We also propose an automatic code generation framework that generates the C-program code, dynamic loader and linker script files from the given SDF-modeled blocks and schematic, so we can run or simulate the program immediately without any additional coding effort. Experiments report that we can reduce the SRAM size significantly with a reasonable amount of time overhead for several real applications Hae-woo Park, Kyoungjoo Oh, Myoung-min Sim, Soonhoi Ha |
DATE | 5 |
| 2006 | Hardware-Software Codesign of Multimedia Embedded Systems: the PeaCEabstractHardware/software codesign involves various design problems including system specification, design space exploration, hardware/software co-verification, and system synthesis. A codesign environment is a software tool that facilitates capabilities to solve these design problems. This paper presents the peace codesign environment mainly targeting for multimedia applications with real-time constraints. Peace specifies the system behavior with a heterogeneous composition of three models of computation. The Peace environment provides seamless co-design flow from functional simulation to system synthesis, utilizing the features of the formal models maximally during the whole design process. Preliminary experiments with real examples prove the viability of the proposed technique Soonhoi Ha, Choonseung Lee, Youngmin Yi, Seongnam Kwon, Youngpyo Joo |
RTCSA | 1 |
| 2006 | Efficient exploration of bus-based system-on-chip architecturesabstractSeparation between computation and communication in system design allows system designers to explore the communication architecture independently after component selection and mapping decision is made. In this paper, we present an iterative two-step exploration methodology for bus-based on-chip communication architecture for multitask applications. We assume that the memory traces from the processing components are given. The proposed methodology uses a static performance estimation technique extended for multitask applications to reduce the design space quickly and drastically and applies a trace-driven simulation to the reduced set of design candidates for accurate performance estimation. For the case that local memory traffics as well as shared memory traffics are involved in bus contention, memory allocation is considered as an important axis of the design space in our technique. Experimental results show that the proposed methodology achieves significant performance gain by optimizing on-chip communication only, up to almost 100% compared with an initial single shared bus architecture, in both two real-life examples, a four-Channel digital video recorder and an equalizer for OFDM DVB-T receiver. Sungchan Kim, Soonhoi Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Static analysis and automatic code synthesis of flexible FSM modelabstractTo describe complex control modules, the following four features are requested for extended FSM models: concurrency, compositionality, static analyzability, and automatic code synthesis capability. In our codesign environment we use a new FSM extension called flexible FSM model. It extends the expression capabilities by concurrency, hierarchy, and state variable while it maintains formal property. Because of formality and the structured nature of fFSM model, we can apply a static analysis method to find ambiguous behavior and synthesize software/hardware automatically, which is the main focus of this paper. We expect that the proposed technique can be applied to other compositional FSM extensions. Dohyung Kim 0007, Soonhoi Ha |
ASP-DAC | 2 |
| 2005 | Embedded software generation from system level specification for multi-tasking embedded systemsabstractAbstract- In this paper we present a new design flow in which embedded software code is generated from system level specification of multi-tasking embedded system, both for simulation and implementation. The generated software has a layered structure using virtual OS APIs and OS wrapper implementations to make it reconfigurable for multiple target platforms. Implementation of the OS wrapper is explained in details. With a Divx play example, we show some experimental results about the real-time performance comparison between two different platforms I. KiSeun Kwon, Youngmin Yi, Dohyung Kim 0007, Soonhoi Ha |
ASP-DAC | 4 |
| 2005 | Single appearance schedule with dynamic loop count for minimum data buffer from synchronous dataflow graphsabstractIn this paper, we propose a new single appearance schedule for synchronous dataflow programs to minimize data memory and code memory size at the same time. When the software code is automatically synthesized from the dataflow program graphs, a single appearance schedule promises only one appearance of each node definition in the generated code. While several heuristics have been developed to find a single appearance schedule, they all have to pay significant amount of data memory overhead compared with a buffer optimal schedule. The key idea of the proposed technique is to make a dynamic decision of loop count to make a schedule quasi-static. The proposed quasi-static static schedule produces a single appearance schedule code with minimum data memory requirement. We prove that the proposed scheduling technique is optimal for a chain-structured graph in terms of data memory requirement while maintaining the single appearance schedule. The only penalty for the proposed technique is slight performance overhead of computing loop counts dynamically. Experimental results show that the proposed algorithm reduces 20% total memory with less than 1% performance overhead compared with the previous single appearance schedule algorithms for CD2DAT and non uniform filter bank applications. Hyunok Oh, Nikil Dutt, Soonhoi Ha |
CASES | 3 |
| 2005 | Trace-driven HW/SW cosimulation using virtual synchronization techniqueabstractPoor performance of HW/SW cosimulation is mainly caused by synchronization requirement between component simulators. Virtual synchronization technique was proposed to remove the need of synchronization in cycle accurate cosimulation. But the previous execution-driven simulation based on virtual synchronization has limitations in the application area. In this paper, we propose a novel trace-driven HW/SW cosimulation using virtual synchronization technique. Through OS modeling and channel modeling, the proposed cosimulation technique could be applied more widely while improving the simulation performance further. Experiments with a DIVX player example prove the viability of the proposed technique. Dohyung Kim 0007, Youngmin Yi, Soonhoi Ha |
DAC | 3 |
| 2005 | Schedule-aware performance estimation of communication architecture for efficient design space explorationabstractIn this paper, we are concerned about performance estimation of bus-based communication architectures assuming that task partitioning and scheduling on processing elements are already determined. Since communication overhead is dynamic and unpredictable due to bus contention, a simulation-based approach seems inevitable for accurate performance estimation. However, it is too time-consuming to be used for exploring the wide design space of bus architectures. We propose a static performance-estimation technique based on a queueing analysis assuming that the memory traces and the task schedule information are given. We use this static estimation technique as the first step in our design space exploration framework to prune the design space drastically before applying a simulation-based approach to the reduced design space. Experimental results show that the proposed technique is several orders of magnitude faster than a trace-driven simulation while keeping the estimation error within 10% consistently in various communication architecture configurations. Sungchan Kim, Chaeseok Im, Soonhoi Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2004 | Many-to-Many Core-Switch Mapping in 2-D Mesh NoC ArchitecturesabstractIn this paper, we investigate the core-switch mapping (CSM) problem that optimally maps cores onto an NoC architecture such that either the energy consumption or the congestion is minimized. We propose a many-to-many core-switch mapping (mCSM) that allows a switch (core) to have multiple connections to its adjacent cores (switches). We also present decomposition methods that can obtain the suboptimal solutions with enhanced computational efficiency. Our work is the first to provide an exact mixed-integer linear programming (MILP) formulation for the complete CSM problems, including the optimal choice of core placements, switches for each core, and network interfaces for communication flows. Experiments with four random benchmarks show that 4:4 mCSM achieves 81.2% of energy savings and 2.5% of bandwidth savings compared with one-to-one mapping. They also show that, for one-to-one mapping, our optimal solutions obtained by the full MILP save 34.8% of energy consumption and 34.4% of bandwidth requirement compared with those from the existing algorithms. Chan-Eun Rhee, Han-You Jeong, Soonhoi Ha |
ICCD | 3 |
| 2004 | Dynamic voltage scaling for real-time multi-task scheduling using buffersabstractThis paper proposes energy efficient real-time multi-task scheduling (EDF and RM) algorithms by using buffers. The buffering technique overcomes a drawback of previous approaches by utilizing the slack time of a system fully. It increases the CPU utilization and averages the workload of a system, so it enhances the effectiveness of the DVS technique. We target multimedia applications where a slight buffering delay is tolerable within a latency constraint. We modify the state transition and queue handling mechanism of multi-task scheduling in the kernel. In experiments, our algorithms achieve up to 44% of energy consumption saving for EDF scheduling and 49% for RM scheduling with realistic task set configurations and reasonable machine specifications. Chaeseok Im, Soonhoi Ha |
LCTES | 2 |
| 2004 | Memory management for multi-threaded software DSM systems
Yang-Suk Kee, Jin-Soo Kim 0001, Soonhoi Ha |
Parallel Comput. | 3 |
| 2004 | Dynamic voltage scheduling with buffers in low-power multimedia applicationsabstractPower-efficient design of multimedia applications becomes more important as they are used increasingly in many embedded systems. We propose a simple dynamic voltage scheduling (DVS) technique, which suits multimedia applications well and, in case of soft real-time applications, allows all idle intervals of the processor to be fully exploited by using buffers. Our main theme is to determine the minimum buffer size to maximize energy saving in three cases: (i) single task, (ii) multiple subtask, and (iii) multitask. We also present a technique of adjusting task deadlines for further reducing energy consumption in the multiple-subtask and multitask cases. Unlike other DVS techniques using buffers, we guarantee to meet the real-time latency constraint. Experimental results show that the proposed technique does indeed achieve significant power reduction in real-world multimedia applications. Chaeseok Im, Soonhoi Ha, Huiseok Kim |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2003 | Memory access pattern analysis and stream cache design for multimedia applicationsabstractMemory system is a major performance and power bottleneck in embedded systems especially for multimedia applications. Most multimedia applications access stream type of data structures with regular access patterns. It is observed that conventional caches behave poorly for stream-type data structure. Therefore, prediction-based prefetching techniques have been extensively researched to exploit the regular access patterns. Prefetching, however, may pollute the cache if the prediction is not accurate and needs extra hardware prediction logic. To overcome these problems, we propose a novel hardware prefetching technique that is assisted by static analysis of data access pattern with stream caches. With the proposed stream cache architecture, we could achieve significant performance improvement compared with the conventional cache architecture. Junghee Lee 0004, Chanik Park, Soonhoi Ha |
ASP-DAC | 3 |
| 2003 | ParADE: An OpenMP Programming Environment for SMP Cluster SystemsabstractDemand for programming environments to exploit clusters of symmetric multiprocessors (SMPs) is increasing. In this paper, we present a new programming environment, called ParADE, to enable easy, portable, and high-performance programming on SMP clusters. It is an OpenMP programming environment on top of a multi-threaded software distributed shared memory (SDSM) system with a variant of home-based lazy release consistency protocol. To boost performance, the runtime system provides explicit message-passing primitives to make it a hybrid-programming environment. Collective communication primitives are used for the synchronization and work-sharing directives associated with small data structures, lessening the synchronization overhead and avoiding the implicit barriers of work-sharing directives. The OpenMP translator bridges the gap between the OpenMP abstraction and the hybrid programming interfaces of the runtime system. The experiments with several NAS benchmarks and applications on a Linux-based cluster show promising results that ParADE overcomes the performance problem of the conventional SDSM-based OpenMP environment. Yang-Suk Kee, Jin-Soo Kim 0001, Soonhoi Ha |
SC | 3 |
| 2003 | Design and implementation of a user-level Sockets layer over Virtual Interface ArchitectureabstractAbstract The Virtual Interface Architecture (VIA) is an industry standard user‐level communication architecture for system area networks. The VIA provides a protected, directly‐accessible interface to a network hardware, removing the operating system from the critical communication path. In this paper, we design and implement a user‐level Sockets layer over VIA, named SOVIA (Sockets Over VIA). Our objective is to use the SOVIA layer to accelerate the existing Sockets‐based applications with a reasonable effort and to provide a portable and high‐performance communication library based on VIA to application developers. SOVIA realizes comparable performance to native VIA, showing a minimum one‐way latency of 10.5 $\mu$ s and a peak bandwidth of 814 Mbps on Giganet's cLAN. We have shown the functional compatibility with the existing Sockets API by porting File Transfer Protocol (FTP) and Remote Procedure Call (RPC) applications over the SOVIA layer. Compared to the Giganet's LAN Emulation (LANE) driver which emulates TCP/IP inside the kernel, SOVIA easily doubles the file transfer bandwidth in FTP and reduces the latency of calling an empty remote procedure by 77% in RPC applications. Copyright © 2003 John Wiley & Sons, Ltd. Kangho Kim, Sung-In Jung, Soonhoi Ha |
Concurr. Comput. Pract. Exp. | 4 |
| 2002 | Efficient code synthesis from extended dataflow graphs for multimedia applicationsabstractThis paper presents efficient automatic code synthesis techniques from dataflow graphs for multimedia applications. Since multimedia applications require large size buffers containing composite type data, we aim to reduce the buffer sizes with fractional rate dataflow extension and buffer sharing technique. In an H.263 encoder experiment, the FRDF extension and buffer sharing technique enable us to reduce the buffer size by 67%. The final buffer size is no more than in a manual reference code. Hyunok Oh, Soonhoi Ha |
DAC | 2 |
| 2002 | Efficient hardware controller synthesis for synchronous dataflow graph in system level designabstractThis paper concerns automatic hardware synthesis from data flow graph (DFG) specification in system level design. In the presented design methodology, each node of a data flow graph represents a hardware library module that contains a synthesizable VHDL code. Our proposed technique automatically synthesizes a clever control structure, cascaded counter controller, that supports asynchronous interaction with outside modules while efficiently implementing the synchronous dataflow semantics of the graph at the same time. Through comparison with previous works with some examples, the novelty of the proposed technique is demonstrated. Hyunuk Jung, Kangnyoung Lee, Soonhoi Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2002 | Combined data-driven and event-driven scheduling technique for fast distributed cosimulationabstractFast distributed cosimulation is a challenging problem for embedded system design. The main theme of this paper is to increase the simulation speed by reducing the frequency of intersimulator communications, reducing the active duration of simulators, and utilizing the parallelism of component simulators. Those enhancements are accomplished by a proposed virtual synchronization technique, which combines event-driven and data-driven simulation methods. Experimental results show that the proposed technique can boost the cosimulation speed significantly compared with previous conservative approaches. Dohyung Kim 0007, Chan-Eun Rhee, Soonhoi Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2001 | A dataflow specification for system level synthesis of 3D graphics applicationsabstractAbstract- 3D graphics is becoming an important application area together with multimedia applications. The dynamic behavior of 3D graphics application brings new challenges for system level specification and synthesis methodologies. Although existing dataflow models are successfully used for DSP system design, they are not sufficient to deal with 3D graphics algorithms since they lack in global state management and dynamic behavior handling. In this paper, we propose an extended synchronous piggybacked dataflow model with dynamic constructs for representing 3D graphics algorithm. We also present implementation techniques for software and hardware synthesis from the specification. With a simple 3D graphics pipeline, we show the novelty and usefulness of the proposed specification model. I. Chanik Park, Sungchan Kim, Soonhoi Ha |
ASP-DAC | 3 |
| 2001 | xBSP: An Efficient BSP Implementation for clanabstractVirtual Interface Architecture (VIA) is a light-weight protocol for protected user-level zero-copy communication. In spite of the high performance of VIA, the previous MPI implementation for GigaNet's cLAN revealed low communication performance. The main sources of the low performance are the discrepancy of communication model between MPI and VIA and multi-threading overhead. We propose a novel implementation of the Bulk Synchronous Parallel (BSP) programming library for VIA called xBSP for overcoming such problems. To the best of our knowledge, xBSP is the first implementation of the BSP library for VIA. xBSP demonstrates that selecting a proper library is important to exploit the features of light-weight protocols. The intensive use of RDMA operation leads to high performance, close to the native VIA performance with respect to round trip delay and bandwidth. Based on the study of the effects of multithreading, memory registration, and completion policy on performance, we could obtain an efficient BSP implementation for cLAN, which is confirmed by experimental results. Yang-Suk Kee, Soonhoi Ha |
CCGRID | 2 |
| 2001 | Dynamic voltage scheduling technique for low-power multimedia applications using buffersabstractAs multimedia applications are used increasingly in many embedded systems, power efficient design for the applications becomes more important than ever. This paper proposes a simple dynamic voltage scheduling technique, which suits the multimedia applications well. The proposed technique fully utilizes the idle intervals with buffers in a variable speed processor. The main theme of this paper is to determine the minimum buffer size to achieve the maximum energy saving in three cases: single-task, multiple subtasks, and multi-task. Experimental results show that the proposed technique is expected to obtain significant power reduction for several real-world multimedia applications. Chaeseok Im, Huiseok Kim, Soonhoi Ha |
ISLPED | 3 |
| 2000 | Data memory minimization by sharing large size buffersabstractThis paper presents software synthesis techniques to deal with non-primitive data type from graphical dataflow programs based on the synchronous dataflow (SDF) model.Non-primitive data types, often used in multimedia and graphics applications, require buffer memory of large size.To minimize the buffer requirement, we separate global data buffers and local pointer buffers.The proposed approach first allocates the minimum size of global buffers and next binds the local buffers to the global buffers by setting the pointers.Static binding and dynamic binding techniques are devised.Experimental results prove the significance of the proposed techniques. Hyunok Oh, Soonhoi Ha |
ASP-DAC | 2 |
| 2000 | Memory efficient software synthesis with mixed coding style from dataflow graphsabstractThis paper presents a set of techniques to reduce the code and data sizes for software synthesis from graphical digital signal-processing programs based on the synchronous dataflow model. By sharing the kernel code among multiple instances of a block with a shared function, we can further reduce the code size below the previous results based on inline coding style. A systematic approach also is devised to give up the single appearance schedule for reducing the data buffer requirement. The proposed techniques have been evaluated with two real-life examples to prove their significance. Wonyong Sung, Soonhoi Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1998 | A Hardware Software Cosimulation Backplane with Automatic Interface GenerationabstractA hardware software cosimulation environment is developed using the backplane approach. This paper defines the backplane protocol for communication and synchronization between client simulators to seamlessly integrate a new simulator without modification. Automatic interface generation facility is also devised for more effective cosimulation environment. The environment is implemented based on Ptolemy and validated with QAM example run on different configurations. Wonyong Sung, Soonhoi Ha |
ASP-DAC | 2 |
| 1998 | Rate Optimal VLSI Design from Data Flow Graph
Moonwook Oh, Soonhoi Ha |
DAC | 2 |
| 1998 | Optimized Timed Hardware Software Cosimulation without Roll-backabstractAn optimized hardware software cosimulation method based on the backplane approach is presented in this paper. To enhance the performance of cosimulation, efforts are focused on reducing control packets between simulators as well as concurrent execution of simulators without roll-back. Wonyong Sung, Soonhoi Ha |
DATE | 2 |
| 1998 | Relaxed Barrier Synchronization for the BSP Model of Computation on Message-Passing Architectures
Soonhoi Ha, Chu Shik Jhon |
Inf. Process. Lett. | 2 |
| 1997 | Quantitative Analysis on Caching Effect of I-Structure Data in Frame-Based Multithreaded ProcessingabstractSince long latency due to remote memory access could be tolerated by rapidly switching to another thread in multithreaded processing, caching I-structure data is expected to have less beneficial effect on the performance than caching ordinary data. In this paper we show that caching I-structure data could improve the overall performance in spite of latency tolerating property of multithreading. Our quantitative analysis reveals that the most important caching effect off-structure data in frame-based multithreading is the enhancement of frame parallelism. It reduces the idle time due to latency by lowering latency sensitivity and at the same time decreases the thread processing time by exploiting more processors. Hyong-Shik Kim, Soonhoi Ha, Chu Shik Jhon |
ICPP | 2 |
| 1997 | Reducing Overheads of Local Communications in Fine-grain Parallel ComputationabstractFor fine-grain computation to be effective, the cost of communications between the large number of subtasks should be minimised. In this paper we present an optimization technique which reduces overheads of communications between local subtasks by bypassing the network interface and transferring data directly from memory or registers to memory. On average, the optimization results in 35.6% improvement in total execution time on instruction-level simulations with six benchmark programs from 1 to 32 nodes. Soonhoi Ha, Chu Shik Jhon |
ICPP | 2 |
| 1996 | COP: a Crosstalk OPtimizer for gridded channel routingabstractThe interwire spacing in a VLSI chip becomes closer as the VLSI fabrication technology rapidly evolves. Accordingly, it becomes important to consider crosstalk caused by the coupling capacitance between adjacent wires in the layout design for the fast and safe VLSI circuits. The upper bounds of the allowable crosstalk for nets, called crosstalk constraints, are usually given in the design specification. This paper proposes a crosstalk minimization technique based on segment rearrangement for gridded channel routing. The technique repeatedly rearranges horizontal wire segments and/or increase the number of tracks to satisfy the crosstalk constraints. With experiments, we observed that the presented technique is more effective than the track permutation technique. Kyoung-Son Jhang, Soonhoi Ha, Chu Shik Jhon |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | An integrated hardware-software cosimulation environment for heterogeneous systems prototypingabstractNo abstract available. Kyuseok Kim, Youngsoo Shin, Taekyoon Ahn, Wonyong Sung, Kiyoung Choi, Soonhoi Ha |
ASP-DAC | 7 |
| 1991 | Multirate signal processing in PtolemyabstractThe use of two models of computation, synchronous dataflow (SDF) and dynamic dataflow (DDF), to design and implement signal processing applications with multiple sample rates is discussed. The SDF model is used for synchronous applications. SDF is amenable to compile-time scheduling, and hence is much more efficient at runtime. The design environment, Ptolemy, can simultaneously support multiple models of computation, so SDF and DDF can be combined in a single application. Hence, the implementation will incur the run-time cost of DDF only for those asynchronous portions that absolutely must incur such cost. As an illustration, the authors detail a synchronous application, sample-rate conversion using polyphase filters, and an asynchronous application, timing recovery for an amplitude-shift-keyed signal.> Joseph T. Buck, Soonhoi Ha, Edward A. Lee, David G. Messerschmitt |
ICASSP | 2 |
| 1991 | Compile-Time Scheduling and Assignment of Data-Flow Program Graphs with Data-Dependent IterationabstractFour scheduling strategies for dataflow graphs onto parallel processors are classified: (1) fully dynamic, (2) static-assignment, (3) self-timed, and (4) fully static. Scheduling techniques valid for strategies (2), (3), and (4) are proposed. The focus is on dataflow graphs representing data-dependent iteration. A known probability mass function for the number of cycles in the data-dependent iteration is assumed, and how a compile-time decision about assignment and/or ordering as well as timing can be made is shown. The criterion used is to minimize the expected total idle time caused by the iteration. In certain cases, this will also minimize the expected makespan of the schedule. How to determine the number of processors that should be assigned to the data-dependent iteration is shown. The method is illustrated with a practical programming example.> Soonhoi Ha, Edward A. Lee |
IEEE Trans. Computers | 1 |