VLDB 2026 Research / reviewers in the wild / expert
Hoeseok Yang
dblp:35/5349
· DBLP profile ↗
33ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-7929-7470ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | oFFN: Outlier and Neuron-aware Structured FFN for Fast yet Accurate LLM InferenceabstractWith the advent of large-scale language models (LLMs), various optimization techniques have been proposed to enable efficient inference. Among these, methods that aggressively exploit output activation sparsity have attracted significant attention, which leverage ReLU-fied LLMs and skip the entire memory accesses as well as the computation for the output element if it was predicted as sparse. Achieving fast and accurate prediction of output activation sparsity is crucial to enhancing inference efficiency. However, in practice, phenomena such as activation outliers and hot and cold neurons, which significantly affect the exploitation of sparsity during LLM inference, have either been addressed individually or not structurally integrated in existing work. Geunsoo Song, Hoeseok Yang, Youngmin Yi |
ASPLOS (2) | 2 |
| 2025 | Grasp: Group-based Prediction of Activation Sparsity for Fast LLM InferenceabstractOptimizing LLM inference has become increasingly important as the demand for efficient on-device deployments grows. To reduce the computational overhead in the MLP components, which account for a significant portion of LLM inference, ReLU-fied LLMs have been introduced to maximize activation sparsity. Several sparsity prediction methods have been developed to efficiently skip unnecessary memory accesses and computations by predicting activation sparsity. In this paper, we propose a novel magnitude-based, training-free sparsity prediction technique called Grasp that builds on the existing sign bitbased method for ReLU-fied LLMs. The proposed method enhances prediction accuracy by grouping values considering the distribution within vectors and explicitly accounting for statistical outliers. This allows us to estimate the impact of each element more accurately yet in an efficient way, improving both activation sparsity prediction accuracy and computational efficiency. Compared to the-state-of-the-art technique, Grasp achieves higher sparsity prediction accuracy and $11 \%$ higher skipping efficiency, which corresponds to $1.85 \times$ speedup against the dense inference. Hoeseok Yang, Youngmin Yi |
DAC | 2 |
| 2025 | SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM InferenceabstractLeveraging sparsity is crucial for optimizing large language model (LLM) inference; however, modern LLMs employing SiLU as their activation function exhibit minimal activation sparsity. Recent research has proposed replacing SiLU with ReLU to induce significant activation sparsity and showed no downstream task accuracy degradation through fine-tuning. However, taking full advantage of it required training a predictor to estimate this sparsity. In this paper, we introduce Sparselnfer, a simple, lightweight, and training-free predictor for activation sparsity of ReLU-fied LLMs, in which activation sparsity is predicted by comparing only the sign bits of inputs and weights. To compensate for possible prediction inaccuracy, an adaptive tuning of the predictor's conservativeness is enabled, which can also serve as a control knob for optimizing LLM inference. The proposed method achieves approximately 21% faster inference speed over the state-of-the-art, with negligible accuracy loss of within 1 % p. Hoeseok Yang, Youngmin Yi |
DATE | 2 |
| 2025 | Accelerating Retrieval Augmented Language Model via PIM and PNM Integration
Je-Woo Jang, Junyong Oh, Youngbae Kong, Jae-Youn Hong, Sung-Hyuk Cho, Jeongyeol Lee, Hoeseok Yang, Joon-Sung Yang |
MICRO | 7 |
| 2024 | Platform Design for Privacy-Preserving Federated Learning using Homomorphic Encryption : Wild-and-Crazy-Idea PaperabstractFederated learning (FL) has been increasingly widely used for distributed and privacy-preserving machine learning (ML) environments, as the raw training data can stay local to clients while leveraging model updates from individual clients. Homomorphic encryption (HE) technologies can provide additional privacy protection for FL by encrypting the model update parameters while allowing model aggregation on a remote server. Although HE-enabled FL seems to be a promising privacy-preserving ML solution, it requires significantly more computational and memory resources, requiring a dedicated hardware and software platform. In this paper, we discuss preliminary but concrete and realizable research ideas for analyzing the requirements for HE-enabled FL and for designing a hardware and software platform. Furthermore, we propose a platform co-design process that considers various design stages and challenges in the platform co-design. Hokeun Kim, Younghyun Kim 0001, Hoeseok Yang |
FDL | 3 |
| 2024 | Machine Learning-Driven Burrowing with a Snake-Like RobotabstractSubterranean burrowing is inherently difficult for robots because of the high forces experienced as well as the high amount of uncertainty in this domain. Because of the difficulty in modeling forces in granular media, we propose the use of a novel machine-learning control strategy to obtain optimal techniques for vertical self-burrowing. In this paper, we realize a snake-like bio-inspired robot that is equipped with an IMU and two triple-axis magnetometers. Utilizing magnetic field strength as an analog for depth, a novel deep learning architecture was proposed based on sinusoidal and random data in order to obtain a more efficient strategy for vertical self-burrowing. This strategy was able to outperform many other standard burrowing techniques and was able to automatically reach targeted burrowing depths. We hope these results will serve as a proof of concept for how optimization can be used to unlock the secrets of navigating in the subterranean world more efficiently. Sean Even, Holden Gordon, Hoeseok Yang, Yasemin Ozkan Aydin |
ICRA | 3 |
| 2021 | A GPU Architecture Aware Fine-Grain Pruning Technique for Deep Neural Networks
Kyusik Choi, Hoeseok Yang |
Euro-Par | 2 |
| 2021 | Scheduling of Iterative Computing Hardware Units for Accuracy and Energy EfficiencyabstractIterative computing, where the output accuracy gradually improves over multiple iterations, enables dynamic reconfiguration of energy-quality trade-offs by adjusting the latency (i.e., number of iterations). In order to take full advantage of the dynamic reconfigurability of iterative computing hardware, an efficient method for determining the optimal latency is crucial. In this paper, we introduce an integer linear programming (ILP)- based scheduling method to determine the optimal latency of iterative computing hardware. We consider the input-dependence of output accuracy of approximate hardware using data-driven error modeling for accurate quality estimation. The proposed method finds optimal or near-optimal latency with a significant speedup compared to exhaustive search and decision tree-based optimization. Setareh Behroozi, Hoeseok Yang, Younghyun Kim 0001 |
ISCAS | 3 |
| 2021 | Real-Time Schedulability Analysis and Enhancement of Transiently Powered Processors With NVMsabstractRecent Internet-of-Things or Wireless Sensor Network devices are often operated with energy harvesters. As there are no energy storages in those devices, power is not consistently provided to the devices at all times. In such transiently powered systems, in order to keep the system reliable without losing any execution contexts, non-volatile memories (NVMs) are typically used for swift backup/restoration of execution contexts. In this article, we perform a real-time schedulability analysis of the transiently powered processors with NVMs. We first quantitatively characterize the charging and discharging behaviors of the energy harvester and extract the compute capability of the system in time interval domain. Then, based on Real-Time Calculus, we determine whether the given multi-task workload is schedulable or not with respect to the earliest deadline first (EDF) or fixed-priority (FP) scheduling policies. In addition, we study how the choice of the threshold voltage parameter affects the schedulability, then propose a feasible threshold selection algorithm to enhance schedulability. We verify the effectiveness of the proposed technique with extensive simulations. Compared to the naive selection method, the proposed technique always shows improvements in schedulability in various workloads. Dasom Lee, Hyeonseok Jung, Hoeseok Yang |
IEEE Trans. Computers | 3 |
| 2020 | Improvement of CNN-Based Road Extraction from Satellite Images via Morphological Image ProcessingabstractIn this paper, we propose to improve the recall of CNN-based road extraction from satellite images by means of label thickening and thinning. With the thickened road labels, the CNN is led to extract roads more aggressively, preserving the topological information of roads. After inference, the predicted segment maps need to be thinned back to the original width. The proposed technique has been evaluated with an existing road extraction dataset in various degrees of thickening. Throughout the experiments, the relaxed recall score has been successfully improved by the proposed technique, reducing the number of false-negative pixels. However, at the same time, it has been observed that the number of false-positive pixels also increases slightly. Overall, it has been visually observed that the topological information of roads such as connectivity is better extracted by the proposed technique. Heeji Im, Hoeseok Yang |
IGARSS | 2 |
| 2019 | Optimization of Fault-Tolerant Mixed-Criticality Multi-Core Systems with Enhanced WCRT AnalysisabstractThis article proposes a novel optimization technique of fault-tolerant mixed-criticality multi-core systems with worst-case response time (WCRT) guarantees. Typically, in fault-tolerant multi-core systems, tasks can be replicated or re-executed in order to enhance the reliability. In addition, based on the policy of mixed-criticality scheduling, low-criticality tasks can be dropped at runtime. Such uncertainties caused by hardening and mixed-criticality scheduling make WCRT analysis very difficult. We show that previous analysis techniques are pessimistic as they consider avoidably extreme cases that can be safely ignored within the given reliability constraint. We improve the analysis in order to tighten the pessimism of WCRT estimates by considering the maximum number of faults to be tolerated. Further, we improve the mixed-criticality scheduling by allowing partial dropping of low-criticality tasks. On top of those, we explore the design space of hardening, task-to-core mapping, and quality-of-service of the multi-core mixed-criticality systems. The effectiveness of the proposed technique is verified by extensive experiments with synthetic and real-life benchmarks. Junchul Choi, Hoeseok Yang, Soonhoi Ha |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2018 | Context-aware dataflow adaptation technique for low-power multi-core embedded systemsabstractToday's embedded systems operate under increasingly dynamic conditions. First, computational workloads can be either fluctuating or adjustable. Moreover, as many devices are battery-powered, it is common to have runtime power management technique, which results in dynamic power budget. This paper presents a design methodology for multi-core systems, based on dataflow specification, that can deal with various contexts. We optimize the original dataflow considering various working conditions, then, autonomously adapt it to a pre-defined optimal form in response to context changes. We show the effectiveness of the proposed technique with a real-life case study and synthetic benchmarks. Hyeonseok Jung, Hoeseok Yang |
DAC | 2 |
| 2017 | Executable dataflow benchmark generation technique for multi-core embedded systemsabstractAs the complexity of multi-core embedded systems continuously grows, the optimization and verification of such systems become non-trivial. Thus, it is important to secure a set of benchmarks of reasonable complexity to validate the design of multi-core embedded systems. Dataflow model has long been considered as a suitable model-of-computation for specifying the behavior of embedded systems. In this paper, we proposes a dataflow benchmark generation technique for multi-core embedded systems, leveraging two existing tools: a random dataflow topology generator and a random C code generator. In the proposed technique, as a preparatory step, a C code database is established by means of a random C code generation tool. Then, a random dataflow graph, with execution time information annotated to each node, is generated by an existing tool. For each node in the generated graph, a number of randomly generated C code segments are properly chosen and accommodated in a single function as per the given execution time information. In doing so, a set of linear equations are derived and solved. Subsequently, using existing model-based embedded system design frameworks, we automatically generate an executable benchmark for the entire dataflow graph. Further, in order to enhance the accuracy of the generated code, a simple calibration technique is applied after the generation and test runs. It is shown that the generated codes assure the diversity and complexity as embedded software benchmark for multi-core embedded systems. Jeonggyu Jang, Hoeseok Yang |
RSP | 2 |
| 2016 | Real-time co-scheduling of multiple dataflow graphs on multi-processor systemsabstractIt is challenging to schedule multiple dataflow applications concurrently on multi-processor embedded systems with processor sharing. As a viable solution, an approach has been proposed recently, in which the dataflow graphs are transformed into a set of independent realtime tasks. However, it may produce poor resource utilization and excessive buffer usage. Alternatively, we propose a novel two-phase scheduling technique. In the first phase, a set of static schedules is produced for each dataflow considering the resource sharing possibility; Then, we use a meta-heuristic to find the combination of per-graph schedules to minimize the resource requirement by processor sharing. We show that the proposed technique exhibits better resource and buffer efficiency. Shin-Haeng Kang, Duseok Kang, Hoeseok Yang, Soonhoi Ha |
DAC | 3 |
| 2016 | A Formal Approach to Power Optimization in CPSs With Delay-Workload Dependence AwarenessabstractThe design of cyber-physical systems (CPSs) faces various new challenges that are unheard of in the design of classical real-time systems. Power optimization is one of the major design goals that is witnessing such new challenges. The presence of interaction between the cyber and physical components of a CPS leads to dependence between the time delay of a computational task and the amount of workload in the next iteration. We demonstrate that it is essential to take this delay-workload dependence into consideration in order to achieve low power consumption. In this paper, we identify this new challenge, and present the first formal and comprehensive model to enable rigorous investigations on this topic. We propose a simple power management policy, and show that this policy achieves a best possible notion of optimality. In fact, we show that the optimal power consumption is attained in a “steady-state” operation and a simple policy of finding and entering this steady state suffices, which can be quite surprising considering the added complexity of this problem. Finally, we validated the efficiency of our policy with experiments. Hyung-Chan An, Hoeseok Yang, Soonhoi Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Modeling and power optimization of cyber-physical systems with energy-workload tradeoffabstractIn this paper, we propose to take the relationship between delay and workload into account in the optimization of cyber-physical systems (CPSs). Since the components at the physical side continuously change their values or properties, a longer delay at the cyber part may result in a bigger workload for the next computation. We formulate this tradeoff and apply it to the power optimization of CPS. In doing so, we examine the schedulability of the given CPS with respect to the given parameters and initial workload. Then, we propose to keep the system operate in the stable state with minimum scaling factor and prove that it is better than any alternating sequences. We verify the validity of the proposed delay-workload model by measuring the execution delay of real-life examples. The effectiveness of the proposed power optimization policy is demonstrated with simulation results. Hoeseok Yang, Soonhoi Ha |
ISLPED | 1 |
| 2014 | AdaPNet: Adapting process networks in response to resource variationsabstractA widely considered strategy to prevent interference issues on multi-processor systems is to isolate the execution of the individual applications by running each of them on a dedicated virtual guest machine. The amount of computing power available to a single application, however, depends on the other applications running on the system and may change over time. A promising approach to maximize the performance under such conditions is to adapt the application's degree of parallelism when the resources allocated to the application are changed. This enables an application to exploit not more parallelism than required, thereby reducing inter-process communication and scheduling overheads. In this paper, we introduce AdaPNet, a run-time system to execute streaming applications, which are modeled as process networks, efficiently on platforms with dynamic resource allocation. AdaPNet responds to changes in the available resources by first calculating a process network that maximizes the performance of the application on the new resources. Then, AdaPNet transparently transforms the application into the alternative network without discarding the program state. Targeting two many-core systems, we demonstrate that AdaPNet outperforms comparable run-time systems, which do not adapt the degree of parallelism, in terms of speed-up and memory usage. Lars Schor, Iuliana Bacivarov, Hoeseok Yang, Lothar Thiele |
CASES | 3 |
| 2014 | On the Scheduling of Fault-Tolerant Mixed-Criticality SystemsabstractWe consider in this paper fault-tolerant mixed-criticality scheduling, where heterogeneous safety guarantees must be provided to functionalities (tasks) of varying criticalities (importances). We model explicitly the safety requirements for tasks of different criticalities according to safety standards, assuming hardware transient faults. We further provide analysis techniques to bound the effects of task killing and service degradation on the system safety and schedulability. Based on our model and analysis, we show that our problem can be converted to a conventional mixed-criticality scheduling problem. Thus, we broaden the scope of applicability of the conventional mixed-criticality scheduling techniques. Our proposed techniques are validated with a realistic flight management system application and extensive simulations. Pengcheng Huang 0001, Hoeseok Yang, Lothar Thiele |
DAC | 2 |
| 2014 | Static Mapping of Mixed-Critical Applications for Fault-Tolerant MPSoCsabstractThis paper presents a static mapping optimization technique for fault-tolerant mixed-criticality MPSoCs. The uncertainties imposed by system hardening and mixed criticality algorithms, such as dynamic task dropping, make the worst-case response time analysis difficult for such systems. We tackle this challenge and propose a worst-case analysis framework that considers both reliability and mixed-criticality concerns. On top of that, we build up a design space exploration engine that optimizes fault-tolerant mixed-criticality MPSoCs and provides worst-case guarantees. We study the mapping optimization considering judicious task dropping, that may impose a certain service degradation. Extensive experiments with real-life and synthetic benchmarks confirm the effectiveness of the proposed technique. Shin-Haeng Kang, Hoeseok Yang, Sungchan Kim, Iuliana Bacivarov, Soonhoi Ha, Lothar Thiele |
DAC | 2 |
| 2014 | Reliability-aware mapping optimization of multi-core systems with mixed-criticalityabstractThis paper presents a novel mapping optimization technique for mixed critical multi-core systems with different reliability requirements. For this scope, we derived a quantitative reliability metric and presented a scheduling analysis that certifies given mixed-criticality constraints. Our framework is capable of investigating re-execution, passive replication, and modular redundancy with optimized voter placement, while typical hardening approaches consider only one or two of these techniques. The proposed technique complies with existing safety standards and is power-efficient, as demonstrated by our experiments. Shin-Haeng Kang, Hoeseok Yang, Sungchan Kim, Iuliana Bacivarov, Soonhoi Ha, Lothar Thiele |
DATE | 2 |
| 2014 | COOLIP: Simple yet effective job allocation for distributed thermally-throttled processorsabstractThermal constraints limit the time for which a processor can run at high frequency. Such thermal-throttling complicates the computation of response times of jobs. For multiple processors, a key decision is where to allocate the next job. For distributed thermally-throttled procesosrs, we present COOLIP with a simple allocation policy: a job is allocated to the earliest available processor, and if there are several available simultaneously, to the coolest one. For Poisson distribution of inter-arrival times and Gaussian distribution of execution demand of jobs, COOLIP matches the 95-percentile response time of Earliest Finish-Time (EFT) policy which minimizes response time with full knowledge of execution demand of unfinished jobs and thermal models of processors. We argue that COOLIP performs well because it directs the processors into states such that a defined sufficient condition of optimality holds. Hoeseok Yang, Iuliana Bacivarov, Lothar Thiele |
DATE | 2 |
| 2013 | Expandable process networks to efficiently specify and explore task, data, and pipeline parallelismabstractRunning each application of a many-core system on an isolated (virtual) guest machine is a widely considered solution for performance and reliability issues. When a new application is started, the guest machine is assigned with an amount of computing resources that depends on the overall workload of the system and is not known to the designer at specification time. For instance, the computing resources might consist of many slow or a few fast processing elements. If the application is statically specified, as, for example, with Kahn process networks, the number of processing elements usable by an application is upper bounded by its number of processes. Similarly, the inter-process communication overhead might limit the maximum performance if the number of processing elements is significantly smaller than the number of processes. In this paper, we propose a formal extension for streaming programming models called expandable process networks (EPNs) that tackles this challenge by abstracting several possible granularities in a single specification. This enables the automatic exploration of task, data, and pipeline parallelism by two basic design transformation techniques, namely replication and unfolding. Then, the EPN semantics facilitates the synthesis of multiple design implementations that are all derived from one high-level specification. At runtime, the best fitting implementation for the given computing resources is selected to maximize the performance. Finally, we demonstrate the effectiveness of the proposed model on Intel's 48-core SCC processor. Lars Schor, Hoeseok Yang, Iuliana Bacivarov, Lothar Thiele |
CASES | 2 |
| 2013 | Efficient Worst-Case Temperature Evaluation for Thermal-Aware Assignment of Real-Time Applications on MPSoCs
Lars Schor, Iuliana Bacivarov, Hoeseok Yang, Lothar Thiele |
J. Electron. Test. | 3 |
| 2013 | Real-time worst-case temperature analysis with temperature-dependent parameters
Hoeseok Yang, Iuliana Bacivarov, Devendra Rai, Jian-Jia Chen, Lothar Thiele |
Real Time Syst. | 1 |
| 2013 | Predictability for timing and temperature in multiprocessor system-on-chip platformsabstractHigh computational performance in multiprocessor system-on-chips (MPSoCs) is constrained by the ever-increasing power densities in integrated circuits, so that nowadays MPSoCs face various thermal issues. For instance, high chip temperatures may lead to long-term reliability concerns and short-term functional errors. Therefore, the new challenge in designing embedded real-time MPSoCs is to guarantee the final performance and correct function of the system, considering both functional and non-functional properties. One way to achieve this is by ruling out mapping alternatives that do not fulfill requirements on performance or peak temperature already in early design stages. In this article, we propose a thermal-aware optimization framework for mapping real-time applications onto MPSoC platforms. The performance and temperature of mapping candidates are evaluated by formal temporal and thermal analysis models. To this end, analysis models are automatically generated during design space exploration, based on the same specifications as used for software synthesis. The analysis models are automatically calibrated with performance data reflecting the execution of the system on the target platform. The data is automatically obtained prior to design space exploration based on a set of benchmark mappings. Case studies show that the performance and temperature requirements are often conflicting goals and optimizing them together leads to major benefits in terms of a guaranteed and predictable high performance. Lothar Thiele, Lars Schor, Iuliana Bacivarov, Hoeseok Yang |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2012 | Power agnostic technique for efficient temperature estimation of multicore embedded systemsabstractTemperature plays an increasingly important role in the overall performance and reliability of a computing system. Multi- and many-core systems provide an opportunity to manage the overall temperature profile by cleverly designing the application-to-core mapping and the associated scheduling policies. An uncontrolled temperature profile may lead to an unplanned performance loss, since the system activates protective mechanisms such as voltage and/or frequency scaling to cool itself. Similarly, deep thermal cycles with high frequency lead to severe deterioration in the overall reliability of the system. Design space exploration tools are often used to optimize binding and scheduling choices based on a given set of constraints and objectives, thus motivating the need for fast and accurate temperature estimation techniques. We argue that the currently available techniques are not an ideal fit to design space exploration tools, and suggest a system level technique which is based on application fingerprinting. It does not need any information about the processor floorplan, the physical and thermal structure, or about power consumption. Instead, its temperature estimation is based on a set of application-specific calibration runs and associated temperature measurements using available built-in sensors. We show that a given application possesses a unique thermal signature on the system it executes on, which provides a computationally fast method to calculate accurate temperature traces. Extensive experimental studies show that our technique can estimate temperature on all cores of a system to within $5^{o}C$, and is three orders of magnitude faster than state of the art numerical simulators like \emph{Hotspot.} Devendra Rai, Hoeseok Yang, Iuliana Bacivarov, Lothar Thiele |
CASES | 2 |
| 2012 | Scenario-based design flow for mapping streaming applications onto on-chip many-core systemsabstractThe next generation of embedded software has high performance requirements and is increasingly dynamic. Multiple applications are typically sharing the system, running in parallel in different combinations, starting and stopping their individual execution at different moments in time. The different combinations of applications are forming system execution scenarios. In this paper, we present the distributed application layer, a scenario-based design flow for mapping a set of applications onto heterogeneous on-chip many-core systems. Applications are specified as Kahn process networks and the execution scenarios are combined into a finite state machine. Transitions between scenarios are triggered by behavioral events generated by either running applications or the run-time system. A set of optimal mappings are precalculated during design-time analysis. Later, at run-time, hierarchically organized controllers monitor behavioral events, and apply the precalculated mappings when starting new applications. To handle architectural failures, spare cores are allocated at design-time. At run-time, the controllers have the ability to move all processes assigned to a faulty physical core to a spare core. Finally, we apply the proposed design flow to design and optimize a picture-in-picture software. Lars Schor, Iuliana Bacivarov, Devendra Rai, Hoeseok Yang, Shin-Haeng Kang, Lothar Thiele |
CASES | 4 |
| 2012 | Worst-Case Temperature Guarantees for Real-Time Applications on Multi-core SystemsabstractDue to increased on-chip power density, multi-core systems face various thermal issues. In particular, exceeding a certain threshold temperature can reduce the system's performance and reliability. Therefore, when designing a real-time application with non-deterministic workload, the designer has to be aware of the maximum possible temperature of the system. This paper proposes an analytic method to calculate an upper bound on the worst-case peak temperature of a real-time system with multiple cores generated under all possible scenarios of task executions. In order to handle a broad range of uncertainties, task arrivals are modeled as periodic event streams with jitter and delay. Finally, the proposed method is applied to a multi-core ARM platform and our results are validated in various case studies. Lars Schor, Iuliana Bacivarov, Hoeseok Yang, Lothar Thiele |
IEEE Real-Time and Embedded Technology and Applications Symposium | 3 |
| 2011 | Thermal-aware system analysis and software synthesis for embedded multi-processorsabstractNowadays, the reliability and performance of modern embedded multi-processor systems is threaten by the ever-increasing power densities in integrated circuits, and a new additional goal of software synthesis is to reduce the peak temperature of the system. However, in order to perform thermal-aware mapping optimization, the timing and thermal characteristics of every candidate mapping have to be analyzed. While the task of analyzing timing characteristics of design alternatives has been extensively investigated in recent years, there is still a lack of methods for accurate and fast thermal analysis. In order to obtain desired evaluation times, the system has to be simulated at a high abstraction level. This often results in a loss of accuracy, mainly due to missing knowledge of system's characteristics. This paper addresses this challenge and presents methods to automatically calibrate high-level thermal evaluation methods. Furthermore, the viability of the methods for automated model calibration is illustrated by means of a novel high-level thermal evaluation method. Lothar Thiele, Lars Schor, Hoeseok Yang, Iuliana Bacivarov |
DAC | 3 |
| 2011 | Worst-case temperature analysis for real-time systemsabstractWith the evolution of today's semiconductor technology, chip temperature increases rapidly mainly due to the growth in power density. For modern embedded real-time systems, it is crucial to estimate maximal temperatures in order to take mapping or other design decisions to avoid burnout, and still be able to guarantee meeting real-time constraints. This paper provides answers to the question: When work-conserving scheduling algorithms, such as earliest-deadline-first (EDF), rate-monotonie (RM), deadline-monotonic (DM), are applied, what is the worst-case peak temperature of a real-time embedded system under all possible scenarios of task executions? We propose an analytic framework, which considers a general event model based on network and real-time calculus. This analysis framework has the capability to handle a broad range of uncertainties in terms of task execution times, task invocation periods, and jitter in task arrivals. Simulations show that our framework is a cornerstone to design real-time systems that have guarantees on both schedulability and maximal temperatures. Devendra Rai, Hoeseok Yang, Iuliana Bacivarov, Jian-Jia Chen, Lothar Thiele |
DATE | 2 |
| 2010 | An MILP-Based Performance Analysis Technique for Non-Preemptive Multitasking MPSoCabstractFor real-time applications, it is necessary to estimate the worst-case performance early in the design process without actual hardware implementation. While the non-preemptive task scheduling is pertinent to multi-core platforms because of easy implementation and high performance, its scheduling anomaly behavior makes the worst-case performance estimation extremely difficult. In this paper, we propose an analysis technique based on mixed integer linear programming (MILP) to estimate the worst-case performance of each task in a non-preemptive multitask application on multi-processor system-on-chip architecture. MILP provides a systematic way to describe the complex interaction among task scheduling, communication architecture, and task execution, which affects the worst-case behavior dynamically. The proposed analysis technique overcomes several limitations that previous work usually has; it allows multiple tasks with different periods and models contention on the communication architecture. We show that the proposed analysis takes affordable computation time to make it of practical value even though it has exponential complexity in theory. The proposed technique estimates a safe bound on task latency statistically, which is demonstrated by extensive random simulations. Hoeseok Yang, Sungchan Kim, Soonhoi Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | Pipelined data parallel task mapping/scheduling technique for MPSoCabstractIn this paper, we propose a multi-task mapping/scheduling technique for heterogeneous and scalable MPSoC. To utilize the large number of cores embedded in MPSoC, the proposed technique considers temporal and data parallelisms as well as task parallelism. We define a multi-task mapping/scheduling problem with all these parallelisms and propose a QEA (quantum-inspired evolutionary algorithm)-based heuristic. Compared with an ILP (Integer Linear Programming) approach, experiments with real-life examples show the feasibility and the efficiency of the proposed technique. Hoeseok Yang, Soonhoi Ha |
DATE | 1 |
| 2007 | Performance evaluation and optimization of dual-port SDRAM architecture for mobile embedded systemsabstractRecently dual-port SDRAM (DPSDRAM) architecture tailored for dual-processor based mobile embedded systems has been announced where a single memory chip plays the role of the local memories and the shared memory for both processors. In order to keep memory consistency from simultaneous accesses of both ports, every access to the shared memory should be protected by a synchronization mechanism, which can result in substantial access latency. We propose two optimization techniques by exploiting the communication patterns of target application: lock-priority scheme and static-copy scheme. Further, by dividing the shared bank into multiple blocks, we enable simultaneous accesses to different blocks and achieve considerable performance gain. Experiments on a virtual prototyping system show a promising result that we achieve about 20-50% performance gain compared to the base DPSDRAM architecture. Hoeseok Yang, Sungchan Kim, Hae-woo Park, Soonhoi Ha |
CASES | 1 |