EDBT 2026 Demo / reviewers in the wild / expert
Dongkun Shin
dblp:05/3861
· DBLP profile ↗
36ranked-venue papers
10as first author
7since 2021 · last 2025
0000-0001-7235-7787ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 9 first-author · 4 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | D2FS: Device-Driven Filesystem Garbage Collection
Joontaek Oh, Dongkun Shin, Youjip Won |
FAST | 4 |
| 2024 | Proxyformer: Nyström-Based Linear Transformer with Trainable Proxy TokensabstractTransformer-based models have demonstrated remarkable performance in various domains, including natural language processing, image processing and generative modeling. The most significant contributor to the successful performance of Transformer models is the self-attention mechanism, which allows for a comprehensive understanding of the interactions between tokens in the input sequence. However, there is a well-known scalability issue, the quadratic dependency (i.e. O(n^2)) of self-attention operations on the input sequence length n, making the handling of lengthy sequences challenging. To address this limitation, there has been a surge of research on efficient transformers, aiming to alleviate the quadratic dependency on the input sequence length. Among these, the Nyströmformer, which utilizes the Nyström method to decompose the attention matrix, achieves superior performance in both accuracy and throughput. However, its landmark selection exhibits redundancy, and the model incurs computational overhead when calculating the pseudo-inverse matrix. We propose a novel Nyström method-based transformer, called Proxyformer. Unlike the traditional approach of selecting landmarks from input tokens, the Proxyformer utilizes trainable neural memory, called proxy tokens, for landmarks. By integrating contrastive learning, input injection, and a specialized dropout for the decomposed matrix, Proxyformer achieves top-tier performance for long sequence tasks in the Long Range Arena benchmark. Hayun Lee, Dongkun Shin |
AAAI | 3 |
| 2022 | Structured Pruning for Deep Convolutional Neural Networks via Adaptive Sparsity RegularizationabstractStructured pruning is a promising method to reduce the computational cost and memory load, and then accelerate the inference process of deep neural networks. Therefore, it facilitates the deep convolutional model application on resource-constrained devices, such as IoT devices and embedded systems. To make a compact deep convolutional neural network model, many network pruning methods enforce the sparsity by imposing sparse constraints on weight parameters during training and pruning some insignificant weights. However, those methods usually impose sparse constraints without any guidance, while simply limiting the compression rate. This paper proposes a dynamic scheme to impose the sparse constraints according to the filter weights, which can guide sparsity-induced training to choose import channels in a deep convolutional neural network. The extensive experiments present that our adaptive sparsity-induced training is more efficient than a static training scheme. Compared to the existing techniques, the proposed method achieved a superior pruning performance on CIFAR10 with 91.5% parameters reduction and 61.6% floating point operations (FLOPs). On the object detection task, our method achieved 91.2% parameter reductions and 65.8% FLOPs reductions on YOLOV5s. Furthermore, we observed a significant acceleration on the inference process of the pruned models on real resource-constrained devices. Tuanjie Shao, Dongkun Shin |
COMPSAC | 2 |
| 2022 | Lifetime-leveling LSM-tree compaction for ZNS SSDabstractThe Log-Structured Merge (LSM) tree is considered well-suited to zoned namespace (ZNS) storage devices since the write requests to LSM-tree is sequential. However, zones can be partially invalidated and be fragmented during LSM-tree compaction. The partially-invalid zones cannot be utilized and thus space amplification becomes significant. To reclaim the invalid space, host-managed garbage collection (GC) is required, which increases the write amplification of ZNS storage and degrades I/O performance. We introduce a lifetime-leveling compaction (LL-compaction) tailored for ZNS SSD, which can alleviate space amplification without GC by making the sorted string tables in a zone have similar lifetimes. In our experiments using LevelDB, the LL-compaction achieved 1.7x better performance by removing GCs. Jeeyoon Jung, Dongkun Shin |
HotStorage | 2 |
| 2022 | When F2FS meets address remappingabstractWhile gaining popularity in mobile devices, F2FS, a flash-friendly variation of log-structured file system, reveals three drawbacks: segment cleaning overhead, metadata update overhead, and file fragmentation, which becomes conspicuous under random update workloads. This paper suggests for the first time to leverage the address-remap technique in flash storage to remedy such pitfalls in F2FS. Our approach can, while preserving the benefit of log-structured writes, achieve the eventual effect of in-place update, completely preventing three drawbacks of F2FS. It can thus significantly outperform ext4 as well as vanilla F2FS under random update workloads. Armed with another write mode, F2FS will become competitive for a wider range of applications. Yongmyung Lee, Jonggyu Park, Hyunho Gwak, Dongkun Shin, Young Ik Eom, Sang-Won Lee 0001 |
HotStorage | 5 |
| 2022 | Remap-Based Inter-Partition Copy for Arrayed Solid-State DrivesabstractThe internal copy (IC) is a simple, yet powerful in-storage processing function, which changes the locations of data blocks without invoking any data transfer between the host and storage. Owing to the out-of-place update constraint of flash memory, solid-state disks (SSDs) employ a flash translation layer (FTL) to manage the logical-to-physical address translation. By leveraging the address indirection feature of SSDs, the IC can be processed only by remapping flash pages to new logical addresses without flash read/write operations. In the existing studies on remap-based IC, SSDs were assumed to have only a single FTL instance. However, recent large-capacity SSDs adopt an arrayed architecture including multiple FTL controllers, where each controller runs an FTL instance to manage its own partitioned address space. For the arrayed SSDs, inter-partition copy requests cannot be handled by address remapping because each partition is managed by a different FTL instance. In this study, we propose an inter-partition remap technique for IC-enabled arrayed SSDs. Additionally, we present the block allocation technique to minimize the number of inter-partition copy requests. Our proposed IC techniques were implemented on an actual arrayed SSD, and showed significant performance improvements compared to the previous remap techniques in several use cases. Kyuhwa Han, Dongkun Shin |
IEEE Trans. Computers | 2 |
| 2021 | ZNS+: Advanced Zoned Namespace Interface for Supporting In-Storage Zone Compaction
Kyuhwa Han, Hyunho Gwak, Dongkun Shin, Jooyoung Hwang |
OSDI | 3 |
| 2020 | Flexible Group-Level Pruning of Deep Neural Networks for On-Device Machine LearningabstractNetwork pruning is a promising compression technique to reduce computation and memory access cost of deep neural networks. Pruning techniques are classified into two types: fine-grained pruning and coarse-grained pruning. Fine-grained pruning eliminates individual connections if they are insignificant and thus usually generates irregular networks. Therefore, it is hard to reduce model execution time. Coarse-grained pruning such as filter-level and channel-level techniques can make hardware-friendly networks. However, it can suffer from low accuracy. In this paper, we focus on the group-level pruning method to accelerate deep neural networks on mobile GPUs, where several adjacent weights are pruned in a group to mitigate the irregularity of pruned networks while providing high accuracy. Although several group-level pruning techniques have been proposed, the previous techniques select weight groups to be pruned at group-size-aligned locations. In this paper, we propose a more flexible approach, called unaligned group-level pruning, to improve the accuracy of the compressed model. We can find the optimal solution of the unaligned group selection problem with dynamic programming. Our technique also generates balanced sparse networks to get load balance at parallel computing units. Experiments demonstrate that the 2D unaligned group-level pruning shows 3.12% a lower error rate at ResNet-20 network on CIFAR-10 compared to the previous 2D aligned group-level pruning under 95% of sparsity. Kwangbae Lee, Hoseung Kim, Hayun Lee, Dongkun Shin |
DATE | 4 |
| 2020 | Reinforcement Learning-Based SLC Cache Technique for Enhancing SSD Write Performance
Sangjin Yoo, Dongkun Shin |
HotStorage | 2 |
| 2020 | Command queue-aware host I/O stack for mobile flash storage
Kyuhwa Han, Dongkun Shin |
J. Syst. Archit. | 2 |
| 2020 | WAL-SSD: Address Remapping-Based Write-Ahead-Logging Solid-State DisksabstractRecent advances in flash memory technology have reduced the cost-per-bit of flash storage devices such as solid-state drives (SSDs), thereby enabling the development of large-capacity SSDs for enterprise-scale storage. However, two major concerns arise in designing SSDs. First, the size of the address mapping table is increasing in proportion to the capacity of the SSD. The SSD-internal firmware, called flash translation layer (FTL), must maintain the address mapping table in the internal DRAM. Although the previously proposed demand map loading technique uses a small size of cached map table, the technique aggravates poor random performance. Second, there are many redundant writes in storage workloads, which have an adverse effect on the performance and lifetime of the SSD. For example, many transaction-supporting applications use the write-ahead-log (WAL) scheme, which writes the same data twice. To resolve these problems, we propose a novel transaction-supporting SSD, called WAL-SSD, which logs transaction data at the internally-managed WAL area and relocates the data atomically via the FTL-level remap operation at the transaction checkpointing. It can also be used to transform random write requests to sequential requests. We implemented a prototype of WAL-SSD with a real SSD device. Experiments demonstrate the performance improvement by WAL-SSD with three use cases: remap-journaling, atomic multi-block update, and random write logging. Kyuhwa Han, Hyukjoong Kim, Dongkun Shin |
IEEE Trans. Computers | 3 |
| 2018 | StreamFTL: Stream-level address translation scheme for memory constrained flash storageabstractAlthough much research efforts have been devoted to reducing the size of address mapping table which consumes DRAM space in solid state drives (SSDs), most SSDs still use page-level mapping for high performance in their firmware called flash translation layer (FTL). In this paper, we propose a novel FTL scheme, called StreamFTL. In order to reduce the size of the mapping table in SSDs, StreamFTL maintains a mapping entry for each stream, which consists of several logical pages written at contiguous physical pages. Unlike extent, which is used by previous FTL schemes, the logical pages in a stream do not need to be contiguous. We show that StreamFTL can reduce the size of the mapping table by up to 90% compared to page-level mapping scheme. Hyukjoong Kim, Kyuhwa Han, Dongkun Shin |
DATE | 3 |
| 2017 | SHRD: Improving Spatial Locality in Flash Storage Accesses by Sequentializing in Host and Randomizing in Device
Hyukjoong Kim, Dongkun Shin, Yunho Jeong, Kyung Ho Kim |
FAST | 2 |
| 2017 | iJournaling: Fine-Grained Journaling for Improving the Latency of Fsync System Call
Daejun Park 0002, Dongkun Shin |
USENIX ATC | 2 |
| 2017 | Priority-driven spatial resource sharing scheduling for embedded graphics processing units
Yunji Kang, Woo Hyun Joo, Sungkil Lee 0002, Dongkun Shin |
J. Syst. Archit. | 4 |
| 2017 | Reinforcement Learning-Assisted Garbage Collection to Mitigate Long-Tail Latency in SSDabstractNAND flash memory is widely used in various systems, ranging from real-time embedded systems to enterprise server systems. Because the flash memory has erase-before-write characteristics, we need flash-memory management methods, i.e., address translation and garbage collection. In particular, garbage collection (GC) incurs long-tail latency, e.g., 100 times higher latency than the average latency at the 99 th percentile. Thus, real-time and quality-critical systems fail to meet the given requirements such as deadline and QoS constraints. In this study, we propose a novel method of GC based on reinforcement learning. The objective is to reduce the long-tail latency by exploiting the idle time in the storage system. To improve the efficiency of the reinforcement learning-assisted GC scheme, we present new optimization methods that exploit fine-grained GC to further reduce the long-tail latency. The experimental results with real workloads show that our technique significantly reduces the long-tail latency by 29--36% at the 99.99 th percentile compared to state-of-the-art schemes. Won-Kyung Kang, Dongkun Shin, Sungjoo Yoo |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | Performance and Resource Analysis on the JavaScript Runtime for IoT Devices
Dongig Sin, Dongkun Shin |
ICCSA (1) | 2 |
| 2016 | Exploiting Sequential and Temporal Localities to Improve Performance of NAND Flash-Based SSDsabstractNAND flash-based Solid-State Drives (SSDs) are becoming a viable alternative as a secondary storage solution for many computing systems. Since the physical characteristics of NAND flash memory are different from conventional Hard-Disk Drives (HDDs), flash-based SSDs usually employ an intermediate software layer, called a Flash Translation Layer (FTL). The FTL runs several firmware algorithms for logical-to-physical mapping, I/O interleaving, garbage collection, wear-leveling, and so on. These FTL algorithms not only have a great effect on storage performance and lifetime, but also determine hardware cost and data integrity. In general, a hybrid FTL scheme has been widely used in mobile devices because it exhibits high performance and high data integrity at a low hardware cost. Recently, a demand-based FTL based on page-level mapping has been rapidly adopted in high-performance SSDs. The demand-based FTL more effectively exploits the device-level parallelism than the hybrid FTL and requires a small amount of memory by keeping only popular mapping entries in DRAM. Because of this caching mechanism, however, the demand-based FTL is not robust enough for power failures and requires extra reads to fetch missing mapping entries from NAND flash. In this article, we propose a new flash translation layer called LAST++. The proposed LAST++ scheme is based on the hybrid FTL, thus it has the inherent benefits of the hybrid FTL, including low resource requirements, strong robustness for power failures, and high read performance. By effectively exploiting the locality of I/O references, LAST++ increases device-level parallelism and reduces garbage collection overheads. This leads to a great improvement of I/O performance and makes it possible to overcome the limitations of the hybrid FTL. Our experimental results show that LAST++ outperforms the demand-based FTL by 27% for writes and 7% for reads, on average, while offering higher robustness against sudden power failures. LAST++ also improves write performance by 39%, on average, over the existing hybrid FTL. Sungjin Lee 0001, Dongkun Shin, Young-Jin Kim 0002, Jihong Kim 0001 |
ACM Trans. Storage | 2 |
| 2014 | Adaptive Paired Page Prebackup Scheme for MLC NAND Flash MemoryabstractMultilevel cell (MLC) NAND flash memory is more cost effective compared with single-level cell NAND flash memory as it can store two or more bits in a memory cell. However, in MLC flash memory, a programming operation can corrupt the paired page under abnormal termination. In order to solve the paired page problem, a backup scheme is generally used, which inevitably causes performance degradation and shortens the lifespan of flash memory. In this paper, we propose a more efficient paired page prebackup scheme for MLC flash memory. It adaptively exploits interleaving, copyback operations, and parity data to reduce the prebackup overhead. In experiments, the proposed scheme reduced the backup overhead by up to 78%. Jaeil Lee, Dongkun Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Semantic-Aware Hot Data Selection Policy for Flash File System in Android-Based SmartphonesabstractFlash memory has different characteristics from traditional hard disk drives. Therefore, the traditional file systems such as EXT4 are not well-optimized for flash memory storage. Recently, a flash memory-aware file system, called F2FS, is announced, which is based on the log-structured file system considering the poor random write performance of flash memory. Although F2FS uses a heuristic for separating hot and cold data, the heuristic is not aware of file type. In this paper, we observe the lifetime of different types of files in Android platform and propose a semantic-aware hot data selection policy for F2FS. Experimental results show that the proposed technique reduces the garbage collection cost by up to 31%. Dongsoo Choi, Dongkun Shin |
ICPADS | 2 |
| 2013 | BAGC: Buffer-Aware Garbage Collection for Flash-Based Storage SystemsabstractNAND flash-based storage device is becoming a viable storage solution for mobile and desktop systems. Because of the erase-before-write nature, flash-based storage devices require garbage collection that causes significant performance degradation, incurring a large number of page migrations and block erasures. To improve I/O performance, therefore, it is important to develop an efficient garbage collection algorithm. In this paper, we propose a novel garbage collection technique, called buffer-aware garbage collection (BAGC), for flash-based storage devices. The BAGC improves the efficiency of two main steps of garbage collection, a block merge step and a victim block selection step, by taking account of the contents of a buffer cache, which is typically used to enhance I/O performance. The buffer-aware block merge (BABM) scheme eliminates unnecessary page migrations by evicting dirty data from a buffer cache during a block merge step. The buffer-aware victim block selection (BAVBS) scheme, on the other hand, selects a victim block so that the benefit of the buffer-aware block merge is maximized. Our experimental results show that BAGC improves I/O performance by up to 43 percent over existing buffer-unaware schemes for various benchmarks. Sungjin Lee 0001, Dongkun Shin, Jihong Kim 0001 |
IEEE Trans. Computers | 2 |
| 2011 | Flash-Aware RAID Techniques for Dependable and High-Performance Flash Memory SSDabstractSolid-state disks (SSDs), which are composed of multiple NAND flash chips, are replacing hard disk drives (HDDs) in the mass storage market. The performances of SSDs are increasing due to the exploitation of parallel I/O architectures. However, reliability remains as a critical issue when designing a large-scale flash storage. For both high performance and reliability, Redundant Arrays of Inexpensive Disks (RAID) storage architecture is essential to flash memory SSD. However, the parity handling overhead for reliable storage is significant. We propose a novel RAID technique for flash memory SSD for reducing the parity updating cost. To reduce the number of write operations for the parity updates, the proposed scheme delays the parity update which must accompany each data write in the original RAID technique. In addition, by exploiting the characteristics of flash memory, the proposed scheme uses the partial parity technique to reduce the number of read operations required to calculate a parity. We evaluated the performance improvements using a RAID-5 SSD simulator. The proposed techniques improved the performance of the RAID-5 SSD by 47 percent and 38 percent on average in comparison to the original RAID-5 technique and the previous delayed parity updating technique, respectively. Soojun Im, Dongkun Shin |
IEEE Trans. Computers | 2 |
| 2010 | Delayed partial parity scheme for reliable and high-performance flash memory SSDabstractThe I/O performances of flash memory solid-state disks (SSDs) are increasing by exploiting parallel I/O architectures. However, the reliability problem is a critical issue in building a large-scale flash storage. We propose a novel Redundant Arrays of Inexpensive Disks (RAID) architecture which uses the delayed parity update and partial parity caching techniques for reliable and high-performance flash memory SSDs. The proposed techniques improve the performance of the RAID-5 SSD by 38% and 30% on average in comparison to the original RAID-5 technique and the previous delayed parity update technique, respectively. Soojun Im, Dongkun Shin |
MSST | 2 |
| 2010 | ComboFTL: Improving performance and lifespan of MLC flash memory using SLC flash buffer
Soojun Im, Dongkun Shin |
J. Syst. Archit. | 2 |
| 2010 | Buffer flush and address mapping scheme for flash memory solid-state disk
Hyunchul Park 0002, Dongkun Shin |
J. Syst. Archit. | 2 |
| 2009 | KAST: K-associative sector translation for NAND flash memory in real-time systemsabstractFlash memory is a good candidate for the storage device in real-time systems due to its non-fluctuating performance, low power consumption and high shock resistance. However, the garbage collection for invalid pages in flash memory can invoke a long blocking time. Moreover, the worst-case blocking time is significantly long compared to the best-case blocking time under the current flash management techniques. In this paper, we propose a novel Flash Translation Layer (FTL), called KAST, where user can configure the maximum log block associativity to control the worst-case blocking time. Performance evaluation using simulations shows that the overall performance of KAST is better than the current FTL schemes as well as KAST guarantees the longest block time is shorter than the specified value. Hyun-jin Cho, Dongkun Shin, Young Ik Eom |
DATE | 2 |
| 2007 | Optimizing Intratask Voltage Scheduling Using Profile and Data-Flow InformationabstractIntratask dynamic-voltage scheduling (IntraDVS), which adjusts the supply voltage within an individual-task boundary, has been introduced as an effective technique for developing low-power single-task applications or low-power multitask applications, where a small number of tasks are dominant in total execution time. The original IntraDVS technique used the remaining worst case execution cycles, and the control-flow information to identify the voltage-scaling points (VSPs) of a program. In this paper, two kinds of improvement techniques enhancing the energy performance of the IntraDVS are proposed. One is to use profile information to optimize the voltage schedule for the remaining average-case execution path (RAEP-IntraDVS). The other is to use data-flow information to optimize the locations of VSPs [look-ahead IntraDVS (LaIntraDVS)]. The experimental results show that the RAEP-IntraDVS can reduce the energy consumption by 20% on average and the LaIntraDVS can reduce the energy consumption by 40%-45% compared with the original IntraDVS Dongkun Shin, Jihong Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Dynamic voltage scaling of mixed task sets in priority-driven systemsabstractThis paper describes dynamic voltage scaling (DVS) algorithms for real-time systems with both periodic and aperiodic tasks. Although many DVS algorithms have been developed for real-time systems with periodic tasks, none of them can be used for a system with both periodic and aperiodic tasks because of the arbitrary temporal behaviors of aperiodic tasks. This paper proposes off-line and on-line DVS algorithms that are based on existing DVS algorithms. The proposed algorithms utilize the execution behaviors of scheduling servers for aperiodic tasks. Since there is a tradeoff between the energy consumption and the response time of aperiodic tasks, the proposed algorithms focus on bounding the response time degradation of aperiodic tasks although they delay the response time by stretching the task execution to get high energy savings in mixed task sets. Experimental results show that the proposed algorithms reduce the energy consumption by 48% and 35% over the non-DVS scheme under rate monotonic (RM) scheduling and earliest deadline first (EDF) scheduling, respectively. Dongkun Shin, Jihong Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Optimizing intra-task voltage scheduling using data flow analysisabstractIntra-task voltage scheduling (IntraDVS), which adjusts the supply voltage within an individual task boundary, is an effective technique for developing low-power applications. In IntraDVS, slack times are estimated by analyzing program's control flow information. In this paper, we propose an optimization technique for IntraDVS using data flow information. The proposed algorithm improves the energy efficiency by moving the voltage scaling points to earlier instructions based on the analysis results of program's data flow. The experimental results using an MPEG-4 encoder program show that the proposed algorithm reduces the energy consumption by 40-45% over the original IntraDVS algorithm. Dongkun Shin, Jihong Kim 0001 |
ASP-DAC | 1 |
| 2005 | Intra-task voltage scheduling on DVS-enabled hard real-time systemsabstractThis paper proposes a novel intra-task dynamic voltage scheduling (IntraDVS) framework for low-energy hard real-time applications. Based on a static timing analysis technique, the proposed approach controls the supply voltage within an individual task boundary. By fully exploiting all the slack times, a scheduled program by the proposed technique always completes its execution near the deadline, thus achieving a high energy reduction ratio. The problem formulation of IntraDVS is first presented and two heuristics are proposed: one based on worst-case execution information and the other on average-case execution information. In order to validate the effectiveness of the proposed heuristics, a software tool that automatically converts a DVS-unaware program into an equivalent low-energy program was built. In an experiment on a DVS-enabled system, the low-energy version of a Moving Pictures Expert Group (MPEG)-4 encoder/decoder consumed only 35%-51% of the energy consumption of the original program running on a fixed-voltage system with a power-down mode. The energy efficiency of the IntraDVS algorithms was also compared with that of task-level voltage scheduling algorithms. The experimental results show that the IntraDVS algorithm can be useful in multitask environments as well. Dongkun Shin, Jihong Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | Dynamic voltage scaling of periodic and aperiodic tasks in priority-driven systems
Dongkun Shin, Jihong Kim 0001 |
ASP-DAC | 1 |
| 2004 | Power-Aware Scheduling of Mixed Task Sets in Priority-Driven Systems
Dongkun Shin, Jihong Kim 0001 |
EUC | 1 |
| 2003 | Power-aware scheduling of conditional task graphs in real-time multiprocessor systemsabstractWe propose a novel power-aware task scheduling algorithm for DVS-enabled real-time multiprocessor systems. Unlike the existing algorithms, the proposed DVS algorithm can handle conditional task graphs (CTGs) which model more complex precedence constraints. We first propose a condition-unaware task scheduling algorithm integrating the task ordering algorithm for CTGs and the task stretching algorithm for unconditional task graphs. We then describe a condition-aware task scheduling algorithm which assigns to each task the start time and the clock speed, taking account of the condition matching and task execution profiles. Experimental results show that the proposed condition-aware task scheduling algorithm can reduce the energy consumption by 50% on average over the non-DVS task scheduling algorithm. Dongkun Shin, Jihong Kim 0001 |
ISLPED | 1 |
| 2001 | Low-Energy Intra-Task Voltage Scheduling Using Static Timing AnalysisabstractWe propose an intra-task voltage scheduling algorithm for low-energy hard real-time applications. Based on a static timing analysis technique, the proposed algorithm controls the supply voltage within an individual task boundary. By fully exploiting all the slack times, a scheduled program by the proposed algorithm always complete its execution near the deadline, thus achieving a high energy reduction ratio. In order to validate the effectiveness of the proposed algorithm, we built a software tool that automatically converts a DVS-unaware program into an equivalent low-energy program. Experimental results show that the low-energy version of an MPEG-4 encoder/decoder (converted by the software tool) consumes less than 7$\sim$25% of the original program running on a fixed-voltage system with a power-down mode. Dongkun Shin, Jihong Kim 0001, Seongsoo Lee |
DAC | 1 |
| 2001 | An operation rearrangement technique for power optimization in VLIM instruction fetchabstractIn VLIW machines where a single instruction contains multiple operations, the power consumption during instruction fetches varies significantly depending on how the operations are arranged within the instruction. In this paper we describe a post-pass operation rearrangement method that reduces the power consumption from the instruction-fetch datapath. The proposed method modifies operation placement orders within VLIW instructions so that the switching activity between successive instruction fetches is minimized. Our experiment shows that the switching activity can be reduced by 34% on average for benchmark programs. Dongkun Shin, Jihong Kim 0001, Naehyuck Chang |
DATE | 1 |
| 2001 | A profile-based energy-efficient intra-task voltage scheduling algorithm for real-time applicationsabstractArticle Share on A profile-based energy-efficient intra-task voltage scheduling algorithm for real-time applications Authors: Dongkun Shin School of Computer Science and Engineering, Seoul National University School of Computer Science and Engineering, Seoul National UniversityView Profile , Jihong Kim School of Computer Science and Engineering, Seoul National University School of Computer Science and Engineering, Seoul National UniversityView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 271–274https://doi.org/10.1145/383082.383162Online:06 August 2001Publication History 24citation267DownloadsMetricsTotal Citations24Total Downloads267Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Dongkun Shin, Jihong Kim 0001 |
ISLPED | 1 |