Kyoungwoo Lee

dblp:20/1203 · DBLP profile ↗
← Back
40ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0001-5082-3775ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8Software engineering, systems software and programming languages · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FluidTrack: Investigating Child-Parent Collaborative Tracking for Pediatric Voiding Dysfunction Management
Junhyung Moon 0001, Sukhyun Lee, Juhee Go, Han Mo Ku, Yeohyun Jung, Seonyeong Hwang, Bongshin Lee, Yong Seung Lee, Hyun-Kyung Lee, Kyoungwoo Lee, Eun Kyoung Choe
CHI11
2025 ProGIP: Protecting Gradient-based Input Perturbation Approaches for OOD Detection From Soft Errors
abstract
Undetected out-of-distribution (OOD) inputs pose a significant threat to the reliability of deep learning models, as they may lead to unexpected behaviors during inference. Several studies have proposed effective OOD input detection methods. However, soft errors—another significant threat to reliability—can impact both the classification results of neural network models and the ID/OOD detections of OOD detection methods. To provide a resilient OOD detection solution against soft errors, we analyze the effect of soft errors on neural network models with gradient-based input perturbation (GIP) approaches, which are representative methods for OOD detection. Building on our analysis, we propose ProGIP, which incorporates two software-level range-based fault detectors to protect all execution phases of GIP approaches, including two forward passes and one backward pass. Because it is purely software‑based and adds just two scalar comparisons, ProGIP is readily deployable even on resource‑constrained embedded platforms. Our ProGIP solution enables GIP approaches to distinguish between ID, OOD, and fault-affected inferences, detecting 97.7% of critical faults with a negligible runtime overhead of only 0.84%. Experimental results with 2.4 million fault injections across various neural networks and OOD detection methods demonstrate ProGIP’s effectiveness in ensuring comprehensive reliability against non-malicious threats.
Sumedh Shridhar Joshi, Hwisoo So, Soyeong Park, Woobin Ko, Jinhyo Jung, Yohan Ko, Uiwon Hwang, Kyoungwoo Lee, Aviral Shrivastava
ACM Trans. Embed. Comput. Syst.8
2024 Passive Indoor Localization using IR-UWB Radars without Location Information
abstract
In the field of indoor localization research, privacy protection is a critical concern. These studies focus on developing technologies that can obtain location information without identifying individuals, not only in public spaces but also in privacy-sensitive areas. However, conventional radio frequency (RF) signal-based localization technologies have mostly been developed in laboratory environments and face limitations when applied in actual privacy-sensitive spaces. These technologies involve direct observation of human forms for tracking locations, raising privacy infringement concerns. To address this, we propose a new framework that does not require identifying an individual’s location at any point in the process, utilizing weakly supervised learning techniques. Employing IR-UWB radars, our framework detects the position of a child in a kindergarten restroom and has halved the mean distance error to 82.56 cm from 177.97 cm compared to the conventional geometric method of trilateration. It also reduces the annotation cost for location information.
Kyungphil Ryoo, Suchun Park, Wonjong Lee, Aejin Park, Kyoungwoo Lee
AVSS6
2024 Maintaining Sanity: Algorithm-based Comprehensive Fault Tolerance for CNNs
abstract
As the deployment of neural networks in safety-critical applications proliferates, it becomes imperative that they exhibit consistent and dependable performance amidst hardware malfunctions. Several protection schemes have been proposed to protect neural networks, but they suffer from huge overheads or insufficient fault coverage. This paper presents Maintaining Sanity, a comprehensive and efficient protection technique for CNNs. Maintaining Sanity extends the state-of-the-art algorithm-based fault tolerance for CNN, utilizing hamming codes and checkpointing to correct over 99.6% of critical faults with about 72% runtime overhead and minimal memory overhead compared to traditional triple modular redundancy (TMR) techniques.
Jinhyo Jung, Hwisoo So, Woobin Ko, Sumedh Shridhar Joshi, Yebon Kim, Yohan Ko, Aviral Shrivastava, Kyoungwoo Lee
DAC8
2024 Generic Soft Error Data and Control Flow Error Detection by Instruction Duplication
abstract
Transient faults or soft errors are considered one of the most daunting reliability challenges for microprocessors. Software solutions for soft error protection are attractive because they can provide flexible and effective error protection. For instance, nZDC (Didehban and Shrivastava 2016) state-of-the-art instruction duplication error protection scheme achieves a high degree of error detection by verifying the results of memory write operations and utilizes an effective control-flow checking mechanism. However, nZDC control-flow checking mechanism is architecture-dependent and suffers from some vulnerability holes. In this work, we address these issues by substituting nZDC control-flow checking mechanism with a general (ISA-independent) scheme and propose two transformations, coarse-grained scheduling, and asymmetric control-flow signatures, for hard-to-detect control flow errors. Fault injection experiments on different hardware components of synthesizable Verilog description of an OpenRISC-based microprocessor reveal that the proposed transformation shows 85% less silent data corruptions compared to nZDC. In addition, programs protected by the proposed scheme run on average around 37% faster than nZDC-protected programs.
Moslem Didehban, Hwisoo So, Prudhvi Gali, Aviral Shrivastava, Kyoungwoo Lee
IEEE Trans. Dependable Secur. Comput.5
2022 Root cause analysis of soft-error-induced failures from hardware and software perspectives
Jinhyo Jung, Yohan Ko, Hwisoo So, Kyoungwoo Lee, Aviral Shrivastava
J. Syst. Archit.4
2022 EXPERTISE: An Effective Software-level Redundant Multithreading Scheme against Hardware Faults
abstract
Error resilience is the primary design concern for safety- and mission-critical applications. Redundant MultiThreading (RMT) is one of the most promising soft and hard error resilience strategies because it does not require additional hardware modification. While the state-of-the-art software RMT scheme can achieve a high degree of error protection, our detailed investigation revealed that it suffers from performance overhead and insufficient fault coverage. This paper proposes EXPERTISE, a compiler-level RMT scheme that can detect the manifestation of hardware faults in all processor components. EXPERTISE transformation generates a checker-thread for the main execution thread. These redundant threads are executed simultaneously on two physically different cores of a multicore processor and perform almost the same computations. After each memory write operation is committed by the main-thread, the checker-thread loads back the written data from the memory and checks it against its own locally computed values. If they match, the execution continues. Otherwise, the error flag is raised. In order to evaluate the effectiveness of the proposed solution, we performed soft and hard error injection experiments on all the different hardware components of an ARM Cortex53-like μ-architecturally simulated microprocessor. Based on statistical fault injection campaigns, we have found that EXPERTISE provides 188× better fault coverage with 27% faster performance as compared to the state-of-the-art scheme.
Hwisoo So, Moslem Didehban, Yohan Ko, Aviral Shrivastava, Kyoungwoo Lee
ACM Trans. Archit. Code Optim.5
2021 CHITIN: A Comprehensive In-thread Instruction Replication Technique Against Transient Faults
abstract
Soft errors have become one of the most important design concerns due to drastic technology scaling. Software-based error detection techniques are attractive, due to their flexibility and hardware independence. However, our in-depth analysis reveals that the state-of-the-art techniques in the area cannot provide comprehensive fault coverage: i) their control-flow protection schemes provide incomplete redundancy of original instructions, ii) they do not protect function calls and returns, and iii) their instruction scheduling leaves many vulnerabilities open. In this paper, we propose CHITIN - code transformations for soft error resilience that adopts the load-back checking scheme of nZDC, an improved version of SWIFT-like control-flow protection scheme, and a contiguous scheduling of the original and redundant instructions to dramatically improve the vulnerability from soft errors that disrupt the control-flow. Our fault injection experiments demonstrate that CHITIN can reduce more than 89% of the silent data corruptions in the state-of-the-art solutions.
Hwisoo So, Moslem Didehban, Jinhyo Jung, Aviral Shrivastava, Kyoungwoo Lee
DATE5
2021 Comprehensive Failure Analysis against Soft Errors from Hardware and Software Perspectives
abstract
With technology scaling, reliability against soft errors is becoming an important design concern for modern embedded systems. To avoid the high cost and performance overheads of full protection techniques, several researches have therefore turned their focus to selective protection techniques. This increases the need to accurately identify the most vulnerable components or instructions in a system. In this paper, we analyze the vulnerability of a system from both the hardware and software perspectives through intensive fault injection trials. From the hardware perspective, we find the most vulnerable hardware components by calculating component-wise failure rates. From the software perspective, we identify the most vulnerable instructions by using the novel root cause instruction analysis. With our results, we show that it is possible to reduce the failure rate of a system to only 12.40% with minimal protection.
Yohan Ko, Hwisoo So, Jinhyo Jung, Kyoungwoo Lee, Aviral Shrivastava
ICCD4
2020 dMazeRunner: Optimizing Convolutions on Dataflow Accelerators
abstract
Convolution neural networks (CNNs) can be efficiently executed on dataflow accelerators. However, the vast space of executing convolutions on computational and memory resources of accelerators makes difficult for programmers to automatically and efficiently accelerate the convolutions and for architects to achieve efficient accelerator designs. We propose dMazeRunner framework, which allows users to optimize execution methods for accelerating convolution and matrix multiplication on a given architecture and to explore dataflow accelerator designs for efficiently executing CNN models. dMazeRunner determines efficient dataflows tailored for CNN layers and achieves efficient execution methods for CNN models within several seconds.
Shail Dave, Aviral Shrivastava, Sasikanth Avancha, Kyoungwoo Lee
ICASSP5
2019 Stress Recognition with State Classification Considering Temporal Variation of Stress Responses
abstract
To avoid and manage stress-related problems, several works have investigated how to recognize stress states of people by exploiting the machine learning. Most of works utilize the supervised learning which requires labels for data to be trained and tested. Thus, labels which represent stress state of each data such as stressed or not become definitely important. Conventional stress state classification methods assign labels to unit data per a certain experimental period by applying stressor appearances or self-evaluation scores. However, those methods ignore temporal variations in stress responses which are involuntarily triggered inside the body within a shorter period of time than an experimental period. Therefore, we propose a stress state classification method by considering not only user's subjective evaluations but also temporal changes of stress responses in short periods. For the demonstration, we label our experimental data of 40 subjects by using our proposed classification method and conventional ones, respectively. Then, we train and test stress recognition models with 6 machine learning algorithms and our implemented neural network ones based on the labeled data. Finally, binary stress recognition with our proposed classification method improves the recognition accuracy by up to 31.6% as compared to those with conventional techniques.
Junhyung Moon 0001, Juneil Lee, Dongmi Cheon, Munhee Lee, Kyoungwoo Lee
BIBM5
2019 A software-level Redundant MultiThreading for Soft/Hard Error Detection and Recovery
abstract
In this work, we investigate the potential of software-only RMT (Redundant MultiThreading) schemes for soft and hard error detection and recovery. We first implement and evaluate the error protection capability of basic software level triple redundant multithreading (STRMT) and analyze its vulnerability. Then we introduce FISHER (FlexIble Soft and Hard Error Resiliency) as a software RMT scheme which can achieve high degree of error resiliency and does not suffer from STRMT vulnerability holes. FISHER executes three threads and rather than having a centralized voting mechanism, it distributes and intertwines error detection and recovery operations between redundant threads. We performed 135,000 soft/hard error injection experiments on different hardware components of an ARM cortex53-like μ-architecturally simulated microprocessor. The results demonstrate that FISHER can reduce programs failure rate by around 42× and 26× compared to original and basic STRMT-protected versions of programs, respectively.
Hwisoo So, Moslem Didehban, Aviral Shrivastava, Kyoungwoo Lee
DATE4
2019 Static Function Prefetching for Efficient Code Management on Scratchpad Memory
abstract
As cache-based memory hierarchy is becoming a primary factor which limits the scalability and power efficiency of multi-core systems, scratchpad memory (SPM) has been studied as an alternative to cache. When SPM is used as an instruction memory, code management techniques are required to load code blocks on SPM using DMAs. In these techniques, code blocks are generally loaded on-demand to avoid loading incorrect block unlike cache (e.g. tag arrays), SPM does not have mechanism to detect and recover from faults. While on-demand loading guarantees no fault, it leads to considerable performance overhead since it serializes the execution of DMA and CPU. This paper presents a technique to insert prefetching instructions for function-level code management to enable overlapping execution between DMA engine and CPU. Our technique inserts DMA instructions statically at compile time and does not rely on any profiling or run-time resources. Our evaluation shows that static prefetching can reduce CPU idle time due to DMAs by 58.5% and achieves 14.7% of average performance improvement on the benchmarks showing high overhead due to DMAs.
Kyoungwoo Lee, Aviral Shrivastava
ICCD2
2019 dMazeRunner: Executing Perfectly Nested Loops on Dataflow Accelerators
abstract
Dataflow accelerators feature simplicity, programmability, and energy-efficiency and are visualized as a promising architecture for accelerating perfectly nested loops that dominate several important applications, including image and media processing and deep learning. Although numerous accelerator designs are being proposed, how to discover the most efficient way to execute the perfectly nested loop of an application onto computational and memory resources of a given dataflow accelerator ( execution method ) remains an essential and yet unsolved challenge. In this paper, we propose dMazeRunner -- to efficiently and accurately explore the vast space of the different ways to spatiotemporally execute a perfectly nested loop on dataflow accelerators (execution methods). The novelty of dMazeRunner framework is in: i) a holistic representation of the loop nests, that can succinctly capture the various execution methods, ii) accurate energy and performance models that explicitly capture the computation and communication patterns, data movement, and data buffering of the different execution methods, and iii) drastic pruning of the vast search space by discarding invalid solutions and the solutions that lead to the same cost. Our experiments on various convolution layers (perfectly nested loops) of popular deep learning applications demonstrate that the solutions discovered by dMazeRunner are on average 9.16× better in Energy-Delay-Product (EDP) and 5.83× better in execution time, as compared to prior approaches. With additional pruning heuristics, dMazeRunner reduces the search time from days to seconds with a mere 2.56% increase in EDP, as compared to the optimal solution.
Shail Dave, Sasikanth Avancha, Kyoungwoo Lee, Aviral Shrivastava
ACM Trans. Embed. Comput. Syst.4
2018 EXPERT: Effective and flexible error protection by redundant multithreading
abstract
Resiliency is a first-order design concern in modern microprocessor design. Compiler-level Redundant MultiThreading (RMT) schemes are promising because of their capability to detect the manifestation of hardware transient and permanent faults. In this work, we propose EXPERT, a compiler-level RMT scheme which can detect the manifestation of hardware faults in all hardware components. EXPERT transformation generates a checker thread for program main execution thread. These redundant threads execute simultaneously on two physically different cores of a multi-core processor. They perform mostly same computations, however, after each memory write operation committed by the main thread, the checker thread loads back the written data from the memory and checks it against its own locally computed values. If they match, execution continues. Otherwise, the error flag will be raised. Our processor-wide statistical transient and permanent fault injection experiments show that EXPERT error coverage is ~65x better than the state-of-the-art scheme.
Hwisoo So, Moslem Didehban, Yohan Ko, Aviral Shrivastava, Kyoungwoo Lee
DATE5
2017 Reducing code management overhead in software-managed multicores
abstract
Software-managed architectures, which use scratch-pad memories (SPMs), are a promising alternative to cached-based architectures for multicores. SPMs provide scalability but require explicit management. For example, to use an instruction SPM, explicit management code needs to be inserted around every call site to load functions to the SPM. such management code would check the state of the SPM and perform loading operations if necessary, which can cause considerable overhead at runtime. In this paper, we propose a compiler-based approach to reduce this overhead by identifying management code that can be removed or simplified. Our experiments with various benchmarks show that our approach reduces the execution time by 14% on average. In addition, compared to hardware caching, using our approach on an SPM-based architecture can reduce the execution times of the benchmarks by up to 15%.
Jian Cai 0001, Yooseong Kim, Aviral Shrivastava, Kyoungwoo Lee
DATE5
2017 Detecting periodic limb movements in sleep using motion sensor embedded wearable band
abstract
Monitoring periodic limb movements in sleep (PLMS) is important since it is correlated with people's quality of sleep and several other sleep disorders. The clinically approved method of examining PLMS is polysomnography (PSG) where the sleep of patients are examined in a laboratory with various sensors attached to their body. However, PSG is time-consuming and expensive for patients and the need for cost-effective and comfortable PLMS detection method has not been fulfilled. Accordingly, we propose a PLMS detection framework which utilizes a wearable motion-sensor-embedded band. In this work, we study the location to comfortably wear the device and accurately collect data on a foot. Further, to increase the accuracy of classifying PLMS, we propose the Motion Synchronized Windowing technique which segments the intervals where movements occur. Finally, we classify PLMS by using various machine learning algorithms typically used in the human activity recognition. Our proposed system achieves the accuracy of up to 96.92% in detecting PLMS. Therefore, our system is a cost-effective and convenient method of monitoring PLMS.
Saewon Kye, Junhyung Moon 0001, Kyoungwoo Lee, Seung-Chul Shin, Yong Seung Lee
SMC5
2017 Sleep stage classification for managing nocturnal enuresis through effective configuration
abstract
Various studies have examined the quality of one's sleep and further investigated several sleep disorders. In those investigations, accurately classifying one's sleep into the standardized sleep stages is important. The conventional classification heavily depends on the manual examination of each expert on one's physiological signals during the sleep. Therefore, various automatic classification models have been proposed using the machine learning. Although they properly classify the sleep stages on average, there have been few investigations to specifically improve the classification accuracy of certain stages. Accurate determination of several stages considerably correlating with a disorder gives us a more effective hint to conquer the disorder. Accordingly, we propose a configured classification model focusing on the interesting sleep stages related to a challenging sleep disorder, the nocturnal enuresis. We consider the deterministic physiological signals of the interesting stages when training the classifiers. Further, the proposed system utilizes recurrent neural network to effectively learn the sequential feature of the physiological data. Our proposed system achieves the classification accuracy by 83.6% over the data. In particular, technique presents up to 15.5% higher accuracy to differentiate interesting stages than the support vector machine approach for the nocturnal enuresis.
Junhyung Moon 0001, Saewon Kye, Kyoungwoo Lee, Yong Seung Lee, Seung-Chul Shin
SMC5
2017 Indoor localization in home environments using appearance frequency information
abstract
Recognizing the location of an individual in a home environment is crucial in order to enable various context-aware home applications such as elderly health monitoring and in home appliance automation. However, due to the limited number of dedicated Wi-Fi access points (APs), it is challenging to guarantee the reliable localization performance in a home environment by using the traditional Wi-Fi fingerprinting (WF) technique. In this paper, we propose a room-level localization system for the typical residential home environments which comprise of a living room, a kitchen, a bathroom, and a bedroom. Specifically, we make use of appearance frequency (AF) information of APs at each location in order to narrow down the number of candidate locations before performing the Wi-Fi Fingerprinting scheme. Our system improves the localization performance by up to 17.5 % (11.29 % on average) over that of the traditional WF-based approach which does not exploit AF information. We achieved the room-level positioning accuracy of 84.76% on the dataset of 6 home environments.
Jonghoon Shin, Hyunchoong Kim, Dayoung Lee, Yohan Ko, Kyoungwoo Lee, Seong-il Hahm, TaeJun Kwon
SMC5
2017 Protecting Caches from Soft Errors: A Microarchitect's Perspective
abstract
Soft error is one of the most important design concerns in modern embedded systems with aggressive technology scaling. Among various microarchitectural components in a processor, cache is the most susceptible component to soft errors. Error detection and correction codes are common protection techniques for cache memory due to their design simplicity. In order to design effective protection techniques for caches, it is important to quantitatively estimate the susceptibility of caches without and even with protections. At the architectural level, vulnerability is the metric to quantify the susceptibility of data in caches. However, existing tools and techniques calculate the vulnerability of data in caches through coarse-grained block-level estimation. Further, they ignore common cache protection techniques such as error detection and correction codes. In this article, we demonstrate that our word-level vulnerability estimation is accurate through intensive fault injection campaigns as compared to block-level one. Further, our extensive experiments over benchmark suites reveal several counter-intuitive and interesting results. Parity checking when performed over just reads provides reliable and power-efficient protection than that when performed over both reads and writes. On the other hand, checking error correcting codes only at reads alone can be vulnerable even for single-bit soft errors, while that at both reads and writes provides the perfect reliability.
Yohan Ko, Reiley Jeyapaul, Kyoungwoo Lee, Aviral Shrivastava
ACM Trans. Embed. Comput. Syst.4
2016 gemV: A validated toolset for the early exploration of system reliability
abstract
Decades of technology scaling has brought the threat of soft errors to modern embedded processors. Though several methods have been proposed to protect systems from soft errors, their effectiveness in ensuring error-free computing cannot be guaranteed; without accurate and quantitative estimation of system reliability. The metric vulnerability - which defines the likelihood of device failure by accurately evaluating the time it is exposed to soft errors - provides the most effective means to perform early design space explorations to estimate system reliability in the presence of transient soft errors. In this paper, we present gemV - the first accurate and comprehensive vulnerability estimation toolset, which is configurable and extendible to analyse future/novel architecture and microarchitecture designs. Some of the key features of gemV are: (1) all possible microarchitecture components that store bits, even temporarily, are modeled for their vulnerability in the gem5 cycle-accurate simulation platform, (2) its models have been validated (<3% correlation error with 90% statistical confidence) through exhaustive bit-level fault injection experiments, (3) the analytical models have incorporated microarchitecture-level masking effects like speculative executions, flushes, and etc. (4) the modular design of the vulnerability models make it easy to be extended and integrated when novel microarchitecture designs are explored. In addition to microarchitecture-level evaluation of system reliability, gemV provides a means to perform software-level design space explorations - that explore performance-vulnerability trade-offs of algorithm choices, compilers used, compiler optimization levels, etc. A system designer can further use gemV to explore the performance-vulnerability trade-offs of choosing different ISAs.
Karthik Tanikella, Yohan Ko, Reiley Jeyapaul, Kyoungwoo Lee, Aviral Shrivastava
ASAP4
2016 Splitting functions in code management on scratchpad memories
abstract
As the number of cores increases, cache-based memory hierarchy is becoming a major problem in terms of the scalability and energy consumption. Software-managed scratchpad memories (SPM) is a scalable alternative to caches, but the benefit comes at the cost of explicit management of data. For instance, an instruction SPM needs a code management techniques to load code blocks to the SPM. This paper presents a technique to split functions into smaller functions, to break away with the fundamental limitations of function-level code management. Our function-splitting technique is able to generate more efficient mappings by modifying the characteristics of functions to be more suitable for function-level code management. We propose two optimization policies to improve performance and reduce size respectively. The performance optimization policy improves performance by 16% on average, which can only be achieved by using 20% more SPM space if without function-splitting The size optimization policy can reduce the minimum SPM size requirement by 31% while increasing only 7% execution time.
Jian Cai 0001, Yooseong Kim, Kyoungwoo Lee, Aviral Shrivastava
ICCAD4
2016 EPOC aware energy expenditure estimation with machine learning
abstract
In 2014, 39 % of adults were overweight, and 13 % were obese. Clearly, knowing exact energy expenditure (EE) is important for sports training and weight control. Furthermore, excess post-exercise oxygen consumption (EPOC) must be included in the total EE. This paper presents a machine learning-based EE estimation approach with EPOC for aerobic exercise using a heart rate sensor. On a dataset acquired from 33 subjects, we apply machine learning algorithms using Weka machine learning toolkit. We could achieve 0.88 correlation and 0.23 kcal/min root mean square error (RMSE) with linear regression. The proposed model could be applied to various wearable devices such as a smartwatch.
Soljee Kim, Kyoungwoo Lee, Junga Lee, Justin Y. Jeon
SMC2
2016 Collaborative classification for daily activity recognition with a smartwatch
abstract
Research of daily activity recognition has been extensively conducted in the field of ubiquitous computing. However, previous daily activity recognition schemes are either obtrusive or inaccurate since they use just special-purpose devices. In this paper, we propose the collaborative classification for recognizing daily activities with a smartwatch. We exploit a single off-the-shelf smartwatch to distinguish 5 different daily activities such as eating, vacuuming, sleeping, showering, and TV watching. More precisely, we conduct experiments for collecting sensor data from accelerometer and acoustic sensor which are embedded in a smartwatch. However, the simple combination of the raw acceleration and acoustic data does not deliver accurate recognition accuracy. In order to achieve high accuracy, we propose a collaborative classification algorithm which integrates sensor data and ground-truth label for improving recognition accuracy by constructing a mapping table. We evaluate accuracies using single-sensor based approach, multi-sensor based approach, and our collaborative classification approach. The results from activity recognition for about 20 hours data collected by subjects show reliable accuracies for all 5 activities, and the overall accuracy of our collaborative approach is about 91.5%. Experimental results reveal that our approach improves the recall rate of each activity by up to 21.5% as compared to that of the simply combined multi-sensor based approach.
Hyunchoong Kim, Jonghoon Shin, Soohwan Kim, Yohan Ko, Kyoungwoo Lee, Hojung Cha, Seong-il Hahm, TaeJun Kwon
SMC5
2016 Multi-level cache vulnerability estimation: The first step to protect memory
abstract
Cache is one of the most susceptible microarchitectural components against soft errors since cache memory not only takes up the majority of chip area but also is frequently accessed by other microarchitectural components. Several protection techniques have been proposed in order to improve the cache reliability. These cache protections can significantly affect the overall performance of the entire processor. Thus, it is extremely important to quantify the reliability of cache memory with and without protections in order to choose appropriate protection techniques. In this paper, we model the vulnerability estimation with considering generally used protection techniques, such as parity and error correction code, on multi-level cache memory. In common processors, level 1 and 2 caches are protected by parity and error correction code, respectively, but our experimental results reveal several interesting results. First off, parity protection for level 1 instruction cache can be good way to decrease the vulnerability, but it is inefficient for level 1 data cache. In special cases, parity protection for level 1 data cache can worsen the reliability as compared to unprotected cache. Secondly, parity protection for level 2 cache can decrease the vulnerability almost by half with the comparable overheads. For some benchmarks, parity protection for level 2 cache can be as reliable as error correcting code with much less overheads.
Yohan Ko, Kyoungwoo Lee
SMC2
2016 Configurable privacy management for secure video surveillance in energy-constrained systems
abstract
Nowadays, lots of people concern about the privacy violation in the video surveillance systems widely utilized in order to protect people and their properties. Various protection techniques have been proposed to protect the privacy-sensitive information in the video surveillance and they properly provide the protection. Unfortunately, the video compression rate and the energy consumption have not been thoroughly investigated in the video surveillance system providing the privacy protection. However, the advance from the wired system to the wireless system which has the limited resources makes the video compression rate and the energy consumption as important as the degree of the perceptual protection and the recognition accuracy in the recovered video. In this paper, we propose the configurable privacy management technique in order to enhance both the compression rate and the energy efficiency of the privacy-protected video surveillance system. Satisfying the highest 10% degree of the perceptual protection in our experiments, the proposed method reduces the compressed bitstream size and the energy consumption by up to 66.0% and 27.3% each as compared to no protection, that is, the compression only while achieving the fine recognition accuracy by 82.1% in the recovered video.
Junhyung Moon 0001, Hwisoo So, Kyoungwoo Lee
SMC3
2016 Software-Based Selective Validation Techniques for Robust CGRAs Against Soft Errors
abstract
Coarse-Grained Reconfigurable Architectures (CGRAs) are drawing significant attention since they promise both performances with parallelism and flexibility with reconfiguration. Soft errors (or transient faults) are becoming a serious design concern in embedded systems including CGRAs since the soft error rate is increasing exponentially as technology is scaling. A recently proposed software-based technique with TMR (Triple Modular Redundancy) implemented on CGRAs incurs extreme overheads in terms of runtime and energy consumption mainly due to expensive voting mechanisms for the outputs from the triplication of every operation. In this article, we propose selective validation mechanisms for efficient modular redundancy techniques in the datapaths on CGRAs. Our techniques selectively validate the results at synchronous operations rather than every operation in order to reduce the expensive performance overhead from the validation mechanism. We also present an optimization technique to further improve the runtime and the energy consumption by minimizing synchronous operations where a validating mechanism needs to be applied. Our experimental results demonstrate that our selective validation-based TMR technique with our optimization on CGRAs can improve the runtime by 41.0% and the energy consumption by 26.2% on average over benchmarks as compared to the recently proposed software-based TMR technique with the full validation.
Yohan Ko, Jihoon Kang, Joonhyun Kim, Hwisoo So, Kyoungwoo Lee, Yunheung Paek
ACM Trans. Embed. Comput. Syst.7
2015 Guidelines to design parity protected write-back L1 data cache
abstract
Several decades of technology scaling has brought the challenge of soft errors to modern computing systems, and caches are most susceptible to soft errors. While it is straightforward to protect L2 and other lower level caches using error correcting coding (ECC), protecting the L1 data caches poses a challenge. Parity-based protection of L1 data cache is a more power-efficient alternative, however, some questions still linger -- How effective is parity protection for caches? How can we design a parity-based L1 data cache so as to maximize the protection achieved? The goal of this paper is to perform a quantitative evaluation of the protection afforded by various parity-protected cache design alternatives, and formulate guidelines for the design of power-efficient and reliable L1 data caches. Towards this goal, this paper develops an algorithm to accurately model the vulnerability of data in caches, in the presence of various configurations of parity protection, and validate it against extensive fault injection campaigns. We find that, (i) checking parity at reads only (and not at writes) provides 11% more protection with 30% lesser power overheads as compared to that at both reads and writes; and (ii) when implementing parity at the word-level granularity for 53% improved protection as compared to block-level parity implementation, the dirty-bits in the cache should also be implemented at the same granularity, otherwise, there is no improvement in protection. We find several popular commercial processors -- even the ones specifically designed for reliability -- not following these design guidelines, and resulting in sub-optimial designs.
Yohan Ko, Reiley Jeyapaul, Kyoungwoo Lee, Aviral Shrivastava
DAC4
2015 Non-obstructive room-level locating system in home environments using activity fingerprints from smartwatch
abstract
Many smart home applications, such as monitoring for the elderly and home automation, require location information for individual occupants. Several techniques have been proposed for tracking occupants in a home environment. However, the current techniques do not provide a seamless in-home locating system owing to the occupants' device-free movement and the lack of cost-effective infrastructure for home location tracking. In this paper, we propose a home occupant tracking system that uses a smartphone and an off-the-shelf smartwatch without additional infrastructure. In our system, activity fingerprints are automatically generated from the microphone and the inertial sensors of the smartwatch, and location information is periodically obtained from the smartphone. We designed a hidden Markov model using the relationship between home activities and the room's location. Extensive experiments showed that our system tracks the location of users with 87% accuracy, even when there is no manual training for activities.
Yungeun Kim, Daye Ahn, Rhan Ha, Kyoungwoo Lee, Hojung Cha
UbiComp5
2015 Performance improvement in ZigBee-based home networks with coexisting WLANs
Kun-Ho Hong, Kyoungwoo Lee
Pervasive Mob. Comput.3
2014 UnSync-CMP: Multicore CMP Architecture for Energy-Efficient Soft-Error Reliability
abstract
Reducing device dimensions, increasing transistor densities, and smaller timing windows, expose the vulnerability of processors to soft errors induced by charge carrying particles. Since these factors are only consequences of the inevitable advancement in processor technology, the industry has been forced to improve reliability on general purpose chip multiprocessors (CMPs). With the availability of increased hardware resources, redundancy-based techniques are the most promising methods to eradicate soft-error failures in CMP systems. In this work, we propose a novel customizable and redundant CMP architecture (UnSync) that utilizes hardware-based detection mechanisms (most of which are readily available in the processor), to reduce overheads during error-free executions. In the presence of errors (which are infrequent), the always forward execution enabled recovery mechanism provides for resilience in the system. The inherent nature of our architecture framework supports customization of the redundancy, and thereby provides means to achieve possible performance-reliability tradeoffs in many-core systems. We provide a redundancy-based soft-error resilient CMP architecture for both write-through and write-back cache configurations. We design a detailed RTL model of our UnSync architecture and perform hardware synthesis to compare the hardware (power/area) overheads incurred. We compare the same with those of the Reunion technique, a state-of-the-art redundant multicore architecture. We also perform cycle-accurate simulations over a wide range of SPEC2000, and MiBench benchmarks to evaluate the performance efficiency achieved over that of the Reunion architecture. Experimental results show that, our UnSync architecture reduces power consumption by 34.5 percent and improves performance by up to 20 percent with 13.3 percent less area overhead, when compared to the Reunion architecture for the same level of reliability achieved.
Reiley Jeyapaul, Fei Hong, Abhishek Rhisheekesan, Aviral Shrivastava, Kyoungwoo Lee
IEEE Trans. Parallel Distributed Syst.5
2013 Selective validations for efficient protections on Coarse-Grained Reconfigurable Architectures
abstract
Coarse-Grained Reconfigurable Architectures or CGRAs are drawing significant attention since they promise both performance with parallelism and flexibility with reconfiguration. Soft errors or transient faults are becoming a serious design concern in embedded systems including CGRAs since soft error rate is increasing exponentially as technology scaling. A recently proposed software-based technique with TMR (Triple Modular Redundancy) implemented on CGRAs incurs extreme performance overhead mainly due to expensive voting mechanisms for outputs from triplication of every operation. In this paper, we propose selective validation mechanisms for efficient modular redundancy techniques in the datapaths on CGRAs. Our techniques selectively validate results at synchronous operations rather than every operation in order to reduce the expensive performance overhead from the validation mechanism. Our experimental results demonstrate that our selective validation based TMR technique can improve the performance by 38.3% on average over benchmarks as compared to the recently proposed software-based TMR technique with the full validation.
Jihoon Kang, Yohan Ko, Hwisoo So, Kyoungwoo Lee, Yunheung Paek
ASAP6
2013 Dynamic code duplication with vulnerability awareness for soft error detection on VLIW architectures
abstract
Soft errors are becoming a critical concern in embedded system designs. Code duplication techniques have been proposed to increase the reliability in multi-issue embedded systems such as VLIW by exploiting empty slots for duplicated instructions. However, they increase code size, another important concern, and ignore vulnerability differences in instructions, causing unnecessary or inefficient protection when selecting instructions to be duplicated under constraints. In this article, we propose a compiler-assisted dynamic code duplication method to minimize the code size overhead, and present vulnerability-aware duplication algorithms to maximize the effectiveness of instruction duplication with least overheads for VLIW architecture. Our experimental results with SoarGen and Synopsys simulation environments demonstrate that our proposals can reduce the code size by up to 40% and detect more soft errors by up to 82% via fault injection experiments over benchmarks from DSPstone and Livermore Loops as compared to the previously proposed instruction duplication technique.
Yohan Ko, Kyoungwoo Lee, Jonghee M. Youn, Yunheung Paek
ACM Trans. Archit. Code Optim.3
2012 EAVE: Error-Aware Video Encoding Supporting Extended Energy/QoS Trade-offs for Mobile Embedded Systems
abstract
Energy/QoS provisioning is challenging for video applications over lossy wireless network with power-constrained mobile handheld devices. In this work, we exploit the inherent error tolerance of video data to generate a range of acceptable operating points by controlling the amount of errors in the system. In particular, we propose an error-aware video encoding technique, EAVE , that intentionally injects errors while ensuring acceptable QoS. The expanded trade-off space generated by EAVE allows system designers to comparatively evaluate different operating points with varying QoS and energy consumption by aggressively exploiting error-resilience attributes, and could potentially result in significant energy savings. The novelty of our approach resides in active exploitation of errors to vary the operating conditions for further optimization of system parameters. Moreover, we present the adaptivity of our approach by incorporating the feedback from the decoding side to achieve the QoS requirement under the dynamic network status. Our experiments show that EAVE can reduce the energy consumption for an encoding device by up to 37% for a video conferencing application over a wireless network without quality degradation, compared to a standard video encoding technique over test video streams. Further, our experimental results demonstrate that EAVE can expand the design space by 14 times with respect to energy consumption and by 13 times with respect to video quality (compared to a traditional approach without active error exploitation) on average, over test video streams.
Kyoungwoo Lee, Nikil Dutt, Nalini Venkatasubramanian
ACM Trans. Embed. Comput. Syst.1
2011 UnSync: A Soft Error Resilient Redundant Multicore Architecture
abstract
Reducing device dimensions, increasing transistor densities, and smaller timing windows, expose the vulnerability of processors to soft errors induced by charge carrying particles. Since these factors are only consequences of the inevitable advancement in processor technology, the industry has been forced to improve reliability on general purpose Chip Multiprocessors (CMPs). With the availability of increased hardware resources, redundancy based techniques are the most promising methods to eradicate soft error failures in CMP systems. In this work, we propose a novel redundant CMP architecture (UnSync) that utilizes hardware based detection mechanisms (most of which are readily available in the processor), to reduce overheads during error free executions. In the presence of errors (which are infrequent), the "always forward execution" enabled recovery mechanism provides for resilience in the system. We design a detailed RTL model of our UnSync architecture and perform hardware synthesis to compare the hardware (power/area) overheads incurred. We compare the same with those of the Reunion technique, a state-of-the-art redundant multi-core architecture. We also perform cycle-accurate simulations over a wide range of SPEC2000, and MiBench benchmarks to evaluate the performance efficiency achieved over that of the Reunion architecture. Experimental results show that, our UnSync architecture reduces power consumption by 34.5% and improves performance by up to 20% with 13.3% less area overhead, when compared to Reunion architecture for the same level of reliability achieved.
Reiley Jeyapaul, Fei Hong, Abhishek Rhisheekesan, Aviral Shrivastava, Kyoungwoo Lee
ICPP5
2010 Partitioning techniques for partially protected caches in resource-constrained embedded systems
abstract
Increasing exponentially with technology scaling, the soft error rate even in earth-bound embedded systems manufactured in deep subnanometer technology is projected to become a serious design consideration. Partially protected cache (PPC) is a promising microarchitectural feature to mitigate failures due to soft errors in power, performance, and cost sensitive embedded processors. A processor with PPC maintains two caches, one protected and the other unprotected, both at the same level of memory hierarchy. The intuition behind PPCs is that not all data in the application is equally prone to soft errors. By finding and mapping the data that is more prone to soft errors to the protected cache, and error-resilient data to the unprotected cache, failures induced by soft errors can be significantly reduced at a minimal power and performance penalty. Consequently, the effectiveness of PPCs critically hinges on the compiler's ability to partition application data into error-prone and error-resilient data. The effectiveness of PPCs has previously been demonstrated on multimedia applications—where an obvious partitioning of data exists, the multimedia data is inherently resilient to soft errors, and the rest of the data and the entire code is assumed to be error-prone. Since the amount of multimedia data is a quite significant component of the entire application data, this obvious partitioning is quite effective. However, no such obvious data and code partitioning exists for general applications. This severely restricts the applicability of PPCs to data caches and instruction caches in general. This article investigates vulnerability-based partitioning schemes that are applicable to applications in general and effectively reduce failures due to soft errors at minimal power and performance overheads. Our experimental results on an HP iPAQ-like processor enhanced with PPC architecture, running benchmarks from the MiBench suite demonstrate that our partitioning heuristic efficiently finds page partitions for data PPCs that can reduce the failure rate by 48% at only 2% performance and 7% energy overhead, and finds page partitions for instruction PPCs that reduce the failure rate by 50% at only 2% performance and 8% energy overhead, on average.
Kyoungwoo Lee, Aviral Shrivastava, Nikil Dutt, Nalini Venkatasubramanian
ACM Trans. Design Autom. Electr. Syst.1
2009 Partially Protected Caches to Reduce Failures Due to Soft Errors in Multimedia Applications
abstract
With advances in process technology, soft errors are becoming an increasingly critical design concern. Owing to their large area, high density, and low operating voltages, caches are worst hit by soft errors. Based on the observation that in multimedia applications, not all data require the same amount of protection from soft errors, we propose a partially protected cache (PPC) architecture, in which there are two caches, one protected and the other unprotected at the same level of memory hierarchy. We demonstrate that as compared to the existing unprotected cache architectures, PPC architectures can provide 47 times reduction in failure rate, at only 1% runtime and 3% power overheads. In addition, the failure rate reduction obtained by PPCs is very sensitive to the PPC cache configuration. Therefore, this observation provides an opportunity for further improvement of the solution by correctly parameterizing the PPC configurations. Consequently, we develop design space exploration (DSE) strategies to discover the best PPC configuration. Our DSE technique can reduce the exploration time by more than six times as compared to an exhaustive approach.
Kyoungwoo Lee, Aviral Shrivastava, Ilya Issenin, Nikil Dutt, Nalini Venkatasubramanian
IEEE Trans. Very Large Scale Integr. Syst.1
2008 Mitigating the impact of hardware defects on multimedia applications: a cross-layer approach
abstract
Increasing exponentially with each technology generation, hardware-induced soft errors pose a significant threat for the reliability of mobile multimedia devices. Since traditional hardware error protection techniques incur significant power and performance overheads, this paper proposes a cooperative cross-layer approach that exploits existing error control schemes at the application layer to mitigate the impact of hardware defects. Specifically, we propose error detection codes in hardware, drop and forward recovery in middleware, and error-resilient video encoding at the application level to effectively and efficiently combat soft errors with minimal overheads. Experimental evaluation on standard test video streams demonstrates that our cooperative error-aware method for video encoding improves performance by 60% and energy consumption by 58% with even better reliability at the cost of only 3% quality degradation on average, as compared to an error correction code based hardware protection technique. Combining intelligent schemes to select a recovery mechanism can guide system designers to trade off multiple constraints such as performance, power, reliability, and QoS.
Kyoungwoo Lee, Aviral Shrivastava, Minyoung Kim 0002, Nikil Dutt, Nalini Venkatasubramanian
ACM Multimedia1
2006 Mitigating soft error failures for multimedia applications by selective data protection
abstract
With advances in process technology, soft errors(SE)are becoming an increasingly critical design concern. Due to their large area and high density, caches are worst hit by soft errors. Although Error Correction Code based mechanisms protect the data in caches, they have high performance and power overheads. Since multimedia applications are increasingly being used in mission-critical embedded systems where both reliability and energy are a major concern, there is a de?nite need to improve reliability in embedded systems, without too much energy overhead. We observe that while a soft error in multimedia data may only result in a minor loss in QoS, a soft error in avariable that controls the execution ?ow of the program may be fatal. Consequently, we propose to partition the data space into failure critical and failure non-critical data, and provide a high-degree of soft error protection only to the failure critical data in Horizontally Partitioned Caches. Experimental results demonstrate that our selective data protection can achieve the failure rate close to that of a soft error protected cache system, while retaining the performance and energy consumption similar to those of a traditional cache system, with some degradation in QoS. For example, for conventional con?guration as in IntelXScale, our approach achieves the same failure rate, while improving performance by 28% and reducing energy consumption by 29%in comparison with a soft error protected cache.
Kyoungwoo Lee, Aviral Shrivastava, Ilya Issenin, Nikil Dutt, Nalini Venkatasubramanian
CASES1
2005 An Experimental Study on Energy Consumption of Video Encryption for Mobile Handheld Devices
abstract
Secure video communication on mobile handheld devices is challenging mainly due to (a) the significant computational needs of both video coding and encryption algorithms and (b) the limited battery capacity of handheld devices. In this paper, we evaluate several video encryption schemes from the perspective of energy consumption both analytically and experimentally. Specifically, we implement video encryption schemes on mobile handhelds to support a H.263 based secure video application, and extensively measure the energy consumption due to encoding and encryption for several classes of video clips. Contrary to popular belief, our experiments show that energy overhead of full video encryption is insignificant compared to the energy consumed for video encoding (between 2% and 4% of total energy cost) in most cases
Kyoungwoo Lee, Nikil Dutt, Nalini Venkatasubramanian
ICME1