EDBT 2026 Demo / reviewers in the wild / expert
Sunggu Lee
dblp:91/5410
· DBLP profile ↗
63ranked-venue papers
10as first author
4since 2021 · last 2024
0000-0003-3858-0779ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 51 · 8 first-author · 2 since 2021Software engineering, systems software and programming languages · 13Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 2Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Memory systems · 26% Energy-efficient computing · 20% Hardware reliability and fault tolerance · 15% | |
| Artificial intelligence
1 paper |
Image recognition and object detection · 62% Deep learning architectures and training · 38% |
Topics — the 30 heaviest of 44, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
image classification |
0.6 | 1 | 2022 | Convolutional Neural Networks With Discrete Cosine Transform Features · IEEE Trans. Computers 2022 |
Hardware reliability and fault tolerance › fault-tolerant design
error-tolerant design |
0.5 | 1 | 2021 | Layerwise Buffer Voltage Scaling for Energy-Efficient Convolutional Neural Network · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Integrated circuit design
low-power circuit design |
0.5 | 1 | 2021 | Layerwise Buffer Voltage Scaling for Energy-Efficient Convolutional Neural Network · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.5 | 1 | 2021 | Layerwise Buffer Voltage Scaling for Energy-Efficient Convolutional Neural Network · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Energy-efficient computing
voltage scaling |
0.5 | 1 | 2021 | Layerwise Buffer Voltage Scaling for Energy-Efficient Convolutional Neural Network · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Memory systems
cache |
0.3 | 2 | 2012 | A Multistep Tag Comparison Method for a Low-Power L2 Cache · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012 Matching cache access behavior and bit error pattern for high performance low Vcc L1 cache · DAC 2011 |
Memory systems
main memory |
0.3 | 2 | 2012 | Write performance improvement by hiding R drift latency in phase-change RAM · DAC 2012 Power management of hybrid DRAM/PRAM-based main memory · DAC 2011 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.2 | 1 | 2022 | Convolutional Neural Networks With Discrete Cosine Transform Features · IEEE Trans. Computers 2022 |
Machine learning › Deep learning architectures and training
feature fusion |
0.2 | 1 | 2022 | Convolutional Neural Networks With Discrete Cosine Transform Features · IEEE Trans. Computers 2022 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.1 | 1 | 2012 | A Multistep Tag Comparison Method for a Low-Power L2 Cache · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012 |
Memory systems › hybrid memory
hybrid main memory |
0.1 | 1 | 2012 | Hybrid DRAM/PRAM-based main memory for single-chip CPU/GPU · DAC 2012 |
Memory systems
non-volatile memory |
0.1 | 1 | 2012 | Write performance improvement by hiding R drift latency in phase-change RAM · DAC 2012 |
Memory systems › non-volatile memory
phase change memory |
0.1 | 1 | 2012 | Write performance improvement by hiding R drift latency in phase-change RAM · DAC 2012 |
Memory systems › cache › cache technology
fault-tolerant cache |
0.1 | 1 | 2011 | Matching cache access behavior and bit error pattern for high performance low Vcc L1 cache · DAC 2011 |
Energy-efficient computing › power management
memory power management |
0.1 | 1 | 2011 | Power management of hybrid DRAM/PRAM-based main memory · DAC 2011 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 1 | 2011 | Matching cache access behavior and bit error pattern for high performance low Vcc L1 cache · DAC 2011 |
Parallel and multicore computing
task scheduling |
0.1 | 2 | 2007 | Push-Pull: Deterministic Search-Based DAG Scheduling for Heterogeneous Cluster Systems · IEEE Trans. Parallel Distributed Syst. 2007 Processor Allocation and Task Scheduling of Matrix Chain Products on Parallel Systems · IEEE Trans. Parallel Distributed Syst. 2003 |
Parallel and multicore computing › task scheduling
DAG scheduling |
0.1 | 1 | 2007 | Push-Pull: Deterministic Search-Based DAG Scheduling for Heterogeneous Cluster Systems · IEEE Trans. Parallel Distributed Syst. 2007 |
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
heterogeneous cluster scheduling |
0.1 | 1 | 2007 | Push-Pull: Deterministic Search-Based DAG Scheduling for Heterogeneous Cluster Systems · IEEE Trans. Parallel Distributed Syst. 2007 |
Energy-efficient computing
power management |
0.0 | 1 | 2012 | A Multistep Tag Comparison Method for a Low-Power L2 Cache · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2012 |
Parallel and multicore computing › parallel algorithms
parallel algorithm design |
0.0 | 1 | 2003 | Processor Allocation and Task Scheduling of Matrix Chain Products on Parallel Systems · IEEE Trans. Parallel Distributed Syst. 2003 |
Parallel and multicore computing
parallel scheduling |
0.0 | 1 | 2003 | Processor Allocation and Task Scheduling of Matrix Chain Products on Parallel Systems · IEEE Trans. Parallel Distributed Syst. 2003 |
Parallel and multicore computing
processor allocation |
0.0 | 1 | 2003 | Processor Allocation and Task Scheduling of Matrix Chain Products on Parallel Systems · IEEE Trans. Parallel Distributed Syst. 2003 |
Distributed systems
fault tolerance |
0.0 | 3 | 1997 | Replicated Process Allocation for Load Distribution in Fault-Tolerant Multicomputers · IEEE Trans. Computers 1997 Optimal and Efficient Probabilistic Distributed Diagnosis Schemes · IEEE Trans. Computers 1993 On Probabilistic Diagnosis of Multiprocessor Systems Using Multiple Syndromes · IEEE Trans. Parallel Distributed Syst. 1994 |
Electronic design automation › hardware verification and test
fault diagnosis |
0.0 | 2 | 1994 | On Probabilistic Diagnosis of Multiprocessor Systems Using Multiple Syndromes · IEEE Trans. Parallel Distributed Syst. 1994 Optimal and Efficient Probabilistic Distributed Diagnosis Schemes · IEEE Trans. Computers 1993 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 2007 | Push-Pull: Deterministic Search-Based DAG Scheduling for Heterogeneous Cluster Systems · IEEE Trans. Parallel Distributed Syst. 2007 |
Distributed systems › replication
process replication |
0.0 | 1 | 1997 | Replicated Process Allocation for Load Distribution in Fault-Tolerant Multicomputers · IEEE Trans. Computers 1997 |
Interconnection networks and networks-on-chip › broadcasting
all-to-all broadcast |
0.0 | 1 | 1994 | Interleaved All-to-All Reliable Broadcast on Meshes and Hypercubes · IEEE Trans. Parallel Distributed Syst. 1994 |
Distributed systems › fault tolerance › fault-tolerant protocols
reliable broadcast |
0.0 | 1 | 1994 | Interleaved All-to-All Reliable Broadcast on Meshes and Hypercubes · IEEE Trans. Parallel Distributed Syst. 1994 |
Electronic design automation › physical design
routing |
0.0 | 1 | 1994 | Interleaved All-to-All Reliable Broadcast on Meshes and Hypercubes · IEEE Trans. Parallel Distributed Syst. 1994 |
Methods — techniques the papers use, named apart from their topics
discrete cosine transform · 0.6convolutional neural network · 0.6error-resilience analysis · 0.5error injection · 0.5write buffer management · 0.1hot data management · 0.1cache miss prediction · 0.1cache hit prediction · 0.1bloom filter · 0.1access scheduling · 0.1remapping · 0.1access behavior history · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Small-Footprint Convolutional Neural Network with Reduced Feature Map for Voice Activity DetectionabstractBy using Voice Activity Detection (VAD) as a preprocessing step, hardware-efficient implementations are possible for speech applications that need to run continuously in severely resource-constrained environments. For this purpose, we propose TinyVAD, which is a new convolutional neural network (CNN) model that executes extremely efficiently with a small memory footprint. TinyVAD uses an input pixel matrix partitioning method, termed patchify, to downscale the resolution of the input spectrogram. The hidden layers use a sequence of special convolutional structures with bypass links, referred to as CSPTiny layers. The proposed model is evaluated and compared with previous VAD methods using a diverse set of noisy environmental datasets. TinyVAD executes 3.13 times faster, utilizes only 12.5% as many multiplications, and requires only 13.0% as many parameters when compared to the previous state-of-the-art. Hwabyeong Chae, Sunggu Lee |
ICASSP | 2 |
| 2022 | Energy-Efficient Image Processing Using Binary Neural Networks with Hadamard Transform
Jaeyoon Park, Sunggu Lee |
ACCV (5) | 2 |
| 2022 | Convolutional Neural Networks With Discrete Cosine Transform FeaturesabstractThe Discrete Cosine Transform (DCT) exposes features of an image that are not evident in the original image's spatial domain. This brief contribution proposes a Convolutional Neural Network architecture that combines features from the spatial domainandthe DCT domain to improve image classification performance with negligible overhead. Sanghyeon Ju, Youngjoo Lee 0002, Sunggu Lee |
IEEE Trans. Computers | 3 |
| 2021 | Layerwise Buffer Voltage Scaling for Energy-Efficient Convolutional Neural NetworkabstractIn order to effectively reduce buffer energy consumption, which constitutes a significant part of the total energy consumption in a convolutional neural network (CNN), it is useful to apply different amounts of energy conservation effort to the different levels of a CNN as the buffer energy to total energy usage ratios can differ quite substantially across the layers of a CNN. This article proposes layerwise buffer voltage scaling as an effective technique for reducing buffer access energy. Error-resilience analysis, including interlayer effects, conducted during design-time is used to determine the specific buffer supply voltage to be used for each layer of a CNN. Then these layer-specific buffer supply voltages are used in the CNN for image classification inference. Error injection experiments with three different types of CNN architectures show that, with this technique, the buffer access energy and overall system energy can be reduced by up to 68.41% and 33.68%, respectively, without sacrificing image classification accuracy. Minho Ha, Younghoon Byun, Seungsik Moon, Youngjoo Lee 0002, Sunggu Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | Low-Complexity Dynamic Channel Scaling of Noise-Resilient CNN for Intelligent Edge DevicesabstractIn this paper, we present a novel channel scaling scheme for convolutional neural networks (CNNs), which can improve the recognition accuracy for the practical distorted images without increasing the network complexity. During the training phase, the proposed work first prepares multiple filters under the same CNN architecture by taking account of different noise models and strengths. We then newly introduce an FFT-based noise classifier, which determines the noise property in the received input image by calculating the partial sum of the frequency-domain values. Based on the detected noise class, we dynamically change the filters of each CNN layer to provide the dedicated recognition. Furthermore, we propose a channel scaling technique to reduce the number of active filter parameters if the input data is relatively clean. Experimental results show that the proposed dynamic channel scaling reduces the computational complexity as well as the energy consumption, still providing the acceptable accuracy for intelligent edge devices. Younghoon Byun, Minho Ha, Sunggu Lee, Youngjoo Lee 0002 |
DATE | 4 |
| 2019 | Similarity-Based LSTM Architecture for Energy-Efficient Edge-Level Speech RecognitionabstractTargeting the resource-limited edge devices, we present a novel processing architecture of long short-term memory (LSTM) networks for low-power speech recognition. The proposed scheme newly defines the similarity score between two inputs of adjacent LSTM cells, and then the processing mode of the current LSTM cell is dynamically determined to reduce the energy while providing the accurate recognition. If the similarity is high, more precisely, the current cell is disabled and the outputs are directly copied from the prior vectors, totally eliminating complex LSTM operations. To maximize the skipping ratio without degrading the accuracy, for the first time, we analyze the effects of skipping the consecutive cells and set the upper limit of the number of consecutive skips. When two adjacent inputs are weakly similar, in addition, we modify the concept of the previous delta-computing, which approximately activate the LSTM cell with low computational resolution, further reducing the energy consumption. Compared to the previous state-of-the-art solutions, as a result, the proposed LSTM architecture remarkably saves the energy consumed for the accurate speech recognition, which is suitable to the resource-limited embedded edges. Junseo Jo, Jaeha Kung 0001, Sunggu Lee, Youngjoo Lee 0002 |
ISLPED | 3 |
| 2016 | Iterative Localization of Network Nodes Using Absence of Distance Measurement InformationabstractThis paper considers the wireless sensor network iterative localization problem, which involves determining the 2D or 3D position of all nodes in a sequential manner. By utilizing information about the absence of distance measurements as well as all available distance measurements, it is possible to localize a significantly larger number of nodes than previous methods. When considering possible candidate locations for a node, the proposed method utilizes the fact that certain candidates are impossible because those candidates would require distance measurements that are absent (missing). A pseudocode solution is proposed, and simulation results are used to demonstrate the level of localizability achievable with this type of scheme. Sunggu Lee, Myunghoon Kang |
ISORC | 1 |
| 2016 | Memory Access Scheduling for a Smart TVabstractA smart TV system-on-chip (SoC) has very heavy computation and memory demands that must be met with low-cost components. As a result, there is potentially an extremely high utilization of the channel between the SoC and its memory chip. This paper presents the design of a new memory access scheduler customized for the type of memory traffic typically encountered with smart TVs. This includes special accumulated hard real-time graphics requirements, user response-sensitive soft real-time requirements, and the need to provide high memory throughput and priority-handling capabilities even under extremely heavy memory traffic conditions. The simulation results show that the proposed memory access scheduler is able to achieve up to 98% of the ideal upper bound memory throughput when faced with extremely heavy memory traffic-this is a significant improvement over previous schedulers. Novel future prediction and light-handed priority handling methods are used to achieve these results while satisfying the unique real-time requirements of smart TVs. Cheul-Hee Hahm, Sunggu Lee, Sungjoo Yoo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Improving Write Performance by Controlling Target Resistance Distributions in MLC PRAMabstractMulti-level cell (MLC) phase change RAM (PRAM) is expected to offer lower cost main memory than DRAM. However, poor write performance is one of the most critical problems for practical applications of MLC PRAM. In this article, we present two schemes to improve write performance by controlling the target resistance distribution of MLC PRAM cells. First, we propose multiple RESET/SET operations that relax the target resistance bands of intermediate logic levels with additional RESET/SET operations, which reduces the program time of intermediate logic levels, thereby improving write performance. Second, we propose a two-step write scheme consisting of lightweight write and idle-time completion write that exploits the fact that hot dirty data tend to be overwritten in a short time period and the MLC PRAM often has long idle times. Experimental results show that the multiple RESET/SET and two-step write schemes result in an average IPC improvement of 15.7% and 10.4%, respectively, on a hybrid DRAM/PRAM main memory subsystem. Furthermore, their integrated solution results in an average IPC improvement of 23.2% (up to 46.4%). Youngsik Kim, Sungjoo Yoo, Sunggu Lee |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2016 | Differential Write-Conscious Software Design on Phase-Change Memory: An SQLite Case StudyabstractPhase-change memory (PCM) has several benefits including low cost, non-volatility, byte-addressability, etc., and limitations such as write endurance. There have been several hardware approaches to exploit the benefits while minimizing the negative impact of limitations. Software approaches could give further improvements, when used together with hardware approaches, by taking advantage of write behavior present in the program, e.g., write behavior on dynamically allocated data, which is hardly captured by hardware approaches. This work proposes a software design methodology to reduce costly PCM writes. First, on top of existing hardware approach such as Flip-N-Write, we advocate exploiting the capability of PCM bit-level differential write in the software by judiciously reusing previously allocated memory resource. In order to avoid wear-out incurred by the reuse, we present software-based wear-leveling methods that distribute writes across PCM cells. In order to further reduce PCM writes, we propose identifying data, the loss of which does not affect the functionality of the underlying software, and then diverting write traffic for those data items to volatile memory. To evaluate the effectiveness of these methods, as a case study, we applied the proposed methods to the design of journaling in SQLite, which is an important database application commonly used in smartphones. For the experiments, we used an in-house PCM-based prototype board. Our experiments with four representative mobile applications show that the proposed design methods, which is applied on top of the hardware approach, Flip-N-Write, result in 75.2% further reduction in total bit updates in PCM, on average, without aggravating wear-out compared with the baseline of PCM-based journaling, which is based only on the hardware approach. Also, the proposed design methods result in 49.4% reduction in energy consumption and 52.3% reduction in runtime compared to a typical FIFO management of free resources. Sungkwang Lee, Taemin Lee, Hyunsun Park, Junwhan Ahn, Sungjoo Yoo, Youjip Won, Sunggu Lee |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2015 | Memory fast-forward: a low cost special function unit to enhance energy efficiency in GPU for big data processing
Eunhyeok Park, Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Sunggu Lee |
DATE | 5 |
| 2015 | A small non-volatile write buffer to reduce storage writes in smartphones
Mungyu Son, Sungkwang Lee, Kyungho Kim, Sungjoo Yoo, Sunggu Lee |
DATE | 5 |
| 2015 | Time slot assignment for convergecast in wireless sensor networks
Sunggu Lee, Sungjoo Yoo |
J. Parallel Distributed Comput. | 2 |
| 2015 | Dynamic Wear Leveling for Phase-Change Memories With Endurance VariationsabstractPhase change memory (PCM) has a write endurance problem. This problem is exacerbated due to endurance variations (EVs) when using advanced process technology (e.g., sub-20 nm), where PCM is expected to provide scaling benefits over dynamic random access memory (RAM). Wear leveling can solve this problem by dynamically changing the mapping from memory addresses to PCM physical addresses such that all PCM cells are evenly written, thereby extending the effective lifetime of such devices. PCM permits fine-grained writes, i.e., even bit level updates are allowed. To allow fine-grained wear leveling, this capability must be exploited. However, previous wear leveling approaches do not fully exploit fine-grained writes since fine-grained writes cause them to suffer from high data copy (called swap) overhead for address remapping, and/or high area and runtime overhead for the management of write frequency and address mapping information. This paper proposes a dynamic wear leveling method for PCMs that addresses all of these issues. The method: 1) uses bloom filters to enable low-cost write counters for fine-grained writes and 2) exploits the EV of PCM cells to avoid mapping hot data onto weak cells. To improve the effectiveness of the bloom filters, dynamic bloom filter management (write counts, hash functions, and write counter thresholds) and hot-cold address lists are used. The proposed method was evaluated using simulations and a hardware implementation. Using a small amount of PCM capacity overhead (0.3%), the proposed method extended the lifetime of a PCM device by 2.8-4.6 times over the existing methods when there were significant EVs. Joosung Yun, Sunggu Lee, Sungjoo Yoo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Coarse-grained Bubble Razor to exploit the potential of two-phase transparent latch designsabstractTiming margin to cover process variation is one of the most critical factors that limit the amount of supply voltage reduction thereby power consumption. To remove too conservative timing margin, Bubble Razor was introduced to dynamically detect and correct errors in two-phase transparent latch designs [13]. However, it does not fully exploit the potential of two-phase transparent latch design, e.g. time borrowing. Thus, especially at low supply voltage where the effect of process variation becomes significant, the existing Bubble Razor can suffer from significant overhead in performance and power consumption due to too frequent occurrence of bubble generations. We present a design methodology for coarse-grained Bubble Razor which exploits the time-borrowing characteristic of two-phase transparent latch design. By selectively inserting error checkpoints, i.e., shadow latches and error management logic, in the circuit, time borrowing can be applied between error checkpoints thereby avoiding bubbles which could occur in the existing Bubble Razor design with a checkpoint at every latch on the critical path. We present a methodology to choose the grain size (the number of stages between error checkpoints) based on 3-sigma delay distribution. We also verify the benefits of coarse-grained Bubble Razor with a real microprocessor, Core-A design [15] using 20nm Predictive Technology Model (PTM) [16]. The proposed methodology offers 62% improvement in performance (MIPS) and 49% less energy consumption (per instruction) at 0.6V operation (zero frequency margin) over the original Bubble Razor scheme. In addition, it gives 25% area reduction in core design. Hayoung Kim, Dongyoung Kim, Jae-Joon Kim, Sungjoo Yoo, Sunggu Lee |
DATE | 5 |
| 2014 | Accelerating graph computation with racetrack memory and pointer-assisted graph representationabstractThe poor performance of NAND Flash memory, such as long access latency and large granularity access, is the major bottleneck of graph processing. This paper proposes an intelligent storage for graph processing which is based on fast and low cost racetrack memory and a pointer-assisted graph representation. Our experiments show that the proposed intelligent storage based on racetrack memory reduces total processing time of three representative graph computations by 40.2%~86.9% compared to the graph processing, GraphChi, which exploits sequential accesses based on normal NAND Flash memory-based SSD. Faster execution also reduces energy consumption by 39.6%~90.0%. The in-storage processing capability gives additional 10.5%~16.4% performance improvements and 12.0%~14.4% reduction of energy consumption. Eunhyuk Park, Sungjoo Yoo, Sunggu Lee, Hai Li 0001 |
DATE | 3 |
| 2014 | FPGA-based prototyping systems for emerging memory technologiesabstractAs DRAM faces scaling limit, several new memory technologies are considered as candidates for replacing or complementing DRAM main memory. Compared to DRAM, the new memories have two major differences, non-volatility and write overhead in terms of endurance, latency and power. We built two different FPGA-based evaluation boards to evaluate hardware and software designs for new-memory based main memory; one with a DRAM subsystem having parameterizable latency and non-volatility emulation, and the other with the real chips of new memory namely phase-change RAM (PRAM). We experimented primitive functions and SQLite-based benchmarks on Linux, verifying the workings of new functionalities, e.g., nonvolatility and evaluating the impacts of new memory on software performance. In our experiments, we also demonstrated the impact of new memory-aware software/hardware designs on program performance on a DRAM/PRAM hybrid memory. Taemin Lee, Dongki Kim, Hyunsun Park, Sungjoo Yoo, Sunggu Lee |
RSP | 5 |
| 2013 | A network congestion-aware memory subsystem for manycoreabstractThe network-on-chip (NoC) plays a crucial role in memory performance due to the fact that it can handle the majority of traffics from/to the DRAM memory controllers. However, there has been little work on the interplay between the NoC and memory controllers. In this article, we address a problem called network congestion-induced memory blocking and propose a novel memory controller, which performs memory access scheduling and network entry control in a network congestion-aware manner. In case of network congestion, in order to avoid performance degradation due to the blocking caused by data bound for congested regions in the NoC, the proposed memory controller favors requests and data associated with uncongested regions. In addition, in order to avoid the fairness problem of such a policy, we also propose a gradual method, which enables a trade-off between performance (in memory utilization) and fairness (in memory access latency). Experimental results show that the proposed method can offer up to 1.76 ∼ 2.99 times improvement in memory utilization in the latency-tolerant designs. Dongki Kim, Sungjoo Yoo, Sunggu Lee |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2013 | MAEPER: Matching Access and Error Patterns With Error-Free Resource for Low Vcc L1 CacheabstractLarge SRAMs are the practical bottleneck to achieve a low supply voltage, because they suffer from process variation-induced bit errors at a low supply voltage. In this paper, we present an error-resilient cache architecture that resolves the drawback of previous approaches, i.e., the performance degradation at a low supply voltage which is caused by cache misses in accesses to faulty resources. We utilize cache access locality and error-free resources in a cost-effective manner. First, we classify cache lines into fully and partially accessed groups and apply appropriate methods to each group. For the partially accessed group, we propose a method of matching memory access behavior and error locations with intra-cache line word-level remapping. In order to reduce the area overhead used to store the cache access information history, we present an access pattern-learning line-fill buffer (LFB). For the fully accessed group, we propose the utilization of error-free assist functions in the cache, i.e., a LFB and victim cache with no process variation-induced error at the target minimum supply voltage. We also present an error-aware prefetch method that allows us to utilize the error-free victim cache to achieve a further reduction in cache misses due to faulty resources. Experimental results show that the proposed method gives an average 32.6% reduction in cycles per instruction at an error rate of 0.2% with a small area overhead of 8.2%. Young-Geun Choi, Sungjoo Yoo, Sunggu Lee, Jung Ho Ahn, Kangmin Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Hybrid DRAM/PRAM-based main memory for single-chip CPU/GPUabstractSingle-chip CPU/GPU architecture is being adopted in high-end (embedded) systems, e.g., smartphones and tablet PCs. Main memory subsystem is expected to consist of hybrid DRAM and phase-change RAM (PRAM) due to the difficulties in DRAM scaling. In this work, we address the performance optimization of the hybrid DRAM/PRAM main memory for single chip CPU/GPU. Based on the tight requirements of low latency from CPU and the relative tolerance to long latency from GPU, DRAM is first allocated to CPU while PRAM with longer write latency is allocated to GPU. Then, in order to improve the write performance of GPU traffic, we propose (1) an in-DRAM write buffer to accommodate GPU write traffics, (2) dynamic hot data management to improve the efficiency of write buffer, (3) runtime-adaptive adjustment of write buffer size to meet the given CPU performance bound, and (4) CPU-aware DRAM access scheduling to give low latency to CPU traffics. The experiments show that the proposed method gives 1.02~44.2 times performance improvement in GPU performance with modest (negligible) CPU performance overhead (when compute-intensive CPU programs run). Dongki Kim, Sungkwang Lee, Jaewoong Chung, Daehyun Kim 0001, Dong Hyuk Woo, Sungjoo Yoo, Sunggu Lee |
DAC | 7 |
| 2012 | Write performance improvement by hiding R drift latency in phase-change RAMabstractPhase-change RAM (PRAM) is considered to be one of the most promising candidates to complement or replace DRAM in the near future. However, it is imperative to overcome the limitations of PRAM, especially, long write latency for its widespread applications. R drift latency occupies a significant portion in PRAM write latency thereby adversely affecting system performance. In this paper, we propose a novel method called write status holding register (WSHR) to reduce the write latency due to R drift latency. The WSHR allows for non-blocking accesses to PRAM during R drift latency thereby improving system performance. Our experiments with SPEC benchmarks show that the proposed WSHR gives 53.6%~0% performance improvements in the hybrid DRAM/PRAM main memory (256MB DRAM and 14nm PRAM). Youngsik Kim, Sungjoo Yoo, Sunggu Lee |
DAC | 3 |
| 2012 | A case study on the application of real phase-change RAM to main memory subsystemabstractPhase-change RAM (PCM) has the advantages of better scaling and non-volatility compared with the DRAM which is expected to face its scaling limit in the near future. There have been many studies on applying the PCM to main memory in order to complement or replace the DRAM. One common limitation of these studies is that they are based on synthetic PCM models. In our study, we investigate the feasibility and issues of applying a real PCM to main memory. In this paper, we report our case study of characterizing the PCM and evaluating its usefulness in the main memory. Our results show that the PCM/DRAM hybrid main memory with a modest DRAM size can give comparable performance to that of the DRAM only main memory. However, the hybrid memory with small DRAMs or large footprint programs can suffer from performance degradation due to the long latency of both PCM writes and write preemption penalty, which requires architectural innovations for exploiting the full potential of PCM write performance. Suknam Kwon, Dongki Kim, Youngsik Kim, Sungjoo Yoo, Sunggu Lee |
DATE | 5 |
| 2012 | Bloom filter-based dynamic wear leveling for phase-change RAMabstractPhase Change RAM (PCM) is a promising candidate of emerging memory technology to complement or replace existing DRAM and NAND Flash memory. A key drawback of PCMs is limited write endurance. To address this problem, several static wear-leveling methods that change logical to physical address mapping periodically have been proposed. Although these methods have low space overhead, they suffer from unnecessary data migrations thereby failing to exploit the full lifetime potential of PCMs. This paper proposes a new dynamic wear-leveling method that reduces unnecessary data migrations by adopting a hot/cold swapping-based dynamic method. Compared with the conventional hot/cold swapping-based dynamic method, the proposed method requires only a small amount of space overhead by applying Bloom filters to the identification of hot and cold data. We simulate our method using SPEC2000 benchmark traces and compare with previous methods. Simulation results show that the proposed method reduces unnecessary data migrations by 58~92% and extends the memory lifetime by 2.18~2.30 times over previous methods with a negligible area overhead of 0.3%. Joosung Yun, Sunggu Lee, Sungjoo Yoo |
DATE | 2 |
| 2012 | Optimal wake-up scheduling of data gathering trees for wireless sensor networks
Ungjin Jang, Sunggu Lee, Sungjoo Yoo |
J. Parallel Distributed Comput. | 2 |
| 2012 | A Multistep Tag Comparison Method for a Low-Power L2 CacheabstractTag comparison in a highly associative cache consumes a significant portion of the cache energy. Existing methods for tag comparison reduction are based on predicting either cache hits or cache misses. In this paper, we present novel ideas for both cache hit and miss predictions. We present a partial tag-enhanced Bloom filter to improve the accuracy of the cache miss prediction method and hot/cold checks that control data liveness to reduce the tag comparisons of the cache hit prediction method. We also combine both methods so that their order of application can be dynamically adjusted to adapt to changing cache access behavior, which further reduces tag comparisons. To overcome the common limitation of multistep tag comparison methods, we propose a method that reduces tag comparisons while meeting the given performance bound. Experimental results showed that the proposed method reduces the energy consumption of tag comparison by an average of 88.40%, which translates to an average reduction of 35.34% (40.19% with low-power data access) in the total energy consumption of the L2 cache and a further reduction of 8.86% (10.07% with low-power data access) when compared with existing methods. Hyunsun Park, Sungjoo Yoo, Sunggu Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2012 | Optimizing Video Application Design for Phase-Change RAM-Based Main MemoryabstractVideo applications including video codecs place a large traffic demand on main memory. Emerging memory technology, such as phase-change RAM (PRAM) tends to suffer from the write endurance problem, in which the maximum number of writes is limited. Thus, it is required to improve video application designs to adapt to the new requirements of emerging memory technology, i.e., to minimize the number of writes in terms of bit updates. In this paper, we present a way to optimize video application design for PRAM-based main memory. We propose two methods to resolve the write endurance problem: inter-block differential data encoding and inter-frame multiple experts. Experimental results show an average of 18.4% reduction in bit updates when compared to the best existing data encoding methods for PRAM. Suknam Kwon, Sungjoo Yoo, Sunggu Lee, Jinpyo Park |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Matching cache access behavior and bit error pattern for high performance low Vcc L1 cacheabstractCache is a roadblock towards low supply voltage (Vcc). It is mainly because low Vcc incurs process variation-induced bit errors in large SRAM in cache. Existing approaches for low Vcc cache suffer from low performance due to reduced effective capacity, long latency to correct errors, and increased misses due to accesses to faulty words. In our work, we propose a word-level sub-block disable-based method which increases the utilization of available cache capacity. Our key idea is to minimize accesses to faulty words. To do that, we propose utilizing access behavior history in allocating cache resource with faulty words. In addition, we propose remapping cache words inside of cache line in order to better match both access and error patterns. Experimental results show that the proposed method gives average 21.8% (up to 34.0%) performance improvement with a small area overhead in L1 and L2 caches. Young-Geun Choi, Sungjoo Yoo, Sunggu Lee, Jung Ho Ahn |
DAC | 3 |
| 2011 | Power management of hybrid DRAM/PRAM-based main memoryabstractHybrid main memory consisting of DRAM and non-volatile memory is attractive since the non-volatile memory can give the advantage of low standby power while DRAM provides high performance and better active power. In this work, we address the power management of such a hybrid main memory consisting of DRAM and phase-change RAM (PRAM). In order to reduce DRAM refresh energy which occupies a significant portion of total memory energy, we present a runtime-adaptive method of DRAM decay. In addition, we present two methods, DRAM bypass and dirty data keeping, for further reduction in refresh energy and memory access latency, respectively. The experiments show that by reducing DRAM refreshes, we can obtain 23.5%~94.7% reduction in the energy consumption with negligible performance overhead compared with the conventional DRAM-only main memory. Hyunsun Park, Sungjoo Yoo, Sunggu Lee |
DAC | 3 |
| 2011 | A quantitative analysis of performance benefits of 3D die stacking on mobile and embedded SoCabstract3D stacked DRAM improves peak memory performance. However, its effective performance is often limited by the constraints of row-to-row activation delay (tRRD), four active bank window (tFAW), etc. In this paper, we present a quantitative analysis of the performance impact of such constraints. In order to resolve the problem, we propose balancing the budget of DRAM row activation across DRAM channels. In the proposed method, an inter-memory controller coordinator receives the current demand of row activation from memory controllers and re-distributes the budget to the memory controllers in order to improve DRAM performance. Experimental results show that sharing the budget of row activation between memory channels can give average 4.72% improvement in the utilization of 3D stacked DRAM. Dongki Kim, Sungjoo Yoo, Sunggu Lee, Jung Ho Ahn, Hyunuk Jung |
DATE | 3 |
| 2011 | A novel tag access scheme for low power L2 cacheabstractTag comparisons occupy a significant portion of cache power consumption in the highly associative cache such as L2 cache. In our work, we propose a novel tag access scheme which applies a partial tag-enhanced Bloom filter to reduce tag comparisons by detecting per-way cache misses. The proposed scheme also classifies cache data into hot and cold data and the tags of hot data are compared earlier than those of cold data exploiting the fact that most of cache hits go to hot data. In addition, the power consumption of each tag comparison can be further reduced by dividing the tag comparison into two micro-steps where a partial tag comparison is performed first and, only if the partial tag comparison gives a partial hit, then the remaining tag bits are compared. We applied the proposed scheme to an L2 cache with 10 programs from SPEC2000 and SPEC2006. Experimental results show average 23.69% and 8.58% reduction in cache energy consumption compared with the conventional serial tag-data access and the other existing methods, respectively. Hyunsun Park, Sungjoo Yoo, Sunggu Lee |
DATE | 3 |
| 2010 | A Network Congestion-Aware Memory ControllerabstractNetwork-on-chip and memory controller become correlated with each other in case of high network congestion since the network port of memory controller can be blocked due to the (back-propagated) network congestion. We call such a problem network congestion-induced memory blocking. In order to resolve the problem, we present a novel idea of network congestion-aware memory controller. Based on the global information of network congestion, the memory controller performs (1) congestion-aware memory access scheduling and (2) congestion-aware network entry control of read data. The experimental results obtained from a 5×5 tile architecture show that the proposed memory controller presents up to 18.9% improvement in memory utilization. Dongki Kim, Sungjoo Yoo, Sunggu Lee |
NOCS | 3 |
| 2008 | Fast Fault-Tolerant Time Synchronization for Wireless Sensor NetworksabstractA wireless sensor network (WSN) typically consists of a large number of small-sized devices that have very low computational capability, small amounts of memory and the need to conserve energy as much as possible (most commonly by entering suspended mode for extended periods of time). Previous approaches for WSN time synchronization do not satisfactorily address all of the requirements of WSN environments. Thus, this paper proposes a new fault-tolerant WSN time sychronization algorithm that is extremely fast (when compared to previous algorithms), achieves a guaranteed level of time synchronization for all non-faulty nodes, can accommodate nodes that enter suspended mode and then wake up, utilizes very little communication and computation resources (thereby leaving those resources available for use by other applications), operates in a completely decentralized manner and tolerates up to f faulty nodes. The efficacy of the proposed algorithm is shown using analysis and experimental results. Sunggu Lee, Ungjin Jang |
ISORC | 1 |
| 2007 | Data Dissemination for Wireless Sensor NetworksabstractDue to the special characteristics (limited battery power, limited computing capability, low bandwidth, need to collect sensor data from multiple fixed-location source nodes to a sink node that may be mobile, etc.) of wireless sensor networks, routing algorithms designed for general mobile ad hoc networks may not be directly applicable to wireless sensor networks. In one possible routing scheme for wireless sensor networks, each node maintains up-to-date hop-distances and next-hop nodes to the mobile sink node (or multiple mobile sink nodes). However, this type of method may require too much control overhead in order to maintain up-to-date and consistent hop-distances and next-hop nodes for all of the sensor nodes in the network. Therefore, we propose a new low-control-overhead data dissemination scheme, referred to as pseudo-distance data dissemination, for efficiently disseminating data packets from all sensor nodes to mobile sink nodes in a wireless sensor network Min-Gu Lee, Sunggu Lee |
ISORC | 2 |
| 2007 | Editorial
Uwe Brinkschulte, Sunggu Lee |
Real Time Syst. | 2 |
| 2007 | Push-Pull: Deterministic Search-Based DAG Scheduling for Heterogeneous Cluster SystemsabstractConsider directed acyclic graph (DAG) scheduling for a large heterogeneous system, which consists of processors with varying processing capabilities and network links with varying bandwidths. The search space of possible task schedules for this problem is immense. One possible approach for this optimization problem, which is NP-hard, is to start with the best task schedule found by a fast deterministic task scheduling algorithm and then iteratively attempt to improve the task schedule by employing a general random guided search method. However, such an approach can lead to extremely long search times, and the solutions found are sometimes not significantly better than those found by the original deterministic task scheduling algorithm. In this paper, we propose an alternative strategy, termed Push-Pull, which starts with the best task schedule found by a fast deterministic task scheduling algorithm and then iteratively attempts to improve the current best solution using a deterministic guided search method. Our simulation results show that given similar runtimes, the Push-Pull algorithm performs well, achieving results similar to or better than all of the other algorithms being compared. Sang Cheol Kim, Sunggu Lee, Jaegyoon Hahm |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | A Link Stability Model and Stable Routing for Mobile Ad-Hoc Networks
Min-Gu Lee, Sunggu Lee |
EUC | 2 |
| 2006 | QoS Support for Mobile Ad-Hoc Networks Based on a Reservation PoolabstractMany interesting applications using mobile ad-hoc networks are possible if quality-of-service (QoS) can be effectively supported. Towards that end, this paper proposes a method based on providing a pool of backup paths that can be used if the primary path can no longer support the required level of QoS. Such a capability is especially important for mobile ad-hoc networks because intermediate nodes (within routing paths) can move around in such networks. Route maintenance and switchover should also be performed in an efficient manner. This is accomplished with the help of a method referred to as pseudo-distance routing. Min-Gu Lee, Sunggu Lee |
ISORC | 2 |
| 2005 | Push-Pull: Guided Search DAG Scheduling for Heterogeneous ClustersabstractConsider a heterogeneous cluster system, consisting of processors with varying processing capabilities and network links with varying bandwidths. Given a DAG application to be scheduled on such a system, the search space of possible task schedules is immense. One possible approach for this type of NP-complete problem, which has been proposed by previous researchers, starts with the best task schedule found by a fast deterministic task scheduling algorithm, and then iteratively improves the task schedule using a random search method such as genetic algorithm search. However, such an approach can lead to extremely long search times, and the solutions found are sometimes not significantly better than those found by the original deterministic task scheduling algorithm. In this paper, we propose an alternative strategy, termed push-pull, which starts with the best task schedule found by a fast deterministic task scheduling algorithm, and then iteratively attempts to improve the current best solution using a deterministic guided search method. Our simulation results show that this new method performs quite well, executing faster and finding better solutions than a genetic algorithm-based method, which in turn has been shown to perform better than previous algorithms. Sang Cheol Kim, Sunggu Lee |
ICPP | 2 |
| 2004 | Implementation of a TMO-structured real-time airplane-landing simulator on a distributed computing environmentabstractIn real-time simulation, the simulated system should display the same (or very close) timing behavior as the target system. The simulation accuracy is increased as the simulation time unit is decreased. Although there are several models for such systems, the TMO model is particularly appropriate due to its natural support for real-time distributed object-oriented programming. This paper discusses the results of the implementation of a real-time airplane-landing simulator on a distributed computing environment using the TMO model. Copyright (C) 2004 John Wiley Sons, Ltd. Min-Gu Lee, Sunggu Lee, K. H. (Kane) Kim |
Softw. Pract. Exp. | 2 |
| 2003 | Real-time wormhole channels
Sunggu Lee |
J. Parallel Distributed Comput. | 1 |
| 2003 | Dynamic load balancing for switch-based networks
Wan Yeon Lee, Sung Je Hong, Jong Kim 0001, Sunggu Lee |
J. Parallel Distributed Comput. | 4 |
| 2003 | Secure checkpointing
Hyo-Chang Nam, Jong Kim 0001, Sung Je Hong, Sunggu Lee |
J. Syst. Archit. | 4 |
| 2003 | Task scheduling using a block dependency DAG for block-oriented sparse Cholesky factorization
Heejo Lee, Jong Kim 0001, Sung Je Hong, Sunggu Lee |
Parallel Comput. | 4 |
| 2003 | Processor Allocation and Task Scheduling of Matrix Chain Products on Parallel SystemsabstractThe problem of finding an optimal product sequence for sequential multiplication of a chain of matrices (the matrix chain ordering problem, MCOP) is well-known. We consider the problem of finding an optimal product schedule for evaluating a chain of matrix products on a parallel computer (the matrix chain scheduling problem, MCSP). The difference between MCSP and MCOP is that MCOP pertains to a product sequence for single processor systems and MCSP pertains to a sequence of concurrent matrix products for parallel systems. The approach of parallelizing each matrix product after finding an optimal product sequence for single processor systems does not always guarantee minimum evaluation time on parallel systems since each parallelized matrix product may use processors inefficiently. We introduce a new processor scheduling algorithm for MCSP which reduces the evaluation time of a chain of matrix products on a parallel computer, even at the expense of a slight increase in the total number of operations. Given a chain of n matrices and a matrix product utilizing at most P/k processors in a P-processor system, the proposed algorithm approaches k(n-1)/(n+klog(k)-k) times the performance of parallel evaluation using the optimal sequence found for MCOP. Experiments performed on a Fujitsu AP1000 multicomputer also show that the proposed algorithm significantly decreases the time required to evaluate a chain of matrix products in parallel systems. Heejo Lee, Jong Kim 0001, Sung Je Hong, Sunggu Lee |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2001 | Path Selection Algorithms for Real-Time CommunicationabstractReal-time messages, which have user defined end-to-end delay requirements, can be delivered using a real-time communication method. This paper proposes two path selection algorithms for real time communication and conducts a simulation study to evaluate these two algorithms. On the basis of this simulation study, it can be seen that, for effective real-time communication, it is important to consider the number of hops utilized by each path and the link utilization of all effected links. In particular, the number of hops required for each path was found to be the most important factor. The proposed path selection algorithms are found to perform as well or better than previously proposed algorithms while using fewer network resources. Yun-kyung Lee, Sunggu Lee |
ICPADS | 2 |
| 2001 | A Secure Checkpointing SystemabstractFault-tolerant computer systems are being used increasingly in such applications as e-commerce, banking, and stock trading, where privacy and integrity of data are as important as the uninterrupted operation of the service provided. While much attention has been paid to the protection of data explicitly communicated over the Internet, there are also other sources of information leakage that must be addressed. This paper addresses one such source of information leakage caused by checkpointing, which is a common method used to provide continued operation in the presence of faults. Checkpointing requires communication of memory state information, which may contain sensitive data, over the network to a reliable backing store. Although the method of encrypting all of this memory state information can protect the data, such a simplistic method is an overkill that can result in a significant slowdown of the target application. This paper examines ways to combine the operations required to perform incremental checkpointing with those required to encrypt this memory state data. Analysis and experimentation on an actual system are used to show that the proposed secure checkpointing schemes are feasible and require a relatively low level of overhead. Hyo-Chang Nam, Jong Kim 0001, Sung Je Hong, Sunggu Lee |
PRDC | 4 |
| 2001 | Measurement and Prediction of Communication Delays in Myrinet Networks
Sang Cheol Kim, Sunggu Lee |
J. Parallel Distributed Comput. | 2 |
| 1999 | Reliable Probabilistic CheckpointingabstractRecently proposed probabilistic checkpointing has one drawback, naming aliasing. When analyzed, 64-bit signatures show negligible possibility of aliasing. But in practice, the shift-XOR signature generation function used with probabilistic checkpointing shows a high aliasing rate, which limits the practicality of probabilistic checkpointing. In this paper, two enhancements are considered to make probabilistic checkpointing more reliable. One is the signature generation function and the other is the recovery scheme. In the signature generation function part, we propose two signature generation functions: HALF for small block sizes (less than or equal to 256 bytes) and C-HALF(CRC combined HALF) for large block sizes (larger than 256 bytes), which have an aliasing probability similar to analytic results and small overhead. In the recovery scheme part, we propose a recovery scheme which ensures the safety of probabilistic checkpointing. To examine the correctness of previous checkpoints at recovery time, the proposed recovery scheme uses a spare node. We analyze the recovery scheme using a mathematical model. Also an optimal checkpoint interval is derived using the model. Hyo-Chang Nam, Jong Kim 0001, Sung Je Hong, Sunggu Lee |
PRDC | 4 |
| 1999 | Synchronous Load Balancing in Hypercube Multicomputers with Faulty Nodes
Kyungwan Nam, Jaewon Seo, Sunggu Lee, Jong Kim 0001 |
J. Parallel Distributed Comput. | 3 |
| 1998 | A Real-Time Communication Method for Wormhole Switching NetworksabstractIn this paper we propose a real-time communication scheme that can be used in general point-to-point real-time multicomputer systems with wormhole switching. Real-time communication should satisfy the two requirements of predictability and priority handling. Since traditional wormhole switching does not support priority handling which is essential in real-time computing, flit-level preemption is adopted in our wormhole switching. Also, we develop an algorithm to determine the message transmission delay upper bound to predict worst-case message delay. Simulation results show that the delay upper bounds calculated using the proposed algorithm are very close to actual average message transmission delays for messages with high priorities. Byungjae Kim, Jong Kim 0001, Sung Je Hong, Sunggu Lee |
ICPP | 4 |
| 1998 | Adaptive Virtual Cut-Through as a Viable Routing Method
Howon Kim 0001, Sunggu Lee, Jong Kim 0001 |
J. Parallel Distributed Comput. | 3 |
| 1997 | Synchronous Load Balancing in Hypercube Multicomputers with Faulty NodesabstractThis paper presents a new dynamic load balancing algorithm for hypercube multicomputers with faulty nodes. The emphasis in our method is on obtaining global load information and performing task migration using "short paths" in a synchronous manner so that a minimal amount of communication overhead is required. To accomplish this, we present an algorithm for constructing a new logical topology from a hypercube topology with faulty nodes. This new topology is used to obtain the global load information and to perform task migration. Simulation results are used to evaluate the performance of our dynamic load balancing method. The proposed strategy shows good performance in the case of a small number of faulty nodes when compared with previous methods. Jaewon Seo, Sunggu Lee, Jong Kim 0001 |
ICPADS | 2 |
| 1997 | Real-Time Job Scheduling in Hypercube SystemsabstractIn this paper, we present the problem of scheduling real-time jobs in a hypercube system and propose a scheduling algorithm. The goals of the proposed scheduling algorithm are to determine whether all jobs can complete their processing before their fixed deadlines in a hypercube system and to find such a schedule. Each job is associated with a computation time, a deadline, and a dimensional requirement. Determining a schedule such that all jobs meet before their respective fixed deadlines in a hypercube system when preemption is not allowed is an NP-complete problem. Hence, we present a heuristic scheduling algorithm for scheduling non-preemptable real-time jobs in a hypercube system. Finally, we evaluate the proposed algorithm using simulation. O-Hoon Kwon, Jong Kim 0001, Sung Je Hong, Sunggu Lee |
ICPP | 4 |
| 1997 | Replicated Process Allocation for Load Distribution in Fault-Tolerant MulticomputersabstractIn this paper, we consider a load-balancing process allocation method for fault-tolerant multicomputer systems that balances the load before as well as after faults start to degrade the performance of the system. In order to be able to tolerate a single fault, each process (primary process) is duplicated (i.e., has a backup process). The backup process executes on a different processor from the primary, checkpointing the primary process and recovering the process in the primary process fails. In this paper, we formalize the problem of load-balancing process allocation and propose a new process allocation method and analyze the performance of the proposed method. Simulations are used to compare the proposed method with a process allocation method that does not take into account the different load characteristics of the primary and backup processes. While both methods perform well before the occurrence of a fault, only the proposed method maintains a balanced load after the occurrence of such a fault. Jong Kim 0001, Heejo Lee, Sunggu Lee |
IEEE Trans. Computers | 3 |
| 1996 | Path Selection for Message Passing in a Circuit-Switched Multicomputer
Sunggu Lee, Jong Kim 0001 |
J. Parallel Distributed Comput. | 1 |
| 1995 | DTN: A New Partitionable Torus Topology
Sangho Chae, Jong Kim 0001, Dongseung Kim, Sung Je Hong, Sunggu Lee |
ICPP (1) | 5 |
| 1995 | Adaptive Virutal Cut-through as an Alternative to Wormhole Routing
Howon Kim 0001, Jong Kim 0001, Sunggu Lee |
ICPP (1) | 4 |
| 1994 | Path Selection for Communicating Tasks in a Wormhole-Routed MulticomputerabstractIn a multicomputer that uses wormhole routing or virtual cut-through circuit switching, the communication delay in sending a message between two processors along a path A increases significantly if the message is "blocked" by another message using one of the channels in path A. Such "blocking" can only occur if there is contention for a common channel by two or more paths. In this paper, we consider the problem of selecting contention-free or minimum-contention paths for a set of communicating tasks that have been mapped onto nodes in a multicomputer, This problem is formalized and shown to be an NPhard problem. thus, a heuristic solution is proposed for general static interconnection networks. Simulations results show that our method performs significantly better than alternative methods for this problem. Sunggu Lee, Jong Kim 0001 |
ICPP (3) | 1 |
| 1994 | Interleaved All-to-All Reliable Broadcast on Meshes and HypercubesabstractAll-to-all (ATA) reliable broadcast is the problem of reliably distributing information from every node to every other node in point-to-point interconnection networks. A good solution to this problem is essential for clock synchronization, distributed agreement, etc. We propose a novel solution in which the reliable broadcasts from individual nodes are interleaved in such a manner that no two packets contend for the same link at any given time-this type of method is particularly suited for systems which use virtual cut-through or wormhole routing for fast communication between nodes. Our solution, called the IHC Algorithm, can be used on a large class of regular interconnection networks including regular meshes and hypercubes. By adjusting a parameter /spl eta/ referred to as the interleaving distance, we can flexibly decrease the link utilization of the IHC algorithm (for normal traffic) at the expense of an increase in the time required for ATA reliable broadcast. We compare the IHC algorithm to several other possible virtual cut-through solutions and a store-and-forward solution. The IHC algorithm with the minimum value of /spl eta/ is shown to be optimal in minimizing the execution time of ATA reliable broadcast when used in a dedicated mode (with no other network traffic).> Sunggu Lee, Kang G. Shin |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1994 | On Probabilistic Diagnosis of Multiprocessor Systems Using Multiple SyndromesabstractThis paper addresses the distributed self-diagnosis of a multiprocessor/multicomputer system based on fault syndromes formed by comparison testing. The authors show that by using multiple fault syndromes, it is possible to achieve significantly better diagnosis than by using a single fault syndrome, even when the amount of time devoted to testing is the same. They derive a multiple syndrome diagnosis algorithm that in terms of the level of diagnostic accuracy achieved, is globally suboptimal, but optimal among all diagnosis algorithms of a certain type to be defined. The diagnosis algorithm produces good results, even with sparse interconnection networks and interprocessor tests with low fault coverage. It is also proven that the diagnosis algorithm produces 100% correct diagnosis as N, the number of nodes in the system, approaches /spl infin/, provided that the interconnection network has connectivity greater than or equal to 2 and that the number of syndromes produced grows faster than log N. This solution and another multiple syndrome diagnosis solution by Fussell and Rangarajan (1989) are comparatively evaluated, both analytically and with simulations.> Sunggu Lee, Kang G. Shin |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1993 | Optimal and Efficient Probabilistic Distributed Diagnosis SchemesabstractThe distributed self-diagnosis of a multiprocessor/multicomputer system based on interprocessor tests with imperfect fault coverage that permits intermittently faulty processors is addressed. Focusing on probabilistic diagnosis methods, the authors define several different categories of probabilistic diagnosis based on the type of fault syndrome information used in the diagnosis. Rigorous probabilistic analysis is then used to derive diagnosis algorithms optimal in terms of diagnostic accuracy for the diagnosis categories introduced. Analysis and simulations are used to evaluate the performance of the diagnosis algorithms introduced.> Sunggu Lee, Kang G. Shin |
IEEE Trans. Computers | 1 |
| 1990 | Interleaved All-to-All Reliable Broadcast on Meshes and Hypercubes
Sunggu Lee, Kang G. Shin |
ICPP (3) | 1 |
| 1990 | Design for test using partial parallel scanabstractTraditional scan design techniques such as level-sensitive scan design, scan path, and random-access scan suffer from the drawback that the extra test application effort (which includes both time and memory) required is directly proportional to the number of latches and can become quite significant. A scan design technique termed partial parallel scan which reduces test application effort by one to two orders of magnitude is presented. Theoretical and practical aspects of the design method are discussed. The practical use of the partial parallel scan technique has been demonstrated with an LSI circuit and a VLSI circuit designed using silicon compiler tools.> Sunggu Lee, Kang G. Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |