Dong-Wan Kim

dblp:88/8064 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware reliability and fault tolerance · 42% Memory systems · 24% Distributed systems · 24%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.422016
RelaxFault Memory Repair · ISCA 2016
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Memory systems
cache
0.322016
Balancing reliability, cost, and performance tradeoffs with FreeFault · HPCA 2015
RelaxFault Memory Repair · ISCA 2016
Memory systems › memory management
address remapping
0.212016
RelaxFault Memory Repair · ISCA 2016
Hardware reliability and fault tolerance
error correction
0.212016
RelaxFault Memory Repair · ISCA 2016
Hardware reliability and fault tolerance › memory reliability
fault repair
0.212016
RelaxFault Memory Repair · ISCA 2016
Hardware reliability and fault tolerance
memory reliability
0.212016
RelaxFault Memory Repair · ISCA 2016
Hardware reliability and fault tolerance › memory fault tolerance
DRAM fault tolerance
0.212015
Balancing reliability, cost, and performance tradeoffs with FreeFault · HPCA 2015
Hardware reliability and fault tolerance
memory fault tolerance
0.212015
Balancing reliability, cost, and performance tradeoffs with FreeFault · HPCA 2015
Distributed systems › fault tolerance
checkpointing
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Interconnection networks and networks-on-chip › error control
error detection and recovery
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
High-performance computing › supercomputing
exascale computing
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Distributed systems › fault tolerance
resilience
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.112016
RelaxFault Memory Repair · ISCA 2016
Memory systems
DRAM
0.112015
Balancing reliability, cost, and performance tradeoffs with FreeFault · HPCA 2015

Methods — techniques the papers use, named apart from their topics

memory scrubbing · 0.2cache associativity · 0.2trace-driven simulation · 0.1analytical modeling · 0.1
YearPublicationVenuePosition
2016 RelaxFault Memory Repair
abstract
Memory system reliability is a serious concern in many systems today, and is becoming more worrisome as technology scales and system size grows. Stronger fault tolerance capability is therefore desirable, but often comes at high cost. In this paper, we propose a low-cost, fault-aware, hardware-only resilience mechanism, RelaxFault, that repairs the vast majority of memory faults using a small amount of the LLC to remap faulty memory locations. RelaxFault requires less than 100KiB of LLC capacity, has near-zero impact on performance and power. By repairing faults, RelaxFault relaxes the requirement for high fault tolerance of other mechanisms, such as ECC. A better tradeoff between resilience and overhead is made by exploiting an understanding of memory system architecture and fault characteristics. We show that RelaxFault provides better repair capability than prior work of similar cost, improves memory reliability to a greater extent, and significantly reduces the number of maintenance events and memory module replacements. We also propose a more refined memory fault model than prior work and demonstrate its importance.
Dong-Wan Kim, Mattan Erez
ISCA1
2015 Stay Alive, Don't Give Up: DUE and SDC Reduction with Memory Repair
abstract
Memory faults are the most common fault in current systems. Strong memory error protection techniques prevent most memory faults from immediately affecting applications and the system, but come at a high and possibly increasing cost as system size increases and fabrication technology continues to scale. To effectively protect against memory faults and errors, a combination of online error-checking and correcting codes are used in conjunction with module replacement. This paper makes two contributions. First, we study the magnitude of the silent data corruption problem when faulty modules are replaced in a timely manner. Second, we show that currently rarely-used fine-grained repair mechanisms drastically reduce both the module replacement rate and the risk of silent data corruption in large systems.
Dong-Wan Kim, Mattan Erez
CLUSTER1
2015 Balancing reliability, cost, and performance tradeoffs with FreeFault
abstract
Memory errors have been a major source of system failures and fault rates may rise even further as memory continues to scale. This increasing fault rate, especially when combined with advent of integrated on-package memories, may exceed the capabilities of traditional fault tolerance mechanisms or significantly increase their overhead. In this paper, we present FreeFault as a hardware-only, transparent, and nearly-free resilience mechanism that is implemented entirely within a processor and can tolerate the majority of DRAM faults. FreeFault repurposes portions of the last-level cache for storing retired memory regions and augments a hardware memory scrubber to monitor memory health and aid retirement decisions. Because it relies on existing structures (cache associativity) for retirement/remapping type repair, FreeFault has essentially no hardware overhead. Because it requires a very modest portion of the cache (as small as 8KB) to cover a large fraction of DRAM faults, FreeFault has almost no impact on performance. We explain how FreeFault adds an attractive layer in an overall resilience scheme of highly-reliable and highly-available systems by delaying, and even entirely avoiding, calling upon software to make tradeoff decisions between memory capacity, performance, and reliability.
Dong-Wan Kim, Mattan Erez
HPCA1
2012 Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems
abstract
This paper describes and evaluates a scalable and efficient resilience scheme based on the concept of containment domains. Containment domains are a programming construct that enable applications to express resilience needs and to interact with the system to tune and specialize error detection, state preservation and restoration, and recovery schemes. Containment domains have weak transactional semantics and are nested to take advantage of the machine and application hierarchies and to enable hierarchical state preservation, restoration, and recovery. We evaluate the scalability and efficiency of containment domains using generalized trace-driven simulation and analytical analysis and show that containment domains are superior to both checkpoint restart and redundant execution approaches.
Jinsuk Chung, Ikhwan Lee, Michael B. Sullivan 0001, Jeeho Ryoo, Dong-Wan Kim, Doe Hyun Yoon, Larry Kaplan, Mattan Erez
SC5
2008 Design of fuzzy power system stabilizer using adaptive evolutionary algorithm
Gi-Hyun Hwang, Dong-Wan Kim, Young-Joo An
Eng. Appl. Artif. Intell.2