Seongil O

dblp:148/9828 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Memory systems · 44% Hardware reliability and fault tolerance · 20% Hardware accelerators and domain-specific architectures · 13%

Topics — the 25 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
DRAM
1.052021
Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017
CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015
Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014
Hardware reliability and fault tolerance › memory fault tolerance
DRAM resilience
0.522017
Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017
CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015
Memory systems › processing-in-memory
DRAM-PIM
0.512021
Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial Product · ISCA 2021
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.512021
Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial Product · ISCA 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.512021
Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial Product · ISCA 2021
Memory systems
processing-in-memory
0.512021
Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial Product · ISCA 2021
Storage systems › key-value storage
in-memory key-value store
0.522016
Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform · ACM Trans. Comput. Syst. 2016
Architecting to achieve a billion requests per second throughput on a single key-value store server platform · ISCA 2015
Storage systems
key-value storage
0.522016
Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform · ACM Trans. Comput. Syst. 2016
Architecting to achieve a billion requests per second throughput on a single key-value store server platform · ISCA 2015
Hardware reliability and fault tolerance › error-correcting codes for memory
rank-level ECC
0.322017
CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015
Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017
Memory systems › DRAM
DRAM scaling
0.312017
Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017
Hardware reliability and fault tolerance › error correction › error-correcting codes
in-DRAM ECC
0.312017
Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017
Performance modeling and evaluation
workload characterization
0.212016
Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform · ACM Trans. Comput. Syst. 2016
Cloud and datacenter computing › datacenter architecture
datacenter server architecture
0.212015
Architecting to achieve a billion requests per second throughput on a single key-value store server platform · ISCA 2015
Hardware reliability and fault tolerance › error-correcting codes for memory
DRAM error correction
0.212015
CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015
Hardware reliability and fault tolerance › error correction
error-correcting codes
0.212015
CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015
Memory systems › DRAM › DRAM microarchitecture
bank-level parallelism
0.212014
Microbank: Architecting Through-Silicon Interposer-Based Main Memory Systems · SC 2014
Memory systems › DRAM
DRAM architecture
0.212014
Microbank: Architecting Through-Silicon Interposer-Based Main Memory Systems · SC 2014
Memory systems › DRAM
DRAM microarchitecture
0.212014
Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014
Memory systems › DRAM › DRAM microarchitecture
row buffer management
0.212014
Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014
Memory systems
access latency reduction
0.212013
Reducing memory access latency with asymmetric DRAM bank organizations · ISCA 2013
Memory systems
memory access latency
0.212013
Reducing memory access latency with asymmetric DRAM bank organizations · ISCA 2013
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA-based key-value store
0.122016
Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform · ACM Trans. Comput. Syst. 2016
Architecting to achieve a billion requests per second throughput on a single key-value store server platform · ISCA 2015
Energy-efficient computing
memory energy efficiency
0.122014
Microbank: Architecting Through-Silicon Interposer-Based Main Memory Systems · SC 2014
Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014
Memory systems
memory controller
0.112014
Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014
Performance modeling and evaluation
simulation
0.012013
Reducing memory access latency with asymmetric DRAM bank organizations · ISCA 2013

Methods — techniques the papers use, named apart from their topics

simulation · 0.9software stack development · 0.52.5d/3d stacking · 0.5measurement · 0.3full-system characterization · 0.2design principles · 0.2hardware-software co-design · 0.2concurrency control · 0.2bloom filter · 0.2SRAM cache · 0.2
YearPublicationVenuePosition
2021 Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial Product
abstract
Emerging applications such as deep neural network demand high off-chip memory bandwidth. However, under stringent physical constraints of chip packages and system boards, it becomes very expensive to further increase the bandwidth of off-chip memory. Besides, transferring data across the memory hierarchy constitutes a large fraction of total energy consumption of systems, and the fraction has steadily increased with the stagnant technology scaling and poor data reuse characteristics of such emerging applications. To cost-effectively increase the bandwidth and energy efficiency, researchers began to reconsider the past processing-in-memory (PIM) architectures and advance them further, especially exploiting recent integration technologies such as 2.5D/3D stacking. Albeit the recent advances, no major memory manufacturer has developed even a proof-of-concept silicon yet, not to mention a product. This is because the past PIM architectures often require changes in host processors and/or application code which memory manufacturers cannot easily govern. In this paper, elegantly tackling the aforementioned challenges, we propose an innovative yet practical PIM architecture. To demonstrate its practicality and effectiveness at the system level, we implement it with a 20nm DRAM technology, integrate it with an unmodified commercial processor, develop the necessary software stack, and run existing applications without changing their source code. Our evaluation at the system level shows that our PIM improves the performance of memory-bound neural network kernels and applications by 11.2× and 3.5×, respectively. Atop the performance improvement, PIM also reduces the energy per bit transfer by 3.5×, and the overall energy efficiency of the system running the applications by 3.2×.
Sukhan Lee 0002, Shinhaeng Kang, Jaehoon Lee 0005, Eojin Lee, Seungwoo Seo, Hosang Yoon, Seungwon Lee 0006, Kyounghwan Lim, Hyunsung Shin, Jinhyun Kim, Seongil O, Anand Iyer, David Wang 0003, Kyomin Sohn, Nam Sung Kim
ISCA12
2018 3D-Xpath: high-density managed DRAM architecture with cost-effective alternative paths for memory transactions
abstract
The advance of DRAM manufacturing technology slows down, whereas the density and performance needs of DRAM continue to increase. This desire has motivated the industry to explore emerging Non-Volatile Memory (e.g., 3D XPoint) and the high-density DRAM (e.g., Managed DRAM Solution). Since such memory technologies increase the density at the cost of longer latency, lower bandwidth, or both, it is essential to use them with fast memory (e.g., conventional DRAM) to which hot pages are transferred at runtime. Nonetheless, we observe that page transfers to fast memory often block memory channels from servicing memory requests from applications for a long period. This in turn significantly increases the high-percentile response time of latency-sensitive applications. In this paper, we propose a high-density managed DRAM architecture, dubbed 3D-XPath for applications demanding both low latency and high capacity for memory. 3D-XPath DRAM stacks conventional DRAM dies with high-density DRAM dies explored in this paper and connects these DRAM dies with 3D-XPath. Especially, 3D-XPath allows unused memory channels to service memory requests from applications when primary channels supposed to handle the memory requests are blocked by page transfers at given moments, considerably increasing the high-percentile response time. This can also improve the throughput of applications frequently copying memory blocks between kernel and user memory spaces. Our evaluation shows that 3D-XPath DRAM decreases high-percentile response time of latency-sensitive applications by ~30% while improving the throughput of an I/O-intensive applications by ~39%, compared with DRAM without 3D-XPath.
Sukhan Lee 0002, Kiwon Lee, Min-Chul Sung, Mohammad Alian, Chankyung Kim, Wooyeong Cho, Reum Oh, Seongil O, Jung Ho Ahn, Nam Sung Kim
PACT8
2017 Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices
abstract
Technology scaling has continuously improved the density, performance, energy efficiency, and cost of DRAM-based main memory systems. Starting from sub-20nm processes, however, the industry began to pay considerably higher costs to screen and manage notably increasing defective cells. The traditional technique, which replaces the rows/columns containing faulty cells with spare rows/columns, has been able to cost-effectively repair the defective cells so far, but it will become unaffordable soon because an excessive number of spare rows/columns are required to manage the increasing number of defective cells. This necessitates a synergistic application of an alternative resilience technique such as In-DRAM ECC with the traditional one. Through extensive measurement and simulation, we first identify that aggressive miniaturization makes DRAM cells more sensitive to random telegraph noise or variable retention time, which is dominantly manifested as a surge in randomly scattered single-cell faults. Second, we advocate using InDRAM ECC to overcome the DRAM scaling challenges and architect In-DRAM ECC to accomplish high area efficiency and minimal performance degradation. Moreover, we show that advancement in process technology reduces decoding/correction time to a small fraction of DRAM access time, and that the throughput penalty of a write operation due to an additional read for a parity update is mostly overcome by the multi-bank structure and long burst writes that span an entire In-DRAM ECC codeword. Lastly, we demonstrate that system reliability with modern rank-level ECC schemes such as single device data correction is further improved by hundred million times with the proposed In-DRAM ECC architecture.
Sang-uhn Cha, Seongil O, Hyunsung Shin, Sangjoon Hwang 0001, Kwang-Il Park, Seong-Jin Jang 0002, Joo-Sun Choi, Gyo-Young Jin, Young Hoon Son, Hyunyoon Cho, Jung Ho Ahn, Nam Sung Kim
HPCA2
2016 Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform
abstract
Distributed in-memory key-value stores (KVSs), such as memcached, have become a critical data serving layer in modern Internet-oriented data center infrastructure. Their performance and efficiency directly affect the QoS of web services and the efficiency of data centers. Traditionally, these systems have had significant overheads from inefficient network processing, OS kernel involvement, and concurrency control. Two recent research thrusts have focused on improving key-value performance. Hardware-centric research has started to explore specialized platforms including FPGAs for KVSs; results demonstrated an order of magnitude increase in throughput and energy efficiency over stock memcached. Software-centric research revisited the KVS application to address fundamental software bottlenecks and to exploit the full potential of modern commodity hardware; these efforts also showed orders of magnitude improvement over stock memcached. We aim at architecting high-performance and efficient KVS platforms, and start with a rigorous architectural characterization across system stacks over a collection of representative KVS implementations. Our detailed full-system characterization not only identifies the critical hardware/software ingredients for high-performance KVS systems but also leads to guided optimizations atop a recent design to achieve a record-setting throughput of 120 million requests per second (MRPS) (167MRPS with client-side batching) on a single commodity server. Our system delivers the best performance and energy efficiency (RPS/watt) demonstrated to date over existing KVSs including the best-published FPGA-based and GPU-based claims. We craft a set of design principles for future platform architectures, and via detailed simulations demonstrate the capability of achieving a billion RPS with a single server constructed following our principles.
Sheng Li 0007, Hyeontaek Lim, Victor W. Lee, Jung Ho Ahn, Anuj Kalia, Michael Kaminsky, David G. Andersen, Seongil O, Sukhan Lee 0002, Pradeep Dubey
ACM Trans. Comput. Syst.8
2015 CiDRA: A cache-inspired DRAM resilience architecture
abstract
Although aggressive technology scaling has allowed manufacturers to integrate Giga bits of cells into a cost-sensitive main memory DRAM device, these cells have become more defect-prone. With increased cell failure rates, conventional solutions such as populating spare DRAM rows and relying on error-correcting codes (ECCs) have shown limited success due to high area overhead, the latency penalties of data coding, and interference between ECC within a device (in-DRAM ECC) and other ECC across devices (rank-level ECC). In this paper, we propose CiDRA, a cache-inspired DRAM resilience architecture, which substantially reduces the area and latency overheads of correcting bit errors on random locations due to these faulty cells. We put a small SRAM cache within a DRAM device to replace accesses to the addresses including the faulty cells with ones that correspond to the cache data array. This CiDRA cache is paired with a Bloom filter to minimize the energy overhead of accessing the cache tags for every DRAM access and is also partitioned into small pieces, each being associated with the I/O pads for better area efficiency. Both the cache and DRAM banks are accessed in parallel while the banks are much slower. Consequently, the cache and filter are not in the critical path for normal DRAM accesses and incur no latency overhead. We also enhance the traditional in-DRAM ECC with error position bits and the appropriate error detecting capability while preventing interference with the traditional rank-level ECC scheme. By combining this enhanced in-DRAM ECC with the cache and Bloom filter, CiDRA becomes more area efficient because the in-DRAM ECC corrects most bit errors that are sporadic while the cache deals with the remaining relatively few pathological cases.
Young Hoon Son, Sukhan Lee 0002, Seongil O, Sanghyuk Kwon, Nam Sung Kim, Jung Ho Ahn
HPCA3
2015 History-Assisted Adaptive-Granularity Caches (HAAG$) for High Performance 3D DRAM Architectures
abstract
3D-stacked DRAM has the potential to provide high performance and large capacity memory for future high performance computing systems and datacenters, and the integration of a dedicated logic die opens up opportunities for architectural enhancements such as DRAM row-buffer caches. However, high performance and cost-effective row-buffer cache designs remain challenging for 3D memory systems. In this paper, we propose History-Assisted Adaptive-Granularity Cache (HAAG$) that employs an adaptive caching scheme to support full associativity at various granularities, and an intelligent history-assisted predictor to support a large number of banks in 3D memory systems. By increasing the row-buffer cache hit rate and avoiding unnecessary data caching, HAAG$ significantly reduces memory access latency and dynamic power. Our design works particularly well for manycore CPUs running (irregular) memory intensive applications where memory locality is hard to exploit. Evaluation results show that with memory-intensive CPU workloads, HAAG$ can outperform the state-of-the-art row buffer cache by 33.5%.
Ke Chen 0020, Sheng Li 0007, Jung Ho Ahn, Naveen Muralimanohar, Jishen Zhao, Cong Xu 0002, Seongil O, Yuan Xie 0001, Jay B. Brockman, Norman P. Jouppi
ICS7
2015 Architecting to achieve a billion requests per second throughput on a single key-value store server platform
abstract
Distributed in-memory key-value stores (KVSs), such as memcached, have become a critical data serving layer in modern Internet-oriented datacenter infrastructure. Their performance and efficiency directly affect the QoS of web services and the efficiency of datacenters. Traditionally, these systems have had significant overheads from inefficient network processing, OS kernel involvement, and concurrency control. Two recent research thrusts have focused upon improving key-value performance. Hardware-centric research has started to explore specialized platforms including FPGAs for KVSs; results demonstrated an order of magnitude increase in throughput and energy efficiency over stock memcached. Software-centric research revisited the KVS application to address fundamental software bottlenecks and to exploit the full potential of modern commodity hardware; these efforts too showed orders of magnitude improvement over stock memcached.
Sheng Li 0007, Hyeontaek Lim, Victor W. Lee, Jung Ho Ahn, Anuj Kalia, Michael Kaminsky, David G. Andersen, Seongil O, Sukhan Lee 0002, Pradeep Dubey
ISCA8
2014 Row-buffer decoupling: A case for low-latency DRAM microarchitecture
abstract
Modern DRAM devices for the main memory are structured to have multiple banks to satisfy ever-increasing throughput, energy-efficiency, and capacity demands. Due to tight cost constraints, only one row can be buffered (opened) per bank and actively service requests at a time, while the row must be deactivated (closed) before a new row is stored into the row buffers. Hasty deactivation unnecessarily re-opens rows for otherwise row-buffer hits while hindsight accompanies the deactivation process on the critical path of accessing data for row-buffer misses. The time to (de)activate a row is comparable to the time to read an open row while applications are often sensitive to DRAM latency. Hence, it is critical to make the right decision on when to close a row. However, the increasing number of banks per DRAM device over generations reduces the number of requests per bank. This forces a memory controller to frequently predict when to close a row due to a lack of information on future requests, while the dynamic nature of memory access patterns limits the prediction accuracy. In this paper, we propose a novel DRAM microarchitecture that can eliminate the need for any prediction. First, we identify that precharging the bitlines dominates the deactivate time, while sense amplifiers that work as a row buffer are physically coupled with the bitlines such that a single command precharges both bitlines and sense amplifiers simultaneously. By decoupling the bitlines from the row buffers using isolation transistors, the bitlines can be precharged right after a row becomes activated. Therefore, only the sense amplifiers need to be precharged for a miss in most cases, taking an order of magnitude shorter time than the conventional deactivation process. Second, we show that this row-buffer decoupling enables internal DRAM μ-operations to be separated and recombined, which can be exploited by memory controllers to make the main memory system more energy efficient. Our experiments demonstrate that row-buffer decoupling improves the geometric mean of the instructions per cycle and MIPS2/W by 14% and 29%, respectively, for memory-intensive SPEC CPU2006 applications.
Seongil O, Young Hoon Son, Nam Sung Kim, Jung Ho Ahn
ISCA1
2014 Microbank: Architecting Through-Silicon Interposer-Based Main Memory Systems
abstract
Through-Silicon Interposer (TSI) has recently been proposed to provide high memory bandwidth and improve energy efficiency of the main memory system. However, the impact of TSI on main memory system architecture has not been well explored. While TSI improves the I/O energy efficiency, we show that it results in an unbalanced memory system design in terms of energy efficiency as the core DRAM dominates overall energy consumption. To balance and enhance the energy efficiency of a TSI-based memory system, we propose μbank, a novel DRAM device organization in which each bank is partitioned into multiple smaller banks (or μbanks) that operate independently like conventional banks with minimal area overhead. The μbank organization significantly increases the amount of bank-level parallelism to improve the performance and energy efficiency of the TSI-based memory system. The massive number of μbanks reduces bank conflicts, hence simplifying the memory system design. We evaluated a sophisticated prediction-based DRAM page-management policy, which can improve performance by up to 20.5% in a conventional memory system without μbanks. However, a μbank-based design does not require such a complex page-management policy and a simple open-page policy is often sufficient -- achieving within 5% of a perfect predictor. Our proposed μbank-based memory system improves the IPC and system energy-delay product by 1.62× and 4.80×, respectively, for memory-intensive SPEC 2006 benchmarks on average, over the baseline DDR3-based memory system.
Young Hoon Son, Seongil O, Hyunggyun Yang, Daejin Jung, Jung Ho Ahn, John Kim 0001, Jangwoo Kim, Jae W. Lee
SC2
2013 Reducing memory access latency with asymmetric DRAM bank organizations
abstract
DRAM has been a de facto standard for main memory, and advances in process technology have led to a rapid increase in its capacity and bandwidth. In contrast, its random access latency has remained relatively stagnant, as it is still around 100 CPU clock cycles. Modern computer systems rely on caches or other latency tolerance techniques to lower the average access latency. However, not all applications have ample parallelism or locality that would help hide or reduce the latency. Moreover, applications' demands for memory space continue to grow, while the capacity gap between last-level caches and main memory is unlikely to shrink. Consequently, reducing the main-memory latency is important for application performance. Unfortunately, previous proposals have not adequately addressed this problem, as they have focused only on improving the bandwidth and capacity or reduced the latency at the cost of significant area overhead.
Young Hoon Son, Seongil O, Yuhwan Ro, Jae W. Lee, Jung Ho Ahn
ISCA2
2013 McSimA+: A manycore simulator with application-level+ simulation and detailed microarchitecture modeling
abstract
With their significant performance and energy advantages, emerging manycore processors have also brought new challenges to the architecture research community. Manycore processors are highly integrated complex system-on-chips with complicated core and uncore subsystems. The core subsystems can consist of a large number of traditional and asymmetric cores. The uncore subsystems have also become unprecedentedly powerful and complex with deeper cache hierarchies, advanced on-chip interconnects, and high-performance memory controllers. In order to conduct research for emerging manycore processor systems, a microarchitecture-level and cycle-level manycore simulation infrastructure is needed. This paper introduces McSimA+, a new timing simulation infrastructure, to meet these needs. McSimA+ models x86based asymmetric manycore microarchitectures in detail for both core and uncore subsystems, including a full spectrum of asymmetric cores from single-threaded to multithreaded and from in-order to out-of-order, sophisticated cache hierarchies, coherence hardware, on-chip interconnects, memory controllers, and main memory. McSimA+ is an application-level+ simulator, offering a middle ground between a full-system simulator and an application-level simulator. Therefore, it enjoys the light weight of an application-level simulator and the full control of threads and processes as in a full-system simulator. This paper also explores an asymmetric clustered manycore architecture that can reduce the thread migration cost to achieve a noticeable performance improvement compared to a state-of-the-art asymmetric manycore architecture.
Jung Ho Ahn, Sheng Li 0007, Seongil O, Norman P. Jouppi
ISPASS3