EDBT 2026 Demo / reviewers in the wild / expert
Seongil O
dblp:148/9828
· DBLP profile ↗
11ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Memory systems · 44% Hardware reliability and fault tolerance · 20% Hardware accelerators and domain-specific architectures · 13% |
Topics — the 25 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
DRAM |
1.0 | 5 | 2021 | Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017 CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015 Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014 |
Hardware reliability and fault tolerance › memory fault tolerance
DRAM resilience |
0.5 | 2 | 2017 | Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017 CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015 |
Memory systems › processing-in-memory
DRAM-PIM |
0.5 | 1 | 2021 | Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial Product · ISCA 2021 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.5 | 1 | 2021 | Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial Product · ISCA 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.5 | 1 | 2021 | Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial Product · ISCA 2021 |
Memory systems
processing-in-memory |
0.5 | 1 | 2021 | Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial Product · ISCA 2021 |
Storage systems › key-value storage
in-memory key-value store |
0.5 | 2 | 2016 | Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform · ACM Trans. Comput. Syst. 2016 Architecting to achieve a billion requests per second throughput on a single key-value store server platform · ISCA 2015 |
Storage systems
key-value storage |
0.5 | 2 | 2016 | Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform · ACM Trans. Comput. Syst. 2016 Architecting to achieve a billion requests per second throughput on a single key-value store server platform · ISCA 2015 |
Hardware reliability and fault tolerance › error-correcting codes for memory
rank-level ECC |
0.3 | 2 | 2017 | CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015 Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017 |
Memory systems › DRAM
DRAM scaling |
0.3 | 1 | 2017 | Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017 |
Hardware reliability and fault tolerance › error correction › error-correcting codes
in-DRAM ECC |
0.3 | 1 | 2017 | Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM Devices · HPCA 2017 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2016 | Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform · ACM Trans. Comput. Syst. 2016 |
Cloud and datacenter computing › datacenter architecture
datacenter server architecture |
0.2 | 1 | 2015 | Architecting to achieve a billion requests per second throughput on a single key-value store server platform · ISCA 2015 |
Hardware reliability and fault tolerance › error-correcting codes for memory
DRAM error correction |
0.2 | 1 | 2015 | CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015 |
Hardware reliability and fault tolerance › error correction
error-correcting codes |
0.2 | 1 | 2015 | CiDRA: A cache-inspired DRAM resilience architecture · HPCA 2015 |
Memory systems › DRAM › DRAM microarchitecture
bank-level parallelism |
0.2 | 1 | 2014 | Microbank: Architecting Through-Silicon Interposer-Based Main Memory Systems · SC 2014 |
Memory systems › DRAM
DRAM architecture |
0.2 | 1 | 2014 | Microbank: Architecting Through-Silicon Interposer-Based Main Memory Systems · SC 2014 |
Memory systems › DRAM
DRAM microarchitecture |
0.2 | 1 | 2014 | Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014 |
Memory systems › DRAM › DRAM microarchitecture
row buffer management |
0.2 | 1 | 2014 | Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014 |
Memory systems
access latency reduction |
0.2 | 1 | 2013 | Reducing memory access latency with asymmetric DRAM bank organizations · ISCA 2013 |
Memory systems
memory access latency |
0.2 | 1 | 2013 | Reducing memory access latency with asymmetric DRAM bank organizations · ISCA 2013 |
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA-based key-value store |
0.1 | 2 | 2016 | Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform · ACM Trans. Comput. Syst. 2016 Architecting to achieve a billion requests per second throughput on a single key-value store server platform · ISCA 2015 |
Energy-efficient computing
memory energy efficiency |
0.1 | 2 | 2014 | Microbank: Architecting Through-Silicon Interposer-Based Main Memory Systems · SC 2014 Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014 |
Memory systems
memory controller |
0.1 | 1 | 2014 | Row-buffer decoupling: A case for low-latency DRAM microarchitecture · ISCA 2014 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 2013 | Reducing memory access latency with asymmetric DRAM bank organizations · ISCA 2013 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.9software stack development · 0.52.5d/3d stacking · 0.5measurement · 0.3full-system characterization · 0.2design principles · 0.2hardware-software co-design · 0.2concurrency control · 0.2bloom filter · 0.2SRAM cache · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial ProductabstractEmerging applications such as deep neural network demand high off-chip memory bandwidth. However, under stringent physical constraints of chip packages and system boards, it becomes very expensive to further increase the bandwidth of off-chip memory. Besides, transferring data across the memory hierarchy constitutes a large fraction of total energy consumption of systems, and the fraction has steadily increased with the stagnant technology scaling and poor data reuse characteristics of such emerging applications. To cost-effectively increase the bandwidth and energy efficiency, researchers began to reconsider the past processing-in-memory (PIM) architectures and advance them further, especially exploiting recent integration technologies such as 2.5D/3D stacking. Albeit the recent advances, no major memory manufacturer has developed even a proof-of-concept silicon yet, not to mention a product. This is because the past PIM architectures often require changes in host processors and/or application code which memory manufacturers cannot easily govern. In this paper, elegantly tackling the aforementioned challenges, we propose an innovative yet practical PIM architecture. To demonstrate its practicality and effectiveness at the system level, we implement it with a 20nm DRAM technology, integrate it with an unmodified commercial processor, develop the necessary software stack, and run existing applications without changing their source code. Our evaluation at the system level shows that our PIM improves the performance of memory-bound neural network kernels and applications by 11.2× and 3.5×, respectively. Atop the performance improvement, PIM also reduces the energy per bit transfer by 3.5×, and the overall energy efficiency of the system running the applications by 3.2×. Sukhan Lee 0002, Shinhaeng Kang, Jaehoon Lee 0005, Eojin Lee, Seungwoo Seo, Hosang Yoon, Seungwon Lee 0006, Kyounghwan Lim, Hyunsung Shin, Jinhyun Kim, Seongil O, Anand Iyer, David Wang 0003, Kyomin Sohn, Nam Sung Kim |
ISCA | 12 |
| 2018 | 3D-Xpath: high-density managed DRAM architecture with cost-effective alternative paths for memory transactionsabstractThe advance of DRAM manufacturing technology slows down, whereas the density and performance needs of DRAM continue to increase. This desire has motivated the industry to explore emerging Non-Volatile Memory (e.g., 3D XPoint) and the high-density DRAM (e.g., Managed DRAM Solution). Since such memory technologies increase the density at the cost of longer latency, lower bandwidth, or both, it is essential to use them with fast memory (e.g., conventional DRAM) to which hot pages are transferred at runtime. Nonetheless, we observe that page transfers to fast memory often block memory channels from servicing memory requests from applications for a long period. This in turn significantly increases the high-percentile response time of latency-sensitive applications. In this paper, we propose a high-density managed DRAM architecture, dubbed 3D-XPath for applications demanding both low latency and high capacity for memory. 3D-XPath DRAM stacks conventional DRAM dies with high-density DRAM dies explored in this paper and connects these DRAM dies with 3D-XPath. Especially, 3D-XPath allows unused memory channels to service memory requests from applications when primary channels supposed to handle the memory requests are blocked by page transfers at given moments, considerably increasing the high-percentile response time. This can also improve the throughput of applications frequently copying memory blocks between kernel and user memory spaces. Our evaluation shows that 3D-XPath DRAM decreases high-percentile response time of latency-sensitive applications by ~30% while improving the throughput of an I/O-intensive applications by ~39%, compared with DRAM without 3D-XPath. Sukhan Lee 0002, Kiwon Lee, Min-Chul Sung, Mohammad Alian, Chankyung Kim, Wooyeong Cho, Reum Oh, Seongil O, Jung Ho Ahn, Nam Sung Kim |
PACT | 8 |
| 2017 | Defect Analysis and Cost-Effective Resilience Architecture for Future DRAM DevicesabstractTechnology scaling has continuously improved the density, performance, energy efficiency, and cost of DRAM-based main memory systems. Starting from sub-20nm processes, however, the industry began to pay considerably higher costs to screen and manage notably increasing defective cells. The traditional technique, which replaces the rows/columns containing faulty cells with spare rows/columns, has been able to cost-effectively repair the defective cells so far, but it will become unaffordable soon because an excessive number of spare rows/columns are required to manage the increasing number of defective cells. This necessitates a synergistic application of an alternative resilience technique such as In-DRAM ECC with the traditional one. Through extensive measurement and simulation, we first identify that aggressive miniaturization makes DRAM cells more sensitive to random telegraph noise or variable retention time, which is dominantly manifested as a surge in randomly scattered single-cell faults. Second, we advocate using InDRAM ECC to overcome the DRAM scaling challenges and architect In-DRAM ECC to accomplish high area efficiency and minimal performance degradation. Moreover, we show that advancement in process technology reduces decoding/correction time to a small fraction of DRAM access time, and that the throughput penalty of a write operation due to an additional read for a parity update is mostly overcome by the multi-bank structure and long burst writes that span an entire In-DRAM ECC codeword. Lastly, we demonstrate that system reliability with modern rank-level ECC schemes such as single device data correction is further improved by hundred million times with the proposed In-DRAM ECC architecture. Sang-uhn Cha, Seongil O, Hyunsung Shin, Sangjoon Hwang 0001, Kwang-Il Park, Seong-Jin Jang 0002, Joo-Sun Choi, Gyo-Young Jin, Young Hoon Son, Hyunyoon Cho, Jung Ho Ahn, Nam Sung Kim |
HPCA | 2 |
| 2016 | Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server PlatformabstractDistributed in-memory key-value stores (KVSs), such as memcached, have become a critical data serving layer in modern Internet-oriented data center infrastructure. Their performance and efficiency directly affect the QoS of web services and the efficiency of data centers. Traditionally, these systems have had significant overheads from inefficient network processing, OS kernel involvement, and concurrency control. Two recent research thrusts have focused on improving key-value performance. Hardware-centric research has started to explore specialized platforms including FPGAs for KVSs; results demonstrated an order of magnitude increase in throughput and energy efficiency over stock memcached. Software-centric research revisited the KVS application to address fundamental software bottlenecks and to exploit the full potential of modern commodity hardware; these efforts also showed orders of magnitude improvement over stock memcached. We aim at architecting high-performance and efficient KVS platforms, and start with a rigorous architectural characterization across system stacks over a collection of representative KVS implementations. Our detailed full-system characterization not only identifies the critical hardware/software ingredients for high-performance KVS systems but also leads to guided optimizations atop a recent design to achieve a record-setting throughput of 120 million requests per second (MRPS) (167MRPS with client-side batching) on a single commodity server. Our system delivers the best performance and energy efficiency (RPS/watt) demonstrated to date over existing KVSs including the best-published FPGA-based and GPU-based claims. We craft a set of design principles for future platform architectures, and via detailed simulations demonstrate the capability of achieving a billion RPS with a single server constructed following our principles. Sheng Li 0007, Hyeontaek Lim, Victor W. Lee, Jung Ho Ahn, Anuj Kalia, Michael Kaminsky, David G. Andersen, Seongil O, Sukhan Lee 0002, Pradeep Dubey |
ACM Trans. Comput. Syst. | 8 |
| 2015 | CiDRA: A cache-inspired DRAM resilience architectureabstractAlthough aggressive technology scaling has allowed manufacturers to integrate Giga bits of cells into a cost-sensitive main memory DRAM device, these cells have become more defect-prone. With increased cell failure rates, conventional solutions such as populating spare DRAM rows and relying on error-correcting codes (ECCs) have shown limited success due to high area overhead, the latency penalties of data coding, and interference between ECC within a device (in-DRAM ECC) and other ECC across devices (rank-level ECC). In this paper, we propose CiDRA, a cache-inspired DRAM resilience architecture, which substantially reduces the area and latency overheads of correcting bit errors on random locations due to these faulty cells. We put a small SRAM cache within a DRAM device to replace accesses to the addresses including the faulty cells with ones that correspond to the cache data array. This CiDRA cache is paired with a Bloom filter to minimize the energy overhead of accessing the cache tags for every DRAM access and is also partitioned into small pieces, each being associated with the I/O pads for better area efficiency. Both the cache and DRAM banks are accessed in parallel while the banks are much slower. Consequently, the cache and filter are not in the critical path for normal DRAM accesses and incur no latency overhead. We also enhance the traditional in-DRAM ECC with error position bits and the appropriate error detecting capability while preventing interference with the traditional rank-level ECC scheme. By combining this enhanced in-DRAM ECC with the cache and Bloom filter, CiDRA becomes more area efficient because the in-DRAM ECC corrects most bit errors that are sporadic while the cache deals with the remaining relatively few pathological cases. Young Hoon Son, Sukhan Lee 0002, Seongil O, Sanghyuk Kwon, Nam Sung Kim, Jung Ho Ahn |
HPCA | 3 |
| 2015 | History-Assisted Adaptive-Granularity Caches (HAAG$) for High Performance 3D DRAM Architecturesabstract3D-stacked DRAM has the potential to provide high performance and large capacity memory for future high performance computing systems and datacenters, and the integration of a dedicated logic die opens up opportunities for architectural enhancements such as DRAM row-buffer caches. However, high performance and cost-effective row-buffer cache designs remain challenging for 3D memory systems. In this paper, we propose History-Assisted Adaptive-Granularity Cache (HAAG$) that employs an adaptive caching scheme to support full associativity at various granularities, and an intelligent history-assisted predictor to support a large number of banks in 3D memory systems. By increasing the row-buffer cache hit rate and avoiding unnecessary data caching, HAAG$ significantly reduces memory access latency and dynamic power. Our design works particularly well for manycore CPUs running (irregular) memory intensive applications where memory locality is hard to exploit. Evaluation results show that with memory-intensive CPU workloads, HAAG$ can outperform the state-of-the-art row buffer cache by 33.5%. Ke Chen 0020, Sheng Li 0007, Jung Ho Ahn, Naveen Muralimanohar, Jishen Zhao, Cong Xu 0002, Seongil O, Yuan Xie 0001, Jay B. Brockman, Norman P. Jouppi |
ICS | 7 |
| 2015 | Architecting to achieve a billion requests per second throughput on a single key-value store server platformabstractDistributed in-memory key-value stores (KVSs), such as memcached, have become a critical data serving layer in modern Internet-oriented datacenter infrastructure. Their performance and efficiency directly affect the QoS of web services and the efficiency of datacenters. Traditionally, these systems have had significant overheads from inefficient network processing, OS kernel involvement, and concurrency control. Two recent research thrusts have focused upon improving key-value performance. Hardware-centric research has started to explore specialized platforms including FPGAs for KVSs; results demonstrated an order of magnitude increase in throughput and energy efficiency over stock memcached. Software-centric research revisited the KVS application to address fundamental software bottlenecks and to exploit the full potential of modern commodity hardware; these efforts too showed orders of magnitude improvement over stock memcached. Sheng Li 0007, Hyeontaek Lim, Victor W. Lee, Jung Ho Ahn, Anuj Kalia, Michael Kaminsky, David G. Andersen, Seongil O, Sukhan Lee 0002, Pradeep Dubey |
ISCA | 8 |
| 2014 | Row-buffer decoupling: A case for low-latency DRAM microarchitectureabstractModern DRAM devices for the main memory are structured to have multiple banks to satisfy ever-increasing throughput, energy-efficiency, and capacity demands. Due to tight cost constraints, only one row can be buffered (opened) per bank and actively service requests at a time, while the row must be deactivated (closed) before a new row is stored into the row buffers. Hasty deactivation unnecessarily re-opens rows for otherwise row-buffer hits while hindsight accompanies the deactivation process on the critical path of accessing data for row-buffer misses. The time to (de)activate a row is comparable to the time to read an open row while applications are often sensitive to DRAM latency. Hence, it is critical to make the right decision on when to close a row. However, the increasing number of banks per DRAM device over generations reduces the number of requests per bank. This forces a memory controller to frequently predict when to close a row due to a lack of information on future requests, while the dynamic nature of memory access patterns limits the prediction accuracy. In this paper, we propose a novel DRAM microarchitecture that can eliminate the need for any prediction. First, we identify that precharging the bitlines dominates the deactivate time, while sense amplifiers that work as a row buffer are physically coupled with the bitlines such that a single command precharges both bitlines and sense amplifiers simultaneously. By decoupling the bitlines from the row buffers using isolation transistors, the bitlines can be precharged right after a row becomes activated. Therefore, only the sense amplifiers need to be precharged for a miss in most cases, taking an order of magnitude shorter time than the conventional deactivation process. Second, we show that this row-buffer decoupling enables internal DRAM μ-operations to be separated and recombined, which can be exploited by memory controllers to make the main memory system more energy efficient. Our experiments demonstrate that row-buffer decoupling improves the geometric mean of the instructions per cycle and MIPS2/W by 14% and 29%, respectively, for memory-intensive SPEC CPU2006 applications. Seongil O, Young Hoon Son, Nam Sung Kim, Jung Ho Ahn |
ISCA | 1 |
| 2014 | Microbank: Architecting Through-Silicon Interposer-Based Main Memory SystemsabstractThrough-Silicon Interposer (TSI) has recently been proposed to provide high memory bandwidth and improve energy efficiency of the main memory system. However, the impact of TSI on main memory system architecture has not been well explored. While TSI improves the I/O energy efficiency, we show that it results in an unbalanced memory system design in terms of energy efficiency as the core DRAM dominates overall energy consumption. To balance and enhance the energy efficiency of a TSI-based memory system, we propose μbank, a novel DRAM device organization in which each bank is partitioned into multiple smaller banks (or μbanks) that operate independently like conventional banks with minimal area overhead. The μbank organization significantly increases the amount of bank-level parallelism to improve the performance and energy efficiency of the TSI-based memory system. The massive number of μbanks reduces bank conflicts, hence simplifying the memory system design. We evaluated a sophisticated prediction-based DRAM page-management policy, which can improve performance by up to 20.5% in a conventional memory system without μbanks. However, a μbank-based design does not require such a complex page-management policy and a simple open-page policy is often sufficient -- achieving within 5% of a perfect predictor. Our proposed μbank-based memory system improves the IPC and system energy-delay product by 1.62× and 4.80×, respectively, for memory-intensive SPEC 2006 benchmarks on average, over the baseline DDR3-based memory system. Young Hoon Son, Seongil O, Hyunggyun Yang, Daejin Jung, Jung Ho Ahn, John Kim 0001, Jangwoo Kim, Jae W. Lee |
SC | 2 |
| 2013 | Reducing memory access latency with asymmetric DRAM bank organizationsabstractDRAM has been a de facto standard for main memory, and advances in process technology have led to a rapid increase in its capacity and bandwidth. In contrast, its random access latency has remained relatively stagnant, as it is still around 100 CPU clock cycles. Modern computer systems rely on caches or other latency tolerance techniques to lower the average access latency. However, not all applications have ample parallelism or locality that would help hide or reduce the latency. Moreover, applications' demands for memory space continue to grow, while the capacity gap between last-level caches and main memory is unlikely to shrink. Consequently, reducing the main-memory latency is important for application performance. Unfortunately, previous proposals have not adequately addressed this problem, as they have focused only on improving the bandwidth and capacity or reduced the latency at the cost of significant area overhead. Young Hoon Son, Seongil O, Yuhwan Ro, Jae W. Lee, Jung Ho Ahn |
ISCA | 2 |
| 2013 | McSimA+: A manycore simulator with application-level+ simulation and detailed microarchitecture modelingabstractWith their significant performance and energy advantages, emerging manycore processors have also brought new challenges to the architecture research community. Manycore processors are highly integrated complex system-on-chips with complicated core and uncore subsystems. The core subsystems can consist of a large number of traditional and asymmetric cores. The uncore subsystems have also become unprecedentedly powerful and complex with deeper cache hierarchies, advanced on-chip interconnects, and high-performance memory controllers. In order to conduct research for emerging manycore processor systems, a microarchitecture-level and cycle-level manycore simulation infrastructure is needed. This paper introduces McSimA+, a new timing simulation infrastructure, to meet these needs. McSimA+ models x86based asymmetric manycore microarchitectures in detail for both core and uncore subsystems, including a full spectrum of asymmetric cores from single-threaded to multithreaded and from in-order to out-of-order, sophisticated cache hierarchies, coherence hardware, on-chip interconnects, memory controllers, and main memory. McSimA+ is an application-level+ simulator, offering a middle ground between a full-system simulator and an application-level simulator. Therefore, it enjoys the light weight of an application-level simulator and the full control of threads and processes as in a full-system simulator. This paper also explores an asymmetric clustered manycore architecture that can reduce the thread migration cost to achieve a noticeable performance improvement compared to a state-of-the-art asymmetric manycore architecture. Jung Ho Ahn, Sheng Li 0007, Seongil O, Norman P. Jouppi |
ISPASS | 3 |