EDBT 2026 Demo / reviewers in the wild / expert
Dongmoon Min
dblp:242/9233
· DBLP profile ↗
15ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-4503-0823ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 10 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SuperSFQ: A Hardware Design to Realize High-Frequency Superconducting ProcessorsabstractSuperconducting computing using single flux quantum (SFQ) technology has been recognized as a promising post-Moore's law era technology thanks to its extremely low power and high performance.Therefore, many researchers have proposed various SFQbased circuits (e.g., ALU, register file) and architectures (e.g., NPU, CPU) to exploit the potential.However, due to the absence of a reliable and high-frequency clocking scheme, general SFQ circuits cannot operate at high frequencies, making all architectural efforts for high-performance SFQ computing ineffective.In this paper, we propose SuperSFQ, a new design methodology for SFQ hardware that unlocks the high-frequency potential of SFQ technology by co-designing the clocking scheme, circuitry, and architecture.First, we propose SuperClocking, a new clocking scheme that enables high frequency in general SFQ hardware.Second, we implement an SFQ-based synchronizer to realize the reliable operation of SuperClocking.Finally, we provide two architectural design guidelines and corresponding solutions to ensure the functional correctness of SuperClocking in general SFQ devices.By applying our clocking scheme, synchronizer, and guidelines to the latest general-purpose SFQ CPU, SuperSFQ achieves up to 62.5 times higher frequency and improves single-thread and multithread performance by 17 and 62.5 times, respectively, compared to conventional designs, with only 34.4% Josephson junction overhead.In addition, to demonstrate the generality of SuperSFQ, we apply SuperSFQ to 48 different benchmark circuits, achieving 88.5 times higher frequency compared to conventional designs, on average. Junhyuk Choi, Juwon Hong, Junpyo Kim, Jungmin Cho, Hyeonseong Jeong, Dongmoon Min, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
MICRO | 6 |
| 2025 | LANCER: Low-Overhead, Accurate, and Non-Destructive Calibration for Real-World Fault-Tolerant Quantum ApplicationsabstractThe ultimate goal of fault-tolerant quantum computing (FTQC) is to run practical applications.Due to the long execution time of practical workloads, an FTQC system must operate reliably for multiple days by correcting the errors of noisy qubits.However, drifts of error sources increase qubit error rates during execution (i.e., error drift), limiting the reliable execution time.Even worse, existing error-drift-handling methods cannot execute long-running workloads as they collapse the qubit states or fail to suppress errors.In this paper, we propose LANCER, a novel accurate and nondestructive calibration method for reliable execution under error drifts.We observe that only a subset of qubits store quantum states during execution.Based on the observation, we periodically stall the program and migrate the quantum states temporarily to idle qubits, enabling accurate calibrations without losing the states.However, this idea faces two major challenges: (1) crosstalk between running and calibrating qubits and (2) huge latency overhead due to stalls when the quantum states are migrated to idle qubits.We propose three solutions to resolve these challenges.First, we mitigate the crosstalk by toggling the frequencies of running qubits to separate them from the frequencies of calibrating qubits.Second, we reduce the latency overhead by utilizing the inherent idle times in the fault-tolerant quantum gate.Lastly, we further reduce the latency overhead by re-designing qubit layout to enable the execution even when the quantum states are migrated to idle qubits.The evaluation shows that LANCER enables the execution of 95 times larger programs (i.e., larger number of gates) compared to the baseline, with negligible latency and qubit overhead (4.3% and 4.0%, respectively). Junpyo Kim, Jungmin Cho, Hyeonseong Jeong, Dongmoon Min, Junhyuk Choi, Juwon Hong, Jangwoo Kim |
MICRO | 4 |
| 2024 | A Fault-Tolerant Million Qubit-Scale Distributed Quantum ComputerabstractA million qubit-scale quantum computer is essential to realize the quantum supremacy. Modern large-scale quantum computers integrate multiple quantum computers located in dilution refrigerators (DR) to overcome each DR's unscaling cooling budget. However, a large-scale multi-DR quantum computer introduces its unique challenges (i.e., slow and erroneous inter-DR entanglement, increased qubit scale), and they make the baseline error handling mechanism ineffective by increasing the number of gate operations and the inter-DR communication latency to decode and correct errors. Without resolving these challenges, it is impossible to realize a fault-tolerant large-scale multi-DR quantum computer. Junpyo Kim, Dongmoon Min, Jungmin Cho, Hyeonseong Jeong, Ilkwon Byun, Junhyuk Choi, Juwon Hong, Jangwoo Kim |
ASPLOS (2) | 2 |
| 2024 | SuperCore: An Ultra-Fast Superconducting Processor for Cryogenic ApplicationsabstractSuperconductor single-flux-quantum (SFQ) logic family has been recognized as a promising technology for cryogenic applications (e.g., quantum computing, astronomy, metrology) thanks to its ultra-fast and low-energy characteristics. Therefore, recent efforts in SFQ-based computing have focused on developing fast and low-power SFQ processors for cryogenic applications. However, there still has been little progress toward a convincing SFQ processor design due to the critical performance challenges originating from its extremely deep pipeline. In this paper, we propose a super-fast and low-power in-order SFQ processor by tackling the challenges from the deep pipeline. First, we develop a minimal-depth SFQ processor pipeline with novel architecture-level ideas. Next, we conduct in-depth performance analyses and identify three real performance bottlenecks in the deeply pipelined SFQ processors (i.e., stall/flush logic, RAW stall, fetch unit). Finally, we propose SuperCore, our super-fast SFQ-based processor architecture, with three SFQ-friendly solutions that effectively resolve the identified bottlenecks. With our solutions applied, SuperCore achieves 11 times speed-up over the SFQ processor baseline. In addition, SuperCore achieves six times speed-up and consumes up to 193 times less power compared to in-order CMOS processors running at 4K. Junhyuk Choi, Ilkwon Byun, Juwon Hong, Dongmoon Min, Junpyo Kim, Jungmin Cho, Hyeonseong Jeong, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
MICRO | 4 |
| 2024 | CoolDC: A Cost-Effective Immersion-Cooled Datacenter with Workload-Aware Temperature ScalingabstractFor datacenter architects, it is the most important goal to minimize the datacenter’s total cost of ownership for the target performance (i.e., TCO/performance). As the major component of a datacenter is a server farm, the most effective way of reducing TCO/performance is to improve the server’s performance and power efficiency. To achieve the goal, we claim that it is highly promising to reduce each server’s temperature to its most cost-effective point (or temperature scaling). In this article, we propose CoolDC , a novel and immediately applicable low-temperature cooling method to minimize the datacenter’s TCO. The key idea is to find and apply the most cost-effective sub-freezing temperature to target servers and workloads. For that purpose, we first apply the immersion cooling method to the entire servers to maintain a stable low temperature with little extra cooling and maintenance costs. Second, we define the TCO-optimal temperature for datacenter operation (e.g., 248K~273K (-25℃~0℃)) by carefully estimating all the costs and benefits at low temperatures. Finally, we propose CoolDC, our immersion-cooling datacenter architecture to run every workload at its own TCO-optimal temperature. By incorporating our low-temperature workload-aware temperature scaling, CoolDC achieves 12.7% and 13.4% lower TCO/performance than the conventional air-cooled and immersion-cooled datacenters, respectively, without any modification to existing computers. Dongmoon Min, Ilkwon Byun, Gyu-hyeon Lee, Jangwoo Kim |
ACM Trans. Archit. Code Optim. | 1 |
| 2023 | QIsim: Architecting 10+K Qubit QC Interfaces Toward Quantum SupremacyabstractA 10+K qubit Quantum-Classical Interface (QCI) is essential to realize the quantum supremacy. However, it is extremely challenging to architect scalable QCIs due to the complex scalability trade-offs regarding operating temperatures, device and wire technologies, and microarchitecture designs. Therefore, architects need a modeling tool to evaluate various QCI design choices and lead to an optimal scalable QCI architecture. Dongmoon Min, Junpyo Kim, Junhyuk Choi, Ilkwon Byun, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
ISCA | 1 |
| 2023 | Is the Future Cold or Tall? Design Space Exploration of Cryogenic and 3D Embedded Cache MemoryabstractMemory latency, density, and power efficiency are key bottlenecks in a variety of computing systems, and the need for efficient and dense memory solutions is exacerbated by the continued importance of data-intensive applications such as machine learning, graph processing, and scientific computing. A myriad of emerging technologies and approaches aim to address the limitations of current systems. For example, 3D integration can enable highly dense memory structures, and multiple alternative device technologies such as STT and PCM have emerged as compelling solutions to improve memory system density and efficiency. Additionally, cryogenic operation of computing systems (i.e., ultra-low temperature cooling) is becoming a compelling solution as thermal hotspots have become a primary roadblock to conventional transistor scaling. This work probes, evaluates, and compares the potential capabilities of 3D integration, embedded non-volatile memories (eNVMs), and cryogenic operation towards improving future memory systems by presenting the first design space exploration of cryogenic operation and 3D integration applied towards the largest on-chip memory structure, the last level cache, as well as presenting and providing open-source tools for future, related design studies. This work specffically evaluates the applicationlevel benefits or limitations of such proposals by leveraging a cross-computing-stack simulation approach. Our studies reveal that the most compelling solution varies depending on the expected memory traffic patterns and workloads of interest, which in turn exposes several opportunities for future optimization and customization. For example, due to potentially high costs of cooling to cryogenic operation, we find that SRAM or 3T-eDRAM operating at 77K is sub-optimal compared to room-temperature SRAM and eNVM solutions, but exhibits advantages for relatively low-traffic workloads. Alexander Hankin, Lillian Pentecost, Dongmoon Min, David Brooks 0001, Gu-Yeon Wei |
ISPASS | 3 |
| 2023 | Fast, Light-weight, and Accurate Performance Evaluation using Representative Datacenter BehaviorsabstractDatacenters rapidly evolve by adopting new features such as new hardware deployment and software patches. Adopting a new feature requires an accurate evaluation of its impact to minimize the risk to the multi-million dollar computing infrastructure. However, a comprehensive performance analysis of a datacenter is extremely challenging due to its cost and multitenancy. Evaluating the performance in a live datacenter is accurate but prohibitive to prevent any damage to production services. Using conventional load-testing benchmarks on small-scale testbeds is imprecise as they do not consider the effect of other co-located jobs. Dongmoon Min, Ilkwon Byun, Hanhwi Jang, Jangwoo Kim |
Middleware | 2 |
| 2022 | CryoWire: wire-driven microarchitecture designs for cryogenic computingabstractCryogenic computing, which runs a computer device at an extremely low temperature, is promising thanks to its significant reduction of wire resistance as well as leakage current. Recent studies on cryogenic computing have focused on various architectural units including the main memory, cache, and CPU core running at 77K. However, little research has been conducted to fully exploit the fast cryogenic wires, even though the slow wires are becoming more serious performance bottleneck in modern processors. In this paper, we propose a CPU microarchitecture which extensively exploits the fast wires at 77K. For this goal, we first introduce our validated cryogenic-performance models for the CPU pipeline and network on chip (NoC), whose performance can be significantly limited by the slow wires. Next, based on the analysis with the models, we architect CryoSP and CryoBus as our pipeline and NoC designs to fully exploit the fast wires. Our evaluation shows that our cryogenic computer equipped with both microarchitectures achieves 3.82 times higher system-level performance compared to the conventional computer system thanks to the 96% higher clock frequency of CryoSP and five times lower NoC latency of CryoBus. Dongmoon Min, Yujin Chung, Ilkwon Byun, Junpyo Kim, Jangwoo Kim |
ASPLOS | 1 |
| 2022 | XQsim: modeling cross-technology control processors for 10+K qubit quantum computersabstract10+K qubit quantum computer is essential to achieve a true sense of quantum supremacy. With the recent effort towards the large-scale quantum computer, architects have revealed various scalability issues including the constraints in a quantum control processor, which should be holistically analyzed to design a future scalable control processor. However, it has been impossible to identify and resolve the processor's scalability bottleneck due to the absence of a reliable tool to explore an extensive design space including microarchitecture, device technology, and operating temperature. Ilkwon Byun, Junpyo Kim, Dongmoon Min, Ikki Nagaoka, Kosuke Fukumitsu, Iori Ishikawa, Teruo Tanimoto, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
ISCA | 3 |
| 2022 | 3D-FPIM: An Extreme Energy-Efficient DNN Acceleration System Using 3D NAND Flash-Based In-Situ PIM UnitabstractThe crossbar structure of the nonvolatile memory enables highly parallel and energy-efficient analog matrix-vector-multiply (MVM) operations. To exploit its efficiency, existing works design a mixed-signal deep neural network (DNN) accelerator, which offloads low-precision MVM operations to the memory array. However, they fail to accurately and efficiently support the low-precision networks due to their naive ADC designs. In addition, they cannot be applied to the latest technology nodes due to their premature RRAM-based memory array.In this work, we present 3D-FPIM, an energy-efficient and robust mixed-signal DNN acceleration system. 3D-FPIM is a full-stack 3D NAND flash-based architecture to accurately deploy low-precision networks. We design the hardware stack by carefully architecting a specialized analog-to-digital conversion method and utilizing the three-dimensional structure to achieve high accuracy, energy efficiency, and robustness. To accurately and efficiently deploy the networks, we provide a DNN retraining framework and a customized compiler. For evaluation, we implement an industry-validated circuit-level simulator. The result shows that 3D-FPIM achieves an average of 2.09x higher performance per area and 13.18x higher energy efficiency compared to the baseline 2D RRAM-based accelerator. Hunjun Lee, Minseop Kim, Dongmoon Min, Joonsung Kim 0001, Jongwon Back, Honam Yoo, Jong-Ho Lee 0002, Jangwoo Kim |
MICRO | 3 |
| 2021 | CryoGuard: A Near Refresh-Free Robust DRAM Design for Cryogenic ComputingabstractCryogenic computing, which runs a computer device at an extremely low temperature, is highly promising thanks to the significant reduction of the wire latency and leakage current. A recently proposed cryogenic DRAM design achieved the promising performance improvement, but it also reveals that it must reduce the DRAM’s dynamic power to overcome the huge cooling cost at 77 K. Therefore, researchers now target to reduce the cryogenic DRAM’s refresh power by utilizing its significantly increased retention time driven by the reduced leakage current. To achieve the goal, however, architects should first answer many fundamental questions regarding the reliability and then design a refresh-free, but still robust cryogenic DRAM by utilizing the analysis result.In this work, we propose a near refresh-free, but robust cryogenic DRAM (NRFC-DRAM), which can almost eliminate its refresh overhead while ensuring reliable operations at 77 K. For the purpose, we first evaluate various DRAM samples of multiple vendors by conducting a thorough analysis to accurately estimate the cryogenic DRAM’s retention time and reliability. Our analysis identifies a new critical challenge such that reducing DRAM’s refresh rate can make the memory highly unreliable because normal memory operations can now appear as row-hammer attacks at 77 K. Therefore, NRFC-DRAM requires a cost-effective, cryogenic-friendly protection mechanism against the new row-hammer-like "faults" at 77 K.To resolve the challenge, we present CryoGuard, our cryogenic-friendly row-hammer protection method to ensure the NRFC-DRAM’s reliable operations at 77 K. With CryoGuard applied, NRFC-DRAM reduces the overall power consumption by 25.9 % even with its cooling cost included, whereas the existing cryogenic DRAM fails to reduce the power consumption. Gyu-hyeon Lee, Seongmin Na, Ilkwon Byun, Dongmoon Min, Jangwoo Kim |
ISCA | 4 |
| 2020 | CryoCache: A Fast, Large, and Cost-Effective Cache Architecture for Cryogenic ComputingabstractCryogenic computing, which is to run a computer at extremely low temperatures (e.g., 77K), is a highly promising solution to dramatically improve the computer's performance and power efficiency thanks to the significantly reduced leakage power and wire resistance. However, computer architects are facing fundamental challenges in developing and deploying cryogenic-optimal architectural units due to the lack of understanding about its cost-effectiveness and feasibility (e.g., device and cooling costs vs. speedup, energy and area saving) and thus how to architect such cryogenic-optimal units. Dongmoon Min, Ilkwon Byun, Gyu-hyeon Lee, Seongmin Na, Jangwoo Kim |
ASPLOS | 1 |
| 2020 | CryoCore: A Fast and Dense Processor Architecture for Cryogenic ComputingabstractCryogenic computing can achieve high performance and power efficiency by dramatically reducing the device's leakage power and wire resistance at low temperatures. Recent advances towards cryogenic computing focus on developing cryogenic-optimal cache and memory devices to overcome memory capacity, latency, and power walls. However, little research has been conducted to develop a cryogenic-optimal core architecture despite its high potentials in performance, power, and area efficiency. Once a cryogenic-optimal core becomes available, it will also take full advantage of the cryogenic-optimal cache and memory devices, which leads to a cryogenic-optimal computer.In this paper, we first develop CryoCore-Model (CC-Modet), a cryogenic processor modeling framework which can accurately estimate the maximum clock frequency of processor models running at 77K. Next, driven by the modeling tool, we design CryoCore, a 77K-optimal core microarchitecture to maximize the core's performance and area efficiency while minimizing the cooling cost. The key idea of CryoCore is to architect a core in a way to reduce the size and number of cooling-unfriendly microarchitecture units and maximize the potential of a voltage and frequency scaling at 77K. Finally, we propose two halfsized, but differently voltage-scaled CryoCore designs aiming for either the maximum performance or power efficiency. With both conventional and our design integrated with cryogenic memories, our high-performance CryoCore design achieves 41% higher single-thread performance for the same power budget and 2x higher multi-thread performance for the same die area. Our lowpower CryoCore design reduces the power cost by 38% without sacrificing the single-thread performance. Ilkwon Byun, Dongmoon Min, Gyu-hyeon Lee, Seongmin Na, Jangwoo Kim |
ISCA | 2 |
| 2019 | Cryogenic computer architecture modeling with memory-side case studiesabstractModern computer architectures suffer from lack of architectural innovations, mainly due to the power wall and the memory wall. That is, architectural innovations become infeasible because they can prohibitively increase power consumption and their performance impacts are eventually bounded by slow memory accesses. To address the challenges, making computer systems run at ultra-low temperatures (or cryogenic computer systems) has emerged as a highly promising solution as both power consumption and wire resistivity are expected to significantly reduce at ultra-low temperatures. However, cryogenic computers have not been yet realized as computer architects do not fully understand the behaviors of existing computer systems and their cost effectiveness at such ultra-low temperatures. Gyu-hyeon Lee, Dongmoon Min, Ilkwon Byun, Jangwoo Kim |
ISCA | 2 |