EDBT 2026 Demo / reviewers in the wild / expert
Junpyo Kim
dblp:314/7079
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Computation in Quantum LDPC CodeabstractFault-tolerant quantum computing (FTQC) uses quantum error correction (QEC) codes to execute large-scale quantum programs on noisy quantum computers. Quantum low-density parity-check (qLDPC) codes are promising as they use an order of magnitude fewer qubits than widely used surface codes. However, as qLDPC codes support only a limited set of operations, they require programs to be decomposed into many qLDPC-supported operations. This prohibitively increases the execution time of FTQC applications to tens of days, hindering the practical viability of qLDPC codes. Jungmin Cho, Hyeonseong Jeong, Junpyo Kim, Junhyuk Choi, Juwon Hong, Jangwoo Kim |
ASPLOS (2) | 3 |
| 2026 | Leveraging Retrieval-Augmented Language Models for Accurate Item/Feature Selection in Conversational Recommender SystemsabstractConversational recommender systems (CRSs) aim to provide personalized item recommendations along with explanations based on the conversations with users. While advancements in language models (LMs) have facilitated CRSs, limitations remain when LMs lack sufficient knowledge about item features that are essential for accurate recommendations and appropriate explanations. To alleviate this issue, retrieval-augmented language models (RALMs) have been introduced; however, they introduce a new challenge: the inclusion of less-relevant knowledge in retrieved passages. To address this limitation, we propose a novel CRS framework, MOCHA, which enhances RALMs through a multi-stage item/feature selection with Chain-of-Thought (CoT) reasoning. Specifically, MOCHA systematically identifies relevant knowledge by first selecting the item to recommend and then selecting its features to explain; each selection is performed via CoT reasoning. Experimental results on two public CRS datasets demonstrate that MOCHA significantly improves the recommendation accuracy, and provides informative and factually-correct explanations for the recommended items. Taeho Kim 0003, Junpyo Kim, Won-Yong Shin, Sang-Wook Kim |
WSDM | 2 |
| 2025 | SuperSFQ: A Hardware Design to Realize High-Frequency Superconducting ProcessorsabstractSuperconducting computing using single flux quantum (SFQ) technology has been recognized as a promising post-Moore's law era technology thanks to its extremely low power and high performance.Therefore, many researchers have proposed various SFQbased circuits (e.g., ALU, register file) and architectures (e.g., NPU, CPU) to exploit the potential.However, due to the absence of a reliable and high-frequency clocking scheme, general SFQ circuits cannot operate at high frequencies, making all architectural efforts for high-performance SFQ computing ineffective.In this paper, we propose SuperSFQ, a new design methodology for SFQ hardware that unlocks the high-frequency potential of SFQ technology by co-designing the clocking scheme, circuitry, and architecture.First, we propose SuperClocking, a new clocking scheme that enables high frequency in general SFQ hardware.Second, we implement an SFQ-based synchronizer to realize the reliable operation of SuperClocking.Finally, we provide two architectural design guidelines and corresponding solutions to ensure the functional correctness of SuperClocking in general SFQ devices.By applying our clocking scheme, synchronizer, and guidelines to the latest general-purpose SFQ CPU, SuperSFQ achieves up to 62.5 times higher frequency and improves single-thread and multithread performance by 17 and 62.5 times, respectively, compared to conventional designs, with only 34.4% Josephson junction overhead.In addition, to demonstrate the generality of SuperSFQ, we apply SuperSFQ to 48 different benchmark circuits, achieving 88.5 times higher frequency compared to conventional designs, on average. Junhyuk Choi, Juwon Hong, Junpyo Kim, Jungmin Cho, Hyeonseong Jeong, Dongmoon Min, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
MICRO | 3 |
| 2025 | LANCER: Low-Overhead, Accurate, and Non-Destructive Calibration for Real-World Fault-Tolerant Quantum ApplicationsabstractThe ultimate goal of fault-tolerant quantum computing (FTQC) is to run practical applications.Due to the long execution time of practical workloads, an FTQC system must operate reliably for multiple days by correcting the errors of noisy qubits.However, drifts of error sources increase qubit error rates during execution (i.e., error drift), limiting the reliable execution time.Even worse, existing error-drift-handling methods cannot execute long-running workloads as they collapse the qubit states or fail to suppress errors.In this paper, we propose LANCER, a novel accurate and nondestructive calibration method for reliable execution under error drifts.We observe that only a subset of qubits store quantum states during execution.Based on the observation, we periodically stall the program and migrate the quantum states temporarily to idle qubits, enabling accurate calibrations without losing the states.However, this idea faces two major challenges: (1) crosstalk between running and calibrating qubits and (2) huge latency overhead due to stalls when the quantum states are migrated to idle qubits.We propose three solutions to resolve these challenges.First, we mitigate the crosstalk by toggling the frequencies of running qubits to separate them from the frequencies of calibrating qubits.Second, we reduce the latency overhead by utilizing the inherent idle times in the fault-tolerant quantum gate.Lastly, we further reduce the latency overhead by re-designing qubit layout to enable the execution even when the quantum states are migrated to idle qubits.The evaluation shows that LANCER enables the execution of 95 times larger programs (i.e., larger number of gates) compared to the baseline, with negligible latency and qubit overhead (4.3% and 4.0%, respectively). Junpyo Kim, Jungmin Cho, Hyeonseong Jeong, Dongmoon Min, Junhyuk Choi, Juwon Hong, Jangwoo Kim |
MICRO | 1 |
| 2024 | A Fault-Tolerant Million Qubit-Scale Distributed Quantum ComputerabstractA million qubit-scale quantum computer is essential to realize the quantum supremacy. Modern large-scale quantum computers integrate multiple quantum computers located in dilution refrigerators (DR) to overcome each DR's unscaling cooling budget. However, a large-scale multi-DR quantum computer introduces its unique challenges (i.e., slow and erroneous inter-DR entanglement, increased qubit scale), and they make the baseline error handling mechanism ineffective by increasing the number of gate operations and the inter-DR communication latency to decode and correct errors. Without resolving these challenges, it is impossible to realize a fault-tolerant large-scale multi-DR quantum computer. Junpyo Kim, Dongmoon Min, Jungmin Cho, Hyeonseong Jeong, Ilkwon Byun, Junhyuk Choi, Juwon Hong, Jangwoo Kim |
ASPLOS (2) | 1 |
| 2024 | SuperCore: An Ultra-Fast Superconducting Processor for Cryogenic ApplicationsabstractSuperconductor single-flux-quantum (SFQ) logic family has been recognized as a promising technology for cryogenic applications (e.g., quantum computing, astronomy, metrology) thanks to its ultra-fast and low-energy characteristics. Therefore, recent efforts in SFQ-based computing have focused on developing fast and low-power SFQ processors for cryogenic applications. However, there still has been little progress toward a convincing SFQ processor design due to the critical performance challenges originating from its extremely deep pipeline. In this paper, we propose a super-fast and low-power in-order SFQ processor by tackling the challenges from the deep pipeline. First, we develop a minimal-depth SFQ processor pipeline with novel architecture-level ideas. Next, we conduct in-depth performance analyses and identify three real performance bottlenecks in the deeply pipelined SFQ processors (i.e., stall/flush logic, RAW stall, fetch unit). Finally, we propose SuperCore, our super-fast SFQ-based processor architecture, with three SFQ-friendly solutions that effectively resolve the identified bottlenecks. With our solutions applied, SuperCore achieves 11 times speed-up over the SFQ processor baseline. In addition, SuperCore achieves six times speed-up and consumes up to 193 times less power compared to in-order CMOS processors running at 4K. Junhyuk Choi, Ilkwon Byun, Juwon Hong, Dongmoon Min, Junpyo Kim, Jungmin Cho, Hyeonseong Jeong, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
MICRO | 5 |
| 2023 | QIsim: Architecting 10+K Qubit QC Interfaces Toward Quantum SupremacyabstractA 10+K qubit Quantum-Classical Interface (QCI) is essential to realize the quantum supremacy. However, it is extremely challenging to architect scalable QCIs due to the complex scalability trade-offs regarding operating temperatures, device and wire technologies, and microarchitecture designs. Therefore, architects need a modeling tool to evaluate various QCI design choices and lead to an optimal scalable QCI architecture. Dongmoon Min, Junpyo Kim, Junhyuk Choi, Ilkwon Byun, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
ISCA | 2 |
| 2022 | CryoWire: wire-driven microarchitecture designs for cryogenic computingabstractCryogenic computing, which runs a computer device at an extremely low temperature, is promising thanks to its significant reduction of wire resistance as well as leakage current. Recent studies on cryogenic computing have focused on various architectural units including the main memory, cache, and CPU core running at 77K. However, little research has been conducted to fully exploit the fast cryogenic wires, even though the slow wires are becoming more serious performance bottleneck in modern processors. In this paper, we propose a CPU microarchitecture which extensively exploits the fast wires at 77K. For this goal, we first introduce our validated cryogenic-performance models for the CPU pipeline and network on chip (NoC), whose performance can be significantly limited by the slow wires. Next, based on the analysis with the models, we architect CryoSP and CryoBus as our pipeline and NoC designs to fully exploit the fast wires. Our evaluation shows that our cryogenic computer equipped with both microarchitectures achieves 3.82 times higher system-level performance compared to the conventional computer system thanks to the 96% higher clock frequency of CryoSP and five times lower NoC latency of CryoBus. Dongmoon Min, Yujin Chung, Ilkwon Byun, Junpyo Kim, Jangwoo Kim |
ASPLOS | 4 |
| 2022 | XQsim: modeling cross-technology control processors for 10+K qubit quantum computersabstract10+K qubit quantum computer is essential to achieve a true sense of quantum supremacy. With the recent effort towards the large-scale quantum computer, architects have revealed various scalability issues including the constraints in a quantum control processor, which should be holistically analyzed to design a future scalable control processor. However, it has been impossible to identify and resolve the processor's scalability bottleneck due to the absence of a reliable tool to explore an extensive design space including microarchitecture, device technology, and operating temperature. Ilkwon Byun, Junpyo Kim, Dongmoon Min, Ikki Nagaoka, Kosuke Fukumitsu, Iori Ishikawa, Teruo Tanimoto, Masamitsu Tanaka, Koji Inoue, Jangwoo Kim |
ISCA | 2 |