EDBT 2026 Demo / reviewers in the wild / expert
Joonseop Sim
dblp:211/0052
· DBLP profile ↗
7ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-3417-0289ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Deployable CXL-PNM: The CMM-Ax Prototype and Software StackabstractProcessing-Near-Memory (PNM) over Compute Express Link (CXL) has strong architectural appeal. However, most existing CXL-attached PNM prototypes remain inflexible or simulation-based and lack integration with full software stacks, which limits their practical deployment. This work introduces CMM-Ax, a deployable hybrid CXL-PNM system that combines streaming near-memory pipelines with device-side programmability and integrates tightly with the Heterogeneous Memory Software Development Kit (HMSDK). Through HMSDK, CMM-Ax exposes CXL-attached memory via standard Linux allocation paths (e.g., malloc and mmap). This allows existing vector databases and search engines to adopt heterogeneous memory with minimal application-level changes, such as selecting the CMM-Ax FAISS backend. CMM-Ax further provides a comprehensive software stack—FAISS integration, a domain-specific compiler for user-defined operators, and a Kubernetes device plugin supporting multi-tenant slicing—capabilities not demonstrated in prior CXL-PNM. Using an FPGA-based CXL Type-3 prototype, CMM-Ax achieves 4.54× higher throughput and 5.56× lower energy per query than a CXL memory-only baseline on the exact k-nearest neighbor (kNN) search. For approximate nearest neighbor (ANN) inverted-file (IVF) workloads, CMM-Ax sustains bandwidth-proportional efficiency at small batches (batch=1: 37% vs. 12% CPU, 7% GPU). In Kubernetes deployments, CMM-Ax–equipped servers can replace a significantly larger CPU-only cluster at equivalent memory capacity and throughput, reducing node count and system energy. Kwangsik Shin, KangKyu Park, Joonseop Sim, Thomas Won Ha Choi, Youngpyo Joo, Hoshik Kim |
IEEE Trans. Computers | 3 |
| 2024 | Computational CXL-Memory Solution for Accelerating Memory-Intensive ApplicationsabstractCXL interface is the up-to-date technology that enables effective memory expansion by providing a memory-sharing protocol in configuring heterogeneous devices. However, its limited physical bandwidth can be a significant bottleneck for emerging data-intensive applications. In this work, we propose a novel CXL-based memory disaggregation architecture with a real-world prototype demonstration, which overcomes the bandwidth limitation of the CXL interface using near-data processing. The experimental results demonstrate that our design achieves up to 1.9× better performance/power efficiency than the existing CPU system. Joonseop Sim, Soohong Ahn, Taeyoung Ahn, Seungyong Lee 0005, Myunghyun Rhee, Kwangsik Shin, Donguk Moon, Euiseok Kim, Kyoung Park |
HPCA | 1 |
| 2022 | Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory SystemsabstractModern deep learning (DL) training is memory-consuming, constrained by the memory capacity of each computation component and cross-device communication bandwidth. In response to such constraints, current approaches include increasing parallelism in distributed training and optimizing inter-device communication. However, model parameter communication is becoming a key performance bottleneck in distributed DL training. To improve parameter communication performance, we propose COARSE, a disaggregated memory extension for distributed DL training. COARSE is built on modern cache-coherent interconnect (CCI) protocols and MPI-like collective communication for synchronization, to allow low-latency and parallel access to training data and model parameters shared among worker GPUs. To enable high bandwidth transfers between GPUs and the disaggregated memory system, we propose a decentralized parameter communication scheme to decouple and localize parameter synchronization traffic. Furthermore, we propose dynamic tensor routing and partitioning to fully utilize the non-uniform serial bus bandwidth varied across different cloud computing systems. Finally, we design a deadlock avoidance and dual synchronization to ensure high-performance parameter synchronization. Our evaluation shows that COARSE achieves up to 48.3% faster DL training compared to the state-of-the-art MPI AllReduce communication. Zixuan Wang 0027, Joonseop Sim, Eui-Cheol Lim, Jishen Zhao |
HPCA | 2 |
| 2022 | COSMO: Computing with Stochastic Numbers in MemoryabstractStochastic computing (SC) reduces the complexity of computation by representing numbers with long streams of independent bits. However, increasing performance in SC comes with either an increase in area or a loss in accuracy. Processing in memory (PIM) computes data in-place while having high memory density and supporting bit-parallel operations with low energy consumption. In this article, we propose COSMO, an architecture for co mputing with s tochastic numbers in me mo ry, which enables SC in memory. The proposed architecture is general and can be used for a wide range of applications. It is a highly dense and parallel architecture that supports most SC encodings and operations in memory. It maximizes the performance and energy efficiency of SC by introducing several innovations: (i) in-memory parallel stochastic number generation, (ii) efficient implication-based logic in memory, (iii) novel memory bit line segmenting, (iv) a new memory-compatible SC addition operation, and (v) enabling flexible block allocation. To show the generality and efficiency of our stochastic architecture, we implement image processing, deep neural networks (DNNs), and hyperdimensional (HD) computing on the proposed hardware. Our evaluations show that running DNN inference on COSMO is 141× faster and 80× more energy efficient as compared to GPU. Saransh Gupta, Mohsen Imani, Joonseop Sim, Andrew Huang 0001, Jaeyoung Kang 0001, Yeseong Kim, Tajana Rosing |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2020 | SCRIMP: A General Stochastic Computing Architecture using ReRAM in-Memory ProcessingabstractStochastic computing (SC) reduces the complexity of computation by representing numbers with long independent bit-streams. However, increasing performance in SC comes with increase in area and loss in accuracy. Processing in memory (PIM) with non-volatile memories (NVMs) computes data inplace, while having high memory density and supporting bitparallel operations with low energy. In this paper, we propose SCRIMP for stochastic computing acceleration with resistive RAM (ReRAM) in-memory processing, which enables SC in memory. SCRIMP can be used for a wide range of applications. It supports all SC encodings and operations in memory. It maximizes the performance and energy efficiency of implementing SC by introducing novel in-memory parallel stochastic number generation and efficient implication-based logic in memory. To show the efficiency of our stochastic architecture, we implement image processing on the proposed hardware. Saransh Gupta, Mohsen Imani, Joonseop Sim, Andrew Huang 0001, M. Hassan Najafi, Tajana Rosing |
DATE | 3 |
| 2019 | UPIM: Unipolar Switching Logic for High Density Processing-in-Memory ApplicationsabstractInternet of Things (IoT) has built a network with billions of connected devices which generate massive volumes of data. Processing large data on existing systems requires significant costs for data movements between processors and memory due to limited cache capacity and memory bandwidth. Processing-In-Memory (PIM) is a promising solution to address the issue. Prior techniques that enable the computation in non-volatile memory (NVM) are designed on a bipolar switching mode, which suffers from a high sneak current in a crossbar array (CBA) structure. In this paper, we propose a unipolar-switching logic for high-density PIM applications, called UPIM. Our design exploits a unipolar-switching mode of memristor devices which can be operated in 1D1R structure hence suppresses the sneak current that exists in prior PIM technologies. Moreover, UPIM takes advantages of a 3D vertical crossbar array (CBA) structure to increase memory utilization per unit area for high-density applications. Our evaluation on a wide range of applications shows that the UPIM achieves up to 31.3× energy saving and 113.8× energy-delay product (EDP) improvement as compared to a recent GPGPU architecture. As compared to the state-of-the-art PIM design based on the bipolar switching mode, our design achieves 3.1× lower energy consumption. Joonseop Sim, Saransh Gupta, Mohsen Imani, Yeseong Kim, Tajana Rosing |
ACM Great Lakes Symposium on VLSI | 1 |
| 2017 | Enabling efficient system design using vertical nanowire transistor current mode logicabstractVertical Nanowire-FET (VNFET) is a promising candidate to succeed in industry mainstream due to its superior suppression of short-channel-effects and area efficiency. However, to design logic gates, CMOS is not an appropriate solution due to the process incompatibility with VNFET, which creates a technical challenge for mass production. In this work, we propose a novel VNFET-based logic design, called VnanoCML (Vertical Nanowire Transistor-based Current Mode Logic), which addresses the process issue while significantly improving power and performance of diverse logic designs. Unlike the CMOS-based logic, our design exploits current mode logic to overcome the fabrication issue. Furthermore, we reduce drain-to-source resistance of VnanoCML, which results in higher performance improvement without compromising the subthreshold swing. In order to show the impact of the proposed VnanoCML, we present key logic designs which are SRAM, full adder and multiplier, and also evaluate the application-level effectiveness of digital designs for image processing and mathematical computation. Our proposed design improves the fundamental circuit characteristics including output swing, delay time and power consumption compared to conventional planar MOSFET (PFET)-based circuits. Consequentially our architecture-level results show that VnanoCML can enhance the performance and power by 16.4× and 1.15×, respectively. Furthermore, we show that VnanoCML improves the energy-delay product by 38.5× on average compared to PFET-based designs. Joonseop Sim, Mohsen Imani, Yeseong Kim, Tajana Rosing |
VLSI-SoC | 1 |