EDBT 2026 Demo / reviewers in the wild / expert
Hyunjun Cho
dblp:206/3662
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Anisotropic Cross-View Texture Transfer With Multi-Reference Non-Local Attention for CT Slice InterpolationabstractComputed tomography (CT) is one of the most widely used non-invasive imaging modalities for medical diagnosis. In clinical practice, CT images are usually acquired with large slice thicknesses due to the high cost of memory storage and operation time, resulting in an anisotropic CT volume with much lower inter-slice resolution than in-plane resolution. Since such inconsistent resolution may lead to difficulties in disease diagnosis, deep learning-based volumetric super-resolution methods have been developed to improve inter-slice resolution. Most existing methods conduct single-image super-resolution on the through-plane or synthesize intermediate slices from adjacent slices; however, the anisotropic characteristic of 3D CT volume has not been well explored. In this paper, we propose a novel cross-view texture transfer approach for CT slice interpolation by fully utilizing the anisotropic nature of 3D CT volume. Specifically, we design a unique framework that takes high-resolution in-plane texture details as a reference and transfers them to low-resolution through-plane images. To this end, we introduce a multi-reference non-local attention module that extracts meaningful features for reconstructing through-plane high-frequency details from multiple in-plane images. Through extensive experiments, we demonstrate that our method performs significantly better in CT slice interpolation than existing competing methods on public CT datasets including a real-paired benchmark, verifying the effectiveness of the proposed framework. The source code of this work is available at https://github.com/khuhm/ACVTT. Kwang-Hyun Uhm, Hyunjun Cho, Sung-Hoo Hong, Seung-Won Jung |
IEEE Trans. Medical Imaging | 2 |
| 2025 | ABC-FHE: A Resource-Efficient Accelerator Enabling Bootstrappable Parameters for Client-Side Fully Homomorphic EncryptionabstractAs the demand for privacy-preserving computation continues to grow, fully homomorphic encryption (FHE)—which enables continuous computation on encrypted data—has become a critical solution. However, its adoption is hindered by significant computational overhead, requiring 10000-fold more computation compared to plaintext processing. Recent advancements in FHE accelerators have successfully improved server-side performance, but client-side computations remain a bottleneck, particularly under bootstrappable parameter configurations, which involve combinations of encoding, encrypt, decoding, and decrypt for large-sized parameters. To address this challenge, we propose ABC-FHE, an area- and power-efficient FHE accelerator that supports bootstrappable parameters on the client side. ABCFHE employs a streaming architecture to maximize performance density, minimize area usage, and reduce off-chip memory access. Key innovations include a reconfigurable Fourier engine capable of switching between NTT and FFT modes. Additionally, an onchip pseudo-random number generator and a unified on-the-fly twiddle factor generator significantly reduce memory demands, while optimized task scheduling enhances the CKKS clientside processing, achieving reduced latency. Overall, ABC-FHE occupies a die area of 28.638 mm2 and consumes 5.654 W of power in 28 nm technology. It delivers significant performance improvements, achieving a 1112× speed-up in encoding and encryption execution time compared to a CPU, and 214× over the state-of-the-art client-side accelerator. For decoding and decryption, it achieves a 963× speed-up over the CPU and 82× over the state-of-the-art accelerator. Sungwoong Yune, Adiwena Putra, Hyunjun Cho, Cuong Duong Manh, Joo-Young Kim 0001 |
DAC | 4 |
| 2025 | SAL-PIM: A Subarray-Level Processing-in-Memory Architecture With LUT-Based Linear Interpolation for Transformer-Based Text GenerationabstractText generation is a compelling sub-field of natural language processing, aiming to generate human-readable text from input words. Although many deep learning models have been proposed, the recent emergence of transformer-based large language models advances its academic research and industry development, showing remarkable qualitative results in text generation. In particular, the decoder-only generative models, such as generative pre-trained transformer (GPT), are widely used for text generation, with two major computational stages: summarization and generation. Unlike the summarization stage, which can process the input tokens in parallel, the generation stage is difficult to accelerate due to its sequential generation of output tokens through iteration. Moreover, each iteration requires reading a whole model with little data reuse opportunity. Therefore, the workload of transformer-based text generation is severely memory-bound, making the external memory bandwidth system bottleneck. In this paper, we propose a subarray-level processing-in-memory (PIM) architecture named SAL-PIM, the first HBM-based PIM architecture for the end-to-end acceleration of transformer-based text generation. With optimized data mapping schemes for different operations, SAL-PIM utilizes higher internal bandwidth by integrating multiple subarray-level arithmetic logic units (S-ALUs) next to memory subarrays. To minimize the area overhead for S-ALU, it uses shared MACs leveraging slow clock frequency of commands for the same bank. In addition, a few subarrays in the bank are used as look-up tables (LUTs) to handle non-linear functions in PIM, supporting multiple addressing to select sections for linear interpolation. Lastly, the channel-level arithmetic logic unit (C-ALU) is added in the buffer die of HBM to perform the accumulation and reduce-sum operations of data across multiple banks, completing end-to-end inference on PIM. To validate the SAL-PIM architecture, we built a cycle-accurate simulator based on Ramulator. We also implemented the SAL-PIM’s logic units in 28-nm CMOS technology and scaled the results to DRAM technology to verify its feasibility. We measured the end-to-end latency of SAL-PIM when it runs various text generation workloads on the GPT-2 medium model (with 345 million parameters), in which the input and output token numbers vary from 32 to 128 and from 1 to 256, respectively. As a result, with 4.81% area overhead, SAL-PIM achieves up to 4.72× speedup (1.83× on average) over the Nvidia Titan RTX GPU running FasterTransformer Framework. Wontak Han, Hyunjun Cho, Joo-Young Kim 0001 |
IEEE Trans. Computers | 2 |
| 2025 | A Nuclei-Focused Strategy for Automated Histopathology Grading of Renal Cell CarcinomaabstractThe rising incidence of kidney cancer underscores the need for precise and reproducible diagnostic methods. In particular, renal cell carcinoma (RCC), the most prevalent type of kidney cancer, requires accurate nuclear grading for better prognostic prediction. Recent advances in deep learning have facilitated end-to-end diagnostic methods using contextual features in histopathological images. However, most existing methods focus only on image-level features or lack an effective process for aggregating nuclei prediction results, limiting their diagnostic accuracy. In this paper, we introduce a novel framework, Nuclei feature Assisted Patch-level RCC grading (NuAP-RCC), that leverages nuclei-level features for enhanced patch-level RCC grading. Our approach employs a nuclei-level RCC grading network to extract grade-aware features, which serve as node features in a graph. These node features are aggregated using graph neural networks to capture the morphological characteristics and distributions of the nuclei. The aggregated features are then combined with global image-level features extracted by convolutional neural networks, resulting in a final feature for accurate RCC grading. In addition, we present a new dataset for patch-level RCC grading. Experimental results demonstrate the superior accuracy and generalizability of NuAP-RCC across datasets from different medical institutions, achieving a 6.15% improvement in accuracy over the second-best model on the USM-RCC dataset. Hyunjun Cho, Dongjin Shin, Kwang-Hyun Uhm, Sung-Jea Ko, Yosep Chong, Seung-Won Jung |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | APINT: A Full-Stack Framework for Acceleration of Privacy-Preserving Inference of Transformers based on Garbled CircuitsabstractAs the importance of Privacy-Preserving Inference of Transformers (PiT) increases, a hybrid protocol that integrates Garbled Circuits (GC) and Homomorphic Encryption (HE) is emerging for its implementation. While this protocol is preferred for its ability to maintain accuracy, it has a severe drawback of excessive latency. To address this, existing protocols primarily focused on reducing HE latency, thus making GC the new latency bottleneck. Furthermore, previous studies only focused on individual computing layers, such as protocol or hardware accelerator, lacking a comprehensive solution at the system level. Hyunjun Cho, Jaehoon Heo, Joo-Young Kim 0001 |
ICCAD | 1 |
| 2024 | PD-CR: Patch-Based Diffusion Using Constrained Refinement for Image RestorationabstractDiffusion models, which are state-of-the-art generative models, have been widely applied to image restoration tasks. However, most image restoration methods based on diffusion models require a large amount of computational memory, making it difficult to use them with high-resolution images. Although patch-based diffusion models have emerged to address this problem, these models are limited in effectively mitigating boundary artifacts and producing results close to the ground truth. In this paper, we propose Patch-based Diffusion using Constrained Refinement (PD-CR) that refines the noise estimated by patch-based diffusion models to produce a restored image while keeping the luminance of the input degraded image. Leveraging patch-based diffusion models, the proposed method can handle a high-resolution image as input with minimal memory requirements. Our experiments on various image restoration tasks, such as image denoising and raindrop removal, demonstrate that the proposed method is better than or on par with the state-of-the-art methods. Hyunjun Cho, Hong-Kyu Shin, Yurim Jang, Sung-Jea Ko, Seung-Won Jung |
IEEE Signal Process. Lett. | 1 |
| 2024 | Revisiting PID Control for Power-Constrained Video DisplayabstractDue to the significant power consumption of emissive display panels and the limited battery capacity of electronic devices, it is necessary to implement power-constrained display operations for video content. One of the effective power-constrained display techniques is power-constrained contrast enhancement (PCCE), which aims to reduce the power demands of the display while maintaining the quality of the content. However, PCCE has been mainly studied in the image domain. In this letter, we explore how PCCE can be applied to the video domain by revisiting the proportional-integral-derivative (PID) control. The proposed method consists of three key components: 1) The proportional term determines a power-saving ratio for each video frame based on its luminance, resulting in aggressive power-saving in bright video frames; 2) the integral term ensures that the target power-saving ratio is achieved over the entire video; 3) the differential term prevents abrupt temporal fluctuations in the power-saving ratio of consecutive video frames. The experimental results demonstrate that the proposed PID-based PCCE method achieves superior performance on two public video datasets. In particular, the proposed method achieves 5.8% improvements in VMAF compared to the previous method on the TVSum dataset. Yurim Jang, Hyunjun Cho, Geon-Ho Park, Min-Jae Yoo, Seung-Won Jung |
IEEE Signal Process. Lett. | 2 |
| 2022 | Sound-Guided Semantic Video Generation
Gyeongrok Oh, Wonmin Byeon, Chanyoung Kim 0001, Wonjeong Ryoo, Sang Ho Yoon, Hyunjun Cho, Jihyun Bae, Jinkyu Kim 0001, Sangpil Kim |
ECCV (17) | 7 |