EDBT 2026 Demo / reviewers in the wild / expert
Sung-Joon Jang
dblp:48/2886
· DBLP profile ↗
5ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-6992-408XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High-throughput Point-Cloud Accelerator with Sparsity-aware Hierarchical Neighbor Voxel Search and SkippingabstractPoint cloud-based 3D sparse convolution networks are widely employed to process voxel features efficiently. However, the irregularity of voxel sparsity poses significant challenges, leading to increased hardware complexity and inefficiencies. We propose an algorithm-hardware co-design for sparse 3D convolution. At the algorithm level, an on-the-fly thresholdbased voxel skipping is adopted, enhancing efficiency. At the hardware level, a hierarchical 3-stage Voxel Search and Skipping is developed to systematically narrow down the non-zero search space, enhancing both performance and hardware utilization. We implemented the proposed accelerator in a 65 nm process to demonstrate a 77.7% reduction in delay compared to a baseline design, which does not support the proposed sparsity adaptations. The proposed system also achieved the $1.34 \times$ and $2.22 \times$ higher energy efficiency and throughput as compared to the state-ofarts. Yun-Chia Yu, Suraj Pn Reddy, Aryan Devrani, Anirudh Srinivasan, Saianudeep Reddy Nayini, Sohyeon Kim, Sung-Joon Jang, Sang-Seol Lee, Mingu Kang |
DAC | 7 |
| 2024 | HAIL-DIMM: Host Access Interleaved with Near-Data Processing on DIMM-based Memory SystemabstractNear-data processing (NDP), a solution to reduce data movement overhead between host and memory, should not interfere with host access to ensure system fairness. We propose a cost-effective and energy-efficient LRDIMM-based NDP architecture (HAIL-DIMM) that can seamlessly interleave NDP and host access and is a drop-in replacement for existing main memory modules. The proposed NDP exploits the interleaving capability of the memory controller to interleave NDP and host access naturally. To take advantage of bank interleaving, an atomic operation of the proposed NDP, which consists of data movement and computation, is recognized by the memory controller as a DDR READ/WRITE but by the HAIL-DIMM as NDP based on the request's address. We implement a prototype of the proposed NDP architecture on an FPGA platform as proof of concept. The evaluation results show that the NDP system achieves up to 2.19x speedup in latency and up to 45.4% energy saving for data movement over the baseline system for memory-bound workloads. Sang-Seol Lee, Kyungho Kim, Eunchong Lee, Sung-Joon Jang |
DAC | 5 |
| 2024 | Skew-CIM: Process-Variation-Resilient and Energy-Efficient Computation-in-Memory Design Technique With Skewed WeightsabstractIn analog-mixed-signal (AMS) compute-in-memory (CIM) systems, the two’s-complement (2SC) format provides better area efficiency than the sign-and-magnitude (SNM) one. However, the 2SC format exacerbates the challenges of AMS-CIM systems, suffering from significant DNN accuracy drop under process variations and high computation currents from activating multiple WLs. In the 2SC format, ‘0’ and ‘1’ are nearly balanced for all logical-order bits, unlike ‘0’-skewed higher-order bits in the SNM format. Consequently, the 2SC-based AMS-CIM systems have much more on-cells than the SNM-based counterpart, deteriorating the above challenges. We propose Skew-CIM, a software-hardware co-design technique to relax these challenges. Our proposed weight skewing (WESK) breaks the ‘0’ and ‘1’ balance at the software level. The potential accuracy drops resulting from WESK are successfully compensated by retraining DNNs. The offsets caused by WESK can be easily corrected using online hardware-level processing. Our Skew-CIM technique can be applied to most AMS-CIM systems with memories showing large on-off cell current ratios. As an example, we use it in a custom-designed 8T-SRAM-based CIM device, demonstrating a significant reduction in the DNN classification error by 7.6 times compared to the 2SC-based AMS-CIM without our Skew-CIM technique. Furthermore, our Skew-CIM markedly enhances energy efficiency by up to 39.9%, outperforming conventional SNM-based AMS-CIM systems. Donghyeon Yi, Injun Choi, Gichan Yun, Edward Choi 0001, Jonghee Park, Jonghoon Kwak, Sung-Joon Jang, Sohmyung Ha, Ik Joon Chang, Minkyu Je |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2020 | Resource-Efficient and High-Throughput VLSI Design of Global Optical Flow Method for Mobile SystemsabstractMost very-large-scale integration (VLSI) designs of the global optical flow method focus on reducing external memory accesses due to global and iterative processes for low-power operation in mobile systems. However, to achieve this goal, considerable amounts of resources are used and the throughput is decreased. To address these issues, we propose a multirow-based propagation approach and an efficient VLSI architecture. This approach divides an image into multiple small subimages. The flow of each subimage is then estimated sequentially using a small amount of internal memory without accessing any external memory. To avoid discontinuity artifacts stemming from the boundaries of the subimages, boundary conditions are imposed based on the smoothness of the optical flow. In addition, the main equations of the global optical flow method are simplified to boost the processing speed. The experimental results demonstrate that the external memory bandwidth is reduced by 96.7% with negligible accuracy degradation. Furthermore, the proposed design uses 63.4%-98% less internal memory and achieves $1.9\times $ -$18.8\times $ higher throughput-per-watt compared to state-of-the-art designs of the global optical flow method. Sung-Joon Jang, Chong-Min Kyung |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | Cache Miss-Aware Dynamic Stack AllocationabstractReducing cache misses without increasing cache associativity is critical for reducing the power consumption and cache access time. This paper has focused on the stack of a program which often occupies more than half of total memory accesses (Calder, et. al., 1998). This paper, as a result, proposes so-called dynamic stack allocation where the stack pointer is shifted at run time to a memory location which is expected to cause least number of cache misses. We implemented the proposed scheme using so-called dynamic stack allocator (DSA) which consists of cache miss predictor (CMP) to compute cache miss probability based on least recently used (LRU) policy and stack pointer manager (SPM) to manage multiple stack locations. We also verified the proposed scheme with both FPGA and ASIC by using iNCITE and Dong-Bu electronics 0.18mum process, respectively. Experimental results show that dynamic stack allocation significantly reduces cache misses from 4% to 42% in various benchmarks with relatively small power consumption and no extra delay. Sung-Joon Jang, Moo-Kyoung Chung, Jaemoon Kim, Chong-Min Kyung |
ISCAS | 1 |