Jungju Oh

dblp:16/9657 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 46% Processor architecture and microarchitecture · 23% Interconnection networks and networks-on-chip · 23%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › synchronization
barrier synchronization
0.112011
TLSync: support for multiple fast barriers using on-chip transmission lines · ISCA 2011
Processor architecture and microarchitecture
multicore design
0.112011
TLSync: support for multiple fast barriers using on-chip transmission lines · ISCA 2011
Parallel and multicore computing › parallel computing
parallel program analysis
0.112011
LIME: a framework for debugging load imbalance in multi-threaded execution · ICSE 2011
Performance modeling and evaluation › profiling
statistical profiling
0.012011
LIME: a framework for debugging load imbalance in multi-threaded execution · ICSE 2011

Methods — techniques the papers use, named apart from their topics

tournament barrier · 0.1statistical analysis · 0.1reduction tree · 0.1notification tree · 0.1
YearPublicationVenuePosition
2025 CTDM: Resource-Efficient FPGA-Accelerated Simulation of Large-Scale NPU Designs
abstract
This paper proposes a novel approach to accelerate large Neural Processing Unit (NPU) simulations on FPGA through Chain-based Time-Division Multiplexing (CTDM) and its automatic compiler. CTDM replaces repeated logic patterns with a single logic pattern and register chains, which can take advantage of built-in shift register primitives. It reduces FPGA resource utilization more effectively than conventional multiplexer-based Time-Division Multiplexing (TDM) approaches by minimizing logic overhead and routing congestion. The automated CTDM compiler supports various hardware design languages (HDL) including Verilog, VHDL, high-level synthesis (HLS), and Chisel, as well as a wide range of FPGA devices—from small on-premise boards to server-grade hardware simulators like Synopsys ZeBu. To extend the applicability of CTDM to multi-FPGA systems, we propose a block interleaving technique that hides inter-FPGA link latency and fully utilizes the pipeline in a high-speed serial I/O channel. When applied to NVIDIA Deep Learning Accelerator (NVDLA), CTDM achieved a 66% and 82% reduction in LUT and FF utilization, respectively, and enabled the successful deployment of the largest variant of NVDLA on a single AMD U250 FPGA device. This demonstrated a 3,653× acceleration in NVDLA simulation time over the Synopsys VCS simulator on a CPU. This method has already been implemented for the simulation and verification of our proprietary NPUs. Notably, it enabled the simulation of a 4-die 1024 TFLOPS chiplet using 144 FPGAs on ZeBu 5 server.
Hyunje Jo, Han-Sok Suh, Hyun-Seok Heo, Jinseok Kim 0006, Hyunsung Kim 0003, Boeui Hong, Jungju Oh, Sunghyun Park 0006, Jinwook Oh, Sunghwan Jo, Kangwook Lee 0008, Jae-sun Seo
ICCAD7
2013 Traffic steering between a low-latency unswitched TL ring and a high-throughput switched on-chip interconnect
abstract
Growth in core count creates an increasing demand for interconnect bandwidth, driving a change from shared buses to packet-switched on-chip interconnects. However, this increases the latency between cores separated by many links and switches. In this paper, we show that a low-latency unswitched interconnect built with transmission lines can be synergistically used with a high-throughput switched interconnect. First, we design a broadcast ring as a chain of unidirectional transmission line structures with very low latency but limited throughput. Then, we create a new adaptive packet steering policy that judiciously uses the limited throughput of this ring by balancing expected latency benefit and ring utilization. Although the ring uses 1.3% of the on-chip metal area, our experimental results show that, in combination with our steering, it provides an execution time reduction of 12.4% over a mesh-only baseline.
Jungju Oh, Alenka G. Zajic, Milos Prvulovic
PACT1
2011 LIME: a framework for debugging load imbalance in multi-threaded execution
abstract
With the ubiquity of multi-core processors, software must make effective use of multiple cores to obtain good performance on modern hardware. One of the biggest roadblocks to this is load imbalance, or the uneven distribution of work across cores. We propose LIME, a framework for analyzing parallel programs and reporting the cause of load imbalance in application source code. This framework uses statistical techniques to pinpoint load imbalance problems stemming from both control flow issues (e.g., unequal iteration counts) and interactions between the application and hardware (e.g., unequal cache miss counts). We evaluate LIME on applications from widely used parallel benchmark suites, and show that LIME accurately reports the causes of load imbalance, their nature and origin in the code, and their relative importance.
Jungju Oh, Christopher J. Hughes, Guru Venkataramani, Milos Prvulovic
ICSE1
2011 TLSync: support for multiple fast barriers using on-chip transmission lines
abstract
As the number of cores on a single-chip grows, scalable barrier synchronization becomes increasingly difficult to implement. In software implementations, such as the tournament barrier, a larger number of cores results in a longer latency for each round and a larger number of rounds. Hardware barrier implementations require significant dedicated wiring, e.g., using a reduction (arrival) tree and a notification (release) tree, and multiple instances of this wiring are needed to support multiple barriers (e.g., when concurrently executing multiple parallel applications).
Jungju Oh, Milos Prvulovic, Alenka G. Zajic
ISCA1