Joungmin Park

dblp:340/6226 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2026 WPU: A Pipelined WebAssembly Processing Unit for Embedded IoT Systems
abstract
WebAssembly (WASM) is a portable, stack-based virtual machine designed to execute high-level languages such as C, C++, and Rust in modern computing platforms. However, its performance on embedded IoT systems remains significantly limited due to software interpretation overhead and resource constraints. This paper presents the WebAssembly Processing Unit (WPU), a dedicated hardware accelerator tailored for WASM execution in embedded environments. The WPU features a five-stage pipelined processor integrated with a Cortex-M0 core, supporting 94 WASM instructions across integer, floating-point, and control flow operations. The architecture supports both 32-bit and 64-bit data types, and includes LEB128 decoding with a specialized register file structure to match WASM’s stack-based execution model. To address pipeline hazards, a custom forwarding mechanism and branch resolution logic are introduced, optimized for WASM’s control structure. The design is implemented on an FPGA and achieves an average of $2.9 \times$ speedup over interpreter-based runtimes, and up to $95 \times$ acceleration over JIT-based runtimes in selected benchmarks. Evaluation across diverse benchmark applications confirms that the WPU provides reliable, low-latency performance. This work establishes WPU as a scalable and practical solution for accelerating WASM workloads in embedded systems.
Jinyeol Kim, Chaebin Lee, Jongwon Oh, Joungmin Park
ASP-DAC4
2025 An Accelerated Block Searching Approach in A* for Autonomous Mobile Robots
abstract
Path planning is crucial to ensure the safe navigation of autonomous mobile robots (AMRs). However, as the scale of maps and paths increases, achieving faster and more efficient computations becomes challenging due to constraints in memory and computational resources. In this paper, we present a hardware architecture for global path planning utilizing block searching A* (BSA*) for AMRs. The BSA* designates surrounding nodes as blocks and searching nodes for expansion with collision detection based on block-based jump point search (JPS(B)). This reduces the amount of stored data, enables faster map scanning, and provides advantages in parallelized architecture. The BSA* accelerator, consisting of a memory controller based on heap sorting and a parallelized collision detector, was implemented on a field-programmable gate array (FPGA). In the benchmarks for grid-based pathfinding, experimental results showed that BSA* stored 83.2% less data compared to A* and BSA* accelerator demonstrated real-time performance, ranging from 5.498 ms (181 Hz) to 6.126 ms (164 Hz), successfully verifying the feasibility of real-time path planning with a block searching approach.
Jinyoung Shin, Joungmin Park, Jinyeol Kim, Yue Ri Jeong, Seongmo An
ISCAS2
2025 SEAM: A synergetic energy-efficient approximate multiplier for application demanding substantial computational resources
Youngwoo Jeong, Joungmin Park, Raehyeong Kim
Integr.2
2025 An Accelerated Block Searching Approach in A* for Autonomous Mobile Robots
abstract
Path planning is crucial to ensure the safe navigation of autonomous mobile robots (AMRs). However, as the scale of maps and paths increases, achieving faster and more efficient computations becomes challenging due to constraints in memory and computational resources. In this paper, we present a hardware architecture for global path planning utilizing block searching A* (BSA*) for AMRs. The BSA* designates surrounding nodes as blocks and searching nodes for expansion with collision detection based on block-based jump point search (JPS(B)). This reduces the amount of stored data, enables faster map scanning, and provides advantages in parallelized architecture. The BSA* accelerator, consisting of a memory controller based on heap sorting and a parallelized collision detector, was implemented on a field-programmable gate array (FPGA). In the benchmarks for grid-based pathfinding, experimental results showed that BSA* stored 83.2% less data compared to A* and BSA* accelerator demonstrated real-time performance, ranging from 0.528 ms to 2.983 ms, successfully verifying the feasibility of real-time path planning with a block searching approach.
Jinyoung Shin, Joungmin Park, Jinyeol Kim, Yue Ri Jeong, Seongmo An
IEEE Trans. Circuits Syst. I Regul. Pap.2