Seungwoo Hong

dblp:94/1560 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice Sparsity
abstract
Low bit-precisions and their bit-slice sparsity have recently been studied to accelerate general matrix-multiplications (GEMM) during large-scale deep neural network (DNN) inferences. While the conventional symmetric quantization facilitates low-resolution processing with bit-slice sparsity for both weight and activation, its accuracy loss caused by the activation’s asymmetric distributions cannot be acceptable, especially for largescale DNNs. In efforts to mitigate this accuracy loss, recent studies have actively utilized asymmetric quantization for activations without requiring additional operations. However, the cuttingedge asymmetric quantization produces numerous nonzero slices that cannot be compressed and skipped by recent bit-slice GEMM accelerators, naturally consuming more processing energy to handle the quantized DNN models.To simultaneously achieve high accuracy and hardware efficiency for large-scale DNN inferences, this paper proposes an Asymmetrically-Quantized bit-Slice GEMM (AQS-GEMM) for the first time. In contrast to the previous bit-slice computing, which only skips operations of zero slices, the AQS-GEMM compresses frequent nonzero slices, generated by asymmetric quantization, and skips their operations. To increase the slicelevel sparsity of activations, we also introduce two algorithm-hardware co-optimization methods: a zero-point manipulation and a distribution-based bit-slicing. To support the proposed AQS-GEMM and optimizations at the hardware-level, we newly introduce a DNN accelerator, Panacea, which efficiently handles sparse/dense workloads of the tiled AQS-GEMM to increase data reuse and utilization. Panacea supports a specialized dataflow and run-length encoding to maximize data reuse and minimize external memory accesses, significantly improving its hardware efficiency. Numerous benchmark evaluations show that Panacea outperforms existing DNN accelerators, e.g., $1.97 \times$ and $3.26 \times$ higher energy efficiency, and $1.88 \times$ and $2.41 \times$ higher throughput than the recent bit-slice accelerator Sibia and the SIMD design, respectively, on OPT-2.7B, while providing better algorithm performance with asymmetric quantization.
Dongyun Kam, Myeongji Yun, Sunwoo Yoo, Seungwoo Hong, Zhengya Zhang, Youngjoo Lee 0002
HPCA4
2025 A Lightweight ML-Based ECG Classification System Using Self-Personalized Anomaly Detector
abstract
Targeting the real-time arrhythmia diagnosis on resource-limited edge devices, in this paper, we present a lightweight electrocardiogram classification system using event-driven machine learning processing. A self-personalized anomaly detector based on signal processing is newly developed to dynamically update internal decision criteria from each patient's recent electrocardiogram history, that activates the following machine learning model only for the abnormal cases. A Siamese neural network is adopted to identify detailed arrhythmia classes by comparing features from the self-personalized normal data and the current abnormal input, increasing the classification accuracy. We also develop a simple version of our Siamese model to reduce the number of trainable parameters while preserving the end-to-end classification accuracy. Experimental results show that the proposed event-driven system reduces ML model activations by 74% for normal beats, achieving a classification accuracy of 96.9% comparable to leading solutions. Additionally, it consumes three times less energy and achieves 3.6 times faster processing latency compared to cost-aware method on a mobile GPU platform, enabling extended battery life and real-time analysis on edge devices.
Sunwoo Yoo, Seungwoo Hong, Dongyun Kam, Youngjoo Lee 0002
IEEE J. Biomed. Health Informatics2
2024 Integrating Model-Based Footstep Planning with Model-Free Reinforcement Learning for Dynamic Legged Locomotion
abstract
In this work, we introduce a control framework that combines model-based footstep planning with Reinforcement Learning (RL), leveraging desired footstep patterns derived from the Linear Inverted Pendulum (LIP) dynamics. Utilizing the LIP model, our method forward predicts robot states and determines the desired foot placement given the velocity commands. We then train an RL policy to track the foot placements without following the full reference motions derived from the LIP model. This partial guidance from the physics model allows the RL policy to integrate the predictive capabilities of the physics-informed dynamics and the adaptability characteristics of the RL controller without overfitting the policy to the template model. Our approach is validated on the MIT Humanoid, demonstrating that our policy can achieve stable yet dynamic locomotion for walking and turning. We further validate the adaptability and generalizability of our policy by extending the locomotion task to unseen, uneven terrain. During the hardware deployment, we have achieved forward walking speeds of up to 1.5 m/s on a treadmill and have successfully performed dynamic locomotion maneuvers such as 90-degree and 180-degree turns.
Ho Jae Lee, Seungwoo Hong, Sangbae Kim
IROS2
2024 Rethinking Explicit Congestion Notification: A Multilevel Congestion Feedback Perspective
abstract
Congestion in the network is a persistent issue that is becoming more challenging with the advent of technologies such as metaverse and immersive AR/VR applications. Therefore, we propose Enhanced ECN (EECN), a novel network-assisted congestion feedback protocol that redesigns the legacy ECN. EECN notifies two congestion levels encoded in the existing two ECN bits, unlike ECN which delivers only Boolean information about the congestion. Additionally, we propose a congestion control algorithm that leverages this multilevel congestion feedback from the network using EECN. Our proposed EECN feedback mechanism can coexist with the legacy ECN and requires minimal changes in the end hosts. Moreover, the proposed congestion control mechanism reduces packet drops by 74% compared to ECN with TCP New Reno and 96% compared to TCP New Reno without ECN. The marked packets are reduced by 30% compared to ECN. Furthermore, the proposed approach enhances the flow completion time of short-lived network flows.
Inayat Ali, Seungwoo Hong, PyungKoo Park, Tae-Yeon Kim 0003
NOSSDAV2
2022 Design of KAIST HOUND, a Quadruped Robot Platform for Fast and Efficient Locomotion with Mixed-Integer Nonlinear Optimization of a Gear Train
abstract
This paper introduces a design method for an efficient and agile quadruped robot. A mixed-integer optimization formulation including the number of gear teeth is derived to obtain the optimal gear ratio that minimizes cost for a running-trot with the target speed of 3 m/s. With the inclusion of integer constraints related to the number of gear teeth, detailed design considerations of gear trains can be included in the optimization process. Thermal dissipation of the motor controller is also taken into account in the optimization to consider heat generation during high-speed running. KAIST Hound, a 45 kg robot, designed with the obtained design parameters has successfully demonstrated a 3 m/s running-trot using a nonlinear model predictive controller (NMPC). Furthermore, the robot has proved its robustness by the demonstration of additional experiments such as 22° slope climbing, 3.2 km walking, and traversing a 35 cm obstacle.
Young-Ha Shin, Seungwoo Hong, Sangyoung Woo, Jonghun Choe, Harim Son, Gijeong Kim, Joon-Ha Kim, Kang Kyu Lee, Jemin Hwangbo, Hae-Won Park 0002
ICRA2
2022 DRPD, Dual Reduction Ratio Planetary Drive for Articulated Robot Actuators
abstract
This paper presents a reduction mechanism for robot actuators that can switch between two types of reduction ratio. By fixing the carrier or ring gear of the proposed actuator which is based on the 3K compound planetary drive, the actuator can shift its reduction ratio. For compact design with reduced weight of the actuator, unique pawl brake mechanism interacting with cams and micro servos for switching mechanism is designed. The resulting prototype module has a reduction ratio of 6.91 and 44.93 for ‘low-reduction’ and ‘high-reduction’ ratios, respectively. Reduction ratios can be easily adjusted by modifying the pitch diameters of gears. Experimental results demonstrate that the proposed actuator could extend its operation region via two reduction modes that are interchangeable with gear shifting.
Tae-Gyu Song, Young-Ha Shin, Seungwoo Hong, Hyungho Chris Choi, Joon-Ha Kim, Hae-Won Park 0002
IROS3
2022 Low-Complexity and Low-Latency SVC Decoding Architecture Using Modified MAP-SP Algorithm
abstract
The compressive sensing (CS) based sparse vector coding (SVC) method is one of the promising ways for the next-generation ultra-reliable and low-latency communications. In this paper, we present advanced algorithm-hardware co-optimization schemes for realizing a cost-effective SVC decoding architecture. The previous maximum a posteriori subspace pursuit (MAP-SP) algorithm is newly modified to relax the computational overheads by applying novel residual forwarding and LLR approximation schemes. A fully-pipelined parallel hardware is also developed to support the modified decoding algorithm, reducing the overall processing latency, especially at the support identification step. In addition, an advanced least-square-problem solver is presented by utilizing the parallel Cholesky decomposer design, further reducing the decoding latency with parallel updates of support values. The implementation results from a 22nm FinFET technology showed that the fully-optimized design is 9.6 times faster while improving the area efficiency by 12 times compared to the baseline realization.
Seungwoo Hong, Dongyun Kam, Sangbu Yun, Jeongwon Choe, Namyoon Lee, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 Real-Time Constrained Nonlinear Model Predictive Control on SO(3) for Dynamic Legged Locomotion
abstract
This paper presents a constrained nonlinear model predictive control (NMPC) framework for legged locomotion. The framework assumes a legged robot as a floating base single rigid body with contact forces being applied to the body as external forces. With consideration of orientation dynamics evolving on the rotation manifold SO(3), analytic Jacobians which are necessary for constructing the gradient and the Gauss-Newton Hessian approximation of the objective function are derived. This procedure also includes the reparameterization of the robot orientation on SO(3) to orientation error in the tangent space of that manifold. Obtained gradient and Gauss-Newton Hessian approximation are utilized to solve nonlinear least squares problems formulated from NMPC in a computationally efficient manner. The proposed algorithm is verified on various types of legged robots and gaits in a simulation environment.
Seungwoo Hong, Joon-Ha Kim, Hae-Won Park 0002
IROS1
2018 Implementing Full-body Torque Control in Humanoid Robot with High Gear Ratio Using Pulse Width Modulation Voltage
abstract
Most state-of-the-art torque control-based legged robots show excellent performance, exceeding that of conventional position control-based robots. Many conventional position control-based legged robots have high gear ratios, but do not have joint torque sensors. In addition, some robots cannot generate current for controlling the motor torque. To apply torque control-based walking algorithms to a position control-based humanoid robot, we proposed current control using a motor thermal model and realized joint torque control by compensating for the joint dynamics and robot dynamics. We conducted experiments to verify the performance of the Hubo2 platform developed in 2008 by applying a full-body dynamics control framework. The results confirmed the possibility of using torque control algorithms with existing position-based robots.
Kang Kyu Lee, Okkee Sim, Hyobin Jeong, Jaesung Oh, Hyoin Bae, Seungwoo Hong, Jun-Ho Oh
IROS6