Jae-Ho Choi 0004

dblp:81/3843-4 · also Jaeho Choi 0004 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0001-9484-4869ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ReFineVQA: Iterative Refinement of Video Description via Feedback Generation for Video Question Answering
abstract
Video question answering is a non-trivial task that demands joint understanding of visual contents and linguistic questions as well as temporal reasoning across video frames. Recent agent-based approaches address this by conducting multi-step reasoning with large language models (LLMs) across frame-level captions generated by vision-language models, but encounter limited temporal coherence across frames. A possible direction based on video language models (VideoLMs) directly captures temporal dynamics via video-level descriptions, but often lacks fine-grained visual cues due to a restricted number of input frames and a large dependency on input prompts. To tackle these challenges, we propose RefineVQA, a training-free framework that can easily be plugged into existing VideoLMs with iterative, LLM-guided description refinements. Specifically, the VideoLM produces an initial description, followed by LLM feedback determining whether the description suffices for the question and guiding further visual extraction, which in turn enhances the description quality while preserving temporal context. Plugged into state-of-the-art VideoLMs, ReFineVQA yields consistent gains across diverse benchmarks–NExT-QA, EgoSchema, VideoMME, ActivityNet, and StreamingBench–even with a small external LLM of 3.8B parameters.
Jeongwan Shin, Chan Hur, Seongmin Cho, Jae-Ho Choi 0004, Hyeyoung Park
WACV4
2025 MVDoppler-Pose: Multi-Modal Multi-View mmWave Sensing for Long-Distance Self-Occluded Human Walking Pose Estimation
abstract
One of the main challenges in reliable camera-based 3D pose estimation for walking subjects is to deal with self-occlusions, especially in the case of using low-resolution cameras or at longer distance scenarios. In recent years, millimeter-wave (mmWave) radar has emerged as a promising alternative, offering inherent resilience to the effect of occlusions and distance variations. However, mmWave-based human walking pose estimation (HWPE) is still in the nascent development stages, primarily due to its unique set of practical challenges including the quality of the observed radar signal dependent on the subject’s motion direction. This paper introduces the first comprehensive study comparing mmWave radar to camera systems for HWPE, highlighting its utility for distance-agnostic and occlusion-resilient pose estimation. Building upon mmWave’s unique advantages, we address its intrinsic directionality issue through a new approach—the synergetic integration of multi-modal, multi-view mmWave signals, achieving robust HWPE against variations both in distance and walking direction. Extensive experiments on a newly curated dataset not only demonstrate the superior potential of mmWave technology over traditional camera-based HWPE systems, but also validate the effectiveness of our approach in over-coming the core limitations of mmWave HWPE.
Jae-Ho Choi 0004, Soheil Hor, Shubo Yang 0002, Amin Arbabian
CVPR1
2025 High-Resolution Gait Micro-Doppler Synthesis from Videos Over Diverse Trajectories
abstract
In recent years, there has been increasing interest in human body motion analysis, with applications in activity classification, human gait analysis, and human intent recognition. Among the various methods, millimeter-wave (mmWave)-based approaches have become popular due to their inherent contactless and privacy-preserving properties, as well as their resilience to lighting, weather conditions, and measurement distance. However, mmWave-based human motion analysis is still in its early stages, primarily due to the limited availability of large-scale datasets. This is particularly true for tasks that require high-resolution motion measurements. In this work, we introduce a novel method for synthesizing high-resolution mmWave datasets directly from videos. The proposed approach is well-suited for applications requiring very high-resolution Doppler signature simulations, such as analyzing subtle hand motions of walking pedestrians. We achieve this by employing an adversarial training strategy combined with custom task-specific loss functions that enhance the micro-motion signatures in the hands and legs of walking pedestrians. This work is the first to design and validate a high-resolution synthesized Doppler dataset of walking activities across multiple trajectories and subjects.
Shubo Yang 0002, Soheil Hor, Jae-Ho Choi 0004, Amin Arbabian
ICASSP3
2024 Fusion-Vital: Video-RF Fusion Transformer for Advanced Remote Physiological Measurement
abstract
Remote physiology, which involves monitoring vital signs without the need for physical contact, has great potential for various applications. Current remote physiology methods rely only on a single camera or radio frequency (RF) sensor to capture the microscopic signatures from vital movements. However, our study shows that fusing deep RGB and RF features from both sensor streams can further improve performance. Because these multimodal features are defined in distinct dimensions and have varying contextual importance, the main challenge in the fusion process lies in the effective alignment of them and adaptive integration of features under dynamic scenarios. To address this challenge, we propose a novel vital sensing model, named Fusion-Vital, that combines the RGB and RF modalities through the new introduction of pairwise input formats and transformer-based fusion strategies. We also perform comprehensive experiments based on a newly collected and released remote vital dataset comprising synchronized video-RF sensors, showing the superiority of the fusion approach over the previous single-sensor baselines in various aspects.
Jae-Ho Choi 0004, Ki-Bong Kang, Kyung-Tae Kim
AAAI1
2024 RF-Vital: Radio-Based Contactless Respiration Monitoring for a Moving Individual
abstract
The noncontact respiration rate measurement (nRRM) method allows a system to monitor the breathing patterns of an individual without physical contact, which is crucial for regular health monitoring. Current nRRM approaches primarily depend on detecting minor variations in RGB profiles reflected from a camera to remotely extract respiration signals. However, these methods require continuous pixel-level tracking, which restricts their use on individuals in quasi-stationary sitting positions. To address this limitation, we propose a radiofrequency (RF)-Vital model, which leverages RF signals to extend the applicability of nRRM methods to individuals who exhibit global motions (GMs) and even walk around. The core idea of the RF-Vital model lies in the unique characteristics of RF signals: the RF signals received from a moving individual capture both their respiratory motions (RMs) and GMs through linear superposition while simultaneously providing the reflections of GM alone. To fully utilize such unique properties, we introduce a new RF modality that allows stable inclusion of micro-level respiration signatures, even when GMs are present. Additionally, we optimize the RF-Vital model using a novel multitask adversarial learning framework combined with a new loss function, which facilitates the direct mapping of the desired RMs as well as the self-supervised removal of GMs, thereby effectively filtering out RMs from mixtures of GMs and RMs. The proposed RFvital model was evaluated using newly published data sets. It demonstrated state-of-the-art performance in static conditions and achieved the significant milestone of enabling nRRM under moving conditions.
Jae-Ho Choi 0004, Ki-Bong Kang, Kyung-Tae Kim
IEEE Internet Things J.1
2024 Radar-Based Crowd Counting in Real-World Environments With Spatiotemporal Transformer
abstract
With the advent of deep learning (DL) for signal processing, the deployment of DL for radar-based crowd counting has yielded significant performance enhancement. Despite these advancements, current methodologies predominantly undergo validation in controlled conditions with limited subject movement variability, posing a challenge for practical usage. Addressing this gap, this letter first attempts the application of radar-based crowd counting in an unregulated and dense setting, capturing the radar reflections of up to 31 subjects in real-world scenarios, such as queues at restaurant kiosks. Furthermore, to address the complexities of such a challenging condition, we introduce a novel radar crowd counting model that utilizes a spatiotemporal transformer. The expremental results demonstrate the potentiality of the proposed model as a robust crowd counting system under the full realistic scenarios, as well as establish its superiority over the conventional radar-based crowd counting models.
Jae-Ho Choi 0004, Kyung-Tae Kim
IEEE Signal Process. Lett.1
2023 MVDoppler: Unleashing the Power of Multi-View Doppler for MicroMotion-based Gait Classification
abstract
Modern perception systems rely heavily on high-resolution cameras, LiDARs, and advanced deep neural networks, enabling exceptional performance across various applications. However, these optical systems predominantly depend on geometric features and shapes of objects, which can be challenging to capture in long-range perception applications. To overcome this limitation, alternative approaches such as Doppler-based perception using high-resolution radars have been proposed. Doppler-based systems are capable of measuring micro-motions of targets remotely and with very high precision. When compared to geometric features, the resolution of micro-motion features exhibits significantly greater resilience to the influence of distance. However, the true potential of Doppler-based perception has yet to be fully realized due to several factors. These include the unintuitive nature of Doppler signals, the limited availability of public Doppler datasets, and the current datasets' inability to capture the specific co-factors that are unique to Doppler-based perception, such as the effect of the radar's observation angle and the target's motion trajectory.This paper introduces a new large multi-view Doppler dataset together with baseline perception models for micro-motion-based gait analysis and classification. The dataset captures the impact of the subject's walking trajectory and radar's observation angle on the classification performance. Additionally, baseline multi-view data fusion techniques are provided to mitigate these effects. This work demonstrates that sub-second micro-motion snapshots can be sufficient for reliable detection of hand movement patterns and even changes in a pedestrian's walking behavior when distracted by their phone. Overall, this research not only showcases the potential of Doppler-based perception, but also offers valuable solutions to tackle its fundamental challenges.
Soheil Hor, Shubo Yang 0002, Jae-Ho Choi 0004, Amin Arbabian
NeurIPS3
2022 Remote Respiration Monitoring of Moving Person Using Radio Signals
Jae-Ho Choi 0004, Ki-Bong Kang, Kyung-Tae Kim
ECCV (37)1
2022 Deep Learning Approach for Radar-Based People Counting
abstract
With the development of deep learning (DL) frameworks in the field of pattern recognition, DL-based algorithms have outperformed handcrafted feature (HF)-based ones in various applications. However, there still exist several challenges in applying the DL framework to a radar-based people counting (RPC) task: The powerful representation capacity of a deep neural network (DNN) learns not only the desired human-induced components but also unwanted nuisance factors, and available data for RPC is usually insufficient to train a huge-sized DNN, leading to an increased possibility of overfitting. To tackle this problem, we propose novel solutions for the successful application of the DL framework to the RPC task from various perspectives. First, we newly formulate the preprocessing pipelines to transform the raw received radar echoes into a better-matched form for a DNN. Second, we devise a novel backbone architecture that reflects the spatiotemporal characteristics of the radar signals, while relieving the burden on training through a parameter efficient design. Finally, an unsupervised pretraining process and a newly defined loss function are proposed for further stabilized network convergence. Several experimental results using real measured data show that the proposed scheme enables an effective utilization of DL for RPC, achieving a significant performance improvement compared to conventional RPC methods.
Jae-Ho Choi 0004, Kyung-Tae Kim
IEEE Internet Things J.1
2022 Fusion of Target and Shadow Regions for Improved SAR ATR
abstract
Synthetic aperture radar (SAR) systems, which operate under a slant-viewing geometry, inevitably entail shadow regions in the resulting radar image. Such shadow profiles contain backprojected signatures of an object’s configuration as with target profiles; however, they are rarely utilized in current SAR-based recognition techniques. A major challenge in leveraging shadow information together lies in the intrinsic limitation of current single-pathway approaches, in which the target and shadow cannot be addressed simultaneously because of their incompatible domain properties. Hence, we herein propose novel solutions that enable the successful fusion of target and shadow regions within SAR for the first time. First, we devise new image preprocessing techniques specifically customized for shadows to compensate for their unique domain characteristics, which are distinct from the target. Second, we introduce a parallelized SAR processing mechanism such that a network can independently extract features oriented toward each conflicting modality. Third, adaptive fusion strategies are proposed for the optimal integration of features from each region while considering their relative significance layer by layer. Extensive experiments on public benchmark datasets demonstrate that the proposed framework allows a network to effectively employ shadow signatures and targets, thereby outperforming previous methods significantly for all setups.
Jae-Ho Choi 0004, Myung-Jun Lee, Nam-Hoon Jeong, Kyung-Tae Kim
IEEE Trans. Geosci. Remote. Sens.1
2021 People Counting Using IR-UWB Radar Sensor in a Wide Area
abstract
Conventional radar-based people counting systems are designed mainly for dense spatial distributions in a small region of interest (ROI). Therefore, a system with only conventional energy-based features, which are effective for a small ROI with a limited spatial distribution of individuals, generally fails to cope with the diverse and complex spatial distributions that arise from the freer movements of individuals as the ROI widens. To address this problem, a novel approach that achieves robust people counting in both wide and small ROIs is presented in this study. The proposed technique incorporates modified CLEAN-based features in the range domain and energy-based features in the frequency domain to efficiently address both dense and dispersed distributions of individuals. Subsequently, principal component analysis and an appropriate normalization of the proposed features are performed for improving the people counting system further. Based on several experiments in practical environments with wide ROIs and severe multipath effects, we observed that the proposed approach yields significantly improved performance compared with traditional people counting systems.
Jae-Ho Choi 0004, Kyung-Tae Kim
IEEE Internet Things J.1