Jangwon Lee 0002

dblp:52/4915-2 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-6601-7302ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Face, body and person analysis · 90% Video understanding and tracking · 10%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis › face alignment
heatmap regression
0.812024
Motion-Aware Heatmap Regression for Human Pose Estimation in Videos · IJCAI 2024
Computer vision › Face, body and person analysis
human pose estimation
0.812024
Motion-Aware Heatmap Regression for Human Pose Estimation in Videos · IJCAI 2024
Computer vision › Face, body and person analysis › human pose estimation
video pose estimation
0.812024
Motion-Aware Heatmap Regression for Human Pose Estimation in Videos · IJCAI 2024
Computer vision › Video understanding and tracking › multi-camera video analysis
cross-view video analysis
0.312017
Identifying First-Person Camera Wearers in Third-Person Videos · CVPR 2017
Computer vision › Face, body and person analysis
person re-identification
0.312017
Identifying First-Person Camera Wearers in Third-Person Videos · CVPR 2017
Wearable and physiological sensing › wearable camera
egocentric video
0.112017
Identifying First-Person Camera Wearers in Third-Person Videos · CVPR 2017

Methods — techniques the papers use, named apart from their topics

motion-aware regression · 0.8triplet loss · 0.6siamese network · 0.6joint embedding · 0.6
YearPublicationVenuePosition
2025 Anomaly Detection for People with Visual Impairments Using an Egocentric 360-Degree Camera
abstract
Recent advancements in computer vision have led to a renewed interest in developing assistive technologies for individuals with visual impairments. Although extensive research has been conducted in the field of computer vision-based assistive technologies, most of the focus has been on understanding contexts in images, rather than addressing their physical safety and security concerns. To address this challenge, we propose the first step towards detecting anomalous situations for visually impaired people by observing their entire surroundings using an egocentric 360-degree camera. We first introduce a novel egocentric 360-degree video dataset called VIEW360 (Visually Impaired Equipped with Wearable 360-degree camera), which contains abnormal activities that visually impaired individuals may encounter, such as shoulder surfing and pickpocketing. Furthermore, we propose a new architecture called the FDPN (Frame and Direction Prediction Network), which facilitates frame-level prediction of abnormal events and identifying of their directions. Finally, we evaluate our approach on our VIEW360 dataset and the publicly available UCF-Crime and Shanghaitech datasets, demonstrating state-of-the-art performance. Code and dataset are available at https://github.com/Songinpyo/VIEW360.
Inpyo Song, Minjun Joo, Jangwon Lee 0002
WACV4
2024 Motion-Aware Heatmap Regression for Human Pose Estimation in Videos
Inpyo Song, Moonwook Ryu, Jangwon Lee 0002
IJCAI4
2024 SFTrack: A Robust Scale and Motion Adaptive Algorithm for Tracking Small and Fast Moving Objects
abstract
This paper addresses the problem of multi-object tracking in Unmanned Aerial Vehicle (UAV) footage. It plays a critical role in various UAV applications, including traffic monitoring systems and real-time suspect tracking by the police. However, this task is highly challenging due to the fast motion of UAVs, as well as the small size of target objects in the videos caused by the high-altitude and wide-angle views of drones. In this study, we thus introduce a simple yet more effective method compared to previous work to overcome these challenges. Our approach involves a new tracking strategy, which initiates the tracking of target objects from low-confidence detections commonly encountered in UAV application scenarios. Additionally, we propose revisiting traditional appearance-based matching algorithms to improve the association of low-confidence detections. To evaluate the effectiveness of our method, we conducted benchmark evaluations on two UAV-specific datasets (VisDrone2019, UAVDT) and one general object tracking dataset (MOT17). The results demonstrate that our approach surpasses current state-of-the-art methodologies, highlighting its robustness and adaptability in diverse tracking environments. Furthermore, we have improved the annotation of the UAVDT dataset by rectifying several errors and addressing omissions found in the original annotations. We will provide this refined version of the dataset to facilitate better benchmarking in the field.
Inpyo Song, Jangwon Lee 0002
IROS2
2024 Action-conditioned contrastive learning for 3D human pose and shape estimation in videos
Inpyo Song, Moonwook Ryu, Jangwon Lee 0002
Comput. Vis. Image Underst.3
2021 A Sample Weighting and Score Aggregation Method for Multi-query Object Matching
abstract
In this paper, we propose a simple and effective method to properly assign weights to the query samples and compute aggregated matching scores using these weights in multi-query object matching. Multi-query object matching commonly exists in many real-life problems such as finding suspicious objects in surveillance videos. In this problem, a query object is represented by multiple samples and the matching candidates in a database are ranked according to their similarities to these query samples. In this context, query samples are not equally effective to find the target object in the database, thus one of the key challenges is how to measure the effectiveness of each query to find the correct matching object. So far, however, very little attention has been paid to address this issue. Therefore, we propose a simple but effective way, Inverse Model Frequency (IMF), to measure of matching effectiveness of query samples. Furthermore, we introduce a new score aggregation method to boost the object matching performance given multiple queries. We tested the proposed method for vehicle re-identification and image retrieval tasks. Our proposed approach achieves state-of-the-art matching accuracy on two vehicle re-identification datasets (VehicleID/VeRi-776) and two image retrieval datasets (the original & revisited Oxford/Paris). The proposed approach can seamlessly plug into many existing multi-query object matching approaches to further boost their performance with minimal effort.
Jangwon Lee 0002, Gang Qian, Allison Beach
AVSS1
2019 Observing Pianist Accuracy and Form with Computer Vision
abstract
We present a first step towards developing an interactive piano tutoring system that can observe a student playing the piano and give feedback about hand movements and musical accuracy. In particular, we have two primary aims: 1) to determine which notes on a piano are being played at any moment in time, 2) to identify which finger is pressing each note. We introduce a novel two-stream convolutional neural network that takes video and audio inputs together for detecting pressed notes and finger presses. We formulate our two problems in terms of multi-task learning and extend a state-of-the-art object detection model to incorporate both audio and visual features. In addition, we introduce a novel finger identification solution based on pressed piano note information. We experimentally confirm that our approach is able to detect pressed piano keys and the piano player's fingers with a high accuracy.
Jangwon Lee 0002, Bardia Doosti, Yupeng Gu, David Cartledge, David Crandall, Christopher Raphael
WACV1
2017 Identifying First-Person Camera Wearers in Third-Person Videos
abstract
We consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in scenarios in which multiple people are wearing body-worn cameras while a third-person static camera also captures the scene. To do this, we need to establish person-level correspondences across first-and third-person videos, which is challenging because the camera wearer is not visible from his/her own egocentric video, preventing the use of direct feature matching. In this paper, we propose a new semi-Siamese Convolutional Neural Network architecture to address this novel challenge. We formulate the problem as learning a joint embedding space for first-and third-person videos that considers both spatial-and motion-domain cues. A new triplet loss function is designed to minimize the distance between correct first-and third-person matches while maximizing the distance between incorrect ones. This end-to-end approach performs significantly better than several baselines, in part by learning the first-and third-person features optimized for matching jointly with the distance measure itself.
Chenyou Fan, Jangwon Lee 0002, Krishna Kumar Singh, Yong Jae Lee, David Crandall, Michael S. Ryoo
CVPR2
2017 Learning robot activities from first-person human videos using convolutional future regression
abstract
We design a new approach that allows robot learning of new activities from unlabeled human example videos. Given videos of humans executing the same activity from a human's viewpoint (i.e., first-person videos), our objective is to make the robot learn the temporal structure of the activity as its future regression network, and learn to transfer such model for its own motor execution. We present a new deep learning model: We extend the state-of-the-art convolutional object detection network for the representation/estimation of human hands in training videos, and newly introduce the concept of using a fully convolutional network to regress (i.e., predict) the intermediate scene representation corresponding to the future frame (e.g., 1-2 seconds later). Combining these allows direct prediction of future locations of human hands and objects, which enables the robot to infer the motor control plan using our manipulation network. We experimentally confirm that our approach makes learning of robot activities from unlabeled human interaction videos possible, and demonstrate that our robot is able to execute the learned collaborative activities in real-time directly based on its camera input.
Jangwon Lee 0002, Michael S. Ryoo
IROS1
2007 Study on Behavioral Personality of a Service Robot to make more Convenient to Customer
abstract
In recent years, a lot of companies and laboratories start to announce their various prototype service robots. It does not exist that a killer application of a service robot in a service robot market yet, however it is expected that customers demand a robot having more variety peculiarity when a robot is supplied widely. In order words, in the future, robot users probably want the robot which is fitted for their individual tastes. In the present paper we shall see the robot's variation of behavioral personality from the experiments. These results show that by adjusting the parameters related to robot perception, the robot can be varied in its personality for user. The robot is implemented by "cognitive robotic engine (CRE)" for the dependable integration of human-robot interaction (HRI) components.
Jangwon Lee 0002, Hun-Sue Lee, Sukhan Lee 0001
RO-MAN1