EDBT 2026 Demo / reviewers in the wild / expert
Jong Taek Lee
dblp:38/7728
· DBLP profile ↗
21ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-6962-3148ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Security and privacy · 2 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring Intrinsic Dimension of Multiobjective Landscapes and Dimensionality-Reduced Neuroevolution of Deep Reinforcement Learning
Oladayo S. Ajani, Jeonggeun Kim, Jong Taek Lee, Byung-Chul Tak |
GECCO | 3 |
| 2026 | A two-stage evolutionary framework with structural-parametric decoupling for sparse large-scale multi-objective optimization
Oladayo S. Ajani, Jeonggeun Kim, Jong Taek Lee, Byung-Chul Tak |
Expert Syst. Appl. | 3 |
| 2025 | Finding Optimal Viewpoints for Monocular 3D Human Pose Estimation in Dynamic 3D Gaussian Splatting SpaceabstractMonocular 3D human pose estimation (HPE) focuses on locating 3D joint positions from a single viewpoint. However, it often suffers from viewpoint-dependent depth ambiguities and occlusions, emphasizing that next-best-viewpoint (NBV) selection becomes crucial. Prior work has explored NBV selection, but typically relies on datasets with fixed, discrete camera setups that do not support continuous viewpoint selection or on synthetic environments that lack photorealism. To address these shortcomings, we propose a novel pipeline for NBV selection in monocular 3D HPE. We construct a continuous and photorealistic environment, using 3D Gaussian Splatting (3DGS), which provides continuous novel-view rendering both spatially and temporally. Within this environment, we enable a reinforcement learning (RL) agent to actively select NBVs that minimize pose estimation error. Our approach is evaluated against fixed, random, and rotating camera trajectories, achieving over 24% improvement in pose estimation error. The code is available at https://github.com/knu-vis/FOV. Wan-Gi Bae, Sanghyeon Lee 0001, Jong Taek Lee |
AVSS | 3 |
| 2025 | Discriminative Skeleton-Based Action Recognition via Co-Learning with Motion Diffusion ModelabstractSkeleton-based action recognition is vital for numerous real-world applications, yet it continues to face challenges due to limited and imbalanced 3D motion data, complex movement dynamics, and sensitivity to viewpoint variations. To tackle these issues, we introduce a novel co-learning framework that unifies generative and discriminative modeling by combining a 3D Motion Diffusion Model (MDM) with its inverse counterpart, I-MDM, for more robust recognition. Through the incorporation of high-quality synthetic motions guided by discriminative features from I-MDM, our method achieves state-of-the-art top-1 accuracy on Hu-manAct12 (94.63%) and NTU-13 (99.25%), while delivering over three times faster inference (5.32 ms) than prior methods. These results highlight the potential of diffusion-based generative augmentation guided by discriminative feedback, paving a new path for efficient and accurate skeleton-based action recognition. The code is available at https://github.com/knu-vis/I-MDM. Sanghyeon Lee 0001, Ahmed Fakhry, Jong Taek Lee |
AVSS | 4 |
| 2025 | Black Hole-Driven Identity Absorbing in Diffusion ModelsabstractRecent advances in diffusion models have positioned them as powerful generative frameworks for high-resolution image synthesis across diverse domains. The emerging "h-space" within these models, defined by bottleneck activations in the denoiser, offers promising pathways for semantic image editing similar to GAN latent spaces. However, as demand grows for content erasure and concept removal, privacy concerns highlight the need for identity disentanglement in the latent space of diffusion models. The high-dimensional latent space poses challenges for identity removal, as traversing with random or orthogonal directions often leads to semantically unvalidated regions, resulting in unrealistic outputs. To address these issues, we propose Black Hole-Driven Identity Absorption (BIA), a novel approach for identity erasure within the latent space of diffusion models. BIA uses a "black hole" metaphor, where the latent region representing a specified identity acts as an attractor, drawing in nearby latent points of surrounding identities to "wrap" the black hole. Instead of relying on random traversals for optimization, BIA employs an identity absorption mechanism by attracting and wrapping nearby validated latent points associated with other identities to achieve a vanishing effect for specified identity. Our method effectively prevents the generation of a specified identity while preserving other attributes, as validated by improved scores on identity similarity (SID), FID metrics, qualitative evaluations, and user studies as compared to SOTA. Muhammad Shaheryar, Jong Taek Lee, Soon Ki Jung |
CVPR | 2 |
| 2025 | Occlusion-aware heatmap generation for enhancing 3D human pose estimation in multi-person environments
Sanghyeon Lee 0001, Jong Taek Lee |
Vis. Comput. | 2 |
| 2024 | ELLAR: An Action Recognition Dataset for Extremely Low-Light Conditions with Dual Gamma Adaptive Modulation
Minse Ha, Wan-Gi Bae, Geunyoung Bae, Jong Taek Lee |
ACCV (6) | 4 |
| 2024 | Learning 2D Human Poses for Better 3D Lifting via Multi-model 3D-Guidance
Sanghyeon Lee 0001, Yoonho Hwang, Jong Taek Lee |
ACCV (1) | 3 |
| 2024 | Enhancing Few-Shot Video Anomaly Detection with Key-Frame Selection and Relational Cross TransformersabstractDetecting illegal activities using video anomaly detection is an enormous challenge in security and surveillance. The lack of labeled instances for anomalous actions poses a significant obstacle to existing learning techniques, and determining the optimal data representation that captures the essential features and patterns vital for detecting anomalies proves to be exceedingly difficult. We have developed a few-shot video anomaly detection method, FewVAD, which employs a key-frame selection module and spatial-temporal relational modeling to extract pertinent features and reduce temporal redundancy from lengthy surveillance recordings. We have evaluated our method on two popular surveillance datasets, UCF-Crime and XD-Violence, and compared its performance against established few-shot models and other unsupervised and weakly supervised learning video anomaly detection models. Our model has attained an accuracy of 41.7% and 54.3% for 5-way 5-shot few-shot configuration on the UCF-Crime and XD-Violence datasets, respectively. Furthermore, it has obtained an AUC score of 86.60% for the 2-way anomaly detection task on the UCFCrime dataset. FewVAD achieves a milestone in few-shot video anomaly detection, competing strongly with current weakly-supervised and unsupervised VAD methods. Ahmed Fakhry, Jong Taek Lee |
AVSS | 2 |
| 2024 | Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term FrequencyabstractScene graph generation (SGG) is an important task in image understanding because it represents the relationships between objects in an image as a graph structure, making it possible to understand the semantic relationships between objects intuitively. Previous SGG studies used a message-passing neural networks (MPNN) to update features, which can effectively reflect information about surrounding objects. However, these studies have failed to reflect the co-occurrence of objects during SGG generation. In addition, they only addressed the long-tail problem of the training dataset from the perspectives of sampling and learning methods. To address these two problems, we propose CooK, which reflects the Co-occurrence Knowledge between objects, and the learnable term frequency-inverse document frequency (TF-$l$-IDF) to solve the long-tail problem. We applied the proposed model to the SGG benchmark dataset, and the results showed a performance improvement of up to 3.8% compared with existing state-of-the-art models in SGGen subtask. The proposed method exhibits generalization ability from the results obtained, showing uniform performance improvement for all MPNN models. Sangwon Kim 0004, Dasom Ahn, Jong Taek Lee, ByoungChul Ko |
ICML | 4 |
| 2024 | Pruning-guided feature distillation for an efficient transformer-based pose estimation modelabstractAbstract The authors propose a compression strategy for a 3D human pose estimation model based on a transformer which yields high accuracy but increases the model size. This approach involves a pruning‐guided determination of the search range to achieve lightweight pose estimation under limited training time and to identify the optimal model size. In addition, the authors propose a transformer‐based feature distillation (TFD) method, which efficiently exploits the pose estimation model in terms of both model size and accuracy by leveraging transformer architecture characteristics. Pruning‐guided TFD is the first approach for 3D human pose estimation that employs transformer architecture. The proposed approach was tested on various extensive data sets, and the results show that it can reduce the model size by 30% compared to the state‐of‐the‐art while ensuring high accuracy. Aro Kim, Jong Taek Lee, Sungjei Kim, Sanghyo Park 0001 |
IET Comput. Vis. | 5 |
| 2022 | IDDiffuse: Dual-Conditional Diffusion Model for Enhanced Facial Image Anonymization
Muhammad Shaheryar, Jong Taek Lee, Soon Ki Jung |
ACCV (4) | 2 |
| 2021 | Acoustics to the Rescue: Physical Key Inference Attack Revisited
Soundarya Ramesh, Rui Xiao 0002, Anindya Maiti, Jong Taek Lee, Harini Ramprasad, Ananda Kumar, Murtuza Jadliwala, Jun Han 0001 |
USENIX Security Symposium | 4 |
| 2020 | Poster Abstract: Don't Wait For Weight: Towards Weight Inference of Passengers and Luggage using Smartphone CameraabstractProposals on weighing passengers and their carry-on luggage prior to flights are gaining traction in the airline industry for fuel efficiency purposes. Adoption of such proposals are difficult in practice, however, as requiring passengers to step on weighing scales would incur significant overhead heavily affecting already busy airports. To solve this problem, we propose CamWeight, a novel vision-based weight inference system that takes video feed of off-the-shelf elastic mat (e.g., yoga mat) placed on the floor as the passenger walks over it while pulling his/her wheeled carry-on luggage. CamWeight makes use of inherent properties including amplitude and recovery time of strain, or mat deformation caused by footsteps and luggage wheels. Due to the inherent design of CamWeight, it incurs no additional time for weighing, while being cost effective. We present a preliminary proof-of-concept evaluation by varying weights in a luggage from 2.5 kg to 20 kg to achieve prediction mean absolute error of approximately 2 kg. Jong Taek Lee, Yu Kai Lim, Jun Han 0001 |
IPSN | 1 |
| 2019 | SurFi: detecting surveillance camera looping attacks with wi-fi channel state informationabstractThe proliferation of surveillance cameras has greatly improved the physical security of many security-critical properties including buildings, stores, and homes. However, recent surveillance camera looping attacks demonstrate new security threats --- adversaries can replay a seemingly benign video feed of a place of interest while trespassing or stealing valuables without getting caught. Unfortunately, such attacks are extremely difficult to detect in real-time due to cost and implementation constraints. In this paper, we propose SurFi to detect these attacks in real-time by utilizing commonly available Wi-Fi signals. In particular, we leverage that channel state information (CSI) from Wi-Fi signals also perceives human activities in the place of interest in addition to surveillance cameras. SurFi processes and correlates the live video feeds and the Wi-Fi CSI signals to detect any mismatches that would identify the presence of the surveillance camera looping attacks. SurFi does not require the deployment of additional infrastructure because Wi-Fi transceivers are easily found in the urban indoor environment. We design and implement the SurFi system and evaluate its effectiveness in detecting surveillance camera looping attacks. Our evaluation demonstrates that SurFi effectively identifies attacks with up to an attack detection accuracy of 98.8% and 0.1% false positive rate. Nitya Lakshmanan, Inkyu Bang, Min Suk Kang, Jun Han 0001, Jong Taek Lee |
WiSec | 5 |
| 2018 | Integrating Multiple Inferences for Vehicle Detection by Focusing on Challenging Test SetsabstractDue to recent advances in object detection with the help of deep convolutional neural networks and region proposal methods, object detection systems have become practical in numerous fields with high accuracy. This paper presents a method for vehicle detection in videos for automatic traffic monitoring. Compared to general object detection datasets such as the PascalVOC and MS-COCO, traffic surveillance datasets such as the UA-DETRAC dataset have different challenging issues: high variation of object size, severe occlusion, dissimilarity between training set and test set. To overcome these difficulties, we employ an unsupervised integration of multiple instances of an image by analyzing video sequences. We applied Faster R-CNN with Neural Architecture Search (NAS) framework as a base network. We achieved 85.76% mAP on the UA-DETRAC detection test set, and outperformed the winner method of the AVSS 2017 challenge on Advanced Traffic Monitoring by 9.19%. Jong Taek Lee, Jang-Woon Baek, Kiyoung Moon, Kil-Taek Lim |
AVSS | 1 |
| 2018 | UA-DETRAC 2018: Report of AVSS2018 & IWT4S Challenge on Advanced Traffic MonitoringabstractA desirable smart traffic-monitoring and street-safety system can elicit and support the intervention of law enforcement agencies or medical staff. Recently, there has been a dramatically higher demand for such smart systems. To this end, the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S) was organized in conjunction with the 15th IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS 2018). Our goal is to advance the state-of-the-art detection and tracking algorithms and provide a comprehensive performance evaluation for them. We evaluate 5 submitted detection and 7 submitted tracking methods on the large-scale UA-DETRAC benchmark, and the results are shared publicly on the website http://detrac-db. rit.albany.edu. We expect this challenge to advance the research and development of new detection and tracking methods for transportation applications. Siwei Lyu, Ming-Ching Chang, Dawei Du, Wenbo Li 0001, Yi Wei 0006, Marco Del Coco, Pierluigi Carcagnì, Arne Schumann, Bharti Munjal, Dinh-Quoc-Trung Dang, Doo-Hyun Choi, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Guna Seetharaman, Jang-Woon Baek, Jong Taek Lee, Kannappan Palaniappan, Kil-Taek Lim, Kiyoung Moon, Kwang-Ju Kim, Lars Wilko Sommer, Meltem Brandlmaier, Minsung Kang, Moongu Jeon, Noor Al-Shakarji, Oliver Acatay, Pyong-Kun Kim, Sikandar Amin, Thomas Sikora, Tien Ba Dinh, Tobias Senst, Vu-Gia-Hy Che, Young-Chul Lim, Yun-Su Chung |
AVSS | 17 |
| 2011 | AVSS 2011 demo session: A large-scale benchmark dataset for event recognition in surveillance videoabstractSummary form only given. We present a concept for automatic construction site monitoring by taking into account 4D information (3D over time), that is acquired from highly-overlapping digital aerial images. On the one hand today's maturity of flying micro aerial vehicles (MAVs) enables a low-cost and an efficient image acquisition of high-quality data that maps construction sites entirely from many varying viewpoints. On the other hand, due to low-noise sensors and high redundancy in the image data, recent developments in 3D reconstruction workflows have benefited the automatic computation of accurate and dense 3D scene information. Having both an inexpensive high-quality image acquisition and an efficient 3D analysis workflow enables monitoring, documentation and visualization of observed sites over time with short intervals. Relating acquired 4D site observations, composed of color, texture, geometry over time, largely supports automated methods toward full scene understanding, the acquisition of both the change and the construction site's progress. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
AVSS | 6 |
| 2011 | A large-scale benchmark dataset for event recognition in surveillance videoabstractWe introduce a new large-scale video dataset designed to assess the performance of diverse visual event recognition algorithms with a focus on continuous visual event recognition (CVER) in outdoor areas with wide coverage. Previous datasets for action recognition are unrealistic for real-world surveillance because they consist of short clips showing one action by one individual [15, 8]. Datasets have been developed for movies [11] and sports [12], but, these actions and scene conditions do not apply effectively to surveillance videos. Our dataset consists of many outdoor scenes with actions occurring naturally by non-actors in continuously captured videos of the real world. The dataset includes large numbers of instances for 23 event types distributed throughout 29 hours of video. This data is accompanied by detailed annotations which include both moving object tracks and event examples, which will provide solid basis for large-scale evaluation. Additionally, we propose different types of evaluation modes for visual recognition tasks and evaluation metrics along with our preliminary experimental results. We believe that this dataset will stimulate diverse aspects of computer vision research and help us to advance the CVER tasks in the years ahead. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
CVPR | 6 |
| 2009 | Real-Time Illegal Parking Detection in Outdoor Environments Using 1-D TransformationabstractWith decreasing costs of high-quality surveillance systems, human activity detection and tracking has become increasingly practical. Accordingly, automated systems have been designed for numerous detection tasks, but the task of detecting illegally parked vehicles has been left largely to the human operators of surveillance systems. We propose a methodology for detecting this event in real time by applying a novel image projection that reduces the dimensionality of the data and, thus, reduces the computational complexity of the segmentation and tracking processes. After event detection, we invert the transformation to recover the original appearance of the vehicle and to allow for further processing that may require 2-D data. We evaluate the performance of our algorithm using the i-LIDS vehicle detection challenge datasets as well as videos we have taken ourselves. These videos test the algorithm in a variety of outdoor conditions, including nighttime video and instances of sudden changes in weather. Jong Taek Lee, Michael S. Ryoo, Matthew Riley, Jake K. Aggarwal |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Real-time detection of illegally parked vehicles using 1-D transformationabstractWith decreasing costs of high quality surveillance systems, human activity detection and tracking has become increasingly practical. Accordingly, automated systems have been designed for numerous detection tasks, but the task of detecting illegally parked vehicles has been left largely to the human operators of surveillance systems. We propose a methodology for detecting this event in realtime by applying a novel image projection that reduces the dimensionality of the image data and thus reduces the computational complexity of the segmentation and tracking processes. After event detection, we invert the transformation to recover the original appearance of the vehicle and to allow for further processing that may require the two dimensional data. The proposed algorithm is able to successfully recognize illegally parked vehicles in real-time in the i-LIDS bag and vehicle detection challenge datasets. Jong Taek Lee, Michael S. Ryoo, Matthew Riley, Jake K. Aggarwal |
AVSS | 1 |