VLDB 2026 Research / reviewers in the wild / expert
Inwoong Lee
dblp:145/2294
· DBLP profile ↗
9ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-4356-7616ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMsabstractVideo large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with token count. To address this, we propose a training-free spatio-temporal token merging method, named STTM. Our key insight is to exploit local spatial and temporal redundancy in video data which has been overlooked in prior work. STTM first transforms each frame into multi-granular spatial tokens using a coarse-to-fine search over a quadtree structure, then performs directed pairwise merging across the temporal dimension. This decomposed merging approach outperforms existing token reduction methods across six video QA benchmarks. Notably, STTM achieves a 2$\times$ speed-up with only a 0.5% accuracy drop under a 50% token budget, and a 3$\times$ speed-up with just a 2% drop under a 30% budget. Moreover, STTM is query-agnostic, allowing KV cache reuse across different questions for the same video. The project page is available at https://www.jshyun.me/projects/sttm. Jeongseok Hyun, Sukjun Hwang, Su Ho Han, Taeoh Kim, Inwoong Lee, Dongyoon Wee, Joon-Young Lee, Seon Joo Kim, Minho Shim |
ICCV | 5 |
| 2025 | Prototypes Are Balanced Units for Efficient and Effective Partially Relevant Video RetrievalabstractIn a retrieval system, simultaneously achieving search accuracy and efficiency is inherently challenging. This challenge is particularly pronounced in partially relevant video retrieval (PRVR), where incorporating more diverse context representations at varying temporal scales for each video enhances accuracy but increases computational and memory costs. To address this dichotomy, we propose a prototypical PRVR framework that encodes diverse contexts within a video into a fixed number of prototypes. We then introduce several strategies to enhance text association and video understanding within the prototypes, along with an orthogonal objective to ensure that the prototypes capture a diverse range of content. To keep the prototypes searchable via text queries while accurately encoding video contexts, we implement cross- and uni-modal reconstruction tasks. The cross-modal reconstruction task aligns the prototypes with textual features within a shared space, while the uni-modal reconstruction task preserves all video contexts during encoding. Additionally, we employ a video mixing technique to provide weak guidance to further align prototypes and associated textual representations. Extensive evaluations on TVR, ActivityNet-Captions, and QVHighlights validate the effectiveness of our approach without sacrificing efficiency. WonJun Moon, Cheol-Ho Cho, Woojin Jun, Taeoh Kim, Inwoong Lee, Dongyoon Wee, Minho Shim, Jae-Pil Heo |
ICCV | 5 |
| 2024 | Classification Matters: Improving Video Action Detection with Class-Specific Attention
Jinsung Lee, Taeoh Kim, Inwoong Lee, Minho Shim, Dongyoon Wee, Minsu Cho, Suha Kwak |
ECCV (20) | 3 |
| 2024 | Kinematic Diversity and Rhythmic Alignment in Choreographic Quality Transformers for Dance Quality AssessmentabstractIn recent years, the dance entertainment industry has experienced significant growth, driven by the desire of consumers to learn and improve their dancing skills. To effectively improve their skills, dancers require evaluation and feedback, which traditionally relies heavily on professional dancers. To address this challenge, researchers have proposed objective assessment methods for dance performance via kinematic data captured by sensors. However, these existing methods primarily focus on assessing the rhythmic accuracy of movements synchronized to music. In this paper, we propose Dance Quality Assessment (DanceQA) Framework to evaluate dance performance, considering choreographic factors that are important criteria in subjective DanceQA. We find that kinematic diversity and rhythmic alignment are significant choreographic factors from human perception perspective. Based on these factors, we design two metrics: kinematic information entropy (KIE) and kinematic-music beat similarity (BSIM). Our study demonstrates that these metrics are closely related to specific body parts in each choreography. To validate the effectiveness of our metrics, we capture dance performance by OptiTrack system providing precise three-dimensional data at very high sampling rate. We then label their dance quality via subjective test. The metrics give strong correlation with subjective opinion, but it is difficult to tell which body part is the most correlated. To comprehensively understand the dance quality, we propose choreographic quality transformers (CQTs), which learn the aforementioned choreographic factors by embedding KIE and BSIM into attention matrices. In numerous experiments, the CQTs outperforms previous methods, graph convolutional networks and multimodal transformers, at least by up to 0.146 in correlation coefficient. Taewan Kim 0002, Inwoong Lee, Sanghoon Lee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | 3-D Human Behavior Understanding Using Generalized TS-LSTM NetworksabstractThis paper addresses the problems of skeleton feature representation and the modeling of temporal dynamics to recognize human actions consisting of poses. In contrast to traditional methods which generally used relative coordinate systems dependent on some joints, or modeled only the long-term dependency, we attempt to understand 3D human behavior with observation by taking temporally different windows. Instead of taking raw skeletons as the input, we transform the skeletons into another coordinate system to obtain the robustness to scale, rotation and translation, and extract motion features between adjacent skeletons, which finally constructs an efficient hybrid-stream combining both pose and motion streams. We propose novel generalized Temporal Sliding Long Short-term Memory (TS-LSTM) networks. The proposed networks are composed of multiple TS-LSTM networks with various hyper-parameters, which can capture various temporal dynamics of actions. We also propose a novel hyper-parameter searching method, which finds decent hyper-parameters of generalized TS-LSTM to handle temporal dynamics of actions. In the experiment, we evaluate the proposed networks to verify the effectiveness of the proposed methods, and compare them with the other methods on three challenging datasets. Additionally, we analyze a relation between the recognized actions and the hyper-parameters, and visualize the layers of the proposed models. Inwoong Lee, Sanghoon Lee 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | Propagating LSTM: 3D Pose Estimation Based on Joint Interdependency
Kyoungoh Lee, Inwoong Lee, Sanghoon Lee 0001 |
ECCV (7) | 2 |
| 2017 | Ensemble Deep Learning for Skeleton-Based Action Recognition Using Temporal Sliding LSTM NetworksabstractThis paper addresses the problems of feature representation of skeleton joints and the modeling of temporal dynamics to recognize human actions. Traditional methods generally use relative coordinate systems dependent on some joints, and model only the long-term dependency, while excluding short-term and medium term dependencies. Instead of taking raw skeletons as the input, we transform the skeletons into another coordinate system to obtain the robustness to scale, rotation and translation, and then extract salient motion features from them. Considering that Long Shortterm Memory (LSTM) networks with various time-step sizes can model various attributes well, we propose novel ensemble Temporal Sliding LSTM (TS-LSTM) networks for skeleton-based action recognition. The proposed network is composed of multiple parts containing short-term, mediumterm and long-term TS-LSTM networks, respectively. In our network, we utilize an average ensemble among multiple parts as a final feature to capture various temporal dependencies. We evaluate the proposed networks and the additional other architectures to verify the effectiveness of the proposed networks, and also compare them with several other methods on five challenging datasets. The experimental results demonstrate that our network models achieve the state-of-the-art performance through various temporal features. Additionally, we analyze a relation between the recognized actions and the multi-term TS-LSTM features by visualizing the softmax features of multiple parts. Inwoong Lee, Seoungyoon Kang, Sanghoon Lee 0001 |
ICCV | 1 |
| 2014 | Optimal phase control for joint transmission and reception with beamformingabstractIn this paper, we propose a joint transmission and reception with phase control for beamforming in a multi-cell environment. For generated transmit weight vectors of multiple base stations (BSs), a mobile station (MS) calculates phases to maximize an achievable rate with low rate feedback. By using the phases, the multiple transmit weight vectors are coordinated to improve the signal to noise ratio (SNR). In order to find optimal phases, we present a phase control method for the effective channel via geometrical approach. Seonghyun Kim, Hojae Lee, Beom Kwon, Inwoong Lee, Sanghoon Lee 0001 |
IPCCC | 4 |
| 2014 | Combinatorial JPT based on orthogonal beamforming for two-cell cooperationabstractIn this paper, we investigate efficient multi-cell cooperation based on CoMP-joint processing and transmission (CoMP-JPT) with orthogonal beamforming. Through the use of a combinatorial optimization algorithm, the optimal user scheduling for joint transmission using multiple transmitters is accomplished. The throughput of the CoMP-JPT can be significantly improved while maintaining fairness among users over a multi-cell environment. Hojae Lee, Beom Kwon, Seonghyun Kim, Inwoong Lee, Sanghoon Lee 0001 |
IPCCC | 4 |