EDBT 2026 Demo / reviewers in the wild / expert
Kyle Min 0001
dblp:218/5498-1
· DBLP profile ↗
14ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-3978-918XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
Jongseo Lee, Kyungho Bae, Kyle Min 0001, Gyeong-Moon Park, Jinwoo Choi 0001 |
ICCV | 3 |
| 2025 | EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven AlignmentabstractErasing harmful or proprietary concepts from powerful text‑to‑image generators is an emerging safety requirement, yet current ``concept erasure'' techniques either collapse image quality, rely on brittle adversarial losses, or demand prohibitive retraining cycles. We trace these limitations to a myopic view of the denoising trajectories that govern diffusion‑based generation. We introduce EraseFlow, the first framework that casts concept unlearning as exploration in the space of denoising paths and optimizes it with a GFlowNets equipped with the trajectory‑balance objective. By sampling entire trajectories rather than single end states, EraseFlow learns a stochastic policy that steers generation away from target concepts while preserving the model’s prior. EraseFlow eliminates the need for carefully crafted reward models and by doing this, it generalizes effectively to unseen concepts and avoids hackable rewards while improving the performance. Extensive empirical results demonstrate that EraseFlow outperforms existing baselines and achieves an optimal trade-off between performance and prior preservation. Abhiram Kusumba, Maitreya Patel, Kyle Min 0001, Changhoon Kim, Chitta Baral, Yezhou Yang |
NeurIPS | 3 |
| 2025 | Deep Geometric Moments Promote Shape Consistency in Text-to-3D GenerationabstractTo address the data scarcity associated with 3D assets, 2D-lifting techniques such as Score Distillation Sampling (SDS) have become a widely adopted practice in text-to-3D generation pipelines. However, the diffusion models used in these techniques are prone to viewpoint bias and thus lead to geometric inconsistencies such as the Janus problem. To counter this, we introduce MT3D, a text-to-3D generative model that leverages a high-fidelity 3D object to overcome viewpoint bias and explicitly infuse geometric understanding into the generation pipeline. Firstly, we employ depth maps derived from a high-quality 3D model as control signals to guarantee that the generated 2D images preserve the funda-mental shape and structure, thereby reducing the inherent viewpoint bias. Next, we utilize deep geometric moments to ensure geometric consistency in the 3D representation explicitly. By incorporating geometric details from a 3D asset, MT3D enables the creation of diverse and geometri-cally consistent objects, thereby improving the quality and usability of our 3D representations. Project page and code: https://moment-3d.github.io/ Utkarsh Nath, Rajeev Goel, Eun Som Jeon, Changhoon Kim, Kyle Min 0001, Yezhou Yang, Yingzhen Yang, Pavan Turaga |
WACV | 5 |
| 2025 | Ego-VPA: Egocentric Video Understanding with Parameter-Efficient AdaptationabstractVideo understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (EgoVFMs) based on video-language pre-training and propose a parameter-efficient adaptation for egocentric video tasks, namely Ego-VPA. It employs a local sparse approximation for each video frame/text feature using the basis prompts, and the selected basis prompts are used to synthesize video/text prompts. Since the basis prompts are shared across frames and modalities, it models context fusion and cross-modal transfer in an efficient fashion. Experiments show that EgoVPA excels in lightweight adaptation (with only 0.84% learnable parameters), largely improving over baselines and reaching the performance of full fine-tuning. Tz-Ying Wu, Kyle Min 0001, Subarna Tripathi, Nuno Vasconcelos |
WACV | 2 |
| 2024 | WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion ModelsabstractThe rapid advancement of generative models, facilitating the creation of hyper-realistic images from textual de-scriptions, has concurrently escalated critical societal con-cerns such as misinformation. Although providing some mitigation, traditional fingerprinting mechanisms fall short in attributing responsibility for the malicious use of syn-thetic images. This paper introduces a novel approach to model fingerprinting that assigns responsibility for the gen-erated images, thereby serving as a potential countermea-sure to model misuse. Our method modifies generative mod-els based on each user's unique digital fingerprint, imprinting a unique identifier onto the resultant content that can be traced back to the user. This approach, incorporating fine-tuning into Text-to-Image (T2I) tasks using the Stable Diffusion Model, demonstrates near-perfect attribution ac-curacy with a minimal impact on output quality. Through extensive evaluation, we show that our method outperforms baseline methods with an average improvement of 11 % in handling image post-processes. Our method presents a promising and novel avenue for accountable model distribution and responsible use. Our code is available in https://github.com/kylemin/WQUAF. Changhoon Kim, Kyle Min 0001, Maitreya Patel, Yezhou Yang |
CVPR | 2 |
| 2024 | Action Scene Graphs for Long-Form Understanding of Egocentric VideosabstractWe present Egocentric Action Scene Graphs (EASGs), a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos, such as verb-noun action labels, by providing a temporally evolving graph-based description of the actions performed by the camera wearer, including interacted objects, their relationships, and how actions unfold in time. Through a novel annotation procedure, we extend the Ego4D dataset adding manually labeled Egocentric Action Scene Graphs which offer a rich set of annotations for long-from egocentric video understanding. We hence define the EASG generation task and provide a baseline approach, establishing preliminary benchmarks. Experiments on two downstream tasks, action anticipation and activity summarization, highlight the effectiveness of EASGs for long-form egocentric video understanding. We will release the dataset and code to replicate experiments and annotations11The code is available at https://github.com/fpv-iplab/EASG. Ivan Rodin, Antonino Furnari, Kyle Min 0001, Subarna Tripathi, Giovanni Maria Farinella |
CVPR | 3 |
| 2024 | R.A.C.E. : Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model
Changhoon Kim, Kyle Min 0001, Yezhou Yang |
ECCV (83) | 2 |
| 2023 | SViTT: Temporal Learning of Sparse Video-Text TransformersabstractDo video-text transformers learn to model temporal relationships across frames? Despite their immense capacity and the abundance of multimodal training data, recent work has revealed the strong tendency of video-text models towards frame-based spatial representations, while temporal reasoning remains largely unsolved. In this work, we identify several key challenges in temporal learning of video-text transformers: the spatiotemporal trade-off from limited network size; the curse of dimensionality for multi-frame modeling; and the diminishing returns of semantic information by extending clip length. Guided by these findings, we propose SVi TT, a sparse video-text architecture that performs multi-frame reasoning with significantly lower cost than naïve transformers with dense attention. Analogous to graph-based networks, SViTT employs two forms of sparsity: edge sparsity that limits the query-key communications between tokens in self-attention, and node sparsity that discards uninformative visual tokens. Trained with a curriculum which increases model sparsity with the clip length, SVi TT outperforms dense transformer baselines on multiple video-text retrieval and question answering benchmarks, with a fraction of computational cost. Project page: http://svcl.ucsd.edu/projects/svitt. Yi Li 0051, Kyle Min 0001, Subarna Tripathi, Nuno Vasconcelos |
CVPR | 2 |
| 2023 | Unbiased Scene Graph Generation in VideosabstractThe task of dynamic scene graph generation (SGG) from videos is complicated and challenging due to the inherent dynamics of a scene, temporal fluctuation of model predictions, and the long-tailed distribution of the visual relationships in addition to the already existing challenges in image-based SGG. Existing methods for dynamic SGG have primarily focused on capturing spatio-temporal context using complex architectures without addressing the challenges mentioned above, especially the long-tailed distribution of relationships. This often leads to the generation of biased scene graphs. To address these challenges, we introduce a new framework called TEMPURA: TEmporal consistency and Memory Prototype guided UnceR-tainty Attenuation for unbiased dynamic SGG. TEMPURA employs object-level temporal consistencies via transformer-based sequence modeling, learns to synthesize unbiased relationship representations using memory-guided training, and attenuates the predictive uncertainty of visual relations using a Gaussian Mixture Model (GMM). Extensive experiments demonstrate that our method achieves significant (up to 10% in some cases) performance gain over existing methods highlighting its superiority in generating more unbiased scene graphs. Code: https://github.com/sayaknag/unbiasedSGG.git Sayak Nag, Kyle Min 0001, Subarna Tripathi, Amit K. Roy-Chowdhury |
CVPR | 2 |
| 2022 | Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection
Kyle Min 0001, Sourya Roy, Subarna Tripathi, Tanaya Guha, Somdeb Majumdar |
ECCV (35) | 1 |
| 2021 | Integrating Human Gaze into Attention for Egocentric Activity RecognitionabstractIt is well known that human gaze carries significant information about visual attention. However, there are three main difficulties in incorporating the gaze data in an attention mechanism of deep neural networks: (i) the gaze fixation points are likely to have measurement errors due to blinking and rapid eye movements; (ii) it is unclear when and how much the gaze data is correlated with visual attention; and (iii) gaze data is not always available in many real-world situations. In this work, we introduce an effective probabilistic approach to integrate human gaze into spatiotemporal attention for egocentric activity recognition. Specifically, we represent the locations of gaze fixation points as structured discrete latent variables to model their uncertainties. In addition, we model the distribution of gaze fixations using a variational method. The gaze distribution is learned during the training process so that the ground-truth annotations of gaze locations are no longer needed in testing situations since they are predicted from the learned gaze distribution. The predicted gaze locations are used to provide informative attentional cues to improve the recognition performance. Our method outperforms all the previous state-of-the-art approaches on EGTEA, which is a large-scale dataset for egocentric activity recognition provided with gaze measurements. We also perform an ablation study and qualitative analysis to demonstrate that our attention mechanism is effective. Kyle Min 0001, Jason J. Corso |
WACV | 1 |
| 2020 | Adversarial Background-Aware Loss for Weakly-Supervised Temporal Activity Localization
Kyle Min 0001, Jason J. Corso |
ECCV (14) | 1 |
| 2019 | TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionabstractTASED-Net is a 3D fully-convolutional network architecture for video saliency detection. It consists of two building blocks: first, the encoder network extracts low-resolution spatiotemporal features from an input clip of several consecutive frames, and then the following prediction network decodes the encoded features spatially while aggregating all the temporal information. As a result, a single prediction map is produced from an input clip of multiple frames. Frame-wise saliency maps can be predicted by applying TASED-Net in a sliding-window fashion to a video. The proposed approach assumes that the saliency map of any frame can be predicted by considering a limited number of past frames. The results of our extensive experiments on video saliency detection validate this assumption and demonstrate that our fully-convolutional model with temporal aggregation method is effective. TASED-Net significantly outperforms previous state-of-the-art approaches on all three major large-scale datasets of video saliency detection: DHF1K, Hollywood2, and UCFSports. After analyzing the results qualitatively, we observe that our model is especially better at attending to salient moving objects. Kyle Min 0001, Jason J. Corso |
ICCV | 1 |
| 2018 | Hierarchical Novelty Detection for Visual Object RecognitionabstractDeep neural networks have achieved impressive success in large-scale visual object recognition tasks with a predefined set of classes. However, recognizing objects of novel classes unseen during training still remains challenging. The problem of detecting such novel classes has been addressed in the literature, but most prior works have focused on providing simple binary or regressive decisions, e.g., the output would be "known," "novel," or corresponding confidence intervals. In this paper, we study more informative novelty detection schemes based on a hierarchical classification framework. For an object of a novel class, we aim for finding its closest super class in the hierarchical taxonomy of known classes. To this end, we propose two different approaches termed top-down and flatten methods, and their combination as well. The essential ingredients of our methods are confidence-calibrated classifiers, data relabeling, and the leave-one-out strategy for modeling novel classes under the hierarchical taxonomy. Furthermore, our method can generate a hierarchical embedding that leads to improved generalized zero-shot learning performance in combination with other commonly-used semantic embeddings. Kibok Lee 0003, Kimin Lee, Kyle Min 0001, Yuting Zhang 0001, Jinwoo Shin, Honglak Lee |
CVPR | 3 |