EDBT 2026 Demo / reviewers in the wild / expert
Dayoung Gong
dblp:321/1839
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Video understanding and tracking · 62% Generative modeling · 24% Vision and language · 8% | |
| Theoretical computer science
1 paper |
Automata and formal languages · 100% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.6 | 2 | 2025 | Generic Event Boundary Detection via Denoising Diffusion · ICCV 2025 ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation · NeurIPS 2024 |
Computer vision › Video understanding and tracking
action segmentation |
1.4 | 2 | 2024 | ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation · NeurIPS 2024 Activity Grammars for Temporal Action Segmentation · NeurIPS 2023 |
Computer vision › Video understanding and tracking
action anticipation |
1.3 | 2 | 2024 | ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation · NeurIPS 2024 Future Transformer for Long-term Action Anticipation · CVPR 2022 |
Computer vision › Video understanding and tracking › action anticipation
long-term action anticipation |
1.3 | 2 | 2024 | ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation · NeurIPS 2024 Future Transformer for Long-term Action Anticipation · CVPR 2022 |
Machine learning › Generative modeling › diffusion model › score-based generative model
denoising diffusion |
0.9 | 1 | 2025 | Generic Event Boundary Detection via Denoising Diffusion · ICCV 2025 |
Computer vision › Video understanding and tracking › video event understanding
event boundary detection |
0.9 | 1 | 2025 | Generic Event Boundary Detection via Denoising Diffusion · ICCV 2025 |
Computer vision › Video understanding and tracking › video event understanding › event boundary detection
generic event boundary detection |
0.9 | 1 | 2025 | Generic Event Boundary Detection via Denoising Diffusion · ICCV 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Video Summarization with Large Language Models · CVPR 2025 |
Multimedia analysis and retrieval
video summarization |
0.9 | 1 | 2025 | Video Summarization with Large Language Models · CVPR 2025 |
Automata and formal languages › formal grammars
context-free grammar |
0.7 | 1 | 2023 | Activity Grammars for Temporal Action Segmentation · NeurIPS 2023 |
Automata and formal languages
grammatical inference |
0.7 | 1 | 2023 | Activity Grammars for Temporal Action Segmentation · NeurIPS 2023 |
Computer vision › Video understanding and tracking
action recognition |
0.6 | 1 | 2022 | Future Transformer for Long-term Action Anticipation · CVPR 2022 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.6 | 1 | 2022 | Future Transformer for Long-term Action Anticipation · CVPR 2022 |
Methods — techniques the papers use, named apart from their topics
global attention · 2.3multimodal large language model · 1.7large language model · 1.7grammar induction · 1.3generalized parser · 1.3temporal self-similarity · 0.9classifier-free guidance · 0.9joint learning · 0.8diffusion model · 0.8anticipative masking · 0.8neural network · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Video Summarization with Large Language ModelsabstractThe exponential increase in video content poses significant challenges in terms of efficient navigation, search, and retrieval, thus requiring advanced video summarization techniques. Existing video summarization methods, which heavily rely on visual features and temporal dynamics, often fail to capture the semantics of video content, resulting in incomplete or incoherent summaries. To tackle the challenge, we propose a new video summarization framework that leverages the capabilities of recent Large Language Models (LLMs), expecting that the knowledge learned from massive data enables LLMs to evaluate video frames in a manner that better aligns with diverse semantics and human judgments, effectively addressing the inherent subjectivity in defining keyframes. Our method, dubbed LLM-based Video Summarization (LLMVS), translates video frames into a sequence of captions using a Muti-modal Large Language Model (M-LLM) and then assesses the importance of each frame using an LLM, based on the captions in its local context. These local importance scores are refined through a global attention mechanism in the entire context of video captions, ensuring that our summaries effectively reflect both the details and the overarching narrative. Our experimental results demonstrate the superiority of the proposed method over existing ones in standard benchmarks, highlighting the potential of LLMs in the processing of multimedia content. Min Jung Lee, Dayoung Gong, Minsu Cho |
CVPR | 2 |
| 2025 | Generic Event Boundary Detection via Denoising DiffusionabstractGeneric event boundary detection (GEBD) aims to identify natural boundaries in a video, segmenting it into distinct and meaningful chunks. Despite the inherent subjectivity of event boundaries, previous methods have focused on deterministic predictions, overlooking the diversity of plausible solutions. In this paper, we introduce a novel diffusion-based boundary detection model, dubbed DiffGEBD, that tackles the problem of GEBD from a generative perspective. The proposed model encodes relevant changes across adjacent frames via temporal self-similarity and then iteratively decodes random noise into plausible event boundaries being conditioned on the encoded features. Classifier-free guidance allows the degree of diversity to be controlled in denoising diffusion. In addition, we introduce a new evaluation metric to assess the quality of predictions considering both diversity and fidelity. Experiments show that our method achieves strong performance on two standard benchmarks, Kinetics-GEBD and TAPOS, generating diverse and plausible event boundaries. Jaejun Hwang, Dayoung Gong, Manjin Kim, Minsu Cho |
ICCV | 2 |
| 2024 | ActFusion: a Unified Diffusion Model for Action Segmentation and AnticipationabstractTemporal action segmentation and long-term action anticipation are two popular vision tasks for the temporal analysis of actions in videos.
Despite apparent relevance and potential complementarity, these two problems have been investigated as separate and distinct tasks. In this work, we tackle these two problems, action segmentation, and action anticipation, jointly using a unified diffusion model dubbed ActFusion.
The key idea to unification is to train the model to effectively handle both visible and invisible parts of the sequence in an integrated manner;
the visible part is for temporal segmentation, and the invisible part is for future anticipation.
To this end, we introduce a new anticipative masking strategy during training in which a late part of the video frames is masked as invisible, and learnable tokens replace these frames to learn to predict the invisible future.
Experimental results demonstrate the bi-directional benefits between action segmentation and anticipation.
ActFusion achieves the state-of-the-art performance across the standard benchmarks of 50 Salads, Breakfast, and GTEA, outperforming task-specific models in both of the two tasks with a single unified model through joint learning. Dayoung Gong, Suha Kwak, Minsu Cho |
NeurIPS | 1 |
| 2023 | Activity Grammars for Temporal Action SegmentationabstractSequence prediction on temporal data requires the ability to understand compositional structures of multi-level semantics beyond individual and contextual properties of parts. The task of temporal action segmentation remains challenging for the reason, aiming at translating an untrimmed activity video into a sequence of action segments.
This paper addresses the problem by introducing an effective activity grammar to guide neural predictions for temporal action segmentation.
We propose a novel grammar induction algorithm, dubbed KARI, that extracts a powerful context-free grammar from action sequence data. We also develop an efficient generalized parser, dubbed BEP, that transforms frame-level probability distributions into a reliable sequence of actions according to the induced grammar with recursive rules.
Our approach can be combined with any neural network for temporal action segmentation to enhance the sequence prediction and discover its compositional structure.
Experimental results demonstrate that our method significantly improves temporal action segmentation in terms of both performance and interpretability on two standard benchmarks, Breakfast and 50 Salads. Dayoung Gong, Joonseok Lee, Deunsol Jung, Suha Kwak, Minsu Cho |
NeurIPS | 1 |
| 2022 | Future Transformer for Long-term Action AnticipationabstractThe task of predicting future actions from a video is crucial for a real-world agent interacting with others. When anticipating actions in the distant future, we humans typically consider long-term relations over the whole sequence of actions, i.e., not only observed actions in the past but also potential actions in the future. In a similar spirit, we propose an end-to-end attention model for action anticipation, dubbed Future Transformer (FUTR), that leverages global attention over all input frames and output tokens to predict a minutes-long sequence of future actions. Unlike the previous autoregressive models, the proposed method learns to predict the whole sequence of future actions in parallel decoding, enabling more accurate and fast inference for long-term anticipation. We evaluate our method on two standard benchmarks for long-term action anticipation, Breakfast and 50 Salads, achieving state-of-the-art results. Dayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha, Minsu Cho |
CVPR | 1 |