VLDB 2026 Research / reviewers in the wild / expert
Guodong Ding
dblp:54/5798
· DBLP profile ↗
18ranked-venue papers
9as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 7 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Condensing Action Segmentation Datasets via Generative Network InversionabstractThis work presents the first condensation approach for procedural video datasets used in temporal action segmentation. We propose a condensation framework that leverages generative prior learned from the dataset and network inversion to condense data into compact latent codes with significant storage reduced across temporal and channel aspects. Orthogonally, we propose sampling diverse and representative action sequences to minimize video-wise redundancy. Our evaluation on standard benchmarks demonstrates consistent effectiveness in condensing TAS datasets and achieving competitive performances. Specifically, on the Breakfast dataset, our approach reduces storage by over 500× while retaining 83% of the performance compared to training with the full dataset. Furthermore, when applied to a downstream incremental learning task, it yields superior performance compared to the state-of-the-art. Guodong Ding, Rongyu Chen, Angela Yao |
CVPR | 1 |
| 2025 | A Temporal Modeling Framework for Video Pre-Training on Video Instance SegmentationabstractContemporary Video Instance Segmentation (VIS) methods typically adhere to a pre-train then fine-tune regime, where a segmentation model trained on images is fine-tuned on videos. However, the lack of temporal knowledge in the pre-trained model introduces a domain gap which may adversely affect the VIS performance. To effectively bridge this gap, we present a novel "video pre-training" approach to enhance VIS models, especially for videos with intricate instance relationships. Our crucial innovation focuses on reducing disparities between the pre-training and fine-tuning stages. Specifically, we first introduce consistent pseudo-video augmentations to create diverse pseudo-video samples for pre-training while maintaining the instance consistency across frames. Then, we incorporate a multi-scale temporal module to enhance the model’s ability to model temporal relations through self- and cross-attention at short- and long-term temporal spans. Our approach does not set constraints on model architecture and can integrate seamlessly with various VIS methods. Experiment results on commonly adopted VIS benchmarks show that our method consistently outperforms state-of-the-art methods. Our approach achieves a notable 4.0% increase in average precision on the challenging OVIS dataset Peng-Tao Jiang, Guodong Ding, Kaiqi Huang |
ICME | 4 |
| 2025 | Contrastive Masked Video Modeling for Coronary Angiography Diagnosis
Zhiming Shao, Yingqian Zhang 0003, Zechen Wei, Guodong Ding, Yundai Chen, Jie Tian 0001, Hui Hui |
MICCAI (9) | 6 |
| 2025 | Spatial and temporal beliefs for mistake detection in assembly tasksabstractAssembly tasks, as an integral part of daily routines and activities, involve a series of sequential steps that are prone to error. This paper proposes a novel method for identifying ordering mistakes in assembly tasks based on knowledge-grounded beliefs. The beliefs comprise spatial and temporal aspects, each serving a unique role. Spatial beliefs capture the structural relationships among assembly components and indicate their topological feasibility. Temporal beliefs model the action preconditions and enforce sequencing constraints. Furthermore, we introduce a learning algorithm that dynamically updates and augments the belief sets online. To evaluate, we first test our approach in deducing predefined rules on synthetic data based on industry assembly. We also verify our approach on the real-world Assembly101 dataset, enhanced with annotations of component information. Our framework achieves superior performance in detecting ordering mistakes under both synthetic and real-world settings, highlighting the effectiveness of our approach. • We present two belief sets for the assembly tasks. • We propose a novel mistake detection framework for assembly tasks. • The framework can better detect ordering mistakes and can be integrated with perception modules. Guodong Ding, Fadime Sener, Shugao Ma, Angela Yao |
Comput. Vis. Image Underst. | 1 |
| 2024 | Coherent Temporal Synthesis for Incremental Action SegmentationabstractData replay is a successful incremental learning technique for images. It prevents catastrophic forgetting by keeping a reservoir of previous data, original or synthe-sized, to ensure the model retains past knowledge while adapting to novel concepts. However, its application in the video domain is rudimentary, as it simply stores frame ex-emplars for action recognition. This paper presents the first exploration of video data replay techniques for incremen-tal action segmentation, focusing on action temporal modeling. We propose a Temporally Coherent Action (TCA) model, which represents actions using a generative model instead of storing individual frames. The integration of a conditioning variable that captures temporal coherence al-lows our model to understand the evolution of action features over time. Therefore, action segments generated by TCA for replay are diverse and temporally coherent. In a 10-task incremental setup on the Breakfast dataset, our approach achieves significant increases in accuracy for up to 22% compared to the baselines. Guodong Ding, Hans Golong, Angela Yao |
CVPR | 1 |
| 2024 | LMTextSpotter: Towards Better Scene Text Spotting with Language Modeling in Transformer
Guodong Ding |
ICDAR (5) | 2 |
| 2024 | OnlineTAS: An Online Baseline for Temporal Action SegmentationabstractTemporal context plays a significant role in temporal action segmentation. In an offline setting, the context is typically captured by the segmentation network after observing the entire sequence. However, capturing and using such context information in an online setting remains an under-explored problem. This work presents the first online framework for temporal action segmentation. At the core of the framework is an adaptive memory designed to accommodate dynamic changes in context over time, alongside a feature augmentation module that enhances the frames with the memory. In addition, we propose a post-processing approach to mitigate the severe over-segmentation in the online setting. On three common segmentation benchmarks, our approach achieves state-of-the-art performance. Guodong Ding, Angela Yao |
NeurIPS | 2 |
| 2024 | Temporal Action Segmentation: An Analysis of Modern TechniquesabstractTemporal action segmentation (TAS) in videos aims at densely identifying video frames in minutes-long videos with multiple action classes. As a long-range video understanding task, researchers have developed an extended collection of methods and examined their performance using various benchmarks. Despite the rapid growth of TAS techniques in recent years, no systematic survey has been conducted in these sectors. This survey analyzes and summarizes the most significant contributions and trends. In particular, we first examine the task definition, common benchmarks, types of supervision, and prevalent evaluation measures. In addition, we systematically investigate two essential techniques of this topic, i.e., frame representation and temporal modeling, which have been studied extensively in the literature. We then conduct a thorough review of existing TAS works categorized by their levels of supervision and conclude our survey by identifying and emphasizing several research gaps. Guodong Ding, Fadime Sener, Angela Yao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Rapid Person Re-Identification via Sub-space Consistency Regularization
Qingze Yin, Guan'an Wang, Guodong Ding, Qilei Li, Shaogang Gong, Zhenmin Tang |
Neural Process. Lett. | 3 |
| 2023 | Temporal Action Segmentation With High-Level Complex Activity LabelsabstractThe temporal action segmentation task segments videos temporally and predicts action labels for all frames. Fully supervising such a segmentation model requires dense frame-wise action annotations, which are expensive and tedious to collect. This work is the first to propose a Constituent Action Discovery (CAD) framework that only requires the video-wise high-level complex activity label as supervision for temporal action segmentation. The proposed approach automatically discovers constituent video actions using an activity classification task. Specifically, we define a finite number of latent action prototypes to construct video-level dual representations with which these prototypes are learned collectively through the activity classification training. This setting endows our approach with the capability to discover potentially shared actions across multiple complex activities. Due to the lack of action-level supervision, we adopt the Hungarian matching algorithm to relate latent action prototypes to ground truth semantic classes for evaluation. We show that with the high-level supervision, the Hungarian matching can be extended from the existing video and activity levels to the global level. The global-level matching allows for action sharing across activities, which has never been considered in the literature before. Extensive experiments demonstrate that our discovered actions can help perform temporal action segmentation and activity recognition tasks. Guodong Ding, Angela Yao |
IEEE Trans. Multim. | 1 |
| 2022 | Leveraging Action Affinity and Continuity for Semi-supervised Temporal Action Segmentation
Guodong Ding, Angela Yao |
ECCV (35) | 1 |
| 2021 | Multi-View Label Prediction for Unsupervised Learning Person Re-IdentificationabstractPerson re-identification (ReID) aims to match pedestrian images across disjoint cameras. Existing supervised ReID methods utilize deep networks and train them with identity-labeled images, which suffer from limited annotations. Recently, clustering-based unsupervised ReID attracts more and more attention. It first clusters unlabeled images and assigns cluster index to the pseudo-identity-labels, then trains a ReID model with the pseudo-identity-labels. However, considering the slight inter-class variations and significant intra-class variations, pseudo-identity-labels learned from clustering algorithms are usually noisy and coarse. To alleviate the problems above, besides clustering pseudo-identity-labels, we propose to learn pseudo-patch-labels, which brings two advantages: (1) Patch naturally alleviates the effect of backgrounds, occlusions, and carryings since they usually occupy small parts in images, thus overcome noisy labels. (2) It is plausible that patches from different pedestrians belong to the same pseudo-identity-label. For example, pedestrians have a high probability of wearing either the same shoes or pants but a low possibility of wearing both. The experiments demonstrate our proposed method achieves the best performance by a large margin on both image- and video-based datasets. Qingze Yin, Guan'an Wang, Guodong Ding, Shaogang Gong, Zhenmin Tang |
IEEE Signal Process. Lett. | 3 |
| 2020 | Feature mask network for person re-identification
Guodong Ding, Salman Khan 0001, Zhenmin Tang, Fatih Porikli |
Pattern Recognit. Lett. | 1 |
| 2019 | Dispersion based Clustering for Unsupervised Person Re-identification
Guodong Ding, Salman Khan 0001, Zhenmin Tang |
BMVC | 1 |
| 2019 | Feature Affinity-Based Pseudo Labeling for Semi-Supervised Person Re-IdentificationabstractVision-based person re-identification aims to match a person's identity across multiple images, which is a fundamental task in multimedia content analysis and retrieval. Deep neural networks have recently manifested great potential in this task. However, a major bottleneck of existing supervised deep networks is their reliance on a large amount of annotated training data. Manual labeling for person identities in large-scale surveillance camera systems is quite challenging and incurs significant costs. Some recent studies adopt generative model outputs as training data augmentation. To more effectively use these synthetic data for an improved feature learning and re-identification performance, this paper proposes a novel feature affinity-based pseudo labeling method with two possible label encodings. To the best of our knowledge, this is the first study that employs pseudo-labeling by measuring the affinity of unlabeled samples with the underlying clusters of labeled data samples using the intermediate feature representations from deep networks. We propose training the network with the joint supervision of cross-entropy loss together with a center regularization term, which not only ensures discriminative feature representation learning but also simultaneously predicts pseudo-labels for unlabeled data. We show that both label encodings can be learned in a unified manner and help improve the overall performance. Our extensive experiments on three person re-identification datasets: Market-1501, DukeMTMC-reID, and CUHK03, demonstrate significant performance boost over the state-of-the-art person re-identification approaches. Guodong Ding, Shanshan Zhang 0001, Salman Khan 0001, Zhenmin Tang, Jian Zhang 0002, Fatih Porikli |
IEEE Trans. Multim. | 1 |
| 2010 | ECON: An Approach to Extract Content from Web News PageabstractThis paper provides a simple but effective approach, named ECON, to fully-automatically extract content from Web news page. ECON uses a DOM tree to represent the Web news page and leverages the substantial features of the DOM tree. ECON finds a snippet-node by which a part of the content of news is wrapped firstly, then backtracks from the snippet-node until a summary-node is found, and the entire content of news is wrapped by the summary-node. During the process of backtracking, ECON removes noise. Experimental results showed that ECON can achieve high accuracy and fully satisfy the requirements for scalable extraction. Moreover, ECON can be applied to Web news page written in many popular languages such as Chinese, English, French, German, Italian, Japanese, Portuguese, Russian, Spanish, Arabic. ECON can be implemented much easily. Yan Guo 0001, Huifeng Tang, Linhai Song, Yu Wang 0089, Guodong Ding |
APWeb | 5 |
| 2009 | ContentEx: A framework for automatic content extraction programsabstractWeb pages are often decorated with extraneous information (such as navigation bars, branding banners, JavaScript and advertisements). This kind of information may distract users from actual content they are really interested in and may reduce effects of many advanced Web applications. Automatic content extraction has many applications ranging from providing data for Web mining to realizing better accessing the Web over mobile devices. In this paper, we propose ContentEx, a framework for automatic content extraction programs, which we use to organize codes of automatic content extraction programs and to facilitate the development of related solutions. We also introduce how we extract content from forum pages in this framework to fulfill the requirement from our actual application. Linhai Song, Xueqi Cheng 0001, Yan Guo 0001, Guodong Ding |
ISI | 5 |
| 2009 | Facilitating wrapper generation with page analysisabstractCurrent approaches for generating wrappers for web page extraction suffer from the requirement of huge amount of labeled training pages to obtain satisfying results. On the other hand, the quality of data extracted by fully automatic methods is not reliable. In this paper, we propose a novel method to facilitate wrapper generation by combining wrapper induction and page analysis approaches. In addition to manually labeled data, we also take advantage of a set of unlabeled pages to improve the quality of induced wrappers. Our experiments demonstrate that our system achieves a satisfying result with fewer manually labeled training pages. Xueqi Cheng 0001, Yu Wang 0009, Guodong Ding |
ISI | 5 |