EDBT 2026 Demo / reviewers in the wild / expert
Zhaobo Qi
dblp:276/3128
· DBLP profile ↗
23ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0001-9196-9818ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery DetectionabstractAs forgery types continue to emerge consistently, Incremental Face Forgery Detection (IFFD) has become a crucial paradigm. However, existing methods typically rely on data replay or coarse binary supervision, which fails to explicitly constrain the feature space, leading to severe feature drift and catastrophic forgetting. To address this, we propose AIFIND, Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection, which leverages semantic anchors to stabilize incremental learning. We design the Artifact-Driven Semantic Prior Generator to instantiate invariant semantic anchors, establishing a fixed coordinate system from low-level artifact cues. These anchors are injected into the image encoder via Artifact-Probe Attention, which explicitly constrains volatile visual features to align with stable semantic anchors. Adaptive Decision Harmonizer harmonizes the classifiers by preserving angular relationships of semantic anchors, maintaining geometric consistency across tasks. Extensive experiments on multiple incremental protocols validate the superiority of AIFIND. Hao Wang 0035, Beichen Zhang 0006, Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu 0008, Weigang Zhang |
ICMR | 5 |
| 2026 | Distinguishing semantically similar queries in temporal video grounding via LLM-generated query
Yibo Dang, Zhaobo Qi, Xinyan Liu 0008, Xinzhe Han, Weigang Zhang |
Multim. Syst. | 2 |
| 2026 | Multimodal-guided mixture-of-experts bias removal strategy for natural language video localization
Xiaowen Ruan, Zhaobo Qi, Ruisi Chen, Yuanrong Xu, Beichen Zhang 0006, Weigang Zhang |
Multim. Syst. | 2 |
| 2025 | Procedure Knowledge Decoupled Distillation Strategy for Procedure Planning in Instructional VideosabstractProcedure planning in instructional videos, producing a structured and plannable action sequence facilitating the transition from the start to the goal states, has achieved significant progress. The dominant single-branch non-autoregressive planning paradigm guides action sequence generation through action labels, overlooking the limitation of the absence of intermediate visual information. Hence, we introduce the procedure knowledge decoupled distillation strategy to address the above issue. This innovative strategy deliberately lets the teacher model see the real visual information among the start and goal states to enhance its action semantic understanding and relationship modeling ability, producing the potential probability distribution containing the real action class and other action classes that may occur. Accordingly, we introduce a decoupled intermediate information knowledge distillation loss, which comprises single action knowledge distillation and sequence distribution knowledge distillation for the student model. The former improves the student model's precise inference ability for individual actions by transferring knowledge of a single action target category using binary classification loss. Conversely, the latter uses MSE loss to constrain the student model to learn the action sequence probability distribution from the teacher model, thereby enhancing the student model's global planning capability. Extensive experiments on three datasets demonstrate that our strategy can improve the performance of multiple weakly supervised models, achieving promising procedure knowledge modeling ability and plug-and-play flexibility. Xiaotian Pan, Zhaobo Qi, Yuanrong Xu, Weigang Zhang |
AAAI | 2 |
| 2025 | Video Language Model Pretraining with Spatio-temporal MaskingabstractThe development of self-supervised video-language models based on mask learning has significantly advanced downstream video tasks. These models leverage masked reconstruction to facilitate joint learning of visual and linguistic information. However, recent study reveals that reconstructing image features yields superior downstream performance compared to video feature reconstruction. We hypothesize that this performance gap stems from the way how masking strategies influence the model’s attention to temporal dynamics. To validate this hypothesis, we performed two sets of experiments that demonstrate that alignment between the masked target and the reconstruction target is crucial for self-supervised video-language learning. Based on these findings, we propose a spatio-temporal masking strategy (STM) for video-language model pretraining that operates across adjacent frames, and a decoder leverages semantic information to enhance the spatio-temporal representations of masked tokens. Thanks to the combination of masking strategy and reconstruction decoder, STM enforces the model to learn spatio-temporal feature representation comprehensively. Experiments in three video understanding downstream tasks validate the superiority of our method. Codes are available here. Zhaobo Qi, Junshu Sun, Yaowei Wang 0001, Qingming Huang, Shuhui Wang |
CVPR | 2 |
| 2025 | Enhancing Pre-trained Representation Classifiability can Boost its InterpretabilityabstractThe visual representation of a pre-trained model prioritizes the classifiability on downstream tasks, while the widespread applications for pre-trained visual models have posed new requirements for representation interpretability. However, it remains unclear whether the pre-trained representations can achieve high interpretability and classifiability simultaneously. To answer this question, we quantify the representation interpretability by leveraging its correlation with the ratio of interpretable semantics within the representations. Given the pre-trained representations, only the interpretable semantics can be captured by interpretations, whereas the uninterpretable part leads to information loss. Based on this fact, we propose the Inherent Interpretability Score (IIS) that evaluates the information loss, measures the ratio of interpretable semantics, and quantifies the representation interpretability. In the evaluation of the representation interpretability with different classifiability, we surprisingly discover that the interpretability and classifiability are positively correlated, i.e., representations with higher classifiability provide more interpretable semantics that can be captured in the interpretations. This observation further supports two benefits to the pre-trained representations. First, the classifiability of representations can be further improved by fine-tuning with interpretability maximization. Second, with the classifiability improvement for the representations, we obtain predictions based on their interpretations with less accuracy degradation. The discovered positive correlation and corresponding applications show that practitioners can unify the improvements in interpretability and classifiability for pre-trained vision models. Codes are available at https://github.com/ssfgunner/IIS. Shufan Shen, Zhaobo Qi, Junshu Sun, Qingming Huang, Qi Tian 0001, Shuhui Wang |
ICLR | 2 |
| 2025 | Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video RetrievalabstractWith the explosive growth of video data, finding videos that meet detailed requirements in large datasets has become a challenge. To address this, the composed video retrieval task has been introduced, enabling users to retrieve videos using complex queries that involve both visual and textual information. However, the inherent heterogeneity between the modalities poses significant challenges. Textual data are highly abstract, while video content contains substantial redundancy. The modality gap in information representation makes existing methods struggle with the modality fusion and alignment required for fine-grained composed retrieval. To overcome these challenges, we first introduce FineCVR-1M, a fine-grained composed video retrieval dataset containing 1,010,071 video-text triplets with detailed textual descriptions. This dataset is constructed through an automated process that identifies key concept changes between video pairs to generate textual descriptions for both static and action concepts. For fine-grained retrieval methods, the key challenge lies in understanding the detailed requirements. Text description serves as clear expressions of intent, but it requires models to distinguish subtle differences in the description of video semantics. Therefore, we propose a textual Feature Disentanglement and Cross-modal Alignment framework (FDCA) that disentangles features at both the sentence and token levels. At the sequence level, we separate text features into retained and injected features. At the token level, an Auxiliary Token Disentangling mechanism is proposed to disentangle texts into retained, injected, and excluded tokens. The disentanglement at both levels extracts fine-grained features, which are aligned and fused with the reference video to extract global representations for video retrieval. Experiments on FineCVR-1M dataset demonstrate the superior performance of FDCA. Our code and dataset are available at: https://may2333.github.io/FineCVR/. Zhaobo Qi, Yiling Wu, Junshu Sun, Yaowei Wang 0001, Shuhui Wang |
ICLR | 2 |
| 2025 | Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional VideosabstractIn this paper, we address the challenge of procedure planning in instructional videos, aiming to generate coherent and task-aligned action sequences from start and end visual observations. Previous work has mainly relied on text-level supervision to bridge the gap between observed states and unobserved actions, but it struggles with capturing intricate temporal relationships among actions. Building on these efforts, we propose the Masked Temporal Interpolation Diffusion (MTID) model that introduces a latent space temporal interpolation module within the diffusion model. This module leverages a learnable interpolation matrix to generate intermediate latent features, thereby augmenting visual supervision with richer mid-state details. By integrating this enriched supervision into the model, we enable end-to-end training tailored to task-specific requirements, significantly enhancing the model's capacity to predict temporally coherent action sequences. Additionally, we introduce an action-aware mask projection mechanism to restrict the action generation space, combined with a task-adaptive masked proximity loss to prioritize more accurate reasoning results close to the given start and end states over those in intermediate steps. Simultaneously, it filters out task-irrelevant action predictions, leading to contextually aware action sequences. Experimental results across three widely used benchmark datasets demonstrate that our MTID achieves promising action planning performance on most metrics. Zhaobo Qi, Lingshuai Lin, Junqi Jing, Tingting Chai, Beichen Zhang 0006, Shuhui Wang, Weigang Zhang |
ICLR | 2 |
| 2025 | Combatting Data Imbalance and Noise in Micro-Action Recognition
Weidong Chen 0010, Zhaobo Qi, Pengqi Huang, Xinyan Liu 0008, Weigang Zhang |
ACM Multimedia | 5 |
| 2025 | Dual-guided multi-modal bias removal strategy for temporal sentence grounding in video
Xiaowen Ruan, Zhaobo Qi, Yuanrong Xu, Weigang Zhang |
Multim. Syst. | 2 |
| 2025 | KN-VLM: KNowledge-guided Vision-and-Language Model for visual abductive reasoning
Kuo Tan, Zhaobo Qi, Jianping Zhong, Yuanrong Xu, Weigang Zhang |
Multim. Syst. | 2 |
| 2025 | Uncertainty-Aware Mixture of Experts for Video Action AnticipationabstractAnticipating future actions in daily life videos is crucial for seamless human-machine collaboration. However, accurately predicting these actions is challenging due to the inherent uncertainty and non-determinism of future events. To address this, we propose the uncertainty-aware mixture-of-experts framework for action anticipation (AntMoE), which employs multiple anticipation experts to model diverse video evolution patterns through learnable expert embeddings. These anticipation experts generate diverse predictions by integrating the top-k semantically similar observed video frames related to the current predicted feature representation, along with their corresponding expert embeddings. An anticipation router then aggregates these predictions based on the relationship between the current feature representation and all expert embeddings. To enhance the effectiveness of AntMoE, we introduce an expert regularization loss with three components: orthogonal loss promotes orthogonality among expert embeddings; expert balance loss ensures equal activation of all experts during training; and stability loss encourages the generation of numerically stable aggregation weights. Additionally, we incorporate an anticipation ranking loss function that aligns the model’s confidence across varying anticipation time durations with the ground-truth ranking order, where a shorter anticipation time length corresponds to a higher confidence level. Experimental results across multiple benchmarks demonstrate that our method achieves remarkable anticipation performance. Zhaobo Qi, Shuhui Wang, Weigang Zhang, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | VPA: Multi-Modal Virtual Point Augmentation for 3D Object DetectionabstractIntegrating LiDAR and camera data is crucial for precise 3D object detection. Existing methods resort to augmenting virtual points from 2D image space in a random manner to complete the appearance of 3D objects with sparse points. However, these augmented virtual points have unreasonable 3D positions and representations, which brings serious negative effects on accurate detection. To this end, we introduce a general 3D object detection framework called Virtual Point Augmenting (VPA) to enrich the 3D point cloud by controllably generating virtual points with accurate depth and position information as well as domain-gap-eliminated multi-modal representations from image and point cloud spaces. VPA contains two core designs, namely Hybrid Sampling Method (HSM) and Fine-Grained Cross-modal Fusion (FGCF). HSM uses the constructed seed point distribution map based on the edge score and mask score map to sample high-quality seed points, and employs a feature similarity function to sample withkneighbors’ depth to obtain more accurate depth for the seed points, thereby enhancing the quality of the virtual points’ 3D positions. FGCF fuses the multi-modal features,i.e., the semantic feature, the geometric feature from the image space, and the 3D position feature in an adaptive manner using self-attention mechanism, thereby further improving the representation of the virtual points. We apply VPA to the LiDAR-based method CenterPoint and fusion-based method Cross-modal transformer. Experimental results on the nuScenes, KITTI, and Waymo benchmarks validate the efficiency of our VPA, which achieves promising performance with 72.9% mAP and 74.8% NDS without using test-time augmentation and model ensemble techniques on the nuScenes test set. Code is available at https://github.com/jianpingZhonggit/vpa.git. Jianping Zhong, Zhaobo Qi, Kaiwen Duan, Yuanrong Xu, Weigang Zhang, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Multi-Modal 3D Object Detector with Object-Guided Fusion and Hierarchical Sample SelectionabstractAccurately detecting objects in 3D scenes is crucial for autonomous driving. Although existing voxel-based methods have achieved remarkable progress, their performance on tail objects remains unsatisfactory. We identify two core issues contributing to this phenomenon: the detectors frequently misidentify some background elements as foreground objects, and there is a misalignment between the classification score and detection quality. To tackle these challenges, we introduce an object-level – guided multi-modal 3D object detector with an object-guided feature fusion (OFF) module and a hierarchical sample selection (HSS) strategy, named OGMMDet. Specifically, OFF introduces rich image features to enhance the representation of objects while using an object distribution heatmap to suppress the background. This approach provides geometry clues for tail objects while providing category priors to filter out the background. HSS uses a local-to-global ranking approach to calculate the relative classification loss weights of all proposals. It assigns higher weights to proposals with higher IoU when optimizing classification branches. This ensures that the model focuses its optimization on these higher-quality proposals. Consequently, there is a positive correlation between the classification score and IoU. This method alleviates the misalignment between the classification score and detection quality. Extensive experiments on the KITTI and nuScenes benchmarks demonstrate the effectiveness of our OGMMDet, which achieves 45.61% and 68.96% mean average precision (mAP) on pedestrians and cyclists on the KITTI benchmark, respectively. Code is available at https://github.com/ZhongJianPing1/ogmmdet.git . Jianping Zhong, Zhaobo Qi, Kaiwen Duan, Yuanrong Xu, Weigang Zhang, Qingming Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in VideoabstractTemporal Sentence Grounding in Video (TSGV) is troubled by dataset bias issue, which is caused by the uneven temporal distribution of the target moments for samples with similar semantic components in input videos or query texts. Existing methods resort to utilizing prior knowledge about bias to artificially break this uneven distribution, which only removes a limited amount of significant language biases. In this work, we propose the bias-conflict sample synthesis and adversarial removal debias strategy (BSSARD), which dynamically generates bias-conflict samples by explicitly leveraging potentially spurious correlations between single-modality features and the temporal position of the target moments. Through adversarial training, its bias generators continuously introduce biases and generate bias-conflict samples to deceive its grounding model. Meanwhile, the grounding model continuously eliminates the introduced biases, which requires it to model multi-modality alignment information. BSSARD will cover most kinds of coupling relationships and disrupt language and visual biases simultaneously. Extensive experiments on Charades-CD and ActivityNet-CD demonstrate the promising debiasing capability of BSSARD. Source codes are available at https://github.com/qzhb/BSSARD. Zhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang, Weigang Zhang, Qingming Huang |
AAAI | 1 |
| 2024 | Improving Sequential DeepFake Detection with Local information enhancement
Longyun Dong, Yuanrong Xu, Jianping Zhong, Zhaobo Qi, Weigang Zhang |
MMAsia | 4 |
| 2024 | Uncertainty-Boosted Robust Video Activity AnticipationabstractVideo activity anticipation aims to predict what will happen in the future, embracing a broad application prospect ranging from robot vision and autonomous driving. Despite the recent progress, the data uncertainty issue, reflected as the content evolution process and dynamic correlation in event labels, has been somehow ignored. This reduces the model generalization ability and deep understanding on video content, leading to serious error accumulation and degraded performance. In this paper, we address the uncertainty learning problem and propose an uncertainty-boosted robust video activity anticipation framework, which generates uncertainty values to indicate the credibility of the anticipation results. The uncertainty value is used to derive a temperature parameter in the softmax function to modulate the predicted target activity distribution. To guarantee the distribution adjustment, we construct a reasonable target activity label representation by incorporating the activity evolution from the temporal class correlation and the semantic relationship. Moreover, we quantify the uncertainty into relative values by comparing the uncertainty among sample pairs and their temporal-lengths. This relative strategy provides a more accessible way in uncertainty modeling than quantifying the absolute uncertainty values on the whole dataset. Experiments on multiple backbones and benchmarks show our framework achieves promising performance and better robustness/interpretability. Zhaobo Qi, Shuhui Wang, Weigang Zhang, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Collaborative Debias Strategy for Temporal Sentence Grounding in VideoabstractTemporal sentence grounding in video has witnessed significant advancements, but suffers from substantial dataset bias, which undermines its generalization ability. Existing debias approaches primarily concentrate on well-known distribution and linguistic biases, while overlooking the relationship among different biases, limiting their debias capability. In this work, we delve into the existence of visual bias and combinatorial bias in the widely used datasets, and introduce a collaborative debias structure that can be seamlessly integrated into present methods. It encompasses four low-capacity models, a re-label module, and a main model. Each biased model deliberately leverages bias as shortcut information to accurately perform grounding, achieved by customizing the appropriate model structure and input data format to align with the bias characteristics. During the training phase, the gradient descent direction for optimizing the main model should align with the negative gradient descent direction of the biased model that is optimized by utilizing ground truth labels. Subsequently, the re-label module introduces a gradient aggregation function, consolidating the gradient descent direction from these biased models and constructing new labels to compel the main model to effectively capture multi-modality alignment features instead of relying on shortcut contents for grounding. Finally, we design two debias structures, P-Debias and C-Debias, to exploit the independence and inclusion relationships between different types of biases. Extensive experiments on multiple span-based models over Charades-CD and ActivityNet-CD demonstrate the exceptional debias capability of our strategy (https://github.com/qzhb/CDS). Zhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang, Weigang Zhang, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Semantic-Aware Dynamic Feature Selection and Fusion for Object Detection in UAV VideosabstractKeypoint-based detectors perform well in surveillance videos but face challenges in detecting objects in UAV videos due to missed corners and mismatches. To address this, we propose a semantic-aware module with a feature fusion sub-module and a feature selection sub-module. The feature fusion module adaptively combines low-level and high-level features, enhancing corner recall. The feature selection module determines spatial location importance, improving discriminative capabilities and reducing background interference, resulting in better precision. Experiments on the UAVDT benchmark show our method achieves competitive results. Notably, our method improves corner recall by 4.0% and reduces the mismatch rate by 2.9% compared to the baseline. Code is available at https://github.com/jianpingZhonggit/SemanticAwareModule. Jianping Zhong, Zhaobo Qi, Weigang Zhang, Qingming Huang |
MMAsia | 2 |
| 2023 | Self-Regulated Learning for Egocentric Video Activity AnticipationabstractFuture activity anticipation is a challenging problem in egocentric vision. As a standard future activity anticipation paradigm, recursive sequence prediction suffers from the accumulation of errors. To address this problem, we propose a simple and effective Self-Regulated Learning framework, which aims to regulate the intermediate representation consecutively to produce representation that (a) emphasizes the novel information in the frame of the current time-stamp in contrast to previously observed content, and (b) reflects its correlation with previously observed frames. The former is achieved by minimizing a contrastive loss, and the latter can be achieved by a dynamic reweighing mechanism to attend to informative frames in the observed content with a similarity comparison between feature of the current frame and observed frames. The learned final video representation can be further enhanced by multi-task learning which performs joint feature learning on the target activity labels and the automatically detected action and object class tokens. SRL sharply outperforms existing state-of-the-art in most cases on two egocentric video datasets and two third-person video datasets. Its effectiveness is also verified by the experimental fact that the action and object concepts that support the activity semantics can be accurately identified. Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Temporal Dynamic Concept Modeling Network for Explainable Video Event RecognitionabstractRecently, with the vigorous development of deep learning and multimedia technology, intelligent urban computing has received more and more extensive attention from academia and industry. Unfortunately, most of the related technologies are black-box paradigms that lack interpretability. Among them, video event recognition is a basic technology. Event contains multiple concepts and their rich interactions, which can assist us to construct explainable event recognition methods. However, the crucial concepts needed to recognize events have various temporal existing patterns, and the relationship between events and the temporal characteristics of concepts has not been fully exploited. This brings great challenges for concept-based event categorization. To address the above issues, we introduce the temporal concept receptive field, which is the length of the temporal window size required to capture key concepts for concept-based event recognition methods. Accordingly, we introduce the temporal dynamic convolution (TDC) to model the temporal concept receptive field dynamically according to different events. Its core idea is to combine the results of multiple convolution layers with the learned coefficients from two complementary perspectives. These convolution layers contain a variety of kernel sizes, which can provide temporal concept receptive fields of different lengths. Similarly, we also propose the cross-domain temporal dynamic convolution (CrTDC) with the help of the rich relationship between different concepts. Different coefficients can help us to capture suitable temporal concept receptive field sizes and highlight crucial concepts to obtain accurate and complete concept representations for event analysis. Based on the TDC and CrTDC, we introduce the temporal dynamic concept modeling network (TDCMN) for explainable video event recognition. We evaluate TDCMN on large-scale and challenging datasets FCVID, ActivityNet, and CCV. Experimental results show that TDCMN significantly improves the event recognition performance of concept-based methods, and the explainability of our method inspires us to construct more explainable models from the perspective of the temporal concept receptive field. Weigang Zhang, Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Towards More Explainability: Concept Knowledge Mining Network for Event RecognitionabstractEvent recognition of untrimmed video is a challenging task due to the big gap between low level visual features and event semantics. Beyond feature learning via deep neural networks, some recent works focus on analyzing event videos using concept-based representation. However, these methods simply aggregate the concept representation vectors of frames or segments, which inevitably introduces information loss on video-level concept knowledge. Moreover, the diversified relation between different concept domains (e.g., scene, object and action) has not been fully explored. To address the above issues, we propose a concept knowledge mining network (CKMN) for event recognition. CKMN is composed of an intra-domain concept knowledge mining subnetwork (IaCKM) and an inter-domain concept knowledge mining subnetwork~(IrCKM). IaCKM aims to obtain a complete concept representation by mining the existing pattern of each concept at different time granularities with dilated temporal pyramid convolution and temporal self-attention, while IrCKM explores the interaction between different types of concepts with co-attention style learning. We evaluate our method on FCVID and ActivityNet datasets. Experimental results show the effectiveness and better interpretability of our model on event analytics. Code is available at https://github.com/qzhb/CKMN. Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2020 | Modeling Temporal Concept Receptive Field Dynamically for Untrimmed Video AnalysisabstractEvent analysis in untrimmed videos has attracted increasing attention due to the application of cutting-edge techniques such as CNN. As a well studied property for CNN-based models, the receptive field is a measurement for measuring the spatial range covered by a single feature response, which is crucial in improving the image categorization accuracy. In video domain, video event semantics are actually described by complex interaction among different concepts, while their behaviors vary drastically from one video to another, leading to the difficulty in concept-based analytics for accurate event categorization. To model the concept behavior, we study temporal concept receptive field of concept-based event representation, which encodes the temporal occurrence pattern of different mid-level concepts. Accordingly, we introduce temporal dynamic convolution (TDC) to give stronger flexibility to concept-based event analytics. TDC can adjust the temporal concept receptive field size dynamically according to different inputs. Notably, a set of coefficients are learned to fuse the results of multiple convolutions with different kernel widths that provide various temporal concept receptive field sizes. Different coefficients can generate appropriate and accurate temporal concept receptive field size according to input videos and highlight crucial concepts. Based on TDC, we propose the temporal dynamic concept modeling network~(TDCMN) to learn an accurate and complete concept representation for efficient untrimmed video analysis. Experiment results on FCVID and ActivityNet show that TDCMN demonstrates adaptive event recognition ability conditioned on different inputs, and improve the event recognition performance of Concept-based methods by a large margin. Code is available at https://github.com/qzhb/TDCMN. Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Weigang Zhang, Qingming Huang |
ACM Multimedia | 1 |