EDBT 2026 Demo / reviewers in the wild / expert
Junyu Gao 0002
dblp:153/4522-2
· DBLP profile ↗
66ranked-venue papers
20as first author
57since 2021 · last 2026
0000-0002-8105-5497ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 15 first-author · 45 since 2021Artificial intelligence and machine learning · 23 · 10 first-author · 20 since 2021Computer networks · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Comprehensive Survey on Evidential Deep Learning and its ApplicationsabstractReliable uncertainty estimation has become a crucial requirement for the industrial deployment of deep learning algorithms, particularly in high-risk applications such as autonomous driving and medical diagnosis. However, uncertainty estimation methods relying on deep ensembling or Bayesian neural networks typically entail significant computational overhead. To address this challenge, a novel paradigm called Evidential Deep Learning (EDL) has emerged, providing high-quality uncertainty estimation with minimal additional computation in a single forward pass. This survey provides a comprehensive overview of the current research on EDL, designed to offer readers a broad introduction to the field without assuming prior knowledge. Specifically, we first delve into the theoretical foundation of EDL, the subjective logic theory, and discuss its distinctions from other uncertainty estimation frameworks. We further present existing theoretical advancements in EDL from four perspectives: reformulating the evidence collection process, improving uncertainty estimation via OOD samples, delving into various training strategies, and evidential regression networks. Thereafter, we elaborate on its extensive applications across various machine learning paradigms and downstream tasks. In the end, an outlook on future directions for better performances and broader adoption of EDL is provided, highlighting potential research avenues. Junyu Gao 0002, Liangyu Xiang, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Occlusion-Aware Multi-Object Tracking via Joint Diffusion Motion Prediction and Appearance PurificationabstractOcclusion, a pervasive challenge in Multi-Object Tracking (MOT) within complex scenes, severely degrades tracking performance. Current methods still face numerous problems when handling occlusions. Motion prediction struggles to accommodate diverse motion patterns, resulting in failure during short-term occlusions. Concurrently, appearance features possess insufficient discriminative power under occluded conditions, which frequently leads to identity switches following long-term occlusion. To enhance the performance of MOT under such challenging conditions, we propose an innovative Occlusion-Aware Multi-Object Tracking via Joint Diffusion Motion Prediction and Appearance Purification (OAMOT). For short-term occlusions, a Diffusion Motion Recovery Model (DMRM) is developed, which integrates residual and noise diffusion branches to recover trajectories precisely under varied motion patterns. For long-term occlusions, the SAM Appearance Purification for ReID (SAPR) module is proposed; the module leverages the mask generation mechanism of SAM as foreground attention to enhance feature discriminability. Furthermore, a Lightweight Attention Predictor (LAP) is integrated into the ReID network to achieve SAM-quality attention during inference without significant computational overhead. Experimental results on the MOT17, MOT20, and DanceTrack datasets demonstrate that the proposed OAMOT method outperforms current state-of-the-art multi-object tracking techniques across multiple evaluation metrics. The code can be available on https://github.com/wangtuo111/OAMOT. Wen Guo 0003, Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | History-Guided Prompt Generation for Vision-and-Language NavigationabstractVision-and-language navigation (VLN) has garnered extensive attention in the field of embodied artificial intelligence. VLN involves time series information, where historical observations contain rich contextual knowledge and play a crucial role in navigation. However, current methods do not explicitly excavate the connection between rich contextual information in history and the current environment, and ignore adaptive learning of clues related to the current environment. Therefore, we explore a Prompt Learning-based strategy which adaptively mines information in history that is highly relevant to the current environment to enhance the agent's perception of the current environment and propose a history-guided prompt generation (HGPG) framework. Specifically, HGPG includes two parts, one is an entropy-based history acquisition module that assesses the uncertainty of the action probability distribution from the preceding step to determine whether historical information should be used at the current time step. The other part is the prompt generation module that transforms historical context into prompt vectors by sampling from an end-to-end learned token library. These prompt tokens serve as discrete, knowledge-rich representations that encode semantic cues from historical observations in a compact form, making them easier for the decision network to understand and utilize. In addition, we share the token library across various navigation tasks, mining common features between different tasks to improve generalization to unknown environments. Extensive experimental results on four mainstream VLN benchmarks (R2R, REVERIE, SOON, R2R-CE) demonstrate the effectiveness of our proposed method. Code is available at https://github.com/Wzmshdong/HGPG. Wen Guo 0003, Zongmeng Wang, Yufan Hu, Junyu Gao 0002 |
IEEE Trans. Cybern. | 4 |
| 2025 | Evidential Knowledge Distillation
Liangyu Xiang, Junyu Gao 0002, Changsheng Xu |
ICCV | 2 |
| 2025 | NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous EnvironmentsabstractVision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to execute sequential navigation actions in complex environments guided by natural language instructions. Current approaches often struggle with generalizing to novel environments and adapting to ongoing changes during navigation. Inspired by human cognition, we present NavMorph, a self-evolving world model framework that enhances environmental understanding and decision-making in VLN-CE tasks. NavMorph employs compact latent representations to model environmental dynamics, equipping agents with foresight for adaptive planning and policy refinement. By integrating a novel Contextual Evolution Memory, NavMorph leverages scene-contextual information to support effective navigation while maintaining online adaptability. Extensive experiments demonstrate that our method achieves notable performance improvements on popular VLN-CE benchmarks. Code is available at https://github.com/Feliciaxyao/NavMorph. Xuan Yao 0001, Junyu Gao 0002, Changsheng Xu |
ICCV | 2 |
| 2025 | Building Embodied EvoAgent: A Brain-inspired Paradigm for Bridging Multimodal Large Models and World ModelsabstractEmbodied artificial intelligence has rapidly developed under the impetus of multimodal learning, robotics, and cognitive science, demonstrating great potential in fields such as navigation and manipulation. However, building embodied agents that can robustly operate in diverse and dynamic environments still faces challenges, such as handling partial observability and environmental adaptability. Multimodal large language models (MLLMs) are vital for embodied intelligence due to their ability to process multimodal information, but they encounter difficulties in understanding spatial environments and performing dynamic decisions and evolution. Inspired by the functional specialization of the left and right hemispheres of the human brain, this paper proposes a brain-inspired learning and evolution paradigm for embodied agents. The method designs an embodied context-augmented MLLM to simulate the language processing and logical analysis capabilities of the left hemisphere, responsible for understanding instructions and visual scenes. At the same time, it constructs a perceptual context-guided world model based on the recurrent state space model to simulate the spatial perception and holistic thinking functions of the right hemisphere, capturing environmental dynamics and predicting future states. By simulating the communication function of the corpus callosum, we propose dynamic communication slots for efficient information exchange between MLLMs and the world model, which also allows the agent to quickly adapt to dynamic environments without requiring extensive computational resources. Experiments show that the proposed paradigm significantly improves the performance of embodied agents in a series of tasks and enhances their generalization ability in zero-shot tasks through embodied exploration experience and online evolution. Our project page is available at https://feliciaxyao.github.io/EvoAgent/. Junyu Gao 0002, Xuan Yao 0001, Yong Rui, Changsheng Xu |
ACM Multimedia | 1 |
| 2025 | Learning Evidential Delta Denoising Scores for Video Editing
Yufan Hu, Junyu Gao 0002, Bin Fan 0001, Hongmin Liu 0001 |
ACM Multimedia | 3 |
| 2025 | Cross-Modal Dual-Causal Learning for Long-Term Action RecognitionabstractLong-term action recognition (LTAR) is challenging due to extended temporal spans with complex atomic action correlations and visual confounders. Although vision-language models (VLMs) have shown promise, they often rely on statistical correlations instead of causal mechanisms. Moreover, existing causality-based methods address modal-specific biases but lack cross-modal causal modeling, limiting their utility in VLM-based LTAR. This paper proposes Cross-Modal Dual-Causal Learning (CMDCL), which introduces a structural causal model to uncover causal relationships between videos and label texts. CMDCL addresses cross-modal biases in text embeddings via textual causal intervention and removes confounders inherent in the visual modality through visual causal intervention guided by the debiased text. These dual-causal interventions enable robust action representations to address LTAR challenges. Experimental results on three benchmarks including Charades, Breakfast and COIN, demonstrate the effectiveness of the proposed model. Our code is available at https://github.com/xushaowu/CMDCL. Shaowu Xu, Xibin Jia, Junyu Gao 0002, Qianmei Sun, Jing Chang 0006, Chao Fan 0001 |
ACM Multimedia | 3 |
| 2025 | R2A2-MoE: Ridge Regression-Based Analytic Adaptation with Mixture of Experts for Continual Learning with Vision-Language Models
Quan Fang, Junyu Gao 0002, Yang Yang 0122 |
PRCV (1) | 3 |
| 2025 | Revisiting Essential and Nonessential Settings of Evidential Deep LearningabstractEvidential Deep Learning (EDL) is an emerging method for uncertainty estimation that provides reliable predictive uncertainty in a single forward pass, attracting significant attention. Grounded in subjective logic, EDL derives Dirichlet concentration parameters from neural networks to construct a Dirichlet probability density function (PDF), modeling the distribution of class probabilities. Despite its success, EDL incorporates several nonessential settings: In model construction, (1) a commonly ignored prior weight parameter is fixed to the number of classes, while its value actually impacts the balance between the proportion of evidence and its magnitude in deriving predictive scores. In model optimization, (2) the empirical risk features a variance-minimizing optimization term that biases the PDF towards a Dirac delta function, potentially exacerbating overconfidence. (3) Additionally, the structural risk typically includes a KL-divergence-minimizing regularization, whose optimization direction extends beyond the intended purpose and contradicts common sense, diminishing the information carried by the evidence magnitude. Therefore, we propose Re-EDL, a simplified yet more effective variant of EDL, by relaxing the nonessential settings and retaining the essential one, namely, the adoption of projected probability from subjective logic. Specifically, Re-EDL treats the prior weight as an adjustable hyperparameter rather than a fixed scalar, and directly optimizes the expectation of the Dirichlet PDF provided by deprecating both the variance-minimizing optimization term and the divergence regularization term. Extensive experiments and state-of-the-art performance validate the effectiveness of our method. Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Learning Probabilistic Presence-Absence Evidence for Weakly-Supervised Audio-Visual Event PerceptionabstractWith only video-level event labels, this paper targets at the task of weakly-supervised audio-visual event perception (WS-AVEP), which aims to temporally localize and categorize events that belong to each modality. Despite the recent progress, most existing approaches either ignore the unsynchronized property of audio-visual tracks or discount the complementary modality for explicit enhancement. We argue that, a modality should provide ample presence evidence for an event, while the complementary modality offers absence evidence as a reference. However, to learn reliable evidence, we face challenging uncertainties caused by weak supervision and the complicated audio-visual data itself. To this end, we propose to collect Probabilistic Presence-Absence Evidence (PPAE) in a unified framework. Specifically, by leveraging uni-modal and cross-modal representations, a probabilistic presence-absence evidence collector (PAEC) is designed. To learn the evidence in a reliable range, we propose a joint-modal mutual learning (JML) process, which calibrates the evidence of diverse audible, visible, and audi-visible events adaptively and dynamically. Extensive experiments show that our method surpasses state-of-the-arts (e.g., absolute gains of 3.1% and 4.2% in terms of event-level audio and visual metrics on the LLP dataset). Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Active Cross-Modal Domain AdaptationabstractMost cross-modal methods assume that training and testing data come from the same domain, which is often not the case in real-world scenarios due to cross-modal domain shifts and potential unknown concepts. Moreover, cross-modal shifts hinder the capture of unknown concepts, and the presence of unknown concepts can in turn exacerbate the cross-modal shifts. To address these challenges, this paper proposes a new paradigm called Active Cross-Modal Domain Adaptation (ACM-DA), wherein only cross-modal data from the source domain and uni-modal data from the target domain are utilized. To concurrently mitigate the adverse effects of both cross-modal domain shifts and unknown concepts, we propose a Curiosity-Driven Active Adaptation Network (CD-A2N), selectively annotating samples to maximize performance gain. First, we present Curiosity Arousal within Cross-modal Domain Adaptation (CA-CDA) to explore the complexity and novelty characteristics of target samples, while reducing cross-modal discrepancy and aligning source and target domains. Second, Curiosity-driven Active Learning (CAL) is devised to strategically select a subset of target samples for annotation, aiming to achieve more valuable data selection at a small labeling cost. Finally, we jointly train CA-CDA and CAL with the newly labeled target domain sub-dataset to alleviate the above issues. Extensive experiments demonstrate that CD-A2N provides an effective solution for achieving ACM-DA. Code will be available athttps://github.com/Feliciaxyao/ACM-DA. Xuan Yao 0001, Junyu Gao 0002, Zhaoquan Yuan, Xiao Wu 0001, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2024 | R-EDL: Relaxing Nonessential Settings of Evidential Deep LearningabstractA newly-arising uncertainty estimation method named Evidential Deep Learning (EDL), which can obtain reliable predictive uncertainty in a single forward pass, has garnered increasing interest. Guided by the subjective logic theory, EDL obtains Dirichlet concentration parameters from deep neural networks, thus constructing a Dirichlet probability density function (PDF) to model the distribution of class probabilities. Despite its great success, we argue that EDL keeps nonessential settings in both stages of model construction and optimization.
In this work, our analysis indicates that (1) in the construction of the Dirichlet PDF, a commonly ignored parameter termed prior weight governs the balance between leveraging the proportion of evidence and its magnitude in deriving predictive scores, and (2) in model optimization, a variance-minimized regularization term adopted by traditional EDL encourages the Dirichlet PDF to approach a Dirac delta function, potentially exacerbating overconfidence. Therefore, we propose the R-EDL (Relaxed-EDL) method by relaxing these nonessential settings. Specifically, R-EDL treats the prior weight as an adjustable hyper-parameter instead of a fixed scalar, and directly optimizes the expectation of the Dirichlet PDF provided to deprecate the variance-minimized regularization term. Extensive experiments and SOTA performances demonstrate the effectiveness of our method. Source codes are provided in Appendix E. Junyu Gao 0002, Changsheng Xu |
ICLR | 2 |
| 2024 | Fast-Slow Test-Time Adaptation for Online Vision-and-Language NavigationabstractThe ability to accurately comprehend natural language instructions and navigate to the target location is essential for an embodied agent. Such agents are typically required to execute user instructions in an online manner, leading us to explore the use of unlabeled test samples for effective online model adaptation. However, for online Vision-and-Language Navigation (VLN), due to the intrinsic nature of inter-sample online instruction execution and intra-sample multi-step action decision, frequent updates can result in drastic changes in model parameters, while occasional updates can make the model ill-equipped to handle dynamically changing environments. Therefore, we propose a Fast-Slow Test-Time Adaptation (FSTTA) approach for online VLN by performing joint decomposition-accumulation analysis for both gradients and parameters in a unified framework. Extensive experiments show that our method obtains impressive performance gains on four popular benchmarks. Code is available at https://github.com/Feliciaxyao/ICML2024-FSTTA. Junyu Gao 0002, Xuan Yao 0001, Changsheng Xu |
ICML | 1 |
| 2024 | Open-Vocabulary Video Scene Graph Generation via Union-aware Semantic AlignmentabstractVideo Scene Graph Generation (VidSGG) plays a crucial role in various visual-language tasks by providing accessible structured visual relation knowledge. However, the requirement of annotating all categories of prevailing VidSGG methods limits their application in real-world scenarios. Despite the popular VLMs facilitating preliminary exploration of open-vocabulary VidSGG tasks, the correspondence between visual union regions and relation predicates is usually ignored. Therefore, we propose an Open-vocabulary VidSGG framework named Union-Aware Semantic Alignment Network (UASAN) to explore the alignment between visual union regions and relation predicate concepts in the same semantic space. Specifically, a visual refiner is designed to acquire open-vocabulary knowledge and the ability to bridge different modalities. To achieve better alignment, we first design a semantic-aware context encoder to achieve a comprehensive semantic interaction between object trajectories, visual union regions, and trajectory motion information to obtain semantic-aware union region representations. Then, a union-relation alignment decoder is utilized to generate the discriminative relation token for each union region for final relation prediction. Ziyue Wu, Junyu Gao 0002, Changsheng Xu |
ACM Multimedia | 2 |
| 2024 | Conjugated Semantic Pool Improves OOD Detection with Pre-trained Vision-Language ModelsabstractA straightforward pipeline for zero-shot out-of-distribution (OOD) detection involves selecting potential OOD labels from an extensive semantic pool and then leveraging a pre-trained vision-language model to perform classification on both in-distribution (ID) and OOD labels. In this paper, we theorize that enhancing performance requires expanding the semantic pool, while increasing the expected probability of selected OOD labels being activated by OOD samples, and ensuring low mutual dependence among the activations of these OOD labels. A natural expansion manner is to adopt a larger lexicon; however, the inevitable introduction of numerous synonyms and uncommon words fails to meet the above requirements, indicating that viable expansion manners move beyond merely selecting words from a lexicon. Since OOD detection aims to correctly classify input images into ID/OOD class groups, we can "make up" OOD label candidates which are not standard class names but beneficial for the process. Observing that the original semantic pool is comprised of unmodified specific class names, we correspondingly construct a conjugated semantic pool (CSP) consisting of modified superclass names, each serving as a cluster center for samples sharing similar properties across different categories. Consistent with our established theory, expanding OOD label candidates with the CSP satisfies the requirements and outperforms existing works by 7.89% in FPR95. Codes are available in https://github.com/MengyuanChen21/NeurIPS2024-CSP. Junyu Gao 0002, Changsheng Xu |
NeurIPS | 2 |
| 2024 | Multimodal Imbalance-Aware Gradient Modulation for Weakly-Supervised Audio-Visual Video ParsingabstractWeakly-supervised audio-visual video parsing (WS-AVVP) aims to localize the temporal extents of audio, visual and audio-visual event instances as well as identify the corresponding event categories with only video-level category labels for training. Most previous efforts have been devoted to refining the supervision for each modality or extracting fruitful cross-modality information for more reliable feature learning. None of them have noticed the imbalanced feature learning between different modalities in the task. In this paper, to balance the feature learning processes of different modalities, a dynamic gradient modulation (DGM) mechanism is explored, where a novel and effective metric function is designed to measure the imbalanced feature learning between audio and visual modalities. Furthermore, by going in depth into the principle of traditional WS-AVVP pipelines, two additional challenges are identified: confusing multimodal calculation will hamper the precise measurement of audio-visual imbalanced feature learning, as well as the global supervision provided by video-level labels can not provide explicit guidance for robust semantic feature learning in each action subspace. To cope with the above issues, the modality-separated decision unit (MSDU) and semantic-aware feature extractor (SAFE) are designed for precise measurement of imbalanced feature learning and unambiguous semantic-aware feature extraction separately. Comprehensive experiments are conducted on public benchmarks and the corresponding experimental results demonstrate the effectiveness of our proposed method. Jie Fu 0004, Junyu Gao 0002, Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning Proposal-Aware Re-Ranking for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) aims to localize and classify action instances in untrimmed videos with only video-level labels available. Despite the remarkable success of existing methods, whose generated proposals are commonly far more than the ground-truth action instances, it still makes sense to improve the ranking accuracy of the generated proposals since users in real-world scenarios usually prioritize the action proposals with the highest confidence scores. The inaccuracy of the proposal ranking mainly comes from two aspects: For one thing, the traditional proposal generation manner entirely relies on snippet-level perception, resulting in a significant yet unnoticed gap with the target of proposal-level localization. For another, existing methods commonly employ a hand-crafted proposal generation manner, a post-process that does not participate in model optimization. To address the above issues, we propose an end-to-end trained two-stage method, termed as Learning Proposal-aware Re-ranking (LPR) for WTAL. In the first stage, we design a proposal-aware feature learning module to inject the proposal-aware contextual information into each snippet, and then the enhanced features are utilized for predicting initial proposals. Furthermore, to perform effective and efficient proposal re-ranking, in the second stage, we contrast the proposals attached with high confidence scores with our constructed multi-scale foreground/background prototypes for further optimization. Evaluated by both the vanilla and Top-$k$mAP metrics, results of extensive experiments on two popular benchmarks demonstrate the effectiveness of our proposed method. Yufan Hu, Jie Fu 0004, Junyu Gao 0002, Jianfeng Dong, Bin Fan 0001, Hongmin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Learning Transferable Conceptual Prototypes for Interpretable Unsupervised Domain AdaptationabstractDespite the great progress of unsupervised domain adaptation (UDA) with the deep neural networks, current UDA models are opaque and cannot provide promising explanations, limiting their applications in the scenarios that require safe and controllable model decisions. At present, a surge of work focuses on designing deep interpretable methods with adequate data annotations and only a few methods consider the distributional shift problem. Most existing interpretable UDA methods are post-hoc ones, which cannot facilitate the model learning process for performance enhancement. In this paper, we propose an inherently interpretable method, named Transferable Conceptual Prototype Learning (TCPL), which could simultaneously interpret and improve the processes of knowledge transfer and decision-making in UDA. To achieve this goal, we design a hierarchically prototypical module that transfers categorical basic concepts from the source domain to the target domain and learns domain-shared prototypes for explaining the underlying reasoning process. With the learned transferable prototypes, a self-predictive consistent pseudo-label strategy that fuses confidence, predictions, and prototype information, is designed for selecting suitable target samples for pseudo annotations and gradually narrowing down the domain gap. Comprehensive experiments show that the proposed method can not only provide effective and intuitive explanations but also outperform previous state-of-the-arts. Code is available at https://drive.google.com/file/d/1b1EHFghiF1ExD-Cn1HYg75VutfkXWp60/view?usp=sharing. Junyu Gao 0002, Xinhong Ma, Changsheng Xu |
IEEE Trans. Image Process. | 1 |
| 2024 | Spatiotemporal Orthogonal Projection Capsule Network for Incremental Few-Shot Action RecognitionabstractIn this paper, we propose a new task named incremental few-shot action recognition (IFSAR), which aims to learn new action classes incrementally with limited samples. Existing few-shot class incremental learning methods are mainly designed for image datasets and cannot be directly applied to action recognition due to the complicated temporal evolution and spatial structure in videos. Besides, because of the incremental and fewshot setting, the catastrophic forgetting and overfitting problems are further intensified in the video domain. To address the above issues, we propose a spatiotemporal orthogonal projection capsule network (STOP), which employs a spatiotemporal attention routing mechanism and an orthogonal projection capsule layer for effective IFSAR. The former can effectively encode spatial and temporal transformation information and explore the action partwhole relationships to prevent catastrophic forgetting, while the latter is further designed to maintain a sufficient distance between the prototypes of old and novel classes to avoid overfitting by considering spatial-temporal features. Extensive experimental results demonstrate that the proposed method outperforms a series of state-of-the-art approaches on UCF-101, Kinetics-100, and HMDB-51 datasets. Yangbo Feng, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2024 | Exploring Rich Semantics for Open-Set Action RecognitionabstractOpen-set action recognition (OSAR) aims to learn a recognition framework capable of both classifying known classes and identifying unknown actions in open-set scenarios. Existing OSAR methods typically reside in a data-driven paradigm, which ignore the rich semantics in both known and unknown categories. In fact, we humans have the capability of leveraging the captured semantic information, i.e., knowledge and experience, to incisively distinguish samples from known and unknown classes. Motivated by this observation, in this paper, we propose a Unified Semantic Exploration (USE) framework for recognizing actions in open-set scenarios. Specifically, we explore the explicit knowledge semantics by simulating the unknown classes with knowledge-guided virtual classes based on an external knowledge graph, which enables the model to simulate open-set perception during model training. Besides, we propose to learn the implicit data semantics by transferring the knowledge structure of action categories to the visual prototype space for semantic structure preservation. Extensive experiments on several action recognition benchmarks validate the effectiveness of our proposed method. Yufan Hu, Junyu Gao 0002, Jianfeng Dong, Bin Fan 0001, Hongmin Liu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Learning Multi-Expert Distribution Calibration for Long-Tailed Video ClassificationabstractMost existing state-of-the-art video classification methods assume that the training data obey a uniform distribution. However, video data in the real world typically exhibit an imbalanced long-tailed class distribution, resulting in a model bias towards head class and relatively low performance on tail class. While the current long-tailed classification methods usually focus on image classification, adapting them to video data is not a trivial extension. We propose an end-to-end multi-expert distribution calibration method to address these challenges based on two-level distribution information. The method jointly considers the distribution of samples in each class (intra-class distribution) and the overall distribution of diverse data (inter-class distribution) to solve the issue of imbalanced data under long-tailed distribution. By modeling the two-level distribution information, the model can jointly consider the head classes and the tail classes and significantly transfer the knowledge from the head classes to improve the performance of the tail classes. Extensive experiments verify that our method achieves state-of-the-art performance on the long-tailed video classification task. Yufan Hu, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2024 | Feature Disentanglement Network: Multi-Object Tracking Needs More Differentiated FeaturesabstractTo reduce computational redundancies, a common approach is to integrate detection and re-identification (Re-ID) into a single network in multi-object tracking (MOT), referred to as “tracking by detection.” Most of the previous research has focused on resolving the conflict between the detection and Re-ID branches, considering it a simple coupling. In our work, we uncover that the entangled state between the detection and Re-ID tasks is much more complex than previous idea, resulting in a form of competition that degrades performance. To address the preceding issue, we propose a feature disentanglement network that deeply disentangles the intricately interwoven latent space of features and provides differentiated feature maps for each individual task. Furthermore, considering the demand for shallow semantic features in the feature re-ID branch, we also introduce a feature re-globalization module to enrich the shallow semantics. By integrating two distinct networks into a one-shot online MOT method, we develop a robust MOT tracker (named HDGTrack ). We conduct extensive experiments on a number of benchmarks, and our experimental results demonstrate that our method significantly outperforms state-of-the-art MOT methods. Besides, HDGTrack is efficient and can run at 13.9 (MOT17) and 8.7 (MOT20) frames per second. Wen Guo 0003, Wuzhou Quan, Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Learning Commonsense-aware Moment-Text Alignment for Fast Video Temporal GroundingabstractGrounding temporal video segments described in natural language queries effectively and efficiently is a crucial capability needed in vision-and-language fields. In this article, we deal with the fast video temporal grounding (FVTG) task, aiming at localizing the target segment with high speed and favorable accuracy. Most existing approaches adopt elaborately designed cross-modal interaction modules to improve the grounding performance, which suffer from the test-time bottleneck. Although several common space-based methods enjoy the high-speed merit during inference, they can hardly capture the comprehensive and explicit relations between visual and textual modalities. In this article, to tackle the dilemma of the speed–accuracy tradeoff, we propose a commonsense-aware cross-modal alignment network (C 2 AN) that incorporates commonsense-guided visual and text representations into a complementary common space for fast video temporal grounding. Specifically, the commonsense concepts are explored and exploited by extracting the structural semantic information from a language corpus. Then, a commonsense-aware interaction module is designed to obtain bridged visual and text features by utilizing the learned commonsense concepts. Finally, to maintain the original semantic information of textual queries, a cross-modal complementary common space is optimized to obtain matching scores for performing FVTG. Extensive results on two challenging benchmarks show that our C 2 AN method performs favorably against states of the art while running at high speed. Our code is available at https://github.com/ZiyueWu59/CCA Ziyue Wu, Junyu Gao 0002, Shucheng Huang, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Cascade Evidential Learning for Open-world Weakly-supervised Temporal Action LocalizationabstractTargeting at recognizing and localizing action instances with only video-level labels during training, Weakly-supervised Temporal Action Localization (WTAL) has achieved significant progress in recent years. However, living in the dynamically changing open world where unknown actions constantly spring up, the closed-set assumption of existing WTAL methods is invalid. Compared with traditional open-set recognition tasks, Open-world WTAL (OW-TAL) is challenging since not only are the annotations of unknown samples unavailable, but also the fine-grained annotations of known action instances can only be inferred ambiguously from the video category labels. To address this problem, we propose a Cascade Evidential Learning framework at an evidence level, which targets at OWTAL for the first time. Our method jointly leverages multi-scale temporal contexts and knowledge-guided prototype information to progressively collect cascade and enhanced evidence for known action, unknown action, and background separation. Extensive experiments conducted on THUMOS-14 and ActivityNet-v1.3 verify the effectiveness of our method. Besides the classification metrics adopted by previous open-set recognition methods, we also evaluate our method on localization metrics which are more reasonable for OWTAL. Junyu Gao 0002, Changsheng Xu |
CVPR | 2 |
| 2023 | Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event PerceptionabstractWith only video-level event labels, this paper targets at the task of weakly-supervised audio-visual event perception (WS-AVEP), which aims to temporally localize and categorize events belonging to each modality. Despite the recent progress, most existing approaches either ignore the unsynchronized property of audio-visual tracks or discount the complementary modality for explicit enhancement. We argue that, for an event residing in one modality, the modality itself should provide ample presence evidence of this event, while the other complementary modality is encouraged to afford the absence evidence as a reference signal. To this end, we propose to collect Cross-Modal Presence-Absence Evidence (CMPAE) in a unified framework. Specifically, by leveraging uni-modal and cross-modal representations, a presence-absence evidence collector (PAEC) is designed under Subjective Logic theory. To learn the evidence in a reliable range, we propose a joint-modal mutual learning (IML) process, which calibrates the evidence of diverse audible, visible, and audi-visible events adaptively and dynamically. Extensive experiments show that our method surpasses state-of-the-arts (e.g., absolute gains of 3.6% and 6.1 % in terms of event-level visual and audio metrics). Code is available in github.com/MengyuanChen21/CVPR2023-CMPAE. Junyu Gao 0002, Changsheng Xu |
CVPR | 1 |
| 2023 | Leveraging Attribute Knowledge for Open-set Action RecognitionabstractOpen-set action recognition(OSAR) aims to recognize known classes and reject unknown classes. Most OSAR methods focus on learning a favorable threshold to distinguish known and unknown samples in a pure data-driven manner. However, these methods do not utilize the prior knowledge of action classes. In this paper, we propose to Leverage Attribute Knowledge (LAK) for OSAR. Specifically, the class-attribute knowledge learning is designed to integrate attribute knowledge into the model based on spatial-temporal features. Here, attributes are used as a bridge, linking known and unknown classes implicitly to make up the knowledge gap. Furthermore, a learnable relation matrix is adaptively adjusted during training to obtain the class-attribute relations that are expected to be generalized in open-set settings. Extensive experiments on three popular datasets show that the proposed method achieves state-of-the-art performance. Junyu Gao 0002, Yangbo Feng, Changsheng Xu |
ICME | 2 |
| 2023 | Lite-MKD: A Multi-modal Knowledge Distillation Framework for Lightweight Few-shot Action RecognitionabstractExisting few-shot action recognition methods have placed primary focus on improving the recognition accuracy while neglecting another important indicator in practical scenarios, i.e., model efficiency. In this paper, we make the first attempt and propose a Lightweight Multi-modal Knowledge Distillation framework (Lite-MKD) for few-shot action recognition. In this framework, the teacher model conducts multi-modal learning to achieve a comprehensive fusion of the optical flow, depth, and appearance features of human movements, thus achieving a more robust representation of actions. The student model is utilized to learn to recognize actions from the single RGB modality at a lower computational cost under the guidance of the teacher. To fully explore and integrate multi-modal information, a hierarchical Multi-modal Fusion Module (MFM) is introduced in the teacher model. Besides, a multi-level Distinguish-to-Mimic (D2M) knowledge distillation component is proposed for the student model. D2M improves the ability of the student model to mimic the action classification probabilities of the teacher model by enhancing the distinguishability of the student model for different video categories in the support set. Extensive experiments on three action recognition datasets Kinetics, HMDB51, and UCF101 demonstrate our framework's effectiveness and stable generalization ability. With a much more lightweight network for inference, we achieve comparable performance to previous state-of-the-art methods. Our source code is available at https://github.com/HuiGuanLab/Lite-MKD Daizong Liu, Xiaoye Qu, Junyu Gao 0002, Jianfeng Dong, Xun Wang 0007 |
ACM Multimedia | 6 |
| 2023 | Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action UnderstandingabstractUnsupervised pre-training has shown great success in skeleton-based action understanding recently. Existing works typically train separate modality-specific models (i.e., joint, bone, and motion), then integrate the multi-modal information for action understanding by a late-fusion strategy. Although these approaches have achieved significant performance, they suffer from the complex yet redundant multi-stream model designs, each of which is also limited to the fixed input skeleton modality. To alleviate these issues, in this paper, we propose a Unified Multimodal Unsupervised Representation Learning framework, called UmURL, which exploits an efficient early-fusion strategy to jointly encode the multi-modal features in a single-stream manner. Specifically, instead of designing separate modality-specific optimization processes for uni-modal unsupervised learning, we feed different modality inputs into the same stream with an early-fusion strategy to learn their multi-modal features for reducing model complexity. To ensure that the fused multi-modal features do not exhibit modality bias, i.e., being dominated by a certain modality input, we further propose both intra- and inter-modal consistency learning to guarantee that the multi-modal features contain the complete semantics of each modal via feature decomposition and distinct alignment. In this manner, our framework is able to learn the unified representations of uni-modal or multi-modal skeleton input, which is flexible to different kinds of modality input for robust action understanding in practical cases. Extensive experiments conducted on three large-scale datasets, i.e., NTU-60, NTU-120, and PKU-MMD II, demonstrate that UmURL is highly efficient, possessing the approximate complexity with the uni-modal methods, while achieving new state-of-the-art performance across various downstream task scenarios in skeleton-based action representation learning. Our source code is available at https://github.com/HuiGuanLab/UmURL. Shengkai Sun, Daizong Liu, Jianfeng Dong, Xiaoye Qu, Junyu Gao 0002, Xun Yang 0001, Xun Wang 0007, Meng Wang 0001 |
ACM Multimedia | 5 |
| 2023 | Weakly-supervised Video Scene Graph Generation via Unbiased Cross-modal LearningabstractVideo Scene Graph Generation (VidSGG), which aims to detect the relations between objects in a continuous spatio-temporal environment, has shown great potential in video understanding. Almost all prevailing VidSGG approaches are in a fully-supervised manner where expensive manual annotations are required. Therefore, we introduce a novel and challenging task named Weakly-supervised Video Scene Graph Generation (WS-VidSGG), in which a model is trained with only unlocalized scene graphs as supervisory information. Due to the imbalanced data distribution and the lack of fine-grained annotations, models learned in this setting is prone to be biased. Therefore, we propose an Unbiased Cross-Modal Learning (UCML) framework to address the WS-VidSGG task. Specifically, a cross-modal alignment module is firstly designed for allocating pseudo labels to unlabeled visual objects. We then extract unbiased knowledge from dataset statistics, and utilize prompt to make our model finely comprehend semantic concepts. The learned features that from the prompts and unbiased knowledge reinforced each other, resulting in discriminative textual representations. In order to better explore the relations between visual entities, we design a knowledge-guided attention graph to capture the cross-modal relations. Finally, the learned textual and visual features are integrated into a unified framework for relation prediction. Extensive ablation studies verify the effectiveness of our framework. Moreover, the comparison with state-of-the-art fully-supervised methods shows that our proposed framework also achieves comparable performance. Code https://github.com/ZiyueWu59/UCML is available. Ziyue Wu, Junyu Gao 0002, Changsheng Xu |
ACM Multimedia | 2 |
| 2023 | Video Entailment via Reaching a Structure-Aware Cross-modal ConsensusabstractThis paper targets at the task of video entailment, which aims to achieve a thorough comprehension and draw inferences on whether a natural language statement entails or contradicts a given multi-modal video. Despite the recent progress, most existing methods focus on designing a vision-language encoder for multi-modal feature extraction in video entailment, which ignore the underlying consensus knowledge between two modalities, hindering the reasoning performance. As human beings, we make sense of the world by synthesizing information from different sense perceptions, which can acquire consensus among multiple modalities to form a more thorough and coherent representation of the surroundings, as well as to perform complicated understanding tasks. In this paper, we attempt to recreate this ability to infer the truthfulness of a given statement in the context of video entailment. To this end, we propose a unified structure-aware cross-modal consensus method to excavate the consensus semantics shared between video and language modalities, thereby incorporating which into video entailment as statement-related clues. Specifically, the consensus information is achieved by filtering away redundant information by utilizing the global information from one modality and the local complementary information from the other one. Moreover, a consensus-guided graph reasoning method is designed to explore inter-modality consistency and emphasize the significant features related to the judged statement, generating the inference results. Extensive experiments on two benchmarks demonstrate the accurate and robust performance of our approach compared to state-of-the-arts. Code is available at https://github.com/Feliciaxyao/MM2023-SACCN. Xuan Yao 0001, Junyu Gao 0002, Changsheng Xu |
ACM Multimedia | 2 |
| 2023 | Uncertainty-Aware Dual-Evidential Learning for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) aims to localize the action instances and recognize their categories with only video-level labels. Despite great progress, existing methods suffer from severe action-background ambiguity, which mainly arises from background noise and neglect of non-salient action snippets. To address this issue, we propose a generalized evidential deep learning (EDL) framework for WTAL, called Uncertainty-aware Dual-Evidential Learning (UDEL), which extends the traditional paradigm of EDL to adapt to the weakly-supervised multi-label classification goal with the guidance of epistemic and aleatoric uncertainties, of which the former comes from models lacking knowledge, while the latter comes from the inherent properties of samples themselves. Specifically, targeting excluding the undesirable background snippets, we fuse the video-level epistemic and aleatoric uncertainties to measure the interference of background noise to video-level prediction. Then, the snippet-level aleatoric uncertainty is further deduced for progressive mutual learning, which gradually focuses on the entire action instances in an "easy-to-hard" manner and encourages the snippet-level epistemic uncertainty to be complementary with the foreground attention scores. Extensive experiments show that UDEL achieves state-of-the-art performance on four public benchmarks. Our code is available in github/mengyuanchen2021/UDEL. Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Semantic and Temporal Contextual Correlation Learning for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WSTAL) aims to automatically identify and localize action instances in untrimmed videos with only video-level labels as supervision. In this task, there exist two challenges: (1) how to accurately discover the action categories in an untrimmed video (what to discover); (2) how to elaborately focus on the integral temporal interval of each action instance (where to focus). Empirically, to discover the action categories, discriminative semantic information should be extracted, while robust temporal contextual information is beneficial for complete action localization. However, most existing WSTAL methods ignore to explicitly and jointly model the semantic and temporal contextual correlation information for the above two challenges. In this article, a Semantic and Temporal Contextual Correlation Learning Network (STCL-Net) with the semantic (SCL) and temporal contextual correlation learning (TCL) modules is proposed, which achieves both accurate action discovery and complete action localization by modeling the semantic and temporal contextual correlation information for each snippet in the inter- and intra-video manners respectively. It is noteworthy that the two proposed modules are both designed in a unified dynamic correlation-embedding paradigm. Extensive experiments are performed on different benchmarks. On all the benchmarks, our proposed method exhibits superior or comparable performance in comparison to the existing state-of-the-art models, especially achieving gains as high as 7.2% in terms of the average mAP on THUMOS-14. In addition, comprehensive ablation studies also verify the effectiveness and robustness of each component in our model. Jie Fu 0004, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Vectorized Evidential Learning for Weakly-Supervised Temporal Action LocalizationabstractWith the explosive growth of videos, weakly-supervised temporal action localization (WS-TAL) task has become a promising research direction in pattern analysis and machine learning. WS-TAL aims to detect and localize action instances with only video-level labels during training. Modern approaches have achieved impressive progress via powerful deep neural networks. However, robust and reliable WS-TAL remains challenging and underexplored due to considerable uncertainty caused by weak supervision, noisy evaluation environment, and unknown categories in the open world. To this end, we propose a new paradigm, named vectorized evidential learning (VEL), to explore local-to-global evidence collection for facilitating model performance. Specifically, a series of learnable meta-action units (MAUs) are automatically constructed, which serve as fundamental elements constituting diverse action categories. Since the same meta-action unit can manifest as distinct action components within different action categories, we leverage MAUs and category representations to dynamically and adaptively learn action components and action-component relations. After performing uncertainty estimation at both category-level and unit-level, the local evidence from action components is accumulated and optimized under the Subject Logic theory. Extensive experiments on the regular, noisy, and open-set settings of three popular benchmarks show that VEL consistently obtains more robust and reliable action localization performance than state-of-the-arts. Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Weakly-Supervised Video Object Grounding via Causal InterventionabstractWe target at the task of weakly-supervised video object grounding (WSVOG), where only video-sentence annotations are available during model learning. It aims to localize objects described in the sentence to visual regions in the video, which is a fundamental capability needed in pattern analysis and machine learning. Despite the recent progress, existing methods all suffer from the severe problem of spurious association, which will harm the grounding performance. In this paper, we start from the definition of WSVOG and pinpoint the spurious association from two aspects: (1) the association itself is not object-relevant but extremely ambiguous due to weak supervision; and (2) the association is unavoidably confounded by the observational bias when taking the statistics-based matching strategy in existing methods. With this in mind, we design a unified causal framework to learn the deconfounded object-relevant association for more accurate and robust video object grounding. Specifically, we learn the object-relevant association by causal intervention from the perspective of video data generation process. To overcome the problems of lacking fine-grained supervision in terms of intervention, we propose a novel spatial-temporal adversarial contrastive learning paradigm. To further remove the accompanying confounding effect within the object-relevant association, we pursue the true causality by conducting causal intervention via backdoor adjustment. Finally, the deconfounded object-relevant association is learned and optimized under a unified causal framework in an end-to-end manner. Extensive experiments on both IID and OOD testing sets of three benchmarks demonstrate its accurate and robust grounding performance against state-of-the-arts. Wei Wang 0354, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Learning Dual-Routing Capsule Graph Neural Network for Few-Shot Video ClassificationabstractFew-shot video classification (video FSL), which learns classifiers for novel concepts, has gained increasing attention in the last few years from only a few samples. The existing methods rarely consider the local-global relation for video feature learning, which would ultimately result in low discriminative ability. Recently, the capsule network (CapsNet) has shown considerable potential in local-global relation learning in the image analysis field. However, CapsNet cannot be directly applied in video FSL since it ignores the interaction between videos and has high computational complexity. In this paper, a dual-routing capsule graph neural network (DR-CapsGNN) is proposed to solve the above issues. The DR-CapsGNN leverages CapsNet and a graph neural network (GNN) to explore local-global relations and to preserve the detailed properties. Specifically, the CapsGNN is used to learn video relations and structural information to generate high-quality hierarchical capsules. Furthermore, a novel dual-routing mechanism is designed to filter low-discriminative capsules from a holistic perspective and achieves high efficiency, which consists of inter-video and intra-video routing. Extensive experimental results demonstrate that our proposed approach performs favorably compared to state-of-the-art methods on two popular benchmarks. Yangbo Feng, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2023 | Spatial-Temporal Exclusive Capsule Network for Open Set Action RecognitionabstractOpen set action recognition (OSAR) is a rising research domain that simultaneously identifies all videos from known classes and rejects videos from unknown classes. Existing methods rarely consider the open set data distribution and the spatial-temporal relations of video subsequence. Recently proposed Capsule Network (CapsNet) has shown robust performance in many fields, especially image recognition. However, the current CapsNet has not been directly applied to the OSAR task since it cannot explicitly consider the data distribution of known and unknown classes along with the spatial-temporal relations for videos. This paper proposes the Spatial-Temporal Exclusive Capsule Network (STE-CapsNet) to solve the problems in the OSAR task. The STE-CapsNet designs the temporal-spatial routing mechanism to jointly capture the spatial-temporal information of the videos. Furthermore, the exclusive capsules are learned with dot product routing mechanism to limit the data distribution of closed set and open set and reduce the open set risk for OSAR. Extensive experimental results demonstrate that our proposed approach performs favorably compared with state-of-the-art methods on three standard datasets, which verifies its effectiveness and generalization ability. Yangbo Feng, Junyu Gao 0002, Shicai Yang, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2023 | Learning Scene-Aware Spatio-Temporal GNNs for Few-Shot Early Action PredictionabstractWe aim to address a new task named few-shot early action prediction (FS-EAP) that learns classifiers for novel actions from only a few partially observed videos. We argue that the task is extremely challenging since the partially observed videos do not contain enough action information in a few-shot environment. To tackle this task, in this paper, we propose a scene-aware spatio-temporal graph neural network (SA-STGNN) by leveraging the fine-grained spatio-temporal interactions in the video scenes. Specifically, we first generate a spatio-temporal graph corresponding to the partially observed video to capture comprehensive spatio-temporal correlations. Then we utilize the spatio-temporal graph as the input of our SA-STGNN and predict the augmented video features corresponding to the complete video. The architecture uses several scene-aware learning blocks, which are a combination of edge fusion graph neural layers and temporal gated convolutional layers to jointly model spatial and temporal dependencies. Finally, we employ an early action predictor to exploit the learned video features for predicting actions in the few-shot setting. Extensive experimental results on two widely adopted video datasets demonstrate the effectiveness of our approach and its superior performance over the state-of-the-art approaches. Yufan Hu, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2023 | Weakly-Supervised Video Object Grounding via Learning Uni-Modal AssociationsabstractGrounding objects described in natural language to visual regions in the video is a crucial capability needed in vision-and-language fields. In this paper, we deal with the weakly-supervised video object grounding (WSVOG) task, where only video-sentence pairs are provided for learning. The essence of this task is to learn the cross-modal associations between words in textual modality and regions in visual modality. Despite the recent progress, we find that most existing methods focus on the association learning for cross-modal samples, while the rich and complementary information within uni-modal samples has not been fully exploited. To this end, we propose to explicitly learn uni-modal associations on both textual and visual sides, so as to fully exploit the useful uni-modal information for accurate video object grounding. Specifically, (1) we learn textual prototypes by considering rich contextual information of the same object in different sentences, and (2) we estimate visual prototypes in an adaptive manner so as to overcome the uncertainties in selecting object-relevant visual regions. Besides, a cross-modal correspondence is learned which not only bridges the visual and textual modalities for WSVOG task, but also tightly cooperates with the uni-modal association learning process. We conduct extensive experiments on three popular datasets, and the favorable results demonstrate the effectiveness of our method. Wei Wang 0354, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2023 | Many Hands Make Light Work: Transferring Knowledge From Auxiliary Tasks for Video-Text RetrievalabstractThe problem of video-text retrieval, which searches videos via natural language descriptions or vice versa, has attracted growing attention due to the explosive scale of videos produced every day. The dominant approaches for this problem follow the pipeline that firstly learns compact feature representations of videos and texts, and then jointly embeds them into a common feature space where matched video-text pairs are close and unmatched pairs are far away. However, most of them neither consider the structural similarities among cross-modal samples in a global view, nor leverage useful information from other relevant retrieval processes. We argue that both information has great potential for video-text retrieval. In this paper, we treat the relevant retrieval processes as auxiliary tasks and we extract useful knowledge from them by exploiting structural similarities via Graph Neural Networks (GNNs). We then progressively transfer the knowledge from auxiliary tasks in a general-to-specific manner to assist the main task of the current retrieval process. Specifically, for the retrieval of the given query, we first construct a sequence of query-graphs whose central queries are chosen from distant to close to the given query. Then we conduct knowledge-guided message passing in each query-graph to exploit regional structural similarities and gather knowledge of different levels from the updated query-graphs with a knowledge-based attention mechanism. Finally, we transfer the extracted useful knowledge from general to specific to assist the current retrieval process. Extensive experimental results show that our model outperforms the state-of-the-arts on four benchmarks. Wei Wang 0354, Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2022 | Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationabstractWe target at the task of weakly-supervised action localization (WSAL), where only video-level action labels are available during model training. Despite the recent progress, existing methods mainly embrace a localization-by-classification paradigm and overlook the fruitful fine-grained temporal distinctions between video sequences, thus suffering from severe ambiguity in classification learning and classification-to-localization adaption. This paper argues that learning by contextually comparing sequence-to-sequence distinctions offers an essential inductive bias in WSAL and helps identify coherent action instances. Specifically, under a differentiable dynamic programming formulation, two complementary contrastive objectives are designed, including Fine-grained Sequence Distance (FSD) contrasting and Longest Common Subsequence (LCS) contrasting, where the first one considers the relations of various action/background proposals by using match, insert, and delete operators and the second one mines the longest common subsequences between two videos. Both contrasting modules can enhance each other and jointly enjoy the merits of discriminative action-background separation and alleviated task gap between classification and localization. Extensive experiments show that our method achieves state-of-the-art performance on two popular benchmarks. Our code is available at https://github.com/MengyuanChen21/CVPR2022-FTCL. Junyu Gao 0002, Changsheng Xu |
CVPR | 1 |
| 2022 | Dual-Evidential Learning for Weakly-supervised Temporal Action Localization
Junyu Gao 0002, Shicai Yang, Changsheng Xu |
ECCV (4) | 2 |
| 2022 | Learning Semantic-Aware Spatial-Temporal Attention for Interpretable Action RecognitionabstractHuman beings can concentrate on the most semantically relevant visual information when performing action recognition, so as to make reasonable and interpretable predictions. However, most existing approaches, which are applied to address visual tasks, neglect to explicitly imitate such ability for improving the performance and reliability of models. In this paper, we propose an interpretable action recognition framework that can not only improve the performance but also enhance the visual interpretability of 3D CNNs. Specifically, we design a semantic-aware attention module to learn correlative spatial-temporal attention for different action categories. To further leverage the rich semantics of features extracted from different layers, we design a hierarchical semantic fusion module with the help of the learned attention. The proposed two modules can enhance and complement each other, meanwhile, the semantic-aware attention module enjoys the plug-and-play merit. We evaluate our method on different benchmarks with comprehensive ablation studies and visualization analysis. Experimental results demonstrate the effectiveness of our method, showing favorable accuracy against state-of-the-arts while enhancing the semantic interpretability (Code will be available at this linkhttps://github.com/PHDJieFu). Jie Fu 0004, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Learning Video Moment Retrieval Without a Single Annotated VideoabstractVideo moment retrieval has progressed significantly over the past few years, aiming to search the moment that is most relevant to a given natural language query. Most existing methods are trained in a fully-supervised or a weakly-supervised manner, which requires a time-consuming and expensive manually labeling process. In this work, we propose an alternative approach to achieving video moment retrieval that requires no textual annotations of videos and instead leverages the existing visual concept detectors and a pre-trained image-sentence embedding space. Specifically, we design a video-conditioned sentence generator to produce a suitable sentence representation by utilizing the mined visual concepts in videos. We then design a GNN-based relation-aware moment localizer to reasonably select a portion of video clips under the guidance of the generated sentence. Finally, the pre-trained image-sentence embedding space is adopted to evaluate the matching scores between the generated sentence and moment representations with the knowledge transferred from the image domain. By maximizing these scores, the sentence generator and moment localizer can enhance and complement each other to achieve the moment retrieval task. Experimental results on the Charades-STA and ActivityNet Captions datasets demonstrate the effectiveness of our proposed method. Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Compact Representation and Reliable Classification Learning for Point-Level Weakly-Supervised Action LocalizationabstractPoint-level weakly-supervised temporal action localization (P-WSTAL) aims to localize temporal extents of action instances and identify the corresponding categories with only a single point label for each action instance for training. Due to the sparse frame-level annotations, most existing models are in the localization-by-classification pipeline. However, there exist two major issues in this pipeline: large intra-action variation due to task gap between classification and localization and noisy classification learning caused by unreliable pseudo training samples. In this paper, we propose a novel framework CRRC-Net, which introduces a co-supervised feature learning module and a probabilistic pseudo label mining module, to simultaneously address the above two issues. Specifically, the co-supervised feature learning module is applied to exploit the complementary information in different modalities for learning more compact feature representations. Furthermore, the probabilistic pseudo label mining module utilizes the feature distances from action prototypes to estimate the likelihood of pseudo samples and rectify their corresponding labels for more reliable classification learning. Comprehensive experiments are conducted on different benchmarks and the experimental results show that our method achieves favorable performance with the state-of-the-art. Jie Fu 0004, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Image Process. | 2 |
| 2022 | The Model May Fit You: User-Generalized Cross-Modal RetrievalabstractIn real-world applications, a cross-model retrieval model trained on multimodal instances without considering differences in data distributions among users, termed as user domain shift, usually cannot generalize well to unknown user domains. In this paper, we define a new task of user-generalized cross-modal retrieval, and propose a novel Meta-Learning Multimodal User Generalization (MLMUG) method to solve it. MLMUG simulates the user domain shift with meta-optimization, which aims to embed multimodal data effectively and generalize the cross-modal retrieval model to any unknown user domains. We design a cross-modal embedding network with a learnable meta covariant attention module to encode transferable knowledge among different user domains. A user-adaptive metaoptimization scheme is proposed to adaptively aggregate gradients and meta-gradients for fast and stable meta-optimization.We build two benchmarks for user-generalized cross-modal retrieval evaluation. Experiments on the proposed benchmarks validate the generalization of our method compared with several stateof-the-art methods. Xinhong Ma, Xiaoshan Yang, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2022 | Learning Hierarchical Video Graph Networks for One-Stop Video DeliveryabstractThe explosive growth of video data has brought great challenges to video retrieval, which aims to find out related videos from a video collection. Most users are usually not interested in all the content of retrieved videos but have a more fine-grained need. In the meantime, most existing methods can only return a ranked list of retrieved videos lacking a proper way to present the video content. In this paper, we introduce a distinctively new task, namely One-Stop Video Delivery (OSVD) aiming to realize a comprehensive retrieval system with the following merits: it not only retrieves the relevant videos but also filters out irrelevant information and presents compact video content to users, given a natural language query and video collection. To solve this task, we propose an end-to-end Hierarchical Video Graph Reasoning framework (HVGR) , which considers relations of different video levels and jointly accomplishes the one-stop delivery task. Specifically, we decompose the video into three levels, namely the video-level, moment-level, and the clip-level in a coarse-to-fine manner, and apply Graph Neural Networks (GNNs) on the hierarchical graph to model the relations. Furthermore, a pairwise ranking loss named Progressively Refined Loss is proposed based on prior knowledge that there is a relative order of the similarity of query-video, query-moment, and query-clip due to the different granularity of matched information. Extensive experimental results on benchmark datasets demonstrate that the proposed method achieves superior performance compared with baseline methods. Yaguang Song, Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Fast Video Moment RetrievalabstractThis paper targets at fast video moment retrieval (fast VMR), aiming to localize the target moment efficiently and accurately as queried by a given natural language sentence. We argue that most existing VMR approaches can be divided into three modules namely video encoder, text encoder, and cross-modal interaction module, where the last module is the test-time computational bottleneck. To tackle this issue, we replace the cross-modal interaction module with a cross-modal common space, in which moment-query alignment is learned and efficient moment search can be performed. For the sake of robustness in the learned space, we propose a fine-grained semantic distillation framework to transfer knowledge from additional semantic structures. Specifically, we build a semantic role tree that decomposes a query sentence into different phrases (subtrees). A hierarchical semantic-guided attention module is designed to perform message propagation across the whole tree and yield discriminative features. Finally, the important and discriminative semantics are transferred to the common space by a matching-score distillation process. Extensive experimental results on three popular VMR benchmarks demonstrate that our proposed method enjoys the merits of high speed and significant performance. Junyu Gao 0002, Changsheng Xu |
ICCV | 1 |
| 2021 | Active Universal Domain AdaptationabstractMost unsupervised domain adaptation methods rely on rich prior knowledge about the source-target label set relationship, and they cannot recognize categories beyond the source classes, which limits their applicability in practical scenarios. This paper proposes a new paradigm for unsupervised domain adaptation, termed as Active Universal Domain Adaptation (AUDA), which removes all label set assumptions and aims for not only recognizing target samples from source classes but also inferring those from target-private classes by using active learning to annotate a small budget of target data. For AUDA, it is challenging to jointly adapt the model to the target domain and select informative target samples for annotations under a large domain gap and significant semantic shift. To address the problems, we propose an Active Universal Adaptation Network (AUAN). Specifically, we first introduce Adversarial and Diverse Curriculum Learning (ADCL), which progressively aligns source and target domains to classify whether target samples are from source classes. Then, we propose a Clustering Non-transferable Gradient Embedding (CNTGE) strategy, which utilizes the clues of transferability, diversity, and uncertainty to annotate target informative sample, making it possible to infer labels for target samples of target-private classes. Finally, we propose to jointly train ADCL and CNTGE with target supervision to promote domain adaptation and target-private class recognition. Extensive experiments demonstrate that the proposed AUDA model equipped with ADCL and CNTGE achieves significant results on four popular benchmarks. Xinhong Ma, Junyu Gao 0002, Changsheng Xu |
ICCV | 2 |
| 2021 | Diving Into The Relations: Leveraging Semantic and Visual Structures For Video Moment RetrievalabstractExisting dominant approaches for video moment retrieval task are to learn semantic correlation between a given query and the video. However, these methods rarely explore the fine-grained semantic structure and comprehensive visual structure, leading to insufficient utilization of textual and visual relations. In this paper, we propose a unified framework for video moment retrieval, which considers to simultaneously encode semantic and visual structures. Specifically, a semantic role tree is built to reveal the fine-grained semantic information by generating hierarchical textual embeddings. Then the semantic structure is adopted to facilitate the visual structure learning with a contextual attention-based proposal interaction module. Finally, we adaptively aggregate and obtain the visual-semantic matching information through a multi-level fusion strategy to select the best matching moment proposal. Extensive experiments on two popular benchmarks (Charades-STA and ActivityNet Captions) show that our proposed method achieves state-of-the-art performance. Codes are available in the Supplementary Material. Ziyue Wu, Junyu Gao 0002, Shucheng Huang, Changsheng Xu |
ICME | 2 |
| 2021 | Weakly-Supervised Video Object Grounding via Stable Context LearningabstractWe investigate the problem of weakly-supervised video object grounding (WSVOG), where only the video-sentence annotations are provided for training. It aims at localizing the queried objects described in the sentence to visual regions in the video. Despite the recent progress, existing approaches have not fully exploited the potential of the description sentences for cross-modal alignment in two aspects: (1) Most of them extract objects from the description sentences and represent them with fixed textual representations. While achieving promising results, they do not make full use of the contextual information in the sentence. (2) A few works have attempted to utilize contextual information to learn object representations, but found a significant decrease in performance due to the unstable training in cross-modal alignment. To address the above issues, in this paper, we propose a Stable Context Learning (SCL) framework for WSVOG which jointly enjoys the merits of stable learning and rich contextual information. Specifically, we design two modules named Context-Aware Object Stabilizer module and Cross-Modal Alignment Knowledge Transfer module, which are cooperated together to inject contextual information to stable object concepts in text modality and transfer contextualized knowledge in cross-modal alignment. Our approach is finally optimized under a frame-level MIL paradigm. Extensive experiments on three popular benchmarks demonstrate its significant effectiveness. Wei Wang 0354, Junyu Gao 0002, Changsheng Xu |
ACM Multimedia | 2 |
| 2021 | Learning to Model Relationships for Zero-Shot Video ClassificationabstractWith the explosive growth of video categories, zero-shot learning (ZSL) in video classification has become a promising research direction in pattern analysis and machine learning. Based on some auxiliary information such as word embeddings and attributes, the key to a robust ZSL method is to transfer the learned knowledge from seen classes to unseen classes, which requires relationship modeling between these concepts (e.g., categories and attributes). However, most existing approaches ignore to model the explicit relationships in an end-to-end manner, resulting in low effectiveness of knowledge transfer. To tackle this problem, we reconsider the video ZSL task as a task-driven message passing process to jointly enjoy several merits including alleviated heterogeneity gap, low domain shift, and robust temporal modeling. Specifically, we propose a prototype-sample GNN (PS-GNN) consisting of a prototype branch and a sample branch to directly and adaptively model all the relationships between category-attribute, category-category, and attribute-attribute. The prototype branch aims to learn robust representations of video categories, which takes as input a set of word-embedding vectors corresponding to the concepts. The sample branch is designed to generate features of a video sample by leveraging its object semantics. With the co-adaption and cooperation between both branches, a unified and robust ZSL framework is achieved. Extensive experiments strongly evidence that PS-GNN obtains favorable performance on five popular video benchmarks consistently. Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Unsupervised Video Summarization via Relation-Aware Assignment LearningabstractWe address the problem of unsupervised video summarization that automatically selects key video clips. Most state-of-the-art approaches suffer from two issues: (1) they model video clips without explicitly exploiting their relations, and (2) they learn soft importance scores over all the video clips to generate the summary representation. However, a meaningful video summary should be inferred by taking the relation-aware context of the original video into consideration, and directly selecting a subset of clips with a hard assignment. In this paper, we propose to exploit clip-clip relations to learn relation-aware hard assignments for selecting key clips in an unsupervised manner. First, we consider the clips as graph nodes to construct an assignment-learning graph. Then, we utilize the magnitude of the node features to generate hard assignments as the summary selection. Finally, we optimize the whole framework via a proposed multi-task loss including a reconstruction constraint, and a contrastive constraint. Extensive experimental results on three popular benchmarks demonstrate the favourable performance of our approach. Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2021 | Learning Dual-Pooling Graph Neural Networks for Few-Shot Video ClassificationabstractWe address the problem of few-shot video classification that learns classifiers for novel concepts from only a few examples. Most current methods ignore to explicitly consider the relations in both intra-video and inter-video domains, thus cannot take full advantage of the structural information in few-shot learning. In this paper, we propose to exploit the comprehensive intra-video and inter-video relations via Graph Neural Networks (GNNs). To improve the discriminative ability for accurately selecting the representative video content and refining video relations, a Dual-Pooling GNN (DPGNN) is constructed, which stacks customized graph pooling layers in a hierarchical fashion. Specifically, to select the most representative frames in a video, we build intra-video graphs and utilize a node pooling module to extract robust video-level features. We construct an inter-video graph by taking the video-level features as nodes. By designing an edge pooling module, the proposed method can adaptively eliminate the negative relations in the inter-video graph. Extensive experimental results show that our method consistently outperforms the state-of-the-art on two benchmarks. Yufan Hu, Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2021 | Learning Coarse-to-Fine Graph Neural Networks for Video-Text RetrievalabstractWe address the problem of video-text retrieval that searches videos via natural language description or vice versa. Most state-of-the-art methods only consider cross-modal learning for two or three data points in isolation, ignoring to get benefit from the structural information of other data points from a global view. In this paper, we propose to exploit the comprehensive relationships among cross-modal samples via Graph Neural Networks (GNN). To improve the discriminative ability for accurately finding the positive sample, a Coarse-to-Fine GNN is constructed, which can progressively optimize the retrieval results via multi-step reasoning. Specifically, we first adopt heuristic edge features to represent relationships. Then we design a scoring module in each layer to rank the edges connected to the query node and drop the edges with lower scores. Finally, to alleviate the class imbalance issue, we propose a random-drop focal loss to optimize the whole framework. Extensive experimental results show that our method consistently outperforms the state-of-the-arts on four benchmarks. Wei Wang 0354, Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2021 | Knowledge-driven Egocentric Multimodal Activity RecognitionabstractRecognizing activities from egocentric multimodal data collected by wearable cameras and sensors, is gaining interest, as multimodal methods always benefit from the complementarity of different modalities. However, since high-dimensional videos contain rich high-level semantic information while low-dimensional sensor signals describe simple motion patterns of the wearer, the large modality gap between the videos and the sensor signals raises a challenge for fusing the raw data. Moreover, the lack of large-scale egocentric multimodal datasets due to the cost of data collection and annotation processes makes another challenge for employing complex deep learning models. To jointly deal with the above two challenges, we propose a knowledge-driven multimodal activity recognition framework that exploits external knowledge to fuse multimodal data and reduce the dependence on large-scale training samples. Specifically, we design a dual-GCLSTM (Graph Convolutional LSTM) and a multi-layer GCN (Graph Convolutional Network) to collectively model the relations among activities and intermediate objects. The dual-GCLSTM is designed to fuse temporal multimodal features with top-down relation-aware guidance. In addition, we apply a co-attention mechanism to adaptively attend to the features of different modalities at different timesteps. The multi-layer GCN aims to learn relation-aware classifiers of activity categories. Experimental results on three publicly available egocentric multimodal datasets show the effectiveness of the proposed model. Yi Huang 0037, Xiaoshan Yang, Junyu Gao 0002, Jitao Sang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Health Status Prediction with Local-Global Heterogeneous Behavior GraphabstractHealth management is getting increasing attention all over the world. However, existing health management mainly relies on hospital examination and treatment, which are complicated and untimely. The emergence of mobile devices provides the possibility to manage people’s health status in a convenient and instant way. Estimation of health status can be achieved with various kinds of data streams continuously collected from wearable sensors. However, these data streams are multi-source and heterogeneous, containing complex temporal structures with local contextual and global temporal aspects, which makes the feature learning and data joint utilization challenging. We propose to model the behavior-related multi-source data streams with a local-global graph, which contains multiple local context sub-graphs to learn short-term local context information with heterogeneous graph neural networks and a global temporal sub-graph to learn long-term dependency with self-attention networks. Then health status is predicted based on the structure-aware representation learned from the local-global behavior graph. We take experiments on the StudentLife dataset, and extensive results demonstrate the effectiveness of our proposed model. Xiaoshan Yang, Junyu Gao 0002, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Find Objects and Focus on Highlights: Mining Object Semantics for Video Highlight Detection via Graph Neural NetworksabstractWith the increasing prevalence of portable computing devices, browsing unedited videos is time-consuming and tedious. Video highlight detection has the potential to significantly ease this situation, which discoveries moments of user's major or special interest in a video. Existing methods suffer from two problems. Firstly, most existing approaches only focus on learning holistic visual representations of videos but ignore object semantics for inferring video highlights. Secondly, current state-of-the-art approaches often adopt the pairwise ranking-based strategy, which cannot enjoy the global information to infer highlights. Therefore, we propose a novel video highlight framework, named VH-GNN, to construct an object-aware graph and model the relationships between objects from a global view. To reduce computational cost, we decompose the whole graph into two types of graphs: a spatial graph to capture the complex interactions of object within each frame, and a temporal graph to obtain object-aware representation of each frame and capture the global information. In addition, we optimize the framework via a proposed multi-stage loss, where the first stage aims to determine the highlight-probability and the second stage leverage the relationships between frames and focus on hard examples from the former stage. Extensive experiments on two standard datasets strongly evidence that VH-GNN obtains significant performance compared with state-of-the-arts. Junyu Gao 0002, Xiaoshan Yang, Yan Li 0068, Changsheng Xu |
AAAI | 2 |
| 2020 | CI-GNN: Building a Category-Instance Graph for Zero-Shot Video ClassificationabstractWith the ever-growing video categories, Zero-Shot Learning (ZSL) in video classification has drawn considerable attention in recent years. To transfer the learned knowledge from seen categories to unseen categories, most existing methods resort to an implicit model that learns a projection between visual features and semantic category-representations. However, such methods ignore the explicit relationships among video instances and categories, which impede the direct information propagation in a Category-Instance graph (CI-graph) consisting of both instances and categories. In fact, exploring the structure of the CI-graph can capture the invariances of the ZSL task with good generality for unseen instances. Inspired by these observations, we propose an end-to-end framework to directly and collectively model the relationships between category-instance, category-category, and instance-instance in the CI-graph. Specifically, to construct node features of this graph, we adopt object semantics as a bridge to generate unified representations for both videos and categories. Motivated by the favorable performance of Graph Neural Networks (GNNs), we design a Category-Instance GNN (CI-GNN) to adaptively model the structure of the CI-graph and propagate information among categories and videos. With the task-driven message passing process, the learned model is able to transfer label information from categories towards unseen videos. Extensive experiments on four video datasets demonstrate the favorable performance of the proposed framework. Junyu Gao 0002, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2019 | I Know the Relationships: Zero-Shot Action Recognition via Two-Stream Graph Convolutional Networks and Knowledge GraphsabstractRecently, with the ever-growing action categories, zero-shot action recognition (ZSAR) has been achieved by automatically mining the underlying concepts (e.g., actions, attributes) in videos. However, most existing methods only exploit the visual cues of these concepts but ignore external knowledge information for modeling explicit relationships between them. In fact, humans have remarkable ability to transfer knowledge learned from familiar classes to recognize unfamiliar classes. To narrow the knowledge gap between existing methods and humans, we propose an end-to-end ZSAR framework based on a structured knowledge graph, which can jointly model the relationships between action-attribute, action-action, and attribute-attribute. To effectively leverage the knowledge graph, we design a novel Two-Stream Graph Convolutional Network (TS-GCN) consisting of a classifier branch and an instance branch. Specifically, the classifier branch takes the semantic-embedding vectors of all the concepts as input, then generates the classifiers for action categories. The instance branch maps the attribute embeddings and scores of each video instance into an attribute-feature space. Finally, the generated classifiers are evaluated on the attribute features of each video, and a classification loss is adopted for optimizing the whole network. In addition, a self-attention module is utilized to model the temporal information of videos. Extensive experimental results on three realistic action benchmarks Olympic Sports, HMDB51 and UCF101 demonstrate the favorable performance of our proposed framework. Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu |
AAAI | 1 |
| 2019 | Graph Convolutional TrackingabstractTracking by siamese networks has achieved favorable performance in recent years. However, most of existing siamese methods do not take full advantage of spatial-temporal target appearance modeling under different contextual situations. In fact, the spatial-temporal information can provide diverse features to enhance the target representation, and the context information is important for online adaption of target localization. To comprehensively leverage the spatial-temporal structure of historical target exemplars and get benefit from the context information, in this work, we present a novel Graph Convolutional Tracking (GCT) method for high-performance visual tracking. Specifically, the GCT jointly incorporates two types of Graph Convolutional Networks (GCNs) into a siamese framework for target appearance modeling. Here, we adopt a spatial-temporal GCN to model the structured representation of historical target exemplars. Furthermore, a context GCN is designed to utilize the context of the current frame to learn adaptive features for target localization. Extensive results on 4 challenging benchmarks show that our GCT method performs favorably against state-of-the-art trackers while running around 50 frames per second. Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu |
CVPR | 1 |
| 2019 | SMART: Joint Sampling and Regression for Visual TrackingabstractMost existing trackers are either sampling-based or regression-based methods. Sampling-based methods estimate the target state by sampling many target candidates. Although these methods achieve significant performance, they often suffer from a high computational burden. Regression-based methods often learn a computationally efficient regression function to directly predict the geometric distortion between frames. However, most of these methods require large-scale external training videos and are still not very impressive in terms of accuracy. To make both types of methods enhance and complement each other, in this paper, we propose a joint sampling and regression scheme for visual tracking, which leverages the region proposal network by a novel design. Specifically, our method can jointly exploit discriminative target proposal generation and structural target regression to predict target location in a simple feedforward propagation. We evaluate the proposed method on five challenging benchmarks, and extensive experimental results demonstrate that our method performs favorably compared with state-of-the-art trackers with respect to both accuracy and speed. Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu |
IEEE Trans. Image Process. | 1 |
| 2018 | Watch, Think and Attend: End-to-End Video Classification via Dynamic Knowledge Evolution ModelingabstractVideo classification has been achieved by automatically mining the underlying concepts (\eg actions, events) in videos, which plays an essential role in intelligent video analysis. However, most existing algorithms only exploit the visual cues of these concepts but ignore external knowledge information for modeling their relationships during the evolution of videos. In fact, humans have remarkable ability to utilize acquired knowledge to reason about the dynamically changing world. To narrow the knowledge gap between existing methods and humans, we propose an end-to-end video classification framework based on a structured knowledge graph, which can model the dynamic knowledge evolution in videos overtime. Here, we map the concepts of videos to the nodes of the knowledge graph. To effectively leverage the knowledge graph, we adopt a graph convLSTM model to not only identify local knowledge structures in each video shot but also model dynamic patterns of knowledge evolution across these shots. Furthermore, a novel knowledge-based attention model is designed by considering the importance of each video shot and relationships between concepts. We show that by using knowledge graphs, our framework is able to improve the performance of various existing methods. Extensive experimental results on two video classification benchmarks UCF101 and Youtube-8M demonstrate the favorable performance of the proposed framework. Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu |
ACM Multimedia | 1 |
| 2018 | P2T: Part-to-Target Tracking via Deep Regression LearningabstractMost existing part based tracking methods are part-to-part trackers, which usually have two separated steps including part matching and target localization. Different from existing methods, in this paper, we propose a novel part-totarget (P2T) tracker in a unified fashion by inferring target location from parts directly. To achieve this goal, we propose a novel deep regression model for part to target regression in an end-to-end framework via Convolutional Neural Networks. The proposed model is able to not only exploit part context information to preserve object spatial layout structure, but also learn part reliability to emphasize part importance for robust part to target regression. We evaluate the proposed tracker on 4 challenging benchmark sequences, and extensive experimental results demonstrate that our method performs favorably against state-of-the-art trackers because of the powerful capacity of the proposed deep regression model. Junyu Gao 0002, Tianzhu Zhang 0001, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Image Process. | 1 |
| 2017 | A Unified Personalized Video Recommendation via Dynamic Recurrent Neural NetworksabstractPersonalized video recommender systems play an essential role in bridging users and videos. However, most existing video recommendation methods assume that user profiles (interests) are static. In fact, the static assumption is inadequate to reflect users' dynamic interests as time goes by, especially in the online video recommendation scenarios with dramatic changes of video contents and frequent drift of users' interests over different topics. To overcome the above issue, we propose a dynamic recurrent neural network to model users' dynamic interests over time in a unified framework for personalized video recommendation. Furthermore, to build a much more comprehensive recommendation system, the proposed model is designed to exploit video semantic embedding, user interest modeling, and user relevance mining jointly to model users' preferences. By considering these three factors, the RNN model becomes an interest network which can capture users' high level interests effectively. Extensive experimental results on both single-network and cross-network video recommendation scenarios demonstrate the superior performance of the proposed model compared with other state-of-the-art algorithms. Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu |
ACM Multimedia | 1 |
| 2017 | Deep Relative TrackingabstractMost existing tracking methods are direct trackers, which directly exploit foreground or/and background information for object appearance modeling and decide whether an image patch is target object or not. As a result, these trackers cannot perform well when target appearance changes heavily and becomes different from its model. To deal with this issue, we propose a novel relative tracker, which can effectively exploit the relative relationship among image patches from both foreground and background for object appearance modeling. Different from direct trackers, the proposed relative tracker is robust to localize target object by use of the best image patch with the highest relative score to target appearance model. To model relative relationship among large-scale image patch pairs, we propose a novel and effective deep relative learning algorithm via Convolutional Neural Network. We test the proposed approach on challenging sequences involving heavy occlusion, drastic illumination changes, and large pose variations. Experimental results show that our method consistently outperforms state-of-the-art trackers due to the powerful capacity of the proposed deep relative model. Junyu Gao 0002, Tianzhu Zhang 0001, Xiaoshan Yang, Changsheng Xu |
IEEE Trans. Image Process. | 1 |