EDBT 2026 Demo / reviewers in the wild / expert
Xun Jiang 0001
dblp:181/7509-1
· DBLP profile ↗
33ranked-venue papers
15as first author
33since 2021 · last 2026
0000-0003-2209-651XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 11 first-author · 28 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hyper-Opinion Vagueness Quantification for Robust Multimodal LearningabstractRobust Multimodal Learning (RML) aims to address the issues of unreliable predictions of multimodal models. Nevertheless, previous RML works often struggle to distinguish between different categories that rely on identical intra-modal cues, making ambiguous predictions. We defined this degree of ``uncertain'' in extracting discriminative features of a multimodal model as vagueness. Neglecting such vagueness, as previous RML works commonly do, will undermine the ability to extract unique semantics of each category in multimodal models, further resulting in worse robustness under disturbances that affect semantic representations. Additionally, this vagueness will lead the parameter updating processes towards unreliable fusion, thus diverting the learning processes of the multimodal model from learning unique features of each category. Based on the above insight, we propose a novel robust multimodal learning approach, termed Hyper-Opinion Quantifying Vagueness (HOQV). Specifically, we first introduce hyper-opinion to capture and quantify the vagueness of multimodal learning in discriminating representations of different categories. Moreover, to mitigate the interference in parameter updating of unreliable representations with high vagueness, we also design the Hyper-Opinion Gradient Modulation to guide the optimization processes. We evaluate our HOQV on six datasets with different disturbances, including noise and adversarial attack, and demonstrate that our proposed method achieves state-of-the-art performance consistently. Disen Hu, Xun Jiang 0001, Xiaofeng Cao 0002, Zheng Wang 0044, Jingkuan Song, Heng Tao Shen, Xing Xu 0001 |
AAAI | 2 |
| 2026 | De-biased Natural Language Egocentric Task Verification via Prototypical Evidence LearningabstractNatural Language-based Egocentric Task Verification (NLETV) aims to verify the alignment between action sequences in egocentric videos and their corresponding textual descriptions. However, existing NLETV approaches are still facing two critical challenges: (1) These methods are designed for simulating environments, ignoring the domain gap between synthetic and realistic data. (2) The matching processes are regarded as a simple binary classification problem, which undermines model reliability due to evaluation bias and uncalibrated decision settings. To address these challenges, we propose a novel method termed Prototypical Evidential Learning (PEL), which can be adapted to existing NLETV approaches and boost the model generalization and mitigate prediction bias. Our method leverages prototypes to guide cross-domain alignment and evidence collection. Specifically, PEL consists of two key components: (1) Prototypical Domain Adaptation module enabling cross-domain feature alignment and intra-domain prototype preservation between synthetic and realistic domains; (2) Matching Evidence Collector module, which quantifies prediction uncertainty on the prototypical representations through evidential deep learning. It enforces the model to collect the vision-text consistency and discrepancy evidence, thus addressing the issues of biased decisions in binary classification. Extensive experiments on two public datasets demonstrate that our PEL method outperforms existing state-of-the-art NLETV methods and shows remarkable generalizability. Xun Jiang 0001, Fumin Shen, Lei Zhu 0002, Jingkuan Song, Heng Tao Shen, Xing Xu 0001 |
AAAI | 2 |
| 2026 | Generalizable Egocentric Task Verification via Cross-Modal Hybrid Hypergraph MatchingabstractEgocentric Task Verification (ETV) aims to determine if the operation flows of procedural tasks in egocentric videos align with the logic of given rules. Early works adopt the video-based verification paradigm that compares a reference video to the testing video, which limits the flexibility of model deployment. Recent researches incorporate reference textual rules instead of videos, describing the operational logic with natural language, but also raises the challenges of cross-modal heterogeneity and hierarchical misalignment between the two modalities. While previous works mainly address the cross-modal heterogeneity between vision and text modalities, they inevitably suffer from two additional key challenges: (1) Existing methods are mostly developed in synthetic domains, yet have not considered the issues of synthetic-to-realistic generalization challenges in real-world applications. (2) The intricate relations between visual content and textual rule involve multiple matching correlations, indicating high-order matching interactions. To address these issues, we proposed the Generalizable Egocentric Task Verification (GETV), and construct a cross-domain ETV benchmark dataset, EgoCross. It features synthetic-to-real cross-domain evaluation, covering both synthetic datasets for training and realistic datasets for testing, across three different types of tasks. Furthermore, we also propose a novel method for this challenge, termed Cross-modal Hybrid Hypergraph Matching (CHHM), which models the logical cross-modal matching in the GETV challenge as a heterogeneous hybrid hypergraph learning process, thus addressing intrinsic multiple matching correlations. Additionally, to tackle the problems of synthetic-to-realistic generalization, we enhance the cross-modal matching process with prototype-based graph representation alignment, which effectively mitigates the cross-domain gap. Extensive experiments on the existing two ETV benchmark datasets, i.e., EgoTV and CSV-NL, and our proposed GETV dataset EgoCross, demonstrate our approach establishes new state-of-the-art performance on both intra-domain and cross-domain challenges. Xun Jiang 0001, Xing Xu 0001, Zheng Wang 0044, Jingkuan Song, Fumin Shen, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Egocentric Online Action Segmentation via Parametric Context Memory LearningabstractTo facilitate smart wearable devices or human-like robotics with real-time first-person perspective perception ability, recent researchers proposed the Egocentric Online Action Segmentation (EOAS) task. It requires models to recognize what is happening in egocentric streaming videos and discriminate the starting and ending times of an activity in a real-time manner. However, compared with offline-recorded exocentric videos, egocentric streaming videos cannot provide equivalent sufficient temporal-spatial cues due to the limited perspective and unknown coming frames. Hence, it raises a high demand for the long-term episodic memory ability of models. To this end, most previous approaches work on compressing long-term memory into feature representations. In this paper, we propose a novel EOAS paradigm, termed Parametric Context Memory Learning (PCML), which integrates episodic memory into learnable parameters and keeps dynamic updates according to real-time frames. Concretely, we design the Parametric Context Perception layer and construct a novel Episodic Semantic Memorization Network (ESMN) based on it, which integrates episodic memory into learnable parameters and keeps dynamic updates with real-time frames. We evaluate our proposed method on three public egocentric streaming video benchmarks including EgoPER, EgoProceL, and GTEA. Extensive experiments demonstrate the ESMN model significantly outperforms recent state-of-the-art methods. Our code is available at https://github.com/XunCHN/PCML. Xun Jiang 0001, Xing Xu 0001, Zheng Wang 0044, Jingkuan Song, Zhe Sun 0009, Andrzej Cichocki, Heng Tao Shen |
IEEE Trans. Image Process. | 1 |
| 2025 | PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric VideosabstractNatural Language-based Egocentric Task Verification (NLETV) aims to equip agents to determine if operation flows of procedural tasks in egocentric videos align with natural language instructions. Describing rules with natural language provides generalizable applications, but also raises cross-modal heterogeneity and hierarchical misalignment challenges. In this paper, we proposed a novel approach termed Procedural Heterogeneous Graph Completion (PHGC), which addresses these challenges with heterogeneous graphs representing the logic in rules and operation flows. Specifically, our PHGC method mainly consists of three key components: (1) Heterogeneous Graph Construction module that defines objective states and operation flows as vertices, with temporal and sequential relations as edges. (2) Cross-Modal Path Finding module that aligns semantic relations between hierarchical video and text elements. (3) Discriminative Entity Representation module excavates hidden entities that integrate general logical relations and discriminative cues to reveal final verification results. Additionally, we further constructed a new dataset called CSV-NL comprised of realistic videos. Extensive experiments on the two benchmark datasets covering both digital and physical scenarios, i.e., EgoTV and CSV-NL, demonstrate that our proposed PHGC establishes state-of-the-art performance across different settings. Our code and dataset are available at https://github.com/XunCHN/PHGC. Xun Jiang 0001, Xing Xu 0001, Jingkuan Song, Fumin Shen, Heng Tao Shen |
CVPR | 1 |
| 2025 | So Far Yet So Near: Time Series Data Augmentation with Exploring non-Semantic Boundaries based on Reinforcement LearningabstractData augmentation effectively expands feature distribution in time series classification, enhancing downstream task performance. However, existing techniques often fail to maintain semantic consistency between augmented and original time series data, causing label noise and thereby degrading downstream task performance. We argue that data augmentation should preserve time series semantic consistency and expand the non-semantic information space. In this paper, we reformulate data augmentation as a semantic path planning problem between original data and augmented data, modeled as a Markov Decision Process (MDP). We propose a reinforcement learning-based algorithm (RL) named FreqSYN, where the action space is defined by a set of learnable Gaussian kernels that perturbs the frequency domain of the original data to generate augmented samples. The confidence coefficients of augmented data in semantically relevant classification tasks are used as a reward to iteratively refine the FreqSYN. Our method is validated across four datasets, achieving state-of-the-art performance, with a 2% improvement in F1 score over the SimPSI method. The code and models are available at https://github.com/NKU-EmbeddedSystem/FreqSYN. Haoran Li 0014, Jiarong Kang, Xun Jiang 0001, Xiaoli Gong, Jin Zhang 0003, Zhe Sun 0009, Andrzej Cichocki |
ICASSP | 4 |
| 2025 | Egocentric Online Action Segmentation with Behavior-Centred Feature AugmentationabstractThe Egocentric Online Action Segmentation (EOAS) task aims to sequentially segment untrimmed egocentric videos into distinct action segments in a streaming manner. Previous methods primarily focused on improving contextual information utilization, which highly relied on leveraging the prior context. However, under online constraints, the absence of post context limits the effectiveness of the prior context in learning the action semantics. Excessive reliance on prior context may lead to insufficient feature representations of current presented behavior. To tackle this problem, we propose a novel EOAS method, termed Behavior-Centred Feature Augmentation (BCFA), which consists of two key modules: (1) Behavior Prototype Learning models the common sense of each action across different surroundings, enhancing the model’s ability to capture the shared characteristics of behaviors. (2) Presented Behavior Enhancement leverages both the intrinsic characteristics of the current presented behavior itself and the common sense captured by BPL for feature enhancement, mitigating the absence of post contextualization. We evaluate our proposed BCFA method on three public EOAS benchmark datasets, GTEA, EgoProceL, and EgoPER, and demonstrate that our proposed BCFA approach outperforms recent state-of-the-art methods. Zhangye Han, Xun Jiang 0001, Zheng Wang 0044, Xin Liu 0011, Fumin Shen, Xing Xu 0001 |
ICME | 2 |
| 2025 | Probabilistic Embeddings with Causal Constraint for Error Detection in Egocentric Procedural VideosabstractError detection in egocentric procedural task videos aims to identify deviations to support intelligent monitoring and task automation. Despite making significant progress, existing methods that leverage prototypes for egocentric error detection have two drawbacks: (1) The neglect of inherent data traits, i.e., large intra-class variance and minimal inter-class distinction. (2) The absence of causal consistency in temporal modeling. To address these challenges, we introduce a novel framework termed Probabilistic Embeddings with Causal Constraint (PECC) for error detection in egocentric procedural videos. Specifically, we first integrated a causal dilated convolution module in temporal action segmentation model to capture temporal causal consistency. We then train Gaussian Mixture Models (GMMs) for each action class to get frame-level probabilistic embeddings. Finally, We evaluate test frames using log-likelihood values to detect erroneous actions. Extensive experiments conducted on EgoPER and HoloAssist demonstrate that our method achieves state-of-the-art performance, significantly surpassing existing methods in error detection. Our code is available at https://github.com/HouTong-s/PECC-for-Error-Detection-in-Egocentric-Videos. Tong Hou, Shenshen Li, Xun Jiang 0001, Zheng Wang 0044, Fumin Shen, Xing Xu 0001 |
ICME | 3 |
| 2025 | Social Optimum Assisted Gradient Modulation for Imbalanced Multimodal LearningabstractThe imbalanced modality problem in multimodal learning is a vicious phenomenon that leads to the sub-optimization of the modalities due to the gradient conflicts among the modalities and the model’s preference for the easier learning modalities. Recent studies have been dedicated to modulating the gradient from the uni-modal perspective, boosting the learning of the single modality. Nevertheless, they overlook the similarity between multimodal gradient optimization and multi-objective learning, while overemphasizing competition between modalities’ gradients to improve optimization. Therefore, we perceive the imbalanced multimodal optimization as Multi-Objective Optimization, and propose a novel training method: Social Optimum Assisted Gradient Modulation (SOA-GM). In detail, we implemented social Optimum to guide the multimodal model to achieve the trade-off state in which the collaborations of modalities can be maximized. We also proposed envy to reduce the model’s preference to a particular modality. Finally, experiments across multiple datasets indicate our superior and extendable method performance on multimodal learning. Disen Hu, Xun Jiang 0001, Zhe Sun 0009, Hao Yang 0015, Xing Xu 0001 |
ICME | 2 |
| 2025 | Cross-Modal Task Verification via Hypergraph-based Sequential MatchingabstractCross-Modal Task Verification (CMTV) assesses whether a procedural task is executed accurately according to language-based rules, presenting challenges due to its multi-modal and chronological nature. Existing methods using graph or neuro-symbolic approaches face two issues: (1) Conventional methods only model linear sequential relationships among intra-modal nodes, ignoring implicit relationships between non-neighboring nodes. (2) Directed graphs model pairwise cross-modal relationships but overlook cases where a step corresponds to multiple video segments. To address these issues, we propose Hypergraph-based Sequential Matching (HSM) with two components: (1) Temporal Complementary Hypergraph Module (TCHM), a hierarchical sequential hyperedge construction method that focuses on both sequential connections and implicit relationships across nodes. (2) Step-wise Hypergraph Modeling (SHM), a novel hypergraph-based alignment mechanism that better aligns an action description with multiple video segments, improving task verification accuracy. We evaluate HSM on EgoTV and CTV datasets, demonstrating its superiority over state-of-the-art methods. Xun Jiang 0001, Zheng Wang 0044, Fumin Shen, Jingkuan Song, Xing Xu 0001 |
ICME | 2 |
| 2025 | Heterogeneous Graph Embedding for Multimodal Multi-Label Emotion RecognitionabstractMultimodal Multi-label Emotion Recognition (MMER) aims to identify human emotions through various modalities. Previous studies mainly focus on aligning cross-modal data to extract discriminative emotion-dependent features using attention or reconstruction-based strategies, while omitting the fact that the MMER task is also subjected to the multi-label noises that exist in the multi-label classifications, which disturb the modality-to-label correlations. Besides, most of the research also failed to balance the strategy of finding internal label correlations and label dependency of modalities in noisy conditions. In this paper, we proposed a novel Heterogeneous Graph Embedding (HGE) method for the MMER task, which exploits heterogeneous graphs to extract emotional commonality in the modal and temporal levels, and explicitly models cross-modal correlations among heterogeneous modalities. Additionally, it also captures uncertainty brought by multi-label noise and leverages the unevenness of multi-label to overcome potential data issues. Experimental results demonstrate that our HGE method achieves state-of-the-art performance on two widely used multimodal multi-label emotion recognition datasets under both noise-free and noisy circumstances. Disen Hu, Xun Jiang 0001, Zhe Sun 0009, Fumin Shen, Xing Xu 0001 |
ICMR | 2 |
| 2025 | Composed Query-Based Event Retrieval in Video Corpus with Multimodal Episodic PerceptronabstractEvent retrieval involves searching for specific events from untrimmed video galleries and has garnered significant attention in recent years. However, most existing works follow a text-based video retrieval paradigm only, limited by two main drawbacks: (1) The episodic information presented in described events is not fully perceived, leading to declines in retrieval performance facing variable query intentions. (2) Current models are prone to returning false positive results with similar semantics, as simple text queries can hardly accurately describe the target video content users seek. In this paper, we propose a novel event retrieval framework termed Composed Query-Based Event Retrieval (CQBER). Specifically, we first construct two CQBER benchmark datasets, namely ActivityNet-CQ and TVR-CQ, which cover TV shows and open-world scenarios, respectively. Additionally, we propose an initial CQBER method, termed Multimodal Episodic Perceptron (MEP), which excavates complete query semantics from both observed static visual cues and various descriptions. Extensive experiments demonstrate that our proposed framework significantly boosts event retrieval accuracy across different existing methods. Our code and datasets are available at https://github.com/VincentVanNF/CQBER. Fan Ni, Xun Jiang 0001, Hao Yang 0015, Zheng Wang 0044, Fumin Shen, Xing Xu 0001 |
ICMR | 2 |
| 2025 | Multimodal Time Series Alignment for Error Detection in Human Robot InteractionsabstractErrors often occur during human-robot interactions, such as failing to respond, interrupting users, or providing answers that do not meet user expectations. Detecting these issues on time is crucial for making human-robot communication more natural and user-friendly. In this technical report, we present the approach proposed by our team, CFM-HRI, for the ERR@HRI 2.0 Challenge 2025, targeting the task of interaction error detection. Specifically, we propose a lightweight and efficient time-series classification approach, empowered by cross-modal alignment, to detect interaction errors more accurately and promptly. To mitigate temporal misalignment across modalities, we adopt an upsampling alignment strategy, followed by feature fusion to obtain unified representations. A sliding-window voting mechanism is then introduced to construct training samples along with corresponding ground truth labels. Several machine learning models are employed to detect errors based on the fused features. Experimental results demonstrate the effectiveness of our approach in capturing cross-modal inconsistencies and improving detection accuracy. Our approach won first place in Sub-Challenge 2 and second place in Sub-Challenge 1 of the ERR@HRI 2.0 Challenge, held in conjunction with ACM MM 2025. We provide detailed descriptions of our data processing and experimental setup, along with an analysis of the limitations of our approach and potential directions for future work. Our code is available on https://github.com/setsaile/CFM-HRI-ERR-HRI2.0. Xun Jiang 0001, Shuangle Li, Xing Xu 0001 |
ACM Multimedia | 1 |
| 2025 | Geometric Gradient Divergence Modulation for Imbalanced Multimodal LearningabstractMultimodal learning, which has been given great significance recently, may face the challenge of the imbalanced multimodal phenomenon, which leads to the insufficient optimization of both multimodal and unimodal objectives. The core problem lies in the optimization conflicts between the above optimization objectives, resulting in the diverse updating directions and strengths that cause antagonism between them. In this paper, we mathematically analyze the optimization processes of imbalanced multimodal learning in the hyperspaces from a novel geometric perspective. Additionally, based on our theoretical analysis, we defined the volumes of the gradients constructed parallel polyhedron in the hyperspace to quantify the misalignment between the optimization objectives. Subsequently, we proposed the Geometric Gradient Divergence Modulation (GGDM), which leverages the volumes of gradient polyhedron to perform gradient modulation, encouraging alignment among gradients and promoting a synergistic optimization effect. Lastly, we evaluate our GGDM on five widely used multimodal benchmarks, where RGB image, optical flow, text, image, video and audio are involved. Our method achieved state-of-the-art performance compared to other imbalanced multimodal learning methods. Our code is available at: https://github.com/ConstantineWayne/GGDM. Disen Hu, Xun Jiang 0001, Zhe Sun 0009, Hao Yang 0015, Heng Tao Shen, Xing Xu 0001 |
ACM Multimedia | 2 |
| 2025 | Query as Supervision: Toward Low-Cost and Robust Video Moment and Highlight RetrievalabstractVideo Moment and Highlight Retrieval (VMHR) aims at retrieving video events with a text query in a long untrimmed video and selecting the most related video highlights by assigning the worthiness scores. However, we observed existing methods mostly have two unavoidable defects: 1) The temporal annotations of highlight scores are extremely labor-cost and subjective, thus it is very hard and expensive to gather qualified annotated training data. 2) The previous VMHR methods would fit the temporal distributions instead of learning vision-language relevance, which reveals the limitations of the conventional paradigm on model robustness towards biased training data from open-world scenarios. In this paper, we propose a novel method termed Query as Supervision (QaS), which jointly tackles the annotation cost and model robustness in the VMHR task. Specifically, instead of learning from the distributions of temporal annotations, our QaS method completely learns multimodal alignments within semantic space via our proposed Hybrid Ranking Learning scheme for retrieving moments and highlights. In this way, it only requires low-cost annotations and also provides much better robustness towards Out-Of-Distribution test samples. We evaluate our proposed QaS method on three benchmark datasets, i.e., QVHighlights, BLiSS, and Charades-STA and their biased training version. Extensive experiments demonstrate that the QaS outperforms existing state-of-the-art methods under the same low-cost annotation settings and reveals better robustness against biased training data. Our code is available athttps://github.com/CFM-MSG/Code_QaS. Xun Jiang 0001, Liqing Zhu, Xing Xu 0001, Fumin Shen, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Joint Objective and Subjective Fuzziness Denoising for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) aims at teaching computers or robotics to understand human sentiment with diverse multimodal signals, including audio, vision, and text. Current MSA approaches primarily concentrate on devising fusion strategies for multimodal signals and trying to learn better multimodal joint representations. However, employing multimodal signals directly is not appropriate since the human psychological states are fuzzy and can not be categorized easily, which undermines the effectiveness of existing methods. In this paper, we regard the natural fuzziness of human sentiments can be observed as two types: objective fuzziness introduced by human expression and subjective fuzziness caused by the complexity of human affection. Based on the assumption, we proposed a novel method termedJoint Objective and Subjective Fuzziness Denoising (JOSFD), which introduced fuzzy logic into the multimodal fusion process and sentiment decision process to overcome the objective and subjective fuzziness. Specifically, our JOSFD method contains two key modules: (1) Modality-Specific Fuzzification Module leveraging uncertainty estimation and fuzzy logic to overcome the influence of objective fuzziness in different modalities in multimodal fusion. (2) Attitude-Intensity Representation Disentangling that learns joint representations for human attitude and sentiment strength separately and further employs fuzzy logic to decide the sentiment analysis results. We evaluate our proposed JOSFD method on three widely used MSA benchmark datasets, CMU-MOSI, CMU-MOSEI, and CH-SIMS. Extensive experiments demonstrate our proposed JOSFD method outperforms recent state-of-the-art methods. Xun Jiang 0001, Xing Xu 0001, Huimin Lu 0001, Lianghua He, Heng Tao Shen |
IEEE Trans. Fuzzy Syst. | 1 |
| 2025 | Resisting Noise in Pseudo Labels: Audible Video Event Parsing With Evidential LearningabstractPerceiving temporal events and discriminating their modality types in audible videos, which is also called audio-visual video parsing (AVVP), is becoming a research hotspot in multimodal video understanding. The AVVP task generally follows weakly supervised learning settings, since only video-level labels are provided. Most existing works usually generate modalitywise pseudo labels (PLs) first and then learn to parse audio or visual events from the audible videos. However, this paradigm inevitably results in two defects: 1) the generated PLs for each modality are not fully reliable, which may confuse models if they are adopted as supervision signals for discriminating modalities; and 2) the absence of temporal annotations increases the ambiguities in localizing foregrounds in videos, furtherly causing models prone to being disturbed by noisy labels. To tackle these problems, we propose a novel AVVP framework termed noise-resistant event parsing (NREP), which introduces evidential deep learning (EDL) to overcome the limitations of noisy pseudo supervision. Specifically, our NREP framework consists of three key components: 1) modalitywise evidential learning (MEL) that discriminates the modality-class dependency; 2) temporalwise evidential learning (TEL) that explores meaningful foregrounds; and 3) foreground-background consistency learning (FBCL) for collaborating two evidential learning branches above. Through perceiving meaningful video content and learning evidence for modality dependencies, our method suppresses the disturbance of noise in generated PLs thus achieving remarkable performance with different PL generation strategies. We evaluate our NREP method on two AVVP benchmark datasets and demonstrate it consistently to establish new state-of-the-art. Our implementation codes are available at https://github.com/CFM-MSG/NREP. Xun Jiang 0001, Xing Xu 0001, Liqing Zhu, Zhe Sun 0009, Andrzej Cichocki, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Embracing Unimodal Aleatoric Uncertainty for Robust Multimodal FusionabstractAs a fundamental problem in multimodal learning, multimodal fusion aims to compensate for the inherent limitations of a single modality. One challenge of multimodal fusion is that the unimodal data in their unique embedding space mostly contains potential noise, which leads to corrupted cross-modal interactions. However, in this paper, we show that the potential noise in unimodal data could be well quantified and further employed to enhance more stable unimodal embeddings via contrastive learning. Specifically, we propose a novel generic and robust multimodal fusion strategy, termed Embracing Aleatoric Uncertainty (EAU), which is simple and can be applied to kinds of modalities. It consists of two key steps: (1) the Stable Unimodal Feature Augmentation (SUFA) that learns a stable unimodal representation by incorporating the aleatoric uncertainty into self-supervised contrastive learning. (2) Robust Multimodal Feature Integration (RMFI) leveraging an information-theoretic strategy to learn a robust compact joint representation. We evaluate our proposed EAU method on five multimodal datasets, where the video, RGB image, text, audio, and depth image are involved. Extensive experiments demonstrate the EAU method is more noise-resistant than existing multimodal fusion strategies and establishes new state-of-the-art on several benchmarks. Zixian Gao, Xun Jiang 0001, Xing Xu 0001, Fumin Shen, Yujie Li 0001, Heng Tao Shen |
CVPR | 2 |
| 2024 | Temporal Self-Paced Proposal Learning for Weakly-Supervised Video Moment Retrieval and Highlight DetectionabstractThe Weakly-Supervised Moment Retrieval and Highlight Detection (WS-MRHD) task aims at retrieving target moments and highlights in an untrimmed video with a semantic relevant text query. One of the most challenging problems in this task is the absence of reliable temporal supervision signals. In this paper, we propose a Temporal Self-paced Proposal Learning (TSPL) method to perform a progressive temporal proposal selection mechanism. It productively improves the effectiveness of contrastive learning even when the frame-level annotations are inaccessible. Specifically, our proposed TSPL method consists of three key components: (1) The Variance-Based Instance Selection (VBIS) module leverages self-paced learning for dynamic temporal proposal selection. (2) A Highlight Broadcasting (HB) module to combine reliable time spans and assign frame-level pseudo labels. (3) A Negative Sample Learning (NSL) module to align the text query with relevant video segments. By dynamically selecting the appropriate temporal proposals for training, our TSPL method conducts more reliable cross-modal alignment thus remarkably boosting retrieval performance. The extensive experiments on two WS-MRHD public benchmarks verify our proposed TSPL method substantially outperforms current state-of-the-art methods. Liqing Zhu, Xun Jiang 0001, Fumin Shen, Guoqing Wang 0001, Yang Yang 0002, Xing Xu 0001 |
ICME | 2 |
| 2024 | PTAN: Principal Token-aware Adjacent Network for Compositional Temporal GroundingabstractCompositional temporal grounding (CTG) aims to localize the most relevant segment from an untrimmed video based on a given natural language sentence, and the test samples for this task contain novel components not seen in training. However, existing CTG methods suffer from two shortcomings: (1) Most methods adopt transformers to model global video information only, thus failing to balance the long-range perception and regional representation of video sequences; (2) Due to the lack of aligning videos and sentences at a fine-grained level, the model's capacity for compositional generalization is limited, particularly when query sentences contain novel components. To address these problems, we propose a novel method called Principal Token-aware Adjacent Network (PTAN), which consists of three parts: (1) Principal Temporal Token Recomposition combining video clip-level features obtained from the transformer backbone to capture more significant local features while retaining enough contextual information. (2) Regional Semantic-Aware Learning, which exploits regional representations of videos for cross-modal semantic alignment on the feature space. (3) Principal Semantic-Aware Learning that facilitates fine-grained alignment between visual and textual by sensing principal visual and textual tokens in a self-supervised manner. Extensive experiments on two widely used benchmarks (i.e., Charades-CG and ActivityNet-CG) show that our PTAN method outperforms recent CTG state-of-the-art methods, achieving remarkable improvements in compositional generalization. Our code is available at https://github.com/rushzy/PTAN. Zhuoyuan Wei, Xun Jiang 0001, Zheng Wang 0044, Fumin Shen, Xing Xu 0001 |
ICMR | 2 |
| 2024 | Counterfactually Augmented Event Matching for De-biased Temporal Sentence GroundingabstractTemporal Sentence Grounding (TSG), which aims to localize events in untrimmed videos with a given language query, has been widely studied in the last decades. However, recently researchers have demonstrated that previous approaches are severely limited in out-of-distribution generalization, thus proposing the De-biased TSG challenge which requires models to overcome weakness towards outlier test samples. In this paper, we design a novel framework, termed Counterfactually-Augmented Event Matching (CAEM), which incorporates counterfactual data augmentation to learn event-query joint representations to resist the training bias. Specifically, it consists of three components: (1) A Temporal Counterfactual Augmentation module that generates counterfactual video-text pairs by temporally delaying events in the untrimmed video, enhancing the model's capacity for counterfactual thinking. (2) An Event-Query Matching model that is used to learn joint representations and predict corresponding matching scores for each event candidate. (3) A Counterfact-Adaptive Framework (CAF) that incorporates the counterfactual consistency rules on the matching process of the same event-query pairs, furtherly mitigating the bias learned from training sets. We conduct thorough experiments on two widely used DTSG datasets, i.e., Charades-CD and ActivityNet-CD, to evaluate our proposed CAEM method. Extensive experimental results show our proposed CAEM method outperforms recent state-of-the-art methods on all datasets. Our implementation code is available at https://github.com/CFM-MSG/CAEM_Code. Xun Jiang 0001, Zhuoyuan Wei, Shenshen Li, Xing Xu 0001, Jingkuan Song, Heng Tao Shen |
ACM Multimedia | 1 |
| 2024 | Enhanced Experts with Uncertainty-Aware Routing for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis, which has garnered widespread attention in recent years, aims to predict human emotional states using multimodal data. Previous studies have primarily focused on enhancing multimodal fusion and integrating information across different modalities, while overlooking the impact of noisy data on the internal features of each single modality. In this paper, we propose the Enhanced experts with Uncertainty-Aware Routing (EUAR) method to address the influence of noisy data on multimodal sentiment analysis by capturing uncertainty and dynamically altering the network. Specifically, we introduce the Mixture of Experts approach into multimodal sentiment analysis for the first time, leveraging its properties under conditional computation to dynamically alter the network in response to different types of noisy data. Particularly, we refine the experts within the MoE framework to capture uncertainty in the data and extract clearer features. Additionally, a novel routing mechanism is introduced. Through our proposed U-loss, which utilizes the quantified uncertainty by experts, the network learns to route different samples to experts with lower uncertainty for processing, thus obtaining clearer, noise-free features. Experimental results demonstrate that our method achieves state-of-the-art performance on three widely used multimodal sentiment analysis datasets. Moreover, experiments on noisy datasets show that our approach outperforms existing methods in handling noisy data. Zixian Gao, Disen Hu, Xun Jiang 0001, Huimin Lu 0001, Heng Tao Shen, Xing Xu 0001 |
ACM Multimedia | 3 |
| 2024 | Multi-Grained Attention Network With Mutual Exclusion for Composed Query-Based Image RetrievalabstractTheComposed Query-Based Image Retrieval (CQBIR)task aims to precisely obtain the preserved and modified parts, based on the multi-grained semantics learned from the composed query. Since the composed query includes a reference image and the modification text, not just a single modality, this task is more challenging than the general image retrieval tasks. Most previous methods attempt to learn preserved and modified parts via different attention modules and fuse them as a unified representation. However, these methods have two intrinsic drawbacks: 1) The different granular semantic information of the composed query is neglected, which results in the fact that learned preserved and modified parts are irrelevant to correct semantics. 2) The preserved and modified parts learned by previous methods have obvious overlaps, which may lead the model to obtain sub-optimal preserved and modified regions. To this end, we propose a novel method termedMulti-Grained Attention Network with Mutual Exclusion (MANME)to address the above problems. Our MANME method mainly consists of two components: 1) A multi-grained semantic construction for obtaining various textual and visual semantic information. 2) An attention with mutual exclusion constraint for reducing the degree of overlap between preserved and modified parts. It adequately utilizes the various granular semantic information and effectively refines the learned preserved and modified parts. Extensive experiments and further analyses on three widely used CQBIR datasets demonstrate that our proposed MANME method achieves new state-of-the-art performance on the CQBIR task. Shenshen Li, Xing Xu 0001, Xun Jiang 0001, Fumin Shen, Xin Liu 0011, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Fuzzy Multimodal Graph Reasoning for Human-Centric Instructional Video GroundingabstractHuman-centric instructional videos provide opportunities for users to learn real-world multistep tasks, such as cooking, makeup, and using professional tools. However, these lengthy videos always lead to a tedious learning experience, making it challenging for learners to catch specific guidance efficiently. In this article, we present a novel approach, namedfuzzy multimodal graph reasoning (FMGR), to extract target events in long untrimmed human-centric instructional videos using natural language. Specifically, we devise a fuzzy multimodal graph learning layers in our method, which encompass first contextual graph reasoning that transforms the individual features into contextualized features, second cross-modal relation fuzzifier that models the fine-grained matching relationships between two modalities, and third fuzzy graph reasoning that conducts massage passing among cross-modal matching node pairs. Particularly, we integrate fuzzy theory into the cross-modal relation fuzzifier to amplify potential matching pairs, while simultaneously mitigating the interference from ambiguous matches. To validate our method, we conducted evaluations on two human-centric instructional video datasets, i.e., MedVidQA and YouMakeUp. Moreover, we also take further analysis on the impacts of interrogative and declarative queries. Extensive experimental results and further analysis reveal the effectiveness of our proposed FMGR method. Yujie Li 0001, Xun Jiang 0001, Xing Xu 0001, Huimin Lu 0001, Heng Tao Shen |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Zero-Shot Video Moment Retrieval With Angular Reconstructive Text EmbeddingsabstractGiven an untrimmed video and a text query, Video Moment Retrieval (VMR) aims at retrieving a specific moment where the video content is semantically related to the text query. Conventional VMR methods rely on video-text paired data or specific temporal annotations for each target event. However, the subjectivity and time-consuming nature of the labeling process limit their practicality in multimedia applications. To address this issue, recently researchers proposed a Zero-Shot Learning setting for VMR (ZS-VMR) that trains VMR models without manual supervision signals, thereby reducing the data cost. In this paper, we tackle the challenging ZS-VMR problem withAngular Reconstructive Text embeddings (ART), generalizing the image-text matching pre-trained model CLIP to the VMR task. Specifically, assuming that visual embeddings are close to their semantically related text embeddings in angular space, our ART method generates pseudo-text embeddings of video event proposals through the hypersphere of CLIP. Moreover, to address the temporal nature of videos, we also design local multimodal fusion learning to narrow the gaps between image-text matching and video-text matching. Our experimental results on two widely used VMR benchmarks, Charades-STA and ActivityNet-Captions, show that our method outperforms current state-of-the-art ZS-VMR methods. It also achieves competitive performance compared to recent weakly-supervised VMR methods. Xun Jiang 0001, Xing Xu 0001, Zailei Zhou, Yang Yang 0002, Fumin Shen, Heng Tao Shen |
IEEE Trans. Multim. | 1 |
| 2024 | SDN: Semantic Decoupling Network for Temporal Language GroundingabstractTemporal language grounding (TLG) is one of the most challenging cross-modal video understanding tasks, which aims at retrieving the most relevant video segment from an untrimmed video according to a natural language sentence. The existing methods can be separated into two dominant types: 1) proposal-based and 2) proposal-free methods, where the former conduct contextual interactions and the latter localizes timestamps flexibly. However, the constant-scale candidates in proposal-based methods limit the localization precision and bring extra computational costs. In contrast, the proposal-free methods perform well on high-precision metrics-based on the fine-grained features but suffer from a lack of coarse-grained interactions, which cause degeneration when the video becomes complex. In this article, we propose a novel framework termed semantic decoupling network (SDN) that combines the advantages of proposal-based and proposal-free methods and overcomes their defects. It contains three key components: 1) semantic decoupling module (SDM); 2) context modeling block (CMB); and 3) semantic cross-level aggregation module (SCAM). By capturing the video-text contexts in multilevel semantics, the SDM and CMB effectively utilize the benefits of proposal-based methods. Meanwhile, the SCAM maintains the merit of proposal-free methods in that it localizes timestamps precisely. The experiments on three challenge datasets, i.e., Charades-STA, TACoS, and ActivityNet-Caption, show that our proposed SDN method significantly outperforms recent state-of-the-art methods, especially the proposal-free methods. Extensive analyses, as well as the implementation code of the proposed SDN method, are provided at https://github.com/CFM-MSG/Code_SDN. Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Fumin Shen, Zuo Cao, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Cross-Modal Attention Preservation with Self-Contrastive Learning for Composed Query-Based Image RetrievalabstractIn this article, we study the challenging cross-modal image retrieval task,Composed Query-Based Image Retrieval (CQBIR), in which the query is not a single text query but a composed query, i.e., a reference image, and a modification text. Compared with the conventional cross-modal image-text retrieval task, the CQBIR is more challenging as it requires properly preserving and modifying the specific image region according to the multi-level semantic information learned from the multi-modal query. Most recent works focus on extracting preserved and modified information and compositing it into a unified representation. However, we observe that the preserved regions learned by the existing methods contain redundant modified information, inevitably degrading the overall retrieval performance. To this end, we propose a novel method termedCross-ModalAttentionPreservation (CMAP). Specifically, we first leverage the cross-level interaction to fully account for multi-granular semantic information, which aims to supplement the high-level semantics for effective image retrieval. Furthermore, different from conventional contrastive learning, our method introduces self-contrastive learning into learning preserved information, to prevent the model from confusing the attention for the preserved part with the modified part. Extensive experiments on three widely used CQBIR datasets, i.e., FashionIQ, Shoes, and Fashion200k, demonstrate that our proposed CMAP method significantly outperforms the current state-of-the-art methods on all the datasets. The anonymous implementation code of our CMAP method is available at https://github.com/CFM-MSG/Code_CMAP. Shenshen Li, Xing Xu 0001, Xun Jiang 0001, Fumin Shen, Zhe Sun 0009, Andrzej Cichocki |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Progressive Event Alignment Network for Partial Relevant Video RetrievalabstractCurrently, most existing text-based video retrieval methods are only adapted to trimmed videos. However, more complicated untrimmed videos are common in multimedia applications nowadays. In this paper, we focus on the Partially Relevant Video Retrieval (PRVR) task that retrieves untrimmed long videos with partial text descriptions. To tackle this challenging problem, we propose a novel method termed Progressive Event Alignment Network (PEAN) to align text queries with local video content progressively. Specifically, it consists of three key components: (1) A Multimodal Representation Module (MRM) that extracts text representations and hierarchical video representations. (2) An Event Searching Module (ESM) that localizes the described video content roughly. (3) An Event Aligning Module (EAM) that aligns text queries and local video content at a fine-grained level. Additionally, we also design a Gaussian-based pooling strategy in both the ESM and EAM, which thoroughly mines the semantic information in representative video frames. The extensive experiments on three PRVR benchmarks demonstrate our proposed PEAN method significantly outperforms current state-of-the-art methods. Xun Jiang 0001, Xing Xu 0001, Fumin Shen, Zuo Cao |
ICME | 1 |
| 2023 | Joint Searching and Grounding: Multi-Granularity Video Content RetrievalabstractText-based video retrieval is a well-studied task aimed at retrieving relevant videos from a large collection in response to a given text query. Most existing TVR works assume that videos are already trimmed and fully relevant to the query thus ignoring that most videos in real-world scenarios are untrimmed and contain massive irrelevant video content. Moreover, as users' queries are only relevant to video events rather than complete videos, it is also more practical to provide specific video events rather than an untrimmed video list. In this paper, we introduce a challenging but more realistic task called Multi-Granularity Video Content Retrieval (MGVCR), which involves retrieving both video files and specific video content with their temporal locations. This task presents significant challenges since it requires identifying and ranking the partial relevance between long videos and text queries under the lack of temporal alignment supervision between the query and relevant moments. To this end, we propose a novel unified framework, termed, Joint Searching and Grounding (JSG). It consists of two branches: (1) a glance branch that coarsely aligns the query and moment proposals using inter-video contrastive learning, and (2) a gaze branch that finely aligns two modalities using both inter- and intra-video contrastive learning. Based on the glance-to-gaze design, our JSG method learns two separate joint embedding spaces for moments and text queries using a hybrid synergistic contrastive learning strategy. Extensive experiments on three public benchmarks, i.e., Charades-STA, DiDeMo, and ActivityNet-Captions demonstrate the superior performance of our JSG method on both video-level retrieval and event-level retrieval subtasks. Our open-source implementation code is available at https://github.com/CFM-MSG/Code_JSG. Xun Jiang 0001, Xing Xu 0001, Zuo Cao, Yijun Mo, Heng Tao Shen |
ACM Multimedia | 2 |
| 2023 | Faster Video Moment Retrieval with Point-Level SupervisionabstractVideo Moment Retrieval (VMR) aims at retrieving the most relevant events from an untrimmed video with natural language queries. Existing VMR methods suffer from two defects: (1) massive expensive temporal annotations are required to obtain satisfying performance; (2) complicated cross-modal interaction modules are deployed, which lead to high computational cost and low efficiency for the retrieval process. To address these issues, we propose a novel method termed Cheaper and Faster Moment Retrieval (CFMR), which balances the retrieval accuracy, efficiency, and annotation cost for VMR. Specifically, our proposed CFMR method learns from point-level supervision where each annotation is a single frame randomly located within the target moment. Such a labeling strategy achieves 6 times cheaper than the conventional annotations of event boundaries. Furthermore, we also design a concept-based multimodal alignment mechanism to bypass the usage of cross-modal interaction modules during the inference process, remarkably improving retrieval efficiency. The experimental results on three widely used VMR benchmarks demonstrate our proposed CFMR method achieves superior comprehensive performance to current state-of-the-art methods. Moreover, it significantly accelerates the retrieval speed with more than 100 times FLOPs compared to existing approaches with point-level supervision. Our open-source implementation is available at https://github.com/CFM-MSG/Code_CFMR. Xun Jiang 0001, Zailei Zhou, Xing Xu 0001, Yang Yang 0002, Guoqing Wang 0001, Heng Tao Shen |
ACM Multimedia | 1 |
| 2022 | Semi-supervised Video Paragraph Grounding with Contrastive EncoderabstractVideo events grounding aims at retrieving the most relevant moments from an untrimmed video in terms of a given natural language query. Most previous works focus on Video Sentence Grounding (VSG), which localizes the moment with a sentence query. Recently, researchers extended this task to Video Paragraph Grounding (VPG) by retrieving multiple events with a paragraph. However, we find the existing VPG methods may not perform well on context modeling and highly rely on video-paragraph annotations. To tackle this problem, we propose a novel VPG method termed Semi-supervised Video-Paragraph TRansformer (SVPTR), which can more effectively exploit contextual information in paragraphs and significantly reduce the dependency on annotated data. Our SVPTR method consists of two key components: (1) a base model VPTR that learns the video-paragraph alignment with contrastive encoders and tackles the lack of sentence-level contextual interactions and (2) a semi-supervised learning framework with multimodal feature perturbations that reduces the requirements of annotated training data. We evaluate our model on three widely-used video grounding datasets, i.e., ActivityNet-Caption, Charades-CD-OOD, and TACoS. The experimental results show that our SVPTR method establishes the new state-of-the-art performance on all datasets. Even under the conditions of fewer annotations, it can also achieve competitive results compared with recent VPG methods. Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Fumin Shen, Zuo Cao, Heng Tao Shen |
CVPR | 1 |
| 2022 | GTLR: Graph-Based Transformer with Language Reconstruction for Video Paragraph GroundingabstractVideo Paragraph Grounding aims at retrieving multiple relevant moments from an untrimmed video with a given natural language paragraph query. However, the complex paragraph query brings more challenges to the multimodal fusion and context modeling, which limited the performance of existing VPG methods. To this end, we propose a novel framework for VPG in this paper, termed Graph-based Transformer with Language Reconstruction (GTLR). It consists of three components: (1) Multimodal Graph Encoder conducting the graph reasoning for video-text fusion. (2) Event-wise Decoder predicting the timestamps based on multiple sentence-level features. (3) Language Reconstructor rebuilding the paragraph queries and making our model explainable. We adopt two benchmarks, i.e., ActivityNet-Caption and Charades-STA, to evaluate our model and conduct comprehensive experiments to analyze the effectiveness of each component. The experimental results show that our GTLR method outperforms recent state-of-the-art methods. Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Fumin Shen, Zuo Cao |
ICME | 1 |
| 2022 | DHHN: Dual Hierarchical Hybrid Network for Weakly-Supervised Audio-Visual Video ParsingabstractThe Weakly-Supervised Audio-Visual Video Parsing (AVVP) task aims to parse a video into temporal segments and predict their event categories in terms of modalities, labeling them as either audible, visible, or both. Since the temporal boundaries and modalities annotations are not provided, only video-level event labels are available, this task is more challenging than conventional video understanding tasks.Most previous works attempt to analyze videos by jointly modeling the audio and video data and then learning information from the segment-level features with fixed lengths. However, such a design exist two defects: 1) The various semantic information hidden in temporal lengths is neglected, which may lead the models to learn incorrect information; 2) Due to the joint context modeling, the unique features of different modalities are not fully explored. In this paper, we propose a novel AVVP framework termedDual Hierarchical Hybrid Network (DHHN) to tackle the above two problems. Our DHHN method consists of three components: 1) A hierarchical context modeling network for extracting different semantics in multiple temporal lengths; 2) A modality-wise guiding network for learning unique information from different modalities; 3) A dual-stream framework generating audio and visual predictions separately. It maintains the best adaptions on different modalities, further boosting the video parsing performance. Extensive quantitative and qualitative experiments demonstrate that our proposed method establishes the new state-of-the-art performance on the AVVP task. Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Jingkuan Song, Fumin Shen, Huimin Lu 0001, Heng Tao Shen |
ACM Multimedia | 1 |