Huaiwen Zhang

dblp:180/5703 · DBLP profile ↗
← Back
43ranked-venue papers
13as first author
40since 2021 · last 2026
0000-0002-3183-9218ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 10 first-author · 28 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Synergistic Effects of Fascia Knife and TCM on Hamstring Fatigue Recovery in Track Athletes: An Electromyographic Evaluation
abstract
Purpose: Sports fatigue is common among track and field athletes. Among them, the recovery of hamstring fatigue is particularly important for athletic performance and injury prevention. However, current rehabilitation methods often have limitations. This study aims to compare the effects of fascia knife therapy, traditional Chinese medicine (TCM), and their combined intervention on hamstring fatigue recovery, and to establish a multidimensional evaluation system through electromyographic signals. Methods: A total of 20 track and field athletes were recruited for this study using a randomized experimental design. A total of 17 eligible subjects ([Formula: see text]) were ultimately enrolled for statistical analysis, all in the experimental group, further subdivided into fascia knife, TCM, and combined intervention subgroups. The intervention lasted for six weeks, with electromyographic signal data collected immediately after training and 24[Formula: see text]h post-intervention. Results: The combined intervention of fascia knife and TCM worked much better than the single interventions in recovering from hamstring fatigue in track and field athletes. Two-way repeated measures ANOVA indicated significant interaction effects between time and group factors on RMS (p-value of 0.022, which is less than 0.05) and integrated electromyographic value (iEMG) (p-value of 0.043, which is less than 0.05), with a significant main effect of the MF time factor (p-value of 0.019, which is less than 0.05); 24[Formula: see text]h post-intervention, the combined intervention group showed significant differences in RMS (p-value of 0.0014, which is less than 0.01) and iEMG (p-value of 0.0002, which is less than 0.01) compared with the fascia knife group, and the TCM group showed significant differences in iEMG compared with the fascia knife group (p-value of 0.033, which is less than 0.05), indicating a synergistic effect of the combined intervention, which can effectively alleviate hamstring fatigue and enhance athletic ability. Conclusions: The combined intervention of fascia knife and TCM shows significant advantages, followed by the intervention of fascia knife alone, and the effect of TCM alone is relatively weak. It is recommended to prioritize the promotion of combined intervention in the track and field athlete group, which can shorten the recovery period and reduce the risk of sports injuries, ensuring the stable performance of athletes in their competitive form.
Xuxia Guo, Hengrui Yu, Huaiwen Zhang, Sichuang Yang
Int. J. Pattern Recognit. Artif. Intell.5
2026 Hierarchical Semantics Interaction for Compressed Video Action Recognition
abstract
Directly recognizing action based on compressed video shows significant advantages, such as low storage demands, efficient decoding, and fast inference speeds. Existing methods on compressed video achieve promising performance by separately modeling spatial and motion cues and directly fusing recognition results of I-frames and P-frames. However, these approaches overlook the following inherent attributes of compressed videos: 1) Temporal misalignment between I-frames and P-frames impairs the accuracy of action recognition. 2) Spatiotemporal sparsity of compressed video frames severely hinders the semantic modeling of complex actions. 3) Semantic discrepancy between I-frames and P-frames, which capture appearance and motion information respectively, leads to suboptimal performance when they are fused directly. To address these challenges, we propose a Hierarchical Semantics Interaction Network (HSINet) that ensures refined semantic modeling of compressed video through alignment, interaction, and calibration within a hierarchical fusion framework. Specifically, to resolve temporal misalignment, we propose an efficient cross-modal temporal alignment module that fully combines the spatiotemporal information of P-frames and the spatial information of I-frames, and includes both pre-alignment and fine alignment stages. To mitigate semantic degradation caused by sparsity, we propose a cross-modal semantics interaction module to provide the multi-scale semantics interaction between spatial and temporal representation learning and enhance representations’ spatiotemporal awareness. To calibrate the semantic imbalance between I-frames and P-frames, we propose a modal imbalance calibration module that optimizing directional differences via cosine similarity in hyperspherical space. Experiments on HMDB-51, UCF-101, and Kinetics-400 benchmarks, demonstrate the effectiveness of hierarchical semantic interaction for compressed video action recognition.
Jinxin Guo, Yang Yang 0121, Huaiwen Zhang, Shengsheng Qian, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.3
2025 Synergizing LLMs with Global Label Propagation for Multimodal Fake News Detection
abstract
Large Language Models (LLMs) can assist multimodal fake news detection by predicting pseudo labels.However, LLM-generated pseudo labels alone demonstrate poor performance compared to traditional detection methods, making their effective integration nontrivial.In this paper, we propose Global Label Propagation Network with LLM-based Pseudo Labeling (GLPN-LLM) for multimodal fake news detection, which integrates LLM capabilities via label propagation techniques.The global label propagation can utilize LLMgenerated pseudo labels, enhancing prediction accuracy by propagating label information among all samples.For label propagation, a mask-based mechanism is designed to prevent label leakage during training by ensuring that training nodes do not propagate their own labels back to themselves.Experimental results on benchmark datasets show that by synergizing LLMs with label propagation, our model achieves superior performance over state-ofthe-art baselines.Our code is available online 1 .
Shuguo Hu, Huaiwen Zhang
ACL (1)3
2025 Cross-domain Rumor Detection via Test-Time Adaptation and Large Language Models
abstract
Rumor detection on social media has become crucial due to the rapid spread of misinformation.Existing approaches primarily focus on within-domain tasks, resulting in suboptimal performance in cross-domain scenarios due to domain shift.To address this limitation, we draw inspiration from the strong generalization capabilities of Test-Time Adaptation (TTA) and propose a novel framework to enhance rumor detection performance across different domains.Specifically, we introduce Test-Time Adaptation for Rumor Detection (T 2 ARD), which incorporates both single-domain model and target graph adaptation strategies tailored to the unique requirements of cross-domain rumor detection.T 2 ARD utilizes a graph adaptation module that updates the graph structure and node attributes through multi-level self-supervised contrastive learning, aiming to derive invariant graph representations.To mitigate the impact of significant distribution shifts on self-supervised signals, T 2 ARD performs model adaptation by using annotations from Large Language Models (LLMs) on target graph to produce pseudo-labels as supervised signals.Experiments conducted on four widely used cross-domain datasets demonstrate that T 2 ARD achieves state-of-the-art performance, surpassing existing methods in rumor detection.
Yuxia Gong, Shuguo Hu, Huaiwen Zhang
EMNLP3
2025 VCSA: Video Copy Localization Via Single Frame Annotation
abstract
Video Copy Localization (VCL) aims to identify all copied segments within untrimmed video pairs. Fully supervised methods, which require annotating the boundaries of copied segments, are labor-intensive and susceptible to distraction from non-copied segments. To address this problem, we propose a more efficient annotation paradigm called "single frame supervision", which only requires two randomly selected timestamps within the copied segments. With single frame supervision, we introduce a method called VCSA. This method combines spatial-temporal perception and fusion modules to extract and fuse features, which are then segmented for contrastive learning. Additionally, we propose grid alignment loss to enhance the ability of model to distinguish copied from non-copied segments. Extensive experiments demonstrate that VCSA outperforms other methods within our paradigm, sometimes even reaching performance levels comparable to fully supervised methods.
Shuguo Hu, Yang Yang 0121, Huaiwen Zhang
ICASSP4
2025 EarlyMix: Hierarchical Mixing for Early Time Series Classification
abstract
Early Time Series Classification (ETSC) aims to predict class labels using only prefixes (partial sequences) of time series, which is crucial for applications requiring prompt decision-making. However, time series prefixes, with their limited data length, pose significant challenges for models to recognize patterns and provide reliable predictions based on incomplete data. Enhancing ETSC by addressing this limitation is essential. Existing ETSC methods primarily focus on improving feature extractors or refining stopping strategies, yet they often fail to overcome the core issue of data insufficiency in early segments. We present EarlyMix, which employs a hierarchical mixing strategy comprising Sample Mixing and Latent Mixing to augment data and enhance feature representation. This method significantly improves model performance by enabling the extraction of more informative and robust features from time series prefixes. Extensive experiments on widely-used benchmark datasets demonstrate the superior performance of EarlyMix over baselines.
Shuguo Hu, Jun Hu 0016, Junwei Lv, Huaiwen Zhang
ICME4
2025 UTSC: Uncertainty-Aware Framework for Early Time Series Classification
abstract
Early Time Series Classification (ETSC) is pivotal in time-sensitive real-world applications. However, existing methods face two significant challenges: (1) the lack of explicit representation of uncertainty, which leads to inaccurate decision-making in the early stages, and (2) the difficulty in handling heteroscedastic uncertainty, where uncertainty dynamically changes as more data is observed, especially in early stages of time series classification. Traditional ETSC methods assume that the predicted probability distributions at each time step are completely reliable and accurate. This assumption neglects the inherent uncertainty in the classification process, leading to cumulative errors that ultimately affect the overall accuracy and reliability of the classification decision. To address these challenges, we propose a novel Uncertainty-aware framework for early Time Series Classification (UTSC), designed to model and adapt to uncertainty during early classification dynamically. UTSC incorporates an Uncertainty Probability Decoder (UPD), which captures the inherent randomness and uncertainty in early-stage data by leveraging hidden-layer information to model variability adaptively. Additionally, UTSC employs an enhanced probability reweighting method to adjust the probability distributions dynamically, enabling the model to account for uncertainty and make more informed decisions. We evaluate UTSC on 45 diverse datasets, comparing its performance against eight baseline models. The results demonstrate that UTSC significantly outperforms existing methods. Our code is publicly available1.
Naixin Yan, Jun Hu 0016, Yuqi Chu, Shuguo Hu, Huaiwen Zhang
IJCNN5
2025 Event Consistency-aware Robust Fake News Detection
abstract
With the rapid development of short video platforms (such as Kuaishou and TikTok), these platforms have increasingly become important channels for the spread of fake news.Therefore, multi-modal fake news detection has attracted extensive attention.Existing studies mainly focus on directly integrating multi-modal information or discovering implicit clues in posts to improve detection performance.However, due to the abuse of video editing techniques, event-irrelevant segments (e.g., advertisements) are frequently mixed into videos, introducing noise information, thereby weakening models' ability to learn crucial information.Moreover, video creators often inject personal tampered information into original news content through audio modality manipulation, potentially distorting the factual. To address these challenges, we propose a novel Event Consistency-aware Robust Fake News Detection (ECR-FND) framework, comprising two key components: an Event-aware Video Denoising Learning (EVDL) and an Audio Tampering-information Capturing Module (ATCM).Specifically, the EVDL filters out the event-irrelevant segments within video modality to focus on core news events. The ATCM adaptively amplifies tampering information in audio modality, enhancing the model's capacity to detect manipulation attempts.Extensive experiments on two benchmark datasets (FakeSV and FakeTT) demonstrate ECR-FND's effectiveness.Our source code is available at https://github.com/immc-lab/ECR-FND.
Zihang Guo, Huaiwen Zhang
ACM Multimedia3
2025 Evaluating and Mitigating Sycophancy in Large Vision-Language Models
abstract
Large vision-language models (LVLMs) have recently achieved significant advancements, demonstrating powerful capabilities in understanding and reasoning about visual information. However, LVLMs may generate biased responses that reflect the user beliefs rather than the facts, a phenomenon known as sycophancy. Sycophancy can pose serious challenges to the performance, trustworthiness, and security of LVLMs, raising concerns about their practical applications. We note that there is limited work on the evaluation and mitigation of sycophancy in LVLMs. In this paper, we introduce SyEval-VL, a benchmark specifically designed to evaluate sycophancy in LVLMs. SyEval-VL offers a comprehensive evaluation of sycophancy in visual understanding and reasoning across various scenarios with a multi-round dialogue format. We evaluate sycophancy in several popular LVLMs, providing an in-depth analysis of various sycophantic behaviors and their consequential impacts. Additionally, we propose a novel framework, Human Feedback-based Retrieval-Augmented Generation (HFRAG), to mitigate sycophancy in LVLMs by determining the appropriate timing of retrieval, profiling the proper retrieval target, and augmenting the decoding of LVLMs. Extensive experiments demonstrate that the proposed method significantly mitigates sycophancy in LVLMs without requiring additional training. Our code is available at: https://github.com/immc-lab/SyEval-VL
Jiayi Gao, Huaiwen Zhang
ACM Multimedia2
2025 Fine-tuning Bias Neurons for Fair Text-to-Image Generation
abstract
Diffusion Models (DMs) have revolutionized Text-to-Image (T2I) generation, yet inherent dataset biases often result in skewed representations across demographics, perpetuating stereotypes and social inequities. Existing debiasing approaches primarily focus on the text processing component, overlooking the intricate biases in the diffusion model's U-Net architecture. This paper presents a novel approach to addressing these biases through a causal analysis of bias disentanglement within the U-Net architecture. We introduce the Contrast Neuron Sensitivity Metric, which enables precise identification of neurons sensitive to bias, allowing for targeted interventions. Our debiasing paradigm fine-tunes these identified neurons with a combination of distribution and semantic loss, requiring only 0.2M parameters to be adjusted, which is far less than prior methods. Experiments show that our method effectively removes gender and race biases and maintains the diversity distribution of images. It enables both absolute fairness and relative adjustments by modifying target attribute distributions (e.g., young:old = 7:3). Furthermore, our approach is scalable, allowing simultaneous fine-tuning across multiple biases, and achieves good bias reduction even with non-templated prompts. The code is available on https://github.com/FanQi-AI/Debias.
Fan Qi, Changsheng Xu, Huaiwen Zhang
ACM Multimedia4
2025 Sequence-Event Semantic Consistent Learning for Text-to-Motion Retrieval
abstract
Text-to-Motion Retrieval (TMR) is a challenging task to retrieve relevant motion sequences with the natural language description. Existing TMR methods primarily utilize single embeddings to represent and align text and motion sequences. However, real-world motion sequences typically contain multiple sequential actions with intricate semantics, which are hard to precisely capture by single embedding. Additionally, relying solely on naive contrastive training to capture high-level semantics may struggle to perceive and capture fine-grained action details necessary for precise text-motion alignment. In this work, we propose a novel Sequence-Event Semantic Consistent Learning (SECL) framework for 3D human motion retrieval. Specifically, we introduce a self-supervised learning strategy to incorporate fine-grained action details into the motion representations via the generative feedback from the diffusion model. We design a parameter-free sequence-level interaction to explore coarse-grained alignment and an event-level interaction that utilizes several learnable queries to capture event semantics in a shared learning manner for fine-grained alignment. Furthermore, an inter-consistency loss is introduced to align the event semantics between the motion and corresponding text, and an intra-diversity loss is designed to encourage event features to attend to different contents, effectively capturing the rich action information. Finally, we modify the traditional contrastive alignment objective and propose an importance-sampling strategy to emphasize harder negatives for discriminative representation learning. Extensive experiments show that our method significantly outperforms existing methods in text-to-motion retrieval and other challenging tasks, e.g., human interaction recognition and motion temporal localization.
Haoyu Shi, Huaiwen Zhang
ACM Multimedia2
2025 One-shot Multimodal Federated Learning via Diverse Synthetic Feature Optimization
abstract
One-shot Federated Learning (FL) enables collaborative model training through a single round of communication, thereby significantly reducing communication overhead. However, in multimodal settings, client data heterogeneity and modality heterogeneity make it difficult to accurately learn the global multimodal representation within a single communication round. We propose Diverse Synthetic Feature Optimization (DSFO) for one-shot MFL framework that achieves global modality consensus via a single communication round. DSFO introduces: (1) Dynamic Probabilistic Scheduling Model Queue (DPSMQ), which ensures local multimodal feature-level distillation quality through diversity-aware model sampling from the queue, directly transmitting synthetic features can significantly reduces the risk of misleading knowledge transfer and (2) Pareto-Optimal Global Feature Consensus (POGFC), a server-side multi-objective optimization that extracts maximally representative synthetic features per modality-category pair to mitigate data heterogeneity, supported by theoretical guarantees in the appendix. Experiments on multimodal datasets demonstrate that DSFO outperforms current methods while significantly reducing communication overhead.
Fan Qi, Zixin Zhang 0004, Huaiwen Zhang
MMAsia4
2025 Controlled Cluster Separation for Class Incremental Learning
Huaiwen Zhang
PRICAI2
2025 L2-GNN: Graph neural networks with fast spectral filters using twice linear parameterization
abstract
To improve learning on irregular 3D shapes, such as meshes with varying discretizations and point clouds with different samplings, we propose L 2 -GNN, a new graph neural network that approximates the spectral filters using twice linear parameterization. First, we parameterize the spectral filters using wavelet filter basis functions. The parameterization allows for an enlarged receptive field of graph convolutions, which can simultaneously capture low-frequency and high-frequency information. Second, we parameterize the wavelet filter basis functions using Chebyshev polynomial basis functions. This parameterization reduces the computational complexity of graph convolutions while maintaining robustness to the change of mesh discretization and point cloud sampling. Our L 2 -GNN based on the fast spectral filter can be used for shape correspondence, classification, and segmentation tasks on non-regular mesh or point cloud data. Experimental results show that our method outperforms the current state of the art in terms of both quality and efficiency.
Siying Huang, Zhengda Lu, Hongxing Qin, Huaiwen Zhang, Yiqun Wang 0001
Graph. Model.5
2025 Enhancing target speaker extraction with Hierarchical Speaker Representation Learning
abstract
Target speaker extraction aims to obtain the speech of the specific speaker from a mixture of multiple voices. The conventional approach exploits the target speaker embeddings from a pre-recorded speech segment as auxiliary information, providing prior for extraction. However, the naive single-vector embedding may lack attention to the subtle acoustic features such as pitch and harmonic distribution in the auxiliary speech, leading to an unsatisfying performance. Furthermore, traditional speaker embeddings are trained by speaker verification system and do not leverage the semantics of the auxiliary speech which may facilitate the extraction. To address these challenges, we propose a simple yet effective Hierarchical Speaker Representation Learning (HSRL). The proposed method comprises three modules: a Local Speaker Feature Extractor (LSFE), a Global Speaker Feature Extractor (GSFE), and a Hierarchical Cascading Input Strategy (HCIS). Specifically, the LSFE utilizes the fine-grained acoustic information in the anchor speech. In GSFE, we utilize ECAPA-TDNN to obtain the speaker embeddings of the target speaker, enhancing extraction performance with this global speaker information. In additional, a novel HCIS is proposed to integrate the output of the LSFE module to the input of the GSFE, which enables the global speaker features to focus on the semantic content of the pre-recorded speech. Experimental results on the Libri-2talker dataset demonstrate that our HSRL has achieved significant performance improvements and established new optimal benchmarks.
Shulin He, Huaiwen Zhang
Neural Networks4
2025 Active Supervised Cross-Modal Retrieval
abstract
Supervised Cross-Modal Retrieval (SCMR) achieves significant performance with the supervision provided by substantial label annotations of multi-modal data. However, the requirement for large annotated multi-modal datasets restricts the use of supervised cross-modal retrieval in many practical scenarios. Active Learning (AL) has been proposed to reduce labeling costs while improving performance in various label-dependent tasks, in which the most informative unlabeled samples are selected for labeling and training. Directly exploiting the existing AL methods for supervised cross-modal retrieval may not be a good idea since they only focus on the uncertainty within each modality, ignoring the inter-modality relationship within the text-image pairs. Furthermore, existing methods focus exclusively on the informativeness of data during sample selection, leading to a biased, homogenized set where selected samples often contain nearly identical semantics and are densely distributed in a region of the feature space. Persistent training with such biased data selections can disturb multi-modal representation learning and substantially degrade the retrieval performance of SCMR. In this work, we propose an Active Supervised Cross-Modal Retrieval (ASCMR) framework, which effectively identifies informative multi-modal samples and generates unbiased sample selections. In particular, we propose a probabilistic multi-modal informativeness estimation that captures both the intra-modality and inter-modality uncertainty of multi-modal pairs within a unified representation. To ensure unbiased sample selection, we introduce a density-aware budget allocation strategy that constrains the active learning objective of maximizing the informativeness of selection with a novel semantic density regularization term. The proposed methods are evaluated on three widely used benchmark datasets, MS-COCO, NUS-WIDE, and MIRFlickr, demonstrating our effectiveness in significantly reducing the annotation cost while outperforming other baselines of active learning strategies. We could achieve over 95% of the fully supervised model's performance by only utilizing 6%, 3%, and 4% active selected samples for MS-COCO, NUS-WIDE, and MIRFlickr, respectively.
Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Learning Temporal Event Knowledge for Continual Social Event Classification
abstract
With the rapid development of Internet and the burgeoning scale of social media, Social Event Classification (SEC) has garnered increasing attention. The existing study of SEC focuses on recognizing a fixed set of social events. However, in real-world scenarios, new social events continually emerge on social media, which suggests the necessity for a practical SEC model that can swiftly adapt to the evolving environment with incremental social events. Therefore, in this paper, we study a new yet crucial problem defined as Continual Social Event Classification (C-SEC), where new events continually emerge in the sequentially collected social data. Accordingly, we propose a novel Temporal Event Knowledge Network (TEKNet) to continually learn temporal event knowledge for C-SEC with temporally incremental events. First, we conduct present event knowledge learning to learn the classification of newly emerging events in the presently incoming data. Second, we design past event knowledge replay with self-knowledge distillation to consolidate the learned knowledge of past events and prevent catastrophic forgetting. Finally, we propose future event knowledge pretraining with a modality mixture mechanism to pretrain the classifiers for events that occur in the future. Comprehensive experiments on real-world social event datasets demonstrate the superiority of our proposed TEKNet for C-SEC.
Shengsheng Qian, Dizhan Xue, Huaiwen Zhang, Changsheng Xu
IEEE Trans. Knowl. Data Eng.4
2025 Multi-View User Preference Modeling for Personalized Text-to-Image Generation
abstract
Personalized text-to-image generation aims to synthesize images tailored to individual user preferences. Existing methods primarily generate customized content using a few reference images, which often struggle to mine user preferences from historical records, and thus fail to synthesize truly personalized content. In addition, it is difficult to directly incorporate the extracted feature of user preferences into the feature space of the generation model, since there exists a considerable gap between them. In this paper, we propose a novel multi-view personalized text-to-image generation method based on the diffusion model, named MVP-Diffusion, which learns instance- and user-level preferences from historical records and integrates them into the generation model. For instance-level user preference modeling, we employ a chain-of-thought prompting strategy to deduce preference keywords and integrate them into input prompts with the aid of a large language model. For user-level preference modeling, we construct a learnable embedding for each user to capture more comprehensive preferences by analyzing their historical records. An adaptive user preference fusion module is proposed to inject user preferences into the generation model via a set of learnable parameters. Experimental results demonstrate that the proposed method significantly enhances the personalization of the generated images compared to the other personalized text-to-image generation methods.
Huaiwen Zhang, Tianci Wu, Yinwei Wei
IEEE Trans. Multim.1
2024 SignGen: End-to-End Sign Language Video Generation with Latent Diffusion
Fan Qi, Huaiwen Zhang, Changsheng Xu
ECCV (53)3
2024 FedVAD: Enhancing Federated Video Anomaly Detection with GPT-Driven Semantic Distillation
Fan Qi, Ruijie Pan, Huaiwen Zhang, Changsheng Xu
ECCV (53)3
2024 Hierarchical Speaker Representation for Target Speaker Extraction
abstract
Target speaker extraction aims to isolate a specific speaker’s voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and integrate them into the separation network to separate the voice of the target speaker. However, the representation of the speaker embedding is too simplistic, often being merely a 1×1024 vector. This dense information makes it difficult for the separation network to harness effectively. To address this limitation, we introduce a pioneering methodology called Hierarchical Representation (HR) that seamlessly fuses anchor data across granular and overarching 5 layers of the separation network, enhancing the precision of target extraction. HR amplifies the efficacy of anchors to improve target speaker isolation. On the Libri-2talker dataset, HR substantially outperforms state-of-the-art time-frequency domain techniques. Further demonstrating HR’s capabilities, we achieved first place in the prestigious ICASSP 2023 Deep Noise Suppression Challenge. The proposed HR methodology shows great promise for advancing target speaker extraction through enhanced anchor utilization.
Shulin He, Huaiwen Zhang, Wei Rao 0002, Kanghao Zhang, Yukai Jv, Yang Yang 0121, Xueliang Zhang 0001
ICASSP2
2024 FEEL: A Framework for Evaluating Emotional Support Capability with Large Language Models
Huaiwen Zhang, Ming Wang 0006, Shi Feng 0001
ICIC (13)1
2024 Multi-Instance Multi-Label Learning for Text-motion Retrieval
abstract
Text-motion retrieval (TMR) is a significant cross-modal task that retrieves motion sequences semantically similar to a given query text. Existing TMR methods primarily utilize single embeddings to represent and align text and motion sequences. However, real-world motion sequences typically contain multiple atomic motions with complex semantics, which is hard to precisely capture by single embeddings. Additionally, the common co-occurring and coupling of atomic motions further post significant challenges in effective modeling and aligning text and motion sequences. In this paper, we regard TMR as a Multi-Instance Multi-Label (MIML) learning problem, where the motion sequence is viewed as a bag of atomic motions and the text is the bag of corresponding phrases. To address the MIML problem, we propose a novel Multi-Granularity Semantics Interaction (MGSI) approach, which effectively captures and aligns the semantics of text and motion sequences across various levels. Specifically, the MGSI approach initially decomposes both the query and motion sequences into three hierarchical levels: token, instance, and bag. Then, we utilize graph neural networks to explicitly model their semantics correlation and perform semantics interaction at these respective levels, precisely capturing the semantics at multiple granularities. To identify and model co-occurring atomic motions, we measure the frame-wise semantic consistency between motions and then fuse and interact the accordant ones to refine their representations. Finally, we exploit token, instance, and bag-wise semantics interaction to comprehensively align text and motion sequence. We evaluated our methods on two widely-used benchmark datasets, HumanML3D and KIT-ML. The proposed method achieves significant improvements, outperforming the state-of-the-art with a 23.09% increase in Rsum on HumanML3D and a 21.84% increase on KIT-ML.
Yang Yang 0121, Haoyu Shi, Huaiwen Zhang
ACM Multimedia4
2024 Modal-Enhanced Semantic Modeling for Fine-Grained 3D Human Motion Retrieval
abstract
Text to Motion Retrieval (TMR) is an emerging task to retrieve relevant motion sequences with the nature language description. The dominant approach learns a joint embedding space to measure global-level similarities. However, simple global embeddings are insufficient to represent complicated motion and textual details, such as the movement of specific body parts and the coordination among these body parts. In addition, most of the motion variations occur subtly and locally, resulting in semantic vagueness among these motions, which further presents considerable challenges in precisely aligning motion sequences with texts. To address these challenges, we propose a novel Modal-Enhanced Semantic Modeling (MESM) method, focusing on fine-grained alignment through enhanced modal semantics. Specifically, we develop a prompt-enhanced textual module (PTM) to generate detailed descriptions of specific body part movements, which comprehensively captures the fine-grained textual semantics for precise matching. We employ a skeleton-enhanced motion module (SMM) to effectively enhance the model's capability to represent intricate motions. This module leverages a graph convolutional network to meticulously model the intricate spatial dependencies among relevant body parts. To improve the sensitivity to the subtle motions, we further propose a text-driven semantics interaction module (TSIM). The TSIM assigns motion features into a set of aggregated descriptors and employs cross-attention to aggregate discriminative motion embeddings guided by text, enabling precise semantic alignment between subtle motions and corresponding texts. Extensive experiments conducted on two widely used benchmark datasets, HumanML3D and KIT-ML, demonstrate the effectiveness of our proposed method. Our approach outperforms existing state-of-the-art retrieval methods, achieving significant Rsum improvements of 24.28% on HumanML3D and 25.80% on KIT-ML.
Haoyu Shi, Huaiwen Zhang
ACM Multimedia2
2024 Hierarchical Semantics Alignment for 3D Human Motion Retrieval
abstract
Text to 3D human Motion Retrieval (TMR) is a challenging task in information retrieval, aiming to query relevant motion sequences with the natural language description. The conventional approach for TMR is to represent the data instances as point embeddings for alignment. However, in real-world scenarios, multiple motions often co-occur and superimpose on a single avatar. Simply aggregating text and motion sequences into a single global embedding may be inadequate for capturing the intricate semantics of superimposing motions. In addition, most of the motion variations occur locally and subtly, which further presents considerable challenges in precisely aligning motion sequences with their corresponding text. To address the aforementioned challenges, we propose a novel Hierarchical Semantics Alignment (HSA) framework for text-to-3D human motion retrieval. Beyond global alignment, we propose the Probabilistic-based Distribution Alignment (PDA) and a Descriptors-based Fine-grained Alignment (DFA) to achieve precise semantic matching. Specifically, the PDA encodes the text and motion sequences into multidimensional probabilistic distributions, effectively capturing the semantics of superimposing motions. By optimizing the problem of probabilistic distribution alignment, PDA achieves a precise match between superimposing motions and their corresponding text. The DFA first adopts a fine-grained feature gating by selectively filtering to the significant and representative local representations and meanwhile excluding the interferences of meaningless features. Then we adaptively assign local representations from text and motion into a set of cross-modal local aggregated descriptors, enabling local comparison and interaction between fine-grained text and motion features. Extensive experiments on two widely used benchmark datasets, HumanML3D and KIT-ML, demonstrate the effectiveness of the proposed method. It significantly outperforms existing state-of-the-art retrieval methods, achieving Rsum improvements of 24.74% on HumanML3D and 23.08% on KIT-ML.
Yang Yang 0121, Haoyu Shi, Huaiwen Zhang
SIGIR3
2024 T3RD: Test-Time Training for Rumor Detection on Social Media
abstract
With the increasing number of news uploaded to the internet daily, rumor detection has garnered significant attention in recent years. Existing rumor detection methods excel on familiar topics with sufficient training data (high resource) collected from the same domain. However, when facing emergent events or rumors propagated in different languages, the performance of these models is significantly degraded, due to the lack of training data and prior knowledge (low resource). To tackle this challenge, we introduce the Test-Time Training for Rumor Detection (T^3RD) to enhance the performance of rumor detection models on low-resource datasets. Specifically, we introduce self-supervised learning (SSL) as an auxiliary task in the test-time training. It consists of global and local contrastive learning, in which the global contrastive learning focuses on obtaining invariant graph representations and the local one focuses on acquiring invariant node representations. We employ the auxiliary SSL tasks for both the training and test-time training phases to mine the intrinsic traits of test samples and further calibrate the trained model for these test samples. To mitigate the risk of distribution distortion in test-time training, we introduce feature alignment constraints aimed at achieving a balanced synergy between the knowledge derived from the training set and the test samples. The experiments conducted on the two widely used cross-domain datasets demonstrate that the proposed model achieves a new state-of-the-art in performance. Our code is available at https://github.com/social-rumors/T3RD.
Huaiwen Zhang, Xinxin Liu 0016, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
WWW1
2024 A Versatile Multimodal Learning Framework for Zero-Shot Emotion Recognition
abstract
Multi-modal Emotion Recognition (MER) aims to identify various human emotions from heterogeneous modalities. With the development of emotional theories, there are more and more novel and fine-grained concepts to describe human emotional feelings. Real-world recognition systems often encounter unseen emotion labels. To address this challenge, we propose a versatile zero-shot MER framework to refine emotion label embeddings for capturing inter-label relationships and improving discrimination between labels. We integrate prior knowledge into a novel affective graph space that generates tailored label embeddings capturing inter-label relationships. To obtain multimodal representations, we disentangle the features of each modality into egocentric and altruistic components using adversarial learning. These components are then hierarchically fused using a hybrid co-attention mechanism. Furthermore, an emotion-guided decoder exploits label-modal dependencies to generate adaptive multimodal representations guided by emotion embeddings. We conduct extensive experiments with different multimodal combinations, including visual-acoustic and visual-textual inputs, on four datasets in both single-label and multi-label zero-shot settings. Results demonstrate the superiority of our proposed framework over state-of-the-art methods.
Fan Qi, Huaiwen Zhang, Xiaoshan Yang, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Nonparametric Clustering-Guided Cross-View Contrastive Learning for Partially View-Aligned Representation Learning
abstract
With the increasing availability of multi-view data, multi-view representation learning has emerged as a prominent research area. However, collecting strictly view-aligned data is usually expensive, and learning from both aligned and unaligned data can be more practicable. Therefore, Partially View-aligned Representation Learning (PVRL) has recently attracted increasing attention. After aligning multi-view representations based on their semantic similarity, the aligned representations can be utilized to facilitate downstream tasks, such as clustering. However, existing methods may be constrained by the following limitations: 1) They learn semantic relations across views using the known correspondences, which is incomplete and the existence of false negative pairs (FNP) can significantly impact the learning effectiveness; 2) Existing strategies for alleviating the impact of FNP are too intuitive and lack a theoretical explanation of their applicable conditions; 3) They attempt to find FNP based on distance in the common space and fail to explore semantic relations between multi-view data. In this paper, we propose a Nonparametric Clustering-guided Cross-view Contrastive Learning (NC3L) for PVRL, in order to address the above issues. Firstly, we propose to estimate the similarity matrix between multi-view data in the marginal cross-view contrastive loss to approximate the similarity matrix of supervised contrastive learning (CL). Secondly, we establish the theoretical foundation for our proposed method by analyzing the error bounds of the loss function and its derivatives between our method and supervised CL. Thirdly, we propose a Deep Variational Nonparametric Clustering (DeepVNC) by designing a deep reparameterized variational inference for Dirichlet process Gaussian mixture models to construct cluster-level similarity between multi-view data and discover FNP. Additionally, we propose a reparameterization trick to improve the robustness and the performance of our proposed CL method. Extensive experiments on four widely used benchmark datasets show the superiority of our proposed method compared with state-of-the-art methods.
Shengsheng Qian, Dizhan Xue, Jun Hu 0016, Huaiwen Zhang, Changsheng Xu
IEEE Trans. Image Process.4
2023 C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language Recognition
abstract
Continuous Sign Language Recognition (CSLR) aims to transcribe the signs of an untrimmed video into written words or glosses. The mainstream framework for CSLR consists of a spatial module for visual representation learning, a temporal module aggregating the local and global temporal information of frame sequence, and the connectionist temporal classification (CTC) loss, which aligns video features with gloss sequence. Unfortunately, the language prior implicit in the gloss sequence is ignored throughout the modeling process. Furthermore, the contextualization of glosses is further ignored in alignment learning, as CTC makes an independence assumption between glosses. In this paper, we propose a Cross-modal Contextualized Sequence Transduction (C2ST) for CSLR, which effectively incorporates the knowledge of gloss sequence into the process of video representation learning and sequence transduction. Specifically, we introduce a cross-modal context learning framework for CSLR, in which the linguistic features of gloss sequences are extracted by a language model, and recurrently integrate with visual features for video modelling. Moreover, we introduce the contextualized sequence transduction loss that incorporates the contextual information of gloss sequences in label prediction, without making any independence assumptions between the glosses. Our method sets the new state of the art on three widely used large-scale sign language recognition datasets: Phoenix-2014, Phoenix-2014-T, and CSL-Daily. On CSL-Daily, our approach achieves an absolute gain of 4.9% WER compared to the best published results.
Huaiwen Zhang, Zihang Guo, Yang Yang 0121, De Hu
ICCV1
2023 C2MR: Continual Cross-Modal Retrieval for Streaming Multi-modal Data
abstract
Massive numbers of new images are uploaded to the internet every day. However, existing cross-modal retrieval (CMR) approaches struggle to accommodate this continuously growing data. The prevalent practice involves periodically retraining or fine-tuning a new model based on the accumulated data, which in turn invalidates billions of indexed features extracted by the previous model and incurs another substantial computational cost to extract new features for the entire data archive. Is it possible to develop a retrieval model that effectively captures the knowledge of upcoming sessions while preserving the discriminative power of features extracted in previous sessions? In this paper, we propose an online continual learning setup, OC-CMR, to formalize the data-incremental growth challenge faced by cross-modal retrieval systems. It consists of two key settings: 1) Similar to the real-world scenarios, the streaming multi-modal data arrives once per session; 2) Consider the computational costs, each instance of archived data has its feature extracted only once and by its corresponding model in its session. Based on our OC-CMR, we perform in-depth evaluations of state-of-the-art cross-modal retrieval methods and observe that they suffer from representational shift and collapse due to the catastrophic forgetting. To address this issue, we propose the Continual Cross-Modal Retrieval (C2MR) approach, which learns a shared common space not only across modalities but also sessions and maintains relationships between samples from distinct sessions via cross-modal relational coherence and semantic representation coordination. We construct two new benchmarks by adapting MS-COCO and Flickr30K datasets to the OC-CMR setting, providing a more challenging evaluation framework for CMR tasks. Experimental results demonstrate that our method effectively alleviates forgetting and significantly outperforms combinations of previous arts in cross-modal retrieval and continual learning.
Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
ACM Multimedia1
2023 Distributed Sampling Rate Offset Estimation Over Acoustic Sensor Networks Based on Asynchronous Network Newton Optimization
abstract
Sampling rate synchronization is an inevitable issue in distributed acoustic sensor networks. In this paper, an analytical sampling rate offset (SRO) estimation approach is first proposed, and then, it is extended to a distributed method that suitable for acoustic sensor networks with arbitrary communication graphs. Specifically, a linear-phase drift model in the short-time Fourier transform domain is used to approximate the SRO between each pair of microphone nodes. Next, after unwrapping the temporally averaged phase information, SROs are recovered analytically via a new weighted-sum criterion. Based on this, a distributed cost function is established at each node to obtain the SROs of all nodes simultaneously in a distributed manner. Finally, a state-of-the-art distributed algorithm named asynchronous network Newton optimization is adopted to carry out the distributed SRO estimation. The proposed method can effectively estimate the SROs among acoustic sensor nodes in noisy and reverberant environments. Compared with the existing approaches, it does not require an external central processor, and only local communications among nodes are needed. Experimental results confirm the validity of the proposed method.
De Hu, Huaiwen Zhang, Feilong Bao, Rui Wang 0046
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Debiased Video-Text Retrieval via Soft Positive Sample Calibration
abstract
With the emergence of enormous videos on various video apps, semantic video-text retrieval has become a critical task for improving the user experience. The primary paradigm for video-text retrieval learns the semantic video-text representations in a common space by pulling the positive samples close to the query and pushing the negative samples away. However, in practice, the video-text datasets contain only the annotations of positive samples. The negative samples are randomly drawn from the entire dataset. There may exist soft positive samples, which are sampled as negatives but share the same semantics as positive samples. Indiscriminately enforcing the model to push all the negative samples away from the query leads to inaccurate supervision and then misleads the video-text feature representation learning. In this paper, we introduce debiased video-text retrieval objectives that calibrate the punishment of soft positive samples. In particular, we propose a novel uncertainty measure framework to estimate the credibility of negative samples for each instance. Then, the reliability of negative samples is used to find the soft positive samples and rescale their contribution within video-text retrieval losses, including triplet loss and contrastive loss. Experimental results on five widely used datasets demonstrate that our debiased video-text retrieval objectives achieve significant performance improvements and establish a new state-of-the-art.
Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.1
2023 Robust Video-Text Retrieval Via Noisy Pair Calibration
abstract
Video-text retrieval is a fundamental task in managing the emerging massive amounts of video data. The main challenge focuses on learning a common representation space for videos and queries where the similarity measurement can reflect the semantic closeness. However, existing video-text retrieval models may suffer from the following noise in the common space learning procedure: First, the video-text correspondences in positive pairs may not be exact matches. The crowdsourcing annotation for existing datasets leads to inevitable tagging noise for non-expert annotators. Second, the learning of video-text representation is based on the negative samples randomly sampled. Instances that are semantically similar to the query may be incorrectly categorized as negative samples. To alleviate the adverse impact of these noisy pairs, we propose a novel robust video-text retrieval method that protects the model from noisy positive and negative pairs by identifying and calibrating noisy pairs with their uncertainty score. In particular, we propose a noisy pair identifier, which divides the training dataset into noisy and clean subsets based on the estimated uncertainty of each pair. Then, with the help of uncertainties, we calibrate the two types of noisy pairs with an adaptive margin triplet loss and a weighted triplet loss function, respectively. To verify the effectiveness of our methods, we conduct extensive experiments on three widely used datasets. Experimental results show that the proposed robust video-text retrieval methods successfully identify and calibrate the noisy pairs and improve retrieval performance.
Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu
IEEE Trans. Multim.1
2022 Alleviating the Loss-Metric Mismatch in Supervised Single-Channel Speech Enhancement
abstract
In this paper, we study the loss-metric mismatch problem of supervised single-channel speech enhancement system. Most of the existing speech enhancement systems achieve unsatisfying performance since their empirically selected loss functions have semantic gaps with the non-differentiable evaluation metrics, a.k.a., the loss-metric mismatch problem. In this work, we propose a simple yet efficient method to generate suitable loss functions for the real front-end speech enhancement scenarios to alleviate the loss-metric mismatch problem. Specifically, we adopt the function smoothing technique and approximate the non-differentiable evaluation metrics by a set of basis functions and their linear combination. Experimental results demonstrate that the loss function generated by our method helps the speech enhancement system achieve remarkable performance in most evaluation metrics than the traditional empirically selected ones.
Yang Yang 0121, Hui Zhang 0031, Xueliang Zhang 0001, Huaiwen Zhang
ICASSP4
2022 Feeling Without Sharing: A Federated Video Emotion Recognition Framework Via Privacy-Agnostic Hybrid Aggregation
abstract
The explosion of video data brings new opportunities and challenges for emotion recognition. Video emotion applications have great commercial value, but the potential to involve illegal snooping on personal feelings has led to controversy over privacy protection. The federated learning (FL) paradigm can substantially address the growing public concerns about data privacy in video emotion recognition. However, conventional FL methods perform poorly due to the uniqueness of the task: the data are heterogeneous across clients induced by emotional label skew and cross-culture expression differences. To mitigate the heterogeneous data, we propose EmoFed, a practical framework of federated learning video-based emotion recognition via multi-group clustering and privacy-agnostic hybrid aggregation. It yields a generically applicable and improved model while protecting privacy, which trains local models under group-aware personalized aggregation. To further encourage communicating comprehensive and privacy-agnostic information among clients, we upload model parameters of both the global layers and personalization layers to the server. We utilize the homomorphically encrypted method for personalization layers, which incurs no learning accuracy loss since no noise is added to the model updates during the encryption/decryption process. The proposed method works on video-based emotion recognition tasks to predict actors' emotional expressions and induced emotion by viewers. Extensive experiments and ablation studies on four benchmarks have demonstrated the efficacy and practicability of our method.
Fan Qi, Zixin Zhang 0004, Xianshan Yang, Huaiwen Zhang, Changsheng Xu
ACM Multimedia4
2022 Multi-Modal Meta Multi-Task Learning for Social Media Rumor Detection
abstract
With the rapid development of social media platforms and the increasing scale of the social media data, the rumor detection task has become vitally important since the authenticity of posts cannot be guaranteed. To date, Many approaches have been proposed to facilitate the rumor detection process by utilizing the multi-task learning mechanism, which aims to improve the performance of rumor detection task by leveraging the useful information in the stance detection task. However, most of the existing approaches suffer from three limitations: (1) only focus on the textual content and ignore the multi-modal information which is key component contained in social media data; (2) ignore the difference of feature space between the stance detection task and rumor detection task, resulting in the unsatisfactory usage of stance information; (3) largely neglect the semantic information hidden in the fine-grained stance labels. Therefore, in this paper, we design a Multi-modal Meta Multi-Task Learning (MM-MTL) framework for social media rumor detection. To make use of multiple modalities, we design a multi-modal post embedding layer which considers both textual and visual content. To overcome the feature-sharing problem of the stance detection task and rumor detection task, we propose a meta knowledge-sharing scheme to share some higher meta network-layers and capture the meta knowledge behind the multi-modal post. To better utilize the semantic information hidden in the fine-grained stance labels, we employ the attention mechanism to estimate the weight of each reply. Extensive experiments on two Twitter benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance.
Huaiwen Zhang, Shengsheng Qian, Quan Fang, Changsheng Xu
IEEE Trans. Multim.1
2021 Dual Adversarial Graph Neural Networks for Multi-label Cross-modal Retrieval
abstract
Cross-modal retrieval has become an active study field with the expanding scale of multimodal data. To date, most existing methods transform multimodal data into a common representation space where semantic similarities between items can be directly measured across different modalities. However, these methods typically suffer from following limitations: 1) They usually attempt to bridge the modality gap by designing losses in the common representation space which may not be sufficient to eliminate potential heterogeneity of different modalities in the common space. 2) They typically treat labels as independent individuals and ignore label relationships which are important for constructing semantic links between multimodal data. In this work, we propose a novel Dual Adversarial Graph Neural Networks (DAGNN) composed of the dual generative adversarial networks and the multi-hop graph neural networks, which learn modality-invariant and discriminative common representations for cross-modal retrieval. Firstly, we construct the dual generative adversarial networks to project multimodal data into a common representation space. Secondly, we leverage the multi-hop graph neural networks, in which a layer aggregation mechanism is proposed to exploit multi-hop propagation information, to capture the label correlation dependency and learn inter-dependent classifiers. Comprehensive experiments conducted on two cross-modal retrieval benchmark datasets, NUS-WIDE and MIRFlickr, indicate the superiority of DAGNN.
Shengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang, Changsheng Xu
AAAI3
2021 Global Relation-Aware Attention Network for Image-Text Retrieval
abstract
The cross-modal image-text retrieval has attracted extensive attention in recent years, which contributes to the development of search engine. Fine-grained features and cross-attention have been widely used in past researches to reach the goal of cross-modal image-text matching. Although cross-related methods have achieved remarkable results, the features must be encoded again in evaluation phase due to the interaction of the two modalities, which is unsuitable for actual scenarios of search engine development. In addition, the aggregated feature does not contain sufficient semantics since it is merely obtained by simple mean pooling. Furthermore, connecting weights of self-attention blocks are target position invariant, which lacks the expected adaptability. To tackle these limitations, in this paper, we propose a novel Global Relation-aware Attention Network (GRAN) for image-text retrieval by designing Global Attention Module (GAM) and Relation-aware Attention Module (RAM) which play an important role in modeling the global feature and the relationships of local fragments. Firstly, we propose Global Attention Module (GAM) followed the fine-grained features to obtain meaningful global feature. Secondly, we use several stacked transformer encoders to further encode features separately. Finally, we propose Relation-aware Attention Module (RAM) to generate a vector which represents the relation information to infer the attention intensity of pairwise fragments. The local features, the global feature, and their relations are considered jointly to conduct an efficient image-text retrieval. Extensive experiments are conducted on the benchmark datasets of Flickr30K and MSCOCO, demonstrating the superiority of our method. On the Flickr30K, compared to the state-of-the-art method TERAN, we improve [email protected](K=1) metric by 5.8% and 4.0 on the image and text retrieval tasks, respectively.
Jie Cao 0002, Shengsheng Qian, Huaiwen Zhang, Quan Fang, Changsheng Xu
ICMR3
2021 Efficient Graph Deep Learning in TensorFlow with tf_geometric
abstract
We introduce tf_geometric1, an efficient and friendly library for graph deep learning, which is compatible with both TensorFlow 1.x and 2.x. It provides kernel libraries for building Graph Neural Networks (GNNs) as well as implementations of popular GNNs. The kernel libraries consist of infrastructures for building efficient GNNs, including graph data structures, graph map-reduce framework, graph mini-batch strategy, etc. These infrastructures enable tf_geometric to support single-graph computation, multi-graph computation, graph mini-batch, distributed training, etc.; therefore, tf_geometric can be used for a variety of graph deep learning tasks, such as node classification, link prediction, and graph classification. Based on the kernel libraries, tf_geometric implements a variety of popular GNN models. To facilitate the implementation of GNNs, tf_geometric also provides some other libraries for dataset management, graph sampling, etc. Different from existing popular GNN libraries, tf_geometric provides not only Object-Oriented Programming (OOP) APIs, but also Functional APIs, which enable tf_geometric to handle advanced tasks such as graph meta-learning. The APIs are friendly and suitable for both beginners and experts.
Jun Hu 0016, Shengsheng Qian, Quan Fang, Youze Wang, Huaiwen Zhang, Changsheng Xu
ACM Multimedia6
2021 Multimodal Disentangled Domain Adaption for Social Media Event Rumor Detection
abstract
With the rapid development of social media and the increasing scale of social media data, the rumor detection on social media platforms has become vitally crucial. The key challenges for rumor detection on social media platforms are how to identify rumors deeply entangled with the specific content and how to detect rumors for the emerging social media events without labeled data. Unfortunately, most of the existing approaches can hardly handle these challenges since they tend to learn event-specific features and cannot transfer the learned features to newly emerged events. To tackle the above challenges, we propose a novel Multimodal Disentangled Domain Adaption (MDDA) method which can derive event-invariant features and thus benefit the detection of rumors on emerging social media events. The model consists of two components: the multimodal disentangled representation learning and the unsupervised domain adaptation. The multimodal disentangled representation learning is responsible for disentangling the multimedia posts into the content features and the rumor style features, and removing the content-specific features from post representation. The unsupervised domain adaptation aims to filter out the event-specific features and keep shared rumor style features among events. Based on the final event-invariant rumor style features, we train a robust social media rumor detector that can transfer knowledge from source events to the target events, which can perform well on the newly emerged events. Extensive experiments on two Twitter benchmark datasets demonstrate that our rumor detection model outperforms state-of-the-art methods.
Huaiwen Zhang, Shengsheng Qian, Quan Fang, Changsheng Xu
IEEE Trans. Multim.1
2019 Multi-modal Knowledge-aware Event Memory Network for Social Media Rumor Detection
abstract
The wide dissemination and misleading effects of online rumors on social media have become a critical issue concerning the public and government. Detecting and regulating social media rumors is important for ensuring users receive truthful information and maintaining social harmony. Most of the existing rumor detection methods focus on inferring clues from media content and social context, which largely ignores the rich knowledge information behind the highly condensed text which is useful for rumor verification. Furthermore, existing rumor detection models underperform on unseen events because they tend to capture lots of event-specific features in seen data which cannot be transferred to newly emerged events. In order to address these issues, we propose a novel Multimodal Knowledge-aware Event Memory Network (MKEMN) which utilizes the Multi-modal Knowledge-aware Network (MKN) and Event Memory Network (EMN) as building blocks for social media rumor detection. Specifically, the MKN learns the multi-modal representation of the post on social media and retrieves external knowledge from real-world knowledge graph to complement the semantic representation of short texts of posts and takes conceptual knowledge as additional evidence to improve rumor detection. The EMN extracts event-invariant features of events and stores them into global memory. Given an event representation, the EMN takes it as a query to retrieve the memory network and output the corresponding features shared among events. With the additional information provided by EMN, our model can learn robust representations of events and consistently perform well on the newly emerged events. Extensive experiments on two Twitter benchmark datasets demonstrate that our rumor detection method achieves much better results than state-of-the-art methods.
Huaiwen Zhang, Quan Fang, Shengsheng Qian, Changsheng Xu
ACM Multimedia1
2018 Learning Multimodal Taxonomy via Variational Deep Graph Embedding and Clustering
abstract
Taxonomy learning is an important problem and facilitates various applications such as semantic understanding and information retrieval. Previous work for building semantic taxonomies has primarily relied on labor-intensive human contributions or focused on text-based extraction. In this paper, we investigate the problem of automatically learning multimodal taxonomies from the multimedia data on the Web. A systematic framework called Variational Deep Graph Embedding and Clustering (VDGEC) is proposed consisting of two stages as concept graph construction and taxonomy induction via variational deep graph embedding and clustering. VDGEC discovers hierarchical concept relationships by exploiting the semantic textual-visual correspondences and contextual co-occurrences in an unsupervised manner. The unstructured semantics and noisy issues of multimedia documents are carefully addressed by VDGEC for high quality taxonomy induction. We conduct extensive experiments on the real-world datasets. Experimental results demonstrate the effectiveness of the proposed framework, where VDGEC outperforms previous unsupervised approaches by a large gap.
Huaiwen Zhang, Quan Fang, Shengsheng Qian, Changsheng Xu
ACM Multimedia1
2017 A Demo for Image-Based Personality Test
Huaiwen Zhang, Jiaming Zhang 0006, Jitao Sang 0001, Changsheng Xu
MMM (2)1