Chaoyue Ding

dblp:269/4517 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2025
0009-0000-0161-4838ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Counterfactual Debiasing for Physical Audiovisual Commonsense Reasoning
abstract
Physical commonsense is an essential aspect of human cognition, involving an intuitive understanding of the physical properties and interactions of everyday objects and materials. Though physical commonsense reasoning should inherently be a multisensory task, integrating both video and audio signals, existing physical audiovisual commonsense reasoning (PACR) models predominantly rely on visual information. This reliance leads to spurious correlations and undermines the models’ reasoning and generalization abilities. To counteract this, we introduce a model-agnostic Counterfactual Physical Audiovisual Commonsense Reasoning (CF-PACR) framework aimed at mitigating visual bias-induced spurious effects. Specifically, we construct a traditional PACR model using both audio and visual information as the factual reasoning model. Subsequently, in the counterfactual reasoning model, we isolate visual information to estimate direct effects. Finally, we subtract the direct effects from the total effects across modalities to derive indirect effects, thereby mitigating visual biases. Extensive experiments validate the effectiveness and generalizability of CF-PACR in alleviating the spurious correlations between visual modality and model predictions.
Daoming Zong, Chaoyue Ding, Kaitao Chen, Shuaiyu Wang
AAAI2
2025 Enhance the old representations' adaptability dynamically for exemplar-free continual learning
Kunchi Li, Chaoyue Ding, Jun Wan 0001
Neurocomputing2
2024 Balancing Multimodal Learning via Online Logit Modulation
Daoming Zong, Chaoyue Ding, Baoxiang Li, Jiakui Li, Ken Zheng
IJCAI2
2024 Toward Explainable Physical Audiovisual Commonsense Reasoning
abstract
For AI systems to be safely and reliably grounded in the real world, they should possess the ability of physical commonsense reasoning. Physical commonsense reasoning is essentially a multisensory task as physical properties of objects are manifested through multiple perception modalities, including both visual and auditory. In this study, we constructed two new benchmarks, called PACS-Reason and PACS-Reason+, for explainable physical audiovisual commonsense reasoning (EPACS), in which each datapoint is accompanied by a golden detailed rationale (intermediate reasoning path) to explain the answer selection. Moreover, we present PAVC-Reasoner, a multimodal large language model (LLM) designed to reason about physical commonsense attributes. The model aligns different modalities with the language modality by integrating three different perceivers for cross-modal pretraining and instruction finetuning at multiple granularities. It utilizes an LLM as a cognitive engine to process multimodal inputs and output convincing intermediate reasoning paths as justification for inferring answers. Numerous experiments have demonstrated the effectiveness and superiority of PAVC-Reasoner as a baseline model for studying EPACS. Most attractively, PAVC-Reasoner is capable of reasoning and obtaining strong interpretable explicit reasoning paths, signifying a significant stride towards real-world physical commonsense reasoning.
Daoming Zong, Chaoyue Ding, Kaitao Chen
ACM Multimedia2
2023 Stable Speech Emotion Recognition with Head-k-Pooling Loss
Chaoyue Ding, Jiakui Li, Daoming Zong, Baoxiang Li, Tian-Hao Zhang, Qunyan Zhou 0002
INTERSPEECH1
2023 AcFormer: An Aligned and Compact Transformer for Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis (MSA) is a popular research topic aimed at utilizing multimodal signals for understanding human emotions. The primary approach to solving this task is to develop complex fusion techniques. However, the heterogeneity and unaligned nature between modalities pose significant challenges to fusion. Additionally, existing methods lack consideration for the efficiency of modal fusion. To tackle these issues, we propose AcFormer, which contains two core ingredients: i) contrastive learning within and across modalities to explicitly align different modality streams before fusion; and ii) pivot attention for multimodal interaction/fusion. The former encourages positive triplets of image-audio-text to have similar representations in contrast to negative ones. The latter introduces attention pivots that can serve as cross-modal information bridges and limit cross-modal attention to a certain number of fusion pivot tokens. We evaluate AcFormer on multiple MSA tasks, including multimodal emotion recognition, humor detection, and sarcasm detection. Empirical evidence shows that AcFormer achieves the optimal performance with minimal computation cost compared to previous state-of-the-art methods. Our code is publicly available at https://github.com/dingchaoyue/AcFormer.
Daoming Zong, Chaoyue Ding, Baoxiang Li, Jiakui Li, Ken Zheng, Qunyan Zhou 0002
ACM Multimedia2
2023 Building Robust Multimodal Sentiment Recognition via a Simple yet Effective Multimodal Transformer
abstract
In this paper, we present the solutions to the MER-MULTI and MER-NOISE sub-challenges of the Multimodal Emotion Recognition Challenge (MER 2023). For the tasks MER-MULTI and MER-NOISE, participants are required to recognize both discrete and dimensional emotions. Particularly, in MER-NOISE, the test videos are corrupted with noise, necessitating the consideration of modality robustness. Our empirical findings indicate that different modalities contribute differently to the tasks, with a significant impact from the audio and visual modalities, while the text modality plays a weaker role in emotion prediction. To facilitate subsequent multimodal fusion, and considering that language information is implicitly embedded in large pre-trained speech models, we have made the deliberate choice to abandon the text modality and solely utilize visual and acoustic modalities for these sub-challenges. To address the potential underfitting of individual modalities during multimodal training, we propose to jointly train all modalities via a weighted blending of supervision signals. Furthermore, to enhance the robustness of our model, we employ a range of data augmentation techniques at the image level, waveform level, and spectrogram level. Experimental results show that our model ranks 1st in both MER-MULTI (0.7005) and MER-NOISE (0.6846) sub-challenges, validating the effectiveness of our method. Our code is publicly available at https://github.com/dingchaoyue/Multimodal-Emotion-Recognition-MER-and-MuSe-2023-Challenges.
Daoming Zong, Chaoyue Ding, Baoxiang Li, Dinghao Zhou, Jiakui Li, Ken Zheng, Qunyan Zhou 0002
ACM Multimedia2
2023 Concept Drift Adaptation for Time Series Anomaly Detection via Transformer
Chaoyue Ding, Jing Zhao 0015, Shiliang Sun
Neural Process. Lett.1
2022 Multi-Modal Adversarial Example Detection with Transformer
abstract
Although deep neural networks have shown great potential for many tasks, they are vulnerable to adversarial examples, which are generated by adding small perturbations to natural examples. Recently, many studies have proved that making full use of different modalities can effectively enhance the representational ability of deep neural networks. We propose a multi-modal deep fusion Transformer, termed MDFT. First, the audio feature and the rich semantic text features are extracted by audio encoders and text encoders, respectively. Then, multi-modal attention mechanisms are established to capture the high-level interactions between the audio and linguistic domains to obtain joint multi-modal representation. Finally, the representation is propagated to a dense layer to generate the detection result. The accuracy of this model compared with its unimodal variant on WiAd dataset and BlAd dataset are improved by 0.12 % and 0.19 %, respectively. Experimental results on the two datasets show that MDFT outperforms its unimodal variant model.
Chaoyue Ding, Shiliang Sun, Jing Zhao 0015
IJCNN1
2022 Speed-Robust Keyword Spotting Via Soft Self-Attention on Multi-Scale Features
abstract
In this work, we focus on the robustness of keyword spotting (KWS) at various speech speeds. First, to enable small-footprint KWS, we graft a depthwise separable convolution and a dilated temporal convolution to build our basic model block. Second, to make KWS desensitized to speech rate, a simple yet effective soft self-attention is proposed to operate between different hierarchical features, offering our model the ability to be aware of the speaker's speech rate and to dynamically integrate multi-scale features from varying sizes of receptive fields. Besides, we construct two frame-level annotated Chinese intelligent in-car speech commands datasets, termed Car-C1 and Car-C2, for model evaluation. Experimental results show that our model achieves the highest accuracy on both datasets with a comparable number of model parameters and computational cost. Meanwhile, the sensitivity analysis suggests that our model outperforms baselines when tested on a fast corpus and a slow corpus.
Chaoyue Ding, Jiakui Li, Martin Zong, Baoxiang Li
SLT1
2021 Multi-Task Transformer with Input Feature Reconstruction for Dysarthric Speech Recognition
abstract
Dysarthria is a motor speech disorder caused by damage to the part of the nervous system that controls the physical production of speech. It poses great challenges in building robust dysarthric speech recognition (DSR) due to the high inter- and intra-speaker variability. To this end, we propose a multi-task Transformer with input feature reconstruction as an auxiliary task, where the main task of DSR and the auxiliary reconstruction task share the same encoder network. The auxiliary task aims to reconstruct clear speech features from corrupted speech of healthy speakers (intra-domain) or dysarthric speakers (cross-domain). Further, to alleviate the imbalanced distribution of dysarthria data sets, we devise an adaptive rebalance sampling scheme to improve the utterance sampling frequency of dysarthric speech. Experimental results show that the proposed model considerably outperforms other baselines across speakers with varying severity of dysarthria.
Chaoyue Ding, Shiliang Sun, Jing Zhao 0015
ICASSP1
2020 SiamBOMB: A Real-time AI-based System for Home-cage Animal Tracking, Segmentation and Behavioral Analysis
abstract
Biologists often need to handle numerous video-based home-cage animal behavior analysis tasks that require massive workloads. Therefore, we develop an AI-based multi-species tracking and segmentation system, SiamBOMB, for real-time and automatic home-cage animal behavioral analysis. In this system, a background-enhanced Siamese-based network with replaceable modular design ensures the flexibility and generalizability of the system, and a user-friendly interface makes it convenient to use for biologists. This real-time AI system will effectively reduce the burden on biologists.
Xi Chen 0031, Hao Zhai 0003, Danqian Liu, Weifu Li, Chaoyue Ding, Qiwei Xie, Hua Han 0001
IJCAI5