VLDB 2026 Research / reviewers in the wild / expert
Siyang Song
dblp:220/3096
· DBLP profile ↗
81ranked-venue papers
10as first author
76since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 7 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 5 first-author · 46 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Personalised Human Internal Cognition from External Expressive Behaviours for Real Personality RecognitionabstractAutomatic real personality recognition (RPR) aims to evaluate human real personality traits from their expressive behaviours. However, most existing solutions generally act as external observers to infer observers' personality impressions based on target individuals' expressive behaviours, which significantly deviate from their real personalities and consistently lead to inferior recognition performance. Inspired by the association between real personality and human internal cognition underlying the generation of expressive behaviours, we propose a novel RPR approach that efficiently simulates personalised internal cognition from external short audio-visual behaviours expressed by target individual. The simulated personalised cognition, represented as a set of network weights that enforce the personalised network to reproduce the individual-specific facial reactions, is further encoded as a graph containing two-dimensional node and edge feature matrices, with a novel 2D Graph Neural Network (2D-GNN) proposed for inferring real personality traits from it. To simulate real personality-related cognition, an end-to-end (E2E) strategy is designed to jointly train our cognition simulation, 2D graph construction, and personality recognition modules. Experiments show our approach’s effectiveness in capturing real personality traits with superior computational efficiency. Xiangyu Kong 0001, Hengde Zhu, Haoqin Sun, Jiayan Gu, Xinyi Ni, Wei Zhang 0243, Shizhe Liu, Siyang Song |
AAAI | 9 |
| 2026 | MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label GenerationabstractThe lack of large-scale, demographically diverse face images with precise Action Unit (AU) occurrence and intensity annotations has long been recognized as a fundamental bottleneck in developing generalizable facial AU recognition systems. In this paper, we propose MAUGen, a diffusion-based multi-modal framework that jointly generates a large collection of photorealistic facial expressions and anatomically consistent AU labels, including both occurrence and intensity, conditioned on a single descriptive text prompt. Our MAUGen involves two key modules: (1) a Multi-modal Representation Learning (MRL) module that captures the relationships among the paired facial textual description, facial identity, facial expression image, and AU activations within a unified latent space; and (2) a Diffusion-based Image-label Generator (DIG) that decodes the obtained joint representation into aligned facial image-label pairs across diverse identities. Under this framework, we introduce the Multi-Identity Facial Action (MIFA), a large-scale multi-modal (i.e., text descriptions, face images with labels) synthetic dataset that features comprehensive AU annotations and identity variations. Extensive experiments demonstrate that MAUGen outperforms existing methods in synthesizing photorealistic, demographically diverse facial images, along with semantically aligned AU labels. Ye Lou, Ao Gao, Wei Zhang 0243, Siyang Song |
AAAI | 5 |
| 2026 | Explainable Depression Assessment from Face Videos by Weakly Supervised LearningabstractExisting video-based automatic depression assessment (ADA) approaches frequently achieve video-level depression assessment by aggregating features or predictions of individual frames or equal-length segments within the given video. While their performances have been largely enhanced by recent advanced deep learning models, they typically fail to explicitly consider the varied importance of depression-related behavioural cues across different video segments, i.e., segments within one video may contain behaviours reflecting varying levels of depression. Underestimating segment-level variations can obscure the detection of facial behaviour cues associated with depression, thereby undermining the accuracy and interpretability of video-based depression detection systems. In this paper, we propose a novel video-based ADA approach that specifically identifies and differentiates video segments that exhibit depression-related facial behaviours across varying temporal durations, providing clear insights into how each segment contributes to the video-level depression prediction. To achieve this, a novel weakly supervised strategy is proposed to compare segment-level behaviours with video-level depression label, enabling the model to assign depression-relevant scores to multiple temporal scale video segments and attend selectively to those most indicative of depressive states. Extensive experiments on the AVEC 2013 and AVEC 2014 face video depression datasets demonstrate the effectiveness of our approach. Rongfan Liao, Xiangyu Kong 0001, Shiqing Tang, Changzeng Fu, Weicheng Xie 0001, Lu Liu 0001, Siyang Song |
AAAI | 9 |
| 2026 | RA3-FDA: Resource-adaptive federated domain adaptation with dual heterogeneity awareness for EEG-based depression detection
Siyang Song, Huaning Wang, Jiewei Jiang, Dongmei Jiang, Jie Zhang 0028, Prayag Tiwari, Jiaqing Liu |
Expert Syst. Appl. | 4 |
| 2026 | Graph in Graph Neural NetworkabstractAbstract Existing Graph Neural Networks (GNNs) are limited to process graphs each of whose vertices is represented by a vector or a single value, limited their representing capability to describe complex objects. In this paper, we propose a novel GNN (called Graph in Graph Neural (GIG) Network) which can process graph-style data (called GIG sample) whose vertices are further represented by graphs. Given a set of graphs or a data sample whose components can be represented by a set of graphs (called multi-graph data sample), our GIG network starts with a GIG sample generation (GSG) module which encodes the input as a GIG sample , where each GIG vertex includes a graph. Then, a set of GIG hidden layers are stacked, with each consisting of: (1) a GIG vertex-level updating (GVU) module that individually updates the graph in every GIG vertex based on its internal information; and (2) a global-level GIG sample updating (GGU) module that updates graphs in all GIG vertices based on their relationships, making the updated GIG vertices become global context-aware. This way, both internal cues within the graph contained in each GIG vertex and the relationships among GIG vertices could be utilized for down-stream tasks. Experimental results demonstrate that our GIG network generalizes well for not only various generic graph analysis tasks but also real-world multi-graph data analysis (e.g., human skeleton video-based action recognition), which achieved the new state-of-the-art results on 15 out of 16 evaluated datasets. Our code is publicly available at https://github.com/wangjs96/Graph-in-Graph-Neural-Network . Jiongshu Wang, Jing Yang 0038, Jiankang Deng, Hatice Gunes, Siyang Song |
Int. J. Comput. Vis. | 5 |
| 2026 | SAFRG: Speech aligned multiple appropriate facial reaction generation
Shizhe Liu, Xiangyu Kong 0001, Junan Long, Jiayan Gu, Siyang Song |
Neurocomputing | 5 |
| 2026 | Compressed video-driven multimodal modeling and interaction for dynamic expression recognition
Weicheng Xie 0001, Junliang Zhang, Haijian Liang, LinLin Shen, Zhihui Lai 0001, Siyang Song, Zitong Yu |
Knowl. Based Syst. | 6 |
| 2026 | CPG: Contrastive Patch-Graph learning for 3D point cloud
Junjie Zhou 0001, Yingde Song, Chinwai Chiu, Yongping Xiong, Yuxin Luo, Siyang Song |
Pattern Recognit. | 6 |
| 2025 | MSAmba: Exploring Multimodal Sentiment Analysis with State Space ModelsabstractMultimodal sentiment analysis, which learns a model to process multiple modalities simultaneously and predict a sentiment value, is an important area of affective computing. Modeling sequential intra-modal information and enhancing cross-modal interactions are crucial to multimodal sentiment analysis. In this paper, we propose MSAmba, a novel hybrid Mamba-based architecture for multimodal sentiment analysis, consisting of two core blocks: Intra-Modal Sequential Mamba (ISM) block and Cross-Modal Hybrid Mamba (CHM) block, to comprehensively address the above-mentioned challenges with hybrid state space models. Firstly, the ISM block models the sequential information within each modality in a bi-directional manner with the assistance of global information. Subsequently, the CHM blocks explicitly model centralized cross-modal interaction with a hybrid combination of Mamba and attention mechanism to facilitate information fusion across modalities. Finally, joint learning of the intra-modal tokens and cross-modal tokens is utilized to predict the sentiment values. This paper serves as one of the pioneering works to unravel the outstanding performances and great research potential of Mamba-based methods in the task of multimodal sentiment analysis. Experiments on CMU-MOSI, CMU-MOSEI and CH-SIMS demonstrate the superior performance of the proposed MSAmba over prior Transformer-based and CNN-based methods. Xilin He, Haijian Liang, Boyi Peng, Weicheng Xie 0001, Muhammad Haris Khan, Siyang Song, Zitong Yu |
AAAI | 6 |
| 2025 | DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression AssessmentabstractDepression can be reflected by long-term human spatio-temporal facial behaviours. While human face videos recorded in real-world usually have long and variable lengths, existing video-based depression assessment approaches frequently re-sample/down-sample such videos to short and equal-length videos, or split each video into several equal-length segments, where segment-level spatio-temporal facial behaviours are suppressed as a vector-style representations for RNN-based long-term (video-level) modelling. Both strategies lead to crucial information loss and distortion. In this paper, we propose a novel graph-style data structure called Matrixial Graph and an effective Matrixial Graph Neural Network (MGNN) for face video-based depression assessment, which can directly and end-to-end model long-term depression-specific spatio-temporal facial cues from variable-length videos without resampling/splitting videos or suppressing video segments to vectors. Importantly, the nodes in our matrixial graph are capable of including matrices of different shapes, and thus nodes of a matrix graph can directly represent all frame-level 2D facial feature maps (or images themselves) of an entire video regardless of its length. Then, our MGNN is the first GNN that can jointly process matrixial graphs containing varying numbers of nodes, which further learns matrix-style edge features, thereby facilitating to explicit model video-level multi-scale spatio-temporal facial behaviours among matrixial graph nodes for depression assessment. Experiments show that the explicit spatio-temporal modeling on 2D facial feature maps, facilitated by our matrixial graph/MGNN, provided significant benefits, leading our approach to achieve new state-of-the-art performances on AVEC2013 and AVEC2014 datasets with large advantages. Leijing Zhou, Shuanglin Li, Changzeng Fu, Jun Lu 0006, Jing Han 0009, Yi Zhang 0036, Siyang Song |
AAAI | 9 |
| 2025 | CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute EditingabstractFor efficient and high-fidelity local facial attribute editing, most existing editing methods either require additional fine-tuning for different editing effects or tend to affect beyond the editing regions. Alternatively, inpainting methods can edit the target image region while preserving external areas. However, current inpainting methods still suffer from the generation misalignment with facial attributes description and the loss of facial skin details. To address these challenges, (i) a novel data utilization strategy is introduced to construct datasets consisting of attribute-text-image triples from a data-driven perspective, (ii) a Causality-Aware Condition Adapter is proposed to enhance the contextual causality modeling of specific details, which encodes the skin details from the original image while preventing conflicts between these cues and textual conditions. In addition, a Skin Transition Frequency Guidance technique is introduced for the local modeling of contextual causality via sampling guidance driven by low-frequency alignment. Extensive quantitative and qualitative experiments demonstrate the effectiveness of our method in boosting both fidelity and editability for localized attribute editing. Our codes will be made publicly available. Xiaole Xian, Xilin He, Zenghao Niu, Junliang Zhang, Weicheng Xie 0001, Siyang Song, Zitong Yu, LinLin Shen |
AAAI | 6 |
| 2025 | PerReactor: Offline Personalised Multiple Appropriate Facial Reaction GenerationabstractIn dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strategies lead the training of previous automatic facial reaction generation (FRG) models to not only suffer from the 'one-to-many mapping' problem, but also fail to comprehensively consider the quality of the generated facial reactions. Besides, none of them considered such personalised behaviour patterns in generating facial reactions. In this paper, we propose the first adversarial FRG model training strategy which jointly learns appropriateness and realism discriminators to provide comprehensive task-specific supervision for training the target facial reaction generators, and reformulates the 'one-to-many (facial reactions) mapping' training problem as a 'one-to-one (distribution) mapping' training task, i.e., the FRG model is trained to output a distribution representing multiple appropriate/plausible facial reaction from each input human behaviour. In addition, our approach also serves as the first offline FRG approach that considers personalised behaviour patterns in generating of target individuals' facial reactions. Experiments show that our PerReactor not only largely outperformed all existing offline solutions for generating more appropriate, diverse and realistic facial reactions, but also is the first approach that can effectively generate personalised appropriate facial reactions. Hengde Zhu, Xiangyu Kong 0001, Weicheng Xie 0001, Xilin He, Lu Liu 0001, LinLin Shen, Wei Zhang 0243, Hatice Gunes, Siyang Song |
AAAI | 10 |
| 2025 | D3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head SynthesisabstractA key challenge in 3D talking head synthesis lies in the reliance on a long-duration talking head video to train a new model for each target identity from scratch. Recent methods have attempted to address this issue by extracting general features from audio through pre-training models. However, since audio contains information irrelevant to lip motion, existing approaches typically struggle to map the given audio to realistic lip behaviors in the target face when trained on only a few frames, causing poor lip synchronization and talking head image quality. This paper proposes D3-Talker, a novel approach that constructs a static 3D Gaussian attribute field and employs audio and Facial Motion signals to independently control two distinct Gaussian attribute deformation fields, effectively decoupling the predictions of general and personalized deformations. We design a novel similarity contrastive loss function during pre-training to achieve more thorough decoupling. Furthermore, we integrate a Coarse-to-Fine module to refine the rendered images, alleviating blurriness caused by head movements and enhancing overall image quality. Extensive experiments demonstrate that D3-Talker outperforms state-of-the-art methods in both high-fidelity rendering and accurate audio-lip synchronization with limited training data. Kaijun Deng, Siyang Song, Jindong Xie, Wenhui Ma, LinLin Shen |
ECAI | 3 |
| 2025 | MERMAID: Multi-perspective Self-reflective Agents with Generative Augmentation for Emotion Recognitionabstract… "Amusement" "Contentment" "Anger" "Excitement" "Fear" "Amusement" "Sadness" " F e a r " " S a d n e s s " MERMAID Figure 1: Demonstration of MERMAID.Given natural or facial images depicting various emotions, MERMAID produces precise emotion classifications by integrating multimodal self-reflection, generative augmentation to amplify and enrich subtle emotional cues, and cross-modal verification. Zhongyu Yang, Junhao Song 0001, Siyang Song, Wei Pang 0001, Yingfang Yuan |
EMNLP | 3 |
| 2025 | CauSkelNet: Causal Representation Learning for Human Behaviour AnalysisabstractTraditional machine learning methods for movement recognition often struggle with limited model interpretability and a lack of insight into human movement dynamics. This study introduces a novel representation learning framework based on causal inference to address these challenges. Our twostage approach combines the Peter-Clark (PC) algorithm and Kullback-Leibler (KL) divergence to identify and quantify causal relationships between human joints. By capturing joint interactions, the proposed causal Graph Convolutional Network (GCN) produces interpretable and robust representations. Experimental results on the EmoPain dataset demonstrate that the causal GCN outperforms traditional GCNs in accuracy, F1 score, and recall, particularly in detecting protective behaviors. This work contributes to advancing human motion analysis and lays a foundation for adaptive and intelligent healthcare solutions. Xingrui Gu, Chuyi Jiang, Erte Wang, Zekun Wu 0003, Leimin Tian, Lianlong Wu, Siyang Song, Chuang Yu 0001 |
FG | 8 |
| 2025 | DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face SynthesisabstractAccurately synthesizing talking face videos and capturing fine facial features for individuals with long hair presents a significant challenge. To tackle these challenges in existing methods, we propose a decomposed per-embedding Gaussian fields (DEGSTalk), a 3D Gaussian Splatting (3DGS)-based talking face synthesis method for generating realistic talking faces with long hairs. Our DEGSTalk employs Deformable Pre-Embedding Gaussian Fields, which dynamically adjust pre-embedding Gaussian primitives using implicit expression coefficients. This enables precise capture of dynamic facial regions and subtle expressions. Additionally, we propose a Dynamic Hair-Preserving Portrait Rendering technique to enhance the realism of long hair motions in the synthesized videos. Results show that DEGSTalk achieves improved realism and synthesis quality compared to existing approaches, particularly in handling complex facial dynamics and hair preservation. Our code is available at https://github.com/CVI-SZU/DEGSTalk. Kaijun Deng, Dezhi Zheng, Jindong Xie, Jinbao Wang 0001, Weicheng Xie 0001, LinLin Shen, Siyang Song |
ICASSP | 7 |
| 2025 | M3ADD: A Novel Benchmark for Physiology Signal-based Automatic Depression Detection with Multimodal Multitask Multievent FrameworkabstractThe prevalence of depression is escalating, especially among youth, which has become a critical mental health concern. Current assessment methods, relying heavily on questionnaires, clinical observations, and AI-driven analyses, are limited by their focus on single-event data, failing to encapsulate the nuanced expressions of depressive symptoms. Moreover, a significant oversight in existing research is the underutilization of electromyogram (EMG) alongside electroencephalogram (EEG) data, which could provide a more holistic view of unconscious body behaviors. Additionally, given the high variability of depression among individuals, traditional analysis models are in urgent need of refinement to accommodate the personality of different individuals. To address these limitations, we propose M3ADD, a novel benchmark for Automatic Depression Detection that employs a Multimodal, Multitask, and Multievent framework. We collected EEG and EMG data from 97 participants across varied events (interview, reading tasks, walking), coupled with standardized questionnaires assessing depression, wellbeing, and personality, enriching our multitask learning approach. Our benchmark recognition algorithm leverages multitask learning, channel and interactive attention mechanisms to synthesize event-specific and modal-specific features, enhancing adaptability to individual differences and improving data utilization efficiency. M3ADD surpasses existing models by achieving 87% accuracy in detecting depression and 95% accuracy in assessing wellbeing, providing a promising avenue for early identification. Changzeng Fu, Kaifeng Su, Yikai Su, Fengkui Qian, Siyang Song, Le Yang 0004, Xiaoyong Lv, Yuliang Zhao |
ICASSP | 7 |
| 2025 | A Frequency-aware Augmentation Network for Mental Disorders Assessment from AudioabstractDepression and Attention Deficit Hyperactivity Disorder (ADHD) stand out as the common mental health challenges today. In affective computing, speech signals serve as effective biomarkers for mental disorder assessment. Current research, relying on labor-intensive hand-crafted features or simplistic time-frequency representations, often overlooks critical details by not accounting for the differential impacts of various frequency bands and temporal fluctuations. Therefore, we propose a frequency-aware augmentation network with dynamic convolution for depression and ADHD assessment. In the proposed method, the spectrogram is used as the input feature and adopts a multi-scale convolution to help the network focus on discriminative frequency bands related to mental disorders. A dynamic convolution is also designed to aggregate multiple convolution kernels dynamically based upon their attentions which are input-independent to capture dynamic information. Finally, a feature augmentation block is proposed to enhance the feature representation ability and make full use of the captured information. Experimental results on AVEC 2014 and self-recorded ADHD dataset prove the robustness of our method, an RMSE of 9.23 was attained for estimating depression severity, along with an accuracy of 89.8% in detecting ADHD. Shuanglin Li, Siyang Song, Rajesh Nair, Syed M. Naqvi |
ICASSP | 2 |
| 2025 | Hierarchical Multimodal Decoupling-Fusion Framework for offline Multiple Appropriate Facial Reaction GenerationabstractFacial reactions convey crucial emotional information and coordinating interpersonal relationships in human dyadic interactions. While existing Multiple Appropriate Facial Reaction Generation (MAFRG) methods focus on generating multiple reasonable facial reactions, none of these approaches combines 2D and 3D facial behaviour information nor account for the influence of individuals’ facial identities, leading to inconsistencies in the generated facial reactions and limited capability in capturing subtle variations in facial depth and expression dynamics. This paper proposes a novel Hierarchical Multimodal Decoupling-Fusion (HMDF) framework that decouples 3D facial identity from expression behaviors, eliminating identity-based interference in the reaction generation process, which are integrated with audio-visual features through a cross-attention mechanism. Experiments show that our framework achieved the enhanced diversity and synchrony in the generated facial reactions. Qincheng Lv, Xiaofeng Liu 0006, Jie Li 0009, Pujun Xue, Siyang Song |
ICASSP | 6 |
| 2025 | SynFER: Towards Boosting Facial Expression Recognition With Synthetic DataabstractFacial expression datasets remain limited in scale due to the subjectivity of annotations and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foundation models, that rely on large-scale data for optimal performance. To tackle the overarching and complex challenge, instead of introducing a new large-scale dataset, we introduce SynFER (Synthesis of Facial Expressions with Refined Control), a novel synthetic framework for synthesizing facial expression image data based on high-level textual descriptions as well as more fine-grained and precise control through facial action units. To ensure the quality and reliability of the synthetic data, we propose a semantic guidance technique to steer the generation process and a pseudo-label generator to help rectify the facial expression labels for the synthetic images. To demonstrate the generation fidelity and the effectiveness of the synthetic data from SynFER, we conduct extensive experiments on representation learning using both synthetic data and real-world data. Results validate the efficacy of our approach and the synthetic data. Notably, our approach achieves a 67.23% classification accuracy on AffectNet when training solely with synthetic data equivalent to the AffectNet training set size, which increases to 69.84% when scaling up to five times the original size. Code is available here. Xilin He, Xiaole Xian, Bing Li 0024, Muhammad Haris Khan, ZongYuan Ge, Weicheng Xie 0001, Siyang Song, LinLin Shen, Bernard Ghanem, Xiangyu Yue 0001 |
ICCV | 8 |
| 2025 | Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided SamplingabstractAdversarial attacks present a critical challenge to deep neural networks' robustness, particularly in transfer scenarios across different model architectures. However, the transferability of adversarial attacks faces a fundamental dilemma between Exploitation (maximizing attack potency) and Exploration (enhancing cross-model generalization). Traditional momentum-based methods over-prioritize Exploitation, i.e., higher loss maxima for attack potency but weakened generalization (narrow loss surface). Conversely, recent methods with inner-iteration sampling over-prioritize Exploration, i.e., flatter loss surfaces for cross-model generalization but weakened attack potency (suboptimal local maxima). To resolve this dilemma, we propose a simple yet effective Gradient-Guided Sampling (GGS), which harmonizes both objectives through guiding sampling along the gradient ascent direction to improve both sampling efficiency and stability. Specifically, based on MI-FGSM, GGS introduces inner-iteration random sampling and guides the sampling direction using the gradient from the previous inner-iteration (the sampling's magnitude is determined by a random distribution). This mechanism encourages adversarial examples to reside in balanced regions with both flatness for cross-model generalization and higher local maxima for strong attack potency. Comprehensive experiments across multiple DNN architectures and multimodal large language models (MLLMs) demonstrate the superiority of our method over state-of-the-art transfer attacks. Code is made available at https://github.com/anuin-cat/GGS. Zenghao Niu, Weicheng Xie 0001, Siyang Song, Zitong Yu, Feng Liu 0013, LinLin Shen |
ICCV | 3 |
| 2025 | Learning from Human Conversations: A Seq2Seq based Multi-modal Robot Facial Expression Reaction Framework in HRIabstractNonverbal communication plays a crucial role in both human-human and human-robot interactions (HRIs), where facial expressions convey emotions, intentions and trust. Enabling humanoid robots to generate human-like facial reactions in response to human speech and facial behaviours remains significant challenges. In this work, we leverage human-human interaction (HHI) datasets to train a humanoid robot, allowing it to learn and imitate facial reactions to both speech and facial expression inputs. Specifically, we extend a sequence-to-sequence (Seq2Seq)-based framework that enables robots to simulate human-like virtual facial expressions that are appropriate for responding to the perceived human user behaviours. Then, we propose a deep neural network-based motor mapping model to translate these expressions into physical robot movements. Experiments demonstrate that our facial reaction–motor mapping framework successfully enables robotic self-reactions to various human behaviours, where our model can best predict 50 frames (two seconds) of facial reactions in response to the input user behaviour of the same duration, aligning with human cognitive and neuromuscular processes. Our code is provided at https://github.com/mrsgzg/Robot_Face_Reaction. Zhegong Shangguan, Xiaoxuan Hei, Fangjun Li, Chuang Yu 0001, Siyang Song, Jianzhuang Zhao, Angelo Cangelosi, Adriana Tapus |
IROS | 5 |
| 2025 | Smooth Online Multiple Appropriate Facial Reaction GenerationabstractIn dyadic interactions, facial reactions are crucial for conveying an individuals' responses to their conversational partners. Individuals may exhibit varied but appropriate facial reactions (AFRs) when perceiving the same behavioral expression. Although some recent methods can already respond multiple appropriate facial reactions to the given human speaker behaviors, the AFRs generated by these methods often fail to adequately preserve crucial head motions, leading to visual jitter and unnatural transitions between generated AFR segments. In this paper, we propose a novel and generic PFLPosNet framework which addresses the aforementioned problems at both pre-processing and post-processing stages, where a new pose-aware face behavior localization method PFL is introduced to retain the head pose displacement information from the source data. In addition, the framework proposes a real-time head pose adjustment method, PosNet, to ensure continuity and smoothness in the visual output of the model when using data with correct head pose displacement. Experimental results demonstrate that our approach not only generates more coherent and natural facial reaction sequences but also significantly outperforms existing online MAFRG methods in terms of continuity and smoothness. Our code is made available at https://github.com/rainforcetime/PFLPosNet. Weicheng Xie 0001, Chunlin Yan, Siyang Song, Zitong Yu, LinLin Shen, Laizhong Cui |
ACM Multimedia | 3 |
| 2025 | The First MPDD Challenge: Multimodal Personality-aware Depression DetectionabstractDepression is a widespread mental health issue affecting diverse age groups, with notable prevalence among college students and the elderly. However, existing datasets and detection methods primarily focus on young adults, neglecting the broader age spectrum and individual differences that influence depression manifestation. Current approaches often establish a direct mapping between multimodal data and depression indicators, failing to capture the complexity and diversity of depression across individuals. This challenge includes two tracks based on age-specific subsets: Track 1 uses the MPDD-Elderly dataset for detecting depression in older adults, and Track 2 uses the MPDD-Young dataset for detecting depression in younger participants. The Multimodal Personality-aware Depression Detection (MPDD) Challenge aims to address this gap by incorporating multimodal data alongside individual difference factors. We provide a baseline model that fuses audio and video modalities with individual difference information to detect depression manifestations in diverse populations. This challenge aims to promote the development of more personalized and accurate de pression detection methods, advancing mental health research and fostering inclusive detection systems. More details are available on the official challenge website: https://hacilab.github.io/MPDDChallenge.github.io. Changzeng Fu, Zelin Fu, Qi Zhang 0124, Xinhe Kuang, Jiacheng Dong, Kaifeng Su, Yikai Su, Junfeng Yao, Yuliang Zhao, Shiqi Zhao 0001, Siyang Song, Yuichiro Yoshikawa, Björn W. Schuller, Hiroshi Ishiguro |
ACM Multimedia | 13 |
| 2025 | ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion ModelabstractThe automatic generation of diverse and human-like facial reactions in dyadic dialogue remains a critical challenge for human-computer interaction systems. Existing methods fail to model the stochasticity and dynamics inherent in real human reactions. To address this, we propose ReactDiff, a novel temporal diffusion framework for generating diverse facial reactions that are appropriate for responding to any given dialogue context. Our key insight is that plausible human reactions demonstrate smoothness, and coherence over time, and conform to constraints imposed by human facial anatomy. To achieve this, ReactDiff incorporates two vital priors (spatio-temporal facial kinematics) into the diffusion process: i) temporal facial behavioral kinematics and ii) facial action unit dependencies. These two constraints guide the model toward realistic human reaction manifolds, avoiding visually unrealistic jitters, unstable transitions, unnatural expressions, and other artifacts. Extensive experiments on the REACT2024 dataset demonstrate that our approach not only achieves state-of-the-art reaction quality but also excels in diversity and reaction appropriateness. Our code is publicly available at https://github.com/lingjivoo/ReactDiff. Siyang Song, Siyuan Yan, ZongYuan Ge |
ACM Multimedia | 2 |
| 2025 | REACT 2025: the Third Multiple Appropriate Facial Reaction Generation ChallengeabstractIn dyadic interactions, a broad spectrum of human facial reactions might be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023 and REACT 2024 challenges, we are proposing the REACT 2025 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can be used to generate multiple appropriate, diverse, realistic and synchronised human-style facial reactions expressed by human listeners in response to an input stimulus (i.e., audio-visual behaviours expressed by their corresponding speakers). As a key of the challenge, we provide challenge participants with the first natural and large-scale multi-modal Multiple Appropriate Facial Reaction Generation (MAFRG) dataset (called MARS) recording 136 human-human dyadic interactions containing a total of 2856 interaction sessions covering five different topics. In addition, this paper also presents the challenge guidelines and the performance of our baselines on the two proposed sub-challenges: Offline MAFRG and Online MAFRG, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2025 Siyang Song, Micol Spitale, Xiangyu Kong 0001, Hengde Zhu, Cristina Palmero, Germán Barquero, Sergio Escalera, Michel F. Valstar, Mohamed Daoudi, Tobias Baur 0001, Fabien Ringeval, Andrew Howes 0001, Elisabeth André, Hatice Gunes |
ACM Multimedia | 1 |
| 2025 | OmniResponse: Online Multimodal Conversational Response Generation in Dyadic InteractionsabstractIn this paper, we introduce Online Multimodal Conversational Response Generation (OMCRG), a novel task designed to produce synchronized verbal and non-verbal listener feedback online, based on the speaker's multimodal inputs. OMCRG captures natural dyadic interactions and introduces new challenges in aligning generated audio with listeners' facial responses. To tackle these challenges, we incorporate text as an intermediate modality to connect audio and facial responses. We propose OmniResponse, a Multimodal Large Language Model (MLLM) that autoregressively generates accurate multimodal listener responses. OmniResponse leverages a pretrained LLM enhanced with two core components: Chrono-Text Markup, which precisely timestamps generated text tokens, and TempoVoice, a controllable online text-to-speech (TTS) module that outputs speech synchronized with facial responses. To advance OMCRG research, we offer ResponseNet, a dataset of 696 detailed dyadic interactions featuring synchronized split-screen videos, multichannel audio, transcripts, and annotated facial behaviors. Comprehensive evaluations on ResponseNet demonstrate that OmniResponse outperforms baseline models in terms of semantic speech content, audio-visual synchronization, and generation quality. Our dataset, code, and models are publicly available at https://omniresponse.github.io/. Jianghui Wang, Bing Li 0024, Siyang Song, Bernard Ghanem |
NeurIPS | 4 |
| 2025 | Towards Robust Training via Gradient-Diversified BackpropagationabstractNeural networks are prone to be vulnerable to adversarial attacks and domain shifts. Adversarial-driven methods including adversarial training and adversarial augmentation, have been frequently proposed to improve the model's robustness against adversarial attacks and distribution-shifted samples. Nonetheless, recent research on adversarial attacks has cast a spotlight on the robustness lacuna against attacks targeted at deep semantic layers. Our analysis reveals that previous adversarial-driven methods tend to generate overpowering perturbations in deep semantic layers, leading to distortion of the training for these layers. This can be primarily attributed to the exclusive utilization of loss functions on the output layer for adversarial gradient generation. This inherent practice projects an excessive adversarial impact on the deep semantic layers, elevating the difficulty of training such layers. Therefore, from the standing point of relaxing the excessive perturbations in the deep semantic layer and diversifying the adversarial gradients to ensure robust training for deep semantic layers, this paper proposes a novel Stochastic Loss Integration Method (SLIM), which can be instantiated into the existing adversarial-driven methods in a plug-and-play manner. Experimental results across diverse tasks, including classification and segmentation, as well as various areas such as adversarial robustness and domain generalization, validate the effectiveness of our proposed method. Furthermore, we provide an in-depth analysis to offer a comprehensive understanding of layer-wise training involving various loss terms. Xilin He, Qinliang Lin, Weicheng Xie 0001, Muhammad Haris Khan, Siyang Song, LinLin Shen |
WACV | 6 |
| 2025 | Facial action units guided graph representation learning for multimodal depression detection
Changzeng Fu, Fengkui Qian, Yikai Su, Kaifeng Su, Siyang Song, Mingyue Niu, Zhigang Liu 0014, Carlos Toshinori Ishi, Hiroshi Ishiguro |
Neurocomputing | 5 |
| 2025 | SymGraphAU: Prior knowledge based symbolic graph for action unit recognition
Weicheng Xie 0001, Junliang Zhang, Siyang Song, LinLin Shen, Zitong Yu |
Pattern Recognit. | 4 |
| 2025 | A Multi-Scale Feature Refinement and Dual-Attention Enhanced Dynamic Convolutional Network for Speech-Based Depression and ADHD AssessmentabstractIn the area of affective computing, speech has been identified as a promising biomarker for assessing depression and attention deficit hyperactivity disorder (ADHD). These disorders manifest as abnormalities in speech across various frequency bands and exhibit temporal variations. Most existing work on speech features relies on the magnitude spectrogram, which discards phase information and also does not consider the impact of different frequency bands on depression and ADHD detection. Inspired by these, we propose a novel multi-scale complex feature refinement and dynamic convolution attention-aware network to enhance speech-based assessment of depression and ADHD. Our approach incorporates three key components: multi-scale complex feature refinement (MSFR), dynamic convolutional neural network (Dy-CNN), and dual-attention feature enhancement (DAFE) module. The MSFR module utilizes depth-wise convolutional networks to process both magnitude and phase input, selectively emphasizing frequency bands associated with depression and ADHD. Importantly, the Dy-CNN module employs an attention mechanism to autonomously generate multiple convolution kernels that adapt to input features and capture relevant temporal dynamics linked to depression and ADHD. Additionally, the DAFE module enhances feature representation and detection performance by incorporating channel shuffle attention (CSA) and spatial axial attention (SAA) mechanisms, which leverage both inter- and intra-channel relationships and examine time-frequency characteristics of the feature map. Extensive experiments conducted on four publicly available datasets, i.e., AVEC2013, AVEC2014, E-DAIC, and a self-collected authentic ADHD dataset demonstrated that the proposed method outperforms previous approaches and exhibits superior generalization capabilities across different language settings (i.e., English, German) for speech-based depression and ADHD assessment. Shuanglin Li, Siyang Song, Syed M. Naqvi |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Two-Stage Temporal Modelling Framework for Video-Based Depression Recognition Using Graph RepresentationabstractVideo-based automatic depression analysis provides a fast, objective and repeatable self-assessment solution, which has been widely developed in recent years. While depression cues may be reflected by human facial behaviours of various temporal scales, most existing approaches either focused on modelling depression from short-term or video-level facial behaviours. In this sense, we propose a two-stage framework that models depression severity from multi-scale short-term and video-level facial behaviours. The short-term depressive behaviour modelling stage first deep learns depression-related facial behavioural features from multiple short temporal scales, where a Depression Feature Enhancement (DFE) module is proposed to enhance the depression-related cues for all temporal scales and remove non-depression related noise. Two novel graph encoding strategies are proposed in the video-level depressive behavior modeling stage, i.e., Sequential Graph Representation (SEG) and Spectral Graph Representation (SPG), to re-encode all short-term features of the target video into a video-level graph representation, summarizing depression-related multi-scale video-level temporal information. As a result, the produced graph representations predict depression severity using both short-term and long-term facial behaviour patterns. The experimental results on AVEC 2013, AVEC 2014 and AVEC 2019 datasets show that the proposed DFE module constantly enhanced the depression severity estimation performance for various CNN models while the SPG is superior than other video-level modelling methods. More importantly, the result achieved for the proposed two-stage framework shows its promising and solid performance compared to widely-used one-stage modelling approaches. Hatice Gunes, Keerthy Kusumam, Michel F. Valstar, Siyang Song |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | Multimodal Depression Assessment Framework Integrating Personality and Gait for Older Adults With Medical ConditionsabstractElderly individuals often suffer from underlying medical conditions, resulting in a significant decline in quality of life and a heightened susceptibility to depression. Presently, AI screening tools based on behavioral indicators offer an objective and effective approach to diagnosing depression. However, current AI depression screening tools are primarily tailored to adolescents and adults, exhibiting shortcomings in their applicability and accuracy for elderly individuals with underlying medical conditions. To address the above issues, first, this paper constructs a depression dataset for elderly people with underlying diseases by using semi-structured interviews. Second, based on cognitive science insights, it is recognized that personality factors significantly influence behavioral expressions and also determine the attitudes of elderly individuals toward current life circumstances/health issues. Therefore, besides annotating depression severity, the Big Five-10 personality scale was utilized to annotate participant personalities. Finally, a late fusion-based multi-task learning framework was proposed, and the effects of introducing gait information and personality annotation on the performance of depression assessment were investigated. The experimental findings affirm the importance of integrating gait information and personality assessment in improving depression detection effectiveness. This study provides valuable foundational resources, as well as beneficial references and insights, for the research on depression in the elderly. Yuliang Zhao, Jian Li 0063, Siyang Song, Chao Lian, Yinghao Liu, Changzeng Fu |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Unleashing Fourier-Domain Potential: Spatial-Spectral Reconstruction Framework for Remote Sensing PansharpeningabstractPansharpening aims to generate high-resolution multispectral (HR-MS) images by fusing the corresponding low-resolution multispectral (LR-MS) and high-resolution panchromatic (PAN) images. While the spatial-spectral properties modeling plays crucial roles in generating high-quality HR-MS images, existing approaches suffer from: 1) directly modeling the entangled spatial-spectral properties; and 2) lacking task-specific priors for spatial and spectral properties modeling. This paper proposes a novel Fourier domain-based approach (Fourier-SSR) to address these problems, where phase and amplitude components of PAN and LR-MS are considered to individually model spatial and spectral properties for the target HR-MS image. Our Fourier-SSR is motivated by the crucial findings in Fourier domain: (i) manifestation of spatial-spectral properties, i.e., the spatial and spectral properties of remote sensing images can be individually manifested in their phase and amplitude components in Fourier domain; (ii) spatial property-related prior, i.e., only reconstructing the phase component of PAN image in Fourier domain, could generate spatial property required for target HR-MS image; and (iii) spectral property-related prior, i.e., jointly models the amplitude components of PAN and LR-MS images could generate the required spectral property for target HR-MS image. Based on the aforementioned findings, we design aFourier-guided spatial Mixerand aFourier-guided spectral Mixer, which innovatively employ complex feature interaction strategies to individually reconstruct the phase and amplitude components for target HR-MS. Experiments show that our methods unleash Fourier domain potential in individually modeling spatial and spectral properties for the target HR-MS image, leading to superior performance over previous state-of-the-art. Our code is provided in Supplementary Material. Mengting Ma, Yizhen Jiang, Mengjiao Zhao, Wei Zhang 0243, Siyang Song |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | LOGCAN++: Adaptive Local-Global Class-Aware Network for Semantic Segmentation of Remote Sensing ImagesabstractRemote sensing images are usually characterized by complex backgrounds, scale and orientation variations, and large intraclass variance. General semantic segmentation methods usually fail to fully investigate the above issues, and thus their performances on remote sensing image segmentation are limited. In this article, we propose our LOGCAN++, a semantic segmentation model customized for remote sensing images, which is made up of a global class-aware (GCA) module and several local class-aware (LCA) modules. The GCA module captures global representations for class-level context modeling to reduce the interference of background noise. The LCA module generates local class representations as intermediate perceptual elements to indirectly associate pixels with the global class representations, targeting dealing with the large intraclass variance problem. In particular, we introduce affine transformations in the LCA module for adaptive extraction of local class representations to effectively tolerate scale and orientation variations in remote sensing images. Extensive experiments on three benchmark datasets show that our LOGCAN++ outperforms current mainstream general and remote sensing semantic segmentation methods and achieves a better trade-off between speed and accuracy. Rongrong Lian, Zhenkai Wu, Fan Yang 0100, Mengting Ma, Sensen Wu, Zhenhong Du, Wei Zhang 0243, Siyang Song |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2025 | Supervised Detail-Guided Multiscale State-Space Model for Pan-SharpeningabstractPan-sharpening reconstructs the high-resolution multispectral (HR-MS) image from its corresponding panchromatic (PAN) image and low-resolution multispectral (LR-MS) image. However, existing deep learning (DL)-based pan-sharpening methods typically suffer from three challenges: 1) the vanilla LR-MS image upsampling employed by them fails to consider domain knowledge, thereby disregarding crucial information; 2) while remote sensing images exhibit multiscale complex land features, existing methods fail to fully exploit crucial multiscale spatial information, that is, scale transformation layers in their models are not effective; and 3) existing convolutional neural network (CNN) and transformer-based pan-sharpening backbones are constrained by inherent local receptive fields or quadratic computational complexity, making them difficult to balance their effectiveness and efficiency. To address these issues, we propose a novel supervised detail-guided multiscale state-space model for pan-sharpening, namely SDMSPan. Our SDMSPan consists of three residual state-space modules (Res-SSMs) that are responsible for handling image information at three spatial scales, where each Res-SSM aims to model both local and long-range dependencies between PAN and LR-MS images at a specific spatial scale with lower computational cost. Between each pair of Res-SSMs, a novel detail-guided upsampling block (DGUB) is proposed to apply spatial details of the PAN image to guide effective and task-aware intermediate feature upsampling, where a novel multiscale intermediate spatial-spectral supervision strategy is also proposed to supervise the training of every DGUB. Experimental results demonstrate that our proposed approach significantly outperforms other state-of-the-art methods in performance. Our code is provided athttps://github.com/zhaomengjiao123/SDMSPan. Mengjiao Zhao, Mengting Ma, Wei Zhang 0243, Siyang Song |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Frequency Restoration and Modality Enforcement towards Resisting-corruption Multimodal Sentiment AnalysisabstractFor Multimodal Sentiment Analysis (MSA), previous methods concentrate on designing sophisticated fusion strategies and performing representation learning across heterogeneous modalities, aiming to leverage multimodal signals to detect human sentiment. However, these approaches fail to address the long-standing issue of corrupted modal details in videos, which may be caused by the challenge of the excessive loss of emotionally relevant semantics resulted from the degradation of detailed information. In this work, we aim to improve the robustness capacity of resisting corruption in MSA, by introducing a Hierarchical Frequency Restoration and Adaptive Modality Enforcement (HFR-AME) approach. The HFR-AME progressively recovers blurred detailed cues in each modality while enhancing the discriminative power of modal representations. Specifically, to reconstruct distinct frequency band features, we propose to equip the HFR module with a key component called the Frequency Multimodal UNet (FM-UNet), so as to utilize complementary modal features as conditions. This meticulous restoration process, performed from low to high frequency, facilitates the comprehensive recovery of intricate details. Meanwhile, to adaptively integrate these diverse frequency features, we introduce the AME module to enhance the beneficial modal frequencies while suppressing irrelevant ones, with the goal of strengthening the restored modal representations. Extensive experiments show our HFR-AME outperforms state-of-the-art methods on the CMU-MOSI and CMU-MOSEI datasets, improving 7-class accuracy by 0.5% and 0.6%, respectively. Further analysis also confirms its cross-lingual generalization and competitive computational efficiency. Our code is made available at https://github.com/nianhua20/HFR-AME . Weicheng Xie 0001, Haijian Liang, Zenghao Niu, Xianxu Hou, Siyang Song, Zitong Yu, LinLin Shen |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | ReactFace: Online Multiple Appropriate Facial Reaction Generation in Dyadic InteractionsabstractIn dyadic interaction, predicting the listener's facial reactions is challenging as different reactions could be appropriate in response to the same speaker's behaviour. Previous approaches predominantly treated this task as an interpolation or fitting problem, emphasizing deterministic outcomes but ignoring the diversity and uncertainty of human facial reactions. Furthermore, these methods often failed to model short-range and long-range dependencies within the interaction context, leading to issues in the synchrony and appropriateness of the generated facial reactions. To address these limitations, this paper reformulates the task as an extrapolation or prediction problem, and proposes an novel framework (called ReactFace) to generate multiple different but appropriate facial reactions from a speaker behaviour rather than merely replicating the corresponding listener facial behaviours. Our ReactFace generates multiple different but appropriate photo-realistic human facial reactions by: (i) learning an appropriate facial reaction distribution representing multiple different but appropriate facial reactions; and (ii) synchronizing the generated facial reactions with the speaker verbal and non-verbal behaviours at each time stamp, resulting in realistic 2D facial reaction sequences. Experimental results demonstrate the effectiveness of our approach in generating multiple diverse, synchronized, and appropriate facial reactions from each speaker's behaviour. The quality of the generated facial reactions is intimately tied to the speaker's speech and facial expressions, achieved through our novel speaker-listener interaction modules. Siyang Song, Weicheng Xie 0001, Micol Spitale, ZongYuan Ge, LinLin Shen, Hatice Gunes |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Boosting Adversarial Transferability across Model Genus by Deformation-Constrained WarpingabstractAdversarial examples generated by a surrogate model typically exhibit limited transferability to unknown target systems. To address this problem, many transferability enhancement approaches (e.g., input transformation and model augmentation) have been proposed. However, they show poor performances in attacking systems having different model genera from the surrogate model. In this paper, we propose a novel and generic attacking strategy, called Deformation-Constrained Warping Attack (DeCoWA), that can be effectively applied to cross model genus attack. Specifically, DeCoWA firstly augments input examples via an elastic deformation, namely Deformation-Constrained Warping (DeCoW), to obtain rich local details of the augmented input. To avoid severe distortion of global semantics led by random deformation, DeCoW further constrains the strength and direction of the warping transformation by a novel adaptive control strategy. Extensive experiments demonstrate that the transferable examples crafted by our DeCoWA on CNN surrogates can significantly hinder the performance of Transformers (and vice versa) on various tasks, including image classification, video action recognition, and audio recognition. Code is made available at https://github.com/LinQinLiang/DeCoWA. Qinliang Lin, Zenghao Niu, Xilin He, Weicheng Xie 0001, Yuanbo Hou, LinLin Shen, Siyang Song |
AAAI | 8 |
| 2024 | Deep Unfolding Network with Spatial-spectral Perception Enhanced for Pan-sharpening
Mengjiao Zhao, Mengting Ma, Ao Gao, Siyang Song, Wei Zhang 0243 |
BMVC | 5 |
| 2024 | Multi-Scale Dynamic and Hierarchical Relationship Modeling for Facial Action Units RecognitionabstractHuman facial action units (AUs) are mutually related in a hierarchical manner, as not only they are associated with each other in both spatial and temporal domains but also AUs located in the same/close facial regions show stronger relationships than those of different facial regions. While none of existing approach thoroughly model such hi-erarchical inter-dependencies among AUs, this paper proposes to comprehensively model multi-scale AU-related dynamic and hierarchical spatiotemporal relationship among AUs for their occurrences recognition. Specifically, we first propose a novel multi-scale temporal differencing network with an adaptive weighting block to explicitly capture facial dynamics across frames at different spatial scales, which specifically considers the heterogeneity of range and mag-nitude in different AUs' activation. Then, a two-stage strategy is introduced to hierarchically model the relationship among AUs based on their spatial distribution (i.e., local and cross-region AU relationship modelling). Experimental results achieved on BP4D and DISFA show that our approach is the new state-of-the-art in the field of AU occurrence recognition. Our code is publicly available at https://github.com/CVI-SZU/MDHR. Zihan Wang 0005, Siyang Song, Songhe Deng, Weicheng Xie 0001, LinLin Shen |
CVPR | 2 |
| 2024 | Advancing Saliency Ranking with Human Fixations: Dataset, Models and BenchmarksabstractSaliency ranking detection (SRD) has emerged as a challenging task in computer vision, aiming not only to identify salient objects within images but also to rank them based on their degree of saliency. Existing SRD datasets have been created primarily using mouse-trajectory data, which inadequately captures the intricacies of human visual perception. Addressing this gap, this paper introduces the first large-scale SRD dataset, SIFR, constructed using genuine human fixation data, thereby aligning more closely with real visual perceptual processes. To establish a baseline for this dataset, we propose QAGNet, a novel model that leverages salient instance query features from a transformer detector within a tri-tiered nested graph. Through extensive experiments, we demonstrate that our approach outperforms existing state-of-the-art methods across two widely used SRD datasets and our newly proposed dataset. Code and dataset are available at https://github.com/EricDengbowen/QAGNet. Bowen Deng 0006, Siyang Song, Andrew P. French, Denis Schluppeck, Michael P. Pound |
CVPR | 2 |
| 2024 | Domain Separation Graph Neural Networks for Saliency Object RankingabstractSaliency object ranking (SOR) has attracted significant attention recently. Previous methods usually failed to ex-plicitly explore the saliency degree-related relationships between objects. In this paper, we propose a novel Domain Separation Graph Neural Network (DSGNN), which starts with separately extracting the shape and texture cues from each object, and builds an shape graph as well as a texture graph for all objects in the given image. Then, we propose a Shape-Texture Graph Domain Separation (STGDS) module to separate the task-relevant and irrelevant information of target objects by explicitly modelling the relationship between each pair of objects in terms of their shapes and textures, respectively. Furthermore, a Cross Image Graph Domain Separation (CIGDS) module is introduced to explore the saliency degree subspace that is robust to different scenes, aiming to create a unified representation for targets with the same saliency levels in different images. Importantly, our DSGNN automatically learns a multi-dimensional feature to represent each graph edge, allowing complex, diverse and ranking-related relationships to be modelled. Experimental results show that our DS-GNN achieved the new state-of-the-art performance on both ASSR and IRSR datasets, with large improvements of 5.2% and 4.1% SA-SOR, respectively. Our code is provided in https://github.com/Wu-ZJ/DSGNN. Jun Lu 0006, Jing Han 0009, Lianfa Bai, Yi Zhang 0036, Siyang Song |
CVPR | 7 |
| 2024 | MTaDCS: Moving Trace and Feature Density-Based Confidence Sample Selection Under Label Noise
Qingzheng Huang, Xilin He, Xiaole Xian, Qinliang Lin, Weicheng Xie 0001, Siyang Song, LinLin Shen, Zitong Yu |
ECCV (71) | 6 |
| 2024 | Multi-modal Human Behaviour Graph Representation Learning for Automatic Depression AssessmentabstractAutomatic depression assessment (ADA) often relies on crucial cues embedded in human verbal and non-verbal behaviors, which exists in video, audio, and text modalities. Although these modalities often show in time-series forms, current research offers limited exploration of the complex intra-modal temporal dynamics inherent to each modality, failing to extract the depression-related cues in a global view. While many methodologies attempt to exploit the multifaceted information encoded across modalities via decision-level or feature-level fusion techniques, they often fall short in effectively representing pairwise inter-modal relationships, which is the key to utilize the distinct complementary relationship between each modality pair. This paper presents a novel graph-based multimodal fusion approach, which can model intra-modal and inter-modal dynamics conveniently using a graph representation. It adopts undirected edges to link not only temporally continuous, pre-extracted features of each modality, but also temporally aligned features across each pair of modalities. This ensures the seamless propagation of global information across temporal dimensions and helps capture the pairwise inter-modal dynamics. We conduct experiments on the E-DAIC dataset to prove our approach's effectiveness, with an RMSE of 4.80 and a CCC value of 0.563, which rival the top-performing method. We also experiment on the AFAR-BSFP dataset to show the generality of our approach. Our code will be made publicly available. Haotian Shen, Siyang Song, Hatice Gunes |
FG | 2 |
| 2024 | REACT 2024: the Second Multiple Appropriate Facial Reaction Generation ChallengeabstractIn dyadic interactions, humans communicate their intentions and state of mind using verbal and non-verbal cues, where multiple different facial reactions might be appropriate in response to a specific speaker behaviour. Then, how to develop a machine learning (ML) model that can automatically generate multiple appropriate, diverse, realistic and synchronised human facial reactions from an previously unseen speaker behaviour is a challenging task. Following the successful organisation of the first REACT challenge (REACT 2023), this edition of the challenge (REACT 2024) employs a subset used by the previous challenge, which contains segmented 30-secs dyadic interaction clips originally recorded as part of the NOXI and RECOLA datasets, encouraging participants to develop and benchmark Machine Learning (ML) models that can generate multiple appropriate facial reactions (including facial image sequences and their attributes) given an input conversational partner's stimulus under various dyadic video conference scenarios. This paper presents: (i) the guidelines of the REACT 2024 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2024. Siyang Song, Micol Spitale, Cristina Palmero, Germán Barquero, Hengde Zhu, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes |
FG | 1 |
| 2024 | Scale-Free And Task-Generic Attack: Generating Photo-Realistic Adversarial Patterns With Patch Quilting GeneratorabstractRecent CNN generator-based attack approaches can synthe-size unrestricted and semantically meaningful entities to the image, which are able to improve the transferability and robustness. However, such methods attack images by either synthesizing local adversarial entities, which are only suitable for attacking specific contents, or performing global attacks, which are only applicable to a specific image scale. In this paper, we propose a novel Patch Quilting Generative Adversarial Networks (PQ-GAN) to learn the first scale-free CNN generator that can be applied to attack images with arbitrary scales for various computer vision tasks. The principal investigation on transferability of the generated adversarial examples, robustness to defense frameworks, and visual quality assessment show that the proposed PQG-based attack framework outperforms the other nine state-of-the-art adversarial attack approaches when attacking the neural networks trained on two standard evaluation datasets (i.e., ImageNet and CityScapes). Our code is made available at https://github.com/XiangboGaoBarry/PQAttack. Xiangbo Gao, Qinliang Lin, Weicheng Xie 0001, LinLin Shen, Keerthy Kusumam, Siyang Song |
ICASSP | 7 |
| 2024 | Multi-Level Graph Learning For Audio Event Classification And Human-Perceived Annoyance Rating PredictionabstractWHO’s report on environmental noise estimates that 22 M people suffer from chronic annoyance related to noise caused by audio events (AEs) from various sources. Annoyance may lead to health issues and adverse effects on metabolic and cognitive systems. In cities, monitoring noise levels does not provide insights into noticeable AEs, let alone their relations to annoyance. To create annoyance-related monitoring, this paper proposes a graph-based model to identify AEs in a sound-scape, and explore relations between diverse AEs and human-perceived annoyance rating (AR). Specifically, this paper proposes a lightweight multi-level graph learning (MLGL) based on local and global semantic graphs to simultaneously perform audio event classification (AEC) and human annoyance rating prediction (ARP). Experiments show that: 1) MLGL with 4.1 M parameters improves AEC and ARP results by using semantic node information in local and global context-aware graphs; 2) MLGL captures relations between coarse-and fine-grained AEs and AR well; 3) Statistical analysis of MLGL results shows that some AEs from different sources significantly correlate with AR, which is consistent with previous research on human perception of these sound sources. Yuanbo Hou, Qiaoqiao Ren, Siyang Song, Wenwu Wang 0001, Dick Botteldooren |
ICASSP | 3 |
| 2024 | Circular Decomposition and Cross-Modal Recombination for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis is a burgeoning research area, leveraging various modalities to predict the sentiment score. Nevertheless, previous studies have disregarded the impact of noise interference on specific modal sentiments during video recording, thereby compromising the accuracy of sentiment prediction. In this paper, we propose the Guided Circular Decomposition and Cross-Modal Recombination (GCD-CMR) model, which aims to eliminate contaminated sentiment features in a fine-grained way. To achieve this, we utilize tailored global information specific to each modality to guide the circular decomposing process in the GCD module, to produce a set of sentiment prototypes. Subsequently, in the CMR module, we align cross-modal sentiment prototypes and remove the contaminated prototypes for recombination. Experimental results on two publicly available datasets demonstrate that our model surpasses state-of-the-art models, confirming the effectiveness of our proposed method. We release the code at: https://github.com/nianhua20/GCD-CMR. Haijian Liang, Weicheng Xie 0001, Xilin He, Siyang Song, LinLin Shen |
ICASSP | 4 |
| 2024 | MERG: Multi-Dimensional Edge Representation Generation Layer for Graph Neural NetworksabstractEdges are essential in describing relationships among nodes. While existing graphs frequently use a single-value edge to describe association between each pair of node vectors, crucial relationships may be disregarded if they are not linearly correlated, which may limit graph analysis performance. Although some recent Graph Neural Networks (GNNs) can process graphs containing multi-dimensional edge features, they cannot convert single-value edge graphs to multi-dimensional edge graphs during propagation. This paper proposes a generic Multi-dimensional Edge Representation Generation (MERG) layer that can be inserted into any GNNs for heterogeneous graph analysis. It assigns multi-dimensional edge features for the input single-value edge graph, describing multiple task-specific and global context-aware relationship cues between each connected node pair. Results on eight graph benchmark datasets demonstrate that inserting the MERG layer into widely-used GNNs (e.g., GatedGCN and GAT) leads to major performance improvements, resulting in state-of-the-art (SOTA) results on seven out of eight evaluated datasets. Our code is publicly available at1. YuXin Song 0001, Aaron S. Jackson, Xi Jia, Weicheng Xie 0001, LinLin Shen, Hatice Gunes, Siyang Song |
ICASSP | 8 |
| 2024 | Robust Facial Reactions Generation: An Emotion-Aware Framework with Modality CompensationabstractThe objective of the Multiple Appropriate Facial Reaction Generation (MAFRG) task is to produce contextually appropriate and diverse listener facial behavioural responses based on the multimodal behavioural data of the conversational partner (i.e., the speaker). Current methodologies typically assume continuous availability of speech and facial modality data, neglecting real-world scenarios where these data may be intermittently unavailable, which often results in model failures. Furthermore, despite utilising advanced deep learning models to extract information from the speaker’s multimodal inputs, these models fail to adequately leverage the speaker’s emotional context, which is vital for eliciting appropriate facial reactions from human listeners. To address these limitations, we propose an Emotion-aware Modality Compensatory (EMC) framework. This versatile solution can be seamlessly integrated into existing models, thereby preserving their advantages while significantly enhancing performance and robustness in scenarios with missing modalities. Our framework ensures resilience when faced with missing modality data through the Compensatory Modality Alignment (CMA) module. It also generates more appropriate emotion-aware reactions via the Emotion-aware Attention (EA) module, which incorporates the speaker’s emotional information throughout the entire encoding and decoding process. Experimental results demonstrate that our framework improves the appropriateness metric FRCorr by an average of 57.2% compared to the original model structure. In scenarios where speech modality data is missing, the performance of appropriate generation shows an improvement, and when facial data is missing, it only exhibits minimal degradation. Guanyu Hu 0003, Siyang Song, Dimitris Kollias, Xinyu Yang 0001, Zhonglin Sun, Odysseus Kaloidas |
IJCB | 3 |
| 2024 | CLIP-Guided Bidirectional Prompt and Semantic Supervision for Dynamic Facial Expression RecognitionabstractDue to the insufficient semantic information supervision in existing works for dynamic facial expression recognition (DFER), videos with similar facial changes but different expressions may be easily confused. Thanks to the potential textual information for semantic supervision, contrastive language-image pretraining (CLIP) model provides a new direction for DFER. However, pre-trained CLIP based on image-text pairs has difficulty in capturing temporal features in the video domain. Therefore, we propose a novel visual language model that captures and aggregates dynamic features of expressions in semantic supervision via Inter-Frame Interaction Transformer (Inter-FIT) and Multi-Scale Temporal Aggregation (MSTA). Furthermore, though prompt learning is often used in CLIP to enhance semantic supervision, previous studies have only focused on the role of textual prompts, ignoring the importance of visual prompts in facilitating the relationality between the two. Therefore, we designed a Bidirectional Enhanced Prompt (BiEhPro) to facilitate the learning of this relationality between text and visual cues in enhancing semantic supervision. Extensive experiments and ablation studies on three benchmark datasets, i.e., DFEW, FERV39K, and MAFW, validate the effectiveness of our modules and algorithm. Code is publicly available at https://github.com/JunLiangZ/CLIP-Guided-DFER. Junliang Zhang, Xiaole Xian, Weicheng Xie 0001, LinLin Shen, Siyang Song |
IJCB | 7 |
| 2024 | Community Hiding Algorithm Based on Likelihood Analysis Method in Link PredictionabstractIn recent years, the rapid development and widespread applications of non-overlapping community discovery have led to a growing issue of privacy leakage. The core of this problem lies in the community discovery algorithms themselves. Researchers have shifted their focus toward developing community hiding algorithms. However, existing non-overlapping community hiding algorithms have mainly been studied in static networks, overlooking the prevalence of dynamic networks in real life. In this paper, we introduce a community hiding algorithm called LALH, which employs the link prediction likelihood analysis method. The algorithm consists of two parts: first, the OLPL algorithm calculates a list of link probabilities in the subsequent time scale of the network, allowing observation of important link distribution. Second, the LALH algorithm selects these crucial links for the community hiding operation. We validate the effectiveness of the LALH algorithm through experiments on three public datasets and a real Twitter dataset. Ya-Si Wang, Zixuan Han, Siyang Song |
ISPA | 5 |
| 2024 | PerFRDiff: Personalised Weight Editing for Multiple Appropriate Facial Reaction GenerationabstractHuman facial reactions play crucial roles in dyadic human-human interactions, where individuals (i.e., listeners) with varying cognitive process styles may display different but appropriate facial reactions in response to an identical behaviour expressed by their conversational partners. While several existing facial reaction generation approaches are capable of generating multiple appropriate facial reactions (AFRs) in response to each given human behaviour, they fail to take human's personalised cognitive process in AFRs generation. In this paper, we propose the first online personalised multiple appropriate facial reaction generation (MAFRG) approach which learns a unique personalised cognitive style from the target human listener's previous facial behaviours and represents it as a set of network weight shifts. These personalised weight shifts are then applied to edit the weights of a pre-trained generic MAFRG model, allowing the obtained personalised model to naturally mimic the target human listener's cognitive process in its reasoning for multiple AFRs generations. Experimental results show that our approach not only largely outperformed all existing approaches in generating more appropriate and diverse generic AFRs, but also serves as the first reliable personalised MAFRG solution. Our code is made available at https://github.com/xk0720/PerFRDiff. Hengde Zhu, Xiangyu Kong 0001, Weicheng Xie 0001, LinLin Shen, Lu Liu 0001, Hatice Gunes, Siyang Song |
ACM Multimedia | 8 |
| 2024 | Towards Combating Frequency Simplicity-biased Learning for Domain GeneralizationabstractDomain generalization methods aim to learn transferable knowledge from source domains that can generalize well to unseen target domains.
Recent studies show that neural networks frequently suffer from a simplicity-biased learning behavior which leads to over-reliance on specific frequency sets, namely as frequency shortcuts, instead of semantic information, resulting in poor generalization performance.
Despite previous data augmentation techniques successfully enhancing generalization performances, they intend to apply more frequency shortcuts, thereby causing hallucinations of generalization improvement.
In this paper, we aim to prevent such learning behavior of applying frequency shortcuts from a data-driven perspective. Given the theoretical justification of models' biased learning behavior on different spatial frequency components, which is based on the dataset frequency properties, we argue that the learning behavior on various frequency components could be manipulated by changing the dataset statistical structure in the Fourier domain.
Intuitively, as frequency shortcuts are hidden in the dominant and highly dependent frequencies of dataset structure, dynamically perturbating the over-reliance frequency components could prevent the application of frequency shortcuts.
To this end, we propose two effective data augmentation modules designed to collaboratively and adaptively adjust the frequency characteristic of the dataset, aiming to dynamically influence the learning behavior of the model and ultimately serving as a strategy to mitigate shortcut learning. Our code will be made publicly available. Xilin He, Qinliang Lin, Weicheng Xie 0001, Siyang Song, Muhammad Haris Khan, LinLin Shen |
NeurIPS | 6 |
| 2024 | CemiFace: Center-based Semi-hard Synthetic Face Generation for Face RecognitionabstractPrivacy issue is a main concern in developing face recognition techniques. Although synthetic face images can partially mitigate potential legal risks while maintaining effective face recognition (FR) performance, FR models trained by face images synthesized by existing generative approaches frequently suffer from performance degradation problems due to the insufficient discriminative quality of these synthesized samples. In this paper, we systematically investigate what contributes to solid face recognition model training, and reveal that face images with certain degree of similarities to their identity centers show great effectiveness in the performance of trained FR models. Inspired by this, we propose a novel diffusion-based approach (namely **Ce**nter-based Se**mi**-hard Synthetic Face
Generation (**CemiFace**) which produces facial samples with various levels of similarity to the subject center, thus allowing to generate face datasets containing effective discriminative samples for training face recognition. Experimental results show that with a modest degree of similarity, training on the generated dataset can produce competitive performance compared to previous generation methods. The code will be available at:https://github.com/szlbiubiubiu/CemiFace Zhonglin Sun, Siyang Song, Ioannis Patras, Georgios Tzimiropoulos |
NeurIPS | 2 |
| 2024 | Loss Relaxation Strategy for Noisy Facial Video-based Automatic Depression RecognitionabstractAutomatic depression analysis has been widely investigated on face videos that have been carefully collected and annotated in lab conditions. However, videos collected under real-world conditions may suffer from various types of noise due to challenging data acquisition conditions and lack of annotators. Although deep learning (DL) models frequently show excellent depression analysis performances on datasets collected in controlled lab conditions, such noise may degrade their generalization abilities for real-world depression analysis tasks. In this article, we uncovered that noisy facial data and annotations consistently change the distribution of training losses for facial depression DL models; i.e., noisy data–label pairs cause larger loss values compared to clean data–label pairs. Since different loss functions could be applied depending on the employed model and task, we propose a generic loss function relaxation strategy that can jointly reduce the negative impact of various noisy data and annotation problems occurring in both classification and regression loss functions for face video-based depression analysis, where the parameters of the proposed strategy can be automatically adapted during depression model training. The experimental results on 25 different artificially created noisy depression conditions (i.e., five noise types with five different noise levels) show that our loss relaxation strategy can clearly enhance both classification and regression loss functions, enabling the generation of superior face video-based depression analysis models under almost all noisy conditions. Our approach is robust to its main variable settings and can adaptively and automatically obtain its parameters during training. Siyang Song, Tugba Tümer, Changzeng Fu, Michel F. Valstar, Hatice Gunes |
ACM Trans. Comput. Heal. | 1 |
| 2024 | Network characteristics adaption and hierarchical feature exploration for robust object recognition
Weicheng Xie 0001, Gui Wang, LinLin Shen, Zhihui Lai 0001, Siyang Song |
Pattern Recognit. | 6 |
| 2024 | An Open-Source Benchmark of Deep Learning Models for Audio-Visual Apparent and Self-Reported Personality RecognitionabstractPersonality determines a wide variety of human daily and working behaviours, and is crucial for understanding human internal and external states. In recent years, a large number of automatic personality computing approaches have been developed to predict either the apparent personality or self-reported personality of the subject based on non-verbal audio-visual behaviours. However, the majority of them suffer from complex and dataset-specific pre-processing steps and model training tricks. In the absence of a standardized benchmark with consistent experimental settings, it is not only impossible to fairly compare the real performances of these personality computing models but also makes them difficult to be reproduced. In this paper, we present the first reproducible audio-visual benchmarking framework to provide a fair and consistent evaluation of eight existing personality computing models (e.g., audio, visual and audio-visual) and seven standard deep learning models on both self-reported and apparent personality recognition tasks. Building upon a set of benchmarked models, we also investigate the impact of two previously-used long-term modelling strategies for summarising short-term/frame-level predictions on personality computing results. We conduct a comprehensive investigation into all the benchmarked models to demonstrate their capabilities in modelling personality traits on two publicly available datasets, audio-visual apparent personality (ChaLearn First Impression) and self-reported personality (UDIVA) datasets. The experimental results conclude: (i) apparent personality traits, inferred from facial behaviours by most benchmarked deep learning models, show more reliability than self-reported ones; (ii) visual models frequently achieved superior performances than audio models on personality recognition; (iii) non-verbal behaviours contribute differently in predicting different personality traits; and (iv) our reproduced personality computing models generally achieved worse performances than their original reported results. We make all the code and settings of this personality computing benchmark publicly available athttps://github.com/liaorongfan/DeepPersonality. Rongfan Liao, Siyang Song, Hatice Gunes |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Generative Imperceptible Attack With Feature Learning Bias Reduction and Multi-Scale Variance RegularizationabstractExisting studies have shown that malicious and imperceptible adversarial samples may significantly weaken the reliability and validity of deep learning systems. Since gradient-based attack algorithms may result in higher generation latency or demand large computation overhead, generative attack methods are frequently considered. However, the effectiveness and imperceptibility are still the main concerns for these generative attacks, 1) biased feature learning may occur, i.e., these algorithms may generate undesirable feature perturbations for samples that are less likely to be successfully attacked; 2) the produced perturbation noises may be easily perceived by human eyes. To this end, we propose a novel generative attack by manipulating the feature update. The proposed algorithm has two main merits, 1) our Bias-reduced Feature Manipulation (BrFM) that differentiates the hard-to-attack (Hard2Attack) and easy-to-attack (Easy2Attack) features, can avoid the possible learning shortcut for different difficulties of features in attack process, by customizing perturbations for Hard2Attack features to make them behave oppositely to those of benign features; 2) our Multi-scale Variance Regularization (MsVR) can reduce the unnatural transitions of perturbations in mask edges and flat areas with low contrast, while simultaneously trading off a reasonable attack capacity. Extensive experiments on the datasets of Caltech-101 and Imagenette in terms of the attack success rate and four imperceptibility metrics, show the effectiveness of our attack paradigm over the related state-of-the-art generative attack methods. Our codes will be made publicly available. Weicheng Xie 0001, Zenghao Niu, Qinliang Lin, Siyang Song, LinLin Shen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Cross-Layer Contrastive Learning of Latent Semantics for Facial Expression RecognitionabstractConvolutional neural networks (CNNs) have achieved significant improvement for the task of facial expression recognition. However, current training still suffers from the inconsistent learning intensities among different layers, i.e., the feature representations in the shallow layers are not sufficiently learned compared with those in deep layers. To this end, this work proposes a contrastive learning framework to align the feature semantics of shallow and deep layers, followed by an attention module for representing the multi-scale features in the weight-adaptive manner. The proposed algorithm has three main merits. First, the learning intensity, defined as the magnitude of the backpropagation gradient, of the features on the shallow layer is enhanced by cross-layer contrastive learning. Second, the latent semantics in the shallow-layer and deep-layer features are explored and aligned in the contrastive learning, and thus the fine-grained characteristics of expressions can be taken into account for the feature representation learning. Third, by integrating the multi-scale features from multiple layers with an attention module, our algorithm achieved the state-of-the-art performances, i.e. 92.21%, 89.50%, 62.82%, on three in-the-wild expression databases, i.e. RAF-DB, FERPlus, SFEW, and the second best performance, i.e. 65.29% on AffectNet dataset. Our codes will be made publicly available. Weicheng Xie 0001, Zhibin Peng, LinLin Shen, Wenya Lu, Yang Zhang 0012, Siyang Song |
IEEE Trans. Image Process. | 6 |
| 2024 | Unlocking Human-Like Facial Expressions in Humanoid Robots: A Novel Approach for Action Unit Driven Facial Expression Disentangled SynthesisabstractHumanoid robots often struggle to express the intricate and authentic facial expressions characteristic of humans, potentially hampering user engagement. To address this challenge, we introduce a comprehensive two-stage methodology to empower our autonomous affective robot with the capacity to exhibit rich and natural facial expressions. In the initial stage, we present an innovative action unit (AU) driven facial expression disentangled synthesis method, enabling the generation of nuanced robot facial expression images guided by AUs. By harnessing facial AUs within a framework of weakly supervised learning, we effectively surmount the scarcity of paired training data (comprising source and target facial expression images). To preserve the integrity of AUs while mitigating identity interference, we leverage a latent facial attribute space to disentangle expression-related and expression-unrelated cues, employing solely the former for expression synthesis. In the subsequent phase, we actualize an affective robot endowed with multifaceted degrees of freedom for facial movements, facilitating the embodiment of the synthesized fine-grained facial expressions. We devise a specialized motor command mapping network that serves as a conduit between the generated expression images and the robot's realistic facial responses. By utilizing the physical motor positions as constraints, we refine the prediction of precise motor commands from the robot's generated facial expressions. This refinement process ensures that the robot's facial movements authentically express accurate and natural expressions. Finally, qualitative and quantitative evaluations on the benchmarking Emotionet dataset verify the effectiveness of the proposed generation method. Results on the self-developed affective robot indicate that our method achieves a promising generation of specific facial expressions with given AUs, significantly enhancing the affective human–robot interaction. Xiaofeng Liu 0006, Siyang Song, Angelo Cangelosi |
IEEE Trans. Robotics | 4 |
| 2023 | Fourier-Net: Fast Image Registration with Band-Limited DeformationabstractUnsupervised image registration commonly adopts U-Net style networks to predict dense displacement fields in the full-resolution spatial domain. For high-resolution volumetric image data, this process is however resource-intensive and time-consuming. To tackle this problem, we propose the Fourier-Net, replacing the expansive path in a U-Net style network with a parameter-free model-driven decoder. Specifically, instead of our Fourier-Net learning to output a full-resolution displacement field in the spatial domain, we learn its low-dimensional representation in a band-limited Fourier domain. This representation is then decoded by our devised model-driven decoder (consisting of a zero padding layer and an inverse discrete Fourier transform layer) to the dense, full-resolution displacement field in the spatial domain. These changes allow our unsupervised Fourier-Net to contain fewer parameters and computational operations, resulting in faster inference speeds. Fourier-Net is then evaluated on two public 3D brain datasets against various state-of-the-art approaches. For example, when compared to a recent transformer-based method, named TransMorph, our Fourier-Net, which only uses 2.2% of its parameters and 6.66% of the multiply-add operations, achieves a 0.5% higher Dice score and an 11.48 times faster inference speed. Code is available at https://github.com/xi-jia/Fourier-Net. Xi Jia, Joseph Bartlett, Wei Chen 0092, Siyang Song, Tianyang Miller, Xinxing Cheng, Wenqi Lu 0001, Zhaowen Qiu, Jinming Duan 0001 |
AAAI | 4 |
| 2023 | Shift from Texture-bias to Shape-bias: Edge Deformation-based Augmentation for Robust Object RecognitionabstractRecent studies have shown the vulnerability of CNNs under perturbation noises, which is partially caused by the reason that the well-trained CNNs are too biased toward the object texture, i.e., they make predictions mainly based on texture cues. To reduce this texture-bias, current studies resort to learning augmented samples with heavily perturbed texture to make networks be more biased toward relatively stable shape cues. However, such methods usually fail to achieve real shape-biased networks due to the insufficient diversity of the shape cues. In this paper, we propose to augment the training dataset by generating semantically meaningful shapes and samples, via a shape deformation-based online augmentation, namely as SDbOA. The samples generated by our SDbOA have two main merits. First, the augmented samples with more diverse shape variations enable networks to learn the shape cues more elaborately, which encourages the network to be shape-biased. Second, semantic-meaningful shape-augmentation samples could be produced by jointly regularizing the generator with object texture and edge-guidance soft constraint, where the edges are represented more robustly with a self information guided map to better against the noises on them. Extensive experiments under various perturbation noises demonstrate the obvious superiority of our shape-bias-motivated model over the state of the arts in terms of robustness performance. Code is available at https://github.com/C0notSilly/-ICCV-23-Edge-Deformation-based-Online-Augmentation. Xilin He, Qinliang Lin, Weicheng Xie 0001, Siyang Song, Feng Liu 0013, LinLin Shen |
ICCV | 5 |
| 2023 | Joint Prediction of Audio Event and Annoyance Rating in an Urban Soundscape by Hierarchical Graph Representation LearningabstractSound events in daily life carry rich information about the objective world. The composition of these sounds affects the mood of people in a soundscape. Most previous approaches only focus on classifying and detecting audio events and scenes, but may ignore their perceptual quality that may impact humans' listening mood for the environment, e.g. annoyance. To this end, this paper proposes a novel hierarchical graph representation learning (HGRL) approach which links objective audio events (AE) with subjective annoyance ratings (AR) of the soundscape perceived by humans. The hierarchical graph consists of fine-grained event (fAE) embeddings with single-class event semantics, coarse-grained event (cAE) embeddings with multi-class event semantics, and AR embeddings. Experiments show the proposed HGRL successfully integrates AE with AR for AEC and ARP tasks, while coordinating the relations between cAE and fAE and further aligning the two different grains of AE information with the AR. Yuanbo Hou, Siyang Song, Qiaoqiao Ren, Weicheng Xie 0001, Jian Kang 0002, Wenwu Wang 0001, Dick Botteldooren |
INTERSPEECH | 2 |
| 2023 | REACT2023: The First Multiple Appropriate Facial Reaction Generation ChallengeabstractThe Multiple Appropriate Facial Reaction Generation Challenge (REACT2023) is the first competition event focused on evaluating multimedia processing and machine learning techniques for generating human-appropriate facial reactions in various dyadic interaction scenarios, with all participants competing strictly under the same conditions. The goal of the challenge is to provide the first benchmark test set for multi-modal information processing and to foster collaboration among the audio, visual, and audio-visual behaviour analysis and behaviour generation (a.k.a generative AI) communities, to compare the relative merits of the approaches to automatic appropriate facial reaction generation under different spontaneous dyadic interaction conditions. This paper presents: (i) the novelties, contributions and guidelines of the REACT2023 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2023. Siyang Song, Micol Spitale, Germán Barquero, Cristina Palmero, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes |
ACM Multimedia | 1 |
| 2023 | Diverse local facial behaviors learning from enhanced expression flow for microexpression recognition
Xu Zhou 0002, Siyang Song, Xiaofeng Liu 0006 |
Knowl. Based Syst. | 4 |
| 2023 | Audio Event-Relational Graph Representation Learning for Acoustic Scene ClassificationabstractMost deep learning-based acoustic scene classification (ASC) approaches identify scenes based on acoustic features converted from audio clips containing mixed information entangled by polyphonic audio events (AEs). However, these approaches have difficulties in explaining what cues they use to identify scenes. This paper conducts the first study on disclosing the relationship between real-life acoustic scenes and semantic embeddings from the most relevant AEs. Specifically, we propose an event-relational graph representation learning (ERGL) framework for ASC to classify scenes, and simultaneously answer clearly and straightly which cues are used in classifying. In the event-relational graph, embeddings of each event are treated as nodes, while relationship cues derived from each pair of nodes are described by multi-dimensional edge features. Experiments on a real-life ASC dataset show that the proposed ERGL achieves competitive performance on ASC by learning embeddings of only a limited number of AEs. The results show the feasibility of recognizing diverse acoustic scenes based on the audio event-relational graph. Visualizations of graph representations learned by ERGL are available here(https://github.com/Yuanbo2020/ERGL). Yuanbo Hou, Siyang Song, Chuang Yu 0001, Wenwu Wang 0001, Dick Botteldooren |
IEEE Signal Process. Lett. | 2 |
| 2023 | Self-Supervised Learning of Person-Specific Facial Dynamics for Automatic Personality RecognitionabstractThis article aims to solve two important issues that frequently occur in existing automatic personality analysis systems: 1. Attempting to use very short video segments or even single frames, rather than long-term behaviour, to infer personality traits; 2. Lack of methods to encode person-specific facial dynamics for personality recognition. To deal with these issues, this paper first proposes a novel Rank Loss which utilizes the natural temporal evolution of facial actions, rather than personality labels, for self-supervised learning of facial dynamics. Our approach first trains a generic U-net style model that can infer general facial dynamics learned from a set of unlabelled face videos. Then, the generic model is frozen, and a set of intermediate filters are incorporated into this architecture. The self-supervised learning is then resumed with only person-specific videos. This way, the learned filters’ weights are person-specific, making them a valuable source for modeling person-specific facial dynamics. We then propose to concatenate the weights of the learned filters as a person-specific representation, which can be directly used to predict the personality traits without needing other parts of the network. We evaluate the proposed approach on both self-reported personality and apparent personality datasets. In addition to achieving promising results in the estimation of personality trait scores from videos, we show that the tasks conducted by the subject in the video matters, that fusion of a combination of tasks reaches highest accuracy, and that multi-scale dynamics are more informative than single-scale dynamics. Siyang Song, Shashank Jaiswal, Enrique Sánchez-Lozano, Georgios Tzimiropoulos, LinLin Shen, Michel F. Valstar |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Learning Person-Specific Cognition From Facial Reactions for Automatic Personality RecognitionabstractThis article proposes to recognise the true (self-reported) personality traits from the target subject's cognition simulated from facial reactions. This approach builds on the following two findings in cognitive science: (i) human cognition partially determines expressed behaviour and is directly linked to true personality traits; and (ii) in dyadic interactions, individuals’ nonverbal behaviours are influenced by their conversational partner's behaviours. In this context, we hypothesise that during a dyadic interaction, a target subject's facial reactions are driven by two main factors: their internal (person-specific) cognitive process, and the externalised nonverbal behaviours of their conversational partner. Consequently, we propose to represent the target subject's (defined as the listener) person-specific cognition in the form of a person-specific CNN architecture that has unique architectural parameters and depth, which takes audio-visual non-verbal cues displayed by the conversational partner (defined as the speaker) as input, and is able to reproduce the target subject's facial reactions. Each person-specific CNN is explored by the Neural Architecture Search (NAS) and a novel adaptive loss function, which is then represented as a graph representation for recognising the target subject's true personality. Experimental results not only show that the produced graph representations are well associated with target subjects’ personality traits in both human-human and human-machine interaction scenarios, and outperform the existing approaches with significant advantages, but also demonstrate that the proposed novel strategies help in learning more reliable personality representations. Siyang Song, Zilong Shao, Shashank Jaiswal, LinLin Shen, Michel F. Valstar, Hatice Gunes |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Emotion Recognition Through Combining EEG and EOG Over Relevant Channels With Optimal WindowingabstractFor dimensional emotion recognition, electroencephalography (EEG) signals and electrooculogram (EOG) signals are often combined to improve the performance of classifiers, as each of them provides complementary features to the other. In this article, we combine the EEG signal on the relevant channels with the EOG signal to boost the recognition accuracy. We first explore the mutual information (MI) of all EEG channels and only select emotion-related channels, i.e., channels with more MI are retained, since the emotion recognition performance can be degraded by the interference between uncorrelated channels, while the computational complexity is significant if all EEG channels are used for recognition. While the optimal lengths of EEG and EOG signals for emotion recognition are still uncertain, we systematically investigate the effects of time-window size on emotion recognition. This strategy not only increases the number of training samples, but also reduces the feature redundancy. At this stage, we not only extract multiple statistical features but also employ the increment entropy to find abrupt changes in EEG signals. The experimental results show that 13 out of 32 EEG channels were selected by the proposed channel selection algorithm, and these selected channels can already produce accurate emotion predictions. We found that using optimal time-windows to split EEG and EOG signals into several thin slices and then combine them can further enhance the emotion recognition performance, where the time-windows of 4, 5, 6, and 10 s allow the combined signals to achieve very high accuracy. Huili Cai, Xiaofeng Liu 0006, Siyang Song, Angelo Cangelosi |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2022 | Statistical, Spectral and Graph Representations for Video-Based Facial Expression Recognition in ChildrenabstractChild facial expression recognition is a relatively less investigated area within affective computing. Children’s facial expressions differ significantly from adults; thus, it is necessary to develop emotion recognition frameworks that are more objective, descriptive and specific to this target user group. In this paper we propose the first approach that (i) constructs video-level heterogeneous graph representation for facial expression recognition in children, and (ii) predicts children’s facial expressions using the automatically detected Action Units (AUs). To this aim, we construct three separate length-independent representations, namely, statistical, spectral and graph at video-level for detailed multi-level facial behaviour decoding (AU activation status, AU temporal dynamics and spatio-temporal AU activation patterns, respectively). Our experimental results on the LIRIS Children Spontaneous Facial Expression Video Database demonstrate that combining these three feature representations provides the highest accuracy for expression recognition in children. Nida Itrat Abbasi, Siyang Song, Hatice Gunes |
ICASSP | 2 |
| 2022 | Learning Multi-dimensional Edge Feature-based AU Relation Graph for Facial Action Unit RecognitionabstractThe activations of Facial Action Units (AUs) mutually influence one another. While the relationship between a pair of AUs can be complex and unique, existing approaches fail to specifically and explicitly represent such cues for each pair of AUs in each facial display. This paper proposes an AU relationship modelling approach that deep learns a unique graph to explicitly describe the relationship between each pair of AUs of the target facial display. Our approach first encodes each AU's activation status and its association with other AUs into a node feature. Then, it learns a pair of multi-dimensional edge features to describe multiple task-specific relationship cues between each pair of AUs. During both node and edge feature learning, our approach also considers the influence of the unique facial display on AUs' relationship by taking the full face representation as an input. Experimental results on BP4D and DISFA datasets show that both node and edge feature learning modules provide large performance improvements for CNN and transformer-based backbones, with our best systems achieving the state-of-the-art AU recognition results. Our approach not only has a strong capability in modelling relationship cues for AU recognition but also can be easily incorporated into various backbones. Our PyTorch code is made available at https://github.com/CVI-SZU/ME-GraphAU. Siyang Song, Weicheng Xie 0001, LinLin Shen, Hatice Gunes |
IJCAI | 2 |
| 2022 | Spectral Representation of Behaviour Primitives for Depression AnalysisabstractDepression is a serious mental disorder affecting millions of people all over the world. Traditional clinical diagnosis methods are subjective, complicated and require extensive participation of clinicians. Recent advances in automatic depression analysis systems promise a future where these shortcomings are addressed by objective, repeatable, and readily available diagnostic tools to aid health professionals in their work. Yet there remain a number of barriers to the development of such tools. One barrier is that existing automatic depression analysis algorithms base their predictions on very brief sequential segments, sometimes as little as one frame. Another barrier is that existing methods do not take into account what the context of the measured behaviour is. In this article, we extract multi-scale video-level features for video-based automatic depression analysis. We propose to use automatically detected human behaviour primitives as the low-dimensional descriptor for each frame. We also propose two novel spectral representations, i.e., spectral heatmaps and spectral vectors, to represent video-level multi-scale temporal dynamics of expressive behaviour. Constructed spectral representations are fed to Convolution Neural Networks (CNNs) and Artificial Neural Networks (ANNs) for depression analysis. We conducted experiments on the AVEC 2013 and AVEC 2014 benchmark datasets to investigate the influence of interview tasks on depression analysis. In addition to achieving state of the art accuracy in severity of depression estimation, we show that the task conducted by the user matters, that fusion of a combination of tasks reaches highest accuracy, and that longer tasks are more informative than shorter tasks, up to a point. Siyang Song, Shashank Jaiswal, LinLin Shen, Michel F. Valstar |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | A Tool to Facilitate the Cross-Cultural Design Process Using Deep LearningabstractCross-cultural design requires designers to understand other foreign cultures, selecting suitable cultural elements, and finally incorporate them into product design. Traditionally, this process is time-consuming and relies to a significant extent on designers’ cultural awareness and design skills. This article proposes a new tool for designers to select and integrate cultural elements in the cross-cultural design process. The proposed approach utilizes state-of-the-art deep learning techniques, which begins by automatically selecting the most suitable style image from all cultural image candidates. Then, the deep-learning-based style transfer technique is introduced to automatically produce a design image that has the same content as the uploaded design content image, and also has the cultural style of the selected style image. To the best of our knowledge, this is the first work that extends deep learning techniques to facilitate cross-cultural design. The tool received positive feedback in a usability evaluation. The empirical results show that our approach can effectively increase designers’ cultural awareness in respect of four cultural element dimensions (color, material, pattern and form). It is an innovative and efficient tool to help designers with idea generation and fast prototyping, although some participants argued that the tool would only assist designers, rather than replace humans. Leijing Zhou, Xu Sun 0002, Guannan Mu, Jiayi Wu 0003, Jiangping Zhou, Qiuning Wu, Yaorun Zhang, Yufan Xi, Nesrin Dilber Günes, Siyang Song |
IEEE Trans. Hum. Mach. Syst. | 10 |
| 2021 | Personality Recognition by Modelling Person-specific Cognitive Processes using Graph RepresentationabstractRecent research shows that in dyadic and group interactions individuals' nonverbal behaviours are influenced by the behaviours of their conversational partner(s). Therefore, in this work we hypothesise that during a dyadic interaction, the target subject's facial reactions are driven by two main factors: (i) their internal (person-specific) cognition, and (ii) the externalised nonverbal behaviours of their conversational partner. Subsequently, our novel proposition is to simulate and represent the target subject's (i.e., the listener) cognitive process in the form of a person-specific CNN architecture whose input is the audio-visual non-verbal cues displayed by the conversational partner (i.e., the speaker), and the output is the target subject's (i.e., the listener) facial reactions. We then undertake a search for the optimal CNN architecture whose results are used to create a person-specific graph representation for recognising the target subject's personality. The graph representation, fortified with a novel end-to-end edge feature learning strategy, helps with retaining both the unique parameters of the person-specific CNN and the geometrical relationship between its layers. Consequently, the proposed approach is the first work that aims to recognize the true (self-reported) personality of a target subject (i.e., the listener) from the learned simulation of their cognitive process (i.e., parameters of the person-specific CNN). The experimental results show that the CNN architectures are well associated with target subjects' personality traits and the proposed approach clearly outperforms multiple existing approaches that predict personality directly from non-verbal behaviours. In light of these findings, this work opens up a new avenue of research for predicting and recognizing socio-emotional phenomena (personality, affect, engagement etc.) from simulations of person-specific cognitive processes. Zilong Shao, Siyang Song, Shashank Jaiswal, LinLin Shen, Michel F. Valstar, Hatice Gunes |
ACM Multimedia | 2 |
| 2020 | EMOPAIN Challenge 2020: Multimodal Pain Evaluation from Facial and Bodily ExpressionsabstractThe EmoPain 2020 Challenge is the first international competition aimed at creating a uniform platform for the comparison of multi-modal machine learning and multimedia processing methods of chronic pain assessment from human expressive behaviour, and also the identification of pain-related behaviours. The objective of the challenge is to promote research in the development of assistive technologies that help improve the quality of life for people with chronic pain via real-time monitoring and feedback to help manage their condition and remain physically active. The challenge also aims to encourage the use of the relatively underutilised, albeit vital bodily expression signals for automatic pain and pain-related emotion recognition. This paper presents a description of the challenge, competition guidelines, bench-marking dataset, and the baseline systems' architecture and performance on the Challenge's three sub-tasks: pain estimation from facial expressions, pain recognition from multimodal movement, and protective movement behaviour detection. Joy Egede, Siyang Song, Temitayo A. Olugbade, Amanda C. de C. Williams, Hongying Meng, M. S. Hane Aung, Nicholas D. Lane, Michel F. Valstar, Nadia Bianchi-Berthouze |
FG | 2 |
| 2020 | Self-supervised learning of Dynamic Representations for Static ImagesabstractFacial actions are spatio-temporal signals by nature, and therefore their modeling is crucially dependent on the availability of temporal information. In this paper, we focus on inferring such temporal dynamics of facial actions when no explicit temporal information is available, i.e. from still images. We present a novel self-supervised learning approach to capture multiple scales of temporal dynamics, with an application to facial Action Unit (AU) intensity estimation and dimensional affect estimation. In particular: 1. We propose a framework that infers a dynamic representation (DR) from a still image, capturing the bi-directional flow of time within a short time-window centered at the input image; 2. We show that the proposed rank loss can apply facial temporal evolution to self-supervise the training process without using target representations, allowing the network to represent dynamics more broadly; 3. We propose a multiple temporal scale approach that infers DRs for different window lengths (MDR) from a still image. We empirically validate the value of our approach on the task of frame ranking, and show how our proposed MDR attains state of the art results on BP4D for AU intensity estimation and on SEMAINE for dimensional affect estimation, using only still images at test time. Siyang Song, Enrique Sanchez, LinLin Shen, Michel F. Valstar |
ICPR | 1 |
| 2019 | Automatic prediction of Depression and Anxiety from behaviour and personality attributesabstractCurrent vision based approaches for automatic prediction of mental health conditions like depression and anxiety, rely on models that use behavioural features (usually extracted from faces) only and do not take personality into account. However, there is a considerable amount of evidence that people with certain personality traits are more prone to depression and anxiety disorders. In order to exploit the underlying relationship between personality and these mental health conditions, we propose to use a combination of features consisting of observed facial behaviour and self-reported personality scores. This combination of features is employed for training deep neural networks for predicting depression and anxiety scores. The proposed method was evaluated on a new dataset consisting of personality scores and interview videos from a total of 55 people. The results show that the proposed combination of features significantly improves the prediction performance compared to using the behavioural features alone. This improvement in performance is shown for 2 different kinds of video feature extraction methods and 3 different kinds of interview questionnaires. Shashank Jaiswal, Siyang Song, Michel F. Valstar |
ACII | 2 |
| 2018 | Human Behaviour-Based Automatic Depression Analysis Using Hand-Crafted Statistics and Deep Learned Spectral FeaturesabstractDepression is a serious mental disorder that affects millions of people all over the world. Traditional clinical diagnosis methods are subjective, complicated and need extensive participation of experts. Audio-visual automatic depression analysis systems predominantly base their predictions on very brief sequential segments, sometimes as little as one frame. Such data contains much redundant information, causes a high computational load, and negatively affects the detection accuracy. Final decision making at the sequence level is then based on the fusion of frame or segment level predictions. However, this approach loses longer term behavioural correlations, as the behaviours themselves are abstracted away by the frame-level predictions. We propose to on the one hand use automatically detected human behaviour primitives such as Gaze directions, Facial action units (AU), etc. as low-dimensional multi-channel time series data, which can then be used to create two sequence descriptors. The first calculates the sequence-level statistics of the behaviour primitives and the second casts the problem as a Convolutional Neural Network problem operating on a spectral representation of the multichannel behaviour signals. The results of depression detection (binary classification) and severity estimation (regression) experiments conducted on the AVEC 2016 DAIC-WOZ database show that both methods achieved significant improvement compared to the previous state of the art in terms of the depression severity estimation. Siyang Song, LinLin Shen, Michel F. Valstar |
FG | 1 |
| 2018 | Noise Invariant Frame Selection: A Simple Method to Address the Background Noise Problem for Text-independent Speaker VerificationabstractThe performance of speaker-related systems usually degrades heavily in practical applications largely due to the presence of background noise. To improve the robustness of such systems in unknown noisy environments, this paper proposes a simple pre-processing method called Noise Invariant Frame Selection (NIFS). Based on several noisy constraints, it selects noise invariant frames from utterances to represent speakers. Experiments conducted on the TIMIT database showed that the NIFS can significantly improve the performance of Vector Quantization (VQ), Gaussian Mixture Model-Universal Background Model (GMM-UBM) and i-vector-based speaker verification systems in different unknown noisy environments with different SNRs, in comparison to their baselines. Meanwhile, the proposed NIFS-based speaker verification systems achieves similar performance when we change the constraints (hyper-parameters) or features, which indicates that it is robust and easy to reproduce. Since NIFS is designed as a general algorithm, it could be further applied to other similar tasks. Siyang Song, Shuimei Zhang, Björn W. Schuller, LinLin Shen, Michel F. Valstar |
IJCNN | 1 |