VLDB 2026 Research / reviewers in the wild / expert
Songlong Xing
dblp:245/8732
· DBLP profile ↗
14ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-2734-1695ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIPabstractDespite its prevalent use in image-text matching tasks in a zero-shot manner, CLIP has been shown to be highly vulnerable to adversarial perturbations added onto images. Recent studies propose to finetune the vision encoder of CLIP with adversarial samples generated on the fly, and show improved robustness against adversarial attacks on a spectrum of downstream datasets, a property termed as zero-shot robustness. In this paper, we show that malicious perturbations that seek to maximise the classification loss lead to ‘falsely stable’ images, and propose to leverage the pre-trained vision encoder of CLIP to counterattack such adversarial images during inference to achieve robustness. Our paradigm is simple and training-free, providing the first method to defend CLIP from adversarial attacks at test time, which is orthogonal to existing methods aiming to boost zero-shot adversarial robustness of CLIP. We conduct experiments across 16 classification datasets, and demonstrate stable and consistent gains compared to test-time defence methods adapted from existing adversarial robustness studies that do not rely on external networks, without noticeably impairing performance on clean images. We also show that our paradigm can be employed on CLIP models that have been adversarially finetuned to further enhance their robustness at test time. Our code is available here. Songlong Xing, Zhengyu Zhao 0001, Nicu Sebe |
CVPR | 1 |
| 2024 | From One to Many Lorikeets: Discovering Image Analogies in the CLIP Space
Songlong Xing, Elia Peruzzo, Enver Sangineto, Nicu Sebe |
ICPR (9) | 1 |
| 2024 | SpectralCLIP: Preventing Artifacts in Text-Guided Style Transfer from a Spectral PerspectiveabstractOwing to the power of vision-language foundation models, e.g., CLIP, the area of image synthesis has seen recent important advances. Particularly, for style transfer, CLIP enables transferring more general and abstract styles without collecting the style images in advance, as the style can be efficiently described with natural language, and the result is optimized by minimizing the CLIP similarity between the text description and the stylized image. However, directly using CLIP to guide style transfer leads to undesirable artifacts (mainly written words and unrelated visual entities) spread over the image. In this paper, we propose SpectralCLIP, which is based on a spectral representation of the CLIP embedding sequence, where most of the common artifacts occupy specific frequencies. By masking the band including these frequencies, we can condition the generation process to adhere to the target style properties (e.g., color, texture, paint stroke, etc.) while excluding the generation of larger-scale structures corresponding to the artifacts. Experimental results show that SpectralCLIP prevents the generation of artifacts effectively in quantitative and qualitative terms, without impairing the stylisation quality. We also apply SpectralCLIP to text-conditioned image generation and show that it prevents written words in the generated images. Our code is available at https://github.com/zipengxuc/SpectralCLIP. Zipeng Xu, Songlong Xing, Enver Sangineto, Nicu Sebe |
WACV | 2 |
| 2024 | Dual Learning for Conversational Emotion Recognition and Emotional Response GenerationabstractEmotion recognition in conversation (ERC) and emotional response generation (ERG) are two important NLP tasks. ERC aims to detect the utterance-level emotion from a dialogue, while ERG focuses on expressing a desired emotion. Essentially, ERC is a classification task, with its input and output domains being the utterance text and emotion labels, respectively. On the other hand, ERG is a generation task with its input and output domains being the opposite. These two tasks are highly related, but surprisingly, they are addressed independently without making use of their duality in prior works. Therefore, in this paper, we propose to solve these two tasks in a dual learning framework. Our contributions are fourfold: (1) We propose a dual learning framework for ERC and ERG. (2) Within the proposed framework, two models can be trained jointly, so that the duality between them can be utilised. (3) Instead of a symmetric framework that deals with two tasks of the same data domain, we propose a dual learning framework that performs on a pair of asymmetric input and output spaces, i.e., the natural language space and the emotion labels. (4) Experiments are conducted on benchmark datasets to demonstrate the effectiveness of our framework. Shuhe Zhang, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Multimodal Graph for Unaligned Multimodal Sequence Analysis via Graph Convolution and Graph PoolingabstractMultimodal sequence analysis aims to draw inferences from visual, language, and acoustic sequences. A majority of existing works focus on the aligned fusion of three modalities to explore inter-modal interactions, which is impractical in real-world scenarios. To overcome this issue, we seek to focus on analyzing unaligned sequences, which is still relatively underexplored and also more challenging. We propose Multimodal Graph, whose novelty mainly lies in transforming the sequential learning problem into graph learning problem. The graph-based structure enables parallel computation in time dimension (as opposed to recurrent neural network) and can effectively learn longer intra- and inter-modal temporal dependency in unaligned sequences. First, we propose multiple ways to construct the adjacency matrix for sequence to perform sequence to graph transformation. To learn intra-modal dynamics, a graph convolution network is employed for each modality based on the defined adjacency matrix. To learn inter-modal dynamics, given that the unimodal sequences are unaligned, the commonly considered word-level fusion does not pertain. To this end, we innovatively devise graph pooling algorithms to automatically explore the associations between various time slices from different modalities and learn high-level graph representation hierarchically. Multimodal Graph outperforms state-of-the-art models on three datasets under the same experimental setting. Sijie Mai, Songlong Xing, Jiaxuan He, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Multi-Fusion Residual Memory Network for Multimodal Human Sentiment ComprehensionabstractMultimodal human sentiment comprehension refers to recognizing human affection from multiple modalities. There exist two key issues for this problem. First, it is difficult to explore time-dependent interactions between modalities and focus on the important time steps. Second, processing the long fused sequence of utterances is susceptible to the forgetting problem due to the long-term temporal dependency. In this article, we introduce a hierarchical learning architecture to classify utterance-level sentiment. To address the first issue, we perform time-step level fusion to generate fused features for each time step, which explicitly models time-restricted interactions by incorporating information across modalities at the same time step. Furthermore, based on the assumption that acoustic features directly reflect emotional intensity, we pioneer emotion intensity attention to focus on the time steps where emotion changes or intense affections take place. To handle the second issue, we propose Residual Memory Network (RMN) to process the fused sequence. RMN utilizes some techniques such as directly passing the previous state into the next time step, which helps to retain the information from many time steps ago. We show that our method achieves state-of-the-art performance on multiple datasets. Results also suggest that RMN yields competitive performance on sequence modeling tasks. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Adapted Dynamic Memory Network for Emotion Recognition in ConversationabstractIn this article, we address Emotion Recognition in Conversation (ERC) where conversational data are presented in a multimodal setting. Psychological evidence shows that self and inter-speaker influence are two central factors to emotion dynamics in conversation. State-of-the-art models do not effectively synthesise these two factors. Therefore, we propose an Adapted Dynamic Memory Network (A-DMN) where self and inter-speaker influences are modelled individually and further synthesised oriented towards the current utterance. Specifically, we model the dependency of the constituent utterances in a dialogue video using a global RNN to capture inter-speaker influence. Likewise, each speaker is assigned an RNN to capture their self influence. Afterwards, an Episodic Memory Module is devised to extract contexts for self and inter-speaker influence and synthesise them to update the memory. This process repeats itself for multiple passes until a refined representation is obtained and used for final prediction. Additionally, we explore cross-modal fusion in the context of multimodal ERC, and propose a convolution-based method which proves effective in extracting local interactions and computationally efficient. Extensive experiments demonstrate that A-DMN outperforms the state-of-the-art models on benchmark datasets. Songlong Xing, Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | A Unimodal Representation Learning and Recurrent Decomposition Fusion Structure for Utterance-Level Multimodal Embedding LearningabstractLearning a unified embedding for utterance-level video attracts significant attention recently due to the rapid development of social media and its broad applications. An utterance normally contains not only spoken language but also the nonverbal behaviors such as facial expressions and vocal patterns. Instead of directly learning utterance embedding based on low-level features, we firstly explore high-level representation for each modality separately via an unimodal representation learning gyroscope structure. In this way, the learnt unimodal representations are more representative and contain more abstract semantic information. In the gyroscope structure, we introduce multi-scale kernel learning, ‘channel expansion’ and ‘channel fusion’ operations to explore high-level features both spatially and channelwise. Another insight of our method lies in that we fuse representations of all modalities to obtain a unified embedding by interpreting fusion procedure as the flow of inter-modality information between various modalities, which is more specialized in terms of the information to be fused and the fusion process. Specifically, considering that each modality carries modality-specific and cross-modality interactions, we innovate to decompose unimodal representations into intra- and inter-modality dynamics using gating mechanism, and further fuse the inter-modality dynamics by passing them from previous modalities to the following one using a recurrent neural fusion architecture. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple benchmark datasets. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Multim. | 3 |
| 2021 | Analyzing Multimodal Sentiment Via Acoustic- and Visual-LSTM With Channel-Aware Temporal Convolution NetworkabstractThe emotion of human is always expressed in a multimodal perspective. Analyzing multimodal human sentiment remains challenging due to the difficulties of the interpretation in inter-modality dynamics. Mainstream multimodal learning architectures tend to design various fusion strategies to learn inter-modality interactions, which barely consider the fact that the language modality is far more important than the acoustic and visual modalities. In contrast, we learn inter-modality dynamics in a different perspective via acoustic- and visual-LSTMs where language features play dominant role. Specifically, inside each LSTM variant, a well-designed gating mechanism is introduced to enhance the language representation via the corresponding auxiliary modality. Furthermore, in the unimodal representation learning stage, instead of using RNNs, we introduce `channel-aware' temporal convolution network to extract high-level representations for each modality to explore both temporal and channel-wise interdependencies. Extensive experiments demonstrate that our approach achieves very competitive performance compared to the state-of-the-art methods on three widely-used benchmarks for multimodal sentiment analysis and emotion recognition. Sijie Mai, Songlong Xing, Haifeng Hu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal FusionabstractLearning joint embedding space for various modalities is of vital importance for multimodal fusion. Mainstream modality fusion approaches fail to achieve this goal, leaving a modality gap which heavily affects cross-modal fusion. In this paper, we propose a novel adversarial encoder-decoder-classifier framework to learn a modality-invariant embedding space. Since the distributions of various modalities vary in nature, to reduce the modality gap, we translate the distributions of source modalities into that of target modality via their respective encoders using adversarial training. Furthermore, we exert additional constraints on embedding space by introducing reconstruction loss and classification loss. Then we fuse the encoded representations using hierarchical graph neural network which explicitly explores unimodal, bimodal and trimodal interactions in multi-stage. Our method achieves state-of-the-art performance on multiple datasets. Visualization of the learned embeddings suggests that the joint embedding space learned by our method is discriminative. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
AAAI | 3 |
| 2020 | Efficient and Fast Real-World Noisy Image Denoising by Combining Pyramid Neural Network and Two-Pathway Unscented Kalman FilterabstractRecently, image prior learning has emerged as an effective tool for image denoising, which exploits prior knowledge to obtain sparse coding models and utilize them to reconstruct the clean image from the noisy one. Albeit promising, these prior-learning based methods suffer from some limitations such as lack of adaptivity and failed attempts to improve performance and efficiency simultaneously. With the purpose of addressing these problems, in this paper, we propose a Pyramid Guided Filter Network (PGF-Net) integrated with pyramid-based neural network and Two-Pathway Unscented Kalman Filter (TP-UKF). The combination of pyramid network and TP-UKF is based on the consideration that the former enables our model to better exploit hierarchical and multi-scale features, while the latter can guide the network to produce an improved (a posteriori) estimation of the denoising results with fine-scale image details. Through synthesizing the respective advantages of pyramid network and TP-UKF, our proposed architecture, in stark contrast to prior learning methods, is able to decompose the image denoising task into a series of more manageable stages and adaptively eliminate the noise on real images in an efficient manner. We conduct extensive experiments and show that our PGF-Net achieves notable improvement on visual perceptual quality and higher computational efficiency compared to state-of-the-art methods. Ruijun Ma 0001, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Image Process. | 3 |
| 2020 | Locally Confined Modality Fusion Network With a Global Perspective for Multimodal Human Affective ComputingabstractIn this paper, we propose a novel multimodal fusion framework, called the locally confined modality fusion network (LMFN), that contains a bidirectional multiconnected LSTM (BM-LSTM) to address the multimodal human affective computing problem. In the LMFN, we introduce a generic fusion structure that explores both local and global fusion to obtain an integral comprehension of information. Specifically, we partition the feature vector corresponding to each modality into multiple segments and learn every local interaction through a tensor fusion procedure. Global interaction is then modeled by learning the dependence between local tensors via an originally designed BM-LSTM architecture, establishing a direct connection of cells and states of local tensors that are several time steps apart. With the LMFN, we achieve advantages over other methods in the following aspects: 1) local interactions are successfully modeled using a feasible vector segmentation procedure that can explore cross-modal dynamics in a more specialized manner; 2) global interactions are modeled to obtain an integral view of multimodal information using BM-LSTM, which guarantees an adequate flow of information; and 3) our general fusion structure is highly extendable by applying other local and global fusion methods. Experiments show that the LMFN yields state-of-the-art results. Moreover, the LMFN achieves higher efficiency compared to other models by applying the outer product as the fusion method. Sijie Mai, Songlong Xing, Haifeng Hu 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Constrained LSTM and Residual Attention for Image CaptioningabstractVisual structure and syntactic structure are essential in images and texts, respectively. Visual structure depicts both entities in an image and their interactions, whereas syntactic structure in texts can reflect the part-of-speech constraints between adjacent words. Most existing methods either use visual global representation to guide the language model or generate captions without considering the relationships of different entities or adjacent words. Thus, their language models lack relevance in both visual and syntactic structure. To solve this problem, we propose a model that aligns the language model to certain visual structure and also constrains it with a specific part-of-speech template. In addition, most methods exploit the latent relationship between words in a sentence and pre-extracted visual regions in an image yet ignore the effects of unextracted regions on predicted words. We develop a residual attention mechanism to simultaneously focus on the pre-extracted visual objects and unextracted regions in an image. Residual attention is capable of capturing precise regions of an image corresponding to the predicted words considering both the effects of visual objects and unextracted regions. The effectiveness of our entire framework and each proposed module are verified on two classical datasets: MSCOCO and Flickr30k. Our framework is on par with or even better than the state-of-the-art methods and achieves superior performance on COCO captioning Leaderboard. Haifeng Hu 0001, Songlong Xing, Xinlong Lu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Divide, Conquer and Combine: Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affective ComputingabstractWe propose a general strategy named 'divide, conquer and combine' for multimodal fusion.Instead of directly fusing features at holistic level, we conduct fusion hierarchically so that both local and global interactions are considered for a comprehensive interpretation of multimodal embeddings.In the 'divide' and 'conquer' stages, we conduct local fusion by exploring the interaction of a portion of the aligned feature vectors across various modalities lying within a sliding window, which ensures that each part of multimodal embeddings are explored sufficiently.On its basis, global fusion is conducted in the 'combine' stage to explore the interconnection across local interactions, via an Attentive Bi-directional Skipconnected LSTM that directly connects distant local interactions and integrates two levels of attention mechanism.In this way, local interactions can exchange information sufficiently and thus obtain an overall view of multimodal information.Our method achieves state-ofthe-art performance on multimodal affective computing with higher efficiency. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
ACL (1) | 3 |