Xiaobao Guo

dblp:246/5846 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-3427-8540ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment
abstract
Propelled by the breakthrough in deep generative models, audio-to-image generation has emerged as a pivotal cross-modal task that converts complex auditory signals into rich visual representations. However, previous works only focus on single-source audio inputs for image generation, ignoring the multi-source characteristic in natural auditory scenes, thus limiting the performance in generating comprehensive visual content. To bridge this gap, we propose a method called MACS to conduct multi-source audio-to-image generation. To our best knowledge, this is the first work that explicitly separates multi-source audio to capture the rich audio components before image generation. MACS is a two-stage method. In the first stage, multi-source audio inputs are separated by a weakly supervised method, where the audio and text labels are semantically aligned by casting into a common space using the large pre-trained CLAP model. We introduce a ranking loss to consider the contextual significance of the separated audio signals. In the second stage, effective image generation is achieved by mapping the separated audio signals to the generation condition using only a trainable adapter and a MLP layer. We preprocess the LLP dataset as the first full multi-source audio-to-image generation benchmark. The experiments are conducted on multi-source, mixed-source, and single-source audio-to-image generation tasks. The proposed MACS outperforms the current state-of-the-art methods in 17 out of the 21 evaluation indexes on all tasks and delivers superior visual quality.
Xiaobao Guo, Yuzhe Zhu, Adams Wai-Kin Kong
AAAI2
2024 Flexible-Modal Deception Detection with Audio-Visual Adapter
abstract
Deception detection within audio-visual modalities is vital across diverse sectors, notably in customs security and multimedia anti-fraud. However, this notable efficacy is lost by the necessity to train and deploy separate models for each conceivable modality scenario, leading to redundancy and inefficiency. Moreover, real-world environments where multi-modal models are deployed often fail to meet these idealized conditions. To overcome these challenges and further elevate performance levels, we propose an advanced Transformer-based framework complemented by an Audio-Visual Adapter (AVA) integrating temporal features from both audio and visual modalities. In addition, we introduce an innovative multi-modal contrastive learning method that is designed to enhance the correlation between uni-modal features and their integrated counterparts within a consistent feature space. Our designed method can deal with the flexible-model scenario instead of deploying different models for various modalities. Empirical evaluations conducted on two benchmark datasets have validated the superiority of our proposed model over other multi-modal fusion techniques, particularly in scenarios characterized by varying and missing modalities. This strongly affirms the effectiveness of our approach in significantly boosting the accuracy of deception detection in complex, real-world multi-modal scenarios. The codes will be released soon.
Zhaoxu Li, Zitong Yu, Xun Lin, Nithish Muthuchamy Selvaraj, Xiaobao Guo, Bingquan Shen, Adams Wai-Kin Kong, Alex Chichung Kot
IJCB5
2023 Audio-Visual Deception Detection: DOLOS Dataset and Parameter-Efficient Crossmodal Learning
abstract
Deception detection in conversations is a challenging yet important task, having pivotal applications in many fields such as credibility assessment in business, multimedia anti-frauds, and custom security. Despite this, deception detection research is hindered by the lack of high-quality deception datasets, as well as the difficulties of learning multimodal features effectively. To address this issue, we introduce DOLOS1, the largest gameshow deception detection dataset with rich deceptive conversations. DOLOS includes 1, 675 video clips featuring 213 subjects, and it has been labeled with audio-visual feature annotations. We provide train-test, duration, and gender protocols to investigate the impact of different factors. We benchmark our dataset on previously proposed deception detection approaches. To further improve the performance by fine-tuning fewer parameters, we propose Parameter-Efficient Crossmodal Learning (PECL), where a Uniform Temporal Adapter (UT-Adapter) explores temporal attention in transformer-based architectures, and a crossmodal fusion module, Plug-in Audio-Visual Fusion (PAVF), combines crossmodal information from audio-visual features. Based on the rich fine-grained audio-visual annotations on DOLOS, we also exploit multi-task learning to enhance performance by concurrently predicting deception and audiovisual features. Experimental results demonstrate the desired quality of the DOLOS dataset and the effectiveness of the PECL. The DOLOS dataset and the source codes are available at here.
Xiaobao Guo, Nithish Muthuchamy Selvaraj, Zitong Yu, Adams Wai-Kin Kong, Bingquan Shen, Alex Chichung Kot
ICCV1
2023 Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
Nithish Muthuchamy Selvaraj, Xiaobao Guo, Adams Wai-Kin Kong, Bingquan Shen, Alex Chichung Kot
INTERSPEECH2
2023 Deep Multimodal Sequence Fusion by Regularized Expressive Representation Distillation
abstract
Multimodal sequence learning aims to utilize information from different modalities to enhance overall performance. Mainstream works often follow an intermediate-fusion pipeline, which explores both modality-specific and modality-supplementary information for fusion. However, the unaligned and heterogeneously distributed multimodal sequences pose significant challenges to the fusion task: 1) to extract both effective unimodal and crossmodal representations and 2) to overcome the overfitting issue in joint multimodal sequence optimization. In this work, we propose regularized expressive representation distillation (RERD) that aims to seek effective multimodal representations and to enhance the generalization of fusion. First, to improve unimodal representation learning, unimodal representations are assigned to multi-head distillation encoders, where the unimodal representations are iteratively updated through distillation attention layers. Second, to alleviate the overfitting issue in joint crossmodal optimization, a multimodal sinkhorn distance regularizer is proposed to reinforce the expressive representation extraction and to reduce the modality gap before fusion adaptively. These representations produce a comprehensive view of the multimodal sequences, which are utilized for downstream fusion tasks. Experimental results on several popular benchmarks demonstrate that the proposed method achieves state-of-the-art performance, compared with widely used baselines for deep multimodal sequence fusion, as shown inhttps://github.com/Redaimao/RERD.
Xiaobao Guo, Adams Wai-Kin Kong, Alex Chichung Kot
IEEE Trans. Multim.1
2023 Pace-Adaptive and Noise-Resistant Contrastive Learning for Multimodal Feature Fusion
abstract
Multimodal feature fusion aims to draw complementary information from different modalities to achieve better performance. Contrastive learning is effective at discriminating coexisting semantic features (positive) from irrelative ones (negative) in multimodal signals. However, positive and negative pairs learn at separate rates, which undermines the overall performance of multimodal contrastive learning (MCL). Moreover, the learned representation model is not robust, as MCL utilizes supervision signals from potentially noisy modalities. To address these issues, a novel multimodal contrastive learning objective, Pace-adaptive and Noise-resistant Noise-Contrastive Estimation (PN-NCE), is proposed for multimodal fusion by directly using unimodal features. PN-NCE encourages the positive and negative pairs reaching to their optimal similarity scores adaptively and shows less susceptibility to noisy inputs during training. A theoretical analysis is performed on its robustness. Maximizing modality invariance information in the fused representation is expected to benefit the overall performance and therefore, an estimator that measures the difference between the fused representation and its unimodal representations is integrated into MCL to obtain a more modality-invariant fusion output. The proposed method is model-agnostic and can be adapted to various multimodal tasks. It also bears less performance degradation when reducing the number of training samples at the linear probing stage. With different networks and modality inputs from three multimodal datasets, experimental results show that PN-NCE achieves consistent enhancements compared with previous state-of-the-art approaches.
Xiaobao Guo, Alex Chichung Kot, Adams Wai-Kin Kong
IEEE Trans. Multim.1
2021 Unimodal and Crossmodal Refinement Network for Multimodal Sequence Fusion
abstract
Effective unimodal representation and complementary crossmodal representation fusion are both important in multimodal representation learning.Prior works often modulate one modal feature to another straightforwardly and thus, underutilizing both unimodal and crossmodal representation refinements, which incurs a bottleneck of performance improvement.In this paper, Unimodal and Crossmodal Refinement Network (UCRN) is proposed to enhance both unimodal and crossmodal representations.Specifically, to improve unimodal representations, a unimodal refinement module is designed to refine modality-specific learning via iteratively updating the distribution with transformer-based attention layers.Self-quality improvement layers are followed to generate the desired weighted representations progressively.Subsequently, those unimodal representations are projected into a common latent space, regularized by a multimodal Jensen-Shannon divergence loss for better crossmodal refinement.Lastly, a crossmodal refinement module is employed to integrate all information.By hierarchical explorations on unimodal, bimodal, and trimodal interactions, UCRN is highly robust against missing modality and noisy data.Experimental results on MOSI and MOSEI datasets illustrated that the proposed UCRN outperforms recent state-of-the-art techniques and its robustness is highly preferred in real multimodal sequence fusion scenarios.Codes will be shared publicly 1 .
Xiaobao Guo, Adams Wai-Kin Kong
EMNLP (1)1
2020 Towards Photo-Realistic Virtual Try-On by Adaptively Generating↔Preserving Image Content
abstract
Image visual try-on aims at transferring a target clothes image onto a reference person, and has become a hot topic in recent years. Prior arts usually focus on preserving the character of a clothes image (e.g. texture, logo, embroidery) when warping it to arbitrary human pose. However, it remains a big challenge to generate photo-realistic try-on images when large occlusions and human poses are presented in the reference person. To address this issue, we propose a novel visual try-on network, namely Adaptive Content Generating and Preserving Network (ACGPN). In particular, ACGPN first predicts semantic layout of the reference image that will be changed after try-on (e.g.long sleeve shirt→arm, arm→jacket), and then determines whether its image content needs to be generated or preserved according to the predicted semantic layout, leading to photo-realistic try-on and rich clothes details. ACGPN generally involves three major modules. First, a semantic layout generation module utilizes semantic segmentation of the reference image to progressively predict the desired semantic layout after try-on. Second, a clothes warping module warps clothes image according to the generated semantic layout, where a second-order difference constraint is introduced to stabilize the warping process during training.Third, an inpainting module for content fusion integrates all information (e.g. reference image, semantic layout, warped clothes) to adaptively produce each semantic part of human body. In comparison to the state-of-the-art methods, ACGPN can generate photo-realistic images with much better perceptual quality and richer fine-details.
Ruimao Zhang, Xiaobao Guo, Wei Liu 0005, Wangmeng Zuo, Ping Luo 0002
CVPR3
2020 DRPL: Deep Regression Pair Learning for Multi-Focus Image Fusion
abstract
In this paper, a novel deep network is proposed for multi-focus image fusion, named Deep Regression Pair Learning (DRPL). In contrast to existing deep fusion methods which divide the input image into small patches and apply a classifier to judge whether the patch is in focus or not, DRPL directly converts the whole image into a binary mask without any patch operation, subsequently tackling the difficulty of the blur level estimation around the focused/defocused boundary. Simultaneously, a pair learning strategy, which takes a pair of complementary source images as inputs and generates two corresponding binary masks, is introduced into the model, greatly imposing the complementary constraint on each pair and making a large contribution to the performance improvement. Furthermore, as the edge or gradient does exist in the focus part while there is no similar property for the defocus part, we also embed a gradient loss to ensure the generated image to be all-in-focus. Then the structural similarity index (SSIM) is utilized to make a trade-off between the reference and fused images. Experimental results conducted on the synthetic and real-world datasets substantiate the effectiveness and superiority of DRPL compared with other state-of-the-art approaches. The testing code can be found in https://github.com/sasky1/DPRL.
Jinxing Li 0003, Xiaobao Guo, Guangming Lu 0002, Bob Zhang 0001, Yong Xu 0001, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.2
2019 Mask-Most Net: Mask Approximation Based Multi-oriented Scene Text Detection Network
abstract
In this paper, a novel multi-task cascade framework, which jointly takes the detection and the segmentation into account, is presented for the scene text detection. To address the issue of multi-oriented scene text detection, we propose an instance-level mask approximation method through the auxiliary regression task on center and corner points. Specifically, the text instance in the image is first coarsely detected, followed by a contextual module which can capture more accurate instances. To cope with the scale variation existing in these detected instances, a combination of high-level semantic and low-level features is further exploited, achieving more robust and better performance. A series of experiments conducted on different benchmark datasets demonstrate the effectiveness of the proposed method.
Xiaobao Guo, Jinxing Li 0003, Bingzhi Chen, Guangming Lu 0002
ICME1
2019 Single Image Reflection Removal Based on Deep Residual Learning
Zhixin Xu, Xiaobao Guo, Guangming Lu 0002
PRCV (2)2