Guibo Luo

dblp:160/7201 · DBLP profile ↗
← Back
55ranked-venue papers
5as first author
43since 2021 · last 2026
0000-0002-1709-1207ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 2 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Feature-Aware One-Shot Federated Learning via Hierarchical Token Sequences
abstract
One-shot federated learning (OSFL) reduces the communication cost and privacy risks of iterative federated learning by constructing a global model with a single round of communication. However, most existing methods struggle to achieve robust performance on real-world domains such as medical imaging, or are inefficient when handling non-IID (Independent and Identically Distributed) data. To address these limitations, we introduce FALCON, a novel framework that enhances the effectiveness of OSFL over non-IID image data. The core idea of FALCON is to leverage the feature-aware hierarchical token sequences generation and knowledge distillation into OSFL. First, each client leverages a pretrained visual encoder with hierarchical scale encoding to compress images into hierarchical token sequences, which capture multi-scale semantics. Second, a multi-scale autoregressive transformer generator is used to model the distribution of these token sequences and generate the synthetic sequences. Third, clients upload the synthetic sequences along with the local classifier trained on the real token sequences to the server. Finally, the server incorporates knowledge distillation into global training to reduce reliance on precise distribution modeling. Experiments on medical and natural image datasets validate the effectiveness of FALCON in diverse non-IID scenarios, outperforming the best OSFL baselines by 9.58\% in average accuracy.
Shudong Liu 0008, Hanwen Zhang 0007, Yuesheng Zhu, Guibo Luo
AAAI5
2026 DSFedMed: Dual-Scale Federated Medical Image Segmentation via Mutual Distillation Between Foundation and Lightweight Models
abstract
Foundation Models (FMs) have demonstrated strong generalization across diverse vision tasks. However, their deployment in federated settings is hindered by high computational demands, substantial communication overhead, and significant inference costs. We propose DSFedMed, a dual-scale federated framework that enables mutual knowledge distillation between a centralized foundation model and lightweight client models for medical image segmentation. To support knowledge distillation, a set of high-quality medical images is generated to replace real public datasets, and a learnability-guided sample selection strategy is proposed to enhance efficiency and effectiveness in dual-scale distillation. This mutual distillation enables the foundation model to transfer general knowledge to lightweight clients, while also incorporating client-specific insights to refine the foundation model. Evaluations on five medical imaging segmentation datasets show that DSFedMed achieves an average 2 percent improvement in Dice score while reducing communication costs and inference time by nearly 90 percent compared to existing federated foundation model baselines. These results demonstrate significant efficiency gains and scalability for resource-limited federated deployments.
Hanwen Zhang 0007, Qiaojin Shen, Yuxi Liu 0005, Yuesheng Zhu, Guibo Luo
AAAI5
2026 SeLaR: Selective Latent Reasoning in Large Language Models
abstract
Chain-of-Thought (CoT) has become a cornerstone of reasoning in large language models, yet its effectiveness is constrained by the limited expressiveness of discrete token sampling.Recent latent reasoning approaches attempt to alleviate this limitation by replacing discrete tokens with soft embeddings (probabilityweighted mixtures of token embeddings) or hidden states, but they commonly suffer from two issues: (1) global activation injects perturbations into high-confidence steps, impairing reasoning stability; and (2) soft embeddings quickly collapse toward the highest-probability token, limiting exploration of alternative trajectories.To address these challenges, we propose SeLaR (Selective Latent Reasoning), a lightweight and training-free framework.Se-LaR introduces an entropy-gated mechanism that activates soft embeddings only at lowconfidence steps, while preserving discrete decoding at high-confidence steps.Additionally, we propose an entropy-aware contrastive regularization that pushes soft embeddings away from the highest-probability token's direction, encouraging sustained exploration of multiple latent reasoning paths.Experiments on five reasoning benchmarks demonstrate that SeLaR consistently outperforms standard CoT and state-of-the-art training-free methods.
Renyu Fu, Guibo Luo
ACL (1)2
2026 LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model
abstract
Large Language Models (LLMs) often produce unnecessarily lengthy reasoning traces, which significantly increase computational cost and latency.Existing approaches typically rely on fixed length penalties, but such penalties are hard to tune and fail to adapt to the evolving reasoning abilities of LLMs, leading to suboptimal trade-offs between accuracy and conciseness.To address this challenge, we propose LEASH (adaptive LEngth penAlty and reward SHaping), a reinforcement learning framework for efficient reasoning in LLMs.We formulate length control as a constrained optimization problem and employ a Lagrangian primal-dual method to dynamically adjust the penalty coefficient.When generations exceed the target length, the penalty is intensified; when they are shorter, it is relaxed.This adaptive mechanism guides models toward producing concise reasoning without sacrificing task performance.Experiments on Deepseek-R1-Distill-Qwen-1.5B and Qwen3-4B-Thinking-2507 show that LEASH reduces the average reasoning length by 60% across diverse tasks-including indistribution mathematical reasoning and outof-distribution domains such as coding and instruction following-while maintaining competitive performance.Our work thus presents a practical and effective paradigm for developing controllable and efficient LLMs that balance reasoning capabilities with computational budgets.
Yanhao Li, Jiaran Zhang, Lexiang Tang, Guibo Luo
ACL (1)6
2026 HAR-SCP: Fine-Grained Advertising Law Auditing from Platform URLs with Multimodal Evidence Centralization and Statutory Constraint Propagation
Huayang Hsu, Yuesheng Zhu, Guibo Luo, Hongye Liu
ICIC (16)3
2026 Generalized Face Recognition With Occlusion
abstract
Face recognition under occlusion remains challenging due to masks, glasses, and other real-world obstructions that partially conceal facial information. Existing approaches typically rely on training with one or more predefined occlusion types, which limits their ability to generalize to unseen scenarios. In this letter, we propose Generalized Face Recognition with Occlusion (GFRO), an occlusion-robust framework that generalizes to diverse occlusion patterns without requiring occlusion-specific training data. GFRO is trained on partial facial views cropped from complete face images using a cross-entropy loss to learn generic representations across different partial views. A dual mixture-of-experts aggregator is then introduced to refine and integrate features from multiple branches, each handling a specific partial-view representation. Optimized with a clean–noisy contrastive loss, the aggregator aligns partial-face features with complete-face features, where each branch contains experts specializing in complementary partial-view information. Extensive experiments on multiple datasets demonstrate that GFRO generalizes effectively to both real and synthetic occlusion scenarios and achieves state-of-the-art performance compared with methods trained on occlusion-specific data under the same occlusion conditions.
Dengwen Zhang, Yuxi Liu 0005, Guibo Luo, Zhenyu Weng
IEEE Signal Process. Lett.3
2026 Federated Learning for Medical Image Classification: A Comprehensive Benchmark
abstract
The federated learning (FL) paradigm is well-suited for the field of medical image analysis, as it can effectively cope with machine learning on isolated multi-center data while protecting the privacy of participating parties. However, current research on optimization algorithms in FL often focuses on limited datasets and scenarios, primarily centered around natural images, with insufficient comparative experiments in medical contexts. In this work, we conduct a comprehensive evaluation of several state-of-the-art FL algorithms in the context of medical imaging. We conduct a fair comparison of classification models trained using various FL algorithms across multiple medical imaging datasets. Additionally, we evaluate system performance metrics, such as communication cost and computational efficiency, while considering different FL architectures. Our findings show that medical imaging datasets pose substantial challenges for current FL optimization algorithms. No single algorithm consistently delivers optimal performance across all medical FL scenarios, and many optimization algorithms may under-perform when applied to these datasets. Our experiments provide a benchmark and guidance for future research and application of FL in medical imaging contexts. Furthermore, we propose an efficient and robust method that combines generative techniques using denoising diffusion probabilistic models with label smoothing to augment datasets, widely enhancing the performance of FL on classification tasks across various medical imaging datasets. Our codes are released on GitHub, offering a reliable and comprehensive benchmark for future FL studies in medical imaging.
Zhekai Zhou, Guibo Luo, Zhenyu Weng, Yuesheng Zhu
IEEE J. Biomed. Health Informatics2
2025 Robust Image Hashing Based on Contrastive Masked Autoencoder with Weak-Strong Augmentation Alignment
abstract
Recently, numerous robust image hashing schemes have been developed for content identification. However, many of these schemes face the challenges of maintaining discrimination while simultaneously resisting large-scale attacks. In this paper, we propose a robust image hashing scheme based on Contrastive Masked Autoencoder with weak-strong augmentation Alignment (CMAA). Leveraging contrastive learning, CMAA is designed to learn features that are robust to large-scale and hybrid attacks while maintaining the discrimination of those features. Specifically, it utilizes distribution divergence to align weak attack augmented features with strong attack augmented features, namely weak-strong augmentation alignment, to enhance the robustness to strong attacks. In addition, a masked vision transformer is incorporated to further enhance content identification performance. CMAA also includes a parameter-free quantization layer to mitigate the loss induced by binarization. Experimental results demonstrate that our method exhibits remarkable robustness against various attacks, including challenging ones such as rotation and hybrid attacks, and delivers excellent identification performance with a F1 score close to 1.0. Our code and supplementary materials are available on Github.
Cundian Yang, Guibo Luo, Yuesheng Zhu, Xiyao Liu 0001
AAAI2
2025 Unraveling the Mystery: Defending Against Jailbreak Attacks Via Unearthing Real Intention
abstract
As Large Language Models (LLMs) become more advanced, the security risks they pose also increase. Ensuring that LLM behavior aligns with human values, particularly in mitigating jailbreak attacks with elusive and implicit intentions, has become a significant challenge. To address this issue, we propose a jailbreak defense method called Real Intentions Defense (RID), which involves two phases: soft extraction and hard deletion. In the soft extraction phase, LLMs are leveraged to extract unbiased, genuine intentions, while in the hard deletion phase, a greedy gradient-based algorithm is used to remove the least important parts of a sentence, based on the insight that words with smaller gradients have less impact on its meaning. We conduct extensive experiments on Vicuna and Llama2 models using eight state-of-the-art jailbreak attacks and six benchmark datasets. Our results show a significant reduction in both Attack Success Rate (ASR) and Harmful Score of jailbreak attacks, while maintaining overall model performance. Further analysis sheds light on the underlying mechanisms of our approach.
Yanhao Li, Hongshen Chen, Zhiwei Ge, Sulong Xu, Guibo Luo
COLING7
2025 A Physical Backdoor Attack Against Practical Federated Learning
abstract
Federated learning (FL) is a distributed machine learning paradigm designed to build a global model while safeguarding data privacy. However, the decentralized nature and data heterogeneity of FL increase its susceptibility to backdoor attacks. Existing works on backdoor attacks and defenses typically focus on digital attacks that use digitally generated patterns as backdoor triggers. While using digitally generated patterns as triggers is convenient, its practicality remains questionable. In this work, a physical backdoor attack is proposed to enhance the practicality of backdoor attacks in FL. We employ real-world objects as backdoor triggers, integrating Backdoor-Robust Training Set Selection and Prox Attack, to execute backdoor attacks across various scenarios. Extensive experiments show that the physical backdoor attack in FL pose a serious real-world threat. This finding underscores the urgent need for more robust backdoor defenses in the physical world.
Yang Li 0034, Xiangtao Lu, Guibo Luo, Yuesheng Zhu
CSCWD3
2025 MFDF-IML: Multi-Feature Dynamic Fusion for Image Manipulation Localization
abstract
With the rapid advancement of image tampering techniques, the authenticity of multimedia content is increasingly challenged, necessitating the development of robust Image Manipulation Localization (IML) technologies. This paper introduces a novel approach, Multi-Feature Dynamic Fusion for Image Manipulation Localization (MFDF-IML), which ad-dresses the limitations of existing methods by integrating Error Level Analysis (ELA) as a new forgery feature. ELA enhances feature diversity and robustness by capturing subtle variations through analyzing discrepancies at different compression levels. Additionally, MFDF-IML employs a dynamic gating mechanism to adaptively fuse multiple features, including SRM, Bayar, Noiseprint++, and ELA, adjusting their weights according to various forgery scenarios. This method also integrates features from Convolutional Neural Networks (CNN) and Vision Transformers (ViT), leveraging CNN's local feature extraction and ViT's global dependency modeling to significantly improve forgery localization precision. Extensive experiments demonstrate that MFDF-IML outperforms existing methods across diverse forgery scenarios, highlighting its potential in image forensics.
Xiangtao Lu, Fuyuan Cheng, Guibo Luo, Yuesheng Zhu
CSCWD3
2025 RGB Single-Channel Frequency Domain Backdoor Attack on Image Manipulation Localization
abstract
In the contemporary digital era, the advancement and pervasive adoption of image manipulation technologies present unprecedented challenges to the authenticity of information. The extensive application of deep learning techniques in image processing has made detecting forged images a critical task for ensuring information security. Nonetheless, research in this field often neglects a potential risk: the security of training datasets. Previous backdoor attacks typically introduce triggers through semantic or noise manipulation, causing models to focus on these elements. Models designed to locate manipulated regions, however, often need to focus on high-frequency edges or artifacts, leading to suboptimal performance when distracted by these triggers. This paper introduces a novel Single-Channel RGB Frequency-Domain Backdoor Attack (SC-FDI), designed to manipulate the RGB channels by introducing triggers into the training data. This manipulation hinders the trained model from accurately identifying forged regions in images. Specifically, the original image is first separated into its RGB channels, then the Discrete Cosine Transform (DCT) is applied to one of the channels, the DCT coefficients in the high-frequency region are amplified, and the image is reassembled after inverse transformation. Experimental results on multiple datasets indicate that our method surpasses existing approaches in both stealthiness and attack efficacy.
Xiangtao Lu, Yang Li 0034, Guibo Luo, Yuesheng Zhu
CSCWD3
2025 DeformSleepNet: Adaptive Multi-Level Feature Extraction for Enhanced Sleep Stage Classification
abstract
Sleep quality plays a crucial role in human health, with sleep stages serving as key indicators of overall well-being. Despite the advancements in sleep stage classification, existing methods encounter several limitations: 1) the use of fixed-size convolution kernels restricts the adaptive extraction of multi-scale features, and 2) redundancy in multi-modal signals results in inadequate attention to highly correlated features. In this paper, we introduce DeformSleepNet, a novel architecture that combines standard and deformable convolutions to enable adaptive multi-scale feature extraction across various frequency bands. The Multi-level Feature Encoder (MFE) update different features in different manner, achieving a trade-off between accuracy and efficiency, effectively balancing accuracy and computational efficiency. Additionally, the Content-Aware Fusion Module (CAFM) dynamically adjusts the importance of different features during fusion, thereby enhancing the overall classification performance. Experimental evaluations conducted on three publicly available datasets demonstrate that DeformSleepNet outperforms main-stream methods. Ablation studies further validate the contribution of each individual module, highlighting the robustness and effectiveness of the proposed approach.
Haoyang Xu, Yuesheng Zhu, Guibo Luo
CSCWD3
2025 Robust Image Hashing Based on Mixture-of-Experts with Hard-Sample Mining Contrastive Learning
abstract
In collaborative system, a large amount of digital image transmission requires a reliable content identification system to ensure the authenticity and copyright of these images. Robust image hashing schemes could efficiently extract robust features of images without manipulating the original image, making them an effective solution for content identification. However, existing schemes face the challenge of resisting complex attacks, such as diverse attack types and their sophisticated combinations, in the real world while maintaining discrimination. The complexity of attacks is reflected in the diversity of attack types, the variation in attack intensity and the combination of multiple attacks. In this study, we propose a robust image hashing approach based on Mixture-of-experts with Hard-sample mining contrastive learning (MiHa). In particular, MiHa utilizes a mixture-of-experts architecture to adaptively extract features from the image under various attacks, including hybrid attacks, thereby enhancing the robustness. Additionally, large-scale attacks can generate hard samples for contrastive learning. To address this, we propose a hard-sample mining contrastive loss that assigns greater weight to these hard samples, thereby further improving the performance of MiHa. Extensive experiments demonstrate that our method achieves superior robustness against various attacks while maintaining competitive discrimination.
Cundian Yang, Guibo Luo, Yuesheng Zhu, Xiyao Liu 0001
CSCWD2
2025 A New One-Shot Federated Learning Framework for Medical Imaging Classification with Feature-Guided Rectified Flow and Knowledge Distillation
abstract
In multi-center scenarios, One-Shot Federated Learning (OSFL) has attracted increasing attention due to its low communication overhead, requiring only a single round of transmission. However, existing generative model-based OSFL methods suffer from low training efficiency and potential privacy leakage in the healthcare domain. Additionally, achieving convergence within a single round of model aggregation is challenging under non-Independent and Identically Distributed (non-IID) data. To address these challenges, in this paper a modified OSFL framework is proposed, in which a new Feature-Guided Rectified Flow Model (FG-RF) and Dual-Layer Knowledge Distillation (DLKD) aggregation method are developed. FG-RF on the client side accelerates generative modeling in medical imaging scenarios while preserving privacy by synthesizing feature-level images rather than pixel-level images. To handle non-IID distributions, DLKD enables the global student model to simultaneously mimic the output logits and align the intermediate-layer features of client-side teacher models during aggregation. Experimental results on three non-IID medical imaging datasets show that our new framework and method outperform multi-round federated learning approaches, achieving up to 21.73% improvement, and exceed the baseline FedISCA by an average of 21.75%. Furthermore, our experiments demonstrate that feature-level synthetic images significantly reduce privacy leakage risks compared to pixel-level synthetic images. The code is available at https://github.com/LMIAPC/one-shot-fl-medical.
Hanwen Zhang 0007, Qiya Yang, Guibo Luo, Yuesheng Zhu
ECAI4
2025 FLIDS-Mamba: Multi-Center Network Intrusion Detection Based on Federated Learning and Bidirectional State Space Model
Denghua Li, Yuesheng Zhu, Guibo Luo
ICA3PP (7)3
2025 SpikingPoint: Rethinking Point as Spike for Efficient 3D Point Cloud Analysis
abstract
Spiking Neural Networks (SNNs), due to their unique spike-based inference mechanism, offer low power consumption and biological plausibility. As a fundamental technology for various real-world applications, 3D point cloud analysis faces significant challenges related to high computational overhead and energy-intensive. In fact, each point can be viewed as a specialized spike data containing positional information in 3D space. Therefore, utilizing the spiking features of SNNs to represent point clouds holds significant potential. However, this exploration faces two main challenges: SNN-compatible architecture design and spike-based 3D spatial modeling. In this work, we introduce the SpikingPoint, a pure Spiking Multi-layer Perceptron (MLP) Architecture that leverages the low power consumption of SNNs and the computational efficiency of linear layers. Furthermore, we propose the spiking 3D Position Embedding (SPE) that effectively models the 3D spatial feature into spike-form features. SpikingPoint achieves competitive performance with a small number of parameters while reducing energy consumption. For example, when achieving similar accuracy, the theoretical energy consumption of our method is reduced by 97.1% compared to the mainstream ANN-based KPConv. With fewer parameters, SpikingPoint surpasses the state-of-the-art SNN-based P2SResLNet on ModelNet40 by 1.42% in accuracy.
Zhaokun Zhou, Yijie Lu, Jiaqiyu Zhan, Guibo Luo, Yuesheng Zhu
ICASSP4
2025 OTTER: Optimized Training with Trustworthy Enhanced Replication via Diffusion and Federated VMUNet for Privacy-Aware Medical Segmentation
Haocheng Kan, Yuesheng Zhu, Guibo Luo, Hanwen Zhang 0007
ICICS (2)3
2025 MedDiff-FT: Data-Efficient Diffusion Model Fine-Tuning with Structural Guidance for Controllable Medical Image Synthesis
Jianhao Xie, Zhenyu Weng, Yuesheng Zhu, Guibo Luo
MICCAI (4)5
2025 AgentFactory: Towards Automated Agentic System Design and Optimization
Enci Zhang, Yuesheng Zhu, Xiaole Cui, Guibo Luo
PRICAI5
2025 Non-IID Medical Image Segmentation Based on Cascaded Diffusion Model for Diverse Multi-Center Scenarios
abstract
Learning from multi-center medical datasets to obtain a high-performance global model is challenging due to the privacy protection and data heterogeneity in healthcare systems. Current federated learning approaches are not efficient enough to learn Non-Independent and Identically Distributed (Non-IID) data and require high communication costs. In this work, a practical privacy computing framework is proposed to train a Non-IID medical image segmentation model under various multi-center setting in low communication cost. Specifically, an efficient cascaded diffusion model is trained to generate image-mask pairs that have similar distribution to the training data of clients, providing rich labeled data on client side to mitigate heterogeneity. Also, a label construction module is developed to improve the quality of generated image-mask pairs. Moreover, a set of aggregation methods is proposed to achieve global model from data generated from Cascaded Diffusion model for diverse scenarios: CD-Syn, CD-Ens and its extension CD-KD. CD-Syn is a one-shot method that trains segmentation model solely on public generated datasets while CD-Ens and CD-KD maximize the utilization of local original data by an extra communication round of ensemble or knowledge distillation. In this way, the setting of our proposed framework is highly practical, providing multiple aggregation methods which can flexibly adapt to varying demands for efficiency, privacy, and accuracy. We systematically evaluated the effectiveness of our proposed framework on five Non-IID medical datasets and observe 5.38% improvement in Dice score compared with baseline method (FednnU-Net) on average.
Hanwen Zhang 0007, Yuxi Liu 0005, Guibo Luo, Yuesheng Zhu
IEEE J. Biomed. Health Informatics4
2024 Cross-Stage Transfer in Multi-Stage Cascade Ranking and Filtering Systems
abstract
Unsupervised Domain Adaptation (UDA) aims to transfer a model from a labeled source domain to an unlabeled target domain, addressing challenges of distinct data distributions, termed domain shift. Existing UDA research primarily focuses on classification-like tasks, but neglects ranking and filtering tasks essential for applications like medical diagnosis and search engines. This paper is the first to notice and identify a new real-world transfer problem: cross-stage transfer in multi-stage cascade ranking and filtering systems, a common issue in diverse applications, including information retrieval systems, medical diagnosis, and other real-world ranking/filtering systems. In this problem, we emphasize the crucial assumption of order-invariance and address the key issue named Cross-stage Class Concept Conflict (C4), highlighting potential inconsistencies in class concepts for the same sample at different stages. To tackle these challenges, we propose a novel method, Unsupervised Rank Adaptation (URA), comprising two key components: order-conditional distribution alignment, characterizing the order-conditional distribution intra-stage and aligning them across stages; and principal projection alignment, aligning the principal component’s projection matrix with classifier parameters to ensure order-invariance without guessing pseudo-labels, mitigating the influence of C4. Experimental results show that our approach reaches state-of-the-art performance in various cross-stage transfer tasks.
Yifan Pan, Guibo Luo, Yuesheng Zhu
ECAI2
2024 AdaIFL: Adaptive Image Forgery Localization via a Dynamic and Importance-Aware Transformer Network
Yuxi Li 0008, Fuyuan Cheng, Wangbo Yu, Guangshuo Wang, Guibo Luo, Yuesheng Zhu
ECCV (43)5
2024 Generalizable Deepfake Detection with Unbiased Feature Extraction and Low-Level Forgery Enhancement
Zhihan Yu, Guangshuo Wang, Yuesheng Zhu, Guibo Luo
ICANN (2)5
2024 Enhanced Unsupervised Domain Adaptation with Dual-Attention Between Classification and Domain Alignment
abstract
Unsupervised Domain Adaptation (UDA) deals with transferring knowledge from labeled source domains to unlabeled target domains. This addresses the challenge of different distributions across domains, commonly known as domain shift. Numerous methods attempt to align distributions across domains while learning the core tasks (e.g., classification) on source domain separately. However, limited research has explored the mutual influence between classification and domain alignment. In this paper, we discuss the conflicting optimization between domain alignment and classification tasks, emphasizing the risk of negative transfer due to conflicting optimization directions. For better optimization consistency, these tasks should concentrate on the common information of features. To address this issue, we propose an innovative framework Dual-attention between classification and Domain Alignment (DuDA). DuDA employs gradient-based saliency maps to generate interpretable attentions, concurrently enhancing both classification and domain alignment through a dual-attention mechanism. Experimental results verify the effectiveness of DuDA in mitigating negative transfer and its strong adaptability and promising performance.
Yifan Pan, Guibo Luo, Bairong Li, Yuesheng Zhu
ICASSP2
2024 LoTraNet: Locality-Guided Transformer Network for Image Manipulation Localization
Fuyuan Cheng, Yuxi Li 0008, Xiangtao Lu, Guibo Luo, Yuesheng Zhu
ICONIP (7)4
2024 TFCM: Tuning-Free Facial Concept-Erasure in Text-to-Image Models Through Attention and Sample Modulation
Guibo Luo, Zhiqiang Bai, Yuesheng Zhu
ICONIP (7)2
2024 Fine-Grained Transformer Encoder of Image-Text Retrieval for Streaming Data in Cross-Modal Continual Learning
Jiaqiyu Zhan, Yuesheng Zhu, Guibo Luo
ICONIP (9)4
2024 MambaFuse: Fusing Multi-scale Mamba and CNN Features for Seizure Prediction
Ding Pan, Guibo Luo, Yuesheng Zhu
ICONIP (4)2
2024 Seizure Prediction Based on Multi-scale Fusion-Attention Transformer
Ding Pan, Guibo Luo, Yuesheng Zhu
ICONIP (5)2
2024 OKey: Towards More Controllable, Secure and Robust Diffusion Model Image Steganography Using Optimized Key
Dongxu Yue, Guibo Luo, Yuesheng Zhu
ICONIP (6)2
2024 DACOA: Diffusion-Aligned Coherent Augmentation and Consistency Constraint Strategies for Federated Domain Generalization
Guangshuo Wang, Yuesheng Zhu, Guibo Luo
ICPR (27)3
2024 Multi-Scale Image Stitching Based on Unsupervised Controllable Fusion Structure
abstract
Traditional feature-based image stitching techniques rely heavily on the quality of feature detection and often fail to get a high quality solution for images with fewer features. Generating high-quality stitching images with a natural structure is still a challenging task in computer vision. The current distortion-based image alignment relies on the division of the grid, which is currently mostly equally divided without reference to the information in the image itself, and thus there are some problems such as straight line retention in the alignment stage. On this basis, if the deep learning approach is added in the later reconstruction stage to limit the generation of artifacts at the pixel level, it will be greatly improved, and we can ensure the high quality of the final splicing result of the image through two stages. In this paper, we succeed in proposing a novelty model called multi-scale image stitching based on unsupervised controllable fusion module which can be divided in two stages: Multi-scale grids alignment structure based on feature points density and image reconstruction structure based on unsupervised controllable fusion structure. We develop new energy terms to adapt to limit the transformations of the new grids. Extensive experiments demonstrate that the proposed method outperforms most state-of-the-arts by effectively preserving the linear structure in the image and improving the robustness.
Wangbo Yu, Yuesheng Zhu, Guibo Luo
IJCNN4
2024 High Resolution Face Privacy-Enhancing method based on Latent Optimization with Identity-Preserving Facial Masking
abstract
The widespread application of face recognition technology easily leads to the leakage of facial soft biometric privacy. Therefore, image-level face privacy-enhancing methods for face recognition scenarios aim to remove soft biometric privacy information while preserving individual identity. Existing methods suffer from performance degradation in face recognition and poor quality in privacy-enhanced images when processing high-resolution images, often resulting in artifacts and complexity in model optimization. To address these issues, we propose an identity-preserving face privacy-enhancing method based on latent optimization and face masking for high-resolution images. The proposed method generates privacy-enhanced faces by optimizing latent codes and mask blending, without relying on a well-defined encoder-decoder structure, making it more suitable for processing high-resolution images. We introduce a face masking strategy and design a facial elements identity influence prioritization module to adaptively select privacy-enhancing facial parts and attributes. As a result, the proposed method can protect multiple facial privacy attributes while improving the quality of generated images and maintaining face recognizability. Extensive experiments on two benchmark datasets demonstrate the effectiveness of the proposed method in anonymizing visual appearances for high-resolution face images while maintaining superior face recognition performance and image quality.
Haoxuan Tai, Bofei Guo, Guibo Luo, Yuesheng Zhu
IJCNN4
2024 Data Augmentation in Class-Conditional Diffusion Model for Semi-Supervised Medical Image Segmentation
abstract
Accurate segmentation of specific organs or diseased tissues in medical images is crucial for precise diagnosis and effective treatment planning. While fully-supervised deep learning methods have demonstrated remarkable performance, their effectiveness heavily relies on the availability of a substantial number of labeled images. Unfortunately, acquiring and manually labeling a large medical dataset is often expensive and impractical, especially for rare diseases, due to challenges related to data sharing and privacy. To address this limitation, a class-conditional diffusion model is proposed to synthesize realistic medical images, thereby augmenting small-scale datasets with high-quality samples. The diffusion-based method generates an unlimited number of authentic medical images, each conditioned on specific class labels, offering a valuable contribution to dataset augmentation strategies. To evaluate the utility of the generated data, synthetic images are incorporated into the genuine labeled dataset, thereby creating augmented datasets. By utilizing the generated data as unlabeled augmented data for the original dataset, our synthetic data is effectively integrated with semi-supervised medical image segmentation algorithms. This integration successfully combines the diversity of synthetic images with the advantages of semi-supervised learning, resulting in a more comprehensive utilization of limited small-scale data. Our experimental results on two public datasets have demonstrated that our class-conditional diffusion model can generate high-fidelity synthetic medical images. Furthermore, it enables semi-supervised methods to achieve superior segmentation performance in comparison to the methods using only original real datasets.
Guibo Luo, Yuesheng Zhu
IJCNN2
2024 FedFMS: Exploring Federated Foundation Models for Medical Image Segmentation
Yuxi Liu 0005, Guibo Luo, Yuesheng Zhu
MICCAI (8)2
2024 CodeDetector: Revealing Forgery Traces with Codebook for Generalized Deepfake Detection
abstract
The malicious use of deepfake technologies poses a significant threat to social security, emphasizing the urgent necessity to advance deepfake detection. Existing detection models tend to overfit specific forgery traces within the training set, resulting in a weak generalization performance on unseen data. Considering that the feature distribution of real images is limited compared to the diverse forgery patterns, capturing the real feature space and treating features outside the distribution as potential forgery traces can mitigate overfitting and enhance the detector's generalization ability. Furthermore, since facial manipulation generates artifacts in pixels and disrupts the global consistency of the data, potential forgery traces at both pixel and feature levels can contribute to detection. In this paper, we propose a novel two-stage deepfake detection model named CodeDetector, which utilizes a codebook to capture the feature space of real faces and obtain potential forgery traces by facial reconstruction for detection. In the codebook learning stage, a codebook is used to capture the real feature distribution. In the detector training stage, we obtain pixel-level and feature-level residuals as potential forgery traces through facial reconstruction on real and fake faces to guide the model's attention to forgery clues. Specifically, we propose a Quantized Residual-Guided Attention Module and a Dual Residual Attention Module, calculating residuals and utilizing the attention mechanism to enhance global feature representation. Additionally, an Indices Prediction Module is introduced to ensure the accuracy of residual guidance that enhances the robustness of reconstruction during detection. Extensive experiments have demonstrated that CodeDetector outperforms state-of-the-art in deepfake detection cross-dataset benchmark.
Zhihan Yu, Guibo Luo, Yuesheng Zhu
ICMR3
2024 Class Incremental Learning With Deep Contrastive Learning and Attention Distillation
abstract
Class incremental learning can solve the issue of catastrophic forgetting when the trained model is used to learn a new task in which the model may forget part of previous knowledge learned. The key issue is to alleviate the stability-plasticity dilemma and maintain the balance between preventing old knowledge from being forgotten and learning new knowledge. In this paper, a new class incremental learning method with deep contrastive learning and attention distillation is proposed. The deep contrastive learning can optimize last layer of the model to learn strong-semantic information and the intermediate layers to learn weak-semantic information when new data is coming so that the plasticity can be improved. To ensure a balance between plasticity and stability, a new attention distillation approach is developed to learn and distill the latent attention information of the feature maps and the similarity between different samples. Our analysis and experimental results indicate that the proposed method can improve the learning and memorization ability and the average accuracy of the model compared with the previous methods, and obtain competitive results in class increment scenarios for image classification.
Jitao Zhu, Guibo Luo, Baishan Duan, Yuesheng Zhu
IEEE Signal Process. Lett.2
2024 Adaptive Face Recognition for Multi-Type Occlusions
abstract
Due to the prevalence of influenza outbreaks and outdoor scenarios with various obstructing decorations, recognizing faces with occlusions has become a pressing challenge to address. However, current research mainly focuses on facial recognition with one kind of occlusion and does not provide compatible solutions for different kinds of common occlusions like glasses, sunglasses, and masks. Therefore, an Adaptive Multi-Type Occluded Face Recognition Model (AMOFR) is proposed to effectively handle multiple occlusion types simultaneously in this paper. In AMOFR, a generator is developed to produce diverse occluded face images for training, achieved by simulating various occlusion types on unoccluded face images. Subsequently, an occlusion type-based adapter is formulated to address a range of occlusion scenarios, guided by prompts from a Visual-Language model. To enhance overall performance by leveraging complete facial information, a feature-level knowledge distillation loss function is implemented, facilitating joint learning of unoccluded-face and occluded-face features. Furthermore, a new sunglasses-wearing dataset (CALFW-SUNGLASSES) is generated for more comprehensive test for AMOFR and further occlusion recognition research. Experimental results on datasets containing different types of occlusions have demonstrated that AMOFR achieves significantly higher accuracy compared to other advanced face recognition models. The implementation codes of AMOFR is available athttps://github.com/LIU-YUXI/Adaptive-Multi-occlusion-Face-Recognition.
Yuxi Liu 0005, Guibo Luo, Zhenyu Weng, Yuesheng Zhu
IEEE Trans. Circuits Syst. Video Technol.2
2023 Machine Unlearning with Affine Hyperplane Shifting and Maintaining for Image Classification
Mengda Liu, Guibo Luo, Yuesheng Zhu
ICONIP (13)2
2023 bt-vMF Contrastive and Collaborative Learning for Long-Tailed Visual Recognition
abstract
Real-world data often exhibit long tail distributions with heavy class imbalance, where the majority (head) classes can dominate the training process and alter the decision boundaries of the minority (tail) classes, leading to biased feature spaces. Recently, researchers have investigated the potential of contrastive learning for long-tailed visual recognition and introduced a class-balanced factor in loss function engineering. Although this method can help improve performance, it harms head performance due to undesirable bias, resulting in poor separability of minority samples in feature spaces. In this paper, we target the logit adjustment and propose balanced student-t von Mises-Fisher (bt-vMF) contrastive learning, encouraging a large margin between the head and tail classes and providing better generalization. In addition, the network trained on long-tailed datasets suffers from great uncertainty in predictions. To alleviate this issue, we build mutual supervision among multiple experts via proposed bilateral collaborative learning (BCL), in which the collaboration is conducted from both bt-vMF similarity and relationship distillation. Simply put, our designs focus on the generalization power of a single expert and the knowledge transfer among multiple experts to alleviate the biased feature space and uncertainty in long-tailed learning, respectively. Experiments on multiple datasets show that our method achieves competitive performance on long-tailed visual recognition task.
Jinhao Du, Guibo Luo, Yuesheng Zhu, Zhiqiang Bai
ICTAI2
2023 Robust Steganography without Embedding Based on Secure Container Synthesis and Iterative Message Recovery
abstract
Synthesis-based steganography without embedding (SWE) methods transform secret messages to container images synthesised by generative networks, which eliminates distortions of container images and thus can fundamentally resist typical steganalysis tools. However, existing methods suffer from weak message recovery robustness, synthesis fidelity, and the risk of message leakage. To address these problems, we propose a novel robust steganography without embedding method in this paper. In particular, we design a secure weight modulation-based generator by introducing secure factors to hide secret messages in synthesised container images. In this manner, the synthesised results are modulated by secure factors and thus the secret messages are inaccessible when using fake factors, thus reducing the risk of message leakage. Furthermore, we design a difference predictor via the reconstruction of tampered container images together with an adversarial training strategy to iteratively update the estimation of hidden messages. This ensures robustness of recovering hidden messages, while degradation of synthesis fidelity is reduced since the generator is not included in the adversarial training. Extensive experimental results convincingly demonstrate that our proposed method is effective in avoiding message leakage and superior to other existing methods in terms of recovery robustness and synthesis fidelity.
Ziping Ma 0002, Yuesheng Zhu, Guibo Luo, Xiyao Liu 0001, Gerald Schaefer, Hui Fang 0003
IJCAI3
2021 Deep-Cleansing: Deep-Learning Based Electronic Cleansing in Dual-Energy CT Colonography
Guibo Luo, Michael E. Zalis, Wenli Cai
MICCAI (7)1
2020 A Disocclusion Inpainting Framework for Depth-Based View Synthesis
abstract
This paper proposes a disocclusion inpainting framework for depth-based view synthesis. It consists of four modules: foreground extraction, motion compensation, improved background reconstruction, and inpainting. The foreground extraction module detects the foreground objects and removes them from both depth map and rendered video; the motion compensation module guarantees the background reconstruction model to suit for moving camera scenarios; the improved background reconstruction module constructs a stable background video by exploiting the temporal correlation information in both 2D video and its corresponding depth map; and the constructed background video and inpainting module are used to eliminate the holes in the synthesized view. The analysis and experiment indicate that the proposed framework has good generality, scalability and effectiveness, which means most of the existing background reconstruction methods and image inpainting methods can be employed or extended as the modules in our framework. Our comparison results have demonstrated that the proposed framework achieves better synthesized quality, temporal consistency, and has lower running time compared to the other methods.
Guibo Luo, Yuesheng Zhu, Zhenyu Weng, Zhaotian Li
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Annular Sector Model for tracking multiple indistinguishable and deformable objects in occlusions
Guibo Luo, Zhenyu Weng, Yuesheng Zhu
Neurocomputing2
2018 A New Temporal Deconvolutional Pyramid Network for Action Detection
Xiangli Ji, Guibo Luo, Yuesheng Zhu
ACCV (4)2
2018 Classification of Bone Tumor on CT Images Using Deep Convolutional Neural Network
Yang Li 0034, Wenyu Zhou, Guiwen Lv, Guibo Luo, Yuesheng Zhu
ICANN (2)4
2018 A New Sparse Subspace Clustering by Rotated Orthogonal Matching Pursuit
abstract
Sparse Subspace Clustering (SSC) is one of the most popular clustering methods in computer vision. However, SSC solved by convex programming tools may suffer from noise and outliers. Orthogonal Matching Pursuit (OMP) can solve these problems effectively but may lose correct information representation and accuracy. To overcome these problems, a new Sparse Subspace Clustering method by Rotated Orthogonal Matching Pursuit (SSC-ROMP) is proposed in this paper, in which once some vector chooses another vector as its neighbor, it has to be rotated to avoid it being chosen when the selected vector chooses neighbors. Also, Nonnegative Matrix Factorization is used for dimension reduction in SSC-ROMP. The analysis and experiment results on synthetic data and several open datasets have demonstrated that our new approach can achieve better performances in terms of accuracy and information representation than other subspace clustering algorithms and keep a good sparsity of clustering representation.
Yuesheng Zhu, Guibo Luo
ICIP3
2018 Fast MRF-Based Hole Filling for View Synthesis
abstract
Hole filling is one of the key issues in generating virtual view from video-plus-depth sequence by depth-image-based rendering. Hole filling method based on Markov random fields (MRF) is a practical way for view synthesis, but the traditional ones might introduce some foreground textures to the hole regions, and suffer from high computational complexity. In this letter, a fast MRF-based hole filling method is proposed for view synthesis, which is formulated as an energy minimization problem and is solved with loopy belief propagation (LBP). The energy function is optimized by employing the depth information to prevent the foreground textures filling holes. Furthermore, the LBP process maintains the visual consistency in the synthesized view by reserving all useful candidate labels. In addition, efficient belief propagation strategy is developed to optimize the LBP process, whose computational complexity is reduced to be linear with the number of candidate labels. Experimental results demonstrate the effectiveness of the proposed method with low running time and good visual consistency.
Guibo Luo, Yuesheng Zhu
IEEE Signal Process. Lett.1
2017 Depth estimation for outdoor image using couple dictionary learning and region detection
abstract
Depth estimation from a single image is a significant and challenging task in computer vision. It is difficult to represent the correspondence between depth and RGB image without any prior information. Unlike previous approaches that only map the RGB images to the corresponding depth locally, we propose to combine the local depth with region-level and global scene structures. Firstly, the global layout is retrieved from similar images. Secondly, local depth estimated by the coupled dictionary learning (DL) formulation is combined with the global layout to maintain the global result. Finally, in order to further refine the depth estimated, we propose to detect the sky region in outdoor scene. In addition, several edge-preserving strategies are taken to clearly distinguish different objects. The experimental results demonstrate that our method represents a more realistic depth map than other methods on the popular public dataset Make3D.
Qiqi Yao, Guibo Luo, Yuesheng Zhu
VCIP2
2017 A near-duplicate 3D video detection algorithm by using hypercomplex representations
Ziqiang Sun, Yuesheng Zhu, Xiaomei Xing, Guibo Luo, Xiyao Liu 0001
Multim. Tools Appl.4
2017 Foreground Removal Approach for Hole Filling in 3D Video and FVV Synthesis
abstract
The depth-image-based rendering is a key technique for 3D video and free viewpoint video synthesis. One of the critical problems in current synthesis methods is that the background (BG) occluded by the foreground objects might be exposed in the new view, and some holes are produced in the synthesized video. However, most of the traditional hole-filling approaches may bring some blurry effect or artifacts in the virtual view. In this paper, a foreground removal approach for hole filling is proposed, in which the foreground objects are removed from both the 2D video and its corresponding depth map, and then a BG video and its depth map are generated before the 3D warping and used to eliminate the holes in the synthesized video. Moreover, a BG extension method is applied in the reference view to prevent the large holes occurring along the border areas in the virtual view. Our analysis and experimental results have indicated that the proposed approach has better performance compared with the other methods in terms of the quality of synthesized video, computational complexity, and running time in multiview synthesis or multiframe synthesis.
Guibo Luo, Yuesheng Zhu
IEEE Trans. Circuits Syst. Video Technol.1
2016 A Hole Filling Approach Based on Background Reconstruction for View Synthesis in 3D Video
abstract
The depth image based rendering (DIBR) plays a key role in 3D video synthesis, by which other virtual views can be generated from a 2D video and its depth map. However, in the synthesis process, the background occluded by the foreground objects might be exposed in the new view, resulting in some holes in the synthetized video. In this paper, a hole filling approach based on background reconstruction is proposed, in which the temporal correlation information in both the 2D video and its corresponding depth map are exploited to construct a background video. To construct a clean background video, the foreground objects are detected and removed. Also motion compensation is applied to make the background reconstruction model suitable for moving camera scenario. Each frame is projected to the current plane where a modified Gaussian mixture model is performed. The constructed background video is used to eliminate the holes in the synthetized video. Our experimental results have indicated that the proposed approach has better quality of the synthetized 3D video compared with the other methods.
Guibo Luo, Yuesheng Zhu, Zhaotian Li, Liming Zhang 0002
CVPR1
2014 A low-complexity visual tracking approach with single hidden layer neural networks
abstract
Visual tracking algorithms based on deep learning have robust performance against variations in a complex environment because deep learning can learn generic features from numerous unlabeled images. However, due to the multilayer architecture, the deep learning trackers suffer from expensive computational costs and are not suitable for real-time applications. In this paper, a low-complexity visual tracking scheme with single hidden layer neural network is proposed based on denoising autoencoder. To further reduce the computational costs, feature selection is applied to simplify the networks and two optimization methods are used during the online tracking process. The experimental results have demonstrated that the proposed algorithm is about six times faster than the trackers based on deep nets and rapid enough for real-time applications with encouraging accuracy.
Yuesheng Zhu, Guibo Luo
ICARCV3
2014 A digital blind watermarking scheme based on quantization index modulation in depth map for 3D video
abstract
3D video provides an immersive experience to viewers and is getting more and more popular. The solution to create 3D video from 2D video is low-cost compared with that captures 3D video directly, and the generation of depth map from 2D video is a key in the 2D-3D video conversion systems. Therefore, protection of depth map is vital for 3D video. In this paper, a digital blind watermarking scheme based on Quantization Index Modulation (QIM) algorithm is proposed in which the copyright information is embedded in the DCT coefficients of depth map imperceptibly. The experimental results show that the proposed scheme has good robustness against video attacks such as salt noise, median filtering, wiener filtering, and scaling. In the meanwhile, the stereo video embedded watermarking can accomplish zero distortion in comparison with the original one.
Yang Guan, Yuesheng Zhu, Xiyao Liu 0001, Guibo Luo, Ziqiang Sun, Liming Zhang 0002
ICARCV4