EDBT 2026 Demo / reviewers in the wild / expert
Xiaochen Yuan
dblp:97/7626
· DBLP profile ↗
101ranked-venue papers
5as first author
83since 2021 · last 2026
0000-0002-7490-6695ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 4 first-author · 35 since 2021Artificial intelligence and machine learning · 31 · 28 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 15 since 2021Security and privacy · 8 · 3 since 2021Computer networks · 7 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PASA: Progressive-Adaptive Spectral Augmentation for Automated Auscultation in Data-Scarce EnvironmentsabstractAutomated auscultation advances the detection of respiratory diseases, especially in areas with limited resources where traditional diagnostic methods are unavailable. On the other hand, the scarcity of auscultation datasets limits the automation performance, prompting the needs for data augmentation methods. However, most of the existing methods neglect the difference in acoustic sounds that requires personalized augmentation strategies. To address this, we propose a Progressive-Adaptive Spectral Augmentation (PASA), which is one of the first paradigms to adaptively select the best augmentation strategy for each sample. The PASA innovatively treats augmentation selection problem as a Markov Decision Process (MDP), creating an alternating loop between the diagnostic model and the augmentation selection. The agent selects the optimal augmentation operations and magnitudes via a task-specific design, including state construction, action sampling, Hybrid Batch-Sample (HBS) strategy execution, and reward guidance. The HBS strategy initially applies uniform augmentation across mini-batches while collecting sample-specific performance statistics. When model performance stabilizes, it transits to sample-level augmentation based on accumulated difficulty assessments. This two-phase design balances computational complexity with personalization. Extensive experiments across three benchmark datasets demonstrate that the PASA outperforms the state-of-the-art methods, pioneering a transformative paradigm for adaptive data augmentation in automated auscultation. Ying Wang 0097, Guoheng Huang, Xueyuan Gong, Xinxin Wang 0003, Xiaochen Yuan |
AAAI | 5 |
| 2026 | FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language ModelsabstractFace anti-spoofing (FAS) is crucial for protecting facial recognition systems from presentation attacks. Previous methods approached this task as a classification problem, lacking interpretability and reasoning behind the predicted results. Recently, multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and decision-making in visual tasks. However, there is currently no universal and comprehensive MLLM and dataset specifically designed for FAS task. To address this gap, we propose FaceShield, a MLLM for FAS, along with the corresponding pre-training and supervised fine-tuning (SFT) datasets, FaceShield-pre10K and FaceShield-sft45K. FaceShield is capable of determining the authenticity of faces, identifying types of spoofing attacks, providing reasoning for its judgments, and detecting attack areas. Specifically, we employ spoof-aware vision perception (SAVP) that incorporates both the original image and auxiliary information based on prior knowledge. We then use an prompt-guided vision token masking (PVTM) strategy to random mask vision tokens, thereby improving the model's generalization ability. We conducted extensive experiments on three benchmark datasets, demonstrating that FaceShield significantly outperforms previous deep learning models and general MLLMs on four FAS tasks, i.e., coarse-grained classification, fine-grained classification, reasoning, and attack localization. Hongyang Wang 0001, Zhuofu Tao, Yuhao Gao, Liepiao Zhang, Xun Lin, Xiaochen Yuan, Zitong Yu, Xiaochun Cao |
AAAI | 8 |
| 2026 | Cross-view Anchor Graph Learning and Factorization for Incomplete Multi-view ClusteringabstractGraph-based incomplete multi-view clustering algorithms have gathered much attention due to their impressive clustering performance. However, existing methods primarily leverage intra-view correlation from observed views, while ignoring the exploration of explicit compensation relationships between different views. Moreover, these methods need post-processing to get labels, and the separate steps lack negotiation, which may lead to sub-optimal solutions. To address these issues, we propose a Cross-view Anchor Graph Learning and Factorization (AGLF) method. AGLF develops an Anchor Graph Completion (AGC) framework that explicitly learn the missing subgraph structures. Instead of requiring post-processing, AGC directly produces soft labels. By establishing a third-order tensor of soft labels, it employs the tensor Schatten p-norm to enhance anchor graph learning and factorization. To significantly improve the quality of subgraph learning, AGLF incorporates compensation subgraphs from supplementary views into the AGC framework, enabling the construction of a better anchor graph for label learning. An optimization algorithm is devised to solve the objective function. Experimental results across various datasets demonstrate the effectiveness of our method. Xinxin Wang 0003, Yongshan Zhang, Xiaochen Yuan, Yicong Zhou |
AAAI | 3 |
| 2026 | SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action RecognitionabstractLarge Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs understand the skeleton? 2) How can LLMs distinguish among actions? To address these problems, we introduce a novel paradigm named learning Skeleton representation with visual-motion knowledge for Action Recognition (SUGAR). In our pipeline, we first utilize off-the-shelf large-scale video models as a knowledge base to generate visual, motion information related to actions. Then, we propose to supervise skeleton learning through this prior knowledge to yield discrete representations. Finally, we use the LLM with untouched pre-training weights to understand these representations and generate the desired action targets and descriptions. Notably, we present a Temporal Query Projection (TQP) module to continuously model the skeleton signals with long sequences. Experiments on several skeleton-based action classification benchmarks demonstrate the efficacy of our SUGAR. Moreover, experiments on zero-shot scenarios show that SUGAR is more versatile than linear-based methods. Qilang Ye, Yu Zhou 0015, Jie Zhang 0081, Xuanming Guo, Mingkui Tan, Weicheng Xie 0001, Yue Sun 0001, Tao Tan 0002, Xiaochen Yuan, Ghada Khoriba, Zitong Yu |
AAAI | 11 |
| 2026 | BiRNet: Bilateral-Attentive Refinement Network for Tampering Detection Across Manipulation ParadigmsabstractImage tampering detection faces persistent challenges in localizing manipulation boundaries across diverse forgery techniques, from traditional manipulation to AI-driven diffusion-based synthesis. Both manipulation types exhibit subtle boundary inconsistencies that demand robust feature extraction. However, existing attention mechanisms widely adopted in tampering detection suffer from limitations. Spatial attention produces diluted activations on weak traces, while channel attention amplifies noise from heterogeneous forensic patterns. To refine these widely-used attention mechanisms, we propose BiRNet, a Bilateral-Attentive Refinement Network featuring two key innovations: (1) a Bilateral Cross-Attention Module that enables bidirectional interaction between complementary features with multi-level fusion, and (2) an Attentive Refinement Module that enhances discrimination through parallel local-global pathways. We conduct extensive experiments across seven benchmarks spanning traditional datasets and diffusion-based manipulations. BiRNet achieves an average F1-score of 0.618 across all datasets. These results validate strong cross-domain generalization from conventional to AI-generated forgeries without requiring AIGC-specific training. Zhiyao Xie, Tong Liu 0021, Nuno Lourenço 0002, Xiaochen Yuan |
ICMR | 5 |
| 2026 | Mission-preserving protocol defense: An in-situ MAVLink honeypot with benchmark-calibrated thresholds and constant-space logging
Zigang Chen, Fulin Yu, Xiaochen Yuan, Haihua Zhu 0004 |
Comput. Commun. | 4 |
| 2026 | Decoding coefficients recovery based on modified Gauss-Jordan elimination for tampered content reconstruction
Tong Liu 0021, Lihao Zhuang, Guoheng Huang, Chi-Man Pun, Xiaochen Yuan |
Expert Syst. Appl. | 5 |
| 2026 | A noise-assistant network for tampering detection via inconspicuous feature enhancement and multi-perspective perception
Zhiyao Xie, Xiaochen Yuan, Chan-Tong Lam, Guoheng Huang, Nuno Lourenço 0002 |
Expert Syst. Appl. | 2 |
| 2026 | LAPWF-IoD: Lightweight Authentication Protocol Combining Passwords and Wireless Fingerprints for Internet of Drones in Smart City Environments
Zigang Chen, Xiaochen Yuan, Liuxin Chen, Linlin Liang |
IEEE Internet Things J. | 3 |
| 2026 | CCSFusion: A Hierarchical Semantic Chain-of-Thought Reasoning Architecture for Infrared-Visible Image Fusion and CaptioningabstractInfrared-Visible Image Fusion (IVIF) aims to generate a single, information-rich image for downstream tasks. However, prevailing methods exhibit two key limitations. First, many approaches lack explicit hierarchical semantic decoupling, failing to effectively integrate semantic features across different levels, which restricts their ability to capture complex scene structures. Second, task-driven fusion frameworks typically adopt a cascaded design, with unidirectional supervision provided by geometry-centric downstream tasks like detection. This architecture not only limits mutual reinforcement between the fusion and task networks, but also creates a ”supervision bottleneck” by lacking interaction with the linguistic modality that captures richer scene relationships. To tackle these challenges, we propose CCSFusion, the first framework that leverages Chain-of-Thought captioning as supervision, redirecting IVIF optimization from narrow geometric accuracy to multimodal scene comprehension. It establishes a mutually reinforcing coupling between the fusion network and the captioning task. Specifically, we introduce a Segmentation Mask Calibration Unit (SMCU) to refine coarse semantic priors, providing precise pixel-level guidance. Subsequently, the calibrated features are fed into Chained Semantic Fusion Module (CSFM) which explicitly decomposes the semantic priors into three hierarchical levels, and then feeds them into the Hierarchical Semantic Attention module. Finally, a bidirectional knowledge distillation mechanism transfers the reasoning ability of the teacher network to the student. Experiments show that CCSFusion achieves superior fusion performance and generates more semantically coherent images for high-level cognitive tasks. The code is available at: https://github.com/Snaillms/CCSFusion. Miaoshan Lin, Guoheng Huang, Jietao Yang, Jiehao Zheng, Xiaochen Yuan, Yan Li 0122, Xiaofeng Zhang 0006, Kim Fung Tsang, Chi-Man Pun |
IEEE Internet Things J. | 5 |
| 2026 | WiSACL: A Subdomain Adaptive Wi-Fi-Based Gesture Recognition via Contrastive LearningabstractWi-Fi-based gesture recognition faces significant performance degradation in cross-domain scenarios due to the distribution shifts between various environments, user locations, and facing orientations. To address this challenge, we propose WiSACL, a novel Wi-Fi-based gesture recognition framework that integrates innovative CSI signal processing with advanced subdomain adaptation techniques. We propose a novel Wi-Fi CSI phase difference representation that suppresses the environmental noise and highlights the motion features via phase ratio computation, wavelet denoising, and temporal differencing, which are then encoded into discriminative image representations for enhanced gesture recognition. Building upon this robust signal representation, we develop an Adaptive Distributioncalibrated Pseudo-Label (ADPL) module that progressively refines target labels through subdomain distribution alignment and confidence-aware selection. To fully exploit these pseudo-labels while mitigating the adverse effects of domain shift and label noise, we further introduce a contrastive learning (CL) model that explicitly enhances intra-class compactness and inter-class separation in the feature space. Extensive experiments on the Widar3.0 dataset demonstrate that WiSACL achieves an average accuracy of 98.62% across cross-location, cross-orientation, and cross-environment scenarios, significantly outperforming state-of-the-art methods and validating the effectiveness of our approach for robust cross-domain wireless gesture recognition. Yangjing Zhou, Xiaochen Yuan, Yue Liu 0001, Yuanwei Liu |
IEEE Internet Things J. | 3 |
| 2026 | The Latens Patronus: Seamless Model Watermarking for Latent Diffusion Model in IoT EnvironmentsabstractWith the rapid development of the Internet of Things (IoT), generative artificial intelligence has been widely applied in various IoT applications. However, the wide adoption of latent diffusion model (LDM) in such IoT scenarios raises severe risks of copyright infringement and model theft due to the lack of effective protection mechanisms. To address this challenge, we propose Latens Patronus, a seamless model watermarking technique for copyright protection of LDM in IoT environments. Unlike existing model watermarking methods, our method does not require additional watermark input and additional parameters for a specialized embedding network, making it more suitable for deployment in real-world IoT applications. Specifically, we design a Watermark Encoder to integrate image watermark into latent features during the generation process and aWatermark Decoder to accordingly extract the watermark from suspicious images accurately. We further introduce a diminishing training strategy that gradually fades out auxiliary supervision signals, eliminating the need for persistent watermark guidance, and adding no extra overhead to the original model. Extensive experiments on multiple LDM variants demonstrate that Latens Patronus outperforms existing watermarking methods in both invisibility and robustness against image-level and model-level attacks. Tong Liu 0021, Wei Ke 0001, Guoheng Huang, Xueyuan Gong, Xiaochen Yuan |
IEEE Internet Things J. | 6 |
| 2026 | Mutual sample-center interaction with hard queue mining for face recognition
Jianqing Li 0001, Xiaochen Yuan, Guanghua Yang, Xiaofan Li 0001, Xueyuan Gong |
Inf. Sci. | 3 |
| 2026 | MADAT: Missing-aware dynamic adaptive transformer model for medical prognosis prediction with incomplete multimodal data
Jianbin He, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Guo Zhong, Bai Ying Lei, Haojiang Li |
Medical Image Anal. | 3 |
| 2026 | QWNet: A quaternion wavelet network for spatial-frequency aware multi-modal image fusion
Jietao Yang, Miaoshan Lin, Guoheng Huang, Xuhang Chen 0002, Xiaofeng Zhang 0006, Xiaochen Yuan, Chi-Man Pun, Bingo Wing-Kuen Ling |
Neural Networks | 6 |
| 2026 | HINTS: Hierarchically Disentangling Subregional Heterogeneity With Structural Priors for Multi-Modal Survival Analysis
Biyun Chen, Guoheng Huang, Xiaochen Yuan, Yan Li 0122, Chi-Man Pun, Bai Ying Lei, Haojiang Li |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Interpretation Before Integration: LLM-Guided Multimodal Completion and Fusion Network for Survival Analysis With Incomplete Data
Feng Ling 0002, Haoming Zeng, Ming Li 0065, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Xianglian Liao, Jiong-Lin Liang, Haojiang Li |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Reversible Unlearnable Examples: Toward the Copyright Protection in Deep Learning EraabstractSignificant advancements in deep learning have been made possible by the utilization of large datasets, underscoring the critical importance of copyright protection. Adding meticulously designed perturbations to examples, making them unlearnable has become a crucial approach for safeguarding data copyright. Existing methods for creating unlearnable examples overlook the risk of data leakage, which can threaten data ownership. Thus, copyright protection in deep learning faces two main threats: illegal model training and malicious data leakage. We investigate that these two threats cannot be solved by straightforwardly combining existing availability attacks and watermarking techniques as their negative interaction effects. Therefore, in this paper, we propose a novel copyright protection mechanism for the aforementioned security concerns. Considering that the prevention of unauthorized model training requires powerful generalizability of unlearnable perturbations, we generate perturbations to induce the model to learn uncorrelated features of input images. It works by minimizing the mutual information of the input and output of the model. On the other hand, to eliminate the side impact of unlearnable perturbations on the watermark extraction, we design a dual extraction strategy by using two distinct watermark extractors. Extensive experiments on the image datasets ImageNet, CIFAR10, and Pets show that our proposed method could provide comprehensive copyright protection to images. The code is available at https://github.com/Yeah21/ReversibleUnlearnableExamples. Binze Wang, Jinyu Tian 0001, Xingrun Wang, Xiaochen Yuan, Jianqing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | DFFormer: Capturing Dynamic Frequency Features to Locate Image Manipulation Through Adaptive Frequency Transformer and Prototype LearningabstractThe proliferation of modern image editing tools has raised concerns about image manipulation, particularly regarding the potential to mislead the public and compromise privacy and security. Consequently, detecting and localizing tampered regions has become a critical research challenge. Traditional methods struggle with subtle manipulations, such as splicing, copy-move, and removal, which are often more discernible in the frequency domain than in the spatial domain. Additionally, the size imbalance between the tampered and background regions further complicates the detection process. To address these challenges, we propose DFFormer, an end-to-end network that leverages frequency feature differences and a dynamic token strategy for precise manipulation localization. DFFormer combines the Conventional Neural Network (CNN) and Transformer in a hybrid architecture with three key modules: the Adaptive Frequency Transformer (AFT), the Prototype Learning Module (PLM), and the Cascaded Progressive Token Fusion Head (CPTF-Head). AFT integrates high- and low-frequency components into self-attention via the Parallel Adaptive Frequency Attention (PAFA) block, enhancing tampering feature representation while preserving fine details. PLM employs KNN-based density peak clustering (DPC-KNN) and weighted token aggregation to optimize dynamic token reduction. The CPTF-Head adopts a hierarchical coarse-to-fine strategy to integrate multiscale features, thereby improving localization accuracy and edge refinement. Experiments demonstrate that DFFormer outperforms state-of-the-art models across four benchmark datasets and one real-world dataset, exhibiting superior generalization and robustness. The source code is publicly available at https://github.com/XiangGD/DFFormer.git. Kaiqi Zhao 0004, Zhenghong Yu, Xiaochen Yuan, Guoheng Huang, Jinyu Tian 0001, Jianqing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | SEM-UCSNet: A Novel Semantic Maps-Guided Compressive Sensing Framework for Underwater ImagesabstractUnderwater images (UWIs) captured by underwater detectors are essential for underwater detection and exploration. The compressive sensing theory (CS) provides a method for recovering images from few measurements, and it has been proven to be suitable for underwater environments with narrow bandwidth and limited communication channel resources, which may have a significant negative impact on the quality of captured UWIs. However, most existing state-of-art CS methods do not take the characteristics of UWIs into account, so their performance is limited in underwater applications. Compared with on-land images, UWIs have the following characteristics: 1) UWIs contain relatively few semantics, with a large amount of similar feature within the same semantics; 2) The importance of different semantics in UWIs is closely related to the underwater imaging model. In this paper, we combine the underwater imaging model and semantic of UWIs with CS task and propose a novel semantic maps-guided CS framework for UWIs, dubbed SEM-UCSNet, which can improve the performance of sampling and reconstruction, especially under extremely low sampling rate. In the sampling stage, a semantic importance analysis module combined with the imaging model is designed to guide the sampling. In the reconstruction process, a graph-based reconstruction strategy guided by semantic maps is proposed to model all features under the same semantic and mine complementarity between them to improve the reconstruction quality. Simultaneously, we introduce GAN into the underwater CS reconstruction task and use sampled features as conditions to make the reconstructed UWIs have richer details. Experimental results on some real-world UWIs datasets have demonstrated the superiority of our SEM-UCSNet on both objective and subjective metrics. Lihao Zhuang, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Fre-QNet: Quaternion Progressive Perception Mechanism with Frequency-Guided Prompt for Blind Image Quality AssessmentabstractBlind Image Quality Assessment faces challenges in enabling computational models to mimic the hierarchical progressive perception mechanisms of the Human Visual System (HVS). Existing methods often neglect the two-stage process of HVS—global distortion identification followed by local quality evaluation—and its distinct sensitivity to distortion types. To address this, we propose Fre-QNet, a novel framework integrating two key components: (1) A Quaternion Progressive Perception (QPP) module that hierarchically extracts multi-scale spatial features using quaternion convolution, explicitly simulating the global-to-local observation process of HVS while enhancing cross-scale interactions; (2) A Frequency Prompting (FP) module that quantifies distortion types and severity in the Fourier domain by leveraging frequency patterns of common distortions and the sensitivity variations of HVS. The QPP and FP modules collaboratively embed biological vision principles into computational modeling through dual-domain feature learning, with the QPP module directly anchoring the core logic of progressive perception. Experiments on TID2013 and CSIQ benchmarks demonstrate Fre-QNet’s superiority over state-of-the-art methods, validating its effectiveness in matching human perceptual quality judgments. Our source code is available at: https://github.com/hhsda/Fre-QNet . Shize Li, Guoheng Huang, Yisen Zheng, Xiaochen Yuan, Xuhang Chen 0002, Lianglun Cheng, Chi-Man Pun |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2026 | WAQNIQA: Wavelet-Augmented Quaternion Network for No-Reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) plays a pivotal role in computer vision by enabling image quality evaluation without reference images. While recent CNN and Transformer-based methods have advanced feature extraction, they face significant limitations. CNNs exhibit local feature bias, limiting their ability to capture global dependencies and complex structures critical to understanding diverse distortions. Transformers, despite modeling nonlocal dependencies through multihead attention, suffer from quadratic computational complexity with spatial dimensions, hindering efficient multiscale analysis. Moreover, their attention mechanisms frequently overlook critical interchannel dependencies, which are vital for capturing fine details in texture-rich images. Coupled with difficulties in handling high-noise environments and complex textures, this results in limited real-world accuracy and poor generalization across diverse datasets and unknown distortions. To bridge these gaps, we propose WAQNIQA, a novel wavelet-augmented quaternion network for NR-IQA. Distinct from conventional architectures, WAQNIQA integrates two synergistic modules: the wavelet-infused adaptive attention (WIAA) module, which leverages wavelet transforms (WTs) to achieve robust multiscale spatial-frequency analysis with linear complexity, and the quaternion collaborative feature enhancement (QCFE) module, which holistically models interchannel correlations to preserve fine texture details. Furthermore, we introduce PowerGridIQ, the first NR-IQA dataset specifically tailored for power grid scenarios. Extensive experiments demonstrate that WAQNIQA consistently surpasses state-of-the-art CNN and Transformer-based methods on PowerGridIQ and six public benchmarks. Notably, WAQNIQA exhibits superior cross-domain generalization, achieving competitive performance on the AGIQA-1K dataset for AI-generated content (AIGC) without explicit semantic alignment training, thereby validating its robustness against diverse and unknown distortions. Our code is available athttps://github.com/king-huoye/WAQNIQA Yejing Huo, Guoheng Huang, Zhiwen Yu 0002, Xiaochen Yuan, Chi-Man Pun, Lianglun Cheng, Xuhang Chen 0002, Zehong Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | DADM: Dual Alignment of Domain and Modality for Face Anti-SpoofingabstractWith the availability of diverse sensor modalities (i.e., RGB, Depth, Infrared) and the success of multi-modal learning, multi-modal face anti-spoofing (FAS) has emerged as a prominent research focus. The intuition behind it is that leveraging multiple modalities can uncover more intrinsic spoofing traces. However, this approach presents more risk of misalignment. We identify two main types of misalignment: (1) \textbf{Intra-domain modality misalignment}, where the importance of each modality varies across different attacks. For instance, certain modalities (e.g., Depth) may be non-defensive against specific attacks (e.g., 3D mask), indicating that each modality has unique strengths and weaknesses in countering particular attacks. Consequently, simple fusion strategies may fall short. (2) \textbf{Inter-domain modality misalignment}, where the introduction of additional modalities exacerbates domain shifts, potentially overshadowing the benefits of complementary fusion. To tackle (1), we propose a alignment module between modalities based on mutual information, which adaptively enhances favorable modalities while suppressing unfavorable ones. To address (2), we employ a dual alignment optimization method that aligns both sub-domain hyperplanes and modality angle margins, thereby mitigating domain gaps. Our method, dubbed \textbf{D}ual \textbf{A}lignment of \textbf{D}omain and \textbf{M}odality (DADM), achieves state-of-the-art performance in extensive experiments across four challenging protocols demonstrating its robustness in multi-modal domain generalization scenarios. The codes will be released soon. Xun Lin, Zitong Yu, Liepiao Zhang, Xin Liu 0012, Hui Li 0089, Xiaochen Yuan, Xiaochun Cao |
ICCV | 7 |
| 2025 | Lorentz Transformation Neural NetworkabstractWe propose a novel neural network architecture, the Lorentz Transformation Neural Network (LTNN), which utilizes Lorentz transformations to generate a complex computation matrix that enhances the network’s expressive power. Furthermore, LTNN is lightweight due to the shared weight matrices in the computation matrix. LTNN treats the input and output as coordinates in high-dimensional spacetime, with the weight matrices in each layer representing the velocity components of a spacetime reference frame. During training, these weight matrices are transformed into a computation matrix via Lorentz transformations, describing the coordinate transformations between different reference frames. We evaluate LTNN on four datasets: California Housing Prices, Iris, MNIST, and Fashion-MNIST. Experimental results demonstrate that LTNN outperforms conventional neural networks and quaternion neural networks in terms of both accuracy and parameter efficiency. Wenyuan Li 0007, Jingchao Wang 0002, Guoheng Huang, Tongxu Lin, Guo Zhong, Xiaochen Yuan, Chi-Man Pun, An Zeng |
ICIP | 6 |
| 2025 | SAM-FE: Segment Anything Model Guided Feature Enhancement for Semantic Change Detection of Remote Sensing ImagesabstractSemantic change detection (SCD) is a crucial research topic in remote sensing. To achieve high-precision semantic segmentation results, Segment Anything Model-Guided Feature Enhancement (SAM-FE) is proposed. SAM-FE utilizes Mobile-SAM to extract features from bi-temporal remote sensing images (RSIs). In addition, the cross-temporal feature aggregation module (CTFA), the multiscale contextual information fusion module (MCIF), and the change feature enhancement module (CFE) are utilized to enhance the general features and the representation of the change information of the RSIs, thus improving the accuracy of change detection. Experimental results indicate that SAM-FE significantly outperforms the existing methods in both the Second datasets and MusSCD, with F1 of 0.6241 and 0.8316, respectively. Meanwhile, SAM-FE maintains lower parameters, demonstrating its superiority and practicality. Junqing Huang, Tong Liu 0021, Chan-Tong Lam, Xiaochen Yuan |
ICME | 4 |
| 2025 | ISSD-NET: Intra-student Self-distillation with Adaptive q-vMF Loss for Enhanced Semi-supervised Medical Segmentation
Guoheng Huang, Xiaochen Yuan, Yan Li 0122, Chi-Man Pun, Bai Ying Lei |
ICONIP (2) | 3 |
| 2025 | Superpixel-Enhanced Quaternion Feature Fusion and Contextualization Graph Contrastive Learning for Cervical Cancer Diagnosis
Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Lianglun Cheng, Chi-Man Pun, Guo Zhong, Qingjian Ye |
ICONIP (2) | 3 |
| 2025 | Adaptive Query Prompting for Multi-Domain Landmark DetectionabstractMedical landmark detection is crucial in various medical imaging modalities and procedures. Although deep learning-based methods have achieve promising performance, they are mostly designed for specific anatomical regions or tasks. In this work, we propose a universal model for multi-domain landmark detection by leveraging transformer architecture and developing a prompting component, named as Adaptive Query Prompting (AQP). Transformer backbone architecture is suitable for our study due to its ability to capture long-range dependencies crucial in multi-domain landmark detection. Specifically, transformers excel at understanding global anatomical relationships, which span across entire images. Instead of embedding additional modules in the backbone network, we design a separate module to generate prompts that can be effectively extended to any other transformer network. In our proposed AQP, prompts are learnable parameters maintained in a memory space called prompt pool. The central idea is to keep the backbone frozen and then optimize prompts to instruct the model inference process. Furthermore, we employ a lightweight decoder to decode landmarks from the extracted features, namely Light-MLD. Thanks to the lightweight nature of the decoder and AQP, we can handle multiple datasets by sharing the backbone encoder and then only perform partial parameter tuning without incurring much additional cost. It has the potential to be extended to more landmark detection tasks. We conduct experiments on three widely used X-ray datasets for different medical landmark detection tasks. Our proposed Light-MLD coupled with AQP achieves SOTA performance on many metrics even without the use of elaborate structural designs or complex frameworks. Qiusen Wei, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Jianwen Huang |
IJCNN | 4 |
| 2025 | Code Retrieval with Mixture of Experts Prototype Learning Based on ClassificationabstractThe semantic connection between code and queries is crucial for code retrieval, but many human-written queries fail to accurately capture the code's core intent, leading to ambiguity.This ambiguity complicates the code search process, as the queries do not provide a clear overview of the code's purpose.Our analysis reveals that while ambiguous queries may not precisely summarize the intent of the code, they often share the same general topics as the corresponding code.In light of this discovery, we propose Code Retrieval with Mixture of Experts Prototype Learning Based on Classification (CRME), a novel approach that combines classification for prototype-based representation learning and result ensembling.CRME utilizes specialized pre-trained models focused on the specific domains of ambiguous queries.It consists of two key components: Multiple Classification Prototype and Representation Learning with a Prototype-based Multi-model Contrastive (PMC) Loss during training, and Multi-Prototype Mixture of Experts Integration (MP-MoE) module for fine-grained ensemble inference.Our method can effectively address the issue of query ambiguity and improves search precision.Experimental results on the CodeSearchNet dataset, covering six sub-datasets, show that CRME outperforms existing methods, achieving an average MRR score of * Corresponding authors. Feng Ling 0002, Guoheng Huang, Jingchao Wang 0002, Xiaochen Yuan, Xuhang Chen 0002, XueYong Zhang, Fanlong Zhang, Chi-Man Pun |
Internetware | 4 |
| 2025 | Cross-View Geo-Localization via Learning Correspondence Semantic Similarity Knowledge
Guanli Chen, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun |
MMM (1) | 3 |
| 2025 | The Structure-sharing Hypergraph Reasoning Attention Module for CNNs
Jingchao Wang 0002, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Tongxu Lin, Chi-Man Pun, Fenfang Xie |
Expert Syst. Appl. | 3 |
| 2025 | Dual-channel hypergraph networks in the time-frequency domain for learning advanced spatiotemporal dependencies in multivariate time series
Jianjian Jiang, Xiangmin Luo, Fangyuan Lei, Xiaochen Yuan, Jin Zhan |
Neurocomputing | 5 |
| 2025 | Using 3D-LMM-Based Encryption to Secure Digital Images With 3-D S-Box and Fibonacci Q-MatrixabstractThe rapid development of communication technology has significantly improved information transmission and increased capacity. In the context of the Internet of Things (IoT), where massive visual data transmission faces challenges of real-time processing and security threats. To address this problem, this paper proposes a novel encryption algorithm based on a 3D S-box combined with Fractal-Sort-Matrix (FSM) and Fibonacci Q-Matrix (3DSFF). To address the shortcomings of traditional S-box encryption, this paper integrates the 3D S-box with the FSM, thus enhancing the uncertainty of permutation and transformation. Additionally, to overcome the limitation of using a single value inciting Fibonacci Q-Matrix (FQM) applications, this paper combines FQM with chaotic sequences to strengthen its resistance against exhaustive attacks. Through comprehensive experimental tests on the algorithm’s outcomes, it demonstrates notable improvements over previous algorithms, achieving an average information entropy of 7.9993 and a correlation coefficient close to 0.01 after encryption. These tests indicate that the scheme can withstand common attacks, making it a sufficiently secure solution for the confidentiality of private images. Moreover, these advances provide a secure and efficient visual data protection framework for IoT applications involving anti-theft surveillance, data acquisition, and communication transmission. Yunlong Liao, Qiutong Li, Guoheng Huang, Donald Donglong Chen, Xiaochen Yuan |
IEEE Internet Things J. | 7 |
| 2025 | Visual-linguistic Diagnostic Semantic Enhancement for medical report generation
Jiahong Chen, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Zhe Tan, Chi-Man Pun |
J. Biomed. Informatics | 3 |
| 2025 | LED-Net: A lightweight edge detection network
Shucheng Ji, Xiaochen Yuan, Junqi Bao, Tong Liu 0021 |
Pattern Recognit. Lett. | 2 |
| 2025 | FastFace: Fast-Converging Scheduler for Large-Scale Face Recognition Training With One GPUabstractComputing power has evolved into a foundational and indispensable resource in the area of deep learning, particularly in tasks such as Face Recognition (FR) model training on large-scale datasets, where multiple GPUs are often a necessity. Recognizing this challenge, some FR methods have started exploring ways to compress the fully-connected layer in FR models. Unlike other approaches, our observations reveal that without prompt scheduling of the learning rate (LR) during FR model training, the loss curve tends to exhibit numerous stationary subsequences. To address this issue, we introduce a novel LR scheduler leveraging Exponential Moving Average (EMA) and Haar Convolutional Kernel (HCK) to eliminate stationary subsequences, resulting in a significant reduction in converging time. However, the proposed scheduler incurs a considerable computational overhead due to its time complexity. To overcome this limitation, we propose FastFace, a fast-converging scheduler with negligible time complexity, i.e.O(1) per iteration, during training. In practice, FastFace is able to accelerate FR model training to a quarter of its original time without sacrificing more than 1% accuracy, making large-scale FR training feasible even with just one single GPU in terms of both time and space complexity. Extensive experiments validate the efficiency and effectiveness of FastFace. The code is publicly available at: https://github.com/amoonfana/FastFace. Xueyuan Gong, Zhiquan Liu 0001, Yain-Whar Si, Xiaochen Yuan, Ke Wang 0068, Xiaoxiang Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | TransHFC: Joints Hypergraph Filtering Convolution and Transformer Framework for TemporalForgery LocalizationabstractThe authenticity of audio-visual content is being challenged by advanced multimedia editing technologies inspired by Artificial Intelligence-Generated Content (AIGC). Temporal forgery localization aims to detect suspicious contents by locating forged segments. So far, most of the existing methods are based on Convolutional Neural Networks (CNNs) or Transformers, yet neither of them has fully considered the complex relationships within forged audio-visual content. To address this issue, in this paper, we propose a novel method, named TransHFC, which innovatively introduces hypergraphs to model group relationships among segments while considering point-to-point relationships through Transformers. Through its dual hypergraph filtering convolution branch, TransHFC captures both temporal and spatial level group relationships, enhancing the representation of forged segment features. Furthermore, we propose a new hypergraph filtering convolution Auto-Encoder that uses a multi-frequency filter bank for adaptive signal capture. This design compensates for the limitation of a single hypergraph filter. Our extensive experiments on Lav-DF, TVIL, Psynd, and HAD datasets demonstrate that TransHFC achieves state-of-the-art performance. Xiaochen Yuan, Chan-Tong Lam, Sio Kei Im, Fangyuan Lei, Xiuli Bi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Let Images Speak More: An Efficient Method for Detecting Image Manipulation HistoryabstractDigital image forensics aims to verify the authenticity of digital images, which has emerged as a prominent research area. To reveal the manipulation history of an image, the existing methods can only detect specific image operations or are based on a general forensic feature with high dimensions. Moreover, these methods perform well only when the operation chain length is no greater than 2. However, their detection accuracy drops significantly for images with longer operation chains that are more representative of real-world scenarios. To break these limitations, we proposed a novel forensics frequency Feature based on Histogram and Detail Map (FHDM(79D)), which can distinguish various operation chains containing different numbers of operations. Specifically, compared to the traces left by image manipulation in the spatial domain, we have discovered that they are more distinct in the frequency domain. This observation has prompted us to extract features from the frequency domain of images by analyzing their histograms and detail maps to capture the manipulation traces of the images. Notably, the proposed feature extracted in the frequency domain has almost 90% fewer dimensions than the commonly used general forensic features, such as SRM(714D), which greatly reduces the computational complexity. Meanwhile, compared to deep learning-based methods, the experiments show that the proposed method achieves a detection accuracy of over 95% for image operations across multiple datasets, while other deep learning-based methods do not exceed 90% accuracy. Extensive experimental results show that the proposed method is more versatile and effective, showing good performance in complex operation chain detection and local forgery detection. The code is available at https://github.com/CherishL-J/Op-detection. Yang Wei 0002, Xiaochen Yuan, Xiuli Bi, Bin Xiao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | PFL-ALP: Personalized Federated Learning Against Backdoor Attacks via Attention-Based Local PurificationabstractFederated learning (FL) enables collaborative model training with local data privacy preserving, but is vulnerable to backdoor attacks from malicious clients. These attacks can manipulate the global model to produce malicious output when encountering specific triggers. Existing defenses, categorized as server-side and client-side approaches, have limitations such as reliance on auxiliary data availability, susceptibility to inference attacks, and instability under non-independent and identically distributed (Non-IID) data. In response to these challenges, we propose a Personalized Federated Learning via Attention-based Local Purification (PFL-ALP) algorithm, a hybrid defense mechanism integrating server-side dynamic clustering and client-side purification enhanced with personalized model knowledge. This approach effectively mitigates bias introduced by Non-IID data on the server side and further purifies the backdoored model on the client side. Specifically, we employ neural attention distillation (NAD) for model purification and enhance it with personalized model knowledge, extending the effectiveness of NAD in Non-IID FL settings. This design makes PFL-ALP compatible with privacy protocols to mitigate inference attacks. Moreover, we establish a convergence guarantee for PFL-ALP and experimentally validate its superior performance in defending against various backdoor attacks compared to multiple state-of-the-art (SOTA) defenses across three datasets. The results show that even with malicious rates ranging from 30% to 90%, PFL-ALP can reduce the attack success rate by more than 69.4 percentage points, with the reduction in main task accuracy less than 12.4 percentage points. Yifeng Jiang 0008, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | A Symmetric Self-Embedding Mechanism for High-Fidelity Image Recovery Against TamperingabstractDigital images are inherently fragile and vulnerable to malicious tampering, significantly compromising their authenticity and integrity. Image recovery is crucial for restoring altered content and preserving the reliability of digital images. Traditional fragile watermarking methods achieve high-quality recovery but fail under post-processing attacks, while existing deep learning-based approaches offer some robustness, yet often produce lower-quality recovered images, typically with a PSNR of around 28 dB. To address these challenges, we propose a novel Symmetric Self-embedding Mechanism for High-Fidelity Image Recovery against tampering (SSEM-HIR), which is capable of restoring tampered images with high quality while maintaining some robustness against common attacks. Unlike existing methods that use the fragility of watermarking solely for tampering localization, SSEM-HIR is the first work to integrate fragility with spatial symmetry, enabling high-quality tampering recovery. Specifically, our SSEM-HIR employs a hierarchical watermark embedding module to embed an inverted version of the original image, utilizing spatial symmetry to retrieve lost information from the extracted watermark. To further improve recovery quality, we design a Dual-branch Region-based Self-Recovery module, where a Spatial-based Watermark Extraction block restores tampered regions using embedded watermark information, while a Frequency-assisted Image Repair block compensates for quality degradation in the untampered area. Extensive experiments show that our method achieves an average PSNR of 34.14 dB under common attack scenarios, including noise addition, image scaling, Gaussian blurring, and no post-processing. This represents an improvement of over 5 dB and 18% in recovered image quality compared to state-of-the-art approaches. Tong Liu 0021, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im, Pedro Martins 0003 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | A Two-Phase Scheme by Integration of Deep and Corner Feature for Balanced Copy-Move Forgery LocalizationabstractIn the era of Industry 4.0, the widespread application of digitization, automation, and Internet technology in industrial production has led to a significant increase in image data. Image security has become crucial because images are at risk of being tampered with at any time. To protect its authenticity, this article proposes a two-phase scheme to achieve balanced performance between accuracy and speed for copy-move forgery detection. Our scheme is divided into detection and localization phases. In the detection phase, the deep features are utilized to calculate the inner similarity. To improve the accuracy, a corner point matching technique is performed on the localization phase as a refinement step. The experimental results demonstrate the average$F1$-score is 0.6334 on CASIA2.0, making a 14.16% improvement. The computation time for each image is only 0.791 s in average. It has great significance in protecting the reliability and authenticity of industrial data. Tong Liu 0021, Xiaochen Yuan, Zhiyao Xie, Kaiqi Zhao 0004, Guoheng Huang, Chi-Man Pun |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Multi-granularity hypergraph-guided transformer learning framework for visual classification
Jianjian Jiang, Fangyuan Lei, Xiaochen Yuan |
Vis. Comput. | 6 |
| 2025 | Psanet: prototype-guided salient attention for few-shot segmentation
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun |
Vis. Comput. | 3 |
| 2025 | Privacy Image Secrecy Scheme Based on Chaos-Driven Fractal Sorting Matrix and Fibonacci Q-Matrix
Yunlong Liao, Xiaochen Yuan |
Vis. Comput. | 4 |
| 2025 | Weakly supervised semantic segmentation via saliency perception with uncertainty-guided noise suppression
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Guo Zhong, Xuhang Chen 0002, Chi-Man Pun |
Vis. Comput. | 3 |
| 2024 | PVALane: Prior-Guided 3D Lane Detection with View-Agnostic Feature AlignmentabstractMonocular 3D lane detection is essential for a reliable autonomous driving system and has recently been rapidly developing. Existing popular methods mainly employ a predefined 3D anchor for lane detection based on front-viewed (FV) space, aiming to mitigate the effects of view transformations. However, the perspective geometric distortion between FV and 3D space in this FV-based approach introduces extremely dense anchor designs, which ultimately leads to confusing lane representations. In this paper, we introduce a novel prior-guided perspective on lane detection and propose an end-to-end framework named PVALane, which utilizes 2D prior knowledge to achieve precise and efficient 3D lane detection. Since 2D lane predictions can provide strong priors for lane existence, PVALane exploits FV features to generate sparse prior anchors with potential lanes in 2D space. These dynamic prior anchors help PVALane to achieve distinct lane representations and effectively improve the precision of PVALane due to the reduced lane search space. Additionally, by leveraging these prior anchors and representing lanes in both FV and bird-eye-viewed (BEV) spaces, we effectively align and merge semantic and geometric information from FV and BEV features. Extensive experiments conducted on the OpenLane and ONCE-3DLanes datasets demonstrate the superior performance of our method compared to existing state-of-the-art approaches and exhibit excellent robustness. Zewen Zheng, Yongqiang Mou, Guoheng Huang, Chi-Man Pun, Xiaochen Yuan |
AAAI | 8 |
| 2024 | IMAN: An Adaptive Network for Robust NPC Mortality Prediction with Missing ModalitiesabstractAccurate prediction of mortality in nasopharyngeal carcinoma (NPC), a complex malignancy particularly challenging in advanced stages, is crucial for optimizing treatment strategies and improving patient outcomes. However, this predictive process is often compromised by the high-dimensional and heterogeneous nature of NPC-related data, coupled with the pervasive issue of incomplete multi-modal data, manifesting as missing radiological images or incomplete diagnostic reports. Traditional machine learning approaches suffer significant performance degradation when faced with such incomplete data, as they fail to effectively handle the high-dimensionality and intricate correlations across modalities. Even advanced multi-modal learning techniques like Transformers struggle to maintain robust performance in the presence of missing modalities, as they lack specialized mechanisms to adaptively integrate and align the diverse data types, while also capturing nuanced patterns and contextual relationships within the complex NPC data. To address these problem, we introduce IMAN: an adaptive network for robust NPC mortality prediction with missing modalities. IMAN features three integrated modules: the Dynamic Cross-Modal Calibration (DCMC) module employs adaptive, learnable parameters to scale and align medical images and field data; the Spatial-Contextual Attention Integration (SCAI) module enhances traditional Transformers by incorporating positional information within the self-attention mechanism, improving multi-modal feature integration; and the Context-Aware Feature Acquisition (CAFA) module adjusts convolution kernel positions through learnable offsets, allowing for adaptive feature capture across various scales and orientations in medical image modalities. Extensive experiments on our proprietary NPC dataset demonstrate IMAN’s robustness and high predictive accuracy, even with missing data. Compared to existing methods, IMAN consistently outperforms in scenarios with incomplete data, representing a significant advancement in mortality prediction for medical diagnostics and treatment planning. Our code is available at https://github.com/king-huoye/BIBM-2024/tree/master. Yejing Huo, Guoheng Huang, Lianglun Cheng, Jianbin He, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun |
BIBM | 6 |
| 2024 | FAQNet: Frequency-Aware Quaternion Network for Endoscopic Highlight RemovalabstractDue to the built-in light source within the endoscope, the illumination of bodily mucous can cause the formation of highlight regions due to reflection. This not only interferes with the diagnosis conducted by doctors but also poses a challenge to subsequent computer vision tasks. To tackle this issue, we introduce FAQNet, a network specifically designed for endoscopic image highlight removal. FAQNet seamlessly integrates multi-channel information leveraging quaternion convolution and spatial channel attention within our Quaternion Multi-Channel Fusion (QMCF) Module. This allows it to capture intricate details of color, texture, spatial information, and highlight characteristics within the imaged organ. Additionally, by employing frequency domain transformation and dilated convolution, the Contextual Information Integration (CII) Module effectively enlarges the receptive field, organizing contextual information between highlight regions and their surrounding areas. Lastly, the PixelShuffle Upsampling (PSU) Module generates the restored image. We validate our model’s performance on two benchmark datasets, demonstrating its superiority over existing highlight removal methodologies. Dingzhou Zhu, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun |
BIBM | 3 |
| 2024 | PDGC: Properly Disentangle by Gating and Contrasting for Cross-Domain Few-Shot Classification
Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Yan Li 0122, Chi-Man Pun, Junbing Quan |
CGI (2) | 3 |
| 2024 | Quantum Robust Coding for Quantum Image Watermarking
Xiaochen Yuan, Chan-Tong Lam |
ICIC (10) | 2 |
| 2024 | MSFGNet: Multi-Scale Features Gathering Network for Change Detection of Remote Sensing ImagesabstractChange detection is an important research area in remote sensing. To achieve accurate results, it is essential to extract multi-scale spatial information from images while filtering out noise. However, existing models lack this capability. Therefore, Multi-Scale Feature Gathering Network (MSFGNet) is proposed. Within MSFGNet, Bi-Temporal Image Multi-Level Fusion Module (BMF) is utilized to fuse bi-temporal remote sensing images. Additionally, Multi-Receptive Field Features Extraction Module (MRFE) is utilized to extract deep features. Within MRFE, Large Receptive Field Features Extraction Module (LRFE) and Multi-Scale Information Fusion Module (MSIF) are designed, which use large kernel convolution and dilated convolution respectively to capture spatial information with large receptive fields. Furthermore, Cross-Dimension Feature Sifting Fusion Module (CDFSF) is designed to sift noise from various dimensions, fusing valuable information. Across multiple public datasets, MSFGNet consistently achieves the best experimental results. The code can be accessed at https://github.com/juncyan/msfgnet.git. Junqing Huang, Xiaochen Yuan, Chan-Tong Lam, Wei Ke 0001 |
ICME | 2 |
| 2024 | FOPS-V: Feature-Aware Optimization and Parallel Scale Fusion for 3D Human Reconstruction in Video
Guoheng Huang, Lianglun Cheng, Yejing Huo, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun |
ICONIP (8) | 6 |
| 2024 | ROSAL: Semi-supervised Active Learning with Representation Aggregation and Outlier for Endoscopy Image Classification
Xiaocong Huang, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Xuhang Chen 0002, Chi-Man Pun, Jianwu Chen |
ICONIP (11) | 4 |
| 2024 | Dual Hypergraph Convolution Networks for Image Forgery Localization
Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam |
ICPR (22) | 2 |
| 2024 | DSTNet: Distinguishing Source and Target Areas for Image Copy-Move Forgery Detection
Kaiqi Zhao 0004, Xiaochen Yuan, Guoheng Huang |
ICPR (22) | 2 |
| 2024 | Cross-Modality Disentangled Information Bottleneck Strategy for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) has been a pivotal domain in current research area which utilizes diverse information carriers such as videos containing multiple modal-ities to understand the user's sentiment. With the success of multimodal fusion techniques, lots of fusion strategies have been proposed to obtain a favorable multimodal joint representation for MSA. However, existing studies hardly consider the problem of redundant information in unimodal, resulting in the joint representation may contain much redundant information from different modalities, thus limiting the accuracy of sentiment prediction. In this work, we propose a Cross-Modality Disentangled Information Bottleneck Strategy (CMDIBS), which consists of a Cross-Modality Knowledge Awareness (CMKA) module and a Multimodal Disentangled Information Bottleneck (MDIB) mechanism. Specifically, the CMKA module encourages in-teractions among different modalities to learn the sentiment embedding relevant to the predicted goals. In particular, MDIB mechanism aims to maximize the mutual information (MI) between the multimodal joint representation and the predicted label, and maximize the MI between the style embedding with the label and the input data while constraining the MI between the multimodal joint representation and the style embedding to obtain a succinct and efficient multimodal joint representation. Experimental results on the benchmark datasets, namely CMU-MOSI and CMU-MOSEI, indicated that the proposed method surpasses existing approaches and attains SOTA performance. Zhengnan Deng, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Lian Huang, Chi-Man Pun |
SMC | 4 |
| 2024 | IAMS-Net: An Illumination-Adaptive Multi-Scale Lesion Segmentation NetworkabstractIn recent years, many Lesion segmentation (LS) models based on UNet have been proposed. However, existing researches rarely consider the influence of illumination change leads to the weak boundary area. Such as melanomas and polyps, the demarcation of the boundary between the diseased area and the surrounding tissue remains particularly challenging. To overcome these challenges, we propose an IlluminationAdaptive Multi-scale Lesion Segmentation Network (IAMS-Net). In IAMS-Net, we integrate Illumination-Adaptive MultiStream Attention (IAMA) and Contour Perception Module (CPM). In the decoding stage, the IAMA is used as a bridge between the encoder and the decoder to solve the adverse effects of illumination changes on the segmentation of weak boundary lesions. In order to further enhance the boundary features lost due to illumination change in the low-contrast lesion area, we introduce the CPM to improve the perception of the integrity of the lesion area. Subsequently, we performed comparison and ablation experiments using the publicly available ISIC2018 dataset and the individually collected data set BoreIllumination(BI). Yisen Zheng, Guoheng Huang, Lianglun Cheng, Xiaochen Yuan, Guo Zhong, Shenghong Luo |
SMC | 5 |
| 2024 | CAMU-Net: Copy-move forgery detection utilizing coordinate attention and multi-scale feature fusion-based up-sampling
Kaiqi Zhao 0004, Xiaochen Yuan, Tong Liu 0021, Zhiyao Xie, Guoheng Huang, Li Feng 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Black-box reversible adversarial examples with invertible neural network
Jielun Huang, Guoheng Huang, Xiaochen Yuan, Fenfang Xie, Chi-Man Pun, Guo Zhong |
Image Vis. Comput. | 4 |
| 2024 | Progressive normalizing flow with learnable spectrum transform for style transfer
Guoheng Huang, Xiaochen Yuan, Guo Zhong, Chi-Man Pun, Yiwen Zeng |
Knowl. Based Syst. | 3 |
| 2024 | Recognition of score words in freestyle kayaking using improved DTW matching
Xiaochen Yuan, Chan-Tong Lam |
Multim. Tools Appl. | 2 |
| 2024 | GDN-CMCF: A Gated Disentangled Network With Cross-Modality Consensus Fusion for Multimodal Named Entity RecognitionabstractMultimodal named entity recognition (MNER) is a crucial task in social systems of artificial intelligence that requires precise identification of named entities in sentences using both visual and textual information. Previous methods have focused on capturing fine-grained visual features and developing complex fusion procedures. However, these approaches overlook the heterogeneity gap and loss of original modality uniqueness that may occur during fusion, leading to incorrect entity identification. This article proposes a novel approach for MNER called a gated disentangled network with cross-modality consensus fusion (GDN-CMCF) to address the above challenges. Specifically, to eliminate cross-modality variation, we propose a cross-modality consensus fusion module that generates a consensus representation by learning inter-and intramodality interactions with a designed commonality constraint. We then introduce a gated disentanglement module to separate modality-relevant features from support and auxiliary modalities, which further filters out extraneous information while retaining the uniqueness of unimodal features. Experimental results on two real public datasets are provided to verify the effectiveness of our proposed GDN-CMCF. The source code of this article can be found at https://github.com/HaoDavis/ GDN-CMCF. Guoheng Huang, Zihao Dai, Guo Zhong, Xiaochen Yuan, Chi-Man Pun |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | MMQW: Multi-Modal Quantum Watermarking SchemeabstractTo address the problem that existing quantum image watermarking schemes have only a single watermarking mode with weak robustness, in this paper we propose a novel multi-modal quantum watermarking (MMQW) scheme using the generalized model of novel enhanced quantum representation. Our scheme provides four quantum watermarking modes (G_G, G_C, C_C, C_G), covering both types of grayscale and color images for the watermark and the carrier image. To enhance the robustness, we propose the Block Bit-plane Centrosymmetric Expansion (BBCE) method, which utilizes controlled quantum gates to extend the watermark, making our method resistant to noise and geometric attacks. Moreover, we propose a Brightness-based Watermarking Mechanism (BWM) for embedding and extraction. By uniform embedding, BWM not only minimizes the impact on the carrier image but also reduces the visual distortion of the extracted watermark. In the proposed MMQW, we implement three adaptive embedding strategies using controlled quantum gates, each of which is adaptively triggered according to the corresponding modalities. Detailed quantum circuits for quantum computing are provided. To evaluate imperceptibility and robustness of the MMQW, we conduct experiments using high-resolution images from the USC-SIPI dataset. The results show that PSNR of the watermarked image ranges from 36 dB to 56 dB, indicating the high visual quality. The PSNR of the extracted watermark is about 34 dB when the noise density is 0.05, while the PSNR is higher than 48 dB under common quantum rotation attacks, which indicate the high robustness against noise addition and geometric attacks. In addition, the proposed MMQW can resist to cropping attack with cropping percentage up to 55%. A comprehensive comparison with existing state-of-the-art works shows that our method has significant advantages. Chan-Tong Lam, Xiaochen Yuan, Sio Kei Im, Penousal Machado |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Quaternion Cross-Modality Spatial Learning for Multi-Modal Medical Image SegmentationabstractRecently, the Deep Neural Networks (DNNs) have had a large impact on imaging process including medical image segmentation, and the real-valued convolution of DNN has been extensively utilized in multi-modal medical image segmentation to accurately segment lesions via learning data information. However, the weighted summation operation in such convolution limits the ability to maintain spatial dependence that is crucial for identifying different lesion distributions. In this paper, we propose a novel Quaternion Cross-modality Spatial Learning (Q-CSL) which explores the spatial information while considering the linkage between multi-modal images. Specifically, we introduce to quaternion to represent data and coordinates that contain spatial information. Additionally, we propose Quaternion Spatial-association Convolution to learn the spatial information. Subsequently, the proposed De-level Quaternion Cross-modality Fusion (De-QCF) module excavates inner space features and fuses cross-modality spatial dependency. Our experimental results demonstrate that our approach compared to the competitive methods perform well with only 0.01061 M parameters and 9.95G FLOPs. Junyang Chen 0001, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Zewen Zheng, Chi-Man Pun, Jian Zhu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Learning From Incorrectness: Active Learning With Negative Pre-Training and Curriculum Querying for Histological Tissue ClassificationabstractPatch-level histological tissue classification is an effective pre-processing method for histological slide analysis. However, the classification of tissue with deep learning requires expensive annotation costs. To alleviate the limitations of annotation budgets, the application of active learning (AL) to histological tissue classification is a promising solution. Nevertheless, there is a large imbalance in performance between categories during application, and the tissue corresponding to the categories with relatively insufficient performance are equally important for cancer diagnosis. In this paper, we propose an active learning framework called ICAL, which contains Incorrectness Negative Pre-training (INP) and Category-wise Curriculum Querying (CCQ) to address the above problem from the perspective of category-to-category and from the perspective of categories themselves, respectively. In particular, INP incorporates the unique mechanism of active learning to treat the incorrect prediction results that obtained from CCQ as complementary labels for negative pre-training, in order to better distinguish similar categories during the training process. CCQ adjusts the query weights based on the learning status on each category by the model trained by INP, and utilizes uncertainty to evaluate and compensate for query bias caused by inadequate category performance. Experimental results on two histological tissue classification datasets demonstrate that ICAL achieves performance approaching that of fully supervised learning with less than 16% of the labeled data. In comparison to the state-of-the-art active learning algorithms, ICAL achieved better and more balanced performance in all categories and maintained robustness with extremely low annotation budgets. The source code will be released at https://github.com/LactorHwt/ICAL. Lianglun Cheng, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Chi-Man Pun, Muyan Cai |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Multi-Label Chest X-Ray Image Classification With Single Positive LabelsabstractDeep learning approaches for multi-label Chest X-ray (CXR) images classification usually require large-scale datasets. However, acquiring such datasets with full annotations is costly, time-consuming, and prone to noisy labels. Therefore, we introduce a weakly supervised learning problem called Single Positive Multi-label Learning (SPML) into CXR images classification (abbreviated as SPML-CXR), in which only one positive label is annotated per image. A simple solution to SPML-CXR problem is to assume that all the unannotated pathological labels are negative, however, it might introduce false negative labels and decrease the model performance. To this end, we present a Multi-level Pseudo-label Consistency (MPC) framework for SPML-CXR. First, inspired by the pseudo-labeling and consistency regularization in semi-supervised learning, we construct a weak-to-strong consistency framework, where the model prediction on weakly-augmented image is treated as the pseudo label for supervising the model prediction on a strongly-augmented version of the same image, and define an Image-level Perturbation-based Consistency (IPC) regularization to recover the potential mislabeled positive labels. Besides, we incorporate Random Elastic Deformation (RED) as an additional strong augmentation to enhance the perturbation. Second, aiming to expand the perturbation space, we design a perturbation stream to the consistency framework at the feature-level and introduce a Feature-level Perturbation-based Consistency (FPC) regularization as a supplement. Third, we design a Transformer-based encoder module to explore the sample relationship within each mini-batch by a Batch-level Transformer-based Correlation (BTC) regularization. Extensive experiments on the CheXpert and MIMIC-CXR datasets have shown the effectiveness of our MPC framework for solving the SPML-CXR problem. Jiayin Xiao, Si Li 0005, Tongxu Lin, Jian Zhu 0001, Xiaochen Yuan, David Dagan Feng, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | SCDet: decoupling discriminative representation for dark object detection via supervised contrastive learning
Tongxu Lin, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Xiaocong Huang, Chi-Man Pun |
Vis. Comput. | 3 |
| 2024 | Automatic detection of breast lesions in automated 3D breast ultrasound with cross-organ transfer learningabstractDeep convolutional neural networks have garnered considerable attention in numerous machine learning applications, particularly in visual recognition tasks such as image and video analyses. There is a growing interest in applying this technology to diverse applications in medical image analysis. Automated three-dimensional Breast Ultrasound is a vital tool for detecting breast cancer, and computer-assisted diagnosis software, developed based on deep learning, can effectively assist radiologists in diagnosis. However, the network model is prone to overfitting during training, owing to challenges such as insufficient training data. This study attempts to solve the problem caused by small datasets and improve model detection performance. We propose a breast cancer detection framework based on deep learning (a transfer learning method based on cross-organ cancer detection) and a contrastive learning method based on breast imaging reporting and data systems (BI-RADS). When using cross organ transfer learning and BIRADS based contrastive learning, the average sensitivity of the model increased by a maximum of 16.05%. Our experiments have demonstrated that the parameters and experiences of cross-organ cancer detection can be mutually referenced, and contrastive learning method based on BI-RADS can improve the detection performance of the model. B. A. O. Lingyun, Zhengrui Huang, Yue Sun 0001, Hui Chen 0020, Xiaochen Yuan, Tao Tan 0002 |
Virtual Real. Intell. Hardw. | 8 |
| 2023 | A Noise Convolution Network for Tampering Detection
Zhiyao Xie, Xiaochen Yuan, Chan-Tong Lam, Guoheng Huang |
ICANN (10) | 2 |
| 2023 | RA-Net: A Deep Learning Approach Based on Residual Structure and Attention Mechanism for Image Copy-Move Forgery Detection
Kaiqi Zhao 0004, Xiaochen Yuan, Zhiyao Xie, Guoheng Huang, Li Feng 0001 |
ICANN (10) | 2 |
| 2023 | Single Cross-domain Semantic Guidance Network for Multimodal Unsupervised Image Translation
Jiaying Lan, Lianglun Cheng, Guoheng Huang, Chi-Man Pun, Xiaochen Yuan, Shangyu Lai, Bingo Wing-Kuen Ling |
MMM (1) | 5 |
| 2023 | Parallel multiple watermarking using adaptive Inter-Block correlation
Xingrun Wang, Xiaochen Yuan, Mianjie Li, Jinyu Tian 0001, Hongfei Guo, Jianqing Li 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Tampering localization and self-recovery using block labeling and adaptive significance
Xiaochen Yuan, Tong Liu 0021, Chan-Tong Lam, Guoheng Huang, Di Lin 0002, Ping Li 0016 |
Expert Syst. Appl. | 2 |
| 2023 | Power normalized cepstral robust features of deep neural networks in a cloud computing data privacy protection scheme
Mianjie Li, Zhihong Tian 0001, Xiaojiang Du, Xiaochen Yuan, Chun Shan, Mohsen Guizani |
Neurocomputing | 4 |
| 2023 | Reversible multi-watermarking for color images with grayscale invariance
Xiaochen Yuan, Xingrun Wang, Jianqing Li 0001 |
Multim. Tools Appl. | 2 |
| 2023 | RBA-GCN: Relational Bilevel Aggregation Graph Convolutional Network for Emotion RecognitionabstractEmotion recognition in conversation (ERC) has received increasing attention from researchers due to its wide range of applications. As conversation has a natural graph structure, numerous approaches used to model ERC based on graph convolutional networks (GCNs) have yielded significant results. However, the aggregation approach of traditional GCNs suffers from the node information redundancy problem, leading to node discriminant information loss. Additionally, single-layer GCNs lack the capacity to capture long-range contextual information from the graph. Furthermore, the majority of approaches are based on textual modality or stitching together different modalities, resulting in a weak ability to capture interactions between modalities. To address these problems, we present the relational bilevel aggregation graph convolutional network (RBA-GCN), which consists of three modules: the graph generation module (GGM), similarity-based cluster building module (SCBM) and bilevel aggregation module (BiAM). First, GGM constructs a novel graph to reduce the redundancy of target node information. Then, SCBM calculates the node similarity in the target node and its structural neighborhood, where noisy information with low similarity is filtered out to preserve the discriminant information of the node. Meanwhile, BiAM is a novel aggregation method that can preserve the information of nodes during the aggregation process. This module can construct the interaction between different modalities and capture long-range contextual information based on similarity clusters. On both the IEMOCAP and MELD datasets, the weighted average F1 score of RBA-GCN has a 2.17$\sim$5.21% improvement over that of the most advanced method. Guoheng Huang, Fenghuan Li, Xiaochen Yuan, Chi-Man Pun, Guo Zhong |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Quaternion-Valued Correlation Learning for Few-Shot Semantic SegmentationabstractFew-shot segmentation (FSS) aims to segment unseen classes given only a few annotated samples. Encouraging progress has been made for FSS by leveraging semantic features learned from base classes with sufficient training samples to represent novel classes. The correlation-based methods lack the ability to consider interaction of the two subspace matching scores due to the inherent nature of the real-valued 2D convolutions. In this paper, we introduce a quaternion perspective on correlation learning and propose a novel Quaternion-valued Correlation Learning Network (QCLNet), with the aim to alleviate the computational burden of high-dimensional correlation tensor and explore internal latent interaction between query and support images by leveraging operations defined by the established quaternion algebra. Specifically, our QCLNet is formulated as a hyper-complex valued network and represents correlation tensors in the quaternion domain, which uses quaternion-valued convolution to explore the external relations of query subspace when considering the hidden relationship of the support sub-dimension in the quaternion space. Extensive experiments on the PASCAL-$5^{i}$and COCO-$20^{i}$datasets demonstrate that our method outperforms the existing state-of-the-art methods effectively. Zewen Zheng, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Bingo Wing-Kuen Ling |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | QGD-Net: A Lightweight Model Utilizing Pixels of Affinity in Feature Layer for Dermoscopic Lesion SegmentationabstractRESPONSE: Pixels with location affinity, which can be also called "pixels of affinity," have similar semantic information. Group convolution and dilated convolution can utilize them to improve the capability of the model. However, for group convolution, it does not utilize pixels of affinity between layers. For dilated convolution, after multiple convolutions with the same dilated rate, the pixels utilized within each layer do not possess location affinity with each other. To solve the problem of group convolution, our proposed quaternion group convolution uses the quaternion convolution, which promotes the communication between to promote utilizing pixels of affinity between channels. In quaternion group convolution, the feature layers are divided into 4 layers per group, ensuring the quaternion convolution can be performed. To solve the problem of dilated convolution, we propose the quaternion sawtooth wave-like dilated convolutions module (QS module). QS module utilizes quaternion convolution with sawtooth wave-like dilated rates to effectively leverage the pixels that share the location affinity both between and within layers. This allows for an expanded receptive field, ultimately enhancing the performance of the model. In particular, we perform our quaternion group convolution in QS module to design the quaternion group dilated neutral network (QGD-Net). Extensive experiments on Dermoscopic Lesion Segmentation based on ISIC 2016 and ISIC 2017 indicate that our method has significantly reduced the model parameters and highly promoted the precision of the model in Dermoscopic Lesion Segmentation. And our method also shows generalizability in retinal vessel segmentation. Jingchao Wang 0002, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Chi-Man Pun |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Detection of weak electromagnetic interference attacks based on fingerprint in IIoT systems
Kai Fang 0001, Tingting Wang 0006, Xiaochen Yuan, Chunyu Miao, Yuanyuan Pan, Jianqing Li 0001 |
Future Gener. Comput. Syst. | 3 |
| 2021 | FD-TR: feature detector based on scale invariant feature transform and bidirectional feature regionalization for digital image watermarking
Mianjie Li, Xiaochen Yuan |
Multim. Tools Appl. | 2 |
| 2021 | A dual-tamper-detection method for digital image authentication and content self-recovery
Tong Liu 0021, Xiaochen Yuan |
Multim. Tools Appl. | 2 |
| 2021 | Gauss-Jordan elimination-based image tampering detection and self-recovery
Xiaochen Yuan, Tong Liu 0021 |
Signal Process. Image Commun. | 1 |
| 2021 | Adaptive Feature Calculation and Diagonal Mapping for Successive Recovery of Tampered RegionsabstractThis article proposes an adaptive scheme for image tampered region localization and content recovery. To generate the watermark information comprised of the authentication data and recovery data, we firstly propose the Adaptive Authentication Feature Calculation algorithm to obtain the authentication data, which includes the information of block location and block feature. The DWT-based Block Feature Calculation method is then proposed to calculate the block feature, and the quantization method is employed to calculate the block location. The recovery data is composed of self-recovery bits and mapped-recovery bits. The self-recovery bits are obtained by the Set Partitioning in Hierarchical Trees encoding algorithm. For retrieving the damaged codes caused by tampering, we propose the Diagonal Mapping algorithm and apply it to the self-recovery bits, thus generating the mapped-recovery bits, to provide a guarantee of recovery data. Experimental results show the superior performance of the proposed scheme in terms of tamper detection and image recovery, by comparing with the state-of-the-art works. The results demonstrate that the proposed method shows efficiency in the adaptiveness, well localization, strong capability for image recovery, and the effectiveness of attack resistance. Tong Liu 0021, Xiaochen Yuan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Image Self-Recovery Based on Authentication Feature ExtractionabstractThis paper proposes a novel image self-recovery scheme based on authentication feature extraction. The Authentication Feature Extraction method is proposed to calculate the authentication information. The Set Partitioning in Hierarchical Trees encoding algorithm is employed to calculate the recovery information. Moreover, in order to retrieve the damaged information caused by tampering, we propose to map each block into another position and generate the mapped-recovery information accordingly. In this way, a double assurance of recovery information can be provided. Experimental results show the superior performance of the proposed scheme in terms of image self-recovery. Comparison with the state-of-the-art works demonstrate that the proposed scheme shows efficiency in strong capability for image recovery, and effectiveness of attack resistance. Tong Liu 0021, Xiaochen Yuan |
TrustCom | 2 |
| 2020 | Adaptive segmentation-based feature extraction and S-STDM watermarking method for color image
Mianjie Li, Xiaochen Yuan |
Neural Comput. Appl. | 2 |
| 2018 | Multi-scale feature extraction and adaptive matching for copy-move forgery detection
Xiuli Bi, Chi-Man Pun, Xiaochen Yuan |
Multim. Tools Appl. | 3 |
| 2018 | Robust image hashing using progressive feature selection for tampering detection
Chi-Man Pun, Cai-Ping Yan, Xiaochen Yuan |
Multim. Tools Appl. | 3 |
| 2018 | RegFrame: fast recognition of simple human actions on a stand-alone mobile device
Di Han 0002, Jianqing Li 0001, Zihua Zeng, Xiaochen Yuan |
Neural Comput. Appl. | 4 |
| 2018 | Local multi-watermarking method based on robust and adaptive feature extraction
Xiaochen Yuan, Mianjie Li |
Signal Process. | 1 |
| 2017 | Image Alignment-Based Multi-Region Matching for Object-Level Tampering DetectionabstractTampering detection methods based on image hashing have been widely studied with continuous advancements. However, most existing models cannot generate object-level tampering localization results, because the forensic hashes attached to the image lack contour information. In this paper, we present a novel tampering detection model that can generate an accurate, object-level tampering localization result. First, an adaptive image segmentation method is proposed to segment the image into closed regions based on strong edges. Then, the color and position features of the closed regions are extracted as a forensic hash. Furthermore, a geometric invariant tampering localization model named image alignment-based multi-region matching (IAMRM) is proposed to establish the region correspondence between the received and forensic images by exploiting their intrinsic structure information. The model estimates the parameters of geometric transformations via a robust image alignment method based on triangle similarity; in addition, it matches multiple regions simultaneously by utilizing manifold ranking based on different graph structures and features. Experimental results demonstrate that the proposed IAMRM is a promising method for object-level tampering detection compared with the state-of-the-art methods. Chi-Man Pun, Cai-Ping Yan, Xiaochen Yuan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | Multi-Level Dense Descriptor and Hierarchical Feature Matching for Copy-Move Forgery Detection
Xiuli Bi, Chi-Man Pun, Xiaochen Yuan |
Inf. Sci. | 3 |
| 2016 | Multi-scale noise estimation for image splicing forgery detection
Chi-Man Pun, Bo Liu 0047, Xiaochen Yuan |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Multi-scale image hashing using adaptive local feature extraction for robust tampering detection
Cai-Ping Yan, Chi-Man Pun, Xiaochen Yuan |
Signal Process. | 3 |
| 2016 | Quaternion-Based Image Hashing for Adaptive Tampering LocalizationabstractImage-hashing-based tampering detection methods have been widely studied with continuous advancements. However, most of existing models are designed for a specific tampering. In this paper, we propose a novel quaternion-based image hashing to detect almost all types of tampering, including color changing, copy move, splicing, and so on. First, the quaternion Fourier-Mellin transform is used to calculate the geometric hash to eliminate the influence of geometric distortions. Then, a new quaternion image construction method, which combines advantages of both color and structural features, is proposed to implement the quaternion Fourier transform to calculate the image feature hash to locate the tampered regions. The objective is to provide a reasonably short image hashing with good performance, i.e., being perceptually robust against various content-preserving attacks while capable of detecting and locating almost all types of tampering. Furthermore, an adaptive tampering localization algorithm is proposed based on clustering analysis to improve the detection accuracy. The experimental results show that the proposed tampering detection model outperforms the existing state-of-the-art models and is very robust against various content-preserving attacks. Cai-Ping Yan, Chi-Man Pun, Xiaochen Yuan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2015 | Robust Mel-Frequency Cepstral coefficients feature detection and dual-tree complex wavelet transform for digital audio watermarking
Xiaochen Yuan, Chi-Man Pun, C. L. Philip Chen |
Inf. Sci. | 1 |
| 2015 | Histogram modification based image watermarking resistant to geometric distortions
Chi-Man Pun, Xiaochen Yuan |
Multim. Tools Appl. | 2 |
| 2015 | Image Forgery Detection Using Adaptive Oversegmentation and Feature Point MatchingabstractA novel copy-move forgery detection scheme using adaptive oversegmentation and feature point matching is proposed in this paper. The proposed scheme integrates both block-based and keypoint-based forgery detection methods. First, the proposed adaptive oversegmentation algorithm segments the host image into nonoverlapping and irregular blocks adaptively. Then, the feature points are extracted from each block as block features, and the block features are matched with one another to locate the labeled feature points; this procedure can approximately indicate the suspected forgery regions. To detect the forgery regions more accurately, we propose the forgery region extraction algorithm, which replaces the feature points with small superpixels as feature blocks and then merges the neighboring blocks that have similar local color features into the feature blocks to generate the merged regions. Finally, it applies the morphological operation to the merged regions to generate the detected forgery regions. The experimental results indicate that the proposed copy-move forgery detection scheme can achieve much better detection results even under various challenging conditions compared with the existing state-of-the-art copy-move forgery detection methods. Chi-Man Pun, Xiaochen Yuan, Xiuli Bi |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Feature extraction and local Zernike moments based geometric invariant watermarking
Xiaochen Yuan, Chi-Man Pun |
Multim. Tools Appl. | 1 |
| 2013 | Geometric invariant watermarking by local Zernike moments of binary image patches
Xiaochen Yuan, Chi-Man Pun, C. L. Philip Chen |
Signal Process. | 1 |
| 2013 | Robust Segments Detector for De-Synchronization Resilient Audio WatermarkingabstractA robust feature points detector for invariant audio watermarking is proposed in this paper. The audio segments centering at the detected feature points are extracted for both watermark embedding and extraction. These feature points are invariant to various attacks and will not be changed much for maintaining high auditory quality. Besides, high robustness and inaudibility can be achieved by embedding the watermark into the approximation coefficients of Stationary Wavelet Transform (SWT) domain, which is shift invariant. The spread spectrum communication technique is adopted to embed the watermark. Experimental results show that the proposed Robust Audio Segments Extractor (RASE) and the watermarking scheme are not only robust against common audio signal processing, such as low-pass filtering, MP3 compression, echo addition, volume change, and normalization; and distortions introduced in Stir-mark benchmark for Audio; but also robust against synchronization geometric distortions simultaneously, such as resample time-scale modification (TSM) with scaling factors up to ±50%, pitch invariant TSM by ±50%, and tempo invariant pitch shifting by ±50%. In general, the proposed scheme can well resist various attacks by the joint RASE and SWT approach, which performs much better comparing with the existing state-of-the art methods. Chi-Man Pun, Xiaochen Yuan |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Geometric Invariant Digital Image Watermarking Scheme Based on Feature Points Detector and Histogram DistributionabstractA robust and geometric invariant digital image watermarking scheme based on SIFT Based Feature Points Detector (SIFTFPD) and histogram distribution is proposed in this paper. The SIFTFPD is proposed to extract geometric invariant feature points from the host image for watermark embedding; and the descriptor is generated subsequently. With the feature extraction procedure, the circular regions centered at the extracted feature points and with the given radius are defined as embedding regions. For watermark embedding, some pixels are moved to form a specific pattern in the intensity-level histogram distribution in each embedding region, to indicate the watermark. For watermark extraction, the embedded regions are generated with the descriptor and according to the intensity-level histogram distribution in each region, the watermark can be extracted. Experimental results show that the proposed scheme is very robust against geometric distortion such as rotation, scaling, cropping, and affine transformation; and common signal processing, such as JPEG compression, median filtering, and Gaussian low-pass filtering. Chi-Man Pun, Xiaochen Yuan, C. L. Philip Chen |
TrustCom | 2 |