Chi-Man Pun

dblp:p/ChiManPun · also Chi Man Pun · DBLP profile ↗
← Back
267ranked-venue papers
17as first author
165since 2021 · last 2026
0000-0003-1788-3746ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 115 · 8 first-author · 68 since 2021Artificial intelligence and machine learning · 83 · 4 first-author · 57 since 2021Applied, interdisciplinary, general and emerging computing · 38 · 35 since 2021Databases, data management, data science and information retrieval · 28 · 1 first-author · 12 since 2021Security and privacy · 16 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 since 2021Computer networks · 6 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 OTI: A Model-free and Visually Interpretable Measure of Image Attackability
abstract
Despite the tremendous success of neural networks, benign images can be corrupted by adversarial perturbations to deceive these models. Intriguingly, images differ in their attackability. Specifically, given an attack configuration, some images are easily corrupted, whereas others are more resistant. Evaluating image attackability has important applications in active learning, adversarial training, and attack enhancement. This prompts a growing interest in developing attackability measures. However, existing methods are scarce and suffer from two major limitations: (1) They rely on a model proxy to provide prior knowledge (e.g., gradients or minimal perturbation) to extract model-dependent image features. Unfortunately, in practice, many task-specific models are not readily accessible. (2) Extracted features characterizing image attackability lack visual interpretability, obscuring their direct relationship with the images. To address these, we propose a novel Object Texture Intensity (OTI), a model-free and visually interpretable measure of image attackability, which measures image attackability as the texture intensity of the image's semantic object. Theoretically, we describe the principles of OTI from the perspectives of decision boundaries as well as the mid- and high-frequency characteristics of adversarial perturbations. Comprehensive experiments demonstrate that OTI is effective and computationally efficient. In addition, our OTI provides the adversarial machine learning community with a visual understanding of attackability.
Chi-Man Pun
AAAI3
2026 SNS-Grasp: Semantic-guided Noise Scaling for Grasp Generation
abstract
While diffusion models show promise for intent-based grasp generation, their isotropic noise schedules struggle with joint-specific sensitivity and task-aware variability. This limitation leads to grasps with suboptimal semantic alignment or physical feasibility. To address this challenge, we propose Semantic-guided Noise Scaling for grasp generation (SNS-Grasp), a novel framework that integrates two key innovations. First, the Semantic-guided Noise Scaling Diffusion (SNS-Diff) module generates intent-aware grasps by replacing isotropic noise with anisotropic modulation, dynamically adapting to task semantics and joint-specific sensitivity. Specifically, SNS-Diff leverages a pretrained Intent Recognizer to extract task-aware confidence scores and joint-specific gradient sensitivities from the interaction context. These signals adjust the noise scaling during denoising, downweighting perturbations for semantically critical joints to ensure semantic alignment. Second, the Fine-grained Grasp Refinement (FGR) module establishes dynamic joint-vertex coupling through fine-grained hand-object spatial relationships, enabling iterative optimization of physically executable grasps. Extensive experiments on OakInk and GRAB demonstrate SNS-Grasp's superior performance in semantic accuracy and physical feasibility, with robust generalization to unseen objects.
Zhenhua Tang 0001, Yudian Zheng, Yuzhang Zhong, Haolun Li 0001, Yanbin Hao, Chi-Man Pun
AAAI6
2026 IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection
abstract
The rapid advancements in artificial intelligence have significantly accelerated the adoption of speech recognition technology, leading to its widespread integration across various applications. However, this surge in usage also highlights a critical issue: audio data is highly vulnerable to unauthorized exposure and analysis, posing significant privacy risks for businesses and individuals. This paper introduces an Information-Obfuscation Reversible Adversarial Example (IO-RAE) framework, the pioneering method designed to safeguard audio privacy using reversible adversarial examples. IO-RAE leverages large language models to generate misleading yet contextually coherent content, effectively preventing unauthorized eavesdropping by humans and Automatic Speech Recognition (ASR) systems. Additionally, we propose the Cumulative Signal Attack technique, which mitigates high-frequency noise and enhances attack efficacy by targeting low-frequency signals. Our approach ensures the protection of audio data without degrading its quality or usability. Experimental evaluations demonstrate the superiority of our method, achieving a targeted misguidance rate of 96.5% and a remarkable 100% untargeted misguidance rate in obfuscating target keywords across multiple ASR models, including a commercial black-box system from Google. Furthermore, the quality of the recovered audio, measured by the Perceptual Evaluation of Speech Quality score, reached 4.45, comparable to high-quality original recordings. Notably, the recovered audio processed by ASR systems exhibited an error rate of 0%, indicating nearly lossless recovery. These results highlight the practical applicability and effectiveness of our IO-RAE framework in protecting sensitive audio privacy.
Xia Du, Jizhe Zhou 0001, Qizhen Xu, Zheng Lin 0001, Chi-Man Pun
AAAI7
2026 Decoupling Vocal and Rhythmic Conditioning for Music-Driven Singing Avatar Animation
abstract
Synthetic media generation is a burgeoning field in multimedia research. While audio-driven avatar animation has garnered significant attention in digital entertainment, yet music-driven singing avatar animation remains relatively underexplored due to its unique challenges. Distinct from speech, singing animation necessitates the simultaneous modeling of lip articulation governed by singing vocal, and global facial dynamics synchronized with musical rhythm. Existing methods typically rely on 3D intermediate representations, which impose geometric constraints and often degrade visual details. Furthermore, some approaches that simply concatenate vocal and BGM features fail to capture the distinct roles of these signals in driving specific facial regions. To address these limitations, we propose MusicAvatar, a diffusion-based framework that directly synthesizes 2D singing avatars without relying on 3D priors. Moreover, we design a dual-stream music attention module that decouples the roles of singing voice and BGM. Specifically, one cross-attention stream extracts vocal cues from the singing track to drive lip movements, while a parallel stream captures rhythmic patterns from the BGM to modulate facial motion. This parallel yet synergistic design ensures that precise lip movement and rhythmic facial motion are modeled explicitly without interference. Extensive experiments demonstrate that MusicAvatar generates highly natural, expressive, and rhythmically synchronized singing avatars, outperforming state-of-the-art approaches.
Yiguo Jiang, Xiaodong Cun, Chen-Bin Feng, Jian Sun 0038, Chi-Man Pun
ICMR5
2026 ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis
abstract
High-fidelity character voice synthesis is a cornerstone of immersive multimedia applications, particularly for interacting with anime avatars and digital humans. However, existing systems struggle to maintain consistent persona traits across diverse emotional contexts. To bridge this gap, we present ATRIE, a unified framework utilizing a Persona-Prosody Dual-Track (P2-DT) architecture. Our system disentangles generation into a static Timbre Track (via Scalar Quantization) and a dynamic Prosody Track (via Hierarchical Flow-Matching), distilled from a 14B LLM teacher. This design enables robust identity preservation (Zero-Shot Speaker Verification EER: 0.04) and rich emotional expression. Evaluated on our extended AnimeTTS-Bench (50 characters), ATRIE achieves state-of-the-art performance in both generation and cross-modal retrieval (mAP: 0.75), establishing a new paradigm for persona-driven multimedia content creation. The code is available at Github.
Aoduo Li, Hongjian Xu, Shengmin Li, Sihao Qin, Zimeng Li 0001, Chi-Man Pun, Xuhang Chen 0002
ICMR7
2026 Decoding coefficients recovery based on modified Gauss-Jordan elimination for tampered content reconstruction
Tong Liu 0021, Lihao Zhuang, Guoheng Huang, Chi-Man Pun, Xiaochen Yuan
Expert Syst. Appl.4
2026 CCSFusion: A Hierarchical Semantic Chain-of-Thought Reasoning Architecture for Infrared-Visible Image Fusion and Captioning
abstract
Infrared-Visible Image Fusion (IVIF) aims to generate a single, information-rich image for downstream tasks. However, prevailing methods exhibit two key limitations. First, many approaches lack explicit hierarchical semantic decoupling, failing to effectively integrate semantic features across different levels, which restricts their ability to capture complex scene structures. Second, task-driven fusion frameworks typically adopt a cascaded design, with unidirectional supervision provided by geometry-centric downstream tasks like detection. This architecture not only limits mutual reinforcement between the fusion and task networks, but also creates a ”supervision bottleneck” by lacking interaction with the linguistic modality that captures richer scene relationships. To tackle these challenges, we propose CCSFusion, the first framework that leverages Chain-of-Thought captioning as supervision, redirecting IVIF optimization from narrow geometric accuracy to multimodal scene comprehension. It establishes a mutually reinforcing coupling between the fusion network and the captioning task. Specifically, we introduce a Segmentation Mask Calibration Unit (SMCU) to refine coarse semantic priors, providing precise pixel-level guidance. Subsequently, the calibrated features are fed into Chained Semantic Fusion Module (CSFM) which explicitly decomposes the semantic priors into three hierarchical levels, and then feeds them into the Hierarchical Semantic Attention module. Finally, a bidirectional knowledge distillation mechanism transfers the reasoning ability of the teacher network to the student. Experiments show that CCSFusion achieves superior fusion performance and generates more semantically coherent images for high-level cognitive tasks. The code is available at: https://github.com/Snaillms/CCSFusion.
Miaoshan Lin, Guoheng Huang, Jietao Yang, Jiehao Zheng, Xiaochen Yuan, Yan Li 0122, Xiaofeng Zhang 0006, Kim Fung Tsang, Chi-Man Pun
IEEE Internet Things J.9
2026 A two-stage sign language generation framework with self-supervised latent representation learning
Qiguang Miao, Guanwen Feng, Junwei Jing, Yilin Zhang 0007, Yunan Li 0001, Chi-Man Pun
Knowl. Based Syst.7
2026 MADAT: Missing-aware dynamic adaptive transformer model for medical prognosis prediction with incomplete multimodal data
Jianbin He, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Guo Zhong, Bai Ying Lei, Haojiang Li
Medical Image Anal.4
2026 QWNet: A quaternion wavelet network for spatial-frequency aware multi-modal image fusion
Jietao Yang, Miaoshan Lin, Guoheng Huang, Xuhang Chen 0002, Xiaofeng Zhang 0006, Xiaochen Yuan, Chi-Man Pun, Bingo Wing-Kuen Ling
Neural Networks7
2026 Explicit Visual Prompting for Universal Foreground Segmentations
abstract
Foreground segmentation is a fundamental problem in computer vision, which includes salient object detection, forgery detection, defocus blur detection, shadow detection, and camouflage object detection. Previous works have typically relied on domain-specific solutions to address accuracy and robustness issues in those applications. In this paper, we present a unified framework for a number of foreground segmentation tasks without any task-specific designs. We take inspiration from the widely-used pre-training and then prompt tuning protocols in NLP and propose a new visual prompting model, named Explicit Visual Prompting (EVP). Different from the previous visual prompting which is typically a dataset-level implicit embedding, our key insight is to enforce the tunable parameters focusing on the explicit visual content from each individual image, i.e., the features from frozen patch embeddings and high-frequency components. Our method freezes a pre-trained model and then learns task-specific knowledge using a few extra parameters. Despite introducing only a small number of tunable parameters, EVP achieves superior performance than full fine-tuning and other parameter-efficient fine-tuning methods. Experiments in fourteen datasets across five tasks show the proposed method outperforms other task-specific methods while being considerably simple. The proposed method demonstrates the scalability in different architectures, pre-trained weights, and tasks.
Weihuang Liu, Xi Shen 0001, Chi-Man Pun, Xiaodong Cun
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Towards structural transformation-based attack for boosting transferability of adversarial examples
Yatie Xiao, Chi-Man Pun, Fei Peng 0001, Kongyang Chen, Qingxiao Guan
Pattern Recognit.2
2026 LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space
abstract
While existing one-shot talking head generation models have achieved progress in coarse-grained emotion editing, there is still a lack of fine-grained emotion editing models with high interpretability. We argue that for an approach to be considered fine-grained, it needs to provide clear definitions and sufficiently detailed differentiation. We present LES-Talker, a novel one-shot talking head generation model with high interpretability, to achieve fine-grained emotion editing across emotion types, emotion levels, and facial units. We propose a Linear Emotion Space (LES) definition based on Facial Action Units to characterize emotion transformations as vector transformations. We design the Cross-Dimension Attention Net (CDAN) to deeply mine the correlation between LES representation and 3D model representation. Through mining multiple relationships across different feature and structure dimensions, we enable LES representation to guide the controllable deformation of 3D model. In order to adapt the multimodal data with deviations to the LES and enhance visual quality, we utilize specialized network design and training strategies. Experiments show that our method provides high visual quality along with multilevel and inter pretable fine-grained emotion editing, outperforming mainstream methods. Project page: https://peterfanfan.github.io/LES-Talker/
Guanwen Feng, Zhihao Qian 0001, Yunan Li 0001, Qiguang Miao, Chi-Man Pun
IEEE Trans. Affect. Comput.6
2026 HINTS: Hierarchically Disentangling Subregional Heterogeneity With Structural Priors for Multi-Modal Survival Analysis
Biyun Chen, Guoheng Huang, Xiaochen Yuan, Yan Li 0122, Chi-Man Pun, Bai Ying Lei, Haojiang Li
IEEE Trans Autom. Sci. Eng.8
2026 Interpretation Before Integration: LLM-Guided Multimodal Completion and Fusion Network for Survival Analysis With Incomplete Data
Feng Ling 0002, Haoming Zeng, Ming Li 0065, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Xianglian Liao, Jiong-Lin Liang, Haojiang Li
IEEE Trans. Comput. Soc. Syst.6
2026 Defensive Adversarial CAPTCHA: A Semantics-Driven Framework for Natural Adversarial Example Generation
abstract
Traditional CAPTCHA (Completely Automated Public Turing Test to Tell Computers and Humans Apart) schemes are increasingly vulnerable to automated attacks powered by deep neural networks (DNNs). Existing adversarial attack methods often rely on the original image characteristics, resulting in distortions that hinder human interpretation and limit their applicability in scenarios where no initial input images are available. To address these challenges, we propose the Unsourced Adversarial CAPTCHA (DAC), a novel framework that generates high-fidelity adversarial examples guided by attacker-specified semantics information. Leveraging a Large Language Model (LLM), DAC enhances CAPTCHA diversity and enriches the semantic information. To address various application scenarios, we examine the white-box targeted attack scenario and the black-box untargeted attack scenario. For target attacks, we introduce two latent noise variables that are alternately guided in the diffusion step to achieve robust inversion. The synergy between gradient guidance and latent variable optimization achieved in this way ensures that the generated adversarial examples not only accurately align with the target conditions but also achieve optimal performance in terms of distributional consistency and attack effectiveness. In untargeted attacks, especially for black-box scenarios, we introduce bi-path unsourced adversarial CAPTCHA (BP-DAC), a two-step optimization strategy employing multimodal gradients and bi-path optimization for efficient misclassification. Experiments show that the defensive adversarial CAPTCHA generated by BP-DAC is able to defend against most of the unknown models, and the generated CAPTCHA is indistinguishable to both humans and DNNs.
Xia Du, Jizhe Zhou 0001, Zheng Lin 0001, Chi-Man Pun, Cong Wu 0003, Tao Li 0001, Zhe Chen 0015, Wei Ni 0001, Jun Luo 0001
IEEE Trans. Dependable Secur. Comput.5
2026 EmoSpeaker: One-Shot Fine-Grained Emotion-Controlled Talking Face Generation
abstract
Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced emotional states, thereby improving the emotional quality and personalization of generated content. Generating fine-grained facial animations that accurately portray emotional expressions using only a portrait and an audio recording presents a challenge. In order to address this challenge, we propose a visual attribute-guided audio decoupler. This enables the obtention of content vectors solely related to the audio content, enhancing the stability of subsequent lip movement coefficient predictions. To achieve more precise emotional expression, we introduce a fine-grained emotion coefficient prediction module. Additionally, we propose an emotion intensity control method using a fine-grained emotion matrix. Through these, effective control over emotional expression in the generated videos and finer classification of emotion intensity are accomplished. Subsequently, a series of 3DMM coefficient generation networks are designed to predict 3D coefficients, followed by the utilization of a rendering network to generate the final video. Our experimental results demonstrate that our proposed method, EmoSpeaker, outperforms existing emotional talking face generation methods in terms of expression variation and lip synchronization. Project page:https://peterfanfan.github.io/EmoSpeaker/
Guanwen Feng, Yunan Li 0001, Chaoneng Li, Zhihao Qian 0001, Qiguang Miao, Chi-Man Pun
IEEE Trans. Multim.8
2026 PASK: Sparse Framework for Crafting Natural Adversarial Example
abstract
As audio adversarial attacks continue to evolve, Automatic Speech Recognition (ASR) models have emerged as a significant target. Traditional audio attack methods often focus on minimizing perturbation magnitude and frequency, overlooking the importance of perturbation location. However, certain audio regions hold lower importance for ASR models, making attacks on these regions less effective and more perceptible as noise. Additionally, the human ear perceives noise differently depending on its placement within the audio sequence, with noise in silent segments being more noticeable. To address these challenges, this paper proposes Pitch Sparse Audio Attack (PASK), an innovative framework designed to enhance adversarial imperceptibility through sparse perturbations. PASK introduces two key techniques: Pitch Mapping, which provides a strategic starting point for perturbation, and an adaptive grouped selective mask that achieves targeted sparsity, focusing perturbations on high-impact audio regions. Experimental results demonstrate that PASK outperforms existing methods in both effectiveness and imperceptibility. Furthermore, a human study confirms that silent-segment perturbations are more easily detected, underscoring the perceptual advantages of our approach.
Xia Du, Jizhe Zhou 0001, Qizhen Xu, Chi-Man Pun
IEEE Trans. Multim.6
2026 Fre-QNet: Quaternion Progressive Perception Mechanism with Frequency-Guided Prompt for Blind Image Quality Assessment
abstract
Blind Image Quality Assessment faces challenges in enabling computational models to mimic the hierarchical progressive perception mechanisms of the Human Visual System (HVS). Existing methods often neglect the two-stage process of HVS—global distortion identification followed by local quality evaluation—and its distinct sensitivity to distortion types. To address this, we propose Fre-QNet, a novel framework integrating two key components: (1) A Quaternion Progressive Perception (QPP) module that hierarchically extracts multi-scale spatial features using quaternion convolution, explicitly simulating the global-to-local observation process of HVS while enhancing cross-scale interactions; (2) A Frequency Prompting (FP) module that quantifies distortion types and severity in the Fourier domain by leveraging frequency patterns of common distortions and the sensitivity variations of HVS. The QPP and FP modules collaboratively embed biological vision principles into computational modeling through dual-domain feature learning, with the QPP module directly anchoring the core logic of progressive perception. Experiments on TID2013 and CSIQ benchmarks demonstrate Fre-QNet’s superiority over state-of-the-art methods, validating its effectiveness in matching human perceptual quality judgments. Our source code is available at: https://github.com/hhsda/Fre-QNet .
Shize Li, Guoheng Huang, Yisen Zheng, Xiaochen Yuan, Xuhang Chen 0002, Lianglun Cheng, Chi-Man Pun
ACM Trans. Multim. Comput. Commun. Appl.8
2026 Harnessing Transferable Adversarial Examples via Multilayer Attention-Guided Spatial Transformations
abstract
Transfer-based adversarial attacks are key for evaluating the robustness of deep neural networks (DNNs) in black-box settings, yet their effectiveness is often constrained by limited cross-model transferability. Existing feature-level approaches typically rely on single-layer attention guidance or static perturbation patterns, which restrict adaptability across diverse architectures. In this work, we introduce a unified adversarial framework, named Multi-layer Attention-guided Spatial Transformations (MAT), to exploit class-discriminative cues from multiple feature layers to craft highly transferable adversarial examples. MAT integrates Multi-layer Attention Fusion (MAF) to capture complementary low-level and high-level semantics from multiple intermediate layers, Attention-guided Augmentation (AGA) to selectively perturb non-critical regions while preserving semantic integrity, and Spatial Random Transformation (SRT) to introduce stochastic spatial augmentations to diversify patterns during optimization. Unlike prior methods that use static or layer-specific attention, MAT dynamically adapts feature guidance to the architecture and task, which enhances generalization. We evaluate MAT against eleven state-of-the-art (SOTA) transfer-based attacks across nine CNN-based and Transformer-based architectures on ImageNet. Comprehensive experiments demonstrate that MAT consistently outperforms eleven state-of-the-art transfer-based attacks in both white-box and black-box settings, including against adversarially trained and input preprocessing-based defensive models, while maintaining higher semantic similarity to the original inputs. It highlights the superior adversarial robustness and excellent adaptability of MAT in adversarial machine learning. Our code is available athttps://github.com/dislab-gzhu/MAT.
Pengfei Dong, Yatie Xiao, Chi-Man Pun, Fei Peng 0001, Kongyang Chen, Qingxian Guan, Siyuan Chen 0005, Xiangyu Ye, Zhenbang Liu
IEEE Trans. Reliab.3
2026 WAQNIQA: Wavelet-Augmented Quaternion Network for No-Reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) plays a pivotal role in computer vision by enabling image quality evaluation without reference images. While recent CNN and Transformer-based methods have advanced feature extraction, they face significant limitations. CNNs exhibit local feature bias, limiting their ability to capture global dependencies and complex structures critical to understanding diverse distortions. Transformers, despite modeling nonlocal dependencies through multihead attention, suffer from quadratic computational complexity with spatial dimensions, hindering efficient multiscale analysis. Moreover, their attention mechanisms frequently overlook critical interchannel dependencies, which are vital for capturing fine details in texture-rich images. Coupled with difficulties in handling high-noise environments and complex textures, this results in limited real-world accuracy and poor generalization across diverse datasets and unknown distortions. To bridge these gaps, we propose WAQNIQA, a novel wavelet-augmented quaternion network for NR-IQA. Distinct from conventional architectures, WAQNIQA integrates two synergistic modules: the wavelet-infused adaptive attention (WIAA) module, which leverages wavelet transforms (WTs) to achieve robust multiscale spatial-frequency analysis with linear complexity, and the quaternion collaborative feature enhancement (QCFE) module, which holistically models interchannel correlations to preserve fine texture details. Furthermore, we introduce PowerGridIQ, the first NR-IQA dataset specifically tailored for power grid scenarios. Extensive experiments demonstrate that WAQNIQA consistently surpasses state-of-the-art CNN and Transformer-based methods on PowerGridIQ and six public benchmarks. Notably, WAQNIQA exhibits superior cross-domain generalization, achieving competitive performance on the AGIQA-1K dataset for AI-generated content (AIGC) without explicit semantic alignment training, thereby validating its robustness against diverse and unknown distortions. Our code is available athttps://github.com/king-huoye/WAQNIQA
Yejing Huo, Guoheng Huang, Zhiwen Yu 0002, Xiaochen Yuan, Chi-Man Pun, Lianglun Cheng, Xuhang Chen 0002, Zehong Chen
IEEE Trans. Syst. Man Cybern. Syst.5
2026 Effective Gaussian Management for High-Fidelity Scene Reconstruction
abstract
This paper proposes an effective Gaussian management framework for high-fidelity scene reconstruction of both appearance and geometry. Unlike recent Gaussian Splatting (GS) pipelines that treat all primitives uniformly during optimization, our framework explicitly manages the attribute activation, representation and pruning of Gaussian. Specifically, our framework first introduces GauSep, a novel densification strategy that selectively activates Gaussian color or normal attributes to alleviate destructive gradient conflicts arising from dual supervision. We further propose GauRep, an adaptive Gaussian representation that dynamically adjusts spherical harmonics (SHs) orders and performs task-decoupled pruning to reduce redundancy at both the individual and global levels. To provide reliable geometric supervision for above mangement process, we additionally introduce CoRe, an regularized surface reconstruction module that distills robust normal fields from an SDF branch to the Gaussian representation through a confidence mechanism. Notably, the proposed Gaussian management is compatible with various reconstruction architectures and can be seamlessly integrated to improve performance while reducing size of the model. Extensive experiments demonstrate that our approach achieves superior or comparable performance in appearance and geometry reconstruction compared with state-of-the-art methods, while using significantly fewer parameters.
Jiateng Liu, Hao Gao 0005, Jiucheng Xie, Chi-Man Pun, Jian Xiong 0005, Haolun Li 0001, Junxin Chen 0001, Feng Xu 0005
IEEE Trans. Vis. Comput. Graph.4
2025 Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation Localization
abstract
The mesoscopic level serves as a bridge between the macroscopic and microscopic worlds, addressing gaps overlooked by both. Image manipulation localization (IML), a crucial technique to pursue truth from fake images, has long relied on low-level (microscopic-level) traces. However, in practice, most tampering aims to deceive the audience by altering image semantics. As a result, manipulation commonly occurs at the object level (macroscopic level), which is equally important as microscopic traces. Therefore, integrating these two levels into the mesoscopic level presents a new perspective for IML research. Inspired by this, our paper explores how to simultaneously construct mesoscopic representations of micro and macro information for IML and introduces the Mesorch architecture to orchestrate both. Specifically, this architecture i) combines Transformers and CNNs in parallel, with Transformers extracting macro information and CNNs capturing micro details, and ii) explores across different scales, assessing micro and macro information seamlessly. Additionally, based on the Mesorch architecture, the paper introduces two baseline models aimed at solving IML tasks through mesoscopic representation. Extensive experiments across four datasets have demonstrated that our models surpass the current state-of-the-art in terms of performance, computational complexity, and robustness.
Xuekang Zhu, Xiaochen Ma 0001, Zhuohang Jiang, Xiwen Wang 0002, Zeyu Lei, Wentao Feng, Chi-Man Pun, Jizhe Zhou 0001
AAAI9
2025 Refined Mamba-Based Lower Limbs Motor Estimator for Parkinson's Disease Diagnostic
abstract
Bradykinesia, a key Parkinson's disease (PD) symptom, requires accurate lower limbs assessment, yet current clinical assessments are subjective and biased, while computer vision methods lack precision in skeleton extraction and PD-specific movement analysis. Furthermore, clothing-induced foot occlusion further aggravates keypoint localization errors. To address these gaps, we propose a vision-assisted diagnostic framework for PD lower limbs assessment. Our approach incorporates the CSDPose model, which employs CNN for local feature extraction, SSM (State Space Model) for global feature capture, and DCT (Discrete Cosine Transform) for frequency domain analysis, to enhance 2D pose estimation accuracy. These keypoints are used to compute objective PD motor indicators that we proposed, which are subsequently analyzed by a classification model to grade lower limbs dysfunction severity. Experiments demonstrate the algorithm achieves over$90\%$accuracy in classifying lower limbs motions. Validated clinical trials confirm that the automated severity ratings and motor indicators effectively support diagnostic decision-making.
Xinyuan Dong, Hao Gao 0005, Yikang He, Yiqin Yao, Chi-Man Pun, Haolun Li 0001, Feng Xu 0005
BIBM6
2025 DGLL: A Hybrid Global-Local Feature Learning Network for Precise Tooth Landmark Detection
abstract
The precise identification of key landmarks on three-dimensional tooth mesh models is paramount for computer-aided orthodontic treatment. However, existing methodologies exhibit limitations with respect to the integration of global and local features, which undermines accuracy in complex scenarios and excessively emphasizes relative landmark positions, resulting in displacement errors. To mitigate these issues, this study introduces DGLL, a hybrid feature learning network characterized by a dual-branch architecture that amalgamates global and local features. DGLL integrates a Cascaded Topological Relation Module (CTRM) to stabilize the extraction of global features and a Pan-scale Feature Modulation Module (PFMM) to balance relative and absolute positional accuracy. Empirical evaluations across various tooth types demonstrate that DGLL consistently enhances the accuracy of landmark localization. This research provides an effective approach to the automated analysis of tooth data, thereby improving the precision and efficacy of orthodontic treatment.
Jianwen Huang, Guoheng Huang, Fuchen Zheng, Chi-Man Pun, Ka-Cheng Choi, Lianglun Cheng, Guanghui Yue 0001
BIBM4
2025 FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI Translation
abstract
Ultra-high-field 7T MRI offers enhanced spatial resolution and tissue contrast that enable the detection of subtle pathological changes in neurological disorders. However, the limited availability of 7T scanners restricts widespread clinical adoption due to substantial infrastructure costs and technical demands. Computational approaches for synthesizing 7T-quality images from accessible 3T acquisitions present a viable solution to this accessibility challenge. Existing CNN approaches suffer from limited spatial coverage, while Transformer models demand excessive computational overhead. RWKV architectures offer an efficient alternative for global feature modeling in medical image synthesis, combining linear computational complexity with strong long-range dependency capture. Building on this foundation, we propose Frequency Spatial-RWKV (FS-RWKV), an RWKV-based framework for 3T-to-7T MRI translation. To better address the challenges of anatomical detail preservation and global tissue contrast recovery, FS- RWKV incorporates two key modules: (1) Frequency-Spatial Omnidirectional Shift (FSO-Shift), which performs discrete wavelet decomposition followed by omnidirectional spatial shifting on the low-frequency branch to enhance global contextual representation while preserving high-frequency anatomical details; and (2) Structural Fidelity Enhancement Block (SFEB), a module that adaptively reinforces anatomical structure through frequency-aware feature fusion. Comprehensive experiments on UNC and BNU datasets demonstrate that FS- RWKV consistently outperforms existing CNN-, Transformer-, GAN-, and RWKV-based baselines across both T1 wand T2w modalities, achieving superior anatomical fidelity and perceptual quality.
Yingtie Lei, Zimeng Li 0001, Yupeng Liu 0003, Xuhang Chen 0002, Chi-Man Pun
BIBM5
2025 DTEA: Dynamic Topology Weaving and Instability-Driven Entropic Attenuation for Medical Image Segmentation
abstract
In medical image segmentation, skip connections are used to merge global context and reduce the semantic gap between encoder and decoder. Current methods often struggle with limited structural representation and insufficient contextual modeling, affecting generalization in complex clinical scenarios. We propose the DTEA model, featuring a new skip connection framework with the Semantic Topology Reconfiguration (STR) and Entropic Perturbation Gating (EPG) modules. STR reorganizes multi-scale semantic features into a dynamic hypergraph to better model cross-resolution anatomical dependencies, enhancing structural and semantic representation. EPG assesses channel stability after perturbation and filters high-entropy channels to emphasize clinically important regions and improve spatial attention. Extensive experiments on three benchmark datasets show our framework achieves superior segmentation accuracy and better generalization across various clinical settings. The code is available at https://github.com/LWX-Research/DTEA.
Quanjun Li, Zimeng Li 0001, Chi-Man Pun, Yupeng Liu 0003, Xuhang Chen 0002
BIBM6
2025 A Noise-Resistant 3D Hand Motion Estimator Framework for Parkinson's Tremor Assessment
abstract
Parkinson's disease (PD) is a progressive neurodegenerative disorder, with tremor being one of its representative motor symptoms. Current clinical evaluations primarily rely on subjective scales such as the MDS-UPDRS, which often introduce significant inter-rater variability. Although vision-based evaluation offers objective motion analysis, existing pose tracking frameworks struggle with accurate tremor quantification due to inter-frame jitter and limited precision, failing to capture fine-grained spatiotemporal dynamics. To address this, we propose Motion-aware Hierarchical Grouping Mamba Network (MHG-Mamba), a non-contact video-based framework for automated evaluation of the 'finger-to-nose' task. Our approach first employs feature extraction to decouple high-degree-of-freedom finger joint movements from stable palm joint motions, enabling accurate finger pose estimation. Second, by incorporating with the hierarchical spatiotemporal scanning mechanism in Mamba's SSM, the model captures global motion features while preserving anatomical constraints, resulting in a temporally smooth and plausible skeletal sequence. Finally, based on the predicted skeletal sequences, we introduce several objective metrics to quantify motion features and apply a classifier for precise objective severity rating. Experimental results demonstrate that MHG-Mamba significantly improves the accuracy of 3D hand pose estimation and reduces noise in the motion sequences. The system achieved a classification accuracy of 93.2 % on the 'finger-to-nose' task. Moreover, clinicians using our system exhibited reduced variability in their assessments, highlighting its high clinical value.
Yixing Ye, Hao Gao 0005, Yikang He, Haolun Li 0001, Chi-Man Pun, Feng Xu 0005
BIBM6
2025 LASGA: Lesion-Aware Saliency-Guided Adversarial Attack for Diabetic Retinopathy Grading
abstract
Diabetic retinopathy (DR) is a leading cause of blindness worldwide, requiring accurate lesion classification for proactive intervention. While deep learning models have improved diagnostic accuracy, their sensitivity to subtle noise challenges their clinical reliability. This paper examines the robustness and interpretability of DR grading models under adversarial attacks. We propose a novel Lesion-Aware Saliency-Guided Adversarial Attack (LASGA) that restricts perturbations to lesion regions. Our approach introduces a Lesion-aware Saliency Map Generation (LSMG) module that incorporates lesion awareness into static saliency maps to guide perturbation optimization, allowing perturbations to focus on key regions crucial for DR decision-making while reducing artifacts in smooth backgrounds. We also design a Dual Strategic Perturbation Mechanism (DSPM) that combines prediction loss and consistency loss to generate adversarial examples that mislead diagnosis while remaining imperceptible. To systematically evaluate model robustness and interpretability, we establish the first adversarial robustness benchmark for DR grading, integrating traditional attacks, LASGA, and hybrid approaches. Extensive experiments on four models across three datasets demonstrate that our method outperforms baseline attacks in both white-box and black-box settings, providing valuable insights for enhancing the security of DR grading in clinical practice.
Yijia Zeng, Chi-Man Pun
BIBM2
2025 HBFormer: A Hybrid-Bridge Transformer for Microtumor and Miniature Organ Segmentation
abstract
Medical image segmentation is a cornerstone of modern clinical diagnostics. While Vision Transformers that leverage shifted window-based self-attention have established new benchmarks in this field, they are often hampered by a critical limitation: their localized attention mechanism struggles to effectively fuse local details with global context. This deficiency is particularly detrimental to challenging tasks such as the segmentation of microtumors and miniature organs, where both finegrained boundary definition and broad contextual understanding are paramount. To address this gap, we propose HBFormer, a novel Hybrid-Bridge Transformer architecture. The 'Hybrid' design of HBFormer synergizes a classic U-shaped encoder-decoder framework with a powerful Swin Transformer backbone for robust hierarchical feature extraction. The core innovation lies in its 'Bridge' mechanism, a sophisticated nexus for multi-scale feature integration. This bridge is architecturally embodied by our novel Multi-Scale Feature Fusion (MFF) decoder. Departing from conventional symmetric designs, the MFF decoder is engineered to fuse multi-scale features from the encoder with global contextual information. It achieves this through a synergistic combination of channel and spatial attention modules, which are constructed from a series of dilated and depth-wise convolutions. These components work in concert to create a powerful feature bridge that explicitly captures long-range dependencies and refines object boundaries with exceptional precision. Comprehensive experiments on challenging medical image segmentation datasets, including multi-organ, liver tumor, and bladder tumor benchmarks, demonstrate that HBFormer achieves state-of-theart results, showcasing its outstanding capabilities in microtumor and miniature organ segmentation. Code and models are available at: https://github.com/lzeeorno/HBFormer.
Fuchen Zheng, Quanjun Li, Junhua Zhou, Xiaojiao Guo, Xuhang Chen 0002, Chi-Man Pun, Shoujun Zhou
BIBM8
2025 Enhancing the Transferability of Adversarial Examples Against No-Reference Image Quality Assessment Models
Yijia Zeng, Xuhang Chen 0002, Chi-Man Pun
CGI (3)3
2025 An Adaptive Framework for Multi-View Clustering Leveraging Conditional Entropy Optimization
abstract
Multi-view clustering (MVC) has emerged as a powerful technique for extracting valuable insights from data characterized by multiple perspectives or modalities. Despite significant advancements, existing MVC methods struggle with effectively quantifying the consistency and complementarity among views, and are particularly susceptible to the adverse effects of noisy views, known as the Noisy-View Drawback (NVD). To address these challenges, we propose CE-MVC, a novel framework that integrates an adaptive weighting algorithm with a parameter-decoupled deep model. Leveraging the concept of conditional entropy and normalized mutual information, CE-MVC quantitatively assesses and weights the informative contribution of each view, facilitating the construction of robust unified representations. The parameter-decoupled design enables independent processing of each view, effectively mitigating the influence of noise and enhancing overall clustering performance. Extensive experiments demonstrate that CE-MVC outperforms existing approaches, offering a more resilient and accurate solution for multi-view clustering tasks.
Lijian Li 0003, Yuanpeng He, Chi-Man Pun
ICASSP3
2025 Underwater Image Restoration via Polymorphic Large Kernel CNNs
abstract
Underwater Image Restoration (UIR) remains a challenging task in computer vision due to the complex degradation of images in underwater environments. While recent approaches have leveraged various deep learning techniques, including Transformers and complex, parameter-heavy models to achieve significant improvements in restoration effects, we demonstrate that pure CNN architectures with lightweight parameters can achieve comparable results. In this paper, we introduce UIR-PolyKernel, a novel method for underwater image restoration that leverages Polymorphic Large Kernel CNNs. Our approach uniquely combines large kernel convolutions of diverse sizes and shapes to effectively capture long-range dependencies within underwater imagery. Additionally, we introduce a Hybrid Domain Attention module that integrates frequency and spatial domain attention mechanisms to enhance feature importance. By leveraging the frequency domain, we can capture hidden features that may not be perceptible to humans but are crucial for identifying patterns in both underwater and on-air images. This approach enhances the generalization and robustness of our UIR model. Extensive experiments on benchmark datasets demonstrate that UIR-PolyKernel achieves state-of-the-art performance in underwater image restoration tasks, both quantitatively and qualitatively. Our results show that well-designed pure CNN architectures can effectively compete with more complex models, offering a balance between performance and computational efficiency. This work provides new insights into the potential of CNN-based approaches for challenging image restoration tasks in underwater environments. The code is available at https://github.com/CXH-Research/UIR-PolyKernel.
Xiaojiao Guo, Yihang Dong, Xuhang Chen 0002, Weiwen Chen, Zimeng Li 0001, Fuchen Zheng, Chi-Man Pun
ICASSP7
2025 Lorentz Transformation Neural Network
abstract
We propose a novel neural network architecture, the Lorentz Transformation Neural Network (LTNN), which utilizes Lorentz transformations to generate a complex computation matrix that enhances the network’s expressive power. Furthermore, LTNN is lightweight due to the shared weight matrices in the computation matrix. LTNN treats the input and output as coordinates in high-dimensional spacetime, with the weight matrices in each layer representing the velocity components of a spacetime reference frame. During training, these weight matrices are transformed into a computation matrix via Lorentz transformations, describing the coordinate transformations between different reference frames. We evaluate LTNN on four datasets: California Housing Prices, Iris, MNIST, and Fashion-MNIST. Experimental results demonstrate that LTNN outperforms conventional neural networks and quaternion neural networks in terms of both accuracy and parameter efficiency.
Wenyuan Li 0007, Jingchao Wang 0002, Guoheng Huang, Tongxu Lin, Guo Zhong, Xiaochen Yuan, Chi-Man Pun, An Zeng
ICIP7
2025 ISSD-NET: Intra-student Self-distillation with Adaptive q-vMF Loss for Enhanced Semi-supervised Medical Segmentation
Guoheng Huang, Xiaochen Yuan, Yan Li 0122, Chi-Man Pun, Bai Ying Lei
ICONIP (2)5
2025 Superpixel-Enhanced Quaternion Feature Fusion and Contextualization Graph Contrastive Learning for Cervical Cancer Diagnosis
Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Lianglun Cheng, Chi-Man Pun, Guo Zhong, Qingjian Ye
ICONIP (2)7
2025 LensNet: An End-to-End Learning Framework for Empirical Point Spread Function Modeling and Lensless Imaging Reconstruction
abstract
Lensless imaging stands out as a promising alternative to conventional lens-based systems, particularly in scenarios demanding ultracompact form factors and cost-effective architectures. However, such systems are fundamentally governed by the Point Spread Function (PSF), which dictates how a point source contributes to the final captured signal. Traditional lensless techniques often require explicit calibrations and extensive pre-processing, relying on static or approximate PSF models. These rigid strategies can result in limited adaptability to real-world challenges, including noise, system imperfections, and dynamic scene variations, thus impeding high-fidelity reconstruction. In this paper, we propose LensNet, an end-to-end deep learning framework that integrates spatial-domain and frequency-domain representations in a unified pipeline. Central to our approach is a learnable Coded Mask Simulator (CMS) that enables dynamic, data-driven estimation of the PSF during training, effectively mitigating the shortcomings of fixed or sparsely calibrated kernels. By embedding a Wiener filtering component, LensNet refines global structure and restores fine-scale details, thus alleviating the dependency on multiple handcrafted pre-processing steps. Extensive experiments demonstrate LensNet's robust performance and superior reconstruction quality compared to state-of-the-art methods, particularly in preserving high-frequency details and attenuating noise. The proposed framework establishes a novel convergence between physics-based modeling and data-driven learning, paving the way for more accurate, flexible, and practical lensless imaging solutions for applications ranging from miniature sensors to medical diagnostics. The link of code is https://github.com/baijiesong/Lensnet.
Jiesong Bai, Yuhao Yin, Yihang Dong, Xiaofeng Zhang 0006, Chi-Man Pun, Xuhang Chen 0002
IJCAI5
2025 Code Retrieval with Mixture of Experts Prototype Learning Based on Classification
abstract
The semantic connection between code and queries is crucial for code retrieval, but many human-written queries fail to accurately capture the code's core intent, leading to ambiguity.This ambiguity complicates the code search process, as the queries do not provide a clear overview of the code's purpose.Our analysis reveals that while ambiguous queries may not precisely summarize the intent of the code, they often share the same general topics as the corresponding code.In light of this discovery, we propose Code Retrieval with Mixture of Experts Prototype Learning Based on Classification (CRME), a novel approach that combines classification for prototype-based representation learning and result ensembling.CRME utilizes specialized pre-trained models focused on the specific domains of ambiguous queries.It consists of two key components: Multiple Classification Prototype and Representation Learning with a Prototype-based Multi-model Contrastive (PMC) Loss during training, and Multi-Prototype Mixture of Experts Integration (MP-MoE) module for fine-grained ensemble inference.Our method can effectively address the issue of query ambiguity and improves search precision.Experimental results on the CodeSearchNet dataset, covering six sub-datasets, show that CRME outperforms existing methods, achieving an average MRR score of * Corresponding authors.
Feng Ling 0002, Guoheng Huang, Jingchao Wang 0002, Xiaochen Yuan, Xuhang Chen 0002, XueYong Zhang, Fanlong Zhang, Chi-Man Pun
Internetware8
2025 Evidential Prototype Learning for Semi-supervised Medical Image Segmentation
abstract
Although current semi-supervised medical segmentation methods can achieve decent performance, they are still affected by the uncertainty in unlabeled data and model predictions, and there is currently a lack of effective strategies that can explore the uncertain aspects of both simultaneously. To address the aforementioned issues, we propose Evidential Prototype Learning (EPL), which utilizes an extended probabilistic framework to effectively fuse voxel-level evidential predictions from different classifiers and achieves prototype fusion utilization of labeled and unlabeled data under a generalized evidential framework, leveraging voxel-level dual uncertainty masking. The uncertainty measure not only enables the model to self-correct predictions but also improves the guided learning process with pseudo-labels and is able to feed back into the construction of hidden features. The method proposed in this paper has been experimented on LA, Pancreas-CT and TBAD datasets, achieving the state-of-the-art performance in three different labeled ratios, which strongly demonstrates the effectiveness of our strategy. The source code will be made publicly available.
Yuanpeng He, Lijian Li 0003, Tianxiang Zhan, Chi-Man Pun, Wenpin Jiao, Zhi Jin 0001
KDD (2)4
2025 MotionRefineNet: Fine-Grained Pose Sequence Smoothing and Refinement
abstract
Capturing human motion with existing monocular estimators often results in large errors when dealing with rare poses, occlusions, truncations, and frame blurring, leading to jitter and long-term drift. Although previous methods have introduced post-processing networks for pose refinement, they struggle to balance global smoothing and fine-grained correction. In this work, we propose MotionRefineNet, which leverages the synergy and complementarity between long- and short-term features in the temporal domain and high- and low-frequency features in the frequency domain to address these challenges. The temporal branch is designed as a hierarchical motion structure to learn multi-time scale features, where long-term features learn motion smoothness, and short-term features capture local rapid changes. The frequency branch employs different frequency band learning strategies based on the degrees of freedom (DoF) of body parts. For body parts with low DoF, the focus is on low-frequency features that represent overall motion trends and regular actions. For body parts with high DoF, we design a filter to adaptively extract useful information from all frequency bands, including subtle motion changes in the high-frequency bands. Extensive experiments on multiple datasets and estimators demonstrate that MotionRefineNet outperforms existing methods in refining 2D, 3D, and SMPL poses, achieving superior pose smoothing and deviation correction. Our code is available at: https://github.com/Wheels319/MotionRefineNet.
Haolun Li 0001, Weihuang Liu, Jiateng Liu, Zhenhua Tang 0001, Chi-Man Pun, Qiguang Miao, Feng Xu 0005, Hao Gao 0005
ACM Multimedia5
2025 I-C Attack: In-place and Cross-pixel Augmentations for Highly Transferable Transformation-based Attacks
Chi-Man Pun
ACM Multimedia2
2025 FGRFlow: Learning Fine-Grained Rigidity Scene Flow from 4D Radar Point Cloud
abstract
Scene flow estimation using 4D millimeter-wave radar has emerged as a prominent research focus for 3D dynamic perception. However, compared to LiDAR point clouds, the drastic sparsity of radar point clouds poses challenges in enforcing local rigidity constraints, which are crucial for accurate 3D motion estimation. To address this issue, we propose a novel Gaussian-based pseudo-point generation method that fully leverages two distinct yet complementary data modalities, 3D coordinates and Doppler velocity, to support multi-body rigidity assumptions, effectively capturing fine-grained and structured motion patterns from highly sparse radar point clouds. Furthermore, a velocity calibration mechanism is designed to improve the reliability of fine-grained rigid motion velocity estimation. In addition, a progressive fusion strategy is introduced to systematically integrate fine-grained rigid motion priors at multiple levels, enhancing the robustness of matching costs and motion features while effectively compensating for coarse flows. Experimental results on real-world radar scans from the View-of-Delft (VoD) dataset demonstrate the promising performance of our FGRFlow compared to other leading 4D radar-based approaches, validating the advantages of our design choices.
Mingliang Zhai, Haidong Hu, Chi-Man Pun, Hao Gao 0005
ACM Multimedia4
2025 Cross-View Geo-Localization via Learning Correspondence Semantic Similarity Knowledge
Guanli Chen, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun
MMM (1)6
2025 ForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization
abstract
The field of Fake Image Detection and Localization (FIDL) is highly fragmented, encompassing four domains: deepfake detection (Deepfake), image manipulation detection and localization (IMDL), artificial intelligence-generated image detection (AIGC), and document image manipulation localization (Doc). Although individual benchmarks exist in some domains, a unified benchmark for all domains in FIDL remains blank. The absence of a unified benchmark results in significant domain silos, where each domain independently constructs its datasets, models, and evaluation protocols without interoperability, preventing cross-domain comparisons and hindering the development of the entire FIDL field. To close the domain silo barrier, we propose ForensicHub, the first unified benchmark & codebase for all-domain fake image detection and localization. Considering drastic variations on dataset, model, and evaluation configurations across all domains, as well as the scarcity of open-sourced baseline models and the lack of individual benchmarks in some domains, ForensicHub: i) proposes a modular and configuration-driven architecture that decomposes forensic pipelines into interchangeable components across datasets, transforms, models, and evaluators, allowing flexible composition across all domains; ii) fully implements 10 baseline models (3 of which are reproduced from scratch), 6 backbones, 2 new benchmarks for AIGC and Doc, and integrates 2 existing benchmarks of DeepfakeBench and IMDLBenCo through an adapter-based design; iii) establishes an image forensic fusion protocol evaluation mechanism that supports unified training and testing of diverse forensic models across tasks; iv) conducts indepth analysis based on the ForensicHub, offering 8 key actionable insights into FIDL model architecture, dataset characteristics, and evaluation standards. Specifically, ForensicHub includes 4 forensic tasks, 23 datasets, 42 baseline models, 6 backbones, 11 GPU-accelerated pixel- and image-level evaluation metrics, and realizes 16 kinds of cross-domain evaluations. ForensicHub represents a significant leap forward in breaking the domain silos in the FIDL field and inspiring future breakthroughs. Code is available at: https://github.com/scu-zjz/ForensicHub.
Xuekang Zhu, Xiaochen Ma 0001, Chenfan Qu, Kaiwen Feng, Chi-Man Pun, Jizhe Zhou 0001
NeurIPS7
2025 SFormer: SNR-Guided Transformer for Underwater Image Enhancement from the Frequency Domain
Yingtie Lei, Zimeng Li 0001, Chi-Man Pun, Xuhang Chen 0002
PRICAI (5)5
2025 The Structure-sharing Hypergraph Reasoning Attention Module for CNNs
Jingchao Wang 0002, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Tongxu Lin, Chi-Man Pun, Fenfang Xie
Expert Syst. Appl.6
2025 Privacy-Aware Secure Data Auditing for Cloud-Based Intelligence of Things Environment
abstract
Cloud-based Intelligence of Things is significant for Augmented Enterprise Management Systems. Data integrity auditing is challenging in the intelligence of things environment, mainly when the newer versions in the public cloud environment update existing encrypted data. The related literature on cloud-based intelligence relies on encrypted data uploading or locally handling encryption and decryption using user keys. Considering the security risk, storage constraints at the edge, and realtime environment, both approaches have limited applicability in the intelligence of things environment. This paper presents the Privacy-Aware Secure Data Auditing (PASDA) framework at the cluster head for online data integrity verification. Specifically, the users hide data files by the blinding process with a generation of their corresponding signatures, which achieves data auditing by utilizing homomorphic techniques. A novel automated self-triggering/ Self-auditing-based data integrity auditing system is proposed, which detects the changes made in the cloud-stored data and sends alert messages to the trusted primary cloud server and users. A data dynamics method is developed containing a timestamp with a pointer to store multiple versions of the same file without signatures re-generation for the whole same file. The user is revoked due to prolonged absence or detection of the missed behaviour with system or service expiry. With these data dynamics, the proposed PASDA framework allows CH to regenerate signatures of the revoked user using its membership key for cloud-based stored data access and data integrity auditing. In-depth security analysis and extensive simulations based on comparative performance evaluation attest to the benefits of the proposed PA
Fasee Ullah, Chi-Man Pun, Muhammad Ismail Mohmand, Rakesh Kumar Mahendran, Arfat Ahmad Khan, Sarah M. Alhammad, Joel J. P. C. Rodrigues, Ahmed Farouk
IEEE Internet Things J.2
2025 Visual-linguistic Diagnostic Semantic Enhancement for medical report generation
Jiahong Chen, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Zhe Tan, Chi-Man Pun
J. Biomed. Informatics6
2025 MPCM-RRG: Multi-modal Prompt Collaboration Mechanism for Radiology Report Generation
Yumian Yu, Guoheng Huang, Zhe Tan, Ming Li 0065, Chi-Man Pun, Fuchen Zheng, Shiqiang Ma, Shuqiang Wang
J. Biomed. Informatics6
2025 Co-evidential fusion with information volume for semi-supervised medical image segmentation
abstract
Although existing semi-supervised image segmentation methods have achieved good performance, they cannot effectively utilize multiple sources of voxel-level uncertainty for targeted learning. Therefore, we propose two main improvements. First, we introduce a novel pignistic co-evidential fusion strategy using generalized evidential deep learning , extended by traditional D–S evidence theory, to obtain a more precise uncertainty measure for each voxel in medical samples. This assists the model in learning mixed labeled information and establishing semantic associations between labeled and unlabeled data. Second, we introduce the concept of information volume of mass function (IVUM) to evaluate the constructed evidence, implementing two evidential learning schemes. One optimizes evidential deep learning by combining the information volume of the mass function with original uncertainty measures. The other integrates the learning pattern based on the co-evidential fusion strategy, using IVUM to design a new optimization objective. Experiments on four datasets demonstrate the competitive performance of our method.
Yuanpeng He, Lijian Li 0003, Tianxiang Zhan, Chi-Man Pun, Wenpin Jiao, Zhi Jin 0001
Pattern Recognit.4
2025 Learning hyperspectral noisy label with global and local hypergraph laplacian energy
Cheng Shi 0002, Linfeng Lu, Minghua Zhao, Xinhong Hei 0001, Chi-Man Pun, Qiguang Miao
Pattern Recognit.5
2025 Lifespan age synthesis on human faces with decorrelation constraints and geometry guidance
Jiucheng Xie, Lingqing Zhang, Hao Gao 0005, Chi-Man Pun
Pattern Recognit. Lett.4
2025 Underwater Image Restoration Through a Prior Guided Hybrid Sense Approach and Extensive Benchmark Analysis
abstract
Underwater imaging grapples with challenges from light-water interactions, leading to color distortions and reduced clarity. In response to these challenges, we propose a novel Color Balance Prior Guided Hybrid Sense Underwater Image Restoration framework (GuidedHybSensUIR). This framework operates on multiple scales, employing the proposed Detail Restorer module to restore low-level detailed features at finer scales and utilizing the proposed Feature Contextualizer module to capture long-range contextual relations of high-level general features at a broader scale. The hybridization of these different scales of sensing results effectively addresses color casts and restores blurry details. In order to effectively point out the evolutionary direction for the model, we propose a novel Color Balance Prior as a strong guide in the feature contextualization step and as a weak guide in the final decoding phase. We construct a comprehensive benchmark using paired training data from three real-world underwater datasets and evaluate on six test sets, including three paired and three unpaired, sourced from four real-world underwater datasets. Subsequently, we tested 14 traditional and retrained 23 deep learning existing underwater image restoration methods on this benchmark, obtaining metric results for each approach. This effort aims to furnish a valuable benchmarking dataset for standard basis for comparison. The extensive experiment results demonstrate that our method outperforms 37 other state-of-the-art methods overall on various benchmark datasets and metrics, despite not achieving the best results in certain individual cases. The code and dataset are available at https://github.com/CXH-Research/GuidedHybSensUIR.
Xiaojiao Guo, Xuhang Chen 0002, Shuqiang Wang, Chi-Man Pun
IEEE Trans. Circuits Syst. Video Technol.4
2025 Boundary-Aware Sentence-Gloss Alignment With Semantic Similarity Measurement for Continuous Sign Language Recognition
abstract
Continuous sign language recognition (CSLR) plays a crucial role in facilitating communication between deaf and hearing individuals. A key aspect of achieving precise CSLR is the alignment of the video segment of each sign with its gloss, namely its corresponding text representation in natural language. However, the coarticulation phenomenon, where contextual dependencies between adjacent signs blur the boundaries of individual signs, poses a significant challenge to this task. In this paper, we propose a novel boundary-aware sentence-gloss alignment network for CSLR to address this challenge. Our network first designs a task-relevant boundary-aware similarity measurement, evaluating sign frames by both appearance and their recognition contribution, mitigating coarticulation-induced transition noise to restore precise boundaries. For enhanced alignment, we propose a hierarchical sentence-gloss alignment: coarse sentence-level alignment reduces cross-modal disparity, while fine-grained gloss-level alignment refines video-to-token mapping. Finally, an adaptive class-divergence loss sharpens gloss decoding by maximizing inter-class discrimination. Our proposed framework provides a simple and effective solution to mitigate the boundary ambiguity caused by coarticulation, optimizing continuous sign language recognition algorithms from a new perspective. Extensive experiments conducted on four public sign language recognition (SLR) datasets demonstrate that our proposed boundary-aware sentence-gloss alignment network learns precise alignments and achieves state-of-the-art performance.
Yunan Li 0001, Xi Geng, Zhuoqi Ma, Qiguang Miao, Chi-Man Pun
IEEE Trans. Circuits Syst. Video Technol.5
2025 Spatial-Aware Conformal Prediction for Trustworthy Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification involves assigning unique labels to each pixel to identify various land cover categories. While deep classifiers have achieved high predictive accuracy in this field, they lack the ability to rigorously quantify confidence in their predictions. This limitation restricts their application in critical contexts where the cost of prediction errors is significant, as quantifying the uncertainty of model predictions is crucial for the safe deployment of predictive models. To address this limitation, a rigorous theoretical proof is presented first, which demonstrates the validity of Conformal Prediction, an emerging uncertainty quantification technique, in the context of HSI classification. Building on this foundation, a conformal procedure is designed to equip any pre-trained HSI classifier with trustworthy prediction sets, ensuring that the true labels are included with a user-defined probability (e.g., 95%). Furthermore, a novel framework of Conformal Prediction specifically designed for HSI data, called Spatial-Aware Conformal Prediction (SACP), is proposed. This framework integrates essential spatial information of HSI by aggregating the non-conformity scores of pixels with high spatial correlation, effectively improving the statistical efficiency of prediction sets. Both theoretical and empirical results validate the effectiveness of the proposed approaches. The source code is available at https://github.com/J4ckLiu/SACP.
Kangdao Liu, Tianhao Sun, Hao Zeng 0005, Yongshan Zhang, Chi-Man Pun, Chi-Man Vong
IEEE Trans. Circuits Syst. Video Technol.5
2025 DP-TRAE: A Dual-Phase Merging Transferable Reversible Adversarial Example for Image Privacy Protection
abstract
In the field of digital security, Reversible Adversarial Examples (RAE) combine adversarial attacks with reversible data hiding techniques to effectively protect sensitive data and prevent unauthorized analysis by malicious Deep Neural Networks (DNNs). However, existing RAE techniques primarily focus on white-box attacks, lacking a comprehensive evaluation of their effectiveness in black-box scenarios. This limitation impedes their broader deployment in complex, dynamic environments. Furthermore, traditional black-box attacks are often characterized by poor transferability and high query costs, significantly limiting their practical applicability. To address these challenges, we propose the Dual-Phase Merging Transferable Reversible Attack method, which generates highly transferable initial adversarial perturbations in a white-box model and employs a memory-augmented black-box strategy to effectively mislead target models. Experimental results demonstrate the superiority of our approach, achieving a 99.0% attack success rate and 100% recovery rate in black-box scenarios with the DN-121 target model and 1000 attack iterations, highlighting its robustness in privacy protection. Moreover, we successfully implemented a black-box attack on a commercial model, further substantiating the potential of this approach for practical use.
Xia Du, Jizhe Zhou 0001, Chi-Man Pun, Zheng Lin 0001, Cong Wu 0003, Zhe Chen 0015, Jun Luo 0001
IEEE Trans. Dependable Secur. Comput.4
2025 Adaptive Multitype Contrastive Views Generation for Remote Sensing Image Semantic Segmentation
abstract
Self-supervised contrastive learning is a powerful pre-training framework for learning the invariant features from the different views of remote sensing images, therefore, the performance of contrastive learning heavily depends on the generation of views. Current view generation is primarily accomplished through different transformations, and the types and parameters of the transformations are require hand-crafted. Hence, the diversity and discriminability of generated views cannot be guaranteed. To address this, we propose a multi-type views optimization method to optimize these transformations. We formulate contrastive learning as a min-max optimization problem, and transformation parameters are optimized by maximizing the contrastive loss. The optimized transformations encourage the negative sample pairs to be close and the positive sample pairs to be far apart. Different from the current adversarial view generation methods, our method can optimize both photometric transformations and geometric transformations. For remote sensing images, the geometric transformation is more critical for view generation, while the existing view optimization methods fail to achieve this. We consider the hue, saturation, brightness, contrast, and geometric rotation transformations in contrastive learning, and evaluate the optimized views on the downstream remote sensing images semantic segmentation task. Extensive experiments are carried on the three remote sensing image segmentation datasets, including ISPRS Potsdam dataset, ISPRS Vaihingen dataset, and LoveDA dataset. Results show that the learned views obtain highly advantages compared to the hand-crafted views and other optimized views. The code associated with this paper has been released and can be accessed at https://github.com/AAAA-CS/AMView.
Cheng Shi 0002, Peiwen Han, Minghua Zhao, Qiguang Miao, Chi-Man Pun
IEEE Trans. Geosci. Remote. Sens.6
2025 A Two-Phase Scheme by Integration of Deep and Corner Feature for Balanced Copy-Move Forgery Localization
abstract
In the era of Industry 4.0, the widespread application of digitization, automation, and Internet technology in industrial production has led to a significant increase in image data. Image security has become crucial because images are at risk of being tampered with at any time. To protect its authenticity, this article proposes a two-phase scheme to achieve balanced performance between accuracy and speed for copy-move forgery detection. Our scheme is divided into detection and localization phases. In the detection phase, the deep features are utilized to calculate the inner similarity. To improve the accuracy, a corner point matching technique is performed on the localization phase as a refinement step. The experimental results demonstrate the average$F1$-score is 0.6334 on CASIA2.0, making a 14.16% improvement. The computation time for each image is only 0.791 s in average. It has great significance in protecting the reliability and authenticity of industrial data.
Tong Liu 0021, Xiaochen Yuan, Zhiyao Xie, Kaiqi Zhao 0004, Guoheng Huang, Chi-Man Pun
IEEE Trans. Ind. Informatics6
2025 Adaptive Pitfall: Exploring the Effectiveness of Adaptation in Skeleton-Based Action Recognition
abstract
Graph convolution networks (GCNs) have achieved remarkable performance in skeleton-based action recognition by exploiting the adjacency topology of body representation. However, the adaptive strategy adopted by the previous methods to construct the adjacency matrix is not balanced between the performance and the computational cost. We assume this concept ofAdaptive Trap, which can be replaced by multiple autonomous submodules, thereby simultaneously enhancing the dynamic joint representation and effectively reducing network resources. To effectuate the substitution of the adaptive model, we unveil two distinct strategies, both yielding comparable effects. (1) Optimization.Individuality and Commonality GCNs (IC-GCNs)is proposed to specifically optimize the construction method of the associativity adjacency matrix for adaptive processing. The uniqueness and co-occurrence between different joint points and frames in the skeleton topology are effectively captured through methodologies like preferential fusion of physical information, extreme compression of multi-dimensional channels, and simplification of self-attention mechanism. (2) Replacement.Auto-Learning GCNs (AL-GCNs)is proposed to boldly remove popular adaptive modules and cleverly utilize human key points as motion compensation to provide dynamic correlation support. AL-GCNs construct a fully learnable group adjacency matrix in both spatial and temporal dimensions, resulting in an elegant and efficient GCN-based model. In addition, three effective tricks for skeleton-based action recognition (Skip-Block, Bayesian Weight Selection Algorithm, and Simplified Dimensional Attention) are exposed and analyzed in this paper. Finally, we employ the variable channel and grouping method to explore the hardware resource bound of the two proposed models. IC-GCN and AL-GCN exhibit impressive performance across NTU-RGB+D 60, NTU-RGB+D 120, NW-UCLA, and UAV-Human datasets, with an exceptional parameter-cost ratio.
Qiguang Miao, Wentian Xin, Ruyi Liu 0001, Cheng Shi 0002, Chi-Man Pun
IEEE Trans. Multim.7
2025 GaussianHead: High-Fidelity Head Avatars With Learnable Gaussian Derivation
abstract
Creating lifelike 3D head avatars and generating compelling animations for diverse subjects remain challenging in computer vision. This paper presents GaussianHead, which models the active head based on anisotropic 3D Gaussians. Our method integrates a motion deformation field and a single-resolution tri-plane to capture the head's intricate dynamics and detailed texture. Notably, we introduce a customized derivation scheme for each 3D Gaussian, facilitating the generation of multiple "doppelgangers" through learnable parameters for precise position transformation. This approach enables efficient representation of diverse Gaussian attributes and ensures their precision. Additionally, we propose an inherited derivation strategy for newly added Gaussians to expedite training. Extensive experiments demonstrate GaussianHead's efficacy, achieving high-fidelity visual results with a remarkably compact model size ($\approx 12$≈12 MB). Our method outperforms state-of-the-art alternatives in tasks such as reconstruction, cross-identity reenactment, and novel view synthesis.
Jie Wang 0137, Jiucheng Xie, Xianyan Li, Feng Xu 0005, Chi-Man Pun, Hao Gao 0005
IEEE Trans. Vis. Comput. Graph.5
2025 Psanet: prototype-guided salient attention for few-shot segmentation
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun
Vis. Comput.7
2025 Weakly supervised semantic segmentation via saliency perception with uncertainty-guided noise suppression
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Guo Zhong, Xuhang Chen 0002, Chi-Man Pun
Vis. Comput.7
2024 COMMA: Co-articulated Multi-Modal Learning
abstract
Pretrained large-scale vision-language models such as CLIP have demonstrated excellent generalizability over a series of downstream tasks. However, they are sensitive to the variation of input text prompts and need a selection of prompt templates to achieve satisfactory performance. Recently, various methods have been proposed to dynamically learn the prompts as the textual inputs to avoid the requirements of laboring hand-crafted prompt engineering in the fine-tuning process. We notice that these methods are suboptimal in two aspects. First, the prompts of the vision and language branches in these methods are usually separated or uni-directionally correlated. Thus, the prompts of both branches are not fully correlated and may not provide enough guidance to align the representations of both branches. Second, it's observed that most previous methods usually achieve better performance on seen classes but cause performance degeneration on unseen classes compared to CLIP. This is because the essential generic knowledge learned in the pretraining stage is partly forgotten in the fine-tuning process. In this paper, we propose Co-Articulated Multi-Modal Learning (COMMA) to handle the above limitations. Especially, our method considers prompts from both branches to generate the prompts to enhance the representation alignment of both branches. Besides, to alleviate forgetting about the essential knowledge, we minimize the feature discrepancy between the learned prompts and the embeddings of hand-crafted prompts in the pre-trained CLIP in the late transformer layers. We evaluate our method across three representative tasks of generalization to novel classes, new target datasets and unseen domain shifts. Experimental results demonstrate the superiority of our method by exhibiting a favorable performance boost upon all tasks with high efficiency. Code is available at https://github.com/hulianyuyy/COMMA.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Chi-Man Pun, Wei Feng 0005
AAAI4
2024 Devignet: High-Resolution Vignetting Removal via a Dual Aggregated Fusion Transformer with Adaptive Channel Expansion
abstract
Vignetting commonly occurs as a degradation in images resulting from factors such as lens design, improper lens hood usage, and limitations in camera sensors. This degradation affects image details, color accuracy, and presents challenges in computational photography. Existing vignetting removal algorithms predominantly rely on ideal physics assumptions and hand-crafted parameters, resulting in the ineffective removal of irregular vignetting and suboptimal results. Moreover, the substantial lack of real-world vignetting datasets hinders the objective and comprehensive evaluation of vignetting removal. To address these challenges, we present VigSet, a pioneering dataset for vignetting removal. VigSet includes 983 pairs of both vignetting and vignetting-free high-resolution (over 4k) real-world images under various conditions. In addition, We introduce DeVigNet, a novel frequency-aware Transformer architecture designed for vignetting removal. Through the Laplacian Pyramid decomposition, we propose the Dual Aggregated Fusion Transformer to handle global features and remove vignetting in the low-frequency domain. Additionally, we propose the Adaptive Channel Expansion Module to enhance details in the high-frequency domain. The experiments demonstrate that the proposed model outperforms existing state-of-the-art methods. The code, models, and dataset are available at https://github.com/CXH-Research/DeVigNet.
Shenghong Luo, Xuhang Chen 0002, Weiwen Chen, Zinuo Li, Shuqiang Wang, Chi-Man Pun
AAAI6
2024 PVALane: Prior-Guided 3D Lane Detection with View-Agnostic Feature Alignment
abstract
Monocular 3D lane detection is essential for a reliable autonomous driving system and has recently been rapidly developing. Existing popular methods mainly employ a predefined 3D anchor for lane detection based on front-viewed (FV) space, aiming to mitigate the effects of view transformations. However, the perspective geometric distortion between FV and 3D space in this FV-based approach introduces extremely dense anchor designs, which ultimately leads to confusing lane representations. In this paper, we introduce a novel prior-guided perspective on lane detection and propose an end-to-end framework named PVALane, which utilizes 2D prior knowledge to achieve precise and efficient 3D lane detection. Since 2D lane predictions can provide strong priors for lane existence, PVALane exploits FV features to generate sparse prior anchors with potential lanes in 2D space. These dynamic prior anchors help PVALane to achieve distinct lane representations and effectively improve the precision of PVALane due to the reduced lane search space. Additionally, by leveraging these prior anchors and representing lanes in both FV and bird-eye-viewed (BEV) spaces, we effectively align and merge semantic and geometric information from FV and BEV features. Extensive experiments conducted on the OpenLane and ONCE-3DLanes datasets demonstrate the superior performance of our method compared to existing state-of-the-art approaches and exhibit excellent robustness.
Zewen Zheng, Yongqiang Mou, Guoheng Huang, Chi-Man Pun, Xiaochen Yuan
AAAI7
2024 Efficient Prototype Consistency Learning in Semi-Supervised Medical Image Segmentation via Joint Uncertainty and Data Augmentation
abstract
Recently, prototype learning has emerged in semi-supervised medical image segmentation and achieved remarkable performance. However, the scarcity of labeled data limits the expressiveness of prototypes in previous methods, potentially hindering the complete representation of prototypes for class embedding. To overcome this issue, we propose an efficient prototype consistency learning via joint uncertainty quantification and data augmentation (EPCL-JUDA) to enhance the semantic expression of prototypes based on the framework of Mean-Teacher. The concatenation of original and augmented labeled data is fed into student network to generate expressive prototypes. Then, a joint uncertainty quantification method is devised to optimize pseudo-labels and generate reliable prototypes for original and augmented unlabeled data separately. High-quality global prototypes for each class are formed by fusing labeled and unlabeled prototypes, which are utilized to generate prototype-to-features to conduct consistency learning. Notably, a prototype network is proposed to reduce high memory requirements brought by the introduction of augmented data. Extensive experiments on Left Atrium, Pancreas-NIH, Type B Aortic Dissection datasets demonstrate EPCL-JUDA’s superiority over previous state-of-the-art approaches, confirming the effectiveness of our framework. The code will be released soon.
Lijian Li 0003, Yuanpeng He, Chi-Man Pun
BIBM3
2024 Mutual Evidential Deep Learning for Semi-supervised Medical Image Segmentation
abstract
Existing semi-supervised medical segmentation co-learning frameworks have realized that model performance can be diminished by the biases in model recognition caused by low-quality pseudo-labels. Due to the averaging nature of their pseudo-label integration strategy, they fail to explore the reliability of pseudo-labels from different sources. In this paper, we propose a mutual evidential deep learning (MEDL) framework that offers a potentially viable solution for pseudo-label generation in semi-supervised learning from two perspectives. First, we introduce networks with different architectures to generate complementary evidence for unlabeled samples and adopt an improved class-aware evidential fusion to guide the confident synthesis of evidential predictions sourced from diverse architectural networks. Second, utilizing the uncertainty in the fused evidence, we design an asymptotic Fisher information-based evidential learning strategy. This strategy enables the model to initially focus on unlabeled samples with more reliable pseudo-labels, gradually shifting attention to samples with lower-quality pseudo-labels while avoiding over-penalization of mislabeled classes in high data uncertainty samples. Additionally, for labeled data, we continue to adopt an uncertainty-driven asymptotic learning strategy, gradually guiding the model to focus on challenging voxels. Extensive experiments on five mainstream datasets have demonstrated that MEDL achieves state-of-the-art performance.
Yuanpeng He, Yali Bi, Lijian Li 0003, Chi-Man Pun, Wenpin Jiao, Zhi Jin 0001
BIBM4
2024 IMAN: An Adaptive Network for Robust NPC Mortality Prediction with Missing Modalities
abstract
Accurate prediction of mortality in nasopharyngeal carcinoma (NPC), a complex malignancy particularly challenging in advanced stages, is crucial for optimizing treatment strategies and improving patient outcomes. However, this predictive process is often compromised by the high-dimensional and heterogeneous nature of NPC-related data, coupled with the pervasive issue of incomplete multi-modal data, manifesting as missing radiological images or incomplete diagnostic reports. Traditional machine learning approaches suffer significant performance degradation when faced with such incomplete data, as they fail to effectively handle the high-dimensionality and intricate correlations across modalities. Even advanced multi-modal learning techniques like Transformers struggle to maintain robust performance in the presence of missing modalities, as they lack specialized mechanisms to adaptively integrate and align the diverse data types, while also capturing nuanced patterns and contextual relationships within the complex NPC data. To address these problem, we introduce IMAN: an adaptive network for robust NPC mortality prediction with missing modalities. IMAN features three integrated modules: the Dynamic Cross-Modal Calibration (DCMC) module employs adaptive, learnable parameters to scale and align medical images and field data; the Spatial-Contextual Attention Integration (SCAI) module enhances traditional Transformers by incorporating positional information within the self-attention mechanism, improving multi-modal feature integration; and the Context-Aware Feature Acquisition (CAFA) module adjusts convolution kernel positions through learnable offsets, allowing for adaptive feature capture across various scales and orientations in medical image modalities. Extensive experiments on our proprietary NPC dataset demonstrate IMAN’s robustness and high predictive accuracy, even with missing data. Compared to existing methods, IMAN consistently outperforms in scenarios with incomplete data, representing a significant advancement in mortality prediction for medical diagnostics and treatment planning. Our code is available at https://github.com/king-huoye/BIBM-2024/tree/master.
Yejing Huo, Guoheng Huang, Lianglun Cheng, Jianbin He, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun
BIBM8
2024 An Automatic Assessment of Parkinson's Disease in Arising from Chair Task via Refined Diffusion-based Pose Estimator
abstract
Parkinson’s disease (PD) is a progressively common neurodegenerative disorder characterized by a decline in motor function. The diagnosis of PD typically relies on the Movement Disorder Society-Unified Parkinson’s Disease Rating Scale (MDS-UPDRS), which involves subjective scoring through observation of targeted movements. However, this objective method heavily depends on professional experience and has relatively high misdiagnosis rates. In this paper, we introduce a novel vision-based architecture for automated assessment of the ‘arising from chair’ task, which is one of the key MDS-UPDRS components. First, a diffusion-based 2D pose estimator is proposed to enhance keypoint accuracy by iteratively learning the distribution of ground-truth data and then denoising noisy poses. Second, a keypoint trajectory refinement network is introduced to eliminate the jitter error by considering motion information such as position, velocity, acceleration, and jerk. Finally, based on the predicted skeleton keypoint trajectories, we propose several objective indicators to assess the movement characteristics and perform the final rating using the classifier. The experiment substantiates the proposed algorithm, achieving a precision of 98.7% and an accuracy of 95.8% in classifying the ‘arising from chair’ task. Furthermore, the classification results and the proposed objective indicators have been validated as effective aids for neurologists to provide more precise diagnoses.
Chi-Man Pun, Haolun Li 0001, Mingliang Zhai, Feng Xu 0005, Hao Gao 0005
BIBM2
2024 SMAFormer: Synergistic Multi-Attention Transformer for Medical Image Segmentation
abstract
In medical image segmentation, specialized computer vision techniques, notably transformers grounded in attention mechanisms and residual networks employing skip connections, have been instrumental in advancing performance. Nonetheless, previous models often falter when segmenting small, irregularly shaped tumors. To this end, we introduce SMAFormer, an efficient, Transformer-based architecture that fuses multiple attention mechanisms for enhanced segmentation of small tumors and organs. SMAFormer can capture both local and global features for medical image segmentation. The architecture comprises two pivotal components. First, a Synergistic Multi-Attention (SMA) Transformer block is proposed, which has the benefits of Pixel Attention, Channel Attention, and Spatial Attention for feature enrichment. Second, addressing the challenge of information loss incurred during attention mechanism transitions and feature fusion, we design a Feature Fusion Modulator. This module bolsters the integration between the channel and spatial attention by mitigating reshaping-induced information attrition. To evaluate our method, we conduct extensive experiments on various medical image segmentation tasks, including multi-organ, liver tumor, and bladder tumor segmentation, achieving state-of-the-art results. Code and models are available at: https://github.com/lzeeorno/SMAFormer.
Fuchen Zheng, Xuhang Chen 0002, Weihuang Liu, Haolun Li 0001, Yingtie Lei, Chi-Man Pun, Shoujun Zhou
BIBM7
2024 FAQNet: Frequency-Aware Quaternion Network for Endoscopic Highlight Removal
abstract
Due to the built-in light source within the endoscope, the illumination of bodily mucous can cause the formation of highlight regions due to reflection. This not only interferes with the diagnosis conducted by doctors but also poses a challenge to subsequent computer vision tasks. To tackle this issue, we introduce FAQNet, a network specifically designed for endoscopic image highlight removal. FAQNet seamlessly integrates multi-channel information leveraging quaternion convolution and spatial channel attention within our Quaternion Multi-Channel Fusion (QMCF) Module. This allows it to capture intricate details of color, texture, spatial information, and highlight characteristics within the imaged organ. Additionally, by employing frequency domain transformation and dilated convolution, the Contextual Information Integration (CII) Module effectively enlarges the receptive field, organizing contextual information between highlight regions and their surrounding areas. Lastly, the PixelShuffle Upsampling (PSU) Module generates the restored image. We validate our model’s performance on two benchmark datasets, demonstrating its superiority over existing highlight removal methodologies.
Dingzhou Zhu, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun
BIBM6
2024 PDGC: Properly Disentangle by Gating and Contrasting for Cross-Domain Few-Shot Classification
Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Yan Li 0122, Chi-Man Pun, Junbing Quan
CGI (2)6
2024 Depth-Aware Test-Time Training for Zero-Shot Video Object Segmentation
abstract
Zero-shot Video Object Segmentation (ZSVOS) aims at segmenting the primary moving object without any human annotations. Mainstream solutions mainly focus on learning a single model on large-scale video datasets, which struggle to generalize to unseen videos. In this work, we introduce a test-time training (TTT) strategy to address the problem. Our key insight is to enforce the model to predict consistent depth during the TTT process. In detail, we first train a single network to perform both segmentation and depth prediction tasks. This can be effectively learned with our specifically designed depth modulation layer. Then, for the TTT process, the model is updated by predicting consistent depth maps for the same frame under different data augmentations. In addition, we explore different TTT weight updating strategies. Our empirical results suggest that the momentum-based weight initialization and looping-based training scheme lead to more stable improvements. Experiments show that the proposed method achieves clear improvements on ZSVOS. Our proposed video TTT strategy provides significant superiority over state-of-the-art TTT methods. Our code is available at: https://nifangbaage.github.io/DATTT/.
Weihuang Liu, Xi Shen 0001, Haolun Li 0001, Xiuli Bi, Bo Liu 0047, Chi-Man Pun, Xiaodong Cun
CVPR6
2024 UIE-UnFold: Deep Unfolding Network with Color Priors and Vision Transformer for Underwater Image Enhancement
abstract
Underwater image enhancement (UIE) plays a crucial role in various marine applications, but it remains challenging due to the complex underwater environment. Current learning-based approaches frequently lack explicit incorporation of prior knowledge about the physical processes involved in underwater image formation, resulting in limited optimization despite their impressive enhancement results. This paper proposes a novel deep unfolding network (DUN) for UIE that integrates color priors and inter-stage feature transformation to improve enhancement performance. The proposed DUN model combines the iterative optimization and reliability of model-based methods with the flexibility and representational power of deep learning, offering a more explainable and stable solution compared to existing learning-based UIE approaches. The proposed model consists of three key components: a Color Prior Guidance Block (CPGB) that establishes a mapping between color channels of degraded and original images, a Nonlinear Activation Gradient Descent Module (NAGDM) that simulates the underwater image degradation process, and an Inter Stage Feature Transformer (ISF-Former) that facilitates feature exchange between different network stages. By explicitly incorporating color priors and modeling the physical characteristics of underwater image for-mation, the proposed DUN model achieves more accurate and reliable enhancement results. Extensive experiments on multiple underwater image datasets demonstrate the superiority of the proposed model over state-of-the-art methods in both quantitative and qualitative evaluations. The proposed DUN-based approach offers a promising solution for UIE, enabling more accurate and reliable scientific analysis in marine research. The code is available at https://github.com/CXH-Research/UIE-UnFold.
Yingtie Lei, Yihang Dong, Changwei Gong, Ziyang Zhou 0001, Chi-Man Pun
DSAA6
2024 Generalized Uncertainty-Based Evidential Fusion with Hybrid Multi-Head Attention for Weak-Supervised Temporal Action Localization
abstract
Weakly supervised temporal action localization (WS-TAL) is a task of targeting at localizing complete action instances and categorizing them with video-level labels. Action-background ambiguity, primarily caused by background noise resulting from aggregation and intra-action variation, is a significant challenge for existing WS-TAL methods. In this paper, we introduce a hybrid multi-head attention (HMHA) module and generalized uncertainty-based evidential fusion (GUEF) module to address the problem. The proposed HMHA effectively enhances RGB and optical flow features by filtering redundant information and adjusting their feature distribution to better align with the WS-TAL task. Additionally, the proposed GUEF adaptively eliminates the interference of background noise by fusing snippet-level evidences to refine uncertainty measurement and select superior foreground feature information, which enables the model to concentrate on integral action instances to achieve better action localization and classification performance. Experimental results conducted on the THUMOS14 dataset demonstrate that our method outperforms state-of-the-art methods. Our code is available in https://github.com/heyuanpengpku/GUEF/tree/main.
Yuanpeng He, Lijian Li 0003, Tianxiang Zhan, Wenpin Jiao, Chi-Man Pun
ICASSP5
2024 DeformMLP: Dynamic Large-Scale Receptive Field MLP Networks for Human Motion Prediction
abstract
Predicting human motion requires addressing dependencies and errors for pose forecasting from sequences. The transformer’s self-attention aids this, but its complexity poses computational challenges. We present an efficient DeformMLP network without self-attention, using fully connected layers. DeformMLP includes DeformFCs, DeformFCt, and DeformFCst layers for spatial temporal modeling and calibration. DeformFCs capture semantics, DeformFCt learns relationships by summarizing time tokens, and DeformFCst assigns significance to dimensions to reduce computation. Our method balances efficiency and accuracy through decomposition and weight allocation. Evaluation on Human3.6M, 3DPW, CMU-MoCap datasets shows state-of-the-art prediction performance by benchmarks. The code is publicly available at https://github.com/HHT-98/DeformMLP.
Chi-Man Pun, Haolun Li 0001, Jian Xiong 0005, Hao Gao 0005
ICASSP2
2024 Local Optimization Networks for Multi-View Multi-Person Human Posture Estimation
abstract
With the growing applicability of multi-view multi-person 3D human pose estimation across diverse scenarios, the impact of external environmental factors and occlusion on accuracy has garnered substantial attention. In this research, we introduce a novel approach to multi-view multi-person 3D human pose estimation, leveraging a localized optimization strategy. Specifically, our method enhances the interplay of feature information from different channels and fine-tunes the optimal feature weights to capture intricate dependencies among joints. This refinement leads to improved accuracy in handling external environmental factors. Experimental evaluations were conducted on two prominent benchmark datasets, namely Campus and Shelf. The proposed method achieved a remarkable performance, with a Percentage of Correct Parts (PCP) score of 97.4% and 98.2% for the Campus and Shelf datasets, respectively.
Jucheng Song, Chi-Man Pun, Haolun Li 0001, Rushi Lan, Jiucheng Xie, Hao Gao 0005
ICASSP2
2024 Hierarchical Local Temporal Feature Enhancing for Transformer-Based 3D Human Pose Estimation
abstract
Recent advancements in transformer-based methods have yielded substantial success in 2D-to-3D human pose estimation. Transformer-based estimators have their inherent advantages like global receptive field. Nevertheless, existing transformer approaches ignore the differences among local contexts, resulting in insufficient learning of local information. To address this issue, we introduce non-uniform graph convolution to extract spatial local relationships in skeletons, remedying the limitations of traditional transformers in learning human body topology effectively. Additionally, our proposed Hierarchical Local Temporal Network (HLTN) models local temporal associations across three hierarchical levels: joints, body-parts and poses, effectively addressing the constraint of traditional transformers in learning localized human movements. We connect these two modules in parallel with the spatial and temporal transformer to obtain better features of skeleton sequences. Compared with the latest methods, our method achieves state-of-the-art performance on multiple datasets.
Chi-Man Pun, Haolun Li 0001, Hao Gao 0005
ICME2
2024 FOPS-V: Feature-Aware Optimization and Parallel Scale Fusion for 3D Human Reconstruction in Video
Guoheng Huang, Lianglun Cheng, Yejing Huo, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun
ICONIP (8)8
2024 ROSAL: Semi-supervised Active Learning with Representation Aggregation and Outlier for Endoscopy Image Classification
Xiaocong Huang, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Xuhang Chen 0002, Chi-Man Pun, Jianwu Chen
ICONIP (11)6
2024 Test-Time Intensity Consistency Adaptation for Shadow Detection
Leyi Zhu, Weihuang Liu, Zimeng Li 0001, Xuhang Chen 0002, Chi-Man Pun
ICONIP (7)7
2024 ShaDocFormer: A Shadow-Attentive Threshold Detector With Cascaded Fusion Refiner for Document Shadow Removal
abstract
Document shadow is a common issue that arises when capturing documents using mobile devices, which significantly impacts readability. Current methods encounter various challenges, including inaccurate detection of shadow masks and estimation of illumination. In this paper, we propose ShaDoc-Former, a Transformer-based architecture that integrates traditional methodologies and deep learning techniques to tackle the problem of document shadow removal. The ShaDocFormer architecture comprises two components: the Shadow-attentive Threshold Detector (STD) and the Cascaded Fusion Refiner (CFR). The STD module employs a traditional thresholding technique and leverages the attention mechanism of the Transformer to gather global information, thereby enabling precise detection of shadow masks. The cascaded and aggregative structure of the CFR module facilitates a coarse-to-fine restoration process for the entire image. As a result, ShaDocFormer excels in accurately detecting and capturing variations in both shadow and illumination, thereby enabling effective removal of shadows. Extensive experiments demonstrate that ShaDocFormer outperforms current state-of-the-art methods in both qualitative and quantitative measurements. The code is available at https://github.com/kilito777/ShaDocFormer.
Weiwen Chen, Yingtie Lei, Shenghong Luo, Ziyang Zhou 0001, Mingxian Li, Chi-Man Pun
IJCNN6
2024 UWFormer: Underwater Image Enhancement via a Semi-Supervised Multi-Scale Transformer
abstract
Underwater images often exhibit poor quality, distorted color balance and low contrast due to the complex and intricate interplay of light, water, and objects. Despite the significant contributions of previous underwater enhancement techniques, there exist several problems that demand further improvement: (i) The current deep learning methods rely on Convolutional Neural Networks (CNNs) that lack the multi-scale enhancement, and global perception field is also limited. (ii) The scarcity of paired real-world underwater datasets poses a significant challenge, and the utilization of synthetic image pairs could lead to overfitting. To address the aforementioned problems, this paper introduces a Multi-scale Transformer-based Network called UWFormer for enhancing images at multiple frequencies via semi-supervised learning, in which we propose a Nonlinear Frequency-aware Attention mechanism and a Multi-Scale Fusion Feed-forward Network for low-frequency enhancement. Besides, we introduce a special underwater semi-supervised training strategy, where we propose a Subaqueous Perceptual Loss function to generate reliable pseudo labels. Experiments using full-reference and non-reference underwater benchmarks demonstrate that our method outperforms state-of-the-art methods in terms of both quantity and visual quality. The code is available at https://github.com/leiyingtie/UWFormer.
Weiwen Chen, Yingtie Lei, Shenghong Luo, Ziyang Zhou 0001, Mingxian Li, Chi-Man Pun
IJCNN6
2024 Dual-Hybrid Attention Network for Specular Highlight Removal
Xiaojiao Guo, Xuhang Chen 0002, Shenghong Luo, Shuqiang Wang, Chi-Man Pun
ACM Multimedia5
2024 DP-RAE: A Dual-Phase Merging Reversible Adversarial Example for Image Privacy Protection
abstract
In digital security, Reversible Adversarial Examples (RAE) blend adversarial attacks with Reversible Data Hiding (RDH) within images to thwart unauthorized access. Traditional RAE methods, however, compromise attack efficiency for the sake of perturbation concealment, diminishing the protective capacity of valuable perturbations and limiting applications to white-box scenarios. This paper proposes a novel Dual-Phase merging Reversible Adversarial Example (DP-RAE) generation framework, combining a heuristic black-box attack and RDH with Grayscale Invariance (RDH-GI) technology. This dual strategy not only evaluates and harnesses the adversarial potential of past perturbations more effectively but also guarantees flawless embedding of perturbation information and complete recovery of the original image. Experimental validation reveals our method's superiority, secured an impressive 96.9% success rate and 100% recovery rate in compromising black-box models. In particular, it achieved a 90% misdirection rate against commercial models under a constrained number of queries. This marks the first successful attempt at targeted black-box reversible adversarial attacks for commercial recognition models. This achievement highlights our framework's capability to enhance security measures without sacrificing attack performance. Moreover, our attack framework is flexible, allowing the interchangeable use of different attack and RDH modules to meet advanced technological requirements.
Xia Du, Jizhe Zhou 0001, Chi-Man Pun, Qizhen Xu
ACM Multimedia4
2024 IMDL-BenCo: A Comprehensive Benchmark and Codebase for Image Manipulation Detection & Localization
abstract
A comprehensive benchmark is yet to be established in the Image Manipulation Detection & Localization (IMDL) field. The absence of such a benchmark leads to insufficient and misleading model evaluations, severely undermining the development of this field. However, the scarcity of open-sourced baseline models and inconsistent training and evaluation protocols make conducting rigorous experiments and faithful comparisons among IMDL models challenging. To address these challenges, we introduce IMDL-BenCo, the first comprehensive IMDL benchmark and modular codebase. IMDL-BenCo: i) decomposes the IMDL framework into standardized, reusable components and revises the model construction pipeline, improving coding efficiency and customization flexibility; ii) fully implements or incorporates training code for state-of-the-art models to establish a comprehensive IMDL benchmark; and iii) conducts deep analysis based on the established benchmark and codebase, offering new insights into IMDL model architecture, dataset characteristics, and evaluation standards.Specifically, IMDL-BenCo includes common processing algorithms, 8 state-of-the-art IMDL models (1 of which are reproduced from scratch), 2 sets of standard training and evaluation protocols, 15 GPU-accelerated evaluation metrics, and 3 kinds of robustness evaluation. This benchmark and codebase represent a significant leap forward in calibrating the current progress in the IMDL field and inspiring future breakthroughs.Code is available at: https://github.com/scu-zjz/IMDLBenCo
Xiaochen Ma 0001, Xuekang Zhu, Zhuohang Jiang, Bingkui Tong, Zeyu Lei, Chi-Man Pun, Jiancheng Lv 0001, Jizhe Zhou 0001
NeurIPS9
2024 MedPrompt: Cross-modal Prompting for Multi-task Medical Image Translation
Xuhang Chen 0002, Shenghong Luo, Chi-Man Pun, Shuqiang Wang
PRCV (14)3
2024 DocDeshadower: Frequency-Aware Transformer for Document Shadow Removal
abstract
Shadows in scanned documents pose significant challenges for document analysis and recognition tasks due to their negative impact on visual quality and readability. Current shadow removal techniques, including traditional methods and deep learning approaches, face limitations in handling varying shadow intensities and preserving document details. To address these issues, we propose DocDeshadower, a novel multi-frequency Transformer-based model built upon the Laplacian Pyramid. By decomposing the shadow image into multiple frequency bands and employing two critical modules: the Attention-Aggregation Network for low-frequency shadow removal and the Gated Multi-scale Fusion Transformer for global refinement. DocDeshadower effectively removes shadows at different scales while preserving document content. Extensive experiments demonstrate DocDe-shadower's superior performance compared to state-of-the-art methods, highlighting its potential to significantly improve document shadow removal techniques. The code is available at https://github.com/leiyingtie/DocDeshadower.
Ziyang Zhou 0001, Yingtie Lei, Xuhang Chen 0002, Shenghong Luo, Chi-Man Pun
SMC6
2024 Cross-Modality Disentangled Information Bottleneck Strategy for Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis (MSA) has been a pivotal domain in current research area which utilizes diverse information carriers such as videos containing multiple modal-ities to understand the user's sentiment. With the success of multimodal fusion techniques, lots of fusion strategies have been proposed to obtain a favorable multimodal joint representation for MSA. However, existing studies hardly consider the problem of redundant information in unimodal, resulting in the joint representation may contain much redundant information from different modalities, thus limiting the accuracy of sentiment prediction. In this work, we propose a Cross-Modality Disentangled Information Bottleneck Strategy (CMDIBS), which consists of a Cross-Modality Knowledge Awareness (CMKA) module and a Multimodal Disentangled Information Bottleneck (MDIB) mechanism. Specifically, the CMKA module encourages in-teractions among different modalities to learn the sentiment embedding relevant to the predicted goals. In particular, MDIB mechanism aims to maximize the mutual information (MI) between the multimodal joint representation and the predicted label, and maximize the MI between the style embedding with the label and the input data while constraining the MI between the multimodal joint representation and the style embedding to obtain a succinct and efficient multimodal joint representation. Experimental results on the benchmark datasets, namely CMU-MOSI and CMU-MOSEI, indicated that the proposed method surpasses existing approaches and attains SOTA performance.
Zhengnan Deng, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Lian Huang, Chi-Man Pun
SMC6
2024 Sketch Video Synthesis
abstract
Abstract Understanding semantic intricacies and high‐level concepts is essential in image sketch generation, and this challenge becomes even more formidable when applied to the domain of videos. To address this, we propose a novel optimization‐based framework for sketching videos represented by the frame‐wise Bézier Curves. In detail, we first propose a cross‐frame stroke initialization approach to warm up the location and the width of each curve. Then, we optimize the locations of these curves by utilizing a semantic loss based on CLIP features and a newly designed consistency loss using the self‐decomposed 2D atlas network. Built upon these design elements, the resulting sketch video showcases notable visual abstraction and temporal coherence. Furthermore, by transforming a video into vector lines through the sketching process, our method unlocks applications in sketch‐based video editing and video doodling, enabled through video composition.
Yudian Zheng, Xiaodong Cun, Menghan Xia, Chi-Man Pun
Comput. Graph. Forum4
2024 Multi-task subspace clustering
Guo Zhong, Chi-Man Pun
Inf. Sci.2
2024 Black-box reversible adversarial examples with invertible neural network
Jielun Huang, Guoheng Huang, Xiaochen Yuan, Fenfang Xie, Chi-Man Pun, Guo Zhong
Image Vis. Comput.6
2024 WavEnhancer: Unifying Wavelet and Transformer for Image Enhancement
Zinuo Li, Xuhang Chen 0002, Shu-Na Guo, Shuqiang Wang, Chi-Man Pun
J. Comput. Sci. Technol.5
2024 Efficient physical image attacks using adversarial fast autoaugmentation methods
Xia Du, Chi-Man Pun, Jizhe Zhou 0001
Knowl. Based Syst.2
2024 Progressive normalizing flow with learnable spectrum transform for style transfer
Guoheng Huang, Xiaochen Yuan, Guo Zhong, Chi-Man Pun, Yiwen Zeng
Knowl. Based Syst.5
2024 DH-GAN: Image manipulation localization via a dual homology-aware generative adversarial network
Weihuang Liu, Xiaodong Cun, Chi-Man Pun
Pattern Recognit.3
2024 Attack-invariant attention feature for adversarial defense in hyperspectral image classification
Cheng Shi 0002, Minghua Zhao, Chi-Man Pun, Qiguang Miao
Pattern Recognit.4
2024 Adaptive Spatial-Temporal Graph-Mixer for Human Motion Prediction
abstract
The Graph Convolutional Network (GCN) has recently achieved promising performance in human motion prediction by modeling the nodes and edges of the human skeleton. However, most previous methods still suffer from two unaddressed drawbacks. First, in the inference stage, their graph topologies are static and fixed, resulting in dependencies between nodes that cannot be dynamically adjusted for different actions. Second, the implicit relationships between pose sequences are ignored, which makes the prior advantages of the graph structure invalid in temporal feature fusion. To address these limitations, we propose an adaptive spatial-temporal graph-mixer (GraphMixer) for human motion prediction, which consists of a series of fully separated spatial-temporal graph convolution structures. In spatial GCN, we construct an additional adaptive skeleton graph to capture the node features of action-specific poses. In temporal GCN, we introduce a variety of graph topologies to enhance feature fusion between pose sequences. Comparing state-of-the-art algorithms on the Human3.6 M and the 3DPW datasets and ablation studies shows that our GraphMixer and the proposed multiple graph topologies are effective and critical. The code is publicly available athttps://github.com/young0304/Adaptive-Spatial-Temporal-Graph-Mixer.
Haolun Li 0001, Chi-Man Pun, Chun Du, Hao Gao 0005
IEEE Signal Process. Lett.3
2024 A Parkinson's Auxiliary Diagnosis Algorithm Based on a Hyperparameter Optimization Method of Deep Learning
abstract
Parkinson's disease is a common mental disease in the world, especially in the middle-aged and elderly groups. Today, clinical diagnosis is the main diagnostic method of Parkinson's disease, but the diagnosis results are not ideal, especially in the early stage of the disease. In this paper, a Parkinson's auxiliary diagnosis algorithm based on a hyperparameter optimization method of deep learning is proposed for the Parkinson's diagnosis. The diagnosis system uses ResNet50 to achieve feature extraction and Parkinson's classification, mainly including speech signal processing part, algorithm improvement part based on Artificial Bee Colony algorithm (ABC) and optimizing the hyperparameters of ResNet50 part. The improved algorithm is called Gbest Dimension Artificial Bee Colony algorithm (GDABC), proposing "Range pruning strategy" which aims at narrowing the scope of search and "Dimension adjustment strategy" which is to adjust gbest dimension by dimension. The accuracy of the diagnosis system in the verification set of Mobile Device Voice Recordings at King's College London (MDVR-CKL) dataset can reach more than 96%. Compared with current Parkinson's sound diagnosis methods and other optimization algorithms, our auxiliary diagnosis system shows better classification performance on the dataset within limited time and resources.
Shujuan Li, Chi-Man Pun, Yijing Guo, Feng Xu 0005, Hao Gao 0005, Huimin Lu 0001
IEEE Trans. Comput. Biol. Bioinform.3
2024 GDN-CMCF: A Gated Disentangled Network With Cross-Modality Consensus Fusion for Multimodal Named Entity Recognition
abstract
Multimodal named entity recognition (MNER) is a crucial task in social systems of artificial intelligence that requires precise identification of named entities in sentences using both visual and textual information. Previous methods have focused on capturing fine-grained visual features and developing complex fusion procedures. However, these approaches overlook the heterogeneity gap and loss of original modality uniqueness that may occur during fusion, leading to incorrect entity identification. This article proposes a novel approach for MNER called a gated disentangled network with cross-modality consensus fusion (GDN-CMCF) to address the above challenges. Specifically, to eliminate cross-modality variation, we propose a cross-modality consensus fusion module that generates a consensus representation by learning inter-and intramodality interactions with a designed commonality constraint. We then introduce a gated disentanglement module to separate modality-relevant features from support and auxiliary modalities, which further filters out extraneous information while retaining the uniqueness of unimodal features. Experimental results on two real public datasets are provided to verify the effectiveness of our proposed GDN-CMCF. The source code of this article can be found at https://github.com/HaoDavis/ GDN-CMCF.
Guoheng Huang, Zihao Dai, Guo Zhong, Xiaochen Yuan, Chi-Man Pun
IEEE Trans. Comput. Soc. Syst.6
2024 Enabling Parity Authenticator-Based Public Auditing With Protection of a Valid User Revocation in Cloud
abstract
The significance of the cloud enables the data owners (DOs) to store data remotely in cloud server (CS). The external and internal attacks on the stored data at CS can deliberately remove data. Furthermore, the CS removes the stored data to make empty location for the user's upcoming new data. However, it is a legal expectation of DOs to know whether their data are correctly stored or altered in CS. In this article, we propose a novel privacy-aware and hash-parity-bits-based public auditing (PA-HPPA) framework to secure full data, left half of the data, and the right half of the data, generated by a DO. DO generates two private key pairs with the assistance of a virtual key and a user ID (IP). The virtual key is the sequence number of DO who is registered and provided by the trusted data manager (TDM) while IP is the sequence number of DO working in an organization. Subsequently, DO blinds the categorized data and generates their signatures and hashes. In addition, DO generates the parity bits using xor and assigns to each hard drive (HD) in CS, which assistants to TDM in public auditing. Second, how to identify the error in the stored data and how to securely recover the error/missed data? Extension to the framework, the novel proposed data error identification and secure data recovery produce tags for installed HDs of CS using truth table and recover the altered/missed data via a authenticator, which is produced using xor function. Third, how to protect a valid user from revocation and, in case a user has revoked on merit basis, then how to securely access the stored data of it? This novel work has proposed three conditions to meet the validity of the valid user from revocation and securely generating the public-private key pairs to access the stored data of the revoked user securely from CS. Fourth, there is an efficient novel proposed dynamic operation scheme to insert, update, or delete the stored data at CS without regenerating the signatures, hashes, and tags for the whole stored data in cloud. The security analysis and the performance evaluation of the proposed solutions are provably efficient and secure with reduced communication costs.
Fasee Ullah, Chi-Man Pun
IEEE Trans. Comput. Soc. Syst.2
2024 CTNet: Contrastive Transformer Network for Polyp Segmentation
abstract
Segmenting polyps from colonoscopy images is very important in clinical practice since it provides valuable information for colorectal cancer. However, polyp segmentation remains a challenging task as polyps have camouflage properties and vary greatly in size. Although many polyp segmentation methods have been recently proposed and produced remarkable results, most of them cannot yield stable results due to the lack of features with distinguishing properties and those with high-level semantic details. Therefore, we proposed a novel polyp segmentation framework called contrastive Transformer network (CTNet), with three key components of contrastive Transformer backbone, self-multiscale interaction module (SMIM), and collection information module (CIM), which has excellent learning and generalization abilities. The long-range dependence and highly structured feature map space obtained by CTNet through contrastive Transformer can effectively localize polyps with camouflage properties. CTNet benefits from the multiscale information and high-resolution feature maps with high-level semantic obtained by SMIM and CIM, respectively, and thus can obtain accurate segmentation results for polyps of different sizes. Without bells and whistles, CTNet yields significant gains of 2.3%, 3.7%, 3.7%, 18.2%, and 10.1% over classical method PraNet on Kvasir-SEG, CVC-ClinicDB, Endoscene, ETIS-LaribPolypDB, and CVC-ColonDB respectively. In addition, CTNet has advantages in camouflaged object detection and defect detection. The code is available at https://github.com/Fhujinwu/CTNet.
Bin Xiao 0002, Jinwu Hu, Weisheng Li 0001, Chi-Man Pun, Xiuli Bi
IEEE Trans. Cybern.4
2024 Quaternion Cross-Modality Spatial Learning for Multi-Modal Medical Image Segmentation
abstract
Recently, the Deep Neural Networks (DNNs) have had a large impact on imaging process including medical image segmentation, and the real-valued convolution of DNN has been extensively utilized in multi-modal medical image segmentation to accurately segment lesions via learning data information. However, the weighted summation operation in such convolution limits the ability to maintain spatial dependence that is crucial for identifying different lesion distributions. In this paper, we propose a novel Quaternion Cross-modality Spatial Learning (Q-CSL) which explores the spatial information while considering the linkage between multi-modal images. Specifically, we introduce to quaternion to represent data and coordinates that contain spatial information. Additionally, we propose Quaternion Spatial-association Convolution to learn the spatial information. Subsequently, the proposed De-level Quaternion Cross-modality Fusion (De-QCF) module excavates inner space features and fuses cross-modality spatial dependency. Our experimental results demonstrate that our approach compared to the competitive methods perform well with only 0.01061 M parameters and 9.95G FLOPs.
Junyang Chen 0001, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Zewen Zheng, Chi-Man Pun, Jian Zhu 0001
IEEE J. Biomed. Health Informatics6
2024 Learning From Incorrectness: Active Learning With Negative Pre-Training and Curriculum Querying for Histological Tissue Classification
abstract
Patch-level histological tissue classification is an effective pre-processing method for histological slide analysis. However, the classification of tissue with deep learning requires expensive annotation costs. To alleviate the limitations of annotation budgets, the application of active learning (AL) to histological tissue classification is a promising solution. Nevertheless, there is a large imbalance in performance between categories during application, and the tissue corresponding to the categories with relatively insufficient performance are equally important for cancer diagnosis. In this paper, we propose an active learning framework called ICAL, which contains Incorrectness Negative Pre-training (INP) and Category-wise Curriculum Querying (CCQ) to address the above problem from the perspective of category-to-category and from the perspective of categories themselves, respectively. In particular, INP incorporates the unique mechanism of active learning to treat the incorrect prediction results that obtained from CCQ as complementary labels for negative pre-training, in order to better distinguish similar categories during the training process. CCQ adjusts the query weights based on the learning status on each category by the model trained by INP, and utilizes uncertainty to evaluate and compensate for query bias caused by inadequate category performance. Experimental results on two histological tissue classification datasets demonstrate that ICAL achieves performance approaching that of fully supervised learning with less than 16% of the labeled data. In comparison to the state-of-the-art active learning algorithms, ICAL achieved better and more balanced performance in all categories and maintained robustness with extremely low annotation budgets. The source code will be released at https://github.com/LactorHwt/ICAL.
Lianglun Cheng, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Chi-Man Pun, Muyan Cai
IEEE Trans. Medical Imaging6
2024 MFDNet: Multi-Frequency Deflare Network for efficient nighttime flare removal
Yiguo Jiang, Xuhang Chen 0002, Chi-Man Pun, Shuqiang Wang, Wei Feng 0005
Vis. Comput.3
2024 SCDet: decoupling discriminative representation for dark object detection via supervised contrastive learning
Tongxu Lin, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Xiaocong Huang, Chi-Man Pun
Vis. Comput.6
2024 Frequency-constrained transferable adversarial attack on image manipulation detection and localization
Yijia Zeng, Chi-Man Pun
Vis. Comput.2
2023 CEE-Net: Complementary End-to-End Network for 3D Human Pose Generation and Estimation
abstract
The limited number of actors and actions in existing datasets make 3D pose estimators tend to overfit, which can be seen from the performance degradation of the algorithm on cross-datasets, especially for rare and complex poses. Although previous data augmentation works have increased the diversity of the training set, the changes in camera viewpoint and position play a dominant role in improving the accuracy of the estimator, while the generated 3D poses are limited and still heavily rely on the source dataset. In addition, these works do not consider the adaptability of the pose estimator to generated data, and complex poses will cause training collapse. In this paper, we propose the CEE-Net, a Complementary End-to-End Network for 3D human pose generation and estimation. The generator extremely expands the distribution of each joint-angle in the existing dataset and limits them to a reasonable range. By learning the correlations within and between the torso and limbs, the estimator can combine different body-parts more effectively and weaken the influence of specific joint-angle changes on the global pose, improving the generalization ability. Extensive ablation studies show that our pose generator greatly strengthens the joint-angle distribution, and our pose estimator can utilize these poses positively. Compared with the state-of-the-art methods, our method can achieve much better performance on various cross-datasets, rare and complex poses.
Haolun Li 0001, Chi-Man Pun
AAAI2
2023 CoordFill: Efficient High-Resolution Image Inpainting via Parameterized Coordinate Querying
abstract
Image inpainting aims to fill the missing hole of the input. It is hard to solve this task efficiently when facing high-resolution images due to two reasons: (1) Large reception field needs to be handled for high-resolution image inpainting. (2) The general encoder and decoder network synthesizes many background pixels synchronously due to the form of the image matrix. In this paper, we try to break the above limitations for the first time thanks to the recent development of continuous implicit representation. In detail, we down-sample and encode the degraded image to produce the spatial-adaptive parameters for each spatial patch via an attentional Fast Fourier Convolution (FFC)-based parameter generation network. Then, we take these parameters as the weights and biases of a series of multi-layer perceptron (MLP), where the input is the encoded continuous coordinates and the output is the synthesized color value. Thanks to the proposed structure, we only encode the high-resolution image in a relatively low resolution for larger reception field capturing. Then, the continuous position encoding will be helpful to synthesize the photo-realistic high-frequency textures by re-sampling the coordinate in a higher resolution. Also, our framework enables us to query the coordinates of missing pixels only in parallel, yielding a more efficient solution than the previous methods. Experiments show that the proposed method achieves real-time performance on the 2048X2048 images using a single GTX 2080 Ti GPU and can handle 4096X4096 images, with much better performance than existing state-of-the-art methods visually and numerically. The code is available at: https://github.com/NiFangBaAGe/CoordFill.
Weihuang Liu, Xiaodong Cun, Chi-Man Pun, Menghan Xia, Yong Zhang 0034, Jue Wang 0001
AAAI3
2023 Supervised Discriminative Discrete Hashing for Cross-Modal Retrieval
Chi-Man Pun
ADMA (2)2
2023 ELF: An End-to-end Local and Global Multimodal Fusion Framework for Glaucoma Grading
abstract
Glaucoma is a chronic neurodegenerative condition that can lead to blindness. Early detection and curing are very important in stopping the disease from getting worse for glaucoma patients. The 2D fundus images and optical coherence tomography(OCT) are useful for ophthalmologists in diagnosing glaucoma. There are many methods based on the fundus images or 3D OCT volumes; however, the mining for multi-modality, including both fundus images and data, is less studied. In this work, we propose an end-to-end local and global multi-modal fusion framework for glaucoma grading, named ELF for short. ELF can fully utilize the complementary information between fundus and OCT. In addition, unlike previous methods that concatenate the multi-modal features together, which lack exploring the mutual information between different modalities, ELF can take advantage of local-wise and global-wise mutual information. The extensive experiment conducted on the multi-modal glaucoma grading GAMMA dataset can prove the effiectness of ELF when compared with other state-of-the-art methods.
Wenyun Li 0001, Chi-Man Pun
BIBM2
2023 An Optimized-Skeleton-Based Parkinsonian Gait Auxiliary Diagnosis Method with Both Monitoring Indicators and Assisted Ratings
abstract
Abnormal gait is one of the indispensable diagnostic sources of Parkinson’s disease (PD) diagnosis, typically presenting as small shuffling steps and gait bradykinesia. However, its diagnostic accuracy is lower due to the subjective judgments of doctors. To assist in improving the accuracy and reducing the subjectivity of the doctors, we propose an optimized-skeleton-based Parkinsonian gait auxiliary diagnosis method with both monitoring indicators and assisted ratings. By inputting a patient gait video captured from the side, our PD symptom-applicable pose trajectory model will extract a more precise and stable 2D skeleton sequence of patients. Next, the sequence will be used to calculate our proposed five monitoring indicators: gait frequency, ankle speed, whole speed, ankle angle speed, ankle acceleration, and previous work indicators: arm swing angle, leg angle, two feet x-axis distance to record the patient’s gait details at every moment. The extracted gait frequency will then be input into a random forest model to obtain the gait rating. Lastly, doctors can make more accurate judgments by referring to our objective monitoring indicators and assisted ratings. Experimental results show that our monitored indicators improve the doctors’ diagnosis accuracy by 16%, the skeleton speed and acceleration error of our optimized-skeleton extraction method achieve 4.18 cm/s and 5.71 cm/s2, and our random forest model has reached a classification accuracy of 95.8%.
Gaoqi Li, Chi-Man Pun, Haolun Li 0001, Jian Xiong 0005, Feng Xu 0005, Hao Gao 0005
BIBM2
2023 Explicit Visual Prompting for Low-Level Structure Segmentations
abstract
We consider the generic problem of detecting low-level structures in images, which includes segmenting the manipulated parts, identifying out-of-focus pixels, separating shadow regions, and detecting concealed objects. Whereas each such topic has been typically addressed with a domain-specific solution, we show that a unified approach performs well across all of them. We take inspiration from the widely-used pre-training and then prompt tuning protocols in NLP and propose a new visual prompting model, named Explicit Visual Prompting (EVP). Different from the previous visual prompting which is typically a dataset-level implicit embedding, our key insight is to enforce the tunable parameters focusing on the explicit visual content from each individual image, i.e., the features from frozen patch embeddings and the input's high-frequency components. The proposed EVP significantly outperforms other parameter-efficient tuning protocols under the same amount of tunable parameters (5.7% extra trainable parameters of each task). EVP also achieves state-of-the-art performances on diverse low-level structure segmentation tasks compared to task-specific solutions. Our code is available at: https://github.com/NiFangBaAGe/Explicit-Visual-Prompt.
Weihuang Liu, Xi Shen 0001, Chi-Man Pun, Xiaodong Cun
CVPR3
2023 Shadocnet: Learning Spatial-Aware Tokens in Transformer for Document Shadow Removal
abstract
Shadow removal improves the visual quality and legibility of digital copies of documents. However, document shadow removal remains an unresolved subject. Traditional techniques rely on heuristics that vary from situation to situation. Given the quality and quantity of current public datasets, the majority of neural network models are ill-equipped for this task. In this paper, we propose a Transformer-based model for document shadow removal that utilizes shadow context encoding and decoding in both shadow and shadow-free regions. Additionally, shadow detection and pixel-level enhancement are included in the whole coarse-to-fine process. On the basis of comprehensive benchmark evaluations, it is competitive with state-of-the-art methods.
Xuhang Chen 0002, Xiaodong Cun, Chi-Man Pun, Shuqiang Wang
ICASSP3
2023 Locality Preserving Multiview Graph Hashing For Large Scale Remote Sensing Image Search
abstract
Hashing is very popular for remote sensing image search. This article proposes a multiview hashing with learnable parameters to retrieve the queried images for a large-scale remote sensing dataset. Existing methods always neglect that real-world remote sensing data lies on a low- dimensional manifold embedded in high-dimensional ambient space. Unlike previous methods, this article proposes to learn the consensus compact codes in a view-specific low-dimensional subspace. Furthermore, we have added a hyperparameter learnable module to avoid complex parameter tuning. In order to prove the effectiveness of our method, we carried out experiments on three widely used remote sensing data sets and compared them with seven state-of-the-art methods. Extensive experiments show that the proposed method can achieve competitive results compared to the other method.
Wenyun Li 0001, Guo Zhong, Chi-Man Pun
ICASSP4
2023 Boosting Face Recognition Performance with Synthetic Data and Limited Real Data
abstract
Face recognition is one of the most precise and straightforward methods to establish individual identity, and is important in our daily life. To solve the issues of privacy, bias, and collection difficulty caused by face recognition relying heavily on collecting a huge number of real face images from the Internet, a seemingly promising idea is to employ GAN-generated synthetic faces as the training data. However, there are obvious surface gaps and domain gaps between real and synthetic face images, and cannot be replaced directly. In this paper, we attempt to boost face recognition simultaneously using synthetic data and limited real data. Specifically, we first design an augmented space for auto augmentation methods to augment synthetic images to alleviate the surface gap, then propose to disentangle the underlying style distributions through dual batch normalization layers so that both synthetic and real images can be learned jointly by convolution layers without mixing across domains. Extensive experiments demonstrate our method can achieve better results than training with large quantities of real data.
Wenqing Wang 0002, Lingqing Zhang, Chi-Man Pun, Jiucheng Xie
ICASSP3
2023 High-Resolution Document Shadow Removal via A Large-Scale Real-World Dataset and A Frequency-Aware Shadow Erasing Net
abstract
Shadows often occur when we capture the document with casual equipment, which influences the visual quality and readability of the digital copies. Different from the algorithms for natural shadow removal, the algorithms in document shadow removal need to preserve the details of fonts and figures in high-resolution input. Previous works ignore this problem and remove the shadows via approximate attention and small datasets, which might not work in real-world situations. We handle high-resolution document shadow removal directly via a larger-scale real-world dataset and a carefully-designed frequency-aware network. As for the dataset, we acquire over 7k couples of high-resolution (2462 × 3699) images of real-world documents pairs with various samples under different lighting circumstances, which is 10 times larger than existing datasets. As for the design of the network, we decouple the high-resolution images in the frequency domain, where the low-frequency details and high-frequency boundaries can be effectively learned via the carefully designed network structure. Powered by our network and dataset, the proposed method shows a clearly better performance than previous methods in terms of visual quality and numerical results. The code, models, and dataset are available at https://github.com/CXH-Research/DocShadow-SD7K.
Zinuo Li, Xuhang Chen 0002, Chi-Man Pun, Xiaodong Cun
ICCV3
2023 Data Representation by Joint Hypergraph Embedding and Sparse Coding (Extended Abstract)
abstract
Matrix factorization (MF), a popular unsupervised learning technique for data representation, has been widely applied in data mining and machine learning. According to different application scenarios, one can impose different constraints on the factorization to find the desired basis, which captures high-level semantics for the given data, and learns the compact representation corresponding to the basis. We note that almost all previous work on MF in data mining has ignored to find such a basis, which can carry high-order semantics in the data. In this work, we propose a novel MF framework called Joint Hypergraph Embedding and Sparse Coding, in which the obtained basis captures high-order semantic information in data. Experimental results on data clustering demonstrate that the proposed method consistently outperforms the other state-of-the-art matrix factorization methods.
Guo Zhong, Chi-Man Pun
ICDE2
2023 Asymmetric Scalable Cross-Modal Hashing
abstract
Cross-modal hashing is a practical approach to solving the problem of large-scale multimedia retrieval. However, there are still specific issues that the current methods cannot solve, such as how to construct binary codes rather than relax them to continuity effectively and how to prevent n × n problem. This paper proposes a novel Asymmetric Scalable Cross-Modal Hashing (ASCMH) to address these issues. It learns a common latent space from the kernelized features of different modalities. It then transforms the similarity matrix optimization to a distance-distance difference minimization problem with the help of semantic labels and common latent space. Additionally, we use an orthogonal constraint of label information to construct hash codes necessary for search accuracy. Extensive experiments on three benchmark datasets show that our ASCMH outperforms the SOTA cross-modal hashing methods.
Wenyun Li 0001, Chi-Man Pun
ICIP2
2023 A Large-Scale Film Style Dataset for Learning Multi-frequency Driven Film Enhancement
abstract
Film, a classic image style, is culturally significant to the whole photographic industry since it marks the birth of photography. However, film photography is time-consuming and expensive, necessitating a more efficient method for collecting film-style photographs. Numerous datasets that have emerged in the field of image enhancement so far are not film-specific. In order to facilitate film-based image stylization research, we construct FilmSet, a large-scale and high-quality film style dataset. Our dataset includes three different film types and more than 5000 in-the-wild high resolution images. Inspired by the features of FilmSet images, we propose a novel framework called FilmNet based on Laplacian Pyramid for stylizing images across frequency bands and achieving film style outcomes. Experiments reveal that the performance of our model is superior than state-of-the-art techniques. The link of our dataset and code is https://github.com/CXH-Research/FilmNet.
Zinuo Li, Xuhang Chen 0002, Shuqiang Wang, Chi-Man Pun
IJCAI4
2023 AdaBrowse: Adaptive Video Browser for Efficient Continuous Sign Language Recognition
abstract
Raw videos have been proven to own considerable feature redundancy where in many cases only a portion of frames can already meet the requirements for accurate recognition. In this paper, we are interested in whether such redundancy can be effectively leveraged to facilitate efficient inference in continuous sign language recognition (CSLR). We propose a novel adaptive model (AdaBrowse) to dynamically select a most informative subsequence from input video sequences by modelling this problem as a sequential decision task. In specific, we first utilize a lightweight network to quickly scan input videos to extract coarse features. Then these features are fed into a policy network to intelligently select a subsequence to process. The corresponding subsequence is finally inferred by a normal CSLR model for sentence prediction. As only a portion of frames are processed in this procedure, the total computations can be considerably saved. Besides temporal redundancy, we are also interested in whether the inherent spatial redundancy can be seamlessly integrated together to achieve further efficiency, i.e., dynamically selecting a lowest input resolution for each sample, whose model is referred to as AdaBrowse+. Extensive experimental results on four large-scale CSLR datasets, i.e., PHOENIX14, PHOENIX14-T, CSL-Daily and CSL, demonstrate the effectiveness of AdaBrowse and AdaBrowse+ by achieving comparable accuracy with state-of-the-art methods with 1.44X throughput and 2.12X fewer FLOPs. Comparisons with other commonly-used 2D CNNs and adaptive efficient methods verify the effectiveness of AdaBrowse. Code is available at https://github.com/hulianyuyy/AdaBrowse.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Chi-Man Pun, Wei Feng 0005
ACM Multimedia4
2023 Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action Recognition
abstract
Vision Transformer, which performs well in various vision tasks, encounters a bottleneck in skeleton-based action recognition and falls short of advanced GCN-based methods. The root cause is that the current skeleton transformer depends on the self-attention mechanism of the complete channel of the global joint, ignoring the highly discriminative differential correlation within the channel, so it is challenging to learn the expression of the multivariate topology dynamically. To tackle this, we present Skeleton MixFormer, an innovative spatio-temporal architecture to effectively represent the physical correlations and temporal interactivity of the compact skeleton data. Two essential components make up the proposed framework: 1) Spatial MixFormer. The channel-grouping and mix-attention are utilized to calculate the dynamic multivariate topological relationships. Compared with the full-channel self-attention method, Spatial MixFormer better highlights the channel groups' discriminative differences and the joint adjacency's interpretable learning. 2) Temporal MixFormer, which consists of Multiscale Convolution, Temporal Transformer and Sequential Holding Module. The multivariate temporal models ensure the richness of global difference expression and realize the discrimination of crucial intervals in the sequence, thereby enabling more effective learning of long and short-term dependencies in actions. Our Skeleton MixFormer demonstrates state-of-the-art (SOTA) performance across seven different settings on four standard datasets, namely NTU-60, NTU-120, NW-UCLA, and UAV-Human. Related code will be available on https://github.com/ElricXin/Skeleton-MixFormer.
Wentian Xin, Qiguang Miao, Ruyi Liu 0001, Chi-Man Pun, Cheng Shi 0002
ACM Multimedia5
2023 Single Cross-domain Semantic Guidance Network for Multimodal Unsupervised Image Translation
Jiaying Lan, Lianglun Cheng, Guoheng Huang, Chi-Man Pun, Xiaochen Yuan, Shangyu Lai, Bingo Wing-Kuen Ling
MMM (1)4
2023 Brain Diffuser: An End-to-End Brain Image to Brain Network Pipeline
Xuhang Chen 0002, Bai Ying Lei, Chi-Man Pun, Shuqiang Wang
PRCV (13)3
2023 Auto-Learning-GCN: An Ingenious Framework for Skeleton-Based Action Recognition
Wentian Xin, Ruyi Liu 0001, Qiguang Miao, Cheng Shi 0002, Chi-Man Pun
PRCV (1)6
2023 TriView-ParNet: parallel network for hybrid recognition of touching printed and handwritten strings based on feature fusion and three-view co-training
Junhao Qiu, Shangyu Lai, Guoheng Huang, Junhui Mai, Chi-Man Pun, Bingo Wing-Kuen Ling
Appl. Intell.6
2023 HIDE-Healthcare IoT Data Trust ManagEment: Attribute centric intelligent privacy approach
abstract
The cloud-based Internet of Things (IoTs) storage enables patients to monitor their health remotely and offers services for physicians of various Medical Institutions (MIs) to diagnose and treat them on time. As a matter of trust, patients are legally expected to hide their real identity and ensure data privacy in the cross-domain of IoT-healthcare, whether it is stored correctly or modified due to external and internal attacks in the cloud. Additionally, physicians treat patients and continuously store duplicated data in cloud storage, which increases the cost of computing. In this context, this paper presents HIDE-Healthcare IoT Data privacy trust management framework, focusing on attributes. Patients’ attributes are used to encrypt and decrypt sensory data between patients and different entities by incorporating the idea of trustworthy and secure shared keys. HIDE uses an intelligent object’s pointer to store the same patient’s sensory data in various versions to prevent data duplication, which will help track MIs that treat patients. An intelligent content-based emergency data access control is developed to monitor multiple patient health criticalities in HIDE. The security analysis and experimental evaluation attest to the benefits of the proposed HIDE framework, considering security and privacy metrics.
Fasee Ullah, Chi-Man Pun, Omprakash Kaiwartya, Ali Safa Sadiq, Jaime Lloret Mauri
Future Gener. Comput. Syst.2
2023 Deep self-learning based dynamic secret key generation for novel secure and efficient hashing algorithm
Fasee Ullah, Chi-Man Pun
Inf. Sci.2
2023 Towards evaluating the robustness of deep neural semantic segmentation networks with Feature-Guided Method
Yatie Xiao, Chi-Man Pun, Kongyang Chen
Knowl. Based Syst.2
2023 Simultaneous Laplacian embedding and subspace clustering for incomplete multi-view data
Guo Zhong, Chi-Man Pun
Knowl. Based Syst.2
2023 Self-taught Multi-view Spectral Clustering
Guo Zhong, Chi-Man Pun
Pattern Recognit.2
2023 RBA-GCN: Relational Bilevel Aggregation Graph Convolutional Network for Emotion Recognition
abstract
Emotion recognition in conversation (ERC) has received increasing attention from researchers due to its wide range of applications. As conversation has a natural graph structure, numerous approaches used to model ERC based on graph convolutional networks (GCNs) have yielded significant results. However, the aggregation approach of traditional GCNs suffers from the node information redundancy problem, leading to node discriminant information loss. Additionally, single-layer GCNs lack the capacity to capture long-range contextual information from the graph. Furthermore, the majority of approaches are based on textual modality or stitching together different modalities, resulting in a weak ability to capture interactions between modalities. To address these problems, we present the relational bilevel aggregation graph convolutional network (RBA-GCN), which consists of three modules: the graph generation module (GGM), similarity-based cluster building module (SCBM) and bilevel aggregation module (BiAM). First, GGM constructs a novel graph to reduce the redundancy of target node information. Then, SCBM calculates the node similarity in the target node and its structural neighborhood, where noisy information with low similarity is filtered out to preserve the discriminant information of the node. Meanwhile, BiAM is a novel aggregation method that can preserve the information of nodes during the aggregation process. This module can construct the interaction between different modalities and capture long-range contextual information based on similarity clusters. On both the IEMOCAP and MELD datasets, the weighted average F1 score of RBA-GCN has a 2.17$\sim$5.21% improvement over that of the most advanced method.
Guoheng Huang, Fenghuan Li, Xiaochen Yuan, Chi-Man Pun, Guo Zhong
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Quaternion-Valued Correlation Learning for Few-Shot Semantic Segmentation
abstract
Few-shot segmentation (FSS) aims to segment unseen classes given only a few annotated samples. Encouraging progress has been made for FSS by leveraging semantic features learned from base classes with sufficient training samples to represent novel classes. The correlation-based methods lack the ability to consider interaction of the two subspace matching scores due to the inherent nature of the real-valued 2D convolutions. In this paper, we introduce a quaternion perspective on correlation learning and propose a novel Quaternion-valued Correlation Learning Network (QCLNet), with the aim to alleviate the computational burden of high-dimensional correlation tensor and explore internal latent interaction between query and support images by leveraging operations defined by the established quaternion algebra. Specifically, our QCLNet is formulated as a hyper-complex valued network and represents correlation tensors in the quaternion domain, which uses quaternion-valued convolution to explore the external relations of query subspace when considering the hidden relationship of the support sub-dimension in the quaternion space. Extensive experiments on the PASCAL-$5^{i}$and COCO-$20^{i}$datasets demonstrate that our method outperforms the existing state-of-the-art methods effectively.
Zewen Zheng, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Bingo Wing-Kuen Ling
IEEE Trans. Circuits Syst. Video Technol.4
2023 Universal Object-Level Adversarial Attack in Hyperspectral Image Classification
abstract
The vulnerability of deep neural networks has garnered significant attention. Various advanced adversarial attack methods have been proposed. However, these methods exhibit higher attack performance on three-band natural images while struggling to handle high-dimensional attacks in terms of attack transferability and robustness. Hyperspectral images, unlike natural images, possess high-dimensional and redundant spectral information. On one hand, different classification models focus on distinct discriminative spectral bands, leading to poor transferability. On the other hand, most existing attack methods are implemented at the pixel-level, making them less resilient to image processing-based defenses. In this paper, we address the improvement of transferability and robustness in high-dimensional attacks and introduce a universal object-level adversarial attack method in hyperspectral image classification. We found that perturbations with higher similarity in a local region can decrease the sensitivity of adversarial attacks to various discriminative spectral patterns and enhance resistance to image processing-based defenses. Consequently, we construct spatial and spectral oversegmented templates by utilizing the local smooth properties of hyperspectral images, aiming to promote similarity among perturbations within a local region. Extensive experiments conducted on two real hyperspectral image datasets validate that our method enhances the attack transferability and robustness of several existing attack methods. By incorporating the object-level adversarial attack with baseline fast gradient sign method (FGSM), momentum iterative FGSM (MI-FGSM), and variance tuning MI-FGSM (VMI-FGSM), the average transferability success rate of the proposed method has increased by 7.38% on the PaviaU dataset and 9.30% on the HoustonU 2018 dataset than the baselines, respectively. Meanwhile, the proposed method outperforms the baselines by an average of 6.19% on the PaviaU dataset and 10.05% on the HoustonU 2018 dataset in attacking image processing-based defense models. The code is available at https://github.com/AAAA-CS/SS_FGSM_HyperspectralAdversarialAttack.
Cheng Shi 0002, Mengxin Zhang, Zhiyong Lv, Qiguang Miao, Chi-Man Pun
IEEE Trans. Geosci. Remote. Sens.5
2023 QGD-Net: A Lightweight Model Utilizing Pixels of Affinity in Feature Layer for Dermoscopic Lesion Segmentation
abstract
RESPONSE: Pixels with location affinity, which can be also called "pixels of affinity," have similar semantic information. Group convolution and dilated convolution can utilize them to improve the capability of the model. However, for group convolution, it does not utilize pixels of affinity between layers. For dilated convolution, after multiple convolutions with the same dilated rate, the pixels utilized within each layer do not possess location affinity with each other. To solve the problem of group convolution, our proposed quaternion group convolution uses the quaternion convolution, which promotes the communication between to promote utilizing pixels of affinity between channels. In quaternion group convolution, the feature layers are divided into 4 layers per group, ensuring the quaternion convolution can be performed. To solve the problem of dilated convolution, we propose the quaternion sawtooth wave-like dilated convolutions module (QS module). QS module utilizes quaternion convolution with sawtooth wave-like dilated rates to effectively leverage the pixels that share the location affinity both between and within layers. This allows for an expanded receptive field, ultimately enhancing the performance of the model. In particular, we perform our quaternion group convolution in QS module to design the quaternion group dilated neutral network (QGD-Net). Extensive experiments on Dermoscopic Lesion Segmentation based on ISIC 2016 and ISIC 2017 indicate that our method has significantly reduced the model parameters and highly promoted the precision of the model in Dermoscopic Lesion Segmentation. And our method also shows generalizability in retinal vessel segmentation.
Jingchao Wang 0002, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Chi-Man Pun
IEEE J. Biomed. Health Informatics5
2023 Bi-deformation-UNet: recombination of differential channels for printed surface defect detection
Guoheng Huang, Ying Wang 0097, Junhao Qiu, Zhiwen Yu 0002, Chi-Man Pun, Bingo Wing-Kuen Ling
Vis. Comput.7
2023 Difference-guided multi-scale spatial-temporal representation for sign language recognition
Liqing Gao, Lianyu Hu 0003, Fan Lyu, Lei Zhu 0003, Chi-Man Pun, Wei Feng 0005
Vis. Comput.6
2023 Video action recognition with Key-detail Motion Capturing based on motion spectrum analysis and multiscale feature fusion
Ganghan Zhang, Guoheng Huang, Haiyuan Chen, Chi-Man Pun, Zhiwen Yu 0002, Bingo Wing-Kuen Ling
Vis. Comput.4
2022 Spatial-Separated Curve Rendering Network for Efficient and High-Resolution Image Harmonization
Jingtang Liang, Xiaodong Cun, Chi-Man Pun, Jue Wang 0001
ECCV (7)3
2022 Fine-grained visual classification with multi-scale features based on self-supervised attention filtering mechanism
Haiyuan Chen, Lianglun Cheng, Guoheng Huang, Ganghan Zhang, Jiaying Lan, Zhiwen Yu 0002, Chi-Man Pun, Bingo Wing-Kuen Ling
Appl. Intell.7
2022 MIVCN: Multimodal interaction video captioning network based on semantic association graph
Ying Wang 0097, Guoheng Huang, Yuming Lin 0005, Chi-Man Pun, Bingo Wing-Kuen Ling, Lianglun Cheng
Appl. Intell.5
2022 Learning ordinal constraint binary codes for fast similarity search
Zheng Zhang 0006, Chi-Man Pun
Inf. Process. Manag.2
2022 Local Learning-based Multi-task Clustering
Guo Zhong, Chi-Man Pun
Knowl. Based Syst.2
2022 Improved Normalized Cut for Multi-View Clustering
abstract
Spectral clustering (SC) algorithms have been successful in discovering meaningful patterns since they can group arbitrarily shaped data structures. Traditional SC approaches typically consist of two sequential stages, i.e., performing spectral decomposition of an affinity matrix and then rounding the relaxed continuous clustering result into a binary indicator matrix. However, such a two-stage process could make the obtained binary indicator matrix severely deviate from the ground true one. This is because the former step is not devoted to achieving an optimal clustering result. To alleviate this issue, this paper presents a general joint framework to simultaneously learn the optimal continuous and binary indicator matrices for multi-view clustering, which also has the ability to tackle the conventional single-view case. Specially, we provide theoretical proof for the proposed method. Furthermore, an effective alternate updating algorithm is developed to optimize the corresponding complex objective. A number of empirical results on different benchmark datasets demonstrate that the proposed method outperforms several state-of-the-arts in terms of six clustering metrics.
Guo Zhong, Chi-Man Pun
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Temporal Relation Inference Network for Multimodal Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is a non-trivial task for humans, while it remains challenging for automatic SER due to the linguistic complexity and contextual distortion. Notably, previous automatic SER systems always regarded multi-modal information and temporal relations of speech as two independent tasks, ignoring their association. We argue that the valid semantic features and temporal relations of speech are both meaningful event relationships. This paper proposes a novel temporal relation inference network (TRIN) to help tackle multi-modal SER, which fully considers the underlying hierarchy of phonetic structure and its associations between various modalities under the sequential temporal guidance. Mainly, we design a temporal reasoning calibration module to imitate real and abundant contextual conditions. Unlike the previous works, which assume all multiple modalities are related, it infers the dependency relationship between the semantic information from the temporal level and learns to handle the multi-modal interaction sequence with a flexible order. To enhance the feature representation, an innovative temporal attentive fusion unit is developed to magnify the details embedded in a single modality from semantic level. Meanwhile, it aggregates the feature representation from both the temporal and semantic levels to maximize the integrity of feature representation by an adaptive feature fusion mechanism to selectively collect the implicit complementary information to strengthen the dependencies between different information subspaces. Extensive experiments conducted on two benchmark datasets demonstrate the superiority of our TRIN method against some state-of-the-art SER methods.
Guannan Dong, Chi-Man Pun, Zheng Zhang 0006
IEEE Trans. Circuits Syst. Video Technol.2
2022 Monocular Robust 3D Human Localization by Global and Body-Parts Depth Awareness
abstract
Learning the human depth localization in camera coordinate space plays a crucial role in understanding the behavior and activities of multi-person in 3D scenes. However, existing monocular-based methods rarely combine the global image features and the human body-parts features effectively, resulting in a large gap from the actual location in some cases, e.g., the special body-sized persons and mutual occlusion between humans in the image. This paper presents a novel Robust 3D Human Localization (R3HL) network consisting of two stages: global depth awareness and body-parts depth awareness, to significantly improve the robustness and accuracy of the 3D location. In the first stage, the front-back and far-near relationship estimation module based on multi-person are proposed to make the network extract depth features from the global perspective. In the second stage, the network focuses on the target human. We propose a Pose-guided Multi-person Repulsion (PMR) module to enhance the target human’s features and reduce the interference features produced by the background and other people. In addition, an Adaptive Body-parts Attention (ABA) module is designed to assign different feature weights to each joint. Finally, the human’s absolute depth is obtained through global pooling and fully connected layers. The experimental results show that the attention from the whole image to a single person helps find the absolute location of different body-sized and poses people from diverse scenes. Our method can achieve better performance than other state-of-the-art methods on both indoor and outdoor 3D multi-person datasets.
Haolun Li 0001, Chi-Man Pun
IEEE Trans. Circuits Syst. Video Technol.2
2022 An Efficient Artificial Bee Colony Algorithm With an Improved Linkage Identification Method
abstract
The artificial colony (ABC) algorithm shows a relatively powerful exploration search capability but is constrained by the curse of dimensionality, especially on nonseparable functions, where its convergence speed slows dramatically. In this article, based on an analysis of the difference between updating mechanisms that include both all-variable and one-variable updating mechanisms, we find that when equipped with the former strategy, the algorithm rapidly converges to an optimal region, while with the latter strategy, it searches the solution space thoroughly. To utilize multivariable and one-variable updating mechanisms on nonseparable and separable functions, respectively, we embed an improved linkage identification strategy into the ABC by detecting the linkage between variables more effectively. Then, we propose three common strategies for ABC to improve its performance. First, a new approach that considers the historic experiences of the population is proposed to balance exploration and exploitation. Second, a new strategy for initializing scout bees is used to reduce the number of function evaluations. Finally, the individual with the worst performance is updated with a defined probability on multiple dimensions instead of one dimension, causing it to follow the population steps on nonseparable functions. This article is the first to propose all these concepts, which could be adopted for other ABC variants. The effectiveness of our algorithm is validated through basic, CEC2010, CEC2013, and CEC2014 functions and real-world problems.
Hao Gao 0005, Zheng Fu, Chi-Man Pun, Jun Zhang 0003, Sam Kwong
IEEE Trans. Cybern.3
2022 Reversible Data Hiding in Encrypted Images using Chunk Encryption and Redundancy Matrix Representation
abstract
For reversible data hiding in encrypted images (RDHEI), the private information in the original image content is protected and the embedded secret data can be used to manage the encrypted image. Up to now, many RDHEI algorithms have been reported, however, these reported RDHEI algorithms cannot achieve a high embedding capacity. Therefore, this article proposes a new RDHEI scheme, which will have improved performance. The main contributions of the proposed scheme include two points: chunk encryption (CE) and redundancy matrix representation (RMR). First, different from the encryption method of pixel-by-pixel encryption used in the previous RDHEI algorithms, CE is the method that uses the traditional exclusive-or (XOR) method to encrypt the original image in chunks. By using the CE method, redundancy will be retained in the encrypted chunks. Second, the RMR method is designed to generate available room for accommodating secret data in redundancy matrices, which are existed in some encrypted chunks. Experimental results show that the proposed scheme achieves a high embedding capacity and the directly decrypted images retain high quality.
Chi-Man Pun
IEEE Trans. Dependable Secur. Comput.2
2022 Multifeature Collaborative Adversarial Attack in Multimodal Remote Sensing Image Classification
abstract
Deep neural networks have strong feature learning ability, but their vulnerability cannot be ignored. Current research shows that deep learning models are threatened by adversarial examples in remote sensing (RS) classification tasks, and their robustness drops sharply in the face of adversarial attacks. Therefore, many adversarial attack methods have been studied to predict the risks faced by a network. However, the existing adversarial attack methods mainly focus on single-modal image classification networks, and the rapid growth of RS data makes multimodal RS image classification a research hotspot. Generating multimodal adversarial examples needs to consider a high attack success rate, subtle perturbation, and collaborative attack ability between different modalities. In this article, we investigate the vulnerability of multimodal RS classification networks and propose a multifeature collaborative adversarial network (MFCANet) for generating multimodal adversarial examples. Two modality-specific generators are designed to generate the multimodal collaborative perturbations with strong attack ability, and two modality-specific discriminators make the generated multimodal adversarial examples closer to the real instances. In addition, a modality-specific generative loss and a modality-specific discriminative loss are proposed, and an alternating optimization strategy is designed for training the proposed MFCANet. Extensive experiments are carried out on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen 2D dataset and ISPRS Potsdam 2D dataset. The results show that the attack performance of the proposed method is stronger than that of the fast gradient sign method (FGSM), project gradient descent (PGD), and Carlini and Wagner (C&W) attack methods.
Cheng Shi 0002, Yenan Dang, Minghua Zhao, Zhiyong Lv, Qiguang Miao, Chi-Man Pun
IEEE Trans. Geosci. Remote. Sens.7
2022 Implicit and Explicit Feature Purification for Age-Invariant Facial Representation Learning
abstract
This paper presents a new method, named implicit and explicit feature purification (IEFP), for age-invariant face recognition. Facial features extracted from a face image contain the information about the identity, age, and other attributes. For age-invariant face recognition, it is important to remove the irrelevant information, and retain the identity information only, in the facial features. Through the two proposed feature purification mechanisms, our framework can produce facial-feature embeddings that preserve identity information as much as possible and are insensitive to age variations. Specifically, on the one hand, a special network module is devised to implicitly purify the original facial features obtained from a face encoder. On the other hand, to obtain purer facial feature representations for age-invariant face recognition, irrelevant information within the implicitly purified features, such as the age, is further removed. This is realized by using a regularizer, based on information theory, to explicitly minimize the correlation between identity-related features and age-related features. Comprehensive ablation studies show that these two feature purification schemes can work independently, as well as collaboratively, to achieve better performance. Extensive evaluations on several benchmark data sets show that the IEFP method is on par with those competitors learned on far more favorable training samples, and it achieves the best performance in a fair comparison. Furthermore, we provide mathematical interpretation to explain the effectiveness of our approach, and find that it tends to generate low-rank, yet high-dimensional, representations for age-invariant face recognition.
Jiucheng Xie, Chi-Man Pun, Kin-Man Lam 0001
IEEE Trans. Inf. Forensics Secur.2
2022 Action Recognition Framework in Traffic Scene for Autonomous Driving System
abstract
For the autonomous driving system, accurately recognizing the actions of different roles in the traffic scene is the prerequisite for realizing this kind of human-vehicle information interaction. In this paper, we propose a complete framework based on 3D human pose estimation to recognize the actions of different roles on the road. The main objects recognized include traffic police, cyclists, and some passersby in need. We perform action recognition based on a dynamic adaptive graph convolutional network, which can realize the action recognition of objects based on 3D human pose. In addition to the action recognition module, we have optimized both the object detection module and the human pose estimation module in the framework so that the framework can handle multiple objects at the same time, which can be closer to the real traffic scene. To realize complex and changeable human action recognition, we built a multi-view camera system to collect responsible 3D human pose datasets containing traffic police gestures, cyclist gestures, and pedestrians’ body movements. In the experiments, compared to other state-of-the-art researches, the proposed framework can achieve comparable results with the same dataset. Satisfactory performance has also been obtained on the real data we collected, which can handle a variety of different action recognition tasks at the same time.
Feiyi Xu, Feng Xu 0005, Jiucheng Xie, Chi-Man Pun, Huimin Lu 0001, Hao Gao 0005
IEEE Trans. Intell. Transp. Syst.4
2022 Data Representation by Joint Hypergraph Embedding and Sparse Coding
abstract
Matrix factorization (MF), a popular unsupervised learning technique for data representation, has been widely applied in data mining and machine learning. According to different application scenarios, one can impose different constraints on the factorization to find the desired basis, which captures high-level semantics for the given data, and learns the compact representation corresponding to the basis. We note that almost all previous work on MF in data mining has ignored to find such a basis, which can carry high-order semantics in the data. In this article, we propose a novel MF framework called Joint Hypergraph Embedding and Sparse Coding (JHESC), in which the obtained basis captures high-order semantic information in data. Specifically, we first propose a new hypergraph learning model to obtain a more discriminative basis by hypergraph-based Laplacian Eigenmap, then sparse coding is conducted on the learned basis such that the new representation has stronger identification capability. In addition, we extend the proposed method to the reproducing kernel Hilbert space for dealing with nonlinear data more effectively. Extensive experimental results on data clustering demonstrate that the proposed method consistently outperforms the other state-of-the-art matrix factorization methods.
Guo Zhong, Chi-Man Pun
IEEE Trans. Knowl. Data Eng.2
2022 Robust Audio Patch Attacks Using Physical Sample Simulation and Adversarial Patch Noise Generation
abstract
Deep neural network (DNNs) based Automatic Speech Recognition (ASR) systems are known vulnerable to adversarial attacks that are maliciously implemented by adding small but powerful distortions to the original audio input. However, most existing methods that generate audio adversarial examples targeting ASR models cannot achieve successful robust attacks against defense methods. This paper proposes a novel framework for robust audio patch attacks using Physical Sample Simulation (PSS) and Adversarial Patch Noise Generation (APNG). First, the proposed PSS simulated real-audio with selected room impulse response for training the adversarial patches. Second, the proposed APNG generates the imperceptible audio adversarial patch examples using the voice activity detector to hide the adversarial patch noise into the non-silent locations of the input audio. Furthermore, the design Sounds Pressure Level-based adaptive noise minimization algorithm helps us further reduce the perturbation during the attack. The experimental results show that our proposed method can achieve the highest attack success rates and SNRs in various cases, comparing with other state-of-the-art attacks.
Xia Du, Chi-Man Pun
IEEE Trans. Multim.2
2022 Virtual Reality Aided High-Quality 3D Reconstruction by Remote Drones
abstract
Artificial intelligence including deep learning and 3D reconstruction methods is changing the daily life of people. Now, an unmanned aerial vehicle that can move freely in the air and avoid harsh ground conditions has been commonly adopted as a suitable tool for 3D reconstruction. The traditional 3D reconstruction mission based on drones usually consists of two steps: image collection and offline post-processing. But there are two problems: one is the uncertainty of whether all parts of the target object are covered, and another is the tedious post-processing time. Inspired by modern deep learning methods, we build a telexistence drone system with an onboard deep learning computation module and a wireless data transmission module that perform incremental real-time dense reconstruction of urban cities by itself. Two technical contributions are proposed to solve the preceding issues. First, based on the popular depth fusion surface reconstruction framework, we combine it with a visual-inertial odometry estimator that integrates the inertial measurement unit and allows for robust camera tracking as well as high-accuracy online 3D scan. Second, the capability of real-time 3D reconstruction enables a new rendering technique that can visualize the reconstructed geometry of the target as navigation guidance in the HMD. Therefore, it turns the traditional path-planning-based modeling process into an interactive one, leading to a higher level of scan completeness. The experiments in the simulation system and our real prototype demonstrate an improved quality of the 3D model using our artificial intelligence leveraged drone system.
Feng Xu 0005, Chi-Man Pun, Yang Yang 0002, Rushi Lan, Yujie Li 0001, Hao Gao 0005
ACM Trans. Internet Techn.3
2021 Split then Refine: Stacked Attention-guided ResUNets for Blind Single Image Visible Watermark Removal
abstract
Digital watermark is a commonly used technique to protect the copyright of medias. Simultaneously, to increase the robustness of watermark, attacking technique, such as watermark removal, also gets the attention from the community. Previous watermark removal methods require to gain the watermark location from users or train a multi-task network to recover the background indiscriminately. However, when jointly learning, the network performs better on watermark detection than recovering the texture. Inspired by this observation and to erase the visible watermarks blindly, we propose a novel two-stage framework with a stacked attention-guided ResUNets to simulate the process of detection, removal and refinement. In the first stage, we design a multi-task network called SplitNet. It learns the basis features for three sub-tasks altogether while the task-specific features separately use multiple channel attentions. Then, with the predicted mask and coarser restored image, we design RefineNet to smooth the watermarked region with a mask-guided spatial attention. Besides network structure, the proposed algorithm also combines multiple perceptual losses for better quality both visually and numerically. We extensively evaluate our algorithm over four different datasets under various settings and the experiments show that our approach outperforms other state-of-the-art methods by a large margin.
Xiaodong Cun, Chi-Man Pun
AAAI2
2021 Latent Low-rank Graph Learning for Multimodal Clustering
abstract
Multimodal clustering has become a fundamental and important problem in the data mining community since the development of multimedia technology over the last two decades has led to a tremendous increase in unlabeled multimodal data. Although a panoply of multimodal subspace clustering methods shows promising performance via fusing information from different views of multimodal data, most of them consist of two sequential steps, i.e., learning a consensus affinity matrix from the original data and then feeding the resulting affinity matrix into the framework of spectral clustering. However, this leads to the suboptimal clustering performance due to the following limitations: 1) the two steps of learning the affinity matrix and clustering are carried out independently; 2) the affinity matrix may be unreliable; 3) the post-processing requirement, such as K-means. To address these issues, we propose a novel multimodal subspace clustering method via adaptively learning a similarity graph on a latent low-rank representation space. In particular, the number of connected components of the learned graph is precisely equal to the number of clusters, i.e., the optimal solution of the associated problem directly reveals the clustering structure of data. Extensive evaluations on several benchmark multimodal datasets demonstrate that the proposed approach outperforms state-of-the-art methods.
Guo Zhong, Chi-Man Pun
ICDE2
2021 A Rate-based Drone Control with Adaptive Origin Update in Telexistence
abstract
A new form of telexistence is achieved by recording videos with a camera on an Uncrewed aerial vehicle (UAV) and playing the videos to a user via a head-mounted display (HMD). One key problem here is how to let the user freely and naturally control the UAV and thus the viewpoint. In this paper, we develop an HMD-based telexistence technique that achieves full 6- DOF control of the viewpoint. The core of our technique is an improved rate-based control technique with our adaptive origin update (AOU), in which the origin of the coordinate system of the user changes adaptively. This makes the user naturally perceive the origin and thus easily perform the control motion to get his/her desired viewpoint changing. As a consequence, without the aid of any auxiliary equipment, the AOU scheme handles the well known self-centering problem in the rate-based control methods. A real prototype is also built to evaluate this feature of our technique. To explore the advantage of our telexistence technique, we further use it as an interactive tool to perform the task of 3D scene reconstruction. User studies demonstrate that comparing with other telexistence solutions and the widely used joystick-based solutions, our solution largely reduces the workload and saves time and moving distance for the user.
Chi-Man Pun, Yang Yang 0002, Hao Gao 0005, Feng Xu 0005
VR2
2021 An improved artificial bee colony algorithm based on elite search strategy with segmentation application on robot vision system
abstract
Summary Aiming at accelerating the convergence speed and enhancing relative poor local search ability of the traditional artificial bee colony algorithm (ABC), this article introduces an ABC with a new elite search strategy. First, we propose a strategy of recording individuals with high performance. Then bees have more chances to learn from a real elite. In the onlooked bee phase, its updating equation is changed for having more opportunities to search in a valuable area. Furthermore, for saving the value of function evaluations, a new learning equation for the best onlooked bee is proposed. The image segmentation of a robot binocular stereo vision system is a key problem in mechanical robot vision system, but the computation time limits its application. The experimental results show that the proposed algorithm achieves better performance on 10 benchmark functions and the image segmentation problem of mechanical robot in comparison with several other state of the art algorithms.
Chuyi Gao, Maolong Xi, Jian Xiong 0005, Chi-Man Pun, Hao Gao 0005
Concurr. Comput. Pract. Exp.7
2021 RPCA-induced self-representation for subspace clustering
Guo Zhong, Chi-Man Pun
Neurocomputing2
2021 Improving adversarial attacks on deep neural networks via constricted gradient-based perturbations
Yatie Xiao, Chi-Man Pun
Inf. Sci.2
2021 Fooling deep neural detection networks with adaptive object-oriented adversarial perturbation
Yatie Xiao, Chi-Man Pun, Bo Liu 0047
Pattern Recognit.2
2021 Deep Collaborative Multi-Modal Learning for Unsupervised Kinship Estimation
abstract
Kinship verification is a long-standing research challenge in computer vision. The visual differences presented to the face have a significant effect on the recognition capabilities of the kinship systems. We argue that aggregating multiple visual knowledge can better describe the characteristics of the subject for precise kinship identification. Typically, the age-invariant features can represent more natural facial details. Such age-related transformations are essential for face recognition due to the biological effects of aging. However, the existing methods mainly focus on employing the single-view image features for kinship identification, while more meaningful visual properties such as race and age are directly ignored in the feature learning step. To this end, we propose a novel deep collaborative multi-modal learning (DCML) to integrate the underlying information presented in facial properties in an adaptive manner to strengthen the facial details for effective unsupervised kinship verification. Specifically, we construct a well-designed adaptive feature fusion mechanism, which can jointly leverage the complementary properties from different visual perspectives to produce composite features and draw greater attention to the most informative components of spatial feature maps. Particularly, an adaptive weighting strategy is developed based on a novel attention mechanism, which can enhance the dependencies between different properties by decreasing the information redundancy in channels in a self-adaptive manner. Moreover, we propose to use self-supervised learning to further explore the intrinsic semantics embedded in raw data and enrich the diversity of samples. As such, we could further improve the representation capabilities of kinship feature learning and mitigate the multiple variations from original visual images. To validate the effectiveness of the proposed method, extensive experimental evaluations conducted on four widely-used datasets show that our DCML method is always superior to some state-of-the-art kinship verification methods.
Guannan Dong, Chi-Man Pun, Zheng Zhang 0006
IEEE Trans. Inf. Forensics Secur.2
2021 Personal Privacy Protection via Irrelevant Faces Tracking and Pixelation in Video Live Streaming
abstract
To date, the privacy-protection intended pixelation tasks are still labor-intensive and yet to be studied. With the prevailing of video live streaming, establishing an online face pixelation mechanism during streaming is an urgency. In this paper, we develop a new method called Face Pixelation in Video Live Streaming (FPVLS) to generate automatic personal privacy filtering during unconstrained streaming activities. Simply applying multi-face trackers will encounter problems in target drifting, computing efficiency, and over-pixelation. Therefore, for fast and accurate pixelation of irrelevant people's faces, FPVLS is organized in a frame-to-video structure of two core stages. On individual frames, FPVLS utilizes image-based face detection and embedding networks to yield face vectors. In the raw trajectories generation stage, the proposed Positioned Incremental Affinity Propagation (PIAP) clustering algorithm leverages face vectors and positioned information to quickly associate the same person's faces across frames. Such frame-wise accumulated raw trajectories are likely to be intermittent and unreliable on video level. Hence, we further introduce the trajectory refinement stage that merges a proposal network with the two-sample test based on the Empirical Likelihood Ratio (ELR) statistic to refine the raw trajectories. A Gaussian filter is laid on the refined trajectories for final pixelation. On the video live streaming dataset we collected, FPVLS obtains satisfying accuracy, real-time efficiency, and contains the over-pixelation problems.
Jizhe Zhou 0001, Chi-Man Pun
IEEE Trans. Inf. Forensics Secur.2
2021 Kinship Verification Based on Cross-Generation Feature Interaction Learning
abstract
Kinship verification from facial images has been recognized as an emerging yet challenging technique in many potential computer vision applications. In this paper, we propose a novel cross-generation feature interaction learning (CFIL) framework for robust kinship verification. Particularly, an effective collaborative weighting strategy is constructed to explore the characteristics of cross-generation relations by corporately extracting features of both parents and children image pairs. Specifically, we take parents and children as a whole to extract the expressive local and non-local features. Different from the traditional works measuring similarity by distance, we interpolate the similarity calculations as the interior auxiliary weights into the deep CNN architecture to learn the whole and natural features. These similarity weights not only involve corresponding single points but also excavate the multiple relationships cross points, where local and non-local features are calculated by using these two kinds of distance measurements. Importantly, instead of separately conducting similarity computation and feature extraction, we integrate similarity learning and feature extraction into one unified learning process. The integrated representations deduced from local and non-local features can comprehensively express the informative semantics embedded in images and preserve abundant correlation knowledge from image pairs. Extensive experiments demonstrate the efficiency and superiority of the proposed model compared to some state-of-the-art kinship verification methods.
Guannan Dong, Chi-Man Pun, Zheng Zhang 0006
IEEE Trans. Image Process.2
2021 A Hybrid Feature Selection Algorithm Based on a Discrete Artificial Bee Colony for Parkinson's Diagnosis
abstract
Parkinson's disease is a neurodegenerative disease that affects millions of people around the world and cannot be cured fundamentally. Automatic identification of early Parkinson's disease on feature data sets is one of the most challenging medical tasks today. Many features in these datasets are useless or suffering from problems like noise, which affect the learning process and increase the computational burden. To ensure the optimal classification performance, this article proposes a hybrid feature selection algorithm based on an improved discrete artificial bee colony algorithm to improve the efficiency of feature selection. The algorithm combines the advantages of filters and wrappers to eliminate most of the uncorrelated or noisy features and determine the optimal subset of features. In the filter, three different variable ranking methods are employed to pre-rank the candidate features, then the population of artificial bee colony is initialized based on the significance degree of the re-rank features. In the wrapper part, the artificial bee colony algorithm evaluates individuals (feature subsets) based on the classification accuracy of the classifier to achieve the optimal feature subset. In addition, for the first time, we introduce a strategy that can automatically select the best classifier in the search framework more quickly. By comparing with several publicly available datasets, the proposed method achieves better performance than other state-of-the-art algorithms and can extract fewer effective features.
Haolun Li 0001, Chi-Man Pun, Feng Xu 0005, Longsheng Pan, Rui Zong, Hao Gao 0005, Huimin Lu 0001
ACM Trans. Internet Techn.2
2020 Towards Ghost-Free Shadow Removal via Dual Hierarchical Aggregation Network and Shadow Matting GAN
abstract
Shadow removal is an essential task for scene understanding. Many studies consider only matching the image contents, which often causes two types of ghosts: color in-consistencies in shadow regions or artifacts on shadow boundaries (as shown in Figure. 1). In this paper, we tackle these issues in two ways. First, to carefully learn the border artifacts-free image, we propose a novel network structure named the dual hierarchically aggregation network (DHAN). It contains a series of growth dilated convolutions as the backbone without any down-samplings, and we hierarchically aggregate multi-context features for attention and prediction, respectively. Second, we argue that training on a limited dataset restricts the textural understanding of the network, which leads to the shadow region color in-consistencies. Currently, the largest dataset contains 2k+ shadow/shadow-free image pairs. However, it has only 0.1k+ unique scenes since many samples share exactly the same background with different shadow positions. Thus, we design a shadow matting generative adversarial network (SMGAN) to synthesize realistic shadow mattings from a given shadow mask and shadow-free image. With the help of novel masks or scenes, we enhance the current datasets using synthesized shadow images. Experiments show that our DHAN can erase the shadows and produce high-quality ghost-free images. After training on the synthesized and real datasets, our network outperforms other state-of-the-art methods by a large margin. The code is available: http://github.com/vinthony/ghost-free-shadow-removal/
Xiaodong Cun, Chi-Man Pun, Cheng Shi 0002
AAAI2
2020 Defocus Blur Detection via Depth Distillation
Xiaodong Cun, Chi-Man Pun
ECCV (13)2
2020 A Unified Framework for Multi-view Spectral Clustering
abstract
In the era of big data, multi-view clustering has drawn considerable attention in machine learning and data mining communities due to the existence of a large number of unlabeled multi-view data in reality. Traditional spectral graph theoretic methods have recently been extended to multi-view clustering and shown outstanding performance. However, most of them still consist of two separate stages: learning a fixed common real matrix (i.e., continuous labels) of all the views from original data, and then applying K-means to the resulting common label matrix to obtain the final clustering results. To address these, we design a unified multi-view spectral clustering scheme to learn the discrete cluster indicator matrix in one stage. Specifically, the proposed framework directly obtain clustering results without performing K-means clustering. Experimental results on several famous benchmark datasets verify the effectiveness and superiority of the proposed method compared to the state-of-the-arts.
Guo Zhong, Chi-Man Pun
ICDE2
2020 Adversarial Image Attacks Using Multi-Sample and Most-Likely Ensemble Methods
abstract
Many studies on deep neural networks have shown very promising results for most image recognition tasks. However, these networks can often be fooled by adversarial examples that simply add small but powerful distortions to the original input. Recent works have demonstrated the vulnerability of deep learning systems to adversarial examples, but most such works directly manipulate and attack the digital images for a specific classifier only, and cannot attack the physical images in real world. In this paper, we propose the multi-sample ensemble method (MSEM) and most-likely ensemble method (MLEM) to generate adversarial attacks that successfully fool the classifier for images in both the digital and real worlds. The proposed adaptive norm algorithm can craft faster and smaller perturbation than other state-of-the-art attack methods. Besides, the proposed MLEM extended with weighted objective function can generate robust adversarial attacks that can mislead multiple classifiers (Inception-v3, Inception-v4, Resnet-v2, Ince-res-v2) simultaneously for physical images in real world. Compared with other methods, experiments show that our adversarial attack methods not only can achieve higher success rates but also can survive in the multi-model defense tests.
Xia Du, Chi-Man Pun
ACM Multimedia2
2020 A Unified Framework for Detecting Audio Adversarial Examples
abstract
Adversarial attacks have been widely recognized as the security vulnerability of deep neural networks, especially in deep automatic speech recognition (ASR) systems. The advanced detection methods against adversarial attacks mainly focus on pre-processing the input audio to alleviate the threat of adversarial noise. Although these methods could detect some simplex adversarial attacks, they fail to handle robust complex attacks especially when the attacker knows the detection details. In this paper, we propose a unified adversarial detection framework for detecting adaptive audio adversarial examples, which combines noise padding with sound reverberation. Specifically, a well-designed adaptive artificial utterances generator is proposed to balance the design complexity, such that the artificial utterances (speech with reverberation) are efficiently determined to reduce the false positive rate and false negative rate of detection results. Moreover, to destroy the continuity of the adversarial noise, we develop a novel multi-noise padding strategy, which implants the Gaussian noises in the silent fragments of the input speech by the voice activity detector. Furthermore, our proposed method can effectively tackle the robust adaptive attacks in an adaptive learning manner. Importantly, the conceived system is easily embedded into any ASR models without requiring additional retraining or modification. The experimental results show that our method consistently outperforms the state-of-the-art audio defense methods, even for the adaptive and robust attacks.
Xia Du, Chi-Man Pun, Zheng Zhang 0006
ACM Multimedia2
2020 Privacy-sensitive Objects Pixelation for Live Video Streaming
abstract
With the prevailing of live video streaming, establishing an online pixelation method for privacy-sensitive objects is an urgency. Caused by the inaccurate detection of privacy-sensitive objects, simply migrating the tracking-by-detection structure applied in offline pixelation into the online form will incur problems in target initialization, drifting, and over-pixelation. To cope with the inevitable but impacting detection issue, we propose a novel Privacy-sensitive Objects Pixelation (PsOP) framework for automatic personal privacy filtering during live video streaming. Leveraging pre-trained detection networks, our PsOP is extendable to any potential privacy-sensitive objects pixelation. Employing the embedding networks and the proposed Positioned Incremental Affinity Propagation (PIAP) clustering algorithm as the backbone, our PsOP unifies the pixelation of discriminating and indiscriminating pixelation objects through trajectories generation. In addition to the pixelation accuracy boosting, experiment results on the streaming video data we built show that the proposed PsOP can significantly reduce the over-pixelation ratio in privacy-sensitive object pixelation.
Jizhe Zhou 0001, Chi-Man Pun, Yu Tong 0003
ACM Multimedia2
2020 News Image Steganography: A Novel Architecture Facilitates the Fake News Identification
abstract
A larger portion of fake news quotes untampered images from other sources with ulterior motives rather than conducting image forgery. Such elaborate engraftments keep the inconsistency between images and text reports stealthy, thereby, palm off the spurious for the genuine. This paper proposes an architecture named News Image Steganography (NIS) to reveal the aforementioned inconsistency through image steganography based on GAN. Extractive summarization about a news image is generated based on its source texts, and a learned steganographic algorithm encodes and decodes the summarization of the image in a manner that approaches perceptual invisibility. Once an encoded image is quoted, its source summarization can be decoded and further presented as the ground truth to verify the quoting news. The pairwise encoder and decoder endow images of the capability to carry along their imperceptible summarization. Our NIS reveals the underlying inconsistency, thereby, according to our experiments and investigations, contributes to the identification accuracy of fake news that engrafts untampered images.
Jizhe Zhou 0001, Chi-Man Pun, Yu Tong 0003
VCIP2
2020 Robust image hashing with visual attention model and invariant moments
abstract
Image hashing is an efficient technique of multimedia processing for many applications, such as image copy detection, image authentication, and social event detection. In this study, the authors propose a novel image hashing with visual attention model and invariant moments. An important contribution is the weighted DWT (discrete wavelet transform) representation by incorporating a visual attention model called Itti saliency model into LL sub‐band. Since the Itti saliency model can efficiently extract saliency map reflecting regions of attention focus, perceptual robustness of the proposed hashing is achieved. In addition, as invariant moments are robust and discriminative features, hash construction with invariant moments extracted from the weighted DWT representation ensures good classification performance between robustness and discrimination. Extensive experiments with open image datasets are done to validate the performances of the proposed hashing. The results demonstrate that the proposed hashing is robust and discriminative. Performance comparisons with some hashing algorithms are also conducted, and the receiver operating characteristic results illustrate that the proposed hashing outperforms the compared hashing algorithms in classification performance between robustness and discrimination.
Zhenjun Tang, Hanyun Zhang, Chi-Man Pun, Mengzhu Yu, Chunqiang Yu, Xianquan Zhang
IET Image Process.3
2020 New multi-view human motion capture framework
abstract
Estimating human pose and shape without markers is a challenging problem. This study proposes a multiple‐view markerless human motion capture framework. Firstly, a multi‐view camera system is built for capturing real‐time images of moving humans on multiple views. Secondly, by employing the OpenPose method, the authors calculate robust 3D key points from 2D key points of the human body, which are estimated from the multi‐view images. And dense 3D point cloud is reconstructed from images. Thirdly, they propose a novel SMPL‐based method to represent human motion by fitting the SMPL model to 3D key points and 3D point clouds. In order to achieve a more accurate human pose, a penalty term is utilised to solve the problem of error accumulation in the process of human motion capture. In addition, they present a dense mesh template‐based SMPL that can be deformed to point cloud to recover a real human body shape. Finally, they map multi‐view colour images onto the human mesh model to acquire rendered mesh. The experimental results show that the proposed method improves the accuracy of human pose and realises the 3D human body model more realistic.
Feiyi Xu, Chi-Man Pun, Wenqi Xiao, Jianhui Nie, Jian Xiong 0005, Hao Gao 0005, Feng Xu 0005
IET Image Process.3
2020 Locating splicing forgery by adaptive-SVD noise estimation and vicinity noise descriptor
Bo Liu 0047, Chi-Man Pun
Neurocomputing2
2020 Crafting adversarial example with adaptive root mean square gradient on deep neural networks
Yatie Xiao, Chi-Man Pun, Bo Liu 0047
Neurocomputing2
2020 Training Feed-Forward Artificial Neural Networks with a modified artificial bee colony algorithm
Feiyi Xu, Chi-Man Pun, Haolun Li 0001, Yushu Zhang 0001, Yurong Song, Hao Gao 0005
Neurocomputing2
2020 Revisiting Nyström extension for hypergraph clustering
Guo Zhong, Chi-Man Pun
Neurocomputing2
2020 Exposing splicing forgery in realistic scenes using deep fusion network
Bo Liu 0047, Chi-Man Pun
Inf. Sci.2
2020 Adversarial example generation with adaptive gradient search for single and ensemble deep neural network
Yatie Xiao, Chi-Man Pun, Bo Liu 0047
Inf. Sci.2
2020 Two-pass hashing feature representation and searching method for copy-move forgery detection
Chi-Man Pun
Inf. Sci.2
2020 Nonnegative self-representation with a fixed rank constraint for subspace clustering
Guo Zhong, Chi-Man Pun
Inf. Sci.2
2020 Dense moment feature index and best match algorithms for video copy-move forgery detection
Chi-Man Pun
Inf. Sci.2
2020 Subspace clustering by simultaneously feature selection and similarity learning
Guo Zhong, Chi-Man Pun
Knowl. Based Syst.2
2020 Endmember Extraction of Hyperspectral Remote Sensing Images Based on an Improved Discrete Artificial Bee Colony Algorithm and Genetic Algorithm
Zheng Fu, Chi-Man Pun, Hao Gao 0005, Huimin Lu 0001
Mob. Networks Appl.2
2020 Rapid facial expression recognition under part occlusion based on symmetric SURF and heterogeneous soft partition network
Guoheng Huang, Chi-Man Pun, Bingo Wing-Kuen Ling, Lianglun Cheng
Multim. Tools Appl.4
2020 An artificial bee algorithm with a leading group and its application into image registration
Haidong Hu, Chi-Man Pun, Ye Liu 0005, Xiangjing Lai, Hao Gao 0005
Multim. Tools Appl.2
2020 Person re-identification based on multi-level feature complementarity of cross-attention with part metric learning
Zeng Lu, Guoheng Huang, Chi-Man Pun, Lianglun Cheng
Multim. Tools Appl.3
2020 Vehicle power train optimization using multi-objective bird swarm algorithm
Dongmei Wu, Chi-Man Pun, Bin Xu 0014, Hao Gao 0005, Zhenghua Wu
Multim. Tools Appl.2
2020 High-quality-guided artificial bee colony algorithm for designing loudspeaker
Hao Gao 0005, Haolun Li 0001, Ye Liu 0005, Huimin Lu 0001, Hyoungseop Kim, Chi-Man Pun
Neural Comput. Appl.6
2020 Audio Replay Spoof Attack Detection by Joint Segment-Based Linear Filter Bank Feature Extraction and Attention-Enhanced DenseNet-BiLSTM Network
abstract
Most automatic speaker verification (ASV) systems are vulnerable to various spoofing attacks. In recent years, there have been many methods were proposed for detecting spoofing attacks in ASV, and significant progress has been made. However, current methods have shown little improvements in replay spoof attack detection as they lack a more suitable model for replay detection. To address this issue, in this article, we propose a novel model based on attention-enhanced DenseNet-BiLSTM network and segment-based linear filter bank features. First, silent segments are selected from each speech signal by using a short-term zero-crossing rate and energy. If the total duration of silent segments only contains a very limited amount of data, the decaying tails will be selected instead. Second, the linear filter bank features are extracted from the selected segments in the relatively high-frequency domain. Finally, an attention-enhanced DenseNet-BiLSTM architecture which can avoid the problems of overfitting is built. To validate this model, we used two datasets, including BTAS2016 and ASVspoof2017. Experiments show that using the attention-enhanced DenseNet-BiLSTM model with the segment-based linear filter bank feature achieves the best performance. Compared with the baseline system based on constant Q cepstral coefficient and Gaussian mixture model (GMM), the proposed model can produce a relative improvement of 91.68% and 74.04% on the two data sets respectively.
Lian Huang, Chi-Man Pun
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Deep and Ordinal Ensemble Learning for Human Age Estimation From Facial Images
abstract
Some recent work treats age estimation as an ordinal ranking task and decomposes it into multiple binary classifications. However, a theoretical defect lies in this type of methods: the ignorance of possible contradictions in individual ranking results. In this paper, we partially embrace the decomposition idea and propose the Deep and Ordinal Ensemble Learning with Two Groups Classification (DOEL2groups) for age prediction. An important advantage of our approach is that it theoretically allows the prediction even when the contradictory cases occur. The proposed method is characterized by a deep and ordinal ensemble and a two-stage aggregation strategy. Specifically, we first set up the ensemble based on Convolutional Neural Network (CNN) techniques, while the ordinal relationship is implicitly constructed among its base learners. Each base learner will classify the target face into one of two specific age groups. After achieving probability predictions of different age groups, then we make aggregation by transforming them into counting value distributions of whole age classes and getting the final age estimation from their votes. Moreover, to further improve the estimation performance, we suggest to regard the age class at the boundary of original two age groups as another age group and this modified version is named the Deep and Ordinal Ensemble Learning with Three Groups Classification (DOEL3groups). Effectiveness of this new grouping scheme is validated in theory and practice. Finally, we evaluate the proposed two ensemble methods on controlled and wild aging databases, and both of them produce competitive results. Note that the DOEL3groupsshows the state-of-the-art performance in most cases.
Jiucheng Xie, Chi-Man Pun
IEEE Trans. Inf. Forensics Secur.2
2020 An End-to-End Dense-InceptionNet for Image Copy-Move Forgery Detection
abstract
A novel image copy-move forgery detection scheme using a Dense-InceptionNet is proposed in this paper. Dense-InceptionNet is an end-to-end, multi-dimensional dense-feature connection, Deep Neural Network (DNN). It is the first DNN model to autonomously learn the feature correlations and search the possible forgery snippets through the matching clues. The proposed Dense-InceptionNet consists of Pyramid Feature Extractor (PFE), Feature Correlation Matching (FCM), and Hierarchical Post-Processing (HPP) modules. The PFE module is proposed to extract multi-dimensional and multi-scale dense-features. The features of each layer in this extractor module are directly connected to the preceding layers. The FCM module is proposed to learn the high correlations of deep features and obtain three candidate matching maps. Finally, the HPP module which makes use of three matching maps to obtain a combination of cross-entropies is amenable to better training via backpropagation. Experiments demonstrate that the efficiency of the proposed Dense-InceptionNet is much better than the other state-of-the-art methods while achieving the relative best performance against most known attacks.
Chi-Man Pun
IEEE Trans. Inf. Forensics Secur.2
2020 Improving the Harmony of the Composite Image by Spatial-Separated Attention Module
abstract
Image composition is one of the most important applications in image processing. However, the inharmonious appearance between the spliced region and background degrade the quality of the image. Thus, we address the problem of Image Harmonization: Given a spliced image and the mask of the spliced region, we try to harmonize the "style" of the pasted region with the background (non-spliced region). Previous approaches have been focusing on learning directly by the neural network. In this work, we start from an empirical observation: the differences can only be found in the spliced region between the spliced image and the harmonized result while they share the same semantic information and the appearance in the nonspliced region. Thus, in order to learn the feature map in the masked region and the others individually, we propose a novel attention module named Spatial-Separated Attention Module (S2AM). Furthermore, we design a novel image harmonization framework by inserting the S2AM in the coarser low-level features of the Unet structure by two different ways. Besides image harmonization, we make a big step for harmonizing the composite image without the specific mask under previous observation. The experiments show that the proposed S2AM performs better than other state-of-the-art attention modules in our task. Moreover, we demonstrate the advantages of our model against other state-of-the-art image harmonization methods via criteria from multiple points of view.
Xiaodong Cun, Chi-Man Pun
IEEE Trans. Image Process.2
2020 Multiscale Superpixel-Based Hyperspectral Image Classification Using Recurrent Neural Networks With Stacked Autoencoders
abstract
This paper develops a novel hyperspectral image (HSI) classification framework by exploiting the spectral-spatial features of multiscale superpixels via recurrent neural networks with stacked autoencoders. The superpixels can be used to segment an HSI into shape-adaptive regions, and multiscale superpixels can capture the object information more accurately. Therefore, the superpixel-based classification methods have been studied by many researchers. In this paper, we propose a multiscale superpixel-based classification method. In contrast to current research, the proposed method not only captures the features of each scale but also considers the correlation among different scales via recurrent neural networks. In this way, the spectral-spatial information within a superpixel is more efficiently exploited. In this paper, we first segment the HSI from coarse to fine scales using the superpixels. Then, the spatial features within each superpixel and among superpixels are sufficiently exploited by the local and nonlocal similarity measure. Finally, recurrent neural networks with stacked autoencoders are proposed to learn the high-level multiscale spectral-spatial features. Experiments are conducted on real HSI datasets. The results demonstrate the superiority of the proposed method over several well-known methods in both visual appearance and classification accuracy.
Cheng Shi 0002, Chi-Man Pun
IEEE Trans. Multim.2
2019 Audio Replay Spoof Attack Detection Using Segment-based Hybrid Feature and DenseNet-LSTM Network
abstract
At present, most automatic speaker verification (ASV) systems are vulnerable to replay spoof attacks. Therefore, this paper proposes a new approach for the detection of audio replay spoof attacks. Here, a segment-based hybrid feature extraction method is used, which includes the Mel-frequency cepstral coefficient (MFCC) features and Constant-Q cepstral coefficients (CQCC) features. Then, hybrid features are trained using a variety of deep learning networks, including DenseNet, LSTM, and DenseNet-LSTM hybrid architectures. Experiments using the DenseNet-LSTM model with mixed features framework achieves the best performance. Compared to the baseline system built on the CQCC and Gaussian mixture model (GMM), the proposed method achieved 64.31% relative improvement.
Lian Huang, Chi-Man Pun
ICASSP2
2019 A 6-DOF Telexistence Drone Controlled by a Head Mounted Display
abstract
Recently, a new form of telexistence is achieved by recording images with cameras on an unmanned aerial vehicle (UAV) and displaying them to the user via a head mounted display (HMD). A key problem here is how to provide a free and natural mechanism for the user to control the viewpoint and watch a scene. To this end, we propose an improved rate-control method with an adaptive origin update (AOU) scheme. Without the aid of any auxiliary equipment, our scheme handles the self-centering problem. In addition, we present a full 6-DOF viewpoint control method to manipulate the motion of a stereo camera, and we build a real prototype to realize this by utilizing a pan-tilt-zoom (PTZ) which not only provides 2-DOF to the camera but also compensates the jittering motion of the UAV to record more stable image streams.
Xingyu Xia, Chi-Man Pun, Yang Yang 0002, Huimin Lu 0001, Hao Gao 0005, Feng Xu 0005
VR2
2019 Adaptive multi-scale deep neural networks with perceptual loss for panchromatic and multispectral images classification
Cheng Shi 0002, Chi-Man Pun
Inf. Sci.2
2019 Copy-move forgery detection using adaptive keypoint filtering and iterative region merging
Chi-Man Pun
Multim. Tools Appl.2
2019 Automatic Medical Image Registration Based on an Integrated Method Combining Feature and Area Information
Jiucheng Xie, Chi-Man Pun, Zhaoqing Pan, Hao Gao 0005, Baoyun Wang
Neural Process. Lett.2
2019 Reversible image reconstruction for reversible data hiding in encrypted images
Chi-Man Pun
Signal Process.2
2019 Chronological Age Estimation Under the Guidance of Age-Related Facial Attributes
abstract
Although the researches of facial attributes' analysis have been launched for decades, the estimation of chronological age attribute remains a big challenge. Previous researchers have found that some facial attributes (e.g., gender and race attributes) have close connections with the age attribute and make age estimation under a specific condition decided by various combinations of those age-related attributes which should be more reasonable. In this paper, we propose a generic framework based on a convolutional neural network, which can consider different conditions for age estimation and jointly output age and age-related facial attributes in the end. Compared with conventional methods, it is more efficient and universal. Besides, we view age estimation as a special multi-class ordinal classification problem and use a losses combination function to optimize the predicted probability distribution of individual age classes. These operations further improve the performance of age estimation. Finally, the proposed method achieves state-of-the-art results on both controlled and wild face datasets.
Jiucheng Xie, Chi-Man Pun
IEEE Trans. Inf. Forensics Secur.2
2019 An Improved Artificial Bee Colony Algorithm With its Application
abstract
The artificial bee colony is a popular evolutionary algorithm that exhibits strong exploration ability but slow convergence. This paper proposes two new updating equations to boost the performances of employed and onlooker bees, respectively. In the new updating equations, two intelligent learning strategies give bees a chance to learn from individuals with better performances. New control operators are also utilized to balance global and local searches. Second, we define a new search direction mechanism to overcome the oscillation phenomenon in employed bees. Finally, an intelligent learning mechanism is proposed to accelerate the convergence rate of the worst employed bee. To test the effectiveness of our algorithm, a series of benchmark functions and two industrial problems are utilized. Experimental results demonstrate that our proposed algorithm performs more favorably on both theoretical and practical problems.
Hao Gao 0005, Yujiao Shi 0002, Chi-Man Pun, Sam Kwong
IEEE Trans. Ind. Informatics3
2018 Perceptual Loss for Superpixel-Level Multispectral and Panchromatic Image Classification
abstract
Convolutional neural networks (CNNs) have proven to be an effective way for deep feature extraction. However, multispectral and panchromatic images are susceptible to illumination unevenness and noise, and the default cross entropy loss function consider only the local information, resulting in misclassification. In this paper, we propose a novel super-pixel-level deep neural networks for multispectral and panchromatic images classification, and define a novel percep-tualloss function via non-local spectral and structure similarity to suppress the interference of unbalanced light and noise. We also propose the corresponding iteration optimization algorithm in this paper. Experimental results show that the proposed method performs better than the state-of-the-art methods.
Cheng Shi 0002, Chi-Man Pun
ICASSP2
2018 Multi-scale hierarchical recurrent neural networks for hyperspectral image classification
Cheng Shi 0002, Chi-Man Pun
Neurocomputing2
2018 Reversible data-hiding in encrypted images by redundant space transfer
Chi-Man Pun
Inf. Sci.2
2018 A two-stage localization for copy-move forgery detection
Chi-Man Pun, Jim-Lee Chung
Inf. Sci.1
2018 Multi-scale feature extraction and adaptive matching for copy-move forgery detection
Xiuli Bi, Chi-Man Pun, Xiaochen Yuan
Multim. Tools Appl.2
2018 On-line video multi-object segmentation based on skeleton model and occlusion detection
Guoheng Huang, Chi-Man Pun
Multim. Tools Appl.2
2018 Robust image hashing using progressive feature selection for tampering detection
Chi-Man Pun, Cai-Ping Yan, Xiaochen Yuan
Multim. Tools Appl.1
2018 Object tracking using distribution fields with correlation coefficients
Peng Qin 0001, Chi-Man Pun
Multim. Tools Appl.2
2018 Fast copy-move forgery detection using local bidirectional coherency error refinement
Xiuli Bi, Chi-Man Pun
Pattern Recognit.2
2018 Superpixel-based 3D deep neural networks for hyperspectral image classification
Cheng Shi 0002, Chi-Man Pun
Pattern Recognit.2
2018 Applying stochastic second-order entropy images to multi-modal image registration
Xiaodong Cun, Chi-Man Pun, Hao Gao 0005
Signal Process. Image Commun.2
2018 Locating splicing forgery by fully convolutional networks and conditional random field
Bo Liu 0047, Chi-Man Pun
Signal Process. Image Commun.2
2017 Fast reflective offset-guided searching method for copy-move forgery detection
Xiuli Bi, Chi-Man Pun
Inf. Sci.2
2017 3D multi-resolution wavelet convolutional neural networks for hyperspectral image classification
Cheng Shi 0002, Chi-Man Pun
Inf. Sci.2
2017 Unsupervised video co-segmentation based on superpixel co-saliency and region merging
Guoheng Huang, Chi-Man Pun, Cong Lin 0001
Multim. Tools Appl.2
2017 Highly non-rigid video object tracking using segment-based object candidates
Cong Lin 0001, Chi-Man Pun, Guoheng Huang
Multim. Tools Appl.2
2017 Efficient shape classification using region descriptors
Cong Lin 0001, Chi-Man Pun, Chi-Man Vong, Donald A. Adjeroh
Multim. Tools Appl.2
2017 Post-boosting of classification boundary for imbalanced data using geometric mean
Jie Du 0001, Chi-Man Vong, Chi-Man Pun, Pak-Kin Wong 0001, Weng-Fai Ip
Neural Networks3
2017 Weighted Joint Sparse Representation for Removing Mixed Noise in Image
abstract
Joint sparse representation (JSR) has shown great potential in various image processing and computer vision tasks. Nevertheless, the conventional JSR is fragile to outliers. In this paper, we propose a weighted JSR (WJSR) model to simultaneously encode a set of data samples that are drawn from the same subspace but corrupted with noise and outliers. Our model is desirable to exploit the common information shared by these data samples while reducing the influence of outliers. To solve the WJSR model, we further introduce a greedy algorithm called weighted simultaneous orthogonal matching pursuit to efficiently approximate the global optimal solution. Then, we apply the WJSR for mixed noise removal by jointly coding the grouped nonlocal similar image patches. The denoising performance is further improved by incorporating it with the global prior and the sparse errors into a unified framework. Experimental results show that our denoising method is superior to several state-of-the-art mixed noise removal methods.
Licheng Liu, Long Chen 0001, C. L. Philip Chen, Yuan Yan Tang, Chi-Man Pun
IEEE Trans. Cybern.5
2017 Image Alignment-Based Multi-Region Matching for Object-Level Tampering Detection
abstract
Tampering detection methods based on image hashing have been widely studied with continuous advancements. However, most existing models cannot generate object-level tampering localization results, because the forensic hashes attached to the image lack contour information. In this paper, we present a novel tampering detection model that can generate an accurate, object-level tampering localization result. First, an adaptive image segmentation method is proposed to segment the image into closed regions based on strong edges. Then, the color and position features of the closed regions are extracted as a forensic hash. Furthermore, a geometric invariant tampering localization model named image alignment-based multi-region matching (IAMRM) is proposed to establish the region correspondence between the received and forensic images by exploiting their intrinsic structure information. The model estimates the parameters of geometric transformations via a robust image alignment method based on triangle similarity; in addition, it matches multiple regions simultaneously by utilizing manifold ranking based on different graph structures and features. Experimental results demonstrate that the proposed IAMRM is a promising method for object-level tampering detection compared with the state-of-the-art methods.
Chi-Man Pun, Cai-Ping Yan, Xiaochen Yuan
IEEE Trans. Inf. Forensics Secur.1
2017 Multi-Scale Difference Map Fusion for Tamper Localization Using Binary Ranking Hashing
abstract
The block-based analysis for tamper localization is a prevailing mechanism in hash-based forgery detection algorithm. One of the main problems with a block-based analysis is its rough localization stemming from the demand to use relatively large blocks to reduce hash length. While decreasing the block size can improve the localization resolution, the hash length tends to become too long to be practical. In this paper, we propose a binary ranking hashing approach that satisfies both the requirements of compact hash length and small block size, to obtain a binary map combined with spatial information. Meanwhile, we investigate a multiscale difference map fusion approach that fuses multiple candidate difference maps, resulting from the analysis of the subtraction between two binary maps with different sliding windows, to obtain a single, more reliable tampering map with better localization resolution. We use manifold ranking to model this multiscale difference map fusion problem and propose a two-stage scheme, namely, ranking with tampering queries and nontampering queries. Our results indicate that the proposed tamper detection method can improve the tamper localization resolution compared with state-of-the-art methods.
Cai-Ping Yan, Chi-Man Pun
IEEE Trans. Inf. Forensics Secur.2
2017 Structure-Regularized Compressive Tracking With Online Data-Driven Sampling
abstract
Being a powerful appearance model, compressive random projection derives effective Haar-like features from non-rotated 4-D-parameterized rectangles, thus supporting fast and reliable object tracking. In this paper, we show that such successful fast compressive tracking scheme can be further significantly improved by structural regularization and online data-driven sampling. Our major contribution is threefold. First, we find that superpixel-guided compressive projection can generate more discriminative features by sufficiently capturing rich local structural information of images. Second, we propose fast directional integration that enables low-cost extraction of feasible Haar-like features from arbitrarily rotated 5-D-parameterized rectangles to realize more accurate object localization. Third, beyond naive dense uniform sampling, we present two practical online data-driven sampling strategies to produce less yet more effective candidate and training samples for object detection and classifier updating, respectively. Extensive experiments on real-world benchmark data sets validate the superior performance, i.e., much better object localization ability and robustness, of the proposed approach over state-of-the-art trackers.
Qing Guo 0005, Wei Feng 0005, Ce Zhou, Chi-Man Pun
IEEE Trans. Image Process.4
2016 Multi-Level Dense Descriptor and Hierarchical Feature Matching for Copy-Move Forgery Detection
Xiuli Bi, Chi-Man Pun, Xiaochen Yuan
Inf. Sci.2
2016 An efficient image segmentation method based on a hybrid particle swarm algorithm with learning strategy
Hao Gao 0005, Chi-Man Pun, Sam Kwong
Inf. Sci.2
2016 On-line video object segmentation using illumination-invariant color-texture feature extraction and marker prediction
Chi-Man Pun, Guoheng Huang
J. Vis. Commun. Image Represent.1
2016 A real-time detector for parked vehicles based on hybrid background modeling
Chi-Man Pun, Cong Lin 0001
J. Vis. Commun. Image Represent.1
2016 Multi-scale noise estimation for image splicing forgery detection
Chi-Man Pun, Bo Liu 0047, Xiaochen Yuan
J. Vis. Commun. Image Represent.1
2016 An improved artificial bee colony and its application
Yujiao Shi 0002, Chi-Man Pun, Haidong Hu, Hao Gao 0005
Knowl. Based Syst.2
2016 Robust lossless digital watermarking using integer transform with Bit plane manipulation
Ka-Cheng Choi, Chi-Man Pun
Multim. Tools Appl.2
2016 Non-rigid visual object tracking using user-defined marker and Gaussian kernel
Guoheng Huang, Chi-Man Pun, Cong Lin 0001, Yicong Zhou
Multim. Tools Appl.2
2016 Multi-scale image hashing using adaptive local feature extraction for robust tampering detection
Cai-Ping Yan, Chi-Man Pun, Xiaochen Yuan
Signal Process.2
2016 Quaternion-Based Image Hashing for Adaptive Tampering Localization
abstract
Image-hashing-based tampering detection methods have been widely studied with continuous advancements. However, most of existing models are designed for a specific tampering. In this paper, we propose a novel quaternion-based image hashing to detect almost all types of tampering, including color changing, copy move, splicing, and so on. First, the quaternion Fourier-Mellin transform is used to calculate the geometric hash to eliminate the influence of geometric distortions. Then, a new quaternion image construction method, which combines advantages of both color and structural features, is proposed to implement the quaternion Fourier transform to calculate the image feature hash to locate the tampered regions. The objective is to provide a reasonably short image hashing with good performance, i.e., being perceptually robust against various content-preserving attacks while capable of detecting and locating almost all types of tampering. Furthermore, an adaptive tampering localization algorithm is proposed based on clustering analysis to improve the detection accuracy. The experimental results show that the proposed tampering detection model outperforms the existing state-of-the-art models and is very robust against various content-preserving attacks.
Cai-Ping Yan, Chi-Man Pun, Xiaochen Yuan
IEEE Trans. Inf. Forensics Secur.2
2015 Video Object Tracking Using Interactive Segmentation and Superpixel Based Gaussian Kernel
abstract
A novel non-rigid video object tracking based on interactive segmentation and super pixel Gaussian kernel is proposed in this paper. In the initialization stage, instead of using the traditional bounding box to locate the targeted object, we employed an interactive segmentation with user-defined marker to segment the object accurately in the first frame of the input video to avoid the background influence in the traditional bounding box. During the tracking stage, using a Gaussian kernel as movement constraint, each super pixel is tracked independently to locate the object in the next frame. Experimental results show that the proposed method compared to state of the art methods can achieve better robustness and accuracy for various challenging video clips.
Guoheng Huang, Chi-Man Pun, Cong Lin 0001
IV2
2015 2D Sine Logistic modulation map for image encryption
Zhongyun Hua, Yicong Zhou, Chi-Man Pun, C. L. Philip Chen
Inf. Sci.3
2015 Robust Mel-Frequency Cepstral coefficients feature detection and dual-tree complex wavelet transform for digital audio watermarking
Xiaochen Yuan, Chi-Man Pun, C. L. Philip Chen
Inf. Sci.2
2015 Application of a generalized difference expansion based reversible audio data hiding algorithm
Ka-Cheng Choi, Chi-Man Pun, C. L. Philip Chen
Multim. Tools Appl.2
2015 Histogram modification based image watermarking resistant to geometric distortions
Chi-Man Pun, Xiaochen Yuan
Multim. Tools Appl.1
2015 Fast and accurate face detection by sparse Bayesian extreme learning machine
Chi-Man Vong, Keng Iam Tai, Chi-Man Pun, Pak-Kin Wong 0001
Neural Comput. Appl.3
2015 Cascade Chaotic System With Applications
abstract
Chaotic maps are widely used in different applications. Motivated by the cascade structure in electronic circuits, this paper introduces a general chaotic framework called the cascade chaotic system (CCS). Using two 1-D chaotic maps as seed maps, CCS is able to generate a huge number of new chaotic maps. Examples and evaluations show the CCS's robustness. Compared with corresponding seed maps, newly generated chaotic maps are more unpredictable and have better chaotic performance, more parameters, and complex chaotic properties. To investigate applications of CCS, we introduce a pseudo-random number generator (PRNG) and a data encryption system using a chaotic map generated by CCS. Simulation and analysis demonstrate that the proposed PRNG has high quality of randomness and that the data encryption system is able to protect different types of data with a high-security level.
Yicong Zhou, Zhongyun Hua, Chi-Man Pun, C. L. Philip Chen
IEEE Trans. Cybern.3
2015 Image Forgery Detection Using Adaptive Oversegmentation and Feature Point Matching
abstract
A novel copy-move forgery detection scheme using adaptive oversegmentation and feature point matching is proposed in this paper. The proposed scheme integrates both block-based and keypoint-based forgery detection methods. First, the proposed adaptive oversegmentation algorithm segments the host image into nonoverlapping and irregular blocks adaptively. Then, the feature points are extracted from each block as block features, and the block features are matched with one another to locate the labeled feature points; this procedure can approximately indicate the suspected forgery regions. To detect the forgery regions more accurately, we propose the forgery region extraction algorithm, which replaces the feature points with small superpixels as feature blocks and then merges the neighboring blocks that have similar local color features into the feature blocks to generate the merged regions. Finally, it applies the morphological operation to the merged regions to generate the detected forgery regions. The experimental results indicate that the proposed copy-move forgery detection scheme can achieve much better detection results even under various challenging conditions compared with the existing state-of-the-art copy-move forgery detection methods.
Chi-Man Pun, Xiaochen Yuan, Xiuli Bi
IEEE Trans. Inf. Forensics Secur.1
2014 Intrinsic image decomposition by hierarchical L0 sparsity
abstract
This paper presents a hierarchical approach to single image intrinsic decomposition based on non-local L0sparsity. In contrast to previous studies using heuristic methods to well-define the ill-posed problem, our approach is able to effectively construct sparse, non-local and multiscale reflectance dependencies in an unsupervised manner, thus is less dependent on the chromaticity feature and more accurately captures the global reflectance correlations. Besides, we impose homogenous smoothness prior and scale constraint in our model to further improve the decomposition accuracy. We formulate the decomposition as a quadratic minimization problem, which can be efficiently solved in closed form. Extensive experiments show that our approach can successfully extract the shading and reflectance components from a single image, and outperforms state-of-the-art methods on benchmark dataset. Besides, our approach can achieve comparable results with user-assisted methods on natural scenes.
Xuecheng Nie, Wei Feng 0005, Haipeng Dai 0002, Chi-Man Pun
ICME5
2014 Bag of squares: A reliable model of measuring superpixel similarity
abstract
As the increasing popularity of superpixel-based applications, measuring superpixel-level similarity becomes an important and commonly required problem. In this paper, we propose a general bag of squares (BoS) model for such particular purpose. Compared to existing methods, our approach provides a full scheme to both invariantly represent superpixels and accurately measure their pairwise similarities. In order to handle the split-and-merge variety of superpixels of same objects in different scenes, our model is based on superpixel pyramid. As a result, the BoS model of a superpixel is built upon a group of subregions consisting of the superpixel itself and its children subregions in the pyramid. For each subregion, we extract a proper number of maximum squares via distance transform, and then use a fast self-validated approach to clustering them into a small number of dominant squares, which together with a rotation and scale invariant square descriptor, jointly compose the BoS model for the particular superpixel. Finally, we measure the similarity between a pair of superpixels by the closeness of their BoS models. Experiments on interactive object segmentation and co-saliency detection show that the proposed BoS model can reliably capture the delicate differences among superpixels, thus always producing better segmentation results, especially for segmenting highly variant objects in clutter scenes.
Wei Feng 0005, Jiawan Zhang, Chi-Man Pun
ICME4
2014 Image encryption using 2D Logistic-Sine chaotic map
abstract
This paper introduces a new two-dimensional Logistic-Sine map (2D-LSM). It has excellent chaotic performance and its outputs are difficult to predict. Using 2D-LSM, this paper proposes a new image encryption algorithm. Simulation results and security analysis demonstrate that the proposed algorithm is able to protect different kinds of images with a high security level.
Zhongyun Hua, Yicong Zhou, Chi-Man Pun, C. L. Philip Chen
SMC3
2014 A new reversible data hiding algorithm in the encryption domain
abstract
This paper introduces a new reversible data hiding algorithm in the encryption domain. It integrates data hiding into the image encryption process to achieve different level of access right and security. Computer simulations and comparisons demonstrate that the proposed algorithm can withstand the differential attack and outperforms other existing methods in terms of security and the message embedding capacity that is 52% larger than the state-of-the-art method in the best scenario. The marked decrypted images of our proposed method show the best visual quality according to the PSNR results.
Yicong Zhou, Chi-Man Pun, C. L. Philip Chen
SMC3
2014 Feature extraction and local Zernike moments based geometric invariant watermarking
Xiaochen Yuan, Chi-Man Pun
Multim. Tools Appl.2
2014 Nonnegative class-specific entropy component analysis with adaptive step search criterion
Chi-Man Pun, Yuan Yan Tang
Pattern Anal. Appl.2
2013 Image co-saliency detection by propagating superpixel affinities
abstract
Image co-saliency detection is a valuable technique to highlight perceptually salient regions in image pairs. In this paper, we propose a self-contained co-saliency detection algorithm based on superpixel affinity matrix. We first compute both intra and inter similarities of superpixels of image pairs. Bipartite graph matching is applied to determine most reliable inter similarities. To update the similarity score between every two superpixels, we next employ a GPU-based all-pair SimRank algorithm to do propagation on the affinity matrix. Based on the inter superpixel affinities we derive a co-saliency measure that evaluates the foreground cohesiveness and locality compactness of superpixels within one image. The effectiveness of our method is demonstrated in experimental evaluation.
Zhiyu Tan, Wei Feng 0005, Chi-Man Pun
ICASSP4
2013 A spectral-multiplicity-tolerant approach to robust graph matching
Wei Feng 0005, Chi-Man Pun, Jianmin Jiang
Pattern Recognit.4
2013 Geometric invariant watermarking by local Zernike moments of binary image patches
Xiaochen Yuan, Chi-Man Pun, C. L. Philip Chen
Signal Process.2
2013 Robust Segments Detector for De-Synchronization Resilient Audio Watermarking
abstract
A robust feature points detector for invariant audio watermarking is proposed in this paper. The audio segments centering at the detected feature points are extracted for both watermark embedding and extraction. These feature points are invariant to various attacks and will not be changed much for maintaining high auditory quality. Besides, high robustness and inaudibility can be achieved by embedding the watermark into the approximation coefficients of Stationary Wavelet Transform (SWT) domain, which is shift invariant. The spread spectrum communication technique is adopted to embed the watermark. Experimental results show that the proposed Robust Audio Segments Extractor (RASE) and the watermarking scheme are not only robust against common audio signal processing, such as low-pass filtering, MP3 compression, echo addition, volume change, and normalization; and distortions introduced in Stir-mark benchmark for Audio; but also robust against synchronization geometric distortions simultaneously, such as resample time-scale modification (TSM) with scaling factors up to ±50%, pitch invariant TSM by ±50%, and tempo invariant pitch shifting by ±50%. In general, the proposed scheme can well resist various attacks by the joint RASE and SWT approach, which performs much better comparing with the existing state-of-the art methods.
Chi-Man Pun, Xiaochen Yuan
IEEE Trans. Speech Audio Process.1
2011 A Robust Anti-tamper Protection Scheme
abstract
This paper proposes a robust anti-tamper protection scheme to protect any critical regions of a program from being modified, using possibly a large number of lightweight protection units, called protectors, installed among the program code. A protector would cause an incorrect execution if the code protected by it has been tempered. The protectors are organized in the form of a protection tree. The root node is a critical region, and other nodes are protectors. The protection scheme also supports non-deterministic execution of functions. Modifying any critical region in the protected program has been shown to require an exponential time. Experiment results show that the proposed scheme would not increase noticeably the program execution time.
Hing-Chung Tsang, Moon-Chuen Lee, Chi-Man Pun
ARES3
2011 Adaptive Client-Side LUT-Based Digital Watermarking
abstract
In this paper we proposed an adaptive secure client-side Look-Up-Table (LUT) based digital audio watermarking, which adaptively choose an optimal embedding strength for individual segment of host signal according to specific properties of each segment. It is clear that different segments of host signal featured different characteristics, so using the embedding strength adaptively would definitely improve the inaudibility of watermarking while maintaining the robustness level nearly unchanged. Consequently, the proposed approach features the combination properties of both security property derived from Secure Client-side Embedding and better robustness and inaudibility from the adaptive embedding strength of individual segment. Simulation results show that the watermark in this scheme can survive robustly after being subjected to various common audio attacks, such as filtering, resample and so on.
Chi-Man Pun, Jing-Jing Jiang, C. L. Philip Chen
TrustCom1
2011 Geometric Invariant Digital Image Watermarking Scheme Based on Feature Points Detector and Histogram Distribution
abstract
A robust and geometric invariant digital image watermarking scheme based on SIFT Based Feature Points Detector (SIFTFPD) and histogram distribution is proposed in this paper. The SIFTFPD is proposed to extract geometric invariant feature points from the host image for watermark embedding; and the descriptor is generated subsequently. With the feature extraction procedure, the circular regions centered at the extracted feature points and with the given radius are defined as embedding regions. For watermark embedding, some pixels are moved to form a specific pattern in the intensity-level histogram distribution in each embedding region, to indicate the watermark. For watermark extraction, the embedded regions are generated with the descriptor and according to the intensity-level histogram distribution in each region, the watermark can be extracted. Experimental results show that the proposed scheme is very robust against geometric distortion such as rotation, scaling, cropping, and affine transformation; and common signal processing, such as JPEG compression, median filtering, and Gaussian low-pass filtering.
Chi-Man Pun, Xiaochen Yuan, C. L. Philip Chen
TrustCom1
2011 Kernel-view based discriminant approach for embedded feature extraction in high-dimensional space
Bin Fang 0001, Chi-Man Pun, Yuan Yan Tang
Neurocomputing3
2009 Complex Zernike Moments Features for Shape-Based Image Retrieval
abstract
Shape is a fundamental image feature used in content-based image-retrieval systems. This paper proposes a robust and effective shape feature, which is based on a set of orthogonal complex moments of images known as Zernike moments (ZMs). As the rotation of an image has an impact on the ZM phase coefficients of the image, existing proposals normally use magnitude-only ZM as the image feature. In this paper, we compare, by using a mathematical form of analysis, the amount of visual information captured by ZM phase and the amount captured by ZM magnitude. This analysis shows that the ZM phase captures significant information for image reconstruction. We therefore propose combining both the magnitude and phase coefficients to form a new shape descriptor, referred to as invariant ZM descriptor (IZMD). The scale and translation invariance of IZMD could be obtained by prenormalizing the image using the geometrical moments. To make the phase invariant to rotation, we perform a phase correction while extracting the IZMD features. Experiment results show that the proposed shape feature is, in general, robust to changes caused by image shape rotation, translation, and/or scaling. The proposed IZMD feature also outperforms the commonly used magnitude-only ZMD in terms of noise robustness and object discriminability.
Shan Li 0003, Moon-Chuen Lee, Chi-Man Pun
IEEE Trans. Syst. Man Cybern. Part A3
2004 Extraction of Shift Invariant Wavelet Features for Classification of Images with Different Sizes
abstract
An effective shift invariant wavelet feature extraction method for classification of images with different sizes is proposed. The feature extraction process involves a normalization followed by an adaptive shift invariant wavelet packet transform. An energy signature is computed for each subband of these invariant wavelet coefficients. A reduced subset of energy signatures is selected as the feature vector for classification of images with different sizes. Experimental results show that the proposed method can achieve high classification accuracy of 98.5 percent and outperforms the other two image classification methods.
Chi-Man Pun, Moon-Chuen Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Invariant content-based image retrieval by wavelet energy signatures
abstract
An effective rotation and scale invariant log-polar wavelet texture feature for image retrieval was proposed. The feature extraction process involves a log-polar transform followed by an adaptive row shift invariant wavelet packet transform. The log-polar transform converts a given image into a rotation and scale invariant but row-shifted image, which is then passed to the adaptive row shift invariant wavelet packet transform to generate adaptively some subbands of rotation and scale invariant wavelet coefficients with respect to an information cost function. An energy signature is computed for each subband of these wavelet coefficients. In order to reduce feature dimensionality, only the most dominant log-polar wavelet energy signatures are selected as feature vector for image retrieval. The whole feature extraction process is quite efficient and involves only O(n/spl middot/log n) complexity. Experimental results show that this rotation and scale invariant texture feature is effective and outperforms the traditional wavelet packet signatures.
Chi-Man Pun
ICASSP (3)1
2003 Invariant content-based image retrieval by wavelet energy signatures
abstract
An effective rotation and scale invariant log-polar wavelet texture feature for image retrieval was proposed. The feature extraction process involves a log-polar transform followed by an adaptive row shift invariant wavelet packet transform. The log-polar transform converts a given image into a rotation and scale invariant but row- shifted image, which is then passed to the adaptive row shift invariant wavelet packet transform to generate adaptively some subbands of rotation and scale invariant wavelet coefficients with respect to an information cost function. An energy signature is computed for each subband of these wavelet coefficients. In order to reduce feature dimensionality, only the most dominant log-polar wavelet energy signatures are selected as feature vector for image retrieval. The whole feature extraction process is quite efficient and involves only O (n/spl middot/log n) complexity. Experimental results show that this rotation and scale invariant texture feature is effective and outperforms the traditional wavelet packet signatures.
Chi-Man Pun
ICME1
2003 Rotation-invariant texture feature for image retrieval
Chi-Man Pun
Comput. Vis. Image Underst.1
2003 Rotation and Scale Invariant Wavelet Feature for Content-based Texture Image Retrieval
abstract
Abstract This article introduces an effective rotation and scale invariant log‐polar wavelet texture feature for image retrieval. The proposed feature is an attempt to enhance the existing content‐based image retrieval systems that largely present difficulty in coping with images with changes in orientations and scales. The underlying feature extraction process involves a log‐polar transform followed by an adaptive row shift invariant wavelet packet transform. The log‐polar transform converts a given image into a rotation and scale invariant but row‐shifted image, which is then further processed through an adaptive row‐shift invariant wavelet packet transform operation to generate adaptively selected subbands of rotation and scale invariant wavelet coefficients, based on an information cost function. An energy signature is computed for each subband of these wavelet coefficients. To reduce feature dimensionality, only the most dominant log‐polar wavelet energy signatures are selected for the feature vector for image retrieval. The overall feature extraction process is quite efficient and involves only O(n · log n) complexity. Experimental results show that this rotation and scale invariant wavelet feature is quite effective for image retrieval and outperforms the traditional wavelet packet signatures.
Moon-Chuen Lee, Chi-Man Pun
J. Assoc. Inf. Sci. Technol.2
2003 Log-Polar Wavelet Energy Signatures for Rotation and Scale Invariant Texture Classification
abstract
Classification of texture images is important in image analysis and classification. This paper proposes an effective scheme for rotation and scale invariant texture classification using log-polar wavelet signatures. The rotation and scale invariant feature extraction for a given image involves applying a log-polar transform to eliminate the rotation and scale effects, but at same time produce a row shifted log-polar image, which is then passed to an adaptive row shift invariant wavelet packet transform to eliminate the row shift effects. So, the output wavelet coefficients are rotation and scale invariant. The adaptive row shift invariant wavelet packet transform is quite efficient with only O(n /spl middot/ log n) complexity. A feature vector of the most dominant log-polar wavelet energy signatures extracted from each subband of wavelet coefficients is constructed for rotation and scale invariant texture classification. In the experiments, we employed a Mahalanobis classifier to classify a set of 25 distinct natural textures selected from the Brodatz album. The experimental results, based on different testing data sets for images with different orientations and scales, show that the proposed classification scheme using log-polar wavelet signatures outperforms two other texture classification methods, its overall accuracy rate for joint rotation and scale invariance being 90.8 percent, demonstrating that the extracted energy signatures are effective rotation and scale invariant features. Concerning its robustness to noise, the classification scheme also performs better than the other methods.
Chi-Man Pun, Moon-Chuen Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Rotation invariant texture feature for content based image retrieval
abstract
An effective rotation invariant polar-wavelet texture feature for content based image retrieval is proposed. The feature extraction process involves a polar transform followed by an adaptive row shift invariant wavelet packet transform. The polar transform converts a given image into a rotation-invariant but row-shifted image, which is then passed to the adaptive row shift invariant wavelet packet transform to generate adaptively some subbands of rotation invariant wavelet coefficients with respect to an information cost function. An energy signature is computed for each sub-band of these wavelet coefficients. In order to reduce feature dimensionality, only the most dominant polar-wavelet energy signatures are selected as feature vectors for image retrieval. The whole feature extraction process is quite efficient and involves only O(n/spl middot/log n) complexity. Experimental results show that this rotation-invariant texture feature is effective and outperforms traditional wavelet packet signatures.
Chi-Man Pun, Moon-Chuen Lee
ICME (1)1
2000 Classifying rotated textures using wavelet packet signatures
abstract
This paper proposes a novel approach to the classification of rotated texture images. The proposed classification method involves decomposing a texture image with a family of real orthonormal wavelet bases for different levels, computing the wavelet packet coefficients, and computing the energy signatures using the wavelet packet coefficients. Such energy signatures are sorted and used as a feature for texture image classification. We employ a Mahalanobis distance classifier to classify a set of twenty distinct natural textures selected from the Brodatz album. Experimental results, based on a large sample data set having different orientations, show that the proposed method outperforms other methods which may perform well in the classification of texture images having the same orientation.
Moon-Chuen Lee, Chi-Man Pun
ICASSP2
1999 Portuguese-Chinese machine translation in Macao
abstract
There have been substantial changes in computing practices in the cyberspace, mainly as a result of the proliferation of low priced under-utilized powerfully heterogeneous computers are connected by high-speed links. In this paper we reminisce the vicissitude of computing platform and introduce our Portuguese-Chinese corpus-based machine translation (CBMT) system which employs a statistical approach with automatic bilingual alignment support. Our improved algorithm for aligning bilingual parallel texts can achieve 97% of accuracy. At the same time, we broach the “distributed translation computing” concept to construct a uniform distributed shared-object technical term retrieving workstation and achieve high computing performance balance of network where heterogeneous computers inherently root and are intermittently under-utilized. Whereby it, we can expedite to retrieve technical terms from noisy bilingual web text and build up the Portuguese-Chinese corpus-base.
Yi-Pmg Li, Chi-Man Pun
MTSummit2