EDBT 2026 Demo / reviewers in the wild / expert
Guoheng Huang
dblp:140/0792
· DBLP profile ↗
78ranked-venue papers
5as first author
68since 2021 · last 2026
0000-0002-3640-3229ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 27 · 26 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PASA: Progressive-Adaptive Spectral Augmentation for Automated Auscultation in Data-Scarce EnvironmentsabstractAutomated auscultation advances the detection of respiratory diseases, especially in areas with limited resources where traditional diagnostic methods are unavailable. On the other hand, the scarcity of auscultation datasets limits the automation performance, prompting the needs for data augmentation methods. However, most of the existing methods neglect the difference in acoustic sounds that requires personalized augmentation strategies. To address this, we propose a Progressive-Adaptive Spectral Augmentation (PASA), which is one of the first paradigms to adaptively select the best augmentation strategy for each sample. The PASA innovatively treats augmentation selection problem as a Markov Decision Process (MDP), creating an alternating loop between the diagnostic model and the augmentation selection. The agent selects the optimal augmentation operations and magnitudes via a task-specific design, including state construction, action sampling, Hybrid Batch-Sample (HBS) strategy execution, and reward guidance. The HBS strategy initially applies uniform augmentation across mini-batches while collecting sample-specific performance statistics. When model performance stabilizes, it transits to sample-level augmentation based on accumulated difficulty assessments. This two-phase design balances computational complexity with personalization. Extensive experiments across three benchmark datasets demonstrate that the PASA outperforms the state-of-the-art methods, pioneering a transformative paradigm for adaptive data augmentation in automated auscultation. Ying Wang 0097, Guoheng Huang, Xueyuan Gong, Xinxin Wang 0003, Xiaochen Yuan |
AAAI | 2 |
| 2026 | DDG-MoE: A Dual-Driven Mixture-of-Experts Framework for Boundary-Ambiguous Radiological Pattern Disentanglement
Siyue Xie, Chun Peng, Qingyu Zhuang, Guoheng Huang, Xuhang Chen 0002 |
ICIC (1) | 4 |
| 2026 | AHRS-Net: Anatomical Hierarchical Refining Semantic Network for Anatomy-Aware Spatially Grounded Radiology Report Generation
Siyue Xie, Kaijun Shen, Qingyu Zhuang, Guoheng Huang, Xuhang Chen 0002 |
ICIC (20) | 4 |
| 2026 | Decoding coefficients recovery based on modified Gauss-Jordan elimination for tampered content reconstruction
Tong Liu 0021, Lihao Zhuang, Guoheng Huang, Chi-Man Pun, Xiaochen Yuan |
Expert Syst. Appl. | 3 |
| 2026 | A noise-assistant network for tampering detection via inconspicuous feature enhancement and multi-perspective perception
Zhiyao Xie, Xiaochen Yuan, Chan-Tong Lam, Guoheng Huang, Nuno Lourenço 0002 |
Expert Syst. Appl. | 4 |
| 2026 | CCSFusion: A Hierarchical Semantic Chain-of-Thought Reasoning Architecture for Infrared-Visible Image Fusion and CaptioningabstractInfrared-Visible Image Fusion (IVIF) aims to generate a single, information-rich image for downstream tasks. However, prevailing methods exhibit two key limitations. First, many approaches lack explicit hierarchical semantic decoupling, failing to effectively integrate semantic features across different levels, which restricts their ability to capture complex scene structures. Second, task-driven fusion frameworks typically adopt a cascaded design, with unidirectional supervision provided by geometry-centric downstream tasks like detection. This architecture not only limits mutual reinforcement between the fusion and task networks, but also creates a ”supervision bottleneck” by lacking interaction with the linguistic modality that captures richer scene relationships. To tackle these challenges, we propose CCSFusion, the first framework that leverages Chain-of-Thought captioning as supervision, redirecting IVIF optimization from narrow geometric accuracy to multimodal scene comprehension. It establishes a mutually reinforcing coupling between the fusion network and the captioning task. Specifically, we introduce a Segmentation Mask Calibration Unit (SMCU) to refine coarse semantic priors, providing precise pixel-level guidance. Subsequently, the calibrated features are fed into Chained Semantic Fusion Module (CSFM) which explicitly decomposes the semantic priors into three hierarchical levels, and then feeds them into the Hierarchical Semantic Attention module. Finally, a bidirectional knowledge distillation mechanism transfers the reasoning ability of the teacher network to the student. Experiments show that CCSFusion achieves superior fusion performance and generates more semantically coherent images for high-level cognitive tasks. The code is available at: https://github.com/Snaillms/CCSFusion. Miaoshan Lin, Guoheng Huang, Jietao Yang, Jiehao Zheng, Xiaochen Yuan, Yan Li 0122, Xiaofeng Zhang 0006, Kim Fung Tsang, Chi-Man Pun |
IEEE Internet Things J. | 2 |
| 2026 | The Latens Patronus: Seamless Model Watermarking for Latent Diffusion Model in IoT EnvironmentsabstractWith the rapid development of the Internet of Things (IoT), generative artificial intelligence has been widely applied in various IoT applications. However, the wide adoption of latent diffusion model (LDM) in such IoT scenarios raises severe risks of copyright infringement and model theft due to the lack of effective protection mechanisms. To address this challenge, we propose Latens Patronus, a seamless model watermarking technique for copyright protection of LDM in IoT environments. Unlike existing model watermarking methods, our method does not require additional watermark input and additional parameters for a specialized embedding network, making it more suitable for deployment in real-world IoT applications. Specifically, we design a Watermark Encoder to integrate image watermark into latent features during the generation process and aWatermark Decoder to accordingly extract the watermark from suspicious images accurately. We further introduce a diminishing training strategy that gradually fades out auxiliary supervision signals, eliminating the need for persistent watermark guidance, and adding no extra overhead to the original model. Extensive experiments on multiple LDM variants demonstrate that Latens Patronus outperforms existing watermarking methods in both invisibility and robustness against image-level and model-level attacks. Tong Liu 0021, Wei Ke 0001, Guoheng Huang, Xueyuan Gong, Xiaochen Yuan |
IEEE Internet Things J. | 4 |
| 2026 | MADAT: Missing-aware dynamic adaptive transformer model for medical prognosis prediction with incomplete multimodal data
Jianbin He, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Guo Zhong, Bai Ying Lei, Haojiang Li |
Medical Image Anal. | 2 |
| 2026 | QWNet: A quaternion wavelet network for spatial-frequency aware multi-modal image fusion
Jietao Yang, Miaoshan Lin, Guoheng Huang, Xuhang Chen 0002, Xiaofeng Zhang 0006, Xiaochen Yuan, Chi-Man Pun, Bingo Wing-Kuen Ling |
Neural Networks | 3 |
| 2026 | HINTS: Hierarchically Disentangling Subregional Heterogeneity With Structural Priors for Multi-Modal Survival Analysis
Biyun Chen, Guoheng Huang, Xiaochen Yuan, Yan Li 0122, Chi-Man Pun, Bai Ying Lei, Haojiang Li |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | Interpretation Before Integration: LLM-Guided Multimodal Completion and Fusion Network for Survival Analysis With Incomplete Data
Feng Ling 0002, Haoming Zeng, Ming Li 0065, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Xianglian Liao, Jiong-Lin Liang, Haojiang Li |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | DFFormer: Capturing Dynamic Frequency Features to Locate Image Manipulation Through Adaptive Frequency Transformer and Prototype LearningabstractThe proliferation of modern image editing tools has raised concerns about image manipulation, particularly regarding the potential to mislead the public and compromise privacy and security. Consequently, detecting and localizing tampered regions has become a critical research challenge. Traditional methods struggle with subtle manipulations, such as splicing, copy-move, and removal, which are often more discernible in the frequency domain than in the spatial domain. Additionally, the size imbalance between the tampered and background regions further complicates the detection process. To address these challenges, we propose DFFormer, an end-to-end network that leverages frequency feature differences and a dynamic token strategy for precise manipulation localization. DFFormer combines the Conventional Neural Network (CNN) and Transformer in a hybrid architecture with three key modules: the Adaptive Frequency Transformer (AFT), the Prototype Learning Module (PLM), and the Cascaded Progressive Token Fusion Head (CPTF-Head). AFT integrates high- and low-frequency components into self-attention via the Parallel Adaptive Frequency Attention (PAFA) block, enhancing tampering feature representation while preserving fine details. PLM employs KNN-based density peak clustering (DPC-KNN) and weighted token aggregation to optimize dynamic token reduction. The CPTF-Head adopts a hierarchical coarse-to-fine strategy to integrate multiscale features, thereby improving localization accuracy and edge refinement. Experiments demonstrate that DFFormer outperforms state-of-the-art models across four benchmark datasets and one real-world dataset, exhibiting superior generalization and robustness. The source code is publicly available at https://github.com/XiangGD/DFFormer.git. Kaiqi Zhao 0004, Zhenghong Yu, Xiaochen Yuan, Guoheng Huang, Jinyu Tian 0001, Jianqing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Fre-QNet: Quaternion Progressive Perception Mechanism with Frequency-Guided Prompt for Blind Image Quality AssessmentabstractBlind Image Quality Assessment faces challenges in enabling computational models to mimic the hierarchical progressive perception mechanisms of the Human Visual System (HVS). Existing methods often neglect the two-stage process of HVS—global distortion identification followed by local quality evaluation—and its distinct sensitivity to distortion types. To address this, we propose Fre-QNet, a novel framework integrating two key components: (1) A Quaternion Progressive Perception (QPP) module that hierarchically extracts multi-scale spatial features using quaternion convolution, explicitly simulating the global-to-local observation process of HVS while enhancing cross-scale interactions; (2) A Frequency Prompting (FP) module that quantifies distortion types and severity in the Fourier domain by leveraging frequency patterns of common distortions and the sensitivity variations of HVS. The QPP and FP modules collaboratively embed biological vision principles into computational modeling through dual-domain feature learning, with the QPP module directly anchoring the core logic of progressive perception. Experiments on TID2013 and CSIQ benchmarks demonstrate Fre-QNet’s superiority over state-of-the-art methods, validating its effectiveness in matching human perceptual quality judgments. Our source code is available at: https://github.com/hhsda/Fre-QNet . Shize Li, Guoheng Huang, Yisen Zheng, Xiaochen Yuan, Xuhang Chen 0002, Lianglun Cheng, Chi-Man Pun |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | WAQNIQA: Wavelet-Augmented Quaternion Network for No-Reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) plays a pivotal role in computer vision by enabling image quality evaluation without reference images. While recent CNN and Transformer-based methods have advanced feature extraction, they face significant limitations. CNNs exhibit local feature bias, limiting their ability to capture global dependencies and complex structures critical to understanding diverse distortions. Transformers, despite modeling nonlocal dependencies through multihead attention, suffer from quadratic computational complexity with spatial dimensions, hindering efficient multiscale analysis. Moreover, their attention mechanisms frequently overlook critical interchannel dependencies, which are vital for capturing fine details in texture-rich images. Coupled with difficulties in handling high-noise environments and complex textures, this results in limited real-world accuracy and poor generalization across diverse datasets and unknown distortions. To bridge these gaps, we propose WAQNIQA, a novel wavelet-augmented quaternion network for NR-IQA. Distinct from conventional architectures, WAQNIQA integrates two synergistic modules: the wavelet-infused adaptive attention (WIAA) module, which leverages wavelet transforms (WTs) to achieve robust multiscale spatial-frequency analysis with linear complexity, and the quaternion collaborative feature enhancement (QCFE) module, which holistically models interchannel correlations to preserve fine texture details. Furthermore, we introduce PowerGridIQ, the first NR-IQA dataset specifically tailored for power grid scenarios. Extensive experiments demonstrate that WAQNIQA consistently surpasses state-of-the-art CNN and Transformer-based methods on PowerGridIQ and six public benchmarks. Notably, WAQNIQA exhibits superior cross-domain generalization, achieving competitive performance on the AGIQA-1K dataset for AI-generated content (AIGC) without explicit semantic alignment training, thereby validating its robustness against diverse and unknown distortions. Our code is available athttps://github.com/king-huoye/WAQNIQA Yejing Huo, Guoheng Huang, Zhiwen Yu 0002, Xiaochen Yuan, Chi-Man Pun, Lianglun Cheng, Xuhang Chen 0002, Zehong Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | DGLL: A Hybrid Global-Local Feature Learning Network for Precise Tooth Landmark DetectionabstractThe precise identification of key landmarks on three-dimensional tooth mesh models is paramount for computer-aided orthodontic treatment. However, existing methodologies exhibit limitations with respect to the integration of global and local features, which undermines accuracy in complex scenarios and excessively emphasizes relative landmark positions, resulting in displacement errors. To mitigate these issues, this study introduces DGLL, a hybrid feature learning network characterized by a dual-branch architecture that amalgamates global and local features. DGLL integrates a Cascaded Topological Relation Module (CTRM) to stabilize the extraction of global features and a Pan-scale Feature Modulation Module (PFMM) to balance relative and absolute positional accuracy. Empirical evaluations across various tooth types demonstrate that DGLL consistently enhances the accuracy of landmark localization. This research provides an effective approach to the automated analysis of tooth data, thereby improving the precision and efficacy of orthodontic treatment. Jianwen Huang, Guoheng Huang, Fuchen Zheng, Chi-Man Pun, Ka-Cheng Choi, Lianglun Cheng, Guanghui Yue 0001 |
BIBM | 2 |
| 2025 | Lorentz Transformation Neural NetworkabstractWe propose a novel neural network architecture, the Lorentz Transformation Neural Network (LTNN), which utilizes Lorentz transformations to generate a complex computation matrix that enhances the network’s expressive power. Furthermore, LTNN is lightweight due to the shared weight matrices in the computation matrix. LTNN treats the input and output as coordinates in high-dimensional spacetime, with the weight matrices in each layer representing the velocity components of a spacetime reference frame. During training, these weight matrices are transformed into a computation matrix via Lorentz transformations, describing the coordinate transformations between different reference frames. We evaluate LTNN on four datasets: California Housing Prices, Iris, MNIST, and Fashion-MNIST. Experimental results demonstrate that LTNN outperforms conventional neural networks and quaternion neural networks in terms of both accuracy and parameter efficiency. Wenyuan Li 0007, Jingchao Wang 0002, Guoheng Huang, Tongxu Lin, Guo Zhong, Xiaochen Yuan, Chi-Man Pun, An Zeng |
ICIP | 3 |
| 2025 | Hierarchical Graph Learning Framework for Multimodal Conversational Emotion RecognitionabstractAccurate emotion detection in conversations using multimodal features is essential for effective human-computer interaction. There are three pivotal aspects in multimodal emotion recognition in conversation (MERC), i.e., intricate temporal information, modality interactions (both intra- and inter-modal) and implicit high-order linguistic cues within dialogues. Existing approaches are limited to the first two, hindering the generation of effective emotional representations. To this end, we propose HIGH, a hierarchical graph learning framework designed for MERC. HIGH enhances the perception of low-level information (temporal and intra-modal information) by constructing directed dialogue graphs for each modality. A dynamic multimodal filtering mechanism and a modality-aligned contrastive learning approach further refine the semantic nuances. Additionally, the constructed speaker-centered hypergraph yields high-level information like cross-modal interactions and high-order linguistic cues between utterances. HIGH effectively integrates fundamental low-level information with high-order details in a hierarchical manner, considering all three key factors simultaneously. Extensive experiments on two benchmark datasets demonstrate the effectiveness and superiority of HIGH. Jiandong Shi, Ming Li 0065, Guoheng Huang, Yongchun Gu, Zhanle Zhu |
ICME | 3 |
| 2025 | ISSD-NET: Intra-student Self-distillation with Adaptive q-vMF Loss for Enhanced Semi-supervised Medical Segmentation
Guoheng Huang, Xiaochen Yuan, Yan Li 0122, Chi-Man Pun, Bai Ying Lei |
ICONIP (2) | 2 |
| 2025 | Superpixel-Enhanced Quaternion Feature Fusion and Contextualization Graph Contrastive Learning for Cervical Cancer Diagnosis
Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Lianglun Cheng, Chi-Man Pun, Guo Zhong, Qingjian Ye |
ICONIP (2) | 2 |
| 2025 | SimCNet: Leveraging Similarity-Aware and Multi-Scale Features for Robust Cell SegmentationabstractAccurate segmentation and tracking of cells are crucial in biomedical research. The latest methods utilize multilayer convolution or block based region similarity measurement to process foreground background relationships, but under low signal-to-noise ratio conditions, there is insufficient target pattern recognition and boundary detail characterization, which can easily lead to segmentation errors and boundary blurring. To address this issue, we propose SimCNet, a segmentation network based on similarity measurement and texture enhancement. The network consists of three modules: the Global Pattern Sorting and Focusing Module (GPRF), which uses pixel level feature histogram aggregation and a Fourier based feedforward network to create similarity fingerprints for different target patterns, enhancing the recognition ability of different regions. The Multi Scale Feature Collaborative Capture Module (MFCM) enhances the global and inter channel information exchange of GPRF through the fusion of multi-scale fusion mechanisms, eliminating semantic gaps. Texture Feature Enhancement Module (TFEM), which uses quaternion convolution to enrich texture representations in similarity feature maps and enhance boundary perception. Experiments have shown that SimCNet exhibits excellent performance and generalization on multiple publicly available datasets such as MoNuSeg. Guoheng Huang, Xiaoxue Ling, Yumian Yu, Yihang Dong, Yuancong Feng, Lianglun Cheng |
IJCNN | 2 |
| 2025 | Multi-Scale Adaptively-Aware and Recalibration Network for Brain Tumor Segmentation with Missing ModalitiesabstractAccurate segmentation of brain tumor regions from multi-modal magnetic resonance imaging (MRI) is critical for clinical diagnosis. However, missing modalities is a common issue in clinical practice, where the unavailability of certain imaging modalities complicates the extraction and integration of complementary information across multiple modalities, leading to a decline in segmentation accuracy. Many existing models fail to adequately address subtle structural changes and boundary information within tumor regions when faced with missing modalities, limiting their ability to effectively adapt to complex tumor morphologies. To tackle these issues, The Multi-Scale Adaptively-Aware and Recalibration Network (MARNet) proposed in this paper can adaptively and fully explore the potential of multi-modal data under different combinations in the presence of missing modalities. MARNet incorporates a Feature Recalibration and Enhancement Module (FREM) that recalibrates and enhances the three-dimensional feature representation, emphasizing important fine-grained features of brain tumors. Subsequently, the Adaptive Shape-Aware Fusion Module (ASFM) fully exploits available modality information, achieving adaptive feature fusion for varying tumor locations and shapes, thereby compensating for information loss due to missing modalities. Furthermore, the Global and Multiscale Feature Integration Module (GMFIM) is designed to effectively capture long-range dependencies of tumors, particularly under conditions of missing modalities, aiding in the restoration and reconstruction of complete tumor structures. Extensive experiments on the BraTS2020, BraTS2018 and BraTS2015 datasets demonstrate that the proposed method surpasses several advanced brain tumor segmentation approaches in the context of missing modalities. Guoheng Huang, Zhipeng Zheng, Xuhang Chen 0002, Lianglun Cheng |
IJCNN | 2 |
| 2025 | Adaptive Query Prompting for Multi-Domain Landmark DetectionabstractMedical landmark detection is crucial in various medical imaging modalities and procedures. Although deep learning-based methods have achieve promising performance, they are mostly designed for specific anatomical regions or tasks. In this work, we propose a universal model for multi-domain landmark detection by leveraging transformer architecture and developing a prompting component, named as Adaptive Query Prompting (AQP). Transformer backbone architecture is suitable for our study due to its ability to capture long-range dependencies crucial in multi-domain landmark detection. Specifically, transformers excel at understanding global anatomical relationships, which span across entire images. Instead of embedding additional modules in the backbone network, we design a separate module to generate prompts that can be effectively extended to any other transformer network. In our proposed AQP, prompts are learnable parameters maintained in a memory space called prompt pool. The central idea is to keep the backbone frozen and then optimize prompts to instruct the model inference process. Furthermore, we employ a lightweight decoder to decode landmarks from the extracted features, namely Light-MLD. Thanks to the lightweight nature of the decoder and AQP, we can handle multiple datasets by sharing the backbone encoder and then only perform partial parameter tuning without incurring much additional cost. It has the potential to be extended to more landmark detection tasks. We conduct experiments on three widely used X-ray datasets for different medical landmark detection tasks. Our proposed Light-MLD coupled with AQP achieves SOTA performance on many metrics even without the use of elaborate structural designs or complex frameworks. Qiusen Wei, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Jianwen Huang |
IJCNN | 3 |
| 2025 | Guidance Net: Remote Sensing Image Dehazing with Guidance of Prompt Texture Information EmbeddingabstractCurrent remote sensing image dehazing models often encounter challenges under dense haze conditions due to the significant loss of high-frequency information, impairing accurate scene recovery. In addition, these models typically do not incorporate additional sensor information to guide the generation of dehazed images. However, satellites generally house a variety of image sensors, among which panchromatic sensors are included. Panchromatic (PAN) images, which usually present clear boundary information, could potentially assist models in generating dehazed images. We introduce Guidance Net, an innovative model that utilizes historical PAN images to enhance the dehazing process. Specifically, we introduce the Adaptive Self-Attention Interest Texture Filter (ASAITF) for the effective integration of guidance information from PAN images. Additionally, recognizing the limitations of existing methods that predominantly emphasize low-frequency features, we propose the Redundancy Filtering Mechanism (RFM), aimed at efficient high-frequency feature extraction and seamless integration within Vision Transformer architectures. To ensure a comprehensive evaluation, we also present the PAN Guidance dataset. Experimental results indicate that Guidance-Net surpasses state-of-the-art methods in generating dehazed images guided by prompts. Zhengguang Tan, Guoheng Huang, Lianglun Cheng, Alex Hayman Ng |
IJCNN | 2 |
| 2025 | Code Retrieval with Mixture of Experts Prototype Learning Based on ClassificationabstractThe semantic connection between code and queries is crucial for code retrieval, but many human-written queries fail to accurately capture the code's core intent, leading to ambiguity.This ambiguity complicates the code search process, as the queries do not provide a clear overview of the code's purpose.Our analysis reveals that while ambiguous queries may not precisely summarize the intent of the code, they often share the same general topics as the corresponding code.In light of this discovery, we propose Code Retrieval with Mixture of Experts Prototype Learning Based on Classification (CRME), a novel approach that combines classification for prototype-based representation learning and result ensembling.CRME utilizes specialized pre-trained models focused on the specific domains of ambiguous queries.It consists of two key components: Multiple Classification Prototype and Representation Learning with a Prototype-based Multi-model Contrastive (PMC) Loss during training, and Multi-Prototype Mixture of Experts Integration (MP-MoE) module for fine-grained ensemble inference.Our method can effectively address the issue of query ambiguity and improves search precision.Experimental results on the CodeSearchNet dataset, covering six sub-datasets, show that CRME outperforms existing methods, achieving an average MRR score of * Corresponding authors. Feng Ling 0002, Guoheng Huang, Jingchao Wang 0002, Xiaochen Yuan, Xuhang Chen 0002, XueYong Zhang, Fanlong Zhang, Chi-Man Pun |
Internetware | 2 |
| 2025 | Cross-View Geo-Localization via Learning Correspondence Semantic Similarity Knowledge
Guanli Chen, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun |
MMM (1) | 2 |
| 2025 | The Structure-sharing Hypergraph Reasoning Attention Module for CNNs
Jingchao Wang 0002, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Tongxu Lin, Chi-Man Pun, Fenfang Xie |
Expert Syst. Appl. | 2 |
| 2025 | Using 3D-LMM-Based Encryption to Secure Digital Images With 3-D S-Box and Fibonacci Q-MatrixabstractThe rapid development of communication technology has significantly improved information transmission and increased capacity. In the context of the Internet of Things (IoT), where massive visual data transmission faces challenges of real-time processing and security threats. To address this problem, this paper proposes a novel encryption algorithm based on a 3D S-box combined with Fractal-Sort-Matrix (FSM) and Fibonacci Q-Matrix (3DSFF). To address the shortcomings of traditional S-box encryption, this paper integrates the 3D S-box with the FSM, thus enhancing the uncertainty of permutation and transformation. Additionally, to overcome the limitation of using a single value inciting Fibonacci Q-Matrix (FQM) applications, this paper combines FQM with chaotic sequences to strengthen its resistance against exhaustive attacks. Through comprehensive experimental tests on the algorithm’s outcomes, it demonstrates notable improvements over previous algorithms, achieving an average information entropy of 7.9993 and a correlation coefficient close to 0.01 after encryption. These tests indicate that the scheme can withstand common attacks, making it a sufficiently secure solution for the confidentiality of private images. Moreover, these advances provide a secure and efficient visual data protection framework for IoT applications involving anti-theft surveillance, data acquisition, and communication transmission. Yunlong Liao, Qiutong Li, Guoheng Huang, Donald Donglong Chen, Xiaochen Yuan |
IEEE Internet Things J. | 5 |
| 2025 | Visual-linguistic Diagnostic Semantic Enhancement for medical report generation
Jiahong Chen, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Zhe Tan, Chi-Man Pun |
J. Biomed. Informatics | 2 |
| 2025 | MPCM-RRG: Multi-modal Prompt Collaboration Mechanism for Radiology Report Generation
Yumian Yu, Guoheng Huang, Zhe Tan, Ming Li 0065, Chi-Man Pun, Fuchen Zheng, Shiqiang Ma, Shuqiang Wang |
J. Biomed. Informatics | 2 |
| 2025 | A Two-Phase Scheme by Integration of Deep and Corner Feature for Balanced Copy-Move Forgery LocalizationabstractIn the era of Industry 4.0, the widespread application of digitization, automation, and Internet technology in industrial production has led to a significant increase in image data. Image security has become crucial because images are at risk of being tampered with at any time. To protect its authenticity, this article proposes a two-phase scheme to achieve balanced performance between accuracy and speed for copy-move forgery detection. Our scheme is divided into detection and localization phases. In the detection phase, the deep features are utilized to calculate the inner similarity. To improve the accuracy, a corner point matching technique is performed on the localization phase as a refinement step. The experimental results demonstrate the average$F1$-score is 0.6334 on CASIA2.0, making a 14.16% improvement. The computation time for each image is only 0.791 s in average. It has great significance in protecting the reliability and authenticity of industrial data. Tong Liu 0021, Xiaochen Yuan, Zhiyao Xie, Kaiqi Zhao 0004, Guoheng Huang, Chi-Man Pun |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Psanet: prototype-guided salient attention for few-shot segmentation
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun |
Vis. Comput. | 2 |
| 2025 | Weakly supervised semantic segmentation via saliency perception with uncertainty-guided noise suppression
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Guo Zhong, Xuhang Chen 0002, Chi-Man Pun |
Vis. Comput. | 2 |
| 2025 | QEAN: quaternion-enhanced attention network for visual dance generation
Zhizhen Zhou, Yejing Huo, Guoheng Huang, An Zeng, Xuhang Chen 0002, Lian Huang, Zinuo Li |
Vis. Comput. | 3 |
| 2024 | PVALane: Prior-Guided 3D Lane Detection with View-Agnostic Feature AlignmentabstractMonocular 3D lane detection is essential for a reliable autonomous driving system and has recently been rapidly developing. Existing popular methods mainly employ a predefined 3D anchor for lane detection based on front-viewed (FV) space, aiming to mitigate the effects of view transformations. However, the perspective geometric distortion between FV and 3D space in this FV-based approach introduces extremely dense anchor designs, which ultimately leads to confusing lane representations. In this paper, we introduce a novel prior-guided perspective on lane detection and propose an end-to-end framework named PVALane, which utilizes 2D prior knowledge to achieve precise and efficient 3D lane detection. Since 2D lane predictions can provide strong priors for lane existence, PVALane exploits FV features to generate sparse prior anchors with potential lanes in 2D space. These dynamic prior anchors help PVALane to achieve distinct lane representations and effectively improve the precision of PVALane due to the reduced lane search space. Additionally, by leveraging these prior anchors and representing lanes in both FV and bird-eye-viewed (BEV) spaces, we effectively align and merge semantic and geometric information from FV and BEV features. Extensive experiments conducted on the OpenLane and ONCE-3DLanes datasets demonstrate the superior performance of our method compared to existing state-of-the-art approaches and exhibit excellent robustness. Zewen Zheng, Yongqiang Mou, Guoheng Huang, Chi-Man Pun, Xiaochen Yuan |
AAAI | 6 |
| 2024 | IMAN: An Adaptive Network for Robust NPC Mortality Prediction with Missing ModalitiesabstractAccurate prediction of mortality in nasopharyngeal carcinoma (NPC), a complex malignancy particularly challenging in advanced stages, is crucial for optimizing treatment strategies and improving patient outcomes. However, this predictive process is often compromised by the high-dimensional and heterogeneous nature of NPC-related data, coupled with the pervasive issue of incomplete multi-modal data, manifesting as missing radiological images or incomplete diagnostic reports. Traditional machine learning approaches suffer significant performance degradation when faced with such incomplete data, as they fail to effectively handle the high-dimensionality and intricate correlations across modalities. Even advanced multi-modal learning techniques like Transformers struggle to maintain robust performance in the presence of missing modalities, as they lack specialized mechanisms to adaptively integrate and align the diverse data types, while also capturing nuanced patterns and contextual relationships within the complex NPC data. To address these problem, we introduce IMAN: an adaptive network for robust NPC mortality prediction with missing modalities. IMAN features three integrated modules: the Dynamic Cross-Modal Calibration (DCMC) module employs adaptive, learnable parameters to scale and align medical images and field data; the Spatial-Contextual Attention Integration (SCAI) module enhances traditional Transformers by incorporating positional information within the self-attention mechanism, improving multi-modal feature integration; and the Context-Aware Feature Acquisition (CAFA) module adjusts convolution kernel positions through learnable offsets, allowing for adaptive feature capture across various scales and orientations in medical image modalities. Extensive experiments on our proprietary NPC dataset demonstrate IMAN’s robustness and high predictive accuracy, even with missing data. Compared to existing methods, IMAN consistently outperforms in scenarios with incomplete data, representing a significant advancement in mortality prediction for medical diagnostics and treatment planning. Our code is available at https://github.com/king-huoye/BIBM-2024/tree/master. Yejing Huo, Guoheng Huang, Lianglun Cheng, Jianbin He, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun |
BIBM | 2 |
| 2024 | FAQNet: Frequency-Aware Quaternion Network for Endoscopic Highlight RemovalabstractDue to the built-in light source within the endoscope, the illumination of bodily mucous can cause the formation of highlight regions due to reflection. This not only interferes with the diagnosis conducted by doctors but also poses a challenge to subsequent computer vision tasks. To tackle this issue, we introduce FAQNet, a network specifically designed for endoscopic image highlight removal. FAQNet seamlessly integrates multi-channel information leveraging quaternion convolution and spatial channel attention within our Quaternion Multi-Channel Fusion (QMCF) Module. This allows it to capture intricate details of color, texture, spatial information, and highlight characteristics within the imaged organ. Additionally, by employing frequency domain transformation and dilated convolution, the Contextual Information Integration (CII) Module effectively enlarges the receptive field, organizing contextual information between highlight regions and their surrounding areas. Lastly, the PixelShuffle Upsampling (PSU) Module generates the restored image. We validate our model’s performance on two benchmark datasets, demonstrating its superiority over existing highlight removal methodologies. Dingzhou Zhu, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun |
BIBM | 2 |
| 2024 | PDGC: Properly Disentangle by Gating and Contrasting for Cross-Domain Few-Shot Classification
Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Yan Li 0122, Chi-Man Pun, Junbing Quan |
CGI (2) | 2 |
| 2024 | FOPS-V: Feature-Aware Optimization and Parallel Scale Fusion for 3D Human Reconstruction in Video
Guoheng Huang, Lianglun Cheng, Yejing Huo, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun |
ICONIP (8) | 2 |
| 2024 | ROSAL: Semi-supervised Active Learning with Representation Aggregation and Outlier for Endoscopy Image Classification
Xiaocong Huang, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Xuhang Chen 0002, Chi-Man Pun, Jianwu Chen |
ICONIP (11) | 2 |
| 2024 | DSTNet: Distinguishing Source and Target Areas for Image Copy-Move Forgery Detection
Kaiqi Zhao 0004, Xiaochen Yuan, Guoheng Huang |
ICPR (22) | 3 |
| 2024 | Medical Visual Prompting (MVP): A Unified Framework for Versatile and High-Quality Medical Image SegmentationabstractAccurate segmentation of lesion regions is crucial for clinical diagnosis and treatment across various diseases. While deep convolutional networks have achieved satisfactory results in medical image segmentation, they face challenges such as the loss of lesion shape information due to continuous convolution and downsampling, as well as the high cost of manually labeling lesions with varying shapes and sizes. To address these issues, we propose a novel Medical Visual Prompting (MVP) framework that leverages pre-training and prompting concepts from Natural Language Processing (NLP). The framework utilizes three key components: Super-Pixel Guided Prompting (SPGP) for superpixelating the input image, Image Embedding Guided Prompting (IEGP) for freezing patch embedding and merging with superpixels to provide visual prompts, and Adaptive Attention Mechanism Guided Prompting (AAGP) for pinpointing prompt content and efficiently adapting all layers. By integrating SPGP, IEGP, and AAGP, the MVP framework enables the segmentation network to better learn shape prompting information and facilitates mutual learning across different tasks. Extensive experiments conducted on five datasets demonstrate the superior performance of the proposed method in various challenging medical image tasks while simplifying single-task medical segmentation models. This novel framework offers improved performance with fewer parameters and holds significant potential for accurate segmentation of lesion regions in various medical tasks, making it clinically valuable. Guoheng Huang, Zijin Lin, Guo Zhong, Shenghong Luo |
SMC | 2 |
| 2024 | Cross-Modality Disentangled Information Bottleneck Strategy for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) has been a pivotal domain in current research area which utilizes diverse information carriers such as videos containing multiple modal-ities to understand the user's sentiment. With the success of multimodal fusion techniques, lots of fusion strategies have been proposed to obtain a favorable multimodal joint representation for MSA. However, existing studies hardly consider the problem of redundant information in unimodal, resulting in the joint representation may contain much redundant information from different modalities, thus limiting the accuracy of sentiment prediction. In this work, we propose a Cross-Modality Disentangled Information Bottleneck Strategy (CMDIBS), which consists of a Cross-Modality Knowledge Awareness (CMKA) module and a Multimodal Disentangled Information Bottleneck (MDIB) mechanism. Specifically, the CMKA module encourages in-teractions among different modalities to learn the sentiment embedding relevant to the predicted goals. In particular, MDIB mechanism aims to maximize the mutual information (MI) between the multimodal joint representation and the predicted label, and maximize the MI between the style embedding with the label and the input data while constraining the MI between the multimodal joint representation and the style embedding to obtain a succinct and efficient multimodal joint representation. Experimental results on the benchmark datasets, namely CMU-MOSI and CMU-MOSEI, indicated that the proposed method surpasses existing approaches and attains SOTA performance. Zhengnan Deng, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Lian Huang, Chi-Man Pun |
SMC | 2 |
| 2024 | IAMS-Net: An Illumination-Adaptive Multi-Scale Lesion Segmentation NetworkabstractIn recent years, many Lesion segmentation (LS) models based on UNet have been proposed. However, existing researches rarely consider the influence of illumination change leads to the weak boundary area. Such as melanomas and polyps, the demarcation of the boundary between the diseased area and the surrounding tissue remains particularly challenging. To overcome these challenges, we propose an IlluminationAdaptive Multi-scale Lesion Segmentation Network (IAMS-Net). In IAMS-Net, we integrate Illumination-Adaptive MultiStream Attention (IAMA) and Contour Perception Module (CPM). In the decoding stage, the IAMA is used as a bridge between the encoder and the decoder to solve the adverse effects of illumination changes on the segmentation of weak boundary lesions. In order to further enhance the boundary features lost due to illumination change in the low-contrast lesion area, we introduce the CPM to improve the perception of the integrity of the lesion area. Subsequently, we performed comparison and ablation experiments using the publicly available ISIC2018 dataset and the individually collected data set BoreIllumination(BI). Yisen Zheng, Guoheng Huang, Lianglun Cheng, Xiaochen Yuan, Guo Zhong, Shenghong Luo |
SMC | 2 |
| 2024 | CAMU-Net: Copy-move forgery detection utilizing coordinate attention and multi-scale feature fusion-based up-sampling
Kaiqi Zhao 0004, Xiaochen Yuan, Tong Liu 0021, Zhiyao Xie, Guoheng Huang, Li Feng 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Binary spectral clustering for multi-view data
Xueming Yan, Guo Zhong, Yaochu Jin, Xiaohua Ke, Fenfang Xie, Guoheng Huang |
Inf. Sci. | 6 |
| 2024 | Black-box reversible adversarial examples with invertible neural network
Jielun Huang, Guoheng Huang, Xiaochen Yuan, Fenfang Xie, Chi-Man Pun, Guo Zhong |
Image Vis. Comput. | 2 |
| 2024 | Cross-domain visual prompting with spatial proximity knowledge distillation for histological image classification
Guoheng Huang, Lianglun Cheng, Guo Zhong, Weihuang Liu, Xuhang Chen 0002, Muyan Cai |
J. Biomed. Informatics | 2 |
| 2024 | Progressive normalizing flow with learnable spectrum transform for style transfer
Guoheng Huang, Xiaochen Yuan, Guo Zhong, Chi-Man Pun, Yiwen Zeng |
Knowl. Based Syst. | 2 |
| 2024 | GDN-CMCF: A Gated Disentangled Network With Cross-Modality Consensus Fusion for Multimodal Named Entity RecognitionabstractMultimodal named entity recognition (MNER) is a crucial task in social systems of artificial intelligence that requires precise identification of named entities in sentences using both visual and textual information. Previous methods have focused on capturing fine-grained visual features and developing complex fusion procedures. However, these approaches overlook the heterogeneity gap and loss of original modality uniqueness that may occur during fusion, leading to incorrect entity identification. This article proposes a novel approach for MNER called a gated disentangled network with cross-modality consensus fusion (GDN-CMCF) to address the above challenges. Specifically, to eliminate cross-modality variation, we propose a cross-modality consensus fusion module that generates a consensus representation by learning inter-and intramodality interactions with a designed commonality constraint. We then introduce a gated disentanglement module to separate modality-relevant features from support and auxiliary modalities, which further filters out extraneous information while retaining the uniqueness of unimodal features. Experimental results on two real public datasets are provided to verify the effectiveness of our proposed GDN-CMCF. The source code of this article can be found at https://github.com/HaoDavis/ GDN-CMCF. Guoheng Huang, Zihao Dai, Guo Zhong, Xiaochen Yuan, Chi-Man Pun |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Quaternion Cross-Modality Spatial Learning for Multi-Modal Medical Image SegmentationabstractRecently, the Deep Neural Networks (DNNs) have had a large impact on imaging process including medical image segmentation, and the real-valued convolution of DNN has been extensively utilized in multi-modal medical image segmentation to accurately segment lesions via learning data information. However, the weighted summation operation in such convolution limits the ability to maintain spatial dependence that is crucial for identifying different lesion distributions. In this paper, we propose a novel Quaternion Cross-modality Spatial Learning (Q-CSL) which explores the spatial information while considering the linkage between multi-modal images. Specifically, we introduce to quaternion to represent data and coordinates that contain spatial information. Additionally, we propose Quaternion Spatial-association Convolution to learn the spatial information. Subsequently, the proposed De-level Quaternion Cross-modality Fusion (De-QCF) module excavates inner space features and fuses cross-modality spatial dependency. Our experimental results demonstrate that our approach compared to the competitive methods perform well with only 0.01061 M parameters and 9.95G FLOPs. Junyang Chen 0001, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Zewen Zheng, Chi-Man Pun, Jian Zhu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Learning From Incorrectness: Active Learning With Negative Pre-Training and Curriculum Querying for Histological Tissue ClassificationabstractPatch-level histological tissue classification is an effective pre-processing method for histological slide analysis. However, the classification of tissue with deep learning requires expensive annotation costs. To alleviate the limitations of annotation budgets, the application of active learning (AL) to histological tissue classification is a promising solution. Nevertheless, there is a large imbalance in performance between categories during application, and the tissue corresponding to the categories with relatively insufficient performance are equally important for cancer diagnosis. In this paper, we propose an active learning framework called ICAL, which contains Incorrectness Negative Pre-training (INP) and Category-wise Curriculum Querying (CCQ) to address the above problem from the perspective of category-to-category and from the perspective of categories themselves, respectively. In particular, INP incorporates the unique mechanism of active learning to treat the incorrect prediction results that obtained from CCQ as complementary labels for negative pre-training, in order to better distinguish similar categories during the training process. CCQ adjusts the query weights based on the learning status on each category by the model trained by INP, and utilizes uncertainty to evaluate and compensate for query bias caused by inadequate category performance. Experimental results on two histological tissue classification datasets demonstrate that ICAL achieves performance approaching that of fully supervised learning with less than 16% of the labeled data. In comparison to the state-of-the-art active learning algorithms, ICAL achieved better and more balanced performance in all categories and maintained robustness with extremely low annotation budgets. The source code will be released at https://github.com/LactorHwt/ICAL. Lianglun Cheng, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Chi-Man Pun, Muyan Cai |
IEEE Trans. Medical Imaging | 3 |
| 2024 | SCDet: decoupling discriminative representation for dark object detection via supervised contrastive learning
Tongxu Lin, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Xiaocong Huang, Chi-Man Pun |
Vis. Comput. | 2 |
| 2023 | A Noise Convolution Network for Tampering Detection
Zhiyao Xie, Xiaochen Yuan, Chan-Tong Lam, Guoheng Huang |
ICANN (10) | 4 |
| 2023 | RA-Net: A Deep Learning Approach Based on Residual Structure and Attention Mechanism for Image Copy-Move Forgery Detection
Kaiqi Zhao 0004, Xiaochen Yuan, Zhiyao Xie, Guoheng Huang, Li Feng 0001 |
ICANN (10) | 4 |
| 2023 | Single Cross-domain Semantic Guidance Network for Multimodal Unsupervised Image Translation
Jiaying Lan, Lianglun Cheng, Guoheng Huang, Chi-Man Pun, Xiaochen Yuan, Shangyu Lai, Bingo Wing-Kuen Ling |
MMM (1) | 3 |
| 2023 | TriView-ParNet: parallel network for hybrid recognition of touching printed and handwritten strings based on feature fusion and three-view co-training
Junhao Qiu, Shangyu Lai, Guoheng Huang, Junhui Mai, Chi-Man Pun, Bingo Wing-Kuen Ling |
Appl. Intell. | 3 |
| 2023 | Tampering localization and self-recovery using block labeling and adaptive significance
Xiaochen Yuan, Tong Liu 0021, Chan-Tong Lam, Guoheng Huang, Di Lin 0002, Ping Li 0016 |
Expert Syst. Appl. | 5 |
| 2023 | RBA-GCN: Relational Bilevel Aggregation Graph Convolutional Network for Emotion RecognitionabstractEmotion recognition in conversation (ERC) has received increasing attention from researchers due to its wide range of applications. As conversation has a natural graph structure, numerous approaches used to model ERC based on graph convolutional networks (GCNs) have yielded significant results. However, the aggregation approach of traditional GCNs suffers from the node information redundancy problem, leading to node discriminant information loss. Additionally, single-layer GCNs lack the capacity to capture long-range contextual information from the graph. Furthermore, the majority of approaches are based on textual modality or stitching together different modalities, resulting in a weak ability to capture interactions between modalities. To address these problems, we present the relational bilevel aggregation graph convolutional network (RBA-GCN), which consists of three modules: the graph generation module (GGM), similarity-based cluster building module (SCBM) and bilevel aggregation module (BiAM). First, GGM constructs a novel graph to reduce the redundancy of target node information. Then, SCBM calculates the node similarity in the target node and its structural neighborhood, where noisy information with low similarity is filtered out to preserve the discriminant information of the node. Meanwhile, BiAM is a novel aggregation method that can preserve the information of nodes during the aggregation process. This module can construct the interaction between different modalities and capture long-range contextual information based on similarity clusters. On both the IEMOCAP and MELD datasets, the weighted average F1 score of RBA-GCN has a 2.17$\sim$5.21% improvement over that of the most advanced method. Guoheng Huang, Fenghuan Li, Xiaochen Yuan, Chi-Man Pun, Guo Zhong |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Quaternion-Valued Correlation Learning for Few-Shot Semantic SegmentationabstractFew-shot segmentation (FSS) aims to segment unseen classes given only a few annotated samples. Encouraging progress has been made for FSS by leveraging semantic features learned from base classes with sufficient training samples to represent novel classes. The correlation-based methods lack the ability to consider interaction of the two subspace matching scores due to the inherent nature of the real-valued 2D convolutions. In this paper, we introduce a quaternion perspective on correlation learning and propose a novel Quaternion-valued Correlation Learning Network (QCLNet), with the aim to alleviate the computational burden of high-dimensional correlation tensor and explore internal latent interaction between query and support images by leveraging operations defined by the established quaternion algebra. Specifically, our QCLNet is formulated as a hyper-complex valued network and represents correlation tensors in the quaternion domain, which uses quaternion-valued convolution to explore the external relations of query subspace when considering the hidden relationship of the support sub-dimension in the quaternion space. Extensive experiments on the PASCAL-$5^{i}$and COCO-$20^{i}$datasets demonstrate that our method outperforms the existing state-of-the-art methods effectively. Zewen Zheng, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Bingo Wing-Kuen Ling |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | QGD-Net: A Lightweight Model Utilizing Pixels of Affinity in Feature Layer for Dermoscopic Lesion SegmentationabstractRESPONSE: Pixels with location affinity, which can be also called "pixels of affinity," have similar semantic information. Group convolution and dilated convolution can utilize them to improve the capability of the model. However, for group convolution, it does not utilize pixels of affinity between layers. For dilated convolution, after multiple convolutions with the same dilated rate, the pixels utilized within each layer do not possess location affinity with each other. To solve the problem of group convolution, our proposed quaternion group convolution uses the quaternion convolution, which promotes the communication between to promote utilizing pixels of affinity between channels. In quaternion group convolution, the feature layers are divided into 4 layers per group, ensuring the quaternion convolution can be performed. To solve the problem of dilated convolution, we propose the quaternion sawtooth wave-like dilated convolutions module (QS module). QS module utilizes quaternion convolution with sawtooth wave-like dilated rates to effectively leverage the pixels that share the location affinity both between and within layers. This allows for an expanded receptive field, ultimately enhancing the performance of the model. In particular, we perform our quaternion group convolution in QS module to design the quaternion group dilated neutral network (QGD-Net). Extensive experiments on Dermoscopic Lesion Segmentation based on ISIC 2016 and ISIC 2017 indicate that our method has significantly reduced the model parameters and highly promoted the precision of the model in Dermoscopic Lesion Segmentation. And our method also shows generalizability in retinal vessel segmentation. Jingchao Wang 0002, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Chi-Man Pun |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Bi-deformation-UNet: recombination of differential channels for printed surface defect detection
Guoheng Huang, Ying Wang 0097, Junhao Qiu, Zhiwen Yu 0002, Chi-Man Pun, Bingo Wing-Kuen Ling |
Vis. Comput. | 2 |
| 2023 | Unsupervised style-guided cross-domain adaptation for few-shot stylized face translation
Jiaying Lan, Fenghua Ye, Zhenghua Ye, Pingping Xu, Bingo Wing-Kuen Ling, Guoheng Huang |
Vis. Comput. | 6 |
| 2023 | Video action recognition with Key-detail Motion Capturing based on motion spectrum analysis and multiscale feature fusion
Ganghan Zhang, Guoheng Huang, Haiyuan Chen, Chi-Man Pun, Zhiwen Yu 0002, Bingo Wing-Kuen Ling |
Vis. Comput. | 2 |
| 2022 | Fine-grained visual classification with multi-scale features based on self-supervised attention filtering mechanism
Haiyuan Chen, Lianglun Cheng, Guoheng Huang, Ganghan Zhang, Jiaying Lan, Zhiwen Yu 0002, Chi-Man Pun, Bingo Wing-Kuen Ling |
Appl. Intell. | 3 |
| 2022 | MIVCN: Multimodal interaction video captioning network based on semantic association graph
Ying Wang 0097, Guoheng Huang, Yuming Lin 0005, Chi-Man Pun, Bingo Wing-Kuen Ling, Lianglun Cheng |
Appl. Intell. | 2 |
| 2022 | Multi-view spectral clustering by simultaneous consensus graph learning and discretization
Guo Zhong, Ting Shu 0001, Guoheng Huang, Xueming Yan |
Knowl. Based Syst. | 3 |
| 2022 | LE-MSFE-DDNet: a defect detection network based on low-light enhancement and multi-scale feature extraction
Weihua Hu, Yangsai Wang, Guoheng Huang |
Vis. Comput. | 5 |
| 2021 | Compensating the vorticity loss during advection with an adaptive vorticity confinement forceabstractAbstract The advection step in grid‐based fluid simulation is prone to numerical dissipation, which results in loss of detail. How to improve the advection accuracy to preserve more fluid details is still challenging. On the other hand, a common way to enhance smoke details is to use vorticity confinement. However, most of the previous methods simply used a fine‐tuned scale factor ε to adjust the strength of the confinement force, which can only amplify existing vortex details and is easy to cause instability when ε is large. In this article, we proposed an adaptive vorticity confinement method, which does not suffer from the above problems, to compensate the vorticity loss during advection with little extra cost. The main idea is to first calculate a scale factor whose value depends on the vorticity loss during advection, and then use it to adaptively control the vorticity confinement force for vorticity compensation with high stability. The experiment results show the effectiveness and efficiency of our method. Jian Zhu 0001, Silong Li, Ruichu Cai, Guoheng Huang, Bin Sheng 0001, Enhua Wu |
Comput. Animat. Virtual Worlds | 5 |
| 2020 | Near orthogonal discrete quaternion Fourier transform components via an optimal frequency rescaling approachabstractThe quaternion‐valued signals consist of four signal components. The discrete quaternion Fourier transform is to map these four signal components in the time domain to that in the frequency domain. These four signal components in the frequency domain are called the discrete quaternion Fourier transform components. There are a total of 16 inner products among any two discrete quaternion Fourier transform components. The total orthogonal error among the discrete quaternion Fourier transform components is defined based on these 16 inner products. This study aims to find the optimal quaternion number in the discrete quaternion Fourier transforms so that the total orthogonal errors among the discrete quaternion Fourier transform components are minimised. It is worth noting that finding the optimal quaternion number in the discrete quaternion Fourier transform is equivalent to finding the optimal rescaling factors. Since the discrete quaternion Fourier transform components are expressed in terms of the high‐order polynomials of the trigonometric functions of the rescaling factors, this optimisation problem is non‐convex. To address this problem, a two‐stage approach is employed for finding the solution to the optimisation problem. The comparison results show that the authors proposed method outperforms the existing methods in terms of achieving the low total orthogonal error among the discrete quaternion Fourier transform components. Lingyue Hu, Bingo Wing-Kuen Ling, Charlotte Yuk-Fan Ho, Guoheng Huang |
IET Signal Process. | 4 |
| 2020 | Animating turbulent fluid with a robust and efficient high-order advection methodabstractAbstract The accuracy of advection has a great influence on the visual effect of fluid simulation. Constrained interpolation profile (CIP) method has been an important advection scheme because of its third‐order accuracy and the fact that it only needs to be performed over a compact stencil, but extending it to high‐dimensional advection equations is not easy, because it involves complex calculations and large memory overheads, and is usually unstable. In this article, we propose a stable and efficient three‐dimensional (3D) CIP scheme which can maintain high accuracy but requires low computation and memory cost. We first construct an efficient two‐dimensional (2D) CIP scheme based on dimensional splitting and local Taylor expansions, and then propose an effective way to extend it for 3D applications without decreasing the computational accuracy or affecting the stability. The experimental results show the advantages of our method over the state‐of‐the‐art advection schemes. Jian Zhu 0001, Silong Li, Ruichu Cai, Guoheng Huang, Bin Sheng 0001, Enhua Wu |
Comput. Animat. Virtual Worlds | 5 |
| 2020 | Rapid facial expression recognition under part occlusion based on symmetric SURF and heterogeneous soft partition network
Guoheng Huang, Chi-Man Pun, Bingo Wing-Kuen Ling, Lianglun Cheng |
Multim. Tools Appl. | 2 |
| 2020 | Person re-identification based on multi-level feature complementarity of cross-attention with part metric learning
Zeng Lu, Guoheng Huang, Chi-Man Pun, Lianglun Cheng |
Multim. Tools Appl. | 2 |
| 2018 | On-line video multi-object segmentation based on skeleton model and occlusion detection
Guoheng Huang, Chi-Man Pun |
Multim. Tools Appl. | 1 |
| 2017 | Unsupervised video co-segmentation based on superpixel co-saliency and region merging
Guoheng Huang, Chi-Man Pun, Cong Lin 0001 |
Multim. Tools Appl. | 1 |
| 2017 | Highly non-rigid video object tracking using segment-based object candidates
Cong Lin 0001, Chi-Man Pun, Guoheng Huang |
Multim. Tools Appl. | 3 |
| 2016 | On-line video object segmentation using illumination-invariant color-texture feature extraction and marker prediction
Chi-Man Pun, Guoheng Huang |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Non-rigid visual object tracking using user-defined marker and Gaussian kernel
Guoheng Huang, Chi-Man Pun, Cong Lin 0001, Yicong Zhou |
Multim. Tools Appl. | 1 |
| 2015 | Video Object Tracking Using Interactive Segmentation and Superpixel Based Gaussian KernelabstractA novel non-rigid video object tracking based on interactive segmentation and super pixel Gaussian kernel is proposed in this paper. In the initialization stage, instead of using the traditional bounding box to locate the targeted object, we employed an interactive segmentation with user-defined marker to segment the object accurately in the first frame of the input video to avoid the background influence in the traditional bounding box. During the tracking stage, using a Gaussian kernel as movement constraint, each super pixel is tracked independently to locate the object in the next frame. Experimental results show that the proposed method compared to state of the art methods can achieve better robustness and accuracy for various challenging video clips. Guoheng Huang, Chi-Man Pun, Cong Lin 0001 |
IV | 1 |