Huiyu Zhou 0001

dblp:36/1648 · DBLP profile ↗
← Back
273ranked-venue papers
25as first author
178since 2021 · last 2027
0000-0003-1634-9840ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 120 · 14 first-author · 77 since 2021Graphics, computer vision, multimedia, augmented reality and games · 87 · 10 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 51 · 43 since 2021Databases, data management, data science and information retrieval · 17 · 14 since 2021Computer networks · 16 · 10 since 2021Security and privacy · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2027 SVR-UNet: Frequency-aware view-routed analysis-synthesis sampling for 3D medical image segmentation
Shuanghua Ye, Wenwen Tang, Huiyu Zhou 0001, Jin Liu 0025, Xiaoshuai Zhang, Xingru Huang
Expert Syst. Appl.4
2026 Wavefront-Constrained Passive Obscured Object Detection
abstract
Accurately localizing and segmenting obscured objects from faint light patterns beyond the field of view is highly challenging due to multiple scattering and medium-induced perturbations. Most existing methods, based on real-valued modeling or local convolutional operations, are inadequate for capturing the underlying physics of coherent light propagation. Moreover, under low signal-to-noise conditions, these methods often converge to non-physical solutions, severely compromising the stability and reliability of the observation. To address these challenges, we propose a novel physics-driven Wavefront Propagating Compensation Network (WavePCNet) to simulate wavefront propagation and enhance the perception of obscured objects. This WavePCNet integrates the Tri-Phase Wavefront Complex-Propagation Reprojection (TriWCP) to incorporate complex amplitude transfer operators to precisely constrain coherent propagation behavior, along with a momentum memory mechanism to effectively suppress the accumulation of perturbations. Additionally, a High-frequency Cross-layer Compensation Enhancement is introduced to construct frequency-selective pathways with multi-scale receptive fields and dynamically models structural consistency across layers, further boosting the model’s robustness and interpretability under complex environmental conditions. Extensive experiments conducted on four physically collected datasets demonstrate that WavePCNet consistently outperforms state-of-the-art methods across both accuracy and robustness.
Yiwei Ouyang, Xiaoshuai Zhang, Huiyu Zhou 0001, Wenwen Tang, Shaowei Jiang, Jin Liu 0025, Xingru Huang
AAAI6
2026 Multi-HypPre: a multi-source data fusion framework for drug-induced hypertension prediction
Xueqiang Gao, Chen Ge, Kui Lu, Fengjun Zhou, Huiyu Zhou 0001, Xiaoying Wang 0007
Appl. Intell.7
2026 Multi-Granularity Modal Interaction and Fusion framework for vision-language tasks
Yangshuyi Xu, Guangzhong Liu, Xiang Shen 0002, Xiuying Wang 0001, Huiyu Zhou 0001
Eng. Appl. Artif. Intell.5
2026 Part-aware cross-integration transformer with proximity guided regularization for domain generalizable animal re-identification
Zeyuan Sun, Junyu Dong, Xiaowei Zhou 0003, Huiyu Zhou 0001, Hao Fan 0004
Expert Syst. Appl.4
2026 P3R: Polymodal palpebral progressive refinement via symmetry aware latent diffusion for precision guided prediction of postoperative blepharoptosis morphology
Shuaixuan Zhou, Xingru Huang, Zhaoyang Xu, Huiyu Zhou 0001, Guangyuan Zhang, Wenwen Tang, Wenbin Zhang 0002, Jin Liu 0025, Lixia Lou, Xiaoshuai Zhang
Expert Syst. Appl.6
2026 A fast gray-box adversarial example generation algorithm based on FakeBob
abstract
There are the excessive queries to the targeted model during the generates of gray-box adversarial examples for speaker recognition systems, which result in high costs of attacks. In this paper, a fast generates algorithm of gray-box adversarial example is proposed based on FakeBob, named F-FakeBob. This algorithm introduces a threshold mechanism for optimization to the optimization strategy of gradient. Only when the increasing of the confidence scores of the adversarial example before and after optimizing is less than the threshold, the gradient is recalculated for the next iteration. By reducing the frequency of gradient calculations, the number of queries to the targeted system is decreased. Experiments on three public datasets of speech, TIMIT, Common Voice, and Voxceleb2, are conducted to generate adversarial examples. The targeted speaker recognition models are based on ECAPA-TDNN and TitaNet architectures. The experimental results show that F-FakeBob can achieve a targeted attack success rate of 99.2% and the numbers of queries are effectively reduced in the adversarial example generates, with an average query reduction of 25.71% compared to FakeBob.
Wanjin Hou, Hua Zhang 0001, Ming Lv, Huiyu Zhou 0001
High Confid. Comput.5
2026 Adversarial batch representation augmentation for batch correction in high-content cellular screening
abstract
• Proposes an Adversarial Batch Representation Augmentation for batch correction. • Models uncertainty of biological batch effects in representation learning. • Uses adversarial learning to identify challenges in the objective function. • Presents a synergistic optimization process for stable training. • Comprehensive experiments validate the effectiveness of the proposed method. High-Content Screening routinely generates massive volumes of cell painting images for phenotypic profiling. However, technical variations across experimental executions inevitably induce biological batch (bio-batch) effects. These cause covariate shifts and degrade the generalization of deep learning models on unseen data. Existing batch correction methods typically rely on additional prior knowledge (e.g., treatment or cell culture information) or struggle to generalize to unseen bio-batches. In this work, we frame bio-batch mitigation as a Domain Generalization (DG) problem and propose Adversarial Batch Representation Augmentation (ABRA). ABRA explicitly models batch-wise statistical fluctuations by parameterizing feature statistics as structured uncertainties. Through a min-max optimization framework, it actively synthesizes worst-case bio-batch perturbations in the representation space, guided by a strict angular geometric margin to preserve fine-grained class discriminability. To prevent representation collapse during this adversarial exploration, we introduce a synergistic distribution alignment objective. Extensive evaluations on the large-scale RxRx1 and RxRx1-WILDS benchmarks demonstrate that ABRA establishes a new state-of-the-art for siRNA perturbation classification.
Xujing Yao, Adam Corrigan, Long Chen 0019, Navin Rathna Kumar, Kerry Hallbrook, Jonathan Orme, Yinhai Wang, Huiyu Zhou 0001
Knowl. Based Syst.9
2026 MICCAI STS 2024 challenge: Semi-supervised instance-level tooth segmentation in panoramic X-ray and CBCT images
abstract
Orthopantomogram (OPGs) and Cone-Beam Computed Tomography (CBCT) are vital for dentistry, but creating large datasets for automated tooth segmentation is hindered by the labor-intensive process of manual instance-level annotation. This research aimed to benchmark and advance semi-supervised learning (SSL) as a solution for this data scarcity problem. We organized the 2nd Semi-supervised Teeth Segmentation (STS 2024) Challenge at MICCAI 2024. We provided a large-scale dataset comprising over 90,000 2D images and 3D axial slices, which includes 2380 OPG images and 330 CBCT scans, all featuring detailed instance-level FDI annotations on part of the data. The challenge attracted 114 (OPG) and 106 (CBCT) registered teams. To ensure algorithmic excellence and full transparency, we rigorously evaluated the valid, open-source submissions from the top 10 (OPG) and top 5 (CBCT) teams, respectively. All successful submissions were deep learning-based SSL methods. The winning semi-supervised models demonstrated impressive performance gains over a fully-supervised nnU-Net baseline trained only on the labeled data. For the 2D OPG track, the top method improved the Instance Affinity (IA) score by over 44 percentage points. For the 3D CBCT track, the winning approach boosted the Instance Dice score by 61 percentage points. This challenge demonstrates the potential benefit benefit of SSL for complex, instance-level medical image segmentation tasks where labeled data is scarce. The most effective approaches consistently leveraged hybrid semi-supervised frameworks that combined knowledge from foundational models like SAM with multi-stage, coarse-to-fine refinement pipelines. Both the challenge dataset and the participants' submitted code have been made publicly available on GitHub (https://github.com/ricoleehduu/STS-Challenge-2024), ensuring transparency and reproducibility.
Yaqi Wang 0002, Jun Liu 0027, Jiaxue Ni, Hongyuan Zhang 0002, Jin Liu 0025, Can Han, Kaiwen Fu, Changkai Ji, Xinxu Cai, Junqiang Chen, Qianni Zhang, Dahong Qian, Shuai Wang 0003, Huiyu Zhou 0001
Medical Image Anal.23
2026 TriFTM-Net: Tri-Path Fourier-Temporal Modulation Network for macular edema pathology segmentation and reconstruction in high-precision intraoperative navigation
abstract
Ophthalmic diseases such significantly impair the vision of numerous individuals globally. Accurate and real-time 3D reconstruction of macular edema and retinal tears is crucial for improving surgical efficiency and success rates. However, lesion areas often exhibit considerable noise and high heterogeneity, and the imaging devices employed may introduce electronic noise and artifacts. Current 2D medical image segmentation techniques fail to achieve optimal outcomes. To overcome these challenges, we propose the Tri-Path Fourier-Temporal Modulation Network (TriFTM-Net). TriFTM-Net synergistically integrates spatial, frequency, and spatiotemporal features. This design effectively augments both feature representation and extraction. TriFTM-Net comprises three critical modules: the Tri-Path Spectral Hierarchical Encoder (TPSHE), which amplifies feature representation by integrating tri-path features; the Feature Re-Modulation (FRM), which reduces noise interference and enhances feature extraction; and the Hierarchical Feature Reconstruction Module (HFRM), which improves detail preservation in upsampled images. Comparative analysis with thirteen baseline methods demonstrates that our approach achieves the highest Dice scores, IoU, and Kappa coefficient on the OIMHS dataset.Our code is publicly available at https://github.com/IMOP-lab/TriFTM-Net.
Xingru Huang, Shuaibin Chen, Gaopeng Huang, Zhaoyang Xu, Wenbin Zhang 0002, Jian Huang 0015, Jin Liu 0025, Xiaoshuai Zhang, Shaowei Jiang, Huiyu Zhou 0001, Yaoqi Sun
Neural Networks14
2026 PolyS-Net: A joint learning framework for depth-aware and scale-aware polyp size estimation
Sijia Du, Yaqi Wang 0002, Chen Liu 0026, Jun Wang 0041, Ruilan Wang, Huiyu Zhou 0001, Qingwei Zhang, Dahong Qian
Pattern Recognit.7
2026 INSERTION: From traditional incremental learning to open-world stream learning
Yanchao Li 0001, Hongwei Dou, Guanxiao Li, Guangwei Gao, Huiyu Zhou 0001
Pattern Recognit.5
2026 Unsupervised cross-domain semantic segmentation on multi-modality ovarian tumor ultrasound data
Shuchang Lyu, Qi Zhao 0037, Wenpei Bai, Linghan Cai, Guangxia Cui, Lijiang Chen, Huiyu Zhou 0001
Pattern Recognit.9
2026 MorVess: Morphology-aware pulmonary vessel segmentation network
Fuyou Mao, Yifei Chen 0019, Beining Wu, Lixin Lin, Jinnan Dai, Zhiling Li, Huiyu Zhou 0001, Fei-wei Qin
Pattern Recognit.11
2026 DiffClick: Click-differentiated enhancement network for interactive segmentation
Siqi Song, Siyue Yu, Huiyu Zhou 0001, Xiaowei Huang 0001, Limin Yu, Jimin Xiao
Pattern Recognit.3
2026 MICCAI 2023 STS Challenge: A retrospective study of semi-supervised approaches for teeth segmentation
abstract
Computer-aided diagnosis greatly enhances personalized treatment planning and diagnostic efficiency by providing accurate dental anatomy through teeth segmentation. However, it still constrained by the scarcity of high-quality annotated dental datasets. To address this issue, this paper presents a dataset combining both 2D panoramic X-rays with over 6,500 images and 3D CBCT with over 580 volumes (88,500+ slices) to support the Semi-supervised Teeth Segmentation (STS) Challenge, which includes partially meticulous annotations and covers all age groups. Moreover, multi-phase semi-supervised teeth segmentation algorithms and high-confidence pseudo-labels refinement strategies were proposed by competitors during this challenge. Algorithms were verified on this proposed dataset and good segmentation performance were achieved, over 93+ and 80+ Dice score were obtained for top three 2D and 3D participants, demonstrating the high quality of this proposed dataset. This paper also summarizes the diverse methods employed by the top-ranking teams in the MICCAI 2023 STS Challenge. Our dataset is publicly accessible through Zenodo ( https://zenodo.org/records/10597292 ), and the participants’ code is hosted on GitHub ( https://github.com/ricoleehduu/STS-Challenge ).
Yaqi Wang 0002, Shuai Wang 0003, Dahong Qian, Hongyuan Zhang 0002, Ruilong Dan, Qianni Zhang, Xingru Huang, Jun Liu 0027, Zhean Ma, Weiwei Cui 0003, Shan Luo 0003, Chengkai Wang, Jiaxue Ni, Dongyun Liu, Zhouhao Lin, Chunshi Wang, Qiupu Chen, Mingqian Li, Huiyu Zhou 0001, Qun Jin
Pattern Recognit.36
2026 Boundary-aware shape recognition using dynamic graph convolutional networks
Jinming Zhao, Junyu Dong, Huiyu Zhou 0001, Xinghui Dong
Pattern Recognit.3
2026 Joint Location and Velocity Estimation and Fundamental CRLB Analysis for Cell-Free MIMO-ISAC
abstract
This paper presents a fundamental performance analysis of joint location and velocity estimation in a cell-free (CF) MIMO integrated sensing and communication (ISAC) system. Unlike prior studies that primarily rely on continuous-time signal models, we consider a more practical and challenging scenario in the discrete-time digital domain. Specifically, we first formulate a logarithmic likelihood function (LLF) and corresponding maximum likelihood estimation (MLE) for both single- and multiple-target sensing. Building upon the proposed LLF framework, closed-form Cramer-Rao lower bounds (CRLBs) for joint location and velocity estimation are derived under deterministic, unknown, and spatially varying radar cross-section (RCS) models. These CRLBs can serve as a fundamental performance metric to guide CF MIMO-ISAC system design. To enhance tractability, we also develop a class of simplified closed-form CRLBs, referred to as approximate CRLBs, along with a rigorous analysis of the conditions under which they remain accurate. Furthermore, we investigate how the sampling rate, squared effective bandwidth, and time width influence CRLB performance. For multi-target scenarios, the concepts of safety distance and safety velocity are introduced to characterize the conditions under which the CRLBs converge to their single-target counterparts. Extensive simulations using orthogonal frequency division multiplexing (OFDM) and orthogonal chirp division multiplexing (OCDM) validate the theoretical findings and provide practical insights for CF MIMO-ISAC system design
Guoqing Xia, Pei Xiao 0001, Qu Luo, Bing Ji 0003, Yue Zhang 0011, Huiyu Zhou 0001
IEEE Trans. Commun.6
2026 Closed-Form BER Analysis for Uplink NOMA With Dynamic SIC Decoding
abstract
This paper, for the first time, presents a closed-form error performance analysis of uplink power-domain non-orthogonal multiple access (PD-NOMA) with dynamic successive interference cancellation (SIC) decoding, where the decoding order is adapted to the instantaneous channel conditions. We first develop an analytical framework that characterizes how dynamic ordering affects error probabilities in uplink PD-NOMA systems. For a two-user system over independent and non-identically distributed Rayleigh fading channels, we derive closed-form probability density functions (PDFs) of ordered channel gains and the corresponding unconditional pairwise error probabilities (PEPs). To address the mathematical complexity of characterizing ordered channel distributions, we employ a Gaussian fitting to approximate truncated distributions while maintaining analytical tractability. Finally, we extend the bit error rate analysis for various $M$-quadrature amplitude modulation schemes (QAM) in both homogeneous and heterogeneous scenarios. Numerical results validate the theoretical analysis and demonstrate that dynamic SIC eliminates the error floor issue observed in fixed-order SIC, achieving significantly improved performance in high signal-to-noise ratio regions. Our findings also highlight that larger power differences are essential for higher-order modulations, offering concrete guidance for practical uplink PD-NOMA deployment.
Hequn Zhang, Qu Luo, Pei Xiao 0001, Yue Zhang 0011, Huiyu Zhou 0001
IEEE Trans. Commun.5
2026 Double Uncertainty-Aware Learning Network for Multi-Modal Cell Image Segmentation
abstract
To perform segmentation for the cell images of different modalites accurately, we should address issues of over-segmentation or under-segmentation caused by uncertainty variations in modal pixel distribution and cell morphology. Moreover, the problem of limited labeled data is studied in this work. Most previous methods lack global uncertainty information perception ability, can not obtain local uncertainty details, and are limited to the number of labeled data from multi-modal cell images. We introduce a novel framework that can accurately learn valuable information for multi-modal cell segmentation task with the data and modal uncertainty aware abilities. Firstly, an image fusion module is proposed that leverages a multi-branch structure, incorporating dilation convolution, regular convolution, and channel attention mechanism for saving global valuable information. Secondly, to obtain local boundaries from obscure and irregular uncertainty regions, a transformer-based model encoding strategy is developed for performing token selection and enhancement based on the feedback confidence score. This confidence score is computed based on the output of a teaching network that indicates most likely local boundary. Thirdly, a pseudo-label selection strategy is employed to improve the annotation quality of unlabeled images. We evaluated our method on three publicly available datasets with different cell modalites and performed a quantitative comparison with the previous fifteen methods. Our method achieved better performance than others. This study has important implications for the improvements of clinical applications including diagnostic accuracy and decision-making reliability.
Jinzhao Yang, Kuan Li, Weiping Ding 0001, Huiyu Zhou 0001
IEEE J. Biomed. Health Informatics5
2025 Multi-Label Transfer Learning in Non-Stationary Data Streams
abstract
Label concepts in multi-label data streams often experience drift in non-stationary environments, either independently or in relation to other labels. Transferring knowledge between related labels can accelerate adaptation, yet research on multi-label transfer learning for data streams remains limited. To address this, we propose two novel transfer learning methods: BR-MARLENE leverages knowledge from different labels in both source and target streams for multi-label classification; BRPW-MARLENE builds on this by explicitly modelling and transferring pairwise label dependencies to enhance learning performance. Comprehensive experiments show that both methods outperform state-of-the-art multi-label stream approaches in non-stationary environments, demonstrating the effectiveness of inter-label knowledge transfer for improved predictive performance. The implementation is available at https://github.com/nino2222/MARLENE.
Honghui Du, Leandro L. Minku, Aonghus Lawlor, Huiyu Zhou 0001
ICDM4
2025 HRGR: Enhancing Image Manipulation Detection via Hierarchical Region-aware Graph Reasoning
abstract
Image manipulation detection is to identify the authenticity of each pixel in images. One typical approach to uncover manipulation traces is to model image correlations. The previous methods commonly adopt the grids, which are fixed-size squares, as graph nodes to model correlations. However, these grids, being independent of image content, struggle to retain local content coherence, resulting in imprecise detection. To address this issue, we describe a new method named Hierarchical Region-aware Graph Reasoning (HRGR) to enhance image manipulation detection. Unlike existing grid-based methods, we model image correlations based on content-coherence feature regions with irregular shapes, generated by a novel Differentiable Feature Partition strategy. Then we construct a Hierarchical Region-aware Graph based on these regions within and across different feature layers. Subsequently, we describe a structural-agnostic graph reasoning strategy tailored for our graph to enhance the representation of nodes. Our method is fully differentiable and can seamlessly integrate into mainstream networks in an end-to-end manner, without requiring additional supervision. Extensive experiments demonstrate the effectiveness of our method in image manipulation detection, exhibiting its great potential as a plug-and-play component for existing architectures. Codes and models are available at https://github.com/OUC-VAS/HRGR-IMD.
Jiaran Zhou, Huiyu Zhou 0001, Junyu Dong, Yuezun Li
ICME3
2025 Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation
abstract
Open-set image segmentation poses a significant challenge because existing methods often demand extensive training or fine-tuning and generally struggle to segment unified objects consistently across diverse text reference expressions. Motivated by this, we propose Segment Anyword, a novel training-free visual concept prompt learning approach for open-set language grounded segmentation that relies on token-level cross-attention maps from a frozen diffusion model to produce segmentation surrogates or *mask prompts*, which are then refined into targeted object masks. Initial prompts typically lack coherence and consistency as the complexity of the image-text increases, resulting in suboptimal mask fragments. To tackle this issue, we further introduce a novel linguistic-guided visual prompt regularization that binds and clusters visual prompts based on sentence dependency and syntactic structural information, enabling the extraction of robust, noise-tolerant mask prompts, and significant improvements in segmentation accuracy. The proposed approach is effective, generalizes across different open-set segmentation tasks, and achieves state-of-the-art results of 52.5 (+6.8 relative) mIoU on Pascal Context 59, 67.73 (+25.73 relative) cIoU on gRefCOCO, and 67.4 (+1.1 relative to fine-tuned methods) mIoU on GranDf, which is the most complex open-set grounded segmentation task in the field.
Amrutha Saseendran, Xilin He, Fariba Yousefi, Nikolay Burlutskiy, Dino Oglic, Tom Diethe, Philip Teare, Huiyu Zhou 0001
ICML10
2025 APGNet: Adaptive Prior-Guided for Underwater Camouflaged Object Detection
abstract
Detecting camouflaged objects in underwater environments is crucial for marine ecological research and resource exploration. However, existing methods face two key challenges: underwater image degradation, including low contrast and color distortion, and the natural camouflage of marine organisms. Traditional image enhancement techniques struggle to restore critical features in degraded images, while camouflaged object detection (COD) methods developed for terrestrial scenes often fail to adapt to underwater environments due to the lack of consideration for underwater optical characteristics. To address these issues, we propose APGNet, an Adaptive Prior-Guided Network, which integrates a Siamese architecture with a novel prior-guided mechanism to enhance robustness and detection accuracy. First, we employ the Multi-Scale Retinex with Color Restoration (MSRCR) algorithm for data augmentation, generating illumination-invariant images to mitigate degradation effects. Second, we design an Extended Receptive Field (ERF) module combined with a Multi-Scale Progressive Decoder (MPD) to capture multi-scale contextual information and refine feature representations. Furthermore, we propose an adaptive prior-guided mechanism that hierarchically fuses position and boundary priors by embedding spatial attention in high-level features for coarse localization and using deformable convolution to refine contours in low-level features. Extensive experimental results on two public MAS datasets demonstrate that our proposed method APGNet outperforms 15 state-of-art methods under widely used evaluation metrics.
Xinxin Huang, Junmin Cai, Ningzhong Liu, Huiyu Zhou 0001
MMAsia5
2025 SliceSemOcc: Vertical Slice-Based Multimodal 3D Semantic Occupancy Representation
Ningzhong Liu, Huiyu Zhou 0001, Jiaquan Shen
PRCV (10)4
2025 SLENet: A Guidance-Enhanced Network for Underwater Camouflaged Object Detection
Xinxin Huang, Ningzhong Liu, Huiyu Zhou 0001, Yinan Yao
PRCV (16)4
2025 DCFS: Continual Test-Time Adaptation via Dual Consistency of Feature and Sample
Wenting Yin, Xinru Meng, Ningzhong Liu, Huiyu Zhou 0001
PRCV (12)5
2025 Deep Content and Contrastive Perception learning for automatic fetal nuchal translucency image quality assessment
Weiping Ding 0001, Jinzhao Yang, Huiyu Zhou 0001, Yiming Du, Bin Hu 0023, Lichi Zhang, Qian Wang 0001
Eng. Appl. Artif. Intell.6
2025 DUAL: A Dual-Stage Approach for Facial Expression Recognition Based on Contrastive Learning
abstract
Facial expression recognition (FER) remains a challenging task in computer vision. Recent works have shown excellent performance in overall recognition accuracy, but its accuracy significantly decreases when recognizing similar expressions. This is due to interclass homogeneity and intraclass heterogeneity. To address these issues, we propose a novel dual‐stage network called DUAL, inspired by contrastive learning. First, we increase the distance between negative samples while reducing the distance between positive ones. This is achieved by dynamically updating pairs of comparison samples. Second, we introduce a two‐stage network architecture. The first stage uses two branches to extract image features and facial keypoint features. These branches interact to learn coarse‐grained features through mutual guidance. The second stage focuses on fine‐grained features using scale‐specific residual blocks. This allows the model to identify facial regions that are critical for recognizing expressions. We conducted extensive experiments on multiple datasets. The results show that DUAL surpasses state‐of‐the‐art models in items of performance. Additionally, the model shows high accuracy even in noisy conditions, highlighting its robustness.
Anting Zhu, Xingxing Jia, Longfei Yang, Huiyu Zhou 0001, Wei Su 0008
Int. J. Intell. Syst.4
2025 Population-Based Meta-Heuristic Optimization Algorithm Booster: An Evolutionary and Learning Competition Scheme
Jun Wang 0041, Junyu Dong, Huiyu Zhou 0001, Xinghui Dong
Neurocomputing3
2025 Surface Multiple Object Tracking: An Accurate HAT-YOLOv8-ADT Tracking Model
abstract
With the development of artificial intelligence technology, Autonomous aerial vehicles (AAV) have the ability to sense the environment. multiple object tracking (MOT) in AAV video is a very important vision task with a wide variety of applications. However, there are still many challenges in MOT in AAV video. First, the movement of the onboard camera in the three-dimensional (3-D) direction during the tracking process, as well as the unpredictable measurement noise characteristics of AAVs flying at high speeds, can lead to significant deviations in the prediction of the object’s position. Second, the applicability of the traditional detection algorithm decreases when the object is small and dense in the AAV viewpoint during detection. Finally, the traditional intersection over union (IoU) matching approach does not take into account the effects of the height and width of the box, and the matching results are inaccurate for the prediction and detection box. In order to address these challenges, we recommend the adaptive DeepSort (ADT) algorithm to reduce the prediction bias due to camera movement and difficulty in predetermining measurement noise characteristics, the hybrid attention transformer-YOLOv8 (HAT-YOLOv8) algorithm to enhance the detection capability of tiny objects, and the IoU of height and width (HWIoU) matching algorithm, which improves the matching accuracy and thus the tracking accuracy. Experimental results show that our proposed solution outperforms the baseline solution. It outperforms the current mainstream StrongSort in MOTA, HOTA and IDF1 by 2.86%, 0.9%, and 9.36%. Code repository link:https://github.com/networkcommunication/.
Na Lin 0001, Lei Zhang 0036, Tianxiong Wu, Ammar Hawbani, Huiyu Zhou 0001, Liang Zhao 0004
IEEE Internet Things J.5
2025 PricoMS: Prior-coordinated multiscale synthesis network for self-supervised-aided vessel segmentation in intravascular ultrasound image amidst label scarcity
Xingru Huang, Shuaibin Chen, Shaowei Jiang, Retesh Bajaj, Nathan Angelo Lecaros Yap, Murat Çap, Xiaoshuai Zhang, Xingwei He 0007, Anantharaman Ramasamy, Ryo Torii, Jouke Dijkstra, Huiyu Zhou 0001, Christos V. Bourantas, Qianni Zhang
Knowl. Based Syst.13
2025 Binary Banyan tree growth optimization: A practical approach to high-dimensional feature selection
abstract
High-dimensional feature spaces in Scientific and Technical Service Resources (STSR) classification present significant challenges, including increased computational costs and diminished accuracy. Identifying an optimal subset of features from raw text vectors is thus critical for effective data classification . This paper introduces a novel metaheuristic algorithm called Binary Banyan Tree Growth Optimization (BBTGO), specifically designed for high-dimensional feature selection (FS). Inspired by the unique growth patterns of the banyan tree , BBTGO leverages a combination of innovative Boolean vectors, including rooting, multi-trunk, and adjustment operator, along with a perturbation phase to enhance the search efficiency and reduce feature dimensionality. These operators enhance the search for promising regions and reduce features by utilizing the optimal solutions clustered within subgroups. Furthermore, BBTGO incorporates a dynamic adjustment mechanism that periodically activates different growth operators to meet the search demands of high-dimensional space. We rigorously evaluate the exploration and exploitation capabilities of BBTGO through comprehensive statistical analyses of various performance metrics. The proposed method demonstrates superior results on 12 high-dimensional benchmark datasets and is successfully applied to feature selection in STSR text classification tasks . Experimental results show that BBTGO significantly outperforms existing methods in terms of classification accuracy , selected features, convergence speed, and processing time. These results underscore the potential of BBTGO as a robust and versatile solution for high-dimensional FS, with broad applicability to real-world classification challenges.
Minrui Fei, Wenju Zhou, Songlin Du, Zixiang Fei, Huiyu Zhou 0001
Knowl. Based Syst.6
2025 Scattering Characteristics Guided Network for ISAR Space Target Component Segmentation
abstract
Affected by the large dynamic range of gray values, strong scattering point edge effect, noise and clutter, inverse synthetic aperture radar (ISAR) images have problems such as boundary blurring and target discontinuity, which bring great challenges to ISAR space target component segmentation. In this paper, a novel ISAR space target component segmentation method, called scattering characteristics guided network (SCGN), is proposed. First, a cross-scale self-attention module (CSSAM) is proposed, which establishes global relationships in different dimensions during cross-scale feature fusion, refining the detailed features of the target while suppressing high sidelobe scattering points and noise. Second, a novel component scattering center extractor (CSCE) is proposed to combine scattering center distribution with the network via explicit supervision. Finally, a novel scattering characteristics-assisted segmentation head (SCASH) is proposed, which introduces the scattering characteristics of each component into the mask segmentation process and models the semantic interdependencies over long distances through a spatial attention mechanism to achieve fine-grained component segmentation. Experimental results on the ISAR simulation dataset and realistic ISAR images show that SCGN outperforms existing methods.
Fengjun Zhong, Fei Gao 0005, Tianjin Liu, Jun Wang 0041, Jinping Sun, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.6
2025 A triple-branch hybrid dynamic-static alignment strategy for vision-language tasks
Xiang Shen 0002, Chongqing Chen, Dezhi Han, Yangshuyi Xu, Xiuying Wang 0001, Huiyu Zhou 0001
Neural Networks6
2025 From Missing Pieces to Masterpieces: Image Completion With Context-Adaptive Diffusion
abstract
Image completion is a challenging task, particularly when ensuring that generated content seamlessly integrates with existing parts of an image. While recent diffusion models have shown promise, they often struggle with maintaining coherence between known and unknown (missing) regions. This issue arises from the lack of explicit spatial and semantic alignment during the diffusion process, resulting in content that does not smoothly integrate with the original image. Additionally, diffusion models typically rely on global learned distributions rather than localized features, leading to inconsistencies between the generated and existing image parts. In this work, we propose ConFill, a novel framework that introduces a Context-Adaptive Discrepancy (CAD) model to ensure that intermediate distributions of known and unknown regions are closely aligned throughout the diffusion process. By incorporating CAD, our model progressively reduces discrepancies between generated and original images at each diffusion step, leading to contextually aligned completion. Moreover, ConFill uses a new Dynamic Sampling mechanism that adaptively increases the sampling rate in regions with high reconstruction complexity. This approach enables precise adjustments, enhancing detail and integration in restored areas. Extensive experiments demonstrate that ConFill outperforms current methods, setting a new benchmark in image completion.
Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Michael Felsberg, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Energy-based pseudo-label refining for source-free domain adaptation
Xinru Meng, Jiamei Liu, Ningzhong Liu, Huiyu Zhou 0001
Pattern Recognit. Lett.5
2025 CF3d: Category fused 3D point cloud retrieval
Zongyi Xu, Ruicheng Zhang, Zuo Li, Shiyang Cheng 0001, Huiyu Zhou 0001, Weisheng Li 0001, Xinbo Gao 0001
Signal Process.5
2025 Terrain Scene Generation Using a Lightweight Vector Quantized Generative Adversarial Network
abstract
Natural terrain scene images play important roles in the geographical research and application. However, it is challenging to collect a large set of terrain scene images. Recently, great progress has been made in image generation. Although impressive results can be achieved, the efficiency of the state-of-the-art methods, e.g., the Vector Quantized Generative Adversarial Network (VQGAN), is still dissatisfying. The VQGAN confronts two issues, i.e., high space complexity and heavy computational demand. To efficiently fulfill the terrain scene generation task, we first collect a Natural Terrain Scene Data Set (NTSD), which contains 36,672 images divided into 38 classes. Then we propose a Lightweight VQGAN (Lit-VQGAN), which uses the fewer parameters and has the lower computational complexity, compared with the VQGAN. A lightweight super-resolution network is further adopted, to speedily derive a high-resolution image from the image that the Lit-VQGAN generates. The Lit-VQGAN can be trained and tested on the NTSD. To our knowledge, either the NTSD or the Lit-VQGAN has not been exploited before.1Experimental results show that the Lit-VQGAN is more efficient and effective than the VQGAN for the image generation task. These promising results should be due to the lightweight yet effective networks that we design.
Huiyu Zhou 0001, Xinghui Dong
IEEE Trans. Big Data2
2025 TBRNet: Two-Branch Reinforcement Network for Few-Shot Semantic Segmentation in Remote Sensing Images
abstract
Traditional semantic segmentation methods for remote sensing images (RSIs) require abundant labeled data yet falter when samples are scarce. Few-shot semantic segmentation (FSS) innovatively resolves this data scarcity bottleneck. However, existing FSS models face unique challenges in RSIs: natural-to-remote-sensing domain gaps, intraclass variance, and multicategory coexistence-induced generalization collapse. These problems may lead to the loss of discriminative query features in model performance. In this article, we present a novel two-branch reinforcement network called TBRNet to tackle these challenges. Specifically, we first propose a prototype reinforcement module (PRM) to generate enhanced context-adaptive prototypes by dynamically weighting query-support feature contexts, which effectively mitigates intraclass variance via strengthening support images’ perception of discriminative query features. In addition, to deal with the challenge of the coexistence of multiple target categories, we develop a multilevel guidance reinforcement module (MGRM), which provides multilevel guidance maps across resolutions to model cross-level semantic dependencies and emphasize discriminative subregions. Extensive experiments conducted on the iSAID-$5^{i}$and DLRSD-$5^{i}$remote sensing (RS) datasets have shown the superiority of our proposed TBRNet compared with several state-of-the-art approaches. Ablation studies have also verified the effectiveness of the proposed modules. Our source code is made publicly available athttps://github.com/WangXin81/TBRNet
Xin Wang 0068, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Temporal-Feedback Self-Training for Semi-Supervised Object Detection in Remote Sensing Images
abstract
Although modern Remote Sensing Object Detection (RSOD) methods have achieved advanced performance, they heavily rely on a large amount of annotated data. This paper explores semi-supervised RSOD to mitigate annotation costs, leveraging recent extensive research in generic Semi-Supervised Object Detection (SSOD) based on the self-training paradigm. Current SSOD methods encounter challenges in adapting to remote sensing images due to the complexity and variability of RSIs. Two key issues remain underexplored: the noise in pseudo-labels caused by model instability and the difficulty in distinguishing similar categories. This paper introduces the Temporal-Feedback Self-Training (TST) framework, a novel approach to tackle these challenges in semi-supervised RSOD. TST consists of two components: Temporal Consistency Based Pseudo-labels Certainty Estimation (TCE) and Temporal Self-Feedback Feature Refinement (TSF). TCE addresses pseudo-label noise during training by evaluating the stability of pseudo-label classification and localization over time series to assess the quality of pseudo-labels. On the other hand, TSF enhances pseudo-label quality by dynamically identifying the models confusing categories as feedback for feature refinement. Both components facilitate the progression of the self-training-based RSOD during training. We conducted extensive experiments on two challenging public datasets, DOTA and DIOR. The results demonstrate that the proposed TST and TCE components significantly improve the baseline models performance, surpassing the state-of-the-art generic SSOD method. This suggests that our approach is more effective than generic SSOD methods in addressing the challenges posed by remote sensing images.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2025 F2Attack: Two-Factors Scoring Method for Query-Efficient Hard-Label Black-Box Textual Adversarial Attacks
Hua Zhang 0001, Qi Li 0057, Huiyu Zhou 0001
IEEE Trans. Inf. Forensics Secur.6
2025 Cross-Skeleton Interaction Graph Aggregation Network for Representation Learning of Mouse Social Behavior
abstract
Automated social behaviour analysis of mice has become an increasingly popular research area in behavioural neuroscience. Recently, pose information (i.e., locations of keypoints or skeleton) has been used to interpret social behaviours of mice. Nevertheless, effective encoding and decoding of social interaction information underlying the keypoints of mice has been rarely investigated in the existing methods. In particular, it is challenging to model complex social interactions between mice due to highly deformable body shapes and ambiguous movement patterns. To deal with the interaction modelling problem, we here propose a Cross-Skeleton Interaction Graph Aggregation Network (CS-IGANet) to learn abundant dynamics of freely interacting mice, where a Cross-Skeleton Node-level Interaction module (CS-NLI) is used to model multi-level interactions (i.e., intra-, inter- and cross-skeleton interactions). Furthermore, we design a novel Interaction-Aware Transformer (IAT) to dynamically learn the graph-level representation of social behaviours and update the node-level representation, guided by our proposed interaction-aware self-attention mechanism. Finally, to enhance the representation ability of our model, an auxiliary self-supervised learning task is proposed for measuring the similarity between cross-skeleton nodes. Experimental results on the standard CRMI13-Skeleton and our PDMB-Skeleton datasets show that our proposed model outperforms several other state-of-the-art approaches.
Feixiang Zhou, Long Chen 0019, Zheheng Jiang, Reiko Heckel, Haikuan Wang, Minrui Fei, Huiyu Zhou 0001
IEEE Trans. Image Process.10
2025 A Novel Multi-Modal Population-Graph Based Framework for Patients of Esophageal Squamous Cell Cancer Prognostic Risk Prediction
abstract
Prognostic risk prediction is pivotal for clinicians to appraise the patient's esophageal squamous cell cancer (ESCC) progression status precisely and tailor individualized therapy treatment plans. Currently, CT-based multi-modal prognostic risk prediction methods have gradually attracted the attention of researchers for their universality, which is also able to be applied in scenarios of preoperative prognostic risk assessment in the early stages of cancer. However, much of the current work focuses only on CT images of the primary tumor, ignoring the important role that CT images of lymph nodes play in prognostic risk prediction. Additionally, it is important to consider and explore the inter-patient feature similarity in prognosis when developing models. To solve these problems, we proposed a novel multi-modal population-graph based framework leveraging CT images including primary tumor and lymph nodes combined with clinical, hematology, and radiomics data for ESCC prognostic risk prediction. A patient population graph was constructed to excavate the homogeneity and heterogeneity of inter-patient feature embedding. Moreover, a novel node-level multi-task joint loss was proposed for graph model optimization through a supervised-based task and an unsupervised-based task. Sufficient experimental results show that our model achieved state-of-the-art performance compared with other baseline models as well as the gold standard on discriminative ability, risk stratification, and clinical utility.
Shuai Wang 0003, Yaqi Wang 0002, Chengkai Wang, Huiyu Zhou 0001, Yatao Zhang
IEEE J. Biomed. Health Informatics5
2025 MIP-CLIP: Multimodal Independent Prompt CLIP for Action Recognition
abstract
Recently, the Contrastive Language Image Pre-training (CLIP) model has shown significant generalizability by optimizing the distance between visual and text features. The mainstream CLIP-based action recognition methods mitigate the low “zero-shot” generalization of the 1-of-N paradigm but also lead to a significant degradation in supervised performance. Therefore, powerful supervision and competitive “zero-shot” need to be effectively traded off. In this work, a Multimodal Independent Prompt CLIP (MIP-CLIP) model is proposed to address this challenge. On the visual side, we propose novel Video Motion Prompt (VMP) to empower the visual encoder with motion perception, which performs short- and long-term motion modelling via temporal difference operation. Next, the visual classification branch is introduced to improve the discrimination of visual features. Specifically, the temporal difference and visual classification operations of the 1-of-N paradigm are extended to CLIP to satisfy the need for strong supervised performance. On the text side, we design Class-Agnostic text prompt Template (CAT) under the constraint of Semantic Alignment (SA) module to solve the label semantic dependency problem. Finally, a Dual-branch Feature Reconstruction (DFR) module is proposed to complete cross-modal interactions for better feature matching, which uses the class confidence of the visual classification branch as input. The experiments are conducted on four widely used benchmarks (HMDB-51, UCF-101, Jester, and Kinetics-400). The results demonstrate that our method achieves excellent supervised performance while preserving competitive generalizability.
Xiong Gao, Zhaobin Chang, Dongyi Kong, Huiyu Zhou 0001, Yonggang Lu
IEEE Trans. Multim.4
2025 Negative Deterministic Information-Based Multiple Instance Learning for Weakly Supervised Object Detection and Segmentation
abstract
Weakly supervised object detection (WSOD) and semantic segmentation with image-level annotations have attracted extensive attention due to their high label efficiency. Multiple instance learning (MIL) offers a feasible solution for the two tasks by treating each image as a bag with a series of instances (object regions or pixels) and identifying foreground instances that contribute to bag classification. However, conventional MIL paradigms often suffer from issues, e.g., discriminative instance domination and missing instances. In this article, we observe that negative instances usually contain valuable deterministic information, which is the key to solving the two issues. Motivated by this, we propose a novel MIL paradigm based on negative deterministic information (NDI), termed NDI-MIL, which is based on two core designs with a progressive relation: NDI collection and negative contrastive learning (NCL). In NDI collection, we identify and distill NDI from negative instances online by a dynamic feature bank. The collected NDI is then utilized in a NCL mechanism to locate and punish those discriminative regions, by which the discriminative instance domination and missing instances issues are effectively addressed, leading to improved object- and pixel-level localization accuracy and completeness. In addition, we design an NDI-guided instance selection (NGIS) strategy to further enhance the systematic performance. Experimental results on several public benchmarks, including PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO, show that our method achieves satisfactory performance. The code is available at: https://github.com/GC-WSL/NDI.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.6
2025 SeSMR: Secure and Efficient Session-Based Multimedia Recommendation in Edge Computing
abstract
Session-based multimedia recommendation in edge computing remains an important issue for boosting the utilization of services since service composition has increasingly attracted attention. Existing session-based recommendations (SBRs) model the session sequence with multilevel feature extraction in graph neural networks (GNNs). However, multilevel feature extraction in disentangled graph neural networks causes over-smoothing and privacy leakage. To address the aforementioned problems, Secure and Efficient Session-based Multimedia Recommendation (SeSMR) model is proposed. In the proposed SeSMR model, based on BGV homomorphic encryption, a ciphertext training submodel is proposed to address the privacy leakage, ensuring the security in SBR. Furthermore, based on the reinforcement of feature activation, a residual attention mechanism is proposed to mitigate over-smoothing while maintaining the independence of multiple features. Finally, based on location coding, a soft attention mechanism is proposed to improve the recommendation accuracy, by introducing the position difference information between items into intra-session and inter-session scenarios. Experiments demonstrate that both Recall and MRR metrics exhibit nearly 2% to 5% improvement.
Fengyin Li, Hongzhe Liu 0003, Guangshun Li, Huiyu Zhou 0001, Shanshan Cao, Tao Li 0043
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Mumpy: Multilateral Temporal-view Pyramid Transformer for Video Inpainting Detection
Yuezun Li, Bo Peng 0002, Jiaran Zhou, Huiyu Zhou 0001, Junyu Dong
BMVC5
2024 rmD4-VTON: Dynamic Semantics Disentangling for Differential Diffusion Based Virtual Try-On
Zhaotong Yang, Zicheng Jiang, Xinzhe Li 0003, Huiyu Zhou 0001, Junyu Dong, Huaidong Zhang, Yong Du 0003
ECCV (46)4
2024 DecoratingFusion: A LiDAR-Camera Fusion Network with the Combination of Point-Level and Feature-Level Fusion
Zixuan Yin, Ningzhong Liu, Huiyu Zhou 0001, Jiaquan Shen
ICANN (2)4
2024 Online Mouse Behavior Detection by Historical Dependency and Typical Instances
abstract
Mouse behavior analysis plays a pivotal role in the research of numerous neurodegenerative diseases. In this paper, we develop a novel online mouse behavior detection approach, which can recognize mice behaviors in real-time videos and pinpoint the initiation and cessation points of target behaviors. In this architecture, the Long Short-Term Representation Aggregator (LSTRA) employs a designed temporal attention mechanism and integrates temporal dilated convolutions for multi-scale historical dependencies, overcoming the challenge of sparse behavior information due to subtle and brief mouse behaviors. Typical Instance Extractor (TIE) extracts representative frames for each category to calculate category-specific representations, addressing mouse body deformation challenges. The sinkhorn divergence-based constraint in our loss function ensures output congruence of these two modules. Extensive experiments on our PDMB-BD dataset and the public CRIM13 dataset demonstrate our approach achieves superior performance over state-of-the-art approaches.
Feixiang Zhou, Huiyu Zhou 0001
ICASSP3
2024 Adaptive Spatial-Temporal Modelling For Human Motion Prediction
abstract
Human motion prediction is the process of predicting future motion sequences based on past motion sequences. The graph convolution methods currently used for modelling human motion are effective in capturing the interrelationships between joints. However, these works lack skeleton constraints on the learning of graph filters, and DCT-based temporal modeling methods produce overly smooth motion representations and ignore the learning of human motion details. In this paper, we propose a network that uses adaptive spatial graph convolution and temporal self-attention to improve human motion prediction. The adaptive graph convolution effectively enhances cross-scale spatial interaction of joint movements based on different motion patterns. Meanwhile, temporal self-attention, combined with historical motion attention, improves the learning of motion temporal information. Our proposed network achieved state-of-the-art performance on two benchmark datasets, as demonstrated by extensive experiments.
Huiyu Zhou 0001
ICIP2
2024 MMFusion: Multi-modality Diffusion Model for Lymph Node Metastasis Diagnosis in Esophageal Cancer
Chengkai Wang, Huiyu Zhou 0001, Yatao Zhang, Yaqi Wang 0002, Shuai Wang 0003
MICCAI (5)3
2024 Fractional Correspondence Framework in Detection Transformer
abstract
The Detection Transformer (DETR), by incorporating the Hungarian algorithm, has significantly simplified the matching process in object detection tasks. This algorithm facilitates optimal one-to-one matching of predicted bounding boxes to ground-truth annotations during training. While effective, this strict matching process does not inherently account for the varying densities and distributions of objects, leading to suboptimal correspondences such as failing to handle multiple detections of the same object or missing small objects. To address this, we propose the Regularized Transport Plan (RTP). RTP introduces a flexible matching strategy that captures the cost of aligning predictions with ground truths to find the most accurate correspondences between these sets. By utilizing the differentiable Sinkhorn algorithm, RTP allows for soft, fractional matching rather than strict one-to-one assignments. This approach enhances the model's capability to manage varying object densities and distributions effectively. Our extensive evaluations on the MS-COCO and VOC benchmarks demonstrate the effectiveness of our approach. RTP-DETR, surpassing the performance of the Deform-DETR and the recently introduced DINO-DETR, achieving absolute gains in mAP of +3.8% and +1.7%, respectively.
Masoumeh Zareapoor, Pourya Shamsolmoali, Huiyu Zhou 0001, Yue Lu 0001, Salvador García 0001
ACM Multimedia3
2024 TF-FusNet: A Novel Framework for Parkinson's Disease Detection via Time-Frequency Domain Fusion
Weijie Yu 0002, Aite Zhao, Huiyu Zhou 0001
NLPCC (2)5
2024 Low-light image enhancement based on cell vibration energy model and lightness difference
Xiaozhou Lei, Zixiang Fei, Wenju Zhou, Huiyu Zhou 0001, Minrui Fei
Comput. Vis. Image Underst.4
2024 Weighted multi-error information entropy based you only look once network for underwater object detection
Haiping Ma, Shengyi Sun, Minrui Fei, Huiyu Zhou 0001
Eng. Appl. Artif. Intell.6
2024 Enhancing privacy management protection through secure and efficient processing of image information based on the fine-grained thumbnail-preserving encryption
abstract
The increase of image information brings the need for secure storage and management, and people are used to uploading images to cloud servers for storage, but the issue of privacy management and protection has become a great challenge because images may contain some sensitive information. To solve this problem, this paper proposes a novel secure and efficient fine-grained TPE scheme (FG-TPE), specifically, the image pixels are firstly divided into blocks, and multiple rounds of neighboring pixel substitution and permutation fine-grained encryption operations are performed in each block to achieve obfuscated protection of sensitive feature information of the image. Then, the state transfer process of image pixel encryption is reduction to the adversarial detection in a stochastic environment, and the optimal encryption rounds bounds are found by Kalman filtering method. Finally, experiments conducted on two face datasets show that, in qualitative and quantitative comparisons, the average encryption time is decreased remarkably, improved encryption efficiency, and the ciphertext expansion rate is reduced by 19.6% on average, possessing a better image spatiality when compared to the state-of-the-art approaches. Excellent resistance to AI restoration performance has been achieved with only 16 × 16 divided block encryption, and face detection recognition has been fully defended against 32 × 32 divided block encryption, achieving a balance between privacy security and usability management of image information.
Yuling Chen 0002, Chaoyue Tan, Huiyu Zhou 0001
Inf. Process. Manag.5
2024 Long-short diffeomorphism memory network for weakly-supervised ultrasound landmark tracking
abstract
Ultrasound is a promising medical imaging modality benefiting from low-cost and real-time acquisition. Accurate tracking of an anatomical landmark has been of high interest for various clinical workflows such as minimally invasive surgery and ultrasound-guided radiation therapy. However, tracking an anatomical landmark accurately in ultrasound video is very challenging, due to landmark deformation, visual ambiguity and partial observation. In this paper, we propose a long-short diffeomorphism memory network (LSDM), which is a multi-task framework with an auxiliary learnable deformation prior to supporting accurate landmark tracking. Specifically, we design a novel diffeomorphic representation, which contains both long and short temporal information stored in separate memory banks for delineating motion margins and reducing cumulative errors. We further propose an expectation maximization memory alignment (EMMA) algorithm to iteratively optimize both the long and short deformation memory, updating the memory queue for mitigating local anatomical ambiguity. The proposed multi-task system can be trained in a weakly-supervised manner, which only requires few landmark annotations for tracking and zero annotation for deformation learning. We conduct extensive experiments on both public and private ultrasound landmark tracking datasets. Experimental results show that LSDM can achieve better or competitive landmark tracking performance with a strong generalization capability across different scanner types and different ultrasound modalities, compared with other state-of-the-art methods.
Xuejun Ni, Sotirios A. Tsaftaris, Huiyu Zhou 0001
Medical Image Anal.6
2024 CLANet: A comprehensive framework for cross-batch cell line identification using brightfield images
abstract
Cell line authentication plays a crucial role in the biomedical field, ensuring researchers work with accurately identified cells. Supervised deep learning has made remarkable strides in cell line identification by studying cell morphological features through cell imaging. However, biological batch (bio-batch) effects, a significant issue stemming from the different times at which data is generated, lead to substantial shifts in the underlying data distribution, thus complicating reliable differentiation between cell lines from distinct batch cultures. To address this challenge, we introduce CLANet, a pioneering framework for cross-batch cell line identification using brightfield images, specifically designed to tackle three distinct bio-batch effects. We propose a cell cluster-level selection method to efficiently capture cell density variations, and a self-supervised learning strategy to manage image quality variations, thus producing reliable patch representations. Additionally, we adopt multiple instance learning(MIL) for effective aggregation of instance-level features for cell line identification. Our innovative time-series segment sampling module further enhances MIL's feature-learning capabilities, mitigating biases from varying incubation times across batches. We validate CLANet using data from 32 cell lines across 93 experimental bio-batches from the AstraZeneca Global Cell Bank. Our results show that CLANet outperforms related approaches (e.g. domain adaptation, MIL), demonstrating its effectiveness in addressing bio-batch effects in cell line identification.
Adam Corrigan, Navin Rathna Kumar, Kerry Hallbrook, Jonathan Orme, Yinhai Wang, Huiyu Zhou 0001
Medical Image Anal.7
2024 HIDEmarks: hiding multiple marks for robust medical data sharing using IWT-LSB
Om Prakash Singh, Kedar Nath Singh, Naman Baranwal, Amrit Kumar Agrawal, Amit Kumar Singh 0001, Huiyu Zhou 0001
Multim. Tools Appl.6
2024 InA: Inhibition Adaption on pre-trained language models
Cheng Kang, Jindrich Prokop, Huiyu Zhou 0001, Yong Hu 0003, Daniel Novák
Neural Networks4
2024 Deep Learning Methods for Calibrated Photometric Stereo and Beyond
abstract
Photometric stereo recovers the surface normals of an object from multiple images with varying shading cues, i.e., modeling the relationship between surface orientation and intensity at each pixel. Photometric stereo prevails in superior per-pixel resolution and fine reconstruction details. However, it is a complicated problem because of the non-linear relationship caused by non-Lambertian surface reflectance. Recently, various deep learning methods have shown a powerful ability in the context of photometric stereo against non-Lambertian surfaces. This paper provides a comprehensive review of existing deep learning-based calibrated photometric stereo methods utilizing orthographic cameras and directional light sources. We first analyze these methods from different perspectives, including input processing, supervision, and network architecture. We summarize the performance of deep learning photometric stereo models on the most widely-used benchmark data set. This demonstrates the advanced performance of deep learning-based photometric stereo methods. Finally, we give suggestions and propose future research trends based on the limitations of existing models.
Yakun Ju, Kin-Man Lam 0001, Wuyuan Xie, Huiyu Zhou 0001, Junyu Dong, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Underwater object detection in noisy imbalanced datasets
Long Chen 0019, Tengyue Li, Andy Zhou, Shengke Wang, Junyu Dong, Huiyu Zhou 0001
Pattern Recognit.6
2024 WRD-Net: Water Reflection Detection using a parallel attention transformer
Huijie Dong, Hao Qi 0009, Huiyu Zhou 0001, Junyu Dong, Xinghui Dong
Pattern Recognit.3
2024 Deep orientated distance-transform network for geometric-aware centerline detection
Zheheng Jiang, Hossein Rahmani 0001, Plamen Angelov 0001, Ritesh Vyas, Huiyu Zhou 0001, Sue Black 0002, Bryan M. Williams 0001
Pattern Recognit.5
2024 Distance-based Weighted Transformer Network for image completion
Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Xuelong Li 0001, Yue Lu 0001
Pattern Recognit.3
2024 MDD-Enabled Two-Tier Terahertz Fronthaul in Indoor Industrial Cell-Free Massive MIMO
abstract
To liberate indoor industrial cell-free massive multiple-input multiple-output (CF-mMIMO) networks from wired fronthaul, this paper proposes a multicarrier-division duplex (MDD)-enabled two-tier terahertz (THz) fronthaul scheme. In our scheme, two layers of fronthaul links rely on the mutually orthogonal subcarrier sets in the same THz band, while access links are implemented over sub-6G band. However, the proposed scheme leads to a complicated mixed-integer nonconvex optimization problem incorporating access point (AP) clustering, device selection, the assignment of subcarrier sets and the resource allocation at both the central processing unit (CPU) and APs. Hence, in order to address the formulated problem, we first resort to the low-complexity but efficient heuristic methods thereby relaxing the involved binary variables. Then, the overall end-to-end optimization is implemented by iteratively optimizing the assignment of subcarrier sets and the number of AP clusters. Furthermore, an advanced MDD frame structure consisting of three parallel data streams is tailored for the proposed scheme. Simulation results demonstrate the effectiveness of the proposed dynamic AP clustering approach in dealing with the networks of varying sizes. Moreover, benefiting from the well-designed frame structure, MDD is capable of outperforming TDD in the two-tier fronthaul networks. Additionally, the effect of the THz bandwidth on system performance is analyzed, and it is shown that empowered by sufficient bandwidth, our proposed two-tier fully-wireless fronthaul scheme can achieve a comparable performance to the fiber-optic based systems. Finally, the superiority of the proposed MDD-enabled fronthaul scheme is verified in a practical scenario with realistic ray-tracing simulations.
Bohan Li 0005, Diego Dupleich, Guoqing Xia, Huiyu Zhou 0001, Yue Zhang 0011, Pei Xiao 0001, Lie-Liang Yang
IEEE Trans. Commun.4
2024 DRNet: Disentanglement and Recombination Network for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment novel classes with only a few annotated samples. Existing methods to FSS generally combine the annotated mask and the corresponding support image to generate the class-specific representation, and perform the segmentation for the query image by matching the features of the query image to these representations. However, the segmentation performance could be fragile for the lack of an effective method to handle the inappropriate use of query features and the neglection of correlation between features in support and query images. In this work, we propose a novel Disentanglement and Recombination Network (DRNet) to alleviate this problem. Concretely, we first apply the self-attention on both support foreground features and query foreground features. Then, the foreground features of the support and query branches are recombined using the cross-attention after self-attention computation, which can encourage the foreground feature alignment between branches. Finally, the prototypes are generated from the recombined foreground features and support background features, and are utilized to guide the segmentation for given images. Considering the sensitivity of prototypes related to the subtle differences among objects from different classes and the same class, we further introduce a joint learning strategy to derive accurate segmentation of both seen and unseen objects in the support image and the query image respectively. Extensive experiments on the PASCAL-5iand COCO-20idatasets demonstrate the superiority of our DRNet comparing with the recent popular methods. The code is released on https://github.com/GS-Chang-Hn/DRNet-fss.
Zhaobin Chang, Xiong Gao, Huiyu Zhou 0001, Yonggang Lu
IEEE Trans. Circuits Syst. Video Technol.4
2024 Maximizing Contrast in XOR-Based Visual Cryptography Schemes
abstract
Visual cryptography (VC) schemes provide a distinguished image encryption technique to protect image security since it can visually decrypt the secret image by superimposing the encrypted shadows. VC schemes for both threshold access structures and general access structures are generally constructed based on the OR operation to minimize the pixel expansion. However, VC schemes with optimal pixel expansion typically have low contrast. Stacking operation OR frequently produces recovered images with poor visual quality and are never able to deliver flawless recovery for secret images. Therefore, we studied XOR-based VC (XVC) schemes that employ linear programming to maximize their contrast. Three schemes for general access structures and three schemes for threshold access structures are designed to maximize their contrast. The proposed schemes’ construction is reduced to a linear programming to maximize the contrast by determining the ideal combinations of basis matrices in terms of primary column matrices and unit matrices, respectively. The comparison study and experimental results demonstrate that the contrast of the previous VC and XVC schemes can be further improved.
Xingxing Jia, Xiangyang Luo 0001, Daoshun Wang, Huiyu Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Small Sample Image Segmentation by Coupling Convolutions and Transformers
abstract
Compared with natural image segmentation, small sample image segmentation tasks, such as medical image segmentation and defect detection, have been less studied. Recent studies made efforts on bringing together Convolutional Neural Networks (CNNs) and Transformers in a serial or interleaved architecture in order to incorporate long-range dependencies into the features extracted using CNNs. In this study, we argue that these architectures limit the capability of the combination of CNNs and Transformers. To this end, we propose a dual-stream small sample image segmentation network, namely, the Interactive Coupling of Convolutions and Transformers Based UNet (ICCT-UNet)1, motivated by the success achieved using the UNet in the scenario of small sample image segmentation. Within this network, a CNN stream is paralleled with a Transformer stream while maintaining feature exchange inside each block through the proposed Window-Based Multi-head Cross-Attention (W-MHCA) mechanism. To derive an overall segmentation, the features learned by both the streams are further fused using a Residual Fusion Module (RFM). Experimental results show that the ICCT-UNet outperforms, or at least performs comparably to, its counterparts on eight sets of medical and defective images. These promising results should be attributed to the effective combination of the local and global features fulfilled by the proposed interactive coupling method.
Hao Qi 0009, Huiyu Zhou 0001, Junyu Dong, Xinghui Dong
IEEE Trans. Circuits Syst. Video Technol.2
2024 Stealthy Measurement-Aided Pole-Dynamics Attacks With Nominal Models
abstract
When traditional pole-dynamics attacks (TPDAs) are implemented with nominal models, model mismatch between exact and nominal models often affects their stealthiness, or even makes the stealthiness lost. To solve this problem, this article presents a novel stealthy measurement-aided pole-dynamics attacks (MAPDAs) method with model mismatch. First, the limitations of TPDAs using exact models are revealed. Second, to handle the limitations, the proposed MAPDAs method is designed by using an adaptive control strategy, which can keep the stealthiness. Moreover, it is easier to implement as only the measurements are needed in comparison with the existing methods requiring both measurements and control inputs. Third, the performance of the proposed MAPDAs method is explored using convergence of multivariate measurements, and MAPDAs with model mismatch have the same stealthiness and similar destructiveness as TPDAs. Finally, experimental results from a networked inverted pendulum system confirm the feasibility and effectiveness of the proposed method.
Dajun Du, Changda Zhang, Chen Peng 0001, Minrui Fei, Huiyu Zhou 0001
IEEE Trans. Cybern.5
2024 SRCNet: Seminal Image Representation Collaborative Network for Oil Spill Segmentation in SAR Imagery
abstract
Effective oil spill segmentation in synthetic aperture radar (SAR) images is critical for marine oil pollution cleanup, and proper image representation contributes to effective learning for accurate oil spill segmentation. In this article, we propose an effective oil spill segmentation network named SRCNet, which is constructed by leveraging seminal SAR image representation to empower the learning capability of the proposed segmentation network for accurate oil spill segmentation. Specifically, the image representation utilized in our proposed SRCNet originates from SAR imagery, modeling with the internal characteristics of oil spill SAR image data, which therefore promotes effective learning for accurate oil spill segmentation in the training process. Besides, to conduct enhanced oil spill segmentation, we construct the proposed SRCNet with a pair of deep neural nets that work in a competition manner, where one neural net strives to produce accurate oil spill segmentation maps by drawing samples from the collaborated seminal image representation, while the other tries its best to distinguish between the produced and the true segmentations. It is the competition and the image representation collaborated that drives the proposed SRCNet to operate accurate oil spill segmentation efficiently with small amount of training data. This establishes an economical and efficient way for oil spill segmentation. Additionally, to further improve the segmentation performance of the proposed SRCNet, a regularization term that penalizes the segmentation loss is devised, which encourages the produced segmentation to approach the ground-truth segmentation, promoting the segmentation capability of the proposed SRCNet for accurate oil spill segmentation. Experimental evaluations from different metrics validate the effectiveness of the proposed SRCNet for oil spill segmentation.
Heiko Balzter, Peng Ren 0001, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 SAR Target Incremental Recognition Based on Features With Strong Separability
abstract
With the rapid development of deep learning technology, many synthetic aperture radar (SAR) target recognition algorithms based on convolutional neural networks have achieved exceptional performance on various datasets. However, conventional neural networks are repeatedly iterated on a fixed dataset until convergence, and once they learn new tasks, a large amount of previously learned knowledge is forgotten, leading to a significant decline in performance on old tasks. This article presents an incremental learning method based on strong separability features (SSF-IL) to address the model’s forgetting of previously learned knowledge. The SSF-IL employs both intraclass and interclass scatter to compute the feature separability loss, in order to enhance the linear separability of features during incremental learning. In the process of learning new classes, an intraclass clustering loss is proposed to replace the conventional knowledge distillation. This loss function constrains the old class features to cluster around the saved class centers, maintaining the separability among the old class features. Finally, a classifier bias correction method based on boundary features is designed to reinforce the classifier’s decision boundary and reduce classification errors. SAR target incremental recognition experiments are conducted on the MSTAR dataset, and the results are compared with several existing incremental learning algorithms to demonstrate the effectiveness of the proposed algorithm.
Fei Gao 0005, Lingzhe Kong, Rongling Lang, Jinping Sun, Jun Wang 0041, Amir Hussain 0001, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.7
2024 BBox-Free SAR Ship Instance Segmentation Method Based on Gaussian Heatmap
abstract
Recently, deep learning methods have been widely adopted for ship detection in synthetic aperture radar (SAR) images. However, many of the existing methods miss adjacent ship instances when detecting densely arranged ship targets in inshore scenes. Besides, they suffer from the lack of precision in the instance indication information and the confusion of multiple instances by a single mask head. In this paper, we propose a novel center point prediction algorithm, which detects the center points by finding a long distance variation relationship between two points. The whole prediction process is anchor-free and does not require additional bounding box (BBox) predictions for non-maximum suppression (NMS). Therefore, our algorithm is BBox-free and NMS-free, solving the problem of low recall rates when conducting NMS for densely arranged targets. Furthermore, to tackle the deficiency of position indication information in localization tasks, we introduce a feature fusion module with feature decoupling (FD). This module uses classification branch to provide guidance information for localization branch, while suppressing the influence of the gradient flow mixing, effectively improving the algorithm’s segmentation performance of ship contours. Finally, through principal component analysis (PCA) of the Gaussian distribution covariance matrix, we propose a loss function based on the distance between centroids and the difference of angle, called centroid and angle constraint (CAC). CAC guides the network in learning the criterion that a single dynamic mask head is only valid for a single instance. Experiments conducted on PSeg-SSDD and HRSID demonstrate the effectiveness and robustness of our method.
Fei Gao 0005, Fengjun Zhong, Jinping Sun, Amir Hussain 0001, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Deep Color-Corrected Multiscale Retinex Network for Underwater Image Enhancement
abstract
The acquisition of high-quality underwater images is of great importance to ocean exploration activities. However, images captured in the underwater environment often suffer from degradation due to complex imaging conditions, leading to various issues, such as color cast, low contrast, and low visibility. Although many traditional methods have been used to address these issues, they usually lack robustness in diverse underwater scenes. On the other hand, deep learning techniques struggle to generalize to unseen images, due to the challenge of learning the complicated degradation process. Inspired by the success achieved by the Retinex-based methods, we decompose the underwater image enhancement (UIE) task into two consecutive procedures, including color correction and visibility enhancement, and introduce a novel deep color-corrected multiscale retinex network (CCMSR-Net) (code and models are available athttps://indtlab.github.io/projects/CCMSRNet). With regard to the two procedures, this network comprises a color correction subnetwork (CC-Net) and an MSR subnetwork (MSR-Net), which are built on top of the hybrid convolution–axial attention block (HCAAB) that we design. Thanks to this block, the CCMSR-Net is able to efficiently capture local characteristics and the global context. Experimental results show that the CCMSR-Net outperforms, or at least performs comparably to, 11 baselines across five test sets. We believe that these promising results are due to the effective combination of color correction methods and the MSR model, achieved by jointly exploiting convolutional neural networks (CNNs) and transformers.
Hao Qi 0009, Huiyu Zhou 0001, Junyu Dong, Xinghui Dong
IEEE Trans. Geosci. Remote. Sens.2
2024 Hierarchical Knowledge Graph for Multilabel Classification of Remote Sensing Images
abstract
Multilabel classification in remote sensing (RS) images aims to correctly predict multiple object labels in an RS image with the primary challenge of mining correlations among multiple labels. In this context, we argue that a scene can be treated as a high-level depiction of the interactions among multiple interconnected objects within the image. However, hierarchical relationships between the scene and local objects are often neglected in other state-of-the-art approaches. In this article, we consider multilabel classification as a global-to-local prediction process, whereas the scene of an image is first identified, followed by recognition of local objects in the image. To achieve this, we propose a novel hierarchical knowledge graph (HKG)-based framework for multilabel classification in RS images (ML-HKG). Specifically, we first construct a hierarchical KG to depict label correlations between scenes and objects and represent the hierarchical knowledge as interrelated scene- and object-level label embeddings. Subsequently, we generate a scene-aware enhanced feature map by recognizing scene categories in an image under the guidance of scene-level knowledge embeddings. Afterward, object-level embeddings are used to derive category-specific visual representations for final multilabel prediction. Extensive experiments on the UCM and AID datasets demonstrate the effectiveness of our framework.
Xiangrong Zhang, Xina Cheng, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2024 DomainForensics: Exposing Face Forgery Across Domains via Bi-Directional Adaptation
abstract
Recent DeepFake detection methods have shown excellent performance on public datasets but are significantly degraded on new forgeries. Solving this problem is important, as new forgeries emerge daily with the continuously evolving generative techniques. Many efforts have been made for this issue by seeking the commonly existing traces empirically on data level. In this paper, we rethink this problem and propose a new solution from the unsupervised domain adaptation perspective. Our solution, called DomainForensics, aims to transfer the forgery knowledge from known forgeries (fully labeled source domain) to new forgeries (label-free target domain). Unlike recent efforts, our solution does not focus on data view but on learning strategies of DeepFake detectors to capture the knowledge of new forgeries through the alignment of domain discrepancies. In particular, unlike the general domain adaptation methods which consider the knowledge transfer in the semantic class category, thus having limited application, our approach captures the subtle forgery traces. We describe a new bi-directional adaptation strategy dedicated to capturing the forgery knowledge across domains. Specifically, our strategy considers both forward and backward adaptation, to transfer the forgery knowledge from the source domain to the target domain in forward adaptation and then reverse the adaptation from the target domain to the source domain in backward adaptation. In forward adaptation, we perform supervised training for the DeepFake detector in the source domain and jointly employ adversarial feature adaptation to transfer the ability to detect manipulated faces from known forgeries to new forgeries. In backward adaptation, we further improve the knowledge transfer by coupling adversarial adaptation with self-distillation on new forgeries. This enables the detector to expose new forgery features from unlabeled data and avoid forgetting the known knowledge of known forgery. Extensive experiments demonstrate that our method is surprisingly effective in exposing new forgeries, and can be plug-and-play on other DeepFake detection architectures.
Qingxuan Lv, Yuezun Li, Junyu Dong, Sheng Chen 0001, Hui Yu 0001, Huiyu Zhou 0001, Shu Zhang 0002
IEEE Trans. Inf. Forensics Secur.6
2024 Sketch-Supervised Histopathology Tumour Segmentation: Dual CNN-Transformer With Global Normalised CAM
abstract
Deep learning methods are frequently used in segmenting histopathology images with high-quality annotations nowadays. Compared with well-annotated data, coarse, scribbling-like labelling is more cost-effective and easier to obtain in clinical practice. The coarse annotations provide limited supervision, so employing them directly for segmentation network training remains challenging. We present a sketch-supervised method, called DCTGN-CAM, based on a dual CNN-Transformer network and a modified global normalised class activation map. By modelling global and local tumour features simultaneously, the dual CNN-Transformer network produces accurate patch-based tumour classification probabilities by training only on lightly annotated data. With the global normalised class activation map, more descriptive gradient-based representations of the histopathology images can be obtained, and inference of tumour segmentation can be performed with high accuracy. Additionally, we collect a private skin cancer dataset named BSS, which contains fine and coarse annotations for three types of cancer. To facilitate reproducible performance comparison, experts are also invited to label coarse annotations on the public liver cancer dataset PAIP2019. On the BSS dataset, our DCTGN-CAM segmentation outperforms the state-of-the-art methods and achieves 76.68 % IOU and 86.69 % Dice scores on the sketch-based tumour segmentation task. On the PAIP2019 dataset, our method achieves a Dice gain of 8.37 % compared with U-Net as the baseline network.
Yilong Li 0002, Linyan Wang, Xingru Huang, Yaqi Wang 0002, Ruiquan Ge, Huiyu Zhou 0001, Juan Ye, Qianni Zhang
IEEE J. Biomed. Health Informatics7
2024 Single Traffic Image Deraining via Similarity-Diversity Model
abstract
Single traffic image deraining technology based on deep learning is a vital branch of image preprocessing, which is of great help to intelligent monitoring systems and driving navigation system. It is well understood that established deraining methods are derived based on one specific imaging model, neglecting the underlying correlations between different weather models and thereby limiting the applicability of these standard methods in real scenarios. To ameliorate this issue, in this work, we first explore the inherent relationship between a rain model and the haze one established up to date. We discover that these two models experience similar degradations in the low-frequency components (i.e., similarity) but diverse degradations in the high-frequency areas (i.e., diversity). Based on these observations, we develop a Similarity-Diversity model to describe these characteristics. Afterwards, we introduce a novel deep neural network to restore the rain-free background embedding the similarity-diversity model, namely deep similarity-diversity network (DSDNet). Extensive experiments have been conducted to evaluate our proposed method that outperforms the other state of the art deraining techniques. On the other hand, we deploy the proposed algorithm with Google Vision API for object recognition, which also obtains satisfactory results both qualitatively and quantitatively.
Youxing Li, Rushi Lan, Huiwen Huang, Huiyu Zhou 0001, Zhenbing Liu
IEEE Trans. Intell. Transp. Syst.4
2024 SMC-NCA: Semantic-Guided Multi-Level Contrast for Semi-Supervised Temporal Action Segmentation
abstract
Semi-supervised temporal action segmentation (SS-TAS) aims to perform frame-wise classification in long untrimmed videos, where only a fraction of videos in the training set have labels. Recent studies have shown the potential of contrastive learning in unsupervised representation learning using unlabelled data. However, learning the representation of each frame by unsupervised contrastive learning for action segmentation remains an open and challenging problem. In this paper, we propose a novel Semantic-guided Multi-level Contrast scheme with a Neighbourhood-Consistency-Aware unit (SMC-NCA) to extract strong frame-wise representations for SS-TAS. Specifically, for representation learning, SMC is first used to explore intra- and inter-information variations in a unified and contrastive way, based on action-specific semantic information and temporal information highlighting relations between actions. Then, the NCA module, which is responsible for enforcing spatial consistency between neighbourhoods centered at different frames to alleviate over-segmentation issues, works alongside SMC for semi-supervised learning (SSL). Our SMC outperforms the other state-of-the-art methods on three benchmarks, offering improvements of up to 17.8$\%$and 12.6$\%$in terms of Edit distance and accuracy, respectively. Additionally, the NCA unit results in significantly better segmentation performance in the presence of only 5$\%$labelled videos. We also demonstrate the generalizability and effectiveness of the proposed method on our Parkinson's Disease Mouse Behaviour (PDMB) dataset.
Feixiang Zhou, Zheheng Jiang, Huiyu Zhou 0001, Xuelong Li 0001
IEEE Trans. Multim.3
2024 Multiview Subspace Clustering via Low-Rank Symmetric Affinity Graph
abstract
Multiview subspace clustering (MVSC) has been used to explore the internal structure of multiview datasets by revealing unique information from different views. Most existing methods ignore the consistent information and angular information of different views. In this article, we propose a novel MVSC via low-rank symmetric affinity graph (LSGMC) to tackle these problems. Specifically, considering the consistent information, we pursue a consistent low-rank structure across views by decomposing the coefficient matrix into three factors. Then, the symmetry constraint is utilized to guarantee weight consistency for each pair of data samples. In addition, considering the angular information, we utilize the fusion mechanism to capture the inherent structure of data. Furthermore, to alleviate the effect brought by the noise and the high redundant data, the Schatten p-norm is employed to obtain a low-rank coefficient matrix. Finally, an adaptive information reduction strategy is designed to generate a high-quality similarity matrix for spectral clustering. Experimental results on 11 datasets demonstrate the superiority of LSGMC in clustering performance compared with ten state-of-the-art multiview clustering methods.
Wei Lan 0001, Tianchuan Yang, Qingfeng Chen, Shichao Zhang 0001, Huiyu Zhou 0001, Yi Pan 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 GR-PSN: Learning to Estimate Surface Normal and Reconstruct Photometric Stereo Images
abstract
In this paper, we propose a novel method, namely GR-PSN, which learns surface normals from photometric stereo images and generates the photometric images under distant illumination from different lighting directions and surface materials. The framework is composed of two subnetworks, named GeometryNet and ReconstructNet, which are cascaded to perform shape reconstruction and image rendering in an end-to-end manner. ReconstructNet introduces additional supervision for surface-normal recovery, forming a closed-loop structure with GeometryNet. We also encode lighting and surface reflectance in ReconstructNet, to achieve arbitrary rendering. In training, we set up a parallel framework to simultaneously learn two arbitrary materials for an object, providing an additional transform loss. Therefore, our method is trained based on the supervision by three different loss functions, namely the surface-normal loss, reconstruction loss, and transform loss. We alternately input the predicted surface-normal map and the ground-truth into ReconstructNet, to achieve stable training for ReconstructNet. Experiments show that our method can accurately recover the surface normals of an object with an arbitrary number of inputs, and can re-render images of the object with arbitrary surface materials. Extensive experimental results show that our proposed method outperforms those methods based on a single surface recovery network and shows realistic rendering results on 100 different materials.
Yakun Ju, Boxin Shi, Yang Chen 0036, Huiyu Zhou 0001, Junyu Dong, Kin-Man Lam 0001
IEEE Trans. Vis. Comput. Graph.4
2024 Joint Beamforming and Compressed Sensing for Uplink Grant-Free Access
abstract
Compressed sensing (CS)-based techniques have been widely applied in the grant-free non-orthogonal multiple access (NOMA) to a single-antenna base station (BS). In this paper, we consider the multi-antenna reception at the BS for uplink grant-free access for the massive machine type communication (mMTC) with limited channel resources. To enhance the overloading performance of the BS, we develop a general framework for the synergistic amalgamation of the spatial division multiple access (SDMA) technique with the CS-based grant-free NOMA. We derive a closed-form statistical beamforming and a dynamic beamforming scheme for the inter-cluster interference suppression when applying SDMA. Based on this, we further develop a joint adaptive beamforming and subspace pursuit (J-ABF-SP) algorithm for the multiuser detection and data recovery, with a novel sparsity level decision method without the accurate knowledge of the noise level. To further improve the data recovery performance, we propose an interference cancellation-based J-ABF-SP scheme (J-ABF-SP-IC) by using the initial signal estimates generated from the J-ABF-SP algorithm. Illustrative simulations verify the superior user detection and signal recovery performance of our proposed algorithms in comparison with existing CS-based grant-free NOMA techniques.
Guoqing Xia, Pei Xiao 0001, Bohan Li 0005, Yue Zhang 0011, Huiyu Zhou 0001
IEEE Trans. Wirel. Commun.5
2023 FGFusion: Fine-Grained Lidar-Camera Fusion for 3D Object Detection
Zixuan Yin, Ningzhong Liu, Huiyu Zhou 0001, Jiaquan Shen
PRCV (3)4
2023 Asymmetric similarity-preserving discrete hashing for image retrieval
Xiuxiu Ren, Xiangwei Zheng 0001, Li-Zhen Cui 0001, Gang Wang 0008, Huiyu Zhou 0001
Appl. Intell.5
2023 Attention guided domain alignment for conditional face image generation
Zonglin Li 0004, Shengping Zhang, Quanling Meng, Qinglin Liu, Huiyu Zhou 0001
Comput. Vis. Image Underst.6
2023 Unsupervised image-to-image translation in multi-parametric MRI of bladder cancer
Zhiying Chen, Lingkai Cai, Chunxiao Chen, Xue Fu, Xiao Yang 0023, Baorui Yuan, Huiyu Zhou 0001
Eng. Appl. Artif. Intell.8
2023 Learning Geometric Transformation for Point Cloud Completion
Shengping Zhang, Xianzhu Liu, Haozhe Xie, Liqiang Nie, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001
Int. J. Comput. Vis.5
2023 Driver Drowsiness EEG Detection Based on Tree Federated Learning and Interpretable Network
abstract
Accurate identification of driver's drowsiness state through Electroencephalogram (EEG) signals can effectively reduce traffic accidents, but EEG signals are usually stored in various clients in the form of small samples. This study attempts to construct an efficient and accurate privacy-preserving drowsiness monitoring system, and proposes a fusion model based on tree Federated Learning (FL) and Convolutional Neural Network (CNN), which can not only identify and explain the driver's drowsiness state, but also integrate the information of different clients under the premise of privacy protection. Each client uses CNN with the Global Average Pooling (GAP) layer and shares model parameters. The tree FL transforms communication relationships into a graph structure, and model parameters are transmitted in parallel along connected branches of the graph. Moreover, the Class Activation Mapping (CAM) is used to find distinctive EEG features for representing specific classes. On EEG data of 11 subjects, it is found that this method has higher average accuracy, F1-score and AUC than the traditional classification method, reaching 73.56%, 73.26% and 78.23%, respectively. Compared with the traditional FL algorithm, this method better protects the driver's privacy and improves communication efficiency.
Huiyu Zhou 0001, Weikuan Jia, Yuanjie Zheng
Int. J. Neural Syst.3
2023 GAN-based watermarking for encrypted images in healthcare scenarios
Himanshu Kumar Singh 0002, Naman Baranwal, Kedar Nath Singh, Amit Kumar Singh 0001, Huiyu Zhou 0001
Neurocomputing5
2023 Multifactor Incentive Mechanism for Federated Learning in IoT: A Stackelberg Game Approach
abstract
In the era of the Internet of Things (IoT), remote sensors and endpoint appliances generate vast amounts of data. Decentralized and collaborative learning builds on these IoT data to enable classification and recognition tasks by inviting multiple data owners. Federated learning (FL), as a popular collaborative learning framework, can significantly improve the performance of models without collecting the original data. To invite data owners to participate in FL, various incentive mechanisms are designed to address this issue by researchers. However, existing solutions still face high costs and low utility due to information asymmetry, where the reputation, computation power, and data quantity of the data owners are not known in advance. Therefore, we propose a Stackelberg game-based multifactor incentive mechanism for FL (SGMFIFL). First, we design the Top-$K$cost selection algorithm based on reverse auction, which can reduce the cost of selecting data owners. Next, we devise a multifactor reward function based on reputation, accuracy, and reward rate, the data owners with high reputation and high accuracy will be of more reward. In particular, to ensure that SGMFIFL can provide reliable incentives in IoT, we use blockchain to provide a secure and trusted environment. Finally, we construct a two-stage Stackelberg game model for the task publisher and the data owners and derive an optimal Equilibrium solution for both stages of the whole game. Experiments conducted on two well-known data sets, MNIST and CIFAR10, demonstrate the significant performance of the proposed mechanism.
Yuling Chen 0002, Hui Zhou 0014, Tao Li 0043, Jin Li 0002, Huiyu Zhou 0001
IEEE Internet Things J.5
2023 Enabling scalable and unlinkable payment channel hubs with oblivious puzzle transfer
Huawei Ma, Shuyu Fan, Huiyu Zhou 0001, Siqi Ju, Xiaoying Wang 0007, Qintai Yang
Inf. Sci.5
2023 Deep learning-based biometric image feature extraction for securing medical images through data hiding and joint encryption-compression
Monu Singh, Naman Baranwal, Kedar Nath Singh, Amit Kumar Singh 0001, Huiyu Zhou 0001
J. Inf. Secur. Appl.5
2023 How effective are current population-based metaheuristic algorithms for variance-based multi-level image thresholding?
Seyed Jalaleddin Mousavirad, Gerald Schaefer, Huiyu Zhou 0001, Mahshid Helali Moghadam
Knowl. Based Syst.3
2023 An Incremental SAR Target Recognition Framework via Memory-Augmented Weight Alignment and Enhancement Discrimination
abstract
Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) is one of the most important research directions in SAR image interpretation. While much existing research into SAR ATR has focused on deep learning technology, an equally important yet underexplored problem is its deployment in incremental learning scenarios. This letter proposes a new benchmark approach, termed Memory augmented weights alignment and Enhancement Discrimination Incremental Learning (MEDIL) algorithm to address this issue. Firstly, the attention mechanism is employed as part of the benchmark. Next, we discuss the problem of height deviation of weights at the fully connected layer and design a more suitable alignment of weights by guiding the memory module for contextual data processing. In addition, we leverage the incremental progressive sampling strategy to alleviate the imbalance between old and new classes during the training period. Finally, we propose to enhance the distinction among various classes with an angular penalty loss function to ensure the diversity of incremental instances. The proposed method is evaluated on MSTAR and OpenSARShip under different experimental settings. Experimental results demonstrate that our proposed approach can effectively solve catastrophic forgetting in SAR multiclass recognition problems.
Fei Gao 0005, Jun Wang 0041, Amir Hussain 0001, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.5
2023 Attention-Aware Three-Branch Network for Salient Object Detection in Remote Sensing Images
abstract
Although remarkable advances have been achieved on salient object detection (SOD) for natural scene images (NSIs), SOD for optical remote sensing images (RSIs) still remains a big challenge due to the unique imaging conditions and various scene patterns. To enable effective SOD for RSIs, this letter proposes a novel end-to-end network, called attention-aware three-branch network (AATBNet). First, an attention feature encoding branch is constructed for learning more discriminative features. Then, a hierarchical feature decoding branch, equipped with three streams, i.e., a decoding stream, a dilated reverse attention stream, and a fusion dense up-sampling convolution stream, is proposed to effectively and robustly compute saliency maps and salient edge maps. Third, a two losses computation branch is designed to further boost SOD performance. Comprehensive evaluations on two well-known RSIs benchmarks, as well as comparisons with 20 state-of-the-art technologies validate the superiority of our AATBNet. The code of our method is publicly available at: https://github.com/WangXin81/AATBNet.
Xin Wang 0068, Zhilu Zhang 0003, Shihan Jing, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.4
2023 POST-IVUS: A perceptual organisation-aware selective transformer framework for intravascular ultrasound segmentation
abstract
Intravascular ultrasound (IVUS) is recommended in guiding coronary intervention. The segmentation of coronary lumen and external elastic membrane (EEM) borders in IVUS images is a key step, but the manual process is time-consuming and error-prone, and suffers from inter-observer variability. In this paper, we propose a novel perceptual oganisation-aware selective transformer framework that can achieve accurate and robust segmentation of the vessel walls in IVUS images. In this framework, temporal context-based feature encoders extract efficient motion features of vessels. Then, a perceptual oganisation-aware selective transformer module is proposed to extract accurate boundary information, supervised by a dedicated boundary loss. The obtained EEM and lumen segmentation results will be fused in a temporal constraining and fusion module, to determine the most likely correct boundaries with robustness to morphology. Our proposed methods are extensively evaluated in non-selected IVUS sequences, including normal, bifurcated, and calcified vessels with shadow artifacts. The results show that the proposed methods outperform the state-of-the-art, with a Jaccard measure of 0.92 for lumen and 0.94 for EEM on the IVUS 2011 open challenge dataset. This work has been integrated into a software QCU-CMS2 to automatically segment IVUS images in a user-friendly environment.
Xingru Huang, Retesh Bajaj, Yilong Li 0002, Xin Ye 0006, Ji Lin 0004, Francesca Pugliese, Anantharaman Ramasamy, Yaqi Wang 0002, Ryo Torii, Jouke Dijkstra, Huiyu Zhou 0001, Christos V. Bourantas, Qianni Zhang
Medical Image Anal.12
2023 Stimulus-guided adaptive transformer network for retinal blood vessel segmentation in fundus images
abstract
Automated retinal blood vessel segmentation in fundus images provides important evidence to ophthalmologists in coping with prevalent ocular diseases in an efficient and non-invasive way. However, segmenting blood vessels in fundus images is a challenging task, due to the high variety in scale and appearance of blood vessels and the high similarity in visual features between the lesions and retinal vascular. Inspired by the way that the visual cortex adaptively responds to the type of stimulus, we propose a Stimulus-Guided Adaptive Transformer Network (SGAT-Net) for accurate retinal blood vessel segmentation. It entails a Stimulus-Guided Adaptive Module (SGA-Module) that can extract local-global compound features based on inductive bias and self-attention mechanism. Alongside a light-weight residual encoder (ResEncoder) structure capturing the relevant details of appearance, a Stimulus-Guided Adaptive Pooling Transformer (SGAP-Former) is introduced to reweight the maximum and average pooling to enrich the contextual embedding representation while suppressing the redundant information. Moreover, a Stimulus-Guided Adaptive Feature Fusion (SGAFF) module is designed to adaptively emphasize the local details and global context and fuse them in the latent space to adjust the receptive field (RF) based on the task. The evaluation is implemented on the largest fundus image dataset (FIVES) and three popular retinal image datasets (DRIVE, STARE, CHASEDB1). Experimental results show that the proposed method achieves a competitive performance over the other existing method, with a clear advantage in avoiding errors that commonly happen in areas with highly similar visual features. The sourcecode is publicly available at: https://github.com/Gins-07/SGAT.
Ji Lin 0004, Xingru Huang, Huiyu Zhou 0001, Yaqi Wang 0002, Qianni Zhang
Medical Image Anal.3
2023 An explainable autoencoder with multi-paradigm fMRI fusion for identifying differences in dynamic functional connectivity during brain development
Faming Xu, Chen Qiao, Huiyu Zhou 0001, Vince D. Calhoun, Julia M. Stephen, Tony W. Wilson, Yu-Ping Wang 0002
Neural Networks3
2023 Graph-based discriminative features learning for fine-grained image retrieval
Wenxi Lang, Can Xu 0006, Ningzhong Liu, Huiyu Zhou 0001
Signal Process. Image Commun.5
2023 Cost-Sensitive Boosting Pruning Trees for Depression Detection on Twitter
abstract
Depression is one of the most common mental health disorders, and a large number of depressed people commit suicide each year. Potential depression sufferers usually do not consult psychological doctors because they feel ashamed or are unaware of any depression, which may result in severe delay of diagnosis and treatment. In the meantime, evidence shows that social media data provides valuable clues about physical and mental health conditions. In this paper, we argue that it is feasible to identify depression at an early stage by mining online social behaviours. Our approach, which is innovative to the practice of depression detection, does not rely on the extraction of numerous or complicated features to achieve accurate depression detection. Instead, we propose a novel classifier, namely, Cost-sensitive Boosting Pruning Trees (CBPT), which demonstrates a strong classification ability on two publicly accessible Twitter depression detection datasets. To comprehensively evaluate the classification capability of CBPT, we use additional three datasets from the UCI machine learning repository and CBPT obtains appealing classification results against several state of the arts boosting algorithms. Finally, we comprehensively explore the influence factors for the model prediction, and the results manifest that our proposed framework is promising for identifying Twitter users with depression.
Zheheng Jiang, Feixiang Zhou, Long Chen 0019, Jialin Lyu, Xiangrong Zhang, Qianni Zhang, Abdul Hamid Sadka, Yinhai Wang, Ling Li 0010, Huiyu Zhou 0001
IEEE Trans. Affect. Comput.12
2023 DGNet: Distribution Guided Efficient Learning for Oil Spill Image Segmentation
abstract
Successful implementation of oil spill segmentation in synthetic aperture radar (SAR) images is vital for marine environmental protection. In this article, we develop an effective segmentation framework named DGNet, which performs oil spill segmentation by incorporating the intrinsic distribution of backscatter values in SAR images. Specifically, our proposed segmentation network is constructed with two deep neural modules running in an interactive manner, where one is the inference module to achieve latent feature variable inference from SAR images and the other is the generative module to produce oil spill segmentation maps by drawing the latent feature variables as inputs. Thus, to yield accurate segmentation, we take into account the intrinsic distribution of backscatter values in SAR images and embed it in our segmentation model. The intrinsic distribution originates from SAR imagery, describing the physical characteristics of oil spills. In the training process, the formulated intrinsic distribution guides efficient learning of optimal latent feature variable inference for oil spill segmentation. The efficient learning enables the training of our proposed DGNet with a small amount of image data. This is economically beneficial to oil spill segmentation where the availability of oil spill SAR image data is limited in practice. Additionally, benefiting from optimal latent feature variable inference, our proposed DGNet performs accurate oil spill segmentation. We evaluate the segmentation performance of our proposed DGNet with different metrics, and experimental evaluations demonstrate its effective segmentations.
Heiko Balzter, Feixiang Zhou, Peng Ren 0001, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Cross-Modality Features Fusion for Synthetic Aperture Radar Image Segmentation
abstract
Synthetic Aperture Radar (SAR) image segmentation stands as a formidable research frontier within the domain of SAR image interpretation. The fully convolutional network (FCN) methods have recently brought remarkable improvements in SAR image segmentation. Nevertheless, these methods do not utilize the peculiarities of SAR images, leading to suboptimal segmentation accuracy. To address this issue, we rethink SAR image segmentation in terms of sequential information of transformers and cross-modal features. We first discuss the peculiarities of SAR images and extract the mean and texture features utilized as auxiliary features. The extraction of auxiliary features helps unearth the distinctive information in the SAR images. Afterward, a feature-enhanced FCN with the transformer encoder structure, termed FE-FCN, which can be extracted to context-level and pixel-level features. In FE-FCN, the features of a single-mode encoder are aligned and inserted into the model to explore the potential correspondence between modes. We also employ long skip connections to share each modality’s distinguishing and particular features. Finally, we present the connection-enhanced conditional random field (CE-CRF) to capture the connection information of the image pixels. Since the CE-CRF utilizes the auxiliary features to enhance the reliability of the connection information, the segmentation results of FE-FCN are further optimized. Comparative experiments conducted on the Fangchenggang (FCG), Pucheng (PC), and Gaofen (GF) SAR datasets. Our method demonstrates superior segmentation accuracy compared to other conventional image segmentation methods, as confirmed by the experimental results.
Fei Gao 0005, Dongyu Li, Shuzhi Sam Ge, Tong Heng Lee, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 Efficient Object Detection in Optical Remote Sensing Imagery via Attention-Based Feature Distillation
abstract
Efficient object detection methods have recently received great attention in remote sensing. Although deep convolutional networks often have excellent detection accuracy, their deployment on resource-limited edge devices is difficult. Knowledge distillation (KD) is a strategy for addressing this issue since it makes models lightweight while maintaining accuracy. However, existing KD methods for object detection have encountered two constraints. First, they discard potentially important background information and only distill nearby foreground regions. Second, they only rely on the global context, which limits the student detector’s ability to acquire local information from the teacher detector. To address the aforementioned challenges, we propose Attention-based Feature Distillation (AFD), a new KD approach that distills both local and global information from the teacher detector. To enhance local distillation, we introduce a multi-instance attention mechanism that effectively distinguishes between background and foreground elements. This approach prompts the student detector to focus on the pertinent channels and pixels, as identified by the teacher detector. Local distillation lacks global information, thus attention global distillation is proposed to reconstruct the relationship between various pixels and pass it from teacher to student detector. The performance of AFD is evaluated on two public aerial image benchmarks, and the evaluation results demonstrate that AFD in object detection can attain the performance of other state-of-the-art models while being efficient.
Pourya Shamsolmoali, Jocelyn Chanussot, Huiyu Zhou 0001, Yue Lu 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 High-Quality Angle Prediction for Oriented Object Detection in Remote Sensing Images
abstract
Oriented object detection is a challenging task in remote sensing, where the detected objects can be represented by oriented bounding boxes (OBBs). Angle prediction in oriented object detection has been widely studied, due to its crucial role in object detection. However, the precision of angle prediction is severely limited by misalignments in most of the existing methods, including representation-, evaluation-, and optimization-based misalignments. To alleviate these misalignments, this paper presents a novel angle prediction method, called Angle Quality Estimation (AQE). Specifically, our proposed AQE transforms the angle prediction task into a distribution estimation task to address the representation misalignment problem and implicitly measure the quality of the predicted angles. Based on the estimated angle quality, we then propose a new metric to comprehensively evaluate the quality of OBBs. Then we propose an object aspect ratio based loss function to optimize angle prediction for addressing the optimization misalignment. Our proposed AQE is a plug-and-play method, which can be embedded on any existing oriented object detector. Experimental results on three public benchmarks, including DOTA, HRSC2016, and ICDAR2015 datasets, show that our method achieves better performance than the other state-of-the-art.
Guanchun Wang, Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Puhua Chen, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 Semantics and Contour Based Interactive Learning Network for Building Footprint Extraction
abstract
Building footprint extraction plays an important role in the analysis of remote sensing images and has an extensive range of applications. Obtaining precise boundaries of buildings remains a challenge in existing building extraction methods. Some previous works have made notable efforts to address this concern. However, most of these methods require cumbersome and expensive post-processing steps. Moreover, they ignored the correlation between building semantics and contours, which we believe is crucial for building footprint extraction. To mitigate this issue, our paper presents an intuitive and effective framework that explores semantic and contour cues of buildings and fully excavates their correlation. Specifically, we construct an interactive dual-stream decoder. The Intermediate connections within this decoder interactively transmit features between branches, contributing to learning correlations between semantics and contours. We propose the Semantic Collaboration Module (SCM) to strengthen the connection between the two branches. To further boost performance, we build the Multi-Scale Semantic Context Fusion Module (MSCF) to fuse semantic information from the higher and lower layers of the network, allowing the network to obtain superior feature representations. The experimental results on the WHU, INRIA, and Massachusetts building datasets demonstrate the superior performance of our method.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2023 ViMDH: Visible-Imperceptible Medical Data Hiding for Internet of Medical Things
abstract
Over the recent years, volume of medical images and related digital records, called electronic medical records, generated, shared, and stored by different intelligent devices, sensors, and Internet of medical things networks, to name a few, has drastically increased. Such records are shared by cloud providers for storage and further processing. However, an increasingly serious concern is the illegal copying, modification, and forgery of medical records. This article presents a visible and imperceptible medical data hiding technique, namely ViMDH, which can prevent to intellectual property theft of medical records. The carrier image is visibly marked with logo mark, which is suitable for owner identification and avoid illegal duplication, and then an imperceptible data hiding based on nonsubsampled shearlet transform (NSST), redundant discrete wavelet transform (RDWT), and multiresolution singular value decomposition is introduced. Finally, key-based encryption scheme designed by RDWT-RSVD ensure the security of the watermarking system. Under the experimental evaluation, our ViMDH is not only visible and imperceptible, but also has a satisfactory advantage in robustness and security compared with the traditional watermarking schemes.
Ashima Anand, Amit Kumar Singh 0001, Huiyu Zhou 0001
IEEE Trans. Ind. Informatics3
2023 Attack Detection for Networked Control Systems Using Event-Triggered Dynamic Watermarking
abstract
Dynamic watermarking schemes can enhance the cyberattack detection capability of networked control systems (NCSs). This article presents a linear event-triggered solution to conventional dynamic watermarking (CDW) schemes. First, the limitations of CDW schemes for event-triggered state estimation-based NCSs are investigated. Second, a new event-triggered dynamic watermarking (ETDW) scheme is designed by treating watermarking as symmetric key encryption, based on the limit convergence theorem in probability. Its security property against the generalized replay attacks (GRAs) is also discussed in the form of bounded asymptotic attack power. Third, finite sample ETDW tests are designed with matrix concentration inequalities. Finally, experimental results of a networked inverted pendulum system demonstrate the validity of our proposed scheme.
Dajun Du, Changda Zhang, Xue Li 0028, Minrui Fei, Huiyu Zhou 0001
IEEE Trans. Ind. Informatics5
2023 VTAE: Variational Transformer Autoencoder With Manifolds Learning
abstract
Deep generative models have demonstrated successful applications in learning non-linear data distributions through a number of latent variables and these models use a non-linear function (generator) to map latent samples into the data space. On the other hand, the non-linearity of the generator implies that the latent space shows an unsatisfactory projection of the data space, which results in poor representation learning. This weak projection, however, can be addressed by a Riemannian metric, and we show that geodesics computation and accurate interpolations between data samples on the Riemannian manifold can substantially improve the performance of deep generative models. In this paper, a Variational spatial-Transformer AutoEncoder (VTAE) is proposed to minimize geodesics on a Riemannian manifold and improve representation learning. In particular, we carefully design the variational autoencoder with an encoded spatial-Transformer to explicitly expand the latent variable model to data on a Riemannian manifold, and obtain global context modelling. Moreover, to have smooth and plausible interpolations while traversing between two different objects' latent representations, we propose a geodesic interpolation network different from the existing models that use linear interpolation with inferior performance. Experiments on benchmarks show that our proposed model can improve predictive accuracy and versatility over a range of computer vision tasks, including image interpolations, and reconstructions.
Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Image Process.3
2023 Cluster-Re-Supervision: Bridging the Gap Between Image-Level and Pixel-Wise Labels for Weakly Supervised Medical Image Segmentation
abstract
Weakly supervised learning, releasing deep learning from highly labor-intensive pixel-wise annotations, has gained great attention, especially for medical image segmentation. With only image-level labels, pixel-wise segmentation/localization usually is achieved based on class activation maps (CAMs) containing the most discriminative regions. One common consequence of CAM-based approaches is incomplete foreground segmentation, i.e. under-segmentation/false negatives. Meanwhile, suffering from relatively limited medical imaging data, class-irrelevant tissues can hardly be suppressed during classification, resulting in incorrect background identification, i.e. over-segmentation/false positives. The above two issues are determined by the loose-constraint nature of image-level labels penalizing on the entire image space, and thus how to develop pixel-wise constraints based on image-level labels is the key for performance improvement which is under-explored. In this paper, based on unsupervised clustering, we propose a new paradigm called cluster-re-supervision to evaluate the contribution of each pixel in CAMs to final classification and thus generate pixel-wise supervision (i.e., clustering maps) for CAMs refinement on both over- and under-segmentation reduction. Furthermore, based on self-supervised learning, an inter-modality image reconstruction module, together with random masking, is designed to complement local information in feature learning which helps stabilize clustering. Experimental results on two popular public datasets demonstrate the superior performance of the proposed weakly-supervised framework for medical image segmentation. More importantly, cluster-re-supervision is independent of specific tasks and highly extendable to other applications.
Zhuo Kuang, Zengqiang Yan, Huiyu Zhou 0001, Li Yu 0003
IEEE J. Biomed. Health Informatics3
2023 Low-Light Image Enhancement Using the Cell Vibration Model
abstract
Low light very likely leads to the degradation of an image’s quality and even causes visual task failures. Existing image enhancement technologies are prone to overenhancement, color distortion or time consumption, and their adaptability is fairly limited. Therefore, we propose a new single low-light image lightness enhancement method. First, an energy model is presented based on the analysis of membrane vibrations induced by photon stimulations. Then, based on the unique mathematical properties of the energy model and combined with the gamma correction model, a new global lightness enhancement model is proposed. Furthermore, a special relationship between image lightness and gamma intensity is found. Finally, a local fusion strategy, including segmentation, filtering and fusion, is proposed to optimize the local details of the global lightness enhancement images. Experimental results show that the proposed algorithm is superior to nine state-of-the-art methods in avoiding color distortion, restoring the textures of dark areas, reproducing natural colors and reducing time cost.
Xiaozhou Lei, Zixiang Fei, Wenju Zhou, Huiyu Zhou 0001, Minrui Fei
IEEE Trans. Multim.4
2023 Detecting and Tracking of Multiple Mice Using Part Proposal Networks
abstract
The study of mouse social behaviors has been increasingly undertaken in neuroscience research. However, automated quantification of mouse behaviors from the videos of interacting mice is still a challenging problem, where object tracking plays a key role in locating mice in their living spaces. Artificial markers are often applied for multiple mice tracking, which are intrusive and consequently interfere with the movements of mice in a dynamic environment. In this article, we propose a novel method to continuously track several mice and individual parts without requiring any specific tagging. First, we propose an efficient and robust deep-learning-based mouse part detection scheme to generate part candidates. Subsequently, we propose a novel Bayesian-inference integer linear programming (BILP) model that jointly assigns the part candidates to individual targets with necessary geometric constraints while establishing pair-wise association between the detected parts. There is no publicly available dataset in the research community that provides a quantitative test bed for part detection and tracking of multiple mice, and we here introduce a new challenging Multi-Mice PartsTrack dataset that is made of complex behaviors. Finally, we evaluate our proposed approach against several baselines on our new datasets, where the results show that our method outperforms the other state-of-the-art approaches in terms of accuracy. We also demonstrate the generalization ability of the proposed approach on tracking zebra and locust.
Zheheng Jiang, Long Chen 0019, Xiangrong Zhang, Xiangyuan Lan, Danny Crookes, Ming-Hsuan Yang 0001, Huiyu Zhou 0001
IEEE Trans. Neural Networks Learn. Syst.9
2022 Semantically Contrastive Learning for Low-Light Image Enhancement
abstract
Low-light image enhancement (LLE) remains challenging due to the unfavorable prevailing low-contrast and weak-visibility problems of single RGB images. In this paper, we respond to the intriguing learning-related question -- if leveraging both accessible unpaired over/underexposed images and high-level semantic guidance, can improve the performance of cutting-edge LLE models? Here, we propose an effective semantically contrastive learning paradigm for LLE (namely SCL-LLE). Beyond the existing LLE wisdom, it casts the image enhancement task as multi-task joint learning, where LLE is converted into three constraints of contrastive learning, semantic brightness consistency, and feature preservation for simultaneously ensuring the exposure, texture, and color consistency. SCL-LLE allows the LLE model to learn from unpaired positives (normal-light)/negatives (over/underexposed), and enables it to interact with the scene semantics to regularize the image enhancement network, yet the interaction of high-level semantic knowledge and the low-level signal prior is seldom investigated in previous methods. Training on readily available open data, extensive experiments demonstrate that our method surpasses the state-of-the-arts LLE models over six independent cross-scenes datasets. Moreover, SCL-LLE's potential to benefit the downstream semantic segmentation under extremely dark conditions is discussed. Source Code: https://github.com/LingLIx/SCL-LLE.
Dong Liang 0008, Ling Li 0010, Mingqiang Wei, Wenhan Yang, Huiyu Zhou 0001
AAAI8
2022 Polycentric Clustering and Structural Regularization for Source-free Unsupervised Domain Adaptation
Ningzhong Liu, Huiyu Zhou 0001
BMVC4
2022 Attention Guided Network for Salient Object Detection in Optical Remote Sensing Images
Ningzhong Liu, Yetong Bian, Jun Cen, Huiyu Zhou 0001
ICANN (1)6
2022 A lightweight multi-scale context network for salient object detection in optical remote sensing images
abstract
Due to the more dramatic multi-scale variations and more complicated foregrounds and backgrounds in optical remote sensing images (RSIs), the salient object detection (SOD) for optical RSIs becomes a huge challenge. However, different from natural scene images (NSIs), the discussion on the optical RSI SOD task still remains scarce. In this paper, we propose a multi-scale context network, namely MSCNet, for SOD in optical RSIs. Specifically, a multi-scale context extraction module is adopted to address the scale variation of salient objects by effectively learning multi-scale contextual information. Meanwhile, in order to accurately detect complete salient objects in complex backgrounds, we design an attention-based pyramid feature aggregation mechanism for gradually aggregating and refining the salient regions from the multi-scale context extraction module. Extensive experiments on two benchmarks demonstrate that MSCNet achieves competitive performance with only 3.26M parameters. The code will be available at https://github.com/NuaaYH/MSCNet.
Ningzhong Liu, Yetong Bian, Jun Cen, Huiyu Zhou 0001
ICPR6
2022 Salient Skin Lesion Segmentation via Dilated Scale-Wise Feature Fusion Network
abstract
Skin lesion detection in dermoscopic images is essential in the accurate and early diagnosis of skin cancer by a computerized apparatus. Current skin lesion segmentation approaches show poor performance in challenging circumstances such as indistinct lesion boundaries, low contrast between the lesion and the surrounding area, or heterogeneous background that causes over/under segmentation of the skin lesion. To accurately recognize the lesion from the neighboring regions, we propose a dilated scale-wise feature fusion network based on convolution factorization. Our network is designed to simultaneously extract features at different scales which are systematically fused for better detection. The proposed model has satisfactory accuracy and efficiency. Various experiments for lesion segmentation are performed along with comparisons with the state-of-the-art models. Our proposed model consistently showcases state-of-the-art results.
Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002, Eric Granger, Huiyu Zhou 0001
ICPR5
2022 Absolute Wrong Makes Better: Boosting Weakly Supervised Object Detection via Negative Deterministic Information
abstract
Weakly supervised object detection (WSOD) is a challenging task, in which image-level labels (e.g., categories of the instances in the whole image) are used to train an object detector. Many existing methods follow the standard multiple instance learning (MIL) paradigm and have achieved promising performance. However, the lack of deterministic information leads to part domination and missing instances. To address these issues, this paper focuses on identifying and fully exploiting the deterministic information in WSOD. We discover that negative instances (i.e. absolutely wrong instances), ignored in most of the previous studies, normally contain valuable deterministic information. Based on this observation, we here propose a negative deterministic information (NDI) based method for improving WSOD, namely NDI-WSOD. Specifically, our method consists of two stages: NDI collecting and exploiting. In the collecting stage, we design several processes to identify and distill the NDI from negative instances online. In the exploiting stage, we utilize the extracted NDI to construct a novel negative contrastive learning mechanism and a negative guided instance selection strategy for dealing with the issues of part domination and missing instances, respectively. Experimental results on several public benchmarks including VOC 2007, VOC 2012 and MS COCO show that our method achieves satisfactory performance.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao
IJCAI5
2022 A Task-Aware Dual Similarity Network for Fine-Grained Few-Shot Learning
Ningzhong Liu, Huiyu Zhou 0001
PRICAI (3)4
2022 An anonymous authentication and key agreement protocol in smart living
Fengyin Li, Xinying Yu, Yuhong Sun, Huiyu Zhou 0001
Comput. Commun.7
2022 SecDH: Security of COVID-19 images based on data hiding with PCA
Om Prakash Singh, Amit Kumar Singh 0001, Amrit Kumar Agrawal, Huiyu Zhou 0001
Comput. Commun.4
2022 Tesia: A trusted efficient service evaluation model in Internet of things based on improved aggregation signature
abstract
Summary Service evaluation model is an essential ingredient in service‐oriented Internet of things (IoT) architecture. Generally, traditional models allow each user to submit their comments with respect to IoT services individually. However, these kind of models are fragile to resist various attacks, like comment denial attacks, and Sybil attacks, which may decrease the comments submission rate. In this article, we propose a new aggregation digital signature scheme to resolve the problem of comments aggregation, which may aggregate different comments into one with high efficiency and security level. Based on the new aggregation digital signature scheme, we further put forward a new service evaluation model named Tesia allowing specific users to submit the comments as a group in IoT networks. More specifically, they aggregate comments and assign one user as a submitter to submit these comments. In addition, we introduce the synchronization token mechanism into the new service evaluation model, to assure that all users in the group may sign their comments one by one, and the last one who receives the token is assigned as the final submitter. Tesia has more acceptable robustness and can greatly improve the comments submission rate with rather lower submission delay time.
Fengyin Li, Rui Ge 0004, Huiyu Zhou 0001, Zhongxing Liu, Xiaomei Yu
Concurr. Comput. Pract. Exp.3
2022 BacklitNet: A dataset and network for backlit image enhancement
Xiaoqian Lv, Shengping Zhang, Qinglin Liu, Haozhe Xie, Bineng Zhong 0001, Huiyu Zhou 0001
Comput. Vis. Image Underst.6
2022 DE-RSTC: A rational secure two-party computation protocol based on direction entropy
abstract
Rational secure multi-party computation means two or more rational parties complete a function on private inputs. Unfortunately, players sending false information can prevent the protocol from executing correctly, which will destroy the fairness of the protocol. To ensure the fairness of the protocol, the existing works on achieving fairness by specific utility functions. In this paper, we leverage game theory to propose the direction entropy-based solution. To this end, we utilize the direction entropy to examine the player's strategy uncertainty and quantify its strategy from different dimensions. Then, we provide mutual information to construct a new utility for the players. What's more, we measure the mutual information of players to appraise their strategies. By analyzing and proofing of protocol, we show that the protocol reaches a Nash equilibrium when players choose a cooperative strategy. Furthermore, we solve the fairness of the protocol. Compared to the previous approaches, our protocol is not required deposits and design-specific utility functions.
Yuling Chen 0002, Xianmin Wang, Huiyu Zhou 0001
Int. J. Intell. Syst.5
2022 PSSPR: A source location privacy protection scheme based on sector phantom routing in WSNs
abstract
Source location privacy (SLP) protection is an emerging research topic in wireless sensor networks. Because the source location represents the valuable information of the target being monitored and tracked, it is of great practical significance to achieve a high degree of privacy of the source location. Although many studies based on phantom nodes have alleviates the protection of SLP to some extent. It is urgent to solve the problems, such as complicate the ac path between nodes, improve the centralized distribution of phantom nodes near the source nodes and reduce the network communication overhead. In this paper, protection scheme based on sector phantom routing (PSSPR) routing is proposed as a visible approach to address SLP issues. We use the coordinates of the center node V to divide sector domain, which act an important role in generating a new phantom node. The phantom nodes perform specified routing policies to ensure that they can choose various locations. In addition, the directed random route can ensure that data packets avoid the visible range when they move to the sink node hop by hop. Thus, the source location is protected. Theoretical analysis and simulation experiments show that this protocol achieves higher security of source node location with less communication overhead.
Yuling Chen 0002, Yixian Yang, Tao Li 0043, Xinxin Niu, Huiyu Zhou 0001
Int. J. Intell. Syst.6
2022 Lattice-based batch authentication scheme with dynamic identity revocation in VANET
abstract
Aggregate signatures allow someone to aggregate multiple signatures into one signature, which is suitable for resource-constrained and computationally inefficient environments. Identify-based aggregate signature can solve the storage problem of public key certificates while achieving efficient signature verification. However, in most of the identity-based aggregate signature schemes, the user identity revocation process is time-consuming and cannot resist quantum attacks. To solve above problems, this paper proposes a lattice-based aggregate signature scheme with dynamic identity revocation by combining lattice-based cryptography and an aggregate signature scheme. The security of the proposed lattice-based aggregate signature scheme with dynamic identity revocation has been proved in the random oracle model. In addition, the verification efficiency of the aggregate signature has been improved compared with multiple different signatures. Much of the data transfer in Vehicular Ad Hoc Network (VANET) is carried out wirelessly, which makes VANET vulnerable to identity spoofing attacks. Identity authentication technology can prevent attackers from impersonating legitimate users, thus ensuring the security of VANET. Based on the proposed lattice-based aggregate signature scheme with dynamic identity revocation, this paper proposes a lattice-based batch authentication scheme with dynamic identity revocation in VANET. Through the proposed batch authentication scheme, we can effectively resist the impersonation attack of VANET in the quantum computer environment, and the efficiency of authentication is improved.
Fengyin Li, Huiyu Zhou 0001, Xiaoying Wang 0007, Qintai Yang
Int. J. Intell. Syst.4
2022 Privacy-aware PKI model with strong forward security
abstract
With the development of network technology, privacy protection and users anonymity become a new research hotspot. The existing blockchain privacy-aware public key infrastructure (PKI) model can ensure the privacy of users in the authentication process to a certain extent, but there are still problems of the storage and leakage of users' keys. This paper first proposes a strong forward-secure ring signature scheme based on RSA, which ensures the anonymity of the signing users and the forward-backward security of the keys. Then, by introducing the ring signature technology into the privacy-aware PKI model, this paper proposes a privacy-aware PKI model with strong forward security based on block chains, which not only ensures the users' identity privacy, but also solves the problem of the storage and leakage of the users' keys, greatly improving the success rate and security of the users' identity authentication. Finally, this paper applies the proposed PKI model to anonymous transactions, designs a privacy-aware anonymous transaction model with strong forward security, realizing anonymous transactions without relying on trusted third parties, and implementing users' privacy protection.
Fengyin Li, Zhongxing Liu, Tao Li 0043, Hongwei Ju, Huiyu Zhou 0001
Int. J. Intell. Syst.6
2022 Intelligent federated learning on lattice-based efficient heterogeneous signcryption
abstract
Signcryption technology combines signature and encryption operations in a single step to achieve message authentication and confidentiality. The ordinary signcryption technology cannot realize communication between two different cryptographic systems. Therefore, to implement efficient communication between different cryptosystems and resist quantum attacks, this paper proposes a lattice-based efficient heterogeneous signcryption scheme. The heterogeneous signcryption scheme is proved to be secure assuming the hardness of small integer solution and learning with errors problems. Then this paper applies the lattice-based efficient heterogeneous signcryption scheme to the federated learning system to achieve the transmission of confidential information, and designs an intelligent federated learning system on lattice-based efficient heterogeneous signcryption. This system realizes federated learning and the quantum security of data transmission while preserving private data.
Fengyin Li, Guangshun Li, Mengjiao Yang 0003, Huiyu Zhou 0001
Int. J. Intell. Syst.5
2022 IPSadas: Identity-privacy-aware secure and anonymous data aggregation scheme
abstract
Intelligent systems are technologically advanced machines that can sense and respond to the surrounding environment. They have been widely used in medicine, military, transportation, automation, and other fields. However, when these systems deal with their environments, problems such as leakage of identities may occur. The adversary can damage the system communication and attack important nodes. To handle resource-constrained wireless sensor network environments, we propose a secure and anonymous data aggregation scheme. First, based on the bilinear mapping operation and onion routing concepts, we propose a key negotiation and secure information transmission scheme, which conducts confidential transmission and anonymous forwarding of messages in data aggregation. Second, an aggregation routing scheme based on link direction and residual energy is proposed to pledge messages that can arrive the base station without passing through many nodes, which saves network resources to a certain extent. Third, on the basis of the first two contributions, we propose an identity-privacy-aware secure and anonymous data aggregation scheme that protects the identity's privacy. This scheme can conceal the real identity of important nodes and protect the anonymity of messages and link relationships. In addition, an anonymous identity update and synchronization scheme is also proposed to ensure the reliability and security of communication. Meanwhile, our performance evaluations and simulations show that the proposed framework is more effective than several standard schemes with respect to the ability against various attacks, security, and overhead.
Pei Ren, Fengyin Li, Ying Wang 0124, Huiyu Zhou 0001, Peiyu Liu 0001
Int. J. Intell. Syst.4
2022 Contrastive hashing with vision transformer for image retrieval
abstract
Hashing techniques have attracted considerable attention owing to their advantages of efficient computation and economical storage. However, it is still a challenging problem to generate more compact binary codes for promising performance. In this paper, we propose a novel contrastive vision transformer hashing method, which seamlessly integrates contrastive learning and vision transformers (ViTs) with hash technology into a well-designed model to learn informative features and compact binary codes simultaneously. First, we modify the basic contrastive learning framework by designing several hash layers to meet the specific requirement of hash learning. In our hash network, ViTs are applied as backbones for feature learning, which is rarely performed in existing hash learning methods. Then, we design a multiobjective loss function, in which contrastive loss explores discriminative features by maximizing agreement between different augmented views from the same image, similarity preservation loss performs pairwise semantic preservation to enhance the representative capabilities of hash codes, and quantization loss controls the quantitative error. Hence, we can facilitate end-to-end joint training to improve the retrieval performance. The encouraging experimental results on three widely used benchmark databases demonstrate the superiority of our algorithm compared with several state-of-the-art hashing algorithms.
Xiuxiu Ren, Xiangwei Zheng 0001, Huiyu Zhou 0001
Int. J. Intell. Syst.3
2022 A Novel deep neural network-based emotion analysis system for automatic detection of mild cognitive impairment in the elderly
Zixiang Fei, Erfu Yang, Leijian Yu, Huiyu Zhou 0001, Wenju Zhou
Neurocomputing5
2022 BSM-ether: Bribery selfish mining in blockchain-based healthcare systems
Minghao Zhao 0001, Xueyang Han, Huiyu Zhou 0001, Xiaoying Wang 0007, Arthur Sandor Voundi Koe
Inf. Sci.5
2022 Discriminative feature mining hashing for fine-grained image retrieval
Wenxi Lang, Can Xu 0006, Ningzhong Liu, Huiyu Zhou 0001
J. Vis. Commun. Image Represent.5
2022 Multilevel Feature Fusion Networks With Adaptive Channel Dimensionality Reduction for Remote Sensing Scene Classification
abstract
Scene classification in very high-resolution (VHR) remote sensing (RS) images is a challenging task due to the complex and diverse content of the images. Recently, convolution neural networks (CNNs) have been utilized to tackle this task. However, CNNs cannot fully meet the needs of scene classification due to clutters and small objects in VHR images. To handle these challenges, this letter presents a novel multilevel feature fusion (MLFF) network with adaptive channel dimensionality reduction for RS scene classification. Specifically, an adaptive method is designed for channel dimensionality reduction of high-dimensional features. Then, an MLFF module is introduced to fuse the features in an efficient way. Experiments on three widely used data sets show that our model outperforms several state-of-the-art methods in terms of both accuracy and stability.
Xin Wang 0068, Lin Duan, Aiye Shi, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Dropout-Based Adversarial Training Networks for Remote Sensing Scene Classification
abstract
Scene classification in remote sensing (RS) images is a challenging task due to the lack of well-labeled data. Recently, deep transfer learning (DTL) has been proposed to handle this task. However, the intraclass variations and interclass similarities remain challenges. To handle these challenges, this letter presents a novel dropout-based adversarial training network (DATN) for RS scene classification. Specifically, a dropout-based label classifier (DLC) module is designed to reduce the selection of ambiguous features on class boundaries. Then, a dropout-based domain discriminator (DDD) module is constructed to capture multimodal structures of RS images so as to achieve fine-grained alignment between cross-domain distributions. Third, a joint distribution of features and labels is built to further enhance the performance. Experiments on seven public RS datasets show that our model outperforms several states of the art (SOTAs) under different conditions. The code of our method is publicly available athttps://github.com/WangXin81/DATN-Submitted-to-IEEE-GRSL.
Xin Wang 0068, Zhipeng Mao, Aiye Shi, Zhilu Zhang 0003, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Multichannel SAR Moving Target Detection via RPCA-Net
abstract
Ground moving target indication (GMTI), as a challenging task for synthetic aperture radar (SAR) systems, keeps drawing considerable attention. Robust principal component analysis (RPCA) aiming at separating low-rank and sparse components has been successfully employed in SAR systems for GMTI recently. However, its practical application is limited by the heavy computational burden as well as the requirement of manual parameter modification. To cope with this problem, a fast and free of presetting parameters RPCA network (RPCA-Net) is proposed for SAR-GMTI under strong clutter background. In the proposed method, a novel RPCA model is first introduced, where not only the low-rank and sparse terms but also the errors in practical SAR systems are taken into account. Moreover, the low-rank factorization plus scaled gradient descent (ScaledGD) is also employed to acquire low-rank clutter background rather than singular value decomposition (SVD). Then, we parameterize our proposed RPCA model and unfold it as a feedforward neural network (FNN) to acquire the iterative parameters through backpropagation. Compared to the GMTI methods based on traditional RPCA models, our proposed RPCA-Net can provide a higher detection ability and faster convergence without presetting parameters empirically. Experiments on two groups of measured data collected by airborne SAR systems validate the superior performance of the proposed RPCA-Net.
Xifeng Zhang, Di Wu 0015, Daiyin Zhu, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Inferred box harmonization and aggregation for degraded face detection in crowds
Dong Liang 0008, Qixiang Geng, Huiyu Zhou 0001, Shun'ichi Kaneko
Multim. Tools Appl.4
2022 SWIPENET: Object detection in noisy underwater scenes
abstract
Deep learning based object detection methods have achieved promising performance in controlled environments. However, these methods lack sufficient capabilities to handle underwater object detection due to these challenges: (1) images in the underwater datasets and real applications are blurry whilst accompanying severe noise that confuses the detectors and (2) objects in real applications are usually small. In this paper, we propose a Sample-WeIghted hyPEr Network (SWIPENET), and a novel training paradigm named Curriculum Multi-Class Adaboost (CMA), to address these two problems at the same time. Firstly, the backbone of SWIPENET produces multiple high resolution and semantic-rich Hyper Feature Maps, which significantly improve small object detection. Secondly, inspired by the human education process that drives the learning from easy to hard concepts, we propose the noise-robust CMA training paradigm that learns the clean data first and then move on to learns the diverse noisy data. Experiments on four underwater object detection datasets show that the proposed SWIPENET+CMA framework achieves better or competitive accuracy in object detection against several state-of-the-art approaches.
Long Chen 0019, Feixiang Zhou, Shengke Wang, Junyu Dong, Ning Li 0012, Haiping Ma, Xin Wang 0068, Huiyu Zhou 0001
Pattern Recognit.8
2022 Video-Based Cross-Modal Auxiliary Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis has a wide range of applications due to its information complementarity in multimodal interactions. Previous works focus more on investigating efficient joint representations, but they rarely consider the insufficient unimodal features extraction and data redundancy of multimodal fusion. In this paper, a Video-based Cross-modal Auxiliary Network (VCAN) is proposed, which is comprised of an audio features map module and a cross-modal selection module. The first module is designed to substantially increase feature diversity in audio feature extraction, aiming to improve classification accuracy by providing more comprehensive acoustic representations. To empower the model to handle redundant visual features, the second module is addressed to efficiently filter the redundant visual frames during integrating audiovisual data. Moreover, a classifier group consisting of several image classification networks is introduced to predict sentiment polarities and emotion categories. Extensive experimental results on RAVDESS, CMU-MOSI, and CMU-MOSEI benchmarks indicate that VCAN is significantly superior to the state-of-the-art methods for improving the classification accuracy of multimodal sentiment analysis.
Rongfei Chen, Wenju Zhou, Yang Li 0129, Huiyu Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Wallpaper Texture Generation and Style Transfer Based on Multi-Label Semantics
abstract
Textures contain a wealth of image information and are widely used in various fields such as computer graphics and computer vision. With the development of machine learning, the texture synthesis and generation have been greatly improved. As a very common element in everyday life, wallpapers contain a wealth of texture information, making it difficult to annotate with a simple single label. Moreover, wallpaper designers spend significant time to create different styles of wallpaper. For this purpose, this paper proposes to describe wallpaper texture images by using multi-label semantics. Based on these labels and generative adversarial networks, we present a framework for perception driven wallpaper texture generation and style transfer. In this framework, a perceptual model is trained to recognize whether the wallpapers produced by the generator network are sufficiently realistic and have the attribute designated by given perceptual description; these multi-label semantic attributes are treated as condition variables to generate wallpaper images. The generated wallpaper images can be converted to those with well-known artist styles using CycleGAN. Finally, using the aesthetic evaluation method, the generated wallpaper images are quantitatively measured. The experimental results demonstrate that the proposed method can generate wallpaper textures conforming to human aesthetics and have artistic characteristics.
Ying Gao 0005, Xiaohan Feng, Tiange Zhang, Eric Rigall, Huiyu Zhou 0001, Lin Qi 0004, Junyu Dong
IEEE Trans. Circuits Syst. Video Technol.5
2022 Gaussian Dynamic Convolution for Efficient Single-Image Segmentation
abstract
Interactive single-image segmentation is ubiquitous in the scientific and commercial imaging software. Lightweight neural network is one practical and effective way to accomplish the single-image segmentation task. This work focuses on the single-image segmentation problem only with some seeds such as scribbles. Inspired by the dynamic receptive field in the human being’s visual system, we propose the Gaussian dynamic convolution (GDC) to fast and efficiently aggregate the contextual information for neural networks. The core idea is randomly selecting the spatial sampling area according to the Gaussian distribution offsets. Our GDC can be easily used as a module to build lightweight or complex segmentation networks. We adopt the proposed GDC to address the typical single-image segmentation tasks. Furthermore, we also build a Gaussian dynamic pyramid Pooling to show its potential and generality in common semantic segmentation. Experiments demonstrate that the GDC outperforms other existing convolutions on three benchmark segmentation datasets including Pascal-Context, Pascal-VOC 2012, and Cityscapes. Additional experiments are also conducted to illustrate that the GDC can produce richer and more vivid features compared with other convolutions. In general, our GDC is conducive to the convolutional neural networks to form an overall impression of the image.
Xin Sun 0003, Changrui Chen, Junyu Dong, Huiyu Zhou 0001, Sheng Chen 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Continuous Prediction of Lower-Limb Kinematics From Multi-Modal Biomedical Signals
abstract
The fast-growing techniques of measuring and fusing multi-modal biomedical signals enable advanced motor intent decoding schemes of lower-limb exoskeletons, meeting the increasing demand for rehabilitative or assistive applications of take-home healthcare. Challenges of exoskeletons’ motor intent decoding schemes remain in making a continuous prediction to compensate for the hysteretic response caused by mechanical transmission. In this paper, we solve this problem by proposing an ahead-of-time continuous prediction of lower-limb kinematics, with the prediction of knee angles during level walking as a case study. Firstly, an end-to-end kinematics prediction network(KinPreNet),1consisting of a feature extractor and an angle predictor, is proposed and experimentally compared with features and methods traditionally used in ahead-of-time prediction of gait phases. Secondly, inspired by the electromechanical delay(EMD), we further explore our algorithm’s capability of compensating response delay of mechanical transmission by validating the performance of the different sections of prediction time. And we experimentally reveal the time boundary of compensating the hysteretic response. Thirdly, a comparison of employing EMG signals or not is performed to reveal the EMG and kinematic signals’ collaborated contributions to the continuous prediction. During the experiments, EMG signals of nine muscles and knee angles calculated from inertial measurement unit (IMU) signals are recorded from ten healthy subjects. Our algorithm can predict knee angles with the averaged RMSE of 3.98 deg which is better than the 15.95-deg averaged RMSE of utilizing the traditional methods of ahead-of-time prediction. The best prediction time is in the interval of 27ms and 108ms. To the best of our knowledge, this is the first study of continuously predicting lower-limb kinematics in an ahead-of-time manner based on the electromechanical delay (EMD).
Chunzhi Yi, Feng Jiang 0001, Shengping Zhang, Hao Guo 0015, Chifu Yang, Zhen Ding, Baichun Wei, Xiangyuan Lan, Huiyu Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.9
2022 Structured Context Enhancement Network for Mouse Pose Estimation
abstract
Automated analysis of mouse behaviours is crucial for many applications in neuroscience. However, quantifying mouse behaviours from videos or images remains a challenging problem, where pose estimation plays an important role in describing mouse behaviours. Although deep learning based methods have made promising advances in human pose estimation, they cannot be directly applied to pose estimation of mice due to different physiological natures. Particularly, since mouse body is highly deformable, it is a challenge to accurately locate different keypoints on the mouse body. In this paper, we propose a novel Hourglass network based model, namely Graphical Model based Structured Context Enhancement Network (GM-SCENet) where two effective modules, i.e., Structured Context Mixer (SCM) and Cascaded Multi-level Supervision (CMLS) are subsequently implemented. SCM can adaptively learn and enhance the proposed structured context information of each mouse part by a novel graphical model that takes into account the motion difference between body parts. Then, the CMLS module is designed to jointly train the proposed SCM and the Hourglass network by generating multi-level information, increasing the robustness of the whole network. Using the multi-level prediction information from SCM and CMLS, we develop an inference method to ensure the accuracy of the localisation results. Finally, we evaluate our proposed approach against several baselines on our Parkinson’s Disease Mouse Behaviour (PDMB) and the standard DeepLabCut Mouse Pose datasets. The experimental results show that our method achieves better or competitive performance against the other state-of-the-art approaches.
Feixiang Zhou, Zheheng Jiang, Long Chen 0019, Zhile Yang, Haikuan Wang, Minrui Fei, Ling Li 0010, Huiyu Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.11
2022 Secure Control of Networked Control Systems Using Dynamic Watermarking
abstract
We here investigate the secure control of networked control systems developing a new dynamic watermarking (DW) scheme. First, the weaknesses of the conventional DW scheme are revealed, and the tradeoff between the effectiveness of false data injection attack (FDIA) detection and system performance loss is analyzed. Second, we propose a new DW scheme, and its attack detection capability is interrogated using the additive distortion power of a closed-loop system. Furthermore, the FDIA detection effectiveness of the closed-loop system is analyzed using auto/cross-covariance of the signals, where the positive correlation between the FDIA detection effectiveness and the watermarking intensity is measured. Third, the tolerance capacity of FDIA against the closed-loop system is investigated, and theoretical analysis shows that the system performance can be recovered from FDIA using our new DW scheme. Finally, the experimental results from a networked inverted pendulum system demonstrate the validity of our proposed scheme.
Dajun Du, Changda Zhang, Xue Li 0028, Minrui Fei, Huiyu Zhou 0001
IEEE Trans. Cybern.6
2022 Learning the Precise Feature for Cluster Assignment
abstract
Clustering is one of the fundamental tasks in computer vision and pattern recognition. Recently, deep clustering methods (algorithms based on deep learning) have attracted wide attention with their impressive performance. Most of these algorithms combine deep unsupervised representation learning and standard clustering together. However, the separation of representation learning and clustering will lead to suboptimal solutions because the two-stage strategy prevents representation learning from adapting to subsequent tasks (e.g., clustering according to specific cues). To overcome this issue, efforts have been made in the dynamic adaption of representation and cluster assignment, whereas current state-of-the-art methods suffer from heuristically constructed objectives with the representation and cluster assignment alternatively optimized. To further standardize the clustering problem, we audaciously formulate the objective of clustering as finding a precise feature as the cue for cluster assignment. Based on this, we propose a general-purpose deep clustering framework, which radically integrates representation learning and clustering into a single pipeline for the first time. The proposed framework exploits the powerful ability of recently developed generative models for learning intrinsic features, and imposes an entropy minimization on the distribution of the cluster assignment by a dedicated variational algorithm. The experimental results show that the performance of the proposed method is superior, or at least comparable to, the state-of-the-art methods on the handwritten digit recognition, fashion recognition, face recognition, and object recognition benchmark datasets.
Yanhai Gan, Xinghui Dong, Huiyu Zhou 0001, Feng Gao 0005, Junyu Dong
IEEE Trans. Cybern.3
2022 Distributed Fusion Estimation for Stochastic Uncertain Systems With Network-Induced Complexity and Multiple Noise
abstract
This article investigates an issue of distributed fusion estimation under network-induced complexity and stochastic parameter uncertainties. First, a novel signal selection method based on event trigger is developed to handle network-induced packet dropouts, as well as packet disorders resulting from random transmission delays, where the${H_{2}}/{H_{\infty } }$performance of the system is analyzed in different noise environments. In addition, a linear delay compensation strategy is further employed for solving the complex network-induced problem, which may deteriorate system performance. Moreover, a weighted fusion scheme is used to integrate multiple resources through an error cross-covariance matrix. Several case studies validate the proposed algorithm and demonstrate satisfactory system performance in target tracking.
Li Liu 0023, Wenju Zhou, Minrui Fei, Zhile Yang, Hongyong Yang, Huiyu Zhou 0001
IEEE Trans. Cybern.6
2022 Semantic Attention and Scale Complementary Network for Instance Segmentation in Remote Sensing Images
abstract
In this article, we focus on the challenging multicategory instance segmentation problem in remote sensing images (RSIs), which aims at predicting the categories of all instances and localizing them with pixel-level masks. Although many landmark frameworks have demonstrated promising performance in instance segmentation, the complexity in the background and scale variability instances still remain challenging, for instance, segmentation of RSIs. To address the above problems, we propose an end-to-end multicategory instance segmentation model, namely, the semantic attention (SEA) and scale complementary network, which mainly consists of a SEA module and a scale complementary mask branch (SCMB). The SEA module contains a simple fully convolutional semantic segmentation branch with extra supervision to strengthen the activation of interest instances on the feature map and reduce the background noise's interference. To handle the undersegmentation of geospatial instances with large varying scales, we design the SCMB that extends the original single mask branch to trident mask branches and introduces complementary mask supervision at different scales to sufficiently leverage the multiscale information. We conduct comprehensive experiments to evaluate the effectiveness of our proposed method on the iSAID dataset and the NWPU Instance Segmentation dataset and achieve promising performance.
Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Cybern.7
2022 Multimodal Gait Recognition for Neurodegenerative Diseases
abstract
In recent years, single modality-based gait recognition has been extensively explored in the analysis of medical images or other sensory data, and it is recognized that each of the established approaches has different strengths and weaknesses. As an important motor symptom, gait disturbance is usually used for diagnosis and evaluation of diseases; moreover, the use of multimodality analysis of the patient's walking pattern compensates for the one-sidedness of single modality gait recognition methods that only learn gait changes in a single measurement dimension. The fusion of multiple measurement resources has demonstrated promising performance in the identification of gait patterns associated with individual diseases. In this article, as a useful tool, we propose a novel hybrid model to learn the gait differences between three neurodegenerative diseases, between patients with different severity levels of Parkinson's disease, and between healthy individuals and patients, by fusing and aggregating data from multiple sensors. A spatial feature extractor (SFE) is applied to generating representative features of images or signals. In order to capture temporal information from the two modality data, a new correlative memory neural network (CorrMNN) architecture is designed for extracting temporal features. Afterward, we embed a multiswitch discriminator to associate the observations with individual state estimations. Compared with several state-of-the-art techniques, our proposed framework shows more accurate classification results.
Aite Zhao, Junyu Dong, Lin Qi 0004, Qianni Zhang, Ning Li 0012, Xin Wang 0068, Huiyu Zhou 0001
IEEE Trans. Cybern.8
2022 Ellipse Encoding for Arbitrary-Oriented SAR Ship Detection Based on Dynamic Key Points
abstract
In recent years, there has been growing interest in developing oriented bounding-box (OBB) based deep learning approaches to detect arbitrary-oriented ship targets in synthetic aperture radar (SAR) images. However, most existing OBB-based detection methods suffer from boundary discontinuity problems for bounding box angle prediction and key point regression challenges. In this paper, we present a novel OBB-based detection algorithm that utilizes ellipse encoding to effectively exploit the geometric and scattering properties of ship targets. Specifically, the ship contour is fitted by an OBB inscribed ellipse that is encoded as a set of distances between dynamic key points on the bow and target center. By combining the bow angle interval and the decoding process, the negative impact of the boundary discontinuity problem is avoided. In addition, we propose an elliptical Gaussian distribution heatmap and a pooling strategy termed double peaks max-pooling (DPM), to deal with the challenge of separating densely distributed ships in inshore scenes. The former can enhance the heatmap’s ship-side score gap between neighboring ship targets, while the latter can solve the problem of target center responses being suppressed after max-pooling. Simulation experiments conducted on the benchmark Rotating SAR Ship Detection Dataset (RSSDD) and Rotated Ship Detection Dataset in SAR Images (RSDD-SAR) demonstrate the superior performance of our method for ship target detection compared to several state-of-the-art OBB-based algorithms. Ablation experiments show that elliptical Gaussian distribution heatmap and DPM can further improve the inshore detection performance.
Fei Gao 0005, Yiyang Huo, Jinping Sun, Amir Hussain 0001, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Anchor Retouching via Model Interaction for Robust Object Detection in Aerial Images
abstract
Object detection has made tremendous strides in computer vision. Small object detection with appearance degradation is a prominent challenge, especially for aerial observations. To collect sufficient positive/negative samples for heuristic training, most object detectors preset region anchors in order to calculate intersection-over-union (IoU) against the ground-truth data. In this case, small objects are frequently abandoned or mislabeled. In this article, we present an effective dynamic enhancement anchor network (DEA-Net) to construct a novel training sample generator. Different from the other state-of-the-art (SOTA) techniques, the proposed network leverages a sample discriminator to realize interactive sample screening between an anchor-based unit and an anchor-free unit to generate eligible samples. Besides, multi-task joint training with a conservative anchor-based inference scheme enhances the performance of the proposed model while reducing computational complexity. The proposed scheme supports both oriented and horizontal object detection tasks. Extensive experiments on two challenging aerial benchmarks (i.e., Dataset of Object deTection in Aerial images (DOTA) and HRSC2016) indicate that our method achieves SOTA performance in accuracy with moderate inference speed and computational overhead for training. On DOTA, our DEA-Net which integrated with the baseline of RoI-transformer surpasses the advanced method by 0.40% mean-average-precision (mAP) for oriented object detection with a weaker backbone network (ResNet-101 vs. ResNet-152) and 3.08% mAP for horizontal object detection with the same backbone. Besides, our DEA-Net which integrated with the baseline of ReDet achieves the SOTA performance by 80.37%. On HRSC2016, it surpasses the previous best model by 1.1% using only three horizontal anchors. The source code and the training set are made publicly available athttps://github.com/QxGeng/DEA-Net.
Dong Liang 0008, Qixiang Geng, Zongqi Wei, Dmitry A. Vorontsov, Ekaterina L. Kim, Mingqiang Wei, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Multipatch Feature Pyramid Network for Weakly Supervised Object Detection in Optical Remote Sensing Images
abstract
Object detection is a challenging task in remote sensing because objects only occupy a few pixels in the images, and the models are required to simultaneously learn object locations and detection. Even though the established approaches well perform for the objects of regular sizes, they achieve weak performance when analyzing small ones or getting stuck in the local minima (e.g. false object parts). Two possible issues stand in their way. First, the existing methods struggle to perform stably on the detection of small objects because of the complicated background. Second, most of the standard methods used hand-crafted features, and do not work well on the detection of objects parts of which are missing. We here address the above issues and propose a new architecture with a multiple patch feature pyramid network (MPFP-Net). Different from the current models that during training only pursue the most discriminative patches, in MPFPNet the patches are divided into class-affiliated subsets, in which the patches are related and based on the primary loss function, a sequence of smooth loss functions are determined for the subsets to improve the model for collecting small object parts. To enhance the feature representation for patch selection, we introduce an effective method to regularize the residual values and make the fusion transition layers strictly norm-preserving. The network contains bottom-up and crosswise connections to fuse the features of different scales to achieve better accuracy, compared to several state-of-the-art object detection models. Also, the developed architecture is more efficient than the baselines.
Pourya Shamsolmoali, Jocelyn Chanussot, Masoumeh Zareapoor, Huiyu Zhou 0001, Jie Yang 0002
IEEE Trans. Geosci. Remote. Sens.4
2022 Rotation Equivariant Feature Image Pyramid Network for Object Detection in Optical Remote Sensing Imagery
abstract
Detection of objects is extremely important in various aerial vision-based applications. Over the last few years, the methods based on convolution neural networks (CNNs) have made substantial progress. However, because of the large variety of object scales, densities, and arbitrary orientations, the current detectors struggle with the extraction of semantically strong features for small-scale objects by a predefined convolution kernel. To address this problem, we propose the rotation equivariant feature image pyramid network (REFIPN), an image pyramid network based on rotation equivariance convolution. The proposed model adopts single-shot detector in parallel with a lightweight image pyramid module (LIPM) to extract representative features and generate regions of interest in an optimization approach. The proposed network extracts feature in a wide range of scales and orientations by using novel convolution filters. These features are used to generate vector fields and determine the weight and angle of the highest-scoring orientation for all spatial locations on an image. By this approach, the performance for small-sized object detection is enhanced without sacrificing the performance for large-sized object detection. The performance of the proposed model is validated on two commonly used aerial benchmarks and the results show our proposed model can achieve state-of-the-art performance with satisfactory efficiency.
Pourya Shamsolmoali, Masoumeh Zareapoor, Jocelyn Chanussot, Huiyu Zhou 0001, Jie Yang 0002
IEEE Trans. Geosci. Remote. Sens.4
2022 Clutter Suppression for Wideband Radar STAP
abstract
Traditional space-time (ST) adaptive processing (STAP) theory is based on the assumption of narrowband or “zero-bandwidth,” where the decorrelation within the ST snapshot is ignored. However, with radar bandwidths increasing, this assumption becomes invalid due to the deteriorated decorrelation of the received signals within the ST snapshot. The decorrelation directly causes the dispersion of the received signals in both spatial and temporal domains, leading to the spreading of the clutter spectrum in the 2-D frequency (Doppler-spatial frequency) domain. With the spreading of the clutter spectrum, the clutter suppression notch in the traditional STAP filters is widened, resulting in a relative poor ability to detect slow-moving targets. In this article, we focus on the clutter suppression for wideband radar STAP. A generalized signal model of the ground clutter is first established for the wideband array radar. Using this outcome, we analyze the influence of bandwidth on the characteristics of the ground clutter and quantitatively describe the 2-D spreading of the ground clutter on the Doppler-spatial frequency plane. Moreover, the model of clutter covariance matrix for wideband STAP (W-STAP) is established. Finally, a 2-D keystone transform (KT) algorithm, referred to as ST KT (ST-KT), is proposed to eliminate the spreading of the ground clutter in the 2-D frequency domain caused by increasing bandwidths. Simulation results are employed to validate the theoretical analysis and verify the overperformance of the ST-KT based W-STAP method in terms of the output signal-to-clutter-plus-noise ratio (SCNR) of moving targets.
Di Wu 0015, Daiyin Zhu, Mingwei Shen 0002, Ning Li 0012, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Guest Editorial: Medical Data Security Solution for Healthcare Industries
abstract
Since smart healthcare systems are highly connected to advanced wearable devices, internet of things (IoT) and mobile internet, valuable patient information and other significant medical records are easily transmitted over the public network. The patient information and clinical records are also stored on the existing databases and local servers of hospitals and healthcare centres. These materials not only provide a reference for healthcare professionals to make correct decisions on the patients, but also provide a strong basis for other professionals to undertake effective treatment and develop plans for correct diagnosis. Furthermore, the databases may be used by various research communities for research, without any possibility of privacy violations. However, leaking of healthcare data is highly likely. Therefore, medical data security is becoming very important in smart healthcare.
Amit Kumar Singh 0001, Huiyu Zhou 0001, Stefano Berretti
IEEE Trans. Ind. Informatics2
2022 Color Alignment for Relative Color Constancy via Non-Standard References
abstract
Relative colour constancy is an essential requirement for many scientific imaging applications. However, most digital cameras differ in their image formations and native sensor output is usually inaccessible, e.g., in smartphone camera applications. This makes it hard to achieve consistent colour assessment across a range of devices, and that undermines the performance of computer vision algorithms. To resolve this issue, we propose a colour alignment model that considers the camera image formation as a black-box and formulates colour alignment as a three-step process: camera response calibration, response linearisation, and colour matching. The proposed model works with non-standard colour references, i.e., colour patches without knowing the true colour values, by utilising a novel balance-of-linear-distances feature. It is equivalent to determining the camera parameters through an unsupervised process. It also works with a minimum number of corresponding colour patches across the images to be colour aligned to deliver the applicable processing. Three challenging image datasets collected by multiple cameras under various illumination and exposure conditions, including one that imitates uncommon scenes such as scientific imaging, were used to evaluate the model. Performance benchmarks demonstrated that our model achieved superior performance compared to other popular and state-of-the-art methods.
Stuart Ferguson, Huiyu Zhou 0001, Chris Elliott 0002, Karen Rafferty
IEEE Trans. Image Process.3
2022 Parallel Complement Network for Real-Time Semantic Segmentation of Road Scenes
abstract
Real-time semantic segmentation is in intense demand for the application of autonomous driving. Most of the semantic segmentation models tend to use large feature maps and complex structures to enhance the representation power for high accuracy. However, these inefficient designs increase the amount of computational costs, which hinders the model to be applied on autonomous driving. In this paper, we propose a lightweight real-time segmentation model, named Parallel Complement Network (PCNet), to address the challenging task with fewer parameters. A Parallel Complement layer is introduced to generate complementary features with a large receptive field. It provides the ability to overcome the problem of similar feature encoding among different classes, and further produces discriminative representations. With the inverted residual structure, we design a Parallel Complement block to construct the proposed PCNet. Extensive experiments are carried out on challenging road scene datasets, i.e., CityScapes and CamVid, to make comparison against several state-of-the-art real-time segmentation models. The results show that our model has promising performance. Specifically, PCNet* achieves 72.9% Mean IoU on CityScapes using only 1.5M parameters and reaches 79.1 FPS with$1024\times 2048$resolution images on GTX 2080Ti. Moreover, our proposed system achieves the best accuracy when being trained from scratch.
Qingxuan Lv, Xin Sun 0003, Changrui Chen, Junyu Dong, Huiyu Zhou 0001
IEEE Trans. Intell. Transp. Syst.5
2022 Associated Spatio-Temporal Capsule Network for Gait Recognition
abstract
It is a challenging task to identify a person based on her/his gait patterns. State-of-the-art approaches rely on the analysis of temporal or spatial characteristics of gait, and gait recognition is usually performed on single modality data (such as images, skeleton joint coordinates, or force signals). Evidence has shown that using multi-modality data is more conducive to gait research. Therefore, we here establish an automated learning system, with an associated spatio-temporal capsule network (ASTCapsNet) trained on multi-sensor datasets, to analyze multimodal information for gait recognition. Specifically, we first design a low-level feature extractor and a high-level feature extractor for spatio-temporal feature extraction of gait with a novel recurrent memory unit and a relationship layer. Subsequently, a Bayesian model is employed for the decision-making of class labels. Extensive experiments on several public datasets (normal and abnormal gait) validate the effectiveness of the proposed ASTCapsNet, compared against several state-of-the-art methods.
Aite Zhao, Junyu Dong, Lin Qi 0004, Huiyu Zhou 0001
IEEE Trans. Multim.5
2022 Binary Representation via Jointly Personalized Sparse Hashing
abstract
Unsupervised hashing has attracted much attention for binary representation learning due to the requirement of economical storage and efficiency of binary codes. It aims to encode high-dimensional features in the Hamming space with similarity preservation between instances. However, most existing methods learn hash functions in manifold-based approaches. Those methods capture the local geometric structures (i.e., pairwise relationships) of data, and lack satisfactory performance in dealing with real-world scenarios that produce similar features (e.g., color and shape) with different semantic information. To address this challenge, in this work, we propose an effective unsupervised method, namely, Jointly Personalized Sparse Hashing (JPSH), for binary representation learning. To be specific, first, we propose a novel personalized hashing module, i.e., Personalized Sparse Hashing (PSH). Different personalized subspaces are constructed to reflect category-specific attributes for different clusters, adaptively mapping instances within the same cluster to the same Hamming space. In addition, we deploy sparse constraints for different personalized subspaces to select important features. We also collect the strengths of the other clusters to build the PSH module with avoiding over-fitting. Then, to simultaneously preserve semantic and pairwise similarities in our proposed JPSH, we incorporate the proposed PSH and manifold-based hash learning into the seamless formulation. As such, JPSH not only distinguishes the instances from different clusters but also preserves local neighborhood structures within the cluster. Finally, an alternating optimization algorithm is adopted to iteratively capture analytical solutions of the JPSH model. We apply the proposed representation learning algorithm JPSH to the similarity search task. Extensive experiments on four benchmark datasets verify that the proposed JPSH outperforms several state-of-the-art unsupervised hashing algorithms.
Chen Chen 0151, Rushi Lan, Licheng Liu, Zhenbing Liu, Huiyu Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2021 Dense Face Detection via High-level Context Mining
abstract
The appearance degradation caused by low resolution is the core problem of small face detection. Therefore, a natural approach is to assemble information from the context. This paper focuses on how to use high-level contextual information to improve the abilities of anchor-based detectors to detect dense and degenerate faces. We tap the spatial contextual information on the overall view based on the density map, and propose the prior of face co-occurrence for inferred bounding-boxes coordination. We also propose score-size-specific non-maximum suppression to replace the traditional non-maximum suppression at the end of anchor-based detectors. According to the inferred face boxes' quantity, score and size, the proposed synthetical solution reduces false positives and increases true positives. Our method does not require additional training, which is model-independent and can be embedded into existing face detectors. We also propose a dataset - Crowd Face for face detection, which is full of challenges. We expect to supply enough samples to highlight the difficulties of detecting dense and degenerate faces. We embed our proposed methods into state-of-the-art face detectors on massively benchmarked face datasets. Compared with the prior art on the WIDER FACE hard set, our method increase an Average Precision of 0.1 %-1.3%. On Crowd Face, it increases an Average Precision of 1 % – 6%. Dataset is available on: https://github.com/QxGeng/Crowd-Face.
Qixiang Geng, Dong Liang 0008, Huiyu Zhou 0001, Liyan Zhang 0001, Ningzhong Liu
FG3
2021 Robust Cross-Scene Foreground Segmentation in Surveillance Video
abstract
1Training only one deep model for large-scale cross-scene video foreground segmentation is challenging due to the off-the-shelf deep learning based segmentor relies on scene-specific structural information. This results in deep models that are scene-biased and evaluations that are scene-influenced. In this paper, we integrate dual modalities (foregrounds’ motion and appearance), and then eliminating features without representativeness of foreground through attention-module-guided selective-connection structures. It is in an end-to-end training manner and to achieve scene adaptation in the plug and play style. Experiments indicate the proposed method significantly outperforms the state-of-the-art deep models and background subtraction methods in un-trained scenes – LIMU and LASIESTA. Source Code is available at: https://github.com/WeiZongqi/HOFAM
Dong Liang 0008, Zongqi Wei, Huiyu Zhou 0001
ICME4
2021 A Dummy Location Selection Algorithm Based on Location Semantics and Physical Distance
Baopeng Ye, Yuling Chen 0002, Huiyu Zhou 0001, Xiaobin Qian
ISPEC4
2021 Adaptive Affinity Loss and Erroneous Pseudo-Label Refinement for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation has been continuously investigated in the last ten years, and majority of the established technologies are based on supervised models. In recent years, image-level weakly supervised semantic segmentation (WSSS), including single- and multi-stage process, has attracted large attention due to data labeling efficiency. In this paper, we propose to embed affinity learning of multi-stage approaches in a single-stage model. To be specific, we introduce an adaptive affinity loss to thoroughly learn the local pairwise affinity. As such, a deep neural network is used to deliver comprehensive semantic information in the training phase, whilst improving the performance of the final prediction module. On the other hand, considering the existence of errors in the pseudo labels, we propose a novel label reassign loss to mitigate over-fitting. Extensive experiments are conducted on the PASCAL VOC 2012 dataset to evaluate the effectiveness of our proposed approach that outperforms other standard single-stage methods and achieves comparable performance against several multi-stage methods.
Xiangrong Zhang, Zelin Peng, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Huiyu Zhou 0001, Licheng Jiao
ACM Multimedia6
2021 Multi-scale Edge-Based U-Shape Network for Salient Object Detection
Yetong Bian, Ningzhong Liu, Huiyu Zhou 0001
PRICAI (2)4
2021 Robust Ensembling Network for Unsupervised Domain Adaptation
Ningzhong Liu, Huiyu Zhou 0001
PRICAI (2)4
2021 MPI: Multi-receptive and parallel integration for salient object detection
abstract
Abstract The semantic representation of deep features is essential for image context understanding, and effective fusion of features with different semantic representations can significantly improve the model's performance on salient object detection. This paper proposes a novel method called multi‐receptive and parallel integration, for salient object detection. Firstly, a multi‐receptive enhancement module is designed to effectively expand the receptive fields of features from different layers and generate features with different receptive fields. Multi‐receptive enhancement module can enhance the semantic representation and improve the model's perception of the image context, which enables the model to locate the salient object accurately. Secondly, in order to reduce the reuse of redundant information in the complex top‐down fusion method and weaken the differences between semantic features, a relatively simple but effective parallel fusion strategy is proposed. It allows multi‐scale features to better interact with each other, thus improving the overall performance of the model. Experimental results on multiple datasets demonstrate that the proposed method outperforms state‐of‐the‐art methods under different evaluation metrics.
Jun Cen, Ningzhong Liu, Dong Liang 0008, Huiyu Zhou 0001
IET Image Process.5
2021 A novel few-shot learning method for synthetic aperture radar image recognition
Fei Gao 0005, Qingxu Xiong, Jinping Sun, Amir Hussain 0001, Huiyu Zhou 0001
Neurocomputing6
2021 Advances in domain adaptation for computer vision
Pourya Shamsolmoali, Salvador García 0001, Huiyu Zhou 0001, M. Emre Celebi 0001
Image Vis. Comput.3
2021 SRPN: similarity-based region proposal networks for nuclei and cells detection in histology images
Yibao Sun, Xingru Huang, Huiyu Zhou 0001, Qianni Zhang
Medical Image Anal.3
2021 An Efficient Anonymous Communication Scheme to Protect the Privacy of the Source Node Location in the Internet of Things
abstract
Advances in machine learning (ML) in recent years have enabled a dizzying array of applications such as data analytics, autonomous systems, and security diagnostics. As an important part of the Internet of Things (IoT), wireless sensor networks (WSNs) have been widely used in military, transportation, medical, and household fields. However, in the applications of wireless sensor networks, the adversary can infer the location of a source node and an event by backtracking attacks and traffic analysis. The location privacy leakage of a source node has become one of the most urgent problems to be solved in wireless sensor networks. To solve the problem of source location privacy leakage, in this paper, we first propose a proxy source node selection mechanism by constructing the candidate region. Secondly, based on the residual energy of the node, we propose a shortest routing algorithm to achieve better forwarding efficiency. Finally, by combining the proposed proxy source node selection mechanism with the proposed shortest routing algorithm based on the residual energy, we further propose a new, anonymous communication scheme. Meanwhile, the performance analysis indicates that the anonymous communication scheme can effectively protect the location privacy of the source nodes and reduce the network overhead.
Fengyin Li, Pei Ren, Guoyu Yang, Yuhong Sun, Siyuan Li 0022, Huiyu Zhou 0001
Secur. Commun. Networks8
2021 Perceptual Underwater Image Enhancement With Deep Learning and Physical Priors
abstract
Underwater image enhancement, as a pre-processing step to support the following object detection task, has drawn considerable attention in the field of underwater navigation and ocean exploration. However, most of the existing underwater image enhancement strategies tend to consider enhancement and detection as two fully independent modules with no interaction, and the practice of separate optimisation does not always help the following object detection task. In this article, we propose two perceptual enhancement models, each of which uses a deep enhancement model with a detection perceptor. The detection perceptor provides feedback information in the form of gradients to guide the enhancement model to generate patch level visually pleasing or detection favourable images. In addition, due to the lack of training data, a hybrid underwater image synthesis model, which fuses physical priors and data-driven cues, is proposed to synthesise training data and generalise our enhancement model for real-world underwater images. Experimental results show the superiority of our proposed method over several state-of-the-art methods on both real-world and synthetic underwater datasets.
Long Chen 0019, Zheheng Jiang, Aite Zhao, Qianni Zhang, Junyu Dong, Huiyu Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.8
2021 Road Segmentation for Remote Sensing Images Using Adversarial Spatial Pyramid Networks
abstract
Road extraction in remote sensing images is of great importance for a wide range of applications. Because of the complex background, and high density, most of the existing methods fail to accurately extract a road network that appears correct and complete. Moreover, they suffer from either insufficient training data or high costs of manual annotation. To address these problems, we introduce a new model to apply structured domain adaption for synthetic image generation and road segmentation. We incorporate a feature pyramid (FP) network into generative adversarial networks to minimize the difference between the source and target domains. A generator is learned to produce quality synthetic images, and the discriminator attempts to distinguish them. We also propose a FP network that improves the performance of the proposed model by extracting effective features from all the layers of the network for describing different scales' objects. Indeed, a novel scale-wise architecture is introduced to learn from the multilevel feature maps and improve the semantics of the features. For optimization, the model is trained by a joint reconstruction loss function, which minimizes the difference between the fake images and the real ones. A wide range of experiments on three data sets prove the superior performance of the proposed approach in terms of accuracy and efficiency. In particular, our model achieves state-of-the-art 78.86 IOU on the Massachusetts data set with 14.89M parameters and 86.78B FLOPs, with 4× fewer FLOPs but higher accuracy (+3.47% IOU) than the top performer among state-of-the-art approaches used in the evaluation.
Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Ruili Wang 0001, Jie Yang 0002
IEEE Trans. Geosci. Remote. Sens.3
2021 Enhanced Feature Pyramid Network With Deep Semantic Embedding for Remote Sensing Scene Classification
abstract
Recent progress on remote sensing (RS) scene classification is substantial, benefiting mostly from the explosive development of convolutional neural networks (CNNs). However, different from the natural images in which the objects occupy most of the space, objects in RS images are usually small and separated. Therefore, there is still a large room for improvement of the vanilla CNNs that extract global image-level features for RS scene classification, ignoring local object-level features. In this article, we propose a novel RS scene classification method via enhanced feature pyramid network (EFPN) with deep semantic embedding (DSE). Our proposed framework extracts multiscale multilevel features using an EFPN. Then, to leverage the complementary advantages of the multilevel and multiscale features, we design a DSE module to generate discriminative features. Third, a feature fusion module, called two-branch deep feature fusion (TDFF), is introduced to aggregate the features at different levels in an effective way. Our method produces state-of-the-art results on two widely used RS scene classification benchmarks, with better effectiveness and accuracy than the existing algorithms. Beyond that, we conduct an exhaustive analysis on the role of each module in the proposed architecture, and the experimental results further verify the merits of the proposed method.
Xin Wang 0068, Chen Ning, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Dynamic Multi-Key FHE in Asymmetric Key Setting From LWE
abstract
Multi-key Fully homomorphic encryption (MFHE) schemes allow computation on the encrypted data under different keys. However, traditional multi-key FHE schemes based on Learning with errors (LWE) have the undesirable property that is the number of keys has to be fixed in advance. A dynamic multi-key FHE scheme is the most versatile variant which the information about the participants is not required before key generation. To support further homomorphic computation on extended ciphertexts and ciphertexts encrypted under additional keys, Peikert and Shiehian (TCC ’16) proposed a leveled dynamic multi-key FHE scheme. Nevertheless, it introduces the circular-security assumption for the LWE parameters to ensure its security, which provides weaker security to the scheme. The problem of how to construct a LWE-based dynamic multi-key FHE scheme is still open. To address the above problem, in this work, we present a dynamic multi-key FHE scheme based on the LWE assumption in public key setting. The ciphertext can be extended and performed homomorphic evaluation with the ciphertexts encrypted under additional keys. Compared with current constructions, our proposed method requires fewer “local” memory and the process of ciphertext extension is distributed. Our proposed method provides a new way to extend the ciphertext such that the ciphertext homomorphism computation is more efficient. Our scheme is proven to be secure under standard LWE assumptions without using the circular-security assumption.
Yuling Chen 0002, Sen Dong, Tao Li 0043, Huiyu Zhou 0001
IEEE Trans. Inf. Forensics Secur.5
2021 Multi-View Mouse Social Behaviour Recognition With Deep Graphic Model
abstract
Home-cage social behaviour analysis of mice is an invaluable tool to assess therapeutic efficacy of neurodegenerative diseases. Despite tremendous efforts made within the research community, single-camera video recordings are mainly used for such analysis. Because of the potential to create rich descriptions for mouse social behaviors, the use of multi-view video recordings for rodent observations is increasingly receiving much attention. However, identifying social behaviours from various views is still challenging due to the lack of correspondence across data sources. To address this problem, we here propose a novel multi-view latent-attention and dynamic discriminative model that jointly learns view-specific and view-shared sub-structures, where the former captures unique dynamics of each view whilst the latter encodes the interaction between the views. Furthermore, a novel multi-view latent-attention variational autoencoder model is introduced in learning the acquired features, enabling us to learn discriminative features in each view. Experimental results on the standard CRMI13 and our multi-view Parkinson's Disease Mouse Behaviour (PDMB) datasets demonstrate that our proposed model outperforms the other state of the arts technologies, has lower computational cost than the other graphical models and effectively deals with the imbalanced data problem.
Zheheng Jiang, Feixiang Zhou, Aite Zhao, Xin Li 0052, Ling Li 0010, Dacheng Tao, Xuelong Li 0001, Huiyu Zhou 0001
IEEE Trans. Image Process.8
2021 Component-Based Feature Saliency for Clustering
abstract
Simultaneous feature selection and clustering is a major challenge in unsupervised learning. In particular, there has been significant research into saliency measures for features that result in good clustering. However, as datasets become larger and more complex, there is a need to adopt a finer-grained approach to saliency by measuring it in relation to a part of a model. Another issue is learning the feature saliency and advanced model parameters. We address the first by presenting a novel Gaussian mixture model, which explicitly models the dependency of individual mixture components on each feature giving a new component-based feature saliency measure. For the second, we use Markov Chain Monte Carlo sampling to estimate the model and hidden variables. Using a synthetic dataset, we demonstrate the superiority of our approach, in terms of clustering accuracy and model parameter estimation, over an approach using a model-based feature saliency with expectation maximisation. We performed an evaluation of our approach with six synthetic trajectory datasets obtaining an average clustering accuracy of 97 percent. To demonstrate the generality of our approach, we applied it to a network traffic flow dataset obtaining an accuracy of 93 percent for intrusion detection. Finally, we performed a comparison with state-of-the-art clustering techniques using three real-world trajectory datasets of vehicle traffic. Our approach achieved an average clustering accuracy of 96 percent compared to 77-95 percent for the other techniques. In conclusion, for the datasets considered, component based feature saliency measures gave improved clustering over those based on whole models.
Hailin Li, Paul Miller 0003, Jianjiang Zhou, Ling Li 0010, Danny Crookes, Yonggang Lu, Xuelong Li 0001, Huiyu Zhou 0001
IEEE Trans. Knowl. Data Eng.9
2021 CANet: Context Aware Network for Brain Glioma Segmentation
abstract
Automated segmentation of brain glioma plays an active role in diagnosis decision, progression monitoring and surgery planning. Based on deep neural networks, previous studies have shown promising technologies for brain glioma segmentation. However, these approaches lack powerful strategies to incorporate contextual information of tumor cells and their surrounding, which has been proven as a fundamental cue to deal with local ambiguity. In this work, we propose a novel approach named Context-Aware Network (CANet) for brain glioma segmentation. CANet captures high dimensional and discriminative features with contexts from both the convolutional space and feature interaction graphs. We further propose context guided attentive conditional random fields which can selectively aggregate features. We evaluate our method using publicly accessible brain glioma segmentation datasets BRATS2017, BRATS2018 and BRATS2019. The experimental results show that the proposed algorithm has better or competitive performance against several State-of-The-Art approaches under different segmentation metrics on the training and validation sets.
Long Chen 0019, Feixiang Zhou, Zheheng Jiang, Qianni Zhang, Yinhai Wang, Caifeng Shan, Ling Li 0010, Huiyu Zhou 0001
IEEE Trans. Medical Imaging10
2021 Building High-Fidelity Human Body Models From User-Generated Data
abstract
We propose a key point-based approach, refers to asKPhub-PC, to estimate high-fidelity human body models from low-quality point clouds acquired with an affordable 3D scanner and a variationKPhub-Ithat can achieve the same purpose based on low-resolution single images taken by smartphones. In KPhub-PC, a sparse set of key points is annotated to guide the deformation of a parametric 3D human body model SMPL and then a high-fidelity human body model that can explain the target point clouds is built. Besides building 3D human body models from point clouds, KPhub-I is designed to estimate accurate 3D human body models from single 2D images. The SMPL model is fitted to 2D joints and the boundary of the human body which are detected using CNN based methods automatically. Considering that people are in stable poses most of the time, a stable pose prior is defined from CMU motion capture dataset for further improving accuracy. Extensive experiments demonstrate that in both types of user-generated data, the proposed approaches can build believable and animatable human body models robustly. Our approach outperforms the state-of-the-arts in the accuracy of both human body shape and pose estimation.
Zongyi Xu, Yindi Zhu, Huiyu Zhou 0001, Qianni Zhang
IEEE Trans. Multim.5
2020 MARLINE: Multi-Source Mapping Transfer Learning for Non-Stationary Environments
abstract
Concept drift is a major problem in online learning due to its impact on the predictive performance of data stream mining systems. Recent studies have started exploring data streams from different sources as a strategy to tackle concept drift in a given target domain. These approaches make the assumption that at least one of the source models represents a concept similar to the target concept, which may not hold in many real-world scenarios. In this paper, we propose a novel approach called Multi-source mApping with tRansfer LearnIng for Non-stationary Environments (MARLINE). MARLINE can benefit from knowledge from multiple data sources in non-stationary environments even when source and target concepts do not match. This is achieved by projecting the target concept to the space of each source concept, enabling multiple source sub-classifiers to contribute towards the prediction of the target concept as part of an ensemble. Experiments on several synthetic and real-world datasets show that MARLINE was more accurate than several state-of-the-art data stream learning approaches.
Honghui Du, Leandro L. Minku, Huiyu Zhou 0001
ICDM3
2020 Underwater object detection using Invert Multi-Class Adaboost with deep learning
abstract
In recent years, deep learning based methods have achieved promising performance in standard object detection. However, these methods lack sufficient capabilities to handle underwater object detection due to these challenges: (1) Objects in real applications are usually small and their images are blurry, and (2) images in the underwater datasets and real applications accompany heterogeneous noise. To address these two problems, we first propose a novel neural network architecture, namely Sample-WeIghted hyPEr Network (SWIPENet), for small object detection. SWIPENet consists of high resolution and semantic-rich Hyper Feature Maps which can significantly improve small object detection accuracy. In addition, we propose a novel sample-weighted loss function which can model sample weights for SWIPENet, which uses a novel sample re-weighting algorithm, namely Invert Multi-Class Adaboost (IMA), to reduce the influence of noise on the proposed SWIPENet. Experiments on two underwater robot picking contest datasets URPC2017 and URPC2018 show that the proposed SWIPENet+IMA framework achieves better performance in detection accuracy against several state-of-the-art object detection approaches.
Long Chen 0019, Zheheng Jiang, Shengke Wang, Junyu Dong, Huiyu Zhou 0001
IJCNN7
2020 Insider Threat Risk Prediction based on Bayesian Network
Nebrase Elmrabit, Shuang-Hua Yang, Lili Yang 0001, Huiyu Zhou 0001
Comput. Secur.4
2020 Deep convolution network based emotion analysis towards mental health care
Zixiang Fei, Erfu Yang, Day-Uei Li, Stephen Butler, Winifred Ijomah, Huiyu Zhou 0001
Neurocomputing7
2020 A novel biologically-inspired target detection method based on saliency analysis for synthetic aperture radar (SAR) imagery
Fei Ma 0001, Fei Gao 0005, Jun Wang 0041, Amir Hussain 0001, Huiyu Zhou 0001
Neurocomputing5
2020 Texture synthesis quality assessment using perceptual texture similarity
Xinghui Dong, Huiyu Zhou 0001
Knowl. Based Syst.2
2020 Recent Advancement in Hybrid Big Data Processing
Shuai Liu 0002, Huiyu Zhou 0001, Xiaochun Cheng
Mob. Networks Appl.2
2020 Perceptual image quality using dual generative adversarial network
Masoumeh Zareapoor, Huiyu Zhou 0001, Jie Yang 0002
Neural Comput. Appl.2
2020 Modality-correlation-aware sparse representation for RGB-infrared object tracking
Xiangyuan Lan, Mang Ye, Shengping Zhang, Huiyu Zhou 0001, Pong C. Yuen
Pattern Recognit. Lett.4
2020 Advanced deep learning for image super-resolution
Pourya Shamsolmoali, Abdul Hamid Sadka, Huiyu Zhou 0001, Wankou Yang
Signal Process. Image Commun.3
2020 A Perception-Inspired Deep Learning Framework for Predicting Perceptual Texture Similarity
abstract
Similarity learning plays a fundamental role in the fields of multimedia retrieval and pattern recognition. Prediction of perceptual similarity is a challenging task as in most cases we lack human labeled ground-truth data and robust models to mimic human visual perception. Although in the literature, some studies have been dedicated to similarity learning, they mainly focus on the evaluation of whether or not two images are similar, rather than prediction of perceptual similarity which is consistent with human perception. Inspired by the human visual perception mechanism, we here propose a novel framework in order to predict perceptual similarity between two texture images. Our proposed framework is built on the top of Convolutional Neural Networks (CNNs). The proposed framework considers both powerful features and perceptual characteristics of contours extracted from the images. The similarity value is computed by aggregating resemblances between the corresponding convolutional layer activations of the two texture maps. Experimental results show that the predicted similarity values are consistent with the human-perceived similarity data.
Ying Gao 0005, Yanhai Gan, Lin Qi 0004, Huiyu Zhou 0001, Xinghui Dong, Junyu Dong
IEEE Trans. Circuits Syst. Video Technol.4
2020 Real-Time H∞ Control of Networked Inverted Pendulum Visual Servo Systems
abstract
Aiming at the challenges of networked visual servo control systems, which rarely consider network communication duration and image processing computational cost simultaneously, we here propose a novel platform for networked inverted pendulum visual servo control using H∞ analysis. Unlike most of the existing methods that usually ignore computational costs involved in measuring, actuating, and controlling, we design a novel event-triggered sampling mechanism that applies a new closed-loop strategy to dealing with networked inverted pendulum visual servo systems of multiple time-varying delays and computational errors. Using the Lyapunov stability theory, we prove that the proposed system can achieve stability whilst compromising image-induced computational and network-induced delays and system performance. In the meantime, we use H∞disturbance attenuation level γ for evaluating the computational errors, whereas the corresponding H∞controller is implemented. Finally, simulation analysis and experimental results demonstrate the proposed system performance in reducing computational errors whilst maintaining system efficiency and robustness.
Dajun Du, Changda Zhang, Yuehua Song, Huiyu Zhou 0001, Xue Li 0028, Minrui Fei, Wangpei Li
IEEE Trans. Cybern.4
2020 A Multipopulation-Based Multiobjective Evolutionary Algorithm
abstract
Multipopulation is an effective optimization component often embedded into evolutionary algorithms to solve optimization problems. In this paper, a new multipopulation-based multiobjective genetic algorithm (MOGA) is proposed, which uses a unique cross-subpopulation migration process inspired by biological processes to share information between subpopulations. Then, a Markov model of the proposed multipopulation MOGA is derived, the first of its kind, which provides an exact mathematical model for each possible population occurring simultaneously with multiple objectives. Simulation results of two multiobjective test problems with multiple subpopulations justify the derived Markov model, and show that the proposed multipopulation method can improve the optimization ability of the MOGA. Also, the proposed multipopulation method is applied to other multiobjective evolutionary algorithms (MOEAs) for evaluating its performance against the IEEE Congress on Evolutionary Computation multiobjective benchmarks. The experimental results show that a single-population MOEA can be extended to a multipopulation version, while obtaining better optimization performance.
Haiping Ma, Minrui Fei, Zheheng Jiang, Ling Li 0010, Huiyu Zhou 0001, Danny Crookes
IEEE Trans. Cybern.5
2020 Texture Classification Using Pair-Wise Difference Pooling-Based Bilinear Convolutional Neural Networks
abstract
Texture is normally represented by aggregating local features based on the assumption of spatial homogeneity. Effective texture features are always the research focus even though both hand-crafted and deep learning approaches have been extensively investigated. Motivated by the success of Bilinear Convolutional Neural Networks (BCNNs) in fine-grained image recognition, we propose to incorporate the BCNN with the Pair-wise Difference Pooling (i.e. BCNN-PDP) for texture classification. The BCNN-PDP is built on top of a set of feature maps extracted at a convolutional layer of the pre-trained CNN. Compared with the outer product used by the original BCNN feature set, the pair-wise difference not only captures the pair-wise relationship between two sets of features but also encodes the difference between each pair of features. Considering the importance of the gradient data to the representation of image structures, we further generalise the BCNN-PDP feature set to two sets of feature maps computed from the original image and its gradient magnitude map respectively, i.e. the Fused BCNN-PDP (F-BCNN-PDP) feature set. In addition, the BCNN-PDP can be applied to two different CNNs and is referred to as the Asymmetric BCNN-PDP (A-BCNN-PDP). The three PDP-based BCNN feature sets can also be extracted at multiple scales. Since the dimensionality of the BCNN feature vectors is very high, we propose a new yet simple Block-wise PCA (BPCA) method in order to derive more compact feature vectors. The proposed methods are tested on seven different datasets along with 21 baseline feature sets. The results show that the proposed feature sets are superior, or at least comparable, to their counterparts across different datasets.
Xinghui Dong, Huiyu Zhou 0001, Junyu Dong
IEEE Trans. Image Process.2
2020 Grayscale-Thermal Tracking via Inverse Sparse Representation-Based Collaborative Encoding
abstract
Grayscale-thermal tracking has attracted a great deal of attention due to its capability of fusing two different yet complementary target observations. Existing methods often consider extracting the discriminative target information and exploring the target correlation among different images as two separate issues, ignoring their interdependence. This may cause tracking drifts in challenging video pairs. This paper presents a collaborative encoding model called joint correlation and discriminant analysis based inver-sparse representation (JCDA-InvSR) to jointly encode the target candidates in the grayscale and thermal video sequences. In particular, we develop a multi-objective programming to integrate the feature selection and the multi-view correlation analysis into a unified optimization problem in JCDA-InvSR, which can simultaneously highlight the special characters of the grayscale and thermal targets through alternately optimizing two aspects: the target discrimination within a given image and the target correlation across different images. For robust grayscale-thermal tracking, we also incorporate the prior knowledge of target candidate codes into the SVM based target classifier to overcome the overfitting caused by limited training labels. Extensive experiments on GTOT and RGBT234 datasets illustrate the promising performance of our tracking framework.
Bin Kang, Dong Liang 0008, Wan Ding, Huiyu Zhou 0001, Wei-Ping Zhu 0001
IEEE Trans. Image Process.4
2020 Siamese Local and Global Networks for Robust Face Tracking
abstract
Convolutional neural networks (CNNs) have achieved great success in several face-related tasks, such as face detection, alignment and recognition. As a fundamental problem in computer vision, face tracking plays a crucial role in various applications, such as video surveillance, human emotion detection and human-computer interaction. However, few CNN-based approaches are proposed for face (bounding box) tracking. In this paper, we propose a face tracking method based on Siamese CNNs, which takes advantages of powerful representations of hierarchical CNN features learned from massive face images. The proposed method captures discriminative face information at both local and global levels. At the local level, representations for attribute patches (i.e:, eyes, nose and mouth) are learned to distinguish a face from another one, which are robust to pose changes and occlusions. At the global level, representations for each whole face are learned, which take into account the spatial relationships among local patches and facial characters, such as skin color and nevus. In addition, we build a new largescale challenging face tracking dataset to evaluate face tracking methods and to facilitate the research forward in this field. Extensive experiments on the collected dataset demonstrate the effectiveness of our method in comparison to several state-of-theart visual tracking methods.
Yuankai Qi, Shengping Zhang, Feng Jiang 0001, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Image Process.4
2020 Monocular Visual-IMU Odometry: A Comparative Evaluation of Detector-Descriptor-Based Methods
abstract
Monocular visual-inertial measurement unit (IMU) odometry has been widely used in various intelligent vehicles. As a popular technique, detector-descriptor-based visual-IMU odometry is effective and efficient due to the fact that local descriptors are robust against occlusions, background clutter, and abrupt content changes. However, to our knowledge, there is not a comprehensive and comparative evaluation study on the performance of different combinations of detectors and descriptors recently developed. In order to bridge this gap, we conduct such a comparative study in a unified framework. In particular, six typical routes with different lengths, shapes, and road scenes are selected from the well-known KITTI dataset. We first evaluate the performance of different combinations of salient point detectors and local descriptors using the six routes. Then, we tune the parameters of the best detector or descriptor obtained for each route, to further augment the results. This paper provides not only comprehensive benchmarks for assessing various algorithms but also instructive guidelines and insights for developing detectors and descriptors to handle different road scenes.
Xingshuai Dong, Xinghui Dong, Junyu Dong, Huiyu Zhou 0001
IEEE Trans. Intell. Transp. Syst.4
2020 AMIL: Adversarial Multi-instance Learning for Human Pose Estimation
abstract
Human pose estimation has an important impact on a wide range of applications, from human-computer interface to surveillance and content-based video retrieval. For human pose estimation, joint obstructions and overlapping upon human bodies result in departed pose estimation. To address these problems, by integrating priors of the structure of human bodies, we present a novel structure-aware network to discreetly consider such priors during the training of the network. Typically, learning such constraints is a challenging task. Instead, we propose generative adversarial networks as our learning model in which we design two residual Multiple-Instance Learning (MIL) models with identical architecture—one is used as the generator, and the other one is used as the discriminator. The discriminator task is to distinguish the actual poses from the fake ones. If the pose generator generates results that the discriminator is not able to distinguish from the real ones, then the model has successfully learned the priors. In the proposed model, the discriminator differentiates the ground-truth heatmaps from the generated ones, and later the adversarial loss back-propagates to the generator. Such procedure assists the generator to learn reasonable body configurations and is proved to be advantageous to improve the pose estimation accuracy. Meanwhile, we propose a novel function for MIL. It is an adjustable structure for both instance selection and modeling to appropriately pass the information between instances in a single bag. In the proposed residual MIL neural network, the pooling action adequately updates the instance contribution to its bag. The proposed adversarial residual multi-instance neural network that is based on pooling has been validated on two datasets for the human pose estimation task and successfully outperforms the other state-of-the-art models. The code will be made available on https://github.com/pshams55/AMIL.
Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Jie Yang 0002
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Introduction to the Special Issue on Multimodal Machine Learning for Human Behavior Analysis
abstract
No abstract available.
Shengping Zhang, Huiyu Zhou 0001, Dong Xu 0001, M. Emre Celebi 0001, Thierry Bouwmans
ACM Trans. Multim. Comput. Commun. Appl.2
2020 Wireless Communications and Mobile Computing Blockchain-Based Trust Management in Distributed Internet of Things
abstract
The development of Internet of Things (IoT) and Mobile Edge Computing (MEC) has led to close cooperation between electronic devices. It requires strong reliability and trustworthiness of the devices involved in the communication. However, current trust mechanisms have the following issues: (1) heavily relying on a trusted third party, which may incur severe security issues if it is corrupted, and (2) malicious evaluations on the involved devices which may bias the trustrank of the devices. By introducing the concepts of risk management and blockchain into the trust mechanism, we here propose a blockchain-based trust mechanism for distributed IoT devices in this paper. In the proposed trust mechanism, trustrank is quantified by normative trust and risk measures, and a new storage structure is designed for the domain administration manager to identify and delete the malicious evaluations of the devices. Evidence shows that the proposed trust mechanism can ensure data sharing and integrity, in addition to its resistance against malicious attacks to the IoT devices.
Fengyin Li, Dongfeng Wang 0003, Xiaomei Yu, Jiguo Yu, Huiyu Zhou 0001
Wirel. Commun. Mob. Comput.7
2019 Score-specific Non-maximum Suppression and Coexistence Prior for Multi-scale Face Detection
abstract
Face detection is an ultimate component to support various visual facial related tasks. However, detecting faces with extremely low resolution or high occlusion is still an open problem. In this paper, we propose a two-step general approach to refine the performance of modern face detectors according to human's high-level context-aware ability. First, we propose Score-specific Non-Maximum Suppression (SNMS) to preserve overlapped faces. Second, we consider the coexistence prior among faces in the scene, which could raise the sensitivity of face detection in the crowd. When integrating our approach to the existing face detectors, most of them have better results on a challenging benchmark (WIDER FACE) and a newly proposed dataset (Faces in Crowd, FIC) made by us. Codes are available on https://github.com/AIoTP/SNMSandCoexistence.
Tianpeng Wu, Dong Liang 0008, Jiaxing Pan, Bin Kang, Shun'ichi Kaneko, Huiyu Zhou 0001
ICASSP7
2019 Multi-Source Transfer Learning for Non-Stationary Environments
abstract
In data stream mining, predictive models typically suffer drops in predictive performance due to concept drift. As enough data representing the new concept must be collected for the new concept to be well learnt, the predictive performance of existing models usually takes some time to recover from concept drift. To speed up recovery from concept drift and improve predictive performance in data stream mining, this work proposes a novel approach called Multi-sourcE onLine TrAnsfer learning for Non-statIonary Environments (Melanie). Melanie is the first approach able to transfer knowledge between multiple data streaming sources in non-stationary environments. It creates several sub-classifiers to learn different aspects from different source and target concepts over time. The sub-classifiers that match the current target concept well are identified, and used to compose an ensemble for predicting examples from the target concept. We evaluate Melanie on several synthetic data streams containing different types of concept drift and on real world data streams. The results indicate that Melanie can deal with a variety drifts and improve predictive performance over existing data stream learning algorithms by making use of multiple sources.
Honghui Du, Leandro L. Minku, Huiyu Zhou 0001
IJCNN3
2019 Cross-layer access control in publish/subscribe middleware over software-defined networks
Yang Zhang 0015, Huiyu Zhou 0001, Junliang Chen 0001
Comput. Commun.2
2019 Editorial: Neural learning in life system and energy system
Chen Peng 0001, Dong Yue 0001, Dajun Du, Huiyu Zhou 0001, Aolei Yang
Neurocomputing4
2019 A Robust Parallel Object Tracking Method for Illumination Variations
Shuai Liu 0002, Gaocheng Liu, Huiyu Zhou 0001
Mob. Networks Appl.3
2019 Cascaded one-vs-rest detection network for fine-grained recognition without part annotations
Long Chen 0019, Shengke Wang, Kin-Man Lam 0001, Huiyu Zhou 0001, Muwei Jian, Junyu Dong
Multim. Tools Appl.4
2019 A benchmark image dataset for industrial tools
Cai Luo, Leijian Yu, Erfu Yang, Huiyu Zhou 0001, Peng Ren 0001
Pattern Recognit. Lett.4
2019 Hierarchical residual learning for image denoising
Wuzhen Shi, Feng Jiang 0001, Shengping Zhang, Rui Wang 0093, Debin Zhao, Huiyu Zhou 0001
Signal Process. Image Commun.6
2019 Fuzzy Optimal Energy Management for Fuel Cell and Supercapacitor Systems Using Neural Network Based Driving Pattern Recognition
abstract
A novel adaptive energy management strategy is proposed for real-time power split between fuel cells (FCs) and supercapacitors (SCs) in a hybrid electric vehicle in view of the fact that driving patterns greatly affect fuel economy. The driving pattern recognition (DPR) is achieved based on the features extracted from the historical velocity window with a multilayer perceptron neural network. After the DPR has been obtained, an adaptive fuzzy energy management controller is utilized for power split according to the required power for vehicle running. In order to prolong the FC lifetime while decreasing the hydrogen consumption, a genetic algorithm is applied to optimize critical factors such as adaptive gains and fuzzy membership function parameters for several standard driving cycles. In the proposed method, the future driving cycles are not required and the current driving pattern can be successfully recognized, demonstrating that less current fluctuations and fuel consumption can be achieved under various driving conditions. Compared with conventional energy management systems, the proposed framework can ensure the state of charge of SCs within the desired limit.
Ridong Zhang, Jili Tao, Huiyu Zhou 0001
IEEE Trans. Fuzzy Syst.3
2019 Hyperspectral Anomaly Detection via Background and Potential Anomaly Dictionaries Construction
abstract
In this paper, we propose a new anomaly detection method for hyperspectral images based on two well-designed dictionaries: background dictionary and potential anomaly dictionary. In order to effectively detect an anomaly and eliminate the influence of noise, the original image is decomposed into three components: background, anomalies, and noise. In this way, the anomaly detection task is regarded as a problem of matrix decomposition. Considering the homogeneity of background and the sparsity of anomalies, the low-rank and sparse constraints are imposed in our model. Then, the background and potential anomaly dictionaries are constructed using the background and anomaly priors. For the background dictionary, a joint sparse representation (JSR)-based dictionary selection strategy is proposed, assuming that the frequently used atoms in the overcomplete dictionary tend to be the background. In order to make full use of the prior information of anomalies hidden in the scene, the potential anomaly dictionary is constructed. We define a criterion, i.e., the anomalous level of a pixel, by using the residual calculated in the JSR model within its local region. Then, it is combined with a weighted term to alleviate the influence of noise and background. Experiments show that our proposed anomaly detection method based on potential anomaly and background dictionaries construction can achieve superior results compared with other state-of-the-art methods.
Ning Huyan, Xiangrong Zhang, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2019 Context-Aware Mouse Behavior Recognition Using Hidden Markov Models
abstract
Automated recognition of mouse behaviors is crucial in studying psychiatric and neurologic diseases. To achieve this objective, it is very important to analyze the temporal dynamics of mouse behaviors. In particular, the change between mouse neighboring actions is swift in a short period. In this paper, we develop and implement a novel hidden Markov model (HMM) algorithm to describe the temporal characteristics of mouse behaviors. In particular, we here propose a hybrid deep learning architecture, where the first unsupervised layer relies on an advanced spatial-temporal segment Fisher vector encoding both visual and contextual features. Subsequent supervised layers based on our segment aggregate network are trained to estimate the state-dependent observation probabilities of the HMM. The proposed architecture shows the ability to discriminate between visually similar behaviors and results in high recognition rates with the strength of processing imbalanced mouse behavior datasets. Finally, we evaluate our approach using JHuang's and our own datasets, and the results show that our method outperforms other state-of-the-art approaches.
Zheheng Jiang, Danny Crookes, Brian Desmond Green, Haiping Ma, Ling Li 0010, Shengping Zhang, Dacheng Tao, Huiyu Zhou 0001
IEEE Trans. Image Process.9
2018 A Novel Method For Unsupervised Scanner-Invariance Using A Dual-Channel Auto-Encoder Model
Andrew D. Moyes, Kun Zhang 0010, Ming Ji, Danny Crookes, Huiyu Zhou 0001
BMVC6
2018 Merging Neurons for Structure Compression of Deep Networks
abstract
Deep neural networks are increasingly used in many fields, such as pattern recognition, computer vision, and natural language processing. However, how to apply deep neural networks in mobile settings has become an urgent issue, as mobile devices are getting more and more popularity. This is mainly due to the fact that mobile devices usually have very limited computation and storage resources, which prevents from running a large-scale deep network. This paper proposes a novel method for structure compression of deep neural networks. The main idea is to merge the neurons and connections of the original network using clustering methods. To the end, the new network after compression possesses much less parameters, which leads to reduced requirements for computation and storage resources. Experiments on benchmark data sets demonstrate that the proposed method can greatly improve the efficiency of deep neural networks, while retain their learning capability.
Guoqiang Zhong 0001, Huiyu Zhou 0001
ICPR3
2018 Transferring deep knowledge for object recognition in Low-quality underwater videos
Xin Sun 0003, Junyu Shi, Lipeng Liu, Junyu Dong, Claudia Plant, Huiyu Zhou 0001
Neurocomputing7
2018 Asymmetric filtering-based dense convolutional neural network for person re-identification combined with Joint Bayesian and re-ranking
Shengke Wang, Long Chen 0019, Huiyu Zhou 0001, Junyu Dong
J. Vis. Commun. Image Represent.4
2018 Incentive-driven attacker for corrupting two-party protocols
abstract
Adversaries in two-party computation may sabotage a protocol, leading to possible collapse of the information security management. In practice, attackers often breach security protocols with specific incentives. For example, attackers manage to reap additional rewards by sabotaging computing tasks between two clouds. Unfortunately, most of the existing research works neglect this aspect when discussing the security of protocols. Furthermore, the construction of corrupting two parties is also missing in two-party computation. In this paper, we propose an incentive-driven attacking model where the attacker leverages corruption costs, benefits and possible consequences. We here formalize the utilities used for two-party protocols and the attacker(s), taking into account both corruption costs and attack benefits. Our proposed model can be considered as the extension of the seminal work presented by Groce and Katz (Annual international conference on the theory and applications of cryptographic techniques, Springer, Berlin, pp 81–98, 2012 ), while making significant contribution in addressing the corruption of two parties in two-party protocols. To the best of our knowledge, this is the first time to model the corruption of both parties in two-party protocols.
Roberto Metere, Huiyu Zhou 0001, Guanghai Cui, Tao Li 0043
Soft Comput.3
2018 BoMW: Bag of Manifold Words for One-Shot Learning Gesture Recognition From Kinect
abstract
In this paper, we study one-shot learning gesture recognition on RGB-D data recorded from Microsoft's Kinect. To this end, we propose a novel bag of manifold words (BoMW)-based feature representation on symmetric positive definite (SPD) manifolds. In particular, we use covariance matrices to extract local features from RGB-D data due to its compact representation ability as well as the convenience of fusing both RGB and depth information. Since covariance matrices are SPD matrices and the space spanned by them is the SPD manifold, traditional learning methods in the Euclidean space, such as sparse coding, cannot be directly applied to them. To overcome this problem, we propose a unified framework to transfer the sparse coding on SPD manifolds to the one on the Euclidean space, which enables any existing learning method to be used. After building BoMW representation on a video from each gesture class, a nearest neighbor classifier is adopted to perform the one-shot learning gesture recognition. Experimental results on the ChaLearn gesture data set demonstrate the outstanding performance of the proposed one-shot learning gesture recognition method compared against the state-of-the-art methods. The effectiveness of the proposed feature extraction method is also validated on a new RGB-D action recognition data set.
Lei Zhang 0036, Shengping Zhang, Feng Jiang 0001, Yuankai Qi, Jun Zhang 0017, Yuliang Guo, Huiyu Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.7
2018 Tensor-Based Low-Rank Graph With Multimanifold Regularization for Dimensionality Reduction of Hyperspectral Images
abstract
Dimensionality reduction is an essential task in hyperspectral image processing. How to preserve the original intrinsic structure information and enhance the discriminant ability is still a challenge in this area. Recently, with the advantage of preserving global intrinsic structure information, low-rank representation has been applied to dimensionality reduction and achieved promising performance. By exploiting the submanifold information of the original data set, multimanifold learning is effective in enhancing the discriminant ability of the processed data set. In addition, due to the ability of preserving the spatial neighborhood structure information, the tensor analysis has become a popular technique for hyperspectral image processing. Motivated by the above-mentioned analysis, a novel tensor-based low-rank graph with multimanifold regularization (T-LGMR) for dimensionality reduction of hyperspectral images is proposed in this paper. In the T-LGMR, a low-rank constraint is employed to preserve the global data structure while multimanifold information is utilized to enhance the discriminant ability, and tensor representation is used to preserve the spatial neighborhood information. Finally, dimensionality reduction is achieved in the graph embedding framework. Experimental results on three real hyperspectral data sets demonstrate the superiority of the proposed method over several state-of-the-art approaches.
Jinliang An, Xiangrong Zhang, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2018 Multifeature Hyperspectral Image Classification With Local and Nonlocal Spatial Information via Markov Random Field in Semantic Space
abstract
Hyperspectral images (HSIs) provide invaluable information in both spectral and spatial domains for image classification tasks. In this paper, we use semantic representation as a middle-level feature to describe image pixels' characteristics. Deriving effective semantic representation is critical for achieving good classification performance. Since different image descriptors depict characteristics from different perspectives, combining multiple features in the same semantic space makes semantic representation more meaningful. First, a probabilistic support vector machine is used to generate semantic representation-based multifeatures. In order to derive better semantic representation, we introduce a new adaptive spatial regularizer that well exploits the local spatial information, while a nonlocal regularizer is also used to search for global patch-pair similarities in the whole image. We combine multiple features with local and nonlocal spatial constraints using an extended Markov random field model in the semantic space. Experimental results on three hyperspectral data sets show that the proposed method provides better performance than several state-of-the-art techniques in terms of region uniformity, overall accuracy, average accuracy, and Kappa statistics.
Xiangrong Zhang, Zeyu Gao 0001, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.4
2018 Hybrid Unmixing Based on Adaptive Region Segmentation for Hyperspectral Imagery
abstract
Unmixing is an important issue of hyperspectral images. Most unmixing methods adopt linear mixing models for simplicity. However, multiple scattering usually occurs between vegetation and soil in a bilinear scene. Thus, nonlinear mixing problems which are difficult to be solved should be taken into consideration under this circumstance. In practice, both linear and nonlinear spectral mixtures exist in hyperspectral scenes. Considering the characteristics of different regions in images, we propose a hybrid unmixing algorithm for hyperspectral images based on region adaptive segmentation. Our method uses a standard K-means clustering algorithm to obtain different regions, including homogeneous regions and detailed regions. The model of the homogeneous regions is assumed to be linear, which will be pursued using the method of sparse-constrained nonnegative matrix factorization (NMF), and the mixing in the detailed regions is assumed to be based on a nonlinear model. We also propose a new nonlinear unmixing method, called graph-regularized semi-NMF, which considers the manifold structure of hyperspectral data as the unmixing method to deal with the detailed regions. Finally, by combining the two regions, we obtain the abundance of the whole hyperspectral image. The proposed method can not only achieve more precise abundance but also be good at keeping the edge information of the bilinear abundance. The experimental results on both synthetic and real data also show that the proposed method is effective for improving the unmixing accuracy of hyperspectral remote-sensing images.
Xiangrong Zhang, Chen Li 0011, Cai Cheng, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.6
2018 Point-to-Set Distance Metric Learning on Deep Representations for Visual Tracking
abstract
For autonomous driving application, a car shall be able to track objects in the scene in order to estimate where and how they will move such that the tracker embedded in the car can efficiently alert the car for effective collision-avoidance. Traditional discriminative object tracking methods usually train a binary classifier via a support vector machine (SVM) scheme to distinguish the target from its background. Despite demonstrated success, the performance of the SVM-based trackers is limited because the classification is carried out only depending on support vectors (SVs) but the target's dynamic appearance may look similar to the training samples that have not been selected as SVs, especially when the training samples are not linearly classifiable. In such cases, the tracker may drift to the background and fail to track the target eventually. To address this problem, in this paper, we propose to integrate the point-to-set/image-to-imageSet distance metric learning (DML) into visual tracking tasks and take full advantage of all the training samples when determining the best target candidate. The point-to-set DML is conducted on convolutional neural network features of the training data extracted from the starting frames. When a new frame comes, target candidates are first projected to the common subspace using the learned mapping functions, and then the candidate having the minimal distance to the target template sets is selected as the tracking result. Extensive experimental results show that even without model update the proposed method is able to achieve favorable performance on challenging image sequences compared with several state-of-the-art trackers.
Shengping Zhang, Yuankai Qi, Feng Jiang 0001, Xiangyuan Lan, Pong C. Yuen, Huiyu Zhou 0001
IEEE Trans. Intell. Transp. Syst.6
2017 Behavior Recognition in Mouse Videos using Contextual Features Encoded by Spatial-temporal Stacked Fisher Vectors
abstract
Manual measurement of mouse behavior is highly labor intensive and prone to error. This investigation aims to efficiently and accurately recognize individual mouse behaviors in action videos and continuous videos. In our system each mouse action video is expressed as the collection of a set of interest points. We extract both appearance and contextual features from the interest points collected from the training datasets, and then obtain two Gaussian Mixture Model (GMM) dictionaries for the visual and contextual features. The two GMM dictionaries are leveraged by our spatial-temporal stacked Fisher Vector (FV) to represent each mouse action video. A neural network is used to classify mouse action and finally applied to annotate continuous video. The novelty of our proposed approach is: (i) our method exploits contextual features from spatiotemporal interest points, leading to enhanced performance, (ii) we encode contextual features and then fuse them with appearance features, and (iii) location information of a mouse is extracted from spatio-temporal interest points to support mouse behavior recognition. We evaluate our method against the database of Jhuang et al. (Jhuang et al., 2010) and the results show that our method outperforms several state-of-the-art approaches.
Zheheng Jiang, Danny Crookes, Brian Desmond Green, Shengping Zhang, Huiyu Zhou 0001
ICPRAM5
2017 Hierarchical Task Network planning with common-sense reasoning for multiple-people behaviour analysis
María J. Santofimia, Jesús Martínez del Rincón, Huiyu Zhou 0001, Paul Miller 0003, David Villa, Juan Carlos López 0001
Expert Syst. Appl.4
2017 A novel target detection method for SAR images based on shadow proposal and saliency analysis
Fei Gao 0005, Jialing You, Jun Wang 0041, Jinping Sun, Erfu Yang, Huiyu Zhou 0001
Neurocomputing6
2017 Recursive Autoencoders-Based Unsupervised Feature Learning for Hyperspectral Image Classification
abstract
For hyperspectral image (HSI) classification, it is very important to learn effective features for the discrimination purpose. Meanwhile, the ability to combine spectral and spatial information together in a deep level is also important for feature learning. In this letter, we propose an unsupervised feature learning method for HSI classification, which is based on recursive autoencoders (RAE) network. RAE utilizes the spatial and spectral information and produces high-level features from the original data. It learns features from the neighborhood of the investigated pixel to represent the whole local homogeneous area of the image. In addition, to obtain more accurate representation of the investigated pixel, a weighting scheme is adopted based on the neighboring pixels, where the weights are determined by the spectral similarity between the neighboring pixels and the investigated pixel. The effectiveness of our method is evaluated by the experiments on two hyperspectral data sets, and the results show that our proposed method has a better performance.
Xiangrong Zhang, Chen Li 0011, Ning Huyan, Licheng Jiao, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.6
2017 Modeling Information Diffusion over Social Networks for Temporal Dynamic Prediction
abstract
Modeling the process of information diffusion is a challenging problem. Although numerous attempts have been made in order to solve this problem, very few studies are actually able to simulate and predict temporal dynamics of the diffusion process. In this paper, we propose a novel information diffusion model, namely GT model, which treats the nodes of a network as intelligent and rational agents and then calculates their corresponding payoffs, given different choices to make strategic decisions. By introducing time-related payoffs based on the diffusion data, the proposed GT model can be used to predict whether or not the user's behaviors will occur in a specific time interval. The user's payoff can be divided into two parts: social payoff from the user's social contacts and preference payoff from the user's idiosyncratic preference. We here exploit the global influence of the user and the social influence between any two users to accurately calculate the social payoff. In addition, we develop a new method of presenting social influence that can fully capture the temporal dynamics of social influence. Experimental results from two different datasets, Sina Weibo and Flickr demonstrate the rationality and effectiveness of the proposed prediction method with different evaluation metrics.
Shengping Zhang, Xin Sun 0003, Huiyu Zhou 0001, Sheng Li 0003, Xuelong Li 0001
IEEE Trans. Knowl. Data Eng.4
2017 A Biologically Inspired Appearance Model for Robust Visual Tracking
abstract
In this paper, we propose a biologically inspired appearance model for robust visual tracking. Motivated in part by the success of the hierarchical organization of the primary visual cortex (area V1), we establish an architecture consisting of five layers: whitening, rectification, normalization, coding, and pooling. The first three layers stem from the models developed for object recognition. In this paper, our attention focuses on the coding and pooling layers. In particular, we use a discriminative sparse coding method in the coding layer along with spatial pyramid representation in the pooling layer, which makes it easier to distinguish the target to be tracked from its background in the presence of appearance variations. An extensive experimental study shows that the proposed method has higher tracking accuracy than several state-of-the-art trackers.
Shengping Zhang, Xiangyuan Lan, Hongxun Yao, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2016 Multi-scale Colorectal Tumour Segmentation Using a Novel Coarse to Fine Strategy
Kun Zhang 0010, Danny Crookes, Jim Diamond, Minrui Fei, Peijian Zhang, Huiyu Zhou 0001
BMVC7
2016 k-fold Subsampling based Sequential Backward Feature Elimination
abstract
We present a new wrapper feature selection algorithm for human detection. This algorithm is a hybrid featureselection approach combining the benefits of filter and wrapper methods. It allows the selection of an optimalfeature vector that well represents the shapes of the subjects in the images. In detail, the proposed featureselection algorithm adopts the k-fold subsampling and sequential backward elimination approach, while thestandard linear support vector machine (SVM) is used as the classifier for human detection. We apply theproposed algorithm to the publicly accessible INRIA and ETH pedestrian full image datasets with the PASCALVOC evaluation criteria. Compared to other state of the arts algorithms, our feature selection based approachcan improve the detection speed of the SVM classifier by over 50% with up to 2% better detection accuracy.Our algorithm also outperforms the equivalent systems introduced in the deformable part model approach witharound 9% improvement in the detection accuracy
Jeonghwan Park 0002, Kang Li 0002, Huiyu Zhou 0001
ICPRAM3
2016 Evidential event inference in transport video surveillance
Wenjun Ma, Sriram Varadarajan, Paul Miller 0003, Weiru Liu, María J. Santofimia, Jesús Martínez del Rincón, Huiyu Zhou 0001
Comput. Vis. Image Underst.9
2016 Editorial: Special issue: Life system modeling and simulation
Huiyu Zhou 0001, Dongbing Gu, Ling Wang 0009, Minrui Fei
Neurocomputing1
2016 An evidential fusion approach for gender profiling
Jianbing Ma, Weiru Liu, Paul Miller 0003, Huiyu Zhou 0001
Inf. Sci.4
2015 RFID network deployment approaches for indoor localisation
abstract
Three RFID reader based network deployment algorithms (grid-covering, diagonal and mixed) were evaluated in this paper. Experimental results show that the grid-covering method can be used to minimize hardware costs, but it leads to many indeterminate positions. The diagonal method can be used to solve the indeterminate problem, however increases the number of readers, especially in a large tracking field. The mixed algorithm can be used to avoid the indeterminate issue and also has the minimum reader number when deployed in a large space. However, it is not suitable for a small tracking field. An optimal deployment algorithm is selected from these three algorithms according to the environmental conditions and the localization requirement. In addition, an optimal RFID reader network deployment combined with a subarea-mapping algorithm can be used to minimize the hardware costs while improving the fine-grained indoor localization accuracy.
Shumei Zhang, Paul J. McCullagh, Huiyu Zhou 0001, Zhe Wen, Zhengcheng Xu
BSN3
2015 Fast convergence of regularised Region-based Mixture of Gaussians for dynamic background modelling
Sriram Varadarajan, Hongbin Wang 0005, Paul Miller 0003, Huiyu Zhou 0001
Comput. Vis. Image Underst.4
2015 Region-based Mixture of Gaussians modelling for foreground detection in dynamic scenes
Sriram Varadarajan, Paul Miller 0003, Huiyu Zhou 0001
Pattern Recognit.3
2015 Machine learning and signal processing for human pose recovery and behavior analysis
Jun Yu 0002, Huiyu Zhou 0001, Xinbo Gao 0001
Signal Process.2
2015 Adaptive NormalHedge for robust visual tracking
Shengping Zhang, Huiyu Zhou 0001, Hongxun Yao, Yanhao Zhang 0001, Kuanquan Wang, Jun Zhang 0017
Signal Process.2
2015 Robust Visual Tracking Using Structurally Random Projection and Weighted Least Squares
abstract
Sparse representation-based visual tracking approaches have attracted increasing interests in the community in recent years. The main idea is to linearly represent each target candidate using a set of target and trivial templates, while imposing a sparsity constraint onto the representation coefficients. After we obtain the coefficients using ℓ1-norm minimization methods, the candidate with the lowest error, when it is reconstructed using only the target templates and the associated coefficients, is considered as the tracking result. In spite of promising system performance widely reported, it is unclear if the performance of these trackers can be maximized. In addition, computational complexity caused by the dimensionality of the feature space limits these algorithms in real-time applications. In this paper, we propose a real-time visual tracking method based on structurally random projection (RP) and weighted least squares (WLS) techniques. In particular, to enhance the discriminative capability of the tracker, we introduce background templates to the linear representation framework. To handle appearance variations over time, we relax the sparsity constraint using a WLS method to obtain the representation coefficients. To further reduce the computational complexity, structurally RP is used to reduce the dimensionality of the feature space, while preserving the pairwise distances between the data points in the feature space. Experimental results show that the proposed approach outperforms several state-of-the-art tracking methods.
Shengping Zhang, Huiyu Zhou 0001, Feng Jiang 0001, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2014 Regularised region-based Mixture of Gaussians for dynamic background modelling
abstract
This paper introduces a momentum-like regularisation term for the region-based Mixture of Gaussians framework. Momentum term has long been used in machine learning, especially in backpropagation algorithms to improve the speed of convergence and subsequently their performance. Here, we prove the convergence of the online gradient method with a momentum term and apply it to background modelling by using it in the update equations of the region-based Mixture of Gaussians algorithm. It is then shown with the help of experimental evaluation on both simulated data and well known video sequences that these regularised updates help improve the performance of the algorithm.
Sriram Varadarajan, Hongbin Wang 0005, Paul Miller 0003, Huiyu Zhou 0001
AVSS4
2014 Video Event Recognition by Dempster-Shafer Theory
abstract
This paper presents an event recognition framework, based on Dempster-Shafer theory, that combines evidence of events from low-level computer vision analytics. The proposed method employing evidential network modelling of composite events, is able to represent uncertainty of event output from low level video analysis and infer high-level events with semantic meaning along with degrees of belief. The method has been evaluated on videos taken of subjects entering and leaving a seated area. This has relevance to a number of transport scenarios, such as onboard buses and trains, and also in train stations and airports. Recognition results of 78% and 100% for four composite events are encouraging.
Wenjun Ma, Paul Miller 0003, Weiru Liu, Huiyu Zhou 0001
ECAI6
2014 Low-rank representation based action recognition
abstract
Human action recognition is an important problem in computer vision, which has been applied to many applications. However, how to learn an accurate and discriminative representation of videos based on the features extracted from videos still remains to be a challenging problem. In this paper, we propose a novel method named low-rank representation based action recognition to recognize human actions. Given a dictionary, low-rank representation aims at finding the lowest-rank representation of all data, which can capture the global data structures. According to its characteristics, low-rank representation is robust against noises. Experimental results demonstrate the effectiveness of the proposed approach on several publicly available datasets.
Xiangrong Zhang, Yang Yang 0062, Hanghua Jia, Huiyu Zhou 0001, Licheng Jiao
IJCNN4
2014 A sparse representation based fast detection method for surface defect detection of bottle caps
Wenju Zhou, Minrui Fei, Huiyu Zhou 0001, Kang Li 0002
Neurocomputing3
2014 Adaptive fusion of particle filtering and spatio-temporal motion energy for human tracking
Huiyu Zhou 0001, Minrui Fei, Abdul Hamid Sadka, Xuelong Li 0001
Pattern Recognit.1
2013 Spatial mixture of Gaussians for dynamic background modelling
abstract
Modelling pixels using mixture of Gaussian distributions is a popular approach for removing background in video sequences. This approach works well for static backgrounds because the pixels are assumed to be independent of each other. However, when the background is dynamic, this is not very effective. In this paper, we propose a generalisation of the algorithm where the spatial relationship between pixels is taken into account. In essence, we model regions as mixture distributions rather than individual pixels. Using experimental verification on various video sequences, we show that our method is able to model and subtract backgrounds effectively in scenes with complex dynamic textures.
Sriram Varadarajan, Paul Miller 0003, Huiyu Zhou 0001
AVSS3
2013 Mean shift based gradient vector flow for image segmentation
Huiyu Zhou 0001, Xuelong Li 0001, Gerald Schaefer, M. Emre Celebi 0001, Paul Miller 0003
Comput. Vis. Image Underst.1
2013 Robust visual tracking based on online learning sparse representation
Shengping Zhang, Hongxun Yao, Huiyu Zhou 0001, Xin Sun 0003, Shaohui Liu
Neurocomputing3
2013 Special issue: Behaviours in video
Huiyu Zhou 0001, Yuan Yuan 0001, Yingzi Du, Pingkun Yan
Neurocomputing1
2013 "Pattern Recognition" special issue: Sparse representation for event recognition in video surveillance
Huiyu Zhou 0001, Jianguo Zhang 0001, Liang Wang 0001, Zhengyou Zhang, Lisa M. Brown
Pattern Recognit.1
2012 Rough C-means and Fuzzy Rough C-means for Colour Quantisation
abstract
Colour quantisation algorithms are essential for displaying true colour images using a limited palette of distinct colours. The choice of a good colour palette is crucial as it directly determines the quality of the resulting image. Colour quantisati
Gerald Schaefer, Qinghua Hu, Huiyu Zhou 0001, James F. Peters, Aboul Ella Hassanien
Fundam. Informaticae3
2012 Nonrigid Structure-From-Motion From 2-D Images Using Markov Chain Monte Carlo
abstract
In this paper we present a new method for simultaneously determining 3-D shape and motion of a nonrigid object from uncalibrated 2-D images without assuming the distribution characteristics. A nonrigid motion can be treated as a combination of a rigid rotation and a nonrigid deformation. To seek accurate recovery of deformable structures, we estimate the probability distribution function of the corresponding features through random sampling, incorporating an established probabilistic model. The fitting between the observation and the projection of the estimated 3-D structure will be evaluated using a Markov chain Monte Carlo based expectation maximization algorithm. Applications of the proposed method to both synthetic and real image sequences are demonstrated with promising results.
Huiyu Zhou 0001, Xuelong Li 0001, Abdul Hamid Sadka
IEEE Trans. Multim.1
2012 Classification of Upper Limb Motion Trajectories Using Shape Features
abstract
To understand and interpret human motion is a very active research area nowadays because of its importance in sports sciences, health care, and video surveillance. However, classification of human motion patterns is still a challenging topic because of the variations in kinetics and kinematics of human movements. In this paper, we present a novel algorithm for automatic classification of motion trajectories of human upper limbs. The proposed scheme starts from transforming 3-D positions and rotations of the shoulder/elbow/wrist joints into 2-D trajectories. Discriminative features of these 2-D trajectories are, then, extracted using a probabilistic shape-context method. Afterward, these features are classified using a k-means clustering algorithm. Experimental results demonstrate the superiority of the proposed method over the state-of-the-art techniques.
Huiyu Zhou 0001, Huosheng Hu, Honghai Liu 0001, Jinshan Tang
IEEE Trans. Syst. Man Cybern. Part C1
2011 Multi-camera detection association for 3D localisation
abstract
A multi-camera system is described for 3-D localisation of subjects within a confined space. In particular, we present a novel neighbourhood association algorithm to solve the problem of associating detections in multiple camera views with subjects. To evaluate our approach, experiments were conducted using multiple view video sequences of up to four subjects simulating typical passenger behaviour on a bus. ROC curves were generated for three different versions which showed that for smaller values of the neighbourhood radius parameter, the system tended to over-estimate the number of subjects. However, increasing the radius reduced the over-estimation from 60% to 5%.
Jiali Shen, Paul Miller 0003, Huiyu Zhou 0001, Michael Loughlin
AVSS3
2011 Age classification using Radon transform and entropy based scaling SVM
abstract
This paper mainly addresses the problem of age classification. Image fe atures can be extracted using a difference of Gaussian filter followed by Radon tran sform. The relevance and importance of these features are determined in a scaling support vector machine classifier, where zero weights are assigned to irrelevant varia bles. To enhance the quality of feature selection, we introduce entropy estimation to the scaling c lassifier. Experimental results demonstrate that the proposed algorithm leads to better recognition accuracy than the state of the art.
Huiyu Zhou 0001, Paul Miller 0003, Jianguo Zhang 0001
BMVC1
2011 Modeling and representing events in multimedia
abstract
This paper presents an overview of the Joint Workshop on Modeling and Representing Events (JMRE), which is held as part of ACM Multimedia 2011. JMRE is concerned with the understanding of events from multimedia, and with using events in order to better organize and consume multimedia.
Vasileios Mezaris, Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli, Huiyu Zhou 0001, Jianguo Zhang 0001, Liang Wang 0001, Zhengyou Zhang
ACM Multimedia5
2011 Combining Perceptual Features With Diffusion Distance for Face Recognition
abstract
Face recognition and identification is a very active research area nowadays due to its importance in both human computer and social interaction. Psychological studies suggest that face recognition by human beings can be featural, configurational, and holistic. In this paper, by incorporating spatially structured features into a histogram-based face-recognition framework, we intend to pursue consistent performance of face recognition. In our proposed approach, while diffusion distance is computed over a pair of human face images, the shape descriptions of these images are built using Gabor filters that consist of a number of scales and levels. It demonstrates that the use of perceptual features by Gabor filtering in combination with diffusion distance enables the system performance to be significantly improved, compared to several classical algorithms. The oriented Gabor filters lead to discriminative image representations that are then used to classify human faces in the database.
Huiyu Zhou 0001, Abdul Hamid Sadka
IEEE Trans. Syst. Man Cybern. Part C1
2010 Intelligent Sensor Information System For Public Transport - To Safely Go
abstract
The Intelligent Sensor Information System (ISIS) is described. ISIS is an active CCTV approach to reducing crime and anti-social behavior on public transport systems such as buses. Key to the system is the idea of event composition, in which directly detected atomic events are combined to infer higher-level events with semantic meaning. Video analytics are described that profile the gender of passengers and track them as they move about a 3-D space. The overall system architecture is described which integrates the on-board event recognition with the control room software over a wireless network to generate a real-time alert. Data from preliminary data-gathering trial is presented.
Paul Miller 0003, Weiru Liu, Chris Fowler, Huiyu Zhou 0001, Jiali Shen, Jianbing Ma, Jianguo Zhang 0001, Wei Qi Yan 0001, Kieran McLaughlin, Sakir Sezer
AVSS4
2010 Human Localization in a Cluttered Space Using Multiple Cameras
abstract
The use of single and dual-camera approaches to locating a subject in a 3-D cluttered space is investigated. Specifically, we investigate the case where the lower portion of the body may be occluded, e.g., by a chair on a bus. Experiments were conducted involving eleven subjects moving along a pre-designated route within a cluttered space. For each time instant the position of each subject was manually estimated and compared to that produced automatically. The dual camera approach was found to give significantly better performance than the single camera approach. It was found that inaccurate bounding of the lowest part of the subject, due to occlusion, led to localisation errors in range as large as 10m for the latter. Using the side bounds of the detected object, which were found to be robust, accurate azimuth estimates can be obtained for a single camera. The dual-camera approach exploits the greater degree of accuracy in azimuth to estimate the range through triangulation, giving average localisation errors of 40cm over the space of interest.
Jiali Shen, Wei Qi Yan 0001, Paul Miller 0003, Huiyu Zhou 0001
AVSS4
2010 A new family of order-statistics based switching vector filters
abstract
In this paper, we present a family of order-statistics based vector filters for the removal of impulsive noise from color images. These filters preserve the edges and fine image details by switching between the identity (no filtering) operation and a robust order-statistics based filter operation based on the univariate median operator. Experiments on a diverse set of images and comparisons with state-of-the-art filters show that the proposed filters combine simplicity, flexibility, good filtering quality, and low computational requirements.
M. Emre Celebi 0001, Gerald Schaefer, Huiyu Zhou 0001
ICIP3
2010 Robust estimation of the fundamental matrix
abstract
Most approaches to estimate the fundamental matrix assume a Gaussian distribution in the errors in view of mathematical tractability. However, this assumption is violated if the distribution computed is not normal. In this paper we propose a robust approach of estimating the fundamental matrix which does not rely on the Gaussian assumption. The proposed technique, weighted least squares (WLS), is the application of linear mixed-effects models considering the correlation between different data sub-samples. It provides an unbiased estimation of the fundamental matrix which is not affected by outlier samples. Experimental results on synthetic and real images confirm the accuracy of our method and its superiority to standard estimation methods.
Huiyu Zhou 0001, Gerald Schaefer
ICIP1
2010 Feature extraction and clustering for dynamic video summarisation
Huiyu Zhou 0001, Abdul Hamid Sadka, Mohammad Rafiq Swash, Jawid Azizi, Umar A. Sadiq
Neurocomputing1
2010 Segmentation of optic disc in retinal images using an improved gradient vector flow algorithm
Huiyu Zhou 0001, Gerald Schaefer, Tangwei Liu, Faquan Lin
Multim. Tools Appl.1
2009 Bayesian image segmentation with mean shift
abstract
Image segmentation plays a key role in many image content analysis applications, and a lot of effort has aimed at improving the performance of established segmentation algorithms. In this paper, we present a mean shift-based combined Dirichlet process mixture (MDP)/Markov Random Field (MRF) image segmentation algorithm. Our method incorporates a mean shift process to iteratively reduce the difference between the mean of cluster centres and image pixels within the standard MDP/MRF procedure. Experimental results show that the proposed segmentation technique outperforms the classical MDP/MRF algorithm.
Huiyu Zhou 0001, Gerald Schaefer, M. Emre Celebi 0001, Minrui Fei
ICIP1
2009 3-D structure recovery from 2-D observations
abstract
In this paper we present a novel method for simultaneously determining three dimensional motion and structure of a non-rigid object from its uncalibrated two dimensional data with Gaussian or non-Gaussian distributions. A non-rigid motion can be treated as a combination of a rigid component and a non-rigid deformation. To reduce the high dimensionality of the deformable structure or shape, we estimate the probability distribution function of the structure through random sampling, integrating an established probabilistic model. The fitting between the observations and the estimated 3-D structure is evaluated using the pooled variance estimator. Applications of the proposed method to both synthetic and real image sequences show promising results.
Huiyu Zhou 0001, Gerald Schaefer, Tangwei Liu, Faquan Lin
ICIP1
2009 Object trajectory clustering via tensor analysis
abstract
In this paper we present a new video object trajectory clustering algorithm1, which allows us to model and analyse the patterns of object behaviors based on the extracted features using tensor analysis. The proposed algorithm consists of three steps as follows: extraction of trajectory features by tensor analysis, non-parametric probabilistic mean shift clustering and clustering correction. The performance of the proposed algorithm is evaluated on standard data-sets and compared with classical techniques.
Huiyu Zhou 0001, Dacheng Tao, Yuan Yuan 0001, Xuelong Li 0001
ICIP1
2009 Object tracking using SIFT features and mean shift
Huiyu Zhou 0001, Yuan Yuan 0001, Chunmei Shi
Comput. Vis. Image Underst.1
2009 Estimation of epipolar geometry by linear mixed-effect modelling
Huiyu Zhou 0001, Patrick R. Green, Andrew M. Wallace
Neurocomputing1
2009 Non-rigid object tracking in complex scenes
Huiyu Zhou 0001, Yuan Yuan 0001, Chunmei Shi
Pattern Recognit. Lett.1
2009 Efficient tracking and ego-motion recovery using gait analysis
Huiyu Zhou 0001, Andrew M. Wallace, Patrick R. Green
Signal Process.1
2008 Level set image segmentation with Bayesian analysis
Huiyu Zhou 0001, Yuan Yuan 0001, Faquan Lin, Tangwei Liu
Neurocomputing1
2008 Recovery of Nonrigid Structures from 2D Observations
abstract
We present a new method for simultaneously determining three-dimensional (3D) motion and structure of a nonrigid object from its uncalibrated two-dimensional (2D) data with Gaussian or non-Gaussian distributions. A nonrigid motion can be treated as a combination of a rigid component and a nonrigid deformation. To reduce the high dimensionality of the deformable structure or shape, we estimate the probability distribution function (PDF) of the structure through random sampling, integrating an established probabilistic model. The fitting between the observations and the estimated 3D structure will be evaluated using the pooled variance estimator. The recovered structure is only available when the 2D feature points have been properly corresponded over two image frames. Applications of the proposed method to both synthetic and real image sequences are demonstrated with promising results.
Huiyu Zhou 0001, Xuelong Li 0001, Tangwei Liu, Faquan Lin, Yusheng Pang, Ji Wu 0010, Junyu Dong
Int. J. Pattern Recognit. Artif. Intell.1
2008 Application of semantic features in face recognition
Huiyu Zhou 0001, Yuan Yuan 0001, Abdul Hamid Sadka
Pattern Recognit.1
2008 Improving image segmentation by gradient vector flow and mean shift
Tangwei Liu, Huiyu Zhou 0001, Faquan Lin, Yusheng Pang, Ji Wu 0010
Pattern Recognit. Lett.2
2005 A hybrid framework for image segmentation
abstract
This paper presents a new approach for image segmentation by combining the classical gradient vector flow (GVF) algorithm with mean shift. Due to the dependence on the gradient vectors of an edge map, the classical GVF is sensitive to the shape irregularities, and hence the snake cannot be ideally located on the concave boundaries. We propose an improved representation of the internal energy force by reducing the Euclidean distance between the guessed centroid and the estimated one of the snake. Experimental work shows the performance of this approach in different tests.
Huiyu Zhou 0001, Tangwei Liu, Huosheng Hu, Yusheng Pang, Faquan Lin, Ji Wu 0010
ICASSP (2)1
2004 Efficient motion tracking using gait analysis
abstract
For navigation and obstacle detection, it is necessary to develop robust and efficient algorithms to compute ego-motion and model the changing scene. These algorithms must cope with the high video data rate from the input sensor. In this paper, we present an approach to achieve improved motion tracking from a monocular image sequence acquired by a camera attached to a pedestrian. The human gait is modelled from the motion history of the camera, and used to predict the feature positions in successive frames. This is encoded within a maximum a posteriori (MAP) framework to seek fast and robust motion estimation. Experimental results show how use of the gait model can reduce the computational load by allowing longer gaps between successive frames, while retaining the robust ability to track features.
Huiyu Zhou 0001, Patrick R. Green, Andrew M. Wallace
ICASSP (3)1
2003 A multistage filtering technique to detect hazards on the ground plane
Huiyu Zhou 0001, Andrew M. Wallace, Patrick R. Green
Pattern Recognit. Lett.1