VLDB 2026 Research / reviewers in the wild / expert
Yicong Zhou
dblp:51/8204
· DBLP profile ↗
275ranked-venue papers
13as first author
144since 2021 · last 2026
0000-0002-4487-6384ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 136 · 4 first-author · 81 since 2021Artificial intelligence and machine learning · 73 · 4 first-author · 40 since 2021Applied, interdisciplinary, general and emerging computing · 51 · 4 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 24 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 16 · 1 first-author · 2 since 2021Computer networks · 12 · 12 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-view Anchor Graph Learning and Factorization for Incomplete Multi-view ClusteringabstractGraph-based incomplete multi-view clustering algorithms have gathered much attention due to their impressive clustering performance. However, existing methods primarily leverage intra-view correlation from observed views, while ignoring the exploration of explicit compensation relationships between different views. Moreover, these methods need post-processing to get labels, and the separate steps lack negotiation, which may lead to sub-optimal solutions. To address these issues, we propose a Cross-view Anchor Graph Learning and Factorization (AGLF) method. AGLF develops an Anchor Graph Completion (AGC) framework that explicitly learn the missing subgraph structures. Instead of requiring post-processing, AGC directly produces soft labels. By establishing a third-order tensor of soft labels, it employs the tensor Schatten p-norm to enhance anchor graph learning and factorization. To significantly improve the quality of subgraph learning, AGLF incorporates compensation subgraphs from supplementary views into the AGC framework, enabling the construction of a better anchor graph for label learning. An optimization algorithm is devised to solve the objective function. Experimental results across various datasets demonstrate the effectiveness of our method. Xinxin Wang 0003, Yongshan Zhang, Xiaochen Yuan, Yicong Zhou |
AAAI | 4 |
| 2026 | Dense Cross-Scale Image Alignment with Fully Spatial Correlation and Just Noticeable Difference GuidanceabstractExisting unsupervised image alignment methods exhibit limited accuracy and high computational complexity. To address these challenges, we propose a dense cross-scale image alignment model. It takes into account the correlations between cross-scale features to decrease the alignment difficulty. Our model supports flexible trade-offs between accuracy and efficiency by adjusting the number of scales utilized. Additionally, we introduce a fully spatial correlation module to further improve accuracy while maintaining low computational costs. We incorporate the just noticeable difference to encourage our model to focus on image regions more sensitive to distortions, eliminating noticeable alignment errors. Extensive quantitative and qualitative experiments demonstrate that our method surpasses state-of-the-art approaches. Jinkun You, Jiaxue Li, Yicong Zhou |
AAAI | 4 |
| 2026 | Dynamic Event-Triggered Stabilization for Parameter-Varying Strict-Feedback Nonlinear Systems: A Two-Level Adaptive Estimation MethodabstractIt is an interesting problem to achieve adaptive asymptotic control for strict-feedback nonlinear systems with fast time-varying parameters, particularly in the presence of substantial parameter uncertainties without the availability of a priori knowledge. In this paper, a solution to this problem is presented with a two-level adaptive estimator and a new form of event-triggered control mechanism. Specifically, three adaptive laws (two for the uncertain parameters in the feedback path and one for the uncertain parameters in the input path), two sets of tuning functions, and a dynamic event-triggering mechanism are integrated and strategically designed within a controller to ensure that higher control precision can be obtained in the presence of time-varying parameters. It is shown that, with the derived dynamic event-triggered adaptive asymptotic control strategy, the closed-loop system is globally uniformly asymptotically stable. Furthermore, the inter-event intervals are guaranteed to be lower-bounded by a positive constant. The benefits and effectiveness of the proposed scheme are validated through two numerical simulations. Hefu Ye, Yicong Zhou |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | CO3+: Improved Collaborative Consortium of Foundation Models for Open-World Few-Shot LearningabstractOpen-World Few-Shot Learning (OFSL) is a critical research domain focused on accurately identifying target samples under conditions where data is scarce and labels are unreliable. This field is highly relevant to real-world scenarios, holding significant practical implications. Currently, the field has only a few solutions, primarily relying on conventional methods such as metric learning and feature aggregation. However, these methods often struggle in more complex scenarios. Recent breakthroughs in foundation models such as CLIP and DINO have demonstrated their strong representational capabilities, even in resource-limited environments. These advancements have led to a shift from “training model from scratch” towards “exploiting the extensive capabilities and expertise of these pre-trained foundation models for OFSL”. Inspired by this shift, we introduce the Improved Collaborative Consortium of Foundation Models (CO+3), an extension of CO3, first presented in AAAI 2024. CO+3significantly improves the accuracy of OFSL by integrating the strengths of four foundational models. It includes three decoupled blocks: (1) The Label Correction Block (LC-Block) rectifies unreliable labels, (2) the Data Augmentation Block (DA-Block) enriches the available data, and (3) the Text-guided Fusion Adapter (TeFu-Adapter) merges various features and reduces the impact of noisy labels through semantic constraints. We evaluate CO+3across eleven benchmark datasets, comparing it against recent state-of-the-art methods. Our thorough evaluations demonstrate that the proposed CO+3consistently surpasses existing methods by a substantial margin, particularly in high-noise scenarios. Shuai Shao 0006, Rui Xu 0012, Bingfeng Zhang, Baodi Liu, Weifeng Liu 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | RWKVSR: Receptance Weighted Key-Value Network for Hyperspectral Image Super-ResolutionabstractDeep learning has achieved significant success in hyperspectral image super-resolution (HSISR) by leveraging advanced feature extraction techniques to reconstruct high-resolution images from low-resolution counterparts. However, existing methods predominantly utilize 2D/3D convolutions or Transformer architectures, which are often hindered by limited receptive fields, quadratic computational complexity, and inadequate fusion of spatial-spectral dependencies. To address these challenges, this paper proposes RWKVSR, a novel lightweight network that integrates a Receptance Weighted Key-Value (RWKV) architecture for efficient HSISR. The proposed RWKVSR comprises of three key components: (1) A linear-complexity RWKV module replacing quadratic self-attention, enabling efficient global spectral-spatial modeling; (2) A Spectral-Spatial Residual Module (SSRM) employing anisotropic, direction-separable 3D convolutions to hierarchically extract multi-scale features while enhancing local-global interactions; and (3) A Hyperspectral Frequency Loss (HFL) optimizing spectral consistency by prioritizing high-frequency structural alignment between reconstructed and ground-truth images in the frequency domain. Extensive experiments conducted on the CAVE and Harvard datasets demonstrate that RWKVSR outperforms the existing state-of-the-art methods, effectively balancing accuracy and efficiency, and providing a practical solution for high-quality HSI reconstruction. Our paper code is publicly available at https://github.com/backy-1/RWKVSR.git. Xiaofei Yang 0002, Sihuan Li, Weijia Cao, Yifang Ban, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Test-Time Adaptation for Detecting Image Inpainting ForgeriesabstractThe rapid development of deep learning-based image inpainting poses serious challenges to image authenticity. As inpainting methods continue to evolve, the inpainted images exhibit extremely high visual fidelity, presenting recognition difficulties to the forgery detection model due to differences in operational mode and forgery traces among methods. In particular, the detection performance tends to drop significantly in the testing phase when the test samples differ from the training data. To address this issue, we propose a test-time adaptive detection framework for image inpainting forgeries. First, we propose an image gradient-based metric that quantifies model uncertainty and orchestrates the entire adaptation process. Integrating this metric with sample-specific batch normalization (BN) statistics enhances the ability of pretrained models in the inference stage. Second, we introduce a cross-attention module as a side-tuning module, enabling the model to adapt dynamically to reliable test samples without altering the backbone network. To validate the effectiveness of the proposed method, we construct a dataset comprising synthetic images of multiple inpainting methods and design experiments under two scenarios of distributional bias. The results demonstrate that our proposed framework outperforms the existing baseline method, enhancing the adaptability and detection performance of the forgery detection model in dynamic environments. Guopu Zhu, Hongli Zhang 0001, Xinpeng Zhang 0001, Yicong Zhou, Ligang Wu 0001 |
IEEE Trans. Cybern. | 5 |
| 2026 | EEG Emotion Recognition With Uncertainty-Aware Contrastive Learning and Frequency-Aware Self-AttentionabstractElectroencephalography (EEG) emotion recognition plays a key role in improving human-machine interactions. Advanced algorithms have been proposed for this task. However, two challenges remain, i.e., unclear decision boundary in the embedded space and noise in physiological signals from various devices. To this end, we develop a novel framework, namely, UACL-Net, for EEG emotion recognition. It is based on uncertainty-aware contrastive learning (UACL) and frequency-aware self-attention (FASA). Specifically, UACL uses a multivariate Gaussian distribution to construct the latent space for different emotions. It is able to highlight interclass differences, thereby improving the robustness of model decisions. In addition, FASA generates learnable weights by applying self-attention (SA) to the real and imaginary components in the frequency domain. This helps adaptively reduce noise and capture global dependencies in temporal sequences. Our model is trained and tested on four benchmark datasets, achieving up to 94.88%, 98.71%, 96.91%, and 99.29% accuracy on SEED, DEAP, DREAMER, and FACED, respectively. Experimental results demonstrate that it is effective and has advantages over peer state-of-the-art (SOTA) methods. Junxin Chen 0001, Qiang He 0002, Yongfei Wu, Yicong Zhou |
IEEE Trans. Cybern. | 5 |
| 2026 | SHL-Net: Semantics-Enhanced Network for Localizing Harmonized Image Splicing
Xiwen Fu, Guopu Zhu, Hongli Zhang 0001, Jiwu Huang, Tao Xiang 0001, Yicong Zhou, Ligang Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | MECI: A Multi-Model Motion Capture-Free Event Dataset Featuring Large-Scale Challenging Indoor EnvironmentsabstractRecently, several event-based datasets have emerged to foster the application of the new event camera to classic vision tasks like Simultaneous Localization and Mapping (SLAM). However, current indoor benchmark datasets depend on the expensive motion capture system to obtain ground-truth trajectories, restricting data acquisition to small object-centric scenes or single-room environments due to infrastructure costs and spatial limitations. Furthermore, these datasets lack sensor diversity, relying solely on a single event camera model that hinders practical cross-device generalization. To address the above limitations, we propose MECI, the first Multi-model Event dataset targeting Challenging Indoor environments, especially including large-scale scenes with complete trajectory ground-truth provided. Specifically, MECI includes 38 posed sequences of visual-inertial-event data from three different models of event cameras with varying resolutions and data frequencies. Apart from typical challenging factors such as illumination changes, motion blur, and dynamics, these sequences involve large-scale indoor scenes (across rooms and floors) with each room occupying approximately 60 m \({}^{2}\) and a maximum trajectory length of 343.8 m. We innovatively leverage ETS (Electronic Total Station) measurements and AprilTag markers to provide complete 6-DoF trajectories with low cost and minimal environmental modifications. Extensive experiments show that our MECI is an effective yet challenging benchmark dataset not only for visual SLAM, but also for event-image reconstruction. Our project page is https://cslinzhang.github.io/MECI_dataset/ . Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Highly Efficient Rotation-Invariant Spectral Embedding for Scalable Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering presents significant challenges due to missing views. Although many existing graph-based methods aim to recover missing instances or complete similarity matrices with promising results, they still face several limitations: (1) Recovered data may be unsuitable for spectral clustering, as these methods often ignore guidance from spectral analysis; (2) Complex optimization processes require high computational burden, hindering scalability to large-scale problems; (3) Most methods do not address the rotational mismatch problem in spectral embeddings. To address these issues, we propose a highly efficient rotation-invariant spectral embedding (RISE) method for scalable incomplete multi-view clustering. RISE learns view-specific embeddings from incomplete bipartite graphs to capture the complementary information. Meanwhile, a complete consensus representation with second-order rotation-invariant property is recovered from these incomplete embeddings in a unified model. Moreover, we design a fast alternating optimization algorithm with linear complexity and promising convergence to solve the proposed formulation. Extensive experiments on multiple datasets demonstrate the effectiveness, scalability, and efficiency of RISE compared to the state-of-the-art methods. Xinxin Wang 0003, Yongshan Zhang, Yicong Zhou |
AAAI | 3 |
| 2025 | Learn Multi-task Anchor: Joint View Imputation and Label Generation for Incomplete Multi-view ClusteringabstractAnchor-based incomplete multi-view clustering methods utilize anchors to uncover clustering structures. However, relying on anchor graphs for producing final indicators is indirect, which can lead to information loss and suboptimal outcomes. Besides, most methods neglect the potential of anchors for imputing missing views. To address these limitations, we propose a Joint View Imputation and Label Generation (JVILG) method. JVILG comprises the Anchor-based tensorized Label Generation (ALG) module for generating clustering labels and the Anchor-based sparse regularized Subspace Correlation (ASC) module for recovering missing views. The ALG module explicitly connects data observations, the fine-grained anchor matrix, and soft label matrices within a reconstruction framework through a membership matrix, while imposing tensor Schatten p-norm regularization on the constructed label tensor to capture spatial correlations among views. Meanwhile, the ASC module directly uses fine-grained anchors to impute missing data in respective views. By integrating the ALG and ASC modules, JVILG enhances synergy between different tasks and mitigates the impact of missing information on clustering. Experimental results on six datasets demonstrate the effectiveness of JVILG compared to both shallow and deep state-of-the art methods.The code is available at https://github.com/W-Xinxin/JVILG. Xinxin Wang 0003, Yongshan Zhang, Yicong Zhou |
IJCAI | 3 |
| 2025 | Application of Artificial Intelligence in Rock Tunnel Engineering: A Survey on Where and HowabstractABSTRACT Rock tunnel engineering (RTE) plays a crucial role in modern infrastructure development. The development of artificial intelligence (AI) is able to drive transformative advances in RTE. This review provides an in‐depth analysis of the AI application in RTE. Through a comprehensive examination of existing literature, we explore how AI technologies have revolutionised various aspects of RTE, including construction methodology, rock parameter estimation, hazard disaster management during construction, and tunnel operation. In addition, we provide an in‐depth study of the synergies between various AI algorithms and related open datasets. This work also outlines promising future research directions for the AI application in RTE, aiming to inspire further advancements in this emerging field. In conclusion, this review underscores the positive influence of AI on RTE, emphasising its capacity to elevate efficiency, accuracy, and safety standards throughout various phases of tunnel projects. The convergence of AI with RTE holds immense promise for advancing the field and ensuring the success and sustainability of future tunnel infrastructure endeavours. Xiaojie Yu, Ben-Guo He, Yicong Zhou, Miguel A. Diaz, Junxin Chen 0001, David Camacho |
Expert Syst. J. Knowl. Eng. | 4 |
| 2025 | Online indoor visual odometry with semantic assistance under implicit epipolar constraints
Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
Pattern Recognit. | 4 |
| 2025 | Feature aggregation and connectivity for object re-identification
Dongchen Han, Baodi Liu, Shuai Shao 0006, Weifeng Liu 0001, Yicong Zhou |
Pattern Recognit. | 5 |
| 2025 | Two-Dimensional Cyclic Chaotic System for Noise-Reduced OFDM-DCSK CommunicationabstractSecure communication techniques can protect data confidentiality during transmission through public channels. Chaotic systems are commonly used in secure communication due to their random-like behavior, unpredictability, and ergodicity. However, existing chaos-based secure communication schemes have some drawbacks concerning the chaotic systems used and the communication structures, so they cannot achieve satisfactory performance to resist transmission channel noise. In light of this, in this paper, we propose a two-dimensional (2D) cyclic chaotic system (2D-CCS) and design a novel chaos-based secure communication scheme called noise-reduced orthogonal frequency division multiplexing based differential chaos shift keying (NR-OFDM-DCSK). The 2D-CCS is a general framework that can generate a large number of new 2D chaotic maps using existing one-dimensional (1D) chaotic maps as seed maps. Theoretical analysis and experiment results demonstrate its robust chaotic behaviors. The NR-OFDM-DCSK employs a new chaotic map generated by 2D-CCS as the chaos generator, and its structure exhibits a strong ability to resist channel noise, as demonstrated by formulaic analysis. Our extensive experiments show that our developed 2D chaotic maps are more suitable for secure communication applications than existing 2D chaotic maps, and our NR-OFDM-DCSK can achieve a lower bit-error-rate (BER) than state-of-the-art secure communication schemes. Zhongyun Hua, Zihua Wu, Yinxing Zhang, Han Bao 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Incomplete Multiview Clustering Using Discriminative Feature Recovery and Tensorized Matrix FactorizationabstractMultiview clustering task groups objects using multiple properties, such as RGB images, infrared images, and texture information. However, incomplete multi-view clustering faces significant challenges due to missing views that hinder clustering performance. This paper proposes a Discriminative Feature Recovery and Tensorized Matrix Factorization method (DFRTMF) that effectively recovers missing views, learns low-dimensional discriminative embeddings, and enables direct clustering. DFRTMF addresses high dimensionality through projection learning and enables the output of soft indicators. To improve projection and facilitate the recovery of missing views, we propose an uncorrelated constraint based on the scatter matrix of the recovered complete data, exploring the correlations between observed and missing views. To capture high-order correlations among views, a low-rank tensor constraint based on tensor Schatten p-norm regularization is applied to a third-order tensor composed of soft indicator matrices. DFRTMF adaptively controls the inter-coordination between these factorizations using view weights to optimally explore complementary information. Furthermore, we propose an alternating optimization algorithm based on the Alternating Direction Method of Multipliers to effectively solve the proposed objective function. Extensive experiments across diverse datasets demonstrate the effectiveness of DFRTMF compared to the state-of-the-art methods. Xinxin Wang 0003, Yongshan Zhang, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | I-DACS: Always Maintaining Consistency Between Poses and the Field for Radiance Field Construction Without Pose PriorabstractThe radiance field, emerging as a novel 3D scene representation, has found widespread application across diverse fields. Standard radiance field construction approaches rely on the ground-truth poses of key-frames, while building the field without pose prior remains a formidable challenge. Recent advancements have made strides in mitigating this challenge, albeit to a limited extent, by jointly optimizing poses and the radiance field. However, in these schemes, the consistency between the radiance field and poses is achieved completely by training. Once the poses of key-frames undergo changes, long-term training is required to readjust the field to fit them. To address such a limitation, we propose a new solution for radiance field construction without pose prior, namely I-DACS (Incremental radiance field construction with Direction-Aware Color Sampling). Diverging from most of the existing global optimization solutions, we choose to incrementally solve the poses and construct a radiance field within a sliding-window framework. The poses are unequivocally retrieved from the radiance field, devoid of any constraints and accompanying noise from other observation models, so as to achieve the consistency of poses to the field. Besides, in the radiance field, the color information is much higher-frequency and more time-consuming to learn compared with the density. To accelerate training, we isolate the color information to a distinct color field, and construct the color field based on an innovative direction-aware color sampling strategy, by which the color field can be derived directly from images without training. The color field obtained in this way is always consistent with the poses, and intricate details of training images can be retained to the utmost extent. Extensive experimental results evidently showcase both the remarkable training speed and the outstanding performance in rendering quality and localization accuracy achieved by I-DACS. To make our results reproducible, the source code has been released athttps://cslinzhang.github.io/I-DACS-MainPage/. Tianjun Zhang, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Seam-Adaptive Structure-Preserving Image Stitching for Drone ImagesabstractDrones have been widely used for remote sensing applications. To perform high-quality drone image stitching, this article first proposes a local and global structure-preserving alignment (LGSPA) method that aligns drone images from local dual feature-based and global pixel-based alignment perspectives, while maintaining local linear and global collinear image structures. To enable an optimal image stitching performance, we then propose a seam-adaptive weighting (SAW) scheme to enhance the local alignment accuracy under the guidance of a seam prior. On the ground of LGSPA and SAW, we further develop a seam-adaptive structure-preserving (SASP) image stitching framework to generate the final stitched drone images. Both qualitative and quantitative experimental results demonstrate that LGSPA and SASP are capable of generating higher quality alignment and stitching results than several state-of-the-art methods over multiple challenging aerial scenarios, including low textures, repetitive textures, large parallax, wide baseline, and occlusions. Jiaxue Li, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Dilated Transformation-Guided Unsupervised Multimodal Learning for Hyperspectral and Multispectral Image FusionabstractMultimodal fusion widely uses convolutional layers to capture local correlations and adjust feature dimensions. However, the progressive expansion of the receptive field in convolutional layers often compromises spatial context retention, leading to the loss of fine details. Furthermore, the fixed-size kernels typically used in standard convolution restrict the network’s ability to capture multiscale contextual details. To address this limitation, this paper develops a dilated transformation-guided unsupervised multimodal learning (DTUML) method to fuse a high-resolution multispectral image (HR-MSI) and a low-resolution hyperspectral image (LR-HSI), thereby generating a high-resolution hyperspectral image (HR-HSI). Our DTUML adopts a dual-stream encoder architecture to conduct multimodal data, where one stream focuses on preserving spectral information from LR-HSIs, while the other emphasizes the acquisition of spatial details from HR-MSIs. These complementary features are subsequently integrated to ensure spectral fidelity and retain spatial detail. Then, a convolutional layer restores dimensional consistency and outputs an HR-HSI. Extensive experiments demonstrate the effectiveness of DTUML, showing superior performance and strong competitiveness compared to state-of-the-art methods. Code:https://github.com/yuanchaosu/TGRS-DTUML. Yuanchao Su, Yicong Zhou, Lianru Gao, Mengying Jiang, Xu Sun 0005, Enke Hou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Multimodal Remote Sensing Image Clustering With Multiscale Spectral-Spatial Anchor GraphsabstractExisting multiview clustering methods have achieved remarkable success for general images (GIs), but still have many limitations for clustering multimodal remote sensing images (RSIs). For example, these methods are sensitive to noise and spectral variability, ignore the diverse spatial structure information across modalities, or are computationally prohibitive for large-scale RSIs, thereby limiting their applications. This article proposes a multiscale spectral-spatial anchor graph fusion (MSSAGF) method for multimodal RSI clustering. MSSAGF develops a superpixel-based nonlinear neighborhood recovery strategy to reduce noise while enhancing spatial smoothness in multimodal RSIs. Using spatial-aware anchors to extract local spatial information for each modality, MSSAGF introduces multiscale local spectral-spatial anchor graphs to capture nonlinear correlations between the pixels and their corresponding local regions. A small number of anchors effectively reduces graph construction and partitioning costs, making the time complexity of MSSAGF nearly linear. This ensures that it is computationally feasible for large-scale RSIs. Finally, MSSAGF develops an adaptive fusion mechanism to fuse multiscale local anchor graphs into a unified global anchor graph, integrating complementary information across multiple modalities while directly obtaining the final clustering results. The experimental results on three multimodal RSI datasets demonstrate the superiority of our proposed method over state-of-the-art methods. Our code is publicly available athttps://github.com/W-Xinxin/MSSAGF. Xinxin Wang 0003, Yongshan Zhang, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | ACTN: Adaptive Coupling Transformer Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) and Transformer networks have shown impressive performance in hyperspectral image (HSI) classification. However, these models usually concentrate on examining either local or global representations of HSI data, frequently falling short of capturing multidimensional representations. Furthermore, these methods fail to fully leverage the strengths of CNNs and Transformers. This article presents the adaptive coupling Transformer network (ACTN), a parallel-hybrid network aiming to improve representation learning for HSI classification. ACTN can capture different types of representation and facilitate mutual learning. Specifically, we introduce a parallel-hybrid module called the adaptive coupling module (ACM), which is designed to capture multifaceted representations from the HSI cube. The ACM consists of two branches: a CNN branch that extracts local contextual representations and a Transformer branch that captures global dependency representations. Our proposal is an adaptive response fusion module (ARFM) that interacts with the hybrid module to merge local and global representations at different resolutions in an adaptive way. In addition, we utilize a cosine similarity function to restrict the loss function in mutual learning, guaranteeing the preservation of both local and global representations to the maximum extent. Extensive experiments conducted on three public HSI datasets demonstrate that ACTN outperforms state-of-the-art methods based on Transformers and CNNs. Xiaofei Yang 0002, Weijia Cao, Yicong Zhou, Yao Lu 0008 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Bidirectional Probabilistic Multi-Graph Learning and Decomposition for Multi-View ClusteringabstractGraph-based multi-view clustering has attracted remarkable attention due to its impressive performance. However, the typical framework consisting of graph learning and indicator generation may fail to align learned graphs with the underlying data structure due to the unidirectional pipeline from refined graphs to indicator generation. Another common problem is the inadequate prior information in graph learning methods. This paper proposes a Bidirectional Probabilistic Multi-graph Learning and Decomposition (BPMLD) method by establishing an explicit bidirectional pipeline between graph learning and indicator generation for multi-view clustering. Specifically, we design a confidence term based on clustering probability indicators and fuse it with graph learning to form clustering confidence driven graph learning. Meanwhile, graph tensor learning is introduced to recover the high-order correlations among the refined graphs. We further propose a multi-graph probability decomposition module to adaptively produce cluster indicators with probability representation from the refined graphs. The seamless integration between graph learning and indicator generation enables them to interact directly and enhance each other. To solve the proposed model, we design an effective optimization algorithm. Extensive experiments demonstrate the effectiveness of our method compared to state-of-the-art methods. The code is available at: https://github.com/W-Xinxin/BPMLD. Xinxin Wang 0003, Yongshan Zhang, Yicong Zhou |
IEEE Trans. Image Process. | 3 |
| 2025 | ATM-NeRF: Accelerating Training for NeRF Rendering on Mobile Devices via Geometric RegularizationabstractRecently, an increasing number of researchers have been dedicated to transferring the impressive novel view synthesis capability of Neural Radiance Fields (NeRF) to resource-constrained mobile devices. One common solution is to pre-train NeRF and bake it into textured meshes which are well supported by mobile graphics hardware. However, the training process of existing methods often requires several hours even with multiple high-end NVIDIA V100 GPUs. The underlying reason is that these schemes mainly rely on photometric rendering loss, neglecting the geometric relationship between the pre-trained NeRF and the baked results. Standing on this point, we presentATM-NeRF(AcceleratingTraining forMobile rendering based onNeRF), which is the first to apply effective geometric regularization constraints during both the pre-training and the baking training stages for faster convergence. Specifically, in the initial NeRF pre-training stage, we enforce consistency of the multi-resolution density grids representing the scene geometry to mitigate the shape-radiance ambiguity problem to some extent, achieving a coarse mesh with smoothness. In the second stage, we utilize the positions and geometric features of 3D points projected from the pre-trained posed depths to provide geometric supervision for joint refinement of geometry and appearance of the coarse mesh. As a result, our ATM-NeRF achieves comparable rendering quality to MobileNeRF with a training speed that is about$30\times \sim 70\times$faster while maintaining finer structure details of the exported mesh. Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Multim. | 4 |
| 2025 | Heterogeneous Domain Adaptation via Correlative and Discriminative Feature LearningabstractHeterogeneous domain adaptation seeks to learn an effective classifier or regression model for unlabeled target samples by using the well-labeled source samples but residing in different feature spaces and lying different distributions. Most recent works have concentrated on learning domain-invariant feature representations to minimize the distribution divergence via target pseudo-labels. However, two critical issues need to be further explored: 1) new feature representations should be not only domain-invariant but also category-correlative and discriminative and 2) alleviating the negative transfer caused by the incorrect pseudo-labeling target samples could boost the adaptation performance during the iterative learning process. To address these issues, in this paper, we put forward a novel heterogeneous domain adaptation method to learn category-correlative and discriminative representations, referred to as correlative and discriminative feature learning (CDFL). Specifically, CDFL aims to learn a feature space where class-specific feature correlations between the source and target domains are maximized, the divergences of marginal and conditional distribution between the source and target domains are minimized, and the distances of inter-class distribution are forced to be maximized to ensure the discriminative ability. Meanwhile, a selective pseudo-labeling procedure based on the correlation coefficient and classifier prediction is introduced to boost class-specific feature correlation and discriminative distribution alignment in an iteration way. Extensive experiments certify that CDFL outperforms the State-of-the-Art algorithms on five standard benchmarks. Yuwu Lu, Dewei Lin, LinLin Shen, Yicong Zhou, Jiahui Pan 0003 |
IEEE Trans. Multim. | 4 |
| 2025 | Learning Orthogonal Latent Representations for Multi-View Clustering
Xiaolin Xiao, Yue-Jiao Gong, Yicong Zhou |
IEEE Trans. Multim. | 3 |
| 2025 | Pseudo-Supervision Affinity Propagation for Efficient and Scalable Multiview ClusteringabstractAnchor graph-based multiview clustering (AGMVC) demonstrates high efficiency and satisfactory performance. However, it still suffers from limitations such as single-structure similarity measurement, high time expenditure for large-scale anchor graph partitioning, and limited generalization ability. To alleviate the instability problem of single-structure information, this article proposes an anchor graph construction method that learns local and global (LG) structures simultaneously. To eliminate the need for graph partitioning and address the out-of-sample problem, we develop a landmark learning method to produce structural anchors, and further propose a pseudo-supervision affinity propagation (PSAP) framework. This framework jointly optimizes graph construction and landmark learning to disentangle the in-cluster distribution between samples and anchors while accelerating convergence. In addition, our framework introduces a clustering inference partition (CIP) strategy to directly output clustering results without the need for time-consuming postprocessing. Extensive experiments validate the efficiency and effectiveness of our framework. Our code is publicly available at https://github.com/W-Xinxin/PSAP. Xinxin Wang 0003, Yongshan Zhang, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Skeleton-Aware Graph-Based Adversarial Networks for Human Pose Estimation from Sparse IMUsabstractRecently, sparse-inertial human pose estimation (SI-HPE) with only a few IMUs has shown great potential in various fields. The most advanced work in this area achieved fairish results using only six IMUs. However, there are still two major issues that remain to be addressed. First, existing methods typically treat SI-HPE as a temporal sequential learning problem and often ignore the important spatial prior of skeletal topology. Second, there are far more synthetic data in their training data than real data, and the data distribution of synthetic data and real data is quite different, which makes it difficult for the model to be applied to more diverse real data. To address these issues, we propose “Graph-based Adversarial Inertial Poser (GAIP),” which tracks body movements using sparse data from six IMUs. To make full use of the spatial prior, we design a multi-stage pose regressor with graph convolution to explicitly learn the skeletal topology. A joint position loss is also introduced to implicitly mine spatial information. To enhance the generalization ability, we propose supervising the pose regression with an adversarial loss from a discriminator, bringing the ability of adversarial networks to learn implicit constraints into full play. Additionally, we construct a real dataset that includes hip support movements and a synthetic dataset containing various motion categories to enrich the diversity of inertial data for SI-HPE. Extensive experiments demonstrate that GAIP produces results with more precise limb movement amplitudes and relative joint positions, accompanied by smaller joint angle and position errors compared to state-of-the-art counterparts. The datasets and codes are publicly available at https://cslinzhang.github.io/GAIP/ . Kaixin Chen 0003, Lin Zhang 0014, Zhong Wang 0009, Shengjie Zhao 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | DeIL: Direct-and-Inverse CLIP for Open-World Few-Shot LearningabstractOpen-World Few-Shot Learning (OFSL) is a critical field of research, concentrating on the precise identification of target samples in environments with scarce data and unre-liable labels, thus possessing substantial practical signif-icance. Recently, the evolution of foundation models like CLIP has revealed their strong capacity for representation, even in settings with restricted resources and data. This development has led to a significant shift in focus, tran-sitioning from the traditional method of “building models from scratch” to a strategy centered on “efficiently utilizing the capabilities of foundation models to extract rele-vant prior knowledge tailored for OFSL and apply it judi-ciously”. Amidst this backdrop, we unveil the Direct-and-Inverse CLIP (DeIL), an innovative method leveraging our proposed “Direct-and-Inverse” concept to activate CLIP-based methods for addressing OFSL. This concept transforms conventional single-step classification into a nuanced two-stage process: initially filtering out less probable cate-gories, followed by accurately determining the specific cat-egory of samples. DeIL comprises two key components: a pretrainer (frozen) for data denoising, and an adapter (tun-able) for achieving precise final classification. In experiments, DeIL achieves SOTA performance on 11 datasets. https://github.com/The-Shuai/DeIL. Shuai Shao 0006, Yan Wang 0076, Baodi Liu, Yicong Zhou |
CVPR | 5 |
| 2024 | Seam Mask Guided Partial Reconstruction with Quantum-Inspired Local Aggregation For Deep Image StitchingabstractIn image stitching, artifacts caused by misalignment affect the visual quality and the performance of subsequent tasks such as segmentation and detection. This paper proposes SMPR, a reconstruction-based aligned image composition method to minimize artifacts. SMPR fuses images in part of the overlapping areas and reconstructs other portions from single images. Specifically, we propose a seam mask generation method to obtain optimal seam masks that pass through minimal misalignment. During training, we use the seam masks to guide the model in detecting optimal fusion areas. In testing, the model can detect fusion areas without seam masks and reconstruct stitching results. We propose a quantum-inspired local aggregation (QILA) module to improve feature reconstruction performance. We develop an encoder-decoder network with QILA and experiment on a real-world dataset. The experiments show that our method outperforms state-of-the-art methods in both qualitative and quantitative aspects. Chen-Bin Feng, Jiaxue Li, Yicong Zhou |
ICASSP | 4 |
| 2024 | Deep Unfolding 3D Non-Local Transformer Network for Hyperspectral Snapshot Compressive ImagingabstractHyperspectral compressive imaging has shown remarkable advancements through the adoption of deep unfolding frameworks, which integrate the proximal mapping prior into the data fidelity term to formulate the reconstruction problem. However, existing technologies still face challenges in effectively capturing spatial-spectral features during the iterative deep prior learning stage, leading to unsatisfactory performance degradation. To address this issue, we propose a deep unfolding 3D non-local transformer (3DNLT) network for hyperspectral compressive imaging. A learnable half-quadratic splitting (HQS) algorithm is utilized to iteratively update the linear projection. Furthermore, a 3D non-local attention ushaped transformer is presented as the deep proximal mapping prior module to obtain the spatial-spectral long-range dependency features, leading to enhance the network’s ability to capture fine-grained hyperspectral and spatial details. Experimental results on both synthetic and real hyperspectral image reconstruction have demonstrated the superior performance of the 3DNLT network compared to state-of-the-art methods. Yongyong Chen, Bingzhi Chen, Yicong Zhou |
ICME | 6 |
| 2024 | Rethinking The Training And Evaluation of Rich-Context Layout-to-Image GenerationabstractRecent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area within generative modeling is layout-to-image (L2I) generation, where predefined layouts of objects guide the generative process. In this study, we introduce a novel regional cross-attention module tailored to enrich layout-to-image generation. This module notably improves the representation of layout regions, particularly in scenarios where existing methods struggle with highly complex and detailed textual descriptions. Moreover, while current open-vocabulary L2I methods are trained in an open-set setting, their evaluations often occur in closed-set environments. To bridge this gap, we propose two metrics to assess L2I performance in open-vocabulary scenarios. Additionally, we conduct a comprehensive user study to validate the consistency of these metrics with human preferences. Jiaxin Cheng, Tong He 0002, Tianjun Xiao, Zheng Zhang 0001, Yicong Zhou |
NeurIPS | 6 |
| 2024 | Fine-Grained Multimodal DeepFake Classification via Heterogeneous Graphs
Qilin Yin, Wei Lu 0001, Xiaochun Cao, Xiangyang Luo 0001, Yicong Zhou, Jiwu Huang |
Int. J. Comput. Vis. | 5 |
| 2024 | Stacked Graph Fusion Denoising Autoencoder for Hyperspectral Anomaly DetectionabstractAnomaly detection for hyperspectral images (HSIs) is a challenging problem to distinguish a few anomalous pixels from a majority of background pixels. Most existing methods cannot simultaneously explore both structural and spatial information from global and local perspectives. In this letter, we propose a stacked graph fusion denoising autoencoder (SGFDAE) for hyperspectral anomaly detection. Specifically, the global and local graphs are constructed from an HSI to explore potential structural and spatial information. With the designed graph fusion strategy, an advanced graph denoising autoencoder with deep architecture is developed in a hierarchical manner. To achieve better reconstruction and detection, a greedy layerwise unsupervised pretraining strategy is presented for network training. Experiments show that SGFDAE achieves 97.17%, 98.43%, and 98.90% detection accuracies by averaging the results of the datasets from three different scenes and outperforms the state-of-the-art methods. Yongshan Zhang, Yijiang Li, Xinxin Wang 0003, Xinwei Jiang, Yicong Zhou |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Few-shot image classification via hybrid representation
Baodi Liu, Shuai Shao 0006, Lei Xing 0005, Weifeng Liu 0001, Weijia Cao, Yicong Zhou |
Pattern Recognit. | 7 |
| 2024 | Deep Reverse Attack on SIFT Features With a Coarse-to-Fine GAN ModelabstractRecently, it has been shown that adversaries can reconstruct images from SIFT features through reverse attacks. However, the images reconstructed by existing reverse attack methods suffer from information loss and are unable to sufficiently reveal the private contents of the original images. In this paper, a two-stage deep reverse attack model called Coarse-to-Fine Generative Adversarial Network (CFGAN) is proposed to more deeply explore the information in SIFT features and further demonstrate the risk of privacy leakage associated with SIFT features. Specifically, the proposed model consists of two sub-networks, namely coarse net and fine net. The coarse net is developed to restore coarse images using SIFT features, while the fine net is responsible for refining the coarse images to obtain better reconstruction results. To effectively leverage the information contained in SIFT features, an efficient fusion strategy based on the AdaIN operation is designed in the fine net. Additionally, we introduce a new loss function called sift loss that enhances the color fidelity of reconstructed images. Extensive experiments conducted on various datasets verify that the proposed CFGAN performs favorably against state-of-the-art methods. The reconstructed images exhibit better visual quality, less texture distortion, and higher color fidelity. Source code is available at https://github.com/HITLiXincodes/CFGAN. Xin Li 0154, Guopu Zhu, Shen Wang 0004, Yicong Zhou, Xinpeng Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Robust Discriminative t-Linear Subspace Learning for Image Feature ExtractionabstractSubspace learning has been widely applied for joint feature extraction and dimensionality reduction, demonstrating significant efficacy. Numerous subspace learning methods with diverse assumptions regarding the criteria for the target subspaces have been developed to obtain compact and interpretable data representations. However, when applied to image data, existing methods fail to fully exploit the inherent correlations within the image set. This paper proposes a Robust Discriminative t-Linear Subspace Learning model (RDtSL) to tackle this issue using t-product. The model mainly has four strengths: 1) Taking advantage of t-product, RDtSL learns the projection basis directly from the image set while fully exploiting its internal correlations; 2) Based on its energy preservation module, RDtSL retains the primary energy of samples in the learned subspace, maintaining satisfactory performance even with low subspace dimensions; 3) Class-distinctive features are effectively preserved in the learned representations due to the incorporation of the classification module; 4) Relying on its graph embedding module, RDtSL learns an affinity graph of samples adaptively to enrich the data representations with locality and similarity information. The harmonious balance maintained between the three proposed modules helps RDtSL learn discriminative and informative data representations. We also develop an iterative algorithm to solve RDtSL. Extensive experiments on benchmark databases demonstrate the superiority of the proposed model. Kangdao Liu, Xiaolin Xiao, Jinkun You, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Enhanced Pseudo-Label Generation With Self-Supervised Training for Weakly- Supervised Semantic SegmentationabstractDue to the high cost of pixel-level labels required for fully-supervised semantic segmentation, weakly-supervised segmentation has emerged as a more viable option recently. Existing weakly-supervised methods tried to generate pseudo-labels without pixel-level labels for semantic segmentation, but a common problem is that the generated pseudo-labels contain insufficient semantic information, resulting in poor accuracy. To address this challenge, a novel method is proposed, which generates class activation/attention maps (CAMs) containing sufficient semantic information as pseudo-labels for the semantic segmentation training without pixel-level labels. In this method, the attention-transfer module is designed to preserve salient regions on CAMs while avoiding the suppression of inconspicuous regions of the targets, which results in the generation of pseudo-labels with sufficient semantic information. A pixel relevance focused-unfocused module has also been developed for better integrating contextual information, with both attention mechanisms employed to extract focused relevant pixels and multi-scale atrous convolution employed to expand receptive field for establishing distant pixel connections. The proposed method has been experimentally demonstrated to achieve competitive performance in weakly-supervised segmentation, and even outperforms many saliency-joined methods. Zhen Qin 0002, Guosong Zhu, Erqiang Zhou, Yingjie Zhou 0001, Yicong Zhou, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Global Localization in Large-Scale Point Clouds via Roll-Pitch-Yaw Invariant Place Recognition and Low-Overlap Global RegistrationabstractFor autonomous ground vehicles, global localization with 3D LiDAR is an indispensable part of tasks such as navigation. Usually, global localization using LiDAR is subdivided into two sub-problems, place recognition and global registration. For place recognition, the recent emerging schemes based on deep learning either rely on 3D convolution with high complexity or need to learn features from various forward perspectives. To mitigate this, we propose a model with roll-pitch-yaw invariance that represents point clouds as probabilistic voxels and generates occupancy grids from a bird’s-eye view, fulfilling robust place recognition by learning aggregated embeddings from a fixed perspective. For low-overlap global registration, the traditional handcraft feature-based methods are mostly limited to dense object-level point clouds, while the state-of-the-art learning-based approaches often rely on complex 3D convolution and additional feature association learning. To fill this gap to some extent, we propose to estimate the relative roll-pitch angles and vertical translation by fitting and aligning the ground plane of the point clouds and to determine the horizontal translations and yaw angle by matching their projected occupancy grids. Extensive experiments corroborate the superior recall and generalization ability of our place recognition model, as well as the advanced success rate and accuracy of our 3D registration approach. Especially in the recognition and registration of hard samples, our results far exceed those of our counterparts by large margins. To ensure full reproducibility, the relevant codes and data are made available online. Zhong Wang 0009, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Ct-LVI: A Framework Toward Continuous-Time Laser-Visual-Inertial Odometry and MappingabstractOwing to the inherent complementarity among LiDAR, camera, and IMU, a growing effort has been paid to laser-visual-inertial SLAM recently. The existing approaches, however, are limited in two aspects. First, at the front-end, they usually employ a discrete-time representation that requires high-precision hardware/software synchronization and are based on geometric laser features, leading to low robustness and scalability. Second, at the backend, visual loop constraints suffer from scale ambiguity and the sparseness of the point cloud deteriorates the scan-to-scan loop detection. To solve these problems, for the front-end, we propose a continuous-time laser-visual-inertial odometry which formulates the carrier trajectory in continuous time, organizes point clouds in probabilistic submaps, and jointly optimizes the loss terms of laser anchors, visual reprojections, and IMU readings, achieving accurate pose estimation even with fast motion or in unstructured scenes where it is difficult to extract meaningful geometric features. At the backend, we propose building 5-DoF laser constraints by matching projected 2D submaps and 6-DoF visual constraints via laser-aided visual relocalization, ensuring mapping consistency in large-scale scenes. Results show that our framework achieves high-precision estimation and is more robust than its counterparts when the carrier works in large scenes or with fast motion. The relevant codes and data are open-sourced at https://cslinzhang.github.io/Ct-LVI/Ct-LVI.html. Zhong Wang 0009, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Tensorial Global-Local Graph Self-Representation for Hyperspectral Band SelectionabstractBand selection aims at selecting a subset of representative bands from original hyperspectral images (HSIs) to alleviate data redundancy. There are at least two issues existing in previous methods. First, most of them ignore global or local structural information without considering both two aspects. Second, the high-order correlations among spectral bands are not explored during learning. In this paper, we propose a tensorial global-local graph self-representation (TGSR) method for hyperspectral band selection. Specifically, we segment the HSI into diverse superpixels to show the inherent spectral-spatial structures. Based on the generated superpixels, we learn the global and local graphs to explore complex structural information from global pixels and local regions. To alleviate the computational burden, a transformation is designed for easy graph convolution of global graph and pixel spectral matrix. With global and local knowledge, we formulate a global-local graph self-representation model to conduct band correlation learning in a self-weighted manner. To explore the high-order correlations among bands, we reorganize the self-representation coefficient matrices into a tensor with low-rank constraint. We design an alternating optimization algorithm to solve the proposed model. The most representative band is selected from each band subset by performing spectral clustering on the constructed affinity matrix. Experiments on HSI datasets verify the effectiveness of our method over the state-of-the-art methods. The source code is released athttps://github.com/ZhangYongshan/TGSR. Yongshan Zhang, Jianwen Qi, Xinxin Wang 0003, Zhihua Cai, Jiangtao Peng, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Self-Completed Bipartite Graph Learning for Fast Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC), excavating diversity and consistency from multiple incomplete views, has aroused widespread research enthusiasm. Nevertheless, most existing methods still encounter the following issues: 1) they generally concentrate on pair-wise instance correlation, which consumes at least a quadratic complexity and precludes them from applying at large scales; 2) they only concentrate on pair-wise instance relevance, whereas ignoring the discriminative correlation hidden across views. To overcome these drawbacks, we propose the Self-Completed Bipartite Graph Learning (SCBGL) method for fast IMVC, which adaptively learns a self-completed consensus bipartite graph with the guidance of global information. Specifically, SCBGL learns the consensus anchor matrix shared among diverse views and further constructs a consensus intra-view bipartite graph with missing instances to explore the diversity and complementarity underlying different views. Meanwhile, we concatenate all the multiple features with projection learning to learn global anchors that would be employed to construct an inter-view bipartite graph. Furthermore, SCBGL dexterously utilizes the abundant inter-view information to tutor the self-completion of the consensus intra-view bipartite graph. By devising an alternatively iterative strategy, we present an efficient algorithm, which enjoys a linear time complexity, to solve the proposed SCBGL model. Numerous experiments conducted on large-scale datasets substantiate the superior performance of the SCBGL beyond the state-of-the-arts. Xiaojia Zhao, Qiangqiang Shen, Yongyong Chen, Yongsheng Liang 0001, Junxin Chen 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Impulse Noise Image Restoration Using Nonconvex Variational Model and Difference of Convex Functions AlgorithmabstractIn this article, the problem of impulse noise image restoration is investigated. A typical way to eliminate impulse noise is to use an$L_{1}$norm data fitting term and a total variation (TV) regularization. However, a convex optimization method designed in this way always yields staircase artifacts. In addition, the$L_{1}$norm fitting term tends to penalize corrupted and noise-free data equally, and is not robust to impulse noise. In order to seek a solution of high recovery quality, we propose a new variational model that integrates the nonconvex data fitting term and the nonconvex TV regularization. The usage of the nonconvex TV regularizer helps to eliminate the staircase artifacts. Moreover, the nonconvex fidelity term can detect impulse noise effectively in the way that it is enforced when the observed data is slightly corrupted, while is less enforced for the severely corrupted pixels. A novel difference of convex functions algorithm is also developed to solve the variational model. Using the variational method, we prove that the sequence generated by the proposed algorithm converges to a stationary point of the nonconvex objective function. Experimental results show that our proposed algorithm is efficient and compares favorably with state-of-the-art methods. Benxin Zhang, Guopu Zhu, Hongli Zhang 0001, Yicong Zhou, Sam Kwong |
IEEE Trans. Cybern. | 5 |
| 2024 | Automatic Quaternion-Domain Color Image StitchingabstractTaking advantages of the quaternion representation of the color image, this paper proposes a quaternion perceptual seamline detection model to generate the seamline in the quaternion domain. It considers seamline detection as a quaternion-domain color image labeling problem and minimizes the local-area quaternion perceptual difference cost to obtain the optimal seamline. To assess seamline quality effectively, we develop a quaternion perceptual seamline quality measure. Based on the proposed quaternion perceptual seamline detection model and quality measure, we further propose a general framework for automatic quaternion-domain color image stitching (AQCIS). To the best of our knowledge, this is the first attempt to perform color image stitching completely in the quaternion domain. Meanwhile, AQCIS introduces the joint optimization strategy of local alignment and seamline in an iterative fashion. Extensive experiments on challenging datasets demonstrate that our AQCIS achieves superior performance for color image stitching in comparison with state-of-the-art methods. Jiaxue Li, Yicong Zhou |
IEEE Trans. Image Process. | 2 |
| 2024 | Lightweight Context-Aware Network Using Partial-Channel Transformation for Real-Time Semantic SegmentationabstractOptimizing the computational efficiency of the artificial neural networks is crucial for resource-constrained platforms like autonomous driving systems. To address this challenge, we proposed a Lightweight Context-aware Network (LCNet) that accelerates semantic segmentation while maintaining a favorable trade-off between inference speed and segmentation accuracy in this paper. The proposed LCNet introduces a partial-channel transformation (PCT) strategy to minimize computing latency and hardware requirements of the basic unit. Within the PCT block, a three-branch context aggregation (TCA) module expands the feature receptive fields, capturing multiscale contextual information. Additionally, a dual-attention-guided decoder (DD) recovers spatial details and enhances pixel prediction accuracy. Extensive experiments on three benchmarks demonstrate the effectiveness and efficiency of the proposed LCNet model. Remarkably, a smaller model LCNet$_{3\_7}$achieves 73.8% mIoU with only 0.51 million parameters, with an impressive inference speed of$\sim$142.5 fps and$\sim$9 fps using a single RTX 3090 GPU and Jetson Xavier NX, respectively, on the Cityscapes test set at$1024\times 1024$resolution. A more accurate version of the LCNet$_{3\_11}$can achieve 75.8% mIoU with 0.74 million parameters at$\sim$117 fps inference speed on Cityscapes at the same resolution. Much faster inference speed can be achieved at smaller image resolutions. LCNet strikes a great balance between computational efficiency and prediction capability for mobile application scenarios. The code is available at https://github.com/lztjy/LCNet. Shaowen Lin, Qingming Yi, Jian Weng 0001, Aiwen Luo, Yicong Zhou |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Two-Stage Watermark Removal Framework for Spread Spectrum WatermarkingabstractSpread spectrum (SS) watermarking has gained significant attention as it prevents attackers from reading, tampering with, or removing watermarks. Secret key estimation can help with the first two unauthorized operations but cannot remove watermarks. Moreover, existing deep-learning watermark removal methods do not consider the characteristics of SS watermarking, thus leading to unsatisfactory results. In this paper, we design a secret key estimation method that treats secret key estimation as a binary classification problem and updates the estimated key via backpropagation and parameter optimization algorithms. We develop a watermark removal network using quaternion convolutional neural networks (QCNNs) to learn watermark features while capturing the relationship between channels to improve image quality. Based on our estimation method and QCNN-based network, we propose a two-stage watermark removal framework that utilizes information of the secret key to train the network. A loss function is introduced to directly prevent watermark extraction, thereby improving removal performance. Extensive experiments demonstrate the superiority of our methods over the state-of-the-art methods. Jinkun You, Yicong Zhou |
IEEE Trans. Multim. | 2 |
| 2024 | Exploiting Substitution Box for Cryptanalyzing Image Encryption Schemes With DNA Coding and Nonlinear DynamicsabstractIn recent years, a number of image encryption schemes based on DNA coding and nonlinear dynamics have been proposed. Generally, these DNA-based schemes first encode plaintext images into DNA sequences and then encrypt them with pseudorandom elements produced by chaotic systems or other nonlinear dynamics. Although ciphertexts can pass some security tests, many image encryption schemes are being shown to have intrinsic flaws and that they cannot guarantee a high level of security. In this article, we cryptanalyze a family of image encryption schemes for which the encryption kernel is DNA coding or its variant. The complex DNA operation can be simplified as a substitution box (S-box). The whole cryptosystem's security level is thus significantly decreased and is vulnerable to the chosen-plaintext attack. Applications of this concept to break five ciphers are theoretically presented and experimentally verified. In addition, some suggestions for resisting similar attacks are also given in this article. Chengrui Zhang, Junxin Chen 0001, Dongming Chen, Wei Wang 0077, Yushu Zhang 0001, Yicong Zhou |
IEEE Trans. Multim. | 6 |
| 2024 | Bipartite Graph-Based Projected Clustering With Local Region Guidance for Hyperspectral ImageryabstractHyperspectral image (HSI) clustering is challenging to divide all pixels into different clusters because of the absent labels, large spectral variability and complex spatial distribution. Anchor strategy provides an attractive solution to the computational bottleneck of graph-based clustering for large HSIs. However, most existing methods require separated learning procedures and ignore noisy as well as spatial information. In this paper, we propose a bipartite graph-based projected clustering (BGPC) method with local region guidance for HSI data. To take full advantage of spatial information, HSI denoising to alleviate noise interference and anchor initialization to construct bipartite graph are conducted within each generated superpixel. With the denoised pixels and initial anchors, projection learning and structured bipartite graph learning are simultaneously performed in a one-step learning model with connectivity constraint to directly provide clustering results. An alternating optimization algorithm is devised to solve the formulated model. The advantage of BGPC is the joint learning of projection and bipartite graph with local region guidance to exploit spatial information and linear time complexity to lessen computational burden. Extensive experiments demonstrate the superiority of the proposed BGPC over the state-of-the-art HSI clustering methods. Yongshan Zhang, Guozhu Jiang, Zhihua Cai, Yicong Zhou |
IEEE Trans. Multim. | 4 |
| 2024 | Gradient Learning With the Mode-Induced Loss: Consistency Analysis and ApplicationsabstractVariable selection methods aim to select the key covariates related to the response variable for learning problems with high-dimensional data. Typical methods of variable selection are formulated in terms of sparse mean regression with a parametric hypothesis class, such as linear functions or additive functions. Despite rapid progress, the existing methods depend heavily on the chosen parametric function class and are incapable of handling variable selection for problems where the data noise is heavy-tailed or skewed. To circumvent these drawbacks, we propose sparse gradient learning with the mode-induced loss (SGLML) for robust model-free (MF) variable selection. The theoretical analysis is established for SGLML on the upper bound of excess risk and the consistency of variable selection, which guarantees its ability for gradient estimation from the lens of gradient risk and informative variable identification under mild conditions. Experimental analysis on the simulated and real data demonstrates the competitive performance of our method over the previous gradient learning (GL) methods. Hong Chen 0004, Youcheng Fu, Weifu Li, Yicong Zhou, Feng Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Tensor Learning Meets Dynamic Anchor Learning: From Complete to Incomplete Multiview ClusteringabstractMultiview clustering (MVC), which can dexterously uncover the underlying intrinsic clustering structures of the data, has been particularly attractive in recent years. However, previous methods are designed for either complete or incomplete multiview only, without a unified framework that handles both tasks simultaneously. To address this issue, we propose a unified framework to efficiently tackle both tasks in approximately linear complexity, which integrates tensor learning to explore the inter-view low-rankness and dynamic anchor learning to explore the intra-view low-rankness for scalable clustering (TDASC). Specifically, TDASC efficiently learns smaller view-specific graphs by anchor learning, which not only explores the diversity embedded in multiview data, but also yields approximately linear complexity. Meanwhile, unlike most current approaches that only focus on pair-wise relationships, the proposed TDASC incorporates multiple graphs into an inter-view low-rank tensor, which elegantly models the high-order correlations across views and further guides the anchor learning. Extensive experiments on both complete and incomplete multiview datasets clearly demonstrate the effectiveness and efficiency of TDASC compared with several state-of-the-art techniques. Yongyong Chen, Xiaojia Zhao, Zheng Zhang 0006, Youfa Liu, Jingyong Su, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Detection of Deepfake Videos Using Long-Distance AttentionabstractWith the rapid progress of deepfake techniques in recent years, facial video forgery can generate highly deceptive video content and bring severe security threats. And detection of such forgery videos is much more urgent and challenging. Most existing detection methods treat the problem as a vanilla binary classification problem. In this article, the problem is treated as a special fine-grained classification problem since the differences between fake and real faces are very subtle. It is observed that most existing face forgery methods left some common artifacts in the spatial domain and time domain, including generative defects in the spatial domain and interframe inconsistencies in the time domain. And a spatial-temporal model is proposed which has two components for capturing spatial and temporal forgery traces from a global perspective, respectively. The two components are designed using a novel long-distance attention mechanism. One component of the spatial domain is used to capture artifacts in a single frame, and the other component of the time domain is used to capture artifacts in consecutive frames. They generate attention maps in the form of patches. The attention method has a broader vision which contributes to better assembling global information and extracting local statistic information. Finally, the attention maps are used to guide the network to focus on pivotal parts of the face, just like other fine-grained classification methods. The experimental results on different public datasets demonstrate that the proposed method achieves state-of-the-art performance, and the proposed long-distance attention method can effectively capture pivotal parts for face forgery. Wei Lu 0001, Lingyi Liu, Xianfeng Zhao, Yicong Zhou, Jiwu Huang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Efficient Harmonic Neural Networks With Compound Discrete Cosine Transform Filters and Shared Reconstruction FiltersabstractThe harmonic neural network (HNN) learns a combination of discrete cosine transform (DCT) filters to obtain an integrated feature from all spectra in the frequency domain. HNN, however, faces two challenges in learning and inference processes. First, the spectrum feature learned by HNN is insufficient and limited because the number of DCT filters is much smaller than that of feature maps. In addition, the number of parameters and the computation costs of HNN are significantly high because the intermediate spectrum layers are expanded multiple times. These two challenges will severely harm the performance and efficiency of HNN. To solve these problems, we first propose the compound DCT (C-DCT) filters integrating the nearest DCT filters to retrieve rich spectrum features to improve the performance. To significantly reduce the model size and computation complexity for improving the efficiency, the shared reconstruction filter is then proposed to share and dynamically drop the meta-filters in every frequency branch. Integrating the C-DCT filters with the shared reconstruction filters, the efficient harmonic network (EH-Net) is introduced. Extensive experiments on different datasets demonstrate that the proposed EH-Nets can effectively reduce the model size and computation complexity while maintaining the model performance. The code has been released at https://github.com/zhangle408/EH-Nets. Yao Lu 0008, Le Zhang 0016, Xiaofei Yang 0002, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Consistent Arbitrary Style Transfer Using Consistency Training and Self-Attention ModuleabstractArbitrary style transfer (AST) has garnered considerable attention for its ability to transfer styles infinitely. Although existing methods have achieved impressive results, they may overlook style consistencies and fail to capture crucial style patterns, leading to inconsistent style transfer (ST) caused by minor disturbances. To tackle this issue, we conduct a mathematical analysis of inconsistent ST and develop a style inconsistency measure (SIM) to quantify the inconsistencies between generated images. Moreover, we propose a consistent AST (CAST) framework that effectively captures and transfers essential style features into content images. The proposed CAST framework incorporates an intersection-of-union-preserving crop (IoUPC) module to obtain style pairs with minor disturbance, a self-attention (SA) module to learn the crucial style features, and a style inconsistency loss regularization (SILR) to facilitate consistent feature learning for consistent stylization. Our proposed framework not only provides an optimal solution for consistent ST but also outperforms existing methods when embedded into the CAST framework. Extensive experiments demonstrate that the proposed CAST framework can effectively transfer style patterns while preserving consistency and achieve the state-of-the-art performance. Yue Wu 0001, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | An Underwater Organism Image Dataset and a Lightweight Module Designed for Object Detection NetworksabstractLong-term monitoring and recognition of underwater organism objects are of great significance in marine ecology, fisheries science and many other disciplines. Traditional techniques in this field, including manual fishing-based ones and sonar-based ones, are usually flawed. Specifically, the method based on manual fishing is time-consuming and unsuitable for scientific researches, while the sonar-based one, has the defects of low acoustic image accuracy and large echo errors. In recent years, the rapid development of deep learning and its excellent performance in computer vision tasks make vision-based solutions feasible. However, the researches in this area are still relatively insufficient in mainly two aspects. First, to our knowledge, there is still a lack of large-scale datasets of underwater organism images with accurate annotations. Second, in consideration of the limitation on hardware resources of underwater devices, an underwater organism detection algorithm that is both accurate and lightweight enough to be able to infer in real time is still lacking. As an attempt to fill in the aforementioned research gaps to some extent, we established the Multiple Kinds of Underwater Organisms (MKUO) dataset with accurate bounding box annotations of taxonomic information, which consists of 10,043 annotated images, covering eighty-four underwater organism categories. Based on our benchmark dataset, we evaluated a series of existing object detection algorithms to obtain their accuracy and complexity indicators as the baseline for future reference. In addition, we also propose a novel lightweight module, namely Sparse Ghost Module, designed especially for object detection networks. By substituting the standard convolution with our proposed one, the network complexity can be significantly reduced and the inference speed can be greatly improved without obvious detection accuracy loss. To make our results reproducible, the dataset and the source code are available online at https://cslinzhang.github.io/MKUO-and-Sparse-Ghost-Module/ . Jiafeng Huang, Tianjun Zhang, Shengjie Zhao 0001, Lin Zhang 0014, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Boosting Diversity in Visual Search with Pareto Non-Dominated Re-RankingabstractThe field of visual search has gained significant attention recently, particularly in the context of web search engines and e-commerce product search platforms. However, the abundance of web images presents a challenge for modern image retrieval systems, as they need to find both relevant and diverse images that maximize users’ satisfaction. In response to this challenge, we propose a non-dominated visual diversity re-ranking (NDVDR) method based on the concept of Pareto optimality. To begin with, we employ a fast binary hashing method as a coarse-grained retrieval procedure. This allows us to efficiently obtain a subset of candidate images for subsequent re-ranking. Fed with this initial retrieved image results, the NDVDR performs a fine-grained re-ranking procedure for boosting both relevance and visual diversity among the top-ranked images. Recognizing the inherent conflict nature between the objectives of relevance and diversity, the re-ranking procedure is simulated as the analytical stage of a multi-criteria decision-making process, seeking the optimal tradeoff between the two conflicting objectives within the initial retrieved images. In particular, a non-dominated sorting mechanism is devised that produces Pareto non-dominated hierarchies among images based on the Pareto dominance relation. Additionally, two novel measures are introduced for the effective characterization of the relevance and diversity scores among different images. We conduct experiments on three popular real-world image datasets and compare our re-ranking method with several state-of-the-art image search re-ranking methods. The experimental results validate that our re-ranking approach guarantees retrieval accuracy while simultaneously boosting diversity among the top-ranked images. Si-chao Lei, Yue-Jiao Gong, Xiaolin Xiao, Yicong Zhou, Jun Zhang 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Tensorial Evolutionary Optimization for Natural Image MattingabstractNatural image matting has garnered increasing attention in various computer vision applications. The matting problem aims to find the optimal foreground/background (F/B) color pair for each unknown pixel and thus obtain an alpha matte indicating the opacity of the foreground object. This problem is typically modeled as a large-scale pixel pair combinatorial optimization (PPCO) problem. Heuristic optimization is widely employed to tackle the PPCO problem owing to its gradient-free property and promising search ability. However, traditional heuristic methods often encode F/B solutions to a one-dimensional (1D) representation and then evolve the solutions in a 1D manner. This 1D representation destroys the intrinsic two-dimensional (2D) structure of images, where the significant spatial correlations among pixels are ignored. Moreover, the 1D representation also brings operation inefficiency. To address the above issues, this article develops a spatial-aware tensorial evolutionary image matting (TEIM) method. Specifically, the matting problem is modeled as a 2D Spatial-PPCO (S-PPCO) problem, and a global tensorial evolutionary optimizer is proposed to tackle the S-PPCO problem. The entire population is represented as a whole by a third-order tensor, in which individuals are classified into two types: F and B individuals for denoting the 2D F/B solutions, respectively. The evolution process, consisting of three tensorial evolutionary operators, is implemented based on pure tensor computation for efficiently seeking F/B solutions. The local spatial smoothness of images is also integrated into the evaluation process for obtaining a high-quality alpha matte. Experimental results compared with state-of-the-art methods validate the effectiveness of TEIM. Si-chao Lei, Yue-Jiao Gong, Xiaolin Xiao, Yicong Zhou, Jun Zhang 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Bi-directional Block Encoding for Reversible Data Hiding over Encrypted ImagesabstractReversible data hiding over encrypted images (RDH-EI) technology is a viable solution for privacy-preserving cloud storage, as it enables the reversible embedding of additional data into images while maintaining image confidentiality. Since the data hiders, e.g., cloud servers, are willing to embed as much data as possible for storage, management, or other processing purposes, a large embedding capacity is desirable in an RDH-EI scheme. In this article, we introduce a novel bi-directional block encoding (BDBE) method, which, for the first time, encodes the distances of values in a binary sequence from both ends. This approach allows for encoding images with smaller sizes compared to traditional and state-of-the-art encoding methods. Leveraging the BDBE technique, we propose a high-capacity RDH-EI scheme. In this scheme, the content owner initially predicts the image pixels and then employs BDBE to encode the prediction errors, creating space for data embedding. The resulting encoded data are subsequently encrypted using a secure stream cipher, such as the Advanced Encryption Standard, before being transmitted to a data hider. The data hider can embed confidential information within the encrypted image for the purposes of storage, management, or other processing. Upon receiving the data, an authorized receiver can accurately recover the original image and the embedded data without any loss. Experimental results demonstrate that our RDH-EI scheme achieves a significantly larger embedding capacity compared to several state-of-the-art schemes. Zhongyun Hua, Yushu Zhang 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | I2P Registration by Learning the Underlying Alignment Feature Space from Pixel-to-Point SimilaritiesabstractEstimating the relative pose between a camera and a LiDAR holds paramount importance in facilitating complex task execution within multi-agent systems. Nonetheless, current methodologies encounter two primary limitations. First, amid the cross-modal feature extraction, they typically employ separate modal branches to extract cross-modal features from images and point clouds. This approach results in the feature spaces of images and point clouds being misaligned, thereby reducing the robustness of establishing correspondences. Second, due to the scale differences between images and point clouds, one-to-many pixel-point correspondences are inevitably encountered, which will mislead the pose optimization. To address these challenges, we propose a framework named I mage-to- P oint cloud registration by learning the underlying alignment feature space from P ixel-to- P oint SIM imilarities (I2P \({}_{\mathbf{ppsim}}\) ) . Central to \(\text{I2P}_{\text{ppsim}}\) is a Shared Feature Alignment Module (SFAM). It is designed under on a coarse-to-fine architecture and uses a weight-sharing network to construct an alignment feature space. Benefiting from SFAM, \(\text{I2P}_{\text{ppsim}}\) can effectively identify the co-view regions between images and point clouds and establish high-reliability 2D-3D correspondences. Moreover, to mitigate the one-to-many correspondence issue, we introduce a similarity maximization strategy termed point-max. This strategy effectively filters out outliers, thereby establishing accurate 2D-3D correspondences. To evaluate the efficacy of our framework, we conduct extensive experiments on KITTI Odometry and Oxford Robotcar. The results corroborate the effectiveness of our framework in improving image-to-point cloud registration. To make our results reproducible, the source codes have been released at https://cslinzhang.github.io/I2P Yunda Sun, Lin Zhang 0014, Zhong Wang 0009, Yang Chen 0037, Shengjie Zhao 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Uformer-ICS: A U-Shaped Transformer for Image Compressive Sensing ServiceabstractMany service computing applications require real-time dataset collection from multiple devices, necessitating efficient sampling techniques to reduce bandwidth and storage pressure. Compressive sensing (CS) has found wide-ranging applications in image acquisition and reconstruction. Recently, numerous deep-learning methods have been introduced for CS tasks. However, the accurate reconstruction of images from measurements remains a significant challenge, especially at low sampling rates. In this paper, we propose Uformer-ICS as a novel U-shaped transformer for image CS tasks by introducing inner characteristics of CS into transformer architecture. To utilize the uneven sparsity distribution of image blocks, we design an adaptive sampling architecture that allocates measurement resources based on the estimated block sparsity, allowing the compressed results to retain maximum information from the original image. Additionally, we introduce a multi-channel projection (MCP) module inspired by traditional CS optimization methods. By integrating the MCP module into the transformer blocks, we construct projection-based transformer blocks, and then form a symmetrical reconstruction model using these blocks and residual convolutional blocks. Therefore, our reconstruction model can simultaneously utilize the local features and long-range dependencies of image, and the prior projection knowledge of CS theory. Experimental results demonstrate its significantly better reconstruction performance than state-of-the-art deep learning-based CS methods. Zhongyun Hua, Yuanman Li, Yushu Zhang 0001, Yicong Zhou |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Quantum-Inspired Spectral-Spatial Pyramid Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification aims at assigning a unique label for every pixel to identify categories of different land covers. Existing deep learning models for HSIs are usually performed in a traditional learning paradigm. Being emerging machines, quantum computers are limited in the noisy intermediate-scale quantum (NISQ) era. The quantum theory offers a new paradigm for designing deep learning models. Motivated by the quantum circuit (QC) model, we propose a quantum-inspired spectral-spatial network (QSSN) for HSI feature extraction. The proposed QSSN consists of a phase-prediction module (PPM) and a measurement-like fusion module (MFM) inspired from quantum theory to dynamically fuse spectral and spatial information. Specifically, QSSN uses a quantum representation to represent an HSI cuboid and extracts joint spectral-spatial features using MFM. An HSI cuboid and its phases predicted by PPM are used in the quantum representation. Using QSSN as the building block, we further propose an end-to-end quantum-inspired spectral-spatial pyramid network (QSSPN) for HSI feature extraction and classification. In this pyramid framework, QSSPN progressively learns feature representations by cascading QSSN blocks and performs classification with a softmax classifier. It is the first attempt to introduce quantum theory in HSI processing model design. Substantial experiments are conducted on three HSI datasets to verify the superiority of the proposed QSSPN framework over the state-of-the-art methods. Yongshan Zhang, Yicong Zhou |
CVPR | 3 |
| 2023 | Strip-Cutmix for Person Re-IdentificationabstractPerson re-identification is a very challenging image retrieval task that aims to match the specific person images from different camera views. Person re-identification model requires a large amount of training data to improve its generalization ability, however the current datasets of person re-identification are not enough that tend to make the model overfit. Therefore, some data augmentation methods are used to increase the amount of training data to improve the generalization ability of the model. Cutmix is a common data augmentation method in the field of deep learning, but it is rarely used in person re-identification task because the triple loss cannot handle the decimal similarity label generated by cutmix. In order to put the cutmix method for data augmentation in person re-identification, we extend the triplet loss that is commonly used in person re-identification to a form which can handle decimal similarity label from the perspective of optimizing image similarity. In addition, we propose Strip-Cutmix data augmentation method, which is more suitable for person re-identification, and discuss the strategies about using Strip-Cutmix in the field of person re-identification. Extensive experiments show that our approach can prevent model overfit and achieve impressive performance on DukeMTMC-ReID, Market-1501 and MSMT17 benchmark datasets. Ke Qi, Yicong Zhou, Yutao Qi |
IJCNN | 3 |
| 2023 | Learning full context feature for human motion prediction
Huiqin Xing, Yicong Zhou, Jianyu Yang 0002, Yang Xiao 0007 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Semi-Fragile Reversible Watermarking for 3D Models Using Spherical Crown Volume DivisionabstractAiming at the large distortion and low tampering localization accuracy of the existing semi-fragile reversible watermarking for 3D mesh models, a novel semi-fragile reversible watermarking for 3D models using spherical crown volume division is proposed. The crown volume of a sphere is divided to reduce the embedding distortion. The possible geometric and topological transformations are separately considered in the watermark generation, and the vertices of the one-ring neighbourhood are grouped to improve the tampering localization accuracy. Experimental results show that the proposed scheme can achieve better localization accuracy and lower embedding distortion than some state-of-the-art algorithms. It has good potential for the applications in integrity authentication for 3D mesh models. Fei Peng 0001, Tongxin Liao, Min Long 0003, Jin Li 0002, Wensheng Zhang 0002, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Attention-Based Multi-View Feature Collaboration for Decoupled Few-Shot LearningabstractDecoupled Few-shot learning (FSL) is an effective methodology that deals with the problem of data-scarce. Its standard paradigm includes two phases: (1) Pre-train. Generating a CNN-based feature extraction model (FEM) via base data. (2) Meta-test. Employing the frozen FEM to obtain the novel data features, then classifying them. Obviously, one crucial factor, the category gap, prevents the development of FSL, i.e., it is challenging for the pre-trained FEM to adapt to the novel class flawlessly. Inspired by a common-sense theory: the FEMs based on different strategies focus on different priorities, we attempt to address this problem from the multi-view feature collaboration (MVFC) perspective. Specifically, we first denoise the multi-view features by subspace learning method, then design three attention blocks (loss-attention block, self-attention block and graph-attention block) to balance the representation between different views. The proposed method is evaluated on four benchmark datasets and achieves significant improvements of 0.9%-5.6% compared with SOTAs. Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Deep and Low-Rank Quaternion Priors for Color Image ProcessingabstractDue to the physical nature of color images, color image processing such as denoising and inpainting has shown extensive and versatile possibilities over grayscale image processing. The monochromatic and the concatenation model have been widely used to process color images by processing each color channel independently or concatenating three color channels as one unified one and then used existing grayscale image processing methods directly without specific operations. These above schemes, however, have some limitations: (1) they would destroy the inherent correlation among three color channels since they cannot represent color images holistically; (2) they usually focus on one specific handcrafted prior such as smoothness, low-rankness, or even deep prior and thus failing to fuse deep and handcrafted priors of color images flexibly. To conquer these limitations, we propose one unified model to integrate deep prior and low-rank quaternion prior (DLRQP) for color image processing under the plug-and-play (PnP) framework. Specifically, the quaternion representation with low-rank constraint is introduced to denote the color image in a holistic way and one advanced denoiser is adopted to explore the deep prior in an iterative process. To tightly approximate the quaternion rank, one nonconvex penalty function is further utilized. We derive an alternate iterative approach to tackle the proposed model. We empirically demonstrate that our model can achieve superior performance over existing methods on both color image denoising and inpainting tasks. Xiaoyu Kong, Qiangqiang Shen, Yongyong Chen, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | QTN: Quaternion Transformer Network for Hyperspectral Image ClassificationabstractNumerous state-of-the-art transformer-based techniques with self-attention mechanisms have recently been demonstrated to be quite effective in the classification of hyperspectral images (HSIs). However, traditional transformer-based methods severely suffer from the following problems when processing HSIs with three dimensions: (1) processing the HSIs using 1D sequences misses the 3D structure information; (2) too expensive numerous parameters for hyperspectral image classification tasks; (3) only capturing spatial information while lacking the spectral information. To solve these problems, we propose a novel Quaternion Transformer Network (QTN) for recovering self-adaptive and long-range correlations in HSIs. Specially, we first develop a band adaptive selection module (BASM) for producing Quaternion data from HSIs. And then, we propose a new and novel quaternion self-attention (QSA) mechanism to capture the local and global representations. Finally, we propose a new and novel transformer method, i.e., QTN by stacking a series of QSA for hyperspectral classification. The proposed QTN could exploit computation using Quaternion algebra in hypercomplex spaces. Extensive experiments on three public datasets demonstrate that the QTN outperforms the state-of-the-art vision transformers and convolution neural networks. Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Anti-Rounding Image Steganography With Separable Fine-Tuned NetworkabstractImage steganographic methods based on encoder-decoder model with end-to-end network architecture recently have been proposed. However, in steganographic applications, the feature map (called stego matrix) generated by the encoder needs to be rounded as a real stego image for the receiver. The loss of precision by rounding stego matrix leads to the decline in the accuracy of extracted secret messages. The challenge of using end-to-end network to preserve robustness against rounding operation is that it is non-differentiable. In this paper, we propose an anti-rounding image steganography method with separable fine-tuning network architecture which includes the joint training stage (JT-stage) and the separable fine-tuning stage (SF-stage). Firstly, in JT-stage, an embedded generator and a stego matrix extractor are jointly learned without rounding operation. Utilizing concatenation in embedded generator can realistically fuse cover image and secret messages. And the multi-scale fusion block and residual dense block in stego matrix extractor can make secret messages more correctly decoded. Moreover, the discriminator is constructed by generative adversarial nets (GAN) in JT-stage to effectively improve the authenticity and steganalysis security. Then, in SF-stage, the embedded generator is frozen, and the stego matrix is obtained and rounded as a stego image. A stego image extractor is constructed by fine-tuning the layers of the stego matrix extractor to improve the accuracy of message extraction. As the loss will not backpropagate in the embedded generator, the non-differentiability of rounding operation can be offset. Experiments show that the proposed separation fine-tuning network is robust to rounding operation, and effectively reduces the degradation of the image quality and steganalysis performance. Xiaolin Yin, Shaowu Wu, Wei Lu 0001, Yicong Zhou, Jiwu Huang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | CNN-Transformer Based Generative Adversarial Network for Copy-Move Source/ Target DistinguishmentabstractCopy-move forgery can be used for hiding certain objects or duplicating meaningful objects in images. Although copy-move forgery detection has been studied extensively in recent years, it is still a challenging task to distinguish between the source and the target regions in copy-move forgery images. In this paper, a convolutional neural network-transformer based generative adversarial network (CNN-T GAN) is proposed to distinguish the source and target regions in a copy-move forged image. A generator is first utilized to generate a mask that is similar to the groundtruth mask. Then, a discriminator is trained to discriminate the true image pairs from the false ones. When the discriminator cannot discriminate the true/false image pairs accurately, the generator can be used to obtain the final localization maps of copy-move forgery. In the generator, convolutional neural network (CNN) and transformer are exploited to extract the local features and global representations in copy-move forgery images, respectively. In addition, feature coupling layers are designed to integrate the features in CNN branch and transformer branch in an interactive way. Finally, a new Pearson correlation layer is introduced to match the similarity features in source and target regions, which can improve the performance of copy-move forgery localization, especially the localization performance on source regions. To the best of our knowledge, this is the first work to utilize transformer for feature extraction in copy-move forgery localization. The proposed method can not only detect the copy-move regions, but also distinguish the source and target regions. Extensive experimental results on several commonly used copy-move datasets have shown that the proposed method outperforms the state-of-the-art methods for copy-move detection. Yulan Zhang, Guopu Zhu, Xiangyang Luo 0001, Yicong Zhou, Hongli Zhang 0001, Ligang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Deep Dynamic Memory Augmented Attentional Dictionary Learning for Image DenoisingabstractMotivated by the advance of deep learning methods, deep unfolding methods such as deep convolutional dictionary learning have achieved great success in image denoising tasks. The main advantages are inheriting both the merits of deep learning (strong learning capacity) and traditional machine learning (powerful interpretable capacity). We observe that the update of dictionaries and coefficients is highly correlated with the previous iterative stage information for deep unfolding-based methods. However, most existing deep convolutional dictionary learning methods deal with each iteration step individually, ignoring the inner-memory within the stage and cross-memory across the stages. To alleviate these issues, we propose a dynamic inner-cross memory augmented attentional dictionary learning (M2ADL) network with attention guided residual connection module, which utilizes the previous important stage features such that better uncovering the inner-cross information. Specifically, the proposed inner-cross memory fully utilizes the previous stage’s hidden and last-layer information to learn the dictionary. In addition, we develop a dual attention-guided residual connection module to well exploit the deep feature learning ability to capture the spatial-spectral attention across the deep tensor-based features. Considerable experiments on both synthetic and real image datasets demonstrate the superiority of the proposed method over other state-of-the-art methods. Yongyong Chen, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | ExS-GAN: Synthesizing Anti-Forensics Images via Extra Supervised GANabstractSo far, researchers have proposed many forensics tools to protect the authenticity and integrity of digital information. However, with the explosive development of machine learning, existing forensics tools may compromise against new attacks anytime. Hence, it is always necessary to investigate anti-forensics to expose the vulnerabilities of forensics tools. It is beneficial for forensics researchers to develop new tools as countermeasures. To date, one of the potential threats is the generative adversarial networks (GANs), which could be employed for fabricating or forging falsified data to attack forensics detectors. In this article, we investigate the anti-forensics performance of GANs by proposing a novel model, the ExS-GAN, which features an extra supervision system. After training, the proposed model could launch anti-forensics attacks on various manipulated images. Evaluated by experiments, the proposed method could achieve high anti-forensics performance while preserving satisfying image quality. We also justify the proposed extra supervision via an ablation study. Feng Ding 0007, Zhangyi Shen, Guopu Zhu, Sam Kwong, Yicong Zhou, Siwei Lyu |
IEEE Trans. Cybern. | 5 |
| 2023 | Refined Prototypical Contrastive Learning for Few-Shot Hyperspectral Image ClassificationabstractRecently, prototypical network based few-shot learning (FSL) has been introduced for small-sample hyperspectral image (HSI) classification and shown good performance. However, existing prototypical-based FSL methods have two problems: prototype instability and domain shift between training and testing datasets. To solve these problems, we propose a refined prototypical contrastive learning network for few-shot learning (RPCL-FSL) in this paper, which incorporates supervised contrastive learning and FSL into an end-to-end network to perform small-sample HSI classification. To stabilize and refine the prototypes, RPCL-FSL imposes triple constraints on prototypes of the support set, i.e., contrastive learning (CL), self-calibration (SC) and cross-calibration (CC) based constraints. The CL module imposes internal constraint on the prototypes aiming to directly improve the prototypes using support set samples in the CL framework, and the SC and CC modules impose external constraints on the prototypes by using the prediction loss of support set samples and the query set prototypes, respectively. To alleviate domain shift in the FSL, a fusion training strategy is designed to reduce the feature differences between training and testing datasets. Experimental results on three HSI datasets demonstrate that the proposed RPCL-FSL outperforms existing state-of-the-art deep learning and FSL methods. Quanyong Liu, Jiangtao Peng, Yujie Ning, Na Chen 0008, Weiwei Sun 0005, Qian Du 0001, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | RAN: Region-Aware Network for Remote Sensing Image Super-ResolutionabstractThe remote sensing (RS) image super-resolution (SR) algorithm aims to reconstruct a high-resolution (HR) image with rich texture details from a given low-resolution (LR) image, improving the spatial resolution. It has been widely concerned in remote sensing image processing and application. Most current deep learning-based methods rely on paired training datasets. However, most datasets are often based on bicubic degradation. This single construction way limits the performance of the pre-trained network. Moreover, SR is an ill-posed problem in that multiple SR images are constructed from a single LR input. This paper proposes a Region-Aware Network (RAN) for remote sensing image super-resolution to alleviate the above issues. First, we introduce the contrastive learning strategy to mine the latent degraded representation of the image and serve as the prior knowledge of the network. Considering the RS images are acquired in specific scenes that have apparent self-similarity. Then, we propose a Region-Aware Module (RAM) based on attention mechanisms and the graph neural network to explore region information and cross-patch self-similarity. Extensive experiments have demonstrated that the proposed RAN adapts to RS image super-resolution tasks with various degradations and performs better in constructing texture information. Baodi Liu, Lifei Zhao, Shuai Shao 0006, Weifeng Liu 0001, Dapeng Tao, Weijia Cao, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Query-Efficient Adversarial Attack With Low Perturbation Against End-to-End Speech Recognition SystemsabstractWith the widespread use of automated speech recognition (ASR) systems in modern consumer devices, attack against ASR systems have become an attractive topic in recent years. Although related white-box attack methods have achieved remarkable success in fooling neural networks, they rely heavily on obtaining full access to the details of the target models. Due to the lack of prior knowledge of the victim model and the inefficiency in utilizing query results, most of the existing black-box attack methods for ASR systems are query-intensive. In this paper, we propose a new black-box attack called the Monte Carlo gradient sign attack (MGSA) to generate adversarial audio samples with substantially fewer queries. It updates an original sample based on the elements obtained by a Monte Carlo tree search. We attribute its high query efficiency to the effective utilization of the dominant gradient phenomenon, which refers to the fact that only a few elements of each origin sample have significant effect on the output of ASR systems. Extensive experiments are performed to evaluate the efficiency of MGSA and the stealthiness of the generated adversarial examples on the DeepSpeech system. The experimental results show that MGSA achieves 98% and 99% attack success rates on the LibriSpeech and Mozilla Common Voice datasets, respectively. Compared with the state-of-the-art methods, the average number of queries is reduced by 27% and the signal-to-noise ratio is increased by 31%. Shen Wang 0004, Zhaoyang Zhang 0002, Guopu Zhu, Xinpeng Zhang 0001, Yicong Zhou, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Spatio-Temporal Graph Attention Network for Sintering Temperature Long-Range Forecasting in Rotary KilnsabstractMonitoring and forecasting of sintering temperature (ST) is vital for safe, stable, and efficient operation of rotary kiln production process. Due to the complex coupling and time-varying characteristics of process data collected by the distributed control system, its long-range prediction remains a challenge. In this article, we propose a multivariate time series forecasting model based on dynamic spatio-temporal graph attention network (GAT) to model time-varying spatio-temporal correlation between the process data and perform long-range forecasting of ST. Aiming at the problem that there is no preset graph structure for multivariate data, we first propose an adaptive adjacency matrix generation algorithm to construct an elementary graph structure for the process data. Then, we design a spatio-temporal graph attention module, which consists of a multihead GAT for extracting time-varying spatial features and a gated dilated convolutional network for temporal features. Finally, considering the different time delay and rhythm of each process variable, we use dynamic system analysis to estimate the delay time and rhythm of each variable to guide the selection of dilation rates in dilated convolutional layers. The application results based on actual data show that the method has high prediction accuracy, and has broad application prospects in industrial processes. Hua Chen 0008, Yu Jiang 0013, Xiaogang Zhang 0002, Yicong Zhou, Lianhong Wang, Jinchao Wei |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Extrinsic Self-Calibration of the Surround-View System: A Weakly Supervised ApproachabstractAn SVS usually consists of four wide-angle fisheye cameras mounted around the vehicle to sense the surrounding environment. From the images synchronously captured by all cameras, a top-down surround-view can be synthesized, on the premise that both intrinsics and extrinsics of the cameras have been calibrated. At present, the intrinsic calibration approach is relatively well-developed and can be pipelined, while the extrinsic calibration is still immature. On one hand, the existing manual calibration schemes are usually reliable, but need to be conducted by professionals in specific sites, which is undoubtedly cumbersome. On the other hand, the majority of the existing self- calibration schemes are based on low-level features and their stability and robustness are usually unsatisfactory. As far as we know, an effective extrinsic self-calibration scheme designed specially for the SVS is still lacking. To fill such a research gap to some extent, we propose a novel self-calibration scheme which follows a weakly supervised framework, namely WESNet (Weakly-supervised Extrinsic Self-calibration Network). The training of WESNet consists of two stages. First, we utilize the corners in a few calibration site images as the weak supervision to roughly optimize the network by minimizing the geometric loss. Then, after the convergency in the first stage, we additionally introduce a self-supervised photometric loss term that can be constructed by the photometric information from natural images for further fine-tuning. Besides, to support training, we totally collected 19,078 groups of synchronously captured fisheye images under various environmental conditions. To our knowledge, thus far this is the largest surround-view dataset containing original fisheye images. By means of learning prior knowledge from the training data, WESNet takes the original fisheye images synchronously collected as the input, and directly yields extrinsics end-to-end with little labor cost. Its efficiency and efficacy have been corroborated by extensive experiments conducted on our collected dataset. To make our results reproducible, source code and the collected dataset have been released at https://cslinzhang.github.io/WESNet/WESNet.html. Yang Chen 0037, Lin Zhang 0014, Ying Shen 0005, Brian Nlong Zhao, Yicong Zhou |
IEEE Trans. Multim. | 5 |
| 2023 | Learning Sparse and Discriminative Multimodal Feature Codes for Finger RecognitionabstractCompared with uni-modal biometrics systems, multimodal biometrics systems using multiple sources of information for establishing an individual’s identity have received considerable attention recently. However, most traditional multimodal biometrics techniques generally extract features from each modality independently, ignoring the implicit associations between different modalities. In addition, most existing work uses hand-crafted descriptors that are difficult to capture the latent semantic structure. This paper proposes to learn the sparse and discriminative multimodal feature codes (SDMFCs) for multimodal finger recognition, which simultaneously takes into account the specific and common information among inter-modality and intra-modality. Specifically, given the multimodal finger images, we first establish the local difference matrix to capture informative texture features in local patches. Then, we aim to jointly learn discriminative and compact binary codes by constraining the observations from multiple modalities. Finally, we develop a novel SDMFC-based multimodal finger recognition framework, which integrates the local histograms of each division block in the learned binary codes together for classification. Experimental results on three commonly used finger databases demonstrate the effectiveness and robustness of the proposed framework in multimodal biometrics tasks. Shuyi Li 0003, Bob Zhang 0001, Lunke Fei, Shuping Zhao, Yicong Zhou |
IEEE Trans. Multim. | 5 |
| 2023 | D-LIOM: Tightly-Coupled Direct LiDAR-Inertial Odometry and MappingabstractSimultaneous localization and mapping via LiDAR-Inertial fusion is a crucial technology in many automation-related applications. Recently, a number of approaches based on geometric features have evolved, yielding impressive results via tightly-coupled estimation. This sort of feature-based techniques, however, are inextricably linked to the scanning mechanism of the LiDAR, relying on stable feature detection, and thus are difficult to adapt to multi-LiDAR systems. A few “direct” solutions, on the other hand, register the raw point cloud with the built probability map, which is more computationally efficient and easy to be extended. But, the existing direct approaches are all loosely-coupled, lacking correction of the IMU biases, and thus only work well in 2D cases. To this end, we present D-LIOM, a tightly-coupled Direct LiDAR-Inertial Odometry and Mapping framework. In D-LIOM, a scan is directly registered to a probability submap, and the LiDAR odometry, the IMU pre-integration, and the gravity constraint are integrated to build a local factor graph in the submap's time window, allowing the system to perform real-time high-precision pose estimation. Furthermore, to eliminate accumulated errors in time, we detect loops and adjust the sparse pose graph based on mutual matching of projected 2D submaps, allowing D-LIOM to run stably in large-scale scenes. In addition, to improve its flexibility to varied sensor combinations, D-LIOM supports multi-LiDAR inputs and facilitates the initialization with a common 6-axis IMU. Extensive experiments demonstrate that D-LIOM largely outperforms the existing state-of-the-art counterparts in mapping effect and localization accuracy as well as with high time efficiency. Lastly, to ensure that our results are entirely reproducible, all necessary data and codes are made open-source available. One introduction video can also be found on the online website. Zhong Wang 0009, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou |
IEEE Trans. Multim. | 4 |
| 2023 | AMS-Net: Adaptive Multi-Scale Network for Image Compressive SensingabstractRecently, deep convolutional neural networks have been applied to image compressive sensing (CS) to improve reconstruction quality while reducing computation cost. Existing deep learning-based CS methods can be divided into two classes: sampling image at single scale and sampling image across multiple scales. However, these existing methods treat the image low-frequency and high-frequency components equally, which is an obstruction to get a high reconstruction quality. This paper proposes an adaptive multi-scale image CS network in wavelet domain called AMS-Net, which fully exploits the different importance of image low-frequency and high-frequency components. First, the discrete wavelet transform is used to decompose an image into four sub-bands, namely the low-low (LL), low-high (LH), high-low (HL), and high-high (HH) sub-bands. Considering that the LL sub-band is more important to the final reconstruction quality, the AMS-Net allocates it a larger sampling ratio, while allocating the other three sub-bands a smaller one. Since different blocks in each sub-band have different sparsity, the sampling ratio is further allocated block-by-block within the four sub-bands. Then a dual-channel scalable sampling model is developed to adaptively sample the LL and the other three sub-bands at arbitrary sampling ratios. Finally, by unfolding the iterative reconstruction process of the traditional multi-scale block CS algorithm, we construct a multi-stage reconstruction model to utilize multi-scale features for further improving the reconstruction quality. Experimental results demonstrate that the proposed model outperforms both the traditional and state-of-the-art deep learning-based methods. Zhongyun Hua, Yuanman Li, Yongyong Chen, Yicong Zhou |
IEEE Trans. Multim. | 5 |
| 2023 | LMFFNet: A Well-Balanced Lightweight Network for Fast and Accurate Semantic SegmentationabstractReal-time semantic segmentation is widely used in autonomous driving and robotics. Most previous networks achieved great accuracy based on a complicated model involving mass computing. The existing lightweight networks generally reduce the parameter sizes by sacrificing the segmentation accuracy. It is critical to balance the parameters and accuracy for real-time semantic segmentation. In this article, we propose a lightweight multiscale-feature-fusion network (LMFFNet) mainly composed of three types of components: split-extract-merge bottleneck (SEM-B) block, feature fusion module (FFM), and multiscale attention decoder (MAD), where the SEM-B block extracts sufficient features with fewer parameters. FFMs fuse multiscale semantic features to effectively improve the segmentation accuracy and the MAD well recovers the details of the input images through the attention mechanism. Without pretraining, LMFFNet-3-8 achieves 75.1% mean intersection over union (mIoU) with 1.4 M parameters at 118.9 frames/s using RTX 3090 GPU. More experiments are investigated extensively on various resolutions on other three datasets of CamVid, KITTI, and WildDash2. The experiments verify that the proposed LMFFNet model makes a decent tradeoff between segmentation accuracy and inference speed for real-time tasks. The source code is publicly available at https://github.com/Greak-1124/LMFFNet. Jialin Shen, Qingming Yi, Jian Weng 0001, Zunkai Huang, Aiwen Luo, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | SLAM for Indoor Parking: A Comprehensive Benchmark Dataset and a Tightly Coupled Semantic FrameworkabstractFor the task of autonomous indoor parking, various Visual-Inertial Simultaneous Localization And Mapping (SLAM) systems are expected to achieve comparable results with the benefit of complementary effects of visual cameras and the Inertial Measurement Units. To compare these competing SLAM systems, it is necessary to have publicly available datasets, offering an objective way to demonstrate the pros/cons of each SLAM system. However, the availability of such high-quality datasets is surprisingly limited due to the profound challenge of the groundtruth trajectory acquisition in the Global Positioning Satellite denied indoor parking environments. In this article, we establish BeVIS, a large-scale Be nchmark dataset with V isual (front-view), I nertial and S urround-view sensors for evaluating the performance of SLAM systems developed for autonomous indoor parking, which is the first of its kind where both the raw data and the groundtruth trajectories are available. In BeVIS, the groundtruth trajectories are obtained by tracking artificial landmarks scattered in the indoor parking environments, whose coordinates are recorded in a surveying manner with a high-precision Electronic Total Station. Moreover, the groundtruth trajectories are comprehensively evaluated in terms of two respects, the reprojection error and the pose volatility, respectively. Apart from BeVIS, we propose a novel tightly coupled semantic SLAM framework, namely VIS SLAM -2, leveraging V isual (front-view), I nertial, and S urround-view sensor modalities, specially for the task of autonomous indoor parking. It is the first work attempting to provide a general form to model various semantic objects on the ground. Experiments on BeVIS demonstrate the effectiveness of the proposed VIS SLAM -2. Our benchmark dataset BeVIS is publicly available at https://shaoxuan92.github.io/BeVIS . Xuan Shao, Ying Shen 0005, Lin Zhang 0014, Shengjie Zhao 0001, Dandan Zhu 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Generation of n-Dimensional Hyperchaotic Maps Using Gershgorin-Type Theorem and its ApplicationabstractHigh-dimensional (HD) chaotic map has wide applications in various research fields such as neural networks and secure communication. Designing HD chaotic maps with expected dynamics and robust hyperchaotic behaviors is an interesting but challenging topic. In this article, we propose an$n$-dimensional hyperchaotic map$(n\text{D}$-HCM) generation method on the basis of the Gershgorin-type theorem. First, the general form of the proposed$n\text{D}$-HCM is built using$n$parametric polynomials. Then, the entity and coefficient parameter matrices are configured according to the Gershorin-type theorem. Theoretical analysis shows that the generated$n\text{D}$-HCM has$n$positive Lyapunov exponents and thus can show robust hyperchaotic behaviors. Two examples of hyperchaotic map with specified equations are provided and their properties are analyzed to show the availability of the proposed method. Performance evaluations display that our$n\text{D}$-HCM possesses abundant properties and complex behaviors, and it can outperform some representative HD chaotic maps. Moreover, to show the application of our$n\text{D}$-HCM, we apply it to a secure communication scheme and the experimental results exhibit that it shows much better performance than these representative HD chaotic maps in resisting transmission noise. Yinxing Zhang, Zhongyun Hua, Han Bao 0001, Hejiao Huang, Yicong Zhou |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2022 | Automatic Color Image Stitching Using Quaternion Rank-1 AlignmentabstractColor image stitching is a challenging task in real-world applications. This paper first proposes a quaternion rank-1 alignment (QR1A) model for high-precision color image alignment. To solve the optimization problem of QR1A, we develop a nested iterative algorithm under the framework of complex-valued alternating direction method of multipliers. To quantitatively evaluate image stitching performance, we propose a perceptual seam quality (PSQ) measure to calculate misalignments of local regions along the seamline. Using QR1A and PSQ, we further propose an automatic color image stitching (ACIS-QR1A) framework. In this frame-work, the automatic strategy and iterative learning strategy are developed to simultaneously learn the optimal seamline and local alignment. Extensive experiments on challenging datasets demonstrate that the proposed ACIS-QR1A is able to obtain high-quality stitched images under several difficult scenarios including large parallax, low textures, moving objects, large occlusions or/and their combinations. Jiaxue Li, Yicong Zhou |
CVPR | 2 |
| 2022 | Towards Controllable and Physical Interpretable Underwater Scene SimulationabstractThe realistic simulation of underwater scenes has important significance for many researches related to underwater vision, such as underwater image restoration, underwater moving object monitoring, etc. To date, however, the existing underwater scene simulation pipelines are either too complicated due to the continuous spectra and camera parameters involved, or difficult to control since the empirically controlled distance based fog effect is usually used by them. In this paper, we try to fill in this research gap by proposing an Underwater Scene Simulation approach, namely USSim, which especially focuses on the influence of ocean water. In USSim, Jerlov water type and depth are regarded as main variables to control the simulation effects. In addition, the spectra of the incident light is decomposed into three primary components and their attenuations are modeled separately, and finally the simulated scene is generated via the hybrid underwater imaging model proposed by us. USSim greatly reduces the computational complexity and enables the fog effect to be controlled by variables with explicit physical meanings. The controllability, physical interpretability and simulation effects of our USSim under different conditions have been verified by extensive experiments. To make our results reproducible, the source code is made online available at https://cslinzhang.github.io/USSim/. Kaixin Chen 0003, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou |
ICASSP | 4 |
| 2022 | Combining Multiple Style Transfer Networks and Transfer Learning For LGE-CMR SegmentationabstractThis paper presents an algorithm for segmenting late gadolinium enhancement cardiac magnetic resonance (LGE-CMR) in the absence of labeled training data. The proposed method includes a data augmentation part and a segmentation network. Multiple style transfer networks are employed for data augmentation to increase the data diversity, and then the synthetic images are used for training an improved U-Net. Finally, the trained model is fine-tuned with a few LGE images and labels. Experiment results demonstrate the effectiveness and advantages of the proposed method. Bo Fang 0005, Junxin Chen 0001, Wei Wang 0077, Yicong Zhou |
ICASSP | 4 |
| 2022 | Chunkfusion: A Learning-Based RGB-D 3D Reconstruction Framework Via Chunk-Wise IntegrationabstractRecent years have witnessed a growing interest in online RGB-D 3D reconstruction. On the premise of ensuring the reconstruction accuracy with noisy depth scans, making the system scalable to various environments is still challenging. In this paper, we devote our efforts to try to fill in this research gap by proposing a scalable and robust RGB-D 3D reconstruction framework, namely Chunk-Fusion. In ChunkFusion, sparse voxel management is exploited to improve the scalability of online reconstruction. Besides, a chunk-wise TSDF (truncated signed distance function) fusion network is designed to perform a robust integration of the noisy depth measurements on the sparsely allocated voxel chunks. The proposed chunk-wise TSDF integration scheme can accurately restore surfaces with superior visual consistency from noisy depth maps and can guarantee the scalability of online reconstruction simultaneously, making our reconstruction framework widely applicable to scenes with various scales and depth scans with strong noises and outliers. The outstanding scalability and efficacy of our ChunkFusion have been corroborated by extensive experiments. To make our results reproducible, the source code is made online available at https://cslinzhang.github.io/ChunkFusion/. Chaozheng Guo, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou |
ICASSP | 4 |
| 2022 | Graph Learning Based Autoencoder for Hyperspectral Band SelectionabstractHyperspectral band selection aims to identify an optimal sub-set of bands from hyperspectral images (HSIs). Most existing methods explore the relationships between pair-wise pixels in a fixed graph. However, the quality of the initial fixed graph may be influenced by noises and user-defined parameters that may not be optimal for HSI analysis. In this paper, we pro-pose a graph learning based autoencoder (GLAE) to achieve unsupervised hyperspectral band selection. Using the relationships of pair-wise pixels within HSIs, GLAE constructs the initial graph to characterize the geometric structures of HSIs and then adjusts the graph to adapt the band selection process. To solve the proposed model, we intoduce an alternative optimization algorithm. Experiments and comparisons on three HSI datasets demonstrate that the proposed GLAE achieves better results over the state-of-the-art methods. Yongshan Zhang, Xinxin Wang 0003, Xinwei Jiang, Yicong Zhou |
ICASSP | 5 |
| 2022 | More Efficient and Locally Enhanced Transformer
Zhefeng Zhu, Ke Qi, Yicong Zhou, Wenbin Chen 0003, Jingdong Zhang 0002 |
ICONIP (5) | 3 |
| 2022 | LVI-ExC: A Target-free LiDAR-Visual-Inertial Extrinsic Calibration FrameworkabstractRecently, the multi-modal fusion with 3D LiDAR, camera, and IMU has shown great potential in applications of automation-related fields. Yet a prerequisite for a successful fusion is that the geometric relationships among the sensors are accurately determined, which is called an extrinsic calibration problem. To date, the existing target-based approaches to deal with this problem rely on sophisticated calibration objects (sites) and well-trained operators, which is time-consuming and inflexible in practical applications. Contrarily, a few target-free methods can overcome these shortcomings, while they only focus on the calibrations of two types of the sensors. Although it is possible to obtain LiDAR-visual-inertial extrinsics by chained calibrations, problems such as cumbersome operations, large cumulative errors, and weak geometric consistency still exist. To this end, we propose LVI-ExC, an integrated LiDAR-Visual-Inertial Extrinsic Calibration framework, which takes natural multi-modal data as input and yields sensor-to-sensor extrinsics end-to-end without any auxiliary object (site) or manual assistance. To fuse multi-modal data, we formulate the LiDAR-visual-inertial extrinsic calibration as a continuous-time simultaneous localization and mapping problem, in which the extrinsics, trajectories, time differences, and map points are jointly estimated by establishing sensor-to-sensor and sensor-to-trajectory constraints. Extensive experiments show that LVI-ExC can produce precise results. With LVI-ExC's outputs, the LiDAR-visual reprojection results and the reconstructed environment map are all highly consistent with the actual natural scenes, demonstrating LVI-ExC's outstanding performance. To ensure that our results are fully reproducible, all the relevant data and codes have been released publicly at https://cslinzhang.github.io/LVI-ExC/. Zhong Wang 0009, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou |
ACM Multimedia | 4 |
| 2022 | An efficient unsupervised image quality metric with application for condition recognition in kiln
Leyuan Wu, Xiaogang Zhang 0002, Hua Chen 0008, Yicong Zhou, Lianhong Wang, Dingxiang Wang |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | A General Loss-Based Nonnegative Matrix Factorization for Hyperspectral UnmixingabstractNonnegative matrix factorization (NMF) is a widely used hyperspectral unmixing model which decomposes a known hyperspectral data matrix into two unknown matrices, i.e., endmember matrix and abundance matrix. Due to the use of least-squares loss, the NMF model is usually sensitive to noise or outliers. To improve its robustness, we introduce a general robust loss function to replace the traditional least-squares loss and propose a general loss-based NMF (GLNMF) model for hyperspectral unmixing in this letter. The general loss function is a superset of many common robust loss functions and is suitable for handling different types of noise. Experimental results on simulated and real hyperspectral data sets demonstrate that our GLNMF model is more accurate and robust than existing NMF methods. Jiangtao Peng, Weiwei Sun 0005, Hong Chen 0004, Yicong Zhou, Qian Du 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Neural Style Transfer With Adaptive Auto-Correlation Alignment LossabstractThe neural style transfer has achieved a significant improvement with deep learning methods. However, the existing methods are susceptible to lack the ability for handling the texture style transfer because of their less consideration of the textural structure from style images. To overcome this drawback, this letter presents a simple method to capture the textural structure by using an adaptive auto-correlation alignment loss function. Furthermore, we also introduce three metrics to quantitatively evaluate the performance. We qualitatively and quantitatively evaluate the proposed methods. The experimental results demonstrate the superiority of the proposed method and our method can synthesize the stylized images with rich texture style patterns. Yue Wu 0001, Xiaofei Yang 0002, Yicong Zhou |
IEEE Signal Process. Lett. | 4 |
| 2022 | n-Dimensional Polynomial Chaotic System With ApplicationsabstractDesigning high-dimensional chaotic maps with expected dynamic properties is an attractive but challenging task. The dynamic properties of a chaotic system can be reflected by the Lyapunov exponents (LEs). Using the inherent relationship between the parameters of a chaotic map and its LEs, this paper proposes an$n$-dimensional polynomial chaotic system ($n\text{D}$-PCS) that can generate$n\text{D}$chaotic maps with any desired LEs. The$n\text{D}$-PCS is constructed from$n$parametric polynomials with arbitrary orders, and its parameter matrix is configured using the preliminaries in linear algebra. Theoretical analysis proves that the$n\text{D}$-PCS can produce high-dimensional chaotic maps with any desired LEs. To show the effects of the$n\text{D}$-PCS, two high-dimensional chaotic maps with hyperchaotic behaviors were generated. A microcontroller-based hardware platform was developed to implement the two chaotic maps, and the test results demonstrated the randomness properties of their chaotic signals. Performance evaluations indicate that the high-dimensional chaotic maps generated from$n\text{D}$-PCS have the desired LEs and more complicated dynamic behaviors compared with other high-dimensional chaotic maps. In addition, to demonstrate the applications of$n\text{D}$-PCS, we developed a chaos-based secure communication scheme. Simulation results show that$n\text{D}$-PCS has a stronger ability to resist channel noise than other high-dimensional chaotic maps. Zhongyun Hua, Yinxing Zhang, Han Bao 0001, Hejiao Huang, Yicong Zhou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | Reversible Data Hiding for Color Images Based on Adaptive 3D Prediction-Error Expansion and Double Deep Q-NetworkabstractReversible data hiding (RDH) for color images has attracted increasing attention in recent years. Due to its effective utilization of the correlation between prediction errors, high-dimensional prediction-error expansion (PEE) can achieve much better performance for color image RDH than low-dimensional PEE. However, existing studies only focus on high-dimensional PEE with nonadaptive embedding. To further improve the embedding performance for color images, we propose a novel three-dimensional PEE method that is adaptive to image content. Double deep Q-network (DDQN), introduced to RDH for the first time, is adopted to find the optimal mapping paths for PEE. In addition, an action selection scheme is presented for DDQN to efficiently find the reversible mapping paths. Extensive experiments show that the proposed method outperforms existing color image RDH methods in image quality. Guopu Zhu, Hongli Zhang 0001, Yicong Zhou, Xiangyang Luo 0001, Ligang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Low-Rank Tensor Graph Learning for Multi-View Subspace ClusteringabstractGraph and subspace clustering methods have become the mainstream of multi-view clustering due to their promising performance. However, (1) since graph clustering methods learn graphs directly from the raw data, when the raw data is distorted by noise and outliers, their performance may seriously decrease; (2) subspace clustering methods use a “two-step” strategy to learn the representation and affinity matrix independently, and thus may fail to explore their high correlation. To address these issues, we propose a novel multi-view clustering method via learning aLow-RankTensorGraph (LRTG). Different from subspace clustering methods, LRTG simultaneously learns the representation and affinity matrix in a single step to preserve their correlation. We apply Tucker decomposition and$l_{2,1}$-norm to the LRTG model to alleviate noise and outliers for learning a “clean” representation. LRTG then learns the affinity matrix from this “clean” representation. Additionally, an adaptive neighbor scheme is proposed to find the$K$largest entries of the affinity matrix to form a flexible graph for clustering. An effective optimization algorithm is designed to solve the LRTG model based on the alternating direction method of multipliers. Extensive experiments on different clustering tasks demonstrate the effectiveness and superiority of LRTG over seventeen state-of-the-art clustering methods. Yongyong Chen, Xiaolin Xiao, Chong Peng 0001, Guangming Lu 0002, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Reversible Data Hiding in Encrypted Images Using Cipher-Feedback Secret SharingabstractReversible data hiding in encrypted images (RDH-EIs) has attracted increasing attention since it can protect the privacy of original images while exactly extracting the embedded data. In this paper, we propose an RDH-EI scheme with multiple data hiders. First, we introduce a cipher-feedback secret sharing (CFSS) technique using the cipher-feedback strategy of the Advanced Encryption Standard. Then, using the CFSS technique, we devise a new$(r,n)$-threshold ($r\leq n$) RDH-EI scheme with multiple data hiders called CFSS-RDHEI. It can encrypt an original image into$n$encrypted images with reduced size using an encryption key and sends each encrypted image to one data hider. Each data hider can independently embed secret data into the encrypted image to obtain a marked encrypted image. The embedded data can be extracted from each marked encrypted image using the data hiding key, and the original image can be completely recovered from$r$marked encrypted images using the encryption key. Performance evaluations show that our CFSS-RDHEI scheme has a higher embedding rate and that its generated encrypted images are much smaller, while still being well protected, compared to existing secret sharing-based RDH-EI schemes. Zhongyun Hua, Yicong Zhou, Xiaohua Jia |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | MOFISSLAM: A Multi-Object Semantic SLAM System With Front-View, Inertial, and Surround-View Sensors for Indoor ParkingabstractThe semantic SLAM (Simultaneous Localization And Mapping) system is a crucial module for autonomous indoor parking. Visual cameras (monocular/binocular) and IMU (Inertial Measurement Unit) constitute the basic configuration to build such a system. The performance of existing SLAM systems typically deteriorates in the presence of dynamically movable objects or objects with little texture. By contrast, semantic objects on the ground embody the most salient and stable features in the indoor parking environment. Due to their inabilities to perceive such features on the ground, existing SLAM systems are prone to tracking inconsistency during navigation. In this paper, we present MOFISSLAM, a novel tightly-coupled${M}$ulti-${O}$bject semantic SLAM system integrating${F}$ront-view,${I}$nertial, and${S}$urround-view sensors for autonomous indoor parking. The proposed system moves beyond existing semantic SLAM systems by complementing the sensor configuration with a surround-view system capturing images from a top-down viewpoint. In MOFISSLAM, apart from low-level visual features and inertial motion data, typical semantic objects (parking-slots, parking-slot IDs and speed bumps) detected in surround-views are also incorporated in optimization, forming robust surround-view constraints. Specifically, each surround-view feature imposes a surround-view constraint that can be split into a contact term and a registration term. The former pre-defines the position of each individual surround-view feature subject to whether it has semantic contact with other surround-view features. Three contact modes, defined ascomplementary,adjacentandcoincident, are identified to guarantee a unified form of all contact terms. The latter further constrains by registering each surround-view observation and its position in the world coordinate system. In parallel, to objectively evaluate SLAM studies for autonomous indoor parking, a large-scale dataset with groundtruth trajectories is collected, which is the first of its kind. Its groundtruth trajectories, commonly unavailable, are obtained by tracking artificial features scattered in the indoor parking environment, whose 3D coordinates are measured with an ETS (Electronic Total Station). The collected dataset has been made publicly available athttps://shaoxuan92.github.io/MOFIS. Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | GCT: Graph Co-Training for Semi-Supervised Few-Shot LearningabstractFew-shot learning (FSL), purposing to resolve the problem of data-scarce, has attracted considerable attention in recent years. A popular FSL framework contains two phases: (i) the pre-train phase employs the base data to train a CNN-based feature extractor. (ii) the meta-test phase applies the frozen feature extractor to novel data (novel data has different categories from base data) and designs a classifier for recognition. To correct few-shot data distribution, researchers propose Semi-Supervised Few-Shot Learning (SSFSL) by introducing unlabeled data. Although SSFSL has been proved to achieve outstanding performances in the FSL community, there still exists a fundamental problem: the pre-trained feature extractor cannot adapt to the novel data flawlessly due to the cross-category setting. Usually, large amounts of noises are introduced to the novel feature. We dub it as Feature-Extractor-Maladaptive (FEM) problem. To tackle FEM, we make two efforts in this paper. First, we propose a novel label prediction method, Isolated Graph Learning (IGL). IGL introduces the Laplacian operator to encode the raw data to graph space, which helps reduce the dependence on features when classifying, and then project graph representation to label space for prediction. The key point is that: IGL can weaken the negative influence of noise from the feature representation perspective, and is also flexible to independently complete training and testing procedures, which is suitable for SSFSL. Second, we propose Graph Co-Training (GCT) to tackle this challenge from a multi-modal fusion perspective by extending the proposed IGL to the co-training framework. GCT is a semi-supervised method that exploits the unlabeled samples with two modal features to crossly strengthen the IGL classifier. We estimate our method on five benchmark few-shot learning datasets and achieve outstanding performances compared with other state-of-the-art methods. It demonstrates the effectiveness of our GCT. Rui Xu 0012, Lei Xing 0005, Shuai Shao 0006, Lifei Zhao, Baodi Liu, Weifeng Liu 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | Forensic Analysis of JPEG-Domain Enhanced Images via Coefficient Likelihood ModelingabstractJPEG-domain enhancement improves the visual quality of JPEG images by directly manipulating the decoded DCT (discrete cosine transform) coefficients, which inevitably leads to mixed compression and enhancement artifacts. Existing forensic methods that merely consider JPEG artifacts are unsuitable to address such mixed artifacts and hence suffer a considerable performance decline in compression parameter estimation and lack the ability to estimate the enhancement parameter. This work attempts to explore the characterization of the mixed artifacts, and to further estimate both the enhancement and compression parameters of JPEG-domain enhanced images. First, a statistical likelihood function is proposed to characterize the periodicity of DCT coefficients, which can measure how well an enhanced image is de-enhanced back to its JPEG compressed version given the compression and enhancement parameters. The proposed likelihood function reaches its maximum if the parameters match their true values. Then, a forensic method of enhancement detection and parameter estimation is developed based on the proposed likelihood function for two kinds of classical JPEG-domain enhancement. Specifically, JPEG-domain enhanced images are detected by thresholding a scalar feature computed upon the likelihoods, and the enhancement and compression parameters are estimated by locating the maximal likelihood. In addition, mathematical proof of the de-enhancement feasibility is provided. Experimental results demonstrate that the proposed method outperforms the compared methods in both enhancement detection and parameter estimation. Jianquan Yang, Guopu Zhu, Yao Luo, Sam Kwong, Xinpeng Zhang 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Spectral-Spatial Feature Extraction With Dual Graph Autoencoder for Hyperspectral Image ClusteringabstractAutoencoder (AE) is an unsupervised neural network framework for efficient and effective feature extraction. Most AE-based methods do not consider spatial information and band correlations for hyperspectral image (HSI) analysis. In addition, graph-based AE methods often learn discriminative representations with the assumption that connected samples share the same label and they cannot directly embed the geometric structure into feature extraction. To address these issues, in this paper, we propose a dual graph autoencoder (DGAE) to learn discriminative representations for HSIs. Utilizing the relationships of pair-wise pixels within homogenous regions and pair-wise spectral bands, DGAE first constructs the superpixel-based similarity graph with spatial information and band-based similarity graph to characterize the geometric structures of HSIs. With the developed dual graph convolution, more discriminative feature representations are learnt from the hidden layer via the encoder-decoder structure of DGAE. The main advantage of DGAE is that it fully exploits both the geometric structures of pixels with spatial information and spectral bands to promote nonlinear feature extraction of HSIs. Experiments on HSI datasets show the superiority of the proposed DGAE over the state-of-the-art methods. The source code of DGAE is available athttps://github.com/ZhangYongshan/DGAE. Yongshan Zhang, Xinwei Jiang, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Multi-Task SE-Network for Image Splicing LocalizationabstractImage splicing can be easily used for illegal activities such as falsifying propaganda for political purposes and reporting false news, which may result in negative impacts on society. Hence, it is highly required to detect spliced images and localize the spliced regions. In this work, we propose a multi-task squeeze and excitation network (SE-Network) for splicing localization. The proposed network consists of two streams, namely label mask stream and edge-guided stream, both of which adopt convolutional encoder-decoder architecture. The information from the edge-guided stream is transmitted to the label mask stream for enhancing the discrimination of features between the spliced and host regions. This work has three main contributions. First, image edges, along with label masks and mask edges, are exploited to supply more comprehensive supervision for the localization of spliced regions. Second, the low-level feature maps extracted from shallow layers are fused with the high-level feature maps from deep layers to provide more reliable feature for splicing localization. Finally, several squeeze and excitation attention modules are incorporated into the network to recalibrate the fused features to enhance the feature expression. Extensive experiments show that the proposed multi-task SE-Network outperforms existing splicing localization methods evidently on two synthetic splicing datasets and four benchmark splicing datasets. Yulan Zhang, Guopu Zhu, Ligang Wu 0001, Sam Kwong, Hongli Zhang 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Secure Halftone Image Steganography Based on Feature Space and Layer EmbeddingabstractSyndrome-trellis codes (STCs) are commonly used in image steganographic schemes, which aim at minimizing the embedding distortion, but most distortion models cannot capture the mutual interaction of embedding modifications (MIEMs). In this article, a secure halftone image steganographic scheme based on a feature space and layer embedding is proposed. First, a feature space is constructed by a characterization method that is designed based on the statistics of 4 ×4 pixel blocks in halftone images. Upon the feature space, a generalized steganalyzer with good classification ability is proposed, which is used to measure the embedding distortion. As a result, a distortion model based on a hybrid feature space is constructed, which outperforms some state-of-the-art models. Then, as the distortion model is established on the statistics of local regions, a layer embedding strategy is proposed to reduce MIEM. It divides the host image into multiple layers according to their relative positions in 4 ×4 blocks, and the embedding procedure is executed layer by layer. In each layer, any two pixels are located at different 4 ×4 blocks in the original image, and the distortion model makes sure that the calculation of pixel distortions is independent. Between layers, the pixel distortions of the current layer are updated according to the previous embedding modifications, thus reducing the total embedding distortion. Comparisons with prior schemes demonstrate that the proposed steganographic scheme achieves high statistical security when resisting the state-of-the-art steganalysis. Wei Lu 0001, Junjia Chen, Junhong Zhang, Jiwu Huang, Jian Weng 0001, Yicong Zhou |
IEEE Trans. Cybern. | 6 |
| 2022 | Robust Estimation of Upscaling Factor on Double JPEG Compressed ImagesabstractAs one of the most important topics in image forensics, resampling detection has developed rapidly in recent years. However, the robustness to JPEG compression is still challenging for most classical spectrum-based methods, since JPEG compression severely degrades the image contents and introduces block artifacts in the boundary of the compression grid. In this article, we propose a method to estimate the upscaling factors on double JPEG compressed images in the presence of image upscaling between the two compressions. We first analyze the spectrum of scaled images and give an overall formulation of how the scaling factors along with the parameters of JPEG compression and image contents influence the appearance of tampering artifacts. The expected positions of five kinds of characteristic peaks are analytically derived. Then, we analyze the features of double JPEG compressed images in the block discrete cosine transform (BDCT) domain and present an inverse scaling strategy for the upscaling factor estimation with a detailed proof. Finally, a fusion method is proposed that through frequency-domain analysis, a candidate set of upscaling factors is given, and through analysis in the BDCT domain, the optimal estimation from all candidates is determined. The experimental results demonstrate that the proposed method outperforms other state-of-the-art methods. Wei Lu 0001, Shangjun Luo, Yicong Zhou, Jiwu Huang, Yun Q. Shi 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | Addi-Reg: A Better Generalization-Optimization Tradeoff Regularization Method for Convolutional Neural NetworksabstractIn convolutional neural networks (CNNs), generating noise for the intermediate feature is a hot research topic in improving generalization. The existing methods usually regularize the CNNs by producing multiplicative noise (regularization weights), called multiplicative regularization (Multi-Reg). However, Multi-Reg methods usually focus on improving generalization but fail to jointly consider optimization, leading to unstable learning with slow convergence. Moreover, Multi-Reg methods are not flexible enough since the regularization weights are generated from a definite manual-design distribution. Besides, most popular methods are not universal enough, because these methods are only designed for the residual networks. In this article, we, for the first time, experimentally and theoretically explore the nature of generating noise in the intermediate features for popular CNNs. We demonstrate that injecting noise in the feature space can be transformed to generating noise in the input space, and these methods regularize the networks in a Mini-batch in Mini-batch (MiM) sampling manner. Based on these observations, this article further discovers that generating multiplicative noise can easily degenerate the optimization due to its high dependence on the intermediate feature. Based on these studies, we propose a novel additional regularization (Addi-Reg) method, which can adaptively produce additional noise with low dependence on intermediate feature in CNNs by employing a series of mechanisms. Particularly, these well-designed mechanisms can stabilize the learning process in training, and our Addi-Reg method can pertinently learn the noise distributions for every layer in CNNs. Extensive experiments demonstrate that the proposed Addi-Reg method is more flexible and universal, and meanwhile achieves better generalization performance with faster convergence against the state-of-the-art Multi-Reg methods. Yao Lu 0008, Zheng Zhang 0006, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | Secret Sharing Based Reversible Data Hiding in Encrypted Images With Multiple Data-HidersabstractThe existing models of reversible data hiding in encrypted images (RDH-EI) are based on single data-hider, where the original image cannot be reconstructed when the data-hider is damaged. To address this issue, this article proposes a novel model with multiple data-hiders for RDH-EI based on secret sharing. It divides the original image into multiple different encrypted images with the same size of the original image and distributes them to multiple different data-hiders for data hiding. Each data-hider can independently embed data into the encrypted image to obtain the corresponding marked encrypted image. The original image can be losslessly recovered by collecting sufficient marked encrypted images from undamaged data-hiders when individual data-hiders are subjected to potential damage. This further protects the security of the original image. We provide four cases of the proposed model, namely, two joint cases and two separable cases. From the proposed model, we derive a separable RDH-EI method with high-capacity. Experimental results are presented to illustrate the effectiveness of the proposed method. Bing Chen 0004, Wei Lu 0001, Jiwu Huang, Jian Weng 0001, Yicong Zhou |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2022 | Hyperspectral Image Transformer Classification NetworksabstractHyperspectral image (HSI) classification is an important task in earth observation missions. Convolution neural networks (CNNs) with the powerful ability of feature extraction have shown prominence in HSI classification tasks. However, existing CNN-based approaches cannot sufficiently mine the sequence attributes of spectral features, hindering the further performance promotion of HSI classification. This article presents a hyperspectral image transformer (HiT) classification network by embedding convolution operations into the transformer structure to capture the subtle spectral discrepancies and convey the local spatial context information. HiT consists of two key modules, i.e., spectral-adaptive 3-D convolution projection module and convolution permutator (ConV-Permutator) to retrieve the subtle spatial–spectral discrepancies. The spectral-adaptive 3-D convolution projection module produces the local spatial–spectral information from HSIs using two spectral-adaptive 3-D convolution layers instead of the linear projection layer. In addition, the Conv-Permutator module utilizes the depthwise convolution operations to separately encode the spatial–spectral representations along the height, width, and spectral dimensions, respectively. Extensive experiments on four benchmark HSI datasets, including Indian Pines, Pavia University, Houston2013, and Xiongan (XA) datasets, show the superiority of the proposed HiT over existing transformers and the state-of-the-art CNN-based methods. Our codes of this work are available athttps://github.com/xiachangxue/DeepHyperXfor the sake of reproducibility. Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Self-Supervised Learning With Prediction of Image Scale and Spectral Order for Hyperspectral Image ClassificationabstractIn recent years, Convolutional Neural Networks (CNNs) have achieved great success in hyperspectral image classification attributed to their unparalleled capacity to extract the local information. However, to successfully learn the high-level semantic image features, they always require massive amounts of manually labeled data during the training process, which is expensive, scarce, and impractical, and severely hinders the improvement of supervised deep learning methods. To alleviate these burdens, we present Self-Supervised Learning methods for hyperspectral image classification by a pre-training model using extensive unlabeled data and fine-tuning the hyperspectral image target classification. In this paper, we propose a new method for learning image characteristics by training a CNN to recognize the image scale that is applied to the hyperspectral images (HSIs). In addition, we propose a multi-pretext task method to learn stable and good feature representations combing two different pretext task methods and contrastive loss function. We evaluate the proposed methods in Self-Supervised Learning benchmarks on four benchmark HSIs datasets. The experiment results demonstrate that the proposed methods outperform the traditional supervised deep learning methods when large amounts of unlabeled HSIs data are used. Moreover, it demonstrates that the Self-Supervised Learning method is promising to alleviate dependence on manually labeled data of hyperspectral image classification. Finally, our research contributes to the creation and refinement of Self-Supervised Learning methods for pretextual tasks within the HSIs community. Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Local Correntropy Matrix Representation for Hyperspectral Image ClassificationabstractThe hyperspectral images (HSIs) classification technique has received widespread attention in the field of remote sensing. However, how to achieve satisfactory classification performance in the presence of a large amount of noise is still a problem worthy of consideration. In this article, a local correntropy matrix (LCEM)-based spatial–spectral feature representation method is proposed for HSI classification. Motivated by the successful application of information-theoretic learning (ITL), we propose to adopt correntropy matrix to represent the spatial–spectral features of HSI. Specifically, the dimension reduction is first performed on the original hyperspectral data. Then, for each pixel, we select its local neighbors within a sliding window using cosine distance for the construction of the LCEM. In this way, each pixel can be characterized as an LCEM. Finally, all the correntropy matrices are fed into a support vector machine (SVM) for final classification. In addition, we also propose a novel way to determine the size of the local window based on standard deviation. Because the LCEM as the feature descriptor can characterize discriminative spatial–spectral features, the proposed method has shown great interclass separability and intraclass compactness. Compared with other advanced approaches, the proposed LCEM method has achieved competitive performance in both evaluation indexes and visual effects, especially when the training size is very small. Xinyu Zhang 0025, Yantao Wei, Weijia Cao, Huang Yao, Jiangtao Peng, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Marginalized Graph Self-Representation for Unsupervised Hyperspectral Band SelectionabstractUnsupervised band selection is an essential step in preprocessing hyperspectral images (HSIs) to select informative bands. Most existing methods exploit the spatial information from the entire HSI while ignoring the difference between diverse homogeneous regions. Moreover, traditional methods utilize the limited size of data for model training that may result in degraded generalization performance. In this article, we propose a marginalized graph self-representation (MGSR) method for unsupervised hyperspectral band selection. To explore the spatial information from diverse homogenous regions, MGSR generates the segmentations of an HSI by superpixel segmentation and records the relationships between adjacent pixels of the same segmentation in a structural graph. Meanwhile, to improve the generalization and robustness, infinite corrupted samples are obtained from the original pixels by introducing noises in spectral bands for model training. To solve the proposed formulation, we design an alternating optimization algorithm to marginalize out the corruption and search for the optimal solution. Experimental studies on HSI datasets demonstrate the effectiveness of the proposed MGSR and the superiority over the state-of-the-art methods. The source code is available athttps://github.com/ZhangYongshan/MGSR. Yongshan Zhang, Xinxin Wang 0003, Xinwei Jiang, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Robust Dual Graph Self-Representation for Unsupervised Hyperspectral Band SelectionabstractUnsupervised band selection aims to select informative spectral bands to preprocess hyperspectral images (HSIs) without using labels. Traditional band selection methods only work well on Euclidean data, but ignore structural information of pixels and spectral bands. Moreover, they treat each HSI as a whole to exploit latent spatial information while ignoring the difference of spatial distribution between diverse homogeneous regions. In this paper, we propose a robust dual graph self-representation (RDGSR) method for unsupervised band selection. RDGSR uses superpixel segmentation technique to generate homogenous regions of each HSI to extract spatial information. Based on the segmentation result, the superpixel-based similarity graph and band-based similarity graph are constructed from HSIs to record spatial and structural information. With this knowledge, the dual graph convolution is developed and thel2,1-norm is introduced in the loss function and regularization term to eliminate the noise in rows for robust and effective band selection. The novelty of RDGSR is the joint utilization of the geometric structure of pixels with spatial consistency and the geometric structure of spectral bands to enhance the performance of band selection in a robustl2,1-norm manner. An iterative optimization algorithm is designed to solve the proposed formulation. Substantial experiments on HSI datasets are conducted to verify the superiority of the proposed RDGSR over the state-of-the-art methods. The source code is available at https://github.com/ZhangYongshan/RDGSR. Yongshan Zhang, Xinxin Wang 0003, Xinwei Jiang, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Combustion Condition Recognition of Coal-Fired Kiln Based on Chaotic Characteristics Analysis of Flame VideoabstractKeeping combustion stable and detecting unstable states in time is crucial for coal-fired furnaces such as rotary kilns, boilers, and oxygen furnaces. Because of the interference and complex conditions in the industrial field, recognition of combustion conditions by vision analysis is difficult. In this article, we propose a robust nonlinear dynamic system analysis-based approach for combustion condition recognition by extracting chaotic characteristics from a flame video. We first discover chaotic characteristics in the intensity sequence extracted from a flame video of coal-fired kilns, and then we further find that the underlying chaos rules differ between combustion conditions. Based on this finding, we design a set of trajectory evolution features and morphology distribution features of chaotic attractors for combustion condition recognition. After reconstructing the chaotic attractors from the intensity sequence of a flame video by phase space reconstruction, the quantified features are extracted from the recurrence plot and morphology distribution and put into a decision tree to recognize the combustion condition. The experimental results on real-world data show that the proposed method can recognize the combustion condition in coal-fired kilns effectively and promptly. Compared with other methods, the recognition accuracy is improved more than 5%. Yu Jiang 0013, Hua Chen 0008, Xiaogang Zhang 0002, Yicong Zhou, Lianhong Wang |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | An $n$-Dimensional Chaotic System Generation Method Using Parametric Pascal MatrixabstractWhen high-dimensional chaotic systems are applied to many practical applications, they are required to have robust and complex hyperchaotic behaviors. In this article, we propose a novel$n$D chaotic system construction method using the Pascal-matrix theory. First, a parametric Pascal matrix is constructed. Then, an$n$D chaotic system can be generated by using the parametric Pascal matrix as the parameter matrix of the system. Theoretical analysis shows that the generated$n$D chaotic systems have robust and complex chaotic behaviors, and they become$n$D Arnold Cat maps by fixing the parameters as some special values. Performance evaluations demonstrate that the$n$D chaotic systems have more complex chaotic behaviors and better distribution of outputs compared with existing HD chaotic systems. A 4-D Arnold Cat map and a 4-D chaotic map with hyperchaotic behaviors are generated as two examples. The two chaotic maps are then simulated on a microcontroller-based hardware platform and the chaotic sequences are tested to show good randomness. Yinxing Zhang, Zhongyun Hua, Han Bao 0001, Hejiao Huang, Yicong Zhou |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | CVIDS: A Collaborative Localization and Dense Mapping Framework for Multi-Agent Based Visual-Inertial SLAMabstractNowadays, visual SLAM (Simultaneous Localization And Mapping) has become a hot research topic due to its low costs and wide application scopes. Traditional visual SLAM frameworks are usually designed for single-agent systems, completing both the localization and the mapping with sensors equipped on a single robot or a mobile device. However, the mobility and work capacity of the single agent are usually limited. In reality, robots or mobile devices sometimes may be deployed in the form of clusters, such as drone formations, wearable motion capture systems, and so on. As far as we know, existing SLAM systems designed for multi-agents are still sporadic, and most of them have non-negligible limitations in functions. Specifically, on one hand, most of the existing multi-agent SLAM systems can only extract some key features and build sparse maps. On the other hand, schemes that can reconstruct the environment densely cannot get rid of the dependence on depth sensors, such as RGBD cameras or LiDARs. Systems that can yield high-density maps just with monocular camera suites are temporarily lacking. As an attempt to fill in the research gap to some extent, we design a novel collaborative SLAM system, namely CVIDS (Collaborative Visual-Inertial Dense SLAM), which follows a centralized and loosely coupled framework and can be integrated with any existing Visual-Inertial Odometry (VIO) to accomplish the co-localization and the dense reconstruction. Integrating our proposed robust loop closure detection module and two-stage pose-graph optimization pipeline, the co-localization module of CVIDS can estimate the poses of different agents in a unified coordinate system efficiently from the packed images and local poses sent by the client-ends of different agents. Besides, our motion-based dense mapping module can effectively recover the 3D structures of selected keyframes and then fuse their depth information to the global map for reconstruction. The superior performance of CVIDS is corroborated by both quantitative and qualitative experimental results. To make our results reproducible, the source code has been released at https://cslinzhang.github.io/CVIDS. Tianjun Zhang, Lin Zhang 0014, Yang Chen 0037, Yicong Zhou |
IEEE Trans. Image Process. | 4 |
| 2022 | Self-Paced Enhanced Low-Rank Tensor Kernelized Multi-View Subspace ClusteringabstractThis paper addresses the multi-view subspace clustering problem and proposes the self-paced enhanced low-rank tensor kernelized multi-view subspace clustering (SETKMC) method, which is based on two motivations: (1) singular values of the representations and multiple instances should be treated differently. The reasons are that larger singular values of the representations usually quantify the major information and should be less penalized; samples with different degrees of noise may have various reliability for clustering. (2) many existing methods may cause the degraded performance when multi-view features reside in different nonlinear subspaces. This is because they usually assumed that multiple features lie within the union of several linear subspaces. SETKMC integrates the nonconvex tensor norm, self-paced learning, and kernel trick into a unified model for multi-view subspace clustering. The nonconvex tensor norm imposes different weights on different singular values. The self-paced learning gradually involves instances from more reliable to less reliable ones while the kernel trick aims to handle the multi-view data in nonlinear subspaces. One iterative algorithm is proposed based on the alternating direction method of multipliers. Extensive results on seven real-world datasets show the effectiveness of the proposed SETKMC compared to fifteen state-of-the-art multi-view clustering methods. Yongyong Chen, Shuqin Wang 0001, Xiaolin Xiao, Youfa Liu, Zhongyun Hua, Yicong Zhou |
IEEE Trans. Multim. | 6 |
| 2022 | Adaptive Transition Probability Matrix Learning for Multiview Spectral ClusteringabstractMultiview clustering as an important unsupervised method has been gathering a great deal of attention. However, most multiview clustering methods exploit theself-representation propertyto capture the relationship among data, resulting in high computation cost in calculating the self-representation coefficients. In addition, they usually employ different regularizers to learn the representation tensor or matrix from which a transition probability matrix is constructed in a separate step, such as the one proposed by Wuet al.. Thus, an optimal transition probability matrix cannot be guaranteed. To solve these issues, we propose a unified model for multiview spectral clustering by directly learning an adaptive transition probability matrix (MCA2M), rather than an individual representation matrix of each view. Different from the one proposed by Wuet al., MCA2M utilizes the one-step strategy to directly learn the transition probability matrix under the robust principal component analysis framework. Unlike existing methods using the absolute symmetrization operation to guarantee the nonnegativity and symmetry of the affinity matrix, the transition probability matrix learned from MCA2M is nonnegative and symmetric without any postprocessing. An alternating optimization algorithm is designed based on the efficient alternating direction method of multipliers. Extensive experiments on several real-world databases demonstrate that the proposed method outperforms the state-of-the-art methods. Yongyong Chen, Xiaolin Xiao, Zhongyun Hua, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | PID Controller-Guided Attention Neural Network Learning for Fast and Effective Real Photographs DenoisingabstractReal photograph denoising is extremely challenging in low-level computer vision since the noise is sophisticated and cannot be fully modeled by explicit distributions. Although deep-learning techniques have been actively explored for this issue and achieved convincing results, most of the networks may cause vanishing or exploding gradients, and usually entail more time and memory to obtain a remarkable performance. This article overcomes these challenges and presents a novel network, namely, PID controller guide attention neural network (PAN-Net), taking advantage of both the proportional-integral-derivative (PID) controller and attention neural network for real photograph denoising. First, a PID-attention network (PID-AN) is built to learn and exploit discriminative image features. Meanwhile, we devise a dynamic learning scheme by linking the neural network and control action, which significantly improves the robustness and adaptability of PID-AN. Second, we explore both the residual structure and share-source skip connections to stack the PID-ANs. Such a framework provides a flexible way to feature residual learning, enabling us to facilitate the network training and boost the denoising performance. Extensive experiments show that our PAN-Net achieves superior denoising results against the state-of-the-art in terms of image quality and efficiency. Ruijun Ma 0001, Bob Zhang 0001, Yicong Zhou, Fangyuan Lei |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Domain-invariant Graph for Adaptive Semi-supervised Domain AdaptationabstractDomain adaptation aims to generalize a model from a source domain to tackle tasks in a related but different target domain. Traditional domain adaptation algorithms assume that enough labeled data, which are treated as the prior knowledge are available in the source domain. However, these algorithms will be infeasible when only a few labeled data exist in the source domain, thus the performance decreases significantly. To address this challenge, we propose a Domain-invariant Graph Learning (DGL) approach for domain adaptation with only a few labeled source samples. Firstly, DGL introduces the Nyström method to construct a plastic graph that shares similar geometric property with the target domain. Then, DGL flexibly employs the Nyström approximation error to measure the divergence between the plastic graph and source graph to formalize the distribution mismatch from the geometric perspective. Through minimizing the approximation error, DGL learns a domain-invariant geometric graph to bridge the source and target domains. Finally, we integrate the learned domain-invariant graph with the semi-supervised learning and further propose an adaptive semi-supervised model to handle the cross-domain problems. The results of extensive experiments on popular datasets verify the superiority of DGL, especially when only a few labeled source samples are available. Weifeng Liu 0001, Yicong Zhou, Jun Yu 0002, Dapeng Tao, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Online Correction of Camera Poses for the Surround-view System: A Sparse Direct ApproachabstractThe surround-view module is an indispensable component of a modern advanced driving assistance system. By calibrating the intrinsics and extrinsics of the surround-view cameras accurately, a top-down surround-view can be generated from raw fisheye images. However, poses of these cameras sometimes may change. At present, how to correct poses of cameras in a surround-view system online without re-calibration is still an open issue. To settle this problem, we introduce the sparse direct framework and propose a novel optimization scheme of a cascade structure. This scheme is actually composed of two levels of optimization and two corresponding photometric error based models are proposed. The model for the first-level optimization is called the ground model, as its photometric errors are measured on the ground plane. For the second level of the optimization, it’s based on the so-called ground-camera model, in which photometric errors are computed on the imaging planes. With these models, the pose correction task is formulated as a nonlinear least-squares problem to minimize photometric errors in overlapping regions of adjacent bird’s-eye-view images. With a cascade structure of these two levels of optimization, an appropriate balance between the speed and the accuracy can be achieved. Experiments show that our method can effectively eliminate the misalignment caused by cameras’ moderate pose changes in the surround-view system. Source code and test cases are available online at https://cslinzhang.github.io/CamPoseCorrection/ . Tianjun Zhang, Hao Deng 0002, Lin Zhang 0014, Shengjie Zhao 0001, Xiao Liu 0030, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | Two-Dimensional Parametric Polynomial Chaotic SystemabstractWhen used in engineering applications, most existing chaotic systems may have many disadvantages, including discontinuous chaotic parameter ranges, lack of robust chaos, and easy occurrence of chaos degradation. In this article, we propose a two-dimensional (2-D) parametric polynomial chaotic system (2D-PPCS) as a general system that can yield many 2-D chaotic maps with different exponent coefficient settings. The 2D-PPCS initializes two parametric polynomials and then applies modular chaotification to the polynomials. Setting different control parameters allows the 2D-PPCS to customize its Lyapunov exponents in order to obtain robust chaos and behaviors with desired complexity. Our theoretical analysis demonstrates the robust chaotic behavior of the 2D-PPCS. Two illustrative examples are provided and tested based on numeral experiments to verify the effectiveness of the 2D-PPCS. A chaos-based pseudorandom number generator is also developed to illustrate the applications of the 2D-PPCS. The experimental results demonstrate that these examples of the 2D-PPCS can achieve robust and desired chaos, have better performance, and generate higher randomness pseudorandom numbers than some representative 2-D chaotic maps. Zhongyun Hua, Yongyong Chen, Han Bao 0001, Yicong Zhou |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | Leveraging GANs via Non-local Features
Xuyang Peng, Weifeng Liu 0001, Baodi Liu, Kai Zhang 0029, Yicong Zhou |
ICANN (2) | 6 |
| 2021 | SauvolaNet: Learning Adaptive Sauvola Network for Degraded Document Binarization
Deng Li 0002, Yue Wu 0001, Yicong Zhou |
ICDAR (4) | 3 |
| 2021 | Linecounter: Learning Handwritten Text Line Segmentation By CountingabstractHandwritten Text Line Segmentation (HTLS) is a low-level but important task for many higher-level document processing tasks like handwritten text recognition. It is often formulated in terms of semantic segmentation or object detection in deep learning. However, both formulations have serious shortcomings. The former requires heavy post-processing of splitting/merging adjacent segments, while the latter may fail on dense or curved texts. In this paper, we propose a novel Line Counting formulation for HTLS – that involves counting the number of text lines from the top at every pixel location. This formulation helps learn an end-to-end HTLS solution that directly predicts per-pixel line number for a given document image. Furthermore, we propose a deep neural network (DNN) model LineCounter to perform HTLS through the Line Counting formulation. Our extensive experiments on the three public datasets (ICDAR2013-HSC [1], HIT-MW [2], and VML-AHTE [3]) demonstrate that LineCounter outperforms state-of-the-art HTLS approaches. Source code is available at https://github.com/Leedeng/LineCounter. Deng Li 0002, Yue Wu 0001, Yicong Zhou |
ICIP | 3 |
| 2021 | Tensor-Based Unsupervised Multi-View Feature Selection for Image RecognitionabstractIn image analysis, image samples from multiple sources may contain noisy features. Due to the difficulty of obtaining label information and complex intrinsic structures, performing unsupervised feature selection on multi-view data is a challenging problem. Most existing unsupervised multi-view feature selection methods may explore only the inter-view correlations at the view-level, and ignore the explicit correlations between features across multiple views. In this paper, we propose a tensor-based unsupervised multi-view feature selection (TUFS) method. Specifically, TUFS efficiently explores the full-order interactions among multi-view data without physically building a tensor. Besides, multiple local geometric structures for different views are constructed to facilitate unsupervised feature selection. To solve the proposed model, we design an alternating optimization algorithm. Experiments and comparisons on three image datasets demonstrate that the proposed TUFS yields better performance over the state-of-the-art methods. Yongshan Zhang, Xinxin Wang 0003, Zhihua Cai, Yicong Zhou, Philip S. Yu |
ICME | 4 |
| 2021 | Unified Batch All Triplet Loss for Visible-Infrared Person Re-identificationabstractVisible-Infrared cross-modality person reidentification (VI-ReID), whose aim is to match person images between visible and infrared modality, is a challenging cross-modality image retrieval task. Batch Hard Triplet loss is widely used in person re-identification tasks, but it does not perform well in the Visible-Infrared person re-identification task. Because it only optimizes the hardest triplet for each anchor image within the mini-batch, samples in the hardest triplet may all belong to the same modality, which will lead to the imbalance problem of modality optimization. To address this problem, we adopt the batch all triplet selection strategy, which selects all the possible triplets among samples to optimize instead of the hardest triplet. Furthermore, we introduce Unified Batch All Triplet loss and Cosine Softmax loss to collaboratively optimize the cosine distance between image vectors. Similarly, we modify the Hetero Center Triplet loss, which is proposed for VI-ReID task, into a batch all form to improve model performance. Extensive experiments indicate the effectiveness of the proposed methods, which outperform state-of-the-art methods by a wide margin. Wenkang Li, Ke Qi, Wenbin Chen 0003, Yicong Zhou |
IJCNN | 4 |
| 2021 | Partial Tubal Nuclear Norm Regularized Multi-view LearningabstractMulti-view clustering and multi-view dimension reduction explore ubiquitous and complementary information between multiple features to enhance the clustering, recognition performance. However, multi-view clustering and multi-view dimension reduction are treated independently, ignoring the underlying correlations between them. In addition, previous methods mainly focus on using the tensor nuclear norm for low-rank representation to explore the high correlation of multi-view features, which often causes the estimation bias of the tensor rank. To overcome these limitations, we propose the partial tubal nuclear norm regularized multi-view learning (PTN2ML) method, in which the partial tubal nuclear norm as a non-convex surrogate of the tensor tubal multi-rank, only minimizes the partial sum of the smaller tubal singular values to preserve the low-rank property of the self-representation tensor. PTN2ML pursues the latent representation from the projection space rather than from the input space to reveal the structural consensus and suppress the disturbance of noisy data. The proposed method can be efficiently optimized by the alternating direction method of multipliers. Extensive experiments, including multi-view clustering and multi-view dimension reduction substantiate the superiority of the proposed methods beyond state-of-the-arts. Yongyong Chen, Shuqin Wang 0001, Chong Peng 0001, Guangming Lu 0002, Yicong Zhou |
ACM Multimedia | 5 |
| 2021 | ROECS: A Robust Semi-direct Pipeline Towards Online Extrinsics Correction of the Surround-view SystemabstractGenerally, a surround-view system (SVS), which is an indispensable component of advanced driving assistant systems (ADAS), consists of four to six wide-angle fisheye cameras. As long as both intrinsics and extrinsics of all cameras have been calibrated, a top-down surround-view with the real scale can be synthesized at runtime from fisheye images captured by these cameras. However, when the vehicle is driving on the road, relative poses between cameras in the SVS may change from the initial calibrated states due to bumps or collisions. In case that extrinsics' representations are not adjusted accordingly, on the surround-view, obvious geometric misalignment will appear. Currently, the researches on correcting the extrinsics of the SVS in an online manner are quite sporadic, and a mature and robust pipeline is still lacking. As an attempt to fill this research gap to some extent, in this work, we present a novel extrinsics correction pipeline designed specially for the SVS, namely ROECS (Robust Online Extrinsics Correction of the Surround-view system). Specifically, a "refined bi-camera error" model is firstly designed. Then, by minimizing the overall "bi-camera error" within a sparse and semi-direct framework, the SVS's extrinsics can be iteratively optimized and become accurate eventually. Besides, an innovative three-step pixel selection strategy is also proposed. The superior robustness and the generalization capability of ROECS are validated by both quantitative and qualitative experimental results. To make the results reproducible, the collected data and the source code have been released at https://cslinzhang.github.io/ROECS/. Tianjun Zhang, Brian Nlong Zhao, Ying Shen 0005, Xuan Shao, Lin Zhang 0014, Yicong Zhou |
ACM Multimedia | 6 |
| 2021 | Highly shared Convolutional Neural Networks
Yao Lu 0008, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2021 | Example-feature graph convolutional networks for semi-supervised classification
Sichao Fu, Weifeng Liu 0001, Kai Zhang 0029, Yicong Zhou |
Neurocomputing | 4 |
| 2021 | Human activity recognition by manifold regularization based dynamic graph convolutional networks
Weifeng Liu 0001, Sichao Fu, Yicong Zhou, Zhengjun Zha, Liqiang Nie |
Neurocomputing | 3 |
| 2021 | Semi-supervised classification by graph p-Laplacian convolutional networks
Sichao Fu, Weifeng Liu 0001, Kai Zhang 0029, Yicong Zhou, Dapeng Tao |
Inf. Sci. | 4 |
| 2021 | Unified Cross-domain Classification via Geometric and Statistical Adaptations
Weifeng Liu 0001, Baodi Liu, Weili Guan, Yicong Zhou, Changsheng Xu |
Pattern Recognit. | 5 |
| 2021 | Visually secure image encryption using adaptive-thresholding sparsification and parallel compressive sensing
Zhongyun Hua, Yuanman Li, Yicong Zhou |
Signal Process. | 4 |
| 2021 | Cryptanalysis of Image Ciphers With Permutation-Substitution Network and ChaosabstractIn recent decades, the introduction of chaos to image encryption has drawn worldwide attention. The permutation-substitution architecture has been widely applied, and chaotic systems are generally employed to produce the required encryption elements. Although many security assessment tests have been conducted, some chaotic image ciphers are being cryptanalyzed. In this article, we evaluate the security of a family of image ciphers whose encryption kernel consists of a bit-level or pixel-level permutation and a bit-wise exclusive OR substitution. After investigating the intrinsic linearity inside the outfitted structures and encryption techniques, we find that each ciphertext-plaintext pair can be represented as a combination of a set of ciphertext-plaintext bases. A chosen-ciphertext attack is proposed to construct the ciphertext-plaintext bases rather than the traditional solution to retrieve equivalent encryption elements. We further reveal that such weakness cannot be remedied by common enhancements such as more chaotic dynamics, complex permutation methods, and random pixel insertion during encryption. In addition, applications of the proposed attack to break 12 ciphers are theoretically presented and experimentally verified. Junxin Chen 0001, Lei Chen 0053, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Re-Evaluation of the Security of a Family of Image Diffusion MechanismsabstractIn recent years, the use of permutation-diffusion architecture for digital image encryption has become increasingly popular. The permutation procedure scrambles the pixel locations, while the diffusion phase modifies the pixel values and gives rise to the avalanche effect. Various diffusion techniques have been developed, and their strength strongly impacts the security of the overall cryptosystem. In this paper, we re-evaluate the security of a family of image diffusion mechanisms that are based on mixing modulo addition with bitwise exclusive OR operations. The recovery of the encryption element of these diffusion mechanisms is comprehensively demonstrated, and the accuracy bounds under various conditions are proved mathematically. Compared to the state-of-the-art methods, our work improves the recovery accuracy of the encryption element while the required prior knowledge is decreased. The proposed analysis of the diffusion mechanisms is further used to cryptanalyze the whole cryptosystem theoretically and experimentally. Junxin Chen 0001, Leo Yu Zhang, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Self-Paced Nonnegative Matrix Factorization for Hyperspectral UnmixingabstractThe presence of mixed pixels in the hyperspectral data makes unmixing to be a key step for many applications. Unsupervised unmixing needs to estimate the number of endmembers, their spectral signatures, and their abundances at each pixel. Since both endmember and abundance matrices are unknown, unsupervised unmixing can be considered as a blind source separation problem and can be solved by nonnegative matrix factorization (NMF). However, most of the existing NMF unmixing methods use a least-squares objective function that is sensitive to the noise and outliers. To deal with different types of noises in hyperspectral data, such as the noise in different bands (band noise), the noise in different pixels (pixel noise), and the noise in different elements of hyperspectral data matrix (element noise), we propose three self-paced learning based NMF (SpNMF) unmixing models in this article. The SpNMF models replace the least-squares loss in the standard NMF model with weighted least-squares losses and adopt a self-paced learning (SPL) strategy to learn the weights adaptively. In each iteration of SPL, atoms (bands or pixels or elements) with weight zero are considered as complex atoms and are excluded, while atoms with nonzero weights are considered as easy atoms and are included in the current unmixing model. By gradually enlarging the size of the current model set, SpNMF can select atoms from easy to complex. Usually, noisy or outlying atoms are complex atoms that are excluded from the unmixing model. Thus, SpNMF models are robust to noise and outliers. Experimental results on the simulated and two real hyperspectral data sets demonstrate that our proposed SpNMF methods are more accurate and robust than the existing NMF methods, especially in the case of heavy noise. Jiangtao Peng, Yicong Zhou, Weiwei Sun 0005, Qian Du 0001, Lekang Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Multivariate Time-Series Modeling for Forecasting Sintering Temperature in Rotary Kilns Using DCGNetabstractThe sintering temperature (ST) is a critical index for condition monitoring and process control of coal-fired equipment and is widely used in the production of cement, aluminum, electricity, steel, and chemicals. The accurate prediction of the ST is important for control systems to anticipate tragedies. In this article, we propose a deep learning model for forecasting the ST using automatic spatiotemporal feature extraction from multivariate thermal time series. A hybrid deep neural network named deep convolutional neural network and gated recurrent unit network (DCGNet) is designed to extract multivariate coupling and nonlinear dynamic characteristics for forecasting the ST. DCGNet uses convolutional neural networks and gated recurrent unit (GRU) to extract the local spatial-temporal dependence patterns among the multivariates, and another parallel GRU using the historical ST data as input is incorporated to more accurately capture the dynamic characteristics of ST time series. Based on the real-world data, application results show that the proposed approach has high forecasting accuracy and robustness, thus having broad application prospects in industrial processes. Xiaogang Zhang 0002, Yanying Lei, Hua Chen 0008, Yicong Zhou |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | Generalized Nonconvex Low-Rank Tensor Approximation for Multi-View Subspace ClusteringabstractThe low-rank tensor representation (LRTR) has become an emerging research direction to boost the multi-view clustering performance. This is because LRTR utilizes not only the pairwise relation between data points, but also the view relation of multiple views. However, there is one significant challenge: LRTR uses the tensor nuclear norm as the convex approximation but provides a biased estimation of the tensor rank function. To address this limitation, we propose the generalized nonconvex low-rank tensor approximation (GNLTA) for multi-view subspace clustering. Instead of the pairwise correlation, GNLTA adopts the low-rank tensor approximation to capture the high-order correlation among multiple views and proposes the generalized nonconvex low-rank tensor norm to well consider the physical meanings of different singular values. We develop a unified solver to solve the GNLTA model and prove that under mild conditions, any accumulation point is a stationary point of GNLTA. Extensive experiments on seven commonly used benchmark databases have demonstrated that the proposed GNLTA achieves better clustering performance over state-of-the-art methods. Yongyong Chen, Shuqin Wang 0001, Chong Peng 0001, Zhongyun Hua, Yicong Zhou |
IEEE Trans. Image Process. | 5 |
| 2021 | Low-Rank Preserving t-Linear Projection for Robust Image Feature ExtractionabstractAs the cornerstone for joint dimension reduction and feature extraction, extensive linear projection algorithms were proposed to fit various requirements. When being applied to image data, however, existing methods suffer from representation deficiency since the multi-way structure of the data is (partially) neglected. To solve this problem, we propose a novel Low-Rank Preserving t-Linear Projection (LRP-tP) model that preserves the intrinsic structure of the image data using t-product-based operations. The proposed model advances in four aspects: 1) LRP-tP learns the t-linear projection directly from the tensorial dataset so as to exploit the correlation among the multi-way data structure simultaneously; 2) to cope with the widely spread data errors, e.g., noise and corruptions, the robustness of LRP-tP is enhanced via self-representation learning; 3) LRP-tP is endowed with good discriminative ability by integrating the empirical classification error into the learning procedure; 4) an adaptive graph considering the similarity and locality of the data is jointly learned to precisely portray the data affinity. We devise an efficient algorithm to solve the proposed LRP-tP model using the alternating direction method of multipliers. Extensive experiments on image feature extraction have demonstrated the superiority of LRP-tP compared to the state-of-the-arts. Xiaolin Xiao, Yongyong Chen, Yue-Jiao Gong, Yicong Zhou |
IEEE Trans. Image Process. | 4 |
| 2021 | Simulation of Atmospheric Visibility ImpairmentabstractChanges in aerosol composition and its proportions can cause changes in atmospheric visibility. Vision systems deployed outdoors must take into account the negative effects brought by visibility impairment. In order to develop vision algorithms that can adapt to low atmospheric visibility conditions, a large-scale dataset containing pairs of clear images and their visibility-impaired versions (along with other annotations if necessary) is usually indispensable. However, it is almost impossible to collect large amounts of such image pairs in a real physical environment. A natural and reasonable solution is to use virtual simulation technologies, which is also the focus of this paper. In this paper, we first deeply analyze the limitations and irrationalities of the existing work specializing on simulation of atmospheric visibility impairment. We point out that many simulation schemes actually even violate the assumptions of the Koschmieder's law. Second, more importantly, based on a thorough investigation of the relevant studies in the field of atmospheric science, we present simulation strategies for five most commonly encountered visibility impairment phenomena, including mist, fog, natural haze, smog, and Asian dust. Our work establishes a direct link between the fields of atmospheric science and computer vision. In addition, as a byproduct, with the proposed simulation schemes, a large-scale synthetic dataset is established, comprising 40,000 clear source images and their 800,000 visibility-impaired versions. To make our work reproducible, source codes and the dataset have been released at https://cslinzhang.github.io/AVID/. Lin Zhang 0014, Shiyu Zhao 0001, Yicong Zhou |
IEEE Trans. Image Process. | 4 |
| 2021 | RefineDNet: A Weakly Supervised Refinement Framework for Single Image DehazingabstractHaze-free images are the prerequisites of many vision systems and algorithms, and thus single image dehazing is of paramount importance in computer vision. In this field, prior-based methods have achieved initial success. However, they often introduce annoying artifacts to outputs because their priors can hardly fit all situations. By contrast, learning-based methods can generate more natural results. Nonetheless, due to the lack of paired foggy and clear outdoor images of the same scenes as training samples, their haze removal abilities are limited. In this work, we attempt to merge the merits of prior-based and learning-based approaches by dividing the dehazing task into two sub-tasks, i.e., visibility restoration and realness improvement. Specifically, we propose a two-stage weakly supervised dehazing framework, RefineDNet. In the first stage, RefineDNet adopts the dark channel prior to restore visibility. Then, in the second stage, it refines preliminary dehazing results of the first stage to improve realness via adversarial learning with unpaired foggy and clear images. To get more qualified results, we also propose an effective perceptual fusion strategy to blend different dehazing outputs. Extensive experiments corroborate that RefineDNet with the perceptual fusion has an outstanding haze removal capability and can also produce visually pleasing results. Even implemented with basic backbone networks, RefineDNet can outperform supervised dehazing approaches as well as other state-of-the-art methods on indoor and outdoor datasets. To make our results reproducible, relevant code and data are available at https://github.com/xiaofeng94/RefineDNet-for-dehazing. Shiyu Zhao 0001, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou |
IEEE Trans. Image Process. | 4 |
| 2021 | Universal Chosen-Ciphertext Attack for a Family of Image Encryption SchemesabstractIn recent decades, there has been considerable popularity in employing nonlinear dynamics and permutation-substitution structures for image encryption. Three procedures generally exist in such image encryption schemes: the key schedule module for producing encryption elements, permutation for image scrambling and substitution for pixel modification. This paper cryptanalyzes a family of image encryption schemes that adopt pixel-level permutation and modular addition-based substitution. The security analysis first reveals a common defect in the studied image encryption schemes. Specifically, the mapping from the differentials of the ciphertexts to those of the plaintexts is found to be linear and independent of the key schedules, permutation techniques and encryption rounds. On this theory basis, a universal chosen-ciphertext attack is further proposed. Experimental results demonstrate that the proposed attack can recover the plaintexts of the studied image encryption schemes without a security key or any encryption elements. Related cryptographic discussions are also given. Junxin Chen 0001, Lei Chen 0053, Yicong Zhou |
IEEE Trans. Multim. | 3 |
| 2021 | Orthogonalization-Guided Feature Fusion Network for Multimodal 2D+3D Facial Expression RecognitionabstractAs 2D and 3D data present different views of the same face, the features extracted from them can be both complementary and redundant. In this paper, we present a novel and efficient orthogonalization-guided feature fusion network, namely OGF$^2$Net, to fuse the features extracted from 2D and 3D faces for facial expression recognition. While 2D texture maps are fed into a 2D feature extraction pipeline (FE2DNet), the attribute maps generated from 3D data are concatenated as input of the 3D feature extraction pipeline (FE3DNet). The two networks are separately trained at the first stage and frozen in the second stage for late feature fusion, which can well address the unavailability of a large number of 3D+2D face pairs. To reduce the redundancies among features extracted from 2D and 3D streams, we design an orthogonal loss-guided feature fusion network to orthogonalize the features before fusing them. Experimental results show that the proposed method significantly outperforms the state-of-the-art algorithms on both the BU-3DFE and Bosphorus databases. While accuracies as high as 89.05% (P1 protocol) and 89.07% (P2 protocol) are achieved on the BU-3DFE database, an accuracy of 89.28% is achieved on the Bosphorus database. The complexity analysis also suggests that our approach achieves a higher processing speed while simultaneously requiring lower memory costs. Shisong Lin, Mengchao Bai, Feng Liu 0013, LinLin Shen, Yicong Zhou |
IEEE Trans. Multim. | 5 |
| 2021 | Prior Knowledge Regularized Multiview Self-Representation and its ApplicationsabstractTo learn the self-representation matrices/tensor that encodes the intrinsic structure of the data, existing multiview self-representation models consider only the multiview features and, thus, impose equal membership preference across samples. However, this is inappropriate in real scenarios since the prior knowledge, e.g., explicit labels, semantic similarities, and weak-domain cues, can provide useful insights into the underlying relationship of samples. Based on this observation, this article proposes a prior knowledge regularized multiview self-representation (P-MVSR) model, in which the prior knowledge, multiview features, and high-order cross-view correlation are jointly considered to obtain an accurate self-representation tensor. The general concept of "prior knowledge" is defined as the complement of multiview features, and the core of P-MVSR is to take advantage of the membership preference, which is derived from the prior knowledge, to purify and refine the discovered membership of the data. Moreover, P-MVSR adopts the same optimization procedure to handle different prior knowledge and, thus, provides a unified framework for weakly supervised clustering and semisupervised classification. Extensive experiments on real-world databases demonstrate the effectiveness of the proposed P-MVSR model. Xiaolin Xiao, Yongyong Chen, Yue-Jiao Gong, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Dynamic Graph Learning Convolutional Networks for Semi-supervised ClassificationabstractOver the past few years, graph representation learning (GRL) has received widespread attention on the feature representations of the non-Euclidean data. As a typical model of GRL, graph convolutional networks (GCN) fuse the graph Laplacian-based static sample structural information. GCN thus generalizes convolutional neural networks to acquire the sample representations with the variously high-order structures. However, most of existing GCN-based variants depend on the static data structural relationships. It will result in the extracted data features lacking of representativeness during the convolution process. To solve this problem, dynamic graph learning convolutional networks (DGLCN) on the application of semi-supervised classification are proposed. First, we introduce a definition of dynamic spectral graph convolution operation. It constantly optimizes the high-order structural relationships between data points according to the loss values of the loss function, and then fits the local geometry information of data exactly. After optimizing our proposed definition with the one-order Chebyshev polynomial, we can obtain a single-layer convolution rule of DGLCN. Due to the fusion of the optimized structural information in the learning process, multi-layer DGLCN can extract richer sample features to improve classification performance. Substantial experiments are conducted on citation network datasets to prove the effectiveness of DGLCN. Experiment results demonstrate that the proposed DGLCN obtains a superior classification performance compared to several existing semi-supervised classification models. Sichao Fu, Weifeng Liu 0001, Weili Guan, Yicong Zhou, Dapeng Tao, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Pedestrian-Aware Panoramic Video Stitching Based on a Structured Camera ArrayabstractThe panorama stitching system is an indispensable module in surveillance or space exploration. Such a system enables the viewer to understand the surroundings instantly by aligning the surrounding images on a plane and fusing them naturally. The bottleneck of existing systems mainly lies in alignment and naturalness of the transition of adjacent images. When facing dynamic foregrounds, they may produce outputs with misaligned semantic objects, which is evident and sensitive to human perception. We solve three key issues in the existing workflow that can affect its efficiency and the quality of the obtained panoramic video and present Pedestrian360, a panoramic video system based on a structured camera array (a spatial surround-view camera system). First, to get a geometrically aligned 360○ view in the horizontal direction, we build a unified multi-camera coordinate system via a novel refinement approach that jointly optimizes camera poses. Second, to eliminate the brightness and color difference of images taken by different cameras, we design a photometric alignment approach by introducing a bias to the baseline linear adjustment model and solving it with two-step least-squares. Third, considering that the human visual system is more sensitive to high-level semantic objects, such as pedestrians and vehicles, we integrate the results of instance segmentation into the framework of dynamic programming in the seam-cutting step. To our knowledge, we are the first to introduce instance segmentation to the seam-cutting problem, which can ensure the integrity of the salient objects in a panorama. Specifically, in our surveillance oriented system, we choose the most significant target, pedestrians, as the seam avoidance target, and this accounts for the name Pedestrian360 . To validate the effectiveness and efficiency of Pedestrian360, a large-scale dataset composed of videos with pedestrians in five scenes is established. The test results on this dataset demonstrate the superiority of Pedestrian360 compared to its competitors. Experimental results show that Pedestrian360 can stitch videos at a speed of 12 to 26 fps, which depends on the number of objects in the shooting scene and their frequencies of movements. To make our reported results reproducible, the relevant code and collected data are publicly available at https://cslinzhang.github.io/Pedestrian360-Homepage/ . Lin Zhang 0014, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Exponential Chaotic Model for Generating Robust ChaosabstractRobust chaos is defined as the inexistence of periodic windows and coexisting attractors in the neighborhood of parameter space. This characteristic is desired because a chaotic system with robust chaos can overcome the chaos disappearance caused by parameter disturbance in practical applications. However, many existing chaotic systems fail to consider the robust chaos. This article introduces an exponential chaotic model (ECM) to produce new one-dimensional (1-D) chaotic maps with robust chaos. ECM is a universal framework and can produce many new chaotic maps employing any two 1-D chaotic maps as base and exponent maps. As examples, we present nine chaotic maps produced by ECM, discuss their bifurcation diagrams and prove their robust chaos. Performance evaluations also show that these nine chaotic maps of ECM can obtain robust chaos in a large parameter space. To show the practical applications of ECM, we employ these nine chaotic maps of ECM in secure communication. Simulation results show their superior performance against various channel noise during data transmission. Zhongyun Hua, Yicong Zhou |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | Geometry Constrained Weakly Supervised Object Localization
Weizeng Lu, Xi Jia, Weicheng Xie 0001, LinLin Shen, Yicong Zhou, Jinming Duan 0001 |
ECCV (26) | 5 |
| 2020 | A Study Of Parking-Slot Detection With The Aid Of Pixel-Level Domain AdaptationabstractThe self-parking system is an important component of self-driving vehicles. Such a system needs to detect and locate the parking-slots from surround-view images, and then guide the vehicle to the designated parking-slot. In the real world, the appearances and environmental conditions of parking-slots can be rich and varied. Thus, to train the parking-slot detection model, it is necessary to collect and label a huge quantity of surround-view images covering as many real cases as possible. Such a process is cumbersome and costly, and will be repeated whenever encountering an unseen parking condition that is quite different from the ones covered by existing training set. To this end, in this paper we propose an extensible pipeline, namely FakePS, to assist parking-slot detection model training by making use of synthetic data. Specifically, with FakePS, we can first build various simulated parking scenes and collect labeled surround-view images automatically. Besides, we resort to pixel-level domain adaptation strategies to enhance the realism of the synthetic images using unlabeled real images while preserving their label information. The efficacy of FakePS has been corroborated by experimental results. Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou |
ICME | 6 |
| 2020 | Oecs: Towards Online Extrinsics Correction For The Surround-View SystemabstractA typical surround-view system consists of four fisheye cameras. By performing an offline calibration that determines both the intrinsics and extrinsics of the system, surround-view images can be synthesized at runtime. However, poses of calibrated cameras sometimes may change. In such a case, if cameras' extrinsics are not updated accordingly, observable geometric misalignment will appear in surround-views. Most existing solutions to this problem resort to re-calibration, which is quite cumbersome. Thus, how to correct cameras' extrinsics in an online manner without using re-calibration is still an open issue. In this paper, we attempt to propose a novel solution to this problem and the proposed solution is referred to as “Online Extrinsics Correction for the Surround-view system OECS for short. We first design a Bi-Camera error model, measuring the photometric discrepancy between two corresponding pixels on images captured by two adjacent cameras. Then, by minimizing the system's overall BiCamera error, cameras' extrinsics can be optimized and the optimization is conducted within a sparse direct framework. The efficacy and efficiency of OECS are validated by experiments. Data and source code used in this work are publicly available at https://z619850002.github.io/OECage/. Tianjun Zhang, Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou |
ICME | 6 |
| 2020 | Zero-Shot Restoration of Underexposed Images via Robust Retinex DecompositionabstractUnderexposed images often suffer from serious quality degradation such as poor visibility and latent noise in the dark. Most previous methods for underexposed images restoration ignore the noise and amplify it during stretching contrast. We predict the noise explicitly to achieve the goal of denoising while restoring the underexposed image. Specifically, a novel three-branch convolution neural network, namely RRDNet (short for Robust Retinex Decomposition Network), is proposed to decompose the input image into three components, illumination, reflectance and noise. As an image-specific network, RRDNet doesn't need any prior image examples or prior training. Instead, the weights of RRDNet will be updated by a zero-shot scheme of iteratively minimizing a specially designed loss function. Such a loss function is devised to evaluate the current decomposition of the test image and guide noise estimation. Experiments demonstrate that RRDNet can achieve robust correction with overall naturalness and pleasing visual quality. To make the results reproducible, the source code has been made publicly available at https://aaaaangel.github.io/RRDNet-Homepage. Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou |
ICME | 6 |
| 2020 | Cauchy NMF for Hyperspectral UnmixingabstractNon-negative matrix factorization (NMF) is a classical hyperspectral unmixing model which minimizes the Euclidean distance between the hyperspectral data matrix and its low rank approximation (i.e., the product of endmember matrix and abundance matrix), and it fails when applied to noisy data because the loss function is sensitive to outliers. In this paper, we propose a Cauchy NMF (CauchyNMF) model for hyperspectral unmixing which uses a Cauchy loss function (CLF) to replace the traditional least-squares loss. Compared with the least-squares loss, CLF can penalize the noise term for suppressing the large noise mixed in the real data and thus is much more robust. Experimental results on simulated and real hyperspectral data sets demonstrate that our proposed CauchyNMF method is more accurate and robust than existing NMF methods, especially in the case of heavy noise. Jiangtao Peng, Weiwei Sun 0005, Yicong Zhou |
IGARSS | 4 |
| 2020 | Improved Local Covariance Matrix Representation for Hyperspectral Image ClassificationabstractThis paper proposes a novel spectral-spatial feature representation method for hyperspectral image (HSI) classification. It combines the advantages of adaptive weighted filtering (AWF) and local covariance matrix representation (L-CMR) to make full use of the spatial similarity and correlation among different spectral bands. Specifically, the proposed method first uses the maximum noise fraction (MNF) to reduce the dimensionality of HSI. Then, multiscale AWF (MAWF) is applied to make use of spatial information. N ext' the spectral-spatial features are obtained by calculating the local covariance matrix of the given pixel and its neighbors. Finally, the learned spectral-spatial features of each pixels are fed into support vector machine (SVM) for classification. Experimental results on two publicly available HSI datasets show that the proposed method is superior to several existing methods in terms of both classification accuracy and classification visual effect, especially when the number of training samples is small. Xinyu Zhang 0025, Yantao Wei, Huang Yao, Yicong Zhou |
IGARSS | 4 |
| 2020 | A Tightly-coupled Semantic SLAM System with Visual, Inertial and Surround-view Sensors for Autonomous Indoor ParkingabstractThe semantic SLAM (simultaneous localization and mapping) system is an indispensable module for autonomous indoor parking. Monocular and binocular visual cameras constitute the basic configuration to build such a system. Features used in existing SLAM systems are often dynamically movable, blurred and repetitively textured. By contrast, semantic features on the ground are more stable and consistent in the indoor parking environment. Due to their inabilities to perceive salient features on the ground, existing SLAM systems are prone to tracking loss during navigation. Therefore, a surround-view camera system capturing images from a top-down viewpoint is necessarily called for. To this end, this paper proposes a novel tightly-coupled semantic SLAM system by integrating Visual, Inertial, and Surround-view sensors, VIS SLAM for short, for autonomous indoor parking. In VIS SLAM, apart from low-level visual features and IMU (inertial measurement unit) motion data, parking-slots in surround-view images are also detected and geometrically associated, forming semantic constraints. Specifically, each parking-slot can impose a surround-view constraint that can be split into an adjacency term and a registration term. The former pre-defines the position of each individual parking-slot subject to whether it has an adjacent neighbor. The latter further constrains by registering between each observed parking-slot and its position in the world coordinate system. To validate the effectiveness and efficiency of VIS SLAM, a large-scale dataset composed of synchronous multi-sensor data collected from typical indoor parking sites is established, which is the first of its kind. The collected dataset has been made publicly available at https://cslinzhang.github.io/VISSLAM/. Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Hongyu Li 0001, Yicong Zhou |
ACM Multimedia | 6 |
| 2020 | Cryptanalysis of a DNA-based image encryption scheme
Junxin Chen 0001, Lei Chen 0053, Yicong Zhou |
Inf. Sci. | 3 |
| 2020 | HesGCN: Hessian graph convolutional networks for semi-supervised classification
Sichao Fu, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou, Liqiang Nie |
Inf. Sci. | 4 |
| 2020 | An LBP encoding scheme jointly using quaternionic representation and angular information
Rushi Lan, Huimin Lu 0001, Yicong Zhou, Zhenbing Liu |
Neural Comput. Appl. | 3 |
| 2020 | Domain Adaptation with Few Labeled Source Samples by Graph Regularization
Weifeng Liu 0001, Yicong Zhou, Dapeng Tao, Liqiang Nie |
Neural Process. Lett. | 3 |
| 2020 | Multi-view subspace clustering via simultaneously learning the representation tensor and affinity matrix
Yongyong Chen, Xiaolin Xiao, Yicong Zhou |
Pattern Recognit. | 3 |
| 2020 | Designing a 2D infinite collapse map for image encryption
Weijia Cao, Yujun Mao, Yicong Zhou |
Signal Process. | 3 |
| 2020 | Hyperspectral image denoising by total variation-regularized bilinear factorization
Yongyong Chen, Jiaxue Li, Yicong Zhou |
Signal Process. | 3 |
| 2020 | Unsupervised quaternion model for blind colour image quality assessment
Leyuan Wu, Xiaogang Zhang 0002, Hua Chen 0008, Yicong Zhou |
Signal Process. | 4 |
| 2020 | Guest Editorial Introduction to Special Section on Modern Reversible Data Hiding and WatermarkingabstractThe rapid development and growing of 4G and 5G mobile networks allow people all over the world to efficiently transmit and share data and information while bringing an increasing demand on information security. Data hiding is a general technique to embed secret messages to be protected in an imperceptible way into a cover media like an image, a video stream, or a document. Traditional data hiding intends to achieve high embedding capacity and imperceptibility of hidden secret messages for secure communication. However, it introduces permanent damage or distortion to the cover media when receiver extracts the secret messages from the cover media. To address this problem, reversible data hiding (RDH) was developed to allow the receiver completely extract the hidden secret messages while fully recovering the original cover media without any distortion. RDH has been widely used for many military and medical applications like the access authentication of reconnaissance images and the sharing of medical images in remote diagnosis. According to the format of cover media, RDH can be done in both plaintext and encryption domains. RDH in plaintext domain intends to embed the secret messages into the original cover media in a way that the marked media (the cover media with embedded secret messages) is visually the same as the original cover media and able to withstand the potential analysis. RDH in the encryption domain embeds the secret messages into the encrypted cover media (e.g., encrypted images) such that secret messages and cover media are protected in a high security level during transmission and completely reconstructed in the receiver side. Xiaochun Cao, Yicong Zhou, Jing-Ming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Prior Knowledge-Based Probabilistic Collaborative Representation for Visual RecognitionabstractCollaborative representation is an effective way to design classifiers for many practical applications. In this paper, we propose a novel classifier, called the prior knowledge-based probabilistic collaborative representation-based classifier (PKPCRC), for visual recognition. Compared with existing classifiers which use the collaborative representation strategy, the proposed PKPCRC further includes characteristics of training samples of each class as prior knowledge. Four types of prior knowledge are developed from the perspectives of image distance and representation capacity. They adaptively accommodate the contribution of each class and result in an accurate representation to classify a query sample. Experiments and comparisons on four challenging databases demonstrate that PKPCRC outperforms several state-of-the-art classifiers. Rushi Lan, Yicong Zhou, Zhenbing Liu |
IEEE Trans. Cybern. | 2 |
| 2020 | Two-Dimensional Sine Chaotification System With Hardware ImplementationabstractChaotic systems are widely employed in many practical applications for their significant properties. Existing chaotic systems may suffer from the drawbacks of discontinuous chaotic ranges and frail chaotic behaviors. To solve this issue, this paper proposes a two-dimensional (2D) sine chaotification system (2D-SCS). 2D-SCS can not only significantly enhance the complexity of 2D chaotic maps, but also greatly extend their chaotic ranges. As examples, this paper applies 2D-SCS to two existing 2D chaotic maps to obtain two enhanced chaotic maps. Performance evaluations show that these two enhanced chaotic maps have robust chaotic behaviors in much larger chaotic ranges than existing 2D chaotic maps. A microcontroller-based experiment platform is also designed to implement these enhanced chaotic maps in hardware devices. Furthermore, to investigate the application of 2D-SCS, these two enhanced chaotic maps are applied to design a pseudorandom number generator. Experiment results show that these enhanced chaotic maps can produce better random sequences than the existing 2D and several state-of-the-art one-dimensional (1D) chaotic maps. Zhongyun Hua, Yicong Zhou, Bocheng Bao |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Low-Rank Quaternion Approximation for Color Image ProcessingabstractLow-rank matrix approximation (LRMA)-based methods have made a great success for grayscale image processing. When handling color images, LRMA either restores each color channel independently using the monochromatic model or processes the concatenation of three color channels using the concatenation model. However, these two schemes may not make full use of the high correlation among RGB channels. To address this issue, we propose a novel low-rank quaternion approximation (LRQA) model. It contains two major components: first, instead of modeling a color image pixel as a scalar in conventional sparse representation and LRMA-based methods, the color image is encoded as a pure quaternion matrix, such that the cross-channel correlation of color channels can be well exploited; second, LRQA imposes the low-rank constraint on the constructed quaternion matrix. To better estimate the singular values of the underlying low-rank quaternion matrix from its noisy observation, a general model for LRQA is proposed based on several nonconvex functions. Extensive evaluations for color image denoising and inpainting tasks verify that LRQA achieves better performance over several state-of-the-art sparse representation and LRMA-based methods in terms of both quantitative metrics and visual quality. Yongyong Chen, Xiaolin Xiao, Yicong Zhou |
IEEE Trans. Image Process. | 3 |
| 2020 | Multipatch Unbiased Distance Non-Local Adaptive Means With Wavelet ShrinkageabstractMany existing non-local means (NLM) methods either use Euclidean distance to measure the similarity between patches, or compute weight ωijonly once and keep it unchanged during the subsequent denoising iterations, or use only the structure information of the denoised image to update weight ωij. These may lead to the limited denoising performance. To address these issues, this paper proposes the non-local adaptive means (NLAM) for image denoising. NLAM treats weight ωijas an optimization variable and iteratively updates its value. We then introduce three unbiased distances, namely, pixel-pixel, patch- patch, and coupled unbiased distances. These unbiased distances are more robust to measure the image pixel/patch similarity than Euclidean distance. Using the coupled unbiased distance, we propose the unbiased distance non-local adaptive means (UD-NLAM). Because UD-NLAM uses only a single patch size to compute weight ωij, we introduce multipatch UD-NLAM (MUD-NLAM) to adapt different noise levels. To further improve denoising performance, we then propose a new denoising method called MUD-NLAM with wavelet shrinkage (MUD-NLAM-WS). Experimental results show that the proposed NLAM, UD-NLAM, and MUD-NLAM outperform existing NLM methods, and MUDNLAM-WS achieves a better performance than the state-of-theart denoising methods. Yicong Zhou, Lianhong Wang |
IEEE Trans. Image Process. | 2 |
| 2020 | 2D Quaternion Sparse Discriminant AnalysisabstractLinear discriminant analysis has been incorporated with various representations and measurements for dimension reduction and feature extraction. In this paper, we propose two-dimensional quaternion sparse discriminant analysis (2D-QSDA) that meets the requirements of representing RGB and RGB-D images. 2D-QSDA advances in three aspects: 1) including sparse regularization, 2D-QSDA relies only on the important variables, and thus shows good generalization ability to the out-of-sample data which are unseen during the training phase; 2) benefited from quaternion representation, 2D-QSDA well preserves the high order correlation among different image channels and provides a unified approach to extract features from RGB and RGB-D images; 3) the spatial structure of the input images is retained via the matrix-based processing. We tackle the constrained trace ratio problem of 2D-QSDA by solving a corresponding constrained trace difference problem, which is then transformed into a quaternion sparse regression (QSR) model. Afterward, we reformulate the QSR model to an equivalent complex form to avoid the processing of the complicated structure of quaternions. A nested iterative algorithm is designed to learn the solution of 2D-QSDA in the complex space and then we convert this solution back to the quaternion domain. To improve the separability of 2D-QSDA, we further propose 2D-QSDAw using the weighted pairwise between-class distances. Extensive experiments on RGB and RGB-D databases demonstrate the effectiveness of 2D-QSDA and 2D-QSDAw compared with peer competitors. Xiaolin Xiao, Yongyong Chen, Yue-Jiao Gong, Yicong Zhou |
IEEE Trans. Image Process. | 4 |
| 2020 | Jointly Learning Kernel Representation Tensor and Affinity Matrix for Multi-View ClusteringabstractMulti-view clustering refers to the task of partitioning numerous unlabeled multimedia data into several distinct clusters using multiple features. In this paper, we propose a novel nonlinear method called joint learning multi-view clustering (JLMVC) to jointly learn kernel representation tensor and affinity matrix. The proposed JLMVC has three advantages: (1) unlike existing low-rank representation-based multi-view clustering methods that learn the representation tensor and affinity matrix in two separate steps, JLMVC jointly learns them both; (2) using the “kernel trick,” JLMVC can handle nonlinear data structures for various real applications; and (3) different from most existing methods that treat representations of all views equally, JLMVC automatically learns a reasonable weight for each view. Based on the alternating direction method of multipliers, an effective algorithm is designed to solve the proposed model. Extensive experiments on eight multimedia datasets demonstrate the superiority of the proposed JLMVC over state-of-the-art methods. Yongyong Chen, Xiaolin Xiao, Yicong Zhou |
IEEE Trans. Multim. | 3 |
| 2019 | Color Image Denoising Using Quaternion Adaptive Non-Local Coupled MeansabstractBecause quaternion representation is able to preserve the relationship among RGB channels of a color image, we take advantage of this characteristic and propose a new color image denoising method, named the Quaternion Adaptive Non-local Coupled Means (QANLCM). QANLCM first builds a four-variable optimization model by decomposing the original noisy image into a luminance component and a chromaticity component, and then alternatively updates the four variables in each denoising iteration. Simulations and comparisons demonstrate that QANLCM shows a superiority to other non-local means methods in removing Gaussian noise from color images. Yicong Zhou |
ICIP | 2 |
| 2019 | Multi-view Clustering via Simultaneously Learning Graph Regularized Low-Rank Tensor Representation and Affinity MatrixabstractLow-rank tensor representation-based multi-view clustering has become an efficient method for data clustering due to the robustness to noise and the preservation of the high order correlation. However, existing algorithms may suffer from two common problems: (1) the local view-specific geometrical structures and the various importance of features in different views are neglected; (2) the low-rank representation tensor and the affinity matrix are learned separately. To address these issues, we propose a novel framework to learn the Graph regularized Low-rank Tensor representation and the Affinity matrix (GLTA) in a unified manner. Besides, the manifold regularization is exploited to preserve the view-specific geometrical structures, and the various importance of different features is automatically calculated when constructing the final affinity matrix. An efficient algorithm is designed to solve GLTA using the augmented Lagrangian multiplier. Extensive experiments on six real datasets demonstrate the superiority of GLTA over the state-of-the-arts. Yongyong Chen, Xiaolin Xiao, Yicong Zhou |
ICME | 3 |
| 2019 | Quaternion Non-local Total Variation for Color Image DenoisingabstractMany existing color image denoising methods process color channels individually and fail to consider their cross-channel correlations. To solve this problem, in this paper, we employ the quaternion representation of the color image and propose a novel Quaternion Non-local Total Variation (QNLTV) model to remove Gaussian noise from color images. We first introduce the coupled quaternion distance to measure the color image patch similarity. Decomposing the color image into brightness and chromaticity components in quaternion domain, we then divide the QNLTV model into two quaternion optimizaiton problems and solve them alternatively. Experiment results show that QNLTV has the significantly better denoising performance than competing methods in terms of visual and quantitative evaluations. Yicong Zhou |
SMC | 2 |
| 2019 | Cross-Channel Similarity based Histograms of Oriented Gradients for Color ImagesabstractThe local features have gained prestige as the powerful descriptors, however, when handling color images, most existing descriptors fail to find an efficient strategy to make full use of channel correlation. To tackle this problem, we propose the cross-channel similarity based histograms of oriented gradients (CCS-HOG) model for color images. Different from the existing methods, our model integrates the color correlation with the structure features together in a better way. To find out the inner connection between channels, the cross-channel similarity measure is developed as a suitable approach. Experimental results for face recognition and kinship verification illustrate the the performance of CCS-HOG superior to other state-of-the-arts descriptors. Yicong Zhou |
SMC | 2 |
| 2019 | Two-order graph convolutional networks for semi-supervised classificationabstractCurrently, deep learning (DL) algorithms have achieved great success in many applications including computer vision and natural language processing. Many different kinds of DL models have been reported, such as DeepWalk, LINE, diffusionconvolutional neural networks, graph convolutional networks (GCN), and so on. The GCN algorithm is a variant of convolutional neural network and achieves significant superiority by using a one‐order localised spectral graph filter. However, only a one‐order polynomial in the Laplacian of GCN has been approximated and implemented, which ignores undirect neighbour structure information. The lack of rich structure information reduces the performance of the neural networks in the graph structure data. In this study, the authors deduce and simplify the formula of two‐order spectral graph convolutions to preserve rich local information. Furthermore, they build a layerwise GCN based on this two‐order approximation, i.e. two‐order GCN (TGCN) for semi‐supervised classification. With the two‐order polynomial in the Laplacian, the proposed TGCN model can assimilate abundant localised structure information of graph data and then boosts the classification significantly. To evaluate the proposed solution, extensive experiments are conducted on several popular datasets including the Citeseer, Cora, and PubMed dataset. Experimental results demonstrate that the proposed TGCN outperforms the state‐of‐art methods. Sichao Fu, Weifeng Liu 0001, Yicong Zhou |
IET Image Process. | 4 |
| 2019 | HpLapGCN: Hypergraph p-Laplacian graph convolutional networks
Sichao Fu, Weifeng Liu 0001, Yicong Zhou, Liqiang Nie |
Neurocomputing | 3 |
| 2019 | Local polynomial contrast binary patterns for face recognition
Yinyan Jiang, Yicong Zhou, Weifeng Li 0001, Qingmin Liao |
Neurocomputing | 4 |
| 2019 | Cosine-transform-based chaotic system for image encryptionabstractChaos is known as a natural candidate for cryptography applications owing to its properties such as unpredictability and initial state sensitivity. However, certain chaos-based cryptosystems have been proven to exhibit various security defects because their used chaotic maps do not have complex dynamical behaviors. To address this problem, this paper introduces a cosine-transform-based chaotic system (CTBCS). Using two chaotic maps as seed maps, the CTBCS can produce chaotic maps with complex dynamical behaviors. For illustration, we produce three chaotic maps using the CTBCS and analyze their chaos complexity. Using one of the generated chaotic maps, we further propose an image encryption scheme. The encryption scheme uses high-efficiency scrambling to separate adjacent pixels and employs random order substitution to spread a small change in the plain-image to all pixels of the cipher-image. The performance evaluation demonstrates that the chaotic maps generated by the CTBCS exhibit substantially more complicated chaotic behaviors than the existing ones. The simulation results indicate the reliability of the proposed image encryption scheme. Moreover, the security analysis demonstrates that the proposed image encryption scheme provides a higher level of security than several advanced image encryption schemes. Zhongyun Hua, Yicong Zhou, Hejiao Huang |
Inf. Sci. | 2 |
| 2019 | Hessian-Regularized Multitask Dictionary Learning for Remote Sensing Image RecognitionabstractLearning effective image representations is a vital issue for remote sensing (RS) image recognition tasks. Although numerous algorithms have been proposed, it is still challenging due to the limited labeled data. One representative work is the Laplacian-regularized multitask dictionary learning (LR-MTDL) that employs graph Laplacian regularization terms to fully utilize both the labeled and unlabeled information. However, it probably conduces to poor extrapolating power because Laplacian regularization biases the solution toward a constant function. In this letter, we propose a Hessian-regularized multitask dictionary learning to learn a source-data set-shared but target-data set-biased representation for RS image recognition. Particularly, Hessian can properly exploit the intrinsic local geometry of the data manifold and finally leverage the performance. Extensive experiments on four RS image data sets validate the effectiveness of the proposed method by comparing with baseline algorithms including single-task dictionary learning and LR-MTDL. Guanhua Feng, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2019 | Hessian Regularized Distance Metric Learning for People Re-Identification
Guanhua Feng, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou |
Neural Process. Lett. | 4 |
| 2019 | Kernel modified optimal margin distribution machine for imbalanced data classification
Xiaogang Zhang 0002, Dingxiang Wang, Yicong Zhou, Hua Chen 0008, Fanyong Cheng |
Pattern Recognit. Lett. | 3 |
| 2019 | $p$ -Laplacian Regularization for Scene RecognitionabstractThe explosive growth of multimedia data on the Internet makes it essential to develop innovative machine learning algorithms for practical applications especially where only a small number of labeled samples are available. Manifold regularized semi-supervised learning (MRSSL) thus received intensive attention recently because it successfully exploits the local structure of data distribution including both labeled and unlabeled samples to leverage the generalization ability of a learning model. Although there are many representative works in MRSSL, including Laplacian regularization (LapR) and Hessian regularization, how to explore and exploit the local geometry of data manifold is still a challenging problem. In this paper, we introduce a fully efficient approximation algorithm of graph p -Laplacian, which significantly saving the computing cost. And then we propose p -LapR (pLapR) to preserve the local geometry. Specifically, p -Laplacian is a natural generalization of the standard graph Laplacian and provides convincing theoretical evidence to better preserve the local structure. We apply pLapR to support vector machines and kernel least squares and conduct the implementations for scene recognition. Extensive experiments on the Scene 67 dataset, Scene 15 dataset, and UC-Merced dataset validate the effectiveness of pLapR in comparison to the conventional manifold regularization methods. Weifeng Liu 0001, Xueqi Ma, Yicong Zhou, Dapeng Tao, Jun Cheng 0002 |
IEEE Trans. Cybern. | 3 |
| 2019 | Hypergraph $p$ -Laplacian Regularization for Remotely Sensed Image RecognitionabstractGraph-based and manifold-regularization (MR)-based semisupervised learning, including Laplacian regularization (LapR) and hypergraph LapR (HLapR), have achieved prominent performance in preserving locality and similarity information. However, it is still a great challenge to exactly explore and exploit the local structure of the data distribution. In this paper, we present an efficient and effective approximation algorithm of hypergraph${p}$-Laplacian and then propose hypergraph${p}$-LapR (HpLapR) to preserve the geometry of the probability distribution. In particular, hypergraph is a generalization of a standard graph while hypergraph${p}$-Laplacian is a nonlinear generalization of the standard graph Laplacian. The proposed HpLapR shows great potential to exploit the local structures. We integrate HpLapR with logistic regression for remote sensing image recognition. Experiments on UC-Merced data set demonstrate that the proposed HpLapR has superior performance compared with several popular MR methods including LapR and HLapR. Xueqi Ma, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | RGB-'D' Saliency Detection With Pseudo DepthabstractRecent studies have shown the effectiveness of using depth information in salient object detection. However, the most commonly seen images so far are still RGB images that do not contain the depth data. Meanwhile, the human brain can extract the geometric model of a scene from an RGB-only image and hence provides a 3D perception of the scene. Inspired by this observation, we propose a new concept named RGB-'D' saliency detection, which derives pseudo depth from the RGB images and then performs 3D saliency detection. The pseudo depth can be utilized as image features, prior knowledge, an additional image channel, or independent depth-induced models to boost the performance of traditional RGB saliency models. As an illustration, we develop a new salient object detection algorithm that uses the pseudo depth to derive a depth-driven background prior and a depth contrast feature. Extensive experiments on several standard databases validate the promising performance of the proposed algorithm. In addition, we also adapt two supervised RGB saliency models to our RGB-'D' saliency framework for performance enhancement. The results further demonstrate the generalization ability of the proposed RGB-'D' saliency framework. Xiaolin Xiao, Yicong Zhou, Yue-Jiao Gong |
IEEE Trans. Image Process. | 2 |
| 2019 | Separable and Reversible Data Hiding in Encrypted Images Using Parametric Binary Tree LabelingabstractThis paper first introduces a parametric binary tree labeling scheme (PBTL) to label image pixels in two different categories. Using PBTL, a data embedding method (PBTL-DE) is proposed to embed secret data to an image by exploiting spatial redundancy within small image blocks. We then apply PBTL-DE into the encrypted domain and propose a PBTL-based reversible data hiding method in encrypted images (PBTL-RDHEI). PBTL- RDHEI is a separable and reversible method that both the original image and secret data can be recovered and extracted losslessly and independently. Experiment results and analysis show that PBTL- RDHEI is able to achieve an average embedding rate as large as 1.752 bpp and 2.003 bpp when block size is set to 2 × 2 and 3 × 3, respectively. Yicong Zhou |
IEEE Trans. Multim. | 2 |
| 2019 | Two-Dimensional Quaternion PCA and Sparse PCAabstractBenefited from quaternion representation that is able to encode the cross-channel correlation of color images, quaternion principle component analysis (QPCA) was proposed to extract features from color images while reducing the feature dimension. A quaternion covariance matrix (QCM) of input samples was constructed, and its eigenvectors were derived to find the solution of QPCA. However, eigen-decomposition leads to the fixed solution for the same input. This solution is susceptible to outliers and cannot be further optimized. To solve this problem, this paper proposes a novel quaternion ridge regression (QRR) model for two-dimensional QPCA (2D-QPCA). We mathematically prove that this QRR model is equivalent to the QCM model of 2D-QPCA. The QRR model is a general framework and is flexible to combine 2D-QPCA with other technologies or constraints to adapt different requirements of real-world applications. Including sparsity constraints, we then propose a quaternion sparse regression model for 2D-QSPCA to improve its robustness for classification. An alternating minimization algorithm is developed to iteratively learn the solution of 2D-QSPCA in the equivalent complex domain. In addition, 2D-QPCA and 2D-QSPCA can preserve the spatial structure of color images and have a low computation cost. Experiments on several challenging databases demonstrate that 2D-QPCA and 2D-QSPCA are effective in color face recognition, and 2D-QSPCA outperforms the state of the arts. Xiaolin Xiao, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Discriminant Projection Representation-Based Classification for Vision Recognition
Qingxiang Feng, Yicong Zhou |
AAAI | 2 |
| 2018 | Robust Principal Component Analysis with Matrix FactorizationabstractTraditional robust principle component analysis (RPCA) has a high computational cost because RPCA needs to calculate the singular value decomposition of large matrices. To address this issue, this paper proposes a matrix-factorization-based RPCA (MFRPCA) model. MFRPCA has high computation efficiency while improving the robustness and flexibility of traditional RPCA using a non-convex low-rank approximation. Experiment results on challenging datasets demonstrate superior performance of MFRPCA compared with several advanced low-rank reconstruction methods. Yongyong Chen, Yicong Zhou |
ICASSP | 2 |
| 2018 | Unbiased Distance Based Non-Local Fuzzy MeansabstractThis paper first introduces several unbiased distances to measure the similarity between image pixels and between image patches. Based on these unbiased distances, we propose a new non-local means denoising method, named Unbiased Distance based Non-local Fuzzy Means (UDNLFM). UDNLFM considers the weight as a fuzzy variable and updates its value in each denoising iteration. Experiments and comparisons demonstrate that UDNLFM outperforms state-of-the-art non-local means methods for image denoising. Yicong Zhou, Lianhong Wang |
ICASSP | 2 |
| 2018 | Two-Dimensional Quaternion Sparse Principle Component AnalysisabstractMotivated by the facts that, (1), the spatial structure of images and the correlation among color channels are important for color face recognition, and (2), natural face images may be occluded, in this work, we propose two-dimensional quaternion sparse principle component analysis (2DQSPCA) to extract features for color face recognition. 2DQSPCA inherently takes the advantage of 2DPCA in preserving the structure of two-dimensional data, as well as the strength of quaternion-s in representing color images holistically. Benefited from the sparsity constraints, 2DQSPCA is robust for occlusions. Experiments demonstrate the superior performance of 2DQSP-CA on color face recognition, especially with occlusions. Xiaolin Xiao, Yicong Zhou |
ICASSP | 2 |
| 2018 | Quaternion Sparse Discriminant Analysis for Color Face RecognitionabstractTo reduce feature dimensions while obtaining robust classification, in this paper, we propose quaternion sparse discriminant analysis (QSDA) for color face recognition. QSDA is formulated as a quaternion sparse regression-type model. It employs the quaternion algebra to provide an elegant and holistic way to represent color face images. The succeeding operations are directly applied to two-dimensional quaternion matrices, and hence QSDA is computationally efficient and well preserves the spatial structure of color face images. Benefited from sparsity constraints, QSDA is robust for classification. An alternating minimization algorithm is designed to solve QSDA. Experimental results demonstrate the effectiveness of QSDA for color face recognition, especially for partially occluded color face images. Xiaolin Xiao, Yicong Zhou |
ICME | 2 |
| 2018 | Total Variation Regularized Low-Rank Tensor Approximation for Color Image DenoisingabstractExisting approaches for low-rank approximation either need a rank prior or ignore the spatial smooth characteristic of a color image. To overcome these drawbacks, we propose a total variation regularized low-rank tensor approximation model for color image denoising. The model integrates the strong low-rank prior into a tensor-SVD framework, and introduces the hyper total variation to model the spatial smooth structure of images. Using the alternating direction method of multipliers, we propose a simple algorithm to solve our model. Extensive results on simulated and real noisy color images demonstrate the better performance of the proposed method against state-of-the-art denoising methods. Yongyong Chen, Yicong Zhou |
SMC | 2 |
| 2018 | Reversible Data Hiding in Encrypted Images Using Prediction-Error EncodingabstractIn this paper, we propose a reversible data hiding method in encrypted images (RDHEI) using prediction-error encoding (PE-RDHEI). It uses a weighted checkerboard based prediction to predict 3/4 of the pixels in an original image. The obtained prediction-error values and the unmodified pixels are encrypted separately. The data hider then embed secret data into the encrypted prediction-error values using the prediction-error encoding method. At the receiver side, the secret data and original image can be completely extracted and recovered. Compared with existing RDHEI methods, PERDHEI significantly improves the embedding rate. Experimental results are provided to show the excellent performance of our proposed algorithm. Yicong Zhou |
SMC | 2 |
| 2018 | Medical image encryption using high-speed scrambling and pixel adaptive diffusion
Zhongyun Hua, Yicong Zhou |
Signal Process. | 3 |
| 2018 | Parametric reversible data hiding in encrypted images using adaptive bit-level data embedding and checkerboard based prediction
Yicong Zhou |
Signal Process. | 2 |
| 2018 | Reversible data hiding in encrypted images using adaptive block-level prediction-error expansion
Yicong Zhou, Zhongyun Hua |
Signal Process. Image Commun. | 2 |
| 2018 | Designing Hyperchaotic Cat Maps With Any Desired Number of Positive Lyapunov ExponentsabstractGenerating chaotic maps with expected dynamics of users is a challenging topic. Utilizing the inherent relation between the Lyapunov exponents (LEs) of the Cat map and its associated Cat matrix, this paper proposes a simple but efficient method to construct an -dimensional ( -D) hyperchaotic Cat map (HCM) with any desired number of positive LEs. The method first generates two basic -D Cat matrices iteratively and then constructs the final -D Cat matrix by performing similarity transformation on one basic -D Cat matrix by the other. Given any number of positive LEs, it can generate an -D HCM with desired hyperchaotic complexity. Two illustrative examples of -D HCMs were constructed to show the effectiveness of the proposed method, and to verify the inherent relation between the LEs and Cat matrix. Theoretical analysis proves that the parameter space of the generated HCM is very large. Performance evaluations show that, compared with existing methods, the proposed method can construct -D HCMs with lower computation complexity and their outputs demonstrate strong randomness and complex ergodicity. Zhongyun Hua, Yicong Zhou, Chengqing Li, Yue Wu 0001 |
IEEE Trans. Cybern. | 3 |
| 2018 | Differential Evolutionary Superpixel SegmentationabstractSuperpixel segmentation has been of increasing importance in many computer vision applications recently. To handle the problem, most state-of-the-art algorithms either adopt a local color variance model or a local optimization algorithm. This paper develops a new approach, named differential evolutionary superpixels, which is able to optimize the global properties of segmentation by means of a global optimizer. We design a comprehensive objective function aggregating within-superpixel error, boundary gradient, and a regularization term. Minimizing the within-superpixel error enforces the homogeneity of superpixels. In addition, the introduction of boundary gradient drives the superpixel boundaries to capture the natural image boundaries, so as to make each superpixel overlaps with a single object. The regularizer further encourages producing similarly sized superpixels that are friendly to human vision. The optimization is then accomplished by a powerful global optimizer-differential evolution. The algorithm constantly evolves the superpixels by mimicking the process of natural evolution, while using a linear complexity to the image size. Experimental results and comparisons with eleven state-of-the-art peer algorithms verify the promising performance of our algorithm. Yue-Jiao Gong, Yicong Zhou |
IEEE Trans. Image Process. | 2 |
| 2018 | Content-Adaptive Superpixel SegmentationabstractSuperpixel segmentation targets at grouping pixels in an image into atomic regions whose boundaries align well with the natural object boundaries. This paper first proposes a new feature representation for superpixel segmentation that holistically embraces color, contour, texture, and spatial features. Then, we introduce a clustering-based discriminability measure to iteratively evaluate the importance of different features. Integrating the feature representation and the discriminability measure, we propose a novel content-adaptive superpixel (CAS) segmentation algorithm. CAS is able to automatically and iteratively adjust the weights of different features to fit various properties of image instances. Experiments on several challenging datasets demonstrate that the proposed CAS outperforms the state-of-the-art methods and has a low computational cost. Xiaolin Xiao, Yicong Zhou, Yue-Jiao Gong |
IEEE Trans. Image Process. | 2 |
| 2018 | Learning Multimodal Parameters: A Bare-Bones Niching Differential Evolution ApproachabstractMost learning methods contain optimization as a substep, where the nondifferentiability and multimodality of objectives push forward the interplay of evolutionary optimization algorithms and machine learning models. The recently emerged evolutionary multimodal optimization (MMOP) technique enables the learning of diverse sets of effective parameters for the models simultaneously, providing new opportunities to the applications requiring both accuracy and diversity, such as ensemble, interactive, and interpretive learning. Targeting at locating multiple optima simultaneously in the multimodal landscape, this paper develops an efficient neighborhood-based niching algorithm. Bare-bones differential evolution is used as the baseline. Further, using Gaussian mutation with local mean and standard deviations, the neighborhoods capture niches that match well with the contours of peaks in the landscape. To increase diversity and enhance global exploration, the proposed algorithm embeds a diversity preserving operator to reinitialize converged or overlapped neighborhoods. The experimental results verify that the proposed algorithm has superior and consistent performance for a wide range of MMOP problems. Further, the algorithm has been successfully applied to train neural network ensembles, which validates its effectiveness and benefits of learning multimodal parameters. Yue-Jiao Gong, Jun Zhang 0003, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Adaptive superpixel segmentation aggregating local contour and texture featuresabstractSuperpixel segmentation targets at grouping pixels in an image into atomic regions that align well with the natural object boundaries. In this paper, we propose a novel superpixel segmentation method based on an iterative and adaptive clustering algorithm that embraces color, contour, texture, and spatial features together. The algorithm adjusts the weights of different features automatically in a content-aware way, so as to fit the requirements of various image instances. More specifically, in each iteration, the weights in the aggregation function are adjusted according to the discriminabilities of features in the current working scenario. This way, the algorithm not only possesses improved robustness but also relieves the burden of setting the parameters manually. Experimental verification shows that the algorithm outperforms existing peer algorithms in terms of commonly used evaluation metrics, while using a low computational cost. Xiaolin Xiao, Yue-Jiao Gong, Yicong Zhou |
ICASSP | 3 |
| 2017 | Adaptive code embedding for reversible data hiding in encrypted imagesabstractIn this paper, we propose an adaptive code embedding method for reversible data hiding in encrypted images (ACE-RDHEI). It encrypts the original image into a noise-like one by block permutation and modulation, and no reserving room process is needed before image encryption. The secret data is then embedded into the encrypted image using adaptive code embedding (ACE). At the receiver side, using different security keys, users are able to extract the secret data and recover the original image separately and losslessly. Compared with existing RDHEI methods, ACE-RDHEI significantly improves the embedding rate. Experimental results are provided to show the excellent performance of our proposed algorithm. Yicong Zhou |
ICIP | 2 |
| 2017 | An extended probabilistic collaborative representation based classifier for image classificationabstractCollaborative representation based classifier (CRC) and its probabilistic improvement ProCRC have achieved satisfactory performance in many image classification applications. They, however, do not comprehensively take account of the structure characteristics of the training samples. In this paper, we present an extended probabilistic collaborative representation based classifier (EProCRC) for image classification. Compared with CRC and ProCRC, the proposed EProCRC further considers a prior information that describes the distribution of each class in the training data. This prior information enlarges the margin between different classes to enhance the discriminative capacity of EProCRC. Experiments on two challenging databases, namely CUB200-2011 and Caltech-256, are conducted to evaluate EProCRC, and comparison results demonstrate that it outperforms several state-of-the-art classifiers. Rushi Lan, Yicong Zhou |
ICME | 2 |
| 2017 | 4 × 4 parametric integer discrete cosine transformsabstractOne of the famous transformations is discrete cosine transform (DCT), which is always used in digital image coding standards like JPEG and MPEG. DCT has different types and all of DCTs have excellent energy compaction properties. Meanwhile, matrices of DCTs II and IV are examined by a lot of researchers. However, other types of DCTs are rarely developed. Therefore, this paper presents 4 × 4 parametric integer DCTs, which cannot only represent DCTs II and IV but other types of DCTs (i.e. DCT I, V, VIII). It shows an excellent performance in terms of mean square errors and transform coding gains while it is comparing with state of the art. Weijia Cao, Yicong Zhou |
SMC | 2 |
| 2017 | Focusness guided salient object detectionabstractSalient object detection aims to correctly highlight the most salient object(s) in an image. Combining fine-grained contrast prior with rough-grained object consistency, this paper proposes a Focusness Guided Salient object detection (FGS) algorithm. To obtain clean and precise contrast map, FGS uses the focusness prior to guide the contrast map. Combing different saliency priors, FGS utilizes a unified least-square framework to generate the final optimal salient map. Experiments demonstrate the proposed method outperforms the state-of-the-arts. Xiaolin Xiao, Yicong Zhou |
SMC | 2 |
| 2017 | Design of image cipher using block-based scrambling and image filtering
Zhongyun Hua, Yicong Zhou |
Inf. Sci. | 2 |
| 2017 | Dictionary Learning-Based Hough Transform for Road Detection in Multispectral ImageabstractIt is of great importance to determine the location and orientation of a straight road in multispectral images for remote sensing. One of the classical methods for straight line detection is the Hough transform that is widely used in binary images. Although there are many previous works for straight road detection, it is still in its infancy to extract a straight road in multispectral images for remote sensing. In this letter, we propose a multiview dictionary learning formulation to approximate the Hough transform for straight road detection in multispectral images. Our formulation can exploit the complementary among the multiple spectral channels. Furthermore, it is natural to incorporate regularizations of prior information to significantly leverage the performance. We consider L1-norm regularization as a case study and conduct extensive experiments on RSSCN7 data set to verify the proposed algorithm. The experimental results demonstrate the superiority of our method in comparison with traditional methods. Weifeng Liu 0001, Zhenqing Zhang, Xinghua Chen, Yicong Zhou |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2017 | Medical image encryption using edge maps
Weijia Cao, Yicong Zhou, C. L. Philip Chen, Liming Xia |
Signal Process. | 2 |
| 2017 | Binary-block embedding for reversible data hiding in encrypted images
Yicong Zhou |
Signal Process. | 2 |
| 2017 | Kernel Regularized Data Uncertainty for Action RecognitionabstractThe traditional data uncertainty (DU) classifier fails to encode the importance of each sample for solving the minimum problem. Moreover, it considers only linear information for classification. To overcome these, we propose four classifiers for action recognition. They are called regularized DU (RDU) classifier, RDU coefficient (RDUC) classifier, kernel RDU (KRDU) classifier, and kernel RDUC (KRDUC) classifier, respectively. Extensive experiments on four benchmark action databases demonstrate that the proposed four classifiers achieve better recognition rates than the traditional DU classifier and several state-of-the-art methods. Moreover, the computation costs of the KRDU and KRDUC classifiers are much less than that of the DU classifier. Qingxiang Feng, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Quaternionic Weber Local Descriptor of Color ImagesabstractThis paper proposes a simple but effective framework named quaternionic Weber local descriptor (QWLD) for color image feature extraction. Integrating quaternionic representation (QR) of the color image and Weber's law (WL), QWLD possesses both their superiorities. It uses QR to handle all color channels of the image in a holistic way while preserving their relations, and applies WL to ensure that the derived descriptors are robust and discriminative. Using the QWLD framework, we further develop the quaternionic-increment-based Weber descriptor and quaternionic-distance-based Weber descriptor in terms of different perspectives. Extensive experiments on different color image recognition problems demonstrate that the proposed framework and descriptors outperform state-of-the-art local descriptors. Rushi Lan, Yicong Zhou, Yuan Yan Tang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Superimposed Sparse Parameter Classifiers for Face RecognitionabstractIn this paper, a novel classifier, called superimposed sparse parameter (SSP) classifier is proposed for face recognition. SSP is motivated by two phase test sample sparse representation (TPTSSR) and linear regression classification (LRC), which can be treated as the extended of sparse representation classification (SRC). SRC uses all the train samples to produce the sparse representation vector for classification. The LRC, which can be interpreted as L2-norm sparse representation, uses the distances between the test sample and the class subspaces for classification. TPTSSR is also L2-norm sparse representation and uses two phase to compute the distance for classification. Instead of the distances, the SSP classifier employs the SSPs, which can be expressed as the sum of the linear regression parameters of each class in iterations, is used for face classification. Further, the fast SSP (FSSP) classifier is also suggested to reduce the computation cost. A mass of experiments on Georgia Tech face database, ORL face database, CVL face database, AR face database, and CASIA face database are used to evaluate the proposed algorithms. The experimental results demonstrate that the proposed methods achieve better recognition rate than the LRC, SRC, collaborative representation-based classification, regularized robust coding, relaxed collaborative representation, support vector machine, and TPTSSR for face recognition under various conditions. Qingxiang Feng, Chun Yuan 0003, Jeng-Shyang Pan 0001, Jar-Ferr Yang, Yang-Ting Chou, Yicong Zhou, Weifeng Li 0001 |
IEEE Trans. Cybern. | 6 |
| 2017 | Combination of Sharing Matrix and Image Encryption for Lossless (k, n)-Secret Image Sharingabstractand its generation algorithm. Mathematical analysis is provided to show its potential for secret image sharing. Combining sharing matrix with image encryption, we further propose a lossless (k,n) -secret image sharing scheme (SMIE-SIS). Only with no less than k shares, all the ciphertext information and security key can be reconstructed, which results in a lossless recovery of original information. This can be proved by the correctness and security analysis. Performance evaluation and security analysis demonstrate that the proposed SMIE-SIS with arbitrary settings of k and n has at least five advantages: 1) it is able to fully recover the original image without any distortion; 2) it has much lower pixel expansion than many existing methods; 3) its computation cost is much lower than the polynomial-based secret image sharing methods; 4) it is able to verify and detect a fake share; and 5) even using the same original image with the same initial settings of parameters, every execution of SMIE-SIS is able to generate completely different secret shares that are unpredictable and non-repetitive. This property offers SMIE-SIS a high level of security to withstand many different attacks. Long Bao, Yicong Zhou |
IEEE Trans. Image Process. | 3 |
| 2017 | Medical Image Retrieval via Histogram of Compressed Scattering CoefficientsabstractThe features used in many current medical image retrieval systems are usually low-level hand-crafted features. This limitation may adversely affect the retrieval performance. To address this problem, this paper proposes a simple yet discriminative feature, called histogram of compressed scattering coefficients (HCSC), for medical image retrieval. In the proposed work, the scattering transform, a particular variation of deep convolutional networks, is first performed to yield more abstract representations of a medical image. A projection operation is then conducted to compress the obtained scattering coefficients for efficient processing. Finally, a bag-of-words (BoW) histogram is derived from the compressed scattering coefficients as the features of the medical image. The proposed HCSC takes the advantages of both scattering transform and BoW model. Experiments on three benchmark medical computer tomography image databases demonstrate that HCSC outperforms several state-of-the-art features. Rushi Lan, Yicong Zhou |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Parametric Planning Model for Video Quality Evaluation of IPTV Services Combining Channel and Video CharacteristicsabstractParametric planning models are designed for estimating the video quality, which can be applied to effective planning, implementation, and management of network video applications and communication networks. However, different from the bitstream-based evaluation models, the planning models are not allowed to exploit the video streams, with only limited information available for use, i.e., a few general parameters predetermined by the service providers and network operators. In this paper, a parametric planning model combining channel and video characteristics is proposed to estimate the video distortion caused by packet loss for Internet protocol television (IPTV) services. More specifically, the probability distribution of the channel states is determined by detailed analysis of the channel characteristics. Then, considering the influence of burst packet loss and the temporal dependence between frames, several sequence-level and frame-level parameters for video quality evaluation are derived from the perspective of the probability distribution of the channel states. Utilizing these parameters, the proposed model approximates the video quality considering the effects of direct packet loss and error propagation. Experimental results show that the proposed model has a superior performance for video quality estimation than the three commonly used parametric planning models. Jiarun Song, Fuzheng Yang 0001, Yicong Zhou |
IEEE Trans. Multim. | 3 |
| 2016 | Pairwise Linear Regression Classification for Image Set RetrievalabstractThis paper proposes the pairwise linear regression classification (PLRC) for image set retrieval. In PLRC, we first define a new concept of the unrelated subspace and introduce two strategies to constitute the unrelated subspace. In order to increase the information of maximizing the query set and the unrelated image set, we introduce a combination metric for two new classifiers based on two constitution strategies of the unrelated subspace. Extensive experiments on six well-known databases prove that the performance of PLRC is better than that of DLRC and several state-of-theart classifiers for different vision recognition tasks: clusterbased face recognition, video-based face recognition, object recognition and action recognition. Qingxiang Feng, Yicong Zhou, Rushi Lan |
CVPR | 2 |
| 2016 | Iterative linear regression classification for image recognitionabstractTraditional linear regression classification (LRC) suffers from a small sample size problem that the limited training samples of each class cannot comprehensively reflect different variations of the class. To address the problem, this paper proposes a novel iterative linear regression classification (ILRC) for image recognition. Different from traditional LRC, ILRC not only generates several new subspaces in each iteration but also uses the discrimination idea to optimize the training-set and testing samples. Extensive experiments on five benchmark databases demonstrate that the proposed ILRC classifier achieves better recognition performance than the traditional LRC and several state-of-the-art methods. Qingxiang Feng, Yicong Zhou |
ICASSP | 2 |
| 2016 | PLIP based unsharp masking for medical image enhancementabstractMedical image enhancement is an effective tool to improve visual quality of digital medical images. In this paper, we propose a new unsharp masking scheme for medical image enhancement. It embeds the PLIP multiplication into the unsharp masking framework. Experimental results demonstrate that the proposed method can effectively enhance the overall contrast and edges of medical images while suppressing background noise. Yicong Zhou |
ICASSP | 2 |
| 2016 | A superpixel segmentation algorithm based on differential evolutionabstractThis paper deals with the superpixel segmentation problem using a powerful global optimization technique: Differential Evolution. The algorithm mimics the process of nature evolution to realize efficient optimization, and it poses no restrictions on the form of objective functions. This way, we develop a novel and comprehensive objective function considering both local and global costs in the segmentation, including within-superpixel error, boundary gradient, a regularization term. The proposed method can produce superpixels in a computational time linear to the image size. Experimental results validate the competitive performance of our algorithm in terms of boundary adherence and segmentation capability. Yue-Jiao Gong, Yicong Zhou, Xinglin Zhang 0001 |
ICME | 2 |
| 2016 | Ideal regularized kernel for hyperspectral image classificationabstractThis paper proposes an ideal regularized composite kernel (IRCK) framework for hyperspectral images (HSI) classification. In learning a composite kernel, IRCK exploits spectral information, spatial information, and label information simultaneously. It incorporates the labels into standard spectral and spatial kernels by means of ideal kernel according to a regularization kernel learning framework, which captures both the sample similarity and label similarity and makes the resulting kernel more appropriate for HSI classification tasks. With the ideal regularization, the kernel learning problem has a simple analytical solution and is very easy to implement. The ideal regularization can be used to improve and refine state-of-the-art kernels, including spectral kernels, spatial kernels and spectral-spatial composite kernels. The effectiveness of the proposed IRCK is validated on the benchmark hyperspectral data set: Indian Pines. Experimental results show the superiority of our ideal regularized composite kernel method over the classical kernel methods. Jiangtao Peng, Yicong Zhou |
IGARSS | 2 |
| 2016 | Stacked Tensor Subspace Learning for hyperspectral image classificationabstractIn this paper, we present a hierarchical feature learning method called Stacked Tensor Subspace Learning (STSL). It can jointly learn spectral and spatial features of hyperspectral images (HSIs) by iteratively abstracting neighboring regions. STSL is able to learn discriminative spectral-spatial features of the input HSI at different scales. In STSL, the joint spectral and spatial features are extracted using Marginal Fisher Analysis (MFA) and Tensor Principal Component Analysis (TPCA). Then Kernel-based Extreme Learning Machine (KELM), a shallow neural network, is embedded in the proposed method to classify image pixels. The important contributions to the success of STSL are exploiting local spatial structure of HSI by using tensor method and designing hierarchical architecture. Extensive experimental results on two challenging HSI data sets taken from the Airborne Visible-Infrared Imaging Spectrometer (AVIRIS) and Reflective Optics System Imaging Spectrometer (ROSIS) airborne sensors show that the proposed method can produce good classification accuracy with smaller training sets. Yantao Wei, Yicong Zhou |
IJCNN | 2 |
| 2016 | Quaternion decomposition based discriminant analysis for color face recognitionabstractIn this paper, we propose a novel quaternion decomposition based discriminant analysis (QDDA) method for color face recognition. Unlike traditional approaches that handle color face images by vector representation or by each color channel individually, QDDA makes use of the quaternion to encode all color channels such that we can process all these channels in a holistic way and consider their relations simultaneously. In order to extract more discriminant color information from the image, a decomposition operation is performed to the quaternion matrix. A linear discriminant analysis is finally implemented to the obtained subcomponents for feature extraction. Experimental results have demonstrated the effectiveness of QDDA by comparing with other quaternion based methods. Rushi Lan, Yicong Zhou |
SMC | 2 |
| 2016 | Improved reversible data hiding in encrypted images using histogram modificationabstractInspired by Zhang et al.'s method that applies the integer discrete wavelet transform (DWT) to the original image, and embeds the secret data into the middle (LH, HL) and high (HH) frequency sub-bands of integer DWT coefficients with histogram modification based method, we propose a reversible data hiding method that embeds the secret data into the encrypted prediction error values. Compared with the LH, HL and HH integer DWT coefficients, the prediction error values generated by our proposed method are more concentrated to 0, and thus a high visual quality of the marked decrypted image can be achieved. Experimental results show that our proposed method has a better performance than Zhang's. Yicong Zhou |
SMC | 2 |
| 2016 | Comparative study of logarithmic image processing models for medical image enhancementabstractMedical image enhancement is an effective tool to improve visual quality of digital medical images. However, conventional linear image enhancement methods often suffers from problems such as over-enhancement and noise sensitivity. In this paper, we study nonlinear arithmetic frameworks designed to solve the common problems of linear enhancement methods, namely, LIP, PLIP and GLIP. We also introduce nonlinear unsharp masking algorithms based on the logarithmic image processing models for medical image enhancement. Experiments are conducted to evaluate and compare the performance of the methods. Yicong Zhou |
SMC | 2 |
| 2016 | Image encryption using 2D Logistic-adjusted-Sine map
Zhongyun Hua, Yicong Zhou |
Inf. Sci. | 2 |
| 2016 | 2D Sudoku associated bijections for image scrambling
Yue Wu 0001, Yicong Zhou, Sos S. Agaian, Joseph P. Noonan |
Inf. Sci. | 2 |
| 2016 | Two strategies to optimize the decisions in signature verification with the presence of spoofing attacks
Shilian Yu, Ye Ai, Yicong Zhou, Weifeng Li 0001, Qingmin Liao, Norman Poh |
Inf. Sci. | 4 |
| 2016 | Non-rigid visual object tracking using user-defined marker and Gaussian kernel
Guoheng Huang, Chi-Man Pun, Cong Lin 0001, Yicong Zhou |
Multim. Tools Appl. | 4 |
| 2016 | Genetic Learning Particle Swarm OptimizationabstractSocial learning in particle swarm optimization (PSO) helps collective efficiency, whereas individual reproduction in genetic algorithm (GA) facilitates global effectiveness. This observation recently leads to hybridizing PSO with GA for performance enhancement. However, existing work uses a mechanistic parallel superposition and research has shown that construction of superior exemplars in PSO is more effective. Hence, this paper first develops a new framework so as to organically hybridize PSO with another optimization technique for "learning." This leads to a generalized "learning PSO" paradigm, the *L-PSO. The paradigm is composed of two cascading layers, the first for exemplar generation and the second for particle updates as per a normal PSO algorithm. Using genetic evolution to breed promising exemplars for PSO, a specific novel *L-PSO algorithm is proposed in the paper, termed genetic learning PSO (GL-PSO). In particular, genetic operators are used to generate exemplars from which particles learn and, in turn, historical search information of particles provides guidance to the evolution of the exemplars. By performing crossover, mutation, and selection on the historical information of particles, the constructed exemplars are not only well diversified, but also high qualified. Under such guidance, the global search ability and search efficiency of PSO are both enhanced. The proposed GL-PSO is tested on 42 benchmark functions widely adopted in the literature. Experimental results verify the effectiveness, efficiency, robustness, and scalability of the GL-PSO. Yue-Jiao Gong, Jingjing Li 0002, Yicong Zhou, Yun Li 0002, Henry S. H. Chung, Yu-hui Shi, Jun Zhang 0003 |
IEEE Trans. Cybern. | 3 |
| 2016 | Dynamic Parameter-Control Chaotic SystemabstractThis paper proposes a general framework of 1-D chaotic maps called the dynamic parameter-control chaotic system (DPCCS). It has a simple but effective structure that uses the outputs of a chaotic map (control map) to dynamically control the parameter of another chaotic map (seed map). Using any existing 1-D chaotic map as the control/seed map (or both), DPCCS is able to produce a huge number of new chaotic maps. Evaluations and comparisons show that chaotic maps generated by DPCCS are very sensitive to their initial states, and have wider chaotic ranges, better unpredictability and more complex chaotic behaviors than their seed maps. Using a chaotic map of DPCCS as an example, we provide a field-programmable gate array design of this chaotic map to show the simplicity of DPCCS in hardware implementation, and introduce a new pseudo-random number generator (PRNG) to investigate the applications of DPCCS. Analysis and testing results demonstrate the excellent randomness of the proposed PRNG. Zhongyun Hua, Yicong Zhou |
IEEE Trans. Cybern. | 2 |
| 2016 | n-Dimensional Discrete Cat Map Generation Using Laplace ExpansionsabstractDifferent from existing methods that use matrix multiplications and have high computation complexity, this paper proposes an efficient generation method of${n}$-dimensional (${n}\text{D}$) Cat maps using Laplace expansions. New parameters are also introduced to control the spatial configurations of the${n}\text{D}$Cat matrix. Thus, the proposed method provides an efficient way to mix dynamics of all dimensions at one time. To investigate its implementations and applications, we further introduce a fast implementation algorithm of the proposed method with time complexity${O(n^{4})}$and a pseudorandom number generator using the Cat map generated by the proposed method. The experimental results show that, compared with existing generation methods, the proposed method has a larger parameter space and simpler algorithm complexity, generates${n}\text{D}$Cat matrices with a lower inner correlation, and thus yields more random and unpredictable outputs of${n}\text{D}$Cat maps. Yue Wu 0001, Zhongyun Hua, Yicong Zhou |
IEEE Trans. Cybern. | 3 |
| 2016 | Learning Hierarchical Spectral-Spatial Features for Hyperspectral Image ClassificationabstractThis paper proposes a spectral-spatial feature learning (SSFL) method to obtain robust features of hyperspectral images (HSIs). It combines the spectral feature learning and spatial feature learning in a hierarchical fashion. Stacking a set of SSFL units, a deep hierarchical model called the spectral-spatial networks (SSN) is further proposed for HSI classification. SSN can exploit both discriminative spectral and spatial information simultaneously. Specifically, SSN learns useful high-level features by alternating between spectral and spatial feature learning operations. Then, kernel-based extreme learning machine (KELM), a shallow neural network, is embedded in SSN to classify image pixels. Extensive experiments are performed on two benchmark HSI datasets to verify the effectiveness of SSN. Compared with state-of-the-art methods, SSN with a deep hierarchical architecture obtains higher classification accuracy in terms of the overall accuracy, average accuracy, and kappa ( κ ) coefficient of agreement, especially when the number of the training samples is small. Yicong Zhou, Yantao Wei |
IEEE Trans. Cybern. | 1 |
| 2016 | Quaternion-Michelson Descriptor for Color Image ClassificationabstractIn this paper, we develop a simple yet powerful framework called quaternion-Michelson descriptor (QMD) to extract local features for color image classification. Unlike traditional local descriptors extracted directly from the original (raw) image space, QMD is derived from the Michelson contrast law and the quaternionic representation (QR) of color images. The Michelson contrast is a stable measurement of image contents from the viewpoint of human perception, while QR is able to handle all the color information of the image holisticly and to preserve the interactions among different color channels. In this way, QMD integrates both the merits of Michelson contrast and QR. Based on the QMD framework, we further propose two novel quaternionic Michelson contrast binary pattern descriptors from different perspectives. Experiments and comparisons on different color image classification databases demonstrate that the proposed framework and descriptors outperform several state-of-the-art methods. Rushi Lan, Yicong Zhou |
IEEE Trans. Image Process. | 2 |
| 2016 | Quaternionic Local Ranking Binary Pattern: A Local Descriptor of Color ImagesabstractThis paper proposes a local descriptor called quaternionic local ranking binary pattern (QLRBP) for color images. Different from traditional descriptors that are extracted from each color channel separately or from vector representations, QLRBP works on the quaternionic representation (QR) of the color image that encodes a color pixel using a quaternion. QLRBP is able to handle all color channels directly in the quaternionic domain and include their relations simultaneously. Applying a Clifford translation to QR of the color image, QLRBP uses a reference quaternion to rank QRs of two color pixels, and performs a local binary coding on the phase of the transformed result to generate local descriptors of the color image. Experiments demonstrate that the QLRBP outperforms several state-of-the-art methods. Rushi Lan, Yicong Zhou, Yuan Yan Tang |
IEEE Trans. Image Process. | 2 |
| 2016 | QoE Evaluation of Multimedia Services Based on Audiovisual Quality and User InterestabstractQuality of experience (QoE) has significant influence on whether or not a user will choose a service or product in the competitive era. For multimedia services, there are various factors in a communication ecosystem working together on users, which stimulate their different senses inducing multidimensional perceptions of the services, and inevitably increase the difficulty in measurement and estimation of the user's QoE. In this paper, a user-centric objective QoE evaluation model (QAVIC model for short) is proposed to estimate the user's overall QoE for audiovisual services, which takes account of perceptual audiovisual quality (QAV) and user interest in audiovisual content (IC) amongst influencing factors on QoE such as technology, content, context, and user in the communication ecosystem. To predict the user interest, a number of general viewing behaviors are considered to formulate the IC evaluation model. Subjective tests have been conducted for training and validation of the QAVIC model. The experimental results show that the proposed QAVIC model can estimate the user's QoE reasonably accurately using a 5-point scale absolute category rating scheme. Jiarun Song, Fuzheng Yang 0001, Yicong Zhou, Shuai Wan, Hong Ren Wu |
IEEE Trans. Multim. | 3 |
| 2016 | Kernel Combined Sparse Representation for Disease RecognitionabstractMotivated by the idea that the correlation structure of the entire training set can disclose the relationship between the test sample and the training samples, we propose the combined sparse representation (CSR) classifier for disease recognition. The CSR classifier minimizes the correlation structure of the entire training set multiplied by its transposition and the sparse coefficient together for classification. Including the kernel concept, we propose the kernel combined sparse representation classifier utilizing the high-dimensional nonlinear information instead of the linear information in the CSR classifier. Furthermore, considering the information of the training samples and the class center, we then propose the center-based kernel combined sparse representation (CKCSR) classifier. CKCSR uses the center-based kernel matrix to increase the center-based information that is helpful for classification. The proposed classifiers have been evaluated by extensive experiments on several well-known databases including the EXACT09 database, Emphysema-CT database, mini-MIAS database, Wisconsin breast cancer database, and HD-PECTF database. The experimental results demonstrate that the proposed classifiers achieve better recognition rates than the sparse representation-based classification, collaborative representation based classification, and several state-of-the-art methods. Yicong Zhou |
IEEE Trans. Multim. | 1 |
| 2016 | Generalization Performance of Regularized Ranking With Multiscale KernelsabstractThe regularized kernel method for the ranking problem has attracted increasing attentions in machine learning. The previous regularized ranking algorithms are usually based on reproducing kernel Hilbert spaces with a single kernel. In this paper, we go beyond this framework by investigating the generalization performance of the regularized ranking with multiscale kernels. A novel ranking algorithm with multiscale kernels is proposed and its representer theorem is proved. We establish the upper bound of the generalization error in terms of the complexity of hypothesis spaces. It shows that the multiscale ranking algorithm can achieve satisfactory learning rates under mild conditions. Experiments demonstrate the effectiveness of the proposed method for drug discovery and recommendation tasks. Yicong Zhou, Hong Chen 0004, Rushi Lan, Zhibin Pan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Image Cipher Using a New Interactive Two-Dimensional Chaotic MapabstractIn this paper, a new two-dimensional (2D) Tent cascade-Logistic map (2D-TCLM) is introduced. Analysis results demonstrate that it has complex chaotic behaviors. Using 2D-TCLM, a new image encryption algorithm is also proposed. Simulation results and security analysis show that it can encrypt different kinds of digital images into unrecognized random-like images with a high security level. Zhongyun Hua, Yicong Zhou |
SMC | 3 |
| 2015 | High-speed implementation of rate-distortion optimised quantisation for H.265/HEVCabstractRate‐distortion optimised quantisation (RDOQ) is generally utilised in video coding for achieving higher coding efficiency. To determine the optimal quantised level for each transform coefficient, it requires considerable complexity in practice to calculate rate‐distortion (RD) costs from multiple candidates of quantised levels. This study proposes a fast RDOQ algorithm with low computational complexity for the latest video coding standard, the high‐efficiency video coding. Instead of calculating the RD costs from two candidates of quantised levels separately, their RD cost differences were derived. A bit‐rate estimation method is then used to accurately calculate the length of coding bits in the RD cost function, avoiding the high‐computation‐cost process of context‐adaptive binary arithmetic coding. Experimental results show that, with a negligible degradation of coding performance, the proposed algorithm is faster than RDOQ by 74.6% in average. Jing He 0007, Fuzheng Yang 0001, Yicong Zhou |
IET Image Process. | 3 |
| 2015 | Patterns of Weber magnitude and orientation for uncontrolled face representation and recognition
Yinyan Jiang, Yicong Zhou, Weifeng Li 0001, Qingmin Liao |
Neurocomputing | 3 |
| 2015 | Generalization ability of extreme learning machine with uniformly ergodic Markov chains
Peipei Yuan, Hong Chen 0004, Yicong Zhou, Xiaoyan Deng, Bin Zou 0002 |
Neurocomputing | 3 |
| 2015 | Image encryption: Generating visually meaningful encrypted images
Long Bao, Yicong Zhou |
Inf. Sci. | 2 |
| 2015 | 2D Sine Logistic modulation map for image encryption
Zhongyun Hua, Yicong Zhou, Chi-Man Pun, C. L. Philip Chen |
Inf. Sci. | 2 |
| 2015 | A new weighted mean filter with a two-phase detector for removing impulse noise
Licheng Liu, C. L. Philip Chen, Yicong Zhou, Xinge You |
Inf. Sci. | 3 |
| 2015 | Fast Fourier transform using matrix decomposition
Yicong Zhou, Weijia Cao, Licheng Liu, Sos S. Agaian, C. L. Philip Chen |
Inf. Sci. | 1 |
| 2015 | Cascade Chaotic System With ApplicationsabstractChaotic maps are widely used in different applications. Motivated by the cascade structure in electronic circuits, this paper introduces a general chaotic framework called the cascade chaotic system (CCS). Using two 1-D chaotic maps as seed maps, CCS is able to generate a huge number of new chaotic maps. Examples and evaluations show the CCS's robustness. Compared with corresponding seed maps, newly generated chaotic maps are more unpredictable and have better chaotic performance, more parameters, and complex chaotic properties. To investigate applications of CCS, we introduce a pseudo-random number generator (PRNG) and a data encryption system using a chaotic map generated by CCS. Simulation and analysis demonstrate that the proposed PRNG has high quality of randomness and that the data encryption system is able to protect different types of data with a high-security level. Yicong Zhou, Zhongyun Hua, Chi-Man Pun, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2015 | Region-Kernel-Based Support Vector Machines for Hyperspectral Image ClassificationabstractThis paper proposes a region kernel to measure the region-to-region distance similarity for hyperspectral image (HSI) classification. The region kernel is designed to be a linear combination of multiscale box kernels, which can handle the HSI regions with arbitrary shape and size. Integrating labeled pixels and labeled regions, we further propose a region-kernel-based support vector machine (RKSVM) classification framework. In RKSVM, three different composite kernels are constructed to describe the joint spatial-spectral similarity. Particularly, we design a desirable stack composite kernel that consists of the point-based kernel, the region-based kernel, and the cross point-to-region kernel. The effectiveness of the proposed RKSVM is validated on three benchmark hyperspectral data sets. Experimental results show the superiority of our region kernel method over the classical point kernel methods. Jiangtao Peng, Yicong Zhou, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | Dimension Reduction Using Spatial and Spectral Regularized Local Discriminant Embedding for Hyperspectral Image ClassificationabstractDimension reduction (DR) is a necessary and helpful preprocessing for hyperspectral image (HSI) classification. In this paper, we propose a spatial and spectral regularized local discriminant embedding (SSRLDE) method for DR of hyperspectral data. In SSRLDE, hyperspectral pixels are first smoothed by the multiscale spatial weighted mean filtering. Then, the local similarity information is described by integrating a spectral-domain regularized local preserving scatter matrix and a spatial-domain local pixel neighborhood preserving scatter matrix. Finally, the optimal discriminative projection is learned by minimizing a local spatial-spectral scatter and maximizing a modified total data scatter. Experimental results on benchmark hyperspectral data sets show that the proposed SSRLDE significantly outperforms the state-of-the-art DR methods for HSI classification. Yicong Zhou, Jiangtao Peng, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Weighted Couple Sparse Representation With Classified Regularization for Impulse Noise RemovalabstractMany impulse noise (IN) reduction methods suffer from two obstacles, the improper noise detectors and imperfect filters they used. To address such issue, in this paper, a weighted couple sparse representation model is presented to remove IN. In the proposed model, the complicated relationships between the reconstructed and the noisy images are exploited to make the coding coefficients more appropriate to recover the noise-free image. Moreover, the image pixels are classified into clear, slightly corrupted, and heavily corrupted ones. Different data-fidelity regularizations are then accordingly applied to different pixels to further improve the denoising performance. In our proposed method, the dictionary is directly trained on the noisy raw data by addressing a weighted rank-one minimization problem, which can capture more features of the original data. Experimental results demonstrate that the proposed method is superior to several state-of-the-art denoising methods. C. L. Philip Chen, Licheng Liu, Long Chen 0001, Yuan Yan Tang, Yicong Zhou |
IEEE Trans. Image Process. | 5 |
| 2014 | Person reidentification using quaternionic local binary patternabstractPerson reidentification is to identify the persons observed in nonoverlapping camera networks. Most existing methods usually extract features from the red, green, and blue color channels of images individually. They, however, neglect the connections between each color component in the image. To overcome this problem, a novel quaternionic local binary pattern (QLBP) is proposed for person reidentification in this paper. In the proposed QLBP, each pixel in a color image is represented by a quaternion so that we can handle all color components in a holistic way. A novel pseudo-rotation of quaternion (PRQ) is proposed to rank two quaternions. Some properties of PRQ are also discussed. After a QLBP coding, the local histograms are extracted and used as features. Experiments on two public benchmarking datasets, ETHZ and i-LIDS MCTS, are carried out to evaluate the QLBP performance. Comparison results show that the QLBP outperforms several stat-of-art methods for person reidentification. Rushi Lan, Yicong Zhou, Yuan Yan Tang, C. L. Philip Chen |
ICME | 2 |
| 2014 | A lossless (2, 8)-chaos-based secret image sharing schemeabstractThis paper introduces a new (2, 8)-secret image sharing scheme integrating the chaos-based image encryption with secret image sharing. It divides the secret image into 8 encrypted shares. Combining any two or more shares is able to completely reconstruct the secret image without any distortion. Each image share is only one pixel larger than the secret image in row and column directions. The proposed scheme is able to directly process the secret images with various formats such as the binary, grayscale, and color images. Experimental and comparison results demonstrate the excellent performance of the proposed scheme. Long Bao, Yicong Zhou, C. L. Philip Chen |
SMC | 2 |
| 2014 | Image encryption using 2D Logistic-Sine chaotic mapabstractThis paper introduces a new two-dimensional Logistic-Sine map (2D-LSM). It has excellent chaotic performance and its outputs are difficult to predict. Using 2D-LSM, this paper proposes a new image encryption algorithm. Simulation results and security analysis demonstrate that the proposed algorithm is able to protect different kinds of images with a high security level. Zhongyun Hua, Yicong Zhou, Chi-Man Pun, C. L. Philip Chen |
SMC | 2 |
| 2014 | Impulse noise removal using sparse representation with fuzzy weightsabstractMany impulse noise removal algorithms do not reach good denoising performance mainly due to the imperfect filters they adopted. In this paper, the popular used sparse representation model is extended for impulse noise removal by using a fuzzy weight matrix. This fuzzy weight is used to describe the noise-like level of the current pixel, and to determine how much information of this pixel should be used in the sparse land model. Besides, a regularization term which counts the proximity between the reconstructed image and the noisy image is also added into the sparse model. This makes the proposed model more robust to the noise detector which generates the fuzzy weight matrix. Moreover, unlike other sparse model, the dictionary used in our model is trained from some reference images that keep the similar structure information of the original image. Therefore, it is more suitable for reconstructing the original image. Simulation results show that our method is superior to all the tested state-of-the-art impulse noise removal methods. Licheng Liu, C. L. Philip Chen, Yicong Zhou, Yuan Yan Tang |
SMC | 3 |
| 2014 | Multi-scale patch based box kernels for hyperspectral image classificationabstractIntegrating labeled pixels with prior knowledge of hyperspectral spatial homogeneous regions, we propose a region-based hyperspectral image classification method, called the support vector machine with the multi-scale patch based box kernel (SVM-MPBK). It models the local homogeneous region of each pixel as a box, and measures the similarity between different box regions using box kernel. The box is represented as multidimensional intervals computed band by band in a neighborhood pixel patch. Using multi-scale patches to calculate box, SVM-MPBK fuses the complementary classification results in different scales by a majority voting. Experimental results on benchmark hyperspectral data sets demonstrate the effectiveness of SVM-MPBK. Jiangtao Peng, Yicong Zhou, C. L. Philip Chen |
SMC | 2 |
| 2014 | A new reversible data hiding algorithm in the encryption domainabstractThis paper introduces a new reversible data hiding algorithm in the encryption domain. It integrates data hiding into the image encryption process to achieve different level of access right and security. Computer simulations and comparisons demonstrate that the proposed algorithm can withstand the differential attack and outperforms other existing methods in terms of security and the message embedding capacity that is 52% larger than the state-of-the-art method in the best scenario. The marked decrypted images of our proposed method show the best visual quality according to the PSNR results. Yicong Zhou, Chi-Man Pun, C. L. Philip Chen |
SMC | 2 |
| 2014 | Integration of the saliency-based seed extraction and random walks for image segmentation
Chanchan Qin, Yicong Zhou, Wenbing Tao, Zhiguo Cao 0001 |
Neurocomputing | 3 |
| 2014 | Generalized Weber-face for illumination-robust face recognition
Yong Wu 0003, Yinyan Jiang, Yicong Zhou, Weifeng Li 0001, Zongqing Lu 0001, Qingmin Liao |
Neurocomputing | 3 |
| 2014 | Spatial adjacent bag of features with multiple superpixels for object segmentation and classification
Wenbing Tao, Yicong Zhou, Liman Liu, Kunqian Li, Kun Sun 0002, Zhiguo Zhang 0005 |
Inf. Sci. | 2 |
| 2014 | Design of image cipher using latin squares
Yue Wu 0001, Yicong Zhou, Joseph P. Noonan, Sos S. Agaian |
Inf. Sci. | 2 |
| 2014 | Extreme learning machine for ranking: Generalization analysis and applications
Hong Chen 0004, Jiangtao Peng, Yicong Zhou, Luoqing Li, Zhibin Pan |
Neural Networks | 3 |
| 2014 | Automatic image segmentation using salient key point extraction and star shape prior
Xiangli Liao, Yicong Zhou, Kunqian Li, Wenbing Tao, Qiuju Guo, Liman Liu |
Signal Process. | 3 |
| 2014 | A symmetric image cipher using wave perturbations
Yue Wu 0001, Yicong Zhou, Sos S. Agaian, Joseph P. Noonan |
Signal Process. | 2 |
| 2014 | A new 1D chaotic system for image encryption
Yicong Zhou, Long Bao, C. L. Philip Chen |
Signal Process. | 1 |
| 2014 | Image encryption using binary bitplane
Yicong Zhou, Weijia Cao, C. L. Philip Chen |
Signal Process. | 1 |
| 2014 | Feature mapping of multiple beamformed sources for robust overlapping speech recognition using a microphone arrayabstractThis paper introduces a nonlinear vector-based feature mapping approach to extract robust features for automatic speech recognition (ASR) of overlapping speech using a microphone array. We explore different configurations and additional sources of information to improve the effectiveness of the feature mapping. First, we investigate the full-vector based mapping of different sources in a log mel-filterbank energy (log MFBE) domain, and demonstrate that retraining the acoustic model using the generated training data can help improve the recognition performance. Then we investigate the feature mapping between different domains. Finally in order to improve the qualities of the mapping inputs we propose a nonlinear mapping of the features from multiple beamformed sources, which are directed at the target and interfering speakers, respectively. We demonstrate the effectiveness of the proposed approach through extensive evaluations on the MONC corpus, which includes non-overlapping single speaker and overlapping multi-speaker conditions. Weifeng Li 0001, Longbiao Wang, Yicong Zhou, John Dines, Mathew Magimai-Doss, Hervé Bourlard, Qingmin Liao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Generalization performance of support vector classifiers for density level detection
Hong Chen 0004, Yicong Zhou, Yi Tang 0003, Yuan Yan Tang, Zhibin Pan |
Neurocomputing | 2 |
| 2013 | Local Shannon entropy measure with statistical tests for image randomness
Yue Wu 0001, Yicong Zhou, George Saveriades, Sos S. Agaian, Joseph P. Noonan, Premkumar Natarajan |
Inf. Sci. | 2 |
| 2013 | Convergence rate of the semi-supervised greedy algorithm
Hong Chen 0004, Yicong Zhou, Yuan Yan Tang, Luoqing Li, Zhibin Pan |
Neural Networks | 2 |
| 2013 | Image encryption using a new parametric switching chaotic system
Yicong Zhou, Long Bao, C. L. Philip Chen |
Signal Process. | 1 |
| 2013 | Feature Denoising Using Joint Sparse Representation for In-Car Speech RecognitionabstractWe address reducing the mismatch between training and testing conditions for hands-free in-car speech recognition. It is well known that the distortions caused by background noise, channel effects, etc., are highly nonlinear in the log-spectral or cepstral domain. This letter introduces a joint sparse representation (JSR) to estimate the underlying clean feature vector from a noisy feature vector. Performing a joint dictionary learning by sharing the same representation coefficients, the proposed method intends to capture the complex relationships (or mapping functions) between clean and noisy speech. Speech recognition experiments on realistic in-car data demonstrate that the proposed method shows excellent recognition performance with a relative improvement of 39.4% compared with the “baseline” frontends. Weifeng Li 0001, Yicong Zhou, Norman Poh, Fei Zhou 0001, Qingmin Liao |
IEEE Signal Process. Lett. | 2 |
| 2013 | Robust Log-Energy Estimation and its Dynamic Change Enhancement for In-car Speech RecognitionabstractThe log-energy parameter, typically derived from a full-band spectrum, is a critical feature commonly used in automatic speech recognition (ASR) systems. However, log-energy is difficult to estimate reliably in the presence of background noise. In this paper, we theoretically show that background noise affects the trajectories of not only the “conventional” log-energy, but also its delta parameters. This results in a poor estimation of the actual log-energy and its delta parameters, which no longer describe the speech signal. We thus propose a new method to estimate log-energy from a sub-band spectrum, followed by dynamic change enhancement and mean smoothing. We demonstrate the effectiveness of the proposed log-energy estimation and its post-processing steps through speech recognition experiments conducted on the in-car CENSREC-2 database. The proposed log-energy (together with its corresponding delta parameters) yields an average improvement of 32.8% compared with the baseline front-ends. Moreover, it is also shown that further improvement can be achieved by incorporating the new Mel-Frequency Cepstral Coefficients (MFCCs) obtained by non-linear spectral contrast stretching. Weifeng Li 0001, Longbiao Wang, Yicong Zhou, Hervé Bourlard, Qingmin Liao |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | (n, k, p)-Gray Code for Image SystemsabstractThis paper introduces a new parametric n-ary Gray code, the (n, k, p)-Gray code, which includes several commonly used codes such as the binary-reflected, ternary, and (n, k)-Gray codes. The new (n, k, p)-Gray code has potential applications in digital communications and signal/image processing systems. This paper focuses on three illustrative applications of the (n, k, p)-Gray code, namely, image bit-plane decomposition, image denoising, and encryption. The computer simulations demonstrate that the (n, k, p)-Gray code shows better performance than other traditional Gray codes for these applications in image systems. Yicong Zhou, Karen Panetta, Sos S. Agaian, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2012 | A new image encryption algorithm using Truncated P-Fibonacci Bit-planesabstractImage encryption is an effective approach to protect privacy and security of images. This paper introduces a novel image encryption algorithm using the Truncated P-Fibonacci Bit-planes as security key images to encrypt images. Simulation results and security analysis are provided to show the encryption performance of the proposed algorithm. Weijia Cao, Yicong Zhou, C. L. Philip Chen |
SMC | 2 |
| 2012 | Image encryption algorithm based on a new combined chaotic systemabstractChaotic theory has been applied to image encryption as an effective and robust technique due to its unique properties. In this paper, we introduce a new combined chaotic system, which shows better chaotic behaviors than the traditional ones. Applying this chaotic system to image processing, a new image encryption algorithm is introduced based on the confusion and diffusion in encryption procedure. Experimental results show that the proposed algorithm has a higher security level and excellent performance in image encryption. C. L. Philip Chen, Tong Zhang 0015, Yicong Zhou |
SMC | 3 |
| 2011 | Nonlinear Unsharp Masking for Mammogram EnhancementabstractThis paper introduces a new unsharp masking (UM) scheme, called nonlinear UM (NLUM), for mammogram enhancement. The NLUM offers users the flexibility 1) to embed different types of filters into the nonlinear filtering operator; 2) to choose different linear or nonlinear operations for the fusion processes that combines the enhanced filtered portion of the mammogram with the original mammogram; and 3) to allow the NLUM parameter selection to be performed manually or by using a quantitative enhancement measure to obtain the optimal enhancement parameters. We also introduce a new enhancement measure approach, called the second-derivative-like measure of enhancement, which is shown to have better performance than other measures in evaluating the visual quality of image enhancement. The comparison and evaluation of enhancement performance demonstrate that the NLUM can improve the disease diagnosis by enhancing the fine details in mammograms with no a priori knowledge of the image contents. The human-visual-system-based image decomposition is used for analysis and visualization of mammogram enhancement. Karen Panetta, Yicong Zhou, Sos S. Agaian, Hongwei Jia |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2011 | Parameterized Logarithmic Framework for Image EnhancementabstractImage processing technologies such as image enhancement generally utilize linear arithmetic operations to manipulate images. Recently, Jourlin and Pinoli successfully used the logarithmic image processing (LIP) model for several applications of image processing such as image enhancement and segmentation. In this paper, we introduce a parameterized LIP (PLIP) model that spans both the linear arithmetic and LIP operations and all scenarios in between within a single unified model. We also introduce both frequency- and spatial-domain PLIP-based image enhancement methods, including the PLIP Lee's algorithm, PLIP bihistogram equalization, and the PLIP alpha rooting. Computer simulations and comparisons demonstrate that the new PLIP model allows the user to obtain improved enhancement performance by changing only the PLIP parameters, to yield better image fusion results by utilizing the PLIP addition or image multiplication, to represent a larger span of cases than the LIP and linear arithmetic cases by changing parameters, and to utilize and illustrate the logarithmic exponential operation for image fusion and enhancement. Karen Panetta, Sos S. Agaian, Yicong Zhou, Eric J. Wharton |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2010 | Nonlinear filtering for enhancing prostate MR images via alpha-trimmed Mean SeparationabstractThis paper introduces a new enhancement algorithm for prostate MR images using a new nonlinear filtering operation and an alpha-trimmed Mean Separation. A new enhancement measure is also introduced to measure and assess the enhanced results. Experimental results show that the presented algorithm can significantly improve the contrast of prostate MR images. It has a potential application in prostate cancer detection. Yicong Zhou, Karen Panetta, Sos S. Agaian |
SMC | 1 |
| 2009 | Image Encryption Using Binary Key-imagesabstractThis paper introduces a new concept for image encryption using a binary ¿key-image¿. The key-image is either a bit plane or an edge map generated from another image, which has the same size as the original image to be encrypted. In addition, we introduce two new lossless image encryption algorithms using this key-image technique. The performance of these algorithms is discussed against common attacks such as the brute force attack, ciphertext attacks and plaintext attacks. The analysis and experimental results show that the proposed algorithms can fully encrypt all types of images. This makes them suitable for securing multimedia applications and shows they have the potential to be used to secure communications in a variety of wired/wireless scenarios and real-time application such as mobile phone services. Yicong Zhou, Karen Panetta, Sos S. Agaian |
SMC | 1 |
| 2008 | Comparison of recursive sequence based image scrambling algorithmsabstractImage scrambling is an effective method for providing image security. This paper compares and discusses the effectiveness of some image scrambling algorithms based on recursive sequences such as the fibonacci number, generalized Fibonacci number, gray code, generalized gray code, generalized P-gray code, P-Fibonacci, P-Lucas, P-recursive sequences, and parametric M-sequences for image scrambling. The comparison of these methods for image security is based on three basic types of attacks: data loss attacks, noise attacks and plaintext attacks. The experimental results demonstrate that the scrambling algorithms based on both P-Fibonacci and P-Lucas sequences show better performance when subjected to attacks and also in terms of algorithm execution analysis which shows implementation efficiency and low computational requirements. This makes them suitable for real-time applications. Yicong Zhou, Karen Panetta, Sos S. Agaian |
SMC | 1 |