EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyan Luo
dblp:28/7701
· DBLP profile ↗
56ranked-venue papers
7as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LDFE: Laplacian Decoupled Feature Enhancement block for dual-stream CNN-based RGB-IR object detection
Xiaoyan Luo, Linlin Yang 0001, Haodong Zhu, Xiaorong Shi, Guodong Guo, Baochang Zhang 0001 |
Pattern Recognit. | 2 |
| 2026 | HySaDe-Mamba: A Mamba-Based Network for Hyperspectral Salient Object DetectionabstractHyperspectral salient object detection (HSOD) aims to identify visually and spectrally distinctive regions in hyperspectral images (HSIs). However, existing HSOD methods often suffer from spectral redundancy and inefficient spatial-spectral modeling, which hinder their scalability and accuracy in complex scenes. To tackle these challenges, we propose HySaDe-Mamba, a novel HSOD framework built upon the Mamba architecture. Specifically, to address information redundancy in HSI, we design a spatial-enhanced spectral-embedding (SeSe) module, which maps high-dimensional data into a more compact but effective representation. On the compact SeSe representation features, we further propose a Bi-scale spatial and Bi-directional spectral (BsBd) Mamba module, performing the selective scanning mechanism in a spatial-spectral hybrid, end-to-end way, which not only facilitates comprehensive spatial structural interaction across both global and local scales, but also effectively exploits the underlying spectral semantic correlation. Extensive experiments on two public HSOD datasets demonstrate that our HySaDe-Mamba achieves state-of-the-art detection accuracy across seven metrics, while maintaining an efficient inference speed of 40.22 FPS. The source code is publicly available at https://github.com/Leezl/HySaDe-Mamba. Lei Zhang 0109, Xiaoyan Luo, Liheng Bian, Xiantong Zhen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | SignMask: Structure-aware Masked Modeling for Holistic 3D Sign Language ProductionabstractSign Language Production (SLP) aims to translate spoken textual languages into sign language sequences, which can significantly bridge the communication gap for deaf and hard-of-hearing individuals. Most previous SLP methods typically rely on skeleton-based data, which hinders their realism and expressive capacity. In this work, we address expressive 3D SLP tasks to generate high-quality 3D holistic sign motions driven by spoken language. However, existing 3D SLP methods struggle to accurately capture spatial relationships within intricate 3D structures and overlook the alignment of semantics at the word level. To overcome these limitations, we propose SignMask, a novel generative masked modeling framework that enhances spatial structure awareness and semantic understanding. We first design a structural holistic sign motion tokenizer that hierarchically learns discrete tokens of body and hand movements. This tokenizer adaptively aggregates 3D SMPL-X pose features corresponding to the same semantic parts and dynamically adjusts the weights between pose features of different semantic parts, enhancing spatial awareness and ensuring semantic consistency. Building on these tokenized representations, we introduce a specialized Sign-M Transformer to learn masked token prediction guided by textual input. Our Sign-M Transformer employs a hierarchical masking strategy, alongside spatio-temporal and cross-modal attention mechanisms, to effectively capture complex spatio-temporal relationships among sign tokens and semantic dependencies between sign and text tokens. During inference, our SignMask model parallelly and iteratively fills up the missing motion tokens starting from full-masked token sequences, therefore achieving high-fidelity and efficient 3D sign avatar generation. Extensive experiments demonstrate the superior performance of our approach compared to existing SLP methods across various lingual sign language datasets in generating high-quality and semantically consistent sign language motions. Yibo Xia, Qihui Zhan, Xiaoyan Luo, Yunhong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object DetectionabstractLeveraging the complementary characteristics of visible (RGB) and infrared (IR) imagery offers significant potential for improving object detection. In this paper, we propose WaveMamba, a cross-modality fusion method that efficiently integrates the unique and complementary frequency features of RGB and IR decomposed by Discrete Wavelet Transform (DWT). An improved detection head incorporating the Inverse Discrete Wavelet Transform (IDWT) is also proposed to reduce information loss and produce the final detection results. The core of our approach is the introduction of WaveMamba Fusion Block (WMFB), which facilitates comprehensive fusion across low-/high-frequency sub-bands. Within WMFB, the Low-frequency Mamba Fusion Block (LMFB), built upon the Mamba framework, first performs initial low-frequency feature fusion with channel swapping, followed by deep fusion with an advanced gated attention mechanism for enhanced integration. High-frequency features are enhanced using a strategy that applies an ``absolute maximum" fusion approach. These advancements lead to significant performance gains, with our method surpassing state-of-the-art approaches and achieving average mAP improvements of 4.5% on four benchmarks. Haodong Zhu, Linlin Yang 0001, Hong Li 0016, Yuguang Yang 0007, Yangyang Ren, Qingcheng Zhu, Zichao Feng, Changbai Li, Shaohui Lin, Runqi Wang, Xiaoyan Luo, Baochang Zhang 0001 |
ICCV | 12 |
| 2025 | Structured Instruction Parsing and Scene Alignment For UAV Vision-Language NavigationabstractRecent advances in aerial Vision-and-Language Navigation (VLN) have introduced a more meaningful and practical paradigm of VLN by considering significantly longer paths and more complex spatial reasoning compared to ground-based VLN. However, the larger scale and increased complexity of outdoor environments in aerial VLN present substantial challenges in establishing accurate correspondence between textual instructions and visual scenes. In this work, we propose to incorporate Large Language Models (LLMs) to extract key components from navigation instructions and construct the corresponding subtasks. This structured instruction parsing module ensures the appropriate granularity of navigation instructions, enabling more precise alignment between language and visual cues. To further enhance the integration of multi-modal information and cross-modal understanding, we introduce a scene-based subtask alignment policy that effectively associates each parsed subtask with corresponding visual observations along the navigation path. Combined, the proposed approach significantly outperforms current state-of-the-art methods on the AerialVLN dataset. Liangyu Zhou, Xiaoyan Luo |
ICIP | 3 |
| 2025 | Modal-aware contrastive learning for hyperspectral and LiDAR classification
Liangyu Zhou, Xiaoyan Luo |
Image Vis. Comput. | 2 |
| 2025 | Diffusion Self-Distillation for Remote Sensing Scene ClassificationabstractRemote sensing scene classification, a fundamental task in remote image analysis, has obtained rapid progress due to the powerful capabilities of Convolutional Neural Networks (CNNs). Achieving precise classification performance heavily relies on the feature extraction capacity of the network. However, due to the large variation and severe distortion within the images, extracting robust feature representations is necessary but challenging. Self-distillation could enhance the shallow layers by providing stronger gradients and more accurate supervision from deeper layers, thereby promoting the extraction of spatially detailed features. Nonetheless, due to the limited capacity of shallow layers to learn truly valuable knowledge, shallow layer features can be viewed as the noisy version of deep layer features and contain more disruptive factors, which significantly impedes the effectiveness of self-distillation. To address this issue, in this paper, we establish the Diffusion Self-Distillation Network (DSDNet), which incorporates the conditional diffusion denoising model into the self-distillation framework. Specifically, DSDNet filters noise from shallow features through the diffusion denoising process, enabling more precise and accurate distillation between the refined student features and the teacher features. Extensive experiments on four challenging remote sensing datasets emonstrate that the proposed DSDNet achieves significant performance improvements over various backbone networks with negligible increases in parameters, delivering state-of-the-art classification performance. Our code and dataset are available on https://github.com/toggle1995/DSDNet. Yutao Hu 0002, Lei Zhang 0001, Xiaoyan Luo, Xianbin Cao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | S3A: A Self-Supervised Saliency Analysis Framework for Hyperspectral ImageabstractMost existing visual saliency analysis methods are meticulously crafted for natural RGB images, which typically lean upon foundational spatial visual cues, such as contrast, structure and texture in scenes. However, these methods exhibit limitations in spectral saliency analysis due to the existence of numerous materials that may manifest identical RGB values while harboring disparate spectral characteristics. Consequently, it is an exigency to introduce saliency method specially for hyperspectral images (HSIs). In this paper, we establish a set of hyperspectral saliency principles that incorporate both spectral and spatial attributes, and accordingly present a novel self-supervised saliency analysis (S3A) framework for HSIs. Note that our S3A is designed as a discriminant architecture composed of three interconnected components. More specifically, an unsupervised HSI representive patch sampling (RPS) module is designed to pick up some representative pixels for self-supervised saliency analysis training. Subsequently, we construct a classification-based patch spectral discrimination network (CPSD-Net) to evaluate the HSI saliency. Finally, a patch-to-global spatial diffusion (P2G-SD) module is constructed to diffuse the saliency from few supervised samples to the other HSI pixels. Moreover, to demonstrate the performance of our saliency analysis framework, we apply the HSI saliency to some classic downstream tasks including band selection (BS) and HSI classification. On several popular HSI datasets, the satisfactory quantization results fully verify the rationality and effectiveness of our S3A framework, in terms of entropy value and mean spectral divergence (MSD) of the selected bands in BS task, as well as the accuracy in HSI classification task. The source code is publicly available at https://github.com/Lee-zl/S3A. Xiaoyan Luo, Lei Zhang 0109, Peixin Gan, Liheng Bian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Image Intrinsic-Based Unsupervised Monocular Depth Estimation in EndoscopyabstractUnsupervised monocular depth estimation plays a vital role for endoscopy-based minimally invasive surgery (MIS). However, it remains challenging due to the distinctive imaging characteristics of endoscopy which disrupt the assumption of photometric consistency, a foundation relied upon by conventional methods. Distinct from recent approaches taking image pre-processing strategy, this paper introduces a pioneering solution through intrinsic image decomposition (IID) theory. Specifically, we propose a novel end-to-end intrinsic-based unsupervised monocular depth learning framework that is comprised of an image intrinsic decomposition module and a synthesis reconstruction module. This framework seamlessly integrates IID with unsupervised monocular depth estimation, and dedicated losses are meticulously designed to offer robust supervision for network training based on this novel integration. Noteworthy, we rely on the favorable property of the resulting albedo map of IID to circumvent the challenging images characteristics instead of pre-processing the input frames. The proposed method is extensively validated on SCARED and Hamlyn datasets, and better results are obtained than state-of-the-art techniques. Beside, its generalization ability and the effectiveness of the proposed components are also validated. This innovative method has the potential to elevate the quality of 3D reconstruction in monocular endoscopy, thereby enhancing the accuracy and robustness of augmented reality navigation technology in MIS. Bojian Li, Bo Liu 0027, Xiaoyan Luo, Fugen Zhou |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Fusion-Mamba for Cross-Modality Object DetectionabstractCross-modality object detection aims to fuse complementary information from different modalities to improve model performance, which achieves a wider range of applications. However, traditional cross-modality fusion methods, based on CNN or Transformer, inadequately address the issue of pseudo-target information, which causes model attention dispersion to degrade object detection performance. In this paper, we investigate a novel cross-modality fusion approach by associating cross-modal features in a hidden state space based on an improved Mamba with a gating attention mechanism. We propose theFusion-Mamba Block(FMB), designed to map cross-modal features into a hidden state space for interaction, thereby refining the model’s attention on true target areas and enhancing overall performance. The FMB comprises two key modules: State Space Channel Swapping (SSCS) module, which facilitates the fusion of shallow features, and Dual State Space Fusion (DSSF) module, which enables deep fusion and effectively suppresses pseudo-target information within the hidden state space. Our proposed method outperforms state-of-the-art approaches, achieving improvements of 5.9%, 3.5% and 2.1% mAP on$M^{3}$FD, DroneVehicle and FLIR-Aligned, respectively. To the best of our knowledge, this work establishes a new baseline for cross-modality object detection, providing a robust foundation for future research in this area. Haodong Zhu, Shaohui Lin, Xiaoyan Luo, Yunhang Shen, Guodong Guo, Baochang Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Hierarchical Self-Distilled Feature Learning for Fine-Grained Visual CategorizationabstractFine-grained visual categorization (FGVC) relies on hierarchical features extracted by deep convolutional neural networks (CNNs) to recognize closely alike objects. Particularly, shallow layer features containing rich spatial details are vital for specifying subtle differences between objects but are usually inadequately optimized due to gradient vanishing during backpropagation. In this article, hierarchical self-distillation (HSD) is introduced to generate well-optimized CNNs features for accurate fine-grained categorization. HSD inherits from the widely applied deep supervision and implements multiple intermediate losses for reinforced gradients. Besides that, we observe that the hard (one-hot) labels adopted for intermediate supervision hurt the performance of FGVC by enforcing overstrict supervision. As a solution, HSD seeks self-distillation where soft predictions generated by deeper layers of the network are hierarchically exploited to supervise shallow parts. Moreover, self-information entropy loss (SIELoss) is designed in HSD to adaptively soften intermediate predictions and facilitate better convergence. In addition, the gradient detached fusion (GDF) module is incorporated to produce an ensemble result with multiscale features via effective feature fusion. Extensive experiments on four challenging fine-grained datasets show that, with neglectable parameter increase, the proposed HSD framework and the GDF module both bring significant performance gains over different backbones, which also achieves state-of-the-art classification performance. Yutao Hu 0002, Xuhui Liu, Xiaoyan Luo, Yao Hu 0002, Xianbin Cao 0001, Baochang Zhang 0001, Jun Zhang 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Orthogonal Subspace Representation for Generative Adversarial NetworksabstractDisentanglement learning aims to separate explanatory factors of variation so that different attributes of the data can be well characterized and isolated, which promotes efficient inference for downstream tasks. Mainstream disentanglement approaches based on generative adversarial networks (GANs) learn interpretable data representation. However, most typical GAN-based works lack the discussion of the latent subspace, causing insufficient consideration of the variation of independent factors. Although some recent research analyzes the latent space on pretrained GANs for image editing, they do not emphasize learning representation directly from the subspace perspective. Appropriate subspace properties could facilitate corresponding feature representation learning to satisfy the independent variation requirements of the obtained explanatory factors, which is crucial for better disentanglement. In this work, we propose a unified framework for ensuring disentanglement, which fully investigates latent subspace learning (SL) in GAN. The novel GAN-based architecture explores orthogonal subspace representation (OSR) on vanilla GAN, named OSRGAN. To guide a subspace with strong correlation, less redundancy, and robust distinguishability, our OSR includes three stages, self-latent-aware, orthogonal subspace-aware, and structure representation-aware, respectively. First, the self-latent-aware stage promotes the latent subspace strongly correlated with the data space to discover interpretable factors, but with poor independence of variation. Second, the following orthogonal subspace-aware stage adaptively learns some 1-D linear subspace spanned by a set of orthogonal bases in the latent space. There is less redundancy between them, expressing the corresponding independence. Third, the structure representation-aware stage aligns the projection on the orthogonal subspace and the latent variables. Accordingly, feature representation in each linear subspace can be distinguishable, enhancing the independent expression of interpretable factors. In addition, we design an alternating optimization step, achieving a tradeoff training of OSRGAN on different properties. Despite it strictly constrains orthogonality, the loss weight coefficient of distinguishability induced by orthogonality could be adjusted and balanced with correlation constraint. To elucidate, this tradeoff training prevents our OSRGAN from overemphasizing any property and damaging the expressiveness of the feature representation. It takes into account both interpretable factors and their independent variation characteristics. Meanwhile, alternating optimization could keep the cost and efficiency of forward inference unchanged and will not burden the computational complexity. In theory, we clarify the significance of OSR, which brings better independence of factors, along with interpretability as correlation could converge to a high range faster. Moreover, through the convergence behavior analysis, including the objective functions under different constraints and the evaluation curve with iterations, our model demonstrates enhanced stability and definitely converges toward a higher peak for disentanglement. To depict the performance in downstream tasks, we compared the state-of-the-art GAN-based and even VAE-based approaches on different datasets. Our OSRGAN achieves higher disentanglement scores on FactorVAE, SAP, MIG, and VP metrics. All the experimental results illustrate that our novel GAN-based framework has considerable advantages on disentanglement. Hongxiang Jiang, Xiaoyan Luo, Jihao Yin, Huazhu Fu, Fuxiang Wang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Fault Localization for Novice Programs Combining Static Analysis and Dynamic DetectionabstractIn programming teaching, teachers or teaching assistants often need to spend a lot of energy helping students solve the problems they face when doing programming. It will be helpful to provide students with valuable programming feedback, such as information on faulty lines. However, most existing algorithms do not perform well on novice programs. Therefore, considering the background in programming teaching, we proposed a novel approach combining static analysis with dynamic detection by using the correct programs submitted by previous students and coverage information for the incorrect program. In particular, the core of the static analysis module is to locate specific faulty lines through syntax tree difference comparison, which includes matching similar programs, variable mapping and replacement, and fault localization based on abstract syntax tree differences. The core of the module on dynamic detection is to perform traditional Spectrum-based Fault Localization. To evaluate the effectiveness of our proposed approach, we conducted some empirical studies on 223 student-failure programs in the real world. The experimental results indicate that our approach outperforms other baselines regarding TOP-l and TOP-3. Furthermore, we analyzed the performance of our method on different categories of programming problems as well as the effectiveness of the combination of static analysis and dynamic detection. Han Wan, Wenhao Nie, Shiyang Yue, Xiaoyan Luo |
COMPSAC | 4 |
| 2024 | Learning Foreground Information Bottleneck for few-shot semantic segmentation
Yutao Hu 0002, Xiaoyan Luo, Jungong Han, Xianbin Cao 0001, Jun Zhang 0007 |
Pattern Recognit. | 3 |
| 2024 | An Interactively Motion-Assisted Network for Multiple Object Tracking in Complex Traffic ScenesabstractMultiple object tracking plays a crucial role in intelligent transportation systems. Due to the varying size, fast motion, and occlusion of traffic objects, multiple object tracking in complex traffic scenes is prone to low tracking accuracy and tracking fragmentation. To solve these problems, numerous tracking methods have been proposed based on object discrimination feature. However, these methods neglect the guiding prior role of historical tracking information for detection, making them not applicable to more complex scenes, such as extremely small, fast-moving, and severely obscured objects. In this paper, we propose an Interactively Motion-assisted Network (IMANet) for multiple object tracking in complex traffic scenes. First, to capture the object motion patterns from historical frames, an object motion modeling module considering camera movement is proposed, which particularly has prominent advantages on videos obtained by moving cameras. Next, a multi-scale fusion detection and embedding module is designed to incorporate historical motion information, thereby improving detection performance. Finally, multiple object tracking can be achieved by associating objects detected in different frames based on detection and embedding results. The proposed method combines the detection and tracking in an interactive way, where detection performance is facilitated using historical information provided by the tracking module, and better detection in turn enhances the tracking. Several real-world traffic examples are used to illustrate the performance of the proposed method in both detecting and tracking traffic objects. The results demonstrate that the proposed method outperforms the state-of-the-art methods, especially in complex surveillance videos with varying sizes and occlusions. Zhiqi Shen 0003, Kaiquan Cai, Peng Zhao 0005, Xiaoyan Luo |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Early Prediction of Student Performance with LSTM-Based Deep Neural Network
Han Wan, Mengying Li, Zihao Zhong, Xiaoyan Luo |
COMPSAC | 4 |
| 2023 | Implicit Diffusion Models for Continuous Super-ResolutionabstractImage super-resolution (SR) has attracted increasing attention due to its widespread applications. However, current SR methods generally suffer from over-smoothing and artifacts, and most work only with fixed magnifications. This paper introduces an Implicit Diffusion Model (IDM) for high-fidelity continuous image super-resolution. IDM integrates an implicit neural representation and a denoising diffusion model in a unified end-to-end framework, where the implicit neural representation is adopted in the decoding process to learn continuous-resolution representation. Furthermore, we design a scale-adaptive conditioning mechanism that consists of a low-resolution (LR) conditioning network and a scaling factor. The scaling factor regulates the resolution and accordingly modulates the proportion of the LR information and generated features in the final output, which enables the model to accommodate the continuous-resolution requirement. Extensive experiments validate the effectiveness of our IDM and demonstrate its superior performance over prior arts. The source code will be available at https://github.com/Ree1s/IDM. Sicheng Gao, Xuhui Liu, Bohan Zeng, Sheng Xu 0007, Yanjing Li, Xiaoyan Luo, Jianzhuang Liu, Xiantong Zhen, Baochang Zhang 0001 |
CVPR | 6 |
| 2023 | Learning Path Recommendation Based on Knowledge Tracing and Reinforcement LearningabstractAdaptive and intelligent web-based educational systems are made to automate the adaptation of the system to the learners' behaviors and needs. Personalized e-learning platforms should make adaptive adjustments according to the individual students' interactions and their knowledge states (KS). This study proposes a more effective personalized learning path recommendation algorithm to promote the individualized development of students. First, the Dynamic Key-Value Memory Network (DKVMN) is enhanced by integrating a learning behavior module, which is used to trace student knowledge states. Then, the proposed knowledge tracing model is used to simulate virtual students and train recommendation policy based on reinforcement learning (RL). The experimental results show that our personalized learning path recommendation algorithm increases the average knowledge state of students by 12.11% and 5.38% on two different data sets, respectively. Han Wan, Baoliang Che, Hongzhen Luo, Xiaoyan Luo |
ICALT | 4 |
| 2023 | An End-To-End Generative Classification Model for Hyperspectral ImageabstractCurrent mainstream deep learning-based algorithms for hyperspectral image classification (HSIC) rely on patch-based local learning models, which suffer from limitations such as restricted receptive fields and high computational complexity. In this paper, we propose an end-to-end generative classification model that overcomes these limitations by adopting a patch-free approach, enabling fast training and inference. We use the transformer to effectively capture complex spectral and spatial information in hyperspectral features. Additionally, we introduce a low-rank constraint on spectral features to leverage their strong correlations for improving global feature representation. Unlike traditional discriminative classification models, our model incorporates a generative global constraint, allowing us to fully utilize global contextual information to assist classification, particularly in few-shot scenarios. Experimental results on three publicly available datasets demonstrate that our model achieves state-of-the-art performance while maintaining robustness. Yaling Li, Xiaoyan Luo |
IGARSS | 2 |
| 2023 | R2H-CCD: Hyperspectral Imagery Generation from RGB Images Based on Conditional Cascade Diffusion Probabilistic ModelsabstractHyperspectral imaging can capture detailed spectra of materials, but due to imaging conditions and equipment limitations, the cost of hyperspectral data collection is extremely high. Leveraging low-priced and easy-obtainable RGB images to generate hyperspectral images (HSIs) has gradually become a trend. Benefiting from the advantages of the denoising diffusion probability model (DDPM) in natural image super-resolution tasks, we consider the structural characteristics of HSI to design a novel spectral super-resolution algorithm, named R2H-CCD, which is a hyperspectral imagery generation method from RGB images based on conditional cascade diffusion probabilistic models. More specifically, the algorithm takes RGB image as conditional input and synthesise corresponding HSI from pure noise through a stochastic iterative denoising process. In addition, to obtain high-fidelity HSIs, it adopts U-Net structure to iteratively refine the generated images in an end-to-end training manner. The experiments on HFD100 dataset show the effectiveness and superiority of the proposed method. Xiaoyan Luo |
IGARSS | 2 |
| 2023 | Learning from small data for hyperspectral image classification
Xiaoyan Luo, Jihao Yin |
Signal Process. | 1 |
| 2023 | Boosting Variational Inference With Margin Learning for Few-Shot Scene-Adaptive Anomaly DetectionabstractAnomaly detection in surveillance videos aims to identify frames where abnormal events happen. Existing approaches assume that the training and testing videos are from the same scene, exhibiting poor generalization performance when encountering an unseen scene. In this paper, we propose a Variational Anomaly Detection Network (VADNet), which is characterized by its high scene-adaptation - it can identify abnormal events in a new scene only via referring to a few normal samples without fine-tuning. Our model embodies two major innovations. First, a novel Variational Normal Inference (VNI) module is proposed to formulate image reconstruction in a conditional variational auto-encoder (CVAE) framework, which learns a probabilistic decision model instead of a traditional deterministic one. Secondly, a Margin Learning Embedding (MLE) module is leveraged to boost the variational inference and aid in distinguishing normal events. We theoretically demonstrate that minimizing the triplet loss in MLE module facilitates maximizing the evidence lower bound (ELBO) of CVAE, which promotes the convergence of VNI. By incorporating variational inference with margin learning, VADNet becomes much more generative that is able to handle the uncertainty caused by the changed scene and limited reference data. Extensive experiments on several datasets demonstrate that the proposed VADNet can adapt to a new scene effectively without fine-tuning and achieve remarkable performance, which outperforms other methods significantly and establishes new state-of-the-art in the case of few-shot scene-adaptive anomaly detection. We believe our method is closer to real-world application due to its strong generalization ability. All codes are released inhttps://github.com/huangxx156/VADNet. Yutao Hu 0002, Xiaoyan Luo, Jungong Han, Baochang Zhang 0001, Xianbin Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Multiscale Prototype Contrast Network for High-Resolution Aerial Imagery Semantic SegmentationabstractSemantic segmentation of high-resolution aerial images is a challenging task on account of complex scene-variation and large scale-difference. However, these two issues are inadequately addressed in general semantic segmentation methods. In this paper, we propose a Multi-scale Prototype Contrast Network (MPCNet) to improve the adaptive capability for different scenes and scales. Specifically, a novel multi-scale prototype transformer decoder (MPTD) is designed to extract dynamic scene-specific prototypes as pixel classifier by fusing information of feature maps and learnable class tokens. To exploit cross-scene context information and accommodate the large scale-difference in aerial image, we build a multi-scale prototype memory queue to store these multi-scale prototypes during training. Upon the multi-scale prototype memory queue, a novel multi-scale prototype contrastive loss is proposed to increase object feature discriminability across multiple scale, which brings better consistency of intermediate feature and boosts the convergence of network. Extensive experimental results on three publicly available datasets demonstrate the effectiveness and efficiency of our MPCNet over other state-of-the-art methods. The code is available at https://github.com/qixiong-wang/mmsegmentation-mpcnet. Qixiong Wang, Xiaoyan Luo, Jiaqi Feng 0001, Guangyun Zhang, Xiuping Jia, Jihao Yin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Spectral Transformer with Dynamic Spatial Sampling and Gaussian Positional Embedding for Hyperspectral Image ClassificationabstractOwing to the global information extraction ability, transformers have been tentatively applied to hyperspectral image(HSI) classification. However, the existing transformer-based methods have not made full use of the flexible characteristics of spatial sampling nor considered the importance of the central pixel to the classification of HSI cubes. In order to enhance adaptability of transformers for HSI classification, we have proposed a novel spectral transformer with dynamic spatial sampling and gaussian positional embedding. To improve the effectiveness of spatial neighborhood information, Spatial Sample Selection(3S) mechanism generates image cube from super pixel region, making image cube more pure for classification. To extract long-range information in spectral dimension, Spectral Feature Extraction(SFE) network splits spectral bands into several slices and calculates the attention between them. To stress the importance of the central pixel to the classification of image cube, Gaussian Positional Embedding(GPE) reduces the weight of surrounding pixels during feature embedding stage. Experimental results demonstrate the performance of our proposed method. The code of this work is available at https://github.com/fengjiaqi927/HSI_transformer. Jiaqi Feng 0001, Xiaoyan Luo, Qixiong Wang, Jihao Yin |
IGARSS | 2 |
| 2022 | Learning spatial structures of proteins improves protein-protein interaction predictionabstractSpatial structures of proteins are closely related to protein functions. Integrating protein structures improves the performance of protein-protein interaction (PPI) prediction. However, the limited quantity of known protein structures restricts the application of structure-based prediction methods. Utilizing the predicted protein structure information is a promising method to improve the performance of sequence-based prediction methods. We propose a novel end-to-end framework, TAGPPI, to predict PPIs using protein sequence alone. TAGPPI extracts multi-dimensional features by employing 1D convolution operation on protein sequences and graph learning method on contact maps constructed from AlphaFold. A contact map contains abundant spatial structure information, which is difficult to obtain from 1D sequence data directly. We further demonstrate that the spatial information learned from contact maps improves the ability of TAGPPI in PPI prediction tasks. We compare the performance of TAGPPI with those of nine state-of-the-art sequence-based methods, and TAGPPI outperforms such methods in all metrics. To the best of our knowledge, this is the first method to use the predicted protein topology structure graph for sequence-based PPI prediction. More importantly, our proposed architecture could be extended to other prediction tasks related to proteins. Bosheng Song, Xiaoyan Luo, Xiaoli Luo, Yuansheng Liu, Zhangming Niu, Xiangxiang Zeng |
Briefings Bioinform. | 2 |
| 2022 | Attentive encoder-decoder networks for crowd counting
Xuhui Liu, Yutao Hu 0002, Baochang Zhang 0001, Xiantong Zhen, Xiaoyan Luo, Xianbin Cao 0001 |
Neurocomputing | 5 |
| 2022 | Class-Balanced Contrastive Learning for Fine-Grained Airplane DetectionabstractAirplane detection and fine-grained recognition in remote sensing images are challenging due to class imbalance and high inter-class indistinction. To alleviate these issues, we propose class-balanced contrastive learning (CBCL) approach for airplane detection to exploit the correlation between samples in different images, which is rarely explored in previous research. Specifically, we first dynamically build class-balanced memory queues during training, which mitigates class imbalance by memorizing training samples. Upon class-balanced memory queues, hard triplet contrastive learning is introduced to increase the inter-class discriminability, which enforces the maximum distance of the positive sample pair to be smaller than the minimum distance of the negative sample pair. We integrate the proposed CBCL strategy into oriented object detection frameworks for fine-grained airplane detection. The experimental results on the FAIR1M dataset reveal that several state-of-the-art algorithms with CBCL achieve significantly improvements. Qixiong Wang, Xiaoyan Luo, Jihao Yin |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Cross-Domain Attention Network for Unsupervised Domain Adaptation Crowd CountingabstractUnsupervised domain adaptation crowd counting (UDACC) has been studied with practical research utility by getting rid of the labeling burden on large-scale dense crowds in the target domain. Current methods generalize well within the specific domain gap by directly aligning domain distributions or translating synthetic data to realistic images. However, it is difficult to define domain gaps among complex real-world datasets, in which the images vary greatly in style, density level and/or content. To tackle this problem, in this paper, we propose a Cross-Domain Attention Network (CDANet), which can effectively generalize the model to the unlabeled domain on both unsupervised synthetic-to-realistic and realistic-to-realistic crowd counting. Specifically, we propose a Cross-Domain Attention Module (CDAM) to learn domain-related information between the source and target domain, which extracts relations in cross-domain attentive information, thus enhancing crowd-informative features. Moreover, to make our CDAM invariant to domain shifts, we introduce a consistency penalty to ensure that the attention maps are consistent before and after the domain shifting. Thus our CDANet can pay attention to the shared counting information across domains, while remaining its invariant ability during domain adaptation. Extensive experiments on several common benchmarks for UDACC demonstrate that our CDANet gets competitive results on both unsupervised synthetic-to-realistic and realistic-to-realistic UDACC tasks. Jun Xu 0019, Xiaoyan Luo, Xianbin Cao 0001, Xiantong Zhen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Variational Self-Distillation for Remote Sensing Scene ClassificationabstractSupported by deep learning techniques, remote sensing scene classification, a fundamental task in remote image analysis, has recently obtained remarkable progress. However, due to the severe uncertainty and perturbation within an image, it is still a challenging task and remains many unsolved problems. In this paper, we note that regular one-hot labels cannot precisely describe remote sensing images, and they fail to provide enough information for supervision and limiting the discriminative feature learning of the network. To solve this problem, we propose a Variational Self-Distillation Network (VSDNet), in which the class entanglement information from the prediction vector acts as the supplement to the category information. Then, the exploited information is hierarchically distilled from the deep layers into the shallow parts via a Variational Knowledge Transfer (VKT) module. Notably, the VKT module performs knowledge distillation in a probabilistic way through variational estimation, which enables end-to-end optimization for mutual information and promotes robustness to uncertainty within the image. Extensive experiments on four challenging remote sensing datasets demonstrate that, with a negligible parameter increase, the proposed VSDNet brings a significant performance improvement over different backbone networks and delivers state-of-the-art results. Yutao Hu 0002, Xiaoyan Luo, Jungong Han, Xianbin Cao 0001, Jun Zhang 0007 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | H2AN: Hierarchical Homogeneity-Attention Network for Hyperspectral Image ClassificationabstractRecently, a self-attention network (SAN) is developed as an effective strategy to extract features from attention areas for image classification. However, for hyperspectral image (HSI), the lack of position supervision of object regions and inefficient similarity computation lead to unsatisfactory classification performance on mixed pixels. To alleviate the above two problems for HSI image classification, we propose a novel hierarchical homogeneity-attention network (H2AN) in this article. First, we design a homogeneity-attention block (HAB) to depict the feature correlation with the homogeneous mask. Using the supervision of homogeneity mask, we can calculate the attention guided by the predefined number of homogeneity embeddings, which can reduce the heavy computation instead of the global search in self-attention block (SAB) of SAN. Second, we propose a hierarchical convolutional neural network (HCNN) inserting the HAB into different levels of network cells for highly efficient feature extraction of target regions, named H2AN. Because of the transferring of homogeneity property from shallow layer to deep layer, our H2AN outperforms the state-of-the-art methods in qualitative and quantitative experiments on three typical datasets. Xiaoyan Luo, Qixiong Wang, Jihao Yin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Hyperspectral Classification Using Cooperative Spatial-Spectral Attention Network With Tensor Low-Rank ReconstructionabstractSpatial and spectral attention networks have been both well introduced to Hyperspectral image (HSI) classification. However, in previous works, they are seldom considered jointly. To obtain a 3D spatial-spectral attention map, which is beneficial for extracting discriminative spatial-spectral features, we propose a novel cooperative spatial-spectral attention network with tensor low-rank reconstruction. Firstly, a tensor low-rank reconstruction (TLRR) block is designed to learn a spatial-spectral attention map tensor, which adaptively emphasizes the attention features of the salient spatial positions and informative spectral bands simultaneously. Secondly, these attention features are merged into simple convolutional features which are more discriminative for classification. Finally, the experimental results demonstrate that our proposed method outperforms some state-of-the-art methods on two typical HSI datasets. Xiaoyan Luo, Qixiong Wang, Weifa Shen, Jihao Yin |
ICIP | 2 |
| 2021 | Adaptive Anomaly Detection Network for Unseen Scene Without Fine-Tuning
Yutao Hu 0002, Xiaoyan Luo |
PRCV (2) | 3 |
| 2021 | Multibranch Spatial-Channel Attention for Semantic Labeling of Very High-Resolution Remote Sensing ImagesabstractVery high-resolution (VHR) remote sensing images can provide fine but sometimes trivial ground object details; thus, the semantic labeling of VHR images is a challenging task. To improve the VHR labeling performance, spatial multiscale information and channel attention have been employed recently. However, the exploitation of global object features is still limited, which leads to the loss of capturing within-class variation from location to location. In this letter, we present a multibranch spatial-channel attention (MSCA) model to efficiently extract global dependency and combine it with multiscale and channel attention methods. In the spatial multiscale attention block, a multibranch feature fusion model is established to exploit the global relationship captured by self-attention and the multiscale correlation learned from dilated convolutions. To alleviate the computational cost of pixel-by-pixel self-attention operation, a spatial pyramid compressing method is also designed. In the channel attention block, average and max global pooling strategies are applied, respectively, in two channel attention branches to generalize global information from different perspectives. Those two blocks are then adaptively united by learnable weighting parameters. Experiments on two VHR image data sets demonstrate that the proposed network can yield better performance in comparison with state-of-the-art labeling methods tested. Bingnan Han, Jihao Yin, Xiaoyan Luo, Xiuping Jia |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Phone Keypad Voice Recognition (PKVR): An Integrated Experiment for Digital Signal Processing EducationabstractThis Innovative Practice Work-In-Progress presents an integrated signal processing experiment, which can cover most knowledge points of digital signal processing (DSP) course. Since the DSP course focuses on one dimension signal processing, voice signal has a great advantage. We provide an integrated voice signal processing experiment named as Phone Keypad Voice Recognition (PKVR), including the following parts: phone keypad voice collection, Discrete Fourier Transform (DFT) and analysis, filter design, digital query table establishment, number recognition of any keypad voice. Through the improvement of the practice training, the classroom teaching theory can be better understood in an interesting way for our students. Xiaoyan Luo, Han Wan, Fugen Zhou |
FIE | 1 |
| 2020 | A Comprehensive Experiment to Enhance Multidisciplinary Engineering Ability via UAVs Visual NavigationabstractThis Research to Practice WIP presents a UAVs visual navigation based comprehensive experiment to enhance multidisciplinary engineering ability in Aerospace engineering education. In traditional courses, aerospace-related disciplines are independently distributed in different courses, and there is rarely a hands-on platform which includes signal processing, control theory, and artificial intelligence into Aerospace engineering. Facing this problem, this paper designs a multidisciplinary comprehensive experiment, aiming to provide a hand-on platform and flexible project-based program to students of aerospace engineering professions. First of all, in order to let the students understand actual aerospace problems, a multidisciplinary simulation platform containing UAVs and remote objects scenarios is constructed for them to explore in the experiments. Second, the content of the experiment is designed into three stages including data acquisition and processing, conceptual design and simulation, in-flight validation, during which the multidisciplinary engineering ability runs through the whole process of the activities. Finally, Project Oriented Design Based Learning is also introduced here to combine engineering design education with innovation and creativity. Through the project demonstration and presentation at the end of the experiment, the multidisciplinary engineering ability of each student can be effectively evaluated. The UVN comprehensive experiment enables students to work on real-world aerospace engineering problems through a hardware-software integration framework, which may greatly stimulate their curiosity and interest in autonomously learning. It also provides students unprecedented opportunities to immerse themselves in projects that cross disciplinary boundaries, improve their professional ability and enhance their exploration competence in aerospace areas. Xiaoyan Luo, Han Wan, Chengxi Wu, Yu Zheng 0017, Fugen Zhou |
FIE | 2 |
| 2020 | Neural Network Pruning for Hyperspectral Image Band SelectionabstractNeural network pruning attempts to reduce parameters without hurting original performance by inducing connection matrix sparsity of network. Inspired by this idea, we proposed an effective pruning-based band selection strategy, which is a potent feature extraction tool in hyperspectral image (HSI) classification. At first, we take the whole HSI bands as input to train original network parameters. For each band, all parameters in network are integrated to measure the band importance. With the novel band signification factor constraining, then the convolutional neural network (CNN) is pruned and remains some representative weights to retrain the compact sub-network, which can finally deal with the hyperspectral band selection problem. Experimental results on the real HSI dataset demonstrate that network pruning-based method can outperform the original CNN in classification accuracy. Also, it can achieve the superiority over filter-based and other CNN-based band selection algorithms in classification accuracy. Our code is available at https://github.com/qixiong-wang/Network-pruning-for-HSI-band-selection. Qixiong Wang, Xiaoyan Luo, Jihao Yin |
IGARSS | 2 |
| 2020 | LG: A clustering framework supported by point proximity relations
Hui Qv, Jihao Yin, Xiaoyan Luo |
Pattern Recognit. | 3 |
| 2019 | Look for Saliency In Hyperspectral ImagesabstractImage saliency detection plays an important role in many vision tasks such as object detection, visualization, and classification. Many existing methods are applicable to trichromatic or grayscale images, but most of them are failure to remote sensing hyperspectral images with hundreds of bands. In this paper, a novel purity-aware saliency method is specifically proposed for hyperspectral images. Firstly, an autoencoder model for unmixing is designed to extract the abundance. Then, an abundance fraction vector norm is calculated as pixel spectral purity. On real-world hyperspectral images, our saliency detection method achieves good performance in holding structure information and purity distribution. Zhiqi Shen 0003, Xiaoyan Luo, Rui Xue 0003 |
IGARSS | 2 |
| 2019 | Hyperspectral Image Classification Using CapsNet With Well-Initialized Shallow LayersabstractIn this letter, an alternative data-driven HSI classification model based on CapsNet is proposed rather than recently predominant convolutional neural network (CNN)-based models. To adjust the CapsNet to HSI classification, we tune a new CapsNet architecture with three convolutional layers. The added shallow layer provides higher level features to the primary capsules, which indirectly speeds up the following routing procedure. To guarantee a good convergence of the whole CapsNet, the three shallow layers are initialized by transferring convolutional parameters from a pretrained CNN model. The improved CapsNet-based models with and without vote strategy both achieve significantly superior performance in HSI classification to the state-of-the-art CNN-based methods on real hyperspectral data sets. Jihao Yin, Hongmei Zhu, Xiaoyan Luo |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2018 | Spectral Diversity Enhancement for PansharpeningabstractPansharpening is to generate a synthetic image with high spatial resolution and high spectral resolution via fusing panchromatic (PAN) and multispectral (MS) images. In most traditional pansharpening methods, the original MS image is firstly interpolated to the same size of PAN image by analytical interpolations. However, these interpolated methods could cause false spectral information due to ignoring mixed spectral characteristics in the MS image. To enhance the spectral diversity of upsampled MS image, we use the spatial structure information in PAN image to support the pansharpening in this paper. By introducing the superpixel structure for PAN image, the processible mixed pixels can be screened out in corresponding MS locations, and the other MS locations are considered to be occupied by pure pixels. For the pure and mixed pixels, their upsampling MS results can be obtained via a directly expending manner and a sparse representation manner respectively. Two different detail injection strategies are used for assessing the performance of analytical interpolations and our approach for pansharpening. Experimental results demonstrate that our method achieves the appreciable improvements with respect to analytical interpolations. Liangyu Zhou, Xiaoyan Luo, Jihao Yin |
ICIP | 2 |
| 2018 | Band Dual Density Discrimination Analysis for Hyperspectral Image ClassificationabstractA novel band discrimination analysis framework for hyperspectral image (HSI) supervised classification is proposed based on dual density (DD). Different from the popular supervised band selection (BS) approaches which measure the discrimination among classes under multivariate normal distribution hypothesis, our work infers the class discrimination degree (overlapping extent) for valid extraction of band subset without any assumed distribution. In the proposed framework, it is crucial to find indexes to measure the discrimination degree of each band, and therefore we develop the DD indexes, including the homogeneity density and the heterogeneity density. Viewing each band of the HSI as a data set, i.e., the data points in each data set are 1-D, and we first obtain the DD value pairs for all data points in each data set. Then, for each data set, we determine its discrimination degree using DD-based zone ratio or score quantify strategy. Finally, the bands, which are determined as the nonoverlapped or have high scores, are chosen as the band subset for the subsequent classification. Superiorities of the proposed BS are demonstrated on the three real-world HSIs over several well-known BS algorithms in terms of classification accuracy and speed. Hui Qv, Jihao Yin, Xiaoyan Luo, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Corrections to "Segment-Oriented Depiction and Analysis for Hyperspectral Image Data"abstractIn[1], information regarding the corresponding author is missing. The information is updated here. The updated footnote below shows that Xiaoyan Luo is the corresponding author for this paper. Jihao Yin, Hui Qv, Xiaoyan Luo, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | A hierarchical superpixel aggregation model for hyperspectral imageabstractSuperpixel has been widely applied in hyperspectral image processing as a pre-processing step for over-segmentation. However, most superpixel algorithms are difficult to control the segmentation balance between fragmentation and accuracy. In this paper, we propose a superpixel aggregation model to cluster the over-segmentations. Based on the own importance and interrelationship of superpixels, a two-step merging procedure is designed in the hierarchical wise from local to global comparisons. Aiding by a density peak metric, which is to exploit the spectral correlation in hyperspectral image, the similar neighbor superpixels are merged firstly, and then the similar regions in discontinuous spatial location are gathered. Experimental results show that the proposed model can achieve high accuracy in low region number compared with original superpixel algorithm, and the performance for unsupervised classification application is also remarkable. Bingnan Han, Jihao Yin, Xiaoyan Luo, Hui Qv |
IGARSS | 3 |
| 2017 | Information-Assisted Density Peak Index for Hyperspectral Band SelectionabstractBand selection has become an effective method to reduce hyperspectral dimensionality. In this letter, an information-assisted density peak index (IaDPI) is proposed to prioritize the bands. Based on a clustering method by finding density peaks, IaDPI introduces the intraband information entropy into the local density and intercluster distance to ensure cluster centers with a high quality. Also, the band distance is integrated with channel proximity to control the compactness of local density. Owing to the intraband entropy and the interband weighted dissimilarity, the selected band set with top-ranked IaDPI scores can hold high local density, clear global distinction, and good informative quality. Experimental results on real hyperspectral data indicate the advantages of the proposed IaDPI in good selection quality, robust noise immunity, and high classification accuracy. Xiaoyan Luo, Rui Xue 0003, Jihao Yin |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Segment-Oriented Depiction and Analysis for Hyperspectral Image DataabstractA novel segment-oriented dictionary learning (SeODL) framework for hyperspectral image (HSI) classification is proposed. Differing from existing HSI classification methods which directly process the original whole spectral curves of pixels, our work focuses on local segment analysis to achieve fine depiction and effective exploitation. Viewing the separated segment as a basic processing unit, we first cluster them into two sets with the homogeneity in trend and fluctuation, and then two small dictionaries can be quickly learned. Second, to get meticulous and discriminability enhanced segment-oriented representations (SORs), the segments of the training and test pixels are coded on a novel binary-separated coding strategy. The coding stage for obtaining SORs is sped up by the employment of our proposed enhanced orthogonal matching pursuit technique. A characteristic splicing classifier with high performance can be trained using these SORs of the training pixels. Finally, a spiral searching strategy and a multiple majority-voting method are adopted for fully spatial information incorporation of the test pixels whose final SORs will be embedded into the trained characteristics splicing classifier to ascertain the labels. Experimental results on three real HSI data sets demonstrate the superiority of the proposed SeODL framework over several well-known classification algorithms in terms of classification accuracies. Jihao Yin, Hui Qv, Xiaoyan Luo, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Human indoor localization based on ceiling mounted PIR sensor nodesabstractThis paper presents a human indoor localization system using ceiling mounted pyroelectric infrared (PIR) sensors. The field of views (FOVs) of the PIR sensors is modulated by two degrees of freedom (DOF) of spatial segmentation. The localization algorithm is proposed to fuse the data stream generated from different sensor nodes within the wireless network. The Kalman Filter and Kalman Smoother are utilized to refine the estimation of the human position. We conduct experiments in a real office environment, and the average root-mean-square error (RMSE) of single human target tracking at different speed is about 0.6 meter. The promising results confirm the efficacy of our system. Xiaomu Luo, Tong Liu 0002, Baihua Shen, Qinqun, Liwen Gao, Xiaoyan Luo |
CCNC | 6 |
| 2016 | Planet mineral distribution detection via clustering-aware nonnegative matrix factorizationabstractSpectral unmixing is an important technique to exploit mineral distribution through remote sensing image. In this paper, we propose an unmixing algorithm combining clustering-aware method with the sparsity-constrained nonnegative matrix factorization (SNMF) algorithm. Pixels with similar spectra have high possibility to share similar typical endmembers, therefore we preprocess the image using K-means cluster algorithm and then optimizes the initial endmember spectra by selecting the typical ones of each cluster as the initial endmember value. Due to the local convergence feature of NMF, the optimal initial value can accelerate the convergence of the algorithm and obtain more accurate results. Meanwhile, we use the sparsity-constrained NMF in global unmixing to control the sparse property of abundance distribution. The experiments on synthetic data and Chang'e-1 hyperspectral data show that K-means nonnegative matrix factorization (KNMF) is superior to the other unmixing methods. Jihao Yin, Xiaoyan Luo, Hui Qv, Bingnan Han |
IGARSS | 3 |
| 2016 | No-Reference Assessment on Haze for Remote-Sensing ImagesabstractAssessment on haze can filter out images with dense haze to improve the reliability of remote-sensing image interpretation. In this letter, a novel no-reference haze assessment method based on haze distribution is proposed for remote-sensing images. First, range channel of an image is defined and the haze distribution map (HDM) is extracted from the hazy image. Then, the haze assessment metric HDM-based haze assessment (HDMHA) is designed according to the HDM. Finally, the degree of haze in remote-sensing images is predicted using the proposed metric. In order to objectively verify the effectiveness of the proposed metric HDMHA, a method of simulating hazy remote-sensing images based on the haze imaging model is proposed in this letter, and the simulated hazy images are greatly similar to real ones in vision. A series of experiments are done on both real images and simulated images, and the results show that the proposed metric achieves good consistency when compared with subjective experiments and outperforms typical blind image quality assessment methods. Xiaoxi Pan, Fengying Xie, Zhiguo Jiang 0001, Zhenwei Shi 0001, Xiaoyan Luo |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2015 | Cooperative service model innovation and pilot application in the prevention and control on diabetic retinopathy with web-movable hand-hold fundus camera in regional health and medicine of digitizationabstractObjective: Systems engineering theory, hand-hold fundus camera(FC) and internet cooperative platform were applied to help community staff, long distance specialists' organizations and teams to cooperate on the screening for diabetic fundus pathological change(or diabetic retinopathy, DR) in regional residents. It aims at promoting the ability to prevent and control (PC) the important non-infectious chronic diseases (NCD) to design a proper value-chain for mutually beneficial mode of interactions among regional health services, and the other corresponding social roles via technological and organizational innovation. DongXu Jiang, Xiaoping Lai, XueLian Mo, Qianrong Liang, Xiaoyan Luo, LiWen Zhou, ZhongXun Yuan, JianBiao Xu, JuanHong Zeng, ShouZeng Zheng, CanDong Li |
BIBM | 6 |
| 2015 | An algorithm for the diversity of XML document based on the model of the rights-distribution for ancestral property and its testing application in the field of health and Chinese MedicineabstractTo seek a method and its potential algorithm of the differentiation of the health status and the intelligent decision for the road-line of the transformation among different healthy status, three procedures were designed: 1. Extensible Markup Language (XML) documents were tree-shape-like structured and modeled from a kind of rule for the distribution of ancestral property rights. 2. An algorithm with a kind of nuclear functions, named as HuangDiNeiJing Tree for the model, was established and a program for its application was designed for calculating the different degree between the matching XML documents and a router-algorithm of Dikjstra, was applied to the analysis on transformation among different healthy status. 3. Ten syndromes of Chinese Medicine(CM) named according to GB/T15657 were selected to test the model and algorithms. The result showed that the algorithm for calculating the weight of element and the different value came from the paired-matching XML documents may distinguishably and accurately reflect their diversities of modeled trees or XML documents. It may also apply to the establishment for roadmap of the transformation among different health status, beneficial to education, research and practices with intelligence auxiliary decision in the field of CM, health and biomedicine. Yongyun Wang, Qianrong Liang, Xiaoyan Luo, Lini Wang, CanDong Li |
BIBM | 4 |
| 2015 | Hybrid fusion and interpolation algorithm with near-infrared image
Xiaoyan Luo, Jun Zhang 0007, Qionghai Dai |
Frontiers Comput. Sci. | 1 |
| 2014 | Shadow detection in remote sensing images based on weighted edge gradient ratioabstractThis paper presents a novel shadow detection method in remote sensing images based on edge feature description of candidate regions. Edge gradient ratio is defined and used to represent the inherent properties of shadow regions. To improve the detection result, weighted edge gradient ratio (WEGR) is addressed, where the weight of a region is determined by the number of pixels belonging to shadow in the region, and edge gradient ratio is proposed to describe the edge feature surrounding the region. Experiments and comparisons indicate that our method achieves better accuracy both on high and low quality remote sensing images. Bin Pan, Zhiguo Jiang 0001, Xiaoyan Luo |
IGARSS | 4 |
| 2014 | Novel infrared and visible image fusion method based on independent component analysis
Yin Lu, Fuxiang Wang, Xiaoyan Luo |
Frontiers Comput. Sci. | 3 |
| 2013 | An auscultatory technique of Chinese medicine: Pattern recognition based on timbre of human-voice matching with standardized patterns of sound from Bianzhong of Marquis Yi of Zeng (***)abstractIn order to find a sensitive and stable auscultatory method, recognition algorithm for human-voice matching with 25 patterns for 25Yin, a pattern library with 25 sounds as standard patterns was established in this manuscript. Furthermore, Mel Frequency Cepstrum Coefficient (MFCC) was applied to analyse for sampled sound and then the algorithm of Support Vector Machine (SVM) was used to match the sampled sound, which had been dipt the baseline signal and the signal at the part of consonants in each single pronunciation during voice sampling, with every sound in the pattern library. A table had been listed with 25 nomenclature for 25 Yin described in HDNJ and related to the code-named (CN) of the wave sound from BMYZ. The recognition algorithm based on MFCC plus SVMto match human-voice and the pattern library with 25 sounds offered a well-accuracy well-precision and potential technique to recognize different individuals according to the classic theory. Xiaoyan Luo, Lingli Wang, Qianrong Liang, Yongyun Wang, Xiaoping Lai |
BIBM | 2 |
| 2012 | A regional image fusion based on similarity characteristics
Xiaoyan Luo, Jun Zhang 0007, Qionghai Dai |
Signal Process. | 1 |
| 2009 | Image fusion in compressed sensingabstractThis paper proposes an efficient image fusion scheme for compressed sensing (CS) imaging, in which fusion is performed on the random projections before reconstruction. Specifically, the measurements of multiple input images are fused into composite measurements via weighted average, in which the weights are calculated based on entropy metrics of the original measurements. Then the fused image with transformation coefficients in a selected basis is reconstructed from the composite measurements by the gradient projection for sparse reconstruction (GPSR) algorithm. The proposed scheme is implemented in a block-based CS framework. Simulation results show that our scheme provides promising fusion performance with a low computational complexity. Xiaoyan Luo, Jun Zhang 0007, Jing-Yu Yang 0002, Qionghai Dai |
ICIP | 1 |