EDBT 2026 Demo / reviewers in the wild / expert
Ying Qu 0001
dblp:51/2820-1
· DBLP profile ↗
28ranked-venue papers
11as first author
14since 2021 · last 2023
0000-0002-4613-8625ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Unsupervised Hyperspectral Image Domain Adaptation through Unmixing-Based Domain AlignmentabstractDespite the great progress in hyperspectral image classification, it is still challenging due to the unique characteristic of satellite imagery where the training and test sets may come from different distributions because of the different acquisition conditions. Hence, directly deploying the trained model on the test data may lead to degradation in the performance. In this work, we propose an unsupervised domain adaptation approach that aligns distributions across the training and test domains. It projects the data to a shared embedding space, i.e., the abundance space, that is regularized by physical constraints. The shared abundance space, together with a metricbased distribution alignment approach applied on the abundance space, would largely reduce the domain discrepancy and provide a more representative feature set for classification purpose. Experimental results on hyperspectral benchmarks demonstrate superiority of the proposed method. Razieh Kaviani Baghbaderani, Ying Qu 0001, Hairong Qi 0001 |
IGARSS | 2 |
| 2023 | Topological Structure and Semantic Information Transfer Network for Cross-Scene Hyperspectral Image ClassificationabstractDomain adaptation techniques have been widely applied to the problem of cross-scene hyperspectral image (HSI) classification. Most existing methods use convolutional neural networks (CNNs) to extract statistical features from data and often neglect the potential topological structure information between different land cover classes. CNN-based approaches generally only model the local spatial relationships of the samples, which largely limits their ability to capture the nonlocal topological relationship that would better represent the underlying data structure of HSI. In order to make up for the above shortcomings, a Topological structure and Semantic information Transfer network (TSTnet) is developed. The method employs the graph structure to characterize topological relationships and the graph convolutional network (GCN) that is good at processing for cross-scene HSI classification. In the proposed TSTnet, graph optimal transmission (GOT) is used to align topological relationships to assist distribution alignment between the source domain and the target domain based on the maximum mean difference (MMD). Furthermore, subgraphs from the source domain and the target domain are dynamically constructed based on CNN features to take advantage of the discriminative capacity of CNN models that, in turn, improve the robustness of classification. In addition, to better characterize the correlation between distribution alignment and topological relationship alignment, a consistency constraint is enforced to integrate the output of CNN and GCN. Experimental results on three cross-scene HSI datasets demonstrate that the proposed TSTnet performs significantly better than some state-of-the-art domain-adaptive approaches. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE_TNNLS_TSTnet. Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ying Qu 0001, Ran Tao 0003, Hairong Qi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Cross-Guided Feature Fusion with Intra-Modality Reweighting for Multi-Spectral Pedestrian DetectionabstractMulti-spectral pedestrian detection has gained extensive attention over the past decade. To alleviate the problem of modality imbalance in the multi-spectral tasks, a novel cross-guided feature fusion network based on the auto-encoder framework is proposed using RGB-thermal image pairs as inputs. To obtain the complementary features, a cross-guided loss is designed, so that the output images are balanced with both modalities in an unsupervised manner. An intra-modality reweighting module is implemented to filter the redundant features before the fusion. Finally, YOLOv3 is chosen as the detector fed by the fused features. The proposed method is verified using the public KAIST and VOT-RGBT datasets. Experimental results demonstrate that the proposed method can outperform the state-of-the-art methods, the miss rate of pedestrian detection reaches 48.57% and 4.52% using KAIST and VOT-RGBT datasets, respectively. Zhenzhou Shao, Ying Qu 0001, Jun Zhang 0031, Zhi-Ping Shi 0002 |
ICPR | 5 |
| 2022 | Non-Local Representation Based Mutual Affine-Transfer Network for Photorealistic StylizationabstractPhotorealistic stylization aims to transfer the style of a reference photo onto a content photo in a natural fashion, such that the stylized image looks like a real photo taken by a camera. State-of-the-art methods stylize the image locally within each matched semantic region and are prone to global color inconsistency across semantic objects/parts, making the stylized image less photorealistic. To tackle the challenging issues, we propose a non-local representation scheme, constrained with a mutual affine-transfer network (NL-MAT). Through a dictionary-based decomposition, NL-MAT is able to successfully decouple matched non-local representations and color information of the image pair, such that the context correspondence between the image pair is incorporated naturally, which largely facilitates local style transfer in a global-consistent fashion. To the best of our knowledge, this is the first attempt to address the photorealistic stylization problem with a non-local representation scheme, such that no additional models or steps for semantic matching are required during stylization. Experimental results demonstrate that, the proposed method is able to generate photorealistic results with local style transfer while preserving both the spatial structure and global color consistency of the content image. Ying Qu 0001, Zhenzhou Shao, Hairong Qi 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Multispectral Scene Classification via Cross-Modal Knowledge DistillationabstractScene classification is a fundamental task for numeral remote sensing applications, which aims to assign semantic labels to image patches. Although deep neural networks (DNN) demonstrated unique strength in scene classification, their performances are still limited due to the lack of training samples in the remote sensing field. Recent studies show that the performance of scene classification can be improved by taking advantage of the knowledge transferred from models pre-trained on RGB images. However, the modalities differences between input images hinder the knowledge transfer across models, especially when the input of the models has distinct spectral bands. To tackle the challenges, we propose a cross-modal knowledge distillation framework to improve the performance of multispectral scene classification by transferring the prior knowledge from teacher models pre-trained on RGB images to the student network with limited samples. Moreover, a teacher assistant (TA) network is introduced to further improve the classification performance by bridging the gap between the teacher and student networks. The proposed strategy is evaluated on models with multimodality inputs with distinct spectral bands and demonstrates superior performance as compared to the state-of-the-art methods. Ying Qu 0001, Liqiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Unsupervised and Unregistered Hyperspectral Image Super-Resolution With Mutual Dirichlet-NetabstractHyperspectral images (HSIs) provide rich spectral information that has contributed to the successful performance improvement of numerous computer vision and remote sensing tasks. However, it can only be achieved at the expense of images’ spatial resolution. HSI super-resolution (HSI-SR), thus, addresses this problem by fusing low-resolution (LR) HSI with the multispectral image (MSI) carrying much higher spatial resolution (HR). Existing HSI-SR approaches require the LR HSI and HR MSI to be well registered, and the reconstruction accuracy of the HR HSI relies heavily on the registration accuracy of different modalities. In this article, we propose an unregistered and unsupervised mutual Dirichlet-Net ($u^{2}$-MDN) to exploit the uncharted problem domain of HSI-SRwithout the requirement of multimodality registration. The success of this endeavor would largely facilitate the deployment of HSI-SR since registration requirement is difficult to satisfy in real-world sensing devices. The novelty of this work is threefold. First, to stabilize the fusion procedure of two unregistered modalities, the network is designed to extract spatial information and spectral information of two modalities with different dimensions through a shared encoder–decoder structure. Second, the mutual information (MI) is further adopted to capture the nonlinear statistical dependencies between the representations from two modalities (carrying spatial information) and their raw inputs. By maximizing the MI, spatial correlations between different modalities can be well characterized to further reduce the spectral distortion. We assume that the representations follow a similar Dirichlet distribution for their inherent sum-to-one and nonnegative properties. Third, a collaborative$l_{2,1}$-norm is employed as the reconstruction error instead of the more common$l_{2}$-norm to better preserve the spectral information. Extensive experimental results demonstrate the superior performance of$u^{2}$-MDN as compared to the state of the art. Ying Qu 0001, Hairong Qi 0001, Chiman Kwan, Naoto Yokoya, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Siamese Transformer Network for Hyperspectral Image Target DetectionabstractHyperspectral target detection can be described as locating targets of interest within a hyperspectral image based on prior information of targets. The complexity of actual scenes limits the performance of traditional statistical methods that rely on model assumptions, and traditional machine learning methods rely on mapping functions with limited complexity. To address these problems, we propose a Siamese transformer network for hyperspectral image target detection (STTD). The contribution of this article is threefold. First, we propose a novel method of constructing training samples using only the image itself and the limited prior information, which is suitable for target detection based on the Siamese network framework. Second, the Siamese network framework is utilized to solve the problem of similarity metric learning, i.e., make homogeneous features as close as possible and heterogeneous features as far as possible. Third, the most state-of-the-art network, transformer, is applied as the backbone of our proposed Siamese network to extract global features from spectra with long-range dependencies to achieve target detection. Furthermore, we make adaptive improvements to transformer for hyperspectral images. The proposed method shows its unique advantages in suppressing the background to a low level and highlighting the target with high probability. Experiments on five different datasets demonstrate the superiority of the proposed STTD as compared to the state-of-the-art. Weiqiang Rao, Lianru Gao, Ying Qu 0001, Xu Sun 0005, Bing Zhang 0001, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Ensemble-Based Information Retrieval With Mass Estimation for Hyperspectral Target DetectionabstractGiven the prior information of the target, hyperspectral target detection focuses on exploiting spectral differences to separate objects of interest from the background, which can be treated as information retrieval (IR) task in machine learning (ML). Most traditional detection methods work in the original feature space and rely heavily on specific assumptions, which cannot guarantee effective extraction of features for the target and background in hyperspectral images (HSIs). Mass estimation (ME) is a base modeling mechanism that has been proven to effectively solve problems in IR and is not restricted by specific assumptions. In this article, we propose a novel target detection method through ensemble-based IR with ME (EIRME). By directly deriving the ordering from a sample set to rank data points, ME provides a simple and straightforward ranking measure to ensure that points similar to the given target are far away from dissimilar points. For the estimation of mass distribution, the proposed method utilizes a tree-structured mapping to generate a feature space, in which the separability of the target and background is further improved. In particular, to break through the technical difficulty that the direct migration of IR methods with mass measure cannot specifically meet the high-precision requirements of target detection in HSIs, we develop a specialized measurement, topological mass, which innovatively combines the mass measure with tree topology to quantify the spectral difference for detection output. Moreover, the IR with ME based on parallel measurements through ensemble trees provides a robust solution with better generalization capacity and higher precision for hyperspectral target detection, facilitating practical applications. Experimental results on benchmark HSI datasets prove that the specialized measurement that we developed successfully overcomes the drawbacks of the direct migration of IR methods with ME and exhibits unique advantages. In addition, comparisons with the most classic and advanced detection algorithms demonstrate the superiority of the proposed method. Ying Qu 0001, Lianru Gao, Xu Sun 0005, Hairong Qi 0001, Bing Zhang 0001, Ting Shen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Endmember-Assisted Camera Response Function Learning, Toward Improving Hyperspectral Image Super-Resolution PerformanceabstractThe camera response function (CRF) that projects hyperspectral radiance to the corresponding RGB images is important for most hyperspectral image super-resolution (HSI-SR) models. In contrast to most studies that focus on improving HSI-SR performance through new architectures, we aim to prevent the model performance drop by learning the CRF of any given HSIs and RGB image from the same scene in an unsupervised manner, independent of the HSI-SR network. Accordingly, we first decompose the given RGB image into endmembers and an abundance map using the Dirichlet autoencoder architecture. Thereafter, a linear CRF learning network is optimized to project the reference HSIs to the RGB image that can be similarly decomposed like the given RGB , assuming that objects in both images share the same endmembers and abundance map. The quality of the RGB images generated from the learned CRFs is compared with that of the corresponding ground-truth images based on the true CRFs of two consumer-level cameras, Nikon 700D and Canon 500D. We demonstrate that the effectively learned CRFs can prevent significant performance drop in three popular HSI-SR models on RGB images from different categories of standard datasets of CAVE, ICVL, Chikuei, Cuprite, Salinas, and KSC. The successfully learned CRF using the method proposed in this study would largely promote a wider implementation of HSI-SR models since tremendous performance drop can be prevented practically. Jiangsan Zhao, Ying Qu 0001, Seishi Ninomiya, Wei Guo 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Prior-Based Tensor Approximation for Anomaly Detection in Hyperspectral ImageryabstractThe key to hyperspectral anomaly detection is to effectively distinguish anomalies from the background, especially in the case that background is complex and anomalies are weak. Hyperspectral imagery (HSI) as an image–spectrum merging cube data can be intrinsically represented as a third-order tensor that integrates spectral information and spatial information. In this article, a prior-based tensor approximation (PTA) is proposed for hyperspectral anomaly detection, in which HSI is decomposed into a background tensor and an anomaly tensor. In the background tensor, a low-rank prior is incorporated into spectral dimension by truncated nuclear norm regularization, and a piecewise-smooth prior on spatial dimension can be embedded by a linear total variation-norm regularization. For anomaly tensor, it is unfolded along spectral dimension coupled with spatial group sparse prior that can be represented by the${l}_{2,1}$-norm regularization. In the designed method, all the priors are integrated into a unified convex framework, and the anomalies can be finally determined by the anomaly tensor. Experimental results validated on several real hyperspectral data sets demonstrate that the proposed algorithm outperforms some state-of-the-art anomaly detection methods. Lu Li 0005, Wei Li 0032, Ying Qu 0001, Chunhui Zhao 0003, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Fast and Unsupervised Non-Local Feature Learning for Direct Volume Rendering of 3D Medical ImagesabstractTo improve the efficiency of medical visualization for computer aided surgery, we propose a fast and unsupervised 3D-CNN based non-local feature learning network. The proposed network consists of an encoder structure and a decoder structure. The encoder of the network projects the cube into a high-dimensional feature space, and the decoder of the network reconstructs the cube from the feature space. The decoder of the network serves as a dictionary shared by the cube to enforce the features for similar parts to be similar although they may distribute at disjointed locations. With such structures, the network is able to extract non-local features of the entire data. Moreover, a sparse constraint is incorporated into the network to increase the discriminative of the non-local features. Then the extracted non-local features of each voxel are fused with the corresponding position matrix and Hessian matrix for the voxel classification using Random Forest. Finally, a multidimensional transfer function is designed to enable the volume rendering. Experimental results demonstrate that the proposed method outperforms the state-of-the-art methods with much less training time. Xinmei Fu, Zhenzhou Shao, Ying Qu 0001, Yibo Zou, Zhi-Ping Shi 0002, Jindong Tan |
IROS | 3 |
| 2021 | Physically Constrained Transfer Learning Through Shared Abundance Space for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is one of the most active research topics and has achieved promising results boosted by the recent development of deep learning. However, most state-of-the-art approaches tend to perform poorly when the training and testing images are on different domains, e.g., the source domain and target domain, respectively, due to the spectral variability caused by different acquisition conditions. Transfer learning-based methods address this problem by pretraining in the source domain and fine-tuning on the target domain. Nonetheless, a considerable amount of data on the target domain has to be labeled and nonnegligible computational resources are required to retrain the whole network. In this article, we propose a new transfer learning scheme to bridge the gap between the source and target domains by projecting the HSI data from the source and target domains into a shared abundance space based on their own physical characteristics. In this way, the domain discrepancy would be largely reduced such that the model trained on the source domain could be applied to the target domain without extra efforts for data labeling or network retraining. The proposed method is referred to as physically constrained transfer learning through shared abundance space (PCTL-SAS). Extensive experimental results demonstrate the superiority of the proposed method as compared to the state of the art. The success of this endeavor would largely facilitate the deployment of HSI classification for real-world sensing scenarios. Ying Qu 0001, Razieh Kaviani Baghbaderani, Wei Li 0032, Lianru Gao, Yuxiang Zhang 0005, Hairong Qi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Unsupervised Pansharpening Based on Self-Attention MechanismabstractPansharpening is to fuse a multispectral image (MSI) of low-spatial-resolution (LR) but rich spectral characteristics with a panchromatic image (PAN) of high spatial resolution (HR) but poor spectral characteristics. Traditional methods usually inject the extracted high-frequency details from PAN into the upsampled MSI. Recent deep learning endeavors are mostly supervised assuming that the HR MSI is available, which is unrealistic especially for satellite images. Nonetheless, these methods could not fully exploit the rich spectral characteristics in the MSI. Due to the wide existence of mixed pixels in satellite images where each pixel tends to cover more than one constituent material, pansharpening at the subpixel level becomes essential. In this article, we propose an unsupervised pansharpening (UP) method in a deep-learning framework to address the abovementioned challenges based on the self-attention mechanism (SAM), referred to as UP-SAM. The contribution of this article is threefold. First, the SAM is proposed where the spatial varying detail extraction and injection functions are estimated according to the attention representations indicating spectral characteristics of the MSI with subpixel accuracy. Second, such attention representations are derived from mixed pixels with the proposed stacked attention network powered with a stick-breaking structure to meet the physical constraints of mixed pixel formulations. Third, the detail extraction and injection functions are spatial varying based on the attention representations, which largely improves the reconstruction accuracy. Extensive experimental results demonstrate that the proposed approach is able to reconstruct sharper MSI of different types, with more details and less spectral distortion compared with the state-of-the-art. Ying Qu 0001, Razieh Kaviani Baghbaderani, Hairong Qi 0001, Chiman Kwan |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Target Detection Through Tree-Structured Encoding for Hyperspectral ImagesabstractTarget detection aims to locate targets of interest within a specific scene. The traditional model-driven detectors based on signal processing have proved to be very effective. However, the detection performance of such traditional methods relies heavily on the model assumption, which is limited by the discrepancy with real hyperspectral images (HSIs) data. In this article, a target detection method through tree-structured encoding (TD-TSE) for HSIs is proposed. Instead of modeling the target and the background to extract valid features, we construct a binary tree based on the features of the data itself and segment the HSI to improve the separability of the target and the background. For the purpose of highlighting the target and suppressing the background, a novel measurement of separation, distance on tree, is calculated via binary encoding based on the constructed tree structure, and the detection output can be obtained according to such distance. To further reduce the generalization error resulting from random subsampling, the statistical average of the distances on multiple independent trees is estimated to improve the robustness of TD-TSE. The proposed method is not constrained by any model assumptions, which is fundamentally different from the most widely used hyperspectral target detectors in the field of signal processing. Moreover, the construction of binary trees without any labeled samples and the linear complexity of the proposed method make it highly practical for the hyperspectral data in real scenes. Extensive experiments on three benchmark HSI data sets demonstrate the effectiveness of the proposed TD-TSE for hyperspectral target detection. Ying Qu 0001, Lianru Gao, Xu Sun 0005, Hairong Qi 0001, Bing Zhang 0001, Ting Shen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Representative-Discriminative Learning for Open-Set Land Cover Classification of Satellite Imagery
Razieh Kaviani Baghbaderani, Ying Qu 0001, Hairong Qi 0001, Craig Stutts |
ECCV (30) | 2 |
| 2020 | Hyperspectral Nonlinear Unmixing via Generative Adversarial NetworkabstractHyperspectral nonlinear unmixing (HNU) is an extremely challenging problem as it is very difficult, if possible at all, to derive an explicit model to describe the underlying nonlinear mixing process. This paper gives the first attempt to tackle this problem by taking advantage of recent advances in deep learning, in specific, the development in generative adversarial network (GAN). The biggest contribution of GAN is that upon training, the network can generate samples with the same probabilistic distribution as that of the training samples, without explicitly knowing what the distribution actually is. Hence, we ask a similar question: can we unmix a hyperspectal image without explicitly knowing the nonlinear mixing model? In order to test this hypothesis, this paper proposes a data-driven supervised HNU method as compared to the traditional model-based approaches and uses a specific GAN framework, CycleGAN to solve the challenging nonlinear unmixing problem. We exploit the linkage between the cycle consistency loss used in CycleGAN and the spectral reconstruction loss used in traditional methods. We make the essential discovery that the usage of the cycle consistency loss enables the learning of the mixing and unmixing processes to be dependent on the training data only, without the need of an explicit mixing model. We refer to the proposed approach as CycleGAN unmixing net, or CGU net. Experimental results indicate that the proposed CGU net exhibits stable and competitive performance on different datasets as compared to traditional HNU methods that are model-based. Maofeng Tang, Ying Qu 0001, Hairong Qi 0001 |
IGARSS | 2 |
| 2020 | Batch Normalization Masked Sparse Autoencoder for Robotic Grasping DetectionabstractTo improve the accuracy of the grasping detection, this paper proposes a novel detector with batch normalization masked evaluation model. It is designed with a two-layer sparse autoencoder, and a Batch Normalization based mask is incorporated into the second layer of the model to effectively reduce the features with weak correlation. The extracted features from such model are more distinctive, which guarantees the higher accuracy of the grasping detection. Extensive experiments show that the proposed evaluation model outperforms the state-of- the-art, and the recognition accuracy can reach 95.51% for robotic grasping detection. Zhenzhou Shao, Ying Qu 0001, Guangli Ren, Zhi-Ping Shi 0002, Jindong Tan |
IROS | 2 |
| 2019 | Hybrid Spectral Unmixing in Land-Cover ClassificationabstractIdentifying land-cover and specifically the type of the material that constitutes building roofs in urban areas provides important reference information for later procedures including semantic labeling, bridge masking, and 3D reconstruction. In this paper, we present a hybrid unmixing-based classification framework that integrates both class-wise unsupervised unmixing and supervised unmixing that effectively convert the classification problem from the original spectral space to the abundance space, such that the intrinsic characteristics of each material can be better represented. Experimental results demonstrate competitive performance in terms of classification accuracy. In addition, we show that the proposed approach has the capability of handling new region of interest with similar scene content but different illumination geometry and atmospheric composition, which is crucial in classification of satellite images with a limited amount of training data. Razieh Kaviani Baghbaderani, Fanqi Wang, Craig Stutts, Ying Qu 0001, Hairong Qi 0001 |
IGARSS | 4 |
| 2019 | Inverse Dynamics Modeling of Robotic Manipulator with Hierarchical Recurrent NetworkabstractInverse dynamics modeling is a critical problem for the computed-torque control of robotic manipulator. This paper presents a novel recurrent network based on the modified Simple Recurrent Unit (SRU) with hierarchical memory (SRU-HM), which is achieved by the nested SRU structure. In this way, it enables the capability to retain the long-term information in the distant past, compared with the conventional stacked structure. The hidden state of SRU is able to provide more complete information relevant to current prediction. Experimental results demonstrate that the proposed method can improve the accuracy of dynamics model greatly, and outperforms the state-of-the-art methods. Zhenzhou Shao, Ying Qu 0001, Jindong Tan |
IROS | 3 |
| 2019 | uDAS: An Untied Denoising Autoencoder With Sparsity for Spectral UnmixingabstractLinear spectral unmixing is the practice of decomposing the mixed pixel into a linear combination of the constituent endmembers and the estimated abundances. This paper focuses on unsupervised spectral unmixing where the endmembers are unknown a priori. Conventional approaches use either geometrical- or statistical-based approaches. In this paper, we address the challenges of spectral unmixing with unsupervised deep learning models, in specific, the autoencoder models, where the decoder serves as the endmembers and the hidden layer output serves as the abundances. In several recent attempts, part-based autoencoders have been designed to solve the unsupervised spectral unmixing problem. However, the performance has not been satisfactory. In this paper, we first discuss some important findings we make on issues with part-based autoencoders. By proof of counterexample, we show that all existing part-based autoencoder networks with nonnegative and tied encoder and decoder are inherently defective by making these inappropriate assumptions on the network structure. As a result, they are not suitable for solving the spectral unmixing problem. We propose a so-called untied denoising autoencoder with sparsity, in which the encoder and decoder of the network are independent, and only the decoder of the network is enforced to be nonnegative. Furthermore, we make two critical additions to the network design. First, since denoising is an essential step for spectral unmixing, we propose to incorporate the denoising capacity into the network optimization in the format of a denoising constraint rather than cascading another denoising preprocessor in order to avoid the introduction of additional reconstruction error. Second, to be more robust to the inaccurate estimation of a number of endmembers, we adopt an $l_{21}$ -norm on the encoder of the network to reduce the redundant endmembers while decreasing the reconstruction error simultaneously. The experimental results demonstrate that the proposed approach outperforms several state-of-the-art methods, especially for highly noisy data. Ying Qu 0001, Hairong Qi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Unsupervised Sparse Dirichlet-Net for Hyperspectral Image Super-ResolutionabstractIn many computer vision applications, obtaining images of high resolution in both the spatial and spectral domains are equally important. However, due to hardware limitations, one can only expect to acquire images of high resolution in either the spatial or spectral domains. This paper focuses on hyperspectral image super-resolution (HSI-SR), where a hyperspectral image (HSI) with low spatial resolution (LR) but high spectral resolution is fused with a multispectral image (MSI) with high spatial resolution (HR) but low spectral resolution to obtain HR HSI. Existing deep learning-based solutions are all supervised that would need a large training set and the availability of HR HSI, which is unrealistic. Here, we make the first attempt to solving the HSI-SR problem using an unsupervised encoder-decoder architecture that carries the following uniquenesses. First, it is composed of two encoder-decoder networks, coupled through a shared decoder, in order to preserve the rich spectral information from the HSI network. Second, the network encourages the representations from both modalities to follow a sparse Dirichlet distribution which naturally incorporates the two physical constraints of HSI and MSI. Third, the angular difference between representations are minimized in order to reduce the spectral distortion. We refer to the proposed architecture as unsupervised Sparse Dirichlet-Net, or uSDN. Extensive experimental results demonstrate the superior performance of uSDN as compared to the state-of-the-art. Ying Qu 0001, Hairong Qi 0001, Chiman Kwan |
CVPR | 1 |
| 2018 | Unsupervised Trajectory Segmentation and Promoting of Multi-Modal Surgical DemonstrationsabstractTo improve the efficiency of surgical trajectory segmentation for robot learning in robot-assisted minimally invasive surgery, this paper presents a fast unsupervised method using video and kinematic data, followed by a promoting procedure to address the over-segmentation issue. Unsupervised deep learning network, stacking convolutional auto-encoder, is employed to extract more discriminative features from videos in an effective way. To further improve the accuracy of segmentation, on one hand, wavelet transform is used to filter out the noises existed in the features from video and kinematic data. On the other hand, the segmentation result is promoted by identifying the adjacent segments with no state transition based on the predefined similarity measurements. Extensive experiments on a public dataset JIGSAWS show that our method achieves much higher accuracy of segmentation than state-of-the-art methods in the shorter time. Zhenzhou Shao, Hongfa Zhao, Jiexin Xie, Ying Qu 0001, Jindong Tan |
IROS | 4 |
| 2018 | Hyperspectral Anomaly Detection Through Spectral Unmixing and Dictionary-Based Low-Rank DecompositionabstractAnomaly detection has been known to be a challenging problem due to the uncertainty of anomaly and the interference of noise. In this paper, we focus on anomaly detection in hyperspectral images (HSIs) and propose a novel detection algorithm based on spectral unmixing and dictionary-based low-rank decomposition. The innovation is threefold. First, due to the highly mixed nature of pixels in HSI data, instead of using the raw pixel directly for anomaly detection, the proposed algorithm applies spectral unmixing to obtain the abundance vectors and uses these vectors for anomaly detection. We show that the abundance vectors possess more distinctive features to identify anomaly from background. Second, to better represent the highly correlated background and the sparse anomaly, we construct a dictionary based on the mean shift clustering of the abundance vectors to improve both the discriminative and representative powers of the algorithm. Finally, a low-rank matrix decomposition method based on the constructed dictionary is proposed to encourage the coefficients of the dictionary, instead of the background itself, to be low rank, and the residual matrix to be sparse. Anomalies can then be extracted by summing up the columns of the residual matrix. The proposed algorithm is evaluated on both synthetic and real data sets. Experimental results show that the proposed approach constantly achieves high detection rate, while maintaining low false alarm rate regardless of the type of images tested. Ying Qu 0001, Wei Wang 0063, Bulent Ayhan, Chiman Kwan, Steven Vance, Hairong Qi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Spectral unmixing through part-based non-negative constraint denoising autoencoderabstractSpectral unmixing is to decompose the hyperspectral data into endmembers and abundances. It has been known to be a challenging and ill-posed task due to the corruption of noise as well as complex environmental conditions. In this paper, we propose a part-based denoising autoencoder with unique structure that solves the unmixing challenges. The effective l21norm and denoising constraints are applied on the network to better handle noise, while at the same time reducing the reconstruction error and redundant endmembers simultaneously. A back propagation optimization method powered with the Armijo rule is proposed to project the weights to the non-negativity space that guarantees the sum-to-one constraint. The experimental results demonstrate the proposed approach is able to outperform several state-of-the-art methods for highly noisy data. Ying Qu 0001, Hairong Qi 0001 |
IGARSS | 1 |
| 2017 | DOES multispectral / hyperspectral pansharpening improve the performance of anomaly detection?abstractPansharpening refers to the fusion of a high spatial resolution panchromatic image with high spectral resolution multispectral or hyperspectral images (MSI or HSI) to yield high resolution data in both spectral and spatial domains. It has been widely adopted as a primary preprocessing step for numerous applications. In this paper, we perform a literature survey of various pansharpening algorithms including the most advanced deep learning approaches for both multispectral and hyperspectral images. We further evaluate the effect of the resolution difference on anomaly detection. Synthetic multispectral and hyperspectral images are generated to evaluate the performance of anomaly detection on high resolution images. Eight state-of-the-art MSI and HSI pansharpening methods are compared in this paper. Experimental results show that, performing anomaly detection on high resolution images improves the detection rate, and at the mean time suppresses the false alarm rate. Ying Qu 0001, Hairong Qi 0001, Bulent Ayhan, Chiman Kwan, Richard Kidd |
IGARSS | 1 |
| 2017 | A fast search algorithm based on image pyramid for robotic graspingabstractTo improve the search efficiency of robotic grasping detection, this paper presents a novel search algorithm based on the image pyramid. It significantly reduces the search space for grasping position detection using the coarse-to-fine strategy. The proposed method searches the positions from the top layer of the pyramid, and initializes the search area at the next layer. The sparse automatic encoder is employed to construct the model which is used to evaluate the grasp quality. The experimental results demonstrate that the proposed search algorithm can improve efficiency of the robotic grasping detection with the comparative performance on the grasp quality. Guangli Ren, Zhenzhou Shao, Ying Qu 0001, Jindong Tan, Hongxing Wei, Guofeng Tong |
IROS | 4 |
| 2016 | Anomaly detection in hyperspectral images through spectral unmixing and low rank decompositionabstractAnomaly detection has been known to be a challenging, ill-posed problem due to the uncertainty of anomaly and the interference of noise. In this paper, we propose a novel low rank anomaly detection algorithm in hyperspectral images (HSI), where three components are involved. First, due to the highly mixed nature of pixels in HSI, instead of using the raw pixel directly for anomaly detection, the proposed algorithm applies spectral unmixing algorithms to obtain the abundance vectors and uses these vectors for anomaly detection. Second, for better classification, a dictionary is built based on the mean-shift clustering of the abundance vectors to better represent the highly-correlated background and the sparse anomaly. Finally, a low-rank matrix decomposition is proposed to encourage the sparse coefficients of the dictionary to be low-rank, and the residual matrix to be sparse. Anomalies can then be extracted by summing up the columns of the residual matrix. The proposed algorithm is evaluated on both synthetic and real datasets. Experimental results show that the proposed approach constantly achieves high detection rate while maintaining low false alarm rate regardless of the type of images tested. Ying Qu 0001, Wei Wang 0063, Hairong Qi 0001, Bulent Ayhan, Chiman Kwan, Steven Vance |
IGARSS | 1 |
| 2016 | Robust penalty-weighted deblurring via kernel adaption using single image
Ying Qu 0001, Andreas F. Koschan, Mongi A. Abidi |
J. Vis. Commun. Image Represent. | 1 |