Mingyi He

dblp:59/1181 · DBLP profile ↗
← Back
65ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0003-2051-6955ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 7 since 2021Systems, architecture and hardware · 1Computer networks · 1
YearPublicationVenuePosition
2026 A Hierarchical Optimization Method for Electric Vertical Takeoff and Landing Aircraft Network Design
abstract
Electric vertical takeoff and landing aircraft (eVTOLs) are expected to serve urban air mobility in a station-to-station configuration, which makes the optimal network design of eVTOL stations a critical question to explore. Existing approaches often face limitations, such as the inability to interact station locations with demand or difficulty in finding the optimal solution for large study regions. This paper first proposes a mathematical model to generate optimal eVTOL station locations while considering associated potential eVTOL demand, and then proposes a heuristic algorithm, Hierarchical Optimization MEthod (HOME), to efficiently solve the model. With a case study of Southern California, HOME was compared to 1) directly solving the original integer linear programming-based network design problem, and 2) employing the widely used genetic algorithm. Results suggest that HOME can find optimal solutions with limited computational resources. The proposed framework powered by HOME provides a computationally efficient way to support urban air mobility planning.
Mingyi He, Bingrong Sun, Venu Garikapati, Zhaocai Liu, Joshua Hoshiko, Mingdong Lyu, Yanbo Ge
IEEE Trans. Intell. Transp. Syst.1
2026 Boosting Few-Shot Hyperspectral Image Classification Through Dynamic Fusion and Hierarchical Enhancement
abstract
Few-shot learning has garnered increasing attention in hyperspectral image classification (HSIC) due to its potential to reduce dependency on labor-intensive and costly labeled data. However, most existing methods are constrained to feature extraction using a single image patch of fixed size, and typically neglect the pivotal role of the central pixel in feature fusion, leading to inefficient information utilization. In addition, the correlations among sample features have not been fully explored, thereby weakening feature expressiveness and hindering cross-domain knowledge transfer. To address these issues, we propose a novel few-shot HSIC framework incorporating dynamic fusion and hierarchical enhancement. Specifically, we first introduce a robust feature extraction module, which effectively combines the content concentration of small patches with the noise robustness of large patches, and further captures local spatial correlations through a central-pixel-guided dynamic pooling strategy. Such patch-to-pixel dynamic fusion enables a more comprehensive and robust extraction of ground object information. Then, we develop a support-query hierarchical enhancement module that integrates intraclass self-attention and interclass cross-attention mechanisms. This process not only enhances support-level and query-level feature representation but also facilitates the learning of more informative prior knowledge from the abundantly labeled source domain. Moreover, to further increase feature discriminability, we design an intraclass consistency loss and an interclass orthogonality loss, which collaboratively encourage intraclass samples to be closer together and interclass samples to be more separable in the metric space. Experimental results on four benchmark datasets demonstrate that our method substantially improves classification accuracy and consistently outperforms competing approaches. Code is available at https://github.com/guoying918/DFHE2025.
Ying Guo 0014, Bin Fan 0002, Yuchao Dai, Yan Feng 0005, Mingyi He
IEEE Trans. Neural Networks Learn. Syst.5
2024 Distribution-Aware and Class-Adaptive Aggregation for Few-Shot Hyperspectral Image Classification
abstract
Recently, few-shot learning based on meta-learning has shown great potential in hyperspectral image classification (HSIC) due to its excellent adaptability to limited training samples. Despite achieving promising results, the existing methods ignore the interaction between the source domain (with abundant-labeled base-class samples) and the target domain (with few-labeled novel-class samples), as well as between the support set and the query set. This issue makes the resulting model usually biased toward the source domain and not robust to the sample variance of novel classes, posing a bottleneck to the improvement of HSIC performance. To overcome these limitations, we propose a flexible and effective distribution-aware and class-adaptive aggregation (DA-CAA) method for few-shot HSIC by transferring the class-level distribution information learned from the base classes to the novel classes. Specifically, we first employ a variational autoencoder (VAE), which is pretrained on abundant-labeled base-class samples, to encode the support set samples as class distributions. Subsequently, we sample class-level features from the learned distribution and adaptively aggregate them with sample-specific query features. This operation not only enhances cross-domain information interaction in a distribution-learning manner, but also ensures that the aggregated features across classes inherit both class-level and sample-specific information. Our proposed class-adaptive aggregation (CAA) encourages complementary fusion of features from all classes, which is beneficial for reducing class confusion. Experiments on four benchmark datasets demonstrate the effectiveness and flexibility of our approach.
Ying Guo 0014, Bin Fan 0002, Yan Feng 0005, Xiuping Jia, Mingyi He
IEEE Trans. Geosci. Remote. Sens.5
2023 Masked Representation Learning for Domain Generalized Stereo Matching
abstract
Recently, many deep stereo matching methods have begun to focus on cross-domain performance, achieving impressive achievements. However, these methods did not deal with the significant volatility of generalization performance among different training epochs. Inspired by masked representation learning and multi-task learning, this paper designs a simple and effective masked representation for domain generalized stereo matching. First, we feed the masked left and complete right images as input into the models. Then, we add a lightweight and simple decoder following the feature extraction module to recover the original left image. Finally, we train the models with two tasks (stereo matching and image reconstruction) as a pseudo-multi-task learning framework, promoting models to learn structure information and to improve generalization performance. We implement our method on two well-known architectures (CFNet and LacGwcNet) to demonstrate its effectiveness. Experimental results on multi-datasets show that: (1) our method can be easily plugged into the current various stereo matching models to improve generalization performance; (2) our method can reduce the significant volatility of generalization performance among different training epochs; (3) we find that the current methods prefer to choose the best results among different training epochs as generalization performance, but it is impossible to select the best performance by ground truth in practice.
Zhibo Rao, Mingyi He, Yuchao Dai, Zhelun Shen, Xing Li 0040
CVPR3
2023 Grid-Transformer for Few-Shot Hyperspectral Image Classification
abstract
The application of few-shot learning to hyperspectral image (HSI) classification tasks has gradually become a research hotspot due to the difficulties in acquiring and labeling HSI data. Existing methods tend to cascade a large number of convolutional neural networks. However, such operations can only focus on local information and cannot accurately capture the strong correlation between spectra. To address this problem, we propose Grid-transformer, an efficient spatial-spectral feature extraction model. Specifically, we first introduce a more powerful transformer to compute non-local self-similarity along the spectral dimension, which is beneficial to mine more discriminative spectral features. Then, they are embedded into a grid-like network architecture to fully aggregate multi-scale contextual information, resulting in a more complete spatial-spectral feature representation. Experiments on two benchmark datasets demonstrate that our approach achieves state-of-the-art classification performance.
Ying Guo 0014, Mingyi He, Bin Fan 0002
ICIP2
2023 Zero-Shot SAR Target Recognition Based on Classification Assistance
abstract
As one of main active learning methods, zero-shot target recognition with synthetic aperture radar (SAR) data has received considerable attention in recent years. Its goal is to distinguish new targets from the known classes without requiring additional training data. Existing zero-shot learning (ZSL) methods perform well on optical targets, but they fail to recognize zero-shot SAR targets. Due to the strong similarity of SAR targets between different classes, the ZSL task usually suffers from the distribution concentration problem. To tackle this problem, a Dual Branch Auto-Encoder (DBAE) network is proposed in this letter. DBAE effectively alleviates the distribution concentration problem by adding a classification assistance net. Its dual branch structure is specially designed for further improving the intra-class similarity and inter-class dissimilarity of SAR targets in the embedding space. By training with the defined hybrid loss function, DBAE automatically builds a stable embedding space. Extensive experiments on the public data set of Moving and Stationary Target Acquisition and Recognition (MSTAR) show that DBAE is of rational design and provides better or comparable ZSL recognition results.
Qian-Ru Wei, Mingyi He, Hongmei He
IEEE Geosci. Remote. Sens. Lett.3
2022 End-to-End Learning the Partial Permutation Matrix for Robust 3D Point Cloud Registration
abstract
Even though considerable progress has been made in deep learning-based 3D point cloud processing, how to obtain accurate correspondences for robust registration remains a major challenge because existing hard assignment methods cannot deal with outliers naturally. Alternatively, the soft matching-based methods have been proposed to learn the matching probability rather than hard assignment. However, in this paper, we prove that these methods have an inherent ambiguity causing many deceptive correspondences. To address the above challenges, we propose to learn a partial permutation matching matrix, which does not assign corresponding points to outliers, and implements hard assignment to prevent ambiguity. However, this proposal poses two new problems, i.e. existing hard assignment algorithms can only solve a full rank permutation matrix rather than a partial permutation matrix, and this desired matrix is defined in the discrete space, which is non-differentiable. In response, we design a dedicated soft-to-hard (S2H) matching procedure within the registration pipeline consisting of two steps: solving the soft matching matrix (S-step) and projecting this soft matrix to the partial permutation matrix (H-step). Specifically, we augment the profit matrix before the hard assignment to solve an augmented permutation matrix, which is cropped to achieve the final partial permutation matrix. Moreover, to guarantee end-to-end learning, we supervise the learned partial permutation matrix but propagate the gradient to the soft matrix instead. Our S2H matching procedure can be easily integrated with existing registration frameworks, which has been verified in representative frameworks including DCP, RPMNet, and DGR. Extensive experiments have validated our method, which creates a new state-of-the-art performance.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He
AAAI6
2022 Context-Aware Video Reconstruction for Rolling Shutter Cameras
abstract
With the ubiquity of rolling shutter (RS) cameras, it is becoming increasingly attractive to recover the latent global shutter (GS) video from two consecutive RS frames, which also places a higher demand on realism. Existing solutions, using deep neural networks or optimization, achieve promising performance. However, these methods generate intermediate GS frames through image warping based on the RS model, which inevitably result in black holes and noticeable motion artifacts. In this paper, we alleviate these issues by proposing a context-aware GS video reconstruction architecture. It facilitates the advantages such as occlusion reasoning, motion compensation, and temporal abstraction. Specifically, we first estimate the bilateral motion field so that the pixels of the two RS frames are warped to a common GS frame accordingly. Then, a refinement scheme is proposed to guide the GS frame synthesis along with bilateral occlusion masks to produce high-fidelity GS video frames at arbitrary times. Furthermore, we derive an approximated bilateral motion field model, which can serve as an alternative to provide a simple but effective GS frame initialization for related tasks. Experiments on synthetic and real data show that our approach achieves superior performance over state-of-the-art methods in terms of objective metrics and subjective visual quality. Code is available at https://github.com/GitCVfb/CVR.
Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002, Qi Liu 0054, Mingyi He
CVPR5
2022 Semantic Segmentation of High-Resolution Remote Sensing Images Using an Improved Transformer
abstract
Semantic segmentation has been widely researched for high level analysis of High Spatial Resolution (HSR) remote sensing images, where Convolutional Neural Network (CNN) is the mainstream method. However, the transformer with attention mechanism has its unique capacity of extracting global information which is generally ignored by CNN models. In this paper, a Swin Transformer with UPer head (STUP) is proposed to tackle with semantic segmentation problem on a challenging remote sensing land-cover dataset called LoveDA, which owns complex background samples and inconsistent classes distributions. The proposed STUP combines the Swin Transformer with Uper Head in the form of an encoder-decoder structure, to extract features of HSR images for segmentation. Furthermore, Focal Loss is adopted to handle the unbalanced distribution problem in the training step. Experimental results demonstrate that the proposed STUP clearly outperforms several state-of-the-art models.
Shaohui Mei, Ye Wang 0020, Mingyi He, Qian Du 0001
IGARSS5
2022 Gaussian Information Entropy based band Reduction for Unsupervised Hyperspectral Video Tracking
abstract
Hyperspectral videos, which provide extra spectral characteristics besides spatial and temporal information, can improve the performance of object tracking using spectral signatures. However, there is a lack of labeled hyperspectral videos to support deep learning based model design. On the contrary, object tracking in the color space has been well developed in the past decade with many benchmark tracking models, e.g., SiamBAN. Therefore, how to transfer models designed in the color space to the hyperspectral space is of great importance. In this paper, hyperspectral videos are reduced into 3 bands using a band reduction algorithm, by which the existing well-trained trackers can be directly used. Specifically, Gaussian Information Entropy (GIE) is used to transform a hyperspectral video into a 3-band pseudo-color video, by which hyperspectral object tracking is conducted in an unsupervised mode. Experimental results demonstrate that object trackers designed in the color space can be transferred to hyperspectral videos using band reduction algorithms and the GIE based reduction is more effective than several well-known band reduction algorithms when using SiamBAN.
Yuru Su, Shaohui Mei, Ge Zhang 0006, Ye Wang 0020, Mingyi He, Qian Du 0001
IGARSS5
2022 A Representation Separation Perspective to Correspondence-Free Unsupervised 3-D Point Cloud Registration
abstract
3-D point cloud registration in remote sensing field has been greatly advanced by deep learning-based methods, where the rigid transformation is either directly regressed from the two point clouds (correspondences-free approaches) or computed from the learned correspondences (correspondences-based approaches). Existing correspondence-free methods generally learn the holistic representation of the entire point cloud, which is fragile for partial and noisy point clouds. In this letter, we propose a correspondence-free unsupervised point cloud registration (UPCR) method from the representation separation perspective. First, we model the input point cloud as a combination of pose-invariant representation and pose-related representation. Second, the pose-related representation is used to learn the relative pose w.r.t. a “latent canonical shape” for thesourceandtargetpoint clouds, respectively. Third, the rigid transformation is obtained from the above two learned relative poses. Our method not only filters out the disturbance in pose-invariant representation but also is robust to partial-to-partial point clouds or noise. Experiments on benchmark datasets demonstrate that our unsupervised method achieves comparable if not better performance than state-of-the-art supervised registration methods.The source code will be made public.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He
IEEE Geosci. Remote. Sens. Lett.6
2022 Sliding space-disparity transformer for stereo matching
Zhibo Rao, Mingyi He, Yuchao Dai, Zhelun Shen
Neural Comput. Appl.2
2022 Self-supervised rigid transformation equivariance for accurate 3D point cloud registration
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He
Pattern Recognit.6
2022 Fast and Robust Differential Relative Pose Estimation With Radial Distortion
abstract
In this letter, we address the differential two-view geometry problem of estimating the relative pose between two consecutive frames in the presence of radial distortion. This problem is of both theoretical and practical interests and has not been solved. We derive its parameterization and present an effective and robust generalized eigenvalue solver based on the hidden variable technique. Furthermore, we propose a nonlinear refinement scheme within the maximum likelihood criterion to produce more accurate estimates of the relative pose and radial distortion. Compared with the standard differential solutions without modeling the radial distortion, our approach can recover more geometrically correct point correspondences for a pair of radially distorted images. Moreover, our differential solution runs an order of magnitude faster than the discrete solution in terms of recovering the full camera motion. Experiment results on both synthetic and real data demonstrate the effectiveness of our model and method in dealing with the radial distortion.
Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002, Mingyi He
IEEE Signal Process. Lett.4
2022 Learning a Task-Specific Descriptor for Robust Matching of 3D Point Clouds
abstract
Existing learning-based point feature descriptors are usually task-agnostic, which pursue describing the individual 3D point clouds as accurate as possible. However, the matching task aims at describing the corresponding points consistently across different 3D point clouds. Therefore these too accurate features may play a counterproductive role due to the inconsistent point feature representations of correspondences caused by the unpredictable noise, partiality, deformation, etc., in the local geometry. In this paper, we propose to learn a robust task-specific feature descriptor to consistently describe the correct point correspondence under interference. Born with anEncoder and aDynamicFusion module, our method EDFNet develops from two aspects. First, we augment the matchability of correspondences by utilizing their repetitive local structure. To this end, a special encoder is designed to exploit two input point clouds jointly for each point descriptor. It not only captures the local geometry of each point in the current point cloud by convolution, but also exploits the repetitive structure from paired point cloud by Transformer. Second, we propose a dynamical fusion module to jointly use different scale features. There is an inevitable struggle between robustness and discriminativeness of the single scale feature. Specifically, the small scale feature is robust since little interference exists in this small receptive field. But it is not sufficiently discriminative as there are many repetitive local structures within a point cloud. Thus the resultant descriptors will lead to many incorrect matches. In contrast, the large scale feature is more discriminative by integrating more neighborhood information. But it is easier to be disturbed since there is much more interference in the large receptive field. Compared with the conventional fusion strategy that handles multiple scale features equally, we analyze the consistency of them to judge the clean ones and perform larger aggregation weights on them during fusion. Then, a robust and discriminative feature descriptor is achieved by focusing on multiple clean scale features. Extensive evaluations validate that EDFNet learns a task-specific descriptor, which achieves state-of-the-art or comparable performance for robust matching of 3D point clouds.
Zhiyuan Zhang 0002, Yuchao Dai, Bin Fan 0002, Jiadai Sun, Mingyi He
IEEE Trans. Circuits Syst. Video Technol.5
2022 VRNet: Learning the Rectified Virtual Corresponding Points for 3D Point Cloud Registration
abstract
3D point cloud registration is fragile to outliers, which are labeled as the points without corresponding points. To handle this problem, a widely adopted strategy is to estimate the relative pose based only on some accurate correspondences, which is achieved by building correspondences on the identified inliers or by selecting reliable ones. However, these approaches are usually complicated and time-consuming. By contrast, the virtual point-based methods learn the virtual corresponding points (VCPs) for allsourcepoints uniformly without distinguishing the outliers and the inliers. Although this strategy is time-efficient, the learned VCPs usually exhibit serious collapse degeneration due to insufficient supervision and the inherent distribution limitation. In this paper, we propose to exploit the best of both worlds and present a novel robust 3D point cloud registration framework. We follow the idea of the virtual point-based methods but learn a new type of virtual points called rectified virtual corresponding points (RCPs), which are defined as the point set with the same shape as thesourceand with the same pose as thetarget. Hence, a pair of consistent point clouds,i.e.sourceand RCPs, is formed by rectifying VCPs to RCPs (VRNet), through which reliable correspondences betweensourceand RCPs can be accurately obtained. Since the relative pose betweensourceand RCPs is the same as the relative pose betweensourceandtarget, the input point clouds can be registered naturally. Specifically, we first construct the initial VCPs by using an estimated soft matching matrix to perform a weighted average on thetargetpoints. Then, we design a correction-walk module to learn an offset to rectify VCPs to RCPs, which effectively breaks the distribution limitation of VCPs. Finally, we develop a hybrid loss function to enforce the shape and geometry structure consistency of the learned RCPs and thesourceto provide sufficient supervision. Extensive experiments on several benchmark datasets demonstrate that our method achieves advanced registration performance and time-efficiency simultaneously.The code will be made public.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Mingyi He
IEEE Trans. Circuits Syst. Video Technol.5
2022 Patch attention network with generative adversarial model for semi-supervised binocular disparity prediction
Zhibo Rao, Mingyi He, Yuchao Dai, Zhelun Shen
Vis. Comput.2
2021 SUNet: Symmetric Undistortion Network for Rolling Shutter Correction
abstract
The vast majority of modern consumer-grade cameras employ a rolling shutter mechanism, leading to image distortions if the camera moves during image acquisition. In this paper, we present a novel deep network to solve the generic rolling shutter correction problem with two consecutive frames. Our pipeline is symmetrically designed to predict the global shutter image corresponding to the intermediate time of these two frames, which is difficult for existing methods because it corresponds to a camera pose that differs most from the two frames. First, two time-symmetric dense undistortion flows are estimated by using well-established principles: pyramidal construction, warping, and cost volume processing. Then, both rolling shutter images are warped into a common global shutter one in the feature space, respectively. Finally, a symmetric consistency constraint is constructed in the image decoder to effectively aggregate the contextual cues of two rolling shutter images, thereby recovering the high-quality global shutter image. Extensive experiments with both synthetic and real data from public benchmarks demonstrate the superiority of our proposed approach over the state-of-the-art methods.
Bin Fan 0002, Yuchao Dai, Mingyi He
ICCV3
2021 Him-Net: A New Neural Network Approach for SAR and Optical Image Template Matching
abstract
SAR and optical images provide highly complementary information about observed scenes. The integrated use of these two data is desired in many data fusion tasks. However, traditional similarity methods cannot correctly match SAR and optical images due to the significant non-linear radio-metric difference between them. This paper proposed a template matching neural network based on stereo matching for SAR and optical image matching. Unlike the classical template matching methods doing feature extraction and similarity calculating separately, our network is a complete end-to-end approach, which allows optimizing the matching between the SAR and optical images through training. Moreover, a heatmap loss function is designed for image template matching, and better result is obtained. Our experiments confirmed our proposed network advantage over the state-of-the-art similarity approaches (such as NCC, CARMI, DeepMatch, and QATM) and superior matching performance.
Mingyi He, Zhibo Rao, Wenyao Li 0002
ICIP2
2021 RS-DPSNet: Deep Plane Sweep Network for Rolling Shutter Stereo Images
abstract
Since the rolling shutter (RS) camera successively exposes each scanline, accurately reconstructing scene depth from an RS stereo image pair remains a great challenge. Directly applying the deep-learning-based depth estimation methods tailored for the global shutter (GS) stereo images leads to undesirable RS depth results due to inherent flaws in the network structure. In this letter, we fill this gap by developing an end-to-end RS-stereo-aware plane sweep network to improve the accuracy of the classic GS-based algorithm (i.e.DPSNet) in estimating the RS depth map. Specifically, we derive the RS-stereo-aware plane sweep model and further produce a more accurate and efficient cost volume through the effective incorporation of this model within DPSNet. Furthermore, to enable learning-based approaches to address the depth estimation problem in the context of RS stereo images, we contribute the first RS stereo dataset, CARLA-RSS. Experimental results demonstrate that our proposed pipeline achieves state-of-the-art performance.
Bin Fan 0002, Yuchao Dai, Mingyi He
IEEE Signal Process. Lett.4
2021 Bidirectional Guided Attention Network for 3-D Semantic Detection of Remote Sensing Images
abstract
Semantic segmentation and disparity estimation are in the research frontier of the computer vision and remote sensing (RS) fields. However, existing methods mostly deal with these two problems separately or use a combination of multiple models to solve these two tasks. Due to a lack of sufficient information sharing and fusion, they still have difficulties in coping with seasonal appearance differences in 3-D RS problems. In this article, we propose a novel multitask learning architecture that considers the bottom–up and up–bottom visual attention mechanism for 3-D semantic detection, named bidirectional guided attention network (BGA-Net). BGA-Net consists of five modules: unified backbone module (UBM), bidirectional guided attention module (BGAM), semantic segmentation module (SSM), feature matching module (FMM), and bidirectional fusion module (BFM). First, in UBM, we use a shared backbone to extract unified features and share them with three branches/modules (BGAM, SSM, and FMM). Then, SSM and FMM branches are applied to estimate segmentation and disparity maps, whereas the third branch/module (BGAM) shares the global features to guide the task-specific learning via attention mechanism. Finally, we fuse the results of the two tasks by BFM to improve the final performance. Extensive experiments demonstrate that: 1) our BGA-Net can handle the two tasks simultaneously and can be trained in an end-to-end way; 2) these modules fully take advantage of the two tasks’ information to share features and enhance the scene understanding ability, effectively against seasons change of RS images; and 3) BGA-Net has notable superiority and greater flexibility and also sets a new state of the art on the urban semantic 3-D (US3D) benchmark. Moreover, BGA-Net also provides insights into the intelligent interpretation of RS data images.
Zhibo Rao, Mingyi He, Zhidong Zhu, Yuchao Dai
IEEE Trans. Geosci. Remote. Sens.2
2020 Monocular human pose estimation: A survey of deep learning-based methods
Yingli Tian, Mingyi He
Comput. Vis. Image Underst.3
2019 Input-Perturbation-Sensitivity for Performance Analysis of CNNS on Image Recognition
abstract
Performance assessment is critical to learning systems, but it is tough to explain the relationship between data, model, and performance. In this paper, the Input-Perturbation-Sensitivity (IPS) is proposed to investigate this problem in a class of Convolutional Neural Networks (CNNs) for image recognition and try to explain their relationships. First, IPS is defined and the CNNs parameters are divided into groups according to the layers of the model. Second, the output perturbations of the CNNs caused by input perturbations are analyzed with a group of local sensitivities (LS). Third, global sensitivity (GS) is obtained over all local IPS. Finally, experiments are carried out on a few CNNs with different hyper-parameters on the image recognition datasets. The analytic and experimental results show that the proposed method correlates well with data, model, and performance. Moreover, the IPS can provide a reasonable explanation of the different networks for image classification, showing the potential to evaluate other learning systems.
Zhibo Rao, Mingyi He, Zhidong Zhu
ICIP2
2019 Convolutional Neural Network with PCA and Batch Normalization for Hyperspectral Image Classification
abstract
A new deep learning based spectral-spatial approach for hyperspectral image classification is developed, which uses spectral reduction as preprocessing and batch normalization in every layer of the deep network. The spectral data is reduced by Principal Component Analysis and the spatial dimension is sliced into patches of 9x9. These patches hierarchically deliver discriminative features when feed to the proposed network. The training process is regularized and the overfitting (previously often encountered problem) is avoided by using combination of batch normalization and dropout. Moreover, oversampling and augmentation in training data is used to expand the training data and to create some variation in available training data. Finally the experimental results demonstrated the performance of our method in comparison to other methods especially for hyperspectral classification tasks.
Aamir Naveed Abbasi, Mingyi He
IGARSS2
2018 3D skeleton based action recognition by video-domain translation-scale invariant mapping and multi-scale dilated CNN
Bo Li 0090, Mingyi He, Yuchao Dai, Xuelian Cheng
Multim. Tools Appl.2
2018 Monocular depth estimation with hierarchical fusion of dilated CNNs and soft-weighted-sum inference
Bo Li 0090, Yuchao Dai, Mingyi He
Pattern Recognit.3
2017 Dense non-rigid structure-from-motion made easy - A spatial-temporal smoothness based solution
abstract
This paper proposes a simple spatial-temporal smoothness based method for solving dense non-rigid structure-frommotion (NRSfM). First, we revisit the temporal smoothness and demonstrate that it can be extended to dense case directly. Second, we propose to exploit the spatial smoothness by resorting to the Laplacian of the 3D non-rigid shape. Third, to handle real world noise and outliers in measurements, we robustify the data term by using the L1norm. In this way, our method could robustly exploit both spatial and temporal smoothness effectively and make dense non-rigid reconstruction easy. Our method is very easy to implement, which involves solving a series of least squares problems. Experimental results on both synthetic and real image dense NRSfM tasks show that the proposed method outperforms state-of-the-art dense non-rigid reconstruction methods.
Yuchao Dai, Huizhong Deng, Mingyi He
ICIP3
2017 Multi-scale 3D deep convolutional neural network for hyperspectral image classification
abstract
Research in deep neural network (DNN) and deep learning has great progress for 1D (speech), 2D (image) and 3D (3D-object) recognition/classification problems. As HSI that with 2D spatial and 1D spectral information is quite different from 3D object image, the existing DNN cannot be directly extended to hyperspectral image (HSI) classification. A Multiscale 3D deep convolutional neural network (M3D-DCNN) is proposed for HSI classification, which could jointly learn both 2D Multi-scale spatial feature and 1D spectral feature from HSI data in an end-to-end approach, promising to achieve better results with large-scale dataset. Although without any hand-craft features or pre/post-processing like PCA, sparse coding etc, we achieve the state-of-the-art results on the standard datasets, which shows the technical validity and advancement of our method.
Mingyi He, Bo Li 0090, Huahui Chen 0002
ICIP1
2017 Integrated deep and shallow networks for salient object detection
abstract
Deep convolutional neural network (CNN) based salient object detection methods have achieved state-of-the-art performance and outperform those unsupervised methods with a wide margin. In this paper, we propose to integrate deep and unsupervised saliency for salient object detection under a unified framework. Specifically, our method takes results of unsupervised saliency (Robust Background Detection, RBD) and normalized color images as inputs, and directly learns an end-to-end mapping between inputs and the corresponding saliency maps. The color images are fed into a Fully Convolutional Neural Networks (FCNN) adapted from semantic segmentation to exploit high-level semantic cues for salient object detection. Then the results from deep FCNN and RBD are concatenated to feed into a shallow network to map the concatenated feature maps to saliency maps. Finally, to obtain a spatially consistent saliency map with sharp object boundaries, we fuse superpixel level saliency map at multi-scale. Extensive experimental results on 8 benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches with a margin.
Jing Zhang 0052, Bo Li 0090, Yuchao Dai, Fatih Porikli, Mingyi He
ICIP5
2016 Hyperspectral image classification based on deep stacking network
abstract
Hyperspectral image (HIS) classification is a hot topic in remote sensing community and most of the existing methods extract the features of original Hyperspectral data using shallow layer networks such as neural network (NN) and support vector machine (SVM). As deep learning recently achieves great success in machine learning and pattern recognition area for its ability in deep feature extraction and representations, two deep networks i.e. deep convolutional network (DCN) and deep belief network (DBN) have been used for hyperspectral image classification and better results have been achieved. Differing from those deep networks for HSI classification, in this paper, we propose a new method for hyperspectral image classification based on deep stacking network (DSN), which owns advantages to other deep models for its simplicity when processing in batch-mode learning - not requiring stochastic gradient descent that other DNNs require. The feature extraction is gradually obtained by employing nonlinear activation function on the hidden layer nodes of each module, which is different from those DSNs that usually use linear weights between the hidden layer and the output layer. Experimental results on AVIRIS hyperspectral images show that the proposed method achieves improved classification performance when compared with that via SVM and NN methods.
Mingyi He, Yifan Zhang 0006, Jing Zhang 0052
IGARSS1
2016 Subpixel mapping of hyperspectral images based on collaborative representation
abstract
Subpixel mapping with a low resolution hyperspectral image as the only input is widely applicable due to the fact that auxiliary image with high spatial resolution is not always available in practice. In this paper, to extract spatial information without auxiliary image, the upscaled low resolution hyperspectral image is classified using collaborative representation-based classifier. Another subpixel scale classification map is available by the combination of collaborative representation-based classification, spectral unmixing and subpixel spatial attraction model. To achieve better classification performance, decision fusion is employed to elect approximate class label from these two initial classification maps for each subpixel by the voting of the neighboring subpixels. Experimental results illustrate that the proposed approach is more promising in extracting and utilizing spatial information compared with some state-of-the-art subpixel mapping approaches.
Xiaoqin Xue, Yifan Zhang 0006, Tuo Zhao, Mingyi He
IGARSS4
2016 Hyperspectral and multispectral image fusion using collaborative representation with local adaptive dictionary pair
abstract
In this paper, the spatial resolution of hyperspectral image (HSI) is enhanced by fusing it with multispectral image (MSI) of the same scene with a higher spatial resolution. The w-hole spectrum covered by HSI channels is divided into several regions according to MSI spectral channels. The HSI-MSI fusion problem is then simplified by fusing images of each spectral region one after another. Specifically, a fusion algorithm based on collaborative representation (CR) with local adaptive dictionary pair is proposed. Compared to the classic global dictionary, the scale of the local adaptive one is much smaller such that the related computational cost is also reduced. The employment of CR is capable of reducing the reconstruction error to guarantee an improved fusion performance. Simulative experiments are deployed for illustration and comparison.
Tuo Zhao, Yifan Zhang 0006, Xiaoqin Xue, Mingyi He
IGARSS4
2016 Orthogonal Nonnegative Matrix Factorization Combining Multiple Features for Spectral-Spatial Dimensionality Reduction of Hyperspectral Imagery
abstract
Nonnegative matrix factorization (NMF), which can lead to nonsubtractive parts-based representation, has been demonstrated to be effective for dimensionality reduction of hyperspectral imagery (HSI). However, existing NMF methods applied to HSI use only a single spectral feature and do not take into consideration spatial information, such as texture or morphological features, while it has been widely acknowledged that exploiting multiple features can improve performance. Consequently, a variant of orthogonal NMF, which can not only achieve a nonnegative factorization but also exploit the complementary information that arises among heterogeneous features, is proposed for hyperspectral dimensionality reduction. The proposed method, which couples orthogonal NMF with a previous multiple-features-combining algorithm, yields a discriminative low-dimensional feature representation that matches the intuition that parts should sum to produce a whole. An efficient multiplicative updating procedure is derived, and its local convergence is guaranteed theoretically. Experimental results on two hyperspectral data sets demonstrate the effectiveness of the proposed method.
Jinhuan Wen, James E. Fowler, Mingyi He, Yongqiang Zhao 0001, Chengzhi Deng, Vineetha Menon
IEEE Trans. Geosci. Remote. Sens.3
2015 Depth and surface normal estimation from monocular images using regression on deep features and hierarchical CRFs
abstract
Predicting the depth (or surface normal) of a scene from single monocular color images is a challenging task. This paper tackles this challenging and essentially underdetermined problem by regression on deep convolutional neural network (DCNN) features, combined with a post-processing refining step using conditional random fields (CRF). Our framework works at two levels, super-pixel level and pixel level. First, we design a DCNN model to learn the mapping from multi-scale image patches to depth or surface normal values at the super-pixel level. Second, the estimated super-pixel depth or surface normal is refined to the pixel level by exploiting various potentials on the depth or surface normal map, which includes a data term, a smoothness term among super-pixels and an auto-regression term characterizing the local structure of the estimation map. The inference problem can be efficiently solved because it admits a closed-form solution. Experiments on the Make3D and NYU Depth V2 datasets show competitive results compared with recent state-of-the-art methods.
Bo Li 0090, Chunhua Shen, Yuchao Dai, Anton van den Hengel, Mingyi He
CVPR5
2015 Spatial preprocessing for spectral endmember extraction by local linear embedding
abstract
Endmember extraction (EE) has been widely utilized to identify spectrally unique signatures of pure ground materials in hyperspectral images. Most of existing EE algorithms focus on spectral signature only, denoted as spectral EE (sEE) algorithms in this paper. In order to improve the performance of these sEE algorithms by considering spatial information, a novel spatial preprocessing (SPP) strategy based on Locally Linear Embedding (LLE) is proposed to alleviate the influence of spectral variation. Specifically, the LLE is adopted to revise pixels by smoothing spectral variation in their spatial neighborhood. Furthermore, anomalous pixels, which may be smoothed excessively by many current SPP algorithms, can be well retained by tuning off the spatial preprocessing if their signatures are revised unexpectively. As a result, the anomalous endmembers can be correctly identified by the proposed LLE based SPP algorithm. Experimental results on simulated benchmark dataset have demonstrated that the proposed LLE based SPP algorithm outperforms many state-of-the-art SPP algorithms.
Shaohui Mei, Qian Du 0001, Mingyi He, Yihang Wang 0001
IGARSS3
2015 Hyperspectral and multispectral image fusion using CNMF with minimum endmember simplex volume and abundance sparsity constraints
abstract
Hyperspectral (HS) remote sensing image with finer spectral information has great advantages in feature identification and classification. However, the spatial resolution of HS image is usually low due to practical limitations. In this paper, the low-spatial-resolution HS image is fused with the high-spatial-resolution multispectral (MS) image of the same observation scene to improve its spatial resolution. A novel spectral unmixing based HS and MS image fusion approach (VSC-CNMF) is proposed, in which CNMF with minimum endmember simplex volume and abundance sparsity constraints is employed for coupled unmixing of HS and MS images. Simulative experiments are employed for verification and comparison. The experimental results illustrate that the newly proposed VSC-CNMF based HS and MS fusion algorithm outperforms several state-of-the-art unmixing based fusion approaches in cases with moderate number of endmembers.
Yifan Zhang 0006, Chuwen Zhang, Mingyi He, Shaohui Mei
IGARSS5
2015 Resource restricted on-line Video Summarization with Minimum Sparse Reconstruction
abstract
Video Summarization (VS) techniques have been widely utilized to produce a concise video content representation, such that the video content can be quickly explored and the complexity of video based analysis and retrieval applications can be highly reduced. However, little attention has been paid for on-line applications, especially for resource restricted applications, such as onboard VS. In this paper, our previous on-line Minimum Sparse Reconstruction (OnMSR) based VS algorithm is improved for resources restricted applications by confining the size of keyframes for reconstruction. Specially, an on-line reconstruction keyframe set update strategy is designed to meet the requirement of real-time resource restricted situation. Experimental results on various types of videos demonstrate the performance of OnMSR does not vary much by imposing resource constraint in the proposed resource restricted OnMSR (RR-onMSR) algorithm. As a result, the proposed RR-onMSR is very effective for real-time onboard VS applications.
Shaohui Mei, Zhiyong Wang 0001, Mingyi He, David Dagan Feng
PCS3
2015 Video summarization via minimum sparse reconstruction
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Shuai Wan, Mingyi He, David Dagan Feng
Pattern Recognit.5
2014 Iterative keyframe selection by orthogonal subspace projection
abstract
Recent developments on sparse dictionary selection have demonstrated promising results for Video Summarization (VS). However, the convex relaxation based solution cannot ensure the sparsity of the dictionary directly. In this paper, a selection matrix is proposed to model the VS problem, according to which the L0norm of this selection matrix is imposed to ensure sparsity directly. As a result, a computational efficient Orthogonal Subspace Projection (OSP) based Iterative Keyframe Selection (IKS) algorithm is proposed for VS. In addition, a Percentage Of Reconstruction (POR) criterion is proposed to provide an intuitive and flexible control of the length of final video summaries even without prior knowledge of a given video. Experimental results on a popular benchmark dataset demonstrate that our proposed algorithm outperforms the state-of-the-art methods.
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Mingyi He, Shuai Wan, David Dagan Feng
ICIP4
2014 L2, 0 constrained sparse dictionary selection for video summarization
abstract
The ever increasing volume of video content has created profound challenges for developing efficient video summarization (VS) techniques to access the data. Recent developments on sparse dictionary selection have demonstrated promising results for VS, however, the convex relaxation based solution cannot ensure the sparsity of the dictionary directly and it selects keyframes in a local point of view. In this paper, an L2,0constrained sparse dictionary selection model is proposed to reformulate the problem of VS. In addition, a simultaneous orthogonal matching pursuit (SOMP) based method is proposed to obtain an approximate solution for the proposed model without smoothing the penalty function, and thus selects keyframes in a global point of view. In order to allow for intuitive and flexible configuration of VS process, a percentage of residuals (POR) criterion is also developed to produce video summaries in different lengths. Experimental results demonstrate that our proposed method outperforms the state-of-the-art.
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Mingyi He, Xian-Sheng Hua 0001, David Dagan Feng
ICME4
2014 A Simple Prior-Free Method for Non-rigid Structure-from-Motion Factorization
Yuchao Dai, Hongdong Li, Mingyi He
Int. J. Comput. Vis.3
2014 Optimizing Hopfield Neural Network for Spectral Mixture Unmixing on GPU Platform
abstract
The Hopfield neural network (HNN) has been demonstrated to be an effective tool for the spectral mixture unmixing of hyperspectral images. However, it is extremely time consuming for such per-pixel algorithm to be utilized in real-world applications. In this letter, the implementation of a multichannel structure of HNN (named as MHNN) on a graphics processing unit (GPU) platform is proposed. According to the unmixing procedure of MHNN, three levels of parallelism, including thread, block, and stream, are designed to explore the peak computing capacity of a GPU device. In addition, constant and texture memories are utilized to further improve its computational performance. Experiments on both synthetic and real hyperspectral images demonstrated that the proposed GPU-based implementation works on the peak computing ability of a GPU device and obtains several hundred times of acceleration versus the CPU-based implementation while its unmixing performance remains unchanged.
Shaohui Mei, Mingyi He, Zhiming Shen
IEEE Geosci. Remote. Sens. Lett.2
2014 Classification Based on 3-D DWT and Decision Fusion for Hyperspectral Image Analysis
abstract
In this letter, a fusion-classification system is proposed to alleviate ill-conditioned distributions in hyperspectral image classification. A windowed 3-D discrete wavelet transform is first combined with a feature grouping-a wavelet-coefficient correlation matrix (WCM)-to extract and select spectral-spatial features from the hyperspectral image dataset. The adjacent wavelet-coefficient subspaces (from the WCM) are intelligently grouped such that correlated coefficients are assigned to the same group. Afterwards, a multiclassifier decision-fusion approach is employed for the final classification. The performance of the proposed classification system is assessed with various classifiers, including maximum-likelihood estimation, Gaussian mixture models, and support vector machines. Experimental results show that with the proposed fusion system, independent of the classifier adopted, the proposed classification system substantially outperforms the popular single-classifier classification paradigm under small-sample-size conditions and noisy environments.
Zhen Ye 0007, Saurabh Prasad, Wei Li 0032, James E. Fowler, Mingyi He
IEEE Geosci. Remote. Sens. Lett.5
2014 Hyperspectral Image Resolution Enhancement Using High-Resolution Multispectral Image Based on Spectral Unmixing
abstract
In this paper, a hyperspectral (HS) image resolution enhancement algorithm based on spectral unmixing is proposed for the fusion of the high-spatial-resolution multispectral (MS) image and the low-spatial-resolution HS image (HSI). As a result, a high-spatial-resolution HSI is reconstructed based on the high spectral features of the HSI represented by endmembers and the high spatial features of the MS image represented by abundances. Since the number of endmembers extracted from the MS image cannot exceed the number of bands in least-squares-based spectral unmixing algorithm, large reconstruction errors will occur for the HSI, which degrades the fusion performance of the enhanced HSI. Therefore, in this paper, a novel fusion framework is also proposed by dividing the whole image into several subimages, based on which the performance of the proposed spectral-unmixing-based fusion algorithm can be further improved. Finally, experiments on the Hyperspectral Digital Imagery Collection Experiment and Airborne Visible/Infrared Imaging Spectrometer data demonstrate that the proposed fusion algorithms outperform other famous fusion techniques in both spatial and spectral domains.
Mohamed Amine Bendoumi, Mingyi He, Shaohui Mei
IEEE Trans. Geosci. Remote. Sens.2
2014 A Top-Down Approach for Video Summarization
abstract
While most existing video summarization approaches aim to identify important frames of a video from either a global or local perspective, we propose a top-down approach consisting of scene identification and scene summarization. For scene identification, we represent each frame with global features and utilize a scalable clustering method. We then formulate scene summarization as choosing those frames that best cover a set of local descriptors with minimal redundancy. In addition, we develop a visual word-based approach to make our approach more computationally scalable. Experimental results on two benchmark datasets demonstrate that our proposed approach clearly outperforms the state-of-the-art.
Genliang Guan, Zhiyong Wang 0001, Shaohui Mei, Maximilian Ott, Mingyi He, David Dagan Feng
ACM Trans. Multim. Comput. Commun. Appl.5
2013 Hyperspectral image classification based on iterative Support Vector Machine by integrating spatial-spectral information
abstract
The well-known difficulty in supervised hyperspectral image classification is the limited availability of training data, which are expensive, and quite difficult to access and to obtain in real remote sensing scenarios. The Support Vector Machine (SVM) technique has been proven to be well suited to classify hyperspectral data by using limited number of training samples. In this paper, modifications over Iterative Support Vector Machine algorithm have been proposed incorporating both spatial and spectral information and correcting the training samples at each iteration in order to increase the classification performance over SVM. In order to demonstrate the effectiveness of the proposed framework, experiments on AVIRIS data over Indian Pine Site (IPS) are conducted to compare the performance of the proposed classification approach against some existing classification techniques such as Linear-SVM, SVM-RBF, ISVM and K-NN. Experimental results demonstrate that the proposed method clearly outperform the well-known classification algorithms.
Belkacem Baassou, Mingyi He, Muhammad Imran Farid, Shaohui Mei
IGARSS2
2013 Neighborhood preserving Nonnegative Matrix Factorization for spectral mixture analysis
abstract
Nonnegative Matrix Factorization (NMF) has been successfully employed to address the mixed-pixel problem of hyperspectral remote sensing images. However, minimizing the representation error by NMF is not sufficient for SMA since the unmixing results of NMF are not unique. Therefore, in this paper, a neighborhood preserving regularization, which preserves the local structure of the hyperspectral data on a low-dimensional manifold, is proposed to constrain NMF for unique solution in SMA. As a result, a Neighborhood Preserving constrained NMF (NP-NMF) algorithm is proposed for SMA of highly mixed hyperspectral data. Finally, experimental results on AVIRIS data demonstrate the effectiveness of our proposed NP-NMF algorithm for SMA applications.
Shaohui Mei, Mingyi He, Zhiming Shen, Belkacem Baassou
IGARSS2
2013 Projective Multiview Structure and Motion from Element-Wise Factorization
abstract
The Sturm-Triggs type iteration is a classic approach for solving the projective structure-from-motion (SfM) factorization problem, which iteratively solves the projective depths, scene structure, and camera motions in an alternated fashion. Like many other iterative algorithms, the Sturm-Triggs iteration suffers from common drawbacks, such as requiring a good initialization, the iteration may not converge or may only converge to a local minimum, and so on. In this paper, we formulate the projective SfM problem as a novel and original element-wise factorization (i.e., Hadamard factorization) problem, as opposed to the conventional matrix factorization. Thanks to this formulation, we are able to solve the projective depths, structure, and camera motions simultaneously by convex optimization. To address the scalability issue, we adopt a continuation-based algorithm. Our method is a global method, in the sense that it is guaranteed to obtain a globally optimal solution up to relaxation gap. Another advantage is that our method can handle challenging real-world situations such as missing data and outliers quite easily, and all in a natural and unified manner. Extensive experiments on both synthetic and real images show comparable results compared with the state-of-the-art methods.
Yuchao Dai, Hongdong Li, Mingyi He
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 A simple prior-free method for non-rigid structure-from-motion factorization
abstract
This paper proposes a simple “prior-free” method for solving non-rigid structure-from-motion factorization problems. Other than using the basic low-rank condition, our method does not assume any extra prior knowledge about the nonrigid scene or about the camera motions. Yet, it runs reliably, produces optimal result, and does not suffer from the inherent basis-ambiguity issue which plagued many conventional nonrigid factorization techniques. Our method is easy to implement, which involves solving no more than an SDP (semi-definite programming) of small and fixed size, a linear Least-Squares or trace-norm minimization. Extensive experiments have demonstrated that it outperforms most of the existing linear methods of nonrigid factorization. This paper offers not only new theoretical insight, but also a practical, everyday solution, to non-rigid structure-from-motion.
Yuchao Dai, Hongdong Li, Mingyi He
CVPR3
2012 Unmixing approach for hyperspectral data resolution enhancement using high resolution multispectral image
abstract
In order to enhance the spatial resolution of the hyperspectral images, a novel fast algorithm based on Spectral Mixture Analysis (SMA) techniques is proposed for the fusion of coarse-resolution hyperspectral (HS) image and high-resolution multispectral (MS) image. The high-resolution hyperspectral image is synthesized by integrating high-resolution spectral information of hyperspectral image represented by endmembers and high-resolution spatial information of multispectral image represented by abundance. As a result, a novel SMA based diagram is designed, in which Endmember Extraction (EE) is performed on hyperspectral images while Abundance Estimation is performed on multispectral images, and the unmixing process in these two images are matched by utilizing the spectral response matrix and the spatial spread transform matrix in the observation model. Finally, real HYDICE data experiments are utilized to demonstrate the effectiveness of the proposed fusion algorithm.
Mohamed Amine Bendoumi, Mingyi He, Shaohui Mei, Yifan Zhang 0006
ICARCV2
2012 Unsupervised Spectral Mixture Analysis with Hopfield Neural Network for hyperspectral images
abstract
Spectral Mixture Analysis (SMA) has been widely utilized to address the mixed-pixel problem in the quantitative analysis of hyperspectral remote sensing images. Recently Nonnegative Matrix Factorization (NMF) has been successfully utilized to simultaneously perform endmember extraction (EE) and abundance estimation (AE). In this paper, we formulate the solution of NMF by performing EE and AE iteratively. Based on our previous Hopfield Neural Network (HNN) based AE algorithm, an HNN is also constructed for EE to solve the multiplicative updating problem of NMF for SMA. As a result, SMA is conducted in an unsupervised manner and our algorithm is able to extract virtual endmembers without assuming the presence of spectrally pure constituents in hyperspectral scenes. We further extend such strategy to solve the constrained NMF (cNMF) models for SMA, where extra constraints are imposed to better model the mixed-pixel problem. Experimental results on both synthetic and real hyperspectral images demonstrate the effectiveness of our proposed HNN based unsupervised SMA algorithms.
Shaohui Mei, Mingyi He, Zhiyong Wang 0001, David Dagan Feng
ICIP2
2012 Sharp curve lane boundaries projective model and detection
abstract
An effective lane boundaries projective model (LBPM) and improved detection method in the images captured with a vehicle-mounted monocular camera in complex environments, especially for sharp circular curve lane, is proposed in this paper. Firstly, a lane boundaries projective model is deduced. This lane model can not only express straight-line lane boundaries, but also describe the actual sharp circular curve lane boundaries very well. Secondly, the lane posterior probability function is derived by employing the lane model, the gradient direction feature, the lane likelihood function, and the lane prior information. And then the lane maximum posteriori probability is found out by using the improved particle swarm optimization algorithm. Further the lane boundaries is positioned, and the lane geometric structure, such as the lane left and right boundaries curve radiuses, can be calculated accurately through the lane model. The experimental results show that the proposed lane boundaries projective model and the improved detection method are more effective and accurate for sharp curve lane detection.
Mingyi He
INDIN2
2011 Minimum endmember-wise distance constrained Nonnegative Matrix Factorization for Spectral Mixture Analysis of hyperspectral images
abstract
Nonnegative Matrix Factorization (NMF) and its extensions have gained lots of attentions in Spectral Mixture Analysis (SMA) since they can handle highly mixed hyperspectral pixels in an unsupervised way. In order to overcome the non-uniqueness problem in NMF, a minimum endmember-wise distance constraint (MewDC), which optimizes endmember spectra as compact as possible, is imposed for satisfying unmixing results. The proposed constraint works similar to minimum volume constraint (MVC). However, the dimension reduction step and numerical instability problems in MVC can be avoided. As a result, a minimum endmember-wise distance constrained NMF (MewDC-NMF) algorithm is proposed to extract endmembers and estimate their corresponding fractional abundance simultaneously. Both synthetic and real hyperspectral data experiments have demonstrate the effectiveness of the proposed MewDC-NMF algorithm.
Shaohui Mei, Mingyi He
IGARSS2
2011 Bayesian fusion of hyperspectral and multispectral images using Gaussian scale mixture prior
abstract
In this paper, a wavelet-based Bayesian fusion framework is presented, in which a low spatial resolution hyperspectral (HS) image is fused with a high spatial resolution multi-spectral (MS) image by accounting for the joint statistics. Particularly, a zero-mean heavy-tailed model, Gaussian Scale Mixture (GSM) model, is employed as the prior, which is believed to be capable of modelling the distribution of wavelet coefficients more accurately than traditional Gaussian model. To keep the calculations feasible, a practical implementation scheme is presented. The proposed approach is validated by simulation experiments for both general HS and MS image fusion as well as the specific case of pansharpening. The experimental results of the proposed approach are also compared with its counterpart employing a Gaussian prior for performance evaluation.
Yifan Zhang 0006, Shaohui Mei, Mingyi He
IGARSS3
2011 Improving Spatial-Spectral Endmember Extraction in the Presence of Anomalous Ground Objects
abstract
Endmember extraction (EE) has been widely utilized to extract spectrally unique and singular spectral signatures for spectral mixture analysis of hyperspectral images. Recently, spatial–spectral EE (SSEE) algorithms have been proposed to achieve superior performance over spectral EE (SEE) algorithms by taking both spectral similarity and spatial context into account. However, these algorithms tend to neglect anomalous endmembers that are also of interest. Therefore, in this paper, an improved SSEE (iSSEE) algorithm is proposed to address such limitation of conventional SSEE algorithms by accounting for both anomalous and normal endmembers. By developing simplex projection and simplex complementary projection, all the hyperspectral pixels are projected into a simplex determined by the normal endmembers extracted in conventional SSEE algorithms. As a result, anomalous endmembers are identified iteratively by utilizing the$l_{2}^{\infty}$norm to find the maximum simplex complementary projection. In order to determine how many anomalous endmembers are to be extracted, a novel Residual-be-Noise Probability-based algorithm is also proposed by elegantly utilizing the spatial-purity map generated in the previous SSEE step. Experimental results on both synthetic and real datasets demonstrate that simplex projection errors can be significantly reduced by identifying both anomalous and normal endmembers in the proposed iSSEE algorithm. It is also confirmed that the performance of the proposed iSSEE algorithm clearly outperforms that of SEE algorithms since both spatial context and spectral similarity are utilized.
Shaohui Mei, Mingyi He, Yifan Zhang 0006, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Geosci. Remote. Sens.2
2010 Element-Wise Factorization for N-View Projective Reconstruction
Yuchao Dai, Hongdong Li, Mingyi He
ECCV (4)3
2010 Advance in triangular mesh simplification study
abstract
With the development of modern 3D acquisition facilities and tools, large mesh models with high-precision for the representation of complex geometric objects becomes feasible. However, many considerable difficulties have emerged in these huge mesh models, such as the storage capacity and rendering speed of computers. Triangular mesh, which is one of the most popular polygonal meshes, plays an important role in mesh simplification field. Many different methods and algorithms have been put forward to the research of triangular mesh simplification in the past few years. In this paper, triangular mesh simplification algorithms and their corresponding improved algorithms are overviewed. In addition, some recent achievements of triangular mesh simplification in the author's laboratory (IAP) are presented. Finally, future development and prospects in this area are also discussed and outlined, respectively.
Mingyi He
ICARCV1
2010 Mixture Analysis by Multichannel Hopfield Neural Network
abstract
Due to the spatial-resolution limitation, mixed pixels containing energy reflected from more than one type of ground objects are widely present in remote sensing images, which often results in inefficient quantitative analysis. To effectively decompose such mixtures, a fully constrained linear unmixing algorithm based on a multichannel Hopfield neural network (MHNN) is proposed in this letter. The proposed MHNN algorithm is actually a Hopfield-based architecture which handles all the pixels in an image synchronously, instead of considering a per-pixel procedure. Due to the synchronous unmixing property of MHNN, a noise energy percentage (NEP) stopping criterion which utilizes the signal-to-noise ratio is proposed to obtain optimal results for different applications automatically. Experimental results demonstrate that the proposed multichannel structure makes the Hopfield-based mixture analysis feasible for real-world applications with acceptable time cost. It has also been observed that the proposed MHNN-based mixture-analysis algorithm outperforms the other two popular linear mixture-analysis algorithms and that the NEP stopping criterion can approach optimal unmixing results adaptively and accurately.
Shaohui Mei, Mingyi He, Zhiyong Wang 0001, David Dagan Feng
IEEE Geosci. Remote. Sens. Lett.2
2010 Spatial Purity Based Endmember Extraction for Spectral Mixture Analysis
abstract
Spectral mixture analysis (SMA) has been widely utilized to address the mixed-pixel problem in the quantitative analysis of hyperspectral remote sensing images, in which endmember extraction (EE) plays an extremely important role. In this paper, a novel algorithm is proposed to integrate both spectral similarity and spatial context for EE. The spatial context is exploited from two aspects. At first, initial endmember candidates are identified by determining the spatial purity (SP) of pixels in their spatial neighborhoods (SNs). Several SP measurements are investigated at both intensity level and feature level. In order to alleviate local spectra variability, the average of the pixels in pure SNs are voted as endmember candidates. Then, the spatial connectivity is utilized to merge spatially related endmember candidates by finding connection paths in a graph so that the number of endmember candidates is further reduced, which results in computational efficiency and better performance in SMA by alleviating global spectral variability. Experimental results on both synthetic and real hyperspectral images demonstrate that the proposed SP based EE (SPEE) algorithm outperforms the other popular EE algorithms. It is also observed that feature-level SP measurements are more distinguishable than intensity-level SP measurements to discriminate pure SNs from mixed SNs.
Shaohui Mei, Mingyi He, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Geosci. Remote. Sens.2
2009 Two Efficient Algorithms for Outlier Removal in Multi-view Geometry Using L-Infinity Norm
abstract
L∞ norm has been recently introduced to multi-view geometry computation to achieve globally optimal computation. It however suffers from a serious sensitivity to outliers. A few remedies have been proposed but with high computational complexity. This paper presents two efficient algorithms to overcome these problems. Our first algorithm is based on a cheap and effective local descent method (as opposed to the conventional but expensive SOCP(Second Order Cone Programming)). The second algorithm further improves the first one by using a Depth-first search heuristics. Both algorithms retain the nice property of global optimality of the L∞ scheme, while at cost only a small fraction of the original computation. Experiments on both synthetic data and real images have validated the proposed algorithms.
Yuchao Dai, Mingyi He, Hongdong Li
ICIG2
2007 Feature selection using Double Parallel Feedforward Neural Networks and Particle Swarm Optimization
abstract
In recent years, the Neural Network (NN) based feature selection becomes a promising method for dimensionality reduction. However, Multi-layer Feedforward Neural Network (MFNN) with wide applications has some disadvantages such as local minimal points on the error surface and over-fitting problem. At the same time, the conventional approaches usually fixing teh number of hidden nodes and focusing on the input selection hinder further remove of the redundant information and improvement of network generalization performance. To solve these problems, a feature selection algorithm using Double Parallel Feedforward Neural Network (DPFNN) and Particle Swarm Optimization (PSO) is proposed. The algorithm adopts DPFNN with the merits of Single-layer Feedforward Neural Network (SFNN) and MFNN as the criterion function, synchronously performs optimization of structure and selection of inputs based on a new defined fitness function keeping balance between network performance and complexity. Experimental results show that the algorithm can effectively remove the redundant features while improving the generalization ability of network.
Mingyi He
IEEE Congress on Evolutionary Computation2
2005 Band selection based on feature weighting for classification of hyperspectral data
abstract
A new feature weighting method for band selection is presented, which is based on the pairwise separability criterion and matrix coefficients analysis. Through decorrelation of each class by principal component transformation, the criterion value of any band subset is the summations of the values of individual bands of it for the transformed feature space, and thus the computation amounts of calculating criteria of each band combinations are reduced. Following it, the corresponding matrix coefficients analysis is done to assign weights to original bands. As feature weighting considers little about the spectral correlation, the redundant bands are removed by choosing those with lower correlation coefficients than a preset threshold. Hyperspectral data classification experiments show the effectiveness of the new band selection method.
Mingyi He
IEEE Geosci. Remote. Sens. Lett.2
2004 Multi-resolution surface reconstruction
Mingyi He, Huajing Yu
ICIP1
2004 Wavelet image coding by dilation-run algorithm
abstract
This paper presents a novel wavelet image coder, dilation-run algorithm, which provides an embedded coder based on bit-plane coding technique. At each bit-plane, morphological dilation is used to extract and encode the clustering significant wavelet coefficients, and a run-length coding method is used to encode the positional information of each cluster. The experiment results show that our coder outperforms the zerotree coder SPIHT and is competitive with the morphology coders MRWD and SLCCA. For wavelet images with strong clustering feature, the new coder outperforms both the morphology coders above.
Mingyi He
ICIP2
1999 A parallel-distributed processing for time-domain deconvolution coefficients
Mingyi He
Signal Process.1