Siyuan Hao

dblp:158/8309 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0001-8247-4207ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Planning and Control for Active Morphing Tensegrity Aerial Vehicles in Confined Spaces
abstract
Morphing quadrotors are capable of adapting to constrained environments through geometric reconfiguration. However, existing systems are limited by mechanical complexity and rigid links, which affect both safety and performance in such environments. In this paper, we propose a strut-actuated tensegrity aerial vehicle that integrates shape adaptation with collision resilience. By incorporating deformable struts and a cable network, our vehicle enables real-time morphological adjustments during flight while maintaining stability. We present a hierarchical planning framework that ensures the entire vehicle remains confined within an icosahedral space, thereby guaranteeing full-body safety. An on-manifold Model Predictive Controller (MPC) is employed to track these optimized trajectories and compensate for inertia shifts during shape deformation. Simulation results validate the effectiveness of the proposed framework, demonstrating its capability to navigate in restricted scenarios.
Siyuan Hao, Zichen Tao, Yun Gui, Songyuan Liu, Jiaxu Shi, Qingkai Yang
IROS1
2025 Semantic-Aware Guidance for Blind Super-Resolution of Remote Sensing Images
abstract
Unlike traditional super-resolution (SR) methods that rely on fixed degradation models, blind SR (BSR) methods can capture the complex processes introduced by factors such as sensor noise and platform motion in real-world remote sensing imagery. While most BSR methods effectively remove degradation from low-resolution (LR) images, they often struggle to preserve high-frequency details, leading to reduced reconstruction accuracy. To address this issue, we propose the semantic-aware guidance BSR (SGBSR) network, which leverages semantic information to guide the entire restoration process, enabling more accurate reconstruction. Specifically, we design a semantic extractor that utilizes powerful pretrained visual models to capture rich semantic information from LR images, which is then integrated into the SR network. To further enhance the network’s ability to handle complex degradations, we introduce an implicit estimation method. Subsequently, the semantic information and degradation representations are, respectively, incorporated into the SR network through the semantic-aware block (SaB) and the degradation-aware block (DaB). Experiments on both synthetic and real-world LR images demonstrate that our method achieves superior reconstruction accuracy.
Siyuan Hao, Wei Wang 0108
IEEE Geosci. Remote. Sens. Lett.2
2024 Prime Label Learning From Multilabel Aerial Image: A Novel Weakly Supervised Task
abstract
In the task of multi-label aerial image classification, various objects and land cover in an image are usually represented by multiple labels which are treated equally. However, from a semantic point of view, the importance of multiple labels are different in a specific scene. There is often a prime label in the image that plays a "leading" role. Obtaining the most important label from candidate multiple labels is crucial, because it best represents the semantics of the entire image. In this letter, we attempt to automatically obtain the prime label of each image from several existing multi-label aerial image datasets without additional supervision cost. In other words, the prime labels are only used to evaluate the performance of models during testing and do not participate in the training process. Therefore, it is essentially a weakly supervised learning task. For this novel aerial image classification task, corresponding datasets are provided in this letter firstly, including over head images with multi-labels for training and prime labels for testing. Then the baselines on the above datasets are provided. Finally, a new prime label learning method is proposed, which improves the baseline accuracy by about 14% and reaches the state-of-the-art on current datasets.
Shiwen Zeng, Lijian Zhou, Tingyuan Nie, Siyuan Hao
IEEE Geosci. Remote. Sens. Lett.5
2023 Hybrid Heterogeneous Graph Neural Networks for Fund Performance Prediction
Siyuan Hao, Le Dai, Le Zhang 0010, Chao Wang 0086, Chuan Qin 0002, Hui Xiong 0001
KSEM (2)1
2023 Generative Adversarial Network With Transformer for Hyperspectral Image Classification
abstract
In recent years, generative adversarial networks (GAN) have made great progress in the field of hyperspectral image classification (HIC), which alleviates the problem of insufficient training samples to a large extent. At present, GAN in the field of HIC are all based on Convolutional Neural Network (CNN). But CNN cannot extract sequence information well, and it is difficult to model remote dependencies. However, hyperspectrum is rich in spectral sequence information, and Transformer has been proven to be good at processing sequence information. Therefore, in order to process spectral information and alleviate the problem of insufficient training samples of hyperspectral images (HSI), we put forward a new frame Transformer with residual upscale GAN (TRUG). TRUG includes a generator G and a discriminator D. In the G, we propose residual upscale (RU) to improve the resolution of generated features, while also extracting texture features and capturing context relationships. In addition, we visualized the generated fake images for more intuitive analysis. In the D, we adopt Transformer block with progressively decreasing scale, and use grid self-attention mechanism in the first layer to better extract features for classification. In addition, GAN are prone to the problem of unstable training. In order to solve this problem, we improve the normalization algorithm and add relative position coding. We applied a pure Transformer based GAN to HIC datasets. Experimental results show that the proposed TRUG model has better performance than other models on the three datasets.
Siyuan Hao, Yufeng Xia, Yuanxin Ye
IEEE Geosci. Remote. Sens. Lett.1
2022 Spectral and Spatial Feature Fusion for Hyperspectral Image Classification
abstract
Compared with traditional images, hyperspectral images (HSI) not only have spatial information, but also have rich spectral information. However, the mainstream hyperspectral image classification (HIC) methods are all based on Convolutional Neural Network (CNN), which has great advantages in extracting spatial features, but it has certain limitations in dealing with spectral continuous sequence information. Therefore Transformer which is good at processing sequences, has also been gradually applied to HIC. Besides, Since HSI are typical three-dimensional structures, we believe that the correlation of the three dimensions is also an important information. So in order to fully extract the spectral spatial information, as well as the correlation of the three dimensions. we propose a spectral and spatial feature fusion module (i.e., TransCNN) for HIC. TransCNN consists of CNNs and a Transformer. The former is in charge of mining the spatial and spectral information from different dimensions, while the latter not only undertakes the most critical fusion but also captures the deeper relationship characteristics. We transpose the data to extract features and their correlation through three CNNs branches. we believe that these feature maps still have deep spectral information. Therefore, we have embedded them into one-dimensional vectors and use Transformer’s Encoder to extract features. However, some information will be lost when embedding into one-dimensional vectors. Therefore we use Decoder, which has been ignored in the field of vision, to fuse the features before passing Encoder and the features after extracted by Encoders. Two kinds of features are fused by Decoder, and the obtained information is finally input into the classifier for classification. Experimental results on real HSIs show that the proposed architecture can achieve competitive performance compared with the state-of-the-art methods.
Siyuan Hao, Yufeng Xia, Lijian Zhou, Yuanxin Ye, Wei Wang 0108
IEEE Geosci. Remote. Sens. Lett.1
2022 A Multiscale Framework With Unsupervised Learning for Remote Sensing Image Registration
abstract
Registration for multisensor or multimodal image pairs with a large degree of distortions is a fundamental task for many remote sensing applications. To achieve accurate and low-cost remote sensing image registration, we propose a multiscale framework with unsupervised learning, named MU-Net. Without costly ground truth labels, MU-Net directly learns the end-to-end mapping from the image pairs to their transformation parameters. MU-Net stacks several deep neural network (DNN) models on multiple scales to generate a coarse-to-fine registration pipeline, which prevents the backpropagation from falling into a local extremum and resists significant image distortions. We design a novel loss function paradigm based on structural similarity, which makes MU-Net suitable for various types of multimodal images. MU-Net is compared with traditional feature-based and area-based methods, as well as supervised and other unsupervised learning methods on the optical-optical, optical-infrared, optical-synthetic aperture radar (SAR), and optical-map datasets. Experimental results show that MU-Net achieves more comprehensive and accurate registration performance between these image pairs with geometric and radiometric distortions. We share the code implemented by Pytorch athttps://github.com/yeyuanxin110/MU-Net.
Yuanxin Ye, Tengfeng Tang, Bai Zhu, Chao Yang 0028, Bo Li 0090, Siyuan Hao
IEEE Trans. Geosci. Remote. Sens.6
2022 Bounding Boxes Are All We Need: Street View Image Classification via Context Encoding of Detected Buildings
abstract
Street view image classification aiming at the urban land use analysis is difficult because the class labels (e.g., commercial area) are concepts with higher abstract levels compared to the ones of general visual tasks (e.g., persons and cars). Therefore, classification models using only visual features often fail to achieve satisfactory performance. In this article, a novel approach based on a “bottom-up and top-down” framework is proposed. Instead of using visual features of the whole image directly as common image-level models based on convolutional neural networks (CNNs) do, the proposed framework first obtains low-level semantic, namely, the bounding boxes of buildings in street view images through a bottom-up object discovery process. Their contextual information, such as the co-occurrence patterns of building classes and their layout, is then encoded into metadata by the proposed algorithm “Context encOding of Detected buildINGs” (CODING). Finally, these metadata (low-level semantic encoded with context information) are abstracted to high-level semantic, namely, the land use label of the street view image through a top-down semantic aggregation process implemented by a recurrent neural network (RNN). In addition, in order to effectively discover low-level semantic as the bridge between visual features and higher abstract concepts, we made a dual-labeled data set named “Building dEtection And Urban funcTional-zone portraYing” (BEAUTY) of 19070 street view images and 38857 buildings based on the existing BIC_GSV. The data set can be used not only for street view image classification but also for multiclass building detection. Experiments on “BEAUTY” show that the proposed approach achieves a 12.65% performance improvement on macroprecision and 12% on macrorecall over image-level CNN-based models. Our code and data set are available athttps://github.com/kyle-one/Context-Encoding-of-Detected-Buildings/.
Yongkun Liu, Siyuan Hao, Shaoxing Lu, Lijian Zhou
IEEE Trans. Geosci. Remote. Sens.3
2021 A Novel Keypoint Detector Combining Corners and Blobs for Remote Sensing Image Registration
abstract
Keypoint detection is a crucial step for feature-based image registration. The traditional detectors only extract one type of keypoint such as a corner or a blob, which is not quite beneficial to image registration. Accordingly, this letter presents a novel keypoint detector that aims to simultaneously extract corners and blobs. The proposed detector is named as Harris-Difference of Gaussian (DoG), which combines the advantages of the Harris-Laplace corner detector and the DoG blob detector. In the definition of Harris-DoG, we first build an image scale space and extract the corners by using the multiscale Harris detector. Then, these corners are screened by an automatic scale selection technique based on their DoG responses. This can make the corners robust to scale changes. Meanwhile, DoG also is used to detect the blobs in the image scale space by a nonmaxima suppression scheme. Finally, the scale invariant feature transform (SIFT) descriptors are computed for the detected corners and blobs, and they are applied together for image registration. The proposed Harris-DoG has been tested by using three pairs of multisensor remote sensing images. The experimental results show that Harris-DoG can effectively increase the number of correct matches and improve the registration accuracy compared with the state-of-the-art keypoint detectors.
Yuanxin Ye, Mengmeng Wang 0007, Siyuan Hao, Qing Zhu 0012
IEEE Geosci. Remote. Sens. Lett.3
2021 Geometry-Aware Deep Recurrent Neural Networks for Hyperspectral Image Classification
abstract
Variants of deep networks have been widely used for hyperspectral image (HSI)-classification tasks. Among them, in recent years, recurrent neural networks (RNNs) have attracted considerable attention in the remote sensing community. However, complex geometries cannot be learned easily by the traditional recurrent units [e.g., long short-term memory (LSTM) and gated recurrent unit (GRU)]. In this article, we propose a geometry-aware deep recurrent neural network (Geo-DRNN) for HSI classification. We build this network upon two modules: a U-shaped network (U-Net) and RNNs. We first input the original HSI patches to the U-Net, which can be trained with very few images and obtain a preliminary classification result. We then add RNNs on the top of the U-Net so as to mimic the human brain to refine continuously the output-classification map. However, instead of using the traditional dot product in each gate of the RNNs, we introduce a Net-Gated GRU that increases the nonlinear representation power. Finally, we use a pretrained ResNet as a regularizer to improve further the ability of the proposed network to describe complex geometries. To this end, we construct a geometry-aware ResNet loss, which leverages the pretrained ResNet's knowledge about the different structures in the real world. Our experimental results on real HSIs and road topology images demonstrate that our approach outperforms the state-of-the-art classification methods and can learn complex geometries.
Siyuan Hao, Wei Wang 0108, Mathieu Salzmann
IEEE Trans. Geosci. Remote. Sens.1
2020 Face recognition based on local binary pattern and improved Pairwise-constrained Multiple Metric Learning
Lijian Zhou, Shanshan Lin, Siyuan Hao, Zheming Lu 0001
Multim. Tools Appl.4
2019 Combining multi-wavelet and CNN for palmprint recognition against noise and misalignment
abstract
A palmprint recognition approach based on multi‐wavelet and convolutional neural network (CNN) against noise and misalignment is given. CNN method has high robustness in biometrics, but a large number of training samples are necessary. Moreover, the gathered palmprint images should be cropped to obtain their region of interest (ROI) and noise pollution and misalignment are not well solved. Therefore, the original training database is augmented to reduce the effects of noise and misalignment. First, an original training palmprint image is split into five new images, and every new image is decomposed once by multi‐wavelet. Three lower frequency bands in the low‐frequency multi‐wavelet component corresponding to pre‐filters are extracted as three samples. Furthermore, the split image is downsampled as a new sample. Second, the CNN model is constructed based on the augmented database by experiments. Third, the softmax method is used to classify the test samples. At last, the final result is obtained from 20 results by using the voting method. The experimental results based on PolyU, CASIA, and IIT Delhi Touchless Palmprint Database palmprint databases show that the proposed method can effectively recognise palmprint with high robustness while there is noise and misalignment, and has a generalisation to other palmprint databases.
Lijian Zhou, Shanshan Lin, Siyuan Hao
IET Image Process.4
2018 A Deep Network Architecture for Super-Resolution-Aided Hyperspectral Image Classification With Classwise Loss
abstract
The supervised deep networks have shown great potential in improving the classification performance. However, training these supervised deep networks is very challenging for hyperspectral image given the fact that usually only a small amount of labeled samples are available. In order to overcome this problem and enhance the discriminative ability of the network, in this paper, we propose a deep network architecture for a super-resolution (SR)-aided hyperspectral image classification with classwise loss (SRCL). First, a three-layer SR convolutional neural network (SRCNN) is employed to reconstruct a high-resolution image from a low-resolution image. Second, an unsupervised triplet-pipeline CNN (TCNN) with an improved classwise loss is built to encourage intraclass similarity and interclass dissimilarity. Finally, SRCNN, TCNN, and a classification module are integrated to define the SRCL, which can be fine-tuned in an end-to-end manner with a small amount of training data. Experimental results on real hyperspectral images demonstrate that the proposed SRCL approach outperforms other state-of-the-art classification methods, especially for the task in which only a small amount of training data are available.
Siyuan Hao, Wei Wang 0108, Yuanxin Ye, Enyu Li, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2018 Two-Stream Deep Architecture for Hyperspectral Image Classification
abstract
Most traditional approaches classify hyperspectral image (HSI) pixels relying only on the spectral values of the input channels. However, the spatial context around a pixel is also very important and can enhance the classification performance. In order to effectively exploit and fuse both the spatial context and spectral structure, we propose a novel two-stream deep architecture for HSI classification. The proposed method consists of a two-stream architecture and a novel fusion scheme. In the two-stream architecture, one stream employs the stacked denoising autoencoder to encode the spectral values of each input pixel, and the other stream takes as input the corresponding image patch and deep convolutional neural networks are employed to process the image patch. In the fusion scheme, the prediction probabilities from two streams are fused by adaptive class-specific weights, which can be obtained by a fully connected layer. Finally, a weight regularizer is added to the loss function to alleviate the overfitting of the class-specific fusion weights. Experimental results on real HSIs demonstrate that the proposed two-stream deep architecture can achieve competitive performance compared with the state-of-the-art methods.
Siyuan Hao, Wei Wang 0108, Yuanxin Ye, Tingyuan Nie, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2017 Class-wise dictionary learning for hyperspectral image classification
Siyuan Hao, Wei Wang 0108, Yan Yan 0002, Lorenzo Bruzzone
Neurocomputing1
2016 Spatial-dictionary for collaborative representation classification of hyperspectral images
Siyuan Hao, Liguo Wang 0001, Lorenzo Bruzzone, Qunming Wang
Multim. Tools Appl.1
2015 A Multiple-Mapping Kernel for Hyperspectral Image Classification
abstract
The kernel function plays an important role in machine learning methods such as the support vector machine. In this letter, a new kernel framework is developed for hyperspectral image classification. In contrast to existing composite kernels constructed via a linearly weighted combination, the multiple-mapping kernel proposed in this letter is obtained through repeated nonlinear mappings. Experiments indicate that the proposed multiple-mapping kernel framework (MMKF) is effective for hyperspectral image classification. Compared to the single kernel methods, the MMKF tends to be more advantageous in terms of classification accuracy, particularly for the situation with a small-size training set.
Liguo Wang 0001, Siyuan Hao, Qunming Wang, Peter M. Atkinson
IEEE Geosci. Remote. Sens. Lett.2